Quality measurement of base calling in next generation sequencing

The system uses pre-trained models and look-up tables to predict base calling quality in next-generation sequencing by selecting appropriate predictors for each cycle, addressing inaccuracies in existing methods and enhancing prediction accuracy and efficiency.

WO2026096640A1PCT designated stage Publication Date: 2026-05-07ELEMENT BIOSCIENCES INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ELEMENT BIOSCIENCES INC
Filing Date
2025-10-29
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing methods for predicting the quality of base calling in next-generation sequencing suffer from inaccuracies due to computational burden, contamination by sequencing errors, and the inability to provide predictive measurements before actual base calling, leading to overestimation or underestimation of quality scores.

Method used

A system and method that utilizes pre-trained quality prediction models or look-up tables, trained with flow cell images from individual sequencing cycles, to predict quality scores efficiently and accurately by selecting appropriate predictors for each cycle, avoiding contamination and reducing computational intensity.

Benefits of technology

The system provides more accurate and efficient prediction of base calling quality across multiple channels and within a single channel, improving reliability without increasing computational burden, and allowing real-time processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025053125_07052026_PF_FP_ABST
    Figure US2025053125_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate to a method for predicting quality of base calling in sequencing. A quality score may be determined from a look-up table among a plurality of look-up tables, each of the look-up table is trained using different training datasets. A plurality of predictors may be selected and a corresponding value for each of the plurality of predictors may be determined from one or more flow cell images, and the look-up table comprises a plurality of dimensions corresponding to a plurality of predictors.
Need to check novelty before this filing date? Find Prior Art

Description

QUALITY MEASUREMENT OF BASE CALLING IN NEXT GENERATION SEQUENCINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 713,884, filed October 30, 2024, which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Next generation sequencing-by-synthesis, sequencing by binding, or sequencing by avidity using a flow cell may be used for identifying sequences of DNA. As single-stranded DNA fragments from a sequencing library are flooded across a flow cell, the fragments may attach to the surface of the flow cell. An amplification process is then performed on the DNA fragments, such that copies of a given fragment form a cluster or polony of nucleotide strands. In some aspects, a single cluster may attach to the flow cell at random locations.

[0003] In next-generation sequencing (NGS) or NGS-like applications such as sequencing by synthesis, sequencing by binding, or sequencing by avidity, in order to identify the sequence of a target nucleic acid, a new strand is synthesized one nucleotide base at a time. During each cycle, 3 ’-blocked nucleotides attach at complementary positions on the strands, ensuring that only one base will attach to any given strand during a single cycle. In some aspects, such as in sequencing by synthesis, the blocked nucleotide may also be fluorescently labeled, while in others, such as in sequencing by binding or sequencing by avidity, a label is reversibly or noncovalently bound to the synthesis complex in a separate step that takes place after the blocked nucleotide has been incorporated. During the detection step, the flow cell is exposed to excitation light, exciting the labels and causing them to fluoresce. Because, in most existing aspects, the strands undergoing sequencing are clustered together, the fluorescent signal for any one fragment is amplified by the signal from its clonal counterparts, such that the fluorescence for an entire colony may be recorded by an imager. To initiate subsequent sequencing steps, the blocking groups are then cleaved, the surface is washed, and the cycle repeats. Importantly, at the imaging step of each sequencing cycle, one or more images are recorded. A base-calling algorithm is applied to the recorded images to “read” the successive signals from each cluster or polony and convert the optical signals into an identification of the nucleotide base sequence added to each fragment.Errors may be introduced in base calling due to many reasons including the base-calling algorithm, the sequencing steps prior to it, or during library preparation. It remains a challenge to accurately measure the quality of base calling.SUMMARY

[0004] Provided herein are system, apparatus, method, and / or computer program product aspects, and / or combinations and sub-combinations thereof which enables predictive measurement of quality of base calling obtained using various sequencing methods.

[0005] As a particular application of such, aspects of methods, systems, and media for predicting quality of base calling using various sequencing methods disclosed herein.

[0006] Other aspects of these aspects include corresponding computer systems, apparatus, and computer program product recorded on computer storage device(s), which, alone or in combination, configured to perform the actions of the methods. For a computer system configured or to be configured to perform operations or actions, the computer system has installed on it software, firmware, hardware, or their combinations that in operation cause the computer system to perform the operations or actions. For a computer program product configured or to be configured to perform operations or actions, the computer program product includes instructions that, when executed, by a hardware processor, cause the hardware processor to perform the operations or actions.

[0007] Further aspects, features, and advantages of the present disclosure, as well as the structure and operation of the various aspects of the present disclosure, are described in detail below with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate aspects of the present disclosure and, together with the description, further serve to explain the principles of the disclosure and to enable a person skilled in the art(s) to make and use the aspects.

[0009] FIG. 1 illustrates a block diagram of a system for predicting quality of base calling, according to some aspects.

[0010] FIG. 2A is a flow chart illustrating a method for predicting quality of base calling, according to some aspects.

[0011] FIG. 2B is a flow chart illustrating a method for predicting quality of base calling, according to some aspects.

[0012] FIGS. 3 A-3B show flow charts illustrating a method for generating a look-up table of quality scores for predicting quality of base-calling in sequencing runs, according to some aspects.

[0013] FIG. 4 illustrates a block diagram of a computer system for predicting quality of base calling, according to some aspects.

[0014] FIG. 5 illustrates tables of base calls and quality scores for generating the look-up table of quality scores for predicting quality of base-calling, according to some aspects.

[0015] FIGS. 6A-6B illustrate exemplary quality scores predictions across multiple channels (FIG. 6A) and within a single channel (FIG. 6B) in multiple flow cycles of a sequencing run, according to some aspects.

[0016] FIGS. 7A-7D illustrate the prediction of quality scores across multiple channels (FIG. 7D) and within the same channel (FIG. 7C) in comparison with recalibrated quality scores, according to some aspects.

[0017] FIGS. 8A-8B illustrate accuracy of the prediction of quality scores in comparison with recalibrated quality score across multiple channels and within the same channel, according to some aspects.

[0018] FIG. 9 is a schematic showing an exemplary linear single stranded library molecule, according to some aspects.

[0019] FIG. 10 is a schematic showing an exemplary linear single stranded library molecule, according some aspects.

[0020] FIG. 11 is a schematic of various exemplary configurations of multivalent molecules, according to some aspects.

[0021] FIG. 12 is a schematic of an exemplary multivalent molecule comprising a generic core attached to a plurality of nucleotide-arms, according to some aspects.

[0022] FIG. 13 is a schematic of an exemplary multivalent molecule comprising a dendrimer core attached to a plurality of nucleotide-arms, according to some aspects.

[0023] FIG. 14 shows a schematic of an exemplary multivalent molecule comprising a core attached to a plurality of nucleotide-arms, where the nucleotide arms comprise biotin, spacer, linker and a nucleotide unit, according to some aspects.

[0024] FIG. 15 is a schematic of an exemplary nucleotide-arm comprising a core attachment moiety, spacer, linker and nucleotide unit, according to some aspects.

[0025] FIG. 16 shows the chemical structure of an exemplary spacer (top), and the chemical structures of various exemplary linkers, including an 11 -atom Linker, 16-atom Linker, 23 -atom Linker and an N3 Linker (bottom) , according to some aspects.

[0026] FIG. 17 shows the chemical structures of various exemplary linkers, including Linkers 1-9, according to some aspects.

[0027] FIG. 18 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units, according to some aspects.

[0028] FIG. 19 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units, according to some aspects.

[0029] FIG. 20 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units, according to some aspects.

[0030] FIG. 21 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units, according to some aspects.

[0031] FIG. 22 shows the chemical structure of an exemplary biotinylated nucleotide-arm, according to some aspects.

[0032] FIG. 23 shows a schematic illustration of one embodiment of the flow cell in which the support comprises a glass substrate and alternating layers of hydrophilic coatings which are covalently or non-covalently adhered to the glass, and which further comprises chemicallyreactive functional groups that serve as attachment sites for oligonucleotide primers.

[0033] FIG. 24 shows an exemplary support with multiple tiles for immobilized polonies or clusters, according to some aspects.

[0034] FIG. 25 A shows exemplary quality scores of a sequencing run predicted using the full look-up table; in this case; the quality scores have a correlation with the sequencing cycle number and the predicted quality of base calling decreases gradually as the cycle number increases in the sequencing run.

[0035] FIG. 25B shows correlation between actual quality scores and predicted quality scores using the full look-up table in different cycles of an exemplary sequencing run.

[0036] In the drawings, like reference numbers generally indicate identical or similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears.DETAILED DESCRIPTION

[0037] Provided herein are system, apparatus, method, and / or computer program product aspects, and / or combinations and sub-combinations thereof which enables predictive measurement of quality of base calling obtained using various sequencing methods disclosed herein. The techniques disclosed herein are useful for base-calling in next generation sequencing, and base-calling will be used as the primary example herein for describing the application of these techniques. However, such imaging analysis techniques may also be useful in other applications where spot-detection and / or charged coupled device (CCD) imaging is used.

[0038] In DNA sequencing, primary analysis can include image processing steps including but not limited to identifying the centers of clusters or polonies and generating base calls of clusters or polonies. Primary analysis can also involve the formation of a template (e.g., a polony map) for the flow cell. The template (e.g., polony map) can include the estimated locations of all detected clusters or polonies in a common coordinate system. Templates can be generated by identifying cluster or polony locations in all images (e.g., flow cell images) in the first few sequencing cycles of the sequencing process. The images may be aligned across all the sequencing cycles and / or color channels in the common coordinate system. Cluster or polony locations from different images may be merged based on proximity in the common coordinate system.

[0039] After the actual cluster or polony centers are identified in each flow cell image, base calling may be performed based on the actual clusters or polony centers. Each cluster or polony of signals may be used to generate a single base call in one sequencing cycle. The techniques disclosed herein may be used to predict the quality of base calling before the actual base calling has been performed. In other words, the techniques disclosed herein may be used to predict how reliable and / or accurate the base calls can be before the actual base calling. A variety of algorithms exist for calculating the quality of base calling. These existing algorithms suffer from various shortcomings. For example, predicting quality score using neural networks may add computational burden to the processing hardware, and cause undesired delay in the estimation and downstream operations. As another example, quality score measurement that relies on downstream alignment information may not be able to provide predictive measurement of the base calls before the actual base calling has been made. Additionally, the accuracy and reliability of existing algorithms needs to be improved without sacrificing computational simplicity and processing time.

[0040] In NGS sequencing, certain errors such as library preparation errors may occur more frequently in some sequencing cycles than in other cycles of a sequencing run. Traditional quality prediction models or look-up tables for predicting base calling quality may be contaminated or negatively impacted when trained based on images from all cycles of a sequencing run acquired with such library preparation errors. As a result, prediction using traditional prediction models or look-up tables may cause overestimation of quality in cycles that experience such errors, e.g., library preparation errors or other sequencing errors, and underestimation of quality in cycles that are less affected by such errors, thereby causing inaccurate quality estimation of base calling.

[0041] To solve this inaccuracy problem associated with traditional quality prediction models or look-up tables, quality prediction models or look-up tables that are trained with flow cell images from individual sequencing cycles may significantly increase the computational burden to train a model / look-up table per cycle (e.g., a typical sequencing run may have over 100 cycles). Further, quality prediction models or look-up tables that are trained with flow cell images from only individual sequencing cycles may cause the models or look-up tables to learn characteristics that are unique to individual cycles (e.g., accidental vibration, heating, accidental contamination, leakage, etc. in a cycle) but are not related to image quality or base calling quality. As a result, individual prediction models or look-up tables for each cycle may cause inaccuracy at least in certain cycles with such characteristics.

[0042] The techniques disclosed herein advantageously select a quality prediction model or a look-up table, among a plurality of pretrained quality prediction models or look-up tables, and each prediction model or look-up table is generated or populated by training using a different corresponding training dataset and different corresponding quality predictors based on that training dataset. Each corresponding training dataset may comprise flow cell images from one or more color channels and / or acquired at one or more different z-locations in one or more cycles of the sequencing run, e.g., cycle 1, cycle 2, cycles 2-10, etc. In some embodiments, quality prediction models or look-up tables that are trained with flow cell images from one or more color channels and / or from one or more different z locations may be from a single cycle or multiple different cycles depending on the effect of errors on different cycles.

[0043] For look-up tables, each quality score in the look-up table may be generated based on quality scores calculated using two different methods per polony (e.g., signal spot in the flow cell image(s)), e.g., a standard quality score, and a conservative quality score, without the need of using complicated neural networks. The look-up table may be pre-trained, and the prediction ofquality score using the look-up table may be conveniently and efficiently performed independently from the timing of base calling, e.g., in parallel to the sequencing operations before or during any actual base calling has been made or after base calling

[0044] The techniques disclosed herein advantageously allow selection of different predictors for different cycles during a sequencing run for predicting quality of base calling in such cycles. The techniques disclosed herein also advantageously avoid contamination of quality prediction by errors within some but not all cycles during a sequencing run and provide more accurate prediction of quality score across multiple channels and / or within a single channel comparing with prediction of quality with a full look-up table or full prediction model generated using base calling from all sequencing cycles (e.g., cycles affected by errors differently). The techniques disclosed herein advantageously improve prediction of quality for base calling in cycles that are affected by errors (e.g., sequencing errors and / or library preparation errors) that change rapidly across cycles unlike traditional signal-to-noise deterioration that occurs gradually toward the end of sequencing runs.

[0045] FIG. 1 illustrates a block diagram of a computer-implemented system 100, according to one or more aspects disclosed herein. The system 100 has a sequencing system 110 that includes a flow cell 112, a sequencer 114, an imager 116, data storage 122, and user interface 124. The sequencing system 110 may be connected to a cloud 130. The sequencing system 110 may include one or more of dedicated processors 118, Field-Programmable Gate Array(s) (FPGAs) 120, and a computer system 126.

[0046] In some aspects, the flow cell 112 is configured to capture DNA fragments and form DNA sequences for base-calling on the flow cell. The flow cell 112 may include the support as disclosed herein. The support may be a solid support. The support may include a surface coating thereon as disclosed herein. The surface coating may be a polymer coating as disclosed herein.

[0047] A flow cell 112 may include multiple tiles (e.g., imaging areas) thereon, and each tile may be separated into a grid of subtiles. Each subtile may include a plurality of clusters or polonies thereon. As a non-limiting example, a flow cell may have 424 tiles, and each tile may be divided into a 6*9 grid, therefore 54 subtiles. The flow cell image as disclosed herein may be an image including signals of a plurality of clusters or polonies. The flow cell image may include one or more tiles of signals or one or more subtiles of signals. In some aspects, each tile or subtile may include millions of polonies or clusters. As a nonlimiting example, a tile may include about 1 to 10 million of clusters or polonies. Each polony may be a collection of many copies of DNA fragments.

[0048] In some aspects, a flow cell image may be an image that includes all the tiles and approximately all signals thereon. In some aspects, a flow cell image may include some areas of a tile, e.g., a subtile or a portion of one or more subtiles. The flow cell image may be acquired from a channel during an imaging or sequencing cycle using the imager 116. Depending on the sample(s) immobilized on the support or flow cell, the flow cell images may include multiple z levels which are orthogonal to the image plane of the flow cell images, i.e., the x-y plane in FIG. 24. In particular, for three dimensional (3D) samples, e.g., cells, tissues, or other in situ samples, the flow cell images can include multiple z-levels in order to cover the whole sample(s) in 3D. The z axis can extend from the objective lens of the optical system disclosed herein to the support, e.g., flow cell. The axial axis can be orthogonal to the image plane of the flow cell images. Each z level of flow cell images may be separated from the adjacent z level(s) for a predetermined distance, for example, for about 0.1 um to about 15 urns. Each z level of flow cell images may be separated from the adjacent level(s) for 1 um to 10 urns. At each z-level, a flow cell image can be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include in its field of view at least part of one or more tiles or subtiles of the flow cell. FIG. 24 shows a portion of a flow cell 2212 with multiple tiles 2210. The image plane is defined by the x and y axis. And the z axis is orthogonal to the x-y plane. Although the flow cell images, samples, and the axial axis are described in a Cartesian coordinate system as shown in FIG. 24, any other coordinate systems can be used to define spatial locations and relationships herein. Other coordinate systems can include but are not limited to the polar coordinate system, cylindrical, or spherical coordinate systems.

[0049] In some aspects, the density of signal spots (e.g., polonies or clusters) in flow cell images is determined by the density of polonies immobilized on the flow cell. In some aspects, the density of the polonies or clusters on the flow cell (e.g., nucleic acid template molecules) are immobilized at a plurality of sites with various surface densities. In some aspects, the surface density is in a range of 102- 1015per mm2. In some aspects, the surface density may be of a mixture of nucleic acid template molecules that may be from different samples.

[0050] In some aspects, the polonies or clusters are at a density of more than 104, 105, 106, 107, 108, or 109per mm2. In some aspects, the polonies or clusters are at a high density, e.g., 1010per mm2, such that different polonies may overlap partially or completely with each other on the flow cell, and traditional flow cell imaging methods are unable to differentiate them with satisfactory sequencing analysis accuracy. In some aspects, the polonies or clusters are overloaded on the flow cell at a high density so that existing sequencing analysis methods may beinsufficient to obtain accurate and reliable sequencing analysis results. In some aspects, the polonies or clusters are overloaded at a high density so that they can be sequenced in different batches with different sequencing reagents so that only a portion of the overloaded sample are sequenced and imaged in individual flow cycles thus virtually reducing the density of template molecules in individual flow cell images. Detailed embodiments of sequencing the overloaded samples or template molecules in separate batches and detecting optical signals in such separate batches are disclosed in PCT Application No. PCT / US2023 / 065972, filed April 19, 2023, and are incorporated herein by reference in its entirety.

[0051] The sequencer 114 may be configured to flow a nucleotide mixture onto the flow cell 112, cleave blockers from the nucleotides in between flowing steps, and perform other steps for the formation of the DNA sequences on the flow cell 112. The nucleotides may have fluorescent elements attached that emit light or energy in a wavelength that indicates the type of nucleotide. Each type of fluorescent element may correspond to a particular nucleotide base (e.g., A, G, C, T). The fluorescent elements may emit light in visible wavelengths. In some aspects, the sequencer 114 and the flow cell 112 may be configured to perform various sequencing methods disclosed herein, for example, sequencing-by-avidity. For example, each nucleotide base may be assigned a color. Different types of nucleotide may have different colors. Adenine may be red, cytosine may be blue, guanine may be green, and thymine may be yellow, for example. The color or wavelength of the fluorescent element for each nucleotide may be selected so that the nucleotides are distinguishable from one another based on the wavelengths of light emitted by the fluorescent elements.

[0052] The sequencer 114 may be configured to perform one or more sequencing methods disclosed herein, including but not limited to sequencing-by-avidite.

[0053] The imager 116 may be configured to capture images of the flow cell 112 after each flowing step. In an aspect, the imager 116 is a camera configured to capture digital images, such as an active-pixel sensor (CMOS) or a CCD camera. The camera may be configured to capture images at the wavelengths of the fluorescent elements bound to the nucleotides. The images may be called flow cell images. In some aspects, the imager 116 may include one or more optical systems disclosed herein.

[0054] In an aspect, the images of the flow cell may be captured in groups, where each image in the group is taken at a wavelength or in a spectrum that matches or includes only one of the fluorescent elements. In another aspect, the images may be captured as single images that captures all of the wavelengths of the fluorescent elements.

[0055] The resolution of the imager 116 controls the level of detail in the flow cell images, including pixel size. In existing systems, this resolution is very important, as it controls the accuracy with which a spot-finding algorithm identifies the polony centers. One way to increase the accuracy of spot finding is to improve the resolution of the imager 116 (e.g., by incorporating a higher-resolution camera), or improve the processing performed on images taken by imager 116. The methods described herein may detect polony centers in pixels other than those detected by a spot-finding algorithm. These methods allow for improved accuracy in detection of polony centers without increasing the resolution of the imager 116. The resolution of the imager may even be less than existing systems with comparable performance, which may reduce the cost of the sequencing system 110. In some aspects, the resolution of the imager may be the same as existing systems but achieve improved performance (e.g., better CNR) as compared to those existing systems due to the image processing.

[0056] The image quality of the flow cell images controls the base calling quality. One way to increase the accuracy of base calling is to improve the imager 116, or improve the processing performed on images taken by imager 116 to result in a better image quality. The methods described herein select image quality parameters and signal crowdedness as predictors of quality of base calling. Such predictors may be conveniently and efficiently calculated and made available when flow cell images are acquired. These methods allow for improved accuracy and reliability in predicting quality score of base calling without the trade-off of increased computational burden. Further, since the methods disclosed here are computationally less intensive than traditional methods so that the heat dissipation by the computer / processors may be easier to manage so that it is unlikely to cause undesired change to the proper chemistry of sequencing techniques disclosed herein. These methods may be advantageously performed in parallel in the computer-implemented system 100, without interference with or delay of existing sequencing workflow of the computer-implemented system 100. The results of predicted quality score may be available before actual base calling starts in the sequencing workflow.

[0057] The sequencing system 100 may be configured to predict quality score of base calling from one or more polonies based on the flow cell images. The operations or actions for predicting quality score of one or more polonies may be performed by the dedicated processors 118, the FPGA(s) 120, the computing system 126, or a combination thereof. One or more operations or actions in methods 200 300 disclosed herein may be performed by the dedicated processors 118, the FPGA(s) 120, the computing system 126, or a combination thereof. In some aspects, which operations or actions are to be performed by performed by the dedicatedprocessors 118, the FPGA(s) 120, the computing system 126, or their combinations may be determined based on one or more of: a computation time for predicting quality score, a computation time for generating a look-up table, the complexity of computation in the specific operation(s), the need for data transmission between the hardware devices, or their combinations. Predicting quality score of base calling may be performed after the flow cell images are acquired but before actual base calling of the flow cell images is performed.

[0058] The computing system 126 may include one or more general purpose computers that provide interfaces to run a variety of program in an operating system, such as Windows™ or Linux™. Such an operating system typically provides great flexibility to a user.

[0059] In some aspects, the dedicated processors 118 may be configured to perform operations of predicting quality scores as described herein. They may not be general-purpose processors, but instead custom processors with specific hardware or instructions for performing those steps. Dedicated processors directly run specific software without an operating system. The lack of an operating system reduces overhead, at the cost of the flexibility in what the processor may perform. A dedicated processor may make use of a custom programming language, which may be designed to operate more efficiently than the software run on general purpose computers. This may increase the speed at which the steps are performed and allow for real time processing.

[0060] In some aspects, the FPGA(s) 120 may be configured to perform operations in the methods for predicting quality scores herein. An FPGA is programmed as hardware that will only perform a specific task. A special programming language may be used to transform software steps into hardware componentry. Once an FPGA is programmed, the hardware directly processes digital data that is provided to it without running software. The FPGA instead uses logic gates and registers to process the digital data. Because there is no overhead required for an operating system, an FPGA generally processes data faster than a general-purpose computer.Similar to dedicated processors, this is at the cost of flexibility.

[0061] The lack of software overhead may also allow an FPGA to operate faster than a dedicated processor, although this will depend on the exact processing to be performed and the specific FPGA and dedicated processor.

[0062] A group of FPGA(s) 120 may be configured to perform the steps in parallel. For example, a number of FPGA(s) 120 may be configured to perform a processing step for an image, a set of images, or a polony location in one or more images. Each FPGA(s) 120 may perform its own part of the processing step at the same time, reducing the time needed to processdata. This may allow the processing steps to be completed in real time. Further discussion of the use of FPGAs is provided below.

[0063] Performing the processing steps in real time may allow the system to use less memory, as the data may be processed as it is received. This improves over conventional systems may need to store the data before it may be processed, which may require more memory or accessing a computer system located in the cloud 130.

[0064] In some aspects, the data storage 122 is used to store information used in the predicting the quality scores. This information may include the images themselves or information derived from the images (e.g., pixel intensities, colors, etc.) captured by the imager 116. The DNA sequences determined from the base-calling may be stored in the data storage 122. As an example, the look-up table generated using method 300 may be stored in the data storage 122. Parameters identifying polony locations may also be stored in the data storage 122.

[0065] The user interface 124 may be used by a user to operate the sequencing system or access data stored in the data storage 122 or the computer system 126.

[0066] The computer system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. It may also perform steps in the prediction of quality scores and subsequent operations leading to actual base-calling. In some aspects, the computer system 126 is a computer system 400, as described in more detail in FIG. 4. The computer system 126 may store information regarding the operation of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 may be configured to pass information between the sequencing system 110 and the cloud 130.

[0067] As discussed above, the sequencing system 110 may have dedicated processors 118, FPGA(s) 120, or the computer system 126. The sequencing system may use one, two, or all of these elements to accomplish necessary processing described above. In some aspects, when these elements are present together, the processing tasks are split between them. For example, the FPGA(s) 120 may be used to perform the methods for generating the look-up table, while the computer system 126 may perform other processing functions for the sequencing system 110. Those skilled in the art will understand that various combinations of these elements will allow various system aspects that balance efficiency and speed of processing with cost of processing elements.

[0068] The cloud 130 may be a network, remote storage, or some other remote computing system separate from the sequencing system 110. The connection to cloud 130 may allow accessto data stored externally to the sequencing system 110 or allow for updating of software in the sequencing system 110.

[0069] FIG. 2 A shows a flow chart illustrating a method 200 for predicting quality of base calling in various sequencing applications.

[0070] The method 200 may include some or all of the operations disclosed herein. The operations may be performed in the order that is described herein, but is not limited to the order that has been described herein.

[0071] The method 200 may be performed by one or more processors disclosed herein. In some aspects, the processor may include one or more of: a processing unit, an integrated circuit, or their combinations. For example, the processing unit may include a central processing unit (CPU) and / or a graphic processing unit (GPU). The integrated circuit may include a chip such as a field-programmable gate array (FPGA). In some aspects, the processor may include the computing system 400.

[0072] In some aspects, some or all operations in method 200 may be performed by the FPGAs. In some aspects, when some operations are performed by FPGAs, the data after an operation performed by the FPGA may be communicated by the FPGAs to the CPUs so that CPUs may perform subsequent operation(s) in method 200 using such data. In some aspects, all the operations in method 200 may be performed by CPUs. Alternatively, the operations performed by CPUs may be performed by other processors such as the dedicated processors, or GPUs.

[0073] In some aspects, the method 200 is performed during a cycle N after the flow cell images in cycle N have been acquired. In some aspects, the method 200 is performed in cycle N so that base calling of cycles prior to cycle N (e.g., cycle N-l) has already been performed, while sequencing reactions and subsequent base calling of cycle N (and similarly, cycle N+l, N+2) is yet to be performed. In some aspects, cycle N is the current cycle. While sequencing of the current cycle N is being performed, the flow cell images or information derived therefrom prior to cycle N may have been determined or saved to a memory or a data storage device disclosed herein. Such information may be used in one or more operations in method 200, e.g., operations 220-240. Such information may be loaded from the memory or data storage device disclosed herein. N may be any integer that is greater than 2. For example, for a sequencing run, N may be any integer from 1 to 150 or 1 to 300.

[0074] In some aspects, the method 200 comprises an operation of ranking candidate predictors based on their corresponding effects on quality of base calling. The candidatepredictors may be indicative of the image quality of the one or more flow cell images that base calls are going to be made from. The candidate predictors may be indicative of polony density within the flow cell images or region(s) within the images. The candidate predictors may be indicative of signal errors that may negatively impact image quality for base calling, such errors may be based by color, phasing, pre-phasing, background noise, image distortion or transformation, etc. The candidate predictors may be indicative of characteristics that may influence image quality of the flow cell images including but not limited to the type of flow cells 112, the imager 116, the origin or of the DNA sequences, etc.

[0075] For example, the predictors are indicative of a signal to noise (SNR) ratio of one or more regions in flow cell image(s). As another example, the predictors are indicative of a contrast to noise (CNR) ratio of one or more regions in flow cell image(s). The corresponding sensitiveness of candidate predictors may be determined by the extent of their changes in correspondence to CNR or accuracy changes of base calling of the flow cell images. For example, with an identical CNR change in the image, if the percentage of change in the first candidate predictor is greater than that of a second candidate predictor, the first candidate predictor may be ranked higher than the second candidate predictor. With an identical accuracy change in the base calling, if a percentage change in a first candidate predictor is greater than a second candidate predictor, the first candidate predictor may be ranked higher than the second candidate predictor. As another example, if the accuracy of base calling decreases as the density of polonies within a region of a fixed size increase, a candidate predictor may be selected that corresponds to the density of polonies. If sensitivity of two or more candidate predictors are substantially similar, e.g., with less than 1%, 2 %, 3% or 5% difference, their ranking may be identical.

[0076] In some aspects, the method 200 comprises an operation 220 of selecting the predictors. The predictors may be selected based on the ranking of their corresponding effects on the quality of base calling. In some aspects, the operation 220 of selecting the predictors may be based on flow cell images acquired in the corresponding sequencing run in one or more cycles. For example, the selection may be based on the flow cell images of the current cycle, N, or alternative, on cycles 1 to N or 1 to N-l, of a current sequencing run being performed. As another example, the selection may be based on the flow cell images acquired from completed sequencing run(s) of the same sequencing system. In some aspects, the operation 220 may be performed before the current sequencing run starts. In some aspects, the operation 220 is performed with flow cell images that are not actually acquired using the sequencing system, e.g.,computer-simulated flow cell images. In some aspects, the operation 220 is optional and selection of predictors can be based on user’s input of one or more predictors.

[0077] The quantity of selected predictors may be 2, 3, 4, 5, 6, or even more. Having more predictors may add additional computation time and complexity in generating the look-up table in the method 300 described in FIG. 3. The quantity of predictors may be manually selected based on a user indication received at the system 100. Alternatively, the quantity of selected predictors may be automatically selected based on the characteristics of the system 100 or the specific custom application. For example, the quantity of predictors may be selected to be less than 7 to meet the time constraint in estimating the quality score or in generating the look-up table. Having more predictors may not necessarily improve the estimation of quality of base calling by a predetermined amount. The methods disclosed herein determines the number of predictors based on the trade-off between increasing computational burden and delay in generating the look-up table and quality scores and increased accuracy and sensitivity in quality score prediction.

[0078] In some aspects, for quality score prediction across multiple channels, e.g., 2, 3, 4, or even more color channels, the predictors disclosed herein include 3 different predictors. For example, the predictors include a first predictor of clarity, a second predictor of max intensity, and a third predictor of low intensity median clarity. In some aspects, for quality score prediction across multiple channels, the predictors disclosed herein include 4 different predictors. For example, the predictors include a first predictor of clarity, a second predictor of max intensity, a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

[0079] In some aspects, the predictors disclosed herein include 3 or 4 predictors selected from a list including but is not limited to: a first predictor of clarity, a second predictor of max intensity, a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

[0080] The method 200 may further comprise operation 230 of determining a corresponding value for each predictor from one or more flow cell images. The operation 230 can be performed based on one or mor cycles of the sequencing run. When the prediction of quality score is performed for a particular cycle, the one or more flow cell images are within the same cycle, and the resulting predictors are determined using flow cell images within the particular cycle. In some aspects, the operation can be performed per cycle for a number of cycles in the sequencing run, the number can be any integer no greater than the largest possible number of cycle in the sequencing run.

[0081] Within individual cycles, the value for each of the one or more predictors can be determined for each individual polony or cluster. In some aspects, values of one or more predictors can be determined for a plurality of individual polonies or clusters (e.g., within a subtile or a tile). In some aspects, the method 200 may include, before operation 230, an operation of detecting optical signals, by the optical system herein, from 2D sample(s) or 3D samples located on the flow cell. In some aspects, the method 200 may include an operation of acquiring 2D flow cell images, a z-stack of 2D flow cell images, and / or 3D flow cell images of the sample(s).In some aspects, the method 200 may include an operation of acquiring 2D flow cell images, a z-stack of 2D flow cell images, and / or 3D flow cell images of the sample(s) with low or unbalance diversity. In some aspects, the method 200 may include an operation of acquiring 2D flow cell images, a z-stack of 2D flow cell images, and / or 3D flow cell images of the sample(s) with a polony or cluster density that is greater than 108, 109, or 1010per mm2. In some aspects, the method 200 may include an operation generating, by the sequencing system, flow cell images of one or more sample(s) immobilized on the flow cell by conducting one or more cycles of sequencing reactions, wherein each individual sample is immobilized on a support and comprises nucleotide template molecules therewithin. In some aspects, the flow cell images are acquired at a z-stack of different z-locations.

[0082] In some aspects, the flow cell images used in operation 230 are 2D flow cell images, 2D projection of flow cell images, and / or flow cell images at different z-locations forming a z- stack to cover 3D sequencing sample(s). In some aspects, the method 200 may advantageously allow prediction of quality of base calls of a plurality of polonies or clusters where the polonies or clusters are in 2D flow cell images. In some aspects, the method 200 may allow prediction of quality of base calls in 2D project! on(s) of flow cell images at different z -locations. Such 2D stack of flow cell images may be acquired to cover in situ or otherwise 3D sample(s), such as single cells. In some aspects, the method 200 may allow prediction of quality of base calls in flow cell images at different z-locations. Such 2D stack of flow cell images may be acquired to cover in situ or otherwise 3D sample(s), such as single cells.

[0083] In some aspects, the flow cell images in operation 230 are from cycles with low or unbalanced diversity.

[0084] In some aspects, the flow cell images in operation 230 are acquired from sample(s) overloaded on the flow cell so that existing sequencing systems and methods may not be capable of generating accurate and reliable base calls. The method 200 may allow prediction of quality of base calls of a plurality of polonies or clusters even if the polonies or clusters are of low orunbalanced diversity in sequencing cycle(s). In some aspects, the flow cell images used in operation 230 are of low or unbalanced diversity in one or more sequencing cycles.

[0085] The nucleotide diversity of a population of immobilized polonies or clusters can refer to the relative proportion of nucleotides A, G, C and T / U that are present in each sequencing cycle. An optimal high diversity library can generally include approximately equal proportions of all four types of nucleotides represented in each cycle of a sequencing run. A low diversity library can generally include a high proportion of certain nucleotide types and low proportion of other nucleotide types in one or more sequencing cycles.

[0086] In some aspects, the polonies or clusters being sequenced in a flow cycle may have a certain nucleotide diversity. The nucleotide diversity of a population of nucleotide acid molecules, e.g., polonies or clusters, can refer to the relative proportion of nucleotides A, G, C, and T / U that are present in each flow cycle. An optimally high or balanced diversity data can generally have approximately equal proportions of all four nucleotides represented in each flow cycle of a sequencing run. A low or unbalanced diversity data can generally include a high proportion of certain nucleotides and low proportion of other nucleotides in some flow cycles of a sequencing run, e.g., less than 10% of the total number of all 4 nucleotides. As a result, images corresponding to the high portion of certain nucleotides can have more brighter spots (polonies) than images corresponding to the low portion of certain nucleotides. As an example of low or unbalanced diversity data in a flow cycle, the bases A, T, C, G can be about 1%, about 2%, about 1%, and about 95%, respectively, of the total number of polonies, in a certain flow cycle. As another example of low or unbalanced diversity data, the bases A, T, C, G in polonies at multiple flow cycles can be about 2%, about 5%, about 10%, and about 83%, respectively. In embodiments where low or unbalanced diversity data is present in a sequencing run and is imaged for sequencing analysis, image registration failure may occur because image(s) from one or more channels are too dark (e.g., signal spots of polonies are too sparse and / or dim) comparing with images acquired from other channels.

[0087] The balanced diversity of polonies with nucleotide bases of A, G, C and T / U among nucleic acid template molecules may comprises: a percentage of (1) a number of each individual type of nucleotide bases to (2) a total number of bases in the one or more cycles. The percentage can be more than 10%, 15%, or 20%. For example, the balanced diversity of polonies in one particular cycle may include a number of nucleotide bases A, G, C, T that is 26%, 15%, 27%, and 32% respectively of the total number of all nucleotide bases (or polonies) of the cycle.

[0088] The unbalanced diversity of polonies with nucleotide bases of A, G, C and T / U among nucleic acid template molecules may comprise: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases is less than 20%, 15%, 10%, or 5% in the one or more cycles. For example, the unbalanced diversity of polonies within a particular cycle includes a number of polonies with nucleotide base A that is about 5% of the total number of all nucleotide bases (polonies) of the cycle. As another example, the unbalanced diversity of nucleotide bases includes a number of nucleotide base C that is about 8% and T that is about 1% of the total number of the template molecules of a cycle.

[0089] In addition to the base biases affecting diversity, plexity can also be a factor that when plexity is lower than a number, e.g., 8 or 16, the signal could be of low diversity. For example, in a 2-cycle sequence, all polonies are of AT or TG or GC or CA. It is 25% for every base in every cycle, but its plexity is less than 8, and the sequence is not all random. In some aspects, the methods 500 is configured to register flow cell images even if the polonies are of low diversity or low plexity.

[0090] In general, plexity can indicate source(s) of the sample. A uniplex sample may include DNA fragments or molecules from a same sample region in a genome or a same sample source. A multiplex sample may include DNA fragments or molecules from different sample sources, e.g., liver, kidney, heart, cancerous tissue, etc., or from one or more sample regions in the genome.

[0091] In some aspects, the flow cell images used in operation 230 are acquired from overloaded sample(s) on the flow cell. The polonies or clusters of the overloaded sample may be at a density of more than 103, 104, 105, 106, 107, 108, or 109per mm2. In some aspects, the method 200 may allow prediction of quality of base calls of a plurality of polonies or clusters that are in flow cell images, either 2D or 3D, obtained from flow cell(s) with a density of polonies or clusters that is higher than the maximum polony or cluster density used in existing sequencing analysis. In some aspects, the surface density of polonies or clusters on the flow cell is in a range of 102- 1015per mm2or 105- 1015per mm2. In some aspects, the surface density may be of a mixture of nucleic acid template molecules that may be from different samples. In some aspects, the polonies or clusters are at a density of more than 104, 105, 106, 107, 108, or 109per mm2. In some aspects, the polonies or clusters are at a high density, e.g., 106, 107, 108, or 109per mm2, such that different polonies may overlap partially or completely with each other on the flow cell, and traditional flow cell imaging methods are unable to differentiate them with satisfactory sequencing analysis accuracy. In some aspects, the polonies or clusters are overloaded on theflow cell at a high density so that existing sequencing analysis methods may be insufficient to obtain accurate and reliable sequencing analysis results. In some aspects, the polonies or clusters of the overloaded sample may be at a density that is 2x, 3x, 4x, 5x, 8x, lOx, 20x, 30x, 40x, or more of standard or average polony densities that existing sequencing systems can handle with satisfactory accuracy (e.g., 90% Q30) .

[0092] In some aspects, the operation 230 may include an operation of determining whether a signal spot in the flow cell image(s) is a polony (or cluster) or not. Such determination may be based on a polony map. The polony map may contain approximately all the signal spots that are polonies or clusters within a 2D or 3D field of view. The signal spots that are not included in the polony map may be excluded to avoid false identification of polonies and subsequent quality score error(s) of such falsely identified polonies. For example, in flow cell images of 3D sample(s), a single polony may appear in flow cell image(s) at different z-levels, some version(s) of the single polony may be in-focus, some other version(s) of the polony may be out-of-focus. The polony map may only include the in-focus version of the single polony, and the duplicate(s) are removed. The polony map may facilitate accurate identification of individual polonies or clusters and improve quality score accuracy and reliability than quality score prediction methods without using the polony map.

[0093] In some aspects, the operation 230 may include an operation of determining or obtaining a polony map. In some aspect, the operation 230 may include an operation of determining or obtaining a polony map based on flow cell images of multiple channels in one or more cycles in the sequencing run. The polony map may be for 2D or 3D sample(s) on the flow cell. The polony map may contain information related to individual polonies within a field of view, e.g., a tile or a subtile. The polony map may exclude signal spots that are not polonies but from other sources such as background or artifacts. The polony map may exclude duplicate polonies that may appear in more than one flow cell image at different z locations. Such information in the polony map may include spatial location of individual polonies, intensity of in dividual polonies, size of individual polonies, unique identifier of individual polonies especially there are some overlapped polonies sharing similar spatial locations, etc. In some aspects, the polony map may be in a common coordinate system shared by all the flow cell images. In other words, flow cell images per channel per cycle are registered to the common coordinate system so that the same polony or cluster appear at the same spatial location across different channels and / or different flow cycles. In some aspects, the polony map may include a list of entries, each entry corresponding to an individual polony or cluster. Each entry may include information suchas spatial location of individual polonies, intensity of in dividual polonies, size of individual polonies, unique identifier of individual polonies especially there are some overlapped polonies sharing similar spatial locations, etc. The polony map in the form of a list of information may be computationally simple and convenient to use by the methods disclosed herein, and may advantageously improve computational complexity and reduce computational time in generating quality score predictions. The details of determining a polony map is disclosed in U.S. Patent No. 11,200,446, and is herein incorporated by reference in its entirety. In some aspects, the polony map may be obtained via electronic communication with a processor or a memory device that is external to the sequencing system, e.g., a CPU or a cloud. In some aspects, the operation of determining the polony map may include an operation of removing duplicate polonies or clusters. In some aspects, the operation of removing duplicate polonies or clusters may include an operation of generating a 3D polony map. The details of generating the 3D polony map is disclosed in U.S. Application Nos.18 / 078,797 and 18 / 078,820, and PCT application No. PCT / US23 / 76125, and are incorporated herein by reference in their entireties.

[0094] In some aspects, the operation of determining the polony map may include an operation of removing polonies or clusters that are outside the cell boundaries. The cell boundaries can be determined based on flow cell images acquired using the sequencing system when the flow cell images include some level of background signal from cell boundaries. Alternatively, or in combination, the cell boundaries can be determined by staining the cell membranes after the sequencing run, optionally keeping the polonies or clusters at the spatial locations, e.g., immobilized on the flow cell and then the flow cell is mechanically fixed to the sample stage, and by determining the cell boundaries using various image processing techniques.

[0095] In some aspects, the operation of determining the polony map may include an operation of including only polonies or clusters that are inside the nucleus or the cell membranes.

[0096] The first predictor of clarity may be determined based on a ratio of a max intensity to a second max intensity. For one or more polonies or clusters within a cycle, after determining a brightest channel among multiple channels, e.g., highest average image intensity among flow cell images from 4 channels. The max intensity of the one or more polonies or clusters may be obtained from the brightest channel. The second max intensity of the same one or more polonies or clusters may be obtained from the second brightest channel.

[0097] The max intensity may be normalized. The normalization may be by a predetermined value, e.g., by 90thpercentile brightest intensity within the channel, for flow cell images of each channel. The normalization may be performed during preprocessing operations of the flow cellimages disclosed herein before prediction of quality scores. Before obtaining the max intensity, the flow cell image may be preprocessed using one or more of the preprocessing operations disclosed herein. The normalization may be performed after background subtraction, correction, and phasing and prephasing correction. After normalization, the signal intensity may be scaled to a pre-determined range, e.g., [0, 2000], [0, 2500], or [0, 3000], The predetermined range may be determined to encode the range of normalized intensities in integers. To make the predictor changing inversely and / or negatively with the quality score, the scaled intensity may be negated and added to an intensity offset.

[0098] The second predictor of max intensity may be obtained as c- cl* int / int norm), wherein c and cl may be identical or different integers, int is the intensity of the polony or cluster, and int norm is the intensity used for normalization of intensity int. For example, int norm may be the brightest intensity of the flow cell image after background subtraction and color correction. The max intensity may be obtained for each polony or a group of polonies or clusters within the flow cell image in a cycle. In other words, the max intensity may be per polony-cycle.

[0099] The second max intensity may be determined similarly as the max intensity disclosed herein, but with respect to the second brightest channel among multiple channels, e.g., the second highest average intensity among 4 channels. The second max intensity may be obtained for each polony within the flow cell image in a cycle. In other words, the second max intensity may be per polony-cycle. As such, the first predictor of clarity may be per polonycycle. The second predictor of max intensity may also be per polony-cycle. In some aspects, the first predictor of clarity may be for a group of polonies or clusters per cycle. Similarly, the second predictor of max intensity may also be for a group of polonies or clusters per cycle.

[0100] Alternatively, the max intensity of the one or more polonies or clusters may be obtained from the brightest intensity of the individual polonies or clusters across four different channels, and the second max intensity of the one or more polonies or clusters may be obtained from the second brightest intensity of the individual polonies or clusters across 4 different channels.

[0101] For quality score prediction across multiple channels, the calculation of max intensity for the second predictor of max intensity is the same as disclosed herein in the calculation of the first predictor of clarity.

[0102] In some aspects, the value of one or more predictors and the accuracy of base calling satisfies a similar pattern, which is the smaller the value of a predictor, the more accurate thebase calling. In other words, the value of the predictor is negatively and / or inversely correlated with the accuracy of base calling. For example, the first predictor of clarity comprises an inverse of the ratio of the max intensity to the second max intensity. As another example, the second predictor of max intensity includes a fixed offset value, e.g., about 2000, that is added to the negative max intensity. The fixed offset value may be customized to provide optimal quality score prediction results for different sequencing systems 110 in which one or more subunits of the system 110 is different from a reference sequencing system 110, e.g., a next generation sequencing system-by-avidite as disclosed herein.

[0103] The third predictor of low intensity median clarity may be indicative of density of polonies or clusters in the flow cell images. The third predictor of low intensity median clarity may comprise a median clarity for a selected number of polonies or cluster, e.g., in a subtle. A flow cell image may be separated into multiple tiles, and each tile is an imaging area that may have multiple subtiles, and each subtile may have an equal or different number of polonies. As an example, a flow cell may include 424 tiles, and depending on the density of polonies in the tile, a total number of polonies may be 3 to 5 million of polonies per tile. A flow cell image disclosed herein may be an image of a single tile. Alternatively, a flow cell image disclosed herein may be an image of multiple tiles. To make the third predictor compatible with the inverse or negative pattern described above, it is inversely proportional to the median clarity of selected polonies per tile-cycle in the one or more flow cell images. The selected number of polonies are dim polonies selected from the flow cell image. A dim polony may be a polony with an intensity lower than a pre-determined percentage of the brightest signal intensity in a tile or a subtile. For example, the pre-determined percentage may be 15%, 20%, 25%, 30%, or any other percentages lower than 45%. Alternatively, a dim polony belongs to the darkest population of polonies with a preselected percentage. The pre-selected percentage may be 8%, 10%, 12%, or 15% of the total population of polonies within the subtile or tile. A median of all clarity values from the dim polonies or clusters, subtiles, or tiles may be used to generate the third predictor of low intensity median clarity.

[0104] The median clarity of dim polonies or clusters, subtiles, or tile may be further inverted so that the third predictor is smaller when the quality score estimation is more accurate. The third predictor of low intensity median clarity may be calculated as c*median(max r2 / max j), wherein c is a constant, max r is maximal signal intensity of the dim polonies within a selected region of the flow cell image, and max r 2 is the second maximal signal intensity of the dim polonies within the same selected region. The third predictor of low intensity median clarity may becalculated as c / median(max r / max _r2) , wherein c is a constant, max r is maximal signal intensity of the dim polonies within a selected region of the flow cell image, and max r 2 is the second maximal signal intensity of the dim polonies within the same selected region. In some aspects, the third predictor of low intensity median clarity may be calculated using certain constants, e.g., as c*median[(max r2+a) / (max r+b)] , where a or b can be non-zero numbers.

[0105] The fourth predictor of phasing and prephasing may be determined as an average percentage of phasing and prephasing in a selected number of polonies in a region of the flow cell images. The region of the flow cells are tiles or subtiles as described herein. The corresponding value of the fourth predictor of phasing and prephasing is based on multiple polonies in the flow cell image. In some aspects, the fourth predictor of phasing and prephasing can be per tile-cycle. The average percentage of phasing and prephasing can be indicative about how much image intensity is affected (e.g. decreased) by the phasing and prephasing effect.

[0106] In some aspects, the operation 230 may be performed during a current cycle N. In some aspects, the operation 230 may use flow cell images from current cycle N. In some aspects, the operation 230 may use flow cell images only from current cycle N. In some aspects, the operation 230 may use flow cell images from one or more cycles of cycles 1 to cycle N.

[0107] In some aspects, one or more operations of method 200 is performed while a sequencing run is being performed. For example, operations 230 and 240 can be performed after the flow cell images in cycle N have been acquired, while other flow cell images in subsequent cycles are being acquired or are yet to be acquired. In some aspects, one or more operations of method 200 is performed while a sequencing run is being performed so that the method 200 and the sequencing run are being performed in parallel to speed up the sequencing analysis of the sequencing run.

[0108] The method 200 may further comprise one or more preprocessing operations for preprocessing the flow cell images, before or after operation of ranking the predictors. Such preprocessing operations includes one or more of: (1) background subtraction; (2) color correction; (3) phasing or prephasing correction; and (4) and normalization. Such pre-processing operations may include image registration to align flow cell image across different channels in a same coordinate system so that individual polonies or clusters are aligned spatially across different channels. The preprocessing operations are configured to improve image quality and reduce errors that might be cause by noise and biases introduced during imaging using the sequencer and / or imager 116.

[0109] The flow cell images disclosed herein may be after one or more of the preprocessing steps disclosed herein. The flow cell images may comprise images acquired from multiple channels. For example, the flow cell image may include images from all 4 different channels. In some aspects, operation (4) normalization may include normalization of intensities of flow cell images, e.g., in each channel by 90thpercentile of the corresponding channel.

[0110] One channel disclosed herein corresponds to a wavelength spectrum of a fluorescent element representing a nucleotide base in a set of nucleotide bases.[OHl] The method 200 may further comprises operation 240, which is determining a quality score from a look-up table based on the corresponding value for each of the plurality of predictors. The look-up table may be generated using method 300 in FIG. 3. The look-up table may include a plurality of dimensions, each dimension corresponding to a selected predictor. The look-up table may be prepopulated with values by using a reference base call set and a training dataset.

[0112] The look-up table may be generated based on the selection of predictors in operation 220. In other words, the operation 220 may need to be performed before operation 240 and also before method 300 which generates the look-up table. In some aspects, multiple look-up tables can be generated using method 300, each based on a possible selection of predictors that may likely occur in operation 220 such that operation 220 can be performed after the look-up tables have been generated, and the operation 220 can include selecting a corresponding look-up table that corresponds to the selection of predictors in operation 220.

[0113] FIGS. 6A-6B show predicted quality scores generated using the methods, systems, and media disclosed here. The predicted quality scores are plotted for multiple polonies each cycle from cycle 1 to cycle 150. FIG. 6A shows predicted quality scores across four different channels, and FIG. 6B shows predicted quality scores for each individual channel. The spread of light colors above the dark / black background shows that the predicted quality score in individual channels spread in a range from about 35 to 45 or more in earlier cycles, e.g., cycle 6 to about cycle 120, and the average quality score slightly decreases in later cycles, e.g., cycles 130-150.

[0114] In some aspects, unlike traditional quality score that predict a same quality score for all channels, the method 200 may be used for predicting quality of base calling for each individual channels. Such aspects advantageously remove the need to normalize intensity of flow cell images but instead may work with raw image intensity of a specific channel, so as to provide computationally simplicity and channel-specific estimation over traditional methods. Further, such aspects may work with a larger range of image intensity in the flow cell images fordetermine base calling, thereby providing a larger dynamic range and more accurate quality estimation than traditional methods.

[0115] In such aspects, the preprocessing operations of predicting the quality score within a single channel is different from that of prediction of quality score across multiple channels. In some aspects, preprocessing operations of the flow cell images may comprise one or more of: (1) background subtraction; (2) color correction; and (3) phasing or prephasing correction, but not (4) normalization. The preprocessing operations may also include image registration of flow cell images within a cycle to the template as disclosed herein or a common coordinate system. In some aspects, one or more of the preprocessing operations is performed among flow cell images from the same channel. In some aspects, one or more of the preprocessing operations is performed among flow cell images from different channels. In some aspects, (3) phasing and prephasing correction may be performed with respect to individual channels so that each flow cell image per channel may have its independent phasing and / or prephasing correction. In other aspects, a median or average of the phasing and / or prephasing correction may be applied to all the flow cell images from different channels in one or more cycles. In some aspects, the preprocessing operations may include extracting polonies from normalized flow cell images and register them within a common coordinate system. After normalization, the image intensity of different channels may be in similar intensity ranges. In other words, normalization reduces the variation in the ranges of image intensities across channels. An example technique for extracting the plurality of polonies in a flow cell is described in U.S. Patent No. 11,200,446, which is hereby incorporated by reference in its entirety. The image intensities may be determined as described in, for example, U.S. Patent No 11,200,446, incorporated by reference herein in its entirety.

[0116] Operations 220, 230, and 240 and other operations disclosed herein in relation to method 200 are similar for predicting quality score across multiple channels or within a single channel. However, the second predictor of max intensity is used in the methods 200 for predicting the quality across multiple channels, and it may be replaced by the second predictor of max cc intensity when the prediction is for a single channel. The replacement of max intensity by max cc intensity may be only within the calculation of the second predictor, the calculation of predictors other than the second predictor may stay the same as those disclosed herein for making predictions across multiple channels.

[0117] Alternatively, in some aspects, the replacement of max intensity to max cc intensity may be in two or more predictors, for example, in the first predictor of clarity and the secondpredictor. The other predictors without using max intensity may stay the same as the as those disclosed herein for making predictions across multiple channels.

[0118] In some aspects, for quality score prediction within a single channel, the predictors disclosed herein include 3 different predictors. For example, the predictors include a first predictor of clarity, a second predictor of max cc intensity, and a third predictor of low intensity median clarity. In some aspects, for quality score prediction within a single channel, the predictors disclosed herein include 4 different predictors. For example, the predictors include a first predictor of clarity, a second predictor of max cc intensity, a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing. For quality score estimation within a single channel, the second predictor of max cc intensity is calculated differently from that of the second predictor for multiple channels. In some aspects, the predictors disclosed herein include 3 or 4 predictors selected from a list including but is not limited to: a first predictor of clarity, a second predictor of max cc intensity, a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

[0119] In some aspects, the max cc intensity may be calculated as c’- c *(int raw), wherein c’ and cl’ are identical or different constants, int raw is the raw intensity of the polony without normalization. The raw intensity may be after some preprocessing steps disclosed herein, such as background subtraction, color correction, phasing and prephasing correction, or their combinations. The signal intensity may be scaled to a pre-determined range, e.g., [0, 300], [0, 400], or [0, 500], The predetermined range may be determined to encode the range of raw intensities in integers. To make the predictor changing inversely or negatively with the quality score, the scaled intensity may be negated and added to an intensity offset.

[0120] For quality score within a single channel, when max intensity is replaced by max cc intensity in the first predictor of clarity, the first predictor of clarity may include a ratio of a max cc intensity to a second max cc intensity. After determining a brightest channel among multiple channels, e.g., highest average intensity among 4 channels. The max cc intensity may be obtained from the brightest channel. The max cc intensity may be without normalized by a predetermined value, e.g., by 90thpercentile brightest intensity within the channel, for flow cell images of each channel. Before obtaining the max cc intensity, the flow cell image may be preprocessed using one or more of the preprocessing operations disclosed herein, e.g., background subtraction, correction, and phasing and prephasing correction. After preprocessing, the signal intensity may be scaled to a pre-determined range, e.g., [0, 300], [0, 400], or [0, 500], The predetermined range may be determined to encode the range of intensitiesin integers. To make the first predictor changing inversely or negatively with the quality score, the scaled intensity may be negated and added to an intensity offset. The first predictor of max cc intensity may be obtained as c- c*(inf), wherein c is an integer, int is the intensity of the polony. The max cc intensity may be obtained for each polony or a group of polonies within the flow cell image in a cycle. The max cc intensity may be per polony-cycle. The second max cc intensity may be determined similarly as the max cc intensity disclosed herein, but with respect to the second brightest channel among multiple channels, e.g., the second highest average intensity among 4 channels. The second max cc intensity may be obtained for each polony or the same group of polonies for max cc intensity within the flow cell image in a cycle. The second max cc intensity may be per polony-cycle. As such, the first predictor of clarity may be per polony-cycle. The second predictor of max cc intensity may also be per polony-cycle.

[0121] For quality score within a single channel, when max intensity is replaced by max cc intensity in the second predictor of max intensity. The calculation of max cc intensity may be the same as disclosed in the calculation of the first predictor of clarity.

[0122] In some aspects, the value of predictors and the accuracy of base calling satisfies a similar pattern, which is the smaller the value of a predictor, the more accurate the base calling. In other words, the value of the predictor is negatively or inversely correlated with the accuracy of base calling. For example, the first predictor of clarity comprises an inverse of the ratio of the max cc intensity to the second max cc intensity. As another example, the second predictor of max cc intensity includes a fixed offset value, e.g., about 2000, that is added to the negative max cc intensity. The fixed offset value may be customized to provide optimal quality score prediction results for different sequencing systems 110 in which one or more subunits of the system 110 is different from a reference sequencing system 110, e.g., a next generation sequencing system-by-avidite as disclosed herein.

[0123] FIGS. 7A-7D show the predicted quality score generated using the methods disclosed herein in comparison with recalibrated quality scores. FIGS. 7A-7B show the average recalibrated quality score and the average predicted quality score, respectively. The predicted quality scores are generated using predictions across all 4 channels. FIG. 7B illustrates that the average predicted quality score is substantially identical across different channels. FIG. 7A illustrate that the recalibrated average quality score is different across channels, and the T channel has lowest quality score among the four channels. The average quality score of channel A has relatively higher quality score than the other three channels. FIGS. 7C-7D show the average recalibrated quality score and the average predicted quality score, respectively, obtainedusing the same flow cell images as FIGS. 7A-7B. The predicted quality scores are generated using prediction within a single channel so that the quality score prediction may be channel specific. FIG. 7D illustrates that the average predicted quality scores at different cycles are different in 4 channels. FIG. 7C illustrate that the recalibrated average quality scores are different across channels, and the T channel has lowest quality score among the four channels. Unlike FIG. 7B, such difference across channels has been reflected in the predicted average quality score in FIG. 7D.

[0124] FIGS. 8A-8B show the accuracy of predicted quality score obtained using the technologies disclosed herein. The quality score recalibration is a standard technique for evaluating predicted quality scores. Various recalibration methods may be used. For example, for a specific quality score, it can be based on all the errors and correct base call counts for that specific quality score, and the recalibrated quality score may be determined as -10 * logl0(error rate). After a mapping of the quality score to the recalibrated score, accuracy may be calculated. In this aspect, FIG. 8A show the comparison between predicted quality scores with the recalibrated quality scores. The predicted quality score values are shown along the horizontal axis, and recalibrated quality scores are shown along the vertical axis. The predicted quality score is across all 4 channels. The predicted quality score is substantially aligned with the recalibrated score in all quality score values and in different cycles. The histogram of predicted quality score value shows a peak between Q40 and Q45. FIG. 8B show the comparison between predicted quality scores with the recalibrated quality scores within a single channel. The predicted quality score is substantially aligned with the recalibrated score in all quality score values and in different cycles. The histogram of predicted quality score value shows a peak between Q40 and Q45.

[0125] In some aspects, the method 200 may include an operation of in response to determining that the determined quality score(s) in operation 240 for at least some polonies in the flow cell image(s) satisfy a predetermined quality threshold, e.g., more than 80%, 85% or 90% polonies of the total number of polonies within a predetermined field of view are above Q30, Q35, Q40, or even Q45, continuing performing sequencing reactions and detecting optical signals from the flow cell in one or more flow cycles in the sequencing run. In some aspects, the method 200 may include: in response to determining that the determined quality score(s) in 240 for at least for at least some polonies in the flow cell image(s) fail to satisfy a predetermined quality score or a different quality score (e.g., more than 40%, 50% or more polonies of the total polonies being predicted are below Q20 or QI 5), cease performing sequencing reactions and / orcease detecting optical signals from the flow cell in one or more flow cycles in the sequencing run. In some aspects, the method 200 may include: in response to determining that the determined quality score(s) in 240 for at least for at least some polonies in the flow cell image(s) fail to satisfy the predetermined quality score or a different quality score, forward the determined quality score(s) in operation 240 or otherwise related information for notification to a user and resume with the sequencing run after receiving an input from a user, e.g., sending a visual or audio message to the user, and prompt the user to enter the input about whether to continue the sequencing run or not.

[0126] FIG. 3 shows a flow chart illustrating a method 300 for generating a look-up table for predicting quality of base calling in DNA sequencing. The look-up table may have multiple dimensions and each dimension corresponds to a selected predictor. The method 300 may be performed with a selected number of cycles of sequencing, so that a lock-up table may be generated that may be used for different cycles in the same sequencing run or in different sequencing run(s). The selected number may be 100, 150, 200, or any integer number. The selected number may be 1 so that the look-up table is generated from a specific cycle of a sequencing run.

[0127] The method 300 may include some or all of the operations disclosed herein. The operations may be performed in the order that is described herein, but is not limited to the order that has been described herein.

[0128] The method 300 may be performed by one or more processors disclosed herein. In some aspects, the processor may include one or more of: a processing unit, an integrated circuit, or their combinations. For example, the processing unit may include a central processing unit (CPU) and / or a graphic processing unit (GPU). The integrated circuit may include a chip such as a field-programmable gate array (FPGA). In some aspects, the processor may include the computing system 400.

[0129] In some aspects, some or all operations in method 300 may be performed by the FPGAs. In some aspects, when some operations are performed by FPGAs, the data after an operation performed by the FPGA may be communicated by the FPGAs to the CPUs so that CPUs may perform subsequent operation(s) in method 300 using such data. In some aspects, all the operations in method 300 may be performed by CPUs. Alternatively, the operations performed by CPUs may be performed by other processors such as the dedicated processors, or GPUs.

[0130] The method 300 may comprises an operation 320 of obtaining correct base calls and erroneous bases calls by comparing training base calls to a reference base call set with reference base calls. The training base calls are generated using a training dataset with a reference base call set. The training dataset may be a pre-selected genome or DNA sequence. As a non-limiting example, the training dataset may include HG001, HG005, HG38 consensus human genome, or their combinations.

[0131] Before the operation 320 of obtaining correct base calls and erroneous bases calls, the method may comprise operation 220 of ranking the predictors as disclosed in relation to FIG. 2A as disclosed herein.

[0132] Before operation 320, the method 300 may include an operation of detecting optical signals, by the optical system herein, from 2D sample(s) or 3D samples located on the flow cell, thereby generating training flow cell images in the training dataset. In some aspects, the method 300 may include an operation of generating 2D training flow cell images from 2D sample(s). In some aspects, the method 300 may include an operation of generating 2D projection flow cell image(s) as training flow cell images from 3D sample(s). In some aspects, the method 300 may include an operation of generating a stack of training flow cell images from 3D sample(s), to cover a 3D volumetric sample, e.g., single cell(s). In some aspects, the method 300 may include an operation of generating training flow cell images, either from 2D or 3D sample(s), with low or unbalanced diversity. In some aspects, the method 300 may include an operation of generating training flow cell images, either from 2D or 3D sample(s), with a polony or cluster density of 102-1015per mm2on the flow cell. In some aspects, the method 300 may include an operation of generating training flow cell images, either from 2D or 3D sample(s), with a polony or cluster density on the flow cell that is higher than the maximum density that is used in existing sequencing analysis, e.g., 106, 107, 108or 109per mm2.

[0133] In some aspects, the method 300 may include an operation generating, by the sequencing system, training flow cell images of one or more sample(s) immobilized on the flow cell by conducting one or more cycles of sequencing reactions, wherein each individual sample is immobilized on a support and comprises nucleotide template molecules therewithin. In some aspects, the training flow cell images are acquired at a z-stack of different z-locations.

[0134] In some aspects, the method 300 may include an operation of determining whether a signal spot in the training flow cell image(s) is a polony or cluster or not. In some aspects, such determination may be based on a polony map. The polony map may contain approximately all the signal spots that are polonies or clusters within a 2D or 3D field of view. The signal spots thatare not included in the polony map may be excluded to avoid false identification of polonies and subsequent quality score error(s) of such falsely identified polonies. For example, in training flow cell images of 3D sample(s), a single polony may appear in flow cell image(s) at different z- levels, some version(s) of the single polony may be in-focus, some other version(s) of the polony may be out-of-focus. The polony map may only include the in-focus version of the single polony, and the duplicate(s) are removed. The polony map may facilitate accurate identification of individual polonies or clusters and improve quality score accuracy and reliability than quality score prediction methods without using the polony map.

[0135] In some aspects, before the operation 320, the method 300 may include an operation of determining or obtaining a polony map. The polony map may be for 2D or 3D sample(s) on the flow cell. The polony map may contain information related to individual polonies within a field of view, e.g., a tile or a subtile. The polony map may exclude signal spots that are not polonies but from other sources such as background or artifacts. The polony map may exclude duplicate polonies that may appear in more than one training flow cell image at different z locations. Such information in the polony map may include spatial location of individual polonies, intensity of in dividual polonies, size of individual polonies, unique identifier of individual polonies especially there are some overlapped polonies sharing similar spatial locations, etc. In some aspects, the polony map may be in a common coordinate system shared by all the training flow cell images. In other words, training flow cell images per channel per cycle are registered to the common coordinate system so that the same polony or cluster appear at the same spatial location across different channels and / or different flow cycles. In some aspects, the polony map may include a list of entries, each entry corresponding to an individual polony or cluster. Each entry may include information such as spatial location of individual polonies, intensity of in dividual polonies, size of individual polonies, unique identifier of individual polonies especially there are some overlapped polonies sharing similar spatial locations, etc. The polony map in the form of a list of information may be computationally simple and convenient to use by the methods disclosed herein, and may advantageously improve computational complexity and reduce computational time in generating quality score predictions. The details of determining a polony map is disclosed in U.S. Patent No. 11,200,446, and is herein incorporated by reference in its entirety. In some aspects, the polony map may be obtained via electronic communication with a processor or a memory device that is external to the sequencing system, e.g., a CPU or a cloud. In some aspects, the operation of determining the polony map may include an operation of removing duplicate polonies or clusters.

[0136] In some aspects, the operation of determining the polony map may include an operation of removing polonies or clusters that are outside the cell boundaries. The cell boundaries can be determined based on flow cell images acquired using the sequencing system when the flow cell images include some level of background signal from cell boundaries. Alternatively, or in combination, the cell boundaries can be determined by staining the cell membranes after the sequencing run, optionally keeping the polonies or clusters at the spatial locations, e.g., immobilized on the flow cell and then the flow cell is mechanically fixed to the sample stage, and by determining the cell boundaries using various image processing techniques.

[0137] In some aspects, the operation of determining the polony map may include an operation of including only polonies or clusters that are inside the nucleus or the cell membranes.

[0138] Before operation 320, the method 300 may include an operation of generating training base calls. Th training base calls are generated from training flow cell images acquired using the sequencing system, e.g., 110 in FIG. 1. In some aspects, the training flow cell images are acquired from training sample(s) that include pre-selected genome(s) or DNA sequence(s) using the optical system herein, e.g., 116. In some aspects, the training flow cell images are acquired from training sample(s) that include pre-selected in situ samples using the optical system herein.

[0139] In some aspects, the plurality of polonies or clusters in the training dataset may be extracted from specific regions of one or more tiles, e.g., each subtile, of the training flow cell images. Within each region, the polonies may be extracted with a predetermined spatial pattern or randomly.

[0140] In some aspects, the training flow cell images in the training dataset may be processed by one or more operations disclosed in relation to method 200 which functions to process training flow cell images to generate the values of various predictors, e.g., background subtraction, normalization, image registration, etc.

[0141] The method 300 may include, in the training dataset, flow cell images even if the polonies or clusters are in flow cell images generated from sequencing cycle(s) of low or unbalanced diversity.

[0142] In some aspects, the method 300 may include, in the training dataset as training flow cell images, 2D flow cell images of 2D sample(s). In some aspects, the method 300 may include, in the training dataset as training flow cell images, 2D projection(s) of flow cell images at different z -locations. Such 2D stack of flow cell images may be acquired to cover in situ or otherwise 3D sample(s), such as single cells. In some aspects, the method 300 may include, in the training dataset as training flow cell images, flow cell images at different z-locations. Such 2Dstack of flow cell images from different z-locations may be acquired to cover in situ or otherwise 3D sample(s), such as single cells.

[0143] In some aspects, the method 300 may include, in the training dataset, training flow cell images of sample(s), either 2D or 3D, with a high density of polonies or clusters on the flow cell. In some aspects, the surface density of polonies or clusters is in a range of 102- 1015per mm2. In some aspects, the surface density may be of a mixture of nucleic acid template molecules that may be from different samples. In some aspects, the polonies or clusters are at a density of more than 104, 105, 106, 107, 108, or 109per mm2. In some aspects, the polonies or clusters are at a high density, e.g., IO10per mm2, such that different polonies may overlap partially or completely with each other on the flow cell, and traditional flow cell imaging methods are unable to differentiate them with satisfactory sequencing analysis accuracy. In some aspects, the polonies or clusters are overloaded on the flow cell at a high density so that existing sequencing analysis methods may be insufficient to obtain accurate and reliable sequencing analysis results.

[0144] Before the operation 320 of obtaining correct base calls and erroneous bases calls, the method may further comprise filtering the reference base call set by removing base calls from positions in the genome(s) of the training dataset with pre-determined variants. The method may further comprise filtering the reference base call set by removing base calls from positions in the genome(s) of the training dataset with pre-determined alignment errors. The filtering of the training dataset may advantageously remove error sources that are not intrinsic to the system 100, so that errors in base calling may be more accurately and reliably attributed to the system 100 and the subunits contained therein.

[0145] Before the operation 320 of obtaining correct base calls and erroneous bases calls and after the filtering operation disclosed herein, the method 300 may further comprise obtaining training base calls in the flow cell images of the training dataset. The flow cell images may be acquired using the system 100. And the training base calls are made by using the flow cell images.

[0146] The method 300 may further comprise an operation 330 in which, a set of training regions in each flow cell image is selected, and each training region may contain multiple polonies of signals. The training regions may be obtained from an identical cycle or different cycles of sequencing. The training regions may be from the same sequence run or different sequence runs. The training regions may be from sequence insert Read 1, Read 2, or both. As an example, each flow cell image may include a number of tiles, each tile representing an imagingarea on the flow cell. Each tile may be further divided into a grid of subtiles. For example, a flow cell image may have 424 tiles, a flow cell image may be generated for each tile during a read in a cycle for a channel. The training region may include subtiles that in sum equivalent to 48 tiles of polonies. Each tile may include about 1 to 10 million of polonies. And the flow cell images may be obtained from 150 cycles, 2 reads, and 20 to 30 runs in all 4 channels. The amount of training polonies in training regions may be on the scale IO10or more. For each training polony, its predictors may be determined, and its quality scores may be calculated, the base calling may be made and aggregated in the correct bin of the look-up table. There may be a computational load that requires a computer system 126 disclosed herein, one or more dedicated processors 118, and / or FPGA(s) 120. Each region disclosed herein may be a subtile, and each subtile may include a number of polonies in the range of 5,000 to 200,000. As a non-limiting example, each subtile may include a number of polonies in the range of 10,000 to 80,000. As another nonlimiting example, each subtile may include a number of polonies in the range of 20,000 to 60,000.

[0147] To reduce computation time of performing the methods disclosed herein in a general purpose computer, some or all of the operations may be performed by the dedicated processors 118, and / or FPGA(s) 120. Performing some or all of the operations in the dedicated processors 118, and / or FPGA(s) may advantageously help with the heat dissipated by the general purpose computers which may adversely affect the temperature of the flow cells, thereby causing undesired problems in the chemistry of sequencing disclosed herein.

[0148] For quality scores across multiple channels or quality score of a single channel, the training regions include regions from multiple channels, for example, all 4 channels, 3 channels, or 2 channels.

[0149] The method 300 further comprises operation of determining a corresponding range for each predictor using the training polonies in the set of training regions. The corresponding range may encompass possible values of the predictor in all training polonies of the training regions, either from a single flow cycle, or from multiple different flow cycles. The corresponding range may start from a minimal value of the predictor in all the training polonies and end with the maximal value of the predictor.

[0150] For quality scores across multiple channels, the flow cell images may have been preprocessed using the operations disclosed herein, so that the range may be post-processed range of the training polonies. For quality score within a single channel, the flow cell image fromthe single channel may include raw image intensity, so that the range may be for the raw intensities of the training regions.

[0151] After the corresponding range is determined, the method 300 may further include an operation 350 of dividing the corresponding range for each predictor into a corresponding number of bins. The method 300 may further comprises an operation 360 of initializing of the look-up table so that the look-up table have a number of dimensions determined by the plurality of predictors. In each dimension, the look-up table may have bins corresponding to the number of bins in operation 350.

[0152] In some aspects, the corresponding number of bins do not need to be identical for two or more predictors. The increased number of bins may result in more computational complexity and time consumption in generating the look-up table. At the same time, the increased bins may provide better resolution within the range, thus it may provide more accuracy to the look-up table. As a non-limiting example, the number of bins may be 50 for different predictors. With 4 different predictors, the look-up table will be 50 x 50 x 50 x 50 in size. In some aspects, the number of bins may be 100. In some aspects, the number of bins may be an integer number in the range of 20 to 150. In some aspects, the number of bins may be an integer number in the range of 40 to 100. In some aspects, the number of bins may be an integer number in the range of 30 to 70.

[0153] The method 300 may further comprise operation 370, which is determining a first number of correct base calls and a second number of erroneous base calls in each bin for each of the plurality of predictors. For each polony in the training regions (e.g., for all the cycles or in a cycle), the corresponding values of the predictors for this polony may be obtained, the values may be used to find the corresponding bin. The base call of this polony, either correct or erroneous, by comparing to the reference base call, may be added to the correct or erroneous count of this bin. And the process may be repeated for all the polonies to generate the correct and erroneous counts for each bin. The correct base call counts may be saved in a correct base call table which is the same number of bins in each dimension as the look-up table. Similarly, the erroneous base call counts may be saved in an erroneous base call table, which is also the same size as the look-up table. Exemplary correct and erroneous base call tables are shown in TABLE 2 and TABLE 1, respectively in FIG. 5.

[0154] Using the correct base calls and erroneous based calls determined in operation 370, the method 300 may further comprise determining a first cumulative number of correct base calls and a second cumulative number of erroneous base calls in one or more bins (e.g., each bin) foreach of the plurality of predictors. The cumulative number of correct base calls including a sum of correct base calls from all the bins that has a coordinate smaller than or equal to the coordinate of the current bin. The cumulative number of erroneous base calls may be determined in a similar fashion. For example, if a bin has a coordinate of (i*, j*), then the base calls from all the bins whose coordinate ranges from 1 to i* and from 1 to j* are summed up to generate the cumulative number of correct or erroneous base calls for bin (i*, j*). The cumulative number of correct base calls and erroneous calls may be saved in a separate table, each table having an identical size as the look-up table. Exemplary cumulative correct and erroneous base call tables are shown in TABLE 4 and TABLE 3, respectively in FIG. 5. In this example, each table is with 2 dimensions, and 3 bins in each dimension.

[0155] The method 300 may further comprise operation 380, which is iterating, by the computer and until correct base calls and erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps: step 381 calculating a standard quality score and a conservative quality score for each bin of the look-up table; step 382 finding coordinates of a bin with a maximum conservative quality score in the look-up table; step 383 assigning a selected number of bins with the standard quality score that corresponds to the coordinates in a look-up table; and step 384 setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number. In some aspects, the pre-determined number may be 0. In other aspects, the predetermined number may be a small number, e.g., 2, 3, 4, or 5, that is not greater than 10.

[0156] Assuming there are 4 selected predictors, and each predictor has an identical number of n bins. In step 381, for all the bins in the look-up table, the standard quality score may be calculated using the cumulative correct and erroneous base call counts using equations (1) to (6).The cumulative erroneous base call count may be calculated as:where err () is the cumulative erroneous base call counts, and errCount() is the erroneous count in individual bins, and i ’ is in the range of [0, z], j ’ is in the range of [0, j], A:’ is in the range of [0, k\, and m ’ is in the range of [0, zw], wherein i, j, k, and m may be in the range of [0, ri\.

[0157] Similarly, the cumulative correct base call count may be calculated as:The standard error rate, e(i,j,k,m) may be obtained as:cs + err(i, j, k, m) e(i,j, k, m) = - — — - - - — — - - (3) cs + err(i,j, k, m) + corr(i, j, k, m) The conservative error rate, ec(i,j,k,m) may be obtained asWhere cs is a correction number that may be a positive constant, and cc is a different correction number that may be a different positive constant that is greater than cs. As a non-limiting example, cs may be 1, and cc may be 5. Correspondingly, the standard quality score Q(i,j,k,m) and the conservative quality score Qc(i,j,k,m) may be obtained from the standard error rate, or the conservative error rate using equations (5) or (6) as below:Q(i,j, k, m) = —10 * loglO(e(i,j, k, m) (5)Qc(i,j, k, m) = —10 * loglO(ec(i,j, k,m) (6)

[0158] Exemplary conservative and standard quality scores are shown in TABLE 6 and TABLE 5 in FIG. 5. Conservative quality score in each bin may be greater than the corresponding standard quality score

[0159] In step 382, the coordinates (i*,j*,k*, m*) of the bin with highest conservative quality score may be determined. In step 383, all the bins with coordinates z,j, k, and m less than or equal to k*, and m*, respectively, may be assigned with the standard quality score value of bin (i*,j*,k*, m*). Then in step 384, the erroneous base calls and correct base calls in all the bins with assigned quality values are set to a predetermined value. Exemplary error and correct base calls are in TABLE 1 and TABLE 2 in FIG. 5. A non-limiting example of this predetermined value is 0. In some aspects, the predetermined value is less than 2, 5, 8, 10, 20, or 30. The updated error base calls and correct base calls may be used to update the cumulative erroneous and correct base call counts using equations (1) and (2) for these bins. Subsequently, the corresponding standard and conservative error rates may be updated using equations (3) and (4). The standard and conservative quality scores may then be updated accordingly using equations (5) and (6).

[0160] After the updates, step 381 may be performed thereby starting a new iteration of steps 381-384. When all the error counts and correct counts of base calls in the bins are no greater than the predetermined threshold number, the iteration of operation 380 stops, and the look-up table may be generated with a quality score in each of its bins. The pre-determined threshold number may be customized depending on the characteristics of the sequencing application, flow cell image quality, etc. For example, the pre-determined threshold number may be 0 or 1. Anexemplary look-up table with 2 dimensions, and 3 bins in each dimension is shown in TABLE 7 of FIG. 5. This look-up table is simplified, and the quality score may not correspond to actual quality scores predicted with the system 100.Example generation of the look-up table for prediction of quality score of a single channel

[0161] In an exemplary aspect, the method 300 disclosed herein is used to generate a look-up table for a HG001 human genome. Known positions for common variants are filtered out from the human genome, and regions that present uncommon patterns like CGCGCG, are also removed. Flow cell images are acquired using the system 100. The flow cell images come from 20 runs, 2 reads, and 150 cycles of 4 different channels. In each cycle per channel per read, flow cell images acquired equal the total number of tiles on the flow cell. Among these images, a region (subtile) in some or all of the tiles are selected from each flow cell image. Each subtile may comprise 30k polonies. Three predictors, the max intensity, clarity, and low intensity median clarity are selected. The values of the first two predictors are calculated per colony-cycle. The values of the last predictor are calculated per tile-cycle. The values of the predictors are calculated after preprocessing except for the max intensity. The preprocessing includes: background subtraction, color correction, phasing and prephasing correction, and normalization using 90thpercentile of the brightest intensity of the flow cell image in the corresponding channel. The max intensity is the raw image intensity without normalization. And the range of the values for each predictor is estimated based on the different values in the selected training regions of all the cycles. The range for second predictor of max intensity and other predictors disclosed herein may encompass all the possible values in the training regions of different cycles. In some aspects, since the range is estimated, it may not encompass all the possible values either in the training regions. After the look-up table is generated, when a predictor has a value that is outside the range of the look-up table, the method 200 may select a bin that is closest to the value for locating its quality score. Then, each range is evenly separated into 50 bins. So, the look-up table has 50 x 50 in size. Base calling is performed using the system 100, and the base calls coming from polonies whose predictor values fell into a bin, are counted into either the erroneous or correct base call counts of that bin. For example, bin (25, 25, 25) may have 10% erroneous base calls among the total number base calls. The look-up table has an average quality score of Q40.Example prediction of quality score

[0162] In an exemplary aspect of method 200, flow cell images are acquired using the system 100. The flow cell images come from 30 runs, 2 reads, and 150 cycles with only signal intensities from a single channel. The raw image intensities of the flow cell images without preprocessing are used. In other words, the raw image intensities of the flow cell images without phasing and prephasing correction and normalization are used for calculating the value for each of three predictors, i.e., the max intensity, clarity, and low intensity median clarity. The value is per polony-cycle for the max intensity and clarity. The value is per-tile-cycle for low intensity median clarity. After the value is determined, for each polony, three different values of the predictors are used as coordinate to find a bin in the look-up table, and the quality score for the polony is the quality score at the coordinate of the look-up table. A coordinate of [1000, 0.5, 200] leads to bin [25, 25, 25] of the look-up table, and the quality score for base calling using the polony is Q40.Exemplary primary analysis

[0163] During a sequencing cycle of a sequencing run using the sequencing system 110 in FIG. 1, four flow cell images are generated per field of view, corresponding to the dyes used to label each type of avidite. An analysis pipeline is developed that uses the flow cell images as input to identify the polonies present on the flow cell and to assign to each polony a base call and a quality score for each cycle, representing the accuracy of the underlying base call. Briefly, intensity is extracted for each polony in each color channel based on a polony map that is pregenerated in the first 10 cycles of the same sequencing run. The intensities are corrected for color cross talk, phasing and prephasing. The intensities are normalized to facilitate cross channel comparisons. The highest normalized intensity value for each polony in each cycle across four channels determines the base call. In addition to assigning a base call, a quality score corresponding to the call confidences is also assigned. The Q score definition is defined as Q = —10 * loglO p, where p is the probability that the base call is an error. The Q score generation is based on predictors, and is encoded using the phred+33 ASCII scheme. The predictors used for quality score training are (1) the maximum intensity per polony across color channels, (2) the clarity of each polony (defined as the (4 + 1) / (B + 1), where A is the highest intensity across color channels and B is the second highest intensity), (3) the sum of phasing and prephasing estimates, and (4) the median clarity value taken across the 10% of the lowest intensitypolonies. The sequence of base call assignments and quality scores across the cycles constitute the output of the sequencing cycle run. This data is represented in standard FASTQ format with data from other sequencing cycles for compatibility with downstream analysis tools.Selecting quality prediction models

[0164] In NGS sequencing, certain errors such as library preparation errors may occur more frequently in some sequencing cycles than in other cycles of a sequencing run. Traditional quality prediction models or look-up tables may be contaminated or negatively impacted when trained based on images from all cycles of a sequencing run acquired with such library preparation errors. Further, such errors like library preparation errors or sequencing errors may occur in cycles that may not necessarily overlap with cycles that suffer from deteriorating signal to noise ratio (SNR) or contrast to noise ratio (CNR). In other words, library-prep errors (e.g., truncated adapters, index swaps, low-complexity inserts) may create systematic sequence artifacts before imaging begins, so even with high SNR, base calling may consistently yield wrong bases. Traditional quality prediction models or look-up tables that focus on prediction of base calling due to SNR and / or CNR changes (e.g., using neural networks with different parameters in low SNR and / or CNR cycles) or other causes may not be suitable for predicting quality of base calling affected by errors of different causes such as library preparation errors. As a result, prediction using traditional prediction models or look-up tables may cause overestimation of quality in cycles that experience such errors, e.g., early cycles that with highest impact from library preparation errors or other sequencing errors among all cycles, and underestimation of quality in cycles that are less affected by such errors, thereby causing inaccurate quality estimation of base calling.

[0165] To solve this inaccuracy problem associated with traditional quality prediction models or look-up tables, quality prediction models or look-up tables that are trained with flow cell images from individual sequencing cycles may significantly increase the computational burden to train a model / look-up table per cycle (e.g., a typical sequencing run may have over 100 cycles). Further, quality prediction models or look-up tables that are trained with flow cell images from only individual sequencing cycles may cause the models or look-up tables to learn characteristics that are unique to individual cycles (e.g., accidental vibration, heating, air bubbles, contamination, etc. in a cycle) but are not related to image quality or base calling quality. As a result, individual prediction model or look-up table for each cycle may cause inaccuracy inpredicting base calling quality in at least some cycles with such characteristics. Thus, it may be undesired to generate quality prediction models or look-up tables that are specific to individual cycles, and prediction of base calling using such quality prediction models or look-up tables.

[0166] In some embodiments, multiple quality prediction models or look-up tables are trained for predicting quality of base calling, e.g., quality scores, of polonies or clusters from different groups of cycles. In some embodiments, each group of cycles comprises multiple cycles to avoid learning characteristics that are unique to individual cycles. The cycles that may be included for training each prediction model or look-up table may have comparable level of error(s). In some embodiments, cycles in each group may have comparable level of error(s), e.g., the level of base calling error difference is less than ±10%, 20%, or 30%. For example, a certain library preparation error may decay with cycle increase approximately exponentially. Thus, a separate prediction model or look-up table for cycle 1 alone may allow learning of the error more accurately without mixing additional flow cell images from other cycles. As another example, as the library error decays with cycle increase, the error’s effect on later cycles, e.g., cycles 8 and up, may be comparable, thus a single prediction model or look up table can be trained by including flow cell images from such cycles.

[0167] In some embodiments, the quality prediction models or look-up tables herein are trained with flow cell images from one or more color channels and / or from one or more different z locations. The flow cell images may be from a single cycle or multiple different cycles depending on the effect of errors (e.g., library preparation errors) on different cycles. Various errors may require customization in selecting the cycle numbers for including corresponding flow cell images in such cycles to train each prediction model or look-up table. The cycles that may be included for training an individual prediction model or look-up table may have comparable level of error(s). For example, certain library preparation error may decay exponentially with cycle increase. Thus, a separate prediction model or look-up table for cycle 1 alone may allow learning of the error more accurately without mixing additional flow cell images from other cycles. As another example, as the library error decays with cycle increase, the error’s effect on later cycles, e.g., cycles 8 and up, may be comparable, thus a single prediction model or look up table can be trained by including flow cell images from such cycles.

[0168] The techniques disclosed herein advantageously select a quality prediction model or a look-up table, among a plurality of pretrained quality prediction models or look-up tables, and each prediction model or look-up table is generated or populated by training using a different corresponding training dataset and different corresponding quality predictors based on thattraining dataset. Each corresponding training dataset may comprise flow cell images from one or more color channels and / or acquired at one or more different z-locations in one or more cycles of the sequencing run, e.g., cycle 1, cycle 2, cycle 2-10, etc. In some embodiments, at least one or more corresponding training dataset may only include flow cell images from a single cycle, e.g., cycle 1 or cycle 2. In some embodiments, only 1, 2, or 3 corresponding training datasets may only include flow cell images from a single cycle, e.g., cycle 1 or cycle 2. In some embodiments, multiple corresponding training datasets include flow cell images from multiple but not single sequencing cycles, e.g., cycles 2-10, cycles 8-16, etc. In some embodiments, each corresponding training dataset includes flow cell images from multiple but not single sequencing cycles, e.g., cycles 1-2, cycles 2-10, cycles 8-16, etc.

[0169] For look-up tables, each quality score in the look-up table may be generated based on quality scores calculated using two different methods per polony (e.g., signal spot in the flow cell image(s)), e.g., a standard quality score, and a conservative quality score, without the need of using complicated neural networks. The look-up table may be pre-trained, and the prediction of quality score using the look-up table may be conveniently and efficiently performed in parallel to the sequencing operations before any actual base calling has been made.

[0170] The techniques disclosed herein advantageously allow selection of different predictors for different cycles during a sequencing run for predicting quality of base calling in such cycles. The techniques disclosed herein also advantageously avoid contamination of quality prediction by errors within some but not all cycles during a sequencing run and provide more accurate prediction of quality score across multiple channels and / or within a single channel comparing with prediction of quality with a full look-up table or full prediction model generated using base calling from all sequencing cycles (e.g., cycles affected by errors differently).

[0171] In some embodiments, the techniques disclosed herein advantageously select a quality prediction model or a look-up table, among a plurality of quality prediction models or pretrained look-up tables, and each model or look-up table is generated or populated by training using a different corresponding training dataset and different corresponding quality predictors. Each corresponding training dataset may comprise flow cell images from one or more color channels and / or acquired at one or more different z-locations in one or more cycles of the sequencing run, e.g., cycle 1, cycles 2-10, etc.

[0172] FIG. 2B shows a flow chart illustrating a method 200 for predicting quality of base calling in various sequencing applications using a quality prediction model or look-up table selected from multiple prediction models or look-up tables.

[0173] The method 200 may include some or all of the operations disclosed herein. The operations may be performed in the order that is described herein, but is not limited to the order that has been described herein.

[0174] The method 200 may be performed by one or more processors disclosed herein. In some aspects, the processor may include one or more of: a processing unit, an integrated circuit, or their combinations. For example, the processing unit may include a central processing unit (CPU) and / or a graphic processing unit (GPU). The integrated circuit may include a chip such as a field-programmable gate array (FPGA) and / or an artificial intelligence (Al) chip. In some aspects, the processor may include the computing system 400 disclosed herein.

[0175] In some aspects, some or all operations in method 200 may be performed by the FPGAs and / or Al chips. In some aspects, when some operations are performed by FPGAs and / or Al chips, the data after an operation performed by the FPGA or Al chip may be communicated by the FPGAs or Al chip to the CPUs so that CPUs may perform subsequent operation(s) in method 200 using such data. In some aspects, all the operations in method 200 may be performed by CPUs. Alternatively, the operations performed by CPUs may be performed by other processors such as the dedicated processors, or GPUs.

[0176] In some aspects, the method 200 is performed during a cycle N after the flow cell images in cycle N have been acquired. In some aspects, the method 200 is performed in cycle N so that base calling of cycles prior to cycle N (e.g., cycle N-l) has already been performed, while sequencing reactions and subsequent base calling of cycle N (and similarly, cycle N+l, N+2) is yet to be performed. In some aspects, cycle N is the current cycle. While sequencing of the current cycle N is being performed, the flow cell images or information derived therefrom prior to cycle N may have been determined or saved to a memory or a data storage device disclosed herein. Such information may be used in one or more operations in method 200. Such information may be loaded from the memory or data storage device disclosed herein. N may be any integer that is greater than 2. For example, for a sequencing run, N may be any integer from 1 to 150, 1 to 200, or 1 to 400.

[0177] In some embodiments, the method 200 may include an operation of obtaining a plurality of pre-trained quality prediction models or look-up tables. Such pre-trained quality prediction models or look-up tables may be generated or populated using corresponding training datasets based on method 300 disclosed herein. The corresponding training dataset may be different for each different model or look-up table. For example, flow cell images of different samples from one or more channels and at different z levels of sequencing cycles 1-2 may beused to generate a first look-up table, while flow cell images of the same samples from same channels and at same z levels in sequencing cycles 3-15 may be used to generate a second lookup table. Same or different predictors may be selected to generate the different models or look-up tables.

[0178] In some embodiments, the corresponding training dataset only comprises training flow cell images from a subset of sequencing cycles of all sequencing cycles in a sequencing run, e.g., only cycles 2-10, or only cycles 16 and up. In some embodiments, at least one of the corresponding training datasets only comprises training flow cell images from a single sequencing cycle in a sequencing run (e.g., cycle 1 or cycle 2).

[0179] The operation of obtaining the plurality of models or look-up tables may be performed using a CPU, GPU. FPGA, Al chip or other processor or integrated circuit disclosed herein. The processor or integrated circuit may be part of the sequencing system or external to the sequencing system.

[0180] In some embodiments, the plurality of look-up tables comprises at least three look-up tables. In some embodiments, the plurality of look-up tables comprises the first look-up table, a second look-up table, and a third look-up table. In some embodiments, the first look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a first sequencing cycle, e.g., cycle 1. In some embodiments, the first lookup table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a first sequencing cycle and another sequencing cycle, e.g., cycles 1-2. The second look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a second sequencing cycle, e.g., cycle 2 or 3, and one more sequencing cycles subsequent thereto, e.g., cycles 3-5 or 4 -6. In some embodiments, the third look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a third sequencing cycle, e.g., cycle 18, and one more sequencing cycles subsequent thereto, e.g., cycles 22-250. In some embodiments, the training datasets includes flow cell images from all consecutive cycles that are typical in a sequencing run, e.g., cycles 1- 100 or cycles 1- 150. In some embodiments, the training dataset only include flow cell images from a subset of cycles among all the cycles that are typical in a sequencing run, e.g., cycles 1, 2-5, 8- 10, and 20-50.

[0181] The method 200 may include an operation 291 of selecting a look-up table from the plurality of quality prediction models or look-up tables, wherein each of the plurality of models or look-up tables is generated by training using a corresponding training data set. In someembodiments, the operation 291 is based on at least a current sequencing cycle of the sequence run being performed. In some embodiments, the operation may be based on a current sequencing cycle and a total number of sequencing cycles. In some embodiments, the operation 291 is based on the fact that the likelihood of at least one sequencing cycle being affected by sequencing errors is over a predetermined threshold. In some embodiments, the operation 291 is based on the fact that the likelihood of at least one sequencing cycle being affected by sequencing errors is under a predetermined threshold. In some embodiments, the selected model or look-up table may be a first look-up table when the likelihood of sequencing errors is over a predetermined threshold, and the selected model or look-up table may be a second model or look-up table different from the first one, when the likelihood of sequencing errors is under a predetermined threshold. The predetermined threshold may be customized based on various metrics. For example, for earlier cycles in a sequence run, e.g., cycles 1 -15, the likelihood of library preparation errors may be over the predetermined threshold, while later cycles, e.g., cycles 20- 350, the likelihood of library preparation errors may be under the predetermined threshold. As such, the operation 291 may selected different quality prediction models or look-up tables for cycles 1-15 from cycles 20-350.

[0182] In some embodiments, the at least one sequencing cycle number is less than 10, 15, 20, 25, or 30, and wherein the determined one or more base calls comprises at least an error caused by library preparation of the samples to be sequenced. In some embodiments, the at least one sequencing cycle number is greater than 10, 15, 20, 25, or 30, and wherein the one or more base calls lack an error caused by library preparation of the samples to be sequenced.

[0183] In some embodiments, the method further comprises an operation of determining, by the computer, a total number of sequencing cycles of the sequencing run. In some embodiments, the method further comprises an operation of determining, by the computer, a sequencing cycle number of the sequencing run. The sequencing cycle number may be of a current sequencing cycle. In some embodiments, the operation 291 of selecting the look-up table from the plurality of look-up tables is based on the at least one sequencing cycle number of the one or more flow cell images. In some embodiments, the look-up table from the plurality of look-up tables is based on the at least one sequencing cycle number of the one or more flow cell images and the total number of sequencing cycles. For example, for a sequencing error that is significant only in the first 10% of the sequencing cycles, the operation 291 of selecting the quality estimation model or look-up table may be based on the current sequencing cycle number and the total sequencing cycle number to determine which model or look-up table (e.g., considering such errors) to beselected. Other various information’s of the sequencing run besides the cycle number may be used alone or in combination with the cycle number as metrics for selecting the quality prediction model or look-up table in operation 291.

[0184] In some embodiments, the method 200 further includes an operation 292 of selecting a plurality of quality predictors based on the selected look-up table. Various quality predictors may be used. Some exemplary quality predictors include but are not limited to: phasing, prephasing, max intensity, clarity, low intensity media clarity, etc. Different quality prediction models or look-up tables may be generated based on different combination of quality predictors. For example, for earlier cycles in a sequence run, e.g., cycles 1 - 20, phasing and pre-phasing are negligible, thus look-up tables for such earlier cycles may be generated without selecting such quality predictors.

[0185] In some embodiments, the operation 292 of selecting, by the computer, the plurality of predictors based on the selected look-up table comprises: in response to determining that the selected look-up table is generated based on at least a sequencing cycle number less than a predetermined cycle number, including at least a predictor that is based on phasing or prephasing of the one or more flow cell images; and in response to determining that the selected look-up table is generated based on at least a sequencing cycle number greater than a predetermined cycle number, excluding any predictor of the plurality of predictors that is based on phasing or prephasing of the one or more flow cell images. For example, in response to determining that the selected look-up table is generated based on at least a sequencing cycle number (e.g., cycle 3) less than a predetermined cycle number (e.g., cycle 20), phasing and prephasing may be significant and may not be negligible for estimating quality of base calls, the operation 291 may include an operation of including at least a predictor that is based on phasing or prephasing of the one or more flow cell images. For earlier cycles in a sequence run, e.g., cycles 1 - 20, phasing and pre-phasing are not negligible, thus look-up tables for such earlier cycles may be generated with selecting such quality predictors.

[0186] In some embodiments, the method 200 further includes an operation 230 of determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced. The operation 230 is as what is disclosed with respect to method 200 in FIG. 2A.In some embodiments, the method 200 further includes an operation 240 of determining a first quality score from the selected look-up table based on the corresponding value for each of the plurality of predictors, wherein the selectedlook-up table comprises a plurality of dimensions corresponding to the plurality of predictors. The operation 240 is as what is disclosed with respect to method 200 in FIG. 2A.

[0187] In some embodiments, the method 200 further includes an operation of improving a baseline quality score by replacing the baseline quality score with the first quality score, wherein the baseline quality score is estimated based on a full quality prediction model or full look-up table generated by training using the corresponding datasets of the plurality of look-up tables. In some embodiments, the first quality score is estimated using a full quality estimation model or full look-up table generated based on training using the corresponding datasets of the plurality of look-up tables. In some embodiments, the full look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least the first (e.g., cycle 1), second (e.g., cycle 2), third (e.g., cycle 11) sequencing cycles and sequencing cycles subsequent thereto (e.g., cycles 16 and up). As an example, if a first corresponding dataset includes flow cell images from cycle 1, a second corresponding dataset includes flow cell images from cycles 2-7, a third corresponding dataset includes flow cell images from 8-20, and a fourth corresponding dataset includes flow cell images from cycles 21 -50. The full quality prediction model or full look-up table may be generated by training using all flow cell images from cycles 1 -50, including the first, second, third, and fourth corresponding datasets. The flow cell images in the corresponding datasets may be from one or more color channels and / or one or more z-levels. In some embodiments, the full quality prediction model or look up table may generate a prediction of quality score, e.g., the baseline quality score, which underestimates the quality of base calling for at least some polonies or clusters in some cycles and overestimates the quality of base calling for at least some polonies or clusters in same cycles and / or some other cycles. The operation of improving the baseline quality score by replacing the baseline quality score with the first quality score may include decreasing or increasing the quality prediction of base calling by using the first quality score but not the baseline quality score. The first quality score may be closer to the actual quality of base calling in the sequencing cycle than the baseline quality score.

[0188] In some embodiments, the method 200 may further comprise one or more operations to detect sequencing errors and / or library errors that may affect sequencing results and cause inaccurate and unreliable base calling using the quality prediction generated herein. As such, a sequencing run that has been impacted may be terminated before completion of all sequencing cycles to avoid waste of time, cost, reagent consumption, etc., on such sequencing runs. In some embodiments, the method 200 comprises an operation of determining one or more base calls based on the one or more flow cell images. Such operation occurs only after determining aquality score using operations in methods 200 and the quality score is above a predetermined quality threshold. The predetermined quality threshold may be customized to include various numbers or ranges. For example, the predetermined quality threshold may be in a range from Q25 to Q60. For example, the predetermined quality threshold may Q 18, Q20, Q22, Q25, or Q30.

[0189] In some embodiments, in response to determining that the first quality score is below a quality threshold, the method 200 may include an operation of terminating the sequencing run before completing all sequencing cycles of the sequencing run. In some embodiments, the first quality score may be different from the baseline quality score, e.g., lower, but may be a more accurate prediction than the baseline quality score. In some embodiments, the baseline quality score may be above the quality threshold and providing an overestimation of quality of base calling. However, it is advantageous to determine whether the sequencing run may be terminated or not based on the first quality score but not the baseline quality score.

[0190] In some embodiments, the first quality score or the baseline quality score is determined with respect to a pixel or a subpixel of the one or more flow cell images. In some embodiments, the first quality score or the baseline quality score is determined with respect to a polony or a cluster of the one or more flow cell images, the polony or cluster comprising one or more pixels or subpixels.

[0191] In some embodiments, the method 200 further comprises one or more operations to generate a full look-up table, e.g., using methods 300 disclosed herein. In some embodiments, the method further comprises one or more operations to predict quality of base calling thereby generating the baseline quality score using the full look-up table. Such one or more operations including: selecting, by the computer, the full look-up table; selecting, by the computer, a plurality of predictors based on the full look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a baseline quality score from the full look-up table based on the corresponding value for each of the plurality of predictors, wherein the full look-up table comprises a plurality of dimensions corresponding to the plurality of predictors. Such one or more operations are similar to the operations in methods 200 except that the full look-up table and its corresponding predictors are used instead of the selected look-up table and corresponding predictors thereof.

[0192] In some embodiments, the baseline quality score determined using the full look-up table is greater than Q35, Q40, Q45, Q50, or more. The baseline quality score may underestimatethe quality of base calling in some cycles, e.g., cycles subsequent to cycle 16, and may overestimate the quality of base calling in some other cycles, e.g., cycles 1- 12. Such inaccuracy may be caused by library preparation errors or sequencing errors. In some embodiments, the method 200 herein improves such inaccuracy by estimating the first quality score using the selected look-up table or quality prediction model and its corresponding predictors. In some embodiments, the first quality score is more accurate in estimating the actual quality of base calling than the baseline quality score. In some embodiments, the at least one sequencing cycle number is less than 5, 10, 12, 15, or 20, and the first quality score is lower than the baseline quality score. In some embodiments, the at least one sequencing cycle is greater than 5, 10, 12, 15, or 20, and the first quality score is higher than the baseline quality score. In some embodiments, the first quality score is greater than Q35, Q40, Q45, Q50, Q55, Q60, or more.

[0193] Fig. 25A shows exemplary quality scores of a sequencing run in which the quality scores have a correlation with the sequencing cycle number and decreases gradually as the cycle number increases in the sequencing run. In some embodiments, the quality scores predicted using methods 200 based on the full look-up table or the selected look-up table may have a similar trend of decreasing as the sequencing cycle number increases. FIG. 25B shows exemplary quality score estimated using the full look up table when the samples have library preparation error(s), e.g., caused by gap filler enzymes and correlation with actual quality score in different cycles. The quality score in earlier cycles, e.g., cycle 1, predicted using the full look-up table is inaccurate as the prediction overestimates the quality of base calling in the earlier cycles. Such overestimation decreases later cycles, e.g., cycles 21, 41, 61, and 76. Such overestimation may be improved by predicting the quality score using the method 200 in FIG. 2B with selected look-up table. Similarly, underestimation or overestimation in later cycles of the sequencing run by using the full look-up table may be improved be improved by predicting the quality score using the method 200 in FIG. 2B with selected look-up table.

[0194] In some embodiments, the method 200 further comprise an operation of determining one or more base calls based on the one or more flow cell images as disclosed herein with respect to method 200 in FIG. 2A. In some embodiments, the method 200 further comprise one or more operations of pre-processing the flow cell images before determining one or more base calls as disclosed herein with respect to method 200 in FIG. 2A.Methods for generating multiple quality prediction models

[0195] In some embodiments, disclosed herein are methods 300 for generating the plurality of quality prediction models or look-up tables herein that may be selected for predicting quality of base calling in different sequencing runs or in different cycles. The methods for generating such plurality of quality prediction models or look-up tables are generally similar to those for generating the full prediction model or look-up table. The difference may exist in different training data and / or usage of different predictors.

[0196] In some embodiments, the computer-implemented method 300 for predicting quality of base calling in DNA sequencing comprises an operation of generating a first look-up table or quality prediction model for predicting quality of base calling in DNA sequencing. The operation of generating the first look-up table or quality prediction model may comprise one or more operations including: obtaining, by a computer, first training base calls from one or more first training flow cell images; selecting, by the computer, a first plurality of predictors based on at least a first sequencing cycle number of the one or more first training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the first training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of the first plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a first look-up table having a number of dimensions determined by the first plurality of predictors and each dimension comprising the corresponding number of bins; determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the first look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps: (a) calculating a standard quality score and a conservative quality score for each bin of the first look-up table; (b) determining coordinates of a bin with a maximum conservative quality score in the first look-up table; (c) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the first look-up table; and (d) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

[0197] In some embodiments, the method 300 further comprises an operation of generating a second look-up table for predicting quality of base calling in DNA sequencing. The operation of generating a second look-up table comprises: obtaining, by a computer, second training base callsfrom one or more second training flow cell images; selecting, by the computer, a second plurality of predictors based on at least a second sequencing cycle number of the one or more second training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the second training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of a second plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a second look-up table having a number of dimensions determined by the second plurality of predictors and each dimension comprising the corresponding number of bins; determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the second look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps: (a) calculating a standard quality score and a conservative quality score for each bin of the second look-up table; (b) determining coordinates of a bin with a maximum conservative quality score in the second look-up table; (c) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the second look-up table; and (d) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

[0198] In some embodiments, the second sequencing cycle number is greater than the first sequencing cycle number. In some embodiments, the second sequencing cycle number is greater than 10, 12, 15, 20, 25, or 30. In some embodiments, the first sequencing cycle number is less than 5, 10, 12, 15, or 20.

[0199] In some embodiments, the second plurality of predictors is different from the first plurality of predictors. In some embodiments, the second plurality of predictors comprises more predictors than the first plurality of predictors. In some embodiments, the second plurality of predictors has a different number of predictors from the first plurality of predictors. For example, the first plurality of predictors may include 3 predictors while the second plurality may include 4 predictors. In some embodiments, the second plurality of predictors comprises at least one predictor based on phasing or pre-phasing. In some embodiments, the second plurality of predictors lacks any predictor based on phasing or pre-phasing.

[0200] In some embodiments, the method 300 further comprise an operation of acquiring, by an optical system, the one or more first training flow cell images in at least the first sequencing cycle number. In some embodiments, the method 300 further comprises: acquiring, by an opticalsystem, the one or more first training flow cell images in only the first sequencing cycle number. For example, the first sequencing cycle may be cycle 1, and the first training flow cell images in the training dataset may only include flow cell images from cycle 1 of one or more sequencing runs of various samples. As another example, the first sequencing cycle may be cycle 2, and the first training flow cell images in the training dataset may include flow cell images from cycle 2 and its subsequent cycles, e.g., cycles 3-9, from one or more sequencing runs of various samples.

[0201] In some embodiments, the training flow cell images are prepared using various libraries. In some embodiments, the training flow cell images are prepared using a PCR-free, Covaris sheared library with well-known genome like human or E.coli.

[0202] In some embodiments, the method 300 further comprises: acquiring, by an optical system, the one or more second training flow cell images in the second sequencing cycle number and a plurality of cycles subsequent to the second sequencing cycle number. As an example, the second sequencing cycle may be cycle 16, and the second training flow cell images in the training dataset may include flow cell images from cycle 16 and its subsequent cycles, e.g., cycles 16-120, from one or more sequencing runs of various samples.

[0203] In some embodiments, the method further comprises an operation of generating a third quality prediction model or look-up table for predicting quality of base calling in DNA sequencing using similar operations as for generating the first or second quality prediction model or look-up table but with different training flow cell images. In some embodiments, the one or more first training flow cell images comprise errors caused by library preparation. In some embodiments, the one or more second training flow cell images comprise errors caused by library preparation. As a nonlimiting example, the corresponding training dataset for the first quality prediction model or look-up table may include first flow cell images only from cycle 1. The corresponding training dataset for the second quality prediction model or look-up table may include second flow cell images from cycle 2-9. The corresponding training dataset for the third quality prediction model or look-up table may include third flow cell images from cycle 10-18. The corresponding training dataset for the fourth quality prediction model or look-up table may include fourth flow cell images from cycle 19-300.

[0204] In some embodiments, quality prediction models or look-up tables of cycles separated into intervals (e.g., cycles 1, cycles 2-8, cycles 9-20, etc.) can cause stepping up or down of quality prediction in cycles at the end of each intervals (e.g., between cycles 1 and 2, between cycles 8 and 9, between cycles 20 and 21, etc.). In some embodiments, the methods 200 herein include an operation of smoothing the stepping in quality estimation. In some embodiments, theoperation of smoothing the stepping in quality estimation includes: predicting the quality scores with the corresponding prediction model or look-up table of a first boundary cycle; predicting the quality scores with the corresponding prediction model or look-up table of a second boundary cycle; and determining the predicted quality scores of the first and second boundary cycles as the average of the predictions. For example, cycles 15 and 16 are boundary cycles of two different cycle intervals. The predicted quality score may be an average of the quality score of the same polony at cycles 15 and 16.

[0205] In some embodiments, the quality prediction model or look-up table is trained on training datasets of same cycles as the cycles that the pretrained prediction model or look-up table may predict quality on. For example, a look-up table generated using training dataset including only flow cell images from cycles 2-8 in sequencing runs may be used only to predict quality score of a sequencing run in cycles 2-8 but not other cycles.

[0206] In some embodiments, the quality prediction model or look-up table is trained on training datasets of different cycles as the cycles that the pretrained prediction model or look-up table may predict quality on. In some embodiments, the cycles used for training the model or look-up table and the cycles for prediction may have at least one identical cycles. In some embodiments, the cycles used for training the model or look-up table and the cycles for prediction may have at least one identical cycle. In some embodiments, the number of cycles for training a prediction model or look-up table may be greater than the number of cycles that the pretrained model or look-up table may be used for predicting base calling quality. For examples, a pretrained prediction model using training dataset only including flow cell images in cycles 2- 10 may be used to predict only cycles 2-5, and cycles 2-5 are the overlapped identical cycles. As another example, a pretrained prediction model using training dataset only including flow cell images in cycles 6-15 may be used to predict only cycles 6-10.Samples

[0207] In some embodiments, the prediction model or look-up table and the methods herein are suitable for predicting quality of base calling in various samples being sequenced.

[0208] Disclosed herein are 2D and / or 3D samples. Disclosed herein are sample(s) comprising one or more cell, tissue, and / or organoids. The sample(s) may comprise one or more cells wherein individual cells of the sample(s) comprise a plurality of target polynucleotides and / or a plurality of target polypeptides on the one or more cells or at spatial locations inside theone or more cells. In some embodiments, the target polynucleotides and / or target polypeptides are located on the cellular membrane. In some embodiments, the target polynucleotides and / or target polypeptides are located inside individual cells, for example at a subcellular location including without limits a nucleus, nucleolus, mitochondria, Golgi apparatus, ribosome, endoplasmic reticulum, microtubule, centriole, spindle, actin filament, flagellum, cilium, peroxisome, lysosome, cytoplasm or any combinations thereof.

[0209] In some embodiments, the plurality of target polynucleotides comprises a plurality of target DNA molecules. In some embodiments, the plurality of target DNA molecules comprise at least one target nucleic acid sequence. In some embodiments, the at least one target nucleic acid sequence can be determined by conducting nucleic acid sequencing inside the sample(s).

[0210] In some embodiments, the plurality of target DNA molecules comprise without limitation genomic DNA, mitochondrial DNA, and non-integrated DNA introduced as a vector into the one or more cells. In some embodiments, the plurality of target DNA molecules comprises one or more DNA sequences stably integrated into the genome of the sample(s) via gene editing using CRISPR Cas9 and guide RNA.

[0211] In some embodiments, the plurality of target polynucleotides comprises a plurality of target RNA molecules. In some embodiments, the plurality of target RNA molecules comprise at least one target nucleic acid sequence. In some embodiments, the at least one target sequence of the target RNA molecule or target cDNA molecule can be determined by conducting nucleic acid sequencing inside the sample(s).

[0212] In some embodiments, the plurality of target RNA molecules comprises any type of RNA transcribed from genomic DNA of the sample(s), including without limitation messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), micro RNA (miRNA), IncRNA (long non coding RNA), lincRNA (long intergenic non coding RNA) and / or siRNA (small interfering RNA).

[0213] In some embodiments, the plurality of target RNA molecules comprises RNA transcribed from the non-integrated DNA vectors or the stably-integrated DNA sequence introduced via CRISPR gene editing.

[0214] In some embodiments, the one or more cells comprise a plurality of target polypeptides. In some embodiments, the plurality of target polypeptides comprise at least one target amino acid sequence. In some embodiments, the at least one target amino acid sequence can be determined by conducting barcode sequencing inside the sample(s) where the barcode sequence corresponds to the target amino acid sequence.

[0215] In some embodiments, the target polypeptides comprise without limitation phosphorylated polypeptides, non-phosphorylated polypeptides, enzymes, docking proteins, structural proteins, transport proteins, hormonal proteins, contractile proteins, storage proteins and antibodies. Exemplary enzymes include oxidoreductases, transferases, hydrolases, lyases, isomerases, and ligases.

[0216] In some embodiments, the sample(s) is fixed and permeabilized. In some embodiments, the sample(s) is fixed but is not permeabilized. In some embodiments, the sample(s) is not fixed or permeabilized.

[0217] In some embodiments, the sample(s) comprises without limitation a single cell, multiple cells, a cell suspension, a subculture, adherent cells, a tissue, cells from a blood sample, tumor cells from a patient sample, or portions thereof.

[0218] In some embodiments, the sample(s) comprises one or more cells from any organ including but not limited to brain, breast, ovary, cervix, colon, rectum, endometrium, gallbladder, intestines, bladder, prostate, testicles, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary, thymus, skin, heart, larynx, or other organs. In some embodiments, the sample(s) comprises fibroblasts, stem cells, neurons, or osteosarcoma cells.

[0219] In some embodiments, the sample(s) comprises any mammalian immune cells including B cells, T cells and macrophage. In some embodiments, the B cells comprise B- lymphocytes, including pre-B cells, B cells or pro-B cells. In some embodiments, the T cells comprise T-lymphocytes including primary T cells.

[0220] In some embodiments, the sample(s) comprises any mammalian cells including and without limitation Chinese Hamster Ovary (CHO) cells, baby hamster kidney cells, NSO myeloma cells, monkey kidney COS cells, monkey kidney fibroblast CV-I cells, dog kidney MDCK cells, human embryonic kidney 293 (HEK293) cells, human breast cancer SKBR3 cells, human leukemia Jurkat T cells, human cervical cancer HeLa cells, human neuroblastoma SH- SY5Y cells and immortalized human neural progenitor cells ReNcell VM and ReNcell CX.

[0221] In some embodiments, the CHO cells include without limitation DHFR CHO cells, CHO-kl cells, CHO-S cells, GS-CHO cells, CHO-K1 cells, CHO-DG44 cells, CHO-DUXB 11 cells, CHO-DUKX cells, and CHOK1 SV cells.

[0222] In some embodiments, the sample(s) comprises any mammalian cells including but not limited to BHK cells, MDCK cells, C3H 10T1 / 2 cells, FLY cells, Flp-cells, Psi-2 cells, BOSC 23 cells, P A317 cells, WEHI cells, COS cells, BSC 1 cells, BSC 40 cells, BMT 10 cells, Jurkat cells, VERO cells, MDCK cells, WI38 cells, V79 cells, B14 AF28-G3 cells, BHK cells,HaK cells, NSO cells, SP2 / 0-Agl4 cells, HEK cells, HEK293 cells (e.g., HEK 293-F, HEK293- H, HEK293-T), MRC5 cells, A549 cells, HUVEC cells, HCT116 cells, PC3 cells, MCF7 cells, U2OS cells, HTI080 cells, Hep cells, iHPC cells, 293 cells, 293T cells, B-50 cells, 3T3 cells, NIH3T3 cells, NK cells, HepG2 cells, Saos-2 cells, Huh7 cells, W163 cells, 211 cells or 211 A cells.

[0223] In some embodiments, the sample(s) comprises healthy cells, diseased cells or a mixture of healthy and diseased cells. In some embodiments, the diseased cells comprise any type of cancer cell including without limitation carcinomas, sarcomas, melanomas, leukemia and lymphomas.

[0224] In some embodiments, the sample(s) comprises diseased cells comprising cells infected with any type of virus including without limitation Epstein-Barr virus (EBV), hepatitis B (HBV), hepatitis C (HCV), human papillomavirus (HPV), human immunodeficiency virus (HIV), Kaposi sarcoma-associated herpesvirus (KSHV), human herpesvirus 8 (e.g., HHV-8 or KSHV), human T-cell leukemia virus type 1 (HTLV-1), Merkel cell polyomavirus (MCV), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), Middle East respiratory syndrome coronavirus (MERS-CoV) and respiratory syncytial virus (RSV).

[0225] In some embodiments, the sample(s) comprises diseased cells comprising cells infected with any type of bacteria including without limitation Streptococcus pneumoniae, Haemophilus influenzae, Shigella, Listeria, Salmonella, E. coli, Staphylococcus aureus, Streptococcus pyogenes, Helicobacter pylori, Neisseria gonorrhoeae, Neisseria meningitidis, Streptococcus pneumoniae, Treponema pallidum, Bordetella pertussis, Mycobacterium tuberculosis, Methicillin-resistant Staphylococcus aureus, Borrelia burgdorferi, Bacillus anthracis, Clostridium tetani, Clostridium difficile, Yersinia pestis, Vibrio cholerae, Legionella pneumophila, Coxiella burnetiid, and Burkholderia pseudomallei.

[0226] In some embodiments, the sample(s) comprises organoid(s) comprising a plurality of cells that are cultured under a condition suitable to promote self-organization into a three- dimensional structure that mimics characteristics of an organ including cell organization, cell function, organ structure and organ biology. The organoid comprises a miniaturized organ-like structure. The organoid can originate from one cell. The organoid can originate from a plurality of cells from a tissue, embryonic stem cells, adult stem cells, adult progenitor cells or induced pluripotent stem cells (iPSCs). Under suitable culture conditions, the originating cells undergo differentiation, cell sorting and spatially restricted lineage commitment to generate two or more organ-specific cell types that form the organoid. The cells of an organoid can exhibit cell-to-cellcommuni cation. The organoid comprises different organ-specific cells that are grouped together and spatially organized.

[0227] In some embodiments, the sample(s) comprises any type of organoid including but not limited to a cerebral organoid, intestinal organoid, gastric organoid, lingual organoid, tooth organoid, thyroid organoid, thymic organoid, testicular organoid, prostate organoid, hepatic organoid, pancreatic organoid, lung organoid, kidney organoid, endometrial organoid, cardiac organoid, retinal organoid and skin organoid.

[0228] In some embodiments, the sample(s) comprises a tumor organoid which can be generated by culturing patient primary tumor cells or metastatic tumor cells in a manner similar to the culture conditions used to generate non-tumor organoids. Cells used to generate tumor organoids also undergo self-organization into a three-dimensional structure that mimics characteristics of an organ. The tumor organoid can retain the genetic mutations present in the original tumor cells obtained from the cancer patient.

[0229] In some embodiments, the sample(s) comprises a tumor organoid comprising a patient-derived xenograft comprising patient tumor tissue that is implanted into an immunocompromi sed mouse.

[0230] In some embodiments, the sample(s) comprises a monolayer cell culture or a primary cell three-dimensional culture both of which can originate from one cell, or from a plurality of cells from a tissue, embryonic stem cells, adult stem cells, adult progenitor cells or induced pluripotent stem cells (iPSCs). Monolayer cell cultures grow in a two-dimensional flat single layer and therefore lack three-dimensional structure and lack cells that self-organize into structures of an organoid. Primary cell three-dimensional cultures can be grown in scaffold-based or scaffold-free systems. Primary cell three-dimensional cultures do not form organoids but instead form multicellular spheroid structures via simple cell-cell adhesion when physical or mechanical force is applied during culturing.

[0231] In some embodiments, any of the sample(s)s described herein can be deposited (e.g., seeded) onto a support comprising a planar or non-planar support. In some embodiments, the support comprises a solid or semi-solid support. In some embodiments, the support comprises a porous, semi-porous or non-porous support. The support can be made of any material such as glass, plastic or a polymer material. In some embodiments, the surface of the support can be coated with one or more compounds to produce a passivated layer on the support. In some embodiments, the passivated layer forms a porous or semi-porous layer.

[0232] In some embodiments, the sample(s) can be deposited (e.g., seeded) onto a support comprising a planar surface and having walls to contain the sample(s) and liquids, such as for example, liquid cell culture medium and reagents for sequencing. The support can be configured to include walls that form at least two wells. In some embodiments, individual wells can be loaded with different types of sample(s)s.

[0233] In some embodiments, the sample(s) comprises a plurality of genetically perturbed cells wherein individual genetically perturbed cells comprise at least one perturbed genomic target region that can undergo cellular transcription to generate a plurality of perturbed target RNA molecules. In some embodiments, individual perturbed target RNA molecules comprise (i) a target-specific spacer sequence, (ii) a universal sequencing primer binding sequence, (iii) a universal scaffold sequence and (iv) a reverse transcription primer binding sequence. In some embodiments, the universal sequencing primer binding sequence is proximal or adjacent to the target-specific spacer sequence.

[0234] In some embodiments, the sequence of at least a portion of the perturbed target RNA molecules can be determined by conducting nucleic acid sequencing inside the sample(s). In some embodiments, the target-specific spacer sequence can be determined by conducting nucleic acid sequencing inside the sample(s).

[0235] In some embodiments, individual genetically perturbed cells can be generated by conducting targeted genome editing employing a complex comprising CRISPR Cas9 and guide RNA. The guide RNA comprises a target-specific spacer sequence and a universal scaffold sequence that binds Cas9 protein. The target-specific spacer sequence mediates double-stranded cleavage at the genomic target region and cellular DNA repair to generate a perturbed genomic target region comprising an inserted guide DNA sequence and scaffold sequence.

[0236] The term “CRISPR” refers to Clustered Regularly Interspaced Short Palindromic Repeats. In some embodiments, the target-specific spacer sequence is 20 nucleotides in length and comprises a sequence that is complementary to the genomic target region. In some embodiments, the target-specific spacer sequence and scaffold sequence can be inserted into a particular genomic target region thereby generating the perturbed genomic target region. In some embodiments, the scaffold sequence comprises a sequence for binding a reverse transcription primer and / or a sequence for binding a sequencing primer.

[0237] In some embodiments, individual genetically perturbed cells can transcribe the perturbed genomic target region(s) and generate RNA molecules comprising a sequence that is complementary to at least a portion of the perturbed genomic target region(s) including thetarget-specific spacer and scaffold sequences. In some embodiments, the transcribed RNA molecules comprise the perturbed target RNA molecules. In some embodiments, the transcribed RNA molecules comprise the target polynucleotides.

[0238] In a perturbed target RNA molecule, the target-specific spacer sequence serves as a surrogate identification sequence that can identify the perturbed genomic target region of a given genetically perturbed cell.

[0239] In some embodiments, the genetically perturbed cell can exhibit a morphological change resulting from the genetic perturbation compared to a non-perturbed cell that lacks a CRISPR guide RNA inserted into its genomic target region.

[0240] In some embodiments, the morphological change includes but is not limited to a change in cell size, a change in cell shape, a change in nuclear size, a change in the cellular location of a protein-of-interest, a change in intracellular protein-protein interaction, a change in the presence or absence of a cell surface protein, a change in the structure or arrangement of organelles including mitochondria and / or a change in cell motility.

[0241] In some embodiments, the morphological change includes a change in cell resistance or a change in cell sensitivity to a challenge condition. In some embodiments, the challenge condition includes but is not limited to a temperature change, a pH change, light exposure, a dark condition, nutrient deprivation, nutrient addition, toxin exposure, chemical compound exposure and / or drug exposure. In some embodiments, the genetic perturbation causes no morphological change of the genetically perturbed cell.

[0242] In some embodiments, the genetically perturbed cell can exhibit a change in gene expression resulting from the genetic perturbation compared to a non-perturbed cell that lacks a CRISPR guide RNA inserted into its genomic target region. In some embodiments, the change in gene expression includes but is not limited to increase or decrease in the levels RNA-of-interest, or increase or decrease in the levels of proteins-of-interest.

[0243] In some embodiments, the change in gene expression can increase or decrease protein phosphorylation by altering gene expression of kinase and / or phosphatase enzymes that append or remove phosphate groups from proteins. The activity of certain proteins can increase or decrease with a change in phosphorylation state.

[0244] In some embodiments, a genetically perturbed cell does not exhibit a change in gene expression until the genetically perturbed cell is subjected to a challenge condition. In some embodiments, the challenge condition includes but is not limited to a temperature change, a pHchange, light exposure, a dark condition, nutrient deprivation, nutrient addition, toxin exposure, chemical compound exposure and / or drug exposure.

[0245] In some embodiments, the genetic perturbation causes no change in gene expression of the target RNA or target protein in the genetically perturbed cell.

[0246] In some embodiments, the sample(s) comprises a plurality of genetically perturbed cells wherein individual genetically perturbed cells comprise the same target-specific spacer sequence inserted at the same genomic target region. In some embodiments, individual genetically perturbed cells in the plurality comprise perturbed target RNA molecules having the same target-specific spacer sequence.

[0247] In some embodiments, the sample(s) comprises a mixture of two or more different genetically perturbed cells wherein at least two of the genetically perturbed cells comprise different target-specific spacer sequences inserted at different genomic target regions. In some embodiments, at least two of the genetically perturbed cells comprise perturbed target RNA molecules having different target-specific spacer sequences.

[0248] In some embodiments, any of the cells described herein can be used to generate the plurality of genetically perturbed cells.

[0249] In some embodiments, cells from any of the organs described herein can be used to generate the plurality of genetically perturbed cells.

[0250] In some embodiments, any of the healthy or diseased cells described herein can be used to generate the plurality of genetically perturbed cells.

[0251] In some embodiments, any of the cells described herein can be used to generate a plurality of genetically perturbed cells which can be cultured under conditions suitable for forming an organoid which harbors the genetic perturbation.

[0252] In some embodiments, any of the patient primary tumor cells or metastatic tumor cells described herein can be used to generate a plurality of genetically perturbed tumor cells which can be cultured under conditions suitable for forming a tumor organoid which harbors the genetic perturbation.Computer systems

[0253] Various aspects of the method 200 and 300 may be implemented, for example, using one or more computer systems, such as computer system 400 shown in FIG. 4. One or morecomputer systems 400 may be used, for example, to implement any of the aspects discussed herein, as well as combinations and sub-combinations thereof.

[0254] Computer system 400 may include one or more hardware processors 404. The hardware processor 404 may be central processing unit (CPU), graphic processing units (GPU), or their combination. Processor 404 may be connected to a bus or communication infrastructure 406.

[0255] Computer system 400 may also include user input / output device(s) 403, such as monitors, keyboards, pointing devices, etc., which may communicate with communication infrastructure 406 through user input / output interface(s) 402. The user input / output devices 403 may be coupled to the user interface 124 in FIG. 1.

[0256] One or more units of processors 404 may be a graphics processing unit (GPU). In an aspect, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. The GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common to computer graphics applications, images, videos, vector processing, array processing, etc., as well as cryptography (including brute-force cracking), generating cryptographic hashes or hash sequences, solving partial hash-inversion problems, and / or producing results of other proof-of- work computations for some blockchain-based applications, for example. With capabilities of general-purpose computing on graphics processing units (GPGPU), the GPU may be particularly useful in at least the image recognition and machine learning aspects described herein.

[0257] Additionally, one or more of processors 404 may include a coprocessor or other implementation of logic for accelerating cryptographic calculations or other specialized mathematical functions, including hardware-accelerated cryptographic coprocessors. Such accelerated processors may further include instruction set(s) for acceleration using coprocessors and / or other logic to facilitate such acceleration.

[0258] Computer system 400 may also include a data storage device such as a main or primary memory 408, e.g., random access memory (RAM). Main memory 408 may include one or more levels of cache. Main memory 408 may have stored therein control logic (i.e., computer software) and / or data.

[0259] Computer system 400 may also include one or more secondary data storage devices or secondary memory 410. Secondary memory 410 may include, for example, a main storage drive 412 and / or a removable storage device or drive 414. Main storage drive 412 may be a hard disk drive or solid-state drive, for example. Removable storage drive 414 may be a floppy disk drive,a magnetic tape drive, a compact disk drive, an optical storage device, tape backup device, and / or any other storage device / drive.

[0260] Removable storage drive 414 may interact with a removable storage unit 418.

[0261] Removable storage unit 418 may include a computer usable or readable storage device having stored thereon computer software and / or data. The software may include control logic. The software may include instructions executable by the hardware processor(s) 404. Removable storage unit 418 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / any other computer data storage device. Removable storage drive 414 may read from and / or write to removable storage unit 418.

[0262] Secondary memory 410 may include other means, devices, components, instrumentalities or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 400. Such means, devices, components, instrumentalities or other approaches may include, for example, a removable storage unit 422 and an interface 420. Examples of the removable storage unit 422 and the interface 420 may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.

[0263] Computer system 400 may further include a communication or network interface 424. Communication interface 424 may enable computer system 400 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referenced by reference number 428). For example, communication interface 424 may allow computer system 400 to communicate with external or remote devices 428 over communication path 426, which may be wired and / or wireless (or a combination thereof), and which may include any combination of LANs, WANs, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 400 via communication path 426. In some aspects, communication path 426 is the connection to the cloud 130, as depicted in FIG. 1. The external devices, etc. referred to by reference number 428 may be devices, networks, entities, etc. in the cloud 130.

[0264] Computer system 400 may also be any of a personal digital assistant (PDA), desktop workstation, laptop or notebook computer, netbook, tablet, smart phone, smart watch or other wearable, appliance, part of the Internet of Things (loT), and / or embedded system, to name a few non-limiting examples, or any combination thereof.

[0265] It should be appreciated that the framework described herein may be implemented as a method, process, apparatus, system, or article of manufacture such as a non-transitory computer-readable medium or device. For illustration purposes, the present framework may be described in the context of distributed ledgers being publicly available, or at least available to untrusted third parties. One example as a modem use case is with blockchain-based systems. It should be appreciated, however, that the present framework may also be applied in other settings where sensitive or confidential information may need to pass by or through hands of untrusted third parties, and that this technology is in no way limited to distributed ledgers or blockchain uses.

[0266] Computer system 400 may be a client or server, accessing or hosting any applications and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions; local or on-premises software (e.g., “on-premise” cloud-based solutions); “as a service” models (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (laaS), database as a service (DBaaS), etc.); and / or a hybrid model including any combination of the foregoing examples or other services or delivery paradigms.

[0267] Any applicable data structures, file formats, and schemas may be derived from standards including but not limited to JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations alone or in combination. Alternatively, proprietary data structures, formats or schemas may be used, either exclusively or in combination with known or open standards.

[0268] Any pertinent data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in human-readable formats such as numeric, textual, graphic, or multimedia formats, further including various types of markup language, among other possible formats. Alternatively or in combination with the above formats, the data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted formats, or any other machine-readable formats.

[0269] Interfacing or interconnection among various systems and layers may employ any number of mechanisms, such as any number of protocols, programmatic frameworks, floorplans,or application programming interfaces (API), including but not limited to Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Pattern (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG) HTML5 Web Messaging, Representational State Transfer (REST or RESTful web services), Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other mechanisms, open or proprietary, that may achieve similar functionality and results.

[0270] Such interfacing or interconnection may also make use of uniform resource identifiers (URI), which may further include uniform resource locators (URL) or uniform resource names (URN). Other forms of uniform and / or unique identifiers, locators, or names may be used, either exclusively or in combination with forms such as those set forth above.

[0271] Any of the above protocols or APIs may interface with or be implemented in any programming language, procedural, functional, or object-oriented, and may be compiled or interpreted. Non-limiting examples include C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or virtually any other language, with any other libraries or schemas, in any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, including but not limited to Node.js, V8, Knockout, j Query, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, and so on, among many other non-limiting examples.

[0272] In some aspects, a tangible, non-transitory apparatus or article of manufacture comprising a tangible, non-transitory computer useable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 400, main memory 408, secondary memory 410, and removable storage units 418 and 422, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 400), may cause such data processing devices to operate as described herein.

[0273] Based on the teachings contained in this disclosure, it will be apparent to persons skilled in the relevant art(s) how to make and use aspects of this disclosure using data processing devices, computer systems and / or computer architectures other than that shown in FIG. 4. In particular, aspects may operate with software, hardware, and / or operating system implementations other than those described herein.

[0274] The imager 116 in FIG. 1 may include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system designs provide for larger fields-of-view, increased spatial resolution, improved modulation transfer, contrast-to- noise ratio, and image quality, higher spatial sampling frequency, faster transitions between image capture when repositioning the sample plane to capture a series of images (e.g., of different fields-of-view), and improved imaging system duty cycle, and thus enable higher throughput image acquisition and analysis.

[0275] In some instances, improvements in imaging performance, e.g., for dual-side (flow cell) imaging applications, may be achieved by using an electro-optical phase plate in combination with an objective lens to compensate for the optical aberrations induced by the layer of fluid separating the upper (near) and lower (far) interior surfaces of a flow cell. In some instances, this design approach may also compensate for vibrations introduced by, e.g., a motion- actuated compensator that is moved in or out of the optical path depending on which surface of the flow cell is being images.

[0276] In some instances, improvements in imaging performance, e.g., for dual-side (flow cell) imaging applications comprising the use of thick flow cell walls (e.g., wall (or coverslip) thickness > 700 pm) and fluid channels (e.g., fluid channel height or thickness of 50 - 200 pm) may be achieved even when using commercially-available, off-the-shelf objectives by using a tube lens design that corrects for the optical aberrations induced by the thick flow cell walls and / or intervening fluid layer in combination with the objective.

[0277] In some instances, improvements in imaging performance, e.g., for multichannel (e.g., two-color or four-color) imaging applications, may be achieved by using multiple tube lenses, one for each imaging channel, where each tube lens design has been optimized for the specific wavelength range used in that imaging channel.

[0278] Exemplary aspects disclosed herein may comprise fluorescence imaging systems, said systems comprising: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; b) an objective lens configured to collect fluorescence arising from within a specified field-of-view of a sample plane upon exposure of the sample plane to the excitation light, wherein a numerical aperture of the objective lens is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9 or a numerical aperture value falling within a range defined by any two of the foregoing; whereina working distance of the objective lens is at least 400 gm, at least 500 gm, at least 600 gm, at least 700 gm, at least 800 gm, at least 900 gm, at least 1000 gm, or a working distance falling within a range defined by any two of the foregoing; and wherein the field-of-view has an area of at least 0.1 mm2, at least 0.2 mm2, at least 0.5 mm2, at least 0.7 mm2, at least 1 mm2, at least 2 mm2, at least 3 mm2, at least 5 mm2, or at least 10 mm2, or a field of view falling within a range defined by any two of the foregoing; and c) at least one image sensor, wherein the fluorescence collected by the objective lens is imaged onto the image sensor, and wherein a pixel dimension for the image sensor is chosen such that a spatial sampling frequency for the fluorescence imaging system is at least twice an optical resolution of the fluorescence imaging system.

[0279] In some aspects, the numerical aperture may be at least 0.75. In some aspects, the numerical aperture is at least 1.0. In some aspects, the working distance is at least 850 pm. In some aspects, the working distance is at least 1,000 pm. In some aspects, the field-of-view may have an area of at least 2.5 mm2. In some aspects, the field-of-view may have an area of at least 3 mm2. In some aspects, the spatial sampling frequency may be at least 2.5 times the optical resolution of the fluorescence imaging system. In some aspects, the spatial sampling frequency may be at least 3 times the optical resolution of the fluorescence imaging system. In some aspects, the system may further comprise an X-Y-Z translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated fashion, wherein each image of the series is or may be acquired for a different field-of-view. In some aspects, a position of the sample plane may be simultaneously adjusted in an X direction, a Y direction, and a Z direction to match the position of an objective lens focal plane in between acquiring images for different fields-of-view. In some aspects, the time required for the simultaneous adjustments in the X direction, Y direction, and Z direction may be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second, or a time falling within a range defined by any two of the foregoing. In some aspects, the system further comprises an autofocus mechanism configured to adjust the focal plane position prior to acquiring an image of a different field-of-view if an error signal indicates that a difference in the position of the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some aspects, the specified error threshold is 100 nm or greater. In some aspects, the specified error threshold is 50 nm or less. In some aspects, the system comprises three or more image sensors, and wherein the system is configured to image fluorescence in each of three or more wavelength ranges onto a different image sensor. In some aspects, a difference in the position of a focal plane for each of the three or more image sensors and the sample plane is lessthan 100 nm. In some aspects, a difference in the position of a focal plane for each of the three or more image sensors and the sample plane is less than 50 nm. In some aspects, the total time required to reposition the sample plane, adjust focus if necessary, and acquire an image is less than 0.4 seconds per field-of-view. In some aspects, the total time required to reposition the sample plane, adjust focus if necessary, and acquire an image is less than 0.3 seconds per field- of-view.

[0280] Also discloser herein are fluorescence imaging systems for dual-side imaging of a flow cell comprising: a) an objective lens configured to collect fluorescence arising from within a specified field-of-view of a sample plane within the flow cell; b) at least one tube lens positioned between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to correct an imaging performance metric for a combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an interior surface of the flow cell, and wherein the flow cell has a wall thickness of at least 700 pm and a gap between an upper interior surface and a lower interior surface of at least 50 pm; wherein the imaging performance metric is substantially the same for imaging the upper interior surface or the lower interior surface of the flow cell without moving an optical compensator into or out of an optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of the tube lens along the optical path, and without moving one or more optical elements of the tube lens into or out of the optical path.

[0281] In some aspects, the objective lens may be a commercially-available microscope objective. In some aspects, the commercially-available microscope objective may have a numerical aperture of at least 0.3. In some aspects, the objective lens may have a working distance of at least 700 pm. In some aspects, the objective lens may be corrected to compensate for a cover slip thickness (or flow cell wall thickness) of 0.17 mm or of greater or lesser thickness than 0.17 mm. In some aspects, the optical system may be corrected to compensate for cover slip thickness, flow cell thickness, or distance between desired focal planes. In some aspects, said correction may be made by inserting a corrective optic, such as a lens or optical assembly into the light path of the optical system. In some aspects, said correction may be made without inserting a corrective optic, such as a lens or optical assembly into the light path of the optical system. In some aspects, the fluorescence imaging system may further comprise an electro-optical phase plate positioned adjacent to the objective lens and between the objective lens and the tube lens, wherein the electro-optical phase plate may provide correction for optical aberrations caused by a fluid filling the gap between the upper interior surface and the lowerinterior surface of the flow cell. In some aspects, the at least one tube lens may be a compound lens comprising three or more optical components. In some aspects, the at least one tube lens is a compound lens comprising four optical components, which may comprise one or more of a first asymmetric convex-convex lens, a second convex-piano lens, a third asymmetric concaveconcave lens, and a fourth asymmetric convex-concave lens which may be present in the order as listed above, or in any alternate order. In some aspects, the at least one tube lens is configured to correct an imaging performance metric for a combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an interior surface of a flow cell having a wall thickness of at least 1 mm. In some aspects, the at least one tube lens is configured to correct an imaging performance metric for a combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an interior surface of a flow cell having a gap of at least 100 pm. In some aspects, the at least one tube lens is configured to correct an imaging performance metric for a combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an interior surface of a flow cell having a gap of at least 200 pm. In some aspects, the system comprises a single objective lens, two tube lenses, and two image sensors, and each of the two tube lenses is designed to provide optimal imaging performance at a different fluorescence wavelength. In some aspects, the system comprises a single objective lens, three tube lenses, and three image sensors, and each of the three tube lenses is designed to provide optimal imaging performance at a different fluorescence wavelength. In some aspects, the system comprises a single objective lens, four tube lenses, and four image sensors, and each of the four tube lenses is designed to provide optimal imaging performance at a different fluorescence wavelength. In some aspects, the design of the objective lens or the at least one tube lens is configured to optimize the modulation transfer function in the mid to high spatial frequency range. In some aspects, the imaging performance metric comprises a measurement of modulation transfer function (MTF) at one or more specified spatial frequencies, defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof. In some aspects, the difference in the imaging performance metric for imaging the upper interior surface and the lower interior surface of the flow cell is less than 10%. In some aspects, the difference in imaging performance metric for imaging the upper interior surface and the lower interior surface of the flow cell is less than 5%. In some aspects, the use of the at least one tube lens provides for an at least equivalent or better improvement in the imaging performance metric for dual-side imaging compared to that for a conventional system comprising an objective lens, a motion-actuated compensator, and an image sensor. In some aspects, the use of the at least one tube lens provides for an at least 10% improvement in the imaging performance metric for dual-side imaging compared to that for a conventional system comprising an objective lens, a motion- actuated compensator, and an image sensor.

[0282] Disclosed herein are illumination systems for use in imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light-guide configured to collect light emitted by the light source and deliver it to a specified field-of-illumination on a support surface comprising tethered biological macromolecules.

[0283] In some aspects, the illumination system further comprises a condenser lens. In some aspects, the specified field-of-illumination has an area of at least 2 mm2. In some aspects, the light delivered to the specified field-of-illumination is of uniform intensity across a specified field-of-view for an imaging system used to acquire images of the support surface. In some aspects, the specified field-of-view has an area of at least 2 mm2. In some aspects, the light delivered to the specified field-of-illumination is of uniform intensity across the specified field- of-view when a coefficient of variation (CV) for light intensity is less than 10%. In some aspects, the light delivered to the specified field-of-illumination is of uniform intensity across the specified field-of-view when a coefficient of variation (CV) for light intensity is less than 5%. In some aspects, the light delivered to the specified field-of-illumination has a speckle contrast value of less than 0.1. In some aspects, the light delivered to the specified field-of-illumination has a speckle contrast value of less than 0.05.

[0284] Imaging modules and systems: It will be understood by those of skill in the art that the disclosed optical systems, imaging systems, or modules may, in some instances, be standalone optical systems designed for imaging a sample or substrate surface. In some instances, they may comprise one or more processors or computers. In some instances, they may comprise one or more software packages that provide instrument control functionality and / or image processing functionality. In some instances, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten-halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical bandpass filters, apertures, and image sensors (e.g., complementary metal oxide semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they may also include mechanical and / or optomechanical components, such as an X-Y translation stage, an X-Y-Z translation stage, a piezoelectric focusing mechanism, and the like. In some instances, they mayfunction as modules, components, sub-assemblies, or sub-systems of larger systems designed for genomics applications (e.g., genetic testing and / or nucleic acid sequencing applications). For example, in some instances, they may function as modules, components, sub-assemblies, or subsystems of larger systems that further comprise light-tight and / or other environmental control housings, temperature control modules, fluidics control modules, fluid dispensing robotics, pick- and-place robotics, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, image processing software packages, data analysis software packages), data storage modules, data communication modules (e.g., Bluetooth, WiFi, intranet, or internet communication hardware and associated software), display modules, or any combination thereof.Methods for sequencing

[0285] Aspects of the present disclosure provide methods for sequencing immobilized or non-immobilized template molecules. The methods may be operated in system 100, for example, in sequencer 114. In some aspects, the immobilized template molecules comprise a plurality of nucleic acid template molecules having one copy of a target sequence of interest. In some aspects, nucleic acid template molecules having one copy of a target sequence of interest may be generated by conducting bridge amplification using linear library molecules. In some aspects, the immobilized template molecules comprise a plurality of nucleic acid template molecules each having two or more tandem copies of a target sequence of interest (e.g., concatemers). In some aspects, nucleic acid template molecules comprising concatemer molecules may be generated by conducting rolling circle amplification of circularized linear library molecules. In some aspects, the non-immobilized template molecules comprise circular molecules. In some aspects, methods for sequencing employ soluble (e.g., non-immobilized) sequencing polymerases or sequencing polymerases that are immobilized to a support.

[0286] In some aspects, the sequencing reactions employ detectably labeled nucleotide analogs. In some aspects, the sequencing reactions employ a two-stage sequencing reaction comprising binding detectably labeled multivalent molecules, and incorporating nucleotide analogs. In some aspects, the sequencing reactions employ non-labeled nucleotide analogs. In some aspects, the sequencing reactions employ phosphate chain labeled nucleotides.

[0287] In some aspects, the immobilized concatemers each comprise tandem repeat units of the sequence-of-interest (e.g., insert region) and any adaptor sequences. For example, the tandemrepeat unit comprises: (i) a left universal adaptor sequence having a binding sequence for a first surface primer (920) (e.g., surface pinning primer), (ii) a left universal adaptor sequence having a binding sequence for a first sequencing primer (940) (e.g., forward sequencing primer), (iii) a sequence-of-interest (910), (iv) a right universal adaptor sequence having a binding sequence for a second sequencing primer (950) (e.g., reverse sequencing primer), (v) a right universal adaptor sequence having a binding sequence for a second surface primer (930) (e.g., surface capture primer), and (vii) a left sample index sequence (960) and / or a right sample index sequence (970). In some aspects, the tandem repeat unit further comprises a left unique identification sequence (980) and / or a right unique identification sequence (990). In some aspects, the tandem repeat unit further comprises at least one binding sequence for a compaction oligonucleotide. In some aspects, FIGS. 15 and 16 show linear library molecules for a unit of a concatemer molecule.

[0288] Specifically, FIG. 9 is a schematic showing an exemplary linear single stranded library molecule (900) which includes: a surface pinning primer binding site (920); an optional left unique identification sequence (980); a left index sequence (960); a forward sequencing primer binding site (940); an insert region having a sequence of interest (910); reverse sequencing primer binding site (950); a right index sequence (970); and a surface capture primer binding site (930).

[0289] FIG. 10 is a schematic showing an exemplary linear single stranded library molecule (900) which includes: a surface pinning primer binding site (920); a left index sequence (960); a forward sequencing primer binding site (940); an insert region having a sequence of interest (910); a reverse sequencing primer binding site (950); a right index sequence (970); an optional right unique identification sequence (990); and a surface capture primer binding site (930).

[0290] FIG. 11 is a schematic of various exemplary configurations of multivalent molecules. Left (Class I): schematics of multivalent molecules having a “starburst” or “helter-skelter” configuration. Center (Class II): a schematic of a multivalent molecule having a dendrimer configuration. Right (Class III): a schematic of multiple multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS with biotin and dNTPs. Nucleotide units are designated ‘N’, biotin is designated ‘B’, and streptavidin is designated ‘SA’.

[0291] The immobilized concatemer may self-collapse into a compact nucleic acid nanoball. Inclusion of one or more compaction oligonucleotides during the RCA reaction may further compact the size and / or shape of the nanoball. An increase in the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer for hybridizing to multiple sequencing primers (e.g., sequencing primers having a universal sequence) which serveas multiple initiation sites for polymerase-catalyzed sequencing reactions. When the sequencing reaction employs detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., having nucleotide units), the signals emitted by the nucleotides or nucleotide units that participate in the parallel sequencing reactions along the concatemer yields an increased signal intensity for each concatemer. Multiple portions of a given concatemer may be simultaneously sequenced. Furthermore, a plurality of binding complexes may form along a particular concatemer molecule, each binding complex comprising a sequencing polymerase bound to a template / primer duplex and bound to a multivalent molecule, wherein the plurality of binding complexes remain stable without dissociation resulting in increased persistence time which increases signal intensity and reduces imaging time.Methods for sequencing using nucleotide analogs

[0292] Aspects of the present disclosure provide methods for sequencing any of the immobilized template molecules described herein, the methods comprising step (a): contacting a sequencing polymerase to (i) a nucleic acid template molecule and (ii) a nucleic acid sequencing primer, wherein the contacting is conducted under a condition suitable to bind the sequencing polymerase to the nucleic acid template molecule which is hybridized to the nucleic acid primer, wherein the nucleic acid template molecule hybridized to the nucleic acid primer forms the nucleic acid duplex. In some aspects, the sequencing polymerase comprises a recombinant mutant sequencing polymerase that may bind and incorporate nucleotide analogs.

[0293] In some aspects, in the methods for sequencing template molecules, the sequencing primer comprises a 3’ extendible end or a 3’ non-extendible end. In some aspects, the plurality of nucleic acid template molecules comprise amplified template molecules (e.g., clonally amplified template molecules). In some aspects, the plurality of nucleic acid template molecules comprise one copy of a target sequence of interest. In some aspects, the plurality of nucleic acid molecules comprise two or more tandem copies of a target sequence of interest (e.g., concatemers). In some aspects, the plurality of nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some aspects, the plurality of nucleic acid primers are in solution or are immobilized to a support. In some aspects, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized to a support, the binding with the first sequencing polymerase generates a plurality of immobilized first complexed polymerases. In some aspects, the plurality of nucleic acid template moleculesand / or nucleic acid primers are immobilized to 102- 1015different sites on a support. In some aspects, the binding of the plurality of template molecules and nucleic acid primers with the plurality of first sequencing polymerases generates a plurality of first complexed polymerases immobilized to 102- 1015different sites on the support. In some aspects, the plurality of immobilized first complexed polymerases on the support are immobilized to pre-determined or to random sites on the support. In some aspects, the plurality of immobilized first complexed polymerases are in fluid communication with each other to permit flowing a solution of reagents (e.g., enzymes including sequencing polymerases, multivalent molecules, nucleotides, and / or divalent cations) onto the support so that the plurality of immobilized complexed polymerases on the support are reacted with the solution of reagents in a massively parallel manner.

[0294] In some aspects, the methods for sequencing further comprise step (b): contacting the sequencing polymerase with a plurality of nucleotides under a condition suitable for binding at least one nucleotide to the sequencing polymerase which is bound to the nucleic acid duplex and suitable for polymerase-catalyzed nucleotide incorporation which extends the sequencing primer by one nucleotide. In some aspects, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some aspects, the plurality of nucleotides comprises at least one nucleotide analog having a chain terminating moiety at the sugar 2’ or 3’ position. In some aspects, the chain terminating moiety is removable from the sugar 2’ or 3’ position to convert the chain terminating moiety to an OH or H group. In some aspects, the plurality of nucleotides comprises at least one nucleotide that lacks a chain terminating moiety. In some aspects, at least on nucleotide is labeled with a detectable reporter moiety (e.g., fluorophore) that emits a detectable signal. The detectable reporter moiety comprises a fluorophore. In some aspects, the fluorophore is attached to the nucleo-base. In some aspects, the fluorophore is attached to the nucleo-base with a linker which is cleavable / removable from the base. In some aspects, at least one of the nucleotides in the plurality is not labeled with a detectable reporter moiety. In some aspects, a particular detectable reporter moiety (e.g., fluorophore) that is attached to the nucleotide may correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to permit detection and identification of the nucleo-base. When the incorporated chain terminating nucleotide is detectably labeled, step (b) further comprises detecting the emitted signal from the incorporated chain terminating nucleotide. In some aspects, step (b) further comprises identifying the nucleo-based of the incorporated chain terminating nucleotide.

[0295] In some aspects, the methods for sequencing further comprise step (c): removing the chain terminating moiety from the incorporated chain terminating nucleotide to generate an extendible 3 ’OH group. In some aspects, step (c) further comprises removing the detectable label from the incorporated chain terminating nucleotide. In some aspects, the sequencing polymerase remains bound to the template molecule which is hybridized to the sequencing primer which is extended by one nucleo-base.

[0296] In some aspects, the methods for sequencing further comprise step (d): repeating steps (b) and (c) at least once.Two-Stage Methods for Nucleic Acid Sequencing

[0297] Aspects of the present disclosure provide a two-stage method for sequencing any of the immobilized template molecules described herein. In some aspects, the first stage generally comprises binding multivalent molecules to complexed polymerases to form multivalent- complexed polymerases, and detecting the multivalent-complexed polymerases.

[0298] In some aspects, the first stage comprises step (a): contacting a plurality of a first sequencing polymerase to (i) a plurality of nucleic acid template molecules and (ii) a plurality of nucleic acid sequencing primers, wherein the contacting is conducted under a condition suitable to bind the plurality of first sequencing polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers thereby forming a plurality of first complexed polymerases each comprising a first sequencing polymerase bound to a nucleic acid duplex wherein the nucleic acid duplex comprises a nucleic acid template molecule hybridized to a nucleic acid primer. In some aspects, the first polymerase comprises a recombinant mutant sequencing polymerase.

[0299] In some aspects, in the methods for sequencing template molecules, the sequencing primer comprises an oligonucleotide having a 3’ extendible end or a 3’ non-extendible end. In some aspects, the plurality of nucleic acid template molecules comprise amplified template molecules (e.g., clonally amplified template molecules). In some aspects, the plurality of nucleic acid template molecules comprise one copy of a target sequence of interest. In some aspects, the plurality of nucleic acid molecules comprise two or more tandem copies of a target sequence of interest (e.g., concatemers). In some aspects, the nucleic acid template molecules in the plurality of nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some aspects, the plurality of nucleic acid template moleculesand / or the plurality of nucleic acid primers are in solution or are immobilized to a support. In some aspects, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized to a support, the binding with the first sequencing polymerase generates a plurality of immobilized first complexed polymerases. In some aspects, the plurality of nucleic acid template molecules and / or nucleic acid primers are immobilized to 102- 1015different sites on a support. In some aspects, the binding of the plurality of template molecules and nucleic acid primers with the plurality of first sequencing polymerases generates a plurality of first complexed polymerases immobilized to 102- 1015different sites on the support. In some aspects, the plurality of immobilized first complexed polymerases on the support are immobilized to pre-determined or to random sites on the support. In some aspects, the plurality of immobilized first complexed polymerases are in fluid communication with each other to permit flowing a solution of reagents (e.g., enzymes including sequencing polymerases, multivalent molecules, nucleotides, and / or divalent cations) onto the support so that the plurality of immobilized complexed polymerases on the support are reacted with the solution of reagents in a massively parallel manner.

[0300] In some aspects, the methods for sequencing further comprise step (b): contacting the plurality of first complexed polymerases with a plurality of multivalent molecules to form a plurality of multival ent-complexed polymerases (e.g., binding complexes). In some aspects, individual multivalent molecules in the plurality of multivalent molecules comprise a core attached to multiple nucleotide arms and each nucleotide arm is attached to a nucleotide (e.g., nucleotide unit) (e.g., FIGS. 11-15). In some aspects, the contacting of step (b) is conducted under a condition suitable for binding complementary nucleotide units of the multivalent molecules to at least two of the plurality of first complexed polymerases thereby forming a plurality of multivalent-complexed polymerases. In some aspects, the condition is suitable for inhibiting polymerase-catalyzed incorporation of the complementary nucleotide units into the primers of the plurality of multivalent-complexed polymerases. In some aspects, the plurality of multivalent molecules comprise at least one multivalent molecule having multiple nucleotide arms (e.g., FIGS. 11-14) each attached with a nucleotide analog (e.g., nucleotide analog unit), where the nucleotide analog includes a chain terminating moiety at the sugar 2’ and / or 3’ position. In some aspects, the plurality of multivalent molecules comprises at least one multivalent molecule comprising multiple nucleotide arms each attached with a nucleotide unit that lacks a chain terminating moiety. In some aspects, at least one of the multivalent molecules in the plurality of multivalent molecules is labeled with a detectable reporter moiety that emits asignal. In some aspects, the detectable reporter moiety comprises a fluorophore. In some aspects, the contacting of step (b) is conducted in the presence of at least one non-catalytic cation comprising strontium, barium and / or calcium.

[0301] In some aspects, the methods for sequencing further comprise step (c): detecting the plurality of multivalent-complexed polymerases. In some aspects, the detecting includes detecting the signals emitted by the multivalent molecules that are bound to the complexed polymerases, where the complementary nucleotide units of the multivalent molecules are bound to the primers but incorporation of the complementary nucleotide units is inhibited. In some aspects, the multivalent molecules are labeled with a detectable reporter moiety to permit detection. In some aspects, the labeled multivalent molecules comprise a fluorophore attached to the core, linker and / or nucleotide unit of the multivalent molecules.

[0302] In some aspects, the methods for sequencing further comprise step (d): identifying the nucleo-base of the complementary nucleotide units that are bound to the plurality of first complexed polymerases, thereby determining the sequence of the template molecule. In some aspects, the multivalent molecules are labeled with a detectable reporter moiety that corresponds to the particular nucleotide units attached to the nucleotide arms to permit identification of the complementary nucleotide units (e.g., nucleotide base adenine, guanine, cytosine, thymine or uracil) that are bound to the plurality of first complexed polymerases.

[0303] In some aspects, the methods for sequencing further comprise step (e): dissociating the plurality of multivalent-complexed polymerases and removing the plurality of first sequencing polymerases and their bound multivalent molecules, and retaining the plurality of nucleic acid duplexes.

[0304] In some aspects, the second stage of the two-stage sequencing method generally comprises nucleotide incorporation. In some aspects, the methods for sequencing further comprises step (I): contacting the plurality of the retained nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is conducted under a condition suitable for binding the plurality of second sequencing polymerases to the plurality of the retained nucleic acid duplexes, thereby forming a plurality of second complexed polymerases each comprising a second sequencing polymerase bound to a nucleic acid duplex. In some aspects, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.

[0305] In some aspects, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is 100% identical to the amino acid sequence as the plurality of thesecond sequencing polymerases of step (f). In some aspects, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that differs from the amino acid sequence of the plurality of the second sequencing polymerases of step (f).

[0306] In some aspects, the methods for sequencing further comprise step (g): contacting the plurality of second complexed polymerases with a plurality of nucleotides, wherein the contacting is conducted under a condition suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second complexed polymerases thereby forming a plurality of nucleotide-complexed polymerases. In some aspects, the contacting of step (g) is conducted under a condition that is suitable for promoting polymerase-catalyzed incorporation of the bound complementary nucleotides into the primers of the nucleotide-complexed polymerases thereby extending the sequencing primer by one nucleo-base. In some aspects, the incorporating the nucleotide into the 3’ end of the sequencing primer in step (g) comprises a primer extension reaction. In some aspects, the contacting of step (g) is conducted in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some aspects, the plurality of nucleotides comprise native nucleotides (e.g., non-analog nucleotides) or nucleotide analogs. In some aspects, the plurality of nucleotides comprise a 2’ and / or 3’ chain terminating moiety which is removable or is not removable. In some aspects, at least one of the nucleotides in the plurality is not labeled with a detectable reporter moiety. In some aspects, the plurality of nucleotides are non-labeled. In some aspects, the plurality of nucleotides comprises a plurality of nucleotides labeled with detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some aspects, the fluorophore is attached to the nucleotide base. In some aspects, the fluorophore is attached to the nucleotide base with a linker which is cleavable / removable from the base or is not removable from the base. In some aspects, a particular detectable reporter moiety (e.g., fluorophore) that is attached to the nucleotide may correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to permit detection and identification of the nucleotide base.

[0307] In some aspects, when the plurality of nucleotides in step (g) are detectably labeled, the methods for sequencing further comprise step (h): detecting the complementary nucleotides which are incorporated into the primers of the nucleotide-complexed polymerases. In some aspects, the plurality of nucleotides are labeled with a detectable reporter moiety to permit detection. In some aspects, when the plurality of nucleotides in step (g) are non-labeled, the detecting of step (h) is omitted.- n -

[0308] In some aspects, when the plurality of nucleotides in step (g) are detectably labeled, the methods for sequencing further comprise step (i): identifying the bases of the complementary nucleotides which are incorporated into the primers of the nucleotide-complexed polymerases. In some aspects, the identification of the incorporated complementary nucleotides in step (i) may be used to confirm the identity of the complementary nucleotides of the multivalent molecules that are bound to the plurality of first complexed polymerases in step (d). In some aspects, the identifying of step (i) may be used to determine the sequence of the nucleic acid template molecules. In some aspects, when the plurality of nucleotides in step (g) are non-labeled, the identifying of step (i) is omitted.

[0309] In some aspects, the methods for sequencing further comprise step (j): removing the chain terminating moiety from the incorporated nucleotide when step (g) is conducted by contacting the plurality of second complexed polymerases with a plurality of nucleotides that comprise at least one nucleotide having a 2’ and / or 3’ chain terminating moiety.

[0310] In some aspects, the methods for sequencing further comprise step (k): repeating steps (a) - (j) at least once. In some aspects, the sequence of the nucleic acid template molecules may be determined by detecting and identifying the multivalent molecules that bind the sequencing polymerases but do not incorporate into the 3’ end of the primer at steps (c) and (d). In some aspects, the sequence of the nucleic acid template molecule may be determined (or confirmed) by detecting and identifying the nucleotide that incorporates into the 3’ end of the primer at steps (h) and (i).

[0311] In some aspects, in any of the methods for sequencing nucleic acid molecules, the binding of the plurality of first complexed polymerases with the plurality of multivalent molecules forms at least one avidity complex, the method comprising the steps: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemer template molecule thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemer template molecule thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, wherein the first and second binding complexes which include the same multivalent molecule forms an avidity complex. In some aspects, the first sequencing polymerase comprises any wild type or mutant polymerase described herein. In some aspects, the second sequencing polymerase comprises any wild type ormutant polymerase described herein. The concatemer template molecule comprises tandem repeat sequences of a sequence of interest and at least one universal sequencing primer binding site. The first and second nucleic acid primers may bind to a sequencing primer binding site along the concatemer template molecule. Exemplary multivalent molecules are shown in FIGS. 11-14.

[0312] In some aspects, in any of the methods for sequencing nucleic acid molecules, wherein the method includes binding the plurality of first complexed polymerases with the plurality of multivalent molecules to form at least one avidity complex, the method comprising the steps: (a) contacting the plurality of sequencing polymerases and the plurality of nucleic acid primers with different portions of a concatemer nucleic acid concatemer molecule to form at least first and second complexed polymerases on the same concatemer template molecule; (b) contacting a plurality of multivalent molecules to the at least first and second complexed polymerases on the same concatemer template molecule, under conditions suitable to bind a single multivalent molecule from the plurality to the first and second complexed polymerases, wherein at least a first nucleotide unit of the single multivalent molecule is bound to the first complexed polymerase which includes a first primer hybridized to a first portion of the concatemer template molecule thereby forming a first binding complex (e.g., first ternary complex), and wherein at least a second nucleotide unit of the single multivalent molecule is bound to the second complexed polymerase which includes a second primer hybridized to a second portion of the concatemer template molecule thereby forming a second binding complex (e.g., second ternary complex), wherein the contacting is conducted under a condition suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, and wherein the first and second binding complexes which are bound to the same multivalent molecule forms an avidity complex; and (c) detecting the first and second binding complexes on the same concatemer template molecule, and (d) identifying the first nucleotide unit in the first binding complex thereby determining the sequence of the first portion of the concatemer template molecule, and identifying the second nucleotide unit in the second binding complex thereby determining the sequence of the second portion of the concatemer template molecule. In some aspects, the plurality of sequencing polymerases comprise any wild type or mutant sequencing polymerase described herein. The concatemer template molecule comprises tandem repeat sequences of a sequence of interest and at least one universal sequencing primer binding site. The plurality of nucleic acid primers may bind to asequencing primer binding site along the concatemer template molecule. Exemplary multivalent molecules are shown in FIGS. 11-14.Sequencing-by-Binding

[0313] Aspects of the present disclosure provide methods for sequencing any of the immobilized template molecules described herein, wherein the sequencing methods comprise a sequencing-by-binding (SBB) procedure which employs non-labeled chain-terminating nucleotides. In some aspects, the sequencing-by-binding (SBB) method comprises the steps of(a) sequentially contacting a primed template nucleic acid with at least two separate mixtures under ternary complex stabilizing conditions, wherein the at least two separate mixtures each include a polymerase and a nucleotide, whereby the sequentially contacting results in the primed template nucleic acid being contacted, under the ternary complex stabilizing conditions, with nucleotide cognates for first, second and third base type base types in the template; (b) examining the at least two separate mixtures to determine whether a ternary complex formed; and (c) identifying the next correct nucleotide for the primed template nucleic acid molecule, wherein the next correct nucleotide is identified as a cognate of the first, second or third base type if ternary complex is detected in step (b), and wherein the next correct nucleotide is imputed to be a nucleotide cognate of a fourth base type based on the absence of a ternary complex in step (b); (d) adding a next correct nucleotide to the primer of the primed template nucleic acid after step(b), thereby producing an extended primer; and (e) repeating steps (a) through (d) at least once on the primed template nucleic acid that comprises the extended primer. Exemplary sequencing-by- binding methods are described in U.S. patent Nos. 10,246,744 and 10,731,141 (where the contents of both patents are hereby incorporated by reference in their entireties).Methods for Sequencing using Phosphate-Chain Labeled Nucleotides

[0314] Aspects of the present disclosure provide methods for sequencing using immobilized sequencing polymerases which bind non-immobilized template molecules, wherein the sequencing reactions are conducted with phosphate-chain labeled nucleotides. In some aspects, the sequencing methods comprise step (a): providing a support having a plurality of sequencing polymerases immobilized thereon. In some aspects, the sequencing polymerase comprises a processive DNA polymerase. In some aspects, the sequencing polymerase comprises a wild type or mutant DNA polymerase, including for example a Phi29 DNA polymerase. In some aspects,the support comprise a plurality of separate compartments and a sequencing polymerase is immobilized to the bottom of a compartment. In some aspects, the separate compartments comprise a silica bottom through which light may penetrate. In some aspects, the separate compartments comprise a silica bottom configured with a nanophotonic confinement structure comprising a hole in a metal cladding film (e.g., aluminum cladding film). In some aspects, the hole in the metal cladding has a small aperture, for example, approximately 70 nm. In some aspects, the height of the nanophotonic confinement structure is approximately 100 nm. In some aspects, the nanophotonic confinement structure comprises a zero mode waveguide (ZMW). In some aspects, the nanophotonic confinement structure contains a liquid.

[0315] In some aspects, the sequencing method further comprises step (b): contacting the plurality of immobilized sequencing polymerases with a plurality of single stranded circular nucleic acid template molecules and a plurality of oligonucleotide sequencing primers, under a condition suitable for individual immobilized sequencing polymerases to bind a single stranded circular template molecule, and suitable for individual sequencing primers to hybridize to individual single stranded circular template molecules, thereby generating a plurality of polymerase / template / primer complexes. In some aspects, the individual sequencing primers hybridize to a universal sequencing primer binding site on the single stranded circular template molecule.

[0316] In some aspects, the sequencing method further comprises step (c): contacting the plurality of polymerase / template / primer complexes with a plurality of phosphate chain labeled nucleotides each comprising an aromatic base, a five carbon sugar (e.g., ribose or deoxyribose), and phosphate chain comprising 3-20 phosphate groups, where the terminal phosphate group is linked to a detectable reporter moiety (e.g., a fluorophore). The first, second and third phosphate groups may be referred to as alpha, beta and gamma phosphate groups. In some aspects, a particular detectable reporter moiety which is attached to the terminal phosphate group corresponds to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to permit detection and identification of the nucleo-base. In some aspects, the plurality of polymerase / template / primer complexes are contacted with the plurality of phosphate chain labeled nucleotides under a condition suitable for polymerase-catalyzed nucleotide incorporation. In some aspects, the sequencing polymerases are capable of binding a complementary phosphate chain labeled nucleotide and incorporating the complementary nucleotide opposite a nucleotide in a template molecule. In some aspect, the polymerase-catalyzed nucleotide incorporationreaction cleaves between the alpha and beta phosphate groups thereby releasing a multiphosphate chain linked to a fluorophore.

[0317] In some aspects, the sequencing method further comprises step (d): detecting the fluorescent signal emitted by the phosphate chain labeled nucleotide that is bound by the sequencing polymerase, and incorporated into the terminal end of the sequencing primer. In some aspects, step (d) further comprises identifying the phosphate chain labeled nucleotide that is bound by the sequencing polymerase, and incorporated into the terminal end of the sequencing primer.

[0318] In some aspects, the sequencing method further comprises step (d): repeating steps (c) - (d) at least once. In some aspects, sequencing methods that employ phosphate chain labeled nucleotides may be conducted according to the methods described in U.S. patent Nos. 7,170,050; 7,302,146; and / or 7,405,281.Sequencing Polymerases

[0319] Aspects of the present disclosure provide methods for sequencing nucleic acid molecules, where any of the sequencing methods described herein employ at least one type of sequencing polymerase and a plurality of nucleotides, or employ at least one type of sequencing polymerase and a plurality of nucleotides and a plurality of multivalent molecules. In some aspects, the sequencing polymerase(s) is / are capable of incorporating a complementary nucleotide opposite a nucleotide in a template molecule. In some aspects, the sequencing polymerase(s) is / are capable of binding a complementary nucleotide unit of a multivalent molecule opposite a nucleotide in a template molecule. In some aspects, the plurality of sequencing polymerases comprise recombinant mutant polymerases.

[0320] Examples of suitable polymerases for use in sequencing with nucleotides and / or multivalent molecules include but are not limited to: Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaeon; Candidatus Hadarchaeum Yellowstonense; Hadesarchaea archaeon; Euryarchaeota archaeon; Thermoplasmata archaeon; Thermococcus polymerases such as Thermococcus litoralis, bacteriophage T7 DNA polymerase; human alpha, delta and epsilon DNA polymerases; bacteriophage polymerases such as T4, RB69 and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; E. coli DNA polymerase III alpha and epsilon; 9 degree N polymerase; reverse transcriptasessuch as HIV type M or O reverse transcriptases; avian myeloblastosis virus reverse transcriptase; Moloney Murine Leukemia Virus (MMLV) reverse transcriptase; or telomerase. Further nonlimiting examples of DNA polymerases include those from various Archaea genera, such as, Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Pyrodictium, Staphylothermus, Stetteria, Sulfolobus, Thermococcus, and Vulcanisaeta and the like or variants thereof, including such polymerases as are known in the art such as 9 degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo and RB69 polymerases.Nucleotides

[0321] Aspects of the present disclosure provide methods for sequencing nucleic acid molecules, where any of the sequencing methods described herein employ at least one nucleotide. The nucleotides comprise a base, sugar and at least one phosphate group. In some aspects, at least one nucleotide in the plurality comprises an aromatic base, a five carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1-10 phosphate groups). The plurality of nucleotides may comprise at least one type of nucleotide selected from a group consisting of dATP, dGTP, dCTP, dTTP and dUTP. The plurality of nucleotides may comprise at a mixture of any combination of two or more types of nucleotides selected from a group consisting of dATP, dGTP, dCTP, dTTP and / or dUTP. In some aspects, at least one nucleotide in the plurality is not a nucleotide analog. In some aspects, at least one nucleotide in the plurality comprises a nucleotide analog.

[0322] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide in the plurality of nucleotides comprise a chain of one, two or three phosphorus atoms where the chain is typically attached to the 5’ carbon of the sugar moiety via an ester or phosphoramide linkage. In some aspects, at least one nucleotide in the plurality is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene or ethylene. In some aspects, the phosphorus atoms in the chain include substituted side groups including O, S or BH3. In some aspects, the chain includes phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoroamidite groups.

[0323] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog having a chain terminating moiety (e.g., blocking moiety) at the sugar 2’position, at the sugar 3’ position, or at the sugar 2’ and 3’ position. In some aspects, the chain terminating moiety may inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in a nascent strand during a primer extension reaction. In some aspects, the chain terminating moiety is attached to the 3’ sugar position where the sugar comprises a ribose or deoxyribose sugar moiety. In some aspects, the chain terminating moiety is removable / cleavable from the 3’ sugar position to generate a nucleotide having a 3 ’OH sugar group which is extendible with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some aspects, the chain terminating moiety comprises an alkyl group, alkenyl group, alkynyl group, allyl group, aryl group, benzyl group, azide group, amine group, amide group, keto group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, silyl or acetal group. In some aspects, the chain terminating moiety is cleavable / removable from the nucleotide, for example by reacting the chain terminating moiety with a chemical agent, pH change, light or heat. In some aspects, the chain terminating moieties alkyl, alkenyl, alkynyl and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine, or with 2,3-Dichloro-5,6-dicyano-l,4-benzo-quinone (DDQ). In some aspects, the chain terminating moieties aryl and benzyl are cleavable with H2 Pd / C. In some aspects, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphine or with a thiol group including beta-mercaptoethanol or dithiothritol (DTT). In some aspects, the chain terminating moiety carbonate is cleavable with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some aspects, the chain terminating moieties urea and silyl are cleavable with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride. In some aspects, the chain terminating moiety may be cleavable / removable with nitrous acid. In some aspects, a chain terminating moiety may be cleavable / removable using a solution comprising nitrite, such as, for example, a combination of nitrite with an acid such as acetic acid, sulfuric acid, or nitric acid. In some further aspects, said solution may comprise an organic acid.

[0324] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog having a chain terminating moiety (e.g., blocking moiety) at the sugar 2’ position, at the sugar 3’ position, or at the sugar 2’ and 3’ position. In some aspects, the chain terminating moiety comprises an azide, azido or azidomethyl group. In some aspects, the chain terminating moiety comprises a 3’-O-azido or 3’-O-azidomethyl group. In some aspects, thechain terminating moieties azide, azido and azidomethyl group are cleavable / removable with a phosphine compound. In some aspects, the phosphine compound comprises a derivatized trialkyl phosphine moiety or a derivatized tri-aryl phosphine moiety. In some aspects, the phosphine compound comprises Tris(2-carboxyethyl)phosphine (TCEP) or bis-sulfo triphenyl phosphine (BS-TPP) or Tri(hydroxyproyl)phosphine (THPP). In some aspects, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP). In some aspects, the chain terminating moiety comprising one or more of a 3’-O-amino group, a 3’-O-aminomethyl group, a 3’-O-methylamino group, or derivatives thereof may be cleaved with nitrous acid, through a mechanism utilizing nitrous acid, or using a solution comprising nitrous acid. In some aspects, the chain terminating moiety comprising one or more of a 3’-O-amino group, a 3’-O-aminomethyl group, a 3’-O- methylamino group, or derivatives thereof may be cleaved using a solution comprising nitrite. In some aspects, for example, nitrite may be combined with or contacted with an acid such as acetic acid, sulfuric acid, or nitric acid. In some further aspects, for example, nitrite may be combined with or contacted with an organic acid such as, for example, formic acid, acetic acid, propionic acid, butyric acid, isobutyric acid, or the like. In some aspects, the chain terminating moiety comprises a 3’-acetal moiety which may be cleaved with a palladium deblocking reagent (e.g., Pd(0)).

[0325] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, the nucleotide comprises a chain terminating moiety which is selected from a group consisting of 3’-deoxy nucleotides, 2’, 3 ’-dideoxynucleotides, 3’-methyl, 3’-azido, 3’- azidom ethyl, 3’-O-azidoalkyl, 3’-O-ethynyl, 3’-O-aminoalkyl, 3’-O-fluoroalkyl, 3’-fluoromethyl, 3’-difluoromethyl, 3’-trifluoromethyl, 3 ’-sulfonyl, 3 ’-malonyl, 3 ’-amino, 3’-O-amino, 3’- sulfhydral, 3 ’-aminomethyl, 3’-ethyl, 3’butyl, 3" -tert butyl, 3’- Fluorenylmethyloxy carbonyl, 3’ / C / 7- Butyl oxy carbonyl, 3’-O-alkyl hydroxylamino group, 3’-phosphorothioate, 3-O-benzyl, and 3’-O-benzyl, 3 -acetal moiety or derivatives thereof.

[0326] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, the plurality of nucleotides comprises a plurality of nucleotides labeled with detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some aspects, the fluorophore is attached to the nucleotide base. In some aspects, the fluorophore is attached to the nucleotide base with a linker which is cleavable / removable from the base. In some aspects, at least one of the nucleotides in the plurality is not labeled with a detectable reporter moiety. In some aspects, a particular detectable reporter moiety (e.g., fluorophore) that isattached to the nucleotide may correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to permit detection and identification of the nucleotide base.

[0327] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an alkyl group, alkenyl group, alkynyl group, allyl group, aryl group, benzyl group, azide group, amine group, amide group, keto group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, or silyl group. In some aspects, the cleavable linker on the base is cleavable / removable from the base by reacting the cleavable moiety with a chemical agent, pH change, light or heat. In some aspects, the cleavable moieties alkyl, alkenyl, alkynyl and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine, or with 2,3-Dichloro-5,6-dicyano-l,4-benzo-quinone (DDQ). In some aspects, the cleavable moieties aryl and benzyl are cleavable with H2 Pd / C. In some aspects, the cleavable moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphine or with a thiol group including beta-mercaptoethanol or dithiothritol (DTT). In some aspects, the cleavable moiety carbonate is cleavable with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some aspects, the cleavable moieties urea and silyl are cleavable with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.

[0328] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, the cleavable linker on the nucleotide base comprises cleavable moiety including an azide, azido or azidomethyl group. In some aspects, the cleavable moieties azide, azido and azidomethyl group are cleavable / removable with a phosphine compound. In some aspects, the phosphine compound comprises a derivatized tri-alkyl phosphine moiety or a derivatized tri-aryl phosphine moiety. In some aspects, the phosphine compound comprises Tris(2-carboxyethyl)phosphine (TCEP) or bis-sulfo triphenyl phosphine (BS-TPP) or Tri(hydroxyproyl)phosphine (THPP). In some aspects, the cleaving agent comprises 4- dimethylaminopyridine (4-DMAP).

[0329] In some aspects, in any of the methods for sequencing nucleic acid molecules described herein, the chain terminating moiety (e.g., at the sugar 2’ and / or sugar 3’ position) and the cleavable linker on the nucleotide base have the same or different cleavable moieties. In some aspects, the chain terminating moiety (e.g., at the sugar 2’ and / or sugar 3’ position) and the detectable reporter moiety linked to the base are chemically cleavable / removable with the same chemical agent. In some aspects, the chain terminating moiety (e.g., at the sugar 2’ and / or sugar3’ position) and the detectable reporter moiety linked to the base are chemically cleavable / removable with different chemical agents.Multivalent Molecules

[0330] Aspects of the present disclosure provide methods for sequencing nucleic acid molecules, where any of the sequencing methods described herein employ at least one multivalent molecule. In some aspects, the multivalent molecule comprises a plurality of nucleotide arms attached to a core and having any configuration including a starburst, helter skelter, or bottle brush configuration (e.g., FIG. 11). The multivalent molecule comprises: (1) a core; and (2) a plurality of nucleotide arms which comprise (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, wherein the linker is attached to the nucleotide unit. In some aspects, the nucleotide unit comprises a base, sugar and at least one phosphate group, and the linker is attached to the nucleotide unit through the base. In some aspects, the linker comprises an aliphatic chain or an oligo ethylene glycol chain where both linker chains having 2-6 subunits. In some aspects, the linker also includes an aromatic moiety.

[0331] FIG. 12 is a schematic of an exemplary multivalent molecule comprising a generic core attached to a plurality of nucleotide-arms.

[0332] FIG. 13 is a schematic of an exemplary multivalent molecule comprising a dendrimer core attached to a plurality of nucleotide-arms.

[0333] FIG. 14 shows a schematic of an exemplary multivalent molecule comprising a core attached to a plurality of nucleotide-arms, where the nucleotide arms comprise biotin, spacer, linker and a nucleotide unit.

[0334] FIG. 15 is a schematic of an exemplary nucleotide-arm comprising a core attachment moiety, spacer, linker and nucleotide unit.

[0335] FIG. 16 shows the chemical structure of an exemplary spacer (top), and the chemical structures of various exemplary linkers, including an 11 -atom Linker, 16-atom Linker, 23 -atom Linker and an N3 Linker (bottom).

[0336] FIG. 17 shows the chemical structures of various exemplary linkers, including Linkers 1-9.

[0337] FIG. 18 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.

[0338] FIG. 19 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.

[0339] FIG. 20 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.

[0340] FIG. 21 shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.

[0341] FIG. 22 shows the chemical structure of an exemplary biotinylated nucleotide-arm. In this exemplary example, the nucleotide unit is connected to the linker via a propargyl amine attachment at the 5 position of a pyrimidine base or the 7 position of a purine base.

[0342] In some aspects, a multivalent molecule comprises a core attached to multiple nucleotide arms, and wherein the multiple nucleotide arms have the same type of nucleotide unit which is selected from a group consisting of dATP, dGTP, dCTP, dTTP and dUTP.

[0343] In some aspects, a multivalent molecule comprises a core attached to multiple nucleotide arms, where each arm includes a nucleotide unit. The nucleotide unit comprises an aromatic base, a five carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1-10 phosphate groups). The plurality of multivalent molecules may comprise one type multivalent molecule having one type of nucleotide unit selected from a group consisting of dATP, dGTP, dCTP, dTTP and dUTP. The plurality of multivalent molecules may comprise at a mixture of any combination of two or more types of multivalent molecules, where individual multivalent molecules in the mixture comprise nucleotide units selected from a group consisting of dATP, dGTP, dCTP, dTTP and / or dUTP.

[0344] In some aspects, the nucleotide unit comprises a chain of one, two or three phosphorus atoms where the chain is typically attached to the 5’ carbon of the sugar moiety via an ester or phosphoramide linkage. In some aspects, at least one nucleotide unit is a nucleotide analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene or ethylene. In some aspects, the phosphorus atoms in the chain include substituted side groups including O, S or BH3. In some aspects, the chain includes phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoroamidite groups.

[0345] In some aspects, the multivalent molecule comprises a core attached to multiple nucleotide arms, and wherein individual nucleotide arms comprise a nucleotide unit which is anucleotide analog having a chain terminating moiety (e.g., blocking moiety) at the sugar 2’ position, at the sugar 3’ position, or at the sugar 2’ and 3’ position. In some aspects, the nucleotide unit comprises a chain terminating moiety (e.g., blocking moiety) at the sugar 2’ position, at the sugar 3’ position, or at the sugar 2’ and 3’ position. In some aspects, the chain terminating moiety may inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in a nascent strand during a primer extension reaction. In some aspects, the chain terminating moiety is attached to the 3’ sugar position where the sugar comprises a ribose or deoxyribose sugar moiety. In some aspects, the chain terminating moiety is removable / cleavable from the 3’ sugar position to generate a nucleotide having a 3 ’OH sugar group which is extendible with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some aspects, the chain terminating moiety comprises an alkyl group, alkenyl group, alkynyl group, allyl group, aryl group, benzyl group, azide group, amine group, amide group, keto group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, or silyl group. In some aspects, the chain terminating moiety is cleavable / removable from the nucleotide unit, for example by reacting the chain terminating moiety with a chemical agent, pH change, light or heat. In some aspects, the chain terminating moieties alkyl, alkenyl, alkynyl and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine, or with 2,3-Dichloro-5,6- dicyano-l,4-benzo-quinone (DDQ). In some aspects, the chain terminating moieties aryl and benzyl are cleavable with H2 Pd / C. In some aspects, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphine or with a thiol group including beta-mercaptoethanol or dithiothritol (DTT). In some aspects, the chain terminating moiety carbonate is cleavable with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some aspects, the chain terminating moieties urea and silyl are cleavable with tetrabutylammonium fluoride, pyridine- HF, with ammonium fluoride, or with triethylamine trihydrofluoride.

[0346] In some aspects, the nucleotide unit comprises a chain terminating moiety (e.g., blocking moiety) at the sugar 2’ position, at the sugar 3’ position, or at the sugar 2’ and 3’ position. In some aspects, the chain terminating moiety comprises an azide, azido or azidomethyl group. In some aspects, the chain terminating moiety comprises a 3’-O-azido or 3’-O- azidomethyl group. In some aspects, the chain terminating moieties azide, azido and azidomethyl group are cleavable / removable with a phosphine compound. In some aspects, the phosphine compound comprises a derivatized tri-alkyl phosphine moiety or a derivatized tri-aryl phosphinemoiety. In some aspects, the phosphine compound comprises Tris(2-carboxyethyl)phosphine (TCEP) or bis-sulfo triphenyl phosphine (BS-TPP) or Tri(hydroxyproyl)phosphine (THPP). In some aspects, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).

[0347] In some aspects, the nucleotide unit comprising a chain terminating moiety which is selected from a group consisting of 3’-deoxy nucleotides, 2’, 3 ’-dideoxynucleotides, 3’-methyl, 3 ’-azido, 3 ’-azidomethyl, 3’-O-azidoalkyl, 3’-O-ethynyl, 3’-O-aminoalkyl, 3’-O-fluoroalkyl, 3’- fluorom ethyl, 3’-difluoromethyl, 3’-trifluoromethyl, 3 ’-sulfonyl, 3 ’-malonyl, 3 ’-amino, 3’-O- amino, 3’-sulfhydral, 3 ’-aminomethyl, 3 ’-ethyl, 3 ’butyl, 3 ’-tert butyl, 3’- Fluorenylmethyloxycarbonyl, 3’ tert-Butyloxycarbonyl, 3’-O-alkyl hydroxylamino group, 3’- phosphorothioate, and 3-O-benzyl, or derivatives thereof.

[0348] In some aspects, the multivalent molecule comprises a core attached to multiple nucleotide arms, wherein the nucleotide arms comprise a spacer, linker and nucleotide unit, and wherein the core, linker and / or nucleotide unit is labeled with detectable reporter moiety. In some aspects, the detectable reporter moiety comprises a fluorophore. In some aspects, a particular detectable reporter moiety (e.g., fluorophore) that is attached to the multivalent molecule may correspond to the base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) of the nucleotide unit to permit detection and identification of the nucleotide base.

[0349] In some aspects, at least one nucleotide arm of a multivalent molecule has a nucleotide unit that is attached to a detectable reporter moiety. In some aspects, the detectable reporter moiety is attached to the nucleotide base. In some aspects, the detectable reporter moiety comprises a fluorophore. In some aspects, a particular detectable reporter moiety (e.g., fluorophore) that is attached to the multivalent molecule may correspond to the base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) of the nucleotide unit to permit detection and identification of the nucleotide base.

[0350] In some aspects, the core of a multivalent molecule comprises an avidin-like or streptavidin-like moiety and the core attachment moiety comprises biotin. In some aspects, the core comprises an streptavidin-type or avidin-type moiety which includes an avidin protein, as well as any derivatives, analogs and other non-native forms of avidin that may bind to at least one biotin moiety. Other forms of avidin moieties include native and recombinant avidin and streptavidin as well as derivatized molecules, e.g. nonglycosylated avidin and truncated streptavidins . For example, avidin moiety includes deglycosylated forms of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), as well as derivatized forms, for example, N-acyl avidins, e.g., N-acetyl, N-phthalyland N-succinyl avidin, and the commercially-available products EXTRA VIDIN, CAPTAVIDIN, NEUTRA VIDIN and NEUTRALITE AVIDIN.

[0351] In some aspects, any of the methods for sequencing nucleic acid molecules described herein may include forming a binding complex, where the binding complex comprises (i) a polymerase, a nucleic acid template molecule duplexed with a primer, and a nucleotide, or the binding complex comprises (ii) a polymerase, a nucleic acid template molecule duplexed with a primer, and a nucleotide unit of a multivalent molecule. In some aspects, the binding complex has a persistence time of greater than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or 1 second. The binding complex has a persistence time of greater than about 0.1-0.25 seconds, or about 0.25-0.5 seconds, or about 0.5-0.75 seconds, or about 0.75-1 second, or about 1-2 seconds, or about 2-3 seconds, or about 3-4 second, or about 4-5 seconds, and / or wherein the method is or may be carried out at a temperature of at or above 15 °C, at or above 20 °C, at or above 25 °C, at or above 35 °C, at or above 37 °C, at or above 42 °C at or above 55 °C at or above 60 °C, or at or above 72 °C, or at or above 80 °C, or within a range defined by any of the foregoing. The binding complex (e.g., ternary complex) remains stable until subjected to a condition that causes dissociation of interactions between any of the polymerase, template molecule, primer and / or the nucleotide unit or the nucleotide. For example, a dissociating condition comprises contacting the binding complex with any one or any combination of a detergent, EDTA and / or water. In some aspects, the present disclosure provides said method wherein the binding complex is deposited on, attached to, or hybridized to, a surface showing a contrast to noise ratio in the detecting step of greater than 20. In some aspects, the present disclosure provides said method wherein the contacting is performed under a condition that stabilizes the binding complex when the nucleotide or nucleotide unit is complementary to a next base of the template nucleic acid, and destabilizes the binding complex when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.Compaction Oligonucleotides

[0352] A compaction oligonucleotide comprises a single-stranded linear oligonucleotide having a 5’ region that may hybridize to a first portion of a concatemer molecule and the compaction oligonucleotide having a 3’ region that may hybridize to a second portion of the concatemer molecule (e.g., the same concatemer molecule). In some aspects, hybridization of the compaction oligonucleotides to individual concatemer molecules causes the concatemermolecule to collapse or fold into a DNA nanoball which is more compact in shape and size compared to a non-collapsed DNA molecule. A spot image of a DNA nanoball may be represented as a Gaussian spot and the size may be measured as a full width half maximum (FWHM). A smaller spot size as indicated by a smaller FWHM typically correlates with an improved image of the spot. In some aspects, the FWHM of a DNA nanoball spot may be about 10 um or smaller. The DNA nanoball may be a compact nucleic acid structure having a full width half maximum (FWHM) that is smaller compared to a concatemer that is not collapsed / folded into a DNA nanoball.

[0353] In some aspects, compaction oligonucleotides comprise a single stranded oligonucleotides comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotides may be any length, including 20-150 nucleotides, or 30-100 nucleotides, or 40- 80 nucleotides in length.

[0354] In some aspects, the compaction oligonucleotides comprises a 5’ region and a 3’ region, and optionally an intervening region between the 5’ and 3’ regions. The intervening region may be any length, for example about 2-20 nucleotides in length. The intervening region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT or UUU). The intervening region comprises a non-homopolymer sequence.

[0355] The 5’ region of the compaction oligonucleotides may be wholly complementary or partially complementary along its length to a first portion of a concatemer molecule. The 3’ region of the compaction oligonucleotides may be wholly complementary or partially complementary along its length to a second portion of a concatemer molecule. The 5’ region of the compaction oligonucleotides may hybridize to a first universal sequence portion of a concatemer molecule. The 3’ region of the compaction oligonucleotides may hybridize to a second universal sequence portion of a concatemer molecule. The 5’ and 3’ regions of the compaction oligonucleotide may hybridize to the concatemer to pull together distal portions of the concatemer causing compaction of the concatemer to form a DNA nanoball.

[0356] The 5’ region of the compaction oligonucleotide may have the same sequence as the 3’ region. The 5’ region of the compaction oligonucleotide may have a sequence that is different from the 3’ region. The 3’ region of the compaction oligonucleotide may have a sequence that is a reverse sequence of the 5’ region.

[0357] In some aspects sequence data may be derived through nanopore sequencing, which comprises sequencing of a nucleic acid by translocating said nucleic acid across a membrane, such as through a pore, and wherein sequence reads or base calls are made by measuring one ormore signals during the translocation event, such as impedance, current, voltage, or capacitance. In some aspects, the identity of a nucleotide may be determined by distinctive electrical signatures, such as the timing, duration, extent, or lineshape of a current block, impedance change, voltage change, or capacitance change. Sequencing of nucleic acids by translocation across a membrane and / or through a pore does not foreclose alternative detection methods, such as optical, chemical, biochemical, fluorescent, luminescent, magnetic, electromagnetic, acoustic, or electroacoustic detection.

[0358] It is to be appreciated that the Detailed Description section, and not any other section, is intended to be used to interpret the claims. Other sections may set forth one or more but not all exemplary aspects as contemplated by the inventor(s), and thus, are not intended to limit this disclosure or the appended claims in any way.

[0359] While this disclosure describes exemplary aspects for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other aspects and modifications thereto are possible, and are within the scope and spirit of this disclosure. For example, and without limiting the generality of this paragraph, aspects are not limited to the software, hardware, firmware, and / or entities illustrated in the figures and / or described herein. Further, aspects (whether or not explicitly described herein) have significant utility to fields and applications beyond the examples described herein.

[0360] Aspects have been described herein with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries may be defined as long as the specified functions and relationships (or equivalents thereof) are appropriately performed. Also, alternative aspects may perform functional blocks, steps, operations, methods, etc. using orderings different from those described herein.

[0361] References herein to “one aspect,” “an aspect,” “an example aspect,” “some aspects,” or similar phrases, indicate that the aspect described may include a particular feature, structure, or characteristic, but every aspect may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same aspect. Further, when a particular feature, structure, or characteristic is described in connection with an aspect, it would be within the knowledge of persons skilled in the relevant art(s) to incorporate such feature, structure, or characteristic into other aspects whether or not explicitly mentioned or described herein.

[0362] Additionally, some aspects may be described using the expression “coupled” and “connected” along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some aspects may be described using the terms “connected” and / or “coupled” to indicate that two or more elements are in direct physical or electrical contact with each other. The term “coupled,” however, may also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other.Supports and Low Non-Specific Coatings

[0363] In some aspects, the flow cell 112 in FIG. 1 may include a support, e.g., a solid support as disclosed herein. Aspects of the present disclosure provide pairwise sequencing compositions and methods which employ a support comprising a plurality of oligonucleotide surface primers immobilized thereon. In some aspects, the support is passivated with a low nonspecific binding coating. The surface coatings described herein exhibit very low non-specific binding to reagents typically used for nucleic acid capture, amplification and sequencing workflows, such as dyes, nucleotides, enzymes, and nucleic acid primers. The surface coatings exhibit low background fluorescence signals or high contrast-to-noise (CNR) ratios compared to conventional surface coatings.

[0364] The low non-specific binding coating comprises one layer or multiple layers (e.g., FIG. 23). In some aspects, the plurality of surface primers are immobilized to the low nonspecific binding coating. In some aspects, at least one surface primer is embedded within the low non-specific binding coating. The low non-specific binding coating enables improved nucleic acid hybridization and amplification performance. In general, the supports comprise a substrate (or support structure), one or more layers of a covalently or non-covalently attached low-binding, chemical modification layers, e.g., silane layers, polymer films, and one or more covalently or non-covalently attached surface primers that may be used for tethering single-stranded nucleic acid library molecules to the support. In some aspects, the formulation of the coating, e.g., the chemical composition of one or more layers, the coupling chemistry used to cross-link the one or more layers to the support and / or to each other, and the total number of layers, may be varied such that non-specific binding of proteins, nucleic acid molecules, and other hybridization and amplification reaction components to the coating is minimized or reduced relative to a comparable monolayer. The formulation of the coating described herein may be varied such that non-specific hybridization on the coating is minimized or reduced relative to a comparablemonolayer. The formulation of the coating may be varied such that non-specific amplification on the coating is minimized or reduced relative to a comparable monolayer. The formulation of the coating may be varied such that specific amplification rates and / or yields on the coating are maximized. Amplification levels suitable for detection are achieved in no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more than 30 amplification cycles in some cases disclosed herein.

[0365] The support structure that comprises the one or more chemically-modified layers, e.g., layers of a low non-specific binding polymer, may be independent or integrated into another structure or assembly. For example, in some aspects, the support structure may comprise one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure may comprise one or more surfaces within a microplate format, e.g., the bottom surface of the wells in a microplate. In some aspects, the support structure comprises the interior surface (such as the lumen surface) of a capillary. In some aspects, the support structure comprises the interior surface (such as the lumen surface) of a capillary etched into a planar chip.

[0366] The attachment chemistry used to graft a first chemically-modified layer to the surface of the support will generally be dependent on both the material from which the surface is fabricated and the chemical nature of the layer. In some aspects, the first layer may be covalently attached to the surface. In some aspects, the first layer may be non-covalently attached, e.g., adsorbed to the support through non-covalent interactions such as electrostatic interactions, hydrogen bonding, or van der Waals interactions between the support and the molecular components of the first layer. In either case, the support may be treated prior to attachment or deposition of the first layer. Any of a variety of surface preparation techniques known to those of skill in the art may be used to clean or treat the surface. For example, glass or silicon surfaces may be acid-washed using a Piranha solution (a mixture of sulfuric acid (H2SO4) and hydrogen peroxide (H2O2)), base treatment in KOH and NaOH, and / or cleaned using an oxygen plasma treatment method.

[0367] Silane chemistries constitute non-limiting approaches for covalently modifying the silanol groups on glass or silicon surfaces to attach more reactive functional groups (e.g., amines or carboxyl groups), which may then be used in coupling linker molecules (e.g., linear hydrocarbon molecules of various lengths, such as C6, Cl 2, Cl 8 hydrocarbons, or linear polyethylene glycol (PEG) molecules) or layer molecules (e.g., branched PEG molecules or other polymers) to the surface. Examples of suitable silanes that may be used in creating any of the disclosed low binding coatings include, but are not limited to, (3 -Aminopropyl) trimethoxysilane (APTMS), (3 -Aminopropyl) tri ethoxy silane (APTES), any of a variety of PEG-silanes (e.g.,comprising molecular weights of IK, 2K, 5K, 10K, 20K, etc.), amino-PEG silane (i.e., comprising a free amino functional group), maleimide-PEG silane, biotin-PEG silane, and the like.

[0368] Any of a variety of molecules known to those of skill in the art including, but not limited to, amino acids, peptides, nucleotides, oligonucleotides, other monomers or polymers, or combinations thereof may be used in creating the one or more chemically-modified layers on the support, where the choice of components used may be varied to alter one or more properties of the layers, e.g., the surface density of functional groups and / or tethered oligonucleotide primers, the hydrophilicity / hydrophobicity of the layers, or the three three-dimensional nature (i.e., “thickness”) of the layer. Examples of polymers that may be used to create one or more layers of low non-specific binding material in any of the disclosed coatings include, but are not limited to, polyethylene glycol (PEG) of various molecular weights and branching structures, streptavidin, polyacrylamide, polyester, dextran, poly-lysine, and poly-lysine copolymers, or any combination thereof. Examples of conjugation chemistries that may be used to graft one or more layers of material (e.g. polymer layers) to the surface and / or to cross-link the layers to each other include, but are not limited to, biotin-streptavidin interactions (or variations thereof), his tag - Ni / NTA conjugation chemistries, methoxy ether conjugation chemistries, carboxylate conjugation chemistries, amine conjugation chemistries, NHS esters, maleimides, thiol, epoxy, azide, hydrazide, alkyne, isocyanate, and silane.

[0369] The low non-specific binding surface coating may be applied uniformly across the support. Alternatively, the surface coating may be patterned, such that the chemical modification layers are confined to one or more discrete regions of the support. For example, the coating may be patterned using photolithographic techniques to create an ordered array or random pattern of chemically-modified regions on the support. Alternately or in combination, the coating may be patterned using, e.g., contact printing and / or ink-jet printing techniques. In some aspects, an ordered array or random pattern of chemically-modified regions may comprise at least 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more discrete regions.

[0370] In some aspects, the low nonspecific binding coatings comprise hydrophilic polymers that are non-specifically adsorbed or covalently grafted to the support. Typically, passivation is performed utilizing polyethylene glycol) (PEG, also known as polyethylene oxide (PEO) or polyoxyethylene) or other hydrophilic polymers with different molecular weights and end groups that are linked to a support using, for example, silane chemistry. The end groups distal from thesurface may include, but are not limited to, biotin, methoxy ether, carboxylate, amine, NHS ester, maleimide, and bis-silane. In some aspects, two or more layers of a hydrophilic polymer, e.g., a linear polymer, branched polymer, or multi-branched polymer, may be deposited on the surface. In some aspects, two or more layers may be covalently coupled to each other or internally crosslinked to improve the stability of the resulting coating. In some aspects, surface primers with different nucleotide sequences and / or base modifications (or other biomolecules, e.g., enzymes or antibodies) may be tethered to the resulting layer at various surface densities. In some aspects, for example, both surface functional group density and surface primer concentration may be varied to attain a desired surface primer density range. Additionally, surface primer density may be controlled by diluting the surface primers with other molecules that carry the same functional group. For example, amine-labeled surface primers may be diluted with amine-labeled polyethylene glycol in a reaction with an NHS-ester coated surface to reduce the final primer density. Surface primers with different lengths of linker between the hybridization region and the surface attachment functional group may also be applied to control surface density. Example of suitable linkers include poly-T and poly-A strands at the 5’ end of the primer (e.g., 0 to 20 bases), PEG linkers (e.g., 3 to 20 monomer units), and carbon-chain (e.g., C6, C12, C18, etc.). To measure the primer density, fluorescently-labeled primers may be tethered to the surface and a fluorescence reading then compared with that for a dye solution of known concentration.

[0371] In some aspects, the low nonspecific binding coatings comprise a functionalized polymer coating layer covalently bound at least to a portion of the support via a chemical group on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some aspects, the functionalized polymer coating comprises a poly(N-(5-azidoacetamidylpentyl)acrylamide-co- acrylamide (PAZAM).

[0372] In order to scale primer surface density and add additional dimensionality to hydrophilic or amphoteric coatings, supports comprising multi-layer coatings of PEG and other hydrophilic polymers have been developed. By using hydrophilic and amphoteric surface layering approaches that include, but are not limited to, the polymer / co-polymer materials described below, it is possible to increase primer loading density on the support significantly. Traditional PEG coating approaches use monolayer primer deposition, which have been generally reported for single molecule applications, but do not yield high copy numbers for nucleic acid amplification applications. As described herein “layering” may be accomplished using traditional crosslinking approaches with any compatible polymer or monomer subunitssuch that a surface comprising two or more highly cross-linked layers may be built sequentially. Examples of suitable polymers include, but are not limited to, streptavidin, poly acrylamide, polyester, dextran, poly-lysine, and copolymers of poly-lysine and PEG. In some aspects, the different layers may be attached to each other through any of a variety of conjugation reactions including, but not limited to, biotin-streptavidin binding, azide-alkyne click reaction, amine-NHS ester reaction, thiol-maleimide reaction, and ionic interactions between positively charged polymer and negatively charged polymer. In some aspects, high primer density materials may be constructed in solution and subsequently layered onto the surface in multiple steps.

[0373] Examples of materials from which the support structure may be fabricated include, but are not limited to, glass, fused-silica, silicon, a polymer (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethylmethacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymers (COP), cyclic olefin copolymers (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic support structures are contemplated.

[0374] The support structure may be rendered in any of a variety of geometries and dimensions known to those of skill in the art, and may comprise any of a variety of materials known to those of skill in the art. For example, the support structure may be locally planar (e.g., comprising a microscope slide or the surface of a microscope slide). Globally, the support structure may be cylindrical (e.g., comprising a capillary or the interior surface of a capillary), spherical (e.g., comprising the outer surface of a non-porous bead), or irregular (e.g., comprising the outer surface of an irregularly-shaped, non-porous bead or particle). In some aspects, the surface of the support structure used for nucleic acid hybridization and amplification may be a solid, non-porous surface. In some aspects, the surface of the support structure used for nucleic acid hybridization and amplification may be porous, such that the coatings described herein penetrate the porous surface, and nucleic acid hybridization and amplification reactions performed thereon may occur within the pores.

[0375] The support structure that comprises the one or more chemically-modified layers, e.g., layers of a low non-specific binding polymer, may be independent or integrated into another structure or assembly. For example, the support structure may comprise one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure may comprise one or more surfaces within a microplate format, e.g., the bottom surface of the wells in a microplate. In some aspects, the support structure comprises the interior surface (such as the lumen surface)of a capillary. In some aspects the support structure comprises the interior surface (such as the lumen surface) of a capillary etched into a planar chip.

[0376] As noted, the low non-specific binding supports of the present disclosure exhibit reduced non-specific binding of proteins, nucleic acids, and other components of the hybridization and / or amplification formulation used for solid-phase nucleic acid amplification. The degree of non-specific binding exhibited by a given support surface may be assessed either qualitatively or quantitatively. For example, exposure of the surface to fluorescent dyes (e.g., cyanins such as Cy3, or Cy5, etc., fluoresceins, coumarins, rhodamines, etc. or other dyes disclosed herein), fluorescently-labeled nucleotides, fluorescently-labeled oligonucleotides, and / or fluorescently-labeled proteins (e.g. polymerases) under a standardized set of conditions, followed by a specified rinse protocol and fluorescence imaging may be used as a qualitative tool for comparison of non-specific binding on supports comprising different surface formulations. In some aspects, exposure of the surface to fluorescent dyes, fluorescently-labeled nucleotides, fluorescently-labeled oligonucleotides, and / or fluorescently-labeled proteins (e.g. polymerases) under a standardized set of conditions, followed by a specified rinse protocol and fluorescence imaging may be used as a quantitative tool for comparison of non-specific binding on supports comprising different surface formulations — provided that care has been taken to ensure that the fluorescence imaging is performed under conditions where fluorescence signal is linearly related (or related in a predictable manner) to the number of fluorophores on the support surface (e.g., under conditions where signal saturation and / or self-quenching of the fluorophore is not an issue) and suitable calibration standards are used. In some aspects, other techniques known to those of skill in the art, for example, radioisotope labeling and counting methods may be used for quantitative assessment of the degree to which non-specific binding is exhibited by the different support surface formulations of the present disclosure.

[0377] Some surfaces disclosed herein exhibit a ratio of specific to nonspecific binding of a fluorophore such as Cy3 of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value spanned by the range herein. Some surfaces disclosed herein exhibit a ratio of specific to nonspecific fluorescence of a fluorophore such as Cy3 of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value spanned by the range herein.

[0378] The degree of non-specific binding exhibited by the disclosed low-binding supports may be assessed using a standardized protocol for contacting the surface with a labeled protein(e.g., bovine serum albumin (BSA), streptavidin, a DNA polymerase, a reverse transcriptase, a helicase, a single-stranded binding protein (SSB), etc., or any combination thereof), a labeled nucleotide, a labeled oligonucleotide, etc., under a standardized set of incubation and rinse conditions, followed be detection of the amount of label remaining on the surface and comparison of the signal resulting therefrom to an appropriate calibration standard. In some aspects, the label may comprise a fluorescent label. In some aspects, the label may comprise a radioisotope. In some aspects, the label may comprise any other detectable label known to one of skill in the art. In some aspects, the degree of non-specific binding exhibited by a given support surface formulation may thus be assessed in terms of the number of non-specifically bound protein molecules (or nucleic acid molecules or other molecules) per unit area. In some aspects, the low-binding supports of the present disclosure may exhibit non-specific protein binding (or non-specific binding of other specified molecules, (e.g., cyanins such as Cy3, or Cy5, etc., fluoresceins, coumarins, rhodamines, etc. or other dyes disclosed herein)) of less than 0.001 molecule per pm2, less than 0.01 molecule per pm2, less than 0.1 molecule per pm2, less than 0.25 molecule per pm2, less than 0.5 molecule per pm2, less than 1 molecule per pm2, less than 10 molecules per pm2, less than 100 molecules per pm2, or less than 1,000 molecules per pm2. Those of skill in the art will realize that a given support surface of the present disclosure may exhibit non-specific binding falling anywhere within this range, for example, of less than 86 molecules per pm2. For example, some modified surfaces disclosed herein exhibit nonspecific protein binding of less than 0.5 molecule / pm2following contact with a 1 pM solution of Cy3 labeled streptavidin (GE Amersham) in phosphate buffered saline (PBS) buffer for 15 minutes, followed by 3 rinses with deionized water. Some modified surfaces disclosed herein exhibit nonspecific binding of Cy3 dye molecules of less than 0.25 molecules per pm2. In independent nonspecific binding assays, 1 pM labeled Cy3 SA (ThermoFisher), 1 pM Cy5 SA dye (ThermoFisher), 10 pM Aminoallyl-dUTP-ATTO-647N (Jena Biosciences), 10 pM Aminoallyl- dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 pM Aminoallyl-dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 pM 7-Propargylamino-7-deaza-dGTP-Cy5 (Jena Biosciences, and 10 pM 7- Propargylamino-7-deaza-dGTP-Cy3 (Jena Biosciences) were incubated on the low binding coated supports at 37° C. for 15 minutes in a 384 well plate format. Each well was rinsed 2-3 x with 50 ul deionized RNase / DNase Free water and 2-3 x with 25 mM ACES buffer pH 7.4. The 384 well plates were imaged on a GE Typhoon instrument using the Cy3, AF555, or Cy5 filter sets (according to dye test performed) as specified by the manufacturer at a PMT gain setting of 800 and resolution of 50-100 pm. For higher resolution imaging, images were collected on anOlympus 1X83 microscope (e.g., inverted fluorescence microscope) (Olympus Corp., Center Valley, Pa.) with a total internal reflectance fluorescence (TIRF) objective (100x, 1.5 NA, Olympus), a CCD camera (e.g., an Olympus EM-CCD monochrome camera, Olympus XM-10 monochrome camera, or an Olympus DP80 color and monochrome camera), an illumination source (e.g., an Olympus 100W Hg lamp, an Olympus 75 W Xe lamp, or an Olympus U- HGLGPS fluorescence light source), and excitation wavelengths of 532 nm or 635 nm. Dichroic mirrors were purchased from Semrock (IDEX Health & Science, LLC, Rochester, N. Y.), e.g., 405, 488, 532, or 633 nm dichroic refl ectors / b earn splitters, and band pass filters were chosen as 532 LP or 645 LP concordant with the appropriate excitation wavelength. Some modified surfaces disclosed herein exhibit nonspecific binding of dye molecules of less than 0.25 molecules per pm2. In some aspects, the coated support was immersed in a buffer (e.g., 25 mM ACES, pH 7.4) while the image was acquired.

[0379] In some aspects, the surfaces disclosed herein exhibit a ratio of specific to nonspecific binding of a fluorophore such as Cy3 of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value spanned by the range herein. In some aspects, the surfaces disclosed herein exhibit a ratio of specific to nonspecific fluorescence signals for a fluorophore such as Cy3 of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value spanned by the range herein.

[0380] The low-background surfaces consistent with the disclosure herein may exhibit specific dye attachment (e.g., Cy3 attachment) to non-specific dye adsorption (e.g., Cy3 dye adsorption) ratios of at least 4: 1, 5: 1, 6: 1, 7:1, 8: 1, 9: 1, 10: 1, 15: 1, 20: 1, 30: 1, 40: 1, 50: 1, or more than 50 specific dye molecules attached per molecule nonspecifically adsorbed. Similarly, when subjected to an excitation energy, low-background surfaces consistent with the disclosure herein to which fluorophores, e.g., Cy3, have been attached may exhibit ratios of specific fluorescence signal (e.g., arising from Cy3 -labeled oligonucleotides attached to the surface) to non-specific adsorbed dye fluorescence signals of at least 4: 1, 5: 1, 6: 1, 7: 1, 8: 1, 9: 1, 10: 1, 15:1, 20: 1, 30: 1, 40: 1, 50: 1, or more than 50: 1.

[0381] In some aspects, the degree of hydrophilicity (or “wettability” with aqueous solutions) of the disclosed support surfaces may be assessed, for example, through the measurement of water contact angles in which a small droplet of water is placed on the surface and its angle of contact with the surface is measured using, e.g., an optical tensiometer. In some aspects, a static contact angle may be determined. In some aspects, an advancing or receding contact angle maybe determined. In some aspects, the water contact angle for the hydrophilic, low-binding support surfaced disclosed herein may range from about 0 degrees to about 30 degrees. In some aspects, the water contact angle for the hydrophilic, low-binding support surfaced disclosed herein may no more than 50 degrees, 40 degrees, 30 degrees, 25 degrees, 20 degrees, 18 degrees, 16 degrees, 14 degrees, 12 degrees, 10 degrees, 8 degrees, 6 degrees, 4 degrees, 2 degrees, or 1 degree. In many cases the contact angle is no more than 40 degrees. Those of skill in the art will realize that a given hydrophilic, low-binding support surface of the present disclosure may exhibit a water contact angle having a value of anywhere within this range.

[0382] In some aspects, the hydrophilic surfaces disclosed herein facilitate reduced wash times for bioassays, often due to reduced nonspecific binding of biomolecules to the low-binding surfaces. In some aspects, adequate wash steps may be performed in less than 60, 50, 40, 30, 20, 15, 10, or less than 10 seconds. For example, adequate wash steps may be performed in less than 30 seconds.

[0383] Some low-binding surfaces of the present disclosure exhibit significant improvement in stability or durability to prolonged exposure to solvents and elevated temperatures, or to repeated cycles of solvent exposure or changes in temperature. For example, the stability of the disclosed surfaces may be tested by fluorescently labeling a functional group on the surface, or a tethered biomolecule (e.g., an oligonucleotide primer) on the surface, and monitoring fluorescence signal before, during, and after prolonged exposure to solvents and elevated temperatures, or to repeated cycles of solvent exposure or changes in temperature. In some aspects, the degree of change in the fluorescence used to assess the quality of the surface may be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% over a time period of 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 40 minutes, 50 minutes, 60 minutes, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 15 hours, 20 hours, 25 hours, 30 hours, 35 hours, 40 hours, 45 hours, 50 hours, or 100 hours of exposure to solvents and / or elevated temperatures (or any combination of these percentages as measured over these time periods). In some aspects, the degree of change in the fluorescence used to assess the quality of the surface may be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% over 5 cycles, 10 cycles, 20 cycles, 30 cycles, 40 cycles, 50 cycles, 60 cycles, 70 cycles, 80 cycles, 90 cycles, 100 cycles, 200 cycles, 300 cycles, 400 cycles, 500 cycles, 600 cycles, 700 cycles, 800 cycles, 900 cycles, or 1,000 cycles of repeated exposure to solvent changes and / or changes in temperature (or any combination of these percentages as measured over this range of cycles).

[0384] In some aspects, the surfaces disclosed herein may exhibit a high ratio of specific signal to nonspecific signal or other background. For example, when used for nucleic acid amplification, some surfaces may exhibit an amplification signal that is at least 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 75, 100, or greater than 100 fold greater than a signal of an adjacent unpopulated region of the surface. Similarly, some surfaces exhibit an amplification signal that is at least 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 75, 100, or greater than 100 fold greater than a signal of an adjacent amplified nucleic acid population region of the surface.

[0385] In some aspects, fluorescence images of the disclosed low background surfaces when used in nucleic acid hybridization or amplification applications to create polonies of hybridized or clonally-amplified nucleic acid molecules (e.g., that have been directly or indirectly labeled with a fluorophore) exhibit contrast-to-noise ratios (CNRs) of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 210, 220, 230, 240, 250, or greater than 250.

[0386] One or more types of primer may be attached or tethered to the support surface. In some aspects, the one or more types of adapters or primers may comprise spacer sequences, adapter sequences for hybridization to adapter-ligated target library nucleic acid sequences, forward amplification primers, reverse amplification primers, sequencing primers, and / or molecular barcoding sequences, or any combination thereof. In some aspects, 1 primer or adapter sequence may be tethered to at least one layer of the surface. In some aspects, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 different primer or adapter sequences may be tethered to at least one layer of the surface.

[0387] In some aspects, the tethered adapter and / or primer sequences may range in length from about 10 nucleotides to about 100 nucleotides. In some aspects, the tethered adapter and / or primer sequences may be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 nucleotides in length. In some aspects, the tethered adapter and / or primer sequences may be at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, or at most 10 nucleotides in length. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure, for example, in some aspects the length of the tethered adapter and / or primer sequences may range from about 20 nucleotides to about 80 nucleotides. Those of skill in the art will recognize that the length of the tethered adapter and / or primer sequences may have any value within this range, e.g., about 24 nucleotides.

[0388] In some aspects, the resultant surface density of primers (e.g., capture primers) on the low binding support surfaces of the present disclosure may range from about 100 primer molecules per pm2to about 100,000 primer molecules per pm2. In some aspects, the resultant surface density of primers on the low binding support surfaces of the present disclosure may range from about 1,000 primer molecules per pm2to about 1,000,000 primer molecules per pm2. In some aspects, the surface density of primers may be at least 1,000, at least 10,000, at least 100,000, or at least 1,000,000 molecules per pm2. In some aspects, the surface density of primers may be at most 1,000,000, at most 100,000, at most 10,000, or at most 1,000 molecules per pm2. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure, for example, in some aspects the surface density of primers may range from about 10,000 molecules per pm2to about 100,000 molecules per pm2. Those of skill in the art will recognize that the surface density of primer molecules may have any value within this range, e.g., about 455,000 molecules per pm2. In some aspects, the surface density of target library nucleic acid sequences initially hybridized to adapter or primer sequences on the support surface may be less than or equal to that indicated for the surface density of tethered primers. In some aspects, the surface density of clonally-amplified target library nucleic acid sequences hybridized to adapter or primer sequences on the support surface may span the same range as that indicated for the surface density of tethered primers.

[0389] Local densities as listed above do not preclude variation in density across a surface, such that a surface may comprise a region having an oligo density of, for example, 500,000 / pm2, while also comprising at least a second region having a substantially different local density.

[0390] In some aspects, the performance of nucleic acid hybridization and / or amplification reactions using the disclosed reaction formulations and low-binding supports may be assessed using fluorescence imaging techniques, where the contrast-to-noise ratio (CNR) of the images provides a key metric in assessing amplification specificity and non-specific binding on the support. CNR is commonly defined as: CNR=(Signal-Background) / Noise. The background term is commonly taken to be the signal measured for the interstitial regions surrounding a particular feature (diffraction limited spot, DLS) in a specified region of interest (ROI). While signal-to- noise ratio (SNR) is often considered to be a benchmark of overall signal quality, it may be shown that improved CNR may provide a significant advantage over SNR as a benchmark for signal quality in applications that require rapid image capture (e.g., sequencing applications for which cycle times must be minimized), as shown in the example below. At high CNR the imaging time required to reach accurate discrimination (and thus accurate base-calling in the caseof sequencing applications) may be drastically reduced even with moderate improvements in CNR. Improved CNR in imaging data on the imaging integration time provides a method for more accurately detecting features such as clonally-amplified nucleic acid colonies on the support surface.

[0391] In most ensemble-based sequencing approaches, the background term is typically measured as the signal associated with 'interstitial' regions. In addition to "interstitial" background (Binter), "intrastitial" background (Bintra) exists within the region occupied by an amplified DNA colony. The combination of these two background signals dictates the achievable CNR, and subsequently directly impacts the optical instrument requirements, architecture costs, reagent costs, run-times, cost / genome, and ultimately the accuracy and data quality for cyclic array -based sequencing applications. The Binter background signal arises from a variety of sources; a few examples include auto-fluorescence from consumable flow cells, non-specific adsorption of detection molecules that yield spurious fluorescence signals that may obscure the signal from the ROI, the presence of non-specific DNA amplification products (e.g., those arising from primer dimers). In typical next generation sequencing (NGS) applications, this background signal in the current field-of-view (FOV) is averaged over time and subtracted. The signal arising from individual DNA colonies (i.e., (Signal)-B(interstial) in the FOV) yields a discernable feature that may be classified. In some aspects, the intrastitial background (B(intrastitial)) may contribute a confounding fluorescence signal that is not specific to the target of interest, but is present in the same ROI thus making it far more difficult to average and subtract.

[0392] Nucleic acid amplification on the low-binding coated supports described herein may decrease the B(interstitial) background signal by reducing non-specific binding, may lead to improvements in specific nucleic acid amplification, and may lead to a decrease in non-specific amplification that may impact the background signal arising from both the interstitial and intrastitial regions. In some aspects, the disclosed low-binding coated supports, optionally used in combination with the disclosed hybridization and / or amplification reaction formulations, may lead to improvements in CNR by a factor of 2, 5, 10, 100, 250, 500 or 1000-fold over those achieved using conventional supports and hybridization, amplification, and / or sequencing protocols. Although described here in the context of using fluorescence imaging as the read-out or detection mode, the same principles apply to the use of the disclosed low-binding coated supports and nucleic acid hybridization and amplification formulations for other detection modes as well, including both optical and non-optical detection modes.

[0393] The headings provided herein are not limitations of the various aspects of the disclosure, which aspects may be understood by reference to the specification as a whole.

[0394] Unless defined otherwise, technical and scientific terms used herein have meanings that are commonly understood by those of ordinary skill in the art unless defined otherwise. Generally, terminologies pertaining to techniques of molecular biology, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production, and hybridization described herein are those well-known and commonly used in the art. Techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the instant specification. For example, see Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. 2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclatures utilized in connection with, and the laboratory procedures and techniques described herein are those well-known and commonly used in the art.

[0395] Unless otherwise required by context herein, singular terms shall include pluralities and plural terms shall include the singular. Singular forms “a”, “an” and “the”, and singular use of any word, include plural referents unless expressly and unequivocally limited on one referent.

[0396] It is understood the use of the alternative term (e.g., “or”) is taken to mean either one or both or any combination thereof of the alternatives.

[0397] The term “and / or” used herein is to be taken mean specific disclosure of each of the specified features or components with or without the other. For example, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include: “A and B”; “A or B”; “A” (A alone); and “B” (B alone). In a similar manner, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to encompass each of the following aspects: ...

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method for predicting quality of base calling in DNA sequencing, comprising: selecting, by a computer, a look-up table from a plurality of look-up tables, wherein each of the plurality of look-up tables is generated by training using a corresponding training dataset; selecting, by the computer, a plurality of predictors based on the selected look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a first quality score from the selected look-up table based on the corresponding value for each of the plurality of predictors, wherein the selected look-up table comprises a plurality of dimensions corresponding to the plurality of predictors.

2. A computer-implemented method for predicting quality of base calling in DNA sequencing, comprising: selecting, by a computer, a look-up table from a plurality of look-up tables, wherein each of the plurality of look-up tables is generated by training using a corresponding training dataset; selecting, by the computer, a plurality of predictors based on the selected look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a first quality score from the selected look-up table based on the corresponding value for each of the plurality of predictors, wherein the selected look-up table comprises a plurality of dimensions corresponding to the plurality of predictors; improving, by the computer, a baseline quality score by replacing the baseline quality score with the first quality score, wherein the baseline quality score is estimatedbased on a full look-up table generated by training using the corresponding datasets of the plurality of look-up tables.

3. The computer-implemented method of claim 1, wherein the method further comprising: determining, by the computer, at least one sequencing cycle number of the one or more flow cell images.

4. The computer-implemented method of claim 1, wherein the method further comprising: determining, by the computer, whether a likelihood of at least one sequencing cycle being affected by sequencing errors is over a predetermined threshold; and selecting, by a computer, a look-up table from a plurality of look-up tables based on the determination.

5. The computer-implemented method of any one of the preceding claims, wherein selecting, by the computer, the look-up table from the plurality of look-up tables is based on that the likelihood of at least one sequencing cycle being affected by sequencing errors is over a predetermined threshold.

6. The computer-implemented method of claim 1, wherein the one or more are acquired by a sequencing system during a sequencing run, and wherein the method further comprising: determining, by the computer, a total number of sequencing cycles of the sequencing run.

7. The computer-implemented method of any one of the preceding claims, wherein selecting, by the computer, the look-up table from the plurality of look-up tables is based on the at least one sequencing cycle number of the one or more flow cell images.

8. The computer-implemented method of any one of the preceding claims, wherein selecting, by the computer, the look-up table from the plurality of look-up tables is based on the at least one sequencing cycle number of the one or more flow cell images and the total number of sequencing cycles.

9. The computer-implemented method of any one of the preceding claims, wherein the corresponding training dataset is different for each of the plurality of look-up tables.

10. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: determining, by the computer, one or more base calls based on the one or more flow cell images.

11. The computer-implemented method of claim 9, wherein determining, by the computer, the one or more base calls based on the one or more flow cell images is after determining the first quality score.

12. The computer-implemented method of any one of the preceding claims further comprising: in response to determining that the first quality score is below a quality threshold, terminating the sequencing run before completing all sequencing cycles of the sequencing run.

13. The computer-implemented method of claim 12, wherein the baseline quality score is above the quality threshold.

14. The computer-implemented method of any one of the preceding claims, wherein the first quality score or the baseline quality score is for a pixel or a subpixel of the one or more flow cell images.

15. The computer-implemented method of any one of the preceding claims, wherein the first quality score or the baseline quality score is for a polony or a cluster of the one or more flow cell images, the polony or cluster comprising one or more pixels or subpixels.

16. The computer-implemented method of any one of the preceding claims, wherein the at least one sequencing cycle number is less than 5, 10, 15, 20, 25, or 30, and wherein the determined one or more base calls comprises at least an error caused by library preparation of the samples to be sequenced.

17. The computer-implemented method of any one of the preceding claims, wherein the at least one sequencing cycle number is greater than 5, 10, 15, 20, 25, or 30, and wherein the one or more base calls lack an error caused by library preparation of the samples to be sequenced.

18. The computer-implemented method of any one of the preceding claims, wherein the plurality of look-up tables comprises at least three look-up tables.

19. The computer-implemented method of any one of the preceding claims, wherein the plurality of look-up tables comprises the first look-up table, a second look-up table, and a third look-up table.

20. The computer-implemented method of any one of the preceding claims, wherein the first look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a first sequencing cycle.

21. The computer-implemented method of any one of the preceding claims, wherein the second look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a second sequencing cycle and one more sequencing cycles subsequent thereto.

22. The computer-implemented method of any one of the preceding claims, wherein the third look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least a third sequencing cycle and one more sequencing cycles subsequent thereto.

23. The computer-implemented method of any one of the preceding claims, wherein the full look-up table is generated by training using the corresponding training dataset comprising the training flow cell images from at least the first, second, third sequencing cycles and sequencing cycles subsequent thereto.

24. The computer-implemented method of any one of the preceding claims, wherein selecting, by the computer, the plurality of predictors based on the selected look-up table comprises: in response to determining that the selected look-up table is generated based on at least a sequencing cycle number less than a predetermined cycle number, including at least a predictor that is based on phasing or prephasing of the one or more flow cell images; and in response to determining that the selected look-up table is generated based on at least a sequencing cycle number greater than a predetermined cycle number, excluding any predictor of the plurality of predictors that is based on phasing or prephasing of the one or more flow cell images.

25. The computer-implemented method of any one of the preceding claims, wherein the corresponding training dataset only comprises training flow cell images from a subset of sequencing cycles of all sequencing cycles in a sequencing run.

26. The computer-implemented method of any one of the preceding claims, wherein at least one of the corresponding training datasets only comprises only training flow cell images from a single sequencing cycle in a sequencing run.

27. The computer-implemented method of any one of the preceding claims, wherein only one of the corresponding training datasets only comprises only training flow cell images from a single sequencing cycle in a sequencing run.

28. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: selecting, by the computer, the full look-up table; selecting, by the computer, a plurality of predictors based on the full look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a baseline quality score from the full look-up table based on the corresponding value for each of the plurality of predictors, wherein the fulllook-up table comprises a plurality of dimensions corresponding to the plurality of predictors.

29. The computer-implemented method of any one of the preceding claims, wherein the at least one sequencing cycle number is less than 5, 10, 12, 15, or 20, and the first quality score is lower than the baseline quality score.

30. The computer-implemented method of any one of the preceding claims, wherein the at least one sequencing cycle is greater than 5, 10, 12, 15, or 20, and the first quality score is higher than the baseline quality score.

31. The computer-implemented method of any one of the preceding claims, wherein the first quality score is greater than Q35, Q40, Q45, Q50, Q55, or Q60.

32. The computer-implemented method of any one of the preceding claims, wherein the baseline quality score is greater than Q35, Q40, Q45, Q50, Q55, or Q60.

33. The computer-implemented method of any one of the preceding claims, wherein one or more of the plurality of predictors are indicative of an image quality of the one or more flow cell images.

34. The computer-implemented method of any one of the preceding claims further comprising: performing one or more preprocessing steps to generate the one or more flow cell images from multiple channels, the one or more preprocessing steps comprising:(1) background subtraction;(2) color correction;(3) phasing or prephasing correction; and(4) and normalization.

35. The computer-implemented method of any one of the preceding claims, wherein (4) normalization comprises normalization of intensities of flow cell images across the multiple channels.

36. The computer-implemented method of any one of the preceding claims, wherein (4) normalization comprises normalization of intensities of flow cell images across all channels.

37. The computer-implemented method of any one of the preceding claims, wherein (4) normalization comprises normalization of intensities of flow cell images from a corresponding channel by a predetermined intensity.

38. The computer-implemented method of any one of the preceding claims, wherein one channel corresponds to a wavelength spectrum of a fluorescent element representing a type of a nucleotide base in a set of nucleotide bases.

39. The computer-implemented method of any one of the preceding claims further comprising: ranking, by a computer, candidate predictors based on corresponding effects thereof on quality scores of base calling, and wherein selecting, by the computer, the plurality of predictors is based on the ranking of the candidate predictors;40. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises at least three predictors.

41. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises: a first predictor of clarity, a second predictor of max intensity, and a third predictor of low intensity median clarity.

42. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises four predictors.

43. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises a first predictor of clarity, a second predictor of max intensity, and a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

44. The computer-implemented method of any one of the preceding claims, wherein the first predictor of clarity comprises a ratio of a max intensity to a second max intensity of a polony in a selected imaging cycle.

45. The computer-implemented method of any one of the preceding claims, wherein the second predictor of max intensity comprises an intensity from a brightest channel among multiple channels of a polony in a selected imaging cycle.

46. The computer-implemented method of any one of the preceding claims, wherein the intensity is normalized within the brightest channel.

47. The computer-implemented method of any one of the preceding claims, wherein the max intensity comprises a normalized intensity of a brightest channel among four channels.

48. The computer-implemented method of any one of the preceding claims, wherein the second max intensity comprises an intensity from a second brightest channel among different channels of a polony in a selected imaging cycle.

49. The computer-implemented method of any one of the preceding claims, wherein the second max intensity comprises an intensity from a second brightest channel among 4 channels of a polony in a selected imaging cycle.

50. The computer-implemented method of any one of the preceding claims, wherein the intensity is normalized within the second brightest channel.

51. The computer-implemented method of any one of the preceding claims, wherein the third predictor of low intensity median clarity is indicative of signal density of polonies within a region in a selected imaging cycle in the one or more flow cell images.

52. The computer-implemented method of any one of the preceding claims, wherein the third predictor of low intensity median clarity is inversely proportional to the signal intensity in the one or more flow cell images.

53. The computer-implemented method of any one of the preceding claims, wherein the third predictor of low intensity median clarity comprises a median clarity of the signal intensity of dim polonies in a region of a selected imaging cycle of the one or more flow cell images.

54. The computer-implemented method of any one of the preceding claims, wherein the dim polonies comprises polonies that are about 25% or less of the brightest intensity in the one or more flow cell images.

55. The computer-implemented method of any one of the preceding claims, wherein the dim polonies comprises polonies whose is within about 10% darkest population of polonies of the one or more flow cell images.

56. The computer-implemented method of any one of the preceding claims, wherein the third predictor of low intensity median clarity is a median value of an inverse of clarity of selected dim polonies.

57. The computer-implemented method of any one of the preceding claims, wherein the corresponding value of the second predictor of max intensity is based on the polony of signals in the one or more flow cell images, the one or more flow cell images being within an imaging cycle .

58. The computer-implemented method of any one of the preceding claims, wherein the corresponding value of the third predictor of low intensity median clarity is based on a region with multiple polonies in the one or more flow cell images, the one or more flow cell images being within an imaging cycle.

59. The computer-implemented method of any one of the preceding claims, wherein the corresponding value of the fourth predictor of phasing and prephasing is based on multiple polonies in the one or more flow cell images, the one or more flow cell images being within an imaging cycle.

60. The computer-implemented method of any one of the preceding claims, wherein the computer comprises a general-purpose computer.

61. The computer-implemented method of any one of the preceding claims, wherein the computer comprises one or more central processing units (CPUs).

62. The computer-implemented method of any one of the preceding claims, wherein the one or more dedicated processors comprises field-programmable gate arrays (FPGAs).

63. The computer-implemented method of any one of the preceding claims, wherein the computer comprises FPGAs or artificial intelligence (Al) chips.

64. The computer-implemented method of any one of the preceding claims, wherein the samples comprise in situ samples of cell, tissue, organoid, or their combinations.

65. The computer-implemented method of any one of the preceding claims, wherein a field of view of the one or more flow cell images is greater than 1mm2, 4mm2, 6 mm2, or 10mm2.

66. The computer-implemented method of any one of the preceding claims, wherein the one or more flow cell images comprises multiple flow cell images acquired at different z- locations of the samples.

67. The computer-implemented method of any one of the preceding claims further comprising: acquiring, by an imager of a sequencing system, the one or more flow cell images of the sample immobilized on a flow cell device; and obtaining, by the computer and from the imager, the one or more flow cell images.

68. A computer-implemented method for predicting quality of base calling in DNA sequencing, comprising: generating a first look-up table for predicting quality of base calling in DNA sequencing, comprising: obtaining, by a computer, first training base calls from one or more first training flow cell images;selecting, by the computer, a first plurality of predictors based on at least a first sequencing cycle number of the one or more first training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the first training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of the first plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a first look-up table having a number of dimensions determined by the first plurality of predictors and each dimension comprising the corresponding number of bins; determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the first look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps:(e) calculating a standard quality score and a conservative quality score for each bin of the first look-up table;(f) determining coordinates of a bin with a maximum conservative quality score in the first look-up table;(g) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the first look-up table; and(h) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

69. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: generating a second look-up table for predicting quality of base calling in DNA sequencing, comprising: obtaining, by a computer, second training base calls from one or more second training flow cell images;selecting, by the computer, a second plurality of predictors based on at least a second sequencing cycle number of the one or more second training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the second training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of a second plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a second look-up table having a number of dimensions determined by the second plurality of predictors and each dimension comprising the corresponding number of bins; determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the second look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps:(e) calculating a standard quality score and a conservative quality score for each bin of the second look-up table;(f) determining coordinates of a bin with a maximum conservative quality score in the second look-up table;(g) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the second look-up table; and(h) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

70. The computer-implemented method of any one of the preceding claims, wherein the second sequencing cycle number is greater than the first sequencing cycle number.

71. The computer-implemented method of any one of the preceding claims, wherein the second sequencing cycle number is greater than 10, 12, 15, 20, 25, or 30.

72. The computer-implemented method of any one of the preceding claims, wherein the first sequencing cycle number is less than 5, 10, 12, 15, or 20.

73. The computer-implemented method of any one of the preceding claims, wherein the second plurality of predictors is different from the first plurality of predictors.

74. The computer-implemented method of any one of the preceding claims, wherein the second plurality of predictors comprises more predictors than the first plurality of predictors.

75. The computer-implemented method of any one of the preceding claims, wherein the second plurality of predictors comprises at least one predictor based on phasing or prephasing.

76. The computer-implemented method of any one of the preceding claims, wherein the first plurality of predictors lacks any predictor based on phasing or pre-phasing.

77. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: acquiring, by an optical system, the one or more first training flow cell images in at least the first sequencing cycle number.

78. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: acquiring, by an optical system, the one or more first training flow cell images in only the first sequencing cycle number.

79. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: acquiring, by an optical system, the one or more second training flow cell images in the second sequencing cycle number and a plurality of cycles subsequent to the second sequencing cycle number.

80. The computer-implemented method of any one of the preceding claims, wherein the method further comprising: generating a third look-up table for predicting quality of base calling in DNA sequencing.

81. The computer-implemented method of any one of the preceding claims, wherein the one or more first training flow cell images comprise errors caused by library preparation.

82. The computer-implemented method of any one of the preceding claims, wherein the one or more second training flow cell images comprise errors caused by library preparation.

83. The computer-implemented method of any one of the preceding claims, wherein the reference base calls are obtained from a pre-selected genome with known bases.

84. The computer-implemented method of claim 68, further comprising: filtering, by the computer, the reference base calls by removing base calls from: a first set of positions in the pre-selected genome with pre-determined variants; a second set of positions in the pre-selected genome with pre-determined alignment errors, or both.

85. The computer-implemented method of claim 68, further comprising: determining, by the computer, a first cumulative number of correct base calls and a second cumulative number of erroneous base calls in each bin of the look-up table.

86. The computer-implemented method of any one of the preceding claims, wherein the one or more flow cell images are acquired from multiple channels.

87. The computer-implemented method of any one of the preceding claims, wherein the one or more flow cell images are acquired from all 4 channels.

88. The computer-implemented method of claim 68, further comprising: performing one or more preprocessing steps to generate the one or more flow cell images, the one or more preprocessing steps comprising:(1) background subtraction;(2) color correction;(3) phasing or prephasing correction; and(4) and normalization.

89. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises three predictors.

90. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises: a first predictor of clarity, a second predictor of max intensity, and a third predictor of low intensity median clarity.

91. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises four predictors.

92. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises a first predictor of clarity, a second predictor of max intensity, and a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

93. The computer-implemented method of any one of the preceding claims, wherein the predetermined number is 0.

94. The computer-implemented method of any one of the preceding claims, wherein dividing the corresponding range for each of the plurality of predictors into a corresponding number of bins comprises dividing the corresponding range evenly into the corresponding number of bins.

95. The computer-implemented method of any one of the preceding claims, wherein dividing the corresponding range for each of the plurality of predictors into a corresponding number of bins comprises dividing the corresponding range non-evenly, so at least two bins encompass a different range therewithin.

96. The computer-implemented method of any one of the preceding claims, wherein the corresponding number of bins for each of the plurality of predictors is identical.

97. The computer-implemented method of any one of the preceding claims, wherein the corresponding number of bins for each of the plurality of predictors is 50.

98. The computer-implemented method of any one of the preceding claims, wherein the corresponding number of bins for each of the plurality of predictors is different.

99. The computer-implemented method of any one of the preceding claims, wherein a standard error rate for a bin is calculated as a ratio of (1): a quantity of erroneous base calls in the bin plus a predetermined correction number; and (2): the quantity of erroneous base calls in the bin and the quantity of correct base calls plus a predetermined correction number.

100. The computer-implemented method of any one of the preceding claims, wherein the predetermined correction number is 1.

101. The computer-implemented method of any one of the preceding claims, wherein the standard quality score is proportional to a logarithm of the standard error rate.

102. The computer-implemented method of any one of the preceding claims, wherein the standard quality score is calculated as a negative of a multiplication of 10 with the logarithm of the standard error rate.

103. The computer-implemented method of any one of the preceding claims, wherein a conservative error rate for a bin is calculated as a ratio of (1): a quantity of erroneous base calls in the bin plus a predetermined correction number; and (2): the quantity of erroneous base calls in the bin and the quantity of correct base calls plus a predetermined correction number.

104. The computer-implemented method of any one of the preceding claims, wherein the predetermined correction number is 5.

105. The computer-implemented method of any one of the preceding claims, wherein the conservative quality score is proportional to a logarithm of the conservative error rate.

106. The computer-implemented method of any one of the preceding claims, wherein calculating the standard quality score and the conservative quality score for each of the corresponding number of bins for each of the plurality of predictors further comprises: calculating a cumulative count of erroneous base calls; and calculating a cumulative count of correct base calls, wherein the cumulative count comprises counts from all bins whose coordinates are not greater than the coordinates of the corresponding bin.

107. The computer-implemented method of claim 68 further comprising: selecting, by the computer, a plurality of predictors; and ranking, by the computer, the plurality of predictors based on corresponding effects on a quality score of base calling.

108. The computer-implemented method of any one of the preceding claims, wherein each of the plurality of predictors corresponds to a dimension of the look-up table.

109. The computer-implemented method of any one of the preceding claims, wherein a size of the look-up table in the dimension is determined by number of bins corresponding of the corresponding predictor.

110. The computer-implemented method of any one of the preceding claims, wherein the size of the look-up table in the dimension is 50.

111. The computer-implemented method of claim 68 further comprising: selecting, by the computer, the training regions in one or more flow cell images acquired from a reference dataset, each training region comprising polonies of signals;112. The computer-implemented method of claim 68 further comprising: determining, by the computer, the corresponding range for each of the plurality of predictors in each training polony of the set of training regions.

113. The computer-implemented method of claim 68, further comprising:performing one or more preprocessing steps to generate the one or more flow cell images, the one or more preprocessing steps comprising:(1) background subtraction; and(2) color correction.

114. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises three predictors.

115. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises: a first predictor of clarity, a second predictor of max cc intensity, and a third predictor of low intensity median clarity.

116. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises four predictors.

117. The computer-implemented method of any one of the preceding claims, wherein the plurality of predictors comprises a first predictor of clarity, a second predictor of max cc intensity, and a third predictor of low intensity median clarity, and a fourth predictor of phasing and prephasing.

118. The computer-implemented method of any one of the preceding claims, wherein the second predictor of max cc intensity comprises image intensities without phasing or prephasing correction and without normalization of the image intensities within a single channel.

119. A computer-implemented system for predicting quality of base calling in sequencing, comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising:selecting, by a computer, a look-up table from a plurality of look-up tables, wherein each of the plurality of look-up tables is generated by training using a corresponding training dataset; selecting, by the computer, a plurality of predictors based on the selected look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a first quality score from the selected lookup table based on the corresponding value for each of the plurality of predictors, wherein the selected look-up table comprises a plurality of dimensions corresponding to the plurality of predictors.

120. A computer-implemented system for predicting quality of base calling in sequencing, comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising any one of claims 1 - 67.

121. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for predicting quality of base calling in sequencing, the operations comprising: selecting, by a computer, a look-up table from a plurality of look-up tables, wherein each of the plurality of look-up tables is generated by training using a corresponding training dataset; selecting, by the computer, a plurality of predictors based on the selected look-up table; determining, by the computer, a corresponding value for each of the plurality of predictors from one or more flow cell images of samples to be sequenced; and determining, by the computer, a first quality score from the selected look-up table based on the corresponding value for each of the plurality of predictors, wherein the selected look-up table comprises a plurality of dimensions corresponding to the plurality of predictors.

122. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for predicting quality of base calling in sequencing, the operations comprising any one of claims 1- 67.

123. A computer-implemented system for predicting quality of base calling in sequencing, comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising: generating a first look-up table for predicting quality of base calling in DNA sequencing, comprising: obtaining, by a computer, first training base calls from one or more first training flow cell images; selecting, by the computer, a first plurality of predictors based on at least a first sequencing cycle number of the one or more first training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the first training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of the first plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a first look-up table having a number of dimensions determined by the first plurality of predictors and each dimension comprising the corresponding number of bins; determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the first look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps:(a) calculating a standard quality score and a conservative quality score for each bin of the first look-up table;(b) determining coordinates of a bin with a maximum conservative quality score in the first look-up table;(c) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the first look-up table; and(d) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

124. A computer-implemented system for predicting quality of base calling in sequencing, comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising any one of claims 69 - 118.

125. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for predicting quality of base calling in sequencing, the operations comprising: generating a first look-up table for predicting quality of base calling in DNA sequencing, comprising: obtaining, by a computer, first training base calls from one or more first training flow cell images; selecting, by the computer, a first plurality of predictors based on at least a first sequencing cycle number of the one or more first training flow cell images, wherein each predictor is an indicator of base calling quality; obtaining, by the computer, correct base calls and erroneous bases calls by comparing the first training base calls to reference base calls; dividing, by the computer, a corresponding range of value for each of the first plurality of predictors into training regions corresponding to a corresponding number of bins; initializing, by the computer, a first look-up table having a number of dimensions determined by the first plurality of predictors and each dimension comprising the corresponding number of bins;determining, by the computer, a first number of correct base calls and a second number of erroneous base calls in each bin of the first look-up table; and iterating, by the computer and until the first number of correct base calls and the second number of erroneous base calls in each bin are no greater than a predetermined number, one or more of the following steps:(e) calculating a standard quality score and a conservative quality score for each bin of the first look-up table;(f) determining coordinates of a bin with a maximum conservative quality score in the first look-up table;(g) assigning a selected number of bins with the standard quality score that corresponds to the determined coordinates in the first look-up table; and(h) setting correct base calls and erroneous base calls in the selected number of bins to be a predetermined number.

126. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for predicting quality of base calling in sequencing, the operations comprising any one of claims 69-118.

Citation Information

Patent Citations

  • Methods and systems for analyzing image data

    US20210310065A1

  • Identifying nucleotides by determining phasing

    US20220051407A1

  • Quality measurement of base calling in next generation sequencing

    WO2023230279A1