Separating sequencing data in parallel with sequencing run in next generation sequencing data analysis

By sorting and sequencing data in advance during sequencing operation, the index sequence is automatically determined, and the delay problem of sequencing data classification and separation in the prior art is solved, and more efficient and accurate sequencing results and resource optimization are achieved.

CN120359305APending Publication Date: 2025-07-22ELEMENT BIOSCIENCES INC
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202380084762.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-21
Filing Date
2023-10-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, sequencing data from different samples can only be classified and separated after the sequencing run is completed, resulting in errors in sequencing results and waste of resources, and the incorrect sequencing task cannot be detected and terminated in the sequencing run.

Method used

By sorting and sequencing data in advance during the sequencing operation, and automatically determining the index sequence, the index sequence is read in the sequencing cycle using computer systems and optical imaging technology, the accurate classification and separation of the sequencing data is achieved, and the wrong sequencing operation is detected and terminated in a timely manner.

Benefits of technology

It improves the accuracy and efficiency of sequencing data analysis, avoids resource waste, reduces sequencing time and calculation costs, and ensures that the sequencing system can be released for other tasks in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359305A_ABST
    Figure CN120359305A_ABST
Patent Text Reader

Abstract

Described herein are system, apparatus, method, and / or computer program product embodiments and / or combinations and sub-combinations thereof that are capable of automatically determining index sequences during DNA sequencing data analysis. Embodiments of methods, systems, and media for automatically determining index sequences are disclosed herein as such particular applications such that sequencing results from multiple samples can be accurately classified and separated for downstream analysis, such as secondary analysis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 415,914, filed Oct. 13, 2022, and U.S. Provisional Patent Application No. 63 / 426,922, filed Nov. 21, 2022, which are hereby incorporated by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to separating sequencing data and sequencing data analysis, and particularly to separating sequencing data while a sequencing run is still in progress and index sequence determination during DNA sequence data analysis. Background Art

[0004] In next-generation sequencing (NGS) or NGS-like applications such as sequencing-by-synthesis, sequencing-by-ligation, or affinity sequencing, to identify the sequence of a target nucleic acid, new strands are synthesized one nucleotide base at a time. During each cycle, 3'-blocked nucleotides attach at complementary positions on the strand, ensuring that only one base will attach to any given strand during a single cycle. During the imaging step of each sequencing cycle, one or more images are recorded. A base calling algorithm is applied to the images to "read" the consecutive signals from each cluster or colony and convert the optical signals into an identification of the nucleotide bases added to the nucleotide sequence of each DNA fragment. An index sequence is appended to each DNA fragment to facilitate identification of DNA fragments from different samples.

[0005] Traditionally, in NGS, the sorting and separation (e.g., demultiplexing) of sequencing data from different samples can only be performed after the sequencing run has been completed. Summary of the Invention

[0006] Provided herein are embodiments of systems, devices, methods, and / or computer program products and / or combinations and sub-combinations thereof that are capable of pre-sorting and / or separating partial data from a sequencing run while the sequencing run is still in progress and automatically determining index sequences during DNA sequencing data analysis.

[0007] As such a specific application, embodiments of methods, systems, and media for automatically determining index sequences are disclosed herein such that sequencing results from multiple samples can be accurately sorted and separated for downstream analysis, such as secondary analysis.

[0008] Furthermore, as such a specific application, embodiments of methods, systems, and media for pre-sorting and / or separating (e.g., demultiplexing) a partial sequencing run are disclosed herein such that sequencing results from multiple samples can be accurately separated for downstream analysis, such as secondary analysis.

[0009] Other embodiments of these aspects include corresponding computer systems, devices, and computer program products recorded on computer storage devices, which are configured, alone or in combination, to perform the actions of the method. For a computer system that has been configured or is to be configured to perform an operation or action, the computer system has installed thereon software, firmware, hardware, or a combination thereof that causes the computer system to perform the operation or action during operation. For a computer program product that has been configured or is to be configured to perform an operation or action, the computer program product includes instructions that, when executed by a hardware processor, cause the hardware processor to perform the operation or action.

[0010] Additional embodiments, features, and advantages of the present disclosure, as well as the structure and operation of the embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings incorporated herein and forming a part of this specification illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable those skilled in the art to make and use the embodiments.

[0012] Figure 1 A block diagram of a system for obtaining flow cell images, generating sequencing data, performing early classification and / or separation of sequencing data during a sequencing run, and / or determining index sequences in sequencing data analysis, according to some embodiments, is shown.

[0013] Figure 2 A schematic diagram of a double-stranded DNA fragment, according to some embodiments, is shown, having a read of the fragment (Read 1), a second read of the complementary strand of the DNA fragment (Read 2), a first index sequence, and a second index sequence.

[0014] Figure 3 A table of exemplary first and second index sequences that can be used to uniquely identify different sequencing samples or DNA fragments attached thereto, according to some embodiments, is shown.

[0015] Figure 4 A block diagram of a computer system for performing early classification and / or separation of sequencing data and / or determining index sequences in sequencing data analysis, according to some embodiments, is shown.

[0016] Figure 5A A flowchart of a method for performing early classification and / or separation of sequencing data in DNA sequence data analysis, according to some embodiments, is shown.

[0017] Figure 5BA flowchart of a method for performing automated index sequence determination in DNA sequence data analysis according to some embodiments is shown.

[0018] Figure 6 is a schematic diagram showing an exemplary linear single-stranded library molecule (700), which includes: a surface pinned primer binding site (720); an optional left unique identification sequence (780); a left index sequence (760); a forward sequencing primer binding site (740); an insertion region (710) having a target sequence; a reverse sequencing primer binding site (750); a right index sequence (770); and a surface capture primer binding site (730).

[0019] Figure 7 is a schematic diagram showing an exemplary linear single-stranded library molecule (700), which includes: a surface pinned primer binding site (720); a left index sequence (760); a forward sequencing primer binding site (740); an insertion region (710) having a target sequence; a reverse sequencing primer binding site (750); a right index sequence (770); an optional right unique identification sequence (790); and a surface capture primer binding site (730).

[0020] Figure 8A is a schematic diagram showing hybridization of an exemplary linear single-stranded library molecule ( Figures 6 - 7 700 in) with a double-stranded splint molecule (200), thereby circularizing the library molecule to form a library-splint complex (800) having two nicks. The library molecule ( Figures 6 - 7The 700 in it may include one or more selected from the following: a first additional left universal adapter sequence; a first left universal adapter sequence (720); a first left ligation adapter sequence; a left index sequence (760); a second left ligation adapter sequence; a second left universal adapter sequence (740); a third left ligation adapter sequence; a sequence of interest (710); a third right ligation adapter sequence; a second right universal adapter sequence (740); a second right ligation adapter sequence; a right index sequence (770); a first right ligation adapter sequence; a first right unique identifier sequence; a first right universal adapter sequence (730); and a first additional right universal adapter sequence. The double-stranded splint molecule 200 includes a first splint strand hybridized to a second splint strand. The first splint strand includes a first region (320) hybridized to a sequence at one end of a linear single-stranded library molecule and a second region (330) hybridized to a sequence at the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand is hybridized to the second splint strand. For simplicity, the library-splint complex (800) does not show any ligation adapter sequences or additional universal adapter sequences. Those skilled in the art will recognize that the linear library molecule (700) can include any combination of any one or two or more ligation adapters, with or without one or both of the additional universal adapter sequences. Those skilled in the art will recognize that the library-splint complex (800) can include any combination of any one or two or more ligation adapters, with or without one or both of the additional universal adapter sequences present in the library molecule (700).

[0021] Figure 8B Shows a covalently closed circular library molecule (900) having a sequence of interest (710), a second right universal adapter sequence (750), a right index sequence (770), a first right universal adapter sequence (730), a second splint strand sequence (1400), a first left universal adapter sequence (720), a first left unique identifier sequence (780), a left index sequence (760), and a second left universal adapter sequence (740).

[0022] Figure 9 Is a schematic diagram of various exemplary configurations of multivalent molecules. Left figure (Class I): A schematic diagram of a multivalent molecule with a "starburst" or "helter-skelter" configuration. Middle figure (Class II): A schematic diagram of a multivalent molecule with a dendrimer configuration. Right figure (Class III): A schematic diagram of multiple multivalent molecules formed by the reaction of streptavidin with a 4-arm or 8-arm PEG-NHS having biotin and dNTP. The nucleotide unit is designated as 'N', biotin is designated as 'B', and streptavidin is designated as 'SA'.

[0023] Figure 10 Schematic of an exemplary multivalent molecule comprising a universal core attached to multiple nucleotide arms.

[0024] Figure 11 Schematic of an exemplary multivalent molecule comprising a dendritic core attached to multiple nucleotide arms.

[0025] Figure 12 Schematic of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, where the nucleotide arms comprise biotin, a spacer, a linker, and nucleotide units.

[0026] Figure 13 Schematic of an exemplary nucleotide arm comprising a core attachment portion, a spacer, a linker, and nucleotide units.

[0027] Figure 14 Shows the chemical structure of an exemplary spacer (top) and the chemical structures of various exemplary linkers, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker (bottom).

[0028] Figure 15 Shows the chemical structures of various exemplary linkers, including Linkers 1 to 9.

[0029] Figure 16 Shows the chemical structures of various exemplary linkers that are joined / attached to nucleotide units.

[0030] Figure 17 Shows the chemical structures of various exemplary linkers that are joined / attached to nucleotide units.

[0031] Figure 18 Shows the chemical structures of various exemplary linkers that are joined / linked to nucleotide units.

[0032] Figure 19 Shows the chemical structures of various exemplary linkers that are joined / linked to nucleotide units.

[0033] Figure 20 Shows the chemical structure of an exemplary biotinylated nucleotide arm. In this example, the nucleotide unit is attached to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.

[0034] Figure 21 Provides a schematic of an embodiment of the low-binding solid support of the present disclosure, where the support comprises an alternating layer of a glass substrate and a hydrophilic coating covalently or non-covalently adhered to the glass, and the support further comprises a chemically reactive functional group that serves as an attachment site for an oligonucleotide primer.

[0035] In the accompanying drawings, like reference numerals generally denote the same or similar elements. Additionally, generally, the leftmost digit of a reference numeral may identify the drawing in which the reference numeral first appears. Detailed Description

[0036] Embodiments of systems, devices, methods, and / or computer program products and / or combinations and sub - combinations thereof are provided herein that are capable of pre - sorting and / or separating sequencing data while the corresponding sequencing run is still in progress to ensure accurate and reliable sequencing analysis and to avoid wasting time and resources on problematic sequencing runs. Embodiments of systems, devices, methods, and / or computer program products and / or combinations and sub - combinations thereof are provided herein that are capable of automatically determining index sequences during sequencing data analysis to ensure accurate and reliable sequencing results. The techniques herein can be used for sequencing data, such as sequencing reads obtained by various imaging and / or sequencing techniques. The techniques disclosed herein can be used for data analysis in next - generation sequencing (NGS), and NGS will be used as the main example herein for describing the applications of these techniques. However, such techniques can also be used in other applications that use index sequences in sequencing data analysis.

[0037] In existing NGS data analysis, different samples can be sequenced together in a sequence run, and after completion of the sequence run, the samples are sorted and / or separated into different data files, such as FastQ files. Identifiers of the samples (e.g., index sequences) are attached to library molecules or sequences of interest, enabling the separation of sequencing reads based on such identifiers. The sorting and / or separation of sequencing data can also rely on the manual determination and input of index sequences. Errors in the index sequences or other regions of library molecules can cause incorrect sequencing results and lead to a significant waste of time and resources when performing any downstream analysis on the incorrect sequencing data, such as for hours or more. Additionally, regenerating data files to remove errors is also time - consuming and resource - intensive. The techniques disclosed herein can be used for the pre - sorting and / or separation of sequencing data and the pre - determination of sequencing errors while the sequence run is still in progress. Thus, problematic runs can be detected and terminated early to avoid wasting time and resources and to free up the sequencing system for other sequencing applications. The techniques disclosed herein can also speed up sequencing data analysis by enabling the pre - sorting and / or separation of partial sequencing results (e.g., index sequences) before completion of the sequencing run.

[0038] The techniques disclosed herein are rooted in the characteristics of index-based sequencing data analysis. Specifically, the techniques disclosed herein are advantageously applicable to the read order in which one or more index sequences (or functional equivalents thereof) are read before reading a sequence of interest (e.g., a DNA fragment and / or its optional complementary strand). After completing the sequencing cycles corresponding to the index sequences, the index sequences can be classified and / or separated, and statistical information can be computed to evaluate whether there are errors in the index sequences. Such errors may be caused by inaccuracies, including but not limited to index sequence errors and / or library problems that may cause problems in the sequencing results. If desired, the techniques herein advantageously enable early detection of such errors and early termination of the sequencing run to prevent wasting sequencing time, sequencing resources, and computational time in performing a problematic sequencing run and analyzing its sequencing results. Thus, the sequencing system can be freed up to perform other sequencing tasks.

[0039] Sequencing System

[0040] Figure 1 FIG. shows a block diagram of a computer-implemented system 100 in accordance with one or more embodiments disclosed herein. System 100 has a sequencing system 110, which includes a flow cell 112, a sequencer 114, an imager 116, a data memory 122, and a user interface 124. Sequencing system 110 can be connected to a cloud 130. Sequencing system 110 can include one or more of the following: a dedicated processor 118, a field programmable gate array (FPGA) 120, and a computer system 126.

[0041] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form DNA sequences for base calling on the flow cell. The flow cell 112 can include the scaffolds described herein. The carrier can be a solid-phase carrier. As disclosed herein, the carrier can include a surface coating thereon. The surface coating can be a polymer coating as disclosed herein.

[0042] Flow cell 112 may include a plurality of tiles or imaging regions thereon, and each tile may be divided into a grid of sub-tiles. Each sub-tile may include a plurality of clusters or colonies thereon. As a non-limiting example, the flow cell may have 424 tiles, and each tile may be divided into a 6x9 grid, thus having 54 sub-tiles. The flow cell images disclosed herein may be images of signals including a plurality of clusters or colonies. The flow cell image may include one or more signal tiles or one or more signal sub-tiles. In some embodiments, the flow cell image may be an image including all tiles and substantially all signals thereon. The flow cell image may be acquired from the channel using imager 116 during an imaging or sequencing cycle. In some embodiments, each tile may include millions of colonies or clusters. As a non-limiting example, a tile may include from about 1 to 10 million clusters or colonies. Each colony may be a collection of many copies of DNA fragments. In some aspects, the flow cell images herein may be images including at least a portion of the tiles and substantially all signals thereon. The flow cell image may be acquired from the channel using imager 116 during an imaging or sequencing cycle.

[0043] In the case where a 3D sample (e.g., a cell or tissue) is fixed on the flow cell, the flow cell image may be at multiple z-levels that are orthogonal to the image plane of the flow cell image. Thus, the flow cell image may include multiple z-levels in order to cover the entire 3D sample. The z-axis may extend from the objective lens of the optical system disclosed herein to the carrier, e.g., the flow cell. Each z-level of the flow cell image may be separated from an adjacent z-level by a predetermined distance, e.g., from about 0.1 um to about 15 um. Each z-level of the flow cell image may be separated from an adjacent level by 1 um to 10 um. At each z-level, the flow cell image may be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include at least a portion of one or more tiles or sub-tiles of the flow cell in its field of view. The image plane is defined by the x-axis and the y-axis. And the z-axis is orthogonal to the x-y plane. Although the flow cell image, the sample, and the z-axis are described in a Cartesian coordinate system, any other coordinate system may be used to define the spatial positions and relationships of the colonies or clusters and their images herein. Other coordinate systems may include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.

[0044] Sequencer 114 can be configured to flow a nucleotide mixture onto flow cell 112, cleave blockers from the nucleotides between flow steps, and perform other steps for forming a DNA sequence on flow cell 112. The nucleotides can have a fluorescent element attached thereto that emits light or energy at a wavelength indicative of the nucleotide type. Each type of fluorescent element can correspond to a specific nucleobase (e.g., A, G, C, T). The fluorescent element can emit light at a visible wavelength. In some embodiments, sequencer 114 and flow cell 112 can be configured to perform the various sequencing methods disclosed herein, e.g., affinity sequencing.

[0045] For example, each nucleobase can be assigned a color. Different types of nucleotides can have different colors. For example, adenine (A) can be red, cytosine (C) can be blue, guanine (G) can be green, and thymine (T) can be yellow. The color or wavelength of the fluorescent element for each nucleotide can be selected such that the nucleotides can be distinguished from one another based on the wavelength of the light emitted by the fluorescent element.

[0046] Imager 116 can be configured to capture an image of flow cell 112 after each flow step. In one embodiment, imager 116 is a camera configured to capture a digital image, such as a CMOS or CCD camera. The camera can be configured to capture an image of the wavelength of the fluorescent element bound to the nucleotide. These images can be referred to as flow cell images.

[0047] In some embodiments, imager 116 can include one or more of the optical systems disclosed herein. The optical system can be configured to capture an optical signal from the flow cell and generate a corresponding digital image. The digital image can then be used for base calling.

[0048] In some embodiments, images of the flow cell can be captured in groups, where each image in the group is taken at a wavelength or spectrum that matches or includes only one of the fluorescent elements. In another embodiment, the image can be acquired as a single image that captures all of the wavelengths of the fluorescent elements.

[0049] The resolution of imager 116 can control the level of detail in the flow cell image, including pixel size. In existing systems, this resolution is very important because it controls the accuracy of the dot-finding algorithm in identifying the center of the community. In some embodiments, the image resolution of the flow cell image disclosed herein can be from about 10 nanometers (nm) to several hundred nm or higher. One way to improve the accuracy of dot finding is to improve the resolution of imager 116 or to improve the processing of the images captured by imager 116. It is possible to perform detection of the center of the community in pixels other than those detected by the dot-finding algorithm. These methods can allow for an improvement in the accuracy of community center detection without increasing the resolution of imager 116. The resolution of the imager can even be lower than that of existing systems with comparable performance, which can reduce the cost of sequencing system 110.

[0050] The image quality of the flow cell image can control the base calling quality. One way to improve the accuracy of base calling is to improve imager 116 or to improve the processing of the images captured by imager 116 to obtain better image quality.

[0051] Sequencing system 110 can be configured to perform the various operations disclosed herein. Sequencing system 110 can be configured to perform imaging, preliminary analysis steps, and / or other operations disclosed herein. The operations or actions disclosed herein can be performed by a dedicated processor 118, FPGA 120, computing system 126, or a combination thereof. One or more operations or actions in method 500 disclosed herein can be performed by a dedicated processor 118, FPGA 120, computing system 126, or a combination thereof. In some embodiments, which operations or actions will be performed by a dedicated processor 118, FPGA 120, computing system 126, or a combination thereof can be determined based on, but not limited to, one or more of the following: the computational time for a particular operation, the complexity of the computation in a particular operation, the need for data transfer between hardware devices, the energy consumption required for the computation, the heat dissipation for performing the computation, or a combination thereof.

[0052] Computing system 126 can include one or more general-purpose computers that provide an interface for running various programs in an operating system such as Windows TM or Linux TM and the like. Such operating systems typically provide a great deal of flexibility to the user.

[0053] In some embodiments, the dedicated processor 118 may be configured to perform the operations in the methods herein. The dedicated processor may not be a general-purpose processor, but rather a customized processor with specific hardware or instructions for performing these steps. The dedicated processor runs specific software directly without an operating system. The lack of an operating system reduces overhead but at the cost of the flexibility of what the processor can execute. The dedicated processor may use a customized programming language, which may be designed to operate more efficiently than software running on a general-purpose computer. This can increase the speed of performing the steps and allow for real-time processing.

[0054] In some embodiments, the dedicated processor 118 or the computing system 126 may include a reconfigurable logic device, such as an artificial intelligence (AI) chip, a neural processing unit (NPU), an application-specific integrated circuit (ASIC), or a combination thereof. The reconfigurable logic device may be configured to perform one or more of the operations herein. The reconfigurable logic device may be configured to perform one or more of the operations herein and accelerate the operations by allowing parallel data processing compared to a CPU.

[0055] In some embodiments, the FPGA 120 may be configured to perform the operations of the methods herein. The FPGA is programmed to be hardware that will only perform specific tasks. Special programming languages may be used to convert software steps into hardware components. Once the FPGA is programmed, the hardware directly processes the digital data provided to it without running software. Instead, the FPGA can use logic gates and registers to process digital data. Since no operating system overhead is required, the FPGA generally processes data faster than a general-purpose computer. Similar to a dedicated processor, this is at the cost of flexibility.

[0056] The lack of software overhead can also allow the FPGA to operate faster than a dedicated processor, although this will depend on the exact processing to be performed and the specific FPGA and dedicated processor.

[0057] A group of FPGAs 120 may be configured to perform these steps in parallel. For example, many FPGAs 120 may be configured to perform processing steps or operations on an image, a collection of images, sub-blocks, or selected regions in one or more images. Each FPGA 120 can simultaneously perform its own part of the processing steps or operations, thereby reducing the time required to process the data. This can allow the processing steps or operations to be completed in real time. Further discussion of the use of FPGAs is provided below.

[0058] Performing processing steps or operations in real time can allow the system to use less storage because the data can be processed as it is received. This is an improvement over conventional systems that may need to store the data before processing it, which may require more storage or access to a computer system located in the cloud 130.

[0059] In some embodiments, the data storage 122 is used to store information used in the methods herein. This information may include the image itself or information derived from the images captured by the imager 116. The DNA sequences determined according to base determination may be stored in the data storage device 122. The parameters for identifying the community locations may also be stored in the data storage device 122. The raw and / or processed image intensities of each community may be stored in the data storage device. The regions and / or sub-tiles corresponding to each community may also be stored in the data storage device 122. The transformation matrices for each region and / or sub-tile of different cycles and / or channels may also be stored in the data storage device 122. The mapped index sequences, reference index sequences, and / or hash tables or hash maps may be stored in the data storage 122. The values of the statistical parameters may be stored in the data storage 122. The mapped index sequences, reference index sequences, and / or hash tables or hash maps may be stored in the data storage 122. The values of the statistical parameters may be stored in the data storage 122.

[0060] The user interface 124 may be used by the user to operate the sequencing system or access data stored in the data storage device 122 or the computer system 126.

[0061] The computer system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. The computer system may also perform the pre-classification and / or separation of the index sequences and the steps in the foregoing operations and / or subsequent operations including but not limited to secondary analysis. The computer system may also perform the index sequence determination and the steps in the foregoing operations and / or subsequent operations including but not limited to secondary analysis. In some embodiments, the computer system 126 is the computer system 400, as Figure 4 more specifically described. The computer system 126 may store information about the operation of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 may be configured to transfer information between the sequencing system 110 and the cloud 130.

[0062] As discussed above, sequencing system 110 may have a dedicated processor 118, FPGA 120, or computer system 126. The sequencing system may use one, both, or all of these components to perform the necessary processing described above. In some embodiments, when these components are present together, the processing tasks are divided among them. For example, FPGA 120 may be used to perform some or all of the following operations: operations prior to the early separation of index sequences herein, early classification and / or separation of index sequences, and subsequent operations, while computer system 126 may perform other processing functions on sequencing system 110, such as base calling. As another example, FPGA 120 may be used to perform some or all of the following operations: preprocessing operations, index sequence determination, and subsequent operations, while computer system 126 may perform other processing functions on sequencing system 110, such as base calling. Those skilled in the art will understand that various combinations of these components will allow for various system embodiments that balance the efficiency and speed of processing with the cost of the processing elements.

[0063] Cloud 130 may be a network, remote storage device, or some other remote computing system separate from sequencing system 110. The connection to cloud 130 may allow access to data stored external to sequencing system 110 or allow for the updating of software in sequencing system 110.

[0064] Method for Early Classification and Separation of Index Sequences

[0065] In some embodiments, each index sequence herein is a sequence of two or more nucleotide bases that, alone or in combination with other parts of the library molecule, functions to uniquely identify the library molecule. Each index sequence may comprise a contiguous sequence of multiple nucleotide bases. The number of nucleotide bases may be from 2 to 200. The number of nucleotide bases may be from 4 to 100. The number of nucleotide bases may be from 6 to 50. Each library molecule may comprise one or more index sequences. Two index sequences may be different from each other. The identification of a library molecule may associate the library molecule with a single sample from a list of samples. The identification of a library molecule may associate the library molecule with one or more samples from a list of samples. The classification and / or separation of index sequences (or functional equivalents thereof) as disclosed herein may be performed after a flow cell image corresponding to the index sequence has been acquired, but before other flow cell images of subsequent cycles have been acquired or are being acquired. For example, the methods disclosed herein may be performed after a flow cell image of an index sequence pair has been acquired, which may be about 8 to 12 of the first 20 to 30 cycles of a sequencing run. The methods herein may be performed in parallel while the sequence run is still in progress and flow cell images of subsequent cycles (e.g., cycles 31 to 151) are being acquired.

[0066] Figure 5A FIG. 1 shows a flowchart of an exemplary embodiment of a computer-implemented method 500 for pre-classifying and / or separating index sequences during NGS sequencing. Method 500 may include some or all of the operations disclosed herein. The operations may be performed in an order that is not limited to that described herein.

[0067] Method 500 may be executed by one or more processors disclosed herein. In some embodiments, the processor may include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU), a graphics processing unit (GPU), and / or an NPU. The integrated circuit may include a chip such as a field programmable gate array (FPGA), an ASIC, and an AI chip. In some embodiments, the processor may include a computing system 400.

[0068] In some embodiments, some or all of the operations in method 500 may be executed by an FPGA and / or other devices (e.g., an AI chip or an NPU). In embodiments where some operations are executed by an FPGA, the data after the operations are executed by the FPGA may be transmitted by the FPGA to other devices (e.g., a CPU) such that the other devices (e.g., a CPU) may use such data to perform subsequent operations in method 500. Similarly, data may also be transmitted from other devices (e.g., a CPU) to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 500 may be executed by a CPU. Alternatively, the operations executed by the CPU may be executed by other processors (such as a dedicated processor) or an NPU. In some embodiments, all of the operations in method 500 may be executed by an FPGA. In some embodiments, some of the operations in method 500 may be executed by an FPGA, while some other operations in method 500 are executed by an AI chip or an NPU to improve the energy consumption, heat dissipation, and / or computing time required for sequencing analysis.

[0069] In some embodiments, method 500 is performed during or after cycle N, which is different from a reference cycle. A template image (e.g., polymerase colony map) can be generated in the reference cycle, and polymerase colonies from one or more channels within the reference cycle can be included in the template image in a reference coordinate system while the flow cell image for cycle N or after cycle N has not been captured or is currently being captured. In some embodiments, cycle N is the current cycle. N can be any non - zero integer. For example, for short read sequencing, N can be any integer from 1 to 150. In some embodiments, N can be any number after some or all of the sequencing cycles for the index sequence. For example, when each library molecule has a single index of 9 bases sequenced during the first 15 cycles, N can be 16, 20, 30, or 40. As another example, N can be any integer from 1 to 300 or from 1 to 400.

[0070] In some embodiments, method 500 is performed during cycle N while sequencing and image acquisition in a subsequent cycle (e.g., cycle N + 1) are being performed or have not been performed. In some embodiments, method 500 is performed in parallel with the sequence run to advantageously reduce the total time for sequencing and preliminary analysis. In some embodiments, method 500 is performed in parallel with the sequence run to advantageously reduce the storage space required to store the flow cell images.

[0071] The reference coordinate system can be the common coordinate system disclosed herein. The common coordinate system can be predefined. The common coordinate system can be a Cartesian coordinate system. Various other coordinate systems can also be used. Other coordinate systems can include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.

[0072] The flow cell images herein can be obtained from 1, 2, 3, 4, or more channels of imager 116 using the optical systems disclosed herein. In some embodiments, multiple flow cell images are obtained in a single flow cycle or multiple flow cycles during the sequence run. In some embodiments, the flow cell images are obtained in the first 5, 10, 15, 20, 30, 50, 80, or 100 cycles of the sequence run. Each flow cell image can include one or more tiles (imaging regions), and each tile can be divided into multiple sub - tiles. Each sub - tile can include multiple colonies. Each sub - tile can include multiple regions, where each region includes a number of colonies. For example, polymerase colonies can be extracted from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, colonies can be extracted from a flow cell image from a single channel. The flow cell images as disclosed herein can be images obtained using the flow cell 112 as Figure 1 shown.

[0073] Flow cell 112 may contain a sample immobilized thereon. The sample may include a plurality of nucleic acid template molecules. The sample may include a two-dimensional (2D) sample or a three-dimensional (3D) volume sample. The nucleic acid template molecules may be randomly or distributed in various patterns on the flow cell 112. In some embodiments, a plurality of communities or clusters herein may be extracted from a specific region of a tile (e.g., each sub-tile). In the case of each sub-tile, communities may be extracted in a predetermined pattern or randomly.

[0074] In some embodiments, the communities or clusters being sequenced in the flow cycle can have a certain nucleotide diversity, e.g., in base calling. Method 500 can allow for the early separation and classification of nucleotide base sequences even if the polymerase communities or clusters in the sequencing cycle are of low diversity or unbalanced diversity. The nucleotide diversity of a population of nucleotide acid molecules (e.g., a community or a cluster) can refer to the relative proportions of nucleotides A, G, C, and T / U present in each flow cycle. The relative proportions of nucleotides can be within the field of view or within the entire flow cell image. Optimal high or balanced diversity data typically can have approximately equal proportions of all four nucleotides represented in each flow cycle of the sequencing run. Low or unbalanced diversity data typically can include high proportions of certain nucleotides and low proportions of other nucleotides in some flow cycles of the sequencing run, e.g., less than 10% of the total of all 4 nucleotides. Thus, an image corresponding to a high proportion of certain nucleotides can have more signal points (communities or clusters) compared to an image corresponding to a low proportion of certain nucleotides. As an example of low or unbalanced diversity data, in a certain flow cycle, bases A, T, C, and G can be approximately 1%, 2%, 1%, and 95% of the total community, respectively. Subsequently, in that particular flow cycle, the flow cell image from the channels corresponding to A, T, and C is darker than the flow cell image corresponding to nucleotide G, and the number of polymerase chain reaction (polony) or clusters is much less. As another example of low or unbalanced diversity data, bases A, T, C, and G in the polymerase chain reaction in multiple flow cycles may be approximately 2%, 5%, 10%, and 83%, respectively. In embodiments where low or unbalanced diversity data is present in a particular cycle and the data is imaged for sequencing analysis, using prior art image registration may fail because the images from one or more channels are too dark compared to the images obtained from other channels (e.g., the signal points of the polymerase communities are too sparse and / or dim), thus causing problems in subsequent color correction. Further, in embodiments where low or unbalanced diversity data is present in a particular cycle, using prior art image registration, color correction, and subsequent base identification may fail because the images from one or more channels are too dark (e.g., the signal points of the polymerase communities are too sparse and / or dim). In some embodiments, method 500 is configured to perform early classification and separation of index sequences based on the flow cell image even if the polymerase communities or clusters in the flow cell image are of unbalanced nucleotide diversity. In some embodiments, method 500 is configured to perform early classification and separation of index sequences based on the flow cell image at a pre-determined quality level even if the polymerase communities or clusters in the flow cell image are of unbalanced nucleotide diversity. In terms of base identification in one or more cycles, the pre-determined quality level can be no less than Q20, Q25, Q28, Q30, Q35, Q38, Q40, or higher.In terms of base identification in one or more cycles, the predetermined quality level can be an error of no greater than 2%, 1%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01% or less.

[0075] In addition to base bias that affects diversity, complexity is also a factor affecting the existing preliminary analysis methods that perform base identification and the foregoing steps leading to base identification. The methods herein allow for accurate and reliable pre-classification of sequencing data and its pre-separation from low-complexity data. Generally, complexity can indicate the source of a sample. A singleplex sample can contain DNA fragments or molecules from the same sample region or the same sample source in a genome. A multiplex sample can contain DNA fragments or molecules from different sample sources (e.g., liver, kidney, heart, cancerous tissue, etc.) or from one or more sample regions in a genome. When the complexity is below a certain number (e.g., 8 or 16), the signal may be of low complexity. For example, in a 2-cycle sequence, the entire polymerase population is AT, TG, GC, or CA in two adjacent cycles. Each of the bases A, T, C, and G accounts for 25% of the total number of bases in that cycle, but its complexity is less than 8, and the sequence is not all random. In some embodiments, method 500 is configured to perform pre-classification and separation of sequencing data at a predetermined quality level, even if the polymerase population or cluster is of low complexity.

[0076] In some embodiments, method 500 can include operation 510 of generating one or more mapped index sequences and mapping the mapped index sequences to multiple samples. Operation 510 can be performed by a processor as disclosed herein. The mapping can be based on a first error tolerance rate. The mapped index sequences can include an error-free reference index sequence and its variants within the first error tolerance rate.

[0077] In some embodiments, operation 510 of generating one or more mapped index sequences and mapping them to multiple samples is performed by using a mapping table (e.g., a hash table or a hash map). The mapping table or its functional equivalent herein can use a mapping function to map a key to a value. In the instance of a hash table, a hash function can be calculated for each key (e.g., the mapped index sequence) to determine a unique location for storing the value (e.g., the sample identification) associated with that key. During a lookup, the index sequence to be mapped is hashed using the hash function, and the resulting hash value can indicate the storage location of the sample identification associated with that index sequence. The mapping function (e.g., the hash function) can generate a unique result value for each different key (e.g., each index sequence), such that the sample identifications associated with each different index sequence are stored in different locations in the hash table.

[0078] The mapping of the mapped index sequences (e.g., the mapped index sequences can include the reference index sequences and their variants within the error tolerance rate) to different samples can be pre-computed before the sequencing run. If one or more factors associated with the sequencing run change, the mapping function herein can be updated. For example, if the number of samples, the error tolerance rate, the length of the index sequences, or the diversity of the index sequences change, the mapping function herein can be updated.

[0079] In some embodiments, a large number of samples (e.g., 100, 200, or more) can be combined and loaded onto the same support of the flow cell device and sequenced simultaneously during a single sequencing run. Two different library molecules or different DNA fragments (with or without their complementary strands) can be from the same or different samples. In some embodiments, each DNA fragment (alone or in combination with its complementary strand) can be uniquely identified by one or more index sequences. One or more index sequences can serve as the unique identifier for each set of sequencing reads of the polymerase colonies or clusters. In some embodiments, even if there are errors within the error tolerance rate in one or more index sequences, e.g., 1 or 2 incorrect out of 10 bases in each index sequence, the index sequence can still be used to uniquely identify the DNA fragment, i.e., one or more index sequences within the predetermined error tolerance rate can still be used to uniquely identify the DNA fragment.

[0080] In some embodiments, the index sequences can have a set of mutually exclusive variants within the error tolerance rate. As an example, if the error tolerance rate is 1 mismatch in 9 bases, the entire set of index sequences within one mismatch for sample A can still be different from the entire set of index sequences within one mismatch for sample B. Figure 3 Table 1 in shows 96 exemplary index sequence pairs that can be used to uniquely identify 96 different samples or DNA fragments.

[0081] Each library molecule can have a single index sequence, a pair of index sequences, or even more index sequences. Figure 2 、 6 and 7 show exemplary embodiments of index sequence pairs located at the opposite ends of the reads of the DNA fragment.

[0082] When there are two or more index sequences in the set of sequencing reads, the lengths of two different index sequences can be the same or different. In some embodiments, each of the one or more index sequences contains any number of nucleotide bases less than 100. For example, the index sequences herein can contain 1 to 100, 2 to 80, or 2 to 75 bases. In some embodiments, the index sequences can contain 6 to 12, 7 to 13, 8 to 12, 6 to 16, or 8 to 16 nucleotide bases.

[0083] In some embodiments, there is no limitation on the diversity of nucleotide bases in a single index sequence and / or multiple index sequences for identifying different numbers of samples in the index sequences herein. In some embodiments, the index sequences can have balanced or unbalanced nucleotide diversity. In other words, a single index sequence can contain any number of two, three, or four different types of nucleotide bases, and one or more types of nucleotide bases can be less than 10%, 8%, 5%, 2%, or even 1% of the total number of bases in some single index sequences or all index sequences considered together. For example, the index sequence of sample A is AGAGGAAGG, which has only two bases, and another index sequence of sample B is TCTCTCTGC, which has only two other different bases. These two index sequences have unbalanced nucleotide diversity. In some embodiments, the diversity of image data in a sequencing cycle (e.g., different types of bases across multiple polymerase communities in a flow cell image) can be balanced diversity or unbalanced diversity as disclosed herein.

[0084] In some embodiments, method 500 can include an operation of generating one or more reference index sequences for each set of sequencing reads.

[0085] Generating the reference index sequences can be based on the number of samples or DNA fragments being sequenced in a sequencing run, e.g., 120 different samples. Generating the reference index sequences can be before operation 510. Generating the reference index sequences can include determining the percentage of occurrence of each of the 4 nucleotide bases (i.e., A, T, C, G) in all the index sequences. In other words, generating the reference index sequences can include determining the diversity of the index sequences. For example, the reference index sequences can have high diversity such that when combined together, each of the 4 bases appears in each of the reference index sequences or in all the index sequences at about 25%. In some embodiments, the reference index sequences herein can have unbalanced diversity such that the total occurrence rate of one or more of the 4 bases in all the index sequences is less than 10%. Generating the reference index sequences can include determining the length of each index sequence. The length of the reference index sequences can be based on the number of samples and / or the diversity of the index sequences. In some embodiments, generating the reference index sequences can include determining a plurality of ordered sequences, each sequence having a preset number of nucleotide bases. In the case where there is 1, 2, or even more nucleotide errors in an ordered sequence, the ordered sequence may still be different from all other determined sequences. For example, as Figure 3 shown, when there is an error in index 1 of example_sample_1, instead of the correct sequence "GGCTCCTAC", it becomes "AGCTCCTAC". Compared with other index 1 sequences, the incorrect index 1 is still unique.

[0086] After generating the reference index sequence(s), method 500 may include an operation of storing one or more reference index sequences. In some embodiments, each of the reference index sequences may be stored corresponding to a unique identification number, e.g., "sample_1" as shown in Figure 3 . The stored reference index sequence(s) used as a reference may be retrieved later to calculate the statistical parameters for sequencing as disclosed herein.

[0087] In some embodiments, method 500 herein may include an operation of generating one or more reference index sequences for each of a plurality of samples. The one or more reference index sequences are index sequences without errors and can be used to compare with the index sequences obtained from the sequencing analysis of a sequencing run. Figure 3 Table 1 in shows 96 reference index sequences. The operation of generating one or more mapping index sequences is based on the one or more reference index sequences and an error tolerance rate.

[0088] The error tolerance rate herein, i.e., the first error tolerance rate and / or the second error tolerance rate, may be pre-determined. The error tolerance rate herein may be customized. The error tolerance rate may be customized based on various factors (such as the length of the index sequence and / or the characteristics of the sample). The errors herein may include, but are not limited to, one or more of the following: mismatch, unassigned, missing, insertion, mis-association, mis-pairing, or a combination thereof. For example, the error tolerance rate may be about 1 missing base in 7, 8, 9, 10, 11, or 12 bases. The error tolerance rate may be about 2 mismatched bases in 7, 8, 9, 10, 11, or 12 bases. The error tolerance rate may be 2%, 5%, 10%, 15%, 20%, 21%, 22%, 23%, 24%, 25%, or 30%.

[0089] As an example, in the case where the error tolerance rate is 1 mismatched base in 2 bases, the reference index sequence of AG can have the following mapped index sequences: AG, TG, GG, CG, AT, AC, and AA. As another example, in the case where the error tolerance rate is 1 deleted base in 4 bases, the reference index sequence of AGTC can have at least the following mapped index sequences: AGTC, GTC, ATC, AGC, and AGT. In other words, the mapped index sequences can include some or all of the possible sequences that differ from the reference index sequence by less than or equal to the errors within the error tolerance rate. If the reference index sequence (e.g., AG) maps to sample A, then all the mapped index sequences of AG also map to sample A. The mapped index sequences within the tolerance rate of the reference index sequence cannot map to other samples. The mapping of the mapped index sequences to their corresponding samples can be a many-to-one mapping relationship. Each of the mapped index sequences can be mapped to a single sample to avoid mapping collisions.

[0090] In some embodiments, method 500 can include operation 520 of generating sequencing data of the sequencing run while the sequencing run is still in progress. The sequencing data can include a set of sequencing reads for each polymerase community in a plurality of samples. Each set of sequencing reads can include one or more of the plurality of index sequences.

[0091] For example, there can be 80 samples fixed on a flow cell, and in each flow cell image of the same field of view of the flow cell, there can be 1,000 polymerase communities, and each polymerase community has a set of sequencing reads that includes an index sequence pair or a single index sequence that uniquely maps the polymerase community to one of the 80 different samples.

[0092] In some embodiments, the sequencing data for each sequencing run can include multiple sets of sequencing reads. Each set of sequencing reads can correspond to one or more polymerase communities or clusters. Each set of sequencing reads can correspond to the library molecules or library-clamp complexes disclosed herein.

[0093] Each set of sequencing reads can include sequencing data from multiple sequencing cycles and multiple channels. For example, in a short-read sequencing run, the total number of cycles can be any non-zero integer less than 150 or 200. The total number of cycles in a run can be any non-zero integer. Each set of sequencing reads can include the total number of cycles in the sequencing run. However, while the sequencing run is still in progress, each set of sequencing reads only includes the sequencing results corresponding to the sequencing cycles that have been completed in the sequencing run. For example, after the sequencing cycles corresponding to the first index sequence (e.g., index 1) have been completed and the cycles corresponding to the second index sequence (e.g., index 2) have not been executed, the set of sequencing reads can include the first index sequence but not the second index sequence.

[0094] In some embodiments, each set of sequencing reads may include, but is not limited to: reads of a fragment of a DNA sequence (i.e., read 1); reads of the complementary strand of the fragment (i.e., read 2); one or more index sequences (e.g., index 1 and / or index 2) or combinations thereof. In some embodiments, the fragment of the DNA sequence may be single-stranded such that the set of sequencing reads does not include any reads of the complementary strand of the fragment. In other embodiments, the fragment may be double-stranded and the set of sequencing reads further includes reads of the complementary strand of the fragment.

[0095] In some embodiments, the reads of a fragment of a DNA sequence and the reads of the complementary strand of the fragment of the DNA sequence may include a non-zero number of nucleotide bases. For example, the number may be any integer in the range of 1 to 300. Each set of sequencing reads may include different DNA fragments (alone or in combination with their complementary strands). Two different DNA fragments (optionally, with their corresponding complementary strands) may be from the same sample or two different samples.

[0096] One or more index sequences may include index sequences attached or appended to the 5' or 3' end of the DNA fragment. The attachment may be adjacent to the end of the DNA fragment. Alternatively, the insertion sequence may be attached to the fragment with some nucleotide spacing therebetween, as Figures 6 - 7 shown.

[0097] One or more index sequences may include a first and a second index sequence, each attached to the end of the DNA fragment. The attachment may be adjacent to the end of the fragment. Alternatively, the insertion sequence may be attached to the fragment with some nucleotide spacing therebetween, as Figures 6 - 7 shown.

[0098] In some embodiments, more than one index sequence may be attached to a single end of the fragment, spaced apart from or adjacent to each other.

[0099] Figure 2 Exemplary paired-end sequencing reads are shown. In this particular embodiment, read 1 is the forward read of a DNA fragment (i.e., the insert fragment) from the 5' end to the 3' end. Index 1 is attached to the 3' end of the fragment and index 2 is attached to the 5' end. Read 2 is the reverse read of the complementary strand of the DNA fragment.

[0100] In some embodiments, operation 520 may include an operation of determining whether a first plurality of flow cell images of a plurality of samples in a first plurality of sequencing cycles corresponding to a plurality of index sequences in a sequencing run have been acquired. In response to determining that the first plurality of flow cell images have not been acquired, operation 520 of generating sequencing data is not performed. In response to determining that the first plurality of flow cell images have been acquired, operation 520 of generating sequencing data may be performed.

[0101] In some embodiments, operation 520 may include determining whether the first plurality of sequencing cycles corresponding to a plurality of index sequences in a sequencing run have been completed. In response to determining that the first plurality of cycles have not been completed, the operation 520 of generating sequencing data is not performed. In response to determining that the first plurality of cycles have been completed, the operation 520 of generating sequencing data may be performed.

[0102] In some embodiments, determining whether the first plurality of sequencing cycles corresponding to a plurality of index sequences in a sequencing run have been completed may include actively requesting information from or passively receiving information from a processor of the sequencing system, the information being useful for determining whether the first plurality of sequencing cycles corresponding to a plurality of index sequences have been completed. Such information may include, but is not limited to: the number of cycles of sequencing that have been completed, the current sequencing cycle number, the number of cycles of sequencing corresponding to the index sequences, and the number of index sequences in the library molecules. For example, if the current cycle number is 50, and the cycles corresponding to the first index sequence are cycles 10 to 22, and there is only one index sequence, it may be determined that the first plurality of sequencing cycles corresponding to a plurality of index sequences in the sequencing run have been completed.

[0103] In some embodiments, operation 520 may include determining whether base calling for the first plurality of flow cell images has been performed. Determining whether base calling for the first plurality of flow cell images has been performed may include actively requesting or passively receiving information from a processor of the sequencing system, the information being useful for determining whether the first plurality of sequencing cycles corresponding to a plurality of index sequences have been completed. Such information may include, but is not limited to, the total number of base calls that have been saved or stored. In response to determining that such base calling has not been completed, the operation 520 of generating sequencing data is not performed. In response to determining that base calling for the first plurality of flow cell images has been completed, the operation 520 of generating sequencing data may be performed.

[0104] The operation 520 of generating sequencing data is performed while the sequence run is still in progress and before the sequence run has been completed to enable early sorting and / or separation of the index sequences.

[0105] In some embodiments, operation 520 may include obtaining a first plurality of flow cell images of one or more samples positioned on a flow cell. The first plurality of flow cell images may be obtained by the optical system of the sequencing system 110 disclosed herein. The first plurality of flow cell images may be obtained during a first plurality of sequencing cycles. In some embodiments, the first plurality of sequencing cycles may include cycles that cover at least some or all of the length of one or more index sequences. For example, the first plurality of flow cell images are obtained during at least the first 40 of approximately 150 total sequencing cycles. The first 40 cycles are sufficient to cover at least the total length of index 1 and index 2, which precede the cycles corresponding to DNA fragments (e.g., read 1 or read 2).

[0106] Flow cell images herein may be obtained from one of 1, 2, 3, 4, or more channels of imager 116 using the optical system disclosed herein. Each flow cell image may include one or more tiles (imaging regions), and each tile may be divided into multiple sub-tiles. Each sub-tile may include multiple colonies or clusters. Each sub-tile may include multiple regions, where each region includes a number of colonies. For example, polymerase colonies may be extracted from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, colonies may be extracted from a flow cell image from a single channel. Flow cell images as disclosed herein may be images obtained using the flow cell 112 as Figure 1 shown.

[0107] In some embodiments, the operation of obtaining the first plurality of flow cell images may include passively receiving the flow cell images from or actively requesting the flow cell images from the optical system disclosed herein after the optical system disclosed herein generates or captures the flow cell images. In some embodiments, the optical system is included in the Figure 1 imager 116 therein.

[0108] In some embodiments, the operation of obtaining the first plurality of flow cell images may include obtaining the flow cell images using the optical system.

[0109] Each flow cell image may contain multiple polymerase colonies or clusters as bright spots of different intensities, and each polymerase colony may include a size and / or shape. The flow cell image may contain at least a portion of a sub-tile or tile (imaging region) of the flow cell. The flow cell image may be obtained from two or more channels.

[0110] In some embodiments, each of the plurality of flow cell images may cover at least a portion of a sample immobilized on a support of a flow cell device. Each of the plurality of flow cell images may contain optical signals from a polymerase population of a sample immobilized on the support. In some embodiments, the plurality of flow cell images may contain optical signals emitted from nucleotide reagents that bind to an imbalanced diversity of nucleobases A, G, C, and T / U in a plurality of nucleic acid template molecules in a sample immobilized on the support. During one or more cycles of a sequencing run, an imbalanced diversity of nucleobases A, G, C, and T / U may occur in at least some regions of the flow cell images.

[0111] In some embodiments, method 500 herein advantageously processes optical signals from a sample that may have an imbalanced diversity of nucleobases A, G, C, and T / U during one or more cycles. In some embodiments, method 500 herein advantageously generates base calls from a sample that may have an imbalanced diversity of nucleobases A, G, C, and T / U at a pre-determined quality level. In some embodiments, the imbalanced diversity of the sample comprises the percentage of: (1) the number of one or more types of nucleobases (e.g., the number of polymerase populations or clusters corresponding to nucleobase A in a base call) to (2) the total number of nucleobases in a region of the sample immobilized on the flow cell device (e.g., the number of polymerase populations or clusters corresponding to A, G, C, and T in a base call). During one or more cycles, this percentage may be less than 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or 5%. In some embodiments, a region herein may be any pre-determined area within the field of view of a flow cell image.

[0112] In some embodiments, a region of the sample comprises at least a portion of a sub-tile of the flow cell device. In some embodiments, a region of the sample may comprise the entire field of view of a flow cell image. In some embodiments, a region may be selected from the sample based on pre-determined selection rules. For example, a region may be selected to be a pre-determined size (e.g., 256×256 pixels or 128×128 pixels), and the region contains the central pixel of the flow cell image. In some embodiments, the region may comprise one microfluidic channel of the flow cell device but not other microfluidic channels of the same flow cell device. In some embodiments, the region may comprise a region of various numbers of pixels.

[0113] In some embodiments, the operation of obtaining a first plurality of flow cell images comprises obtaining a plurality of flow cell images from two or more channels at different z-levels. The plurality of flow cell images from different z-levels may be configured to cover some or all of the 3D sample along the z-axis.

[0114] In some embodiments, method 500 includes the operation of aligning or registering flow cell images across different sequencing cycles and / or from different channels to a common coordinate system for subsequent image analysis that can yield base calling or other sequencing results. The common coordinate system can be a reference coordinate system disclosed herein. The common coordinate system can be predetermined. In some embodiments, method 500 includes the operation of registering the flow cell images to one or more template images. After registering the flow cell images, the 2D or 3D coordinates of the polymerase population can be determined, thereby determining the polymerase population or clusters.

[0115] Various methods can be used to register the flow cell images herein, e.g., images from different channels and / or different flow cycles. Exemplary image registration methods are described in PCT patent application No. PCT / US2023 / 067931, the content of which is incorporated herein by reference in its entirety.

[0116] In some embodiments, method 500 includes the operation of performing color correction on flow cell images across different sequencing cycles and / or from different channels. Various methods can be used to perform color correction on the flow cell images herein (e.g., images from different channels and / or different flow cycles). Exemplary color correction methods are described in PCT patent application No. PCT / US23 / 74486, the content of which is incorporated herein by reference in its entirety. In some embodiments, method 500 is configured to process flow cell images across different sequencing cycles and / or from different channels such that base calling can be performed based on the processed image intensities in the flow cell images.

[0117] In some embodiments, operation 520 can include performing one or more preliminary analysis steps on the first plurality of flow cell images and / or the second plurality of flow cell images.

[0118] In some embodiments, one of the preliminary analysis steps can include generating base calling for polymerase populations or clusters in the flow cell images. In some embodiments, the base calling generated using the methods herein can have a predetermined quality level. The predetermined quality level can be not less than Q20, Q25, Q28, Q30, Q35, Q38, Q40, Q45, Q50 or higher in terms of at least one or more cycles of the sequencing run. The predetermined quality level can be an error of not more than 2%, 1%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01% or less in terms of base calling in at least one or more cycles.

[0119] Multiple methods can be used to perform 2D or 3D base calling using the flow cell images herein (e.g., images from different channels and / or different flow cycles). Exemplary 3D base calling methods are described in PCT patent application number PCT / US2023 / 076125, the content of which is incorporated herein by reference in its entirety.

[0120] Each polymerase colony can have base calling of nucleobases (e.g., A, T, C, or G) in a single cycle. Preliminary analysis including but not limited to the steps disclosed herein can be used to generate base calling for a specific cycle. Base calling can be generated after acquiring the flow cell image in that specific cycle. In some embodiments, base calling for a specific cycle may also depend on the foregoing and / or subsequent cycles adjacent thereto, and thus base calling can be generated after capturing flow cell images from such cycles.

[0121] In some embodiments, some preliminary analysis steps can be performed before generating base calling. In some embodiments, some preliminary analysis steps can be performed to ensure the quality of base calling. In some embodiments, performing one or more preliminary analysis steps includes determining a quality score for base calling of multiple polymerase colonies of multiple samples in a first or second plurality of flow cell images.

[0122] In some embodiments, one or more preliminary analysis steps for multiple flow cell images include: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization quality score estimation; adapter trimming; or combinations thereof.

[0123] In some embodiments, one of the preliminary analysis steps can include identifying the centers of clusters or polymerase colonies (which typically form on beads). In some embodiments, the preliminary analysis involves forming a template for the flow cell image, e.g., a polymerase colony map. The template can include the estimated positions of all detected clusters or polymerase colonies in a common coordinate system. The template is generated by identifying the positions of clusters or polymerase colonies in all images in the first few cycles of the sequencing process. Exemplary methods for generating a template image or polymerase colony map are described in U.S. patent application numbers 18 / 078,797 and 18 / 078,820, the content of which is incorporated herein by reference in its entirety.

[0124] In some embodiments, operation 520 may include generating sequencing data based on one or more preliminary analysis steps of the first or second plurality of flow cell images. For example, generating sequencing data may be based on base calling performed in the preliminary analysis. Sequencing data may be generated for any number of sequencing cycles within a run in operation 520. For example, sequencing data may be generated for some or all of the cycles for which base calling has been performed. Additional preliminary analysis steps, such as adapter trimming, may be performed after base calling but before sorting and / or separating the sequencing data for the corresponding cycles.

[0125] Traditionally, sequencing data has been generated after substantially all of the cycles of a sequencing run have been completed (e.g., more than 70%, 80%, or 90% of the total number of cycles). However, it is advantageous to generate sequencing data in parallel while the sequencing run is still being performed to speed up the data analysis process. More importantly, it is advantageous to detect sequencing run problems early and stop the problematic sequencing run before it is completed to reduce or minimize waste of time and resources.

[0126] In some embodiments, method 500 may include an operation of determining an order for reading each set of sequencing reads. In some embodiments, each set of sequencing reads may include: reads of fragments of a DNA sequence, i.e., read 1; reads of the complementary strand of the fragment of the DNA sequence, i.e., read 2; a first index sequence, i.e., index 1; a second index sequence, i.e., index 2, or a combination thereof. Determining the order for reading the sequencing reads may include, when there are at least two index sequences, determining the order for reading the combination of read 1, read 2, index 1, and index 2. In some embodiments, when there are reads 1 and 2 with the same reading order, read 1 is before read 2. In some embodiments, the sequence reads (e.g., the sequence reads stored in a data file) may include one or more contiguous nucleotide base sequences. The order for reading each set of sequencing reads may determine which portion of the nucleotide bases in the contiguous sequence belongs to read 1, read 2, index 1, and / or index 2.

[0127] In some embodiments, the operation of determining the order for reading each set of sequencing reads includes determining the position to place a turn in the reading order such that only the nucleotide bases after the turn are reversed and complemented. An exemplary order for reading sequencing reads having two index sequences may be: index 1, index 2, read 1, turn, and read 2. Another exemplary order is: read 1, read 2, turn, index 1, and index 2. Figure 2 An exemplary reading order is shown as: read 1, turn, read 2, index 1, and index 2.

[0128] In some embodiments, the operation of determining the order in which each set of sequencing reads is read includes determining whether one or more index sequences are located before the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster and / or the reads of the complementary strand of the fragment of the DNA sequence.

[0129] In some embodiments, method 500 includes the operation of determining whether one or more index sequences are located before the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster and / or the reads of the complementary strand of the fragment of the DNA sequence, in the read order of the set of sequencing reads.

[0130] In response to determining that one or more index sequences are located before the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster or the reads of the complementary strand of the fragment of the DNA sequence, method 500 may include performing one or more of operations 520 to 560.

[0131] In some embodiments, method 500 may include the operation of determining whether a first plurality of flow cell images in a first plurality of sequencing cycles corresponding to a plurality of index sequences in a sequencing run have been acquired. In response to determining that the first plurality of flow cell images have been acquired, method 500 may include performing one or more of operations 520 to 560.

[0132] In some embodiments, method 500 herein is at least configured to perform a sequencing run in the following read order: read one or more index sequences (index 1 and / or index 2) before reading the reads of the fragment of the DNA sequence (i.e., read 1) and the reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2).

[0133] In some embodiments, one or more index sequences of a polymerase cluster are separate and different from the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster and / or the reads of the complementary strand of the fragment of the DNA sequence. In some embodiments, each of the one or more index sequences is not a continuous portion of the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster and / or the reads of the complementary strand of the fragment of the DNA sequence. In some embodiments, each of the one or more index sequences is not the same as or similar (within the error tolerance rate) to the continuous portion of the reads of the fragments of the DNA sequence of the set of reads for each polymerase cluster and / or the reads of the complementary strand of the fragment of the DNA sequence.

[0134] In some embodiments, each of one or more index sequences is not a sequence of nucleotide bases that is separate or distinct from a read of a fragment of a DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. Instead, each index sequence can be included in a read of a fragment of a DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. In some embodiments, each index sequence comprises a contiguous portion of a read of a fragment of a DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. In some embodiments, each of one or more index sequences is a part, e.g., a contiguous part, of a read of a fragment of a DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. For example, a read of a fragment of a DNA sequence is GGCTCCTACAATTCCGGAATGAGTGG. The first 9 bases of the read are index 1, and the last 9 bases of the read are index 2. Index 1 and index 2 are each a contiguous part of the read (e.g., R1 of the sequence of interest).

[0135] In some embodiments, method 500 herein is configured to perform a sequencing run on the library molecules with any ordered sequence of nucleotide bases that serves as a unique identifier to associate the library molecules with the sample. For example, method 500 herein is configured to perform a sequencing run where a part of a read of a fragment of a DNA sequence (i.e., read 1) or a part of a read of the complementary strand of the fragment of the DNA sequence (i.e., read 2) can similarly or equivalently serve as an index sequence that acts as a unique identifier.

[0136] In some embodiments, method 500 herein endeavors to further accelerate the sequencing data analysis by, while the sequencing run is still in progress, classifying and / or separating read 1 into a separate data file when the sequencing cycle of read 1 has been completed for each polymerase community, and, while the sequencing run is still in progress, classifying and / or separating read 2 into a separate data file when the sequencing cycle of R2 has been completed. Thus, the classification and / or separation of the sequencing data after the sequencing run is greatly reduced to a minimum.

[0137] In some embodiments, the order of reading sequencing read data can be based on various factors that define a sequencing run. In some embodiments, the order of reading can be based on a particular sequencing instrument 110 and its parameters for running a sequencing analysis. The order of reading can also be based on whether the DNA fragment is paired-end or single-end. In some embodiments, the order of reading can be based on the specific composition of the library molecules and their corresponding hybridization methods. In some embodiments, the order of reading can also depend on the kit configuration used, such as 300 or 150 cycles. In some embodiments, the order of reading can also be based on the way the sequencing run data is processed and stored, e.g., software version.

[0138] Manual determination of the order of reading may require the user to know all the factors that affect the order of reading and derive the order of reading based on all the factors. However, the various factors affecting the order of reading listed herein are not exclusive or fixed. With the use of new kits, instruments, chemicals, and software, it becomes increasingly complex for the user to determine the order of reading in a timely and accurate manner. Errors in determining the order of reading can lead to significant errors in downstream sequencing analysis, and these errors take time and resources to discover and correct. The techniques herein enable automatic determination of the order of reading to ensure faster speed and higher accuracy that cannot be achieved by traditional methods.

[0139] In some embodiments, the operation of determining the order of reading each set of sequencing reads includes: extracting at least one value for each parameter from a parameter file. Each of the parameters can correspond to one or more of the one or more factors that control the order of reading. Some of the factors are disclosed herein. In some embodiments, the operation 520 of determining the order of reading each set of sequencing reads includes: determining the order of reading based on the extracted parameter values. In some embodiments, determining the order of reading based on the extracted parameter values includes: searching for the order of reading in a pre-generated lookup table. In some embodiments, each set of extracted parameter values can be linked to the order of reading in the pre-generated lookup table. In some embodiments, determining the order of reading based on the extracted parameter values includes: determining the order of reading of at least two elements of the sequencing reads based on one or more of the extracted parameter values. For example, the software version can determine whether the index sequence is before read 1. As another example, the prerequisite for the order of reading regarding a particular surface chemistry may require that index 1 be before index 2 in the order of reading. As yet another example, a "turn" may not be placed at the very end of the order of reading after all other elements.

[0140] In some embodiments, method 500 may include an operation of determining whether one or more index sequences are to be reverse complemented. Such determination may be performed by a processor disclosed herein. Such determination may be based on the read order that has been determined in operation 520. For example, in a read order of "read segment 1, read segment 2, turn, index 1, and index 2", both index 1 and index 2 need to be reverse complemented because they are read after the "turn". As another example, in a read order of "index 1, index 2, read segment 1, turn, read segment 2", none of the index sequences need to be reverse complemented.

[0141] In some embodiments, operation 530 may include an operation of determining a value of a mismatch parameter for a polymerase population. Operation 530 may be performed for each polymerase population among a plurality of polymerase populations on a flow cell.

[0142] Operation 530 may include determining whether each of one or more index sequences matches one of a plurality of mapped index sequences and whether it maps to one of a plurality of samples (e.g., a single sample). In response to determining that each of the one or more index sequences matches one of the mapped index sequences, operation 530 may include determining a value of the mismatch parameter for the polymerase population. The value of the mismatch parameter may be in a predetermined unit, e.g., bases. For example, the value of the mismatch parameter may be 1 mismatched base, 2 deleted bases, 15% mismatched bases, or 10% deleted bases.

[0143] In response to determining that there is no match with any of the mapped index sequences, the corresponding polymerase population may be determined as an unassigned polymerase population that cannot be classified into any of the samples. The polymerase populations may be counted and the identities of the polymerase populations may be recorded to determine the unassigned rate of the sequence run.

[0144] Method 500 may further include operation 540 of assigning polymerase populations to the mapped samples. Operation 540 may further include determining whether the value of the mismatch parameter satisfies an error tolerance rate herein, e.g., a second error tolerance rate. The second error tolerance rate may be predetermined. It may be the same as or different from the first error tolerance rate. In response to determining that the value of the mismatch parameter satisfies the second error tolerance rate, operation 540 may include assigning to one of the mapped samples. In response to determining that the value of the mismatch parameter does not satisfy the second tolerance rate, operation 540 may include determining the polymerase population as an unassigned polymerase population. The polymerase populations may be counted and the identities of the polymerase populations may be recorded to determine the unassigned rate of the sequence run.

[0145] In some embodiments, the operation of determining whether the value of the mismatch parameter satisfies the second error tolerance rate includes: determining one or more positions in each index sequence, where each position corresponds to a mismatched base and / or a quality score for base identification that is below a predetermined quality threshold as compared to a reference index sequence. The value of the mismatch parameter for the polymerase population can be the total number of one or more positions.

[0146] For example, sample 17 can be associated with the reference index sequence ACTC, which is represented in binary as: 01001100. Exemplary bit encodings are: C = 00, A = 01, G = 10, T = 11. The index sequence CCTC = 00001100 from the sequencing data meets the matching criteria for sample 17, and it matches one of the mapped index sequences of the reference index sequence ACTC. The mismatched base is in the first position, so the mismatch mask for this sequence is 1000. The corresponding quality scores for the index sequence CCTC are [13, 28, 0, 33]. A quality score of 0 indicates that there is no base identification at the corresponding position in the index sequence. The quality score of 13 is below the quality score threshold of 25. Thus, the quality mask for this particular index sequence is 1010. The total number of positions with mismatched bases or with quality below the threshold is 2, because overlapping positions are only counted once in the total number of positions. In this particular embodiment, considering the mismatched base and the bases with low quality scores, the value of the mismatch parameter is 2. Assuming there are 4 bases in the index sequence, the mismatch percentage is 50%. The index sequence CCTC does not meet the error tolerance rate of not exceeding 25%, so the corresponding polymerase population cannot be recorded as corresponding to sample 17. Instead, this particular polymerase population is considered an unassigned polymerase population to determine the unassigned rate and / or the assigned rate of the total number of polymerase populations.

[0147] In some embodiments, the method 500 disclosed herein can advantageously process sequencing data in binary format rather than processing sequencing reads of nucleobases in formats such as A, T, C, G, etc. In some embodiments, each base is represented by a pair of binary bits to significantly speed up the operations herein. In some embodiments, each base can be a bitwise integer with multiple binary bits. For example, each base can be represented as a 4-bit binary number. In some embodiments, with bases being multi-bit binary numbers, the operations disclosed herein can advantageously be performed using bitwise arithmetic. Compared to the same operations but using different formats of bases, bitwise operations can significantly speed up the operations disclosed herein. Compared to existing methods for sorting and / or separating sequence data (e.g., demultiplexing), bitwise operations can also significantly reduce the computation time.

[0148] In some embodiments, the method 500 herein can include an operation 550 of evaluating one or more statistical parameters of the sequencing data while the sequencing run of multiple samples is in progress.

[0149] Operation 550 may include an operation of calculating at least one value of one or more statistical parameters for a plurality of polymerase communities based on values of mismatch parameters for each polymerase community. In some embodiments, one or more statistical parameters of the sequencing results include one or more of the following: mismatch rate; unassigned rate; error association rate; assignment rate; match rate; deletion rate; insertion rate; and mixed pairing rate.

[0150] The unassigned rate may be the percentage of polymerase communities that do not match any of the mapped index sequences and / or the percentage of polymerase communities that match two mapped index sequences of two or more different samples. In other words, within the error tolerance rate, the percentage of polymerase communities whose index sequences do not match the reference index sequence of a single sample relative to the total number of polymerase communities.

[0151] The assignment rate may be the percentage of polymerase communities whose index sequences match the mapped index sequence of a single sample. In other words, within the error tolerance rate, the percentage of polymerase communities that match the reference index sequence of a single sample relative to the total number of polymerase communities.

[0152] The mismatch rate may be the percentage of polymerase communities that match (within the error tolerance rate) the reference index sequence of a single sample with some mismatched bases relative to the total number of polymerase communities.

[0153] The match rate may be the percentage of polymerase communities that match the reference index sequence of a single sample without errors relative to the total number of polymerase communities.

[0154] The deletion rate may be the percentage of polymerase clones that match (within the error tolerance rate) the reference index sequence of a single sample with a certain deletion relative to the total number of polymerase clones.

[0155] The insertion rate may be the percentage of polymerase communities that match (within the error tolerance rate) the reference index sequence of a single sample with a certain insertion relative to the total number of polymerase communities.

[0156] The mixed pairing rate may be the percentage of polymerase communities that match (within the error tolerance rate) the reference index sequences of more than one sample relative to the total number of polymerase communities.

[0157] Operation 550 may include an operation of comparing a value of a statistical parameter with a predetermined corresponding threshold to determine whether the value meets the predetermined threshold. Two or more statistical parameters may have different or the same thresholds. The threshold may be customized based on a sequence run, characteristics of a sample, or various other factors associated with sequencing analysis. As an example, the mismatch rate may be from about 1% to about 20%. As another example, the deletion rate may be about 1 to 2 bases out of 8 bases, 9 bases, 10 bases or even more bases.

[0158] In some embodiments, operation 550 is performed in parallel with the sequence run while the sequence run is still in progress.

[0159] In some embodiments, operation 550 is performed in parallel with the sequence run while the sequencing system 110 is still acquiring a second plurality of flow cell images of a plurality of samples in a second plurality of sequencing cycles.

[0160] The second plurality of sequencing cycles herein may correspond to the second plurality of flow cell images. The second plurality of sequencing cycles may correspond to at least a portion of the reads of a fragment of the DNA sequence of the set of sequencing reads for each polymerase cluster or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence. The second plurality of flow cell images correspond to the reads of a fragment of the DNA sequence of the set of sequencing reads for each polymerase cluster or the reads of the complementary strand of the fragment of the DNA sequence. In some embodiments, the second plurality of sequencing cycles are after the first plurality of sequencing cycles. In some embodiments, whether or not the index sequence is part of the reads of a fragment of the DNA sequence (i.e., the insert), the second plurality of sequencing cycles or the second plurality of flow cell images do not correspond to one or more index sequences. In some embodiments, the second plurality of sequence cycles correspond only to the reads of a fragment of the DNA sequence of the set of reads for a polymerase cluster. In some embodiments, the second plurality of sequence cycles correspond only to the reads of the complementary strand of a fragment of the DNA sequence of the set of reads for a polymerase cluster.

[0161] The first and / or second plurality of flow cell images may be acquired from one or more color channels. Each flow cell image may contain a plurality of polymerase clusters or clusters. Each flow cell image may contain a sub-tile or tile of one or more flow cells on which one or more samples are held.

[0162] In some embodiments, the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the reads of a fragment of the DNA sequence of the set of sequencing reads for each polymerase cluster or the reads of the complementary strand of the fragment of the DNA sequence. In some embodiments, the first plurality of sequence cycles correspond only to one or more index sequences.

[0163] In some embodiments, operation 550 is performed in parallel with the sequence run while the sequencing system 110 or the processor is still generating at least a portion of the reads of the DNA sequence fragments corresponding to each polymerase colony for the plurality of samples or at least a portion of the reads of the complementary strand of the DNA sequence fragments.

[0164] In some embodiments, operation 540 may include evaluating one or more statistical parameters of the sequencing results while generating multiple sets of sequencing reads, the multiple sets of sequencing reads including reads of DNA sequence fragments (i.e., read 1) and / or reads of the complementary strand of the DNA sequence fragments (i.e., read 2).

[0165] In some embodiments, operation 550 may include evaluating one or more statistical parameters of the sequencing results while the operation of generating multiple sets of sequencing reads is still being performed. Specifically, at least a portion of the multiple sets of sequencing reads has been generated. Such a portion may correspond to some or all of the index sequences in the multiple sets of sequencing reads.

[0166] In some embodiments, operation 550 may include evaluating one or more statistical parameters of the sequencing results while the operation of acquiring a second plurality of flow cell images is still being performed. In other words, the flow cell images in the sequencing cycles corresponding to sequencing some or all of the index sequences have been completed, while the sequencing cycles corresponding to sequencing the DNA fragments have not been completed.

[0167] In some embodiments, operation 550 may include evaluating one or more statistical parameters of the sequencing results after capturing flow cell images corresponding to only some or all of the index sequences in the sequencing cycle, after a preliminary analysis of these flow cell images, and / or after sequencing data corresponding to these flow cell images has been generated.

[0168] In some embodiments, operation 550 of evaluating one or more statistical parameters of the sequencing results includes: calculating at least one value of each statistical parameter using one or more reference index sequences. The reference index sequences may be in their forward or forward complementary direction with minimal or no errors. In some embodiments, evaluating one or more statistical parameters of the sequencing results includes: comparing at least one value of each statistical parameter with a pre-determined threshold. In some embodiments, the statistical parameters of the sequencing results may include, but are not limited to: a mismatch parameter and / or an unassigned parameter.

[0169] The mismatch parameter can indicate the percentage or number of bases that do not match the retrieved reference index sequence out of the total number of bases. For example, a 10% mismatch rate indicates that on average 1 out of every 10 bases does not match across all index sequences. The unassigned parameter can indicate how many index sequences do not match any of the stored reference index sequences at all. For example, 1% unassigned may be due to reasonable sequencing errors. However, 15% unassigned across all index sequences can indicate that some of the index sequences were incorrectly entered or were incorrectly reverse complemented.

[0170] In response to this assessment, if one or more statistical parameters do not meet a predetermined threshold, there may be a problem, e.g., an indexing error or a library molecule problem, which may render the remaining sequence run invalid. Thus, the faulty sequence run can be stopped before it is completed, and the instrument can be made available for other runs.

[0171] In some embodiments, method 500 can include operation 560 of determining whether to terminate a sequence run before it is completed. The determination in operation 560 can be based on the assessment of one or more statistical parameters in operation 550. When the values of one or more statistical parameters meet a stop criterion, the sequence run may be problematic, e.g., there may be a problem with the index sequence and / or library molecules such that completion of the sequence run may waste time or resources. Thus, reading the index sequence before reading any sequence of interest can advantageously save COGS costs, e.g., the cost of reagents used to dehybridize the sequence of interest in a faulty sequence run. Conversely, if the values of one or more statistical parameters do not meet the stop criterion, the sequence run can be continued and completed.

[0172] In some embodiments, operation 560 of determining whether to terminate a sequence run before it is completed includes determining whether the values of one or more statistical parameters meet a stop criterion. Operation 560 includes, in response to determining that the values do not meet the stop criterion, continuing the sequencing run until it is completed. In response to determining that the values do meet the stop criterion, method 500 can include automatically terminating the sequencing run before it is completed, or sending a notification (e.g., an audio or video notification) to the user informing that the sequencing run can be terminated early before it is completed. In some embodiments, the operation of automatically terminating the sequencing run before it is completed includes triggering a termination instruction executable by a processor to be sent to the processor of sequencing system 110. Such an operation can be performed by the processor disclosed herein. Such an operation can be performed by the processor of sequencing system 110 or a processor external to the sequencing system (e.g., on a user's computer).

[0173] In some embodiments, operation 560 of determining whether to terminate a sequence run before completion includes generating sequencing data while the sequencing run is still in progress in response to determining that one or more values of one or more statistical parameters do not meet a stop criterion, where the sequencing data includes reads of fragments of the DNA sequence of the group sequencing reads for each polymerase community, reads of the complementary strand of the fragment of the DNA sequence, or both.

[0174] In some embodiments, the stop criterion can be customized according to various factors, including but not limited to: the characteristics of the sample, the parameters of the sequencing run, the characteristics of the sequencing system and its structural elements, and the underlying chemistry used to generate the library molecules. The stop criterion can be predetermined to optimize the reliable early detection of problematic sequencing runs and minimize premature termination of sequencing runs due to errors. For example, using two or more statistical parameters instead of a single statistical parameter as the stop criterion may be more reliable in the early detection of problematic sequencing runs. As another example, the stop criterion can be that the value of the mismatch rate is not less than 20%, and the value of the deletion rate or insertion rate is not less than 11%.

[0175] In some embodiments, in response to determining that the sequence run can be terminated, method 500 can include the operation of automatically stopping generating sequencing reads. Specifically, method 500 can include automatically stopping generating reads of fragments of the DNA sequence (i.e., read 1), reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2), or both. Method 500 can include automatically terminating the operation of capturing flow cell images in any subsequent cycles that have not yet been completed. Specifically, method 500 can include terminating the capture of flow cell images corresponding to reads of fragments of the DNA sequence (i.e., read 1), reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2), or both.

[0176] As an example, a sequencing run has 120 cycles, and the first 20 cycles contain all cycles corresponding to index 1. Cycles 21 through 30 correspond to index 2 in each sequencing read. After the flow cell images in the first 20 or 30 cycles of the sequencing run have been captured, the flow cell images in the 20 or 30 cycles can be processed, and sequencing data corresponding to index 1 (alone or in combination with index 2) can be generated while still capturing the flow cell images in the subsequent 90 cycles corresponding to DNA fragments (e.g., read 1, alone or in combination with read 2). Sequencing statistics can be calculated as disclosed herein. After determining that the statistical parameters do not meet a predetermined threshold, the user can decide to prematurely terminate the sequencing run at the 50th sequencing cycle to address library preparation issues or index sequence conflicts. After the issues are resolved, the sequence run can be executed again starting from the first sequencing cycle. In this example, time, computational resources, and cost of goods sold (COGS) for running the subsequent 70 sequencing cycles and processing the problematic sequencing data are saved.

[0177] In response to the evaluation, if one or more statistical parameters meet a predetermined threshold, method 500 can include continuing to generate sequencing reads, e.g., read 1 and / or read 2, while concurrently sorting and / or separating index sequences in parallel to save total sequence data analysis time. In some embodiments, method 500 herein includes performing demultiplexing of the portion of the sequencing run that has already been completed while concurrently performing the remaining sequencing run.

[0178] In some embodiments, the stop criteria include: a value of the mismatch rate that is not less than about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

[0179] In some embodiments, the stop criteria include: a value of the unassigned rate that is not less than about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

[0180] In some embodiments, method 500 may include operations of generating sequencing results in multiple data files in a pre-determined data format based on corresponding mapped sample classifications and / or separating polymerase communities. The operation of separating polymerase communities may be based on one or more index sequences corresponding to the polymerase communities and the assignment of the polymerase communities to samples. In some embodiments, a polymerase community may be assigned to only a single sample. The assignment of a single polymerase community to multiple samples may be regarded as a collision error, and the corresponding polymerase community may be regarded as an unassigned polymerase community and not assigned to any of the multiple samples.

[0181] In some embodiments, the sequencing data generated using method 500 herein may be used for informational purposes. For example, in operation 560, the sequencing data of the index sequences is used to determine whether to terminate the sequence run. Conversely, in some embodiments, the sequencing data containing index sequences (alone or in combination with its quality information) generated using method 500 herein may be regarded as part of the sequencing data that is conventionally to be sorted and separated after the sequence run is completed. Thus, when the sequence run is still in progress, data of the index sequences is generated using method 500 herein and / or classification and separation of such data is performed. In some embodiments, compared with conventional methods, classification and / or separation of the index sequences do not need to be repeated after the sequence run ends to speed up sequencing data analysis.

[0182] In some embodiments, method 500 may include an operation of separating, upon completion of a sequencing run, the reads of the fragments of the DNA sequence of the set of sequencing reads for each polymerase community, the reads of the complementary strand of the fragment of the DNA sequence, or both, without separating one or more of the multiple index sequences in the set of sequencing reads.

[0183] In some embodiments, method 500 may include an operation of generating sequencing results in a pre-determined data format using one or more index sequences. For example, the index sequence may be determined based on: 1) determining what the reference index sequence is according to retrieving the reference index sequence; and also based on 2) whether the reference index sequence needs to be reverse complemented, depending on whether the index sequence is before the "turn" of the read order. The determined index sequence may then be used to extract the reads of the DNA fragment (e.g., read 1 and / or read 2) and save them in a single data file or multiple data files. The data files may be in a pre-determined format.

[0184] In some embodiments, method 500 may include separating multiple sets of sequencing reads based on one or more index sequences and their determination, to generate sequencing results in a predetermined data format in multiple files. In some embodiments, the method may include demultiplexing the set of sequencing reads and saving them separately into different data files.

[0185] In some embodiments, the predetermined data format includes the FastQ format. In some embodiments, the predetermined data format includes a text-based format. In some embodiments, the predetermined data format includes the ASCII format. In some embodiments, the predetermined data format includes an 8-bit encoding format.

[0186] In some embodiments, some or all of the operations herein may be performed on the sequencing data of a sequencing run after only a part of the specific sequencing run is completed. Specifically, the part that is only completed may correspond to at least some or all of the cycles of one or more index sequences.

[0187] In some embodiments, method 500 may include the operation of generating library molecules 700 using one or more index sequences. Each library molecule 700 may contain an insert fragment 710, that is, the sequence of interest from the sample, for example, R1 or R2. Figures 6 - 7 An exemplary linear library molecule 700 disclosed herein is shown. The linear library molecule may hybridize with a splint molecule. The splint molecule herein may be single-stranded or double-stranded. The splint molecule can be used as an amplification primer for a rolling circle amplification reaction. In some embodiments, the linear library module herein may perform a ligation reaction to close a single nick, thereby forming a covalently closed circular library molecule, which may hybridize with a splint molecule (for example, a single-stranded splint molecule (not shown)), where the splint molecule is used as an amplification primer to perform a rolling circle amplification reaction. In some embodiments, a linear single-stranded library molecule (700) may hybridize with a double-stranded splint molecule, thereby circularizing the library molecule to form a library-splint complex (800) with two nicks. Figures 8A - 8B An exemplary library-splint complex 800 is shown, which performs a ligation reaction to close the nicks, thereby forming a covalently closed circular library molecule (900), which may hybridize with a double-stranded splint molecule. The splint molecule can be used as an amplification primer for a rolling circle amplification reaction. The dashed line represents the nascent extension product.

[0188] Method for Index Sequence Determination

[0189] Traditionally, the manual determination and entry of index sequences and / or their corresponding reverse complementary sequences have been a prerequisite for sorting and separating sequencing data from different samples in a single sequence run. There is a need to automatically determine index sequences at a faster rate and with higher accuracy, especially when sequencing a large number of samples in each sequence run. With the emergence of new kits, sequencing instruments, and sequencing chemistries, there is an urgent need to eliminate the manual determination of index sequences, which is based on increasingly complex factors (including, for example, kits, instruments, sequencing software versions, and specific sequencing chemistries) and is highly error-prone.

[0190] In NGS data analysis, it may be necessary for the user to manually enter an index sequence or its corresponding reverse complementary sequence based on various factors in order to sort and separate sequencing data from different samples into data files. Errors made by the user when entering the index sequence can cause errors when generating sequencing results based on that index sequence and waste a significant amount of time and resources, such as hours or more, when performing any downstream analysis based on the incorrect data file. Additionally, regenerating the data file to remove the error is also time-consuming and resource-intensive. As new kits, sequencing instruments, and sequencing chemistries become available, the manual determination of index sequences based on increasingly complex factors becomes increasingly challenging for the user. The manual entry of index sequences may require a high level of proficiency or expertise in sequencing data analysis and additional time and effort to ensure accuracy and avoid errors that propagate in any subsequent analysis. The techniques disclosed herein can be used for the automatic determination of index sequences such that sequencing data can be sorted and separated for downstream analysis at a faster rate and with higher accuracy. These techniques can be compatible with new kits, sequencing instruments, and sequencing chemistries.

[0191] The techniques disclosed herein advantageously determine the order of reading DNA fragments, their optional complementary strands, and one or more index sequences. Based on the read order, the techniques herein can determine whether an index sequence needs to be reverse-complemented. This determination is based on various factors that can be increasingly complex and time-consuming for the user to handle and thus highly error-prone when performed manually. The techniques disclosed herein are rooted in the characteristics of index-based sequencing analysis. Additionally, the automatic determination of index sequences prevents waste of computational time and resources by advantageously enabling early detection of index sequence errors or library problems. Specifically, the techniques herein can allow for parallel checking of sequencing statistics while a sequencing run is being performed. Thus, the techniques herein can stop a problematic sequence run before it is completed, thereby preventing waste of computational time and resources and freeing up the sequencing system for other sequencing tasks.

[0192] Figure 5BA flowchart of a computer-implemented method 501 for automatic index sequence determination in NGS data analysis is shown. Method 501 may include some or all of the operations disclosed herein. The operations may be performed in an order that is not limited to that described herein.

[0193] Method 501 may be executed by one or more processors disclosed herein. In some embodiments, the processor may include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU), a graphics processing unit (GPU), or a neural processing unit (NPU). The integrated circuit may include a chip such as a field-programmable gate array (FPGA). In some embodiments, the processor may include computing system 400.

[0194] In some embodiments, some or all of the operations in method 501 may be executed by an FPGA. In an embodiment, when some operations are executed by an FPGA, the data after the operations are executed by the FPGA may be transmitted by the FPGA to other devices, such as a CPU or an NPU, so that the CPU or NPU may use such data to execute subsequent operations in method 501. Similarly, data may also be transmitted from other devices (such as a CPU) to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 501 may be executed by a CPU. Alternatively, the operations executed by the CPU may be executed by other processors (such as a dedicated processor) or an NPU. In some embodiments, all of the operations in method 501 may be executed by an FPGA.

[0195] In some embodiments, method 501 is configured to align or register flow cell images across different sequencing cycles and / or from different channels with a common coordinate system. The common coordinate system may be the reference coordinate system disclosed herein. The common coordinate system may be predefined.

[0196] The flow cell image may be obtained from one of one, two, three, four, or more channels of imager 116 using the optical system disclosed herein. Each flow cell image may include one or more tiles (imaging regions), and each tile may be divided into multiple sub-tiles. Each sub-tile may include multiple clusters. Each sub-tile may include multiple regions, where each region includes a number of clusters. For example, polymerase clusters may be extracted from corresponding regions of flow cell images of four different channels in a given cycle. As another example, clusters may be extracted from a flow cell image from a single channel. The flow cell image as disclosed herein may be an image obtained using a flow cell 112 as Figure 1 shown.

[0197] In some embodiments, method 501 is configured to process flow cell images across different sequencing cycles and / or from different channels for base calling. Method 501 can be configured to process flow cell images even if the polymerase populations are diversity imbalanced in one or more cycles and / or fields of view.

[0198] In some embodiments, method 501 is performed during or after cycle N that is different from a reference cycle. A template image (e.g., polymerase population map) can be generated in the reference cycle, and polymerase populations from one or more channels within the reference cycle can be included in the template image in a reference coordinate system while the flow cell image for cycle N or after cycle N has not been captured or is currently being captured. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150. In some embodiments, to facilitate early determination of index sequences and possible errors, N can be any number after some or all of the index sequence sequencing cycles. For example, N can be 20, 30, or 40. As another example, N can be any integer from 1 to 300 or from 1 to 400.

[0199] In some embodiments, method 501 is performed during cycle N while sequencing and image acquisition in a subsequent cycle (e.g., cycle N+1) are being performed or have not been performed. In some embodiments, method 501 is performed in parallel with the sequence run to advantageously reduce the total time for sequencing and preliminary analysis. In some embodiments, method 501 is performed in parallel with the sequence run to advantageously reduce the storage space required to store flow cell images.

[0200] The reference coordinate system can be the common coordinate system disclosed herein. The common coordinate system can be predefined. The common coordinate system can be a Cartesian coordinate system. A variety of other coordinate systems can also be used. Other coordinate systems can include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.

[0201] The flow cell images in this document can be acquired from one, two, three, four, or more channels of imager 116 using the optical system disclosed herein. In some embodiments, multiple flow cell images are acquired during a single flow cycle or multiple flow cycles in a sequence run. In some embodiments, the flow cell images are acquired during the first 5, 10, 15, 20, 30, 50, 80, or 100 cycles of a sequence run. Each flow cell image can include one or more tiles (imaging regions), and each tile can be divided into multiple sub-tiles. Each sub-tile can include multiple communities. Each sub-tile can include multiple regions, where each region includes a number of communities. For example, polymerase communities can be extracted from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, communities can be extracted from a flow cell image from a single channel. The flow cell images disclosed herein can be images acquired using flow cell 112 as shown in Figure 1 The image obtained using flow cell 112 as shown.

[0202] Flow cell 112 can contain a sample immobilized thereon. The sample can include multiple nucleic acid template molecules. The sample can include a two-dimensional (2D) sample or a three-dimensional (3D) volume sample. The nucleic acid template molecules can be randomly distributed or distributed in various patterns on flow cell 112. In some embodiments, multiple communities or clusters in this document can be extracted from a specific region (e.g., each sub-tile) of a tile. In the case of each sub-tile, the communities can be extracted according to a predetermined pattern or randomly.

[0203] In some embodiments, the communities or clusters being sequenced in the flow cycle can have a certain nucleotide diversity, e.g., in base calling. Method 500 can allow for the early separation and classification of nucleotide base sequences even if the polymerase communities or clusters in the sequencing cycle are of low diversity or unbalanced diversity. The nucleotide diversity of a nucleotide acid molecule population (e.g., a community or a cluster) can refer to the relative proportions of nucleotides A, G, C, and T / U present in each flow cycle. The relative proportions of nucleotides can be within the field of view or across the entire flow cell image. Optimal high or balanced diversity data typically can have roughly equal proportions of all four nucleotides represented in each flow cycle of a sequencing run. Low or unbalanced diversity data typically can include high proportions of certain nucleotides and low proportions of other nucleotides in some flow cycles of a sequencing run, e.g., less than 10% of the total of all 4 nucleotides. Thus, an image corresponding to a high proportion of certain nucleotides can have more signal points (communities or clusters) compared to an image corresponding to a low proportion of certain nucleotides. As an example of low or unbalanced diversity data, in a certain flow cycle, bases A, T, C, G can be approximately 1%, 2%, 1%, and 95% of the total community, respectively. Subsequently, in that particular flow cycle, the flow cell image from the channels corresponding to A, T, and C is darker than the flow cell image corresponding to nucleotide G, and the number of polymerase chain reaction (polony) or clusters is much less. As another example of low diversity or unbalanced diversity data, bases A, T, C, G in the polymerase chain reaction over multiple flow cycles can be approximately 2%, 5%, 10%, and 83%, respectively. In embodiments where low or unbalanced diversity data is present in a particular cycle and the data is imaged for sequencing analysis, the use of prior art image registration may fail because the images from one or more channels are too dark compared to the images obtained from other channels (e.g., the signal points of the polymerase communities are too sparse and / or dim), causing problems in subsequent color correction. Additionally, in embodiments where low or unbalanced diversity data is present in a particular cycle, the use of prior art image registration, color correction, and subsequent base identification may fail because the images from one or more channels are too dark (e.g., the signal points of the polymerase communities are too sparse and / or dim). In some embodiments, method 501 is configured to perform automatic index sequence determination even if the polymerase communities or clusters in the flow cell image are of unbalanced nucleotide diversity. In some embodiments, method 501 is configured to perform early classification and separation of index sequences based on the flow cell image at a pre-determined quality level even if the polymerase communities or clusters in the flow cell image are of unbalanced nucleotide diversity. In terms of base identification in one or more cycles, the pre-determined quality level can be no less than Q20, Q25, Q28, Q30, Q35, Q38, Q40 or higher.In terms of base calling in one or more cycles, a predetermined quality level can be an error of no greater than 2%, 1%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01% or less.

[0204] In addition to base bias that affects diversity, complexity is also a factor affecting the existing preliminary analysis methods that perform base calling and the foregoing steps leading to base calling. The methods herein allow for accurate and reliable pre-classification of sequencing data and its pre-separation from low-complexity data. Generally, complexity can indicate the source of the sample. A singleplex sample can contain DNA fragments or molecules from the same sample region or the same sample source in the genome. A multiplex sample can contain DNA fragments or molecules from different sample sources (e.g., liver, kidney, heart, cancerous tissue, etc.) or from one or more sample regions in the genome. When the complexity is below a certain number (e.g., 8 or 16), the signal may be of low complexity. For example, in a 2-cycle sequence, the entire polymerase population is AT, TG, GC, or CA in two adjacent cycles. Each of the bases A, T, C, and G accounts for 25% of the total number of bases in the cycle, but its complexity is less than 8, and the sequence is not all random. In some embodiments, method 501 is configured to perform automatic index sequence determination at a predetermined quality level even if the polymerase population or cluster is of low complexity.

[0205] In some embodiments, method 501 can include operation 511 of generating sequencing data from one or more sequencing runs. Operation 511 can be performed by sequencing system 110 or some of its elements, as disclosed herein. In some embodiments, the sequencing data for each sequencing run can contain multiple sets of sequencing reads. Each set of sequencing reads can correspond to one or more polymerase populations or clusters.

[0206] Each sequencing run herein can contain multiple sequencing cycles. For example, in a short-read sequencing run, the total number of cycles can be any non-zero integer less than 150 or 200. The total number of cycles in a run can be any non-zero integer.

[0207] Each set of sequencing reads can include: reads of a fragment of a DNA sequence (i.e., read 1); reads of the complementary strand of the fragment (i.e., read 2); one or more index sequences (e.g., index 1, index 2) or combinations thereof. In some embodiments, the fragment of the DNA sequence can be single-stranded such that the set of sequencing reads does not include any reads of the complementary strand of the fragment. In other embodiments, the fragment can be double-stranded and the set of sequencing reads includes reads of the complementary strand of the fragment. In some embodiments, the reads of the fragment of the DNA sequence and the reads of the complementary strand of the fragment of the DNA sequence can contain a non-zero number of nucleotide bases. For example, the number can be any integer in the range of 1 to 300. Each set of sequencing reads can contain different DNA fragments (alone or in combination with their complementary strands). Two different DNA fragments (optionally, with their corresponding complementary strands) can be from the same sample or two different samples.

[0208] One or more index sequences can include a single index sequence attached to the 5' or 3' end of the DNA fragment. The attachment can be adjacent to the end of the fragment. Alternatively, a spacer sequence can be attached to the fragment with some nucleotide spacing therebetween.

[0209] One or more index sequences can contain a first and a second index sequence, each attached to the end of the fragment. The attachment can be adjacent to the end of the fragment. Alternatively, a spacer sequence can be attached to the fragment with some nucleotide spacing therebetween, as Figure 2 shown.

[0210] In some embodiments, there can be more than one index sequence attached to a single end of the fragment, spaced apart from or adjacent to each other.

[0211] Figure 2 Exemplary paired-end sequencing reads are shown. In this particular embodiment, read 1 is the forward read of a DNA fragment (i.e., the insert fragment) from the 5' end to the 3' end. Index 1 is attached to the 3' end of the fragment and index 2 is attached to the 5' end. Read 2 is the reverse read of the complementary strand of the DNA fragment.

[0212] In some embodiments, operation 511 can include obtaining an image of the flow cell with one or more samples located thereon. The flow cell image can be obtained by the optical system of the sequencing system 110 disclosed herein. The flow cell image can be obtained during multiple sequencing cycles. In some embodiments, the number of cycles can include cycles that cover at least some or all of the lengths of one or more index sequences. For example, the flow cell image is obtained during at least the first 40 of about 150 total sequencing cycles. The first 40 cycles are sufficient to cover at least the total lengths of index 1 and index 2, which precede the cycles corresponding to the DNA fragments.

[0213] The flow cell images herein can be obtained from one of one, two, three, four, or more channels of imager 116 using the optical system disclosed herein. Each flow cell image can include one or more tiles (imaging regions), and each tile can be divided into multiple sub-tiles. Each sub-tile can include multiple colonies or clusters. Each sub-tile can include multiple regions, where each region includes a number of colonies. For example, polymerase colonies can be extracted from corresponding regions of flow cell images of four different channels in a given cycle. As another example, colonies can be extracted from a flow cell image from a single channel. The flow cell images disclosed herein can be images obtained using a flow cell 112 as Figure 1 shown.

[0214] In some embodiments, the operation of obtaining the first plurality of flow cell images can include passively receiving the flow cell images from or actively requesting the flow cell images from the optical system disclosed herein after the flow cell images are generated or captured by the optical system disclosed herein. In some embodiments, the optical system is included in the Figure 1 imager 116 in.

[0215] In some embodiments, the operation of obtaining the first plurality of flow cell images can include obtaining the flow cell images using the optical system.

[0216] Each flow cell image can contain multiple polymerase colonies or clusters as bright spots of different intensities, and each polymerase colony can include a size and / or shape. The flow cell image can contain at least a portion of a sub-tile or tile (imaging region) of the flow cell. The flow cell image can be obtained from two or more channels.

[0217] In some embodiments, each of the plurality of flow cell images can cover at least a portion of a sample immobilized on a support of a flow cell device. Each of the plurality of flow cell images can contain an optical signal from polymerase colonies of a sample immobilized on the support. In some embodiments, the plurality of flow cell images can contain optical signals emitted from nucleotide reagents that bind to an imbalanced diversity of nucleobases A, G, C, and T / U in multiple nucleic acid template molecules in a sample immobilized on the support. During one or more cycles of a sequence run, an imbalanced diversity of nucleobases A, G, C, and T / U may occur in at least some regions of the flow cell image.

[0218] In some embodiments, the method 501 herein advantageously processes optical signals from samples that can have an imbalanced diversity of nucleotide bases A, G, C, and T / U in one or more cycles. In some embodiments, the method 501 herein advantageously generates base calls from samples that can have an imbalanced diversity of nucleotide bases A, G, C, and T / U in one or more cycles at a pre-determined quality level. In some embodiments, the imbalanced diversity of the sample comprises the percentage of the following items: (1) the number of one or more types of nucleotide bases (e.g., the number of polymerase colonies or clusters corresponding to nucleotide base A in a base call) and (2) the total number of nucleotide bases in the region of the sample immobilized on the flow cell device (e.g., the number of polymerase colonies or clusters corresponding to A, G, C, and T in a base call). In one or more cycles, this percentage can be less than 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or 5%. In some embodiments, the region herein can be any pre-determined region within the field of view of the flow cell image.

[0219] In some embodiments, the region of the sample comprises at least a portion of a sub-tile of the flow cell device. In some embodiments, the region of the sample can comprise the entire field of view of the flow cell image. In some embodiments, the region can be selected from the sample based on pre-determined selection rules. For example, the region can be selected to be a pre-determined size (e.g., 256×256 pixels or 128×128 pixels), and the region comprises the central pixel of the flow cell image. In some embodiments, the region can comprise one microfluidic channel of the flow cell device but not other microfluidic channels of the same flow cell device. In some embodiments, the region can comprise regions of various numbers of pixels.

[0220] In some embodiments, the operation of obtaining the first plurality of flow cell images comprises obtaining a plurality of flow cell images from two or more channels at different z levels. The plurality of flow cell images from different z levels can be configured to cover some or all of the 3D sample along the z-axis.

[0221] In some embodiments, method 500 includes the operation of aligning or registering flow cell images across different sequencing cycles and / or from different channels to a common coordinate system for subsequent image analysis that can yield base calls or other sequencing results. The common coordinate system can be the reference coordinate system disclosed herein. The common coordinate system can be pre-determined. In some embodiments, method 500 includes the operation of registering the flow cell images to one or more template images. After registering the flow cell images, the 2D or 3D coordinates of the polymerase colonies can be determined, thereby determining the polymerase colonies or clusters.

[0222] Various methods can be used to register the flow cell images herein, e.g., images from different channels and / or different flow cycles. Exemplary image registration methods are described in PCT patent application No. PCT / US2023 / 067931, the content of which is incorporated herein by reference in its entirety.

[0223] In some embodiments, method 501 includes an operation of performing color correction on flow cell images across different sequencing cycles and / or from different channels. Various methods can be used to perform color correction on the flow cell images herein (e.g., images from different channels and / or different flow cycles). Exemplary color correction methods are described in PCT patent application No. PCT / US23 / 74486, the content of which is incorporated herein by reference in its entirety. In some embodiments, method 501 is configured to process flow cell images across different sequencing cycles and / or from different channels such that base calling can be performed based on the processed image intensities in the flow cell images.

[0224] In some embodiments, operation 511 may include performing one or more preliminary analysis steps on the flow cell image.

[0225] In some embodiments, the base calling generated using the methods herein may have a predetermined quality level. In terms of at least one or more cycles of a sequencing run, the predetermined quality level may be not less than Q20, Q25, Q28, Q30, Q35, Q38, Q40, Q45, Q50 or higher. In terms of base calling in at least one or more cycles, the predetermined quality level may be an error of not more than 2%, 1%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01% or less.

[0226] Multiple methods can be used to perform 2D or 3D base calling using the flow cell images herein (e.g., images from different channels and / or different flow cycles). Exemplary 3D base calling methods are described in PCT patent application No. PCT / US2023 / 076125, the content of which is incorporated herein by reference in its entirety.

[0227] In some embodiments, one of the preliminary analysis steps may include generating base calling for polymerase colonies or clusters in the flow cell image. Each polymerase colony can have a base call of a nucleotide base (e.g., A, T, C, or G) in a single cycle. Preliminary analysis including but not limited to the steps disclosed herein can be used to generate base calling for a specific cycle. The base calling for a specific cycle can be generated after acquiring the flow cell image in that specific cycle. In some embodiments, the base calling for a specific cycle may also depend on its immediately preceding and / or subsequent cycles, and thus the base calling can be generated after capturing the flow cell images from such cycles.

[0228] In some embodiments, some preliminary analysis steps may be performed before base calling. In some embodiments, some preliminary analysis steps may be performed to ensure the quality of base calling. In some embodiments, one or more preliminary analysis steps on multiple flow cell images include: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization quality score estimation; adapter trimming; or combinations thereof.

[0229] In some embodiments, one of the preliminary analysis steps may include identifying the centers of clusters or polymerase colonies (which typically form on beads). In some embodiments, the preliminary analysis involves forming a template for the flow cell image, e.g., a polymerase colony map. The template may include the estimated positions of all detected clusters or polymerase colonies in a common coordinate system. The template is generated by identifying the positions of clusters or polymerase colonies in all images in the first few cycles of the sequencing process. Exemplary methods for generating template images and / or polymerase colony maps are described in U.S. Patent Application Nos. 18 / 078,797 and 18 / 078,820, the contents of which are incorporated herein by reference in their entirety.

[0230] In some embodiments, operation 511 may include generating sequencing data based on one or more preliminary analysis steps on multiple flow cell images. For example, generating sequencing data may be based on base calling performed in the preliminary analysis. Sequencing data may be generated for nearly all cycles of a run. Sequencing data may be generated for some or all cycles for which base calling has been performed. Additional preliminary analysis steps, such as adapter trimming, may be performed after base calling but before determining the index sequence.

[0231] Traditionally, sequencing data has been generated after nearly all cycles of a sequencing run have been completed. However, it is advantageous to generate sequencing data in parallel while the sequencing run is still being performed to speed up the data analysis process. More importantly, it is advantageous to detect sequencing run problems early and stop a problematic sequencing run before it is completed to reduce or minimize waste of time and resources.

[0232] In some embodiments, a large number of libraries or samples (e.g., 100, 200, or more samples) can be pooled and sequenced simultaneously during a single sequencing run. Each DNA fragment, whether or not it has its complementary strand, can be from a different sample. In some embodiments, each DNA fragment (alone or in combination with its complementary strand) can be uniquely identified by one or more index sequences. The one or more index sequences serve as a unique identifier for each set of sequencing reads. In some embodiments, even if there are some errors in one or more of the index sequences, e.g., 1 or 2 incorrect out of 10 bases in each index sequence, the index sequence can still be used to uniquely identify the DNA fragment, i.e., the one or more index sequences can uniquely identify the DNA fragment within a pre-determined error tolerance rate. The pre-determined error tolerance rate can be about 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. The error tolerance rate can be customized based on various factors such as the length of the index sequence and / or the characteristics of the sample. Figure 3 Table 1 in Figure 3 shows 20 exemplary index sequence pairs that can be used to uniquely identify 20 different samples or DNA fragments.

[0233] When there are two or more index sequences, the lengths of two different index sequences can be the same or different. In some embodiments, each of the one or more index sequences contains any number of nucleotide bases less than 100. For example, the index sequences herein can contain 1 to 100, 2 to 80, or 2 to 75 bases. In some embodiments, the index sequence can contain 6 to 12, 7 to 13, 8 to 12, 6 to 16, or 8 to 16 nucleotide bases.

[0234] In some embodiments, method 501 may include an operation of generating one or more index sequences for each set of sequencing reads. Generating the index sequences may be based on the number of samples or DNA fragments being sequenced in a sequencing run, e.g., 120 different samples. Generating the index sequences may include determining the percentage of occurrences of each of the four nucleotide bases (i.e., A, T, C, G) in all of the index sequences. In other words, generating the index sequences may include determining the diversity of the index sequences. For example, the index sequences may have high diversity such that when combined together, each of the four bases appears in each of the index sequences or in all of the index sequences at approximately 25%. Generating reference index sequences may include determining the length of each index sequence. The length of the index sequences may be based on the number of samples and / or the diversity of the index sequences. In some embodiments, generating the index sequences may include determining a plurality of ordered sequences, each sequence having a preset number of nucleotide bases. In the case where there is one, two, or even more nucleotide errors in an ordered sequence, the ordered sequence is still different from all of the other determined sequences. For example, as Figure 3 shown, when there is an error in index 1 of example_sample_1, instead of the correct sequence "GGCTCCTAC", it becomes "AGCTCCTAC". The incorrect index 1 is still unique compared to other index 1 sequences.

[0235] After generating the index sequences, method 501 may include an operation of storing one or more index sequences. In some embodiments, each of the index sequences may be stored corresponding to a unique identification number, e.g., as Figure 3 shown, "example_sample 1". The stored index sequences may later be retrieved as references to calculate statistical parameters for sequencing analysis as disclosed herein.

[0236] In some embodiments, one or more index sequences of a polymerase population are separated and different from the reads of the DNA sequence fragments in the set of reads for each polymerase population or the reads of the complementary strand of the DNA sequence fragment. In some embodiments, each of the one or more index sequences is not a continuous portion of the reads of the DNA sequence fragments in the set of reads for each polymerase population or the reads of the complementary strand of the DNA sequence fragment. In some embodiments, each of the one or more index sequences is not identical or similar (within an error tolerance rate) to a continuous portion of the reads of the DNA sequence fragments in the set of reads for each polymerase population or the reads of the complementary strand of the DNA sequence fragment.

[0237] In some embodiments, each of one or more index sequences is not a sequence of nucleotide bases that is separate or different from a read of a fragment of the DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. Instead, each index sequence can be included in a read of a fragment of the DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. In some embodiments, each index sequence comprises a contiguous portion of a read of a fragment of the DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. In some embodiments, each of one or more index sequences is a part, e.g., a contiguous portion, of a read of a fragment of the DNA sequence in the set of reads for each polymerase community or a read of the complementary strand of the fragment of the DNA sequence. For example, the read of the fragment of the DNA sequence is GGCTCCTACAATTCCGGAATGAGTGG. The first 9 bases of the read are Index 1, and the last 9 bases of the read are Index 2. Index 1 and Index 2 are each a contiguous portion of the read (e.g., R1 of the sequence of interest).

[0238] In some embodiments, method 501 may include operation 521 of determining the order of reading each set of sequencing reads. In some embodiments, each set of sequencing reads may include: a read of a fragment of the DNA sequence, i.e., Read 1; a read of the complementary strand of the fragment of the DNA sequence, i.e., Read 2; a first index sequence, i.e., Index 1; a second index sequence, i.e., Index 2; or a combination thereof. Determining the order of reading the sequencing reads may include, when there are at least two index sequences, determining the order of reading the combination of Read 1, Read 2, Index 1, and Index 2. In some embodiments, when Read 1 and Read 2 have the same reading order, Read 1 is before Read 2. In some embodiments, the sequence reads (e.g., the sequence reads stored in a data file) may include one or more contiguous nucleotide base sequences. The order of reading each set of sequencing reads may determine which portion of the nucleotide bases in the contiguous sequence belongs to Read 1, Read 2, Index 1, and / or Index 2.

[0239] In some embodiments, operation 521 includes determining the position to place a turn in the reading order such that only the nucleotide bases after the turn are reversed and complemented. An exemplary order of reading the sequencing reads having two index sequences can be: Index 1, Index 2, Read 1, turn, and Read 2. Another exemplary order is: Read 1, Read 2, turn, Index 1, and Index 2. Figure 2 An exemplary reading order is shown as: Read 1, turn, Read 2, Index 1, and Index 2.

[0240] In some embodiments, the order of reading sequencing read data can be based on various factors that define a sequencing run. In some embodiments, the order of reading can be based on a particular sequencing instrument 110 and its parameters for running a sequencing analysis. The order of reading can also be based on whether the DNA fragments are paired-end or single-end. In some embodiments, the order of reading can be based on the specific composition of the library molecules and their corresponding hybridization methods. In some embodiments, the order of reading can also depend on the kit configuration used, such as 300 or 150 cycles. In some embodiments, the order of reading can also be based on the way the sequencing run data is processed and stored, e.g., the software version.

[0241] Manual determination of the order of reading may require the user to know all the factors that affect the order of reading and derive the order of reading based on all the factors. However, the various factors affecting the order of reading listed herein are not exclusive or fixed. With the use of new kits, instruments, chemicals, and software, it becomes increasingly complex for the user to determine the order of reading in a timely and accurate manner. Errors in determining the order of reading can lead to significant errors in downstream sequencing analysis, and these errors take time and resources to discover and correct. The techniques herein enable the automatic determination of the order of reading to ensure faster speed and higher accuracy that cannot be achieved by traditional methods.

[0242] In some embodiments, operation 521 of determining the order of reading each set of sequencing reads includes: extracting at least one value for each parameter from a parameter file. Each of the parameters can correspond to one or more of the one or more factors that control the order of reading. Some of the factors are disclosed herein. In some embodiments, operation 521 of determining the order of reading each set of sequencing reads includes: determining the order of reading based on the extracted parameter values. In some embodiments, determining the order of reading based on the extracted parameter values includes: searching for the order of reading in a pre-generated lookup table. In some embodiments, each set of extracted parameter values can be linked to the order of reading in the pre-generated lookup table. In some embodiments, determining the order of reading based on the extracted parameter values includes: determining the order of reading of at least two elements of the sequencing reads based on one or more of the extracted parameter values. For example, the software version can determine whether the index sequence is before read 1. As another example, the prerequisite for the order of reading regarding a particular surface chemistry may require index 1 to be before index 2 in the order of reading. As yet another example, a "turn" may not be placed at the very end of the order of reading after all other elements.

[0243] In some embodiments, method 501 may include an operation 531 of determining whether one or more index sequences need to be reverse complemented. This determination may be performed by a processor disclosed herein. This determination may be based on the read order determined in operation 521. For example, in a read order of "read segment 1, read segment 2, turn, index 1, and index 2", both index 1 and index 2 need to be reverse complemented because they are read after the "turn". As another example, in a read order of "index 1, index 2, read segment 1, turn, read segment 2", none of the index sequences need to be reverse complemented.

[0244] In some embodiments, the method may include an operation 541 of generating sequencing results in a predetermined data format using the one or more index sequences based on the determination of the one or more index sequences. For example, the index sequences may be determined based on: 1) determining what the reference index sequence is according to the retrieved generated index sequence as a reference; and also based on 2) whether the reference index sequence needs to be reverse complemented, which depends on whether the index sequence is before the "turn" in the read order. The determined index sequences may then be used to extract read segments of DNA fragments and save them in a single data file or multiple data files. The data files may be in a predetermined format.

[0245] In some embodiments, operation 541 may include separating multiple sets of sequencing reads based on the one or more index sequences and the determination thereof to generate sequencing results in a predetermined data format in multiple files. In some embodiments, operation 541 may include demultiplexing the set of sequencing reads and saving them separately into different data files.

[0246] In some embodiments, the predetermined data format includes the FastQ format. In some embodiments, the predetermined data format includes a text-based format. In some embodiments, the predetermined data format includes the ASCII format. In some embodiments, the predetermined data format includes an 8-bit encoding format.

[0247] The accurate determination of index sequences and whether some or all of the index sequences need to be reverse complemented is crucial for the error-free classification, separation, and storage of sequencing data from different samples. Any errors at this stage can propagate to downstream analyses and make them problematic and unreliable. When a large number of samples' index sequences need to be reverse complemented, manual processing is time-consuming and highly error-prone. For example, for 100 samples and a pair of index sequences of 12 bases each per sample, the user may need to manually determine the reverse complement sequences of 2400 bases and input them into the correct positions. Currently, this manual determination and input of index sequences remains a major pain point for industry users and is also one of the main reasons for re-running sequencing runs. The techniques disclosed herein eliminate all manual determination and input of index sequences and are capable of performing the determination and input much faster and with higher accuracy than existing methods.

[0248] In some embodiments, after an entity of a particular sequencing run is completed, operation 541 can be performed on the data of the sequencing run.

[0249] In some embodiments, after a portion of a particular sequencing run is completed, operation 541 can be performed on the data of the sequencing run. Specifically, only the completed portion can correspond to at least some or all of the cycles of one or more index sequences.

[0250] In some embodiments, operation 541 can include evaluating one or more statistical parameters of the sequencing results while generating multiple sets of sequencing reads, where the multiple sets of sequencing reads include reads of fragments of the DNA sequence (i.e., read 1) and / or reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2).

[0251] In some embodiments, operation 541 can include evaluating one or more statistical parameters of the sequencing results while the operation 511 of generating multiple sets of sequencing reads is still being performed. Specifically, at least a portion of the operation 511 of generating multiple sets of sequencing reads has been completed. The completed portion of operation 511 can correspond to some or all of the index sequences in the multiple sets of sequencing reads.

[0252] In some embodiments, operation 541 can include evaluating one or more statistical parameters of the sequencing results while the operation of acquiring a flow cell image is still being performed. Specifically, at least a portion of the operation of acquiring a flow cell image has been completed. The completed portion of the operation can correspond to some or all of the index sequences in the multiple sets of sequencing reads. In other words, the flow cell images corresponding to the sequencing cycles for some or all of the index sequences have been completed, while the sequencing cycles corresponding to sequencing DNA fragments have not been completed.

[0253] In some embodiments, operation 541 may include evaluating one or more statistical parameters of the sequencing results after capturing flow cell images corresponding to some or all of the index sequences in a sequencing cycle, after a preliminary analysis of these flow cell images, and / or after sequencing data corresponding to these flow cell images has been generated.

[0254] In some embodiments, evaluating one or more statistical parameters of the sequencing results includes: calculating at least one value of each statistical parameter based on one or more retrieved index sequences used as a reference. The one or more retrieved index sequences are assumed to be reference index sequences, and they can be in their forward or forward-complementary direction with minimal errors (if any). In some embodiments, evaluating one or more statistical parameters of the sequencing results includes: comparing at least one value of each statistical parameter with a pre-determined threshold. In some embodiments, the statistical parameters of the sequencing results may include, but are not limited to: a mismatch parameter and / or an unassigned parameter.

[0255] The mismatch parameter may indicate the percentage or number of bases that do not match the retrieved reference index sequence out of the total number of bases. For example, a mismatch rate of 10% indicates that on average 1 out of every 10 bases does not match in all index sequences. The unassigned parameter may indicate how many index sequences do not match any stored reference index sequence at all. For example, an unassigned rate of 1% may be caused by reasonable sequencing errors. However, an unassigned rate of 15% in all index sequences may indicate that some of the index sequences were incorrectly entered or were incorrectly reverse-complemented.

[0256] In response to this evaluation, if one or more statistical parameters do not meet the pre-determined threshold, there may be a problem, such as an indexing error or a library problem, which may render the remaining sequence runs invalid. Therefore, the incorrect sequence run can be stopped before completion, and the instrument can be made available for other runs. Thus, method 501 may include an operation of stopping the generation of sequencing reads. Specifically, method 501 may include stopping the generation of reads of a fragment of a DNA sequence (i.e., read 1), reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2), or both. Method 501 may include an operation of terminating the capture of flow cell images in any subsequent cycles that have not been completed. Specifically, method 501 may include terminating the capture of flow cell images corresponding to reads of a fragment of a DNA sequence (i.e., read 1), reads of the complementary strand of the fragment of the DNA sequence (i.e., read 2), or both.

[0257] As an example, a sequencing run has 120 cycles, and the first 20 cycles contain all cycles corresponding to index 1. Cycles 21 to 30 correspond to index 2 in each sequencing read. After the flow cell images in the first 20 or 30 cycles of the sequencing run have been captured, these flow cell images in the 20 or 30 cycles can be processed, and sequencing data corresponding to index 1 (alone or in combination with index 2) can be generated, while still capturing the flow cell images in the subsequent 90 cycles corresponding to DNA fragments (e.g., read 1, alone or in combination with read 2). Sequencing statistics can be calculated as disclosed herein. After determining that the statistical parameters do not meet a predetermined threshold, the user can decide to prematurely terminate the sequencing run at the 50th sequencing cycle to address library preparation issues or index sequence conflicts. After the issues are resolved, the sequence run can be executed again starting from the first sequencing cycle. In this example, time, computational resources, and COGS used for running the subsequent 70 sequencing cycles and processing the problematic sequencing data are saved.

[0258] In response to this assessment, if one or more statistical parameters meet a predetermined threshold, method 501 can include continuing to generate sequencing reads, e.g., read 1 and / or read 2, while concurrently sorting and separating index sequences in parallel to save total sequence data analysis time. In some embodiments, method 501 herein includes performing demultiplexing of the completed portion of the sequencing run while concurrently performing the remaining sequencing run.

[0259] In some embodiments, method 501 can include an operation of generating library molecules 700 using one or more index sequences. Each library molecule 700 can include an insert fragment 710, i.e., the sequence of interest from the sample, e.g., R1 or R2. Figures 6 - 7 An exemplary linear library molecule 700 disclosed herein is shown. The linear library molecule can hybridize with a splint molecule. The splint molecule herein can be single-stranded or double-stranded. The splint molecule can be used as an amplification primer for a rolling circle amplification reaction. In some embodiments, the linear library module herein can perform a ligation reaction to close a single nick, thereby forming a covalently closed circular library molecule, which can hybridize with a splint molecule (e.g., a single-stranded splint molecule (not shown)), where the splint molecule is used as an amplification primer to perform a rolling circle amplification reaction. In some embodiments, a linear single-stranded library molecule (700) can hybridize with a double-stranded splint molecule, thereby circularizing the library molecule to form a library-splint complex (800) with two nicks. Figures 8A - 8B An exemplary library-splint complex 800 is shown, which performs a ligation reaction to close the nicks, thereby forming a covalently closed circular library molecule (900), which can hybridize with a double-stranded splint molecule. The splint molecule can be used as an amplification primer for a rolling circle amplification reaction. The dashed lines represent nascent extension products.

[0260] In some embodiments, methods 500 and 501 herein may include an operation of providing a sample having a plurality of tandem molecules immobilized on a support prior to operation 510 or 511, wherein each tandem molecule corresponds to a target RNA of a cell sample. In some embodiments, methods 500 and 501 herein may include obtaining a first plurality of flow cell images of a sample immobilized on a support prior to operation 510 or 511.

[0261] In some embodiments, the operation of obtaining a first plurality of flow cell images of a sample includes: generating, by a sequencing system, the first plurality of flow cell images by performing one or more sequencing reaction cycles on a sample immobilized on a support, wherein the plurality of flow cell images are generated along an axial axis from two or more color channels at two or more different z levels. In some embodiments, the operation of obtaining a first plurality of flow cell images of a sample includes: generating, by a sequencing system, a plurality of flow cell images by performing one or more sequencing reaction cycles on a plurality of tandem molecules of a sample immobilized on a support.

[0262] A sample herein may contain a polymerase community or cluster immobilized thereon. The polymerase community or cluster may correspond to a plurality of nucleotide template molecules or tandem molecules. In some embodiments, method 500 or 501 may include an operation of generating, by a sequencing system, a first plurality of flow cell images by performing one or more sequencing reaction cycles on a plurality of tandem molecules immobilized on a support prior to operation 510 or 511. In some embodiments, performing one or more sequencing reaction cycles includes contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U. In some embodiments, performing one or more sequencing reaction cycles includes contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity agents. A single affinity agent in the mixture may include a nucleus attached with a plurality of nucleotide arms, and each arm of the single affinity agent includes the same type of nucleobase. Performing one or more sequencing reaction cycles may include imaging, in each of one or more cycles, an optical color signal emitted by a nucleotide reagent bound to the plurality of tandem molecules by an optical system. Performing one or more sequencing reaction cycles may include obtaining, in each of one or more cycles, a first plurality of flow cell images containing an optical color signal emitted by a nucleotide reagent bound to the plurality of tandem molecules by an optical system.

[0263] In some embodiments, the first plurality of flow cell images include optical signals emitted from nucleotide reagents that incorporate into nucleobases A, G, C, and T / U of a plurality of templates or concatemer molecules immobilized on a support in one or more cycles with an unbalanced diversity. The plurality of polymerase colonies or clusters may include an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity includes the percentage of (1) the number of one or more types of nucleobases in a zone of the first plurality of flow cell images to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the zone.

[0264] In some embodiments, methods 500 and 501 may further include, prior to operation 510 or 511, providing a sample containing a plurality of RNAs that includes at least a first target RNA molecule and a second target RNA molecule. In some embodiments, methods 500 and 501 may further include generating a plurality of cDNA molecules within the sample prior to operation 510 or 511, the plurality of cDNA molecules including at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules including a second target cDNA molecule corresponding to the second target RNA molecule. In some embodiments, methods 500 and 501 may further include contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes prior to operation 510 or 511, the plurality of padlock probes including at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, methods 500 and 501 may further include, prior to operation 510 or 511, closing the nicks or gaps in at least the first and second circularized target-specific padlock probes by performing an enzymatic reaction to generate at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample. In some embodiments, methods 500 and 501 may further include, prior to operation 510 or 511, performing a rolling circle amplification reaction within the sample using the first and second covalently closed circular padlock probes as template molecules to generate a plurality of concatemer molecules, the plurality of concatemer molecules including at least a first concatemer molecule corresponding to the first target RNA molecule, and the plurality of concatemer molecules including at least a second concatemer molecule corresponding to the second target RNA molecule. In some embodiments, methods 500 and 501 may further include sequencing the plurality of concatemer molecules within the sample prior to operation 510 or 511, the sequencing including: sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products; and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.

[0265] In some embodiments, operations 510 and 511 may further include an operation of sequencing a plurality of concatemer molecules within a sample. The operation of sequencing a plurality of concatemer molecules within a sample may include: contacting the plurality of concatemer molecules within the sample with (i) a plurality of universal sequencing primers; (ii) a plurality of sequencing polymerases; and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers. In some embodiments, the nucleotide reagents include one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs. In some embodiments, operations 510 and 511 may further include an operation of removing a plurality of first sequencing read products from a first concatemer molecule and retaining the first concatemer molecule in the sample, and an operation of removing a plurality of second sequencing read products from a second concatemer molecule and retaining the second concatemer molecule in the sample.

[0266] Computer System

[0267] Various embodiments of the method may be implemented, for example, using one or more computer systems (such as the computer system 400 shown in Figure 4 ). For example, one or more computer systems 400 may be used to implement any of the various embodiments described herein, as well as combinations and sub - combinations thereof.

[0268] The computer system 400 may include one or more hardware processors 404. The hardware processor 404 may be a central processing unit (CPU), a graphics processing unit (GPU), or a combination thereof. The processor 404 may be connected to a bus or communication infrastructure 406.

[0269] The computer system 400 may further include a user input / output device 403, such as a display, a keyboard, a pointing device, etc., and the computer system may communicate with the communication infrastructure 406 through a user input / output interface 402. The user input / output device 403 may be coupled to the Figure 1 user interface 124 in

[0270] One or more of the processors 404 can be a graphics processing unit (GPU). In one embodiment, the GPU can be a processor that is a specialized electronic circuit designed to process math-intensive applications. The GPU can have a parallel architecture that is effective for parallel processing of large blocks of data, such as the math-intensive data common in computer graphics applications, images, video, vector processing, array processing, etc., as well as cryptography (including brute-force cracking), generating cryptographic hashes or hash sequences, solving partial hash inversion problems, and / or producing results of other proof-of-work calculations for some blockchain-based applications, for example. By virtue of the capabilities of general-purpose computing on graphics processing units (GPUs), GPUs can be particularly useful at least in the image recognition and machine learning described herein.

[0271] Additionally, one or more of the processors 404 can include a coprocessor or other logical implementations for accelerating cryptographic computations or other specialized math functions, including a hardware-accelerated cryptographic coprocessor. Such an acceleration processor can further include an instruction set for using the coprocessor and / or other logic to facilitate such acceleration.

[0272] The computer system 400 can also include a data storage device, such as a main memory 408, e.g., random access memory (RAM). The main memory 408 can include one or multiple levels of cache. The main memory 408 can store control logic (i.e., computer software) and / or data therein.

[0273] The computer system 400 can also include one or more secondary data storage devices or secondary memories 410. The secondary memory 410 can include, for example, a main storage drive 412 and / or a removable storage device or drive 414. The main storage drive 412 can be, for example, a hard disk drive or a solid-state drive. The removable storage drive 414 can be a floppy disk drive, a tape drive, an optical disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.

[0274] The removable storage drive 414 can interact with a removable storage unit 418.

[0275] The removable storage unit 418 may include a computer-usable or readable storage device on which computer software and / or data is stored. The software may include control logic. The software may include instructions executable by the hardware processor 404. The removable storage unit 418 may be a floppy disk, magnetic tape, optical disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 414 may read from and / or write to the removable storage unit 418.

[0276] The secondary storage 410 may include other components, devices, assemblies, tools, or other means for allowing the computer system 400 to access computer programs and / or other instructions and / or data. Such components, devices, assemblies, tools, or other means may include, for example, a removable storage unit 422 and an interface 420. Examples of the removable storage unit 422 and the interface 420 may include a program cartridge and a cartridge interface (such as the interface found in a video game device), a removable memory chip (such as an EPROM or PROM) and an associated socket interface, a memory stick and a USB port, a memory card and an associated memory card slot, and / or any other removable storage unit and an associated interface.

[0277] The computer system 400 may also include a communication or network interface 424. The communication interface 424 may enable the computer system 400 to communicate and interact with any combination of external devices, external networks, external entities, etc. (collectively and individually referred to by the reference numeral 428). For example, the communication interface 424 may allow the computer system 400 to communicate with an external or remote device 428 via a communication path 426, which may be wired and / or wireless (or a combination thereof), and may include any combination of a LAN, a WAN, a network, etc. Control logic and / or data may be transmitted to and from the computer system 400 via the communication path 426. In some embodiments, the communication path 426 is a connection to the cloud 130, as Figure 1 depicted. The external devices, etc. referred to by the reference numeral 428 may be devices, networks, entities, etc. in the cloud 130.

[0278] The computer system 400 may also be any one or any combination of the following: a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet computer, a smartphone, a smartwatch or other wearable device, an appliance, a part of the Internet of Things (IoT), and / or an embedded system, to name just a few non-limiting examples.

[0279] It should be understood that the frameworks described herein can be implemented as a method, process, apparatus, system, or article of manufacture, such as a non-transitory computer-readable medium or device. For illustrative purposes, the frameworks can be described in the context of a distributed ledger that is publicly available or at least accessible to untrusted third parties. An example of a modern use case is a blockchain-based system. However, it should be understood that the frameworks can also be applied to other settings where sensitive or confidential information may need to pass through the hands of untrusted third parties, and this technology is in no way limited to distributed ledger or blockchain use.

[0280] Computer system 400 can be a client or a server that accesses or hosts any application and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions; local or on-premises software (e.g., "on-premises" cloud-based solutions); "as-a-service" models (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS), Database as a Service (DBaaS), etc.); and / or a hybrid model that includes any combination of the foregoing examples or other services or delivery paradigms.

[0281] Any applicable data structures, file formats, and schemas can be derived from standards including but not limited to: JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible HyperText Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representation, either alone or in combination. Alternatively, proprietary data structures, formats, or schemas can be used alone or in combination with known or open standards.

[0282] Any relevant data, files, and / or databases can be stored, retrieved, accessed, and / or transmitted in a human-readable format (such as numeric, text, graphic, or multimedia formats, further including various types of markup languages and other possible formats). Alternatively, or in combination with the above formats, data, files, and / or databases can be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted formats or any other machine-readable format.

[0283] The interface connection or interconnection between various systems and layers can adopt any number of mechanisms, such as any number of protocols, programming frameworks, layout plans, or application programming interfaces (APIs), including but not limited to the Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Pattern (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG) HTML5 Web Messaging, Representational State Transfer (REST or RESTful web services), eXtensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other open or proprietary mechanism that can achieve similar functions and results.

[0284] Such interface connection or interconnection can also utilize Uniform Resource Identifiers (URIs), which can further include Uniform Resource Locators (URLs) or Uniform Resource Names (URNs). Other forms of uniform and / or unique identifiers, locators, or names can be used, either alone or in combination with those forms such as the above.

[0285] Any one of the above protocols or APIs can interface with or be implemented in any programming language (procedural, functional, or object-oriented) and can be compiled or interpreted. Non-limiting examples include C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or almost any other language, as well as any other library or pattern in any type of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, including but not limited to Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, etc., and many other non-limiting examples.

[0286] In some embodiments, a tangible non-transitory apparatus or article of manufacture that includes a tangible non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or a program storage device. This includes, but is not limited to, computer system 400, main memory 408, auxiliary memory 410, and removable storage units 418 and 422, as well as tangible articles embodying any combination of the foregoing. When executed by one or more data processing devices, such as computer system 400, such control logic may cause such data processing devices to operate as described herein.

[0287] Based on the teachings contained in this disclosure, it will be apparent to those skilled in the relevant art how to make and use embodiments of this disclosure using data processing devices, computer systems, and / or computer architectures different from those Figure 4 shown. Specifically, the embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.

[0288] Optical System

[0289] Figure 1 The imager 116 in may include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system design provides a larger field of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio, and image quality, a higher spatial sampling frequency, faster switching between image captures when repositioning the sample plane to capture a series of images (e.g., images of different fields of view), and an improved imaging system duty cycle, and thus enables higher throughput image acquisition and analysis.

[0290] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, may be achieved by using an electro-optic phase plate in combination with an objective lens to compensate for optical aberrations caused by fluid layers separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach may also compensate for vibrations introduced by, for example, a motion-actuated compensator that moves into or out of the optical path depending on which surface of the flow cell is being imaged.

[0291] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, involve using a thick flow cell wall (e.g., wall (or coverslip) thickness > 700 μm) and a fluid channel (e.g., fluid channel height or thickness of 50 to 200 μm), and can be achieved by using a tube lens design that corrects for optical aberrations caused by the thick flow cell wall and / or intervening fluid layer in combination with the objective lens, even when using an off-the-shelf commercially available objective lens.

[0292] In some cases, improvements in imaging performance, such as for multi-channel (e.g., two-color or four-color) imaging applications, can be achieved by using multiple tube lenses (one tube lens per imaging channel), where each tube lens design has been optimized for the specific wavelength range used in this imaging channel.

[0293] Exemplary embodiments disclosed herein may include a fluorescence imaging system, the system comprising: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; b) an objective lens configured to collect fluorescence generated from within a specified field of view of a sample plane when the sample plane is exposed to the excitation light, wherein the numerical aperture of the objective lens is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9, or the numerical aperture value falls within the range defined by any two of the foregoing; wherein the working distance of the objective lens is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least 1000 μm, or a working distance that falls within the range defined by any two of the foregoing; and wherein the field of view has an area of at least 0.1 mm 2 、at least 0.2 mm 2 、at least 0.5 mm 2 、at least 0.7 mm 2 、at least 1 mm 2 、at least 2 mm 2 、at least 3 mm 2 、at least 5 mm 2 or at least 10 mm 2 or a field of view that falls within the range defined by any two of the foregoing; and c) at least one image sensor, wherein the fluorescence collected by the objective lens is imaged onto the image sensor, and wherein the pixel size of the image sensor is selected such that the spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.

[0294] In some embodiments, the numerical aperture can be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the area of the field of view can be at least 2.5 mm2. In some embodiments, the area of the field of view can be at least 3 mm2. In some embodiments, the spatial sampling frequency can be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency can be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system can further include an X-Y-Z translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, where each image in the series is or can be acquired for a different field of view. In some embodiments, the positioning of the sample plane can be adjusted simultaneously in the X direction, Y direction, and Z direction to match the positioning of the objective focal plane between acquiring images of different fields of view. In some embodiments, the time required for simultaneous adjustment in the X direction, Y direction, and Z direction can be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second or a time falling within the range defined by any two of the foregoing. In some embodiments, the system further includes an autofocus mechanism configured to adjust the focal plane positioning before acquiring images of different fields of view if an error signal indicates that the positioning difference between the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or greater. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system includes three or more image sensors, and wherein the system is configured to image fluorescence in each of three or more wavelength ranges onto different image sensors. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 50 nm. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.4 seconds per field of view. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.3 seconds per field of view.

[0295] The present disclosure also discloses a fluorescence imaging system for dual-sided imaging of a flow cell, the fluorescence imaging system comprising: a) an objective lens configured to collect fluorescence generated within a specified field of view of a sample plane within the flow cell; b) at least one tube lens positioned between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to correct an imaging performance metric of a combination of the objective lens, the at least two tube lenses, and the at least one image sensor when imaging an inner surface of the flow cell, and wherein a wall thickness of the flow cell is at least 700 μm and a gap between an upper inner surface and a lower inner surface is at least 50 μm; wherein an imaging performance metric is substantially the same for imaging an upper inner surface or a lower inner surface of the flow cell without moving an optical compensator into or out of an optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of the tube lens along the optical path, and without moving one or more optical elements of the tube lens into or out of the optical path.

[0296] In some embodiments, the objective lens can be a commercially available microscope objective lens. In some embodiments, the numerical aperture of the commercially available microscope objective lens can be at least 0.3. In some embodiments, the working distance of the objective lens can be at least 700 μm. In some embodiments, the objective lens can be corrected to compensate for a cover glass thickness (or flow cell wall thickness) of 0.17 mm or greater than or less than 0.17 mm. In some embodiments, the optical system can be corrected to compensate for the distance between the cover glass thickness, the flow cell thickness, or the desired focal plane. In some embodiments, the correction can be performed by inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the correction can be performed without inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the fluorescence imaging system can further include an electro-optic phase plate, which is positioned adjacent to the objective lens and between the objective lens and the tube lens, wherein the electro-optic phase plate can provide correction for optical aberrations caused by a fluid filling the gap between the upper inner surface and the lower inner surface of the flow cell. In some embodiments, at least one tube lens can be a compound lens including three or more optical components. In some embodiments, at least one tube lens is a compound lens including four optical components, and the four optical components can include one or more of the following: a first asymmetric convex-convex lens, a second convex-plano lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, which can be present in the order listed above or in any alternative order. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a wall thickness of at least 1 mm. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 100 μm. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, and each of the two tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, and each of the three tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, and each of the four tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths.In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the medium to high spatial frequency range. In some embodiments, the imaging performance metric includes measurements of the modulation transfer function (MTF), defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof at one or more specified spatial frequencies. In some embodiments, the difference in the imaging performance metric for imaging the upper inner surface and the lower inner surface of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metric for imaging the upper inner surface and the lower inner surface of the flow cell is less than 5%. In some embodiments, compared with a conventional system including an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least equivalent or better improvement in the imaging performance metric for bilateral imaging. In some embodiments, compared with a conventional system including an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least a 10% improvement in the imaging performance metric for bilateral imaging.

[0297] Disclosed herein is an illumination system for imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and transmit it to a specified illumination field on a support surface including immobilized biological macromolecules.

[0298] In some embodiments, the illumination system further includes a condenser lens. In some embodiments, the area of the specified illumination field is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has uniform intensity across a specified field of view of an imaging system for acquiring an image of the support surface. In some embodiments, the area of the specified field of view is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the speckle contrast value of the light transmitted to the specified illumination field is less than 0.1. In some embodiments, the speckle contrast value of the light transmitted to the specified illumination field is less than 0.05.

[0299] Imaging Modules and Systems:

[0300] Those skilled in the art will understand that, in some cases, the disclosed optical systems, imaging systems or modules can be stand-alone optical systems designed to image a sample or the surface of a substrate. In some cases, they can include one or more processors or computers. In some cases, they can include one or more software packages that provide instrument control functions and / or image processing functions. In some cases, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical band-pass filters, apertures and image sensors (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they can also include mechanical and / or optomechanical components such as X-Y translation stages, X-Y-Z translation stages, piezoelectric focusing mechanisms, etc. In some cases, they can serve as modules, components, sub-assemblies or subsystems of a larger system designed for genomics applications (e.g., gene testing and / or nucleic acid sequencing applications). For example, in some cases, they can serve as modules, components, sub-assemblies or subsystems of a larger system that further includes a light-tight and / or other environmental control housing, a temperature control module, a fluid control module, a fluid dispensing robot, a pick-and-place robot, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software package, image processing software package, data analysis software package), a data storage module, a data communication module (e.g., Bluetooth, WiFi, intranet or Internet communication hardware and related software), a display module or any combination thereof.

[0301] Figure 1 The imager 116 therein can include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for genomics applications based on fluorescence imaging. The disclosed optical imaging system design provides a larger field of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio and image quality, higher spatial sampling frequency, faster switching between image captures when repositioning the sample plane to capture a series of images (e.g., images of different fields of view), and improved imaging system duty cycle, and thus enables higher throughput image acquisition and analysis.

[0302] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, can be achieved by using an electro-optic phase plate in combination with an objective lens to compensate for optical aberrations caused by fluid layers separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach can also compensate for vibrations introduced by, for example, a motion actuator compensator that moves into or out of the optical path depending on which surface of the flow cell is being imaged.

[0303] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, involve using a thick flow cell wall (e.g., wall (or cover glass) thickness > 700 μm) and a fluid channel (e.g., fluid channel height or thickness of 50 to 200 μm), and can be achieved by using a tube lens design that corrects for optical aberrations caused by the thick flow cell wall and / or intervening fluid layer in combination with the objective lens, even when using an off-the-shelf commercially available objective lens.

[0304] In some cases, improvements in imaging performance, such as for multi-channel (e.g., two-color or four-color) imaging applications, can be achieved by using multiple tube lenses (one tube lens per imaging channel), where each tube lens design has been optimized for a specific wavelength range used in this imaging channel.

[0305] Exemplary embodiments disclosed herein may include a fluorescence imaging system that includes: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; b) an objective lens configured to collect fluorescence generated from within a specified field of view of a sample plane when the sample plane is exposed to the excitation light, where the numerical aperture of the objective lens is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9, or the numerical aperture value falls within a range defined by any two of the foregoing; where the working distance of the objective lens is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least 1000 μm, or a working distance that falls within a range defined by any two of the foregoing; and where the field of view has at least 0.1 mm 2 、at least 0.2 mm 2 、at least 0.5 mm 2 、at least 0.7 mm 2 、at least 1 mm 2 、at least 2 mm 2 、at least 3 mm 2 、at least 5 mm 2 or at least 10 mm 2the area, or a field of view that falls within the range defined by any two of the foregoing; and c) at least one image sensor onto which the fluorescence collected by the objective lens is imaged, and wherein the pixel size of the image sensor is selected such that the spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.

[0306] In some embodiments, the numerical aperture can be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the area of the field of view can be at least 2.5 mm2. In some embodiments, the area of the field of view can be at least 3 mm2. In some embodiments, the spatial sampling frequency can be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency can be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system can further include an X-Y-Z translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, wherein each image in the series is or can be acquired for a different field of view. In some embodiments, the positioning of the sample plane can be adjusted simultaneously in the X direction, Y direction, and Z direction to match the positioning of the objective lens focal plane between acquiring images of different fields of view. In some embodiments, the time required for simultaneous adjustment in the X direction, Y direction, and Z direction can be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second, or a time that falls within the range defined by any two of the foregoing. In some embodiments, the system further includes an autofocus mechanism configured to adjust the focal plane positioning before acquiring images of different fields of view if an error signal indicates that the positioning difference between the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or greater. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system includes three or more image sensors, and wherein the system is configured to image the fluorescence in each of three or more wavelength ranges onto different image sensors. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 50 nm. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.4 seconds per field of view. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.3 seconds per field of view.

[0307] The present disclosure also discloses a fluorescence imaging system for dual-sided imaging of a flow cell, the fluorescence imaging system comprising: a) an objective lens configured to collect fluorescence generated within a specified field of view of a sample plane within the flow cell; b) at least one tube lens positioned between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to correct an imaging performance metric of a combination of the objective lens, the at least two tube lenses, and the at least one image sensor when imaging an inner surface of the flow cell, and wherein a wall thickness of the flow cell is at least 700 μm and a gap between an upper inner surface and a lower inner surface is at least 50 μm; wherein an imaging performance index is substantially the same for imaging an upper inner surface or a lower inner surface of the flow cell without moving an optical compensator into or out of an optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of the tube lens along the optical path, and without moving one or more optical elements of the tube lens into or out of the optical path.

[0308] In some embodiments, the objective lens can be a commercially available microscope objective lens. In some embodiments, the numerical aperture of the commercially available microscope objective lens can be at least 0.3. In some embodiments, the working distance of the objective lens can be at least 700 μm. In some embodiments, the objective lens can be corrected to compensate for a coverslip thickness (or flow cell wall thickness) of 0.17 mm or greater than or less than 0.17 mm. In some embodiments, the optical system can be corrected to compensate for the distance between the coverslip thickness, the flow cell thickness, or the desired focal plane. In some embodiments, the correction can be performed by inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the correction can be performed without inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the fluorescence imaging system can further include an electro-optic phase plate that is positioned adjacent to the objective lens and between the objective lens and the tube lens, wherein the electro-optic phase plate can provide correction for optical aberrations caused by the fluid filling the gap between the upper and lower inner surfaces of the flow cell. In some embodiments, at least one tube lens can be a compound lens that includes three or more optical components. In some embodiments, at least one tube lens is a compound lens that includes four optical components, and the four optical components can include one or more of the following: a first asymmetric convex-convex lens, a second convex-plano lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, which can be present in the order listed above or in any alternative order. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a wall thickness of at least 1 mm. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 100 μm. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, and each of the two tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, and each of the three tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, and each of the four tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths.In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the medium to high spatial frequency range. In some embodiments, the imaging performance metrics include measurements of the modulation transfer function (MTF), defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof, at one or more specified spatial frequencies. In some embodiments, the difference in the imaging performance metrics for imaging the upper and lower inner surfaces of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metrics for imaging the upper and lower inner surfaces of the flow cell is less than 5%. In some embodiments, compared to a conventional system comprising an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least equivalent or better improvement in the imaging performance metrics for bilateral imaging. In some embodiments, compared to a conventional system comprising an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least a 10% improvement in the imaging performance metrics for bilateral imaging.

[0309] Disclosed herein is an illumination system for imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and transmit it to a specified illumination field on a support surface comprising immobilized biological macromolecules.

[0310] In some embodiments, the illumination system further comprises a condenser lens. In some embodiments, the area of the specified illumination field is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has a uniform intensity across a specified field of view of the imaging system used to acquire an image of the support surface. In some embodiments, the area of the specified field of view is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has a uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the specified illumination field has a uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the speckle contrast value of the light delivered to the specified illumination field is less than 0.1. In some embodiments, the speckle contrast value of the light delivered to the specified illumination field is less than 0.05.

[0311] Method for Sequencing

[0312] The present disclosure provides methods for sequencing immobilized or non-immobilized template molecules. The methods can be operated in system 100, e.g., in sequencer 114. In some embodiments, the immobilized template molecules include multiple nucleic acid template molecules having one copy of a target sequence of interest. In some embodiments, nucleic acid template molecules having one copy of a target sequence of interest can be generated by bridge amplification using linear library molecules. In some embodiments, the immobilized template molecules include multiple nucleic acid template molecules, each nucleic acid template molecule having two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, nucleic acid template molecules containing concatemer molecules can be generated by performing rolling circle amplification of circularized linear library molecules. In some embodiments, the non-immobilized template molecules include circular molecules. In some embodiments, the methods for sequencing employ soluble (e.g., non-immobilized) sequencing polymerases or sequencing polymerases immobilized to a support.

[0313] In some embodiments, the sequencing reaction employs detectably labeled nucleotide analogs. In some embodiments, the sequencing reaction employs a two-stage sequencing reaction, including binding a multivalent molecule labeled with a detectable label and incorporating nucleotide analogs. In some embodiments, the sequencing reaction employs unlabeled nucleotide analogs. In some embodiments, the sequencing reaction employs phosphate-linked nucleotides.

[0314] In some embodiments, each immobilized concatemer includes a tandem repeat unit (e.g., an insert region) of a sequence of interest and any linker sequences. For example, the tandem repeat unit includes: (i) a left universal linker sequence having a binding sequence for a first surface primer (720) (e.g., a surface pinned primer), (ii) a left universal linker sequence having a binding sequence for a first sequencing primer (740) (e.g., a forward sequencing primer), (iii) the sequence of interest (710), (iv) a right universal linker sequence having a binding sequence for a second sequencing primer (750) (e.g., a reverse sequencing primer), (v) a right universal linker sequence having a binding sequence for a second surface primer (730) (e.g., a surface capture primer), and (vii) a left sample index sequence (760) and / or a right sample index sequence (770). In some embodiments, the tandem repeat unit further includes a left unique identifier sequence (780) and / or a right unique identifier sequence (790). In some embodiments, the tandem repeat unit further includes a binding sequence for at least one compaction oligonucleotide. In some embodiments, Figure 6 and 7 illustrates a unit of a linear library molecule or a concatemer molecule.

[0315] The immobilized concatemer can self-collapse into a compact nucleic acid nanosphere. Inclusion of one or more compaction oligonucleotides during the RCA reaction can further compact the size and / or shape of the nanosphere. An increase in the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer for hybridization with a variety of sequencing primers (e.g., sequencing primers having a universal sequence), which serve as multiple initiation sites for polymerase-catalyzed sequencing reactions. When the sequencing reaction employs detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., having nucleotide units), the signals emitted by the nucleotides or nucleotide units participating in parallel sequencing reactions along the concatemer produce an increased signal intensity for each concatemer. Multiple portions of a given concatemer can be sequenced simultaneously. In addition, multiple binding complexes can form along a particular concatemer molecule, each binding complex comprising a sequencing polymerase that binds to a template / primer duplex and to a multivalent molecule, wherein the multiple binding complexes remain stable and do not dissociate, resulting in an increased dwell time, which increases the signal intensity and reduces the imaging time.

[0316] Method for Sequencing Using Nucleotide Analogs

[0317] The present disclosure provides a method for sequencing any of the immobilized template molecules described herein, the method comprising step (a): contacting a sequencing polymerase with (i) a nucleic acid template molecule; and (ii) a nucleic acid sequencing primer, wherein the contacting is carried out under conditions suitable for binding the sequencing polymerase to the nucleic acid template molecule hybridized to the nucleic acid primer, wherein the nucleic acid template molecule hybridized to the nucleotide primer forms a nucleic acid duplex. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase that can bind and incorporate nucleotide analogs.

[0318] In some embodiments of the method for sequencing a template molecule, the sequencing primer comprises a 3'-extendable end or a 3'-non-extendable end. In some embodiments, the multiple nucleic acid template molecules include amplified template molecules (e.g., template molecules amplified in a clonal manner). In some embodiments, the multiple nucleic acid template molecules include one copy of a target sequence of interest. In some embodiments, the multiple nucleic acid molecules include two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, the multiple nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the multiple nucleic acid primers are in solution or immobilized to a support. In some embodiments, when the multiple nucleic acid template molecules and / or the multiple nucleic acid primers are immobilized to a support, binding to a first sequencing polymerase produces multiple immobilized first complex polymerases. In some embodiments, the multiple nucleic acid template molecules and / or nucleic acid primers are immobilized to the support at 10 2 -10 15different sites. In some embodiments, the binding of multiple template molecules and nucleic acid primers to multiple first sequencing polymerases generates multiple first complex polymerases fixed to a support at 10 2 –10 15 different sites on the support. In some embodiments, the multiple immobilized first complex polymerases on the support are fixed to predetermined or random sites on the support. In some embodiments, the multiple immobilized first complex polymerases are in fluid communication with each other to allow a reagent solution (e.g., an enzyme including a sequencing polymerase, a multivalent molecule, nucleotides, and / or divalent cations) to flow onto the support such that the multiple immobilized complex polymerases on the support react with the reagent solution in a massively parallel manner.

[0319] In some embodiments, the method for sequencing further includes step (b): contacting a sequencing polymerase with multiple nucleotides under conditions suitable for binding at least one nucleotide to the sequencing polymerase bound to a nucleic acid duplex and suitable for incorporating the polymerase-catalyzed nucleotide, wherein incorporating the polymerase-catalyzed nucleotide extends the sequencing primer by one nucleotide. In some embodiments, the sequencing polymerase is contacted with multiple nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the multiple nucleotides include at least one nucleotide analog having a chain-terminating moiety at the 2' or 3' position of the sugar. In some embodiments, the chain-terminating moiety can be removed from the 2' or 3' position of the sugar to convert the chain-terminating moiety into an OH or H group. In some embodiments, the multiple nucleotides include at least one nucleotide lacking a chain-terminating moiety. In some embodiments, at least one nucleotide is labeled with a detectable reporter moiety (e.g., a fluorophore) that emits a detectable signal. The detectable reporter moiety includes a fluorophore. In some embodiments, the fluorophore is linked to a nucleobase. In some embodiments, the fluorophore is linked to the nucleobase with a linker that can be cleaved / removed from the base. In some embodiments, at least one nucleotide among the multiple nucleotides is not labeled with a detectable reporter moiety. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) linked to a nucleotide can correspond to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase. When the incorporated chain-terminating nucleotide is detectably labeled, step (b) further comprises detecting the signal emitted from the incorporated chain-terminating nucleotide. In some embodiments, step (b) further comprises identifying the nucleobase of the incorporated chain-terminating nucleotide.

[0320] In some embodiments, the method for sequencing further comprises step (c): removing the chain-terminating moiety from the incorporated chain-terminating nucleotide to generate an extendable 3' OH group. In some embodiments, step (c) further comprises removing a detectable label from the incorporated chain-terminating nucleotide. In some embodiments, the sequencing polymerase remains bound to the template molecule, which hybridizes to a sequencing primer that extends one nucleobase.

[0321] In some embodiments, the method for sequencing further comprises step (d): repeating steps (b) and (c) at least once.

[0322] Two - Stage Method for Sequencing Nucleic Acids

[0323] The present disclosure provides a two-stage method for sequencing any immobilized template molecule described herein. In some embodiments, the first stage generally comprises binding a multivalent molecule to a complex polymerase to form a multivalent-complex polymerase, and detecting the multivalent-complex polymerase.

[0324] In some embodiments, the first stage comprises step (a): contacting a plurality of first sequencing polymerases with (i) a plurality of nucleic acid template molecules; and (ii) a plurality of nucleotide sequencing primers, wherein the contacting is carried out under conditions suitable for binding the plurality of first sequence polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers to form a plurality of first complex polymerases, each first complex polymerase comprising a first sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a nucleic acid template molecule hybridized to a nucleic acid primer. In some embodiments, the first polymerase comprises a recombinant mutant sequencing polymerase.

[0325] In some embodiments, in the method for sequencing a template molecule, the sequencing primer comprises an oligonucleotide having a 3' extendable terminus or a 3' non-extendable terminus. In some embodiments, the plurality of nucleic acid template molecules comprises amplified template molecules (e.g., template molecules amplified in a clonal manner). In some embodiments, the plurality of nucleic acid template molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, the nucleic acid template molecules in the plurality of nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are in solution or immobilized to a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized to a support, binding to the first sequencing polymerase produces a plurality of immobilized first complex polymerases. In some embodiments, the plurality of nucleic acid template molecules and / or nucleic acid primers are immobilized on the support at 10 2 –10 15different sites. In some embodiments, the binding of multiple template molecules and nucleic acid primers to multiple first sequencing polymerases generates multiple first complex polymerases immobilized on a support at 10 2 –10 15 different sites. In some embodiments, the multiple immobilized first complex polymerases on the support are immobilized at predetermined or random sites on the support. In some embodiments, the multiple immobilized first complex polymerases are in fluid communication with each other to allow a reagent solution (e.g., an enzyme including a sequencing polymerase, a multivalent molecule, a nucleotide, and / or a divalent cation) to flow onto the support, such that the multiple immobilized complex polymerases on the support react with the reagent solution in a massively parallel manner.

[0326] In some embodiments, the method for sequencing further includes step (b): contacting the multiple first complex polymerases with multiple multivalent molecules to form multiple multivalent-complex polymerases (e.g., binding complexes). In some embodiments, each multivalent molecule among the multiple multivalent molecules includes a core attached to multiple nucleotide arms, and each nucleotide arm is attached to a nucleotide (e.g., a nucleotide unit) (e.g., Figures 9 - 13 ). In some embodiments, the contacting of step (b) is carried out under conditions suitable for binding complementary nucleotide units of the multivalent molecules to at least two of the multiple first complex polymerases to form multiple multivalent-complex polymerases. In some embodiments, the conditions are suitable for inhibiting the polymerase-catalyzed incorporation of the complementary nucleotide units into the primers of the multiple multivalent-complex polymerases. In some embodiments, the multiple multivalent molecules include at least one multivalent molecule having multiple nucleotide arms (e.g., Figures 9 - 12 ), each nucleotide arm attached to a nucleotide analogue (e.g., a nucleotide analogue unit), wherein the nucleotide analogue includes a chain-terminating moiety at the 2' and / or 3' position of the sugar. In some embodiments, the multiple multivalent molecules include at least one multivalent molecule including multiple nucleotide arms, each nucleotide arm attached to a nucleotide unit lacking a chain-terminating moiety. In some embodiments, at least one of the multivalent molecules among the multiple multivalent molecules is labeled with a signal-emitting detectable reporter gene moiety. In some embodiments, the detectable reporter gene moiety includes a fluorophore. In some embodiments, the contacting of step (b) is carried out in the presence of at least one non-catalytic cation including strontium, barium, and / or calcium.

[0327] In some embodiments, the method for sequencing further comprises step (c): detecting a plurality of multivalent - complex polymerases. In some embodiments, the detection comprises detecting a signal emitted by a multivalent molecule that binds to the complex polymerase, wherein the complementary nucleotide units of the multivalent molecule bind to the primer but inhibit the incorporation of the complementary nucleotide units. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety to allow detection. In some embodiments, the labeled multivalent molecule comprises a fluorophore attached to the nucleus, linker, and / or nucleotide units of the multivalent molecule.

[0328] In some embodiments, the method for sequencing further comprises step (d): identifying the nucleobases of the complementary nucleotide units that bind to the plurality of first complex polymerases, thereby determining the sequence of the template molecule. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety that corresponds to a specific nucleotide unit linked to a nucleotide arm to allow identification of the complementary nucleotide units (e.g., nucleobases adenine, guanine, cytosine, thymine, or uracil) that bind to the plurality of first complex polymerases.

[0329] In some embodiments, the method for sequencing further comprises step (e): dissociating the plurality of multivalent - complex polymerases and removing the plurality of first sequencing polymerases and their bound multivalent molecules, and retaining the plurality of nucleic acid duplexes.

[0330] In some embodiments, the second stage of the two - stage sequencing method generally comprises nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (f): contacting the plurality of retained nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is carried out under conditions suitable for the plurality of second sequencing polymerases to bind to the plurality of retained nucleic acid duplexes, thereby forming a plurality of second complex polymerases, each second complex polymerase comprising a second sequencing polymerase that binds to the nucleic acid duplex. In some embodiments, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.

[0331] In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is 100% identical to the amino acid sequence of the plurality of second sequencing polymerases of step (f). In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is different from the amino acid sequence of the plurality of second sequencing polymerases of step (f).

[0332] In some embodiments, the method for sequencing further comprises step (g): contacting a plurality of second complex polymerases with a plurality of nucleotides, wherein the contacting is carried out under conditions suitable for binding at least two complementary nucleotides from the plurality of nucleotides to the complex polymerase, thereby forming a plurality of nucleotide complex polymerases. In some embodiments, the contacting of step (g) is carried out under conditions suitable for promoting the catalytic incorporation of the bound complementary nucleotide polymerase into the primer of the nucleotide complex polymerase, thereby extending the sequencing primer by one nucleobase. In some embodiments, the incorporation of the nucleotide into the 3'-end of the sequencing primer in step (g) comprises a primer extension reaction. In some embodiments, the contacting of step (g) is carried out in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises natural nucleotides (e.g., non-analogue nucleotides) or nucleotide analogues. In some embodiments, the plurality of nucleotides comprises removable or non-removable 2' and / or 3' chain terminating moieties. In some embodiments, at least one nucleotide among the nucleotides of the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the plurality of nucleotides is unlabeled. In some embodiments, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is linked to the nucleobase. In some embodiments, the fluorophore is attached to the nucleobase with a linker that can be cleaved / removed from the base or cannot be removed from the base. In some embodiments, a particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleobase (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to allow detection and identification of the nucleobase.

[0333] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further comprises step (h): detecting the complementary nucleotides incorporated into the primer of the nucleotide complex polymerase. In some embodiments, the plurality of nucleotides are labeled with a detectable reporter moiety to allow detection. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the detection of step (h) is omitted.

[0334] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further comprises step (i): identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide complex polymerase. In some embodiments, identifying the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the plurality of first complex polymerases in step (d). In some embodiments, the identification of step (i) can be used to determine the sequence of the nucleic acid template molecule. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the identification of step (i) is omitted.

[0335] In some embodiments, the method for sequencing further comprises step (j): removing the chain-terminating moiety from the incorporated nucleotide when step (g) is carried out by contacting a plurality of second complex polymerases with a plurality of nucleotides comprising at least one nucleotide having a 2' and / or 3' chain-terminating moiety.

[0336] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) to (j) at least once. In some embodiments, the sequence of the nucleic acid template molecule can be determined by detecting and identifying the multivalent molecule that binds to the sequencing polymerase but is not incorporated into the 3'-end of the primer at steps (c) and (d). In some embodiments, the sequence of the nucleic acid template molecule can be determined (or confirmed) by detecting and identifying the nucleotide incorporated into the 3'-end of the primer at steps (h) and (i).

[0337] In some embodiments, in any method for sequencing a nucleic acid molecule, the binding of a plurality of first complex polymerases to a plurality of multivalent molecules forms at least one affinity complex, and the method comprises the following steps: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemer template molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemer template molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, and the first binding complex and the second binding complex comprising the same multivalent molecule form an affinity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemer template molecule comprises a tandem repeat sequence of a target sequence and at least one universal sequencing primer binding site. The first nucleic acid primer and the second nucleic acid primer can bind to the sequencing primer binding site along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 9 - 12 in.

[0338] In some embodiments, in any method for sequencing nucleic acid molecules, wherein the method comprises binding a plurality of first complex polymerases to a plurality of multivalent molecules to form at least one affinity complex, the method comprises the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid concatemer molecule to form at least first and second complex polymerases on the same concatemeric template molecule; (b) contacting the plurality of multivalent molecules with at least the first and second complex polymerases on the same concatemeric template molecule under conditions suitable for binding a single multivalent molecule of the plurality of multivalent molecules to the first and second complex polymerases, wherein at least a first nucleotide unit of the single multivalent molecule binds to the first complex polymerase, the first complex polymerase comprising a first primer hybridized to a first portion of the concatemeric template molecule, thereby forming a first binding complex (e.g., a first ternary complex), and wherein at least a second nucleotide unit of the single multivalent molecule binds to the second complex polymerase, the second complex polymerase comprising a second primer hybridized to a second portion of the concatemeric template molecule, thereby forming a second binding complex (e.g., a second ternary complex), wherein the contacting is carried out under conditions suitable for inhibiting polymerase-catalyzed incorporation of the first and second nucleotide units bound in the first and second binding complexes, and wherein the first and second binding complexes bound to the same multivalent molecule form an affinity complex; and (c) detecting the first and second binding complexes on the same concatemeric template molecule; and (d) identifying the first nucleotide unit in the first binding complex to determine the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second binding complex to determine the sequence of the second portion of the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises any wild-type or mutant sequencing polymerase described herein. The concatemeric template molecule comprises a tandem repeat of a target sequence and at least one universal sequencing primer binding site. The plurality of nucleic acid primers can bind to the sequencing primer binding sites along the concatemeric template molecule. Exemplary multivalent molecules are shown in Figures 10 - 13 in.

[0339] Combined Sequencing

[0340] The present disclosure provides methods for sequencing any of the immobilized template molecules described herein, wherein the sequencing method includes a Sequencing by Binding (SBB) procedure using unlabeled chain-terminating nucleotides. In some embodiments, the Sequencing by Binding (SBB) method includes the steps of: (a) contacting the primed template nucleic acid with at least two separate mixtures sequentially under ternary complex stabilizing conditions, wherein each of the at least two separate mixtures contains a polymerase and nucleotides, whereby the sequential contact causes the primed template nucleic acid to contact homologs of nucleotides of the first, second, and third base types in the template under ternary complex stabilizing conditions; (b) examining the at least two separate mixtures to determine whether a ternary complex has formed; and (c) identifying the next correct nucleotide of the primed template nucleic acid molecule, wherein if a ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first, second, or third base type, and wherein based on the absence of a ternary complex in step (b), the next correct nucleotide is inferred to be a nucleotide homolog of the fourth base type; (d) after step (b), adding the next correct nucleotide to the primer of the primed template nucleic acid to produce an extended primer; and (e) repeating steps (a) through (d) at least once on the primed template nucleic acid containing the extended primer. Exemplary methods of sequencing while binding are described in U.S. Patent Nos. 10,246,744 and 10,731,141 (the contents of the two patents are hereby incorporated by reference in their entirety).

[0341] Method for Sequencing Using Phosphate - Linked Nucleotides

[0342] The present disclosure provides methods for sequencing using an immobilized sequencing polymerase that binds to an unfixed template molecule, wherein the sequencing reaction is performed with nucleotides labeled with a phosphodiester chain. In some embodiments, the sequencing method includes step (a): providing a support having a plurality of sequencing polymerases immobilized thereon. In some embodiments, the sequencing polymerase includes a processive DNA polymerase. In some embodiments, the sequencing polymerase comprises a wild-type or mutant DNA polymerase, including, for example, Phi29 DNA polymerase. In some embodiments, the support includes a plurality of separate compartments, and the sequencing polymerase is immobilized to the bottom of the compartments. In some embodiments, the separate compartments include a silica bottom through which light can penetrate. In some embodiments, the separate compartments include a silica bottom configured with a nanophotonic confinement structure including pores in a metal-coated membrane (e.g., an aluminum-coated membrane). In some embodiments, the small pore diameter of the pores in the metal coating is, for example, about 70 nm. In some embodiments, the height of the nanophotonic confinement structure is about 100 nm. In some embodiments, the nanophotonic confinement structure includes a zero-mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure contains liquid.

[0343] In some embodiments, the sequencing method further comprises step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules and a plurality of oligonucleotide sequencing primers under conditions suitable for binding of individual immobilized sequencing polymerases to single-stranded circular template molecules and for hybridization of individual sequencing primers to individual single-stranded circular template molecules, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, the individual sequencing primers hybridize to a universal sequencing primer binding site on the single-stranded circular template molecule.

[0344] In some embodiments, the sequencing method further comprises step (c): contacting the plurality of polymerase / template / primer complexes with a plurality of phosphate-chain-labeled nucleotides, each phosphate-chain-labeled nucleotide comprising an aromatic base, a pentose sugar (e.g., ribose or deoxyribose), and a phosphate chain comprising 3 - 20 phosphate groups, wherein the terminal phosphate group is linked to a detectable reporter moiety (e.g., a fluorophore). The first phosphate group, the second phosphate group, and the third phosphate group may be referred to as the α, β, and γ phosphate groups. In some embodiments, the particular detectable reporter moiety attached to the terminal phosphate group corresponds to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase. In some embodiments, the plurality of polymerase / template / primer complexes are contacted with the plurality of phosphate-chain-labeled nucleotides under conditions suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is capable of binding to a complementary phosphate-chain-labeled nucleotide and incorporating a complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, the polymerase-catalyzed nucleotide incorporation reaction cleaves between the α phosphate group and the β phosphate group, thereby releasing the polyphosphate chain linked to the fluorophore.

[0345] In some embodiments, the sequencing method further comprises step (d): detecting the fluorescence signal emitted by the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated into the end of the sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated into the end of the sequencing primer.

[0346] In some embodiments, the sequencing method further comprises step (d): repeating steps (c) to (d) at least once. In some embodiments, the sequencing method using phosphate-chain-labeled nucleotides may be according to U.S. Patent Nos. 7,170,050; 7,302,146; and / or 7,405,281.

[0347] Sequencing Polymerase

[0348] The present disclosure provides methods for sequencing nucleic acid molecules, wherein any of the sequencing methods described herein employ at least one type of sequencing polymerase and a plurality of nucleotides, or employ at least one type of sequencing polymerase, a plurality of nucleotides, and a plurality of multivalent molecules. In some embodiments, the sequencing polymerase is capable of incorporating complementary nucleotides opposite nucleotides in a template molecule. In some embodiments, the sequencing polymerase is capable of binding to complementary nucleotide units of a multivalent molecule opposite nucleotides in a template molecule. In some embodiments, the plurality of sequencing polymerases includes recombinant mutant polymerases.

[0349] Examples of suitable polymerases for sequencing with nucleotides and / or multivalent molecules include, but are not limited to: Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales; Candidatus subterraneus; Candidatus hades; Euryarchaeota; Thermoplasmata; Thermococcus polymerase, such as Thermococcus litoralis, bacteriophage T7 DNA polymerase; human α, δ, and ε DNA polymerases; bacteriophage polymerases, such as T4, RB69, and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; Escherichia coli DNA polymerase IIIα and ε; 9°N polymerase; reverse transcriptases, such as HIV type M or O reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Additional non-limiting examples of DNA polymerases include those from various archaea genera (such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Thermofilum, Thermococcus, Thermoproteus, Sulfolobus, Thermococcus, and Vulcanisaeta, etc. or variants thereof), including such polymerases known in the art, such as 9°N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.

[0350] Nucleotide

[0351] The present disclosure provides methods for sequencing nucleic acid molecules, wherein any of the sequencing methods described herein employs at least one nucleotide. A nucleotide includes a base, a sugar, and at least one phosphate group. In some embodiments, at least one nucleotide among a plurality of nucleotides comprises an aromatic base, a pentose sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 - 10 phosphate groups). The plurality of nucleotides may comprise at least one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of nucleotides may comprise a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one nucleotide among the plurality of nucleotides is not a nucleotide analogue. In some embodiments, at least one nucleotide among the plurality of nucleotides comprises a nucleotide analogue.

[0352] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide among the plurality of nucleotides comprises a chain of one, two, or three phosphorus atoms, wherein the chain is typically attached to the 5'-carbon of the sugar moiety via an ester bond or a phosphoramide bond. In some embodiments, at least one nucleotide among the plurality of nucleotides is an analogue having a phosphorus chain, wherein the phosphorus atoms are linked together by intervening O, S, NH, methylene, or ethylene groups. In some embodiments, the phosphorus atoms in the chain include substituted side groups (including O, S, or BH3). In some embodiments, the chain includes phosphate groups substituted with analogues, the analogues including phosphoramide, thiophosphate, dithiophosphate, and O-methylphosphoramidite groups.

[0353] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, at least one nucleotide of the plurality of nucleotides includes a terminator nucleotide analogue having a chain termination moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the chain termination moiety may inhibit polymerase-catalyzed incorporation of subsequent nucleotide units or free nucleotides into the nascent strand during primer extension reactions. In some embodiments, the chain termination moiety is linked to the 3'-sugar moiety, wherein the sugar comprises a ribose or deoxyribose moiety. In some embodiments, the chain termination moiety can be removed / cleaved from the 3'-sugar moiety to produce a nucleotide having a 3'-OH sugar group that can be extended with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain termination moiety comprises an alkyl, alkenyl, alkynyl, allyl, aryl, benzyl, azide group, amine group, amide group, ketone group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, silyl group, or acetal group. In some embodiments, the chain termination moiety can be cleaved / removed from the nucleotide, e.g., by reacting the chain termination moiety with a chemical agent, pH change, light, or heat. In some embodiments, the chain termination moieties alkyl, alkenyl, alkynyl, and allyl can be cleaved with tetrakis(triphenylphosphine)-palladium(0) (Pd(PPh3)4) and piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the chain termination moieties aryl and benzyl can be cleaved with H2 Pd / C. In some embodiments, the chain termination moieties amine, amide, ketone, isocyanate, phosphate, thio, disulfide can be cleaved with a phosphine or thiol group (including β-mercaptoethanol or dithiothreitol (DTT)). In some embodiments, the chain termination moiety carbonate can be cleaved with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the chain termination moieties urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride. In some embodiments, the chain termination moiety can be cleaved / removed with nitrous acid. In some embodiments, the chain termination moiety can be cleaved / removed using a solution comprising nitrite (e.g., a combination of nitrite with an acid such as acetic acid, sulfuric acid, or nitric acid). In some additional embodiments, the solution can comprise an organic acid.

[0354] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide of the plurality of nucleotides includes a terminator nucleotide analog that has a chain termination moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the chain termination moiety comprises an azide, an azido group, or an azidomethyl group. In some embodiments, the chain termination moiety comprises 3'-O-azido or 3'-O-azidomethyl. In some embodiments, the chain termination moieties azide, azido, and azidomethyl can be cleaved / removed with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bissulfonated triphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP). In some embodiments, a chain termination moiety comprising one or more of 3'-O-amino, 3'-O-aminomethyl, 3'-O-methylamino, or a derivative thereof can be cleaved with nitrous acid by a mechanism utilizing nitrous acid or using a solution comprising nitrous acid. In some embodiments, a chain termination moiety comprising one or more of 3'-O-amino, 3'-O-aminomethyl, 3'-O-methylamino, or a derivative thereof can be cleaved using a solution comprising nitrite. In some embodiments, for example, nitrite can be combined or contacted with an acid such as acetic acid, sulfuric acid, or nitric acid. In some additional embodiments, for example, nitrite can be combined or contacted with an organic acid (e.g., formic acid, acetic acid, propionic acid, butyric acid, isobutyric acid, etc.). In some embodiments, the chain termination moiety comprises a 3'-acetal moiety that can be cleaved with a palladium deblocking reagent (e.g., Pd(0)).

[0355] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, the nucleotide comprises a chain termination moiety selected from the group consisting of: 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-mercapto, 3'-aminomethyl, 3'-ethyl, 3'-butyl, 3'-tert-butyl, 3'-fluorenylmethoxycarbonyl, 3'-tert-butoxycarbonyl, 3'-O-alkylhydroxyamino, 3'-thiophosphate, 3-O-benzyl and 3'-O-benzyl, 3-acetal moiety, or a derivative thereof.

[0356] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, the plurality of nucleotides includes a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety includes a fluorophore. In some embodiments, the fluorophore is linked to a nucleobase. In some embodiments, the fluorophore is linked to the nucleobase with a linker that is cleavable / removable from the base. In some embodiments, at least one nucleotide of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) attached to a nucleotide can correspond to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase.

[0357] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, the cleavable linker on the nucleobase includes a cleavable moiety that includes an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base can be cleaved / removed from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and allyl can be cleaved with tetrakis(triphenylphosphine)-palladium(0) (Pd(PPh3)4) and piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the cleavable moieties aryl and benzyl can be cleaved with H2 Pd / C. In some embodiments, the cleavable moieties amine, amide, keto, isocyanate, phosphate, sulfur, disulfide can be cleaved with a phosphine or with a thiol group including β-mercaptoethanol or dithiothreitol (DTT). In some embodiments, the cleavable moiety carbonate can be cleaved with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the cleavable moieties urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.

[0358] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, the cleavable linker on the nucleobase comprises a cleavable moiety that includes an azide, azido group, or azidomethyl group. In some embodiments, the cleavable moieties azide, azido group, and azidomethyl group can be cleaved / removed with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfonated triphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP).

[0359] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the cleavable linker on the nucleobase have the same or different cleavable moieties. In some embodiments, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) attached to the base and the detectable reporter moiety can be chemically cleaved / removed with the same chemical agent. In some embodiments, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) attached to the base and the detectable reporter moiety can be chemically cleaved / removed with different chemical agents.

[0360] Multivalent Molecule

[0361] The present disclosure provides methods for sequencing nucleic acid molecules, wherein any of the sequencing methods described herein employs at least one multivalent molecule. In some embodiments, the multivalent molecule comprises a plurality of nucleotide arms attached to a core and having any configuration, including starburst, spiral ladder, or bottlebrush configurations (e.g., Figure 9 ). The multivalent molecule includes: (1) a core; and (2) a plurality of nucleotide arms, which comprise: (i) a core attachment portion, (ii) a spacer comprising a PEG portion, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, and wherein the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit through the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, wherein both linker chains have 2 to 6 subunits. In some embodiments, the linker also comprises an aromatic moiety. Exemplary nucleotide arms are shown in Figure 13 . Exemplary multivalent molecules are shown in Figures 9 - 12 . Exemplary spacers are shown in Figure 14 (top), and exemplary linkers are shown in Figure 15 (bottom) and Figure 15 . Exemplary nucleotides attached to the linker are shown in Figures 16 - 19in. Figure 20 An exemplary biotinylated nucleotide arm is shown therein.

[0362] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, and wherein the plurality of nucleotide arms have the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.

[0363] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein each arm comprises a nucleotide unit. The nucleotide unit comprises an aromatic base, a pentose sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 - 10 phosphate groups). The plurality of multivalent molecules can include a type of multivalent molecule having one type of nucleotide unit selected from the group consisting of: dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of multivalent molecules can include a mixture of any combination of two or more types of multivalent molecules, wherein the individual multivalent molecules in the mixture include nucleotide units selected from the group consisting of: dATP, dGTP, dCTP, dTTP, and / or dUTP.

[0364] In some embodiments, the nucleotide unit comprises a chain of one, two, or three phosphorus atoms, which chain is typically linked to the 5'-carbon of the sugar moiety by an ester or phosphoramide bond. In some embodiments, at least one nucleotide unit is a nucleotide analogue having a phosphorus chain, wherein the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene groups. In some embodiments, the phosphorus atoms in the chain include substituted side groups (including O, S, or BH3). In some embodiments, the chain includes phosphate groups substituted with analogues, the analogues including phosphoramide, thiophosphate, dithiophosphate, and O-methylphosphoramidite groups.

[0365] In some embodiments, the multivalent molecule comprises a core linked to a plurality of nucleotide arms, and wherein each nucleotide arm comprises a nucleotide unit that is a nucleotide analogue having a chain terminating moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the nucleotide unit comprises a chain terminating moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the chain terminating moiety may inhibit the polymerase-catalyzed incorporation of subsequent nucleotide units or free nucleotides into the nascent strand during primer extension reactions. In some embodiments, the chain terminating moiety is linked to the 3'-sugar position, wherein the sugar comprises a ribose or deoxyribose moiety. In some embodiments, the chain terminating moiety may be removed / cleaved from the 3'-sugar position to produce a nucleotide having a 3'-OH sugar group that may be extended with subsequent nucleotides in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain terminating moiety comprises an alkyl, alkenyl, alkynyl, allyl, aryl, benzyl, azide group, amine group, amide group, ketone group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, or silyl group. In some embodiments, the chain terminating moiety may be cleaved / removed from the nucleotide unit, e.g., by reacting the chain terminating moiety with a chemical agent, pH change, light, or heat. In some embodiments, the chain terminating moieties alkyl, alkenyl, alkynyl, and allyl may be cleaved with tetrakis(triphenylphosphine)-palladium(0) (Pd(PPh3)4) and piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the chain terminating moieties aryl and benzyl may be cleaved with H2 Pd / C. In some embodiments, the chain terminating moieties amine, amide, ketone, isocyanate, phosphate, thio, disulfide may be cleaved with a phosphine or thiol group (including β-mercaptoethanol or dithiothreitol (DTT)). In some embodiments, the chain terminating moiety carbonate may be cleaved with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the chain terminating moieties urea and silyl may be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.

[0366] In some embodiments, the nucleotide unit comprises a chain-terminating moiety (e.g., a blocking moiety) at the 2'-position of the sugar, at the 3'-position of the sugar, or at both the 2'- and 3'-positions of the sugar. In some embodiments, the chain-terminating moiety comprises an azide, an azido group, or an azidomethyl group. In some embodiments, the chain-terminating moiety comprises 3'-O-azido or 3'-O-azidomethyl. In some embodiments, the chain-terminating moieties azide, azido, and azidomethyl can be cleaved / removed with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bissulfotriphenylphosphine (BS-TPP) or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP).

[0367] In some embodiments, the nucleotide unit comprises a chain-terminating moiety selected from the group consisting of: 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-mercapto, 3'-aminomethyl, 3'-ethyl, 3'-butyl, 3'-tert-butyl, 3'-fluorenylmethoxycarbonyl, 3'-tert-butoxycarbonyl, 3'-O-alkylhydroxyamino, 3'-thiophosphate, and 3-O-benzyl or derivatives thereof.

[0368] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein the nucleotide arms comprise a spacer, a linker, and a nucleotide unit, and wherein the core, linker, and / or nucleotide unit are labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) attached to the multivalent molecule can correspond to the base of the nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase.

[0369] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter moiety. In some embodiments, the detectable reporter moiety is attached to the nucleobase. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) attached to the multivalent molecule can correspond to the base of the nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase.

[0370] In some embodiments, the core of the multivalent molecule comprises an avidin-like or streptavidin-like moiety, and the core attachment moiety comprises biotin. In some embodiments, the core comprises a streptavidin-type or avidin-type moiety that includes avidin protein, as well as any derivatives, analogs, and other non-natural forms of avidin that can bind to at least one biotin moiety. Other forms of the avidin moiety include native and recombinant avidin and streptavidin, and derived molecules, such as, for example, deglycosylated avidin and truncated streptavidin. For example, the avidin moiety includes a deglycosylated form of avidin, bacterial streptavidin produced by Streptomyces species (e.g., Streptomyces avidinii), and derived forms, such as N-acyl avidin, such as N-acetyl, N-phthaloyl, and N-succinyl avidin, as well as the commercially available products EXTRAVIDIN, CAPTAVIDIN, NEUTRAVIDIN, and NEUTRALITE AVIDIN.

[0371] In some embodiments, any method described herein for sequencing a nucleic acid molecule can include forming a binding complex, where the binding complex comprises (i) a polymerase, a nucleic acid template molecule that forms a duplex with a primer, and a nucleotide, or the binding complex comprises (ii) a polymerase, a nucleic acid template molecule that forms a duplex with a primer, and a nucleotide unit of a multivalent molecule. In some embodiments, the dwell time of the binding complex is greater than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. The dwell time of the binding complex is greater than about 0.1 to 0.25 seconds, or about 0.25 to 0.5 seconds, or about 0.5 to 0.75 seconds, or about 0.75 to 1 second, or about 1 to 2 seconds, or about 2 to 3 seconds, or about 3 to 4 seconds, or about 4 to 5 seconds, and / or wherein the method is or can be performed at a temperature that is at or above 15°C, at or above 20°C, at or above 25°C, at or above 35°C, at or above 37°C, at or above 42°C, at or above 55°C, at or above 60°C, or at or above 72°C, or at or above 80°C, or within a range defined by any one of the foregoing. The binding complex (e.g., ternary complex) remains stable prior to being subjected to conditions that cause dissociation of the interactions between any of the polymerase, template molecule, primer, and / or nucleotide unit or nucleotide. For example, the dissociation conditions include contacting the binding complex with any one of a detergent, EDTA, and / or water, or any combination thereof. In some embodiments, the present disclosure provides the method, wherein the binding complex is deposited on, attached to, or hybridized to a surface that exhibits a contrast-to-noise ratio greater than 20 in the detection step. In some embodiments, the present disclosure provides the method, wherein the contacting is performed under conditions where the binding complex is stabilized when the nucleotide or nucleotide unit is complementary to the next base of the template nucleic acid and destabilized when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.

[0372] Compacted Oligonucleotide

[0373] The compacted oligonucleotide comprises a single-stranded linear oligonucleotide having a 5' region that can hybridize to a first portion of a concatemer molecule and a compacted oligonucleotide having a 3' region that can hybridize to a second portion of the concatemer molecule (e.g., the same concatemer molecule). In some embodiments, hybridization of the compacted oligonucleotide to a separate concatemer molecule causes the concatemer molecule to collapse or fold into a DNA nanosphere, which is more compact in both shape and size compared to the non-collapsed DNA molecule. The point image of the DNA nanosphere can be represented as a Gaussian spot, and the size can be measured as the full width at half maximum (FWHM). A smaller point size, as indicated by a smaller FWHM, generally correlates with an improved image of the point. In some embodiments, the FWHM of the DNA nanosphere point can be about 10 μm or less. The DNA nanosphere can be a compact nucleic acid structure that has a smaller full width at half maximum (FWHM) compared to the concatemer that has not collapsed / folded into a DNA nanosphere.

[0374] In some embodiments, the compacted oligonucleotide comprises a single-stranded oligonucleotide comprising DNA, RNA, or a combination of DNA and RNA. The compacted oligonucleotide can be of any length, including lengths of 20 - 150 nucleotides, or 30 - 100 nucleotides, or 40 - 80 nucleotides.

[0375] In some embodiments, the compacted oligonucleotide comprises a 5' region and a 3' region, and optionally an intermediate region between the 5' region and the 3' region. The intermediate region can be of any length, e.g., a length of about 2 - 20 nucleotides. The intermediate region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT, or UUU). The intermediate region comprises a non-homopolymer sequence.

[0376] The 5' region of the compacted oligonucleotide can be fully or partially complementary to the first portion of the concatemer molecule along its length. The 3' region of the compacted oligonucleotide can be fully or partially complementary to the second portion of the concatemer molecule along its length. The 5' region of the compacted oligonucleotide can hybridize to a first universal sequence portion of the concatemer molecule. The 3' region of the compacted oligonucleotide can hybridize to a second universal sequence portion of the concatemer molecule. The 5' and 3' regions of the compacted oligonucleotide can hybridize to the concatemer to pull the distal portions of the concatemer together, thereby causing the concatemer to compact to form a DNA nanosphere.

[0377] The 5' region of the compacted oligonucleotide can have the same sequence as the 3' region. The 5' region of the compacted oligonucleotide can have a different sequence from the 3' region. The 3' region of the compacted oligonucleotide can have a sequence opposite to the 5' region.

[0378] In some embodiments, sequence data may be obtained by nanopore sequencing, which includes sequencing a nucleic acid by translocation across a membrane (e.g., through a pore), and wherein sequence reads or base calling are performed by measuring one or more signals (e.g., impedance, current, voltage, or capacitance) during the translocation event. In some embodiments, the identity of a nucleotide may be determined by unique electrical signatures such as the timing, duration, magnitude, or shape of a current block, impedance changes, voltage changes, or capacitance changes. Sequencing of nucleic acids by translocation across a membrane and / or through a pore does not exclude alternative detection methods such as optical, chemical, biochemical, fluorescent, luminescent, magnetic, electromagnetic, acoustic, or electroacoustic detection.

[0379] Supports and Low Nonspecific Coatings

[0380] In some embodiments, Figure 1 the flow cell 112 in

[0381] The low non-specific binding coating comprises one or more layers ( Figure 20)。In some embodiments, multiple surface primers are immobilized to a low non-specific binding coating. In some embodiments, at least one surface primer is embedded within the low non-specific binding coating. The low non-specific binding coating enables improved nucleic acid hybridization and amplification performance. Generally, a support comprises a substrate (or support structure), one or more layers of a covalently or non-covalently attached low-binding chemical modification layer (e.g., a silane layer, a polymer film), and one or more covalently or non-covalently attached surface primers that can be used to tether single-stranded nucleic acid library molecules to the support. In some embodiments, the formulation of the coating (e.g., the chemical composition of one or more layers, the coupling chemistry used to crosslink one or more layers to the support and / or to each other, and the total number of layers) can vary such that non-specific binding of proteins, nucleic acid molecules, and other hybridization and amplification reaction components to the coating is minimized or reduced relative to a comparable single layer. The formulations of the coatings described herein can be varied such that non-specific hybridization to the coating is minimized or reduced relative to a comparable single layer. The formulation of the coating can be varied such that non-specific amplification on the coating is minimized or reduced relative to a comparable single layer. The formulation of the coating can be varied such that the specific amplification rate and / or yield on the coating is maximized. In some cases disclosed herein, amplification levels suitable for detection are achieved in no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more than 30 amplification cycles.

[0382] A support structure comprising one or more chemically modified layers (e.g., a layer of a low non-specific binding polymer) can be standalone or can be integrated into another structure or component. For example, in some embodiments, the support structure can include one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure can comprise one or more surfaces within a microplate format (e.g., the bottom surface of a well in a microplate). In some embodiments, the support structure includes the inner surface of a capillary (such as the lumen surface). In some embodiments, the support structure includes the inner surface of a capillary (such as the lumen surface) etched into a planar chip.

[0383] The attachment chemistry for grafting a first chemically modified layer to the surface of a support will generally depend on both the material from which the surface is made and the chemistry of the layer. In some embodiments, the first layer may be covalently attached to the surface. In some embodiments, the first layer may be non-covalently attached (e.g., adsorbed to the support) to the support through non-covalent interactions (such as electrostatic interactions, hydrogen bonding, or van der Waals interactions) between a carrier and the molecular components of the first layer. In either case, the support may be treated prior to the attachment or deposition of the first layer. Any of a variety of surface preparation techniques known to those skilled in the art may be used to clean or treat the surface. For example, a glass or silicon surface may be acid washed using a Piranha solution (a mixture of sulfuric acid (H2SO4) and hydrogen peroxide (H2O2)), base treated in KOH and NaOH, and / or cleaned using an oxygen plasma treatment method.

[0384] Silane chemistry constitutes a non-limiting method for covalently modifying silanol groups on a glass or silicon surface to attach more reactive functional groups (e.g., amine or carboxyl), which can then be used to couple linker molecules (e.g., linear hydrocarbon molecules of various lengths (such as C6, C12, C18 hydrocarbons) or linear polyethylene glycol (PEG) molecules) or layer molecules (e.g., branched PEG molecules or other polymers) to the surface. Examples of suitable silanes that can be used to produce any of the disclosed low-binding coatings include, but are not limited to: (3-aminopropyl)trimethoxysilane (APTMS), (3-aminopropyl)triethoxysilane (APTES), any of a variety of PEG-silanes (e.g., having molecular weights including 1K, 2K, 5K, 10K, 20K, etc.), amino-PEG silane (i.e., containing a free amino functional group), maleimide-PEG silane, biotin-PEG silane, and the like.

[0385] Any one of a variety of molecules known to those skilled in the art (including but not limited to amino acids, peptides, nucleotides, oligonucleotides, other monomers or polymers, or combinations thereof) can be used to generate one or more chemically modified layers on a support, where the choice of components used can vary to alter one or more properties of the layer (e.g., the surface density of functional groups and / or tethered oligonucleotide primers, the hydrophilicity / hydrophobicity of the layer, or the three-dimensional properties of the layer (i.e., "thickness")). Examples of polymers that can be used to generate one or more layers of low non-specific binding material in any of the disclosed coatings include but are not limited to: polyethylene glycol (PEG) of various molecular weights and branching structures, streptavidin, polyacrylamide, polyester, dextran, polylysine and polylysine copolymers, or any combination thereof. Examples of conjugation chemistries that can be used to graft one or more layers of material (e.g., a polymer layer) to a surface and / or crosslink layers to each other include but are not limited to: biotin-streptavidin interaction (or variants thereof), his-tag-Ni / NTA conjugation chemistry, methoxy ether conjugation chemistry, carboxylate conjugation chemistry, amine conjugation chemistry, NHS ester, maleimide, thiol, epoxy resin, azide, hydrazide, alkyne, isocyanate, and silane.

[0386] The low non-specific binding surface coating can be applied uniformly across the support. Alternatively, the surface coating can be patterned such that the chemically modified layer is restricted to one or more discrete regions of the support. For example, lithographic techniques can be used to pattern the coating to produce an ordered array or random pattern of chemically modified regions on the support. Alternatively or in combination, techniques such as contact printing and / or inkjet printing can be used to pattern the coating. In some embodiments, the ordered array or random pattern of chemically modified regions can include at least 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more discrete regions.

[0387] In some embodiments, the low non-specific binding coating comprises a hydrophilic polymer that is non-specifically adsorbed or covalently grafted to the support. Typically, passivation is performed using poly(ethylene glycol) (PEG, also known as poly(ethylene oxide) (PEO) or poly(oxyethylene)) or other hydrophilic polymers having different molecular weights and end groups that are attached to the support using, for example, silane chemistry. End groups away from the surface can include, but are not limited to: biotin, methoxy ether, carboxylate, amine, NHS ester, maleimide, and disilane. In some embodiments, two or more layers of hydrophilic polymers (e.g., linear polymers, branched polymers, or multi-branched polymers) can be deposited on the surface. In some embodiments, the two or more layers can be covalently coupled to each other or internally crosslinked to improve the stability of the resulting coating. In some embodiments, surface primers (or other biomolecules, e.g., enzymes or antibodies) having different nucleotide sequences and / or base modifications can be tethered to the resulting layer at various surface densities. In some embodiments, for example, both the surface functional group density and the surface primer concentration can be varied to obtain a desired range of surface primer densities. Additionally, the surface primer density can be controlled by diluting the surface primer with other molecules carrying the same functional group. For example, amine-labeled surface primers can be diluted with amine-labeled polyethylene glycol in a reaction with an NHS-ester coated surface to reduce the final primer density. Surface primers of different lengths with a linker between the hybridization region and the surface attachment functional group can also be applied to control the surface density. Examples of suitable linkers include poly-T and poly-A chains (e.g., 0 to 20 bases) at the 5' end of the primer, PEG linkers (e.g., 3 to 20 monomer units), and carbon chains (e.g., C6, C12, C18, etc.). To measure the primer density, fluorescently labeled primers can be tethered to the surface and then the fluorescence readings are compared to the fluorescence readings for a dye solution of known concentration.

[0388] In some embodiments, the low non-specific binding coating includes a functionalized polymer coating that is at least covalently bound to a portion of the support via a chemical group on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamido pentyl)acrylamide-co-acrylamide (PAZAM).

[0389] To scale primer surface density and add an additional dimension to hydrophilic or amphiphilic coatings, supports with multilayer coatings including PEG and other hydrophilic polymers have been developed. By using hydrophilic and amphiphilic surface layering methods (including but not limited to the polymer / copolymer materials described below), the primer loading density on the support can be significantly increased. Traditional PEG coating methods use single-layer primer deposition, which has generally been reported for single-molecule applications but does not produce high copy numbers for nucleic acid amplification applications. As described herein, "layering" can be achieved using traditional crosslinking methods with any compatible polymer or monomer subunit, such that surfaces containing two or more highly crosslinked layers can be sequentially constructed. Examples of suitable polymers include but are not limited to: streptavidin, polyacrylamide, polyester, dextran, polylysine, and copolymers of polylysine and PEG. In some embodiments, the different layers can be attached to each other by any of a variety of conjugation reactions (including but not limited to: biotin-streptavidin binding, azide-alkyne click reaction, amine-NHS ester reaction, thiol-maleimide reaction, and ionic interactions between positively charged and negatively charged polymers). In some embodiments, high primer density materials can be constructed in solution and then layered onto the surface in multiple steps.

[0390] Examples of materials from which the support structure can be fabricated include but are not limited to: glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), cycloolefin polymer (COP), cycloolefin copolymer (COC), polyethylene terephthalate (PET)) or any combination thereof. A variety of compositions of both glass and plastic support structures are envisioned.

[0391] The support structure can be presented in any of a variety of geometries and sizes known to those skilled in the art and can include any of a variety of materials known to those skilled in the art. For example, the support structure can be locally planar (e.g., comprising the surface of a microscope slide or microscope coverslip). Generally speaking, the support structure can be cylindrical (e.g., including the inner surface of a capillary or capillary tube), spherical (e.g., including the outer surface of a non-porous bead), or irregular (e.g., including the outer surface of an irregularly shaped non-porous bead or particle). In some embodiments, the surface of the support structure for nucleic acid hybridization and amplification can be a solid non-porous surface. In some embodiments, the surface of the support structure for nucleic acid hybridization and amplification can be porous such that the coatings described herein penetrate the porous surface and nucleic acid hybridization and amplification reactions performed thereon can occur within the pores.

[0392] A support structure including one or more chemically modified layers (e.g., a layer of a low non-specific binding polymer) can be standalone or can be integrated into another structure or component. For example, the support structure can include one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure can include one or more surfaces within a microplate format (e.g., the bottom surface of a well in a microplate). In some embodiments, the support structure includes the inner surface of a capillary (such as the lumen surface). In some embodiments, the support structure includes the inner surface of a capillary (such as the lumen surface) etched into a planar chip.

[0393] As noted, the low non-specific binding supports of the present disclosure exhibit reduced non-specific binding of proteins, nucleic acids, and other components of hybridization and / or amplification formulations for solid-phase nucleic acid amplification. The degree of non-specific binding exhibited by a given support surface can be evaluated qualitatively or quantitatively. For example, the surface is exposed to a fluorescent dye (e.g., cyanine (such as Cy3 or Cy5, etc.), fluorescein, coumarin, rhodamine, etc. or other dyes disclosed herein), a fluorescently labeled nucleotide, a fluorescently labeled oligonucleotide, and / or a fluorescently labeled protein (e.g., polymerase) under a set of standardized conditions, and then a specified wash protocol and fluorescence imaging can be used as a qualitative tool to compare non-specific binding on supports containing different surface formulations. In some embodiments, the surface is exposed to a fluorescent dye, a fluorescently labeled nucleotide, a fluorescently labeled oligonucleotide, and / or a fluorescently labeled protein (e.g., polymerase) under a set of standardized conditions, and then a specified wash protocol and fluorescence imaging can be used as a qualitative tool to compare non-specific binding on carriers including different surface formulations—provided that care is taken to ensure that the fluorescence imaging is performed under conditions where the fluorescence signal is linearly related (or related in a predictable manner) to the number of fluorophores on the carrier surface and using appropriate calibration standards (e.g., conditions where signal saturation and / or self-quenching of the fluorophores is not a problem). In some embodiments, other techniques known to those skilled in the art (e.g., radioisotope labeling and counting methods) can be used to quantitatively evaluate the degree of non-specific binding exhibited by different support surface formulations of the present disclosure.

[0394] Some of the surfaces disclosed herein exhibit a ratio of specific binding to non - specific binding of a fluorophore (such as Cy3) of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value covered by the ranges herein. Some of the surfaces disclosed herein exhibit a ratio of specific fluorescence to non - specific fluorescence of a fluorophore (such as Cy3) of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value covered by the ranges herein.

[0395] A normalization scheme for the following can be used to evaluate the extent of non - specific binding exhibited by the disclosed low - binding supports: Under a set of standardized incubation and washing conditions, contact the surface with a labeled protein (e.g., bovine serum albumin (BSA), streptavidin, DNA polymerase, reverse transcriptase, helicase, single - stranded binding protein (SSB), etc., or any combination thereof), a labeled nucleotide, a labeled oligonucleotide, etc., and then detect the amount of label remaining on the surface and compare the signal obtained therefrom with an appropriate calibration standard. In some embodiments, the label can include a fluorescent label. In some embodiments, the label can include a radioisotope. In some embodiments, the label can include any other detectable label known to those skilled in the art. In some embodiments, thus, the extent of non - specific binding exhibited by a given support surface formulation can be evaluated based on the number of protein molecules (or nucleic acid molecules or other molecules) non - specifically bound per unit area. In some embodiments, the low - binding supports of the present disclosure can exhibit the following non - specific protein binding (or non - specific binding of other specified molecules (e.g., cyanines (such as Cy3 or Cy5, etc.), fluorescein, coumarin, rhodamine, etc., or other dyes disclosed herein)): less than 0.001 molecules / μm 2 less than 0.01 molecules / μm 2 less than 0.1 molecules / μm 2 less than 0.25 molecules / μm 2 less than 0.5 molecules / μm 2 less than 1 molecule / μm 2 less than 10 molecules / μm 2 less than 100 molecules / μm 2 or less than 1,000 molecules / μm 2 . Those skilled in the art will recognize that a given support surface of the present disclosure can exhibit any value within this range (e.g., less than 86 molecules per μm2 (of) non-specific binding. For example, after contacting with a solution of 1 μM Cy3-labeled streptavidin (GE Amersham) in phosphate buffered saline (PBS) buffer for 15 minutes and then rinsing three times with deionized water, some of the modified surfaces disclosed herein exhibit less than 0.5 molecules / μm 2 of non-specific protein binding. Some of the modified surfaces disclosed herein exhibit less than 0.25 molecules / μm 2Non-specific binding of Cy3 dye molecules. In an independent non-specific binding assay, 1 μM labeled Cy3 SA (ThermoFisher), 1 μM Cy5 SA dye (ThermoFisher), 10 μM aminoallyl-dUTP-ATTO-647N (Jena Biosciences), 10 μM aminoallyl-dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 μM aminoallyl-dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 μM 7-propynylamino-7-deaza-dGTP-Cy5 (Jena Biosciences), and 10 μM 7-propynylamino-7-deaza-dGTP-Cy3 (Jena Biosciences) were incubated at 37 °C for 15 minutes in a 384-well plate format on a low-binding coated support. Each well was rinsed 2 - 3 times with 50 μl of deionized RNase / DNase-free water and 2 - 3 times with 25 mM ACES buffer (pH 7.4). The 384-well plate was imaged on a GE Typhoon instrument with a PMT gain setting of 800 and a resolution of 50 μm to 100 μm using Cy3, AF555, or Cy5 filter sets as specified by the manufacturer (depending on the dye test performed). For higher resolution imaging, images were collected on an Olympus IX83 microscope (e.g., inverted fluorescence microscope) (Olympus Corp., Center Valley, Pa.) with a total internal reflection fluorescence (TIRF) objective (100×, 1.5 NA, Olympus), a CCD camera (e.g., Olympus EM-CCD monochrome camera, Olympus XM-10 monochrome camera, or Olympus DP80 color and monochrome camera), an illumination source (e.g., Olympus 100W mercury lamp, Olympus 75W xenon lamp, or Olympus U-HGLGPS fluorescence light source), and an excitation wavelength of 532 nm or 635 nm. Dichroic mirrors were purchased from Semrock (IDEX Health & Science, LLC, Rochester, N.Y.), e.g., 405 nm, 488 nm, 532 nm, or 633 nm dichroic reflectors / beam splitters, and bandpass filters were selected as 532LP or 645LP integrated with the appropriate excitation wavelength. Some of the modified surfaces disclosed herein exhibited less than 0.25 molecules / μm 2 of non-specific binding of dye molecules. In some embodiments, the coated carrier was immersed in a buffer (e.g., 25 mM ACES, pH 7.4) while acquiring images.

[0396] In some embodiments, the surfaces disclosed herein exhibit a ratio of specific to non-specific binding of a fluorophore (such as Cy3) of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges herein. In some embodiments, the surfaces disclosed herein exhibit a ratio of specific to non-specific fluorescent signals for a fluorophore (such as Cy3) of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges herein.

[0397] A low-background surface in accordance with the disclosure herein can exhibit a ratio of specific dye attachment (e.g., Cy3 attachment) to non-specific dye adsorption (e.g., Cy3 dye adsorption) of at least 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 30:1, 40:1, 50:1 or more than 50 attached specific dye molecules / non-specifically adsorbed molecules. Similarly, when subjected to excitation energy, a low-background surface in accordance with the disclosure herein that has been attached to a fluorophore (e.g., Cy3) can exhibit a ratio of specific fluorescent signals (e.g., generated from Cy3-labeled oligonucleotides attached to the surface) to non-specifically adsorbed dye fluorescent signals of at least 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 30:1, 40:1, 50:1 or more than 50:1.

[0398] In some embodiments, the degree of hydrophilicity (or “wettability” with an aqueous solution) of the disclosed support surface can be evaluated, for example, by measuring the water contact angle, where a small drop of water is placed on the surface and its contact angle with the surface is measured using, for example, an optical tensiometer. In some embodiments, the static contact angle can be determined. In some embodiments, the advancing or receding contact angle can be determined. In some embodiments, the range of water contact angles for the hydrophilic low-binding support surface disclosed herein can be from about 0 degrees to about 30 degrees. In some embodiments, the water contact angle for the hydrophilic low-binding support surface disclosed herein can be no more than 50 degrees, 40 degrees, 30 degrees, 25 degrees, 20 degrees, 18 degrees, 16 degrees, 14 degrees, 12 degrees, 10 degrees, 8 degrees, 6 degrees, 4 degrees, 2 degrees or 1 degree. In many cases, the contact angle is no more than 40 degrees. Those skilled in the art will recognize that a given hydrophilic low-binding support surface of the present disclosure can exhibit a water contact angle having a value at any position within this range.

[0399] In some embodiments, the hydrophilic surfaces disclosed herein help reduce the wash time for a biometric assay, typically due to reduced non-specific binding of biomolecules to the low-binding surface. In some embodiments, a sufficient wash step can be performed in less than 60 seconds, 50 seconds, 40 seconds, 30 seconds, 20 seconds, 15 seconds, 10 seconds, or less than 10 seconds. For example, a sufficient wash step can be performed in less than 30 seconds.

[0400] Some of the low-binding surfaces of the present disclosure exhibit significant improvements in terms of stability or persistence upon prolonged exposure to solvents and elevated temperatures or for repeated cycles of solvent exposure or temperature variations. For example, the stability of the disclosed surfaces can be tested by fluorescently labeling functional groups on the surface or tethered biomolecules on the surface (e.g., oligonucleotide primers) and monitoring the fluorescent signal before, during, and after prolonged exposure to solvents and elevated temperatures or repeated cycles of solvent exposure or temperature variations. In some embodiments, the degree of change in fluorescence used to evaluate the quality of the surface can be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% (or any combination of these percentages as measured over these time periods) over a time period of 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 40 minutes, 50 minutes, 60 minutes, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 15 hours, 20 hours, 25 hours, 30 hours, 35 hours, 40 hours, 45 hours, 50 hours, or 100 hours of exposure to solvents and / or elevated temperatures. In some embodiments, the degree of change in fluorescence used to evaluate the quality of the surface can be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% (or any combination of these percentages as measured over this range of cycles) over 5 cycles, 10 cycles, 20 cycles, 30 cycles, 40 cycles, 50 cycles, 60 cycles, 70 cycles, 80 cycles, 90 cycles, 100 cycles, 200 cycles, 300 cycles, 400 cycles, 500 cycles, 600 cycles, 700 cycles, 800 cycles, 900 cycles, or 1,000 cycles of repeated exposure to changes in solvents and / or changes in temperature.

[0401] In some embodiments, the surfaces disclosed herein may exhibit a high ratio of specific signal to non-specific signal or other background. For example, when used in nucleic acid amplification, some surfaces may exhibit an amplification signal that is at least 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold, 75-fold, 100-fold, or greater than 100-fold greater than the signal of an adjacent population-free region of the surface. Similarly, some surfaces exhibit an amplification signal that is at least 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold, 75-fold, 100-fold, or greater than 100-fold greater than the signal of an adjacent amplified nucleic acid population region of the surface.

[0402] In some embodiments, when used in nucleic acid hybridization or amplification applications to generate a polymerase community of hybridized or clonally amplified nucleic acid molecules (e.g., directly or indirectly labeled with a fluorophore), the fluorescence image of the disclosed low-background surface exhibits a contrast-to-noise ratio (CNR) of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 1...

Claims

1. A computer-implemented method for evaluating index sequences of different samples, comprising: generating, by a processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; in response to determining that first flow cell images of the plurality of samples in a first plurality of sequencing cycles corresponding to the plurality of index sequences in a sequencing run have been acquired, generating, by a sequencing system while the sequencing run is in progress, sequencing data comprising a set of sequencing reads for each polymerase colony of the plurality of samples, the set of sequencing reads comprising one or more of the plurality of index sequences; for the set of sequencing reads of each polymerase colony: determining, in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, a value of a mismatch parameter for the polymerase colony; and assigning, in response to determining that the value of the mismatch parameter satisfies a second error tolerance rate, the polymerase colony to the one mapped sample; and evaluating, by the processor, one or more statistical parameters of the sequencing data while acquiring second flow cell images of the plurality of samples in a second plurality of sequencing cycles, wherein the second plurality of sequencing cycles or the second plurality of flow cell images correspond to at least a portion of the reads of a fragment of the DNA sequence of the set of sequencing reads for each polymerase colony or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence; and determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion.

2. A computer-implemented method for evaluating index sequences of different samples, comprising: generating, by a processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; generating, by a sequencing system while the sequencing run of the plurality of samples is in progress, sequencing data comprising a set of sequencing reads for each polymerase colony of the plurality of samples, the set of sequencing reads comprising one or more of the plurality of index sequences; for the set of sequencing reads of each polymerase colony: determining, in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, a value of a mismatch parameter for the polymerase colony; and assigning, in response to determining that the value of the mismatch parameter satisfies a second error tolerance rate, the polymerase colony to the one mapped sample; and evaluating, by the processor, one or more statistical parameters of the sequencing data while the sequencing run of the plurality of samples is in progress; and determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion.

3. A computer-implemented method for evaluating index sequences of different samples, comprising: generating, by a processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; While the sequencing runs of the plurality of samples are in progress, sequencing data is generated by a sequencing system, the sequencing data comprising a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads comprising one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase community: In response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, determining a value of a mismatch parameter for the polymerase community; And In response to determining that the value of the mismatch parameter satisfies a second error tolerance rate, assigning the polymerase community to the one mapped sample; And Evaluating, by the processor, one or more statistical parameters of the sequencing data while generating at least a portion of the reads of the fragment of the DNA sequence or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence for the set of sequencing reads for each polymerase community of the plurality of samples; And Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion of the sequence run.

4. The computer-implemented method according to any one of the preceding claims, further comprising: Generating, by the processor, one or more reference index sequences for each of the plurality of samples, and Wherein generating the one or more mapped index sequences and mapping them to the plurality of samples is based on the one or more reference index sequences for each of the plurality of samples.

5. The computer-implemented method according to any one of the preceding claims, further comprising: Determining, by the processor, a read order for reading the set of sequencing reads for each polymerase community; Determining, by the processor and based on the read order, whether the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence for the set of reads for each polymerase community.

6. The computer-implemented method according to any one of the preceding claims, further comprising: Determining, by the processor and based on the read order, that the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence for the set of reads for each polymerase community; And Determining that the first plurality of flow cell images corresponding to the first plurality of sequencing cycles of the plurality of index sequences in the sequencing run have been acquired.

7. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more mapped index sequences maps to only one of the plurality of samples.

8. The computer-implemented method according to any one of the preceding claims, wherein generating the one or more mapped index sequences and mapping them to the plurality of samples is based on a hash table that applies a hash function to each of the one or more mapped index sequences.

9. The computer-implemented method according to any one of the preceding claims, wherein determining that the value of the mismatch parameter satisfies the second error tolerance rate comprises: determining one or more positions in each of the one or more index sequences, each of the one or more positions corresponding to a mismatched base, a quality score for base identification below a pre-determined quality threshold, or both, wherein the value of the mismatch parameter for the polymerase population is the total number of the one or more positions.

10. The computer-implemented method according to any one of the preceding claims, further comprising: separating, by the processor, each polymerase population based on the corresponding mapped sample to generate sequencing results in a plurality of data files in a pre-determined data format.

11. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.

12. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles or the second plurality of flow cell images do not correspond to the one or more index sequences.

13. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion comprises: automatically terminating the sequencing run before completion in response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion.

14. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion comprises: automatically terminating the generation of the sequencing data before the reads of the fragment of the DNA sequence, the reads of the complementary strand of the fragment of the DNA sequence, or both, in the set of reads for each polymerase population have been completed, in response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion.

15. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion comprises: continuing the sequencing run until it is completed in response to determining that one or more values of the one or more statistical parameters do not satisfy a stop criterion.

16. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion comprises: generating the sequencing data while the sequencing run is still in progress, the sequencing data comprising the reads of the fragment of the DNA sequence, the reads of the complementary strand of the fragment of the DNA sequence, or both, in the set of sequencing reads for each polymerase population, in response to determining that one or more values of the one or more statistical parameters do not satisfy a stop criterion.

17. The computer-implemented method according to any one of the preceding claims, further comprising: Upon completion of the sequencing run, the reads of the fragments of the DNA sequences of the set of sequencing reads for each polymerase community, the reads of the complementary strands of the fragments of the DNA sequences, or both are separated without separating one or more of the plurality of index sequences in the set of sequencing reads.

18. The computer-implemented method according to any one of the preceding claims, wherein evaluating the one or more statistical parameters of the sequencing data comprises: Calculating at least one value of each of the one or more statistical parameters for a plurality of polymerase communities based on the value of the mismatch parameter for each polymerase community; Comparing the at least one value of each of the one or more statistical parameters with a predetermined threshold; And Determining whether the at least one value of each of the one or more statistical parameters meets the predetermined threshold.

19. The computer-implemented method according to any one of the preceding claims, wherein the one or more statistical parameters of the sequencing results comprise one or more of the following: Mismatch rate; Unassigned rate; Error correlation rate; Assignment rate; Match rate; Deletion rate; and Mixed pairing rate.

20. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and different.

21. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and the same.

22. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is 2 mismatched bases out of 9 bases.

23. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 1 deleted base out of 9 bases.

24. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 5%, 10%, 15%, 20%, 21%, 22%, 23%, 24%, 25% or 30%.

25. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences are separated from the reads of the fragments of the DNA sequences of the set of reads for each polymerase community or the reads of the complementary strands of the fragments of the DNA sequences.

26. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is not part of the reads of the fragments of the DNA sequences of the set of reads for each polymerase community or the reads of the complementary strands of the fragments of the DNA sequences.

27. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises a consecutive portion of the reads of the fragments of the DNA sequence in the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

28. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is a part of the reads of the fragments of the DNA sequence in the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

29. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is a consecutive portion of the reads of the fragments of the DNA sequence in the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

30. The computer-implemented method according to any one of the preceding claims, wherein the plurality of samples are located on one or more subtiles or one or more tiles of the flow cell.

31. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragments of the DNA sequence are forward reads, forward complementary reads, or reverse complementary reads.

32. The computer-implemented method according to any one of the preceding claims, wherein the reads of the complementary strand of the fragments of the DNA sequence are forward or reverse reads.

33. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data comprises: generating the sequencing data by processing based on one or more preliminary analysis steps of the first plurality of flow cell images.

34. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the reads of the fragments of the DNA sequence in each set of sequencing reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

35. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further comprises: acquiring, by the optical system of the sequencing system, the first plurality of flow cell images of one or more samples located on the flow cell; performing, by the processor, one or more preliminary analysis steps on the first plurality of flow cell images; or both.

36. The computer-implemented method according to any one of the preceding claims, wherein acquiring the first plurality of flow cell images of one or more samples located on the flow cell is during the first plurality of sequencing cycles.

37. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles correspond only to the one or more index sequences.

38. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequence cycles corresponds only to the reads of the fragments of the DNA sequence for the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

39. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images include: generating base identification of the first plurality of flow cell images.

40. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further includes: determining a quality score for the base identification for each polymerase community of the plurality of samples in the first plurality of flow cell images.

41. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images include: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization; quality score estimation; or a combination thereof.

42. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.

43. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); or a combination thereof.

44. The computer-implemented method according to any one of the preceding claims, further comprising: generating, by the processor, the one or more index sequences for the set of sequencing reads, wherein the one or more index sequences are unique identifiers of the set of sequencing reads within a predetermined error tolerance rate.

45. The computer-implemented method according to any one of the preceding claims, further comprising: storing, by the processor, the one or more index sequences determined for the set of sequencing reads; and retrieving, by the processor, the one or more index sequences.

46. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The mismatch rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%.

47. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The unassigned rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%.

48. The computer-implemented method according to any one of the preceding claims, further comprising: generating one or more library molecules using the one or more index sequences.

49. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only a first index sequence.

50. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise a first index sequence and a second index sequence.

51. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 1 to 100 nucleotide bases.

52. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 3 to 16 nucleotide bases.

53. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise a plurality of nucleotide bases.

54. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the FastQ format.

55. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises a text-based format.

56. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the ASCII format.

57. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises an 8-bit encoding format.

58. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the positions in the order at which turns are placed such that only the nucleotide bases after the turns are reversed and complementary.

59. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the order of reading: the reads of the fragment of the DNA sequence; the reads of the complementary strand of the fragment of the DNA sequence; the first index sequence; and the second index sequence.

60. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the order of reading: the reads of the fragment of the DNA sequence; the reads of the complementary strand of the fragment of the DNA sequence; and the first index sequence.

61. The computer-implemented method according to any one of the preceding claims, further comprising: providing a sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.

62. The computer-implemented method according to any one of the preceding claims, further comprising: Obtain the first plurality of flow cell images of the sample immobilized on the support.

63. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the first plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z-levels along the axial axis.

64. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support.

65. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase colony or cluster immobilized thereon.

66. The computer-implemented method according to any one of the preceding claims, wherein the polymerase colony or cluster corresponds to the plurality of nucleotide template molecules or tandem molecules.

67. The computer-implemented method according to any one of the preceding claims, further comprising: Generating the first plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the plurality of tandem molecules immobilized on the support.

68. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U.

69. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity bodies.

70. The computer-implemented method according to any one of the preceding claims, wherein the individual affinity body in the mixture comprises a nucleus attached with a plurality of nucleotide arms, and each arm of the individual affinity body comprises the same type of nucleobase.

71. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, imaging an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules by an optical system.

72. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, acquiring the first plurality of flow cell images by an optical system, the first plurality of flow cell images comprising an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules.

73. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents that incorporate into the nucleotide bases A, G, C, and T / U of the plurality of templates or concatemer molecules immobilized on the support with unbalanced diversity in one or more cycles.

74. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase colonies comprises unbalanced diversity of nucleotide bases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleotide bases in the region of the first plurality of flow cell images to (2) the total number of nucleotide bases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region.

75. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule.

76. The computer-implemented method according to any one of the preceding claims, further comprising: Generating a plurality of cDNA molecules within the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule.

77. The computer-implemented method according to any one of the preceding claims, further comprising: Contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.

78. The computer-implemented method according to any one of the preceding claims, further comprising: Closing the nicks or gaps in at least the first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample.

79. The computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules to perform a rolling circle amplification reaction within the sample, thereby generating a plurality of concatemer molecules, the plurality of concatemer molecules comprising at least a first concatemer molecule corresponding to the first target RNA molecule, and the plurality of concatemer molecules comprising at least a second concatemer molecule corresponding to the second target RNA molecule.

80. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the plurality of concatemer molecules within the sample, which includes: sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.

81. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of tandem molecules within the sample comprises: Contacting the plurality of concatemer molecules within the sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers.

82. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent comprises one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.

83. The computer-implemented method according to any one of the preceding claims, further comprising: Removing the plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing the plurality of second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the sample.

84. A computer-implemented system comprising: One or more hardware processors; One or more data storage devices storing instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapping index sequences and mapping them to a plurality of samples; In response to determining that first plurality of flow cell images of the plurality of samples in a first plurality of sequencing cycles corresponding to the plurality of index sequences have been acquired during the ongoing sequencing run, generating, by the sequencing system, sequencing data that includes a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase community: In response to determining that each of the one or more index sequences matches one of the one or more mapping index sequences and maps to one of the plurality of samples, determining a value of a mismatch parameter for the polymerase community; and Determining that the value of the mismatch parameter satisfies a second error tolerance rate and assigning the polymerase community to the one mapped sample; and Evaluating, by the processor, one or more statistical parameters of the sequencing data while acquiring second plurality of flow cell images of the plurality of samples in a second plurality of sequencing cycles, wherein the second plurality of sequencing cycles or the second plurality of flow cell images correspond to at least a portion of the reads of a fragment of the DNA sequence of the set of sequencing reads for each polymerase community or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence; and Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion.

85. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices that store instructions that are executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the preceding claims.

86. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapping index sequences and mapping them to a plurality of samples; In response to determining that first flow cell images of the plurality of samples in a first plurality of sequencing cycles corresponding to the plurality of index sequences in a sequencing run have been acquired, generating, while the sequencing run is in progress, sequencing data by a sequencing system, the sequencing data including a set of sequencing reads for each polymerase colony of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase colony: Determining a value of a mismatch parameter for the polymerase colony in response to determining that each of the one or more index sequences matches one of the one or more mapping index sequences and maps to one of the plurality of samples; And Determining that the value of the mismatch parameter satisfies a second error tolerance rate and assigning the polymerase colony to the one mapped sample; And Evaluating, by the processor, one or more statistical parameters of the sequencing data while acquiring second flow cell images of the plurality of samples in a second plurality of sequencing cycles, wherein the second plurality of sequencing cycles or the second plurality of flow cell images correspond to at least a portion of the reads of a fragment of the DNA sequence or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence for the set of sequencing reads for each polymerase colony; And Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion.

87. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including any one of the preceding claims.

88. A computer-implemented system comprising: One or more hardware processors; One or more data storage devices that store instructions that are executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapping index sequences and mapping them to a plurality of samples; Generating, while the sequencing run of the plurality of samples is in progress, sequencing data by a sequencing system, the sequencing data including a set of sequencing reads for each polymerase colony of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase colony: Determining a value of a mismatch parameter of the polymerase community in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples; and Determining that the value of the mismatch parameter satisfies a second tolerance rate and assigning the polymerase community to the one mapped sample; and While the sequencing run of the plurality of samples is in progress, evaluating, by the processor, one or more statistical parameters of the sequencing data; and Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion thereof.

89. A computer-implemented system, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the foregoing claims.

90. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; While the sequencing run of the plurality of samples is in progress, generating, by a sequencing system, sequencing data comprising a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads of each polymerase community: Determining a value of a mismatch parameter of the polymerase community in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples; And Determining that the value of the mismatch parameter satisfies a second tolerance rate and assigning the polymerase community to the one mapped sample; And While the sequencing run of the plurality of samples is in progress, evaluating, by the processor, one or more statistical parameters of the sequencing data; And Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion thereof.

91. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations in sequencing data analysis including any one of the foregoing claims.

92. A computer-implemented system, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; While the sequencing runs of the plurality of samples are in progress, sequencing data is generated by a sequencing system, the sequencing data including a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase community: In response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, determining a value of a mismatch parameter for the polymerase community; and Determining that the value of the mismatch parameter satisfies a second tolerance rate, assigning the polymerase community to the one mapped sample; and Evaluating, by the processor, one or more statistical parameters of the sequencing data while generating at least a portion of the reads of the fragment of the DNA sequence or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence for the set of sequencing reads for each polymerase community of the plurality of samples; and Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion of the sequence run.

93. A computer-implemented system, comprising: One or more hardware processors; One or more data storage devices that store instructions that are executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the preceding claims.

94. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations including: Generating, by the processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; While the sequencing runs of the plurality of samples are in progress, generating sequencing data by a sequencing system, the sequencing data including a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads including one or more of the plurality of index sequences; For the set of sequencing reads for each polymerase community: In response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, determining a value of a mismatch parameter for the polymerase community; And Determining that the value of the mismatch parameter satisfies a second tolerance rate, assigning the polymerase community to the one mapped sample; And Evaluating, by the processor, one or more statistical parameters of the sequencing data while generating at least a portion of the reads of the fragment of the DNA sequence or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence for the set of sequencing reads for each polymerase community of the plurality of samples; And Determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion of the sequence run.

95. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations in sequencing data analysis, the operations including any of the foregoing claims.

96. A computer-implemented method for evaluating index sequences of different samples, comprising: generating, by a processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; generating, by a sequencing system while the sequencing run of the plurality of samples is in progress, sequencing data comprising a set of sequencing reads for each polymerase colony of the plurality of samples, the set of sequencing reads comprising one or more of the plurality of index sequences; for the set of sequencing reads for each polymerase colony: determining, in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples, a value of a mismatch parameter for the polymerase colony; and assigning, in response to determining that the value of the mismatch parameter satisfies a second error tolerance rate, the polymerase colony to the one mapped sample; and evaluating, by the processor, one or more statistical parameters of the sequencing data while the sequencing run of the plurality of samples is in progress; and determining, by the processor and based on the evaluation, whether to terminate the sequence run before completion.

97. The computer-implemented method according to any of the foregoing claims, further comprising: generating, by the processor, one or more reference index sequences for each of the plurality of samples, and wherein generating the one or more mapped index sequences and mapping them to the plurality of samples is based on the one or more reference index sequences for each of the plurality of samples.

98. The computer-implemented method according to any of the foregoing claims, further comprising: determining, by the processor, a read order for reading the set of sequencing reads for each polymerase colony; determining, by the processor and based on the read order, whether the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence of the set of reads for each polymerase colony.

99. The computer-implemented method according to any of the foregoing claims, further comprising: determining, by the processor and based on the read order, that the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence of the set of reads for each polymerase colony; and determining that the first plurality of flow cell images corresponding to the first plurality of sequencing cycles of the plurality of index sequences in the sequencing run have been acquired.

100. The computer-implemented method according to any one of the preceding claims, wherein the mapping of each of the one or more mapping index sequences is to only one of the plurality of samples.

101. The computer-implemented method according to any one of the preceding claims, wherein generating the one or more mapping index sequences and mapping them to a plurality of samples is based on a hash table that applies a hash function to each of the one or more mapping index sequences.

102. The computer-implemented method according to any one of the preceding claims, wherein determining that the value of the mismatch parameter satisfies the second error tolerance rate includes: Determining one or more positions in each of the one or more index sequences, each of the one or more positions corresponding to a mismatched base, a quality score for base identification below a pre-determined quality threshold, or both, wherein the value of the mismatch parameter for the polymerase population is the total number of the one or more positions.

103. The computer-implemented method according to any one of the preceding claims, further comprising: By the processor, separating each polymerase population based on the corresponding mapped sample to generate sequencing results in a plurality of data files in a pre-determined data format.

104. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.

105. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles or the second plurality of flow cell images do not correspond to the one or more index sequences.

106. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion includes: In response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion, automatically terminating the sequencing run before completion.

107. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion includes: In response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion, automatically terminating the generation of the sequencing data before the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both in the set of reads for each polymerase population have been completed.

108. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion includes: In response to determining that one or more values of the one or more statistical parameters do not satisfy a stop criterion, continuing the sequencing run until it is completed.

109. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion includes: In response to determining that one or more values of the one or more statistical parameters do not meet the stopping criteria, while the sequencing run is still in progress, generating the sequencing data, the sequencing data including the reads of the fragments of the DNA sequence of the set of sequencing reads for each polymerase community, the reads of the complementary strand of the fragments of the DNA sequence, or both.

110. The computer-implemented method according to any one of the preceding claims, further comprising: Upon completion of the sequencing run, separating the reads of the fragments of the DNA sequence of the set of sequencing reads for each polymerase community, the reads of the complementary strand of the fragments of the DNA sequence, or both, without separating one or more of the plurality of index sequences in the set of sequencing reads.

111. The computer-implemented method according to any one of the preceding claims, wherein evaluating the one or more statistical parameters of the sequencing data comprises: Calculating at least one value of each of the one or more statistical parameters for a plurality of polymerase communities based on the value of the mismatch parameter for each polymerase community; Comparing the at least one value of each of the one or more statistical parameters with a predetermined threshold; And Determining whether the at least one value of each of the one or more statistical parameters meets the predetermined threshold.

112. The computer-implemented method according to any one of the preceding claims, wherein the one or more statistical parameters of the sequencing result comprise one or more of the following: Mismatch rate; Unassigned rate; Error correlation rate; Assignment rate; Match rate; Deletion rate; and Hybrid pairing rate.

113. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and different.

114. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and the same.

115. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is 2 mismatched bases out of 9 bases.

116. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 1 deleted base out of 9 bases.

117. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 5%, 10%, 15%, 20%, 21%, 22%, 23%, 24%, 25%, or 30%.

118. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences are separated from the reads of the fragments of the DNA sequence of the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

119. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is not part of the read of the fragment of the DNA sequence or the read of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

120. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises a consecutive portion of the read of the fragment of the DNA sequence or the read of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

121. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is part of the read of the fragment of the DNA sequence or the read of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

122. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is a consecutive portion of the read of the fragment of the DNA sequence or the read of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

123. The computer-implemented method according to any one of the preceding claims, wherein the plurality of samples are located on one or more subtiles or one or more tiles of the flow cell.

124. The computer-implemented method according to any one of the preceding claims, wherein the read of the fragment of the DNA sequence is a forward read, a forward complementary read, or a reverse complementary read.

125. The computer-implemented method according to any one of the preceding claims, wherein the read of the complementary strand of the fragment of the DNA sequence is a forward or reverse read.

126. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data comprises: generating the sequencing data by processing based on one or more preliminary analysis steps of the first plurality of flow cell images.

127. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the read of the fragment of the DNA sequence or the read of the complementary strand of the fragment of the DNA sequence in each set of sequencing reads for each polymerase community.

128. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further comprises: acquiring, by the optical system of the sequencing system, the first plurality of flow cell images of one or more samples located on the flow cell; performing, by the processor, one or more preliminary analysis steps on the first plurality of flow cell images; or both.

129. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of one or more samples located on the flow cell is during the first plurality of sequencing cycles.

130. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles corresponds only to the one or more index sequences.

131. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequence cycles corresponds only to the reads of the fragments of the DNA sequence of the set of reads for each polymerase community or the reads of the complementary strand of the fragments of the DNA sequence.

132. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images include: Generating base calling for the first plurality of flow cell images.

133. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further includes: Determining a quality score for the base calling for each polymerase community of the plurality of samples in the first plurality of flow cell images.

134. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images include: Background subtraction; Image sharpening; Intensity offset adjustment; Color correction; Intensity normalization; Phasing and pre-phasing correction; Image registration; Intensity normalization; Quality score estimation; or A combination thereof.

135. The computer-implemented method according to any one of the preceding claims, wherein the processor includes: One or more processing units; One or more integrated circuits; Or A combination thereof.

136. The computer-implemented method according to any one of the preceding claims, wherein the processor includes: One or more central processing units (CPUs); One or more field programmable gate arrays (FPGAs); Or a combination thereof.

137. The computer-implemented method according to any one of the preceding claims, further comprising: Generating, by the processor, the one or more index sequences for each set of sequencing reads, wherein the one or more index sequences are unique identifiers of each set of sequencing reads within a pre-determined error tolerance rate.

138. The computer-implemented method according to any one of the preceding claims, further comprising: Storing, by the processor, the one or more index sequences determined for each set of sequencing reads; And Retrieving, by the processor, the one or more index sequences.

139. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The mismatch rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%.

140. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The unassigned rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

141. The computer-implemented method according to any one of the preceding claims, further comprising: generating one or more library molecules using the one or more index sequences.

142. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only a first index sequence.

143. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise a first index sequence and a second index sequence.

144. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 1 to 100 nucleotide bases.

145. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 3 to 16 nucleotide bases.

146. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise a plurality of nucleotide bases.

147. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the FastQ format.

148. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises a text-based format.

149. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the ASCII format.

150. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises an 8-bit encoding format.

151. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the positions in the order where turnarounds are placed such that only the nucleotide bases after the turnarounds are reversed and complementary.

152. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the reading order: the reads of the fragment of the DNA sequence; the reads of the complementary strand of the fragment of the DNA sequence; the first index sequence; and the second index sequence.

153. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the reading order: reads of the fragment of the DNA sequence reads of the complementary strand of the fragment of the DNA sequence and the first index sequence 154. The computer-implemented method according to any one of the preceding claims, further comprising: providing a sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample 155. The computer-implemented method according to any one of the preceding claims, further comprising: obtaining the first plurality of flow cell images of the sample immobilized on the support 156. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z-levels along an axial axis 157. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system through one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support 158. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community or cluster immobilized thereon 159. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community or cluster corresponds to the plurality of nucleotide template molecules or tandem molecules 160. The computer-implemented method according to any one of the preceding claims, further comprising: generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support 161. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleotide bases A, G, C, and T / U 162. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of adaptors 163. The computer-implemented method according to any one of the preceding claims, wherein each individual adaptor in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the individual adaptor comprises the same type of nucleotide base 164. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules is imaged by an optical system.

165. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, the optical system acquires the first plurality of flow cell images, the first plurality of flow cell images comprising an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules.

166. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents that bind to nucleobases A, G, C, and T / U of the plurality of templates or tandem molecules immobilized on the support with an unbalanced diversity in one or more cycles.

167. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase communities comprises an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases in the region of the first plurality of flow cell images to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region.

168. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule.

169. The computer-implemented method according to any one of the preceding claims, further comprising: Generating a plurality of cDNA molecules within the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule.

170. The computer-implemented method according to any one of the preceding claims, further comprising: Contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.

171. The computer-implemented method according to any one of the preceding claims, further comprising: Closing the nicks or gaps in the at least first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample.

172. The computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules, a rolling circle amplification reaction is carried out inside the sample to generate a plurality of concatemer molecules, the plurality of concatemer molecules comprising at least a first concatemer molecule corresponding to a first target RNA molecule, and the plurality of concatemer molecules comprising at least a second concatemer molecule corresponding to a second target RNA molecule.

173. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the plurality of concatemer molecules inside the sample, comprising: sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.

174. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of concatenated molecules inside the sample comprises: Contacting the plurality of concatemer molecules inside the sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemer.

175. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent comprises one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.

176. The computer-implemented method according to any one of the preceding claims, further comprising: Removing the plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing the plurality of second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the sample.

177. A computer-implemented method for evaluating index sequences of different samples, comprising: Generating, by a processor and based on a first error tolerance rate, one or more mapped index sequences and mapping them to a plurality of samples; Generating sequencing data by a sequencing system while the sequencing run of the plurality of samples is in progress, the sequencing data comprising a set of sequencing reads for each polymerase community of the plurality of samples, the set of sequencing reads comprising one or more of the plurality of index sequences; For the set of sequencing reads of each polymerase community: Determining a value of a mismatch parameter of the polymerase community in response to determining that each of the one or more index sequences matches one of the one or more mapped index sequences and maps to one of the plurality of samples; And Assigning the polymerase community to the one mapped sample in response to determining that the value of the mismatch parameter meets a second error tolerance rate; And Evaluating, by the processor, one or more statistical parameters of the sequencing data while generating at least a portion of the reads of the fragment of the DNA sequence or at least a portion of the reads of the complementary strand of the fragment of the DNA sequence for the set of sequencing reads of each polymerase community of the plurality of samples; And Determine, by the processor and based on the evaluation, whether to terminate it before completing the sequence run.

178. The computer-implemented method according to any one of the preceding claims, further comprising: generating, by the processor, one or more reference index sequences for each of the plurality of samples, and wherein generating the one or more mapped index sequences and mapping them to the plurality of samples is based on the one or more reference index sequences for each of the plurality of samples.

179. The computer-implemented method according to any one of the preceding claims, further comprising: determining, by the processor, a read order for reading the set of sequencing reads for each polymerase community; determining, by the processor and based on the read order, whether the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence of the set of reads for each polymerase community.

180. The computer-implemented method according to any one of the preceding claims, further comprising: determining, by the processor and based on the read order, that the one or more index sequences are located before the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence of the set of reads for each polymerase community; and determining that the first plurality of flow cell images in the first plurality of sequencing cycles corresponding to the plurality of index sequences in the sequencing run have been acquired.

181. The computer-implemented method according to any one of the preceding claims, wherein the mapping of each of the one or more mapped index sequences is to only one of the plurality of samples.

182. The computer-implemented method according to any one of the preceding claims, wherein generating the one or more mapped index sequences and mapping them to the plurality of samples is based on a hash table that applies a hash function to each of the one or more mapped index sequences.

183. The computer-implemented method according to any one of the preceding claims, wherein determining that the value of the mismatch parameter satisfies the second error tolerance rate includes: determining one or more positions in each of the one or more index sequences, each of the one or more positions corresponding to a mismatched base, a quality score for base identification below a pre-determined quality threshold, or both, wherein the value of the mismatch parameter for the polymerase community is the total number of the one or more positions.

184. The computer-implemented method according to any one of the preceding claims, further comprising: separating, by the processor and based on the corresponding mapped sample, each polymerase community to generate sequencing results in a plurality of data files in a pre-determined data format.

185. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.

186. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequencing cycles or the second plurality of flow cell images do not correspond to the one or more index sequences.

187. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion thereof comprises: automatically terminating the sequencing run before completion thereof in response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion.

188. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion thereof comprises: automatically terminating generation of the sequencing data before the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both, in the set of reads for each polymerase colony have been completed, in response to determining that one or more values of the one or more statistical parameters satisfy a stop criterion.

189. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion thereof comprises: continuing the sequencing run until completion thereof in response to determining that one or more values of the one or more statistical parameters do not satisfy a stop criterion.

190. The computer-implemented method according to any one of the preceding claims, wherein determining whether to terminate the sequence run before completion thereof comprises: generating the sequencing data, which comprises the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both, in the set of sequencing reads for each polymerase colony, while the sequencing run is still in progress, in response to determining that one or more values of the one or more statistical parameters do not satisfy a stop criterion.

191. The computer-implemented method according to any one of the preceding claims, further comprising: at the completion of the sequencing run, separating the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both, in the set of sequencing reads for each polymerase colony, without separating one or more of the plurality of index sequences in the set of sequencing reads.

192. The computer-implemented method according to any one of the preceding claims, wherein evaluating the one or more statistical parameters of the sequencing data comprises: calculating at least one value of each of the one or more statistical parameters for the plurality of polymerase colonies based on the value of the mismatch parameter for each polymerase colony; comparing the at least one value of each of the one or more statistical parameters with a predetermined threshold; and determining whether the at least one value of each of the one or more statistical parameters satisfies the predetermined threshold.

193. The computer-implemented method according to any one of the preceding claims, wherein the one or more statistical parameters of the sequencing result include one or more of the following: Mismatch rate; Unassigned rate; Error correlation rate; Assignment rate; Match rate; Deletion rate; and Hybrid pairing rate.

194. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and different.

195. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is predetermined and the same.

196. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is 2 mismatched bases out of 9 bases.

197. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 1 deleted base out of 9 bases.

198. The computer-implemented method according to any one of the preceding claims, wherein the first or second error tolerance rate is approximately 5%, 10%, 15%, 20%, 21%, 22%, 23%, 24%, 25% or 30%.

199. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences are separated from the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

200. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is not a part of the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

201. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises a consecutive portion of the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

202. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is a part of the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

203. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences is a consecutive portion of the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence in the set of reads for each polymerase community.

204. The computer-implemented method according to any one of the preceding claims, wherein the plurality of samples are located on one or more sub-tiles or one or more tiles of a flow cell.

205. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragments of the DNA sequence are forward reads, forward complementary reads, or reverse complementary reads.

206. The computer-implemented method according to any one of the preceding claims, wherein the reads of the complementary strand of the fragments of the DNA sequence are forward or reverse reads.

207. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data comprises: generating the sequencing data by processing, based on one or more preliminary analysis steps of the first plurality of flow cell images.

208. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the reads of the fragments of the DNA sequence or the reads of the complementary strand of the fragments of the DNA sequence in each set of sequencing reads of each polymerase colony.

209. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further comprises: acquiring, by an optical system of the sequencing system, the first plurality of flow cell images of one or more samples located on a flow cell; performing, by the processor, one or more preliminary analysis steps on the first plurality of flow cell images; or both.

210. The computer-implemented method according to any one of the preceding claims, wherein acquiring the first plurality of flow cell images of one or more samples located on a flow cell is within the first plurality of sequencing cycles.

211. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles correspond only to the one or more index sequences.

212. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of sequence cycles correspond only to the reads of the fragments of the DNA sequence or the reads of the complementary strand of the fragments of the DNA sequence for the set of reads for each polymerase colony.

213. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: generating base calling for the first plurality of flow cell images.

214. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data further comprises: determining a quality score for the base calling for each polymerase colony of the plurality of samples in the first plurality of flow cell images.

215. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; Image registration; Intensity normalization; Quality score estimation; or A combination thereof.

216. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: One or more processing units; One or more integrated circuits; Or A combination thereof.

217. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: One or more central processing units (CPUs); One or more field-programmable gate arrays (FPGAs); Or a combination thereof.

218. The computer-implemented method according to any one of the preceding claims, further comprising: Generating, by the processor, the one or more index sequences for each set of sequencing reads, wherein the one or more index sequences are unique identifiers of each set of sequencing reads within a predetermined error tolerance rate.

219. The computer-implemented method according to any one of the preceding claims, further comprising: Storing, by the processor, the one or more index sequences determined for each set of sequencing reads; And Retrieving, by the processor, the one or more index sequences. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The mismatch rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%.

221. The computer-implemented method according to any one of the preceding claims, wherein the stopping criterion comprises: The unassigned rate is about 25%, 24%, 23%, 22%, 21%, 20%, 18%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%.

222. The computer-implemented method according to any one of the preceding claims, further comprising: Generating one or more library molecules using the one or more index sequences.

223. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only a first index sequence.

224. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise a first index sequence and a second index sequence.

225. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 1 to 100 nucleotide bases.

226. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 3 to 16 nucleotide bases.

227. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise a plurality of nucleotide bases.

228. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format includes the FastQ format.

229. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises a text-based format.

230. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the ASCII format.

231. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises an 8-bit encoding format.

232. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining positions in the order at which to place turns such that only nucleotide bases after the turn are reversed and complementary.

233. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the order of reading: the read of the fragment of the DNA sequence; the read of the complementary strand of the fragment of the DNA sequence; the first index sequence; and the second index sequence.

234. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragment of the DNA sequence and the one or more index sequences comprises: determining the order of reading: the read of the fragment of the DNA sequence; the read of the complementary strand of the fragment of the DNA sequence; and the first index sequence.

235. The computer-implemented method according to any one of the preceding claims, further comprising: providing a sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.

236. The computer-implemented method according to any one of the preceding claims, further comprising: obtaining the first plurality of flow cell images of the sample immobilized on the support.

237. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels along an axial axis at two or more different z levels.

238. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system through one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support.

239. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community or cluster immobilized thereon.

240. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community or cluster corresponds to the plurality of nucleotide template molecules or concatemer molecules.

241. The computer-implemented method according to any one of the preceding claims, further comprising: generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on the plurality of concatemer molecules immobilized on the support.

242. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of concatemer molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U.

243. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of concatemer molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of adaptors.

244. The computer-implemented method according to any one of the preceding claims, wherein an individual adaptor in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the individual adaptor comprises the same type of nucleobase.

245. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: in each of the one or more cycles, imaging an optical color signal emitted from a nucleotide reagent bound to the plurality of concatemer molecules by an optical system.

246. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: in each of the one or more cycles, acquiring the first plurality of flow cell images by an optical system, the first plurality of flow cell images comprising optical color signals emitted from nucleotide reagents bound to the plurality of concatemer molecules.

247. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents with an unbalanced diversity of nucleobases A, G, C, and T / U that bind to the plurality of template or concatemer molecules immobilized on the support in one or more cycles.

248. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase communities comprises an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases in a region of the first plurality of flow cell images to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule.

250. The computer-implemented method according to any one of the preceding claims, further comprising: generating a plurality of cDNA molecules within the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule.

251. The computer-implemented method according to any one of the preceding claims, further comprising: contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.

252. The computer-implemented method according to any one of the preceding claims, further comprising: generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample by closing the nicks or gaps in the at least first and second circularized target-specific padlock probes by performing an enzymatic reaction.

253. The computer-implemented method according to any one of the preceding claims, further comprising: performing a rolling circle amplification reaction within the sample using the first and second covalently closed circular padlock probes as template molecules, thereby generating a plurality of tandem molecules, the plurality of tandem molecules comprising at least a first tandem molecule corresponding to the first target RNA molecule, and the plurality of tandem molecules comprising at least a second tandem molecule corresponding to the second target RNA molecule.

254. The computer-implemented method according to any one of the preceding claims, further comprising: sequencing the plurality of tandem molecules within the sample, comprising: sequencing the first tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of tandem molecules within the sample comprises: contacting the plurality of tandem molecules within the sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the tandem.

256. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagents comprise one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.

257. The computer-implemented method according to any one of the preceding claims, further comprising: removing the plurality of first sequencing read products from the first tandem molecule and retaining the first tandem molecule in the sample, and removing the plurality of second sequencing read products from the second tandem molecule and retaining the second tandem molecule in the sample.

258. A computer-implemented method for determining an index sequence in sequencing data analysis, comprising: Sequencing data is generated from one or more sequencing runs by a sequencing system, and the sequencing data for each of the one or more sequencing runs includes multiple sets of sequencing reads, each set of sequencing reads including reads of fragments of a DNA sequence and one or more index sequences; A processor determines an order of reading the reads of the fragments of the DNA sequence and the one or more index sequences; The processor determines, based on the reading order, whether the one or more index sequences are to be reverse complemented; and The processor generates a predicted sequence result in a predetermined data format using the one or more index sequences based on the determination of the one or more index sequences.

259. The computer-implemented method according to any one of the preceding claims, wherein each set of the sequencing reads further includes reads of the complementary strand of the fragment of the DNA sequence.

260. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragments of the DNA sequence are forward reads or reverse-complemented reads.

261. The computer-implemented method according to any one of the preceding claims, wherein each set of the sequencing reads further includes forward or reverse reads of the complementary strand of the fragment of the DNA sequence.

262. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs includes: generating the sequencing data through processing based on one or more preliminary analysis steps of multiple flow cell images.

263. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs further includes: an optical system of the sequencing system acquiring the multiple flow cell images of one or more samples positioned on a flow cell; the processor performing the one or more preliminary analysis steps on the multiple flow cell images; or both.

264. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the multiple flow cell images include: generating base calling using the multiple flow cell images.

265. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the multiple flow cell images include: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization; quality score estimation; or a combination thereof.

266. The computer-implemented method according to any one of the preceding claims, wherein the processor includes: one or more processing units; one or more integrated circuits; or a combination thereof.

267. The computer-implemented method according to any one of the preceding claims, wherein the processor includes: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); or a combination thereof.

268. The computer-implemented method according to any one of the preceding claims, wherein the sequencing system comprises: The optical system and the processor.

269. The computer-implemented method according to any one of the preceding claims, further comprising: generating, by the processor, the one or more index sequences for each set of sequencing reads, wherein the one or more index sequences are unique identifiers for each set of sequencing reads within a pre-determined error tolerance rate.

270. The computer-implemented method according to any one of the preceding claims, further comprising: storing, by the processor, the one or more index sequences determined for each set of sequencing reads; and retrieving, by the processor, the one or more index sequences.

271. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined error tolerance rate is about 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or 5%.

272. The computer-implemented method according to any one of the preceding claims, further comprising: preparing one or more samples using the generated one or more index sequences.

273. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing result in the pre-determined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: separating, by the processor, the multiple sets of sequencing reads based on the one or more index sequences and the determination thereof to generate the sequencing result in the pre-determined data format in multiple files.

274. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only a first index sequence.

275. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise a first index sequence and a second index sequence.

276. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 1 to 100 nucleotide bases.

277. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises from 6 to 16 nucleotide bases.

278. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise multiple nucleotide bases.

279. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined data format comprises the FastQ format.

280. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined data format comprises a text-based format.

281. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined data format comprises the ASCII format.

282. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined data format comprises an 8-bit encoding format.

283. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragment of the DNA sequence and the one or more index sequences comprises: Determining positions in the order where turns are placed such that only nucleotide bases after the turn are reversed and complementary.

284. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragment of the DNA sequence and the one or more index sequences comprises: Determining the read order: The reads of the fragment of the DNA sequence; The reads of the complementary strand of the fragment of the DNA sequence; The first index sequence; and The second index sequence.

285. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: Generating sequencing data from one or more sequencing runs by a sequencing system, the sequencing data for each of the one or more sequencing runs containing multiple sets of sequencing reads, each set of sequencing reads containing reads of a fragment of a DNA sequence and one or more index sequences; Determining, by a processor, the order of the reads of the fragment of the DNA sequence and the one or more index sequences; Determining, by the processor and based on the read order, whether the one or more index sequences are to be reverse-complemented; and Generating, by the processor, based on the determination of the one or more index sequences, a predicted sequence result in a pre-determined data format using the one or more index sequences.

286. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including any one of the preceding claims.

287. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including: Generating sequencing data from one or more sequencing runs by a sequencing system, the sequencing data for each of the one or more sequencing runs containing multiple sets of sequencing reads, each set of sequencing reads containing reads of a fragment of a DNA sequence and one or more index sequences; Determining, by a processor, the order of the reads of the fragment of the DNA sequence and the one or more index sequences; Determining, by the processor and based on the read order, whether the one or more index sequences are to be reverse-complemented; And Based on the determination of the one or more index sequences, the processor generates a predicted sequence result using the one or more index sequences in a predetermined data format. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including any one of the preceding claims.

289. A computer-implemented method for determining index sequences in sequencing data analysis, comprising: The sequencing system generates sequencing data from one or more sequencing runs, and the sequencing data for each of the one or more sequencing runs includes multiple sets of sequencing reads, and each set of sequencing reads includes one or more index sequences; The processor determines the order of reading each set of sequencing reads; Based on multiple factors, the processor determines whether the one or more index sequences are to be reverse complemented; And Based on the determination of the one or more index sequences, the processor generates a predicted sequence result using the one or more index sequences in a predetermined data format.

290. The computer-implemented method according to any one of the preceding claims, wherein generating a sequencing result using the one or more index sequences in the predetermined data format based on the determination of the one or more index sequences includes: Based on the one or more index sequences and the determination thereof, the processor separates the multiple sets of sequencing reads to generate the sequencing result in multiple files in the predetermined data format.

291. The computer-implemented method according to any one of the preceding claims, wherein generating a sequencing result using the one or more index sequences in the predetermined data format based on the determination of the one or more index sequences includes: The processor evaluates one or more statistical parameters of the sequencing result while the sequencing system generates the multiple sets of sequencing reads, and each set of sequencing reads further includes the reads of the fragment of the DNA sequence.

292. The computer-implemented method according to any one of the preceding claims, wherein generating a sequencing result using the one or more index sequences in the predetermined data format based on the determination of the one or more index sequences includes: The processor evaluates one or more statistical parameters of the sequencing result while the sequencing system generates the multiple sets of sequencing reads, and each set of sequencing reads further includes the reads of the fragment of the DNA sequence and the reads of the complementary strand of the fragment of the DNA sequence.

293. The computer-implemented method according to any one of the preceding claims, wherein generating a sequencing result using the one or more index sequences in the predetermined data format based on the determination of the one or more index sequences includes: The processor evaluates one or more statistical parameters of the sequencing result while obtaining a second plurality of flow cell images of one or more samples located on a flow cell in a second plurality of sequencing cycles, wherein the second plurality of sequencing cycles or the second plurality of flow cell images correspond to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragments of the DNA sequence.

294. The computer-implemented method according to any one of the preceding claims, wherein based on the determination of the one or more index sequences, generating the sequencing result in the predetermined data format using the one or more index sequences comprises: The processor evaluates one or more statistical parameters of the sequencing result while obtaining a second plurality of flow cell images of one or more samples located on a flow cell in a second plurality of sequencing cycles, wherein the second plurality of sequencing cycles or the second plurality of flow cell images do not correspond to the one or more index sequences.

295. The computer-implemented method according to any one of the preceding claims, wherein based on the determination of the one or more index sequences, generating the sequencing result in the predetermined data format using the one or more index sequences further comprises: In response to the evaluation, stop generating the plurality of sets of sequencing reads, each set of sequencing reads further comprising the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both.

296. The computer-implemented method according to any one of the preceding claims, wherein based on the determination of the one or more index sequences, generating the sequencing result in the predetermined data format using the one or more index sequences further comprises: In response to the evaluation, continue generating the plurality of sets of sequencing reads, each set of sequencing reads further comprising the reads of the fragments of the DNA sequence, the reads of the complementary strand of the fragments of the DNA sequence, or both.

297. The computer-implemented method according to any one of the preceding claims, wherein evaluating the one or more statistical parameters of the sequencing result comprises: Based on the one or more retrieved index sequences, calculating at least one value of each of the one or more statistical parameters; Comparing the at least one value of each of the one or more statistical parameters with a predetermined threshold; And Determining whether the at least one value of each of the one or more statistical parameters meets the predetermined threshold.

298. The computer-implemented method according to any one of the preceding claims, wherein the one or more statistical parameters of the sequencing result comprise: A mismatch parameter; An unassigned parameter; Or both.

299. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragments of the DNA sequence are forward reads, forward complementary reads, or reverse complementary reads.

300. The computer-implemented method according to any one of the preceding claims, wherein the reads of the complementary strand of the fragment of the DNA sequence are forward or reverse reads.

301. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs comprises: generating the sequencing data by processing based on one or more preliminary analysis steps of a first plurality of flow cell images.

302. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired during a first plurality of sequence cycles.

303. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles correspond to the one or more index sequences.

304. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the reads of the fragments of the DNA sequence in each sequence read or the reads of the complementary strand of the fragment of the DNA sequence.

305. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs further comprises: acquiring, by an optical system of the sequencing system, the first plurality of flow cell images of one or more samples located on a flow cell; performing, by the processor, one or more preliminary analysis steps on the first plurality of flow cell images; or both.

306. The computer-implemented method according to any one of the preceding claims, wherein acquiring the first plurality of flow cell images of one or more samples located on a flow cell is during the first plurality of sequencing cycles.

307. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles correspond to the one or more index sequences.

308. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles correspond only to the one or more index sequences.

309. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: generating base calling of the first plurality of flow cell images.

310. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization; quality score estimation; or a combination thereof.

311. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.

312. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); or a combination thereof. The computer-implemented method according to any one of the preceding claims, wherein the sequencing system comprises: the optical system and the processor.

314. The computer-implemented method according to any one of the preceding claims, further comprising: generating, by the processor, the one or more index sequences for each set of sequencing reads, wherein the one or more index sequences are unique identifiers of each set of sequencing reads within a predetermined error tolerance rate.

315. The computer-implemented method according to any one of the preceding claims, further comprising: storing, by the processor, the one or more index sequences determined for each set of sequencing reads; and retrieving, by the processor, the one or more index sequences.

316. The computer-implemented method according to any one of the preceding claims, wherein the predetermined error tolerance rate is about 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or 5%.

317. The computer-implemented method according to any one of the preceding claims, further comprising: preparing one or more samples using the generated one or more index sequences.

318. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing result in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: by the processor, separating the multiple sets of sequencing reads based on the one or more index sequences and the determination thereof to generate the sequencing result in the predetermined data format in multiple files.

319. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only a first index sequence.

320. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise a first index sequence and a second index sequence.

321. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 1 to 100 nucleotide bases.

322. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 6 to 16 nucleotide bases.

323. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise multiple nucleotide bases.

324. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the FastQ format.

325. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises a text-based format.

326. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the ASCII format.

327. The computer-implemented method according to any one of the preceding claims, wherein the pre-determined data format includes an 8-bit encoding format.

328. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragments of the DNA sequence and the one or more index sequences comprises: Determining positions in the order where turns are placed such that only nucleotide bases after the turns are reversed and complementary.

329. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragments of the DNA sequence and the one or more index sequences comprises: Determining the read order: The reads of the fragments of the DNA sequence; The reads of the complementary strand of the fragments of the DNA sequence; The first index sequence; and The second index sequence.

330. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragments of the DNA sequence and the one or more index sequences comprises: Determining the read order: The reads of the fragments of the DNA sequence; The reads of the complementary strand of the fragments of the DNA sequence; And The first index sequence.

331. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: Generating sequencing data from one or more sequencing runs by a sequencing system, the sequencing data for each of the one or more sequencing runs containing multiple sets of sequencing reads, each set of sequencing reads containing one or more index sequences; Determining by the processor the order of reading each set of sequencing reads; Determining by the processor and based on multiple factors whether the one or more index sequences are to be reverse-complemented; And Generating a predicted sequence result in a pre-determined data format by the processor, based on the determination of the one or more index sequences, using the one or more index sequences.

332. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including any one of the preceding claims.

333. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including: Sequencing data is generated from one or more sequencing runs by a sequencing system, and for each of the one or more sequencing runs, the sequencing data includes multiple sets of sequencing reads, each set of sequencing reads including one or more index sequences; A processor determines an order for reading each set of sequencing reads; The processor determines, based on multiple factors, whether the one or more index sequences are to be reverse complemented; and The processor generates a predicted sequencing result in a predetermined data format using the one or more index sequences based on the determination of the one or more index sequences.

334. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including any one of the foregoing claims.

335. A computer-implemented method for determining index sequences in sequencing data analysis, the method including: In response to determining that first multiple flow cell images corresponding to one or more index sequences have been acquired, sequencing data is generated from one or more sequencing runs by a sequencing system, and for each of the one or more sequencing runs, the sequencing data includes multiple sets of sequencing reads, each set of sequencing reads including the one or more index sequences; A processor determines an order for reading each set of sequencing reads; The processor determines, based on multiple factors, whether the one or more index sequences are to be reverse complemented; and The processor generates a sequencing result in a predetermined data format using the one or more index sequences based on the determination of the one or more index sequences, while concurrently acquiring second multiple flow cell images corresponding to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragments of the DNA sequence.

336. A computer-implemented method for determining index sequences in sequencing data analysis, the method including: In response to determining that first multiple flow cell images corresponding to a first sequencing cycle have been acquired, sequencing data is generated from one or more sequencing runs by a sequencing system, and for each of the one or more sequencing runs, the sequencing data includes multiple sets of sequencing reads, each set of sequencing reads including the one or more index sequences; A processor determines an order for reading each set of sequencing reads; The processor determines, based on multiple factors, whether the one or more index sequences are to be reverse complemented; and The processor generates a sequencing result in a predetermined data format using the one or more index sequences based on the determination of the one or more index sequences, while concurrently generating second multiple flow cell images corresponding to second multiple sequencing cycles after the first multiple sequencing cycles.

337. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequencing cycles corresponds to the one or more index sequences, and wherein the second plurality of sequencing cycles corresponds to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragment of the DNA sequence.

338. The computer-implemented method according to any one of the preceding claims, further comprising: determining that the first plurality of flow cell images corresponding to the one or more index sequences have been acquired; and determining that the second plurality of flow cell images not corresponding to the one or more index sequences have not been acquired.

339. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing results in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: separating, by the processor, the plurality of sets of sequencing reads based on the one or more index sequences and the determination thereof to generate the sequencing results in the predetermined data format in a plurality of files.

340. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing results in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: evaluating, by the processor, one or more statistical parameters of the sequencing results while the sequencing system generates the plurality of sets of sequencing reads, each set of sequencing reads further comprising reads of the fragment of the DNA sequence.

341. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing results in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: evaluating, by the processor, one or more statistical parameters of the sequencing results while the sequencing system generates the plurality of sets of sequencing reads, each set of sequencing reads further comprising reads of the fragment of the DNA sequence and reads of the complementary strand of the fragment of the DNA sequence.

342. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing results in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: evaluating, by the processor, one or more statistical parameters of the sequencing results while acquiring second plurality of flow cell images of one or more samples located on the flow cell in the second plurality of sequence cycles, wherein the second plurality of sequence cycles or the second plurality of flow cell images corresponds to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragment of the DNA sequence.

343. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing results in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: One or more statistical parameters of the sequencing result are evaluated by the processor while obtaining a second plurality of flow cell images of one or more samples located on the flow cell in a second plurality of sequence cycles, wherein the second plurality of sequence cycles or the second plurality of flow cell images do not correspond to the one or more index sequences.

344. The computer-implemented method according to any one of the preceding claims, wherein based on the determination of the one or more index sequences, generating the sequencing result in the predetermined data format using the one or more index sequences further comprises: In response to the evaluation, stopping generating the multiple sets of sequencing reads, each set of sequencing reads further comprising the reads of the fragments of the DNA sequence, the reads of the complementary strands of the fragments of the DNA sequence, or both.

345. The computer-implemented method according to any one of the preceding claims, wherein evaluating the one or more statistical parameters of the sequencing result comprises: Calculating at least one value of each of the one or more statistical parameters based on the one or more retrieved index sequences; Comparing the at least one value of each of the one or more statistical parameters with a predetermined threshold; And Determining whether the at least one value of each of the one or more statistical parameters meets the predetermined threshold.

346. The computer-implemented method according to any one of the preceding claims, wherein the one or more statistical parameters of the sequencing result comprise: A mismatch parameter; An unassigned parameter; Or both.

347. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragments of the DNA sequence are forward reads, forward complementary reads, or reverse complementary reads.

348. The computer-implemented method according to any one of the preceding claims, wherein the reads of the complementary strands of the fragments of the DNA sequence are forward or reverse reads.

349. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs comprises: Generating the sequencing data by processing based on one or more preliminary analysis steps of a first plurality of flow cell images.

350. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are obtained in a first plurality of sequence cycles.

351. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles correspond to the one or more index sequences.

352. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the first plurality of sequence cycles do not correspond to the reads of the fragments of the DNA sequence in each sequence read or the reads of the complementary strands of the fragments of the DNA sequence.

353. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing data from the one or more sequencing runs further comprises: obtaining, by an optical system of the sequencing system, the first plurality of flow cell images of one or more samples located on a flow cell; performing, by the processor, one or more preliminary analysis steps on the first plurality of flow cell images; or both.

354. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of one or more samples located on a flow cell is during a first plurality of sequencing cycles.

355. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles corresponds to the one or more index sequences.

356. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequence cycles corresponds only to the one or more index sequences.

357. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: generating base calling using the first plurality of flow cell images.

358. The computer-implemented method according to any one of the preceding claims, wherein the one or more preliminary analysis steps of the first plurality of flow cell images comprise: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; intensity normalization; quality score estimation; or a combination thereof.

359. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.

360. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); or a combination thereof.

361. The computer-implemented method according to any one of the preceding claims, wherein the sequencing system comprises: The optical system and the processor.

362. The computer-implemented method according to any one of the preceding claims, further comprising: generating, by the processor, the one or more index sequences for each set of sequencing reads, wherein the one or more index sequences are unique identifiers of each set of sequencing reads within a predetermined error tolerance rate.

363. The computer-implemented method according to any one of the preceding claims, further comprising: storing, by the processor, the one or more index sequences determined for each set of sequencing reads; and retrieving, by the processor, the one or more index sequences.

364. The computer-implemented method according to any one of the preceding claims, wherein the predetermined error tolerance rate is about 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or 5%.

365. The computer-implemented method according to any one of the preceding claims, further comprising: Prepare one or more samples using the generated one or more index sequences.

366. The computer-implemented method according to any one of the preceding claims, wherein generating the sequencing result in the predetermined data format using the one or more index sequences based on the determination of the one or more index sequences comprises: By the processor, based on the one or more index sequences and the determination thereof, separating the multiple sets of sequencing reads to generate the sequencing result in the predetermined data format in multiple files.

367. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise only the first index sequence.

368. The computer-implemented method according to any one of the preceding claims, wherein the one or more index sequences comprise the first index sequence and the second index sequence.

369. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 1 to 100 nucleotide bases.

370. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more index sequences comprises 6 to 16 nucleotide bases.

371. The computer-implemented method according to any one of the preceding claims, wherein the reads of the fragment of the DNA sequence or the reads of the complementary strand of the fragment of the DNA sequence comprise multiple nucleotide bases.

372. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the FastQ format.

373. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises a text-based format.

374. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises the ASCII format.

375. The computer-implemented method according to any one of the preceding claims, wherein the predetermined data format comprises an 8-bit encoding format.

376. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: Determining the positions of the turns in the order such that only the nucleotide bases after the turns are reversed and complementary.

377. The computer-implemented method according to any one of the preceding claims, wherein determining the order of reading the reads of the fragment of the DNA sequence and the one or more index sequences comprises: Determining the reading order: The reads of the fragment of the DNA sequence; The reads of the complementary strand of the fragment of the DNA sequence; The first index sequence; and The second index sequence.

378. The computer-implemented method according to any one of the preceding claims, wherein determining the order of the reads of the fragments of the DNA sequence and the one or more index sequences comprises: Determining the read order: The reads of the fragments of the DNA sequence; The reads of the complementary strands of the fragments of the DNA sequence; And The first index sequence.

379. The computer-implemented method according to any one of the preceding claims, further comprising: Providing a sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.

380. The computer-implemented method according to any one of the preceding claims, further comprising: Obtaining the first plurality of flow cell images of the sample immobilized on the support.

381. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels along an axial axis at two or more different z levels.

382. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the plurality of flow cell images by the sequencing system through one or more sequencing reaction cycles on the plurality of tandem molecules of the sample immobilized on the support.

383. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community or cluster immobilized thereon.

384. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community or cluster corresponds to the plurality of nucleotide template molecules or tandem molecules.

385. The computer-implemented method according to any one of the preceding claims, further comprising: Generating the first plurality of flow cell images by a sequencing system through one or more sequencing reaction cycles on the plurality of tandem molecules immobilized on the support.

386. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleotide bases A, G, C, and T / U.

387. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of adaptors.

388. The computer-implemented method according to any one of the preceding claims, wherein each individual adaptor in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the individual adaptor comprises the same type of nucleotide base.

389. A computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, imaging, by an optical system, an optical color signal emitted from nucleotide reagents bound to the plurality of tandem molecules.

390. A computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, acquiring, by an optical system, the first plurality of flow cell images, the first plurality of flow cell images comprising optical color signals emitted from nucleotide reagents bound to the plurality of tandem molecules.

391. A computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images comprise optical signals emitted from nucleotide reagents that are bound to nucleobases A, G, C, and T / U of the plurality of templates or tandem molecules immobilized on the support with an unbalanced diversity in one or more cycles.

392. A computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase colonies comprise an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases in a region of the first plurality of flow cell images to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule.

394. A computer-implemented method according to any one of the preceding claims, further comprising: Generating, within the sample, a plurality of cDNA molecules, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule.

395. A computer-implemented method according to any one of the preceding claims, further comprising: Contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.

396. A computer-implemented method according to any one of the preceding claims, further comprising: Closing a gap or space in at least the first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample.

397. A computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules, a rolling circle amplification reaction is performed inside the sample to generate a plurality of tandem molecules, the plurality of tandem molecules including at least a first tandem molecule corresponding to a first target RNA molecule, and the plurality of tandem molecules including at least a second tandem molecule corresponding to a second target RNA molecule.

398. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the plurality of tandem molecules inside the sample, comprising: sequencing the first tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.

399. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of tandem molecules inside the sample comprises: Contacting the plurality of tandem molecules inside the sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their corresponding universal sequencing primer binding sites on the tandem molecules.

400. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent comprises one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.

401. The computer-implemented method according to any one of the preceding claims, further comprising: Removing the plurality of first sequencing read products from the first tandem molecule and retaining the first tandem molecule in the sample, and removing the plurality of second sequencing read products from the second tandem molecule and retaining the second tandem molecule in the sample.

402. The computer-implemented method according to any one of the preceding claims, further comprising: Providing a sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.

403. The computer-implemented method according to any one of the preceding claims, further comprising: Obtaining the first plurality of flow cell images of the sample immobilized on the support.

404. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the first plurality of flow cell images by a sequencing system by performing one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z levels along an axial axis.

405. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample comprises: Generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the plurality of tandem molecules of the sample immobilized on the support.

406. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community or cluster immobilized thereon.

407. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community or cluster corresponds to the plurality of nucleotide template molecules or concatemer molecules.

408. The computer-implemented method according to any one of the preceding claims, further comprising: generating the first plurality of flow cell images by a sequencing system by performing one or more sequencing reaction cycles on the plurality of concatemer molecules immobilized on the support.

409. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of concatemer molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleotide bases A, G, C, and T / U.

410. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of concatemer molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of adaptors.

411. The computer-implemented method according to any one of the preceding claims, wherein a separate adaptor in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the separate adaptor comprises the same type of nucleotide base.

412. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: imaging, by an optical system, an optical color signal emitted from a nucleotide reagent bound to the plurality of concatemer molecules in each of the one or more cycles.

413. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: acquiring, by an optical system, the first plurality of flow cell images in each of the one or more cycles, the first plurality of flow cell images comprising optical color signals emitted from nucleotide reagents bound to the plurality of concatemer molecules.

414. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents with an unbalanced diversity of nucleotide bases A, G, C, and T / U that bind to the plurality of template or concatemer molecules immobilized on the support in one or more cycles.

415. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase communities comprises an unbalanced diversity of nucleotide bases A, G, C, and T / U, and wherein the unbalanced diversity comprises a percentage of (1) the number of one or more types of nucleotide bases in a region of the first plurality of flow cell images to (2) the total number of nucleotide bases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region. The computer-implemented method according to any one of the preceding claims, further comprising: Provide the sample containing multiple RNAs, the multiple RNAs including at least a first target RNA molecule and a second target RNA molecule.

417. The computer-implemented method according to any one of the preceding claims, further comprising: Generating multiple cDNA molecules inside the sample, the multiple cDNA molecules including at least a first target cDNA molecule corresponding to the first target RNA molecule, and the multiple cDNA molecules including a second target cDNA molecule corresponding to the second target RNA molecule.

418. The computer-implemented method according to any one of the preceding claims, further comprising: Contacting the multiple cDNA molecules in the sample with multiple target-specific padlock probes, the target-specific padlock probes including at least a first multiple of target-specific padlock probes and a second multiple of target-specific padlock probes.

419. The computer-implemented method according to any one of the preceding claims, further comprising: Closing the nicks or gaps in the at least first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe inside the sample.

420. The computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules to perform a rolling circle amplification reaction inside the sample, thereby generating multiple concatemer molecules, the multiple concatemer molecules including at least a first concatemer molecule corresponding to the first target RNA molecule, and the multiple concatemer molecules including at least a second concatemer molecule corresponding to the second target RNA molecule.

421. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the multiple concatemer molecules inside the sample, including: sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate multiple first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate multiple second sequencing read products. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of concatenated molecules within the sample comprises: Contacting the multiple concatemer molecules inside the sample with (i) multiple universal sequencing primers, (ii) multiple sequencing polymerases, and (iii) multiple nucleotide reagents under conditions suitable for hybridizing the multiple universal sequencing primers to their corresponding universal sequencing primer binding sites on the concatemer.

423. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent includes one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.

424. The computer-implemented method according to any one of the preceding claims, further comprising: Removing the multiple first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing the multiple second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the sample.

425. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including: In response to determining that a first plurality of flow cell images corresponding to one or more index sequences have been acquired, generating sequencing data from one or more sequencing runs by a sequencing system, the sequencing data for each of the one or more sequencing runs including multiple sets of sequencing reads, each set of sequencing reads including the one or more index sequences; Determining, by a processor, an order for reading each set of the sequencing reads; Determining, by the processor and based on multiple factors, whether the one or more index sequences are to be reverse complemented; And Generating, by the processor, based on the determination of the one or more index sequences, a sequencing result in a predetermined data format using the one or more index sequences, while concurrently acquiring a second plurality of flow cell images corresponding to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragment of the DNA sequence.

426. A computer-implemented system for determining index sequences in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the preceding claims.

427. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including: In response to determining that a first plurality of flow cell images corresponding to one or more index sequences have been acquired, generating sequencing data from one or more sequencing runs by a sequencing system, the sequencing data for each of the one or more sequencing runs including multiple sets of sequencing reads, each set of sequencing reads including the one or more index sequences; Determining, by a processor, an order for reading each set of the sequencing reads; Determining, by the processor and based on multiple factors, whether the one or more index sequences are to be reverse complemented; And Generating, by the processor, based on the determination of the one or more index sequences, a sequencing result in a predetermined data format using the one or more index sequences, while concurrently acquiring a second plurality of flow cell images corresponding to reads of fragments of the DNA sequence in each sequence read or reads of the complementary strand of the fragment of the DNA sequence.

428. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for determining index sequences in sequencing data analysis, the operations including any one of the preceding claims.

Citation Information

Patent Citations

  • Method and system for sequencing nucleic acids

    US10246744B2

  • Engineered polymerases for improved sequencing

    US10731141B2

  • Improvement in drawers

    US133138A

  • Primary analysis in next generation sequencing

    US20230326064A1

  • Primary analysis in next generation sequencing

    US20230326065A1