On-flow cell library preparation for single-cell RNA analysis

The two-step process of encapsulating and tagging single cells for on-board linked-read library preparation and sequencing addresses the limitations of short-read technologies in de novo transcriptome assembly, achieving efficient long-range contiguity and gene expression analysis in single cells.

WO2026102281A1PCT designated stage Publication Date: 2026-05-15ILLUMINA INC
View PDF 55 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-11-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

De novo transcriptome assembly using short-read technologies is challenging due to limited ability to determine long-range information, and existing methods for single-cell RNA analysis often require additional user intervention and are not efficient in providing both long-range contiguity and gene expression information.

Method used

A two-step process involving encapsulation of single cells, tagging of molecular content, and on-board linked-read library preparation and sequencing, where unique tags are incorporated during reverse transcription and library preparation to assemble long-range contiguity information, enabling simultaneous determination of gene expression without additional user work.

Benefits of technology

Enables efficient assembly of long-range contiguity information and gene expression counting in single cells using a single workflow, without requiring additional bioinformatic processing, and supports various analytes by leveraging spatial proximity for tag assembly and expression counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025054600_15052026_PF_FP_ABST
    Figure US2025054600_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for characterizing analytes present in a sample of individual cells. The initial sample processing steps take place on a per-cell basis within an individual partition (e.g., droplet) to tag or barcode each cell with a cell-specific barcode (e.g. using capture probe with cell barcode bound to bead). Subsequent sample processing steps take place on a flow cell surface with transposase to generate location-specific information that serves to link individual sequences reads with one another based on intrinsic molecular identifiers, adjacent fragmenting and cluster formation. The disclosed workflow generates expression-level data (based on read count) as well as linked read data to facilitate genome or full transcript assembly.
Need to check novelty before this filing date? Find Prior Art

Description

ON-FLOW CELL LIBRARY PREPARATION FOR SINGLE-CELL RNA ANALYSISCross-Reference to Related Applications

[0001] The present application claims priority to and the benefit of U.S. Provisional Application No. 63 / 718,270, titled “SINGLE CELL SPATIAL ANALYSIS METHODS,'’ filed November 8, 2024, and U.S. Provisional Application No 63 / 912.852, titled “ON-FLOW CELL LIBRARY PREPARATION FOR SINGLECELL RNA ANALYSIS,” filed November 6, 2025, the disclosures of which are hereby incorporated by reference for all purposes.Technical Field

[0002] The invention relates generally to the quantitative detection and analysis of molecules in a single cell or a sample of single cells.Background

[0003] Living organisms store genetic information in DNA. Genes in the coding regions of DNA are transcribed into messenger RNA (mRNA), which is translated into protein. Proteins play critical functional and structural roles in living organisms. For example, most enz mes are proteins, and those enzymes catalyze the metabolic reactions essential to life. It is also enzymes that copy DNA into mRNA. Proteins are also structural, and constitute the essential fibers of muscles, the predominant material of hair, as well as basic structural linkages within the cytoskeleton. Essentially, all such proteins are made by translating an mRNA into the protein. In fact, one mRNA can sen e as the template for synthesizing multiple copies of a protein.

[0004] Because cells need to change in response to different conditions, and need different proteins at different times, it is helpful if any given mRNA is short-lived. Most mRNA molecules have a lifetime measured in seconds or minutes. Nevertheless, the health of a cell, or its response to a pathogen, or a drug, or to age-specific developmental changes may be indicated by the quantities of mRNA molecules present in a cell. As aconsequence, there is interest in measuring levels of different mRNA transcripts present in cells or tissue. See Adil, 2021, Single- cell transcriptomics: cunent methods and challenges in data acquisition and analysis, Front Neurosci 15:a591122, incorporated by reference.Summary

[0005] De novo transcriptome assembly is challenging using short-read technologies alone. Short-read, reference-guided transcriptome analysis is widely utilized. However short read de novo assembly approaches may have limited ability to determine long- range information coming from the same starting template molecules. The presently described techniques provide methods for measuring analytes including RNA, DNA, and / or proteins in a biological sample that can provide characterization information for a particular analyte (e.g., sequence read data, gene expression levels) while also providing contiguity or association information that uses location information as an input to sequence assembly or to providing cell or nucleus content information. Although certain embodiments of the disclosure are discussed in the context of nucleic acid, e.g., RNA, analysis, it should be understood that the disclosure encompasses determination of a presence and / or level of one or more of RNA, DNA. or proteins.

[0006] The present techniques include characterization of analytes within a single cell using a two-step process by which, first, single-cell (or single nuclei) are encapsulated and molecular content (DNA / RNA / protein) is tagged. The tagging may be as generally discussed in W02024 / 077114, which is incorporated by reference in its entirety herein. Protein analytes may be characterized as generally discussed in U.S. Patent No. 11,965,880, which is incorporated by reference in its entirety' herein. The second step of the process provides the tagged molecules, or droplets containing the molecules, to a sequencing device in which on-board linked-read library preparation and sequencing take place. This library preparation and sequencing may be as generally discussed in U.S. Patent Nos. 10,443,087 and 9,683,230 and in WO2023 / 122755, which are incorporated by reference in their entireties herein. This process allows short reads (e.g., 150bp) originating from the same starting tagged molecules to be assembled into long-range contiguity information. This is achieved through bioinformaticallyassociating the unique tags incorporated during reverse transcription and library preparation with linked reads prepared and sequenced on the sequencer. Linked reads and tagged reads are then read back during processing of the sequencing reads from the final libraries and unique tags and linkages and are used to assemble long-range information coming from single cells. The standard approach to assemble single cell long-range contiguity information is by unique tagging of each molecule with a UMI and performing native long read sequencing on the resulting whole transcriptome amplified I i brary or concatenated formats of the library. This method is unique in that it is orthogonal from the UMI and native long-read approach. Instead of uniquely tagging molecules with a unique sequence, the tagging co-occurs both during the reverse transcription or extension step and during on-board library preparation, during which reads coming from the same molecule are spatially adjacent on the flow cell. Furthermore, this process enables simultaneous determination of both long-range information and also gene expression counting information using a single workflow because this spatial proximity serves as a tag both for assembling linked reads as well as for gene expression counting. The process does not require any additional work from the user and can be made transparent to them, although a bioinformatic pipeline to assemble linked reads and associate back to starting unique molecules or gene expression data will be used. This process is not limited to RNA reads and can support any analyte that is captured and converted to a DNA library using that assay.

[0007] In some embodiments, the method includes isolating the cells into partitions and creating sequencing libraries from single cells in the partitions. The solid support may be a bead attached to a plurality of copies of the cellular barcode sequence. The method may include attaching synthetic oligos (such as paired-end sequences, PCR handles, or sequencing adaptors) via transposomes.

[0008] Methods of the presently described techniques may include isolating a cell with a beads in an aqueous partition and lysing the cell to release the transcripts within the partition. A plurality of cells may be isolated into droplets, e.g., either in serial fashion using channels of a microfluidic platform or simultaneously by mixing cells with beads in water under oil and vortexing or shearing the mixture to generate the partitions (droplets). The beads may be decorated with capture oligonucleotides. All captureoligonucleotides on one bead may have a common barcode, which can serve as a cellular barcode. In some embodiments, a short binning index (BI) (e.g., about 3-6 bases) is added to the sequence information via workflow process steps. Each capture oligonucleotide, in certain embodiments, may include a binning index to perform deduplication as discussed herein and then a primer segment at the 3' end that anneals to target templates. After sample preparation and sequencing, the binning index appears in sequence reads.

[0009] Because the sequencing start site for each molecule is essentially random based on the fragmentation site, at least some of the bases in the sequencing information are essentially random and are thus also effectively unique to a specific mRNA or nucleic acid molecule, thereby functioning as a molecular identifier. As used herein a molecular identifier may be understood to include bases present in the naturally occurring sequence or its complement (i.e., intrinsic molecular identifiers (IMI)), bases added or otherwise not present in the naturally occurring sequence or its complement (i.e., extrinsic molecular identifiers), or a combination of intrinsic and extrinsic molecular identifiers. With this in mind, such a molecular identifier may be understood to comprise a unique (or functionally unique) sub-sequence in a source nucleic acid molecule. One or more such molecular identifiers may, alone or in conjunction with other information, uniquely identify a source nucleic acid molecule. Thus, in the context of intrinsic molecular identifiers, bases that are naturally and intrinsically present in each molecule are used to associate sequence data with the molecule from which that sequence data was read. Counts of unique sequence reads (based on IMI) are indexed by their binning indexes. The presently described techniques provide correction factors that apply to sequence read counts to correct for bias potentially introduced during sample preparation.

[0010] The sequence reads include the aforementioned effectively unique portions (e.g., the IMIs), and the method includes determining counts of the IMIs per genomic region and assigning the counts to associated binning indexes. Those assignments may be performed by writing in memory one or more files that include the counts indexed by the binning indexes. In certain implementations, after the assigning step, for each gene, only the binning indexes, the counts, and the correction factor are used to for theapplying and summing step. A correction factor, as discussed herein, is applied to the counts to reduce bias introduced during sample preparation, which may be an empirically derived measure of over- or under-representation of unique molecules when relying on IMIs to uniquely label molecules. The correction factor may be a divisor that reduces the counts by an expected factor if the molecules were subject to limited amplification prior to priming or fragmenting to generate the IMIs.

[0011] In related aspects, the presently described techniques provide a method of measuring expression levels. The method includes sequencing transcripts from random start sites, as discussed herein, to generate sequence reads, wherein each sequence read includes a binning index added by an oligonucleotide during sample preparation. Each sequence read is mapped to a gene, and the method includes obtaining counts per gene of unique intrinsic sequences (e.g., IMIs) defined by the random start sites; assigning the counts to associated binning indexes; and applying a correction factor to each count, to correct for bias introduced in the sample preparation. The corrected counts are summed across the binning indexes for each gene to provide an estimated number of the transcripts per gene in the sample. However, while certain aspects of the disclosure are discussed in the context of IMIs, it should be noted that the sequencing library fragments may additionally or alternatively make use of an exogenous unique molecular identifier (UMI) that may be incorporated via the capture oligonucleotide.

[0012] Disclosed embodiments include a method having steps of providing a sample comprising a plurality of cells, each cell of the plurality of cells comprising sample nucleic acids and isolating the plurality of cells into partitions such that an individual partition comprises: only one individual cell of the plurality of cells and a solid support comprising linked capture oligonucleotides. For the individual partition, the steps include capturing the sample nucleic acids of the individual cell onto the solid support using capture sequences of the capture oligonucleotides, the capture sequences of each individual partition comprising a unique identifier that is associated with the individual partition and distinguishable from other partitions of the partitions; and extending the capture sequences to form duplexes with the sample nucleic acids, wherein the duplexes comprise a sequence of the unique identifier. The method also includes amplifying the duplexes of the partitions using a primer pair, wherein one primer of the primer paircomprises a first adaptor sequence such that amplification products of the duplexes comprise the first adaptor sequence and its complement at one end of an individual amplification product; and contacting the amplification products with surface-linked transposome complexes of a flow cell surface to generate fragments of a sequencing library. The surface-linked transposome complexes include a transposase; and a transposon comprising a 3' portion comprising a first transposon end sequence and a second adaptor sequence at the 5' end of the first transposon end sequence; a second transposon comprising a second transposon end sequence complementary to at least a portion of the first transposon end sequence, wherein generating the fragments of the sequencing library comprises incorporating the second adaptor sequence into the amplification products such that the fragments comprise the first adaptor sequence and its complement at one end of an individual fragment and the second adaptor sequence and its complement at the other end of the individual fragment. The method includes generating sequencing data from the sequencing library', wherein individual sequencing reads comprise respective unique identifiers from respective partitions.

[0013] The various methods and techniques discussed herein, may be embodied and implemented as processor-executable code or stored routine, such as may be stored on a processor-accessible memory and executed via one or more hardware or virtual processors or dedicated circuitry. As such, it should be understood that actions or steps described a part of a method or process herein may be implemented as processor executable code, routines, or algorithms stored on one or more tangible computer- readable media.Brief Description of the Drawings

[0014] FIG. 1 shows a method for using location information from a flow cell surface as part of single-cell analyte characterization, in accordance with aspects of the present disclosure.

[0015] FIG. 2 shows a workflow for generating sequencing data from a single cell, in accordance with aspects of the present disclosure.

[0016] FIG. 3 is a schematic illustration of direct droplet loading or pooled droplet content loading, in accordance with aspects of the present disclosure.

[0017] FIG. 4 shows a method for sequencing single-cell analytes, in accordance with aspects of the present disclosure.

[0018] FIG. 5 shows example partitioning, in accordance with aspects of the present disclosure.

[0019] FIG. 6 diagrams RNA capture and library preparation with binning indexes, in accordance with aspects of the present disclosure.

[0020] FIG. 7 shows the generation of double-stranded cDNA with template switching oligos (TSOs) and tagmentation of double-stranded cDNA with transposons, in accordance with aspects of the present disclosure.

[0021] FIG. 8 shows the use of IMIs using capture oligos with binning indexes, in accordance with aspects of the present disclosure.

[0022] FIG. 9 shows an example workflow for single-cell expression analysis, in accordance with aspects of the present disclosure.

[0023] FIG. 10 show s partition contents before cell lysis in the workflow' of FIG. 9, in accordance with aspects of the present disclosure.

[0024] FIG. 11 shows first adaptor incorporation via amplification in the workflow of FIG. 9, in accordance with aspects of the present disclosure.

[0025] FIG. 12 shows flow cell loading to incorporate the second adaptor in the workflow of FIG. 9, in accordance with aspects of the present disclosure.

[0026] FIG. 13 illustrates a workflow' within a system performing aspects of the presently described techniques, in accordance with aspects of the present disclosure.

[0027] FIG. 14 depicts a sample process flow for estimating an inflation factor based on amplification-based inflation and selecting a correction factor, in accordance with aspects of the present disclosure.

[0028] FIG. 15 depicts a sample process flow for estimating inflation probabilities based on sample-specific observed count data, in accordance with aspects of the present disclosure.

[0029] FIG. 16 depicts a further view of a sample process flow for estimating inflation probabilities based on sample-specific observed count data, in accordance with aspects of the present disclosure.

[0030] FIG. 17 is a block diagram of an exemplary computing device, in accordance with aspects of the present disclosure.

[0031] FIG. 18 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.

[0032] FIG. 19 depicts an example of the transposome complexes that can be used in different examples of the methods and kits disclosed herein.

[0033] FIG. 20 depicts an example of the transposome complexes that can be used in different examples of the methods and kits disclosed herein.

[0034] FIG. 21 depicts an example of the transposome complexes that can be used in different examples of the methods and kits disclosed herein.

[0035] FIG. 22 depicts an example of the transposome complexes that can be used in different examples of the methods and kits disclosed herein.

[0036] FIG. 23A is a top view of a flow cell, in accordance with aspects of the present disclosure.

[0037] FIG. 23B is an enlarged, perspective, and cut-away of one architecture of the flow cell of FIG. 23 A where library preparation, amplification, and sequencing can take place.

[0038] FIG. 23C is an enlarged, perspective, and cut-away of another architecture of the flow cell of FIG. 23 A where library preparation, amplification, and sequencing can take place.

[0039] FIG. 23D is an enlarged, perspective, and cut-away of yet architecture of the flow cell of FIG. 23 A where off-target RNA depletion can take place.Detailed Description

[0040] The disclosure, in certain embodiments, provides methods of characterizing analytes that can be used for single cell / single nucleus analysis. FIG. 1 is a method 10 that includes a step of generating partitions (e.g., droplets) at block 12 to isolate individual cells of a sample that includes multiple cells. Within each partition is at most one cell (e.g., an individual cell) and one solid support (e.g., a bead). Accordingly, some partitions may be empty or may have one cell but no solid support or vice versa. In certain embodiments, within the partition, the solid support may be temporarily isolated from the single cell in a reversible manner (e.g., via a hydrogel material that can be heated to permit contact between the cell and the solid support).

[0041] The cell can be lysed to permit the contents of the cell to contact the solid support. This in turn allows the solid support to capture analytes of the cell and tag the analytes with a cell-specific tag at block 14. In certain cases, the solid support is coupled to a lawn of capture oligonucleotides that all carry7a same cell-specific barcode. Different solid supports in different partitions have different barcodes. In this manner, each individual cell can be associated with a unique barcode. Tagging may occur as generally discussed herein, such as via capture of cell nucleic acids and extension from 3’ ends of one or both of the hybridized oligonucleotides to form an at least partially duplex structure.

[0042] The tagged cell nucleic acids, or copied strands thereof, can be contacted with a flow cell surface at block 16. In an embodiment, the intact partition can be loaded onto the flow cell surface and subsequently disrupted to permit the contents of the partition to contact surface-linked transposomes. The transposomes cause fragmentation of the duplex structure such that previously contiguously regions of cell nucleic acids tend to associate on the flow cell surface relatively close to one another. Stated another way, there is a relationship between location on the flow cell surface and sequence contiguity. This spatial information can be used to assemble a sequence, such as an mRNA transcript or DNA sequence at block 18.

[0043] In certain embodiments described herein, the libraries may be created with emulsions and template particles that segregate individual cells into droplets upon vortexing. FIG. 2 is an overview of a sample processing workflow that uses particle templating to compartmentalize cells in monodispersed water-in-oil droplets. Rapid emulsification with a standard vortexer allows cells to be encapsulated at the bench or point of collection in minutes. The cells may be lysed inside the droplets, to release RNA. The RNA may be captured by bead-bound capture oligos that include a beadspecific, and thus a cell-specific barcode while in the droplets. The cells are lysed by increasing the temperature to 65 °C, which activates proteinase K (PK), releasing cellular mRNA that is captured on polyacrylamide bead solid supports decorated with capture oligonucleotides that include barcoded poly(T) sequences that can capture full- length mRNA transcripts. In the illustrated example, all of the capture oligonucleotides on a bead can have a same sequence. As generally discussed herein, while the illustrated example uses polyT capture, other capture sequences may also be used, including randomers, target-specific sequences, and / or aptamer sequences.

[0044] The capture oligos may be extended by a reverse transcriptase, copying the RNA to yield cDNA, which is provided with PCR primer binding sites (‘ PCR handles”). Preferably, at least one PCR handle is attached at an essentially random location in the cDNA such that a segment of the cDNA adjacent the random location provides a identifier sequence that is unique to that molecule. Those cDNAs may be amplified and sequenced. Sequence reads from the cells may be mapped to a reference and deduplicated, providing for identification and quantification of RNA from amultitude of single cells in one experiment, in which each cell was isolated in its own aqueous partition. Accordingly, methods of the invention provide a massively parallel, analytical workflow for preparing single-cell sequencing libraries. The methods are inexpensive, scalable, and accurate, and do not require UMIs in certain cases, because the molecular indexes can be generated via the fragmenting step performed on the flow cell.

[0045] In certain embodiments, the presently described techniques may be used to create single-cell sequencing libraries and, in particular, libraries useful in single-cell RNA-sequencing (scRNA- Seq). Some scRNA-Seq protocols involve sequencing RNA from a cell and, in certain embodiments, providing a measure of gene expression levels from the sequence data. Some approaches to scRNA-Seq rely on isolating cells into droplets with the potential to assay a large number of cells per experiment. Popular droplet-based protocols include Drop-seq (described in Macosko, 2015, Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets, Cell 161(5): 1202-14, incorporated by reference) and inDrop (see Klein, 2015, Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells, Cell 161(5): 1187-201, incorporated by reference). As discussed herein, the presently described techniques provide intrinsic molecular identifiers that may be used in such droplet-based protocols such that those protocols do not require UMIs (although the presently described technique may be implemented with or in addition to the use of UMIs).

[0046] The sample preparation may include tagmentation using Tn5 transposase linked to a flow cell surface to attach primer binding sites or sequencing adaptor at essentially random sites in the molecules. The sample preparation may include cleaving the templates with mechanical force, heat, chemicals such as detergents, or enzymes (e.g., endonucleases) followed by ligation of PCR handles or adaptors.

[0047] FIG. 3 shows different sample loading conditions for providing bead-captured analytes to the flow cell surface for downstream sequencing steps. In one embodiment, each contained or intact droplet can be loaded directly on the flow cell surface and subsequently broken to release analyte contents. Such an embodiment may permitdifferent analyte types (e.g. protein, nucleic acid) in the lysed cell contents, captured and uncaptured, to associate with the flow cell surface in proximity to one another to undergo transposome-mediated fragmentation and, in certain cases, IMI incorporation to generate molecular indexes via random fragmentation. In another embodiment, the droplets can be broken and the contents from different individual droplets and their associated single cells pooled (e g., a pool of different beads with respective captured cell nucleic acids forming duplex structures with capture oligonucleotides) prior to loading the beads onto the flow cell surface. Such an embodiment permits a pooled extension step for captured or bead-bound cell nucleic acids, which may be more efficient. Once duplexes are formed, the duplex structures can be loaded onto the flow cell surface (e.g., while in bead-bound form in an embodiment or after separation / cleavage from the bead in an embodiment) to undergo transposome- mediated fragmentation. In both embodiments, when a duplex structure undergoes transposome-mediated fragmentation, the now-separated but previously contiguous fragments are nonetheless in proximity to one another on the flow cell surface. The flow cell surface can also include a lawn of sequencing primers that bind strands of the separated fragments to initiate cluster formation for sequencing. The complementary sequences that permit binding to the surface-associated primers can be incorporated via the capture oligonucleotides and / or via the transposome-mediated fragmentation.

[0048] FIG. 4 shows a block diagram of a method 101 for preparing a sequencing library. The method 101 includes reverse transcribing 103 RNA into cDNA. Each cDNAs is cleaved 109 at a random location, or "random cut site’', and a synthetic oligonucleotide that includes a binning index is attached 115 at the random cut site. Notably, the method 101 includes indexing 116 the template molecules, by virtue of having added the binning index. Optionally or alternatively, the random site may be defined by random priming, e.g., using a random hexamer. The cleavage and attachment may be carried out by any suitable methods. For example, fragmenting 109 may be performed by physical methods, such as acoustic shearing or sonication, or by enzymatic methods, such as with a restriction enzyme, or by exposing the RNA to high temperatures, e.g., about 95 degrees Celsius, in the presence of multivalent cations, such as, metal ions, for example, Mg2+, Mn2+. or Zn2+. For example, the RNA maybe incubated in a solution comprising MgC 12, at 95 degrees Celsius, for a few minutes. In certain embodiments, the cleavage 109 and attachment 115 are performed by a transposase such as a Tn5 transposase.

[0049] Cleavage 109, according to aspects of the presently described techniques, generate cut sites at substantially random positions in the cDNA. Because cleavage of the cDNA is at a random cut site, the cleaved ends of the cDNA molecules are random and essentially unique. A downstream (i.e., later) step of the method involves reading sequence from the cleaved ends of the cDNA molecules. Those sequences can be treated as unique if enough bases are read from the cleaved end. That is, if there are hundreds of thousands of cDNA molecules, and only a 3-base intrinsic label is read, then there are only 64 possible unique labels. However, if 10 bases are read (assuming random use of bases) then there are greater than 1 million unique labels. Reading 12 bases gives more than 16 million labels. Fifteen bases give more than 1 billion labels. Reading 17 bases provides more than 17 billion labels; 18 bases give> 68B labels; 19 b give> 274 B labels; and 20 b give> 1 trillion labels. For many applications, it is not necessary7that each intrinsic molecular identifier (IMI) be unique. For example, in various applications, having two, or three, or even ten IMIs that have a duplicate among the set will yield scRNA-seq quantitative results that are useful and significant, i.e., not statistically significantly different than without the duplicates for many end-uses.

[0050] A slight variant of method 101 does not use cleavage of the cDNA but instead uses primer binding at a random site such that bases that are intrinsic within the target nucleic acid adjacent the random primer binding site are copied into new DNA and come to serve as a unique intrinsic molecule identifier. These versions may use random hexamers, which are suitable as primers for capturing volumes of RNA such as mRNA from a single cell. In such embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by a random hexamer. For example, an RNA may be captured by a random primer that is extended to create a cDNA. Because random priming (e.g., using a random hexamer) binds at effectively random sites within nucleic acid, each cDNA will include a segment of bases in the cDNA adjacent the priming site that is useful as a unique identifier sequence. Random priming works similarly to random cleavage by a transpose, random mechanical or chemicalcleavage, or restriction enzyme cleavage. What is in common among those techniques is that, as far as the sequences of the target nucleic acids are concerned, the binding sites or cut sites are effectively random. Each target nucleic acid will be bound or cut in manner that is unpredictable or inconsistent enough, for the purposes of techniques such as scRNA-seq, that downstream amplicons will have binning indexes and substantially unique intrinsic molecular identifiers (noting again that nearly unique is sufficient for most purposes) at one end that get sequenced and appear in sequence reads.

[0051] The intrinsic molecular identifiers may be used to establish, or contribute to. the unique molecular identity of nucleic acids of a library. For example, RNA transcribed from the same genomic loci may have sequences that are substantially identical. Here, by randomly cleaving the cDNA, each cDNA is made unique by virtue of the bases adjacent the random cut site left by the cleavage 109.

[0052] A synthetic oligonucleotide that includes a binning index is attached 115 to the cDNA at the random cut site to create a construct that includes at least a portion of the cDNA and the synthetic oligonucleotide. Remembering that the cDNA was created by extending a capture oligonucleotide that annealed to an RNA, any sequence in the capture oligonucleotide will be present in the construct. Any sequence present in the synthetic oligonucleotide will also be present in the construct. The capture oligonucleotide and the synthetic oligonucleotide may either or both have a PCR handle (i.e., a primer binding site, a "‘universal primer binding site”, a capture tag, a sequencing adaptor, or similar). Thus, in certain embodiments, the attachment 115 creates a construct that includes a first PCR handle, an optional functional sequence such as a sample barcode and / or a cell barcode, ahybrid capture portion of the capture oligo (e.g., a poly-T region), a portion of the cDNA, the location of the random cut site, and a second PCR handle. The attachment also indexes 116 the template (by labeling with a binning index).

[0053] Because the construct is, in certain embodiment, a contiguous DNA molecule with PCR handles at both ends, it is amenable to amplification 123 by, for example, polymerase chain reaction (PCR). The optional functional sequence may be a cellbarcode. More specifically, the capture oligonucleotide may be one of a plurality of capture oligonucleotides that are attached to a solid support such as a bead, e.g., a hydrogel bead. All capture oligonucleotides may share one common barcode (which may be referred to as a “bead barcode”). If the bead is isolated in an aqueous partition with a single cell and the capture oligonucleotides are used to anneal to, and capture, RNA molecules from the single cell, then the common barcode of the bead becomes a cell barcode. This is because, downstream, after sequencing 127, the presence of the cell barcode sequence in a sequence read is useful to map that sequence read back to the single cell associated with that bead. Thus, a barcode in the construct may be used as a cell barcode. All such constructs from a single cell may be amplified 123.

[0054] The constructs may include sequence platform specific primers (e.g., P5 and P7) or those may be added by a round of amplification, e.g., PCR. Amplification 123 produces amplicons which may be sequenced 127. Due to the sample preparation, the sequencing is effectively initiated from random start sites to generate sequence reads. Each sequence read is indexed 116 by a binning index added by an oligonucleotide during sample preparation.

[0055] The method 101 may further include mapping each sequence read to a gene; obtaining counts per gene of unique intrinsic sequences defined by the random start sites; assigning the counts to associated binning indexes; applying a correction factor to each count, to correct for bias introduced in the sample preparation; and summing corrected counts across the binning indexes for each gene to provide an estimated number of the transcripts per gene in the sample. The correction factor may: adjust for a probability of the random start sites being duplicated among the transcripts; adjust for a probability of multiple random start sites per transcript within the sample; include an estimate of a number of copies of each transcript resulting from the amplification and the applying step includes dividing each count by the correction factor; or provide any other suitable adjustment or correction to the read counts. Informatically, after read counts are assigned to their binning indexes, it is possible that for each gene, only the binning indexes, the counts, and the correction factor are used to for the applying and summing step.

[0056] As discussed, methods of the presently described techniques are useful for scRNA-Seq and specifically for expression analysis. In certain embodiments, cells are isolated into, and lysed within, aqueous partitions with capture oligonucleotides that include binning indexes. The capture oligonucleotides anneal to RNAs released from the cells. The capture oligonucleotides may include partition-specific barcodes, binning indexes, and PCR handles. Once the capture oligonucleotides have hybridized to the RNAs, those duplexes may be released from partitions and pooled at any subsequent stage. Because capture oligonucleotides with partition-specific barcodes are used to capture and tag RNA from cells isolated in the partition, any arbitrary number of cells may be captured in parallel (i.e., simultaneously). Because the RNAs are tagged with a cell barcode during hybrid capture (e.g., a partition-specific barcode or a bead barcode), if those duplexes are pooled and ultimately sequenced, the cell barcodes in the sequencing data can be used to “bin” the sequence data by original cell, i.e., assign each sequence read (or assembled contigs or sequences therefrom) back to originating cells.

[0057] Multiplexing in the present context may involve isolating cells and the capture oligonucleotides into partitions. Any suitable partitions may be used. The partitions may be any suitable partition in a pico-, nano-, or microtiter plate or substrate, or fluidic harbors (see, e.g.. US Pub 2010 / 0041046 Al, incorporated by reference), chambers (see, e.g., 20210178395 Al, incorporated by reference), regions defined within a fluidic device (see, e.g., 20200269248 Al, incorporated by reference), others, or combinations thereof. In certain embodiments, the partitions are aqueous partitions in an immiscible liquid, e.g., slugs or droplets surrounded or separated by oil within a microfluidic device. A microfluidic device may use channels to mix samples and reagents and form droplets in an immiscible carrier fluid. In certain embodiments, the partitions are a plurality of droplets that are formed essentially simultaneously. Methods may be performed with a sample comprising a mixture with cells, and, in certain embodiments, template particles. The mixture may comprise two immiscible fluids such as an aqueous fluid and oil. The mixture is sheared, e.g., vortexed, to generate an emulsion with template particles that sen e to template the formation of droplets and segregate individual cells into the droplets. Because the cells are individually segregated into droplets, the cells may be individually profiled in parallel. This method provides amassively parallel, analytical workflow for analyzing single cells that is inexpensive, scalable, and accurate.

[0058] For example, methods of the presently described techniques may include combining template particles with cells in a first fluid and then adding a second fluid that is immiscible with the first fluid to the mixture. The first fluid is, in certain embodiments, an aqueous fluid. While any suitable order may be used, in some instances, a tube may be provided comprising the template particles. The tube can be any type of tube, such as a sample preparation tube sold under the trade name Eppendorf, or a blood collection tube, sold under the trade name Vacutainer. The sample may be a blood sample and may be added directly to the tube using a pipette.

[0059] The fluids can be sheared to generate a monodisperse emulsion with droplets. To generate a monodisperse emulsion, methods may include a step of shearing the mixture provided by combining cells and template particles in an aqueous fluid with the immiscible fluid. Any suitable method or technique may be utilized to apply a sufficient shear force to the mixture. For example, the mixture may be sheared by flowing the second mixture through a pipette tip. Other methods include, but are not limited to, shaking the mixture with a homogenizer (e.g.. vortex er), or shaking the mixture with a bead beater. In some embodiments, vortex may be performed for example for 30 seconds, or in the range of 30 seconds to 5 minutes. The application of a sufficient shear force breaks the mixture into monodisperse droplets that encapsulate one of a plurality of template particles.

[0060] After vortexing, a plurality (e.g., thousands, tens of thousands, hundreds of thousands, one million, two million, ten million, or more) of aqueous partitions is formed essentially simultaneously. Vortexing causes the fluids to partition into a plurality of monodisperse droplets. A substantial portion of droplets will contain a single template particle and a single target cell. Droplets containing more than one or none of a template particle or target cell can be removed, destroyed, or otherwise ignored.

[0061] The next step of the method is to lyse the cells. Cell lysis may be induced by a stimulus, such as, for example, lytic reagents, detergents, or enzymes. Reagents to induce cell lysis may be provided by the template particles via internal compartments. In certain embodiments, lysing involves heating the monodisperse droplets to a temperature sufficient to release lytic reagents contained inside the template particles into the monodisperse droplets. This accomplishes cell lysis of the target cells, thereby releasing nucleic acids, such as RNA, such as mRNA, inside of the droplets that contained the target cells.

[0062] After lysing target cells inside the droplets. mRNA is released. The mRNA may be used to create a sequencing library. Methods and systems of the presently described techniques may use template particles to template the formation of monodisperse droplets and isolate single target cells. The disclosed template particles and methods for targeted library preparation thereof leverage the particle-templated emulsification technology described in Hatori. 2018. Particle-templated emulsification for microfluidics-free digital biology. Anal Chem 90(16):9813-9820, incorporated by reference. Essentially, micron-scale beads (such as hydrogels) or “template particles” are used to define an isolated fluid volume surrounded by an immiscible partitioning fluid and stabilized by temperature insensitive surfactants.

[0063] In practicing the methods as described herein, the composition and nature of the template particles may vary'. For instance, in certain aspects, the template particles maybe microgel particles that are micron-scale spheres of gel matrix. In some embodiments, the microgels are composed of a hydrophilic polymer that is soluble in water, including alginate or agarose. In other embodiments, the microgels are composed of a lipophilic microgel.

[0064] FIG. 5 illustrates a sample prep tube 229 compnsing droplets 201. In particular, the sample prep tube 229 comprises a plurality of monodisperse droplets generated by shearing a mixture 239 according to certain implementations of the present techniques. In certain embodiments, each of the droplets 201 includes, on average, one template particle 213 and zero or one single target cell 209. The template particles 213 may comprise crater-like depressions to facilitate capture of single cells 209. The templateparticles 213 may further comprise an internal compartment 221 to deliver one or more reagents into the droplets 201 upon stimulus. Each template particle 213 may be decorated with capture oligonucleotides that include a binning index and optionally a cell barcode and a 3' hybrid capture portion.

[0065] In some embodiments, the template particles 213 contain internal compartments 221. The internal compartments 221 of the template particles 213 may be used to encapsulate reagents that can be triggered to release a desired compound, e.g., a substrate for an enzymatic reaction, or induce a certain result, e.g., lysis of an associated target cell. Reagents encapsulated in the template particles' compartment may be without limitation reagents selected from buffers, salts, lytic enzymes (e.g., proteinase k), other lytic reagents (e. g. Triton X-100, Tween-20, IGEPAL), nucleic acid synthesis reagents, or combinations thereof.

[0066] Lysis of single target cells occurs within the monodisperse droplets and may be induced by a stimulus such as heat, osmotic pressure, lytic reagents (e.g., DTT, betamercaptoethanol), detergents (e.g., SDS, Triton X-100, Tween-20), enzymes (e.g., proteinase K), or combinations thereof. In some embodiments, one or more of the reagents (e.g., lytic reagents, detergents, enzymes) is compartmentalized within the template particle 213. In other embodiments, one or more of the reagents is present in the mixture. In some other embodiments, one or more of the reagents is added to the solution comprising the monodisperse droplets, as desired.

[0067] In certain embodiments, template particles 213 comprise a plurality of capture probes.

[0068] Generally, the capture probe of the present disclosure is an oligonucleotide. In some embodiments, the capture probes are attached to the template particle's material, e.g., hydrogel material, via covalent acrylic linkages. In some embodiments, the capture probes are acrydite-modified on their 5' end (linker region). Generally, acrydite- modified oligonucleotides can be incorporated, stoichiometrically, into hydrogels such as polyacrylamide, using standard free radical polymerization chemistry, where the double bond in the acrydite group reacts with other activated double bond containingcompounds such as acrylamide. Specifically, copolymerization of the aery di te- modified capture probes with acrylamide including a crosslinker. e.g.. N,N'- methylenebis, will result in a crosslinked gel material comprising covalently attached capture probes. In some other embodiments, the capture probes comprise acry late terminated hydrocarbon linker and combining the capture probes with a template particle 213 causes their attachment to the template particle 213.

[0069] In some embodiments, after cell suspensions are introduced to template particles 213 in a pre-equilibrated buffer, droplets are generated by vortexing the mixture to capture single cells with individual template particles. The resulting emulsion may be heated on a thermocycler to induce cell lysis. Cell lysis releases the contents of the cell and exposes those contents, including mRNA, to the template particle 213. The presently described techniques provide steps for RNA capture and library7preparation using those particles and released mRNA.

[0070] FIG. 6 diagrams RNA capture and library7preparation according to certain embodiments of the presently described techniques. As shown, particle 213 is linked to a capture oligonucleotide 305. The capture oligonucleotide 305 anneals to an mRNA 311. The capture oligonucleotide 305 includes a binning index 316. Optionally, methods include capturing transcripts with capture oligonucleotides linked to beads, in which each capture oligonucleotide includes 5'-linkage to bead, cell barcode, binning index 316, annealing primer section-2'. The binning index 116 may comprise, in some implementations, three or fewer bases. In certain embodiments, the binning indexes are 3 bases in length. In other embodiments, the binning indexes are either 1 or 2 or 3 bases or are absent, as if an equimolar mixture of 0, 1 , 2, and 3 bases among all of the capture oligonucleotides on a bead.

[0071] In some embodiments, poly-T tails of the capture oligonucleotides 305 anneal to and capture RNA released by7lysis. Particle-bound capture oligonucleotides in this application may comprise an acrydite linker, a PEI priming sequence, a particle barcode, optionally a random sequence, and a poly-T capture moiety. A polymerase (e.g., a reverse transcriptase) extends the capture oligonucleotide 305 to form a cDNA 315. The cDNA 315 and capture oligonucleotide 305 in combination with the mRNA311 form a duplex 323. This duplex is stably linked to the bead 213. At this stage, it is suitable to break the droplets (e.g., using perfluorooctanol (PFO) or chloroform). Depending on the desired workflow, the droplets may be applied intact to a flow cell surface and then broken. In one embodiment, the droplets are broken / their contents pooled, and then washed in buffer before proceeding in library preparation on a flow cell surface as generally discussed herein.

[0072] A surface-linked transposase complex 325 (sometimes called a transposome) is introduced. The transposase complex 325 includes a dimer that includes two of a transposase 327 and two transposon end sequences 329. Here, the transposon end sequences 329 are depicted as both being paired-end 2 end (PE2) sequences, which will cooperate with paired-end 1 (PEI) sequences in the capture oligonucleotide 305 in subsequent amplification and sequence steps. In the depicted method, the transposase randomly cuts the cDNA / mRNA duplex 323 thereby defining a random cut site 333. In a downstream step, read 2 of paired-end sequencing will include the first segment of bases in the cDNA 315 adjacent the random cut site 333.

[0073] Attachment 115 of the end sequence 329 to the cDNA 315 at the random cut site 333 produces a construct 337. The construct 337 is a contiguous DNA molecule that includes a first PCR handle (PEI), a cell barcode, a capture segment, a portion of the cDNA 315 terminating at the random cut site 333, and a second PCR handle (PE2).

[0074] Amplification 123 of the construct 337 yields amplicons 341. In some embodiments, constructs are amplified with a P5-PE1 hybrid oligonucleotide and P7 index primer directly into a sequencing library. The library may be sequenced to assess RNA expression, for example, as described in Hrdlickova, 2017, RNA-Seq methods for transcriptome analysis. Wiley Interdisc Rev RNA 8(1): I 0.1002, incorporated by reference.

[0075] Constructs or amplicons may include certain primer and index sequences or copies thereof, such as, P5s and P7s. Those sequences may be any arbitrary sequence useful in downstream analysis. For example, they may be additional universal primer binding sites or sequencing adaptors. For example, either or both of the P5s and P7smay be arbitrary universal priming sequence (universal meaning that the sequence information is not specific to the naturally occurring genomic sequence being studied, but is instead suited to being amplified using a pair of cognate universal primers, by design). The index segment may be any suitable barcode or index such as may be useful in downstream information processing. It is contemplated that the P5 sequences, the P7 sequence, and the index segment may be the sequences use in NGS indexed sequences such as performed on an NGS instrument sold under the trademark ILLUMINA, and as described in Bowman, 2013, Multiplexed Illumina sequencing libraries from picogram quantities of DNA, BMC Genomics 14:466 (esp. in Figure 2), incorporated by reference.

[0076] A transposase 327 is used to randomly cut the cDNA 315. This may be performed using a transposase such as Tn5. See Lin, 2020, RNA sequencing by direct tagmentation of RNA / DNA hybrids, PNAS1 17 (6) 2886-2893, incorporated by reference. In brief, the Tn5 transposase randomly binds and cuts double-stranded RNA / DNA and attaches its end sequence to the random cut site.

[0077] Accordingly, some embodiments of the presently described techniques use Tn5 transposase to directly tagment RNA / DNA hybrids and form polynucleotide libraries with intrinsic molecular identifiers (essentially unique sequences of bases originating in genetic material of the organism or biological system being studied). In particular, Tn5, a RNase H superfamily member, binds to RNA / DNA hybrids similarly as to dsDNA and effectively cuts randomly and then ligates a desired oligonucleotide onto the hybrid. The desired oligonucleotide is, in certain embodiments, a PCR handle (aka a universal primer binding site, a sequencing adaptor, a synthetic oligonucleotide of known sequence to which a PCR primer anneals, etc.). Methods of the presently described techniques may be used w ith various amounts of input sample, from single cells to large numbers of cells, with a dynamic range spanning numerous orders of magnitude.

[0078] FIG. 7 shows a workflow for directional tagmentation that w orks with template switching oligonucleotides (TSOs). The illustrated technique may be employed in hybrid workflows where one is using TSOs for some other benefit and one also want touse IMIs. This tagmentation approach is useful for 3' end capture and analysis of mRNAs. The steps of the method are shown. In brief. mRNA or total RNA from lysed cells are mixed with an oligonucleotide and incubated (such as at 65° Celsius for 3 minutes). The oligonucleotide may include specific primers for amplifying final libraries, such as an adapter-B sequence complementary to an i7 primer. The oligonucleotide may further include a poly-T sequence of, for example. 30 nucleotides that hybridizes with poly -A tails of mRNA. Importantly, the use of this oligonucleotide to prime a first strand cDNA synthesis may result in libraries enriched for the 3' end of mRNA.

[0079] In this figure, notably, the binning index could be added at any of several different steps. The binning index could be part of the oligonucleotide in step 1. The binning index could be part of the template switch oligonucleotide in step 2. The binning index could be part of the adaptor added by Tn5 in step 4. The binning index could be part of either sequence linker in step 5.

[0080] Reverse transcription can be performed using a reverse transcriptase such as the reverse transcriptase sold under the trade name SMARTSCRIBE by Takara Bio optionally in the presence of a template switching oligonucleotide (TSO). The template switching oligonucleotide allows for template switching at the 5' end of the mRNA molecule to incorporate an oligonucleotide such as a universal 3' sequence during first strand cDNA synthesis. Synthesis of the first cDNA strand may be performed using a thermocycler at 42 degrees Celsius for 1 h, followed by 15 minutes at 70 degrees Celsius to inactivate the reverse transcriptase. Afterwards, the cDNA may be amplified. The cDNA may be amplified by PCR using commercially available kits such as the kit sold under the trade name OneTaq HS by New7England Biolabs. After amplification, the RNA / DNA duplexes may be subjected to tagmentation and adapter ligation.

[0081] During tagmentation and adapter ligation, Tn5 bound adapter (adapter-A) complexes bind with the double RNA / DNA duplexes. The duplexes are cut by the enzymatic activity of the Tn5 complexes and the adapters ("Adaptor A"’) are ligated. Tn5 cuts at a random site. Afterwards, the products of the tagmentation reaction may be amplified using the adapters. In certain embodiments, each adaptor includes abinning index. As shown, an i7 primer anneals to Adapter B and an i5 primer anneals to Adapter A. In this depicted embodiment (as drawn) the read 1 primer will read into a segment of an amplicon adjacent the random cut site. Because the cut site is random, a sequence of bases in that segment is essentially unique. Because the segment is in the amplicon copy of the cDNA, the sequence of the bases is intrinsic to the mRNA, i.e., is a sequence from genetic material of the organism being studied. Because the sequence of bases is essentially unique, a read 1 sequence read will include a unique, intrinsic molecular identifier. More specifically, all sequence reads from the read 1 primer from this library7member will include the identical copies of that unique, intrinsic molecular identifier (IMI). Thus, the figure illustrates that IMIs are compatible with workflows that include or use TSOs.

[0082] After sequencing, reads with identical gene-mapping and identical IMIs can be “collapsed"’, and a count of only unduplicated reads is a quantitative measure of gene transcripts in the sample, i.e., the single cell.

[0083] As shown in FIG. 7, the use of IMIs is compatible with RNA capture without necessarily requiring any bead-linked capture oligonucleotides. That is, capture oligonucleotides may be free in solution (as opposed to linked to a solid support). Methods of the presently described techniques are also compatible with the use of capture oligonucleotides that are linked to a solid support such as a bead as discussed herein.

[0084] FIG. 8 shows a method for making libraries that include IMIs using capture oligonucleotides linked to a solid support. The depicted embodiment show s the creation of a sequencing library' that includes certain next-generation sequencing (NGS) adaptors. The solid support may be a bead and library preparation may be performed using a microfluidic device (e.g., to encapsulate beads, cells, and reagents into droplets). In some embodiments, beads decorated with capture oligonucleotides are used to simultaneously form a monodisperse emulsion that includes a plurality’ of droplets. Each droplet includes, on average, one bead and one or zero cells. Because the beads are particles that serve as templates that cause the droplets (or aqueous partitions) to form (e g., when a mixture is vortexed), the droplets may be referred to asparticle-templated instant partitions (PIPs), the beads maybe referred to as template particles, and sequencing from such libraries may be referred to as PIP-seq. In the illustrated embodiment, a template particle 1301 is linked to a capture oligonucleotide 1305. As shown, the template particle 1301 is linked to (among other things) mRNA capture oligonucleotides 1305 that include a 3' poly-T region 1309 (although sequencespecific primers or random N-mers may be used). Where the sample includes cell-free RNA, the capture oligonucleotide hybridizes by Watson-Crick base-pairing to a target in the RNA and serves as a primer for reverse transcriptase, which makes a cDNA copy of the RNA. Where the initial sample includes intact cells, the same logic applies but the hybridizing and reverse transcription occurs once acell releases RNA (e.g., by being lysed).

[0085] In certain embodiments, the target RNAs are mRNAs 1313. Where the target RNAs are mRNAs, the particles 1301 may include mRNA capture oligonucleotides 1305 used to at least synthesize cDNA 1317 as a copy of an mRNA 1313. The particles 1301 may further include cDNA capture oligonucleotides with 3' portions that hybridize to cDNA copies of the mRNA. For the cDNA capture oligonucleotides, the 3' portions may include gene-specific sequences or hexamers. As shown, each of the mRNA capture oligonucleotides 1305 may include, from 5' to 3', a SMART site 1319, a PEI sequence 1321, a cell or droplet barcode 1323, and a poly-T segment 1309. The capture oligos may also include a binning index 1316 in certain implementations.

[0086] As shown, the capture oligonucleotide 1305 hybridizes to the mRNA 1313. A reverse transcriptase binds and initiates synthesis of a cDNA copy 1317 of the mRNA 1313 to make an RNA / DNA hybrid. Note that the mRNA 1313 is connected to the particle 1301 non-covalently, by complementary base-pairing. The cDNA 1317 that is synthesized may be covalently linked to the particle by virtue of the phosphodiester bonds formed by the reverse transcriptase.

[0087] A transposase 1401 binds to the RNA / DNA hybrid. The transposase 1401, which may be a Tn5 transposase in certain embodiments, is attached with adapters 1406 for attaching onto the 5' end of the cDNA 1317. The Tn5 cuts the RNA / DNA hybrid at a random cut site 1351, and the adapters 1406 are ligated onto the random cut site ofthe cDNA 1317. In certain embodiments the adapter 1406 includes a primer handle 1403 for copying / amplification. At this stage, RNaseH may be introduced to degrade the mRNA 1313. The adaptor 1406 may include a binning index.

[0088] In some embodiments, sequencing adapter 1501 is extended to create a dsDNA 1409. The adapter 1501 includes a first sequence 1503 complementary to the primer handle 1403 and a sequencing primer 1505, such as P7. The adapter 1501 will hybridize to, and prime the copying of, cDNA to create a dsDNA 1409 with the sequencing adapter. Afterwards, the polynucleotide can be separated from the particle and made into a final library product. As discussed herein, the cDNA 1317, and thus also the dsDNA 1409, has a segment adjacent the random cut site 1351 with sequence intrinsic to the mRNA 1313. When that segment is sequenced from primer handle 1403, the resultant sequence reads include the intrinsic sequence.

[0089] Amplification produces a final library product 1601. In this example, the final library product 1601 is formed by the PCR-based extension of a P5-PE1 primer 1505 that is complementary to the PEI 1509 of the released polynucleotide 1409. Extension of the P5-PE1 primer 1505 by PCR creates the final library product 1601. In some embodiments, the P5-PE1 primer 1505 may include indexes, such as an 15 index, and a P5 index. The final library product may be amplified by PCR in advance of sequencing.

[0090] FIG. 9 shows an example workflow for single-cell expression analysis. One of the challenges for single-cell workflows is the relatively long sequencing library preparation workflow. In standard operation, this may take up to two days to complete library preparation prior to sequencing. By performing a distributed adaptorization in which one adaptor sequence is incorporated during amplification and a second adaptor sequence is incorporated via surface-linked transposome complexes, the overall workflow is simplified. In the illustrated workflow, individual cells are partitioned such that each partition includes a single cell or no cells (see FIG. 5 for example partitioning via emulsification).

[0091] In certain embodiments, cells are isolated into, and lysed within, aqueous partitions with capture oligonucleotides that include a capture sequence, such as a polyTcapture sequence. However, it should be understood that other capture sequences may be used. Cell partitioning and lysis may be performed as discussed herein, and, in an embodiment, this may be a relatively rapid workflow step (e.g., 1-2 hrs).

[0092] The capture oligonucleotides anneal to RNAs released from the cells. The capture oligonucleotides may include partition-specific barcodes, optionally binning indexes, and PCR handles (e.g., Illumina PEI by way of example, but any PCR or primer binding sequence may be used). The PCR handle may be common between partitions, but the partition-specific barcode or partition unique identifier is unique and distinguishes one partition from another. Once the capture oligonucleotides have hybridized to the RNAs, reverse transcnption occurs to form duplexes. Because capture oligonucleotides with partition-specific barcodes are used to capture and tag RNA from cells isolated in the partition, any arbitrary number of cells from one sample may be captured in parallel (i.e., simultaneously). Because the RNAs are tagged with a cell barcode during hybrid capture (e.g., a partition-specific barcode or a bead barcode), if those partitions are broken and the duplexes are combined and ultimately sequenced, the cell barcodes in the sequencing data can be used to “bin” the sequence data by original cell, i.e., assign each sequence read (or assembled contigs or sequences therefrom) back to originating cells. The cell barcodes may be as generally discussed herein.

[0093] As discussed herein, any suitable partitions may be used. The partitions may be any suitable partition in a pico-, nano-, or microtiter plate or substrate, or fluidic harbors (see, e.g., US Pub 2010 / 0041046 Al, incorporated by reference), chambers (see, e.g., 20210178395 Al, incorporated by reference), regions defined within a fluidic device (see, e.g., 20200269248 Al, incorporated by reference), others, or combinations thereof. In certain embodiments, the partitions are aqueous partitions in an immiscible liquid, e.g., slugs or droplets surrounded or separated by oil within a microfluidic device. A microfluidic device may use channels to mix samples and reagents and form droplets in an immiscible carrier fluid. In certain embodiments, the partitions are a plurality of droplets that are formed essentially simultaneously. Methods may be performed with a sample comprising a mixture with cells, and, in certain embodiments, template particles. The mixture may comprise two immiscible fluids such as an aqueous fluidand oil. The mixture is sheared, e.g., vortexed, to generate an emulsion with template particles that serve to template the formation of droplets and segregate individual cells into the droplets. Because the cells are individually segregated into droplets, the cells may be individually profiled in parallel. This method provides a massively parallel, analytical workflow for analyzing single cells that is inexpensive, scalable, and accurate.

[0094] For example, methods of the presently described techniques may include combining template particles with cells in a first fluid and then adding a second fluid that is immiscible with the first fluid to the mixture. The first fluid is, in certain embodiments, an aqueous fluid. While any suitable order may be used, in some instances, a tube may be provided comprising the template particles. The tube can be any type of tube, such as a sample preparation tube sold under the trade name Eppendorf, or a blood collection tube, sold under the trade name Vacutainer. The sample may be a blood sample and may be added directly to the tube using a pipette.

[0095] The fluids can be sheared to generate a monodisperse emulsion with droplets. To generate a monodisperse emulsion, methods may include a step of shearing the mixture provided by combining cells and template particles in an aqueous fluid with the immiscible fluid. Any suitable method or technique may be utilized to apply a sufficient shear force to the mixture. For example, the mixture may be sheared by flowing the second mixture through a pipette tip. Other methods include, but are not limited to, shaking the mixture with a homogenizer (e.g.. vortexer), or shaking the mixture with a bead beater. In some embodiments, vortex may be performed for example for 30 seconds, or in the range of 30 seconds to 5 minutes. The application of a sufficient shear force breaks the mixture into monodisperse droplets that encapsulate one of a plurality of template particles.

[0096] After vortexing, a plurality (e.g., thousands, tens of thousands, hundreds of thousands, one million, two million, ten million, or more) of aqueous partitions is formed essentially simultaneously. Vortexing causes the fluids to partition into a plurality of monodisperse droplets. A substantial portion of droplets will contain a single template particle and a single target cell. Droplets containing more than one ornone of a template particle or target cell can be removed, destroyed, or otherwise ignored.

[0097] Methods and systems of the presently described techniques may use template particles to template the formation of monodisperse droplets and isolate single target cells. The disclosed template particles and methods for targeted library preparation thereof leverage the particle-templated emulsification technology described in Hatori, 2018, Particle-templated emulsification for microfluidics-free digital biology, Anal Chem 90(16):9813-9820, incorporated by reference. Essentially, micron-scale beads (such as hydrogels) or “template particles"’ are used to define an isolated fluid volume surrounded by an immiscible partitioning fluid and stabilized by temperature insensitive surfactants.

[0098] In practicing the methods as described herein, the composition and nature of the template particles may vary. For instance, in certain aspects, the template particles may be microgel particles that are micron-scale spheres of gel matrix. In some embodiments, the microgels are composed of a hydrophilic polymer that is soluble in water, including alginate or agarose. In other embodiments, the microgels are composed of a lipophilic microgel.

[0099] The next step of the method is to lyse the cells within the partition. FIG. 10 shows an example partition before cell lysis. Cell lysis may be induced by a stimulus, such as, for example, lytic reagents, detergents, or enzymes. Reagents to induce cell lysis may be provided by the template particles via internal compartments. In certain embodiments, lysing involves heating the monodisperse droplets to a temperature sufficient to release lytic reagents contained inside the template particles into the monodisperse droplets. This accomplishes cell lysis of the target cells, thereby releasing nucleic acids, such as RNA, such as mRNA, inside of the droplets that contained the target cells.

[0100] After lysing target cells inside the droplets, mRNA is released to interact with a bead or other structure in the partition (see FIG. 1 1). The mRNA may be used to create a sequencing library after reverse transcription and amplification. In the illustratedembodiment, a template particle is linked to a capture oligonucleotide. As shown, the template particle is linked to (among other things) mRNA capture oligonucleotides that include a 3' poly-T region (although sequence-specific primers or random N-mers may be used). After lysis, the capture oligonucleotide hybridizes by Watson-Crick basepairing to a target in the RNA and serves as a primer for reverse transcriptase, which makes a cDNA copy of the RNA.

[0101] As shown, each of the mRNA capture oligonucleotides may include, from a PEI sequence, a cell or droplet barcode, and a poly-T segment at a 3’ end. The capture oligos may also include a binning index in certain implementations. As shown, the capture oligonucleotide hybridizes to the mRNA.

[0102] In certain implementations of the workflow if FIG. 9, the partitions may be disrupted (e.g., the emulsification may be broken) such that all of the beads with hybridized mRNAs are combined in a reaction mixture for reverse transcription. Because the mRNAs are hybridized, the mRNAs are still associated with their cellspecific beads. A reverse transcriptase binds and initiates synthesis of a cDNA copy of the mRNA to make an RNA / DNA hybrid. The cDNA copy extends from the capture oligonucleotide and includes the cell barcode However, in other implementations, the reverse transcription is performed within the partitions and the partitions are disrupted for the subsequent amplification step.

[0103] Note that the mRNA is connected to the particle non-covalently, by complementary base-pairing. The cDNA that is synthesized may be covalently linked to the particle by virtue of the phosphodiester bonds formed by the reverse transcriptase.

[0104] All capture oligonucleotides on one bead in the partition may have a common barcode, which can serve as a cellular barcode (e.g., unique identifier, unique partition identifier, unique cellular identifier). The barcode may be 6-30 nucleotides in length in an embodiment. The barcodes for each partition are sequentially distinct from one another. The barcode is incorporated via reverse transcription.

[0105] After reverse transcription, an amplification step is used that includes an extended forward primer that fully integrates p5, index, and PEI sequences. Thisallows for single index differentiation of multiple samples that may be combined for on-cartridge library preparation. That is. the amplification products, after amplification, include the first adaptor sequence (P5, index sequence) and its complement on one end as well as the cell barcode, cDNA of interest and both paired- end sequencing primers (PEI, PE2). The paired-end sequencing primers may be any Illumina or other sequencing platform sequences. For example, the PEI and PE2 may be read 1 and read 2 primers from Illumina. Thus, each amplification product can include a cell-specific barcode and a sample-specific index.

[0106] The amplification may be a limited amplification (e.g., 10 or fewer cycles, 5 or fewer cycles) in an embodiment.

[0107] As shown in FIG. 12, amplified DNA is introduced to a compatible flow cell. A small DNA retain may be re-amplified for quality control of cDNA fragment length. The introduced DNA is processed on flow cell with tethered transposome complex. The transposome complex is a one-way or direction complex that incorporates a second adaptor sequence via the transposase, e.g., TN5. The first adaptor sequence is incorporated in the amplification step (FIG. 11), and the second adaptor sequence is incorporated via the transposome complex. Thus, the transposome complex can be a single type or a homodimer as opposed to workflows in which both adaptors are added via transposome.

[0108] In certain embodiments, only a single index is used rather than dualindexing. Thus, the transposome complexes need not carry a separate index, which may be more efficient.

[0109] The transposon accommodates directional library preparation such that only fragments with intact cell barcoding information may be amplified on flow cell. A given flow cell may accommodate a maximum number of paired end reads. This may be apportioned to a desired number of target reads per cell loaded. Experimentally, numbers of cells may be apportioned into multiple sample preparations that may be identified by included sample index. Sequencing efficiency may be calibrated to cells with high transcript content (cancer cell lines for example). Low content cells in hiscase may result in fewer overall reads per cell, but the representation of transcripts form cells may be normalized in a given experiment.

[0110] These workflows are compatible with magnetic beads, and may provide an opportunity for direct to flowcell library loading with automation. On flowcell tagmentation provides the random fragmentation required for IMI based analysis as disclosed herein. This may have competing advantages and disadvantages regarding efficiency in fragment cutting and diversity of cut sites. On flow cell processing eliminates library PCR amplification and may reduce needs for molecular deduplication processes.

[0111] The disclosed workflows may be used to estimate gene expression levels by' counting aligned reads for an individual cell. That is, sequencing reads sharing a cell barcode and that align to a reference sequence may provide indications of overall expression.

[0112] Certain workflow steps for library7preparation may be completed relatively rapidly (e.g., in the order of 1-5 hours) and eliminate certain quality checks, thus permitting more efficient transcriptome analysis.

[0113] Any one of the above-described strategies and methods, or combinations thereof, may be used in conjunction with particle-templated emulsions. For example, methods may be used for single cell expression profiling, which may include combining target cells with a plurality7of template particles in a first fluid to provide a mixture in a reaction tube. The mixture may be incubated to allow association of the plurality7of the template particles with target cells. A portion of the plurality of template particles may become associated with the target cells. The mixture is then combined with a second fluid which is immiscible with the first fluid. The fluid and the mixture are then sheared so that a plurality7of monodisperse droplets is generated within the reaction tube. The monodisperse droplets generated comprise (i) at least a portion of the mixture, (ii) a single template particle, and (iii) a single target particle. Of note, in practicing methods of the techniques described herein, a substantial number of the monodisperse droplets generated will comprise a single template particle and a single target particle,however, in some instances, a portion of the monodisperse droplets may comprise none or more than one template particle or target cell.

[0114] In some aspects, generating the template parti cles-based monodisperse droplets involves shearing two liquid phases. The mixture is the aqueous phase and, in some embodiments, comprises reagents selected from, for example, buffers, salts, lytic enzymes (e.g., proteinase k) and / or other lytic reagents (e.g., Triton X-100, Tween-20, IGEPAL, bm 135, or combinations thereof), nucleic acid synthesis reagents (e.g., nucleic acid amplification reagents or reverse transcription mix), or combinations thereof. The fluid is the continuous phase and may be an immiscible oil such as fluorocarbon oil, a silicone oil, or a hydrocarbon oil, or a combination thereof. In some embodiments, the fluid may comprise reagents such as surfactants (e.g., octylphenol ethoxylate and / or octylphenoxypolyethoxyethanol), reducing agents (e.g., DTT, beta mercaptoethanol, or combinations thereof).

[0115] Some methods of the presently described techniques use oligonucleotides. Oligonucleotides, sometimes referred to as oligos, are sequences of contiguous nucleotides of DNA, RNA, or a mixture thereof. In certain implementations oligonucleotides comprise DNA. However, in other embodiments, oligonucleotides may comprise RNA. In additional embodiments, oligonucleotides may comprise a mixture of DNA and RNA. Oligonucleotides may comprise noncanonical nucleotides, such as, synthetic nucleotides that have been modified to incorporate certain biomolecular properties. The length of the oligonucleotide is usually denoted by ”- mer”. For example, an oligonucleotide of six nucleotides is a hexamer, or 6-mer, while one of 25 nucleotides may be referred to as a 25-mer. An oligonucleotide may include other features such one or more conformationally-restricted nucleic acids or a locked nucleic acid (LNA) bases or phosphorothioate inter-base linkages, to improve binding stability or residence times.

[0116] In particular, and with reference to FIG. 4, certain implementations of the presently described techniques are useful to create sequence libraries. Those libraries may be sequenced 127 to identify transcript abundances, or gene expression levels, of single cells. In general, the product of amplification step 123 may be understood asproducing a sequencing library. However, depending on the sequencing technology being used (e.g.. single- molecule long-read sequencing versus short-read ensemble sequencing), the distribution of steps, and the desired storage times or conditions for certain library preparation products, the product of attachment 115 may also or instead be considered a sequencing library'. Also, a sequencing library produced by amplification 123 may be subject to further rounds of amplification (e.g.. at the discretion of a user and / or after being shipped to a different location). For example, RNA capture, cDNA synthesis, and a first round of amplification may be performed at a research or clinical services laboratory' to create a sequencing library', which may be stored in a tube such as a microcentrifuge tube. The sequencing library may be shipped (e.g., on dry ice) to a genomics core facility for sequencing. The genomics core facility may provide sequence data via a server or data room. The research or clinical services laboratory or another party7may access the sequence data to initiate mapping and / or deduplication, which may occur in an online server, in the cloud, or on a local computer. In general, a sequencing library includes DNA copies of target nucleic acids from a sample of interest with PCR handles or adaptors attached at ends. The amplicons may be stored, for example, at -20 degrees Celsius, or may be analyzed. Analyzing amplicons may involve sequencing. The sequencing library may be sequenced 127.

[0117] Sequencing 127 may be performed by any suitable method. An example of a sequencing technology7that can be used is Illumina sequencing. Illumina sequencing is based on the amplification of DNA on a solid surface using fold-back PCR and anchored primers. Genomic DNA is fragmented and attached to the surface of flow cell channels. Four fluorophore-labeled, reversibly terminating nucleotides are used to perform sequencing. After nucleotide incorporation, a laser is used to excite the fluorophores, and an image is captured, and the identity7of the first base is recorded. Sequencing according to this technology is described in U.S. Pub. 2011 / 0009278, U.S. Pub. 2007 / 0114362. U.S. Pub. 2006 / 0024681, U.S. Pub. 2006 / 0292611, U.S. Pat. 7,960,120, U.S. Pat. 7,835,871, U.S. Pat. 7,232,656, U.S. Pat. 7,598,035, U.S. Pat. 6,306,597, U.S. Pat. 6,210,891, U.S. Pat. 6,828,100, U.S. Pat. 6,833,246, and U.S. Pat. 6,911,345, each incorporated by reference. In certain embodiments, an Illumina Mi- Seq sequencer may be used.

[0118] Sequencing 127 creates sequence reads, i.e., a record of a sequence of bases from at least a part of a nucleic acid strand. The sequence reads may be analyzed to determine expression of RNA associated with genes based on unique reads that correspond to those genes. Analyzing the sequence reads may be performed using software and following multistep procedures that are known in the art. For example, first, the quality of each sequence read, i.e., FASTQ sequence, may be assessed using the software FASTQC. Next, the reads may be trimmed using, for example, using Trimmomatic software. See Bolger, 2014, Trimmomatic: a flexible trimmer for Illumina sequence data, Bioinformatics 30(15):2114-2120, incorporated by reference. The trimmed sequence reads may then be mapped to a human genome using, for example. HISAT2 software. HISAT2 output files in a SAM (sequence alignment / map format), which may be compressed to binary sequence alignment / map files. Other methods useful for processing and analyzing sequence reads are discussed in U.S. Pat. No. 8,209,130, which is incorporated by reference. Determining gene expression generally involves counting numbers of unique sequence reads that uniquely map to a human reference genome. Mapping reads to a reference to identify genes may be performed using computer software packages known in the art.

[0119] An aspect of the presently described techniques is that mapping reads to a reference and identifying genes gives a quantitative result when reads are deduplicated by IMI to yield one read per mRNA from which those reads originated.

[0120] Because each mRNA is typically copied into cDNA and each cDNA is typically copied into an unpredictably large number of amplicons in the sequencing library, and because each library member is often amplified or read redundantly as part of a sequencing technique, a number of raw sequence reads does not necessarily correlate to numbers of input molecules from the single cells. Nevertheless, one cell may include abundant transcripts that map to one gene. Here, compositions and methods of the presently described techniques give each cDNA a unique intrinsic identifier that can be identified and used to deduplicate sequence reads. After sequence reads are identified by gene and deduplicated, counts of those reads are associated with their binning indexes.

[0121] The disclosed methods include performing nucleic acid sequencing to obtain sequence information and quantitation of mRNA in the sample. According to such approaches, sample preparation and sequencing are performed so that the sequence information for each mRNA (or its corresponding cDNA) is read from what is essentially a random start site within the molecule. As used herein, such a random start site may be understood to be “random” with respect to a whole or larger sequence corresponding to or derived from an mRNA or cDNA strand or reference. That is, a “random start site” in this context may be understood to not necessarily correspond to a sequencing operation that starts at a random location along an intact strand but also to a sequencing operation that starts at an end of a fragment of a larger or whole strand that has undergone random fragmentation as part of the technique (e.g., via enzymatic fragmentation) or via a naturally occurring process (e.g., damage associated with a disease or disorder causing fragmentation or cleavage of a larger strand or sequence and which may be indicative of spatial information related to the sample or subject). Thus, as used herein a “random start site” does not necessarily imply that the initiation of the sequencing operation occurs randomly with respect to a whole or intact sequence, but that the whole or intact sequence may undergo some form of naturally occurring or induced fragmentation at random locations prior to sequencing such that the fragments, when sequenced, are effectively random in terms of their start location with respect to the whole or intact sequence.

[0122] There are some sequencing techniques that involve making an abundant number of copies of nucleic acid and generating a corresponding number of sequence reads from those copies. Using embodiments of the presently disclosed techniques, each sequence read includes an intrinsic molecular identifier that associates the read with one original molecule. The intrinsic molecular identifier is unique because it is adjacent a random cleavage or break site within the original molecule or a complementary copy of the original molecule and proximate to sequencing initiation (hence, as used herein, a “random start site”). Thus, even when a sample has numerous transcripts that would otherwise appear to be exact duplicates of each other, and would otherwise generate identical sequence reads, by starting the sequencing at random start sites for each original molecule (such as in based on random cleavage or breakage of with respect toan original or reference molecule), the reads from otherwise identical transcripts are unique at least because those reads began at different places within those transcripts. Thus, some portion of each sequence read (e.g., the first 10 to 20 bases or so) may be used to identify the transcript, from which the read was derived. The identify ing portion of the sequence read, and the corresponding portion of the molecule from which the sequence read came, is referred to as an intrinsic molecular identifier. After obtaining sequence information by methods and techniques as discussed herein, any portions of the sequence information (e.g., any sequence reads) that are identical and include the same intrinsic molecular identifier, are considered duplicate sequences taken from the same original molecule, or transcript, in the sample. A count of unique (i.e., deduplicated) sequences from the sample provides a count of the molecules present in the sample. For example, if there are 10,000 transcripts from gene 1 and 7 transcripts from gene 2, even after exponential amplification, those likely millions of sequence reads are deduplicated using the intrinsic molecular identifiers to yield 10,000 deduplicated sequence reads of gene 1 and 7 reads of gene 2.

[0123] Methods of the presently described techniques further add a small segment of extrinsic bases, the aforementioned “binning index”, to the molecules during preparation for sequencing. The added bases appear in the sequence information and are used as an index. Such binning indexes may be useful when the sequence reads are mapped to reference information to identify genes and assigned to bins according to gene, intrinsic molecular identifiers, and other optional barcode information such as cell-specific “cellular barcode” that may be used when methods of the presently described techniques are applied to single-cell RNA sequencing (scRNA-Seq). The binning index, which refers to both the small segment of bases added to each molecule during sample preparation and also the corresponding segment of base information in each sequence read, is a useful informational tool for assigning counts of deduplicated sequence reads to bins and correcting those counts to adjust for bias that may arise during sample preparation. In an embodiment, the binding index may be introduced to a bead-bound duplexes (extended capture oligonucleotide hybridized to cell nucleic acid) via a transposition step that can cut at one or more random sites on the bead-bound duplex to simultaneously 1) add a PCR handle (e.g.. adapterization) and 2) introduce arandom cut end having a random and unique sequence that forms the IMI. Thus, the surface-bound transposition in one step can efficiently add a universal sequence as part of a sequencing workflow as well as introduce a variable sequence that serves as a unique identifier for the sequenced molecule. This improves sequencing efficiency by eliminating complex indexing on a per-molecule basis via a separate step using an exogenous indexing molecule. Further, introduction of different indexes via distributed transposome complexes may be complex. In the disclosed arrangement, all or most of the transposome complexes can carry a universal transposon sequence or set of sequences (e.g., in a heterodimer with one transposase being associated with a first universal nucleic acid and a second transposase being associated with a second universal nucleic acid). Transposase complexes may be as generally discussed with respect to FIGS. 15-18 by way of example.

[0124] For example, if target molecules undergo limited amplification (e.g., three or four rounds of polymerase chain reaction) prior to fragmentation, there may be overrepresentation of those molecules in sequence reads counts. In such cases, the counts can be divided by a correction factor proportional to the expected amplification by PCR. In another example, if a transcript is short or very highly expressed (e.g., millions of copies in a cell), then even random fragmentation will generate some identical cut sites, yielding some limited number of identical intrinsic molecular identifiers. Those duplicate intrinsic molecular identifiers will lead to under-representation of the transcripts in the sequence read counts. In those instances, the sequence read counts can be multiplied by a correction factor (e.g., that has been derived experimentally) to provide an accurate measure of expression levels in a cell. Thus, the binning index, which is typically provided by a capture oligonucleotide or adaptor added and used during sample preparation, in combination with the intrinsic molecular identifiers, provide for the accurate measurement of expression levels in cells.

[0125] A transcript may further be labeled with a short oligonucleotide tag referred to herein as a molecular diversity7enhancer (MDE). An MDE may be two, three, four, or so bases and may be used to ensure unduplicated intrinsic molecular identifiers (IMIs). In practice the MDE may not be. by itself, long enough to function as a molecule-specific barcode or unique molecular identifier. That role is performed by theIMI and the MDE may be added to supplement the information of the IMI, such as in the event the IMI is based on a non-unique identifier due to a fragmentation or breakage site shared by two original molecules, and thereby a shared sequencing start site. As used herein, the MDE (which is optional) supplements the IMI to ensure that each molecule is uniquely labeled. The binning index plays a different role and is introduced as a molecule and then used during bioinformatics to characterize or sort sequence read counts into bins, where each bin may be specifically associated with a cell (via cell barcode introduced during sample preparation), a gene (identified by mapping sequence reads to reference information), and a molecule (shown the IMI and optional MDE). Read counts are collected based on binning index (i.e., bins), allowing those counts to be corrected to adjust for bias that may be introduced during sample preparation. By those means, methods of the presently described techniques are useful for quantifying expression levels of single cells, including for single cells that have been isolated, such as in aqueous partitions (e.g., droplets or wells of a plate).

[0126] With respect to the correction of read counts as discussed herein, in certain embodiments a correction approach is taught which is dynamic in nature in that the correction factor derived is sample specific (e.g., based upon cell type, level of gene expression, sequencing depth, and / or other factors specific to a given run or sample). In certain described implementations, the correction is based at least in part on, or otherwise corrects for, inflation attributable to whole transcriptome amplification (WTA). By way of example, a correction factor determination routine, in one embodiment may be used to derive a correction factor that may be used to adjust the observed molecule counts within a given experimental dataset (i.e.. for a respective sample), thus adjusting counts dynamically based upon the observed level of WTA- based inflation (e.g., amplification-based inflation). In one such embodiment, the analytical workflow begins with read mapping to identify genes or other chromosome regions of interest and to enable all reads to be segregated by relevant identifier (e.g., gene identity) in an output file (e.g., a BAM format file). In such an embodiment, within the same cell barcode and gene, the output file is parsed to identify IMIs without regard for the BI. IMIs are segregated into a number of bins (e.g., 64 bins) determined based on the length of the BI sequence (e.g., 3-bases) in each read. Within each bin. thenumber of unique IMIs is determined by collapsing duplicate IMIs into a single count. This step mitigates error attributable to obtaining identical IMIs from distinct transcripts from the same gene as well as improving computational and process efficiency.

[0127] After obtaining counts of unique IMIs in each bin, a subset of the barcode+gene combinations (BGCs) may be used to estimate the inflation distribution, i.e., the probabilities of a single parent molecule producing any number of copies from one to x during WTA. This is done by selecting a subset of BGCs in which the number of unique IMIs is lower than or equal to x (e.g., < 15 in one embodiment). In certain embodiments a routine may be employed to derive the inflation probabilities based on observed IMI counts. BGCs where the number of molecules can be determined unequivocally may be used to estimate the true number of IMIs that can be attributed to inflation from a single molecule in a single bin. Finally, each of the x counts in single bins is divided by the sum of counts to produce the probability' of one molecule producing this many IMIs.

[0128] Once the inflation probabilities are established, a correction factor can be selected. In one implementation, the correction factor is an integer and may be chosen such that no more than 1% of molecules will be over-counted, i.e., the cumulative inflation probability is at least 0.99. Within each BGC, the number of unique IMIs in each bin is divided by the correction factor and rounded up to the nearest integer. The divided counts are then summed across bins to produce the final count for this BGC, which may be written into the raw count matrix.

[0129] In certain aspects, the embodiments of the presently described techniques relate to methods for measuring gene expression, including at the single cell level of granularity. Certain of the described methods include sequencing nucleic acids (e.g., mRNA or cDNA) from random start sites (e.g.. after random fragmentation by chemical or enzymatic agents or after natural breakage due to a disease or disorder) of the nucleic acids to generate sequence reads having an effectively unique portion (e.g., an intrinsic molecular identifier (IMI)), attaching a binning index (BI) to the sequence reads, and mapping each sequence read to a gene or genomic region. Methods include determining counts of the unique portions per gene or genomic region, assigning the counts toassociated binning indexes, and correcting counts to reduce bias introduced during sample preparation. By summing corrected counts across the binning indexes for each gene or genomic region, methods of the presently described techniques provide an estimated number of the transcripts per gene or genomic region in the sample, and correspondingly one or more measures of gene expression for the sample, including in some implementations measures of gene expression at the single cell level of granularity.

[0130] The binning index may comprise, in certain embodiments, six or fewer bases, and in certain implementations comprises three or fewer bases. The sample preparation in certain such embodiments may include fragmenting mRNA at random sites (such as by the application of chemical or enzymatic agents), annealing oligonucleotides to the fragments, and extending the oligonucleotides to make cDNA. In other embodiments, the sample preparation includes annealing oligonucleotides to the mRNA, extending the oligonucleotides to make cDNA copies of the transcripts, and fragmenting the cDNA copies at random sites. The correcting step may be performed to account for a probability of the random sites, from which sequencing is initiated, being duplicated among the transcripts. In this manner, the correction factor may account for a probability of multiple random start sites per transcript within the sample.

[0131] In certain embodiments, cDNA is prepared by capturing the transcripts with capture oligonucleotides linked to beads (e.g., capture beads), wherein each capture oligonucleotide (also referred to as an "oligo" herein) includes a 5'-linkage to a respective bead, a cell barcode, a binning index, and an annealing primer section-3'. The annealing primer section may include a poly-T region and random segment (e.g., a hexamer), or a gene-specific primer.

[0132] In certain implementations, the binning index is variable in length to improve sequencing quality. For example, a bead decorated with capture oligonucleotides may have a mixture of binning indexes with some being 3 bases, some 2, some only one, and some oligonucleotides having no binning index (i.e., zero bases long). In some embodiments the capture oligonucleotides have a conserved sequence 3' of the mixed- length binning indexes, which may be useful to identify the start sites in the moleculesin the sequence read data. For such embodiments, a first portion of the capture oligonucleotides linked to the beads include no binning index and a second portion of the capture oligonucleotides linked to the beads each include a binning index that each independently consists of 1, 2, or 3 bases. Due to the variable length binning indexes, sequencing library material (e.g., amplicons) attached to the flow cell of a sequencer may be sequenced out of phase with each other. Methods of the presently described techniques improve the quality of sequence data by avoiding problems of sequencing conserved molecules “in phase”, or in lock-step with each other.

[0133] More generally, the disclosure provides methods of counting molecules present in a sample so as to reduce or eliminate duplicates. The disclosure makes use of intrinsic sequences present near random fragmentation or priming sites to identify individual molecules. The disclosure further provides a binning index that is added during sample preparation and is used during read deduplication and counting as part of a method of correcting for biases that arise during sample preparation. According to methods of the disclosure, an individual molecule is randomly fragmented, and the information encoded in the first N bases of the sequence proximate or adjacent the fragmentation site encodes information about both the gene identity and the unique fragmentation position within that sequence. This provides intrinsic molecular identification (i.e., the identifying information is intrinsic to the original or natural sequence itself). In certain embodiments, it may be useful to also append a short randomer (e.g., NNN) at the cut end of the molecule. This serves as a “molecular diversity enhancer” (MDE) and effectively expands the number of potential molecules that may be resolved for a given gene. The MDE addresses potential concerns regarding undercounting due to limited cut site diversify by increasing the available IMI space.

[0134] As a tool for read counting and correction, the disclosure provides for the addition of a diversify of short (e g., N < 6 bases) random sequences that may be introduced with the cell barcode sequence, e.g., among the capture oligonucleotides decorating a bead such as may be used in scRNA-Seq applications. This “binning index” is distinct from a UMI as conventionally described and implemented. In one embodiment, the binning indexes are 3 or fewer bases in length and comprise fewerthan 100 unique identities. After sequencing, sequence reads are grouped by cell barcode, by gene to which the read was mapped, and by IMI (possibly augmented by an MDE).

[0135] Each cell barcode, gene, and IMI combination is associated with a number of reads, and with one or more different binning indices, one coming from each read. To resolve exact PCR duplicates (identical sequences), all the reads from this combination are treated as a single molecular count, and that count of reads are associated with a single binning index. If multiple binning indexes are associated with such a combination, one index may be chosen arbitrarily.

[0136] To illustrate the correction, the binning index may be used as part of a method to resolve multiple fragments derived from the same transcript during whole transcriptome amplification (WTA) and fragmented at different cut sites. For each binning index, a system will count the number of molecules assigned to that binning index after the deduplication step, and divide that number by a correction factor. The correction factor may be the “worst-case scenario”, based on the number of WTA cycles. For example, with 4 WTA cycles the “worst-case” factor would be 7. The system may sum the corrected counts across binning indices to get the final count for the associated cell barcode and gene. Additional aspects of an implementation of a suitable correction are discussed herein in greater detail.

[0137] The presently described techniques are useful for creating sequencing libraries that can be sequenced to quantify molecules such as mRNA transcripts in a sample, such as the mRNAs of a single cell. Those molecules can be quantified in accordance with the presently described approaches due to library preparation methods described herein that provide each molecule with a binning index and an effectively unique, intrinsic identifier, a sequence within the molecule that could be referred to as an intrinsic molecular identifier (IMI), and optionally with an MDE.

[0138] As discussed herein, the molecular identifier is intrinsic in that it comprises bases that are copied from the genetic material being studied, as opposed to extrinsically added bases that are not initially part of the genetic material being studied. For example,where single-cell RNA-seq (scRNA-Seq) is being performed to quantify messenger RNA (mRNA) transcripts present in a single cell, those mRNA transcripts are copied into cDNAs and a segment of, or sequence of bases from, each cDNA is used as the intrinsic molecular identifier. The sequence of bases is intrinsic in that the sequence originates as part of the genome of the organism (or virus, or other biological source material) and is produced as a cut site in the cDNA. In accordance with aspects of the presently described techniques, the molecule identifier is useful to identify each cDNA because each cDNA has an intrinsic molecular identifier that is “unique”, “nearly unique”, or “essentially unique”. One important feature is that, across all RNA molecules from a cell, those RNA molecules are copied into cDNA molecules that can be mapped to their genes of origin and in which substantially most of the cDNA molecules have a unique intrinsic molecular identifier.

[0139] That level of essentially unique labeling is achieved by cleaving each cDNA molecule at a random site and attaching a PCR handle to the random site or by priming sample nucleic acids at random locations using, e.g., random hexamers, where the primers include a 5' tail with a PCR handle. Typically, the cDNA molecule will have a first PCR handle that has been provided as part of a capture oligonucleotide that annealed, or hybridized, to the mRNA. The capture oligonucleotide, which includes the binning index, is extended by a polymerase, copying the mRNA to form the cDNA. The cDNA is then cleaved at a random cut site, and a second PCR handle is attached at the random cut site. Because the cut site or priming site is random, a segment of the cDNA adjacent to that site will include a sequence of bases that is effectively unique for that cDNA molecule.

[0140] The cDNA molecules can be amplified from the PCR handles and the amplicons can be sequenced. Sequence reads that are generated by sequencing into the cDNA from the random site will include the binning index and a sequence of bases unique to that molecule, i.e., the intrinsic molecular identifier. Sequence reads can be deduplicated and / or mapped to a reference (e.g., a human genome or a gene atlas) to identify genes. After deduplication and mapping, a count of de-duplicated (or unique) reads mapping to each gene is associated with the binning index. The counts may be corrected by a correction factor, and the corrected counts provide a measure oftranscripts of that gene from that cell. Thus, the corrected counts provide a measure of expression levels for the cell.

[0141] Specifically, reads with identical gene-mapping and identical IMIs can be “collapsed'’ to produce a count of only unduplicated reads that may be used as a quantitative measure of gene transcripts in the sample, i.e., the single cell.

[0142] As shown in FIG. 7, the use of IMIs is compatible with RNA capture without necessarily requiring any bead-linked capture oligonucleotides. That is, capture oligonucleotides may be free in solution (as opposed to linked to a solid support). Aspects of the presently described techniques are also compatible with the use of capture oligonucleotides that are linked to a solid support such as a bead.

[0143] As implemented, sequencing reads are clustered in the following order: by cell barcode; by the gene to which the read was mapped; and by IMI (possibly augmented by an MDE).

[0144] Each cell barcode, gene and IMI combination is associated with a number of reads, and with one or more different binning indices, one coming from each read. To resolve exact PCR duplicates (identical sequences), all reads from this combination are regarded as a single molecular count, and the count is associated with a single binning index. If multiple binning indexes are associated with such a combination, one index is chosen arbitrarily.

[0145] In the next step, the system may resolve multiple fragments derived from the same transcript during WTA and fragmented at different cut sites. For each binning index, the system may count the number of molecules assigned to that binning index in the previous step, and divide, multiply, shift, or scale that number by a correction factor, as discussed in greater detail below. This correction factor may be the “worst-case scenario’', based on the number of WTA cycles. For example, with 4 WTA cycles the “worst-case” factor would be 7. The system may sum the corrected counts across binning indices to get the final count for the cell barcode and gene.

[0146] In a further implementation, binning tag sequences may be utilized for instrument phasing. Methods have been implemented with a 0. 1, 2, or 3 base stagger in the capture oligonucleotide, optionally embodied within the binning indexes, as a tool to disrupt alignment of conserved sequences in sequencing. This is performed to improve color balance and avoid loss of sequencer registry' on certain sequencing instruments. In one implementation of the binning tag strategy, the binning indexes have 0, 1, 2, or 3 N bases (it is understood that having zero bases means that the binning index is not there, which is what is intended here: as a set, the capture oligonucleotides have a mixture of different binning index lengths). This results in a diversity of 85 potential bins without addition of any additional sequenced bases.

[0147] FIG. 13 illustrates a workflow within a system configured to perform embodiments of the presently described techniques. In one such implementation the system brings in sequence reads, e.g., as a FASTQ file from a sequencing instrument. The system deduplicates the reads by IMI (and optionally MDE) to obtain counts 605. Read counts are indexed 116, i.e., associated with their binning indexes, such as in a read count file 607, which is written to tangible, non-transitory memory. For each binning index, read counts are summed together to provide binning index read counts 609. The read counts are corrected to provided corrected read counts 611. In one embodiment, read counts 609 are divided by 7 and rounded up to provided corrected read counts 611, but other corrections, such as may be based on an estimate of WTA- based inflation as discussed herein (or more generally, “amplification-based inflation”), are within the scope of the disclosure.

[0148] In particular, with respect to correction of read counts and to the determination of a correction factor, certain implementations and explanations are further provided by way of example. By way of context, to evaluate the benefits of such a dynamic correction approach, UMI-containing core particles (UCPs) were used to perform PIP-seq experiments using multiple sample input types for direct comparison of UMI-based and IMI -based molecule counts for quantitative estimation of WTA-based inflation (e.g., amplification-based inflation). Across all sample types, when evaluating the observed distribution of IMIs per UMI, there were fewer instances of 15 unique IMIs than would be expected assuming 100% WTA efficiency of a 5-cyclePCR. Such data demonstrate that WTA-based inflation varies across sample ty pes, as is expected for samples of vary ing intrinsic RNA content. With this in mind, attempts to correct molecule counts based upon an assumption of 100% WTA efficiency would result in significant deflation of the molecule counts, which will vary across samples and experiments. This analysis suggests that a dynamic correction factor that accounts for the observed variability in the IMI count distribution could be used to adjust the observed molecule counts within a sample to be equivalent to UMI-based counts. With this in mind, presently contemplated automated routines are described that may be used in determining such a correction factor that may be used to adjust the observed molecule counts within a given experimental dataset (e.g., a sample), thereby adjusting counts dynamically based upon the observed level of WTA-based inflation.

[0149] Turning to FIG. 14, an example analytical workflow begins with read mapping to identify genes and to enable all reads to be segregated by gene identity in an output file in BAM format. Within the same cell barcode and gene, the BAM file is parsed to identify IMIs without regard for the BI. The IMI sequences are corrected for sequencing errors by merging IMIs within a Hamming distance of 1. When such similarity is detected, the less frequent IMI is merged into the more frequent IMI in certain implementations. Reads with IMIs that contain anbase may be discarded. Next, IMIs are segregated into 64 bins, based on the 3-base BI sequence in each read. Within each bin, the number of unique IMIs is determined by collapsing duplicate IMIs into a single count. This step mitigates error arising from obtaining identical IMIs from distinct transcripts from the same gene, owing to a combination of limited-cycle PCR amplification and subsequent fragmentation.

[0150] After obtaining counts of unique IMIs in each bin, a subset of the barcode+gene combinations (BGCs) is used to estimate the inflation distribution, i.e., the probabilities of a single parent molecule producing any number of copies from one to fifteen during WTA. This is done by' selecting a subset of BGCs in which the number of unique IMIs is lower than or equal to fifteen. Some IMI configurations definitively point to a certain number of molecules. For example, for a given cell barcode, three unique IMIs observed across three bins can only arise from three unique molecules. In other cases, such as three unique IMIs observed in two bins, the number of originalmolecules is ambiguous, since multiple IMIs in the same bin can represent distinct molecules or copies of the same molecule fragmented at different locations. This approach assumes that IMIs from the same parent molecule will have the same BGC. As discussed below, this or another algorithm may use observed IMI counts to derive the inflation probabilities, as discussed herein. The BGCs where the number of molecules can be determined unequivocally are used to estimate the true number of IMIs that can be attributed to inflation from a single molecule in a single bin. Finally, each of the fifteen counts in single bins is divided by the sum of counts to produce the probability of one molecule producing this many IMIs.

[0151] After establishing the inflation probabilities, a correction factor is selected. In one embodiment, this is an integer, chosen such that no more than 1% of molecules will be over-counted; the cumulative inflation probability is at least 0.99. Within each BGC, the number of unique IMIs in each bin is divided by the correction factor and rounded up to the nearest integer. The divided counts are then summed across bins to produce the final count for this BGC, which is written into the raw count matrix.

[0152] By way of further example and explanation, further description of a process for determining inflation probabilities from counts of IMIs, and the corresponding derivation of a correction factor based on the inflation probabilities is provided. In particular, as discussed herein a correction factor may be selected by estimating the IMI inflation probability from a single molecule using observed IMIs and binning indexes with the same barcode+gene combinations (BGCs). The magnitude of IMI inflation for a given dataset depends on factors such as cell type and sequencing depth. However, the accuracy of the described algorithm is independent of IMI inflation because it is based on the combinatorics of the binning indexes (e.g., 64 binning indexes). Across multiple cell types and sequencing depths, the described algorithm produces accurate estimates of IMI inflation. In practice, inflation probability estimates may be most accurate for low numbers of inflated IMIs, where IMIs are more likely to have distinct binning indexes. For example, when estimating the probability' of a molecule having four inflated IMIs, the result is based on the initial known observations of four uninflated IMIs which each have different binning indexes. Four molecules will be distributed in different bins -91% of the time, so this estimate will be highly accuratebecause the majority of the data is observed. With this in mind, the described algorithm offers a robust, computationally-efficient, and data-driven approach for correcting amplification bias in scRNAseq data, independent of cell type or sequencing depth.

[0153] With this in mind, and turning to FIGS. 15 and 16 (wherein FIG. 15 depicts a process flow diagram and FIG. 16 depicts a corresponding process flow in the context of sample data representations), further aspects of this approach as they pertain to estimation of inflation probabilities are described. In particular, as discussed herein, the presently contemplated algorithmic estimation of inflation probability estimates inflation from a single parent cDNA molecule using the observed counts (shown here as count data 800) for each distribution of IMIs in bins, from 1 IMI to the maximum possible number (threshold 812) of IMIs that can be obtained from one parent molecule (e.g., 15 for 5 WTA cycles) (step 808). In embodiments where the binning index consists of 3 bases, there are 64 (43) possible bins. For each number of IMIs and each distribution of IMIs (step 816), all potential distributions (shown as remaining counts 820) of parent molecules that can create the observed IMI distribution are assessed. First, the known uninflated distribution 804 is considered, where each IMI is in its own bin and the number of parent molecules (m) is equal to the number of IMIs ( / ). The total pool of i IMIs derived from m molecules where m = i (shown as estimated total counts 832) is estimated (step 824) by dividing (step 828) the known observations by the expected fraction of m molecules as follows: For each number of molecules m, let Tmbe the number of instances of uninflated m molecules, and Omthe number observed uninflated m molecule in m bins. I'm can be derived by:(1)For example:> 64202> 6420264O2?2 -~63For every other possible distribution of m molecules (in fewer than m bins), the estimated number of occurrences (estimated counts 840) of the distribution is calculated by multiplying (step 836) Tmby the probability of getting that molecule distribution when m molecules are distributed among 64 bins. Then, the expected number of observations of each distribution of i IMIs that can be created by this molecule distribution is taken by proportionally splitting the estimated occurrences of the molecule distribution by the probability of getting each IMI distribution from the molecule distribution (the product of inflation probabilities from each molecule to each IMI for each configuration of i IMIs from this distribution of m molecules). For the first pass where z equals m. the molecule distribution is equivalent to the IMI distribution, so there is only one possibility. For subsequent values of m, there are i-m IMIs which are inflated, and the total pool of molecules Tmwhich has i-m inflated IMIs is estimated as before. The estimated occurrences are subtracted from the remaining observed counts 820 for each IMI distribution in order to obtain the estimated (uninflated) IMI counts for each of the 15 IMI distributions.

[0154] Subtracting the estimated occurrences from m molecules for each configuration of i IMIs allows estimation of the occurrences of the same configuration from m-1 molecules, since the remaining counts for each distribution of i IMIs no longer include observations from i to m molecules. For example, the distribution of z IMIs with three IMIs in one bin and two IMIs in another bin (Is, 2) can result from at least two molecules (three IMIs are inflated) and at most five molecules (uninflated). In the first pass, where i=5 and m=5, the estimated occurrences of Is, 2 from m molecules is calculated by obtaining T5 from the distribution of five IMIs in five bins (05). 15 is then multiplied by the probability of the getting the molecule distribution Ms, 2: 0.0000376when 5 molecules are distributed among 64 bins. Since the molecules are uninflated, this gives the estimated occurrences of Is, 2 from five molecules. This estimate is then subtracted from the observed counts of Is, 2. In the next pass, where i=5 and m=4, T4 with one inflated IMI is estimated by using the remaining counts of the IMI distribution h, 1,1,1 with five IMIs in four bins (04). The remaining counts for h, 1,1,1 now only contain the estimated occurrences from four molecules, since estimated occurrences from five molecules were subtracted in the previous pass. The possible molecule distributions which can produce Is, 2 include M34 and M2, 2. The estimated occurrences of each of these molecule distributions is obtained by multiplying T4 by the probability of getting the distribution when four molecules are distributed among 64 bins. Adding an inflated IMI to M2.2 will always give the IMI distribution I32. while M3.1 will produce Is, 2 one in four times (and I4.1 otherwise). The estimated occurrences of M3,I are split proportionally between Is, 2 and h.i. The total estimated occurrences of I32 for m-4 (the sum of the estimates of I3.2 from M34 and M2, 2) is subtracted from the remaining counts of Is,2. The remaining counts can then be used to estimate the occurrences of I3.2 when m=3, and so on. Note that when m=3, there are two inflated IMIs, so the inflation probabilities must be included in the estimation. The molecule distribution M , 1 can produce Is, 2 either by having two molecules each with one inflated IMI (one in each bin) or having two inflated IMIs from one molecule in its own bin. The total occurrences of M .i from T3 with two inflated IMIs is first calculated by multiplying T3 by the probability of M2.1, then splitting this estimate by the relative inflation probabilities (four ways to select one inflated molecule multiplied by the probability of two uninflated molecules and one inflated with two IMIs. and C(4, 2)=6to select two inflated molecules multiplied by the probability of one uninflated molecule and two molecules with one inflated IMI). The estimates within these two inflation configurations are then further split betw een the possible resulting IMI distributions in each case (Is, 2 and .i), and subtracted from the remaining counts of the IMI distributions. For values of m and z with more than one inflated IMI, the probabilities of each potential configuration of inflated molecules that could be assigned to m molecules is thereby accounted for.

[0155] This process of estimating the number of observations for each IMI distribution is performed for each number of molecules m from m=i down to m=l. When m=l, the remaining observed count (shown as reference number 880) of i IMIs in 1 bin becomes the count for i IMIs from 1 molecule in the inflation distribution. Because this process starts at z=I and goes up to the maximum possible i, for any number of IMIs i, the inflation distribution has already been calculated for values from 1 to M, giving all the probabilities (shown by reference number 888) necessary to estimate the counts from each molecule distribution for each distribution of i IMIs. If the remaining count when m= is 0 at any point (i. e. , there are no instances of i IMIs from a single molecule), the algorithm terminates. As discussed herein, the inflation probabilities 888, in turn, may be used to derive (step 892) a correction factor (896) as discussed herein.

[0156] Aspects of the presently described techniques provide a system for nucleic acid analysis that includes a solid support and a nucleic acid construct attached to the solid support, such as a bead. The nucleic acid construct includes a linker for attachment to the solid support, a cell-identification barcode, a binning index, a capture region, and a region of cDNA comprising a portion in which a unique identifier sequence that is intrinsic to the cDNA has been generated. In certain embodiments each bead is linked to a plurality of the nucleic acid constructs and the binning indexes are about 2 to 6 bases in length, such as 3 bases. In some embodiments, the region of cDNA has been randomly cleaved at a cut site, and a synthetic oligonucleotide has been attached at the cut site. The system may include a transposase that functions to cleave the cDNA or a primer that primes at an essentially random location, thereby generating the unique identifier sequence, and a paired-end (PE) sequence (or similar synthetic oligonucleotide) for hybridization to a sequencing surface. In certain embodiments the transpose cleaves the cDNA at a cut site that is random or cannot be predicted and attaches the paired-end sequence to the cDNA at the cut site. The unique identifier sequence is defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The system may include a plurality' of paired-end sequence-ligated cDNAs, e.g., all linked to the solid support (through oligonucleotides that each include a binning index) and each comprising an identifier sequence in the cDNA adjacent a random cutsite. In certain embodiments, sequence reads from the plurality of cDNAs can be deduplicated to quantify RNAs captured on the solid support.

[0157] In some embodiments, the cDNA has been randomly cleaved by a restriction enzyme or sonication, and the synthetic oligonucleotide has been attached by a ligase. In embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by a random hexamer. For example, an RNA may have been captured by a random primer that was extended to create the cDNA such that the primer-binding site is random and a segment of bases in the cDNA adjacent the priming site is useful as a unique identifier sequence. In certain embodiments, the cDNA has been randomly cleaved by, and the synthetic oligonucleotide has been attached by, a ligase.

[0158] In certain implementations the unique identifier sequence is defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The plurality of bases may be intrinsic, e.g., copied from genetic material of an organism. The system may include a plurality of the solid supports (e.g., beads), each solid support attached to cDNA copies of RNAs from a single cell, in which each cDNA copy has a unique identifier defined by bases in a segment of that cDNA copy adjacent a random cut site, such that the RNAs from a single cell can be quantified by sequencing the RNAs and deduplicating sequence reads by the unique identifier. In certain embodiments, the deduplicated reads are counted and each count is associated with its binning index. For each binning index, the count is corrected (e.g., divided or multiplied by a correction factor to correct for bias introduced during sample preparation) and then, for each binning index, the counts are added together to provide a quantitative measure of transcripts in a sample from which the cDNAs were prepared. The solid supports may comprise hydrogel beads linked to a plurality of capture oligonucleotides that each include a binning index, e.g., of about three bases. The system may include a plurality of the hydrogel beads, each isolated in an aqueous partition.

[0159] Other aspects of the presently described techniques provide a method for generating nucleic acid library. The method includes providing a sample comprising a plurality of cells, each comprising sample nucleic acids; hybridizing sample nucleicacids to a construct comprising a solid support to which is attached, via a linker, a cellular barcode sequence, a binning index of fewer than about eight bases (e.g. fewer than five bases), and a capture sequence; extending said construct from said capture sequence to form a duplex comprising an extended construct; exposing said duplex to a transposase, thereby to generate a unique identifier sequence at a 3' end of said construct; and amplifying said extended construct; thereby to create a nucleic acid library. The extending step may reverse transcribe a sample nucleic acid into a cDNA in the construct. In certain embodiments the transposase cuts the cDNA at a random cut site. The unique identifier sequence may be provided by a segment of the cDNA adjacent or near the random cut site. The method may include sequencing the library to generate sequence reads, mapping the sequence reads to genes in a reference, collapsing (i.e., de-duplicating) reads that include the same unique identifier sequence leaving only unique reads, counting unique reads, and associating each count with its binning index.

[0160] In some embodiments, the method includes isolating the cells into partitions and creating sequencing libraries from single cells in the partitions. The solid support may be a bead attached to a plurality of copies of the cellular barcode sequence. The method may include attaching synthetic oligonucleotides (such as paired-end (PE) sequences, PCR handles, or sequencing adaptors) at random cut sites in mRNA molecules.

[0161] In some aspects, the presently described techniques provide a method for generating a nucleic acid library. The method includes capturing RNA molecules from a single cell with capture oligonucleotides that include a first PCR handle; extending the capture oligos to form duplexes comprising the RNA molecules and cDNA; and cleaving the duplexes at, and attaching second PCR handles to, random cut sites to thereby form constructs that each include a label defined by intrinsic sequence of a cDNA segment adjacent the random cut site wherein at least the first or second PCR handle is provided by an oligonucleotide that includes a binning index. The method may further include amplifying the constructs to form amplicons; sequencing the amplicons to produce sequence reads; counting sequence reads with duplicate intrinsic sequences as one RNA molecule from the single cell, and storing the resultant countunder its binning index. The capture oligonucleotides may be linked to a solid support in an aqueous partition that includes the single cell.

[0162] The solid support may be a bead and the aqueous partition may be a droplet. The method may include forming a plurality of droplets that each include, on average, one bead decorated with capture oligonucleotides and zero or one single cell. In some embodiments, the droplets are formed in channels of a microfluidic device. In certain embodiments the plurality of droplets are formed substantially simultaneously by shearing or vortexing a vessel comprising an aqueous phase, an immiscible phase, oligo-linked beads, and cells. The capture oligonucleotides may further include cell barcodes.

[0163] In certain embodiments, the constructs include at least the first PCR handles, the cell barcodes, cDNAs, and the second PCR handles, in which one of the PCR handles includes the binning index. The amplicons may include copies of the binning index and the first and second PCR handles such that the copies of the first and second PCR handles anneal to sequencing adaptors. The method may include mapping the sequence reads to a reference to identify genes from which one or more of the RNA molecules were transcribed. The method may include correcting the indexed counts by a correction factor (optionally summing corrected counts per binning index) to give estimated transcription levels and providing a report with transcription levels of the genes in the single cells based on the counted sequence reads and identified genes.

[0164] In some embodiments, the cleaving and / or the attaching steps are performed by an enzy me such as a transposase that creates the random cut sites. The intrinsic sequence of the cDNA may be copied from genetic material of the single cell such that, due to the random cut site, each duplex includes a label useful to uniquely identify the cDNA is sequencing data.

[0165] Other aspects of the presently described techniques provide a method that includes cleaving a cDNA at, and attaching an oligonucleotide that includes a binning index of 3 bases or less to, a random cut site; copying the cDNA to generate copies that include the binning index and an intrinsic label copied from a segment of the cDNAadjacent the cut site; sequencing the copies to generate sequence reads; and collapsing duplicate sequence reads that contain the same intrinsic label. Deduplicated read counts are stored by binning index and optionally corrected. The intrinsic label may be some number of bases, e.g., between about 5 and about 30, from a segment near or adjacent the random cut site. The method may include capturing an mRNA w ith a capture oligo and extending the capture oligo to synthesize the cDNA. The cDNA may be linked to a bead. The method may include isolating single cells into droplets, lysing the cells to release RNA into the droplets, capturing an mRNA within one of the droplets, and making the cDNA from the mRNA. The droplets may be formed simultaneously in a technique that also includes isolating the cells into the droplets (e.g.. shearing a mixture that includes beads and cells in an aqueous phase plus an oil). The cleaving and the attaching at the random cut site may be performed using an enzyme such as a transposase. The oligo that is attached at the random cut site may include a PCR handle used in the copying step. The attaching step may yield a DNA construct including a first priming site, a cell barcode, the binning index, a portion of the cDNA, the random cut site, and a second priming site. The copying step may comprise amplification by polymerase chain reaction (PCR).

[0166] Overview of System for Biological or Chemical Analysis

[0167] With the preceding discussion in mind, examples and embodiments described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction. Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No. 7,741,463, entitled 'Method of Preparing Libraries of Template Polynucleotides.” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety; and / orU.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007. the disclosure of which is incorporated by reference herein, in its entirety.

[0168] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.

[0169] FIG. 17 is a block diagram of an exemplary server device 1002 that may be used in connection with the system of FIG. 18. The server device 1002 may be configured to determine a copy number variant genotype in a nucleic acid sample. The general architecture of the server device 1002 depicted in FIG. 17 includes an arrangement of computer hardware and software components. The server device 1002 may include many more (or fewer) elements than those shown in FIG. 17. It is not necessary’, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 1002 includes a processing unit 1019, a network interface 1020, a computer readable medium drive 1030, an input / output device interface 1040, a display 1050, and an input device 1060, ah of w hich may communicate with one another by w ay of a communication bus. The network interface 1020 may provide connectivity7to one or more networks or computing systems. The processing unit 1010 may thus receive information and instructions from other computing systems or services via a network. The processing unit 1010 may also communicate to and from memory71070 and further provide output information for an optional display 1050 via the input, / output device interface 1040. The input / output device interface 1040 may also accept input from the optional input device 1060, such as a keyboard, mouse, digital pen. microphone, touch screen, gesturerecognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0170] The memory 1070 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 1010 executes in order to implement one or more embodiments. The memory’ 1070 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 1070 may store an operating system 1072 that provides computer program instructions for use by the processing unit 1010 in the general administration and operation of the server device 1002. The memory 1070 may store a reference genome 10710, such as for use by the sequencing application 1010. The memory 1070 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0171] For example, in one embodiment, the memory 1070 includes a sequencing application 1010, which may include copy number detection system 1016. The system 1016 can perform the methods disclosed herein. In addition, memory’ 1070 may include or communicate with the data store 1090 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining a copy number variant genotype in a nucleic acid sample of the present disclosure, such the sequencing reads, the estimated copy number(s), and the variant call (for example, the detection of a copy number variant) determined.

[0172] In some embodiments, the disclosed sy stems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data byusers. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0173] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0174] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0175] Example of System with Higher V olume Throughput

[0176] With the preceding in mind, FIG. 18 illustrates a schematic diagram of an example of a system (1100) that may be used to perform an analysis on one or more samples of interest. A controller ( 1114) of the present example includes a user interface (1206), a communication interface (1208), one or more processors (1210), and a memory (1212) storing instructions executable by the one or more processors (1210) to perform various functions including the disclosed implementations. User interface (1206). communication interface (1208), and memory (1212) are electrically and / or communicatively coupled to the one or more processors (1210). User interface (1206) may be adapted to receive input from a user and to provide information to the user associated with the operation of system (1100) and / or an analysis taking place. User interface (1206) may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.

[0177] Communication interface (1208) is adapted to enable communication between system (1100) and a remote system(s) (e.g., computers) via a network(s) (e.g.. the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system (1100). Some of the communications provided to system (1100) may be associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system (1100).

[0178] The one or more processors (1210) and / or system (1100) may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors (1210) and / or system (1100) includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.

[0179] Memory (1212) may include one or more of a semiconductor memoiy, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memoi7(EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).

[0180] In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation show n, system ( 1100) is configured to receive a flow cell cartridge assembly (1102) including a flow cell assembly (1103) and a sample cartridge (1104). System (1100) includes a flow cell receptacle (1122) that receives flow cell cartridge assembly (1102), a vacuum chuck (1124) that supports flow cell assembly (1103), and a flow cell interface (1126) that is used to establish a fluidic coupling between system (1100) and flow cell assembly (1103). Flow cell interface (1126) may include one or more manifolds. System (1100) further includes a sipper manifold assembly (1106), a sample loading manifold assembly (1108), and a pump manifold assembly (1110). System (1100) also includes a drive assembly (1112). the controller (1114), an imaging system (1 116), and a waste reservoir (1 118). Controller (1114) is electrically and / or communicatively coupled to drive assembly (1112) and to imaging system (1116); and is configured to cause drive assembly (1112) and / or the imaging system (1116) to perform various functions for performing the techniques disclosed herein.

[0181] In the present example, flow cell assembly (1103) includes a flow cell (1128) having a channel (1130) and defining a plurality of first openings (1132), which are fluidically coupled to the channel (1130) and arranged on a first side (1134) of the channel (1130). Flow cell (1128) further includes a plurality of second openings (1136) fluidically coupled to the channel (1130) and arranged on a second side (1138) of the channel (1130). Fluid may thus flow through flow cell (1128) via the channel (1130). While the flow cell (1128) is shown including one channel (1130). flow cell (1 128) may include two or more channels (1130). Flow cell assembly (1103) also includes a flow cell manifold assembly (1140) coupled to flow' cell (1128) and having a first manifold fluidic line (1142) and a second manifold fluidic line (1144). Flow7cell manifold assembly (1140) may be in the form of a laminate including a plurality of layers as discussed in more detail below'.

[0182] In the implementation shown, first manifold fluidic line (1142) has a first fluidic line opening (1146) and is fluidically coupled to each of the first openings (1132) of flow cell (1128); and second manifold fluidic line (1144) has a second fluidic lineopening (1148) and is fluidically coupled to each of the second openings (1136). As shown, flow cell assembly (1103) includes gaskets (1150) coupled to flow cell manifold assembly (1140) and fluidically coupled to fluidic line openings (1146, 1148). In some implementations where flow cell (1128) includes a plurality of channels (1130), flow cell manifold assembly (1140) may include additional fluidic lines (1152) that couple first fluidic line openings (1146) to a single manifold port (1154). In such implementations, a single gasket (1150) may be coupled to flow cell manifold assembly (1140) that surrounds the manifold port (1154) and is in fluidic communication with a plurality of channels (1130). In operation, flow cell interface (1126) engages with corresponding gaskets (1150) to establish a fluidic coupling between system (1100) and flow cell (1128). The engagement between flow cell interface (1126) and gaskets (1 150) reduces or eliminates fluid leakage between flow cell interface ( 1126) and flow cell (1128).

[0183] In the implementation shown, first manifold fluidic line (1142) has a portion (1156) that is substantially parallel to a longitudinal axis (1158) of channel (1 130); and second manifold fluidic line (1144) has a portion (1160) that is substantially parallel to longitudinal axis (1158) of channel (1130). Additionally, first manifold fluidic line (1142) is shown being at least partially adjacent a first end (1162) of flow cell (1128) and spaced from a second end (1164) of flow cell (1128); and second manifold fluidic line (1144) is shown being at least partially adjacent second end (1164) of flow cell (1128) and spaced from first end (1162). Other arrangements of manifold fluidic lines (1142, 1144) may prove suitable, however.

[0184] In the implementation shown, system (1100) includes a sample cartridge receptacle (1166) that receives sample cartridge (1104) that carries one or more samples of interest (e.g., an analyte). System (1100) also includes a sample cartridge interface (1168) that establishes a fluidic connection with sample cartridge (1104). Sample loading manifold assembly (1 108) includes one or more sample valves (1170). Pump manifold assembly (1110) includes one or more pumps (1172), one or more pump valves (1174), and a cache (1176). Valves (1170, 1174) and pumps (1172) may take any suitable form. Cache (1176) may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations ofthe system (1100). While cache (1176) is shown being included in pump manifold assembly (1110). cache (1176) may alternatively be located elsewhere (e.g.. in sipper manifold assembly (1106) or in another manifold downstream of a bypass fluidic line (1178), etc.).

[0185] Sample loading manifold assembly (1108) and pump manifold assembly (1110) flow one or more samples of interest from sample cartridge (1104) through a fluidic line (1180) toward flow cell cartridge assembly (1102). In some implementations, sample loading manifold assembly (1108) may individually load or address each channel (1130) of flow cell (1128) with a respective sample of interest. The process of loading channel (1130) with a sample of interest may occur automatically using system (1100). As shown in FIG. 14, sample cartridge (1104) and sample loading manifold assembly (1108) are positioned downstream of flow cell cartridge assembly (1102). In the implementation shown, sample loading manifold assembly (1108) is coupled between flow cell cartridge assembly (1102) and pump manifold assembly (11 10). To draw a sample of interest from sample cartridge (1104) and toward pump manifold assembly (1110), sample valves (1170), pump valves (1174), and / or pumps (1172) may be selectively actuated to urge the sample of interest toward pump manifold assembly (1110). Sample cartridge (1104) may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves (1170). To individually flow the sample of interest toward channel (1130) of flow' cell (1128) and away from pump manifold assembly (1110), sample valves (1170), pump valves (1174). and / or pumps (1172) may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly (1102) and into respective channels (1 130) of flow cell (1128).

[0186] Drive assembly (1112) interfaces with sipper manifold assembly (1106) and pump manifold assembly (1110) to flow one or more reagents that interact with the sample within flow cell (1 128). In some scenarios, a reversible terminator is attached to the reagent to allow' a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system(1116) excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system (1100).

[0187] After image data is obtained, drive assembly (1112) interfaces with sipper manifold assembly (1106) and pump manifold assembly (1110) to flow another reaction component (e.g., a reagent) through flow cell (1128) that is thereafter received by waste reservoir (1118) via a primary waste fluidic line (1182) and / or otherwise exhausted by system (1100). Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.

[0188] The primary waste fluidic line (1182) is coupled between pump manifold assembly (1110) and waste reservoir (1118). In some implementations, pumps (1172) and / or pump valves (1174) of pump manifold assembly (1110) selectively flow the reaction components from flow cell cartridge assembly (1102), through fluidic line (1180) and sample loading manifold assembly (1108) to primary waste fluidic line (1182). Flow cell cartridge assembly (1102) is coupled to a central valve (1184) via flow cell interface (1126). Central valve (1184) is coupled with flow cell interface (1126) via a fluidic line (1185). An auxiliary waste fluidic line (1186) is coupled to central valve (1184) and to waste reservoir (1118). In some implementations, auxiliary waste fluidic line (1186) receives excess fluid of a sample of interest from flow cell cartridge assembly (1102), via central valve (1184), and flows the excess fluid of the sample of interest to waste reservoir (1118) when back loading the sample of interest into flow cell (1128).

[0189] Sipper manifold assembly (1106) includes a shared line valve (1188) and a bypass valve (1190). Shared line valve (1188) may be referred to as a reagent selector valve. Central valve (1184) and the valves (1188, 1190) of sipper manifold assembly (1106) may be selectively actuated to control the flow of fluid through fluidic lines (1192, 1194, 1196). Sipper manifold assembly (1 106) may be coupled to acorresponding number of reagent reservoirs (1198) via reagent sippers (1200). Reagent reservoirs (1198) may contain fluid (e.g., reagent and / or another reaction component). In some implementations, sipper manifold assembly (1 106) includes a plurality of ports. Each port of sipper manifold assembly (1106) may receive one of the reagent sippers (1200). Reagent sippers (1200) may be referred to as fluidic lines. Some forms of reagent sippers (1200) may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly (1106). Reagent reservoirs (1198) may be provided in a cartridge, and the tubes of reagent sippers (1200) may be configured to be inserted into corresponding reagent reservoirs (1198) in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir (1198) into the sipper manifold assembly (1106).

[0190] Shared line valve (1188) of sipper manifold assembly (1106) is coupled to central valve (1184) via shared reagent fluidic line (1192). Different reagents may flow through shared reagent fluidic line (1192) at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly (1110) may draw wash buffer through shared reagent fluidic line (1192), central valve (1184), and flow cell cartridge assembly (1102).

[0191] Bypass valve (1190) of sipper manifold assembly (1106) is coupled to central valve (1184) via dedicated reagent fluidic lines (1194, 196). Each of the dedicated reagent fluidic lines (1194, 1196) may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines (1194. 1196) may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.

[0192] Bypass valve (1190) is also coupled to cache (1176) of pump manifold assembly (1110) via bypass fluidic line (1178). One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line (1178). The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly (1102). Thus, the operations using bypass fluidic line (1178) may occur during, for example, incubation of one or more samplesof interest within flow cell cartridge assembly (1102). That is, shared line valve (1188) may be utilized independently of bypass valve (1190) such that bypass valve (1190) may utilize bypass fluidic line (1 178) and / or cache (1176) to perform one or more operations while shared line valve (1188) and / or central valve (1184) simultaneously, substantially simultaneously, or offset synchronously perform other operations.

[0193] Drive assembly (1112) includes a pump drive assembly (1202) and a valve drive assembly (1204). Pump drive assembly (1202) may be adapted to interface with one or more pumps (1172) to pump fluid through flow cell (1128) and / or to load one or more samples of interest into flow cell (1128). Valve drive assembly (1204) may be adapted to interface with one or more of the valves (1170, 1174, 1184, 1188, 1190) to control the position of the corresponding valves (1170, 1174, 1184, 1188, 1190).

[0194] While the foregoing examples are provided in the context of a system (1100) that may be used in nucleotide sequencing processes, the teachings herein may also be readily applied in other contexts, including in systems that perform other processes (i.e., other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes.

[0195] Definitions

[0196] Terms used herein will be understood to take on their ordinary meaning in the relevant art unless specified otherwise. Several terms used herein and their meanings are set forth below.

[0197] As used herein, the singular forms “a,” “an,’" and ‘‘the” refer to both the singular as well as plural, unless the context clearly indicates otherwise. The term ■‘comprising” as used herein is synonymous with ’‘including,” '‘containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps.

[0198] Reference throughout the specification to “one example,” “another example,” “an example,” and so forth, means that a particular element (e.g., feature, structure, composition, configuration, and / or characteristic) described in connectionwith the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.

[0199] The terms "‘substantially" and “about" used throughout this disclosure, including the claims, are used to describe and account for small fluctuations, such as those due to variations in processing. For example, these terms can refer to less than or equal to ±5% from a stated value, such as less than or equal to ±2% from a stated value, such as less than or equal to ±1% from a stated value, such as less than or equal to ±0.5% from a stated value, such as less than or equal to ±0.2% from a stated value, such as less than or equal to ±0.1% from a stated value, such as less than or equal to ±0.05% from a stated value.

[0200] Adapter'. An oligonucleotide sequence that can be fused to a nucleic acid molecule (e.g., nucleotide), for example, by ligation or tagmentation. Suitable adapter lengths may range from about 10 nucleotides to about 100 nucleotides, or from about 12 nucleotides to about 60 nucleotides, or from about 15 nucleotides to about 50 nucleotides. The adapter may include any combination of nucleotides and / or nucleic acids. In some examples, the adapter can include an amplification domain, e.g., having a universal nucleotide sequence, such as a P5 or P7 sequence, that can serve as a starting point for template amplification and cluster generation. In other examples, the adapter can include a sequence that is complementary to at least a portion of a flow cell surface bound primer (which includes the universal nucleotide sequence). In the latter example, the adapter sequence can hybridize to the complementary flow cell surface bound primer during amplification and cluster generation. In some examples, the adapter can also include a sequencing primer sequence (i.e., sequencing binding site) or a sequencing sample index (i.e., a barcode sequence). Combinations of different adapters may be incorporated into the nucleic acid molecule, such as the cDNA fragments generated via tagmentation.

[0201] Amplification-. Some embodiments further comprise amplifying and / or replicating one or more nucleic acid templates, including fragments thereof. Theamplifying and / or replicating comprises use of one or more of a bridge amplification reaction, an isothermal bridge amplification reaction, a rolling circle amplification (RCA) reaction, a modified rolling circle multiple displacement amplification, a helicase-dependent amplification reaction, a recombinase-dependent amplification reaction, a single-stranded DNA binding (SSB) protein mediated Isothermal amplification, a PCR reaction, a strand-displacement reaction, a ligase chain reaction, a transcription-mediated reaction, a loop-mediated amplification reaction, other suitable reactions, and combinations thereof.

[0202] Amplification Domain'. A portion of an adapter having a universal nucleotide sequence, such as a P5 or P7 sequence or a complement thereof, that can serve as a starting point for template amplification and cluster generation.

[0203] Clusters: Moreover, as used herein, the term "‘cluster of oligonucleotides’' (or “cluster” or “oligonucleotide cluster” or “colony”) refers to a localized group or collection of DNA or RNA on a nucleotide-sample support, such as a flow cell, particle, polymer scaffold, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, the cluster can comprise one or more concatemers, such as, for example, a polony or a nanoball. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-pattemed flow cell. In typical embodiments, a cluster is the product of an amplification reaction. A cluster of oligonucleotides can be imaged utilizing one or more light signals, changes in pH, changes in conductance, and other signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent labeled nucleotides incorporated into oligonucleotides, fluorescent labeled nucleotides bound but not incorporated into oligonucleotides, and other fluorescent labeled complexes associated with incorporated or bound nucleotides from one or more clusters on a flow cell. Examples of other sequencing procedures are set forth herein. In some embodiments, a cluster can be monoclonal or polyclonal.

[0204] Complementary DNA (cDNA): A synthetic deoxyribonucleic acid strand made from an RNA strand using a reverse transcriptase.

[0205] Corresponds with: When one primer “corresponds with” an amplification domain, the primer and amplification domain may have the same sequence, so that a copy of the amplification domain generates a sequence complementary to the primer.

[0206] Depositing'. Any suitable application technique, which may be manual or automated, and, in some instances, results in modification of the surface properties. Generally, depositing may be performed using vapor deposition techniques, coating techniques, grafting techniques, or the like. Some specific examples include chemical vapor deposition (CVD), spray coating (e.g., ultrasonic spray coating), spin coating, dunk or dip coating, doctor blade coating, puddle dispensing, flow through coating, aerosol printing, screen printing, microcontact printing, inkjet printing, or the like.

[0207] Depression: A discrete concave feature in a substrate or a layer of a substrate (e.g., a patterned resin) having a surface opening that is at least partially surrounded by interstitial region(s) of the substrate or the layer. Depressions can have any of a variety of shapes at their opening in a surface including, as examples, round, elliptical, square, polygonal, star shaped (with any number of vertices), etc. The cross-section of a depression taken orthogonally with the surface can be curved, square, polygonal, hyperbolic, conical, angular, etc. The depression may also have more complex architectures, such as ridges, step features, etc.

[0208] Each'. When used in reference to a collection of items, each identifies an individual item in the collection, but does not necessarily refer to every item in the collection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise.

[0209] Flow Cell'. A vessel having an enclosed flow channel where a reaction can be carried out, or a vessel having a channel that is open to a surrounding environment and in which a reaction can be carried out. The vessel with an open flow channel may be referred to herein as an open wafer flow cell. Any example of the flow cell may include an inlet for delivering reagent(s) to the channel, and an outlet for removingreagent(s) from the channel. In some examples, the flow cell enables the detection of the reaction that occurs therein. For example, the flow cell can include one or more transparent surfaces allowing for the optical detection of arrays, optically labeled molecules, or the like.

[0210] Flow channel’. An area that is defined between two bonded or otherwise attached components or that is defined within a lane so that it is open to the surrounding environment. The flow channel can selectively receive a liquid sample. In some examples, the flow channel may be defined between two patterned sequencing surfaces or a patterned sequencing surface and a lid. and thus may be in fluid communication with one or more components of the sequencing surface(s).

[0211] Fragment'. A portion or piece of cDNA generated from the RNA sample. A “partially adapted fragment” is a portion or piece of the cDNA that has been tagmented, and thus includes an adapter ligated to the 5’ end of the cDNA fragment. A “fully adapted fragment” is a portion or piece of the cDNA sample that has adapters incorporated at both the 3’ and 5’ ends of the cDNA fragment.

[0212] Fragmentation'. “Fragmentation” as described herein refers to the shearing or fragmenting of nucleic acid into shorter lengths. Fragmentation methods include enzymatic, physical (including sonication, nebulization, needle shearing, microwave, etc.), and chemical (including depurination, hydrolysis, oxidation, etc.). The terms “fragmenting enzymes” or "enzyme-based fragmentation” or “enzyme fragmentation” as used herein refers to enzymes that fragment nucleic acid. The enzymes can be a single enzyme or two or more enzymes that work together to fragment the nucleic acid. Some enzymes work on single stranded nucleic acid whereas others work on double stranded nucleic acid and yet others work on one strand of a double stranded nucleic acid. Fragmenting enzymes can cut randomly or specifically. Non-limiting examples of fragmenting enzymes include transposase, restriction enzymes, Argonaute, CRISPR -associated nuclease (Cas), endonucleases, exonuclease, topoisomerase, Fragmentase™ (New England Biolabs, Ipswich, MA). Preferred fragmentation embodiments include methods that fragment while retaining proximity information of the fragments.

[0213] Gene Isoform: Messenger RNAs (mRNAs) that are produced from the same locus but are different in their transcription start sites, protein coding DNA sequences, and / or untranslated regions.

[0214] Immobilization'. The term “immobilized”, “affixed” and “attached” are used interchangeably herein and both terms are intended to encompass direct or indirect, covalent or non-covalent attachment unless indicated otherwise, either explicitly or by context.

[0215] Non-limiting exemplary’ covalent attachment includes, for example, those that result from the use of click chemistry techniques. Exemplary non-covalent attachment includes, but are not limited to, non-specific interactions (e.g. hydrogen bonding, ionic bonding, van der Waals interactions etc.) or specific interactions (e.g. affinity interactions, receptor-ligand interactions, antibody-epitope interactions, avidinbiotin interactions, streptavidin-biotin interactions, lectin-carbohydrate interactions, etc.). Exemplary attachments are set forth in U.S. Pat. Nos. 6,737,236 Bl ; 7,259,258 B2; 7,375,234 B2and 7,427,678 B2; and U.S. Pat. Pub. 2011 / 0059865 Al, each of which is incorporated herein by reference in its entirety.

[0216] In certain embodiments, the molecules (e.g. nucleic acids, enzymes) remain immobilized or attached to the solid support under the conditions in which it is intended to use the solid support, for example in applications requiring nucleic acid amplification and / or sequencing. In other embodiments, the molecules are reversibly immobilized and can be removed from the solid support through the use of cleavable sites, linkers, and the like.

[0217] Immobilized on: When one component (e.g.. primers, transposome complexes, etc.) is attached, either directly or indirectly through an intervening component, to another component (e.g., a substrate). As an example, the primers described herein may be immobilized on the anchoring layer and the substrate. The primers may be directly attached to the anchoring layer, and indirectly attached to the substrate through the anchoring layer.

[0218] Nanoballs: Some embodiments further comprise rolling circle amplification / replication used to form nucleic acid nanoballs. The term “nucleic acid nanoball’’ may be a concatemer comprising multiple copies of a target nucleic acid molecule. These nucleic acid copies may be arranged one after another in a continuous linear strand of nucleotides. These nucleic acid copies may result in a nanoball folding configuration. The multiple copies of a target nucleic acid molecule in a nucleic acid nanoball may each contain an adaptor sequence of known sequence to facilitate amplification or sequencing. The adaptor sequence of each target nucleic acid molecule may be the same or different. The nucleic acid nanoball can be loaded on the surface of solid support. The nanoball can be attached to the surface of solid support by any suitable method. Non-limiting examples of such methods include nucleic acid hybridization, biotin streptavidin binding, thiol binding, photoactive binding, covalent binding, antibody-antigen, physical constraints via hydrogels or other porous polymers, etc., or combinations thereof. In some cases, the nanoball can be digested with an enzyme (nuclease, etc.) to produce a smaller nanoball or a fragment from the nanoball.

[0219] Nucleic Acid Sample: A sample, typically derived from any organism, including but not limited to animals, plants, fungi, and microbes. For example, such samples may be derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc ), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. Alternatively, the sample may be microbial such as bacteria, viral, or fungal. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation ofinterfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated"’ or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0220] Nucleotide: A compound consisting of a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. In RNA (ribonucleic acid), the sugar is a ribose, and in DNA (deoxyribonucleic acid), the sugar is a deoxyribose, i.e.. a sugar lacking a hydroxyl group that is present at the 2' position in ribose. The nitrogen containing heterocyclic base (i.e., nucleobase) can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives (i.e., inosine) or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-l atom of deoxyribose is bonded to N-l of a pyrimidine or N-9 of a purine. A nucleic acid analog may have any of the phosphate backbone, the sugar, or the nucleobase modified. Examples of nucleic acid analogs include, for example, universal bases or phosphate-sugar backbone analogs, such as peptide nucleic acid (PNA) or locked nucleic acid (LNA).

[0221] Patterned / Random'. In some embodiments, the solid support comprises a patterned surface suitable for immobilization of molecules, such as enz mes, nucleic acids, and complexes thereof, in an ordered pattern. A “patterned surface” refers to an arrangement of different regions in or on an exposed layer of a solid support. The features can be separated by interstitial regions that contribute to the pattern. In some embodiments, the interstitial regions can be a different height, creating wells or raised platform patterns. In other embodiments, the interstitial regions can have a different surface charges. In yet other embodiments, the interstitial regions can have a different attachment moieties. In some embodiments, the pattern can be any suitable pattern,such as a grid patterns, radial patterns, and combinations thereof. In some embodiments, a patterned surface can contain pre-determined locations of features but the features are not arrayed in a repetitive pattern. Examples of grid patterns include rectangular patterns, hexagonal patterns, triangular, and other suitable grid patterns. The regions for immobilization of molecules may be depressed regions, elevated regions, or planar regions relative to the interstitial regions. The regions may be fabricated as is generally known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, microetching techniques, and combinations thereof. As will be appreciated by those in the art, the technique used will depend on the composition and shape of the regions. For example, the regions for immobilization of molecules of a patterned surface may be wells, pits, channels, posts, pillars, ridges, stripes, swirls, lines, and other suitable topographies. For example, the wells may have any opening in any shape, such as circular, oval, polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.). Exemplary patterned surfaces that can be used in the methods and compositions set forth herein are described in U.S. Pat. No. 8,778,849 B2, which is incorporated herein by reference in its entirety.

[0222] In some embodiments, the solid support comprises a surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in a random distribution over the solid support. Exemplary random distribution over a solid support is described in U.S. Pat. No. 8,241,573 B2, which is incorporated herein by reference in its entirety.

[0223] Polonies'. Some embodiments further comprise rolling circle amplification / replication used to form polonies. The term “polony” or “polonies” used herein refers to a nucleic acid library molecule clonally amplified in-solution or on- support to generate an amplicon that can serve as a template molecule for sequencing. In some aspects, a linear library molecule can be circularized to generate a circularized library molecule, and the circularized library molecule can be clonally amplified insolution or on-support to generate a concatemer. In some aspects, the concatemer can sen e as a nucleic acid template molecule which can be sequenced. The concatemer issometimes referred to as a polony. In some aspects, a polony includes nucleotide strands.

[0224] Primer'. A single stranded nucleic acid molecule that can hybridize to a complementary sequence, such as an adapter attached to a fragment. As one example, a flow cell surface bound primer can sen e as a starting point for fragment amplification and cluster generation. As another example, a primer (e.g., a sequencing primer) may be introduced that can hybridize to fragments or fragment amplicons in order to prime synthesis of a new strand that is complementary to the fragments or fragment amplicons. Any primer can include any combination of nucleotides or analogs thereof. In some examples, the primer is a single-stranded oligonucleotide or polynucleotide. The primer length can be any number of bases long. In an example, each of the flow cell surface bound primer and the sequencing primer is a short strand, ranging from 10 to 60 bases, or from 20 to 40 bases.

[0225] PolyA and Poly . A nucleic acid sequence (e g., a capture nucleotide sequence) including a series of two or more adenine (A) bases and thymine (T) bases, respectively. A polyA or polyT can include at least about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, or more of the A or T bases, respectively. Alternatively or additionally, a polyA or polyT can include at most about 40, 38, 35, 32, 30, 28, 25, 22, 20, 18, 15, 12, 10, 8, 5, or 2 of the A or T bases, respectively.

[0226] Reverse Transcriptase'. An enzyme that converts RNA into DNA. In certain embodiments, a reverse transcriptase may be compatible with template switching oligonucleotides. For example, the reverse transcriptase may add nucleotides (CCC) to a terminus of a cDNA.

[0227] RNA Nickase'. An enzyme that can nick RNA.

[0228] RNA Sample: Multiple strands of a polymeric form of nucleotides of any length that includes ribonucleotides or ribonucleotide analogs. The RNA strands are single stranded. The RNA sample may include naturally occurring RNA, including microRNAs (miRNAs), messenger RNA (mRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA). Each ribonucleotide includes a nitrogen containing heterocyclicbase (a nucleobase such as adenine, thymine, cytosine and / or guanine), a sugar (specifically ribose), and a backbone containing phosphodiester bonds. An analog structure can have an alternate backbone linkage including any of a variety known in the art.

[0229] The RNA sample can be isolated from one or more cells, bodily fluids (e.g., whole blood, blood spots, saliva) or tissues. RNA can be prepared by lysing a cell that contains the RNA. The cell may be lysed under conditions that substantially preserve the integrity' of the cell's RNA. In one particular example, thermal lysis may be used to lyse a cell. In another particular example, exposure of a cell to alkaline pH can be used to lyse a cell while causing relatively little damage to RNA. Any of a variety of basic compounds can be used for lysis including, for example, potassium hydroxide, sodium hydroxide, and the like. Additionally, relatively undamaged RNA can be obtained from a cell lysed by an enzyme that degrades the cell wall. Cells lacking a cell wall either naturally or due to enzymatic removal can also be lysed by exposure to osmotic stress. Other conditions that can be used to lyse a cell include exposure to detergents, mechanical disruption, sonication heat, pressure differential such as in a French press device, or Dounce homogenization. Agents that stabilize RNA can be included in a cell lysate or isolated RNA sample including, for example, nuclease inhibitors, chelating agents, salts, buffers and the like. Examples of a crude cell lysate as disclosed herein containing RNA may be used without further isolation of the RNA.

[0230] Sequencing Procedures'. The term "read" or ‘“sequence read" (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A. T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be alignedand specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques.

[0231] Embodiments described herein can be used with any suitable sequencing chemistry, such as sequencing by synthesis (SBS), sequencing by binding, sequencing by ligation, or nanopore sequencing.

[0232] SBS can be with or without the use of reversible terminators. For example, SBS can be initiated by contacting the target nucleic acids with one or more nucleotides (e.g., labelled, synthetic, modified, or a combination thereof), DNA polymerase, etc. Those features where a primer is extended using the target nucleic acid as template will incorporate a labeled nucleotide that can be detected. The incorporation time used in a sequencing run can be significantly reduced using the altered polymerases described herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery' steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems, and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Pat. Nos. 7,057,026 B2, 7,329,492 B2, 7.211,414 B2, 7,315,019 B2, 7.405,281 B2, and 8,343,746 B2. Sequence reads can be generated using instruments such as MiniSeq™, MiSeq™,NextSeq™, HiSeq™, and NovaSeq™ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0233] One example of SBS is termed sequencing by binding. One implementation of sequencing by binding includes cycles of initiating sequencing of a template with a reversible blocker on the 3’ end to prevent additional bases from incorporating, interrogating the template by flooding the flow cell with fluorescently tagged bases that do not include a blocker and measuring an emitted signal of bound bases, activating the 3’ end via removal of the reversible blocker, and incorporating the complementary base from unlabeled, blocked nucleotides. Reads using sequencing by binding can be generated from using instruments such as Onso™ sequencing instruments from Pacific Biosciences of California, Inc. (Menlo Park, CA). Another implementation of sequencing by binding could be sequencing by avidity. In sequencing by avidity, fluorescent dye labeled cores termed avidites are used. One potential cycle of sequencing by avidity includes providing a reagent of polymerase and reversibly terminated nucleotides to templates immobilized on a solid surface, de-blocking the incorporated nucleotides, flowing a set of four types of avidites, washing away unbound avidites, detecting the incorporated bases / nucleotides, and removing the bound avidites. The steps in the cycle of sequencing by avidity may be performed in other orders. Sequencing by avidity is described in Arslan, S., Garcia, F.J., Guo, M. et al. Sequencing by avidity enables high accuracy with low reagent consumption. Nat Biotechnol 42, 132-138 (2024). https: / / doi.org / 10.1038 / s41587-023-01750-7, which is incorporated by reference in its entirety. Reads using sequencing by avidity can be generated using instruments such as Aviti™ sequencing instruments from Element Biosciences (San Diego).

[0234] One example of SBS using an open flow cell and without using reversible terminators is disclosed in Almogy. G.(2022) “Cost-efficient whole genomesequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform” https: / / doi.org / 10.1101 / 2022.05.29.493900, which is incorporation by reference in its entirety. Sequence reads using an open flow' cell can be generated using instruments such as UG 100TM Sequencer from Ultima Genomics. Inc. (Fremont, CA).

[0235] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are described in U.S. Pat. Nos. 8,262,900 B2, 7,948,015 B2, 8,349,167 B2, and U.S. Pat. Pub. 2010 / 0137143 Al, which are incorporated by reference in its entirety'.

[0236] Sequence reads can be generated using instruments such as DNBSEQTM sequencing instruments from MGI Tech Co., Ltd. (Shenzhen, China) and as SURFSeq™, FASTASeq™, and GenoLab™ sequencing instruments from GeneMind Biosciences Co., Ltd. (Shenzhen, China).

[0237] Some embodiments can use methods involving the real-time monitoring of DNA polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate-labeled nucleotides, or with zeromode waveguides. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33. 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176- 1181 (2008), which are incorporated by reference in its entirety. Techniques sequencing using zeromode waveguides is described in U.S. Pat. No. 6,917,726 B2, which is incorporated by reference in its entirety.

[0238] Solid Support. The terms "‘solid support ' “solid surface." and other grammatical equivalents herein refer to any substrate that is appropriate for or can be modified to be appropriate for the attachment of enzymes, nucleic acids, and complexes thereof. As will be appreciated by those in the art, the number of possible substrates is very large. Possible substrates include, but are not limited to, glass and modified or functionalized glass, polymers (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, plastics, optical fiber bundles, quartz, metal oxides, inorganic oxides, other suitable transparent materials, other suitable non-transparent materials, other suitabletranslucent materials, and combinations thereof. The composition and geometry' of the solid support can vary with its use.

[0239] In some embodiments, the solid support or solid surface is a planar structure, such as a flowcell, slide, chip, microchip, array, microarray, wafer, panel, charge pad, and / or web. The planar structure can be a single surface structure having a single surface of sample / reach on sites. The planar structure can be a dual surface structure. One example of a dual surface structure includes a top substrate having a top surface of sample / reactions sites, a bottom substrate having a bottom surface of sample / reactions sites, and a spacer layer separating the top substrate and the bottom substrate. The solid support or solid surface can be open to direct application of a fluid. One example of an open solid support or open solid surface is an open flow cell having a single surface structure without an inlet port. In some embodiments, the solid support is not necessarily planar, such as, for example, the surface of a well, tube, or other vessel. Nonlimiting examples include the surface of a microcentrifuge tube, a well of a multiwell plate, and the like.

[0240] In some embodiments, the solid support comprises one or more surfaces of a flowcell or flow cell. The term “flowcell” or “flow cell” as used herein refers to a solid surface across which one or more fluid reagents can be flowed. Examples of flowcells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. 7,057,026 B2; WO 91 / 06678; WO 07 / 123744; U.S. 7,329,492 B2; U.S. 7,211,414 B2; U.S. 7,315,019 B2; U.S. 7,405,281 B2, and U.S. Pat. Pub. 2008 / 0108082 Al, each of which is incorporated herein by reference in its entirety7. In some embodiments, the flowcells can be one or more flow lanes. For flow cells having a plurality of flow lanes, each of the flow lanes can be independently accessed or two or more flow lanes can be accessed as a group.

[0241] In some embodiments, the solid support or solid surface is a non-planar structure, such beads, microspheres, and / or inner and / or outer surface of a tube or vessel. The terms “beads”, “microspheres,” or “particles” or grammatical equivalents herein is refer to small discrete particles. Suitable bead compositions include, but arenot limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, thoria sol, carbon graphite, titanium dioxide, latex, polysaccharide (e g. Dextran™ , Sepharose™, cellulose, nylon, cross-linked micelles, Teflon™, as well as any other materials outlined herein for solid supports may all be used. “Microsphere Detection Guide’' from Bangs Laboratories, Fishers Ind. is a helpful guide. In certain embodiments, the microspheres are magnetic microspheres or beads. The beads need not be spherical; irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead sizes range from nanometers, i.e. 100 nm, to millimeters, i.e. 1 mm, with beads from about 0.2 micron to about 200 microns being preferred, and from about 0.5 to about 5 micron being particularly preferred, although in some embodiments smaller or larger beads may be used.

[0242] Tagmentation'. A process in which the cDNA sample is cleaved / fragmented and tagged (e g., with the adapters) for analysis. Tagmentation is an in vitro transposition reaction.

[0243] Transferred and Non-Transferred Strands'. The term “transferred strand” refers to a sequence that includes a transferred portion of a transposon end. Similarly, the term “non-transferred strand” refers to a sequence that includes the non-transferred portion of a transposon end. The 3’-end of a transferred strand is joined or transferred to a double-stranded fragment during tagmentation. The non-transferred strand is not joined or transferred to the double-stranded fragment during tagmentation. In an example, the transferred and non-transferred strands include at least partially complementary portions that are covalently bound together.

[0244] Transposase or Transposase Enzyme'. An enzyme that is capable of forming a functional complex with a transposon end-containing composition (e.g.. transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon end-containing composition into the double-stranded cDNA sample with which it is incubated, for example, in the in vitro transposition reaction (i.e., tagmentation). A transposase as presented herein can also include integrases from retrotransposons and retroviruses. Although many examples described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that anytransposase that is capable of inserting a transposon end with sufficient efficiency to 5 '-tag and fragment the cDNA sample for its intended purpose can be used.

[0245] Transposome Complex'. An entity formed between a transposase and a double stranded nucleic acid including a transposase integration recognition site. For example, the transposome complex can be formed when a transposase enzyme is preincubated with partially hybridized transferred and non-transferred strands under conditions that support non-covalent complex formation. The complementary portion includes double-stranded transposon DNA, for example, Tn5 DNA, a portion of Tn5 DNA, a transposon end composition, a mixture of transposon end compositions or other double-stranded DNAs capable of interacting with a transposase, such as the hyperactive Tn5 transposase. As will be described in more detail herein, transposome complexes can form dimers.

[0246] Transposon End. A double-stranded nucleic acid strand that exhibits only the nucleotide sequences (the "‘transposon end sequences”) that are necessary to form the complex with the transposase that is functional in tagmentation. The double-stranded nucleic acid strand of the transposon end can include any nucleic acid or nucleic acid analogue suitable for forming the functional complex with the transposase. For example, the transposon end can include natural DNA or DNA analogs (with modified bases and / or backbones), and can include nicks in one or both strands. Transposon ends are often referred to as mosaic ends.

[0247] Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U.S. Pat. App. Pub. 2010 / 0120098, which is incorporated herein by reference in its entirety. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is doublestranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non- covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting w ith a transposase such as the hyperactive Tn5 transposase.

[0248] Transposome Complexes

[0249] FIG. 12, FIG. 19, FIG. 20, FIG. 21, and FIG. 22 illustrate monomers of different examples of the transposome complexes (e.g., 10A, 10B, 10C and 10D) that may be used in the methods and kits disclosed herein. When a plurality of any one ty pe of the transposome complexes 10A, 10B, 10C, or 10D is incorporated into a liquid carrier, the transposome complexes 10A, 10B, 10C, or 10D are capable of forming homodimers. These homodimers may then be introduced into any of the flow cells 32 (see Fig. 23A) disclosed herein. If two different types of the complexes, e.g., 10A and 10B, are included in solution, homodimers and heterodimers will form. In any of the solutions, some transposome complexes 10A through 10D may not dimerize, and these individual transposome complexes 10A, 10B, IOC, or 10D can attach to the flow cell surface. The monomeric transposome complex(es) 10A, 10B, IOC, or 10D will not participate in tagmentation. During some of the methods set forth herein, two types of transposome complex(es) 10A and 10B. or 10A and IOC are used together. It is to be understood that homodimers of these complexes 10A and 10B, or 10A and IOC, may be formed separately in solution and then added to the flow cell 32 for attachment thereto.

[0250] In some of the figures, the transposome complex(es) 10A, 10B, IOC, 10D are collectively referred to with reference numeral “10.”

[0251] Referring specifically to Fig. 19 and Fig. 20, each of the transposome complexes 10A, 10B includes a transposase enzyme 12A, 12B non-covalently bound to a transposon end 14A, 14B. Each transposon end 14A, 14B is a double-strandednucleic acid strand, one strand MEA or MEB of which is part of a transferred strand 16A, 16B and the other strand ME’ or ME’B of which is the non-transferred strand 18A, 18B. In other words, the transposon end 14A, 14B includes a portion of the respective transferred strand 16A, 16B that is hybridized to the non-transferred strand 18A, 18B.

[0252] The transferred strand 16A includes a 5’ end functional group 20A, a first amplification domain 26, and a sequencing primer sequence 28A that is attached to the strand MEA of the transposon end 14 A. The strand MEA of the transposon end 14A is positioned at the 3 ’ end of the transferred strand 16A. In some examples, the transferred strand 16A further includes an index sequence 30A positioned between the first amplification domain 26 and the sequencing primer sequence 28A.

[0253] Similar to the transferred strand 16A, the transferred strand 16B includes a 5’ end functional group 20B, a second amplification domain 38, and a sequencing primer sequence 28B that is attached to the strand ME’B of the transposon end 14B. The strand ME’B of the transposon end 14B is positioned at the 3’ end of the transferred strand 16B. In some examples, the transferred strand 16B further includes an index sequence 30B positioned between the second amplification domain 38 and the sequencing primer sequence 28B.

[0254] The 5 ’ end functional groups 20A, 20B may be any functional group that is capable of covalently or non-covalently attaching to an anchoring layer 22 in the flow cell 32. In one example, the surface functional groups of the anchoring layer 22 include azide or tetrazine surface groups, and the 5’ end functional groups 20 A, 20B respectively include a terminal alkyne (e.g., hexynyl) or an internal alkyne, where the alkyne is part of a cyclic compound (e.g., bicyclo[6.1.0]nonyne (BCN) or dibenzocyclooctyne (DBCO)). In another example, the anchoring layer 22 includes biotin surface groups, and the 5’ end functional groups 20A, 20B are each biotin. In these examples, additional streptavidin or avidin is added to indirectly attach the biotin groups to one another. In another example, the anchoring layer 22 is streptavidin, and the 5' end functional groups 20A. 20B are each biotin. In still another example, theanchoring layer 22 is functionalized with one member of a binding pair, and the second member of the binding pair is used for each of the 5' end functional group 20A, 20B.

[0255] The first and second amplification domains 26, 38 have different sequences from each other, but have the same sequence, respectively, as first and second primers 34, 36 attached to the anchoring layer 22 (e.g., in the flow cell 32). The first amplification domain 26, its complement, and the primer 34, together with the second amplification domain 38, its complement, and the primer 36 enable the amplification of the cDNA sample fragments generated during tagmentation.

[0256] Examples of suitable sequences for the first amplification domain 26 / primer 34 and for the second amplification domain 38 / primer 36 include P5 and P7 primer sequences; P15 and P7 primer sequences; or any combination of the PA primer sequences, the PB primer sequences, the PC primer sequences, and the PD primer sequences set forth herein. Examples of P5 and P7 primer sequences are used on the surface of commercial flow cells sold by Illumina Inc. for sequencing, for example, on HISEQ™, HISEQX™, MISEQ™, MISEQDX™, MINISEQ™, NEXTSEQ™, NEXTSEQDX™, NOVASEQ™, ISEQ™, GENOME ANALYZER™, and other instrument platforms.

[0257] The P5 amplification domain / primer sequence is one of:P5 #I: 5’ — >- 3’AATGATACGGCGACCACCGAGAUCTACAC (SEQ. ID. NO. 1);P5 #2: 5’ ^ 3’AATGATACGGCGACCACCGAGAnCTACAC (SEQ. ID. NO. 2) where “n” is inosine in SEQ. ID. NO. 2; orP5 #3: 5’ — >- 3’AATGATACGGCGACCACCGAGAnCTACAC (SEQ. ID. NO. 3)where “n” is alkene-thymidine (i. e. , alkene-dT) in SEQ. ID. NO. 3.The P5’ sequence is the complement of any of the P5 examples.The P7 amplification domain / primer sequence may be any of the following:P7 #1: 5’ — 3’CAAGCAGAAGACGGCATACGAnAT (SEQ. ID. NO. 4)P7 #2: 5’ — > 3’CAAGCAGAAGACGGCATACnAGAT (SEQ. ID. NO. 5)P7 #3: 5’ — * 3’CAAGCAGAAGACGGCATACnAnAT (SEQ. ID. NO. 6) where “n” is 8-oxoguanine in each of SEQ. ID. NOS. 4-6. The P7’ sequence is the complement of any of the P7 examples.

[0258] In the examples of P5 and P7, the uracil, inosine, orc’n’’ is a cleavage site 40A, 40B. The cleavage sites are orthogonal (i.e., not susceptible to the same cleaving agent), so that after fully adapted fragments are generated and amplified, forward or reverse strands can be removed, for example, from a flow cell surface, while the other of the reverse or forward strands remains atached for exposure to sequencing.

[0259] It is to be understood that other sequences may be used for the amplification domains 26, 38 and for the primers 34, 36, as long as the combination enables the desired amplification. As other examples, a P15, PA, PB. PC, or PD primer may be used.The P15 amplification domain / primer sequence is:P15: 5’ 3’AATGATACGGCGACCACCGAGAnCTACAC (SEQ. ID. NO. 7) where “n” is allyl-T (i.e., a thymine nucleotide analog having an allyl functionality).The other amplification domain / primer sequences (PA-PD) mentioned above include:PA 5‘ ^ 3‘GCTGGCACGTCCGAACGCTTCGTTAATCCGTTGAG (SEQ. ID. NO. 8)PB 5 ' — > 3'CGTCGTCTGCCATGGCGCTTCGGTGGATATGAACT (SEQ. ID. NO. 9)PC 5' ^ 3'ACGGCCGCTAATATCAACGCGTCGAATCCGCAACT (SEQ. ID. NO. 10)PD 5 ’ ^ 3’GCCGCGTTACGTTAGCCGGACTATTCGATGCAGC (SEQ. ID. NO. 11)

[0260] While not shown in the example sequences for PA-PD, it is to be understood that any of these sequences may include a cleavage site 40A, 40B, such as uracil, 8- oxoguanine, allyl-T, diols, etc. at any point in the strand. As previously mentioned, the sequences for the first amplification domain 26 / primer 34 and for the second amplification domain 38 / primer 36 may be selected to have orthogonal cleavage sites (i.e., one cleavage site is not susceptible to the cleaving agent used for the other cleavage site), so that after amplification, forward or reverse strands can be cleaved, leaving the other of the reverse or forw ard strands for sequencing.

[0261] Referring briefly to the primers 34, 36 used in the methods disclosed herein, each of the primers 34, 36 may also include a polyT sequence at the 5’ end of the primer sequence. In some examples, the polyT region includes from 2 T bases to 20 T bases. As specific examples, the polyT region may include 3, 4, 5, 6, 7, or 10 T bases.

[0262] Referring back to Fig. 19 and Fig. 20, the sequencing primer sequences 28A, 28B have different sequences from each other that respectively bind to sequencing primers that are introduced, e.g., to a flow cell surface after amplification have been performed. As examples, the sequencing primer sequences 28A may bind a sequencing primer that primes synthesis of a new strand that is complementary, e.g., to forward strand fragments / fragment amplicons, and the sequencing primer sequence 28B may bind a sequencing primer that primes synthesis of a new strand that is complementary, e.g., to reverse strand fragments / fragment amplicons.

[0263] The transposon ends 14 A, 14B of each transposome complex 10A, 10B include the strands MEA, MEB respectively hybridized to the strands ME’A, ME’B. AS such, the strands MEA, ME’A are complementary and the strands MEB, ME’B are complementary. The double-stranded transposon ends 14A, 14B are respectively capable of complexing with the transposases 12A, 12B. As examples, the strands MEA, ME'A and MEB, ME’B of the transposon ends 14A, 14B may be the related but nonidentical 19-base pair (bp) outer end (e.g., strands MEA, ME’A) and inner end (e.g., strands MEB, ME’B) sequences that serve as the substrate for the activity of the Tn5 transposase, or the mosaic ends recognized by a wild-type or mutant Tn5 transposase, or the R1 end (e.g., strands MEA, ME’A) and the R2 end (strands MEB, ME’B) recognized by the MuA transposase.

[0264] When included, the index sequences 30A, 30B of the transposome complexes 10A, 10B are the same, and include a particular nucleic acid sequence that function as a barcode for the cDNA sample tagmented with the transposome complexes 10A, 10B. The unique indexes can be used for sample identification. Index sequences 30A, 30B may range from 7 bases to 15 bases long.

[0265] In the examples shown in Fig. 19 and Fig. 20, the non-transferred strands 18A, 18B are made up of the strands MEB, ME’B.

[0266] When used together, homodimers of the respective transposome complexes 10A, 10B are generated, and the homodimers are attached to a flow cell surface viatheir respective 5’ end functional groups 20A, 20B. This is an example of homodimers that are symmetrically attached.

[0267] Alternatively, the transposome complex 10A can be used with the transposome complex IOC shown in Fig. 21. When used together, homodimers of the respective transposome complexes 10A, IOC are generated, and the homodimers are attached to a flow cell surface. The homodimers of the complexes 10A attach via their 5’ end functional groups 20A, while the homodimers of the complexes IOC attach via their 3’ end functional groups 42. This is an example of homodimers that are capable of asymmetric attachment to a surface.

[0268] As shown in Fig. 21, the transposome complex IOC includes a transposase enzy me 12C non-covalently bound to the transposon end 14C. The transposon end 14C is a double-stranded nucleic acid strand, one strand (e.g., MEc) of which is part of the transferred strand 16C and the other strand (e.g., ME’c) of which is a part of the nontransferred strand 18C. Any of the example strands for the transposon end 14A, 14B (e.g., MEA and ME’A or MEB and ME'B) described herein may be used.

[0269] In the transposome complex IOC. the transferred strand 16C includes a second amplification domain 38 and a sequencing primer sequence 28C that is attached to one strand MEc of the transposon end 14C. In some examples, the transferred strand IOC includes an index sequence 30C between the second amplification domain 38 and the sequencing primer sequence 28C. The strand MEc of the transposon end 14C is positioned at the 3’ end of the transferred strand 16C.

[0270] Similar to the transposome complexes 10A, 10B, the first and second amplification domains 26, 38 of the transposome complexes 10A, IOC have different sequences from each other, but have the same sequence, respectively, as first and second primers 34, 36 described herein.

[0271] Similar to the sequencing primer sequences 28A, 28B, the sequencing primer sequences 28A, 28C have different sequences from each other that respectively bind to sequencing primers introduced during sequencing.

[0272] As mentioned, the homodimers of the respective transposome complexes 10A, IOC are configured for asymmetric attachment to the flow cell surface. As such, one of the complexes, e.g., complex IOC, includes a 3’ end functional group 42 for attachment to the surface, and the other of the complexes, e.g., complex 10A, includes the 5' end functional group 20A for attachment to the surface. The 3’ end functional group 42 and the 5' end functional group 20A may be any functional group that is capable of covalently or non-covalently attaching, directly or indirectly, to the anchoring layer 22. Any of the groups 20A, 20B may be used for the 3 ’ end functional group 42.

[0273] Referring now to FIG. 22, still another example of the transposome complex 10D is shown. This transposome complex 10D is a forked transposome complex.

[0274] The transferred strand 16D of the transposome complex 10D is similar to the transferred strand 16 A, and includes a 5’ end functional group 20D, the first amplification domain 26, and a sequencing primer sequence 28D that is attached to the strand MED of the transposon end 14D. The strand MED of the transposon end 14D is positioned at the 3 ’ end of the transferred strand 16D. In some examples, the transferred strand 16D further includes an index sequence (not shown) positioned between the first amplification domain 26 and the sequencing primer sequence 28D. The 5’ end functional group 20D and the sequencing primer sequence 28D may be any of the examples set forth herein, respectively, for the 5’ end functional group 20 A and the sequencing primer sequence 28A.

[0275] The non-transferred strands 18D of the transposome complexes 10D are made up of the strands ME’D and adapter segments 24 that create a forked adapter when hybridized to the transferred strands 16D. When the non-transferred strand 18D includes the adapter segments 24 and creates the forked adapter, it is to be understood that homodimers of this one type of transposome complex 10D are grafted to the anchoring layer 22, without homodimers of another type of transposome complex 10A or 10B or IOC. This is because the individual complex 10D includes both amplification domains 26, 38' to create fully adapted cDNA fragments, and thus the second complexes, e.g., 10B, IOC, with the second amplification domain 38 are not needed.As one example, the transposome complex 10D includes the transferred strand 16D as described herein with the first amplification domain 26 and the non-transferred strand 18D includes the strand ME’D, a complement of the sequencing primer sequence 28E, and a complement of the second amplification domain 38 (shown as 38’).

[0276] Flow?Cells

[0277] Each of the methods disclosed herein uses an example of a flow' cell 32. Fig. 19A depicts an example of the flow cell 32 from a top view, and different patterned structures 44A, 44B that may be included within individual flow channels 46 of the flow cell 32 are respectively shown in Fig. 19B and Fig. 19C.

[0278] In the examples shown in Fig. 19A and Fig. 19B. the flow cells 32 include a patterned structure 44A, 44B. While not shown, it is to be understood that a lid or a second patterned structure may be attached to the patterned structure 44A, 44B (e.g., at regions 58). Alternatively, the patterned structure 44 A or 44B is not bonded to another component, but rather, is open to the surrounding environment.

[0279] The patterned structure 44 A, 44B may include a single-layer substrate 50 or a multi-layer substrate 52.

[0280] Examples of suitable materials for the single-layer substrate 50 include epoxy siloxane, glass, modified or functionalized glass, polymeric materials (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, polytetrafluoroethylene (such as TEFLON® from Chemours). cyclic olefins / cyclo-olefin polymers (COP) (such as ZEONOR® from Zeon), polyimides, nylon (polyamides), etc.), ceramics / ceramic oxides, silica, fused silica, or silica-based materials, aluminum silicate, silicon and modified silicon (e.g., boron doped p+ silicon), silicon nitride (SisNr), silicon oxide (SiO2), tantalum pentoxide (TazOs) or other tantalum oxide(s) (TaOx), hafnium oxide (HIO2). carbon, metals, or the like.

[0281] When the single-layer substrate 50 is used, plurality of depressions 48 (Fig. 19B) or a single lane 54 (Fig. 19C) is defined at a surface of the single-layer substrate 50. Interstitial regions 56 surround each depression 48, and a perimeter region 58 surrounds the lane 54. The patterned structure 44A may also include a perimeter region 58. In these examples, the surface of the single-layer substrate 50 defines the interstitial regions 56 and / or the perimeter region 58.

[0282] Examples of the multi-layer substrate 52 include a base support 60 and a patterned material 62 positioned over the base support 60. The base support 60 may be any of the examples set forth herein for the single-layer substrate 50. The patterned material 62 may be any material that is capable of being patterned with the depressions 48 or the lane 54.

[0283] In an example, the patterned material 62 may be an inorganic oxide that is selectively applied to the base support 60, e.g., via vapor deposition, aerosol printing, or inkjet printing, in the desired pattern. Examples of suitable inorganic oxides include tantalum oxide (e.g., TazCh), aluminum oxide (e.g., AI2O3), silicon oxide (e.g., SiCh), hafnium oxide (e.g., HfCh), etc. In another example, the patterned material 62 may be a resin matrix material that is applied to the base support 60 and then patterned. Suitable deposition techniques include chemical vapor deposition, dip coating, dunk coating, spin coating, spray coating, puddle dispensing, ultrasonic spray coating, doctor blade coating, aerosol printing, screen printing, microcontact printing, etc. Suitable patterning techniques include photolithography, nanoimprint lithography (NIL), stamping techniques, embossing techniques, molding techniques, microetching techniques, printing techniques, etc. Some examples of suitable resins include a polyhedral oligomeric silsesquioxane-based resin, a non-polyhedral oligomeric silsesquioxane epoxy resin, a poly(ethylene glycol) resin, a polyether resin (e.g., ring opened epoxies). an acrylic resin, an acrylate resin, a methacrylate resin, an amorphous fluoropolymer resin (e.g., CYTOP® from Bellex), and combinations thereof.

[0284] In an example, the substrate 50 or 52 may be round and have a diameter ranging from about 2 mm to about 300 mm, or may be a rectangular, having its largest dimension up to about 10 feet (~ 3 meters). In an example, the substrate 50 or 52 maybe formed from a wafer having a diameter ranging from about 200 mm to about 300 mm. Wafers may subsequently be diced to form the individual substrate 50 or 52. In another example, the substrate 50 or 52 is a die having a width ranging from about 0. 1 mm to about 10 mm. While example dimensions have been provided, it is to be understood that a substrate 50 or 52 with any suitable dimensions may be used. For another example, a rectangular panel may be used, which has a greater surface area than a 300 mm round wafer. These panels may subsequently be diced to form individual substrates 50 or 52.

[0285] Each flow cell 32 also includes a flow channel 46 (Fig. 19A) or lane. The flow channel 46 may be an enclosed channel that is defined between the patterned structures 44A or 44B and a lid. In an alternate example, the flow channel 46 may be defined between two patterned structures 44A or 44B that are bonded together. In enclosed versions of the flow cell 32, a separate material (not shown) may attach the perimeter regions 58 of the patterned structure 44A or 44B to the lid or other patterned structure so that the separate material defines at least a portion of the w alls of the flow channel 46.

[0286] When the between the patterned structure 44A or 44B is open to the surrounding environment, the flow channel 46 may be defined by a lane in which the depressions 48 are formed, or by the lane 54.

[0287] The flow cell 32 shown in Fig. 19A includes eight flow channels 46. It is to be understood, however, that any example of the flow cell 32 may include any number of How channels 36 (e.g., one channel, four channels, etc.). With multiple channels 46, it is to be understood that each flow channel 46 may be isolated from each other flow channel 46 so that fluid introduced into any particular flow channel 46 does not flow into any adjacent flow channel 46. Separation may be obtained through the separate material, which can be applied at the perimeter of each flow channel 46 and at the perimeter of the entire flow cell 32.

[0288] The length and width of the flow channel 46 may be smaller, respectively, than the length and width of the patterned structure 44A or 44B so that a portion of thestructure surface surrounds the flow channel 46 and is available for attachment to another patterned structure or to the lid, or is available to define the perimeter of the open flow channel 46. In some instances, the width of each flow channel 46 can be at least about 1 mm, at least about 2.5 mm, at least about 5 mm, at least about 7 mm, at least about 10 mm, or more. In some instances, the length of each flow channel 46 can be at least about 10 mm, at least about 25 mm. at least about 50 mm, at least about 100 mm, or more. The width and / or length of each flow channel 46 can be greater than, less than or between the values specified above. In another example, the flow channel 46 is square (e.g., 10 mm x 10 mm).

[0289] The depth / height of each flow channel 46 can be as small as a few monolayers thick, for example, when microcontact, aerosol, or inkjet printing is used to deposit the separate material that partially defines the flow channel walls. In other examples, the depth / height of each flow channel 46 can be about 1 pm, about 10 pm, about 50 pm. about 100 pm, or more. In an example, the depth / height may range from about 10 pm to about 100 pm. In another example, the depth / height is about 5 pm or less. It is to be understood that the depth / height of each flow channel 46 can also be greater than, less than or between the values specified above. The depth / height of the flow channel 46 also varies along the length and width of the flow cell 32. e.g., because of the depressions 48.

[0290] Referring now to Fig. 19B, the patterned structure 44A includes the depressions 48, which are defined in the single-layer substrate 50 or in the patterned material 62 of the multi-layer substrate 52, and that are separated by interstitial regions 56. Many different layouts of the depressions 48 may be envisaged, including regular, repeating, and non-regular patterns. In an example, the depressions 48 are disposed in a hexagonal grid for close packing and improved density. Other layouts may include, for example, rectangular layouts, triangular layouts, and so forth. In some examples, the layout or pattern can be an x-y format in rows and columns. In some other examples, the layout or pattern can be a repeating arrangement of the depressions 48 and the interstitial regions 56.

[0291] The layout or pattern may be characterized with respect to the density (number) of the depressions 48 in a defined area. For example, the depressions 48 may be present at a density of approximately 2 million per mm2. The density may be tuned to different densities including, for example, a density of about 100 per mm2, about 1,000 per mm2, about 0.1 million per mm2, about 1 million per mm2, about 2 million per mm2, about 5 million per mm2, about 10 million per mm2, about 50 million per mm2, or more, or less. It is to be further understood that the density can be between one of the lower values and one of the upper values selected from the ranges above, or that other densities (outside of the given ranges) may be used. As examples, a high density7array may be characterized as having depressions 48 separated by less than about 100 nm, a medium density array may be characterized as having the depressions 48 separated by about 400 nm to about 1 pm, and a low density array may be characterized as having the depressions 48 separated by greater than about 1 pm.

[0292] The layout or pattern of the depressions 48 may also or alternatively be characterized in terms of the average pitch, or the spacing from the center of one depression 48 to the center of an adjacent depression 48 (center-to-center spacing) or from the right edge of one depression 48 to the left edge of an adjacent depression 48 (edge-to-edge spacing). The pattern can be regular, such that the coefficient of variation around the average pitch is small, or the pattern can be non-regular in which case the coefficient of variation can be relatively large. In either case, the average pitch can be, for example, about 50 nm, about 0.1 pm, about 0.5 pm, about 1 pm, about 5 pm, about 10 pm, about 100 pm, or more or less. The average pitch for a particular pattern can be between one of the lower values and one of the upper values selected from the ranges above. In an example, the depressions 48 have a pitch (center-to-center spacing) of about 1.5 pm. While example average pitch values have been provided, it is to be understood that other average pitch values may be used.

[0293] The size of each depression 48 may be characterized by its volume, opening area, depth, and / or diameter or length and width. For example, the volume can range from about I O3pm3to about 100 pm3, e.g., about I O2pm3, about 0.1 pm3, about 1 pm3, about 10 pm3, or more, or less. For another example, the opening area can range from about 1 x 103pm2to about 100 pm2, e. g.. about P I ()2pm2, about 0.1 pm2, about1 pm2, at least about 10 pm2. or more, or less. For still another example, the depth can range from about 0.1 pm to about 100 pm, e.g., about 0.5 pm. about 1 pm. about 10 pm, or more, or less. For yet another example, the diameter or each of the length and width can range from about 0.1 pm to about 100 pm, e.g., about 0.5 pm, about 1 pm, about 10 pm, or more, or less.

[0294] In the patterned structure 44A, the anchoring layer 22 is positioned within each of the depressions 48. The anchoring layer 22 is selected so that it can bind with a 5’ end of each of the primers 34, 36, and in some instances, the functional groups 20 A, 20B or 20 A, 42, or 20D of the transposome complexes 10A. 10B or 10A, 10C, or 10D.

[0295] In one example, the anchoring layer 22 is streptavidin. In another example, the anchoring layer 22 is a biotinylated polymeric hydrogel. In one example, the biotinylated polymeric hydrogel is a copolymer including a first recurring unit of formula (I):wherein:R1is selected from the group consisting of -H, a halogen, an alkyl, an alkoxy, an alkenyl, an alkynyl. a cycloalkyl, an aryl, a heteroaryl, a heterocycle, and optionally substituted variants thereof;R2is selected from the group consisting of an azide and a tetrazine; each (CH2)Pcan be optionally substituted; and p is an integer from 1 to 50;

[0296] and a second recurring unit of formula (II): wherein: each of R3, R3, R4, R4is independently selected from the group consisting of -H, R5, - OR5, -C(O)OR5, -C(O)R5, -OC(O)R5, -C(O)NR6R7, and -NR6R7; R5is selected from the group consisting of -H, -OH, an alkyl, a cycloalkyl, a hydroxyalkyl, an aryl, a heteroaryl, a heterocycle, and optionally substituted variants thereof; and each of R6and R7is independently selected from the group consisting of -H and an alkyl. The biotin of the biotinylated polymeric hydrogel is attached to a linker, such as bicyclo[6.1.0]nonyne (BCN), which can covalently attach to the R2groups. In still another example, the anchoring layer 22 is the polymeric hydrogel without the additional biotin groups. In these instances, the R2groups can be selected to covalently attach to the 5’ end of each of the primers 34, 36, and in some instances, the functional groups 20 A, 20B or 20A, 42, or 20D of the transposome complexes 10A, 10B or 10A, IOC, or 10D.

[0297] The patterned structure 44A includes the primers 34, 36 attached to the anchoring layer 22 within each of the depressions 48. The primers 34, 36 respectively have the same sequence as the first and second amplification domains 26, 38, or the primers 34 have the same sequence as the first amplification domain 26 and the primers 36 have the complementary sequence to the second amplification domain 38’. Any of the sequences set forth herein for the amplification domains 26, 38 may be used for the primers 34, 36.

[0298] As show n in Fig. 19B, the patterned structure 44A also includes any of the transposome complexes 10 (e.g., in dimer form) attached to the anchoring layer 22 within each of the depressions 48. In other examples, the transposome complexes 10 may be attached to the interstitial regions 56 rather than in the depressions 48. In still other examples, the transposome complexes 10 may be attached to the interstitial regions 56 and in the depressions 48. Thus, some examples of the patterned structure44A (and the flow cell 32) include the substrate 50 or 52 having depressions 48 separated by interstitial regions 56, the anchoring layer 22 positioned in each of the depressions 48, the first and second primers 34, 36 attached to the anchoring layer 22, and the first and second transposome complexes 10A, 10B or 10A, IOC attached to the anchoring layer 22, to the interstitial regions 56, or to both the anchoring layer 22 and the interstitial regions 56. Similarly, other examples of the paterned structure 44A (and the flow cell 32) include the substrate 50 or 52 having depressions 48 separated by interstitial regions 56, the anchoring layer 22 positioned in each of the depressions 48, the first and second primers 34, 36 atached to the anchoring layer 22, and a plurality of the transposome complexes 10D attached to the anchoring layer 22, to the interstitial regions 56. or to both the anchoring layer 22 and the interstitial regions 56.

[0299] Referring now to Fig. 19C, the paterned structure 44B includes the lane 54, which is defined in the single-layer substrate 50 or in the paterned material 62 of the multi-layer substrate 52, and is surrounded by the perimeter regions 58. The depth of lane 54 is large enough to house at least the anchoring layer 22. In one example, the lane 54 may be filled with the anchoring layer 22. In an example, the depth may be at least about 0.1 pm, at least about 0.5 pm, at least about 1 pm, at least about 10 pm, at least about 100 pm, or more. Altematively or additionally, the depth can be at most about 1 x 103pm, at most about 100 pm, at most about 10 pm, or less. In some examples, the depth is about 0.4 pm. The depth of the lane 56, 56’ can be greater than, less than or between the values specified above.

[0300] As shown in Fig. 19C, the paterned structure 44B includes the lane 54, the anchoring layer 22 positioned in the lane 54, and the primers 34, 36 and the transposome complexes 10A, 10B or 10A, 10C or 10D atached to the anchoring layer 22.

[0301] Referring now to Fig. 19D, another example of the architecture within a channel 46 of the flow- cell 32 is depicted. This architecture is suitable for depleting ribosomal RNA (rRNA) from a sample containing RNA before the sample is transported to another change 46 that includes surface chemistry for tagmentation, amplification, and sequencing.

[0302] As shown in Fig. 19D, this example of the architecture includes a flow cell lane 54', the anchoring layer 22 positioned in the lane 54’, and surface bound DNA probes 66 that are at least partially complementary to rRNA strands in the sample containing RNA. When the sample containing RNA is introduced into this example of the architecture, the rRNA strands in the sample respectively hybridize to the DNA probes 66.

[0303] Examples of the rRNA strands that are in high-abundance and for which the DNA probes 66 can be designed are disclosed in WO2021 / 127191 and WO 2020 / 132304, each of which is incorporated by reference herein in its entirety.

[0304] Each DNA probe 66 is from 10 nucleotides to 100 nucleotides long. As other examples, the DNA probes 66 may each range from 20 nucleotides to 100 nucleotides long, from 20 nucleotides to 80 nucleotides long, from 40 nucleotides to 60 nucleotides long, or from 45 nucleotides to 55 nucleotides long. In one specific examples, the DNA probes 66 are each 50 nucleotides long. In some embodiments, the at least one immobilized oligonucleotide is 45-55 bases in length.

[0305] Examples of suitable sequences for the DNA probes 66 include those (i.e.. SEQ. ID. Nos. 1-1131) in Table 2 of WO2023 / 056328, which is incorporated by reference herein in its entirety.

[0306] While the DNA probes 66 described herein are at least partially complementary to rRNA strands, it is to be understood that the DNA probes 66 may be designed to hybridize any off-target or unwanted RNA molecule. The off-target or unwanted RNA molecule may be any RNA sequence that a user does not want to analyze.

[0307] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in the description herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology7used herein is for the purpose of description and should not be regarded aslimiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or "an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0308] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.

[0309] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.

[0310] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departing from its scope. While the dimensions, typesof materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms '’including" and “in which" are used as the plain-English equivalents of the respective terms ’’comprising" and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U.S.C. §1 12(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.

[0311] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.

Claims

What is claimed is:1 . A method of single-cell analyte detection, the method comprising the steps of: providing a sample comprising a plurality of cells, each cell of the plurality7of cells comprising sample nucleic acids; isolating the plurality of cells into partitions such that an individual partition comprises: only one individual cell of the plurality of cells and a solid support comprising linked capture oligonucleotides; for the individual partition: capturing the sample nucleic acids of the individual cell onto the solid support using capture sequences of the capture oligonucleotides; extending the capture sequences to form duplexes with the sample nucleic acids; and exposing the duplexes of the partitions to surface-linked transposases of a flow cell surface to generate fragments of a sequencing library; generating sequencing data from the sequencing library, wherein individual sequence reads of the sequencing data are associated with different locations on the flow cell surface: and using the different locations on the flow cell surface to identity linked sequence reads of the sequencing data.

2. The method of claim 1, wherein the exposing comprises: loading the partitions onto the flow cell; and disrupting the partitions subsequent to the loading.

3. The method of claim 1, wherein the exposing comprises: disrupting the partitions; and loading contents of the disrupted partitions onto the flow cell.

4. The method of claim 2 or 3, wherein the contents of the partitions further comprise protein analytes.

5. The method of claim 2 or 3, wherein the partitions comprise droplets.

6. The method of claim 1, comprising lysing the individual cell within the individual partition to capture the sample nucleic acids on the solid support.

7. The method of claim 1, wherein the extending step reverse transcribes a sample nucleic acid into a cDNA.

8. The method of claim 5. wherein the transposase cuts the cDNA at a random cut site.

9. The method of claim 6, wherein a unique identifier sequence for each fragment of the sequencing library is provided by a segment of the cDNA adjacent the random cut site.

10. The method of claim 7, further comprising mapping the sequence reads to genes in a reference, and collapsing reads that include the same unique identifier sequence.

11. The method of claim 1 , wherein the solid support is a bead and wherein the capture oligonucleotides of the bead comprise a same cellular barcode sequence.

12. The method of claim 1, wherein the solid support is a bead and wherein the capture sequences comprise a polyT sequence.

13. The method of claim 1, wherein the solid support is a bead and wherein the capture sequences comprise gene-specific sequences.

14. The method of claim 1, wherein the solid support is a bead and wherein the bead further comprises aptamers configured to capture protein analytes.

15. The method of claim 14, further comprising: contacting the solid support with one or more reporter probes configured to bind to the aptamer, wherein the one or more reporter probes comprises aptamer binding sequences; washing unbound reporter probes from the solid support; and amplifying the one or more reporter probes to generate additional fragments of the library.

16. The method of claim 1, comprising amplifying the fragments on the flow cell surface to generate clusters and generate the sequencing data from sequencing one or more strands of the clusters.

17. The method of claim 1 , wherein using the different locations on the flow cell surface to identify' linked sequence reads of the sequencing data comprises obtaining location information for the clusters on the flow cell surface.

18. The method of claim 1, comprising assigning linked sequence reads to a target sequence region or mRNA transcript using the location information.

19. The method of claim 1, wherein the plurality of partitions is formed substantially simultaneously by shearing or vortexing a vessel comprising an aqueous phase, an immiscible phase, oligo-linked beads, and the plurality of cells.

20. The method of claim 1, wherein the capture oligonucleotides comprise one or more sequences identical to or complementary to adaptor sequences linked to the flow cell surface.

21. The method of claim 1. comprising: aligning the sequence reads to a reference; and determining a read count of aligned sequence reads per region of the reference.

22. The method of claim 21, comprising providing a report with transcription levels of genes in the single cells based on the read count per region.

23. The method of claim 1, wherein the flow cell surface comprises a plurality of lanes, such that a flow cell comprising the flow cell surface is loaded in a one cell per lane basis.

24. A kit for performing the method of claim 1.

25. A device for performing the method of claim 1.

26. A method of single-cell analyte detection, the method comprising the steps of: providing a sample comprising a plurality of cells, each cell of the plurality7of cells comprising sample nucleic acids; isolating the plurality of cells into partitions such that an individual partition comprises: only one individual cell of the plurality of cells and a solid support comprising linked capture oligonucleotides; for the individual partition: capturing the sample nucleic acids of the individual cell onto the solid support using capture sequences of the capture oligonucleotides, the capture sequences of each individual partition comprising a unique identifier that is associated with the individual partition and distinguishable from other partitions of the partitions; extending the capture sequences to form duplexes with the sample nucleic acids, wherein the duplexes comprise a sequence of the unique identifier; amplifying the duplexes of the partitions using a primer pair, wherein one primer of the primer pair comprises a first adaptor sequence such that amplification products of the duplexes comprise the first adaptor sequence and its complement at one end of an individual amplification product; andcontacting the amplification products with surface-linked transposome complexes of a flow cell surface to generate fragments of a sequencing library, the surface-linked transposome complexes comprising: a transposase; and a transposon comprising a 3' portion comprising a first transposon end sequence and a second adaptor sequence at the 5' end of the first transposon end sequence; a second transposon comprising a second transposon end sequence complementary to at least a portion of the first transposon end sequence, wherein generating the fragments of the sequencing library comprises incorporating the second adaptor sequence into the amplification products such that the fragments comprise the first adaptor sequence and its complement at one end of an individual fragment and the second adaptor sequence and its complement at the other end of the individual fragment; and generating sequencing data from the sequencing library, wherein individual sequencing reads comprise respective unique identifiers from respective partitions.

27. The method of claim 26, wherein the sample nucleic acids comprise messenger RNAs.

28. The method of claim 26 or claim 27, comprising: disrupting the partitions before the amplifying.

29. The method of any of claims 26-28, comprising: disrupting the partitions before the extending.

30. The method of any of claims 26-29, wherein the partitions comprise droplets.

31. The method of any of claims 26-30, comprising lysing the individual cell within the individual partition to capture the sample nucleic acids on the solid support.

32. The method of any of claims 26-31, wherein the extending step reverse transcribes a sample nucleic acid into a cDNA.

33. The method of any of claims 26-32, comprising using sequencing reads comprising the unique identifier and aligned to a same sequence of interest to determine an expression level of the sequence of interest in the individual cell.

34. The method of any of claims 26-33, wherein the flow cell surface does not comprise transposome complexes having the first adaptor sequence such that the first adaptor sequence is only provided amplification and the second adaptor sequence is only provided via the surface-linked transposome complexes.

35. The method of any of claims 26-34, wherein the transposome complex is a homodimer.

36. The method of any of claims 26-34, wherein the first adaptor sequence comprises a sample-specific index sequence.

37. The method of claim 36, wherein the sequencing library is single-indexed.

38. The method of any of claims 26-37, wherein the amplification comprises 10 or fewer or 5 or fewer amplification cycles.

39. The method of any of claims 26-37, comprising determining a number of sequencing reads per individual cell based on the unique identifier.

40. The method of claim 39, comprising normalizing sequencing reads for individual cells across the sample.

41. The method of any of claims 26-40, wherein the transposase cuts the amplification products at random cut sites.

42. The method of claim 41, wherein an internal identifier for each fragment of the sequencing library is provided by a segment of an amplification product adjacent a random cut site.