Molecular deduplication analysis methods

Intrinsic molecular identifiers and binning indexes in nucleic acid sequencing techniques address duplication issues in mRNA measurement, ensuring accurate gene expression analysis.

WO2026059922A1PCT designated stage Publication Date: 2026-03-19ILLUMINA INC
View PDF 20 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing methods for measuring mRNA levels in cells are inaccurate and inefficient due to duplication of sequence reads during sample preparation, leading to biased expression level estimates.

Method used

The use of intrinsic molecular identifiers (IMIs) and binning indexes during nucleic acid sequencing to uniquely identify and correct for sequence read duplicates, combined with optional molecular diversity enhancers (MDEs), to accurately quantify mRNA expression levels.

Benefits of technology

Provides accurate and efficient measurement of mRNA levels by correcting for duplication biases, enabling precise gene expression analysis at the single-cell level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045602_19032026_PF_FP_ABST
    Figure US2025045602_19032026_PF_FP_ABST
Patent Text Reader

Abstract

The presently described techniques provide methods for measuring quantities of mRNA transcripts (311, 1313) present in a sample, where sequence information for each molecule is read from what is essentially a random start site within that molecule and in which a short binning index (e.g., about 3 bases) (316, 1316) is added (step 115) to the sequence information. The binning index (316, 1316) is useful to resolve any bias arising with the use of the intrinsic sequences to uniquely identify and count the molecules.
Need to check novelty before this filing date? Find Prior Art

Description

MOLECULAR DEDUPLICATION ANALYSIS METHODS

[0001] REFERENCE TO ELECTRONIC SEQUENCE LISTING The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on September 8, 2025, is named “ILUM0197PCT.xml” and is 13,949 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety. Technical Field

[0002] The invention relates generally to the quantitative detection and analysis of molecules in a sample. Background

[0003] Aspects of the present disclosure relate generally to devices, systems, and methods providing biological or chemical analysis. Various protocols in biological or chemical research involve performing a large number of controlled reactions on local support surfaces or within predefined reaction chambers. The designated reactions may then be observed or detected, and subsequent analysis may help identify or reveal properties of chemicals involved in the reaction. For example, in some multiplex assays, an unknown analyte having an identifiable label (e.g., fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited into a corresponding well of a flow cell channel. Observing any chemical reactions that occur between the known probes and the unknown analyte within the wells may help identify or reveal properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis (SBS) or cyclic-array sequencing.

[0004] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventor(s) has made or used the devices and techniques described herein.

[0005] With the preceding in mind, it may be appreciated that living organisms store genetic information in DNA. Genes in the coding regions of DNA are transcribed into messenger RNA (mRNA), which is translated into protein. Proteins play critical functional and structural roles in living organisms. For example, most enzymes are proteins, and those enzymes catalyze the metabolic reactions essential to life. It is also enzymes that copy DNA into mRNA. Proteins are also structural, and constitute the essential fibers of muscles, the predominant material of hair, as well as basic structural linkages within the cytoskeleton. Essentially, all such proteins are made by translating an mRNA into the protein. In fact, one mRNA can serve as the template for synthesizing multiple copies of a protein.

[0006] Because cells need to change in response to different conditions, and need different proteins at different times, it is helpful if any given mRNA is short-lived. Most mRNA molecules have a lifetime measured in seconds or minutes. Nevertheless, the health of a cell, or its response to a pathogen, to a drug, or to age-specific developmental changes may be indicated by the quantities of mRNA molecules present in a given cell or cell type. As a consequence, it may be useful to be able to measure levels of different mRNA transcripts present in cells, cell types, or tissue in an accurate and rapid manner. Summary

[0007] The presently described techniques provide methods for measuring mRNA in a biological sample. Such methods include performing nucleic acid sequencing to obtain sequence information and quantitation of mRNA in the sample. According to such approaches, sample preparation and sequencing are performed so that the sequence information for each mRNA (or its corresponding cDNA) is read from what is essentially a random start site within the molecule. As used herein, such a random start site may be understood to be “random” with respect to a whole or larger sequence corresponding to or derived from an mRNA or cDNA strand or reference. That is, a “random start site” in this context may be understood to not necessarily correspond to a sequencing operation that starts at a random location along an intact strand but also to a sequencing operation that starts at an end of a fragment of a larger or whole strand that has undergone random fragmentation as part of the technique (e.g., via enzymatic fragmentation) or via a naturally occurring process (e.g., damage associated with a disease or disorder causing fragmentation orcleavage of a larger strand or sequence and which may be indicative of spatial information related to the sample or subject). Thus, as used herein a “random start site” does not necessarily imply that the initiation of the sequencing operation occurs randomly with respect to a whole or intact sequence, but that the whole or intact sequence may undergo some form of naturally occurring or induced fragmentation at random locations prior to sequencing such that the fragments, when sequenced, are effectively random in terms of their start location with respect to the whole or intact sequence.

[0008] In addition, as discussed herein, a short binning index (BI) (e.g., about 3-6 bases) may be added to the sequence information. Because the sequencing start site for each molecule is essentially random, at least some of the bases in the sequencing information are essentially random and are thus also effectively unique to a specific mRNA, thereby functioning as a molecular identifier. As used herein, a molecular identifier may be understood to include bases present in the naturally occurring sequence or its complement (i.e., intrinsic molecular identifiers (IMI)), to include bases added or otherwise not present in the naturally occurring sequence or its complement (i.e., extrinsic molecular identifiers), or to a combination of intrinsic and extrinsic molecular identifiers. With this in mind, such a molecular identifier may be understood to comprise a unique (or functionally unique) sub-sequence in a source nucleic acid molecule. One or more such molecular identifiers may, alone or in conjunction with other information, uniquely identify a source nucleic acid molecule. Thus, in the context of intrinsic molecular identifiers, bases that are naturally and intrinsically present in each molecule are used to associate sequence data with the molecule from which that sequence data was read. Counts of unique sequence reads (based on IMI) are indexed by their binning indexes (BI). The presently described techniques, among other disclosed concepts, provide corrections that apply to sequence read counts to correct for bias potentially introduced during sample preparation.

[0009] In particular, various sequencing techniques involve making an abundant number of copies of nucleic acid and generating a corresponding number of sequence reads from those copies. Using embodiments of the presently disclosed techniques, each sequence read includes an intrinsic molecular identifier that associates a given read with one original molecule. The intrinsic molecular identifier is unique for the purpose of associating such a given read with the original molecule because it is adjacent a random cleavage or break site within the original molecule or acomplementary copy of the original molecule and proximate to sequencing initiation location (hence, as used herein, a “random start site”). Thus, even when a sample has numerous transcripts that would otherwise appear to be exact duplicates of each other, and would otherwise generate identical sequence reads, by starting the sequencing at random start sites for each original molecule (such as in based on random cleavage or breakage of with respect to an original or reference molecule), the reads from otherwise identical transcripts are effectively unique at least because those reads began at different places within those transcripts. Thus, some portion of each sequence read (e.g., the first 10 to 20 bases or so) may be used to identify the transcript, from which the read was derived. As noted above and used herein, the identifying portion of the sequence read, and the corresponding portion of the molecule from which the sequence read came, is referred to as an intrinsic molecular identifier. After obtaining sequence information by methods and techniques as discussed herein, any portions of the sequence information (e.g., any sequence reads) that are identical and include the same intrinsic molecular identifier, are considered duplicate sequences taken from the same original molecule, or transcript, in the sample. A count of unique (i.e., deduplicated) sequences from the sample provides a count of the molecules present in the sample. For example, if there are 10,000 transcripts from gene 1 and 7 transcripts from gene 2, even after exponential amplification, those likely millions of sequence reads are deduplicated using the intrinsic molecular identifiers to yield 10,000 deduplicated sequence reads of gene 1 and 7 reads of gene 2.

[0010] Certain implementations of the presently described techniques further add a small segment of extrinsic bases, the aforementioned “binning index”, to the molecules during preparation for sequencing. The added bases appear in the sequence information and are used as an index in such implementations. Such binning indexes may be useful when the sequence reads are mapped to reference information to identify genes and are assigned to bins according to the gene, to intrinsic molecular identifiers, and to other optional barcode information, such as a cell-specific “cellular barcode” that may be used when methods of the presently described techniques are applied to single-cell RNA sequencing (scRNA-Seq). The binning index, which may refer to both the small segment of bases added to each molecule during sample preparation and also to the corresponding segment of base information in each sequence read, is a useful informational tool for assigningcounts of deduplicated sequence reads to bins and correcting those counts to adjust for bias that may arise during sample preparation.

[0011] For example, if target molecules undergo limited amplification (e.g., three or four rounds of polymerase chain reaction) prior to fragmentation, there may be over-representation of those molecules in sequence reads counts. In such cases, the counts can be divided by a correction factor proportional to the expected amplification by PCR. In another example, if a transcript is short or very highly expressed (e.g., millions of copies in a cell), then even random fragmentation will generate some identical cut sites, yielding some limited number of identical intrinsic molecular identifiers. Such duplicate intrinsic molecular identifiers may lead to under-representation of the transcripts in the sequence read counts. In those instances, the sequence read counts can be multiplied by a correction factor (e.g., that has been derived experimentally) to provide an accurate measure of expression levels in a cell. Thus, the binning index, which may be applied by a capture oligonucleotide or adaptor added and used during sample preparation, in combination with the intrinsic molecular identifiers, provides for the accurate measurement of expression levels in cells.

[0012] A transcript may further be labeled with a short oligonucleotide tag referred to herein as a molecular diversity enhancer (MDE). An MDE may be two, three, four, or so bases and may be used to ensure unduplicated intrinsic molecular identifiers (IMIs). In practice the MDE may not be, by itself, long enough to function as a molecule-specific barcode or unique molecular identifier. That role is, as described herein, instead performed by the IMI. However the MDE may be added to supplement the information of t. IMI, such as in the event the IMI is based on a non-unique identifier due to a fragmentation or breakage site shared by two original molecules, and thereby a shared sequencing start site. As used herein, the MDE (which is optional) supplements the IMI to ensure that each molecule is uniquely labeled. As discussed herein, the binning index plays a different role and is introduced as a molecule and then used during bioinformatics to characterize or sort sequence read counts into bins, where each bin may be specifically associated with a cell (via cell barcode introduced during sample preparation), a gene (identified by mapping sequence reads to reference information), and a molecule (via the IMI and optional MDE). Read counts are collected based on binning index (i.e., bins), allowing those counts to be corrected to adjust for bias that may be introduced during sample preparation. By those means, methods of the presentlydescribed techniques are useful for quantifying expression levels of single cells, including for single cells that have been isolated, such as in aqueous partitions (e.g., droplets or wells of a plate).

[0013] With respect to the correction of read counts as discussed herein, in certain embodiments a correction approach is taught which is dynamic in nature in that the correction derived or applied is sample specific (e.g., based upon cell type, level of gene expression, sequencing depth, and / or other factors specific to a given run or sample). In certain described implementations, the correction is based at least in part on, or otherwise corrects for, inflation attributable to whole transcriptome amplification (WTA). By way of example, a correction determination routine, in one embodiment, may be used to derive a correction that may be used to adjust the observed molecule counts within a given experimental dataset (i.e., for a respective sample), thus adjusting counts dynamically based upon the observed level of WTA-based inflation (e.g., amplification- based inflation). In one such embodiment, the analytical workflow begins with read mapping to identify genes or other chromosome regions of interest and to enable all reads to be segregated by relevant identifier (e.g., gene identity) in an output file (e.g., a BAM format file). In such an embodiment, within the same cell barcode and gene, the output file is parsed to identify IMIs without regard for the BI. IMIs are segregated into a number of bins (e.g., 64 bins) determined based on the length of the BI sequence (e.g., 3-bases) in each read. Within each bin, the number of unique IMIs is determined by collapsing duplicate IMIs into a single count. This step mitigates error attributable to obtaining identical IMIs from distinct transcripts from the same gene as well as improving computational and process efficiency.

[0014] In one such implementation, after obtaining counts of unique IMIs in each bin, a subset of the barcode+gene combinations (BGCs) may be used to estimate the inflation distribution, i.e., the probabilities of a single parent molecule producing any number of copies from one to x during WTA. This may be done by selecting a subset of BGCs in which the number of unique IMIs is lower than or equal to x (e.g., ≤ 15 in one embodiment). In certain embodiments a routine may be employed to derive the inflation probabilities based on observed IMI counts. BGCs where the number of molecules can be determined unequivocally may be used to estimate the true number of IMIs that can be attributed to inflation from a single molecule in a single bin. Finally, each of the x counts in single bins is divided by the sum of counts to produce the probability of one molecule producing the observed number of IMIs.

[0015] Once the inflation probabilities are established in such an approach, a correction factor can be selected. In one implementation, the correction factor is an integer and may be chosen such that no more than 1% of molecules will be over-counted, i.e., the cumulative inflation probability is at least 0.99. Within each BGC, the number of unique IMIs in each bin is divided by the correction factor and rounded up to the nearest integer. The divided counts are then summed across bins to produce the final count for this BGC, which may be written into the raw count matrix.

[0016] While the preceding approach may be referred to herein as a dynamic IMI correction or dynamic correction factor approach, other techniques are also discussed and disclosed herein. By way of example, a second approach discussed herein may be described or referred to as an average IMI per molecule (IPM) correction. As discussed herein, such an approach may be based on estimating the average IMIs from a single parent molecule. This average IPM value may then be used to correct IMI counts in bins to molecule counts. Such an approach may be more computationally efficient and faster than other techniques. In both approaches, however, a correction is estimated that reflects the level of IMI inflation across the entire sample and, further, a correction is applied to each BGC to estimate molecular counts from observed IMI counts.

[0017] In certain aspects, the embodiments of the presently described techniques relate to methods for measuring gene expression, including at the single cell level of granularity. Certain of the described methods include sequencing nucleic acids (e.g., mRNA or cDNA) from random start sites (e.g., after random fragmentation by chemical or enzymatic agents or after natural breakage due to a disease or disorder) of the nucleic acids to generate sequence reads having an effectively unique portion (e.g., an intrinsic molecular identifier (IMI)), attaching a binning index (BI) to the sequence reads, and mapping each sequence read to a gene or genomic region. Methods include determining counts of the unique portions per gene or genomic region, assigning the counts to associated binning indexes, and correcting counts to reduce bias introduced during sample preparation. By summing corrected counts across the binning indexes for each gene or genomic region, methods of the presently described techniques provide an estimated number of the transcripts per gene or genomic region in the sample, and correspondingly one or more measures of gene expression for the sample, including in some implementations measures of gene expression at the single cell level of granularity.

[0018] The binning index may comprise, in certain embodiments, six or fewer bases, and in certain implementations comprises three or fewer bases. The sample preparation in certain such embodiments may include fragmenting mRNA at random sites (such as by the application of chemical or enzymatic agents), annealing oligonucleotides to the fragments, and extending the oligonucleotides to make cDNA. In other embodiments, the sample preparation includes annealing oligonucleotides to the mRNA, extending the oligonucleotides to make cDNA copies of the transcripts, and fragmenting the cDNA copies at random sites. The correcting step may be performed to account for a probability of the random sites, from which sequencing is initiated, being duplicated among the transcripts. In this manner, the correction factor may account for a probability of multiple random start sites per transcript within the sample.

[0019] In certain embodiments, cDNA is prepared by capturing the transcripts with capture oligonucleotides linked to beads (e.g., capture beads), wherein each capture oligonucleotide (also referred to as an “oligo” herein) includes a 5'-linkage to a respective bead, a cell barcode, a binning index, and an annealing primer section-3'. The annealing primer section may include a poly-T region and random segment (e.g., a hexamer), or a gene-specific primer.

[0020] In certain implementations, the binning index is variable in length to improve sequencing quality. For example, a bead decorated with capture oligonucleotides may have a mixture of binning indexes with some being 3 bases, some 2, some only one, and some oligonucleotides having no binning index (i.e., zero bases long). In some embodiments the capture oligonucleotides have a conserved sequence 3' of the mixed-length binning indexes, which may be useful to identify the start sites in the molecules in the sequence read data. For such embodiments, a first portion of the capture oligonucleotides linked to the beads include no binning index and a second portion of the capture oligonucleotides linked to the beads each include a binning index that each independently consists of 1, 2, or 3 bases. Due to the variable length binning indexes, sequencing library material (e.g., amplicons) attached to the flow cell of a sequencer may be sequenced out of phase with each other. Methods of the presently described techniques improve the quality of sequence data by avoiding problems of sequencing conserved molecules “in phase”, or in lock- step with each other.

[0021] Methods of the presently described techniques may include isolating a cell with one of the beads in an aqueous partition and lysing the cell to release the transcripts within the partition. A plurality of cells may be isolated into droplets, e.g., either in serial fashion using channels of a microfluidic platform or simultaneously by mixing cells with beads in water under oil and vortexing or shearing the mixture to generate the partitions (droplets). The beads may be decorated with capture oligonucleotides. All capture oligonucleotides on one bead may have a common barcode, which can serve as a cellular barcode. Each capture oligonucleotide, in certain embodiments, includes the binning index and a primer segment at the 3' end that anneals to target templates. After sample preparation and sequencing, the binning index appears in sequence reads.

[0022] The sequence reads include the aforementioned effectively unique portions (e.g., the IMIs), and the method includes determining counts of the IMIs per genomic region and assigning the counts to associated binning indexes. Those assignments may be performed by storing (e.g., writing in memory) one or more files that include the counts indexed by the binning indexes. In certain implementations, after the assigning step, for each gene, only the binning indexes, the counts, and the correction factor are used for the applying and summing step. A correction factor, as discussed herein, is applied to the counts to reduce bias introduced during sample preparation, which may be an empirically derived measure of over- or under-representation of unique molecules when relying on IMIs to uniquely label molecules. The correction factor may be a divisor that reduces the counts by an expected factor if the molecules were subject to limited amplification prior to priming or fragmenting to generate the IMIs.

[0023] In related aspects, the presently described techniques provide a method of measuring expression levels. The method includes sequencing transcripts from random start sites, as discussed herein, to generate sequence reads, wherein each sequence read includes a binning index added by an oligonucleotide during sample preparation. Each sequence read is mapped to a gene, and the method includes obtaining counts per gene of unique intrinsic sequences (e.g., IMIs) defined by the random start sites; assigning the counts to associated binning indexes; and applying a correction factor to each count, to correct for bias introduced in the sample preparation. The corrected counts are summed across the binning indexes for each gene to provide an estimated number of the transcripts per gene in the sample.

[0024] Sample preparation, as used herein, may include (i) fragmenting the transcripts at random sites, annealing oligonucleotides to the fragments, and extending the oligonucleotides to make cDNA copies of the transcripts; or (ii) annealing oligonucleotides to the transcripts, extending the oligonucleotides to make cDNA copies of the transcripts, and fragmenting the cDNA copies at random sites. The correction factor may adjust for a probability of the random sites, from which sequencing is initiated or otherwise started, being duplicated among the transcripts, may adjust for a probability of multiple random start sites per transcript within the sample, or both. That is, the correction factor may include an estimate of a number of copies of each transcript resulting from the amplification and the applying step may include dividing each count by the correction factor.

[0025] Processor-implementable aspects of the various methods and techniques discussed herein may be embodied and implemented as processor-executable code or stored routines, such as may be stored on a processor-accessible memory and executed via one or more hardware or virtual processors or dedicated circuitry. As such, it should be understood that certain actions or steps described as part of a method or process herein may be implemented as processor executable code, routines, or algorithms stored on one or more tangible computer-readable media. Brief Description of the Drawings

[0026] FIG.1 shows a method for indexing transcripts with binning indexes, in accordance with aspects of the present disclosure.

[0027] FIG. 2 shows beads decorated with capture oligos that include binning indexes, in accordance with aspects of the present disclosure.

[0028] FIG.3 diagrams RNA capture and library preparation with binning indexes, in accordance with aspects of the present disclosure.

[0029] FIG. 4 shows the generation of double-stranded cDNA with template switching oligos (TSOs) and tagmentation of double-stranded cDNA with transposons, in accordance with aspects of the present disclosure.

[0030] FIG.5 shows the use of IMIs using capture oligos with binning indexes, in accordance with aspects of the present disclosure.

[0031] FIG.6 illustrates a workflow within a system performing aspects of the presently described techniques, in accordance with aspects of the present disclosure.

[0032] FIGS.7A (SEQ ID NOS: 1 – 15) and 7B in combination depict a sample process flow for estimating an inflation factor based on amplification-based inflation and selecting a correction factor, in accordance with aspects of the present disclosure.

[0033] FIG.8 depicts a sample process flow for estimating inflation probabilities based on sample- specific observed count data, in accordance with aspects of the present disclosure.

[0034] FIG.9 depicts a further view of a sample process flow for estimating inflation probabilities based on sample-specific observed count data, in accordance with aspects of the present disclosure.

[0035] FIG.10 depicts in tabular form aspects of an IPM calculation strategy, in accordance with aspects of the present disclosure.

[0036] FIG.11 depicts an example of a correction process based on average IPM, in accordance with aspects of the present disclosure.

[0037] FIG. 12 depicts a further example of a correction process based on average IPM, in accordance with aspects of the present disclosure.

[0038] FIG. 13 depicts a comparison of correction results based on a dynamic correction factor approach, a worst-case correction factor approach and an average IPM correction approach relative to a UMI reference, in accordance with aspects of the present disclosure.

[0039] FIG.14 depicts a comparison of metrics derived for a dynamic correction factor approach, a worst-case correction factor approach and an average IPM correction approach, in accordance with aspects of the present disclosure.

[0040] FIG.15 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure. Detailed description

[0041] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components and / or software modules. Thus, for example, one or more of the functional blocks (e.g., processors or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or random access memory, hard disk, or the like). Similarly, the programs may be stand-alone programs, may be incorporated as subroutines in an operating system, may be functions in an installed software package, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.

[0042] The disclosure, in certain examples and embodiments, provides methods of measuring gene expression levels that are applicable to single cell data. More generally, the disclosure provides methods of counting molecules present in a sample so as to reduce or eliminate duplicates. The disclosure makes use of intrinsic sequences present near random fragmentation or priming sites to identify individual molecules. The disclosure further provides a binning index that is added during sample preparation and is used during read deduplication and counting as part of a method of correcting for biases that arise during sample preparation. According to methods of the disclosure, an individual molecule is randomly fragmented, and the information encoded in the first N bases of the sequence proximate or adjacent the fragmentation site encodes information about both the gene identity and the unique fragmentation position within that sequence. This provides intrinsic molecular identification (i.e., the identifying information is intrinsic to the original or natural sequence itself). In certain embodiments, it may be useful to also append a short randomer (e.g., NNN) at the cut end of the molecule. This serves as a “molecular diversity enhancer” (MDE) and effectively expands the number of potential molecules that may be resolved for a given gene. The MDE addresses potential concerns regarding undercounting due to limited cut site diversity by increasing the available intrinsic molecular identifier (IMI) space.

[0043] As a tool for read counting and correction, the present description provides for the addition of a diversity of short (e.g., N < 6 bases) random sequences as binning indexes that may be introduced with the cell barcode sequence, e.g., among the capture oligonucleotides decorating a bead such as may be used in scRNA-Seq applications. This “binning index” is distinct from a UMIas conventionally described and implemented. In one embodiment, the binning indexes are 3 or fewer bases in length and comprise fewer than 100 unique identities. After sequencing, sequence reads are grouped by cell barcode, by gene to which the read was mapped, and by IMI (possibly augmented by an MDE).

[0044] Each cell barcode, gene, and IMI combination is associated with a number of reads and with one or more different binning indices, one coming from each read. To resolve exact PCR duplicates (identical sequences), all the reads from this combination are treated as a single molecular count, and that count of reads is associated with a single binning index. If multiple binning indexes are associated with such a combination, one index may be chosen arbitrarily.

[0045] To illustrate the correction, the binning index (BI) may be used as part of a method to resolve multiple fragments derived from the same transcript during whole transcriptome amplification (WTA) and fragmented at different cut sites. In one embodiment, for each binning index, a system will count the number of molecules assigned to that binning index after the deduplication step and divide that number by a correction factor. As discussed herein, in certain implementations, to account for inflation the read counts which have the same gene and barcode (referred to herein as a barcode+gene combination (BGC)) and binning index are divided by the correction factor. The correction factor may be the “worst-case scenario”, based on the number of WTA cycles. For example, with 4 WTA cycles the “worst-case” factor would be 7. The system may sum the corrected counts across binning indices to get the final count for the associated cell barcode and gene. Additional aspects of implementations of suitable corrections are discussed herein in greater detail.

[0046] As discussed in greater detail herein, different implementations and techniques for calculating a suitable correction are described. By way of example, in accordance with a dynamic correction factor selection technique, an integer correction factor is derived that tailors (i.e., is sample-specific) IMI correction to the conversion efficiency in an individual sample. In such an implementation, the correction factor is derived based at least in part by calculating the probability distribution of molecule conversion to IMIs. In an alternative or supplemental approach, a technique employing molecular counting with IMIs that adjusts to the conversion efficiency in an individual sample is also described. Both approaches are believed to be more suitable thanapproaches that employ a single default correction factor (e.g., a default correction factor based on “worst-case” scenarios or maximum correction factor), which may be inconsistent from sample to sample. In particular, the outputs of the deduplication and correction approaches discussed herein may be understood to transform the input data (e.g., single-cell mRNA count data) into valuable and easily accessible insights into accurate expression data at the single-cell level for a given subject or sample.

[0047] With the preceding in mind and by way of further context, high-throughput sequencing technologies yield vast numbers of short sequence reads from a pool of nucleic acid fragments. The presently described techniques provide sequencing applications (e.g., processor-implemented code or routines) that estimate the abundance of a particular fragment by the number of reads obtained in a sequencing experiment (read counting) using intrinsic molecular identifiers. Sequencing and read counting approaches are useful in RNA sequencing (RNA-seq), which may be used to quantify transcript abundance in a sample such as a single cell. Typical workflows involve copying RNA into cDNA, amplifying the cDNA into amplicons that include a molecular identifier copied from the RNA, and sequencing the amplicons to yield sequence reads. The molecular identifier is useful because PCR is non-uniform and neither the abundance of sequence reads nor amplicons is a measure of transcript abundance in a sample. Nucleic acid barcodes known as unique molecular identifiers (UMIs) were previously proposed as a method to count the number of mRNA molecules in a sample, e.g., by labeling PCR duplicates with a common UMI. By incorporating a UMI into each fragment during library preparation, but prior to PCR amplification, the concept pf the UMI-based approach is to identify PCR duplicates because such duplicates would have both identical alignment coordinates and identical UMI sequences. Unfortunately, UMIs require additional steps during sample preparation and consume sequencing real estate. For example, some short-read technologies give sequence reads that can be as short as 35 bases, and some UMIs can be as long as 30 bases. As discussed herein, the presently described techniques provide similar results without requiring the use of UMIs.

[0048] The presently described techniques are useful for creating sequencing libraries that can be sequenced to quantify molecules such as mRNA transcripts in a sample, such as the mRNAs of a single cell. Those molecules can be quantified in accordance with the presently described approaches due to library preparation methods described herein that provide each molecule with abinning index (BI) and an effectively unique, intrinsic identifier, a sequence within the molecule referred to as an intrinsic molecular identifier (IMI), which may optionally also be associated with an appended at the cut end of the molecule, referred to herein as a molecular diversity enhancer (MDE).

[0049] As discussed herein, the molecular identifier is intrinsic in that it comprises bases that are copied from the genetic material being studied, as opposed to extrinsically added bases that are not initially part of the genetic material being studied. For example, where single-cell RNA sequencing (scRNA-Seq) is being performed to quantify messenger RNA (mRNA) transcripts present in a single cell, those mRNA transcripts are copied into cDNAs and a segment of, or sequence of bases from, each cDNA is used as the intrinsic molecular identifier. The sequence of bases is intrinsic in that the sequence originates as part of the genome of the organism (or virus, or other biological source material) and is produced as a cut site in the cDNA. In accordance with aspects of the presently described techniques, the molecule identifier is useful to identify each cDNA because each cDNA has an intrinsic molecular identifier that is effectively unique (e.g., “nearly unique”, or “essentially unique”) in terms of being able to statistically identify a given cDNA from which it is derived. One important feature is that, across all RNA molecules from a cell, those RNA molecules are copied into cDNA molecules that can be mapped to their genes of origin and in which substantially most of the cDNA molecules have a unique intrinsic molecular identifier.

[0050] That level of essentially unique labeling is achieved by cleaving each cDNA molecule at a random site and attaching a PCR handle to the random site or by priming sample nucleic acids at random locations using, e.g., random hexamers, where the primers include a 5' tail with a PCR handle. Typically, the cDNA molecule will have a first PCR handle that has been provided as part of a capture oligonucleotide that annealed, or hybridized, to the mRNA. The capture oligonucleotide, which includes the binning index, is extended by a polymerase, copying the mRNA to form the cDNA. The cDNA is then cleaved at a random cut site, and a second PCR handle is attached at the random cut site. Because the cut site or priming site is random, a segment of the cDNA adjacent to that site will include a sequence of bases that is effectively unique for that cDNA molecule.

[0051] The cDNA molecules can be amplified from the PCR handles and the amplicons can be sequenced. Sequence reads that are generated by sequencing into the cDNA from the random site will include the binning index and a sequence of bases unique to that molecule, i.e., the intrinsic molecular identifier. Sequence reads can be deduplicated and / or mapped to a reference (e.g., a human genome or a gene atlas) to identify genes. After deduplication and mapping, a count of de- duplicated (or unique) reads mapping to each gene is associated with the binning index. The counts may be corrected by a correction factor, and the corrected counts provide a measure of transcripts of that gene from that cell. Thus, the corrected counts provide a measure of expression levels for the cell.

[0052] In certain embodiments, the presently described techniques may be used to create single- cell sequencing libraries and, in particular, libraries useful in single-cell RNA-sequencing (scRNA- Seq). Some scRNA-Seq protocols involve sequencing RNA from a cell and, in certain embodiments, providing a measure of gene expression levels from the sequence data. Some approaches to scRNA-Seq rely on isolating cells into droplets with the potential to assay a large number of cells per experiment. Popular droplet-based protocols include Drop-seq (described in Macosko, 2015, Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets, Cell 161(5):1202-14, incorporated by reference) and inDrop (see Klein, 2015, Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells, Cell 161(5):1187-201, incorporated by reference). As discussed herein, the presently described techniques provide intrinsic molecular identifiers that may be used in such droplet-based protocols in that those protocols do not require UMIs (although the presently described technique may be implemented with or in addition to the use of UMIs).

[0053] In certain embodiments described herein, the libraries may be created with emulsions and template particles that segregate individual cells into droplets upon vortexing. The cells may be lysed inside the droplets, to release RNA. The RNA may be captured by bead-bound capture oligonucleotides that include a bead-specific, and thus a cell-specific, barcode while in the droplets. The capture oligonucleotides may be extended by a reverse transcriptase, copying the RNA to yield cDNA, which is provided with PCR primer binding sites (e.g., “PCR handles”). In certain embodiments, at least one PCR handle is attached at an essentially random location with respect to the cDNA such that a segment of the cDNA adjacent the random location provides anidentifier sequence that is essentially unique to that molecule. Those cDNAs may be amplified and sequenced. Sequence reads from the cells may be mapped to a reference and deduplicated, and deduplicated reads may be counted. The counts are assigned to the associated binning index and optionally corrected, to provide for identification and quantification of RNA from a multitude of single cells in one experiment, in which each cell was isolated in its own aqueous partition. Accordingly, methods of the presently described techniques provide a massively parallel, analytical workflow for preparing single-cell sequencing libraries. The described methods are inexpensive, scalable, and accurate, and do not require UMIs, although such UMIs may be optionally used in addition to the techniques discussed herein.

[0054] Specifically sample preparation workflows are described to illustrate principles of the present techniques, but the binning index as discussed may be used with different sample preparation workflows. For example, the step of sample preparation may include (i) fragmenting the transcripts at the random sites, annealing oligonucleotides to the fragments, and extending the oligonucleotides to make cDNA copies of the transcripts; or (ii) annealing oligonucleotides to the transcripts, extending the oligonucleotides to make cDNA copies of the transcripts, and fragmenting the cDNA copies at the random sites. The sample preparation may include tagmentation using Tn5 transposase to attach primer binding sites or sequencing adaptor at essentially random sites in the molecules. The sample preparation may include cleaving the templates with mechanical force, heat, chemicals such as detergents, or enzymes (e.g., endonucleases) followed by ligation of PCR handles or adaptors.

[0055] Turning to the figures, FIG. 1 shows a block diagram of a method 101 for preparing a sequencing library. The method 101 includes reverse transcribing (step 103) RNA into cDNA. Each cDNAs is cleaved (step 109) at a random location, or “random cut site”, and a synthetic oligonucleotide that includes a binning index is attached (step 115) at the random cut site. Notably, the method 101 includes indexing (step 116) the template molecules, by virtue of having added the binning index. Optionally or alternatively, the random site may be defined by random priming, e.g., using a random hexamer. The cleavage (step 109) and attachment (step 115) may be carried out by any suitable methods. For example, fragmenting may be performed by physical methods, such as acoustic shearing or sonication, or by enzymatic methods, such as with a restriction enzyme, or by exposing the RNA to high temperatures, e.g., about 95 degrees Celsius, in thepresence of multivalent cations, such as, metal ions, for example, Mg2+, Mn2+, or Zn2+. For example, the RNA may be incubated in a solution comprising MgC12, at 95 degrees Celsius, for a few minutes. In certain embodiments, the cleavage (step 109) and attachment (step 115) are performed by a transposase, such as a Tn5 transposase.

[0056] Cleavage (step 109), according to aspects of the presently described techniques, generates cut sites at substantially random positions in the cDNA. Because cleavage of the cDNA is at a random cut site, the cleaved ends of the cDNA molecules are random and essentially unique for the purpose of identification. A downstream (i.e., later) step of the method involves reading sequence from the cleaved ends of the cDNA molecules. Those sequences can be treated as unique if enough bases are read from the cleaved end. That is, if there are hundreds of thousands of cDNA molecules, and only a 3-base intrinsic label is read, then there are only 64 possible unique labels. However, if 10 bases are read (assuming random use of bases) then there are greater than 1 million unique labels. Reading 12 bases gives more than 16 million labels. Fifteen bases give more than 1 billion labels. Reading 17 bases provides more than 17 billion labels; 18 bases give greater than 68 billion labels; 19 bases give greater than 274 billion labels; and 20 bases give greater than 1 trillion labels. For many applications, it is not necessary that each intrinsic molecular identifier (IMI) be unique. For example, in various applications, having two, or three, or even up to ten IMIs that have a duplicate among the set will yield scRNA-seq quantitative results that are useful and significant, i.e., not statistically significantly different than without the duplicates for many end- uses.

[0057] A variation of method 101 does not use cleavage of the cDNA to provide the noted randomness, but instead uses primer binding at a random site such that bases that are intrinsic within the target nucleic acid adjacent the random primer binding site are copied into new DNA and come to serve as a unique intrinsic molecule identifier. These versions may use random hexamers, which are suitable as primers for capturing volumes of RNA such as mRNA from a single cell. In such embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by a random hexamer. For example, an RNA may be captured by a random primer that is extended to create a cDNA. Because random priming (e.g., using a random hexamer) binds at effectively random sites within nucleic acid, each cDNA will include a segment of bases in the cDNA adjacent the priming site that is useful as a unique identifiersequence. Random priming works similarly to random cleavage by a transpose, random mechanical or chemical cleavage, or restriction enzyme cleavage. What is in common among these techniques is that, as far as the sequences of the target nucleic acids are concerned, the binding sites or cut sites are effectively random. Each target nucleic acid will be bound or cut in manner that is unpredictable or inconsistent enough, for the purposes of techniques such as scRNA-seq, that downstream amplicons will have binning indexes and substantially unique intrinsic molecular identifiers (noting again that nearly unique is sufficient for most purposes) at one end that get sequenced and appear in sequence reads.

[0058] The intrinsic molecular identifiers may be used to establish, or contribute to, the unique molecular identity of nucleic acids of a library. For example, RNA transcribed from the same genomic loci may have sequences that are substantially identical. That is, by randomly cleaving the cDNA, each cDNA is made unique by virtue of the bases adjacent the random cut site left by the cleavage (step 109).

[0059] A synthetic oligonucleotide that includes a binning index is attached (step 115) to the cDNA at the random cut site to create a construct that includes at least a portion of the cDNA and the synthetic oligonucleotide. Remembering that the cDNA was created by extending a capture oligonucleotide that annealed to an RNA, any sequence in the capture oligonucleotide will be present in the construct. Any sequence present in the synthetic oligonucleotide will also be present in the construct. The capture oligonucleotide and the synthetic oligonucleotide may either or both have a PCR handle (i.e., a primer binding site, a “universal primer binding site”, a capture tag, a sequencing adaptor, or similar). Thus, in certain embodiments, the attachment (step 115) creates a construct that includes a first PCR handle, an optional functional sequence such as a sample barcode and / or a cell barcode, a hybrid capture portion of the capture oligo (e.g., a poly-T region), a portion of the cDNA, the location of the random cut site, and a second PCR handle. The attachment also indexes (step 116) the template (by labeling with a binning index).

[0060] Because the construct is, in certain embodiments, a contiguous DNA molecule with PCR handles at both ends, it is amenable to amplification (step 123) by, for example, polymerase chain reaction (PCR). The optional functional sequence may be a cell barcode. More specifically, the capture oligonucleotide may be one of a plurality of capture oligonucleotides that are attached toa solid support such as a bead, e.g., a hydrogel bead. All capture oligonucleotides may share one common barcode (which may be referred to as a “bead barcode”). If the bead is isolated in an aqueous partition with a single cell and the capture oligonucleotides are used to anneal to, and capture, RNA molecules from the single cell, then the common barcode of the bead becomes a cell barcode. This is because, downstream, after sequencing (step 127), the presence of the cell barcode sequence in a sequence read is useful to map that sequence read back to the single cell associated with that bead. Thus, a barcode in the construct may be used as a cell barcode. All such constructs from a single cell may be amplified (step 123).

[0061] The constructs may include sequence platform specific primers (e.g., P5 and P7) or those may be added by a round of amplification, e.g., PCR. Amplification (step 123) produces amplicons which may be sequenced (step 127). Due to the sample preparation, the sequencing is effectively initiated from random start sites to generate sequence reads. Each sequence read is indexed (step 116) by a binning index added by an oligonucleotide during sample preparation.

[0062] The method 101 may, in certain embodiments, further include mapping each sequence read to a gene. Counts per gene are obtained of unique intrinsic sequences defined by the random start sites. The counts are assigned to associated binning indexes. A correction factor may be applied to each count, such as to correct for bias introduced in the sample preparation. The corrected counts are summed across the binning indexes for each gene to provide an estimated number of the transcripts per gene in the sample. In such an example, the correction factor may: (1) adjust for a probability of the random start sites being duplicated among the transcripts; (2) adjust for a probability of multiple random start sites per transcript within the sample; (3) include an estimate of a number of copies of each transcript resulting from the amplification, in which case the applying step may include dividing each count by the correction factor; or (4) provide any other suitable adjustment or correction to the read counts. Informatically, after read counts are assigned to their binning indexes, it is possible that for each gene, only the binning indexes, the counts, and the correction factor are used for the applying and summing steps.

[0063] As discussed, methods of the presently described techniques are useful for scRNA-Seq, including but not limited to expression analysis. In certain embodiments, cells are isolated into, and lysed within, aqueous partitions with capture oligonucleotides that include binning indexes.The capture oligonucleotides anneal to RNAs released from the cells. The capture oligonucleotides may include partition-specific barcodes, binning indexes, and PCR handles. Once the capture oligonucleotides have hybridized to the RNAs, those duplexes may be released from partitions and pooled at any subsequent stage. Because capture oligonucleotides with partition-specific barcodes are used to capture and tag RNA from cells isolated in the partition, any arbitrary number of cells may be captured in parallel (i.e., simultaneously). Because the RNAs are tagged with a cell barcode during hybrid capture (e.g., a partition-specific barcode or a bead barcode), if those duplexes are pooled and ultimately sequenced, the cell barcodes in the sequencing data can be used to “bin” the sequence data by original cell, i.e., assign each sequence read (or assembled contigs or sequences therefrom) back to originating cells.

[0064] Multiplexing in the present context may involve isolating cells and the capture oligonucleotides into partitions. Any suitable partitions may be used. The partitions may be any suitable partition in a pico-, nano-, or microtiter plate or substrate, or fluidic harbors (see, e.g., U.S. Publication No.2010 / 0041046 Al, incorporated herein by reference in its entirety for all purposes), chambers (see, e.g., U.S. Publication No. 2021 / 0178395 Al, incorporated herein by reference in its entirety for all purposes), regions defined within a fluidic device (see, e.g., U.S. Publication No. 2020 / 0269248 Al, incorporated herein by reference in its entirety for all purposes), or combinations thereof. In certain embodiments, the partitions are aqueous partitions in an immiscible liquid, e.g., slugs or droplets surrounded or separated by oil within a microfluidic device. A microfluidic device may use channels to mix samples and reagents and form droplets in an immiscible carrier fluid. In certain embodiments, the partitions are a plurality of droplets that are formed essentially simultaneously. Certain implementations may be performed with a sample comprising a mixture with cells, and, in certain embodiments, template particles. The mixture may comprise two immiscible fluids such as an aqueous fluid and oil. The mixture is sheared, e.g., vortexed, to generate an emulsion with template particles that serve to template the formation of droplets and segregate individual cells into the droplets. Because the cells are individually segregated into droplets, the cells may be individually profiled in parallel. Such a method provides a massively parallel, analytical workflow for analyzing single cells that is inexpensive, scalable, and accurate.

[0065] For example, methods of the presently described techniques may include combining template particles with cells in a first fluid and then adding a second fluid that is immiscible with the first fluid to the mixture. The first fluid is, in certain embodiments, an aqueous fluid. While any suitable order may be used, in some instances, a tube may be provided comprising the template particles. The tube can be any type of tube, such as a sample preparation tube sold under the trade name Eppendorf, or a blood collection tube, sold under the trade name Vacutainer. The sample may be a blood sample and may be added directly to the tube using a pipette.

[0066] The fluids can be sheared to generate a monodisperse emulsion with droplets. To generate a monodisperse emulsion, methods may include a step of shearing the mixture provided by combining cells and template particles in an aqueous fluid with the immiscible fluid. Any suitable method or technique may be utilized to apply a sufficient shear force to the mixture. For example, the mixture may be sheared by flowing the second mixture through a pipette tip. Other methods include, but are not limited to, shaking the mixture with a homogenizer (e.g., vortexer), or shaking the mixture with a bead beater. In some embodiments, vortex may be performed for example for 30 seconds, or in the range of 30 seconds to 5 minutes. The application of a sufficient shear force breaks the mixture into monodisperse droplets that encapsulate one of a plurality of template particles.

[0067] After vortexing, a plurality (e.g., thousands, tens of thousands, hundreds of thousands, one million, two million, ten million, or more) of aqueous partitions is formed essentially simultaneously. Vortexing causes the fluids to partition into a plurality of monodisperse droplets. A substantial portion of droplets will contain a single template particle and a single target cell. Droplets containing more than one or none of a template particle or target cell can be removed, destroyed, or otherwise ignored.

[0068] The next step of the method is to lyse the cells. Cell lysis may be induced by a stimulus, such as, for example, lytic reagents, detergents, or enzymes. Reagents to induce cell lysis may be provided by the template particles via internal compartments. In certain embodiments, lysing involves heating the monodisperse droplets to a temperature sufficient to release lytic reagents contained inside the template particles into the monodisperse droplets. This accomplishes cell lysisof the target cells, thereby releasing nucleic acids, such as RNA, such as mRNA, inside of the droplets that contained the target cells.

[0069] After lysing target cells inside the droplets, mRNA is released. The mRNA may be used to create a sequencing library. Methods and systems of the presently described techniques may use template particles to template the formation of monodisperse droplets and isolate single target cells. The disclosed template particles and methods for targeted library preparation thereof may, in one implementation, leverage the particle-templated emulsification technology described in Hatori, 2018, Particle-templated emulsification for microfluidics-free digital biology, Anal Chem 90(16):9813-9820, incorporated herein by reference in its entirety for all purposes. Essentially, micron-scale beads (such as hydrogels) or “template particles” are used to define an isolated fluid volume surrounded by an immiscible partitioning fluid and stabilized by temperature insensitive surfactants.

[0070] In practicing the methods as described herein, the composition and nature of the template particles may vary. For instance, in certain aspects, the template particles may be microgel particles that are micron-scale spheres of gel matrix. In some embodiments, the microgels are composed of a hydrophilic polymer that is soluble in water, including alginate or agarose. In other embodiments, the microgels are composed of a lipophilic microgel.

[0071] With the preceding in mind, FIG.2 illustrates a sample prep tube 229 comprising droplets 201. In particular, the sample prep tube 229 comprises a plurality of monodisperse droplets generated by shearing a mixture 239 according to certain implementations of the present techniques. In certain embodiments, each of the droplets 201 includes, on average, one template particle 213 and zero or one single target cell 209. The template particles 213 may comprise crater- like depressions to facilitate capture of single cells 209. The template particles 213 may further comprise an internal compartment 221 to deliver one or more reagents into the droplets 201 upon stimulus. Each template particle 213 may be decorated with capture oligonucleotides that include a binning index and optionally a cell barcode and a 3' hybrid capture portion.

[0072] In some embodiments, the template particles 213 contain internal compartments 221. The internal compartments 221 of the template particles 213 may be used to encapsulate reagents that can be triggered to release a desired compound, e.g., a substrate for an enzymatic reaction, orinduce a certain result, e.g., lysis of an associated target cell. Reagents encapsulated in the template particles' compartment may be without limitation reagents selected from buffers, salts, lytic enzymes (e.g., proteinase k), other lytic reagents (e. g. Triton X-100, Tween-20, IGEPAL), nucleic acid synthesis reagents, or combinations thereof.

[0073] Lysis of single target cells occurs within the monodisperse droplets and may be induced by a stimulus such as heat, osmotic pressure, lytic reagents (e.g., DTT, beta-mercaptoethanol), detergents (e.g., SDS, Triton X-100, Tween-20), enzymes (e.g., proteinase K), or combinations thereof. In some embodiments, one or more of the reagents (e.g., lytic reagents, detergents, enzymes) is compartmentalized within the template particle 213. In other embodiments, one or more of the reagents is present in the mixture. In some other embodiments, one or more of the reagents is added to the solution comprising the monodisperse droplets, as desired.

[0074] In certain embodiments, template particles 213 comprise a plurality of capture probes. Generally, a capture probe as used herein is an oligonucleotide. In some embodiments, the capture probes are attached to the template particle's material, e.g., hydrogel material, via covalent acrylic linkages. In some embodiments, the capture probes are acrydite- modified on their 5' end (linker region). Generally, acrydite-modified oligonucleotides can be incorporated, stoichiometrically, into hydrogels such as polyacrylamide, using standard free radical polymerization chemistry, where the double bond in the acrydite group reacts with other activated double bond containing compounds such as acrylamide. Specifically, copolymerization of the acrydite-modified capture probes with acrylamide including a crosslinker, e.g., N,N'-methylenebis, will result in a crosslinked gel material comprising covalently attached capture probes. In some other embodiments, the capture probes comprise acrylate terminated hydrocarbon linker and combining the capture probes with a template particle 213 causes their attachment to the template particle 213.

[0075] In some embodiments, after cell suspensions are introduced to template particles 213 in a pre-equilibrated buffer, droplets are generated by vortexing the mixture to capture single cells with individual template particles. The resulting emulsion may be heated on a thermocycler to induce cell lysis. Cell lysis releases the contents of the cell and exposes those contents, including mRNA,to the template particle 213. The presently described techniques provide steps for RNA capture and library preparation using those particles and released mRNA.

[0076] FIG. 3 diagrams RNA capture and library preparation according to certain embodiments of the presently described techniques. As shown, particle 213 is linked to a capture oligonucleotide 305. The capture oligonucleotide 305 anneals to an mRNA 311. The capture oligonucleotide 305 includes a binning index 316. Optionally, methods include capturing transcripts with capture oligonucleotides linked to beads, in which each capture oligonucleotide includes 5'-linkage to bead, cell barcode, binning index 316, annealing primer section-2'. The binning index 116 may comprise, in some implementations, three or fewer bases. In certain embodiments, the binning indexes are 3 bases in length. In other embodiments, the binning indexes are either 1 or 2 or 3 bases or are absent, as if an equimolar mixture of 0, 1, 2, and 3 bases among all of the capture oligonucleotides on a bead.

[0077] In some embodiments, poly-T tails of the capture oligonucleotides 305 anneal to and capture RNA released by lysis. Particle-bound capture oligonucleotides in this application may comprise an acrydite linker, a PE1 priming sequence, a particle barcode, optionally a random sequence, and a poly-T capture moiety. A polymerase (e.g., a reverse transcriptase) extends the capture oligonucleotide 305 to form a cDNA 315. The cDNA 315 and capture oligonucleotide 305 in combination with the mRNA 311 form a duplex 323. This duplex is stably linked to the bead 213. At this stage, it is suitable to break the droplets and pool their contents, wash in buffer, and proceed in library preparation.

[0078] A transposase complex 325 (sometimes called a transposome) is introduced. The transposase complex 325 includes a dimer that includes two of a transposase 327 and two transposon end sequences 329. Here, the transposon end sequences 329 are depicted as both being paired-end 2 end (PE2) sequences, which will cooperate with paired-end 1 (PE1) sequences in the capture oligonucleotide 305 in subsequent amplification and sequence steps. In the depicted method, the transposase randomly cuts the cDNA / mRNA duplex 323 thereby defining a random cut site 333. In a downstream step, read 2 of paired-end sequencing will include the first segment of bases in the cDNA 315 adjacent the random cut site 333.

[0079] Attachment (step 115, FIG.1) of the end sequence 329 to the cDNA 315 at the random cut site 333 produces a construct 337. The construct 337 is a contiguous DNA molecule that includes a first PCR handle (PE1), a cell barcode, a capture segment, a portion of the cDNA 315 terminating at the random cut site 333, and a second PCR handle (PE2).

[0080] Amplification (step 123, FIG. 1) of the construct 337 yields amplicons 341. In some embodiments, constructs are amplified with a P5-PE1 hybrid oligonucleotide and P7 index primer directly into a sequencing library. The library may be sequenced to assess RNA expression, for example, as described in Hrdlickova, 2017, RNA-Seq methods for transcriptome analysis, Wiley Interdisc Rev RNA 8(1): I 0.1002, incorporated herein by reference in its entirety for all purposes.

[0081] Constructs 337 or amplicons 341 may include certain primer and index sequences or copies thereof, such as, P5s and P7s. Those sequences may be any arbitrary sequence useful in downstream analysis. For example, they may be additional universal primer binding sites or sequencing adaptors. For example, either or both of the P5s and P7s may be arbitrary universal priming sequence (universal meaning that the sequence information is not specific to the naturally occurring genomic sequence being studied, but is instead suited to being amplified using a pair of cognate universal primers, by design). The index segment may be any suitable barcode or index such as may be useful in downstream information processing. It is contemplated that the P5 sequences, the P7 sequence, and the index segment may be the sequences use in NGS indexed sequences such as performed on an NGS instrument sold under the trademark ILLUMINA, and as described in Bowman, 2013, Multiplexed Illumina sequencing libraries from picogram quantities of DNA, BMC Genomics 14:466 (esp. in Figure 2), incorporated herein by reference in its entirety for all purposes.

[0082] A transposase 327 is used to randomly cut the cDNA 315. This may be performed using a transposase such as Tn5. See Lin, 2020, RNA sequencing by direct tagmentation of RNA / DNA hybrids, PNASl 17 (6) 2886-2893, incorporated herein by reference in its entirety for all purposes. In brief, the Tn5 transposase randomly binds and cuts double-stranded RNA / DNA and attaches its end sequence to the random cut site.

[0083] Accordingly, some embodiments of the presently described techniques use Tn5 transposase to directly tagment RNA / DNA hybrids and form polynucleotide libraries with intrinsicmolecular identifiers (essentially unique sequences of bases originating in genetic material of the organism or biological system being studied). In particular, Tn5, a RNase H superfamily member, binds to RNA / DNA hybrids similarly as to dsDNA and effectively cuts randomly and then ligates a desired oligonucleotide onto the hybrid. The desired oligonucleotide is, in certain embodiments, a PCR handle (aka a universal primer binding site, a sequencing adaptor, a synthetic oligonucleotide of known sequence to which a PCR primer anneals, etc.). Methods of the presently described techniques may be used with various amounts of input sample, from single cells to large numbers of cells, with a dynamic range spanning numerous orders of magnitude.

[0084] FIG.4 shows a workflow for directional tagmentation that works with template switching oligonucleotides (TSOs). The illustrated technique may be employed in hybrid workflows where one is using TSOs for some other benefit and one also want to use IMIs. This tagmentation approach is useful for 3' end capture and analysis of mRNAs. The steps of the method are shown. In brief, mRNA or total RNA from lysed cells are mixed with an oligonucleotide and incubated (such as at 65° for 3 minutes). The oligonucleotide may include specific primers for amplifying final libraries, such as an adapter-B sequence complementary to an i7 primer. The oligonucleotide may further include a poly-T sequence of, for example, 30 nucleotides that hybridizes with poly- A tails of mRNA. Importantly, the use of this oligonucleotide to prime a first strand cDNA synthesis may result in libraries enriched for the 3' end of mRNA.

[0085] In this figure, notably, the binning index may be added at any of several different steps. By way of example, the binning index may be part of the oligonucleotide in step 1. Alternatively, the binning index may be part of the template switch oligonucleotide in step 2. Similarly, the binning index may be part of the adaptor added by Tn5 in step 4. Further, the binning index may be part of either sequence linker in step 5.

[0086] Reverse transcription can be performed using a reverse transcriptase, such as the reverse transcriptase sold under the trade name SMARTSCRIBE by Takara Bio, optionally in the presence of a template switching oligonucleotide (TSO). The template switching oligonucleotide allows for template switching at the 5' end of the mRNA molecule to incorporate an oligonucleotide such as a universal 3' sequence during first strand cDNA synthesis. Synthesis of the first cDNA strand may be performed using a thermocycler at 42 degrees Celsius for 1 h, followed by 15 minutes at 70degrees Celsius to inactivate the reverse transcriptase. Afterwards, the cDNA may be amplified. The cDNA may be amplified by PCR using commercially available kits such as the kit sold under the trade name OneTaq HS by New England Biolabs. After amplification, the RNA / DNA duplexes may be subjected to tagmentation and adapter ligation.

[0087] During tagmentation and adapter ligation, Tn5 bound adapter (adapter-A) complexes bind with the double RNA / DNA duplexes. The duplexes are cut by the enzymatic activity of the Tn5 complexes and the adapters (“Adaptor A”) are ligated. Tn5 cuts at a random site. Afterwards, the products of the tagmentation reaction may be amplified using the adapters. In certain embodiments, each adaptor includes a binning index. As shown, an i7 primer anneals to Adapter B and an i5 primer anneals to Adapter A. In this depicted embodiment, the read 1 primer will read into a segment of an amplicon adjacent the random cut site. Because the cut site is random, a sequence of bases in that segment is essentially unique. Because the segment is in the amplicon copy of the cDNA, the sequence of the bases is intrinsic to the mRNA, i.e., is a sequence from genetic material of the organism being studied. Because the sequence of bases is essentially unique, a read 1 sequence read will include a unique, intrinsic molecular identifier. More specifically, all sequence reads from the read 1 primer from this library member will include the identical copies of that unique, intrinsic molecular identifier (IMI). Thus, the figure illustrates that IMIs are compatible with workflows that include or use TSOs.

[0088] After sequencing, reads with identical gene-mapping and identical IMIs can be “collapsed”, and a count of only unduplicated reads may be obtained as a quantitative measure of gene transcripts in the sample, i.e., the single cell.

[0089] As shown in FIG.4, the use of IMIs is compatible with RNA capture without necessarily requiring any bead-linked capture oligonucleotides. That is, capture oligonucleotides may be free in solution (as opposed to linked to a solid support). Methods of the presently described techniques are also compatible with the use of capture oligonucleotides that are linked to a solid support such as a bead.

[0090] FIG. 5 shows a method for making libraries that include IMIs using capture oligonucleotides linked to a solid support. The depicted embodiment shows the creation of a sequencing library that includes certain next-generation sequencing (NGS) adaptors. The solidsupport may be a bead and library preparation may be performed using a microfluidic device (e.g., to encapsulate beads, cells, and reagents into droplets). In some embodiments, beads decorated with capture oligonucleotides are used to simultaneously form a monodisperse emulsion that includes a plurality of droplets. Each droplet includes, on average, one bead and one or zero cells. Because the beads are particles that serve as templates that cause the droplets (or aqueous partitions) to form (e.g., when a mixture is vortexed), the droplets may be referred to as particle- templated instant partitions (PIPs), the beads maybe referred to as template particles, and sequencing from such libraries may be referred to as PIP-seq. In the illustrated embodiment, a template particle 1301 is linked to a capture oligonucleotide 1305. As shown, the template particle 1301 is linked to (among other things) mRNA capture oligonucleotides 1305 that include a 3' poly- T region 1309 (although sequence-specific primers or random N-mers may be used). Where the sample includes cell-free RNA, the capture oligonucleotide hybridizes by Watson-Crick base- pairing to a target in the RNA and serves as a primer for reverse transcriptase, which makes a cDNA copy of the RNA. Where the initial sample includes intact cells, the same logic applies but the hybridizing and reverse transcription occurs once a cell releases RNA (e.g., by being lysed).

[0091] In certain embodiments, the target RNAs are mRNAs 1313. Where the target RNAs are mRNAs, the particles 1301 may include mRNA capture oligonucleotides 1305 used to at least synthesize cDNA 1317 as a copy of an mRNA 1313. The particles 1301 may further include cDNA capture oligonucleotides with 3' portions that hybridize to cDNA copies of the mRNA. For the cDNA capture oligonucleotides, the 3' portions may include gene-specific sequences or hexamers. As shown, each of the mRNA capture oligonucleotides 1305 may include, from 5' to 3', a SMART site 1319, a PE1 sequence 1321, a cell or droplet barcode 1323, and a poly-T segment 1309. The capture oligos may also include a binning index 1316 in certain implementations.

[0092] As shown, the capture oligonucleotide 1305 hybridizes to the mRNA 1313. A reverse transcriptase binds and initiates synthesis of a cDNA copy 1317 of the mRNA 1313 to make an RNA / DNA hybrid. Note that the mRNA 1313 is connected to the particle 1301 non-covalently by complementary base-pairing. The cDNA 1317 that is synthesized may be covalently linked to the particle by virtue of the phosphodiester bonds formed by the reverse transcriptase.

[0093] A transposase 1401 binds to the RNA / DNA hybrid. The transposase 1401, which may be a Tn5 transposase in certain embodiments, is attached with adapters 1406 for attaching onto the 5' end of the cDNA 1317. The Tn5 cuts the RNA / DNA hybrid at a random cut site 1351, and the adapters 1406 are ligated onto the random cut site of the cDNA 1317. In certain embodiments the adapter 1406 includes a primer handle 1403 for copying / amplification. At this stage, RNaseH may be introduced to degrade the mRNA 1313. The adaptor 1406 may include a binning index.

[0094] In some embodiments, sequencing adapter 1501 is extended to create a dsDNA 1409. The adapter 1501 includes a first sequence 1503 complementary to the primer handle 1403 and a sequencing primer 1505, such as P7. The adapter 1501 will hybridize to, and prime the copying of, cDNA to create a dsDNA 1409 with the sequencing adapter. Afterwards, the polynucleotide can be separated from the particle and made into a final library product. As discussed herein, the cDNA 1317, and thus also the dsDNA 1409, has a segment adjacent the random cut site 1351 with sequence intrinsic to the mRNA 1313. When that segment is sequenced from primer handle 1403, the resultant sequence reads include the intrinsic sequence.

[0095] Amplification produces a final library product 1601. In this example, the final library product 1601 is formed by the PCR-based extension of a P5-PE1 primer 1505 that is complementary to the PE11509 of the released polynucleotide 1409. Extension of the P5-PE1 primer 1505 by PCR creates the final library product 1601. In some embodiments, the P5-PE1 primer 1505 may include indexes, such as an I5 index, and a P5 index. The final library product may be amplified by PCR in advance of sequencing.

[0096] Any one of the above-described strategies and methods, or combinations thereof, may be used in conjunction with particle-templated emulsions. For example, methods may be used for single cell expression profiling, which may include combining target cells with a plurality of template particles in a first fluid to provide a mixture in a reaction tube. The mixture may be incubated to allow association of the plurality of the template particles with target cells. A portion of the plurality of template particles may become associated with the target cells. The mixture is then combined with a second fluid which is immiscible with the first fluid. The fluid and the mixture are then sheared so that a plurality of monodisperse droplets is generated within the reaction tube. The monodisperse droplets generated comprise (i) at least a portion of the mixture,(ii) a single template particle, and (iii) a single target particle. Of note, in practicing methods of the techniques described herein, a substantial number of the monodisperse droplets generated will comprise a single template particle and a single target particle, however, in some instances, a portion of the monodisperse droplets may comprise none or more than one template particle or target cell.

[0097] In some aspects, generating the template particles-based monodisperse droplets involves shearing two liquid phases. The mixture is the aqueous phase and, in some embodiments, comprises reagents selected from, for example, buffers, salts, lytic enzymes (e.g., proteinase k) and / or other lytic reagents (e.g., Triton X-100, Tween-20, IGEPAL, bm 135, or combinations thereof), nucleic acid synthesis reagents (e.g., nucleic acid amplification reagents or reverse transcription mix), or combinations thereof. The fluid is the continuous phase and may be an immiscible oil such as fluorocarbon oil, a silicone oil, or a hydrocarbon oil, or a combination thereof. In some embodiments, the fluid may comprise reagents such as surfactants (e.g., octylphenol ethoxylate and / or octylphenoxypolyethoxyethanol), reducing agents (e.g., DTT, beta mercaptoethanol), or combinations thereof.

[0098] Some implementations of the presently described techniques use oligonucleotides. Oligonucleotides, sometimes referred to as oligos, are sequences of contiguous nucleotides of DNA, RNA, or a mixture thereof. In certain implementations oligonucleotides comprise DNA. However, in other embodiments, oligonucleotides may comprise RNA. In additional embodiments, oligonucleotides may comprise a mixture of DNA and RNA. Oligonucleotides may comprise noncanonical nucleotides, such as synthetic nucleotides that have been modified to incorporate certain biomolecular properties. The length of the oligonucleotide is usually denoted by “-mer”. For example, an oligonucleotide of six nucleotides is a hexamer, or 6-mer, while one of 25 nucleotides may be referred to as a 25-mer. An oligonucleotide may include other features such as one or more conformationally-restricted nucleic acids or a locked nucleic acid (LNA) base(s) or phosphorothioate inter-base linkages, to improve binding stability or residence times.

[0099] In particular, and with reference to FIG. 1, certain implementations of the presently described techniques are useful to create sequence libraries. Those libraries may be sequenced (step 127, FIG.1) to identify transcript abundances, or gene expression levels, of single cells. Ingeneral, the product of amplification (step 123, FIG. 1) may be understood as producing a sequencing library. However, depending on the sequencing technology being used (e.g., single- molecule long-read sequencing versus short-read ensemble sequencing), the distribution of steps, and the desired storage times or conditions for certain library preparation products, the product of attachment (step 115. FIG. 1) may also or instead be considered a sequencing library. Also, a sequencing library produced by amplification (step 123, FIG.1) may be subject to further rounds of amplification (e.g., at the discretion of a user and / or after being shipped to a different location). For example, RNA capture, cDNA synthesis, and a first round of amplification may be performed at a research or clinical services laboratory to create a sequencing library, which may be stored in a tube such as a microcentrifuge tube. The sequencing library may be shipped (e.g., on dry ice) to a genomics core facility for sequencing. The genomics core facility may provide sequence data via a server or data room. The research or clinical services laboratory or another party may access the sequence data to initiate mapping and / or deduplication, which may occur in an online server, in the cloud, or on a local computer. In general, a sequencing library includes DNA copies of target nucleic acids from a sample of interest with PCR handles or adaptors attached at ends. The amplicons may be stored, for example, at -20 degrees Celsius, or may be analyzed.

[0100] Analyzing amplicons may involve sequencing. The sequencing library may be sequenced (step 127, FIG.1). Sequencing (step 127, FIG.1) may be performed by any suitable method. An example of a sequencing technology that can be used is Illumina sequencing. Illumina sequencing is based on the amplification of DNA on a solid surface using fold-back PCR and anchored primers. Genomic DNA is fragmented and attached to the surface of flow cell channels. Four fluorophore-labeled, reversibly terminating nucleotides are used to perform sequencing. After nucleotide incorporation, a laser is used to excite the fluorophores, and an image is captured, and the identity of the first base is recorded. Sequencing according to this technology is described in U.S. Pub. 2011 / 0009278, U.S. Pub. 2007 / 0114362, U.S. Pub. 2006 / 0024681, U.S. Pub. 2006 / 0292611, U.S. Pat.7,960,120, U.S. Pat.7,835,871, U.S. Pat.7,232,656, U.S. Pat.7,598,035, U.S. Pat.6,306,597, U.S. Pat.6,210,891, U.S. Pat.6,828,100, U.S. Pat.6,833,246, and U.S. Pat. 6,911,345, each incorporated by reference in their entirety for all purposes. In certain embodiments, an Illumina Mi-Seq sequencer may be used, though other suitable nucleic acid sequencers, including next generation sequencers (NGSs), may be employed in generating the dataprocessed in accordance with the techniques described herein. As may be appreciated such sequencing devices and systems may be based on or otherwise employ specialized, non- conventional circuitry, including specialized processing circuitry, specialized memory circuitry or structures, application-specific integrated circuitry, specialized bus or communication structures, and so forth, that are beneficial in addressing the particular issues and problems arising from both the acquisition of and processing of vast amounts of nucleic acid sequence data.

[0101] Sequencing (step 127, FIG. 1) creates sequence reads, i.e., a record of a sequence of bases from at least a part of a nucleic acid strand. The sequence reads may be analyzed to determine expression of RNA associated with genes based on unique reads that correspond to those genes. Analyzing the sequence reads may be performed using software and following multistep procedures. For example, first, the quality of each sequence read, i.e., FASTQ sequence, may be assessed using the software FASTQC. Next, the reads may be trimmed using, for example, Trimmomatic software. See Bolger, 2014, Trimmomatic: a flexible trimmer for Illumina sequence data, Bioinformatics 30(15):2114-2120, incorporated herein by reference in its entirety for all purposes. The trimmed sequence reads may then be mapped to a human genome using, for example, HISAT2 software. HISAT2 outputs files in a SAM (sequence alignment / map format), which may be compressed to binary sequence alignment / map files. Other methods useful for processing and analyzing sequence reads are discussed in U.S. Pat. No. 8,209,130, which is incorporated herein by reference in its entirety for all purposes. Determining gene expression generally involves counting numbers of unique sequence reads that uniquely map to a human reference genome. Mapping reads to a reference to identify genes may be performed using computer software packages known in the art.

[0102] An aspect of the presently described techniques is that mapping reads to a reference and identifying genes gives a quantitative result when reads are deduplicated by IMI to yield one read per mRNA from which those reads originated.

[0103] Because each mRNA is typically copied into cDNA and each cDNA is typically copied into an unpredictably large number of amplicons in the sequencing library, and because each library member is often amplified or read redundantly as part of a sequencing technique, a number of raw sequence reads does not necessarily correlate to numbers of input molecules from the singlecells. Nevertheless, one cell may include abundant transcripts that map to one gene. Here, compositions and methods of the presently described techniques give each cDNA a unique intrinsic identifier that can be identified and used to deduplicate sequence reads. After sequence reads are identified by gene and deduplicated, counts of those reads are associated with their binning indexes. Specifically, reads with identical gene-mapping and identical IMIs can be “collapsed” to produce a count of only unduplicated reads that may be used as a quantitative measure of gene transcripts in the sample, i.e., the single cell.

[0104] As shown in FIG.4, the use of IMIs is compatible with RNA capture without requiring bead-linked capture oligonucleotides. That is, capture oligonucleotides may be free in solution (as opposed to linked to a solid support). However, as also discussed, aspects of the presently described techniques are also compatible with the use of capture oligonucleotides that are linked to a solid support such as a bead.

[0105] As implemented, sequencing reads are clustered in the following order: by cell barcode; by the gene to which the read was mapped; and by IMI (possibly augmented by an MDE). Each cell barcode, gene, and IMI combination is associated with a number of reads, and with one or more different binning indices, one coming from each read. To resolve exact PCR duplicates (identical sequences), all reads from this combination are regarded as a single molecular count, and the count is associated with a single binning index. If multiple binning indexes are associated with such a combination, one index is chosen arbitrarily or by a pre-defined selection rule.

[0106] In the next step, the system may resolve multiple fragments derived from the same transcript during WTA and fragmented at different cut sites. For each binning index, the system may count the number of molecules assigned to that binning index in the previous step, and divide, multiply, shift, or scale that number by a correction factor, as discussed in greater detail below. This correction factor may be the “worst-case scenario”, based on the number of WTA cycles. For example, with 4 WTA cycles the “worst-case” factor would be 7. The system may sum the corrected counts across binning indices to get the final count for the cell barcode and gene.

[0107] In a further implementation, binning tag sequences may be utilized for instrument phasing. Approaches have been implemented with a 0, 1, 2, or 3 base stagger in the capture oligonucleotide, optionally embodied within the binning indexes, as a tool to disrupt alignment ofconserved sequences in sequencing. This may be performed to improve color balance and avoid loss of sequencer registry on certain sequencing instruments. In one implementation of the binning tag strategy, the binning indexes have 0, 1, 2, or 3 N bases. It is understood that having zero bases means that the binning index is not present (i.e., has zero length), which is what is intended here: as a set, the capture oligonucleotides have a mixture of different binning index lengths. This results in a diversity of 85 potential bins without addition of any additional sequenced bases.

[0108] FIG. 6 illustrates a workflow within a system configured to perform embodiments of the presently described techniques. In one such implementation the system brings in sequence reads, e.g., as a FASTQ file from a sequencing instrument. The system deduplicates identical reads (i.e., true PCR duplicates) using IMI (and optionally MDE) as discussed above to obtain counts 605. By way of example, each IMI is “collapsed” into a single count. In this manner, library preparation PCR duplicates are resolved using IMIs (with or without MDEs) and gene ID. In this example barcode 6, gene 2 is collapsed in this manner.

[0109] Read counts are indexed (step 116, FIG.1), i.e., associated with their binning indexes, such as in a read count file 607, which is written to tangible, non-transitory memory. In this example, for each binning index, read counts are summed together to provide binning index read counts 609 in which read counts are grouped by index. The read counts are corrected and summed across indexes to provided corrected read counts 611. In one example, total IMIs for each binning index are counted, the total count for each binning index divided by the “worst case scenario” (e.g., 7 for 4 WTA cycles) and the result rounded up to provided corrected read counts 611, but other corrections, such as may be based on an estimate of WTA-based inflation as discussed herein (or more generally, “amplification-based inflation”), are within the scope of the disclosure. In this manner, and in one implementation, on-bead indexes may be used to resolve molecules that map to the same gene and have different sequences from one another (e.g., due to different WTA cut sites of the same molecule).

[0110] In particular, with respect to correction of read counts and to the determination of a correction factor, certain implementations and explanations are further provided by way of example. By way of context, to evaluate the benefits of such a dynamic correction approach, UMI- containing core particles (UCPs) were used to perform PIP-seq experiments using multiple sampleinput types for direct comparison of UMI-based and IMI-based molecule counts for quantitative estimation of WTA-based inflation (e.g., amplification-based inflation). As used herein, it may be understood that IMI inflation occurs when more than one IMI is created from a single molecule. Across all sample types, when evaluating the observed distribution of IMIs per UMI, there were fewer instances of 15 unique IMIs than would be expected assuming 100% WTA efficiency of a 5-cycle PCR. Such data may be used to infer that WTA-based inflation varies across sample types, as is expected for samples of varying intrinsic RNA content. With this in mind, attempts to correct molecule counts based upon an assumption of 100% WTA efficiency would result in significant deflation of the molecule counts, which will vary across samples and experiments. This analysis suggests that a dynamic correction factor that accounts for the observed variability in the IMI count distribution could be used to adjust the observed molecule counts within a sample to be equivalent to UMI-based counts. Presently contemplated automated routines are described that may be used in determining such a correction factor that may be used to adjust the observed molecule counts within a given experimental dataset (e.g., a sample), thereby adjusting counts dynamically based upon the observed level of WTA-based inflation.

[0111] Turning to FIGS.7A and 7B, these figures depict, in combination, an example analytical workflow is described which may be referred to herein as a dynamic integer correction factor approach or, more simply, as a dynamic correction factor approach. Such an approach may be used for identifying the parent molecule count from discrete or ambiguous IMI distributions and may begin with read mapping to identify genes and to enable all reads to be segregated by gene identity in an output sequence file, such as may be in a binary alignment map (BAM) format. Within the same cell barcode and gene, the BAM file (or corresponding file type) is parsed to identify IMIs without regard for the binning index (BI). As may be appreciated, all IMIs from the same parent molecule share the same BI. In one implementation, the IMI sequences are corrected for sequencing errors by merging IMIs within a Hamming distance of 1. When such similarity is detected, the less frequent IMI is merged into the more frequent IMI in certain implementations. Reads with IMIs that contain an “N” base may be discarded. Next, as illustrated at block 700, IMIs are segregated into 64 bins, based on the 3-base BI sequence in each read. Within each bin, the number of unique IMIs is determined by collapsing duplicate IMIs into a single count, as discussed herein. This step mitigates error arising from obtaining identical IMIs from distinct transcriptsfrom the same gene, owing to a combination of limited-cycle PCR amplification and subsequent fragmentation.

[0112] Turning to block 704, after obtaining counts of unique IMIs in each bin, a subset of the barcode+gene combinations (BGCs) is used to estimate the inflation distribution, i.e., the probabilities of a single parent molecule producing any number of copies from 1 to 15 during WTA (i.e., with 5 WTA cycles, fragmentation can crate between 1 and 15 IMIs from the same parent molecule). This is done by selecting a subset of BGCs in which the number of unique IMIs is lower than or equal to fifteen. Some IMI configurations definitively point to a certain number of molecules. For example, for a given cell barcode, three unique IMIs observed across three bins can only arise from three unique molecules. In other cases, such as three unique IMIs observed in two bins, the number of original molecules is ambiguous, since multiple IMIs in the same bin can represent distinct molecules or copies of the same molecule fragmented at different locations. This approach assumes that IMIs from the same parent molecule will have the same BGC. As discussed below, this or another algorithm may use observed IMI counts to derive the inflation probabilities, as discussed herein. The BGCs where the number of molecules can be determined unequivocally are used to estimate the true number of IMIs that can be attributed to inflation from a single molecule in a single bin. Finally, each of the fifteen counts in single bins is divided by the sum of counts to produce the probability (middle column of the table shown in block 704) of one molecule producing this many IMIs. After establishing the inflation probabilities, a correction factor 708 is selected. In one embodiment, this is an integer, chosen such that no more than 1% of molecules will be over-counted; the cumulative inflation probability is at least 0.99.

[0113] Turning to blocks 712 and 716, within each BGC, IMIs are sorted into bins and collapsed (i.e., deduplicated) into unique counts. Use of unique counts for each IMI in this context eliminates duplication introduced by PCR amplification during library preparation. The number of unique IMI counts in each bin is divided by the correction factor 708 and rounded up to the nearest integer. The estimated count for the barcode and gene (BGC) is the sum of the counts from all of the bins. In one implementation, this final count is written into the raw count matrix. In this context, the final count corresponds to the lower bound on the true molecular count for the barcode and gene combination in question.

[0114] By way of further example and explanation, further description of a process for determining inflation probabilities from counts of IMIs, and the corresponding derivation of a correction factor 708 based on the inflation probabilities is provided. In particular, as discussed herein a correction factor 708 may be selected by estimating the IMI inflation probability from a single molecule using observed IMIs and binning indexes with the same barcode+gene combinations (BGCs). The magnitude of IMI inflation for a given dataset depends on factors such as cell type and sequencing depth. However, the accuracy of the described algorithm is independent of IMI inflation because it is based on the combinatorics of the binning indexes (e.g., 64 binning indexes). Across multiple cell types and sequencing depths, the described algorithm produces accurate estimates of IMI inflation. In practice, inflation probability estimates may be most accurate for low numbers of inflated IMIs, where IMIs are more likely to have distinct binning indexes. For example, when estimating the probability of a molecule having four inflated IMIs, the result is based on the initial known observations of four uninflated IMIs which each have different binning indexes. Four molecules will be distributed in different bins ~91% of the time, so this estimate will be highly accurate because the majority of the data is observed. With this in mind, the described approach offers a robust, computationally efficient, and data-driven approach for correcting amplification bias in scRNA-Seq data, independent of cell type or sequencing depth.

[0115] With this in mind, and turning to FIGS.8 and 9 (wherein FIG.8 depicts a process flow diagram and FIG. 9 depicts a corresponding process flow in the context of sample data representations), further aspects of this approach as they pertain to estimation of inflation probabilities are described. In particular, as discussed herein, the presently contemplated algorithmic estimation of inflation probability estimates inflation from a single parent cDNA molecule using the observed counts (shown here as count data 800) for each distribution of IMIs in bins, from 1 IMI to the maximum possible number (threshold 812) of IMIs that can be obtained from one parent molecule (e.g., 15 for 5 WTA cycles) (step 808). In embodiments where the binning index consists of 3 bases, there are 64 (43) possible bins. For each number of IMIs and each distribution of IMIs (step 816), all potential distributions (shown as remaining counts 820) of parent molecules that can create the observed IMI distribution are assessed. First, the known uninflated distribution 804 is considered, where each IMI is in its own bin and the number of parent molecules (m) is equal to the number of IMIs (i). The total pool of i IMIs derived from m moleculeswhere m = i (shown as estimated total counts 832) is estimated (step 824) by dividing (step 828) the known observations by the expected fraction of m molecules as follows: For each number of molecules m, let Tmbe the number of instances of uninflated m molecules, and Omthe number observed uninflated m molecule in m bins. Tm can be derived by: (1) 64^^^ ^^ ൌ^^లర! ^లరష^^! For example: 64^^^^64^^^64^^ ^^ ൌ ర ൌ ! ൌ^^ల ! లర 64ൌ ^^^ଷFor every other possible distribution of m molecules (in fewer than m bins), the estimated number of occurrences (estimated counts 840) of the distribution is calculated by multiplying (step 836) Tmby the probability of getting that molecule distribution when m molecules are distributed among 64 bins. Then, the expected number of observations of each distribution of i IMIs that can be created by this molecule distribution is taken by proportionally splitting the estimated occurrences of the molecule distribution by the probability of getting each IMI distribution from the molecule distribution (the product of inflation probabilities from each molecule to each IMI for each configuration of i IMIs from this distribution of m molecules). For the first pass where i equals m, the molecule distribution is equivalent to the IMI distribution, so there is only one possibility. For subsequent values of m, there are i-m IMIs which are inflated, and the total pool of molecules Tmwhich has i-m inflated IMIs is estimated as before. The estimated occurrences are subtracted from the remaining observed counts 820 for each IMI distribution in order to obtain the estimated (uninflated) IMI counts for each of the 15 IMI distributions.

[0116] Subtracting the estimated occurrences from m molecules for each configuration of i IMIs allows estimation of the occurrences of the same configuration from m-1 molecules, since the remaining counts for each distribution of i IMIs no longer include observations from i to m molecules. For example, the distribution of i IMIs with three IMIs in one bin and two IMIs in another bin (I3,2) can result from at least two molecules (three IMIs are inflated) and at most five molecules (uninflated). In the first pass, where i=5 and m=5, the estimated occurrences of I3,2from m molecules is calculated by obtaining T5 from the distribution of five IMIs in five bins (O5). T5 is then multiplied by the probability of the getting the molecule distribution M3,2: 64 ∗ 63 ∗ ^^^5,3^ ∗ ^^^2,2^ 64 ∗ 63 ∗ 10 ∗ 1^^൫^^ଷ,ଶ൯ ൌൌ 64ହ64ହൌ 0.0000376when 5 molecules are distributed among 64 bins. Since the molecules are uninflated, this gives the estimated occurrences of I3,2from five molecules. This estimate is then subtracted from the observed counts of I3,2. In the next pass, where i=5 and m=4, T4 with one inflated IMI is estimated by using the remaining counts of the IMI distribution I2,1,1,1 with five IMIs in four bins (O4). The remaining counts for I2,1,1,1now only contain the estimated occurrences from four molecules, since estimated occurrences from five molecules were subtracted in the previous pass. The possible molecule distributions which can produce I3,2 include M3,1 and M2,2. The estimated occurrences of each of these molecule distributions is obtained by multiplying T4by the probability of getting the distribution when four molecules are distributed among 64 bins. Adding an inflated IMI to M2,2will always give the IMI distribution I3,2, while M3,1 will produce I3,2 one in four times (and I4,1 otherwise). The estimated occurrences of M3,1are split proportionally between I3,2and I4,1. The total estimated occurrences of I3,2for m=4 (the sum of the estimates of I3,2from M3,1and M2,2) is subtracted from the remaining counts of I3,2. The remaining counts can then be used to estimate the occurrences of I3,2 when m=3, and so on. Note that when m=3, there are two inflated IMIs, so the inflation probabilities must be included in the estimation. The molecule distribution M2,1can produce I3,2 either by having two molecules each with one inflated IMI (one in each bin) or having two inflated IMIs from one molecule in its own bin. The total occurrences of M2,1 from T3 with two inflated IMIs is first calculated by multiplying T3by the probability of M2,1, then splitting this estimate by the relative inflation probabilities (four ways to select one inflated molecule multiplied by the probability of two uninflated molecules and one inflated with two IMIs, and C(4, 2)=6 waysto select two inflated molecules multiplied by the probability of one uninflated molecule and two molecules with one inflated IMI). The estimates within these two inflation configurations are then further split between the possible resulting IMI distributions in each case (I3,2and I4,1), and subtracted from the remaining counts of the IMI distributions. For values of m and i with more than one inflated IMI, the probabilities of each potential configuration of inflated molecules that could be assigned to m molecules is thereby accounted for.

[0117] This process of estimating the number of observations for each IMI distribution is performed for each number of molecules m from m=i down to m=1. When m=1, the remaining observed count (shown as reference number 880) of i IMIs in 1 bin becomes the count for i IMIs from 1 molecule in the inflation distribution. Because this process starts at i=1 and goes up to the maximum possible i, for any number of IMIs i, the inflation distribution has already been calculated for values from 1 to i-1, giving all the probabilities (shown by reference number 888) necessary to estimate the counts from each molecule distribution for each distribution of i IMIs. If the remaining count when m=1 is 0 at any point (i.e., there are no instances of i IMIs from a single molecule), the algorithm terminates. As discussed herein, the inflation probabilities 888, in turn, may be used to derive (step 892) a correction factor (708) as discussed herein.

[0118] While the preceding describes one approach, referred to herein as a dynamic integer correction factor approach or dynamic correction factor approach, for implementing the present inflation correction techniques, other approaches are also contemplated. In the discussion of the preceding approach, the process uses IMI distributions among bin indexes in barcode+gene combinations (BGCs) with 15 or fewer IMIs to estimate a probability distribution for each molecular conversion outcome (i.e., 1 molecule converted to 1, 2, 3, … 15 IMIs).

[0119] Alternatively, in a different implementation referred to herein as an IPM correction approach or an average IPM correction approach, molecule capture is conceptualized as a series of random samples from the 64 possible binning indexes (BIs). As a result, the observed outcome is the number of distinct BIs. This process may be analogized to the statistical premise characterized as the “coupon collector problem” (CCP), which gives the probability distribution of the number of selections that would need to be made from a set of coupons to see some distinct number of those coupons. In such a context, the expected value (average) of that probabilitydistribution for a certain number of unique (BIs) may be used as the average number of molecules among the set of barcode+gene combinations with a certain number of BIs.

[0120] For a subset of the number of BIs where molecule capture is most similar to the CCP (e.g. 5-32 or 5-30, where the lower threshold is set to avoid noisy data which is concentrated in low BI numbers and where the upper threshold corresponds to the chance of selecting a new BI versus one that has been seen already is equal), the average conversion rate of molecules to IMIs for the sample is approximated by taking the average of the total number of IMIs in each of these barcodes+genes divided by the expected average number of molecules from CCP given the number of unique BIs observed. In certain implementations, the correction is then applied so that BGCs with few molecules (e.g., BIs ≤10 as the threshold), which are expected to have high variance around the average sample IMIs per molecule (IPM or IMI / molecule), get corrected counts based on the number of BIs (where the count is the number of distinct BIs), which is a better predictor of molecule count that IMIs are at low expression. In such an implementation, the remainder of the BGCs may have a corrected count equal to the maximum of the: (1) number of distinct BIs or (2) the result of floor dividing the total number of IMIs by the previously calculated IPM value, as described in greater detail below.

[0121] In further implementations the threshold at which the number of BIs or the number of IMIs are used as the molecular count predictor may be optimized, such as based on one or more of variance, a weighting of the two which adjusts with conversion efficiency and so forth. Further, in certain implementations a hybrid approach may be employed which includes aspects of the combinatorial probabilities in the molecule conversion approach and the approach based on average IPM as discussed below.

[0122] While the above discussion provides a general overview of the IPM correction approach, the following discussion provides a detailed discussion of such an approach with relation back to concepts previously discussed and described. In particular, as discussed herein BI selection occurs at the level of the molecule. As may be appreciated, when 3 molecules are captured, the probability of each outcome (e.g., 1, 2, or 3 unique BIs being selected) is the same in any sample. Further, all IMIs from the same molecule will have the same BI, though conversion efficiency of IMIs to molecules varies from sample to sample. With this in mind, among BGCswith a certain number of molecules, there is a most probable number of unique BIs that may be observed. Further, among BGCs with a certain number of BIs, there is a most probable number of selections (i.e., molecules). Correspondingly, given sufficient observations of BGCs with the same number of unique BIs ) and with sufficiently low variance in the expected molecules) the average rate of conversion (i.e., the IMIs per molecule or IPM) can be estimated by taking the average of the total IMIs divided by the expected number of molecules across the BGCs.

[0123] With respect to estimating the average number of true molecules from the observed number of unique bins, given some number k of unique bins observed in a BGC (where the population of possible bins includes 64 options), the question of what is the average number of molecules that must have been captured to see this many unique bins may be framed. In this context, the expected value for the number of molecules needed to see k distinct bins is: (2) ^ ^^ 1which is related, this model holds true, particularly when k is low.

[0124] With this model in mind, each collection of m captured molecules with the same BGC may be considered a random experiment where m samples are selected with replacement (randomly choosing BIs for molecules) from b total types (i.e., b = 64 BIs). Each molecule is one of a sequence of independent random variables (i.e., x = (X1, X2, …, Xm)) that are uniformlydistributed on D, the population of BIs (i.e., D = {1, 2, …, b}). In this context, ^^^ ∈ ^^ is the BI ofthe ithmolecule (1 ≤ i ≤ m).

[0125] In this context, one may let Vndenote the number of distinct values (e.g., bins) in thefirst n selections. The sample size needed to get k distinct BIs (Wk) where ^^ ∈ ^1, 2, … , ^^^ is:(3) ^^^ ൌ ^^^^^^^^^ ∈ ^^ା: ^^^ ൌ ^^^and the set of possible values for W is {k, k+1. …}. The probability density function of W is:k k(4) ^^^^^^ ൌ ^^^ ൌ ^ ^^ െ 1^ ∑^ି^ ^െ1^^ ൬ ^^ െ 1^ ^^ି^ି^ ^ି^ ^^ െ 1 ^ୀ^ ^^^ ^, ^^ ∈ ^^^, ^^ ^ 1, … ^density function of W. Then:k^ି^ ^ି^ା^^ ^ ^ ^^ ^(5) ^^ ^^ ^ 1 ൌ ^^ ^^ ^^^ ^^^ ^ ^ି^^ ^

[0126] In the present context, W can be decomposed as a sum of k independent, geometricallyk ^^ distributed random variables. For ^^ ∈ ^^ 1, 2, … , ^^ , let Z denote the number of additional samplesi(e.g., molecules) needed to go from i-1 distinct bins to i distinct bins. Then z = (Z, Z, …, Z) is1 2 ba sequence of independent random variables, and Z has the geometric distribution on N withi +^ି^ା^ parameter ^^ ൌ , resulting in:^^ ^∑ ^ ^ (7) ^^ ൌ ^^ , ^^ ∈ ^^ 1, 2, … , ^^.^ ^ ^ୀ^the probability of getting the next new selection (i+1 distinct bins) decreases, so it takes increasingly more additional molecules to get i+1 distinct bins. With this in mind, the mean of the number of molecules (e.g., sample size) needed to get k distinct BIs is: ^ ^^ ^ ∑ (8) ^^ ^^ ൌ^ ^ୀ^^ି^ା^^ସ^∑ which in the present context and example may be equated to.^ୀ^^ସି^ା^ The variance of the number of molecules is: ^ ^^ି^^^^ ^ ∑ (9) ^^^^^^ ^^ ൌ.^

[0127] With the preceding in mind, the presently described correction approach based on average IPM may be further described. Turning to FIG.10, and by way of example, aspects of an IPM calculation strategy are illustrated. In this example, calculation is performed for BGC with between 5 and 32 unique BIs, thus avoiding low and high level of expression. In practice, other suitable low expression level thresholds and / or high expression level thresholds may be employed. In this example, the average number of expected molecules is estimated (column 4 of FIG. 10). This estimate should be constant for each number of BIs. Next, the total number of IMIs per estimated molecules is calculated (column 5 of FIG.10) for each BGC. The mean of the IMIs per molecule (i.e., the mean of the IPM) is then calculated.

[0128] In terms of performing the correction step, for BGCs with ≤10 BIs (or another suitable threshold for cutoff, such as 8, 9, 11, or 12), the corrected count is the number of BIs. For BGCs with >10 BIs, the corrected count is the maximum between the number of BIs or the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule (i.e., the IPM).

[0129] Examples of this correction process are illustrated in FIGS.11 and 12. Turning to FIG 11, in this example the average IPM is 1.99 and the correction rules as outlined for the above- described implementation were employed. Thus, for BGCs having 10 or less BIs (rows 1, 2, and 5 of the depicted table) the corrected count is the number of BIs (column 3). For rows 3, 4, and 6, for which the BIs are greater than 10, a second calculation is performed to determine which is greater, the number of BIs (column 3) or the result (column 6) of floor dividing the total number of IMIs in the BGC by the IMIs per molecule (i.e., the IPM). In the depicted example, the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule exceeds the number of BIs in each of rows 3, 4, and 6, so for these rows the corrected count corresponds to the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule.

[0130] Turning to FIG 12, in this example the average IPM is 2.32 and the correction rules as outlined for the above-described implementation were employed. As in the preceding example, for BGCs having 10 or less BIs (rows 1, 2, and 5 of the depicted table) the corrected count is the number of BIs (column 3). As described above, for rows 3, 4, and 6, for which the BIs are greater than 10, a second calculation is performed to determine which is greater, the number of BIs (column 3) or the result (column 6) of floor dividing the total number of IMIs in the BGC by theIMIs per molecule (i.e., the IPM). In the depicted example, the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule exceeds the number of BIs in each of rows 3 and 4, so for these rows the corrected count corresponds to the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule. However, in this example, for row 6 the number of BIs exceeds the results of the floor division step (i.e., the value in row 6, column 3 (the number of unique BIs) exceeds value in row 6, column 6 (the result of floor dividing the total number of IMIs in the BGC by the IMIs per molecule). As a result, the number of Bis is the corrected count for row 6.

[0131] As may be appreciated, in practice the above-described correction approaches would be implemented using a processor-based system, such as a next generation sequencing (NGS) system or a workstation, computer, or server in direct or indirect communication with such a system or a data repository storing results from such a sequencing system. As may be appreciated, such NGS systems or other downstream analytic devices may constitute or include specialized circuitry and components optimized for the acquisition and / or processing of nucleic acid sequence data, including in large quantities. Correspondingly, the above described approaches may be stored as processor-executable code or routines on a tangible computer-readable medium that may be accessed by a suitable processor or circuitry (e.g., one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and so forth). With respect to such processor-executable code or routines, and example of pseudocode that may be implemented in a suitable language is provided for the purpose of illustration.

[0132] The following pseudocode, for example, illustrates logic for implementing the calculation of an IPM value for a sample using Python-Y pseudocode: # expected average molecules given the number of non-empty bins (num_bins) expected_mols = {} expected = 0 For each num_bins from 1 to 32: expected += 64 / (64-num_bins+1) expected_mols[num_bins] = expected# accumulate sum of total_imis divided by expected molecules from each BGC with between 2 and 10 non-empty bins (num_bins) ipm_total = 0 ipm_count = 0 for each BGC: # num_bins is unique BIs, total_imis is sum of unique IMIs from each BI if 5 <= num_bins <= 32: current_ipm = total_imis / expected_mols[num_bins] ipm_total += current_ipm ipm_count += 1 # final IPM value for correction is the mean of the ipm values for BGCs with 5 to 32 nonempty bins IPM = ipm_total / ipm_count

[0133] With the preceding in mind, it may be appreciated that the approach based on combinatorial probabilities in the molecule conversion (which may be referred to herein as the dynamically selected correction factor (CF) approach or dynamic correction factor approach) and the IPM correction approach (which may also be referred to herein as and average IPM correction approach) are based on a shared basic strategy, namely estimation of a correction that corresponds to the level of IMI inflation across an entire sample followed by application of the correction to some or all of the barcode+gene combination (BC) to estimate molecular counts from observed IMI counts.

[0134] By way of summary and comparison, as described herein the dynamic correction factor approach first estimates molecular conversion efficiency, such as via execution of suitable algorithms or routines as outlined herein. Based on a probability distribution of each possible outcome (e.g., 1, 2, …, 15 IMIs from one parent molecule) a correction factor is selected as the first value where the cumulative probability is greater than or equal to a threshold probability (e.g., ≥ 0.99, ≥ 0.98, ≥ 0.95, and so forth). However, in some instances correcting with a dynamically- selected correction factor may overcorrect BGCs with more non-empty bins (i.e., BGCs with higher expression). Use of such a higher correction factor may have the greatest impact on high- expression BGCs, where high IMI counts in bins are due to true bin collisions among molecules.

[0135] By way of comparison, the average IPM correction approach estimates or calculates IMIs per molecule, such as directly from the data, thus potentially bypassing the complexity that may be associated with application of a molecular conversion efficiency estimation algorithm.This allows determination of an average IMIs per molecule (IPM). The IPM can then be used in conjunction with a rule-based correction approach (e.g., a hierarchical rules-based approach) as outlined herein to broadly or selectively apply a correction.

[0136] In view of the differing focus of the two approaches, it may be appreciated that a point of distinction between the two is that the dynamic correction factor approach provides a conservative correction factor directed to ensuring that at least a threshold percentage of counts (e.g., 99% of counts) are not inflated. The average IPM correction approach is, instead, a global conversion rate from molecules to IMIs, which helps preserve counts in the context of high expression. Further, the use of an average conversion rate, rather than modeling the full conversion probability breakdown, makes the correction faster and more computationally efficient. As may be appreciated, turning IMI counts into estimated molecule counts with the average conversion rate requires enough molecules that the true IMIs / molecule (IPM) in a given individual barcode+gene is expected to be relatively close to the average IPM in the sample. Such an approach may be further enhanced by having an element of sample-specific optimization of the combination approach where the number of BIs where one estimator is favored over the other can be adjusted according to the conversion efficiency.

[0137] Results and Comparison

[0138] Based on the preceding discussion and described techniques various comparisons were performed to evaluate the presently described techniques. In a first comparison, a single-cell protocol was run with no transcriptomic mapping. In this example, IPM is the mean of IPMS from BGCs with 5-30 bins. For BGCs with ≤ 10 unique bins, the corrected count was the number of unique bins. For BGCs with > 10 unique bins, the IPM correction was applied by floor dividing the total IMIs in the BGC by IPM. The corrected count was determined by taking the maximum of this value (i.e., the division result) or the number of unique bins. Results are shown in the table provided in FIG.13.

[0139] As may be observed from these results, compared to UMI metrics, IMI analysis based on the worst-case integer correction factor exhibited an ~8-10% drop in median transcripts. IMI analysis based on the dynamic integer correction factor approach discussed herein exhibited a ~5-7% drop in medial transcripts. IMI analysis based on the IPM correction approach discussed herein exhibited a ± ~2% change in medial transcripts.

[0140] In a further metrics comparison example, a single-cell protocol was run with no transcriptomic mapping. In this example, IPM is the mean of IPMS from BGCs with 5-30 bins. For BGCs with ≤ 10 unique bins, the corrected count was the number of unique bins. For BGCs with > 10 unique bins, the IPM correction was applied by floor dividing the total IMIs in the BGC by IPM. The corrected count was determined by taking the maximum of this value (i.e., the division result) or the number of unique bins. Correction factors were calculated with the dynamic selection approach or worst-case approach (max IMIs from one molecule given the number of WTA cycles). Results are shown in the table provided in FIG.14.

[0141] Overview of System for Biological or Chemical Analysis

[0142] With the preceding discussion in mind, examples and embodiments described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction. Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No.7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety and for all purposes; and / or U.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety and for all purposes.

[0143] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from thereaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.

[0144] Examples of Fluidics Devices and Fluid Flow Paths

[0145] Example of System with Higher Volume Throughput

[0146] With the preceding in mind, FIG.10 illustrates a schematic diagram of an example of a system (1100) that may be used to perform an analysis on one or more samples of interest. A controller (1114) of the present example includes a user interface (1206), a communication interface (1208), one or more processors (1210), and a memory (1212) storing instructions executable by the one or more processors (1210) to perform various functions including the disclosed implementations. User interface (1206), communication interface (1208), and memory (1212) are electrically and / or communicatively coupled to the one or more processors (1210). User interface (1206) may be adapted to receive input from a user and to provide information to the user associated with the operation of system (1100) and / or an analysis taking place. User interface (1206) may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.

[0147] Communication interface (1208) is adapted to enable communication between system (1100) and a remote system(s) (e.g., computers) via a network(s) (e.g., the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system (1100). Some of the communications provided to system (1100) may be associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system (1100).

[0148] The one or more processors (1210) and / or system (1100) may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the oneor more processors (1210) and / or system (1100) includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.

[0149] Memory (1212) may include one or more of a semiconductor memory, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid- state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read- only memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).

[0150] In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shown, system (1100) is configured to receive a flow cell cartridge assembly (1102) including a flow cell assembly (1103) and a sample cartridge (1104). System (1100) includes a flow cell receptacle (1122) that receives flow cell cartridge assembly (1102), a vacuum chuck (1124) that supports flow cell assembly (1103), and a flow cell interface (1126) that is used to establish a fluidic coupling between system (1100) and flow cell assembly (1103). Flow cell interface (1126) may include one or more manifolds. System (1100) further includes a sipper manifold assembly (1106), a sample loading manifold assembly (1108), and a pump manifold assembly (1110). System (1100) also includes a drive assembly (1112), the controller (1114), an imaging system (1116), and a waste reservoir (1118). Controller (1114) is electrically and / or communicatively coupled to drive assembly (1112) and to imaging system (1116); and is configured to cause drive assembly (1112) and / or the imaging system (1116) to perform various functions for performing the techniques disclosed herein.

[0151] In the present example, flow cell assembly (1103) includes a flow cell (1128) having a channel (1130) and defining a plurality of first openings (1132), which are fluidically coupled to the channel (1130) and arranged on a first side (1134) of the channel (1130). Flow cell (1128) further includes a plurality of second openings (1136) fluidically coupled to the channel (1130) and arranged on a second side (1138) of the channel (1130). Fluid may thus flow through flow cell (1128) via the channel (1130). While the flow cell (1128) is shown including one channel (1130), flow cell (1128) may include two or more channels (1130). Flow cell assembly (1103) also includes a flow cell manifold assembly (1140) coupled to flow cell (1128) and having a first manifold fluidic line (1142) and a second manifold fluidic line (1144). Flow cell manifold assembly (1140) may be in the form of a laminate including a plurality of layers as discussed in more detail below.

[0152] In the implementation shown, first manifold fluidic line (1142) has a first fluidic line opening (1146) and is fluidically coupled to each of the first openings (1132) of flow cell (1128); and second manifold fluidic line (1144) has a second fluidic line opening (1148) and is fluidically coupled to each of the second openings (1136). As shown, flow cell assembly (1103) includes gaskets (1150) coupled to flow cell manifold assembly (1140) and fluidically coupled to fluidic line openings (1146, 1148). In some implementations where flow cell (1128) includes a plurality of channels (1130), flow cell manifold assembly (1140) may include additional fluidic lines (1152) that couple first fluidic line openings (1146) to a single manifold port (1154). In such implementations, a single gasket (1150) may be coupled to flow cell manifold assembly (1140) that surrounds the manifold port (1154) and is in fluidic communication with a plurality of channels (1130). In operation, flow cell interface (1126) engages with corresponding gaskets (1150) to establish a fluidic coupling between system (1100) and flow cell (1128). The engagement between flow cell interface (1126) and gaskets (1150) reduces or eliminates fluid leakage between flow cell interface (1126) and flow cell (1128).

[0153] In the implementation shown, first manifold fluidic line (1142) has a portion (1156) that is substantially parallel to a longitudinal axis (1158) of channel (1130); and second manifold fluidic line (1144) has a portion (1160) that is substantially parallel to longitudinal axis (1158) of channel (1130). Additionally, first manifold fluidic line (1142) is shown being at least partially adjacent a first end (1162) of flow cell (1128) and spaced from a second end (1164) of flow cell(1128); and second manifold fluidic line (1144) is shown being at least partially adjacent second end (1164) of flow cell (1128) and spaced from first end (1162). Other arrangements of manifold fluidic lines (1142, 1144) may prove suitable, however.

[0154] In the implementation shown, system (1100) includes a sample cartridge receptacle (1166) that receives sample cartridge (1104) that carries one or more samples of interest (e.g., an analyte). System (1100) also includes a sample cartridge interface (1168) that establishes a fluidic connection with sample cartridge (1104). Sample loading manifold assembly (1108) includes one or more sample valves (1170). Pump manifold assembly (1110) includes one or more pumps (1172), one or more pump valves (1174), and a cache (1176). Valves (1170, 1174) and pumps (1172) may take any suitable form. Cache (1176) may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system (1100). While cache (1176) is shown being included in pump manifold assembly (1110), cache (1176) may alternatively be located elsewhere (e.g., in sipper manifold assembly (1106) or in another manifold downstream of a bypass fluidic line (1178), etc.).

[0155] Sample loading manifold assembly (1108) and pump manifold assembly (1110) flow one or more samples of interest from sample cartridge (1104) through a fluidic line (1180) toward flow cell cartridge assembly (1102). In some implementations, sample loading manifold assembly (1108) may individually load or address each channel (1130) of flow cell (1128) with a respective sample of interest. The process of loading channel (1130) with a sample of interest may occur automatically using system (1100). As shown in FIG. 10, sample cartridge (1104) and sample loading manifold assembly (1108) are positioned downstream of flow cell cartridge assembly (1102). In the implementation shown, sample loading manifold assembly (1108) is coupled between flow cell cartridge assembly (1102) and pump manifold assembly (1110). To draw a sample of interest from sample cartridge (1104) and toward pump manifold assembly (1110), sample valves (1170), pump valves (1174), and / or pumps (1172) may be selectively actuated to urge the sample of interest toward pump manifold assembly (1110). Sample cartridge (1104) may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves (1170). To individually flow the sample of interest toward channel (1130) of flow cell (1128) and away from pump manifold assembly (1110), sample valves (1170), pump valves (1174), and / or pumps (1172) may be selectively actuated to urge the sample ofinterest toward flow cell cartridge assembly (1102) and into respective channels (1130) of flow cell (1128).

[0156] Drive assembly (1112) interfaces with sipper manifold assembly (1106) and pump manifold assembly (1110) to flow one or more reagents that interact with the sample within flow cell (1128). In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system (1116) excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system (1100).

[0157] After image data is obtained, drive assembly (1112) interfaces with sipper manifold assembly (1106) and pump manifold assembly (1110) to flow another reaction component (e.g., a reagent) through flow cell (1128) that is thereafter received by waste reservoir (1118) via a primary waste fluidic line (1182) and / or otherwise exhausted by system (1100). Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.

[0158] The primary waste fluidic line (1182) is coupled between pump manifold assembly (1110) and waste reservoir (1118). In some implementations, pumps (1172) and / or pump valves (1174) of pump manifold assembly (1110) selectively flow the reaction components from flow cell cartridge assembly (1102), through fluidic line (1180) and sample loading manifold assembly (1108) to primary waste fluidic line (1182). Flow cell cartridge assembly (1102) is coupled to a central valve (1184) via flow cell interface (1126). Central valve (1184) is coupled with flow cell interface (1126) via a fluidic line (1185). An auxiliary waste fluidic line (1186) is coupled to central valve (1184) and to waste reservoir (1118). In some implementations, auxiliary waste fluidic line (1186) receives excess fluid of a sample of interest from flow cell cartridge assembly(1102), via central valve (1184), and flows the excess fluid of the sample of interest to waste reservoir (1118) when back loading the sample of interest into flow cell (1128).

[0159] Sipper manifold assembly (1106) includes a shared line valve (1188) and a bypass valve (1190). Shared line valve (1188) may be referred to as a reagent selector valve. Central valve (1184) and the valves (1188, 1190) of sipper manifold assembly (1106) may be selectively actuated to control the flow of fluid through fluidic lines (1192, 1194, 1196). Sipper manifold assembly (1106) may be coupled to a corresponding number of reagent reservoirs (1198) via reagent sippers (1200). Reagent reservoirs (1198) may contain fluid (e.g., reagent and / or another reaction component). In some implementations, sipper manifold assembly (1106) includes a plurality of ports. Each port of sipper manifold assembly (1106) may receive one of the reagent sippers (1200). Reagent sippers (1200) may be referred to as fluidic lines. Some forms of reagent sippers (1200) may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly (1106). Reagent reservoirs (1198) may be provided in a cartridge, and the tubes of reagent sippers (1200) may be configured to be inserted into corresponding reagent reservoirs (1198) in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir (1198) into the sipper manifold assembly (1106).

[0160] Shared line valve (1188) of sipper manifold assembly (1106) is coupled to central valve (1184) via shared reagent fluidic line (1192). Different reagents may flow through shared reagent fluidic line (1192) at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly (1110) may draw wash buffer through shared reagent fluidic line (1192), central valve (1184), and flow cell cartridge assembly (1102).

[0161] Bypass valve (1190) of sipper manifold assembly (1106) is coupled to central valve (1184) via dedicated reagent fluidic lines (1194, 196). Each of the dedicated reagent fluidic lines (1194, 1196) may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines (1194, 1196) may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.

[0162] Bypass valve (1190) is also coupled to cache (1176) of pump manifold assembly (1110) via bypass fluidic line (1178). One or more reagent priming operations, hydration operations,mixing operations, and / or transfer operations may be performed using bypass fluidic line (1178). The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly (1102). Thus, the operations using bypass fluidic line (1178) may occur during, for example, incubation of one or more samples of interest within flow cell cartridge assembly (1102). That is, shared line valve (1188) may be utilized independently of bypass valve (1190) such that bypass valve (1190) may utilize bypass fluidic line (1178) and / or cache (1176) to perform one or more operations while shared line valve (1188) and / or central valve (1184) simultaneously, substantially simultaneously, or offset synchronously perform other operations.

[0163] Drive assembly (1112) includes a pump drive assembly (1202) and a valve drive assembly (1204). Pump drive assembly (1202) may be adapted to interface with one or more pumps (1172) to pump fluid through flow cell (1128) and / or to load one or more samples of interest into flow cell (1128). Valve drive assembly (1204) may be adapted to interface with one or more of the valves (1170, 1174, 1184, 1188, 1190) to control the position of the corresponding valves (1170, 1174, 1184, 1188, 1190).

[0164] Examples of Combinations

[0165] The following examples relate to various non-exhaustive ways in which the teachings herein may be combined or applied. The following examples are not intended to restrict the coverage of any claims that may be presented at any time in this application or in subsequent filings of this application. No disclaimer is intended. The following examples are being provided for nothing more than merely illustrative purposes. It is contemplated that the various teachings herein may be arranged and applied in numerous other ways. It is also contemplated that some variations may omit certain features referred to in the below examples. Therefore, none of the aspects or features referred to below should be deemed critical unless otherwise explicitly indicated as such at a later date by the inventors or by a successor in interest to the inventors. If any claims are presented in this application or in subsequent filings related to this application that include additional features beyond those referred to below, those additional features shall not be presumed to have been added for any reason relating to patentability.

[0166] Example 1

[0167] An apparatus comprising [COPY AND PASTE CLAIMS HERE, WITH EACH CLAIM BEING LISTED AS AN “EXAMPLE” (INCLUDING IN THE DEPENDENCIES). ALL “EXAMPLES” COPIED FROM DEPENDENT CLAIMS SHOULD BE IN MULTIPLE- DEPENDENT FORMAT].

[0168]

[0169]

[0170] Aspects of the preceding examples may be understood in view of the preceding discussion as well as the following context material. In particular, aspects of the presently described techniques provide a system for nucleic acid analysis that includes a solid support and a nucleic acid construct attached to the solid support, such as a bead. The nucleic acid construct may include a linker for attachment to the solid support, a cell-identification barcode, a binning index, a capture region, and a region of cDNA comprising a portion in which a unique identifier sequence that is intrinsic to the cDNA has been generated. In certain embodiments each bead is linked to a plurality of the nucleic acid constructs and the binning indexes are about 2 to 6 bases in length, such as 3 bases. In some embodiments, the region of cDNA has been randomly cleaved at a cut site, and a synthetic oligonucleotide has been attached at the cut site. The system may include a transposase that functions to cleave the cDNA or a primer that primes at an essentially random location, thereby generating the unique identifier sequence, and a paired-end sequence (or similar synthetic oligonucleotide) for hybridization to a sequencing surface. In certain embodiments the transpose cleaves the cDNA at a cut site that is random or cannot be predicted and attaches the paired-end sequence to the cDNA at the cut site. The unique identifier sequence is defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The system may include a plurality of paired-end sequence-ligated cDNAs, e.g., all linked to the solid support (through oligonucleotides that each include a binning index) and each comprising an identifier sequence in the cDNA adjacent a random cut site. In certain embodiments, sequence reads from the plurality of cDNAs can be deduplicated to quantify RNAs captured on the solid support.

[0171] In some embodiments, the cDNA has been randomly cleaved by a restriction enzyme or sonication, and the synthetic oligonucleotide has been attached by a ligase. In embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by arandom hexamer. For example, an RNA may have been captured by a random primer that was extended to create the cDNA such that the primer-binding site is random and a segment of bases in the cDNA adjacent the priming site is useful as a unique identifier sequence. In certain embodiments, the cDNA has been randomly cleaved by, and the synthetic oligonucleotide has been attached by, a ligase.

[0172] In certain implementations the unique identifier sequence is defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The plurality of bases may be intrinsic, e.g., copied from genetic material of an organism. The system may include a plurality of the solid supports (e.g., beads), each solid support attached to cDNA copies of RNAs from a single cell, in which each cDNA copy has a unique identifier defined by bases in a segment of that cDNA copy adjacent a random cut site, such that the RNAs from a single cell can be quantified by sequencing the RNAs and deduplicating sequence reads by the unique identifier. In certain embodiments, the deduplicated reads are counted and each count is associated with its binning index. For each binning index, the count is corrected (e.g., divided or multiplied by a correction factor to correct for bias introduced during sample preparation) and then, for each binning index, the counts are added together to provide a quantitative measure of transcripts in a sample from which the cDNAs were prepared. The solid supports may comprise hydrogel beads linked to a plurality of capture oligonucleotides that each include a binning index, e.g., of about three bases. The system may include a plurality of the hydrogel beads, each isolated in an aqueous partition.

[0173] Other aspects of the presently described techniques provide a method for generating nucleic acid library. The method includes providing a sample comprising a plurality of cells, each comprising sample nucleic acids; hybridizing sample nucleic acids to a construct comprising a solid support to which is attached, via a linker, a cellular barcode sequence, a binning index of fewer than about eight bases (e.g. fewer than five bases), and a capture sequence; extending said construct from said capture sequence to form a duplex comprising an extended construct; exposing said duplex to a transposase, thereby to generate a unique identifier sequence at a 3' end of said construct; and amplifying said extended construct; thereby to create a nucleic acid library. The extending step may reverse transcribe a sample nucleic acid into a cDNA in the construct. In certain embodiments the transposase cuts the cDNA at a random cut site. The unique identifier sequence may be provided by a segment of the cDNA adjacent or near the random cut site. Themethod may include sequencing the library to generate sequence reads, mapping the sequence reads to genes in a reference, collapsing (i.e., de-duplicating) reads that include the same unique identifier sequence leaving only unique reads, counting unique reads, and associating each count with its binning index.

[0174] In some embodiments, the method includes isolating the cells into partitions and creating sequencing libraries from single cells in the partitions. The solid support may be a bead attached to a plurality of copies of the cellular barcode sequence. The method may include attaching synthetic oligonucleotides (such as paired-end sequences, PCR handles, or sequencing adaptors) at random cut sites in mRNA molecules.

[0175] In some aspects, the presently described techniques provide a method for generating a nucleic acid library. The method includes capturing RNA molecules from a single cell with capture oligonucleotides that include a first PCR handle; extending the capture oligos to form duplexes comprising the RNA molecules and cDNA; and cleaving the duplexes at, and attaching second PCR handles to, random cut sites to thereby form constructs that each include a label defined by intrinsic sequence of a cDNA segment adjacent the random cut site wherein at least the first or second PCR handle is provided by an oligonucleotide that includes a binning index. The method may further include amplifying the constructs to form amplicons; sequencing the amplicons to produce sequence reads; counting sequence reads with duplicate intrinsic sequences as one RNA molecule from the single cell, and storing the resultant count under its binning index. The capture oligonucleotides may be linked to a solid support in an aqueous partition that includes the single cell.

[0176] The solid support may be a bead and the aqueous partition may be a droplet. The method may include forming a plurality of droplets that each include, on average, one bead decorated with capture oligonucleotides and zero or one single cell. In some embodiments, the droplets are formed in channels of a microfluidic device. In certain embodiments the plurality of droplets are formed substantially simultaneously by shearing or vortexing a vessel comprising an aqueous phase, an immiscible phase, oligo-linked beads, and cells. The capture oligonucleotides may further include cell barcodes.

[0177] In certain embodiments, the constructs include at least the first PCR handles, the cell barcodes, cDNAs, and the second PCR handles, in which one of the PCR handles includes the binning index. The amplicons may include copies of the binning index and the first and second PCR handles such that the copies of the first and second PCR handles anneal to sequencing adaptors. The method may include mapping the sequence reads to a reference to identify genes from which one or more of the RNA molecules were transcribed. The method may include correcting the indexed counts by a correction factor (optionally summing corrected counts per binning index) to give estimated transcription levels and providing a report with transcription levels of the genes in the single cells based on the counted sequence reads and identified genes.

[0178] In some embodiments, the cleaving and / or the attaching steps are performed by an enzyme such as a transposase that creates the random cut sites. The intrinsic sequence of the cDNA may be copied from genetic material of the single cell such that, due to the random cut site, each duplex includes a label useful to uniquely identify the cDNA is sequencing data.

[0179] Other aspects of the presently described techniques provide a method that includes cleaving a cDNA at, and attaching an oligonucleotide that includes a binning index of 3 bases or less to, a random cut site; copying the cDNA to generate copies that include the binning index and an intrinsic label copied from a segment of the cDNA adjacent the cut site; sequencing the copies to generate sequence reads; and collapsing duplicate sequence reads that contain the same intrinsic label. Deduplicated read counts are stored by binning index and optionally corrected. The intrinsic label may be some number of bases, e.g., between about 5 and about 30, from a segment near or adjacent the random cut site. The method may include capturing an mRNA with a capture oligo and extending the capture oligo to synthesize the cDNA. The cDNA may be linked to a bead. The method may include isolating single cells into droplets, lysing the cells to release RNA into the droplets, capturing an mRNA within one of the droplets, and making the cDNA from the mRNA. The droplets may be formed simultaneously in a technique that also includes isolating the cells into the droplets (e.g., shearing a mixture that includes beads and cells in an aqueous phase plus an oil). The cleaving and the attaching at the random cut site may be performed using an enzyme such as a transposase. The oligo that is attached at the random cut site may include a PCR handle used in the copying step. The attaching step may yield a DNA construct including a first primingsite, a cell barcode, the binning index, a portion of the cDNA, the random cut site, and a second priming site. The copying step may comprise amplification by polymerase chain reaction (PCR).

[0180] Miscellaneous

[0181] While the foregoing examples are provided in the context of a system (1100) that may be used in nucleotide sequencing processes, the teachings herein may also be readily applied in other contexts, including in systems that perform other processes (i.e., other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes.

[0182] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in the description herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0183] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.

[0184] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are notrestricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.

[0185] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departing from its scope. While the dimensions, types of materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain- English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means—plus-function format and are not intended to be interpreted based on 35 U.S.C. §112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.

[0186] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.

Claims

What is claimed is:

1. A method of measuring gene expression for a sample, the method comprising: for the sample, sequencing DNA molecules derivative of respective mRNA molecules, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise a molecular identifier sequence; attaching a binning index to the sequence reads; mapping each sequence to a genomic region; determining counts based at least in part on the molecular identifier sequence for each sequence read; assigning the counts to associated binning indexes; determining a measure of amplification-based inflation using observed counts for each distribution of molecular identifier sequences associated with each binning index; determining a correction factor based on the measure of amplification-based inflation, wherein the correction factor is specific to the sample; applying the correction factor to correct the counts within the sample so as to reduce bias; and determining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

2. The method of claim 1, wherein the DNA molecules comprise cDNA molecules.

3. The method of claim 1, wherein the random break points comprise random fragmentation sites introduced by use of a chemical or enzymatic agent.

4. The method of claim 1, wherein the random break points correspond to cellular or genetic damage associated with a spatial location within a subject.

5. The method of claim 1, wherein the bias is introduced during sample preparation.

6. The method of claim 1, wherein the amplification-based inflation is estimated from a single cDNA molecule.

7. The method of claim 1, wherein the measure of amplification-based inflation is on a scale 0 to 15.

8. The method of claim 1, wherein deriving the measure of amplification-based inflation comprises iterating from 1 unique portion to the maximum possible number of unique portions that can be obtained for one parent molecule based on a number of amplification cycles performed.

9. The method of claim 1, wherein the binning index comprises 2-6 bases.

10. The method of claim 1, wherein the counts are corrected dynamically based upon the measure of amplification-based inflation.

11. The method of claim 10, wherein the counts are corrected dynamically based on one or more of sample specific factors, cell type, level of gene expression, or sequencing depth.

12. The method of claim 1, wherein the correction factor is derived based upon the measure of amplification-based inflation using observed molecular identifier sequences and binning indexes having the same barcode + gene combinations (BGCs).

13. The method of claim 1, wherein deriving the measure of amplification-based inflation comprises: generating distribution data from observed count data; determining a number of remaining counts for each unique portion from 1 to a threshold determined based upon a number of amplification cycles performed; and based on the remaining counts, estimating a distribution {i} corresponding to a single molecule.

14. The method of claim 13, further comprising: for a number of parent molecules M greater than 1, determining an estimated total counts; determining an estimated counts based on the estimated total counts; and subtracting the estimated counts from the remaining counts.

15. The method of claim 14, wherein the estimated total counts are determined by dividing an observed counts by a fraction corresponding to an observed counts distribution fraction of total distributions with M molecules, wherein the observed counts correspond to the remaining counts where a distribution length corresponding to the number of binning indices equals M.

16. The method of claim 13, further comprising: for a number of parent molecules M equal to 1, determining non-zero counts from remaining counts in each iterative loop; and dividing non-zero counts by a sum of counts to generate the measure of amplification- based inflation.

17. The method of claim 1, further comprising: fragmenting mRNA transcripts at the random break points; annealing oligonucleotides to the mRNA fragments; and extending the oligonucleotides to generate cDNA.

18. The method of claim 1, further comprising: annealing oligonucleotide to the mRNA molecules; extending the oligonucleotides to generate cDNA of the mRNA molecules; and fragmenting the cDNA at the random break points.

19. The method of claim 1, wherein the correcting step accounts for a probability of the random break points being duplicated among the mRNA molecules.

20. The method of claim 1, wherein the correcting step accounts for a probability of multiple random break points per mRNA molecule.

21. One or more computer readable-media encoding processor-executable code, wherein the encoded code comprises: code for sequencing DNA molecules derivative of respective mRNA molecules present in a sample, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise a molecular identifier sequence; code for mapping each sequence to a genomic region; code for determining counts based at least in part on the molecular identifier sequence for each sequence read; code for assigning the counts to associated binning indexes based on respective binning indexes attached to the sequence reads; code for determining a measure of amplification-based inflation using observed counts for each distribution of molecular identifier sequences associated with each binning index; code for determining a correction factor based on the measure of amplification-based inflation, wherein the correction factor is specific to the sample; code for applying the correction factor to correct the counts within the sample so as to reduce bias; and code for determining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

22. A nucleic acid sequencing system comprising: one or more processing components configured to execute stored routines for executing operations of the nucleic acid sequencing system; and one or more memory components encoding routines for execution by the one or more processing components, wherein the routines, when executed by the one or more processing components, cause the nucleic acid sequencing system to performing acts comprising:for a sample, sequencing DNA molecules derivative of respective mRNA molecules, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise a molecular identifier sequence; attaching a binning index to the sequence reads; mapping each sequence to a genomic region; determining counts based at least in part on the molecular identifier sequence for each sequence read; assigning the counts to associated binning indexes; determining a measure of amplification-based inflation using observed counts for each distribution of molecular identifier sequences associated with each binning index; determining a correction factor based on the measure of amplification-based inflation, wherein the correction factor is specific to the sample; applying the correction factor to correct the counts within the sample so as to reduce bias; and determining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

23. A method of measuring gene expression for a sample, the method comprising: for the sample, sequencing DNA molecules derivative of respective mRNA molecules, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise an intrinsic molecular identifier (IMI); attaching a binning index to the sequence reads; mapping each sequence to a genomic region; determining counts based at least in part on the molecular identifier sequence for each sequence read; assigning the counts to associated binning indexes; determining an average IMI per molecule (IPM) for each parent molecule; correcting the counts for each parent molecule within the sample based on the respective average IPM so as to reduce bias; anddetermining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

24. The method of claim 23, wherein determining the average IPM for each parent molecule is based on a subset of barcode+gene combinations (BGCs) for the respective parent molecule.

25. The method of claim 24, wherein the subset is determined at least in part by a lower threshold selected to reduce or eliminate noisy data.

26. The method of claim 24, wherein the subset is determined at least in part by an upper threshold selected to correspond to the chance of selecting a new binning index versus one that has been observed.

27. The method of claim 24, wherein correcting the counts for each parent molecule within the sample based on the respective average IPM is based on one or more rules.

28. The method of claim 27, wherein the one or more rules comprise a first rule by which BGCs having binning indexes at or below a threshold are assigned corrected counts based on their binning indexes.

29. The method of claim 28, wherein the one or more rules comprise a second rule by which BGCs having binning indexes above the threshold are assigned corrected counts equal to the maximum of: (1) the number of distinct BIs or (2) the result of floor dividing the total number of IMIs by the IPM.

30. The method of claim 23, wherein the DNA molecules comprise cDNA molecules.

31. The method of claim 23, wherein the random break points comprise random fragmentation sites introduced by use of a chemical or enzymatic agent.

32. The method of claim 23, wherein the random break points correspond to cellular or genetic damage associated with a spatial location within a subject.

33. The method of claim 23, wherein the bias is introduced during sample preparation.

34. The method of claim 23, further comprising: fragmenting mRNA transcripts at the random break points; annealing oligonucleotides to the mRNA fragments; and extending the oligonucleotides to generate cDNA.

35. The method of claim 23, further comprising: annealing oligonucleotide to the mRNA molecules; extending the oligonucleotides to generate cDNA of the mRNA molecules; and fragmenting the cDNA at the random break points.

36. One or more computer readable-media encoding processor-executable code, wherein the encoded code comprises: code for sequencing DNA molecules derivative of respective mRNA molecules present in a sample, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise a molecular identifier sequence; code for mapping each sequence to a genomic region; code for determining counts based at least in part on the molecular identifier sequence for each sequence read; code for assigning the counts to associated binning indexes based on respective binning indexes attached to the sequence reads; code for determining a measure of amplification-based inflation using observed counts for each distribution of molecular identifier sequences associated with each binning index; code for determining a correction factor based on the measure of amplification-based inflation, wherein the correction factor is specific to the sample; code for applying the correction factor to correct the counts within the sample so as to reduce bias; andcode for determining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

37. A nucleic acid sequencing system comprising: one or more processing components configured to execute stored routines for executing operations of the nucleic acid sequencing system; and one or more memory components encoding routines for execution by the one or more processing components, wherein the routines, when executed by the one or more processing components, cause the nucleic acid sequencing system to performing acts comprising: for a sample, sequencing DNA molecules derivative of respective mRNA molecules, wherein sequencing starts at or proximate to random break points such that sequence reads are generated that each comprise a molecular identifier sequence; attaching a binning index to the sequence reads; mapping each sequence to a genomic region; determining counts based at least in part on the molecular identifier sequence for each sequence read; assigning the counts to associated binning indexes; determining an average IMI per molecule (IPM) for each parent molecule; correcting the counts for each parent molecule within the sample based on the respective average IPM so as to reduce bias; and determining corrected counts across the binning indexes for each genomic region to provide a corrected estimate of the number of mRNA transcripts for each genomic region in the sample.

Citation Information

Patent Citations

  • Methods for producing a paired tag from a nucleic acid sequence and methods of use thereof

    US20060024681A1

  • Paired end sequencing

    US20060292611A1

  • Confocal imaging methods and apparatus

    US20070114362A1

  • Method and apparatus for the discretization and manipulation of sample volumes

    US20100041046A1

  • Nucleic acid sequencing system and method

    US20110009278A1