Thresholding techniques for supplementary nucleotide tag assignment in single-cell analysis
Patent Information
- Application Number
- PCT/US2026/015531
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-21
- Filing Date
- 2026-02-17
- Publication Date
- 2026-08-27
Smart Images

Figure US2026015531_27082026_PF_FP_ABST
Abstract
Description
THRESHOLDING TECHNIQUES FOR SUPPLEMENTARY NUCLEOTIDE TAG ASSIGNMENT IN SINGLE-CELL ANALYSISCross-Reference to Related ApplicationsThe present application claims priority to and the benefit of U.S. Provisional Application No.63 / 761,651, titled “THRESHOLDING TECHNIQUES FOR SUPPLEMENTARY NUCLEOTIDE TAG ASSIGNMENT IN SINGLE-CELL ANALYSIS,” filed February 21, 2025, the disclosure of which is hereby incorporated by reference for all purposes.Technical Field
[0001] The presently described techniques relate to the quantitative detection and analysis of molecules in a sample, such as a single cell.Background
[0002] Living organisms store genetic information in DNA. Genes in the coding regions of DNA are transcribed into messenger RNA (mRNA), which is translated into protein. Proteins play critical functional and structural roles in living organisms. For example, most enzymes are proteins, and those enzymes catalyze the metabolic reactions essential to life. It is also enzymes that copy DNA into mRNA. Proteins are also structural, and constitute the essential fibers of muscles, the predominant material of hair, as well as basic structural linkages within the cytoskeleton. Essentially, all such proteins are made by translating an mRNA into the protein. In fact, one mRNA can serve as the template for synthesizing multiple copies of a protein.
[0003] Because cells need to change in response to different conditions, and need different proteins at different times, it is helpful if any given mRNA is short-lived. Most mRNA molecules have a lifetime measured in seconds or minutes. Nevertheless, the health of a cell, or its response to a pathogen, or a drug, or to age-specific developmental changes may be indicated by the quantities of mRNA molecules present in a cell. As a consequence, there is a need for methods to sequence, quantitate, and evaluate cellular RNA. That is usually accomplished by creating so-called sequencing libraries, arrays of complementary DNA (cDNA) reversetranscribed from cellular RNA. The cDNA is sequenced via one of several procedures and the whole process can take days or even weeks to complete. There is, therefore, a continuing need for improved methods for identifying and evaluating cellular RNA in a rapid, sensitive, and specific fashion.Summary
[0004] The presently described techniques provide systems and methods for making sequencing libraries that are useful for analyzing and determining quantities of nucleic acids in a sample including for measuring mRNA transcript levels in a single cell. Methods and systems of the present techniques provide each nucleic acid in a sample with a unique label during preparation of a sequencing library. When the nucleic acids are sequenced, sequence data includes the unique label, allowing sequence reads to be deduplicated by those labels, leaving one unduplicated sequence for each molecule in the original sample. The unduplicated sequences represent a quantitative measure of nucleic acids in the starting sample. Where, for example, the sample includes messenger RNA transcripts from a cell, the sequence reads can be mapped to a reference to identify genes being expressed in a cell. But the sequence reads can also be deduplicated by the unique label to give a measure of transcripts from each gene, i.e., giving gene expression levels for the cell.
[0005] According to the presently described techniques, the unique label is provided as sequence that is intrinsic to the nucleic acid being analyzed. Compared to prior methods that required the attachment of large numbers of synthetic unique molecular identifiers (UMIs) to all of the nucleic acid molecules, the present techniques uses random cleavage or priming of sample nucleic acids, at a random cutting or priming site, leaving those nucleic acids with essentially unique segments adjacent that random cut site. Adaptors or PCR handles are attached at the random sites and library preparation steps, such as amplification or sequencing proceed by annealing primers to those PCR handles. Amplification provides sequencing libraries in which a sequencing primer anneals in the adaptor or PCR handle and generates a sequence read from adjacent the random site. With paired-end sequencing, only one sequence read need begin at the random site. The sequence reads can be mapped to a reference, and they will include a unique identifier sequence that comes from within the nucleic acid molecule being analyzed, i.e., anintrinsic molecular identifier (IMI). The IMI is unique for each molecule and can thus be used to deduplicate sequence reads originating from the same molecule.
[0006] In certain aspects, the presently described techniques provide systems for nucleic acid analysis. Systems of the present techniques include a solid support having attached thereto a plurality of DNA molecules, each member of the plurality comprising a capture sequence, a segment of DNA containing an identifier sequence that is intrinsic to the segment of DNA and unique with respect to other DNA molecules attached to the same or a different solid support, and a primer binding sequence.
[0007] The intrinsic identifier sequence may be from about 5 to about 20 base pairs in length. The DNA may come from a coding region. Systems may include a linker for attachment (e.g., covalent) to the solid support. In systems of the presently described techniques, the primer binding sequence may be ligated to the DNA using transposase such as Tn5. In certain embodiments, the capture sequence hybridizes with a portion of an RNA, e.g., such as a poly-T region when the DNA is a cDNA. The transposase cleaves the DNA (e.g., at a random site) to generate the identifier sequence and may insert a paired-end sequence or similar primer binding site or sequencing adaptor (e.g., at the random site.
[0008] Aspects of the presently described techniques provide a system for nucleic acid analysis that includes a solid support and a nucleic acid construct attached to the solid support, such as a bead. The nucleic acid construct includes a linker for attachment to the solid support, a cell-identification barcode, a capture region, and a region of cDNA comprising a portion in which a unique identifier sequence that is intrinsic to the cDNA has been generated. In certain embodiments the region of cDNA has been randomly cleaved at a cut site, and a synthetic oligomer (also referred to as an “oligo” herein) has been attached at the cut site. The system may include a transposase that functions to cleave the cDNA or a primer that primes at an essentially random location, thereby generating the unique identifier sequence, and a paired-end sequence (or similar synthetic oligo) for hybridization to a sequencing surface. In certain embodiments the transpose cleaves the cDNA at a cut site that is random or cannot be predicted and attaches the paired-end sequence to the cDNA at the cut site. The unique identifier sequence is defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The system may include a plurality of paired-end sequence-ligated cDNAs, e.g., all linked to the solid support and each comprising an identifier sequence in the cDNA adjacent a random cut site. In certainembodiments, sequence reads from the plurality of cDNAs can be deduplicated to quantify RNAs captured on the solid support.
[0009] In some embodiments, the cDNA has been randomly cleaved by a restriction enzyme or sonication, and the synthetic oligo has been attached by a ligase. In embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by a random hexamer. For example, an RNA may have been captured by a random primer that was extended to create the cDNA such that the primer-binding site is random and a segment of bases in the cDNA adjacent the priming site is useful as a unique identifier sequence. In certain embodiments, the cDNA has been randomly cleaved by, and the synthetic oligo has been attached by, a ligase.
[0010] The unique identifier sequence may be defined by a plurality of bases in a segment of the cDNA adjacent the cut site. The plurality of bases may be intrinsic, e.g., copied from genetic material of an organism. The system may include a plurality of the solid supports (e.g., beads), each solid support attached to cDNA copies of RNAs from a single cell, in which each cDNA copy has a unique identifier defined by bases in a segment of that cDNA copy adjacent a random cut site, such that the RNAs from a single cell can be quantified by sequencing the RNAs and deduplicating sequence reads by the unique identifier. In certain embodiments, the solid support comprises a hydrogel bead linked to a plurality of capture oligos, e.g., the system may include a plurality of the hydrogel beads, each isolated in an aqueous partition.
[0011] Other aspects of the present techniques provide a method for generating nucleic acid libraries. The method includes providing a sample comprising a plurality of cells (such as cells determined to be positive for the presence of one or more supplementary nucleotide tags (SNTs), as discussed herein), each comprising sample nucleic acids; hybridizing sample nucleic acids to a construct comprising a solid support to which is attached, via a linker, a cellular barcode sequence, and a capture sequence; extending said construct from said capture sequence to form a duplex comprising an extended construct; exposing said duplex to a transposase, thereby to generate a unique identifier sequence at a 3’ end of said construct; and amplifying said extended construct; thereby to create a nucleic acid library. The extending step may reverse transcribe a sample nucleic acid into a cDNA in the construct. In certain embodiments the transposase cuts the cDNA at a random cut site. The unique identifier sequence may be provided by a segment of the cDNA adjacent the random cut site. The method may include sequencing the library togenerate sequence reads, mapping the sequence reads to genes in a reference, and collapsing reads that include the same unique identifier sequence.
[0012] In some embodiments, the method includes isolating the cells into partitions and creating sequencing libraries from single cells in the partitions. The solid support may be a bead attached to a plurality of copies of the cellular barcode sequence. The method may include attaching synthetic oligos (such as paired-end sequences, PCR handles, or sequencing adaptors) at random cut sites in mRNA molecules.
[0013] In some aspects, the presently described techniques provide a method for generating a nucleic acid library. The method includes capturing RNA molecules from a single cell with capture oligos that include a first PCR handle; extending the capture oligos to form duplexes comprising the RNA molecules and cDNA; and cleaving the duplexes at, and attaching second PCR handles to, random cut sites to thereby form constructs that each include a label defined by intrinsic sequence of a cDNA segment adjacent the random cut site. The method may further include amplifying the constructs to form amplicons; sequencing the amplicons to produce sequence reads; and counting sequence reads with duplicate intrinsic sequences as one RNA molecule from the single cell. The capture oligos may be linked to a solid support in an aqueous partition that includes the single cell.
[0014] The solid support may be a bead and the aqueous partition may be a droplet. The method may include forming a plurality of droplets that each include, on average, one bead decorated with capture oligos and zero or one single cell. In some embodiments, the droplets are formed in channels of a microfluidic device. In certain embodiments the plurality of droplets are formed substantially simultaneously by shearing or vortexing a vessel comprising an aqueous phase, an immiscible phase, oligo-linked beads, and cells. The capture oligos may further include cell barcodes.
[0015] In certain embodiments, the constructs include at least the first PCR handles, the cell barcodes, cDNAs, and the second PCR handles. The amplicons may include copies of the first and second PCR handles such that the copies of the first and second PCR handles anneal to sequencing adaptors. The method may include mapping the sequence reads to a reference to identify genes from which one or more of the RNA molecules were transcribed. The method may include providing a report with transcription levels of the genes in the single cells based on the counted sequence reads and identified genes.
[0016] In some embodiments, the cleaving and / or the attaching steps are performed by an enzyme such as a transposase that creates the random cut sites. In certain embodiments the intrinsic sequence of the cDNA is copied from genetic material of the single cell such that, due to the random cut site, each duplex includes a label useful to uniquely identify the cDNA is sequencing data.
[0017] Other aspects of the presently described techniques provide a method that includes cleaving a cDNA at, and attaching an oligo to, a random cut site; copying the cDNA to generate copies that include an intrinsic label copied from a segment of the cDNA adjacent the cut site; sequencing the copies to generate sequence reads; and collapsing duplicate sequence reads that contain the same intrinsic label. The intrinsic label may be some number of bases, e.g., between about 5 and about 30, from a segment near or adjacent the random cut site. The method may include capturing an mRNA with a capture oligo and extending the capture oligo to synthesize the cDNA. In certain embodiments the cDNA is linked to a bead. The method may include isolating single cells into droplets, lysing the cells to release RNA into the droplets, capturing an mRNA within one of the droplets, and making the cDNA from the mRNA. The droplets may be formed simultaneously in a technique that also includes isolating the cells into the droplets (e g., shearing a mixture that includes beads and cells in an aqueous phase plus an oil). In certain embodiments the cleaving and the attaching at the random cut site are performed using an enzyme such as a transposase. The oligo that is attached at the random cut site may include a PCR handle used in the copying step. The attaching step may yield a DNA construct including a first priming site, a cell barcode, a portion of the cDNA, the random cut site, and a second priming site. In certain embodiments the copying step comprises amplification by polymerase chain reaction (PCR).
[0018] In certain implementations calling of cell-associated features within single-cell data may be practiced as part of the sample preparation and / or analysis. By way of example, such cell-associated features may be characterized herein as supplementary nucleotide tags (SNTs), which may include, but are not limited to: CRISPR guide RNA (gRNA), gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces. By way of example, in certain implementations the presence of gRNA (or other SNTs) in single-cell data may be assessed by determining a threshold for each class of gRNA byapplying k-means clustering to a respective guide purity fraction for each class of gRNA, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of gRNA. In such an example, the guide purity fraction for each class of gRNA may be determined for each respective single-cell by dividing read counts for each respective gRNA by total gRNA counts for the respective single-cell. The cells determined to be positive for a respective class or classes of gRNA may then be further processed, such as in the library preparation operations described herein.Brief Description of the Drawings
[0019] FIG. 1 diagrams a method of preparing a sequencing library.
[0020] FIG. 2 illustrates cell encapsulation in droplets for scRNA-seq.
[0021] FIG. 3 diagrams RNA capture and library preparation by methods of the present disclosure.
[0022] FIG. 4 shows IMI libraries made in workflows using template switching oligos (TSOs).
[0023] FIG. 5 shows IMI libraries with beads and next-generation sequencing (NGS) adaptors.
[0024] FIG. 6 shows an exemplary polynucleotide construct for obtaining sequence reads from paired-end sequencing.
[0025] FIG. 7 shows an exemplary polynucleotide construct for obtaining sequence reads from single-directional sequencing.
[0026] FIG. 8 shows a polynucleotide construct for use in enriching a target nucleic acid through magnetic bead pulldown.
[0027] FIGS. 9A and 9B graphically depict comparative results of a fractional scaling approach (right) as disclosed herein versus a conventional log-transformed counts approach (left).
[0028] FIGS. 10A, 10B, 10C, and 10D depict plots of results for presently disclosed techniques as applied to different classes of gRNAs.
[0029] FIGS. 11 A and 1 IB depict plots of results for presently disclosed techniques as applied to different classes of gRNAs and for different read depths.
[0030] FIG. 12 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.Detailed description
[0031] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components and / or software modules. Thus, for example, one or more of the functional blocks (e.g., processors or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or randomaccess memory, hard disk, or the like). Similarly, the programs may be stand-alone programs encoded on one or more computer-readable media, may be incorporated as subroutines in an operating system, may be functions in an installed software package stored on computer-readable media, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.
[0032] High-throughput sequencing technologies yield vast numbers of short sequence reads from a pool of nucleic acid fragments. The presently described techniques provide sequencing applications that estimate the abundance of a particular fragment by the number of reads obtained in a sequencing experiment (read counting) using intrinsic molecular identifiers.Sequencing and read counting approaches are useful in RNA sequencing (RNA-seq), which may be used to quantify transcript abundance in a sample such as a single cell. Typical workflows involve copying RNA into cDNA, amplifying the cDNA into amplicons that include a molecular identifier copied from the RNA, and sequencing the amplicons to yield sequence reads. The molecular identifier is useful because PCR is non-uniform and neither the abundance of sequence reads nor amplicons is a measure of transcript abundance in a sample. Nucleic acid barcodes known as unique molecular identifiers (UMIs) were previously proposed as a method to count the number of mRNA molecules in a sample, e.g., by labeling PCR duplicates with a common UMI. By incorporating a UMI into each fragment during library preparation, but prior to PCR amplification, the idea was to identify PCR duplicates because they would have both identical alignment coordinates and identical UMI sequences. Unfortunately, UMIs require additional steps during sample preparation and consume “sequencing real estate”. Some short-read technologies give sequence reads that can be as short as 35 bases, and some UMIs can be as long as 30 bases. Methods of the presently described techniques provide similar results without requiring the use of UMIs.
[0033] Further, the presently described techniques facilitate distinguishing cells that are positive for one or more supplementary nucleotide tags (SNTs) from non-positive cells. Such SNTs may include, but are not limited to: CRISPR guide RNA (gRNA), gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces. By way of example, in one such implementation the described techniques may facilitate the separation of gRNA targeted cells from cells without gRNA perturbations. In such implementations the presence of gRNA (or, more broadly, any SNTs) in single-cell data may be assessed by determining a threshold for each class of gRNA by applying k-means clustering to a respective guide purity fraction for each class of gRNA, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of gRNA.
[0034] Methods of the presently described techniques are useful for creating sequencing libraries that can be sequenced to quantify molecules such as mRNA transcripts in a sample, such as the mRNAs of a single cell. Those molecules can be quantified by methods of the present techniques due to library preparation methods given herein that provide each molecule with a unique, intrinsic identifier, a sequence within the molecule that could be referred to as an intrinsic molecular identifier (IMI).
[0035] The molecular identifier is intrinsic in that it is made of bases that are copied from the genetic material being studied. For example, where single-cell RNA-seq (scRNA-Seq) is being performed to quantify messenger RNA (mRNA) transcripts present in a single cell, those mRNA transcripts are copied into cDNAs and a segment of, or sequence of bases from, each cDNA is used as the intrinsic molecular identifier. The sequence of bases is intrinsic in that the sequence originates as part of the genome of the organism (or virus, or other biological source material) and is produced as a cut site in the cDNA. For certain implementations, the molecule identifier is useful to identify each cDNA because each cDNA has an intrinsic molecular identifier that is “unique”, “nearly unique”, or “essentially unique”. One important feature is that, across all RNA molecules from a cell, those RNA molecules are copied into cDNA molecules that can bemapped to their genes of origin and in which substantially most of the cDNA molecules have a unique intrinsic molecular identifier.
[0036] That level of unique labeling is achieved by cleaving each cDNA molecule at a random site and attaching a PCR handle to the random site or by priming sample nucleic acids at random locations using, e g., random hexamers, where the primers include a 5’ tail with a PCR handle. Typically, the cDNA molecule will have a first PCR handle that has been provided as part of a capture oligo that annealed, or hybridized, to the mRNA. The capture oligo is extended by a polymerase, copying the mRNA to form the cDNA. The cDNA is then cleaved at a random cut site, and a second PCR handle is attached at the random cut site. Because the cut site or priming site is random, a segment of the cDNA adjacent to that site will include a sequence of bases that is effectively unique for that cDNA molecule.
[0037] The cDNA molecules can be amplified from the PCR handles and the amplicons can be sequenced. Sequence reads that are generated by sequencing into the cDNA from the random site will include a sequence of bases unique to that molecule, i.e., the intrinsic molecular identifier. Sequence reads can be deduplicated and / or mapped to a reference (e.g., a human genome or a gene atlas) to identify genes. After deduplication and mapping, a count of deduplicated (or unique) reads mapping to each gene is a measure of transcripts of that gene from that cell. Thus the de-duplicated and happed reads provide a measure of expression levels for the cell.
[0038] In certain embodiments, methods of the presently described techniques are used to create single-cell sequencing libraries and, in particular, libraries useful in single-cell RNA-sequencing (scRNA-Seq). Some scRNA-Seq protocols involve sequencing RNA from a cell and, in most embodiments, providing a measure of gene expression levels from the sequence data. Some approaches to scRNA-Seq rely on isolating cells into droplets with the potential to assay a large number of cells per experiment. Popular droplet-based protocols include Drop-seq (described in Macosko, 2015, Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets, Cell 161(5): 1202-14, incorporated by reference) and inDrop (see Klein, 2015, Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells, Cell 161(5): 1187-201, incorporated by reference). The presently described techniques provide intrinsic molecular identifiers that may be used in such droplet-based protocols such that thoseprotocols do not require UMIs (although, to be clear, methods of present techniques are perfectly compatible with the use of UMIs).
[0039] In certain embodiments described herein, the libraries may be created with emulsions and template particles that segregate individual cells into droplets upon vortexing. The cells may be lysed inside the droplets, to release RNA. The RNA may be captured by bead-bound capture oligos that include a bead-specific, and thus a cell-specific barcode while in the droplets. The capture oligos may be extended by a reverse transcriptase, copying the RNA to yield cDNA, which is provided with PCR primer binding sites (“PCR handles”). In certain embodiments at least one PCR handle is attached at an essentially random location in the cDNA such that a segment of the cDNA adjacent the random location provides an identifier sequence that is unique to that molecule. Those cDNAs may be amplified and sequenced. Sequence reads from the cells may be mapped to a reference and deduplicated, providing for identification and quantification of RNA from a multitude of single cells in one experiment, in which each cell was isolated in its own aqueous partition. Accordingly, methods of the present techniques provide a massively parallel, analytical workflow for preparing single-cell sequencing libraries. The methods are inexpensive, scalable, and accurate, and do not require UMIs.
[0040] FIG. 1 shows a block diagram of a method 101 for preparing a sequencing library. The method 101 includes reverse transcribing 103 RNA into cDNA. Each cDNAs is cleaved 109 at a random location, or “random cut site”, and a synthetic oligo is attached 115 at the random cut site. Optionally or alternatively, the random site may be defined by random priming, e.g., using a random hexamer. The cleavage and attachment may be carried out by any suitable methods known in the art. For example, fragmenting 109 may be performed by physical methods, such as acoustic shearing or sonication, or by enzymatic methods, such as with a restriction enzyme, or by exposing the RNA to high temperatures, e.g., about 95 degrees Celsius, in the presence of multivalent cations, such as, metal ions, for example, Mg2+, Mn2+, or Zn2+. For example, the RNA may be incubated in a solution comprising MgC12, at 95 degrees Celsius, for a few minutes. In certain embodiments, the cleavage 109 and attachment 115 are performed by a transposase such as a Tn5 transposase.
[0041] Cleavage 109, according to methods of the present techniques, generates cut sites at substantially random positions in the cDNA. Because cleavage of the cDNA is at a random cut site, the cleaved ends of the cDNA molecules are random and essentially unique. A downstream(i.e., later) step of the method involves reading sequence from the cleaved ends of the cDNA molecules. Those sequences can be treated as unique if enough bases are read from the cleaved end. That is, if there are hundreds of thousands of cDNA molecules, and only a 3-base intrinsic label is read, then there are only 64 possible unique labels. However, if 10 bases are read (assuming random use of bases) then there are greater than 1 million unique labels. Reading 12 bases gives more than 16 million labels. Fifteen bases gives more than 1 billion labels. Reading 17 bases provides more than 17 billion labels; 18 bases give > 68B labels; 19 b give > 274 B labels; and 20 b give > 1 trillion labels. For many applications, it is not necessary that each intrinsic molecular identifier (IMI) be unique. For example, in various applications, having two, or three, or even ten IMIs that have a duplicate among the set will yield scRNA-seq quantitative results that are useful and significant, i.e., not statistically significantly different than without the duplicates for many end-uses.
[0042] A slight variant of method 101 does not use cleavage of the cDNA but instead uses primer binding at a random site such that bases that are intrinsic within the target nucleic acid adjacent the random primer binding site are copied into new DNA and come to serve as a unique intrinsic molecule identifier. These versions may use random hexamers, which are suitable as primers for capturing volumes of RNA such as mRNA from a single cell. In such embodiments, the unique identifier sequence intrinsic to the cDNA has been defined by random priming, e.g., by a random hexamer. For example, an RNA may be captured by a random primer that is extended to create a cDNA. Because random priming (e g., using a random hexamer) binds at effectively random sites within nucleic acid, each cDNA will include a segment of bases in the cDNA adjacent the priming site that is useful as a unique identifier sequence. Random priming works similarly to random cleavage by a transpose, random mechanical or chemical cleavage, or restriction enzyme cleavage. What is in common among those techniques is that, as far as the sequences of the target nucleic acids are concerned, the binding sites or cut sites are effectively random. Each target nucleic acid will be bound or cut in manner that is unpredictable or inconsistent enough, for the purposes of techniques such as scRNA-seq, that downstream amplicons will have unique intrinsic molecular identifiers (noting again that nearly unique is sufficient for most purposes) at one end that get sequenced and appear in sequence reads.
[0043] Those intrinsic molecular identifiers of the presently described techniques may be used to establish, or contribute to, the unique molecular identity of nucleic acids of a library. Forexample, RNA transcribed from the same genomic loci may have sequences that are substantially identical. Here, by randomly cleaving the cDNA, each cDNA is made unique by virtue of the bases adjacent the random cut site left by the cleavage 109.
[0044] A synthetic oligo is attached 115 to the cDNA at the random cut site to create a construct that includes at least a portion of the cDNA and the synthetic oligo. Remembering that the cDNA was created by extending a capture oligo that annealed to an RNA, any sequence in the capture oligo will be present in the construct. Any sequence present in the synthetic oligo will also be present in the construct. The capture oligo and the synthetic oligo may either or both have a PCR handle (i.e., a primer binding site, a “universal primer binding site”, a capture tag, a sequencing adaptor, or similar). Thus, in certain embodiments, the attachment 115 creates a construct that includes a first PCR handle, an optional functional sequence such as a sample barcode and / or a cell barcode, a hybrid capture portion of the capture oligo (e.g., a poly-T region), a portion of the cDNA, the location of the random cut site, and a second PCR handle.
[0045] Because the construct may be a contiguous DNA molecule with PCR handles at both ends, it is amenable to amplification 123 by, for example, polymerase chain reaction (PCR). The optional functional sequence may be a cell barcode. More specifically, the capture oligo may be one of a plurality of capture oligos that are attached to a solid support such as a bead, e.g., a hydrogel bead. All capture oligos may share one common barcode (reasonably referred to as a “bead barcode”). If the bead is isolated in an aqueous partition with a single cell and the capture oligos are used to anneal to, and capture, RNA molecules from the single cell, then the common barcode of the bead becomes a cell barcode. This is because, downstream, after sequencing 127, the presence of the cell barcode sequence in a sequence read is useful to map that sequence read back to the single cell associated with that bead. Thus, a barcode in the construct may be used as a cell barcode. All such constructs from a single cell may be amplified 123.
[0046] The constructs may include sequence platform specific primers (e.g., P5 and P7) or those may be added by a round of amplification, e.g., PCR. Amplification 123 produces amplicons which may be sequenced 127.
[0047] In particular, methods of the present techniques are useful to create sequence libraries. Those libraries may be sequenced 127 to identify transcript abundances, or gene expression levels, of single cells. In general, it is understood that the product of amplification step 123 produces a sequencing library. However, depending on the sequencing technologybeing used (e.g., single-molecule long-read sequencing versus short-read ensemble sequencing), the distribution of steps, and the desired storage times or conditions for certain library prep products, the product of attachment 115 could be considered a sequencing library. Also, a library produced by amplification 123 may freely be subject to further rounds of amplification (e.g., entirely at the preference of a user), e.g., after being shipped to a different location. For example, RNA capture, cDNA synthesis, and a first round of amplification may be performed at a research or clinical services laboratory to create a sequencing library, which may be stored in a tube such as a microcentrifuge tube. The sequencing library may be shipped (e.g., on dry ice) to a genomics core facility for sequencing. The genomics core facility may provide sequence data via a server or data room. The research or clinical services laboratory or another party may access the sequence data to initiate mapping and / or deduplication, which may occur on an online server, in the cloud, or on a local computer. In general, a sequencing library includes DNA copies of target nucleic acids from a sample of interest with PCR handles or adaptors attached at ends. The amplicons may be stored, for example, at -20 degrees Celsius, or may be analyzed. Analyzing amplicons may involve sequencing. The sequencing library may be sequenced 127.
[0048] Sequencing 127 may be performed by any method known in the art. An example of a sequencing technology that can be used is Illumina sequencing. Illumina sequencing is based on the amplification of DNA on a solid surface using fold-back PCR and anchored primers.Genomic DNA is fragmented and attached to the surface of flow cell channels. Four fluorophore-labeled, reversibly terminating nucleotides are used to perform sequencing. After nucleotide incorporation, a laser is used to excite the fluorophores, and an image is captured, and the identity of the first base is recorded. Sequencing according to this technology is described in U.S. Pub. 2011 / 0009278, U.S. Pub. 2007 / 0114362, U.S. Pub. 2006 / 0024681, U.S. Pub.2006 / 0292611, U.S. Pat. 7,960,120, U.S. Pat. 7,835,871, U.S. Pat. 7,232,656, U.S. Pat.7,598,035, U.S. Pat. 6,306,597, U.S. Pat. 6,210,891, U.S. Pat. 6,828,100, U.S. Pat. 6,833,246, and U.S. Pat. 6,911,345, each incorporated by reference. In certain embodiments, an Illumina Mi-Seq sequencer is used.
[0049] Sequencing 127 creates sequence reads, i.e., a record of a sequence of bases from at least a part of a nucleic acid. The sequence reads may be analyzed to determine expression of RNA associated with genes based on unique reads that correspond to those genes. Analyzing the sequence reads may be performed using known software and following multistep procedures thatare known in the art. For example, first, the quality of each sequence read, i.e., FASTQ sequence, may be assessed using the software FASTQC. Next, the reads may be trimmed using, for example, using Trimmomatic software. See Bolger, 2014, Trimmomatic: a flexible trimmer for Illumina sequence data, Bioinformatics 30(15):2114-2120, incorporated by reference. The trimmed sequence reads may then be mapped to a human genome using with, for example, HISAT2 software. HISAT2 output files in a SAM (sequence alignment / map format), which may be compressed to binary sequence alignment / map files. Other methods useful for processing and analyzing sequence reads are discussed in U.S. Pat. No. 8,209,130, which is incorporated by reference. Determining gene expression generally involves counting numbers of unique sequence reads that uniquely map to a human reference genome. Mapping reads to a reference to identify genes may be performed using computer software packages known in the art.
[0050] An important benefit of the presently described techniques is that mapping reads to a reference and identifying genes gives a quantitative result when reads are deduplicated to yield one read per mRNA from which those reads originated. Because each mRNA is typically copied into cDNA and each cDNA is typically copied into an unpredictably large number of amplicons in the sequencing library, and because each library member is often amplified or read redundantly as part of a sequencing technique, a number of raw sequence reads does not necessarily correlate to numbers of input molecules from the single cells. Nevertheless, one cell may include abundant transcripts that map to one gene. Here, compositions and methods of the presently described techniques give each cDNA a unique intrinsic identifier that can be identified within, and used to deduplicate, sequence reads. After those sequence reads are identified by gene and deduplicated, then counts of those reads provide a quantitative measure of gene expression levels.
[0051] As discussed, methods of the presently described techniques are useful for scRNA-Seq and specifically for expression analysis. In certain embodiments, cells are isolated into, and lysed within, aqueous partitions with capture oligos. The capture oligos anneal to RNAs released from the cells. The capture oligos may include partition-specific barcodes and PCR handles. Once the capture oligos have hybridized to the RNAs, those duplexes may be released from partitions and pooled at any subsequent stage. Because capture oligos with partition-specific barcodes are used to capture and tag RNA from cells isolated in the partition, any arbitrary number of cells may be captured in parallel (simultaneously). Because the RNAs are tagged witha cell barcode during hybrid capture (e.g., aka a partition-specific barcode or a bead barcode), if those duplexes are pooled and ultimately sequenced, the cell barcodes in the sequencing data can be used to “bin” the sequence data by original cell, i.e., assign each sequence read (or assembled contigs or sequences therefrom) back to originating cells.
[0052] Multiplexing may involve isolating cells and the capture oligos into partitions. Any suitable partitions may be used. The partitions may be any suitable partition in a pico-, nano-, or microtiter plate or substrate, or fluidic harbors (see, e.g., US Pub 2010 / 0041046 Al, incorporated by reference), chambers (see, e g., 20210178395 Al, incorporated by reference), regions defined within a fluidic device (see, e.g., 20200269248 Al, incorporated by reference), others, or combinations thereof. In certain embodiments, the partitions are aqueous partitions in an immiscible liquid, e.g., slugs or droplets surrounded or separated by oil within a microfluidic device. A microfluidic device may use channels to mix samples and reagents and form droplets in an immiscible carrier fluid. In certain embodiments, the partitions are a plurality of droplets that are formed essentially simultaneously. Methods may be performed with a sample comprising a mixture with cells, and in certain embodiments template particles. The mixture may include two immiscible fluids such as an aqueous fluid and oil. The mixture is sheared, e.g., vortexed, to generate an emulsion with template particles that serve to template the formation of droplets and segregate individual cells into the droplets. Because the cells are individually segregated into droplets, the cells may be individually profded in parallel. This method provides a massively parallel, analytical workflow for analyzing single cells that is inexpensive, scalable, and accurate.
[0053] For example, methods of the presently described techniques may include combining template particles with cells in a first fluid and then adding a second fluid that is immiscible with the first fluid to the mixture. The first fluid may be an aqueous fluid. While any suitable order may be used, in some instances, a tube may be provided comprising the template particles. The tube can be any type of tube, such as a sample preparation tube sold under the trade name Eppendorf, or a blood collection tube, sold under the trade name Vacutainer. The sample may be a blood sample and may be added directly to the tube using a pipette.
[0054] The fluids can be sheared to generate a monodisperse emulsion with droplets. To generate a monodisperse emulsion, methods includes a step of shearing the mixture provided by combining cells and template particles in an aqueous fluid with the immiscible fluid. Anysuitable method or technique may be utilized to apply a sufficient shear force to the mixture. For example, the mixture may be sheared by flowing the second mixture through a pipette tip. Other methods include, but are not limited to, shaking the mixture with a homogenizer (e.g., vortexer), or shaking the mixture with a bead beater. In some embodiments, vortex may be performed for example for 30 seconds, or in the range of 30 seconds to 5 minutes. The application of a sufficient shear force breaks the mixture into monodisperse droplets that encapsulate one of a plurality of template particles.
[0055] After vortexing, a plurality (e.g., thousands, tens of thousands, hundreds of thousands, one million, two million, ten million, or more) of aqueous partitions is formed essentially simultaneously. Vortexing causes the fluids to partition into a plurality of monodisperse droplets. A substantial portion of droplets will contain a single template particle and a single target cell. Droplets containing more than one or none of a template particle or target cell can be removed, destroyed, or otherwise ignored.
[0056] The next step of the method is to lyse the cells. Cell lysis may be induced by a stimulus, such as, for example, lytic reagents, detergents, or enzymes. Reagents to induce cell lysis may be provided by the template particles via internal compartments. In certain embodiments a, lysing involves heating the monodisperse droplets to a temperature sufficient to release lytic reagents contained inside the template particles into the monodisperse droplets. This accomplishes cell lysis of the target cells, thereby releasing nucleic acids, such as RNA, including mRNA, inside of the droplets that contained the target cells.
[0057] After lysing target cells inside the droplets, mRNA is released. The mRNA may be used to create a sequencing library. Methods and systems of the presently described techniques may use template particles to template the formation of monodisperse droplets and isolate single target cells. The disclosed template particles and methods for targeted library preparation thereof leverage the parti cle-templated emulsification technology described in Hatori, 2018, Parti cle-templated emulsification for microfluidics-free digital biology, Anal Chem 90(16):9813-9820, incorporated by reference. Essentially, micron-scale beads (such as hydrogels) or “template particles” are used to define an isolated fluid volume surrounded by an immiscible partitioning fluid and stabilized by temperature insensitive surfactants.
[0058] In practicing the methods as described herein, the composition and nature of the template particles may vary. For instance, in certain aspects, the template particles may bemicrogel particles that are micron-scale spheres of gel matrix. In some embodiments, the microgels are composed of a hydrophilic polymer that is soluble in water, including alginate or agarose. In other embodiments, the microgels are composed of a lipophilic microgel.
[0059] FIG. 2 illustrates a sample prep tube 229 comprising droplets 201. In particular, the sample prep tube 229 comprises a plurality of monodisperse droplets generated by shearing a mixture 239 according to certain embodiments of the presently described techniques. In certain embodiments, each of the droplets 201 includes, on average, one template particle 213 and zero or one single target cell 209. The template particles 213 may comprise crater-like depressions (not shown) to facilitate capture of single cells 209. The template particles 213 may further comprise an internal compartment 221 to deliver one or more reagents into the droplets 201 upon stimulus.
[0060] In some embodiments, the template particles contain internal compartments. The internal compartments of the template particles may be used to encapsulate reagents that can be triggered to release a desired compound, e.g., a substrate for an enzymatic reaction, or induce a certain result, e.g. lysis of an associated target cell. Reagents encapsulated in the template particles’ compartment may be without limitation reagents selected from buffers, salts, lytic enzymes (e.g. proteinase k), other lytic reagents (e. g. Triton X-100, Tween-20, IGEPAL), nucleic acid synthesis reagents, or combinations thereof.
[0061] Lysis of single target cells occurs within the monodisperse droplets and may be induced by a stimulus such as heat, osmotic pressure, lytic reagents (e.g., DTT, betamercaptoethanol), detergents (e.g., SDS, Triton X-100, Tween-20), enzymes (e.g., proteinase K), or combinations thereof. In some embodiments, one or more of the said reagents (e.g., lytic reagents, detergents, enzymes) is compartmentalized within the template particle. In other embodiments, one or more of the said reagents is present in the mixture. In some other embodiments, one or more of the said reagents is added to the solution comprising the monodisperse droplets, as desired.
[0062] In certain embodiments, template particles 213 comprise a plurality of capture probes. Generally, the capture probe of the present disclosure is an oligonucleotide. In some embodiments, the capture probes are attached to the template particle’s material, e.g. hydrogel material, via covalent acrylic linkages. In some embodiments, the capture probes are acrydite-modified on their 5’ end (linker region). Generally, acrydite-modified oligonucleotides can beincorporated, stoichiometrically, into hydrogels such as polyacrylamide, using standard free radical polymerization chemistry, where the double bond in the acrydite group reacts with other activated double bond containing compounds such as acrylamide. Specifically, copolymerization of the acrydite-modified capture probes with acrylamide including a crosslinker, e.g. N,N'-methylenebis, will result in a crosslinked gel material comprising covalently attached capture probes. In some other embodiments, the capture probes comprise acrylate terminated hydrocarbon linker and combining the said capture probes with a template particle will cause their attachment to the template particle.
[0063] In some embodiments, after cell suspensions are introduced to template particles in a pre-equilibrated buffer, droplets are generated by vortexing the mixture to capture single cells with individual template particles. The resulting emulsion may be heated on a thermocycler to induce cell lysis. Cell lysis releases the contents of the cell and exposes those contents, including mRNA, to the particle 213. The presently described techniques provide steps for RNA capture and library preparation using those particles and released mRNA.
[0064] FIG. 3 diagrams RNA capture and library preparation according to methods of the presently described techniques. As shown, particle 213 is linked to a capture oligo 305. The capture oligo 305 anneals to an mRNA 311. In some embodiments, poly-T tails of the capture oligos anneal to and capture RNA released by lysis. Particle-bound capture oligos in this application may comprise an acrydite linker, a PEI priming sequence, a particle barcode, optionally a random sequence, and a poly-T capture moiety. A polymerase (not pictured) extends the capture oligo 305 to form a cDNA 315. The cDNA 315 and capture oligo 305 in combination with the mRNA 311 form a duplex 323. This duplex is stably linked to the bead 213. At this stage, it is suitable to break the droplets and pool their contents, wash in buffer, and proceed in library preparation.
[0065] A transposase complex 325 (sometimes called a transposasome) is introduced. The transposase complex 325 includes a dimer that includes two of a transposase 327 and two transposon end sequences 329. Here, the transposon end sequences 329 are depicted as both being paired-end 2 end (PE2) sequences, which will cooperate with paired-end 1 (PEI) sequences in the capture oligo 305 in subsequent amplification and sequence steps. In the depicted method, the transposase randomly cuts the cDNA / mRNA duplex 323 thereby defining arandom cut site 333. Tn a downstream step, read 2 of paired-end sequencing will include the first segment of bases in the cDNA 315 adjacent the random cut site 333.
[0066] Attachment 115 of the end sequence 329 to the cDNA 315 at the random cut site 333 produces a construct 337. The construct 337 is a contiguous DNA molecule that includes a first PCR handle (PEI), a cell barcode, a capture segment, a portion of the cDNA 315 terminating at the random cut site 333, and a second PCR handle (PE2).
[0067] Amplification 123 of the construct 337 yields amplicons 341. In some embodiments, constructs are amplified with a P5-PE1 hybrid oligo and P7 index primer directly into a sequencing library. The library may be sequenced to assess RNA expression, for example, as described in Hrdlickova, 2017, RNA-Seq methods for transcriptome analysis, Wiley Interdisc Rev RNA 8(1): 10.1002, incorporated by reference.
[0068] Constructs or amplicons may include certain primer and index sequences or copies thereof, such as, P5s and P7s. Those sequences may be any arbitrary sequence useful in downstream analysis. For example, there may be additional universal primer binding sites or sequencing adaptors. For example, either or both of the P5s and P7s may be arbitrary universal priming sequence (universal meaning that the sequence information is not specific to the naturally occurring genomic sequence being studied, but is instead suited to being amplified using a pair of cognate universal primers, by design). The index segment may be any suitable barcode or index such as may be useful in downstream information processing. It is contemplated that the P5 sequences, the P7 sequence, and the index segment may be the sequences use in NGS indexed sequences such as performed on an NGS instrument sold under the trademark ILLUMINA, and as described in Bowman, 2013, Multiplexed Illumina sequencing libraries from picogram quantities of DNA, BMC Genomics 14:466 (esp. in Figure 2), incorporated by reference.
[0069] Importantly, a transposase 327 is used to randomly cut the cDNA 315. This may be performed using a transposase such as Tn5. See Lin, 2020, RNA sequencing by direct tagmentation of RNA / DNA hybrids, PNAS117 (6) 2886-2893, incorporated by reference. In brief, the Tn5 transposase randomly binds and cuts double-stranded RNA / DNA and attaches its end sequence to the random cut site.
[0070] Accordingly, some embodiments of the presently described techniques use Tn5 transposase to directly tagment RNA / DNA hybrids and form polynucleotide libraries withintrinsic molecular identifiers (essentially unique sequences of bases originating in genetic material of the organism or biological system being studied). In particular, Tn5, a RNase H superfamily member, binds to RNA / DNA hybrids similarly as to dsDNA and effectively cuts randomly and then ligates a desired oligo onto the hybrid. The desired oligo is, in certain embodiments, a PCR handle (aka a universal primer binding site, a sequencing adaptor, a synthetic oligo of known sequence to which a PCR primer anneals, etc.). Methods of the present techniques may be used with various amounts of input sample, from single cells to large numbers of cells, with a dynamic range spanning numerous orders of magnitude.
[0071] FIG. 4 shows a workflow for directional tagmentation that works with template switching oligos (TSOs). The illustrated technique may be employed in hybrid workflows where one is using TSOs for some other benefit and one also want to use IMIs. This tagmentation approach is useful for 3’ end capture and analysis of mRNAs. The steps of the method are shown. In brief, mRNA or total RNA from lysed cells are mixed with an oligo and incubated at 65° for 3 min. The oligo may include specific primers for amplifying final libraries, such as an adapter-B sequence complementary to an i7 primer. The oligo may further include a poly-T sequence of, for example, 30 nucleotides that hybridizes with poly-A tails of mRNA.Importantly, the use of this oligo to prime a first strand cDNA synthesis may result in libraries enriched for the 3' end of mRNA.
[0072] Reverse transcription can be performed using a reverse transcriptase such as the reverse transcriptase sold under the trade name SMARTSCRIBE by Takara Bio optionally in the presence of a template switching oligo (TSO). The template switching oligo allows for template switching at the 5' end of the mRNA molecule to incorporate an oligo such as a universal 3' sequence during first strand cDNA synthesis. Synthesis of the first cDNA strand may be performed using a thermocycler at 42 degrees Celsius for Ih, followed by 15 minutes at 70 degrees Celsius to inactivate the reverse transcriptase. Afterwards, the cDNA may be amplified. The cDNA may be amplified by PCR using commercially available kits such as the kit sold under the trade name OneTaq HS by New England Biolabs. After amplification, the RNA / DNA duplexes may be subjected to tagmentation and adapter ligation.
[0073] During tagmentation and adapter ligation, Tn5 bound adapter (adapter- A) complexes bind with the double RNA / DNA duplexes. The duplexes are cut by the enzymatic activity of the Tn5 complexes and the adapters (“Adaptor A”) are ligated. Importantly, Tn5 cuts at a randomsite. Afterwards, the products of the tagmentation reaction may be amplified using the adapters. As shown, an i7 primer anneals to Adapter B and an i5 primer anneals to Adapter A. In this depicted embodiment (as drawn) the read 1 primer will read into a segment of an amplicon adjacent the random cut site. Because the cut site is random, a sequence of bases in that segment is essentially unique. Because the segment is in the amplicon copy of the cDNA, itself a copy of the mRNA, the sequence of the basis is intrinsic to the mRNA, i.e., is a sequence from genetic material of the organism being studied. Because the sequence of bases is essentially unique, a read 1 sequence read will include a unique, intrinsic molecular identifier. More specifically, all sequence reads from the read 1 primer from this library member will include the identical copies of that unique, intrinsic molecular identifier (IMI). Thus the figure illustrates that IMIs are compatible with workflows that include or use TSOs.
[0074] After sequencing, reads with identical gene-mapping and identical IMIs can be “collapsed”, and a count of only unduplicated such reads is a quantitative measure of gene transcripts in the sample, i.e., the single cell.
[0075] As shown in FIG. 4, the use of IMIs is compatible with RNA capture without necessarily requiring any bead-linked capture oligos. That is, capture oligos may be free in solution (as opposed to linked to a solid support). Methods of the present techniques are also compatible with the use of capture oligos that are linked to a solid support such as a bead.
[0076] FIG. 5 shows a method for making libraries that include IMIs using capture oligos linked to a solid support. This embodiment shows the creating of a sequencing library that includes certain next-generation sequencing (NGS) adaptors. The solid support may be bead and library preparation may be performed using a microfluidic device (e.g., to encapsulate beads, cells, and reagents into droplets). In some embodiments, beads decorated with capture oligos are used to simultaneously form a monodisperse emulsion that includes a plurality of droplets. Each droplet includes, on average, one bead and one or zero cells. Because the beads are particles that serve as templates cause the droplets (or aqueous partitions) to form (e.g., when a mixture is vortexed), the droplets may be referred to as particle-templated instant partitions (PIPs), the beads may be referred to template particles, and sequencing from such libraries may be referred to as PIP-seq. In the illustrated embodiment, a template particle 1301 is linked to a capture oligo 1305. As shown, the particle 1301 is linked to (among other things) mRNA capture oligos 1305 that include a 3’ poly-T region 1309 (although sequence-specific primers or random N-mers maybe used). Where the sample includes cell-free RNA, the capture oligo hybridizes by Watson-Crick base-pairing to a target in the RNA and serves as a primer for reverse transcriptase, which makes a cDNA copy of the RNA. Where the initial sample includes intact cells, the same logic applies but the hybridizing and reverse transcription occurs once a cell releases RNA (e.g., by being lysed).
[0077] In certain embodiments, the target RNAs are mRNAs 1313. Where the target RNAs are mRNAs, the particles 1301 may include mRNA capture oligos 1305 used to at least synthesize cDNA 1317 as a copy of an mRNA 1313. The particles 1301 may further include cDNA capture oligos with 3’ portions that hybridize to cDNA copies of the mRNA. For the cDNA capture oligos, the 3’ portions may include gene-specific sequences or hexamers. As shown, each of the mRNA capture oligos 1305 may include, from 5’ to 3’, a SMART site 1319, a PEI sequence 1321, a cell or droplet barcode 1323, and a poly-T segment 1309. Optionally, the capture oligos may include a UMI 1311.
[0078] As shown, the capture oligo 1305 hybridizes to the mRNA 1313. A reverse transcriptase binds and initiates synthesis of a cDNA copy 1317 of the mRNA 1313 to make an RNA / DNA hybrid. Note that the mRNA 1313 is connected to the particle 1301 non-covalently, by complementary base-pairing. The cDNA 1317 that is synthesized may be covalently linked to the particle 1317 by virtue of the phosphodiester bonds formed by the reverse transcriptase.
[0079] A transposase 1401 binds to the RNA / DNA hybrid. The transposase 1401, which may be a Tn5 transposase, is attached with adapters 1406 for attaching onto the 5’ end of the cDNA 1317. The Tn5 cuts the RNA / DNA hybrid at a random cut site and the adapters 1406 are ligated onto the random cut site of the cDNA 1317. In certain embodiments the adapter 1406 includes a primer handle 1403 for copying / amplification. At this stage, RNaseH may be introduced to degrade the mRNA 1313.
[0080] In some embodiments, sequencing adapter 1501 is extended to create a dsDNA 1409. The adapter 1501 includes a first sequence 1503 complementary to the primer handle 1403 and a sequencing primer 1505, such as P7. The adapter 1501 will hybridize to, and prime the copying of, cDNA to create a dsDNA 1409 with the sequencing adapter. Afterwards, the polynucleotide can be separated from the particle and made into a final library product. What is important is that the cDNA 1317 and thus also the dsDNA 1409 has a segment adjacent the random cut site 1351with sequence intrinsic to the mRNA 1313. When that segment is sequenced from primer handle 1403, the resultant sequence reads include the intrinsic sequence.
[0081] Amplification produces a final library product 1601. In this example, the final library product 1601 is formed by the PCR-based extension a P5-PE1 primer 1505 that is complementary to the PEI 1509 of the released polynucleotide 1409. Extension of the P5-PE1 primer 1505 by PCR creates the final library product 1601. In some embodiments, the P5-PE1 primer 1505 may include indexes, such as an 15 index, and a P5 index. The final library product may be amplified by PCR in advance of sequencing.
[0082] FIG. 6 shows a non-limiting example of a polynucleotide construct or an amplicon as described above, that allows for four individual sequence reads from paired-end sequencing. The polynucleotide construct 615 may include a first and second sequence platform specific primer binding sites (e.g., P5 and P7), a first and second index, a read 1 primer binding site, a read 2 primer binding site, barcodes, a poly-T sequence, and the transcript comprising the IMI. During paired-end sequencing, sequencing primer 607 may bind to the read 2 binding site, and generate a sequence read 2 that includes the second index. A sequencing primer 609 may also bind to the read 2 binding site and generate an alternative sequence read 2 that includes the transcript comprising the IMI. A sequence primer 611 may bind to the read 1 primer binding site and generate a sequence read 1 that includes the first index. A sequence primer 613 may also bind to the read 1 primer binding site and generate an alternative sequence read 1 that includes the barcode.
[0083] FIG. 7 shows an alternative non-limiting example of a polynucleotide construct or an amplicon that allows for two individual sequence reads from single-directional sequencing. The polynucleotide construct 619 include a first and second index, a read 1 primer binding site, a read 2 primer binding site, barcodes, a poly-T sequence, and a transcript comprising the IMI. During single-directional sequencing, a sequencing primer 615 may bind to the read 1 primer binding site and generate a sequence read 1 that includes the first index and the transcript comprising the IMI. A sequencing primer 617 may bind to read 2 primer binding site and generate a sequence read 2 that includes the barcodes and the second index. In such methods involving single directional sequencing, the IMIs are read first in sequence read 1 and preserved to prevent running out of sequencing length. Further, (i) there is no longer a need to read through long anddifficult to resolve poly A or poly T sequences and (ii) barcode information is preserved as first bases read in sequence read 2.
[0084] In certain embodiments, methods of the disclosure include using additional configurations of polynucleotide constructs and amplicons, as shown in the alternative example above, for IMI based analysis via single-directional sequencing. The alternative polynucleotide or amplicon has advantages for single direction sequencing technology such as the sequencing technology sold under the trademark ULTIMA by Ultima Genomics.
[0085] Single direction sequencing technology may include sequencing technology that produces single-read, flow-based data by flowing one nucleotide at a time in order, iteratively. Such single direction sequencing technology is in contrast to traditional technologies that flow all four nucleotides at once. The iterative approach of single direction sequencing ensures that only one dNTP is responsible for the signal and it does not require the blocking of dNTPs. Such a sequencing platform is described in Almogy G, Pratt M, Oberstrass F, Lee L, Mazur D, Beckett N, et al. Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform, [preprint]. Genomics (2022), incorporated by reference.
[0086] FIG. 8 shows a non-limiting example of a method of the disclosure further comprising enriching a target RNA. In some embodiments, methods of the disclosure comprising enriching a target nucleic acid (e.g., RNA) further comprise magnetic bead pulldown. For the method further comprising targeted RNA enrichment, IMI based PIP-seq library preparation is performed as described above to generate a full sequencing-ready library comprising the whole 3’ transcriptome library with polynucleotide constructs or amplicons, each comprising an IMI.
[0087] The library pool comprising polynucleotide construct 621 may then be incubated with one or more gene specific primers 623 (e.g., fishing primer). Each gene specific primer 623 (e.g., fishing primer) is designed to bind a specific region of the target gene desired to be enriched. The gene specific primer 623 (e.g., fishing primer) is designed to achieve a desired annealing temperature (e.g., about 50°C) and include a 5’ biotinylation modification and 3’ non-extensible bases. The gene specific primer 623 annealed to the target gene may then be captured onto streptavidin modified magnetic beads.
[0088] After the gene specific primers annealed to the target genes are bound to the streptavidin modified magnetic beads, the magnetic beads are retained by a magnetic field to allow removal of unbound non-target genes and washed to remove the non-specific binders. The final material is re-suspended in a small volume. The magnetic beads are then isolated from the supernatant, and a population of specifically enriched target genes are captured. The resulting bead bound population of specifically enriched target genes may be then further amplified by PCR, using universal amplification primers designed to target sequence platform specific primer binding sites 625 and 627 or amplification handles (e.g., P5 and P7).
[0089] Any remaining gene specific primers are isolated from the PCR enrichment via 2 mechanisms: (i) the primers do not have extensible 3’ bases, and (ii) the primers are designed to anneal at lower temperatures than the P5 and P7 specific amplification handles. This mechanism of universal amplification preserves all necessary barcoding information without risk of cross talk or barcode swapping. Indexing barcodes, cell barcodes, and IMIs are all preserved in target gene enriched libraries.
[0090] The preceding offers guidance as to the use of intrinsic molecular identifiers (or other suitable unique identifiers) in the context of deduplication, library preparation, and nucleic acid sequence analysis. As discussed herein, in additional embodiments techniques for assessing the presence or absence of supplementary nucleotide tags may be incorporated to further improve performance and efficiency. By way of example, such supplementary nucleotide tags may be more generally characterized as cell-associated features and may correspond to features or characteristics that may vary at the single-cell level and for which such information may be useful, such as for disease diagnosis, treatment or pharmaceutical development, patient counseling, and so forth. By way of example, such supplementary nucleotide tags may include, but are not limited to: CRISPR guide RNA (gRNA), gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
[0091] By way of context, various single-cell analyses involve associating transcriptomic data with supplementary nucleotide tags. In general, correct identification of the cell of origin for any supplementary nucleotide tag (SNT) is necessary to assign functional biology to that cell. In this context, such supplementary nucleotide tags may correspond to a variety of cellmodifying and / or cell-associated features and may be relevant to a multitude of multiomicapplications. By way of a representative example, CRISPR guide RNA (gRNA) represents one such supplementary nucleotide tag (SNT) and may be used in massively multiplexed perturbational analysis at single-cell resolutions. In such contexts, CRISPR screening SNTs may be analyzed as polyadenylated guide reports (as in CROP-Seq assays) or as direct captured guides (as may be implemented in certain single-cell CRISPR screens). More generally, presence or absence of an SNT based on a threshold may be useful in evaluating the function of individual gene edits in CRISPR screens, quantification of surface epitopes in CITE-Seq applications, or disambiguation of sample of origin for multiplexes sample hashing workflows, to list a handful of suitable representative applications. However, it should be understood that other applications where cell-associated features at the single-cell level are to be evaluated are also contemplated and may be facilitated in accordance with the present techniques.
[0092] In single-cell sequencing applications, one aspect of assay performance is sequencing efficiency, such as the number of information-containing sequencing reads that may be confidently associated with specific cells, as opposed to sequences that cannot be confidently associated with cells and are therefore attributed to background noise. Such background characterizations of sequencing reads may reflect various assay inefficiencies, including premature cell lysis, carryover of ambient SNT labels into the reactions, or redistribution of SNT reads among cell barcodes during the sequencing library preparation. Ideally, sequencing library workflows are optimized to minimize such sources of background noise.
[0093] Typically, sequencing analysis processes are designed or optimized to distinguish, often in accordance with a threshold, confidence between cell-associated and background (e.g., cell-free or cell-independent) sequencing reads. Such processes are particularly useful if the number of cells associated with individual SNTs are sparse and / or the number of SNTs associated with an individual cell are sparse relative to the number of cells per sample or the number of sequencing reads per cell, respectively.
[0094] With the preceding in mind, in certain embodiments described herein techniques are described for facilitating cell assignment (e.g., positive or negative cell characterization for association with an SNT). Such techniques as described herein may be performed using operations performed in parallel or other suitable multi-threaded implementations and may be performed on data acquired form a nucleic acid sequencing instrument. Further, the presently described techniques offer technical improvement over conventional approaches due to theimproved sensitivity afforded in low-expression or limited-sequencing regimes (i.e., low read or sequencing depth). In this manner, the presently described approach extends beyond the use case of existing conventional approaches.
[0095] Such cell assignment calls may be made, in certain implementations, with a specified or threshold degree of confidence. In certain embodiments suitable statistical clustering techniques, such as k-means clustering, may be employed. The described techniques may be generally applicable across SNT applications in single-cell or spatial applications, such as, but not limited to, cell assignment based on gRNA presence in single-cells in CRISPR Perturb-Seq applications. In certain embodiments, read counts may be corrected or de-noised (such as in accordance with the techniques taught in Characterization and bioinformatic filtering of ambient gRNAs in single-cell CRISPR screens using CLEANSER, Liu, Siyan et al.: Cell Genomics, Volume 5, Issue 2, 100766, incorporated by reference herein in its entirety for all purposes) as part of or prior to implementation of the described techniques. With respect to the presently described techniques, discussion is provided in the context of gRNA characterization in singlecell contexts for the purpose of illustration and to provide a useful, real-world context to facilitate explanation. As noted above, however, such discussion in the gRNA context is solely for the purpose of illustration and providing a real-world context. It will be appreciated that the techniques as described are readily transferrable to the other SNT contexts described herein as well as to other cell-associated feature analyses.
[0096] In the context of a CRISPR assay, work was done using conventional approaches to perform gRNA assessment of single-cells. Based on this work it was observed that such conventional approaches were unable to reliably threshold gRNA at read depths below 5,000 reads per cell for data with low signal moise (e.g., approximately 1 : 1 or less, or cases where the counts distribution of signal and noise largely overlaps and cannot be reliably separated by gaussian mixture modeling, or other common algorithms). Further, conventional approaches were unable to facilitate the detection of multiple gRNA classes per cell (e.g., 2 or more gRNA classes per cell). Conversely, the present technique, as described below, is able to reliably threshold gRNA at read depths below 5,000 reads per cell at low signal moise levels and is also able to facilitate detection of single or multiple gRNA classes in individual cells.
[0097] Certain conventional techniques perform thresholding using log-transformed read counts. Such approaches, however, result in mixing signals of cells with a high-purity guidesignal with those having lower purity guide signals but a high abundance of counts. This is addressed in the techniques described herein by taking the read counts from gRNAs (or other cell-associated features) in each cell and dividing this value by the sum of all gRNA read counts for the respective cell, resulting in what is hereafter referred to as a “guide purity fraction” or a “feature purity fraction” in more general contexts.
[0098] With this in mind, for each class of gRNA a guide purity fraction (which may be a ratio in certain embodiments) is obtained for each cell in the population of cells. In other, nonlimiting embodiments this may be understood to be a feature purity fraction for a given cell-associated feature, such as an SNT. Determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell may, in certain embodiments, be achieved by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell. In general, this may be a ratio of the number of reads associated with a true positive SNT divided by the total number of SNT counts. It may be expected that SNT counts originating from background contributions would have a low ratio, whereas true cell associated counts would have a high ratio by this metric. In the context of gRNA, if a respective cell expresses a single class of gRNA with high purity, it may have a guide purity (or feature purity) fraction of 1. Conversely, the lowest possible value of 0 may be assigned for 0 counts of a given gRNA class (or other cell-associated feature).
[0099] A statistical clustering technique, such as k-means clustering or other data segmentation approaches, may then be applied to the guide purity fractions (or, more generally, the feature purity fractions) for each respective class of guides or features. A k of 2 may be used by default, though other values of k, such as 3, may be used in other embodiments. For a £ of 2, two populations will be called by the k-means clustering and a threshold value may be determined as the boundary separating the two populations. Cells having guide purity fractions (or, more generally feature purity fractions) above the threshold may be called as positive for expression of the respective class of gRNA (or cell-associated feature). Cells below the threshold may be classified as negative or otherwise characterized as background noise. As this approach is performed per class of gRNA or cell-associated feature and per individual cell, respective single cells may be called as positive for more than one respective cell associated feature. By way of example, a respective single cell can be called as positive for 1, 2 or more respective classes of gRNA in a gRNA characterization context.
[0100] It should also be appreciated that additional parameters may be incorporated into this approach, such as a normalized scaled expression for each cell-associated feature (e.g., gRNA). Further, in certain embodiments constraints may be added to prevent or reduce cell-associated feature signal below a threshold (e.g., low guide signal) from being included in the positive cell fraction in the event there are no true positive cells for a particular cell-associated feature.
[0101] With respect to implementation specific terminology and by way of an example calculation, a further gRNA example is provided by way of elaboration of the presently disclosed techniques. In this example, usable gRNA reads may be understood to be:(1) Usable gRNA reads = Input Reads * Barcoding Rate * gRNA Tag-Matching Rate * gRNA Reads in Cells PercentWhere “Barcoding Rate” is the percentage of cell barcodes that match a predetermined white-listed barcode and “gRNA Tag-Matching Rate” is equivalent to “mapping” or matching to a feature barcode in a given read. The barcoding and tag-matching rates can vary widely between different cell types and / or experiments. Further, gRNA reads / cell may, in certain embodiments, be given by:(2) Usable gRNA reads per cell = (Usable gRNA reads in Called Cells) / (# of Called Cells)where “Called Cells” are the number of cells in an assay used to generate the single-cell sequencing data. In practice, the number of cells in the assay may be determined based on a paired gene expression sample or a gene-expression-free assay. In this example, gRNA reads per cell may be dependent on assay kit size, the number of cells input by the user, and the cell capture rate.
[0102] In terms of more generalized guidance, barcoding rate and / or gRNA tag matching rate may be disregarded and an assumption instead made that -5,000 gRNA reads / called cell generally gives sufficient data to accommodate for read loss via the above discussed techniques. In such a context, the effective usable gRNA reads per cell value will typically be around 25% of the original input gRNA reads / called cell (i.e., closer to 1250 usable gRNA reads / called cell). Equations (1) and (2) generally provide further specificity as to what data is most useful and why in the present analytic construct. Correspondingly, equations (1) and (2) technically describewhat are essentially the “usable guide reads per cell” for the sake of explanation and clarity. It may be understood, though, that as used herein when a given number of reads per cell (e.g., gRNA reads per cell) is referenced, this number actually refers to the input reads / cell (e.g., gRNA reads per cell) rather than the usable gRNA reads / cell unless usable reads are otherwise specified.
[0103] Turning to FIGS. 9A and 9B, a comparison of results for the presently described fractional scaling approach (FIG. 9B) versus a conventional log-transformed counts approach (FIG. 9A) is depicted. Distributions corresponding to gRNA signal for individual cells are denoted by a check mark if the positive-negative population subsets can be distinguished from one another (i.e., if the bimodality of the populations is discernible). As can be seen, the presently disclosed approach using fractional scaling yields greater discrimination between the positive-negative population subsets.
[0104] Similarly, and turning to FIGS. 10A, 10B, 10C, and 10D, plots depicting results of the presently disclosed techniques are depicted for different classes of gRNAs. In these examples, counts ratios (from 0 to 1) are depicted along the y-axis for each gRNA class and represent the fraction of total counts belonging to each gRNA class as described herein. The x-axis depicts the log-scaled counts. Individual clusters or dots in each plot are sorted into high expression (above the depicted horizontal threshold line) and low expression (below the depicted horizontal threshold line) for the respective gRNA class. As illustrated, each gRNA class, based on the k-means cluster analysis, has a different respective derived threshold (i.e., a dynamic threshold that may be calculated on a per sample or experimental run basis) which, once derived may be used to separate or otherwise distinguish positive expressing cells from negative. In this manner, and based upon the derived respective threshold, the high expression single-cells can be distinguished from the low expression single-cells, which may be considered as noise or background for the purpose of analysis.
[0105] Further, the presently described techniques, unlike conventional approaches, continue to perform well at below full read depth (i.e., the number of gRNA reads / cell). For the purpose of this discussion, full read depth may be equated to approximately 5,000 gRNA reads per cell, though in practice full read depth may be higher. Turning to FIGS. 11 A and 1 IB, this is illustrated by plots generated using the presently described approaches for four different classesof gRNA. In FIG. 11 A, the depicted plots illustrate the separation of the high- and low-expression population subsets (and corresponding dynamic thresholds separating high- and low-expression population subsets) for 75% depth coverage (i.e., approximately 3,750 reads / cell). In FIG. 1 IB, the depicted plots illustrate the separation of the high- and low-expression population subsets (and corresponding dynamic thresholds) for 1% depth coverage (i.e., approximately 50 reads / cell). As may be observed, good separation is observed at 75% depth of coverage and is even useful at extremely low depth of coverage. In particular, while resolution is worse at lower read depth, the dynamic threshold obtained was approximately 0.50, which is a conservative threshold. In practice, and in accordance with the presently disclosed techniques, read depths at low signal to noise may be 3,000 reads / cell or less, 2,500 reads / cell or less, 2,000 reads / cell or less, 1,500 reads / cell or less, 1,000 reads / cell or less, 750 reads / cell or less, or 500 reads / cell or less. Alternatively, in accordance with the presently disclosed techniques, read depths at low signal to noise may be 50 reads / cell or greater, 100 reads / cell or greater, 150 reads / cell or greater, 200 reads / cell or greater, 250 reads / cell or greater, 300 reads / cell or greater, 350 reads / cell or greater, 400 reads / cell or greater, or 450 reads / cell or greater. Further, in accordance with the presently disclosed techniques, read depths at low signal to noise may be between approximately 100 to 4,000 reads per cell, 200 to 3,000 reads per cell, or 500 to 2,000 reads per cell.
[0106] With respect to comparison to conventional approaches, the presently described k-means thresholding approach was found (using intentionally low-quality data) to provide an intermediately low CV% for recovered singlets (i.e., a cell that is solely positive for a single class of gRNA) across read depths (i.e., downsampling rate) while still recovering more (e.g. -2,000 more) singlets than conventional approaches. In a further comparative study, the presently described k-means thresholding approach was found (again, using intentionally low-quality data) to provide the lowest CV% for recovered singlets across all read depths. In general, the presently described k-means thresholding approach was found to reliably recover -30% to -50% more singlets than conventional approaches. Further, it was observed that the presently described k-means thresholding approach provided singlet gRNA assignments that were generally stable across all read depths. Though variability may be observed for gRNAs with low expression.
[0107] By way of further context, certain prior conventional approaches (e g., Datlinger’s SNR approach) are gnostic to guide class and may mix all of the cells in the sample together and consider only whether a cell is above or below an arbitrary threshold for all guide classes. Such approaches mean that one is dealing with a difficult optimization curve and ultimately the user has to empirically test multiple threshold for every single sample or has to pick an arbitrary threshold, which is sub-optimal. Further, such approaches only allow for one guide assignment per cell.
[0108] Overview of System for Biological or Chemical Analysis:
[0109] As may be appreciated, in practice the above-described correction approaches would be implemented using a processor-based system, such as a next generation sequencing (NGS) system or a workstation, computer, or server in direct or indirect communication with such a system or a data repository storing results from such a sequencing system. As may be appreciated, such NGS systems or other downstream analytic devices may constitute or include specialized circuitry and components optimized for the acquisition and / or processing of nucleic acid sequence data, including in large quantities. Correspondingly, the above described operations and steps may be embodied as processor-executable code or routines on a tangible computer-readable medium that may be accessed by a suitable processor or circuitry (e.g., one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and so forth).
[0110] With the preceding discussion in mind, examples and embodiments described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction. Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No. 7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of whichis incorporated by reference herein, in its entirety and for all purposes; and / or U.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety and for all purposes.
[0111] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.
[0112] Examples of Fluidics Devices and Fluid Flow Paths - Example of System with Higher Volume Throughput
[0113] With the preceding in mind, FIG. 12 illustrates a schematic diagram of an example of a system 1100 that may be used to perform an analysis on one or more samples of interest. In practice, the presently described techniques may be performed in a parallel or multi-threaded manner using a device as described in FIG. 12 or by a processor-based device in communication downstream from a system 1100. A controller 1114 of the present example includes a user interface 1206, a communication interface 1208, one or more processors 1210, and a memory 1212 storing instructions executable by the one or more processors 1210 to perform various functions including the disclosed implementations. User interface 1206, communication interface 1208, and memory 1212 are electrically and / or communicatively coupled to the one or more processors 1210. User interface 1206 may be adapted to receive input from a user and to provide information to the user associated with the operation of system 1100 and / or an analysis taking place. User interface 1206 may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.
[0114] Communication interface 1208 is adapted to enable communication between system 1100 and a remote system(s) (e.g., computers) via a network(s) (e.g., the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellularnetwork, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system 1100. Some of the communications provided to system 1100 may be associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system 1100.
[0115] The one or more processors 1210 and / or system 1100 may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors 1210 and / or system 1100 includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.
[0116] Memory 1212 may include one or more of a semiconductor memory, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable readonly memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).
[0117] In some implementations, the sample may include one or more clusters of nucleotides (e g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shown, system 1100 is configured to receive a flow cell cartridge assembly 1102 including a flow cell assembly 1103 and a sample cartridge 1104. System 1100 includes a flow cell receptacle 1122 that receives flow cell cartridge assembly 1102, a vacuum chuck 1124 that supports flow cell assembly 1103, and a flow cell interface 1126 that is used to establish a fluidic coupling between system 1100 and flow cell assembly 1103. Flow cell interface 1126 may include one or more manifolds. System 1100 further includes a sipper manifold assembly 1106, a sample loading manifold assembly 1108, and a pump manifold assembly 1110. System 1100also includes a drive assembly 1112, the controller 1114, an imaging system 1116, and a waste reservoir 1118. Controller 1114 is electrically and / or communicatively coupled to drive assembly 1112 and to imaging system 1116; and is configured to cause drive assembly 1112 and / or the imaging system 1116 to perform various functions for performing the techniques disclosed herein.
[0118] In the present example, flow cell assembly 1103 includes a flow cell 1128 having a channel 1130 and defining a plurality of first openings 1132, which are fluidically coupled to the channel 1130 and arranged on a first side 1134 of the channel 1130. Flow cell 1128 further includes a plurality of second openings 1136 fluidically coupled to the channel 1130 and arranged on a second side 1138 of the channel 1130. Fluid may thus flow through flow cell 1128 via the channel 1130. While the flow cell 1128 is shown including one channel 1130, flow cell 1128 may include two or more channels 1130. Flow cell assembly 1103 also includes a flow cell manifold assembly 1140 coupled to flow cell 1128 and having a first manifold fluidic line 1142 and a second manifold fluidic line 1144. Flow cell manifold assembly 1140 may be in the form of a laminate including a plurality of layers as discussed in more detail below.
[0119] In the implementation shown, first manifold fluidic line 1142 has a first fluidic line opening 1146 and is fluidically coupled to each of the first openings 1132 of flow cell 1128; and second manifold fluidic line 1144 has a second fluidic line opening 1148 and is fluidically coupled to each of the second openings 1136. As shown, flow cell assembly 1103 includes gaskets 1150 coupled to flow cell manifold assembly 1140 and fluidically coupled to fluidic line openings 1146, 1148. In some implementations where flow cell 1128 includes a plurality of channels 1130, flow cell manifold assembly 1140 may include additional fluidic lines 1152 that couple first fluidic line openings 1146 to a single manifold port 1154. In such implementations, a single gasket 1150 may be coupled to flow cell manifold assembly 1140 that surrounds the manifold port 1154 and is in fluidic communication with a plurality of channels 1130. In operation, flow cell interface 1126 engages with corresponding gaskets 1150 to establish a fluidic coupling between system 1100 and flow cell 1128. The engagement between flow cell interface 1126 and gaskets 1150 reduces or eliminates fluid leakage between flow cell interface 1126 and flow cell 1128.
[0120] In the implementation shown, first manifold fluidic line 1142 has a portion 1156 that is substantially parallel to a longitudinal axis 1158 of channel 1130; and second manifold fluidicline 1144 has a portion 1160 that is substantially parallel to longitudinal axis 1158 of channel 1130. Additionally, first manifold fluidic line 1142 is shown being at least partially adjacent a first end 1162 of flow cell 1128 and spaced from a second end 1164 of flow cell 1128; and second manifold fluidic line 1144 is shown being at least partially adjacent second end 1164 of flow cell 1128 and spaced from first end 1162. Other arrangements of manifold fluidic lines 1142, 1144 may prove suitable, however.
[0121] In the implementation shown, system 1100 includes a sample cartridge receptacle 1166 that receives sample cartridge 1104 that carries one or more samples of interest (e.g., an analyte). System 1100 also includes a sample cartridge interface 1168 that establishes a fluidic connection with sample cartridge 1104. Sample loading manifold assembly 1108 includes one or more sample valves 1170. Pump manifold assembly 1110 includes one or more pumps 1172, one or more pump valves 1174, and a cache 1176. Valves 1170, 1174 and pumps 1172 may take any suitable form. Cache 1176 may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system 1100. While cache 1176 is shown being included in pump manifold assembly 1110, cache 1176 may alternatively be located elsewhere (e.g., in sipper manifold assembly 1106 or in another manifold downstream of a bypass fluidic line 1178, etc.).
[0122] Sample loading manifold assembly 1108 and pump manifold assembly 1110 flow one or more samples of interest from sample cartridge 1104 through a fluidic line 1180 toward flow cell cartridge assembly 1102. In some implementations, sample loading manifold assembly 1108 may individually load or address each channel 1130 of flow cell 1128 with a respective sample of interest. The process of loading channel 1130 with a sample of interest may occur automatically using system 1100. As shown in FIG. 12, sample cartridge 1104 and sample loading manifold assembly 1108 are positioned downstream of flow cell cartridge assembly 1102. In the implementation shown, sample loading manifold assembly 1108 is coupled between flow cell cartridge assembly 1102 and pump manifold assembly 1110. To draw a sample of interest from sample cartridge 1104 and toward pump manifold assembly 1110, sample valves 1170, pump valves 1174, and / or pumps 1172 may be selectively actuated to urge the sample of interest toward pump manifold assembly 1110. Sample cartridge 1104 may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves 1170. To individually flow the sample of interest toward channel 1130 of flow cell 1128 andaway from pump manifold assembly 1110, sample valves 1170, pump valves 1174, and / or pumps 1172 may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly 1102 and into respective channels 1130 of flow cell 1128.
[0123] Drive assembly 1112 interfaces with sipper manifold assembly 1106 and pump manifold assembly 1110 to flow one or more reagents that interact with the sample within flow cell 1128. In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system 1116 excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system 1100.
[0124] After image data is obtained, drive assembly 1112 interfaces with sipper manifold assembly 1106 and pump manifold assembly 1110 to flow another reaction component (e.g., a reagent) through flow cell 1128 that is thereafter received by waste reservoir 1118 via a primary waste fluidic line 1182 and / or otherwise exhausted by system 1100. Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.
[0125] The primary waste fluidic line 1182 is coupled between pump manifold assembly 1110 and waste reservoir 1118. In some implementations, pumps 1172 and / or pump valves 1174 of pump manifold assembly 1110 selectively flow the reaction components from flow cell cartridge assembly 1102, through fluidic line 1180 and sample loading manifold assembly 1108 to primary waste fluidic line 1182. Flow cell cartridge assembly 1102 is coupled to a central valve 1184 via flow cell interface 1126. Central valve 1184 is coupled with flow cell interface 1126 via a fluidic line 1185. An auxiliary waste fluidic line 1186 is coupled to central valve 1184 and to waste reservoir 1118. In some implementations, auxiliary waste fluidic line 1186 receives excess fluid of a sample of interest from flow cell cartridge assembly 1102, via central valve 1184, and flows the excess fluid of the sample of interest to waste reservoir 1118 when back loading the sample of interest into flow cell 1128.
[0126] Sipper manifold assembly 1106 includes a shared line valve 1188 and a bypass valve 1190. Shared line valve 1188 may be referred to as a reagent selector valve. Central valve 1184 and the valves 1188, 1190 of sipper manifold assembly 1106 may be selectively actuated to control the flow of fluid through fluidic lines 1192, 1194, 1196. Sipper manifold assembly 1106 may be coupled to a corresponding number of reagent reservoirs 1198 via reagent sippers 1200. Reagent reservoirs 1198 may contain fluid (e.g., reagent and / or another reaction component). In some implementations, sipper manifold assembly 1106 includes a plurality of ports. Each port of sipper manifold assembly 1106 may receive one of the reagent sippers 1200. Reagent sippers 1200 may be referred to as fluidic lines. Some forms of reagent sippers 1200 may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly 1106. Reagent reservoirs 1198 may be provided in a cartridge, and the tubes of reagent sippers 1200 may be configured to be inserted into corresponding reagent reservoirs 1198 in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir 1198 into the sipper manifold assembly 1106.
[0127] Shared line valve 1188 of sipper manifold assembly 1106 is coupled to central valve 1184 via shared reagent fluidic line 1192. Different reagents may flow through shared reagent fluidic line 1192 at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly 1110 may draw wash buffer through shared reagent fluidic line 1192, central valve 1184, and flow cell cartridge assembly 1102.
[0128] Bypass valve 1190 of sipper manifold assembly 1106 is coupled to central valve 1184 via dedicated reagent fluidic lines 1194, 196. Each of the dedicated reagent fluidic lines 1194, 1196 may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines 1194, 1196 may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.
[0129] Bypass valve 1190 is also coupled to cache 1176 of pump manifold assembly 1110 via bypass fluidic line 1178. One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line 1178. The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly 1102. Thus, the operations using bypass fluidic line 1178 may occur during, for example, incubation of one ormore samples of interest within flow cell cartridge assembly 1102. That is, shared line valve 1188 may be utilized independently of bypass valve 1190 such that bypass valve 1190 may utilize bypass fluidic line 1178 and / or cache 1176 to perform one or more operations while shared line valve 1188 and / or central valve 1184 simultaneously, substantially simultaneously, or offset synchronously perform other operations.
[0130] Drive assembly 1112 includes a pump drive assembly 1202 and a valve drive assembly 1204. Pump drive assembly 1202 may be adapted to interface with one or more pumps 1172 to pump fluid through flow cell 1128 and / or to load one or more samples of interest into flow cell 1128. Valve drive assembly 1204 may be adapted to interface with one or more of the valves 1170, 1174, 1184, 1188, 1190 to control the position of the corresponding valves 1170, 1174, 1184, 1188, 1190.
[0131] Examples of Combinations
[0132] The following examples relate to various non-exhaustive ways in which the teachings herein may be combined or applied. The following examples are not intended to restrict the coverage of any claims that may be presented at any time in this application or in subsequent filings of this application. No disclaimer is intended. The following examples are being provided for nothing more than merely illustrative purposes. It is contemplated that the various teachings herein may be arranged and applied in numerous other ways. It is also contemplated that some variations may omit certain features referred to in the below examples. Therefore, none of the aspects or features referred to below should be deemed critical unless otherwise explicitly indicated as such at a later date by the inventors or by a successor in interest to the inventors. If any claims are presented in this application or in subsequent filings related to this application that include additional features beyond those referred to below, those additional features shall not be presumed to have been added for any reason relating to patentability.
[0133] Example 1 : A computer-implemented method of assessing the presence of supplementary nucleotide tags (SNTs) within single-cell data from background noise, the method comprising: receiving, from a nucleotide sequencing device, single-cell sequencing data for a plurality of cells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features; determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell bydividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell; determining a threshold for each class of cell-associated feature by applying k-means clustering to the respective feature purity fraction for each class of cell-associated feature, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of cell-associated feature; and generating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
[0134] Example 2. The computer-implemented method of Example 1, wherein supplementary nucleotide tags comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
[0135] Example 3. The computer-implemented method of Examples 1 or 2, wherein determining the guide purity fraction for each class of gRNA molecules and for each respective single-cell comprises dividing read counts for each respective gRNA by total gRNA counts for the respective single-cell; and determining a threshold for each class of gRNA by applying k-means clustering to the respective guide purity fraction for each class of gRNA, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of gRNA.
[0136] Example 4. The computer-implemented method of Examples 1, 2, or 3, wherein the read counts are deduplicated.
[0137] Example 5. The computer-implemented method of Examples 1, 2, 3, or 4, wherein the read counts contain duplicate counts due to amplification during sample processing.
[0138] Example 6. The computer-implemented method of Examples 1, 2, 3, 4, or 5, wherein the reads counts are corrected or de-noised,
[0139] Example 7. The computer-implemented method of Examples 1, 2, 3, 4, 5, or 6, wherein supplementary nucleotide tags (SNTs) comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
[0140] Example 8. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, or 7, wherein a number of reads per cell is determined as usable reads divided by called cells, wherein called cells correspond to the number of cells in an assay used to generate the single-cell sequencing data.
[0141] Example 9. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, or 8, wherein the number of cells in the assay is determined based on a paired gene expression sample or a gene-expression-free assay.
[0142] Example 10. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, or 9, wherein usable reads of the plurality of cells are determined based at least in part on a white-listed barcode.
[0143] Example 11. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7.8, 9, or 10, wherein usable reads of the plurality of cells are determined based at least in part on matching to a feature barcode in a given read.
[0144] Example 12. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, wherein read depths for some or all of the single-cell sequencing data is less than 3,000 reads per cell at low signal to noise ratios.
[0145] Example 13. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12, wherein read depths for some or all of the single-cell sequencing data is approximately 250 reads per cell or greater.
[0146] Example 14. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, wherein read depths for some or all of the single-cell sequencing data is between approximately 500 to approximately 2000 reads per cell.
[0147] Example 15. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, wherein single-cells are determined to be at or above the respective thresholds of one or more classes of cell-associated features and are determined to be positive for the one or more respective cell-associated features for which they are above the respective thresholds.
[0148] Example 16. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, wherein k-means clustering is applied using k = 2 or 3.
[0149] Example 17. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16, wherein k-means clustering is applied with an additional parameter corresponding to normalized counts of cell-associated feature class or total cell-associated feature counts.
[0150] Example 18. The computer-implemented method of Examples 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17, wherein determining the feature purity fraction further comprises excluding single-cells having cell-associated feature signal below an exclusion threshold.
[0151] Example 19. One or more computer-readable media, the one or more computer-readable media encoding processor-executable routines which, when executed by a processor, cause acts to be performed comprising: receiving, from a nucleotide sequencing device, singlecell sequencing data for a plurality of cells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features; determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell; determining a threshold for each class of cell-associated feature by applying k-means clustering to the respective feature purity fraction for each class of cell-associated feature, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of cell-associated feature; and generating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
[0012] Example 20. The one or more computer-readable media of Example 19, wherein cell-associated features comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
[0153] Example 21. The one or more computer-readable media of Examples 19 or 20, wherein cell-associated features comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
[0154] Example 22. A nucleic acid sequence processing system, comprising: one or more processors; and one or more memory structures in communication with the one or more processors and encoding processor-executable routines, wherein the processor-executable routines, when executed by the one or more processors, cause act to be performed comprising: receiving, from a nucleotide sequencing device, single-cell sequencing data for a plurality ofcells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features; determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell; determining a threshold for each class of cell-associated feature by applying k-means clustering to the respective feature purity fraction for each class of cell-associated feature, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of cell-associated feature; and generating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
[0155] Example 23. The nucleic acid sequence processing system of Example 22, wherein cell-associated features comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
[0156] Example 24. The nucleic acid sequence processing system of Examples 22 or 23, wherein cell-associated features comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
[0157] Miscellaneous
[0158] Any one of the above-described strategies and methods, or combinations thereof may be used in the conjunction particle-templated emulsions. For example, methods may be used for single cell expression profiling, which may include combining target cells with a plurality of template particles in a first fluid to provide a mixture in a reaction tube. The mixture may be incubated to allow association of the plurality of the template particles with target cells. A portion of the plurality of template particles may become associated with the target cells. The mixture is then combined with a second fluid which is immiscible with the first fluid. The fluidand the mixture are then sheared so that a plurality of monodisperse droplets is generated within the reaction tube. The monodisperse droplets generated comprise (i) at least a portion of the mixture, (ii) a single template particle, and (iii) a single target particle. Of note, in practicing methods of the techniques provided by this disclosure a substantial number of the monodisperse droplets generated will comprise a single template particle and a single target particle, however, in some instances, a portion of the monodisperse droplets may comprise none or more than one template particle or target cell.
[0159] In some aspects, generating the template particles-based monodisperse droplets involves shearing two liquid phases. The mixture is the aqueous phase and, in some embodiments, comprises reagents selected from, for example, buffers, salts, lytic enzymes (e.g. proteinase k) and / or other lytic reagents (e.g., Triton X-100, Tween-20, IGEPAL, bm 135, or combinations thereof), nucleic acid synthesis reagents e.g. nucleic acid amplification reagents or reverse transcription mix, or combinations thereof. The fluid is the continuous phase and may be an immiscible oil such as fluorocarbon oil, a silicone oil, or a hydrocarbon oil, or a combination thereof. In some embodiments, the fluid may comprise reagents such as surfactants (e.g. octylphenol ethoxylate and / or octylphenoxypolyethoxyethanol), reducing agents (e.g. DTT, beta mercaptoethanol, or combinations thereof).
[0160] Some methods of the disclosure use oligos. Oligos, sometimes referred to as oligonucleotides, are sequences of contiguous nucleotides of DNA, RNA, or a mixture thereof. In certain embodiments oligos comprise DNA. However, in certain embodiments, oligos may comprise RNA. In other embodiments, oligos may comprise a mixture of DNA and RNA. Oligos may comprise noncanonical nucleotides, such as, synthetic nucleotides that have been modified to incorporate certain biomolecular properties. The length of the oligo is usually denoted by mer". For example, an oligo of six nucleotides is a hexamer, or 6-mer, while one of 25 nucleotides may be referred to as a 25-mer. An oligo may include other features such one or more conformationally-restricted nucleic acid or a locked nucleic acid (LNA) bases or phosphorothioate inter-base linkages, to improve binding stability or residence times.Incorporation by reference
[0161] References and citations to other documents, such as patents, patent applications, patent publications, journals, books, papers, web contents, have been made throughout this disclosure. All such documents are hereby incorporated herein by reference in their entirety for all purposes.Equivalents
[0162] Various modifications of the invention and many further embodiments thereof, in addition to those shown and described herein, will become apparent to those skilled in the art from the full contents of this document, including references to the scientific and patent literature cited herein. The subject matter herein contains important information, exemplification and guidance that can be adapted to the practice of this invention in its various embodiments and equivalents thereof.
Claims
What is claimed is:
1. A computer-implemented method of assessing the presence of supplementary nucleotide tags (SNTs) within single-cell data from background noise, the method comprising:receiving, from a nucleotide sequencing device, single-cell sequencing data for a plurality of cells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features;determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell;determining a threshold for each class of cell-associated feature by applying k-means clustering to the respective feature purity fraction for each class of cell-associated feature, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of cell-associated feature; andgenerating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
2. The computer-implemented method of claim 1, wherein supplementary nucleotide tags comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
3. The computer-implemented method of claim 2, wherein determining the guide purity fraction for each class of gRNA molecules and for each respective single-cell comprises dividing read counts for each respective gRNA by total gRNA counts for the respective single-cell; and determining a threshold for each class of gRNA by applying k-means clustering to the respective guide purity fraction for each class of gRNA, wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of gRNA.
4. The computer-implemented method of claim 1, wherein the read counts are deduplicated.
5. The computer-implemented method of claim 1, wherein the read counts contain duplicate counts due to amplification during sample processing.
6. The computer-implemented method of claim 1, wherein the reads counts are corrected or de-noised.
7. The computer-implemented method of claim 1, wherein supplementary nucleotide tags (SNTs) comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
8. The computer-implemented method of claim 1, wherein a number of reads per cell is determined as usable reads divided by called cells, wherein called cells correspond to the number of cells in an assay used to generate the single-cell sequencing data.
9. The computer-implemented method of claim 8, wherein the number of cells in the assay is determined based on a paired gene expression sample or a gene-expression-free assay.
10. The computer-implemented method of claim 8, wherein usable reads of the plurality of cells are determined based at least in part on a white-listed barcode.
11. The computer-implemented method of claim 8, wherein usable reads of the plurality of cells are determined based at least in part on matching to a feature barcode in a given read.
12. The computer-implemented method of claim 1, wherein read depths for some or all of the single-cell sequencing data is less than 3,000 reads per cell at low signal to noise ratios.
13. The computer-implemented method of claim 1, wherein read depths for some or all of the single-cell sequencing data is approximately 250 reads per cell or greater.
14. The computer-implemented method of claim 1, wherein read depths for some or all of the single-cell sequencing data is between approximately 500 to approximately 2000 reads per cell.
15. The computer-implemented method of claim 1, wherein single-cells are determined to be at or above the respective thresholds of one or more classes of cell-associated features and are determined to be positive for the one or more respective cell-associated features for which they are above the respective thresholds.
16. The computer-implemented method of claim 1, wherein k-means clustering is applied using k = 2 or 3.
17. The computer-implemented method of claim 1, wherein k-means clustering is applied with an additional parameter corresponding to normalized counts of cell-associated feature class or total cell-associated feature counts.
18. The computer-implemented method of claim 1, wherein determining the feature purity fraction further comprises excluding single-cells having cell-associated feature signal below an exclusion threshold.
19. One or more computer-readable media, the one or more computer-readable media encoding processor-executable routines which, when executed by a processor, cause acts to be performed comprising:receiving, from a nucleotide sequencing device, single-cell sequencing data for a plurality of cells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features;determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell;determining a threshold for each class of cell-associated feature by applying k-means clustering to the respective feature purity fraction for each class of cell-associated feature,wherein respective single-cells determined to be at or above the threshold are determined to be positive for the respective class of cell-associated feature; andgenerating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
20. The one or more computer-readable media of claim 19, wherein cell-associated features comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
21. The one or more computer-readable media of claim 19, wherein cell-associated features comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody-derived tags, cellfeature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.
22. A nucleic acid sequence processing system, comprising:one or more processors; andone or more memory structures in communication with the one or more processors and encoding processor-executable routines, wherein the processor-executable routines, when executed by the one or more processors, cause act to be performed comprising:receiving, from a nucleotide sequencing device, single-cell sequencing data for a plurality of cells, a subset of which exhibit one or more cell-associated features, wherein each cell of the subset comprises one or more classes of cell-associated features;determining a feature purity fraction for each class of cell-associated feature and for each respective single-cell by dividing read counts for each respective cell-associated feature by total cell-associated feature counts for the respective single-cell;determining a threshold for each class of cell-associated feature by applying k- means clustering to the respective feature purity fraction for each class of cell-associated feature, wherein respective single-cells determined to be ator above the threshold are determined to be positive for the respective class of cell-associated feature; andgenerating an output comprising a set of one or more feature assignments for each cell in the form of a matrix, table, or plot.
23. The nucleic acid sequence processing system of claim 22, wherein cell-associated features comprise CRISPR gRNA and the plurality of cells comprise CRISPR edited cells, wherein the classes of cell-associated features comprise classes of gRNA molecules, wherein the feature purity fraction comprises a guide purity fraction for each respective class of gRNA molecules.
24. The nucleic acid sequence processing system of claim 22, wherein cell-associated features comprise one or more of gRNA CROP-Seq, spatial tags, CITE-Seq, antibody -derived tags, cell-feature specific aptamers, or hashing oligonucleotides delivered by one or more of antibody, lipid association, or direct cross-linking to cell surfaces.