Methods and compositions to enrich target RNAS for perturb-seq
The method enhances Perturb-Seq by converting RNAs to cDNA, labeling, and circularizing for efficient guide RNA detection, addressing high costs and recombination issues, achieving cost-effective and versatile RNA library construction.
Patent Information
- Application Number
- PCT/US2025/029546
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-15
- Filing Date
- 2025-05-15
- Publication Date
- 2025-11-20
AI Technical Summary
Existing methods for Perturb-Seq experiments face challenges such as high costs, limitations in delivering multiple sgRNAs, lentiviral recombination leading to unintended sgRNA combinations, and difficulties in constructing libraries for multiplexing, which complicate accurate guide RNA identification and control of vector particles delivery.
A method involving converting RNAs of interest into cDNA, labeling with a 5'-phosphate, circularizing using a ligase, and amplifying with specific primers to enrich guide RNAs and mRNAs, utilizing a split-pool barcoding approach and sequencing to create a single-cell RNA library.
This method significantly reduces costs by over 20 times, enables multiplexing without limits, improves guide RNA detection efficiency by 7 times, and ensures stable capture across cell types, allowing for versatile application in various Perturb-Seq experiments.
Smart Images

Figure US2025029546_20112025_PF_FP_ABST
Abstract
Description
PATENT Attorney Docket No.079445-014910PC-1492697 Client Ref. No. S24-115 METHODS AND COMPOSITIONS TO ENRICH TARGET RNAS FOR PERTURB-SEQ CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No.63 / 647,822, filed May 15, 2024, the disclosure of which is herein incorporated by reference in its entirety for all purposes. STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT
[0002] This invention was made with Government support under contracts HG011972 and HL164811 awarded by the National Institutes of Health. The Government has certain rights in the invention. BACKGROUND
[0003] Cellular systems possess complex regulatory mechanisms that modulate signals, make decisions, and synchronize responses in diverse conditions. Diseases arise from the malfunction of these systems, manifested when components are absent, faulty, or overly active. To decipher disease mechanisms and improve treatments, it is essential to gain a thorough understanding of cellular components, recognize the cellular pathways they engage in, and comprehend the collaborative functioning of these components and pathways in generating cellular responses. Recent technological progress has empowered researchers to transition from mere observational studies to extensive perturbation of individual components, aiming to establish the causal impact of each component on the entire cellular system.
[0004] Perturb-Seq involves perturbing numerous genes in a pooled screen using CRISPR and evaluating outcomes through single-cell RNA-seq (Adamson et al.2016; Dixit et al.2016). This method provides an unbiased and comprehensive insight into cellular programs as reflected in gene expression. It enables the simultaneous examination of gene perturbations on multiple cellular pathways or phenotypes. However, previous methods used in Perturb-Seq experiments were expensive, with recent genome-wide screens (Replogle et al. 2022) costing an estimated ~$1.6M for 10x Genomics scRNA-seq kits and Illumina sequencing.
[0005] The key challenge in extending combinatorial indexing from single-cell RNA-Seq to single-cell CRISPR screens lies in assigning perturbation identities to single-cell phenotypes. Existing techniques depend on the use of the CROP-Seq vector (Datlinger et al.2017), which replicates the sgRNA sequence during lentiviral transduction. This duplication results in two expression units within the same construct: one for functional sgRNA expression transcribed by RNA Pol III and another for generating a chimeric transcript transcribed by RNA Pol II withthe sgRNA sequence at its 3 end. By employing this approach, CROP-Seq guarantees a preciseassignment between sgRNAs and polyadenylated single-cell barcodes. Nonetheless, the CROP-Seq vector has limitations: it cannot deliver multiple sgRNAs in tandem, and its chimeric Pol II transcript expression may be very low in some cell types. Recent studies have made efforts to address these challenges while also keeping costs low. Nevertheless, each approach has its own set of drawbacks. For instance, Xu et al. (Xu et al.2023) and Jiang et al. (Jiang et al. 2024) made improvements to the split-pool protocol to enhance the detection of unique molecular identifiers (UMIs) for each cell, yet their method still depends on the CROP- Seq vector, which is designed for capturing Pol II transcripts. Meanwhile, the CROPseq-multi method (Walton et al. 2024) has been adapted from the original CROP-Seq to enable the delivery of two sgRNAs concurrently. However, expanding beyond two sgRNAs presents significant challenges.
[0006] In addition, some methods do not sequence guide RNAs directly to identify them but instead determine the identity through the sequencing of barcodes adjacent to the guide RNA on the delivery vector. A significant challenge with this strategy is the gap between the sgRNA spacer sequences and the barcodes, which can result in lentiviral recombination. This recombination causes about 30% of cells to exhibit unintended sgRNA combinations, complicating accurate guide RNA identification within those cells.
[0007] Furthermore, constructing libraries for multiplexing that include three separate sequence components (such as two sgRNAs and one barcode) presents considerable difficulties. To overcome the issue of lentiviral recombination, a recent study (Chardon et al. 2023) has introduced an alternative method using a piggyBac transposon-based gRNA expression vector. Despite this innovation, controlling the number of vector particles delivered to each cell (known as the multiplicity of infection or MOI) remains a challenge in these systems.SUMMARY
[0008] The terms “invention,” “the invention,” “this invention,” and “the present invention,” as used in this document, are intended to refer broadly to all of the subject matter of this patent application and the claims below. Statements containing these terms should be understood not to limit the subject matter described herein or to limit the meaning or scope of the patent claims below. Covered embodiments of the invention are defined by the claims, not this summary. This summary is a high-level overview of various aspects of the invention and introduces some of the concepts that are described and illustrated in the present document and the accompanying figures. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all figures, and each claim. Some of the exemplary embodiments of the present invention are discussed below.
[0009] In general, provided herein a method of detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell library. In some embodiments, the method comprises: i) converting each RNA of interest into a cDNA through reverse transcription; ii) labeling each cDNA with a 5’-phosphate; iii) circularizing each 5’-phosphate labeled cDNA using a ligase; and iv) amplifying each circularized cDNA using at least two primers specific to the RNAs / cDNAs of interest, thereby detecting or enriching one or more RNA of interest in the cell.
[0010] In some embodiments, the single-cell library is a single-cell RNA library. In some embodiments, the 5’ phosphate is added to the cDNA using a 5’-phosphorylated probe. In some embodiments, the 5’ phosphate is added through use of a kinase. In some embodiments, the 5’ phosphate is added through PCR using a 5’-phosphorylated primer.
[0011] In some embodiments, the RNAs of interest are one or more guide RNAs (gRNAs). In some embodiments, the two primers comprise a P5-gRNA primer and a P7-gRNA primer. In some embodiments, the P5-gRNA primer comprises a sequence having at least 80% identityto SEQ ID NO: 15. In some embodiments, the P7-gRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 16.
[0012] In some embodiments, the RNAs of interests are one or more mRNAs. In some embodiments, the two primers comprise a P5-mRNA primer and a P7-mRNA primer. In some embodiments, the P5-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 1. In some embodiments, the P7-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 2. In some embodiments, the ligase is a ssDNA ligase.
[0013] In some embodiments, step i) comprises converting all the RNAs in the cell into cDNAs through reverse transcription. In some embodiments, prior step ii), the method further comprises barcoding each cDNA. In some embodiments, each cDNA is barcoded through a “split-pool” approach. In some embodiments, the “split-pool” approach comprises multiple “split and pool” cycles. In some embodiments, each “split and pool” cycle comprises a “split” step that the plurality of cells is distributed individually to attach a barcode followed by a “pool” step that the cells are mixed together. In some embodiments, the “split-pool” approach comprises at least two “split and pool” cycles. In some embodiments, the “split-pool” approach comprises three “split and pool” cycles. In some embodiments, the first cycle comprises adding a Barcode 1 (BC1) through reverse transcription. In some embodiments, the second cycle comprises adding a Barcode 2 (BC2) at the end of the BC1. In some embodiments, the third cycle comprises adding a Barcode 3 (BC3) at the end of the BC2.
[0014] In some embodiments, the BC1 comprises a sequence having at least 80% identity to a sequence selected from Table 10. In some embodiments, the BC2 comprises a sequence having at least 80% identity to a sequence selected from Table 11. In some embodiments, the BC3 comprises a sequence having at least 80% identity to a sequence selected from Table 12. In some embodiments, the BC1 and the BC2 are linked via a linker comprising a sequence of SEQ ID NO: 6. In some embodiments, the BC2 and the BC3 are linked via a linker comprising a sequence of SEQ ID NO: 7. In some embodiments, each cDNA is barcoded through reverse transcription with a barcoded primer enclosed in a physical space, such as a droplet or a well.
[0015] In some embodiments, prior step iv), the method further comprises linearizing the circularized cDNAs using a restriction enzyme. In some embodiments, the restriction enzyme is a DraI enzyme.
[0016] In some embodiments, after step iv), the method further comprises v) sequencing the enriched RNAs of interest using one or more sequencing primers. In some embodiments, theone or more sequencing primers are Illumina sequencing primers. In some embodiments, the one or more sequencing primers are adapters for Ultima sequencing.
[0017] The present disclosure further provides a kit for detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell RNA library via circular capture, comprising a) a ligase; and b) a 5’ primer and a 3’ primer to amplify one or more RNAs of interest. In some embodiments, the kit further comprises c) one or more barcode oligonucleotides. In some embodiments, the one or more barcode oligonucleotides comprise a sequence having at least 80% identity to a sequence from Tables 4-9. In some embodiments, the ligase is a ssDNA ligase.
[0018] In some embodiments, the one or more RNAs of interest are gRNAs. In some embodiments, the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 15 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 16.
[0019] In some embodiments, the one or more RNAs of interest are mRNAs. In some embodiments, the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 1 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 2.
[0020] In some embodiments, the one or more RNAs of interest are gRNAs. In some embodiments, the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 987 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 988.
[0021] These and other aspects of the disclosure are described in detail below. For example, other aspects are directed to systems, devices, and computer readable media associated with methods described herein.
[0022] Reference to the remaining portions of the specification, including the drawings and claims, will realize other features and advantages of the present disclosure. Further features and advantages of the present disclosure, as well as the structure and operation of various aspects of the present disclosure, are described in detail below with respect to the accompanying drawings. In the drawings, like reference numbers can indicate identical or functionally similar elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Fig.1 presents an overview of circularization-based split-pool Perturb-Seq.
[0024] Fig.2 presents boxplot of number of sgRNA UMIs per cell in hiPSCs for CROP-Seq vs Circular Capture (our method). Human induced pluripotent stem cells (hiPSCs) engineered with CRISPRi systems were transduced with a lentiviral vector carrying an sgRNA targeting the CDH5 gene, at an MOI of 1. We concurrently employed CROP-Seq to isolate Pol II gRNA transcripts and our protocol for Circular Capture to isolate Pol III gRNA transcripts. These transcripts from ~500 cells were then processed into sequencing libraries and sequenced to saturation. Afterward, we quantified the unique molecular identifiers (UMIs) in both libraries." The figure shows that we were able to capture >6x more guide UMIs using Circular Capture compared to CROP-Seq. Higher UMI capture allows for higher signal-to-noise determination of guide assignment in a Perturb-Seq experiment.
[0025] Fig. 3A presents violin plot of mRNA UMIs per cell at 200K read depth in hiPSCs using Circularization protocol. From 500 cells in the experiment in Figure 2, we also extracted the mRNA and converted it into sequencing libraries, which were then sequenced to saturation. The figure illustrates the distribution of total transcriptome UMIs across the sampled cells (each dot represents a single cell). This highlights the high complexity of the transcriptome we captured.
[0026] Fig. 3B presents plot of the median number of UMIs detected per cell as a function of sequencing depth. This plot shows the median number of UMIs per cell as a function sequencing depth, utilizing the same dataset as in Fig.3B. Here, we depict the number of UMIs in relation to sequencing depth, illustrating the complexity of our mRNA library. At 50K reads, the library yields approximately 12.5K UMIs, which is on par with or surpasses existing technologies such as 10x Genomics.
[0027] Fig. 4 presents knockdown efficiency in a Perturb-Seq experiment with 14 genes at MOI of 3. In a Perturb-Seq experiment targeting 14 genes at an MOI of 3, hiPSCs engineered with CRISPRi systems were transduced with lentiviral vectors. These vectors carried a pool of 100 sgRNAs aimed at genes such as ROCK1, SMARCE1, GSK3B, NANOG, SMAD5, ACVR2A, PLCG1, CDKN1A, KDR, TP53, SMAD1, LEMD3, SOX2, and PRKD1, with others serving as negative control sgRNAs. We conducted Circular-Capture Perturb-Seq to create single-cell gRNA and mRNA libraries. The gRNA library facilitated the identification of specific guides in each cell, and the mRNA library enabled the evaluation of knockdown efficiency on the target genes and the examination of their effects on the transcriptome. We employed the SCEPTRE package (https: / / katsevich-lab.github.io / sceptre / ) for both gRNAassignment and knockdown efficiency calculation. Knockdown efficiency is influenced not only by the precision of sgRNA targeting but also by our capability to accurately assign the correct sgRNA identities to the corresponding cells. Our observations show that knockdown efficiency generally exceeds 50%, except when targeting genes essential to hiPSCs, such as NANOG and SOX2. In these instances, cells containing highly efficient sgRNAs targeting these crucial genes were lost due to cell death, and thus went undetected. The surviving cells typically carried less effective sgRNAs, resulting in lower observed knockdown efficiencies for NANOG and SOX2.
[0028] Fig. 5 presents a schematic diagram of split-pool approach with four barcode combination to index both mRNA and gRNA in single cells.
[0029] Fig. 6 presents a molecular process for preparing mRNA and gRNA libraries in Circular-Capture Perturb-Seq.
[0030] Fig. 7 presents an exemplary diagram of full-length cDNA on Bioanalyzer, with a peak between 300-4000 bp.
[0031] Fig.8 presents a comparison between CC-Perturb-Seq (Circular Capture) and CROP- seq, focusing on their performance in TeloHAEC cells. Fig. 8A shows violin plots of the number of gene expression UMIs per cell. Fig. 8B shows violin plots of the number of guide RNA expression UMIs per cell. Fig.8C shows quantile-quantile (QQ) plot.
[0032] Fig.9 presents histograms showing the distribution of sgRNA UMI counts per cell in hiPSCs (Fig.9A) and endothelial cells (ECs) (Fig.9B) with mean at 193 UMIs and 190 UMIs, respectively, demonstrating efficient guide capture with CC-Perturb-Seq.
[0033] Fig. 10 presents accurate guide assignment with SCEPTRE despite high MOI. Fig. 10A shows panels with clear separation in gRNA counts between cells with (pert) and without (unpert) the guide across multiple examples. Even under high multiplicity of infection (MOI = 5.81), where individual cells receive multiple sgRNAs (as shown in Fig. 10B), SCEPTRE maintains clean classification.
[0034] Fig. 11 displays the distribution of gene knockdown efficiencies in human induced pluripotent stem cells (hiPSCs) (Fig. 11A) and their differentiated endothelial cell (EC) derivatives (Fig.11B) using the CC-Perturb-Seq platform.DETAILED DESCRIPTION I. General
[0035] The present disclosure generally provides innovative methods and compositions for enriching one or more RNAs of interest, such as CRISPR guide RNAs or specific mRNAs of interest), in a cell for preparing a single-cell library such a single-cell RNA library. Such RNA enrichment reduces the complexity of a single-cell RNA / cDNA library to facilitate further processing and genetic analysis. For example, RNA enrichment may provide a means for obtaining size selected sequencing library molecules that include barcode sequences and the target RNAs. RNA enrichment may also provide for a cost-effective implementation and analysis of large-scale sequencing reads, such as in a Perturb-seq.
[0036] In one aspect, the present disclosure provides a method of detecting or enriching one or more RNAs of interest in a single-cell library. In some embodiments, such RNA detection / enrichment is for an optimized Perturb-Seq. In some embodiments, such RNA detection is from RNA, ATAC-seq, CITE-seq, or probe hybridization-based single-cell libraries. In some embodiments, such RNA detection / enrichment can significantly reduce costs of a Perturb-Seq by over 20 times. First, inventors replaced the traditional 10x scRNA-seq method with a “split-pool” combinatorial barcoding (Split-seq) method (Rosenberg et al. 2018). With such amendment, the methods have no limit on the number of sgRNAs that can be multiplexed together. Second, inventors developed a more versatile and scalable tool for single-cell Perturb-Seq that involves capturing the sgRNA using a novel approach. This approach directly captures and reads-out the functional sgRNA transcribed by RNA Pol III to (1) ensure stable capture efficiency across cell types. Since our approach doesn’t rely on cell- type dependent Pol II promoter, any cell type can be applied, including smooth muscle cells (SMCs), endothelial cells, hepatocytes, etc.; and (2) enable multiplexed delivery of distinct sgRNAs targeting the same gene within each lentiviral particle, thereby maximizing CRISPRi efficacy. By using this new method of gRNA detection and assignment, inventors were able to improve the efficiency of guide RNA detection by about 7 times through an innovative circularized capture approach.
[0037] As disclosed herein, this novel RNA detection and enrichment technology is versatile, making it applicable across various Perturb-Seq experiment variations. For example, this technology can be applied at both low (e.g. MOI of 1) and high Multiplicity of Infection (MOIs) (e.g. MOI of 5, 10, 20). This technology can capture both gRNAs and other target RNAs ofinterest, such as RNA barcodes, miRNAs, small RNAs, mRNAs of interests, and any endogenously or exogenously expressed RNAs. This technology, as opposed to 10x Genomics’ Chromium Single Cell 3’ Gene Expression and CRISPR Screening, can be also applied to any sgRNA cloning vector (e.g. CROP-seq vector or sgOpti vector) with or without an engineered Capture Sequence. This technology, as opposed to Parse Bioscience’s CRISPR Detect kit, can be applied to any sgRNA expression vector with or without expression of a Pol II “CROP-seq” transcript.
[0038] In another aspect, the present disclosure provides a kit for detecting or enriching one or more RNAs of interest in a single-cell RNA library. In some embodiments, the kit comprises a list of the components (e.g. oligo sequences, enzymes) for a circular capture. In some embodiments, the kit comprises a list of the components (e.g. oligo sequences, enzymes) for a Perturb-seq. II. Definitions
[0039] It is to be understood that this disclosure is not strictly limited to particular embodiments described, as such may of course vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the claims.
[0040] It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. It should further be understood that as used herein, the term “a” entity or “an” entity refers to one or more of that entity. For example, a nucleic acid molecule refers to one or more nucleic acid molecules. As such, the terms “a”, “an”, “one or more” and “at least one” can be used interchangeably. Similarly, the terms “comprising”, “including” and “having” can be used interchangeably.
[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which thepublications are cited. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0042] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub- combination. All combinations of the embodiments are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0043] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements or use of a “negative” limitation.
[0044] As used herein, the term “about” means a range of values including the specified value, which a person of ordinary skill in the art would consider reasonably similar to the specified value. In some embodiments, about means within a standard deviation using measurements generally acceptable in the art. In embodiments, about means a range extending to + / - 10% of the specified value (e.g., + / - 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the specified value). In embodiments, about means the specified value.
[0045] The term “RNAs of interest” may also be interchangeably referred to as “cDNAs of interest”, “target RNAs”, or “target cDNAs”. An RNA of interest can refer to any RNA polynucleotides in a single-cell RNA library. RNAs of interest may be derived from the nucleus or cytoplasm of a cell, and may include nucleic acids in or from mitochondrial, organelles, vesicles, liposomes or particles present within the cell.
[0046] As used herein, a “barcode” refers to a short sequence of nucleotides (for example, DNA or RNA) that is used as an identifier for an associated molecule, such as a target molecule and / or target nucleic acid, or as an identifier of the source of an associated molecule, such as acell-of-origin. A barcode may also refer to any unique, non-naturally occurring, nucleic acid sequence that may be used to identify the originating source of a nucleic acid fragment.
[0047] As used herein, “unique molecular identifier” (UMI) refers to sequences of nucleotides present in DNA molecules that may be used to distinguish individual DNA molecules from one another. See, e.g., Kivioja, Nature Methods 9, 72-74 (2012). UMIs may be sequenced along with the DNA sequences with which they are associated to identify sequencing reads that are from the same source nucleic acid. The term “UMI” is used herein to refer to both the nucleotide sequence of the UMI and the physical nucleotides, as will be apparent from context. UMIs may be random, pseudo-random, or partially random, or nonrandom nucleotide sequences that are inserted into adapters or otherwise incorporated in source nucleic acid (e.g., DNA or RNA) molecules to be sequenced. In some embodiments, each UMI is expected to uniquely identify any given source molecule present in a sample. In some embodiments, a probe can further include a cleavage domain and / or a functional domain (e.g., a primer-binding site, such as for next generation sequencing (NGS).
[0048] The term “Perturb-seq” refers to a high-throughput method of performing single cell RNA sequencing (scRNA-seq) on pooled genetic perturbation screens. Perturb-seq is also known as CRISP-seq and CROP-seq. Perturb-seq combines multiplexed CRISPR mediated gene inactivations with single cell RNA sequencing to assess comprehensive gene expression phenotypes for each perturbation.
[0049] The term “MOI” stands for the multiplicity of infection, representing the ratio of the numbers of virus particles to the numbers of the host cells in a given infection medium. A value of MOI = 1 implies that on an average there is a single host cell for a single viral particle. In some embodiments. MOI may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more. In general, an MOI of 1 indicates a low viral infection ratio, and an MOI over 1, e.g.2, 5, 10, or 20, indicates a high viral infection ratio.
[0050] As used herein, “reverse transcription” is the synthesis of DNA from an RNA template. This process is driven by RNA-dependent DNA polymerases, also known as reverse transcriptases (RTs).
[0051] The terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” “nucleic acid fragment,” “oligonucleotide” and “polynucleotide,” as used herein, generally refer to a polynucleotide that may have various lengths of bases, comprising, for example, deoxyribonucleotide, deoxyribonucleic acid (DNA), ribonucleotide, or ribonucleic acid(RNA), or analogs thereof. A nucleic acid may be single-stranded. A nucleic acid may be double-stranded. A nucleic acid may be partially double-stranded, such as to have at least one double-stranded region and at least one single-stranded region. A partially double-stranded nucleic acid may have one or more overhanging regions. An “overhang,” as used herein, generally refers to a single-stranded portion of a nucleic acid that extends from or is contiguous with a double-stranded portion of a same nucleic acid molecule and where the single-stranded portion is at a 3’ or 5’ end of the same nucleic acid molecule. Non-limiting examples of nucleic acids include DNA, RNA, genomic DNA or synthetic DNA / RNA or coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, and isolated RNA of any sequence. A nucleic acid can comprise a sequence of four natural nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (or uracil (U) instead of thymine (T) when the nucleic acid is RNA). A nucleic acid may include one or more nonstandard nucleotide(s), nucleotide analog(s) and / or modified nucleotide(s). Polynucleotide sequences, when provided, are listed in the 5' to 3' direction, unless stated otherwise. Nucleic acid(s) can be derived from a completely chemical synthesis process, such as a solid phase- mediated chemical synthesis, from a biological source, such as through isolation from any species that produces nucleic acid, or from processes that involve the manipulation of nucleic acids by molecular biology tools, such as DNA replication, PCR amplification, reverse transcription, or from a combination of those processes.
[0052] The terms "hybridize" and "hybridization" refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form duplexes via Watson-Crick base pairing.
[0053] As used herein, the terms "complementary" or "complementarity" refers to polynucleotides that are able to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in an anti-parallel orientation between polynucleotide strands. Complementary polynucleotide strands can base pair in a Watson- Crick manner (e.g., A to T, A to U, C to G), or in any other manner that allows for the formation of duplexes. As persons skilled in the art are aware, when using RNA as opposed to DNA, uracil (U) rather than thymine (T) is the base that is considered to be complementary to adenosine. However, when a uracil is denoted in the context of the present disclosure, theability to substitute a thymine is implied, unless otherwise stated. "Complementarity" may exist between two RNA strands, two DNA strands, or between a RNA strand and a DNA strand. It is generally understood that two or more polynucleotides may be "complementary" and able to form a duplex despite having less than perfect or less than 100% complementarity. Two sequences are "perfectly complementary" or "100% complementary" if at least a contiguous portion of each polynucleotide sequence, comprising a region of complementarity, perfectly base pairs with the other polynucleotide without any mismatches or interruptions within such region. Two or more sequences are considered "perfectly complementary" or "100% complementary" even if either or both polynucleotides contain additional non-complementary sequences as long as the contiguous region of complementarity within each polynucleotide is able to perfectly hybridize with the other. "Less than perfect" complementarity refers to situations where less than all of the contiguous nucleotides within such region of complementarity are able to base pair with each other. Determining the percentage of complementarity between two polynucleotide sequences is a matter of ordinary skill in the art. For purposes of Cas9 targeting, a gRNA may comprise a sequence "complementary" to a target sequence (e.g., major or minor allele), capable of sufficient base-pairing to form a duplex (i.e., the gRNA hybridizes with the target sequence). Additionally, the gRNA may comprise a sequence complementary to a sequence adjacent to a PAM sequence, wherein the gRNA also hybridizes with the sequence adjacent to a PAM sequence in a target DNA.
[0054] The abbreviation “bp” refers to base pairs. In some instances, “bp” may be used to denote a length of a DNA fragment, even though the DNA fragment may be single stranded and does not include a base pair. In the context of single-stranded DNA, “bp” may be interpreted as providing the length in nucleotides.
[0055] The term “sequencing,” as used herein, generally refers to a process for generating or identifying a sequence of a biological molecule, such as a nucleic acid. The sequence may be a nucleic acid sequence which comprises a sequence of nucleic acid bases. As used herein, the term “template nucleic acid” generally refers to the nucleic acid to be sequenced. The template nucleic acid may be an analyte or be associated with an analyte. For example, the analyte can be a mRNA, and the template nucleic acid is the mRNA or a cDNA derived from the mRNA, or other derivative thereof. In another example, the analyte can be a protein, and the template nucleic acid is an oligonucleotide that is conjugated to an antibody that binds to the protein, or derivative thereof. Examples of sequencing include single molecule sequencing or sequencing by synthesis, for example. Sequencing may comprise generating sequencing signals and / orsequencing reads. Sequencing may be performed on template nucleic acids immobilized on a support, such as a flow cell, substrate, and / or one or more beads. In some cases, a template nucleic acid may be amplified to produce a colony of nucleic acid molecules attached to the support to produce amplified sequencing signals. In one example, (i) a template nucleic acid is subjected to a nucleic acid reaction, e.g., amplification, to produce a clonal population of the nucleic acid attached to a bead, the bead immobilized to a substrate, (ii) amplified sequencing signals from the immobilized bead are detected from the substrate surface during or following one or more nucleotide flows, and (iii) the sequencing signals are processed to generate sequencing reads. The substrate surface may immobilize multiple beads at distinct locations, each bead containing distinct colonies of nucleic acids, and upon detecting the substrate surface, multiple sequencing signals may be simultaneously or substantially simultaneously processed from the different immobilized beads at the distinct locations to generate multiple sequencing reads. In some sequencing methods, the nucleotide flows comprise non-terminated nucleotides. In some sequencing methods, the nucleotide flows comprise terminated nucleotides.
[0056] The terms “amplifying,” “amplification,” and “nucleic acid amplification” are used interchangeably and generally refer to generating one or more copies of a nucleic acid or a template. For example, “amplification” of DNA generally refers to generating one or more copies of a DNA molecule. Amplification of a nucleic acid may be linear, exponential, or a combination thereof. Amplification may be emulsion based or non-emulsion based. Non- limiting examples of nucleic acid amplification methods include reverse transcription, primer extension, polymerase chain reaction (PCR), ligase chain reaction (LCR), helicase-dependent amplification, asymmetric amplification, rolling circle amplification (RCA), recombinase polymerase reaction (RPA), loop mediated isothermal amplification (LAMP), nucleic acid sequence based amplification (NASBA), self-sustained sequence replication (3SR), and multiple displacement amplification (MDA). Where PCR is used, any form of PCR may be used, with non-limiting examples that include real-time PCR, allele-specific PCR, assembly PCR, asymmetric PCR, digital PCR, emulsion PCR (ePCR or emPCR), dial-out PCR, helicase- dependent PCR, nested PCR, hot start PCR, inverse PCR, methylation-specific PCR, miniprimer PCR, multiplex PCR, nested PCR, overlap-extension PCR, thermal asymmetric interlaced PCR, and touchdown PCR. Amplification can be conducted in a reaction mixture comprising various components (e.g., a primer(s), template, nucleotides, a polymerase, buffer components, co-factors, etc.) that participate or facilitate amplification. In some cases, thereaction mixture comprises a buffer that permits context independent incorporation of nucleotides. Non-limiting examples include magnesium-ion, manganese-ion and isocitrate buffers. Additional examples of such buffers are described in Tabor, S. et al. C.C. PNAS, 1989, 86, 4076-4080 and U.S. Pat. Nos. 5,409,811 and 5,674,716, each of which is herein incorporated by reference in its entirety. Useful methods for clonal amplification from single molecules include rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998), which is incorporated herein by reference), bridge PCR (Adams and Kron, Method for Performing Amplification of Nucleic Acid with Two Primers Bound to a Single Solid Support, Mosaic Technologies, Inc. (Winter Hill, Mass.); Whitehead Institute for Biomedical Research, Cambridge, Mass., (1997); Adessi et al., Nucl. Acids Res.28:E87 (2000); Pemov et al., Nucl. Acids Res. 33:e11(2005); or U.S. Pat. No.5,641,658, each of which is incorporated herein by reference), polony generation (Mitra et al., Proc. Natl. Acad. Sci. USA 100:5926-5931 (2003); Mitra et al., Anal. Biochem. 320:55-65(2003), each of which is incorporated herein by reference), and clonal amplification on beads using emulsions (Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), which is incorporated herein by reference) or ligation to bead-based adapter libraries (Brenner et al., Nat. Biotechnol.18:630-634 (2000); Brenner et al., Proc. Natl. Acad. Sci. USA 97:1665-1670 (2000)); Reinartz, et al., Brief Funct. Genomic Proteomic 1:95-104 (2002), each of which is incorporated herein by reference). Amplification products from a nucleic acid may be identical or substantially identical. A nucleic acid colony resulting from amplification may have identical or substantially identical sequences.
[0057] As used herein, the term “sequence read” refers to a string of nucleotides obtained from any part or all of a nucleic acid molecule. For example, a sequence read may be a short string of nucleotides (e.g., 20-150 nucleotides) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. A sequence read may be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes as may be used in microarrays, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Example sequencing techniques include massively parallel sequencing, targeted sequencing, Sanger sequencing, sequencing by ligation, ion semiconductor sequencing, and single molecule sequencing (e.g., using a nanopore, or single- molecule real-time sequencing (e.g., from Pacific Biosciences)). Such sequencing can be random sequencing or targeted sequencing (e.g., by using capture probes hybridizing tospecific regions or by amplifying certain region, both of which enrich such regions). Example PCR techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). As part of an analysis of a biological sample, a statistically significant number of sequence reads can be analyzed, e.g., at least 1,000 sequence reads can be analyzed. As other examples, at least 5,000, 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 sequence reads, or more, can be analyzed.
[0058] As used herein, the terms “identity,” “substantial identity,” “similarity,” “substantial similarity,” “homology” and the related terms and expressions used in the context of describing nucleic acid or amino acid sequences refer to a sequence that has at least 60% sequence identity to a reference sequence. Examples include at least: 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, sequence identity, as compared to a reference sequence using the programs for comparison of nucleic acid or amino acid sequences, such as BLAST using standard parameters. For sequence comparison, typically one sequence acts as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default (standard) program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters. A “comparison window” includes reference to a segment of any one of the number of contiguous positions (from 20 to 600, usually about 50 to about 200, more commonly about 100 to about 150), in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well-known. Optimal alignment of sequences for comparison may be conducted, for example, by the local homology algorithm of Smith and Waterman (Smith and Waterman “Identification of common molecular subsequences.” J Mol Biol. 147(1):195-7 (1981)) by the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch “A general method applicable to the search for similarities in the amino acid sequence of two proteins.” J Mol Biol.48(3):443- 53 (1970)), by the search for similarity method of Pearson and Lipman (Pearson and Lipman “Improved tools for biological sequence comparison.” Proc Natl Acad Sci USA.85(8):2444-8 (1988)), by computerized implementations of these algorithms (for example, BLAST), or by manual alignment and visual inspection.
[0059] Algorithms that are suitable for determining percent sequence identity and sequence similarity include BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. “Basic local alignment search tool.” J. Mol. Biol. 215:403-410 (1990), and Altschul et al., “Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.” Nucleic Acids Res. 25:3389-3402 (1997), respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (NCBI) web site. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold. These initial neighborhood word hits acts as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=1, N=-2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (Henikoff and Henikoff, “Amino acid substitution matrices from protein blocks.” Proc. Natl. Acad. Sci. USA 89:10915-10919 (1989)). The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (Karlin and Altschul “Applications and statistics for multiple high-scoring segments in molecular sequences.” Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.01, more preferably less than about 10-5, and most preferably less than about 10-20.
[0060] BLAST nucleotide searches may be performed with the BLASTN program (nucleotide query searched against nucleotide sequences) to obtain nucleotide sequences homologous to nucleic acid molecules of this disclosure, or with the BLASTX program (translated nucleotide query searched against protein sequences) to obtain protein sequences homologous to nucleic acid molecules of the invention. BLAST protein searches may be performed with the BLASTP program (protein query searched against protein sequences) to obtain amino acid sequences homologous to protein molecules of this disclosure, or with the TBLASTN program (protein query searched against translated nucleotide sequences) to obtain nucleotide sequences homologous to protein molecules of this disclosure. To obtain gapped alignments for comparison purposes, Gapped BLAST (in BLAST 2.0) may be utilized as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389. Alternatively, PSI-Blast may be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) supra. When utilizing BLAST, Gapped BLAST, and PSI-Blast programs, the default parameters of the respective programs (e.g., BLASTX and BLASTN) may be used.
[0061] As used herein, the terms “including,” “comprising,” “having,” “containing,” and variations thereof, are inclusive and open-ended and do not exclude additional, unrecited elements or method steps beyond those explicitly recited. As used herein, the phrase “consisting of” is closed and excludes any element, step, or ingredient not explicitly specified. As used herein, the phrase “consisting essentially of” limits the scope of the described feature to the specified materials or steps and those that do not materially affect the basic and novel characteristics of the disclosed feature.
[0062] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within embodiments of the present disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither, or both limits are included in the smaller ranges is also encompassed within the present disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the present disclosure.
[0063] Standard abbreviations may be used, e.g., bp, base pair(s); kb, kilobase(s); pi, picoliter(s); s or sec, second(s); min, minute(s); h or hr, hour(s); aa, amino acid(s); nt, nucleotide(s); and the like.
[0064] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments of the present disclosure, some potential and exemplary methods and materials may now be described. III. Methods for Enriching RNAs of Interest
[0065] In one aspect, the present disclosure provides a method of detecting or enriching one or more RNAs of interest in a cell. In some embodiments, the detection or enrichment of RNAs from a single cell is useful for preparation of a single-cell library, such as a single-cell RNA library. In some embodiments, the single-cell RNA library comprises a plurality of cDNAs that are reverse-transcribed from total RNAs in a single cell. Circularization-based RNA capture
[0066] In some embodiments, the method comprises a circularization-based RNA capture. In particular, the method comprises: i) converting the one or more RNAs of interest (optionally, along with other RNAs) into cDNAs via reverse transcription, ii) labeling each cDNA with a 5’-phosphate; iii) circularizing each 5’-phosphate labeled cDNA using a ligase; and iv) amplifying each circularized cDNA using two primers specific to the RNAs / cDNAs of interest, thereby detecting or enriching one or more RNA of interest in the cell. An exemplary protocol of the circularization-based guide RNA (gRNA) capture is illustrated in Example 6. An exemplary molecular process for preparing mRNA and gRNA libraries in Circular-Capture Perturb-Seq is shown in Fig.6. RNAs of interest
[0067] As disclosed herein, RNAs of interest can refer to any RNA polynucleotides in a single-cell RNA library. In some instances, the RNAs of interest are one or more guide RNAs (gRNAs). In other instances, the RNAs of interest are one or more message RNAs (mRNAs). In yet other instances, the RNAs of interest are a combination of gRNAs and mRNAs. In some embodiments, an RNA of interests can be endogenously or exogenously expressed RNA in thelibrary, such as a barcode, a miRNA, a small RNA or a mRNA of interest. In some embodiments, not only RNAs of interest but also other RNAs in the cell are converted to cDNAs through reverse transcription in step i) of the methods described herein. Primers
[0068] In some embodiments, the RNAs of interest comprise one or more guide RNAs (gRNAs). In such embodiments, the two primers comprise a 5’ gRNA primer and a 3’ gRNA primer to amplify the cDNA of interest.
[0069] In some embodiments, the 5’ gRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 17. In some embodiments, the 5’ gRNA primer further comprises a barcode (BC) sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. In some embodiments, the 5’ gRNA primer further comprises a P5 sequence of SEQ ID NO: 21, herein referred to as a P5-gRNA primer. In some embodiments, the 5’ gRNA primer is a P5-gRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 15 or 987. In a particular embodiment, the 5’ gRNA primer is a P5- gRNA primer comprising a sequence of SEQ ID NO: 15 or 987.
[0070] In some embodiments, the 3’ gRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 18. In some embodiments, the 3’ gRNA primer further comprises a BC sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. The BC sequence in the 3’ gRNA primer can be the same or different to the BC sequence in the 5’ gRNA primer. In some embodiments, the 3’ gRNA primer further comprises a P7 sequence of SEQ ID NO: 22, herein referred to as a P7-gRNA primer. In some embodiments, the 3’ gRNA primer is a P7-gRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or100% identity to SEQ ID NO: 16 or 988. In a particular embodiment, the 3’ gRNA primer is a P7-gRNA primer comprising a sequence of SEQ ID NO: 16 or 988.
[0071] In some embodiments, the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 15 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 16. In some embodiments, the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 987 and the 3’ primer comprises a sequence of having at least 80% identity to SEQ ID NO: 988.
[0072] In some embodiments, the RNAs of interest comprise one or more message RNAs (mRNAs). In such embodiments, the two primers comprise a 5’ mRNA primer and a 3’ mRNA primer to amplify the cDNA of interest.
[0073] In some embodiments, the 5’ mRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 19. In some embodiments, the 5’ mRNA primer further comprises a barcode (BC) sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. In some embodiments, the 5’ mRNA primer further comprises a P5 sequence of SEQ ID NO: 21, herein referred to as a P5-mRNA primer. In some embodiments, the 5’ mRNA primer is a P5-mRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 1. In a particular embodiment, the 5’ mRNA primer is a P5-mRNA primer comprising a sequence of SEQ ID NO: 1.
[0074] In some embodiments, the 3’ mRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 20. In some embodiments, the 3’ mRNA primer further comprises a BC sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. The BCsequence in the 3’ mRNA primer can be the same or different to the BC sequence in the 5’ mRNA primer. In some embodiments, the 3’ mRNA primer further comprises a P7 sequence of SEQ ID NO: 22, herein referred to as a P7-mRNA primer. In some embodiments, the 3’ mRNA primer is a P7-mRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 2. In a particular embodiment, the 3’ mRNA primer is a P7-mRNA primer comprising a sequence of SEQ ID NO: 2.
[0075] In some embodiments, the RNAs of interest comprise both gRNAs and mRNAs. In such embodiments, more than two primers are used to amplify both the gRNAs and mRNAs of interest. In such embodiments, four primers are used in the method. In some embodiments, the four primers comprise a 5’ gRNA primer, a 3’ gRNA primer, a 5’ mRNA primer, and a 3’ mRNA primer described above. Table 1 shows exemplary sequences for amplifying RNAs of interest from the single-cell library. Table 1. Exemplary sequences for amplifying gRNA and mRNA from the library.N6 refers to 6 nucleotides, N8 refers to 8 nucleotides, and N15 refers to 15 nucleotides, wherein N is any nucleotide. Ligases
[0076] DNA ligase is a type of enzyme that facilitates the joining of DNA strands together by catalyzing the formation of a phosphodiester bond. Some DNA ligases can repair single- strand breaks in duplex DNA, and some other ligases forms (such as DNA ligase IV) can specifically repair double-strand breaks (i.e. a break in both complementary strands of DNA). Single-strand breaks are repaired by DNA ligase using the complementary strand of the double helix as a template, with DNA ligase creating the final phosphodiester bond to fully repair the DNA. As disclosed herein, the ligase used in the method can circularize each 5’-phosphate labeled cDNA in the library. In some embodiments, the ligase is a ssDNA ligase, such as CircLigase™ Ligase. CircLigase™ ssDNA Ligase is a thermostable ATP-dependent ligase that catalyses intramolecular ligation (i.e. circularization) of ssDNA templates having a 5´- phosphate and a 3´-hydroxyl group. In contrast to T4 DNA Ligase and Ampligase™ DNA Ligase, which ligate DNA ends that are annealed adjacent to each other on a complementary DNA sequence, CircLigase ssDNA Ligase ligates ends of ssDNA in the absence of a complementary sequence.Linearization
[0077] In some embodiments, after the cDNAs are circularized (step iii), the circularized cDNAs are linearized to facilitate the subsequent amplification step (iv). In some embodiments, the circularized cDNAs are linearized by a restriction enzyme. In some embodiments, the restriction enzyme is a DraI enzyme. Barcoding
[0078] In some embodiments, prior step ii), the method further comprises barcoding each cDNA. In some embodiments, each cDNA is barcoded through a “split-pool” approach. In some embodiments, the “split-pool” approach comprises multiple “split and pool” cycles, wherein each “split and pool” cycle comprises a “split” step that the plurality of cells is distributed individually to attach a barcode followed by a “pool” step that the cells are mixed together, wherein the “split-pool” approach comprises at least two “split and pool” cycles. In some embodiments, the “split-pool” approach comprises three “split and pool” cycles. In some embodiments, the first cycle comprises adding a Barcode 1 (BC1) at the 3’ end of each cDNAs, wherein the second cycle comprises adding a Barcode 2 (BC2) at the 3’ end of the BC1, and wherein the third cycle comprises adding a Barcode 3 (BC3) at the 3’ end of the BC2. An exemplary protocol of barcoding cDNAs is illustrated in Examples 3 and 4. Exemplary BC1 sequences are listed in Table 10. As persons skilled in the art are aware, such barcode sets (e.g., BC1, BC2, and / or BC3) can be generated using certain software in public domain. In general, such barcode sets (e.g., BC1, BC2, and BC3) need to satisfy certain conditions, such as. each barcode set is different from each other by at least 2 nucleotides difference. Fig. 5 presents a schematic diagram of split-pool approach with four barcode combination to index both mRNA and gRNA in single cells. The first “split and pool” cycle
[0079] In some embodiments, the first “split and pool” cycle comprises a reverse transcription process from total RNAs to cDNAs in the library. In some embodiments, four primers are used in a reaction, such as in a 96 well plate. In some embodiments, the four primers include a Splitseq Round1 polyT primer, a hexamer primer, a DC1 primer, and a DC2 primer. In some embodiments, the four primers include a Splitseq Round1 polyT primer, a hexamer primer, a DC1 primer, a DC2 primer, a CS1 primer, and a CS2 primer. The Splitseq Round1 polyT primer and the hexamer primer can be used to reverse transcribe mRNAs. TheDC1 primer and the DC2 primer are used to reverse transcribe gRNAs with native gRNA scaffolds (e.g. CROP-Seq). The DC1 primer and the CS1 primer are used to reverse transcribe gRNAs with engineered gRNA scaffolds containing a CS1 sequence (e.g. 10x CRISPR scaffold). The DC1 primer and the CS2 primer are used to reverse transcribe gRNAs of interest with engineered gRNA scaffolds containing a CS2 sequence (e.g.10x CRISPR scaffold).
[0080] In some embodiments, the Splitseq Round1 polyT primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 23-118. In some embodiments, the Splitseq Round1 polyT primer can be selected from any one of SEQ ID Nos: 23-118 (Table 2).
[0081] In some embodiments, the hexamer primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 119-214. In some embodiments, the hexamer primer can be selected from any one of SEQ ID Nos: 119-214 (Table 3).
[0082] In some embodiments, the DC1 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 215-310. In some embodiments, the DC1 primer can be selected from any one of SEQ ID Nos: 215-310 (Table 4).
[0083] In some embodiments, the DC2 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 311-406. In some embodiments, the DC2 primer can be selected from any one of SEQ ID Nos: 311-406 (Table 5).
[0084] In some embodiments, the CS1 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 407-502. In some embodiments, the CS1 primer can be selected from any one of SEQ ID Nos: 407-502 (Table 6).
[0085] In some embodiments, the DC2 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 697-792. In some embodiments, the CS1 primer can be selected from any one of SEQ ID Nos: 697-792 (Table 7).
[0086] In some embodiments, the CS1 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 793-888. In someembodiments, the CS1 primer can be selected from any one of SEQ ID Nos: 793-888 (Table 8).
[0087] In some embodiments, the CS2 primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 889-984. In some embodiments, the CS1 primer can be selected from any one of SEQ ID Nos: 889-984 (Table 9). Table 2. Exemplary sequences of Splitseq Round1 polyT primersTable 3. Exemplary sequences of hexamer primersTable 4. Exemplary sequences of DC1 primersTable 5. Exemplary sequences of DC2 primersTable 6. Exemplary sequences of CS1 primersTable 7. Exemplary sequences of DC2 primers compatible with v3 kitTable 8. Exemplary sequences of CS1 primers compatible with v3 kitTable 9. Exemplary sequences of CS2 primers compatible with v3 kitTable 10. Exemplary barcode sequences.ACTTAGCTĴĵCATTCTACĴĶThe second “split and pool” cycle
[0088] In some embodiments, the second “split and pool” cycle comprises adding a Barcode 2 (BC2) at the 3’ end of the BC1 linked with a target cDNA. In some embodiments, a BC2 segment comprising the BC2 is add to the 3’ end of the BC1 through a ligation. In some embodiments, the ligation is completed using a ligase. In some embodiments, the ligase is a T4 DNA Ligase.
[0089] In some embodiments, the BC2 segment comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 503-576, and 673- 694. In some embodiments, the BC2 segment can be selected from any one of SEQ ID Nos: 503-576, and 673-694 (Table 11). Table 11. Exemplary sequences of the BC2 segmentsThe third “split and cycle
[0090] In some embodiments, the third “split and pool” cycle comprises adding a Barcode 3 (BC3) at the 3’ end of a BC2 linked with a BC1 and a target cDNA. In some embodiments, a BC3 segment comprising the BC3 is add to the 3’ end of the BC2 through a ligation. In someembodiments, the ligation is completed using a ligase. In some embodiments, the ligase is a T4 DNA Ligase.
[0091] In some embodiments, the BC3 segment comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to any one of SEQ ID Nos: 577-672. In some embodiments, the BC3 segment can be selected from any one of SEQ ID Nos: 577-672 (Table 12). Table 12. Exemplary sequences of the BC3 segmentsLinkers
[0092] In some embodiments, BCs are linked via a linker. The linkers in the barcoded cDNAs can be same or different. In some embodiments, the BC1 and the BC2 are linked via a linker comprising a sequence of SEQ ID NO: 6 or 7. In some embodiments, the BC2 and the BC3 are linked via a linker comprising a sequence of SEQ ID NO: 6 or 7. Sequencing
[0093] In some embodiments, the method further comprises a step of sequencing the enriched RNAs of interest using one or more sequencing primers. In some embodiments, the one or more sequencing primers are Illumina sequencing primers. In some embodiments, the Ultima sequencing primer comprises a PS primer having at least 80% identity to SEQ ID NO: 989. Table 13 shows exemplary sequencing parameters. Table 13. Exemplary sequencing parameters.Optimized Perturb-seq
[0094] The present disclosure provides an optimized Perturb-seq method with enriched RNAs of interest in a single-cell RNA library. The method comprises: a) reverse-transcribing the isolated total RNAs to cDNAs; b) barcoding the cDNAs; c) amplifying the barcoded cDNAs to generate a cDNAs library, d) enriching the RNAs of interest through a circularization-based RNA capture, ande) conducting a Perturb-seq.
[0095] In some embodiments, the single-cell RNA library comprises a plurality of cDNAs that are reverse-transcribed from total RNAs in a single cell. In some embodiments, the total RNAs are isolated form a fixed cell. In some embodiments, the cell is fixed by formaldehyde. In some embodiments, the cell is fixed by formaldehyde in combination with a crosslinker. In some embodiments, the crosslinker is a Bissulfosuccinimidyl suberate (BS3). An exemplary protocol of the cell fixation is illustrated in Example 2. IV. Compositions
[0096] In another aspect, the present disclosure also provides a kit for detecting or enriching one or more RNAs of interest in a single-cell RNA library. In some embodiments, the kit comprises a list of the components (e.g. oligo sequences, enzymes) for a circularization-based RNA capture. In some embodiments, the kit comprises a list of the components (e.g. oligo sequences, enzymes) for a Perturb-seq.
[0097] In some embodiments, the kit for a circularization-based RNA capture comprises a ligase; and a 5’ primer and a 3’ primer to amplify one or more RNAs of interest. In some embodiments, the ligase is a ssDNA ligase.
[0098] In some embodiments, the kit further comprises one or more barcode oligonucleotides. In some embodiments, the one or more barcode oligonucleotides comprise a sequence having at least 80% identity to a sequence from Tables 4-9. In some embodiments, the one or more barcode oligonucleotides comprise a DC1 primer selected from any one of SEQ ID Nos: 215-310 (Table 4). In some embodiments, the one or more barcode oligonucleotides comprise a DC2 primer selected from any one of SEQ ID Nos: 311-406 and 697-792 (Tables 5 and 7). In some embodiments, the one or more barcode oligonucleotides comprise a CS1 primer selected from any one of SEQ ID Nos: 407-502 and 793-888 (Tables 6 and 8). In some embodiments, the one or more barcode oligonucleotides comprise a CS2 primer selected from any one of SEQ ID Nos: 889-984 (Table 9).
[0099] In some embodiments, the kit is for detecting or enriching one or more gRNAs of interest in a single-cell RNA library. In some embodiments, the kit further comprises a 5’ gRNA primer and a 3’ gRNA primer.
[0100] In some embodiments, the 5’ gRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 17. In some embodiments, the 5’ gRNA primer further comprises a barcode (BC) sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. In some embodiments, the 5’ gRNA primer further comprises a P5 sequence of SEQ ID NO: 21. In some embodiments, the 5’ gRNA primer is a P5-gRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 15 or 987. In a particular embodiment, the 5’ gRNA primer is a P5-gRNA primer comprising a sequence of SEQ ID NO: 15 or 987.
[0101] In some embodiments, the 3’ gRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 18. In some embodiments, the 3’ gRNA primer further comprises a BC sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. The BC sequence in the 3’ gRNA primer can be the same or different to the BC sequence in the 5’ gRNA primer. In some embodiments, the 3’ gRNA primer further comprises a P7 sequence of SEQ ID NO: 22. In some embodiments, the 3’ gRNA primer is a P7-gRNA primer comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 16 or 988. In a particular embodiment, the 3’ gRNA primer is a P7-gRNA primer comprising a sequence of SEQ ID NO: 16 or 988.
[0102] In some embodiments, the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 15 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 16. In some embodiments, the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 987 and the 3’ primer comprises a sequence of having at least 80% identity to SEQ ID NO: 988.
[0103] In some embodiments, the kit is for detecting or enriching one or more mRNAs of interest in a single-cell RNA library. In some embodiments, the kit further comprises a 5’ mRNA primer and a 3’ mRNA primer.
[0104] In some embodiments, the 5’ mRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 19. In some embodiments, the 5’ mRNA primer further comprises a barcode (BC) sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. In some embodiments, the 5’ mRNA primer further comprises a P5 sequence of SEQ ID NO: 21. In a particular embodiment, the 5’ mRNA primer is a P5-mRNA primer comprising a sequence of SEQ ID NO: 1.
[0105] In some embodiments, the 3’ mRNA primer comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identity to SEQ ID NO: 20. In some embodiments, the 3’ mRNA primer further comprises a BC sequence. In some embodiments, the BC sequence comprises a random sequence of variable length. In some embodiments, the BC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the BC sequence comprises the 6 nucleotides as NNNNNN (N6), where N is any nucleotide. The BC sequence in the 3’ mRNA primer can be the same or different to the BC sequence in the 5’ mRNA primer. In some embodiments, the 3’ mRNA primer further comprises a P7 sequence of SEQ ID NO: 22. In a particular embodiment, the 3’ mRNA primer is a P7-mRNA primer comprising a sequence of SEQ ID NO: 2.
[0106] In some embodiments, the kit can detect or enrich both mRNAs of interest in a single- cell RNA library. In such embodiments, the kit comprises four primers. In some embodiments, the four primers comprise a 5’ gRNA primer, a 3’ gRNA primer, a 5’ mRNA primer, and a 3’ mRNA primer described above.
[0107] In another aspect, the present disclosure also provides a system for detecting or enriching one or more RNAs of interest in a single-cell RNA library. In some embodiments,the system comprises a list of the components (e.g. oligo sequences, enzymes) for a circularization-based RNA capture. In some embodiments, the system comprises a list of the components (e.g. oligo sequences, enzymes) for a Perturb-seq. V. Adaptability of Circular Capture to Other Single-Cell Assays
[0108] The circular capture strategy designed for CC-Perturb-Seq is highly versatile and can be readily integrated into a variety of single-cell transcriptomic platforms. To adapt to other single-cell combinatorial indexing protocols beyond the specific sequences included here, the adapter sequences in the RT primer can be changed in a straightforward manner. Beyond split- pool combinatorial indexing, circular capture is compatible with droplet-based systems such as the 10x Genomics Chromium 3' gene expression assay, as well as microwell-based technologies like BD Rhapsody and Seq-Well. In these workflows, circularization can be performed after cDNA synthesis by incorporating scaffold-targeting capture sequences into the sgRNA design. For instance, in the 10x 3' workflow or other 3’ capture single-cell RNA-seq protocols, non-polyadenylated RNAs such as Pol III-transcribed sgRNAs can be efficiently recovered by using custom reverse transcription primers (or direct capture primers) that bind to conserved regions of the sgRNA scaffold, followed by enrichment and circularization after barcode incorporation. In microwell systems, where transcripts are bulk-amplified after cell lysis, the circularization step can be inserted before library prep to specifically enrich guide RNA-derived cDNA. This modularity makes the circular capture method broadly applicable for recovering small RNAs and perturbation barcodes across a wide range of single-cell sequencing platforms. VI. Primer Compatibility Across sgRNA Vectors in CC-Perturb-seq
[0109] As shown in Table 14, the primer set used in CC-Perturb-seq has been experimentally confirmed to be compatible with several widely adopted sgRNA scaffold designs, including, but not limited to, the original spCas9 scaffold (Cong LE, Ran F. Ann et al. Science 2013), opti-scaffold (Chen et al. Cell 2013), and CS1 vectors (Replogle et al. Nat Biotechnol. 2020). This demonstrates the method’s broad applicability across common CRISPR screening platforms without the need for substantial primer re-engineering. Table 14. Primer Compatibility Across sgRNA Vectors
[0110] Importantly, the modular nature of the circular capture approach also enables the design of customized primers tailored to novel scaffold variants, ensuring extensibility to future vector systems. Table 15 summarizes the compatibility of CC-Perturb-seq primers with various sgRNA scaffold types. Table 15. the compatibility of CC-Perturb-seq primers with various sgRNA scaffold types*XXXXXXXX indicates 8-bp 96 barcode set EXAMPLES
[0111] The present disclosure will be better understood in view of the following non-limiting examples. The following examples are intended for illustrative purposes only and do not limit in any way the scope of the present invention. Example 1. Overview of the Technology
[0112] This example illustrates an overview of “split-pool” Perturb-seq with a circularization-based sgRNA capture.
[0113] This application comes with a comprehensive protocol for conducting split-pool Perturb-seq with a circularization-based sgRNA capture. Here, we provide a summary of the essential steps required to carry out the protocol (Fig. 1). The protocol starts with sgRNA positive cells that are ready for generating sequencing-ready single cell libraries for assessing gene expression profiles along with CRISPR-mediated perturbation. Step 1: Cell Preparation and Fixation
[0114] Cells were prepared by suspending in a suitable buffer, ensuring a single-cell suspension before fixation to prevent clumping and ensure uniform treatment. The fixation process included a combination of formaldehyde and bissulfosuccinimidyl suberate (BS3) crosslinker, to stabilize RNA and proteins and allow for the physical separation of cells while preserving cellular context. This method was selected for its superior efficiency in crosslinking compared to formaldehyde alone. Step 2: Combinatorial Barcoding
[0115] Previous Perturb-Seq was conducted using commercial kits by 10x Genomics that enable profiling both the transcriptome and identifying corresponding sgRNAs in thousands of individual cells, making it possible to investigate tens of perturbations concurrently. However, studying thousands of perturbations remains cost prohibitive. To reduce the expense of single- cell Perturb-Seq, we have developed a method that involves combinatorial barcoding. The core of this barcoding scheme revolves around subjecting cells to multiple cycles of a "split-and- pool" process. During the "split" phase, a large number of cells are distributed among numerous wells to attach a section of a barcode. After cells reunite in the "pool" phase, potentiallyundergoing remixing and re-division during the subsequent split stage of the next round, the intended barcoding result is attained. This repetitive "split-and-pool" process significantly augments the maximum number of barcodes. In our protocol, commencing with fewer than 96 barcodes in each cycle, after four rounds (involving cDNA synthesis from mRNA, two ligation rounds, and one PCR amplification round), the ability to barcode millions of cells becomes readily attainable. The barcoding rounds are performed in fixed cells.
[0116] In the first cDNA synthesis round, we use two types of primers: (1) a mixture of poly(dT) and random hexamer primer sequences that enable the production of cDNA from poly-adenylated mRNA for assessing gene expression; (2) a mixture of primers that priming from the scaffold of sgRNA transcripts produced by Pol III RNA polymerase.
[0117] The combinatorial barcoding process involved distributing cells across a 96-well plate where each well contained unique barcode primers. This "split-and-pool" technique was repeated over multiple cycles to attach a unique barcode combination to each cell. Each round of barcoding was designed to attach additional barcode sequences, exponentially increasing the number of unique identifiers possible with each subsequent round. The process includes: Round 1: cDNA synthesis using a mix of poly-dT, random hexamers, and primers specific to the scaffold of sgRNAs to initiate transcription from polyadenylated mRNA and RNA Pol III transcribed sgRNAs. Rounds 2 and 3: Two successive ligation steps to attach additional barcode sequences. Final Round: Cells are divided into sub-libraries. Each library will undergo PCR amplification at the library construction step with barcoded primers to add a 4thbarcode. Step 3: Library Construction and Size Selection
[0118] Following Round 3 barcoding, cells were lysed and the BS3-formaldehyde crosslinks were reversed using a combination of heat and proteinase K treatment, which released the RNA and allowed for the synthesis of cDNA. Separate libraries for gene expression analysis (GEX library) and sgRNA identification (sgRNA library) were constructed from this cDNA. We attach the 4thbarcode during the PCR amplification at this step. Another key step involved using SPRIselect beads to perform size selection, effectively enriching for smaller cDNA fragments that contain sgRNA sequences. Step 4: Circularization approach
[0119] Following the barcoding rounds, each cell is linked with a distinct combination of barcodes. Subsequently, the cells undergo lysis and reverse-crosslinking (we apply both heat treatment in combination with proteinase digestion to remove bound protein and release free RNAs) to release the cDNAs, which will undergo conversion into sequencing-ready libraries. At this stage, two libraries are generated from the same cDNA pool: one for evaluating gene expression (GEX library) and another for identifying the sgRNAs expressed in individual cells (sgRNA library). Here, we have developed a technique enabling the enrichment of each specific library from the shared cDNA pool. After converting mRNA and gRNA into cDNA, both sub-libraries, equipped with common PCR handles, are co-amplified initially. Subsequent to amplification, we utilize SPRIselect beads (Beckman Coulter) for size selection to isolate cDNAs under 250 bp. These smaller fragments potentially include gRNAs as well as other small RNA types such as tRNAs, miRNAs, and other non-coding RNAs. These are subsequently converted to single-stranded DNAs (ssDNAs), which are then circularized. Finally, we apply PCR primers specifically targeting the gRNA's scaffold to selectively amplify the gRNAs.
[0120] A frequently used method for constructing full-length cDNA libraries from RNA samples obtained from individual cells involves template-switching reverse transcription. This process results in cDNAs that differ solely at one end, attributable to specific reverse transcription primers, while the other end remains constant and carries the sequence of the template-switching oligo. Given that poly-A RNAs (which contribute to the GEX library) and sgRNAs (which contribute to the sgRNA library) are transcribed from distinct priming oligos, cDNAs can possess a different PCR handle at one end for each type of library. Nevertheless, having only one differing PCR handle is insufficient to selectively enrich one library type over the other from the shared cDNA pool. Therefore, we have devised the Circularization approach to selectively amplify the sgRNA library over the GEX library through PCR from the shared cDNA pool.
[0121] To differentiate sgRNA sequences from other small RNAs within the smaller size cDNA pool, we implemented a circularization strategy. This involved converting the selected small ssDNAs into circular DNA templates, which were then specifically amplified using primers targeting the sgRNA scaffold. This method ensures that only sgRNA sequences were amplified, significantly increasing the specificity and efficiency of sgRNA library preparation. Step 5: Sequencing
[0122] mRNA (GEX) and gRNA libraries can be pooled and sequenced together using any Illumina sequencing platform. Recommend cycle number: 70 cycles for Read 1, 6-10 cycles for index 1 (depending on the length of the barcode), 6-10 cycles for index 2, 86 cycles for Read 2. Example 2. Cell Fixation Protocols
[0123] This example illustrates two cell fixation protocols.
[0124] In some embodiments, cells can be fixed using a parse fixation kit. A Cell Fixation (Parse Fixation kit) protocol is described below. Cell Fixation (Parse Fixation kit) protocol. - set swing bucket centrifuge to 4ºC - thaw and keep on ice (from Parse Fixation v2 kit) o Cell Prefixation Buffer, Cell Buffer, Cell Fixation Solution, Cell Fixation Additive, Cell Permeabilization Solution, Cell Neutralization Buffer, RNase Inhibitor - gather from 4ºC: o 7.5% BSA (Thermo 15260037, low RNase, Parse doesn’t recommend substitutions) - keep at room temp: o DMSO, 40μm strainer (2 per sample) 1. Add 550μL Cell Fixation Additive directly into Cell Fixation Solution, mix thoroughly by pipetting 5x (with P1000 set to 750μL), mark on cap and keep on ice 2. Add 50μL RNase Inhibitor directly into Cell Prefixation Buffer, mix thoroughly by pipetting 5x (with P1000 set to 750μL), mark on cap and keep on ice 3. Add 17μL RNase Inhibitor directly into Cell Buffer, mix thoroughly by pipetting 5x (with P1000 set to 750μL), mark on cap and keep on ice - after mixing reagents, should only be freeze-thawed once and stored for up to 1 month at -20ºC 4. If sample is cell-limited or prone to clumping, prepare Cell Prefixation Buffer + BSA fresh and use the same day, mix thoroughly by pipetting 5x, store on ice - 200μL per sample: 187.5μL Cell Prefixation Buffer (with RNase Inhibitor) + 12.5μL 7.5% BSA 5. Prepare cells in single cell suspension, put on ice, count cells 6. Take 100k - 1M cells to a tube, centrifuge at 200g at 4ºC for 10min, remove supernatant 7. Fully resuspend pellet in 187.5μL cold Cell Prefixation Buffer, or with BSA - failure to fully resuspend cells may result in elevated doublet 8. Pipette cells through a 40μm strainer into a new tube, keep on ice - for cells larger than 40μm, can use 70μm or 100μm strainer instead - press pipette tip directly against strainer with force 9. Add 62.5μL cold Cell Fixation Solution, mix immediately by pipetting exactly 3x (with pipette set to 62.5μL), incubate on ice for 10min - additional mixing will lead to more doublets 10. Add 20μL cold Cell Permeabilization Solution, mix by pipetting 3x (with pipette set to 62.5μL), incubate on ice for 3min11. Add 1mL cold Cell Neutralization Buffer, gently invert the tube once to mix, keep on ice - do NOT vortex Cell Neutralization Buffer, invert the tube five times to mix 12. Centrifuge at 200g at 4ºC for 10min, remove supernatant 13. Fully resuspend cells in 150μL cold Cell Buffer (with P1000 set to 150μL), keep on ice 14. Pass cells through a 40μm strainer into a new tube with P1000 pipette, keep on ice 15. Count the number of cells, keep cells on ice during counting and proceed quickly to minimize the time fixed cells are out 16. To freeze cells: 1) add 2.5μL DMSO, gently flick tube 3x to mix, incubate on ice for 1min 2) repeat twice to add a total of 7.5μL DMSO 3) mix the final suspension by gently pipetting 5x with P200 pipette without creating bubbles (set to 75μL) o do NOT vortex cells 4) split into aliquots (no more than 500k cells per aliquot) 5) store in Mr. Frosty (cooling at -1ºC / min) at -80ºC - Stop: can store at -80ºC
[0125] In some embodiments, cells can also be fixed using Formaldehyde combined with BS3. A Cell Fixation (Formaldehyde with BS3) protocol is described below. Cell Fixation (Formaldehyde with BS3) - set swing bucket centrifuge to 4ºC - thaw BS3 powder (Thermo A39266) to room temp for at least 45min - prepare per sample, keep on ice o 1X PBS + RNase Inhibitor: 4mL PBS + 10μL Protector (Sigma 3335399001) + 10μL Enzymatics (Y9240L) o 0.5X PBS + RNase Inhibitor: 0.5mL 1X PBS+RI + 0.5mL H2O 1. Prepare cells in single cell suspension, put on ice, count cells 2. Centrifuge at 300g at 4ºC for 5min, remove supernatant 3. Resuspend cells in 750μL cold 1X PBS+RI 4. Pipette cells through a 40μm strainer into a new tube, keep on ice - for cells larger than 40μm, can use 70μm or 100μm strainer instead - press pipette tip directly against strainer with force 5. Add 50μL fresh 16% Formaldehyde (for final 1%, Thermo 28906) to fix, mix by pipetting 3x (with pipette set to 500μL), incubate on ice for 10min - additional mixing will lead to more doublets 6. Add 16μL 10% Triton (for final 0.2%, Sigma T8787) to permeabilize, mix by pipetting 5x, incubate on ice for 3min 7. Add 160μL 1M Tris (Thermo AM9856) 8. Centrifuge at 500g at 4ºC for 5min, remove supernatant 9. Wash with 1mL 1X PBS+RI, centrifuge at 500g at 4ºC for 5min, remove supernatant 10. Resuspend cells in 960μL 1X PBS+RI 11. Immediately before use, resuspend BS3 in water for 50mM final: 2mg BS3 + 70μL H2O - reconstituted BS3 quickly hydrolyzes in water, use immediately 12. Add 40μL 50mM BS3 to resuspended cells, incubate on ice for 20min13. Add 20μL 10% Triton, incubate on ice for 3min 14. Add 200μL 1M Tris, centrifuge at 500g at 4ºC for 5min, remove supernatant 15. Resuspend in 300μL 0.5X PBS+RI 16. Pass cells through a 40μm strainer into a new tube, keep on ice 17. Count the number of cells, keep cells on ice during counting and proceed quickly to minimize the time fixed cells are out 18. To freeze cells: 6) add 5μL DMSO, gently flick tube 3x to mix, incubate on ice for 1min 7) repeat twice to add a total of 15μL DMSO 8) mix the final suspension by gently pipetting 5x with P200 pipette without creating bubbles (set to 75μL) o do NOT vortex cells 9) split into aliquots (no more than 500k cells per aliquot) 10) store in Mr. Frosty (cooling at -1ºC / min) at -80ºC - Stop: can store at -80ºC Example 3. Preparation of Barcodes and Reagents
[0126] This example illustrates the materials and methods in preparation for “split-pool” barcoding.
[0127] Four rounds of barcoding: 48x96x96 wells x 8 PCR reactions = 3.5M barcode combinations, which are enough to uniquely label up to 100k cells while avoiding doublets. We loaded ~400k cells (~4X) for Round1 RT. The method described below can be used to barcode 100k cells. Preparation of Barcodes and Reagents 1. Generate 3 barcoding stock plates - oligos o RT Barcode Plate (100μM, / 5Phos / ACTGTGG-N8-polyT or -random_hexamer or -direct_capture) o Round2 Barcode Plate (100μM, / 5Phos / CATCGGCGTACGACT-N8- ATCCACGTGCTTGAG (SEQ ID NO: 695)) o Round3 Barcode Plate (100μM, / 5Biosg / CAGACGTGTGCTCTTCCGATCT- N10-N8-GTGGCCGATGTTTCG (SEQ ID NO: 696)) o Round2 Linker (1mM, CCACAGTCTCAAGCACG (SEQ ID NO: 6)) o Round2 Blocking (1mM, CGTGCTTGAGACTGTGG (SEQ ID NO: 8)) o Round3 Linker (1mM, TACGCCGATGCGAAACATCG (SEQ ID NO: 7)) o Round3 Blocking (1mM, GTGGCCGATGTTTCGCATCGGCGTACGACT (SEQ ID NO: 9)) - prepare: o 25mM NaCl to help annealing: 250μL 5M NaCl (Thermo J60434) + 49.75mL H2O 1) Round1 RT barcode stock plate: 12.5μM polyT + 12.5μM random hexamer +6μM direct capture per wello 12.5μL polyT + 12.5μL random hexamer + 6μL direct capture + (fill to 100μL) H2O 2) Round2 ligation stock plate: 12μM Round2 Barcode + 11μM Round2 Linker per well o overshot for 120 wells: 132μL (1.1μLx120) Round2 Linker (1mM) + 10.428mL (86.9μLx120) H2O (with 25mM NaCl) o 88μL diluted Round2 Linker + 12μL Round2 Barcode (100μM) per well 3) Round3 ligation stock plate: 14μM Round3 Barcode + 13μM Round3 Linker per well o overshot for 120 wells: 156μL (1.3μLx120) Round3 Linker (1mM) + 10.164mL (84.7μLx120) H2O (with 25mM NaCl) o 86μL diluted Round2 Linker + 14μL Round3 Barcode (100μM) per well 4) Anneal Round2 and Round3 ligation plates of Barcode+Linker- each Split-seq experiment only requires 8μL / well of RT Barcode and 10μL / well of Round2 and Round3 Barcode+Linker, store stock plates in -20ºC 2. Prepare Reverse Transcription Mix: Table 17.3. Prepare 3 barcoding plates for experiments - can prepare in advance, then store at -20ºC 1) Round1 plate: o 18μL Reverse Transcription Mix + 8μL primer mix from Round1 RT barcode stock plate o final volume should be 26μL per well, will add 14μL cells for total 40μL RT reaction 2) Round2 and Round3 plates: o 10μL Barcode+Linker from Round2 and Round3 ligation stock plates 4. Prepare Spin Additive: 10% Triton X-100- 1mL Triton X-100 (Sigma T8787, room temp) + 9mL H2O, store at 4ºC 5. Prepare Dilution Buffer: 0.5X PBS + RNase Inhibitor - can be prepared in advance, then store at -20ºC, keep on ice after thaw. Table 18.6. Prepare Resuspension Buffer: NEBuffer r3.1 + RNase Inhibitor - can be prepared in advance, then store at -20ºC, keep on ice after thaw. Table 19.7. Prepare Ligation Mix: T4 Ligase + BSA + RNase Inhibitor - for 96 wells, 2.04mL total - if want to prepare in advance, do NOT add T4 DNA Ligase, store at -20ºC Table 20.8. Prepare Round2 Stop Mix: 26.4μM Round2 Blocking - can be prepared in advance, then store at -20ºC, keep on ice after thaw - 37μL Round2 Blocking (1mM) + 350μL 10X Ligase Buffer + 1013μL H2O -> total 1.4mL9. Prepare Round3 Stop Mix: Round3 Blocking, 125mM EDTA - can be prepared in advance, then store at -20ºC, keep on ice after thaw - 37μL Round3 Blocking (1mM) + 800μL 0.5M EDTA + 2363μL H2O -> total 3.2mL 10. Prepare Pre-Lyse Wash Buffer: 0.1% Triton X-100 / PBS + RNase Inhibitor - can be prepared in advance, then store at -20ºC, keep on ice after thaw Table 21.Barcoding Single Cells
[0128] In this process, a swing bucket centrifuge can be used for all spins. A fixed-angle centrifuge will lead to substantial cell loss. Polypropylene tubes can be used instead of polystyrene tubes because the latter will lead to substantial cell loss. To maximize cell retention during pooling, pipetting up and down several times (in the middle and on the front & back sides) in each well before pooling is recommended. In addition, to avoid excess bubble formation (not affect quality), one can pipette up and down with pipette set to 10μL less volume, pooling, then collecting any remaining liquid. Below is the protocol of barcoding single cells. 1. Thaw Round1, Round2, Round3 plates in a thermocycler o Lid Temperature: 70ºC o Reaction Volume: 26μL for Round1, 10μL for Round2 & Round3 Table 22.2. Sample Counting and Loading Setup 1) thaw fixed cell samples at 37ºC until all ice crystals dissolve, then put on ice o important to fully thaw samples before placing on ice 2) count the number of cells in each sample 3) dilute samples with Dilution Buffer (0.5X PBS + RNase Inhibitor), put on ice 3. Round1 Reverse Transcription Barcoding 1) centrifuge Round1 plate at 100g for 1min, put on ice) add 14μL diluted cells to each well of Round1 plate according to the Loading Table; immediately after adding cells, mix gently by pipetting up and down exactly 3x o when pipetting same sample into many wells, periodically mix sample by gentle pipetting to avoid cell settling; do NOT vortex cells o use different tips when pipetting cells into each well ) reverse transcription (~40min) o Lid Temperature: 70ºC o Reaction Volume: 40μL (14μL cells + 18μL RT Mix + 8μL Primer Mix) Table 23.) pool all wells from Round1 Plate into a 2mL tube on ice o set pipette to 30μL, mix up and down at least 3x on each side to retain settled cells o both Round1 Plate and 2mL tube with pooled cells should be kept on ice during pooling ) discard Round1 Plate ) add 9.6μL (for 24-well) or 19.2μL (for 48-well) Spin Additive (10% Triton X-100) for final concentration of 0.1% to the pooled cells, gently invert once to mix) centrifuge the pooled cells at 200g at 4ºC for 10min in a swinging bucket o proceed to the next step immediately, avoid dislodging the cell pellet ) remove supernatant with P1000 and P200 pipettes, leave ~40μL of liquid o do not disturb the pellet ) gently resuspend with 1mL Resuspension Buffer (NEBuffer r3.1 + RNase Inhibitor), once cells are fully resuspended, add an additional 1mL for 2mL total, keep on ice ound2 Ligation Barcoding ) make sure 20μL T4 DNA Ligase (2,000U / μL) have been added to the Ligation Mix o do not vortex ) add 2mL of cells in Resuspension Buffer into 2.04mL Ligation Mix, mix 10x with P1000, keep on ice ) centrifuge Round2 Plate at 100g for 1min, keep at room temp ) add entirety of cells to a basin, add 40μL cell mix to each well of Round2 plate; as adding cells, pipetting up and down exactly 2x to ensure proper mixing o avoid cells settling in basin by gently pipetting up and down 2x before transferring cells o if volume is insufficient to fill every well, a few wells can be left empty o use different tips when pipetting cells into each well ) incubate at 37ºC for 30min while shaking (15s at 900RPM, 1min45s pause) o Lid Temperature: 50ºC o Reaction Volume: 50μL (40μL cell + 10μL R2 primer)6) vortex Round2 Stop Mix briefly, add all to a basin 7) transfer Round2 Plate from thermocycler, keep at room temp 8) add 10μL Round2 Stop Mix to each well of Round2 Plate, pipette up and down exactly 3x to ensure proper mixing o use different tips when pipetting Stop Mix into each well 9) incubate at 37ºC for 30min while shaking (15s at 900RPM, 1min45s pause) o Lid Temperature: 50ºC o Reaction Volume: 60μL (40μL cell + 10μL R2 primer + 10μL R2 Stop) 10) transfer Round2 Plate from thermocycler, keep at room temp 11) pool all wells from Round2 Plate into a new basin o set pipette to 50μL, mix up and down at least 3x on each side to retain settled cells 12) discard Round2 Plate 13) pass all cells through a 40μm strainer into a new basin o press the tip of pipette against the filter to ensure all liquid passes 5. Round3 Ligation Barcoding - keep on ice: T4 DNA Ligase (2,000U / μL, NEB M0202M, store at -20ºC) 1) add 20μL T4 DNA Ligase (2,000U / μL) to the basin with strained cells, mix by gently pipetting up and down ~20x 2) centrifuge Round3 Plate at 100g for 1min, keep at room temp 3) add 50μL cell mix to each well of Round3 plate; as adding cells, pipetting up and down exactly 2x to ensure proper mixing o avoid cells settling in basin by gently pipetting up and down 2x before transferring cells o if volume is insufficient to fill every well, a few wells can be left empty o use different tips when pipetting cells into each well 4) incubate at 37ºC for 30min while shaking (15s at 900RPM, 1min45s pause) o Lid Temperature: 50ºC o Reaction Volume: 60μL (50μL cell + 10μL R3 primer) 5) vortex Round3 Stop Mix briefly, add all to a basin 6) transfer Round3 Plate from thermocycler, keep at room temp 7) add 20μL Round3 Stop Mix to each well of Round3 Plate, pipette up and down exactly 3x to ensure proper mixing o use different tips when pipetting Stop Mix into each well o no incubation required, proceed directly to pooling 8) pool all wells from Round3 Plate into a new basin o set pipette to 70μL, mix up and down at least 3x on each side to retain settled cells 9) discard Round3 Plate 10) pass all cells through a 40μm strainer into a new 15mL tube o press the tip of pipette against the filter to ensure all liquid passes 6. Lysis and Sublibrary Generation - keep on ice: Proteinase K (20mg / mL, Thermo 25530049, store at -20ºC) 1) prepare 2X Lysis Buffer: 20mM Tris, 400mM NaCl, 100mM EDTA, 4.4% SDS o can be prepared in advance, store at 4ºC, keep at 37ºC until use to dissolve precipitate Table 24.2) add 60μL Spin Additive (10% Triton X-100) for final concentration of 0.1% to the pooled cells, gently invert once to mix 3) centrifuge the pooled cells at 200g at 4ºC for 10min in a swinging bucket 4) remove supernatant with P1000 and P200 pipettes, leave ~40μL of liquid 5) gently resuspend with 1mL Pre-Lyse Wash Buffer (0.1% Triton X-100 / PBS + RNase Inhibitor), pipette slowly to prevent mechanical damage to cells; once cells are fully resuspended, add an additional 3mL for 4mL total 6) centrifuge the pooled cells at 200g at 4ºC for 10min in a swinging bucket 7) remove supernatant with P1000 and P200 pipettes, leave ~50μL of liquid 8) gently resuspend pellet with 100μL Dilution Buffer to bring the volume to ~150μL, transfer to a 1.5mL tube, keep on ice 9) set pipette to 80μL, gently pipette up and down 5x, immediately use 6μL to count o 6μL cells + 6μL 2.4nM YOYO-1 (Thermo Y3601, 2.4μL 1mM stock + 997.6μL PBS, 4ºC) o some level of debris is normal 10) aliquot sublibraries, record sublibrary sizes and labels, add Dilution Buffer to 25μL, keep on ice o each sublibrary will have separate sequencing index o useful to have at least one sublibrary with few cells (200-500) with deep sequence (>50,000 reads per cell), this sublibrary provides good estimate of transcript detection per cell that would be expected if other sublibraries were also sequenced deeply o do not overload a sublibrary, 12,500 cells / sublibrary is the maximum 11) per sublibrary, add 25μL 2X Lysis Buffer and 5μL Proteinase K, keep at room temp 12) vortex for 10s to initiate lysis, briefly centrifuge 13) incubate at 65ºC for 60min while shaking (15s at 900RPM, 1min45s pause) o Lid Temperature: 80ºC o Reaction Volume: 55μL (25μL cell + 25μL Lysis Buffer + 5μL Proteinase K) - Stop: can store sublibrary lysates at -80ºC for up to 6 months Example 4: Amplification of Barcoded cDNAs
[0129] This example illustrates a protocol to amplify barcoded cDNAs. 1. Prepare reagents 1) prepare 2X Binding&Washing Buffer: 10mM Tris, 2M NaCl, 1mM EDTA o keep at room temp Table 25. Storage Volumeo keep at room temp Table 26.3) prepare Bind Buffer A: 2X B&W + RNase Inhibitor o store at -20ºC, keep on ice Table 27.4) prepare Bind Buffer B: 0.05% Tween in 1X B&W + RNase Inhibitor o store at -20ºC, keep on ice Table 28.5) prepare Bead Storage Buffer: 0.1% Tween in 10mM Tris + RNase Inhibitor o store at -20ºC, keep on ice Table 29.Prepare Streptavidin beads obtain: o Streptavidin C1 beads (Thermo 65001, 4ºC) 1) vortex Streptavidin beads, take (44μL x # sublibrary) beads to a 1.5mL tube 2) put on magnet until liquid becomes clear (~2min), discard supernatant3) remove from magnet, resuspend with (100μL x # sublibrary) Bead Wash Buffer, ensure all beads are fully resuspended, not stuck to the side 4) put on magnet until liquid becomes clear (~2min), discard supernatant 5) repeat wash twice more for a total of three washes 6) resuspend with (55μL x # sublibrary) Bind Buffer A, keep at room temp 3. Apply Streptavidin beads to sublibrary lysates obtain: o Lysis Neutralizer: 100mM PMSF (Thermo 36978, -20ºC): add isopropanol to dissolve, make sure at saturate level by still having white undissolved in the bottom 1) remove sublibrary lysates from -80ºC, incubate at 37ºC for 5min, ensure no precipitate before proceeding, quick centrifuge sublibrary lysates 2) per sublibrary, add 2.5μL Lysis Neutralizer (100mM PMSF), mix 5x (set to 40μL); quick centrifuge, incubate at room temp for 10min 3) per sublibrary, add 50μL Streptavidin beads in Bind Buffer A, mix 5x (set to 90μL) 4) incubate at 25ºC for 60min while shaking (30s at 1400RPM, 30s pause) 5) quick centrifuge, put on magnet High until liquid becomes clear, discard supernatant o cDNA is unamplified, discarding any beads will result in reduction of transcripts detected 6) resuspend beads with 125μL Bind Buffer B, keep at room temp for 1min 7) put on magnet High until liquid becomes clear, discard supernatant 8) repeat wash with 125μL Bind Buffer B 9) resuspend beads with 125μL Bead Storage Buffer, keep at room temp for 1min 4. Template Switch - oligo: o Split_TSO (100μM, / 5dSp / AAGCAGTGGTATCAACGCAGAGTGAATrGrGrG (SEQ ID NO: 3), order as RNA with HPLC purification) - obtain: o Maxima H Minus RT (200U / μL) and 5X RT Buffer (Thermo EP0753, -20ºC) o dNTP (10mM, NEB N0447L, -20ºC) o PEG 8000 (25% at -20ºC): from 5g powder (Sigma P5413) in H2O (~15mL) for total 20mL 1) prepare Template Switching Mix: o keep on ice Table 30.2) put on magnet High until liquid becomes clear, discard supernatant o cDNA is unamplified, discarding any beads will result in reduction of transcripts detected 3) without resuspending beads, add 125μL H2O, wait 1min, discard supernatant 4) resuspend with 100μL Template Switch Mix, quick centrifuge o Template Switch Mix is viscous, ensure beads are fully resuspended before proceeding 5) incubate at 25ºC for 30min while shaking (30s at 1400RPM, 30s pause) 6) mix by pipetting 5x to resuspend settled beads, incubate at 42ºC for 90min while shaking (30s at 1400RPM, 30s pause) 5. cDNA Amplification - oligo: o Split_Partial-TSO (100μM, AAGCAGTGGTATCAACGCAGAGT (SEQ ID NO: 10)) o Split_TruSeq-Read2 (100μM, CAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 11)) - obtain: o KAPA HiFi 2X Master Mix (Roche KK2602, -20ºC) 1) prepare Amplification Reaction Solution: o keep on ice Table 31.2) mix by pipetting to resuspend settled beads, put on magnet High until liquid becomes clear, discard supernatant 3) without resuspending beads, add 125μL H2O, wait 1min, discard supernatant 4) resuspend with 100μL Amplification Reaction Solution, keep on ice 5) incubate (~70-90min) o adjust the number of 2ndcycles based on number of cells in each sublibrary o Lid Temperature: 105ºC o Reaction Volume: 100μL Table 32.- Stop: can store sublibrary at 4ºC overnight 6. Post-Amplifcation SPRI Clean Up (0.7X for GEX, 0.7X-1.2X for gRNA).
[0130] This stage involves size selection using Solid Phase Reversible Immobilization (SPRI) beads. Each bead is made of polystyrene coated with a magnetite layer and further covered with carboxyl molecules, which reversibly bind DNA when combined with polyethylene glycol (PEG) and salt (20% PEG and 2.5M NaCl). The crowding effect of PEG causes negatively charged DNA to adhere to the carboxyl groups on the beads. The DNA immobilization relies on the concentrations of PEG and salt, making the volume ratio of beads to DNA critical. The size of DNA fragments that bind to or are released from the beads depends on the PEG concentration. For example, a mix of 50ul of DNA with 50ul of beads sets a SPRI:DNA ratio of 1. Changing this ratio influences the DNA fragment sizes that bind or remain in solution; with lower SPRI:DNA ratios, only larger DNA fragments bind to the beads. We employ a 0.7X bead ratio to select for cDNA fragments larger than ~300bp, retaining them for the mRNA sublibrary. Meanwhile, a bead ratio between 0.7X and 1.2X selects for cDNA fragments approximately 180bp to 300bp, which are kept for further downstream enrichment of gRNA. - prepare fresh 85% EtOH 1) put on magnet High until liquid becomes clear o do NOT discard supernatant 2) transfer 90μL clear supernatant to a new tube, discard original tubes with beads 3) vortex SPRI, add 63μL SPRI (0.7X), brief vortex, incubate at room temp for 5min 4) put on magnet High until liquid becomes clear 5) transfer and save 150μL clear supernatant to a new tube (for gRNA) 6) do NOT discard the pellet (for GEX) 7. Pellet Clean Up for GEX 1) without resuspending beads, add 180μL 85% EtOH, wait 1min, discard supernatant 2) repeat wash with 85% EtOH 3) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 2min o do NOT over-dry beads as this will loss yield, cracking of beads is over-drying 4) resuspend beads in 25μL H2O, incubate at 37ºC for 10min to maximize elution 5) put on magnet Low until liquid becomes clear6) transfer 25μL eluted GEX cDNA to a new tube, discard original tubes with beads 7) measure the concentration of cDNA using Qubit dsDNA HS, record for Index PCR 8) run 1μL cDNA (1:10 dilution) on Bioanalyzer - Stop: can store at 4ºC for up to 2 days or -20ºC for up to 3 months 8. Supernatant Clean Up for gRNA 1) add 45μL SPRI (1.2X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear, discard supernatant 3) without resuspending beads, add 180μL 85% EtOH, wait 1min, discard supernatant 4) repeat wash with 85% EtOH 5) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 30s 6) resuspend beads in 50μL H2O, incubate at room temp for 5min to elute 7) put on magnet Low until liquid becomes clear 8) transfer 50μL eluted gRNA cDNA to a new tube, discard original tubes with beads - Stop: can store at 4ºC for up to 2 days or -20ºC for up to 3 months Example 5: Preparing Gene Expression Libraries for Sequencing
[0131] This example illustrates the materials and methods in preparation of gene expression libraries for sequencing. 1. Gene Expression Fragmentation, End Repair, A-Tailing - obtain: o Fragmentation Enzyme: 5X WGS Fragmentation Mix (Enzymatics Y9410L) o Fragmentation Buffer: 10X Fragmentation Buffer (Enzymatics B0330L) 1) brief vortex cDNA, quick centrifuge, take 10μL cDNA (at least 125ng in total), add 25μL H2O to bring total volume to 35μL, keep on ice o concentration based on Qubit, not Bioanalyzer o remaining cDNA can be stored at -20ºC 2) prepare thermal cycler (40min) o Lid Temperature: 70ºC o Reaction Volume: 50μL Table 33.o pre-cool block prior to preparing Fragmentation Mix 3) prepare Fragmentation Mix: o confirm reagents fully thawed and mixed well before using o mix well by pipetting 10x, keep on ice Table 34.4) add 15μL Fragmentation Mix, mix 10x (set to 40μL) on ice, quick centrifuge 5) transfer to pre-cooled thermal cycler and press skip to initiate protocol 2. Post-Fragmentation Double-Sided SPRI Selection (0.6X-0.8X) 1) vortex SPRI, add 30μL SPRI (0.6X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear o do NOT discard supernatant 3) transfer 75μL clear supernatant to a new tube, discard original tubes with beads 4) add 10μL SPRI (0.8X), brief vortex, incubate at room temp for 5min 5) put on magnet High until liquid becomes clear, discard supernatant o this may take longer due to low volume of beads 6) without resuspending beads, add 180μL 85% EtOH, wait 1min, discard supernatant 7) repeat wash with 85% EtOH 8) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 30s o only 30s due to small amount of beads, do NOT over-dry beads as this will loss yield 9) resuspend beads in 50μL H2O, incubate at room temp for 5min to elute 10) put on magnet High until liquid becomes clear 11) transfer 50μL fragmented DNA to a new tube, discard original tubes with beads - Stop: can store at 4ºC overnight or -20ºC for up to 2 weeks 3. Gene Expression Adapter Ligation - oligo: o Split_Adaptor_Top (100μM, ACACTCTTTCCCTACACGACGCTCTTCCGATC*T (SEQ ID NO: 4), *phosphorothioated) o Split_Adaptor_Bottom (100μM, GATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 5)) - obtain: o Adaptor Ligase: WGS Ligase (Enzymatics L6030-W-L) o Adaptor Ligation Buffer: 5X Rapid Ligation Buffer (Enzymatics B9020L) 1) anneal Adaptor Duplex: o per sublibrary: 2.5μL Split_Adaptor_Top + 2.5μL o Split_Adaptor_Bottom (100μM) + 0.125μL 1M NaCl (final 25mM NaCl to help annealing) Table 35.2) prepare Adaptor Ligation Mix:o confirm reagents fully thawed and mixed well before using o mix well by pipetting, keep on ice Table 36.3) add 50μL Adaptor Ligation Mix to 50μL fragmented DNA, mix 10x (set to 80μL), quick centrifuge 4) incubate (15min) o Lid Temperature: 30ºC o Reaction Volume: 100μL Table 37.o proceed directly to next step, do NOT leave in thermal cycler for longer than indicated 4. Post-Ligation SPRI Clean Up (0.8X) 1) vortex SPRI, add 80μL SPRI (0.8X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear, discard supernatant 3) without resuspending beads, add 180μL 85% EtOH, wait 1min, discard supernatant 4) repeat wash with 85% EtOH 5) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 3min o do NOT over-dry beads as this will loss yield 6) resuspend beads in 23μL H2O, incubate at room temp for 5min to elute 7) put on magnet Low until liquid becomes clear 8) transfer 21μL eluted DNA to a new tube, discard original tubes with beads 5. Gene Expression Sublibrary Index PCR - oligo: o Split_P5 (10μM, AATGATACGGCGACCACCGAGATCTACAC-N6- ACACTCTTTCCCTACACGACGC (SEQ ID NO: 1)) o Split_P7 (10μM, CAAGCAGAAGACGGCATACGAGAT-N6- GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 2)) - obtain: o KAPA HiFi 2X Master Mix (Roche KK2602, -20ºC) 1) set up Index PCR per sublibrary on ice, record P5 & P7 index primer pairs used o confirm reagents fully thawed and mixed well before using o ensure no two sublibraries contain the same index primer pair o mix well by pipetting, keep on ice Table 38.2) incubate (~30min) o Lid Temperature: 105ºC o Reaction Volume: 50μL o adjust cycles depending on the amount of cDNA during fragment based on Qubit. Table 39.- Stop: can store at 4ºC overnight 6. Post-Amplification Double-Sided Size Selection (0.6X-0.8X) 1) vortex SPRI, add 30μL SPRI (0.6X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear o do NOT discard supernatant 3) transfer 75μL clear supernatant to a new tube, discard original tubes with beads 4) add 10μL SPRI (0.8X), brief vortex, incubate at room temp for 5min 5) put on magnet High until liquid becomes clear, discard supernatant o this may take longer due to low volume of beads 6) without resuspending beads, add 180μL 85% EtOH, wait 1min, discard supernatant 7) repeat wash with 85% EtOH 8) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 30s o only 30s due to small amount of beads, do NOT over-dry beads as this will loss yield 9) resuspend beads in 20μL H2O, incubate at room temp for 5min to elute 10) put on magnet Low until liquid becomes clear 11) transfer 20μL eluted DNA to a new tube, discard original tubes with beads 12) measure the concentration of GEX sequencing library using Qubit dsDNA HS 13) run 1μL cDNA (1:10 dilution) on Bioanalyzer o there should be a peak between 400-500bp (Fig.7) - Stop: can store at -20ºC for up to 3 monthsExample 6: Preparing gRNA Libraries for Sequencing
[0132] This example illustrates the materials and methods in preparation of gRNA Libraries for Sequencing. 1. gRNA Feature PCR to add 5’-Phosphate and 5’-Biotin - oligo: o Split_5Phos_TSO (100μM, / 5Phos / / 5Phos / NNNN-N8- ACACTCTTTCCCTACACGACGCTCTTCCGATCT-NNNNNNN- AAGCAGTGGTATCAACGCAGAGT (SEQ ID NO: 12)) (N8 is a barcode for each gRNA sublibrary while string of Ns are degenerate / random DNA bases) o Split_5Biotin_Read2 (100μM, / 5Biosg / NNNN-N6- GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 13)) - obtain: o KAPA HiFi 2X Master Mix (Roche KK2602, -20ºC) 1) obtain amplified gRNA cDNA, quick centrifuge, aliquot out 10μL cDNA (1 / 5 of total), add H2O to bring total volume to 188μL (for 1 sublibrary), 94μL (for 2 sublibraries), or 47μL (for 4 sublibraries) 2) set up Feature PCR per sublibrary on ice o confirm reagents fully thawed and mixed well before using o ensure no two sublibraries contain the same index primer pair o expected size: 254bp (scaffold), 279bp (CS1) o aliquot to 100μL per tube Table 40.3) incubate (~30min) o Lid Temperature: 105ºC o Reaction Volume: 100μL Table 41.- Stop: can store at 4ºC overnight2. Post- SPRI Clean Up (1.2X) and pool1) 120μL SPRI (1.2X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear 3) without resuspending beads, add 250μL 85% EtOH, wait 1min, discard supernatant 4) repeat wash with 85% EtOH 5) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 3min o do NOT over-dry beads as this will loss yield 6) resuspend beads in 25μL H2O, incubate at room temp for 5min to elute 7) put on magnet Low until liquid becomes clear 8) transfer 25μL eluted DNA to a new tube, discard original tubes with beads 9) measure the concentration of gRNA Feature PCR DNA using Qubit dsDNA HS 10) combine eluted DNA from sublibraries to be sequenced together - take ~1μg PCR product for circularization - Stop: can store at 4ºC for up to 2 days or -20ºC for up to 3 months 3. Prepare Streptavidin beads (Thermo 65001, 4ºC) 1) vortex Streptavidin beads, take (100μL x μg PCR DNA) beads to a 1.5mL tube 2) put on magnet until liquid becomes clear (~2min), discard supernatant 3) remove from magnet, resuspend with (200μL x μg PCR DNA) Bead Wash Buffer, ensure all beads are fully resuspended, not stuck to the side 4) put on magnet until liquid becomes clear (~2min), discard supernatant 5) repeat wash twice more for a total of three washes 6) resuspend with (50μL x μg PCR DNA) 2X B&W Buffer, keep at room temp 4. Apply Streptavidin beads to biotinylated gRNA cDNA and release of ssDNA 1) per 1μg biotinylated DNA, add 50μL Streptavidin beads in 2X B&W Buffer, mix 5x 2) incubate at 25ºC for 60min while shaking (30s at 1400RPM, 30s pause) 3) quick centrifuge, put on magnet High until liquid becomes clear, discard supernatant 4) resuspend beads with 125μL Bead Wash Buffer, keep at room temp for 1min 5) put on magnet High until liquid becomes clear, discard supernatant 6) repeat wash with 125μL Bead Wash Buffer 7) resuspend beads with 125μL Bead Storage Buffer (0.1% Tween in 10mM Tris), keep at room temp for 1min 8) put on magnet High until liquid becomes clear, discard supernatant 9) elute with 50μL freshly made 0.2M NaOH, incubate at 37ºC for 5min 10) put on magnet Low until liquid becomes clear, transfer supernatant to a new tube 11) repeat elute with another 50μL 0.2M NaOH for 100μL total eluate 12) neutralize by adding 20μL 1M HCl and 50μL 1M Tris for 170μL total o 1M HCl: 0.4mL 12.1M HCl (Sigma H1758) + 4.44mL H2O 5. Purify gRNA ssDNA with Zymos DNA concentrator-5 (D4013) 1) for ssDNA, add 7x170μL = 1190μL DNA Binding Buffer, mix by vortexing 2) transfer mixture to Zymo-Spin Column in Collection Tube o the column holds 800μL3) centrifuge at 10,000-16,000g for 30s, discard flow-through 4) add 200μL DNA Wash Buffer, centrifuge at 10,000-16,000g for 30s 5) repeat wash 6) transfer column to a new tube, add 40μL H2O, incubate at 40ºC for 2min 7) centrifuge at 16,000g for 1min to elute 6. Circularize gRNA ssDNA with CircLigase ssDNA Ligase (BioSearch CL4115K) 1) set up circularization reaction on ice Table 42.2) incubate (~10h) o Lid Temperature: 100ºC o Reaction Volume: 40μL Table 43.7. (Optional) Relinearization gRNA ssDNA by restriction enzyme - oligo: o Split_circ_scaffold (100μM, GTTGATAACGGACTAGCCTTATTTAAACTTGCTATGCTGTTTCCAGC ATAGCTCT (SEQ ID NO: 14)) - obtain: o DraI (NEB R0129L) 1) set up annealing reaction: Table 44.2) anneal reverse complement of scaffold Table 45.3) add 3μL DraI and 1.5μL H2O to annealed reaction, for total 55μL 4) incubate (~1.5h) o Lid Temperature: 80ºC o Reaction Volume: 55μL Table 46.8. Purify gRNA ssDNA with Zymos DNA concentrator-5 (D4013) 1) for ssDNA, add 7 volumes of DNA Binding Buffer, mix by vortexing 2) transfer mixture to Zymo-Spin Column in Collection Tube o the column holds 800μL 3) centrifuge at 10,000-16,000g for 30s, discard flow-through 4) add 200μL DNA Wash Buffer, centrifuge at 10,000-16,000g for 30s 5) repeat wash 6) transfer column to a new tube, add 40μL H2O, incubate at 40ºC for 2min 7) centrifuge at 16,000g for 1min to elute 9a. gRNA Index PCR for Illumina sequencing
[0133] We use a 6bp barcodes here to “index” gRNA sublibrary in a particular sample. We index it so that we can mix gRNA sublibraries from different samples together in one Illumina sequencing lane. - oligo: o Split_circ_gRNA_P5 (100μM, AATGATACGGCGACCACCGAGATCTACAC-N6- AATAAGGCTAGTCCGTTATCAACTTG (SEQ ID NO: 15)) o Split_circ_gRNA_P7 (100μM, CAAGCAGAAGACGGCATACGAGAT- N6-TTTAAACTTGCTATGCTGTTTCCAG (SEQ ID NO: 16)) - obtain: o KAPA HiFi 2X Master Mix (Roche KK2602, -20ºC) 1) set up gRNA Index PCR o expected size: 309bp (scaffold), 334bp (CS1) Table 47.2) incubate (~30min) o Lid Temperature: 105ºC o Reaction Volume: 100μL Table 48.9b. Alternate gRNA Index PCR for sequencing with Ultima Genomics o
[0134] Alternatively, a gRNA library can be prepared to be sequenced with Ultima Genomics’ technologies. To make the gRNA library compatible with Ultima sequencing, a final PCR can be performed using the following primer pair:Ultima_circ_gRNA_PB (100μM, CTGTGTGCCTTGGCAGTCTCAGCTAATAAGGCTAGTCCGTTATCAA CTTG (SEQ ID NO: 987)) o Ultima_circ_gRNA_PS (100μM, CCATCTCATCCCTGCGTGTCTCCGACTGCA-N15- TTAAACTTGCTATGCTGTTTCCAG (SEQ ID NO: 988); where N15 denotes a 15-nucleotide sample-specific barcode.
[0135] Following this PCR step, the resulting gRNA library can be sequenced using the Ultima sequencing primer: PS primer: CCATCTCATCCCTGCGTGTCTCCGACTGCA (SEQ ID NO: 989). - obtain: o KAPA HiFi 2X Master Mix (Roche KK2602, -20ºC) 3) set up gRNA Index PCR o expected size: 309bp (scaffold), 334bp (CS1) Table 49.4) incubate (~30min)o Lid Temperature: 105ºC o Reaction Volume: 100μL Table 50.9. Post-Amplification Double-Sided Size Selection (0.9X) 1) vortex SPRI, add 90μL SPRI (0.9X), brief vortex, incubate at room temp for 5min 2) put on magnet High until liquid becomes clear 3) without resuspending beads, add 250μL 85% EtOH, wait 1min, discard supernatant 4) repeat wash with 85% EtOH 5) centrifuge briefly, put on magnet Low, remove remaining EtOH, dry for 3min o do NOT over-dry beads as this will loss yield 6) resuspend beads in 25μL H2O, incubate at room temp for 5min to elute 7) put on magnet Low until liquid becomes clear 8) transfer 25μL eluted DNA to a new tube, discard original tubes with beads 9) measure the concentration of CRISPR sequencing library using Qubit dsDNA HS 10) run 1μL cDNA (1:10 dilution) on Bioanalyzer - Stop: can store at -20ºC for up to 3 months Example 7: Materials
[0136] This example illustrates the materials used in the technology.
[0137] Consumable Reagents are listed below. Thermo Maxima H Minus RT (200U / μL): Thermo EP0753 Proteinase K: Thermo 25530049 YOYO-1: Thermo Y3601 Streptavidin C1 beads: Thermo 65001 NEB dNTP (10mM): NEB N0447L 10X NEBuffer r3.1: NEB B6003S T4 DNA Ligase (2,000U / μL): NEB M0202M 10X T4 Ligase Buffer: NEB B0202S BSA (Recombinant Albumin, 20mg / mL): NEB B9200SSigma Protector RNase Inhibitor (40U / μL): Sigma 3335399001 Enzymatic Enzymatics RNase Inhibitor (40U / μL): Enzymatic Y9240L 5X WGS Fragmentation Mix: Enzymatics Y9410L WGS Ligase: Enzymatics L6030-W-L Roche - KAPA HiFi 2X Master Mix: Roche KK2602 BioSearch - CircLigase ssDNA Ligase kit: CL4115K Example 8. Preliminary Data
[0138] This example illustrates the materials and methods in preparation of gRNA Libraries for Sequencing.
[0139] With this approach, the capture efficiency of the transcriptome with our approach was very high. For instance, during pilot experiments conducted in hiPSCs, we observed the presence of correct sgRNAs in as much as 95% of the cells, with a median count of 197 unique molecular identifiers (UMIs) per cell when using circular capture. In contrast, when employing CROP-Seq, the median UMI count per cell was 26 (Fig. 2). Notably, this increase in sgRNA detection did not compromise the transcriptome quality, as our mRNA capturing sensitivities reached a median of 26,000 UMIs per cell at 200K read-depth (Fig.3A and 3B).
[0140] We examined additional essential metrics and demonstrated that the circularization approach enables highly targeted enrichment of the gRNA library compared to the GEX library (Table 51). Table 51. Representative CRISPR Application metrics.
[0141] Fraction Reads with Putative Protospacer Sequence: in the first filter only reads in which a predefined constant region (scaffold sequence) of the guide RNA can be found are retained. These reads are termed as "Reads with Putative Protospacer Sequence".
[0142] Fraction Guide Reads: after removing reads without a constant sequence, reads that contain expected protospacer sequences are retained. These reads are termed as Fraction Guide Reads.
[0143] Fraction Guide Reads in Cells: The mRNA sublibrary was used to determine whether a barcode is associated with real cells. Then, fraction of gRNA reads associated with barcodes labeled as cells in the parent whole transcriptome library divided by the total number of mapped reads.
[0144] Guide Reads Usable per Cell: Number of reads passed three filters above (but this number also depends on how deep we sequence the gRNA library)
[0145] Table 51 demonstrates that our enrichment method is highly specific as most of the reads are associated with expected gRNA sequences and there was very little of cross- contamination from mRNA sublibrary.
[0146] We performed a pilot Perturb-Seq experiment using our method, which targets 14 genes (ROCK1, SMARCE1, GSK3B, NANOG, SMAD5, ACVR2A, PLCG1, CDKN1A, KDR, TP53, SMAD1, LEMD3, SOX2, PRKD1) in iPSCs. We assessed knockdown efficiency of those genes using Sceptre and found 13 / 14 genes to have significant knockdown efficiency (Fig.4). Example 9. CC-Perturb-Seq applicability in multiple cell types 1. Application in TeloHAEC
[0147] We conducted a comprehensive comparison between CC-Perturb-Seq (Circular Capture) and CROP-seq, focusing on their performance in TeloHAEC cells, as shown in Fig. 8. Fig. 8A shows violin plots of the number of gene expression UMIs per cell, revealing that both methods achieve comparable transcriptome capture. This confirms that CC-Perturb-Seq does not compromise the ability to profile gene expression while implementing a novel capture strategy. In contrast, Fig. 8B demonstrates a stark difference in guide RNA detection: CC- Perturb-Seq captures substantially more sgRNA UMIs per cell than CROP-seq, validating its enhanced sensitivity and efficiency in capturing Pol III-transcribed guide RNAs. This improved guide recovery translates to higher resolution in detecting gene perturbation effects,as seen in the Fig. 8C quantile-quantile (QQ) plot. Here, CC-Perturb-Seq shows a marked deviation from the null expectation, with a higher number of statistically significant p-values compared to CROP-seq, indicating increased power to resolve true biological effects. To illustrate this improvement, Fig. 8C includes insets showing ITGB1 expression in cells with and without ITGB1-targeting guides. In CC-Perturb-Seq, ITGB1 expression is clearly reduced in cells with the guide, confirming efficient knockdown. In CROP-seq, the effect is much weaker, likely due to poorer guide capture or assignment. Altogether, Fig. 8 highlights that CC-Perturb-Seq not only maintains high transcriptome complexity but also enables more reliable guide detection and knockdown resolution, making it a superior platform for high- throughput CRISPR screening in single-cell contexts. 2. Application in iPSC-EC differentiation
[0148] We conducted large-scale CRISPR perturbation experiments targeting 250 genes across four timepoints (D0-D3) in hiPSCs undergoing EC differentiation. Using 1,926 total guides (6 guides / gene, 15% controls), we captured over 90,000 cells. Circular capture enabled accurate assignment of guides with >90% sgRNA detection per cell, and knockdown (KD) efficiency exceeded 50% on average in both pluripotent and differentiated states. Notably, CROP-seq alternatives showed lower guide recovery and more variability in transcriptomic readouts.
[0149] Fig.9 presents histograms illustrating the distribution of sgRNA UMI counts per cell in hiPSCs (Fig. 9A) and endothelial cells (ECs) (Fig. 9B) with mean at 193 UMIs and 190 UMIs, respectively, demonstrating efficient guide capture with CC-Perturb-Seq. This indicates robust recovery of guide information in the majority of cells. This strong guide recovery underlines the sensitivity of the circular capture approach for detecting perturbations in single- cell CRISPR screens.
[0150] Furthermore, as shown in Fig.10, the high guide RNA capture efficiency enabled by CC-Perturb-Seq allows the SCEPTRE tool to reliably distinguish between perturbed and unperturbed cells. The top panels show clear separation in gRNA counts between cells with (pert) and without (unpert) the guide across multiple examples (Fig. 10A). Even under high multiplicity of infection (MOI = 5.81), where individual cells receive multiple sgRNAs (as shown in Fig. 10B), SCEPTRE maintains clean classification. This indicates robust guide detection and accurate assignment despite the complexity introduced by high MOI.
[0151] Fig. 11 displays the distribution of gene knockdown efficiencies in human induced pluripotent stem cells (hiPSCs) (Fig. 11A) and their differentiated endothelial cell (EC) derivatives (Fig. 11B) using the CC-Perturb-Seq platform. As shown in Fig. 11A, the cumulative curve illustrates that a substantial proportion of gene targets in hiPSCs exhibit high knockdown efficiency, with many surpassing 60-80% repression. This demonstrates that the circular capture-based perturbation method is highly effective in the pluripotent context. Fig. 11B shows the same analysis in ECs, revealing a similar overall pattern. This suggests a modest increase in variability of knockdown efficiency. CC-Perturb-Seq consistently achieves strong knockdown for a wide range of targets in both cell types. Overall, these results confirm that CC-Perturb-Seq is a robust and scalable approach for performing CRISPR interference in both stem and lineage-committed cells, enabling high-resolution functional genomics across diverse cellular contexts. REFERENCES 1. Adamson, Britt, Thomas M. Norman, Marco Jost, Min Y. Cho, James K. Nuñez, Yuwen Chen, Jacqueline E. Villalta, et al. 2016. “A Multiplexed Single-Cell CRISPR Screening Platform Enables Systematic Dissection of the Unfolded Protein Response.” Cell 167 (7): 1867–82.e21. 2. Datlinger, Paul, André F. Rendeiro, Christian Schmidl, Thomas Krausgruber, Peter Traxler, Johanna Klughammer, Linda C. Schuster, Amelie Kuchler, Donat Alpar, and Christoph Bock.2017. “Pooled CRISPR Screening with Single-Cell Transcriptome Readout.” Nature Methods 14 (3): 297–301. 3. Xu, Z., Sziraki, A., Lee, J. et al. Dissecting key regulators of transcriptome kinetics through scalable single-cell RNA profiling of pooled CRISPR screens. Nat Biotechnol (2023) 4. Jiang, L., Dalgarno, C., Papalexi, E., Mascio, I., Wessels, H.-H., Yun, H., Iremadze, N., Lithwick-Yanai, G., Lipson, D., & Satija, R. bioRxiv 2024 5. Walton, R. T., Qin, Y., & Blainey, P. C. bioRxiv 2024 6. Chardon, F. M., McDiarmid, T. A., Page, N. F., Martin, B., Domcke, S., Regalado, S. G., Lalanne, J.-B., Calderon, D., Starita, L. M., Sanders, S. J., Ahituv, N., & Shendure, J. bioRxiv 2023 7. Dixit, Atray, Oren Parnas, Biyu Li, Jenny Chen, Charles P. Fulco, Livnat Jerby-Arnon, Nemanja D. Marjanovic, et al. 2016. “Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens.” Cell 167 (7): 1853– 66.e17.8. Replogle, Joseph M., Reuben A. Saunders, Angela N. Pogson, Jeffrey A. Hussmann, Alexander Lenail, Alina Guna, Lauren Mascibroda, et al. 2022. “Mapping Information- Rich Genotype-Phenotype Landscapes with Genome-Scale Perturb-Seq.” Cell, June. https: / / doi.org / 10.1016 / j.cell.2022.05.013. 9. Rosenberg, Alexander B., Charles M. Roco, Richard A. Muscat, Anna Kuchina, Paul Sample, Zizhen Yao, Lucas T. Graybuck, et al. 2018. “Single-Cell Profiling of the Developing Mouse Brain and Spinal Cord with Split-Pool Barcoding.” Science (New York, N.Y.) 360 (6385): 176–82. EXEMPLARY EMBODIMENTS
[0152] Exemplary embodiments provided in accordance with the presently disclosed subject matter include, but are not limited to, the claims and the following embodiments:
[0153] Embodiment 1. A method of detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell library, the method comprising: i) converting each RNA of interest into a cDNA through reverse transcription; ii) labeling each cDNA with a 5’-phosphate; iii) circularizing each 5’-phosphate labeled cDNA using a ligase; and iv) amplifying each circularized cDNA using at least two primers specific to the RNAs / cDNAs of interest, thereby detecting or enriching one or more RNAs of interest in the cell.
[0154] Embodiment 2. The method of embodiment 1, wherein the single-cell library is a single-cell RNA library.
[0155] Embodiment 3. The method of embodiment 1 or 2, wherein the cDNA is labeled with the 5’-phosphate using a 5’-phosphorylated probe, a kinase, or a polymerase chain reaction (PCR) with a 5’-phosphorylated primer.
[0156] Embodiment 4. The method of any one of embodiments 1-3, wherein the RNAs of interest are one or more guide RNAs (gRNAs).
[0157] Embodiment 5. The method of embodiment 4, wherein the two primers comprise a P5-gRNA primer and a P7-gRNA primer.
[0158] Embodiment 6. The method of embodiment 5, wherein the P5-gRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 15.
[0159] Embodiment 7. The method of embodiment 5, wherein the P7-gRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 16.
[0160] Embodiment 8. The method of any one of embodiments 1-3, wherein the RNAs of interests are one or more mRNAs.
[0161] Embodiment 9. The method of embodiment 8, wherein the two primers comprise a P5-mRNA primer and a P7-mRNA primer.
[0162] Embodiment 10. The method of embodiment 9, wherein the P5-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 1.
[0163] Embodiment 11. The method of embodiment 9, wherein the P7-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 2.
[0164] Embodiment 12. The method of any one of embodiments 1-11, wherein the ligase is a ssDNA ligase.
[0165] Embodiment 13. The method of any one of embodiments 1-12, wherein step i) comprises converting all the RNAs in the cell into cDNAs through reverse transcription.
[0166] Embodiment 14. The method of any one of embodiments 1-13, wherein prior step ii), further comprising barcoding each cDNA.
[0167] Embodiment 15. The method of embodiment 14, wherein each cDNA is barcoded through a “split-pool” approach.
[0168] Embodiment 16. The method of embodiment 15, wherein the “split-pool” approach comprises multiple “split and pool” cycles, wherein each “split and pool” cycle comprises a “split” step that the plurality of cells is distributed individually to attach a barcode followed by a “pool” step that the cells are mixed together, wherein the “split-pool” approach comprises at least two “split and pool” cycles.
[0169] Embodiment 17. The method of embodiment 16, wherein the “split-pool” approach comprises three “split and pool” cycles.
[0170] Embodiment 18. The method of embodiment 17, wherein the first cycle comprises adding a Barcode 1 (BC1) at the 3’ end of each cDNAs, wherein the second cycle comprisesadding a Barcode 2 (BC2) at the 3’ end of the BC1, and wherein the third cycle comprises adding a Barcode 3 (BC3) at the 3’ end of the BC2.
[0171] Embodiment 19. The method of embodiment 18, wherein the BC1 comprises a sequence having at least 80% identity to a sequence selected from Table 10.
[0172] Embodiment 20. The method of embodiment 18, wherein the BC2 comprises a sequence having at least 80% identity to a sequence selected from Table 11.
[0173] Embodiment 21. The method of embodiment 18, wherein the BC3 comprises a sequence having at least 80% identity to a sequence selected from Table 12.
[0174] Embodiment 22. The method of any one of embodiments 18-21, wherein the BC1 and the BC2 are linked via a linker comprising a sequence of SEQ ID NO: 6.
[0175] Embodiment 23. The method of any one of embodiments 18-22, wherein the BC2 and the BC3 are linked via a linker comprising a sequence of SEQ ID NO: 7.
[0176] Embodiment 24. The method of any one of embodiments 1-23, wherein prior step iv), further comprising linearizing the circularized cDNAs using a restriction enzyme.
[0177] Embodiment 25. The method of embodiment 24, wherein the restriction enzyme is a DraI enzyme.
[0178] Embodiment 26. The method of any one of embodiments 1-25, wherein after step iv), further comprising sequencing the enriched RNAs of interest using one or more sequencing primers.
[0179] Embodiment 27. The method of embodiment 26, wherein the one or more sequencing primers are Illumina sequencing primers.
[0180] Embodiment 28. The method of embodiment 26, wherein the one or more sequencing primers are Ultima sequencing primers.
[0181] Embodiment 29. A kit for detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell RNA library via circular capture, comprising: a) a ligase; and b) a 5’ primer and a 3’ primer to amplify one or more RNAs of interest.
[0182] Embodiment 30. The kit of embodiment 29, further comprising c) one or more barcode oligonucleotides.
[0183] Embodiment 31. The kit of embodiment 30, wherein the one or more barcode oligonucleotides comprise a sequence having at least 80% identity to a sequence from Tables 4-9.
[0184] Embodiment 32. The kit of any one of embodiments 29-31, wherein the ligase is a ssDNA ligase.
[0185] Embodiment 33. The kit of any one of embodiments 29-32, wherein the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 15 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 16.
[0186] Embodiment 34. The kit of any one of embodiments 29-32, wherein the one or more RNAs of interest are mRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 1 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 2.
[0187] Embodiment 35. The kit of any one of embodiments 29-32, wherein the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 987 and the 3’ primer comprises a sequence of having at least 80% identity to SEQ ID NO: 988.
[0188] Although the foregoing disclosure has been described in some detail by way of illustration and example for purpose of clarity of understanding, one of skill in the art will appreciate that certain changes and modifications within the spirit and scope of the disclosure may be practiced, e.g., within the scope of the appended claims. It should also be understood that aspects of the disclosure and portions of various recited embodiments and features can be combined or interchanged either in whole or in part. In the foregoing descriptions of the various embodiments, those embodiments which refer to another embodiment may be appropriately combined with other embodiments as will be appreciated by one of skill in the art. Furthermore, those of ordinary skill in the art will appreciate that the foregoing description is by way of example only, and is not intended to limit the disclosure. In addition, each reference provided herein is incorporated by reference in its entirety for all purposes to the same extent as if each reference was individually incorporated by reference.
Claims
WHAT IS CLAIMED IS:
1. A method of detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell library, the method comprising: i) converting each RNA of interest into a cDNA through reverse transcription; ii) labeling each cDNA with a 5’-phosphate; iii) circularizing each 5’-phosphate labeled cDNA using a ligase; and iv) amplifying each circularized cDNA using at least two primers specific to the RNAs / cDNAs of interest, thereby detecting or enriching one or more RNAs of interest in the cell.
2. The method of claim 1, wherein the single-cell library is a single-cell RNA library.
3. The method of claim 1 or 2, wherein the cDNA is labeled with the 5’- phosphate using a 5’-phosphorylated probe, a kinase, or a polymerase chain reaction (PCR) with a 5’-phosphorylated primer.
4. The method of claim 1, wherein the RNAs of interest are one or more guide RNAs (gRNAs).
5. The method of claim 4, wherein the two primers comprise a P5-gRNA primer and a P7-gRNA primer.
6. The method of claim 5, wherein the P5-gRNA primer comprises a sequence having at least 80% identity to SEQ ID NO:
15.
7. The method of claim 5, wherein the P7-gRNA primer comprises a sequence having at least 80% identity to SEQ ID NO:
16.
8. The method of claim 1, wherein the RNAs of interests are one or more mRNAs.
9. The method of claim 8, wherein the two primers comprise a P5-mRNA primer and a P7-mRNA primer.
10. The method of claim 9, wherein the P5-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO: 1.
11. The method of claim 9, wherein the P7-mRNA primer comprises a sequence having at least 80% identity to SEQ ID NO:
2.
12. The method of claim 1, wherein the ligase is a ssDNA ligase.
13. The method of claim 1, wherein step i) comprises converting all the RNAs in the cell into cDNAs through reverse transcription.
14. The method of claim 1, wherein prior step ii), further comprising barcoding each cDNA.
15. The method of claim 14, wherein each cDNA is barcoded through a “split-pool” approach.
16. The method of claim 15, wherein the “split-pool” approach comprises multiple “split and pool” cycles, wherein each “split and pool” cycle comprises a “split” step that the plurality of cells is distributed individually to attach a barcode followed by a “pool” step that the cells are mixed together, wherein the “split-pool” approach comprises at least two “split and pool” cycles.
17. The method of claim 16, wherein the “split-pool” approach comprises three “split and pool” cycles.
18. The method of claim 17, wherein the first cycle comprises adding a Barcode 1 (BC1) at the 3’ end of each cDNAs, wherein the second cycle comprises adding a Barcode 2 (BC2) at the 3’ end of the BC1, and wherein the third cycle comprises adding a Barcode 3 (BC3) at the 3’ end of the BC2.
19. The method of claim 18, wherein the BC1 comprises a sequence having at least 80% identity to a sequence selected from Table 10.
20. The method of claim 18, wherein the BC2 comprises a sequence having at least 80% identity to a sequence selected from Table 11.
21. The method of claim 18, wherein the BC3 comprises a sequence having at least 80% identity to a sequence selected from Table 12.The method of claim 18, wherein the BC1 and the BC2 are linked via a linker comprising a sequence of SEQ ID NO:
6.
23. The method of claim 18, wherein the BC2 and the BC3 are linked via a linker comprising a sequence of SEQ ID NO:
7.
24. The method of claim 1, wherein prior step iv), further comprising linearizing the circularized cDNAs using a restriction enzyme.
25. The method of claim 24, wherein the restriction enzyme is a DraI enzyme.
26. The method of claim 1, wherein after step iv), further comprising sequencing the enriched RNAs of interest using one or more sequencing primers.
27. The method of claim 26, wherein the one or more sequencing primers are Illumina sequencing primers.
28. The method of claim 26, wherein the one or more sequencing primers are Ultima sequencing primers.
29. A kit for detecting or enriching one or more RNAs of interest in a cell for preparing a single-cell RNA library via circular capture, comprising: a) a ligase; and b) a 5’ primer and a 3’ primer to amplify one or more RNAs of interest.
30. The kit of claim 29, further comprising c) one or more barcode oligonucleotides.
31. The kit of claim 30, wherein the one or more barcode oligonucleotides comprise a sequence having at least 80% identity to a sequence from Tables 4-9.
32. The kit of claim 29, wherein the ligase is a ssDNA ligase.
33. The kit of claim 29, wherein the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQID NO: 15 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO:
16.
34. The kit of claim 29, wherein the one or more RNAs of interest are mRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 1 and the 3’ primer comprises a sequence having at least 80% identity to SEQ ID NO:
2.
35. The kit of claim 29, wherein the one or more RNAs of interest are gRNAs, and wherein the 5’ primer comprises a sequence having at least 80% identity to SEQ ID NO: 987 and the 3’ primer comprises a sequence of having at least 80% identity to SEQ ID NO: 988.
Citation Information
Patent Citations
COMPOSITIONS AND METHODS FOR MAKING cDNA LIBRARIES FROM SMALL RNAs
US20150087556A1
Methods and compositions for detection of small rnas
US20170159106A1
Digital counting of individual molecules by stochastic attachment of diverse labels
US20200354788A1
High-throughput single-cell transcriptome libraries and methods of making and of using
US20210102194A1
Imaging system hardware
US20220241780A1