Whole-cell transcriptome analysis in single cells

The split-pooling method with reverse transcriptase and barcode assembly addresses inefficiencies in single-cell analysis by enabling efficient detection of nucleic acid targets in individual cells, enhancing whole-transcriptome analysis without cell isolation.

JP7856650B2Active Publication Date: 2026-05-11F HOFFMANN LA ROCHE & CO AG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
F HOFFMANN LA ROCHE & CO AG
Filing Date
2021-12-01
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing methods for single-cell analysis, particularly whole-transcriptome analysis, require physical isolation of cells and are inefficient for detecting rare nucleic acid targets, necessitating a more robust method for barcoding and detection.

Method used

A method involving split-pooling and reverse transcriptase to assemble compound barcodes onto nucleic acid targets in individual cells, using oligonucleotide primers and barcode subunits, with polymerases having terminal transferase activity to extend primers and add non-template nucleotides, allowing for unique cell-specific barcoding without cell isolation.

Benefits of technology

Enables efficient and robust detection of multiple nucleic acid targets in individual cells without isolation, improving the accuracy and efficiency of whole-transcriptome analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007856650000008
    Figure 0007856650000008
  • Figure 0007856650000009
    Figure 0007856650000009
  • Figure 0007856650000010
    Figure 0007856650000010
Patent Text Reader

Abstract

The present invention is a method for single-cell transcriptome analysis, which involves detecting multiple transcripts in each individual cell of a plurality of cells by barcoding the transcripts with cell-specific compound barcodes formed using DNA polymerase and terminal transferase, optionally in a single enzyme such as reverse transcriptase.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention generally relates to single-cell analysis. More specifically, it relates to detecting multiple nucleic acid targets in individual cells without the need to isolate or separate individual cells. [Background technology]

[0002] Single-cell analysis is gaining increasing importance in understanding biology and disease. Gene expression studies, including whole-transcriptome analysis, reveal variations in gene expression between cells. Previous methods required the physical isolation of each cell under a microscope (see Tang et al. (2009) mRNA-Seq, whole-transcriptome analysis of a single cell, Nature Methods 6:377). U.S. Patent No. 10,392,662 describes a technical solution involving encapsulating individual cells in droplets and treating a water-oil emulsion so that multiple individual reactions can occur simultaneously. An alternative, simpler, and more sophisticated method is described in U.S. Patent No. 10,144,950. This novel approach, called "Quantum Barcoding" or "QBC," does not require the separation of individual cells in droplets or by any other means. Instead, QBC involves passing multiple whole cells through a series of split-pool rounds, as a result of which a unique compound barcode can be assembled on each cell. In each round, a solution containing multiple cells is divided into several reaction volumes, each containing a barcode sequence. After the barcodes bind to the target of each cell, the reaction volumes are pooled and divided again, resulting in each cell receiving a barcode for the next round. Each cell follows a unique pathway through a series of barcode-containing wells to acquire a unique cell-associated compound barcode. The QBC method allows these unique cell-associated compound barcodes to attach to any target of interest within the cell, i.e., DNA, RNA, protein, or other cellular target.

[0003] The assembly of compound barcodes on nucleic acids typically involves a ligation step; see Rosenberg, et al. (2018) Single cell profiling of the developing mouse brain and spinal cord by split-pool barcoding, Science, 360:176. Ligation requires additional reagents and reaction conditions and is less efficient than primer extension. Low efficiency is unacceptable in applications where nucleic acid targets are rare, such as whole transcriptome analysis. There is an unmet need for a more robust method for barcoding and detecting nucleic acid targets in individual cells. [Overview of the project]

[0004] The present invention includes a method for assembling compound barcodes onto nucleic acid targets in individual cells. Each of multiple targets in an individual cell is to be labeled with the same cell-specific barcode. The barcodes are assembled from barcode subunits via a split-pooling process. The barcode subunits are copied and the compound barcodes are assembled by utilizing the unique properties of reverse transcriptase.

[0005] In one embodiment, the present invention is a method for detecting multiple target nucleic acids in multiple cells, comprising: contacting multiple cells in a sample with oligonucleotide primers for each target nucleic acid in the presence of nucleic acid polymerase having terminal transferase activity; extending the oligonucleotide primers to form copy strands having one or more non-template nucleotides at the 3' end of the copy strand; and distributing the sample into a first reaction volume set, each volume containing the first barcode subunit and the 3' end of the copy strand to copy the first barcode subunit. The method comprises forming a cell-characterizing compound barcode on a copy strand by sequentially combining a first reaction volume set into a pool and distributing a nucleic acid polymerase for terminal extension, each volume comprising a second barcode subunit and a nucleic acid polymerase for further extending the 3' end of the copy strand to copy the second barcode subunit, through a split-pool process comprising one or more rounds of the distributing process, and determining the sequence of the extended copy strand containing the cell-characterizing compound barcode, thereby detecting multiple target nucleic acids in multiple cells. In some embodiments, the method further comprises a step of amplifying the extended copy strand containing the cell-characterizing compound barcode before sequencing.

[0006] In some embodiments, the oligonucleotide primer includes a barcode and / or a universal amplification primer binding site. In some embodiments, the barcode subunit to be copied includes a universal amplification primer binding site.

[0007] In some embodiments, the target nucleic acid is DNA. In some embodiments, the target nucleic acid is RNA, such as messenger RNA, and the oligonucleotide primer includes a poly-dT sequence or a target-specific sequence. In some embodiments, the oligonucleotide primer includes a barcode.

[0008] In some embodiments, the nucleic acid polymerase is a reverse transcriptase. RT may have reduced RNaseH activity.

[0009] In some embodiments, the non-template nucleotide is deoxycytosine.

[0010] In some embodiments, the barcode subunit includes a moiety complementary to one or more non-template nucleotides. In some embodiments, the barcode subunit includes one or more modified nucleotides, such as an isonucleotide located at the 5' end of the barcode subunit, which reduces the stability of subunit hybridization.

[0011] In some embodiments, the present invention is a kit for detecting multiple target nucleic acids in multiple cells, the kit comprising oligonucleotide primers, reverse transcriptase, and multiple barcode subunits, each barcode subunit comprising a poly-dG sequence. The barcode subunits may be present in multiple separate reaction volumes. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is a schematic diagram of the barcoding workflow. [Figure 2] Figure 2 is a schematic diagram of another embodiment of the barcoding workflow. [Figure 3] Figure 3 is a schematic diagram of yet another embodiment of the barcode workflow. [Modes for carrying out the invention]

[0013] definition The following definitions will help you understand this disclosure.

[0014] The term "adaptor" refers to a nucleotide sequence that can be added to another sequence to impart additional properties to that sequence. Adapters may be single-stranded, double-stranded, or have both single-stranded and double-stranded portions.

[0015] The term "barcode" refers to a nucleotide sequence that confers identity to molecules or groups of molecules that share common characteristics or origins. A barcode can confer unique identity to individual molecules (and their copies). Such a barcode is a unique ID (UID) or unique molecular identifier (UMI). A barcode can confer identity to an entire group of molecules (and their copies) originating from the same source (e.g., a sample). This barcode is a multiplex ID (MID) or sample ID (SID). For nucleic acid molecules to be identified, the barcode does not need to be unique. Since nucleic acids are also identified by their sequence, two nucleic acids with different sequences can share the same barcode, and such a barcode functions as a unique molecular barcode (UID or UMI).

[0016] The term "compound barcode" refers to a barcode assembled from barcode subunits. Each compound barcode is unique, however, two or more compound barcodes may share one or more barcode subunits.

[0017] The term "nucleic acid" refers to polymers of nucleotides (e.g., ribonucleotides and deoxyribonucleotides, both natural and unnatural), including DNA, RNA, and their subcategories such as cDNA and mRNA. Nucleic acids can be single-stranded or double-stranded and generally contain a 5'-3' phosphodiester bond, although in some cases nucleotide analogs may have other bonds available in the art, as well as linkers, spacers, and labels. Nucleic acids may contain natural bases (adenosine, guanosine, cytosine, uracil, and thymidine) and unnatural bases. Some examples of unnatural bases include, for example, those described in Seela et al., (1999) Helv. Chim. Acta 82:1640. Unnatural bases may have specific functions, such as increasing the stability of nucleic acid double helixes, inhibiting nuclease digestion, or blocking primer elongation or chain polymerization.

[0018] The term "DNA polymerase" refers to an enzyme that performs template-directed synthesis of polynucleotides from deoxyribonucleotides. Examples of DNA polymerases include Pol I, Pol II, Pol III, Pol IV, and Pol V in prokaryotes, DNA polymerases in eukaryotes, and DNA polymerases, telomerases, and reverse transcriptases in archaea. The term "thermally stable polymerase" refers to an enzyme that is stable and heat-resistant, and that, when exposed to high temperatures for the time required to denature double-stranded nucleic acids, retains sufficient activity to subsequently carry out the polynucleotide elongation reaction without being irreversibly denatured (inactivated). In some embodiments, the following thermostable polymerases can be used: Thermococcus litoralis (Vent, GenBank: AAA72101), Pyrococcus furiosus (Pfu, GenBank: D12983, BAA02362), Pyrococcus woesii, Pyrococcus GB-D (Deep Vent, GenBank: AAA67131), Thermococcus kodakaraensis KODI (KOD, GenBank: BD175553, BAA06142, Thermococcus species KOD strain (Pfx, GenBank: AAE68738)), Thermococcus gorgonarius (Thermococcus gorgonarius) (Tgo, Pdb:4699806), Sulfolobus solataricus (GenBank:NC002754, P26811), Aeropyrum pernix (GenBank:BAA81109), Archaeglobus fulgidus (GenBank:029753), Pyrobaculum aerophilum (GenBank:AAL63952), Pyrodictium occultum* occultum* (GenBank:BAA07579, BAA07580), * Thermococcus 9°Nm* (GenBank:AAA88769, Q56366), * Thermococcus fumicolans* (GenBank:CAA93738, P74918), * Thermococcus hydrothermalis* (GenBank:CAC18555), * Thermococcus GE8* (GenBank:CAC12850), * Thermococcus JDF-3* (GenBank:AX135456, WO0132887), * Thermococcus TY* (GenBank:CAA73475), * Pyrococcus abyssi* abyssi) (GenBank:P77916), Pyrococcus glycovorans (GenBank:CAC12849), Pyrococcus horikoshii (GenBank:NP 143776), Pyrococcus species GE23 (GenBank:CAA90887), Pyrococcus species ST700 (GenBank:CAC 12847), Thermococcus pacificus (GenBank:AX411312.1), Thermococcus zilligii (GenBank:DQ3366890), Thermococcus aggregans, Thermococcus barosii (Thermococcus Thermococcus barossii), Thermococcus celer (GenBank:DD259850.1), Thermococcus profundus (GenBank:E14137), Thermococcus siculi (GenBank:DD259857.1), Thermococcus thioreducens, Thermococcus onnulineus (Thermococcusonnurineus)NA1, Sulfolobus acidocaldarium, Sulfolobus tokodaii, Pyrobaculum calidifontis, Pyrobaculum islandicum (GenBank:AAF27815), Methanococcus yannaskii B polymerases of *Desulfurococcus jannaschii* (GenBank:Q58295), *Desulfurococcus* species TOK, genera *Desulfurococcus*, *Pyrolobus*, *Pyrodictium*, *Staphylothermus*, *Vulcanisaetta*, *Methanococcus* (GenBank:P52025), and other individual bacteria, e.g., GenBank AAC62712, P956901, BAAA07579), thermophilic bacteria Thermus species (e.g., flavus, ruber, thermophilus, lacteus, rubens, aquaticus), Bacillus stearothermophilus, Thermotoga maritima, Methanothermus fervidus, KOD polymerase, TNA1 polymerase, Thermococcus species 9°N-7, T4, T7, phi29, Pyrococcus friosusPolymerases of Thermococcus species 9N-7, including *T. furiosus*, *P. abyssi*, *T. goorgonarius*, *T. litoralis*, *T. zilligii*, *T. GT*, *P. GB-D*, *KOD*, *Pfu*, *T. goorgonarius*, *T. zilligii*, *T. litoralis*, and *Thermococcus* species 9N-7. In some cases, nucleic acid (e.g., DNA or RNA) polymerases may be modified, naturally occurring type A polymerases. Further embodiments of the present invention generally relate to a method by which a modified type A polymerase in a primer extension, terminal modification (e.g., terminal transferase, degradation or polishing), or amplification reaction can be selected from any species of the genera Meiothermus, Thermotoga, or Thermomicrobium. Another embodiment of the present invention generally relates to a method by which a polymerase can be isolated from any of Thermus aquaticus (Taq), Thermus thermophilus, Thermus caldophilus, or Thermus filiformis in a primer extension, terminal modification (e.g., terminal transferase, degradation or polishing), or amplification reaction. Further embodiments of the present invention generally involve modified type A polymerases in reactions such as primer extension, terminal modification (e.g., terminal transferase, degradation or polishing) or amplification reactions, such as those involving Bacillus stearothermophilus, Sphaerobacter thermophilus, Dictoglomus thermophilum, or Escherichia coli.The present invention encompasses methods for isolating a mutant Taq-E507K polymerase from coli. In another embodiment, the present invention generally relates to methods for which a modified type A polymerase in a reaction such as primer extension, terminal modification (e.g., terminal transferase, degradation, or polishing) or amplification reaction may be a mutant Taq-E507K polymerase. Another embodiment of the present invention generally relates to methods for amplifying a target nucleic acid using a thermostable polymerase.

[0019] The term "reverse transcriptase" refers to an RNA-dependent DNA polymerase that synthesizes cDNA from an RNA template in the presence of appropriate primers. Wild-type reverse transcriptase has RNaseH activity, which degrades RNA within the RNA-DNA hybrid. Some wild-type and engineered reverse transcriptases have reduced RNaseH activity. While wild-type RNaseH+RT can be used, enzymes with reduced RNaseH activity are particularly suitable for the methods of the present invention. Reverse transcriptase also has template switch activity, which requires a template switch oligonucleotide (TSO), and when it reaches the end of the template, the reverse transcriptase switches from a copy of the template to a copy of the TSO.

[0020] The terms "polynucleotide" and "oligonucleotide" are used interchangeably. A polynucleotide is a single-stranded or double-stranded nucleic acid. Oligonucleotide is a term sometimes used to describe shorter polynucleotides. Oligonucleotides may consist of at least 6 nucleotides or about 15 to 30 nucleotides. Oligonucleotides are prepared by any suitable method known in the art, including direct chemical synthesis as described, for example, Narang et al. (1979) Meth. Enzymol. 68:90-99; Brown et al. (1979) Meth. Enzymol. 68:109-151; Beaucage et al. (1981) Tetrahedron Lett. 22:1859-1862; Matteucci et al. (1981) J. Am. Chem. Soc. 103:3185-3191.

[0021] The term "primer" refers to a single-stranded oligonucleotide that hybridizes to a sequence in a target nucleic acid and can act as a starting point for synthesis along the complementary strand of the nucleic acid under conditions suitable for such synthesis. A primer can be partially or fully complementary to the target nucleic acid as long as it forms a stable hybrid with the target and can be extended by a nucleic acid polymerase. The term "forward and reverse primers" refers to a primer pair that is complementary to opposite strands of the target nucleic acid at sites adjacent to the target sequence. Forward and reverse primers can exponentially amplify the target by polymerase chain reaction (PCR).

[0022] The term "sample" refers to any composition that contains or is presumed to contain a target nucleic acid. This term includes samples of tissue or fluid isolated from an individual, such as skin, plasma, serum, cerebrospinal fluid, lymph, synovial fluid, urine, tears, blood cells, organs, and tumors, as well as samples of in vitro cultures established from cells obtained from an individual that include formalin-fixed paraffin-embedded tissue (FFPET) and nucleic acids isolated therefrom. Samples may also include cell-free materials such as cell-free blood fractions containing cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). In some embodiments, as will be apparent to those skilled in the art from the context, the term "sample" refers to a preparation obtained by processing a primary sample obtained from a patient (e.g., by removing or adding one or more components). For example, such processing can include removing some tissue material (including blood components) from the sample and lysing any intact cells to release nucleic acids.

[0023] The term "sequencing" refers to any method for determining the sequence of nucleotides in a target nucleic acid.

[0024] The term "solid support" refers to any solid material that can interact with the capture portion. A solid support may be a solution-phase support (e.g., glass beads, magnetic beads, or other similar particles) or a solid-phase support (e.g., a silicon wafer, glass slide, etc.) that can be suspended in a solution. Examples of solution-phase supports include superparamagnetic spherical polymer particles such as DYNABEADS magnetic beads or Beckman Coulter AMPure Solid-Phase Reversible Immobilization (SPRI) paramagnetic beads (ThermoFisher Scientific, Waltham, Massachusetts), or magnetic glass particles as described in U.S. Patents 6,274386, 7371830, 6870047, 6255477, 6746874, and 6258531.

[0025] The terms “target sequence,” “target nucleic acid,” or “target” refer to a portion of a nucleic acid sequence in a sample to be detected or analyzed. The term “target” includes all variants of the target sequence, e.g., one or more mutant variants and wild-type variants.

[0026] The terms "universal primer" and "universal priming site" refer to primers and priming sites that do not naturally exist in the target sequence. Typically, universal priming sites are located at the adapter or tail end of a target-specific primer. Universal primers can bind to universal priming sites and direct primer extension from there.

[0027] Whole-transcriptome analysis of single cells is a useful new tool in developmental biology and disease research. U.S. Patent No. 10,144,950 describes a method for labeling multiple targets, including multiple nucleic acid targets (DNA and RNA) in individual cells, with unique cell-associated compound barcodes (referred to as “quantum barcoding” or “QBC”). Compound barcodes may be formed from nucleic acid subunits, such as oligonucleotides, that contain or consist of nucleic acid barcodes. These barcode subunits are linked together, for example, for reading by sequencing. One embodiment of QBC (U.S. Patent No. 10,144,950) involves assembling compound barcodes from oligonucleotide subunits on nucleic acid targets. One method of linking oligonucleotides is by ligation. Alternatively, polymerase can copy barcode subunits annealed to an antibody-bound template, as described in U.S. Patent Application No. 16 / 250,974, “Identifying Multiple Epitopes in Cells,” published January 17, 2019. This invention provides an improved method for nucleic acid barcoding using QBC. Specifically, this method improves existing QBC technology by utilizing the same polymerase enzyme for target enrichment and target barcoding when analyzing multiple nucleic acid targets in individual cells.

[0028] This invention provides a novel method for analyzing gene expression (including the whole transcriptome) in individual cells without the need to separate or encapsulate them. This invention simplifies and improves upon existing methods for single-cell whole transcriptome analysis.

[0029] The sample used in the method of the present invention includes any individual (e.g., a non-human mammal, a human subject, or a patient). The sample may also be an environmental sample or a plant sample containing cells. The sample may consist of any tissue or cell-containing fluid, including a cell culture. For example, the sample may be an organ (e.g., a lymph node) or tumor biopsy, or a blood or plasma sample. The sample may also consist of a subset of tissue-derived cells, such as immune cells isolated from a tumor, tumor-infiltrating lymphocytes (TILs). In some embodiments, the sample is a formalin-fixed and paraffin-embedded (FFPE) sample.

[0030] In some embodiments, the sample includes a cell compartment or intracellular compartment containing the target nucleic acid. In such embodiments, the method may include a step of permeabilizing the cell or intracellular compartment to allow access to the target nucleic acid. The cells may be fixed in methanol buffer and resuspended in a suitable buffer containing a nuclease inhibitor (e.g., an RNase inhibitor). See Chen et al., (2018) PBMC Fixation and Processing for Chromium Single-Cell RNA Sequencing, J.Transl Med. 16:198.

[0031] The primers used in this method may include target-specific sequences. Target-specific sequences may be gene-specific sequences, motif-specific sequences (e.g., kinase domain-specific sequences, RAS family-specific sequences, trinucleotide repeat sequences, etc.), or random sequences (e.g., random hexamer sequences). Target-specific sequences may also be poly-dT sequences (e.g., dT 12-18 ) may be. In some embodiments, the primer is a mixture of a random primer and a poly-dT primer.

[0032] Preferably, the target-specific sequence is located in the 3' portion of the primer. In addition to the target-specific region, the primer may include additional sequences. Preferably, these additional sequences are located at the 5' end of the target-specific region. In other embodiments, additional sequences may be included elsewhere in the primer, as long as the target-specific region hybridizes to the target and drives the primer extension reaction as described below. Additional sequences in the primer may include one or more barcode sequences, such as a unique molecular identification sequence (UID) or a multiplex sample identification sequence (MID). The barcode sequences may exist as a single sequence or as two or more sequences.

[0033] In some embodiments, additional sequences include sequences that facilitate ligation to the 5' end of the primer. The primer may include a universal ligation sequence that enables ligation of the adapter.

[0034] In some embodiments, the additional sequence includes one or more binding sites for one or more universal amplification primers. In some embodiments, the primer includes a universal capture sequence that allows capture of the primer and primer extension product by hybridization to a capture oligonucleotide.

[0035] The present invention provides a method for simultaneously evaluating one or more target nucleic acid sequences in multiple cells. The target sequences may include clinically relevant biomarkers. The present invention also includes a method for simultaneously evaluating one or more mRNA transcripts in multiple cells. In this embodiment, the target includes a gene (or gene fragment) whose expression level is a biomarker for a disease or symptom.

[0036] In some embodiments, the target corresponds to the type of cells used in the method. For example, when applying to multiple immune cells, genes characteristic of different cell types and states are used. In this embodiment, the target includes one or more RNA transcripts from CD45, CD3, CD8, CD39, CD25, IL-7R, CD4, CXCR3, CCR6, CD3G, CD3D, CD3E, CD2, CD8A, GZMA FOXP3, CD19, CD79A, PDCD1, HAVCR2, IFNG, TNF, ITGAE, and CXCR6.

[0037] In some embodiments, the method is applied to multiple tumor cells. In this embodiment, the target includes one or more RNA transcripts of one or more fusion genes common in cancer, such as ALK, NTRK1, FGFR2, FGFR3, RET, ROS1, and FIP1L1-PDGFRA. The target may also include RNA transcripts of genes that are commonly mutated in cancer. The target may include one or more genes listed in Tables 3(a) to (c) below. [Table 1] [Table 2] [Table 3-1] [Table 3-2]

[0038] The primer extension step is carried out by a nucleic acid polymerase. Depending on the type of nucleic acid being analyzed, the polymerase may be a DNA-dependent DNA polymerase ("DNA polymerase") or an RNA-dependent DNA polymerase ("reverse transcriptase"). In some embodiments, a mixture of two or more polymerases is used to provide the desired combination of enzyme activity.

[0039] In some embodiments, the polymerase is a reverse transcriptase with reduced RNaseH activity (e.g., Moloney's mouse leukemia virus (MMLV) RT). In some embodiments, the reverse transcriptase is active at high temperatures (up to 55°C). In some embodiments, the polymerase has terminal nucleotide transferase (terminal transferase) activity. Reverse transcriptase (e.g., MMLV RT) naturally possesses this activity. In some embodiments, in the presence of manganese (Mn) ions and dCTP, MMLV or MMLV-derived RT adds a dC nucleotide stretch to the 3' end of the copy strand. Under other conditions, MMLV or MMLV-derived RT adds a combination of dC and dA nucleotides.

[0040] In some embodiments, polymerases possess terminal transferase activity and template switch activity. For example, reverse transcriptase, e.g., MMLV or MMLV-derived RT during first-chain synthesis, exhibits innate terminal transferase activity upon reaching the 5' end of the copy strand. The added non-template nucleotide stretch acts as an anchoring site for the template switch oligo (TSO). When the TSO hybridizes to the non-template nucleotide stretch, the polymerase switches the strand from a copy of the target (e.g., RNA) to a copy of the TSO.

[0041] In some embodiments, the DNA polymerase is a type A DNA polymerase (DNA-dependent DNA polymerase). Some DNA polymerases have limited terminal transferase activity (Taq polymerases that add a single dA to the 3' end of the copy strand). Other DNA polymerases do not have detectable terminal transferase activity. In such embodiments, separate terminal transferase enzymes are used to add non-template nucleotides to the 3' end of the copy strand.

[0042] In some embodiments, the DNA polymerase is a hot-start polymerase or a similarly conditionally activated polymerase. For the amplification step, a heat-stable DNA polymerase is used, for example, the polymerase is Taq or a Taq-derived polymerase (e.g., KAPA 2G polymerase from KAPA Biosystems, Wilmington, Massachusetts).

[0043] The present invention includes a method for assembling compound barcodes from barcode subunits. A barcode subunit is an oligonucleotide containing a nucleic acid barcode. When joined together, the barcode subunits form an ordered combination called a compound barcode. The compound barcode provides information necessary to identify a tagged entity, such as a cell within a group of cells. A barcode subunit may further contain a nucleic acid sequence necessary to form a hybrid with a 3' portion of the copy strand sufficient to allow copying of the barcode subunit by nucleic acid polymerase. A barcode subunit may contain a sequence complementary to the non-template nucleotide at the 3' end of the copy strand. In some embodiments, the barcode subunit contains a polydG sequence at its 3' end.

[0044] Each barcode subunit can have 2 to 50 nucleotides. Each barcode subunit may contain a predetermined sequence, a random sequence, or a combination thereof (see Figures 1, 2, and 3). The use of random sequence barcodes (UMI, Figures 1, 2, and 3) can increase the diversity of the additional barcodes. For predetermined sequences, sets of barcodes can be designed for experimental purposes.

[0045] In the workflows described below (Figures 1, 2, and 3), two or more barcoded entities (e.g., cells) may share one or more barcode subunits at the same or different positions within the compound barcode. However, each barcoded entity has a unique ordered combination of barcoded subunits that forms its own compound barcode.

[0046] In some embodiments, to enhance the stability of the hybrid formed between the copy strand and the barcode subunit, the subunit contains modified nucleotides that increase the stability of the hybrid, as revealed by a higher melting temperature (Tm). To increase Tm, the following bases may be used instead of conventional bases. [Table 4]

[0047] In some embodiments, the barcode subunit includes a modified nucleotide that inhibits the annealing of a second subunit in one round of assembly. In some embodiments, the modified nucleotide is one or more 5'-nucleotide isomers (isodC or isodG).

[0048] The present invention provides a method comprising a sufficient number of split pool rounds to generate enough unique compound barcodes to label each cell in a plurality of cells with a unique barcode. An exemplary calculation provided in U.S. Patent No. 10144950 can be summarized as follows: The number of unique compound barcodes (B) can be determined according to the following formula: B = ln(1-C) / ln(1-1 / N), ≡≡≡(in the equation, B is the number of compound barcodes, C is the certainty of barcode over-presentation to cells, N is the number of cells.

[0049] For example, N=10 6 Starting with a 1 / 10 probability that one cell (1 million) and two cells have the same barcode, the certainty of over-presentation C = 0.9999999. The required number of unique compound barcodes (BC) is as follows: B = ln(0.0000001) / (1-10) -6 )≡16×10 6

[0050] 16 million tags (16 x 10 6 For ), the number of rounds (X) of split pool synthesis of barcodes (B) from subunits (S) is determined by the following formula: X = In(B) / ln(S) (X is the number of rounds in split pool synthesis, B is the number of compound barcodes, S is the number of subunits available to synthesize the compound barcode.

[0051] For example, to obtain 16 million compound barcodes from 20 subunits, X = ln(16 × 10) 6 ) / ln20≡6

[0052] In some embodiments, the present invention utilizes adapters attached to one or both ends of a target nucleic acid or copy strand. Various shapes and functions of adapters are known in the art (see, for example, PCT / EP2019 / 05515, filed February 28, 2019, U.S. Patent No. 8,822,150 and U.S. Patent No. 8,455,193).

[0053] The adapter may be double-stranded, partially single-stranded, or single-stranded. In some embodiments, a Y-shaped, hairpin adapter, or stem-loop adapter is used, and the double-stranded portion of the adapter is ligated to a double-stranded nucleic acid formed as described herein.

[0054] In some embodiments, the adapter molecule is an in vitro synthesized artificial sequence. In other embodiments, the adapter molecule is an in vitro synthesized natural sequence. In yet another embodiment, the adapter molecule is an isolated natural molecule or an isolated non-natural molecule.

[0055] The adapter further includes a primer binding site for at least one universal primer.

[0056] Double-stranded or partially double-stranded adapter oligonucleotides can have overhangs or blunt ends. In some embodiments, the double-stranded DNA formed by the methods described herein includes blunt ends to which blunt-end ligation can be applied to ligate blunt-end adapters. In other embodiments, the blunt-end DNA undergoes A-tailing and is adapted to an adapter designed to have a single T nucleotide extending from the blunt end such that a single A nucleotide is added to the blunt end, facilitating ligation between the DNA and the adapter. Commercially available kits for adapter ligation include the AVENIO ctDNA Library Prep kit or the KAPA HyperPrep, and HyperPlus kits (Roche Sequencing Solutions, Pleasanton, CA). In some embodiments, adapter-ligated (adapted) DNA can be separated from excess adapter and unligated DNA.

[0057] In one aspect, the universal blocking oligonucleotide includes a non-specific region adjacent to the first and second specific regions. The non-specific region includes, for example, a series of inosines that align with the sample index sequence when the universal blocking oligonucleotide hybridizes to the target adapter sequence. A specific region of the universal blocking oligonucleotide is complementary to the invariant portion of the adapter sequence and contains one or more melting temperature (T m ) modified bases to increase the T m of the blocking oligonucleotide-adapter duplex. Examples of T m modified base substitutions are shown in Table 1.

[0058]

Table 5

[0059] In another embodiment, a non-amplified nucleic acid library prepared using two different adapter sequences can be processed without blocking oligonucleotides if the adapter ends do not hybridize with each other. Suitable adapter types for this technique include fork-shaped and Y-shaped adapters.

[0060] Adapter ligation may be by single-strand ligation or double-strand ligation. In the case of single-strand ligation, the RT primer may include a universal ligation site. In such embodiments, an adapter having a double-strand region and a single-strand overhang complementary to the universal ligation site in the primer may be annealed and ligated. Annealing of the adapter's single-strand 3' overhang to the universal ligation site at the 5' end of the primer creates a double-strand region with a nick in the strand (copy strand) containing the RT primer. The two strands can be ligated at the nick with a DNA ligase or another enzyme, or a non-enzymatic reagent, that can catalyze the reaction between the 5'-phosphate of the primer extension product and the 3'-OH of the adapter.

[0061] A universal priming site can be added to the opposite end of a copy strand using a single-strand ligation method. In such embodiments, a sequence-independent single-strand ligation method is used. An exemplary method is described in U.S. Patent Application Publication 20140193860. Essentially, this method uses a population of adapters in which the single-strand 3' end overhang has a random sequence, such as a random hexamer sequence, instead of having a universal ligation site. In some embodiments of the method, the adapters also have a hairpin structure. Another example is the method enabled by the ACCEL-NGS 1S DNA Library Kit (Swift Biosciences, Ann Arbor, Michigan).

[0062] The ligation step of this method utilizes a ligase or another enzyme or non-enzymatic reagent with similar activity. The ligase may be a DNA or RNA ligase, such as T4 or E. coli ligase of viral or bacterial origin, or a heat-stable ligase such as Afu, Taq, Tfl, or Tth. In some embodiments, an alternative enzyme, such as a topoisomerase, can be used. Furthermore, a non-enzymatic reagent can be used to form a phosphate-diester bond between the 5'-phosphate of the primer extension product and the 3'-OH of the adapter, as described and referenced in U.S. Patent Application Publication No. 20140193860.

[0063] The present invention relates to assembling cell-specific barcodes for target sequences in individual cells. The resulting barcoded target nucleic acids are subjected to nucleic acid sequencing, preferably large-scale parallel single-molecule sequencing. Analysis of individual molecules by large-scale parallel sequencing typically requires separate levels of barcoding for sample identification and error correction. Use of molecular barcodes as described in U.S. Patents 7,393,665, 8,168,385, 8,481,292, 8,685,678 and 8,722,368. A unique molecular barcode is added to each molecule to be sequenced to mark the molecule and its offspring (e.g., the original molecule and its amplicons generated by PCR). Unique molecular barcodes (UIDs) have multiple applications, including counting the number of original target molecules in a sample and error correction (Newman, A., et al., (2014) An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage, Nature Medicine doi:10.1038 / nm.3519).

[0064] In some embodiments, unique molecular barcodes (UIDs) are used for sequencing error correction. The entire lineage of a single target molecule is labeled with the same barcode, forming a barcoded family. Sequence variations not shared by all members of the barcoded family are artifacts and discarded. Since the entire family represents a single molecule in the original sample, the barcode can also be used for positional deduplication and target quantification (Newman, A., et al., (2016) Integrated digital error suppression for improved detection of circulating tumor DNA, Nature Biotechnology 34:547).

[0065] In some embodiments of the present invention, adapters ligated to one or both ends of a barcode-to-be-labeled target nucleic acid contain one or more barcodes used for sequencing. The barcodes may be UIDs or multiplexed sample IDs (MIDs or SIDs) used to identify the source of the sample from which the sample is mixed (multiplexed). The barcodes may also be a combination of UIDs and MIDs. In some embodiments, a single barcode is used as both the UID and the MID. In some embodiments, each barcode contains a predetermined sequence. In other embodiments, the barcodes contain a random sequence. In some embodiments of the present invention, the barcodes are approximately 4 to 20 nucleotides long, and therefore 96 to 384 different adapters (each having different pairs of identical barcodes) are attached to a human genome sample. In some embodiments, the number of UIDs in the reaction may exceed the number of molecules to be labeled. Those skilled in the art will recognize that the number of barcodes depends on the complexity of the sample (i.e., the expected number of unique target molecules) and that a suitable number of barcodes can be created for each experiment.

[0066] The present invention relates to a method for detecting multiple target nucleic acids in multiple cells. In some embodiments, the method includes contacting multiple cells in a sample with oligonucleotide primers for each target nucleic acid in the presence of nucleic acid polymerase. The nucleic acid polymerase has terminal transferase activity and polymerase activity. The polymerase extends the primers to form a copy strand and adds one or more non-template nucleotides to the 3' end of the copy strand. The method further includes forming a cellular characteristic compound barcode on the copy strand (for each target nucleic acid in each cell of the multiple cells) by sequentially attaching a series of barcode subunits to the 3' end of the copy strand. The barcodes are added by a split-pool process. In some embodiments, the sample is distributed into a first reaction volume set (e.g., tubes or wells in a microwell plate) of each volume containing a first barcode subunit so that the nucleic acid polymerase can extend the 3' end of the copy strand to copy the first barcode subunit. Next, the first set of reaction volumes is pooled and redistributed into a second set of reaction volumes, each volume containing a second barcode subunit so that the nucleic acid polymerase can further extend the 3' end of the copy strand to copy the second barcode subunit. These split-pooling steps are repeated until the desired number of barcode subunits are added to the copy strand of each target nucleic acid. The barcoded strands are then sequenced by any available method. Optionally, the barcoded strands are amplified before sequencing.

[0067] In another embodiment, the disclosure provides a kit for detecting multiple target nucleic acids in multiple cells. In some embodiments, the kit comprises oligonucleotide primers, a reverse transcriptase, and a plurality of barcode subunits, each barcode subunit comprising a poly-dG sequence. The kit may optionally include a buffer containing Mn ions for the terminal transferase activity of the reverse transcriptase. The kit may also include any reagents for DNA capture and purification, for example, for separating copy strands or barcoded copy strands from excess primers or excess barcode subunits.

[0068] In some embodiments, the method includes the step of forming an expression library from multiple cells. The library consists of multiple barcoded copy strands generated from polynucleotides present in the multiple cells. The library can be stored and used multiple times for further processing, such as amplification or sequencing of nucleic acids in the library.

[0069] In some embodiments, the library is subjected to a targeted enrichment process to select a subset of target nucleic acids for further analysis. For example, in embodiments in which oligonucleotide primers include poly-dT sequences for capturing multiple mRNA targets, the targeted enrichment step may be used to capture mRNA transcripts of one or more genes of interest.

[0070] In some embodiments, the method further includes one or more purification steps. In some embodiments, the purification step removes unused primers or unused barcode subunits before the next step of the method. In some embodiments, the primers and barcode subunits are separated from larger copy strands by size exclusion methods (e.g., gel electrophoresis, chromatography, or isophalic or epitacoelectrophoresis).

[0071] In some embodiments, purification is performed by affinity binding. In variations of this embodiment, affinity is for a specific target sequence (sequence capture). In other embodiments, the primer includes an affinity tag. Any affinity tag known in the art, such as biotin, or an antibody or antigen on which a specific antibody exists, can be used. The affinity partner for the affinity tag may be present in solution, for example on a solution-phase solid support such as suspended particles or beads, or it may be bound to a solid-phase support. In the affinity purification process, unbound components of the reaction mixture are washed away. In some embodiments, an additional step is taken to remove unused primers. In some embodiments, affinity capture alters the charge of the primer extension product. For example, inclusion and binding of one or more biotinylated nucleotides or streptavidin generates an altered charge on the nascent nucleic acid chain. The altered charge can be used to separate the nascent chain (primer extension product) by isophatic or epitope electrophoresis.

[0072] In some embodiments, the present invention includes an amplification step prior to the sequencing step. Amplification utilizes upstream and downstream primers. In some embodiments, at least one primer is a target-specific primer, i.e., has a sequence complementary to the target sequence. In some embodiments, one or both primers are universal primers. A universal primer may be paired with another universal primer (of the same or different sequence). In other embodiments, a universal primer is paired with a target-specific primer.

[0073] The universal primer binding site can be attached to the target nucleic acid in various ways. In some embodiments, the universal primer binding site is located on the 5' portion of the RT primer. Similarly, the universal primer binding site may be located on the 5' end of the last barcode subunit attached to the compound barcode. In other embodiments, the universal primer binding site is located in the adapter and is attached to one or both ends of the target nucleic acid by ligation.

[0074] In some embodiments, barcoded copy strands or libraries of barcoded copy strands from multiple cells are sequenced. Many sequencing techniques or sequencing assays are available. As used herein, the term “next-generation sequencing (NGS)” refers to sequencing methods that enable massively parallel sequencing of cloned amplified molecules and single nucleic acid molecules.

[0075] Non-limiting examples of sequence assays suitable for use in the methods disclosed herein include nanopore sequencing (U.S. Patent Applications Publications 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycle sequencing (Sears et al., Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol., 3:39-42 (1992)), and matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nature). This includes sequencing by mass spectrometry such as Biotech., 16:381-384 (1998), sequencing by hybridization (Drmanac et al., Nature Biotech., 16:54-58 (1998)), NGS methods including but not limited to sequencing by synthesis (e.g., HiSeq®, MiSeq®, or Genome Analyzer, all available from Illumina), sequencing by ligation (e.g., SOLiD®, Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent®, Life Technologies), and SMRT® sequencing (e.g., Pacific Biosciences).

[0076] Commercial sequencing technologies include the hybridization sequencing platform from Affymetrix Inc. (Sunnyvale, California), the synthesis sequencing platforms from Illumina / Solexa (San Diego, California) and Helicos Biosciences (Cambridge, Massachusetts), and the ligation sequencing platform from Applied Biosystems (Foster City, California). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisher Scientific), and nanopore sequencing (Genia Technology, Roche Sequencing Solutions, Santa Clara, California) and Oxford Nanopore Technologies (Oxford, UK).

[0077] In some embodiments, the sequencing step involves sequence alignment. In some embodiments, alignment is used to determine a consensus sequence from multiple sequences, for example, multiple sequences having the same unique molecular barcode (UID). A molecular ID is a barcode that can be added to each molecule before sequencing, or before the amplification step if an amplification step is included. In some embodiments, the UID is located in the 5' portion of the RT primer. Similarly, the UID can be located at the 5' end of the last barcode subunit added to the compound barcode. In other embodiments, the UID is located in an adapter and is added to one or both ends of the target nucleic acid by ligation.

[0078] In some embodiments, a consensus sequence is determined from multiple sequences, all having the same UID. Sequences with the same UID are presumed to originate from the same original molecule through amplification. In other embodiments, the UID is used to eliminate artifacts, i.e., variations present in the offspring of a single molecule (characterized by a particular UID). Such artifacts resulting from PCR errors or sequencing errors can be removed.

[0079] In some embodiments, the number of each sequence in a sample can be determined by quantifying the relative number of sequences with each UID in a population having the same multiple sample ID (MID). Each UID is a single molecule in the original sample, and by counting the different UIDs associated with each sequence variant, the fraction of each sequence variant in the original sample, where all molecules share the same MID, can be determined. Those skilled in the art can determine the number of sequence reads required to determine a consensus sequence. In some embodiments, the number in question is the number of reads per UID ("sequence depth") required for accurate quantification results. In some embodiments, the desired depth is 5 to 50 reads per UID.

[0080] The present invention will be described in more detail with reference to Figures 1, 2, and 3. The figures illustrate several embodiments of the present invention and should not be construed as limiting any features or the order or number of steps in the claimed method. The embodiments shown in Figures 1, 2, and 3 have the same steps but differ in the number of any features. The workflow will be described in relation to Figure 1, and where appropriate, Figures 2 and 3 will be referenced to illustrate any features.

[0081] Referring to Figure 1, the first step is to contact a single-stranded target nucleic acid (any type of DNA or RNA) with a primer. The primer has a 3' portion substantially complementary to the target nucleic acid to ensure hybridization and extension. In the example in Figure 1, the complementary sequence is polydT for hybridizing to the polydA end of the mRNA. Other examples of complementary sequences include gene-specific sequences, motif-specific sequences (e.g., kinase domain-specific sequences, RAS family-specific sequences, trinucleotide repeat sequences, etc.), or random sequences (e.g., random hexamer sequences). In some embodiments, the primer sequence is a combination of a random sequence and a polydT sequence.

[0082] Primers may include a 5' portion with additional sequences. The primer in Figure 1 is shown to have a primer binding site, e.g., an amplified primer binding site. The primer in Figure 2 lacks any additional sequences, while the primer in Figure 3 has both a primer binding site and a barcode. The barcode includes a random unique molecular identifier barcode (UMI) (also known as UID) that uniquely identifies a copy strand within the cell. The barcode also includes a defined barcode (BC). Optionally, the sample can be distributed into a series of reaction volumes, each containing a primer with a different barcode (BC). The diversity of barcodes (BC) in different reaction volumes adds to the diversity created by compound barcodes, as described below, to create unique identifiers for individual cells.

[0083] Referring further to Figure 1, the next step is the extension of the primer for copying the target nucleic acid using nucleic acid polymerase. The polymerase is shown to form a copy strand and add a non-template nucleotide to the 3' end of the copy strand. Figure 1 shows the non-template addition of dC3, characteristic of certain types of reverse transcriptase in the presence of manganese ions. Other nucleotide sequences are also possible at the 3' end of the copy strand. For example, reverse transcriptase can add other nucleotides (such as dA or a dA / dC combination) under different buffer conditions. Taq polymerase adds a single dA to the 3' end of the copy strand. It is further possible to separate the functions of polymerase and terminal transferase so that the two functions are performed by separate enzymes.

[0084] Referring further to Figure 1, the next step is a split-pool procedure, in which the reaction mixture is divided into several reaction volumes (e.g., wells of a microwell plate, a tube strip, or other kind of distinct volume), each volume containing a unique barcode subunit oligonucleotide. As shown in Figure 1, the barcode subunit has a complementary sequence that allows it to anneal to the stretch of non-template nucleotides at the ends of the copy strand. Annealing allows the barcode subunit to be copied by nucleic acid polymerase, which extends the copy strand. As shown in Figure 1, in addition to the complementary sequence to the non-template portion of the copy strand, the barcode subunit has a barcode sequence (BC1). The barcode can be a predetermined sequence as shown in Figures 1 and 2. As shown in Figure 3, the barcode is further enhanced with a random sequence (UMI).

[0085] Referring further to Figure 1, the process of adding barcode subunits is repeated until a compound barcode of the desired length is assembled at the end of the copy strand. Each addition of a barcode subunit is performed by a split-pool process. After the barcode subunits are copied by the polymerase, the reactants are pooled and divided into reaction volumes for the next round, each volume containing a unique barcode subunit oligonucleotide for the next series of barcode assemblies. In each step, the polymerase adds a chain of non-template nucleotides to the end of the extended copy strand to enable annealing of the barcode subunit oligonucleotides for the next round.

[0086] In some embodiments, there is an optional purification step to remove unused oligonucleotides, such as unused barcode subunits and primers. In some embodiments, purification utilizes an exonuclease, such as exonuclease I. Removing unused oligonucleotides, and optionally exonucleases, may require buffer exchange, including that by affinity binding as described in the previous section.

[0087] In some embodiments, the use of additional barcodes (e.g., 5' barcodes on oligonucleotide primers) and UMIs minimizes the number of rounds in compound barcode assembly. Fewer rounds, especially a single round, can avoid the purification step.

[0088] The process of annealing to a copy strand and extending the copy strand to copy the next barcode subunit may be repeated until a barcode of the desired length is assembled. Those skilled in the art can calculate the number of assembly rounds required to match the length of the barcode subunit and the number of entities (e.g., cells) tagged with the unique compound barcode.

[0089] The template switching activity of RT is publicly known and commercially used (see Zhu et al., (2001) Reverse transcriptase template switching: a SMART approach for full-length cDNA library construction, Biotechniques, 30:892, and U.S. Patents 5,962,271 and 5,962,272). This property of RT is used to enrich full-length cDNA (SMARTER PCR cDNA synthesis kit (Clontech, Mountain View, California)). More recently, this property of RT has been found to have applications in next-generation sequencing (NGS), where template switch oligonucleotides (TSOs) act as sequencing adapters, i.e., containing NGS-specific sequence elements (see Ohtsubo, et al., (2018) Optimization of single strand DNA incorporation reaction by Moloney murine leukaemia virus reverse transcriptase, DNA Research, 25:477, and SMARTer® smRNA Sequencing kit for Illumina (Takara Bio., Mountain View, Cal.)).

[0090] The present invention comprises the step of ligating (adding multiple TSOs) to a single-stranded copy of a target nucleic acid. In the art, the addition of multiple TSOs is considered an undesirable artifact to be minimized or avoided. For example, Kapteyn et al. ((2010) Incorporation of non-natural nucleotides into template-switching oligonucleotides reduces background and improves cDNA synthesis from very small RNA samples, BMC Genomics, 11:413) teach that a modified nucleotide is incorporated into the 5' end of the TSO to prevent the occurrence of a second template switch. Ohtsubo et al. (previously cited) use 5'-biotinylated oligonucleotides to prevent ligation. The inventors have devised a novel use of this enzymatic property. The present invention includes the use of a 5'-nucleotide isomer (isodC) in the barcode subunit and a shortened incubation time to minimize ligation in the same round while enabling the addition of the barcode subunit in the next round. This method utilizes polymerase to assemble compound barcodes from multiple TSOs, thereby demonstrating the usefulness of polymerase's undesirable binding ability.

[0091] Referring further to Figure 1, the barcode subunit oligonucleotide added in the final assembly round may have an additional sequence in the 5' portion of the oligonucleotide. As shown in Figures 1 and 3, the final barcode subunit has an amplification primer binding site. The barcode subunit may also include a sequencing primer binding site, a unique molecular barcode, and a sample barcode. In Figures 1 and 3, sequencing method-specific or instrument-specific elements are added during pre-sequencing amplification via the 5' portion of the amplification primer. Figure 2 shows the addition of sequencing instrument-specific elements by ligating an adapter to one or both ends of the nucleic acid to be sequenced.

[0092] Referring further to Figures 1 and 2, the workflow yields a double-stranded nucleic acid fragment containing a copy strand of the target nucleic acid and a cell-specific compound barcode assembled at one end of the copy strand by extending the copy strand. In Figure 3, the double-stranded nucleic acid fragment further includes a barcode at the opposite end of the compound barcode.

[0093] In some embodiments, the method further includes evaluating the condition of the subject (e.g., a patient). In some embodiments, the method further includes determining the sequence and optionally the amount of RNA transcripts of multiple disease biomarkers in a patient sample. In one example, the present invention includes evaluating the condition of a patient by determining the expression of immune cell biomarkers in individual immune cells from an immune cell population isolated from the patient. Immune cell biomarkers may be selected from T cell type markers, T cell exhaustion, T cell activation, tissue-resident memory cell markers, tumor-responsive T cell markers, and markers(s). These markers may include one or more RNA transcripts of CD45, CD3, CD8, CD39, CD25, IL-7R, CD4, CXCR3, CCR6, CD3G, CD3D, CD3E, CD2, CD8A, GZMA FOXP3, CD19, CD79A, PDCD1, HAVCR2, IFNG, TNF, ITGAE, and CXCR6 (see Yost et al. (2019) Clonal replacement of tumor specific T-cells following PD-1 blockade, Nature Medicine, doi.org / 10.1038 / s41591-019-0522-3). The method further includes selecting or modifying the treatment based on the diagnosis of the patient's disease or the expression of T-cell biomarkers and disease biomarkers determined in individual cells from a cell population isolated from the patient. [Examples]

[0094] Example 1: Proof of the principle of barcode assembly on cellular mRNA In this example, compound barcodes were constructed on isolated nucleic acids. The target sequences were rearranged immunosequences: T cell receptor genes alpha and beta (TCRA and TCRB), and rearranged immunoglobulin genes (IgG). RNA was isolated from peripheral blood mononuclear cells (PBMCs). The primers contained poly-T sequences and universal PCR primer binding sites. The first-strand cDNA synthesis protocol (Single Cell and cDNA Synthesis) and primer sequences were obtained from New England BioLab (Ipswich, Massachusetts). Three samples were processed as follows:

[0095] Sample 1: Barcode subunit 1 (SEQ ID NO: 1) + Barcode subunit 2 (SEQ ID NO: 2); Sample 2: Barcode subunit 1 (SEQ ID NO: 1) + Barcode subunit 2 (SEQ ID NO: 2) + Enzyme mixture; Sample 3: Barcode subunit 1 only (SEQ ID NO: 1). Sample 2 was processed by sequencing using TCR A / B or IgG gene-specific primers.

[0096] The following barcode subunits were used: Barcode subunit 1 (TSO_1_01) (Sequence ID 1) WNNWCACACTGCTGACNWNrGrGrG (rG is riboguanosine) Barcode subunit 2 (TSO_nano_i7_11) (Sequence ID 2) / 5Me-iso-dC / GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNWCACACTGCTGACWGrGrGrG 5-methylisodeoxycytosine The reaction was subjected to the following temperature profile: 4℃ 1 minute 42℃ 75 minutes 4℃ ~ Add a second barcode subunit 42℃ 10 minutes 72℃ 10 minutes 4℃~ After reverse transcription and barcode assembly, the reaction product was subjected to PCR amplification using gene-specific primers for TCRA, TCRB, and IgG (SEQ ID NOs: 3, 4, and 5) opposed to the universal primer (SEQ ID NO: 6). The amplicon had the following structure: [Universal primer - BC2 - BC1 - gene sequence - gene-specific primer] The amplification primers contained gene-specific sequences, barcode sequences (UMI), and Illumina-specific adapter sequences. TCRA primer (TCRA_i501) (SEQ ID NO: 3) 5'-AATGATACGGCGACCACCGAGATCTACAC TATAGCCTACACTCTTTCCCTACACGACGCTCTTCCGATCT NNNNNNGGCAGGGTCAGGGTTCTGGATA-3' TCRB primer (TCRB_i501) (SEQ ID NO: 4) 5'-AATGATACGGCGACCACCGAGATCTACAC TATAGCCTACACTCTTTCCCTACACGACGCTCTTCCGATCT NNNNNNTGCTTCTGATGGCTCAAACACAGCG-3' IgG primer (IgG_503) (SEQ ID NO: 5) AATGATACGGCGACCACCGAGATCTACACCCTATCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNGTAGTCCTTGACCAGGCAGCC-3' Universal Primer (i_701) (Sequence ID 6) CAAGCAGAAGACGGCATACGAGATCGAGTAATGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC*T (*T is a phosphorothioate nucleotide) PCR amplification was performed using Q5® Hot Start High-Fidelity 2X Master Mix (New England BioLabs, Ipswich, Massachusetts) with the temperature profile recommended by the manufacturer. The amplified product was purified using Kapa Pure Beads (Kapa Biosystems, Wilmington, Massachusetts). The purified amplified product was sequenced using Illumina MiSeq for sequencing. Read lengths were 2 × 300 nucleotides. Reads were analyzed to detect barcode subunits 1 and 2 (BC1, BC2). The results are shown in the table below. [Table 6]

Claims

1. A method for detecting multiple target nucleic acids in multiple cells, wherein the method is a. Contacting multiple target nucleic acids in each of the multiple cells in the sample with oligonucleotide primers in the presence of nucleic acid polymerase having terminal transferase activity and template switch activity. b. For each of the multiple target nucleic acids in each of the multiple cells, the oligonucleotide primer is extended with the nucleic acid polymerase to form a copy chain having one or more non-template nucleotides at its 3' end, where the one or more non-template nucleotides at the 3' end of the copy chain can function as anchoring sites for a barcode subunit. c. Forming a cell characteristic compound barcode containing an attached barcode subunit for each of the multiple target nucleic acids in each of the multiple cells, wherein the multiple target nucleic acids in different cells of the multiple cells contain different cell characteristic compound barcodes, and the multiple target nucleic acids in any one individual cell of the multiple cells contain the same cell characteristic compound barcode, and the formation of the cell characteristic compound barcode in each of the multiple cells in the sample comprises repeatedly performing split pool synthesis rounds until each of the multiple target nucleic acids in each of the different cells of the multiple cells contains the different cell characteristic compound barcodes, and each split pool synthesis round is as follows: i. Randomly distributing the plurality of cells in the sample into a plurality of reaction volume sets, wherein each of the plurality of reaction volumes comprises a unique barcode subunit and the nucleic acid polymerase, and each of the unique barcode subunits comprises a barcode and a nucleic acid sequence complementary to the one or more non-template nucleic acids at the 3' end of the copy strand. ii. The 3' end of the copy strand is extended with the nucleic acid polymerase, wherein the extension of the 3' end of the copy strand copies the unique barcode subunit and adds a non-template nucleic acid strand to the 3' end of the copy strand. iii. Combining the aforementioned multiple reaction volumes into a pool, Includes, The sequence of the extended copy strands obtained above of the plurality of target nucleic acids in each of the plurality of cells containing the formed cell characteristic compound barcode is determined, and thereby the plurality of target nucleic acids in the plurality of cells is detected. Methods that include...

2. The method according to claim 1, further comprising the step of amplifying the extended copy strand containing the cell-characterizing compound barcode before sequencing.

3. The method according to claim 1, wherein the oligonucleotide primer includes a barcode.

4. The method according to claim 2, wherein the oligonucleotide primer includes a universal amplification primer binding site.

5. The method according to claim 2, wherein the last barcode subunit to be copied includes a universal amplification primer binding site.

6. The method according to claim 1, wherein the barcode subunit comprises one or more modified nucleotides that reduce the stability of subunit hybridization.

7. The method according to claim 6, wherein the modified nucleotide is an isonucleotide located at the 5' end of the barcode subunit.

8. The method according to claim 1, wherein the target nucleic acid is RNA.

9. The method according to claim 1, wherein the oligonucleotide primer comprises a target-specific sequence.

10. The method according to claim 1, wherein the oligonucleotide primer includes a barcode.

11. The method according to claim 1, wherein the target nucleic acid is DNA.

12. The method according to claim 1, wherein the nucleic acid polymerase is a reverse transcriptase.

13. The method according to claim 1, wherein the one or more non-template nucleotides are deoxycytosine.

14. The method according to claim 1, wherein the barcode subunit includes a portion complementary to the one or more non-template nucleotides.