Method for barcoding molecules in single cell experiments
By using oligonucleotide co-partitioning and amplification steps with different barcode sequences in single-cell analysis, the high cost and time consumption problems of existing technologies are solved, realizing an economical and efficient combination of target nucleic acids with barcodes and partition-specific combinations, thus improving the accuracy and efficiency of molecular allocation.
Patent Information
- Application Number
- CN202510730482.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-05
AI Technical Summary
Existing techniques for barcoding target nucleic acids in single-cell analysis are costly and time-consuming. Bead-based methods are not only expensive but also complex, making it difficult to achieve efficient and reliable multiplexing.
A single cell or biological particle is used to co-partition with a first oligonucleotide and a second oligonucleotide containing different barcode sequences. The first oligonucleotide is provided in excess, and a target nucleic acid carrying different barcode sequences is generated through an amplification step, which then binds to the second oligonucleotide to achieve partition-specific combination.
This approach enables cost-effective and reliable barcoding of target nucleic acids, reducing material requirements and costs while improving the accuracy and efficiency of molecular allocation.
Smart Images

Figure HDA0005431806620000011 
Figure HDA0005431806620000021 
Figure HDA0005431806620000031
Abstract
Description
Technical Field
[0001] This invention belongs to the field of nucleic acid library preparation and barcoding. Key applications include single-cell sequencing and / or antibody sequencing. Background Technology
[0002] Methods for analyzing nucleic acids at the single-cell level are increasingly being used in biological and biomedical research. These methods allow for understanding the heterogeneity of tissue or cell populations and can be used to identify cell subpopulations involved in diseases.
[0003] To facilitate single-cell analysis and enable multiplexing, current technology workflows include two fundamental features: partitioning each single cell and barcoding cell-specific target nucleic acids.
[0004] Different types of barcodes are known in the art, such as cell barcodes, sample-specific barcodes, and unique molecular identifiers (UMIs). For example, cell barcodes are used to assign target nucleic acids to their cellular origin. When using cell barcodes, all target nucleic acids from one type of cell carry the same barcode sequence, while all target nucleic acids from another type of cell carry different barcodes. Based on this, target nucleic acids can be identified in a cell-specific manner. In contrast, unique molecular identifiers are used to identify individual target nucleic acids. In experiments, each target nucleic acid may carry a different barcode. Furthermore, depending on the application, sample-specific barcodes may be used to identify which target nucleic acids originated from which sample. In standard sequencing experiments, different types of barcodes are combined to achieve multiplexing. During the analysis steps of sequencing data, these barcodes can be used to distinguish different target molecules (based on unique molecular identifiers) to assign each reading to the cell from which it originated and / or a specific sample.
[0005] Current methods for analyzing individual biological particles, including single cells, use droplets (water-in-oil emulsions) or small containers to partition these particles, and then incorporate specific barcodes into target nucleic acid molecules within these partitions. Based on this, in cell analysis steps such as next-generation sequencing, barcoded target nucleic acid molecules from one cell can be distinguished from those from another. One example is the highly parallel analysis of different mRNA transcripts expressed in various cells; this principle can also be applied to the analysis of the presence of genomic DNA, proteins, or other biomolecules using barcode binding agents.
[0006] During sequence analysis, barcodes can be used to associate nucleic acids derived from the same cell.
[0007] Exemplary workflows for these methods include, for example, the 10x Genomics (chromium-based) single-cell product portfolio; InDrop single-cell next-generation transcript sequencing (Klein et al., 2015); and Drop-seq single-cell next-generation transcript sequencing (Macosco et al., 2015).
[0008] One of the main drawbacks of current techniques is that barcoded molecules (or oligonucleotides containing specific barcoded sequences or combinations) must be provided to the target nucleic acid at high concentrations and in multiple copies for the cellular barcode; otherwise, not every target nucleic acid is barcoded. This, in turn, means a large amount of material (barcoded oligonucleotides) is required, making it very costly. Furthermore, most methods for barcoding are bead-based. Bead-based methods are not only expensive but also time-consuming because bead synthesis is complex, error-prone, and time-consuming.
[0009] Therefore, there is a need in the field for improved or alternative, inexpensive and reliable methods for barcoding target nucleic acids. Summary of the Invention
[0010] This invention provides an economical, efficient, and reliable method for barcoding target nucleic acids from a single cell.
[0011] The key element of this method is the co-regioning of a single cell or bioparticle with a single copy of a first oligonucleotide and a second oligonucleotide containing different barcode sequences. The barcode sequences are distinct in each region. Importantly, the first oligonucleotide may be supplied in excess compared to the second oligonucleotide. After releasing the target nucleic acid from the bioparticle, the first oligonucleotide is attached to the target nucleic acid. Based on this, each target nucleic acid carries a different oligonucleotide containing a different barcode sequence (similar to UMI sequences). This is followed by an amplification step, in which the barcode-encoded target nucleic acid and the second oligonucleotide are amplified, resulting in copies of the second oligonucleotide and the barcode-encoded target nucleic acid. After amplification, the copy of the second oligonucleotide is attached to the copy of the barcode-encoded target nucleic acid. Finally, the target nucleic acid carries both the first and second barcode sequences, constituting a cell- or region-specific combination. After sequencing, these combinations can be redistributed to the cell source or region.
[0012] In a first aspect, the present invention provides a method for barcoding nucleic acids, comprising the steps of: A) providing a plurality of biological particles containing a target nucleic acid, a first plurality of oligonucleotides containing a first barcoding sequence, wherein each of the first plurality of oligonucleotides contains a different barcoding sequence, and a second plurality of oligonucleotides containing a second barcoding sequence, wherein each of the second plurality of oligonucleotides contains a different barcoding sequence. B) partitioning the plurality of biological particles such that each partition contains one biological particle and a subset of the first and second oligonucleotides. C) (Optionally) releasing the target nucleic acid of the biological particles into the partitions. D) attaching a first oligonucleotide to each target nucleic acid to generate a barcoded target nucleic acid, wherein each target nucleic acid contains the first barcoding sequence, wherein the first barcoding sequence is different for each target nucleic acid. E) amplifying the second plurality of oligonucleotides and the barcoded target nucleic acid to generate multiple copies of the second plurality of oligonucleotides and the target nucleic acid. F) Attaching a second plurality of oligonucleotides to a target nucleic acid to generate a combined barcoded target nucleic acid, wherein each target nucleic acid contains a specific combination of a first specific barcode sequence and a second specific barcode sequence, wherein the specific combination of the first barcode and the second barcode acts as a partition-specific barcode. Attached Figure Description
[0013] Figure 1: Possible structures of the first and second oligonucleotides. (Legend) Figure 1A -C: P1 (PS1) and P2 (PS2): Specific primer binding sites; BC1.x: Barcode for the first oligonucleotide; BC2.x: Barcode for the second oligonucleotide; TSO: Template-changing oligonucleotide. Dashed lines depict RNA, solid lines or boxes depict DNA. Figure 1A In –B, the target nucleic acid is depicted as a dashed line, and the barcoded cDNA (a copy of the target nucleic acid) is depicted as a solid line. Figure 1C In the diagram, the target nucleic acid is double-stranded and depicted as a solid line.
[0014] Figure 1A The first oligonucleotide serves as a barcode for the target nucleic acid mRNA / cDNA. It contains a first barcode (BC1.x) and an oligomeric (dT) sequence. This first oligonucleotide acts as a primer for cDNA synthesis, generating barcoded cDNA (a copy of the target nucleic acid) in step d). A primer-binding sequence (PS1 or P1) is added to the barcoded nucleic acid (barcoded cDNA) using template conversion of the template-converting oligonucleotide. Both oligonucleotides (first and second) contain an intermediate sequence (IM), which serves as a primer-binding site for amplification in step e), and in step f), the barcoded copy of the target nucleic acid is linked to the second oligonucleotide containing the barcode.
[0015] Figure 1B Barcoding of mRNA / cDNA as target nucleic acid. mRNA is converted to cDNA using an oligonucleotide containing an oligomeric (dT) sequence and a primer binding site (P1 or PS1). The barcode is incorporated into the cDNA through template conversion using a first oligonucleotide containing a first barcode (BC1.x), an intermediate sequence (IM), and a template conversion oligonucleotide sequence (step D).
[0016] Figure 1C dsDNA is barcoded via ligation. The first oligonucleotide contains a first barcode (BC1.x), an intermediate sequence (IM), and a double-stranded region to facilitate ligation via a T4 ligase (or a similar enzyme) (step D). A second (at least partially double-stranded) adaptor is added to the reaction to generate a barcoded nucleic acid containing two primer-binding sites (IM and PS1).
[0017] Figure 2: Principle of the method, example cDNA workflow. Input: Multiple copies of a first oligonucleotide (containing a first random barcode [depicted as BC1.x, where each number with respect to x indicates a different sequence], a primer binding site P2 (PS2), and a sequence that promotes template transition [TS]), and multiple copies of a second oligonucleotide (containing a second random barcode flanked by two primer binding sites P1 (PS1) and P2 (PS2) [depicted as BC2.x, where each number with respect to x indicates a different sequence]), and cells. Figure 2A Step 1: Partition the input (statistical partitioning or active partitioning) to generate partitions containing subsets of a first oligonucleotide, a second oligonucleotide, and cells. Ideally, each partition contains a single cell. Within each partition, the cells are lysed, and the mRNA is converted into cDNA using the first and third oligonucleotides, the third oligonucleotide containing a primer for cDNA synthesis and a third primer binding site P3 (PS3). Figure 2B Step 2 (Initial amplification of oligonucleotides containing the first barcode and cDNA containing the second barcode): Copies of the second oligonucleotide and the barcoded cDNA are amplified in the partition (using primer pairs with primer binding sites P1 / P2 (PS1 / PS2) and P2 / P3 (PS2 / PS3)). Figure 2CStep 3 (Amplicon denaturation and random hybridization at the second primer binding site → extension to generate ds hybrid molecules): The amplicon generated in Step 2 is randomly combined via a common primer binding site P2 (PS2). After initial extension, the hybrid molecules can optionally be amplified using primer pairs with primer binding sites P1 / P3 (PS2 / PS3). As a result, partition-specific combinations of the first and second barcodes will be generated. By analyzing enough molecules, sufficient partition-specific combinations will be obtained to allocate cDNA to separate partitions.
[0018] Figure 3: An example of a workflow using three barcodes.
[0019] Figure 3A Input: Multiple copies of a first oligonucleotide (containing a first random barcode [described as BC1.x, where each number with respect to x indicates a different sequence], a capture sequence [this figure uses an example of an oligomeric (dT) primer], and a third primer-binding site P3 (PS3)); multiple copies of a second oligonucleotide (containing a second random barcode [described as BC2.x, where each number with respect to x indicates a different sequence], a primer-binding site P2 (PS2), and a sequence that promotes template conversion [TSO]); and multiple copies of a third oligonucleotide (containing a third random barcode [described as BC3.x, where each number with respect to x indicates a different sequence] flanked by two primer-binding sites P1 (PS1) and P2 (PS2); multiple biological particles, such as cells. Note: In Figure 3B The example barcode using cells is shown in ff. In the first step, the input is divided into multiple compartments (zones).
[0020] Figure 3B After partitioning, each partition contains multiple copies of the first oligonucleotide, multiple copies of the second oligonucleotide, multiple copies of the third oligonucleotide, and a random subset of biological particles (cells).
[0021] Figure 3C –3H: Reaction cascade within the partition. Note: In Figure 3C In ff, this process uses Figure 3B The instance of the first partition continues.
[0022] Figure 3C Step 1: Lyse the cell and release its mRNA into the compartment.
[0023] Figure 3D Step 2: mRNA is converted into cDNA by reverse transcription using a barcoded first oligonucleotide as a primer and a barcoded second oligonucleotide as a template. The resulting cDNA will contain specific barcodes at the 5' and 3' ends.
[0024] Figure 3E / 3F: Step 3: Separately amplify the third oligonucleotide and the barcoded cDNA molecule generated in Step 2. Figure 3E The nucleic acid after annealing of the amplification primers was depicted, and Figure 3F The nucleic acid after the first amplification cycle is depicted.
[0025] Figure 3G / 3H: Step 4 (denaturation, followed by combinations of different constructs hybridizing via the second primer binding site): After step 3, the amplicon is prepared as a single strand (e.g., by heating). Both the resulting single-stranded copy of the third oligonucleotide and the copy of the barcoded cDNA molecule contain the second primer binding site (complete or partial sequence), which can lead to hybridization between the copy of the oligonucleotide and the copy of the barcoded cDNA molecule. Figure 3H These hybrids are extended using DNA polymerase. Note: For simplicity, only one strand of each amplicon (the strand that can form the extended hybrid molecule) is depicted.
[0026] Figure 4 This diagram illustrates a method for separately amplifying barcode-containing oligonucleotides (or their derivatives), followed by a combination of hybridization and extension (a method for amplifying barcode-containing fragments, followed by random combination – a method for constructing structures). The top of the diagram shows the structures of the first and second oligonucleotides, and the bottom shows the structure of the molecule obtained after combination. The first and second oligonucleotides contain only partial sequences of the second primer site (depicted as P2-part 1 and P2-part 2). Both P2-part 1 and P2-part 2 overlap, but only partially.
[0027] First, the barcode-containing constructs are amplified separately using specific primer pairs (P1+P2-part 1 and P2-part 2+P3, respectively). The annealing temperature is chosen so that P2-part 1 and P2-part 2 can hybridize with both barcode-containing constructs, but below the annealing temperatures of P2-part 1 and P2-part 2 relative to each other (because they only partially overlap). After initial amplification, the annealing temperature is lowered to allow the nucleic acid containing P2-part 1 to hybridize with the nucleic acid containing P2-part 2. This promotes hybridization of the separately amplified constructs, which can be extended to generate double-stranded hybrid constructs. The desired product is then generated (a random combination of oligonucleotides containing the first barcode and barcoded nucleic acids containing the second barcode).
[0028] Figure 5: Method for adding barcodes to nucleic acids. Figures 5A to 5EDifferent methods for linking barcodes to nucleic acids are shown. In each figure, the barcode is depicted as BC1.x, where each digit with respect to x indicates a different sequence. The barcode may be linked ( Figure 5A (A) Directly linked to nucleic acids, or linked to nucleic acids during the synthesis of target nucleic acid copies.
[0029] Figure 5A Barcoding via ligation: Barcoding can be added by attaching a barcode-containing adaptor to one side (shown here) or both sides of a nucleic acid fragment. These barcoded fragments can be amplified using appropriate primers (universal primers specific to the adaptor and / or primers specific to the barcoded nucleic acid) during the initial amplification step (step 2 of the proposed method).
[0030] Figure 5B Fragmentation followed by barcoding via ligation: Some samples may require fragmentation before ligating barcoded adaptors. For example, in methods for analyzing genomic DNA, genomic DNA is often fragmented before ligation of adaptors. Barcoding may be added by ligating a barcoded adaptor to one side (shown here) or both sides of the nucleic acid fragment. These barcoded fragments can be amplified using appropriate primers (universal primers specific to adaptors and / or primers specific to the barcoded nucleic acid) during the initial amplification step (step 2 of the proposed method).
[0031] Figure 5C Barcoding via primer extension (through limited pre-amplification): The adaptor may also be added via a primer extension step. For example, primers containing a primer-binding site specific to the target nucleic acid (delineated in black), a barcode sequence (BC1.x), and a universal primer-binding site (PX) may be used. Such primers may be used for the primer extension reaction (single cycle or multiple cycles). The resulting barcoded nucleic acid then serves as a template in the initial amplification step (step 2) of the proposed method.
[0032] Figure 5D and 5E Barcoding via template conversion: mRNA molecules may be indirectly barcoded during cDNA synthesis. In one approach, barcoding may be introduced through template conversion of oligonucleotides containing barcodes (Zhu et al., 2001). Figure 5D In another approach, barcodes might be introduced using barcoded primers to initiate the reverse transcription reaction. Figure 5E These primers may be target-specific or more general, such as barcoded random primers or barcoded oligo(dT) primers.
[0033] Figure 6: Example of barcoding cDNA within a partition. The example uses a cell as the biological particle, with a template oligonucleotide containing the first barcode sequence and additional oligonucleotides containing the second barcode.
[0034] Figure 6A Top: Yield after final amplification (after step f): The amplified barcoded target nucleic acid successfully attached to the amplified barcoded oligonucleotide (higher yield for samples 1–6 corresponding to different partitions [wells] compared to control samples 7 and 8).
[0035] Figure 6A Bottom: Sequencing results for samples 1-6. Most of the obtained readings contain two barcodes (second column) and show typical mapping statistics for the cDNA library. Detailed Implementation
[0036] In a first aspect, the present invention provides a method for barcoding nucleic acids, comprising the following steps:
[0037] a) provided
[0038] i. Multiple biological particles containing target nucleic acids
[0039] ii. A first plurality of oligonucleotides comprising a first barcode sequence, wherein each of the first plurality of oligonucleotides comprises a different barcode sequence, and
[0040] iii. A second plurality of oligonucleotides comprising a second barcode sequence, wherein each of the second plurality of oligonucleotides comprises a different barcode sequence.
[0041] b) Divide the multiple biological particles into partitions such that each partition contains one biological particle and a subset of the first and second oligonucleotides.
[0042] c) Optionally, release the target nucleic acid of the bioparticle into the partition.
[0043] d) Attaching a first oligonucleotide to each target nucleic acid to generate barcoded target nucleic acids, wherein each (barcoded) target nucleic acid contains a first barcode sequence, wherein the first barcode sequence is different for each target nucleic acid.
[0044] e) Amplify a second set of oligonucleotides and a barcoded target nucleic acid, thereby generating multiple copies of the second set of oligonucleotides and the target nucleic acid.
[0045] f) Attaching a second set of oligonucleotides to the target nucleic acid to generate a combined barcoded target nucleic acid, wherein each (combined barcoded) target nucleic acid comprises a specific combination of a first specific barcode sequence and a second specific barcode sequence, wherein the specific combination of the first barcode and the second barcode acts as a partition-specific barcode.
[0046] Step c) is optional, meaning that the method of the present invention as defined above may only include steps a), b), d), e), and f), while one embodiment of the method of the present invention as defined above includes steps a), b), c), d), e), and f).
[0047] The method may further include the following steps:
[0048] g. Disrupt the partitions to obtain a mixture of target nucleic acids from the partitions, combined with barcodes.
[0049] h. Sequencing the target nucleic acid with the combined barcodes yields sequencing data containing the sequence of the combination of the first and second barcodes, as well as the sequence of the target nucleic acid.
[0050] i. Use combined barcodes to assign target nucleic acids to each partition.
[0051] Step a
[0052] In the first step of the method of the present invention, a plurality of biological particles comprising a target nucleic acid, a first plurality of oligonucleotides comprising a first barcode sequence, and a second plurality of oligonucleotides comprising a second barcode sequence are provided.
[0053] The biological particles according to the invention provided in step a) may be selected from: single cells, bacterial cells, yeast cells, plant cells and virus particles, organelles (nucleus, mitochondria, chloroplasts).
[0054] In one embodiment of the invention, the bioparticles provided in step a) may be single cells. The single cell may be contained in a sample, preferably in a biological sample. The single cell may be a naturally occurring single cell, such as circulating cells contained in PBMCs, umbilical cord blood, or bone marrow (sample). Alternatively, the single cell may be derived from a tissue sample, such as an organ of the lymphatic system or a biopsy isolated from a diseased patient or parts therein suspected to be diseased. The tissue sample may have undergone a dissociation procedure to disrupt cell binding and release cells from the tissue, thereby generating a population of single cells. Standard dissociation procedures are generally known in the art.
[0055] Exemplary tissue samples may be blood, tonsils, lymph nodes, colon, pancreas, skin, or any tissue from the human body.
[0056] In one embodiment of the invention, the single-cell suspension (or single-cell population) may be a “pure” cell population or a mixed cell population. In other words, the single-cell suspension may contain a single cell type, or it may be a mixture of different cell types.
[0057] In one embodiment of the invention, the individual cells included in the single cell population may be selected from cells of hematopoietic cell lineages such as B cells, T cells or NK cells, or solid tissue composed of epithelial cells and mesenchymal cells.
[0058] If the individual cells are a mixed cell population, the cells may be separated or isolated prior to step A) to obtain a specific cell population. Cell separation or isolation can be performed according to standard procedures or using commercially available kits generally known in the art.
[0059] In one embodiment of the invention, the bioparticles may be immune cells, and the method may be used to determine the clonality of an individual cell by profiling B cell receptor (BCR) or T cell receptor (TCR) mRNA, or by analyzing the genomic loci of B cells or T cell receptors.
[0060] In another embodiment of the invention, the bioparticles may be derived from cells obtained from a tissue biopsy, and the method is used, for example, to determine genomic integrity in individual cells by analyzing known or suspected mutations involved in cancer.
[0061] In one embodiment of the invention, the bioparticle may be an organelle. Exemplary organelles are the cell nucleus, mitochondria, or chloroplasts. The organelle may be derived and / or isolated from, for example, a single cell (e.g., a mammalian cell or a plant cell) using methods known in the art prior to step a).
[0062] According to the method for which protection is sought, the biological particle contains a target nucleic acid.
[0063] The target nucleic acid may be contained within a biological particle. In other words, the target nucleic acid may be located within a biological particle (e.g., a single cell or organelle). Therefore, it must be released from the cell's interior (e.g., through cell lysis) before being barcoded.
[0064] In another embodiment of the invention, the target nucleic acid may attach to or bind to the cell surface, for example, via a nucleic acid-labeled antibody (e.g., as used in the CITE-Seq method, Stoeckius et al., 2017). Such nucleic acids may also be referred to as target nucleic acids. The target nucleic acid must be released from the cell surface before barcoding.
[0065] Furthermore, in one embodiment of the invention, the target nucleic acid (DNA, RNA, or chromosome) may be a biological particle. In this embodiment of the invention, the release step (c) is unnecessary.
[0066] As disclosed in this article, the bioparticles may contain a certain amount of target nucleic acid (nx).
[0067] The target nucleic acid may be DNA or RNA (molecule). The target nucleic acid may be genomic DNA, mRNA, a vector (e.g., plasmid, granule, and yeast artificial chromosome), or viral DNA or viral RNA. In one embodiment of the invention, the target nucleic acid may be genomic DNA and mRNA. In a preferred embodiment of the invention, the target nucleic acid is either genomic DNA or mRNA.
[0068] Depending on the application, the target nucleic acid and / or bioparticle to be analyzed may differ. An exemplary application is gene expression profiling or mutation analysis of the entire transcriptome or a subset thereof contained within a single cell (bioparticle). In such applications, the preferred target nucleic acid may be mRNA. Another application may be copy number variation analysis or mutation analysis of genomic DNA (as the target nucleic acid). Another example is the analysis of (cell) surface markers. In such examples, the target nucleic acid may be an oligonucleotide that binds to a binding agent such as an antibody.
[0069] The key elements of this invention are First Multiple oligonucleotides And second Multiple oligonucleotides.
[0070] The first and second oligonucleotides may be double-stranded or single-stranded, with single-stranded being preferred.
[0071] Each oligonucleotide contains at least one barcode sequence. The barcode sequences contained in each oligonucleotide are different from one another. In other words, the first oligonucleotide contains or is composed of a first barcode sequence (BC1mx), wherein each oligonucleotide in the first oligonucleotide contains a different barcode sequence (BC1mx); and the second oligonucleotide contains or is composed of a second barcode sequence (BC2mx), wherein each oligonucleotide in the second oligonucleotide contains a different barcode sequence (BC2mx).
[0072] In other words, the first type of oligonucleotide contains the barcode sequence BC1mx, while the second type of oligonucleotide contains the barcode sequence BC2mx. (m = specific partition; x = barcode sequence number)
[0073] The first plurality of oligonucleotides (comprising the first barcode sequence BC1mx) may be provided at a concentration high enough to attach the barcode sequence to all target nucleic acids in step d). Thus, the oligonucleotides comprised in the first plurality of oligonucleotides (comprising the barcode sequence BC1mx) may be provided in an amount nx < BC1mx.
[0074] In contrast, the second plurality of oligonucleotides may not be provided at a concentration high enough to attach the barcode sequence to all target nucleic acids. In amplification step e), multiple copies of the second barcode oligonucleotides are generated. Based on this, the initial concentration of the second plurality of oligonucleotides (comprising the second barcode sequence BC2mx) need not be as high as that of the first barcode oligonucleotides (comprising the first barcode sequence BC1mx). The concentration may be significantly reduced. As a result, materials and costs can be reduced.
[0075] Thus, the first plurality of oligonucleotides (comprising the first barcode sequence BC1mx) may be provided in excess compared to the second plurality of oligonucleotides (comprising the second barcode sequence BC2mx). In other words, the amount of oligonucleotides comprised in the first plurality of oligonucleotides (comprising the first barcode sequence BC1mx) may be provided in excess compared to the amount of oligonucleotides comprised in the second plurality of oligonucleotides (comprising the second barcode sequence BC2mx). The first plurality of oligonucleotides (comprising the first barcode sequence BC1mx) may be provided in an amount BC1mx >>> BC2mx.
[0076] In one embodiment of the invention, the ratio of the first plurality of oligonucleotides (comprising the first barcode sequence BC1mx) to the second plurality of oligonucleotides (comprising the second barcode sequence BC2mx) may be (at least) 2:1 or (at least) 5:1.
[0077] The concentrations of the first plurality of oligonucleotides and the second plurality of oligonucleotides depend, for example, on the amount and partitioning of the biological particles (e.g., target cells). In one embodiment of the invention, the first plurality of oligonucleotides (comprised within one partition) may comprise at least 1, at least 5, at least 10, at least 100, at least 1000, at least 1000, at least 10000 or at least 100000 first oligonucleotides (molecules) / target nucleic acid (molecules). In addition, the second plurality of oligonucleotides (comprised within one partition) may comprise (at least) 2, (at least) 5, (at least) 10 or (at least) 1 hundred second oligonucleotides (molecules / partition).
[0078] In one embodiment of the invention, the first plurality of oligonucleotides and / or the second plurality of oligonucleotides may additionally contain at least one primer-binding sequence (Figure 1). Based on this, the primer-binding sequence can be added to the target nucleic acid. The primer-binding sequences contained in the first plurality of oligonucleotides and / or the second plurality of oligonucleotides may be the same or different, preferably different.
[0079] In addition, the first and / or second oligonucleotides may further include an intermediate sequence (IM, PS2). The intermediate sequence is identical or complementary in each oligonucleotide (first and second oligonucleotides). This is necessary for the second oligonucleotide to attach to the barcoded target nucleic acid via hybridization. Furthermore, the intermediate sequence (IM, PS2) may act as an additional primer binding site.
[0080] In addition, the first and / or second oligonucleotides may contain additional barcode sequences. These additional (third) barcode sequences may be sample-specific barcodes. They may be identical across all partitions generated from a single sample. They may differ between partitions generated from different samples.
[0081] In one embodiment of the invention, the first plurality of oligonucleotides and the second plurality of oligonucleotides are not attached to the solid support.
[0082] The first and second oligonucleotides are made from DNA, preferably synthetic DNA.
[0083] In one embodiment of the invention, the oligonucleotides (first and second) may be provided in solution (in step a).
[0084] In another embodiment of the invention, the oligonucleotides (first and second) may be coupled to a molecule or protein, such as an antibody, that binds to the biological particle. In such embodiments of the invention, the oligonucleotides may be provided by coupling them to a molecule that binds to the biological particle (in step a).
[0085] First plurality of oligonucleotides
[0086] As disclosed herein, the first multiple oligonucleotide contains a barcode sequence (BC1mx). The barcode sequence is different for each of the multiple oligonucleotides. Optionally, the first multiple oligonucleotide may additionally contain at least one primer-binding sequence and / or intermediate sequence (IM, PS2 / P2). Furthermore, the first multiple oligonucleotide may additionally contain a template-transfer oligonucleotide (TSO, also known as a template-transfer sequence) to facilitate template-transfer reactions. To facilitate the incorporation of the first multiple oligonucleotide (containing the first barcode sequence BC1mx) into the target nucleic acid, the first multiple oligonucleotide may additionally contain a target-specific binding site (TBS) that is at least partially complementary to the target nucleic acid. Figure 1 shows an example of the structure of the first multiple oligonucleotide.
[0087] In one embodiment of the invention, the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), an intermediate sequence (IM, PS2), and a template-transfer oligonucleotide (TSO). It should be understood that the template-transfer oligonucleotide (TSO) is located at the 3' end of the oligonucleotide, while the intermediate sequence (IM, PS2) is located at the 5' end of the oligonucleotide. Figure 1B Accordingly, the barcode sequence (BC1mx) is placed between these sequences. Additional sequences encoding, for example, other barcodes or primer sequences may also be placed between the intermediate sequence (IM, PS2) and the template-transfer oligonucleotide (TSO) sequence. In a preferred embodiment of the invention, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – template-transfer oligonucleotide (TSO).
[0088] In another embodiment of the invention, the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx) and a primer-binding sequence (PS3). It should be understood that the primer-binding sequence (PS3) is located at the 5' end of the oligonucleotide, while the first barcode sequence (BC1mx) is positioned at the 3' end of the primer-binding sequence (PS3). Furthermore, additional sequences encoding, for example, another barcode or primer-binding sequence may also be positioned at the 5' end of the primer-binding sequence (PS3). In a preferred embodiment of the invention, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS3) - barcode sequence (BC1mx).
[0089] In another embodiment of the invention, the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), an intermediate sequence (IM, PS2), and a target-specific binding site (TSB). It should be understood that the target-specific binding site is located at the 3' end of the oligonucleotide, while the intermediate sequence (IM, PS2) is located at the 5' end of the oligonucleotide. Accordingly, the barcode sequence (BC1mx) is positioned between these sequences. It should be understood that additional sequences encoding, for example, other barcodes or primer-binding sequences may also be positioned between the target-specific binding site (TSB) and the intermediate sequence (IM, PS2). In a preferred embodiment of the invention, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – target-specific binding site (TSB) Figure 1A ).
[0090] Second plurality of oligonucleotides It includes a second barcode sequence (BC2mx) as disclosed herein. The barcode sequence is different for each oligonucleotide in the second plurality of oligonucleotides. Furthermore, the (second) barcode sequence (BC2mx) contained in the second plurality of oligonucleotides may differ from the (first) barcode sequence (BC1mx) contained in the first plurality of oligonucleotides.
[0091] In addition, the second plurality of oligonucleotides may further include at least one primer-binding sequence (PS1). The at least one primer-binding sequence may be located at the 5' of the barcode sequence (BC2mx). Based on this, the structure of the first plurality of oligonucleotides may be or may include (5'-3'): primer-binding sequence (e.g., PS1) - barcode sequence (BC2mx).
[0092] In addition, the second oligonucleotide may further include an intermediate sequence (IM, PS2). All intermediate sequences are identical in the first oligonucleotide and all intermediate sequences included in the second oligonucleotide are also identical. Furthermore, the intermediate sequence (IM, PS2) may be (at least partially) complementary to the intermediate sequence (IM, PS2) included in the first oligonucleotide. In one embodiment of the invention, this intermediate sequence is used to attach the second oligonucleotide to a barcoded target nucleic acid via hybridization (step f). Furthermore, the intermediate sequence (IM, PS2) may act as an additional primer-binding site / sequence. It should be understood that the intermediate sequence (IM, PS2) is located at the 3' end of the oligonucleotide, while the primer-binding sequence (PS1) is located at the 5' end of the oligonucleotide. Accordingly, the barcoded sequence (BC2mx) is located between these sequences. It should be understood that the oligonucleotide may also have an inverse complementary orientation (i.e., in a 5'-3' orientation: IM–BC2mx–PS2), or the oligonucleotide may be (partially or in full length) double-stranded. It should be understood that additional sequences encoding, for example, additional barcodes or primer binding sites may also be placed between the primer binding sequence (PS1) and the intermediate sequence (IM, PS2). In this embodiment of the invention, the second plurality of oligonucleotides may have or may comprise the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – primer binding sequence (PS1).
[0093] In one embodiment of the invention, the second plurality of oligonucleotides may not contain the intermediate sequence as disclosed herein. In this embodiment of the invention, the second plurality of oligonucleotides may contain another primer-binding sequence (PS2). The primer-binding sequence (PS2) is identical in each plurality of second oligonucleotides. The primer-binding sequence (PS2) may act as an additional primer-binding site. It should be understood that the primer-binding sequence (PS2) is located at the 3' end of the oligonucleotide, while the primer-binding sequence (PS1) is located at the 5' end of the oligonucleotide, or vice versa. Accordingly, a barcode sequence (BC2mx) is located between these sequences. It should be understood that additional sequences encoding, for example, additional barcodes or primer-binding sites may also be located between the primer-binding sequence (PS2) and the primer-binding sequence (PS1). In this embodiment of the invention, the second plurality of oligonucleotides may have or may contain the following structure (5'-3'): primer-binding sequence (PS1) - barcode sequence (BC2mx) - primer-binding sequence (PS2).
[0094] In one embodiment of the invention, the second oligonucleotide may include an additional (third) barcode sequence (BC3). This additional (third) barcode sequence BC3 may be identical for all second oligonucleotides contained within the second oligonucleotide. Such a barcode can serve as, for example, a sample-specific barcode. The additional (third) barcode sequence may be positioned between the primer-binding sequence (PS1) and the barcode sequence (BC2mx) or between the barcode sequence (BC2mx) and the intermediate sequence (IM, PS2). Therefore, in one embodiment of the invention, the second oligonucleotide may have or may include the following structure (5'-3'): primer-binding sequence (PS1) – additional (third) barcode sequence (BC3) – barcode sequence (BC2mx) – intermediate sequence (IM, PS2). In another embodiment of the invention, the second plurality of oligonucleotides may have or may include the following structure (5'-3'): primer binding sequence (PS1) – barcode sequence (BC2mx) – additional (third) barcode sequence (BC3) – intermediate sequence (IM, PS2).
[0095] Step b) partitioning
[0096] In step b), the multiple biological particles are partitioned such that each partition contains one biological particle and a subset of the first and second oligonucleotides. Based on this, multiple partitions are generated.
[0097] In other words, each of the multiple partitions contains a biological particle (containing the target nucleic acid nx), a first oligonucleotide containing the barcode sequence BC1mx, and a subset of a second oligonucleotide containing the barcode sequence BC2mx. For example, the first partition may contain a subset of the first barcode sequence (BC11x) and the second barcode sequence (BC21x) (containing the first and second oligonucleotides), while the second partition may contain another subset of the first barcode sequence (BC12x) and the second barcode sequence (BC22x) (containing the first and second oligonucleotides), and the third partition may contain another subset of the first barcode sequence (BC13x) and the second barcode sequence (BC23x) (containing the first and second oligonucleotides), and so on.
[0098] Standard procedures for partitioning are known in the art. Example procedures for partitioning cells using droplets or micropores are the Drop-seq method (Macosco et al., 2015) or the Smart-Seq method (Ramskold et al., 2012). In addition, commercially available devices such as the 10x Genomics Chromium Controller (10x Genomics, Pleasanton, CA, USA) can also be used to partition components into droplets.
[0099] Depending on the method used to generate the partitions, some partitions may be empty or contain multiple biological particles. This is considered to be within the error range of the method. However, it is desirable to minimize the number of partitions with multiple biological particles. The allocation of biological particles may be statistical or through an active process. As defined in this paper, the statement "each partition contains one biological particle" refers to the statistical distribution of biological particles within the partition (within the error range of the method).
[0100] In one embodiment of the invention, more than 1%, more than 10%, and preferably more than 50% of the partitions in a plurality of partitions have a given composition.
[0101] In one embodiment of the invention, the partition may be a droplet comprising a subset of biological particles, a first plurality of oligonucleotides, and a second plurality of oligonucleotides. The droplets may be generated according to methods known in the art.
[0102] In another embodiment of the invention, the partition may be, for example, micropores contained in a plate. The micropores may contain biological particles, a subset of a first oligonucleotide and a second oligonucleotide.
[0103] It should be understood that the partition further contains reagents required for the reactions of the method and its implementation schemes. The partition may therefore contain additional primer sets, reaction buffers, nucleotides (dNTPs), and enzymes, such as enzymes for amplification, ligation, and fragmentation.
[0104] (Optional) Step c) release of target nucleic acid
[0105] Step c) involves releasing the target nucleic acid from the biological particle into the partition. It should be understood that this step is performed within the partition. Methods for such release are known in the art.
[0106] In one embodiment of the invention, the bioparticle may be a single cell, and the release step may be a cell lysis step. As a result, the target nucleic acid molecule is released into the compartment.
[0107] Depending on the biological particles and the target nucleic acid, non-target nucleic acids may also be released into the partition. Based on this, a mixture containing both target and non-target nucleic acids is obtained within the partition.
[0108] Cell lysis can be accomplished by enzymatic and / or chemical methods using specific buffer conditions. Several methods and compositions are known in the art. An example of a lysis buffer compatible with oligonucleotides for subsequent hybridization with RNA (especially mRNA) is the one used in Dynabeads. TM Lysis / binding buffer for mRNA purification kit (catalog number A33562, ThermoFisherScientific, Waltham, MA, USA).
[0109] Step d) attachment of first oligonucleotide to target nucleic acid
[0110] Step d) of the method of the present invention includes attaching a first oligonucleotide (containing a barcode sequence BC1mx) to each target nucleic acid (e.g., nx) to generate a barcoded target nucleic acid, wherein each target nucleic acid contains a first barcode sequence (nx-BC1mx, not representing a specific construct) and the first barcode sequence is different for each target nucleic acid.
[0111] For example, in one partition, the first target nucleic acid n1 might contain BC1.1.1, the second target nucleic acid n2 might contain BC1.1.2, the third target nucleic acid n3 might contain BC1.1.3, and so on. Based on this, the target nucleic acids for barcoding could be: n1-BC1.1.1, n2-BC1.1.2, n3-BC1.1.3…nx-BC1.1.x. In another partition, the first target nucleic acid n1 might contain BC1.2.1, the second target nucleic acid n2 might contain BC1.2.2, the third target nucleic acid might contain BC1.2.3, and so on. Based on this, the target nucleic acids for barcoding could be: n1-BC1.2.1, n2-BC1.2.2, n3-BC1.2.3…nx-BC1.2.x.
[0112] In one embodiment of the invention, the attachment of the oligonucleotide may be accomplished by ligation using a ligase, hybridization followed by primer extension, template switching during reverse transcription, or other reactions well known in the art. Common examples of such ligases are, for example, T4 DNA ligase, Taq ligase, or equivalent enzymes. Common examples of enzymes that promote reverse transcription are MMLV reverse transcriptase and its derivatives. To facilitate the ligation of the oligonucleotide to the double-stranded target nucleic acid, the oligonucleotide, as disclosed herein, may be double-stranded or partially double-stranded.
[0113] In one embodiment of the invention, the oligonucleotide may be hybridized and subsequently extended to a target nucleic acid as a primer. Based on this, an oligonucleotide derived from a first plurality of oligonucleotides may hybridize with the target nucleic acid and act as a primer. Such primers can then be used as a starting point for extension and / or amplification reactions. One example of this embodiment is the incorporation of barcoded nucleic acids during cDNA synthesis, wherein the barcoded nucleic acids act as primers for cDNA synthesis. Another example is the synthesis of reverse complementary nucleic acids using single-stranded DNA molecules as templates. This synthesis may be linear or part of a polymerase chain reaction or other amplification reaction (such as whole-genome amplification using Phi29 polymerase).
[0114] In another embodiment of the invention, the oligonucleotide (from a first set of oligonucleotides) may act as a primer during the amplification reaction. In one embodiment, the oligonucleotide may act as a primer in an early cycle of the amplification reaction (e.g., polymerase chain reaction) to generate a barcoded target nucleic acid. In subsequent cycles, separate primer sets may be used to amplify the barcoded target nucleic acid (generated in the early amplification cycle).
[0115] In another embodiment of the invention, the first oligonucleotide is attached via reverse transcription and template conversion. In this embodiment, the target nucleic acid may be mRNA, and the first oligonucleotide additionally contains a template conversion sequence (TSO). Based on this, a barcode is attached to the target nucleic acid (or a copy thereof) via a reverse transcription reaction with template conversion, wherein the barcode oligonucleotide acts as a template conversion oligonucleotide. The newly synthesized nucleic acid is referred to as "cDNA". Furthermore, additional primers, such as oligodt primers, may be required to initiate cDNA synthesis.
[0116] The techniques and conditions for reverse transcription reactions with template switching are well known in the art. Examples of protocols for reverse transcription with template switching can be found in Zhu et al., 2001 and Wellenreuther et al., 2004. The key elements are a specific polymerase and a template-switching oligonucleotide. Commonly used polymerases are Moloney murine leukemia virus reverse transcriptase (MMLV-RT) and its derivatives, or thermostable class II intron reverse transcriptase (TIGRT).
[0117] Based on this, a target nucleic acid containing the first oligonucleotide is generated; in other words, a target nucleic acid with a barcode is generated.
[0118] Step e) amplification of second plurality of oligonucleotides and barcoded target nucleic acid
[0119] Step e) encompasses two amplification steps: the first is the amplification of a second set of oligonucleotides, and the second is the amplification of the barcoded target nucleic acid. It should be understood that the order of these amplifications may be reversed, and the amplifications may also occur simultaneously.
[0120] In the first amplification step, a second oligonucleotide is amplified. In this step, the second oligonucleotide serves as a template. As disclosed herein, the oligonucleotide contains at least one primer-binding sequence (e.g., PS1 and / or PS2).
[0121] In a preferred embodiment of the invention, the oligonucleotide comprises two primer-binding sequences with side barcode sequences, or comprises a primer-binding sequence and an intermediate sequence.
[0122] Amplification primers bind to a specific sequence and initiate the amplification process. This generates multiple copies of a second oligonucleotide. Each copy of the oligonucleotide contains the same barcode sequence. The reaction is limited by the amount of amplification primers. Additional primers (included in the partition) may be provided for this amplification reaction.
[0123] This is followed by a second amplification reaction, in which amplified a barcoded target nucleic acid (from step f). In this amplification reaction, the barcoded target nucleic acid serves as a template. Additional primers (contained in the partition) may be provided for the amplification reaction.
[0124] Step f) attachment of second barcode
[0125] Step f) involves attaching a second plurality of oligonucleotides to the target nucleic acid. For this purpose, combinatorial barcoded target nucleic acids are generated. Each combinatorial barcoded target nucleic acid contains a specific combination of a first specific barcode sequence and a second specific barcode sequence. This specific combination of the first and second barcodes acts as a partition-specific barcode. It is important to note that there are multiple different partition-specific barcodes, i.e., not a single partition-specific barcode as in existing methods (e.g., the 10x GenomicsChromium system).
[0126] The second type of oligonucleotide may attach through hybridization, extension, amplification, or ligation.
[0127] In a preferred embodiment of the invention, a second plurality of oligonucleotides may be attached via hybridization followed by an extension reaction. It should be understood that the target nucleic acid exists as a single-stranded nucleic acid during the hybridization step (or is partially single-stranded in the region where hybridization occurs), along with the second plurality of oligonucleotides. In this embodiment of the invention, the second plurality of oligonucleotides may contain a sequence (e.g., an IM sequence) complementary to the sequence (IM sequence) in the barcoded target nucleic acid generated in step d). The oligonucleotide binds complementary to these sites (Figure 5). There is no complete overlap between the barcoded target nucleic acid and the oligonucleotide. An extension reaction follows to generate complementary double-stranded nucleic acids. Here, the target nucleic acid acts as a template and the oligonucleotide acts as a primer, and vice versa.
[0128] As a result, a double-stranded target nucleic acid containing a first barcode sequence (included in a first plurality of oligonucleotides) and a second barcode sequence (included in a second plurality of oligonucleotides) is generated.
[0129] In another embodiment of the invention, a second plurality of oligonucleotides may be attached to a barcoded target nucleic acid via ligation. Common examples of such ligases are, for example, T4 DNA ligase, Taq ligase, or equivalent enzymes.
[0130] The oligonucleotide may be linked to a single-stranded or double-stranded target nucleic acid. Therefore, the second or more oligonucleotides may be single-stranded or double-stranded.
[0131] Independent of the attachment method, a composite barcode containing a specific combination of a first oligonucleotide and a second oligonucleotide (containing barcode sequences BC1mx and BC2mx) is generated.
[0132] Step g) destruction of partition
[0133] Optionally, the method for which protection is sought may additionally include step g). Step g is a step of disrupting the partitions to obtain a mixture of target nucleic acids from different partitions, barcoded in combination. The disruption of the partitions may be accomplished chemically using specific buffer conditions.
[0134] Optionally, after partitioning is disrupted, target nucleic acids may be separated from non-target nucleic acids.
[0135] Optionally, a washing step may be performed after step g) to remove residues of the lysis buffer and cell debris and / or other substances that may interfere with subsequent processing steps. Buffer conditions are generally known in the art.
[0136] Step h) sequencing of combined barcoded target nucleic acid
[0137] After partitioning, the target nucleic acid is then sequenced to combine the barcodes, thereby obtaining sequencing data containing the combination of the first and second barcodes, as well as the sequence of the target nucleic acid.
[0138] Sequencing may be performed using methods known in the art. For example, sequencing-by-synthesis using one of the Illumina sequencing platforms (e.g., MiSeq, NextSeq, or NovaSeq; Illumina, San Diego, CA, USA), semiconductor sequencing using the Ion Torrent next-generation sequencer (Thermo Fisher Scientific, Waltham, MA, USA), long-read sequencing using the PacBio Revio sequencing system (PacBio, Menlo Park, CA, USA), or nanopore sequencing using one of the Oxford Nanopore sequencing systems (MinIon, GridION, PromethION; Oxford Nanopore, Oxford, UK).
[0139] In one embodiment of the invention, additional sequencing adaptors or barcodes may be attached to the target nucleic acid (e.g., as disclosed in https: / / www.illumina.com / techniques / sequencing / ngs-library-prep / ligation.html).
[0140] Step i) data analysis
[0141] Step i) is a data analysis step. Such a data analysis step can be performed according to methods known in the art. In one embodiment of the invention, the sequencing data obtained in step h) may be compared with a sequence database to determine the identity of the target nucleic acid. However, the identity of the target nucleic acid may be assigned using alternative methods.
[0142] In addition, the sequences of the first barcode sequence (BC1mx) and the second barcode sequence (BC2mx) are analyzed, and specific combinations of the first barcode sequence (BC1mx) and the second barcode sequence (BC2mx) are assigned to specific partitions.
[0143] More specifically, sequencing data can be grouped or clustered based on the first barcode sequence (BC1mx) and the second barcode sequence (BC2mx) contained in the target nucleic acid.
[0144] This is based on the fact that each first barcode and each second barcode is unique or specific. As a result, each partition will have a basic number / quantity of unique or specific first and second barcodes. The method generates partition-specific combinations of first and second barcodes throughout.
[0145] This specific barcode combination (the combination of the first barcode and the second barcode) can be used to evaluate which of the first and second barcodes is present in the same partition. This allows the target nucleic acid (which is linked to the first barcode in step d) to be assigned to the partition.
[0146] This is based on the fact that each copy of the target nucleic acid contained within a partition contains the same first oligonucleotide, which contains the same barcode sequence (BC1mx). Therefore, it is clear that they reside in the same partition. Furthermore, some copies of the same target nucleic acid include the same second oligonucleotide containing the barcode sequence (BC2mx), while other copies of the same target nucleic acid contain different second barcode oligonucleotides. However, based on the first oligonucleotide (containing the barcode sequence BC1mx), it is clear that even if they contain different second oligonucleotides (containing the barcode sequence BC2mx), they originate from the same partition. This type of alignment is performed for all sequencing data, and therefore for all target nucleic acids.
[0147] Based on the limited amount of the second oligonucleotide (and therefore the second barcode sequence BC2mx) provided in step a), there are only a limited number of combinations of BC2mx for each partition. Therefore, it is possible to trace which BC2mx belongs to which partition, and thus which target nucleic acid is derived from which partition.
[0148] In steps h) and i), barcode correction may be employed. Most techniques used for amplification and sequencing have an inherent error rate, and therefore may result in sequences with errors. To avoid misassigning one barcode to another, a barcode correction step may be employed. Methods for barcode correction are generally known in the art.
[0149] To minimize the possibility of misassignment, the barcode sequence may not be random, but rather a large number of different barcodes may be selected, with a Hamming distance and / or Levenstein distance greater than 1, to ensure that amplification or sequencing errors do not lead to barcode misassignment.
[0150] Specific embodiments:
[0151] One aspect of the present invention is a method for barcoding (target) nucleic acids, comprising several steps.
[0152] Step a) includes providing the following components:
[0153] i. A plurality of biological particles comprising nx kinds of target nucleic acids;
[0154] ii. A first plurality of oligonucleotides comprising a first barcode sequence (BC1mx), wherein each oligonucleotide in the first plurality of oligonucleotides comprises a different barcode sequence, and
[0155] iii. A second plurality of oligonucleotides comprising a second barcode sequence (BC2mx), wherein each oligonucleotide in the second plurality of oligonucleotides comprises a different barcode sequence,
[0156] wherein the first plurality of oligonucleotides (comprising barcode sequences) are provided in an amount of nx < BC1mx,
[0157] wherein the second plurality of oligonucleotides (comprising barcode sequences) are provided in an amount of BC1mx > BC2mx.
[0158] Step B includes partitioning the plurality of biological particles such that each partition contains a biological particle and a subset of the first oligonucleotides and the second oligonucleotides (comprising barcode sequences BC1mx and BC2mx respectively), thereby generating a plurality of partitions each containing a different subset of the first plurality of oligonucleotides and the second plurality of oligonucleotides. As a result, a plurality of partitions are generated. An exemplary first partition may contain a subset of the first oligonucleotides and the second oligonucleotides (comprising barcode sequences BC11x and BC21x respectively), while a second partition may contain a subset of the first oligonucleotides and the second oligonucleotides (comprising barcode sequences BC12x and BC22x respectively), and a third partition may contain a subset of the first oligonucleotides and the second oligonucleotides (comprising barcode sequences BC13x and BC23x respectively), and so on.
[0159] (Optionally) Step c) includes releasing the target nucleic acid of the biological particle into the partition.
[0160] Step d) includes attaching a first oligonucleotide to each target nucleic acid, thereby generating barcoded target nucleic acids, wherein each target nucleic acid comprises a first barcode sequence (nx - BC11x), and wherein the first barcode sequence is different for each target nucleic acid.
[0161] Next (step e), a second oligonucleotide (containing the barcoded sequence BC2mx) and a barcoded target nucleic acid are amplified, thereby generating multiple copies of the second oligonucleotide and the barcoded target nucleic acid. In this step, multiple copies of the second oligonucleotide (BC2mx'-x') are generated. Exemplarily, after this step, the first partition contains multiple copies of the second oligonucleotide, such as BC21x; BC21x', BC21x”; and multiple copies of the barcoded target nucleic acid (nx-BC11x); (nx'-BC11x'); (nx”-BC11x”).
[0162] The next step is f), in which a second plurality of oligonucleotides are attached to the target nucleic acid to generate a combined barcoded target nucleic acid, wherein each target nucleic acid contains a specific combination of a first specific barcode sequence and a second specific barcode sequence, wherein the specific combination of the first and second barcode sequences acts as a partition-specific barcode. The attachment of the second plurality of oligonucleotides results in a combined barcoded target nucleic acid, for example: nx-BC11x-BC21x. Based on the different amounts of the first and second oligonucleotides provided in step a), statistically each target nucleic acid contains a specific barcode of the first plurality of oligonucleotides. Furthermore, each copy n' of the target nucleic acid generated in step e) contains the same specific barcode sequence as the "original" target nucleic acid.
[0163] For example, if a partition contains three target nucleic acids N1, N2, and N3, then after step e), they will contain BC1.1.1, BC1.1.2, and BC1.1.3. In step e), multiple copies of these barcoded nucleic acids are generated. For example:
[0164] -N1-BC1.1.1, N1'-BC1.1.1; N1”-BC1.1.1; N1-BC1.1.1”’…
[0165] -N2-BC1.1.2, N2'-BC1.1.2; N2”-BC1.1.2; N2”’-BC1.1.2…
[0166] -N3-BC1.1.3; N3'-BC1.1.3; N3”-BC1.1.3, N3”’-BC1.1.3…
[0167] In addition, in step e), multiple copies of a second oligonucleotide (containing the barcode sequence BC2mx) are generated: for example:
[0168] -BC2.1.1.1;BC2.1.1.1';BC2.1.1.1", BC2.1.1.1"'
[0169] -BC2.1.1.2;BC2.1.1.2';BC2.1.1.2", BC2.1.1.2"'
[0170] -BC2.1.1.3;BC2.1.1.3';BC2.1.1.3", BC2.1.1.3"'
[0171] The attachment of multiple oligonucleotides results in the combination of barcoded nucleic acids. Statistically, each copy of the target nucleic acid contains the same or different BC2mx sequences.
[0172] Some copies of the target nucleic acid contain different BC2mx sequences, while some copies of the same target nucleic acid contain the same BC2mx sequence. This generates specific barcode combinations for each pool of each partition / target nucleic acid. Each pool / copy of the target nucleic acid will contain similar combinations with a second BC2mx. Based on the specific combinations of barcode sequences (BC1mx and BC2mx) contained in each pool / copy of the target nucleic acid, it is possible to trace which partition the target nucleic acid originated from / which target nucleic acids reside in the same partition.
[0173] The subsequent analysis step involves clustering the sequencing data based on the first barcode (BCmx) sequence. Each target nucleic acid containing the same barcode sequence (BC1mx) (which is contained within a first plurality of oligonucleotides) is derived from the initial target nucleic acid or a copy thereof. These copies are located in a partition.
[0174] Furthermore, sequencing data can be clustered based on (second) barcode sequences (which are contained in a second plurality of oligonucleotides). This is based on the assumption that each target nucleic acid containing the same first barcode sequence (BC1mx) is derived from the initial target nucleic acid or a copy thereof. Therefore, each target nucleic acid containing the same first barcode sequence (BC1mx) but containing different second barcodes is derived from the initial target nucleic acid / or a copy thereof, and thus these second barcode sequences (BC2mx) are the barcodes of that partition. Therefore, other target nucleic acids contained in the same partition contain similar second barcode sequences (BC2mx), which can be identified / assigned using the first barcode sequence (BC1mx).
[0175] Variant 1 : attachment of second oligonucleotide by hybridisation in step f)
[0176] In one variant of the invention, a second plurality of oligonucleotides may attach to the target nucleic acid via hybridization.
[0177] The prerequisite is the introduction of an intermediate sequence (IM, PS2) that is complementary to a portion of a second oligonucleotide sequence.
[0178] The intermediate sequence ((IM, PS2)) can be introduced into the target nucleic acid using different methods and primer combinations, as described in the following embodiments, depending on the target nucleic acid and primer construct. The composition or structure of the first multiple oligonucleotide sequence (containing the barcode sequence BC1mx) may vary depending on the method and the target nucleic acid.
[0179] Conversely, the second plurality of oligonucleotide sequences (containing the barcode sequence BC2mx) provided in step a) will be identical for all embodiments. The second plurality of oligonucleotides may contain a primer-binding sequence, an intermediate sequence, and a barcode sequence. Preferably, the second plurality of oligonucleotides may have or may contain the following structure (5'-3'): intermediate sequence (IM, PS2) - barcode (BC2mx) - primer-binding sequence (PS1).
[0180] In a first embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., contained in a single cell), and the first plurality of oligonucleotides may comprise a first barcode sequence (BC1mx), an intermediate sequence (IM, PS2), and a template-transfer oligonucleotide (TSO), wherein each of the first plurality of oligonucleotides comprises a different barcode sequence (BC1mx). Preferably, the first plurality of oligonucleotides may have or may comprise a (5'-3') intermediate sequence (IM, PS2) – barcode (BC1mx) – template-transfer oligonucleotide (TSO).
[0181] In addition, the following components may be provided in step a):
[0182] - Primers used for cDNA synthesis and attachment of multiple oligonucleotides. These primers contain an oligodt sequence (at the 3' end) to bind to the multi-A tail of mRNA; or a specific primer sequence (at the 3' end) that is complementary to the sequence within the target nucleic acid (mRNA). Additionally, primers may contain an additional primer-binding sequence (PS3) that is not complementary to the target mRNA but is incorporated during cDNA synthesis.
[0183] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM; PS2 and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0184] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0185] In step d), a first plurality of oligonucleotide sequences (containing the barcode sequence BC1mx) are attached to the target mRNA via template conversion. More specifically, in this process, primers for cDNA synthesis (oligonucleotides or sequence-specific to the target mRNA) bind to the target mRNA and act as the starting point for reverse transcription. A reverse transcriptase with template conversion activity reverse transcribes the target mRNA into cDNA, and the first plurality of oligonucleotides are attached to the target cDNA via template conversion. As a result, the target cDNA contains the first barcode sequence (BC1mx). In addition, the target cDNA contains primer binding sites PS3 and IM (PS2) at the 3' and 5' ends of each target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – template conversion oligonucleotide (TSO) – target nucleic acid (sequence) – primer binding sequence (PS3).
[0186] In a second embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., from a single cell), and the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), a primer-binding sequence (PS3), and a target-specific binding site (TSB, which acts as a primer sequence for cDNA synthesis). The target-specific binding site may comprise an oligodt sequence (at the 3' end) for binding to the multi-A tail of the mRNA; or a specific sequence (at the 3' end) complementary to a sequence within the target nucleic acid (mRNA). Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): target-specific binding site (TSB)-barcode sequence (BC1mx)-primer-binding sequence (PS3).
[0187] In addition, the following components may be provided in step a):
[0188] - Template-transformed oligonucleotides (TSOs) containing an intermediate sequence (IM, PS2), wherein the TSO is at the 3' end of the oligonucleotide and the intermediate sequence is at the 5' end.
[0189] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0190] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0191] Step d) A first oligonucleotide (containing the barcode sequence BC1mx) is attached to each target mRNA to generate a barcoded target cDNA (target nucleic acid), wherein each cDNA contains the first barcode sequence (BC1mx), which is different for each target cDNA. In this embodiment of the invention, the first oligonucleotide is attached to the target cDNA via template conversion. More specifically, in this process, the first oligonucleotide contains a target-specific binding site (TSB), which acts as a primer sequence (oligomeric dt or sequence-specific within the target mRNA) for cDNA synthesis. It binds to the mRNA and acts as the starting point for reverse transcription. A reverse transcriptase with template conversion activity reverse transcribes the target mRNA into cDNA, and the template-converting oligonucleotide is attached to the target cDNA via template conversion. As a result, the target cDNA contains the first barcode sequence (BC1mx). In addition, the target cDNA contains primer binding sites IM / PS2 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): intermediate sequence (IM, PS2) – template switching oligonucleotide (TSO) – target nucleic acid (sequence) – barcode sequence (BC1mx) – primer binding sequence (PS3).
[0192] In a third embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., derived from a single cell), and the first plurality of oligonucleotides provided in step a) may comprise a barcode sequence (BC1mx), an intermediate sequence (IM, PS2), and a target-specific binding site (TSB) as disclosed herein. The target-specific binding site sequence may be specific to the sequence in the target nucleic acid. Preferably, the first plurality of oligonucleotides (BC1mx) may have or may comprise a (5'-3') intermediate sequence (IM, PS2) – barcode (BC1mx) – target-specific binding site (TSB).
[0193] In addition, the following components may be provided in step a):
[0194] - Primers used for cDNA synthesis and attachment of the first multiple oligonucleotides containing the barcode sequence (BC1mx). These primers contain an oligodt sequence (at the 3' end) to bind to the multi-A tail of the mRNA; or a specific primer sequence (at the 3' end) that is complementary to the sequence within the target nucleic acid (mRNA). Additionally, primers may contain an additional primer-binding sequence (PS3) that is not complementary to the target mRNA but is incorporated during cDNA synthesis.
[0195] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0196] - For the amplification of the target nucleic acid with barcode added in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0197] Subsequently, in step d), a first oligonucleotide containing the barcode sequence (BC1mx) is attached to each target mRNA, thereby generating barcoded target cDNA (target nucleic acid), wherein each cDNA contains the first barcode sequence (BC1mx), which is different for each target cDNA. In this embodiment, the first multiple oligonucleotides are attached to the target cDNA during the amplification reaction. During this process, primers for cDNA synthesis (oligomeric dt or sequence-specific within the target mRNA) bind to the target mRNA and act as the starting point for reverse transcription. Reverse transcriptase reverse transcribes the target mRNA into cDNA. The strands separate, and the first multiple oligonucleotide (containing the barcode sequence (BC1mx)) binds to the target cDNA (via TSB) and acts as a primer for second-strand synthesis and thus as the starting point. This subsequently results in the incorporation of the first multiple barcode oligonucleotide (containing the barcode sequence (BC1mx)) into the target cDNA. As a result, the target cDNA contains the first barcode sequence (BC1mx). In addition, the target cDNA contains primer binding sites IM / PS2 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – target nucleic acid (sequence) – primer binding sequence (PS3).
[0198] In another embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., derived from a single cell), and the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), a primer-binding sequence (PS3), and a target-specific binding site (TSB). The target-specific binding site (TSB) sequence may be an oligodt sequence or a sequence complementary to a sequence contained in the target nucleic acid. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS3) - first barcode sequence (BC1mx) - target-specific binding site (TSB).
[0199] In addition, the following components may be provided:
[0200] - The primers contain a specific sequence complementary to the sequence within the target mRNA (cDNA) and an intermediate sequence (IM, PS2). The intermediate sequence is located at the 5' end of the oligonucleotide.
[0201] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0202] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer IM and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0203] Subsequently, in step d), a first oligonucleotide (containing a first barcode sequence (BC1mx)) is attached to each target mRNA to generate barcoded target cDNA (target nucleic acid), wherein each cDNA contains the first barcode sequence (BC1mx), which is different for each target cDNA. In this embodiment, the first multiple oligonucleotides are attached to the target cDNA during the amplification reaction. In this process, the first multiple oligonucleotides act as primers (TSB, oligodt, or sequence-specific within the target mRNA) for cDNA synthesis, binding to the target mRNA and acting as the starting point for reverse transcription. Reverse transcriptase reverse transcribes the target mRNA into cDNA. A primer containing an IM and a complementary sequence binds to the target cDNA, separating the strands and acting as a primer for second-strand synthesis and thus as the starting point. This subsequently results in the incorporation of the first multiple barcode oligonucleotides (containing the first barcode sequence (BC1mx)) into the target cDNA. As a result, the target cDNA contains the first barcode sequence (BC1mx). In addition, the target cDNA contains primer binding sites IM (PS2) and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – target nucleic acid (sequence) – primer binding sequence (PS3).
[0204] In another embodiment of this variant of the invention, the target nucleic acid may be genomic DNA (e.g., from a single cell), and the first plurality of oligonucleotides provided in step a) may comprise a barcode sequence (BC1mx), an intermediate sequence (IM, PS2), and a target-specific binding site (TSB) as disclosed herein. The target-specific binding site (TSB) may be specific to the sequence in the target nucleic acid. Preferably, the first plurality of oligonucleotides (BC1mx) may have or may comprise a structural (5'-3') intermediate sequence (IM, PS2) – barcode (BC1mx) – target-specific binding site (TSB).
[0205] In addition, the following components may be provided:
[0206] - The primer contains a specific primer sequence (at the 3' end) that is complementary to the sequence within the target genomic DNA. In addition, the primer may contain an additional primer-binding sequence (PS3) that is not complementary to the target sequence in the genomic DNA but is incorporated during the attachment of the first oligonucleotide.
[0207] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0208] - For the amplification of the target nucleic acid with barcode added in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0209] In step d), a first oligonucleotide (containing a first barcode sequence (BC1mx)) is attached to each target genomic DNA (or fragment) to generate barcoded target genomic DNA (or fragment), wherein each genomic DNA (or fragment) contains the first barcode sequence, which is different for each target genomic DNA (or fragment). In this embodiment, the first oligonucleotide (containing the first barcode sequence (BC1mx)) is attached to the target genomic DNA (or fragment) during the amplification reaction. During this process, primers for genomic DNA (or fragment) synthesis (sequence-specific primers within the target genomic DNA (or fragment)) bind to the target genomic DNA (or fragment) and act as the starting point for DNA synthesis. The strands are separated, and the first oligonucleotide (containing the first barcode sequence (BC1mx)) binds to the target genomic DNA (or fragment) (via a target-specific binding site (TSB)) and acts as a primer for second-strand synthesis and thus as the starting point. This subsequently results in the incorporation of a first plurality of oligonucleotides (containing a first barcode sequence (BC1mx)) into the target genomic DNA (or a fragment thereof). As a result, the target genomic DNA (or fragment) contains the first barcode sequence (BC1mx). In addition, the target genomic DNA (or fragment) contains primer binding sites IM (PS2) and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): intermediate sequence (IM, PS2) – barcode sequence (BC1mx) – target nucleic acid (sequence) – primer binding sequence (PS3).
[0210] In another embodiment of this variant of the invention, the target nucleic acid is target genomic DNA (e.g., from a single cell), and the first plurality of oligonucleotides comprises a barcode sequence (BC1mx), a primer-binding sequence (PS3), and a target-specific binding site (TSB). The target-specific binding site (TSB) sequence may be a specific sequence complementary to a sequence contained in the target nucleic acid. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS3) -- first barcode sequence (BC1mx) -- target-specific binding site (TSB).
[0211] In addition, the following components may be provided:
[0212] - The primer contains a specific sequence complementary to the sequence within the target genomic DNA and an intermediate sequence (IM, PS2). The intermediate sequence is located at the 5' end of the oligonucleotide / primer.
[0213] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer IM and reverse primer PS1). The forward primer (IM, PS2) may be specific for the intermediate sequence, and the reverse primer may be specific for the primer-binding sequence (PS1), or vice versa.
[0214] - For the amplification of the target nucleic acid with barcode added in step e), there are specific primer pairs (forward primer IM / PS2 and reverse primer PS3). The forward primer may be specific for the intermediate sequence (IM, PS2), and the reverse primer (PS3) may be specific for the primer-binding sequence (PS3), or vice versa.
[0215] Subsequently, in step d), a first oligonucleotide is attached to each target genomic DNA (or target genomic DNA fragment) to generate barcoded target genomic DNA (or target genomic DNA fragment) (target nucleic acid), wherein each genomic DNA (or target genomic DNA fragment) contains a first barcode sequence (BC1mx), which is different for each target genomic DNA (or target genomic DNA fragment). In this embodiment, the first oligonucleotide is attached to the target genomic DNA (or target genomic DNA fragment) during the amplification reaction. In this process, the first oligonucleotide acts as a primer for the synthesis of genomic DNA (or target genomic DNA fragment), binding to the target genomic DNA (or target genomic DNA fragment) and acting as a starting point for nucleic acid synthesis. As a result, the first barcode sequence BC1mx is incorporated into the target nucleic acid. After nucleic acid synthesis, a primer containing a specific sequence (which is complementary to the sequence within the target genomic DNA (and the intermediate sequence (IM, PS2))) is attached to the target genomic DNA (or target genomic DNA fragment) and acts as a starting point for the synthesis of the second strand. This subsequently leads to the incorporation of the intermediate sequence. As a result, the target genomic DNA (or target genomic DNA fragment) comprises the first barcode sequence and the intermediate sequence. In addition, the target genomic DNA (or target genomic DNA fragment) contains primer binding sites IM (PS2) and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target genomic DNA (or target genomic DNA fragment) has the following structure (5'-3'): intermediate sequence (IM, PS2) - target nucleic acid (sequence) - barcode sequence (BC1mx) - primer binding sequence (PS3).
[0216] Regardless of whether the target nucleic acid is cDNA, mRNA, or genomic DNA, all target nucleic acids generated in step d) contain an intermediate sequence at the 5' end, which can be used for hybridization with a second or more barcodes.
[0217] Based on this, in step e), a second plurality of oligonucleotides and barcoded target nucleic acids are amplified, thereby generating multiple copies of the second plurality of oligonucleotides and barcoded target nucleic acids. More specifically, the second plurality of barcodes are amplified using primers specific to the intermediate sequence (IM, PS2) and the primer-binding sequence (PS1). Furthermore, the barcoded target nucleic acid may be amplified using primers specific to the intermediate sequence (IM, PS2) and primers specific to the primer-binding sequence (PS3). The amplification of both (oligonucleotides and barcoded target nucleic acids) may be performed by PCR. The result of this process is multiple copies of the barcoded target cDNA and multiple copies of each oligonucleotide contained in the second plurality of oligonucleotides.
[0218] The next step, f), involves attaching a second plurality of oligonucleotides to the target cDNA, thereby generating a combined barcoded target nucleic acid, wherein each barcoded target nucleic acid contains a specific combination of a first specific barcode sequence (BC1mx) and a second specific barcode sequence (BC2mx), wherein the specific combination of the first and second barcodes acts as a partition-specific barcode. More specifically, the second plurality of oligonucleotides may attach to the barcoded target nucleic acid via hybridization. This is a complementary hybridization between the 3' end of the second plurality of oligonucleotides and the intermediate sequence (IM, PS2) contained at the 3' end of the target barcoded target nucleic acid. Following hybridization, a chain extension reaction occurs, in which the target barcoded target nucleic acid and the oligonucleotides are extended by polymerase, thereby generating a double-stranded target barcoded target nucleic acid. This target nucleic acid contains two barcodes, which can act as partition-specific barcodes.
[0219] Optionally, the method may additionally include a step of disrupting the partitions to obtain a mixture of combined barcoded target nucleic acids from the partitions.
[0220] The target nucleic acid, combined with barcodes, is sequenced to obtain sequencing data containing the combined sequence of the first and second barcodes, as well as the sequence of the target nucleic acid. The combined barcodes are used to assign the target nucleic acid to each partition.
[0221] Variant 2: attachment of second oligonucleotide by ligation
[0222] In another variant of the invention, a second plurality of oligonucleotides may be attached via linkage.
[0223] In contrast to previous variations of the invention, if a second oligonucleotide containing a barcode sequence (BC2mx) is attached via a linker (to the target nucleic acid), there is no need to introduce an intermediate site.
[0224] In this variant of the invention, the attachment and structure of the first plurality of oligonucleotides containing the barcode sequence (BC1mx) to the target nucleic acid may differ and may depend on the target nucleic acid.
[0225] Conversely, the second plurality of oligonucleotides provided in step a) will be the same for all embodiments. The second plurality of oligonucleotides may comprise a primer-binding sequence (PS2), a primer-binding sequence (PS1), and a barcode sequence (BC2mx). Preferably, the second plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS2) – barcode sequence (BC2mx) – primer-binding sequence (PS1).
[0226] In a first embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., contained in a single cell), and the first plurality of oligonucleotides may comprise a first barcode sequence (BC1mx), a primer-binding sequence (PS4), and a template-transfer oligonucleotide (TSO), wherein each of the first plurality of oligonucleotides comprises a different barcode sequence. Preferably, the first plurality of oligonucleotides may have or may comprise a (5'-3') primer-binding sequence (PS4)-barcode sequence (BC1mx)-template-transfer oligonucleotide (TSO).
[0227] In addition, the following components may be provided in step a):
[0228] - Primers used for cDNA synthesis and attachment of multiple oligonucleotides. These primers contain an oligodt sequence (at the 3' end) to bind to the multi-A tail of mRNA; or a specific primer sequence (at the 3' end) that is complementary to the sequence within the target nucleic acid (mRNA). Additionally, primers may contain an additional primer-binding sequence (PS3) that is not complementary to the target mRNA but is incorporated during cDNA synthesis.
[0229] - For the amplification of the second multiple oligonucleotides in step e), there are specific primer pairs (forward primer PS2 and reverse primer PS1).
[0230] - For the amplification of the target nucleic acid with barcode added in step e), there are specific primer pairs (forward primer PS3 and reverse primer PS4).
[0231] In step d), a first plurality of oligonucleotides attach to the target mRNA via template conversion. More specifically, in this process, primers for cDNA synthesis (P1, oligodt, or sequence-specific within the target mRNA) bind to the target mRNA and act as the starting point for reverse transcription. A reverse transcriptase with template conversion activity reverse transcribes the target mRNA into cDNA, and the first plurality of oligonucleotides attach to the target cDNA via template conversion. As a result, the target cDNA contains the first barcode sequence. In addition, the target cDNA contains primer binding sites PS4 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): primer binding sequence (PS4) - barcode sequence (BC1mx) - template conversion oligonucleotide (TSO) - target nucleic acid (sequence) - primer binding sequence (PS3).
[0232] In a second embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., from a single cell), and the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), a primer-binding sequence (PS3), and a target-specific binding site (TSB, which acts as a primer sequence for cDNA synthesis). The target-specific binding site may comprise an oligodt sequence (at the 3' end) for binding to the multi-A tail of the mRNA; or a specific sequence (at the 3' end) complementary to a sequence within the target nucleic acid (mRNA). Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): target-specific binding site (TSB)-barcode sequence (BC1mx)-primer-binding sequence (PS3).
[0233] In addition, the following components may be provided in the steps:
[0234] - A template-converting oligonucleotide (TSO) containing a primer-binding sequence (PS4), wherein the TSO is at the 3' end of the oligonucleotide and the intermediate sequence is at the 5' end.
[0235] - For the amplification of the second multiple oligonucleotides in step e), there are specific primer pairs (forward primer PS1 and reverse primer PS2).
[0236] - For the amplification of the target nucleic acid with barcode added in step e), there are specific primer pairs (forward primer PS3 and reverse primer PS4).
[0237] In step d), a first oligonucleotide is attached to each target mRNA to generate a barcoded target cDNA (target nucleic acid), wherein each cDNA contains a first barcode sequence (BC1mx), which is different for each target cDNA. In this embodiment of the invention, the first oligonucleotide is attached to the target cDNA via template conversion. More specifically, in this process, the first oligonucleotide contains a target-specific binding site (TSB), which acts as a primer sequence (oligomeric dt or sequence-specific within the target mRNA) for cDNA synthesis. It binds to the mRNA and acts as the starting point for reverse transcription. A reverse transcriptase with template conversion activity reverse transcribes the target mRNA into cDNA, and the template-converting oligonucleotide is attached to the target cDNA via template conversion. As a result, the target cDNA contains the first barcode sequence (BC1mx). In addition, the target cDNA contains primer binding sites PS4 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): primer binding sequence (PS4) – template switching oligonucleotide (TSO) – target nucleic acid (sequence) – barcode sequence (BC1mx) – primer binding sequence (PS3).
[0238] In a third embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., derived from a single cell), and the first plurality of oligonucleotides provided in step a) may comprise the barcode sequence (BC1mx), primer-binding sequence (PS4), and target-specific binding site (TSB) disclosed herein. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS4) - barcode (BC) - and target-specific binding site (TSB).
[0239] In addition, the following components may be provided in step a):
[0240] - Primers used for cDNA synthesis and attachment of the first multiple oligonucleotides containing the barcode sequence (BC1mx). These primers contain an oligodt sequence (at the 3' end) to bind to the multi-A tail of the mRNA; or a specific primer sequence (at the 3' end) that is complementary to the sequence within the target nucleic acid (mRNA). Additionally, primers may contain an additional primer-binding sequence (PS3) that is not complementary to the target mRNA but is incorporated during cDNA synthesis.
[0241] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer PS1 and reverse primer PS2). The forward primer (PS1) may be specific to the primer-binding sequence (PS1), and the reverse primer may be specific to the primer-binding sequence (PS2), or vice versa.
[0242] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer PS4 and reverse primer PS3). The forward primer may be specific to the primer-binding sequence (PS4), and the reverse primer (PS3) may be specific to the primer-binding sequence (PS3), or vice versa.
[0243] Subsequently, in step d), a first oligonucleotide is attached to each target mRNA to generate barcoded target cDNA (target nucleic acid), wherein each cDNA contains a first barcoded sequence (BC1mx), which is different for each target cDNA. In this embodiment, the first multiple oligonucleotides are attached to the target cDNA during cDNA synthesis. During this process, primers for cDNA synthesis (oligomeric dt or sequence-specific within the target mRNA) bind to the target mRNA and act as the starting point for reverse transcription. Reverse transcriptase reverse transcribes the target mRNA into cDNA. The strands separate, and the first multiple oligonucleotides bind to the target cDNA and act as primers for the synthesis of the second strand and thus as the starting point. This subsequently results in the incorporation of the first multiple barcoded oligonucleotides into the target cDNA. As a result, the target cDNA contains the first barcoded sequence (BC1mx) ( Figure 1A In addition, the target cDNA contains primer binding sites PS3 and PS4 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): primer binding sequence (PS4) - barcode sequence (BC1mx) - target nucleic acid (sequence) - primer binding sequence (PS3).
[0244] In another embodiment of this variant of the invention, the target nucleic acid may be mRNA (e.g., derived from a single cell), and the first plurality of oligonucleotides may comprise a barcode sequence (BC1mx), a primer-binding sequence (PS3), and a target-specific binding site (TSB). The target-specific binding site (TSB) sequence may be an oligodt sequence or a sequence complementary to a sequence contained in the target nucleic acid. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS3) - first barcode sequence (BC1mx) - target-specific binding site (TSB).
[0245] In addition, the following components may be provided:
[0246] - The primer contains a specific sequence complementary to the sequence within the target mRNA (cDNA) and an additional primer binding site (PS4). PS4 is located at the 5' end of the oligonucleotide.
[0247] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer PS1 and reverse primer PS2). The forward primer (PS1) may be specific for primer-binding sequence 2, and the reverse primer may be specific for primer-binding sequence (PS2), or vice versa.
[0248] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer PS4 and reverse primer PS3). The forward primer may be specific to the primer-binding sequence (PS4), and the reverse primer (PS3) may be specific to the primer-binding sequence (PS3), or vice versa.
[0249] Subsequently, in step d), a first oligonucleotide (containing a first barcode sequence (BC1mx)) is attached to each target mRNA, thereby generating barcoded target cDNA (target nucleic acid), wherein each cDNA contains the first barcode sequence (BC1mx), which is different for each target cDNA. In this embodiment, the first multiple oligonucleotides are attached to the target cDNA during the amplification reaction. In this process, the first multiple oligonucleotides act as primers (TSB, oligodt, or sequence-specific within the target mRNA) for cDNA synthesis, binding to the target mRNA and acting as the starting point for reverse transcription. Reverse transcriptase reverse transcribes the target mRNA into cDNA. The strand-separating primer, containing IM and complementary sequences, binds to the target cDNA and acts as a primer for second-strand synthesis and thus as the starting point. This subsequently results in the incorporation of the first multiple barcode oligonucleotides (containing the first barcode sequence (BC1mx)) into the target cDNA. As a result, the target cDNA contains the first barcode sequence (BC1mx) (… Figure 1A In addition, the target cDNA contains primer binding sites PS4 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): primer binding sequence (PS4) – barcode sequence (BC1mx) – target nucleic acid (sequence) – primer binding sequence (PS3).
[0250] In another embodiment of this variant of the invention, the target nucleic acid may be genomic DNA (e.g., from a single cell), and the first plurality of oligonucleotides provided in step a) may comprise the barcode sequence (BC1mx), target-specific binding site (TSB), and primer-binding sequence (PS4) disclosed herein. The target-specific binding site (TSB) may be specific to the sequence in the target nucleic acid. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): primer-binding sequence (PS4) – barcode (BC) – target-specific binding site (TSB).
[0251] In addition, the following components may be provided in step a):
[0252] a. The primer contains a specific primer sequence (at the 3' end) that is complementary to the sequence within the target genomic DNA. Additionally, the primer may contain a separate primer-binding sequence (PS3) that is not complementary to the target sequence in the genomic DNA but is incorporated during the attachment of the first oligonucleotide.
[0253] b. For the amplification of the second multiple oligonucleotide in step e), there is a specific primer pair (forward primer PS1 and reverse primer PS2). The forward primer (PS1) may be specific to the primer-binding sequence (PS1), and the reverse primer may be specific to the primer-binding sequence (PS2), or vice versa.
[0254] c. For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer PS4 and reverse primer PS3). The forward primer may be specific to the primer-binding sequence (PS4), and the reverse primer (PS3) may be specific to the primer-binding sequence (PS3), or vice versa.
[0255] In step d), a first oligonucleotide (containing a first barcode sequence (BC1mx)) is attached to each target genomic DNA (or fragment) to generate barcoded target genomic DNA (or fragment), wherein each genomic DNA (or fragment) contains the first barcode sequence, which is different for each target genomic DNA (or fragment). In this embodiment, the first oligonucleotide (containing the first barcode sequence (BC1mx)) is attached to the target genomic DNA (or fragment) during the amplification reaction. During this process, primers for genomic DNA (or fragment) synthesis (sequence-specific primers within the target genomic DNA (or fragment)) bind to the target genomic DNA (or fragment) and act as the starting point for DNA synthesis. The strands are separated, and the first oligonucleotide (containing the first barcode sequence (BC1mx)) binds to the target genomic DNA (or fragment) (via a target-specific binding site (TSB)) and acts as a primer for second-strand synthesis and thus as the starting point. This subsequently results in the incorporation of a first plurality of oligonucleotides (containing a first barcode sequence (BC1mx)) into the target genomic DNA (or a fragment thereof). As a result, the target genomic DNA (or fragment) contains the first barcode sequence. In addition, the target genomic DNA (or fragment) contains PS3 and PS4 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target nucleic acid has the following structure (5'-3'): primer-binding sequence (PS4) – barcode 1 (BC1) – target nucleic acid (sequence) – primer-binding sequence (PS3).
[0256] In another embodiment of this variant of the invention, the target nucleic acid is target genomic DNA (e.g., from a single cell), and the first plurality of oligonucleotides comprises a barcode sequence (BC) and a target-specific binding site (TSB). The target-specific binding site (TSB) sequence may be a specific sequence complementary to a sequence contained in the target nucleic acid. Preferably, the first plurality of oligonucleotides may have or may comprise the following structure (5'-3'): target-specific binding site (TSB) – first barcode sequence (BC1mx) – primer binding sequence (PS3).
[0257] In addition, the following components may be provided:
[0258] - Primers contain a specific sequence complementary to a sequence within the target genomic DNA (or fragment) and an additional primer binding site (PS4). PS4 is located at the 5' end of the oligonucleotide.
[0259] - For the amplification of the second multiple oligonucleotide in step e), there are specific primer pairs (forward primer PS1 and reverse primer PS2). The forward primer (PS1) may be specific for primer-binding sequence 2, and the reverse primer may be specific for primer-binding sequence (PS2), or vice versa.
[0260] - For the amplification of the target nucleic acid barcoded in step e), there are specific primer pairs (forward primer PS4 and reverse primer PS3). The forward primer may be specific to the primer-binding sequence (PS4), and the reverse primer (PS3) may be specific to the primer-binding sequence (PS3), or vice versa.
[0261] Subsequently, in step d), a first oligonucleotide is attached to each target genomic DNA (or target genomic DNA fragment) to generate barcoded target genomic DNA (or target genomic DNA fragment) (target nucleic acid), wherein each genomic DNA (or target genomic DNA fragment) contains a first barcode sequence (BC1mx), which is different for each target genomic DNA (or target genomic DNA fragment). In this embodiment, the first oligonucleotide is attached to the target genomic DNA (or target genomic DNA fragment) during the amplification reaction. In this process, the first oligonucleotide acts as a primer for the synthesis of genomic DNA (or target genomic DNA fragment), binding to the target genomic DNA (or target genomic DNA fragment) and acting as a starting point for nucleic acid synthesis. As a result, the first barcode sequence BC1mx is incorporated into the target nucleic acid. After nucleic acid synthesis, the strands are separated, and the TSB sequence binds to the target genomic DNA (or target genomic DNA fragment) and acts as a starting point for the synthesis of the second strand. This subsequently leads to the incorporation of the primer-binding sequence PS4. As a result, the target genomic DNA (or target genomic DNA fragment) contains the first barcode sequence. In addition, the target genomic DNA (or target genomic DNA fragment) contains primer binding sites PS4 and PS3 at the 3' and 5' ends of the target nucleic acid, respectively. Preferably, the target genomic DNA (or target genomic DNA fragment) has the following structure (5'-3'): primer binding sequence (PS4) – target nucleic acid (sequence) – barcode sequence (BC1mx) – primer binding sequence (PS3).
[0262] Regardless of whether the target nucleic acid is cDNA, mRNA, or genomic DNA, none of the target nucleic acids generated in step d) contain regions that overlap with a second type of oligonucleotide.
[0263] In step e), a second plurality of oligonucleotides and barcoded target nucleic acids are amplified, resulting in multiple copies of both. More specifically, the second plurality of barcodes are amplified using primers specific to the first primer-binding sequence (PS1) and the primer-binding sequence (PS4). Furthermore, the barcoded target nucleic acid may be amplified using primers specific to the primer-binding sequence (PS3) and primers specific to the primer-binding sequence (PS4). Both amplification of the oligonucleotides and the barcoded target nucleic acid may be performed by PCR. The result of this process is multiple copies of the barcoded target cDNA and multiple copies of each oligonucleotide contained in the second plurality of oligonucleotides.
[0264] The next step, f), involves attaching a second plurality of oligonucleotides to the target cDNA to generate a combined barcoded target nucleic acid, wherein each barcoded target nucleic acid contains a specific combination of a first specific barcode sequence and a second specific barcode sequence, wherein the specific combination of the first and second barcodes acts as a partition-specific barcode. More specifically, the second plurality of oligonucleotides may be attached to the barcoded target nucleic acid to generate a double-stranded target barcoded target nucleic acid. This target nucleic acid contains two barcodes, which can act as partition-specific barcodes.
[0265] Optionally, the method may additionally include a step of disrupting the partitions to obtain a mixture of combined barcoded target nucleic acids from the partitions.
[0266] The target nucleic acid, combined with barcodes, is sequenced to obtain sequencing data containing the combined sequence of the first and second barcodes, as well as the sequence of the target nucleic acid. The combined barcodes are used to assign the target nucleic acid to each partition.
[0267] All definitions, features, and embodiments defined herein with respect to the first aspect of the invention disclosed herein, with necessary modifications, are also applicable in the context of other aspects of the invention disclosed herein.
[0268] Definitions
[0269] Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0270] As used herein, the terms “comprising” or “comprises” are used when referring to a composition, method, and their respective components that are essential to the method or composition, but may also include unspecified elements, whether or not they are essential.
[0271] This invention discloses a "method for barcoding nucleic acids". It should be understood that this method is an in vitro method. Therefore, the expressions "method for barcoding nucleic acids" and "in vitro method for barcoding nucleic acids" are used interchangeably.
[0272] The terms “combination,” “hybridize,” “hybridization,” and their exuviates may be used interchangeably. Hybridization occurs when two nucleic acid strands are complementary to each other. Hybridization can occur under conditions known in the art.
[0273] As used herein, the term "complementarity" refers to the ability of two nucleotides to pair precisely via Watson-Crick base pairing. Specifically, if a nucleotide at a given position on one nucleic acid strand is able to form a hydrogen bond with a nucleotide on another nucleic acid strand, the two nucleic acids are considered complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial," where only some nucleotides bind, or it can be complete, where complete complementarity exists between the single-stranded molecules.
[0274] As used herein, a “primer” is a single-stranded oligonucleotide composed of nucleotides that can bind to a complementary nucleic acid sequence, referred to as a “primer binding site.” As disclosed herein, the first and second oligonucleotides may contain such primer binding sites. It should be understood that all primers described in this invention can serve as a starting point for nucleic acid synthesis (e.g., amplification or extension reactions).
[0275] The term "nucleic acid synthesis" is well known in the art. Nucleic acid synthesis (reactions) as disclosed herein may be extension or amplification reactions. In short: for nucleic acid synthesis, a template nucleic acid (e.g., a target nucleic acid) is provided, which may be single-stranded or double-stranded. In the case of initially using a double-stranded nucleic acid, the first step is to denature it into a single nucleic acid strand (complement and reverse complement) using techniques known in the art. For single-stranded nucleic acids, no denaturation step is required. In the next step, primers are provided that bind to the complementary regions of the nucleic acid strands. The 3' end of the primers is then extended (or extended) using a polymerase, and a complementary strand is generated by filling it with complementary nucleotides. As a result, a complementary nucleic acid strand is formed. This will be the result of an extension reaction. For further amplification (amplification reaction), denaturation of the double-stranded nucleic acid is required, and then another round of nucleic acid synthesis can be initiated.
[0276] The terms "nucleic acid," "nucleic acid sequence," "nucleic acid molecule," and "oligonucleotide" are used interchangeably and refer to a biopolymer composed of covalently bonded nucleotide monomers in a chain. The amplified nucleic acid may be named an "amplifier." Nucleic acids can be DNA or RNA.
[0277] According to current methods, "target nucleic acids" are barcoded. These target nucleic acids are also referred to as "nx" (n = nucleic acid; x = nucleic acid sequence number).
[0278] The "oligonucleotides" according to the present invention may contain "barcode" sequences. A barcode sequence is a short nucleotide sequence used for identification purposes. As disclosed herein, in the methods of the present invention, a first plurality of oligonucleotides and a second plurality of oligonucleotides containing barcode sequences are provided. As disclosed herein, "BC1mx" defines a first plurality of oligonucleotides containing or consisting of a first barcode sequence; (m = specific partition; x = barcode sequence number), and "BC2mx" defines a second plurality of oligonucleotides containing or consisting of a second barcode sequence; (m = specific partition; x = barcode sequence number). In addition, the first plurality of oligonucleotides and / or the second plurality of oligonucleotides may further contain a third barcode sequence (BC3). Such barcodes can serve, for example, as sample-specific barcodes.
[0279] As disclosed herein, in step a) of the method of the present invention, the first (or second, respectively) plurality of oligonucleotides comprise a first (or second, respectively) barcode sequence, wherein each of the first (or second, respectively) plurality of oligonucleotides comprises a different barcode sequence. In other words, it is assumed that the first oligonucleotide and the second oligonucleotide each comprise a different barcode sequence. It should be understood that this assumption is made within the known error rate of the methods for synthesizing these oligonucleotides. For example, a popular method in the art is to generate unique or nearly unique barcodes during oligonucleotide synthesis. In oligonucleotide synthesis, nucleotides are synthesized stepwise (one nucleotide at a time in a cycle). Ideally, a discrete nucleotide is added in each cycle to generate a specific sequence. However, several methods add multiple (at least two) different nucleosides in a synthesis cycle, resulting in the random incorporation of one of multiple different types of nucleosides in the respective cycles. For example, theoretically, 16 different barcodes can be generated by having two bases containing an "N" base (A, T, G, or C). Similarly, two or three bases can be incorporated in a single cycle. By using this method, a large collection of different barcode sequences can be generated. What all these methods have in common is that barcodes can appear multiple times during the initial synthesis of oligonucleotides. To have a collection of oligonucleotides containing “different” barcodes, it is important to have a pool of oligonucleotides with different barcode sequences that is larger than the number of oligonucleotides used or incorporated in the experiment. This ensures that the vast majority of the barcodes used or incorporated will be different. Throughout this text, the terms “different barcodes” or “specific barcode” sequences therefore refer to a situation where, statistically, the majority of the barcodes of the oligonucleotides will be different. It refers to a situation where at least 50% / 90% or 95% of the barcode sequences utilized are different (within the population of the first or second oligonucleotide).
[0280] Furthermore, the oligonucleotides according to the invention may contain an intermediate sequence, which may be referred to as "IM" or "IM sequence". This sequence may be used to introduce the oligonucleotide via hybridization.
[0281] As used in this article, the term "multiple" refers to two or more of something.
[0282] The terms “template switching” or “template switching reaction” are well known in the art. In short: First, a primer is hybridized to an RNA molecule. This primer acts as the initiation site for cDNA synthesis via an enzyme with reverse transcriptase activity. Once the enzyme reaches the 5' end of the RNA template, a subset of reverse transcriptases (e.g., MMLV reverse transcriptase) are able to add one or more additional nucleotides (primarily deoxycytidines) to the 3' end of the newly synthesized cDNA. This allows template-switching oligonucleotides (TSOs) to bind to these deoxycytidines. Subsequently, the reverse transcriptase can switch the template and continue cDNA synthesis. Examples of schemes using reverse transcription and template switching can be found in Zhu et al., 2001 and Wellenreuther et al., 2004.
[0283] Examples
[0284] The following examples are intended to explain the invention in more detail, but are not intended to limit the invention to these examples.
[0285] Example 1:
[0286] To demonstrate the feasibility of this method, we evaluated the protocol in different wells of a 96-well plate. First, using a modified protocol from Hagemann-Jensen et al., 2020 (cells were not pre-lysed; instead, the final RT buffer was added directly to the cells), cDNA was synthesized using 6000 cells (K562 cell line) in each well. As template-converting oligonucleotides, we used the oligonucleotide corresponding to SEQ02, which contains a random 12-nucleotide sequence (first barcode).
[0287] After cDNA synthesis, the following reagents were added to each well: 1 pmol of SEQ01 (a random sequence containing 16 nucleotides [second barcode]) in 50 μl of reaction, 1 μM of primers SEQ04 and SEQ07, 0.01 μM of primers SEQ05 and SEQ06, and 25 μl of KAPA HiFiHotStartReadyMix (catalog number 07958927001, Roche Molecular Systems, Basel, Switzerland). The reaction was subjected to the following cycling conditions: 12 cycles of 98°C for 45 seconds; 20 seconds at 98°C – 30 seconds at 69°C and 1 minute at 72°C; 12 cycles of 98°C for 20 seconds – 30 seconds at 63°C and 1 minute at 72°C; 1 minute at 72°C; and held at 4°C.
[0288] After cycling, each sample was purified at 0.6x using SPRIselect beads (catalog number B23317, Beckman Coulter, Brea, CA, USA). Only the combination of barcoded cDNA (samples #1-6) with the second oligonucleotide combination resulted in significant amplification, while the controls (samples #7 and #8) did not show amplification. Figure 6A (Yield column in the middle).
[0289] Next, the samples were analyzed on an Agilent 4200 TapeStation System using either a D5000 or High Sensitivity D5000 Screen Tapes (catalog numbers 5067-5588 and 5067-5592, Agilent, Santa Clara, CA, USA). The product sizes of samples 1-6 show a typical distribution of cDNA, strongly suggesting that the barcoded cDNA (the target nucleic acid with the first barcode) was successfully combined with the second oligonucleotide (the barcoded oligonucleotide).
[0290] We then analyzed the obtained products by sequencing. For this purpose, the samples underwent library preparation using the 10x Genomics 5' Library Construction Kit (PN-220111, 10x Genomics, Pleasanton, CA, USA). Sample indexing PCR was performed using the Sample Index PCR Primer (PN-220111, 10x Genomics, Pleasanton, CA, USA) and the Single Index Plate T Set A (PN-2000240, 10x Genomics, Pleasanton, CA, USA). The libraries were subsequently sequenced on an Illumina MiSeq Nano Cartridge (MiSeq Reagent Nano Kit v2, Illumina, San Diego, CA, USA).
[0291] The results of the sequencing run are summarized as follows Figure 6A The table below shows the percentage readings containing two barcodes. The percentage readings column refers to the relative abundance of index 2 readings containing the barcode plus the first four bases. This combination can only occur if the amplified second barcode has successfully combined with the target nucleic acid for which it is barcoded. Such combinations occur in a surprisingly high number (almost 95% of all obtained readings). Furthermore, the obtained readings are typical for gene expression / cDNA analysis, indicated by the number of mapping readings, aligned to genes, exons, and between genes.
Claims
1. A method for barcoding of nucleic acids comprising the steps of: a) providing i. a plurality of biological particles comprising target nucleic acids, ii. a first plurality of oligonucleotides comprising a first barcode sequence, wherein each oligonucleotide of the first plurality of oligonucleotides comprises a different barcode sequence, and iii. a second plurality of oligonucleotides comprising a second barcode sequence, wherein each oligonucleotide of the second plurality of oligonucleotides comprises a different barcode sequence; b) partitioning the plurality of biological particles such that each partition comprises one biological particle and a subset of first and second oligonucleotides; c) optionally, releasing the target nucleic acids of the biological particles into the partitions; d) attaching one first oligonucleotide to each target nucleic acid, thereby generating barcoded target nucleic acids, wherein each of the target nucleic acids comprises a first barcode sequence, wherein the first barcode sequence is different for each target nucleic acid; e) amplifying the second plurality of oligonucleotides and the barcoded target nucleic acids, thereby generating a plurality of copies of the second plurality of oligonucleotides and the target nucleic acids; f) attaching the second plurality of oligonucleotides to the target nucleic acids, thereby generating combinatorially barcoded target nucleic acids, wherein each target nucleic acid comprises a specific combination of a first specific barcode sequence and a second specific barcode sequence, wherein the specific combination of the first and second barcode acts as a partition specific barcode.
2. The method according to claim 1, wherein the method additionally comprises the steps of: g) disrupting the partitions, thereby obtaining a mixture of combinatorially barcoded target nucleic acids from the partitions; h) sequencing the combinatorially barcoded target nucleic acids, thereby obtaining sequencing data comprising the sequence of the first and second barcode combination and the sequence of the target nucleic acids; i) using the combination barcode for assigning target nucleic acids to each partition.
3. The method according to any one of claims 1 or 2, wherein the attachment of the second oligonucleotides in step f) is by hybridization, elongation, amplification or ligation.
4. The method according to any one of claims 1-3, wherein the first and second plurality of oligonucleotides additionally comprise intermediate sequences (IM), wherein the IM sequences are complementary to each other, wherein the IM sequences are incorporated into the barcoded target nucleic acids, and wherein the second plurality of oligonucleotides are attached in step e) by complementary hybridization of the IM sequences, followed by an elongation reaction, wherein the second plurality of oligonucleotides act as primers.
5. The method according to any one of claims 1-4, wherein the biological particles provided in step a) are selected from the group consisting of: single cells, bacteria, viral particles.
6. The method according to any one of claims 1-5, wherein the plurality of biological particles provided in step a) is a population of single cells, and wherein in step c) the single cells are lysed and the target nucleic acids are released into the partitions.
7. The method according to any one of claims 1-6, wherein the target nucleic acids are genomic DNA or mRNA. 8. The method of claim 7, wherein the target nucleic acid is mRNA, and wherein the first plurality of oligonucleotides additionally comprises a template switch sequence, and wherein the first oligonucleotide is attached in step d) by a template switch reaction.
9. The method of claim 7, wherein the target nucleic acid is genomic DNA, wherein the genomic DNA is fragmented after step c), wherein the first oligonucleotide is attached in step d) by ligation.