OPTIMIZED SET OF OLIGONUCLEOTIDS FOR MASS RNA BARCODING AND SEQUENCER
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- ALITHEA GENOMICS SA
- Filing Date
- 2022-06-07
- Publication Date
- 2026-08-05
AI Technical Summary
Existing RNA sequencing methods for bulk samples face challenges in ensuring uniform distribution of sequencing reads across samples due to issues such as secondary structures, sequence similarity, preferential amplification, and mitochondrial transcript bias, leading to inaccurate and costly data analysis.
A set of optimized oligonucleotide molecules, each comprising a sequencing adapter, a 9-15 nucleotide barcode, and an mRNA capture sequence, selected to minimize technical biases and ensure uniform read distribution, including a unique molecular identifier (UMI) for error-tolerance.
The solution provides a reliable and cost-effective method for RNA sequencing by ensuring uniform sequencing read distribution and reducing sample-to-sample variation, enhancing the accuracy and efficiency of bulk RNA sequencing.
Description
FIELD OF THE INVENTION
[0001] The present invention relates generally to the field of nucleic acid sequencing and provides oligonucleotide molecules and barcodes contained therein. These oligonucleotide molecules and barcodes molecules are useful in sequencing to identify and resolve errors.BACKGROUND OF THE INVENTION
[0002] RNA sequencing has become the method of choice for genome-wide transcriptomic analyses as its price has substantially decreased over the last years. Nevertheless, the high cost of standard RNA library preparation and the complexity of the underlying data analysis still prevent this approach from becoming as routine as quantitative PCR (qPCR), especially when many samples need to be analyzed.
[0003] To alleviate this high cost, the emerging single-cell transcriptomics field implemented the sample barcoding / early multiplexing principle. This reduces both the RNA-seq cost and preparation time by allowing the generation of a single sequencing library that contains multiple distinct samples / cells (Ziegenhain et al., 2017).
[0004] Such a strategy could also be of value to reduce the cost and processing time of bulk RNA sequencing of large sets of samples (Kilpinen et al., 2017; Pradhan et al., 2017; Waszak et al., 2015). However, there have been surprisingly few efforts to explicitly adapt and validate the early-stage multiplexing protocols for reliable and affordable profiling of bulk RNA samples. Early multiplexing protocols designed for single-cell RNA profiling (CEL-seq2, SCRB-seq, and STRT-seq) provide a great capacity for transforming large sets of samples into a unique sequencing library (Hashimshony et al., 2016; Islam et al., 2012; Soumillon et al., 2014). This is achieved by introducing a sample-specific barcode during the RT reaction using a "molecular tag" carried by either the oligo-dT or the template switch oligo (TSO). After individual samples have been "tagged", they are pooled together, and the remaining steps are performed in bulk, thus shortening the time and cost of library preparation.
[0005] Since the tag is introduced to the terminal part of the transcript prior to fragmentation, the reads solely cover the 3' or 5' end of the transcripts. The 3'DGE approach for bulk RNA profiling, has been adopted in several recent studies, such as PLATE-seq (Bush et al., 2017), DRUG-seq (Ye et al., 2018), 3'POOL-seq (Sholder et al., 2020), PME-seq (Pandey et al., 2020)) and BRB-seq (Alpern et al., 2019). These techniques have two main commonalities: i) using barcoded DNA oligos used to "tag" poly-adenylated RNA molecules during first strand synthesis and ii) pooling together of all the tagged samples in one tube after the barcoding step.
[0006] The overarching goal of these techniques is to decrease the costs and increase the throughput associated with mRNA sequencing library preparation of bulk samples.
[0007] This is achieved by reducing reagents, consumables and personnel time through pooling in one solution several barcoded samples. In simple terms, it is much more cost-effective and simpler to process e.g. 96 samples in one tube than 96 samples in 96 tubes.
[0008] One of the main challenges of RNA barcoding applied to bulk samples is the ability to guarantee a uniform distribution of sequencing reads across all samples. This challenge is due to the fact that the "molecular barcodes" used during the RNA barcoding step are a functional portion of the "reverse transcription primer" and, as such, different barcodes (i.e. barcodes with different sequences) can have significant effect on the efficiency of the overall workflow. For example, empirical experimental evidence highlights the following potential issues: i) barcodes can lead to unwanted secondary structures that interfere or prevent efficient priming ii) completely random barcodes may end up having very similar or repeating sequences which are then difficult to resolve at the sequencing stage, iii) certain barcodes can be preferentially amplified within the same pool and last, but not least, iv) certain barcodes preferentially bind mitochondrial transcripts, which then appear as an unwanted bias in the sequencing results.
[0009] WO / 2022 / 069039 (PCT / EP2020 / 077437) relates to methods for the preparation of cDNA library based on one or many RNA samples useful for efficient RNA sequencing and uses thereof. This document does neither disclose the present methods nor the barcode sequences of the present invention.
[0010] US 2020 / 157600 discloses oligonucleotides for single cell transcriptional profiling with sequencing adapter, cell label barcode, UMI and polyT for capturing of mRNA and cDNA synthesis. This document does not disclose the barcode sequences of the present invention.
[0011] Therefore, there is still a need for more accurate, dependable sequencing tools, (i) and methods that can eliminate barcoding sample-to-sample variation and (ii) methods that use them to improve various barcoding approaches, including the barcoding-mediated high-accuracy sequencing method.SUMMARY OF THE INVENTION
[0012] The present invention provides a set of oligonucleotide molecules comprising, from 5' to 3', a) a sequencing adapter, b) a barcode sequence consisting of 9 to 15 nucleotides, preferably 14 nucleotides, and c) an mRNA capture sequence, wherein the barcode sequence is selected from the group comprising SEQ ID No. 1 to SEQ ID No. 96, or SEQ ID No. 98 to SEQ ID No. 481, or a combination of one or more thereof.
[0013] Further provided is set of oligonucleotide molecules, each molecule comprising, from 5' to 3', a) a sequencing adaptor, b) a barcode sequence independently selected from the group consisting of SEQ IDNo 1 to No.96 or a barcode sequence independently selected from the group consisting of SEQ ID No. 98 to No. 481, and c) an mRNA capture sequence.
[0014] Further provided is the use of a set of oligonucleotide molecules of the invention in a sequencing method.
[0015] Further provided is a method for providing a cDNA library, the method comprising the steps of a) Providing a plurality of RNA samples obtained from a biological sample; b) Contacting separately each RNA sample with a set of oligonucleotide molecules of the invention, or of a library of the invention, or of a barcode oligonucleotide sequence or a combination of one or more thereof, of the invention, under annealing conditions; c) Incubating separately each sample under reverse transcription reaction conditions; d) Pooling together all the cDNA:RNA sample; e) Proceeding to second strand synthesis under synthesis conditions; and f) Proceeding with tagmentation and amplification under suitable conditions so as to obtain a cDNA library.
[0016] Further provided is a method for sequencing RNA, the method comprising the steps of a) Providing a cDNA library obtained by the method of the invention; and b) Proceeding to the sequencing under suitable conditions.
[0017] Further provided is a kit comprising a) a set of oligonucleotide molecules of the invention, b) a support for sample preparation, and c) reagents for sequencing.DESCRIPTION OF THE FIGURES
[0018] Figure 1 shows the read distribution of the optimal set of 96 barcodes (SEQ ID No. 1 to SEQ ID No. 96) of the invention in which all barcodes are functional and obtain a similar number of sequencing reads. Figure 2 shows an example of sequencing read distribution for a optimal set of 384 barcodes (SEQ ID No. 98 to SEQ ID No. 481) of the invention in which all barcodes are functional and obtain a similar number of sequencing reads. Figure 3 shows an example of sequencing read distribution for a suboptimal set of 96 barcodes. Barcodes marked with an arrow systematically underperform as compared to the others. DESCRIPTION OF THE INVENTION
[0019] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The publications and applications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0020] In the case of conflict, the present specification, including definitions, will control. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of skill in art to which the subject matter herein belongs. As used herein, the following definitions are supplied in order to facilitate the understanding of the present invention.
[0021] The term "comprise / comprising" is generally used in the sense of include / including, that is to say permitting the presence of one or more features or components. The terms "comprise(s)" and "comprising" also encompass the more restricted ones "consist(s)", "consisting" as well as "consist / consisting essentially of", respectively.
[0022] As used in the specification and claims, the singular form "a", "an" and "the" include plural references unless the context clearly dictates otherwise.
[0023] As used herein, "one or more" includes "two or more", "three or more", etc. For example, one or more oligonucleotide molecules refers to one oligonucleotide molecule, two oligonucleotide molecules, three oligonucleotide molecules, etc....
[0024] The present invention is based on the discovery of an optimal set barcoded oligonucleotides for multiplexed RNA sequencing. These oligonucleotides contain barcodes that are 14 base pairs long and have been therefore selected from a pool of 4^14 = 268'435'456 potential candidates. This large pool has been filtered twice, first computationally and then experimentally as disclosed herein. The goal was to obtain an optimized set of barcodes for further being able to 1) uniquely demultiplex the samples, with error-tolerance for sequencing errors, and 2) adapt the barcodes for potential technical bias such as overrepresentation of polyT sequences due to preferential amplification of certain sequences.
[0025] In one aspect, the invention provides a set oligonucleotide molecules comprising, from 5' to 3', a) a sequencing adapter, b) a barcode sequence consisting of 9 to 15 nucleotides, preferably 14 nucleotides, and, c) an mRNA capture sequence.
[0026] In one aspect, the a set of oligonucleotide molecules further comprise d) a unique molecular identifier (UMI). Usually, said UMI consists of 14 nucleotides. "Oligonucleotide" or "polynucleotide," which are used synonymously, means a linear polymer of natural or modified nucleosidic monomers linked by phosphodiester bonds or analogs thereof. The term "oligonucleotide" usually refers to a shorter polymer, e.g., comprising from about 3 to about 100 monomers, and the term "polynucleotide" usually refers to longer polymers, e.g., comprising from about 100 monomers to many thousands of monomers, e.g., 10,000 monomers, or more. Oligonucleotides and polynucleotides may be natural or synthetic. Oligonucleotides and polynucleotides include deoxyribonucleosides, ribonucleosides, and non-natural analogs thereof, such as anomeric forms thereof, peptide nucleic acids (PNAs), and the like, provided that they are capable of specifically binding to a target genome by way of a regular pattern of monomer-to-monomer interactions, such as Watson-Crick type of base pairing, base stacking, Hoogsteen or reverse Hoogsteen types of base pairing, or the like.
[0027] The terms "peptide," "protein," and "polypeptide" are used interchangeably to refer to a natural or synthetic molecule comprising two or more amino acids linked by the carboxyl group of one amino acid to the alpha amino group of another.
[0028] The term "nucleic acid" refers to a natural or synthetic molecule comprising a single nucleotide or two or more nucleotides linked by a phosphate group at the 3' position of one nucleotide to the 5' end of another nucleotide. The nucleic acid is not limited by length, and thus the nucleic acid can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0029] "Sequencing" refers to determining the order of nucleotides (base sequences) in a nucleic acid sample, e.g. DNA or RNA. Many techniques are available such as Sanger sequencing and High Throughput Sequencing technologies (HTS). Sanger sequencing may involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries may be sequence analysed in one run. High throughput sequencing involves the parallel sequencing of thousands or millions or more sequences at once. HTS can be defined as Next Generation sequencing, i.e. techniques based on solid phase pyrosequencing or as Next-Next Generation sequencing based on single nucleotide real time sequencing (SMRT). HTS technologies are available such as offered by Roche, Illumina and Applied Biosystems (Life Technologies). Further high throughput sequencing technologies are described by and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio. Each of these sequencing technologies have their own way of preparing samples prior to the actual sequencing step. Depending on the sequencing technology used, amplification steps may be omitted.
[0030] As used herein, the term "barcode" refers to a unique oligonucleotide sequence that allows a corresponding nucleic acid base and / or nucleic acid sequence to be identified. In certain aspects, the nucleic acid base and / or nucleic acid sequence is located at a specific position on a larger polynucleotide sequence (e.g., a polynucleotide covalently attached to a bead). In certain aspects, barcodes can each have a length within a range of from 4 to 150 nucleotides. The barcode technology (or barcoding) has been a particularly powerful technique for studying the genetic and functional variations of the target pool and for high-accuracy target DNA sequencing. Each barcode can comprise deoxyribonucleotides, optionally all of the nucleotides in a barcode region are deoxyribonucleotides. One or more of the deoxyribonucleotides may be a modified deoxyribonucleotide (e.g. a deoxyribonucleotide modified with a biotin moiety or a deoxyuracil nucleotide). The barcodes may comprise one or more degenerate nucleotides or sequences. The barcode regions may not comprise any degenerate nucleotides or sequences.
[0031] The barcode sequence of the invention consists of 9 to 15 nucleotides, preferably 14 nucleotides.
[0032] In one aspect, the barcode sequence is selected from the group comprising, or consisting of, SEQ ID NO. 1TACGTTATTCCGAASEQ ID NO. 2AACAGGATAACTCCSEQ ID NO. 3ACTCAGGCACCTCCSEQ ID NO. 4ACGAGCAGATGCAGSEQ ID NO. 5TTCAATCTCCTTAGSEQ ID NO. 6CTCGGTTCGAATGCSEQ ID NO. 7TACACTATAGCTAGSEQ ID NO. 8CAAGTATAAGGAACSEQ ID NO. 9CTGATATGCAGCGASEQ ID NO. 10ACAATAGGTGGTCCSEQ ID NO. 11AACCATCGCATCCASEQ ID NO. 12CAGTCTAGTGCGCASEQ ID NO. 13CCGCAAGAGGTTGGSEQ ID NO. 14TAACACCTCCAACGSEQ ID NO. 15CAGACATGATAGAGSEQ ID NO. 16AACAACCGAAGTAASEQ ID NO. 17AACTAGTATCCGGASEQ ID NO. 18ACGCAGCGAGCAGASEQ ID NO. 19TCCAAGGCATTAAGSEQ ID NO. 20TATGACTGCATGCGSEQ ID NO. 21ACCGGATAACGGAASEQ ID NO. 22ATAATATTGCCTCCSEQ ID NO. 23TCTTGAACGACTAASEQ ID NO. 24TTAACTTACGGTCASEQ ID NO. 25CATGCTAGCGATCGSEQ ID NO. 26TAAGCAACAGCCGASEQ ID NO. 27ATGTAGGTAATCGGSEQ ID NO. 28CAGCAGCGCTGACASEQ ID NO. 29TATGTACGTAGAGASEQ ID NO. 30CAGATCGCTCTGCGSEQ ID NO. 31TAACCTAACTTCACSEQ ID NO. 32CTCCTAACACAGCCSEQ ID NO. 33ATACCGTTGGAGGASEQ ID NO. 34CTTATATCGTCGGCSEQ ID NO. 35AAGTACTGATCCAASEQ ID NO. 36TACTGACTATCGAGSEQ ID NO. 37ATTGCGTCCATTAGSEQ ID NO. 38TTGCTACCATTGAASEQ ID NO. 39AAGCGCACCGGAACSEQ ID NO. 40ACCAACTTAGACAASEQ ID NO. 41TCAGCGTGTCGTCASEQ ID NO. 42ATTCCGCTCTGAGGSEQ ID NO. 43ACCGATGTGCCGGASEQ ID NO. 44CTAAGGAAGGACCGSEQ ID NO. 45TTCTGAGGCGATGGSEQ ID NO. 46ACATCTTCTTGAACSEQ ID NO. 47ACAGACGATAGTCASEQ ID NO. 48ACTATCTGCAGTCGSEQ ID NO. 49CATGTGTACTAGCASEQ ID NO. 50CTCTGGCTTGCAGGSEQ ID NO. 51ACGGCCTTCAACCASEQ ID NO. 52CCATATTGAAGCACSEQ ID NO. 53TCGATTGGTAATGGSEQ ID NO. 54TTCATGGTAGAGAGSEQ ID NO. 55ACGCTCCTAACGGCSEQ ID NO. 56CTCAATTAGAGACCSEQ ID NO. 57TCGATCTTCGGCAGSEQ ID NO. 58CACATCGACAGCAGSEQ ID NO. 59AACCTGCTATCCGGSEQ ID NO. 60ATCCGGCGATGACCSEQ ID NO. 61TCCATGTGTCCAGGSEQ ID NO. 62ATCGATGTCATAGCSEQ ID NO. 63CTCCGCATGCTTCCSEQ ID NO. 64CACGTAGCTAGACGSEQ ID NO. 65AACATGGCACGGAASEQ ID NO. 66ATGCTATCTAGACASEQ ID NO. 67ACCAATCCGTATAASEQ ID NO. 68CTCGTGCCTCAGGASEQ ID NO. 69ATGCAATTACACCGSEQ ID NO. 70CCTCCATCCATGGASEQ ID NO. 71ACGGCTCTTAGGCCSEQ ID NO. 72TCCAACTACCGAACSEQ ID NO. 73TATGAGAGCTATGASEQ ID NO. 74CAAGGCGTGGTTACSEQ ID NO. 75ATTATCGGATTCGGSEQ ID NO. 76ATCCTAAGTGGAGGSEQ ID NO. 77ATACGATGCTGCGCSEQ ID NO. 78AAGCCTTAGGCCGCSEQ ID NO. 79AATGTGTGGATACCSEQ ID NO. 80ATGCTGCAACGGCASEQ ID NO. 81CTTGATAAGCCAGGSEQ ID NO. 82CATGGAACTCCGGCSEQ ID NO. 83TAAGGACCTCATAASEQ ID NO. 84TCTGTCTATGCAGCSEQ ID NO. 85CAACCAACCTATAGSEQ ID NO. 86CCGGTAGAATGGCCSEQ ID NO. 87TCCGAAGTTGTAGASEQ ID NO. 88TCGATGCTATATGASEQ ID NO. 89ACAAGGTTCCGCCGSEQ ID NO. 90CCACGTTCCGGTAGSEQ ID NO. 91ACTTAAGGTCCAACSEQ ID NO. 92TCCAGAAGGCAACASEQ ID NO. 93CTGCTCATGAGAGASEQ ID NO. 94ATGAATGTTGCCACSEQ ID NO. 95CTCTCAAGTCCTCCSEQ ID NO. 96ATTCCAGTTCTTGC, or a combination of one or more thereof.
[0033] Preferably, the combination consists of at least 10 oligonucleotides, at least 20 oligonucleotides, at least 30 oligonucleotides, at least 40 oligonucleotides, at least 50 oligonucleotides, at least 60 oligonucleotides, at least 70 oligonucleotides, at least 80 oligonucleotides, at least 90 oligonucleotides, or a combination of all the oligonucleotides of SEQ ID No. 1 to 96.
[0034] In one aspect, the barcode sequence is selected from the group comprising, or consisting of, SEQ ID NO. 98TCGTACATAGCGCASEQ ID NO. 99CAAGGAACCGAACCSEQ ID NO. 100TCAGTCACGCGTGCSEQ ID NO. 101CACACGATACTCACSEQ ID NO. 102AAGGAGAAGAATGCSEQ ID NO. 103CAGTGGCCTAAGCASEQ ID NO. 104ACCGTCTTCCTTGCSEQ ID NO. 105CCTCATTCGTTGAGSEQ ID NO. 106ACGATATGCTCTGCSEQ ID NO. 107CCACCGTTAAGACASEQ ID NO. 108ACTACTCGATCTAGSEQ ID NO. 109AACACTTGGCCACCSEQ ID NO. 110ATCACATCTATCGCSEQ ID NO. 111CTCCAATATTCAACSEQ ID NO. 112CCAATAGACTTCCGSEQ ID NO. 113ACCGTAACGTGGACSEQ ID NO. 114ATAAGGACCGGAGGSEQ ID NO. 115ACCTTCTCACGGCCSEQ ID NO. 116CAACGTGGACCTACSEQ ID NO. 117AACTGCTACGCACCSEQ ID NO. 118AACTTCCGGCGTGCSEQ ID NO. 119CATCGTAGCAGAGCSEQ ID NO. 120AACAACCAAGCGGCSEQ ID NO. 121TAACCGGCCGCAACSEQ ID NO. 122ATAACCATCCGAACSEQ ID NO. 123CTAGCTGTGTCAGCSEQ ID NO. 124CTAACCTTCATTGCSEQ ID NO. 125CTCACCTAGGCCACSEQ ID NO. 126TCATGCGCATGTAGSEQ ID NO. 127TCGCAATTACCTAASEQ ID NO. 128CATCTACGCTACGCSEQ ID NO. 129TTCCACCGATGTGGSEQ ID NO. 130TAGCTTCTCCAGGCSEQ ID NO. 131TTCCAACCGGACAGSEQ ID NO. 132TTGACCTTGCTGGASEQ ID NO. 133ACCTTGTCCAATACSEQ ID NO. 134CTTAGAATTAGTGGSEQ ID NO. 135AAGTGACCGACGGCSEQ ID NO. 136TCGGTTATCTTCGASEQ ID NO. 137CTAGCTTGTTGTAGSEQ ID NO. 138TTAAGCCGAACAAGSEQ ID NO. 139TTCCGATGTCCGGCSEQ ID NO. 140ATGAATCAGCCGCCSEQ ID NO. 141TCCTAACAAGTGCGSEQ ID NO. 142CAATAGCAACACCASEQ ID NO. 143TATACGCGACATCGSEQ ID NO. 144CCGGATGGTGGTACSEQ ID NO. 145CCTACGGTGTTCGGSEQ ID NO. 146CTCGCACCTGGTAASEQ ID NO. 147TCGCCTGAACTACCSEQ ID NO. 148TCACTCATACATAGSEQ ID NO. 149CTTGGTCTCCGGCASEQ ID NO. 150TTGATTACCGTAACSEQ ID NO. 151TTGATCCGTGAAGCSEQ ID NO. 152ATTCTTCAGGAAGASEQ ID NO. 153AAGCCAGGTAAGGCSEQ ID NO. 154CCGCGGAAGGATAASEQ ID NO. 155CATAATTCTATGCCSEQ ID NO. 156ATGAACTTGGCACASEQ ID NO. 157CTTAAGCTTAAGAGSEQ ID NO. 158TCCTCGGAATGACASEQ ID NO. 159AATCTTACCTAAGGSEQ ID NO. 160ACCACTGGTTGCCGSEQ ID NO. 161CTCCGTTCTGTACCSEQ ID NO. 162CCTAACTCGCCAGASEQ ID NO. 163AATAGTGTTCAACCSEQ ID NO. 164CTACGAGTCATAGASEQ ID NO. 165ACACACGACATCGCSEQ ID NO. 166AACTTACACGTTGGSEQ ID NO. 167AACACAACCTTCCGSEQ ID NO. 168ATACCTTAATCTCCSEQ ID NO. 169TAACCGCAATAGGASEQ ID NO. 170CCGAACCATAGCCGSEQ ID NO. 171TTCATATAAGGCGCSEQ ID NO. 172AATACCAGCCGCGGSEQ ID NO. 173CTACGGACTAAGGCSEQ ID NO. 174TAATCGCCTACAGASEQ ID NO. 175TTAGCCTAACTCGCSEQ ID NO. 176TAGTGTGGTAGACCSEQ ID NO. 177CAACGCCATAGACCSEQ ID NO. 178TAATCCATCAACGCSEQ ID NO. 179CCGGTGTGCGCTAASEQ ID NO. 180TACCGGTATGGTCCSEQ ID NO. 181CACGTATAGATCGCSEQ ID NO. 182AACCTAAGCAATCCSEQ ID NO. 183CTAGTGTTCGATCCSEQ ID NO. 184CCTACACTTGAGCCSEQ ID NO. 185CAGCATCATACTGCSEQ ID NO. 186AATACGGAGACGGCSEQ ID NO. 187ATTGAACTCCATCCSEQ ID NO. 188AAGAGTCGGATAGGSEQ ID NO. 189AATGCAAGACAACCSEQ ID NO. 190TTCATTGGAAGAGASEQ ID NO. 191TTCTAGGACGGCCGSEQ ID NO. 192CTCAACCACGTTCCSEQ ID NO. 193CTGGTCTGTCTTCCSEQ ID NO. 194TTCAAGTCAGTGCCSEQ ID NO. 195CACACGACAGAGCGSEQ ID NO. 196CCTAAGCCTGTTCGSEQ ID NO. 197AAGCTGGCGCCTAGSEQ ID NO. 198ACGCCTGTTGGAGGSEQ ID NO. 199TCTCCATTATACAGSEQ ID NO. 200TCCGCCGCACTCAASEQ ID NO. 201TATCGATCATCTACSEQ ID NO. 202TCTTGTACCGGACCSEQ ID NO. 203AACGGCTGAACGACSEQ ID NO. 204TTCCGAAGCCTAAGSEQ ID NO. 205TTCTAGATAGGTACSEQ ID NO. 206AATTCTCTGGCGCCSEQ ID NO. 207ATGACTGAGAGTGASEQ ID NO. 208TAATTAGGTCCTCGSEQ ID NO. 209AACCGTCCAAGCGCSEQ ID NO. 210ACATTATTGGCCAASEQ ID NO. 211TCCACAGAACCGGCSEQ ID NO. 212TAGTGATCTGTAGCSEQ ID NO. 213CATTCACGCTTGAASEQ ID NO. 214CACACACGAGATGCSEQ ID NO. 215CCATGACCACCGGASEQ ID NO. 216CAACAAGACCTAAGSEQ ID NO. 217AAGGAACCTCTCGGSEQ ID NO. 218CATAACCGGCACAASEQ ID NO. 219CCAAGTGGTACAGCSEQ ID NO. 220TTCCTTACCGCGCGSEQ ID NO. 221ATACTAGTAGTCAGSEQ ID NO. 222TATATGCTCGCTCASEQ ID NO. 223TCCGCGATCTCAACSEQ ID NO. 224TATTACCGCTCAGGSEQ ID NO. 225ATCGAGCAGTGTACSEQ ID NO. 226CATGCGTAGCGCACSEQ ID NO. 227TCCAACAGTGTTGGSEQ ID NO. 228CACTTGGACCATCCSEQ ID NO. 229TTACCTGGCCATCCSEQ ID NO. 230CATCATGAGAGCGASEQ ID NO. 231ATTCCGAATGTAACSEQ ID NO. 232CCGGTATATCCGGASEQ ID NO. 233ACTGACGTGGAACCSEQ ID NO. 234TTAAGCTGCCTTAASEQ ID NO. 235CTCCGCAGTTGGAGSEQ ID NO. 236TCCTAATTCTAAGGSEQ ID NO. 237TACGATCGTGATCASEQ ID NO. 238ATAACCAACTAGGASEQ ID NO. 239ATCCTTATACGTGCSEQ ID NO. 240ACTTCCAGGCCTGASEQ ID NO. 241TCTGCACATCACGCSEQ ID NO. 242ACAGACGCATCACGSEQ ID NO. 243TAAGTTGTGTGTCCSEQ ID NO. 244ACTCAACGGCAAGGSEQ ID NO. 245TAGCACATACTCGASEQ ID NO. 246ATATCACAGTCGCASEQ ID NO. 247ATATACACATGAGCSEQ ID NO. 248ACGTACGCGCATGCSEQ ID NO. 249CAAGCCAGTTCCGGSEQ ID NO. 250CCGAACAACCGTGASEQ ID NO. 251CACACACATAGCGASEQ ID NO. 252ACAGCGTACGCGCASEQ ID NO. 253TCAACGCTCGGAGCSEQ ID NO. 254ACTACGATCATAGASEQ ID NO. 255CAGGCGTAAGTGCCSEQ ID NO. 256ACATATCATGAGCASEQ ID NO. 257AACTCCGTACAAGCSEQ ID NO. 258CTGTCGAGTAGCAGSEQ ID NO. 259TAGACGATAGTACASEQ ID NO. 260CATGCTCACGCAGCSEQ ID NO. 261ATCACGAGCGACGCSEQ ID NO. 262CAGCCTGGCTTAGGSEQ ID NO. 263CCAGAGGTTGCCAGSEQ ID NO. 264TACTATGCACGACGSEQ ID NO. 265ACGCTGACGTGCGASEQ ID NO. 266CAGCACATGTACAGSEQ ID NO. 267ACACGCATCATACGSEQ ID NO. 268CTGTACACGTCTACSEQ ID NO. 269CACACAGTGAGTACSEQ ID NO. 270TAGCACGCGATGAGSEQ ID NO. 271TCCAATAGGTCCAGSEQ ID NO. 272CTCCTGCGGTCCAASEQ ID NO. 273ATGGCAACGGCTCASEQ ID NO. 274CCGAGGCTTCCTAGSEQ ID NO. 275ACGACAGTAGTCGCSEQ ID NO. 276TTATAACTGGCAACSEQ ID NO. 277ACTCTGAGTATCAGSEQ ID NO. 278TTATGTCTTAACGGSEQ ID NO. 279ATGCGACTCGAGACSEQ ID NO. 280CCTCAATTCCGGCGSEQ ID NO. 281AATGGTTGGCAGAGSEQ ID NO. 282ATGCAAGCGATTCCSEQ ID NO. 283TATTGAGTCCGAGASEQ ID NO. 284CCGCGTATTGCAACSEQ ID NO. 285TCATTCTACTTGGCSEQ ID NO. 286ATCATCGCGCGACGSEQ ID NO. 287CAGAGACACTCGCASEQ ID NO. 288AAGAAGAGTGGACCSEQ ID NO. 289ACACACGCTGCGACSEQ ID NO. 290TCCGCAAGTCATAGSEQ ID NO. 291AATGTCCTCAGAACSEQ ID NO. 292CCGAGTGTAAGTCCSEQ ID NO. 293ACACGTACTCATACSEQ ID NO. 294TTCTCTCCGGTTCASEQ ID NO. 295AATCCGTGTCTTAASEQ ID NO. 296CCTTGCCAGAGTGGSEQ ID NO. 297CTTACTGACAAGCCSEQ ID NO. 298TCCTCCGATGCCAGSEQ ID NO. 299TTCTTCTCCAGAAGSEQ ID NO. 300CACCTGTTGGTTAASEQ ID NO. 301CTCAGTCACACGAGSEQ ID NO. 302CCTTATCTACGACCSEQ ID NO. 303AACCGAACGCTTAASEQ ID NO. 304ACAGGTAGGAAGGASEQ ID NO. 305CAGCATAGTCTGACSEQ ID NO. 306ATTCGGTCCGATGCSEQ ID NO. 307ACTAACCTGGCCGGSEQ ID NO. 308CATTCCTCTAGCCGSEQ ID NO. 309CTTAATCCATATCCSEQ ID NO. 310AACAAGAGGTGTGGSEQ ID NO. 311TATCATATGAGTCGSEQ ID NO. 312ACGAGTATGTATGCSEQ ID NO. 313TCGCGGATTAGGCASEQ ID NO. 314CTACTGTGACACACSEQ ID NO. 315CAACCGTGTGAACCSEQ ID NO. 316CAGCACTAGCTGCASEQ ID NO. 317ACACGACATGCACGSEQ ID NO. 318ACTCATGTCAGACASEQ ID NO. 319ACACTATACACGAGSEQ ID NO. 320ACTTGAGCAACCGGSEQ ID NO. 321TTAGACGCCGGTAASEQ ID NO. 322CTGTGTAGGATTAASEQ ID NO. 323CACGCATCAGTCAGSEQ ID NO. 324CAACCATTGGCAAGSEQ ID NO. 325TCATCTGCGCACACSEQ ID NO. 326TAGGCCTCTGATGGSEQ ID NO. 327TAGCATGTCAGCACSEQ ID NO. 328TCCTGGACGATGCCSEQ ID NO. 329ATATCACATGACAGSEQ ID NO. 330ACTACGTCGACACGSEQ ID NO. 331AATCCAACTCGAAGSEQ ID NO. 332AAGGATGCCAGTGGSEQ ID NO. 333ACCGCCTTGGCTAGSEQ ID NO. 334CTGTGATTCGGACCSEQ ID NO. 335CATGAGACGTCGCGSEQ ID NO. 336AATCCGCCTAATGGSEQ ID NO. 337AATGCCGTATTGACSEQ ID NO. 338CTCCGAAGAACCGGSEQ ID NO. 339ACTGATTAGAATCGSEQ ID NO. 340ACTCGCCACCGAAGSEQ ID NO. 341CCGGTATTGTTACASEQ ID NO. 342TTGGTGAGAACCGCSEQ ID NO. 343CAGATCTCAGACACSEQ ID NO. 344CCTTAACACGGAGASEQ ID NO. 345ATGTCTAGTATAGCSEQ ID NO. 346CCAGAGGTAATGGASEQ ID NO. 347CTGTCACACAGTCASEQ ID NO. 348TCACTGAGTGCTGASEQ ID NO. 349TCCTATGGTAGGAGSEQ ID NO. 350AAGGAACAATACACSEQ ID NO. 351TAAGCCGGAATTAGSEQ ID NO. 352ATAATGCTCCAGAASEQ ID NO. 353ATTCTTGCGAAGCGSEQ ID NO. 354ACATGCTATACTACSEQ ID NO. 355AAGCCAACAGTTGGSEQ ID NO. 356TATGAGCACACGACSEQ ID NO. 357CACTAGATATCACASEQ ID NO. 358TCAATTAAGGCACGSEQ ID NO. 359TCCTCCTCTTGTGASEQ ID NO. 360TACTACTATGTGACSEQ ID NO. 361TATCTTCCACCTGASEQ ID NO. 362CACTTACCTTACAASEQ ID NO. 363TCTAGCATCGACGASEQ ID NO. 364ACACGAGCACGAGCSEQ ID NO. 365ACAAGGCAACAGCCSEQ ID NO. 366CACTACGCGCTCACSEQ ID NO. 367TAACCGTCTCTGGCSEQ ID NO. 368CCAACTCGCCAGGASEQ ID NO. 369CTTCGGTACACTAASEQ ID NO. 370CCACACCTTCAGGCSEQ ID NO. 371CCACTTCTTGTTGGSEQ ID NO. 372CTAGTCTTGGTAGGSEQ ID NO. 373TCAGAATCCTCGCCSEQ ID NO. 374CTAGGAGCTTGAACSEQ ID NO. 375CCGGCGCAATTCAGSEQ ID NO. 376ACCGGTCGTTCTGCSEQ ID NO. 377CATTAAGGTATCGCSEQ ID NO. 378CTTGGAGTGGCACASEQ ID NO. 379AACCTTCTAGAGAASEQ ID NO. 380CATGGAATGTGTAASEQ ID NO. 381CTGTTCGGAGGTCASEQ ID NO. 382ATCGTGTGTGGCCASEQ ID NO. 383TTGTCGGTTGAGGCSEQ ID NO. 384CTGGCGTTGAGTCGSEQ ID NO. 385AATGACCACGACCGSEQ ID NO. 386ATGTCTACAGACGASEQ ID NO. 387TCTCCGAGGTAGGCSEQ ID NO. 388ACGACGACTGACAGSEQ ID NO. 389ACTATGACTGCGCASEQ ID NO. 390CCGCCTTCTAACACSEQ ID NO. 391CACACGCAGTATCASEQ ID NO. 392CCAGCTTAGGAAGASEQ ID NO. 393ACACGATCTGACGASEQ ID NO. 394AAGGCAAGTTGTGASEQ ID NO. 395AACTCTTCGGTAGGSEQ ID NO. 396CCGTCCTTCTTCACSEQ ID NO. 397CTTGGTTGGTTCCGSEQ ID NO. 398TACCTTGGAGTTGGSEQ ID NO. 399ACCGGCGGTATAGGSEQ ID NO. 400CACGACGACTCAGASEQ ID NO. 401TCTTGTGAGGTTCGSEQ ID NO. 402AACGCTTCCGGTCCSEQ ID NO. 403CTACTGCATAGTAGSEQ ID NO. 404ACACAGCACGATAGSEQ ID NO. 405CCTACAGGTCGAGGSEQ ID NO. 406ATGCACGATCGAGCSEQ ID NO. 407CTTGATACCGGCGCSEQ ID NO. 408CCTCGGTTGAAGCCSEQ ID NO. 409CAACACCGGCTTGGSEQ ID NO. 410CTCTGGACCGATCASEQ ID NO. 411CCGTGAACCTTGCGSEQ ID NO. 412TCATGCTCGAGACASEQ ID NO. 413ACCGGAACCTATGGSEQ ID NO. 414TTCCAAGGTTGACASEQ ID NO. 415ACGAGAGTGCTGAGSEQ ID NO. 416CTCAACCGGACGGASEQ ID NO. 417AATCAAGTGTGAACSEQ ID NO. 418CATACGGCCTCTAASEQ ID NO. 419ACTAGAGCGTGTGASEQ ID NO. 420TCAGAGTAGTGCAGSEQ ID NO. 421CAACAACCTTGAGGSEQ ID NO. 422ACAATTAGCCTTGGSEQ ID NO. 423TCACTATGCGATACSEQ ID NO. 424TACTCTCATCTAGCSEQ ID NO. 425TCTCTCTCAATTGGSEQ ID NO. 426CCAGTGAGGTATACSEQ ID NO. 427CTCCTACTTAATGCSEQ ID NO. 428ATAGGCCTTCGAGGSEQ ID NO. 429CATTGGTGGTGTCCSEQ ID NO. 430CATACTATGTAGACSEQ ID NO. 431ATGGATCGTTAGGASEQ ID NO. 432CCAACCGAACGGCASEQ ID NO. 433ACGTACATCTATCGSEQ ID NO. 434TACGAGTGTCTCAGSEQ ID NO. 435ATAACAACGAAGCCSEQ ID NO. 436TACGATAGTACAGCSEQ ID NO. 437ATTGCGGTTAATCASEQ ID NO. 438CAGTTGCGAAGTGGSEQ ID NO. 439TTAACACCTCAAGGSEQ ID NO. 440CACGACATGCGAGCSEQ ID NO. 441TCCAGAACTGAGGCSEQ ID NO. 442TATTCCTAGTTCGASEQ ID NO. 443AACACCTTCTCCGCSEQ ID NO. 444TAACCTGTTCAAGASEQ ID NO. 445AAGCTATGTGCAACSEQ ID NO. 446CTCTTAATCCTTGASEQ ID NO. 447ATCAGCAGACTAGCSEQ ID NO. 448CAAGGAAGCATTGGSEQ ID NO. 449ACCTCGAAGTTCAASEQ ID NO. 450ATAGCACTCACACGSEQ ID NO. 451ATCTAGCTGCACGASEQ ID NO. 452TCTGCTGCAGAGCGSEQ ID NO. 453TAACAGTCCGTTCASEQ ID NO. 454CATCTATCTGCGCGSEQ ID NO. 455TTGACCAAGATTACSEQ ID NO. 456TCTCACATCACGAGSEQ ID NO. 457ACACATCTCTACGASEQ ID NO. 458AACAGAATAGGCGGSEQ ID NO. 459AACACCTAAGGAAGSEQ ID NO. 460AACACCTGGTTGAASEQ ID NO. 461AAGACGAATAGGAASEQ ID NO. 462CAAGTTCAGAATGGSEQ ID NO. 463CCAAGCTTAGTGCGSEQ ID NO. 464AACGCTCAATGTGGSEQ ID NO. 465TCATCTAGGTGGAASEQ ID NO. 466CCTGTGATTACACCSEQ ID NO. 467TCAGAACAGGTTAASEQ ID NO. 468ATATTGTAAGGTGGSEQ ID NO. 469TTAAGTGACCGCGCSEQ ID NO. 470AATCCTGAATCCAASEQ ID NO. 471CTATAGCGCACTACSEQ ID NO. 472TCTAGACACTATCGSEQ ID NO. 473TCAATGGCAGGCCGSEQ ID NO. 474AAGTGGTATTGACASEQ ID NO. 475TTGGACATTCCGGCSEQ ID NO. 476CAGCATCGAGAGCGSEQ ID NO. 477CAAGACCTTGGCCASEQ ID NO. 478CCGGTTCCTAGTGASEQ ID NO. 479ACAACTTAGCCTAASEQ ID NO. 480TCGTGTCGCATGCASEQ ID NO. 481CACACAGACGCTCG, or a combination of one or more thereof.
[0035] Preferably, the combination consists of at least 10 oligonucleotides, at least 20 oligonucleotides, at least 30 oligonucleotides, at least 40 oligonucleotides, at least 50 oligonucleotides, at least 60 oligonucleotides, at least 70 oligonucleotides, at least 80 oligonucleotides, at least 90 oligonucleotides, at least 100 oligonucleotides, at least 150 oligonucleotides, at least 200 oligonucleotides, at least 250 oligonucleotides, at least 300 oligonucleotides, or a combination of all 384 oligonucleotides of SEQ ID No. 98 to 481.
[0036] As used herein, the term "sequencing adapter" refers to an oligonucleotide sequence that can be used in subsequent sequencing steps (so-called sequencing adapters), or to primers that are used to amplify a subset of fragments prior to sequencing may contain parts within their sequence that introduce sections that can later be used in the sequencing step, for instance by introducing through an amplification step a sequencing adapter or a capturing moiety in an amplicon that can be used in a subsequent sequencing step. Depending also on the sequencing technology used, amplification steps may be omitted.
[0037] Any commercially available sequencing adapter can be used. In one aspect, the sequencing adapter comprises, or consists of, SEQ ID No. 97 (CTA CAC GAC GCT CTT CCG ATC T ).
[0038] As used herein, a "unique molecular identifier" or UMI is a complex indices added to sequencing libraries before any PCR amplification steps, enabling the accurate bioinformatic identification of PCR duplicates thus enabling to remove PCR duplicates. In one aspect, the UMI is an oligonucleotide sequence consisting of a sequence (N)n(V)m, wherein N is any nucleotide selected from A, T, C and G; V is any nucleotide selected from A, C and G; n is an integer selected from 1 to 20, and m is an integer selected from 1 to 20. In a preferred aspect, the UMI is an oligonucleotide sequence consisting of NNNNNNNNNVVVVV (SEQ ID No. 482).
[0039] As used herein, an "mRNA capture sequence" is an oligonucleotide sequence that specifically hybridizes to mRNAs. In one aspect, the mRNA capture sequence is a poly-T sequence. In a preferred aspect, the mRNA capture sequence is a poly-T sequence followed by at least one V and one N, wherein N is any nucleotide selected from A, T, C and G and V is any nucleotide selected from A, C and G. Example of an mRNA capture sequence consists of TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN (SEQ ID No. 483).
[0040] "Multiplex sequencing" refers to a sequencing technique that allows for processing a large number of samples on a high-throughput instrument. For multiplex sequencing, individual "barcode" sequences of the invention are added to each sample so that nucleotide sequences from different samples can be distinguished by the unique barcode sequences embedded in each sample. With this technique, multiple DNA or RNA samples can be pooled, processed, sequenced, and analyzed simultaneously.
[0041] The present invention further provides a set of oligonucleotide molecules, each molecule comprising, from 5' to 3', a) a sequencing adaptor, b) a barcode sequence independently selected from the group consisting of SEQ ID: 1 to 96 or a barcode sequence independently selected from the group consisting of SEQ ID: 98 to 481, and c) an mRNA capture sequence.
[0042] In one aspect, the set of oligonucleotide molecules of the invention, further comprises d) a UMI as described herein.
[0043] Also provided is the use of a set of oligonucleotide molecules of the invention, or of a set of oligonucleotide molecules of the invention, or of a barcode oligonucleotide sequence, or a combination of one or more thereof, of the invention, in a sequencing method.
[0044] The set of oligonucleotide molecules of the invention, or of a barcode oligonucleotide sequence, or a combination of one or more thereof, may be linked by attachment to a solid support (e.g. a bead). A solution of soluble beads (e.g. superparamagnetic beads or styrofoam beads) may be functionalized to enable attachment of two or more oligonucleotide molecules of the invention, set of oligonucleotide molecules of the invention, barcode oligonucleotide sequences, or a combination of one or more thereof. This functionalization may be enabled through chemical moieties (e.g. carboxylated groups), and / or protein-based adapters (e.g. streptavidin) on the beads. The functionalized beads may be brought into contact with a solution of the above-described molecules under conditions which promote the attachment of two or more molecules to each bead in the solution. Optionally, the molecules are attached through a covalent linkage, or through a (stable) non-covalent linkage such as a streptavidin-biotin bond, or a (stable) oligonucleotide hybridization bond.
[0045] The present invention further encompasses a method for providing a cDNA library, the method comprising the step of Providing a plurality of RNA samples obtained from a biological sample (step a).
[0046] As used herein, the term "biological sample" refers to a tissue (e.g., tissue biopsy), organ, cell (including a cell maintained in culture), cell lysate (or lysate fraction), biomolecule derived from a cell or cellular material (e.g. a polypeptide or nucleic acid), or body fluid from a subject. Non-limiting examples of body fluids include blood, urine, plasma, serum, tears, lymph, bile, cerebrospinal fluid, interstitial fluid, aqueous or vitreous humor, colostrum, sputum, amniotic fluid, saliva, anal and vaginal secretions, perspiration, semen, transudate, exudate, and synovial fluid.
[0047] The RNA samples can be obtained from any techniques know in the art. According to a particular aspect, the RNA samples are mRNA samples that can be cell lysates, total DNA / RNA eluate, blood and FFPE tissues.
[0048] The method further comprises a step (b) of contacting separately each RNA sample with a set of oligonucleotide molecules of the invention, or of a library of the invention, or of a barcode oligonucleotide sequence or a combination of one or more thereof, of the invention, under annealing conditions.
[0049] For examples, RNA samples are thawed on ice, transferred to the corresponding wells of the Oligo-dT primer plat, the plate is then sealed with the AluSeal and placed it in a thermocycler at 65°C for 5 min and immediately put on ice.
[0050] The method further comprises a step (c) of incubating separately each sample under reverse transcription reaction conditions.
[0051] For example, the RT reaction mix is prepared according to commercial manufacture instruction. Any RT enzyme and buffer commercially available can be used, such as e.g. Lucigen's ERT12910K, ThermoFisher's 18064014 and NEB's M0368S among others.For example:
[0052] RT Mix Per well (µL) RT reaction buffer5.0RT reaction enzyme0.4ddH2O4.6TOTAL10.0 • Incubate RT reaction mix in thermocycler with the following program:
[0053] Step Temperature, °C Time Incubation4250 minInactivation7010 minKeep4pause
[0054] The method further comprises a step (d) of pooling together all the cDNA:RNA sample.
[0055] The method further comprises a step (e) of proceeding to second strand synthesis under synthesis conditions such as. This second strand synthesis can be generated by any method known in the art. In one aspect, the second strand synthesis method is selected from the group comprising PCR amplification and nick translation, or a combination thereof.
[0056] The method further comprises a step (f) of proceeding with tagmentation, and / or end-repair and ligation and amplification under suitable conditions such as, e.g. the conditions described in the examples, so as to obtain a cDNA library.
[0057] Examples of RNAs include but are not limited to: mRNA, amplicons, rRNA, tRNA, nRNA, siRNA, snRNA, snoRNA, scaRNA, microRNA, dsRNA, ncRNA (e.g. lncRNA), ribozyme, riboswitch and viral RNA (e.g., retroviral RNA).
[0058] Examples of RNAs include but are not limited to: mRNA, amplicons, rRNA, tRNA, nRNA, siRNA, snRNA, snoRNA, scaRNA, microRNA, dsRNA, ncRNA (e.g. lncRNA), ribozyme, riboswitch and viral RNA (e.g., retroviral RNA).
[0059] The present invention further encompasses a method for sequencing RNA, the method comprising the steps of a) Providing a cDNA library obtained by the method described herein; and b) Proceeding to the sequencing under suitable conditions such as, e.g. those defined by NGS sequencing providers, which are also known in the art.
[0060] In one aspect, the set of oligonucleotide molecules of the invention may be linked by attachment to a solid support (e.g. a bead).
[0061] Also contemplated is one or more kits for performing one or more methods according to the invention. The one or more kits comprising i) a set of oligonucleotide molecules of the invention, or of a barcode oligonucleotide sequence, or a combination of one or more thereof of the invention, ii) a support for sample preparation, such as a 96-well plate or 384-well plate or a set of 96-well plates combining 384 oligos with different primers, and iii) reagents for sequencing.
[0062] The kit(s) can comprise various molecular biology reagents, including DNA polymerases, RNA polymerases, reverse-transcriptases, DNA ligases, RNA ligases, transposases, viral integrase, CRISPR / Cas9, zinc finger nucleases, transcription activator-like effector nucleases, exonucleases, endonucleases, polynucleotide kinases, nucleotides, oligonucleotides, modified oligonucleotides, optimized buffers and cell lysis reagents.
[0063] Further contemplated is the use of the kit(s) of the invention, or of the set of oligonucleotide molecules of the invention, or of a library of the invention, or of a barcode oligonucleotide sequence or a combination of one or more thereof, of the invention, in a single-cell RNA profiling method. Any single cell RNA profiling known in the art are considered. Preferably, the single-cell RNA profiling method is the Bulk RNA Barcoding and sequencing (BRB-seq) method described in (Alpern et al., 2019) or an adaptation thereof.
[0064] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications without departing from the spirit or essential characteristics thereof. The invention also includes all of the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations or any two or more of said steps or features. The present disclosure is therefore to be considered as in all aspects illustrated and not restrictive, the scope of the invention being indicated by the appended Claims, and all changes which come within the meaning and range of equivalency are intended to be embraced therein. Various references are cited throughout this Specification, each of which is incorporated herein by reference in its entirety. The foregoing description will be more fully understood with reference to the following Examples.EXAMPLESSecond-strand synthesis
[0065] Double-stranded cDNA was generated by nick translation (Gubler and Hoffman, 1983). For that, a mix containing 2 µL of RNAse H (NEB, #M0297S), 1 µL of Escherichia coli DNA ligase (NEB, #M0205 L), 5 µL of E. coli DNA Polymerase (NEB, #M0209 L), 1 µL of dNTP (0.2mM), 10 µL of 5× Second Stand Buffer (100 mM Tris-HCl (pH 6.9) (AppliChem, #A3452); 25 mM MgCl2 (Sigma, #M2670); 450 mM KCl (AppliChem, #A2939); 0.8 mM β-NAD; 60 mM (NH4)2SO4 (Fisher Scientific Acros, #AC20587); and 11 µL of water was added to 20 µL of ExoI-treated first-strand reaction on ice. The reaction was incubated at 16 °C for 2.5 h or overnight. Full-length double-stranded cDNA was purified with 30 µL (0.6×) of AMPure XP magnetic beads (Beckman Coulter, #A63881) and eluted in 20 µL of water. Alternatively, double stranded cDNA was generated by PCR amplification following template switch oligo assisted first strand synthesis. PCR amplification was performed in 50 µL reaction using NEBNext High-Fidelity 2X PCR Master Mix (NEB, #M0541 L) and 1 µL of primers CTACACGACGCTCTTCCGATCT (SEQ ID No. 484) and AAGCAGTGGTATCAACGCAGAG (SEQ ID No. 485) (10 µM, IDT).Library preparation and sequencing
[0066] The sequencing libraries were prepared by tagmentation of 1-50 ng of full-length double-stranded cDNA. Tagmentation was done either with Illumina Nextera XT kit (Illumina, #FC-131-1024) following the manufacturer's recommendations or with in-house produced Tn5 preloaded with dual (Tn5-A / B) or same adapters (Tn5-B / B) under the following conditions: 1 µL (11 µM) Tn5, 4 µL of 5× TAPS buffer (50 mM TAPS (Sigma, #T5130), and 25 mM MgCl2 (Sigma, #M2670)) in 20 µL total volume. The reaction was incubated 10 min at 55 °C followed by purification with DNA Clean & Concentrator-5 kit (Zymo Research) and elution in 21 µL of water. After that, tagmented library (20 µL) was PCR amplified using 25 µL NEBNext High-Fidelity 2X PCR Master Mix (NEB, #M0541 L), 2.5 µL of each of P5 and P7 indexing primers bearing library indexing seqeuence (5 µM, IDT), using the following program: incubation 72 °C-3 min, denaturation 98 °C-30 s; 10 cycles: 98 °C-10 s, 63 °C-30 s, 72 °C-30 s; final elongation at 72 °C-5 min. The fragments ranging 200-1000 bp were size-selected using AMPure beads (Beckman Coulter, #A63881) (first round 0.5× beads, second 0.7×). The libraries were profiled with High Sensitivity NGS Fragment Analysis Kit (Agilent, #DNF-474) and measured with Qubit dsDNA HS Assay Kit (Invitrogen, #Q32851) prior to pooling and sequencing using the Illumina insturments (Nextseq, Novaseq, Miseq, iSeq or Hiseq4000). The read1 sequencing was performed for 28 cycles and read2 for 60-90 cycles depending on the experiment.REFERENCES
[0067] Alpern, D., Gardeux, V., Russeil, J., Mangeat, B., Meireles-Filho, A.C.A., Breysse, R., Hacker, D., and Deplancke, B. (2019). BRB-seq: ultra-affordable high-throughput transcriptomics enabled by bulk RNA barcoding and sequencing. Genome Biol. 20, 71. https: / / doi.org / 10.1186 / s13059-019-1671-x. Bush, E.C., Ray, F., Alvarez, M.J., Realubit, R., Li, H., Karan, C., Califano, A., and Sims, P.A. (2017). PLATE-Seq for genome-wide regulatory network analysis of high-throughput screens. Nat. Commun. 8. https: / / doi.org / 10.1038 / s41467-017-00136-z. Gubler, U., and Hoffman, B.J. (1983). A simple and very efficient method for generating cDNA libraries. Gene 25, 263-269. https: / / doi.org / 10.1016 / 0378-1119(83)90230-5. Hashimshony, T., Senderovich, N., Avital, G., Klochendler, A., de Leeuw, Y., Anavy, L., Gennert, D., Li, S., Livak, K.J., Orit, R.-R., et al. (2016). CEL-Seq2: sensitive highly-multiplexed single-cell RNA-Seq. Genome Biol 17, 77. https: / / doi.org / 10.1186 / s13059-016-0938-8. Islam, S., Kjällquist, U., Moliner, A., Zajac, P., Fan, J.-B.B., Lönnerberg, P., and Linnarsson, S. (2012). Highly multiplexed and strand-specific single-cell RNA 5' end sequencing. Nat Protoc 7, 813-828. https: / / doi.org / 10.1038 / nprot.2012.022. Kilpinen, H., Goncalves, A., Leha, A., Afzal, V., Alasoo, K., Ashford, S., Bala, S., Bensaddek, D., Casale, F.P., Culley, O.J., et al. (2017). Common genetic variation drives molecular heterogeneity in human iPSCs. Nature 546, 370-375. https: / / doi.org / 10.1038 / nature22403. Pandey, S., Takahama, M., Gruenbaum, A., Zewde, M., Cheronis, K., and Chevrier, N. (2020). A whole-tissue RNA-seq toolkit for organism-wide studies of gene expression with PME-seq. Nat. Protoc. https: / / doi.org / 10.1038 / s41596-019-0291-y. Pradhan, R.N., Bues, J.J., Gardeux, V., Schwalie, P.C., Alpern, D., Chen, W., Russeil, J., Raghav, S.K., and Deplancke, B. (2017). Dissecting the brown adipogenic regulatory network using integrative genomics. Sci. Rep. 7, 42130. https: / / doi.org / 10.1038 / srep42130. Sholder, G., Lanz, T.A., Moccia, R., Quan, J., Aparicio-Prat, E., Stanton, R., and Xi, H.S. (2020). 3'Pool-seq: an optimized cost-efficient and scalable method of whole-transcriptome gene expression profiling. BMC Genomics 21. https: / / doi.org / 10.1186 / s12864-020-6478-3. Soumillon, M., Cacchiarelli, D., and Semrau, S. (2014). Characterization of directed differentiation by high-throughput single-cell RNA-Seq. BioRxiv. Waszak, S.M., Delaneau, O., Gschwind, A.R., Kilpinen, H., Raghav, S.K., Witwicki, R.M., Orioli, A., Wiederkehr, M., Panousis, N.I., Yurovsky, A., et al. (2015). Population Variation and Genetic Control of Modular Chromatin Architecture in Humans. Cell 162, 1039-1050. https: / / doi.org / 10.1016 / j.cell.2015.08.001. Ye, C., Ho, D.J., Neri, M., Yang, C., Kulkarni, T., Randhawa, R., Henault, M., Mostacci, N., Farmer, P., Renner, S., et al. (2018). DRUG-seq for miniaturized high-throughput transcriptome profiling in drug discovery. Nat. Commun. 9. https: / / doi.org / 10.1038 / s41467-018-06500-x. Ziegenhain, C., Vieth, B., Parekh, S., Reinius, B., Guillaumet-Adkins, A., Smets, M., Leonhardt, H., Heyn, H., Hellmann, I., and Enard, W. (2017). Comparative Analysis of Single-Cell RNA Sequencing Methods. Mol. Cell 65, 631-643.e4. https: / / doi.org / 10.1016 / j.molcel.2017.01.023.
Claims
1. A set of oligonucleotide molecules comprising, from 5' to 3', a) a sequencing adapter, b) a barcode sequence consisting of 9 to 15 nucleotides, preferably 14 nucleotides, and c) an mRNA capture sequence, wherein the barcode sequence is selected from the group comprising SEQ ID NO. 1TACGTTATTCCGAASEQ ID NO. 2AACAGGATAACTCCSEQ ID NO. 3ACTCAGGCACCTCCSEQ ID NO. 4ACGAGCAGATGCAGSEQ ID NO. 5TTCAATCTCCTTAGSEQ ID NO. 6CTCGGTTCGAATGCSEQ ID NO. 7TACACTATAGCTAGSEQ ID NO. 8CAAGTATAAGGAACSEQ ID NO. 9CTGATATGCAGCGASEQ ID NO. 10ACAATAGGTGGTCCSEQ ID NO. 11AACCATCGCATCCASEQ ID NO. 12CAGTCTAGTGCGCASEQ ID NO. 13CCGCAAGAGGTTGGSEQ ID NO. 14TAACACCTCCAACGSEQ ID NO. 15CAGACATGATAGAGSEQ ID NO. 16AACAACCGAAGTAASEQ ID NO. 17AACTAGTATCCGGASEQ ID NO. 18ACGCAGCGAGCAGASEQ ID NO. 19TCCAAGGCATTAAGSEQ ID NO. 20TATGACTGCATGCGSEQ ID NO. 21ACCGGATAACGGAASEQ ID NO. 22ATAATATTGCCTCCSEQ ID NO. 23TCTTGAACGACTAASEQ ID NO. 24TTAACTTACGGTCASEQ ID NO. 25CATGCTAGCGATCGSEQ ID NO. 26TAAGCAACAGCCGASEQ ID NO. 27ATGTAGGTAATCGGSEQ ID NO. 28CAGCAGCGCTGACASEQ ID NO. 29TATGTACGTAGAGASEQ ID NO. 30CAGATCGCTCTGCGSEQ ID NO. 31TAACCTAACTTCACSEQ ID NO. 32CTCCTAACACAGCCSEQ ID NO. 33ATACCGTTGGAGGASEQ ID NO. 34CTTATATCGTCGGCSEQ ID NO. 35AAGTACTGATCCAASEQ ID NO. 36TACTGACTATCGAGSEQ ID NO. 37ATTGCGTCCATTAGSEQ ID NO. 38TTGCTACCATTGAASEQ ID NO. 39AAGCGCACCGGAACSEQ ID NO. 40ACCAACTTAGACAASEQ ID NO. 41TCAGCGTGTCGTCASEQ ID NO. 42ATTCCGCTCTGAGGSEQ ID NO. 43ACCGATGTGCCGGASEQ ID NO. 44CTAAGGAAGGACCGSEQ ID NO. 45TTCTGAGGCGATGGSEQ ID NO. 46ACATCTTCTTGAACSEQ ID NO. 47ACAGACGATAGTCASEQ ID NO. 48ACTATCTGCAGTCGSEQ ID NO. 49CATGTGTACTAGCASEQ ID NO. 50CTCTGGCTTGCAGGSEQ ID NO. 51ACGGCCTTCAACCASEQ ID NO. 52CCATATTGAAGCACSEQ ID NO. 53TCGATTGGTAATGGSEQ ID NO. 54TTCATGGTAGAGAGSEQ ID NO. 55ACGCTCCTAACGGCSEQ ID NO. 56CTCAATTAGAGACCSEQ ID NO. 57TCGATCTTCGGCAGSEQ ID NO. 58CACATCGACAGCAGSEQ ID NO. 59AACCTGCTATCCGGSEQ ID NO. 60ATCCGGCGATGACCSEQ ID NO. 61TCCATGTGTCCAGGSEQ ID NO. 62ATCGATGTCATAGCSEQ ID NO. 63CTCCGCATGCTTCCSEQ ID NO. 64CACGTAGCTAGACGSEQ ID NO. 65AACATGGCACGGAASEQ ID NO. 66ATGCTATCTAGACASEQ ID NO. 67ACCAATCCGTATAASEQ ID NO. 68CTCGTGCCTCAGGASEQ ID NO. 69ATGCAATTACACCGSEQ ID NO. 70CCTCCATCCATGGASEQ ID NO. 71ACGGCTCTTAGGCCSEQ ID NO. 72TCCAACTACCGAACSEQ ID NO. 73TATGAGAGCTATGASEQ ID NO. 74CAAGGCGTGGTTACSEQ ID NO. 75ATTATCGGATTCGGSEQ ID NO. 76ATCCTAAGTGGAGGSEQ ID NO. 77ATACGATGCTGCGCSEQ ID NO. 78AAGCCTTAGGCCGCSEQ ID NO. 79AATGTGTGGATACCSEQ ID NO. 80ATGCTGCAACGGCASEQ ID NO. 81CTTGATAAGCCAGGSEQ ID NO. 82CATGGAACTCCGGCSEQ ID NO. 83TAAGGACCTCATAASEQ ID NO. 84TCTGTCTATGCAGCSEQ ID NO. 85CAACCAACCTATAGSEQ ID NO. 86CCGGTAGAATGGCCSEQ ID NO. 87TCCGAAGTTGTAGASEQ ID NO. 88TCGATGCTATATGASEQ ID NO. 89ACAAGGTTCCGCCGSEQ ID NO. 90CCACGTTCCGGTAGSEQ ID NO. 91ACTTAAGGTCCAACSEQ ID NO. 92TCCAGAAGGCAACASEQ ID NO. 93CTGCTCATGAGAGASEQ ID NO. 94ATGAATGTTGCCACSEQ ID NO. 95CTCTCAAGTCCTCCSEQ ID NO. 96ATTCCAGTTCTTGC,SEQ ID NO. 98TCGTACATAGCGCASEQ ID NO. 99CAAGGAACCGAACCSEQ ID NO. 100TCAGTCACGCGTGCSEQ ID NO. 101CACACGATACTCACSEQ ID NO. 102AAGGAGAAGAATGCSEQ ID NO. 103CAGTGGCCTAAGCASEQ ID NO. 104ACCGTCTTCCTTGCSEQ ID NO. 105CCTCATTCGTTGAGSEQ ID NO. 106ACGATATGCTCTGCSEQ ID NO. 107CCACCGTTAAGACASEQ ID NO. 108ACTACTCGATCTAGSEQ ID NO. 109AACACTTGGCCACCSEQ ID NO. 110ATCACATCTATCGCSEQ ID NO. 111CTCCAATATTCAACSEQ ID NO. 112CCAATAGACTTCCGSEQ ID NO. 113ACCGTAACGTGGACSEQ ID NO. 114ATAAGGACCGGAGGSEQ ID NO. 115ACCTTCTCACGGCCSEQ ID NO. 116CAACGTGGACCTACSEQ ID NO. 117AACTGCTACGCACCSEQ ID NO. 118AACTTCCGGCGTGCSEQ ID NO. 119CATCGTAGCAGAGCSEQ ID NO. 120AACAACCAAGCGGCSEQ ID NO. 121TAACCGGCCGCAACSEQ ID NO. 122ATAACCATCCGAACSEQ ID NO. 123CTAGCTGTGTCAGCSEQ ID NO. 124CTAACCTTCATTGCSEQ ID NO. 125CTCACCTAGGCCACSEQ ID NO. 126TCATGCGCATGTAGSEQ ID NO. 127TCGCAATTACCTAASEQ ID NO. 128CATCTACGCTACGCSEQ ID NO. 129TTCCACCGATGTGGSEQ ID NO. 130TAGCTTCTCCAGGCSEQ ID NO. 131TTCCAACCGGACAGSEQ ID NO. 132TTGACCTTGCTGGASEQ ID NO. 133ACCTTGTCCAATACSEQ ID NO. 134CTTAGAATTAGTGGSEQ ID NO. 135AAGTGACCGACGGCSEQ ID NO. 136TCGGTTATCTTCGASEQ ID NO. 137CTAGCTTGTTGTAGSEQ ID NO. 138TTAAGCCGAACAAGSEQ ID NO. 139TTCCGATGTCCGGCSEQ ID NO. 140ATGAATCAGCCGCCSEQ ID NO. 141TCCTAACAAGTGCGSEQ ID NO. 142CAATAGCAACACCASEQ ID NO. 143TATACGCGACATCGSEQ ID NO. 144CCGGATGGTGGTACSEQ ID NO. 145CCTACGGTGTTCGGSEQ ID NO. 146CTCGCACCTGGTAASEQ ID NO. 147TCGCCTGAACTACCSEQ ID NO. 148TCACTCATACATAGSEQ ID NO. 149CTTGGTCTCCGGCASEQ ID NO. 150TTGATTACCGTAACSEQ ID NO. 151TTGATCCGTGAAGCSEQ ID NO. 152ATTCTTCAGGAAGASEQ ID NO. 153AAGCCAGGTAAGGCSEQ ID NO. 154CCGCGGAAGGATAASEQ ID NO. 155CATAATTCTATGCCSEQ ID NO. 156ATGAACTTGGCACASEQ ID NO. 157CTTAAGCTTAAGAGSEQ ID NO. 158TCCTCGGAATGACASEQ ID NO. 159AATCTTACCTAAGGSEQ ID NO. 160ACCACTGGTTGCCGSEQ ID NO. 161CTCCGTTCTGTACCSEQ ID NO. 162CCTAACTCGCCAGASEQ ID NO. 163AATAGTGTTCAACCSEQ ID NO. 164CTACGAGTCATAGASEQ ID NO. 165ACACACGACATCGCSEQ ID NO. 166AACTTACACGTTGGSEQ ID NO. 167AACACAACCTTCCGSEQ ID NO. 168ATACCTTAATCTCCSEQ ID NO. 169TAACCGCAATAGGASEQ ID NO. 170CCGAACCATAGCCGSEQ ID NO. 171TTCATATAAGGCGCSEQ ID NO. 172AATACCAGCCGCGGSEQ ID NO. 173CTACGGACTAAGGCSEQ ID NO. 174TAATCGCCTACAGASEQ ID NO. 175TTAGCCTAACTCGCSEQ ID NO. 176TAGTGTGGTAGACCSEQ ID NO. 177CAACGCCATAGACCSEQ ID NO. 178TAATCCATCAACGCSEQ ID NO. 179CCGGTGTGCGCTAASEQ ID NO. 180TACCGGTATGGTCCSEQ ID NO. 181CACGTATAGATCGCSEQ ID NO. 182AACCTAAGCAATCCSEQ ID NO. 183CTAGTGTTCGATCCSEQ ID NO. 184CCTACACTTGAGCCSEQ ID NO. 185CAGCATCATACTGCSEQ ID NO. 186AATACGGAGACGGCSEQ ID NO. 187ATTGAACTCCATCCSEQ ID NO. 188AAGAGTCGGATAGGSEQ ID NO. 189AATGCAAGACAACCSEQ ID NO. 190TTCATTGGAAGAGASEQ ID NO. 191TTCTAGGACGGCCGSEQ ID NO. 192CTCAACCACGTTCCSEQ ID NO. 193CTGGTCTGTCTTCCSEQ ID NO. 194TTCAAGTCAGTGCCSEQ ID NO. 195CACACGACAGAGCGSEQ ID NO. 196CCTAAGCCTGTTCGSEQ ID NO. 197AAGCTGGCGCCTAGSEQ ID NO. 198ACGCCTGTTGGAGGSEQ ID NO. 199TCTCCATTATACAGSEQ ID NO. 200TCCGCCGCACTCAASEQ ID NO. 201TATCGATCATCTACSEQ ID NO. 202TCTTGTACCGGACCSEQ ID NO. 203AACGGCTGAACGACSEQ ID NO. 204TTCCGAAGCCTAAGSEQ ID NO. 205TTCTAGATAGGTACSEQ ID NO. 206AATTCTCTGGCGCCSEQ ID NO. 207ATGACTGAGAGTGASEQ ID NO. 208TAATTAGGTCCTCGSEQ ID NO. 209AACCGTCCAAGCGCSEQ ID NO. 210ACATTATTGGCCAASEQ ID NO. 211TCCACAGAACCGGCSEQ ID NO. 212TAGTGATCTGTAGCSEQ ID NO. 213CATTCACGCTTGAASEQ ID NO. 214CACACACGAGATGCSEQ ID NO. 215CCATGACCACCGGASEQ ID NO. 216CAACAAGACCTAAGSEQ ID NO. 217AAGGAACCTCTCGGSEQ ID NO. 218CATAACCGGCACAASEQ ID NO. 219CCAAGTGGTACAGCSEQ ID NO. 220TTCCTTACCGCGCGSEQ ID NO. 221ATACTAGTAGTCAGSEQ ID NO. 222TATATGCTCGCTCASEQ ID NO. 223TCCGCGATCTCAACSEQ ID NO. 224TATTACCGCTCAGGSEQ ID NO. 225ATCGAGCAGTGTACSEQ ID NO. 226CATGCGTAGCGCACSEQ ID NO. 227TCCAACAGTGTTGGSEQ ID NO. 228CACTTGGACCATCCSEQ ID NO. 229TTACCTGGCCATCCSEQ ID NO. 230CATCATGAGAGCGASEQ ID NO. 231ATTCCGAATGTAACSEQ ID NO. 232CCGGTATATCCGGASEQ ID NO. 233ACTGACGTGGAACCSEQ ID NO. 234TTAAGCTGCCTTAASEQ ID NO. 235CTCCGCAGTTGGAGSEQ ID NO. 236TCCTAATTCTAAGGSEQ ID NO. 237TACGATCGTGATCASEQ ID NO. 238ATAACCAACTAGGASEQ ID NO. 239ATCCTTATACGTGCSEQ ID NO. 240ACTTCCAGGCCTGASEQ ID NO. 241TCTGCACATCACGCSEQ ID NO. 242ACAGACGCATCACGSEQ ID NO. 243TAAGTTGTGTGTCCSEQ ID NO. 244ACTCAACGGCAAGGSEQ ID NO. 245TAGCACATACTCGASEQ ID NO. 246ATATCACAGTCGCASEQ ID NO. 247ATATACACATGAGCSEQ ID NO. 248ACGTACGCGCATGCSEQ ID NO. 249CAAGCCAGTTCCGGSEQ ID NO. 250CCGAACAACCGTGASEQ ID NO. 251CACACACATAGCGASEQ ID NO. 252ACAGCGTACGCGCASEQ ID NO. 253TCAACGCTCGGAGCSEQ ID NO. 254ACTACGATCATAGASEQ ID NO. 255CAGGCGTAAGTGCCSEQ ID NO. 256ACATATCATGAGCASEQ ID NO. 257AACTCCGTACAAGCSEQ ID NO. 258CTGTCGAGTAGCAGSEQ ID NO. 259TAGACGATAGTACASEQ ID NO. 260CATGCTCACGCAGCSEQ ID NO. 261ATCACGAGCGACGCSEQ ID NO. 262CAGCCTGGCTTAGGSEQ ID NO. 263CCAGAGGTTGCCAGSEQ ID NO. 264TACTATGCACGACGSEQ ID NO. 265ACGCTGACGTGCGASEQ ID NO. 266CAGCACATGTACAGSEQ ID NO. 267ACACGCATCATACGSEQ ID NO. 268CTGTACACGTCTACSEQ ID NO. 269CACACAGTGAGTACSEQ ID NO. 270TAGCACGCGATGAGSEQ ID NO. 271TCCAATAGGTCCAGSEQ ID NO. 272CTCCTGCGGTCCAASEQ ID NO. 273ATGGCAACGGCTCASEQ ID NO. 274CCGAGGCTTCCTAGSEQ ID NO. 275ACGACAGTAGTCGCSEQ ID NO. 276TTATAACTGGCAACSEQ ID NO. 277ACTCTGAGTATCAGSEQ ID NO. 278TTATGTCTTAACGGSEQ ID NO. 279ATGCGACTCGAGACSEQ ID NO. 280CCTCAATTCCGGCGSEQ ID NO. 281AATGGTTGGCAGAGSEQ ID NO. 282ATGCAAGCGATTCCSEQ ID NO. 283TATTGAGTCCGAGASEQ ID NO. 284CCGCGTATTGCAACSEQ ID NO. 285TCATTCTACTTGGCSEQ ID NO. 286ATCATCGCGCGACGSEQ ID NO. 287CAGAGACACTCGCASEQ ID NO. 288AAGAAGAGTGGACCSEQ ID NO. 289ACACACGCTGCGACSEQ ID NO. 290TCCGCAAGTCATAGSEQ ID NO. 291AATGTCCTCAGAACSEQ ID NO. 292CCGAGTGTAAGTCCSEQ ID NO. 293ACACGTACTCATACSEQ ID NO. 294TTCTCTCCGGTTCASEQ ID NO. 295AATCCGTGTCTTAASEQ ID NO. 296CCTTGCCAGAGTGGSEQ ID NO. 297CTTACTGACAAGCCSEQ ID NO. 298TCCTCCGATGCCAGSEQ ID NO. 299TTCTTCTCCAGAAGSEQ ID NO. 300CACCTGTTGGTTAASEQ ID NO. 301CTCAGTCACACGAGSEQ ID NO. 302CCTTATCTACGACCSEQ ID NO. 303AACCGAACGCTTAASEQ ID NO. 304ACAGGTAGGAAGGASEQ ID NO. 305CAGCATAGTCTGACSEQ ID NO. 306ATTCGGTCCGATGCSEQ ID NO. 307ACTAACCTGGCCGGSEQ ID NO. 308CATTCCTCTAGCCGSEQ ID NO. 309CTTAATCCATATCCSEQ ID NO. 310AACAAGAGGTGTGGSEQ ID NO. 311TATCATATGAGTCGSEQ ID NO. 312ACGAGTATGTATGCSEQ ID NO. 313TCGCGGATTAGGCASEQ ID NO. 314CTACTGTGACACACSEQ ID NO. 315CAACCGTGTGAACCSEQ ID NO. 316CAGCACTAGCTGCASEQ ID NO. 317ACACGACATGCACGSEQ ID NO. 318ACTCATGTCAGACASEQ ID NO. 319ACACTATACACGAGSEQ ID NO. 320ACTTGAGCAACCGGSEQ ID NO. 321TTAGACGCCGGTAASEQ ID NO. 322CTGTGTAGGATTAASEQ ID NO. 323CACGCATCAGTCAGSEQ ID NO. 324CAACCATTGGCAAGSEQ ID NO. 325TCATCTGCGCACACSEQ ID NO. 326TAGGCCTCTGATGGSEQ ID NO. 327TAGCATGTCAGCACSEQ ID NO. 328TCCTGGACGATGCCSEQ ID NO. 329ATATCACATGACAGSEQ ID NO. 330ACTACGTCGACACGSEQ ID NO. 331AATCCAACTCGAAGSEQ ID NO. 332AAGGATGCCAGTGGSEQ ID NO. 333ACCGCCTTGGCTAGSEQ ID NO. 334CTGTGATTCGGACCSEQ ID NO. 335CATGAGACGTCGCGSEQ ID NO. 336AATCCGCCTAATGGSEQ ID NO. 337AATGCCGTATTGACSEQ ID NO. 338CTCCGAAGAACCGGSEQ ID NO. 339ACTGATTAGAATCGSEQ ID NO. 340ACTCGCCACCGAAGSEQ ID NO. 341CCGGTATTGTTACASEQ ID NO. 342TTGGTGAGAACCGCSEQ ID NO. 343CAGATCTCAGACACSEQ ID NO. 344CCTTAACACGGAGASEQ ID NO. 345ATGTCTAGTATAGCSEQ ID NO. 346CCAGAGGTAATGGASEQ ID NO. 347CTGTCACACAGTCASEQ ID NO. 348TCACTGAGTGCTGASEQ ID NO. 349TCCTATGGTAGGAGSEQ ID NO. 350AAGGAACAATACACSEQ ID NO. 351TAAGCCGGAATTAGSEQ ID NO. 352ATAATGCTCCAGAASEQ ID NO. 353ATTCTTGCGAAGCGSEQ ID NO. 354ACATGCTATACTACSEQ ID NO. 355AAGCCAACAGTTGGSEQ ID NO. 356TATGAGCACACGACSEQ ID NO. 357CACTAGATATCACASEQ ID NO. 358TCAATTAAGGCACGSEQ ID NO. 359TCCTCCTCTTGTGASEQ ID NO. 360TACTACTATGTGACSEQ ID NO. 361TATCTTCCACCTGASEQ ID NO. 362CACTTACCTTACAASEQ ID NO. 363TCTAGCATCGACGASEQ ID NO. 364ACACGAGCACGAGCSEQ ID NO. 365ACAAGGCAACAGCCSEQ ID NO. 366CACTACGCGCTCACSEQ ID NO. 367TAACCGTCTCTGGCSEQ ID NO. 368CCAACTCGCCAGGASEQ ID NO. 369CTTCGGTACACTAASEQ ID NO. 370CCACACCTTCAGGCSEQ ID NO. 371CCACTTCTTGTTGGSEQ ID NO. 372CTAGTCTTGGTAGGSEQ ID NO. 373TCAGAATCCTCGCCSEQ ID NO. 374CTAGGAGCTTGAACSEQ ID NO. 375CCGGCGCAATTCAGSEQ ID NO. 376ACCGGTCGTTCTGCSEQ ID NO. 377CATTAAGGTATCGCSEQ ID NO. 378CTTGGAGTGGCACASEQ ID NO. 379AACCTTCTAGAGAASEQ ID NO. 380CATGGAATGTGTAASEQ ID NO. 381CTGTTCGGAGGTCASEQ ID NO. 382ATCGTGTGTGGCCASEQ ID NO. 383TTGTCGGTTGAGGCSEQ ID NO. 384CTGGCGTTGAGTCGSEQ ID NO. 385AATGACCACGACCGSEQ ID NO. 386ATGTCTACAGACGASEQ ID NO. 387TCTCCGAGGTAGGCSEQ ID NO. 388ACGACGACTGACAGSEQ ID NO. 389ACTATGACTGCGCASEQ ID NO. 390CCGCCTTCTAACACSEQ ID NO. 391CACACGCAGTATCASEQ ID NO. 392CCAGCTTAGGAAGASEQ ID NO. 393ACACGATCTGACGASEQ ID NO. 394AAGGCAAGTTGTGASEQ ID NO. 395AACTCTTCGGTAGGSEQ ID NO. 396CCGTCCTTCTTCACSEQ ID NO. 397CTTGGTTGGTTCCGSEQ ID NO. 398TACCTTGGAGTTGGSEQ ID NO. 399ACCGGCGGTATAGGSEQ ID NO. 400CACGACGACTCAGASEQ ID NO. 401TCTTGTGAGGTTCGSEQ ID NO. 402AACGCTTCCGGTCCSEQ ID NO. 403CTACTGCATAGTAGSEQ ID NO. 404ACACAGCACGATAGSEQ ID NO. 405CCTACAGGTCGAGGSEQ ID NO. 406ATGCACGATCGAGCSEQ ID NO. 407CTTGATACCGGCGCSEQ ID NO. 408CCTCGGTTGAAGCCSEQ ID NO. 409CAACACCGGCTTGGSEQ ID NO. 410CTCTGGACCGATCASEQ ID NO. 411CCGTGAACCTTGCGSEQ ID NO. 412TCATGCTCGAGACASEQ ID NO. 413ACCGGAACCTATGGSEQ ID NO. 414TTCCAAGGTTGACASEQ ID NO. 415ACGAGAGTGCTGAGSEQ ID NO. 416CTCAACCGGACGGASEQ ID NO. 417AATCAAGTGTGAACSEQ ID NO. 418CATACGGCCTCTAASEQ ID NO. 419ACTAGAGCGTGTGASEQ ID NO. 420TCAGAGTAGTGCAGSEQ ID NO. 421CAACAACCTTGAGGSEQ ID NO. 422ACAATTAGCCTTGGSEQ ID NO. 423TCACTATGCGATACSEQ ID NO. 424TACTCTCATCTAGCSEQ ID NO. 425TCTCTCTCAATTGGSEQ ID NO. 426CCAGTGAGGTATACSEQ ID NO. 427CTCCTACTTAATGCSEQ ID NO. 428ATAGGCCTTCGAGGSEQ ID NO. 429CATTGGTGGTGTCCSEQ ID NO. 430CATACTATGTAGACSEQ ID NO. 431ATGGATCGTTAGGASEQ ID NO. 432CCAACCGAACGGCASEQ ID NO. 433ACGTACATCTATCGSEQ ID NO. 434TACGAGTGTCTCAGSEQ ID NO. 435ATAACAACGAAGCCSEQ ID NO. 436TACGATAGTACAGCSEQ ID NO. 437ATTGCGGTTAATCASEQ ID NO. 438CAGTTGCGAAGTGGSEQ ID NO. 439TTAACACCTCAAGGSEQ ID NO. 440CACGACATGCGAGCSEQ ID NO. 441TCCAGAACTGAGGCSEQ ID NO. 442TATTCCTAGTTCGASEQ ID NO. 443AACACCTTCTCCGCSEQ ID NO. 444TAACCTGTTCAAGASEQ ID NO. 445AAGCTATGTGCAACSEQ ID NO. 446CTCTTAATCCTTGASEQ ID NO. 447ATCAGCAGACTAGCSEQ ID NO. 448CAAGGAAGCATTGGSEQ ID NO. 449ACCTCGAAGTTCAASEQ ID NO. 450ATAGCACTCACACGSEQ ID NO. 451ATCTAGCTGCACGASEQ ID NO. 452TCTGCTGCAGAGCGSEQ ID NO. 453TAACAGTCCGTTCASEQ ID NO. 454CATCTATCTGCGCGSEQ ID NO. 455TTGACCAAGATTACSEQ ID NO. 456TCTCACATCACGAGSEQ ID NO. 457ACACATCTCTACGASEQ ID NO. 458AACAGAATAGGCGGSEQ ID NO. 459AACACCTAAGGAAGSEQ ID NO. 460AACACCTGGTTGAASEQ ID NO. 461AAGACGAATAGGAASEQ ID NO. 462CAAGTTCAGAATGGSEQ ID NO. 463CCAAGCTTAGTGCGSEQ ID NO. 464AACGCTCAATGTGGSEQ ID NO. 465TCATCTAGGTGGAASEQ ID NO. 466CCTGTGATTACACCSEQ ID NO. 467TCAGAACAGGTTAASEQ ID NO. 468ATATTGTAAGGTGGSEQ ID NO. 469TTAAGTGACCGCGCSEQ ID NO. 470AATCCTGAATCCAASEQ ID NO. 471CTATAGCGCACTACSEQ ID NO. 472TCTAGACACTATCGSEQ ID NO. 473TCAATGGCAGGCCGSEQ ID NO. 474AAGTGGTATTGACASEQ ID NO. 475TTGGACATTCCGGCSEQ ID NO. 476CAGCATCGAGAGCGSEQ ID NO. 477CAAGACCTTGGCCASEQ ID NO. 478CCGGTTCCTAGTGASEQ ID NO. 479ACAACTTAGCCTAASEQ ID NO. 480TCGTGTCGCATGCASEQ ID NO. 481CACACAGACGCTCG, or a combination of one or more thereof.
2. The set of oligonucleotide molecules of claim 1, further comprising d) a unique molecular identifier (UMI).
3. The set of oligonucleotide molecules of claim 2, wherein the UMI consists of 14 nucleotides.
4. The set of oligonucleotide molecules of any of the preceding claims, wherein the sequencing adapter comprises, or consists of, CTA CAC GAC GCT CTT CCG ATC T (SEQ ID No. 97).
5. The set of oligonucleotide molecules of claim 1, wherein the combination consists of at least 10 oligonucleotides, at least 20 oligonucleotides, at least 30 oligonucleotides, at least 40 oligonucleotides, at least 50 oligonucleotides, at least 60 oligonucleotides, at least 70 oligonucleotides, at least 80 oligonucleotides, at least 90 oligonucleotides, or a combination of all the oligonucleotides of SEQ ID No. 1 to 96.
6. The set of oligonucleotide molecules of claim 1, wherein the combination consists of at least 10 oligonucleotides, at least 20 oligonucleotides, at least 30 oligonucleotides, at least 40 oligonucleotides, at least 50 oligonucleotides, at least 60 oligonucleotides, at least 70 oligonucleotides, at least 80 oligonucleotides, at least 90 oligonucleotides, at least 100 oligonucleotides, at least 150 oligonucleotides, at least 200 oligonucleotides, at least 250 oligonucleotides, at least 300 oligonucleotides, or a combination of all 384 oligonucleotides of SEQ ID No. 98 to 481.
7. The set of oligonucleotide molecules of any one of claims 2-6, wherein the UMI consists of a sequence (N)n(V)m, wherein N is any nucleotide selected from A, T, C and G; V is any nucleotide selected from A, C and G; n is an integer selected from 1 to 20, and m is an integer selected from 1 to 20.
8. The set of oligonucleotide molecules of any one of the preceding claims, wherein the mRNA capture sequence is a poly-T sequence followed by at least one V and one N, wherein N is any nucleotide selected from A, T, C and G and V is any nucleotide selected from A, C and G.
9. A set of oligonucleotide molecules, each molecule comprising, from 5' to 3', a) a sequencing adaptor, b) a barcode sequence independently selected from the group consisting of SEQ ID: 1 to 96 or a barcode sequence independently selected from the group consisting of SEQ ID: 98 to 481, and c) an mRNA capture sequence.
10. The set of oligonucleotide molecules of claim 9, further comprising d) a UMI.
11. Use of a set of oligonucleotide molecules of any one of claims 1 to 8, or of a set of oligonucleotide molecules of claim 9 or 10, in a sequencing method.
12. A method for providing a cDNA library, the method comprising the steps of a) Providing a plurality of RNA samples obtained from a biological sample ; b) Contacting separately each RNA sample with a set of oligonucleotide molecules of any one of claims 1 to 8, or of a set of oligonucleotide molecules of claim 9 or 10, under annealing conditions; c) Incubating separately each sample under reverse transcription reaction conditions; d) Pooling together all the first strand cDNA or cDNA:RNA sample; e) Proceeding to second strand synthesis under synthesis conditions; and f) Proceeding with tagmentation and amplification under suitable conditions so as to obtain a sequencing compatible cDNA library or g) Proceeding with enzymatic and physical force fragmentation followed by DNA end repair, A-tailing, DNA adapter ligation and amplification under suitable conditions so as to obtain a sequencing compatible cDNA library.
13. The method for providing a cDNA library of claim 12, wherein the second strand synthesis is generated by a method selected from the group comprising PCR amplification and nick translation, or a combination thereof.
14. A method for sequencing RNA, the method comprising the steps of a) Providing a cDNA library obtained by the method of claims 12 to 13; and b) Proceeding to the sequencing under suitable conditions.
15. A kit comprising a) a set of oligonucleotide molecules of any one of claims 1 to 8, or a set of oligonucleotide molecules of claim 9 or 10, b) a support for sample preparation, and c) reagents for sequencing.