Method for determining the copy number or sequence of one or more RNA molecules

JP2024529548A5Pending Publication Date: 2025-06-27BASIC GENOMICS AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024507173
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-03
Filing Date
2022-07-29
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Current short-read single-cell RNA sequencing methods are limited in their ability to count RNA molecules with allelic and isoform resolution, while long-read sequencing is too expensive for large-scale applications, and existing technologies struggle to accurately determine the copy number and sequence of RNA molecules within populations.

Method used

A method involving error-prone reverse transcription to introduce unique base conversion patterns in DNA molecules, allowing for the determination of RNA copy number and sequence by analyzing molecule-specific base conversion patterns in DNA molecules.

Benefits of technology

Enables accurate counting and sequencing of RNA molecules within populations, improving resolution and reducing costs compared to existing methods, and allowing for the reconstruction of RNA sequences with high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method for determining the copy number of one or more RNA molecules in a population of RNA molecules and a method for determining the sequence of one or more RNA molecules in a population of RNA molecules, which comprises converting the population of RNA molecules by error-prone reverse transcription into a population of DNA molecules comprising one or more base alterations. The present invention also relates to a population of DNA molecules obtained or obtainable by the methods disclosed herein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to methods for determining the copy number of one or more RNA molecules in a population of RNA molecules and for determining the sequence of one or more RNA molecules in a population of RNA molecules, which methods comprise converting the population of RNA molecules into a population of DNA molecules by error-prone reverse transcription. The present invention also relates to a population of DNA molecules obtained or obtainable by the methods disclosed herein.

[0002] The application of massively parallel sequencing has transformed biology and medicine. To investigate genetic programs in cell populations or single cells, it is routine today to perform RNA sequencing on thousands to millions of individual single cells or cell populations. Such analyses can reveal patterns of gene, isoform and allele expression across cell types and states. However, current short-read single-cell RNA sequencing (scRNA-seq) methods are limited in their ability to count RNA with allele and isoform resolution, and long-read sequencing techniques are prohibitively expensive for the depth required for large-scale application across cells.

[0003] Most scRNA-seq methods count RNA by sequencing a short portion of the RNA (from either the 5' or 3' end) along with a unique molecular identifier (UMI). Although these RNA end-counting strategies have been effective in estimating gene expression across a large number of cells while controlling for PCR amplification bias, RNA end-sequencing provides limited coverage of transcribed gene variants and transcript isoform expression.

[0004] Most contemporary sequencing protocols are built on short-read sequencing platforms (e.g., those of Illumina or MGI), and these mature platforms are cost-effective for deep large-scale sequencing across cells. Transcriptional genetic analysis using short-read sequencing techniques has a common limitation that individual short reads can only cover a small portion of the body of a particular transcript. Short fragments generated by commonly used droplet methods (e.g., 10× Chromium system) target either end of the RNA transcript (i.e., 3' or 5' end, depending on the protocol used). Alternatively, short reads can be distributed throughout the entire RNA transcript as in Smart-seq2 (Picelli et al, 2013. Nature Methods, 10:1096-1098) or Smart-seq3 (Hagemann-Jensen et al, 2020. Nature Biotechnology, 38:708-714) techniques. However, even in methods with reads distributed across the RNA transcript, multiple short reads (or paired end reads) cannot be assembled individually (e.g., to reconstruct the original sequence of individual molecules). Instead, the total read coverage across the transcript relates to the total number of RNA transcripts sequenced from the cell(s). Importantly, in the above methods, molecular counting through the use of unique molecular identifiers (UMIs) is always limited to the 3' or 5' end of a complementary DNA (cDNA) molecule. Combining reads that cover the same unique molecular barcode can provide a limited RNA sequence reconstruction (Hagemann-Jensen et al, 2020. Nature Biotechnology, 38:708-714), which can theoretically provide up to the maximum fragment length that can be sequenced on a short-read sequencing instrument (e.g., 200-800 base pairs for Illumina sequencing).

[0005] Sequencing full-length RNA transcripts using long-read DNA sequencing technologies (e.g., using Pacific Biosystem reaction reactors or Oxford nanopore sequencing) can directly quantify allele- and isoform-level expression, but their current cost relative to read depth prevents their widespread application across cells, tissues, and organisms. Furthermore, such long-read sequencing platforms are more expensive and do not offer the same level of parallelization in terms of the number of DNA molecules that can be sequenced simultaneously as short-read platforms.

[0006] Thus, there is a need for a method for counting and / or qualitatively sequencing RNA molecules that addresses the shortcomings encountered in prior art methods.

[0007] The present inventors have developed a new approach for counting and / or qualitatively sequencing RNA molecules in a population that addresses the above problems.

[0008] As discussed in detail below, our approach involves introducing unique patterns of base conversion to individual cDNA molecules during reverse transcription of corresponding RNA molecules in the population, and then using those unique patterns to count individual RNA molecules in the population and assemble sequences from short reads. We have surprisingly found that the unique patterns of base conversion can be stably amplified during subsequent DNA amplification and can be used to identify and count individual transcripts present in the RNA molecule population. Due to the unique nature of each base conversion pattern in a given molecule in the starting plurality of cDNA molecules, we can simultaneously sequence and count many more transcripts in the RNA molecule population than is possible using existing short read sequencing techniques. Advantageously, the method of the present invention also identifies the origin of the analyzed sequencing reads as RNA transcribed from the positive strand, RNA transcribed from the negative strand (collectively referred to as "strandedness"), or any DNA source (e.g., genomic DNA).

[0009] In a first aspect, the present invention provides a method for determining the copy number of one or more RNA molecules in a population of RNA molecules, comprising the steps of: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to produce a population of DNA molecules, each DNA molecule containing one or more base conversions relative to a corresponding RNA molecule, each DNA molecule containing a molecule-specific base conversion pattern; and determining the copy number of one or more RNA molecules in the population using the molecule-specific base conversion pattern.

[0010] In some embodiments of the first aspect, the step of using the molecule-specific base conversion pattern to determine the copy number of one or more RNA molecules in the population is performed after step (ii), hereinafter: (iii) determining the sequences of overlapping fragments of DNA molecules in the population; (iv) determining partial or full-length sequences of DNA molecules in the population from the information of step (iii) by assembling sequences of overlapping fragments based on molecule-specific base conversion patterns in the DNA molecules; (v) determining the sequence of an RNA molecule corresponding to the DNA molecule from the information from step (iv); (vi) determining the copy number of one or more RNA molecules in the population from the information of step (v).

[0011] In a second aspect, the present invention provides a method for determining the copy number of one or more RNA molecules in a population of RNA molecules, comprising the steps of: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to produce a population of DNA molecules, each DNA molecule containing one or more base conversions relative to a corresponding RNA molecule, each DNA molecule containing a molecule-specific base conversion pattern; (iii) determining the sequences of overlapping fragments of DNA molecules in the population; (iv) determining partial or full-length sequences of DNA molecules in the population from the information of step (iii) by assembling sequences of overlapping fragments based on molecule-specific base conversion patterns in the DNA molecules; (v) determining the sequence of an RNA molecule corresponding to the DNA molecule from the information from step (iv); (vi) determining the copy number of one or more RNA molecules in the population from the information of step (v).

[0012] In a third aspect, the present invention provides a method for determining the sequence of one or more RNA molecules in a population of RNA molecules, comprising: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to produce a population of DNA molecules, each DNA molecule containing one or more base conversions relative to a corresponding RNA molecule, each DNA molecule containing a molecule-specific base conversion pattern; and determining the sequence of an RNA molecule corresponding to one or more DNA molecules using the molecule-specific base conversion pattern.

[0013] In some embodiments of the third aspect, the step of using the molecule-specific base conversion pattern to determine the copy number of one or more RNA molecules in the population is performed after step (ii), hereinafter: (iii) determining the sequences of overlapping fragments of DNA molecules in the population; (iv) determining the sequence of one or more DNA molecules in the population from the information of step (iii) by assembling the sequences of the overlapping fragments based on the molecular specific base conversion patterns of the DNA molecules; (v) determining the sequence of the RNA molecule that corresponds to the one or more DNA molecules from the information from step (iv).

[0014] In a fourth aspect, the present invention provides a method for determining the sequence of one or more RNA molecules in a population of RNA molecules, comprising: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to produce a population of DNA molecules, each DNA molecule containing one or more base conversions relative to a corresponding RNA molecule, each DNA molecule containing a molecule-specific base conversion pattern; (iii) determining the sequences of overlapping fragments of DNA molecules in the population; (iv) determining the sequence of one or more DNA molecules in the population from the information of step (iii) by assembling the sequences of the overlapping fragments based on the molecular specific base conversion patterns of the DNA molecules; (v) determining the sequence of an RNA molecule that corresponds to the one or more DNA molecules from the information of step (iv).

[0015] "One or more RNA molecules" includes the meaning of RNA molecules having a unique sequence. The sequence of the one or more RNA molecules may differ from the sequences of other RNA molecules in the population of RNA molecules because they are derived from different genes and are sequence variants of the same gene, allelic variants of the same gene, splice variants of the same gene, RNA isoforms resulting from alternative promoter usage in the same gene, RNA isoforms resulting from alternative splice site usage in the same gene, or RNA isoforms resulting from alternative polyadenylation site usage in the same gene.

[0016] "RNA molecule population" includes the meaning of multiple individual RNA molecules with the same or different sequences that are analyzed using the methods disclosed herein. For example, an RNA molecule population may contain multiple copies of the same RNA molecule, or more typically, may contain a mixture of RNA molecules with different sequences, optionally with each RNA sequence present in different copy numbers. Examples of RNA molecule populations include, but are not limited to, total RNA obtained from a single cell, multiple cells, or tissue, nuclear or cytoplasmic RNA obtained from a single cell, multiple cells, or tissue, purified pre-mRNA and / or mRNA, free RNA obtained from bodily fluids such as blood, cerebrospinal fluid, and urine, in vitro transcribed RNA, or combinations thereof. For example, an RNA molecule population may include RNA molecules from different sources that are analyzed together as a single experiment using the methods disclosed herein.

[0017] "DNA molecule population" includes the meaning of a plurality of individual DNA molecules having the same or different sequences.For example, a DNA molecule population may contain multiple copies of the same DNA molecule, or more typically, may contain a mixture of DNA molecules having different sequences, optionally with each DNA sequence present in different copy numbers.In the context of the present invention, such a population may be a plurality of individual cDNA molecules produced by reverse transcription of a population of RNA molecules, such as a population of RNA molecules as defined herein.

[0018] By "the same sequence" we include the meaning of RNA molecules which have the same sequence as one another, or DNA molecules which have the same sequence as one another.

[0019] "Different sequences" includes the meaning of RNA molecules that are different from each other in sequence, or DNA molecules that are different from each other in sequence. For example, RNA molecules may have different sequences because they are produced from different genes, or they are different processed transcription products from the same gene (e.g., splice variants). In the case of DNA molecules, they may have different sequences because they are generated from different RNA molecules during reverse transcription, or amplified from different template DNA molecules (e.g., in a PCR process), or because they are sequence variants of genes or alleles.

[0020] "Error-prone reverse transcription" includes the meaning of a reverse transcription process in which the resulting DNA molecules have sequence changes compared to the template RNA molecule from which they are derived. In the context of the present invention, error-prone reverse transcription is reverse transcription performed to purposely incorporate sequence changes into the DNA molecules produced by reverse transcription. This can be achieved in three main ways: (i) reverse transcriptases that incorporate bases that are not complementary to the RNA template molecule in the first strand cDNA, (ii) reverse transcriptases that incorporate non-canonical bases into the first strand cDNA, thereby resulting in more frequent errors during second strand cDNA synthesis, (iii) reverse transcriptases that incorporate non-canonical bases into the first strand cDNA, which change the sensitivity / resistance to chemical treatments, thereby resulting in a change in the frequency of errors at non-canonical base positions during second strand cDNA synthesis after exposure to such chemical treatments. In each of the above three examples, the double-stranded cDNA generated from the RNA template molecule contains base conversions resulting from errors made during the reverse transcription process.

[0021] "Base conversion" includes the meaning of a change in a DNA molecule produced by reverse transcription that results in a change in the base sequence of a DNA molecule amplified from that DNA molecule compared to the base sequence of the corresponding RNA template molecule in a population of RNA molecules. A change in a DNA molecule can be induced, for example, by an error during reverse transcription (i.e., erroneous incorporation of a base not present in the template RNA molecule during first or second strand cDNA synthesis), chemical modification of the RNA molecule before reverse transcription, or chemical modification of the DNA molecule after reverse transcription (but before amplification). A change in a DNA molecule can also be induced, for example, by an error during reverse transcription or incorporation of a non-canonical base (e.g., incorporation of a base that is not the canonical complementary base to the corresponding base in the template RNA). For example, a chemical modification that deaminates cytosine (which base pairs with guanine) can result in the production of uracil (which base pairs with adenine), inducing a GC to AT transition. In relation to base analogs, the purine analog 2-aminopurine is an analog of guanine or adenine that can base pair with either thymine (as a thymine analog) or cytosine (as a guanine analog), and thus can induce AT to GC or GC to AT transitions, while 5-bromouracil (5-BrU) is an analog of thymine that can base pair with adenine (as 5-BrU keto) or guanine (as 5-BrU enol), and thus can induce AT to GC transitions. In another example, base analogs can be incorporated during reverse transcription such that the resulting cDNA molecule or subsequently amplified molecule contains a base change relative to the RNA template molecule.

[0022] "Base analog" includes the meaning of a molecule that has a structure similar to one of the four canonical nitrogenous bases present in DNA (i.e., guanine, cytosine, adenine, and thymine) and can replace one of these canonical bases by reverse transcriptase during cDNA synthesis or by DNA polymerase enzyme during DNA synthesis. In the context of the present invention, a base analog introduced into a DNA molecule produced during reverse transcription can form an altered base pairing with a canonical base present in an RNA molecule (i.e., guanine, cytosine, adenine, and uracil). During the subsequent amplification of the DNA molecule produced by reverse transcription, the base analog can pair with a base that is different from the base present in the corresponding RNA molecule in the population of RNA molecules, resulting in a stable specific base conversion at that position in the sequence of the DNA molecule amplified from that particular DNA molecule. Different base analogs can form different altered base pairs and therefore induce different base conversions.

[0023] "Molecular specific base conversion pattern" includes the meaning of a pattern of base conversion that is unique to a single individual DNA molecule present in a population of DNA molecules produced by reverse transcription. The molecular specific base conversion pattern is relative to the sequence of the corresponding RNA molecule from which the DNA molecule was derived during reverse transcription and stably propagated with the sequence amplified from that DNA molecule. Thus, the molecular specific base conversion pattern can be used to identify all molecules amplified from individual DNA molecules within the population of DNA molecules produced by reverse transcription.

[0024] It is important that the molecule-specific base conversion pattern is stably associated with all molecules derived from each DNA molecule produced by reverse transcription.For example, if new base conversions occur during amplification and / or sequencing, and / or changes occur to existing base conversion patterns, molecules with new base conversion patterns that did not exist in the DNA molecules produced by reverse transcription will be generated.The production of molecules with new base conversion patterns during amplification and / or sequencing will result in an overestimation of the number of individual molecules of a specific sequence in the DNA molecule population produced by reverse transcription, and as a result, an overestimation of the copy number of the corresponding RNA molecule in the initial population of RNA molecules.

[0025] Therefore, in view of the importance of minimizing and / or preventing new molecule-specific base conversion patterns arising during subsequent steps of reverse transcription (e.g., amplification and sequencing), conditions that induce base conversion during reverse transcription are removed prior to subsequent steps of the methods disclosed herein. For example, conditions that induce base conversion during reverse transcription can be removed from the DNA molecule population by cleaning up and / or purifying the DNA molecule population. For example, conditions that induce base conversion during reverse transcription can be removed from the DNA molecule population by methods such as dilution, phenol-chloroform extraction, bead cleanup, enzyme removal, and / or thermal decomposition.

[0026] In some embodiments of the methods disclosed herein, the methods also allow for the determination of the origin and strandedness of the analyzed sequencing reads, such as RNA transcribed from the positive strand, RNA transcribed from the negative strand (collectively referred to as "strandedness"), or reads derived from DNA (e.g., genomic DNA).

[0027] "Strandedness" refers to whether the sequence of an original RNA molecule is present in the plus or minus strand of the DNA into which it is transcribed.

[0028] Typically, the reverse transcription reaction is carried out using template RNA, reverse transcriptase, dNTPs, and primer molecules. The reverse transcription reaction may also contain associated salts and / or other additives. Examples of commercially available reverse transcriptases are known in the art and include enzymes such as AMV reverse transcriptase (New England Biolabs), SmartScribe II (Takara), Maxima H-minus (Thermofisher), RevertAid (Thermofisher), or any of Superscript I-IV reverse transcriptases (Thermofisher). The reverse transcriptase used may or may not have RNase H activity and / or template switching ability. The concentration of dNTPs used during reverse transcription typically ranges from about 0.5 to about 1 mM per dNTP. Reverse transcription may be carried out using oligo-dT, random hexamer primers, or gene-specific primers. The temperature of the reverse transcription reaction may vary, but is typically between 37°C and 55°C. The amount of RNA that serves as a template in a typical reverse transcription reaction can range from picograms of RNA template to micrograms of RNA template. For example, the amount of RNA template can be less than 1 picogram of RNA.

[0029] In some embodiments of the methods disclosed herein, the population of RNA molecules comprises RNA molecules with different sequences and / or RNA molecules with the same sequence.

[0030] In some embodiments of the methods disclosed herein, the population of RNA molecules analyzed comprises at least 1 individual RNA molecule, 10 individual RNA molecules, 100 individual RNA molecules, at least 1,000 individual RNA molecules, at least 10,000 individual RNA molecules, at least 25,000 individual RNA molecules, at least 50,000 individual RNA molecules, at least 75,000 individual RNA molecules, at least 100,000 individual RNA molecules, at least 250,000 individual RNA molecules. individual RNA molecules, at least 500,000 individual RNA molecules, at least 750,000 individual RNA molecules, at least 1,000,000 individual RNA molecules, at least 10,000,000 individual RNA molecules, at least 100,000,000 individual RNA molecules, at least 1,000,000,000 individual RNA molecules, at least 10,000,000,000 individual RNA molecules, or at least 100,000,000,000 individual RNA molecules. In a preferred embodiment, the population of RNA molecules analyzed comprises at least 100,000 individual RNA molecules.

[0031] In some embodiments of the methods disclosed herein, the population of RNA molecules analyzed comprises between 1 and 1,000 individual RNA molecules, between 1 and 10,000 individual RNA molecules, between 1 and 25,000 individual RNA molecules, between 1 and 50,000 individual RNA molecules, between 1 and 100,000 individual RNA molecules, between 1 and 250,000 individual RNA molecules, between 1 and 500,000 individual RNA molecules, between 1 and 750,000 individual RNA molecules, between 1 and 1,000,000 individual RNA molecules, between 1 and 10,000,000 individual RNA molecules, between 1 and 100,000,000 individual RNA molecules, between 1 and 1,000,000,000 individual RNA molecules, between 1 and 10,000,000,000 individual RNA molecules, or between 1 and 100,000,000,000 individual RNA molecules. Preferably, the population of RNA molecules to be analysed comprises between 100 and 1,000,000,000,000 individual RNA molecules, more preferably between 1,000 and 1,000,000,000 individual RNA molecules, and most preferably between 100,000 and 100,000,000 individual RNA molecules.

[0032] In some embodiments of the methods disclosed herein, the one or more RNA molecules are 1-10 copies, 1-20 copies, 1-30 copies, 1-40 copies, 1-50 copies, 1-60 copies, 1-70 copies, 1-80 copies, 1-90 copies, 1-100 copies, 1-125 copies, 1-150 copies, 1-175 copies, 1-200 copies, 1-225 copies, 1-250 copies, 1-275 copies, 1-300 copies, 1-400 copies, 1-500 copies, 1-600 copies, 1-700 copies, 1- The RNA molecules are present in a population of RNA molecules at a copy number of 800 copies, 1-900 copies, 1-1,000 copies, 1-2,000 copies, 1-3,000 copies, 1-3,000 copies, 1-4,000 copies, 1-5,000 copies, 1-10,000 copies, 1-25,000 copies, 1-50,000 copies, 1-75,000 copies, 1-100,000 copies, 1-200,000 copies, 1-300,000 copies, 1-400,000 copies, 1-500,000 copies, or greater than 500,000 copies. Preferably, the one or more RNA molecules are present in the population of RNA molecules in a copy number of from 1 to 500,000 copies, more preferably from 1 to 250,000 copies, even more preferably from 1 to 100,000 copies, even more preferably from 1 to 50,000 copies, and most preferably from 1 to 5,000 copies.

[0033] In some embodiments of the methods disclosed herein, the size range of the RNA molecules in the population is between 100 base pairs and 1,000 base pairs, between 100 base pairs and 2,000 base pairs, between 100 base pairs and 3,000 base pairs, between 100 base pairs and 4,000 base pairs, between 100 base pairs and 5,000 base pairs, between 100 base pairs and 6,000 base pairs, between 100 base pairs and 7,000 base pairs, between 100 base pairs and 8,000 base pairs, between 100 base pairs and 9,000 base pairs, between 100 base pairs and 10,000 base pairs, between 100 base pairs and 11,000 base pairs, base pairs, 100 base pairs to 12,000 base pairs, 100 base pairs to 13,000 base pairs, 100 base pairs to 14,000 base pairs, 100 base pairs to 15,000 base pairs, 100 base pairs to 16,000 base pairs, 100 base pairs to 17,000 base pairs, 100 base pairs to 18,000 base pairs, 100 base pairs to 19,000 base pairs, 100 base pairs to 20,000 base pairs, 500 base pairs to 20,000 base pairs, 1,000 base pairs to 20,000 base pairs, or 2,000 base pairs to 20,000 base pairs.

[0034] In some embodiments of the methods disclosed herein, the population of RNA molecules may be from a single cell, multiple cells or cell populations, tissue, or bodily fluids such as blood, cerebrospinal fluid, or urine, hi some embodiments, the population of RNA molecules is from viral particles.

[0035] The population of RNA molecules can be from any cell. In some embodiments, the cell is a eukaryotic cell (e.g., from a metazoan, a plant, or a fungus), a bacterial cell (i.e., from a eubacteria), or an archaeal cell (i.e., from an archaea). In some embodiments, the population of RNA molecules is from a subcellular compartment of the cell. For example, in a eukaryotic cell, the population of RNA molecules can be from a compartment such as the nucleus, the cytoplasm, the mitochondria, or the chloroplast.

[0036] In some embodiments of the methods disclosed herein, the population of RNA molecules comprises one or more RNA molecules selected from the group consisting of messenger RNA (mRNA), precursor mRNA (pre-mRNA), antisense RNA (asRNA) and its precursors, enhancer RNA and its precursors, long non-coding RNA (lncRNA) and its precursors, microRNA (miRNA) and its precursors, ribosomal RNA (rRNA) and its precursors, transfer RNA (tRNA) and its precursors, histone RNA and its precursors, small nucleolar RNA (snoRNA) and its precursors, small nuclear RNA (snRNA) and its precursors, mitochondrial RNA and its precursors, viral RNA, transposon RNA, synthetic RNA, in vitro transcribed RNA, or a combination thereof.

[0037] In some embodiments of the methods disclosed herein, prior to step (i), the population of RNA molecules is purified and / or enriched for a particular class of RNA molecules. For example, the population of RNA molecules can be enriched for pre-mRNA and / or mRNA molecules.

[0038] In some embodiments of the methods disclosed herein, step (ii) is performed at a concentration of from about 0.5% to about 99.5%, from about 2% to about 98%, from about 3% to about 97%, from about 4% to about 96%, from about 5% to about 95%, from about 6% to about 94%, from about 7% to about 93%, from about 8% to about 92%, from about 9% to about 91%, from about 10% to about 90%, from about 11% to about 89%, from about 12% to about 88%, from about 13% to about 87%, or from about 14% to about 99%. , about 14% to about 86%, about 15% to about 85%, about 16% to about 84%, about 17% to about 83%, about 18% to about 82%, or about 19% to about 81%, about 20% to about 80%, about 25% to about 75%, about 30% to about 70%, about 35% to about 65%, about 40% to about 60%, about 45% to about 55%, about 50% to about 99.5%, about 55% to about 99.5%, about 60% to about 99.5 %, 65% to 99.5%, 70% to 99.5%, 75% to 99.5%, 80% to 99.5%, 81% to 99.5%, 82% to 99.5%, 83% to 99.5%, 84% to 99.5%, 85% to 99.5%, 86% to 99.5%, 87% to 99.5%, 88% to 99.5%, 89% to 99.5% , about 90% to about 99.5%, about 91% to about 99.5%, about 92% to about 99.5%, about 93% to about 99.5%, about 94% to about 99.5%, about 95% to about 99.5%, about 96% to about 99.5%, about 97% to about 99.5%, about 98% to about 99.5%, or about 99% to about 99.5%. Preferably, step (ii) comprises introducing one or more base substitutions into each DNA molecule at a total ratio of about 0.5% to about 99.5%, more preferably about 2% to about 98%, even more preferably about 5% to about 95%, even more preferably about 5% to about 50%, and even more preferably about 5% to about 20%. Most preferably, step (ii) involves introducing one or more base alterations into each DNA molecule at a total ratio of about 15% to about 30%.

[0039] In some embodiments of the methods disclosed herein, step (ii) comprises introducing one or more base changes into each DNA molecule at a total ratio of at least 0.5%, 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50%. Preferably, step (ii) comprises introducing one or more base changes into each DNA molecule at a total ratio of at least 0.5%, more preferably at least 1%, even more preferably at least 3%, even more preferably at least 5%. Most preferably, step (ii) comprises introducing one or more base changes into each DNA molecule at a total ratio of at least 15%.

[0040] The base conversion rate per molecule is measured as the percentage of the total sequenced bases converted in each DNA molecule (and its amplified progeny DNA molecules) produced by reverse transcription relative to the corresponding RNA molecules in the initial population of RNA molecules. For example, base conversion rates are often used in terms of the conversion rate per eligible base. For example, 50% C to T conversion indicates that 50% of cytosines are converted to thymines.

[0041] In some embodiments of the methods disclosed herein, step (ii) comprises reverse transcription in the presence of one or more base analogues.

[0042] In a preferred embodiment of the methods disclosed herein, the one or more base analogs are selected from the group consisting of: [ka] 2'-Deoxy-P-nucleoside-5'-triphosphate (dPTP) (TriLink: N-2037; Jena Bioscience; NU-1119) [ka] 8-oxo-2'-deoxyguanosine-5'-triphosphate (8-oxo-GTP) (Trilink: N-2034), [ka] 2-thiothymidine-5'-triphosphate (2-thioTTP) (TriLink:N-2035), [ka] 5-Formyl-2'-deoxyuridine-5'-triphosphate (TriLink:N-2067) [ka] 5-Propynyl-2'-deoxycytidine-5'-triphosphate (TriLink:N-2016) [ka] 5-Iodo-2'-deoxycytidine-5'-triphosphate (TriLink:N-2023) [ka] 5-Propargylamino-2'-deoxyuridine-5'-triphosphate (N-2062) Or a combination thereof.

[0043] In some embodiments of the methods disclosed herein, step (ii) comprises reverse transcription in the presence of a suboptimal amount of one or more dNTP bases.

[0044] "Suboptimal amount of one or more dNTP bases" includes the meaning of dNTP bases at a concentration lower than that typically used in reverse transcription reactions. Reverse transcription reactions generally contain dNTPs at a concentration ranging from 0.2 mM to 0.5 mM. It is also possible to use higher concentrations of dNTPs (e.g., 0.5 mM to 1 mM) in reverse transcription reactions. "Suboptimal amount of one or more dNTP bases" also includes the meaning of dNTP bases having a different (i.e., lower or higher) concentration relative to one or more of the other dNTPs in the reaction mixture. It will be understood that performing reverse transcription in the presence of base analogs and with a suboptimal amount of one or more dNTP bases may result in the incorporation of errors into the sequence of the resulting DNA molecule.

[0045] In some embodiments of the methods disclosed herein, step (ii) comprises reverse transcription in the presence of one or more dNTP bases at a concentration of less than 0.5 mM, less than 0.4 mM, less than 0.3 mM, less than 0.2 mM, or less than 0.1 mM. Preferably, step (ii) comprises reverse transcription in the presence of one or more dNTP bases at a concentration of less than 0.3 mM, more preferably less than 0.2 mM, and most preferably less than 0.1 mM.

[0046] In some embodiments of the methods disclosed herein, step (ii) comprises reverse transcription in the presence of one or more dNTP bases at a concentration of at least 0.1 mM, at least 0.2 mM, at least 0.3 mM, at least 0.4 mM, at least 0.5 mM, at least 0.6 mM, at least 0.7 mM, at least 0.8 mM, at least 0.9 mM, at least 1 mM, at least 1.1 mM, at least 1.2 mM, at least 1.3 mM, at least 1.4 mM, or at least 1.5 mM. Preferably, step (ii) comprises reverse transcription in the presence of one or more dNTP bases at a concentration of at least 0.5 mM, more preferably at least less than 1 mM, and most preferably at least less than 1.5 mM.

[0047] In some embodiments, the method comprises incorporating one or more base analogs into one or more RNA molecules in the population of RNA molecules prior to step (i). In some embodiments, the one or more base analogs are 4-thio-uridine.

[0048] In some embodiments of the methods disclosed herein, step (ii) further comprises chemically modifying the RNA molecule population before subjecting the RNA molecule population to reverse transcription. It will be understood that such chemical modification may result in the incorporation of errors into the sequence of the resulting DNA molecule. It is also possible to edit the RNA molecule population before subjecting the RNA molecule population to reverse transcription using an editing enzyme such as APOBEC1, which can deaminate RNA cytosines, resulting in C to T editing (Gruenewald et al, 2019. Nature, 569:433-437). Another possibility is to incorporate base analogs such as 4-thio-uridine into the RNA during transcription of those molecules. For example, compounds such as 4-thio-uridine can be incorporated by cell transcription by introducing them into the cell medium during culture. Alternatively, such compounds can be incorporated during in vitro transcription as part of the process of sequencing library preparation, for example as used in CEL-seq and CEL-seq2 (Hashimshony et al, 2012. Cell Rep., 2(3):666-73; Hashimshony et al, 2016. Genome Biol., 17:77). Incorporation of 4-thio-uridine into RNA provides a target for subsequent chemical modification by alkylation with iodoacetamide (Herzog et al, 2017. Nat. Methods, 14(12):1198-1204). Alternatively, for example, the oxidizing agent NaIO 3Alternatively, 4-thio-uridine-containing RNA can be subjected to oxidative aromatic nucleophilic substitution using mCPBA and the nucleophile 2,2,2-trifluoroethylamine (Schofield et al., 2018. Nat. Methods 15, 221-225). These different modifications of the 4-thio-uridine base mimic cytosine, resulting in the incorporation of guanidine instead of adenosine during reverse transcription, resulting in a unique pattern of errors or base conversions in the cDNA derived from each RNA molecule. After amplification, such patterns can be used to identify the source molecules for short reads corresponding to portions of these RNA molecules.

[0049] "Chemically modifying" includes the meaning of a process that changes the chemical composition and / or structure of an RNA or DNA molecule. In particular, in the context of the present application, chemical modification relates to a process that results in a change in the chemical composition and / or structure of the nitrogenous base components of an RNA or DNA molecule. The frequency of base conversion can be adjusted by the incorporation during reverse transcription of non-canonical bases with altered susceptibility / resistance to chemical modification.

[0050] In some embodiments, chemically modifying the population of RNA molecules comprises alkylating the population of RNA molecules, hi some embodiments, the alkylation is by iodoacetamide treatment or oxidative aromatic nucleophilic substitution.

[0051] In some embodiments of the methods disclosed herein, step (ii) further comprises the substep of chemically modifying the population of DNA molecules produced by reverse transcription.

[0052] In some embodiments, the chemical modification of the population of DNA molecules generated by reverse transcription comprises a deamination reaction. In certain embodiments, the deamination is carried out using one or more selected from the list consisting of bisulfite treatment, reduction of (previously modified) nucleosides with pyridine borane or its derivative 2-picoline-borane (Liu Y. et al. 2019. Nature Biotechnology 37:424-429), or the use of enzymatic deamination strategies such as, for example, APOBEC treatment.

[0053] In the case of bisulfite treatment of DNA produced by reverse transcription, the treatment results in the conversion of unmethylated cytosines to uracil, while methylated cytosines are unaffected. Thus, by incorporating a given percentage of methylated cytosines during reverse transcription and then performing bisulfite treatment, it is possible to obtain a partially converted library with a high percentage of C to T base conversions.

[0054] In some embodiments of the methods disclosed herein, step (ii) comprises reverse transcription using an error-prone reverse transcriptase.

[0055] By "error-prone reverse transcriptase" we include the meaning of a reverse transcriptase that introduces a base alteration into the complementary strand of a DNA molecule produced by reverse transcription against an RNA template sequence.

[0056] In some embodiments of the methods disclosed herein, the error-prone reverse transcriptase has at least 1 error per 100 bases, at least 2 errors per 100 bases, at least 3 errors per 100 bases, at least 4 errors per 100 bases, at least 5 errors per 100 bases, at least 6 errors per 100 bases, at least 7 errors per 100 bases, at least 8 errors per 100 bases, at least 9 errors per 100 bases, at least 10 errors per 100 bases, at least 11 errors per 100 bases, at least 12 errors per 100 bases, at least 13 errors per 100 bases, at least 14 errors per 100 bases, at least 15 errors per 100 bases, at least 16 errors per 100 bases, at least 17 errors per 100 bases, at least 18 errors per 100 bases, at least 19 errors per 100 bases, at least 20 errors per 100 bases, at least 21 errors per 100 bases, at least 22 errors per 100 bases, at least 23 errors per 100 bases, at least 24 errors per 100 bases, at least 25 errors per 100 bases, at least 26 errors per 100 bases, at least 27 errors per 100 bases, at least 28 errors per 100 bases, at least 29 errors per 100 bases, at least 30 errors per 100 bases, at least 31 errors per 100 bases, at least 32 errors per 100 bases, at least 33 errors per 100 bases, at least 34 errors per 100 bases, at least 35 errors per 100 bases, at least 36 errors per 100 base at least 14 errors per 100 bases, at least 15 errors per 100 bases, at least 16 errors per 100 bases, at least 17 errors per 100 bases, at least 18 errors per 100 bases, at least 19 errors per 100 bases, at least 20 errors per 100 bases, at least 25 errors per 100 bases, at least 30 errors per 100 bases, at least 35 errors per 100 bases, at least 40 errors per 100 bases, at least 45 errors per 100 bases, at least 50 errors per 100 bases, at least 55 errors per 100 bases, or at least 60 errors per 100 bases.

[0057] Error-prone reverse transcriptases can be produced using approaches known in the art of molecular biology and protein engineering. The most commonly used strategies for protein engineering are rational protein design (i.e., using knowledge of the protein's function and / or sequence to make defined amino acid changes) and directed evolution (i.e., using rounds of random mutagenesis and selection based on desired characteristics), and combinations of each approach are often used by researchers. Modified reverse transcriptases with increased incorporation of modified bases can also be produced using approaches known in the art of molecular biology and protein engineering (e.g., Zhou et al, 2019. Nat. Methods, 16, 1281-1288).

[0058] In some embodiments of the methods disclosed herein, step (iii) comprises amplifying the population of DNA molecules from step (ii) to generate one or more amplicons for each DNA molecule in the population.

[0059] By "amplicon" we include the meaning of a DNA molecule that is amplified from a DNA template, eg, a PCR product.

[0060] In some embodiments of the methods disclosed herein, amplifying the population of DNA molecules comprises high-fidelity amplification.

[0061] In some embodiments of the methods disclosed herein, amplifying the population of DNA molecules comprises PCR amplification.

[0062] "High fidelity amplification" includes the meaning of amplification that results in amplicons that have very little or no sequence variation relative to the corresponding sequence in the original template molecule (e.g., the original cDNA molecule). Such high fidelity amplification can be performed using commercially available proofreading DNA polymerase enzymes.

[0063] In some embodiments of the method disclosed herein, a non-proofreading DNA polymerase enzyme is used during second strand cDNA synthesis, and then a high-fidelity proofreading DNA polymerase enzyme is used for amplifying the DNA molecule population. A non-proofreading DNA polymerase enzyme (e.g., Taq DNA polymerase) is preferred for second strand cDNA synthesis because it is more likely to tolerate the presence of non-canonical bases in the cDNA first strand, and thus introduce base conversions into the cDNA second strand. A proofreading DNA polymerase is preferred for amplifying the DNA molecule population because it is more likely to maintain the base conversion pattern introduced during the error-prone reverse transcription step.

[0064] In some embodiments of the methods disclosed herein, the step of amplifying the population of DNA molecules is performed in the absence of base analogs.

[0065] In some embodiments of the methods disclosed herein, at least the first cycle of amplifying the population of DNA molecules is performed in the presence of a suboptimal amount of one or more dNTP bases.

[0066] In some embodiments, at least the first cycle of amplifying the population of DNA molecules is performed in the presence of one or more dNTP bases at a concentration of less than 0.5 mM, less than 0.4 mM, less than 0.3 mM, less than 0.2 mM, or less than 0.1 mM. Preferably, at least the first cycle of amplifying the population of DNA molecules is performed in the presence of one or more dNTP bases at a concentration of less than 0.3 mM, more preferably less than 0.2 mM, and most preferably less than 0.1 mM.

[0067] In some embodiments of the methods disclosed herein, at least the first cycle of amplifying the population of DNA molecules is performed in the presence of one or more dNTP bases at a concentration of at least 0.1 mM, at least 0.2 mM, at least 0.3 mM, at least 0.4 mM, at least 0.5 mM, at least 0.6 mM, at least 0.7 mM, at least 0.8 mM, at least 0.9 mM, at least 1 mM, at least 1.1 mM, at least 1.2 mM, at least 1.3 mM, at least 1.4 mM, or at least 1.5 mM. Preferably, at least the first cycle of amplifying the population of DNA molecules is performed in the presence of one or more dNTP bases at a concentration of at least 0.5 mM, more preferably at least 1 mM, and most preferably at least 1.5 mM.

[0068] After incorporation of one or more base analogs into the first strand cDNA, varying amounts of individual dNTPs (e.g., suboptimal or non-uniform amounts of one or more dNTPs) can be used in the first cycle of amplification, in which the first strand cDNA serves as a template for amplification, and by varying the amounts of dNTPs relative to each other in the reaction, it is possible to bias the base analogs in the first strand cDNA towards preferential pairing with one base over another, thereby affecting the identity of the conversion events and / or changing the overall conversion rate at sites in the first strand cDNA that bear the base analog.

[0069] It will be appreciated that the use of suboptimal amounts of one or more dNTP bases during at least the first cycle of the amplification step is applicable to the amplification of any polynucleotide that contains one or more base analogues, and similarly, such base analogues can be biased towards preferential pairing with one base over other bases.

[0070] Accordingly, a further aspect of the invention is a method for generating a base alteration in one or more polynucleotide molecules within a population of polynucleotide molecules, comprising the steps of: (i) providing a population of polynucleotide molecules, one or more of the polynucleotide molecules comprising one or more base analogs; (ii) amplifying the population of polynucleotide molecules from step (i) to produce one or more amplicons for each polynucleotide molecule in the population, wherein at least a first cycle of the amplifying step is performed in the presence of a suboptimal amount of one or more dNTP bases.

[0071] In some embodiments of the methods for generating a base alteration in one or more polynucleotide molecules within a population of polynucleotide molecules, the one or more polynucleotide molecules are cDNA molecules, DNA molecules, or RNA molecules (including double-stranded RNA molecules).

[0072] In some embodiments of the method for generating base conversions in one or more polynucleotide molecules in a population of polynucleotide molecules, at least the first cycle of amplifying the population of polynucleotide molecules is performed in the presence of one or more dNTP bases at a concentration of less than 0.5 mM, less than 0.4 mM, less than 0.3 mM, less than 0.2 mM, or less than 0.1 mM. Preferably, at least the first cycle of amplifying the population of polynucleotide molecules is performed in the presence of one or more dNTP bases at a concentration of less than 0.3 mM, more preferably less than 0.2 mM, and most preferably less than 0.1 mM.

[0073] In some embodiments of the method for generating base conversions in one or more polynucleotide molecules in a population of polynucleotide molecules, at least the first cycle of amplifying the population of polynucleotide molecules is performed in the presence of one or more dNTP bases at a concentration of at least 0.1 mM, at least 0.2 mM, at least 0.3 mM, at least 0.4 mM, at least 0.5 mM, at least 0.6 mM, at least 0.7 mM, at least 0.8 mM, at least 0.9 mM, at least 1 mM, at least 1.1 mM, at least 1.2 mM, at least 1.3 mM, at least 1.4 mM, or at least 1.5 mM. Preferably, at least the first cycle of amplifying the population of polynucleotide molecules is performed in the presence of one or more dNTP bases at a concentration of at least 0.5 mM, more preferably at least 1 mM, and most preferably at least 1.5 mM.

[0074] In some embodiments of the methods for generating base alterations in one or more polynucleotide molecules within a population of polynucleotide molecules, amplifying the population of polynucleotide molecules comprises high-fidelity amplification.

[0075] In some embodiments of the methods for generating a base alteration in one or more polynucleotide molecules within a population of polynucleotide molecules, amplifying the population of polynucleotide molecules comprises PCR amplification.

[0076] In some embodiments of the methods for generating a base alteration in one or more polynucleotide molecules within a population of polynucleotide molecules, the step of amplifying the population of polynucleotide molecules is performed in the absence of base analogs.

[0077] In embodiments in which reverse transcription is performed in the presence of one or more base analogs and chemical modification is performed on the population of RNA molecules prior to reverse transcription or chemical modification is performed on the population of DNA molecules produced by reverse transcription, the conditions or treatments that induce base conversion are removed prior to the amplification step. For example, if base analogs are used in the reverse transcription step, any unincorporated base analog molecules are removed (or degraded) prior to amplification by methods such as dilution, phenol-chloroform extraction, bead cleanup, enzyme removal, and / or thermal degradation.

[0078] In some embodiments of the methods disclosed herein, step (iii) comprises fragmenting the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population to generate overlapping fragments. In some embodiments, the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population are purified prior to fragmentation.

[0079] In some embodiments of the methods disclosed herein, fragmenting the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population comprises tagging, DNA shearing, and / or enzymatic fragmentation.

[0080] By "tagging" we include the meaning of the process of incorporating sequencing adaptors into DNA using a transposase, e.g., the incorporation of partial sequencing adaptors.

[0081] In some embodiments of the methods disclosed herein, the fragments are from about 50 base pairs to about 2000 base pairs in length, from about 50 base pairs to about 1900 base pairs in length, from about 50 base pairs to about 1800 base pairs in length, from about 50 base pairs to about 1700 base pairs in length, from about 50 base pairs to about 1600 base pairs in length, from about 50 base pairs to about 1500 base pairs in length, from about 50 base pairs to about 1400 base pairs in length, from about 50 base pairs to about 1300 base pairs in length, from about 50 base pairs to about 1200 base pairs in length, from about 50 base pairs to about 1100 base pairs in length, from about 50 base pairs to about 1000 base pairs in length, from about 50 base pairs to about 950 base pairs in length, base pairs length, about 50 base pairs to about 900 base pairs length, about 50 base pairs to about 850 base pairs length, about 50 base pairs to about 800 base pairs length, about 50 base pairs to about 750 base pairs length, about 50 base pairs to about 700 base pairs length, about 50 base pairs to about 650 base pairs length, about 50 base pairs to about 600 base pairs length, about 50 base pairs to about 550 base pairs length, about 50 base pairs to about 500 base pairs length, about 50 base pairs to about 450 base pairs length, about 50 base pairs to about 400 base pairs length, about 50 base pairs to about 350 base pairs length, about 50 base pairs to about 300 base pairs length, about 50 base pairs to about 25 0 base pairs in length, about 50 base pairs to about 200 base pairs in length, about 50 base pairs to about 150 base pairs in length, about 50 base pairs to about 100 base pairs in length, about 100 base pairs to about 1500 base pairs in length, about 150 base pairs to about 1400 base pairs in length, about 200 base pairs to about 1300 base pairs in length, about 250 base pairs to about 1200 base pairs in length, about 300 base pairs to about 1100 base pairs in length, about 350 base pairs to about 1000 base pairs in length, about 400 base pairs to about 1000 base pairs in length, about 450 base pairs to about 950 base pairs in length, about 500 base pairs to about 900 base pairs in length, about 550 base pairs to about 8 50 base pairs in length, about 600 base pairs to about 800 base pairs in length, about 650 base pairs to about 750 base pairs in length, about 700 base pairs to about 1500 base pairs in length, about 750 base pairs to about 1500 base pairs in length, about 800 base pairs to about 1500 base pairs in length, about 850 base pairs to about 1500 base pairs in length, about 900 base pairs to about 1500 base pairs in length, about 950 base pairs to about 1500 base pairs in length, about 1000 base pairs to about 1500 base pairs in length, about 1100 base pairs to about 1500 base pairs in length, about 1200 base pairs to about 1500 base pairs in length, about 1300 base pairs to about 1500 base pairs in length,or from about 1400 base pairs to about 1500 base pairs in length. Preferably, the fragment is from about 50 base pairs to about 1500 base pairs in length, more preferably from 50 base pairs to 1200 base pairs in length, even more preferably from 50 base pairs to 1000 base pairs in length, and most preferably from 50 base pairs to 800 base pairs in length.

[0082] "Overlapping fragment" includes the meaning of any overlapping portion of at least two DNA sequences. The sequence containing the overlapping portion can be derived from a sequence obtained directly from a short-read sequencing experiment (i.e., single-end or paired-end reads) or from a partially reconstructed DNA sequence. The partial reconstruction of the DNA sequence can be achieved, for example, using molecular barcodes or iteratively using the methods disclosed herein.

[0083] In some embodiments of the methods disclosed herein, the length of the overlapping sequences required to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern is at least 10 base pairs, at least 15 base pairs, at least 20 base pairs, at least 25 base pairs, at least 30 base pairs, at least 35 base pairs, at least 40 base pairs, at least 45 base pairs, at least 50 base pairs, at least 55 base pairs, at least 60 base pairs, at least 65 base pairs, at least 70 base pairs, at least 75 base pairs, at least 80 base pairs, at least 85 base pairs, at least 90 base pairs, at least 95 base pairs, at least 100 base pairs, at least 125 base pairs, at least 150 base pairs, at least 175 base pairs, or at least 200 base pairs. Preferably, the length of the overlapping sequences required to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern is at least 200 base pairs, more preferably at least 100 base pairs, even more preferably at least 75 base pairs, and most preferably at least 50 base pairs.

[0084] In some embodiments of the methods disclosed herein, the length of overlapping sequences required to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern is less than 500 base pairs, less than 450 base pairs, less than 400 base pairs, less than 350 base pairs, less than 300 base pairs, less than 250 base pairs, less than 200 base pairs, less than 175 base pairs, less than 150 base pairs, less than 125 base pairs, less than 100 base pairs, less than 95 base pairs, less than 90 base pairs, less than 85 base pairs, less than 80 base pairs, less than 75 base pairs, less than 70 base pairs, less than 65 base pairs, less than 60 base pairs, less than 55 base pairs, less than 50 base pairs, less than 45 base pairs, less than 40 base pairs, less than 35 base pairs, less than 30 base pairs, less than 25 base pairs, less than 20 base pairs, less than 15 base pairs, or less than 10 base pairs. Preferably, the length of overlapping sequences required to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern is less than 500 bases, more preferably less than 300 bases, even more preferably less than 200 base pairs, and most preferably less than 100 base pairs.

[0085] In some embodiments of the methods disclosed herein, the length of overlapping sequences required to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern can be between 10 base pairs and 500 base pairs in length, between 15 base pairs and 450 base pairs in length, between 20 base pairs and 400 base pairs in length, between 25 base pairs and 350 base pairs in length, between 30 base pairs and 300 base pairs in length, between 35 base pairs and 250 base pairs in length, between 40 base pairs and 200 base pairs in length, between 45 base pairs and 175 base pairs in length, between 50 base pairs and 150 base pairs in length, between 55 base pairs and 125 base pairs in length, between 60 base pairs and 100 base pairs in length, between 65 base pairs and 95 base pairs in length, The length is 70 base pairs to 90 base pairs, 75 base pairs to 90 base pairs, 80 base pairs to 85 base pairs, 90 base pairs to 500 base pairs, 95 base pairs to 500 base pairs, 100 base pairs to 500 base pairs, 125 base pairs to 500 base pairs, 150 base pairs to 500 base pairs, 175 base pairs to 500 base pairs, 200 base pairs to 500 base pairs, 250 base pairs to 500 base pairs, 300 base pairs to 500 base pairs, 350 base pairs to 500 base pairs, 400 base pairs to 500 base pairs, or 450 base pairs to 500 base pairs. Preferably, the length of overlapping sequences necessary to identify and assemble overlapping fragments having the same molecule-specific base conversion pattern is between 10 base pairs and 500 base pairs in length, more preferably between 25 base pairs and 250 base pairs in length, even more preferably between 50 base pairs and 150 base pairs in length, and most preferably between 50 base pairs and 100 base pairs in length.

[0086] In some embodiments of the methods disclosed herein, step (iii) comprises sequencing overlapping fragments of the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population. In some embodiments, the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population are purified prior to sequencing.

[0087] In some embodiments, the population of DNA molecules and / or one or more amplicons of each DNA molecule within the population are purified prior to fragmentation and / or sequencing.

[0088] If fragmentation is performed, indexing and library amplification or PCR-free ligation is performed. In the context of the methods described herein, indexing involves adding a specific molecular sample barcode to a sequencing library derived from a specific population of RNA molecules. Such sample indexing allows multiple libraries derived from different starting populations of RNA molecules to be sequenced in parallel (e.g., on a flow cell), and is then used to associate sequence reads with the correct population of RNA molecules. Sample barcodes can be added to oligo-dT primers or template switching oligos, and thus are present at the ends of cDNA molecules generated using such oligos. In such a strategy, only a subset of paired end sequence reads covering the 5' or 3' ends of the molecules will have a cell / sample barcode, and internal read pairs will not have a barcode. Alternatively, sample barcodes can be added after tagging (e.g., post-tagging PCR oligos), such that all sequences in the library have a barcode (i.e., both the 5' and 3' terminal fragments and the internal fragments have a barcode).

[0089] "Molecular barcode" includes the meaning of a pool of nucleic acid sequences that can be added to a particular population of RNA or DNA molecules and act as a unique identifier that allows grouping of amplified DNA sequences derived from the same initial RNA or DNA molecule. Molecular barcodes are added prior to cDNA amplification, and they are typically included in oligo or oligo-dT switching templates. Molecular barcodes may also be referred to as unique molecular identifiers (UMIs), and they are often stretches of 4-25 random nucleotides.

[0090] Using a library where all paired end reads have a sample barcode can aid in reconstructing the sequence of the RNA molecules in a population of RNA molecules, as the search space for finding unique base conversion patterns is smaller. However, it is still possible to effectively reconstruct the RNA sequence using a library without sample barcodes on the internal paired end reads. In the present invention, molecular barcodes are not required, as the base conversion patterns introduced in the error-prone reverse transcription step are superior to traditional UMIs. Thus, the methods disclosed herein can be performed using libraries where no molecules have added molecular barcodes, a subset of molecules have added molecular barcodes, or all molecules have added molecular barcodes. Furthermore, the methods disclosed herein can be performed using libraries where no molecules have added sample barcodes, a subset of molecules have added sample barcodes, or all molecules have added molecular barcodes.

[0091] In some embodiments of the methods disclosed herein, the sequencing comprises a short read sequencing method.

[0092] "Short-read sequencing methods" include those that do not cover the entire sequenced molecule in a single sequencing read. Short-read sequencing typically produces sequencing reads that are about 50 base pairs to about 400 base pairs in length.

[0093] In some embodiments, the short-read sequencing method is selected from the list consisting of massively parallel short-read sequencing, DNA nanoball sequencing, Illumina dye sequencing (Solexa sequencing), 454 pyrosequencing, SOLiD sequencing, Helicos single-molecule fluorescent sequencing, combinatorial probe anchor synthesis (cPAS), polony sequencing, electrical sequencing chips (e.g., GenapSys), or combinations thereof.

[0094] In some embodiments of the methods disclosed herein, step (iv) comprises: (a) assigning overlapping segments to RNA molecules present in a population of RNA molecules based on their alignment to some or all of the sequences of that RNA molecule; and / or (b) selecting the assigned fragments based on the positions in the RNA molecule to which they align.

[0095] The number of overlapping DNA fragments (and their respective lengths) required to obtain sequence reads covering the entire length of the RNA present in the initial RNA molecule population depends on the sequencing strategy used. Typically, as the average length of the generated reads increases, the probability of obtaining longer overlaps increases, and vice versa. Thus, there is an interaction between the sequence depth and the short-read sequencing strategy used, and the uniformity of the read pairs obtained over the length of the sequence of a given RNA molecule in the initial RNA molecule population. That interaction ultimately determines the number of paired end reads required to assemble the sequence of a particular RNA molecule.

[0096] The allocation and alignment of overlapping sequence fragments to an RNA molecule and the selection of those fragments based on their alignment position to that RNA molecule can be performed using computational methods. For example, software can be used to map all acquired sequence reads to a database of reference sequences, and then annotate each sequence read (or read pair) based on the population DNA molecule from which the read / read pair originates, for example, using the molecular barcode / UMI present in the read / read pair. The annotated group of sequenced fragments obtained by alignment to the reference sequence can then be selected by the software based on their mapping position on the reference sequence. The position of each base transversion in the aligned fragments is then determined before a probabilistic approach is used to estimate the co-occurrence strength of the pair of base transversions. Based on the co-occurrence information, it is possible to identify groups of fragments that share the same base transversion pattern in a statistically significant manner. The analysis is then repeated until no more reads can be assembled.

[0097] "Reference sequence" includes the meaning of a known sequence, typically from a database, to which sequence reads can be compared and aligned. The reference sequence may or may not be part of a reference genome.

[0098] In some embodiments of the methods disclosed herein, step (v) comprises comparing the sequence information of step (iv) to a reference sequence and identifying mismatches corresponding to one or more base conversions.

[0099] Alignment software can be used to identify the correct alignment position of short reads to a reference sequence despite the presence of many base conversions. Examples of such software include: -STAR(https: / / github.com / alexdobin / STAR), -BWA (https: / / github.com / lh3 / bwa), and -Bowtie (http: / / bowtie-bio.sourceforge.net / index.shtml) is one example.

[0100] Once alignment of sequence reads is performed, base conversions are found based on relative mismatches to the reference sequence. Again, software can be used to "find" induced base conversions. Such software can also distinguish reverse transcription-induced base conversions from mismatches resulting from mutations of RNA molecules in a population of RNA molecules, single nucleotide polymorphisms (SNPs), and PCR / sequencing errors. This is possible because induced base conversions occur at a much higher frequency and are therefore much more common than background sources of mismatches to the reference sequence. Exemplary software that can find induced base conversions uses Samtools and htslib (https: / / github.com / samtools), Pysam (a Python package; https: / / github.com / pysam-developers / pysam), and Rsamtools (an R package; https: / / kasperdanielhansen.github.io / genbioconductor / html / Rsamtools.html) to efficiently load SAM / BAM files and compare reads to a reference sequence to identify read-level mismatches.

[0101] In some embodiments of the methods for determining the copy number of one or more RNA molecules disclosed herein, step (vi) includes determining from the information of step (v) a number of unique molecule-specific base conversion patterns that correspond to RNA molecules having a particular sequence within the population of RNA molecules.

[0102] The first step in the process of determining the number of unique molecule-specific base conversion patterns is pattern imputation. Each sequenced fragment is aligned to a subset of the sequences of the RNA molecules in the RNA molecule population. Thus, each molecule-specific base conversion pattern is incomplete on a read-by-read basis. Therefore, the complete base conversion pattern must be imputed for each read. For example, the reads can be aggregated to construct a matrix of conditional probabilities, where each entry is the estimated probability of observing a base conversion at that position given the known presence of a base conversion at another position. The estimated probability is based on a Bayesian estimator using a beta distribution as a conjugate prior distribution with parameters α=0.1, β=1 for the binomial distribution, and p=(x+α) / (n+α+β), where x is the number of observed base conversions conditioned on n observed reads with base conversions at other positions. In general, α and β can be other values ​​as long as α is small and β is large. Such an estimator is used to account for non-duplicated positions in any read, which results in a small but non-zero probability of observing a base conversion.

[0103] After this matrix is ​​constructed, all base conversions observed in each read are used to impute all positions using this conditional probability matrix. Note that even positions observed in a read can be imputed to account for noise in the sequencing read. We perform two interesting missing value imputations. The first is the most likely value, most likely that a base conversion is present or most likely that a base conversion is absent. The second missing value imputation is the imputed probability of observing a base conversion at that position, which can be used to propagate uncertainty in downstream analyses. The resulting imputed patterns can then be clustered using a preferred clustering algorithm.

[0104] Clustering of the imputed patterns is then performed. The clustering step serves two purposes: (i) to count the number of patterns that effectively count the number of molecules observed, and (ii) to group reads by the molecules used in the full-length reconstruction.

[0105] There are several options for clustering the imputed patterns, including Bernoulli mixture models and density-based clustering.

[0106] Bernoulli mixture model clustering treats each read as a composite of one or more binary patterns found by expectation maximization. Density-based clustering identifies dense regions of binary patterns and connects points in this space by a distance metric. In the context of the methods disclosed herein, any distance metric for binary data is appropriate. For example, Dice dissimilarity, Hamming distance, Jaccard-Needham dissimilarity, Krusinski dissimilarity, Rogers-Tanimoto dissimilarity, Russell-Rao dissimilarity, Sokal-Michener dissimilarity, Sokal-Sneath dissimilarity, or Yule dissimilarity. Examples of algorithms in this category are DBSCAN and OPTICS. Another option is to cluster the imputed probabilities instead of the imputed patterns. The main consideration for the algorithms used in density-based clustering is the distance a point is from a dense region to be part of that cluster. For example, if a point is too far from any dense region, it will not be considered part of any cluster. DBSCAN allows for a tunable ε parameter to adjust this, while OPTICS abstracts this parameter and instead allows you to set the minimum number of points that form a cluster.

[0107] Determining the number of unique molecule-specific base conversion patterns can be achieved by applying a statistical model to all molecule-specific base conversion patterns of sequenced DNA molecules / fragments that align with the sequence of the RNA molecule of interest. The statistical model can be in the form of a python programming language derived from a package such as SciPy (website: www.scipy.org). The main processing steps that such software must perform are: (i) obtaining the base conversion pattern of each DNA molecule / fragment, and (ii) grouping the fragments by their base conversion patterns by a statistical method. Examples of such statistical methods include, but are not limited to, multivariate Bernoulli mixture models, density-based clustering, naive Bayes, and random graph-based methods.

[0108] Another strategy to group sequences by molecule-specific base conversion patterns is to compare each sequence with a set of other sequences using a similarity measure. In the context of this application, the conversion patterns obtained for each sequence or derived from one or more sequences are compared. For example, mutual information or land score metrics can be used as similarity metrics. To avoid false positives, the similarity metric can be adjusted according to the actual number of overlapping eligible positions found in the sequences and using a background model of similarity values ​​that can be attributed to chance alone. As an example, two conversion patterns from two reads with many overlapping eligible positions are easy to statistically assign as arising from the same or different original molecules. However, two conversion patterns from two reads with only three overlapping eligible positions may be perfectly consistent due to chance alone, which needs to be controlled. One such background model is the hypergeometric distribution model, which takes into account the number of overlapping eligible bases and the number of converted positions of both patterns in the overlapping region. Direct examples of adjusted similarity metrics are adjusted mutual information and adjusted rand score, where values ​​close to 0 are consistent with unadjusted similarity scores that occur due to chance, and values ​​close to 1 indicate similarities that do not occur due to chance. Using these adjusted similarity metrics, it is possible to accurately assign sequences according to their base conversion patterns over the entire range of overlapping sequence lengths. More specifically, sequenced fragments are often ordered based on their genomic location and their base conversion patterns. Each fragment can then be compared to all base conversion patterns obtained from a group of previously analyzed sequence fragments (or a merge from previous such comparisons). Often, the threshold used for adjusted similarity metrics ranges from 0.15 to 0.50. Higher values ​​within that range result in a more strict assignment of sequences to each other, while lower thresholds may produce more false positives. A sufficiently good match is often within the range of values ​​from 0.20 to 0.30, with higher values ​​indicating even better matches.The presence of a good adjusted similarity value (i.e., above a set threshold) results in adding the specific fragment to one or more previously grouped sequences, and adding the specific base conversion pattern in that sequence to that group. If there are not enough good matches (i.e., all comparisons result in values ​​below the set threshold), the fragment becomes a new group representing a unique molecule-specific base conversion pattern. The use of such an approach simultaneously provides the number of unique molecule-specific base conversion patterns, the actual patterns themselves, and the sequencing reads that make up each pattern.

[0109] Using the molecule-specific base conversion patterns generated by the methods disclosed herein, it is possible to count RNA molecules after successful (or partial) RNA sequence reconstruction, or to skip the RNA sequence reconstruction (e.g., when sequencing at a lower sequence depth) and locally count RNA molecules based on the molecule-specific base conversion patterns observed around specific base pairs of the DNA / RNA sequence. For example, all reads covering a specific exon-exon junction of a gene may be collected. Then, using the strategy for grouping read sequences by their molecule-specific base conversion patterns described in the previous paragraph, molecules spanning a specific exon-exon junction may be locally reconstructed. Other features of interest may be transcription start sites or polyadenylation sites. Although the counts obtained using the latter strategy may be underestimated due to limited sequencing depth, the approach may be useful for applications such as diagnostics.

[0110] In some embodiments of the methods disclosed herein, one or more of steps (i)-(iii) are performed in a droplet-based environment, a plate-based environment, attached to a bead, or in situ.

[0111] In some embodiments of the methods disclosed herein, the population of RNA molecules comprises one or more sequence variants of the same gene, or one or more allelic variants of the same gene, or one or more splice variants of the same gene, one or more RNA isoforms due to alternative promoter usage, or one or more RNA isoforms due to alternative splice site usage, or one or more RNA isoforms due to alternative polyadenylation site usage.

[0112] In a fifth aspect, the present invention provides the use of an error-prone reverse transcriptase to generate, from a population of RNA molecules, a population of DNA molecules having a molecule-specific base conversion pattern, each DNA molecule comprising one or more base conversions relative to a corresponding RNA molecule, in order to determine the copy number of one or more RNA molecules in the population.

[0113] The first and second aspects disclosed herein provide examples of methods of using error-prone transcription to generate a population of DNA molecules for determining the copy number of one or more RNA molecules in the population.

[0114] In a sixth aspect, the present invention provides the use of an error-prone reverse transcriptase to generate, from a population of RNA molecules, a population of DNA molecules, each DNA molecule containing one or more base conversions relative to a corresponding RNA molecule, and having a molecule-specific base conversion pattern, in order to determine the sequence of one or more RNA molecules in the population.

[0115] The third and fourth aspects disclosed herein provide examples of methods of using error-prone transcription to generate a population of DNA molecules for determining the copy number of one or more RNA molecules in the population.

[0116] In a seventh aspect, the present invention provides a population of DNA molecules obtained or obtainable by a method of the first, second, third or fourth aspect, or by use of the fifth or sixth aspect.

[0117] In an eighth aspect, the present invention provides a kit for performing error-prone reverse transcription, the kit comprising: (i) a reverse transcriptase; (ii) one or more base analogues; and (iii) instructions for use.

[0118] In some embodiments of the kits disclosed herein, the one or more base analogs are selected from the group consisting of 2'-deoxy-P-nucleoside-5'-triphosphate (dPTP), 8-oxo-2'-deoxyguanosine-5'-triphosphate (8-oxo-GTP), 2-thiothymidine-5'-triphosphate (2-thioTTP), 5-formyl-2'-deoxyuridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5-propargylamino-2'-deoxyuridine-5'-triphosphate, or a combination thereof.

[0119] In some embodiments of the kits disclosed herein, the reverse transcriptase is an error-prone reverse transcriptase.

[0120] In some embodiments of the kits disclosed herein, the kit further comprises a composition comprising dNTPs.

[0121] In some embodiments of the kits disclosed herein, the kit further comprises an oligonucleotide primer composition suitable for use in reverse transcription, hi some embodiments, the oligonucleotide primer composition comprises an oligo-dT primer, a random hexamer primer, or a gene specific primer.

[0122] In some embodiments of the kits disclosed herein, the kit further comprises a compound capable of modifying bases on the first strand cDNA, in some embodiments, the compound deaminates nitrogenous bases, for example, using bisulfite.

[0123] In a ninth aspect, the present invention provides a method, or a use, or a population of DNA molecules, or a kit substantially as herein described with reference to the accompanying description, examples, claims and drawings.

[0124] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which: FIG. [Brief description of the drawings]

[0125] [Figure 1] The core techniques that can be used to obtain cDNA with molecular specific conversion patterns are shown below. (A) Direct and incorrect incorporation of a canonical base in the first strand cDNA molecule, for example by an error-prone reverse transcriptase. (B) Incorporation of a wide range of base analogs into the first strand cDNA during reverse transcription. During second strand synthesis, the incorrect canonical base can be incorporated, thus resulting in an error at that position. (C) Incorporation of a protective drug or drug-sensitive base analog into the first strand cDNA during reverse transcription. Subsequent chemical or enzymatic treatments modify either the base analog or the corresponding canonical base. During second strand synthesis, this can lead to incorporation of the incorrect canonical base, which can be detected as an error at this position. [Diagram 2] We present the core steps of the method of the present invention and explain how base conversion patterns can be used to identify sequences from the same initial RNA molecule. By introducing random base conversions during the synthesis of cDNA from an RNA molecule, the RNA molecules can be counted and the RNA sequence can be reconstructed. [Figure 3A-C]Genome browser screenshot of single-cell RNA sequencing data of a representative cell (generated according to Smart-seq3 technology) with induced base conversions in the genes MED27, GUK1 and AP2M1, respectively. In the corresponding experiment, base conversions were induced using 0.5 mM 2'-deoxy-P-nucleoside-5'-triphosphate (dPTP), and the induced base conversion patterns uniquely mark each read originating from the same initial RNA molecule, and reads are grouped into specific molecules based on their 5' molecular barcodes. [Figure 4] We show that reverse transcription in the presence of the base analog dPTP can produce useful levels of base conversions, and that the stability of those base conversions in subsequent steps depends on efficient removal of the base analog after the reverse transcription step. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. (A) Reverse transcription in the presence of the base analog dPTP produces high levels of base conversions as long as the base analog (dPTP) is efficiently removed after reverse transcription, either by bead cleanup or treatment with alkaline phosphatase (FastAP). This panel also shows that the type of base conversion event depends on the strand on which the gene is encoded. (B) We show that the stability of base conversions in sequencing reads corresponding to the same molecule depends on efficient removal of the base analog after reverse transcription, either by bead cleanup or treatment with alkaline phosphatase (FastAP). [Diagram 5]Simulation results of the number of unique base conversion patterns expected (y-axis) in experiments with different base conversion fractions (x-axis) and different overlaps in the DNA fragment (50-200 bp; individual curves in each figure) are shown. The expected number of base conversion patterns was calculated for genes expressed at different RNA copy numbers (10, 100, or 1000; columns) and for experiments where 1-4 of the bases present in the molecule were potentially converted using the same specified individual base conversion fraction (as indicated on the x-axis) applied to 1, 2, 3, or 4 bases (as indicated on the rows) (first row: 1 base; second row: 2 bases, e.g., in the case of dPTP; third row: 3 bases; fourth row: all 4 bases). The dashed line indicates a base conversion fraction of 0.04. [Figure 6] The amount of dPTP-induced base conversion on the plus strand is positively correlated with the applied dose of dPTP during reverse transcription. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. [Figure 7] We show that the base analog dPTP can be incorporated into cDNA on RNA bound to beads captured in droplets using MGI C4. Reverse transcription was performed with the addition of dPTP, and PCR amplification was performed using KAPA HiFi PCR enzyme. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. Please note that in this figure, and in the following figures, unless otherwise specified, the base conversion rates shown are for features on the positive strand. [Figure 8]1 shows base conversions induced by incorporation of different base analogs during reverse transcription. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. (A) Base conversions obtained by incorporation of 2-thioTTP during reverse transcription (performed in biological replicates). Experimental details of the data shown in this figure are described in Example 5 below. (B) Base conversions obtained by incorporation of 5-formyl-2'-deoxyuridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, or 5-propargylamino-2'-deoxyuridine-5'-triphosphate during reverse transcription. [Figure 9] Shown are all induced base conversions for different second strand synthesis approaches performed on cDNA containing dPTP, 5-formyl-dUTP, or only canonical bases (H20 results). It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. [Figure 10] It is shown that different PCR enzymes efficiently incorporate non-canonical bases opposite the canonical dNTP into cDNA. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. [Figure 11] We show that incorporation of a non-canonical base during reverse transcription (here we use a methylated cytosine base) combined with bisulfite treatment of cDNA (resulting in the conversion of unmethylated cytosine to uracil) can result in base conversion in a highly controlled manner. It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. [Figure 12]RNA reconstruction results in the context of single-cell RNA sequencing (see Example 8). (A) Histogram and density plot of the percentage of internal reads that can be assigned to 5' anchor read pairs based on dPTP-induced base conversion. Internal reads are classified as paired end sequenced reads with a first read that does not originate from the RNA 5' end, such that both read fragments capture the internal portion of the RNA. (B) Line plot with the length of reconstructed RNA in experiment 5 (with and without assigning internal reads to 5' anchor reads based on induced base conversion patterns) compared to long-read sequencing of a similar cDNA library (here, sequencing on a Pacific Biosystems Sequel instrument). Reconstruction based on dPTP-induced base conversion allowed us to assign internal reads to 5' anchor RNA reads to reconstruct cDNA of about 1,250 bp with a quality similar to long-read sequencing technology. [Figure 13] We show that dPTP-induced base conversions in single-cell RNA sequencing data can be used to assign sequenced reads to the correct strand. (A) Base conversions observed when separating genes according to their position on the plus or minus strand of a DNA molecule. Two conversions (A to G and G to A) were specifically induced in genes located on the plus strand (and reverse complement conversions for genes located on the minus strand). (B) Log-likelihood ratios of each partially reconstructed sequence being assigned to the correct strand based on base conversions induced by 0.5 mM dPTP. The log-likelihood distributions of reads assigned to genes by plus or minus strand demonstrate that the induced base conversions contain the information necessary to correctly assign the majority of reads to the correct strand. [Figure 14] We provide a schematic of the application of the method of the present invention to count and reconstruct RNA sequences from single cells in the context of Smart-seq3. [Figure 15]In the context of a novel early pooling-based full-length transcriptome sequencing method, a schematic diagram of an application in which the method of the present invention is used to count and reconstruct RNA sequences from single cells is provided. In such an application, the method of the present invention can enable both RNA counting and sequence reconstruction in a highly parallel manner to characterize a large number of single cells. [Figure 16] The cell barcoding approach used in Example 10 is shown, in which not all obtained reads contain cell barcode (and UMI) information, and such experiments rely on molecular pattern identification to link reads to their corresponding cell barcodes. [Figure 17] Using the initial pooling as shown in Figure 16, dPTP-mediated conversions obtained in single cell experiments are shown (see Example 10). It will be understood that the conversion identity describes the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. [Figure 18] 1 shows the cumulative distribution of reconstituted molecule lengths across the entire experiment (n=96 cells) described in Example 10. [Figure 19] For each gene with more than 50 molecules detected across all cells (n=6,684 genes), the percentage of reads without cell barcodes that were successfully linked to a reconstructed molecule is shown. The center line indicates the median, the hinge indicates the first and third quartiles, and the whiskers indicate the 1.5-fold interquartile range (IQR). [Figure 20] Representative screenshots from the Integrated Genome Viewer of individual reads in single cells using mismatches induced by 4-thio-uridine labeling in cell culture are shown, as well as a reconstructed molecule from the mouse gene Psma2. [Figure 21]Adding dATP during second strand synthesis produces a suboptimal and unbalanced mix of dNTP concentrations, thereby favoring one conversion type over another (i.e., G to A conversion over A to G conversion). (A) The ratios of G to A conversions observed with "no dATP added" and "dATP added" repeat. (B) The ratios of A to G conversions observed with "no dATP added" and "dATP added" repeat. It will be understood that the conversion identities describe the original reference base in lower case and the new base in upper case. For example, a G to A conversion can be described as gA. EXAMPLES

[0126] Example 1 Materials and Methods Single human K562 cells were sorted into individual wells of a 384-well plate containing 3 μL Vapor-Lock (Qiagen) and 0.3 μL Smart-seq3 lysis buffer (see Hagemann-Jensen et al, 2020. Nature Biotechnology, 38:708-714) and either 0 or 0.5 mM dPTP was added. A 10-fold reduction in volume, a reduction in dNTP concentration to 0.1 mM, and MgCl adjusted to 1.5 mM were added. 2 Reverse transcription was performed as described in Hagemann-Jensen et al., 2020 (i.e., Smart-seq3 approach), except for the concentration. The final volume of reverse transcription was 0.4 μL. For dilution conditions, 4.6 μL of nuclease-free water was added to the reaction. For wells treated with alkaline phosphatase, 0.1 μL of FastAP (Thermo Scientific 0.2 U / μL) was added to the reaction, which was then incubated at 37°C for 20 min and 75°C for 10 min to inactivate the FastAP enzyme. For bead purification, the volume was adjusted to 10 μL and cleanup was performed with 8 μL of SPRI paramagnetic beads.

[0127] The purified cDNA was eluted in 5 μL. PCR master mix was then added to a final volume of 5 μL, 0.5 μL, 5 μL, and 0.5 μL for the bead cleanup, FastAP, dilution, and no cleanup conditions, respectively. PCR was performed as described in Hagemann-Jensen et al., 2020, except that for the different conditions, there were various amounts of salts and enzymes carried over from the reverse transcription and FastAP reactions. The resulting libraries were tagged and amplified using Illumina Nextera XT chemistry. The resulting libraries were circularized using the MGI App-A conversion kit and then sequenced using the StandardMPS PE100 kit on an MGI DNBSEQ-G400 platform.

[0128] Data were processed using zUMI (Parekh et al, 2018. Gigascience, 2018 Jun 1;7(6):giy059. doi:10.1093 / gigascience / giy059). The option find_pattern ATTGCGCAATG (SEQ ID NO:5) was specified to identify UMI-containing 5' reads, and all reads were mapped using STAR and aligned to the human genome (hg38). STAR settings were modified to allow up to 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0129] result Through the strategy described in this application, the use of error-prone reverse transcription results in the presence of unique patterns in the cDNA molecules produced (Figure 1). These patterns can be used to identify the molecule of origin in downstream applications such as molecular counting or molecular reconstruction (Figure 2). Example 1 shows that these patterns can be created by the use of dPTP during the reverse transcription reaction (Figure 3). Furthermore, this example confirms that these patterns uniquely identify the molecule of origin, as they correspond well with the unique molecular identifiers (UMIs) used in this experiment.

[0130] To avoid incorporation of base analogs during cDNA amplification, residual dPTP and free dPTP must be removed before PCR can be performed. This is important because incorporation of base analogs during PCR results in the production of base conversion patterns that do not uniquely correspond to individual RNA molecules and are therefore not useful for downstream analysis.

[0131] Importantly, dPTP-induced conversions can be easily detected since incorporation of dPTP during reverse transcription can only result in aG and gA conversions for features located on the plus strand, and cT and tC conversions for features located on the minus strand. In the absence of any cleanup strategy, both pairs of possible conversions were found for both plus and minus strand features (Figure 4A), indicating that dPTP was incorporated during PCR instead of reverse transcription. However, cleanup with SPRI paramagnetic beads and treatment with alkaline phosphatase (fastAP) reduce the conversion rate corresponding to incorporation that occurred during amplification of cDNA (i.e., tC and cT conversions in features located on the plus strand, and aG and gA conversions in features located on the minus strand). In addition, the base conversion patterns in samples where free dPTP was efficiently removed (either by fastAP or SPRI paramagnetic bead cleanup) are stable in contrast to samples where no cleanup was performed (Figure 4B).

[0132] Example 2 Materials and Methods Single human K562 cells were sorted into individual wells of a 384-well plate containing 0.3 μL of Smart-seq3 lysis buffer with dNTPs present at 0.1 mM, and various concentrations of dPTP were added. The concentrations of dPTP present in each reverse transcription reaction were 0 mM, 0.25 mM, 0.5 mM, and 1 mM. After reverse transcription (as in Example 1), FastAP (Thermo Scientific) was added to a final concentration of 0.1 U / μL in a total volume of 0.5 μL. The reaction was incubated at 37° C. for 20 minutes, and FastAP was inactivated at 72° C. for 10 minutes. PCR, tagging, and subsequent amplification were performed as described in Example 1 above. The resulting libraries were circularized using the MGI App-A conversion kit and sequenced using the StandardMPS SE100 kit on an MGI DNBSEQ-G400 platform.

[0133] Data were processed using zUMI (Parekh et al, 2018. Gigascience, 2018 Jun 1;7(6):giy059. doi:10.1093 / gigascience / giy059). The option find_pattern ATTGCGCAATG (SEQ ID NO:5) was specified to identify UMI-containing 5' reads, and all reads were mapped using STAR and aligned to the human genome (hg38). STAR settings were modified to allow up to 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0134] result The efficiency of introducing errors into cDNA is important to generate a pattern unique enough to allow identification of molecules derived from a particular original RNA molecule in the methods of the invention (Figure 5). This example shows that the reaction conditions, specifically the concentration of the base analog dPTP, can be adjusted to obtain a high percentage of base conversion events (Figure 6). Furthermore, this example shows that efficient reverse transcription of RNA from a single cell can be performed in the presence of these concentrations of dPTP.

[0135] Example 3 Materials and Methods 120,000 K562 cells were encapsulated and lysed in droplets according to the standard protocol of MGI C4 DNBelab. RNA capture and cleaning were performed according to the standard protocol. The reaction was then split into two and reverse transcription was performed in a 50 μL reaction using the RT primer mix from the MGI C4 DNBelab kit, with a concentration of 0.1 mM for each dNTP, according to the standard Smart-seq3 protocol (Hagemann-Jensen et al., 2020). For one of the two samples, 1 mM dPTP was added. Reverse transcription was performed according to the standard protocol. The resulting reaction was then cleaned up according to the standard MGI C4 DNBelab protocol. PCR amplification was performed using KAPA HiFi in the presence of 10 mM of each dNTP and a total of 4 μL of MGI C4 DNBelab cDNA amplification primer mix per sample. 200 ng of the resulting cDNA library was tagged in 1 / 5 volume using an Illumina Nextera XT. The resulting cDNA library of 200 pg can also be used. The resulting library was circularized using the MGI App-A conversion kit and sequenced on the MGI DNBSEQ-G400 platform using the SE100 kit.

[0136] The data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were modified to allow a maximum of 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0137] result Single-cell transcriptomics methods are broadly divided into plate-based and droplet-based methods. Plate-based methods rely on the separation of cells into separate wells of a multi-well plate, while droplet-based methods instead utilize lipid droplets in which cells are physically separated from each other. This example shows that implementing error-prone reverse transcription by incorporating dPTP into a droplet-based single-cell library preparation protocol (C4 DNBelab, MGI technologies) can result in a high percentage of base conversions (Figure 7). This demonstrates that in addition to plate-based methods (as shown in the examples above), droplet-based methods are also compatible with error-prone reverse transcription as described in this application.

[0138] Example 4 Materials and Methods The purified DNAse-treated RNA was reverse transcribed using modified Smart-seq3 reaction conditions (as in Example 1) in the presence of 2 mM 2-thio-dTTP (TriLink Biotechnologies N-2035). Alkaline phosphatase treatment of the reaction was performed using FastAP (Thermo Scientific) at a final concentration of 0.04 U / μL. The reaction was incubated at 37° C. for 20 minutes, and then the FastAP was inactivated at 75° C. for 10 minutes. PCR, tagging, and indexing PCR were then performed as described in Example 1 above. The resulting library was sequenced on an Illumina NextSeq500 platform using 75 cycles of High Output Kit v2.5.

[0139] The data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were modified to allow a maximum of 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0140] result This example demonstrates that incorporation of 2-thio-dTTP during reverse transcription can generate a high percentage of base conversion events (Figure 8A).

[0141] Example 5 Materials and Methods 4ng of purified DNAse-treated RNA was reverse transcribed using Maxima H-minus reverse transcriptase (5% polyethylene glycol 8000, 0.1% Triton X-100, 5U / μL recombinant RNAse inhibitor, 0.1mM each dNTP, 25mM Tris-HCL, 30mM NaCl, 1.5mM MgCl, 1mM GTP, 8mM DTT, 0.5uM Smart-seq2 oligo-dT, 2μM Smart-seq2 template switch oligo (see Picelli et al., 2013. Nature Methods, 10:1096-1098), Maxima H-minus reverse transcriptase 2U / μL). The names and product numbers of the different analogs tested in this experiment were 5-formyl-2'-deoxyuridine-5'-triphosphate (TriLink Biotechnologies N-2067), 5-propynyl-2'-deoxycytidine-5'-triphosphate (TriLink Biotechnologies N-2016), 5-iodo-2'-deoxycytidine-5'-triphosphate (TriLink Biotechnologies N-2023), and 5-propargylamino-2'-deoxyuridine-5'-triphosphate (TriLink Biotechnologies N-2062). Base analogs were present at a concentration of either 4 mM or 0.25 mM during reverse transcription. Base analogs were dephosphorylated by treatment with 0.12 U of FastAP (Thermo Scientific) for 20 min at 37°C, followed by inactivation of FastAP at 75°C for 10 min. PCR was performed according to the Smart-seq3 standard protocol, except that ISPCR primers were used instead of the standard Smart-seq3 forward and reverse primers (see Hagemann-Jensen et al., 2020). DNA libraries were tagged and indexed as described in Example 1 above. The resulting libraries were circularized using the MGI App-A conversion kit and sequenced using the StandardMPS PE200 kit on an MGI DNBSEQ-G400 platform.

[0142] The data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were modified to allow a maximum of 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0143] result This example shows that four additional base analogs can produce conversions with varying efficiency (Figure 8B). Although the error rates obtained with these base analogs individually are relatively low, utilizing combinations of base analogs can increase the effective overall conversion rate.

[0144] Example 6 Materials and Methods 20 ng of DNAse-treated RNA was reverse transcribed according to Smart-seq2 reaction conditions (Picelli et al., 2013), enriched with 0.1 mM of each dNTP, in the presence of dPTP (0.5 mM), 5-formyl-dUTP (0.25 mM), or in the absence of any base analogues. The resulting cDNA was purified with AMPure SPRI paramagnetic beads (1:1 bead-to-cDNA volume ratio) and eluted in a final volume of 120 μL. For each condition, 2 μL of purified cDNA was used for second strand synthesis with Klenow, T4, or water as negative controls. In addition to the enzyme or water negative controls, reactions consisted of 1× NEB buffer 2, 0.2 mM of each dNTP, and 0.2 μM of ISPCR primers. Reactions were incubated at 37 C for 2 h. Second strand products were then amplified using KAPA according to the Smart-seq2 protocol (Picelli et al., 2013) in the presence of 0.4 μM ISPCR primers and 1 mM of each dNTP for 24 cycles in a total reaction volume of 10 μL. The resulting libraries were tagged according to the Smart-seq3 protocol (Hagemann-Jensen et al., 2020), circularized using the MGI App-A conversion kit, and sequenced using the StandardMPS SE100 kit on an MGI DNBSEQ-G400 platform.

[0145] Data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were changed from the standard to allow up to 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0146] result Incorporation of a base analog into first strand cDNA (Figure 1) directly leads to incorporation of an incorrect canonical base during synthesis of second strand cDNA. This example shows that the method chosen for second strand cDNA synthesis affects the conversion rate achieved (Figure 9). The addition of a separate second strand synthesis step with Klenow DNA polymerase (but not T4 DNA polymerase) (prior to cDNA amplification with KAPA) increases the conversion rate for both dPTP and 5-formyl-dUTP.

[0147] Example 7 Materials and Methods 100 ng of DNAse-treated RNA was reverse transcribed using Maxima H-minus reverse transcriptase (5% polyethylene glycol 8000, 0.1% Triton X-100, 5 U / μL recombinant RNAse inhibitor, 0.1 mM each dNTP, 25 mM Tris-HCL, 30 mM NaCl, 1.5 mM MgCl, 1 mM GTP, 8 mM DTT, Smart-seq2 oligo-dT 0.5 μM, Smart-seq2 template switch oligo 2 μM (see Picelli et al, 2013. Nature Methods, 10:1096-1098), Maxima H-minus reverse transcriptase 2 U / μL) in the presence of 1 mM dPTP base analogue. cDNA was purified using SPRI beads at a ratio of 0.8:1. The cDNA was then amplified using the following PCR enzymes: KAPA HiFi HotStart PCR enzyme (KAPA BioSystems KK2501), Phusion HF HotStart II (Thermo Scientific F459), NEBNext (NEB M0541), Q5 DNA polymerase (NEB M0491), Q5 Ultra II (NEB M0543), Platinum Superfi II (Thermo Scientific 12361010), Platinum II (Thermo Scientific 14966005), Terra polymerase (Takara ST0287), VeriFi polymerase (PB10.45), Amplitaq Gold (8080240), and Taq DNA polymerase (Invitrogen 18038-042). All PCRs were performed according to the manufacturer's protocol using appropriate concentrations of ISPCR primers (Picelli et al., 2013). Then, all DNA libraries were purified using SPRI beads at a ratio of 0.8:1. The resulting DNA was tagged as described in Example 1 above. The resulting library was circularized using the MGI App-A conversion kit and sequenced using the StandardMPS SE100 kit on the MGI DNBSEQ-G400 platform.

[0148] The data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were modified to allow a maximum of 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0149] result Most widely used single-cell RNA-seq library preparation strategies do not perform a dedicated second strand synthesis. The second strand is instead synthesized in the first cycle of the cDNA amplification PCR, thereby effectively streamlining the protocol and increasing sensitivity. Given the importance of second strand synthesis (Figure 10), the choice of PCR enzyme is potentially very important. This example shows that the choice of PCR enzyme is a key factor that affects the rate at which errors can be induced when amplifying cDNA containing dPTP. Furthermore, based on the results discussed in Example 6 above (and shown in Figure 9), this example suggests that cDNA amplification strategies can be tailored to the specific base analogs used.

[0150] Example 8 Materials and Methods 1.1ug of purified DNAse treated RNA was reverse transcribed using Superscript II (Thermofisher) in the presence of various percentages of CTP in the dNTP mix replaced with 5'-methyl-CTP according to the manufacturer's protocol. The percentages of 5'-methyl-CTP used were 0%, 20%, 50%, 80%, and 100%, respectively. The resulting cDNA was converted to bisulfite using the EZ DNA Methylation-Gold Kit (Zymo Research) according to the manufacturer's protocol. Second strand synthesis was performed using Klenow (NEB) according to the manufacturer's protocol using random hexamer primers. The second strand synthesis reaction was terminated by adding EDTA to a final concentration of 10mM, and the resulting double stranded DNA was purified using SPRI beads (1:1 ratio). The resulting DNA library was quantified and tagmentation was performed using an Illumina Nextera XT in 1 / 5 of the total volume using the manufacturer's protocol. The resulting library was circularized using the MGI App-A conversion kit and sequenced using the StandardMPS SE100 kit on an MGI DNBSEQ-G400 platform.

[0151] The data were processed using STAR to perform an alignment with the human genome (hg38) using zUMI (https: / / github.com / sdparekh / zUMIs). STAR settings were modified to allow up to 40% mismatches. Note that mismatches here correspond to all possible mismatches.

[0152] result Bisulfite conversion of unmethylated cytosines in DNA results in cT conversion. However, 5' methylated cytosines are protected against bisulfite conversion. This example shows that incorporating various percentages of 5'-methyl-dCTP into cDNA by reverse transcription and performing bisulfite conversion results in gA conversion and no cT conversion of features located on the plus strand in subsequent sequencing library preparation (Figure 11). Since cDNA is the reverse complement of the original RNA molecule, gA conversion is expected for plus strand features and cT for minus strand features. This shows that the strategy described in this example efficiently produces cDNA molecules with patterns of errors that can be used, for example, for molecule counting, strandedness identification, and RNA molecule sequence reconstruction.

[0153] Example 9 Materials and Methods Single K562 cells were sorted into individual wells of a 384-well plate containing 0.3 μL of Smart-seq3 lysis buffer (see Hagemann-Jensen et al., 2020) with dPTP present at 0.5 mM and each dNTP present at 0.1 mM. Reverse transcription was performed in a 10-fold reduction in volume and with no MgCl 2 The concentration was adjusted to 1.5 mM and run according to the Smart-seq3 protocol (see Hagemann-Jensen et al., 2020). FastAP was added to a final concentration of 0.1 U / μL and a total volume of 0.5 μL. The reaction was incubated at 37°C for 20 min and FastAP was inactivated at 72°C for 10 min. cDNA was amplified as described in Example 1 above. The resulting cDNA library was quadruple tagged as described in Example 1 above to maximize fragment complexity. The resulting library was circularized using the MGI App-A conversion kit and sequenced on an MGI DNBSEQ-G400 platform using the StandardMPS PE200 kit.

[0154] Data were processed using zUMI (https: / / github.com / sdparekh / zUMIs). The option find_pattern ATTGCGCAATG (SEQ ID NO:5) was specified to identify UMI-containing 5' reads, and all reads were mapped using STAR and aligned to the human genome (hg38). STAR settings were modified to allow up to 20% mismatches. Note that mismatches here correspond to all possible mismatches.

[0155] result Smart-seq3 data are usually composed of "UMI reads" and "internal reads". UMI reads contain UMIs and can be linked to individual RNA molecules, and those reads typically correspond to the 5' end of the molecule. Using the patterns introduced during reverse transcription by the method of the present invention, the "internal reads" can be efficiently assigned to the molecule of origin (Figure 12A). The length of the reconstructed molecule is comparable to that obtained from long-read sequencing of the full-length cDNA (Figure 12B). As already shown in the previous examples, the base conversion pattern is unique to the origin strand of the RNA molecule (Figure 13A). Thus, in addition to the reconstruction, the induced base conversion pattern can be used to easily identify the strand from which the corresponding RNA was transcribed (Figure 13B).

[0156] Example 10 Materials and Methods Single K562 cells were sorted into 96-well plates with 0.2 μL of lysis buffer containing 1 mM dATP, 0.2 mM dCTP, 1 mM dGTP, 1 mM dTTP, 10 mM dPTP, 0.08% Triton-X100 (Sigma), 1.6 U / μL recombinant RNAse inhibitor (Takara), cell barcoding and UMI-containing oligo-dT primer (e.g.: TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGAAGTCTGTACTATGGNNNNNNNNTTTTTTTTTTTTTTTTTTTTTTTTTT (SEQ ID NO: 1), 2 μM) and 5 μL of Vapor-Lock (Qiagen). Cells were lysed at 72 °C for 10 min. 0.2 μL of RT reaction mixture (10 mM DTT, 2 M betaine, 12 mM MgCl 2, 0.8U / μL recombinant RNAse inhibitor (Takara), 2× Superscript II RT buffer, and 20U / μL Superscript II enzyme) were added. Reverse transcription was performed at 42° C. for 90 min, followed by 10 cycles of 2 min 50° C. and 2 min 42° C., with a final single hold at 85° C. for 5 min and then a hold at 4° C. RT reactions were pooled and purified using a Zymo Research Clean & Concentrator DNA purification column using 5 volumes of DNA binding buffer, washed twice with DNA wash buffer, and eluted in 20 μL. First strand cDNA was polyadenylated using terminal deoxynucleotidyl transferase (TDT) in a 25 μL reaction containing 0.75 U / μL TDT enzyme (Sigma, 20 U / μL), 1.5 mM dATP, 0.55× ThermoPol buffer (NEB) and 0.02 U / μL RNAse H (Invitrogen, 2 U / uL). The TDT reaction was incubated at 37° C. for 1 min 15 sec and 65° C. for 10 min, followed by a hold at 4° C. 30 μL of second strand synthesis mix (27.5 μL of 2× Terra PCR Direct buffer, 1.76 μL of primer (TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGTTTTTTTTTTTTTTTTT TTTTTTT (SEQ ID NO: 2), 1 μM) and 0.55 μL of Terra PCR Direct polymerase mix (1.25 U / μL, Takara) and 0.19 μL of nuclease-free water) were added to the TDT reaction. The resulting reaction was held at 98° C. for 2 min, then 40° C. for 1 min, then ramped to 68° C. at 0.2° C. / sec and held at this temperature for 6 min. Cleanup was performed using a Zymo Research Clean & Concentrator DNA purification column as above and DNA was eluted in a volume of 20 μL. cDNA amplification was performed in a 50 μL reaction (1× Terra PCR Direct buffer, 0.8 μM amplification primer (TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO: 3)), 0.025 U / μL Terra Direct polymerase mix).PCR was performed by denaturing at 98°C for 2 min, followed by 18 cycles of denaturing at 98°C for 10 s, annealing at 65°C for 15 s, and extending at 68°C for 6 min. After 18 cycles, a 5 min hold at 68°C was followed by a 4°C hold. Amplified cDNA was purified using SPRI beads and tagged as in Example 8. The resulting library was circularized using the MGI App-A conversion kit according to the manufacturer's instructions and sequenced using the StandardMPS PE200 kit on an MGI DNBSEQ-G400RS platform.

[0157] Reads were separated into 3' cell barcoded reads (>16 for reads 1 base 1-24), 5' anchored reads (>16 for reads 1 base 25-48), and internal reads (neither). Each group was processed separately with zUMI v2.9.7 (https: / / github.com / sdparekh / zUMIs) and mapped to hg38 with STAR settings "-outFilterMismatchNmax 80 --outFilterMismatchNoverLmax 0.4 --outSAMattributes MD NH HI AS nM --clip3pAdapterSeq AAAAAAAAAAAAA" (SEQ ID NO: 4) to allow for a large number of mismatches. The resulting bam files were then merged into one bam file. The reads were then used for molecular reconstruction. For each gene, each read was sorted according to the start and end positions of the plus and minus stranded gene, respectively. First, we grouped the barcoded cell reads according to adjusted mutual information, taking into account eligible base overlaps (G in reference) and overlap conversions (G>A). If the base call quality of a read at a given position was below a Phred score of 15, that position was not considered for adjusted mutual information calculation. If the adjusted mutual information of a unique group was above 0.2, the read was added to an existing group. If there were no groups above 0.15, the read formed a new group. If there were multiple matches above 0.2, the read was discarded. The conversion pattern of a molecular group was determined by requiring at least 20% of the reads with a Phred score above 14 to have a conversion at that position.

[0158] Non-barcoded reads were used if all barcoded cell reads were grouped according to the transformation pattern. Each non-barcoded read was compared to the molecular pattern across all cells in the sample. If a read had an adjusted mutual information of >0.3 for one unique molecular group, the read and its corresponding transformation pattern were added to that molecular group. If there was no match or multiple matches (>0.3 adj. mutual information), the read was discarded. This process was repeated twice.

[0159] Once non-barcoded reads were assigned to molecules, all reads were written to a new bam file with their new molecular group as a tag. If a read was not barcoded, an inferred cell of origin was also added. The reads were then merged into one reconstructed molecular read using stitcher.py (https: / / github.com / AntonJMLarsson / stitcher.py).

[0160] result The addition of dPTP to a library preparation strategy that relies on 5'A tailing instead of template switching produces 3' reads with cellular barcodes and UMIs, as well as reads without barcodes or UMIs, all of which are derived from RNA molecules reverse transcribed with primers that introduced both cellular barcodes and UMIs (see FIG. 16). This type of approach results in very high conversion rates (i.e., >20% for the desired G>A conversion) (see FIG. 17), which allows for efficient RNA molecule reconstruction. As shown in FIGS. 18 and 19, even reads without cellular barcodes can be effectively linked to the cellular barcodes added during reverse transcription of the original RNA molecules via molecule-specific base conversion patterns.

[0161] Example 11 Materials and Methods Primary mouse fibroblasts were cultured for 2 h in the presence of 4-thio-uridine (Sigma, 200 μM) and single cells were sorted into 0.3 μL of lysis buffer (2.5 U / μL recombinant RNAse inhibitor (Takara), 0.2% Triton-X100) in 3 μL of Vapor-Lock (Qiagen). 0.3 μL of acylation reaction mixture was added (final reaction concentrations: 50 mM Tris-HCL (pH 8), 45% DMSO, 10 mM iodoacetamide) and the reaction was incubated at 50 °C for 10 min. 0.4 μL of quench mixture was added (final concentrations: 35 mM DTT, 2 mM dNTPs, 2.4 μM Smart-seq3 oligo-dT (Hagemann-Jensen et al., 2020), and 1.6 U / μL recombinant RNAse inhibitor (Takara)). The samples were then incubated at 72° C. for 10 min. 3 μL of reverse transcription mixture (33.3 mM Tris-HCL (pH 8), 46.7 mM NaCl, 1.3 mM GTP, 3.3 mM MgCl 2 , 6.7% PEG (MW8000), 2.7 mM DTT, 0.5 U / μL recombinant RNAse inhibitor (Takara), 2.7 μM Smart-seq3 template switching oligo (Hagemann-Jensen et al., 2020), 2.7 U / μL Maxima H-minus RT enzyme) were added. Reverse transcription and the remaining library preparation were performed as described in Hagemann-Jensen et al., 2020. Library circularization and sequencing were performed as in Example 10.

[0162] Reads were processed using zUMI (https: / / github.com / sdparekh / zUMIs) with option find_pattern ATTGCGCAATG (SEQ ID NO: 5) specified to identify 5' reads containing UMIs and mapped to mm10 with STAR settings "--outFilterMismatchNmax 40 --outFilterMismatchNoverLmax 0.25 --outSAMattributes MD NH HI AS nM XS --outSAMstrandField intronMotif --clip3pAdapterSeq CTGTCTCTTATACACATCT" (SEQ ID NO: 6).

[0163] The reads were then used for molecular reconstruction. For each gene, each read was sorted according to the start and end positions of the plus and minus stranded gene, respectively. First, the barcoded cell reads were grouped according to the adjusted mutual information, taking into account eligible base overlaps (T in the reference) and overlap conversions (T>C). If the base call quality of a read at a given position was below a Phred score of 15, that position was not considered for the adjusted mutual information calculation. If the adjusted mutual information of a unique group was above 0.2, the read was added to the existing group. If there were no groups above 0.15, the read was used to form a new group. If there were multiple matches above 0.2, the read was discarded. The conversion pattern of a molecular group was determined by requiring at least 20% of the reads with a Phred score above 14 to have a conversion at that position. All reads were written to a new bam file with that new molecular group as a tag. If a read was not barcoded, the inferred cell of origin was also added. The reads were then merged into one reconstructed molecular read using stitcher.py (https: / / github.com / AntonJMLarsson / stitcher.py).

[0164] result Newly produced RNA molecules in single mouse fibroblasts were labeled with 4-thiouridine U and read out as base conversions corresponding to the RNA molecule using an updated version of NASC-seq (see Materials and Methods in Hendriks et al. 2019. Nat. Commun., 10(1):3138). The results of this example demonstrate that the base conversion patterns introduced using this method can be used to effectively reconstruct the RNA molecule sequence (Figure 20). This approach shows that by labeling newly produced RNA in cells with 4-thio-uridine, followed by treatment with iodoacetamide and preparation of a sequencing library, a molecular specific pattern was created that can be used to reconstruct the sequence of the original RNA molecule present.

[0165] Example 12 Materials and Methods Single HEK293T cells were sorted into 96-well plates and lysed and reverse-transcribed as described in Example 10. The pooled and purified first strand cDNA was then polyadenylated and cleaned up again using Zymo Research Clean & Concentrator columns before being split into four reactions. Second strand synthesis was then performed using Terra PCR Direct Polymerase Buffer and PCR Direct Polymerase Mix with 0.03 μM primer (TCGTCGGCAGCGTCAGATGTGTATAAG AGACAGTTTTTTTTTTTTTTTTTTTTTTTT) (SEQ ID NO: 2). The concentration of dATP in two reactions was then increased by 1 mM by adding extra dATP. The four reactions were then cleaned up using Zymo Research Clean & Concentrator columns. The remainder of the library preparation process was then performed as in Example 10. Circularization of the library was performed as in Example 10 and sequencing was carried out on a DNBSEQ-G400RS using StandardMPS PE150 chemistry.

[0166] The resulting data was processed as in Example 10, without any reconstruction. Error rates were calculated directly from the output bam file of zUMI. Cells with less than 400,000 bases covered by the sequencing reads were removed from the analysis.

[0167] result Adding extra dATP during second strand synthesis, thereby creating a suboptimal balance of dNTP concentrations, results in favor of G to A conversion instead of A to G conversion, as can be seen in Figure 21. Figure 21 shows significant differences (two-tailed t-test) in conversion rates between both replicates of the two condition groups in response to including additional dATP during second strand synthesis.

Claims

**Claim 1** A method for determining the copy number of one or more RNA molecules within a population of RNA molecules, comprising: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to generate a population of DNA molecules, wherein each DNA molecule contains one or more base conversions relative to the corresponding RNA molecule and each DNA molecule contains a molecule-specific base conversion pattern; determining the copy number of the one or more RNA molecules within the population using the molecule-specific base conversion pattern. **Claim 2** The method according to claim 1, further comprising, after step (ii): (iii) determining the sequences of overlapping fragments of the DNA molecules within the population; (iv) determining the partial or full-length sequences of the DNA molecules within the population from the information of step (iii) by assembling the sequences of the overlapping fragments based on the molecule-specific base conversion patterns in the DNA molecules; (v) determining the sequences of the RNA molecules corresponding to the DNA molecules from the information of step (iv); (vi) determining the copy number of one or more RNA molecules within the population from the information of step (v). **Claim 3** A method for determining the sequence of one or more RNA molecules within a population of RNA molecules, comprising: (i) providing a population of RNA molecules; (ii) subjecting the population of RNA molecules to error-prone reverse transcription to generate a population of DNA molecules, wherein each DNA molecule contains one or more base conversions relative to the corresponding RNA molecule and each DNA molecule contains a molecule-specific base conversion pattern; determining the sequence of the RNA molecules corresponding to the one or more DNA molecules using the molecule-specific base conversion pattern. **Claim 4** The method according to claim 3, further comprising, after step (ii): (iii) determining the sequences of overlapping fragments of the DNA molecules within the population; (iv) determining the sequences of one or more DNA molecules within the population from the information of step (iii) by assembling the sequences of the overlapping fragments based on the molecule-specific base conversion patterns of the DNA molecules; (v) determining the sequences of the RNA molecules corresponding to the one or more DNA molecules from the information of step (iv). **Claim 5** The method according to claim 1 or 3, wherein the RNA molecule population comprises RNA molecules having different sequences and / or RNA molecules having the same sequence.

6. The method according to claim 1 or 3, wherein the RNA molecule population to be analyzed comprises from 1 to 100,000,000,000 individual RNA molecules, preferably from 100 to 1,000,000,000,000 individual RNA molecules, more preferably from 1,000 to 1,000,000,000 individual RNA molecules, and most preferably from 100,000 to 100,000,000 individual RNA molecules.

7. The method according to claim 1 or 3, wherein the RNA molecule population comprises one or more RNA molecules selected from the group consisting of messenger RNA (mRNA), precursor mRNA (pre-mRNA), antisense RNA (asRNA) and its precursor, enhancer RNA and its precursor, long non-coding RNA (lncRNA) and its precursor, microRNA (miRNA) and its precursor, ribosomal RNA (rRNA) and its precursor, transfer RNA (tRNA) and its precursor, histone RNA and its precursor, small nucleolar RNA (snoRNA) and its precursor, small nuclear RNA (snRNA) and its precursor, mitochondrial RNA and its precursor, viral RNA, transposon RNA, synthetic RNA, in vitro transcribed RNA, or a combination thereof.

8. The method according to claim 1 or 3, wherein step (ii) comprises introducing one or more base conversions into each DNA molecule at a total ratio of about 0.5% to about 99.5%, more preferably about 2% to about 98%, even more preferably about 5% to about 95%, even more preferably about 5% to about 50%, even more preferably about 5% to about 20%, and most preferably about 15% to about 30%.

9. The method according to claim 1 or 3, wherein step (ii) comprises reverse transcription in the presence of one or more base analogs.

10. The method according to claim 9, wherein the one or more base analogs are selected from the group consisting of 2'-deoxy-P-nucleoside-5'-triphosphate (dPTP), 8-oxo-2'-deoxyguanosine-5'-triphosphate (8-oxo-GTP), 2-thiothymidine-5'-triphosphate (2-thio-TTP), 5-formyl-2'-deoxyuridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5-propynylamino-2'-deoxyuridine-5'-triphosphate, or combinations thereof.

11. The method according to claim 1 or 3, wherein step (ii) comprises reverse transcription in the presence of a sub-optimal amount of one or more dNTP bases.

12. The method according to claim 1 or 3, wherein the method comprises incorporating one or more base analogs into the one or more RNA molecules within the RNA molecule population prior to step (i).

13. The method according to claim 12, wherein the one or more base analogs are 4-thio-uridine.

14. The method according to claim 1 or 3, wherein step (ii) further comprises chemically modifying the RNA molecule population prior to subjecting the RNA molecule population to reverse transcription.

15. The method according to claim 14, wherein the step of chemically modifying the RNA molecule population comprises alkylating the RNA molecule population, and optionally, the alkylation is performed by iodoacetamide treatment or oxidative aromatic nucleophilic substitution.

16. The method according to claim 1 or 3, wherein step (ii) further comprises chemically modifying the DNA molecule population generated by reverse transcription.

17. The method according to claim 14, wherein the chemical modification comprises a deamination reaction.

18. The method according to claim 16, wherein the chemical modification comprises a deamination reaction.

19. The method according to claim 1 or 3, wherein step (ii) comprises reverse transcription using an error-prone reverse transcriptase.

20. The method according to claim 2 or 4, wherein step (iii) comprises amplifying the DNA molecule population from step (ii) to generate one or more amplicons of each DNA molecule within the population.

21. The method according to claim 20, wherein the step of amplifying the DNA molecule population comprises high-fidelity amplification.

22. The method according to claim 20, wherein the step of amplifying the DNA molecule population comprises PCR amplification.

23. The method according to claim 20, wherein the step of amplifying the DNA molecule population is carried out in the absence of base analogs.

24. The method according to claim 20, wherein at least a first cycle of the step of amplifying the DNA molecule population is carried out in the presence of a sub-optimal amount of one or more dNTP bases.

25. The method according to claim 2 or 4, wherein step (iii) comprises fragmenting the one or more amplicons of the DNA molecule population and / or each DNA molecule within the population to generate overlapping fragments.

26. The method according to claim 25, wherein the step of fragmenting the one or more amplicons of the DNA molecule population and / or each DNA molecule within the population comprises tagging, DNA shearing, and / or enzymatic fragmentation.

27. The method according to claim 25, wherein the fragments are about 50 base pairs to about 1500 base pairs in length.

28. The method according to claim 2 or 4, wherein step (iii) comprises sequencing the overlapping fragments of the one or more amplicons of the DNA molecule population and / or each DNA molecule within the population.

29. The method according to claim 28, wherein sequencing comprises a short-read sequencing method.

30. Step (iv) comprises (a) assigning overlapping fragments to RNA molecules present in the RNA molecule population based on alignment of some or all of their sequences to those RNA molecules, and / or (b) selecting the assigned fragments based on the positions in the RNA molecules where those fragments align, the method according to claim 2 or 4.

31. The method according to claim 2 or 4, wherein step (v) comprises comparing the sequence information of step (iv) with a reference sequence to identify mismatches corresponding to one or more base conversions.

32. The method for determining the copy number of one or more RNA molecules according to claim 2, wherein step (vi) comprises identifying the number of unique molecule-specific base conversion patterns corresponding to RNA molecules having a specific sequence within the RNA molecule population from the information of step (v).

33. The method according to claim 2 or 4, wherein one or more of steps (i) to (iii) are carried out in a droplet-based environment, a plate-based environment, an environment attached to beads, or in situ.

34. The method according to claim 1 or 3, wherein the RNA molecule population comprises one or more sequence variants of the same gene, or one or more allelic variants of the same gene, or one or more splice variants of the same gene, or one or more RNA isoforms resulting from alternative use of a promoter, or one or more RNA isoforms resulting from alternative use of a splice site, or one or more RNA isoforms resulting from alternative use of a polyadenylation site.

35. A method for generating base conversions of one or more polynucleotide molecules within a polynucleotide molecule population, comprising: (i) providing a polynucleotide molecule population, wherein one or more of the polynucleotide molecules comprise one or more base analogs; (ii) amplifying the polynucleotide molecule population from step (i) to generate one or more amplicons of each polynucleotide molecule within the population, wherein the amplifying is carried out in the presence of a sub-optimal amount of one or more dNTP bases.

36. Use of an error-prone reverse transcriptase to generate a population of DNA molecules from an RNA molecule population, wherein each DNA molecule comprises one or more base conversions relative to the corresponding RNA molecule and has a molecule-specific base conversion pattern, for determining the copy number of one or more RNA molecules within the population.

37. Use of an error-prone reverse transcriptase to generate a population of DNA molecules from an RNA molecule population, wherein each DNA molecule comprises one or more base conversions relative to the corresponding RNA molecule and has a molecule-specific base conversion pattern, for determining the sequence of one or more RNA molecules within the population.

38. A population of DNA molecules obtainable or obtainable by the method according to any one of claims 1, 3 and 35, the use according to claim 36, or the use according to claim 37.

39. A kit for performing error-prone reverse transcription, the kit comprising: (i) a reverse transcriptase; (ii) one or more base analogs; (iii) instructions for use.

40. The kit according to claim 39, wherein the one or more base analogs are selected from the group consisting of 2'-deoxy-P-nucleoside-5'-triphosphate (dPTP), 8-oxo-2'-deoxyguanosine-5'-triphosphate (8-oxo-GTP), 2-thiothymidine-5'-triphosphate (2-thio TTP), 5-formyl-2'-deoxyuridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5-propynylamino-2'-deoxyuridine-5'-triphosphate, or combinations thereof.

41. The kit according to claim 39 or 40, wherein the reverse transcriptase is an error-prone reverse transcriptase.