Identifying the origin of amplified RNA molecules
By introducing molecular landmarks during reverse transcription or polymerization, the method addresses amplification-induced variability in RNA molecule counts, ensuring accurate gene expression analysis and computational manageability.
Patent Information
- Application Number
- PCT/US2025/032691
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Existing methods for analyzing gene expression in biological samples, particularly through single-cell RNA sequencing, suffer from inaccuracies due to amplification processes that introduce variability in RNA molecule counts, leading to noisy quantitative measurements and computational challenges with vast amounts of data.
Introduce molecular landmarks, such as mutations, during reverse transcription or polymerization using low bias DNA polymerases or nucleotide analogs, to uniquely identify amplified RNA molecules, allowing accurate association of amplicons with their originating mRNA molecules, thereby reducing data volume and improving computational manageability.
Enables accurate adjustment of read counts to reflect actual mRNA representation, reducing data complexity and enhancing computational efficiency in analyzing gene expression data.
Smart Images

Figure US2025032691_11122025_PF_FP_ABST
Abstract
Description
IDENTIFYING THE ORIGIN OF AMPLIFIED RNA MOLECULESFIELD OF THE TECHNOLOGY DISCLOSED
[0001] The technology disclosed relates to techniques suitable for identifying the origin of amplified analytes in a biological sample, such as for identifying the origin of amplified RNA molecules. Further, the disclosed technology describes techniques suitable for stitching sequences reads to form a longer read, such as stitching two or more short sequence reads to form a longer sequence.BACKGROUND
[0002] The subj ect matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, a problem mentioned in this section or associated with the subject matter provided as background should not be assumed to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which in and of themselves can also correspond to implementations of the claimed technology.
[0003] Aspects of the present disclosure relate generally to devices, systems, and methods for use in biological or chemical analysis. Various protocols in biological or chemical research involve performing a large number of controlled reactions on local support surfaces or within predefined reaction chambers. The designated reactions may then be observed or detected, and subsequent analysis may help identify or reveal properties of chemicals involved in the reaction. For example, in some multiplex assays, an unknown analyte having an identifiable label (e.g., fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited into a corresponding well of a flow cell channel. Observing any chemical reactions that occur between the known probes and the unknown analyte within the wells may help identify or reveal properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis (SBS) or cyclic-arraysequencing.
[0004] In various contexts, as part of a biochemical analysis of a tissue or single-cells within a sample, it may be desirable to obtain count information representative of mRNA molecules within the tissue or cells. By way of example, in assessing gene expression within a tissue or cell, it may be desirable to evaluate what mRNAs are present in the tissue or cell in general, at a given time, or under particular conditions. However, the amplification processes typically employed as part of such analyses may introduce variation in the number of copies made of respective RNA molecules such that the observed number or counts do not accurately reflect the actual proportions between different mRNA molecules present in the tissue or cells. That is, the amplification process may inadvertently make some RNAs appear to be more or less represented than they are in actuality. With this in mind, it is desirable to be able to accurately relate amplicons generated in such processes to the underlying mRNA with which they are associated.
[0005] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventor(s) has made or used the devices and techniques described herein.INCORPORATION BY REFERENCE
[0006] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure. The following references are incorporated by reference in their entirety and for all purposes, such as, but not limited to, subject matter pertaining to stitching sequence reads together to form longer and / or contiguous reads:(1) U.S. Patent No. US 11,008,606 titled “Random nucleotide mutation for nucleotide template counting and assembly”, filed October 9, 2015 to Wigler et al.(2) U.S. Patent Publication No. US2021 / 0174905 titled “Sequencing Algorithm”, filed lune 10, 2021 to Imelfort et al.(3) U.S. Patent Publication No. US2023 / 0044570 titled “Method for Determining a Measure Correlated to the Probability That Two Mutated Sequence Reads Derive from the Same Sequence Comprising Mutations”, filed February 9, 2023 to Darling.(4) U.S. Patent Publication No. US 2021 / 0403991 titled “Sequencing Process”, filed December 30, 2021 to Burke et al.(5) U.S. Patent Publication No. US 2022 / 0348940 titled “Method for Introducing Mutations”, filed November 3, 2022 to Monahan et al.(6) WO 2014 / 013218 titled “Methods and systems for determining haplotypes and phasing of haplotypes”, filed May 20, 2013 to Rigatti et al.(7) WO 2015 / 177570 titled “Sequencing Process”, filed May 22, 2015 to Burke et al.(8) WO 2019 / 162657 titled “Method for Introducing Mutations”, filed February 19, 2019 to Monahan et al.(9) WO 2023 / 230552 titled “Preparation of long read nucleic acid libraries”, filed May 25, 2023 to Meinholz et al.(10) WO 2023 / 230550 titled “Preparation of long read nucleic acid libraries”, filed May 25, 2023 to Meinholz et al.(11) WO 2016 / 057947 titled “Random nucleotide mutation for nucleotide template counting and assembly”, filed October 9, 2015 to Wigler et al.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings,wherein:
[0008] FIG. 1 depicts an example process flow for measuring gene expression, in accordance with aspects of the presently described techniques;
[0009] FIG. 2 depicts a further example process flow for measuring gene expression, in accordance with aspects of the presently described techniques;
[0010] FIG. 3 depicts an additional example process flow for measuring gene expression, in accordance with aspects of the presently described techniques;
[0011] FIG. 4 depicts a table illustrating aspects of a mutagenic nucleotide embodiment, in accordance with aspects of the presently described techniques;
[0012] FIG. 5 depicts a table illustrating aspects of secondary single-cell metrics in an example embodiment, in accordance with aspects of the presently described techniques;
[0013] FIG. 6 graphically illustrates results related to increased load of specific transition mutations for experimental results, in accordance with aspects of the presently described techniques;
[0014] FIG. 7 depicts observed mutation rate differences at different mutagenic nucleotide concentrations, in accordance with aspects of the presently described techniques;
[0015] FIG. 8 depicts additional observed mutation rate differences at different mutagenic nucleotide concentrations, in accordance with aspects of the presently described techniques; and
[0016] FIG. 9 graphically illustrates trend in mutation rate per cycle, in accordance with aspects of the presently described techniques.DET AILED DESCRIPTION
[0017] The following discussion is presented to enable any person skilled in the art to make and use the technology disclosed and is provided in the context of a particular application and its requirements. Various modifications to the disclosed implementations will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to otherimplementations and applications without departing from the spirit and scope of the technology disclosed. Thus, the technology disclosed is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0018] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components. Thus, for example, one or more of the functional blocks (e g., processors or memories) may be implemented in a single piece of hardware (e.g., a general purpose signal processor or random access memory, hard disk, or the like). Similarly, the programs may be stand-alone programs, may be incorporated as subroutines in an operating system, may be functions in an installed software package, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.
[0019] The presently disclosed techniques relate to the identification of the origin of amplified RNA molecules. As discussed herein, mutations or other molecular landmarks may be introduced to a reverse transcription product (e.g., a first strand cDNA) and / or a complement of such a first strand cDNA (e.g., a second strand cDNA). The molecular landmarks in turn may be used to associate a given first strand cDNA and / or second strand cDNA to an mRNA molecule from which the respective first strand cDNA or second strand cDNA are derived. In this manner, amplicons of these cDNA molecules can be related back to the initial mRNA molecule from which the cDNA strands are derived, allowing read counts of such strands to adjusted to accurately correspond to the mRNA representation within the cells or tissue being analyzed. Such accurate count data may in turn be useful in evaluating or determining gene expression within the cells or tissue in question.
[0020] In this manner problems associated with conventional approaches may be addressed. For example, many existing methodologies employ read counts of cDNA molecules generated based off of mRNA to assess gene expression activity. In addition to the variability introduced by the amplification process, the extent of data that may be generated may exceed the practical processing or handling limits of such approaches. For example, as single-cell type data has become available, large amounts of information may be available that allows gene expression to beevaluated down to the level of cellular granularity. However, such approaches may generate a vast amount of data to be analyzed (e.g., amplified expression data for each individual cell in a sample) that exceeds or approaches the limits of computational bandwidth typically available in a laboratory or clinical context. By way of example, in a single-cell context using an amplificationbased approach there may be approximately 25,000, 50,000, or 100,000 reads per cell, with 100’s, 1,000’s, 10,000’s or more cells potentially being analyzed. Such read counts exceed what can efficiently be performed using conventional computational hardware and it is therefore necessary to reduce the effective read counts by relating the counts to the underlying molecule (e g., mRNA) that was amplified, which may effectively reduce the counts to 50 to 100 counts per cell, which is computationally manageable.
[0021] Further, the amplification process itself may introduce variability to the molecule count data being generated. In particular, due to the inherent low sensitivity and yield associated with generating single-cell RNA sequencing read counts, such counts do not correspond well to the number of RNA molecules in a given cell. This lack of correspondence arises due to single-cell RNA sequencing typically involving multiple rounds of amplification to compensate for the low sensitivity and yield. As a result, the number of reads in a gene in each cell is affected by random differences in the amplification of the different molecules. Because each molecule may end up being amplified multiple times, the eventual readout will have a varying number of reads per molecule, resulting in quantitative measurements that are noisy.
[0022] As discussed herein, certain approaches may be employed that allow for the identification of unique molecules using molecular landmarks (which may also be referred to herein as “landmark mutations) introduced prior to or during an amplification step. By way of example, in certain embodiments an error prone reverse transcription step may be employed prior to amplification, with the errors introduced during reverse transcription being reproduced during amplification and serving as an indicator (i.e., a molecular landmark) that is indicative of molecule origin. As discussed herein, by creating such identifiable mutations at an early stage (e.g., reverse transcription) all sequencing reads coming from the same individual RNA molecule will share the mutation(s) in question, thus allowing the amplified RNA molecules to be recognized and related back based on their molecular origin without the use of an exogenous tag, such as a unique molecular identifier (UMI). As discussed herein, the processes and examples described may beutilized with (i.e., are compatible with) any suitable RNA library preparation, single cell or otherwise). Further, the processes described may be combined with other types of tagging (e.g., UMIs) for a more robust readout.
[0023] As discussed herein, various methodologies may be employed in the generation of molecular landmarks in either a first strand cDNA (by use of a mutation inducing transcription process, i.e., a mutagenesis transcription), in a complementary second strand cDNA (by use of a mutation inducing polymerization process, i.e., a mutagenesis polymerization), or both. By way of example, in certain aspects a low bias reverse transcriptase or a low bias polymerase (e.g., DNA polymerase) may be employed to polymerize (e.g., copy or transcribe) respective target nucleic acid molecules while introducing random mutation. In such contexts, a “low bias” reverse transcriptase or DNA polymerase is, respectively, a reverse transcriptase or DNA polymerase that (a) exhibits low mutation bias, and / or (b) exhibits low template amplification bias.
[0024] For simplicity and to reduce the introduction of duplicative concepts, the concept of a low bias DNA polymerase is described, however it will be understood that low bias reverse transcriptase exhibits similar low mutation bias introduction but in the context of introducing mutation into a transcribed DNA strand from a template RNA.
[0025] With this in mind, a low bias DNA polymerase, as used herein, exhibits a low mutation bias in that it mutates adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at similar rates. In an embodiment, the low bias DNA polymerase is able to mutate adenine, thymine, guanine, and cytosine at similar rates.
[0026] Optionally, the low bias DNA polymerase is able to mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7- 1.3, 0.8-1.2:0.8-1.2, or around 1 : 1 respectively. In certain embodiments, the low bias DNA polymerase is able to mutate guanine and adenine at a rate ratio of 0.5- 1.5:0.5-1.5, 0.6- 1.4:0.6- 1.4, 0.7-1.3:0.7- 1.3, 0.8-1.2:0.8- 1.2, or around 1 : 1 respectively. Further, the low bias DNA polymerase may be able to mutate thymine and cytosine at a rate ratio of 0.5- 1.5:0.5-1.5, 0.64.4:0.6-1 .4, 0.7-1.3:0.74.3, 0.84.2:0.8- 1.2, or around 1 : 1 respectively.
[0027] In such embodiments, in a step of polymerizing the at least one target nucleic acid molecule using a low bias DNA polymerase, the DNA polymerase mutates adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5- 1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 07-1.3:07-1.3, 0.8-1.2:0.8- 1.2, or around 1 : 1 respectively. In one embodiment, the low bias DNA polymerase mutates guanine and adenine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0 5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7- 1.3, 0.8-1.2:0.8-1.2, or around 1 : 1 respectively. In further embodiments, the low bias DNA polymerase mutates thymine and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0.5- 1.5, 0.6-1.4:0.6-1.4, 07-1.3:07-1.3, 0.8-1.2:0.8-1.2, or around 1 : 1 respectively.
[0028] Optionally, the low bias DNA polymerase is able to mutate adenine, thymine, guanine, and cytosine at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6- 1.4:0.6-1.4, 0.7- 1 3:0.7-1 .3:0.7-1.3:0.7-1 3, 0.8-1.2:0.8-1.2:0.8-1.2:0 8-1.2, or around 1 : 1 : 1 : 1 respectively. In certain embodiments, the low bias DNA polymerase is able to mutate adenine, thymine, guanine and cytosine at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0029] In such embodiments, in a step of polymerizing the at least one target nucleic acid molecule using a low bias DNA polymerase, the DNA polymerase may mutate adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1 5, 0.6-1.4:0.6-1.4:0.6-1.4:0 6- 1.4, 0.7-1.3:07- 1.3:0.7-1.3:07-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or around 1 : 1 : 1 : 1 respectively. In one embodiment, the low bias DNA polymerase mutates adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.7-1.3:07-1.3:07-1.3:07-1.3. The adenine, thymine, cytosine, and / or guanine may be substituted with another nucleotide. For example, if the low bias DNA polymerase is able to mutate adenine, polymerizing the at least one target nucleic acid molecule in the presence of the low bias DNA polymerase may substitute at least one adenine nucleotide in the nucleic acid molecule with thymine, guanine, or cytosine. Similarly, if the low bias DNA polymerase is able to mutate thymine, polymerizing the at least one target nucleic acid molecule in the presence of the low bias DNA polymerase may substitute at least one thymine nucleotide with adenine, guanine, or cytosine. If the low bias DNA polymerase is able to mutateguanine, copying the at least one target nucleotide in the presence of the low bias DNA polymerase may substitute at least one guanine nucleotide with thymine, adenine, or cytosine. If the low bias DNA polymerase is able to mutate cytosine, amplifying the at least one target nucleotide in the presence of the low bias DNA polymerase may substitute at least one cytosine nucleotide with thymine, guanine, or adenine.
[0030] The low bias DNA polymerase may not be able to substitute a nucleotide directly, but it may still be able to mutate that nucleotide by replacing the corresponding nucleotide on the complementary strand. For example, if the target nucleic acid molecule comprises thymine, there will be an adenine nucleotide present in the corresponding position of the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule. The low bias DNA polymerase may be able to replace the adenine nucleotide of the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule with a guanine and so, when the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule is replicated, this will result in a cytosine being present in the corresponding replicated at least one target nucleic acid molecule where there was originally a thymine (a thymine to cytosine substitution).
[0031] In an embodiment, the low bias DNA polymerase mutates between 1% and 15%, between 2% and 10%, or around 8% of the nucleotides in the at least one target nucleic acid. In such embodiments, the step of amplifying the at least one target nucleic acid molecule using a low bias DNA polymerase is carried out in such a way that between 1% and 15%, between 2% and 10%, or around 8% of the nucleotides in the at least one target nucleic acid are mutated.
[0032] In an embodiment, the low bias DNA polymerase is able to mutate between 0% and 3%, between 0% and 2%, between 0.1% and 5%, between 0.2% and 3%, or around 1.5% of the nucleotides in the at least one target nucleic acid molecule per polymerization event. In an embodiment, the low bias DNA polymerase mutates between 0% and 3%, between 0% and 2%, between 0.1% and 5%, between 0.2% and 3%, or around 1.5% of the nucleotides in the at least one target nucleic acid molecule per polymerization event.
[0033] The low bias DNA polymerase is able to mutate a nucleotide such as adenine, if, when used to amplify a nucleic acid molecule, it provides a nucleic acid molecule in which someinstances of that nucleotide are substituted or deleted. As used herein, the term “mutate” refers to introduction of substitution mutations, and in some embodiments the term “mutate” can be replaced with“ introduces substitutions of .
[0034] The low bias DNA polymerase mutates a nucleotide such as adenine in at least one target nucleic acid molecule if, when the step of polymerizing (e.g., copying or transcribing) the at least one target nucleic acid molecule using a low bias DNA polymerase is carried out, this step results in a mutated at least one target nucleic acid molecule in which instances of that nucleotide are mutated. For example, if the low bias DNA polymerase mutates adenine in the at least one target nucleic acid molecule, when the step of polymerizing the at least one target nucleic acid molecule using a low bias DNA polymerase is carried out, this step results in a mutated at least one target nucleic acid molecule in which at least one adenine has been substituted or deleted.
[0035] Mutation induction in accordance with the presently described techniques may also, or additionally, be induced using nucleotide analogs. By way of example, in certain embodiments the low bias DNA polymerase may not be able to replace nucleotides with other nucleotides directly (at least not with high frequency), but the low bias DNA polymerase may still be able to mutate a nucleic acid molecule using a nucleotide analog. The low bias DNA polymerase may be able to replace nucleotides with other natural nucleotides (i.e. cytosine, guanine, adenine or thymine) or with nucleotide analogs.
[0036] In embodiments employing nucleotide analogs for mutation introduction to a nucleic acid molecule, the DNA polymerase (or reverse transcriptase as noted above) need not be a low- bias DNA polymerase. That is, in the present context and with the introduction of nucleotide analogs, conventional or high-fidelity DNA polymerases may mutate a target nucleic acid molecule as they may be able to introduce nucleotide analogs into a target nucleic acid molecule.
[0037] In certain embodiments, both nucleotide analogs and low bias DNA polymerase may be employed to incorporate nucleotide analogs into a target nucleic acid molecule. In an embodiment, the low bias DNA polymerase incorporates nucleotide analogs into the at least one target nucleic acid molecule. In an embodiment, the low bias DNA polymerase can mutate adenine, thymine, guanine, and / or cytosine using a nucleotide analog. In an embodiment, the low bias DNA polymerase mutates adenine, thymine, guanine, and / or cytosine in the at least one targetnucleic acid molecule using a nucleotide analog. In an embodiment, the DNA polymerase replaces guanine, cytosine, adenine and / or thymine with a nucleotide analog. In an embodiment, the DNA polymerase can replace guanine, cytosine, adenine and / or thymine with a nucleotide analog.
[0038] Incorporating nucleotide analogs into the at least one target nucleic acid molecule can be used to mutate nucleotides, as they may be incorporated in place of existing nucleotides and they may pair with nucleotides in the opposite strand. For example dPTP, as discussed herein, can be incorporated into a nucleic acid molecule in place of a pyrimidine nucleotide (i.e., may replace thymine or cytosine). Once in a nucleic acid strand, it may pair with adenine when in an imino tautomeric form. Thus, when a complementary strand is formed, that complementary strand may have an adenine present at a position complementary to the dPTP. Similarly, once in a nucleic acid strand, it may pair with guanine when in an amino tautomeric form. Thus, when a complementary strand is formed, that complementary strand may have a guanine present at a position complementary to the dPTP.
[0039] For example, if a dPTP is introduced into the at least one target nucleic acid molecule, when an at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule is formed, the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule will comprise an adenine or a guanine at a position complementary to the dPTP in the at least one target nucleic acid molecule (depending on whether the dPTP is in its amino or imino form). When the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule is replicated, the resulting replicate of the at least one target nucleic acid molecule will comprise a thymine or a cytosine in a position corresponding to the dPTP in the at least one target nucleic acid molecule. Thus, a mutation to thymine or cytosine can be introduced into the mutated at least one target nucleic acid molecule.
[0040] Alternatively, if a dPTP is introduced in at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, when a replicate of the at least one target nucleic acid molecule is formed, the replicate of the at least one target nucleic acid molecule will comprise an adenine or a guanine at a position complementary to the dPTP in the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule (depending on the tautomeric form of the dPTP). Thus, a mutation to adenine or guanine can be introduced into the mutated at least one target nucleic acid molecule.
[0041] With this in mind, in an embodiment the low bias DNA polymerase can replace cytosine or thymine with a nucleotide analog. In a further embodiment, the low bias DNA polymerase introduces guanine or adenine nucleotides using a nucleotide analog at a rate ratio of 0.5-1.5:0.5- 1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or around 1 : 1 respectively. The guanine or adenine nucleotides may be introduced by the low bias DNA polymerase pairing them opposite a nucleotide analog such as dPTP. In a further embodiment, the low bias DNA polymerase introduces guanine or adenine nucleotides using a nucleotide analog at a rate ratio of 0.7-1.3:0.7- 1.3 respectively.
[0042] If a user wishes to mutate the at least one target nucleic acid molecule using a nucleotide analog, the method may comprise a step of polymerizing (e.g., copying or transcribing in the reverse transcription context) the at least one target nucleic acid molecule using a low bias DNA polymerase, where the step of polymerizing the at least one target nucleic acid molecule using a low bias DNA polymerase is carried out in the presence of the nucleotide analog.
[0043] Suitable nucleotide analogs include, but are not limited to, dPTP (2’ deoxy -P- nucleoside-5’ -triphosphate), 8- Oxo-dGTP (7,8-dihydro-8-oxoguanine), 5Br-dUTP (5-bromo-2’- deoxy-uridine-5’- triphosphate), 20H-dATP (2-hydroxy-2’ -deoxy adenosine-5’ -triphosphate), dKTP (9- (2 -Deoxy -P-D-ribofuranosyl)-N6-methoxy-2, 6, -diaminopurine-5’ -triphosphate) and diTP (2’-deoxyinosine 5 ’-trisphosphate). The nucleotide analog may be dPTP. The nucleotide analogs may be used to introduce the substitution mutations described in Table 1.Table 1
[0044] The different nucleotide analogs can be used, alone or in combination, to introducedifferent mutations into the at least one target nucleic acid molecule. Accordingly, the low bias DNA polymerase may introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations using a nucleotide analog. The low bias DNA polymerase may be able to introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations, optionally using a nucleotide analog.
[0045] The low bias DNA polymerase may be able to introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.5-1.5:0.5- 1.5:0.5-1.5:0.5-1.5, 0.6-1 .4:0.64.4:0.64.4:0.64.4, 0.7-1.3 :0.7-1.3 :0.7-1.3 :0.7-1.3, 0.8- 1.2:0.8-1.2:0.8-1.2:0.8-1.2, or around 1 : 1 : 1 : 1 respectively. In certain embodiments, the low bias DNA polymerase is able to introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.7-1.3:0.7-1.3:0.7- 1.3 :0.7-1.3 respectively.
[0046] In some methods the low bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.5- 1.5:0.5-1.5:0.5- 1.5:0.5-1 5, 0.6-1.4:0.64.4:0.6-1.4:0.6-1.4, 0.74.3:0.7-1.3:0.7- 1 3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8- 1.2:0.8-1.2, or around 1 : 1 : 1 : 1 respectively. In certain embodiments, the low bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.7-1.3:0.7-1.3:0.7- 1.3 :0.7-1.3 respectively.Example 1
[0047] With this in mind, and turning to FIG. 1, in a first example embodiment a method is provided for identifying an origin of amplified molecules within a biological sample. In accordance with this method, a respective first strand complementary DNA (cDNA) 208 is generated for individual RNA molecules 200 of a sample comprising a plurality of RNA molecules (e g., single-cell gene expression data for a sample comprising 100’s, 1,000’s, 10,000’s, or moreindividual cells). The generated first strand cDNAs 208 comprise one or more molecular landmarks (ML) introduced by a reverse transcription process (step 204) used to generate the respective first strand cDNAs. The first strand cDNAs 208 are distinguishable with respect to origin based on the molecular landmarks (i.e., the molecular landmarks are unique or substantially unique to individual reverse transcribed molecules). An amplification (step 212) of the first strand cDNAs 208 is performed to generate a plurality of amplicons 216 for each first strand cDNA. Each respective amplicon 212 comprises the respective one or more molecular landmarks present on the first strand cDNA 208 from which the respective amplicon 216 is amplified. Read counts 224 are generated (step 220) for each first strand cDNA 208 and amplicon 216. Each amplicon 216 is associated with the first strand cDNA 208 from which the respective amplicon is amplified based on the molecular landmarks. Each first strand cDNA 208 corresponds to an RNA molecule 200 present in the sample.
[0048] With respect to the first example embodiment, various further aspects or refinements may be provided, which may be utilized in isolation or in any suitable combination with one another. By way of example, in certain further aspects of the first example embodiment the sample is from a single cell or tissue sample. In addition, as discussed herein the molecular landmarks associated with respective first strand cDNAs 208 are used to distinguish the first strand cDNAs. The molecular landmarks in such embodiments may be generated by the reverse transcription process (step 204) using one or more of: an error-prone reverse transcriptase (e.g., a mutagenesis transcription in which mutations are introduced during transcription into transcribed polynucleotides); one or more mutagenic nucleotides; one or more nucleotide analogs (e.g., dPTP and / or 8-oxo-dGTP); or reaction conditions selected to introduce transcription errors. By way of example, in certain embodiments the mutagenesis transcription comprises transcribing the plurality of RNA molecules 200 with a low bias reverse transcriptase and / or with a nucleotide analogue. Further, in some embodiments the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
[0049] In addition, with respect to the first example embodiment various further actions may be performed. Such further actions may be performed in any suitable combination with one another. By way of example, in one embodiment a sequencing operation is performed on at least the first strand cDNAs 208 and amplicons 216. In such an embodiment the read counts 224 arebased on an output of the sequencing operation. In a further embodiment the read count 224 is adjusted or corrected (e.g., deduplicated (step 228) for each first strand cDNA based on the association of each amplicon 216 to the first strand cDNA 208 from which the respective amplicon is copied. In such an embodiment quantitative gene expression data 232 may be generated based on the adjusted or corrected read counts.
[0050] Further, with respect to the first example embodiment additional acts may be performed prior to performing the amplification (step 212). By way of example, and with reference to FIG. 2, prior to performing the amplification a plurality of second strand cDNAs 250 may be generated. Each second strand cDNA 250 is generated using a first strand cDNA 208 as a template. The second strand cDNAs 250 may comprise different molecular landmarks introduced by a polymerization process (step 254). Second strand cDNAs 250 are distinguishable with respect to origin based on the different molecular landmarks. Performing the amplification (step 212) generates a plurality of second strand cDNA amplicons for the second strand cDNAs in addition to plurality of first strand cDNA amplicons for the first strand cDNAs, shown collectively by reference number 258. Respective second strand cDNA amplicons comprise the associated different molecular landmarks present on the second strand cDNA 250 from which the corresponding second strand cDNA amplicon is amplified. Generating (step 220) read counts 224 further generates read counts for each second strand cDNA and second strand cDNA amplicon. Each second strand cDNA amplicon is associated with the second strand cDNA from which the respective second strand cDNA amplicon is amplified based on the different molecular landmarks and, further, each second strand cDNA corresponds to an RNA molecule 200 present in the sample. In addition, as discussed herein the molecular landmarks associated with respective second strand cDNAs are used to distinguish the second strand cDNAs. The molecular landmarks in such embodiments may be generated by the polymerization process using one or more of: an error- prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors. In further embodiments the second strand cDNA is generated using template switching or a template switching oligomer.
[0051] By way of further example, in some embodiments the step of generating the molecular landmarks as part of the polymerization process involves utilizing a mutagenesis polymerization such that mutations are introduced into copied polynucleotides. In some embodiments this mayinvolve polymerizing the second strand cDNAs 250 with a low bias DNA polymerase and / or with a nucleotide analogue. In some embodiments, the nucleotide analogue comprises (such as, 6H,8H- 3,4-Dihydro-pyrimido(4,5- c)(l,2)oxazin-7-one-8-u-D-2'-deoxy-ribofuranoside-5'-triphosphate) and / or 8-oxo-dGTP. In some embodiments, the low bias DNA polymerase is a Thermococcal polymerase, or a functional derivative thereof. In some embodiments, the Thermococcal polymerase is derived from a Thermococcal strain selected from the group consisting of T. kodakarensis, T. siculi, T. celer, and T. sp KS-1.Example 2
[0052] In a second example embodiment, and turning to FIG. 3, a method is provided for identifying an origin of amplified molecules within a biological sample. In accordance with this method, a respective first strand complementary DNA (cDNA) 266 is generated via reverse transcription (step 270) for individual RNA molecules 200 of a sample comprising a plurality of RNA molecules. A plurality of second strand cDNAs 250 are generated. Each second strand cDNA 250 is generated using a first strand cDNA 266 as a template. The generated second strand cDNAs 250 comprise one or more molecular landmarks introduced by a polymerization process (step 254) used to generate the respective second strand cDNAs. Second strand cDNAs 250 are distinguishable with respect to origin based on the molecular landmarks. An amplification (step 212) of each second strand cDNA 250 is performed to generate a plurality of amplicons 262 for each second strand cDNA. Each respective amplicon 262 comprises the respective one or more molecular landmarks present on the second strand cDNA 250 from which the respective amplicon is amplified. Read counts 224 are generated (step 220) for each second strand cDNA 250 and amplicons 262. Each amplicon 262 is associated with the second strand cDNA 250 from which the respective amplicon is amplified based on the molecular landmarks. Each second strand cDNA 250 corresponds to an RNA molecule 200 present in the sample.
[0053] With respect to the second example embodiment, various further aspects or refinements may be provided, which may be utilized in isolation or in any suitable combination with one another. By way of example, in certain further aspects of the second example embodiment the sample is from a single cell or tissue sample. In addition, as discussed herein the molecular landmarks associated with respective second strand cDNAs 250 are used to distinguish second strand cDNAs. The molecular landmarks in such embodiments may be generated by thepolymerization process (step 254) using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors. Further, in some embodiments the second strand cDNA is generated using template switching or a template switching oligomer.
[0054] In addition, with respect to the second example embodiment various further actions may be performed. Such further actions may be performed in any suitable combination with one another. By way of example, in one embodiment a sequencing operation is performed on at least the second strand cDNAs 250 and amplicons 262. In such an embodiment the read counts 224 are based on an output of the sequencing operation. In a further embodiment the read count 224 is adjusted or corrected (e.g., deduplicated (step 228) for each second strand cDNA 250 based on the association of each amplicon 262 to the second strand cDNA from which the respective amplicon is copied. In such an embodiment quantitative gene expression data 232 may be generated based on the adjusted or corrected read counts.
[0055] Further, with respect to the second example embodiment additional acts may be performed related to the first strand cDNA. By way of example, and with reference to FIG. 2, the first strand cDNA 208 for each RNA molecule may be generated to comprise different molecular landmarks introduced by a reverse transcription process (step 204). In such an embodiment, first strand cDNAs 208 are distinguishable with respect to origin based on the different molecular landmarks. Performing the amplification (step 212) generates a plurality of first strand cDNA amplicons (collectively with the second strand amplicons referred to herein as amplicons 258) for the first strand cDNAs 208. Respective first strand cDNA amplicons comprise the associated different molecular landmarks present on the first strand cDNA from which the corresponding first strand cDNA amplicon is amplified. Generating read counts 224 further generates read counts for each first strand cDNA and first strand cDNA amplicons. Each first strand cDNA amplicon is associated with the first strand cDNA 208 from which the respective first strand cDNA amplicon is amplified based on the different molecular landmarks and, further, each first strand cDNA 208 corresponds to an RNA molecule 200 present in the sample.
[0056] In addition, as discussed herein the different molecular landmarks associated with respective first strand cDNAs 208 are used to distinguish first strand cDNAs. The molecular landmarks in such embodiments may be generated by the reverse transcription process (step 204)using one or more of: an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors. In further embodiments the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.Example 3
[0057] In a third example embodiment a processor-implemented method is provided for generating gene expression data. In accordance with this method, a read count is generated for a plurality of first strand complementary DNA (cDNA) and corresponding amplicons derived from a sample. Each first strand cDNA is derived from a respective RNA molecule and comprises one or more molecular landmarks introduced by a reverse transcription process used to generate the respective first strand cDNA. Each amplicon comprises the respective one or more molecular landmarks of the first strand cDNA from which the respective amplicon is amplified. Each amplicon is associated with the first strand cDNA from which the respective amplicon is amplified based on the molecular landmarks. The read count for each first strand cDNA is adjusted or corrected based on the association of each amplicon to the first strand cDNA from which the respective amplicon is copied. Quantitative gene expression data is generated based on the adjusted or corrected read counts.
[0058] In addition, with respect to the third example embodiment various further actions may be performed. Such further actions may be performed in any suitable combination with one another. By way of example, in one embodiment an output is received from a sequencer device. The read count for the plurality of first strand cDNA and corresponding amplicons is generated from the output.
[0059] With respect to the third example embodiment, various further aspects or refinements may be provided, which may be utilized in isolation or in any suitable combination with one another. By way of example, in certain further aspects of the third example embodiment the one or more molecular landmarks associated with respective first strand cDNAs are used to distinguish first strand cDNAs. The one or more molecular landmarks in such embodiments may be generated by the reverse transcription process using one or more of an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected tointroduce transcription errors. Further, in some embodiments the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
[0060] Further, with respect to the third example embodiment additional acts may be performed related to the second strand cDNA. By way of example, an additional read count may be generated for a plurality of second strand cDNAs and corresponding second strand cDNA amplicons. Each second strand cDNA is generated using a first strand cDNA as a template. The second strand cDNAs comprise different molecular landmarks introduced by a polymerization process used to generate the respective second strand cDNA. Each second strand cDNA amplicon comprises the respective different molecular landmarks of the second strand cDNA from which the respective second strand cDNA amplicon is amplified. Each second strand cDNA amplicon is associated with the second strand cDNA from which the respective second strand cDNA amplicon is amplified based on the different molecular landmarks. Each second strand cDNA corresponds to an RNA molecule present in the sample. The additional read count for each second strand cDNA is adjusted or corrected based on the association of each second strand cDNA amplicon to the second strand cDNA from which the respective second strand cDNA amplicon is copied. The quantitative gene expression data is further generated based on the adjusted or corrected additional read counts.
[0061] In addition, as discussed herein the different molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs. The different molecular landmarks in such embodiments may be generated by the polymerization process using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors. In further embodiments the second strand cDNA is generated using template switching or a template switching oligomer.Example 4
[0062] In a fourth example embodiment a processor-implemented method is provided for generating gene expression data. In accordance with this method, a read count is generated for a plurality of second strand complementary (cDNA) and corresponding amplicons derived from a sample. Each second strand cDNA is derived from a respective first strand cDNA molecule that is derived from a respective mRNA molecule. Each second strand cDNA comprises one or moremolecular landmarks introduced by a polymerization process. Each amplicon comprises the respective one or more molecular landmarks of the second strand cDNA strand from which the respective amplicon is amplified. Each amplicon is associated with the second strand cDNA from which the respective amplicon is amplified based on the molecular landmarks. Each second strand cDNA corresponds to an RNA molecule present in the sample. The read count for each second strand cDNA is adjusted or corrected based on the association of each amplicon to the second strand cDNA from which the respective amplicon is copied. Quantitative gene expression data is generated based on the adjusted or corrected read counts.
[0063] In addition, with respect to the fourth example embodiment various further actions may be performed. Such further actions may be performed in any suitable combination with one another. By way of example, in one embodiment an output is received from a sequencer device. The read count for the plurality of second strand cDNA and corresponding amplicons is generated from the output.
[0064] With respect to the fourth example embodiment, various further aspects or refinements may be provided, which may be utilized in isolation or in any suitable combination with one another. By way of example, in certain further aspects of the fourth example embodiment the one or more molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs. The one or more molecular landmarks in such embodiments may be generated by the polymerization process using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors. Further, in some embodiments the second strand cDNA is generated using template switching or a template switching oligomer.
[0065] Further, with respect to the fourth example embodiment additional acts may be performed related to the first strand cDNA. By way of example, an additional read count may be generated for a plurality of first strand cDNAs and corresponding first strand cDNA amplicons. Each first strand cDNA is derived from a respective RNA molecule and comprises one or more different molecular landmarks introduced by a reverse transcription process used to generate the respective first strand cDNA. Each first strand cDNA amplicon comprises the respective different molecular landmarks of the first strand cDNA from which the respective first strand cDNA amplicon is amplified. Each first strand cDNA amplicon is associated with the first strand cDNAfrom which the respective first strand cDNA amplicon is amplified based on the different molecular landmarks. The additional read count for each first strand cDNA is adjusted or corrected based on the association of each first strand cDNA amplicon to the first strand cDNA from which the respective first strand cDNA amplicon is copied. The quantitative gene expression data is further generated based on the adjusted or corrected additional read counts.
[0066] In addition, as discussed herein the one or more different molecular landmarks associated with respective first strand cDNAs are used to distinguish first strand cDNAs. The different molecular landmarks in such embodiments may be generated by the reverse transcription process using one or more of an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors. In further embodiments the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.Example 5
[0067] In a fifth example embodiment a kit for quantifying gene expression is provided. In accordance with this embodiment, a kit may comprise one or more solid support structures, each comprising at least one RNA capture sequence. The kit may further comprise a reverse transcriptase and a DNA polymerase, at least one of which is error prone.
[0068] In addition, with respect to the fifth example embodiment various further components may be present in the kit. Such further components may be included in any suitable combination with one another in the kit. By way of example, in certain embodiments the kit may further comprise one or more of a ribosomal RNA (rRNA) depletion agent, an RNA fragmenting enzyme, a random primer mixture, a poly(A) priming mixture or an oligo (dT) priming mixture, or a targeted or gene specific priming mixture.
[0069] Further, various embodiments of the kit may comprise a reverse transcriptase having an error rate of 1% to 5% inclusive, 1% to 10% inclusive, 1% to 20% inclusive, or of at least 1%. Similarly, in certain embodiments the kit may comprise a DNA polymerase having an error rate of 1% to 5% inclusive, 1% to 10% inclusive, 1% to 20% inclusive, or of at least 1%. In certain embodiments one or both of the reverse transcriptase or the DNA polymerase are engineered to have a threshold error-rate. Similarly, in certain embodiments one or both of the reversetranscriptase or the DNA polymerase have a known relationship between error rate and reaction conditions. In addition, in certain embodiments the reverse transcriptase has terminal transferase activity.
[0070] In certain embodiments the RNA capture sequences comprise one of a targeted capture sequence or a universal capture sequence. By way of example, the universal capture sequence, if present, may comprise a random sequence, a semi-random sequence, or a poly(T) sequence. Further, in certain embodiments the at least one RNA capture sequence comprises a cell-specific, tissue-specific, or molecule specific bar code. Similarly, in certain embodiments the at least one RNA capture sequence is at least 3 nucleotides in length.
[0071] In certain embodiments the one or more solid support structures comprise one or more bead structures or one or more planar surfaces. By way of example, the one or more bead structures, if present, may comprise one or more magnetic bead structures and / or solid or hydrogel beads. Similarly, the one or more planar surfaces, if present, may comprise an array substrate, a flow cell, or a glass slide.Example 6
[0072] In a sixth example embodiment a kit for quantifying gene expression is provided. In accordance with this embodiment, a kit may comprise one or more solid support structures each comprising at least one RNA capture sequence. The kit may further comprise a reverse transcriptase, a DNA polymerase, one or both of mutagenic nucleotides or nucleotide analogs.
[0073] In addition, with respect to the sixth example embodiment various further components may be present in the kit. Such further components may be included in any suitable combination with one another in the kit. By way of example, in certain embodiments the kit may further comprise one or more of a ribosomal RNA (rRNA) depletion agent, an RNA fragmenting enzyme, a random primer mixture, a poly(A) priming mixture or an oligo (dT) priming mixture, or a targeted or gene specific priming mixture.
[0074] In certain embodiments the RNA capture sequences comprise one of a targeted capture sequence or a universal capture sequence. By way of example, the universal capture sequence, if present, may comprise a random sequence, a semi-random sequence, or a poly(T) sequence. Further, in certain embodiments the at least one RNA capture sequence comprises a cell-specific,tissue-specific, or molecule specific bar code. Similarly, in certain embodiments the at least one RNA capture sequence is at least 3 nucleotides in length.
[0075] In certain embodiments the one or more solid support structures comprise one or more bead structures or one or more planar surfaces. By way of example, the one or more bead structures, if present, may comprise one or more magnetic bead structures and / or solid or hydrogel beads. Similarly, the one or more planar surfaces, if present, may comprise an array substrate, a flow cell, or a glass slide.
[0076] Further, various embodiments of the kit may comprise a reverse transcriptase having terminal transferase activity. In addition, in certain embodiments one or both of the reverse transcriptase or the DNA polymerase have a known relationship between error rate and reaction conditions.Experimental Results
[0077] With the preceding in mind, aspects of the above-described example approaches were tested using two mutagenic nucleotides: 8-oxo-dGTP and dPTP. The mutagenic nucleotides were incorporated via PCR (Taq Pol). The primary type of mutation induced by 8-oxo-dGTP was a transversion mutation, with the primarily induced mutations being A:T to C:G and T:A to G:C. The primary type of mutation induced by dPTP was a transition mutation, with the primarily induced mutations being A:T to G:C and G:C to A:T. It may be noted that in other experiments other approaches for inducing errors in reverse transcription were employed, such as using the following concentrations of dNTPs: 8 mM 8-oxo-dGTP, 0.25 mM dPTP, 0.05 mM dGTP, 0.075 mM dCTP, and 0.075 mM TTP
[0078] The cancer cell line K-562 was used as the experimental system as it is suitable for single-cell methodologies and provides good (i.e., high) capture of molecules. For the presently described experiment spike-in was limited to <10% of the total reverse transcription reaction and the formulation of the dNTPs was unchanged from the kit reagents. Correspondingly, the mutagenic nucleotides were at a lower concentration than regular dNTPs. Table 2 illustrates the experimental conditions:Table 2All preparations were determined to be successful, as shown in FIG. 4 in a tabular format. Sequencing was performed on a NV6K SP flowcell with a PE300 configuration (50 R1 + 450 R2). Saturation was 0.13 and deemed sufficient for the experiment. Secondary single-cell metrics were not different between the conditions, demonstrating that the additives were tolerated in the preparation. Metrics are shown in FIG. 5 in a tabular format.
[0079] Turning to FIG. 6, results related to increased load of specific transition mutations are illustrated for the experimental results. In the depicted plot, summarized mutation rates across cycles are shown as normalized rate (vertical axis) versus mutation type (horizontal axis). In the depicted results, A > G and T > C mutations show a consistent trend of increasing with concentration of the mutagenic nucleotide. As noted above, dPTP is associated with transition mutations while dOXO-8-GTP is associated with transversion mutations. Accordingly, the observed result likely primarily reflects dPTP induced mutations. In certain embodiments, and in accordance with observed experimental work, a concentration of 10-fold higher dOXO to dPTP may be employed to obtain the observed mutation rates.
[0080] With respect to whether the observed differences are significant, and turning to FIG. 7, each concentration was tested against the 0 mM concentration. As shown via plots in FIG. 7, the increase in A > G at 0.37 mM and 0.71 mM concentrations are highly significant. Further, and turning to FIG. 8, focusing on the first 150 cycles show an increase in mutation rates going from a 0.37 mM concentration to a 0.71 mM concentration. The per-cycle increase is around 0.005 percent from baseline, so 0.01 percent total for the two mutations. As shown in FIG. 9, this trend can be visualized per cycle (horizontal axis), with the first 150 cycles showing a consistent increase in mutation rate with respect to concentration of the mutagenic nucleotides. For later cycles, the 0.37 mM and 0.71 mM concentrations yield similar increases compared to the control (i.e., 0 mM)and 0.18 mM concentrations.
[0081] With the preceding in mind, mutagenic nucleotides do appear to successfully incorporate into cDNA and result in an increased mutation load, specifically A > G and T > C transitions in the experimental data described herein. As may be appreciated, in addition to the mutagenic nucleotides described herein, other possible nucleotides may be employed to induce mutation in cDNA as discussed. By way of further, non-limiting example, Table 3 below illustrates possible mutagenic nucleotides potentially suitable for use with the techniques described herein. As will be appreciated, however, other mutagenic nucleotides may be suitable for use in the present approaches, whether listed in Table 3 or not.Table 3
[0082] While the preceding conveys certain aspects of the present techniques that may be useful in associating amplicons with a molecule of origin, such as an RNA molecule in the context of gene expression evaluation, other aspects of these techniques may be useful in facilitating the generation or assembly of long read sequences from a plurality of shorter reads (i.e., stitching) based on the use of molecular landmarks as described herein. By way of example, in such further embodiments a molecule of interest, such as a DNA or RNA molecule, may be provided in a sample to be analyzed. For the purpose of consistency with the preceding discussion, an RNA molecule may be assumed in this example, though it should be appreciated that the described approach is not limited to RNA processing.
[0083] In this example, as in the preceding discussion, the RNA molecule in question may undergo reverse transcription to generate a first strand cDNA. The first strand cDNA in turn may undergo polymerization to generate a second strand cDNA. As discussed herein one or both of the reverse transcription or polymerization may be error prone, such as by use of an error-prone transcriptase or polymerase, the use of mutagenic nucleotides, or a combination of such approaches. Correspondingly, one or both of the first strand cDNA or second strand cDNA may include distinctive molecular landmarks as discussed herein. Complementary first strand and second strand cDNA may form dual strand complexes, which may in turn undergo fragmentation, such as via the introduction of suitable transposase agents. A library may be prepared using conventional techniques based on the fragments generated in this manner and sequencing may be performed to generate a plurality of short read sequences, some of which will include the respective molecular landmarks present on the initial respective cDNA strands. The molecular landmarks present at different locations in the short read sequence data may in turn be used to digitally associate and assemble (i.e., stitch) the short read data into longer sequence reads up to and including sequence reads corresponding to the full length of the original RNA (or DNA) molecule. In this manner, the presently described molecular landmarks generated using the techniques described herein may be used to stitch together fragment sequence data to construct a more useful long read or full sequence.Miscellaneous
[0084] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in thedescription herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
[0085] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” or “derived from” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on” or “derived from”. Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.
[0086] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.
[0087] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departingfrom its scope. While the dimensions, types of materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain- English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U.S.C. §112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.
[0088] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.
Claims
What is claimed is:
1. A method for identifying an origin of amplified molecules within a biological sample, the method comprising: generating a respective first strand complementary DNA (cDNA) for individual RNA molecules of a sample comprising a plurality of RNA molecules, wherein the generated first strand cDNAs comprise one or more molecular landmarks introduced by a reverse transcription process used to generate the respective first strand cDNAs, wherein first strand cDNAs are distinguishable with respect to origin based on the molecular landmarks; performing an amplification of the first strand cDNAs to generate a plurality of amplicons for each first strand cDNA, wherein each respective amplicon comprises the respective one or more molecular landmarks present on the first strand cDNA from which the respective amplicon is amplified; generating read counts for each first strand cDNA and amplicon; and associating each amplicon with the first strand cDNA from which the respective amplicon is amplified based on the molecular landmarks, wherein each first strand cDNA corresponds to an RNA molecule present in the sample.
2. The method of claim 1, wherein the sample is from a single cell or tissue sample.
3. The method of claim 1, further comprising: performing a sequencing operation on at least the first strand cDNAs and amplicons, wherein the read counts are based on an output of the sequencing operation.
4. The method of claim 1, further comprising: adjusting or correcting the read count for each first strand cDNA based on the association of each amplicon to the first strand cDNA from which the respective amplicon is copied.
5. The method of claim 4, further comprising: generating quantitative gene expression data based on the adjusted or corrected read counts.
6. The method of claim 1, wherein the molecular landmarks associated with respective first strand cDNAs are used to distinguish the first strand cDNAs and are generated by the reverse transcription process using one or more of: an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors.
7. The method of claim 1, wherein the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
8. The method of claim 1, further comprising: prior to performing the amplification, generating a plurality of second strand cDNAs, wherein each second strand cDNA is generated using a first strand cDNA as a template, wherein the second strand cDNAs comprise different molecular landmarks introduced by a polymerization process, wherein second strand cDNAs are distinguishable with respect to origin based on the different molecular landmarks; wherein performing the amplification generates a plurality of second strand cDNA amplicons for the second strand cDNAs, wherein respective second strand cDNA amplicons comprise the associated different molecular landmarks present on the second strand cDNA from which the corresponding second strand cDNA amplicon is amplified;wherein generating read counts further generates read counts for each second strand cDNA and second strand cDNA amplicon; and wherein each second strand cDNA amplicon is associated with the second strand cDNA from which the respective second strand cDNA amplicon is amplified based on the different molecular landmarks, wherein each second strand cDNA corresponds to an RNA molecule present in the sample.
9. The method of claim 8, wherein the different molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs and are generated by the polymerization process using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors.
10. The method of claim 8, wherein the second strand cDNA is generated using template switching or a template switching oligomer.
11. A method for identifying an origin of amplified molecules within a biological sample, the method comprising: generating a respective first strand complementary DNA (cDNA) for individual RNA molecules of a sample comprising a plurality of RNA molecules; generating a plurality of second strand cDNAs, wherein each second strand cDNA is generated using a first strand cDNA as a template, wherein the generated second strand cDNAs comprise one or more molecular landmarks introduced by a polymerization process used to generate the respective second strand cDNAs, wherein second strand cDNAs are distinguishable with respect to origin based on the molecular landmarks;performing an amplification of each second strand cDNA to generate a plurality of amplicons for each second strand cDNA, wherein each respective amplicon comprises the respective one or more molecular landmarks present on the second strand cDNA from which the respective amplicon is amplified; generating read counts for each second strand cDNA and amplicon; and associating each amplicon with the second strand cDNA from which the respective amplicon is amplified based on the molecular landmarks, wherein each second strand cDNA corresponds to an RNA molecule present in the sample.
12. The method of claim 11, wherein the sample is from a single cell or tissue sample.
13. The method of claim 11, further comprising: performing a sequencing operation on at least the second strand cDNAs and amplicons, wherein the read counts are based on an output of the sequencing operation.
14. The method of claim 11, further comprising: adjusting or correcting the read count for each second strand cDNA based on the association of each amplicon to the second strand cDNA from which the respective amplicon is copied.
15. The method of claim 14, further comprising: generating quantitative gene expression data based on the adjusted or corrected read counts.
16. The method of claim 11, wherein the molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs and are generated by the polymerization process using one or more of:an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors.
17. The method of claim 11, wherein the second strand cDNA is generated using template switching or a template switching oligomer.
18. The method of claim 11, further comprising: wherein the first strand cDNA for each RNA molecule is generated to comprise different molecular landmarks introduced by a reverse transcription process, wherein first strand cDNAs are distinguishable with respect to origin based on the different molecular landmarks; wherein performing the amplification generates a plurality of first strand cDNA amplicons for the first strand cDNAs, wherein respective first strand cDNA amplicons comprise the associated different molecular landmarks present on the first strand cDNA from which the corresponding first strand cDNA amplicon is amplified; wherein generating read counts further generates read counts for each first strand cDNA and first strand cDNA amplicon; and wherein each first strand cDNA amplicon is associated with the first strand cDNA from which the respective first strand cDNA amplicon is amplified based on the different molecular landmarks, wherein each first strand cDNA corresponds to an RNA molecule present in the sample.
19. The method of claim 18, wherein the different molecular landmarks associated with respective first strand cDNAs are used to distinguish first strand cDNAs and are generated by the reverse transcription process using one or more of: an error-prone reverse transcriptase;one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors.
20. The method of claim 18, wherein the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
21. A processor-implemented method for generating gene expression data, the method comprising: generating a read count for a plurality of first strand complementary DNA (cDNA) and corresponding amplicons derived from a sample, wherein each first strand cDNA is derived from a respective RNA molecule and comprises one or more molecular landmarks introduced by a reverse transcription process used to generate the respective first strand cDNA and wherein each amplicon comprises the respective one or more molecular landmarks of the first strand cDNA from which the respective amplicon is amplified; associating each amplicon with the first strand cDNA from which the respective amplicon is amplified based on the molecular landmarks; adjusting or correcting the read count for each first strand cDNA based on the association of each amplicon to the first strand cDNA from which the respective amplicon is copied; and generating quantitative gene expression data based on the adjusted or corrected read counts.
22. The processor-implemented method of claim 21, further comprising: receiving an output from a sequencer device, wherein the read count for the plurality of first strand cDNA and corresponding amplicons is generated from the output.
23. The processor-implemented method of claim 21, wherein the one or more molecular landmarks associated with respective first strand cDNAs are used to distinguish first strand cDNAs and are generated by the reverse transcription process using one or more of: an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors.
24. The processor-implemented method of claim 21, wherein the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
25. The processor-implemented method of claim 21, further comprising: generating an additional read count for a plurality of second strand cDNAs and corresponding second strand cDNA amplicons, wherein each second strand cDNA is generated using a first strand cDNA as a template, wherein the second strand cDNAs comprise different molecular landmarks introduced by a polymerization process used to generate the respective second strand cDNA and wherein each second strand cDNA amplicon comprises the respective different molecular landmarks of the second strand cDNA from which the respective second strand cDNA amplicon is amplified; associating each second strand cDNA amplicon with the second strand cDNA from which the respective second strand cDNA amplicon is amplified based on the different molecular landmarks, wherein each second strand cDNA corresponds to an RNA molecule present in the sample; adjusting or correcting the additional read count for each second strand cDNA based on the association of each second strand cDNA amplicon to the second strand cDNA from which the respective second strand cDNA amplicon is copied; and wherein generating the quantitative gene expression data is further based on the adjusted or corrected additional read counts.
26. The processor-implemented method of claim 25, wherein the different molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs and are generated by the polymerization process using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors.
27. The processor-implemented method of claim 25, wherein the second strand cDNA is generated using template switching or a template switching oligomer.
28. A processor-implemented method for generating gene expression data, the method comprising: generating a read count for a plurality of second strand complementary (cDNA) and corresponding amplicons derived from a sample, wherein each second strand cDNA is derived from a respective first strand cDNA molecule that is derived from a respective mRNA molecule, wherein each second strand cDNA comprises one or more molecular landmarks introduced by a polymerization process and wherein each amplicon comprises the respective one or more molecular landmarks of the second strand cDNA strand from which the respective amplicon is amplified; associating each amplicon with the second strand cDNA from which the respective amplicon is amplified based on the molecular landmarks, wherein each second strand cDNA corresponds to an RNA molecule present in the sample; adjusting or correcting the read count for each second strand cDNA based on the association of each amplicon to the second strand cDNA from which the respective amplicon is copied; andgenerating quantitative gene expression data based on the adjusted or corrected read counts.
29. The processor-implemented method of claim 28, further comprising: receiving an output from a sequencer device, wherein the read count for the plurality of second strand cDNA and corresponding amplicons is generated from the output.
30. The processor-implemented method of claim 28, wherein the one or more molecular landmarks associated with respective second strand cDNAs are used to distinguish second strand cDNAs and are generated by the polymerization process using one or more of: an error-prone polymerase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce polymerization errors.
31. The processor-implemented method of claim 28, wherein the second strand cDNA is generated using template switching or a template switching oligomer.
32. The processor-implemented method of claim 28, further comprising: generating an additional read count for a plurality of first strand cDNAs and corresponding first strand cDNA amplicons, wherein each first strand cDNA is derived from a respective RNA molecule and comprises one or more different molecular landmarks introduced by a reverse transcription process used to generate the respective first strand cDNA and wherein each first strand cDNA amplicon comprises the respective different molecular landmarks of the first strand cDNA from which the respective first strand cDNA amplicon is amplified; associating each first strand cDNA amplicon with the first strand cDNA from which the respective first strand cDNA amplicon is amplified based on the different molecular landmarks;adjusting or correcting the additional read count for each first strand cDNA based on the association of each first strand cDNA amplicon to the first strand cDNA from which the respective first strand cDNA amplicon is copied; and wherein generating the quantitative gene expression data is further based on the adjusted or corrected additional read counts.
33. The processor-implemented method of claim 32, wherein the one or more different molecular landmarks associated with respective first strand cDNAs are used to distinguish first strand cDNAs and are generated by the reverse transcription process using one or more of: an error-prone reverse transcriptase; one or more mutagenic nucleotides; one or more nucleotide analogs; or reaction conditions selected to introduce transcription errors.
34. The processor-implemented method of claim 32, wherein the reverse transcription process utilizes a reverse transcriptase having terminal transferase activity.
35. A kit for quantifying gene expression, comprising: one or more solid support structures each comprising at least one RNA capture sequence; and a reverse transcriptase and a DNA polymerase, at least one of which is error prone.
36. The kit of claim 35, wherein the reverse transcriptase has an error rate of 1% to 5% inclusive.
37. The kit of claim 35, wherein the reverse transcriptase has an error rate of 1 % to 10% inclusive.
38. The kit of claim 35, wherein the reverse transcriptase has an error rate of 1% to 20% inclusive.
39. The kit of claim 35, wherein the reverse transcriptase has an error rate of at least 1%.
40. The kit of claim 35, wherein the reverse transcriptase has terminal transferase activity.
41. The kit of claim 35, wherein the DNA polymerase has an error rate of 1% to 5% inclusive.
42. The kit of claim 35, wherein the DNA polymerase has an error rate of 1% to 10% inclusive.
43. The kit of claim 35, wherein the DNA polymerase has an error rate of 1% to 20% inclusive.
44. The kit of claim 35, wherein the DNA polymerase has an error rate of at least 1%.
45. The kit of claim 35, further comprising a ribosomal RNA (rRNA) depletion agent.
46. The kit of claim 35, further comprising an RNA fragmenting enzyme.
47. The kit of claim 35, further comprising a random primer mixture.
48. The kit of claim 35, further comprising a poly(A) priming mixture or an oligo (dT) priming mixture.
49. The kit of claim 35, further comprising a targeted or gene specific priming mixture.
50. The kit of claim 35, wherein the RNA capture sequences comprise one of a targeted capture sequence or a universal capture sequence.
51. The kit of claim 50, wherein the universal capture sequence comprises a random sequence, a semi-random sequence, or a poly(T) sequence.
52. The kit of claim 35, wherein one or more solid support structures comprise one or more bead structures or one or more planar surfaces.
53. The kit of claim 52, wherein the one or more bead structures comprise one or more magnetic bead structures.
54. The kit of claim 52, wherein the one or more bead structures comprise solid or hydrogel beads.
55. The kit of claim 52, wherein the one or more planar surfaces comprise an array substrate, a flow cell, or a glass slide.
56. The kit of claim 35, wherein the at least one RNA capture sequence comprises a cellspecific, tissue-specific, or molecule specific bar code.
57. The kit of claim 35, wherein the at least one RNA capture sequence is at least 3 nucleotides in length.
58. The kit of claim 35, wherein one or both of the reverse transcriptase or the DNA polymerase are engineered to have a threshold error-rate.
59. The kit of claim 35, wherein one or both of the reverse transcriptase or the DNA polymerase have a known relationship between error rate and reaction conditions.
60. A kit for quantifying gene expression, comprising: one or more solid support structures each comprising at least one RNA capture sequence; and a reverse transcriptase; a DNA polymerase; and one or both of mutagenic nucleotides or nucleotide analogs.
61. The kit of claim 60, wherein the reverse transcriptase has terminal transferase activity.
62. The kit of claim 60, further comprising a ribosomal RNA (rRNA) depletion agent.
63. The kit of claim 60, further comprising an RNA fragmenting enzyme.
64. The kit of claim 60, further comprising a random primer mixture.
65. The kit of claim 60, further comprising a poly(A) priming mixture or an oligo (dT) priming mixture.
66. The kit of claim 60, further comprising a targeted or gene specific priming mixture.
67. The kit of claim 60, wherein the RNA capture sequences comprise one of a targeted capture sequence or a universal capture sequence.
68. The kit of claim 67, wherein the universal capture sequence comprises a random sequence, a semi-random sequence, or a poly(T) sequence.
69. The kit of claim 60, wherein one or more solid support structures comprise one or more bead structures or one or more planar surfaces.
70. The kit of claim 69, wherein the one or more bead structures comprise one or more magnetic bead structures.
71. The kit of claim 69, wherein the one or more bead structures comprise solid or hydrogel beads.
72. The kit of claim 69, wherein the one or more planar surfaces comprise an array substrate, a flow cell, or a glass slide.
73. The kit of claim 60, wherein the at least one RNA capture sequence comprises a cellspecific, tissue-specific, or molecule specific bar code.
74. The kit of claim 60, wherein the at least one RNA capture sequence is at least 3 nucleotides in length.
75. The kit of claim 60, wherein one or both of the reverse transcriptase or the DNA polymerase have a known relationship between error rate and reaction conditions.
Citation Information
Patent Citations
Random nucleotide mutation for nucleotide template counting and assembly
US11008606B2
Sequencing Algorithm
US20210174905A1
Sequencing Process
US20210403991A1
Method for introducing mutations
US20220348940A1
Method for determining a measure correlated to the probability that two mutated sequence reads derive from the same sequence comprising mutations
US20230044570A1