Methods and compositions for transcriptome characterization

By engineering a reverse transcriptase to inhibit strand displacement and RNAse-H activities, the method addresses the limitations of existing transcriptomics techniques, allowing for accurate retrieval of RNA sequence variations and comprehensive transcriptome profiling in fixed samples.

WO2025199320A1PCT designated stage Publication Date: 2025-09-25CORNELL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/020706
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-20
Filing Date
2025-03-20
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing single-cell and spatial transcriptomics techniques fail to capture RNA sequence variations such as mutations, deletions, or duplications in target RNA transcripts, limiting the accuracy of gene expression analysis.

Method used

Engineering a reverse transcriptase with mutations to inhibit strand displacement and RNAse-H activities, allowing for the incorporation of RNA sequences between non-contiguous probes, enabling accurate retrieval of RNA transcript sequences, particularly in fixed samples.

Benefits of technology

Enables simultaneous identification of single nucleotide variations and profiling of the entire transcriptome with spatial resolution, retaining information about mutations and duplications in fixed samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025020706_25092025_PF_FP_ABST
    Figure US2025020706_25092025_PF_FP_ABST
Patent Text Reader

Abstract

In some aspects, the present disclosure provides engineered reverse transcriptase variants that prevent strand displacement and / or prevent digestion of DNA:RNA hybrid, and compositions and uses thereof. In some aspects, the present disclosure provides methods for analyzing RNA transcripts, e.g., in fixed samples, using, inter aha, engineered reverse transcriptase enzyme variants.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODSAND COMPOSITIONS FOR TRANSCRIPTOME CHARACTERIZATION

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application No.: 63 / 567,644, filed March 20, 2024, the content of which is hereby incorporated by reference herein in its entirety.

[0004] TECHNICAL FIELD

[0005] This disclosure generally relates to the field of transcriptomics, such as single-cell and spatial transcriptomics.

[0006] BACKGROUND

[0007] Accurately retrieving RNA sequence information is essential for spatially evaluating gene expression, including evaluating how sequence variations affect cellular phenotypes. Single-cell and spatial transcriptomics techniques traditionally use contiguous probe pairs that hybridize to RNA target sequence and are then sequenced to obtain RNA sequence information. However, such techniques do not capture RNA sequence variations such as mutations, deletions, or duplications in the target RNA transcript. Therefore, compositions and methods useful for leveraging the advantages of single-cell spatial transcriptomics and accurate retrieval of RNA transcript sequences are needed.

[0008] BRIEF SUMMARY

[0009] Single-cell and spatial transcriptomics have emerged as powerful techniques that have revolutionized understanding of cellular dynamics and tissue architecture by providing spatially resolved gene expression data.

[0010] Methods and compositions that leverage single-cell and spatial transcriptomics to identify single nucleotide variations, small deletions and / or insertions in RNA transcripts are needed. Methodologies to achieve this are described herein. In particular, described herein are new and improved methods to genotype the transcriptome and / or profile or identify single nucleotide variations, small deletions and / or insertions. The methods described herein can be used in fixed samples (e.g., formalin-fixed samples).

[0011] In some aspects, provided herein is an engineered reverse transcriptase, wherein a wildtype reverse transcriptase is engineered to comprise: a) a first mutation that inhibits, reduces, interferes with, substantially abolishes, or abolishes the strand displacement activity of the reverse transcriptase relative to the wild-type reverse transcriptase, wherein the first mutation is one or more substitutions, deletions or insertions of amino acids (e.g., wherein the first mutation is an amino acid substitution such as any one or more described herein); and / or b) a second mutation that inhibits, reduces, interferes with, substantially abolishes, or abolishes RNAse-H activity of the reverse transcriptase relative to the wild-type reverse transcriptase, wherein the second mutation is one or more substitutions, deletions or insertions of amino acids (e.g., wherein the second mutation is an amino acid substitution such as any one or more described herein). In some embodiments, the inhibiting or reducing enzymatic activity is inhibiting or reducing enzymatic activity by at least or more than 25%, 30%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%. In some embodiments, the reverse transcriptase strand displacement activity is measured by any means known in the art or disclosed herein. In some embodiments, reverse transcriptase RNAse-H activity is measured by any means known in the art or disclosed herein. In some embodiments, one or both of these activities are inhibited by at least or more than 80%, 85%, 90%, 95%, 97%, 98%, or 99%, or to the level such that the activity is no longer detectable by one or more assays of such activity known in the art or described herein.

[0012] In some aspects, provided herein is an engineered reverse transcriptase, wherein a wildtype reverse transcriptase is engineered to comprise: a) a first mutation that inhibits, reduces, interferes with, substantially abolishes, or abolishes the strand displacement activity of the reverse transcriptase relative to the wild-type reverse transcriptase, wherein the first mutation is one or more substitutions, deletions or insertions of amino acids (e.g., wherein the first mutation is an amino acid substitution such as any one or more described herein); and / or b) a second mutation that inhibits, reduces, interferes with, substantially abolishes, or abolishes digestion of DNA:RNA hybrid by the reverse transcriptase relative to the wild-type reverse transcriptase, wherein the second mutation is one or more substitutions, deletions or insertions of amino acids (e.g., wherein the second mutation is an amino acid substitution such as any one or more described herein). In some embodiments, the inhibiting or reducing enzymatic activity is inhibiting or reducing enzymatic activity by at least or more than 25%, 30%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%. In some embodiments, the reverse transcriptase strand displacement activity is measured by any means known in the art or disclosed herein. In some embodiments, the digestion of DNA:RNA hybrid by reverse transcriptase is measured by any means known in the art or disclosed herein. In some embodiments, one or both of these activities are inhibited by at least or more than 80%, 85%, 90%, 95%, 97%, 98%, or 99%, or to the level such that the activity is no longer detectable by one or more assays of such activity known in the art or described herein.

[0013] In some embodiments, the wild-type reverse transcriptase is any natural reverse transcriptase. In some embodiments, the wild-type reverse transcriptase is HIV-1 reverse transcriptase, M-MLV reverse transcriptase, AMV reverse transcriptase, or Telomerase reverse transcriptase. In some embodiments, the wild-type reverse transcriptase is Moloney Murine Leukemia Virus Reverse Transcriptase (M-MLV RT) and / or comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1. In some embodiments, the engineered reverse transcriptase comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 97%, 98%, 99% identical to SEQ ID NO: 2 and further comprises a methionine (M) added at the beginning of the sequence (as the first amino acid).

[0014] In some embodiments, the first mutation is in the first finger domain of the wild-type reverse transcriptase. In some embodiments, the first finger domain comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 3. In some embodiments, the second mutation is in the RNAse-H domain of the wild-type reverse transcriptase. In some embodiments, the RNAse-H domain comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 4.

[0015] In some embodiments, the first mutation is in or corresponds to amino acid Y64 (e.g., Y64A) relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1, and optionally wherein the engineered reverse transcriptase is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1 or comprises a sequence that is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 3. In some embodiments, the first mutation is or corresponds to Y64A relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1. In some embodiments, the engineered reverse transcriptase comprises a sequence identical to SEQ ID NO: 1 or SEQ ID NO: 3 except for the Y64 (e.g., Y64A) mutation.

[0016] In some embodiments, the second mutation is in or corresponds to amino acid D524, D583 or E562 (e.g., D524A, D524G, D583N, D524N, or E562Q) relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1, and optionally wherein the engineered reverse transcriptase is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1 or comprises a sequence that is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 4. In some embodiments, the second mutation is or corresponds to D524A, D524G, D583N, D524N, or E562Q relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1. In some embodiments, the engineered reverse transcriptase comprises a sequence identical to SEQ ID NO: 1 or SEQ ID NO: 4 except for the D524, D583 or E562 (e.g., D524A, D524G, D583N, D524N, or E562Q) mutation.

[0017] In some embodiments, the first mutation is in or corresponds to amino acid Y64 or Y64A of M-MLV RT, and / or the second mutation is in or corresponds to amino acid D524, D583 or E562, or D524A, D524G, D583N, D524N, or E562Q, of M-MLV RT.

[0018] In some embodiments, the first mutation is or corresponds to Y64A (tyrosine-to-alanine at position 64) and / or the second mutation is or corresponds to D524A (aspartic acid-to-alanine at position 524) relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1, optionally wherein the engineered reverse transcriptase is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1 or comprises a sequence that is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 3 and / or SEQ ID NO: 4.

[0019] In some embodiments, the engineered reverse transcriptase has a sequence identical to SEQ ID NO: 1 except for an Y64 (e.g., Y64A) mutation and an D524, D583 or E562 (e.g., D524A, D524G, D583N, D524N, or E562Q) mutation.

[0020] In some embodiments, the engineered reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 2. In some embodiments, the engineered reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 2 and further comprises a methionine (M) added at the beginning of the sequence (as the first amino acid). In some embodiments, and without being bound by any theory, M is added at the beginning of the sequence of reverse transcriptase to improve or facilitate recombinant protein engineering and / or production. In some embodiments, the engineered reverse transcriptase comprises an amino acid sequence that is at least 85%, 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 2, wherein the sequence comprises Y64 (e.g., Y64A) substitution and an D524, D583 or E562 (e.g., D524A, D524G, D583N, D524N, or E562Q) substitution, and optionally wherein the sequence further comprises one or more amino acids at the beginning of the sequence, which amino acids may facilitate protein production (e.g., wherein the sequence further comprises M as the first amino acid at the beginning of the sequence). In some embodiments, the engineered reverse transcriptase comprises the first finger domain having the amino acid sequence of SEQ ID NO: 7. In some embodiments, the engineered reverse transcriptase comprises the RNAse-H domain having the amino acid sequence of SEQ ID NO: 8. In some embodiments, the engineered reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 7 and / or the amino acid sequence of SEQ ID NO: 8.

[0021] In some aspects, this disclosure provides a composition comprising the engineered reverse transcriptase disclosed herein. In some aspects, this disclosure provides a kit comprising the engineered reverse transcriptase disclosed herein.

[0022] In some aspects, this disclosure provides a nucleic acid encoding the engineered reverse transcriptase disclosed herein.

[0023] In some aspects, provided herein is a method for analyzing a target ribonucleic acid sequence (RNA) in a sample comprising a) contacting the sample with a first polynucleotide probe comprising a nucleotide sequence substantially complementary to a first sequence of the target RNA and a second polynucleotide probe comprising a nucleotide sequence substantially complementary to a second sequence of the target RNA, wherein the first sequence and the second sequence are separated by a third sequence of the target RNA comprising at least one nucleotide; b) contacting the sample with the engineered reverse transcriptase described herein (such as any of the engineered reverse transcriptases described herein), whereby the first polynucleotide probe is extended to produce a DNA complementary to the third sequence (cDNA); and c) conducting a ligation reaction to ligate the cDNA to the second polynucleotide probe (e.g., contacting the sample with a ligase). In some embodiments, the third sequence comprises up to 100 nucleotides, up to 75 nucleotides, up to 50 nucleotides, up to 40 nucleotides, up to 30 nucleotides, up to 25 nucleotides, or up to 20 nucleotides. In some embodiments, the third sequence comprises about 1 to 100 nucleotides. In some embodiments, the third sequence comprises about 1 to 75 nucleotides. In some embodiments, the third sequence comprises about 1 to 50, 1 to 40, 1 to 30, or 1 to 20 nucleotides. In some embodiments, the third sequence comprises about 2 to 100 nucleotides. In some embodiments, the third sequence comprises about 2 to 75 nucleotides. In some embodiments, the third sequence comprises about 2 to 50, 2 to 40, 2 to 30, or 2 to 20 nucleotides. In some embodiments, the third sequence comprises at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, or more than 3, 4, 5, 6, 7, or 8 nucleotides. In some embodiments, the third sequence comprises about 3 to 100 nucleotides. In some embodiments, the third sequence comprises about 3 to 75 nucleotides. In some embodiments, the third sequence comprises about 3 to 50, 3 to 40, 3 to 30, or 3 to 20 nucleotides. In some embodiments, the third sequence comprises about 3 to 40 nucleotides. In some embodiments, the third sequence comprises about 4 to 100 nucleotides. In some embodiments, the third sequence comprises about 4 to 75 nucleotides. In some embodiments, the third sequence comprises 4 to 50, 4 to 40, 4 to 35, 4 to 30, 4 to 25, or 4 to 20 nucleotides. In some embodiments, the third sequence comprises about 5 to 100 or 6 to 100 nucleotides. In some embodiments, the third sequence comprises about 5 to 75 or 6 to 75 nucleotides. In some embodiments, the third sequence comprises 5 to 50, 5 to 40, 5 to 35, 5 to 30, 5 to 25, or 5 to 20 nucleotides. In some embodiments, the third sequence comprises 6 to 50, 6 to 40, 6 to 35, 6 to 30, or 6 to 25 nucleotides. In some embodiments, the third sequence comprises 6 to 20 nucleotides. In some embodiments, the third sequence comprises 6 to 15, 6 to 12, 6 to 10 nucleotides. In some embodiments, the third sequence comprises 7 to 50, 7 to 40, 7 to 35, 7 to 30, 7 to 25, 7 to 20, 7 to 15, 7 to 12, or 7 to 10 nucleotides. In some embodiments, the third sequence comprises 8 to 50, 8 to 40, 8 to 35, 8 to 30, 8 to 25, 8 to 20, or 8 to 15 nucleotides. In some embodiments, the third sequence comprises 9 to 50, 9 to 40, 9 to 35, 9 to 30, 9 to 25, 9 to 20, or 9 to 15 nucleotides. In some embodiments, the third sequence comprises 9 to 20 nucleotides. In some embodiments, the target RNA is mRNA.

[0024] In some embodiments, the method provided herein utilizes an engineered reverse transcriptase (e.g., from M-MLV) in which the strand displacement activity is inhibited, reduced, interfered with, substantially abolished, or abolished as described herein, and optionally wherein the RNAse-H activity is inhibited, reduced, interfered with, substantially abolished, or abolished as described herein (e.g., comprising the sequence of SEQ ID NO: 2, or a variant thereof in which such activity or activities are inhibited). In some embodiments, the method provided herein inhibits (e.g., by at least or more than 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 100%) or prevents displacement of the second nucleotide probe from the target RNA. In some embodiments, the displacement of the second polynucleotide probe can be measured using any method known in the art or described herein.

[0025] In some embodiments, the first polynucleotide and / or the second polynucleotide probe comprises 1, 2, 3 or more non-natural nucleotides. In some embodiments, the first polynucleotide and / or the second polynucleotide probe comprises 4, 5, 6 or more non-natural nucleotides. In some embodiments, the non-natural nucleotides are selected from one or more of: locked nucleic acid (LNA), Super T (5-hydroxybutynl-2’-deoxyuridine), and Super G (8-aza- 7-deazaguanosine). In some embodiments, the first polynucleotide and / or the second polynucleotide probe comprises 1, 2, 3 or more locked nucleic acids (LNAs). In some embodiments, the second polynucleotide probe comprises at least one locked nucleic acid. In some embodiments, the first polynucleotide probe comprises at least one locked nucleic acid. In some embodiments, at least the first three positions of the second polynucleotide probe are LNA nucleotides. In some embodiments, at least the first three positions of the first polynucleotide probe are LNA nucleotides.

[0026] In some embodiments, step b) and step c) of the method provided herein occur simultaneously by contacting the sample with the engineered reverse transcriptase disclosed herein and a ligase at the same time. In some embodiments, step b) and step c) of the method provided herein occur sequentially.

[0027] In some embodiments, the method comprises, after step c), a step of binding the first polynucleotide probe and the second polynucleotide probe to a barcoded oligonucleotide. In some embodiments, the barcoded oligonucleotide is a spatially barcoded oligonucleotide.

[0028] In some embodiments, the method comprises after step c), amplifying the cDNA. In some embodiments, the amplifying is performed by real-time quantitative PCR.

[0029] In some embodiments, the method further comprises sequencing the cDNA. In some embodiments, the sequencing is next generation sequencing (NGS).

[0030] In some embodiments, the method comprises permeabilizing the sample before contacting the sample with the first and / or second polynucleotide probes. In some embodiments, the sample used in the method described herein is a fixed sample. In some embodiments, the fixed sample is formalin-fixed. In some embodiments, the fixed sample is formalin-fixed and paraffin-embedded (FFPE).

[0031] In some embodiments, the method described herein does not comprise using a third polynucleotide probe to gap-fill the sequence between the first polynucleotide probe and the second polynucleotide probe.

[0032] In some embodiments, the sample used in the method described herein is a single cell, and / or the target RNA is the target RNA of a single cell. In some embodiments, the sample is a tissue. In some embodiments, the target RNA is the transcriptome of a sample.

[0033] In some embodiments, the method described herein is for sequencing and / or genotyping a transcriptome comprising the target RNA. In some embodiments, the method provides spatial location of the target RNA. In some embodiments, the method described herein is for assessing, detecting the presence of or identifying a mutant allele, a single nucleotide variation, a splice isoform, or a TCR / BCR junction in the target RNA. In some embodiments, the method is for providing spatial location and / or single cell quantification of the mutant allele, the single nucleotide variation, the splice isoform or the TCR / BCR junction.

[0034] In some embodiments, the method described herein is for single cell transcriptome analysis, optionally wherein the sample is a single cell, and the target RNA is transcriptome of the single cell.

[0035] In some embodiments, the single cell transcriptome analysis is single cell CRISPR screening. In some embodiments, the single cell transcriptome analysis is single cell T-cell receptor (TCR) and / or B-cell receptor (BCR) sequencing.

[0036] In some embodiments, the sample used in the method described herein is a cell or population of cells from a tissue or organ associated with or affected by a disease or disorder in a subject.

[0037] BRIEF DESCRIPTION OF THE DRAWINGS

[0038] FIG. 1A is a schematic showing the domains of an engineered reverse transcriptase having two point mutations Y64A and D524A.

[0039] FIG. IB is a series of schematic diagrams showing methods described herein for profiling a target RNA sequence, including identifying mutations present therein.

[0040] FIG. 1C is a schematic showing a wild-type reverse transcriptase displacing a polynucleotide probe, and an engineered reverse transcriptase with no strand displacement activity leaving the second polynucleotide probe annealed to the target RNA.

[0041] FIG. 2 is a schematic showing the extension of a cDNA with a wild-type reverse transcriptase having strand displacement activity compared to extension of cDNA using an engineered reverse transcriptase with no strand displacement activity.

[0042] FIG. 3 is a schematic showing exemplary positions of a pair of primers, External FW and Probe RV.

[0043] FIG. 4A is an image of a protein gel showing production of the Stradivari-RT enzyme.

[0044] FIG. 4B is a graph showing relative fold change of cDNA produced using the engineered reverse transcriptase as described in Fig. 4A compared to using wild-type reverse transcriptase (SuperScriptll or SSII). Compared to commercially available enzymes, Stradivari- RT had reduced performance. FIG. 5 is a graph showing relative fold change of cDNA produced using the engineered reverse transcriptase as described in Fig. 4A where LNA nucleotides were introduced in the RHS probe, as described in Example 1. The addition of a block reduced Stradivari activity.

[0045] FIG. 6 is a schematic showing the extension of a cDNA with a wild-type reverse transcriptase having strand displacement activity compared to extension of cDNA using an engineered reverse transcriptase with a substantially disrupted or abolished strand displacement activity. The design of RHS and LHS probes is shown. Exemplary positions of a pair of primers, External FW and Probe RV, used to amplify a cDNA extended using a wild-type reverse transcriptase are also shown.

[0046] FIG. 7 is a graph showing relative fold change of cDNA produced using the RHS and LHS probes, reverse transcriptases (SuperScriptll or SSII) and primers as described in Example 1.

[0047] FIG. 8 is a schematic showing the extension of a cDNA with an engineered reverse transcriptase and ligation of the cDNA to the second polynucleotide probe using SplintR ligase. Primers used for amplifying produced cDNA in both reaction conditions are shown. The schematic shows testing if gap-filling and ligation can be used in a single step.

[0048] FIG. 9 is a pair of graphs showing fold change production of cDNA using a reverse transcriptase with strand displacement activity, an engineered reverse transcriptase without strand displacement activity, and simultaneous reverse transcription and ligation. qPCR- amplified cDNA product is graphed for each condition. The data demonstrate that gap-filling and ligation could be incorporated in one single step, that SDV-RT elongation is blocked by the presence of the RHS probe and that the two probes can be ligated after gap filling with high efficiency.

[0049] FIG. 10A is a Uniform Manifold Approximation and Projection (UMAP) plot showing clustering of single-cell RNA-sequencing data of two cell lines, HEL (erythroblast cell line) and CCRF (T lymphoblast cell line), in which two highly expressed genes GAPDH and CTCF were detected as described in Example 2. The RNA projected UMAPs successfully separated the cell lines, demonstrating that Stradivari-RT does not interfere with the standard lOx kit for singlecell RNA-seq.

[0050] FIG. 10B is a pair of graphs showing the number of RNA features and RNA counts measured using single-cell RNA-sequencing of two cell lines, HEL and CCRF, in which two highly expressed genes GAPDH and CTCF were detected as described in Example 2. FIG. 10C is a graph showing the expression of specific cell-line marker genes, identifying the pattern of expression of lymphoid genes in CCRF cells and erythroid genes in HEL cells as described in Example 2.

[0051] FIG. 10D is a UMAP plot showing detected cDNA in single cells using polynucleotide probes with complementarity against target GAPDH transcripts, with a 9-nucleotide third sequence space between the first polynucleotide probe and the second polynucleotide probe. Signal was successfully detected in single cell.

[0052] FIG. 10E is a UMAP plot showing detected cDNA in single cells using polynucleotide probes with complementarity against target CTCF transcripts, with a 9-nucleotide third sequence space between the first polynucleotide probe and the second polynucleotide probe. Signal was successfully detected in single cell with 100% genotyping accuracy as shown by the LOGO sequence plots.

[0053] FIG. 10F is a schematic showing the sequences of the two 9-nucleotide GAPDH and CTCF sequences identified using an engineered reverse transcriptase.

[0054] FIG. 11 is a series of UMAP plots showing the expression levels of specific cell-line marker genes (GATA1, GYPA, CD3D, and CD3G), identifying the pattern of expression of lymphoid genes in CCRF cells and erythroid genes in HEL cells as described in Example 2.

[0055] FIG. 12A is a UMAP plot showing clustering of single-cell RNA-sequencing data of two cell lines, HEP2G (a BRAF WT cell line) and SKML (a BRAFV600E heterozygous mutant cell line) as described in Example 3. It was found that Stradivari can be used for accurate genotyping of cell lines. Prior technologies failed to profile this mutant locus. The RNA projected UMAPs successfully separated the cell lines, demonstrating that Stradivari-RT does not interfere with the standard lOx kit for single-cell RNA-seq.

[0056] FIG. 12B is a series of graphs showing the number of RNA features an RNA counts measured using single-cell RNA-sequencing of two cell lines, HEP2G (WT) and SKML (Het), as described in Example 3. The number of RNA features and RNA counts were concordant with lOx genomics guidelines, underscoring the compatibility of Stradivari-RT with lOx application.

[0057] FIG. 12C is a heat map graph showing the pattern of expression of hepatic and melanoma genes measured using single-cell RNA sequencing of two cell lines, HEP2G (WT) and SKML (Het), as described in Example 3. The majority of HEP2G cells were genotyped as BRAF wild-type (>90%). FIG. 12D is a UMAP plot showing genotyping of HEP2G cells and SKML cells, as described in Example 3. With regards to the SKML line, heterozygous, mutant and wild type cells were detected, likely the result of allelic drop out.

[0058] FIG. 13A is a series of UMAP plots showing detected cDNA in single cells using polynucleotide probes with complementarity against target CD5, MME, FCER2, MS4A1, CD22, ITGAX, IL2RA, and IL3RA transcripts from hairy cell leukemia (BRAF V600E mutant) cells, as described in Example 4.

[0059] FIG. 13B is a UMAP plot showing clustering of single-cell RNA-sequencing data of cells from a patient sample with Hairy cell leukemia cells, in which the BRAFV600E mutation was detected as described in Example 4. The vast majority of BRAFV600E mutant cells fell within the hairy cell leukemia cluster.

[0060] FIG. 13C is a pair of UMAP plots showing two distinct clusters (a high CD38 expression / low FMOD expression cluster and a low CD38 / high FMOD expression cluster) identified using single-cell RNA-sequencing data measured in a patient sample chronic lymphocytic leukemia (CLL) as described in Example 4.

[0061] FIG. 13D is a UMAP plot showing that in the two clusters, the high CD38 expression / low FMOD expression cluster and the low CD38 / high FMOD expression cluster, the majority of BTK mutant cells were found within the low CD38 / high FMOD expression cluster as described in Example 4.

[0062] FIG. 14A is a histological image of a bone marrow tissue sample from a patient with NPM1 mutant acute myeloid leukemia (AML). The histological image was stained with an antibody against the mutant NPM1 protein as described in Example 5.

[0063] FIG. 14B is a probe-based spatial transcriptomic image of a bone marrow tissue sample from a patient with NPM1 mutant acute myeloid leukemia (AML) obtained as described in Example 5. The NPM1 mutational landscape as defined by sequencing was compared to the mutational landscape defined by an antibody against the mutant NPM1 protein.

[0064] TERMINOLOGY

[0065] Unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Generally, nomenclatures utilized in connection with, and techniques of, chemistry, biochemistry, molecular biology, pharmacology and toxicology are those well-known and commonly used in the art. In order for the present invention to be more readily understood, certain terms are defined below. Additional definitions for the following terms and other terms are set forth throughout the specification. Any publications and other reference materials referenced herein to describe the background of the invention and to provide additional detail regarding its practice are hereby incorporated by reference.

[0066] Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.

[0067] In this application, unless otherwise clear from context, (i) the terms “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article; (ii) the term “or” may be understood to mean “and / or”; (iii) the terms “comprising” and "including" may be understood to encompass itemized components or steps whether presented by themselves or together with one or more additional components or steps; and (iv) where ranges are provided, endpoints are included.

[0068] As used herein, the term “about” refers to a range of + / - 10% of the stated value. In some embodiments, the range is + / - 5%, + / - 3%, + / - 2%, + / - 1%, + / - 0.5%, or + / - 0.1% of the stated value.

[0069] As used herein, the term “engineered reverse transcriptase” refers to a reverse transcriptase having an amino acid sequence modified by one or more amino acid substitutions, deletions and / or additions such that its amino acid sequence differs from the amino acid sequence of the corresponding wild-type reverse transcriptase.

[0070] As used herein, the term “substantially,” as it refers to “substantially” disrupting or “substantially” abolishing enzyme activity as described herein, refers to disrupting or abolishing of activity of the enzyme such that the activity is reduced or inhibited by at least 70% and by less than 100%, e.g., by at least 70%, 75%, 80%, 85%, 90%, 98%, or 99%.

[0071] As used herein, the term “locked nucleic acid” or “LNA” refers to a modified RNA nucleotide in which the ribose moiety is modified with a bridge connecting the 2’ oxygen and the 4’ carbon.

[0072] DETAILED DESCRIPTION

[0073] Described herein is a new method for genotyping the transcriptome with single-cell and spatial resolution using commercially available probe-based techniques. The new method utilizes a new engineered reverse transcriptase, and it unexpectedly allows simultaneous identification of single nucleotide variations in the RNA as well as profiling the entire transcriptome in fixed samples. As described herein, the inventors designed and engineered a Reverse-Transcriptase enzyme that can copy RNA sequences from a template strand into a probe set. As described herein, when two non-contiguous probes are bound to RNA, the gap between them can be filled and the two molecules can be subsequently ligated, copying the original RNA sequence into the probe, which is then captured and barcoded by using commercially available probe-based single-cell and spatial products, such as lOx Genomics FFPE Visium or lOx Genomics single-cell Flex kits. Probe-based single-cell and spatial products, such as lOx Genomics FFPE Visium or lOx Genomics single-cell Flex kits, are well known in the art.

[0074] Single-cell and Spatial transcriptomics have emerged as powerful techniques that have revolutionized understanding of cellular dynamics and tissue architecture by providing spatially resolved gene expression data. Recently a technology was introduced that is able to profile gene expression in formalin-fixed samples. The Single Cell Gene Expression Flex Fixed RNA Profiling assay and Spatial Gene Expression for FFPE use probes that target protein-coding genes in the human or mouse transcriptome. Each probe consists of a pair of oligonucleotides that hybridize to the targeted transcript and are subsequently ligated. Briefly, human or mouse whole transcriptome probe panels, consisting of a pair of specific probes for each targeted gene, are added to the tissue or permeabilized fixed single cells. These contiguous probe pairs hybridize to their gene target and are then ligated to one another. Subsequently, the ligated probe pairs bind with spatially barcoded oligonucleotides present on the Capture Area for spatial analysis, or are encapsulated in oil droplets and barcoded for single-cell analysis. All the probes captured by primers on a specific spot or droplet share a common barcode. Libraries are generated from the probes and sequenced and the spatial / cell barcodes are used to associate the reads back to the tissue section images for spatial mapping of gene expression or to a single cell. This powerful tool allows for gene expression profiling in fixed cells and tissues. Since the probe set is sequenced instead of the RNA, this approach can profile the transcriptome in samples where the RNA is highly degraded, such as fixed samples. On the other hand, since the synthetic probes are subject to sequencing, any possible information about the original RNA sequence, including point mutations, small deletions, or duplications is lost. Accurately retrieving this information is essential for evaluating how these variants (for example, cancer driver mutations) affect the cellular phenotype. The invention described herein aims to overcome this limitation. As shown in the Examples, the inventors developed a methodology that, using a gap-filling approach, allows to retain all the benefits of this assay, while adding additional layers of information about the original RNA nucleotide composition. In particular, as shown in the Examples, the inventors used non-contiguous locusspecific probes, separated from each other by a small gap, e.g., from 6 to 20 base pairs. Before the ligation step, the inventors used a reverse-transcriptase enzyme to fill this gap, copying the original RNA sequence that is located between the two probes. After ligation, the two probes together with the newly copied sequence form a single DNA molecule, which can be captured and barcoded by single-cell or spatial assays. In this way, when a mutation is present on the original RNA molecule, it will be incorporated in the probe set and sequenced by NGS, without the need to design mutation- specific probes.

[0075] All commercially available reverse transcriptase (RT) enzymes possess very high strand displacement activity, which is a crucial feature in their function during reverse transcription. This feature represents a major limitation for workflow of this methodology, as the RT enzyme synthesizes the DNA strand and fills the gap in between probes. When the enzyme reaches the second probe, it can displace it, continuing to copy the RNA molecule without stopping, resulting in the misincorporation of the copied sequence into the capturable molecule (see attached overview) and loss of the barcode on one of the probes.

[0076] For this reason, as shown in the Examples, the inventors engineered and produced an RT enzyme lacking this strand displacement activity. The inventors introduced a unique combination of mutations in Moloney Murine Leukemia Virus Reverse Transcriptase (M-MLV RT): the first mutation (Y64A) abolished the strand displacement activity of this enzyme, the second mutation (D524A) abolished the RNase H activity of this enzyme, preventing digestion of DNA:RNA hybrid to preserve target integrity. This engineered enzyme is referenced herein as StraDiVari-RT (Strand Displacement Variant Reverse Transcriptase). This enzyme allows to incorporate a copy of the original RNA sequence into a locus- specific probe set, enabling genotyping of the transcriptome in fixed samples, providing spatial location and single-cell quantification of the mutant allele to study phenotypic changes of known driver mutations with high resolution.

[0077] Compositions of the Invention

[0078] In some aspects, provided herein is any engineered reverse transcriptase in which the strand displacement activity is inhibited, substantially abolished or entirely abolished, and / or in which the RNAse-H activity is inhibited, substantially abolished or entirely abolished. In some embodiments, the reverse transcriptase engineered as described herein is derived from any source, e.g., Moloney Murine Leukemia Virus. Various reverse transcriptases that can be engineered as described herein (to interfere with or abolish the strand displacement activity and / or RNAse-H activity), and used in the methods described herein, are known in the art. In some embodiments, the reverse transcriptase is any natural reverse transcriptase. In some embodiments, the reverse transcriptase is HIV-1 reverse transcriptase, M-MLV reverse transcriptase, AMV reverse transcriptase, or Telomerase reverse transcriptase.

[0079] In some embodiments, provided herein is an engineered reverse transcriptase in which the strand displacement activity is inhibited, substantially abolished or entirely abolished, e.g., by introducing a mutation (e.g., amino acid substitution) in the first finger domain of the reverse transcriptase. In some embodiments, provided herein is an engineered reverse transcriptase in which the RNAse-H activity is inhibited, substantially abolished or entirely abolished, e.g., by introducing a mutation (e.g., amino acid substitution) in the RNAse-H domain of the reverse transcriptase. In some embodiments, a mutation is a substitution or substitution of one or more amino acids (in the finger domain and / or the RNAse-H domain of a reverse transcriptase). In some embodiments, a mutation comprises a deletion or insertion of one or more amino acids (in the finger domain and / or the RNAse-H domain of a reverse transcriptase). In some embodiments, multiple mutations can be introduced (in the finger domain and / or the RNAse-H domain of a reverse transcriptase), e.g., for optimal strand displacement activity and / or interference with RNAse-H activity. In some embodiments, a mutation is a deletion or insertion (in the finger domain and / or the RNAse-H domain of a reverse transcriptase). In some embodiments, a mutation is a partial truncation of the first finger domain and / or the RNAse-H domain of a reverse transcriptase. In some embodiments, a mutation of a reverse transcriptase is a mutation in Y64, e.g., Y64A or another mutation in Y64 that can interfere with the strand displacement activity. In some embodiments, a mutation of a reverse transcriptase is a mutation in D524, e.g., D524A, D524G, D524N or another mutation in D524 that can interfere with the RNAse-H activity. In some embodiments, a mutation of a reverse transcriptase is a mutation in D583, e.g., D583N or another mutation in D583 that can interfere with the RNAse-H activity. In some embodiments, a mutation of a reverse transcriptase is a mutation in E562, e.g., E562Q or another mutation in E562 that can interfere with the RNAse-H activity. In some embodiments, the engineered reverse transcriptase comprises more than one (e.g., 2 or 3) mutations that interfere with the RNAse-H activity (such as 2 or 3 mutations described herein). In some embodiments, provided herein is an engineered reverse transcriptase comprising Y64 mutation in Moloney Murine Leukemia Virus reverse transcriptase, or comprising equivalent mutations(s) in a reverse transcriptase from different species. In some embodiments, provided herein is an engineered reverse transcriptase comprising D524, E562, D583 or partial truncation of the RNAse-H domain in Moloney Murine Leukemia Virus reverse transcriptase, or comprising equivalent mutations(s) in a reverse transcriptase from different species. In some embodiments, provided herein is an engineered reverse transcriptase comprising Y64A mutation and / or D524A mutation in Moloney Murine Leukemia Virus reverse transcriptase, or comprising equivalent mutations(s) in a reverse transcriptase from different species. In some embodiments, provided herein is an engineered reverse transcriptase comprising D524A, D524G, D524N, E562Q, D583N or partial truncation of the RNAse-H domain in Moloney Murine Leukemia Virus reverse transcriptase, or comprising equivalent mutations(s) in a reverse transcriptase from different species. In some embodiments, the engineered reverse transcriptase described herein disrupts the ability to displace the second polynucleotide probe by at least 70%, 75%, 80%, 85%, 90%, 98%, 99% or 100%. In some embodiments, the mutation(s) described herein are in the sequence of M-MLV reverse transcriptase of SEQ ID NO: 1. In some embodiments, the engineered reverse transcriptase comprises, essentially consists of, or consists of SEQ ID NO: 2 or a variant thereof (e.g., a variant having one, two or more mutations described herein). In some embodiments, the engineered reverse transcriptase comprises, essentially consists of, or consists of SEQ ID NO: 2 or a variant thereof (e.g., a variant having one, two or more mutations described herein), and further comprises one or more amino acids before SEQ ID NO: 2 that, e.g., may facilitate protein production (e.g., an amino acid methionine or M).

[0080] In some embodiments, provided herein are compositions comprising an engineered reverse transcriptase described herein.

[0081] In some embodiments, provided herein are nucleic acids encoding the engineered reverse transcriptase described herein.

[0082] Methods of the Invention

[0083] In some aspects, provided here is a method for characterizing, analyzing, genotyping, or profiling the transcriptome (“target RNA”), e.g., mRNA, of a sample comprising: a) contacting the sample with a first polynucleotide probe comprising a nucleotide sequence substantially complementary to a first sequence of the target RNA and a second polynucleotide probe comprising a nucleotide sequence substantially complementary to a second sequence of the target RNA, wherein the first sequence and the second sequence are noncontiguous (e.g., separated by a third sequence comprising at least one nucleotide and, e.g., up to 100 nucleotides); b) contacting the sample with a reverse transcriptase (e.g., any engineered reverse transcriptase described herein), capable of extending the first polynucleotide probe to produce a DNA complementary to the third sequence (cDNA); c) contacting the sample with a ligase capable of ligating the cDNA to the second polynucleotide probe (e.g., wherein steps (b) and (c) can occur simultaneously); d) optionally, amplifying the cDNA, e.g., by real-time quantitative PCR (qPCR); and e) optionally, sequencing the cDNA, e.g., by next generation sequencing (NGS).

[0084] In some aspects, provided here is a method for characterizing, analyzing, genotyping, profiling, detecting, or identifying mutations (e.g., single nucleotide variations) in the transcriptome (“target RNA”), e.g., mRNA, of a sample comprising: a) contacting the sample with a first polynucleotide probe comprising a nucleotide sequence substantially complementary to a first sequence of the target RNA and a second polynucleotide probe comprising a nucleotide sequence substantially complementary to a second sequence of the target RNA, wherein the first sequence and the second sequence are noncontiguous (e.g., separated by a third sequence comprising at least one nucleotide and, e.g., up to 100 nucleotides); b) contacting the sample with a reverse transcriptase (e.g., any engineered reverse transcriptase described herein), capable of extending the first polynucleotide probe to produce a DNA complementary to the third sequence (cDNA); c) contacting the sample with a ligase capable of ligating the cDNA to the second polynucleotide probe (wherein steps (b) and (c) can occur simultaneously); d) optionally, amplifying the cDNA, e.g., by real-time quantitative PCR (qPCR); and e) optionally, sequencing the cDNA, e.g., by next generation sequencing (NGS).

[0085] In some embodiments, the ligase is any ligase known in the art that is capable of ligating DNA or cDNA. In some embodiments, the ligase is a DNA ligase capable of ligating singlestranded DNA molecules (e.g., wherein a single-stranded DNA molecule is annealed to an RNA molecule). In some embodiments, the ligase is a SplintR Ligase, also known as PBCV- 1 DNA Ligase or Chlorella virus DNA Ligase. For disclosures relevant to ligases that can be used in the methods described herein, see, e.g., Lohman et al., 2013, Nucleic Acid Research 42(3): 1831-1844, which is incorporated by reference herein in its entirety. In some embodiments, in step c), any means for ligating the cDNA to the second polynucleotide probe can be used.

[0086] In some embodiments, the ligase is a 10X Genomics kit ligase. In some embodiments, the first and second sequence are separated by 1 to 100 nucleotides. In some embodiments, the first and second sequence are separated by 3 to 40 nucleotides, or 6 to 20 nucleotides. In some embodiments, the first and second sequence are separated by 6 to 75 nucleotides, or 8 to 50 nucleotides. In some embodiments, the first and second sequence are separated by 8 to 40 nucleotides, or 9 to 20 nucleotides (e.g., 8, 9, 10, 11, 12, 13, 14 or 15 nucleotides). In some embodiments, these nucleotides separating first and second sequences represent a gap between the right-hand side (RHS) and left-hand side (LHS) probe used in the methods described herein.

[0087] It is known in the art how to design suitable polynucleotide probes using known techniques, and polynucleotide probe design methodology is also described herein (see, e.g., the examples and figures). In some embodiments, the polynucleotide probes are designed in accordance with 10X Genomics guidelines or design criteria for Visium spatial gene expression and / or chromium single cell gene expression flex. See, e.g., https: / / www.10xgenomics.com / support / cytassist-spatial-gene- expression / documentation / steps / experimental-design-and-planning / custom-probe-design-for- visium-spatial-gene-expression-and-chromium-single-cell-gene-expression-flex.

[0088] In some embodiments, the first polynucleotide probe comprises a nucleotide sequence substantially complementary to a first sequence of the target RNA and a second polynucleotide probe comprises a nucleotide sequence substantially complementary to a second sequence of the target RNA, wherein the nucleotide sequence of one or each probe that is substantially complementary to the target RNA is about 15 to about 30 nucleotides in length, e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides, or any number of nucleotides in between any of these values. In some embodiments, “substantially complementary” is at least or more than 80%, 85%, 90%, 95% or 100% complementary. In some embodiments, “substantially complementary” allows for up to 1, 2, 3 or 4 mismatches (e.g., up to 2 mismatches). In some embodiments, “substantially complementary” is entirely complementary such that there are no mismatches between the part of the polynucleotide sequence that recognizes the target RNA and the target RNA.

[0089] In some embodiments, the first polynucleotide and / or the second polynucleotide probe comprises one or more (e.g., 2, 3, 4 or 5 or more) non-natural nucleotides, e.g., locked nucleic acid (LNA), Super T (5-hydroxybutynl-2’-deoxyuridine), and / or Super G (8-aza-7- deazaguanosine). Such non-natural nucleotides may be used to further impair RT strand displacement activity. In some embodiments, the first polynucleotide and / or the second polynucleotide probe comprises 1, 2, 3, 4 or 5 or more locked nucleic acids (LNAs), optionally wherein the second polynucleotide probe comprises 1, 2, 3, 4, 5 or more LNAs.

[0090] In some embodiments, the first polynucleotide probe is referenced as the LHS probe. In some embodiments, the second polynucleotide probe is referenced as the RHS probe.

[0091] In some embodiments, locked nucleic acid (LNA) nucleotides are to be used in the first 1, 2, 3 or 4 positions (e.g., in the first three positions) of the RHS probe.

[0092] In some embodiments, the 3’ most nucleotide of the gap between the probes is a T. In some embodiments, the 3’ most nucleotide of the gap between the first sequence and the second sequence of the target RNA is a T.

[0093] In some embodiments, the RHS probe is 5’ phosphorylated.

[0094] In some embodiments, the GC content of one or each of the polynucleotide probes to be used in the methods described herein is from 44% to 72%. In some embodiments, the GC content of the RHS polynucleotide probe is at least 44%, less than 72%, or from 44% to 72%. In some embodiments, the GC content of the LHS polynucleotide probe is at least 44%, less than 72%, or from 44% to 72%. In some embodiments, the GC content of the first polynucleotide probe is at least 44%, less than 72%, or from 44% to 72%. In some embodiments, the GC content of the second polynucleotide probe is at least 44%, less than 72%, or from 44% to 72%.

[0095] In some embodiments, the sample is permeabilized before contacting the sample with the first and / or second polynucleotide probes. In some embodiments, the sample is a fixed sample (e.g., formalin-fixed and paraffin-embedded (FFPE)).

[0096] In some embodiments, the method described herein does not comprise using a third polynucleotide probe to gap-fill the sequence between first polynucleotide probe and the second polynucleotide probe. In particular, the methods described herein use a reverse transcriptase (such as the engineered reverse transcriptase described herein) to serve the gap-filling function.

[0097] In some embodiments, the first polynucleotide probe and the second polynucleotide probe can be any probes described herein or known in the art for use for similar purposes (such as single-cell and spatial transcriptomics). In some embodiments, the probes can be barcoded or can be bound to a barcoded oligonucleotide (e.g., spatially barcoded).

[0098] In some embodiments, the probes that can be used for single-cell and / or spatial transcriptomics are those made in accordance with the methods known in the art by, e.g., 10X Genomics. For disclosures relevant to the single-cell and / or spatial transcriptomics and / or probes that can be used in the methods described herein, see, e.g., Williams et al., 2022, Genome Medicine 14(68), which is incorporated by reference herein in its entirety. In some embodiments, the methods described herein are used for genotyping a transcriptome comprising the target RNA. In some embodiments, the methods described herein are used for sequencing a transcriptome comprising the target RNA. In some embodiments, the methods describing herein is for genotyping and / or sequencing the entire transcriptome of a single cell. In some embodiments, the methods described herein are used for single cell transcriptome analysis such as single cell CRISPR screening, single cell T-cell receptor (TCR) sequencing, and / or single cel B-cell receptor (BCR) sequencing. In some embodiments, the methods described herein are for identifying a mutation, a single nucleotide variation, a splice isoform, and / or a TCR / BCR junction in the target RNA (e.g., on a single cell level). In some embodiments, the methods described herein are for identifying a mutation such as a single nucleotide variation in the target RNA (e.g., on a single cell level). In some embodiments, the methods described herein provide spatial location of the target RNA and / or mutation, single nucleotide variation, splice isoform or TCR / BCR junction (e.g. on a single cell level).

[0099] EXAMPLES

[0100] Example 1: An engineered reverse transcriptase described herein (Stradivari-RT) enables gap-filling reactions of non-contiguous probes

[0101] Single-cell and spatial transcriptomics have emerged as powerful tools for dissecting spatial cellular heterogeneity. Human disease-relevant tissues are commonly preserved as formalin-fixed and paraffin-embedded (FFPE) blocks, representing an important resource of abundant human tissue specimens. The capability to explore RNA biology in FFPE tissues holds transformative potential for human biological research and clinical histopathology. Many cancers and other somatic diseases are caused by mutations in key genes that dramatically alter cellular phenotype in emerging subsets of mutant clonal cells, highlighting a need for genotype- aware single-cell and spatial analysis methods that can be applied directly to widely available FFPE samples. To address this gap, commercially available probe-based droplet and spatial workflows were adapted to selectively profile point mutations, deletions, or duplications in archival tissues for combined genotyping and phenotypic profiling. The inventors hypothesized that by gap-filling between two non-contiguous probes targeting RNA molecules around a specific mutation, they could copy the original RNA sequence into the probe set for downstream genotyping. To this end, M-MuLV reverse transcriptase was engineered, knocking out its strand displacement function to enable a gap-filling reaction without loss of probe binding (FIG. 2). Such engineered Strand displacement variant Reverse Transcriptase (Stradivari-RT) was used to profile NPM1 mutation in paraformaldehyde (PFA)-fixed cell lines and AML primary sample, resolving the admixture of wild-type and mutant with single-cell resolution. Stradivari-RT was also validated in spatial transcriptomic applications, localizing wild-type and NPM1 mutant cells in bone marrow sections of patients affected by (disease).

[0102] Applied to preserved somatic and tumor tissue, Genotyping of Transcriptomes (GoT) by Stradivari turns the admixture of mutant and wild-type cells from a limitation to an advantage, enabling the direct comparison of the molecular or spatial features of mutant versus wild-type cells in archival samples, within the same individual, overcoming patient-specific confounders in human studies.

[0103] Formalin-fixed paraffin-embedded (FFPE) tissues are crucial for clinical practice, as they form the basis of human disease histopathological diagnoses. This method is used to preserve surgical pathology samples and is conventionally preferred over fresh frozen specimens due to its economic benefits, including lower storage, space, and personnel costs. FFPE processing maintains the tissue's morphology and cellular integrity at room temperature, making it an ideal method for preservation. Clinical pathology researchers have amassed vast collections of FFPE blocks over time, creating a rich yet underutilized compendium of materials that, when accompanied by comprehensive clinical data, stand as a treasure trove for human biology and translational research.

[0104] However, FFPE specimens present challenges when it comes to analyzing RNA. During the paraffin-embedding process, the RNA within these samples is highly fragmented, with susceptibility to further degradation under suboptimal storage conditions. The loss of poly- A tails adds another layer of complexity, which limits the usefulness of oligo-dT primed reverse transcription. While imaging-based platforms like MERSCOPE (Vizgen), CosMx (Nanostring), and Xenium (lOx Genomics) have demonstrated spatial mapping of hundreds or thousands of genes in unfixed samples, they are constrained by their reliance on a targeted panel for gene expression analysis with limited discovery power to profile diverse RNA species or unknown sequences. Similarly, the Visium (lOx Genomics) chemistry for FFPE samples also relies on a predefined panel to target and capture RNA fragments, enabling near transcriptome-level measurement of gene expression, yet still confined to quantifying the expression of known protein-coding genes rather than base-by-base sequencing of RNA.

[0105] The Single Cell Gene Expression Flex Fixed RNA Profiling assay and Spatial Gene Expression for FFPE use probes that target protein-coding genes in the human or mouse transcriptome. Briefly, human or mouse whole transcriptome probe panels, consisting of a pair of specific probes for each targeted gene, are added to the tissue or permeabilized fixed single cells. These contiguous probe pairs hybridize to their gene target and are then ligated to one another. Subsequently, the ligated probe pairs bind with spatially barcoded oligonucleotides present on the Capture Area for spatial analysis, or are encapsulated in oil droplets and barcoded for single-cell analysis. All the probes captured by primers on a specific spot or droplet share a common barcode. Libraries are generated from the probes and sequenced, and the spatial / cell barcodes are used to associate the reads back to the tissue section images for spatial mapping of gene expression or to a single cell. This powerful tool allows for gene expression profiling in fixed cells and tissues. Since the probe set is sequenced instead of the RNA, this approach can profile the transcriptome in samples where the RNA is highly degraded, such as fixed samples. However, since it is the synthetic probes themselves that are subject to sequencing, any possible information about the original RNA sequence, including point mutations, small deletions, or duplications is lost. Accurately retrieving this sequencing information is essential for evaluating how these variants (for example, cancer driver mutations) affect the cellular phenotype.

[0106] Here, the inventors hypothesized that using a gap-filling approach could allow to retain all the benefits of this assay, while adding additional layers of information about the original RNA nucleotide composition.

[0107] The approach described herein uses non-contiguous locus-specific probes, separated from each other by a small gap, from 6 to 20 base pairs. Before the ligation step, a reversetranscriptase enzyme is used to fill this gap, copying the original RNA sequence that is located between the two probes. After ligation, the two probes together with the newly copied sequence form a single DNA molecule, which can be captured and barcoded by single-cell or spatial assays. In this way, when a mutation is present on the original RNA molecule, it will be incorporated in the probe set and sequenced by next generation sequencing (NGS).

[0108] To achieve this, the inventors engineered and produced a reverse transcriptase (RT) enzyme lacking the strand displacement activity. This was achieved by introduced a unique combination of mutations in Moloney Murine Leukemia Virus Reverse Transcriptase (M-MuLV RT): the first mutation (Y64A) abolished the strand displacement activity of this enzyme, the second mutation (D524A) abolished the RNase H activity of this enzyme, preventing digestion of DNA:RNA hybrid to preserve target integrity.

[0109] The Strand displacement variant Reverse Transcriptase (Stradivari-RT) engineered as described herein allowed to incorporate a copy of the original RNA sequence into a locusspecific probe set, enabling genotyping of the transcriptome in fixed samples, providing spatial location and single-cell quantification of the mutant allele to study phenotypic changes of known driver mutations with high resolution.

[0110] A mutant version of M-MLV reverse transcriptase was engineered and produced by introducing two point mutations to knock out 2 specific functions of the enzyme. The Y64A mutation was introduced in the first finger domain of the protein to delete strand displacement activity, while D524A was introduced in the RNAse-H domain of the enzyme to avoid RNA degradation during reverse transcription (FIG. 1A). The inventors hypothesized that these mutations would allow to fill the gap between non-contiguous probes, with the aim of copying a somatic mutation present in the RNA into a single- strand DNA molecule, which can be subsequently captured and barcoded by commercially available probe-based single-cell and spatial transcriptomic workflows (FIG. IB). To verify that the point mutations introduced are not significantly affecting enzyme processivity, Stradivari-RT was tested in a targeted cDNA experiment. Total RNA was obtained from K562 cells and GAPDH- specific reverse transcription was performed with increasing concentrations of Stradivari-RT or SuperScriptll (wild-type control) (FIG. 1C, FIG. 4A and FIG. 4B). The efficiency of the reaction was measured by real-time quantitative PCR with a set of primers designed on the GAPDH transcript (FIG. 3). It was observed that Stradivari-RT has slightly lower performance compared to the commercially available enzyme; at the highest concentration tested, a 50% reduction in GAPDH cDNA level was observed for Stradivari-RT most likely due to its inability to resolve the secondary structure of the RNA (FIG. 4A and FIG. 4B). Since the goal is to amplify small regions around somatic variants and not to obtain whole transcript cDNA, the inventors hypothesized that, although reduced, Stradivari-RT activity was sufficient to fill short gaps in between probes. Next the strand displacement properties of Stradivari-RT was tested. Total RNA from K562 was pre-incubated with 2 non-contiguous probes. The right-hand side and lefthand side probes are referenced as RHS and LHS, respectively. Both RHS and LHS were designed to be the reverse complement of the target gene and to be separated by a gap of 13 nucleotides (FIG. 6). The LHS probe serves as a primer to initiate the reverse transcription, and after gap filling the reverse transcriptase can either displace the RHS probe, elongating the cDNA molecule passing through it, or, with effective knockout of the strand displacement function, stop as soon it reaches the 5’ phosphate nucleotide of the RHS probe and eventually lose the binding with the newly synthesized cDNA (FIG. 1C). cDNA synthesis was performed with a LHS-like probe targeting the NPM1 transcript, in the presence or absence of the blocking RHS probe. The reaction was performed with Stradivari-RT or SuperScript II, as a negative control. To measure the strand displacement capacity of both enzymes, after the reverse transcription step to generate cDNA, a real-time quantitative PCR was performed with a set of primers designed to amplify a downstream region relative to the RHS probe. It was hypothesized that an effective blocking of the cDNA elongation will result in the absence of qPCR amplification. As expected, SusperScript II showed a very strong strand displacement activity, where the presence or absence of RHS probes resulted in a comparable level of cDNA product measured by qPCR. In contrast, Stradivari-RT resulted in significantly less product, demonstrating that the introduced Y64A mutation effectively disrupted the ability to displace the second probe, resulting in 80% less product in the presence of the blocking RHS probe (FIG. 5 and FIG. 6). To further reduce this activity, inventors introduced locked nucleic acid (LNA) nucleotides in the first 3 positions of the RHS probe, in order to create a more stable DNA / RNA hybrid molecule. By repeating the experiment described above, with the LNA probe configuration, 99% of cDNA elongation was blocked using Stradivari-RT (FIG. 5), demonstrating that a gap-filling approach could potentially be used for single-cell and spatial transcriptomic probe-based protocols. Inventors next tested if gap-filling and ligation could be incorporated in one single step (FIG. 8). Two non-contiguous probes targeting GAPDH mRNA, separated by 9 nucleotides, were pre-annealed to total RNA obtained from K562 cells. After annealing, Stradivari-RT or Superscript!! plus SplintR ligase were added to initiate reverse transcription and simultaneously ligate the two probes after the gap-filling reaction. Reverse transcription blocking efficiency was tested as described above, while ligation efficiency was tested by qPCR using primers targeting both probe handles. Inventors confirmed that Stradivari- RT elongation is blocked by the presence of the second probe (FIG. 9) and, more importantly, that the two probes can be ligated after gap filling with high efficiency (FIG. 9).

[0111] Example 2: Single-cell probes gap-filling technology using an engineered reverse transcriptase (Stradivari-RT)

[0112] Next, the compatibility of Stradivari-RT with downstream single-cell applications was tested. A cell line mixing experiment was performed with HEL (erythroblast cell line) and CCRF (T lymphoblast cell line) cells. Cells were fixed and permeabilized, and permeabilized cells were incubated for 24 hours with lOx Genomics human whole-transcriptome probe set supplemented with custom designed probe pairs targeting two highly expressed genes, GAPDH and CTCF. Unlike lOx Genomics gene-specific probe pairs that are designed to be contiguous on the matched mRNA, for this experiment custom probe pairs were designed with a 9- nucleotide gap between the LHS and RHS. After incubation, the excess of unbounded probes was washed out and the cells were encapsulated in lOx Genomics barcoding mix supplemented with Stradivari-RT enzyme to allow in-droplet simultaneous gap-filling reaction and ligation. Finally, NGS libraries of barcoded ligated probe pairs were prepared.

[0113] The RNA projected UMAPs successfully separated the cell lines, demonstrating that Stradivari-RT does not interfere with the standard lOx kit for single-cell RNA-seq (FIG. 10A). Notably, the number of RNA features and RNA counts were concordant with lOx genomics guidelines, underscoring the compatibility of Stradivari-RT with lOx application (FIG. 10B). As a further quality control metric, the expression of specific cell-line marker genes was examined, identifying the expected pattern of high expression of lymphoid genes in CCRF cells and erythroid genes in HEL cells (FIG. IOC and FIG. 11). To test gap-filling and amplification at specific loci, probes were designed against GAPDH or CTCF transcripts, with a 9-nucleotide space between RHS and LHS probes. As a result, a signal was successfully detected in single cells (FIGs. 10D and 10E) and amplified the gap sequence with 100% genotyping accuracy (FIG. 10F). Collectively, these results demonstrate that Stradivari-RT is compatible with lOx single-cell RNA sequencing kits and can accurately genotype gap sequences between probes with high accuracy, at the single-cell level.

[0114] Conclusions

[0115] These results demonstrate the feasibility of the gap-filling technology described herein. It was found that Stradivari-RT enables the characterization of endogenous RNA nucleotide content in fixed samples using commercially available probe-based methods with very minimal modification of the existing workflow.

[0116] Cloning of StraDiVari-RT plasmid constructs

[0117] The DNA sequence coding for M-MLV reverse transcriptase harboring Y64A and D524A mutation was synthesized as a gene fragment (IDT) flanked by restriction enzyme sites Xbal and Spel. pTXBl plasmid and gene fragments were digested with Xbal (NEB, R0145S) and Spel (NEB, R3133S) Ih at 37°C. The digested plasmid backbone was purified with gel extraction (QIAquick Gel Extraction Kit, #28704), and dephosphorylated with Shrimp Alkaline Phosphatase (NEB, M0371S) at 37°C for 30 min. The digested gene fragments were purified with columns (QIAquick PCR Purification Kit, #28106). The purified gene fragments and plasmid backbones were ligated with quick ligase (NEB, M2200) at room temperature for 30 min and subsequently transformed into competent cells per vendor instruction (NEB 5-alpha, NEB C2987H). StraDiVari-RT production

[0118] The pTXBl -StraDiVari-RT vector was transformed into enhanced BL21 -derivative competent Escherichia coli cells (T7 Express lysY / Iq, NEB, C3013I), and StraDiVari-RT was produced via intein purification with an affinity chitin-binding tag (24). Briefly, 400 mL of Luria broth (LB) culture supplemented with ampicillin was grown at 37 °C to optical density (OD600) = 0.6. StraDiVari-RT expression was then induced with isopropyl-P-D-1- thiogalactopyranoside (IPTG) 0.5 mM at 22°C for 24 hours in the presence of 4ml / L of catabolite repression buffer (25% Glycerol, 25% Glucose, 1 mM MgC12, 0.1 mM MnC12). After induction, cells were pelleted and then frozen at -80°C overnight. Cells were then lysed by sonication in 30 mL pf HEGX (20 mM HEPES-KOH pH 7.5, 0.8 M NaCl, 1 mM EDTA, 10% glycerol, 0.2% Triton X-100) with a protease inhibitor cocktail (Roche, no. 04693132001). The lysate was pelleted at 10,000g for 20 min at 4°C. The supernatant was loaded on four 2 mL chitin columns (NEB, no. S665 IS). The columns were washed with 20 mL of HEGX, then 3 mL of HEGX containing 100 mM DTT was added to the column, followed by incubation for 48 h at 4°C to allow for cleavage of StraDiVari-RT from the intein tag. StraDiVari-RT was eluted directly into 30 kDa molecular- weight cutoff (MWCO) spin columns (Millipore, no. UFC903008) by the addition of 2 mL of HEGX. Protein was dialyzed in five dialysis steps using 15 mL of 2x dialysis buffer (100 HEPES-KOH pH 7.2, 0.2 M NaCl, 0.2 mM EDTA, 2 mM DTT, 20% glycerol) and concentrated to 1 mL by centrifugation at 5,000g. The protein concentrate was transferred to a new tube and mixed with an equal volume of 100% glycerol. nb-Tn5 aliquots were stored at -80°C.

[0119] Cell culture

[0120] K562 cells were acquired from the American Type Culture Collection (CCL-243). OCI- AML3 cells were donated. OCI-AML3 cells were maintained at 37 °C and 5% CO2 in alpha- MEM medium (catalog number) supplemented with 10% FBS (Thermo Fisher Scientific, 16000044). K562 cells were maintained at 37C and 5% CO2 in R10 medium (RPMI with stabilized L-glutamine (Thermo Fisher Scientific, 11875119) supplemented with 10% FBS.

[0121] Example 3: Genotyping single cells using an engineered reverse transcriptase (Stradivari- RT) This example shows that Stradivari-RT can accurately genotype single cells. Two cell lines were used: 1) HEP2G, a BRAF WT cell line and 2) SKML, a BRAFV600E heterozygous mutant cell line. Prior technologies developed in the lab failed to profile this mutant locus. The RNA projected UMAPs successfully separated the cell lines, demonstrating that Stradivari-RT does not interfere with the standard lOx kit for single-cell RNA-seq (FIG. 12A). Notably, the number of RNA features and RNA counts were concordant with lOx genomics guidelines, underscoring the compatibility of Stradivari-RT with lOx application (FIG. 12B). The expression of specific cell-line marker genes was examined, identifying a pattern of expression of hepatic and melanoma genes (FIG. 12C). The majority of HEP2G cells were genotyped as BRAF wild-type (>90%) (FIG. 12D and Table 1). With regards to the SKML line, heterozygous, mutant, and wild type cells were detected, likely the result of allelic drop out (FIG. 12D and Table 1).

[0122] Table 1. Genotyping of SKML and HEP2G cells

[0123] Custom Probe Design

[0124] Custom targeted probes were designed so that there was a 9-20 base gap between the right-hand side (RHS) and left-hand side (LHS) probe. Per lOx recommendations for custom probe design, a GC content between 44 - 72% for each probe half was achieved. Gap size and location was adjusted so that when possible the 3' most nucleotide of the gap would be a T. The handle for the RHS probe was constructed depending on whether the intended use was for 10X FLEX single cell or 10X Visium spatial protocol. For the LHS probe, the Nextera read 2 sequence was used for the handle.

[0125] Single-cell fixed cell lines

[0126] Cells were processed according to the lOx protocol for the fixation of cells and nuclei. Briefly, suspensions of up to 10 million cells were spun-down at 400 ref for 5 min at 4C and cells were resuspended in 1 mL of Fixation Buffer (4% Formaldehyde, IX Cone. Fix and Perm Buffer - lOx Genomics PN-2000517). Cells were fixed for 30 min at room temperature. To stop the fixation, cells were spun-down at 850 ref for 5 min at room temperature and quenched with 1 mL of Quenching Buffer (IX Cone. Quench Buffer - lOx Genomics PN-2000516).

[0127] Up to 2 million cells (50% SK-MEL-239 and 50% HEP2G) were processed for hybridization following lOx recommendations. Hybridizations were set up in 80 pl of hybridization mix with 20 pl of Human WTA probes (lOx Genomics PN-2000510 or PN- 2000718). To use custom probes, a spike-in pool containing 40 nM of each LHS and 80 nM of each RHS probe in nuclease-free water was prepared. 5 pl of custom probe mix was added to the hybridization mix. Hybridizations were performed at 42C for 16-24h. After hybridization samples were washed 3 times in Post-Hyb Wash Buffer for 10 min at 42C. After the washes cells were resuspended in Post-Hyb Resuspension Buffer, filtered through a Miltenyi Biotec 30 um filter and measured with the cell counter to determine the amount needed for the Chromium X run.

[0128] For the GEM encapsulation the lOx Genomics protocol and guidelines were followed regarding the volume of cells and reagents required per well according to the targeted cell recovery. lOpl of StraDiVari enzyme was added and the volume of the Post-Hyb Resuspension Buffer was reduced by the same amount in the cell-reagent mix. For the mixing study, 20,000 cells per sample were targeted. After loading the Chip Q and running it on the Chromium X, GEMs were recovered as indicated by lOx Genomics. Reverse transcription and ligation were performed at 25 °C for 90 minutes, followed by extension at 60°C for 45 minutes and enzyme deactivation at 80°C for 20 minutes. After processing the GEMs, the product is pre-amplified; in addition to the Pre-Amp Primers B (lOx Genomics PN-2000529) 5pl of lOpM Nextera partial read 2 primer was added to amplify the targeted library. Sample indexing of gene expression library was performed per lOx protocol. For indexing of the targeted libraries, an additional 15 cycles of pre-amplification PCR was performed using a TruSeq partial R1 primer and a Nextera partial R2 primer followed by 5-10 cycles of indexing using a N7 NY sample indexing primer as well as an indexed P5-TruSeq R1 primer.

[0129] Example 4: Studying the effects of somatic mutations in human disease using an engineered reverse transcriptase (Stradivari-RT)

[0130] This example shows that Stradivari-RT can be utilized to study the effects of somatic mutations in human disease. Two patient samples were processed: 1) Hairy cell leukemia sample (BRAF V600E mutant) (FIGs. 13A-13B) and 2) chronic lymphocytic leukemia (CLL) with subclonal BTK mutation (FIGs. 13C-13D). The cluster of hairy cell leukemia cells (FIG. 13 A) were able to be clearly identified. The vast majority of BRAFV600E mutant cells fell within the hairy cell leukemia cluster (FIG. 13B). In the CLL sample, two distinct clusters of CLL cells were identified based on RNA expression. The majority of BTK mutant cells, from the CLL sample, were found within the CD381oFMODhi cluster.

[0131] Example 5: Determination of mutational landscape using an engineered reverse transcriptase (Stradivari-RT) and a spatial transcriptomic platform

[0132] This example shows that StraDiVari is compatible with probe-based spatial transcriptomic platforms. A bone marrow sample from a patient with NPM1 mutant acute myeloid leukemia (AML) was processed by combining Stradivari with lOx Visium spatial transcriptomic platform (FIGs. 14A and 14B). The NPM1 mutational landscape as defined by sequencing (FIG. 14B) was compared to the mutational landscape defined by an antibody against the mutant NPM1 protein (FIG. 14A).

[0133] Spatial Stradivari

[0134] Targeted and gene expression cDNA libraries were prepared following the guidelines outlined in the Visium CytAssist Spatial Gene Expression for FFPE User Guide with the following modifications. FFPE tissue sections of 5 pm thickness were mounted on Surgipath Apex Superior Adhesive Slides (Leica™). Sections underwent deparaffinization followed by H&E staining. Following imaging, sections were processed for hematoxylin de-staining and decrosslinking. Probe hybridization was set up in 140.1 pL of FFPE Hyb Buffer (lOx Genomics PN 2000423), 20 pL of Human WT Probes v2 - RHS (lOx Genomics PN - 2000657), 20 pL of Human WT Probes v2 - LHS (lOx Genomics PN - 2000658), and 20 pL of custom probe. For the custom probes, a spike-in pool containing 25 nM of each LHS and 50 nM of each RHS probe in nuclease-free water was prepared. 20pL of custom probe mix was added to the hybridization mix. Hybridization was performed at 50°C for 16-24h. After hybridization, reverse transcription was performed using the StraDiVari enzyme at 30°C for 60 minutes. This was followed by probe ligation at 37 °C for 60 minutes. Probes were transferred from the glass slide to the Visium CytAssist Spatial Gene Expression slide using the Visium CytAssist instrument (lOx Genomics). Probe extension and release were carried out according to the standard Visium FFPE workflow. During pre-amplification, a partial Nextera read 2 primer was spiked into the reaction in addition to the TS Primer Mix B (lOx Genomics PN 2000537). Sample indexing of gene expression library was performed per lOx protocol. For indexing of the targeted libraries, an additional 15 cycles of pre-amplification PCR was performed using a TruSeq partial R1 primer and a Nextera partial R2 primer followed by 5-10 cycles of indexing using a N7 NY sample indexing primer as well as an indexed P5-TruSeq R1 primer.

[0135] Sequencing

[0136] Libraries were sent for sequencing to the Weill Cornell Medical School Genomics Resources Core Facility with either the Illumina Nova-Seq 6000 or NextSeq 2000 sequencing platform using paired-end dual-indexing (28 cycles for Read 1, 10 cycles for i7 index, 10 cycles for i5 index, and 90 cycles for Read 2).

[0137] SEQUENCES

[0138] INCORPORATION BY REFERENCE

[0139] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference.

[0140] EQUIVALENTS

[0141] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the following claims.

Claims

What is claimed is:

1. An engineered reverse transcriptase, wherein a wild-type reverse transcriptase is engineered to comprise: a) a first mutation that interferes with or substantially abolishes the strand displacement activity, wherein the first mutation is one or more substitutions, deletions or insertions of amino acids; and / or b) a second mutation that interferes with or substantially abolishes RNAse-H activity, wherein the second mutation is one or more substitutions, deletions or insertions of amino acids.

2. The engineered reverse transcriptase of claim 1 , wherein the wild-type reverse transcriptase is HIV-1 reverse transcriptase, M-MLV reverse transcriptase, AMV reverse transcriptase, or Telomerase reverse transcriptase, or any other natural reverse transcriptase.

3. The engineered reverse transcriptase of claim 1, wherein the wild- type reverse transcriptase is Moloney Murine Leukemia Virus Reverse Transcriptase and / or comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identical to SEQ ID NO: 1.

4. The engineered reverse transcriptase of any one of claims 1-3, wherein: a) the first mutation is in the first finger domain of the wild-type reverse transcriptase; and b) the second mutation is in the RNAse-H domain of the wild-type reverse transcriptase; optionally wherein the first finger domain comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identical to SEQ ID NO: 3 and the RNAse-H domain comprises an amino acid sequence at least 80%, 85%, 90%, 95%, 98%, 99% or 100% identical to SEQ ID NO: 4.

5. The engineered reverse transcriptase of any one of claims 1-4, wherein the first mutation is in or corresponds to amino acid Y 64 and / or the second mutation is in or corresponds to amino acid D524, D583 or E562 relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1.

6. The engineered reverse transcriptase of any one of claims 1-5, wherein the first mutation is or corresponds to Y64A and / or the second mutation is or corresponds to D524A, D524G, D583N, D524N, or E562Q relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1.

7. The engineered reverse transcriptase of any one of claims 1-5, wherein the first mutation is or corresponds to Y 64A (tyrosine-to-alanine at position 64) and / or the second mutation is or corresponds to D524A (aspartic acid-to-alanine at position 524) relative to the amino acid sequence numbering of Moloney Murine Leukemia Virus Reverse Transcriptase or SEQ ID NO: 1.

8. The engineered reverse transcriptase of any one of claims 1-7, wherein the engineered reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 2.

9. A composition or kit comprising the engineered reverse transcriptase of any one of claims 1- 8.

10. A nucleic acid encoding the engineered reverse transcriptase of any one of claim 1-8.

11. A method for analyzing a target ribonucleic acid sequence (RNA) in a sample comprising: a) contacting the sample with a first polynucleotide probe comprising a nucleotide sequence substantially complementary to a first sequence of the target RNA and a second polynucleotide probe comprising a nucleotide sequence substantially complementary to a second sequence of the target RNA, wherein the first sequence and the second sequence are separated by a third sequence comprising at least one nucleotide; b) contacting the sample with the engineered reverse transcriptase of any one of claims 1-8, whereby the first polynucleotide probe is extended to produce a DNA complementary to the third sequence (cDNA); and c) conducting a ligation reaction to ligate the cDNA to the second polynucleotide probe.

12. The method of claim 11, wherein the third sequence comprises about 3 to 40 nucleotides.

13. The method of claim 11, wherein the third sequence comprises 6 to 20 nucleotides.

14. The method of any one of claims 11-13, wherein the target RNA is mRNA.

15. The method of any one of claims 11-14, wherein the first polynucleotide and / or the second polynucleotide probe comprises 1, 2, 3 or more non-natural nucleotides, optionally wherein the nonnatural nucleotides are selected from one or more of: locked nucleic acid (LNA), Super T (5- hydroxybutynl-2’-deoxyuridine), and Super G (8-aza-7-deazaguanosine).

16. The method of any one of claims 11-14, wherein the first polynucleotide and / or the second polynucleotide probe comprises 1, 2, 3 or more locked nucleic acids (LNAs), optionally wherein the second polynucleotide probe comprises at least one locked nucleic acid.

17. The method of any one of claims 11-16, wherein step b) and step c) occur simultaneously by contacting the sample with the engineered reverse transcriptase of any one of claims 1-8 and a ligase at the same time.

18. The method of any one of claims 11-17, which comprises, after step c), a step of binding the first polynucleotide probe and the second polynucleotide probe to a barcoded oligonucleotide, optionally wherein the barcoded oligonucleotide is a spatially barcoded oligonucleotide.

19. The method of any one of claims 11-18, comprising, after step c), amplifying the cDNA, optionally wherein the amplifying is performed by real-time quantitative PCR.

20. The method of claim 19, further comprising sequencing the cDNA, optionally wherein the sequencing is next generation sequencing (NGS).

21. The method of any one of claims 11-20, comprising permeabilizing the sample before contacting the sample with the first and second polynucleotide probes.

22. The method of any one of claims 11-21, wherein the sample is a fixed sample.

23. The method of claim 22, wherein the fixed sample is formalin-fixed and paraffin-embedded (FFPE).

24. The method of any one of claims 11-23, wherein the sample is a single cell, and / or the target RNA is the target RNA of a single cell.

25. The method of any one of claims 11-23, wherein the sample is a tissue.

26. The method of any one of claims 11-25, wherein the target RNA is the transcriptome of a sample.

27. The method of any one of claims 11-26, wherein the method is for sequencing and / or genotyping a transcriptome comprising the target RNA.

28. The method of any one of claims 11-27, wherein the method provides spatial location of the target RNA.

29. The method of any one of claims 11-26, wherein the method is for assessing, detecting the presence of or identifying a mutant allele, a single nucleotide variation, a splice isoform, or a TCR / BCR junction in the target RNA.

30. The method of claim 29, wherein the method is for providing spatial location and / or single cell quantification of the mutant allele, the single nucleotide variation, the splice isoform or the TCR / BCR junction.

31. The method of any one of claims 11-26, wherein the method is for single cell transcriptome analysis, optionally wherein the sample is a single cell, and the target RNA is transcriptome of the single cell.

32. The method of claim 31 , wherein the single cell transcriptome analysis is single cell CRISPR screening.

33. The method of claim 31, wherein the single cell transcrip tome analysis is single cell T-cell receptor (TCR) and / or B-cell receptor (BCR) sequencing.

Citation Information

Patent Citations

  • Strand displacement stop (SDS) ligation

    US20210332355A1

  • High fidelity reverse transcriptases and uses thereof

    WO2001068895A1

  • Reverse transcriptase with increased enzyme activity and application thereof

    WO2020132966A1