Single molecule sequencing of different RNA species including nascent RNA using a virus model system

STREP-Seq addresses the limitations of existing methods by enabling direct sequencing of nascent viral RNA within purified elongation complexes, providing real-time, high-resolution transcription profiling and insights into viral adaptation and evolution.

WO2025118553A1PCT designated stage expired Publication Date: 2025-06-12THE CHINESE UNIVERSITY OF HONG KONG
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/101190
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-06-25
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current methods for studying transcription regulation and genomic mutations in viruses, such as run-on assays, PRO-Seq, CAGE-seq, GRO-seq, RAP-seq, and Pacific Biosciences Iso-Seq, face limitations including lack of real-time monitoring, low resolution, and inability to efficiently capture viral enzymes, leading to biased results and obscured viral signals.

Method used

The development of Strep-tagged RNA Polymerase Elongation Complex Pull-down Sequencing (STREP-Seq) allows for direct sequencing of nascent viral RNA still embedded within purified viral RNA polymerase elongation complexes, providing real-time transcription profiling with single-nucleotide resolution and enabling precise mapping of polymerase positions.

Benefits of technology

STREP-Seq offers a comprehensive and unbiased method for analyzing viral transcription, enabling the identification of transcriptional signatures associated with viral adaptation, evolution, and pathogenesis, and providing insights into viral replication dynamics and gene expression regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024101190-FTAPPB-I100001
    Figure PCTCN2024101190-FTAPPB-I100001
  • Figure 00000032_0000
    Figure 00000032_0000
  • Figure 00000033_0000
    Figure 00000033_0000
Patent Text Reader

Abstract

A method of RNA sequencing using a novel technique, Strep-tagged viral RNA Polymerase Elongation Complex Pull-down Sequencing (STREP-Seq). STREP-Seq involves sequencing nascent RNA molecules still embedded within purified viral RNA polymerase elongation complexes. This allows direct monitoring of transcription across viral genomes at single-nucleotide resolution.
Need to check novelty before this filing date? Find Prior Art

Description

SINGLE MOLECULE SEQUENCING OF DIFFERENT RNA SPECIES INCLUDING NASCENT RNA USING A VIRUS MODEL SYSTEM

[0001] CROSS-REFERENCE TO RELATED APPLICATION

[0002] This application claims the benefit of U.S. Patent Application Serial No. 63 / 606, 365, filed December 5, 2023, which is hereby incorporated by reference in its entirety including any tables, figures, or drawings.FIELD OF THE INVENTION

[0003] The present invention pertains to a novel sequencing method to study transcription regulation and genomic mutations in viruses. The present invention relates to a sequencing method named Strep-tagged RNA Polymerase Elongation Complex Pull-down Sequencing or STREP-Seq for direct sequencing of nascent viral RNA still embedded within purified viral RNA polymerase elongation complexes.BACKGROUND OF THE INVENTION

[0004] There are several methods that can be utilized today to study transcription regulation and detect mutations in viruses, e.g., run-on assays, PRO-seq, CAGE-seq, GRO-seq, RAP-seq, Native elongating transcript (NET) -seq, and Pacific Biosciences Iso-Seq.

[0005] Run-on assays, such as nuclear run-on sequencing, can be utilized to capture population-level active sites. In this assay, cell nuclei are isolated to capture endogenous RNA polymerases still engaged in elongation, natural NTPs are washed out to halt further transcription, and incorporation of labeled NTPs (e.g., Br-UTP) allows newly synthesized RNA to be distinguished. Ths technique provides a population-level snapshot of active transcription sites in a genome. However, run-on methods have some important limitations. They do not monitor the transcription process in real-time since run-on occurs post-isolation in vitro. The labeled RNA population represents an average of multiple polymerases, therefore obscuring single-molecule dynamics. The position of each polymerase can only be approximated through mapping, not pinpointed. Also, these assays lack nucleotide resolution needed to precisely define pause sites or termination. without nucleotide resolution.

[0006] PRO-Seq (Precision nuclear run-on sequencing) is a powerful technique for profiling active transcription at individual genomic loci. However, this approach has some limitations. PRO-seq, requires indispensable nuclear isolation and lacks strand-specificity. The nuclear isolation step is indispensable because PRO-Seq requires physically separating cell nuclei for  washout of the natural NTPs. However, some of virus only transcribe their genome in the cytoplasm. As a result, such viral transcribing polymerases can not be isolated by using PRO-seq. In addtion, biotinylated NTPs must be incorporated into transcribing RNA for viral polymerase. The biontinylated NTPs is crucial for PRO-seq because these nucleotides serve as chain terminator and are further used for nascent RNA isolation from other RRNAs. Pro-seq is designed for mammalian RNA polymerase and whether biotinylated NTPs can be incorporated and further halt the elongation limits the application of PRO-seq in viral polymerase. In fact, our preliminary experiments showed that biotinylated NTPs can be proofread by viral polymerase of SARS-Covid-2. Additionally, PRO-Seq was optimized for mammalian RNAPII and may not efficiently capture viral enzymes. In addition, bias may be generated from RNA library preparation for PRO-Seq. Illumina next-generation sequencing is used for PRO-Seq. Isolated RAN need to be fragmented, reverse transcribed into cDNA and amplified by PCR before being sequenced. Therefore, bias from cDNA synthesis, adaptor tagmentation and PCR amplification may be introduced into the final sequencing result. Further limitations of PRO-seq are that it identifies the last incorporated nucleotide but not full-length nascent RNA, and that PRO-Seq signal may be obscured by host RNA background in virus-infected systems.

[0007] CAGE-seq provides profiling of transcription start sites (TSS) but not full-length transcripts. It involves sequencing the 5' ends of short RNA fragments extracted from capped transcripts. This allows identifying primary TSS locations across an entire genome. However, CAGE-seq provides only TSS information. CAGE-seq pinpoints initiation but does not reveal full-length transcripts or elongation dynamics. It also lacks transcript elongation data because does not track polyerase progression after initiation, thus missing key regulatory information. Another limitation is its lower resolution of alternative TSS. CAGE-seq may cluster closely spaced TSS instead of resolving individual start sites. Furthermore, CAGE-Seq analyzes processed RNA populations without linkage to transcrption machinery. Overall, CAGE-Seq is a powerful tool for TSS profiling but does not monitor full viral transcription programs in their native state.

[0008] GRO-seq detects initiation GRO-seq (Global Run-On Sequencing) by providing snapshots of transcription initiation genome-wide. This method involves labeling newly synthesized RNA during nuclear run-on, similar to PRO-seq. thus being able to detect polymerase distribution shortly after promoter escape, but not along genes. GRO-seq has several limititations. It only detects pause-free transcripts but cannot resolve pausing during elongation. It also offers low positional resolution because signals arise from pooled newly  initiated transcripts rather than single molecules, and has no viral polymerase targeting. In fact, GRO-seq was optimized for RNA Pol II and may not efficiently isolate viral elongation complexes.

[0009] RAP-seq measures promoter-proximal pausing but cannot resolve genome-wide pausing landscape. RAP-seq (RNA Antisense Purification) utilizes DRIP-seq to enrich for promoter-proximal nascent RNA. It targets RNA / DNA hybrids formed within the first 200-300 nucleotides of transcription, and therefore is able to profile polymerase positioning and pausing near gene starts but not throughout full transcription units. It has low positional resolution because can detect average dwell times in proximal regions rather than pinpoint individual pause sites. In addition, RAP-seq enriches hybrids independently from polymerase identity, which can limit viral signal.

[0010] (NET) -seq (Native elongating transcript sequencing) maps distribution and pausing of endogenous RNA polymerase II. (NET) -seq was developed for cellular systems and requires purification of free nascent RNA to profile elongation genome-wide in yeast. It involves purifying free nascent RNA from digested nuclear run-on samples. NET-seq is highly dependent on the specificity of antibodies used for immunoprecipitation of nascent RNA and requires free RNA purification, which disrupts polymerase complexes. In addition, enrichment steps may not efficiently recover polymerase-associated transcripts because it is optimized for cellular RNAPII, not for viral RdRPs. Viral signals can be obscured by host RNA backgorund after purification, resulting in lower coverage of viral transcription. NET-seq also lacks polymerase linkage because it maps RNA positins without correlation to bound elongation complexes and offers low coverage of viral transcription.

[0011] Pacific Biosciences Iso-Seq utilizes single-molecule real-time sequencing on the PacBio (Menlo Park, California, USA) platform to caracterize RNA isoforms but has a low throughput. Iso-Seq constructs full-length, non-amplified cDNA libraries for long-read sequencing, which reveals transcript isoforms, modifications, and splice variants. However, Iso-Seq has limitations for the characterization of viral transcription because it focuses on isoforms and does not reveal polymerase dynamics or pausing landscapes. It also requires cDNA synthesis and library preparation steps prior to sequencing, and is not optimized for viral enrichment.

[0012] Thus, novel methods of single-molecule sequencing are needed to address the limitations of the exisiting methods and provide a more effective and comprehensive method of RNA sequencing and real-time transcription profiling.

[0013] BRIEF SUMMARY OF THE INVENTION

[0014] The present invention provides a novel sequencing method named Strep-tagged RNA Polymerase Elongation Complex Pull-down Sequencing – “STREP-Seq” (FIG. 1) for analyzing transcription regulation and genomic mutations. In some embodiments, the method of the subject invention comprises direct sequencing of nascent viral RNA still embedded within purified viral RNA polymerase elongation complexes.

[0015] In some aspects, the method of the present invention comprises sequencing nascent viral RNA, including: (a) infecting host cells with a recombinant virus, where the recombinant virus may include a viral RNA-dependent RNA polymerase (RdRP) subunit, and where the viral RdRP subunit is engineered with an epitope tag; (b) isolating one or more intact viral RdRP elongation complexes from infected host cells, and (c) directly sequencing the nascent viral RNA contained within the isolated one or more viral RdRP elongation complexes.

[0016] In some embodiments, the RdRP subunit engineered with an epitope tag may include an exogenous polymerase protein fused to an epitope tag, where the exogenous polymerase can be used to infect host cells before live-virus infection.

[0017] In some embodiments, one or more intact viral RdRP elongation complexes are isolated from the nuclei of the infected host cells using affinity purification of the epitope tag, and viral RNA is extracted and purified for sequencing. In preferred embodiments, sequencing is performed using nanopore sequencing and sequencing data are generated.

[0018] In some embodiments, the present invention provides direct visualization of the polymerase position and transcription status in vivo, allowing precise mapping of each polymerase's position at pause / termination sites, nucleotide-level resolution of regulatory events along individual transcripts, and assessment of complex behaviors like backtracking hidden in run-on averages. In preferred embodiments, sequencing includes monitoring viral RdRP positions along the viral genome with single-nucleotide resolution.

[0019] In some examples, monitoring viral RdRP positions comprises detecting features of viral transcription that may include a promoter, a transcription start site, a pause site, and a termination signal.

[0020] In some embodiments, viral RdRP elongation complexes are isolated under non-denaturing conditions where the nascent viral RNA remans embedded within the intact one or more RdRP complexes during RdRP isolation.

[0021] In some embodiments, patterns in the sequences of viral RdRP positions are analyzed to identify transcriptional signatures that may be associated with viral adaptation, evolution, and pathogenesis.

[0022] In some embodiments, the epitope tag can be a Strep tag, Strep II tag, Twin-Strep II tag, Biotin tag, Avi tag, CBP tag, Myc tag, GFP tag, Flag tag, Fc tag, GST tag, HA tag, His tag, MBP tag, SNAP tag, SUMO tag, V5 tag, VSV tag, S protein tag, or any other peptides or labelings capable of being specifically recognized and captured by antibody-or affinity-based purification and enrichment of the target proteins. In further emodiments, targeting the purification of tag engineered viral RdRP complexes may result in the enrichment of viral RNA over host cell RNA background.

[0023] In preferred embodiments, the viral RdRP is a viral RNA polymerase or RNA-dependent RNA polymerase from a family of viruses including Coronaviridae, Flaviviridae, Orthomyxoviridae, Picornaviridae, Retroviridae, and Rhabdoviridae.

[0024] In more preferred embodiments, the method of the present invention includes a custom adaptor for RNA ligation, where the custom adaptor includes a first oligonucleotide A (oligo A) and a second nucleotide B (oligo B) , where both oligo A and oligo B have fixed nucleotide sequence and length, where oligo A is at least 15 nucleotides long and is phosphate adenylated at its 5’ end, where oligo B is at least 15 nucleotides long; where the last nucleotide at its 3’ is a random nucleotide N; and where the random nucleotide N is complementary to the last base at its 3’ end of the RNA sequence to which oligo A ligates.

[0025] In most preferred embodiments, oligo A includes an adenylated 5’ end, and oligo B includes a non-paired 3’ end. In other preferred embodiments, the sequence of oligo A is 30 nucleotides long, where the 5’ phosphate is adenylated, and the sequence of oligo B is 40 nucleotides long, where nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.

[0026] In some aspects, the present invention comprises a system for sequencing a nascent viral RNA using STREP-Seq, including: (a) a recombinant virus including a viral RdRP subunit engineered with an epitope tag; (b) infecting a host cell with the recombinant virus; (c) purifying intact viral RdRP elongation complexes from the infected cell, where the infected cell includes a viral RdRP engineered with an epitope tag; and (d) sequencing the viral RNA embedded in the purified RdRP elongation complexes using a nanopore sequencing device and generating sequencing data.

[0027] In some embodiments, the system further includes a computer processor and computer-readable instructions that execute a bioinformatics workflow for mapping and analyzing viral RdRP positions and transcription profiles utilizing the sequencing data.

[0028] In preferred embodiments, the present invention includes a STREP-Seq kit for sequencing and analyzing nascent viral RNA, including (a) recombinant viruses with an  engineered epitope tag on a viral RdRP subunit for at least one clinically relevant virus; (b) affinity reagents for isolating tagged viral RdRP elongation complexes; (c) buffers for homogenizing and washing infected cells and isolated complexes; (d) a custom adaptor for RNA ligation, where the custom adaptor comprises a first oligonucleotide A (oligo A) and a second nucleeotide B (oligo B) , where both oligo A and oligo B have fixed nucleotide sequence and length, where oligo A is at least 15 nucleotides long and is phosphate adenylated at its 5’ end, where oligo B is at least 15 nucleotides long; where the last nucleotide at its 3’ end of oligo B is a random nucleotide N; and where the random nucleotide N is complementary to the last base at its 3’ end of the RNA sequence to which oligo A ligates.

[0029] In more preferred embodiments, the STREP-Seq kit includes oligo A and oligo B, where oligo A is 30 nucleotides long, and is 5’ phosphate adenylated, and oligo B is 40 nucleotides long, where nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG. 1. Workflow of STREP-Seq. FIG. 1 Left: Virus culture: Strep-tagged recombinant virus was generated by reverse genetics. FIG. 1 Middle: RNA pull-down: After infection with Strep-tagged recombinant virus, the cell nucleus is isolated. The RNA polymerase complex containing nascent viral RNA is pulled down with Strep-tag purification system. Viral RNA is extracted for sequencing. FIG. 1 Right: Sequencing library preparation: Sequencing is performed using Nanopore (Oxford, United Kingdom) platform with custom adaptor to capture nascent RNA lacking poly (A) tail.

[0031] FIG. 2. Generation of the recombinant IAV. FIG. 2 Top: Schematic representation of the pol I-pol II transcription system for synthesis of viral RNA (vRNA) and mRNA in pHW2000 vector. Influenza A / WSN / 33 virus-derived genome was used. Recombinant IAV was generated by using pHW2000 reverse genetic vectors. FIG. 2 Bottom: Construction of PB2-twin-strep segments. A twin-strep tag was designed to be fused to the C terminus of PB2.

[0032] FIG. 3. Workflow for viral RNA extraction. Subcellular fractionation: After 4-hour infection at MOI: 5 in A549 cell with recombinant strep-tagged virus, the cell nucleus was isolated and the nucleoplasmic fraction was subjected to further purification. Strep purification and RNA extraction: The RNA polymerase complex containing nascent viral RNA in nucleoplasmic fraction was purified with Strep-tag system. After elution, DNase digestion was performed to remove DNA and the RNA was extracted with NucleoZOL and purified with RNA clean-up kit.

[0033] FIGs. 4A-4C. Adopting nanopore sequencing platform for nascent RNA sequencing.

[0034] FIG. 4A The Nanopore platform developed for mRNA sequencing involves a poly (T) DNA adaptor capturing the poly (A) tail in mRNA. FIG. 4B Design of custom DNA adaptor that allows direct ligation to 3’-end RNA nucleotides as well as to the sequencing adaptor. FIG. 4C The custom DNA adaptor modified from the original nanopore design. For Oligo A, the 5’ phosphorylated end was replaced by an adenylated end. For Oligo B, the 3’ poly (T) end was replaced by 1N end. Oligo A and B were annealed before direct ligation with 3’-end RNA nucleotide.

[0035] FIG. 5. Denature RNA polyacrylamide gel analysis for ligation efficiency. 1 μM FAM-30nt RNA was used for ligation by using different adaptors. Adaptor 0N refers to blunt-end oligo B adaptor, 1N, 6N and 9N refers to addition of one, six and nine nucleotides respectively on 3’-end of oligo B adaptor that is non-paired with oligo A adaptor. Refer to FIG. 4C for the adaptor design.

[0036] FIG. 6. Enrichment of nascent viral RNA with STREP-Seq. FIG. 6 Top: Pie chart describing the percentage of mapped reads distribution among the references. The bar chart outlines the distribution of IAV reads mapped to IAV segments. FIG. 6 Bottom: Pie chart describing the percentage of reads with poly (A) tail detected by nanopolish (Loman, et al., Nat Methods, 2015, 12 (8) , 733) .

[0037] FIGs. 7A-7B. IAV Defective viral genome. FIG. 7A Gene coverage plots of vRNA. FIG. 7B Gene coverage plots of c / mRNA.

[0038] FIG. 8. Gene features of nascent viral c / mRNA. Last nucleotide position coverage of 8 IAV genome of c / mRNA with gene feature annotation. Annotations for packaging signals from each gene were previeously described (Li, et al., Virol J, 2021, 18 (1) , 36, Fujii, et al., Proc Natl Acad Sci U S A, 2003, 100 (4) , 2002, Muramoto, et al., J Virol, 2006, 80 (5) , 2318, Ozawa, et al., J Virol, 2007, 81 (1) , 30, Ozawa, et al., J Virol, 2009, 83 (7) , 3384, Fujii, et al., J Virol, 2005, 79 (6) , 3766) .

[0039] FIGs. 9A-9B. Detailed analysis on genomic pausing site of c / mRNA and vRNA. Speculated pausing sites derived from the last nucleotide position of nascent FIG. 9A cRNA / mRNA and FIG. 9B vRNA from IAV from 5’ to 3’ direction. Major peaks are identified by calculating local maxima of read coverage throughout the reference position. The top 3 peaks are labelled by the percentage of reads at the position. Significant peak selected based on percentage and distance between 3’ end is labelled with numbers in percentage. Adjacent box plot denotes the last nucleotide distribution at the significant peak position, with black square box highlighting the correct base from reference.

[0040] FIG. 10. Gene features of nascent viral vRNA. Last nucleotide position coverage of 8 IAV genome of vRNA with gene feature annotation. Annotations for packaging signals from each gene were previously described (Li, et al., Virol J, 2021, 18 (1) , 36, Fujii, et al., Proc Natl Acad Sci U S A, 2003, 100 (4) , 2002, Muramoto, et al., J Virol, 2006, 80 (5) , 2318, Ozawa, et al., J Virol, 2007, 81 (1) , 30, Ozawa, et al., J Virol, 2009, 83 (7) , 3384, Fujii, et al., J Virol, 2005, 79(6) , 3766) .

[0041] FIG. 11. SHAPE-MaP analysis. Median SHAPE-MaP value of A / WSN / 1933 (H1N1) obtained from (Dadonaite, et al., Nat Microbiol, 2019, 4 (11) , 1781) . The median SHAPE-MaP value from each vRNA genome segment was calculated by rolling window median with a window of 50 nt, as previously described (Dadonaite, et al., Nat Microbiol, 2019, 4 (11) , 1781) . Position highlighted by verticle lines indicate the significant peak position of cRNA / mRNA to the corresponding vRNA location which suggests a pausing site with the corresponding median value from SHAPE-MaP. Adjacent box plot denotes the base distribution from vRNA at the peak position derived from cRNA / mRNA, with black square box highlighting the correct base from reference, and the total number of reads that are mapped to the location.

[0042] DESCRIPTION OF SEQUENCES

[0043] SEQ ID NO. 1: 5′-pGGCTTCTTCTTGCTCTTAGGTAGTAGGTTC-3′

[0044] SEQ ID NO. 2: 5′-GAGGCGAGCGGTCAATTTTCCTAAGAGCAAGAAGAAGCCN-3‘

[0045] DETAILED DISCLOSURE OF THE INVENTION

[0046] Selected Definitions

[0047] As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including” , “includes” , “having” , “has” , “with” , or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising” . The transitional terms / phrases (and any grammatical variations thereof) “comprising” , “comprises” , “comprise” , “consisting essentially of” , “consists essentially of” , “consisting” and “consists” can be used interchangeably.

[0048] The term “about” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured, i.e., the limitations of the measurement system. In the context of compositions containing amounts of ingredients where the term “about” is used, these compositions contain the stated amount of the ingredient with a variation (error range) of 0-10%around the value (X  ± 10%) . In other contexts, the term “about” is providing a variation (error range) of 0-10%around a given value (X ± 10%) . As is apparent, this variation represents a range that is up to 10%above or below a given value, for example, X ± 1%, X ± 2%, X ± 3%, X ± 4%, X ± 5%, X ± 6%, X ± 7%, X ± 8%, X ± 9%, or X ± 10%.

[0049] In the present disclosure, ranges are stated in shorthand to avoid having to set out at length and describe each and every value within the range. Any appropriate value within the range can be selected, where appropriate, as the upper value, lower value, or the terminus of the range. For example, a range of 0.1-1.0 represents the terminal values of 0.1 and 1.0, as well as the intermediate values of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, and all intermediate ranges encompassed within 0.1-1.0, such as 0.2-0.5, 0.2-0.8, 0.7-1.0, etc. Values having at least two significant digits within a range are envisioned, for example, a range of 5-10 indicates all the values between 5.0 and 10.0 as well as between 5.00 and 10.00 including the terminal values. When ranges are used herein, combinations and subcombinations of ranges (e.g., subranges within the disclosed range) and specific embodiments therein are explicitly included.

[0050] As used herein, the term “amino acid” refers to standard nomenclature, amino acid residue as denominated by either a three letter or a single letter code as indicated as follows: Alanine (Ala, A) , Arginine (Arg, R) , Asparagine (Asn, N) , Aspartic Acid (Asp, D) , Cysteine (Cys, C) , Glutamine (Gln, Q) , Glutamic Acid (Glu, E) , Glycine (Gly, G) , Histidine (His, H) , Isoleucine (Ile, I) , Leucine (Leu, L) , Lysine (Lys, K) , Methionine (Met, M) , Phenylalanine (Phe, F) , Proline (Pro, P) , Serine (Ser, S) , Threonine (Thr, T) , Tryptophan (Trp, W) , Tyrosine (Tyr, Y) , and Valine (Val, V) .

[0051] As used herein, the terms “oligo” , “oligonucleotide” are used interchangeably to describe short single strands of synthetic DNA or RNA, such as, for example, about a 5 nucleic acid base sequence to about a 500 nucleic acid base sequence.

[0052] As used herein, “vector” refers to a DNA molecule such as a plasmid for introducing a nucleotide construct, for example, a DNA construct, into a host cell. Cloning vectors typically contain one or a small number of restriction endonuclease recognition sites at which foreign DNA sequences can be inserted in a determinable fashion without loss of essential biological function of the vector, as well as a marker gene that is suitable for use in the identification and selection of cells transformed with the cloning vector. Marker genes typically include genes that provide a selectable characteristic, such as tetracycline resistance, hygromycin resistance or ampicillin resistance.

[0053] In this application, the terms “peptide” , and “protein” are used interchangeably herein to refer to a polymer of amino acids. The terms apply to amino acid polymers in which one or  more amino acid residues are artificial chemical mimetic of a corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.

[0054] The terms “label” and like terms refer to a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include fluorescent dyes (fluorophores) , luminescent agents, electron-dense reagents, enzymes (e.g., as commonly used in an ELISA) , biotin, enzymes acting on a substrate (e.g., horseradish peroxidase) , digoxigenin, 32P and other isotopes, haptens, and proteins which can be made detectable, e.g., by incorporating a fluorescent label into the peptide or used to detect antibodies specifically reactive with the peptide. The term includes combinations of single labeling agents, e.g., a combination of fluorophores that provides a unique detectable signature, e.g., at a particular wavelength or combination of wavelengths. In the context of detecting nucleic acids (e.g., target sequences) , the probes can, typically, be labeled with radioisotopes, fluorescent labels (fluorophores) , or luminescent agents.

[0055] By “reduces” is meant a negative alteration of at least 1%, 5%, 10%, 25%, 50%, 75%, or 100%.

[0056] By “increases” is meant as a positive alteration of at least 1%, 5%, 10%, 25%, 50%, 75%, or 100%.

[0057] As used herein, “throughput” describes the amount of material or items passing through a system or process. In some embodiments, increased throughput levels are advantageous.

[0058] As used herein, “encapsidation” refers to the process in which a virus’s nucleic acid is enclosed in a capsid. “Capsid” is the protein shell of a virus that encloses its genetic material.

[0059] As used herein, “tag” or “protein tag” are used interchangeably to refer to a peptide sequence that is genetically grafted onto a recombinant protein. Tags can be attached to proteins for different purposes. In some embodiments, tags are added to either end of the target protein, so that they are C-terminus or N-terminus specific, or both. The term “sample” encompasses a variety of sample types containing a nucleotide. The term encompasses bodily fluids such as blood, blood components, saliva, nasal mucous, serum, plasma, cerebrospinal fluid (CSF) , urine and other liquid samples of biological origin, solid tissue biopsy, tissue cultures, or supernatant taken from cultured patient cells. A sample further encompasses samples from the environment, including water, soil, and air. The sample can be processed prior to assay, e.g., to remove cells or cellular debris. The term encompasses samples that have  been manipulated after their procurement, such as by treatment with reagents, solubilization, sedimentation, or enrichment for certain components.

[0060] As used herein, an “isolated” or “purified” compound is substantially free of other compounds. In certain embodiments, purified compounds are at least 60%by weight (dry weight) of the compound of interest. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99%, by weight of the compound of interest. For example, a purified compound is one that is at least 90%, 91%, 92%, 93%, 94%, 95%, 98%, 99%, or 100% (w / w) of the desired compound by weight. Purity is measured by any appropriate standard method, for example, by column chromatography, thin layer chromatography, or high-performance liquid chromatography (HPLC) analysis.

[0061] As used herein, “RdRP” refers to RNA-dependent RNA polymerase. This is an enzyme that catalyzes the replication of RNA from an RNA template.

[0062] As used herein, “tagmentation” describes when unfragmented DNA is cleaved and tagged for analysis.

[0063] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein.

[0064] Other features and advantages of the invention will be apparent from the following description of the preferred embodiments thereof, and from the claims.

[0065] All references cited herein are hereby incorporated by reference in their entirety.

[0066] The present invention relates to a novel technique, STREP-Seq, for studying viral replication and transcription at high resolution. The method of the present invention includes sequencing nascent RNA molecules still embedded within purified viral RNA polymerase elongation complexes, which allows direct monitoring of transcription across viral genomes at single-nucleotide resolution.

[0067] The purpose and practical use of the STREP-Seq invention is to comprehensively understand viral replication and transcription at high resolution by investigating how these processes are regulated at the viral polymerase level and how they contribute to viral adaptation and evolution. Direct sequencing of nascent viral RNA still embedded within purified viral RNA polymerase elongation complexes provides key insights that were previously inaccessible. STREP-Seq has multiple advantages over the methods that are currently used in RNA sequencing. Conventionally, nascent RNA profiling methods require isolating and purifying nascent RNA pools before sequencing. However, this carries risks of loss, degradation or bias during extraction procedures from the host cell environment. STREP-Seq overcomes this challenge by capturing viral transcription in situ.

[0068] The introduction of a highly specific tag (e.g., Twin-Strep tag) allows targeted purification of the viral RNA polymerase elongation complex directly from infected cell nuclei. This preserves the precise position of the actively transcribing polymerase on the viral genome. It also provides a real-time snapshot of ongoing transcription processes as they occur within the intact enzymatic machinery. Sequencing the nascent RNA molecules while still embedded within the purified elongation complexes thus enables direct monitoring of viral transcription across genomes at single-nucleotide resolution without need for mapping or assembly, pinpointing the exact 3' end of each nascent RNA molecule, eliminating potential losses, artifacts or quantification errors introduced during separate RNA handling / processing steps of prior methods, and capturing rare or transient transcriptional events that may otherwise be missed or underrepresented in population-level profiles. STREP-Seq uniquely allows intricate molecular details and dynamics of viral replication programs to be visualized and studied at an unprecedented genomic scale and resolution through direct nascent RNA sequencing within endogenous viral polymerase complexes.

[0069] The present invention provides a method of detection of polymerase behaviors like pausing and backtracking during transcription elongation. This reveals novel regulatory mechanisms governing viral replication and gene expression that were not previously accessible. Detection of polymerase pausing and backtracking during viral transcription elongation provides novel insights into regulatory mechanisms governing viral gene expression.

[0070] By pinpointing the exact 3' end nucleotide position of each captured nascent RNA molecule, STREP-Seq is uniquely capable of identifying prominent pause sites experienced by the polymerase throughout viral genomes, revealing unappreciated genome-wide patterns and determinants of regulated transcriptional stalling, quantifying kinetics of backtracking behaviors that allow proofreading vs. limiting transcript output, linking pausing signatures to specific sequences, structures or expressed factors to map cis-and trans-acting regulators, comparing profiles between variants to uncover relationships between pause modulation and phenotypic changes.

[0071] STREP-Seq Method of Use

[0072] In some embodiments, the method comprises providing unbiased profiling of complete viral transcription including providing data related to viral replication dynamics, control of gene expression programs, and evolution of such programs, e.g., how viral genomes are differentially transcribed across all phases of infection without any target, the relative  abundances of complex transcripts produced, including pervasive antisense RNAs, nested genes, and alternative splice forms.

[0073] In some embodiments, the method further provides a quantitative assessment of how transcriptional output is balanced between structural or non-structural genes and allocation of resources throughout the viral life cycle, including identifying subtle or lowly-expressed genomic elements and regulatory sequences that influence viral fitnes.

[0074] In preferred embodiments, provided is a highly specific twin-Strep tag on RdRP, which allows viral RNA enrichment against host RNA by strep-tag pull-down of transcribing RdRP complex.

[0075] In some examples, the method can detect rare defective or recombined genomes presenting novel gene constellations, temporal changes in genome-wide transcriptional signatures accompanying viral adaptation or immune selection, and cross-comparison of profiles between variants associated with transcriptional determinants of phenotypic alterations.

[0076] In other examples, the method provides insights into viral population dynamics, pathogenesis, and evolution at the level of transcriptional control.

[0077] Identification of Transcriptional Regulatory Factors

[0078] In some embodiments, the method identifies transcriptional regulatory factors pause sites and mutations that influence viral fitness. Identification of transcriptional regulatory factors, pause sites, and mutations that influence viral fitness through STREP-Seq profiling, has significant practical applications. This knowledge can guide rational design of innovative antiviral drugs and therapies targeting these vulnerabilities. Pinpointing prominent pause sequences experienced throughout viral genomes may reveal vulnerabilities that modified variants lose due to altered pause control. Factors found to be important for maintaining critical polymerase blocks may represent promising antiviral targets. Their disruption may interfere with transcriptional regulation and the viral life cycle. Comparison of pause signatures between variants linked to changes in transmission / pathogenicity may uncover sequence / structure determinants modulating pause profiles with phenotypic consequences. Mutations altering regulated pausing at critical genomic junctures may compromise ordered packaging or gene expression balance, attenuating viral fitness.

[0079] In some examples, the method can uncover novel elements / sequences emerging upon viral adaptation that influence transcriptional outputs or fidelity, and characterize rare mutated defective interfering genomes escaping quality control.

[0080] With comprehensive profiling of complete viral transcriptional landscapes and high-resolution mapping of polymerase kinetics, STREP-Seq provides an unmatched resource for illuminating factors driving viral evolution and suitability as drug targets.

[0081] Potential Applications

[0082] In some embodiments, the methods of the present invention can be applied to any virus system by introducing epitope-tagged versions of viral polymerases. This broadens the technique's usefulness compared to more targeted prior methods. The ability to apply STREP-Seq transcriptional profiling to any virus system by introducing epitope-tagged versions of the viral polymerases broadly enhances the technique's usefulness compared to more targeted prior methods. It allows comprehensive investigation of novel or poorly characterized viruses, helping prioritize public health threats through elucidating replication strategies. By engineering epitope tags at different polymerase subunits, insights into modular contributions and complex assembly can be gained.

[0083] In preferred embodiments, STREP-Seq is not restricted by virus taxonomy or genome organization and is suitable for DNA or RNA viruses with single or segmented genomes.

[0084] In some embodiments, tagged polymerases can be introduced to clinical isolates, uncoupling transcriptional signatures from genetic backgrounds to study adaptation mechanisms.

[0085] In some examples, differentially tagged polymerase variants can be multiplexed in the same sample to enable comparative analyses for dissecting cis-acting elements. Insights into conserved and unique features of viral RNA synthesis across diverse families may provide information for broad-spectrum antiviral design.

[0086] In further examples, polymerases of high-risk pathogens like Ebola, Lassa, Marburg, and others, can be tagged.

[0087] In some embodiments, the present invention offers a universally applicable methodological framework to significantly expand the scope of investigative studies on viral transcriptional control mechanisms beyond what was feasible with earlier targeted approaches, including elucidation of resistance mechanisms to guide combinatorial therapies, evaluation of candidate inhibitors' effects on viral fitness, and surveillance of genetic characterization of emerging variants / mutations implicated in outbreaks. The present invention technique offers a powerful new approach with significant practical value for comprehensive dissection of viral transcriptional control mechanisms and adaptation, thus facilitating development of improved antiviral strategies and countermeasures.

[0088] MATERIALS AND METHODS

[0089] It should be understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and the scope of the appended claims. In addition, any elements or limitations of any invention or embodiment thereof disclosed herein can be combined with any and / or all other elements or limitations (individually or in any combination) or any other invention or embodiment thereof disclosed herein, and all such combinations are contemplated with the scope of the invention without limitation thereto.

[0090] Cell Culture

[0091] Human embryonic kidney 293T (HEK293T) cells and human alveolar basal epithelial (A549) cells were cultured in Dulbecco’s modified Eagle’s medium (DMEM) , supplemented with 10%fetal bovine serum, penicillin, and streptomycin. Madin-Darby canine kidney (MDCK) cells were cultured in Advanced DMEM supplemented with 4%fetal bovine serum, L-alanyl-L-glutamine dipeptide, penicillin, and streptomycin. Cells were grown at 37 ℃ in a humidified 5%CO2 atmosphere.

[0092] Plasmid

[0093] The plasmids pHW2000-PB2 / PB1 / PA / NP / HA / NA / NS / M were used for virus rescue. The plasmid pHW-PB2-Strep was constructed by orderly fusing the twin-Strep tag coding sequence and the 5’ terminal 143 nucleotides of the PB2 segment to the last amino acid of the PB2 open reading frame of pHW-PB2 plasmid using NEBuilder HiFi DNA Assembly Cloning Kit (NEB, E5520) .

[0094] Virus

[0095] The wild-type influenza A / WSN / 1933 (H1N1) and recombinant PB2-strep viruses were rescued by reverse genetics. Briefly, MDCK and HEK293T cells were co-transfected by the eight pHW2000-WSN plasmids in 6-well dishes. For one well of the 6-well dish, 80%of confluence cells in the virus infection medium were transfected with each pHW2000 plasmid. The culture media were collected and spun down after 72 hr. The supernatants were flash-frozen in liquid nitrogen and stored at -80℃. The recombinant virus was propagated in A549 cells at low M.O.I. (multiplicity of infection, M.O.I. = 0.005) in transfection medium (DMEM  supplemented with 20mM HEPES) . Viral titer was measured by standard plaque assay in MDCK cell.

[0096] Subcellular Fractionation

[0097] Twenty million A549 cells (in a 15 cm culture dish) were infected by PB2-strep recombinant influenza virus at high M.O.I. (5) , and then washed and pelleted in PBS. The cell pellet was resuspended in 5 ml sucrose buffer (10 mM Tris, pH 8, 2 mM ZnCl2, 5 mM EDTA, 5 mM EGTA, 340 mM sucrose, 1 mM DTT, 1 × protease inhibitor cocktail, 4 units / ml RNase inhibitor) , and cells were incubated on ice. NP-40 was then added to a final concentration of 0.5%, followed by vortexing and centrifugation. Supernatant was saved as the cytoplasmic fraction, and the nuclear pellet was washed once with 3 ml sucrose buffer. The nuclear pellet was then resuspended in 1.5 ml nucleoplasm extraction buffer (50 mM Tris, pH 8, 420 mM sodium chloride, 1.5 mM MgCl2, 0.1%NP-40, 1 mM DTT, 1 × protease inhibitor cocktail, 4 units / ml RNase inhibitor) , transferred to an all-glass Dounce homogenizer with a tight-fitting pestle, and homogenized with 20 strokes. Homogenate was then incubated 10 min on ice, followed by centrifugation for 10 min at 4 ℃ in a microcentrifuge. The supernatant was collected as the nucleoplasm fraction and used for the following purification.

[0098] RNA Pull-down

[0099] 200 μL  50%suspension (IBA Lifescience) was washed with nucleoplasm extraction buffer. Nuclear lysate was added into column and incubated with the bead before lysate flew through the column. Then the beads were washed by 10 column volume of nucleoplasm extraction buffer and eluted with 400 μl elution buffer (50 mM Tris, pH 8, 100 mM biotin, 2 mM EDTA, 420 mM NaCl, 1 mM DTT, 1 × protease inhibitor cocktail, 4 units / ml RNase inhibitor) . 1 ml NucleoZOL was added into eluates and mixed with 0.4 ml water vigorously and incubated for 5 min at room temperature. The supernatant was transferred after centrifugation. 1 ml Ethanol (> 95%) was then added to the supernatant and mixed before using RNA cleanup column (NEB T2030) to purify RNA. The RNA was eluted in Tris buffer.

[0100] Library Preparation and Sequencing

[0101] RNA libraries were prepared by using the ONT SQK-RNA002 kit with some modification from the standard protocol. A custom DNA adapter was used in place of the RTA adapter provided by the kit. To prepare the custom adapter, the top strand of the custom (SEQ ID NO. 1: 5′-pGGCTTCTTCTTGCTCTTAGGTAGTAGGTTC-3′) was adenylated in advance  using 5' DNA Adenylation Kit (NEB, E2610) , and then mixed with bottom strand (SEQ ID NO. 2: 5′-GAGGCGAGCGGTCAATTTTCCTAAGAGCAAGAAGAAGCCN-3 ‘) in annealing buffer (10 mM Tris-HCl 7.5, 50 mM NaCl) to the final concentration of 1.4 μM each in a total volume of 50 μL. RNA (300-500 ng) was ligated to 4 μL custom adapter in a total reaction volume of 40 μl with 15%PEG 8000 (NEB, B1004) , 1× T4 RNA Ligase Buffer (NEB, B0216) , 1 μl of T4 RNA ligase 2 KQ (NEB, M0373) , 0.5 μl ERCC RNA Spike-In Mix (Invitrogen, 4456740) , 0.5 μl calibrant RNA strand from ONT kit, and 1 μl of RNaseOUT (Invitrogen, 10777019) at 16℃ for 16h. The optional reverse transcription step was omitted and a 1.8× volume of room-temperature-equilibrated AMPure RNAClean XP beads (Beckman Coulter, A63987) was used for RNA clean up. The RNA was washed with 150 μL 70%freshly prepared ethanol and eluted by resuspending the beads in 21 μL nuclease-free water. The RNA concentration was determined using RNA HS Qubit Fluorometric Quantification. 6 μL RNA Adapter (RMX) from ONT kit, composed of sequencing adapters with motor protein, was ligated onto the first ligation product by using Quick T4 DNA Ligase (NEB, M2200) and 1×T4 DNA Ligase Buffer (NEB, B6058) in a total volume of 40 μL at room temperature for 15 min. The reaction was stopped by adding 1× volume beads. The RNA library was washed with 150 μL wash buffer (ONT SQK-RNA002 kit) twice, and then eluted in 21 μL elution buffer (ONT SQK-RNA002 kit) and mixed with RNA running buffer (ONT SQK-RNA002 kit) prior to loading onto a primed R9.4.1 flowcell. The samples were run on MinION sequencer for 48 h or less (until all pores were inactive) .

[0102] Evaluation of Custom Adaptor Ligation Efficiency

[0103] The custom linker contains a first oligonucleotide (oligonucleotide A) and a second oligonucleotide (oligonucleotide B) in fixed sequence and length. Oligonucleotide A is 30 nucleotides in length with the 5′phosphate adenylated, and oligonucleotide B is 40 nucleotides in length. Nucleotides 1 to 20 of oligo A are complementary to nucleotides 20 to 39 of oligo B. The last nucleotide at the 3’ end of oligo B is a random nucleotide (N) , complementary to the last base at the 3’ end of the RNA sequence to which the oligo A ligates.

[0104] Sequence Analysis

[0105] Samples were sequenced under the default parameters for Nanopore Direct RNA Sequencing libraries (SQK-RNA002) and output as fast5 files. After converting fast5 files to pod5 files, basecalling were performed using Dorado v0.3.4 (Oxford Nanopore Technologies) to produce fastq files. Basecalled reads were aligned to host, virus, and control references  using minimap2 (Li, Bioinformatics, 2018, 34 (18) , 3094) . In IAV sequencing experiments, human assembly GRCh38. p14, FASTA sequence of IAV genome A / WSN / 1933 (H1N1) , and ERCC sequence (Invitrogen) were used as reference. To evaluate whether the library preparation contains mature RNA, nanopolish (Loman, et al., Nat Methods, 2015, 12 (8) , 733) was used to identify aligned reads with poly (A) tail. After discarding unaligned reads, reads were categorized as cRNA / mRNA or vRNA based on its direction. The coverage and sequencing depth were extracted using samtools (Li, et al., Bioinformatics, 2009, 25 (16) , 2078) . A custom Python script was used to identify the last nucleotide position of each read.

[0106] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification.

[0107] Following are examples that illustrate procedures for practicing the invention. All percentages are by weight and all solvent mixture proportions are by volume unless otherwise noted.

[0108] EXAMPLE 1-RATIONALE AND DESIGN OF NASCENT VIRAL RNA SEQUENCING METHOD

[0109] The method of the present invention was first developed using a BSL2 virus Influenza A virus (IAV) , an enveloped single-strand RNA virus that causes pandemics of human respiratory tract infection and severe clinical diseases. We applied reverse genetics technique to generate recombinant IAV carrying highly-specific twin-Strep tag on RdRP PB2 based on published plasmid design (Li, et al., Veterinary Microbiology, 2021, 254 (108985) (FIG. 2) . Recombinant IAV was generated by co-transfecting MDCK and HEK293T cells with eight pHW2000-WSN plasmids containing the constructed pHW-PB2-Strep plasmid. Propagation of recombinant virus was then performed in A549 cells. After a 4-hour viral infection of propagated recombinant virus in A549 cell, cell nuclei where IAV RdRP populated was isolated. RdRP elongation complexes carrying nascent viral RNA were then pulled down following standard Strep-tag purification protocol, in which the viral RNA is further extracted and purified for sequencing (FIG. 3) .

[0110] We adopted Nanopore platform in our sequencing method which provides several advantages over Illumina. For instance, when using Illumina next-generation sequencing (NGS) system for RNA-sequencing, the sequencing library is subjected to PCR bias introduced during library preparation, such as cDNA synthesis and adaptor tagmentation before sequencing, resulting in a changed ratio of molecules that is different from the original input.  In contrast, Nanopore direct RNA sequencing (DRS) utilizes adaptor ligation to RNA molecules, thereby excluding any bias from amplification and allowing sequencing directly the sample-extracted RNA. In addition, the position of the last nucleotide of the actively transcribing RNA molecules, a highlight in many nascent RNA-Seq technology, can be easily measured and pinpointed through Nanopore DRS, while sequencing through short-read sequencing such as Illumina sequencing require reference mapping that introduces errors, and complex library preparation design to preserve positive and negative sense viral RNAs that could compromise RNA integrity before the sequencing step.

[0111] In Nanopore platform, a sequencing adaptor containing a motor protein is ligated to the poly (T) adaptor to capture the poly (A) tail on mature mRNA. During Nanopore sequencing, the motor protein facilitates the directional passage of the captured RNA through a protein nanopore where electrical current changes is monitored (FIG. 4A) . The 3’-end of RNA is sequenced first and provides information on strand direction. Therefore, Nanopore platform allows direct RNA sequencing without generating cDNA and does not require PCR or other forms of amplification. Since nascent RNA lacks the poly (A) tail targeted by Nanopore platform, we therefore designed a custom DNA adaptor based on the original Nanopore adaptor (FIG. 4B and FIG. 4C) . Our custom DNA adaptor allows direct ligation to 3’-end RNA nucleotides as achieved by a designed adenylated 5’-end in Oligo A and a 1N end in Oligo B (FIG. 5) .

[0112] The ligation efficiency of our custom adaptor to the RNA is critical to prevent loss of genetic materials in subsequent sequencing. The optimization of ligation condition was performed with adaptors with different number of N on 3’-end of Oligo B. Our results showed that RNA can be efficiently ligated to the custom adaptor with 1N on Oligo B adaptor by using RNA T4 ligase2 truncated KQ at 25 ℃ (FIG. 5) .

[0113] EXAMPLE 2-SUCCESSFUL ENRICHMENT OF NASCENT VIRAL RNA WITH PB2-STREP RECOMBINANT IAV IN STREP-SEQ

[0114] With the use of Strep-tag purification system, the viral RNA was successfully enriched (FIG. 6) . A high signal-to-background ratio was shown by a high percentage of IAV reads compared to host, with 54.5%of the reads were IAV reads, 40.9%from host, and 4.6%from the ERCC reference. As a result, 38285 IAV reads were available for our genomic analysis. All the IAV 8 segments were successfully pulled down by purifying PB2-strep RdRP in our STREP-Seq. Except for a higher proportion of PA segment (22%) , all other 7 segments were  evenly distributed (9–13%) . Most importantly, 92.92%of the total mapped IAV reads were nascent RNA, allowing us to perform statistical analysis on genomic regulation mechanism.

[0115] EXAMPLE 3-DEFECTIVE VIRAL GENOME WITH LARGE DELETIONS IN PB2, PB1 AND PA SEGMENTS

[0116] One of our proposed goals is to determine how SARS-CoV-2 polymerase transcription contributes to genetic variations. One such example are defective viral genomes (DVGs) , which are viral subgenomic materials carrying mutations, deletions, or a variety of gene rearrangements that have been observed in a wide range of viruses. (Vignuzzi and López, Nature Microbiology, 2019, 4 (7) , 1075) . DVGs carrying incomplete genetic information were known to compete with full genomes in the form of defective interfering particles (DIPs) . DIPs are viral particles packed with DVGs that compete with the “standard” particles packed with full-length genomes with a faster replication and packaging speed. By controlling the balance of DIPs subpopulation, viruses regulate a range of virus–host interactions, including reducing virulence, promoting antiviral immunity, and increasing host cell survival, to enhance transmission and virus persistence (Vignuzzi and López, Nature Microbiology, 2019, 4 (7) , 1075; Ghorbani, et al., Annual Review of Animal Biosciences, 2020, 8 (247) .

[0117] DVGs were largely observed in our system as indicated by significantly lower internal gene coverages in PB2, PB1 and PA segments in vRNA (FIG. 7A) , and similar defective segments were observed in c / mRNA (FIG. 7B) . IAV DVGs were usually characterized by a large internal deletion of gene segments while retaining the packaging signals at 3’-and 5’-termini and observed most frequently from PB2, PB1 and PA gene (Ghorbani, et al., Annual Review of Animal Biosciences, 2020, 8 (247) ; Jakob, et al., Nucleic Acids Research, 2022, 50(16) , 9023) . These have been known to be generated in virus infection at high M.O.I. (≥ 0.1) and able to interfere with the replication of stranded virus (Xue, et al., Frontiers in Microbiology, 2016, 7 (326) .

[0118] In our current method development settings with the use of low M.O.I. (= 0.005) for propagation and high M.O.I. (= 5) for infection (FIG. 3) , the generation of DVGs was largely promoted during infection stage. Hence, our data set represents the viral genomic fingerprint under highly-activated virus–host regulatory mechanism. The lengths of DVGs in the PB2, PB1 and PA segments were consistent with previous in vitro and clinical studies (Jennings, et al., Cell, 1983, 34 (2) , 619; Noble and Dimmock, Virology, 1995, 210 (1) , 9; Saira, et al., Journal of Virology, 2013, 87 (14) , 8064) . Both 5’ and 3’ packaging signals were maintained  in the DVGs. While PB2 and PB1 DVGs contained comparable lengths of 3’ and 5’ regions, PA DVGs had more contents in 5’ region. Imbalance in PA DVG termini was also observed in IAV from clinical specimens (Saira, et al., Journal of Virology, 2013, 87 (14) , 8064) .

[0119] EXAMPLE 4-TRANSCRIPTION PAUSING SITE IN PB2, PB1 AND PA SEGMENTS

[0120] One of the most valuable pieces of information provided by our STREP-Seq is the viral transcriptional pausing site, which could only be captured by nascent RNA sequencing. Therefore, we investigated the population of the last nucleotide position and observed transcriptional pausing sites appear at the 3’ packaging signal region of all the 8 IAV segments in c / mRNA (FIG. 8) . By a closer inspection on the pausing sites, we discovered that the most populated pausing site in each segment was not mutational pausing. The percentage of correct base at the identified sites was 100%for all the PB2, PB1, PA, HA, NP, NA, M and NS segments, respectively (FIG. 9A) . A comparison study was performed with vRNA, where the results showed that vRNA pausing sites were also not driven by mutation as 100%of the sequence were correct at the identified sites for all the 8 segments (FIG. 9B and FIG. 10) . Hence, we speculated that the transcriptional pausing was due to a local RNA secondary structure barrier that RdRP encountered during transcription. We plotted the SHAPE-MaP reactivities along the vRNA genome previously reported for influenza A / WSN / 1933 (H1N1) (Dadonaite, et al., Nat Microbiol, 2019, 4 (11) , 1781) and compared with our c / mRNA pausing sites (FIG. 11) . In this SHAPE-MaP analysis, more positive value represents less structure and more negative value represents more structure. Since the analysis was performed with vRNA, the 3’-end c / mRNA pausing sites appeared at 5’-end of vRNA genome.

[0121] We discovered that transcriptional pausing was mainly driven by RNA structure but there could be exceptions. For PB2, PB1, HA, NA, M and NS segments, the pausing positions were located in the region with increasing structure (sloping down towards 5’-end) where the local values were -0.17, -0.17, -0.13, -0.13, -0.07 and -0.2 respectively (FIG. 11) . For PA segment the pausing position was located at the local maxima (the least structured) , with a local value of -0.01. The only exception was observed in NP segment, where the pausing site with local value of 0.02 appeared in the region with upward slopping, implying less structure in the nearby region (FIG. 11) . The observed pausing was also not driven by the mutation on vRNA template as for all the 8 segments the percentage of correct vRNA sequence fell within the range of 96.82–100% (FIG. 11) , although this number is slightly lower than the one of vRNA pausing sites (FIG. 9B) .

[0122] Overall, our analysis showed that the viral transcriptional pausing was mainly driven by RNA structure not mutation. However, there could be contributions from other RNA-protein interactions or protein-protein interactions that were not accounted for in our analysis as one exceptional case was identified in NP segment.

[0123] Our results revealed widespread polymerase pausing, predominantly at untranslated regions. This suggests that precise regulation of transcriptional processivity is critical for viral genome encapsidation. STREP-Seq thus provides unique insights into the molecular details and dynamics of viral replication and gene expression. By elucidating transcriptional regulatory mechanisms at high resolution, STREP-Seq has the potential to revolutionize our understanding of how viruses adapt and evolve through fine-tuning of their replication and gene expression programs. Knowledge gained from comprehensive profiling of viral transcription landscapes using this approach will facilitate rational design of antiviral strategies. Targeting transcriptional regulatory factors or pausing sequences identified may interfere with viral replication or alter pathogenic phenotypes. STREP-Seq is a powerful new technique for investigating viral replication and transcription at an unprecedented genomic scale. It offers new opportunities to decipher the intricate mechanisms that drive viral evolution and pathogenesis, enabling development of innovative antiviral therapies.

[0124] EXEMPLARY EMBODIMENTS

[0125] Embodiment 1. A method of sequencing nascent viral RNA, comprising:

[0126] (a) infecting host cells with a recombinant virus, wherein the recombinant virus comprises a viral RNA-dependent RNA polymerase (RdRP) subunit, and wherein the viral RdRP subunit is engineered with an epitope tag;

[0127] (b) isolating one or more intact viral RdRP elongation complexes from said infected host cells, and

[0128] (c) directly sequencing the nascent viral RNA contained within the isolated one or more viral RdRP elongation complexes.

[0129] Embodiment 2. The method of embodiment 1, wherein the RdRP subunit engineered with an epitope tag comprises an exogenous polymerase protein fused to an epitope tag.

[0130] Embodiment 3. The method of embodiment 2, wherein the exogenous polymerase protein fused to an epitope tag is used to infect the host cells before live-virus infection.

[0131] Embodiment 4. The method of embodiment 1, wherein one or more intact viral RdRP elongation complexes are isolated from nuclei of the infected host cells using affinity purification of the epitope tag.

[0132] Embodiment 5. The method of embodiment 1, wherein the sequencing comprises using nanopore sequencing and generating sequencing data.

[0133] Embodiment 6. The method of embodiment 5, comprising mapping viral RdRP positions and transcriptional elongation dynamics within the viral genome using the sequencing data.

[0134] Embodiment 7. The method of embodiment 1, wherein sequencing comprises monitoring viral RdRP positions along the viral genome with single-nucleotide resolution.

[0135] Embodiment 8. The method of embodiment 7, wherein the monitoring viral RdRP positions comprises detecting features of viral transcription of a promoter, a transcription start site, a pause site, or a termination signal.

[0136] Embodiment 9. The method of embodiment 1, wherein viral RdRP elongation complexes are isolated under non-denaturing conditions, wherein the non-denaturing conditions comprise condititons where the nascent viral RNA remains embedded within the intact viral one or more RdRP elongation complexes during said sequencing.

[0137] Embodiment 10. The method of embodiment 7, comprising analizying patterns in the sequenced viral RdRP positions to identify transcriptional signatures, wherein transcriptional signatures are associated with viral adaptation, evolution, or pathogenesis.

[0138] Embodiment 11. The method of embodiment 1, wherein the epitope tag comprises at least one of Strep tag, Strep II tag, Twin-Strep II tag, Biotin tag, Avi tag, CBP tag, Myc tag, GFP tag, Flag tag, Fc tag, GST tag, HA tag, His tag, MBP tag, SNAP tag, SUMO tag, V5 tag, VSV tag, and S protein tag.

[0139] Embodiment 12. The method of embodiment 1, comprising targeting purification of tag engineered viral RdRP complexes, wherein the targeting results in the enrichment of viral RNA over host cell RNA background.

[0140] Embodiment 13. The method of embodiment 1, wherein the viral RdRP comprises a viral RNA polymerase or RNA-dependent RNA polymerase from viruses Coronaviridae, Flaviviridae, Orthomyxoviridae, Picornaviridae, Retroviridae or Rhabdoviridae.

[0141] Embodiment 14. The method of embodiment 1, comprising obtaining a custom adaptor for RNA ligation, wherein the custom adaptor comprises a first oligonucleotide A (oligo A) and a second nucleeotide B (oligo B) , wherein both oligo A and oligo B have fixed nucleotide sequence and length, wherein oligo A comprises a sequence of at least 15 nucleotides and is phosphate adenylated at its 5’ end, and wherein oligo B comprises a sequence of at least 15 nucleotides, wherein a last nucleotide at a 3’ end of oligo B is a random nucleotide N; and wherein the random nucleotide N is complementary to a last base at a 3’ end of an RNA sequence to which oligo A ligates.

[0142] Embodiment 15. The method of embodiment 14, wherein oligo A comprises an adenylated 5’ end, and oligo B comprises a non-paired 3’ end.

[0143] Embodiment 16. The method of embodiment 14, wherein the sequence of oligo A is 30 nucleotides long and is phosphate adenylated at its 5’ end, wherein the sequence of oligo B is 40 nucleotides long; and wherein nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.

[0144] Embodiment 17. A system for sequencing a nascent viral RNA using STREP-Seq, comprising

[0145] (a) a recombinant virus comprising a viral RdRP subunit engineered with an epitope tag;

[0146] (b) means for infecting a host cell with the recombinant virus;

[0147] (c) means for purifying intact viral RdRP elongation complexes from the infected cell, wherein the infected cell comprises a viral RdRP engineered with an epitope tag;

[0148] (d) means for sequencing the viral RNA embedded in the purified RdRP elongation complexes using a nanopore sequencing device and generating sequencing data.

[0149] Embodiment 18. The system of embodiment 17, wherein the system comprises a computer processor and computer-readable instructions that execute a bioinformatics workflow for mapping and analyzing viral RdRP positions and transcription profiles utilizing the sequencing data.

[0150] Embodiment 19. A STREP-Seq kit for sequencing and analyzing nascent viral RNA, comprising

[0151] (a) recombinant viruses comprising an engineered epitope tag on a viral RdRP subunit for at least one target virus;

[0152] (b) affinity reagents for isolating tagged viral RdRP elongation complexes;

[0153] (c) buffers for homogenizing and washing infected cells and isolated complexes;

[0154] (d) a custom adaptor for RNA ligation, wherein the custom adaptor comprises a first oligonucleotide A (oligo A) and a second nucleeotide B (oligo B) , wherein both oligo A and oligo B have fixed nucleotide sequence and length, wherein oligo A comprises a sequence of at least 15 nucleotides and is phosphate adenylated at its 5’ end, and wherein oligo B comprises a sequence of at least 15 nucleotides, and wherein the oligo B comprises a non-paired 3’ end.

[0155] Embodiment 20. The A STREP-Seq kit of embodiment 19, wherein the sequence of oligo A is 30 nucleotides long and is phosphate adenylated at its 5’ end, wherein the sequence of oligo B is 40 nucleotides long; and wherein nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.

[0156] REFERENCES

[0157] 1. Loman, et al., A complete bacterial genome assembled de novo using only nanopore sequencing data. Nat Methods, 2015, 12 (8) , 733.

[0158] 2. Li, et al., Packaging signal of influenza A virus. Virol J, 2021, 18 (1) , 36.

[0159] 3. Fujii, et al., Selective incorporation of influenza virus RNA segments into virions. Proc Natl Acad Sci U S A, 2003, 100 (4) , 2002.

[0160] 4. Muramoto, et al., Hierarchy among viral RNA (vRNA) segments in their role in vRNA incorporation into influenza A virions. J Virol, 2006, 80 (5) , 2318.

[0161] 5. Ozawa, et al., Contributions of two nuclear localization signals of influenza A virus nucleoprotein to viral replication. J Virol, 2007, 81 (1) , 30.

[0162] 6. Ozawa, et al., Nucleotide sequence requirements at the 5′ end of the influenza A virus M RNA segment for efficient virus replication. J Virol, 2009, 83 (7) , 3384.

[0163] 7. Fujii, et al., Importance of both the coding and the segment-specific noncoding regions of the influenza A virus NS segment for its efficient incorporation into virions. J Virol, 2005, 79 (6) , 3766.

[0164] 8. Dadonaite, et al., The structure of the influenza A virus genome. Nat Microbiol, 2019, 4 (11) , 1781.

[0165] 9. Li, Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 2018, 34 (18) , 3094.

[0166] 10. Li, et al., The Sequence Alignment / Map format and SAMtools. Bioinformatics, 2009, 25 (16) , 2078.

[0167] 11. Li, et al., Generation of recombinant influenza virus bearing strep tagged PB2 and effective identification of interactional host factors. Veterinary Microbiology, 2021, 254 (108985.

[0168] 12. Vignuzzi and López, Defective viral genomes are key drivers of the virus–host interaction. Nature Microbiology, 2019, 4 (7) , 1075.

[0169] 13. Ghorbani, et al., Influenza A Virus subpopulations and their implication in pathogenesis and vaccine development. Annual Review of Animal Biosciences, 2020, 8 (247) .

[0170] 14. Jakob, et al., The influenza A virus genome packaging network-complex, flexible and yet unsolved. Nucleic Acids Research, 2022, 50 (16) , 9023.

[0171] 15. Xue, et al., Propagation and Characterization of Influenza Virus Stocks That Lack High Levels of Defective Viral Genomes and Hemagglutinin Mutations. Frontiers in Microbiology, 2016, 7 (326) .

[0172] 16. Jennings, et al., Does the higher order structure of the influenza virus ribonucleoprotein guide sequence rearrangements in influenza viral RNA? Cell, 1983, 34 (2) , 619.

[0173] 17. Noble and Dimmock, Characterization of putative defective interfering (DI) A / WSN RNAs isolated from the lungs of mice protected from an otherwise lethal respiratory infection with influenza virus A / WSN (H1N1) : a subset of the inoculum DI RNAs. Virology, 1995, 210 (1) , 9.

[0174] 18. Saira, et al., Sequence Analysis of In Vivo Defective Interfering-Like RNA of Influenza A H1N1 Pandemic Virus. Journal of Virology, 2013, 87(14) , 8064) .

Claims

1.A method of sequencing nascent viral RNA, comprising:(a) infecting host cells with a recombinant virus, wherein the recombinant virus comprises a viral RNA-dependent RNA polymerase (RdRP) subunit, and wherein the viral RdRP subunit is engineered with an epitope tag;(b) isolating one or more intact viral RdRP elongation complexes from said infected host cells, and(c) directly sequencing the nascent viral RNA contained within the isolated one or more viral RdRP elongation complexes.2.The method of claim 1, wherein the RdRP subunit engineered with an epitope tag comprises an exogenous polymerase protein fused to an epitope tag.3.The method of claim 2, wherein the exogenous polymerase protein fused to an epitope tag is used to infect the host cells before live-virus infection.4.The method of claim 1, wherein one or more intact viral RdRP elongation complexes are isolated from nuclei of the infected host cells using affinity purification of the epitope tag.5.The method of claim 1, wherein the sequencing comprises using nanopore sequencing and generating sequencing data.6.The method of claim 5, comprising mapping viral RdRP positions and transcriptional elongation dynamics within the viral genome using the sequencing data.7.The method of claim 1, wherein sequencing comprises monitoring viral RdRP positions along the viral genome with single-nucleotide resolution.8.The method of claim 7, wherein the monitoring viral RdRP positions comprises detecting features of viral transcription of a promoter, a transcription start site, a pause site, or a termination signal.9.The method of claim 1, wherein viral RdRP elongation complexes are isolated under non-denaturing conditions, wherein the non-denaturing conditions comprise condititons where the nascent viral RNA remains embedded within the intact viral one or more RdRP elongation complexes during said sequencing.10.The method of claim 7, comprising analizying patterns in the sequenced viral RdRP positions to identify transcriptional signatures, wherein transcriptional signatures are associated with viral adaptation, evolution, or pathogenesis.11.The method of claim 1, wherein the epitope tag comprises at least one of Strep tag, Strep II tag, Twin-Strep II tag, Biotin tag, Avi tag, CBP tag, Myc tag, GFP tag, Flag tag, Fc tag, GST tag, HA tag, His tag, MBP tag, SNAP tag, SUMO tag, V5 tag, VSV tag, and S protein tag.12.The method of claim 1, comprising targeting purification of tag engineered viral RdRP complexes, wherein the targeting results in the enrichment of viral RNA over host cell RNA background.13.The method of claim 1, wherein the viral RdRP comprises a viral RNA polymerase or RNA-dependent RNA polymerase from viruses Coronaviridae, Flaviviridae, Orthomyxoviridae, Picornaviridae, Retroviridae, or Rhabdoviridae.14.The method of claim 1, comprising obtaining a custom adaptor for RNA ligation, wherein the custom adaptor comprises a first oligonucleotide A (oligo A) and a second oligonucleotide B (oligo B) , wherein both oligo A and oligo B have fixed nucleotide sequence and length, wherein oligo A comprises a sequence of at least 15 nucleotides and is phosphate adenylated at its 5’ end, and wherein oligo B comprises a sequence of at least 15 nucleotides, wherein a last nucleotide at a 3’ end of oligo B is a random nucleotide N; and wherein the random nucleotide N is complementary to a last base at a 3’ end of an RNA sequence to which oligo A ligates.15.The method of claim 14, wherein oligo A comprises an adenylated 5’ end, and oligo B comprises a non-paired 3’ end.16.The method of claim 14, wherein the sequence of oligo A is 30 nucleotides long and is phosphate adenylated at its 5’ end, wherein the sequence of oligo B is 40 nucleotides long; and wherein nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.17.A system for sequencing a nascent viral RNA using STREP-Seq, comprising(a) a recombinant virus comprising a viral RdRP subunit engineered with an epitope tag;(b) means for infecting a host cell with the recombinant virus;(c) means for purifying intact viral RdRP elongation complexes from the infected cell, wherein the infected cell comprises a viral RdRP engineered with an epitope tag;(d) means for sequencing the viral RNA embedded in the purified RdRP elongation complexes using a nanopore sequencing device and generating sequencing data.18.The system of claim 17, wherein the system comprises a computer processor and computer-readable instructions that execute a bioinformatics workflow for mapping and analyzing viral RdRP positions and transcription profiles utilizing the sequencing data.19.A STREP-Seq kit for sequencing and analyzing nascent viral RNA, comprising(a) recombinant viruses comprising an engineered epitope tag on a viral RdRP subunit for at least one target virus;(b) affinity reagents for isolating tagged viral RdRP elongation complexes;(c) a buffer for homogenizing and washing infected cells and isolated complexes; and(d) a custom adaptor for RNA ligation, wherein the custom adaptor comprises a first oligonucleotide A (oligo A) and a second oligonucleotide B (oligo B) , wherein both oligo A and oligo B have fixed nucleotide sequence and length, wherein oligo A comprises a sequence of at least 15 nucleotides and is phosphate adenylated at its 5’ end, and wherein oligo B comprises a sequence of at least 15 nucleotides, and wherein the oligo B comprises a non-paired 3’ end.20.The A STREP-Seq kit of claim 19, wherein the sequence of oligo A is 30 nucleotides long and is phosphate adenylated at its 5’ end, wherein the sequence of oligo B is  40 nucleotides long; and wherein nucleotides 1-20 of oligo A are complementary to nucleotides 20-39 of oligo B.

Citation Information

Patent Citations

  • Methods and compositions for nucleic acid sample preparation

    US20150072869A1

  • RNA dependent RNA polymerase mediated protein evolution

    WO2003027330A1

  • Polynucleotide adapters and methods of use thereof

    WO2018175798A1

  • Methods of producing ribosomal ribonucleic acid complexes

    WO2018200847A1

  • Recombinant RNA-dependent RNA polymerase of RNA viruses

    WO2019056090A1