Bat-related reverse transcriptases and their usage
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-08-14
AI Technical Summary
主要关注点在于脱靶效应的可能性,其中系统无意中修饰了非预期的DNA序列,导致不可预见的后果
Smart Images

Figure CN122580438A_ABST
Abstract
Description
[0001] References to sequence lists
[0002] This application contains a sequence list that has been electronically submitted in XML format, which is incorporated herein by reference in its entirety. The XML copy created on September 17, 2024, is named “Sequence Listing_084284.00298.xml” and has a size of 162,353 bytes.
[0003] Cross-references to related applications
[0004] This application claims priority to U.S. Provisional Patent Application No. 63 / 583,648, filed September 19, 2023, pursuant to 35 USC §119(e). The foregoing application is incorporated herein by reference in its entirety. Technical Field
[0005] This disclosure generally relates to bat-related reverse transcriptases and methods of using them. Background Technology
[0006] The broader context of this disclosure is the ongoing evolutionary arms race between organisms and viruses. This ongoing struggle has driven the development of diverse and complex antiviral defense mechanisms across all domains of life. The CRISPR-Cas system, an adaptive immune mechanism employed by prokaryotes, has become a breakthrough tool in gene editing. Its ability to precisely target and modify specific DNA sequences has revolutionized biomedical research, enabling scientists to manipulate genes with unprecedented accuracy and efficiency. The advent of CRISPR-Cas has not only accelerated the understanding of gene function but also paved the way for innovative gene therapies, offering potential solutions for a wide range of genetic diseases. While the field of gene editing has been significantly advanced by the CRISPR-Cas system, it is not without its limitations. A major concern is the possibility of off-target effects, where the system inadvertently modifies unintended DNA sequences, leading to unforeseen consequences. Furthermore, delivering these bacterial-derived systems into mammalian cells can be challenging, and their long-term presence may trigger immune responses. The inherent incompatibility between prokaryotic CRISPR-Cas systems and the complex mechanisms of eukaryotic cells further underscores the need for alternative gene-editing tools better suited to the mammalian environment. The development of such tools could potentially enhance the precision and safety of gene editing, paving the way for more effective and ethical therapeutic applications.
[0007] The current limitations of CRISPR-Cas systems, particularly their struggle against rapidly evolving RNA viruses and potential incompatibility with mammalian cellular mechanisms, underscore the urgent need for mammalian equivalents. The discovery of such systems could not only provide more customized and effective methods for gene editing in humans but also reveal unique antiviral strategies employed by mammals, potentially leading to breakthroughs in combating viral infections. Therefore, there is a strong need for improved gene editors to address the shortcomings of existing tools, enabling more precise gene editing with reduced off-target effects and paving the way for safer and more effective gene therapies. Summary of the Invention
[0008] This disclosure addresses the aforementioned needs in several aspects. In one aspect, this disclosure provides a method for modifying a target polynucleotide. In some embodiments, the method includes delivering an enzyme having reverse transcriptase and endonuclease activity, and one or more nucleic acid components to the target polynucleotide, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0009] In another aspect, this disclosure provides a method for modifying the expression of a target polynucleotide. In some embodiments, the method includes: introducing into a cell or a subject an enzyme having reverse transcriptase activity and endonuclease activity, or a nucleic acid molecule encoding said enzyme, and one or more nucleic acid components, wherein said one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide, and wherein said enzyme binds to one or more sites on the target polynucleotide such that the binding of the enzyme increases or decreases the expression level of the target polynucleotide.
[0010] In some embodiments, the enzyme further possesses integrase activity. In some embodiments, one or more nucleic acid components include single-stranded navigator DNA.
[0011] In some embodiments, one or more nucleic acid components further include a payload RNA. In some embodiments, the payload RNA includes a stem-loop structure.
[0012] In some embodiments, the enzyme reverse transcribes the payload RNA into cDNA. In other embodiments, the enzyme integrates the cDNA into the target polynucleotide.
[0013] In some embodiments, the enzyme includes an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain. In some embodiments, the enzyme further includes a C-terminal domain. In some embodiments, the C-terminal domain facilitates interaction between the enzyme and the target polynucleotide. In some embodiments, the C-terminal domain includes a zinc finger. In some embodiments, the zinc finger includes a CCHC motif.
[0014] In some embodiments, the enzyme includes bat-associated reverse transcriptase (BART). In some embodiments, BART includes BART from the horseshoe bat (Rhinolophus ferrum equinum), the rat-eared bat (Myotis myotis), the European badger (Meles meles), the wild yak (Bos mutus), the goat (Capra hircus), Homo sapiens (Homo sapiens), the domestic dog (Canis lupus familiaris), the wild boar (Susscrofa), the common marmoset (Callithrix jacchus), the house mouse (Mus musculus), the brown bear (Ursus arctos), the Indian elephant (Elephas maximus indicus), or variants thereof.
[0015] In some embodiments, the enzyme is provided via one or more polynucleotide molecules encoding the enzyme. In some embodiments, one or more nucleic acid components are provided via one or more polynucleotide molecules encoding or comprising one or more nucleic acid components. In some embodiments, the one or more polynucleotide molecules comprise one or more vectors. In some embodiments, the enzyme and one or more nucleic acid components are provided in a single vector.
[0016] In some embodiments, the target polynucleotide includes a genomic locus. In some embodiments, the target polynucleotide includes RNA or DNA.
[0017] In some embodiments, the RNA includes viral RNA of an RNA virus. In some embodiments, the RNA virus is selected from the group consisting of: Norwalk virus, Rotavirus, Poliovirus, Ebola virus, Marburg virus, Lassa virus, Hantavirus, Rabies virus, Influenza virus, Yellow fever virus, Coronavirus, SARS, SARS-CoV-2, West Nile virus, Hepatitis A, C (HCV) and E viruses, Dengue fever virus, Togavirus, Rhabdovirus, Picornavirus, Myxovirus, Retrovirus, Bunyavirus, Coronavirus and Reovirus.
[0018] In some implementations, the DNA includes genomic DNA or cDNA.
[0019] In some embodiments, modification of the target polynucleotide includes cleavage of the target polynucleotide. In some embodiments, the target polynucleotide is contained in a nucleic acid molecule, either intracellularly or in vitro.
[0020] In some embodiments, the cells include eukaryotic cells. In some embodiments, the eukaryotic cells include mammalian cells. In some embodiments, the eukaryotic cells include non-human animal cells, human cells, or plant cells.
[0021] In another aspect, this disclosure provides a method for treating or preventing viral infection of an RNA virus in a cell or subject. In some embodiments, the method includes delivering to a cell or subject an enzyme having reverse transcriptase and endonuclease activity, or a nucleic acid molecule encoding said enzyme, wherein a single-stranded guiding polynucleotide hybridizes with viral RNA of the RNA virus and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0022] In some implementations, the method includes reverse transcription of viral RNA into cDNA using an enzyme.
[0023] In some embodiments, the enzyme has integrase activity, and the method described therein includes integrating cDNA into the genome of a cell or subject via the enzyme.
[0024] In some implementations, the method includes transcribing cDNA into a single-stranded guide polynucleotide capable of hybridizing with viral RNA.
[0025] In another aspect, this disclosure provides a method for enhancing immunity against viral infection by an RNA virus in cells or a subject. In some embodiments, the method includes delivering to cells or a subject an enzyme or a nucleic acid molecule encoding said enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, wherein the enzyme reverse transcribes the mRNA into a second ssDNA, wherein the second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0026] Also within the scope of this disclosure is a method for enhancing immunity against viral infection by an RNA virus in cells or a subject. In some embodiments, the method includes delivering to cells or a subject an enzyme or a nucleic acid molecule encoding said enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, and wherein the mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
[0027] In another aspect, this disclosure provides a method for generating a cell line with immunity against viral infection by an RNA virus. In some embodiments, the method includes delivering an enzyme or a nucleic acid molecule encoding said enzyme to cells having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or a subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, wherein the enzyme reverse transcribes the mRNA into a second ssDNA, wherein the second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0028] Also within the scope of this disclosure is a method for generating a cell line with immunity against viral infection by an RNA virus. In some embodiments, the method includes delivering an enzyme or a nucleic acid molecule encoding said enzyme to cells having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, and wherein the mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
[0029] In some implementations, the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0030] In some embodiments, the enzyme includes an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain. In some embodiments, the enzyme further includes a C-terminal domain. In some embodiments, the C-terminal domain facilitates interaction between the enzyme and the target polynucleotide. In some embodiments, the C-terminal domain includes a zinc finger. In some embodiments, the zinc finger includes a CCHC motif.
[0031] In some embodiments, the enzyme includes bat-associated reverse transcriptase (BART). In some embodiments, BART includes BART from horseshoe bat, rat eared bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
[0032] In some implementations, the enzyme is provided via one or more polynucleotide molecules that encode the enzyme.
[0033] In another aspect, this disclosure provides a gene editing system for modifying target polynucleotides. In some embodiments, the method includes an enzyme having reverse transcriptase and endonuclease activities, and one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0034] In some implementations, the enzyme further possesses integrase activity.
[0035] In some implementations, one or more nucleic acid components include single-stranded navigation DNA.
[0036] In some embodiments, one or more nucleic acid components further include a payload RNA. In some embodiments, the payload RNA includes a stem-loop structure. In some embodiments, the enzyme reverse transcribes the payload RNA into cDNA. In some embodiments, the enzyme integrates the cDNA into a target polynucleotide.
[0037] In some embodiments, the enzyme includes an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain. In some embodiments, the enzyme further includes a C-terminal domain. In some embodiments, the C-terminal domain facilitates interaction between the enzyme and the target polynucleotide. In some embodiments, the C-terminal domain includes a zinc finger. In some embodiments, the zinc finger includes a CCHC motif.
[0038] In some embodiments, the enzyme includes bat-associated reverse transcriptase (BART). In some embodiments, BART includes BART from horseshoe bat, rat eared bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
[0039] In some embodiments, the gene editing system includes one or more polynucleotide molecules encoding an enzyme. In some embodiments, the gene editing system includes one or more polynucleotide molecules encoding or comprising one or more nucleic acid components. In some embodiments, the one or more polynucleotide molecules include one or more vectors. In some embodiments, the enzyme and one or more nucleic acid components are provided in a single vector.
[0040] In some implementations, the target polynucleotide includes RNA or DNA.
[0041] In some embodiments, RNA includes viral RNA of RNA viruses. In some embodiments, the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E viruses, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0042] In some implementations, the DNA includes genomic DNA or cDNA.
[0043] In some implementations, the modification of the target polynucleotide includes the cleavage of the target polynucleotide.
[0044] In another aspect, this disclosure provides a delivery system including a gene editing system, and said delivery system is adapted to deliver the gene editing system to cells or a subject. In some embodiments, the delivery system includes nanoparticles or vesicles encapsulating the gene editing system.
[0045] In another aspect, this disclosure provides a vector system comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding enzymes having reverse transcriptase activity, endonuclease activity and integrase activity, and comprise one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0046] In another respect, this disclosure provides a kit comprising a gene editing system, delivery system or vector system as described herein.
[0047] In another aspect, this disclosure provides a cell line comprising one or more polynucleotide molecules encoding an enzyme having reverse transcriptase activity, endonuclease activity and integrase activity, and one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0048] In some embodiments, the cell line comprises eukaryotic cells. In some embodiments, the eukaryotic cells comprise mammalian cells. In some embodiments, the eukaryotic cells comprise stem cells or stem cell lines.
[0049] A method for preparing cell lines is also provided. In some embodiments, the method includes introducing into cells an enzyme encoding reverse transcriptase activity, endonuclease activity, and integrase activity, as well as one or more polynucleotide molecules encoding one or more nucleic acid components.
[0050] The foregoing summary is not intended to limit every aspect of this disclosure, and other aspects are described in other sections such as the following detailed description. The entire document is intended to be associated with a unified disclosure, and it should be understood that all combinations of the features described herein are contemplated, even if such combinations of features do not appear together in the same sentence, paragraph, or section of this document. Other features and advantages of the invention will become apparent from the following detailed description. However, it should be understood that while the detailed description and specific embodiments indicate specific implementations of this disclosure, they are given by way of illustration only, as various changes and modifications within the spirit and scope of this disclosure will become apparent to those skilled in the art from these specific embodiments. Attached Figure Description
[0051] Figure 1 The process by which BART is used to integrate RNA payloads via navigation DNA is shown.
[0052] Figure 2A and Figure 2B This figure illustrates BART-mediated immunity against VSV infection. It shows the results of VSV infection in BART-expressing cell lines. Figure 2A The study showed a reduction in viral load in BART-expressing cells compared to control cells. Figure 2B The presence of viral cDNA in the cytoplasm of cells expressing BART was shown, highlighting the role of BART in reverse transcription and viral suppression.
[0053] Figure 3 The structural model of BART is shown, highlighting the APE domain, reverse transcriptase domain, and C-terminal domain (CTD).
[0054] Figure 4 The vector map of the bart cloning construct is shown. The plasmid vector map used to clone the BART enzyme coding sequence is depicted, including the arrangement of the CMV promoter, neomycin resistance gene, and HindIII and NotI restriction sites. Detailed Implementation
[0055] This disclosure is partly based on the unexpected discovery of a novel mammalian defense mechanism similar to CRISPR-Cas systems traditionally associated with prokaryotic biology. As disclosed herein, bat-associated reverse transcriptase (BART) possesses unique characteristics found in bat induced pluripotent stem cells (iPSCs) and the ability to convert viral genomic RNA into DNA. The disclosed BART system represents a next-generation tool for genome and RNA editing, exhibiting high efficiency and specificity, potentially surpassing current CRISPR-Cas systems. This discovery blurs the lines between prokaryotic and eukaryotic defense mechanisms, opening new opportunities for comprehensive viral defense strategies and advanced gene therapy applications. This disclosure also represents a significant advance in addressing the most pressing health challenges, including pandemics caused by RNA viruses and gene editing for therapeutic purposes.
[0056] BART exhibits a remarkable ability to convert viral RNA into DNA (a process typically associated with retroviruses). Furthermore, it possesses endonuclease activity that enables it to cleave nucleic acids and integrase activity that promotes the insertion of genetic material into the host genome. This unique combination of enzymatic activities in mammalian proteins makes BART a groundbreaking innovation with significant implications for both antiviral defense and gene editing.
[0057] BART functions as an innovative antiviral defense mechanism in mammalian cells. It achieves this by targeting and degrading viral RNA, thereby disrupting the viral life cycle and preventing further infection. A unique aspect of BART is its ability to create a “genomic memory” of past viral encounters. By converting viral RNA into DNA and integrating it into the host genome, BART establishes a persistent record of infection. This genomic archive allows for the rapid production of RNA transcripts that guide BART to recognize and neutralize the same virus in subsequent infections, thus providing a form of adaptive immunity in mammalian cells.
[0058] BART's unique mechanism of action distinguishes it from other known antiviral and gene-editing systems. Upon encountering viral RNA, BART uses its reverse transcriptase activity to convert the RNA into complementary DNA (cDNA). This cDNA is then integrated into the host genome at a specific location, creating a library of viral sequences similar to CRISPR arrays found in bacteria. The integrated cDNA is subsequently transcribed to generate RNA transcripts that guide BART to target and cleave homologous viral RNA, thereby disrupting the viral life cycle and providing long-term immunity against future infections. This remarkable process not only confers antiviral defense but also provides a powerful tool for precise gene editing, as specific modifications can be introduced into the host genome using integration and targeting mechanisms.
[0059] Beyond its role in antiviral defense, BART's unique mechanism of action offers a method for precise and effective gene editing. Its ability to introduce specific gene modifications at the DNA and RNA levels, guided by programmable ssDNA oligonucleotides, makes BART the successor to current CRISPR-Cas systems. BART exhibits superior specificity and reduced off-target effects, overcoming major limitations of existing gene editing technologies. This enhanced precision, coupled with its potential for broader applicability across different cell types and therapeutic settings, makes BART a system for next-generation gene editing tools. The ability to target the genome using this naturally evolved mammalian system will revolutionize the field of gene therapy, providing new avenues for treating genetic disorders and other diseases.
[0060] The discovery of BART represents a paradigm shift in the understanding of antiviral immunity and gene editing. By identifying a mammalian system functionally similar to the prokaryotic CRISPR-Cas system, BART challenges the traditional boundaries between these two realms of life. This demonstrates the convergent evolution of adaptive immune mechanisms, highlighting the universality of strategies employed by organisms to combat viral threats. BART's dual function as an antiviral agent and a precise gene-editing tool opens new avenues for innovative therapies against viral infections and genetic diseases. This naturally evolved mammalian system for targeted genome manipulation will revolutionize the field of gene therapy, providing a more compatible and effective alternative to current CRISPR-Cas-based approaches. Therefore, the discovery of BART represents a major leap forward in addressing some of the most pressing health challenges, paving the way for a new era of antiviral and gene-editing therapies.
[0061] BART System and Usage
[0062] This disclosure covers the BART system described herein for modifying target DNA sequences (e.g., chromosomal sequences) or target RNA sequences, such as for methods and uses to alter or manipulate the expression of one or more genes or gene products in prokaryotic or eukaryotic cells, in vitro, in vivo, or ex vivo. The BART protein possesses unique characteristics in that it exhibits both reverse transcriptase and endonuclease activities. In some embodiments, the BART protein may additionally possess integrase activity, thereby allowing it to integrate DNA (e.g., DNA reverse transcribed from viral RNA) into the host genome.
[0063] Publicly available BART systems offer an efficient means of modifying (e.g., deletion, insertion, translocation, inactivation, activation) target RNA or DNA (double-stranded, linear, or supercoiled) in a variety of cell types. Therefore, publicly available BART systems have broad applications, such as gene therapy, drug screening, and disease diagnosis / prognosis.
[0064] Methods of modifying target polynucleotides
[0065] In one aspect, this disclosure provides a method for modifying a target polynucleotide. In some embodiments, the method includes delivering an enzyme (e.g., BART) having reverse transcriptase and endonuclease activity, and one or more nucleic acid components to the target polynucleotide, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0066] In some implementations, the enzyme further possesses integrase activity.
[0067] In some embodiments, the enzyme includes a BART protein. In some embodiments, one or more nucleic acid components include single-stranded navigation DNA (“navigation ssDNA”).
[0068] In some embodiments, one or more nucleic acid components further include a payload RNA. In some embodiments, the payload RNA includes a stem-loop structure.
[0069] In some embodiments, the enzyme reverse transcribes the payload RNA into cDNA. In other embodiments, the enzyme integrates the cDNA into the target polynucleotide.
[0070] On the other hand, this disclosure provides a method for modifying the expression of a target polynucleotide (e.g., a target sequence of interest) in a cell. In some embodiments, the method allows a BART complex (e.g., a BART / navigation ssDNA complex) to bind to the target polynucleotide, resulting in increased or decreased expression of the target polynucleotide or a gene containing the target polynucleotide. In some embodiments, the BART complex comprises BART complexed with an ssDNA sequence that hybridizes to the target sequence within the polynucleotide.
[0071] In some embodiments, methods for modifying the target polynucleotide include delivering the BART system, isolated nucleic acids encoding the BART protein and / or navigation ssDNA, or particles containing the BART protein and / or navigation ssDNA to the target sequence or to a cell containing the target sequence. In some embodiments, after the formation of the BART / navigation ssDNA complex and hybridization of the ssDNA with one or more nucleic acids of the target sequence, the BART protein induces modification (e.g., cleavage) of the target sequence.
[0072] In some embodiments, the modification includes cleaving one or both strands at the location of the target sequence via an enzyme. In some embodiments, the modification results in reduced or increased transcription of the target gene. In some embodiments, the method further includes repairing the cleaved target polynucleotide by homologous recombination with a foreign template polynucleotide, wherein the repair results in a mutation of insertion, deletion, or substitution of one or more nucleotides comprising the target polynucleotide. In some embodiments, the mutation results in an alteration of one or more amino acids in a protein expressed by a gene comprising the target sequence.
[0073] In some implementations, the modification of the target polynucleotide includes the cleavage of the target polynucleotide.
[0074] In some implementations, modification of the target polynucleotide can include modification (e.g., increasing or decreasing) of the target polynucleotide's expression due to the binding of the BART protein to a specific site on the target polynucleotide. For example, if the BART protein binds to a site between the promoter or other regulatory element and the coding sequence of the target polynucleotide, the BART protein can act as a repressor that decreases the expression of the target polynucleotide. On the other hand, if the BART protein binds to a site upstream of the promoter or other regulatory element, the BART protein can act as an activator that increases the expression of the target polynucleotide.
[0075] Therefore, in one aspect, this disclosure also provides a method for modifying the expression of a target polynucleotide. In some embodiments, the method includes introducing into a cell or subject an enzyme (e.g., BART) having reverse transcriptase and endonuclease activity, or a nucleic acid molecule encoding said enzyme, and one or more nucleic acid components (e.g., navigation ssDNA), wherein said one or more nucleic acid components target and hybridize with the target polynucleotide and direct the binding of the enzyme to the target polynucleotide, and wherein said enzyme binds to one or more sites on the target polynucleotide such that the binding of the enzyme increases or decreases the expression level of the target polynucleotide.
[0076] In this application, the enzyme may include one or more mutations that result in an enzyme with no catalytic activity. In some embodiments, the enzyme with no catalytic activity lacks endonuclease activity, such that the enzyme can still bind to the target polynucleotide but does not cleave the target polynucleotide.
[0077] In some embodiments, the nucleotide sequence encoding the BART protein is present in the recombinant expression vector. In some embodiments, the recombinant expression vector is a viral construct, such as a recombinant adeno-associated virus construct, a recombinant adenovirus construct, a recombinant lentivirus construct, etc. For example, the viral vector can be based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, and human immunodeficiency virus, etc. Retroviral vectors can be based on murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rouss sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloid proliferative sarcoma virus, and mammary tumor virus, etc. Useful expression vectors are known to those skilled in the art, and many are commercially available. Examples of vectors provided as eukaryotic host cells include pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40. However, any other vector can be used if it is compatible with the host cell. For example, useful expression vectors containing nucleotide sequences encoding the Cas9 enzyme are commercially available from companies such as Addgene, Life Technologies, Sigma-Aldrich, and Origene.
[0078] Depending on the target cell / expression system used, any of the many transcriptional and translational control elements, including promoters, transcription enhancers, and transcription terminators, can be used in the expression vector. Useful promoters can be derived from viruses or any organism, such as prokaryotes or eukaryotes. Suitable promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoters; cytomegalovirus (CMV) promoters such as the CMV immediate early promoter region (CMVIE), Rouss' sarcoma virus (RSV) promoter, the human U6 small nucleus promoter (U6), the enhanced U6 promoter, and the human H1 promoter (H1), etc.
[0079] In some embodiments, the enzyme is provided via one or more polynucleotide molecules encoding the enzyme. In some embodiments, one or more nucleic acid components are provided via one or more polynucleotide molecules encoding or comprising one or more nucleic acid components. In some embodiments, the one or more polynucleotide molecules comprise one or more vectors. In some embodiments, the enzyme and one or more nucleic acid components are provided in a single vector.
[0080] In some embodiments, the polynucleotide including the first navigation ssDNA and the polynucleotide encoding the BART protein are located on the same vector. In some embodiments, the BART system further includes a polynucleotide containing a second navigation ssDNA. In some embodiments, the polynucleotide encoding the BART protein, the polynucleotide including the first navigation DNA, and the polynucleotide including the second navigation ssDNA are present on the same vector or on two or more different vectors.
[0081] BART protein or its variants / fragments can be introduced into cells (e.g., in vitro cells such as primary cells for ex vivo therapy, or in vivo cells such as in a patient) as BART protein or its variants or fragments, mRNA encoding BART protein or its variants or fragments, or recombinant expression vectors containing nucleotide sequences encoding BART protein or its variants or fragments.
[0082] In some embodiments, the method further includes delivering one or more vectors to host cells. In some embodiments, the vector is delivered to the subject's host cells. In some embodiments, the modification occurs in eukaryotic cells in cell culture. In some embodiments, the method further includes isolating eukaryotic cells from the subject prior to modification. In some embodiments, the method further includes returning cells derived therefrom to the subject.
[0083] The method further includes maintaining cells or embryos under suitable conditions such that ssDNA guides the BART protein to a target site in the target sequence to modify the target sequence. Typically, cells can be maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art and are described, for example, in the following references: “Current Protocols in Molecular Biology” Ausubel et al., John Wiley & Sons, New York, 2003; or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, (2001); Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306. Those skilled in the art will understand that the methods used for culturing cells are known in the art and can and will vary depending on the cell type. In all cases, routine optimization can be used to determine the optimal technique for a particular cell type.
[0084] Embryos can be cultured in vitro (e.g., in cell culture). Typically, embryos are cultured at appropriate temperatures and in suitable media with the necessary O2 / CO2 ratio to allow for the expression of protein and RNA scaffolds (if necessary). Suitable, non-limiting examples of media include M2, M16, KSOM, BMOC, and HTF media. Those skilled in the art will understand that culture conditions can and will be varied depending on the species of the embryo. In all cases, routine optimization can be used to determine the optimal culture conditions for embryos of a particular species. In some cases, cell lines may be derived from embryos cultured in vitro (e.g., embryonic stem cell lines).
[0085] Alternatively, embryos can be cultured in vivo by transferring them into the uterus of a female host. Generally, the female host is from the same or similar species as the embryo. Preferably, the female host is pseudopregnant. Methods for preparing pseudopregnant female hosts are known in the art. Additionally, methods for transferring embryos into female hosts are known. In vivo embryo culture allows for embryonic development and can result in live births in animals derived from the embryo. Such animals will include modified chromosome sequences in every cell of their body.
[0086] BART protein
[0087] In some embodiments, the BART protein may include variants or fragments of BART. In some embodiments, BART includes BART from horseshoe bats, rat-eared bats, European badgers, wild yaks, goats, Homo sapiens, domestic dogs, wild boars, common marmosets, house mice, brown bears, Indian elephants, or variants thereof.
[0088] In some embodiments, the enzyme includes an N-terminal depurinyl-depyrimidine endonuclease (APE) domain optionally linked to a reverse transcriptase domain. The depurinyl-depyrimidine endonuclease (APE) domain is a protein motif that plays a crucial role in DNA repair. These domains are responsible for recognizing and cleaving DNA strands containing depurinyl or depyrimidine sites, which are damage sites that occur when purine or pyrimidine bases are lost from the DNA backbone.
[0089] In some embodiments, the enzyme further includes a C-terminal domain. In some embodiments, the C-terminal domain facilitates the interaction between the enzyme and the target polynucleotide. In some embodiments, the C-terminal domain includes a zinc finger. A zinc finger is characterized by one or more zinc ions (Zn0.05). 2+ These metal ions coordinate with small protein structural motifs. They help stabilize protein structure and facilitate interactions with other molecules.
[0090] In some implementations, the zinc finger includes a CCHC motif. CCHC motifs are zinc finger motifs found in many proteins, particularly those involved in DNA binding and transcriptional regulation. They consist of a conserved sequence of amino acids containing cysteine and histidine residues. These residues coordinate with zinc ions to form a stable structure that facilitates protein-DNA binding.
[0091] In some embodiments, the BART protein includes an amino acid sequence of any of SEQ ID NO:1-12 or includes an amino acid sequence that has at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:1-12.
[0092] Table 1. Representative BART proteins
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102] The terms “protein,” “polypeptide,” and “peptide” are used interchangeably herein to refer to polymers of amino acids of any length. Polymers may be linear or branched, may include modified amino acids, and may be interrupted by non-amino acid components. The term also covers polymers of modified amino acids, such modifications as disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, polyethylene glycolation, or any other manipulation, such as conjugation with a labeled component. As used herein, the term “amino acid” includes natural and / or non-natural or synthetic amino acids, including glycine and its D or L optical isomers, as well as amino acid analogs and peptide mimics.
[0103] As used herein, a peptide or polypeptide “fragment” refers to a peptide, polypeptide, or protein that is shorter than its full length. For example, a peptide or polypeptide fragment may have a length of at least about 3, at least about 4, at least about 5, at least about 10, at least about 20, at least about 30, at least about 40 amino acids, or a single unit length thereof. For example, the length of a fragment may be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or more amino acids. There is no upper limit to the size of a peptide fragment. However, in some embodiments, the length of a peptide fragment may be less than about 500 amino acids, less than about 400 amino acids, less than about 300 amino acids, or less than about 250 amino acids.
[0104] As used herein, the term "variant" refers to a first composition (e.g., a first molecule) associated with a second composition (e.g., a second molecule, also referred to as a "parent" molecule). Variant molecules may be derived from, isolated from, based on, or homologous to the parent molecule. The term variant may be used to describe polynucleotides or polypeptides.
[0105] When applied to polynucleotides, variant molecules may have whole nucleotide sequence identity with the original parent molecule, or alternatively, may have less than 100% nucleotide sequence identity with the parent molecule. For example, a variant of a gene nucleotide sequence may be a second nucleotide sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99% or more identical in nucleotide sequence to the original nucleotide sequence. Polynucleotide variants also include polynucleotides comprising the whole parent polynucleotide and further comprising additional fusion nucleotide sequences. Polynucleotide variants also include polynucleotides that are portions or subsequences of the parent polynucleotide; for example, the invention also covers unique subsequences of the polynucleotides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques).
[0106] On the other hand, polynucleotide variants include nucleotide sequences comprising minor, insignificant, or negligible alterations to the parental nucleotide sequence. For example, minor, insignificant, or negligible alterations include changes to the nucleotide sequence that (i) do not change the amino acid sequence of the corresponding polypeptide, (ii) occur outside the protein-coding open reading frame of the polynucleotide, (iii) result in deletions or insertions that may affect the corresponding amino acid sequence but have little or no effect on the biological activity of the polypeptide, and (iv) result in the substitution of an amino acid by a chemically similar amino acid. Where the polynucleotide does not encode a protein, the variant of the polynucleotide may include nucleotide alterations that do not result in loss of function of the polynucleotide. On the other hand, the present invention covers conserved variants of the disclosed nucleotide sequences that produce functionally identical nucleotide sequences. Those skilled in the art will understand that the present invention covers many variants of the disclosed nucleotide sequences.
[0107] When applied to proteins, the variant polypeptide can have whole amino acid sequence identity with the original parent polypeptide, or alternatively, it can have less than 100% amino acid identity with the parent protein. For example, the variant amino acid sequence can be a second amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99% or more identical in amino acid sequence to the original amino acid sequence.
[0108] Peptide variants include peptides comprising the entire parent peptide and further comprising additional fusion amino acid sequences. Peptide variants also include peptides that are portions or subsequences of the parent peptide; for example, the present invention also covers unique subsequences of the peptides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques).
[0109] On the other hand, peptide variants include peptides comprising minor, negligible, or insignificant changes to the parental amino acid sequence. For example, minor, negligible, or insignificant changes include amino acid alterations (including substitutions, deletions, and insertions) that have little or no effect on the biological activity of the peptide and produce a functionally identical peptide, including the addition of a non-functional peptide sequence. In other aspects, the variant peptides of the present invention alter the biological activity of the parent molecule. Those skilled in the art will understand that the present invention covers many variants of the disclosed peptides.
[0110] In some implementations, polynucleotide or polypeptide variants may include variant molecules in which a small percentage of nucleotide or amino acid positions have been altered, added, or deleted, typically less than about 10%, less than about 5%, less than 4%, less than 2%, or less than 1%.
[0111] As used herein, a “functional variant” of a protein refers to a variant of such a protein that retains at least part of its activity. Functional variants can include mutants (which can be insertion, deletion, or substitution mutants), including polymorphs, etc. Functional variants also include fusion products of such proteins with another normally unrelated nucleic acid, protein, polypeptide, or peptide. Functional variants can be naturally occurring or artificial.
[0112] In some implementations, variants of the BART protein may include one or more conserved modifications. BART protein variants with one or more conserved modifications may retain the desired functional properties, which can be tested using functional assays known in the art.
[0113] As used herein, the term "conserved sequence modification" refers to an amino acid modification that does not significantly affect or alter the binding properties of a protein containing an amino acid sequence. Such conserved modifications include amino acid substitutions, additions, and deletions. Modifications can be introduced using standard techniques known in the art, such as site-directed mutagenesis and PCR-mediated mutagenesis. A conserved amino acid substitution is an amino acid substitution in which an amino acid residue is replaced by an amino acid residue having a similar side chain. Families of amino acid residues with similar side chains have been defined in the art. These families include: amino acids with basic side chains (e.g., lysine, arginine, histidine); amino acids with acidic side chains (e.g., aspartic acid, glutamic acid); amino acids with uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan); amino acids with nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine); amino acids with β-branched side chains (e.g., threonine, valine, isoleucine); and amino acids with aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine), which include one or more conserved modifications. BART proteins with one or more conserved modifications can retain desired functional properties, which can be tested using functional assays known in the art.
[0114] As used herein, the percentage homology between two amino acid sequences is equivalent to the percentage identity between the two sequences. The percentage identity between two sequences is a function of the number of common positions shared by the sequences (i.e., % homology = # of common positions / total number of positions # x 100), taking into account the number of vacancies introduced for optimal alignment of the two sequences and the length of each vacancies. Sequence comparison and determination of the percentage identity between two sequences can be accomplished using mathematical algorithms, as described in the non-limiting examples below.
[0115] The percentage identity between two amino acid sequences can be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4:11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weighted residue table, a vacancy length penalty of 12, and a vacancy penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the algorithm of Needleman and Wunsch (J. Mol. Biol. 48:444-453 (1970)), which has been incorporated into the GAP program of the GCG software package (available at www.gcg.com), using a Blossum62 matrix or a PAM250 matrix, and vacancy weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.
[0116] Alternatively or additionally, the protein sequences of the present invention can be further used as “query sequences” to search public databases, thereby identifying, for example, relevant sequences. Such searches can be performed using the XBLAST program (version 2.0) as described in Altschul et al. (1990) J. Mol. Biol. 215:403-10. BLAST protein searches can be performed using the XBLAST program with a score of 50 and a word length of 3 to obtain amino acid sequences homologous to the molecules of the present invention. For obtaining vacancy alignments for comparative purposes, Gapped BLAST as described in Altschul et al. (1997) Nucleic Acids Res. 25(17):3389-3402 can be used. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used (see www.ncbi.nlm.nih.gov).
[0117] In some embodiments, variants of the BART protein may be conjugated to or linked to a detectable tag or detectable marker (e.g., a radionuclide, fluorescent dye, or MRI-detectable marker). In some embodiments, the detectable tag may be an affinity tag. As used herein, the term "affinity tag" refers to a portion linked to a polypeptide that allows the polypeptide to be purified from a biochemical mixture. Affinity tags may consist of an amino acid sequence or may include an amino acid sequence linked to a chemical group through post-translational modification. Non-limiting examples of affinity tags include His tags, CBP tags (CBP: calmodulin-binding protein), CYD tags (CYD: covalently but dissociable NorpD peptide), Strep tags, StrepII tags, FLAG-tags, HPC tags (HPC: heavy chain of protein C), GST-tags (GST: glutathione S-transferase), Avi tags, biotinylated tags, Myc tags, myc-myc-hexahistidine (mmh) tags, 3xFLAG tags, SUMO tags, and MBP tags (MBP: maltose-binding protein). Further examples of affinity tags can be found in Kimple et al., Curr Protoc Protein Sci. 2013 Sep 24; 73: Unit 9.9.
[0118] In some embodiments, the detectable tag may be conjugated to or linked to the N- and / or C-terminus of a variant of the BART protein. The detectable tag and affinity tag may also be separated by one or more amino acids. In some embodiments, the detectable tag may be conjugated to or linked to the variant via a cleavable element. In the context of this invention, the term "cleavable element" refers to a peptide sequence that is readily cleaved by chemical reagents or enzymatic means such as proteases. The protease may be sequence-specific (e.g., thrombin) or may have limited sequence specificity (e.g., trypsin). Cleavable elements I and II may also be included in the amino acid sequence of the detection tag or polypeptide, particularly where the last amino acid of the detection tag or polypeptide is K or R.
[0119] As used herein, the terms “conjugate,” “conjugate,” or “link” refer to the connection of two or more entities to form a single entity. Conjugates encompass both peptide-small molecule conjugates and peptide-protein / peptide conjugates.
[0120] The terms "fusion polypeptide" or "fusion protein" refer to a protein produced by joining two or more polypeptide sequences together. Fusion polypeptides covered by this invention include the translational product of a chimeric gene construct that joins a nucleic acid sequence encoding a first polypeptide with a nucleic acid sequence encoding a second polypeptide to form a single open reading frame. In other words, a "fusion polypeptide" or "fusion protein" is a recombinant protein of two or more proteins joined by peptide bonds or via several peptides. Fusion proteins may also include a peptide linker between two domains.
[0121] The term "connector" refers to any means, entity, or portion used to join two or more entities. A connector can be covalent or non-covalent. Examples of covalent connectors include connector portions covalently or covalently linked to one or more proteins or domains to be joined. Connectors can also be non-covalent, for example, organometallic bonds through a metal center such as a platinum atom. For covalent bonding, various functionalities can be used, such as amide groups, including carbonate derivatives, ethers, esters (including organic and inorganic esters), amino groups, carbamates, and ureas. To provide a connection, the domain can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide a coupling site. Connecting methods are well known to those skilled in the art and are covered in this invention for use. Connector portions include, but are not limited to, chemical connector portions, or, for example, peptide connector portions (connector sequences).
[0122] In some embodiments, the linker can be a peptide linker or a non-peptide linker. Examples of peptide linkers can include [Ser(Gly)n]m or [Ser(Gly)n]mSer, where n can be an integer between 1 and 20. As used herein, the term "non-peptide linker" refers to a biocompatible polymer consisting of two or more repeating units linked to each other by any non-peptide covalent bond. Such non-peptide linkers can have two or three ends. Examples of non-peptide linkers can include, but are not limited to, polyethylene glycol, polypropylene glycol, copolymers of ethylene glycol and propylene glycol, polyoxyethylene polyols, polyvinyl alcohol, polysaccharides, dextran, polyvinyl ether, biodegradable polymers such as polylactic acid (PLA) and polylactic-glycolic acid (PLGA), lipid polymers, chitin, hyaluronic acid, and combinations thereof. Aptamers can be added as non-peptide linkers.
[0123] In some embodiments, variants of the BART protein can be fused to a fusion partner by crosslinking with a crosslinking agent, such as a crosslinking agent. A crosslinking agent is a reagent having a reactive terminus on a specific functional group (e.g., a primary amine or a thiol group) on a protein or other molecule. A crosslinking agent is capable of covalently bonding two or more molecules. Crosslinking agents include, but are not limited to, amine-p-amine crosslinking agents (e.g., disuccinyl succinyl ester (DSS)), amine-p-thiol crosslinking agents (e.g., N-γ-maleimide butyryl-oxosuccinyl ester (GMBS)), carboxyl-p-amine crosslinking agents (e.g., dicyclohexylcarbodiimide (DCC)), thiol-p-carbohydrate crosslinking agents (e.g., N-β-maleimide propionic acid hydrazide (BMPH)), thiol-p-thiol crosslinking agents (e.g., 1,4-bismaleimide butane (BMB)), photoreactive crosslinking agents (e.g., N-5-azido-2-nitrobenzoyloxysuccinylimide (ANB-NOS)), and chemically selectively linked crosslinking agents (e.g., NHS-PEG4-azide).
[0124] In some implementations, the polynucleotide encoding the BART protein can be codon-optimized. Generally, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in the host cell by replacing at least one codon of the native sequence with a codon that is more commonly or most frequently used in the gene in the host cell while maintaining the native amino acid sequence. Various species exhibit specific preferences for certain codons of specific amino acids. Codon preference (differences in codon use between organisms) is generally associated with the translation efficiency of messenger RNA (mRNA), which is further considered to depend (among other things) on the nature of the codons being translated and the availability of a particular transfer RNA (tRNA) molecule. The dominance of the selected tRNA in a cell is generally a reflection of the most frequently used codons in peptide synthesis. Therefore, genes can be customized for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, the "Codon Usage Database" available at www.kazusa.orjp / codon / , and these tables can be modified in various ways. See Nakamura, Y. et al. “Codon usage tabulated from the international DNA sequence databases: status for the year 2000” Nucl.Acids Res.28:292 (2000). Computer algorithms for codon optimization of specific sequences expressed in specific host cells are also available, such as Gene Forge (Aptagen; Jacobus, Pa.). In some implementations, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in the DNA / RNA sequence encoding the target BART protein correspond to the most commonly used codon for a specific amino acid. For information on codon usage in yeast, refer to the online yeast genome database available at http: / / www.yeastgenome.org / community / codonusage.shtml, or Codon selection in yeast, Bennetzen and Hall, J Biol Chem. 1982 Mar. 25; 257(6):3026-31.Regarding codon usage in plants, including algae, see Codon usage in higher plants, green algae, and cyanobacteria, Campbell and Gowri, Plant Physiol. Jan 1990; 92(1): 1-11.; and Codon usage in plant genes, Murray et al., Nucleic Acids Res. Jan 25 1989; 17(2):477-98; or Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages, Morton BR, J Mol Evol. Apr 1998; 46(4):449-59.
[0125] Navigation ssDNA
[0126] As used herein, the term "navigation ssDNA" generally refers to an ssDNA molecule (or a group of DNA molecules) that can bind to the BART protein and target the BART protein to a specific location within the target RNA or DNA. The targeting fragment includes a nucleotide sequence that is complementary to (or at least can hybridize with) the target sequence under stringent conditions.
[0127] The navigation ssDNA can be any polynucleotide sequence that is sufficiently complementary to the target polynucleotide sequence (e.g., the target DNA or RNA sequence) to hybridize with the target sequence and guide the BART complex to bind sequence-specifically to the target sequence. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the guide sequence of the navigation ssDNA and its corresponding target sequence is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. The optimal alignment can be determined by using any suitable algorithm for aligning sequences. Non-limiting examples of such algorithms include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., Burrows-WheelerAligner), ClustalW, ClustalX, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).
[0128] In some implementations, the length of the navigation ssDNA sequence is approximately 3, 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230. 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500 or more nucleotides. In some embodiments, the guide sequence is about 50 nucleotides in length. In some embodiments, the guide sequence is about 100 nucleotides long. In some embodiments, the guide sequence is about 150 nucleotides long. In some embodiments, the guide sequence is about 200 nucleotides long. In some embodiments, the guide sequence is about 250 nucleotides long. In some embodiments, the guide sequence is about 300 nucleotides long.
[0129] The ability of a guide sequence to direct the sequence-specific binding of the BART complex to a target sequence can be assessed by any suitable assay. For example, components of a BART system sufficient to form the BART complex (including the guide sequence to be tested) can be provided to host cells with the corresponding target sequence, such as by transfection with a vector encoding a component of the BART sequence, and then the preferential cleavage within the target sequence can be assessed. Similarly, the cleavage of the target polynucleotide sequence can be evaluated in vitro by providing the target sequence, components of the BART complex (including the guide sequence to be tested and a control guide sequence different from the test guide sequence), and comparing the binding or cleavage rates at the target sequence between the test and control guide sequence responses.
[0130] Table 2. Representative ssDNA
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138] In some embodiments, the navigation ssDNA includes the polynucleotide sequences of SEQ ID NO:13-96 and 107-112 or includes a polynucleotide sequence having at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:13-96 and 107-112.
[0139] In some embodiments, the navigation ssDNA includes a synthetic nucleic acid sequence (e.g., a synthetic DNA molecule). In some embodiments, the navigation ssDNA includes one or more modifications.
[0140] As used herein, the term “modification” in the context of oligonucleotides or polynucleotides includes, but is not limited to, (a) terminal modifications, such as 5' or 3' terminal modifications, (b) nucleobase (or “base”) modifications, including base substitution or removal, (c) sugar modifications, including modifications at the 2', 3', and / or 4' positions, and (d) backbone modifications, including modifications or substitutions of phosphodiester bonds. The term “modified nucleotide” generally refers to a nucleotide having modifications to one or more of the chemical structures of bases, sugars, and phosphodiester bonds or backbone portions, including nucleotide phosphates. The terms “Z” and “P” refer, for example, to the nucleotide, nucleobase, or nucleobase analogue described in Yang, Z. et al., Nucleic Acids Res., 34, 6095-101 (2006), the disclosure of which is incorporated herein by reference in its entirety.
[0141] In some embodiments, one or more modifications may include a 2'-O-methyl moiety, a Z base, a 2'-deoxynucleotide, an intermolecular phosphate thioester bond, an intermolecular phosphonoacetate (PACE) bond, an intermolecular thiophosphonoacetate (thioPACE) bond, or a combination thereof. In some embodiments, one or more modifications include one or more modifications selected from the group consisting of: a 2'-O-methyl nucleotide having a 3'-phosphate thioester group, a 2'-O-methyl nucleotide having a 3'-phosphonoacetate group, a 2'-O-methyl nucleotide having a 3'-phosphonoacetate group, or a 2'-deoxy nucleotide having a 3'-phosphonoacetate group. In some embodiments, one or more modifications include 2-thiouracil (2-thioU), 4-thiouracil (4-thioU), 2-aminoadenine, 2'-o-methyl, 2'-fluoro, 5-methyluridine, 5-methylcytidine, or locked nucleic acid (LNA) modifications.
[0142] Target polynucleotides
[0143] In the context of BART complex formation, a "target polynucleotide" or "target sequence" refers to a sequence to which the guide sequence is designed to be complementary, wherein hybridization between the target sequence and the guide sequence promotes the formation of the BART complex. The target sequence may include RNA or DNA polynucleotides. A "target nucleic acid strand" refers to a strand of target nucleic acid that base-pairs with the navigation ssDNA disclosed herein. That is, the strand of target nucleic acid that hybridizes with the guide sequence is called the "target nucleic acid strand." The other strand of the target nucleic acid that is not complementary to the guide sequence is called the "non-complementary strand." In the case of double-stranded target nucleic acids (e.g., DNA), each strand can be a "target nucleic acid strand" to design the navigation ssDNA and to practice the disclosed methods.
[0144] As used herein, the term "target RNA" refers to an RNA polynucleotide that is or includes a target sequence. In other words, a target RNA can be an RNA polynucleotide or a portion of an RNA polynucleotide, a portion of a gRNA (i.e., a guide sequence) designed to be complementary to it, and an effector function mediated by a complex including the BART protein and the gRNA will be directed thereto. In some embodiments, the target sequence is located in the cell nucleus or cytoplasm.
[0145] As used herein, the term "target DNA" refers to a DNA polynucleotide that is or includes a target sequence. In other words, the target DNA may be a DNA polynucleotide or a portion thereof, and a portion thereof (i.e., a guide sequence) of a navigation ssRNA is designed to be complementary to the target DNA and to guide effector functions mediated by a complex comprising the BART protein and the navigation ssRNA to the target DNA. In some embodiments, the target sequence is located in the cell nucleus or cytoplasm.
[0146] Target polynucleotides are not sequence-restricted and can be located in coding regions of genes, in introns of genes, or in control regions between genes. In some embodiments, target polynucleotides are contained in nucleic acid molecules, either intracellularly or in vitro. Genes can be coding or non-coding. Target polynucleotides can be any polynucleotide, endogenous or exogenous within the cell. For example, target polynucleotides can be polynucleotides present in the nucleus of eukaryotic cells. Target polynucleotides can be sequences encoding gene products (e.g., proteins) or non-coding sequences (e.g., regulatory polynucleotides).
[0147] In some embodiments, the target polynucleotide includes RNA or DNA. In some embodiments, the RNA includes viral RNA of an RNA virus. In some embodiments, the target polynucleotide includes cDNA reverse-transcribed from viral RNA of an RNA virus.
[0148] In some implementations, the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0149] In some implementations, the DNA includes genomic DNA or cDNA.
[0150] In some embodiments, the target polynucleotide includes the polynucleotide sequences of SEQ ID NO:105, 106 and 113-118 or includes a polynucleotide sequence having at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:105, 106 and 113-118.
[0151] In some embodiments, the BART protein forms a complex with the navigation ssDNA that binds to the target polynucleotide. In some embodiments, the BART-navigation ssDNA complex comprises the BART protein having the amino acid sequence of SEQ ID NO:1 and the navigation ssDNA comprising a polynucleotide sequence having at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:13-96 and 107-112. In some embodiments, the BART-navigation ssDNA complex comprises a BART protein having an amino acid sequence having at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:1-12 and a navigation ssDNA comprising the polynucleotide sequences of SEQ ID NO:13-96 and 107-112.
[0152] In some embodiments, the BART-navigation ssDNA complex comprises a BART protein having the amino acid sequence of SEQ ID NO:1 and a navigation ssDNA having the polynucleotide sequences of SEQ ID NO:13-96 and 107-112.
[0153] In some embodiments, the BART protein forms a complex with a navigation ssDNA operatively linked to the payload RNA. In some embodiments, the payload RNA comprises a polynucleotide sequence of SEQ ID NO:119-133 or comprises a polynucleotide sequence having at least 75% (e.g., 75%, 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with any of SEQ ID NO:119-133.
[0154] Treatment
[0155] The BART system, one or more polynucleotides, or vectors or delivery systems described above can be used in therapeutic treatments. Therapeutic treatments may include gene or genome editing, or gene therapy. In one aspect, this disclosure provides a method of treating a subject in need, comprising inducing gene editing by introducing a polynucleotide or any vector as described herein into the subject's cells. In some embodiments, the method comprises inducing transcriptional activation or repression by introducing a polynucleotide or any vector as described herein into the subject's cells.
[0156] In another aspect, this disclosure further provides a method for treating a disease in a subject caused by a gene defect in a target sequence. The method includes administering the above-described system or composition to cells containing the target sequence in a subject in need, thereby inducing modification of the target sequence. In some embodiments, the target sequence is located at a genomic locus of interest. In some embodiments, the target sequence is part of a gene, and the modification in the target sequence regulates the gene's expression level. In some embodiments, the modification in the target sequence reduces the gene's expression level.
[0157] In some implementations, the method includes reverse transcription of viral RNA into cDNA using an enzyme.
[0158] In some embodiments, the enzyme has integrase activity, and the method described therein includes integrating cDNA into the genome of a cell or subject via the enzyme.
[0159] In some implementations, the method includes transcribing cDNA into a single-stranded guide polynucleotide capable of hybridizing with viral RNA.
[0160] In another aspect, this disclosure provides a method for enhancing immunity against viral infection by an RNA virus in cells or a subject. In some embodiments, the method includes delivering to cells or a subject an enzyme or a nucleic acid molecule encoding said enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, wherein the enzyme reverse transcribes the mRNA into a second ssDNA, wherein the second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0161] As used herein, “treating” or “treatment” for a disease in a subject means (1) preventing the occurrence of symptoms or disease in a subject who is susceptible to the disease or who has not yet shown symptoms of the disease; (2) suppressing the disease or preventing its development; or (3) improving or causing the remission of the disease or its symptoms. As understood in the art, “treatment” is a method for obtaining a beneficial or desired outcome, including clinical outcomes. For the purposes of this art, a beneficial or desired outcome may include, but is not limited to, one or more, reductions or improvements in one or more symptoms, a decrease in the severity of a condition (including disease), a stable (i.e., non-deterioration) state of a condition (including disease), a delay or slowing of a condition (including disease), a progression of a condition (including disease), improvement or mitigation, a state of remission (whether partial or complete), whether detectable or undetectable. In one aspect, the term “treatment” does not include prevention.
[0162] The terms “subject,” “individual,” and “patient” are used interchangeably in this document and refer to vertebrates, such as mammals, including humans. Mammals include, but are not limited to, rodents, apes, humans, livestock, sport animals, and pets. It also encompasses tissues, cells, and their progeny from biological entities obtained in vivo or cultured in vitro.
[0163] Many devastating human diseases share a common cause: genetic alterations or mutations. Pathogenic mutations in patients are inherited from their parents or caused by environmental factors. These diseases include, but are not limited to, the following categories. First, some genetic disorders are caused by germline mutations. One example is cystic fibrosis, caused by a mutation in the CFTR gene inherited from the parent. A second repressor mutation in CFTR can partially restore the function of the CFTR protein in somatic tissues. Other examples of genetic diseases caused by point gene mutations that can be corrected using the techniques disclosed include Gaucher disease, alpha-trypsin deficiency, and sickle cell anemia, to name just a few. Second, some diseases, such as chronic viral infections, are caused by exogenous environmental factors that lead to genetic alterations. One example is AIDS, caused by the insertion of the human HIV genome into the genome of infected T cells. Third, some neurodegenerative diseases involve genetic alterations. One example is Huntington's disease, caused by the amplification of the CAG trinucleotide in the huntingtin protein gene of affected patients. Finally, cancer is caused by various somatic mutations that accumulate in cancer cells. Therefore, correcting pathogenic gene mutations or functional correction sequences offers attractive therapeutic opportunities for treating these diseases.
[0164] Somatic cell gene editing is an attractive strategy for many human diseases. Through precise editing of target DNA or RNA sequences, the BART system can correct mutated genes in genetic disorders, inactivate viral genomes in infected cells, eliminate the expression of pathogenic proteins in neurodegenerative diseases, or silence oncogenic proteins in cancer. Therefore, the systems and methods disclosed herein can be used to correct potential genetic alterations in diseases, including the aforementioned genetic disorders, chronic infectious diseases, neurodegenerative diseases, and cancer.
[0165] genetic diseases
[0166] It is estimated that over six thousand genetic diseases are caused by known gene mutations. Correcting underlying pathogenic mutations in pathological tissues / organs can provide mitigation or cure of the disease. For example, one in 3,000 people in the United States suffers from cystic fibrosis. It is caused by the inheritance of a mutated CFTR gene, and 70% of patients have the same mutation; the deletion of a trinucleotide leads to the deletion of phenylalanine at position 508, which results in the misalignment and degradation of CFTR. The systems and methods disclosed in this invention can be used to convert Val 509 residues (GTT) to Phe 509 (TTT) in affected tissue (lung), thereby functionally correcting the Phe 508 mutation. Furthermore, a second repressive mutation in the mutated Phe 508 CFTR (such as R553Q, R553M, or V510D) can partially restore the function of the CFTR protein in somatic tissues.
[0167] Chronic infectious diseases
[0168] The disclosed systems and methods can also be used to specifically inactivate any gene in a viral genome incorporated into human cells / tissues. For example, the systems and methods disclosed herein allow for the generation of stop codons for the early termination of translation of essential viral genes, thereby rescuing or curing chronic debilitating infectious diseases. For example, current AIDS therapies can reduce viral load but cannot completely eliminate dormant HIV from positive T cells. The systems and methods disclosed herein can be used to permanently inactivate the expression of one or two essential HIV genes integrated into the HIV genome in human T cells by introducing one or two stop codons. Another example is hepatitis B virus (HBV). The systems and methods disclosed herein can be used to specifically inactivate one or two essential HBV genes incorporated into the human genome and to silence the HBV life cycle.
[0169] Neurodegenerative diseases
[0170] Some neurodegenerative diseases are caused by gain-of-function mutations. For example, SOD1G93A contributes to the development of amyotrophic lateral sclerosis (ALS). The systems and methods disclosed in this invention can be used to correct mutations or eliminate mutant protein expression by introducing stop codons or by altering splice sites.
[0171] Muscular system diseases
[0172] Dystrophin is a cytoplasmic protein that provides structural stability to the myoglobin complex of the cell membrane, which is responsible for regulating muscle cell integrity and function. The dystrophin gene, or “DMD gene” as used interchangeably in this article, is 2.2 megabases long at locus Xp21. The primary transcript measures approximately 2,400 kb, and the mature mRNA is approximately 14 kb. 79 exons encode a protein of over 3,500 amino acids. Exon 51 is typically adjacent to the box-breaking deletions in DMD patients and has been targeted in clinical trials using oligonucleotide-based exon skipping. A recent clinical trial of the exon 51 skipping compound eteplirsen reported significant functional benefits over 48 weeks, with an average of 47% dystrophin-positive fibrils compared to baseline. Mutations in exon 51 are ideally suited for permanent correction via NHEJ-based genome editing. The method in U.S. Patent Publication No. 20130145487 can also be modified for use in the nucleic acid targeting system of the present invention, which involves a wide range of nuclease variants that cleave target sequences from the human dystrophin gene (DMD).
[0173] cancer
[0174] Many genes, including tumor suppressor genes, oncogenes, and DNA repair genes, contribute to cancer development. Mutations in these genes commonly lead to various cancers. Using the systems and methods disclosed herein, these mutations can be specifically targeted and corrected. As a result, pathogenic oncogenes can be functionally neutralized, or their expression can be eliminated by introducing point mutations at catalytic or splicing sites. In some embodiments, cancer treatment, prophylaxis, or diagnosis is provided. Targets are preferably one or more of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, TRAC, or TRBC genes. Cancer can be one or more of the following: lymphoma, chronic lymphocytic leukemia (CLL), B-cell acute lymphoblastic leukemia (B-ALL), acute lymphoblastic leukemia, acute myeloid leukemia, non-Hodgkin's lymphoma (NHL), diffuse large cell lymphoma (DLCL), multiple myeloma, renal cell carcinoma (RCC), neuroblastoma, colorectal cancer, breast cancer, ovarian cancer, melanoma, sarcoma, prostate cancer, lung cancer, esophageal cancer, hepatocellular carcinoma, pancreatic cancer, astrocytoma, mesothelioma, head and neck cancer, and medulloblastoma. This can be implemented using engineered chimeric antigen receptor (CAR) T cells. This is described in WO2015161276, the disclosure of which is incorporated herein by reference and described below. In some embodiments, target genes suitable for treating or preventing cancer may include those described in WO2015048577, the disclosure of which is incorporated herein by reference.
[0175] Stem cell gene modification
[0176] In some embodiments, stem cells or progenitor cells can be genetically modified using the systems and methods disclosed in this invention. Suitable cells include, for example, stem cells (adult stem cells, embryonic stem cells, iPS cells, etc.) and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.). Suitable cells include mammalian stem cells and progenitor cells, including, for example, rodent stem cells, rodent progenitor cells, human stem cells, human progenitor cells, etc. Suitable host cells include in vitro host cells, such as isolated host cells.
[0177] In some implementations, the BART system can be used for targeted and precise gene modification of ex vivo tissues to correct underlying genetic defects. After ex vivo correction, the tissue can be returned to the patient. Furthermore, this technology can be widely used in cell-based therapies to correct genetic diseases.
[0178] Gene editing in animals and plants
[0179] The systems and methods described above can be used to generate transgenic non-human animals or plants with one or more gene modifications of interest. In some embodiments, the transgenic non-human animal is homozygous for the gene modification. In some embodiments, the transgenic non-human animal is heterozygous for the gene modification. In some embodiments, the transgenic non-human animal is a vertebrate, such as fish (e.g., zebrafish, goldfish, pufferfish, cave fish, etc.), amphibians (frogs, salamanders, etc.), birds (e.g., chickens, turkeys, etc.), reptiles (e.g., snakes, lizards, etc.), mammals (e.g., ungulates, such as pigs, cattle, goats, sheep, etc.); rabbits (e.g., rabbits); rodents (e.g., rats, mice); or non-human primates.
[0180] This system and method can be used to treat animal diseases in a manner similar to that used to treat human diseases as described above. Optionally, it can be used to generate knock-in animal disease models carrying specific gene mutations for research, drug discovery, and target validation purposes. The system and method described above can also be used to introduce point mutations into ES cells or embryos of various organisms for breeding and improving the quality of animal stock and crops.
[0181] Methods for introducing exogenous nucleic acids into plant cells are well known in the art. Suitable methods include viral infection (such as double-stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, gene gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, and Agrobacterium-mediated transformation. The choice of method usually depends on the type of cells to be transformed and the circumstances under which transformation occurs (i.e., in vitro, ex vivo, or in vivo).
[0182] In another aspect, this disclosure provides a method for treating or preventing viral infection of an RNA virus in a cell or subject. In some embodiments, the method includes delivering to a cell or subject an enzyme having reverse transcriptase and endonuclease activity, or a nucleic acid molecule encoding said enzyme, wherein the enzyme reverse transcribes viral RNA of the RNA virus into ssDNA, wherein the ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0183] In some implementations, the RNA virus is selected from norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0184] Methods to enhance immunity against RNA virus infections
[0185] In another aspect, this disclosure provides a method for enhancing immunity against viral infection by an RNA virus in cells or a subject. In some embodiments, the method includes delivering to cells or a subject an enzyme or a nucleic acid molecule encoding said enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into ssDNA to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, wherein the enzyme reverse transcribes the mRNA into a second ssDNA, wherein the second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0186] In some embodiments, the method includes delivering to a cell or subject an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, or a nucleic acid molecule encoding said enzyme, wherein the enzyme reverse transcribes viral RNA of an RNA virus into single-stranded DNA (ssDNA) to generate cDNA of viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, and wherein the mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
[0187] In some implementations, the RNA virus is selected from norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0188] In another aspect, this disclosure also provides a method for generating a cell line with immunity against viral infection by an RNA virus. In some embodiments, the method includes delivering to cells an enzyme or a nucleic acid molecule encoding said enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of the RNA virus into ssDNA to generate cDNA of the viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or a subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, wherein the enzyme reverse transcribes the mRNA into a second ssDNA, wherein the second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA, and wherein the enzyme cleaves the viral RNA.
[0189] In some embodiments, the method includes delivering an enzyme or a nucleic acid molecule encoding the enzyme to a cell having reverse transcriptase activity, endonuclease activity, and integrase activity, wherein the enzyme reverse transcribes viral RNA of an RNA virus into single-stranded DNA (ssDNA) to generate cDNA of viral RNA, wherein the enzyme integrates the cDNA into the host genome of the cell or subject, wherein the cDNA integrated into the host genome is transcribed into mRNA, and wherein the mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
[0190] In some implementations, the RNA virus is selected from norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0191] Carrier systems and cells
[0192] carrier system
[0193] In some implementations, the BART system described herein can be delivered to a host cell via one or more vectors, such as viral vectors. For example, one or more viral vectors may include adenoviruses, lentiviruses, adeno-associated viruses, or RNA-based viral vectors, which may be capable of replication or may encode only genes for self-amplification; the latter construct will be referred to herein as a replicon.
[0194] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is attached. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that include one or more free ends, or those without free ends (e.g., circular); nucleic acid molecules that include DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which a virally derived DNA or RNA sequence is present in the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, adeno-associated viruses, and / or RNA-based replicons). Viral vectors also include polynucleotides carried by the virus for transfection into host cells. Some vectors are capable of autonomous replication in the host cells into which they are introduced (e.g., RNA vectors that include their own RNA-dependent RNA polymerase, bacterial vectors with bacterial replication initiation sites, and free-type mammalian vectors). Other vectors (e.g., non-free mammalian vectors) integrate into the host cell's genome upon introduction into the host cell, thereby replicating along with the host genome. Furthermore, some vectors can direct the expression of genes operatively linked to them. Such vectors are referred to herein as "expression vectors." Vectors used in eukaryotic cells and resulting in expression in eukaryotic cells may be referred to herein as "eukaryotic expression vectors." Common expression vectors useful in recombinant DNA technologies are typically in the form of plasmids.
[0195] Recombinant expression vectors may include nucleic acids of the present invention in a form suitable for expression in host cells. This means that the recombinant expression vector includes one or more regulatory elements, which may be selected based on the host cell used for expression and are operatively linked to the nucleic acid sequence to be expressed. Within the recombinant expression vector, "operatively linked" is intended to mean that the nucleotide sequence of interest is linked to a regulatory element (one or more) in a manner that allows the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0196] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences, and RNA elements required for recognition by self-coding RNA-dependent RNA polymerases). Such regulatory elements are described, for example, in Goeddel, *GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY* 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of nucleotide sequences in many types of host cells and those that direct expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may direct expression primarily in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). Regulatory elements may also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell type-specific. In some embodiments, the vector comprises one or more pol III promoters (e.g., pol III promoters 1, 2, 3, 4, 5 or more), one or more pol II promoters (e.g., pol II promoters 1, 2, 3, 4, 5 or more), one or more pol I promoters (e.g., pol I promoters 1, 2, 3, 4, 5 or more), or combinations thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retroviral Rouss sarcoma virus (RSV) LTR promoter (optionally having an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally having a CMV enhancer) [see, for example, Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the glycerol phosphokinase (PGK) promoter, and the EF1α promoter. The term “regulatory element” also encompasses enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globulin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981).Those skilled in the art will understand that the design of expression vectors can depend on factors such as the choice of host cells to be transformed and the desired expression level. Vectors can be introduced into host cells to produce transcripts, proteins, or peptides encoded by nucleic acids as described herein, including fusion proteins or peptides (e.g., clustered regularly spaced short palindromic repeat (BART) transcripts, proteins, enzymes, their mutant forms, their fusion proteins, etc.).
[0197] In some embodiments, the vector system may include one or more viral vectors. In some embodiments, the one or more viral vectors include adenovirus-based vectors, lentivirus-based vectors, adeno-associated virus-based vectors, or RNA-based replicons. In some embodiments, upon expression, the BART system can bind and cleave at the target sequence, thereby preventing the formation of functional virions. Therefore, functional virions will not be assembled. Thus, the vector systems described herein include self-replicating RNA (e.g., Nodamurovirus-based replicons) that, in the presence of navigation DNA, produces BART proteins while simultaneously self-inactivating in the performance of its function.
[0198] cell
[0199] On the other hand, this disclosure provides host cells or cell lines or their progeny that include the BART system, vector system, or polynucleotide as described above. In some embodiments, the host cells or cell lines or their progeny include stem cells or stem cell lines. The cells can be eukaryotic cells (e.g., plant, animal, or human cells) or prokaryotic cells.
[0200] It also provides the products of any such cells or any such progeny produced by one or more target loci modified by the BART system. The products can be peptides, polypeptides, or proteins.
[0201] In some embodiments, the host cell or cell line includes one or more polynucleotide molecules encoding enzymes having reverse transcriptase activity, endonuclease activity and integrase activity, and one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0202] As used herein, the term "cell" can refer to prokaryotic or eukaryotic cells optionally obtained from a subject or a commercially available source. In some embodiments, the eukaryotic cell may be a human cell, rodent cell, optionally a mouse cell, yeast cell, or insect cell. In some embodiments, the eukaryotic cell may be a Chinese hamster ovary (CHO) cell.
[0203] "Eukaryotic cells" encompass all life forms except for prokaryotes. They are easily distinguished by their membrane-bound nuclei. Animals, plants, fungi, and protists are eukaryotic organisms or organisms whose cells are organized into complex structures by an inner membrane and cytoskeleton. The most typical membrane-bound structure is the nucleus. Unless otherwise specified, the term "host" includes eukaryotic hosts, including, for example, cells of yeast, higher plants, insects, and mammals. Non-limiting examples of eukaryotic cells or hosts include apes, cattle, pigs, mice, rats, birds, reptiles, and humans, such as HEK293 cells and 293T cells.
[0204] Prokaryotic cells typically lack a nucleus or any other membrane-bound organelles and are divided into two domains: bacteria and archaea. In addition to chromosomal DNA, these cells can also contain genetic information in circular loops called episomes. Bacterial cells are very small, roughly the size of an animal mitochondria. Prokaryotic cells are characterized by three main shapes: rod-shaped, spherical, and spiral. Bacterial cells separate by binary fission, rather than undergoing a complex replication process like eukaryotes. Examples include, but are not limited to, Bacillus, Escherichia coli, and Salmonella.
[0205] As used herein, the term "offspring," such as the offspring of a transgenic plant, refers to offspring produced, reproduced, or derived from a plant or a transgenic plant. Introduced nucleic acid molecules can also be introduced transiently into recipient cells, such that the introduced nucleic acid molecules are not inherited by subsequent offspring, and are therefore not considered "transgenic." Thus, as used herein, a "non-transgenic" plant or plant cell is a plant that does not contain foreign nucleic acids stably integrated into its genome.
[0206] Also within the scope of this disclosure are methods for preparing the aforementioned cell lines. In some embodiments, the method includes introducing into the cells an enzyme encoding reverse transcriptase and endonuclease activities, and one or more polynucleotide molecules encoding one or more nucleic acid components.
[0207] In another aspect, this disclosure also provides a method for generating model eukaryotic cells comprising a mutated disease gene, which can be any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes (a) introducing a BART system into eukaryotic cells; and (b) allowing a BART complex (e.g., a BART / navigation ssDNA complex) to bind to a target polynucleotide to achieve cleavage of the target polynucleotide within the disease gene, wherein the ssDNA comprises a sequence that hybridizes to a target sequence within the target polynucleotide, thereby generating model eukaryotic cells comprising the mutated disease gene.
[0208] In some embodiments, the cleavage involves cutting one or both strands at the location of the target sequence via a BART protein. In some embodiments, the cleavage results in reduced or increased transcription of the target gene. In some embodiments, the method further includes repairing the cleaved target polynucleotide via a non-homologous end joining (NHEJ)-based gene insertion mechanism with a foreign template polynucleotide, wherein the repair results in a mutation of insertion, deletion, or substitution of one or more nucleotides comprising the target polynucleotide. In some embodiments, the mutation results in an alteration of one or more amino acids in a protein expressed by a gene comprising the target sequence.
[0209] Various eukaryotic cells are suitable for use in the method. For example, the cells can be human cells, non-human mammalian cells, non-mammal vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotic organisms. Various embryos are suitable for use in the method. For example, the embryo can be a 1-cell, 2-cell, or 4-cell human or non-human mammalian embryo. Exemplary mammalian embryos include 1-cell embryos such as mouse, rat, hamster, rodent, rabbit, feline, canine, sheep, pig, bovine, equine, and primate embryos. In other embodiments, the cell can be a stem cell. Suitable stem cells include, but are not limited to, embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, multipotent stem cells, oligopotent stem cells, and unipotent stem cells. In exemplary embodiments, the cell is a mammalian cell, or the embryo is a mammalian embryo. In some embodiments, non-human mammalian cells can include, but are not limited to, primate, bovine, sheep, pig, canine, rodent, and lagomorphic cells, such as monkey, bovine, sheep, pig, dog, rabbit, rat, or mouse cells. In some embodiments, the cells can be non-mammal eukaryotic cells such as poultry, birds (e.g., chickens), vertebrates, fish (e.g., salmon), or shellfish (e.g., oysters, clams, lobsters, shrimp) cells. In some embodiments, the non-human eukaryotic cells are plant cells. Plant cells can belong to monocotyledonous or dicotyledonous plants or to crop or cereal plants such as cassava, corn, sorghum, soybeans, wheat, oats, or rice. Plant cells can also be the cells of algae, trees, or producing plants, fruits, or vegetables (e.g., trees such as citrus trees, such as orange, grapefruit, or lemon trees; peach or nectarine trees; apple or pear trees; nut trees such as almond, walnut, or pistachio trees; nightshade plants; brassica plants; lettuce plants; spinach plants; pepper plants; cotton, tobacco, asparagus, carrots, cabbage, broccoli, cauliflower, tomatoes, eggplants, peppers, lettuce, spinach, strawberries, blueberries, raspberries, blackberries, grapes, coffee, cocoa, etc.).
[0210] Gene editing systems and kits
[0211] Gene editing system
[0212] In another aspect, this disclosure provides a gene editing system for modifying target polynucleotides. In some embodiments, the method includes an enzyme having reverse transcriptase and endonuclease activities, and one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0213] In some embodiments, the BART system disclosed herein, or a gene editing system comprising the disclosed BART-based system, can be delivered via liposomes, particles (e.g., nanoparticles), exosomes, microvesicles, lipids, cell-penetrating peptides (CPPs), or gene guns. Delivery vehicles, particles, nanoparticles, formulations, and components thereof used to express one or more elements of the aforementioned BART system are as described in PCT / US2013 / 074667.
[0214] In some implementations, the enzyme further possesses integrase activity.
[0215] In some implementations, one or more nucleic acid components navigate to DNA using a single strand.
[0216] In some embodiments, one or more nucleic acid components further include a payload RNA. In some embodiments, the payload RNA includes a stem-loop structure. In some embodiments, the enzyme reverse transcribes the payload RNA into cDNA. In some embodiments, the enzyme integrates the cDNA into a target polynucleotide.
[0217] In some embodiments, the enzyme includes an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain. In some embodiments, the enzyme further includes a C-terminal domain. In some embodiments, the C-terminal domain facilitates interaction between the enzyme and the target polynucleotide. In some embodiments, the C-terminal domain includes a zinc finger. In some embodiments, the zinc finger includes a CCHC motif.
[0218] In some embodiments, the enzyme includes bat-associated reverse transcriptase (BART). In some embodiments, BART includes BART from horseshoe bat, rat eared bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
[0219] In some embodiments, the gene editing system includes one or more polynucleotide molecules encoding an enzyme. In some embodiments, the gene editing system includes one or more polynucleotide molecules encoding or comprising one or more nucleic acid components. In some embodiments, the one or more polynucleotide molecules include one or more vectors. In some embodiments, the enzyme and one or more nucleic acid components are provided in a single vector.
[0220] In some implementations, the target polynucleotide includes RNA or DNA.
[0221] In some embodiments, RNA includes viral RNA of RNA viruses. In some embodiments, the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E viruses, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
[0222] In some implementations, the DNA includes genomic DNA or cDNA.
[0223] In some implementations, the modification of the target polynucleotide includes the cleavage of the target polynucleotide.
[0224] In one aspect, this disclosure provides a gene editing system comprising one or more vectors, liposomes, particles (e.g., nanoparticles, lipid nanoparticles), exosomes, or microvesicles, comprising one or more components of a BART system and optionally a pharmaceutically acceptable carrier.
[0225] In another aspect, this disclosure provides a vector system comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding enzymes having reverse transcriptase activity, endonuclease activity and integrase activity, and comprise one or more nucleic acid components, wherein the one or more nucleic acid components target and hybridize with the target polynucleotide and direct the binding of the enzyme to the target polynucleotide, and wherein the enzyme modifies the target polynucleotide.
[0226] As used herein, the terms "gene editing system," "composition," or "pharmaceutical composition" refer to a mixture of at least one component useful within this disclosure with other components (such as carriers, stabilizers, diluents, dispersants, suspending agents, thickeners, and / or excipients). The pharmaceutical composition facilitates the administration of one or more components of the BART system to an organism.
[0227] As used herein, the term “pharmaceutically acceptable” means a material, such as a carrier or diluent, that does not eliminate the biological activity or properties of the composition and is relatively non-toxic; that is, the material can be administered to an individual without causing undesirable biological effects or interacting with any component of the composition comprising it in a harmful manner.
[0228] As used herein, the term "pharmaceutically acceptable carrier" includes pharmaceutically acceptable salts, pharmaceutically acceptable materials, compositions, or carriers, such as liquid or solid fillers, diluents, excipients, solvents, or encapsulating materials, relating to carrying or transporting the compounds of the present invention within or to a subject so that the subject can perform its intended function. Typically, such compounds are carried or transported from one organ or part of the body to another. Each salt or carrier must be "acceptable" in the sense of compatibility with other components of the formulation and harmless to the subject. Some examples of materials that can serve as pharmaceutically acceptable carriers include: sugars, such as lactose, glucose, and sucrose; starches, such as corn starch and potato starch; cellulose and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose, and cellulose acetate; powdered tragacanth gum; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols, such as propylene glycol; and polyols, such as glycerin. Sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethanol; phosphate buffer solution; diluents; granulators; lubricants; binders; disintegrants; wetting agents; emulsifiers; colorants; release agents; coating agents; sweeteners; flavoring agents; aroma agents; preservatives; antioxidants; plasticizers; gelling agents; thickeners; hardening agents; setting agents; suspending agents; surfactants; humectants; carriers; stabilizers; and other non-toxic and compatible substances used in pharmaceutical formulations, or any combination thereof. As used herein, "pharmaceutically acceptable carrier" also includes any and all coating agents, antibacterial and antifungal agents, and absorption delay agents that are compatible with the activity of one or more components of the invention and are physiologically acceptable to the subject. Additional active compounds may also be incorporated into the composition.
[0229] In another aspect, this disclosure provides a delivery system including a gene editing system, and said delivery system is adapted to deliver the gene editing system to cells or a subject. In some embodiments, the delivery system includes nanoparticles or vesicles encapsulating the gene editing system.
[0230] "Gene delivery solvent" is defined as any molecule that can carry an inserted polynucleotide into a host cell. Examples of gene delivery solvents are liposomes, micellar biocompatible polymers, including natural and synthetic polymers; lipoproteins; peptides; polysaccharides; lipopolysaccharides; artificial viral envelopes; metal particles; and bacteria or viruses, such as baculoviruses, adenoviruses, and retroviruses, bacteriophages, spores, plasmids, fungal vectors, and other recombinant solvents commonly used in the art, which have been described for expression in a variety of eukaryotic and prokaryotic hosts and can be used for gene therapy as well as for simple protein expression.
[0231] The polynucleotides disclosed herein can be delivered to cells or tissues using gene delivery solvents. As used herein, “gene delivery,” “gene transfer,” and “transduction” are terms used to refer to the introduction of exogenous polynucleotides (sometimes called “transgenic”) into host cells, regardless of the method used for introduction. Such methods include a variety of well-known techniques, such as vector-mediated gene transfer (via, for example, viral infection / transfection, or various other protein- or lipid-based gene delivery complexes) and techniques that facilitate the delivery of “naked” polynucleotides (such as electroporation, “gene gun” delivery, and various other techniques for introducing polynucleotides). The introduced polynucleotides can be maintained stably or transiently in the host cell. Stable maintenance typically requires the introduced polynucleotide to contain a replication origin compatible with the host cell or to integrate into a host cell replicon such as an extrachromosomal replicon (e.g., a plasmid) or into the nucleus or mitochondrial chromosome. Many vectors are known to mediate the transfer of genes into mammalian cells, as known in the art and described herein.
[0232] Reagent test kit
[0233] This disclosure further provides a kit comprising one or more components of the above-described system (e.g., BART protein, navigation ssDNA). In some embodiments, the kit may include one or more other reaction components. In such kits, appropriate amounts of one or more reaction components are provided in one or more containers or held on a substrate.
[0234] Examples of additional components for the kit include, but are not limited to, one or more host cells, one or more reagents for introducing foreign nucleotide sequences into the host cells, one or more reagents (e.g., probes or PCR primers) for detecting RNA or protein expression or verifying the state of target nucleic acids, and buffers or culture media (in 1x or concentrated form) for the reaction. The kit may also include one or more of the following components: loading reagents, termination reagents, modification or digestion reagents, permeabilizers, and devices for detection.
[0235] The reaction components used can be provided in various forms. For example, components (e.g., enzymes, RNA, probes, and / or primers) can be suspended in an aqueous solution or as lyophilized powders, pellets, or beads. In the latter case, the components form a complete mixture of the components used for assay upon reconstruction. The kits of the present invention can be provided at any suitable temperature. For example, for storage purposes, the kits are preferably provided and maintained below 0°C, preferably at -20°C or below, or otherwise frozen.
[0236] The kit or system may contain any combination of the components described herein in an amount sufficient for use in at least one assay. In some applications, one or more reaction components may be provided in a pre-measured single-use amount in separate, typically single-use tubes or equivalent containers. The amount of components supplied in the kit may be any suitable amount and may depend on the target market for the product. The container in which the components are supplied may be any conventional container capable of maintaining the supply form, such as microcentrifuge tubes, microtiter plates, ampoules, bottles, or integral test devices, such as fluid devices, boxes, lateral flow devices, or other similar devices.
[0237] The kit may also include packaging materials for holding the containers or combinations of containers. Typical packaging materials for such kits and systems include solid matrices (e.g., glass, plastic, paper, foil, and microparticles) that hold the reaction components or detection probes in any of a variety of configurations (e.g., in vials, microtiter plate wells, and microarrays). The kit may further include instructions for using the components, documented in tangible form.
[0238] Additional definition
[0239] To aid in understanding the detailed description of the compositions and methods according to this disclosure, several specific definitions are provided to facilitate the explicit disclosure of various aspects of this disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0240] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. The following references provide general definitions for many of the terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walkered., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), SpringerVerlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, unless otherwise stated, the following terms have the meanings assigned to them below.
[0241] As used herein, the term “disease” is intended to be generally synonymous with and interchangeable with the terms “symptom” and “condition” (as in medical condition), since they all reflect an abnormal condition of the human or animal body or part thereof that impairs normal function, usually manifested by distinguishable signs and symptoms, and resulting in a reduced lifespan or quality of life in humans or animals.
[0242] The terms “sample,” “test sample,” and “patient sample” are used interchangeably herein. A sample may be a sample of serum, urine, plasma, amniotic fluid, cerebrospinal fluid, cells, or tissue. Such samples may be used directly as obtained from a patient, or may be pretreated, such as by filtration, distillation, extraction, concentration, centrifugation, inactivation of interfering components, and the addition of reagents, thereby altering the characteristics of the sample in some of the ways discussed herein or as known in the art. As used herein, the terms “sample” and “biological sample” generally refer to biological material to be tested and / or suspected of containing an analyte of interest, such as an antibody. A sample may be any tissue sample from a subject. A sample may include proteins from a subject.
[0243] The terms “reduction,” “lower,” “reduction,” “reduction,” or “inhibition” are generally used throughout this document to mean a reduction in a statistically significant amount. However, to avoid ambiguity, “reduction,” “lower,” “reduction,” or “inhibition” means a reduction of at least 10% compared to a reference level, for example, a reduction of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or up to and including a 100% reduction (e.g., a level that is not present compared to the reference sample), or any reduction between 10% and 100%.
[0244] As used in this article, the term “adjustment” means any change in the biological state, i.e., increase and decrease, etc.
[0245] The terms “increased,” “increased,” “enhanced,” or “activated” are generally used throughout this document to mean an increase in a statistically significant amount; to avoid any ambiguity, the terms “increased,” “increased,” “enhanced,” or “activated” mean an increase of at least 10% compared to a reference level, for example, at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or an increase of up to and including 100%, or any increase between 10 and 100%, or an increase of at least about 2, or at least about 3, or at least about 4, or at least about 5, or at least about 10, or any increase between 2 and 10 or more times compared to a reference level.
[0246] As used herein, the term “in vitro” refers to events that occur in an artificial environment, such as in a test tube or reaction vessel, in cell culture, etc., rather than in a multicellular organism.
[0247] As used in this article, the term "in vivo" refers to events that occur within multicellular organisms such as non-human animals.
[0248] It should be noted here that, as used in this specification and the appended claims, the singular forms “a / an / a type (a)”, “an / a type (an)” and “the” include plural references unless the context clearly specifies otherwise.
[0249] Unless otherwise stated, the terms “including,” “comprising,” “containing,” or “having,” and variations thereof, mean to cover the items listed thereafter and their equivalents, as well as additional subjects.
[0250] The phrases “in one implementation,” “in various implementations,” and “in some implementations,” etc., are used repeatedly. Such phrases do not necessarily refer to the same implementation, but they may refer to the same implementation unless the context otherwise specifies.
[0251] The terms “and / or” or “ / ” mean any one of the items, any combination of items, or all items in connection with the term.
[0252] The term "substantially" does not exclude "completely," for example, a composition that is "substantially free" of Y can be completely free of Y. Where necessary, the term "substantially" can be omitted from the definition of this invention.
[0253] As used herein, the term "about" or "approximately" when applied to one or more values of interest refers to a value similar to the reference value. In some embodiments, the term "about" or "approximately" refers to a range of values falling within a range of 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in any direction (greater or less than) of the reference value, unless otherwise stated or obvious from the context (other than such figures would exceed 100% of possible values). Unless otherwise stated herein, the term "about" is intended to include values close to the range that are functionally equivalent in relation to a single ingredient, composition, or embodiment, such as weight percentage.
[0254] It should be understood that wherever values and ranges are provided herein, all values and ranges covered by these values and ranges are meant to be covered within the scope of this invention. Furthermore, all values falling within these ranges, as well as the upper or lower limits of the value ranges, are also considered in this application.
[0255] As used herein, the term "each / every" when used to refer to a collection of items is intended to indicate a single item in the collection, but not necessarily every single item in the collection. Exceptions may occur if there is clear public knowledge or if the context otherwise specifies otherwise.
[0256] Unless otherwise stated, the use of any and all instances or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate the invention and not to limit the scope of the invention. The language in the specification should not be construed as indicating that any unclaimed element is essential to the practice of the invention. When used in this document, the term “exemplary” is intended to mean “as an example” and is not intended to indicate that a particular exemplary item is preferred or necessary.
[0257] Unless otherwise stated herein or clearly contradicted by the context, all methods described herein are performed in any suitable order. With respect to any method provided, the steps of the method may occur simultaneously or sequentially. When the steps of a method occur sequentially, they may occur in any order unless otherwise stated.
[0258] Where a method comprises combinations of steps, each and every combination or subcombination of steps is covered within the scope of this disclosure unless otherwise stated herein.
[0259] Each publication, patent application, patent, and other reference cited herein is incorporated in its entirety by reference without inconsistency with this disclosure. Publications disclosed herein are provided only because their disclosures predate the filing date of this invention. Nothing herein should be construed as an admission that this invention is not entitled to claim prior invention rights to such publications. Furthermore, the provided publication date may differ from the actual publication date, which may require independent verification.
[0260] It should be understood that the examples and embodiments described herein are for illustrative purposes only, and those skilled in the art will make various modifications or changes based thereon, and such modifications or changes are included within the spirit and scope of this application and the appended claims.
[0261] Example
[0262] Example 1
[0263] This embodiment describes the materials and methods used in the following embodiment 2.
[0264] Cell culture and maintenance
[0265] As previously described (Dejosez et al., 2023, Cell 186, 957-974), bat iPSCs from *Rhinolophus equina* were cultured. Cells were plated on irradiated MEF and cultured using DMEM / F12 (Gibco, 11330-032), 20% KOSR (Life Technologies, 10828-028), 0.1 mM NEAA (Gibco, 11140-050), 2 mM GlutaMAX (Gibco, 35050-061), Pen / Strep (Gibco, 15140122) at 10 U / ml and 10 μg / ml, respectively, 100 ng / ml FGF2 (R&D Systems, 233-FB), 100 ng / ml hSCF (StemCell Technologies, 78062.2), and 10 4U / ml mLIF (Millipore-Sigma, ESG1107), 20 nM laryngin (Sigma, F6886), and 100 μM 2-mercaptoethanol were used for maintenance.
[0266] Bat embryonic fibroblasts (BEFs) were immortalized using SV40-LT and cultured in 0.1% gelatin on DMEM (Gibco, 10569044) supplemented with 10% HI FBS (Sigma, F4135), NEAA, GlutaMAX, and Pen / Strep, similar to HEK293 cells (ATCC, CRL-1573).
[0267] Human ES cells (H9) were cultured on brinetin XF (Stemcell Technologies, 85850) using mTeSR1 (Stemcell Technologies, 100-0763).
[0268] Immunofluorescence staining
[0269] Immunostaining was performed on μ-slides (Ibidi, 80286) where cell lines were cultured until the fixation day. After washing the cells once with DPBS, they were fixed at 4°C for 20 min with Cytofix / Cytoperm solution (BD, BDB554714). The cells were then washed once with 1x Perm / wash buffer (BD, BDB554714) and incubated overnight at 4°C in Perm / wash buffer containing 1:100 diluted primary antibody against anti-ssDNA [F7-26] (Fisher Scientific, MAB3299MI), J2 anti-dsRNA (Scicons, 10010200), anti-HIV1 reverse transcriptase antibody (Abcam, ab63911), and anti-DNA:RNA hybridization antibody [S9.6] (Millipore-Sigma, MABE1095). Cells were washed three times with Perm / wash buffer before incubation for 1 hour at a 1:200 dilution with secondary antibodies goat anti-rabbit-AF488 (Life Technologies, A-10034), goat anti-mouse-AF488 (Life Technologies, A-11029), and donkey anti-goat-AF488 (Life Technologies, double check) at room temperature and in the dark. Finally, after washing twice with DPBS, cells were treated for 5 minutes at room temperature with the minor chemical Hoechst (Millipore-Sigma, 94403) at 5 μg / ml. Fixed cells were treated with 200 mg / ml RNase A (Fisher Scientific, EN0531) at 37°C for 4 hours before staining with anti-ssDNA [F7-26]. Live cells were incubated for 30 minutes at 37°C with the minor chemical SYBR Gold (Invitrogen, S11494) diluted 1:10,000 before visualization.
[0270] Cell lines carrying BART
[0271] Plasmids containing BART were synthesized using GeneScript with a His tag in the main strand and C-terminal region of pcDNA3.1(+) to allow for protein purification. Their 3D structures were predicted using ColabFold (ColabFold v1.5.2-patch: AlphaFold2 with MMseqs2). Following the manufacturer's protocol, they were cloned into H9 hES cells using Lipofectamine 3000 (ThermoFisher, L3000001). Selection was performed 48 hours post-transfection using increased concentrations of genimycin (ThermoFisher, 10131027), with a final concentration of 150 ng / ml.
[0272] Cytoplasmic DNA extraction
[0273] Following the previously described protocol (Cell Rep. 2012 Aug 30;2(2):207-15), cells were harvested by washing with DPBS and treating with Versene solution (Gibco, 15040066) at 37°C for 5 minutes. Cells were then collected and washed twice with ice-cold PBS and centrifuged at 1000G for 3 minutes at 4°C. Cytoplasmic DNA was extracted by vortexing cells for 4 seconds in a hypotonic lysis solution containing 10 mM Tris pH 8 (Invitrogen, AM9855G), 10 mM NaCl, 1.5 mM MgCl (Invitrogen, AM9530G), and 1 mM DTT (Thermo Scientific, R0862). After centrifuging cells at 1000G for 3 minutes at 4°C, the supernatant was collected and maintained at -80°C.
[0274] VSV infection
[0275] Cells were infected with vesicular stomatitis virus carrying the eGFP reporter at an MOI of 0.01 for 24 hours. The supernatant was collected by discarding floating cells and centrifuged at 300G for 5 minutes.
[0276] BART digestion analysis
[0277] For in vitro analysis, BART.0 and BART.1 were purified using a Ni-NTA rotation kit (Qiagen, 31314), with 2 μg of each protein used per reaction. Combined with 122 bp oligonucleotides (Sigma), 100 ng of dsDNA template was added, and genes B2M and CXCR4 were amplified using primers listed in Table 3.
[0278] Table 3. Primer Sequences
[0279]
[0280]
[0281] For in vivo analysis, a 122 bp ssDNA guide was transfected into an H9 cell line containing BART using Lipofectamine 3000. After 24 hours, genomic DNA was extracted using a Blood & Cell Culture DNA Kit (Qiagen, 13323), and potential mutation sites were amplified using GoTaq Master Mix (Promega, M7123). Sequences were purified using a PCR purification kit (Qiagen, 28104) and performed Sanger sequencing (GeneWiz).
[0282] Example 2
[0283] Given the dominance of RNA viruses in infecting mammals (e.g., SARS-CoV-2), a key feature of any mammalian-centric CRISPR-Cas-like system is the ability to convert viral genomic RNA into DNA for subsequent processing in the cell nucleus. Therefore, the presence of reverse transcriptase activity, an enzyme activity previously unreported by any mammalian genome, was investigated. Unique characteristics were observed in bat induced pluripotent stem cell (iPSC) colonies—they exhibited a compact and homogeneous morphology, and their cytoplasm was filled with tiny vesicles not observed in iPSCs from other mammalian species. These vesicles showed a uniform distribution in the cell's cytosol and possessed lipid properties. iPS cells were cultured for 3 days in lipid-deprived medium E8 and in E8 and AlbuMAX to understand how these vesicles persisted only in the presence of lipids in the medium. These vesicles were also absent in bat embryonic fibroblasts cultured in serum-containing medium.
[0284] These vesicles were observed to contain diverse contents, ranging from viral particles to autophagosome systems. To investigate the contents of these vesicles, immunostaining was performed using antibodies against the known reverse transcriptase (RT) (HIV-RT-p51 / p66) found in the HIV genome. The vast majority of vesicles showed strong positive staining, indicating the presence of reverse transcriptase. Subsequently, it was investigated whether these vesicles might also contain single-stranded DNA (ssDNA) (a product of reverse transcription). Immunostaining with antibodies specifically detecting ssDNA revealed abundant ssDNA within the cytoplasmic vesicles. The presence of intracytoplasmic DNA / RNA hybrids was further confirmed using the S9.6 antibody. The presence of cytoplasmic ssDNA was further confirmed using the small molecule SybrGold, which selectively detects ssDNA at low concentrations (low concentration Sybr Gold). Furthermore, this phenotype is only observed in the pluripotent stage in bats. No reverse transcriptase or single-stranded DNA was found in embryonic fibroblasts. However, signals of DNA / RNA hybrids were observed in BEF. In summary, these results strongly confirm the presence of a protein with reverse transcriptase activity in the cytoplasm of bat iPSCs. This protein appears to generate large amounts of ssDNA, providing credibility to the hypothesis that bats possess a unique, mammal-adapted CRISPR-Cas-like system.
[0285] To further understand the nature of ssDNA generated in bat iPSCs, its sequence was determined. The approach was twofold: first, cytoplasmic DNA was isolated, and libraries of these single-stranded DNAs were prepared for subsequent next-generation sequencing analysis. Specifically, cytoplasmic contents were extracted using a mild hypotonic solution without rupturing the nucleus. Next, SMART (5' end conversion mechanism of RNA template) technology was applied to generate cDNA from the extracted single-stranded DNA, effectively avoiding trace amounts of RNA (it is not yet clear how precisely this works during library preparation). This technique ensures that the full length of the ssDNA (including the 5' end) is presented in the final cDNA library, thus providing an accurate depiction of the original sequence. Analysis of the nucleic acids obtained by this method revealed clear bands of ssDNA in the bat iPSC cytoplasm, approximately 122 bp in length. These findings provide further evidence of active reverse transcription in bat iPSCs. Furthermore, the results suggest that the process not only generates a large amount of ssDNA but also appears to involve specific processing steps. Template RNA or reverse-transcribed ssDNA is processed into well-defined DNA stretches, a phenomenon reminiscent of sequence processing observed in bacterial CRISPR-Cas systems.
[0286] To gain a more comprehensive understanding of the nature of the sequences contained within these short ssDNAs, next-generation sequencing was performed on SMART libraries, and the sequence reads were then mapped to the genome of the horseshoe bat (R. ferrumequinum). The resulting data provided several unexpected insights into the characteristics of the ssDNA sequences. Notably, the majority of sequences could not be mapped to the bat genome. Of the total 116,550,396 sequence reads, 111,728,741 (95.86%) did not align with any known bat sequences. Subsequent attempts to map these sequences using BLAST searches against all publicly available databases failed, indicating that these sequences are distinct from any known genetic material. Only a small fraction (4,679,127 sequences (4.01%)) mapped to the bat genome. Of these small percentage of mappable sequences, approximately 10% comprised cellular genes such as POU5F1, X, Y, and Z. Notably, the remaining 90% of these mappable sequences were found to align with unplaced scaffolds in the bat genome, rather than standard chromosomal sequences. These unplaced scaffolds typically do not integrate into the genome assembly, but are usually embedded as 50-100 kb islands within highly repetitive heterochromatin. Upon closer examination of the ssDNA maps, these sequences frequently align with endogenous viral sequences such as transposons or retroviruses. This pattern applies even for the 10% of reads that are indeed successfully mapped to chromosomal sequences. In summary, these results suggest that most short ssDNA sequences correspond to viral sequences embedded within islands of heterochromatin DNA and possess potential antiviral capabilities.
[0287] Next, the pattern of genome integration of short ssDNA fragments was investigated. Notably, sequences within the unlocalized scaffold sequences exhibited a unique alignment pattern—highly aligned islands separated by relatively equidistant regions with no alignment whatsoever. This not only suggests the precise selection of sequences integrated into the genome array but also indicates the possibility that these scattered regions may play a crucial role. This observed pattern is functionally reminiscent of the CRISPR-Cas system, where specific viral sequences are integrated into the genome array at uniform distances to ensure proper expression. Therefore, these findings provide evidence that bat genomes contain sequence arrays composed of viral sequences homologous to cytoplasmic ssDNA sequences. This striking pattern further underscores the potential of these bats as natural and eukaryotic equivalents of the CRISPR-Cas system.
[0288] A key feature of the CRISPR-Cas defense mechanism is the spatial proximity (in cis) of the defense array to the genes encoding the enzymes responsible for the system. Given this, an open reading frame (ORF) search was performed within each island containing the scattered ssDNA array. ORFs of approximately 1300 amino acids in length were identified within each island, exhibiting nearly identical sequences under different conditions. Next, a BLAST search was performed. Although no known gene entries matched the sequence, the overall structure showed similarity to other proteins known to exhibit reverse transcriptase activity. This observation again demonstrates comparability with the CRISPR-Cas system. The discovery of the enzyme within scattered viral sequences in heterochromatin islands adjacent to the bat genome highlights the potential relevance of this newly discovered system.
[0289] To further understand the properties of this enzyme, a more detailed analysis was performed at the protein level. When the sequence was first compared with the genome of the horse horseshoe bat, a second enzyme emerged, differing only in 100 amino acids in the C-terminal region. Due to the high similarity of these enzymes, which span over 1000 amino acids, studies were conducted to determine the relationship between the two enzymes, for example, whether one enzyme is an editing variant and / or a more efficient version of the other. To distinguish between the two enzymes, the initially identified shorter version was named BART.0, and the slightly longer version was named BART.1. Experimental data showed that these enzymes do indeed possess reverse transcriptase activity, the ability crucial for initiating reverse transcription of mRNA targets, a key step in reverse transposition. Notably, these enzymes also exhibit endonuclease function. This activity is essential for introducing a nick into the chromosomal target DNA, thereby facilitating the integration of the new sequence. Specifically, it cleaves DNA in an AT-rich region located between the 5' stretch of purines and the 3' stretch of pyrimidines, which corresponds to the established (LINE-1) genomic integration site. Notably, these enzymatic features are consistent with observations including: specific reverse transcriptase activity for mRNA-to-cDNA conversion, nuclease function for creating nicks in DNA, and integration activity for incorporating new sequences into the genome (AlphaFold 3D structure). Therefore, based on sequence homology, this enzyme presents a robust mechanical framework capable of explaining the unique genomic events that have been identified.
[0290] Next, we investigated whether the system truly possesses specific anti-nucleic acid activity and the ability to encode genomic memory for existing infections, similar to the defense mechanisms bacteria use with CRISPR-Cas systems against bacteriophages. Two variants of the BART enzyme were cloned into human expression vectors to establish stable cell lines expressing these enzymes. These variants correspond to (BART.0) directly observed through mapping of ssDNA fragments and (BART.1) with higher homology found after alignment with the USCS Genome Database. Despite the introduction of this bat-derived enzyme, human cell lines (H9 and 293) did not show adverse effects on their growth properties. Notably, overexpression of BART conferred partial immunity against VSV virus, suggesting the interesting potential of BART-based systems as novel tools for viral resistance. Following BART overexpression, the BART-expressing cell lines were infected with VSV at a low MOI (0.01). The presence of viral cDNA in the cytoplasm was confirmed after extraction of cytoplasmic ssDNA. Not only was it present, but it was also processed and cleaved into shorter fragments. This observation demonstrates that the BART system is operable and effective against viral threats. But the most surprising finding lies in the fate of these short fragments. As genomic PCR and sequencing show, they are not aimlessly floating in the cytoplasm. Instead, they have been integrated into scattered arrays of similar islands, as observed with older genomic sequences. It should be noted that they are not merely idle DNA sequences—these fragments are transcribed into RNA, suggesting they can potentially influence cellular function. Furthermore, when the reverse transcriptase is inhibited, it leads to a loss of the antiviral activity observed in human cell lines. This further confirms the crucial role of BART and its reverse transcriptase activity in providing this novel type of viral resistance. Thus, these findings strongly suggest that enzyme systems, as described in this paper, can initiate defenses against viral infections. This is not limited to immediate acute infection—by integrating viral sequences into the genome, it appears to also provide a long-term memory of the infection, offering a form of durable resistance against future attacks by the same virus. These findings may open new frontiers in the fight against viral pathogens, providing previously unknown approaches to antiviral defense.
[0291] Next, the applicability of this system to genome and RNA editing was investigated. An array of ssDNA oligonucleotides was designed to target specific sites within the genome. 122 bp oligonucleotides were used to target genes described in Saito M., et al. (Saito M., et al. Nature 620, 660–668 (2023)) such as B2M, CXCR4, VEGFA, CA2, KRAS, DYRK1A, HPRT1, and DMD, as these are proto-oncogenes involved in the immune system, some of which have antiviral capabilities and are important for development. Human cells were first transfected with these oligonucleotides along with BART. Unexpectedly, significant nuclease activity was observed not only at the genomic sites targeted by the ssDNA oligonucleotides but also at mRNA sites. For example, cleavage efficiency at the target genomic loci was increased by 60% (p < 0.001) and cleavage at mRNA sites was enhanced by 55% (p < 0.01) compared to control cells. A series of further experiments were conducted to evaluate the fidelity of the system. Off-target effects were examined using whole-genome sequencing and RNA-Seq. These experiments revealed that BART-mediated editing resulted in 90% fewer off-target effects than conventional CRISPR-Cas systems (p < 0.05). This minimal off-target effect unique to BART-mediated editing is crucial for precise gene editing. Subsequently, various ssDNA oligonucleotides were engineered to target specific mutations known to cause hereditary diseases. BART was found to correct these mutations with high efficiency ranging from 70% to 90% (p < 0.01), indicating its potential therapeutic significance. The results of these studies are significant. The reverse transcriptase activity characterized in BART, combined with its ability to integrate specific sequences into the genome, suggests the system's potential application in precision gene therapy. In this case, BART can be used to correct gene mutations at specific genomic sites without the side effects observed with current CRISPR-Cas systems.
[0292] discuss
[0293] The findings described in this article reveal an unexpected feature in mammalian biology that will significantly impact the understanding of antiviral defense mechanisms and the development of novel gene-editing technologies. This research stemmed from the observation of unique vesicles in bat induced pluripotent stem cells (iPSCs). Through rigorous investigation, a novel CRISPR-like system adapted to mammals was discovered, called BART (Bat-Associated Reverse Transcriptase).
[0294] Based on the detection of reverse transcriptase activity and the presence of single-stranded DNA (ssDNA) in bat iPSC vesicles, it is hypothesized that an enzyme similar to the bacterial CRISPR-Cas system exists that can convert viral genomic RNA into DNA, potentially triggering a unique antiviral mechanism. Such reverse transcriptase activity has not been previously reported in mammalian genomes, and its discovery underscores the importance of exploratory research in elucidating novel biological phenomena.
[0295] In-depth examination of the ssDNA sequences present in the cytoplasm revealed that the ssDNA primarily originates from retroviral and transposon sequences. These sequences are mainly aligned with unlocalized scaffold sequences in the bat genome, i.e., genomic islands that are typically embedded in highly repetitive heterochromatin. This unexpected discovery revealed the existence of specific genomic arrays carrying viral sequences, similar to bacterial CRISPR arrays.
[0296] Furthermore, a unique pattern of ssDNA insertion into these genomic islands was identified, resembling the precision and functionality of the CRISPR-Cas system. Coupled with scattered ssDNA arrays to open reading frames (ORFs) encoding reverse transcriptases very close to these arrays—again reminiscent of CRISPR-Cas. These findings strongly suggest that, despite significant evolutionary differences between mammals and bacteria, mammalian defense mechanisms against viral invasion share functional equivalence with bacterial CRISPR-Cas systems.
[0297] Next, the functional implications of BART were investigated, and it was found that its overexpression confers immunity against vesicular stomatitis virus (VSV). During VSV infection, the BART enzyme integrates short fragments of viral cDNA into the host genome. These fragments are not merely inertly integrated; they are also transcribed, suggesting a possible 'genomic memory' of viral infection. This memory enables the host to launch a rapid and effective defense against subsequent infections of the same virus, a feature reminiscent of adaptive immunity seen in bacteria via the CRISPR-Cas system.
[0298] Furthermore, the results show that BART surpasses the accuracy of existing CRISPR-Cas systems, possessing high editing efficiency and fewer off-target effects. This opens up promising prospects for the BART system as a superior tool for gene therapy, endowing it with increased specificity for correcting gene defects.
[0299] In summary, these results reveal an intriguing aspect of mammalian biology, blurring the lines between traditionally distinct prokaryotic and eukaryotic defense mechanisms. By revealing this mammalian-adapted CRISPR-like system, the disclosed BART system could revolutionize antiviral therapies and gene-editing technologies, opening entirely new avenues for addressing the most pressing health challenges.
[0300] Example 3
[0301] The origins of BART can be traced back to research on bat induced pluripotent stem cells (iPSCs), where unique cytoplasmic vesicles not found in other mammalian iPSCs were observed. The reverse transcriptase activity and presence of single-stranded DNA (ssDNA) within these vesicles suggest a potential link to retro-transcriptional elements, mobile genetic elements that utilize reverse transcription throughout their life cycle. Further analysis revealed that BART shares sequence homology with the L1 family of non-LTR retrotransposons, ancient elements that have existed in mammalian genomes for millions of years. However, BART exhibits key modifications in both its structure and function, indicating its domestication and reuse in novel cellular roles, distinct from its retrotransposon ancestors. The discovery of BART in bat stem cells, along with its evolutionary association with retro-transcriptional elements, underscores the potential for viral elements to be co-opted and reused for host functions, driving the evolution of new cellular mechanisms.
[0302] According to this disclosure, the BART protein has been found to exhibit a multi-domain structure supporting its diverse functions. The core of the protein houses a reverse transcriptase (RT) domain, essential for the conversion of RNA into DNA (a hallmark of retroviral activity). As demonstrated by sequence and structural alignment, the presence of key conserved sequence blocks and residues within this domain indicates that BART possesses functional RT activity. Further enhancing its functional repertoire, BART also possesses a nuclease domain capable of cleaving DNA at specific sites. This nuclease activity facilitates the integration of newly synthesized DNA into the host genome. The presence of a DNase I-like motif absent in relevant retrotransposons suggests an evolutionary adaptation for precise DNA cleavage. Additionally, an integrase domain, inferred from sequence homology and structural prediction, enables BART to insert reverse-transcribed DNA into the host genome, reminiscent of retroviral integration. The synergistic effect of these three domains—reverse transcriptase, nuclease, and integrase—equips BART with a significant ability to capture, process, and integrate viral genetic material, forming the basis of its antiviral and gene-editing capabilities.
[0303] The presence of zinc finger domains within the BART protein suggests its potential for direct interaction with specific RNA sequences or structural motifs. Furthermore, stem-loop forming RNAs that bind to BART were identified, acting as guides or scaffolds to facilitate the recognition and binding of viral RNAs. The target sequences preferred by BART in the genome appear to require double-stranded DNA adjacent to single-stranded DNA during the cell cycle, indicating a cell cycle-dependent aspect of the targeting mechanism. The interactions between these elements—the zinc finger domain, stem-loop RNA, and specific DNA background—contribute to BART's ability to selectively recognize and target viral RNAs, enabling its antiviral and gene-editing functions.
[0304] The process by which BART integrates viral cDNA into the host genome involves a series of coordinated, ordered steps echoing the CRISPR-Cas system in prokaryotes. This process begins when BART encounters viral RNA. Using its reverse transcriptase activity, BART meticulously transcribes the RNA into complementary DNA (cDNA). This cDNA then undergoes processing steps, producing a short fragment of approximately 122 base pairs. This precision of processing ensures that only specific segments of the viral genome are tagged for integration, a characteristic that contributes to system specificity. The processed cDNA fragment is then seamlessly woven into the host genome at precise locations within untargeted scaffold sequences, typically located in heterochromatin-rich regions. The integration sites exhibit a unique, scattered array pattern, strikingly similar to CRISPR arrays observed in bacteria. The process culminates in the formation of a cDNA array within the host genome, which acts as a 'genomic memory' of past viral encounters. The integrated cDNA is then transcribed into RNA, which guides BART to target and cleave homologous viral RNA during subsequent infection. This complex mechanism not only contributes to the form of adaptive immunity but also lays the foundation for BART's use as a powerful gene-editing tool.
[0305] The viral cDNA integrated into the host genome serves as a template for RNA molecule transcription, which acts as a guide for BART (Figure 2). These RNA transcripts, with sequences complementary to the original viral RNA, form a complex with the BART enzyme. This complex then actively seeks out and binds to homologous viral RNA within the cell. Upon binding, the endonuclease activity of BART cleaves the viral RNA, effectively neutralizing the virus and preventing its replication. The integration of viral cDNA into the host genome ensures a continuous source of guide RNA, enabling a rapid and targeted response to future infections with the same virus, thereby establishing a form of long-term adaptive immunity in mammalian cells.
[0306] The antiviral potential of BART is demonstrated by its ability to confer resistance against vesicular stomatitis virus (VSV) in human cells. Overexpression of BART in these cells resulted in a significant reduction in VSV replication, demonstrating its direct antiviral effect. Furthermore, the detection of processed viral cDNA fragments integrated into the host genome of BART-expressing cells provides compelling evidence for its role in building adaptive immunity. The observation that inhibition of BART's reverse transcriptase activity eliminated this antiviral effect further solidifies its crucial function in this novel defense mechanism. These findings collectively highlight BART's ability to combat RNA viruses and its potential applications in innovative antiviral therapies.
[0307] BART's inherent ability to reverse transcribe RNA and integrate the resulting cDNA into the genome can be strategically used for targeted gene editing. This process involves designing and delivering specific ssDNA oligonucleotides, which act as 'navigators' guiding BART to the desired genomic loci. These oligonucleotides, designed to be complementary to the target DNA sequence, facilitate the precise targeting and integration of the accompanying RNA payload. Figure 1 The RNA payload encoding the desired gene modification is then reverse transcribed via BART and seamlessly inserted into the genome, producing the desired edit. This method provides a universal platform for introducing various types of gene modifications, including insertions, deletions, and substitutions, with the potential for high efficiency and specificity.
[0308] Gene editing strategies
[0309] BART's gene-editing capabilities focus on its ability to reverse transcription, convert RNA into DNA, and integrate the newly synthesized DNA into the host genome. Unlike traditional gene-editing tools that directly modify DNA, such as CRISPR-Cas9, BART uses RNA as a template, providing a unique approach to gene manipulation. This RNA-adaptive approach expands the possibilities of gene editing, allowing for the insertion of different gene modifications at both the DNA and RNA levels and the regulation of gene expression.
[0310] The BART system comprises three key components: the BART enzyme, navigation ssDNA, and RNA payload. Figure 1 The ssDNA guides BART to a specific genomic locus targeting the modification, while the RNA payload carries the gene sequence to be inserted. This mechanism, independent of double-strand DNA breaks, offers advantages in specificity and reduced off-target effects, particularly due to the precision of ssDNA targeting. The RNA payload acts as a template for reverse transcription, enabling greater flexibility in inserting the desired modification compared to traditional DNA-based methods.
[0311] BART's dual functionality—targeting both genomic DNA and mRNA transcripts—provides a powerful tool for gene correction and transient gene regulation, expanding the scope of therapeutic interventions. Data show that BART achieves higher on-target editing rates and fewer off-target effects compared to CRISPR-Cas9, positioning it as a next-generation gene editing technology.
[0312] application
[0313] BART exhibits several key advantages over existing gene-editing systems, particularly when compared to CRISPR-Cas9. Data show that BART achieves higher on-target editing efficiency under various experimental conditions, resulting in more precise gene modifications. Furthermore, BART demonstrates enhanced specificity with significantly reduced off-target effects, thus minimizing the risk of unintended genomic alterations. One of BART's most notable features is its dual functionality, allowing it to target both DNA and RNA, extending its utility beyond traditional gene editing to include the regulation of gene expression. Moreover, BART's mammalian origin provides better compatibility with human cells, reducing the likelihood of immune responses and other complications typically associated with bacterial CRISPR-Cas systems. These properties collectively position BART as a versatile and highly effective gene-editing tool with the ability to introduce a wide range of gene modifications (including insertions, deletions, and substitutions) at both the DNA and RNA levels. This versatility, combined with its efficiency and specificity, underscores BART's potential for therapeutic applications, particularly in the correction of pathogenic mutations in human cells.
[0314] Example 4
[0315] This section describes the discovery of unique cytoplasmic vesicles in bat iPSCs, which are absent in other mammalian iPSCs and bat embryonic fibroblasts. These vesicles were found to exhibit reverse transcriptase activity and contain single-stranded DNA (ssDNA), suggesting a novel antiviral mechanism potentially similar to a mammalian CRISPR-Cas-like system. A multi-step approach was employed to elucidate the nature of the ssDNA within the cytoplasmic vesicles. First, cytoplasmic DNA was extracted from bat iPSCs using a mild hypotonic solution, carefully preserving the ssDNA while minimizing nuclear contamination. Then, cDNA libraries were created from the extracted ssDNA using SMART (5' end conversion mechanism of RNA template) technology, facilitating comprehensive sequence analysis. Next-generation sequencing of these libraries, subsequently mapped to the *Hippophae rhamnoides* (horse-horned horseshoe bat) genome, revealed that the majority of the ssDNA sequence originated from retroviral and transposon elements, primarily aligned with unlocalized scaffold sequences within the bat genome.
[0316] Mapping ssDNA sequences to the bat genome revealed a striking pattern: sequences primarily aligned with unlocalized scaffold sequences typically located in heterochromatin-rich regions. Within these scaffolds, ssDNA integrates in a unique, sporadic manner, forming arrays with regularly spaced sequences. This pattern is reminiscent of CRISPR arrays found in prokaryotes, suggesting functional parallelism and novel adaptive immune mechanisms in bats.
[0317] BART Classification
[0318] BART is classified within the broader category of non-LTR (long terminal repeat) retrotransposons based on its reverse transcriptase (RT) domain. Reverse transcriptase is a crucial enzyme that allows retroviruses such as HIV to replicate by converting their RNA genome into DNA, which is then integrated into the host genome. The classification of non-LTR retrotransposons is generally based on the concept of a "clade," a term coined by Julian Huxley in 1959 and refined by Malik, Burke, and Eickbush in 1999. In evolutionary biology, a clade refers to a group of genetic elements that share common characteristics, such as structural features and phylogenetic relationships. Specifically, clades within non-LTR retrotransposons are defined by three main criteria: 1) shared structural features, 2) phylogenetic grouping based on RT domain analysis, and 3) an evolutionary history dating back to ancient times, particularly the Precambrian. BART's RT domain exhibits these defining characteristics, indicating that it is included in one of the established clades of non-LTR retrotransposons. This classification emphasizes the evolutionary significance of BART and compares it to a group of elements that play a key role in the evolution of genomes across different species.
[0319] The dataset, comprising 211 RT domain protein sequences spanning 28 identified clades, was used as a comprehensive framework for BART classification. These clades (including well-known groups such as L1, CR1, and Jockey) capture the evolutionary diversity of non-LTR elements, some of which have persisted since the early evolution of eukaryotes (approximately 1-2 billion years ago). For example, the L1 clade is particularly ancient, with evidence suggesting it shares a common ancestor with bacterial class II mobile introns. This ancient lineage highlights the deep evolutionary roots of the L1 clade and its importance in the genome construction of modern organisms. Conversely, other RT enzymes, such as the p51 subunit of HIV RT, represent evolutionary specialization and divergence within RT domains. The p51 subunit is an inactive form, evolutionarily distant from its active counterpart, p66, which underscores the functional diversity that has emerged within retrotranscribed elements over time. This diversity reflects the adaptation of RT domains to a wide range of biological roles across different species, further emphasizing the complexity and evolutionary significance of these elements.
[0320] To classify BART within this evolutionary framework, the RTclass1 tool, a specialized bioinformatics program designed for phylogenetic analysis of RT domains, was employed. The process began with the extraction of RT domains from BART sequences using WU-BLAST / CENSOR, a tool for identifying and annotating repetitive elements in genomic sequences. Once the RT domains were isolated, they were aligned with sequences in the RTclass1 dataset, which contains broad representations of RT domains from 28 identified clades. This alignment is crucial for ensuring accurate comparison of BART RT domains with their evolutionary counterparts.
[0321] Following alignment, the BART RT domain sequences underwent multiple sequence alignment using SEQBOOT, a procedure that generates a repeated dataset for estimating phylogenetic confidence. These bootstrap alignments were then used to compute protein distance matrices, quantifying the evolutionary divergence between sequences. The phylogenetic tree was inferred from these distance matrices using the BIONJ algorithm, an optimized method for constructing a tree of minimal evolutionary change.
[0322] To ensure robustness of classification, RTclass1 generates 1,000 bootstrap trees, each representing a possible evolutionary path. Analysis of these trees allows RTclass1 to construct consensus bootstrap trees that identify the phylogenetic cluster or clade to which BART belongs. This rigorous process of alignment, permutation, and tree inference provides a reliable classification of BART within the established evolutionary framework of non-LTR retrotransposons.
[0323] Phylogenetic analysis revealed that the RT domain of BART is most closely associated with elements within the L1 clade, a group identified for its ancient origin and widespread presence in eukaryotic genomes. A rootless phylogenetic tree generated using neighbor-joining demonstrates the clustering of BART with other non-LTR elements, particularly those within the L1 clade. This clustering is supported by high bootstrap values, indicating strong and reliable phylogenetic localization of BART within this clade. The tree is rooted in RT sequences from class II introns, which act as evolutionary outgroups, highlighting the relationships between various non-LTR elements. BART's location within the L1 clade suggests that it shares an important evolutionary ancestor with these ancient mobile gene elements, emphasizing its potential role in eukaryotic genome evolution.
[0324] In summary, the classification of BART within the L1 clade of non-LTR retrotransposons is strongly supported by the structural and phylogenetic characteristics of its RT domain. The application of advanced phylogenetic tools and comprehensive sequence analysis has validated BART's location within this ancient clade, providing valuable insights into its evolutionary history. This classification also demonstrates the functional similarity between BART and other well-characterized RT elements within the L1 clade, highlighting its importance in the broader context of genome evolution.
[0325] Reverse transcriptase (RT) domain
[0326] Structure prediction of BART proteins was performed using the AlphaFold protein structure database, a deep learning-based tool developed by DeepMind. The process began by retrieving the BART amino acid sequence from sequencing data. The BART sequence was then compared with a comprehensive database of known protein structures and sequences, including those from the associated reverse transcriptase (RT) domains. Figure 3AlphaFold generates a three-dimensional (3D) structural model of BART by predicting the spatial arrangement of amino acids within a protein, focusing on highly conserved RT domains. The AlphaFold model is accompanied by a confidence score for each residue, indicating the reliability of the predicted positions within the structure. Following prediction, the structural model is refined through energy minimization to address any steric hindrance or unfavorable interactions predicted by AlphaFold. The resulting model is visualized and analyzed using PyMOL (a molecular visualization system), where key structural features such as finger, palm, and thumb subdomains are identified. Conserved residues within active sites are mapped, and the overall structural integrity of the BART RT domain is assessed.
[0327] Structural analysis of BART, particularly when compared to the human L1 retrotransposon, reveals a compelling thread of evolutionary adaptation and functional innovation. At the heart of BART lies its RT domain, which exhibits the classic "right-handed" construction characteristic of retroviruses and other retrotranscriptional elements. This construction comprises distinct subdomains: fingers, palm, and thumb, each playing a crucial role in the enzyme's function. In particular, the thumb subdomain extends into the wrist region, which is believed to enhance the enzyme's interaction with the template DNA during reverse transcription. This structural adaptation may contribute to the efficiency and specificity of BART's reverse transcription process.
[0328] The key to the functional integrity of BART lies in the preservation of critical amino acid residues within its RT active site. Specifically, residues at positions 519, 531, 533, 659, 604, 702, 600, 591, 605, 668, 566, 700, and 703 are highly conserved, reflecting essential motifs found in other active RT enzymes. These residues are important for coordinating the binding of the RNA template, the incoming nucleotide triphosphates, and the divalent metal ions necessary for catalysis. The conservation of these residues, aligned with the conserved sequence regions defined by Eickbush et al., underscores BART's functional capacity as a reverse transcriptase. The spatial arrangement of these residues within the active site creates a highly specialized microenvironment that facilitates the precise and efficient conversion of RNA into DNA—a hallmark of BART's function.
[0329] Furthermore, comparisons with the human L1 RT domain highlight both shared evolutionary traits and unique structural adaptations, reflecting the specialization of BART in the bat genome. BART's ability to maintain these conserved traits while potentially acquiring new functional properties suggests a significant evolutionary advantage, possibly related to the unique biological needs of its host organism.
[0330] F605 is a highly conserved residue within the RT domain of BART and plays a crucial role in its enzymatic function. Located within the active site, F605 acts as a gatekeeper, its aromatic side chain providing a structural barrier that selectively excludes ribonucleotides from entering the active site. This exclusion is essential for ensuring that BART functions as a DNA-dependent RNA polymerase, synthesizing DNA solely from an RNA template. The presence of F605 effectively prevents RNA-dependent RNA polymerization, a process that can interfere with the fidelity of reverse transcription. This selective mechanism is particularly important for BART's function, as it ensures that the enzyme does not participate in RNA synthesis, which could lead to undesirable genome integration or the production of unintended RNA transcripts. The conservation of F605 in relevant retrotransposons highlights its evolutionary importance, suggesting that this residue is retained due to its role in maintaining the specificity and accuracy of reverse transcription. In BART, this evolutionary adaptation may be related to its specialized role in transcribing viral RNA into DNA, a process central to its antiviral activity and gene-editing applications.
[0331] Apurinyl-depyrimidine endonuclease (APE) domain
[0332] In addition to the RT domain, BART's unique structural organization further distinguishes it from its retrotransposon ancestor. Figure 3 The N-terminal region of BART contains an apurinyl-depyrimidine endonuclease (APE) domain, which is crucial for DNA cleavage required for integration. Although the APE domain in BART shows some differences from the APE domain found in L1 retrotransposons, it retains key active site residues, indicating the preservation of its DNA cleavage function. This endonuclease activity is important for generating precise cleavages in the target DNA, which act as entry points for the integration of newly synthesized cDNA. Notably, the APE domain of BART also includes a DNase I-like motif, a feature absent in closely related retrotransposons, indicating an evolutionary refinement of BART's DNA cleavage mechanism. This refinement could contribute to BART's enhanced specificity and efficiency in gene editing applications, distinguishing it from other elements in the non-LTR retrotransposon family.
[0333] The localization of the APE domain in BART, located at the N-terminus of the RT domain, reflects the structural organization observed in other retrotransposons, highlighting a common structural theme. The role of the APE domain in retrotransposition is to introduce a targeted nick into the host DNA, facilitating the initiation of the retrotranscription process. The evolutionary origin of the APE domain can be traced back to host DNA repair mechanisms, emphasizing its ancient lineage and its role in the adaptation and function of retrotransposons such as BART. This domain, together with the RT domain, forms an important functional unit for BART's ability to mediate precise genome integration.
[0334] Although the sequence homology between the BART APE domain and its closest related human L1 APE is moderate—showing approximately 60–70% identity at the amino acid level—key residues essential for catalysis and DNA binding remain highly conserved. These residues include Glu43, Asp145, His230, Asn14, and Tyr115, which are important for coordinating catalytic metal ions, activating nucleophilic water, and correctly positioning DNA for cleavage. Despite evolutionary divergence, the retention of these active site residues underscores the functional integrity of the BART APE domain. Furthermore, structural modeling predicts the presence of a β-hairpin loop, a hallmark feature of APEs inserted into the minor groove of DNA substrates, further supporting BART's ability to precisely cleave DNA. Phylogenetic analysis reveals the evolutionary relationships between the BART APE domain and other APE domains, including human L1 APE, highlighting both the conservation of key functional residues and the evolutionary adaptations that differentiate BART.
[0335] BART proteins are characterized by a distinctive “tower” region spanning amino acid 240-440. This region comprises several subdomains: a base plate (residues 254-300), a tower helice (301-370), a tower lock (374-382), and a PIP box (404-419). The presence of this region suggests a potential role in the regulation of protein-protein interactions, RNA binding, and BART activity. The tower region may contribute to the stability and assembly of the BART complex, facilitating its interaction with target DNA and RNA molecules. Notably, the PIP box, known to mediate interactions with proliferating cell nuclear antigen (PCNA)—a key player in DNA replication and repair—suggests a possible link between BART activity and the host cell cycle. This connection further underscores the complex integration of BART with cellular processes, thereby enhancing its efficiency in gene editing and antiviral defense.
[0336] Despite conserved catalytic features, phylogenetic analysis revealed that the BART APE domain clustered distinctly from other APEs within the L1 clade, suggesting a potentially older origin or a different evolutionary trajectory. This distinct clustering suggests that while BART shares functional similarities with L1 APEs, it may have undergone unique adaptations that distinguish it from its related species. The presence of a DNase I-like motif in BART (absent in its L1 APEs) further supports the functional adaptation perspective. This motif could represent evolutionary refinement and may contribute to enhanced DNA cleavage specificity or efficiency in BART. These differences in BART activity and specificity highlight the potential for evolutionary fine-tuning of the APE domain, optimizing its specialization within the BART system.
[0337] C-terminal structural domain (CTD)
[0338] In contrast to the conserved RT and APE domains, the C-terminal region of BART shows significant truncation and modification when compared to retrotransposons. Several amino acid sequences essential for retrotransposition, particularly those involved in RNA binding and protein-protein interactions, are absent or significantly altered in BART. These modifications suggest that BART has evolved to relinquish its autonomous mobilization and replication capabilities, a hallmark of active retrotransposons. This loss of retrotransposition capability aligns with the idea of BART domestication, in which BART has been repurposed for novel cellular functions independent of its movement within the genome. This adaptation may be beneficial to the host, as uncontrolled transposition can be detrimental, potentially leading to genomic instability or insertional mutagenesis.
[0339] The C-terminal domain (CTD) of BART, while retaining some structural features from its retrotransposon ancestor, exhibits a markedly distinct adaptation, highlighting its evolutionary divergence and specialized role in the host cell. Notably, the "wrist" region, immediately downstream of the reverse transcriptase (RT) domain and spanning amino acids 863 to 1061, remains relatively conserved compared to the L1 retrotransposon. This conservation suggests that the wrist region continues to play a crucial role in BART function, potentially facilitating interactions with target DNA or other cytokines during integration. The preservation of this region underscores the evolutionary constraints imposed on this structural element, indicating its role in retrotransposons and their domesticated counterparts such as BART.
[0340] The truncation and modification of other regions within the CTD (including those involved in RNA binding and protein-protein interactions) indicate that BART has undergone significant structural evolution to adapt to its new functions. These changes may reflect a shift from retrotransposon activity to roles in gene regulation, antiviral defense, or other cellular processes, where precise and controlled activity is paramount. The differences between BART's CTD and its ancestral counterparts support the concept of its reuse and specialization, reinforcing the idea that domesticated elements like BART can evolve to achieve new functions within the host genome.
[0341] The CTD of BART, extending from amino acid 1062 to 1275, exhibits significant structural differences from its retrotransposon ancestor, most notably in the absence of the RNase H domain. In typical non-LTR retrotransposons and retroviruses, the RNase H domain is essential for replication, playing a crucial role in post-reverse transcriptional degradation of the RNA template and thus promoting the synthesis of complementary DNA strands. The absence of this domain in BART suggests a deep departure from the typical lifecycle of retrotranscribed elements, indicating that BART may no longer rely on the classical mechanism of RNA template degradation.
[0342] Further analysis of BART's CTD revealed the presence of a conserved CCHC motif (a zinc finger domain known for its ability to interact with nucleic acids). This motif, shared between L1 and BART, demonstrates its conservation from ancestral reverse transcription elements, highlighting its functional importance. Zinc finger domains, such as the CCHC motif, are typically involved in DNA binding, and in the context of BART, this domain may play a crucial role in recognizing and targeting specific genomic loci for integration. The preservation of this motif throughout evolution, even as other parts of the protein evolve and adapt to new functions, points to selective pressures to maintain certain structural elements.
[0343] In addition to the presence of the CCHC motif, BART's CTD exhibits significant modifications at key amino acid residues compared to L1. While the CCHC motif is conserved, other residues crucial for the catalytic activity of the L1 endonuclease domain are absent or altered in BART. These modifications further support the idea that BART's CTD has undergone adaptive changes, potentially fine-tuning its function to better suit its new role within the host cell. The loss of certain catalytic residues may signify a shift away from traditional endonuclease activity, suggesting that BART has evolved and achieved alternative functions in DNA repair, gene regulation, or targeted integration.
[0344] When compared to its L1 retrotransposon ancestor, BART's CTD also exhibits significant truncation, resulting in the loss of several α-helices and key amino acid residues essential for reverse transcription. This truncation is clearly visible in structural analysis, revealing a "stump" that effectively terminates at the zinc finger domain. The loss of these structural elements and key residues significantly impairs BART's ability to move autonomously and replicate, a hallmark of active retrotransposons. This loss of retrotransposability aligns with BART's evolutionary shift toward domestication, in which it has been repurposed for new cellular roles. In this context, the ability to move within the genome may be detrimental, potentially leading to genomic instability or harmful mutations.
[0345] The truncation of the CTD in BART is a compelling example of how evolutionary processes can selectively remove unnecessary or potentially harmful functions while preserving and refining those that offer advantages to the host organism. By removing components required for retrotransposition, BART may mitigate the risks associated with uncontrolled gene movement, allowing it to adopt more specialized and stable functions within the cell. This evolutionary refinement underscores the balance between maintaining essential functions and eliminating those that could threaten genome integrity. The structural and functional evolution of BART's CTD highlights the complex interplay between adaptation and conservation in the ongoing evolution of retrotransposon-derived elements.
[0346] In summary, the structural and functional analysis of the C-terminal domain of BART reveals an evolutionary process of adaptation and specialization. The conservation of the wrist region and the CCHC zinc finger motif highlights the importance of these elements in maintaining BART's function, particularly in DNA binding and structural stability. Conversely, the truncation and modifications observed in other regions of the CTD highlight the significant differences between BART and its retrotransposon ancestors, reflecting its evolutionary shift away from autonomous retrotransposons. The loss of retrotransposability, paired with the preservation of key enzyme functions and the emergence of new structural features, suggests that BART has been repurposed to meet new cellular needs, particularly in the fields of antiviral defense and gene editing.
[0347] This evolutionary refinement demonstrates how BART has adapted to enhance its utility in the host genome while removing redundant or potentially harmful functions.
[0348] Example 5
[0349] BART's antiviral function
[0350] The central hypothesis that the newly discovered BART system possesses specific anti-nucleic acid activity and the ability to encode genomic memory for infection was rigorously tested through a series of experiments. To evaluate this hypothesis, two variants of the BART enzyme (named BART.0 and BART.1) were cloned into human expression vectors and subsequently stably expressed in human cell lines. The expression of these bat-derived enzymes in human cells was closely monitored to assess any potential effects on cell viability and proliferation. Notably, the introduction and sustained expression of BART.0 and BART.1 did not adversely affect the growth or viability of human cells, indicating that the enzymes are well tolerated in the human cellular environment.
[0351] The antiviral potential of the BART system was tested using human cell lines engineered to stably express two variants of the BART enzyme, BART.0 and BART.1. Surprisingly, overexpression of BART conferred partial immunity against vesicular stomatitis virus (VSV) infection, highlighting its potential as a novel antiviral tool. Following VSV infection, cells expressing BART exhibited a significant reduction in viral load, indicating that BART is actively involved in cellular defense against viral pathogens (Figure 2).
[0352] Further analysis revealed that viral cDNA was detected in the cytoplasm of BART-expressing cells after infection, indicating that BART promotes the reverse transcription of viral RNA into cDNA. This viral cDNA is then processed into shorter fragments and integrated into the host genome in a scattered array pattern reminiscent of endogenous reverse transcription elements. This integration pattern suggests that BART not only inhibits viral replication but also incorporates the viral sequence into the host genome as genomic memory.
[0353] The crucial role of BART's reverse transcriptase (RT) activity in this antiviral response was further confirmed by the loss of viral resistance following pharmacological inhibition of the RT domain. These results demonstrate that BART's enzymatic activity is essential for its function in antiviral defense.
[0354] Furthermore, similar protective effects against monkeypox virus (MPV) further highlight BART's broad-spectrum antiviral capabilities. The combined evidence suggests that BART acts as a defense mechanism against viral infection, providing both immediate protection and genomic memory that can confer long-term immunity.
[0355] BART gene editing
[0356] The gene-editing potential of BART was evaluated by introducing BART into human cells along with specifically designed single-stranded DNA (ssDNA) oligonucleotides and an RNA payload. These oligonucleotides act as “navigators” guiding BART to specific genomic loci, while the RNA payload provides a template for precise gene modification. The system’s efficacy and specificity were assessed through sequencing, revealing successful mid-target editing without requiring a CRISPR-Cas system or its components.
[0357] a. In vitro targeted gene editing
[0358] Transgenic BART was extracted, transfected, expressed, and purified from the HEK293 cell line. The process began with transfection of HEK293 cells cultured in DMEM supplemented with 10% FBS, 1% penicillin-streptomycin, and 2 mM L-glutamine using a plasmid encoding BART via Lipofectamine 3000. Cells were transfected with the reagent-complexed plasmid at 70-80% confluency according to the manufacturer's protocol. Following transfection, cells were incubated at 37°C for 24-48 hours to allow protein expression, which was monitored using an eGFP reporter gene. Cells were then allowed to grow for an additional 48-72 hours to maximize BART expression. Cells were then harvested by trypsin-EDTA treatment and centrifugation, followed by washing the cell pellet with phosphate-buffered saline (PBS) to remove residual culture medium and debris.
[0359] For protein purification, BART protein was isolated from cell lysates using affinity chromatography, utilizing the His tag incorporated into the BART expression construct. The lysates were passed through a nickel-NTA agarose column, and after washing away unbound protein, BART was eluted with a buffer containing 250 mM imidazole. The concentration and purity of the purified BART protein were then assessed by SDS-PAGE and Western blotting.
[0360] To evaluate BART cleavage activity, in vitro assays targeting specific sequences within the HPRT and HSF1 loci were performed using two different navigation DNAs. The target and corresponding navigation sequences are detailed in Tables 4 and 5. The experiments began with the preparation of the navigation-target DNA complex. Navigation and target DNA were combined in equimolar amounts (typically 100 nM each) in a reaction buffer containing 10 mM Tris-HCl (pH 7.5), 50 mM NaCl, and 1 mM EDTA. The mixture was heated to 95 °C for 5 minutes to denature the DNA, followed by gradual cooling to room temperature over 30 minutes to promote proper hybridization between the navigation and target DNA strands.
[0361] Following hybridization, the reaction was assembled by adding 1 μL of navigation DNA (100 nM final concentration), 1 μL of target DNA, 0.5 μL of BART enzyme, and 1 μL of 10X BART reaction buffer (100 mM Tris-HCl, 500 mM NaCl, 10 mM MgCl2, pH 7.9) to the annealed DNA mixture, with nuclease-free water added to bring the final volume to 10 μL. The reaction mixture was incubated at 37°C for 30 minutes to allow BART-mediated cleavage at the target site. After the initial cleavage, 1 μL of strand displacement polymerase was added to extend the 3' end of the cleaved strand, followed by the addition of flap endonuclease 1 (FEN1) to induce a double-strand break at the target site.
[0362] To terminate the reaction and degrade the protein, 1 μL of proteinase K (20 mg / mL) was added to the reaction mixture, and the mixture was incubated at 55 °C for 10 min. The reaction products were then analyzed by agarose gel electrophoresis to assess cleavage efficiency, where successful cleavage was indicated by the presence of specific DNA fragments corresponding to the expected size.
[0363] Table 4: Target sequences used for in vitro cleavage assays. A summary of specific DNA sequences within the HPRT and HSF1 loci used as targets for BART-mediated cleavage assays, including the 250 nucleotides upstream and downstream of the cleavage site (underlined).
[0364]
[0365] Table 5: Navigation DNA sequences used for in vitro cleavage assays. A list of navigation DNA sequences designed to hybridize with target sequences and guide BART activity.
[0366]
[0367] These findings highlight the potential of BART as a highly precise gene-editing tool capable of inducing targeted DNA modifications with significant sequence specificity. The stringent experimental conditions and rigorous controls used in this assay further confirm the reliability and effectiveness of BART's cleavage activity under defined conditions, emphasizing its potential for precise genome editing applications.
[0368] b. In vivo targeted gene editing
[0369] The gene editing process using BART begins with the design and selection of navigation DNA (navDNA) specifically targeting the HPRT1 and HSF1 loci. These navDNAs play a crucial role in guiding the BART enzyme to the precise genomic location intended for editing. Each navDNA is engineered with three essential components: a target sequence, a primer-binding site (PBS), and a payload RNA (plRNA). The target sequence (a 20-nucleotide fragment) is selected based on its proximity to the 5'-TTTTT / AA-3' BART target motif within the HPRT1 and HSF1 genes (see Table 6 for specific sequences). The PBS, typically 13–17 nucleotides in length, is designed to anneal upstream of the editing site, facilitating the initiation of reverse transcription by the BART enzyme. The plRNA, detailed in Table 7, is constructed to include a short disruptive sequence followed by an insertion sequence encoding the enhanced green fluorescent protein (EGFP) gene. This design ensures that the BART enzyme accurately targets and modifies the desired locus.
[0370] NavDNA was synthesized using the GenScript biosynthesis platform to maintain high fidelity and prevent any sequence errors that could compromise the editing process. Specific loci within the HPRT1 and HSF1 genes, along with their corresponding target sites and BART target sequences, are presented, providing a visual representation of the precise locations where gene loci and BART-mediated editing occur. This strategy of designing navDNA and plRNA is intended to allow BART to efficiently and accurately modify target sequences, facilitating successful gene editing outcomes.
[0371] Table 6: Target sequences for navDNA. A summary of 20 nucleotide target sequences selected for the HPRT1 and HSF1 loci, including specific BART target sites.
[0372]
[0373]
[0374] Table 7: plRNA sequences for gene editing. Details of the plRNA sequences designed for the BART editing process, including disruptive and EGFP insertion sequences.
[0375]
[0376]
[0377]
[0378]
[0379] The next step involved cloning the BART enzyme coding sequence into a plasmid vector. The BART enzyme was placed under the regulation of a cytomegalovirus (CMV) promoter, known for its strong and constitutive expression in mammalian cells, to ensure robust enzyme production. Additionally, the plasmid included a neomycin resistance gene to allow selection of successfully transfected cells. The cloning process began with digestion of both the vector and the insert DNA for 1 hour at 37°C using EcoRI and HindIII restriction enzymes in NEBuffer 2 (10 mM Tris-HCl, 10 mM MgCl2, 50 mM NaCl, pH 7.9). Ligation was then performed overnight at 16°C using T4 DNA ligase in 1XT4 DNA ligase buffer (50 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP, 10 mM DTT, pH 7.5). The ligation mixture was then transformed into chemically competent DH5α *E. coli* cells via a 42°C heat shock for 45 seconds, followed by recovery in SOC medium (2% tryptone, 0.5% yeast extract, 10 mM NaCl, 2.5 mM KCl, 10 mM MgCl2, 20 mM glucose) at 37°C with shaking at 225 rpm for 1 hour. The transformed cells were plated on LB agar plates containing 50 μg / mL neomycin and incubated overnight at 37°C. Positive clones were selected, and plasmid DNA was extracted using the Qiagen Plasmid PlusMidi kit. The correct insertion of the BART sequence was then verified by Sanger sequencing (see vector map). Figure 4 ).
[0380] After successful cloning and validation of the plasmid construct, mammalian cells were transfected to express BART enzyme and navDNA. HEK293T cells were selected for this study; HEK293T is a human embryonic kidney cell line commonly used in transfection experiments due to its high transfection efficiency and robust growth. Cells were cultured at 2.5 x 102 cells per well. 5 Cells were seeded at a density of 1,000 cells per well in 6-well plates containing 2 mL of Duchenne Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS), 1% penicillin-streptomycin, and 2 mM L-glutamine. Cells were incubated overnight at 37°C in a humidified environment with 5% CO2 to allow them to reach 60-70% confluence, which is ideal for transfection.
[0381] On the day of transfection, BART plasmid DNA (1–2 μg per well) and nav DNA were diluted in 250 μL of Opi-MEM medium, a serum-depleted medium to enhance transfection efficiency. Separately, 2.5 μL of Lipofectamine 3000 reagent was added to the diluted DNA. The DNA and Lipofectamine 3000 mixture was gently mixed and incubated at room temperature for 15–20 minutes to allow for the formation of the DNA-lipid complex, which is essential for efficient DNA delivery into cells.
[0382] Before adding the transfection mixture to the cells, HEK293T cells were pre-washed with 1 mL of sterile phosphate-buffered saline (PBS) to remove any residual growth medium that might interfere with the transfection process. After washing, 2 mL of fresh Opti-MEM medium was added to each well. The DNA-lipid complex was then carefully added to the cells, ensuring even distribution throughout the wells. The cells were gently agitated back and forth to promote uniform exposure to the transfection complex.
[0383] Cells were incubated at 37°C in a humidified incubator with 5% CO2 for 4–6 hours to allow for optimal plasmid DNA uptake. This incubation period is important because it provides sufficient time for the DNA-lipid complex to fuse with the cell membrane, allowing plasmid DNA to enter the cells. After the initial transfection phase, the medium containing the transfection reagent was carefully aspirated and replaced with 2 mL of fresh DMEM supplemented with 10% FBS, 1% penicillin-streptomycin, and 2 mM L-glutamine. Cells were then incubated for an additional 48–72 hours to allow for optimal expression of BART enzymes and nav DNA.
[0384] To select successfully transfected cells, 500 μg / mL neomycin was added to the culture medium 24 hours post-transfection. Neomycin selection was maintained for 5–7 days, during which time cells were monitored daily. Dead cells were removed by gentle washing with PBS, and the medium was replaced with fresh neomycin-containing DMEM every 2–3 days to ensure continuous selection pressure. Viable cells were amplified for further analysis. The efficiency of the transfection and selection process was subsequently confirmed by fluorescence microscopy (for eGFP expression) and PCR analysis of genomic DNA extracted from the selected cell population, validating the presence and expression of BART and navDNA constructs.
[0385] To confirm successful gene editing at the HPRT1 and HSF1 loci, genomic DNA was extracted from transfected HEK293T cells using the Qiagen DNeasy Blood & Tissue Kit, following the manufacturer's protocol. Cells were first harvested by trypsin digestion, and the cell pellet was washed with phosphate-buffered saline (PBS) to remove any remaining culture medium. Genomic DNA was then extracted by lysing the cells with proteinase K and lysis buffer provided in the kit, followed by binding the DNA to a silica membrane in a centrifuge column. The membrane was washed with a series of ethanol-containing buffers to remove contaminants, and the DNA was eluted in nuclease-free water. The concentration and purity of the extracted DNA were measured using a NanoDrop spectrophotometer, with an A260 / A280 ratio of approximately 1.8 indicating high-quality DNA suitable for downstream applications.
[0386] Next, PCR primers were designed to be located flanking the regions within the HPRT1 and HSF1 loci targeted for editing (primer sequences are shown in Table 8). These primers were synthesized to amplify both unmodified and modified sequences, allowing for the detection of successful gene editing events. PCR reactions were established in a 50 μL volume containing 1X ThermoPol reaction buffer (20 mM Tris-HCl, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1% Triton X-100, pH 8.8), 200 μM of each dNTP, 0.2 μM of each primer, 1.25 U of Taq DNA polymerase, and 100 ng of template DNA. PCR cycling conditions consisted of an initial denaturation step at 95 °C for 3 min to completely denature the DNA, followed by 35 cycles of 95 °C for 30 s, 58 °C for 30 s (for primer annealing), and 72 °C for 1 min (for DNA extension). The final extension step at 72°C for 5 minutes ensures the complete synthesis of all PCR products.
[0387] PCR products were analyzed by electrophoresis on a 1.5% agarose gel containing 0.5 μg / mL ethidium bromide in 1X TAE buffer (40 mM Tris-acetate, 1 mM EDTA, pH 8.3). The gel was run at 100 V for 45 min to separate DNA fragments by size. Visualization was performed using a UV transilluminator, where successful insertion of the EGFP sequence was indicated by the presence of a PCR product approximately 720 bp larger than the wild-type amplicon, corresponding to the size of the EGFP insertion. Comparison of band sizes on the gel confirmed the presence of the desired gene modification.
[0388] Following gel electrophoresis, the PCR products were purified using a Qiagen PCR purification kit, which involves binding the DNA fragment to a silica membrane in a centrifuge column, washing with an ethanol-based buffer, and eluting the purified DNA in nuclease-free water. The purified products were then sequenced using the BigDye Terminator v3.1 Cyclic Sequencing Kit on an ABI 3730 DNA Analyzer. Sequencing reactions were prepared by combining the purified PCR products, sequencing primers, and BigDye reagents in a thermal cycler, following the manufacturer's instructions. Sequence analysis was performed using Geneious software, where the edited sequence was aligned to a wild-type reference sequence to confirm precise integration of the EGFP gene. The sequencing results clearly demonstrated successful gene editing, presenting alignments of the edited sequence to the reference genome and validating the accurate insertion of the EGFP sequence.
[0389] Table 8: PCR primer sequences for the HPRT1 and HSF1 loci. A list of PCR primers designed to be located flanking the targeted editing regions within the HPRT1 and HSF1 loci. These primers are used to amplify both unmodified and modified sequences, thereby allowing for the detection and validation of successful gene editing.
[0390]
[0391] To further validate the success of gene editing, EGFP expression in transfected HEK293T cells was analyzed using both fluorescence microscopy and flow cytometry. Following transfection and selection, cells were gently washed three times with phosphate-buffered saline (PBS) to remove any residual culture medium and dead cells. Cells were then fixed with 4% paraformaldehyde in PBS at room temperature for 10 minutes to preserve cell structure and fluorescence signal. After fixation, cells were washed three times with PBS to remove any residual paraformaldehyde.
[0392] The fixed cells were then mounted on slides using a VECTASHIELDHardSet antifluorescence quenching mounting medium containing DAPI (a nuclear counterstain that fluoresces blue under UV light) to allow visualization of the cell nuclei. Fluorescence images were captured using a Zeiss Axio Observer fluorescence microscope equipped with an EGFP filter array (excited at 488 nm, emitted at 530 nm) to specifically detect EGFP expression. Multiple fields of view imaging was performed to ensure representative sampling of the cell population. Image analysis was performed using ImageJ software, where the presence of EGFP-positive cells was quantified to confirm successful gene editing at the targeted locus. Analysis of the intensity of EGFP fluorescence and the number of EGFP-positive cells in the images provided qualitative evidence of successful gene editing.
[0393] To more quantitatively analyze EGFP expression, transfected cells were harvested by trypsin digestion. Cells were incubated with 0.25% trypsin-EDTA at 37°C for 2–3 minutes to separate them from the culture plate, and then neutralized with DMEM containing 10% FBS. The cell suspension was then passed through a 40 μm cell filter to obtain a single-cell suspension, which is crucial for accurate flow cytometry analysis. Cells were resuspended in 500 μL of ice-cold PBS to maintain cell viability and integrity during analysis.
[0394] The prepared cell suspensions were analyzed using a BD FACSCanto II flow cytometer. EGFP expression was detected by exciting cells at 488 nm and measuring emission at 530 nm. Data acquisition was set to collect at least 10,000 events per sample to ensure statistically significant results. The percentage of EGFP-positive cells was determined using FlowJo software, which allows for gating of cell populations and precise quantification of EGFP expression. This analysis provides a robust indicator of the efficiency of the prime editing process, with a high percentage of EGFP-expressing cells confirming the successful introduction of the desired gene modification via the BART enzyme.
[0395] Using this detailed protocol, gene editing with BART was successfully achieved, introducing targeted gene disruption at the HPRT1 and HSF1 loci while simultaneously inserting the EGFP gene. This process relies on the precise design of the navigation DNA (navDNA) and payload RNA (plRNA), which are customized to target specific sequences within the gene of interest. These components are crucial in guiding the BART enzyme to the exact location of the intended gene modification. Transfection of HEK293T cells with these constructs was highly efficient, ensuring a significant proportion of cells were incorporated into the BART and navDNA constructs. Subsequent selection and validation steps confirmed the successful editing of the target loci, as demonstrated by the disruption of the HPRT1 and HSF1 sequences and the precise insertion of the EGFP reporter gene. This protocol not only demonstrates the effectiveness of BART as a gene editing tool but also underscores the importance of meticulous experimental design and optimization for achieving reliable and specific gene modifications in mammalian cells.
[0396] The scope of this disclosure is not limited to the specific embodiments described herein. In fact, various modifications to the invention, in addition to those described herein, will become apparent to those skilled in the art from the foregoing description and drawings. Such modifications are intended to fall within the scope of the appended claims.
[0397] All references cited in this article are incorporated herein by reference in their entirety.
Claims
1. A method for modifying a target polynucleotide, comprising delivering to the target polynucleotide an enzyme having reverse transcriptase activity and endonuclease activity, and one or more nucleic acid components. The one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide. The enzyme described therein modifies the target polynucleotide.
2. A method for modifying the expression of a target polynucleotide, comprising: Introducing an enzyme with reverse transcriptase and endonuclease activity, or a nucleic acid molecule encoding said enzyme, and one or more nucleic acid components into cells or a subject. The one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide. The enzyme binds to one or more sites on the target polynucleotide, such that the binding of the enzyme increases or decreases the expression level of the target polynucleotide.
3. The method according to any one of the preceding claims, wherein the enzyme further has integrase activity.
4. The method according to any one of the preceding claims, wherein the one or more nucleic acid components comprise single-stranded navigation DNA.
5. The method according to any one of the preceding claims, wherein the one or more nucleic acid components further comprise a payload RNA.
6. The method of claim 5, wherein the payload RNA comprises a stem-loop structure.
7. The method of claim 5, wherein the enzyme reverse transcribes the payload RNA into cDNA.
8. The method of claim 7, wherein the enzyme integrates the cDNA into the target polynucleotide.
9. The method according to any one of the preceding claims, wherein the enzyme comprises an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain.
10. The method according to any one of the preceding claims, wherein the enzyme further comprises a C-terminal domain.
11. The method of claim 10, wherein the C-terminal domain facilitates the interaction between the enzyme and the target polynucleotide.
12. The method of claim 10, wherein the C-terminal structural domain includes a zinc finger.
13. The method of claim 12, wherein the zinc finger comprises a CCHC motif.
14. The method according to any one of the preceding claims, wherein the enzyme comprises bat-associated reverse transcriptase (BART).
15. The method of claim 9, wherein the BART comprises horseshoe bat, rat eared bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
16. The method according to any one of the preceding claims, wherein the enzyme is provided by one or more polynucleotide molecules encoding the enzyme.
17. The method according to any one of the preceding claims, wherein the one or more nucleic acid components are provided by one or more polynucleotide molecules encoding or comprising the one or more nucleic acid components.
18. The method according to any one of claims 15 to 17, wherein the one or more polynucleotide molecules comprise one or more carriers.
19. The method according to any one of claims 15 to 18, wherein the enzyme and the one or more nucleic acid components are provided in a single carrier.
20. The method according to any one of the preceding claims, wherein the target polynucleotide comprises a genomic locus.
21. The method according to any one of the preceding claims, wherein the target polynucleotide comprises RNA or DNA.
22. The method of claim 21, wherein the RNA comprises viral RNA of an RNA virus.
23. The method of claim 22, wherein the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
24. The method of claim 21, wherein the DNA comprises genomic DNA or cDNA.
25. The method according to any one of claims 21 to 24, wherein the modification of the target polynucleotide comprises the cleavage of the target polynucleotide.
26. The method according to any one of claims 21 to 25, wherein the target polynucleotide is contained in a nucleic acid molecule intracellularly or in vitro.
27. The method of claim 26, wherein the cell comprises a eukaryotic cell.
28. The method of claim 27, wherein the eukaryotic cell comprises a mammalian cell.
29. The method of claim 27, wherein the eukaryotic cell comprises a non-human animal cell, a human cell, or a plant cell.
30. A method for treating or preventing viral infection of an RNA virus in a cell or a subject, comprising delivering to the cell or the subject an enzyme having reverse transcriptase activity and endonuclease activity, or a nucleic acid molecule encoding the enzyme. The single-stranded guide polynucleotide hybridizes with the viral RNA of the RNA virus and directs the binding of the enzyme to the viral RNA. The enzyme described therein cleaves the viral RNA.
31. The method of claim 30, further comprising reverse transcribing the viral RNA into cDNA using the enzyme.
32. The method of claim 30, wherein the enzyme has integrase activity, and wherein the method comprises integrating the cDNA into the genome of the cell or the subject via the enzyme.
33. The method of claim 30, further comprising transcribing the cDNA into a single-stranded guide polynucleotide capable of hybridizing with the viral RNA.
34. A method for enhancing immunity against viral infection by an RNA virus in cells or a subject, comprising delivering to the cells or the subject an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, or a nucleic acid molecule encoding the enzyme. The enzyme reverse transcribes the viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA. The enzyme described therein integrates the cDNA into the cell or the host genome of the subject. The cDNA integrated into the host genome is transcribed into mRNA. The enzyme described therein reverse transcribes the mRNA into a second ssDNA. The second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA. The enzyme described therein cleaves the viral RNA.
35. A method for enhancing immunity against viral infection by an RNA virus in cells or a subject, comprising delivering to the cells or the subject an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, or a nucleic acid molecule encoding the enzyme. The enzyme reverse transcribes the viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA. The enzyme described therein integrates the cDNA into the cell or the host genome of the subject. The cDNA integrated into the host genome is transcribed into mRNA, and The mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
36. A method for generating a cell line with immunity against viral infection by an RNA virus, comprising delivering to the cells an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, or a nucleic acid molecule encoding said enzyme. The enzyme reverse transcribes the viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA. The enzyme described therein integrates the cDNA into the cell or the host genome of the subject. The cDNA integrated into the host genome is transcribed into mRNA. The enzyme described therein reverse transcribes the mRNA into a second ssDNA. The second ssDNA hybridizes with the viral RNA and directs the binding of the enzyme to the viral RNA. The enzyme described therein cleaves the viral RNA.
37. A method for generating a cell line with immunity against viral infection by an RNA virus, comprising delivering to the cells an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, or a nucleic acid molecule encoding said enzyme. The enzyme reverse transcribes the viral RNA of the RNA virus into single-stranded DNA (ssDNA) to generate cDNA of the viral RNA. The enzyme described therein integrates the cDNA into the cell or the host genome of the subject. The cDNA integrated into the host genome is transcribed into mRNA, and The mRNA forms an RNA hybrid with the viral RNA to silence the viral RNA.
38. The method according to any one of claims 30 to 36, wherein the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
39. The method according to any one of claims 30 to 36, wherein the enzyme comprises an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain.
40. The method according to any one of claims 30 to 38, wherein the enzyme further comprises a C-terminal domain.
41. The method of claim 40, wherein the C-terminal domain facilitates the interaction between the enzyme and the target polynucleotide.
42. The method of claim 40, wherein the C-terminal structural domain comprises a zinc finger.
43. The method of claim 40, wherein the zinc finger comprises a CCHC motif.
44. The method according to any one of claims 30 to 43, wherein the enzyme comprises bat-associated reverse transcriptase (BART).
45. The method of claim 44, wherein the BART comprises horseshoe bat, rat eared bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
46. The method according to any one of claims 30 to 45, wherein the enzyme is provided by one or more polynucleotide molecules encoding the enzyme.
47. A gene editing system for modifying target polynucleotides, comprising an enzyme having reverse transcriptase activity and endonuclease activity, and one or more nucleic acid components. The one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide. The enzyme described therein modifies the target polynucleotide.
48. The gene editing system of claim 47, wherein the enzyme further has integrase activity.
49. The gene editing system of claim 47, wherein one or more nucleic acid components are single-stranded navigation DNA.
50. The gene editing system according to any one of claims 47 to 48, wherein the one or more nucleic acid components further comprise payload RNA.
51. The gene editing system of claim 50, wherein the payload RNA comprises a stem-loop structure.
52. The gene editing system of claim 49, wherein the enzyme reverse transcribes the payload RNA into cDNA.
53. The gene editing system according to any one of claims 47 to 52, wherein the enzyme integrates the cDNA into the target polynucleotide.
54. The gene editing system according to any one of claims 47 to 53, wherein the enzyme comprises an N-terminal depurinyl-depyrimidine endonuclease domain optionally linked to a reverse transcriptase domain.
55. The system according to any one of claims 47 to 54, wherein the enzyme further comprises a C-terminal domain.
56. The gene editing system of claim 55, wherein the C-terminal domain facilitates the interaction between the enzyme and the target polynucleotide.
57. The gene editing system of claim 55, wherein the C-terminal domain comprises a zinc finger.
58. The gene editing system of claim 57, wherein the zinc finger comprises a CCHC motif.
59. The gene editing system according to any one of claims 47 to 58, wherein the enzyme comprises bat-associated reverse transcriptase (BART).
60. The gene editing system of claim 59, wherein the BART comprises BART of horseshoe bat or rat ear bat, European badger, wild yak, goat, Homo sapiens, domestic dog, wild boar, common marmoset, house mouse, brown bear, Indian elephant, or variants thereof.
61. The gene editing system according to any one of claims 47 to 60, comprising one or more polynucleotide molecules encoding the enzyme.
62. The gene editing system according to any one of claims 47 to 60, comprising one or more polynucleotide molecules encoding or including said one or more nucleic acid components.
63. The gene editing system according to any one of claims 61 to 62, wherein the one or more polynucleotide molecules comprise one or more vectors.
64. The gene editing system according to any one of claims 61 to 63, wherein the enzyme and the one or more nucleic acid components are provided in a single vector.
65. The gene editing system according to any one of claims 47 to 64, wherein the target polynucleotide comprises RNA or DNA.
66. The gene editing system of claim 65, wherein the RNA comprises viral RNA of an RNA virus.
67. The gene editing system of claim 66, wherein the RNA virus is selected from the group consisting of: norovirus, rotavirus, poliovirus, Ebola virus, green monkey virus, Lassa virus, Hantavirus, rabies virus, influenza virus, yellow fever virus, coronavirus, SARS, SARS-CoV-2, West Nile virus, hepatitis A, C (HCV) and E virus, dengue virus, cloacal virus, rhabdovirus, piconemavirus, myxovirus, retrovirus, Bunyavirus, coronavirus and reovirus.
68. The gene editing system of claim 65, wherein the DNA comprises genomic DNA or cDNA.
69. The gene editing system according to any one of claims 47 to 68, wherein the modification of the target polynucleotide comprises the cleavage of the target polynucleotide.
70. A delivery system comprising the gene editing system according to any one of claims 47 to 69, wherein the delivery system is adapted to deliver the gene editing system to a cell or a subject.
71. The delivery system of claim 70, comprising nanoparticles or vesicles encapsulating the gene editing system.
72. A vector system comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, and comprises one or more nucleic acid components. The one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide. The enzyme described therein modifies the target polynucleotide.
73. A kit comprising the gene editing system according to any one of claims 47 to 69, the delivery system according to any one of claims 70 to 71, or the vector system according to claim 72.
74. A cell line comprising one or more polynucleotide molecules encoding an enzyme having reverse transcriptase activity, endonuclease activity, and integrase activity, and one or more nucleic acid components. The one or more nucleic acid components target and hybridize with the target polynucleotide, and direct the binding of the enzyme to the target polynucleotide. The enzyme described therein modifies the target polynucleotide.
75. The cell line according to claim 74, comprising eukaryotic cells.
76. The cell line of claim 75, wherein the eukaryotic cells include mammalian cells.
77. The cell line of claim 75, wherein the eukaryotic cells comprise stem cells or stem cell lines.
78. A method for preparing a cell line according to any one of claims 74 to 77, comprising introducing into the cell an enzyme encoding an enzyme having reverse transcriptase activity, endonuclease activity and integrase activity, and one or more polynucleotide molecules encoding one or more nucleic acid components.
Citation Information
Patent Citations
Meganuclease variants cleaving a DNA target sequence from the dystrophin gene and uses thereof
US20130145487A1
Crispr-related methods and compositions
WO2015048577A2
Crispr-CAS-related methods, compositions and components for cancer immunotherapy
WO2015161276A2