Compositions and methods for precise genome editing using retrons

WO2025010350A3PCT designated stage expired Publication Date: 2025-05-30BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/036763
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-03
Filing Date
2024-07-03
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current genome editing technologies face challenges in delivering and integrating exogenous template DNA for precise edits, particularly due to genotoxicity concerns with viral vectors and incompatibility with RNA-based delivery, limiting their efficiency and applicability in multiplexed applications.

Method used

Development of a retron editor system comprising non-coding RNA, reverse transcriptase, and a guide RNA, which generates high copies of long single-stranded DNA in vivo, optimized for use with nucleases like Cas9 or Cas12a to facilitate precise genome editing by templated homology-directed repair.

Benefits of technology

Enhances the precision and efficiency of genome editing by promoting targeted nucleic acid modifications, expanding the genomic target range and reducing off-target errors, while being compatible with RNA-based delivery systems.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Retron editors allow for the insertion of a segment of user-defined DNA in a host cell genome. To maximize the activity of retron editors, they were optimized by improved linkers, nuclear localization sequences, and RNA sequences. The resulting retron editor can programmably insert, delete, or change the genome of a host cell or organism, including human genomes.
Need to check novelty before this filing date? Find Prior Art

Description

COMPOSITIONS AND METHODS FOR PRECISE GENOME EDITING USING RETRONSCROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Applications 63 / 511,699, filed July 3, 2023, the contents of each are hereby incorporated in their entirety.SEQUENCE LISTING

[0002] A Sequence Listing conforming to the rules of WIPO Standard ST.26 is hereby incorporated by reference. Said Sequence Listing has been filed as an electronic document via Patentcenter encoded as XML in UTF-8 text. The electronic document, created on June 30, 2024, is entitled "10046-534W01_ST26.xml", and is 194,966 bytes in sizeINCORPORATION BY REFERENCE

[0003] All documents cited or referenced herein ("herein cited documents"), and all documents cited or referenced in herein cited documents, together with any manufacturer's instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the invention. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.BACKGROUND OF THE INVENTION

[0004] Precise genome editing is a cornerstone of biomedicine and gene therapy. However, installing specific edits via templated homology-directed repair (HDR) is limited by the challenge of delivering and integrating an exogenous template DNA into the genome [1-3], The most efficient delivery methods are viral vectors and synthetic DNA donors [4-7], Viral vectors can induce genotoxicity via insertional mutagenesis, are depleted during the gene editing window, and are not suitable for multiplexed applications [5, 7, 8], Synthetic DNA templates must be transfected or electroporated into cells, do not target the nucleus, and are incompatible with RNA-based delivery in cells and organisms [4, 6], Therefore, recent gene editing technologies have begun to harness reverse transcriptases (RTs) that can continuously generate the template DNA near its edit site. [9-35], Among these approaches,retron-RTs are especially promising tools because they are self-priming and can generate high copies of long single-stranded DNAs (ssDNAs) in vivo.

[0005] Retrons are bacterial anti-phage defense systems that consist of a self-priming reverse transcriptase (retron-RT), a cognate non-coding RNA (ncRNA) that primes and templates reverse transcription, and an accessory protein that participates in anti-viral immunity [36-44], The ncRNA consists of two main regions: the msr (msDNA-specific region) and the msd (multi-copy ssDNA-coding region) (see Fig. 1A). The msr is located at the 5' end of the ncRNA and forms a specific structure that is recognized by the RT [45, 46], This region typically contains one to three stable stem loops with 7-10 base pair stems and 3-10 nucleotide loops. The msr also includes a highly conserved guanosine residue at the 5' end, which serves as the branching point for initiating reverse transcription (see Fig. 1A). The msd is positioned downstream of the msr and can be divided into two parts: a dispensable region that can be replaced with a desired sequence (i.e., the donor DNA for genome editing) and a conserved region that is essential for the proper folding and function of the ncRNA. The RT primes from the msr, and uses the msd as a template [45, 46], The resulting msDNA remains covalently linked to the ncRNA through a 2',5'-phosphodiester bond formed between the branching guanosine residue in the msr and the 5' end of the msDNA [45, 46], The host RNAse H degrades the RNA-DNA hybrid to expose the ssDNA (see Fig. 1A, right)

[0039] , By replacing the dispensable msd region with a user-programmable sequence, retron-RTs can generate templates for homology-directed DNA repair in cells.

[0006] Recent studies have demonstrated the potential of retrons coupled with CRISPR- Cas9 to enhance precise genome editing in bacteria, yeast, plants, and mammalian cells [17, 19, 24, 30-32, 34], Early studies fused the Cas9 single guide RNA (sgRNA) to a retron ncRNA with a modified donor msd

[0032] , This increases the local concentration of the donor template as a double-stranded DNA break (DSB), biasing repair toward templated HDR (see Fig. 1C). Alternative designs fused the retron RT to Cas9 directly with the same goal [30, 34], Despite these advances, the potential of retrons for precise genome editing has yet to be fully realized. To date, only a handful of retron RTs have been benchmarked in mammalian cells.

[0007] What is needed in the art are optimized systems, as well as new retrons, which can be used to conduct gene editing. These DNA can be coupled to a nuclease such as Cas9, Casl2a, Zinc Finger Nucleases (ZFNs), Transcription activator-like effector nucleases(TALENs), or similar to enable genetic editing, targeted genome mutations, and other gene engineering applications.SUMMARY OF THE INVENTION

[0008] Disclosed herein is a retron editor system comprising: a) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence; b) a nucleotide sequence encoding a guide RNA (gRNA); and c) a nucleotide sequence encoding a nuclease.

[0009] Also disclosed herein is a vector comprising a) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence; b) a nucleotide sequence encoding a guide RNA (gRNA); and c) a nucleotide sequence encoding a nuclease.

[0010] Further disclosed herein is a method of using a retron editor system to edit target nucleic acid, the method comprising: a) providing a retron editor system, wherein said system comprises: i) a nucleotide sequence encoding a guide RNA (gRNA); ii) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence, wherein said msd sequence comprises a donor nucleic acid; and iii) a nucleotide sequence encoding a nuclease; b) expressing a product from the retron editor system; and c) placing the retron editor system under conditions such that gene editing of target nucleic acid takes place.

[0011] Also disclosed is a method of modifying one or more target nucleic acids of interest at one or more target loci in a host cell, the method comprising: a) transforming the host cell with a vector encoding a retron editing system; b) culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing a retron editor system from the vector; c) providing conditions suitable for the nuclease of the retron editor system to cut at or near the target loci; and d) providing conditions for the donor nucleic acid insertion sequence to recombine with the one or more target nucleic acid sequences to insert, delete, and / or substitute one or more bases of the sequence of the one or moretarget nucleic acid sequences to induce one or more sequence modifications at the one or more target loci.

[0012] Disclosed is a method of screening for functional retron editors, the method comprising: a) providing a potential retron editor system; b) transforming a cell with the potential retron editor system, wherein said cell has been modified to express a signal upon successful transformation using the gene-editing retron system; and c) detecting the presence of the signal.

[0013] Further disclosed is a retron editor system comprising: a) a first zinc finger nuclease (ZFN) or a ZFN nucleic acid sequence encoding a zinc finger nuclease, that binds a first area of a target region, wherein the first ZFN comprises a cleavage domain and a ZFN protein; b) a second ZFN or nucleic acid encoding a second ZFN that binds a second area of a target region, wherein the second ZFN comprises a cleavage domain and a second ZFN protein; wherein the first and second ZFN are capable of dimerization and cleavage of the target region; and c) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence.

[0014] Also disclosed is a retron editor system comprising: a) a first transcriptional activator-like effect nuclease (TALEN) or a nucleic acid encoding a first TALEN, wherein the first TALEN comprises a target region binding site and a nuclease; b) a second TALEN or a nucleic acid encoding a second TALEN, wherein the second TALEN comprises a target region binding site and a nuclease; and c) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence.BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1A-H shows results of a metagenomic survey reveals highly active RTs in mammalian cells. Figure 1A shows retron RTs self-prime from a non-coding RNA, termed the msr-msd. msr: gray; msd stem: black; variable region: black, in dotted box. Arrow indicates the direction of reverse transcription. Figure IB shows schematic of a retron editor. The RT is linked to Cas9 (shown) or another nuclease. Figure 1C shows reverse transcription of the variable region of the msd generates a ssDNA template for homology-directed repair of the cleavage site. Figure ID shows a plasmid-encoded fluorescent reporter assay. The RFP has a9bp deletion proximal to a Y64L mutation to completely turn off RFP fluorescence. The reporter is co-transfected with a plasmid that encodes the retron editor, along with an msd that repairs the RFP. RFP+ cells are imaged via confocal microscopy and quantified via flow cytometry. Figure IE shows confocal microscopy images of cells transfected with Cas9 + Ecol-RT (left), Cas9 + ssODN (middle), and Cas9-MvalRT (right). Figure IF shows phylogenetic classification of novel retron systems discovered from metagenomic sources. These systems are classified into clades, as described in

[0051] , Figure 1G shows rank ordered list of RFP repair efficiency with 98 metagenomically discovered retron-RTs using flow cytometry. Dashed line: RFP+ repair with Ecol-RT. Inset: flow cytometry data for Cas9 with a scrambled sgRNA (top left); RFP-targeting sgRNA (top right); RFP sgRNA and a ssODN repair template (bottom left); RFP sgRNA and Mval-RT (bottom right). Error bars: mean of three replicates. The three most active RTs are labeled. Figure 1H shows gene editing activity of the six most active retron-RTs, along with Ecol-RT with a cognate (diagonal) or non-cognate msr-msd. Flow cytometry was used to score activity with the transient RFP reporter. Mean of three replicates.

[0016] Figure 2A-B shows an overview of the bioinformatic retron discovery pipeline. Figure 2A shows the pipeline involves five steps: 1: predict open reading frames (ORFs) with Prodigal [1]; 2: annotate reverse transcriptase (RT) genes using HMMER [2]; 3: identify putative msr-msd sequences in non-coding regions using cmfinder and infernal [3, 4]; 4: reannotate adjacent ORFs with HMMER; and 5: manually inspect msr-msd structures with ViennaRNA. Figure 2B shows the analysis was conducted on 2,068,918 reference genomes from the human gut microbiome [5], along with 15,574 bacterial and 531 archaeal genomes from the NCBI database [6], After a 95% de-duplication at the amino acid sequence level and annotating msr-msds; 568 new candidate systems were identified.

[0017] Figure 3A-B shows a comparison of top RTs in transient and genomically-integrated RFP reporter. Figure 3A shows confocal images of the indicated RTs, or Cas9 with a scrambled sgRNA. Scale bar: 100 pm; inset: 50 pm. Figure 3B shows correlation between retron editing activity in a genomic vs. transient RFP reporter system. The genomic RFP reporter was integrated into the AAVSl locus.

[0018] Figure 4A-B shows sequence identity and structural analysis of highly active retron- RTs. Figure 4A shows sequence identity heatmap of RT's amino acid sequences and ncRNA sequences for retron candidates. Figure 4B shows structure of a representative msr-msdtranscript encoded by retron candidates. The structures are predicted by ViennaRNA. The msd region is highlighted by a dashed rectangle.

[0019] Figure 5A-E shows Efel-RT catalyzes precise genomic insertions across multiple loci. Figure 5A shows a schematic of the NGS library preparation strategy. Genomic DNA is first amplified with primers that are outside the homology arms to avoid amplifying the retron- synthesized msDNA. After gel extraction, a second round of PCR amplifies and barcodes the insertion site for deep sequencing. Blue, orange: universal Illumina P5 / I5 and P7 / I7 adapters and indices. Figure 5B shows normalized insertion efficiency for the top 5 retron-RTs at the CFTR and EMX1 loci. Error bars: mean of three replicates. Figure 5C shows the relative insertion frequency of a 10 nt cargo at the EMX1 locus, along with the four most frequent misincorporated sequences (shown, from top to bottom, are SEQ ID NOS: 125, 126, 127, 128, 129, and 130). The most common errors are a deletion or insertion at the periphery of the homology arms. Error bars: mean of three replicates. Dots represent individual replicates. Figure 5D shows Efel-RT substitution errors (left) are less frequent than ssODN insertion (right) at the EMX1. Substitution rates are computed from the insert in the EMX1 locus across three biological replicates. Figure 5E shows schematic (top) and results (bottom) of the insertion efficiency with Efel-RT as a function of the homology arm length. Templated insertion is most active with 50 nt homology arms at five genomic loci. Error bars indicate three replicates.

[0020] Figure 6A-G shows rational engineering of an Efel-RT-based retron editor. Figure 6A shows results of the effect of splitting the sgRNA and msr-msd (left), the identity of the nuclear localization sequences (NLSs), and the linker between the Cas9 and Efel-RT (bottom). Figure 6B shows results of expressing the sgRNA and msr-msd increased gene editing by 60% relative to a fused sgRNA-msr-msd design. Figure 6C shows optimization of the N- and C-terminal nuclear localization sequences (NLSs). Gray: reference design that was used for normalization. Error bars: mean of three replicates. Figure 6D shows optimization of the linker peptide between the Cas9 and Efel-RT. Retron editors tolerate a broad range of flexible (blue) and rigid (orange) linkers. Splitting the two enzymes via a ribosomal skipping peptide (T2A, light blue) also retains most activity. However, multimerization domains abrogated activity (green). Gray: reference design that was used for normalization. Error bars: mean of three replicates. Figure 6E shows schematic (top) and editing activity of a Casl2a-based retron editor at five genomic loci. Error bars: mean of three replicates.Figure 6F shows the relative insertion frequency of a 10 nt cargo at the BRD8 locus, along with the four most frequent misincorporated sequences. In contrast to Cas9-based retron editors, Casl2a editors generate substitution error in the insert. Error bars: mean of three replicates (dots). Shown in order from top to bottom are SEQ ID NOS: 131-136, respectively. Figure 6G shows Efel-RT substitution errors at the BRD8 locus. Substitution rates are computed from the insert in the locus. Error bars indicate three replicates.

[0021] Figure 7 shows Casl2a-based editing outcomes at the indicated loci.

[0022] Figure 8A-G shows inhibiting non-homologous end joining boosts templated insertion. Figure 8A shows schematics of two strategies that boost templated insertionleft: Cas9 is fused to proteins that alter DNA repair pathway choice and right: small molecule inhibition of DNA-dependent protein-kinase catalytic subunit (DNAPKcs) or CDC7. Figure 8B shows the effect of inhibitors (left) and Cas9 fusions (right) on the relative rate of templated insertion at five loci. AZD7648 (left, top) and Cas9-CtlP-dnRNF168 (right, top) both have the strongest effect at all tested loci. Open circles: editing with no inhibitor or DNA repair protein. Closed circle: editing with the indicated inhibitor or DNA repair protein fused to Cas9. All circles indicate the mean across three biological replicates. Arrow: change in editing efficiency with the indicated inhibitor. Fusions increase the relative rates of templated repair across all loci. Figure 8C shows a schematic of experiments with 50 nt homology arms and increasing insert lengths at the EMX1 locus. Figure 8D shows AZD7648 increases the insertion efficiency across all cargo sizes tested in this study. Error bars: mean of three replicates. Figure 8E shows AZD7648 outperforms TAK-931 and M3814 in boosting insertion efficiency at EMX1 without increasing mutational signature or Cas9-generated indels. Error bars: mean of three replicates. Figure 8F shows the insertion efficiency decreases for all Cas9-repair protein fusions at EMX1. Error bars: mean of three replicates. Figure 8G shows Cas9 fused to CtlP-dnRNF168 increased insertion efficiency of 10 nt cargo at EMX1 without increasing mutational signature or Cas9-generated indels compared to no DNA repair fusion, DN1S and hGeml / 110. Error bars indicate three replicates.

[0023] Figure 9A-E shows optimizing inhibitor and DNA repair fusions improves gene editing with nickase Cas9 (nCas9). Figure 9A shows optimization of inhibitor concentrations for optimal retron editing at the EMX1 locus. Dashed line: gene editing without inhibitors. Figure 9B shows the effect of combining DNA repair inhibitors with Cas9-DNA repair domain fusions. Combining the most active inhibitor, AZD7468, with Cas9-CtlP-dnRNF168 reducesoverall insertion activity. In other cases, AZD7468 improves retron editing, likely due to the limited improvement observed with Cas9-DN1S and Cas9-hGeml / 100 fusions. Figure 9C shows illustration of the nickase Cas9(D10A) nickase-DNA repair fusions, and a nickasebased retron editor. Figure 9D shows fusing Cas9(D10A) with CtlP-dnRNF168, hRad51, and Rep-X helicase [7-9] increases the relative rates of templated repair at EMX1. Open circles: editing with no repair factor fusion. Closed circle: editing with the indicated repair factor fusion. All circles indicate the mean across three biological replicates. Arrow: change in editing efficiency with the indicated Cas9-DNA repair protein fusion. Figure 9E shows breakdown of editing outcomes, as reported by NGS read counts. Bottom: retron editor- mediated insertion; middle: nCas9-only edits without any insert, top: unmodified reads. Error bars indicate three biological replicates.

[0024] Figure 10A-C shows in-frame precise epitope insertion with retron editors. Figure 10A shows schematic of the split super-folder GFP system. The 11th GFP -strand (GFP11) is expressed as a fusion to the protein of interest (POI). Reconstitution of GFP11 with GFP1-10 restores fluorescence. Figure 10B shows GFP1-10 is expressed from a genomically- integrated inducible promoter. The POI-GFP11 is in its native genomic locus. Figure 10C shows confocal imaging of three GFPll-protein fusions. Scale bars: 10 pm.

[0025] Figure 11A-G shows retron editor delivery via an all-RNA package. Figure 11A shows an illustration of RNA-based retron editing. Cas9 and Efel-RT are delivered as capped and poly-A tailed mRNAs. The sgRNA and msr-msd are mixed with the mRNAs in a 10:1 molar ratio prior to transfection into HEK293T cells. Figure 11B shows insertion efficiency of a 10 nt insert at the indicated loci with RNA-based delivery. Error bars: mean of three replicates. Dots represent individual replicates. Figure 11C shows the relative insertion frequency of a 10 nt cargo at the EMX1 locus, along with the four most frequent misincorporated sequences. Dots represent individual replicates. Shown from top to bottom, respectively, are SEQ ID NOS: 125, 126, and 139-142. Figure 11D shows a schematic of the three substitutions that are introduced in kif 6 by mRNA injection. The first mutation is silent but abolishes a Bsal cut site. Shown from top to bottom are SEQ ID NOS: 137 and 138. Figure HE shows deep sequencing of embryos 24 hr post RNA injection for the indicated conditions. Each point is a single embryo, p-values are determined by a One-Way Anova test. Figure 11F shows Bsal restriction enzyme digests of the edit site sub-cloned from WT, kif6A, and edited embryos. The Bsal cut site is abolished in edited, but not WT orkif6Aembryos. Figure 11G shows Sanger sequencing of the edited embryo from Figure HE confirms precise editing at the three expected sites (triangles). The sequence shown is SEQ ID NO: 138.

[0026] Figure 12A-D shows a schematic for exemplary embodiments of retron editor cassettes. Figure 12A shows a retron editing cassette comprising a nucleotide sequence encoding Cas9. Figure 12B shows a retron editing cassette comprising a nucleotide sequence encoding Casl2a. Figure 12C shows a retron editing cassette comprising a nucleotide sequence encoding zinc finger nuclease (ZFN). Figure 12D shows a retron editing cassette comprising a nucleotide sequence encoding TALENs. In this figure, "retron" can refer to the reverse transcriptase only.

[0027] Figure 13 shows the structure of Retron Editor 2.0. It includes a nuclear localization signal (NLS), a nuclease, a linker, and sgRNA and msr-msd extension and expression, as well as msr-msd structure / codon juggling.

[0028] Figure 14A-B shows various linker designs that were experimentally tested. Figure 28A shows flexible linkers (5 total) and rigid linkers (6 total). Figure 28B shows dimerization / multimerization polyvalent linkers (5 total).

[0029] Figure 15 shows linker results for NRT-49 with a split ncRNA.

[0030] Figure 16 shows the designs of a series of NLS sequences that were tested experimentally.

[0031] Figure 17 shows an editing summary of NRT-49. Shown are SEQ ID NO: 79 and SEQ ID NO: 80.

[0032] Figure 18A-C shows the use of nickase along with the retron editor. (A) is a schematic. (B) shows proteins which increase the relative rates of templated repair. (C) shows percentage of editing event using various proteins.

[0033] Figure 19 shows results from DNA repair fusion domains and Casl2a.DETAILED DESCRIPTION OF THE INVENTIONDEFINITIONS

[0034] Unless specifically indicated otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. In addition, any method or material similar orequivalent to a method or material described herein can be used in the practice of the present invention. For purposes of the present invention, the following terms are defined.

[0035] The terms "a," "an," or "the" as used herein not only include aspects with one member, but also include aspects with more than one member. For instance, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells and reference to "the agent" includes reference to one or more agents known to those skilled in the art, and so forth.

[0036] The term "about" in relation to a reference numerical value can include a range of values plus or minus 10% from that value. For example, the amount "about 10" includes amounts from 9 to 11, including the reference numbers of 9, 10, and 11. The term "about" in relation to a reference numerical value can also include a range of values plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% from that value.

[0037] As used herein, unless otherwise specified, the terms "5"' and "3"' denote the positions of elements or features relative to the overall arrangement of the retron- guide RNA cassettes, vectors, or retron donor DNA-guide molecules of the present invention in which they are included. Positions are not, unless otherwise specified, referred to in the context of the orientation of a particular element or features. Unless otherwise specified, the term "upstream" refers to a position that is 5' of a point of reference. Conversely, the term "downstream" refers to a position that is 3' of a point of reference.

[0038] The term "gene editing" or "genome editing" refers to a type of genetic engineering in which DNA is inserted, replaced, or removed from a target DNA (e.g., the genome of a cell) using one or more nucleases and / or nickases. The nucleases create specific double-strand breaks (DSBs) at desired locations in the genome, and harness the cell's endogenous mechanisms to repair the induced break by homology-directed repair (HDR) (e.g., homologous recombination) or by nonhomologous end joining (NHEJ). The nickases create specific single-strand breaks at desired locations in the genome. In one non-limiting example, two nickases can be used to create two single-strand breaks on opposite strands of a target DNA, thereby generating a blunt or a sticky end. Any suitable DNA nuclease can be introduced into a cell to induce genome editing of a target DNA sequence. In some embodiments, a nickase can be used in place of a nuclease in the retron editors described herein.

[0039] The term "programmable nuclease" refers to an enzyme capable of cleaving the phosphodiester bonds between the nucleotide subunits of DNA, and may be an endonuclease or an exonuclease. According to the present invention, the programmable nuclease may be an engineered so that it can be used to induce gene editing of a target nucleic acid sequence. Any suitable nuclease can be used including, but not limited to, CRISPR-associated protein (Cas) nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or other endo- or exo-nucleases, variants thereof, fragments thereof, and combinations thereof.

[0040] The term "double-strand break" or "DSB" or "double-strand cut" refers to the severing or cleavage of both strands of the DNA double helix. The DSB may result in cleavage of both stands at the same position leading to "blunt ends" or staggered cleavage resulting in a region of single-stranded DNA at the end of each DNA fragment, or "sticky ends". A DSB may arise from the action of one or more DNA nucleases.

[0041] The term "nonhomologous end joining" or "NHEJ" refers to a pathway that repairs double-strand DNA breaks in which the break ends are directly ligated without the need for a homologous template.

[0042] The term "homology-directed repair" or "HDR" refers to a mechanism in cells to accurately and precisely repair double-strand DNA breaks using a homologous template to guide repair. One common form of HDR is homologous recombination (HR), a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical molecules of DNA.

[0043] The term "nucleic acid," "nucleotide," or "polynucleotide" refers to deoxyribonucleic acids (DNA), ribonucleic acids (RNA) and polymers thereof in either single- , double- or multi-stranded form. The term includes, but is not limited to, single-, double- or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and / or pyrimidine bases or other natural, chemically modified, biochemically modified, non-natural, synthetic or derivatized nucleotide bases. In some embodiments, a nucleic acid can comprise a mixture of DNA, RNA and analogs thereof. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservativelymodified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)).

[0044] The term "single nucleotide polymorphism" or "SNP" refers to a change of a single nucleotide within a polynucleotide, including within an allele. This can include the replacement of one nucleotide by another, as well as the deletion or insertion of a single nucleotide. Most typically, SNPs are biallelic markers although tri- and tetra-a llelic markers can also exist. By way of non-limiting example, a nucleic acid molecule comprising SNP A\C may include a C or A at the polymorphic position.

[0045] The term "gene" means the segment of DNA involved in producing a polypeptide chain. The DNA segment may include regions preceding and following the coding region (leader and trailer) involved in the transcription / translation of the gene product and the regulation of the transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).

[0046] Generally, "host" refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, a "host cell" refers to an in vivo or in vitro eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell), or cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, into which a heterologous polynucleotide or polypeptide has been introduced. In some embodiments, the cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-cell organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, in invertebrate cell, a vertebrate cell, a fish cell, a frog cell, a bird cell, an insect cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some cases, the cell is in vitro. In some cases, the cell is in vivo.

[0047] The term "recombinant" refers to an artificial combination of two otherwise separated segments of sequence, e.g., by chemical synthesis, or manipulation of isolated segments of nucleic acids by genetic engineering techniques.

[0048] The terms "plasmid", "vector" and "cassette" refer to a linear or circular extra chromosomal element often carrying genes that are not part of the central metabolism of the cell, and usually in the form of double-stranded DNA. Such elements may be autonomously replicating sequences, genome integrating sequences, phage, or nucleotide sequences, in linear or circular form, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a polynucleotide of interest into a cell.

[0049] "Transformation cassette" refers to a specific vector comprising a gene and having elements in addition to the gene that facilitates transformation of a particular host cell. "Expression cassette" refers to a specific vector comprising a gene and having elements in addition to the gene that allow for expression of that gene in a host.

[0050] The terms "recombinant DNA molecule", "recombinant DNA construct", "expression construct", "construct", and "recombinant construct" are used interchangeably herein. A recombinant DNA construct comprises an artificial combination of nucleic acid sequences, e.g., regulatory and coding sequences that are not all found together in nature. For example, a recombinant DNA construct may comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source but arranged in a manner different than that found in nature. Such a construct may be used by itself or may be used in conjunction with a vector. If a vector is used, then the choice of vector is dependent upon the method that will be used to introduce the vector into the host cells as is well known to those skilled in the art. For example, a plasmid vector can be used. The skilled artisan is well aware of the genetic elements that must be present on the vector in order to successfully transform, select and propagate host cells. The skilled artisan will also recognize that different independent transformation events may result in different levels and patterns of expression (Jones et al. , (1985) EMBO J 4:2411-2418; De Almeida et al. , (1989 )Mol Gen Genetics 218:78-86), and thus that multiple events are typically screened in order to obtain lines displaying the desired expression level and pattern. Such screeningmay be accomplished standard molecular biological, biochemical, and other assays including Southern analysis of DNA, Northern analysis of mRNA expression, PCR, real time quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting analysis of protein expression, enzyme or activity assays, and / or phenotypic analysis.

[0051] The term "heterologous" refers to the difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. As used herein, "heterologous" in reference to a sequence can refer to a sequence that originates from a different species, variety, foreign species, or, if from the same species, is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, a promoter operably linked to a heterologous polynucleotide is from a species different from the species from which the polynucleotide was derived, or, if from the same / analogous species, one or both are substantially modified from their original form and / or genomic locus, or the promoter is not the native promoter for the operably linked polynucleotide.

[0052] The term "expression", as used herein, refers to the production of a functional end-product (e.g., an mRNA, guide RNA, or a protein) in either precursor or mature form.

[0053] A "mature" protein refers to a post-translationally processed polypeptide (i.e., one from which any pre- or propeptides present in the primary translation product have been removed). "Precursor" protein refers to the primary product of translation of mRNA (i.e., with pre- and propeptides still present). Pre- and propeptides may be but are not limited to intracellular localization signals. The term "operably linked" refers to two or more genetic elements, such as a polynucleotide coding sequence and a promoter, placed in relative positions that permit the proper biological functioning of the elements, such as the promoter directing transcription of the coding sequence.

[0054] The term "inducible promoter" refers to a promoter that responds to environmental factors and / or external stimuli that can be artificially controlled in order to modify the expression of, or the level of expression of, a polynucleotide sequence or refers to a combination of elements, for example an exogenous promoter and an additional element such as a trans-activator operably linked to a separate promoter. An inducible promoter may respond to abiotic factors such as oxygen levels or to chemical or biologicalmolecules. In some embodiments, the chemical or biological molecules may be molecules not naturally present in humans.

[0055] The terms "vector" and "expression vector" refer to a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular polynucleotide sequence in a host cell. An expression vector may be part of a plasmid, viral genome, or nucleic acid fragment. Typically, an expression vector includes a polynucleotide to be transcribed, operably linked to a promoter. The term "promoter" is used herein to refer to an array of nucleic acid control sequences that direct transcription of a nucleic acid. As used herein, a promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of a polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements, which can be located as much as several thousand base pairs from the start site of transcription. Other elements that may be present in an expression vector include those that enhance transcription (e.g., enhancers) and terminate transcription (e.g., terminators).

[0056] The terms "reporter" and "selectable marker" can be used interchangeably and refer to a gene product that permits a cell expressing that gene product to be identified and / or isolated from a mixed population of cells. Such isolation might be achieved through the selective killing of cells not expressing the selectable marker, which may be, as a non-limiting example, an antibiotic resistance gene. Alternatively, the selectable marker may permit identification and / or subsequent isolation of cells expressing the marker as a result of the expression of a fluorescent protein such as GFP or the expression of a cell surface marker which permits isolation of cells by fluorescence- activated cell sorting (FACS), magnetic-activated cell sorting (MACS), or analogous methods. Suitable cell surface markers include CD8, CD19, and truncated CD19. Preferably, cell surface markers used for isolating desired cells are non-signaling molecules, such as subunit or truncated forms of CD8, CD19, or CD20. Suitable markers and techniques are known in the art. Also described herein is "traffic light reporting" (Kawalpreet K Aneja."Traffic Light Reporter for Genome Engineering". Acta Scientific Microbiology 3.9 (2020): 27-28, herein incorporated by reference in its entirety).

[0057] The terms "culture," "culturing," "grow," "growing," "maintain," "maintaining," "expand," "expanding," etc., when referring to cell culture itself or theprocess of culturing, can be used interchangeably to mean that a cell (e.g., human cell) is maintained outside its normal environment under controlled conditions, e.g., under conditions suitable for survival. Cultured cells are allowed to survive, and culturing can result in cell growth, stasis, differentiation or division. The term does not imply that all cells in the culture survive, grow, or divide, as some may naturally die or senesce. Cells are typically cultured in media, which can be changed during the course of the culture.

[0058] The terms "subject," "individual," and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.

[0059] As used herein, the term "administering" includes oral administration, topical contact, administration as a suppository, intravenous, intraperitoneal, intramuscular, intralesional, intrathecal, intranasal, or subcutaneous administration to a subject. Administration is by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, or transdermal). Parenteral administration includes, e.g., intravenous, intramuscular, intra-arteriole, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc.

[0060] The term "treating" refers to an approach for obtaining beneficial or desired results including, but not limited to, a therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant any therapeutically relevant improvement in or effect on one or more diseases, conditions, or symptoms under treatment. For prophylactic benefit, the compositions may be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more of the physiological symptoms of a disease, even though the disease, condition, or symptom may not have yet been manifested.

[0061] The term "effective amount" or "sufficient amount" refers to the amount of an agent that is sufficient to effect beneficial or desired results. The therapeutically effective amount may vary depending upon one or more of: the subject and disease condition being treated, the weight and age of the subject, the severity of the diseasecondition, the manner of administration and the like, which can readily be determined by one of ordinary skill in the art. The specific amount may vary depending on one or more of: the particular agent chosen, the host cell type, the location of the host cell in the subject, the dosing regimen to be followed, whether it is administered in combination with other compounds, timing of administration, and the physical delivery system in which it is carried.

[0062] The term "pharmaceutically acceptable carrier" refers to a substance that aids the administration of an active agent to a cell, an organism, or a subject."Pharmaceutically acceptable carrier" refers to a carrier or excipient that can be included in the compositions of the invention and that causes no significant adverse toxicological effect on the patient. Non-limiting examples of pharmaceutically acceptable carrier include water, NaCI, normal saline solutions, lactated Ringer's, normal sucrose, normal glucose, cell culture media, and the like. One of skill in the art will recognize that other pharmaceutical carriers are useful in the present invention.

[0063] The term "cellular localization tag" refers to an amino acid sequence, also known as a "protein localization signal," that targets a protein for localization to a specific cellular or subcellular region, compartment, or organelle (e.g., nuclear localization sequence, Golgi retention signal). Cellular localization tags are typically located at either the N-terminal or C-terminal end of a protein. For more information regarding cellular localization tags, see, e.g., Negi, et al. Database (Oxford). 2015: bav003 (2015); incorporated herein by reference in its entirety for all purposes.

[0064] "Percent similarity," in the context of polynucleotide or peptide sequences, is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the sequence (e.g., an msr locus sequence) in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence which does not comprise additions or deletions, for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleotide or amino acid occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of similarity (e.g., sequence similarity).

[0065] When a polynucleotide or peptide has at least about 80% similarity (e.g., sequence similarity), preferably at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93, 94%, 95%, 96%, 97%, 98%, 99%, or 100% similarity, to a reference sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection, such sequences are then said to be "substantially similar." With regard to polynucleotide sequences, this definition also refers to the complement of a test sequence.

[0066] For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence similarities for the test sequences relative to the reference sequence, based on the program parameters. For sequence comparison of nucleic acids and proteins, the BLAST and BLAST 2.0 algorithms and the default parameters discussed below are used.

[0067] Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FAST A, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, e.g., Current Protocols in Molecular Biology (Ausubel et al., eds. 1995 supplement)).

[0068] Additional examples of algorithms that are suitable for determining percent sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res. 25: 3389-3402, respectively. Software for performing BLAST analyses is publicly available at the National Center for Biotechnology Information website, ncbi.nlm.nih.gov. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying shortwords of length W in the query sequence, which either match or satisfy some positivevalued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=l, N=-2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see, e.g., Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)).

[0069] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul, Proc. Nat'l. Acad. Sci. USA, 90:5873- 5787 (1993)). One measure of similarity provided by the BLAST algorithm is the smallest sum probability ( P( N )), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001.

[0070] "Binding" refers to a sequence-specific, non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). Not all components of a binding interaction need be sequence-specific (e.g., contacts with phosphate residues in a DNA backbone), as long as the interaction as a whole is sequence-specific. Such interactions are generally characterized by a dissociation constant (Kd) of 10-6 M-l or lower. "Affinity" refers to the strength of binding: increased binding affinity being correlated with a lower Kd.

[0071] A "binding protein" is a protein that is able to bind non-covalently to another molecule. A binding protein can bind to, for example, a DNA molecule (a DNA-binding protein), an RNA molecule (an RNA-binding protein) and / or a protein molecule (a proteinbinding protein). In the case of a protein-binding protein, it can bind to itself (to formhomodimers, homotrimers, etc.) and / or it can bind to one or more molecules of a different protein or proteins. A binding protein can have more than one type of binding activity.

[0072] A "zinc-finger DNA binding protein" (or binding domain) is a protein, or a domain within a larger protein, that binds DNA in a sequence-specific manner through one or more zinc-fingers, which are regions of amino acid sequence within the binding domain whose structure is stabilized through coordination of a zinc ion. A zinc-finger DNA binding protein that is fused to a nuclease (i.e., the nuclease domain of Fokl) is often abbreviated as a zinc-finger nuclease or ZFN.

[0073] As used herein, the term "TALEN" or "TALE-nucleases" refers to an endonuclease comprising a DNA- binding domain, which in one embodiment comprises 14- 20 or 16-22 TAL domain repeats, which can be fused to any portion of the Fokl nuclease domain. An example of this technology can be found in WO2011072246, herein incorporated by reference in its entirety.GENERAL DESCRIPTION

[0074] Disclosed herein are retron editor systems, methods of making and using retron editor systems, as well as novel retrons which can be used with the retron editor systems described herein. As can be seen in Figs. 1 and 2, the retron editor systems disclosed herein comprise two main components: the nuclease component and the retron component. Together, these components can edit nucleic acids and insert a donor sequence. The retron component delivers the donor nucleic acid, and the nuclease component cleaves the target nucleic acid, allowing for insertion of a donor nucleic acid. Fig. IB shows the mechanism by which the retron produces a reverse transcribed single-stranded DNA (ssDNA), which can be inserted into a gene of interest by using a programmable nuclease.

[0075] A retron is a distinct nucleotide sequence found in the genome of many bacteria. These naturally occurring retrons can be engineered to produce single stranded DNA with a donor nucleic acid inserted therein. The retrons disclosed herein comprise the elements needed to produce an edited gene. This includes, but is not limited to, non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT). The ncRNA comprises an msd sequence and a msr sequence. The msr is the is the immediate precursor to the synthesis of msDNA. The retron msr RNA folds into a characteristic secondary structure that contains a conserved guanosine residue at the end of a stem loop. Synthesisof DNA by the retron-encoded reverse transcriptase (RT) results in a DNA / RNA chimera which is composed of small single-stranded DNA linked to small single-stranded RNA. This is referred to herein as ncDNA. The RNA strand is joined to the 5' end of the DNA chain via a 2'-5' phosphodiester linkage that occurs from the 2' position of the conserved internal guanosine residue. A description of retrons and how they function can be found, for example, in Simon et al. (2019), herein incorporated by reference in its entirety for its teaching concerning retrons.

[0076] The nuclease component of the retron editor system can comprise the components needed to edit a gene. The nuclease component is referred to herein as a "programmable nuclease," meaning that it can be programmed to target a nucleic acid of interest. This programmable nuclease can include, for example, a nucleotide sequence encoding a nuclease, as well as a guide RNA (gRNA) sequence. The programmable nuclease component can be any system known in the art which can allow restriction of the target nucleic acid to occur, as well as guidance of the nuclease to the proper location within the target nucleic acid. These programmable nuclease components include, but are not limited to, Cas9, Casl2a, Casl2f, TnpB, TALEN, and ZFN systems, as well as any combination thereof.Generally speaking, the programmable nuclease component of the retron editor system disclosed herein can comprise those components needed to carry out cleavage of the target. This can include, at a minimum, a guide RNA and a nuclease. Those specific components which are needed for specific gene editing systems are known in the art and are discussed in more detail below.

[0077] Programmable nuclease components and retrons are described in more detail below. Fig. IB shows a general schematic of how the nuclease component creates a double strand break in DNA, which can then be "repaired" using retron generated single stranded DNA with an inserted donor nucleic acid sequence. This inserted donor nucleic acid sequence can originate from the msd portion of an engineered retron and is described in further detail herein.

[0078] The retron editor systems disclosed herein can comprise linkers and nuclear localization sequences (NLS) as well. These are described in detail below.Retron Editor Cassettes

[0079] The components of a retron editor system can be encoded in one or more retron editor cassettes, which comprises nucleic acid encoding not only the retron nuclease components, but any other elements or components which can form a fully functional cassette or cassettes. In one embodiment, every element needed to form a retron editor system can be found in the same cassette. In other embodiments, various elements of the retron editor system can be in different cassettes. For example, each of the following can be in separate cassette: the retron, the nucleotide sequence encoding the gRNA, and the nucleotide sequence encoding the nuclease. The elements within the retron itself can also be in separate cassettes. For example, the reverse transcriptase (RT) can be found in separate cassette from the ncRNA. The RT can be coupled to the gRNA, for example. Examples of such cassettes can be found in Figs. 1 and 2.

[0080] When this cassette or cassettes are introduced into a cell, such as in the form of a vector, the cassette(s) are capable of producing a fully functional retron editor system, which can then edit nucleic acids found in the cell. The key benefit of this technology is that retrons can produce intracellular DNA at high concentration in different hosts, including mammalian cells.

[0081] In some embodiments, various components of the retron editor system can be integrated into the host cell genome and need not be present in a vector. For example, the RT component can be encoded in a sequence that has been integrated into the host cell genome.Retrons

[0082] As described above, retrons use reverse transcriptase to form a multicopy singlestranded DNA (msDNA), which is a molecule comprising a single-stranded DNA that is branched out from an internal nucleic acid of an RNA molecule (msdRNA) via a 2', 5'- phosphodiester linkage. A retron, as used herein, comprises the components necessary to form this msDNA: 1) a nucleotide sequence encoding reverse transcriptase (RT); and 2) a nucleotide sequence encoding non-coding RNA. This non-coding RNA comprises two segments, an msr sequence and an msd sequence. The msr sequence is recognized by the RT, and the msd sequence is reverse transcribed by the RT, and can comprise donor nucleic acids. Prior to reverse transcription, the msr and msd form a single highly structuredtranscript (Simon et al. 2019). Reverse transcription requires both an msr-msd sequence and its cognate RT.

[0083] As described above, the msd region of a retron transcript typically codes for the DNA component of msDNA, and the msr region is the RNA component of msDNA. In some retrons, the msr and msd loci have overlapping ends, and may be oriented opposite one another with a promoter located upstream of the msr locus which transcribes through the msr and msd loci (Fig. 1). As described above, the msd sequence can be modified to include a donor nucleic acid, which can be introduced into a target nucleic acid sequence via the nuclease component and gRNA.

[0084] The msd and msr regions of retron transcripts generally contain first and second inverted repeat sequences, which together make up a stable stem structure (Simon et al. 2019). The combined msr-msd region of the retron transcript serves not only as a template for reverse transcription but, by virtue of its secondary structure, also serves as a primer (i.e., self-priming) for msDNA synthesis by a reverse transcriptase. In some embodiments of retron component of the retron editor system of the present invention, the first inverted repeat sequence coding region is located within the 5' end of the msr locus. In other embodiments, the second inverted repeat sequence coding region is located 3' of the msd locus. In some embodiments of retron donor DNA-guide molecules of the present invention, the first inverted repeat sequence is located within the 5' end of the msr region. In other embodiments, the second inverted repeat sequence is located 3' of the msd region.

[0085] Disclosed herein are newly discovered retrons which can be used with retron editor systems as described herein. These can be seen in Table 7. SEQ ID NOS: 1-22 represent the protein sequence of the transcribed retrons, while SEQ ID NOS: 23-44 represent nucleic acids of the associated msr-msd sequences. While in its native form, specific proteins are associated with specific msr-msd sequences (such as SEQ ID NO: 1 and 23, SEQ ID NO: 2 and 24, etc.), it is noted that the protein and the msr-msd sequence can be interchangeable in the case of highly homologous retrons, so, for example, it is contemplated herein that the protein represented by SEQ ID NO: 1 can be used with any of the msr-msd sequences found in SEQ ID NOS: 23-44. Likewise, SEQ ID NO: 2 can be used with any of SEQ ID NOS: 23-44, and so on for all of SEQ ID NOS: 1-22, as each can be used with any of the msr-msd sequences disclosed herein.

[0086] Furthermore, contemplated herein are not only the exact sequences of SEQ ID NOS: 1-22, but proteins with 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% homology to any of SEQ ID NOS: 1-22. Likewise, also contemplated are nucleic acid msr-msd sequences with 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% homology to any of SEQ ID NOS: 23-44.

[0087] Any number of retrons may be used in alternative embodiments of the present invention. If desired, the nucleotide sequence of a native retron may be modified, for example using known codon optimization techniques, so that expression within the desired host is optimized. By codon optimization it is meant the selection of appropriate DNA nucleotides for the synthesis of oligonucleotide building blocks, and their subsequent enzymatic assembly, of a structural gene or fragment thereof in order to approach codon usage within the host. For more information regarding retrons, see, e.g., U.S. Pat. No. 8,932,860 and Lampson, et al. Cytogenet. Res. 110:491-499 (2005); both incorporated herein by reference in their entirety for all purposes.Nuclease Component

[0088] There are many types of programmable nucleases known to those of skill in the art. The retrons described herein can make use of any of those systems. Some of those systems are described in Gonzalez et al. Comparison of the Feasibility, Efficiency, and Safety of Genome Editing Technologies. Int J Mol Sci. 2021 Sep 26;22(19):10355, herein incorporated by reference in its entirety for its teaching concerning types of programmable nucleases. Some exemplary embodiments are given below, but again, it is emphasized that the present invention can be used with any number of programmable nucleases.Cas Endonucleases

[0089] Cas endonucleases, either as single effector proteins or in an effector complex with other components, unwind the DNA duplex at the target sequence and optionally cleave at least one DNA strand, as mediated by recognition of the target sequence by a polynucleotide (such as, but not limited to, a guide RNA (gRNA), which can comprise a CRISPR RNA (crRNA) or an sgRNA) that is in complex with the Cas effector protein. Such recognition and cutting of a target sequence by a Cas endonuclease typically occurs if the correct protospacer-adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. PAM sequences are described in more detail below. Alternatively, a Casendonuclease herein may lack DNA cleavage or nicking activity but can still specifically bind to a DNA target sequence when complexed with a suitable RNA component. (See also U.S. Patent Application US20150082478 published 19 Mar. 2015 and US20150059010 published 26 Feb. 2015). Cas endonucleases may occur as individual effectors (Class 2 CRISPR systems) or as part of larger effector complexes (Class I CRISPR systems).

[0090] Cas endonucleases that have been described include, but are not limited to, for example: Cas3 (a feature of Class 1 type I systems), Cas9 (a feature of Class 2 type II systems) and Casl2-family enzymes (e.g., Cpfl) (a feature of Class 2 type V systems). Cas3 (and its variants Cas3' and Cas3") functions as a single-stranded DNA nuclease (HD domain) and an ATP-dependent helicase. A variant of the Cas3 endonuclease can be obtained by disabling the functional activity of one or both domains of the Cas3 endonuclease poly peptide. Disabling the ATPase dependent helicase activity (by deletion, knockout of the Cas3-helicase domain, or through mutagenesis of critical residues or by assembling the reaction in the absence of ATP as described previously (Sinkunas, T. et al., 2013, EMBO J. 32:385-394) can convert the cleavage ready Cascade comprising the modified Cas3 endonuclease into a nickase (as the HD domain is still functional). Disabling the HD endonuclease activity can be accomplished by any method known in the art, such as but not limited to, mutagenesis of critical residues of the HD domain, can convert the cleavage ready Cascade comprising the modified Cas3 endonuclease into a helicase. Disabling the both the Cas helicase and Cas3 HD endonuclease activity can be accomplished by any method known in the art, such as but not limited to, mutagenesis of critical residues of both the helicase and HD domains, can convert the cleavage ready Cascade comprising the modified Cas3 endonuclease into a binder protein that binds to a target sequence.

[0091] Cas9 (formerly referred to as Cas5, Csnl, or Csxl2) is a Cas endonuclease that forms a complex with a crRNA and a tracrRNA, or with a single guide polynucleotide, for specifically recognizing and cleaving all or part of a DNA target sequence. Some Cas9 endonucleases recognize a 3' GC-rich PAM sequence on the target dsDNA, while other Cas9 endonucleases recognize other PAM sequences. A Cas9 protein comprises a RuvC nuclease with an HNH (H— N— H) nuclease adjacent to the RuvC-ll domain. The RuvC nuclease and HNH nuclease each can cleave a single DNA strand at a target sequence (the concerted action of both domains leads to DNA double-strand cleavage, whereas activity of one domain leads to a nick). In general, the RuvC domain comprises subdomains I, II and III,where domain I is located near the N-terminus of Cas9 and subdomains II and III are located in the middle of the protein, flanking the HNH domain (Hsu et al., 2013, Cell 157:1262- 1278). Cas9 endonucleases are typically derived from a type II CRISPR system, which includes a DNA cleavage system utilizing a Cas9 endonuclease in complex with at least one polynucleotide component. For example, a Cas9 can be in complex with a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). In another example, a Cas9 can be in complex with a single guide RNA (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15).

[0092] Casl2-family enzymes (formerly referred to as Cpfl, and variants c2cl, c2c3, CasX, and CasY) comprise an RuvC nuclease domain and produced staggered, 5' overhangs on the dsDNA target. Some variants do not require a tracrRNA, unlike the functionality of Cas9. Casl2 and its variants recognize a 5' AT-rich PAM sequence on the target dsDNA. An insert domain, called Nuc, of the Casl2a protein has been proposed to be responsible for target strand cleavage (Yamano et al., Cell 2016, 165:949-962). Additional mutation studies demonstrated the Nuc domain contributes to guide and target binding, with the RuvC domain responsible for cleavage of both DNA strands (Swarts et al., Mol Cell 2017, 66:221- 233 e224).

[0093] Cas endonucleases and effector proteins can be used for targeted genome editing (via simplex and multiplex double-strand breaks and nicks) and targeted genome regulation (via tethering of epigenetic effector domains to either the Cas protein or sgRNA. A Cas endonuclease can also be engineered to function as an RNA-guided recombinase, and via RNA tethers could serve as a scaffold for the assembly of multiprotein and nucleic acid complexes (Mali et al., 2013, Nature Methods Vol. 10:957-963).

[0094] Many Cas endonucleases have been described to date that can recognize specific PAM sequences (WO2016186953 published 24 Nov. 2016, WO2016186946 published 24 Nov. 2016, and Zetsche B et al. 2015. Cell 163, 1013) and cleave the target DNA at a specific position. It is understood that based on the methods and embodiments described herein utilizing a novel guided Cas system one skilled in the art can now tailor these methods such that they can utilize any guided endonuclease system.Zinc Finger Proteins

[0095] Zinc-finger binding domains (referred to herein as either ZFN or ZFP) can be engineered to bind to a sequence of choice. See, for example, Beerli et al. (2002) Nature Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nature Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632- 637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416. An engineered zinc-finger binding domain can have a novel binding specificity, compared to a naturally-occurring zinc- finger protein. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using databases comprising triplet (or quadruplet) nucleotide sequences and individual zinc-finger amino acid sequences, in which each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of zinc-fingers which bind the particular triplet or quadruplet sequence.

[0096] Exemplary selection methods, including phage display and two-hybrid systems, are disclosed in U.S. Pat. Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; and 6,242,568; as well as WO 98 / 37.186; WO 98 / 53057; WO 00 / 27878; WO 01 / 88197 and GB 2,338,237.

[0097] Selection of target sites; ZFNs and methods for design and construction of fusion proteins (and polynucleotides encoding same) are known to those of skill in the art and described in detail in U.S. Patent Application Publication Nos. 20050064474 and 20060188987, incorporated by reference in their entireties herein.

[0098] The ZFNs used with the systems disclosed herein can be nucleic acids when encode ZFNs, or can be the ZFN itself. The ZFN can also comprise a nuclease (cleavage domain, cleavage half-domain). The cleavage domain portion of the fusion proteins disclosed herein can be obtained from any endonuclease or exonuclease. Exemplary endonucleases from which a cleavage domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, 2002-2003 Catalogue, New England Biolabs, Beverly, Mass.; and Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Additional enzymes which cleave DNA are known (e.g., SI Nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease; see also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993). One or more of these enzymes (or functional fragments thereof) can be used as a source of cleavage domains and cleavage half-domains.

[0099] Similarly, a cleavage half-domain can be derived from any nuclease or portion thereof, as set forth above, that requires dimerization for cleavage activity. In general, two fusion proteins are required for cleavage if the fusion proteins comprise cleavage halfdomains. Alternatively, a single protein comprising two cleavage half-domains can be used. The two cleavage half-domains can be derived from the same endonuclease (or functional fragments thereof), or each cleavage half-domain can be derived from a different endonuclease (or functional fragments thereof). In addition, the target sites for the two fusion proteins are preferably disposed, with respect to each other, such that binding of the two fusion proteins to their respective target sites places the cleavage half-domains in a spatial orientation to each other that allows the cleavage half-domains to form a functional cleavage domain, e.g., by dimerizing. Thus, in certain embodiments, the near edges of the target sites are separated by 5-8 nucleotides or by 15-18 nucleotides. However any integral number of nucleotides or nucleotide pairs can intervene between two target sites (e.g., from 2 to 50 nucleotide pairs or more). In general, the site of cleavage lies between the target sites.TALENs

[0100] Transcription Activator-Like Effector Nucleases (TALENs) are artificial restriction enzymes generated by fusing the TAL effector DNA binding domain to a DNA cleavage domain. These reagents enable efficient, programmable, and specific DNA cleavage and represent powerful tools for genome editing in situ. Transcription activator-like effectors (TALEs) can be quickly engineered to bind practically any DNA sequence. The term TALEN, as used herein, is broad and includes a monomeric TALEN that can cleave double stranded DNA without assistance from another TALEN. The term TALEN is also used to refer to one or both members of a pair of TALENs that are engineered to work together to cleave DNA at the same site. TALENs that work together may be referred to as a left-TALEN and a right- TALEN, or a first and second TALEN, which references the handedness of DNA. See U.S. Ser. No. 12 / 965,590; U.S. Ser. No. 13 / 426,991 (U.S. Pat. No. 8,450,471); U.S. Ser. No. 13 / 427,040 (U.S. Pat. No. 8,440,431); U.S. Ser. No. 13 / 427,137 (U.S. Pat. No. 8,440,432); and U.S. Ser. No. 13 / 738,381, all of which are incorporated by reference herein in their entirety.

[0101] In some embodiments, the nuclease used with TALEN is selected from a group consisting of Pvull, MutH, Tevl, Fokl, Alwl, Mlyl, Sbfl, Sdal, Stsl, CleDORF, Clo051, andPept071. When Fokl is fused to a TALE domain each member of the TALEN pair binds to the DNA sites flanking a target site, the Fokl monomers dimerize and cause a DSB at the target site.

[0102] Besides the wild-type Fokl cleavage domain, variants of the Fokl cleavage domain with mutations have been designed to improve cleavage specificity and cleavage activity. The Fokl domain functions as a dimer, requiring two constructs with unique DNA binding domains for sites in the target genome with proper orientation and spacing. Both the number of amino acid residues between the TALEN DNA binding domain and the Fokl cleavage domain, and the number of bases between the two individual TALEN binding sites are parameters for achieving high levels of activity. Pvull, MutH, and Tevl cleavage domains are useful alternatives to Fokl and Fokl variants for use with TALEs. Pvull functions as a highly specific cleavage domain when coupled to a TALE (see Yank et al. 2013. PLoS One. 8: e82539). MutH is capable of introducing strand-specific nicks in DNA (see Gabsalilow et al. 2013. Nucleic Acids Research. 41: e83). Tevl introduces double-stranded breaks in DNA at targeted sites (see Beurdeley et al., 2013. Nature Communications. 4: 1762).

[0103] The relationship between amino acid sequence and DNA recognition of the TALE binding domain allows for designable proteins. Software programs such as DNA Works can be used to design TALE constructs. Other methods of designing TALE constructs are known to those of skill in the art. Doyle et al. (2012) TAL Effector-Nucleotide Targeter (TALE-NT) 2.0: tools for TAL effector design and target prediction. Nucleic Acids Res. 4O(W1):W117- W122; Cermak (2011). Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic Acids Res. 39(12):e82.

[0104] Engineered TALEN nucleases of the invention can be delivered into a cell in the form of a protein or, preferably, as a nucleic acid encoding the engineered nuclease. Such nucleic acid can be DNA (e.g., circular or linearized plasmid DNA or PCR products) or RNA or a combination of RNAs. Such RNA may have various stability, various lengths and be delivered in various amounts.

[0105] For embodiments in which the engineered TALEN nuclease coding sequence is delivered in DNA form, it can be operably linked to a promoter to facilitate transcription of the nuclease, (TALEN or meganuclease gene). Mammalian promoters suitable for the invention include constitutive promoters such as the cytomegalovirus early (CMV) promoter (Thomsen et al. (1984), Proc Natl Acad Sci USA. 81(3):659-63) or the SV40 early promoter(Benoist and Chambon (1981), Nature. 290(5804) :304-10) as well as inducible promoters such as the tetracycline-inducible promoter (Dingermann et al. (1992), Mol Cell Biol. 12(9):4038-45).Guide Polynucleotides

[0106] Many of the retron editor systems described herein require a guide polynucleotide. The guide polynucleotide enables target recognition, binding, and optionally cleavage by the nuclease, and can be a single molecule or a double molecule. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence). Optionally, the guide polynucleotide can comprise at least one nucleotide, phosphodiester bond or linkage modification such as, but not limited, to Locked Nucleic Acid (LNA), 5-methyl dC, 2,6-Diaminopurine, 2'-Fluoro A, 2'-Fluoro U, 2'-O-Methyl RNA, phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 (hexaethylene glycol chain) molecule, or 5' to 3' covalent linkage resulting in circularization. A guide polynucleotide that solely comprises ribonucleic acids is also referred to as a "guide RNA" or "gRNA" (US20150082478 published 19 Mar. 2015 and US20150059010 published 26 Feb. 2015). A guide polynucleotide may be engineered or synthetic.

[0107] The guide polynucleotide includes a chimeric non-naturally occurring guide RNA comprising regions that are not found together in nature (i.e., they are heterologous with each other). For example, a chimeric non-naturally occurring guide RNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA, linked to a second nucleotide sequence that can recognize the Cas endonuclease, such that the first and second nucleotide sequence are not found linked together in nature.

[0108] The guide polynucleotide can be a double molecule (also referred to as duplex guide polynucleotide) comprising a crNucleotide sequence (such as a crRNA) and a tracrNucleotide (such as a tracrRNA) sequence. In some cases, there is a linker polynucleotide that connects the crRNA and tracrRNA to form a single guide, for example an sgRNA.

[0109] The crNucleotide includes a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in atarget DNA and a second nucleotide sequence (also referred to as a tracr mate sequence) that is part of a Cas endonuclease recognition (CER) domain. The tracr mate sequence can hybridized to a tracrNucleotide along a region of complementarity and together form the Cas endonuclease recognition domain or CER domain. The CER domain is capable of interacting with a Cas endonuclease polypeptide. The crNucleotide and the tracrNucleotide of the duplex guide polynucleotide can be RNA, DNA, and / or RNA-DNA-combination sequences. In some embodiments, the crNucleotide molecule of the duplex guide polynucleotide is referred to as "crDNA" (when composed of a contiguous stretch of DNA nucleotides) or "crRNA" (when composed of a contiguous stretch of RNA nucleotides), or "crDNA-RNA" (when composed of a combination of DNA and RNA nucleotides). The crNucleotide can comprise a fragment of the crRNA naturally occurring in Bacteria and Archaea. The size of the fragment of the crRNA naturally occurring in Bacteria and Archaea that can be present in a crNucleotide disclosed herein can range from, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments the tracr nucleotide is referred to as "tracrRNA" (when composed of a contiguous stretch of RNA nucleotides) or "tracrDNA" (when composed of a contiguous stretch of DNA nucleotides) or "tracrDNA-RNA" (when composed of a combination of DNA and RNA nucleotides. In one embodiment, the RNA that guides the RNA / Cas9 endonuclease complex is a duplexed RNA comprising a duplex crRNA-tracrRNA. The tracrRNA (transactivating CRISPR RNA) comprises, in the 5'-to-3' direction, (i) a sequence that anneals with the repeat region of CRISPR type II crRNA and (ii) a stem loop-comprising portion (Deltcheva et al., Nature 471:602-607). The duplex guide polynucleotide can form a complex with a Cas endonuclease, wherein said guide polynucleotide / Cas endonuclease complex (also referred to as a guide polynucleotide / Cas endonuclease system) can direct the Cas endonuclease to a genomic target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) into the target site.(US20150082478 published 19 Mar. 2015 and US20150059010 published 26 Feb. 2015).

[0110] The guide RNA includes a dual molecule comprising a chimeric non-naturally occurring crRNA linked to at least one tracrRNA. A chimeric non-naturally occurring crRNA includes a crRNA that comprises regions that are not found together in nature (i.e., they are heterologous with each other. For example, a crRNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to anucleotide sequence in a target DNA, linked to a second nucleotide sequence (also referred to as a tracr mate sequence) such that the first and second sequence are not found linked together in nature.

[0111] The guide polynucleotide can also be a single molecule (also referred to as single guide polynucleotide) comprising a crNucleotide sequence linked to a tracr nucleotide sequence. The single guide polynucleotide comprises a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA and a Cas endonuclease recognition domain (CER domain), that interacts with a Cas endonuclease polypeptide.

[0112] The VT domain and / or the CER domain of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA-combination sequence. The single guide polynucleotide being comprised of sequences from the crNucleotide and the tracrNucleotide may be referred to as "single guide RNA" (when composed of a contiguous stretch of RNA nucleotides) or "single guide DNA" (when composed of a contiguous stretch of DNA nucleotides) or "single guide RNA-DNA" (when composed of a combination of RNA and DNA nucleotides). The single guide polynucleotide can form a complex with a Cas endonuclease, wherein said guide polynucleotide / Cas endonuclease complex (also referred to as a guide polynucleotide / Cas endonuclease system) can direct the Cas endonuclease to a genomic target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) the target site. (US20150082478 published 19 Mar. 2015 and US20150059010 published 26 Feb. 2015).

[0113] A chimeric non-naturally occurring single guide RNA (sgRNA) includes a sgRNA that comprises regions that are not found together in nature (i.e., they are heterologous with each other. For example, a sgRNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA linked to a second nucleotide sequence (also referred to as a tracr mate sequence) that are not found linked together in nature.

[0114] The nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA combination sequence. In one embodiment, the nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide (also referred to as "loop") can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23,24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48,49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73,74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99 or 100 nucleotides in length. In another embodiment, the nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can comprise a tetraloop sequence, such as, but not limiting to a GAAA tetraloop sequence.

[0115] The guide polynucleotide can be produced by any method known in the art, including chemically synthesizing guide polynucleotides (such as but not limiting to Hendel et al. 2015, Nature Biotechnology 33, 985-989), in vitro generated guide polynucleotides, and / or self-splicing guide RNAs (such as but not limited as such).

[0116] In some embodiments, the degree of complementarity between a guide sequence of the gRNA (i.e., crRNA sequence) and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a crRNA sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some instances, a crRNA sequence is about 20 nucleotides in length. In other instances, a crRNA sequence is about 15 nucleotides in length. In other instances, a crRNA sequence is about 25 nucleotides in length.

[0117] The nucleotide sequence of a modified gRNA can be selected using any of the webbased software described above. Considerations for selecting a DNA-targeting RNA include the PAM sequence for the nuclease (e.g., Cas9 or Cpfl) to be used, and strategies for minimizing off-target modifications. Tools, such as the CRISPR Design Tool, can provide sequences for preparing the gRNA, for assessing target modification efficiency, and / or assessing cleavage at off-target sites.

[0118] In some embodiments, the length of the gRNA molecule is about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140,145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, or more nucleotides in length. In some instances, the length of the gRNA is about 100 nucleotides in length. In other instances, the gRNA is about 90 nucleotides in length. In other instances, the gRNA is about 110 nucleotides in length.

[0119] Nucleotide sequence modification of the guide polynucleotide can be selected from, but not limited to, the group consisting of a 5' cap, a 3' polyadenylated tail, a riboswitch sequence, a stability control sequence, a sequence that forms a dsRNA duplex, a modification or sequence that targets the guide poly nucleotide to a subcellular location, a modification or sequence that provides for tracking, a modification or sequence that provides a binding site for proteins, a Locked Nucleic Acid (LNA), a 5-methyl dC nucleotide, a 2,6-Diaminopurine nucleotide, a 2'-Fluoro A nucleotide, a 2'-Fluoro U nucleotide; a 2'-O- Methyl RNA nucleotide, a phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 molecule, a 5' to 3' covalent linkage, or any combination thereof. These modifications can result in at least one additional beneficial feature, wherein the additional beneficial feature is selected from the group of a modified or regulated stability, a subcellular targeting, tracking, a fluorescent label, a binding site for a protein or protein complex, modified binding affinity to complementary target sequence, modified resistance to cellular degradation, and increased cellular permeability.

[0120] The gRNA described herein can be provided within a cassette along with a retron and a nuclease component, or can be provided alone with the retron or alone with the nuclease in separate cassettes. It can also be provided by being encoded directly into a host cell genome. In one particular embodiment, the gRNA can be fused to ncRNA of the retron, as described above.Protospacer Adjacent Motif (PAM)

[0121] A "protospacer adjacent motif" (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that can be recognized (targeted) by a guide polynucleotide / Cas endonuclease system. The Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM herein can differ depending on the Casprotein or Cas protein complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.

[0122] A "randomized PAM" and "randomized protospacer adjacent motif" are used interchangeably herein, and refer to a random DNA sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas endonuclease system. The randomized PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long. A randomized nucleotide includes anyone of the nucleotides A, C, G or T.Nuclear Localization Sequence

[0123] The retron editor systems disclosed herein can include a nuclear localization sequence (NLS). A heterologous NLS amino acid sequence herein may be of sufficient strength to drive accumulation of the nuclease, the RT, or a combination thereof in a detectable amount in the nucleus of a mammalian cell herein, for example. An NLS may comprise one (monopartite) or more (e.g., bipartite) short sequences (e.g., 2 to 30 residues) of basic, positively charged residues (e.g., lysine and / or arginine), and can be located anywhere in the nuclease or retron RT amino acid sequence but such that it is exposed on the protein surface. An NLS may be operably linked to the N-terminus or C-terminus of a Cas protein, the retron RT, or their combination, for example. Two or more NLS sequences can be linked to a Cas protein and the retron RT, for example, such as on both the N- and C- termini of the retron RT. The Cas endonuclease gene and the RT can be operably linked to a SV40 nuclear targeting signal upstream of the coding region and a bipartite VirD2 nuclear localization signal (Tinland et al. (1992) Proc. Natl. Acad. Sci. USA 89:7442-6) downstream of the coding region. Non-limiting examples of suitable NLS sequences herein include those disclosed in U.S. Pat. Nos. 6,660,830 and 7,309,576.Linkers

[0124] In some embodiments, a linker can be used with the retron editor systems described herein. Fig. ID shows an example of where the linker can be placed in the cassette. The linker can be placed such that it connects the C-terminus of the nuclease with the N- terminus of the retron RT, the N-terminus of the nuclease with the C-terminus of the RT, or linkers can connect two internal regions. The linker can be a flexible linker, a rigid linker, an in vivo cleavable linker, a polyvalent linker, or any combination thereof. The linker can forma one-to-one connection between the RT and nuclease, or the linker can connect one nuclease to two or more RTs, or two or more nucleases to a single RT (polyvalent linkers). In some cases, a linker may link functional domains together (as in flexible and rigid linkers) or releasing free functional domain in vivo as in in vivo cleavable linkers.

[0125] Linkers may improve biological activity, increase expression yield, and achieving desirable pharmacokinetic profiles. A linker can also comprise hydrazone, peptide, disulfide, or thioesther.

[0126] In some cases, a linker sequence described herein can include a flexible linker. Flexible linkers can be applied when a joined domain requires a certain degree of movement or interaction. Flexible linkers can be composed of small, non-polar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. A flexible linker can have sequences consisting primarily of stretches of Gly and Ser residues ("GS" linker). An example of a flexible linker can have the sequence of (Gly-Gly-Ser)n. By adjusting the copy number "n", the length of this exemplary GS linker can be optimized to achieve appropriate separation of functional domains, or to maintain necessary inter-domain interactions. Besides GS linkers, other flexible linkers can be utilized for recombinant fusion proteins. In some cases, flexible linkers can also be rich in small or polar amino acids such as Gly and Ser, but can contain additional amino acids such as Thr and Ala to maintain flexibility. In other cases, polar amino acids such as Lys and Glu can be used to improve solubility.

[0127] Flexible linkers included in linker sequences described herein, can be rich in small or polar amino acids such as Gly and Ser to provide good flexibility and solubility. Flexible linkers can be suitable choices when certain movements or interactions are desired for fusion protein domains. In addition, although flexible linkers may not have rigid structures, they can serve as a passive linker to keep a distance between functional domains. The length of flexible linkers can be adjusted to allow for proper folding or to achieve optimal biological activity of the fusion proteins.

[0128] A linker described herein can further include a rigid linker in some cases. A rigid linker may be utilized to maintain a fixed distance between domains of a polypeptide. Examples of rigid linkers can be: Alpha helix-forming linkers, Pro-rich sequence, (XP)n, X-Pro backbone, KL( A / E) n (n=l-6), to name a few. Rigid linkers can exhibit relatively stiff structures by adopting a-helical structures or by containing multiple Pro residues in some cases.

[0129] A linker described herein can include a polyvalent linker that connects one nuclease to two or more RTs or one RT to two or more nucleases. A polyvalent linker may be utilized to increase the copy number of particular enzyme relative to another enzyme (i.e., to bring multiple RTs to the site of a DNA break induced by Cas9). Examples of polyvalent linkers can be: SunTag and SpyTag, to name a few. Polyvalent linkers can link two or more entities to a single polypeptide. Examples of such linkers can be found, for example, in W02016 / 011070A2 and EP3303374B1, both of which are incorporated by reference herein.

[0130] A linker described herein can be cleavable in some cases. In other cases a linker is not cleavable. Linkers that are not cleavable may covalently join functional domains together to act as one molecule throughout in vivo processes or ex vivo processes. A linker can also be cleavable in vivo. A cleavable linker can be introduced to release free functional domains in vivo. A cleavable linker can be cleaved by the presence of reducing reagents, proteases, to name a few. For example, a reduction of a disulfide bond may be utilized to produce a cleavable linker. In the case of a disulfide linker, a cleavage event through disulfide exchange with a thiol, such as glutathione, could produce a cleavage. A cleavable linker can also comprise hydrazone, peptides, disulfide, or thioesther. For example, a hydrazone can confer serum stability. In other cases, a hydrazone can allow for cleavage in an acidic compartment. An acidic compartment can have a pH up to 7. A linker can also include a thioether. A thioether can be nonreducible A thioether can be designed for intracellular proteolytic degradation.

[0131] In some cases, a linker can be engineered. Methods of designing linkers can be computational. In some cases, computational methods can include graphic techniques. Computation methods can be used to search for suitable peptides from libraries of three- dimensional peptide structures derived from databases. For example, a Brookhaven Protein Data Bank (PDB) can be used to span the distance in space between selected amino acids of a linker.Recombinant Constructs for Transformation of Cells

[0132] The retron editor systems disclosed herein can be introduced into a cell via a variety of methods. For example, methods of delivering RNA to a cell can include providing the RNA in a therapeutic payload, such as in a synthetic vehicle like lipid or lipid-based nanoparticles. Examples of such systems can be found in Paunovska et al. (Drug delivery systems for RNAtherapeutics. Nat Rev Genet 23, 265-280, 2022, herein incorporated by reference in its entirety its teaching concerning RNA delivery). Delivery can also be accomplished by way of a vector. In certain embodiments, the delivery system comprises lipid particles as described in Kanasty R, Delivery materials for siRNA therapeutics Nat Mater. 12(ll):967-77 (2013), which is hereby incorporated by reference. In some embodiments, the lipid-based vector is a lipid nanoparticle, which is a lipid particle between about 1 and about 100 nanometers in size.

[0133] Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells as well as plants and seeds produced by the methods described herein.

[0134] Methods for introducing the retron editor system into cells or organisms also include, but are not limited to, microinjection, electroporation, stable transformation methods, transient transformation methods, ballistic particle acceleration (particle bombardment), whiskers mediated transformation, direct gene transfer, viral-mediated introduction, transfection, transduction, cell-penetrating peptides, mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, topical applications, or any combination thereof.

[0135] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described more fully in Sambrook et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory: Cold Spring Harbor, N.Y. (1989). Transformation methods are well known to those skilled in the art and are described infra.

[0136] Vectors and constructs include circular plasmids, and linear polynucleotides, comprising a polynucleotide of interest and optionally other components including linkers, adapters, regulatory or analysis. In some examples a recognition site and / or target site can be comprised within an intron, coding sequence, 5' UTRs, 3' UTRs, and / or regulatory regions.

[0137] The invention further provides expression constructs for expressing in a prokaryotic or eukaryotic cell / organism a gene editing system that is capable of recognizing, binding to, and optionally nicking, unwinding, or cleaving all or part of a target sequence.

[0138] The delivery vehicles (whether viral vector or non-viral vector or RNA conjugate material) may be administered by any method known in the art, including injection, optionally by direct injection to target tissues. Nucleic acid modification can be monitoredover time by, for example, periodic biopsy with PCR amplification and / or sequencing of the target region from genomic DNA, or by RT-PCR and / or sequencing of the expressed transcripts. Alternatively, nucleic acid modification can be monitored by detection of a reporter gene or reporter sequence. Alternatively, nucleic acid modification can be monitored by expression or activity of a corrected gene product or a therapeutic effect in the subject.

[0139] In one embodiment, the expression constructs of the disclosure comprise a promoter operably linked to a nucleotide sequence encoding a Cas gene and a promoter operably linked to a guide RNA of the present disclosure. The promoter is capable of driving expression of an operably linked nucleotide sequence in a prokaryotic or eukaryotic cell / organism. The cassette can further comprise at least one promoter. The promoter can be capable of driving sgRNA transcription. An example of such as promoter is a U6 promoter. In another embodiment, the promoter can be capable of driving msr-msd transcription. An example of such as promoter is the Hl promoter. The promoter can drive nuclease and RT expression, or the expression of their fusion. An example of such a promoter is a CMV protomer. The promoter can be operably linked to non-coding RNA, sgRNA, msr, msd, (or both msr and msd), the nuclease, the RT, or to a nuclease-RT fusion nucleic acid. The promoter can be an RNA polymerase II promoter or an RNA polymerase III promoter.Donor Nucleic Acid

[0140] Various methods and compositions can be employed to obtain a cell or organism having donor nucleic acid inserted in a target site by a retron editor system. Such methods can employ homologous recombination (HR) to provide integration of the polynucleotide of interest at the target site.

[0141] The donor nucleic acid can further comprises a first and a second region of homology that flank the polynucleotide of interest. The first and second regions of homology of the donor nucleic acid can share homology to a first and a second genomic region, respectively, present in or flanking the target site of the cell or organism genome.

[0142] The donor nucleic acid sequence can be tethered to the guide polynucleotide. For example, the ncDNA can be fused to the gRNA. Tethered donor DNAs can allow for colocalizing target and donor DNA, useful in genome editing, gene insertion, and targetedgenome regulation, and can also be useful in targeting post-mitotic cells where function of endogenous HR machinery is expected to be highly diminished (Mali et al., 2013, Nature Methods Vol. 10:957-963).

[0143] The amount of homology or sequence identity shared by a target nucleic acid and a donor nucleic acid can vary and includes total lengths and / or regions having unit integral values in the ranges of about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5-3 kb, 2-4 kb, 2.5-5 kb, 3- 6 kb, 3.5-7 kb, 4-8 kb, 5-10 kb, or up to and including the total length of the target site. These ranges include every integer within the range, for example, the range of 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 and 20 bps. The amount of homology can also be described by percent sequence identity over the full aligned length of the two polynucleotides which includes percent sequence identity at least of about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, between 98% and 99%, 99%, between 99% and 100%, or 100%. Sufficient homology includes any combination of polynucleotide length, global percent sequence identity, and optionally conserved regions of contiguous nucleotides or local percent sequence identity, for example sufficient homology can be described as a region of 75-150 bp having at least 80% sequence identity to a region of the target locus. Sufficient homology can also be described by the predicted ability of two polynucleotides to specifically hybridize under high stringency conditions, see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and, Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology— Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0144] The structural similarity between a given genomic region and the corresponding region of homology found on the donor nucleic acid can be any degree of sequence identity that allows for homologous recombination to occur. For example, the amount of homology or sequence identity shared by the "region of homology" of donor nucleic acid and the "genomic region" of the organism genome can be at least 50%, 55%, 60%, 65%, 70%, 75%,80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%,97%, 98%, 99% or 100% sequence identity, such that the sequences undergo homologous recombination.

[0145] The region of homology on the donor nucleic acid can have homology to any sequence flanking the target site. While in some instances the regions of homology share significant sequence homology to the genomic sequence immediately flanking the target site, it is recognized that the regions of homology can be designed to have sufficient homology to regions that may be further 5' or 3' to the target site. The regions of homology can also have homology with a fragment of the target site along with downstream genomic regions. In one embodiment, the first region of homology further comprises a first fragment of the target site and the second region of homology comprises a second fragment of the target site, wherein the first and second fragments are dissimilar.Gene Editing Methods

[0146] The retron editor systems described herein can be used for gene editing. In general, gene editing can be performed by cleaving one or both strands at a specific polynucleotide sequence in a cell with a nuclease associated with a suitable donor nucleic acid sequence. Once a single or double-strand break is induced in the DNA, the cell's DNA repair mechanism is activated to repair the break via nonhomologous end-joining (NHEJ) or Homology-Directed Repair (HDR) processes which can lead to modifications at the target site. This is illustrated in Fig. 1C.

[0147] The length of the DNA sequence at the target site can vary, and includes, for example, target sites that are at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more than 30 nucleotides in length. It is further possible that the target site can be palindromic, that is, the sequence on one strand reads the same in the opposite direction on the complementary strand. The nick / cleavage site can be within the target sequence or the nick / cleavage site could be outside of the target sequence. In another variation, the cleavage could occur at nucleotide positions immediately opposite each other to produce a blunt end cut or, in other cases, the incisions could be staggered to produce single-stranded overhangs, also called "sticky ends", which can be either 5' overhangs, or 3' overhangs. Active variants of genomic target sites can also be used. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%,95%, 96%, 97%, 98%, 99% or more sequence identity to the given target site, wherein theactive variants retain biological activity and hence are capable of being recognized and cleaved by a gene editing system.

[0148] Assays to measure the single or double-strand break of a target site by an endonuclease are known in the art and generally measure the overall activity and specificity of the agent on DNA substrates comprising recognition sites.

[0149] A targeting method herein can be performed in such a way that two or more DNA target sites are targeted in the method, for example. Such a method can optionally be characterized as a multiplex method. Two, three, four, five, six, seven, eight, nine, ten, or more target sites can be targeted at the same time in certain embodiments. A multiplex method is typically performed by a targeting method herein in which multiple different RNA components are provided, each designed to guide a guide polynucleotide / Cas endonuclease complex to a unique DNA target site.

[0150] Uses for gene editing have been described in the art and are contemplated herein (see for example: US20150082478 Al published 19 Mar. 2015, WO2015026886 published 26 Feb. 2015, and US20150059010 published 26 Feb. 2015) and include but are not limited to modifying or replacing nucleotide sequences of interest (such as a regulatory elements), insertion of polynucleotides of interest, gene knock-out, gene-knock in, modification of splicing sites and / or introducing alternate splicing sites, modifications of nucleotide sequences encoding a protein of interest, amino acid and / or protein fusions, and gene silencing by expressing an inverted repeat into a gene of interest.

[0151] Direct delivery of the retron editors described herein can be accompanied by direct delivery (co-delivery) of other mRNAs that can promote the enrichment and / or visualization of cells receiving the retron editor system. For example, direct co-delivery of the system, together with mRNA encoding phenotypic markers, can enable the selection and enrichment of cells without the use of an exogenous selectable marker by restoring function to a non-functional gene product as described in W02017070032 published 27 Apr. 2017.Methods of Screening

[0152] Disclosed herein are methods of screening for functional retron editors, the method comprising: providing a potential retron editor system; transforming a cell with thepotential retron editor, wherein said cell has been modified to express a signal upon transformation with a retron editor; and detecting the presence of the signal.

[0153] Methods of detecting a signal are known to those of skill in the art. For example, fluorescence can be detected. An example of detecting fluorescence includes the use of "traffic light" reporters. Other examples of reporters can be found in the literature, such as in Stepanenko OV, Verkhusha VV, Kuznetsova IM, Uversky VN, Turoverov KK. Fluorescent proteins as biomarkers and biosensors: throwing color lights on molecular and cellular processes. Curr Protein Pept Sci. 2008 Aug;9(4):338-69, which is herein incorporated by reference in its entirety for its teaching of reporter systems. Detection can also be via sequencing, such as by Sanger or next generation (NGS) DNA sequencing.EXAMPLESEXAMPLE 1: Discovery and Engineering of Retrons for Precise Genome Editing

[0154] Introduction

[0155] Disclosed herein is the development of a highly efficient retron gene editor is reported. More than 500 high-confidence retrons from metagenomic sources were bioinformatically identified. Using a functional reporter system, 98 variants in mammalian cells were screened and 17 RTs were identified that were more active than the previously- established Ecol-RT. Further rational design achieved editing efficiencies that were comparable to conventional single-stranded oligodeoxynucleotide (ssODN) donors. Steering DNA repair outcomes towards HDR via small molecule inhibitors and Cas9-DNA repair protein fusions boosted targeted DNA insertion. Retron editors also function with Casl2a, significantly broadening their genomic target range. The nickase Cas9(D10A) also supports retron editing, and this activity can be improved with DNA repair protein fusions. Retron editors were applied for installing in-frame epitopes for live cell imaging in U2OS cells. Finally, all RNA-based retron editing was demonstrated in cell lines and vertebrates.Broadly, this work establishes retron editors as a powerful tool for templated cargo insertion, opening new gene editing modalities.

[0156] Results

[0157] A metagenomic survey identifies retron-RTs active in mammalian cells. The present study explored whether a metagenomic survey of retron-RTs will uncover variants that improve homology-directed repair in heterologous hosts. Towards this goal, a pipelineto phenotypically assess retron RTs using fluorescent proteins in HEK293T cells was explored (see Fig. ID). The reporter expressed RFP and GFP that were separated by a ribosomal skipping T2A sequence

[0047] , RFP had a 9 basepair (bp) deletion (A9) adjacent to a Y64L mutation

[0048] , These mutations ensured that RFP(A9) was dark until the wild type (WT) sequence was restored via templated HDR following a Cas9-generated double stranded break (DSB). GFP served as a transfection control, and also reported on Cas9-generated insertions and / or deletions (indels) that shifted the open reading frame out of frame. As expected, transfecting Cas9 and a single-stranded oligodeoxynucleotide (ssODN) donor restored RFP fluorescence (see Fig. IE, middle). The well-characterized Ecol-RT also repaired RFP, although HDR activity was substantially lower than the ssODN (see Fig. IE, left). The variable msd region included 29 nt of homology flanking a 9 nt insertion that paired with the target strand and also reverted the Y64L mutation. Having established this assay, additional RTs were tested from diverse microbes.

[0158] Retron-RTs are ubiquitous throughout bacteria, but only a handful have been tested experimentally

[0033] , Therefore, a bioinformatics pipeline was developed to identify new retron-RTs from metagenomic sources (Fig. 2A). First, all RTs in the NCBI database of non- redundant bacterial and archaeal genomes were annotated, as well as 2M partially assembled bacterial genomes from the human microbiome [49, 50], It was considered that human microbiome-derived RTs will also be active at physiological conditions. After identifying likely retron-RTs, the msr-msd non-coding RNA (ncRNA) and accessory proteins were queried. This search identified >500 high confidence, non-redundant retrons with well-annotated msr-msds (see Fig. 2B). The identified systems were classified into a phylogenetic tree and sorted into 11 clades following a prior bioinformatic survey (Fig. IF)

[0051] , The highest-confidence systems across multiple clades were prioritized for experimental characterization in mammalian cells.

[0159] Next, 98 retron RTs were screened using fluorescence cytometry and the RFP(A9) reporter (see Fig. 1G). As expected, transfecting the Cas9-sgRNA plasmid with an ssODN partially restored the RFP signal (see Fig. 1G, inset). Thirty-one RTs (32% of all tested systems) restored RFP fluorescence (defined as >1% RFP+ cells). Ecol-RT restored RFP+ signal in 5% of the cells. Ten RTs had greater than 2-fold higher repair activity than Ecol-RT (see Fig. 1G). Mval-RT, derived from Myxococcus vastator, had an editing efficiency that was 6-fold higher than Ecol-RT in this transient repair assay (see Fig. 1G, inset). The best-performing RTs all belonged to clade 9, indicating that these enzymes were especially active in mammalian cells, and / or that the bioinformatics workflow was most accurate in predicting the msr-msd sequences from this clade. The top systems also showed mostly RFP+ cells via confocal microscopy (see Figs. IE, 3A). To test these RTs in a genomic reporter, the RFP cassette was integrated into the AAVSl locus of HEK293T cells (see Fig. 3B). While Ecol-RT was minimally able to restore RFP signal, the top six RT were all active in the genomically integrated assay. Escherichia fergusonii (Efel) RT was the most active in the genomic reporter, restoring RFP+ signal approximately 10-fold better than Ecol-RT.

[0160] Retron RTs co-evolve with a cognate msr-msd, but their ability to reverse transcribe from the msr-msd of other retrons is unknown. Therefore, the feasibility of using two or more orthogonal RTs for multiplexed retron editing was explored (Fig. 1H). The RFP repair activity of six active novel RTs was tested, along with Ecol-RT, when co-expressed with the msr-msd from other systems. Mval-RT, the most broadly cross-reactive RT, shared only 33- 36% amino acid sequence identity and 49-59% msr-msd nucleotide identity with all other systems (see Fig. 4A). In contrast, Efel-RT shared 44-46% amino acid sequence identity but remained exclusive to its cognate msr-msd. The remaining RTs used a broad range of msr- msds, including those from Ecol-RT. Vibrio rotiferianus (Vrol)- and Vibrio aphrogenes (Vapl)-RTs showed comparable activity with Proteus sp. (Pspl) msr-msd and their native msr-msds. Notably, all RTs shared a similar msr-msd secondary structure, including a palindromic repeat and an extended msd hairpin (see Fig. 4B). It was concluded that retron- RTs are more flexible in msr-msd utilization than previously appreciated, and that caution should be taken when combining multiple systems in a single cell line. Further, Efel-RT was the most active enzyme in the genomic reporter assay and struck an excellent balance between high activity and specificity for gene editing applications.

[0161] Next, the top-performing RTs were tested for their ability to integrate a 10 nt cargo into native genomic loci (see Fig. 5). Two rounds of PCR were used to generate libraries for next generation DNA sequencing (NGS, Fig. 5A). The first PCR reaction primed outside the homology arms to avoid amplifying the reverse transcribed ssDNA. The second PCR reaction barcoded the amplicons for short-read NGS. Cas9 and an ssODN with the same homology arms were used as a positive control, and to benchmark the RT fidelity. Retron editing efficiency was distinguished from Cas9-generated indels by scoring the percentage of modified reads that had the intended insert relative to all modified reads. Editingefficiencies ranged from 8-30%. Efel-RT showed the highest editing activity at EMX1 and CFTR, which was consistent with the genomic RFP(A9) assay (Fig. 5B). Thus Efel-RT was selected for all subsequent experiments.

[0162] Deep sequencing of the insert at the EMX1 locus showed that more than 99% of Efel-RT driven insertion events contained the intended 10 nt cargo (see Fig. 5C,D). The most frequent imperfect insertions included indels and single nucleotide substitutions that were the signature of alternative end joining pathways [52, 53], To determine whether RT fidelity was contributing to this error, a similar analysis was conducted for the Cas9+ssODN experiment at EMX1 (see Fig. 5D). Efel-RT error rates were approximately 10“3to 10“4(errors / nt), consistent with the substitution rates measured for high fidelity RTs (see Fig. 5D) in vitro [54, 55], In contrast, ssODN substitution rates were 10-fold greater at the same locus. These results indicated that RT-based editing fidelity exceeds that of conventional ssODNs, and that repair fidelity may be limited by cell-intrinsic repair pathways, as described in more detail below.

[0163] Next, retron editing was tested across five native loci, and as a function of the homology arm length. In all cases, inserting 10 nt was most efficient with 50 nt homology arms (7-28% insertion rates across five loci) (see Fig. 5E). Editing efficiency was strongly correlated with Cas9 cleavage activity at each locus. Insertion efficiency generally matched, and in the case of EMX1 and HBB, exceeded the edit rates with Cas9 + ssODN. It was concluded that short homology arms supported the highest insertion rates in the system.

[0164] Rational Engineering of Retron Editors Increases Insertion Activity. To further improve insertion activity, ncRNA expression, nuclear localization signals (NLSs), and the nuclease-RT linker were optimized in an Efel-RT-based retron editor (Fig. 6). The integrated RFP(A9) reporter was used for rapid iterative screens. Splitting the sgRNA and msr-msd into two transcripts increased editing efficiency (see Fig. 6A,B). In the "split" design, sgRNA expression was driven by a U6 promoter, and the msr-msd was transcribed via the Hl promoter (Fig. 6A). These results suggest that fusing the sgRNA and the msr-msd may impact Cas9 and / or RT activity, possibly by misfolding the structural elements of each ncRNA, or by imposing additional steric constraints. All subsequent experiments used the "split" design.

[0165] Nuclear import can be a rate limiting factor in mammalian gene editing [56-61], Therefore, 25 N- and C-terminal NLS combinations were tested that were previouslyreported to improve Cas9-based gene editing (see Fig. 6C & Table 5). A combination of island C-terminal bipartite SV40 (BP-SV40) NLSs showed the highest RFP repair activity. Adding cMyc-SV40 to the N-terminus and two or more SV40s to the C-terminus increased Cas9 cleavage 1.6-fold relative to an N-terminal SV40 and C-terminal NLP signal, as measured by a reduction in GFP+ cells. However, this did not increase templated DNA insertion (Fig. 6C). Retron-RTs interacted extensively with their ncRNA and msDNA via their C-terminal domains [41, 62], An extended C-terminal NLS may impair this interaction, reducing overall repair activity but not Cas9-cata lyzed DNA cleavage. It was concluded that nuclear import was likely not the limiting factor for templated insertion.

[0166] The linker between the nuclease and the RT can also impact editing outcomes.Fifteen linkers with various physical properties, including a ribosomal skipping T2A sequence between Cas9 and Efel-RT were tested (see Fig. 6D and Table 6). Separating the two enzymes via a T2A sequence reduced RFP+ cells by 20% relative to the reference linker, (SGGS)2-XTEN-(SGGS)2

[0063] , Next, a panel of flexible (e.g., (GGS)N) and rigid (e.g., (KL(A / E)AA)n) linkers were tested. Retron editors accommodated a broad range of linker designs (Table 6). However, multimerizing the RT on Cas9 via SpyTags or SunTags reduced RFP+ signal by 85-90%

[0064] , As Cas9-Sun / SpyTag-Cas9 fusions were reportedly active in mammalian cells, it was speculate that multimerization disrupts RT activity [64, 65], It was concluded that retron editors can accommodate a broad range of linker geometries, and even split enzyme designs.

[0167] To expand the retron editor target range, Efel-RT with AsCasl2a and AsCasl2a-Ultra nucleases were tested at five genomic loci (see Fig. 6E-G)

[0066] , Casl2a and Efel-RT were fused via an (SGGS)2-XTEN(SGGS)2linker, and the crRNA and msr-msd were expressed from independent promoters. In all cases, Casl2a-Ultra-RT fusions had higher insertion activity than the WT Casl2a (see Fig. 6E). This higher activity was due to the increased cleavage by Casl2a-Ultra relative to WT enzyme (see Fig. 7). Deep sequencing of the BRD8 locus confirmed that 99% of the inserts had the intended cargo sequence (see Fig. 6F). The most frequent error was an indel outside of the immediate insert site, followed by mismatches within the insert. Base substitution rates within the insert closely matched the pattern that was observed at EMX1 and were lower than those with Cas9 + ssODN donor (Figs. 6G, 5D). In sum, retron editors can be assembled from a broad range of ncRNA, NLS, and linkerconfigurations, and can be paired with Cas9 and Casl2a nucleases to expand their target range.

[0168] Channeling repair pathway choice increases insertion efficiency. DNA repair via non-homologous end joining (NHEJ) limits templated DNA insertion in mammalian cells [53, 67], To further increase retron editing efficiency, two approaches were utilized to channel repair away from NHEJ (see Fig. 8A). The focus of the first step was on small molecule inhibitors that inhibit NHEJ or enhance HDR. AZD7648 and M3814 inhibit the DNA- dependent protein-kinase catalytic subunit (DNA-PKcs) to improve templated repair of Cas9 breaks [68-71], TAK-931 is a CDC7-selective inhibitor that arrests cells in S phase, thereby increasing the HDR time window

[0072] , The optimal working concentrations were established for each inhibitor (see Fig. 9A). All three inhibitors improved insertion activity, with the strongest improvements with AZD7648 at all loci (see Fig. 8B, left). In contrast, M3814 showed strong improvements at all loci except HBB, and TAK-931 decreased retron editing at F9. Next, it was tested whether retron editors can insert larger cargos with 50 nt homology arms, and how this is modulated by inhibitors or DNA repair proteins (see Fig.8C). AZD7648 stimulated insertion of 25 and 50 nt inserts by 8.8- and 6.0-fold respectively at EMX1 (Fig. 8D). TAK-931 showed more modest 2.4- and 1.8-fold editing increases for 25 and 50 nt cargos, respectively. A comparison of Cas9-generated indels and retron-driven insertions confirmed that AZD7648 did not increase Cas9 cleavage but increased the utilization of a template ssDNA for genomic repair. Additionally, AZD7648 reduced the mutational signature at the target site, further highlighting the utility of repair pathway modulation in retron editing applications.

[0169] Cas9 fusion proteins can improve templated DNA insertion locally without perturbing repair pathways globally

[0053] , At this step, three fusions that had previously been characterized across multiple loci and cell types were the focus (see Figs. 8A, 8B, right) [73- 75], Fusing Cas9 with the HDR-promoting CtIP and a dominant negative RNF168 (dnRNF168) increased retron editing efficiency 1.8- to 2.5-fold across five loci. A dominant negative mutant of 53BP1 (DN1S) had variable effects across the five loci, with no improvements at EMX1 (see Fig. 8B, right)

[0074] , Fusion of Cas9 to the N-terminal region of human Geminin limited Cas9 expression to the S / G2 / M phases of the cell cycle, when HDR was most active

[0076] , However, Cas9-hGeml / 100 fusions either decreased or had no effect on retron editing activity (see Fig. 8B, right). This result may reflect the mechanistic differences betweenssODN and double-stranded donor DNA repair pathways, which is considered in depth later in the example. Combining Cas9-CtlP-dnRNF168 fusions with AZD7648 and TAK-931 did not improve editing any further at EMX1 (see Fig. 9B). This finding was consistent with the hypothesis that Cas9-CtlP-dnRNF168 and AZD7648 both down-regulate NHEJ, albeit via different mechanisms. In contrast, retron-mediated insertion with Cas9-DN1S and Cas9- hGeml / 110 fusions was higher withAZD7648 and TAK-931 but never exceeded AZD7648 or Cas9-CtlP-dnRNF168 alone. It was concluded that selective inhibition of DNA-PKcs significantly improves retron-mediated gene editing.

[0170] Next, retron editing activity with the nickase Cas9(D10A) was tested, and its fusions with DNA repair proteins (see Fig. 9C-E). On its own, Cas9(D10A) paired with Efel-RT caused an insertion in 1% of all reads at EMX1. However, fusing Cas9(D10A) to CtlP-dnRNF168 and a hyperactive HR recombinase hRad51 (K133R) increased insertion rates more than 20-fold (see Fig. 9D)

[0077] , It was contemplated that fusing nCas9 to a helicase could improve HDR by peeling back the target and non-target DNA strands for the retron-synthesized DNA, which was tested with Rep-X, an engineered helicase with exceptional processivity

[0078] , nCas9- Rep-X helicase fusions increased insertion efficiency more than 20-fold, suggesting that retron editors can be used for DSB-free editing. It is contemplate that nCas9-retron-RT editors and their combinations with DNA repair factors can be further optimized.

[0171] Retron editing in cell lines and vertebrates. As a proof of principle, retron editors were used to insert a split GFP for live cell imaging of endogenously expressed proteins in U2OS cells. GFP1-10, comprised of the first 10 GFP p-strands, was expressed from an integrated and inducible promoter (see Fig. 10). The 11thp-strand was fused to the protein of interest via a short linker. GFP1-10 was not fluorescent until it was complemented by GFP11 because chromophore maturation requires a critical GFPll-encoded glutamic acid

[0079] , Thus, fusing GFP11 to a target protein allowed visualization of sub-cellular localization in live cells [80-82],

[0172] Two proteins were targeted for endogenous GFP11 tagging via retron editors. The dispensable msd region was replaced with a 197 nt DNA template that included 70 nt homology arms, a 9 nt linker encoding Gly-Gly-Gly, and the 48 nt GFP11 epitope. To maximize the integration efficiency, cells were transfected with Cas9-CtlP-dnRNF168, which is discussed in depth in a later section. 72-hours post transfection, GFP1-10 was induced for 48 hours and GFP+ cells were collected using fluorescence-activated cell sorting (FACS).Confocal imaging of the insertion region confirmed the genomic edit (see Figs. 11C). These results demonstrate that Efel-RT can synthesize 200 nt ssDNAs. More broadly, this approach can be readily used for installing epitopes, disease-specific alleles, and other large insertions across the entire proteome from a genetically encoded cassette.

[0173] To expand retron editors beyond plasmid-based expression, delivery in an all-RNA format was tested (see Fig. 11). In these experiments, Cas9 and Efel-RT mRNAs were transcribed in vitro, 5'-capped and 3'-polyA-tailed. The msr-msd was transcribed in vitro, and the sgRNAs was chemically synthesized (see Fig. 11A). The four RNA-cocktail was transfected into HEK293T cells and editing efficiency scored by deep sequencing. Insertion of a 10 nt cargo was 9.9, 11.4, and 9.3% for the EMX1, CFTR, and AAVSl loci, respectively (see Fig. 11B). This lower editing efficiency was likely due to the reduced Cas9 and RT expression from mRNA relative to plasmid-based delivery. Optimizing the UTRs, chemical modifications, and mRNA to ncRNA ration may further boost editing at therapeutically relevant loci. It was concludes that retron editors are compatible with direct RNA delivery.

[0174] Next, the repair of a pathogenic mutation via mRNA-based retron editing in zebrafish was tested. A sgRNA was designed to target a pathogenic mutation in the Kinesin Family Member 6 kif6ut2° gene. Mutations in kif6ut2° cause scoliosis in zebrafish and are linked to neurological defects in humans

[0083] , The msd was designed to correct two base mutations that reverted a pathogenic Pro->Thr substitution. In addition, a silent T->C mutation was introduced that abolished a Bsal cleavage site for downstream restriction enzyme analysis (see Fig. 5E-F). Embryos were injected with the sgRNA, msr-msd, and a fused Cas9-RT or split Cas9 and Efel-RT mRNAs. Genomic DNA was harvested 24 hours post injection and was submitted to NGS, restriction enzyme digestion, and Sanger sequencing (see Fig. 11D-F). As expected, untreated controls did not show any editing. Injecting Cas9 and Efel-RT mRNAs, along with the two ncRNAs, induced edits in up to 10% of all reads from crude genomic preparations. The fused Cas9-RT mRNA showed reduced editing as compared to the split mRNAs (see Fig. 11D). This may be due to the stability of the fusion mRNA in zebrafish. To further confirm retron editing, the genomic DNA was sub-cloned and subjected to restriction enzyme and Sanger sequencing analysis (see Fig. 11E-F). Bsal treatment of the edited, but not WT and kif 6ut2° embryos, resulted in a single band, indicating the expected cleavage pattern. Sanger sequencing of the insert site sub-cloned from an edited embryoalso showed the expected edits. Collectively, these results indicate that retron editors are active in cell lines and zebrafish embryos when delivered as an all-RNA package.

[0175] Materials and Methods

[0176] Oligonucleotides and Plasmids. Retron-RT and ncRNAs gene blocks were ordered from IDT or Twist Biosciences and cloned into a GFP dropout entry vector via Golden Gate assembly. mRNAs were purchased from Cisterna Biologies. Retron editor expression plasmids were assembled by combining the sgRNA, retron msr-msd, the RT, and SpCas9 or AsCasl2a in a ccdB dropout mammalian expression vector using Golden Gate assembly.

[0177] Tissue Culture. HEK293T were generously provided by Professor XiaoluA. Cambronne. HEK293T and U2OS were cultured in Dulbecco's modified Eagle's medium (DMEM) with 10% fetal bovine serum (Gibco) and 1% Penicillin-Streptomycin (Gibco). All cell lines were cultured and maintained at 37 °C and 5% CO2. U2OS Fl p-l n TREx - GFP1-10 were generated by stably integrating a GFP1-10 construct at the single FRT locus through dual transfection of pcDNA5 / FRT / TO / lntron-eGFPl-10 and pOG44 plasmids in a 1:10 ratio and subsequent selection with hygromycin at 200 pg / mL and blasticidin at 15 pg / mL.

[0178] Metagenomic Discovery. To find new retrons, the NCBI genome and metagenomic contigs were searched for the retron-RT gene using hidden Markov models (HMMER, [97, 98]) and a database containing all experimentally validated retron-RT with a permissive e- value threshold of 10“4[33, 49, 50], Next, a pipeline was established to detect structured RNAs near retron RT sequences. Initially, each experimentally validated msr-msd transcript and its close counterparts served as seeds for CMfinder 0.4.1, an RNA motif predictor that leverages both folding energy and sequence covariation

[0099] , Covariate models were then crafted with Infernal suite's embuild and used to search for analogous structures around the start of the RT open reading frame (ORF). The msr-msd regions of all retron candidates were manually inspected, especially those that did not return any hits via the automated pipeline. RNAfold was used to inspect structured regions and to compare them to known msr-msd transcripts

[0100] , Then MAFFT-Q-INS-i was used for multiple alignments, focusing on identifying conserved sequences in related genomes. Subsequently, the alignments was pruned at the al and a2 areas, cycling back to earlier pipeline steps to acquire more sequences from both ends. R-scape was used to pinpoint covarying base pairs in the suggested consensus structures, mitigating the influence of phylogenetic correlations and base composition biases not attributed to conserved RNA structure. Accessory retron genesadjacent to the RT were annotated using the Pfam database as a query and HMMER. High- confidence retron systems were manually inspected to confirm the expected protein domains, catalytic residues, and ncRNA structure. To remove redundant sequences, all putative hits were clustered with CD-HIT

[0101] , setting a sequence identity threshold at 90% and an alignment overlap of 80%. The clustered datasets were then transformed into phylogenetic trees, and candidates for validation were selected based on their positioning in the tree.

[0179] To organize the retron RTs into phylogenetic groups, we employed MAFFT software was employed, and progressive methods were used for multiple sequence alignments (MSAs)

[0102] , An MSA was constructed from the RTO-7 domain of 1,912 sequences, sourced from a dataset of 9,141 entries previously categorized as retron / retron-like RTs and an additional 16 RTs from experimentally verified retrons

[0033] , Phylogenetic trees were generated using FastTree, applying the WAG evolutionary model, combined with a discrete gamma model featuring twenty rate categories. Specifically, the RT tree was crafted using IQ-TREE vl.6.12, incorporating 1000 ultra-fast bootstraps (UFBoot) and the SH-like approximate likelihood ratio test (SH-aLRT) with 1000 iterations

[0103] , The best-fit model, identified by Modelfinder as the LG+F+R10 due to its minimal Bayesian Information Criterion (BIC) value among 546 protein models, was used. The RT Clades' internal nodes exhibited UFBoot and SH-aLRT support values exceeding 85%

[0101] ,

[0180] Plasmid-based fluorescent reporter assays. 1.2xl05HEK293T cells were seeded in a 24-well plate 18-24 hours before transfection. 0.35 pg of the retron editor plasmid and 0.35 pg of the fluorescent reporter plasmid were co-transfected using Lipofectamine 2000 (Invitrogen). Cells were trypsinized and collected for flow analysis 72 hours after transfection. Flow analysis was conducted on a Novocyte flow cytometer (ACEA Biosciences). Cells were gated to exclude dead cells and doublets, and 10,000 cells were analyzed in all samples. Cells were then gated by FITC-A (x-axis) and PE-Texas Red-A (y-axis). The editing efficiency was reported as the percentage of cells in the quadrant of FITC-A and PE-Texas Red-A.

[0181] Genomic reporter assays. The fluorescent reporter was cloned in a plasmid designed for Bxbl recombinase-driven landing pad system

[0104] , This plasmid was transfected into landing pad HEK293T cells followed by doxycycline induction and AP1903 selection to generate stably integrated reporter cells. 1.2xl05of HEK293T reporter cellswere seeded in a 24-well plate 18-24 hours prior to transfection. 1 pg of the retron editor plasmid was transfected using Lipofectamine 2000. 48-72 hours after transfection, cells were treated with doxycycline to induce the expression of the fluorescent reporter. Cells were trypsinized and collected for flow analysis and genomic DNA extraction (Qiagen DNeasy Blood and Tissue kit).

[0182] Confocal imaging. HEK293T cells were seeded and transfected as described for the plasmid-based fluorescent reporter assay (see above). 48 hours post -transfection, cells were seeded into 15 mm glass-bottom cell culture dishes (NEST) and incubated for an additional 24 hours. Cells were then imaged with a Nikon Ti2 Spinning Disk Confocal Microscope at 20x magnification. 8858 x 8858 pixel images were acquired and processed using imageJ.

[0183] U2OS Flpl n TREx GFP1-10 cells were transfected with 1 pg retron editor plasmids that target either the N- or C-terminus of the protein of interest to insert a GFPn fragment. 48-72 hours after transfection, cells were expanded into 10-cm plates. Cells with high GFP intensities were then sorted via a cell sorter (Sony MA900). For confocal imaging, cells were incubated with doxycycline for 48-72 hours before imaging to induce the expression of the GFP1-10 construct. Image acquisition was performed with live cells under spinning-disk confocal microscopy (Olympus).

[0184] Next generation DNA sequencing. Genomic samples were subjected to two rounds of PCR for NGS library preparation 2. The first round of PCR was performed using the KOD One PCR master mix (TOYOBO). Primers were designed about 600 base pairs away from the cut site on each side to avoid amplifying the RT-generated ssDNA. PCR products were gel purified and barcoded via a second round of PCR with Illumina P5 / P7 adapters using Q5 HotStart High Fidelity master mix (NEB). PCR amplicons were sequenced on an Illumina Novaseq. Reads were demultiplexed using NovaSeq Reporter (Illumina). Alignment of amplicon sequences to a reference sequence was performed using CRISPResso2

[0105] , The aligned sequences were first checked for the expected cargo insertion. Next, the homology arms and insertion sequences were further checked for any mismatches or deletion. A "perfect edit" was defined as any sequence that had only the expected insertion, with no additional mismatches and deletions. If these events also had a mismatch or deletion, there were counted as an "imperfect" edit. Any insertions, deletions, or mismatches that did not have the intended cargo were surmised to be from nuclease-only activity.

[0185] In vitro transcription of the msr-msd. The msr-msd was PCR amplified from a plasmid or gene block with a T7 promoter. The ncRNA was generated using the HiScribe T7 High Yield RNA Synthesis Kit (NEB) according to the manufacturer's protocol. RNA products were purified using the RNeasy Mini Kit (Qiagen).

[0186] Zebrafish Maintenance and Gene Editing. All zebrafish experiments were performed according to University of Texas at Austin IACUC standards. Embryos were raised at 28.5 °C in fish water (0.15% Instant Ocean in reverse osmosis water). Wildtype and kif6ut20(P293T)mutants were in-crossed, and one-cell stage embryos were injected with 1 nL of the injection mix using a microinjector-pump system (World Precision Instruments Nanoliter Injector and PV 820 Pneumatic PicoPump). Injected embryos and un-injected sibling controls were incubated at 28.5 °C in fish water until 24 hr post fertilization, at which point, surviving embryos were euthanized in excess Tricaine (0.4% MS-222). Genomic DNA was extracted from individual embryos using the HotSHOT Method

[0106] , Briefly, embryos were transferred into 50 mM NaOH and heated to 95 °C for 20 min. The samples were neutralized with a quarter volume of 1 M Tris-HCI, pH 8, prior to downstream analysis. For NGS sequencing, genomic DNA was PCR amplified to extend the amplicon with Illumina adapters using the KOD One PCR master mix (TOYOBO). PCR amplicons were directly sequenced on an Illumina NovaSeq sequencer.

[0187] Discussion

[0188] Here, retron editors were engineered for precise genome engineering in mammalian cells and zebrafish embryos. The metagenomic survey yielded 17 RTs that outperform Ecol- RT in mammalian cells. Efel-RT, the lead candidate, was highly active, specific for its cognate RNA, and was capable of generating at least ~200 nt ssDNAs in nucleo (Fig. 1). The most active RTs in the survey were all derived from clade 9. A recent bacterial functional screen also concluded that RTs in this clade generate high ssDNA levels in E. coli

[0017] , It is contemplated that additional structure-function studies eludicate the mechanistic basis for this higher activity. Finally, retron editor delivery in an all-RNA format was demonstrated (Fig. 11). Based on the findings, it is contemplated that further optimizations, such as increasing mRNA and ncRNA stability will further improve editing efficiency.

[0189] Prior designs were iteratively improved upon by optimizing the NLS, nuclease-RT linkers, and nuclease combinations (Fig. 6). The linker and NLS are key design components for both base and prime editors [84, 85], In contrast, it was shown that retron editors canaccommodate diverse NLS and linker combinations. Importantly, retron editors are compatible with Cas9, Casl2a, and even the nickase Cas9(D10A) (Fig. 6, 9). Casl2a-based retron editors further expand the potential target range and create opportunities for multiplexed retron editing due to Casl2a's ability to process its crRNA

[0084] , Based on the results, it appears that retron-RTs are compatible with other established, i.e. transcription activator-like effectors (TALEs), and emerging RNA / DNA-guided nucleases. Importantly, further development of nickase-based retron editors can avoid the induction of doublestranded DNA breaks (Fig. 9).

[0190] Retron editors are uniquely capable of synthesizing high copy numbers of the repair template at the edit site [48, 86, 87], Boosting ssDNA synthesis via RT and ncRNA engineering will further improve the processivity, fidelity, and ultimately, edit length and efficiency. For example, rational engineering of a retron RT-based prime editor boosted editing efficiency more than 8-fold

[0011] , The insertion and truncation site of the native msd, homology arm length, target / nontarget strand selection, and overall RNA structure can be modified. RNA circularization, structured RNA pseudoknots, and chemical modifications also increase ncRNA stability in mammalian cells

[0088] , Further optimization in these directions can enhance the overall efficiency of retron editing. In general principles and predictive algorithms for msr-msd and repair template will further improve retron editors.

[0191] The repair of DSBs with a single-stranded DNA donor proceeds via single-strand template repair (SSTR) [89-91], The results are also consistent with an SSTR-based repair mechanism for retron editors. First, NHEJ inhibitors increase retron editing and reduce NHEJ-associated indels for both Cas9 and Casl2a across all tested target sites (Fig. 8). Second, fusing Cas9 with a dominant negative allele of 53BP1, or with the DNA resection promoting CtIP, also increases retron editor efficiency. Third, it was observed that 3650 nt homology arms were optimal for retron editors (Fig. 5). Similarly, SSTR was maximized with 30-60 nt homology arms [92, 93], SSTR is a RAD52-dependent process in yeast and human cells, suggesting that nuclease-RAD52 and / or RT-RAD52 fusions may boost retron editing [75, 94], SSTR competes with two error-prone DSB repair pathways: classical NHEJ and polymerase theta-mediated end joining (TMEJ) [52, 95], Dual inhibition of NHEJ and TMEJ may further synergize with RAD52 fusions [47, 69], Rational design of asymmetric templates, cleavage-blocking mutations, and dual Cas9 nickases can be done to maximizeediting efficiency [93, 96], Mechanistic studies of retron editor-mediated repair will further improve editing outcomes across all domains of life.

[0192] In conclusion, retron editors are emerging as a highly promising gene editing tool. Their unique ability to accurately insert or replace sizeable DNA segments opens up possibilities for correcting complex genetic mutations that were previously challenging to address. Compatibility with an all-RNA formulation opens new avenues for therapeutic delivery into cells and organisms. Additionally, retron editors are poised to broaden the scope of high-throughput functional screens, allowing for the characterization of complex genetic variants with single-base resolution. Integration of retron editors into existing screening pipelines holds great promise for advancing the understanding of gene function and regulation, ultimately paving the way for novel therapeutic interventions and biotechnological applications.EXAMPLE 2: Retron Editors

[0193] A bioinformatics discovery pipeline that identifies genes in the retron operon and annotates the RT-associated non-coding RNAs was developed. Open reading frames in metagenomic contigs are identified using established tools (i.e., Prodigal), while putative retron RTs are identified using Hidden Markov Models (HMMs). To reduce false positives, a bootstrapped phylogeny of retron RTs based on multiple sequence alignments (MSAs) was constructed, with non-retron RTs as outgroups. Accessory retron genes that are adjacent to the RT were annotated using the Pfam database.

[0194] Second, the non-coding RNA that is recognized by the retron RT is annotated, based on conserved features of the ncRNA: it is almost always in an intergenic region between the RT and the next ORF, and the 5'-3' ends encode a long palindromic repeat. After identifying palindromic repeats in the non-coding region adjacent to the RT, covariate models of the msr-msd region were built using the experimentally validated msr-msd sequences. The putative ncRNA is folded using ViennaRNA 2.0 to predict the msr-msd structure and visual verification is conducted on a subset of candidates. This approach ensures that the most likely RT-ncRNA candidates can be identified for downstream experimental validation. The combination of phylogenetic clustering, the presence of expected accessory proteins, andthe identification of plausible, adjacent msr-msd sequences provide a set of high-confidence retron systems for downstream characterization.

[0195] To screen for the most active variants, a "traffic light" HEK293T reporter cell line was used that consists of a red fluorescent protein (RFP) with a deletion that renders it inactive, and a GFP that is separated by a T2A tag. Expression of a retron-Cas9 fusion, along with the ncRNA, will generate a DNA break adjacent to the RFP deletion. This break can then be repaired by the retron-generated msDNA to restore RFP fluorescence in mammalian tissue culture cells. This is detected via fluorescence activated cell sorting.

[0196] Using the bioinformatics approach, about 100 novel retrons were synthesized and screened using the traffic light reporter system, with the reporter on one plasmid and the retron editor on another plasmid. Following this screen, the traffic light reporter was integrated into the mammalian genome to better test genomic DNA editing. The approximately 20 most active retron editors were tested in the genomic reporter assay and the most active variants were selected for further optimization. The most active retron editor was further optimized by testing linkers between the nuclease (i.e., Cas9 and the RT; Fig. 14), as well as the nuclear localization signals (NLS) on the Cas9 and the RT (Fig. 16). Constructs were tested with (1) different nuclear localization signals, (2) flexible and rigid linkers between the Cas nuclease and the RT, and (3) the promoters expressing the polypeptides and ncRNAs. This screen and engineering effort produced a first-generation retron editor that can efficiently and specifically integrate cargo DNA into mammalian cells. This can be seen in Fig. 1.TABLESTable 1. Plasmids used in Example 1Table 2. Oligonucleotides used to amplify genomic loci in this study. [BC] indicates a unique Illumina indexTable 2. (continued)Table 3: Spacers for single guide RNAs (sgRNAs) and CRISPR RNAs (crRNAs) used in Example 1Table 4: Nuclear localization sequences used in Example 1Table 5: Nuclear localization signals used in Example 1Table 6: Linker sequences used in Example 1Table 7. Retrons and msr-msd SequencesREFERENCES1. Fichter, K. M., Setayesh, T. & Malik, P. Strategies for precise gene edits in mammalian cells. Molecular Therapy. Nucleic Acids 32, 536-552. ISSN: 2162- 2531. (2024) (Apr. 2023).2. Jacinto, F. V., Link, W. & Ferreira, B. I. CRISPR / Cas9-mediated genome editing: From basic research to translational medicine, en. Journal of Cellular and Molecular Medicine 24. 14916, 3766-3778. ISSN: 1582-4934 (2024) (2020).3. Jang, H.-K., Song, B., Hwang, G.-H. & Bae, S. Current trends in gene recovery mediated by the CRISPR-Cas system, en. Experimental & Molecular Medicine 52. Publisher: Nature Publishing Group, 1016-1027. ISSN: 2092-6413. (2024) (July 2020).4. Li, L., Hu, S. & Chen, X. Non-viral delivery systems for CRISPR / Cas9-based genome editing: Challenges and opportunities. Biomaterials 171, 207-218. ISSN: 0142-9612. (2024) (July 2018).5. Luther, D., Lee, Y., Nagaraj, H., Scaletti, F. & Rotello, V. Delivery approaches for CRISPR / Cas9 therapeutics in vivo: advances and challenges. Expert Opinion on Drug Delivery 15. Publisher: Taylor & Francis 905-913. ISSN: 1742-5247. (2024) (Sept. 2018).6. Wang, H.-X. et al. CRISPR / Cas9-Based Genome Editing for Disease Modeling and Therapy: Challenges and Opportunities for Nonviral Delivery. Chemical Reviews 117. Publisher: American Chemical Society, 9874-9906. ISSN: 0009- 2665. (2024) (Aug. 2017).7. Yip, B. H. Recent Advances in CRISPR / Cas9 Delivery Strategies, en. Biomolecules 10. Number: 6 Publisher: Multidisciplinary Digital Publishing Institute, 839. ISSN: 2218-273X. (2024) (June 2020).8. David, R. M. & Doherty, A. T. Viral Vectors: The Road to Reducing Genotoxicity. Toxicological Sciences 155, 315-325. ISSN: 1096-6080. ..(2024) (Feb. 2017).9. Anzalone, A. V. et al. Search-and-replace genome editing without doublestrand breaks or donor DNA. eng. Nature 576, 149-157. ISSN: 1476-4687 (Dec. 2019).10. Chen, P. J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. English. Cell 184. Publisher: Elsevier, 5635- 5652. e29. ISSN: 00928674, 1097-4172. (Oct. 2021).Doman, J. L. et al. Phage-assisted evolution and protein engineering yield compact, efficient prime editors, en. Cell 186, 3983-4002. e26. ISSN: 00928674. (2024) (Aug. 2023). Tang, S. & Sternberg, S. H. Genome editing with retroelements. Science (New York, N. Y.) 382, 370-371. ISSN: 0036-8075. (2024) (Oct. 2023). Bhattarai-Kline, S. et al. Recording gene expression order in DNA by CRISPR addition of retron barcodes, en. Nature 608. Publisher: Nature Publishing Group, 217-225. ISSN: 14764687. (2024) (Aug. 2022). Fishman, C. B. et al. Continuous Multiplexed Phage Genome Editing Using Recombitrons en. Pages: 2023.03.24.534024 Section: New Results. Mar. 2023. (2024). Hwang, J., Ye, D.-Y., Jung, G. Y. & Jang, S. Mobile genetic element-based gene editing and genome engineering: Recent advances and applications. Biotechnology Advances 72, 108343. ISSN: 0734-9750. (2024) (May 2024). Kaur, N. & Pati, P. K. Retron Library Recombineering: Next Powerful Tool for Genome Editing after CRISPR / Cas. ACS Synthetic Biology 13. Publisher: American Chemical Society, 1019- 1025. 2024) (Apr. 2024). Khan, A. G. et al. An experimental census of retrons for DNA production and genome editing en. Pages: 2024.01.25.577267 Section: New Results. Jan. 2024. (2024). Lee, G. & Kim, J. Engineered retrons generate genome-independent proteinbinding DNA for cellular control en. Sept. 2023. (2024). Lim, H. et al. Multiplex Generation, Tracking, and Functional Screening of Substitution Mutants Using a CRISPR / Retron System. ACS Synthetic Biology 9. Publisher: American Chemical Society, 1003-1009. (2024) (May 2020). Lin, Q. et al. Prime genome editing in rice and wheat, en. Nature Biotechnology 38. Publisher: Nature Publishing Group, 582-585. ISSN: 1546-1696. (2024) (May 2020). Liu, J. et al. Generation of DNAzyme in Bacterial Cells by a Bacterial Retron System. ACS Synthetic Biology 13. Publisher: American Chemical Society, BOOBOO. (2024) (Jan. 2024).Liu, W. et al. Retron-mediated multiplex genome editing and continuous evolution in Escherichia coli. Nucleic Acids Research 51, 8293-8307. ISSN:0305-1048. (2024) (Aug. 2023). Ramirez-Chamorro, L., Boulanger, P. & Rossier, O. Strategies for Bacteriophage T5 Mutagenesis: Expanding the Toolbox for Phage Genome Engineering. English. Frontiers in Microbiology 12. Publisher: Frontiers. ISSN: 1664-302X. (2024) (Apr. 2021). Roy, K. R. et al. Dissecting guantitative trait nucleotides by saturation genome editing en. Pages: 2024.02.02.577784 Section: New Results. Feb. 2024. (2024). Schubert, M. G. et al. High-throughput functional variant screens via in vivo production of single-stranded DNA. Proceedings of the National Academy of Sciences 118. Publisher: Proceedings of the National Academy of Sciences, e2018181118. (2024) (May 2021). Simon, A. J., Morrow, B. R. & Ellington, A. D. Retroelement-Based Genome Editing and Evolution. ACS Synthetic Biology 7. Publisher: American Chemical Society, 2600-2611. (2024) (Nov. 2018). Ellington, A. J. & Reisch, C. R. Efficient and iterative retron-mediated in vivo recombineering in Escherichia coli. Synthetic Biology 7, ysac007. ISSN: 2397- 7000. (2024) (Jan. 2022). Farzadfard, F. & Lu, T. K. Genomically encoded analog memory with precise in vivo DNA writing in living cell populations, en. Science 346, 1256272. ISSN: 0036-8075, 1095-9203. (2024) (Nov. 2014). Gonzalez-Delgado, A., Lopez, S. C., Rojas-Montero, M., Fishman, C. B. & Shipman, S. L. Simultaneous multi-site editing of individual genomes using retron arrays en. July 2023. (2024). Kong, X. et al. Precise genome editing without exogenous donor DNA via retron editing system in human cells, en. Protein & Cell 12, 899-902. ISSN: 1674-800X, 1674-8018 (2024) (Nov. 2021). Lopez, S. C., Crawford, K. D., Lear, S. K., Bhattarai-Kline, S. & Shipman, S. L. Precise genome editing across kingdoms of life using retron-derived DNA. en. Nature Chemical Biology 18, 199-206. ISSN: 1552-4450, 1552-4469. (2024) (Feb. 2022).Sharon, E. et al. Functional Genetic Variants Revealed by Massively Parallel Precise Genome Editing, en. Cell 175, 544-557.el6. ISSN: 00928674 (2024) (Oct. 2018). Simon, A. J., Ellington, A. D. & Finkelstein, I. J. Retrons and their applications in genome engineering, en. Nucleic Acids Research 47, 11007-11019. ISSN: 0305- 1048, 1362-4962. (2024) (Dec. 2019). Zhao, B., Chen, S.-A. A., Lee, J. & Fraser, H. B. Bacterial Retrons Enable Precise Gene Editing in Human Cells, eng. The CRISPR journal 5, 31-39. ISSN: 2573- 1602 (Feb. 2022). Jiang, W. et al. High-efficiency retron-mediated single-stranded DNA production in plants. Synthetic Biology 7, ysac025. ISSN: 2397-7000. (2024) (Jan. 2022). Bobonis, J. et al. Bacterial retrons encode phage-defending tripartite toxinantitoxin systems, en. Nature 609, 144-150. ISSN: 0028-0836, 1476-4687. (2024) (Sept. 2022). Gao, L. et al. Diverse enzymatic activities mediate antiviral immunity in prokaryotes. Science 369. Publisher: American Association for the Advancement of Science, 1077-1084. (2024) (Aug. 2020). Millman, A. et al. Bacterial Retrons Function In Anti-Phage Defense, en. Cell 183, 1551-1561. el2. ISSN: 00928674. (Dec. 2020). Palka, C., Fishman, C. B., Bhattarai-Kline, S., Myers, S. A. & Shipman, S. L. Retron reverse transcriptase termination and phage defense are dependent on host RNase Hl. en. Nucleic Acids Research 50, 3490-3504. ISSN: 0305-1048, 1362- 4962. (2024) (Apr. 2022). Rychlik, I., Sebkova, A., Gregorova, D. & Karpiskova, R. Low-Molecular-Weight Plasmid of Salmonella enterica Serovar Enteritidis Codes for Retron Reverse Transcriptase and Influences Phage Resistance. Journal of Bacteriology 183, 2852-2858. ISSN: 0021-9193. (2024) (May 2001). Wang, Y. et al. Cryo-EM structures of Escherichia coli Ec86 retron complexes reveal architecture and defence mechanism, en. Nature Microbiology 7. Publisher: Nature Publishing Group, 1480-1489. ISSN: 2058-5276. (2024) (Sept. 2022).Carabias, A. et al. Retron-Ecol assembles NAD+-hydrolyzing filaments that provide immunity against bacteriophages. English. Molecular Cell 0. Publisher: Elsevier. ISSN: 10972765. (May 2024). Wang, Y. et al. Defense mechanism of a bacterial retron supramolecular assembly en. Pages: 2023.08.16.553469 Section: New Results. Aug. 2023. (2024). Azam, A. H. et al. Viruses encode tRNA and anti-retron to evade bacterial immunity en. Pages: 2023.03.15.532788 Section: New Results. Mar. 2023. (2024). Hsu, M. Y., Eagle, S. G., Inouye, M. & Inouye, S. Cell-free synthesis of the branched RNAIinked msDNA from retron-Ec67 of Escherichia coli. eng. The Journal of Biological Chemistry 267, 13823-13829. ISSN: 0021-9258 (July 1992). Shimamoto, T., Inouye, M. & Inouye, S. The formation of the 2', 5'- phosphodiester linkage in the cDNA priming reaction by bacterial reverse transcriptase in a cell-free system, eng. The Journal of Biological Chemistry 270, 581-588. ISSN: 0021-9258 (Jan. 1995). Schimmel, J. et al. Modulating mutational outcomes and improving precise gene editing at CRISPR-Cas9-induced breaks by chemical inhibition of endjoining pathways, eng. Cell Reports 42, 112019. ISSN: 2211-1247 (Feb. 2023). Savic, N. et al. Covalent linkage of the DNA repair template to the CRISPR-Cas9 nuclease enhances homology-directed repair. eLife 7 (ed de Massy, B.) Publisher: eLife Sciences Publications, Ltd, e33761. ISSN: 2050-084X. (2024) (May 2018). Almeida, A. et al. A unified catalog of 204,938 reference genomes from the human gut microbiome, en. Nature Biotechnology 39. Publisher: Nature Publishing Group, 105-114. ISSN: 1546-1696. (2024) (Jan. 2021). Pruitt, K. D., Tatusova, T. & Maglott, D. R. NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. Nucleic Acids Research 35, D61-D65. ISSN: 0305-1048. (2024) (Jan. 2007). Mestre, M. R., Gonzalez-Delgado, A., Gutierrez-Rus, L. I., Martinez-Abarca, F. & Toro, N. Systematic prediction of genes functionally associated with bacterial retrons and classification of the encoded tripartite systems, en.Nucleic Acids Research 48, 12632-12647. ISSN: 0305-1048, 1362-4962. (2024) (Dec. 2020). Hussmann, J. A. et al. Mapping the genetic landscape of DNA double-strand break repair, eng. Cell 184, 5653-5669.e25. ISSN: 1097-4172 (Oct. 2021). Nambiar, T. S., Baudrier, L., Billon, P. & Ciccia, A. CRISPR-based genome editing through the lens of DNArepair. en. Molecular Cell 82, 348-388. ISSN: 10972765. (2024) (Jan. 2022). Potapov, V. et al. Base modifications affecting RNA polymerase and reverse transcriptase fidelity. Nucleic Acids Research 46, 5753-5763. ISSN: 0305-1048. (2024) (June 2018). Yasukawa, K. et al. Next-generation sequencing-based analysis of reverse transcriptase fidelity. Biochemical and Biophysical Research Communications 492, 147-153. ISSN: 0006291X. (Oct. 2017). Goeckel, M. E. et al. Modulating CRISPR gene drive activity through nucleocytoplasmic localization of Cas9 in S. cerevisiae. en. Fungal Biology and Biotechnology 6, 2. ISSN: 20543085. (2024) (Dec. 2019). Liu, P. et al. Improved prime editors enable pathogenic allele correction and cancer modelling in adult mice. en. Nature Communications 12, 2121. ISSN: 2041-1723. (2024) (Apr. 2021). Maggio, I. et al. Integrating gene delivery and gene-editing technologies by adenoviral vector transfer of optimized CRISPR-Cas9 components, en. Gene Therapy 27, 209-225. ISSN: 0969-7128, 1476-5462. (2024) (May 2020). Suzuki, K. et al. In vivo genome editing via CRISPR / Cas9 mediated homologyindependent targeted integration, en. Nature 540, 144-149. ISSN: 0028-0836, 1476-4687. (2024) (Dec. 2016). Wu, Y. et al. Highly efficient therapeutic gene editing of human hematopoietic stem cells, en. Nature Medicine 25, 776-783. ISSN: 1078-8956, 1546-170X. (2024) (May 2019). Luk, K. et al. Optimization of Nuclear Localization Signal Composition Improves CRISPR- Casl2a Editing Rates in Human Primary Cells, en. GEN Biotechnology 1, 271-284. ISSN: 2768-1572, 2768-1556. (2024) (June 2022). Inouye, S., Hsu, M.-Y., Xu, A. & Inouye, M. Highly Specific Recognition of Primer RNA Structures for 2'-OH Priming Reaction by Bacterial ReverseTranscriptases*. Journal of Biological Chemistry 274, 31236-31244. ISSN: 0021-9258. (2024) (Oct. 1999). Tan, J., Zhang, F., Karcher, D. & Bock, R. Engineering of high-precision base editors for sitespecific single nucleotide replacement. Nature Communications 10, 439. ISSN: 2041-1723. (2024) (Jan. 2019). Tanenbaum, M. E., Gilbert, L. A., Qi, L. S., Weissman, J. S. & Vale, R. D. A Protein- Tagging System for Signal Amplification in Gene Expression and Fluorescence Imaging. English. Cell 159. Publisher: Elsevier, 635-646. ISSN: 0092-8674, 1097- 4172. (2024) (Oct. 2014). Hinrichsen, M. et al. A new method for post-translationally labeling proteins in live cells for fluorescence imaging and tracking. Protein Engineering, Design and Selection 30, 771-780. ISSN: 1741-0126. (2024) (Dec. 2017). Zhang, L. et al. AsCaslZa ultra nuclease facilitates the rapid generation of therapeutic cell medicines, en. Nature Communications 12. Publisher: Nature Publishing Group, 3908. ISSN: 2041-1723. (2024) (June 2021). Carusillo, A. & Mussolino, C. DNA Damage: From Threat to Treatment, eng. Cells 9, 1665. ISSN: 2073-4409 (July 2020). Fok, J. H. L. et al. AZD7648 is a potent and selective DNA-PK inhibitor that enhances radiation, chemotherapy and olaparib activity, en. Nature Communications 10. Publisher: Nature Publishing Group, 5065. ISSN: 2041- 1723. (2024) (Nov. 2019). Wimberger, S. et al. Simultaneous inhibition of DNA-PK and Pol© improves integration efficiency and precision of genome editing. Nature Communications 14, 4761. ISSN: 20411723. (2024) (Aug. 2023). Fu, Y.-W. et al. Dynamics and competition of CRISPR-Cas9 ribonucleoproteins and AAV donor-mediated NHEJ, MMEJ and HDR editing, en. Nucleic Acids Research 49. Publisher: Oxford University Press, 969. (2024) (Jan. 2021). Riesenberg, S. et al. Simultaneous precise editing of multiple genes in human cells. Nucleic Acids Research 47, ell6. ISSN: 0305-1048. (2024) (Nov. 2019). Iwai, K. et al. Molecular mechanism and potential target indication of TAK-931, a novel CDC7-selective inhibitor. Science Advances 5, eaav3660. ISSN: 2375- 2548. (2024) (May 2019).Reint, G. et al. Rapid genome editing by CRISPR-Cas9-POLD3 fusion, en. eLife 10, e75415. ISSN: 2050-084X (2024) (Dec. 2021). Jayavaradhan, R. et al. CRISPR-Cas9 fusion to dominant-negative 53BP1 enhances HDR and inhibits NHEJ specifically at Cas9 target sites, en. Nature Communications 10, 2866. ISSN: 2041-1723. (June 2019). Carusillo, A. et al. A novel Cas9 fusion protein promotes targeted genome editing with reduced mutational burden in primary human cells, en. Nucleic Acids Research 51, 4660- 4673. ISSN: 0305-1048, 1362-4962. (2024) (May 2023). Gutschner, T., Haemmerle, M., Genovese, G., Draetta, G. F. & Chin, L. Post- translational Regulation of Cas9 during G1 Enhances Homology-Directed Repair, en. Cell Reports 14, 1555-1566. ISSN: 22111247. (2024) (Feb. 2016). Rees, H. A., Yeh, W.-H. & Liu, D. R. Development of hRad51-Cas9 nickase fusions that mediate HDR without double-stranded breaks, en. Nature Communications 10, 2212. ISSN: 2041-1723. (2024) (May 2019). Arslan, S., Khafizov, R., Thomas, C. D., Chemla, Y. R. & Ha, T. Engineering of a superhelicase through conformational control, en. Science 348, 344-347. ISSN: 0036-8075, 10959203. (2024) (Apr. 2015). Barondeau, D. P., Putnam, C. D., Kassmann, C. J., Tainer, J. A. & Getzoff, E. D. Mechanism and energetics of green fluorescent protein chromophore synthesis revealed by trapped intermediate structures. Proceedings of the National Academy of Sciences 100. Publisher: Proceedings of the National Academy of Sciences, 12111-12116. (2024) (Oct. 2003). Kamiyama, D. et al. Versatile protein tagging in cells with split fluorescent protein, en. Nature Communications 7. Publisher: Nature Publishing Group, 11046. ISSN: 2041-1723. (2024) (Mar. 2016). Pinaud, F. & Dahan, M. Targeting and imaging single biomolecules in living cells by complementation-activated light microscopy with split-fluorescent proteins, eng. Proceedings of the National Academy of Sciences of the United States of America 108, E201-210. ISSN: 1091-6490 (June 2011). Cabantous, S., Terwilliger, T. C. & Waldo, G. S. Protein tagging and detection with engineered self-assembling fragments of green fluorescent protein, en.Nature Biotechnology 23. Publisher: Nature Publishing Group, 102-107. ISSN: 1546-1696. (2024) (Jan. 2005). Konjikusic, M. J. et al. Mutations in Kinesin family member 6 reveal specific role in ependymal cell ciliogenesis and human neurological development, en. PLOS Genetics 14 (ed Pazour, G. J.) el007817. ISSN: 1553-7404. (2024) (Nov. 2018). Rees, H. A. & Liu, D. R. Base editing: precision chemistry on the genome and transcriptome of living cells. Nature reviews. Genetics 19, 770-788. ISSN: 1471- 0056. (2024) (Dec. 2018). Zhao, Z., Shang, P., Mohanraju, P. & Geijsen, N. Prime editing: advances and therapeutic applications. English. Trends in Biotechnology 41. Publisher: Elsevier, 1000-1012. ISSN: 0167-7799, 1879-3096. (2024) (Aug. 2023). Carlson-Stevermer, J. et a / .Assembly of CRISPR ribonucleoproteins with biotinylated oligonucleotides via an RNA aptamer for precise gene editing, en. Nature Communications 8. Publisher: Nature Publishing Group, 1711. ISSN: 2041-1723. (2024) (Nov. 2017). Aird, E. J., Lovendahl, K. N., St. Martin, A., Reuben S. Harris & Gordon, W. R. Increasing Cas9-mediated homology-directed repair efficiency through covalent tethering of DNA repair template, en. Communications Biology 1. Publisher: Nature Publishing Group, 1-6. ISSN: 2399-3642. (2024) (May 2018). Chen, P. J. & Liu, D. R. Prime editing for precise and highly versatile genome manipulation, en. Nature Reviews Genetics 24. Publisher: Nature Publishing Group, 161-177. ISSN: 1471-0064. (2024) (Mar. 2023). Gallagher, D. N. et al. A Rad51-independent pathway promotes single-strand template repair in gene editing, eng. PLoS genetics 16, el008689. ISSN: 1553- 7404 (Oct. 2020). Gallagher, D. N. & Haber, J. E. Single-strand template repair: key insights to increase the efficiency of gene editing, en. Current Genetics 67, 747-753. ISSN: 1432-0983. (2024) (Oct. 2021). Richardson, C. D. et al. CRISPR-Cas9 genome editing in human cells occurs via the Fanconi anemia pathway, en. Nature Genetics 50. Publisher: Nature Publishing Group, 1132- 1139. ISSN: 1546-1718. (2024) (Aug. 2018). Okamoto, S., Amaishi, Y., Maki, L, Enoki, T. & Mineno, J. Highly efficient genome editing for single-base substitutions using optimized ssODNs withCas9-RNPs. en. Scientific Reports 9. Publisher: Nature Publishing Group, 4811.ISSN: 2045-2322. (2024) (Mar. 2019). Schubert, M. S. et al. Optimized design parameters for CRISPR Cas9 and Casl2a homology-directed repair, en. Scientific Reports 11. Publisher: Nature Publishing Group, 19482. ISSN: 2045-2322. (2024) (Sept. 2021). Paulsen, B. S. et al. Ectopic expression of RAD52 and dn53BPl improves homology-directed repair during CRISPR-Cas9 genome editing, en. Nature Biomedical Engineering 1. Publisher: Nature Publishing Group, 878-888. ISSN: 2157-846X. (2024) (Nov. 2017). Wyatt, D. W. et al. Essential Roles for Polymerase 0-Mediated End Joining in the Repair of Chromosome Breaks, eng. Molecular Cell 63, 662-673. ISSN: 1097-4164 (Aug. 2016). Richardson, C. D., Ray, G. J., DeWitt, M. A., Curie, G. L. & Corn, J. E. Enhancing homologydirected genome editing by catalytically active and inactive CRISPR- Cas9 using asymmetric donor DNA. en. Nature Biotechnology 34. Publisher: Nature Publishing Group, 339-344. ISSN: 1546-1696. (2024) (Mar. 2016). Finn, R. D., Clements, J. & Eddy, S. R. HMMER web server: interactive sequence similarity searching, en. Nucleic Acids Research 39, W29-W37. ISSN: 0305-1048, 1362-4962. (2024) (July 2011). Eddy, S. R. Accelerated Profile HMM Searches, en. PLoS Computational Biology 7 (ed Pearson, W. R.) el002195. ISSN: 1553-7358. (2024) (Oct. 2011). Yao, Z., Weinberg, Z. & Ruzzo, W. L. CMfinder— a covariance model based RNA motif finding algorithm, en. Bioinformatics 22, 445-452. ISSN: 1367-4811, 1367-4803. (2024) (Feb. 2006). . Lorenz, R. et al. ViennaRNA Package 2.0. en. Algorithms for Molecular Biology 6, 26. ISSN: 1748-7188.(Dec. 2011). . Fu, L., Niu, B., Zhu, Z., Wu, S. & Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data. en. Bioinformatics 28, 3150-3152. ISSN: 1367-4803, 1367-4811. (2024) (Dec. 2012). . Katoh, K. & Standley, D. M. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability, en. Molecular Biology and Evolution 30, 772-780. ISSN: 0737-4038, 1537-1719. (2024) (Apr. 2013).. Nguyen, L.-T., Schmidt, H. A., Von Haeseler, A. & Minh, B. Q. IQ-TREE: A Fast and Effective Stochastic Algorithm for Estimating Maximum-Likelihood Phylogenies, en. Molecular Biology and Evolution 32, 268-274. ISSN: 1537- 1719, 0737-4038. (2024) (Jan. 2015). . Matreyek, K. A., Stephany, J. J., Chiasson, M. A., Hasle, N. & Fowler, D. M. An improved platform for functional assessment of large protein libraries in mammalian cells. Nucleic Acids Research 48, el. ISSN: 0305-1048. (2024) (Jan. 2020). . Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis, eng. Nature Biotechnology 37, 224-226. ISSN: 1546-1696 (Mar. 2019). . Meeker, N. D., Hutchinson, S. A., Ho, L. & Trede, N. S. Method for Isolation of PCR-Ready Genomic DNA from Zebrafish Tissues. EN. BioTechnigues. Publisher: Taylor & Francis. (2024) (Nov. 2007). . Makarova KS, Wolf Yl, Iranzo J, Shmakov SA, Alkhnbashi OS, Brouns SJJ, Charpentier E, Cheng D, Haft DH, Horvath P, Moineau S, Mojica FJM, Scott D, Shah SA, Siksnys V, Terns MP, Venclovas C, White MF, Yakunin AF, Yan W, Zhang F, Garrett RA, Backofen R, van der Oost J, Barrangou R, Koonin EV. Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nat Rev Microbiol. 2020 Feb;18(2):67-83. doi: 10.1038 / s41579-019- 0299-x. Epub 2019 Dec 19. PMID: 31857715; PMCID: PMC8905525. . Shmakov S, Smargon A, Scott D, Cox D, Pyzocha N, Yan W, Abudayyeh OO, Gootenberg JS, Makarova KS, Wolf Yl, Severinov K, Zhang F, Koonin EV. Diversity and evolution of class 2 CRISPR-Cas systems. Nat Rev Microbiol. 2017 Mar;15(3):169-182. doi: 10.1038 / nrmicro.2016.184. Epub 2017 Jan 23. PMID: 28111461; PMCID: PMC5851899. . Christopher Vassallo, Christopher Doering, Megan L. Littlehale, Gabriella Teodoro, Michael T. Laub. Mapping the landscape of anti-phage defense mechanisms in the E. coli pangenome. bioRxiv 2022.05.12.491691; doi: 491691. Pennisi E. Like CRISPR, mystery gene editor began as a virus fighter.Science. 2020 Nov 20;370(6519):898-899. doi: 10.1126 / science.370.6519.898.PMID: 33214258.. Millman A, Melamed S, Leavitt A, Doron S, Bernheim A, Hbr J, Garb J, Bechon N, Brandis A, Lopatina A, Ofir G, Hochhauser D, Stokar-Avihail A, Tai N, Sharir S, Voichek M, Erez Z, Ferrer JLM, Dar D, Kacen A, Amitai G, Sorek R. An expanded arsenal of immune systems that protect bacteria from phages. Cell Host Microbe. 2022 Nov 9;30(ll):1556-1569.e5. doi: 10.1016 / j.chom.2022.09.017. Epub 2022 Oct 26. PMID: 36302390.SEQUENCESSEQUENCES FROM TABLE 7 (SEQ ID NOS: 1-44):SEQ ID NO: 1NRT-83JAAIYB010000035.1-23385- 24669MSWFDTTLSRLKGLFSRPVTRSTTGLDVPLDAHGRPQDVVRETVSTSGPLKPGHLRQLRRDARLLPKGIRRYTPGRKKWMEAAEARRLFSATLRTRNRNLRDLLPDEAQLARYGLPLWRTEEDVAAALGVSVGVLRHYSIHRPRERVRHYVTFAVPKRSGGVRLLHAPKRRLKALQRRLLALLASKLPVSPQAHGFVSGRSIKTGAMPHVGRRVVLKLDLKDFFPSVTFARVRGLLIALGYGHPVAATLAVLMTESERQPVELEGTVFHVPVGPRVCVQGAPTSPALCNAVLLRLDRRLAGLARRYGYTYTRYADDLTFSGDDVAALERVRALAARYVQEEGFEVNREK TRVQRRGGAQRVTGVTVNTTLGLSREERRRLRAMLHQEGRSGDGEARRAHLDGLLAYVKMLNPEQAE RLARRRKPRGTSEQ ID NO: 2NRT-49CP057918.1-4518082-4519042MDATRITLLALDLFGSPGWSADKEIQRLYALSNHTGRHYRRIILSKRHGGQRLVLAPDYLLKTVQ RNILKNVLSQFPLSSFATAYRPGCPIVSNAQPHCQQPQILKFDIENFFDSISWLQVWRVFRQAQLPRNVV TMLTWLCCYNDALPQGAPTSPAISNLVMCRFDERIGEWCQARGITYTRYCDDMTFSGHFNARLVKNKV CGLLAELGLNLNQRKSCLVAACKRQQVTGIVVNHKPQLAREARRALRQEVHLCQKYGVISHLSHRGELDS SGDLHAQATAYLYALQGRINWLLQINPEDEGFKQAREGVKRMLVAWSEQ ID NO: 3NRT-39CP034105.1-3807920-3808880MTSKNIYSLKNLGLPAMVSVDDFANESRLSSAKIRYLSATADLHYKVFSIPKVTGKSRLISQPSRDL KAIQAWILRNILDKLSSSSCAKGFERGSSILDNARPHIGSNYILTIDLENFFPSIAASKIYGVFSSIGYNKQLSV LLTNLCTYKGGLPQGAPTSPKLANLVCAKLDSRIQGYAGPKGIVYTRYADDMTFSANTALKITKTKQFIGTI ISDEDFKINSQKTVISGTKKQKKVTGLILSENSLGIGRVKYREVRVKLHYLFTDKLSNFSYMNGMLAYIYSVD KKSYNKLLSYIAKLKNKYPNSIAVNFINSKNKSEQ ID NO: 4NRT-42PZOB01000007.1-49956-50916MTSKNIYSLKNIGLPAMLSIDDFANQARLSAGKIRYLSATAESHYKVYPVPKVSGKSRLISQPSREL KAVQAWILRNILDKLSSSPCSKGFEVGTSILDNARPHIGSNYVLTIDLEDFFPNVSASKVYGVFSSIGYNKEL SILLTNLCTYGGGLPQGAPTSPKLANLVCAKLDARIQGYAGPRGIVYTRYADDMTFSSHTASKITKVKHFIG TIISDEGFKINHKKTAICGTKRQKKVTGLILSESSVGIGRVKHREIRAKLHHLFTGKSTEFSHINGLLSYTYSVD RRAYNKLYAYIGGLTQKYPLSNAIKEISPKIVSEQ ID NO: 5NRT-36AP018689.1-1528883-1529843MTSKNIYSLKNIGLPAMASIDDFANEARLSPGKIRYLSATAESHYRIYPIPKVGGKSRLISQPSRELK AVQAWILRNILDKLSSSPCSKGFEIGTSILDNARPHIGSNYVLTIDLEDFFPNVSASKVFGVFSSIGYNKELSIL LTNLCTYSGGLPQGAPTSPKLANLVCAKLDARIQGYTGPRGIVYTRYADDMTFSAHTASKIRKVKHFIGTIISDEGFKINHKKTAICGTKRQKKVTGLILSESSVGIGRVKHREIRAKLHHLFTGKSAEFSHVNGLLSYTYSVDR RSYNKLYSYIAGLSKKYPNSNAIIEISPKIVSEQ ID NO: 6NRT-45UFVY01000003.1-1181704-1182664MKSAEHLNTFRLRHLGLPVMNDLQDMSKATRISVETLRLLIYRADFRYKIYSIKKKDSQSIRTIYQP SRELKALQGWVLRNILDKLSSSPFSIGFEKHQSILNNATPHIGANFILNIDLENFFPSLSAKKVFGVFHSLGY NRTISSSLTKICCYKNLLPQGAPSSPKLANLICSKLDYRIQGYAGSRGLIYTRYADDLTLSAQSMKKVLKAKD FLFSIIPTENLVVNTKKTCISGPRSQKKVTGLVISQEKAGIGRLKHKELRAKIHHLFLGRSKDIEHVRGWLSFV LSVDSKSYRRLLVYIGKLEKKYGWNPLSKTKTSEQ ID NO: 7NRT-46AAWQPW010000019.1-56262-57222MDATRIPLLALDLFGSPGWSADKEIQRLYALSNHAGRHYRRIILSKRHGGQRLVLAPDYLLKTVQ RNILKNVLSQFPLSSFATAYRPGCPIVSNAQPHCQQPQILKLDIENFFDSISWLQVWRVFRQAQLPRNVV TMLTRLCCYNDALPQGAPTSPAISNFVMRRFDERIGEWCQARGITYTRYCDDMTFSGHFNARQVKNKV CGLLAELGLSLNQHKSCLIAASKRQQVTGIVVNDKPQLAREARRALRQEVHLCQKYGVISHLSRRGELDPS CDLHAQATAYLYALQGRINWLLQINPEDVAFKQARESVKRMLLAWSEQ ID NO: 8NRT-58JABELA010000134.1-8905-10351MTAKLESFVPAAPPQAVADTAPAASPNSVAKREAAKAAHDALITRWKAITEAGGADEWVQAQ LVSKGALADEVDVSSLKEKEKTAWKEKKKAEAVERRALKRQAHEAWKATHVNHLGVGIHWNEAGLPDK FDLEHREERARQNGLPTLDSAEDLAKALGLSVSKLRGFAFHRDVDTGSNYVTWSIPKRTGGERTITSPKRE LKEAQRWVLSNVVERLPVHGAAHGFVAGRSILTNALAHRGADVLVKVDLKDFFPSVHWRRVKGLLRKG GLKENTSTLLALMSTEAPRERMSFRGKTLHVAKGPRALPQGAPTSPGITNALCLRLDKRLSALSRKLGFTY TRYADDLTFSWTKAKAPKARRAQGAPVAVLLARVKDIVEAEGFTVHPDKTRVARKGSRQRVTGLVVND AKDGTPAARVPRDVVRRLRAAIHNRVKGKPGREGESLDQLKGMAAFIYMTDPEKGRMYLDQLAKLEAA QPAQASEQ ID NO: 9NRT-59JABUMS010000007.1-136401-137859MTARLESFVPAAAPQAVATPAPQAPAANAVAQREARRAAH EALLI RWKAIVEAGGAESWAQ GQLVSRGLAVADLDFSGASEKEKTAWKEKKKAEAAELRALGRQAHEAWKATHISHLGAGVHWKEDAG SDKFDVAHREERAKSNGLPELGTVDALAKALGLSVSKLRWFAFHREVDTGSHYISWGIPKRDGGTRTITS PKPELKQAQRWVLSNVVERLPVHGAAHGFVAGRSILTNALAHHGADVVVKVDLKDFFPTVTWRRVKGL LRKGGLQENTATLLSLMATEAPRETVQFRGKTLYVAKGPRALPQGAPTSPGITNALCLKLDKRLSALSKRL GFIYTRYADDLTFSWTKTKQPKAKRAQGPAVSVLLARVKEVVEDEGFHVHPNKTRVSRKGTRQQVTGLV VNKARDGVASARVPRDVVRRLRAAIHNRQKGKPGREGESLEQLKGMAAFVYMTDADKGRAFLKDVEA LEAREKETPKVASEQ ID NO: 10NRT-55CP003969.1-3185711-3187187MTAKIDATVPARAPVVLPLATPPAPFAAARGSAEERRLAHEERVARWKAIAEAGGIDAWVSAE LVARGAVATGDPKAMSDRERAQWKERKKVEARERRALRRLAWQAYLATHINHLGAGVHFRDREGPDA FDVPGLAERARSRGLPDLPSAGELARALGLTIPRLRWLAYHREVDAGTHYRRWLIPKRDGSARAISSPKRE LKRAQRWALRNLFEKLPVHAAAHGFLASRSIVTNAAAHAGADTIVKIDIKDFFPTITWRRVRGLLRKAGV AEGPATLVALLATEAPREVVQFRGQTLYVATGPRVLPQGAPTSPAITNAICLRLDRRVSGLARKLGFRYTR YADDLTFSFRAPHAPDAPGLAGAARPRAPVGALLRGVREILSAEGFRLHPGKTVVMRKGSRQKVTGLVV NGAGEAAPAARVPRERVRELRAAIRNRELGRPGKGETLAQLKGLAAFVYMTDPVRGRAFLGRIEALERN QPAPADGESGRSEQ ID NO: 11NRT-75CP002582.1-3047327-3048323MYNYEYTIRLCETLKSLSADDVYISKCCNYAEGLLDKELPVIFDPTHLKQILRLDDISLDEYHIFYIDK KNGGSREINAPSEELKKRQRWILKNILEKISISHNVHGFIKGKSIVSNARKHLNKEYVLNIDIKDFFPSVTKYS VEKIFRRMGYCNSVAQLLARVCCYRGGLPQGAPTSPYLANLAFDEVDQEIINVVRNRDITYTRYADDMTF SANYDLSTFKKEVYKSLGKYRFSPNIMKTHQMSGEKRKLVTGLIVDDKVKVCKKYKRKLRQEIYYCKKFGV TNHLRNCHSEKSINYKEYLYGKAYFIKMVEEIVGEKFLADLDSIDWYSEQ ID NO: 12NRT-56JAAIYA010000005.1-274472-275924MTARLESFVPAASPQAVPTPAPSAPAANAVSQREVRRAAHEALLLRAKAIEESGGADSWVQG QLVSKGLAVGDLDFSKASEKEKTAWKEKKKAEATERRALERQAHEAWKATHIGHLGAGVHWEEDAGA DKFDVADREERAKANGLPELGSAEALAKALGLSVSKLRWFAFHREVDTATHYISWKIPKRDGGSRTITSPK PELKEAQRWVLSNVVERLPVHGAAHGFIAGRSILTNALAHASADVVVKVDLKDFFPSVTWRRVKGLLRK GGLPENTSTLLALMATEAPREAVQFRGKLLHVAKGPRALPQGAPTSPGITNALCLRLDKRLSALAKRLGFT YTRYADDLTFSWTKAKQPKAKRSQRPPVAVLLSRVQTVVEGEGFRLHPDKTRVARKGTRQRVTGLVVNE AKAGQPAARVPRDVVRRLRAAIHNREKGKPAREGESLDQLKGMAAFIHMADPAKGRAFLEQLSKLEAK EPAPAASEQ ID NO: 13NRT-51PDIH01000062.1-144132-145095MDATRITLLALNLFGSPGWSADKEIQRLYALSNHVGRHYRRIILSKRHGGQRMLLAPDYLLKTVQ RNILKNVLSRFPLSPFATAYRPGCPIVSNAQPHCQQSQILKLDIENFFDSISWLQVWRVFSQAQLPRNVVT ILTWLCCYNDALPQGAPTSPAISNLVMRRFDECIGAWCQTRGIIYTRYCDDMTFSGHFNARQVKNKVCG LLAELLGLSLNQRKSCLVAACKRQQVTGIVVNHKPQLAREARRALRQEVHLCQKYGVISHLSHRGELDPS GDLHAQATAYLYALQGRINWLLQINPEDEAFKQARESVKRMLVAWSEQ ID NO: 14NRT-57JAAIYB010000034.1-24820-26278MTARLDPFVPAASPQAVPTPEPTAPAPDAAAKREARRLAHEALLVRAKAIDEAGGADDWVQAQLVSKGLAVEDLDFSSASEKDKKAWKEKKKAEATERRALKRQAHEAWKATHVGHLGAGVHWVEDRPADAFDVPHREERARANGLTELDSAEALAKALGLSVSKLRWFAFHREVDTATHYVSWTIPKRDGSQRTITSP KPELKAAQRWVLSNVVERLPVHGAAHGFVAGRSILTNALAHQGADVVVKVDLKDFFPSVTWRRVKGLL RKGGLPEGTSTLLSLLSTEAPREAVQFRGKLLHVAKGPRALPQGAPTSPGITNALCLKLDKRLSALAKRLGF TYTRYADDLTFSWTKAKQPKARRTQRPPVAMLLSRVQEVVEAEGFRVHPDKTRVARKGTRQRVTGLVV NAAGKDAPAARVPRDVVRQLRAAIHNRKKGKPAREGESLEQLKGMAAFIHMTDPAKGRAFLAQLGELE STASAAPQAESEQ ID NO: 15NRT-82DQHL01000007.1-131724-132741MTTLRHHLLTSPLVAEHAFSLLSSSSSKKVTKSDVEKWLTISLATVMTSGPKMYKVYTIPKKNGGK RTIAHPSKVLKVFQKGLVEFLENKLPVHDAAYAYRKNISIKDNALQHAKNAYFLRMDLADFFNSIDVNLFE KQIEKHKIELATADKALLRRSAFWSPKKMSSGNLILSVGAPSSPMISNFIMFSFDETITSVCDPLGIKYTRYA DDLFFSTGHKNLLFKIPSLIQYVLTSEFGKKIAINEVKTSFSSKAHNRHIAGVTITNNDTLSLGRERKRIISSMI HKYSLGNSNDIELAKLQGLLSFSKHIEPLFIDRMKEKYSISLVENIITGRWRKSEQ ID NO: 16NRT-28MGYG000267465_7-164328-165288MDATRITLLALDLFGSPGWSADKEIQRLYALSNHAARHYRRIILSKRHGGQRLVLAPDYLLKTVQ RNILKNVLSQFPLSSFATAYRPGCPIVSNAQPHCQQPQILKLDIENFFDSISWLQVWRVFCQAQLPRNVV TMLTWLCCYNDALPQGAPTSPAISNLVMRRFDERIGEWCQARGITYTRYCDDMTFSGHFNARLVKNKV CGLLAELGLNLNQRKSCLVAACKRQQVTGIVVNHKPQLAREVRRALRQEVHLCQKYGVISHLSRRGELDP SGDLHAQATAYLYALQGRINWLLQINPEDEGFKQAREGVKRMLVAWSEQ ID NO: 17NRT-31MGYG000097028_ll-91674-92607MELLSYISKGLFLSEEEAKRYIVTIPRRYKIYPIAKRNGNGYRLIAQPARQVKALQKVVINHMLANL KVHENAFAYEADKSIRSNALMHCNNDYLLKMDFKNFFMSIKPHNLLSVLKDYGVELTDTDIFVLKNLFFW KLRRNSPLRLSVGAPSSPIISNTVMYFFDEEISKRCKDLGITYSRYADDLTFSTKVKGALFDVPKLVREVLNV VNLRNIKINHEKTVFTSRKFNRHVTGVTITTDGYLSLGRDRKRLLRSKIHYYVCGVLSEKEILTLKGELGYAKF IEHKFFASMIKKYGDDVISEISKYEISEQ ID NO: 18NRT-47AAXELM010000020.1-29552-30512MDATRITLLALDLFGSPGWSADKEIQRLYALSNHAARHYRRIILSKRHGGQRVVLAPDYLLKTVQ RN I LKN VLSQFPLSPFATAYRQG RPI VSN AQPHCQQPQI LKLDI EN FFDSISWLQVWRVFSQTQLPRN VV TMLTWLCCYNDALPQGAPTSPAISNLVMRRFDECIGEWCQARGITYTRYCDDMTFSGHFNARQVKNK VCGLLAELGLSLNQRKSCLIAACKRQQVTGIVVNHKPQLAREARRALRQEVHLCQKYGVISHLSHRGELD PSGDLHAQATAYLYALQGRINWLLQINPEDEAFKQARVSVKRMLVAWSEQ ID NO: 19NRT-103 UFZI01000001.1-147466-148618MDATRITLLALDLFGSPGWSADKEIQRLYALSNHAGRHYRRIILSKRHGGQRLVLAPDYLLKTVQ RNILKNVLSQFPLSPFATAYRPGCPIVSNAQPHCQQQQILKLDIENFFDSISWLQVWRVFRQAQLPRNVV TMLTWICCYNDALPQGAPTSPAISNLVMRCFDERIGEWCQARGITYTRYCDDMTFSGHFNARQVKNKV CGLLAELGLSLNQRKSCLIAACKRQQVTGIVVNHKPQLAREARRALRQEVHLCQKYGVISHLSHRGELDPS GDLHAQATAYLYALQGRINWLLQINPEDEAFQQARESVKRCWLQGKKSVRQTFLPDRLGENYCNCAAIS GQRASKSSVGRYLNSLRTKRDSIPSQKARIWLASRVSSVASEQ ID NO: 20NRT-52UFZI01000001.1-147466-148618MDATRITLLALDLFGSPGWSADKEIQRLYALSNHAGRHYRRIILSKRHGGQRLVLAPDYLLKTVQ RNILKNVLSQFPLSPFATAYRPGCPIVSNAQPHCQQQQILKLDIENFFDSISWLQVWRVFRQAQLPRNVV TMLTWICCYNDALPQGAPTSPAISNLVMRCFDERIGEWCQARGITYTRYCDDMTFSGHFNARQVKNKV CGLLAELGLSLNQRKSCLIAACKRQQVTGIVVNHKPQLAREARRALRQEVHLCQKYGVISHLSHRGELDPS GDLHAQATAYLYALQGRINWLLQINPEDEAFQQARESVKRCWLQGKKSVRQTFLPDRLGENYCNCAAIS GQRASKSSVGRYLNSLRTKRDSIPSQKARIWLASRVSSVASEQ ID NO: 21NRT-89JAELVU010000001.1-582727-584041MRLLILIALGFLFWRLWKLWRQLRLRHQARHPVRDIARLVDDADSDARGDTVERIDMSGGPLK PGHRRRALRDRRLLPKPKRDWSDPKPPKVMSLDAATRLFAGTLRTRDRHARDLDTDAEQLERHGLPPW RSEAEVAAALGIEEKRLRHYAIHRQRERVVHYIAFAIAKRNGGERVILAPKKELKALQRKLNKLLVDKLPVS DAAHGFRPGRSIASNAAPHAGKAVVLKLDLENFFPSLHVGRVRGLLIALGYAYPVAAGLAALMTESERQP VVVDGETFHVPTSSRHAVQGAPTSPGLANALALRLDRRLTGLARAHGFVYTRYADDLAFSSDDVATARQ LLKRAESVIRAEGFRVNRAKTRVMTQASAQRIAGVTVNAAPGWSRAQRRKLRAELHRARTSGATDPSL WQKLRG KLAFVRM LN RAQG ERM ARADSSEQ ID NO: 22NRT-5 MGYG000192862_l-4890622-4891558MDILQHISDLLLTKKSEIISFSLTAPYRYKIYKIAKRNSDKKRTIAHPSKELKFIQREITEYLTDKLPVHE CAFAYKKGSSIKTNAQVHLHTKYLLKMDFENFFPSITPRLFFSKLRLANIDLTADDKVLLENILFFKSKRNSNL RLSIGAPSSPLISNFVMYFWDIEVQEICSKIGVNYTRYADDLTFSTNNKDVLFDIPDMLENVLPKYSLGRIRI NHEKTVFSSKGHNRHVTGITLTNDNKLSIGRERKRKISAMIHHFINGKLSTDECNKLVGLLAFAKNIEPSFY KSMVIKYGSDNIYKLQKQKDKSEQ ID NO: 23NRT-83JAAIYB010000035.1-23385-24669 cgggagaggccagggctcgcagatgagccatgagtaccgcggtgcttcgccgcgggggtgttctgtccccttctcttcgc cagggtcccagcgtacgcaacgcagggagccccgggtccaacgcctcgcaggtcgtcccctggcctctcccgSEQ ID NO: 24NRT-49CP057918.1-4518082-4519042 cgccagcagtggcaatagcgtttccggccttttgtgccgggagggtcggcgagtcgccgacttaacgccagtagtttgtct atatacccaaagccgcttcattgtacttaagtacgctttgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcgggatggt ttgatggtattgccgctgttggcgSEQ ID NO: 25NRT-39CP034105.1-3807920-3808880 tacggtatacacccttagcgaatgaacataatgttcattggatagcgtttcgctatcctgcacataatctgattttatgccgt atgaaaatgtgcatcaccagaataccgtaSEQ ID NO: 26NRT-42PZQB01000007.1-49956-50916 cacacccttagcgaatgagctaacttagttcattggatagcgtttcgctatcctgcatacaatctgattcaatgccgcatga aaatgtgcagagccagaatacagtcgtttctggaactgcaccttttcatccgcgacttaagacgtaagggtgtgSEQ ID NO: 27NRT-36AP018689.1-1528883-1529843 cgcacccttagcgaatgagcttacttagttcattggatagcgtttcgctatcctgcatacaatctgattcaatgccgcatga aaatgtgcagagccagaatacagtagtttctggaactgcacattttcatccgcgacttaagacgtaagggtgtgSEQ ID NO: 28NRT-45UFVY01000003.1-1181704-1182664 tgcgcacccttagcgagaggcttaccattcgtgggaacctctggatgctgattcggcatcctgcatgtaatctgagttactg tctgctttccttgttggaacggagagcatcgcttgatgctctccgagccaaccaggaaatccgtcttttttgacgtaagggtgcgcaSEQ ID NO: 29NRT-46AAWQPW010000019.1-56262-57222 gccagcagtggcaatagcgtttccggcctttgtgccgggagggtcggcgagtcgccgacttaacgccagtagtatgtcca tatacccaaagtcgcttcattgtacctgagtacgcttcgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcgggatggtt taatggtattgccgctgttggcSEQ ID NO: 30NRT-58JABELA010000134.1-8905-10351 gcgagtggtcgagagaggtctggaccgcatcagccttaacgcctcgagcgtaggaacggcgttgcgccgttctggttga aatgctggacactctccgcaaggtagcctgttcttggctctctccctcccgagcactaccgtcggggtgggaagcggaaccaacgac gcagccgccgttttcccaccccgacggtagtgctcgggaggggagagccggtgaggctaccgtgccccaggtgagctggtggtgccttcctggcctccctcga ccgctcgcSEQ ID NO: 31NRT-59JABUMS010000007.1-136401-137859 agcctgagcgcctcgagcgcgggagcggcgttgcgccgctccggttggaatgcaggacactctccgcaaggtagcctgt tcttggctctctccctcctaggcactacggcctgcgggggcagctgagccaacgacgcgatcgccgtttgcccccgcaggccgtagt gcctaggaggggagagccggtgaggctSEQ ID NO: 32NRT-55CP003969.1-3185711-3187187 tgttgccaggagaggttctgaacgcaccagcctccgcgccttgagcgcaggagcggcgtcgcgccgctctggaggaacc gccagtcactctccgcaaggtagcctgttcttggttctcttcggctagaccagcgctacggcgccgggagagctgagacaacctcgc aaccgccgtttctcccggcgccgtagcgctggtctagccggagaaccggtgaggctaccgtgccccatgggtagaaggtgatgcgct cggggcctctcttggcgacaSEQ ID NO: 33NRT-75CP002582.1-3047327-3048323 aaaagagcaactagattgaggcgattcgcctccttggaaaagggtactaagtttctgtcgcacaccaatttataagcttat aaattggtgtgcgacagaaatgaaataaatagtagttgctcttttSEQ ID NO: 34NRT-56JAAIYA010000005.1-274472-275924 gcgcgagcagccgagagaggtctggagtgcatcagcctgagcgcctcgagcgcgggagcggtgttgcgccgctccggtt ggaatgcaggacactctccgcaaggtagcctgttcttggctctctccctcctaggcactacggcccgggcgggtagctgagccaacg acgcgaacgccgtttacccgcccgggccgtagtgcctaggaggggagagccggtgaggctaccgtgccccaggtaagatggtggt gctttcccggcctccctcga ctgctcgcgcSEQ ID NO: 35NRT-51PDIH01000062.1-144132-145095 cgccagcagtggcaatagcgtttccggccttttgtgccgggagggtcggcgagtcgccgacttaacgccagtagtttgtcc atatacccaaagccgcttcattgtacttaagtacgcttcgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcggatggtt tgatggtattgccgctgttggcgSEQ ID NO: 36NRT-57JAAIYB010000034.1-24820-26278 gcgcgagcagccgagagaggtccggagtgcatcagcctgagcgcctcgagcgcgggagcggcgttgcgccgctccggt tggaatgcaggacactctccgcaaggtagcctgttcttggctctctccctcctaggcactacggccagggtgggtagcggagccaac gacgcgaccgccgtttacccaccccggccgtagtgcctaggaggggagagccggtgaggctaccgtgccccaggtaagatggtggt gctttcccggcctccctcga ctgctcgcgcSEQ ID NO: 37NRT-82DQHL01000007.1-131724-132741 gctctttagcttatggctctgtattaaaagccttgtcgggcgtttcgccagacaccaacttattgaacgacattgtggttgc gaaagtctcgcgacccaagctcgttaactcacttaggtcgcgagactttcgcttcctctagtaaagagtSEQ ID NO: 38NRT-28MGYG000267465_7-164328-165288 tactgtgcgcagcgtgatgcggtttaagatatcgtgttaatctgctttcgccagcagtggcaatagcgtttccggccttttgt gccgggagggtcggcgagtcgccgacttaacgccagtagtttgtccagatacccaaagtcgctccattgtacttaagtacgcttcgc gta cgtcgcgctga cgcgctcagtaSEQ ID NO: 39NRT-31MGYG000097028_ll-91674-92607 taaatgtcaatggtttaggtggttgctggcagccagtacgcttacttattgaaaatctcgcatcttttgctatcctaacccta cctttacgcgcgggataacaggtttatccttgtcgggcgtttcgccagacactaatttattgtggacatttaSEQ ID NO: 40NRT-47AAXELM010000020.1-29552-30512 cgccagcagtggcaatagcgtttccggccttttgtgccgggagggtcggcgagtcgccgacttaacgccagtagtttgtcc agatacccaaagtcgctccattgtacttaagtacgcttcgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcgggatag tttgagggtattgccgctgttggSEQ ID NO: 41NRT-103 UFZI01000001.1-147466-148618 cgccagcagtggcaatagcgtttccggccttttgtgccgggagggtcggcgagtcgctgacttaacgccagtagtatgtcc atatacccaaagtcgcttcattgtacctgagtacgcttcgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcgggatgg tttaatggtattgccgctgttggcgSEQ ID NO: 42NRT-52UFZI01000001.1-147466-148618 cgccagcagtggcaatagcgtttccggccttttgtgccgggagggtcggcgagtcgctgacttaacgccagtagtatgtcc atatacccaaagtcgcttcattgtacctgagtacgcttcgcgtacgtcgcgctgacgcgctcagtacagttacgcgccttcgggatgg tttaatggtattgccgctgttggcgSEQ ID NO: 43NRT-89JAELVU010000001.1-582727-584041 cgcacggccttcgcgatagtcgcgtccgcccggaggggttgcagctagcgacgatgctgcggtgtttcgccgcggtagtttcttgcca ccatca ca tcgcca cggtcccca cgta cgga a cgtggggaggccgtgagSEQ ID NO: 44NRT-5 MGYG000192862_l-4890622-4891558 tcactctttagcgttaggctttgatttatagccttgtcgagcgtttcgccagacactaacttattgagtacttttagggttgcg ctagaaagttttctaccgatcctagaagtctctaggatcggtagaaaactttctagcgcctcctctagtaaagagtaaRetron msr-msd sequences used in Example 1:Efel (SEQ ID NO: 24):CGCCAGCAGTGGCAATAGCGTTTCCGGCCTTTTGTGCCGGGAGGGTCGGCGAGTCGCCGACTTAACGCCAGTAGTTTGTCTATATACCCAAAGCCGCTTCATTGTACTTAAGTACGCTTTGCGTACGTCG CGCTGACGCGCTCAGTACAGTTACGCGCCTTCGGGATGGTTTGATGGTATTGCCGCTGTTGGCGEco8 (SEQ ID NO: 29):GCCAGCAGTG GCAATAGCGT TTCCGGCCTT TGTGCCGGGA GGGTCGGCGA GTCGCCGACTTAACGCCAGT AGTATGTCCA TATACCCAAA GTCGCTTCAT TGTACCTGAG TACGCTTCGCGTACGTCGCG CTGACGCGCT CAGTACAGTT ACGCGCCTTC GGGATGGTTT AATGGTATTGCCGCTGTTGG CVapl (SEQ ID NO: 27):CGCACCCTTA GCGAATGAGC TTACTTAGTT CATTGGATAG CGTTTCGCTA TCCTGCATACAATCTGATTC AATGCCGCAT GAAAATGTGC AGAGCCAGAA TACAGTAGTT TCTGGAACTGCACATTTTCA TCCGCGACTT AAGACGTAAG GGTGTGCexl (SEQ ID NO: 114):AGTGGTCGAG AGAGGTCTGG ACCGCATCAG CCTTAACGCC TCGAGCGTAG GAACGGCGTTGCGCCGTTCT GGTTGAAATG CTGGACACTC TCCGCAAGGT AGCCTGTTCT TGGCTCTCTCCCTCCCGAGC ACTACCGTCG GGGTGGGAAG CGGAACCAAC GACGCAGCCG CCGTTTTCCCACCCCGACGG TAGTGCTCGG GAGGGGAGAG 211 CCGGTGAGGC TACCGTGCCC CAGGTGAGCTGGTGGTGCCT TCCTGGCCTC CCTCGACCGC TCGCVrol (SEQ ID NO: 26):CACACCCTTA GCGAATGAGC TAACTTAGTT CATTGGATAG CGTTTCGCTA TCCTGCATACAATCTGATTC AATGCCGCAT GAAAATGTGC AGAGCCAGAA TACAGTCGTT TCTGGAACTGCACCTTTTCA TCCGCGACTT AAGACGTAAG GGTGTGRFP reporter (SEQ ID NO: 115):TCGGGGATGC CCTGGGTGTG GTTGATGAAG GCTTTGCTGC CGTACATGAA GCTGGTAGCCAGGATGTCGA AGGCGAAGGG GGenomic loci in HEK293T cellsEMX1 (SEQ ID NO: 116):TGGCCTGCTT CGTGGCAATG CGCCACCGGT TGATGTGATG GGAGCCCTTC GCTACATGCTTTCTTCTGCT CGGACTCAGG CCCTTCCTCC TCCAGCTTCT GCCGTTTGTAAAVS1 (SEQ ID NO: 117):GGAGACTAGG AAGGAGGAGG CCTAAGGATG GGGCTTTTCT GTCACCAATC GCTACATGCTCTGTCCCTAG TGGCCCCACT GTGGGGTGGA GGGGACAGAT AAAAGTACCCHBB (SEQ ID NO: 118):CTGCCCAGGG CCTCACCACC AACTTCATCC ACGTTCACCT TGCCCCACAG GCTACATGCTGGCAGTAACG GCAGACTTCT CCTCAGGAGT CAGATGCACC ATGGTGTCTGCFTR (SEQ ID NO: 119):TATAAAAAGA TTCCATAGAA CATAAATCTC CAGAAAAAAC ATCGCCGAAG GCTACATGCTGGCATTAATG AGTTTAGGAT TTTTCTTTGA AGCCAGCTCT CTATCCCATTF9 (SEQ ID NO: 120):GTGAACATGA TCATGGCAGA ATCACCAGGC CTCATCACCA TCTGCCTTTT GCTACATGCTAGGATATCTA CTCAGTGCTG AATGTACAGG TTTGTTTCCT TTTTTAAAATBRD8 (SEQ ID NO: 121):GGAGACTAGG AAGGAGGAGG CCTAAGGATG GGGCTTTTCT GTCACCAATC GCTACATGCTCTGTCCCTAG TGGCCCCACT GTGGGGTGGA GGGGACAGAT AAAAGTACCCGenomic loci in U2OSTOMM70A (SEQ ID NO: 122):AGTGTCTTTA GGGTTCAGTT GAAGAGGGGG TAAACTTTTA AAAAGAGGGT CAGTCTGCTT TCCCCCTGTT TTATGTAATC CCAGCAGCAT TTACATACTC ATGAAGGACC ATGTGGTCAC GGCCGCCACC TAATGTTGGT GGTTTTAATC CGTATTTCTT TGCAACTTCT GTCTGGGCATGGGCGGCATC GCAAAGTGAARAB11 (SEQ ID NO: 123):TAGAGTGCGA GAGCCCATGG CCTCACCTTT AAAGAGGTAG TCGTACTCGT CGTCCCGTGT GCCGCCGCCA CCTGTAATCC CAGCAGCATT TACATACTCA TGAAGGACCA TGTGGTCACG CATTGCGCGG CCGAGGAGCG AAAGGGCGGG AGCAGCAGTG GTATCTGTGG GACCAGGGGGCGTCGCTGCA GGGGTAACCCCorrecting Kif6 mutations (SEQ ID NO: 124):TTACCGCCCA AACTGTCTCT GAGTACAGAG GTCATCATTG AATTCCTATA GGGGATGTGG GACCTATCCT TCTCTGACAG GGCTATGATG ACCTAAACCA

Claims

CLAIMS1. A retron editor system comprising: a) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence; b) a nucleotide sequence encoding a guide RNA (gRNA); and c) a nucleotide sequence encoding a nuclease.

2. The retron editor system of claim 1, wherein the system is provided as a single cassette.

3. The retron editor system of claim 1, wherein at least one of components a), b), or c) are provided as separate cassettes.

4. The retron editor system of any one of claims 1-3, wherein the retron is non-naturally occurring.

5. The retron editor system of claim 4, wherein the msd sequence of the retron has been modified from a naturally occurring msd.

6. The retron editor system of claim 5, wherein a donor nucleic acid insertion sequence is inserted within the msd sequence.

7. The retron editor system of any one of claims 1-6, wherein the msd sequence comprises homology arms.

8. The retron editor system of claim 7, wherein the homology arms comprise at least 80% similarity to genetic locus of interest on either side of a nuclease cleavage site.

9. The retron editor system of any one of claims 1-8, wherein the nuclease is Cas9, Casl2a, Casl2f, TnpB, or a combination thereof.

10. The retron editor system of any one of claims 1-9, where a palindromic repeat sequence of the non-coding region is located within the ncRNA.

11. The retron editor system of any one of claims 1-10, wherein the RT is capable of selfpriming reverse transcription from the ncRNA.

12. The retron editor system of any one of claims 1-11, wherein the gRNA is single guide RNA (sgRNA).

13. The retron editor system of claim 2 or 3, wherein the cassette or cassettes further comprises at least one promoter.

14. The retron editor system of claim 13, wherein the promoter is capable of driving sgRNA transcription.

15. The retron editor system of claim 14, wherein the promoter is a U6 promoter.

16. The retron editor system of any one of claims 13-15, wherein the promoter is capable of driving msr-msd transcription.

17. The retron editor system of claim 16, wherein the promoter is an Hl promoter.

18. The retron editor system of any one of claims 15-17, wherein the promoter drives nuclease-RT fusion.

19. The retron editor system of claim 18, wherein the promoter is a CMV protomer.

20. The retron editor system of claim 2, wherein the nucleotide sequence encoding the nuclease is located at or near a 3' end of the cassette.

21. The retron editor system of claim 2, wherein the nucleotide sequence encoding the RT is located at or near a 5' end of the cassette.

22. The retron editor system of claim 2, wherein the nucleotide sequence encoding the nuclease and RT are fused by a flexible, rigid, or variable linker, thereby forming a nuclease- RT fusion nucleic acid.

23. The retron editor system of claim 22, wherein the nuclease-RT fusion nucleic acid is at or near a 3' end of the cassette.

24. The retron editor system of any one of claims 1-23, wherein the retron editor system further comprises a nucleotide sequence encoding a nuclear localization signal (NLS).

25. The retron editor system of claim 24, wherein the NLS is in a single cassette with a nucleotide sequence encoding both the nuclease and the retron, and further wherein the nucleotide sequence encoding the nuclease and retron are fused.

26. The retron editor system of claim 25, wherein the NLS is 5' of the nuclease-RT fusion nucleotide sequence.

27. The retron editor system of claim 9, wherein the nuclease is Casl2a and the guide RNA is crRNA.

28. The retron editor system of any one of claims 1-27, wherein the retron further comprises a second msr sequence.

29. The retron editor system of any one of claims 1-28, where the system is packaged for delivery.

30. The retron editor system of claim 29, wherein the package is a lipid nanoparticle.

31. The retron editor system of claim 30, wherein the package is an AAV.

32. The retron editor of any one of claims 1-31, wherein the nuclease or reverse transcriptase is fused to a protein that modifies DNA repair pathway choice.

33. The retron editor of claim 32, wherein the fusion protein is a DNA repair protein.

34. The retron editor of claim 32, wherein the fusion protein is a helicase.

35. The retron editor of claim 33, wherein the DNA repair protein comprises CtlP-dnRNF168, hRad51, DN1S, hGeml / 110, and / or Rep-X helicase.

36. The retron editor of any one of claims 1-35, wherein the nuclease in the retron editor is replaced with at least one nickase.

37. The retron editor of claim 36, wherein the retron editor comprises two nickases.

38. The retron editor of claim 36 or 37, wherein the nickase comprises Cas9(D10A).

39. A vector comprising the system of claim 1.

40. The vector of claim 39, wherein the system comprises one or more cassettes.

41. The vector of claim 40, wherein the vector further comprises a promoter, wherein said promoter is not within the one or more cassettes.

42. The vector of claim 41, wherein the promoter is operably linked to the nucleotide sequence encoding the retron.

43. The vector of claim 41, wherein the promoter is operably linked to the gRNA.

44. The vector of claim 41, wherein the promoter is operably linked to the nucleotide sequence encoding the nuclease.

45. The vector of any one of claims 39-44, wherein the promoter is operably linked to a nuclease-RT fusion nucleic acid.

46. The vector of claim 47, wherein the promoter is an RNA polymerase II promoter.

47. The vector of claim 45, wherein the promoter is an RNA polymerase III promoter.

48. A method of using a retron editor system to edit target nucleic acid, the method comprising: a) providing a retron editor system, wherein said system comprises: i) a nucleotide sequence encoding a guide RNA (gRNA); ii) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence, wherein said msd sequence comprises a donor nucleic acid; and iii) a nucleotide sequence encoding a nuclease; b) expressing a product from the retron editor system; and c) placing the retron editor system under conditions such that gene editing of target nucleic acid takes place.

49. The method of claim 48, wherein the system is provided as a single cassette.

50. The method of claim 48, wherein at least one of components i), ii), or iii) are provided as separate cassettes.

51. The method of any one of claims 48-50, wherein the retron editor system is in one or more cassettes.

52. The method of claim 51, wherein the one or more cassettes are within one or more vectors.

53. The method of any one of claims 48-52, wherein appropriate conditions for expression of a retron editor system are provided.

54. The method of any one of claims 48-53, wherein the target nucleic acid is RNA.

55. The method of any one of claims 48-54, wherein the RT carries out reverse transcription of the donor nucleic acid.

56. The method of claim 55, wherein reverse transcription of the donor nucleic acid results in a multicopy single-stranded DNA (msDNA) molecule that comprises RNA and DNA.

57. The method of claim 50, wherein different systems are provided on different cassettes.

58. The method of claim 57, wherein different donor nucleic acids are provided in different cassettes.

59. The method of any one of claims 48-58, wherein the method is carried out in the presence of one or more compositions which enhances editing of the target nucleic acid.

60. The method of claim 59, wherein said one or more composition which enhance editing comprises a small molecule.

61. The method of claim 59 or 60, wherein the composition is an inhibitor of DNA- dependent protein-kinase catalytic subunit (DNAPKcs) or CDC7.

62. The method of any one of claims 59-61, wherein the composition comprises AZD7648.

63. The method of any one of claims 48-62 wherein the retron editor is delivered by microinjection or electroporation.

64. The method of claim 63, wherein the retron editor is delivered in the form of nucleic acid or protein.

65. A method of modifying one or more target nucleic acids of interest at one or more target loci in a host cell, the method comprising: a. Transforming the host cell with a vector of any one of claims 39-47; b. Culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing a retron editor system from the vector; c. Providing conditions suitable for the nuclease of the retron editor system to cut at or near the target loci; andd. Providing conditions for the donor nucleic acid insertion sequence to recombine with the one or more target nucleic acid sequences to inset, delete, and / or substitute one or more bases of the sequence of the one or more target nucleic acid sequences to induce one or more sequence modifications at the one or more target loci.

66. The method of claim 65, wherein a reverse transcript is produced in step b).

67. The method of claim 66, wherein the reverse transcript comprises msd, msr, and reverse transcriptase.

68. The method of any one of claims 65-67, wherein the retron transcript self-primes reverse transcription by a reverse transcriptase expressed by the host cells or the transformed progeny of the host cells.

69. The method of claim 68, wherein at least a portion of the retron transcript is reverse transcribed to produce multicopy single-stranded DNA (msDNA) molecule having one or more donor DNA sequences.

70. The method of claim 69, wherein the donor DNA sequences are homologous to the one or more target loci, but comprise one or more sequence modifications compared to the one or more target nucleic acids.

71. The method of any one of claims 65-70, wherein the target loci is within a genome of the cell.

72. The method of any one of claims 65-71, where the host cell is a eukaryotic cell or plant cell.

73. The method of any one of claims 65-72, where the host cell is part of a population of host cells.

74. The method of any one of claims 65-73, wherein the method is carried out in the presence of one or more compositions which enhances editing of the target nucleic acid.

75. The method of claim 74, wherein said one or more composition which enhance editing comprises a small molecule.

76. The method of claim 74 or 75, wherein the composition is an inhibitor of DNA- dependent protein-kinase catalytic subunit (DNAPKcs) or CDC7.

77. The method of any one of claims 74-76, wherein the composition comprises AZD7648 or Cas9-CtlP-dnRNF168.

78. A method of treating or preventing a disease in a subject, the method comprising administering to the subject the retron editor system of any one of claims 1-38.

79. The method of claim 78, wherein the method is carried out in the presence of one or more compositions which enhances editing of target nucleic acid of the subject.

80. A reverse transcriptase comprising 80% or more identity to any one of SEQ ID NOS: 1-22.

81. A reverse transcriptase comprising 90% or more identity to any one of SEQ ID NOS: 1-22.

82. A reverse transcriptase comprising 95% or more identity to any one of SEQ ID NOS: 1-22.

83. A reverse transcriptase comprising any one of SEQ ID NOS: 1-22.

84. A nucleic acid sequence comprising 80% or more identity to any one of SEQ ID NOS: 23- 44.

85. A nucleic acid sequence comprising 90% or more identity to any one of SEQ ID NOS: 23- 44.

86. A nucleic acid sequence comprising 95% or more identity to any one of SEQ ID NOS: 23- 44.

87. A nucleic acid sequence comprising any one of SEQ ID NOS: 22-44.

88. A cassette comprising a nucleic acid which encodes a protein with 80% or more identity to any one of SEQ ID NOS: 1-22.

89. A cassette comprising a nucleic acid which encodes a protein with 90% or more identity to any one of SEQ ID NOS: 1-22.

90. A cassette comprising a nucleic acid which encodes a protein with 95% or more identity to any one of SEQ ID NOS: 1-22.

91. A cassette comprising a nucleic acid which encodes a protein comprising any one of SEQ ID NOS: 1-22.

92. The cassette of any one of claims 89-91, wherein the cassette further comprises all or part of an msd and an msr sequence.

93. The cassette of claim 92, wherein the msr and / or msd sequence is found, all or in part, in any one of SEQ ID NOS: 23-44.

94. The cassette of claim 93, wherein at least 10 consecutive nucleotides of any one of SEQ ID NOS: 23-44 are contained in the cassette.

95. The cassette of claim 94, wherein at least 20 consecutive nucleotides of any one of SEQ ID NOS: 23-44 are contained in the cassette.

96. The cassette of any one of claims 88-95, wherein the msd sequence comprises a donor nucleic acid sequence.

97. The cassette of claim 96, wherein at least a portion of the msd sequence is replaced with donor nucleic acid.

98. A method of screening for functional retron editors, the method comprising: a) providing a potential retron editor system;b) transforming a cell with the potential retron editor system, wherein said cell has been modified to express a signal upon successful transformation using the geneediting retron system; and c) detecting the presence of the signal.

99. The method of claim 98, wherein the signal is a fluorescent signal.

100. A retron editor system comprising: a) a first zinc finger nuclease (ZFN) or a nucleic acid encoding a first ZFN that binds a first area of a target region, wherein the first ZFN comprises a cleavage domain and a ZFN protein; b) a second ZFN or a nucleic acid encoding a second ZFN that binds a second area of a target region, wherein the second ZFN comprises a cleavage domain and a second ZFN protein; wherein the first and second ZFN are capable of dimerization and cleavage of the target region; and c) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence.

101. A retron editor system comprising: a) a first transcriptional activator-like effect nuclease (TALEN) or a nucleic acid encoding a first TALEN, wherein the first TALEN comprises a target region binding site and a nuclease; b) a second TALEN or a nucleic acid encoding a second TALEN, wherein the second TALEN comprises a target region binding site and a nuclease; and c) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence.

102. A retron editor system comprising: a) a retron, wherein said retron comprises non-coding RNA (ncRNA) and a nucleotide sequence encoding a reverse transcriptase (RT) further wherein the ncRNA comprises an msd sequence and a msr sequence; b) a nucleotide sequence encoding a guide RNA (gRNA); and c) a nucleotide sequence encoding a nickase.

103. The retron editor of claim 102, wherein the nickase is Cas9(D10A).

Citation Information

Patent Citations

  • High-throughput precision genome editing in human cells

    WO2023019164A2