Genome editing using programmable unwinding-annealing helicase and related single-strand annealing proteins
Programmable unwinding-annealing helicases and SSAPs, combined with CRISPR, enable precise and efficient integration of long DNA sequences, addressing safety and efficiency challenges in genome editing for gene therapy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-03-05
AI Technical Summary
Traditional genome editing methods face challenges with off-target effects and limited capacity to integrate long DNA sequences safely and efficiently, posing obstacles for therapeutic applications in gene therapy.
Utilizing programmable unwinding-annealing helicases and single-strand nucleic acid annealing proteins (SSAPs) combined with CRISPR-based targeting systems, specifically employing human-derived helicases like Twinkle or Erf, to achieve precise and efficient integration of long DNA sequences at designated genomic locations, minimizing immune reactions.
Enhances the safety and efficiency of genome editing by reducing immunogenicity and immune-mediated rejection, optimizing compatibility and safety for therapeutic use.
Smart Images

Figure US2025043714_05032026_PF_FP_ABST
Abstract
Description
[0001] STDU2-43469.601
[0002] GENOME EDITING USING PROGRAMMABLE UNWINDING-ANNEALING HELICASE AND RELATED SINGLE-STRAND ANNEALING PROTEINS
[0003] STATEMENT OF RELATED APPLICATIONS
[0004] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 687,496, filed August 27, 2024, the entire contents of which are incorporated herein by reference for all purposes.
[0005] STATEMENT OF GOVERNMENT SUPPORT
[0006] This invention was made with Government support under contracts GM 141627 and HG011316 awarded by the National Institutes of Health. The Government has certain rights in the invention.
[0007] SEQUENCE LISTING
[0008] The text of the computer readable sequence listing filed herewith, titled “STDU2- 43469.601_SQL. xml”, created August 27, 2025, having a file size of 16,822 bytes, is hereby incorporated by reference in its entirety.
[0009] FIELD
[0010] Provided herein are compositions, systems, and methods for genome engineering using programmable unwinding-annealing helicase and related single-strand nucleic acid annealing proteins (SSAPs), for example, from eukaryotic cells.
[0011] BACKGROUND
[0012] The present invention tackles a pivotal challenge in genome engineering and gene therapy: the safe and efficient integration of substantial DNA segments into the human genome. Traditional methods for genome editing often struggle with off-target effects and limited capacity to introduce long DNA sequences without triggering adverse cellular responses. These limitations pose significant hurdles in therapeutic applications, particularly in the development of safe and effective gene therapies. STDU2-43469.601
[0013] SUMMARY
[0014] Provided herein are compositions, systems, and methods for genome engineering using programmable unwinding-annealing helicase and related single-strand nucleic acid annealing proteins (SSAPs), for example, from eukaryotic cells.
[0015] The compositions, systems, and methods utilize the specificity of CRISPR-based DNA- targeting or other targeting system combined with the unique DNA unwinding and annealing capabilities of eukaryotic helicases or other SSAPs to achieve precise and efficient genome editing. By harnessing these natural processes, the system facilitates the integration of long DNA sequences at designated genomic locations, vastly improving the accuracy and efficiency of gene edits compared to existing technologies.
[0016] Moreover, addressing potential adverse immune reactions is paramount for successful gene and cell therapies. To mitigate this risk, the compositions, systems, and methods employ a core gene-editing enzyme that is endogenous to the target cells (e.g., human cells). For example, for use with humans, by using a human-derived helicase and SSAPs, particularly the mitochondria Twinkle helicase or Erf protein, the technology significantly reduces the immunogenicity often associated with foreign gene-editing proteins. This approach not only enhances the safety profile of the technology but also increases its therapeutic capability by minimizing immune-mediated rejection in clinical applications. Thus, the technology not only overcomes major obstacles related to genome engineering of long therapeutic genetic sequences, but also optimizes compatibility and safety for therapeutic use, thereby marking a significant advancement in the field of gene therapy.
[0017] In some embodiments, provided herein are system and methods for modifying a nucleic acid or cell, comprising: contacting a nucleic acid or introducing into a cell, a) a donor (e.g., a DNA or RNA-converted DNA donor); b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same; and, optionally, c) a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system, etc.). In some embodiments, the eukaryotic SSAP (e.g., helicase) is from the same species of organism as the subject. In some embodiments, the nucleic acid or the cell is modified in vitro, ex vivo, or in vivo. In some embodiments, the donor comprises an insertion sequence of 1000 base pairs or more (e.g., 2000 or more, 5000 or more, 10,000 or more, etc.). STDU2-43469.601
[0018] Also provided herein are not editing systems and methods. For example, on some embodiments, provide herein are systems comprising: a) a pegRNA comprising an MS2 hairpin; and b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same. In some embodiments, the system comprises: a) a prime editor or a sequence encoding the same; and b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same. Uses of such system (e.g., for modification of a nucleic acid or cell) are also provided herein.
[0019] In some embodiments, provided herein are methods of editing a nucleic acid molecule comprising: contacting a nucleic acid molecule with: a) a donor (e.g., a DNA or RNA-converted DNA donor); b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same; and, optionally, c) a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system).
[0020] Further provided herein are novel SSAPs (e.g., helicases) and kits, systems, and methods employing the same. For example, in some embodiments, provided herein are non-natural helicases comprising SEQ ID NO:4, or a sequencing having at least 75% sequence identity with SEQ ID NO:4 and having one or more amino acid substitutions at positions selected from positions G24, D127, C247, F289, F3O8, and Q425. In some embodiments, the one or more amino acid substitutions comprise G24K, D127E, C247T, F289Y, F308R, and Q425N. In some embodiments, the non-natural helicase comprises amino acid substitutions: G24K and at least one or all of F289Y, D127E, and C247T or G24K; at least one or all of F289Y, D127E, and C247T; and one or both of F308R and Q425N. In some embodiments, the non-natural helicase comprises SEQ ID NOG, or a sequencing having at least 75% sequence identity with SEQ ID NOG and having one or more amino acid substitutions at positions selected from positions 174, A75, E94, Al 10, S169, and M140. In some embodiments, the one or more amino acid substitutions comprise I74R, A75V, E94R, A110R, S169K or S169Q, and M140K. In some embodiments, the non- natural helicase comprises amino acid substitutions: M140K; S169Q; and one or more of: I74R, A75V, and Al 10R; M140K; S169Q; A75V; and one or both of I74R and Al 10R; M140K; S169K; and one or more of: I74R, A75V, E94R, and Al 10R; M140K; S169K; and one or both of A75V and E94R; or M140K; S169K; one or both of A75V and E94R; and one or both of Al 10R and I74R.
[0021] Also provided herein are compositions (e.g., kits, reactions mixtures, cells, etc.) comprising such non-natural helicase or expression vectors encoding the same. The compositions may further comprise a donor (e.g., DNA or RNA-converted DNA donor) and / or a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing STDU2-43469.601 system). Uses of the helicase, composition, cell, expression vector or system are also provided (e.g., to insert the donor into a heterologous nucleic acid (e.g., a genome of a cell).
[0022] Definitions
[0023] To facilitate an understanding of the present technology, a number of terms and phrases are defined below. Additional definitions are set forth throughout the detailed description.
[0024] The terms “comprisc(s),” “includc(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. The singular forms “a,” “and” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of’ and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.
[0025] For the recitation of numeric ranges herein, each intervening number there between with the same degree of precision is explicitly contemplated. For example, for the range of 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0-7.0, the number 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated.
[0026] Unless otherwise defined herein, scientific, and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. For example, any nomenclature used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those that are well known and commonly used in the ail. The meaning and scope of the terms should be clear; in the event, however of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.
[0027] The terms “complementary” and “complementarity” refer to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick base-paring or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences can be indicated by the percentage of nucleotides in a nucleic acid sequence which can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second STDU2-43469.601 nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary). Two nucleic acid sequences arc “perfectly complementary” if all the contiguous nucleotides of a nucleic acid sequence will hydrogen bond with the same number of contiguous nucleotides in a second nucleic acid sequence. Two nucleic acid sequences are “substantially complementary” if the degree of complementarity between the two nucleic acid sequences is at least 60% (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%. 97%, 98%, 99%, or 100%) over a region of at least 8 nucleotides (e.g., 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides), or if the two nucleic acid sequences hybridize under at least moderate, preferably high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37° C in a solution comprising 20% formamide, 5xSSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5xDenhardt’s solution, 10% dextran sulfate, and 20 mg / ml denatured sheared salmon sperm DNA, followed by washing the filters in IxSSC at about 37-50° C., or substantially similar conditions, e.g., the moderately stringent conditions described in Sambrook et al., infra. High stringency conditions are conditions that use, for example (1) low ionic strength and high temperature for washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50° C, (2) employ a denaturing agent during hybridization, such as formamide, for example, 50% (v / v) formamide with 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42° C., or (3) employ 50% formamide, 5xSSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5xDenhardt’s solution, sonicated salmon sperm DNA (50 pg / ml), 0.1% SDS, and 10% dextran sulfate at 42° C., with washes at (i) 42° C. in 0.2xSSC, (ii) 55° C. in 50% formamide, and (iii) 55° C. in O.lxSSC (preferably in combination with EDTA). Additional details and an explanation of stringency of hybridization reactions are provided in, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, N.Y. (2001); and Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and John Wiley & Sons, New York (1994).
[0028] A cell has been “genetically modified,” “transformed,” or “transfected” by exogenous DNA, e.g., a recombinant expression vector, when such DNA has been introduced inside the cell. The presence of the exogenous DNA results in permanent or transient genetic change. The STDU2-43469.601 transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. In prokaryotes, yeast, and mammalian cells for example, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones that comprise a population of daughter cells containing the transforming DNA. A “clone” is a population of cells derived from a single cell or common ancestor by mitosis. A “cell line” is a clone of a primary cell that is capable of stable growth in vitro for many generations.
[0029] As used herein, a “nucleic acid” or a “nucleic acid sequence” refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively. The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogenous or homogenous in composition and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single- stranded or double- stranded form, including homoduplex, heteroduplex, and hybrid states. In some embodiments, a nucleic acid or nucleic acid sequence comprises other kinds of nucleic acid structures such as, for instance, a DNA / RNA helix, peptide nucleic acid (PNA), morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14): 4503-4510 (2002)) and U.S. Pat. No. 5,034,506, incorporated herein by reference), locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97: 5633-5638 (2000), incorporated herein by reference), cyclohcxcnyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: 8595-8602 (2000), incorporated herein by reference), and / or a ribozyme. Hence, the term “nucleic acid” or “nucleic acid sequence” may also encompass a chain comprising non-natural nucleotides, modified nucleotides, and / or non- nucleotide building blocks that can exhibit the same function as natural nucleotides (e.g., “nucleotide analogs”); further, the term “nucleic acid sequence” as used herein refers to an oligonucleotide, nucleotide or polynucleotide, and fragments or portions thereof, and to DNA or RNA of genomic or synthetic origin, which may be single or double-stranded, and represent the sense or antisense STDU2-43469.601 strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and “oligonucleotide” arc used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.
[0030] A “peptide” or “polypeptide” is a linked sequence of two or more amino acids linked by peptide bonds. The peptide or polypeptide can be natural, synthetic, or a modification or combination of natural and synthetic. Polypeptides include proteins such as binding proteins, receptors, and antibodies. The proteins may be modified by the addition of sugars, lipids or other moieties not included in the amino acid chain. The terms “polypeptide” and “protein,” are used interchangeably herein.
[0031] As used herein, the term “percent sequence identity” refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid sequence, or amino acids in an amino acid sequence, that is identical with the corresponding nucleotides or amino acids in a reference sequence after aligning the two sequences and introducing gaps, if necessary, to achieve the maximum percent identity. Hence, in case a nucleic acid according to the technology is longer than a reference sequence, additional nucleotides in the nucleic acid, that do not align with the reference sequence, are not taken into account for determining sequence identity. Methods and computer programs for alignment are well known in the art, including BLAST, Align 2, and FASTA.
[0032] The term “percent sequence similarity” takes into account conservative amino acid substitutions. Conservative substitution tables providing functionally similar amino acids are well known in the art. The following five groups each contain amino acids that are conservative substitutions for one another: Aliphatic: Glycine (G), Alanine (A), Valine (V), Leucine (L), Isoleucine (I); Aromatic: Phenylalanine (F), Tyrosine (Y), Tryptophan (W); Sulfur containing: Methionine (M), Cysteine (C); Basic: Arginine (R), Lysine (K), Histidine (H); Acidic: Aspartic acid (D), Glutamic acid (E), Asparagine (N), Glutamine (Q). In addition, individual substitutions, deletions, or additions which alter, add, or delete a single amino acid or a small percentage of amino acids in an encoded sequence are also “conservatively modified variations.” Means for making this adjustment are well known to those of skill in the art. Typically this involves scoring a conservative substitution as a partial rather than a full mismatch, thereby increasing the percentage sequence identity. Thus, for example, where an identical amino acid is STDU2-43469.601 given a score of 1 and a non conservative substitution is given a score of zero, a conservative substitution is given a score between zero and 1. The scoring of conservative substitutions is calculated, e.g., as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).
[0033] It will be understood that sequences identified as having greater than a given percent similarity to a reference sequence include as a subset the sequences having greater than the given percent identity to the reference sequence. Thus, recitations herein to sequences having greater than a given percent similarity include the subset of sequences having greater than a given percent identity.
[0034] A “vector” or “expression vector” is a replicon, such as plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an “insert,” may be attached or incorporated so as to bring about the replication of the attached segment in a cell.
[0035] The term “wild-type” refers to a gene or a gene product that has the characteristics of that gene or gene product when isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designated the “normal” or “wild-type” form of the gene. In contrast, the term “modified,” “mutant,” or “polymorphic” refers to a gene or gene product that displays modifications in sequence and or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product.
[0036] A “subject” may be human or non-human and may include, for example, plants or animal strains or species used as “model systems” for research purposes, such a mouse model as described herein. Likewise, subject may include either adults or juveniles (e.g., children). Moreover, subject may mean any living organism, preferably a mammal (e.g., human or non- human) that may benefit from the administration of compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the Mammalian class: humans, non- human primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice, and guinea pigs, and the like. Examples of nonmammals include, but are not limited to, birds, fish and the like. In one embodiment of the STDU2-43469.601 methods and compositions provided herein, the mammal is a human. Plants include without limitation sugar canc, corn, wheat, rice, oil palm fruit, potatoes, soybeans, vegetables, cassava, sugar beets, tomatoes, barley, bananas, watermelon, onions, sweet potatoes, cucumbers, apples, seed cotton, oranges, and the like.
[0037] As used herein, the terms “providing”, “administering,” “introducing,” arc used interchangeably herein and refer to the placement of the systems of the disclosure into a subject by a method or route which results in at least partial localization of the system to a desired site. The systems can be administered by any appropriate route which results in delivery to a desired location in the subject.
[0038] The phrase “altering a DNA sequence,” as used herein, refers to modifying at least one physical feature of a DNA sequence of interest. DNA alterations include, for example, single or double strand DNA breaks, deletion, or insertion of one or more nucleotides, and other modifications that affect the structural integrity or nucleotide sequence of the DNA sequence. The modifications of a target sequence in genomic DNA may lead to, for example, gene correction, gene replacement, gene tagging, transgene insertion, nucleotide deletion, gene disruption, gene mutation, gene knock-down, and the like.
[0039] The term “donor nucleic acid molecule” refers to a nucleotide sequence that is inserted into the target DNA (e.g., genomic DNA). As described above the donor DNA may include, for example, a gene or part of a gene, a sequence encoding a tag or localization sequence, or a regulating element. The donor nucleic acid molecule may be of any length. In some embodiments, the donor nucleic acid molecule is between 10 and 10,000 nucleotides in length. For example, between about 100 and 5,000 nucleotides in length, between about 200 and 2,000 nucleotides in length, between about 500 and 1,000 nucleotides in length, between about 500 and 5,000 nucleotides in length, between about 1,000 and 5,000 nucleotides in length, or between about 1,000 and 10,000 nucleotides in length, or greater than 10,000 nucleotides in length.
[0040] Description of Figures
[0041] FIG. 1 shows exemplary SSAP editor designs: a) construct designs for screening geneediting efficiency of SSAPs; and b) dCas9-SSAP editor designs. In Fig. la, a dCas9-gRNA STDU2-43469.601 vector design is shown having a human U6 promoter linked to an MS2 hairpin guide RNA and further comprising a CBh promoter linked to an NLS-dCas9 (D10A & H840A) and a 2A-BFP reporter gene. A double stranded DNA donor comprises an 800 base pair knock-in construct comprising a sequence encoding mKate fluorescent protein. Also shown is an SSAP screening vector comprising an EF1A promoter and an SSAP of interest. In Fig. lb, an exemplary dCas9- SSAP editor is contacted to a genomic target in mammalian cells leading to SSAP-mediated insertion of the knock-in donor.
[0042] FIG. 2 shows screening strategy and results for newly identified SSAPs and human SSAP-like proteins. FIG. 2a shows a phylogenetic representation of candidate SSAPs. FIG. 2b shows results of editing efficiency studies for candidate SSAPs. FIG 2c shows knock-in % with SSAP16 and Erf, identified in the screen, compared to RecT, Cas only, and Cas + Donor only.
[0043] FIG. 3 shows design of a Twinkle editing system. FIG 3a shows an exemplary construct design. FIG 3b shows knock-in efficiency data.
[0044] FIG. 4 shows exemplary construct designs and data involving SSAP mediated large gene fragments knock-in with single-stranded circular DNA. FIG. 4a shows exemplary dCas9-gRNA, SSAP, and donor designs. FIG. 4b shows knock-in efficiency when SSAPs are combined with cssDNA donors. FIG 4c shows results of dose dependency studies.
[0045] FIG. 5 shows SSAP enhanced large-fragment knock-in in primary cells.
[0046] FIG. 6 shows in vivo knock-in with SSAP-enhanced AAV delivery.
[0047] FIG. 7 shows SSAP-enhanced in vitro prime editing efficiency. FIG. 7a shows exemplary construct designs. FIG. 7b shows editing efficiency data. FIG. 7c shows fold-change data.
[0048] FIG. 8 shows improving gene knock-in efficiency with Cas-free SSAP. FIG 8a shows a knock-in representation. FIG. 8b shows knock-in data.
[0049] FIG. 9 shows an exemplary finetune methodology based on the best sequence-input model.
[0050] FIG. 10 shows fine-tune, model-guided variant Twinkle knock-in results: FIG. 10a (B2M cssDNA); FIG 10b (RABI la cssDNA); and FIG. 10c (ACTB dsDNA). Mutation positions are based on a sequence comprising an N-terminal tag having four amino acids. Mutation numbers STDU2-43469.601 shown, minus 4, correspond to the positions in SEQ ID NOG (e.g., G28K in FIG. 10 corresponds to the G at position 24 in SEQ ID NOG).
[0051] FIG. 11 shows fine-tune, model-guided variant Erf knock-in results: FIG. 1 la (B2M cssDNA); FIG 1 lb (RABI la cssDNA); and FIG. 11c (ACTB dsDNA). Mutation positions are based on a sequence comprising an N-tcrminal tag having one amino acid. Mutation numbers shown, minus 1, correspond to the positions in SEQ ID NOG (e.g., M141K in FIG. 11 corresponds to the M at position 140 in SEQ ID NOG).
[0052] FIG. 12 shows validation data of Al-designed multi-site SSAP variants (Erf & Twinkle). Mutation positions for Twinkle are based on a sequence comprising an N-terminal tag having four amino acids. Mutation numbers shown, minus 4, correspond to the positions in SEQ ID NOG (e.g., G28K in FIG. 12 corresponds to the G at position 24 in SEQ ID NOG). Mutation positions for Erf are based on a sequence comprising an N-terminal tag having one amino acid. Mutation numbers shown, minus 1, correspond to the positions in SEQ ID NOG (e.g., M141K in FIG. 12 corresponds to the M at position 140 in SEQ ID NOG).
[0053] DETAILED DESCRIPTION
[0054] In some embodiments, the technology comprises two or more components: 1) a donor (e.g., a DNA or RNA-converted DNA donor); 2) a helicase or other SSAP (e.g., eukaryotic helicase, e.g., helicase that is species-matched to the target cell, tissue, or organism); and, optionally, 3) a nucleic acid targeting system (e.g., a programmable DNA targeting system) (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system, a base editing system, etc.).
[0055] In some embodiments, the DNA or RNA-converted DNA donor serves as a template, specifying the type of gene editing to be performed. Any of a wide variety of editing approaches can be used including, but not limited to, insertion, deletion, or replacement of DNA sequences. In some embodiments, the technology is used to insert payloads (e.g., therapeutic payloads) into the genome. In some embodiments, this component includes a homology region (e.g., homology arm) that facilitates targeted insertion. In some embodiments, the homology arm is designed to be at least 10, 20, or 50 base pairs in length and shares a sequence similar to the target genomic STDU2-43469.601 site. When used in conjunction with a CRISPR system, the donor nucleic acid may comprise one or more guide RNAs.
[0056] In some embodiments, the SSAP (e.g., helicase) facilitates genome editing by unwinding DNA strands and promoting strand annealing, invasion, and exchange. In some embodiments, the SSAP (e.g., helicase) is a eukaryotic SSAP (e.g., helicase). In some embodiments, the SSAP (e.g., helicase) is a human SSAP (e.g., helicase). In some embodiments, the helicase is SF4 helicase. SF4 helicases are ring shaped and share five conserved helicase motifs: (1) the Hl / Walker A motif, which stabilizes the NTP phosphate; (2) Hl a, which is involved in NTP binding / hydrolysis; (3) the H2 / Walker B motif, which contains an arginine finger and a base stack residue required for positioning and stabilizing the bound NTP as well as stabilizing a bound Mg2+ ion; (4) H3, which with Hla is involved in NTP binding / hydrolysis; and (5) H4, which contributes to DNA binding (see Peter and Falkenberg, Genes (Basel), 1 l(4):408 (2020)). In some embodiments, the helicase is a mitochondrial helicase.
[0057] In some embodiments, the SSAP (e.g., helicase) comprises an amino acid sequence having at least 75% identity (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to any one of SEQ ID NOs: 1-9, as shown in Table 1. Any of the SSAPs described herein may comprise one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or more, etc.) amino acid substitutions as compared to SEQ ID NOs: 1-9. An amino acid “replacement” or “substitution” refers to the replacement of one amino acid at a given position or residue by another amino acid at the same position or residue within a polypeptide sequence. Amino acids are broadly grouped as “aromatic” or “aliphatic.” An aromatic amino acid includes an aromatic ring. Examples of aromatic amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr), and tryptophan (W or Trp). Non-aromatic amino acids are broadly grouped as aliphatic. Examples of aliphatic amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Vai), leucine (L or Leu), isoleucine (I or He ), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine (C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic STDU2-43469.601 acid (D or Asp), asparagine (N or Asn), glutamine (Q or Gin), lysine (K or Lys), and arginine (R or Arg).
[0058] The amino acid replacement or substitution can be conservative, semi-conservative, or non-conservative. The phrase “conservative amino acid substitution” or “conservative mutation” refers to the replacement of one amino acid by another amino acid with a common property. A functional way to define common properties between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schirmer, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analyses, groups of amino acids may be defined where amino acids within a group exchange preferentially with each other and therefore resemble each other most in their impact on the overall protein structure (Schulz and Schirmer). Examples of conservative amino acid substitutions include substitutions of amino acids within the sub-groups described above, for example, lysine for arginine and vice versa such that a positive charge may be maintained, glutamic acid for aspartic acid and vice versa such that a negative charge may be maintained, serine for threonine such that a free -OH can be maintained, and glutamine for asparagine such that a free -NH2 can be maintained. “Semi-conservative mutations” include amino acid substitutions of amino acids within the same groups listed above, but not within the same sub-group. For example, the substitution of aspartic acid for asparagine, or asparagine for lysine, involves amino acids within the same group, but different sub-groups. “Non-conservative mutations” involve amino acid substitutions between different groups, for example, lysine for tryptophan, or phenylalanine for serine, etc.
[0059] In some embodiments, the SSAP is a Twinkle helicase (GenBank AAK69558.1), a member of the SF4 superfamily of helicases, or a homologue, orthologue, or variant thereof. Twinkle, in addition to the other shared SF4 sequence motifs, is hexameric and each 72 kDa monomer is comprised of an N-terminal domain (NTD) and C-terminal domain (CTD) joined by a flexible linker helix. Twinkle has robust unwinding-annealing activities. The Twinkle helicase is the main helicase in mitochondria and is the only helicase required for mtDNA replication.
[0060] In some embodiments, the SSAP (e.g., helicase) comprises an amino acid sequence having at least 75% identity (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least STDU2-43469.601
[0061] 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 4. In some embodiments, the SSAP comprises an amino acid sequence having one or more amino acids substitutions at amino acid positions selected from positions G24, D127, C247, F289, F3O8, and Q425, as compared to SEQ ID NO: 4. In some embodiments, the SSAP comprises at least one amino acid substitution comprising G24K, D127E, C247T, F289Y, F308R, and Q425N, or any combination thereof, as compared to SEQ ID NO: 4. In some embodiments the SSAP comprises the amino acid substitutions: G24K and at least one or all of F289Y, D127E, and C247T, as compared to SEQ ID NO: 4. In some embodiments the SSAP comprises the amino acid substitutions: G24K; at least one or all of F289Y, D127E, and C247T; and one or both of F308R and Q425N, as compared to SEQ ID NO: 4.
[0062] In some embodiments, the SSAP (e.g., helicase) is a Erf family protein, or a homologue, orthologue, or variant thereof. In some embodiments, the SSAP comprises an amino acid sequence having at least 75% identity (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 3. In some embodiments, the SSAP comprises an amino acid sequence having one or more amino acids substitutions at amino acid positions selected from positions 174, A75, E94, Al 10, S169, and M 140, as compared to SEQ ID NO: 3. In some embodiments, the SSAP comprises at least one amino acid substitution comprising I74R, A75V, E94R, A110R, S169K or S169Q, M140K, or any combination thereof, as compared to SEQ ID NO: 3.
[0063] In some embodiments the SSAP (e.g., helicase) comprises the amino acid substitutions: M140K and S169K or S169Q, as compared to SEQ ID NO: 3. In some embodiments the SSAP comprises the amino acid substitutions: M140K; S169Q; and one or more of: 174R, A75V, and Al 10R, as compared to SEQ ID NO: 3. In some embodiments the SSAP comprises the amino acid substitutions: M140K; S169Q; A75V; and one or both of I74R and Al 10R, as compared to SEQ ID NO: 3. In some embodiments the SSAP comprises the amino acid substitutions: M140K; S169K; and one or more of: I74R, A75V, E94R, and Al 10R, as compared to SEQ ID NO: 3. In some embodiments the SSAP comprises the amino acid substitutions: M140K; S169K; STDU2-43469.601 and one or both of A75V and E94R, as compared to SEQ ID NO: 3. In some embodiments, the SSAP comprises the amino acid substitutions: M140K; S169K; one or both of A75V and E94R; and one or both of Al 10R and 174R, as compared to SEQ ID NO: 3.
[0064] Any of the SSAPs (e.g., helicase) disclosed herein may further comprise one or more proteins, polypeptides (e.g., protein domain sequences), or peptides. The one or more proteins, polypeptides (e.g., protein domain sequences), or peptides may be appended at an N-terminus, a C-terminus, internally, or a combination thereof. The one or more proteins, polypeptides (e.g., protein domain sequences), or peptides may be fused in any orientation in relationship to the disclosed protein. The one or more proteins, polypeptides (e.g., protein domain sequences), or peptides may be fused via a linker. For example, the SSAPs disclosed herein may be fused to another protein or protein domain that provides for tagging or visualization (e.g., GFP), localization, stability, and the like.
[0065] In some embodiments, the SSAPs (e.g., helicase) comprise one or more nuclear localization sequences (NLSs). The nuclear localization sequence may be appended, for example, to the N-terminus, a C-terminus, internally, or a combination thereof. The nuclear localization sequence may comprise any amino acid sequence known in the art to functionally tag or direct a protein for import into the nucleus (e.g., for nuclear transport). Usually, a nuclear localization sequence comprises one or more positively charged amino acids, such as lysine and arginine. The NLS may be appended by a linker.
[0066] In some embodiments, the NLS is a monopartite sequence. A monopartite NLS comprises a single cluster of positively charged or basic amino acids. In some embodiments, the monopartite NLS comprises a sequence of K-K / R-X-K / R, wherein X can be any amino acid. Exemplary monopartite NLS sequences include those from the SV40 large T-antigen, c-Myc, and TUS-proteins. In some embodiments, the NLS is a bipartite sequence. Bipartite NLSs comprise two clusters of basic amino acids, separated by a spacer of about 9-12 amino acids. Exemplary bipartite NLSs include the nuclear localization sequences of nucleoplasmin, EGL-12, or bipartite SV40.
[0067] The SSAPs may comprise an epitope tag (e.g., 3xFLAG tag, an HA tag, a Myc tag, and the like). In some embodiments, the epitope tag may be adjacent, either upstream or downstream, STDU2-43469.601 to a nuclear localization sequence. The epitope tags may be at the N-terminus, a C-terminus, or a combination thereof of the corresponding protein or polypeptide.
[0068] In some embodiments, the SSAPs may be fused with one or more (e.g., two, three, four, or more) protein transduction moieties. A protein transduction moiety is a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A protein transduction moiety attached to another molecule facilitates the molecule traversing a membrane, for example going from extracellular space to intracellular space, or cytosol to within an organelle. Examples of protein transduction moieties include but are not limited to a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV- 1 TAT comprising); a poly arginine sequence comprising a number of arginines sufficient to direct entry into a cell (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); a VP22 domain; a Drosophila Antennapedia protein transduction domain; a truncated human calcitonin peptide; polylysine; Transportan; and the like.
[0069] The present disclosure encompasses proteins comprising the disclosed SSAPs, fusion proteins comprising the disclosed SSAPs, nucleic acids encoding the disclosed SSAPs and / or proteins or fusion protein comprising thereof, and cells expressing the disclosed SSAPs and / or proteins or fusion protein comprising thereof or nucleic acids encoding the disclosed SSAPs and / or proteins or fusion protein comprising thereof.
[0070] While use of a programmable DNA targeting system (e.g., CRISPR) is optional, this component significantly enhances the efficiency and specificity of the gene-editing process by directing the helicase to specific genomic loci. A distinctive aspect of this system is the utilization of the core helicase derived from eukaryotic cells, for example, the mitochondria Twinkle protein encoded by human cells. This choice ensures that the gene-editing protein is endogenous to human cells, which minimizes immunogenicity and enhances its suitability for therapeutic applications. The incorporation of, for example, a CRISPR-Cas system for genome targeting and novel DNA modifying enzymes for editing represents a significant advancement in the capability for large-scale genome insertion and complex gene-editing tasks. This system overcomes longstanding barriers in gene therapy, providing a versatile tool for medical research and therapeutic treatments and other applications (e.g., synthetic biology, basic research, etc.). STDU2-43469.601
[0071] In some embodiments, a guide forms a complex with a CRISPR enzyme, is recruited to the target genomic loci for the editing. The enzymes can be cither fully active CRISPR enzymes, nickases, or deactivated CRISPR enzymes (dCas9, dCasl2a, etc.) that only bind to target loci. Cas protein families are described in further detail in, e.g., Haft et al., PLoS Comput. Biol., 1(6): e60 (2005), incorporated herein by reference. The Cas protein may be any Cas endonucleases. In some embodiments, the Cas protein is Cas9 or Cas 12a, otherwise referred to as Cpfl. In one embodiment, the Cas9 protein is a wild-type Cas9 protein. The Cas9 protein can be obtained from any suitable microorganism, and a number of bacteria express Cas9 protein orthologs or variants. In some embodiments, the Cas9 is from Streptococcus pyogenes or Staphylococcus aureus. Cas9 proteins of other species are known in the art (see, e.g., U.S. Patent Application Publication 2017 / 0051312, incorporated herein by reference) and may be used in connection with the present disclosure. The amino acid sequences of Cas proteins from a variety of species are publicly available through the GenBank and UniProt databases.
[0072] Any element of any suitable CRISPR / Cas gene editing system known in the art can be employed in the systems and methods described herein, as appropriate. CRISPR / Cas gene editing technology is described in detail in, for example, U.S. Patent Application Publication 2014 / 0068797; U.S. Patents 8,697,359; 8,771,945; and 8,945,839; US2010 / 0076057;
[0073] US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565; WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US20150045546; US20150031134; US20150024500; US20140377868; US20140357530; US20140349400; US20140335620; US20140335063; US20140315985; US20140310830; US20140310828; US20140309487; US20140304853; US20140298547; US20140295556; US20140294773; US20140287938; US20140273234; US20140273232; US20140273231 ; US20140273230; US20140271987; US20140256046; US20140248702; US20140242702; US20140242700; US20140242699; US20140242664; US20140234972; US20140227787; US20140212869; US20140201857; US20140199767; US20140189896; US20140186958; US20140186919;
[0074] US20140186843; US20140179770; US20140179006; and US20140170753, incorporated herein by reference.
[0075] The genome engineering system leveraging SSAPs (e.g., programmable nucleic acid unwinding-annealing helicases) offers several significant advantages over existing gene-editing STDU2-43469.601 technologies, such as traditional CRISPR-Cas9, TALENs, and ZFNs. These improvements address key challenges in precision, efficiency, immune response, and the scope of editable genetic material.
[0076] The technology provides enhanced precision and efficiency. Traditional CRISPR-Cas9 systems can suffer from off-target effects, where unintended parts of the genome arc modified. The present technology provides targeting capabilities with the fidelity of SSAPs, e.g., eukaryotic helicases, for DNA unwinding and annealing. This dual-action approach significantly minimizes off-target interactions and improves the accuracy of gene insertion, deletion, and replacement events, especially for long DNA sequences.
[0077] The technology provides the capability to integrate long DNA sequences. Methods like CRISPR-Cas9 are typically limited in the length of DNA they can effectively integrate, often struggling with sequences longer than a few thousand base pairs. The present technology can integrate substantially longer sequences of DNA into the genome efficiently, which is crucial for adding complex gene clusters or large therapeutic genes.
[0078] The technology provides reduced immunogenicity. Foreign enzymes used in other geneediting technologies can elicit immune responses in patients, complicating therapeutic applications. By employing a SSAP (e.g., helicase enzyme) that is endogenous to target cells (e.g., human cells), the present technology significantly reduces the risk of immune reactions, enhancing the safety profile for clinical use in gene therapies.
[0079] The technology provides broad applicability across cell types. Some editing tools, like TALENs and ZFNs, show variable efficiency across different cell types and species. The versatility of the SSAP (e.g., helicase) used in the present technology allows for effective use across a diverse range of cell types and organisms, expanding the applications in therapeutic, agricultural, and research settings.
[0080] The technology can also be used in conjunction with existing gene editing technologies which may benefit from the improvements and advantages described above. For example, the disclosed technology may provide improvements and advantages to new generations of CRISPR gene editing or other gene editing methods and tools utilizing Cas proteins, including prime editing and base editing. STDU2-43469.601
[0081] Prime editing ameliorates both transition and transversion mutations in addition to small deletions and insertions. Generally, a prime editing guide RNA (pcgRNA) is used in conjunction with a prime editor, e.g., a H840A Streptococcus pyogenes Cas9 (spCas9) nickase linked to a reverse transcriptase (RT), e.g., optimized Moloney murine leukemia virus (MMLV). pegRNAs are similar to standard single-guide RNAs (sgRNAs) but differ due to a sequence comprising a primer binding site (PBS) and a reverse transcription template (RTT) sequence. The primer binding site hybridizes with the bases upstream of the prime editor generated nick, while the RTT encodes the information of the intended edits and directs reverse transcription. Together, the prime editor and the pegRNA form the prime editing 2 strategy (PE2). The Cas9 nickase is guided to the DNA target site by the pegRNA. After nicking by Cas9, the reverse transcriptase uses the pegRNA to template reverse transcription of the desired edit, directly polymerizing DNA onto the nicked target DNA strand. The edited DNA strand replaces the original DNA strand, creating a heteroduplex containing one edited strand and one unedited strand. Once the prime editor incorporates the edit into one strand, there is a mismatch between the original sequence on one strand and the edited sequence on the other strand. In some embodiments, an additional nicking guide RNA (ngRNA) is used to nick the non-edited strand, directing DNA repair enzymes to use the edited strand as a template to remake the mismatched strand. The prime editor, the pegRNA, and ngRNA form prime editing 3 (PE3) strategies.
[0082] As disclosed here in Example 6 and Figure 7, systems utilizing an SSAP in combination with a prime editor and a pegRNA designed to recruit the SSAP to the prime editing complex enhanced in vitro prime editing efficiency. In some embodiments, the system comprises a prime editor, an SSAP as disclosed herein, and a pegRNA comprising one or more sequences or elements configured to recruit the SSAP to the prime editing complex.
[0083] Base editing facilitates conversion of one base into another without strand cutting, unlike traditional CRISPR-Cas9. Generally, a base editor is used with a guide RNA to target the base for the desired mutation. Base editors comprise fusions between a Cas nuclease, usually catalytically impaired, and a base-modification enzyme (e.g., a deaminase). Upon binding to its target locus, base pairing between the guide RNA and target strand leads to displacement of a small segment of single- stranded DNA in an “R-loop” which is then modified by the deaminase enzyme. Two classes of DNA base editor have been described: cytosine base editors (CBEs) STDU2-43469.601 convert a C*G base pair into a T«A base pair, and adenine base editors (ABEs) convert an A«T base pair to a G*C base pair. Collectively, CBEs and ABEs can mediate four possible transition mutations (C to T, A to G, T to C, and G to A). In RNA, targeted adenosine conversion to inosine has also been developed using both antisense and Casl3-guided RNA-targeting methods.
[0084] Base editing systems can be utilized in conjunction with the SSAP technology described herein. In some embodiments, the system comprises a base editor, an SSAP as disclosed herein, and a gRNA, wherein the gRNA or base editor comprises one or more elements (e.g., sequences or heterodimerization elements) configured to recruit the SSAP to the base editing complex.
[0085] The genome engineering technology described herein opens up a broad spectrum of commercial opportunities across multiple sectors including biotechnology, pharmaceuticals, and personalized medicine. This technology’s ability to precisely and safely integrate long DNA sequences into the human genome addresses critical needs in gene therapy, genetic research, and therapeutic development. Current gene therapies often face limitations due to size constraints and immunogenic reactions, which can compromise treatment efficacy and safety.
[0086] In some embodiments, the technology allows for the development of customized gene therapies that can insert large therapeutic genes or gene clusters directly into genomes of cells or tissues ex vivo or in vivo. These therapies can be used to treat a wide array of genetic disorders such as muscular dystrophy, cystic fibrosis, and hemophilia.
[0087] Researchers require more accurate tools for genetic manipulation to advance understanding of complex genetic diseases and to develop novel treatments. The technology provided herein facilitates these advances. The CRISPR-cnhanccd SSAP (c.g., helicase) methods and kits provided herein allow for high-precision genome editing. These methods and kits find use in academic and commercial research laboratories for model organism development, gene function studies, and drug target validation.
[0088] The technology also finds use in agricultural biotechnology. There is a growing demand for genetically modified organisms that can withstand harsh environments, resist pests, and increase yield. Application of the genome editing technology provided herein finds use in the STDU2-43469.601 development of robust crop strains and livestock with desired traits such as drought resistance, enhanced nutritional content, or improved growth rates.
[0089] The technology also finds use for cell line development, for example, for bioproduction applications. Biomanufacturing processes often require cell lines with enhanced production capabilities or stability, which arc difficult to engineer with current technologies. The present technology allows for custom cell line development services that utilize the SSAP-based (e.g., helicase-based) editing system to insert productivity-enhancing genes into host cells, improving yield and stability for pharmaceutical production.
[0090] One or more of the component parts of the present systems (e.g., DNA or RNA- converted DNA donor; SSAP (e.g., helicase); programmable DNA targeting system (e.g., CRISPR)) may be linked or fused to one another.
[0091] In some embodiments, system components are delivered to cells as functional proteins and nucleic acids. In other embodiments, one or more components are expressed in a target cell. In some such embodiments, the components may be introduced into the cell as an expression vector (e.g., viral vectors such as adenovirus, lentivirus, adeno-associated virus, or retrovirus vectors), whereby the functional proteins and / or nucleic acids are generated within the target cell. Proteins, nucleic acids, or expression systems (e.g., expression vectors) may be introduced into a cell in vitro, ex vivo, or in vivo using any suitable delivery system. Such systems include, but are not limited to, lipid-based nanoparticles (e.g., liposomes, lipid nanoparticles), vectors, electroporation, microinjections, sonoporation, biolistics, dendrimers, peptide-based delivery (cell-penetrating peptides, fusogenic peptides), exosomes, chemical transfer reagents (e.g., calcium phosphate, cationic lipids, cationic polymers), and hydrogel scaffolds.
[0092] In some embodiments, the systems and methods described herein may be used to correct one or more defects or mutations in a gene (referred to as “gene correction”). In such cases, the target genomic DNA sequence encodes a defective version of a gene, and the system further comprises a donor nucleic acid molecule which encodes a wild-type or corrected version of the gene. Thus, in other words, the target genomic DNA sequence is a “disease-associated” gene. The term “disease-associated gene,” refers to any gene or polynucleotide whose gene products are expressed at an abnormal level or in an abnormal form in cells obtained from a disease- affected individual as compared with tissues or cells obtained from an individual not affected by STDU2-43469.601 the disease. A disease-associated gene may be expressed at an abnormally high level or at an abnormally low level, where the altered expression correlates with the occurrence and / or progression of the disease. A disease-associated gene also refers to a gene, the mutation or genetic variation of which is directly responsible or is in linkage disequilibrium with a gene(s) that is responsible for the etiology of a disease. Examples of genes responsible for such “single gene” or “monogenic” diseases include, but are not limited to, adenosine deaminase, a-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), P-hemoglobin (HBB), oculocutaneous albinism II (0CA2), Huntingtin (HTT), dystrophia myotonica-protein kinase (DMPK), low-density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney disease 1 (PKD1), polycystic kidney disease 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate -regulating endopeptidase homologue, X-linked (PHEX), methyl-CpG-binding protein 2 (MECP2), and ubiquitin- specific peptidase 9Y, Y-linked (USP9Y). Other single gene or monogenic diseases are known in the art and described in, e.g., Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1): 192 (2008), incorporated herein by reference; Online Mendelian Inheritance in Man (OMIM); and the Human Gene Mutation Database (HGMD).
[0093] The disclosed systems and methods overcome challenges encountered during conventional gene editing, including low efficiency and off-target events, particularly with kilobase- scale nucleic acids. In some embodiments, the disclosed systems and methods improve the efficiency of gene editing.
[0094] The disclosure further provides kits containing one or more reagents or other components useful, necessary, or sufficient for practicing any of the methods described herein. For example, kits may include one or more SSAPs (e.g., helicases), one or more donor nucleic acids, and / or one or more CRISPR reagents (Cas protein, guide RNA, vectors, compositions, etc.), transfection or administration reagents, negative and positive control samples (e.g., cells, template DNA), cells, containers housing one or more components (e.g., microccntrifugc tubes, boxes), detectable labels, detection and analysis instruments, software, instructions, and the like. STDU2-43469.601
[0095] Although the present invention and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the invention as defined in the appended claims.
[0096] Table 1 STDU2-43469.601
[0097] SpCas9 guideRNA
[0098] Wildtype sgRNA scaffold (SEQ ID NO: 10):
[0099] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAA AGTGGCACCGAGTCGGTGC
[0100] MS2 G-C opt sgRNA scaffold (SEQ ID NO: 11):
[0101] GTTTCAGAGCTAGGCCAACATGAGGATCACCCATGTCAAAAGGCCTAGCAAGTTGA AATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCGCGCACA TGAGGATCACCCATGTGC sgRNA for in vivo study (mouse) mmAlb_sg3: GTAAATATCTACTAAGACAA (SEQ ID NO: 12) mmAlb_sg4: GCATCTTCAGGGAGTAGCTT (SEQ ID NO: 13) STDU2-43469.601
[0102] EXAMPLES
[0103] Example 1
[0104] This example describes the identification of new SSAPs and human SSAP-like proteins. From a collection of over 4,000 candidate SSAP sequences, 380 natural SSAP proteins were selected and synthesized. Each SSAP was cloned into an SSAP-MCP plasmid and co-transfected with dCas9 / gRNA and a double- stranded DNA donor (carrying an mKate reporter) into HEK293T cells (see FIG 1). 72 hours post-transfection, cells were analyzed by FACS to measure knock-in efficiency. The screening was conducted in an arrayed format, with validation at two independent endogenous loci to ensure reproducibility. Among the 380 candidates, novel SSAPs such as the ERF family protein and SSAP 16 demonstrated markedly high efficiency. Compared to the first-generation reference protein RecT, these new SSAPs showed significant improvements in gene integration efficiency across both loci. Importantly, the enhancement was consistently observed across different Cas9 systems (Cas9, dCas9, and nCas9) (see FIG 2).
[0105] Example 2
[0106] This example describes the design and testing of a Twinkle editing system (FIG 3). Twinkle is a human mitochondrial helicase with SSAP-like DNA annealing and strand exchange activity. To adapt Twinkle for genome engineering, its mitochondrial targeting sequence (MTS) was removed. Knock-in efficiency of TwinkleAMTS was tested using arrayed assays in HEK293T cells. To assess functional determinants, Twinkle was fused with MCP or PCP RNA- binding domains and had mutations introduced at key residues (K421A, G575D). TwinkleAMTS demonstrated dose-dependent increases in knock-in efficiency, reaching ~8-9 fold higher than control (FIG 2b). Knock-in activity was significantly reduced when MCP was removed or replaced with PCP. Mutations at key amino acids (K421A, G575D) caused a strong loss of activity.
[0107] Example 3
[0108] This example describes SSAP mediated large gene fragments knock-in with single-stranded circular DNA. Compared to linear double- stranded DNA (dsDNA), singlestranded DNA (ssDNA) elicits a lower immune response (FIG. 4). Circular ssDNA (cssDNA) offers additional advantages, including higher stability and further reduced immunogenicity. It STDU2-43469.601 was tested whether SSAPs (RecT, SSAP16, Erf, Twinkle, etc.) could enhance knock-in efficiency when combined with cssDNA donors. The experimental design used dCas9 / nCas9 / Cas9 with SSAP-MCP and cssDNA donors targeting the RAB11A locus with an mCherry reporter. SSAPs combined with cssDNA donors showed robust knock-in efficiency at the RABI 1A locus across Cas9, nCas9, and dCas9 conditions. ERF, SSAP16, and Twinkle demonstrated higher performance than the RecT benchmark, validating their superior activity. cssDNA donor dosage (30 ng, 90 ng, 180 ng) resulted in dose-dependent increases in knock-in efficiency. Both human and mouse Twinkle variants supported high-efficiency cssDNA- mediated knock-in.
[0109] Example 4
[0110] This example describes SSAP enhanced large-fragment knock-in in primary cells (FIG 5). iPSC (RAB11A, cssDNA donor / mCherry): Electroporated SSAP mRNA, nCas9 mRNA, and gRNA together with a circular ssDNA (cssDNA) donor. Cells were analyzed by FACS 72 h post-electroporation. iPSC (RABI 1 A): Controls: donor only -2%, donor + Cas9 -7% mCherry+. Erf and TwinkleAMTS achieved -14-15% mCherry , a ~2x gain over donor + Cas9.
[0111] Fibroblast (B2M, cssDNA donor / GFP): Transfected SSAP plasmid, nCas9 plasmid, and gRNA with a cssDNA donor using a condensation / polymer-based transfection kit; FACS at 72 h. Tested SSAP candidates Erf and TwinkleAMTS against controls (donor only, donor + Cas9 / dCas9). Fibroblast (B2M): dCas9 + donor -3% GFP' . Erf -11-12% and TwinkleAMTS -10-11% GFP+(~4x increase over control).
[0112] Activity was robust across two primary cell types and two genomic loci, using cssDNA donors.
[0113] Example 5
[0114] This example describes in vivo knock-in with SSAP-enhanced AAV delivery (FIG. 6). Dual-AAVs were injected IV (tail vein) into mice: Cas9 AAV: expression of Cas9. SSAP editor + donor AAV: U6-MS2-dgRNA (Alb target or NT control), TBG-MCP-SSAP (groups: Gen2 STDU2-43469.601
[0115] SSAP = Erf, Gen! SSAP = RecT, or EBFP control), and a NanoLuc donor flanked by Albumin homology arms within the same vector. Animals were imaged by IVIS on Day 14 post-injection.
[0116] With Alb-gRNA, the Gen2 SSAP (Erf) group showed the highest liver bioluminescence, clearly exceeding Genl SSAP (RecT) and EBFP controls (bar graph). With NT-gRNA, all groups were near baseline, confirming on-target dependence. Representative auto- and manualexposure images corroborate the quantitative increase in the Gen2 SSAP cohort.
[0117] Example 6
[0118] This example describes SSAP-enhanced in vitro prime editing efficiency (FIG. 7). An MCP-SSAP fusion (NLS, EF1A promoter) was expressed together with a Prime Editor (FIG 7a). A pegRNA scaffold was engineered to carry 2x MS2 hairpins, permitting MCP-mediated recruitment of SSAP to the PE complex. Compare conditions: NegCtrl (pegRNA only), TJ-PE (PEmax), TJ-PE + eGFP (protein control), TJ-PE + Genl-SSAP (e.g., RecT), and TJ-PE + Gen2- SSAP (Twinkle or erf editor). MCP-SSAP and PE-TJ plasmids were co-transfected into HEK293T cells. Cells were harvested 72 h later and genomic DNA was extracted. Target loci were PCR-amplified, products resolved by agarose gel electrophoresis, and deletion / editing efficiency was calculated by densitometry of DNA bands in Bio-Rad Image Lab.
[0119] Adding Genl-SSAP increases PE efficiency over PEmax and eGFP control. Gen2-SSAP (Twinkle editor) yields the highest editing efficiency (~9-10%), exceeding all other conditions. MS2 hairpins are important for the gain: without MS2 loops, SSAP recruitment provides little / no benefit; with MS2+ pegRNA, RecT / Erf / Twinkle show ~1.7-2.2x fold-improvements versus their MS2- counterparts.
[0120] Recruiting SSAPs to Prime Editing via MS2-pegRNA scaffolds produces a robust, generalizable boost in PE efficiency (best with Twinkle-based SSAPs).
[0121] Example 7
[0122] This example describes improved knock-in efficiency using a Cas-free SSAP system (FIG. 8). Conventional CRISPR systems (Cas9 / dCas9) combined with SSAPs and donors impose delivery challenges due to their large size. To overcome this, a CRISPR-free SSAP system was designed by removing Cas9 / dCas9 and instead using an R-loop guiding RNA (RL- STDU2-43469.601 gRNA). The RL-gRNA recruits SSAP directly to genomic targets, facilitating R-loop formation and donor integration without DNA cleavage. Multiple donor lengths (18 nt, 20 nt, 25 nt homology arms) were tested at two endogenous loci (HSP90AA1 and ACTB). Donor, gRNA and SSAP were co-transfected to HEK293T cell, 72h post transfection cell were harvested by FACS analysis.
[0123] In the absence of SSAP, donor-only conditions showed very low knock-in efficiency (~0.3-0.6%). With SSAP + donor, knock-in efficiency increased dramatically (~6-9%) across different donor arm lengths. Novel SSAPs (Erf, Twinkle, SSAP16) consistently outperformed RecT, achieving efficient knock-in at both HSP90AA1 and ACTB loci. This demonstrates that SSAP proteins alone, guided by RL-gRNA, can drive cleavage-free knock-in.
[0124] Example 8
[0125] This example describes Al fitness prediction and SSAP model fine-tuning. To systematically improve SSAP engineering, Al-based protein fitness prediction benchmarks (ProteinGym) was leveraged, which evaluate models across deep mutational scanning (DMS) datasets and clinical mutation datasets. The ESM-2 protein language model (3B parameters) was selected as the base sequence-input model for fine-tuning. The in-house experimental dataset comprises 350 SSAP valiants screened in array assays, providing quantitative knock-in efficiency data across multiple loci. These experimental results were used to fine-tune ESM-2, creating a model specialized for predicting SSAP activity (see e.g., FIG. 9).
[0126] The in-house dataset comprised 350 SSAP sequences, each measured by array assay for knock-in efficiency. To integrate these experimental data with pre-trained protein language models, data transformation was applied to normalize the values. The goal was to ensure the SSAP dataset aligns with the distribution of large-scale mutational benchmarks (e.g., ProteinGym DMS datasets), making it suitable for fine-tuning. After transformation, the SSAP dataset fits a distribution similar to large mutational datasets, facilitating direct comparability and integration. This normalization step ensures that fine-tuning of the ESM-2 protein language model is robust and unbiased, avoiding artifacts from scale mismatches.
[0127] The pre-trained ESM-2 protein language model was fine-tuned using the 350 SSAP variant dataset with measured knock-in efficiencies. A feed- forward neural network (FNN) was STDU2-43469.601 constructed on top of ESM-2 embeddings to map sequence features to functional outputs. The optimized architecture consisted of: Linear layer (2560 — >■ 256); Dropout (0.2); LcakyRcLU activation; Linear layer (256 — > 20); Total parameters: ~660k.
[0128] A training loss curve showed rapid initial convergence followed by gradual improvement, reaching a stable low-loss regime after ~15M training steps. This indicates that the fine-tuned model successfully learned to associate SSAP sequence variation with functional performance. The compact parameter size (660k) makes the model computationally efficient while maintaining predictive power.
[0129] Based on fine-tuned SSAP prediction models, key residues in Twinkle were selected that are involved in DNA binding and oligomerization.
[0130] A panel of Twinkle mutant (71) variants was synthesized, generated by the Al model. Each mutant was tested by co-transfection with dCas9-gRNA and either circular ssDNA (cssDNA) or dsDNA donors. Knock-in efficiency was measured at three genomic loci: B2M, RAB11A, and ACTB (FIG. 10).
[0131] B2M locus (cssDNA donor): Multiple Twinkle mutants showed dramatic improvements, with the best variants reaching >90% knock-in efficiency, far surpassing wild-type Twinkle.
[0132] RAB11A locus (cssDNA donor): Knock-in efficiency improved moderately, with several variants outperforming wild-type.
[0133] ACTB locus (dsDNA donor): Mutants achieved incremental gains in knock-in efficiency compared to baseline Twinkle.
[0134] Experimental Design
[0135] Mutant Erf variants were designed by fine-tuned SSAP Al model, then synthesized, and cloned into expression constructs. Variants were co-transfected with dCas9-gRNA and ssDNA or dsDNA donors. Validation was performed across three genomic loci: B2M, RAB11A, ACTB (FIG 11).
[0136] B2M locus (cssDNA donor): Multiple Erf variants demonstrated exceptional knock-in efficiencies, with top variants achieving nearly 95% efficiency, outperforming wild-type Erf. STDU2-43469.601
[0137] RABI 1 A locus: Several Erf mutants showed improved efficiencies (6-10%), exceeding baseline levels.
[0138] ACTB locus (dsDNA donor): 1 fold KI efficiency improvements were observed across engineered variants. Using the fine-tuned SSAP model, multi-site mutants of Erf and Twinkle were designed.
[0139] Variants were cloned and co-transfected with dCas9 / gRNA and a circular ssDNA (cssDNA) donor into mammalian cells. Two donor doses were tested: 30 ng and 50 ng cssDNA. 72 h posttransfection, FACS quantified knock-in (% reporter-positive cells).
[0140] Across both doses, multiple Erf and Twinkle variants exceeded the wild-type baseline, confirming the benefit of Al-guided multi-site design. SSAP-mediated gains are more pronounced at the lower donor dose (30 ng).
Claims
STDU2-43469.601CLAIMS1. A method for modifying a cell, comprising: introducing into a cell, a) a DNA or RNA-converted DNA donor; b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same; and, optionally, c) a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system).
2. The method of claim 1, wherein the eukaryotic SSAP (e.g., helicase) is from the same species of organism as the subject.
3. The method of claim 1, wherein the cell is modified in vitro.
4. The method of claim 1, wherein the cell is modified ex vivo.
5. The method of claim 1, wherein the cell is modified in vivo.
6. The method of claim 1, wherein the DNA or RNA-converted DNA donor comprises an insertion sequence of 1000 base pairs or more.
7. The method of claim 1, wherein the DNA or RNA-converted DNA donor comprises an insertion sequence of 5000 base pairs or more.
8. A system comprising: a) a DNA or RNA-converted DNA donor; b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same; and, optionally, c) a nucleic acidSTDU2-43469.601 targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system).
9. A system comprising: a) a pegRNA comprising an MS2 hairpin; and b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same.
10. A system comprising: a) a prime editor or a sequence encoding the same; and b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same.
11. Use of a system of any of claims 8-10.
12. Use of a system of any of claims 8-10, for the modification of a cell.
13. A method of editing a nucleic acid molecule comprising: contacting a nucleic acid molecule with: a) a DNA or RNA-converted DNA donor; b) a eukaryotic SSAP (e.g., helicase) or a sequence encoding the same; and, optionally, c) a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system).
14. A non-natural helicase comprising SEQ ID NOU, or a sequencing having at least 75% sequence identity with SEQ ID NOU and having one or more amino acid substitutions at positions selected from positions G24, D127, C247, F289, F308, and Q425.
15. The non-natural helicase of claim 14, wherein the one or more amino acid substitutions comprise G24K, D127E, C247T, F289Y, F3O8R, and Q425N.STDU2-43469.60116. The non-natural helicase of claim 14, comprising amino acid substitutions: G24K and at least one or all of F289Y, D127E, and C247T.
17. The non-natural helicase of claim 14, comprising amino acid substitutions: G24K; at least one or all of F289Y, D127E, and C247T; and one or both of F308R and Q425N.
18. A non-natural helicase comprising SEQ ID NOG, or a sequencing having at least 75% sequence identity with SEQ ID NOG and having one or more amino acid substitutions at positions selected from positions 174, A75, E94, Al 10, S169, and M140.
19. The non-natural helicase of claim 18, wherein the one or more amino acid substitutions comprise I74R, A75V, E94R, A110R, S169K or S169Q, and M140K.
20. The non-natural helicase of claim 18, comprising amino acid substitutions: M140K; S169Q; and one or more of: I74R, A75V, and Al 10R.
21. The non-natural helicase of claim 18, comprising amino acid substitutions: M140K; S169Q; A75V; and one or both of I74R and A110R.
22. The non-natural helicase of claim 18, comprising amino acid substitutions: M140K; S169K; and one or more of: I74R, A75V, E94R, and Al 10R.
23. The non-natural helicase of claim 18, comprising amino acid substitutions: M140K; S169K; and one or both of A75V and E94R.
24. The non-natural helicase of claim 18, comprising amino acid substitutions: M140K; S169K; one or both of A75V and E94R; and one or both of Al 10R and I74R.STDU2-43469.60125. A composition comprising a non-natural helicase of any of claims 14-24.
26. A cell comprising a non-natural helicase of any of claims 14-24.
27. An expression vector encoding a sequence of a non-natural helicase of any of claims 14-24.
28. A system comprising a non-natural helicase of any of claims 14-24 and a DNA or RNA-converted DNA donor and / or a nucleic acid targeting system (e.g., CRISPR system comprising a Cas protein, a guide, an R-loop guiding RNA, a prime editing system).
29. Use of a helicase, composition, cell, expression vector or system of any of claims 14-28.
30. Use of a helicase, composition, cell, expression vector or system of any of claims 14-28 to insert said donor into a heterologous nucleic acid (e.g., a genome of a cell).
Citation Information
Patent Citations
RNA-guided genome recombineering at kilobase scale
WO2023034925A1
Precise genome editing using retrons
WO2023081756A1
RNA-guided genome recombineering at kilobase scale
WO2023154877A2
Methods and systems for generating nucleic acid diversity in crispr-associated genes
WO2024038003A1
Delivery of an RNA guided recombination system
WO2024168253A1