Cas enabled targeted transposition
A mutated transposase linked to a DNA sequence-targeting protein facilitates the integration of large DNA sequences into the genome, overcoming the limitations of current genome editing technologies by enhancing integration efficiency and reducing off-target effects.
Patent Information
- Application Number
- PCT/US2025/022727
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-03
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-09
AI Technical Summary
Current genome editing technologies are limited by the length of DNA that can undergo homology-directed recombination, typically less than 2 kilobases, making it difficult to insert large DNA sequences, such as those spanning entire genes, which can exceed 2 kilobases in length.
A polypeptide comprising a mutated transposase, linked to a DNA sequence-targeting protein, is used to introduce large DNA sequences into the genome, with specific mutations at positions H165, H187, 1212, and K252 of the transposase improving integration capacity.
Enables the efficient integration of DNA sequences greater than 10 kilobases into the genome, reducing off-target effects and maintaining cell viability.
Smart Images

Figure US2025022727_09102025_PF_FP_ABST
Abstract
Description
CAS ENABLED TARGETED TRANSPOSITIONCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit of priority to U.S. Provisional Application No. 63 / 573,990 filed April 3, 2024. The entire contents of the aforementioned application are incorporated herein by reference in its entirety.STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT
[0002] This invention was made with government support under Grant No. AG068908 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND OF THE INVENTION
[0003] Genome editing technology, particularly CRISPR, has revolutionized biological research, holds therapeutic promise for many diseases, and has broad reaching implications across many disciplines, including agriculture, personalized medicine, ecology, and biodiversity. However, one key limitation of current genome editing technology is the length of DNA (typically less than 2 kilobases) that can undergo homology-directed recombination into a target genome in a guide RNA-dependent, site-directed manner. Often, genetic diseases benefitted by genome editing technology have mutations that span entire genes, which can far exceed 2 kilobases in length. Thus, genome editing tools for insertion of large DNA sequences are desired.BRIEF SUMMARY OF THE INVENTION
[0004] In some embodiments, provided herein is a polypeptide comprising a transposase at least 90% identical to SEQ ID NO:1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype. In some embodiments, the transposase is at least 80, 70, 85, 90, 95, 98, 99, or 100% identical to SEQ ID NO:1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
[0005] In some embodiments, the transposase comprises SEQ ID NO:1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO: 2 is not wildtype.
[0006] In some embodiments, provided herein is a polypeptide comprising a transposase linked to a DNA sequence-targeting protein, wherein the transposase is at least 90% identical to SEQ ID NO: 1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO:1 or SEQ ID NO:2 is not wildtype. In some embodiments, the polypeptide comprises SEQ ID NO:5. In some embodiments, the transposase is at least 80, 70, 85, 90, 95, 98, 99, or 100% identical to SEQ ID NO:1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype. In some embodiments, the transposase comprises SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO:1 or SEQ ID NO:2 is not wildtype.
[0007] In some embodiments, the DNA sequence-targeting protein is a CRISPR / Cas protein, a TALEN protein, a zinc finger protein or a meganuclease protein. In some embodiments, the DNA sequence-targeting protein is a Casl2a or Cas9 protein. In some embodiments, the DNA sequence-targeting protein has nicking activity. In some embodiments, the DNA sequencetargeting protein lacks nicking activity. In some embodiments, the DNA sequence-targeting protein is dCas!2a or dCas9. In some embodiments, the dCas!2a protein comprises SEQ ID NO:3.
[0008] In some embodiments, the polypeptide further comprises one or more nuclear localization signal (NLS) sequence.
[0009] In some embodiments, the transposase linked to the DNA sequence-targeting protein is a translational fusion, optionally linked via a peptide linker. In some embodiments, the transposase is fused to a first affinity agent and the DNA sequence-targeting protein is linked to a second affinity agent and the first and second affinity agents bind. In some embodiments, the first and second affinity agents bind in the presence of a chemical agent but not in the absence of the chemical agent. In some embodiments, the first and second affinity agents are selected from FRB and FKBP and the chemical agent is rapamycin.
[0010] In some embodiments, also provided herein is a polynucleotide encoding the polypeptide described herein. In some embodiments, an expression cassette comprising a promoter operably linked to the polynucleotide is provided. In some embodiments, a vector comprising the expression cassette is provided. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a lentiviral, adenoviral, or adeno-associated viral vector. In some embodiments, the lentiviral vector is a non-integrating lentiviral vector.
[0011] In some embodiments, also provided herein is a method of editing the genome of a eukaryotic cell, the method comprising, delivering the polypeptide described herein to the nucleus of the cell, wherein the polypeptide nicks the genome of the eukaryotic cell in at least one location and introduces a template polynucleotide into the nick, thereby editing the genome of the cell. In some embodiments, delivering the polypeptide comprises contacting the cell with a vector encoding the polypeptide, and wherein the polypeptide is expressed in the cell and enters the nucleus of the cell. In some embodiments, the method is performed in vitro, ex vivo or in vivo.
[0012] In some embodiments, the method further comprises delivering the template polynucleotide to the cell. In some embodiments, the template polynucleotide is delivered to the cell in a viral vector. In some embodiments, the viral vector is a lentiviral, adenoviral, or adeno- associated viral vector. In some embodiments, the template polynucleotide is at least 10, 20, 30, 40, 50, 70, or 100 kb long.
[0013] In some embodiments of the method provided herein, the polypeptide comprises a transposase linked to a DNA sequence-targeting protein, wherein the transposase is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype. In some embodiments, the transposase is at least 80, 70, 85, 90, 95, 98, 99, or 100% identical to SEQ ID NO:1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
[0014] In some embodiments, the DNA sequence-targeting protein is a CRISPR / Cas protein, and the method further comprises delivering one or more guide RNAs (gRNAs) to the cell. In some embodiments, two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) gRNAs are delivered to thecell. In some embodiments, the DNA sequence-targeting protein is a Casl2a or Cas9 protein. In some embodiments, the DNA sequence-targeting protein has nicking activity. In some embodiments, the DNA sequence-targeting protein lacks nicking activity. In some embodiments, the DNA sequence-targeting protein is dCasl2a or dCas9. In some embodiments, the DNA sequence-targeting protein comprises SEQ ID NO:3.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The present application includes the following figures. The figures are intended to illustrate certain embodiments and / or features of the compositions and methods, and to supplement any description(s) of the compositions and methods. The figures do not limit the scope of the compositions and methods, unless the written description expressly indicates that such is the case.
[0016] FIG. 1 shows a schematic of the polypeptide and genome editing system described herein.
[0017] FIG. 2 shows a schematic of the polypeptide comprising a first and second affinity agent described herein.
[0018] FIG. 3 shows modulation of transposition by the polypeptide comprising a first and second affinity agent described herein in the presence and absence of the A / C dimerizer.
[0019] FIG. 4 shows a schematic of genome editing of the pl6 murine locus using the polypeptides and methods described herein.
[0020] FIG. 5A shows the transposase variants discovered by next generation sequencing and confirmed by site saturation mutagenesis.
[0021] FIG. 5B shows the frequencies of all the amino acids following directed evolution.
[0022] FIG. 6 shows the efficiency and specificity of transposition of the polypeptide described herein using puro as a component of the transposon.
[0023] FIG. 7A shows puromycin resistance as a measure of transposition efficiency of the polypeptide described herein.
[0024] FIG. 7B shows 6-TG resistance caused by a disruption in the HRPT locus as a measure of transposition specificity using the polypeptide described herein with a gRNA array targeting the HRPT locus.DEFINITIONS
[0025] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural reference unless the context clearly dictates otherwise.
[0026] The use of any and all examples or exemplary language (e.g., “for example” or “such as”) provided herein, is intended merely to better illustrate the invention, and does not pose a limitation on the scope of the invention unless otherwise claimed.
[0027] The terms “may,” “may be,” “can,” and “can be,” and related terms are intended to convey that the subject matter involved is optional (that is, the subject matter is present in some examples and is not present in other examples), not a reference to a capability of the subject matter or to a probability, unless the context clearly indicates otherwise.
[0028] The use herein of the terms "including," "comprising," or "having," and variations thereof, is meant to encompass the elements listed thereafter and equivalents thereof as well as additional elements. Embodiments recited as "including," "comprising,” or "having" certain elements are also contemplated as "consisting essentially of and "consisting of’ those certain elements. As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations where interpreted in the alternative (“or”).
[0029] “Polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.
[0030] The term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y-carboxy glutamate,and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Naturally encoded amino acids are the 20 common amino acids (alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine) as well as pyrrolysine, pyrroline-carboxy-lysine, and selenocysteine.
[0031] In the present application, amino acid residues are numbered according to their relative positions from the left most residue, which is numbered 1, in an unmodified wildtype polypeptide sequence.
[0032] The term “corresponding to” in the context of nucleic acid sequences means that when the nucleic acid sequences of certain sequences are aligned with each other, the nucleic acids that “correspond to” certain enumerated positions in the present invention are those that align with these positions in a reference sequence. Optimal alignment of sequences for comparison can be conducted by computerized implementations of known methods (e.g., BLAST as described below).
[0033] The term “nucleic acid” or “nucleotide” refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof, herein, “polynucleotides,” in either single- or double-stranded form. The nucleic acid molecule and polynucleotides may be derived from a variety of sources, including DNA, cDNA, synthetic DNA, RNA, or combinations thereof. Such nucleic acid sequences may comprise genomic DNA which may or may not include naturally occurring introns. Moreover, such genomic DNA may be obtained in association with promoter regions, introns, or poly A sequences.
[0034] The term “genome” is used to refer to the entire body of genetic material (i.e., nucleic acids or polynucleotides as described above) contained in each cell of an organism. As used herein, a “gene” refers to a defined region that is located within a genome and that comprises other, primarily regulatory, nucleic acid sequences responsible for the control of the expression,that is to say the transcription and translation, of the coding portion. Genes can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or specific protein, including regulatory sequences. Genes may or may not be capable of being used to produce a functional protein. In some embodiments, a gene refers to only the coding region, (i.e., the region which ultimately results in protein expression).
[0035] The term “heterologous” refers to a polynucleotide sequence or polypeptide that originates from a foreign species, or, if from the same species, is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, polynucleotide sequence is “heterologous to” an organism if it originates from a foreign species, or, if from the same species, is modified from its original form. In the context of fusion proteins, a heterologous fusion partner polypeptide is one which in its native form, does not form a fusion protein with the other partner polypeptide in the fusion protein.
[0036] As used herein, a “cell” can be in vivo, ex vivo or in vitro, and includes any eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.
[0037] The “CRISPR / Cas” system refers to a widespread class of bacterial systems for defense against foreign nucleic acid. CRISPR / Cas systems are found in a wide range of eubacterial and archaeal organisms. CRISPR / Cas systems include type I, II, and III sub-types. Wild-type type II CRISPR / Cas systems utilize an RNA-mediated nuclease, a Cas, in complex with guide RNAs (gRNAs) and activating RNAs to recognize and cleave foreign nucleic acid. gRNAs having the activity of both a gRNA and an activating RNA are also known in the art. In some cases, such dual activity guide RNAs are referred to as small guide RNAs (sgRNAs).
[0038] As used herein, the term “Cas” refers to CRISPR associated proteins that are RNA- mediated nucleases (e.g., of bacterial or archeal orgin, or derived therefrom). Exemplary Cas types include but are not limited to Cas9 proteins and homologs thereof or Cas 12a proteins and homologs thereof. A Cas may be wildtype or in other instances, the Cas is non-natural, artificial, engineered, synthetic, rationally designed, or man-made. A Cas used herein may be mutated (e.g., dCas9 or lbCasl2a RuvC mutant). A Cas may be a nuclease (an enzyme that cleaves both strands of a double-stranded nucleic acid), a nickase (an enzyme that cleaves one strand of adouble-stranded nucleic acid), or a catalytically inactive (or dead) Cas. By way of example, a Cas having nuclease or nickase activity is referred to as a “catalytically active Cas.” A Cas lacking the ability to cleave or nick target nucleic acid is referred to as a “catalytically inactive Cas9 ” or a “dead Cas” (e.g., a dCas9 or lbCasl2a RuvC mutant).
[0039] The Cas may be derived from any species known in the art. For example, Casl2a may be from Acidaminococcus sp. and Lachnospiraceae bacterium (“lbCasl2a”). Cas9 may be derived from bacteria of the following taxonomic groups: Actinobacteria, Aquificae, Bacteroidetes-Chlorobi, Chlamydiae-Verrucomicrobia, Chlroflexi, Cyanobacteria, Firmicutes, Proteobacteria, Spirochaetes, and Thermotogae. An exemplary Cas9 protein is the Streptococcus pyogenes Cas9 protein. Additional Cas9 proteins and homologs thereof are described in, e.g., Chylinksi, et al., RNA Biol., 10(5): 726-737 (2013); Nat. Rev. Microbiol., 9(6): 467-477 (2011); Hou, et al., Proc Natl Acad Sci USA., 110(39): 15644-9 (2013); Sampson et al., Nature, 497(7448):254-7 (2013); and Jinek, et al., Science, 337(6096):816-21 (2012). The Cas nuclease domain can be optimized for efficient activity or enhanced stability in the target cell.
[0040] “Percentage of sequence identity” is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the amino acid sequence or polynucleotide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (e.g., SEQ ID NO: 1 or 2), which does not comprise additions or deletions, for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
[0041] The terms “identical” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same sequences. Two sequences are “substantially identical” if two sequences have a specified percentage of amino acid residues or nucleotides that are the same (i.e., 95% identity, optionally 96%, 97%, 98%, or 99% identity over a specified region, or, when not specified, over the entire sequence), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparisonalgorithms or by manual alignment and visual inspection. For an amino acid sequence, optionally, identity exists over a region that is at least about 50 amino acids in length, or more preferably over a region that is 100 to 150 or 200 or more amino acids in length, or where not indicated over the entire length of the reference sequence.
[0042] For sequence comparison, typically one sequence acts as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
[0043] A “comparison window”, as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 50 to 600, usually about 75 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well known in the art.
[0044] An algorithm for determining percent sequence identity and sequence similarity is the BLAST 2.0 algorithms, e.g., as described in, and Altschul et al. J. Mol. Biol. 215:403-410 (1990) (see also Altschul et al. Nuc. Acids Res. 25:3389-3402 (1977)) . Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrixis used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) or 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1989)) alignments (B) of 50, expectation (E) of 10, M=5, and N=-4.
[0045] The term “affinity agent” as used herein refers to a protein sequence or other agent that has specific affinity (specifically binds) to a second affinity agent such that the first and second affinity agent covalently or non-covalently bind. An affinity agent can be any protein known or selected to have specific affinity for a second amino acid sequence. Examples of affinity agents include but are not limited to FK506 binding protein (FKBP) and FKBP-rapamycin binding (FRB) domain in the FKBP-rapamycin-associated protein. FKBP and FRB bind in the presence of rapamycin (see, e.g., Banaszynski et al., J. Am. Chem. Soc., 127(13):4715-21 (2005)).
[0046] A “promoter” is defined as one or more a nucleic acid control sequences that direct transcription of a nucleic acid. As used herein, a promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of a polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements, which can be located as much as several thousand base pairs from the start site of transcription.
[0047] A promoter is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence, for example, a polynucleotide encoding a polypeptide provided herein. For example, a promoter is operably linked to a polynucleotide if it affects, either positively or negatively, the transcription of the polynucleotide.
[0048] The term "vector," as used herein, refers to a nucleic acid molecule capable of propagating another nucleic acid to which it is linked. The term includes the vector as a selfreplicating nucleic acid structure as well as the vector incorporated into the genome of a targetcell into which it has been optionally introduced. A “vector” as used here refers to a recombinant construct in which a nucleic acid sequence of interest is inserted into the vector.
[0049] The term “template polynucleotide” as used refers to a polynucleotide which is being targeted for insertion, e g., by a polypeptide described herein, into a target genomic region. The template polynucleotide used in the present disclosure may be at least 5, 10, 20, 30, 40, 50, 70, or 100 kilobases (kb) long. In some embodiments, the template polynucleotide is about 100-5,000 kb long. In some embodiments, a template polynucleotide comprises a coding sequence for a polypeptide or RNA, one or more gene, or a promoter or enhancer region operably linked to the one or more genes. In some embodiments, a template polynucleotide may comprise one or more genes, e.g., two genes, three genes, four genes, or five genes, or promoters or enhancer regions operably linked to one or more genes. In some embodiments, a template polynucleotide may comprise a promoter region, or control region, of a gene. In some embodiments, a template polynucleotide may comprise an intron of a gene. In some embodiments, a template polynucleotide may comprise an exon of a gene.
[0050] As used herein, the phrase “editing” in the context of editing of a genome of a cell refers to inducing a structural change in the sequence of the genome at a target genomic region. For example, the editing can take the form of inserting a nucleotide sequence, for example a template polynucleotide, into the genome of the cell. The nucleotide sequence can encode a polypeptide or a fragment thereof. Such editing can be performed by inducing a double stranded break within a target genomic region, or a pair of single stranded nicks on opposite strands and flanking the target genomic region. Methods for inducing single or double stranded breaks at or within a target genomic region include the use of a transposase or a Cas9 or Casl2a nuclease domain, or a derivative thereof, and a guide RNA, or one or more guide RNAs, directed to the target genomic region.
[0051] The term “target” as used herein refers to the recipient cell of the template polynucleotide. The term “target genomic region” is used herein to refer to the genomic site at which integration of the template polynucleotide is intended to occur or does occur.DETAILED DESCRIPTION OF THE INVENTIONI, INTRODUCTION
[0052] Provided herein are compositions and methods for editing the genome of a cell. The polypeptides and / or fusion proteins described herein can comprise a mutated transposase. The mutations at certain residues in the transposase described herein improve integration handling capacity or other activity of the transposase. The fusion proteins described herein can further comprise a DNA sequence targeting protein, which, when combined with the transposase, which can optionally be mutated, results in precise genomic targeting of large template polynucleotides. The inventor has surprisingly discovered that large nucleotide sequences, for example, nucleotide sequences greater than about 10 kilobase nucleotides or base pairs in length, can be inserted into the genome of a cell using the polypeptides or fusion proteins described herein.
[0053] Integration of large nucleic acids, for example nucleic acids greater than 10 kilobase nucleotides or base pairs in size, into cells, can be limited by low efficiency of integration, off- target effects and / or loss of target cell viability. Described herein are methods and compositions for achieving integration of a nucleotide sequence, for example, a nucleotide sequence greater than about 10 kilobase nucleotides or base pairs in size, into the genome of a cell. In some methods the efficiency of integration is increased, off-target effects are reduced and / or loss of cell viability is reduced.
[0054] The following description recites various aspects and embodiments of the present compositions and methods. No particular embodiment is intended to define the scope of the compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that are at least included within the scope of the disclosed compositions and methods. The description is to be read from the perspective of one of ordinary skill in the art; therefore, information well known to the skilled artisan is not necessarily included.II, COMPOSITIONS
[0055] This disclosure provides for polypeptides comprising a transpose, optionally mutated as described herein, and / or optionally linked to a DNA sequence-targeting protein.A. Transposases
[0056] Provided herein are polypeptides comprising a transposase. Transposase refers to a protein having transposase activity. Transposases are a class of enzymes that mediate transposition of a transposon (or genes of interest in genome editing or gene engineering settings) from one location in the genome to another. Transposons are movable DNA elements which encode a transposase enzyme to catalyze their excision and integration in an alternate genomic location. Transposases typically induce double strand breaks to excise the transposon, recognize subterminal repeats at a target site, and integrate the transposon or gene of interest into an alternate location. Transposase activity can therefore include one or more of the aforementioned attributes.
[0057] The Tcl / mariner superfamily is a group of Tc-1 and mariner transposon encoding a transposase. Tcl / mariner transposons are likely the most widespread transposition systems in nature (see, e.g., Dupeyron et al., Mobile DNA, 11(21): 1-14 (2020)). Naturally occurring Tel / mariner-like transposase systems are generally non-functional due to long-accumulated inactivating mutations. Sleeping Beauty (SB) is a synthetic transposition system reconstructed from multiple inactive fish Tc-1 like transposons. SB100X is a hyperactive SB ; The transposon used in SB100X encodes a gene of interest (analogous to a template polynucleotide described herein), instead of a transposase. The transposase used in SBIOOx excises the gene of interest, and aids in its transposition into the target genomic region and ultimate expression by the target cell, (see, e.g., Jin et al., Gene Ther. 18(9): 849-856 (2011). Active fragments or mutants of SBIOOX transposase as described below and elsewhere, may be used in the compositions of the present disclosure. The SBIOOX transposase, as used herein, may be full-length (SEQ ID NO: 1) or truncated (SEQ ID NO: 2). The inventor has discovered novel SBIOOX transposase mutants that improve insertion capacity of, for example, template polynucleotides at least 10 kb in length as described herein. Amino acid residues corresponding to wildtype positions H165, H187, 1212, and K252 of full-length or truncated SBIOOX transposase may be mutated (for example, Hl 65 may be replaced with aspartic acid, leucine, or valine; Hl 87 may be replaced with aspartic acid, leucine, or glutamine, 1212 may be replaced with serine or asparagine, or K252 may be replaced with asparagine). The mutant SBIOOX transposase as described herein may be used alone where transposase activity is desired, or as a fusion with other components, for example, with a DNA sequence-targeting protein as described herein.
[0058] In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype. In some embodiments, the transposase comprises SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
[0059] In some embodiments, at least one position corresponding to the wildtype positions H165 (i.e., histidine at amino acid position 165 of SEQ ID NOs: 1 or 2), H187 (i.e., histidine at amino acid position 187 of SEQ ID NO: 1 or 2), 1212 (i.e., isoleucine at amino acid position 212 of SEQ ID NOs: 1 or 2), and K252 (i.e., lysine at amino acid position 212 of SEQ ID NOs: 1 or 2) of SEQ ID NO: 1 or SEQ ID NO:2 in the transposase described herein is not wildtype. Either 1, 2, 3, or all 4, of the positions corresponding to the wildtype positions Hl 65, Hl 87, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NOs:2 are not wildtype. The at least one position corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 may be substituted with any amino acid (e.g., alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine). In some embodiments, the position corresponding to the wildtype position H165 of SEQ ID NOs: 1 or 2 is aspartic acid, leucine, or valine. In some embodiments, the transposase comprises an H165L mutation relative to SEQ ID NOs: 1 or 2. In some embodiments, the transposase comprises SEQ ID NO:9. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO:9. In some embodiments, the transposase comprises an H165D mutation relative to SEQ ID NOs: 1 or 2. In some embodiments, the transposase comprises SEQ ID NOTO. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NOTO. In some embodiments, the transposase comprises an H165V mutation relative to SEQ ID NOs: 1 or 2. In some embodiments, the transposase comprises SEQ ID NO: 11. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO: 11.
[0060] In some embodiments, the position corresponding to the wildtype position Hl 87 of SEQ ID NOs: 1 or 2 is aspartic acid, leucine, or glutamine. In some embodiments, the positioncorresponding to the wildtype position 1212 of SEQ ID NOs: 1 or 2 is serine or asparagine. In some embodiments, the transposase comprises a K252N mutation relative to SEQ ID NOs: 1 or 2. In some embodiments, the the transposase comprises SEQ ID NO: 12. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO: 12. In some embodiments, the transposase comprises a K252Y mutation relative to SEQ ID NOs: 1 or 2. In some embodiments, the transposase comprises SEQ ID NO: 13. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO: 13. In some embodiments, the position corresponding to the wildtype position K252 of SEQ ID NOs: 1 or 2 is asparagine. In some embodiments, the position corresponding to the wildtype position K252 of SEQ ID NOs: 1 or 2 is an early stop codon.
[0061] The above-described mutations may be present at 0, 1, 2, 3, or 4 of the above-described residues of SEQ ID NOs: 1 or 2 (i.e., H165, H187, 1212, and K252). For example, each of the 4 residues may be wildtype (i.e., not be mutated) or 2 of the residues may be mutated (e.g., H165 may be replaced with leucine and H187 may be replaced with glutamine, but 1212 and K252 may correspond to wildtype residues).B. Fusion Proteins
[0062] In some embodiments, the transposase, for example as described above, is provided as part of a fusion protein with a heterologous fusion partner polypeptide. An exemplary heterologous fusion partner polypeptide can be, for example, a DNA sequence-targeting protein. Thus, also provided herein is a polypeptide comprising a transposase as described herein linked to a DNA sequence-targeting protein. In some embodiments, the polypeptide comprises the transposase linked to the DNA sequence-targeting protein as a single translational fusion protein. i. DNA Sequence-Targeting Proteins
[0063] As used throughout, a DNA sequence-targeting protein is a protein with nucleic acid sequence targeting activity, e.g., a protein which directs a transposase described herein to a DNA target sequence on the target genomic region. In some embodiments, a DNA sequence-targeting protein is capable of targeting a designated nucleotide or region within the target genomic region. In some embodiments, the DNA sequence-targeting protein is capable of targeting a region positioned between the 5' and 3' regions of the target genomic region. In some embodiments, the DNA sequence-targeting protein is capable of targeting a region positionedupstream or downstream of the 5' and 3' regions of the target genomic region (e.g., upstream or downstream of the transcription start site (TSS)). A recognition sequence is a polynucleotide sequence that is specifically recognized and / or bound by the DNA sequence-targeting protein. The length of the recognition site sequence can vary, and includes, for example, nucleotide sequences that are at least 10, 12, 14, 16, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70 or more nucleotides in length. In some embodiments, the recognition sequence is palindromic, i.e., the sequence on one DNA strand reads the same in the opposite direction on the complementary DNA strand. In some embodiments, the target genomic region of the DNA sequence-targeting protein is within the recognition sequence. In some embodiments, a guide RNA, as described below, is used to direct the DNA sequence-targeting protein to the target genomic region.
[0064] In some embodiments the DNA sequence-targeting protein has nicking (i.e., cleavage) activity and can induce single- or double-stranded breaks in the target genomic region. In some embodiments, the DNA sequence-targeting protein lacks nicking activity and single-or doublestranded breaks are induced by another molecule, for example, a transposase described herein. Exemplary DNA sequence-targeting proteins include CRISPR-Cas proteins, TALEN proteins, zinc finger proteins, meganuclease proteins, or promoter repressor elements. a. CRISPR-Cas Proteins
[0065] The DNA sequence-targeting protein used herein can be a CRISPR / Cas protein, i.e., a Cas as described above and elsewhere herein. Exemplary Cas include but are not limited to Cas5, Cas6, Cas7, Cas8, Cas9, Casl2a, Casl2b, Casl2i, Casl2j, Casl2L, Casl2e, Casl2c, Casl2d, Casl2g, Casl2h, TnpB, or Casl4 and those described in, for example, Xi and Li, Comput. Struct. Biotechnol. J., 18:2401-2415 (2020). The Cas used herein may have nicking activity, such that when bound to target nucleic acid as part of a complex with a guide RNA, a single strand break or nick is introduced into the target nucleic acid.
[0066] A Cas used herein can be an active variant, inactive variant, or fragment of a wild-type or modified Cas. A Cas can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wildtype version of the Cas. A Cas can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%sequence identity or sequence similarity to a wild-type exemplary Cas (i.e., the Cas listed above). A Cas can be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas. Variants or fragments can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas or a portion thereof. Variants or fragments can be targeted to a target genomic region in complex with a guide nucleic acid (e.g., a guide RNA) while lacking nucleic acid cleavage activity.
[0067] In some embodiments, a modified Cas has decreased function relative to the unmodified form. In some embodiments, a modified Cas is deficient in a function of the unmodified form. For example, a nuclease deficient Cas retains the ability to bind DNA but lacks or has reduced nucleic acid cleavage activity. A Cas nuclease (e.g., retaining wild-type nuclease activity, having reduced nuclease activity, and / or lacking nuclease activity) can function in a CRISPR / Cas system to regulate the level and / or activity of a target gene or protein (e.g., decrease, increase, or elimination). The Cas can bind to a target polynucleotide and prevent transcription by physical obstruction or edit a nucleic acid sequence to yield non-functional gene products. In some embodiments, the modified Cas has no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the function (e.g., nuclease activity) of the wild-type Cas (e.g., Cas9 from S. pyogenes). In some embodiments, the modified Cas has no substantial function of the wild-type Cas. When a Cas is a modified form that has no substantial nucleic acid-cleaving activity, it can be referred to as enzymatically inactive and / or “dead” (abbreviated by “d”). A dead Cas (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some aspects, a dead Cas is a dead Cas9 or a dead Casl2a (i.e., lbCas!2a).
[0068] In some embodiments, the Cas used herein is a Casl2a or a Cas9. Casl2a (i.e., Cpfl) is a Cas appearing in many bacterial species, e g., Acidaminococcus and Lachnospiraceae. Cas9 is derived from Streptococcus pyogenes. Both Casl2a and Cas9 have endonuclease domains (e.g., RuvC in Cas 12a and HNH and RuVCl in Cas9). In some embodiments, the Cas used herein lacks nicking activity due to, for example, a mutation to the endonuclease domain (e g., dCas9 orlbCasl2a RuvC mutant). The lbCas!2a RuvC mutant has mutations in the RuvC domain causing it to be non-functional. dCas9 likewise has point mutations at H840A in the HNH domain and at D10A in the RuVCl domain causing endonuclease inactivity. In some embodiments, the lbCasl2a used herein comprises SEQ ID NO:3. b. Zine-Finger Nucleases
[0069] In some embodiments, the DNA sequence-targeting protein is a zinc-finger nuclease (ZFN). ZFNs typically comprise a zinc finger DNA binding domain and a nuclease domain. Generally, ZFNs include two zinc finger arrays (ZFAs), each of which is fused to a single subunit of a non-specific endonuclease, such as the nuclease domain from the FokI enzyme, which becomes active upon dimerization. Typically, a single ZFA consists of 3 or 4 zinc finger domains, each of which is designed to recognize a specific nucleotide triplet (GGC, GAT, etc.). A ZFN composed of two "3-finger" ZFAs is therefore capable of recognizing an 18 base pair target genomic region (i.e., recognition sequence); an 18 base pair recognition sequence is generally unique, even within large genomes such as those of humans and plants. By directing the co-localization and dimerization of the two FokI nuclease monomers, ZFNs generate a functional site-specific endonuclease that can target a particular locus (e.g., gene, promoter, or enhancer) on the target genomic region.
[0070] Zinc-finger nucleases useful in the compositions disclosed herein include those that are known and ZFN that are engineered to have specificity for one or more target genomic regions described herein (e.g., promoter or enhancer nucleotide sequence). Zinc finger domains are amenable for designing polypeptides which specifically bind a selected polynucleotide recognition sequence within a target genomic region of the target cell genome. ZFN can comprise an engineered DNA-binding zinc finger domain linked to a non-specific endonuclease domain, for example, a nuclease domain from a Type IIS endonuclease such as HO or FokI.
[0071] In some embodiments, the target genomic region for the zinc finger nuclease is endogenous to the target cell, such as a native locus in the target cell genome. In some embodiments, the target genomic region is selected according to the type of nuclease to be utilized in the method. If the nuclease to be utilized is a zinc finger nuclease, optimal target genomic regions may be selected using a number of publicly available online resources. See, e.g., Reyon et al., BMC Genomics 12:83 (2011), which is hereby incorporated by reference in itsentirety. Publicly available methods for engineering zinc finger nucleases include: (1) Context- dependent Assembly (CoDA), (2) Oligomerized Pool Engineering (OPEN), (3) Modular Assembly, (4) ZiFiT (internet-accessible software for the design of engineered zinc finger arrays), (5) ZiFDB (internet-accessible database of zinc fingers and engineered zinc finger arrays), and (6) ZFNGenome. For example, OPEN is a publicly available protocol for engineering zinc finger arrays with high specificity and in vivo functionality, and has been successfully used to generate ZFNs that function efficiently in plants, zebrafish, and human somatic and pluripotent stem cells. OPEN is a selection-based method in which a pre-constructed randomized pool of candidate ZFAs is screened to identify those with high affinity and specificity for a desired target sequence. Additionally, ZFNGenome is a GBrowse-based tool for identifying and visualizing potential target genomic regions for OPEN-generated ZFNs. ZFNGenome provides a compendium of potential ZFN target genomic regions in sequenced and annotated genomes of model organisms. ZFNGenome includes more than 11 million potential ZFN target genomic regions, mapped within the fully sequenced genomes of seven model organisms; S. cerevisiae, C. reinhardtii, A. thaliana, D. melanogaster, D. rerio, C. elegans, and H. sapiens. ZFNGenome provides information about each potential ZFN target genomic regions, including its chromosomal location and position relative to transcription initiation site(s). Users can query ZFNGenome using several different criteria (e.g., gene ID, transcript ID, target genomic region sequence). c. Transcription Activator-Like Effector Nucleases
[0072] In some embodiments, the DNA sequence-targeting protein is a transcription activatorlike effector nuclease (TALEN). Transcription activator-like effectors (TALEs) are proteins secreted by Xanthomonas bacteria and play an important role in disease or triggering defense mechanisms, by binding target DNA and activating effector-specific target cell genes, see, e.g., Gu et al. Nature 435: 1122-5 (2005). A TALEN comprises a TALE DNA-binding domain fused to a DNA cleavage domain (i.e. a nuclease). The DNA binding domain interacts with target DNA in a sequence-specific manner through one or more tandem repeat domains. The repeated sequence typically comprises 33-34 highly conserved amino acids with divergent 12thand 13thamino acids. These two positions, referred to as the Repeat Variable Diresidue (RVD) are highly variable and show a strong correlation with specific nucleotide recognition (Boch et al., Science 326(5959): 1509-12 (2009); and Moscou and Bogdanove, 326(5959): 1501 (2009)). Thisrelationship between amino acid sequence and DNA recognition sequence has allowed for the engineering of specific DNA-binding domains by selecting a combination of repeat segments containing the appropriate RVDs.
[0073] The TALE DNA-binding domain can be engineered to bind to a target DNA sequence and fused to a nuclease domain, e.g., a Type IIS restriction endonuclease, such as FokI (see e.g., Kim et al. Proc. Natl. Acad. Sci. USA 93: 1156-1160 (1996)). The nuclease domain can comprise one or more mutations (e.g., FokI variants) that improve cleavage specificity (see, Doyon et al., Nature Methods, 8 (1): 74-9 (2011)) and cleavage activity (Guo et al., Journal of Molecular Biology, 400 (1): 96-107 (2010)). Other useful endonucleases that can be used as the nuclease domain include, but are not limited to, Hhal, Hindlll, Nod, BbvCI, EcoRI, Bgll, and AlwI. In some embodiments, the TAKEN can comprise a TAL effector DNA binding domain comprising a plurality of TAL effector repeat sequences that bind to a specific nucleotide sequence (i.e., recognition sequence) in the target DNA
[0074] In some embodiments, the target genomic region for the TALEN is endogenous to the target cell, such as a native locus in the target cell genome. In some embodiments, the target genomic region is selected according to the type of nuclease to be utilized in the method. If the nuclease is a TALEN, optimal target genomic regions may be selected in accordance with the compositions described by Sanjana et al., Nature Protocols,(2012), which is hereby incorporated by reference in its entirety. TALENs function as dimers, and a pair of TALENs, referred to as the left and right TALENs, target sequences on opposite strands of DNA. TALENs are engineered as a fusion of the TALE DNA-binding domain and a monomeric FokI catalytic domain. To facilitate FokI dimerization, the left and right TALEN target genomic regions are generally selected with a spacing of approximately 14-20 bases. d. Meganucleases
[0075] In some embodiments, the DNA sequence-targeting protein used in the compositions herein is a meganuclease. Meganucleases generally refer to rare-cutting endonucleases or homing endonucleases that can be highly specific. Meganucleases can recognize DNA target genomic regions ranging from at least 12 base pairs in length. Meganucleases can be modular DNA-binding nucleases such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA binding domain or protein specifying a nucleic acid targetsequence. The DNA-binding domain can contain at least one motif that recognizes single- or double-stranded DNA. A meganuclease can optionally generate a double-stranded break. A double-strand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing as described in the methods herein). DNA break repair can occur via non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor DNA repair template or template polynucleotide that contains homology arms flanking genomic regions of the target DNA can be provided. The meganuclease can be monomeric or dimeric. In some embodiments, the meganuclease is wildtype type, and in other instances, the meganuclease is non-natural, artificial, engineered, synthetic, rationally designed, or man-made. In some embodiments, the nuclease domain is catalytically inactive. In some embodiments, a meganuclease can bind DNA but cannot cleave the DNA. In some embodiments, the meganuclease of the present disclosure includes an I-Crel meganuclease, I-Ceul meganuclease, I-Msol meganuclease, I-Scel meganuclease, variants thereof, derivatives thereof, and fragments thereof. Detailed descriptions of useful meganucleases and their application in gene editing are found, e.g., in Silva et al., Curr Gene Ther, 11(1): 11-27 (2011); Zaslavoskiy et al., BMC Bioinformatics, 15: 191 (2014), and Takeuchi et al., Proc Natl Acad Sci USA, 11 l(l l):4061-4066 (2014). e. Promotor Repressor Elements
[0076] In some embodiments, the DNA sequence-targeting protein used in the compositions herein is a promoter repressor element. Promoter repressor elements generally refer to proteins which bind to a site on a promoter (i.e., an operator site) and prevent expression of the gene regulated by the promoter. In some embodiments, the promoter repressor element used herein is tetracycline repressor (TetR). ii. Nuclear Localization Signal Sequences
[0077] In some embodiments, the polypeptide comprising the transposase linked to a DNA sequence-targeting protein further comprises one or more nuclear localization signal (NLS) sequences. A NLS is an amino acid sequence that acts as a signal to mediate the transport of the polypeptide described herein from the cytoplasm into the nucleus of a cell, e.g., a eukaryotic cell described herein. For example, the NLS sequence optionally used in the polypeptide described herein is located at two adjacent regions comprising wildtype residues K104-R105 and R118-R119 of SEQ TD NOs: 1 or 2. Other exemplary nuclear localization sequences are described in, for example, Lu et al., Cell Communication and Signaling volume 19:60 (2021).
[0078] In some embodiments, the polypeptide comprising the transposase linked to a DNA sequence-targeting protein lacks a functional NLS sequence. The polypeptide may wholly lack a NLS sequence or may comprise a mutated NLS sequence such that the NLS is dysfunctional. For example, the NLS sequence of the transposase described herein may comprise one or more mutations, for example inserting alanine, at wildtype residues K104, R105, R118, and KI 19 of SEQ ID NOs: 1 or 2. Optionally, the NLS sequence of the transposase described herein comprises R118A and K120A mutations relative to SEQ ID NOs: 1 or 2. iii. Linkers
[0079] In some embodiments, in the polypeptide provided herein, the transposase is linked to the DNA sequence-targeting protein via a peptide linker. The peptide linker may increase the range of orientations that may be adopted by the domains of the fusion protein. The peptide linker may be optimized to produce desired effects in the fusion protein or modified protein. Aspects of peptide linker design and considerations are described, for example, in Chen, X. et al., Adv Drug Deliv Rev. Oct 15; 65(10): 1357-1369 (2013), and Klein, J.S. et al. Protein Eng. Des. Sei. 27(10):325-330 (2014). In some embodiments, the polypeptides provided herein comprise more than one peptide linkers.
[0080] Peptide linkers used herein may be short or long, flexible or rigid. See, e.g., PCT / US2020 / 051383 incorporated herein by reference in its entirety. Flexible linkers provide a certain degree of movement or interaction between the polypeptide domains and are generally rich in small or polar amino acids such as Gly and Ser (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or all of the amino acid residues of the linker are either Gly or Ser). A rigid linker can be used to keep a fixed distance between the domains and to help maintain their independent functions.
[0081] The length of a peptide linker may affect one or more functions of the fusion protein. Selection of linkers to achieve the desired length is within the ability of one skilled in the art. In some embodiments, a peptide linker may be, for example, 5 to 100 or more amino acids in length (e.g., 5 aa, 10 aa, 15 aa, 20 aa, 25 aa, 30 aa, 35 aa, 40 aa, 45 aa, 50 aa, 55 aa, 60 aa, 65 aa, 70 aa, 75 aa, 80 aa, 85 aa, 90 aa, 95 aa, or 100 aa).
[0082] In some embodiments, a peptide linker described herein comprises an amino acid sequence with at least 90%, 95%, 98%, 99%, or 100% identity to SEQ ID NO:4. Other exemplary linkers may be identified, for example, by screening libraries of linkers composed of several amino acid residues that have undergone site saturation mutagenesis (see, e.g., Patriarchi et al., Science, 360(6396): 1-22 (2018)). iv. Affinity Agents
[0083] In some embodiments, the fusion protein described herein comprises a transposase linked to a first affinity agent and a DNA sequence-targeting protein linked to a second affinity agent and the first and second affinity agents bind (e.g., form a dimer). In some embodiments, the first and second affinity agents bind in the presence of a chemical agent, but not in the absence of a chemical agent. In some embodiments the first and second affinity agents are the same (i.e., they form a homodimer). In some embodiments the first and second affinity agents are different (i.e., they form a heterodimer). In some embodiments, the fusion protein comprising the first and second affinity agents comprises SEQ ID NO:8.
[0084] Exemplary pairings of affinity agents and dimerization-inducing chemical agents are described in Dang et al., Front Chem., 10:829312 (2022) and may be selected by one skilled in the art. In some embodiments, the first and second affinity agents are selected from FK506 binding protein (FKBP) and FKBP-rapamycin binding (FRB) domain in the FKBP-rapamycin- associated protein. FKBP and FRB dimerize in the presence of rapamycin. In some embodiments, the chemical agent is rapamycin. In some embodiments, the chemical agent is an analog of rapamycin, for example, C16-(S)-7-methylindolerapamycin (i.e., ligand AP21967).C. Polynucleotides and Expression Cassettes
[0085] Also provided herein are polynucleotides encoding any of the polypeptides described herein. For example, a polynucleotide encoding a polypeptide that has at least at least 90%, 95%, 98%, 99%, or 100% identity to any of SEQ ID NOs: 1, 2, or 5 is also provided.
[0086] Also provided is an expression cassette comprising a promoter operably linked to any of the polynucleotides described above or elsewhere herein. As described above, a promoter is “operably linked” to a polynucleotide when it is placed into a functional relationship with the polynucleotide sequence. Numerous promoters can be used in the constructs described herein. As231described above, the term “promoter” as used herein refers to a nucleotide region or a sequence located upstream and / or downstream from the start of transcription that is involved in recognition and binding of RNA polymerase and other proteins to initiate transcription.
[0087] In some embodiments the promoter is tissue-specific (i.e., it directs transcription at high levels only in particular types of cells or tissues). Exemplary tissue-specific promoters include, inter alia, a synapsin, camKIIa, glial fibrillary acidic protein (GFAP), retinal pigment epithelium (RPE), albumin (ALB), thyroxine binding globulin (TBG), myelin basic proteins (MBP), muscle creatine kinase (MCK), cardiac troponin T (TnT), or alpha-myosin heavy chain (aMHC), and the like. In some embodiments, the promoter is inducible (i.e., it directs transcription only under certain circumstances). For example, an inducible promoter used in the expression cassettes herein may be tetracycline inducible. In some embodiments, the promoter is constitutive (i.e., it directs transcription at relatively similar levels across all cell and tissue types). Exemplary constitutive promoters include, inter alia, a CMV promoter, CAG promoter, CBA promoter, EFla promoter, PGK promoter, and the like.
[0088] The promoter can be a promoter of a template polynucleotide comprising a gene used in the methods described below and elsewhere herein. The promoter can be heterologous to (i.e., not naturally occurring in) the gene of the template polynucleotide used in the methods described below and elsewhere herein.
[0089] The choice of promoters to be included depends upon several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue- preferential expression. It is a routine matter for one of skill in the art to modulate the expression of a sequence by appropriately selecting and positioning promoters and other regulatory regions relative to that sequence. Exemplary promoters useful in the expression cassettes described herein include y-synuclein gene-based promoter, SNCG, and microglia-specific promoters (e.g., promoters from Cluster of Differentiation 68 (CD68), CXC motif chemokine receptor 1 (CXCR1), or Hexosaminidase Subunit Beta (HEXB)).D. Vectors
[0090] Also provided herein are vectors comprising any of the expression cassettes described above or elsewhere herein. In some embodiments, the polynucleotide in the expression cassette may be codon optimized for delivery as a vector. In some embodiments, the vector is a viralvector, wherein the expression cassette comprising a promoter operably linked to any of the polynucleotides described above or elsewhere herein can be contained within a viral vector and administered as a viral particle. Exemplary viral vectors include but are not limited to adenovirus vectors (e.g., Ad2, Ad5, Ad7), adeno-associated viral vectors, herpes simplex viral vectors, retroviral vectors, pox viral vectors (such as vaccinia and avian poxvirus vectors, such as the fowlpox and canarypox vectors), lentiviral vectors, alphavirus vectors, poliovirus vectors, measles vectors, and other positive and negative stranded RNA viruses, viroids, and virusoids, or portions thereof. The vectors herein may or may not have replicative capacity.
[0091] In some embodiments, the viral vector is a lentiviral vector. In some embodiments, the lentiviral vector does not integrate viral genetic material into the target genome, i.e., it is a nonintegrating lentiviral vector.
[0092] In some embodiments, the nucleic acid may be administered as a non-viral vector, including, but not limited to, as a plasmid, in a nanoparticle, (e.g., a lipid nanoparticle), or in a liposome.E. Pharmaceutical Compositions
[0093] The polypeptides, polynucleotides, expression cassettes or vectors described herein may be formulated for pharmaceutical administration, for example, as a pharmaceutical composition further comprising a pharmaceutically-acceptable excipient or carrier. The pharmaceutical composition can additionally contain other therapeutic agents that are suitable for treating or preventing a given disorder. Pharmaceutically carriers and excipients can enhance or stabilize the composition, enhance efficacy of the composition in treating or preventing a given disorder, or facilitate preparation of the composition. Pharmaceutically acceptable carriers include solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible.
[0094] A pharmaceutical composition as described herein can be administered by a variety of methods known in the art. The route and / or mode of administration vary depending upon the desired results. Administration can be intravenous, intramuscular, intraperitoneal, or subcutaneous, or administered proximal to the site of the target. The pharmaceutically acceptable carrier should be suitable for intravenous, intramuscular, subcutaneous, parenteral, spinal or epidermal administration (e.g., by injection or infusion). Depending on the route ofadministration, the active compound, i.e., the polypeptide described herein, may be coated in a material to protect the compound from the action of acids and other natural conditions that may inactivate the compound.
[0095] Typically, a therapeutically effective dose or efficacious dose of the polypeptides, polynucleotides, expression cassettes, or vectors described herein is employed in the pharmaceutical compositions. The polypeptides, polynucleotides, expression cassettes, or vectors can be formulated into pharmaceutically acceptable dosage forms. Dosage regimens are adjusted to provide the desired response (e.g., a therapeutic response). In determining a therapeutically or prophylactically effective dose, a low dose can be administered and then incrementally increased until a desired response is achieved with minimal or no undesired side effects. For example, a single bolus may be administered, several divided doses may be administered over time or the dose may be proportionally reduced or increased as indicated by the exigencies of the therapeutic situation. It is especially advantageous to formulate parenteral compositions in dosage unit form for ease of administration and uniformity of dosage. Dosage unit form as used herein refers to physically discrete units suited as unitary dosages for the subjects to be treated; each unit contains a predetermined quantity of active compound calculated to produce the desired therapeutic effect in association with the required pharmaceutical carrier.
[0096] Actual dosage levels of the active ingredients in the pharmaceutical compositions can be varied so as to obtain an amount of the active ingredient which is effective to achieve the desired therapeutic response for a particular patient, composition, and mode of administration, without being toxic to the patient. The selected dosage level depends upon a variety of pharmacokinetic factors including the activity of the particular compositions employed, or the ester, salt or amide thereof, the route of administration, the time of administration, the rate of excretion of the particular compound being employed, the duration of the treatment, other drugs, compounds and / or materials used in combination with the particular compositions employed, the age, sex, weight, condition, general health and prior medical history of the patient being treated, and like factors.III. METHODS
[0097] Also provided herein are methods of editing the genome of a eukaryotic cell comprising delivering the polypeptide described above or elsewhere herein the nucleus of the eukaryotic cell, such that the polypeptide nicks the genome of the eukaryotic cell in at least one target location and introduces a template polynucleotide into the nick, thereby editing the genome of the eukaryotic cell. The methods provided herein may be for editing the genome of any eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the eukaryotic cell is a human cell. The methods described herein may be performed in vitro, ex vivo, or in vivo.
[0098] In some embodiments, the polypeptide used in the methods of editing described herein comprises a transposase comprising SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype. In some embodiments, the transposase is at least 90%, 95%, 98%, 99%, or 100% identical with SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
[0099] In some embodiments, the polypeptide used in the methods of editing described herein comprises a transposase linked to a DNA sequence-targeting protein (i.e., provided as a fusion protein). Exemplary DNA sequence-targeting proteins are described herein and include CRISPR- Cas proteins, TALEN proteins, zinc finger proteins, meganuclease proteins, or promoter repressor elements. In some embodiments, the DNA sequence-targeting protein is a CRISPR / Cas protein. In some embodiments, the DNA sequence-targeting protein is Casl2a or Cas9. The DNA sequence-targeting protein can have nicking activity or lack nicking activity. In some embodiments, the DNA sequence-targeting protein is dCas!2a or dCas9. In some embodiments, the DNA sequence-targeting protein comprises SEQ ID NO:3.
[0100] The DNA sequence-targeting protein optionally used in the methods herein may be delivered to the cell with one or more guide RNAs (gRNAs). gRNAs generally have a nucleic acid sequence complementary to the nucleic acid sequence of the target genomic region and confer sequence specificity to DNA sequence-targeting protein used herein. In some embodiments, 2 or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9 or 10) gRNAs are delivered to the cell. Insome embodiments, the gRNAs are linked together consecutively and placed under the control of the same promoter, for example, as described in Port et al., Elife, 13(9):e53865 (2020). The gRNAs optionally used herein form a ribonucleoprotein complex with the DNA sequencetargeting protein. The ribonucleoprotein complex may be delivered to the cell through any of the methods described herein (e.g., as a polynucleotide or a vector). In some embodiments, the gRNA and the DNA sequence targeting-protein forms a ribonucleoprotein, and the ribonucleoprotein is delivered to the cell as a polypeptide (e.g., through electroporation).
[0101] The methods described herein comprise delivering the polypeptide as described above or elsewhere herein to the eukaryotic cell. In some embodiments, delivering comprises contacting the cell with a vector comprising a polynucleotide encoding the polypeptide such that the polypeptide is expressed in the cell and enters the nucleus of the cell. In some embodiments, delivering comprises contacting the cell with a polynucleotide (e.g., RNA) encoding the polypeptide such that the polypeptide is expressed in the cell and enters the nucleus of the cell. Delivering may occur in vitro, ex vivo, or in vivo.
[0102] In some cases, the transposase contained in the polypeptide described above or elsewhere herein nicks the genome of the eukaryotic cell in at least one target location. Generally, without intending to be bound by a specific mechanism, the transposase recognizes the ends of template polynucleotide sequence, e.g., at terminal inverted repeats, as described below (analogous to a transposon in nature), and using a “cut and paste” mechanism, the transposase excises the template polynucleotide sequence, and integrates it into the target genomic region at the nick. An exemplary description of transposition is described in Munoz- Lopez and Perez, Curr. Genomics, 11(2): 115-28 (2010). In other cases, the DNA sequencetargeting protein nicks the genome of the eukaryotic cell in at least one target location.
[0103] The template polynucleotide introduced into the genome of the cell in the methods herein template length may be at least 10, 20, 30, 40, 50, 70, or 100 kilobases (kb) long. In some embodiments, the template polynucleotide is about 100-5,000 kb long. Exemplary template polynucleotides used in the methods of editing herein include, but are not limited to, a full bacterial artificial chromosome (e.g., for insertion into the Rosa26 target genomic region in a murine model), the rhodopsin gene (e.g., for insertion into a subject having retinitis pigmentosa), or a synthetic gene placed downstream of a promoter to result in its expression (e.g., YamanakaFactors placed downstream of a senescence-related gene promoter, such as the pl6-INK4a promoter to treat a subject having an aging-related disease).
[0104] The template polynucleotides used in the methods herein may comprise terminal inverted repeats (TIRs). TIRs are a sequence of nucleotides present at either end of a transposon or a template polynucleotide used herein. TIRs are followed downstream by its reverse complement. For example, a transposon or template polynucleotide having a TIR at the 5’ end having the sequence TTA will also have a TIR at the 3’ end with the sequence AAT. The TIRs used in the template polynucleotides herein may be between 10 and 2,000 base pairs in length. The template polynucleotides used herein may have one or more TIRs at both the 5’ and 3’ ends. The TIRs may be placed at varying distances from one another on the template polynucleotide described herein to vary transposition efficiency. In some embodiments, the TIRs are placed less than 10 kb apart. In some embodiments, the TIRs are placed less than 100 kb apart. In some embodiments, the TIRs are placed less than 500 kb apart. In some embodiments, the left TIR comprises SEQ ID NO:6 and the right TIR comprises SEQ ID NO:7. The TIRs used herein may be delivered to the cell in a vector with the template polynucleotide as described below.
[0105] The template polynucleotides used in the methods herein may be delivered to the cell. In some embodiments, delivering comprises contacting the cell with a vector (e.g., a viral vector or a plasmid) comprising the template polynucleotide. In some embodiments, the template polynucleotide is expressed in the nucleus of the cell. In some embodiments, the template polynucleotide is delivered as a viral vector (e.g., a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector). In some embodiments, delivering comprises contacting the cell with the template polynucleotide not contained in a vector (e.g., as RNA), such that the template polynucleotide is expressed in the cell. Delivering may occur in vitro, ex vivo, or in vivo. The template polynucleotide may be delivered together with or separate from the transposase or fusion protein used in the methods herein. For example, two plasmids may be delivered to the cell: one containing the template polynucleotide and one containing the transposase or fusion protein. If delivered separately from the transposase or fusion protein, the template polynucleotide may be delivered before or after delivery of the transposase or fusion protein. In some embodiments, one or more gRNAs for targeting the transposase fusion to a genomicsequence are also encoded by a vector encoding the transposase fusion and / or the template polynucleotide.EXAMPLESMaterials and methods
[0106] Construction of the fusion construct of lbCasl2a to SB100X. The DNA construct encoding lbCasl2a, RuvC domain mutant D832A (obtained from Addgene #113517) was PCR amplified and fused to the N-terminus of SB100X using standard Gibson Assembly cloning methods (NEB). It was established previously (Kovac et al., 2020) that a C-terminal fusion would render SB100X inactive, so we did not attempt this configuration. Between the lbCasl2a D832A and SB100X coding regions was a sequence encoding a short polypeptide linker consisting of 14 amino acids (KLGGGAPAVGGGPK). Only one flexible linker was tested. For the initial cell-based screens of the fusion protein, the lbCasl2a_SB100X fusion sequence was cloned into a mammalian expression vector on a CMV promoter with a Neomycin selection cassette. For subsequent experiments involving lentivirus generation, the fusion was cloned into a pLX lentiviral transfer vector.
[0107] Construction of the gRNA array. The gRNA arrays were constructed following the general protocol of Port et al., 2020. Generally, the pBluescript plasmid was obtained form Addgene (SpCas9_sgRNA_expression_in_pBluescript was a gift from Scot Wolfe (Addgene plasmid # 122089 ) and the Cas9 tRNA was removed and replaced with the pCDF8 Casl2a gBlock (see Table 1 below) downstream of the U6 promoter using standard Gibson Assembly methods. lbCasl2a / Cpfl specific crRNAs were identified using the Benchling gRNA algorithm, which provides scores for ‘on-target’ and ‘off-target’ effects of each guide. 20-mer crRNAs with TTTN PAM sequences were selected in order of optimized score, and also in a manner that would space the PAMs evenly across the genomic target window. The crRNAs were assembled using the ultramers (shown in sequences, in accordance with the Port et al., 2020 protocol) with each crRNA spaced by a stem loop generating sequence. The crRNA was assembled by overlap extension PCR (see primers in Table 1 below) and then cloned into the BbsI site of the pBluescript-scaffold vector using the HiFi assembly system (NEB). The genomic target window ranged from 1600 base pairs in the case of the mouse pl6-INK promoter region to 290 base pairs as in the hHPRT locus. Six crRNAs were selected for each gRNA array in an effort to concentrate the lbCasl2a_SB100x / variant as much as possible to increase the likelihood of an on-target integration, as well as bind up as much apo-tool with RNA spacer as possible to potentially reduce off-target effects. The idea is to essentially flood the system with gRNA, bothin terms of the number of guides per array, as well as in terms of the ratio of gRNA to expressed lbCasl2a_SB100x / variant such that very few tool proteins should reside in RNA unbound form.Table 1. Constructs used to create the gRNA array
[0108] Initial screening of the fusion construct. The lbCasl2a_SB100X fusion plasmid was transfected into C2C12 cells at 1.25 ug / mL (Lipofectamine 2000, final concentration in a 6 or 12 well plate) along with the transposon plasmid pSBtet-GN (Addgene #60501), consisting of eithera CMV promoter or a Dox inducible tight TRE promoter. The transposon encoded EGFP, puromycin selectable marker, and firefly luciferase.
[0109] FIG. 1 shows a fusion protein consisting of lbCas!2a and SB100X composes the RNA guided transposase, which is co-transfected with a transposon. For early cell-based assays, the pSB-TET transposon was used, consisting of a CMV promoter, which drives GFP and puromycin separated by a P2A peptide. The transposon also had a luciferase coding region, driven either by a TRE CMV promoter, or the promoter was removed for studying induction by endogenous promoters upon strategic site directed integration. The transposon carried the Tcl / mariner SB ITRs. A crRNA array with 6 crRNAs was used in targeting experiments, whereas a stem-loop scaffold control was used as an untargeted control. The gRNA vector contained a tRNA downstream of the U6 promoter, followed by the crRNA array, consisting of 6 crRNA spacers with a TTTN PAM, each separated by a stem loop as in Port et al. 2020, described above. The genomic targeting window ranged from 1600bp using 6 guides to 120bp also using 6 guides. Because the efficiency of the SB transposon is dependent largely on the location of the ITRs, rather than the transposon carrying capacity, hypothetically, numerous genes, ORFs, DNA fragments could be constructed into a single genome targeted transposon (represented by multiple ORF fragments).
[0110] For in vitro assays of genomic target selectivity, the CMV tTRE promoter was eliminated so that the luciferase could be driven from an endogenous promoter, such as pl 6- INK4a. A crRNA array consisting of 6 crRNAs spanning a Ikb region of the pINK4A promoter (Table 2 below) was also co-transfected with the fusion construct and the transposon was a crRNA array consisting of 6 crRNAs spanning a Ikb region of the pINK4A promoter (FIG. 4). To assess whether the transposon was integrating into the targeted region, the pl6-INK4A promoter, we assessed the intensity of luminescence production by the firefly luciferase in the transposon upon stimulation by doxorubicin, which is known to drive pl6 expression. These experiments did not yield evidence of directed transposition into the pl 6 promoter region as there was no indication of pl6 dependent luciferase expression following 6 days of doxorubicin (250nM) treatment. Three mutants were selected and generated by single point mutagenesis methods (H165E, H187E, K252E) on the basis of their reported efficiency in SB100X as well as on the proximity of these residues to target DNA in the solved structure (Voigt et al. 2016).Following lipofectamine transfection of the crRNA array targeting the p! 6-INK4A mouse promoter region, the pSBtet-GN transposon, and the vector containing lbCasl2a_SB100X fusion sequence, or the three SB100X mutants as a fusion with lbCasl2a, in a 1 : 1 : 1 ratio, G-418 selection agent was added for 6 days to select for those colonies that had genomic integration of the transposon. Following selection, colonies of resistant C2C12 cells emerged in the conditions that allowed for transposition, the plates were imaged using a Nikon automated image acquisition system (details about the Nikon SRRF / TRF microscope) and EGFP positive colonies were quantified in FIJI using the ‘median,’ ‘threshold,’ and ‘particles’ filters. For the experiments involving quantification of luciferase activity, to test whether the transposon was integrated into the p!6 promoter, C2C12 cells transfected (in a 1: 1: 1 ratio) with the 6x crRNA targeting the pl6 promoter or a gRNA scaffold control, the transposon with the promoter upstream of luciferase removed, and either a mutant or WT lbCasl2a-SB100X fusion construct or a control construct consisting only of WT unfused SB100X, were plated in a white bottom opaque 96 well plate. Cells were selected for 6d with neomycin beginning 72h following transfection since the EGFP and kan / neo were under a CMV promoter, and then luciferase activity was imaged using a luminescent plate reader (BioTek Cytation 3) immediately following the addition of D-Luciferin (Promega).Table 2. pl6 crRNA[OHl] Construction of the site saturation mutagenesis library. Four sites within SB100X were selected on the basis of either their proximity to target DNA, position in the catalytic domain ofSB100X, or prior experiments in our lab and others demonstrating modified efficacy of SB100X (Voigt et al., 2016), these were H165, H187, 1212, K252. A site saturation mutagenesis library was generated, first by designing primers covering single sites with ‘NNK’ in place of the mutagenized position where ‘N’ represents any base and ‘K’ represents G or T, resulting in a generally even distribution of 20 amino acids or 33 codons at each position, and a library size of 160,000 unique variants. The primers were designed to amplify the region between one site for mutagenesis to the last residue before the next mutagenic primer, and then all four fragments were assembled by overlap extension PCR to complete the variable region of SB100X. Primers are shown in Table 3 below; sequences used to generate the library constructs are shown in Table 4 below. The original pLX lentiviral vector (Addgene #162073) was modified to include an adjacent right and left transposable element (TE), specific for SB.Table 3. Primers used to amplify the variant libraryTable 4. Sequences used to generate library construct
[0112] Additionally, 24bp downstream of the Eifla promoter, a puromycin selection cassette and P2A cleavable peptide were assembled into the pLX vector as one single fragment between the TEs and puromycin P2A. Immediately 3’ to the P2A the lbCasl2a fused to the the N-terminal region lacking the variable fragment of SB100X was also assembled (SB100X residues 1-164). The pLX plasmid was amplified by standard methods using e.coli (NEB, Stbl cells). The modified pLX (lug) was digested with Spel (NEB), and the variable region of SB100X was assembled into the modified pLX vector at a 1:2 molar insert to vector ratio, (xng insert, xug vector, HiFi Assembly, NEB) following the NEB recommended protocol. The library was subject to RecBCD Exonuclease V treatment (NEB) to degrade unassembled DNA fragments from the mixture, followed by column cleanup (Zymo, DNA Clean and Concentrator ), yielding a total of 70ng of the assembled pLX_IR_Puro_lbCasl2a_SB100X library plasmid. To avoid reducing the diversity of the library, it was not amplified in e.coli. A small quantity (<0. Ing) was transfected into E. colt and prepared for whole plasmid or Sanger sequencing to ensure the accuracy of the assembled construct, as well as to ensure that the library represented random mutagenesis. Following viral construction, the input library was also subject to NGS (Genewiz) to obtain a complete landscape of the mutagenic library.
[0113] Non-integrating lentivirus generation. HEK293FT cells at below 14 passages were seeded in fibronectin coated (lug / cm2, Thermo) T-175 flasks at 24xlOA6 cells so that they would be 80% confluent on the day of transfection. The pLX complete library plasmid was transfectedinto HEK293FT cells at a ratio of 1 :2000:2000 (pLX library vector: psPAX2 D64V: pMD2.G) using PEI (pH5, lug / mL at 1 ,2ul per lug of helper vector). This low ratio of transfer vector to packaging plasmids reduces the risk that multiple library vectors might be packaged into the same virion (Kumar et al., 2020). The plasmids were added to 1.5mL of OptiMEM, while the PEI was added to another tube of 1 5mL OptiMEM and then added dropwise, % of the PEI OptiMEM mixture to the DNA mixture, followed by 15s of vortexing until all PEI mix had been added. The DNAZPEI mixture was allowed to incubate at room temperature for 15min prior to adding it to the cells. Four T-175 flasks were transfected with the NILV components and allowed to incubate for 48h before the first supernatant was collected. Both media and cells were collected at 60h. Media was stored at 4C for no more than one week before being concentrated using the Trono lab viral preparation protocol (Salmon and Trono, 2006).
[0114] Briefly, supernatants from 48h and 60h were pooled and centrifuged at 500g, then filtered with a 0.22um bottle top filter. The supernatant was pipetted into Beckman 38.5 Ultraclear tubes and then ultracentrifuged for 120min at 50,000g at 16C using (Beckman centrifuge, rotor size). The supernatant was gently discarded by inversion and the pellet was left to air dry for 2-3 min, ensuring that it didn’t dry out. The pellet was resuspended in PBS, 15ul at a time, to a minimal volume of 30ul, first pipetting up and down 15x around the center region, followed by 15x around the circumference of the tube. The volume of the first tube was transferred to the second, and the process of adding 15ul of PBS followed by pipetting around the center and edge was repeated sequentially until all the viral containing PBS was in the final tube, which was then transferred to an Eppendorf tube, centrifuged briefly on a benchtop centrifuge, aliquoted into lOul stocks, and stored at -80C to prevent freeze / thaw cycles of the viral stock. Lentivirus particles were quantified using a p24 ELISA (Abeam), and were estimated to be 9.4x1010 particles per mL. To estimate multiplicity of infection (MOI), C2C12 cells were plated at 3.8x105 cells per well of a 6-well plate and the ratio of viral particles to cells was logarithmically titered from 1000 to 0 vp / cell. lOOvp per cell was found to be an MOI of 0.03 based on the fraction of puromycin resistant cells. This should represent a maximum of a single library variant per cell.
[0115] To test the capacity of library variants to undergo targeted transposition, 10x106 C2C12 cells were electroporated (2M cells per cuvette) were electroporated (AMAXA Nucleofector Vkit) with the 6xcrRNA array (pScaffold_6xcrRNAj)16-INK, lug ) or the pScaffold control (lug) along with EGFP (pMAX GFP, 0.5ug). Based on the percent of GFP positive cells 24h following transfection, the 6xcrRNA array or scaffold control were transfected into at least 50% of cells. One day following transfection with the gRNA or control constructs, NILV containing the lbCasl2a_SB100X library was used to infect the cells as described above at an MOI of 0.036. 72h following NILV infection, puromycin (3ug / mL) was added, selecting for the population of cells that had some global transposition occur somewhere in the genome.
[0116] Screening ofC2C12 cells for site selective transposition. Following selection of NILV infected C2C12 cells with puromycin (6d) were collected using a cell lifter in PBS and genomic DNA was harvested (Qiagen, DNA Easy check kit). Two sets of nested PCR primers were designed. In the first set, primers annealing to genomic regions outside of the farthest 5’ and 3’ PAM sites corresponding to the crRNA targeting the pl6-INK4A promoter region of the mouse genome (see Table 5 below). The second primers in this set annealed to the 5’ and 3’ edges of the 450bp variable region of SB100X. In the second set of nested PCR primers, the first set of primers annealed to the TSS site of pl6-INK4a, about lOObp outside of the most proximal PAM site of the gRNA, and the second primer annealed to a region within the transposon, while the second primer pair in this set was the same as the first set, spanning the variable region of SB100X. The 13.4kb band, indicating a larger amplicon of the size anticipated for a site directed transposition, was cut from a 1% agarose gel, leaving the Ikb WT band), isolated, and then amplified with the second primer pair in the nested PCR set yielding a 450bp band, the expected size for the variable SB100X region amplified. Both sets of nested PCR primers yielded the anticipated 450bp band. This band from the first nested PCR primer set was isolated and subject to NGS (GeneWiz, Illumina miSeq). Read alignment and variant frequency analysis was performed using CRISPResso (Clement et al., 2019)Table 5. Primers to amplify the pl6 promoter region in the mouse genome outside of the putative variable SB100X insert site.
[0117] Screening of individual variants for specificity and efficiency. Individual variants that emerged from the library screen were cloned back into the pLX_puro_lbCasl2a_N- terminal_SB100X vector as single mutants for further cell-based screening, followed by enrichment sequencing analysis of transposition foci within the genome. NILV for each individual lbCasl2a_SB100X variant was generated as described and ratios for the pLX transfer plasmid to psPAX2 D64V to pMD2.G were maintained at 1 :2000:2000. Individual mutants were screened by disrupting the hHPRT gene using a 6xcrRNA array plasmid targeting hHPRT, which allowed for identification of site-specific transposition due to the induction of 6-TG (6- thioguanine) resistance previously demonstrated to occur by HPRT disruption. hHPRT crRNA is described in Table 6 below. For this assay, human HEK293FT cells were electroporated with HPRT targeting 6x crRNA array or control scaffold plasmid as previously described, and then seeded at 38x103 cells per well of a 24 well plate lacking any cell adhesive coating. Cells were infected with NILV 24h following electroporation of the gRNA, and then selected with either 6- TG (6uM) or puromycin (0.5ug / mL) for 6d, fixed in 4% PFA and then antibiotic resistance was assessed by Crystal Violet assay as previously described (Feoktistova et al., 2016). Plates were imaged with a wide field camera (Biorad Gel Dock) and percent remaining cells was quantified using FIJI. The difference in percent of viable cells following puromycin selection vs 6-TG selection, i.e., global genomic transposition vs site directed transposition, gives the selectivity of the lbCasl2a_SB100X, or variant, for a given target vs its overall efficiency for transposition somewhere in the genome.Table 6. hHPRT crRNQA
[0118] Enrichment sequencing was conducted following puromycin selection to assess the percent. gDNa was isolated as previously described, then subject to enzymatic cleavage to generate approximately 500bp fragments (NEB, NEXT), sequencing adapter primers (shown in Table 7 below) were used to amplify the fragments with one primer targeting the enzymatic cut regions and the other primer in the pair was either an adapter targeting the 5’ end of the transposon or the 3’, oriented in the minus or plus directions, respectively. Amplicons represented regions in the genome that had undergone transposition by the lbCasl2_SB100X variant. Enrichment sequencing was used to determine the percent of targeted transposition versus global genomic transposition.Table 7. Primers for sequencing SBIOOx library
[0119] Generation and characterization of the chemogenetic variant. The chemogenetic variant of lbCasl2a_SB100X was generated by first using alanine scanning of possible NLS sites, which were previously described (Yant et al., 2004), and identified to contain the residues: K104, R105, R118, KI 19, KI 20. Each residue was sequentially mutated to alanine, however none of these single mutants resulted in ablation of transposition activity as determined by the number of EGFP positive colonies of C2C12 cells lipofectamine transfected with the mutant SBIOOX (derived from pAF123), together with the EGFP and puromycin containing transposon (pSB- TET), following 6d of puromycin selection. Combinations of alanine mutations were then tested, resulting in the R118A and K120A producing an NLS deficient variant of SBIOOX. As shown in FIG. 2, the chemogenetic FKBP fragment of 321 bases, together with a 7 residue linker, a cleavable component, P2A, a second 7 residue linker, and the FRB heterodimerizer group wasassembled into the linker between lbCasl2a and SB100X to create a cleavable fusion which would bind in the presence of the rapamycin analogue (C16-aiRap, Takara #635055). To determine the percent of genomic transposition events by the chemogenetic variant, lbCasl2a_FKBP / FRB_SB100X, C2C12 cells were lipofectamine transfected as described with the chemogenetic variant or controls (SB100X, lbCasl2_SB100X, or empty vector), along with the transposon (pSB-TET), containing EGFP and puro on CMV promoters. The A / C heterodimerizing rapalogue compound was added (250nM) at different intervals following transfection to test the temporal dependence of nuclear localization of the lbCasl2a_FKBP / FRB_SB100X on efficacy of transposition. Heterodimerizer addition Ih following transfection was selected as the optimal time point for greatest transposition efficacy in this system. Following puromycin selection, plates were imaged on an automated epifluorescence microscope (Nikon SRRF) and EGFP positive colonies were quantified using FIJI as described.Results
[0120] Design of the lbCasl2a SB100X. Catalytically inactive Cas9 has been engineered effectively to serve as a targeting device to bring specific types of proteins within precise proximity to a pre-selected genomic location. These fusion proteins include transcriptional activators (VP64), suppressors (KRAB), single molecule visualization enabling proteins (Knight et al, 2015), which create additional capabilities for the CRISPR-Cas9 system. This work endeavors to add the function of transposition, via a highly optimized transposase, SB100X, to the Cas repertoire. To that end, a fusion protein consisting of lbCas!2a and SB100X, joined by a short flexible linker at the N-terminus of SB 100X was designed.
[0121] lbCas!2a was fused to SB 100X to serve as the transposon targeting mechanism. A crRNA array, consisting of 6x crRNAs on an RNA stem loop scaffold was used in initial experiments to optimize the density of the fusion protein within a window of between 290- 1490bp of the target site. For initial experiments the transposon, pSB-TET, or variations thereof, was used to look for overall genomic transposition, as well as target specific transposition, which consisted of transposable elements (TEs) located 7kb apart, PuroR and EGFP on a CMV promoter and separated by a cleavable linker (P2A), in addition to a firefly luciferase under control of a tight TRE promoter (tTRE). The tTRE promoter was eliminated in someexperiments for the purpose of serving as a reporter system for targeted transposition to determine whether the luciferase could be driven by an inducible endogenous promoter if it had undergone targeted transposition.
[0122] Determination of transposition efficiency of the lbCasl2a SB100X fusion. As depicted in FIGS 6 and 7A, to test whether the fusion of lbCasl2a_SB100X could undergo transposition globally in the genome, GFP positive, puromycin resistant colonies were counted following cotransfection with the fusion construct and a transposon with constitutive GFP and puromycin expression (pSB-TET). The transposon also contained a firefly luciferase with no promoter (pSB-noTET) to test for integration relative to endogenous promoters. The addition of the lbCasl2a to the N-terminal of SB100X reduced its transposition efficiency compared to SB100X alone by approximately 50% when co-transfected with the scaffold gRNA, and approximately 66% when co-transfected with the pl6 promoter targeting 6xcrRNA array. Interestingly, the transposition efficiency of the lbCasl2a_SB100X fusion is slightly reduced when co-transfected with pl 6 6xcrRNA compared to the control scaffold gRNA plasmid based on the firefly luminescence quantification and GFP positive cell colony counts, despite that the SB100X expression alone showed no differences in luminescence or GFP positive colony number between the pl6 6xcrRNA vs the scaffold control gRNA.
[0123] Point mutagenesis of lhCas!2a SB100X to modulate transposition. To determine if the specificity of transposition could be enhanced by mutagenesis of the transposase, we designed and tested 3 point mutations in the catalytic domain of SB100X. These positions corresponded to regions in the solved structure of SB100X likely to interact with target DNA (Voigt et al., 2016). Preliminarily, an increase in the number of puromycin resistant GFP positive colonies, as well as increased luciferase luminescence, in the case of the lbCas l2a_SB100X H165E mutant, relative to the WT SB100X fusion co-transfected with pl 6 6xcrRNA, but not the scaffold control construct, indicates that this mutant may have the capacity to increase transposition when targeted to the genome. H187E and K252E showed lower luciferase activity compared to the WT fusion construct. However, amplification and electrophoretic analysis did not yield a larger band at the 13.4kb size expected for correct placement of the transposon into the p 16 promoter. This could be due to the repetitive nature of the pl6 promoter, or to a deficiency in cells positive for the correctly targeted transposon.
[0124] Directed evolution approach to increase lbCasl2a SB 100X selectivity. Bolstered by the identification of a mutant of SB100X that might modulate transposition efficiency, the next logical step was to extend the mutagenesis to a directed evolution approach by generating a lbCasl2a_SB100X variant library by employing site saturation mutagenesis of the three residue positions tested, plus one additional: H165, H187, 1212, K252, resulting in a library of 1.6x105 variants with one of all 20 amino acids randomly distributed at each position (FIG. 5B). The input library was NGS (Genewiz) verified to ensure the full library diversity (FIG. 5A). Selection of variants competent for targeted transposition relied on the expression of only one variant per cell, and the viral genome was devoid of active integrases, such that the only genomic integration should be performed by a transposition competent variant of lbCasl2a_SB100X. The lbCasl2a_SB100X mutagenic library was packaged into NILV and eukaryotic cells were infected with an MOI of <0.04 to ensure that a maximum of only one variant copy was present per cell.
[0125] Nested PCR results showing targeted genomic integration of the transposon. Nested PCR analysis with one primer pair amplifying the pl6 genomic region outside of the PAM window and the second pair to amplify the variable region of SB 100X, identified a larger DNA amplicon corresponding to the 13.4kb transposon insert size. The band corresponding to the 450bp variable region of the transposase was analyzed by NGS and identified variants . Interestingly, the WT lbCasl2a_SB100X sequence was the major product following the single round of directed evolution. It is likely that the modified methods to include a much higher copy number of the pl6 6xcrRNA by electroporation, in addition to the low MOI NILV infection a day later may have contributed to the identification of WT sequence in the targeted transposon. The variants that came out of the screen were predominately focused at the Hl 65 position of SB100X, suggesting that this position may represent flexibility in SB100X, perhaps leading to greater specificity in exchange for lower efficiency, as demonstrated previously for this position of SB100X. In addition to the Hl 65 variants, there were a couple of very low frequency variants (<2%) at the K252 position. Based on these findings, the H165D and H165V, as well as the K252Q variants, in addition to the WT sequence, were selected for individual characterization. NILV was generated with each mutant or WT transfer vector for the purpose of infecting humanHEK293FT cells with low MOI (0.03) so as to maintain the ratio of gRNA to variant copies as in the previous experiments.
[0126] Disruption of the hHPRT gene by targeted transposition. As shown in FIGS. 6 and 7B, quickly determine the specificity and efficiency of each individual variant and WT of the lbCas!2a_SB100X fusion, hHPRT was targeted using a 6xcrRNA with a targeting window of 290bp. Parallel experiments assessing puromycin resistance, or global genomic transposition efficiency, and 6-TG, or selective hHPRT disruption induced resistance, both in the presence of either the hHPRT 6xcrRNA or scaffold gRNA control, were conducted to determine the specificity versus efficiency of targeted transposition via the variants or WT lbCas!2a_SB 100X. For the variants (WT, H165D and H165V) of the fusion protein, eukaryotic hHPRT targeting experiments showed 3-to-4-fold greater viable cell density compared to either the scaffold gRNA or the wells lacking NILV but did contain hHPRT 6xcrRNA following 6d of 6-TG. Similarly, puromycin selection for overall genomic integration of the puro containing transposon demonstrated a 4-5 fold greater percent of cell viability (Crystal Violet well occupancy quantification). The scaffold gRNA showed a similar effect to the wells lacking the PuroR transposon, which may indicate that the fusion construct may be less functional in the presence of a gRNA scaffold lacking a genomic targeting spacer.
[0127] Testing the chemogenetic variant. Experiments in mouse C2C12 cells demonstrated control of transposition by modulating nuclear localization and fusion of the SB100X to lbCas!2a in the presence of the A / C dimerizer. GFP positive colonies following 8D of puromycin selection indicate positive genomic transposition events (FIG. 3).
[0128] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.INFORMAL SEQUENCE LISTINGSEQ ID NO:1: SB100X mutated transposaseMGKSKEISQDLRKRIVDLHKSGSSLGAISKRLAVPRSSVQTIVRKYKHHGTTQPSYRSGR RRVLSPRDERTLVRKVQINPRTTAKDLVKMLEETGTKVSISTVKRVLYRHNLKGHSARK KPLLQNRHKKARLRFATAHGDKDRTFWRNVLWSDETKIELFGHND(H / D / L / V)RYVWR KKGEACKPKNTIPTVK(H / D / L / Q)GGGSIMLWGCFAAGGTGALHKIDG(I / S / N)MDAVQY VDILKQHLKTSVRKLKLGRKWVFQHDNDPKHTS(K / N)VVAKWLKDNKVKVLEWPSQS PDLNPIENLWAELKKRVRARRPTNLTQLHQLCQEEWAKIHPNYCGKLVEGYPKRLTQV KQFKGNATKYSEQ ID NO:2: truncated SBIOOx mutant transposase resulting from stop codon (*)MGKSKEISQDLRKRIVDLHKSGSSLGAISKRLAVPRSSVQTIVRKYKHHGTTQPSYRSGR RRVLSPRDERTLVRKVQINPRTTAKDLVKMLEETGTKVSISTVKRVLYRHNLKGHSARK KPLLQNRHKKARLRFATAHGDKDRTFWRNVLWSDETKIELFGHND(H / D / L / V)RYVWR KKGEACKPKNTIPTVK(H / D / L / Q)GGGSIMLWGCFAAGGTGALHKIDG(I / S / N)MDAVQY VDILKQHLKT S VRK LKLGRK W VFQI IDNDPK I IT S(K / N * )SEQ ID NO:3: lbCasl2a RuvC mutantSKLEKFTNCYSLSKTLRFKAIPVGKTQENIDNKRLLVEDEKRAEDYKGVKKLLDRYYLS FINDVLHSIKLKNLNNYISLFRKKTRTEKENKELENLEINLRKEIAKAFKGNEGYKSLFKK DIIETILPEFLDDKDEIALVNSFNGFTTAFTGFFDNRENMFSEEAKSTSIAFRCINENLTRYI SNMDIFEKVDAIFDKHEVQEIKEKILNSDYDVEDFFEGEFFNFVLTQEGIDVYNAIIGGFV TESGEKIKGLNEYINLYNQKTKQKLPKFKPLYKQVLSDRESLSFYGEGYTSDEEVLEVFR NTLNKNSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEY DDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEI1IQKVDEIY KVYGSSEKLFDADFVLEKSLKKNDAVVAIMKDLLDSVKSFENYIKAFFGEGKETNRDES FYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSKDKFKLYFQNPQFMGGWDKDKETDYR ATILRYGSKYYLAIMDKKYAKCLQKIDKDDVNGNYEKINYKLLPGPNKMLPKVFFSKK WMAYYNPSEDIQKIYKNGTFKKGDMFNLNDCHKLIDFFKDSISRYPKWSNAYDFNFSET EKYKDIAGFYREVEEQGYKVSFESASKKEVDKLVEEGKLYMFQIYNKDFSDKSHGTPNL HTMYFKLLFDENNHGQIRLSGGAELFMRRASLKKEELVVHPANSPIANKNPDNPKKTTT LSYDVYKDKRFSEDQYELHIPIAINKCPKNIFKINTEVRVLLKHDDNPYVIGIARGERNLL YIVVVDGKGNIVEQYSLNEIINNFNGIRIKTDYHSLLDKKEKERFEARQNWTSIENIKELK AGYISQVVHKICELVEKYDAVIALEDLNSGFKNSRVKVEKQVYQKFEKMLIDKLNYMV DKKSNPCATGGALKGYQITNKFESFKSMSTQNGFIFYIPAWLTSKIDPSTGFVNLLKTKY T SIAD SKKFIS SFDRIMYVPEEDLFEF ALD YKNF SRTD AD YIKKWKLYS YGNRI IFRNPK KNNVFDWEEVCLTSAYKELFNKYGINYQQGDIRALLCEQSDKAFYSSFMALMSLMLQ MRNSITGRTDVDFLISPVKNSDGIFYDSRNYEAQENAILPKNADANGAYNIARKVLWAI GQFKKAEDEKLDKVKIAISNKEWLEYAQTSVKHSEQ ID NO:4: linkerKLGGGAPAVGGGPKSEQ ID NO:5 : fusion protein: lbCasl2a RuvC mutant, linker, and SBIOOX transposase (optionally mutated)SKLEKFTNCYSLSKTLRFKAIPVGKTQENIDNKRLLVEDEKRAEDYKGVKKLLDRYYLS FINDVLHSIKLKNLNNYISLFRKKTRTEKENKELENLEINLRKEIAKAFKGNEGYKSLFKKSNMDIFEKVDAIFDKHEVQEIKEKILNSDYDVEDFFEGEFFNFVLTQEGIDVYNAIIGGFV TESGEKIKGLNEYINLYNQKTKQKLPKFKPLYKQVLSDRESLSFYGEGYTSDEEVLEVFR NTLNKNSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEY DDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEIIIQKVDEIY KVYGSSEKLFDADFVLEKSLKKNDAVVAIMKDLLDSVKSFENYIKAFFGEGKETNRDES FYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSKDKFKLYFQNPQFMGGWDKDKETDYR ATILRYGSKYYLAIMDKKYAKCLQKIDKDDVNGNYEKINYKLLPGPNKMLPKVFFSKK WMAYYNPSEDIQKIYKNGTFKKGDMFNLNDCHKLIDFFKDSISRYPKWSNAYDFNFSET EKYKDIAGFYREVEEQGYKVSFESASKKEVDKLVEEGKLYMFQIYNKDFSDKSHGTPNL HTMYFKLLFDENNHGQIRLSGGAELFMRRASLKKEELVVHPANSPIANKNPDNPKKTTT L S YD V YKDKRF SEDQ YELHIPIAINKCPKNIFKINTE VR VLLK HDDNP Y V I GI ARGERNLL YIVWDGKGNIVEQYSLNEIINNFNGIRIKTDYHSLLDKKEKERFEARQNWTSIENIKELK AGYISQVVHKICELVEKYDAVIALEDLNSGFKNSRVKVEKQVYQKFEKMLIDKLNYMV DKKSNPCATGGALKGYQITNKFESFKSMSTQNGFIFYIPAWLTSKIDPSTGFVNLLKTKY T SIAD SKKFIS SFDRIMYVPEEDLFEF ALD YKNF SRTD AD YIKKWKLYS YGNRIRIFRNPKKNNVFDWEEVCLTSAYKELFNKYGINYQQGDIRALLCEQSDKAFYSSFMALMSLMLQ MRNSITGRTDVDFLISPVKNSDGIFYDSRNYEAQENAILPKNADANGAYNIARKVLWAI GQFKKAEDEKLDKVKIAISNKEWLEYAQTSVKHKLGGGAPAVGGGPKMGKSKEISQDL RKRIVDLHKSGSSLGAISKRLAVPRSSVQTIVRKYKHHGTTQPSYRSGRRRVLSPRDERT LVRKVQINPRTTAKDLVKMLEETGTKVSISTVKRVLYRHNLKGHSARKKPLLQNRHKK ARLRFATAHGDKDRTFWRNVLWSDETKIELFGHND(H / D / L / V)RYVWRKKGEACKPKN TIPTVK(H / D / L / Q)GGGSIMLWGCFAAGGTGALHKIDG(I / S / N)MDAVQYVDILKQHLKTS VRKLKLGRKWVFQHDNDPKHTS(K / N)VVAKWLKDNKVKVLEWPSQSPDLNPIENLWA ELKKRVRARRPTNLTQLHQLCQEEWAKIHPNYCGKLVEGYPKRLTQVKQFKGNATKYSEQ ID NO:6 : left terminal inverted repeatCAGTTGAAGTCGGAAGTTTACATACACTTAAGTTGGAGTCATTAAAACTCGTTTTTC AACTACTCCACAAATTTCTTGTTAACAAACAATAGTTTTGGCAAGTCAGTTAGGACA TCTACTTTGTGCATGACACAAGTCATTTTTCCAACAATTGTTTACAGACAGATTATTT CACTTATAATTCACTGTATCACAATTCCAGTGGGTCAGAAGTTTACATACACTAASEQ ID NO:7: right terminal inverted repeatGAGTGTATGTAAACTTCTGACCCACTGGGAATGTGATGAAAGAAATAAAAGCTGAA ATGAATCATTCTCTCTACTATTATTCTGATATTTCACATTCTTAAAATAAAGTGGTGA TCCTAACTGACCTAAGACAGGGAATTTTTACTAGGATTAAATGTCAGGAATTGTGAA AAAGTGAGTTTAAATGTATTTGGCTAAGGTGTATGTAAACTTCCGACTTCAACTGSEQ ID NO:8 : fusion protein comprising a first and second affinity agentGAACTGGCTTATATCAACACAAAACATCCAGACTTTGGCGGAAGCGGTGGCGTGGGAGTGCAGGTGGAAACCATCTCCCCAGGCGACGGGCGCACCTTCCCCAAGCGCGGCCAGACCTGCGTGGTGCACTACACCGGGATGCTTGAAGATGGAAAGAAATTTGATTCCTCCCGGGACAGAAACAAGCCCTTTAAGTTTATGCTAGGCAAGCAGGAGGTGATCCGAGGCTGGGAAGAAGGGGTTGCCCAGATGAGTGTGGGTCAGAGAGCCAAACTGACTATATCTCCAGATTATGCCTATGGTGCCACTGGGCACCCAGGCATCATCCCACCACATGCCACTCTCGTCTTCGATGTGGAGCTTCTAAAACTGGAACTCGAGAATTCTCACGCGTCTGCCACCAACTTCAGCCTGCTGAAGCAGGCCGGCGACGTGGAGGAGAACCCCGGCCCCGCAGGATATCAAGCTTCCACCATCCTCTGGCATGAGATGTGGCATGAAGGCCTGGAAGAGGCATCTCGTTTGTACTTTGGGGAAAGGAACGTGAAAGGCATGTTTGAGGTGCTGGAGCCCTTGCATGCTATGATGGAACGGGGCCCCCAGACTCTGAAGGAAACATCCTTTAATCAGGCCTATGGTCGAGATTTAATGGAGGCCCAAGAGTGGTGCAGGAAGTACATGAAATCAGGGAATGTCAAGGACCTCCTCCAAGCCTGGGACCTCTATTATCATGTGTTCCGACGAATCTCAAAGCTCGAGGTGGAATTCGCTGATGCTTGTGGGCTAATGAACAATAATATAGAGSEQ ID NO:9: transposase variant (165Leu)TGGTCTGATGAAACAAAAATAGAACTGTTTGGCCATAATGACCTTCGTTATGTTTGGAGGAAGAAGGGGGAGGCTTGCAAGCCGAAGAACACCATCCCAACCGTGAAGCACGGGGGTGGCAGCATCATGTTGTGGGGGTGCTTTGCTGCAGGAGGGACTGGTGCACTTCACAAAATAGATGGCATCATGGACGCGGTGCAGTATGTGGATATATTGAAGCAACATCTCAAGACATCAGTCAGGAAGTTAAAGCTTGGTCGCAAATGGGTCTTCCAACACGACAATGACCCCAAGCATACTTCCAAAGTTGTGGCAAAATGGCTTAAGGACAACAAAGTCAAGGTATTGGAGTGGCCATCACAAAGCCCTGACCTCAATCCTATAGAAAATTTGTGGGCAGAACTGAAAAAGCGTGTGCGAGCAAGGAGGCCTACAAACCTGACTCAGTTACACCAGCTCTGTCAGGAGGAATGGGCCAAAATTCACCCAAATTATTGTGGGAAGCTTGTGGAAGGCTACCCGAAACGTTTGACCCAAGTTAAACAATTTAAAGGCAATGCTACCAAATACTGAGGGCCCTCGCGGGTAATGAACTAGTACCGGTTAAGTCGACAATCAACGCGSEQ ID NO:10: transposase variant (165Asp)TGGTCTGATGAAACAAAAATAGAACTGTTTGGCCATAATGACGATCGTTATGTTTGGAGGAAGAAGGGGGAGGCTTGCAAGCCGAAGAACACCATCCCAACCGTGAAGCACGGGGGTGGCAGCATCATGTTGTGGGGGTGCTTTGCTGCAGGAGGGACTGGTGCACTTCACAAAATAGATGGCATCATGGACGCGGTGCAGTATGTGGATATATTGAAGCAACATCTCAAGACATCAGTCAGGAAGTTAAAGCTTGGTCGCAAATGGGTCTTCCAACACGACAATGACCCCAAGCATACTTCCAAAGTTGTGGCAAAATGGCTTAAGGACAACAAAGTCAAGGTATTGGAGTGGCCATCACAAAGCCCTGACCTCAATCCTATAGAAAATTTGTGGGCAGAACTGAAAAAGCGTGTGCGAGCAAGGAGGCCTACAAACCTGACTCAGTTACACCAGCTCTGTCAGGAGGAATGGGCCAAAATTCACCCAAATTATTGTGGGAAGCTTGTGGAAGGCTACCCGAAACGTTTGACCCAAGTTAAACAATTTAAAGGCAATGCTACCAAATACTGAGGGCCCTCGCGGGTAATGAACTAGTACCGGTTAAGTCGACAATCAACGCGSEQ ID NO:11: transposase variant (165 Vai)TGGTCTGATGAAACAAAAATAGAACTGTTTGGCCATAATGACCGTCGTTATGTTTGGAGGAAGAAGGGGGAGGCTTGCAAGCCGAAGAACACCATCCCAACCGTGAAGCACGGGGGTGGCAGCATCATGTTGTGGGGGTGCTTTGCTGCAGGAGGGACTGGTGCACTTCACAAAATAGATGGCATCATGGACGCGGTGCAGTATGTGGATATATTGAAGCAACATCTCAAGACATCAGTCAGGAAGTTAAAGCTTGGTCGCAAATGGGTCTTCCAACACGACAATGACCCCAAGCATACTTCCAAAGTTGTGGCAAAATGGCTTAAGGACAACAAAGTCAAGGTATTGGAGTGGCCATCACAAAGCCCTGACCTCAATCCTATAGAAAATTTGTGGGCAGAACTGAAAAAGCGTGTGCGAGCAAGGAGGCCTACAAACCTGACTCAGTTACACCAGCTCTGTCAGGAGGAATGGGCCAAAATTCACCCAAATTATTGTGGGAAGCTTGTGGAAGGCTACCCGAAACGTTTGACCCAAGTTAAACAATTTAAAGGCAATGCTACCAAATACTGAGGGCCCTCGCGGGTAATGAACTAGTACCGGTTAAGTCGACAATCAACGCGSEQ ID NO:12: transposase variant (252Asn)TGGTCTGATGAAACAAAAATAGAACTGTTTGGCCATAATGACCATCGTTATGTTTGGAGGAAGAAGGGGGAGGCTTGCAAGCCGAAGAACACCATCCCAACCGTGAAGCACGGGGGTGGCAGCATCATGTTGTGGGGGTGCTTTGCTGCAGGAGGGACTGGTGCACTTCACAAAATAGATGGCATCATGGACGCGGTGCAGTATGTGGATATATTGAAGCAACATCTCAAGACATCAGTCAGGAAGTTAAAGCTTGGTCGCAAATGGGTCTTCCAACACGACAATGACCCCAAGCATACTTCCAATGTTGTGGCAAAATGGCTTAAGGACAACAAAGTCAAGGTATTGGAGTGGCCATCACAAAGCCCTGACCTCAATCCTATAGAAAATTTGTGGGCAGAACTGAAAAAGCGTGTGCGAGCAAGGAGGCCTACAAACCTGACTCAGTTACACCAGCTCTGTCAGGAGGAATGGGCCAAAATTCACCCAAATTATTGTGGGAAGCTTGTGGAAGGCTACCCGAAACGTTTGACCCAAGTTAAACAATTTAAAGGCAATGCTACCAAATACTGAGGGCCCTCGCGGGTAATGAACTAGTACCGGTTAAGTCGACAATCAACGCGSEQ ID NO: 13: transposase variant (252Tyr)TGGTCTGATGAAACAAAAATAGAACTGTTTGGCCATAATGACCATCGTTATGTTTGGAGGAAGAAGGGGGAGGCTTGCAAGCCGAAGAACACCATCCCAACCGTGAAGCACGGGGGTGGCAGCATCATGTTGTGGGGGTGCTTTGCTGCAGGAGGGACTGGTGCACTTCACAAAATAGATGGCATCATGGACGCGGTGCAGTATGTGGATATATTGAAGCAACATCTCAAGACATCAGTCAGGAAGTTAAAGCTTGGTCGCAAATGGGTCTTCCAACACGACAATGACCCCAAGCATACTTCCTATGTTGTGGCAAAATGGCTTAAGGACAACAAAGTCAAGGTATTGGAGTGGCCATCACAAAGCCCTGACCTCAATCCTATAGAAAATTTGTGGGCAGAACTGAAAAAGCGTGTGCGAGCAAGGAGGCCTACAAACCTGACTCAGTTACACCAGCTCTGTCAGGAGGAATGGGCCAAAATTCACCCAAATTATTGTGGGAAGCTTGTGGAAGGCTACCCGAAACGTTTGACCCAAGTTAAACAATTTAAAGGCAATGCTACCAAATACTGAGGGCCCTCGCGGGTAATGAACTAGTACCGGTTAAGTCGACAATCAACGCG
Claims
WHAT IS CLAIMED IS:
1. A polypeptide comprising a transposase at least 90% identical to SEQ ID NO: 1 or SEQ ID NO:2, wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
2. The polypeptide of claim 1, wherein the transposase comprises SEQ ID NO: 1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO: 2 is not wildtype.
3. A polypeptide comprising a transposase linked to a DNA sequencetargeting protein, wherein the transposase is at least 90% identical to SEQ ID NO: 1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO:1 or SEQ ID N0:2 is not wildtype.
4. The polypeptide of claim 3, wherein the transposase comprises SEQ ID NO: 1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
5. The polypeptide of claim 3, wherein the polypeptide comprises SEQ ID NO:5.
6. The polypeptide of any one of claims 3-5, wherein the DNA sequencetargeting protein is a CRISPR / Cas protein, a TALEN protein, a zinc finger protein or a meganuclease protein.
7. The polypeptide of claim 6, wherein the DNA sequence-targeting protein is a Casl2a or Cas9 protein.
8. The polypeptide of any one of claims 6-7, wherein the DNA sequencetargeting protein has nicking activity.
9. The polypeptide of any one of claims 6-7, wherein the DNA sequencetargeting protein lacks nicking activity.
10. The polypeptide of claim 9, wherein the DNA sequence-targeting protein is dCasl2a or dCas9.
11. The polypeptide of claim 10, wherein the dCasl2a protein comprises SEQ ID NO:3.
12. The polypeptide of any one of claims 3-11, wherein the polypeptide further comprises one or more nuclear localization signal (NLS) sequence.
13. The polypeptide of any one of claims 3-12, wherein the transposase linked to the DNA sequence-targeting protein is a translational fusion, optionally linked via a peptide linker.
14. The polypeptide of any one of claims 3-13, wherein the transposase is fused to a first affinity agent and the DNA sequence-targeting protein is linked to a second affinity agent and the first and second affinity agents bind.
15. The polypeptide of claim 14, wherein the first and second affinity agents bind in the presence of a chemical agent but not in the absence of the chemical agent.
16. The polypeptide of claim 15, wherein the first and second affinity agents are selected from FRB and FKBP and the chemical agent is rapamycin.
17. A polynucleotide encoding the polypeptide of any one of claims 1-16.
18. An expression cassette comprising a promoter operably linked to the polynucleotide of claim 17.
19. A vector comprising the expression cassette of claim 18.
20. The vector of claim 19, wherein the vector is a viral vector.
21. The vector of claim 20, wherein the viral vector is a lentiviral, adenoviral, or adeno-associated viral vector.
22. The vector of claim 21, wherein the lentiviral vector is a non-integrating lentiviral vector.
23. A method of editing the genome of a eukaryotic cell, the method comprising, delivering the polypeptide of any one of claims 3-16 to the nucleus of the cell, wherein the polypeptide nicks the genome of the eukaryotic cell in at least one location and introduces a template polynucleotide into the nick, thereby editing the genome of the cell.
24. The method of claim 23, wherein the delivering the polypeptide comprises contacting the cell with a vector encoding the polypeptide, and wherein the polypeptide is expressed in the cell and enters the nucleus of the cell.
25. The method of any one of claims 23-24, wherein the method is performed in vitro, ex vivo or in vivo.
26. The method of any one of claims 23-25, further comprising delivering the template polynucleotide to the cell.
27. The method of claim 26, wherein the template polynucleotide is delivered to the cell in a viral vector.
28. The method of claim 27, wherein the viral vector is a lentiviral, adenoviral, or adeno-associated viral vector.
29. The method of any one of claims 23-28, wherein the template polynucleotide is at least 10, 20, 30, 40, 50, 70, or 100 kb long.
30. The method of any one of claims 23-29, wherein the polypeptide comprises a transposase linked to a DNA sequence-targeting protein, wherein the transposase is at least 90% identical to SEQ ID NO: 1 or SEQ ID NO:2, and wherein at least one of the positions corresponding to wildtype positions H165, H187, 1212, and K252 of SEQ ID NO: 1 or SEQ ID NO:2 is not wildtype.
31. The method of any one of claims 23-30, wherein the DNA sequencetargeting protein is a CRISPR / Cas protein, and the method further comprises delivering one or more guide RNAs (gRNAs) to the cell.
32. The method of claim 31, wherein two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) gRNAs are delivered to the cell.
33. The method of any one of claims 31-32, wherein the DNA sequencetargeting protein is a Casl2a or Cas9 protein.
34. The method of any one of claims 31-33, wherein the DNA sequencetargeting protein has nicking activity.
35. The method of any one of claims 31-34, wherein the DNA sequencetargeting protein lacks nicking activity.
36. The method of claim 35, wherein the DNA sequence-targeting protein is dCasl2a or dCas9.
37. The method of claim 36, wherein the DNA sequence-targeting protein comprises SEQ ID NO:3.