Fusion proteins comprising Cas12a polypeptides and inteins and methods of use thereof
Through the method of binding the Cas12a polypeptide to the inteptic polypeptide fusion protein to guide nucleic acid, the delivery problem of large genome editing agents in adeno-associated virus applications is solved, and effective nucleic acid modification and editing is achieved.
Patent Information
- Application Number
- CN202380085815.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-15
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art is difficult to effectively deliver large genome editing agents to application scenarios requiring the use of adeno-associated viruses, such as delivery to the subject's brain, limiting the choice of delivery methods.
Using a method that contains a fusion protein of Cas12a polypeptide and an intrapeptide polypeptide, combined with a guide nucleic acid and a deaminase, is used to modify or edit the target nucleic acid, and forms a complex with the guide nucleic acid to achieve genome editing.
New delivery routes are provided, which can effectively modify or edit target nucleic acids, and are suitable for application scenarios that require delivery of adeno-associated viruses.
Smart Images

Figure CN120500534A_ABST
Abstract
Description
[0001] Declaration concerning the electronic file of the sequence listing
[0002] The disclosure of the XML-formatted sequence listing named 1499-116_ST26.xml, 502,046 bytes in size, generated on December 14, 2023, and submitted with this application is hereby incorporated by reference in its entirety. Technical Field
[0003] The present invention relates to fusion proteins (e.g., engineered proteins) comprising Cas12a polypeptides and intein polypeptides and methods for using such proteins. The present invention also relates to fusion proteins (e.g., engineered proteins) comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) and an intein polypeptide and methods for using such proteins. The present invention also relates to compositions and systems for modifying or editing target nucleic acids. Background Art
[0004] Large genome editing agents (e.g., editing agents with a length of more than 2000 amino acids) are typically delivered to tissues using means other than adeno-associated virus (AAV) vectors. For example, large genome editing agents can be delivered using lipid nanoparticle-mediated RNP delivery, which has no size restrictions, or using mRNA or DNA. However, these methods prohibit use in applications requiring the use of AAV, such as for delivery to the brain of a subject.
[0005] Therefore, new methods for preparing and / or delivering genome editing agents are needed. Summary of the Invention
[0006] A first aspect of the present invention is directed to a fusion protein comprising an intein polypeptide. In some embodiments, the fusion protein comprises a Cas12a polypeptide fused to an intein polypeptide. In some embodiments, the fusion protein comprises a polypeptide of interest fused to an intein polypeptide. In some embodiments, the fusion protein comprises a reverse transcriptase polypeptide fused to an intein polypeptide. Also provided is a nucleic acid molecule encoding a fusion protein as described herein.
[0007] Another aspect of the present invention is directed to a complex comprising: a Cas12a protein prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide; a guide nucleic acid (e.g., a guide RNA); and optionally a deaminase.
[0008] Another aspect of the invention is directed to a complex comprising: an engineered protein (e.g., a base editor or a templated editor, such as a REDRAW editor) prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA).
[0009] Another aspect of the present invention is directed to a method for modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with: a Cas12a protein prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA), optionally wherein the Cas12a protein and the guide nucleic acid form a complex or are contained in a complex.
[0010] Another aspect of the present invention is directed to a method for modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the following to modify the target nucleic acid: an engineered protein (e.g., a base editor or a templated editor, such as a REDRAW editor) prepared by a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA), optionally wherein the engineered protein and the guide nucleic acid form a complex or are contained in a complex.
[0011] Another aspect of the present invention is directed to a composition comprising: a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide.
[0012] Another aspect of the invention is directed to a composition comprising: a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide.
[0013] Another aspect of the present invention is directed to a composition comprising: a first nucleic acid molecule encoding a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide.
[0014] Another aspect of the present invention is directed to a composition comprising: a first nucleic acid molecule encoding a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide.
[0015] Another aspect of the present invention is directed to a kit comprising: a first nucleic acid molecule encoding a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide.
[0016] Another aspect of the present invention is directed to a kit comprising: a first nucleic acid molecule encoding a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide.
[0017] Another aspect of the present invention is directed to a method for modifying a target nucleic acid, the method comprising: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell, wherein the first nucleic acid molecule encodes a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a protein comprising at least a portion of the first Cas12a polypeptide and at least a portion of the second Cas12a polypeptide, and a guide nucleic acid (e.g., a guide RNA), optionally wherein the protein and the guide nucleic acid form a complex or are contained in a complex, thereby modifying the target nucleic acid.
[0018] Another aspect of the invention is directed to a method for modifying a target nucleic acid, the method comprising: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell, wherein the first nucleic acid molecule encodes a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a protein comprising at least a portion of the polypeptide of interest and at least a portion of the Cas12a polypeptide and a guide nucleic acid (e.g., a guide RNA), optionally wherein the protein and the guide nucleic acid form a complex or are contained in a complex, thereby modifying the target nucleic acid.
[0019] The present invention also provides expression cassettes and / or vectors comprising nucleic acid constructs of the present invention, and cells comprising polypeptides of the present invention, fusion proteins and / or nucleic acid constructs. In addition, the present invention provides test kits comprising nucleic acid constructs of the present invention and expression cassettes, vectors and / or cells comprising said nucleic acid constructs.
[0020] It should be noted that aspects of the invention described with respect to one embodiment may be incorporated into different embodiments, even though no specific description has been made therewith. That is, all embodiments and / or features of any embodiment may be combined in any manner and / or combination. Applicants reserve the right to change any initially filed claim and / or to submit any new claim accordingly, including the right to amend any initially filed claim to be subordinate to and / or incorporate any features of any other claim, even though the claim was not initially made in this manner. These and other objects and / or aspects of the invention will be explained in detail in the specification set forth below. Those skilled in the art will understand further features, advantages and details of the invention from reading the accompanying drawings and the detailed description of the preferred embodiments that follow, such description being merely illustrative of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Schematic diagram of two exemplary split proteins of mCherry.
[0022] Figure 2Graphs showing average mCherry fluorescence for different proteins, which can measure intein splicing activity for mCherry reconstitution assays. (1) Wild-type full-length mCherry expression. (2) Expression of only N-terminal mCherry fused to an N-terminal Npu split intein. (3) Expression of only C-terminal mCherry fused to a C-terminal Npu split intein with a GEP mutation that enhances intein activity and robustness. (4) Expression of only C-terminal mCherry fused to a wild-type (WT) C-terminal Npu split intein. (5) Expression of both N- and C-terminal mCherry portions using a wild-type variant of the Npu intein. (6) Expression of both N- and C-terminal mCherry portions using a GEP variant of the Npu intein.
[0023] Figure 3 is a schematic diagram of two split proteins of Redraw Editor 2 (RE2) comprising Cas12a, showing exemplary split sites according to some embodiments of the present invention.
[0024] Figure 4 is a graph showing the effect of two amino acid insertions (eg, CF or CA residues) on mCherry fluorescence.
[0025] Figure 5 is a graph showing the percentage of inversion-deletions generated using non-split and split Redraw editors using crRNA according to some embodiments of the present invention.
[0026] Figure 6 is a graph showing the percentage of inversion-deletions generated using non-split and split Redraw editors using stagRNA according to some embodiments of the present invention.
[0027] Figure 7 is a graph showing the precise percentage of edits produced using non-split and split Redraw editors using stagRNA according to some embodiments of the present invention. DETAILED DESCRIPTION
[0028] The present invention will now be described hereinafter with reference to the accompanying drawings and examples, in which embodiments of the invention are shown. This detailed description is not intended to be an exhaustive list of all the different ways in which the invention may be implemented or all the features that may be added to the invention. For example, features shown with respect to one embodiment may be incorporated into other embodiments, and features shown with respect to a particular embodiment may be deleted from the embodiment. Therefore, the present invention contemplates that in some embodiments of the invention, any feature or combination of features set forth herein may be excluded or omitted. In addition, in light of this disclosure, many variations and additions to the various embodiments suggested herein will be apparent to those skilled in the art without departing from the invention. Therefore, the following description is intended to illustrate some specific embodiments of the invention rather than to exhaustively describe all permutations, combinations, and variations thereof.
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention pertains. The terms used in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0030] All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentence and / or paragraph in which the reference is presented.
[0031] Unless the context indicates otherwise, it is specifically intended that the various features of the invention described herein may be used in any combination. Furthermore, the present invention contemplates that in some embodiments of the invention, any feature or combination of features set forth herein may be excluded or omitted. For illustration, if the specification states that a composition comprises components A, B, and C, it is specifically intended that any one of A, B, or C, or any combination thereof, may be omitted or disclaimed, individually or in any combination.
[0032] As used in the description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0033] Also as used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0034] As used herein, the term "about" when referring to a measurable value, such as an amount or concentration, is intended to encompass variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value as well as the specified value. For example, "about X," where X is a measurable value, is intended to encompass X as well as variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. The ranges of measurable values provided herein may include any other ranges and / or individual values therein.
[0035] As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y," and phrases such as "from about X to Y" mean "from about X to about Y."
[0036] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a range of 10 to 15 is disclosed, 11, 12, 13, and 14 are also disclosed.
[0037] As used herein, the terms “comprising,” “including,” and “having” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0038] As used herein, the transition phrase "consisting essentially of means that the scope of a claim should be interpreted to encompass the specified materials or steps recited in the claim, as well as materials or steps that do not materially affect the basic and novel characteristics of the claimed invention. Therefore, when used in the claims of the present invention, the term "consisting essentially of is not intended to be interpreted as equivalent to "comprising."
[0039] As used herein, the terms "increase," "increasing," "enhance," "enhancing," and "improve," "improving" (and grammatical variations thereof) describe, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to another measurable property or quantity (e.g., a control value).
[0040] As used herein, the terms "reduce," "reduced," "reducing," "reduction," "diminish," and "decrease" (and grammatical variations thereof) describe, for example, a decrease of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% compared to another measurable property or quantity (e.g., a control value). In some embodiments, the reduction can result in no or substantially no (i.e., a negligible amount, e.g., less than about 10% or even 5%) detectable activity or amount.
[0041] A "heterologous nucleotide sequence" or "recombinant nucleotide sequence" is a nucleotide sequence that is not naturally associated with a host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.
[0042] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a "native nucleic acid" is a nucleic acid that occurs naturally in or is endogenous to the reference organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.
[0043] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleotide sequence," and "polynucleotide" refer to linear or branched, single-stranded or double-stranded RNA or DNA, or hybrids thereof. The terms also encompass RNA / DNA hybrids. When dsRNA is produced synthetically, less common bases such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, and the like may also be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind RNA with high affinity and are potent antisense inhibitors of gene expression. Other modifications, such as modifications to the 2'-hydroxyl group in the phosphodiester backbone or the RNA ribose group, may also be made.
[0044] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5' to 3' ends of a nucleic acid molecule, and includes DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which can be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "recombinant nucleic acid," "oligonucleotide," and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. Nucleic acid molecules and / or nucleotide sequences provided herein are presented in a 5' to 3' direction from left to right in this article, and are represented by the standard code for representing nucleotide characters specified in U.S. sequence rules 37 CFR §§ 1.821-1.825 and World Intellectual Property Organization (WIPO) standard ST.25. As used herein, "5' district" can represent the polynucleotide district closest to the 5' end of a polynucleotide. Therefore, for example, the element in the 5' district of a polynucleotide can be located at any position of the nucleotide from the first nucleotide at the 5' end of the polynucleotide to the nucleotide in the middle of the polynucleotide. As used herein, "3' region" can refer to the region of a polynucleotide closest to the 3' end of a polynucleotide. Thus, for example, elements in the 3' region of a polynucleotide can be located anywhere from the first nucleotide at the 3' end of the polynucleotide to a nucleotide in the middle of the polynucleotide.
[0045] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotides (AMOs), etc. A gene may or may not be capable of producing a functional protein or gene product. A gene may include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions).
[0046] A polynucleotide, gene, or polypeptide can be "isolated," meaning that the nucleic acid or polypeptide is substantially or essentially free from components that normally accompany the nucleic acid or polypeptide, respectively, in nature. In some embodiments, these components include other cellular material, culture medium from recombinant production, and / or various chemicals used to chemically synthesize the nucleic acid or polypeptide.
[0047] The term "mutation" refers to a point mutation (e.g., a missense or nonsense, or an insertion or deletion of a single base pair that causes a frameshift), an insertion, a deletion, and / or a truncation. When the mutation is a substitution of one residue within an amino acid sequence by another, or a deletion or insertion of one or more residues within the sequence, the mutation is typically described by identifying the original residue, followed by identifying the position of the residue within the sequence, and the identity of the newly substituted residue.
[0048] As used herein, the terms "complementary" or "complementarity" refer to the natural binding of polynucleotides through base pairing under permissive salt and temperature conditions. For example, the sequence "AGT" (5' to 3') binds to the complementary sequence "TCA" (3' to 5'). Complementarity between two single-stranded molecules can be "partial," where only some nucleotides bind, or complete, where perfect complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands.
[0049] As used herein, "complementary" can mean 100% complementary to a compared nucleotide sequence, or it can mean less than 100% complementary (e.g., "substantially complementary," e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% complementary, etc.).
[0050] A "portion" or "fragment" of a nucleotide sequence or polypeptide (including domains) is understood to refer to a nucleotide sequence or polypeptide that is reduced in length (e.g., by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more residues (e.g., nucleotides or peptides)) relative to a reference nucleotide sequence or polypeptide, respectively, and comprises, consists essentially of, and / or consists of, a portion of, or a fragment of a nucleotide sequence or polypeptide, respectively, that is reduced in length (e.g., by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more residues (e.g., nucleotides or peptides)). The nucleotide sequence or polypeptide consists of a nucleotide sequence or polypeptide consisting of consecutive residues that are identical or nearly identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical). In some embodiments, a portion of a reference nucleotide sequence or polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or more of the full-length reference nucleotide sequence or polypeptide. Such nucleic acid fragments or portions according to the present invention may, where appropriate, be included in a larger polynucleotide of which they are a component. As an example, the repeat sequence of the guide nucleic acid of the present invention may comprise a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type V-type CRISPR Cas repeat sequence, such as a repeat sequence from a CRISPR Cas system, including but not limited to Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b and / or Cas14c, etc.). Similarly, a portion of a polypeptide may be included in a larger polypeptide of which it is a component.
[0051] Different nucleic acids or proteins with homology are referred to as "homologs" in this article. The term homolog includes homologous sequences from the same species and other species and orthologous sequences from the same species and other species." Homology " refers to the level of similarity between two or more nucleic acids and / or amino acid sequences, expressed as a percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Therefore, the compositions and methods of the present invention further comprise homologs of the nucleotide sequences of the present invention and polypeptides. As used herein, "orthologs" and "orthologs" refer to homologous nucleotide sequences and / or amino acid sequences produced by common ancestral genes in different species during speciation. Homologs or orthologs of the nucleotide sequences of the invention may have substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to the nucleotide sequences of the invention.
[0052] As used herein, "sequence identity" refers to the degree to which two optimally aligned polynucleotide or polypeptide sequences are invariant over the entire component (e.g., nucleotide or amino acid) comparison window. "Identity" can be readily calculated by known methods, including but not limited to those described in Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM and Griffin, HG, eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).
[0053] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides in the linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "percent identity" may refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.
[0054] As used herein, the phrases "substantially identical" or "substantial identity" in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences refers to two or more sequences or subsequences that have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% nucleotide or amino acid residue identity when measured using one of the following sequence comparison algorithms or compared and aligned for maximum correspondence by visual inspection. In some embodiments of the invention, substantial identity exists over a region of contiguous nucleotides of a nucleotide sequence of the invention that is about 10 nucleotides to about 20 nucleotides, about 10 nucleotides to about 25 nucleotides, about 10 nucleotides to about 30 nucleotides, about 15 nucleotides to about 25 nucleotides, about 30 nucleotides to about 40 nucleotides, about 50 nucleotides to about 60 nucleotides, about 70 nucleotides to about 80 nucleotides, about 90 nucleotides to about 100 nucleotides, or more nucleotides in length, and any range therein, up to the full length of the sequence. In some embodiments, the nucleotide sequences may be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, substantially identical nucleotide or protein sequences perform substantially the same function as substantially identical nucleotide sequences (or encoded protein sequences).
[0055] For sequence comparison, typically one sequence acts as a reference sequence to which one or more test sequences are compared. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, subsequence coordinates are specified, if necessary, and sequence algorithm program parameters are specified. The sequence comparison algorithm then calculates the percent sequence identity of the test sequences relative to the reference sequences based on the specified program parameters.
[0056] Optimal alignment of sequences for comparison windows is well known to those skilled in the art and can be performed by tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, and optionally by computerized implementations of these algorithms, e.g., as Wisconsin (Accelrys Inc., San Diego, CA), and web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH2SEQ, EMBOSS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee. In some embodiments, the "best alignment" of two sequences (e.g., two polypeptide sequences) is the highest scoring alignment, optionally from alignments performed by tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, the similarity search method of Wisconsin In certain embodiments, " optimal alignment " of two sequences (for example, two peptide sequences) is the comparison that provides highest sequence identity percentage ratio, optionally allows one or more rooms to be introduced in one or two sequences. " identity score " for the comparison fragment of test sequence and reference sequence is the quantity of the same components shared by two comparison sequences divided by the total number of components in the reference sequence fragment (for example, the less defined part of whole reference sequence or reference sequence). Percent sequence identity is expressed as identity score and multiplied by 100. The comparison of one or more sequences can be the comparison with full-length sequence or its part, or the comparison with longer sequence. For purposes of the present invention, "percent identity" and / or optimal alignment can be determined using the Basic Local Alignment Search Tool (BLAST) provided by the National Center for Biotechnology Information, e.g., BLASTX for translated nucleotide sequences, BLASTN for polynucleotide sequences, and BLASTP for polypeptide sequences.
[0057] When two nucleotide sequences hybridize to each other under stringent conditions, the two sequences may also be considered to be substantially complementary.In some representative embodiments, two nucleotide sequences that are considered to be substantially complementary hybridize to each other under highly stringent conditions.
[0058] In the context of nucleic acid hybridization experiments (e.g., Southern and Northern hybridizations), "stringent hybridization conditions" and "stringent hybridization wash conditions" are sequence-dependent and are different under different environmental parameters. An extensive guide to nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays" Elsevier, New York (1993). In general, highly stringent hybridization and wash conditions are selected to be higher than the thermal melting point (Tf) for the specific sequence at a defined ionic strength and pH. m ) is about 5℃ lower.
[0059] T m It is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are chosen to be equal to the T for a specific probe. m. An example of stringent hybridization conditions for hybridization of complementary nucleotide sequences having more than 100 complementary residues on a filter membrane in a Southern or Northern blot is hybridization overnight with 50% formamide and 1 mg of heparin at 42°C. An example of highly stringent wash conditions is washing with 0.15 M NaCl at 72°C for about 15 minutes. An example of stringent wash conditions is washing with 0.2x SSC at 65°C for 15 minutes (for a description of SSC buffer, see Sambrook below). Typically, a low stringency wash is performed before a high stringency wash to remove background probe signal. For example, an example of a medium stringency wash for a duplex of more than 100 nucleotides is washing with 1x SSC at 45°C for 15 minutes. For example, an example of a low stringency wash for a duplex of more than 100 nucleotides is washing with 4-6x SSC at 40°C for 15 minutes. For short probes (e.g., about 10 to 50 nucleotides), stringent conditions typically involve a salt concentration of less than about 1.0 M Na ion, typically about 0.01 to 1.0 M Na ion concentration (or other salts) at pH 7.0 to 8.3, and a temperature of typically at least about 30°C. Stringent conditions can also be achieved by adding destabilizing agents such as formamide. In general, a signal-to-noise ratio of 2 times (or greater) that observed for an unrelated probe in a particular hybridization assay indicates that specific hybridization has been detected. Nucleotide sequences that do not hybridize to each other under stringent conditions are still substantially identical if the proteins encoded by the nucleotide sequences are substantially identical. This occurs, for example, when copies of a nucleotide sequence are created using the maximum codon degeneracy permitted by the genetic code.
[0060] The polynucleotides and / or recombinant nucleic acid constructs of the present invention can be codon-optimized for expression. In some embodiments, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the present invention (e.g., comprising / encoding a fusion protein, a nucleic acid binding polypeptide (e.g., a DNA binding polypeptide, e.g., a sequence-specific DNA binding domain from a polynucleotide-guided endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), an Argonaute protein, and / or a CRISPR-Cas effector protein), a guide nucleic acid, a cytosine deaminase, and / or an adenine deaminase) can be codon-optimized for expression in an organism (e.g., an animal (e.g., a human), a plant, a fungus, an archaebacteria, or a bacterium). In some embodiments, the codon-optimized nucleic acid constructs, polynucleotides, expression cassettes and / or vectors of the invention are about 70% to about 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100%) identical or more to a reference nucleic acid construct, polynucleotide, expression cassette and / or vector that has not been codon-optimized.
[0061] In any embodiment described herein, the polynucleotides or nucleic acid constructs of the present invention can be operably associated with a variety of promoters and / or other regulatory elements for expression in an organism or cell thereof (e.g., a mammal and / or mammalian cell, a plant and / or plant cell, etc.). Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the present invention may further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter can be operably associated with an intron (e.g., an Ubil promoter and introns). In some embodiments, a promoter associated with an intron can be referred to as a "promoter region" (e.g., an Ubil promoter and introns).
[0062] As used herein, "operably linked" or "operably associated" in reference to a polynucleotide means that the indicated elements are functionally related to each other, and usually also physically related. Thus, as used herein, the terms "operably linked" or "operably associated" refer to functionally related nucleotide sequences on a single nucleic acid molecule. Thus, a first nucleotide sequence that is operably linked to a second nucleotide sequence refers to a situation where the first nucleotide sequence is in a functional relationship with the second nucleotide sequence. For example, if a promoter affects the transcription or expression of a nucleotide sequence, then the promoter is operably associated with the nucleotide sequence. It will be understood by those skilled in the art that a control sequence (e.g., a promoter) need not be adjacent to the nucleotide sequence with which it is operably associated, as long as the function of the control sequence is to direct its expression. Thus, for example, there may be an intervening untranslated but transcribed nucleic acid sequence between a promoter and a nucleotide sequence, and the promoter may still be considered to be "operably linked" to the nucleotide sequence.
[0063] As used herein, the term "connection" or "fusion" in relation to a polypeptide refers to the covalent attachment of one polypeptide to another polypeptide. A polypeptide can be connected or fused to another polypeptide directly (e.g., via a peptide bond) or via a linker (e.g., a peptide linker) (e.g., at the N-terminus or C-terminus). Two polypeptides are directly fused (e.g., directly connected) to covalently attach an amino acid residue of a first polypeptide in the two polypeptides to an amino acid residue of a second polypeptide in the two polypeptides, without intervening elements between the two amino acid residues. For example, a first polypeptide and a second polypeptide can be directly connected via a peptide bond between the first polypeptide and the second polypeptide, without intervening elements (e.g., linkers) between the first polypeptide and the second polypeptide. Two polypeptides are indirectly fused (e.g., indirectly connected) to an insertion element (e.g., a linker, such as a peptide linker) between the two polypeptides, and the insertion element is covalently attached to each polypeptide, optionally wherein the insertion element can attach one end of the first polypeptide in the two polypeptides to one end of the second polypeptide in the two polypeptides.
[0064] As used herein, "fusion protein" refers to two or more polypeptides that are covalently linked (e.g., directly or indirectly) so that they are transcribed and translated as a single unit, thereby producing a single polypeptide comprising the two or more polypeptides. In some embodiments, the two or more polypeptides may be naturally encoded by separate genes, but are encoded by a single gene in the form of a fusion protein.
[0065] The term "connector" is recognized in the art and refers to a chemical group or molecule that connects two molecules or parts, such as two polypeptides or domains that connect a fusion protein, such as connecting a Cas12a polypeptide and an intein polypeptide. The connector may comprise a single connecting molecule (e.g., a single amino acid), or may comprise more than one connecting molecule. In some embodiments, the connector may be an organic molecule, a group, a polymer, or a chemical moiety, such as a divalent organic moiety. In some embodiments, the connector may be an amino acid, or may be a peptide. In some embodiments, the connector is a peptide (e.g., a peptide connector).
[0066] In some embodiments, the peptide linkers useful in the present invention can be from about 2 to about 100 or more amino acids in length, for example, about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, , 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 6 ...6 0, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids (e.g., about 105, 110, 115, 120, 130, 140, 150 or more amino acids in length). In some embodiments, the peptide linker may comprise glycine (G) and serine (S), such as a GS linker. In some embodiments, the peptide linker may comprise cysteine (C) and alanine (A), such as a CA linker. In some embodiments, the peptide linker may comprise cysteine (C) and phenylalanine (F), such as a CF linker. In some embodiments, the peptide linker is a GS linker, a CA linker, or a CF linker having 2, 3, or 4 amino acid residues, optionally 2 or 4 amino acid residues.In some embodiments, the peptide linker has one of the amino acid sequences of SEQ ID NOs: 1-35. In some embodiments, the peptide linker may comprise CA, CF, (GGS). n ,GS,SG,GSSG(SEQ ID NO:31),GSSGSS(SEQ ID NO:32),GSSGSSGS(SEQ ID NO:33),(GSS) n (SEQ ID NO:34), (GSS) n GS (SEQ ID NO: 35), S (GGS) n In some embodiments, the peptide linker comprises an amino acid sequence of (SEQ ID NO: 25), SGGS (SEQ ID NO: 26), (GSS)nG (SEQ ID NO: 191), or (GGGGS)n (SEQ ID NO: 27), wherein n is an integer from 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the peptide linker comprises the following amino acid sequence: SGGSGGSGGS (SEQ ID NO: 28). In some embodiments, the peptide linker comprises the following amino acid sequence: SGSETPGTSESATPES (SEQ ID NO: 29), also known as an XTEN linker. In some embodiments, the peptide linker comprises the following amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 30), also known as a GS-XTEN-GS linker. In some embodiments, the peptide linker has the amino acid sequence of SEQ ID NO: 189 or SEQ ID NO: 190.
[0067] As used herein, the term "connection" or "fusion" relating to polynucleotides refers to the covalent attachment of a polynucleotide to another polynucleotide. In certain embodiments, two or more polynucleotide molecules can be connected by a connexon, which can be an organic molecule, group, polymer, or chemical moiety, such as a divalent organic moiety. Polynucleotides can be connected or fused (at the 5' end or 3' end) to another polynucleotide via direct covalent attachment or via one or more connection nucleotides. In certain embodiments, the polynucleotide motif of a certain structure can be inserted into another polynucleotide sequence (e.g., an extension of a hairpin structure in a guide RNA). In certain embodiments, the connection nucleotides can be naturally occurring nucleotides. In certain embodiments, the connection nucleotides can be non-naturally occurring nucleotides. The direct fusion of two polynucleotides (e.g., directly connected) refers to the covalent attachment of a nucleotide of the first polynucleotide in two polynucleotides to a nucleotide of the second polynucleotide in two polynucleotides, without an insertion element between the two polynucleotides. For example, the first polynucleotide and the second polynucleotide can be directly connected via a phosphodiester bond between the first polynucleotide and the second polynucleotide, without an insertion element (e.g., connexon) between the first polynucleotide and the second polynucleotide. Indirect fusion (e.g., indirect linkage) of two polynucleotides means that there is an intervening element (e.g., a linker, e.g., a polynucleotide linker) between the two polynucleotides, and the intervening element is covalently linked to each polynucleotide, optionally wherein the intervening element links one end of the first of the two polynucleotides to one end of the second of the two polynucleotides.
[0068] "Promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) that is operably associated with the promoter. The coding sequence controlled or regulated by the promoter can encode a polypeptide and / or functional RNA. Generally, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and guides transcription initiation. In general, a promoter is located 5' or upstream of the start of the coding region relative to the corresponding coding sequence. A promoter may include other elements that act as gene expression regulators; for example, a promoter region. These include a TATA box consensus sequence, and generally also include a CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50: 349). In plants, the CAAT box can be replaced by an AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region may comprise at least one intron (eg, SEQ ID NO: 36 or SEQ ID NO: 37).
[0069] Promoters useful in the present invention may include, for example, constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred, and / or tissue-specific promoters for use in preparing recombinant nucleic acid molecules, e.g., "synthetic nucleic acid constructs" or "protein-RNA complexes." These different types of promoters are known in the art.
[0070] The selection of promoter can be different because of the time and space requirement of expression, also can be different because of the host cell that will transform.The promoter of many different organisms is well-known in the art.Based on the extensive knowledge that exists in this area, can be to select suitable promoter for interested specific host organism.Therefore, for example, the promoter of the gene upstream of highly constitutive expression in model organism is known very much, and this knowledge can be obtained at easy speed, and in appropriate time, implement in other system.
[0071] In certain embodiments, a promoter functional in plants can be used together with the construct of the present invention. Non-limiting examples of promoters that can be used to drive expression in plants include the promoter of RubisCo small subunit gene 1 (PrbcS1), the promoter of actin gene (Pactin), the promoter of nitrate reductase gene (Pnr) and the promoter of repeat carbonic anhydrase gene 1 (Pdca1) (see Walker et al., Plant Cell Rep. 23:727-735 (2005); Li et al., Gene 403:132-142 (2007); Li et al., Mol Biol. Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, and Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al., Gene 403: 132-142 (2007)), and Pdca1 is induced by salt (Li et al., Mol Biol. Rep. 37: 1143-1154 (2010)). In some embodiments, the promoter useful in the present invention is an RNA polymerase II (Pol II) promoter. In some embodiments, the U6 promoter or 7SL promoter from corn (Zea mays) can be used in the constructs of the present invention. In some embodiments, the U6c promoter and / or the 7SL promoter from corn can be used to drive expression of the guide nucleic acid. In some embodiments, the U6c promoter, U6i promoter and / or the 7SL promoter from soybean (Glycine max) can be used in the constructs of the present invention. In some embodiments, the U6c promoter, U6i promoter and / or the 7SL promoter from soybean can be used to drive expression of the guide nucleic acid.
[0072] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Pat. No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and U.S. Pat. No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci. USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci. USA 84:6624-6629), sucrose synthase promoter (Yang and Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148) and ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants, for example, sunflower (Binet et al., 1991. Plant Science 79:87-94), corn (Christensen et al., 1989. Plant Molec. Biol. 12:619-632) and Arabidopsis thaliana (Norris et al., 1993. Plant Molec. Biol. 21:895-906). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocot systems, and its sequence and vectors constructed for monocot transformation are disclosed in European Patent Publication EP0342926. The ubiquitin promoter is suitable for expressing the nucleotide sequences of the present invention in transgenic plants, particularly monocotyledons. In addition, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231: 150-160 (1991)) can be easily modified to express the nucleotide sequences of the present invention and is particularly suitable for monocotyledonous hosts.
[0073] In certain embodiments, tissue-specific / tissue-preferred promoters can be used for expressing heterologous polynucleotides in plant cells. Tissue-specific or preferred expression patterns include but are not limited to green tissue-specific or preferred, root-specific or preferred, stem-specific or preferred, flower-specific or preferred, or pollen-specific or preferred. Promoters suitable for expression in green tissue include many promoters of genes involved in photosynthesis, many of which are cloned from monocots and dicots. In one embodiment, the promoter that can be used for the present invention is the corn PEPC promoter (Hudspeth and Grula, Plant Molec.Biol.12:579-589 (1989)) from the phosphoenol carboxylase gene. Non-limiting examples of tissue-specific promoters include promoters associated with genes encoding seed storage proteins (e.g., β-conglycinin, cruciferin, napin, and phaseolin), zein or oil body proteins (e.g., oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad 2-1)), and other nucleic acids expressed during embryo development (e.g., Bce4, see, e.g., Kridl et al. (1991) Seed Sci. Res. 1: 209-219; and EP Patent No. 255378). Tissue-specific or tissue-preferred promoters useful for expressing the nucleotide sequences of the present invention in plants, particularly corn, include, but are not limited to, those that direct expression in roots, pith, leaves, or pollen. These promoters are disclosed, for example, in WO 93 / 07278, the disclosure of which regarding promoters is incorporated herein by reference.Other non-limiting examples of tissue-specific or tissue-preferred promoters that can be used in the present invention are the cotton rubisco promoter disclosed in U.S. Patent 6,040,504; the rice sucrose synthase promoter disclosed in U.S. Patent 5,604,121; the root-specific promoter described by de Framond (FEBS 290:103-106 (1991); European Patent EP 0452269 to Ciba-Geigy); the stem-specific promoter described in U.S. Patent 5,625,136 (to Ciba-Geigy), which drives expression of the maize trpA gene; the tuberose yellow leaf curl virus promoter disclosed in WO 01 / 73087; and pollen-specific or -preferred promoters, including but not limited to ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al., Plant Biotechnol. Reports 2001). 9(5):297-306 (2015)), ZmSTK2_USP from maize (Wang et al., Genome 60(6):485-495 (2017)), LAT52 and LAT59 from tomato (Twell et al., Development 109(3):705-713 (1990)), Zm13 (U.S. Patent No. 10,421,972), PLA2-δ promoter from Arabidopsis thaliana (U.S. Patent No. 7,141,424) and / or ZmC5 promoter from maize (International PCT Publication No. WO 1999 / 042587).
[0074] Additional examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, root hair-specific cis-elements (RHEs) (K IMet al., The Plant Cell 18:2958-2970 (2006)), the root-specific promoters RCc3 (Jeong et al., Plant Physiol. 153:185-197 (2010)) and RB7 (U.S. Pat. No. 5,459,252), the lectin promoter (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), the maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology, 37(8): 1108-1115), the maize light harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89: 3654-3658), the maize heat shock protein promoter (O'Dell et al. (1985) EMBO J. 5: 451-458; and Rochester et al. (1986) EMBO J. 5: 451-458), the pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase," pp. 29-39, in: Genetic Engineering of Plants (Hollaender, ed., Plenum Press, NY). 1983; and Poulsen et al. (1986) Mol. Gen. Genet. 205:193-200), Ti plasmid mannopine synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), Ti plasmid nopaline synthase promoter (Langridge et al. (1989), supra), petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7:1257-1263), legume glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev. 3:1639-1646), truncated CaMV 35S promoter (O'Dell et al. (1985) Nature 313:810-812), potato glycoprotein promoter (Wenzler et al. (1989) Plant Mol.Biol. 13:347-354), root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res. 18:7449), zein promoter (Kriz et al. (1987) Mol. Gen. Genet. 207:90-98; Langridge et al. (1983) Cell 34:1015-1022; Reina et al. (1990) Nucleic Acids Res. 18:6425; Reina et al. (1990) Nucleic Acids Res. 18:7449; and Wandelt et al. (1989) Nucleic Acids Res. 17:2354), globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215:431-440), PEPCase promoter (Hudspeth and Grula (1989) Plant Mol. Biol. 12:579-589), R gene complex-associated promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).
[0075] Useful for seed-specific expression is the vicilin promoter (Czako et al. (1992) Mol. Gen. Genet. 235:33-40; and seed-specific promoters disclosed in U.S. Pat. No. 5,625,136). Useful promoters for expression in mature leaves are those that switch at the onset of senescence, such as the SAG promoter from Arabidopsis thaliana (Gan et al. (1995) Science 270:1986-1988).
[0076] In addition, promoters that are functional in chloroplasts can be used. Non-limiting examples of such promoters include the phage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters that can be used in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).
[0077] Additional regulatory elements useful in the present invention include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.
[0078] Introns that can be used in the present invention can be in plants identified and separated from plants, and then inserted into introns in expression cassettes for plant transformation. As will be appreciated by those skilled in the art, introns can include sequences required for self-excision and be incorporated into nucleic acid constructs / expression cassettes in frame. Introns can be used as spacers to separate multiple protein coding sequences in a nucleic acid construct, or introns can be used in a protein coding sequence, for example, to stabilize mRNA. If they are used in a protein coding sequence, they are inserted "in frame" and include an excision site. Introns can also be associated with promoters to improve or change expression. As an example, promoter / intron combinations that can be used in the present invention include but are not limited to promoter / intron combinations of corn Ubi1 promoter and introns.
[0079] Non-limiting examples of introns that can be used in the present invention include introns from the following genes: ADHI gene (e.g., Adh1-S intron 1, 2 and 6), ubiquitin gene (Ubil), RuBisCO small subunit (rbcS) gene, RuBisCO large subunit (rbcL) gene, actin gene (e.g., actin-1 intron), pyruvate dehydrogenase kinase gene (pdk), nitrate reductase gene (nr), repeated carbonic anhydrase gene 1 (Tdca1), psbA gene, atpA gene or any combination thereof.
[0080] As used herein, an “editing system” refers to any site-specific (e.g., sequence-specific) nucleic acid editing system now known or later developed that can introduce modifications (e.g., mutations) into a nucleic acid in a target-specific manner. For example, an editing system (e.g., a site-specific and / or sequence-specific editing system) may include, but is not limited to, a CRISPR-Cas editing system, a meganuclease editing system, a zinc finger nuclease (ZFN) editing system, a transcription activator-like effector nuclease (TALEN) editing system, a base editing system, and / or a lead editing system, each of which may comprise one or more polypeptides and / or one or more polynucleotides that, when present and / or expressed together (e.g., as a system) in a composition and / or cell, can modify a target nucleic acid in a sequence-specific manner (e.g., mutate a target nucleic acid). In some embodiments, an editing system (e.g., a site-specific and / or sequence-specific editing system) may comprise one or more polynucleotides and / or one or more polypeptides, including but not limited to a nucleic acid binding polypeptide (e.g., a DNA binding domain), a nuclease, another polypeptide, and / or a polynucleotide. In some embodiments, a CRISPR-Cas editing system is provided, wherein the fusion protein of the present invention is used to provide the Cas12a protein of the CRISPR-Cas editing system. The editing system of the present invention can modify a target nucleic acid present in a cell or outside the cell (for example, the method of the present invention can be performed in vitro, in vitro and / or in vivo).
[0081] In some embodiments, the editing system comprises one or more sequence-specific nucleic acid binding polypeptides (e.g., DNA binding domains), which can be derived from, for example, polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, transcription activator-like effector nucleases (TALENs), and / or Argonaute proteins. In some embodiments, the editing system comprises one or more cleavage polypeptides (e.g., nucleases), including but not limited to endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, and / or transcription activator-like effector nucleases (TALENs).
[0082] As used herein, "nucleic acid binding polypeptide" refers to a polypeptide or domain that binds and / or is capable of binding to a nucleic acid (e.g., a target nucleic acid). A DNA binding domain or DNA binding polypeptide is an exemplary nucleic acid binding polypeptide and can be a site and / or sequence-specific nucleic acid binding domain. In some embodiments, the nucleic acid binding polypeptide can be a sequence-specific nucleic acid binding polypeptide, such as, but not limited to, sequence-specific binding domains from, for example, polynucleotide-guided endonucleases, CRISPR-Cas effector proteins (e.g., CRISPR-Cas endonucleases), zinc finger nucleases, transcription activator-like effector nucleases (TALENs), and / or Argonaute proteins. In some embodiments, the nucleic acid binding polypeptide comprises a cleavage domain (e.g., a nuclease domain), such as, but not limited to, endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases, zinc finger nucleases, and / or transcription activator-like effector nucleases (TALENs). In some embodiments, the nucleic acid binding polypeptide is associated with and / or is capable of associating (e.g., forming a complex) with one or more nucleic acid molecules (e.g., forming a complex with a guide nucleic acid as described herein), which can guide and / or direct the nucleic acid binding polypeptide to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof), such that the nucleic acid binding polypeptide binds to the nucleotide sequence at the specific target site. In some embodiments, the nucleic acid binding polypeptide is a CRISPR-Cas effector protein as described herein.
[0083] In some embodiments, the editing system comprises a ribonucleoprotein or is a ribonucleoprotein, such as an assembled ribonucleoprotein complex (e.g., a ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and optionally a deaminase). In some embodiments, the ribonucleoproteins of the editing system can be assembled together (e.g., a preassembled ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and optionally a deaminase), for example, when contacted with a target nucleic acid or when introduced into a cell (e.g., a mammalian cell or a plant cell) (e.g., when a component of the ribonucleoprotein is contacted with a target nucleic acid and / or when a component of the ribonucleoprotein is introduced into a cell). In some embodiments, the ribonucleoproteins of the editing system can be assembled into a complex (e.g., a non-covalently bound complex) when a portion of the ribonucleoprotein contacts the target nucleic acid and / or can be assembled after and / or during introduction into a plant cell. In some embodiments, the editing system can be assembled when introduced into a plant cell (e.g., assembled into a non-covalently bound complex). In some embodiments, the ribonucleoprotein may comprise a fusion protein of the present invention, a guide nucleic acid, and optionally a deaminase. In some embodiments, the ribonucleoprotein of the editing system can be contacted with the target nucleic acid and / or can be introduced into a plant cell. In some embodiments, the editing system can be assembled (e.g., assembled into a non-covalently bound complex) when introduced into a plant cell. In some embodiments, the ribonucleoprotein may comprise a protein of the invention (e.g., a protein prepared using the compositions and / or methods of the invention), a guide nucleic acid, and an optional deaminase and / or reverse transcriptase. In some embodiments, the protein of the invention comprises a CRISPR-Cas effector protein, and the protein is used to replace (e.g., replace) a CRISPR-Cas effector protein (e.g., a composition, complex, kit, method, and / or system described herein, such as a CRISPR-Cas effector protein in an editing system) and / or optionally used as a CRISPR-Cas effector protein, a templated editor, and / or a base editor in the compositions, complexes, ribonucleoproteins, kits, methods, systems, and / or editing systems of the invention.
[0084] As used herein, the term "transgene" or "transgenic" refers to at least one nucleic acid sequence that is obtained from the genome of an organism or synthetically produced and then introduced into a host cell (e.g., a plant cell) or an organism or tissue of interest and subsequently integrated into the host's genome by a "stable" transformation or transfection method. In contrast, the term "transient" transformation or transfection or introduction refers to a manner of introducing a molecular tool comprising at least one nucleic acid (DNA, RNA, single-stranded or double-stranded or a mixture thereof) and / or at least one amino acid sequence, optionally containing a suitable chemical or biological agent, to achieve transfer into at least one compartment of interest of a cell, including but not limited to the cytoplasm, organelles (including nuclei, mitochondria, vacuoles, chloroplasts) or membranes, thereby resulting in transcription and / or translation and / or association and / or activity of the at least one molecule introduced, without achieving stable integration or incorporation into the genome, so that the corresponding at least one molecule introduced into the cellular genome is not inherited. The term "transgene-free" refers to a state in which the transgene is not present or found in the genome of the host cell or tissue or organism of interest.
[0085] In some embodiments, the polynucleotides and / or nucleic acid constructs of the present invention can be "expression cassettes", or can be contained within expression cassettes. As used herein, "expression cassettes" refer to recombinant nucleic acid molecules comprising, for example, nucleic acid constructs of the present invention (e.g., polynucleotides encoding fusion proteins of the present invention, polynucleotides encoding cytosine deaminase, polynucleotides encoding adenine deaminase, polynucleotides encoding deaminase fusion proteins, polynucleotides encoding peptide tags, polynucleotides encoding affinity polypeptides, polynucleotides encoding glycosylases, and / or polynucleotides comprising guide nucleic acids), wherein the nucleic acid constructs are operably associated with at least a control sequence (e.g., a promoter). Thus, some embodiments of the present invention provide expression cassettes designed to express, for example, nucleic acid constructs of the present invention. When the expression cassette comprises more than one polynucleotide, the polynucleotides can be operably linked to a single promoter that drives expression of all polynucleotides, or the polynucleotides can be operably linked to one or more separate promoters (e.g., three polynucleotides can be driven by one, two, or three promoters in any combination). Thus, for example, the polynucleotide encoding the fusion protein, the polynucleotide encoding the deaminase (e.g., adenine deaminase), and the polynucleotide comprising the guide nucleic acid contained in the expression cassette can each be operably associated with a single promoter, or one or more of the polynucleotides can be operably associated with separate promoters (e.g., two or three promoters) in any combination, which promoters can be the same or different from each other.
[0086] In some embodiments, expression cassettes comprising a polynucleotide / nucleic acid construct of the invention can be optimized for expression in an organism (eg, an animal, a plant, a bacterium, etc.).
[0087] The expression cassette comprising the nucleic acid construct of the present invention may be chimeric, meaning that at least one of its components is heterologous with respect to at least one of its other components (e.g., a promoter from a host organism is operably linked to a polynucleotide of interest to be expressed in the host organism, wherein the polynucleotide of interest is from an organism different from the host or is not normally associated with the promoter). The expression cassette may also be naturally occurring but has been obtained in a recombinant form useful for heterologous expression.
[0088] The expression cassette may optionally include a transcriptional and / or translational termination region (i.e., a termination region) and / or an enhancer region that is functional in the selected host cell. A variety of transcriptional terminators and enhancers are known in the art and can be used in the expression cassette. The transcriptional terminator is responsible for terminating transcription and correcting mRNA polyadenylation. The termination region and / or enhancer region may be native to the transcription initiation region, may be native to the gene encoding the CRISPR-Cas effector protein or the gene encoding the deaminase, may be native to the host cell, or may be native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the CRISPR-Cas effector protein or the gene encoding the deaminase, the host cell, or any combination thereof).
[0089] The expression cassette of the present invention may also include a polynucleotide encoding a selectable marker that can be used to select transformed host cells. As used herein, a "selectable marker" refers to a polynucleotide sequence that, when expressed, confers a unique phenotype to the host cell expressing the marker, thereby allowing such transformed cells to be distinguished from cells without the marker. Such a polynucleotide sequence can encode a selectable or screenable marker, depending on whether the marker confers a trait that can be selected by chemical means, such as by using a selection agent (e.g., antibiotics, etc.), or whether the marker is simply a trait that can be identified by observation or testing, such as by screening (e.g., fluorescence). Many examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.
[0090] Expression cassettes, nucleic acid molecules / constructs and polynucleotide sequences described herein can be used in combination with vectors.Term " vector " refers to the composition for transferring, delivering or introducing one or more nucleic acids into a cell.Carrier can comprise nucleic acid construct, and described nucleic acid construct comprises one or more nucleotide sequences to be transferred, delivered or introduced into a cell.Carrier for transforming host organisms is well known in the art.Non-limiting examples of general category carriers include viral vectors (for example, adeno-associated virus (AAV) vectors), plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, Fosmid (fosmid) vectors, phage, artificial chromosomes, mini-rings or double-stranded or single-stranded linear or circular form of Agrobacterium (Agrobacterium) binary vectors, and these carriers can be or not self-propagating or removable.In certain embodiments, viral vectors can include but are not limited to retrovirus, slow virus, adenovirus, adeno-associated virus or herpes simplex virus vectors.Carrier as defined herein can be by being integrated into the cell genome or being present in extrachromosomal (for example, autonomously replicating plasmid with replication origin) and transforming protokaryotic or eukaryotic hosts. In addition, shuttle vectors are also included, and shuttle vectors refer to DNA vectors that can be naturally or intentionally replicated in two different host organisms, and host organisms can be selected from actinomycetes and related species, bacteria and eukaryotes (for example, higher plants, mammals, yeast or fungal cells). In certain embodiments, the nucleic acid in the vector is subject to the control of suitable promoters or other regulatory elements, and is operably connected to suitable promoters or other regulatory elements, so that it is transcribed in the host cell. The vector can be a bifunctional expression vector that works in a variety of hosts. In the case of genomic DNA, this can include its own promoter and / or other regulatory elements, and in the case of cDNA, this can be subject to the control of suitable promoters and / or other regulatory elements, so that it is expressed in the host cell. Therefore, nucleic acid constructs of the present invention and / or expression cassettes comprising it can be included in vectors described herein and known in the art.
[0091] As used herein, "contact," "contacting," "contacted," and grammatical variations thereof, refer to bringing together the components of a desired reaction under conditions suitable for the desired reaction (e.g., transformation, transcriptional control, genome editing, nicking, and / or cleavage). Thus, for example, a target nucleic acid can be contacted with a nucleic acid construct of the invention encoding, for example, a nucleic acid-binding polypeptide (e.g., a DNA-binding polypeptide, e.g., a sequence-specific DNA-binding protein (e.g., a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein (e.g., a CRISPR-Cas endonuclease), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN) and / or an Argonaute protein)), a guide nucleic acid, and optionally a cytosine deaminase and / or an adenine deaminase, under conditions such that the nucleic acid-binding polypeptide (e.g., a CRISPR-Cas effector protein) is expressed, and the nucleic acid-binding polypeptide forms a complex with the guide nucleic acid, the complex hybridizes to the target nucleic acid, and optionally the cytosine deaminase and / or adenine deaminase is recruited to the nucleic acid-binding polypeptide (and thereby to the target nucleic acid), or the cytosine deaminase and / or adenine deaminase is fused to the nucleic acid-binding polypeptide, thereby modifying the target nucleic acid. In some embodiments, the cytosine deaminase and / or adenine deaminase and the nucleic acid binding polypeptide are localized to the target nucleic acid, optionally through covalent and / or non-covalent interactions.
[0092] In some embodiments, the target nucleic acid can be contacted with a nucleic acid construct of the present invention encoding a fusion protein of the present invention, a guide nucleic acid, and optionally a cytosine deaminase and / or adenine deaminase under conditions that produce the fusion protein, or the target nucleic acid can be contacted with a fusion protein of the present invention, a guide nucleic acid, and optionally a cytosine deaminase and / or adenine deaminase. The fusion protein can form a complex with the guide nucleic acid, and the complex can hybridize with the target nucleic acid, and optionally a cytosine deaminase and / or adenine deaminase is recruited to the fusion protein (and therefore recruited to the target nucleic acid), or a cytosine deaminase and / or adenine deaminase is fused to the fusion protein to modify the target nucleic acid. The cytosine deaminase and / or adenine deaminase and the fusion protein can optionally be positioned at the target nucleic acid by covalent and / or non-covalent interactions.
[0093] As used herein, " modification (modifying or modification) " about target nucleic acid includes editing (for example, mutation), covalent modification, exchange / replacement nucleic acid / nucleotide base, deletion, cutting and / or nicking target nucleic acid to thereby provide modified nucleic acid and / or change the transcription control of target nucleic acid to thereby provide modified nucleic acid. In certain embodiments, modification may include insertion and / or deletion and / or any type of single base change (SNP). In certain embodiments, modification includes SNP. In certain embodiments, modification includes exchanging and / or replacing one or more (for example, 1, 2, 3, 4, 5 or more) nucleotides. In some embodiments, the insertion or deletion can be from about 1 base to about 30,000 bases in length or longer (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690 90、700、710、720、730、740、750、760、770、780、790、800、810、820、830、840、850、860、870、880、890、900、910、920、930、940、950、960、970、980、990、1000、1100、1200、1300、1400、1500、1600、1700、1800、1900、2000、2500、3000、3500、4000、4500、5000、5500、6000、6500、7000、7500、8000、8500、9000、9500、10,000、10,500、11,000、11,500、12,000、12,500、13,000、13,500、14,000、14,500、15,000、15,500、16,000、16,500、17,000、17,500、18,000、18, 00, 28,000, 28,500, 29,000, 29,500, 30,000 bases or longer, or any value or range therein). Thus, in some embodiments, the length of the insertion or deletion can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 9, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110 , 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 5 10, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900,910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 bases, or any range or value therein; about 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180 , 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bases to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, , 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases or more, or any value or range thereof; 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases to about 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000,70, 750, 760, 770, 780, 790, or 800 bases in length. 500, 3000, 3500, 4000, 4500 or 5000 bases or longer, or any value or range therein. In some embodiments, the length of the insertion or deletion can be about 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10,000 bases to about 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, or 16,000 bases. 0, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, 20,500, 21,000, 21,500, 22,000, 22,500, 23,000, 23,500, 24,000, 24,500, 25,000, 25,500, 26,000, 26,500, 27,000, 27,500, 28,000, 28,500, 29,000, 29,500 or 30,000 bases or longer, or any value or range therein.
[0094] As used herein, "recruit," "recruiting," or "recruitment" refers to the use of protein-protein interactions, nucleic acid-protein interactions (e.g., RNA-protein interactions), and / or chemical interactions to attract one or more polypeptides or polynucleotides to another polypeptide or polynucleotide (e.g., a specific location in a genome). Protein-protein interactions may include, but are not limited to, peptide tags (epitopes, multimeric epitopes) and corresponding affinity polypeptides, RNA recruitment motifs and corresponding affinity polypeptides, and / or chemical interactions. Exemplary chemical interactions that can be used with polypeptides and polynucleotides for recruitment purposes may include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin interaction; SNAP tag (Hussain et al., Curr Pharm Des. 19(30):5437-42 (2013)); Halo tag (Los et al., ACS Chem Biol. 3(6):373-82 (2008)); CLIP tag (Gautier et al., Chemistry & Biology 15:128-136 (2008)); DmrA-DmrC heterodimer induced by compounds (Tak et al., Nat Methods 14(12):1163-1166 (2017)); bifunctional ligand approach (fusing two protein binding chemistries together) (Voβ et al., Curr Opin Chemical Biology 28:194-201 (2015)) (e.g., dihydrofolate reductase (DHFR) (Kopyteck et al., Cell Cehm Biol 7(5):313-321 (2000)).
[0095] In the context of a polynucleotide or editing system of interest, "introducing," "introduce," "introduced" (and grammatical variations thereof) refers to presenting a nucleotide sequence (e.g., a polynucleotide, a nucleic acid construct, and / or a guide nucleic acid) and / or an editing system (e.g., a polynucleotide, a polypeptide, and / or a ribonucleoprotein) of interest to a host organism or a cell of the organism (e.g., a host cell; e.g., a plant cell) in a manner that allows the nucleotide sequence and / or editing system to gain access to the interior of the cell. Thus, for example, a nucleic acid construct of the invention encoding a fusion protein, a guide nucleic acid, and / or a cytosine deaminase and / or adenine deaminase of the invention can be introduced into a cell of an organism, thereby transforming the cell with the fusion protein, the guide nucleic acid, and / or the cytosine deaminase and / or the adenine deaminase. In some embodiments, the fusion protein and / or the guide nucleic acid of the invention can be introduced into a cell of an organism, optionally wherein the fusion protein and the guide nucleic acid can be contained in a complex (e.g., a ribonucleoprotein). In some embodiments, the organism is a eukaryotic organism (e.g., a mammal, e.g., a human).
[0096] As used herein, the term "transformation" refers to the introduction of nucleic acids, polypeptides and / or ribonucleoproteins (e.g., heterologous nucleic acids, polypeptides and / or ribonucleoproteins) into cells. The transformation of cells can be stable or transient. Thus, in some embodiments, host cells or host organisms can be stably transformed with the polynucleotides / nucleic acid molecules of the present invention. In some embodiments, host cells or host organisms can be transiently transformed with the nucleic acid constructs, polypeptides and / or ribonucleoproteins of the present invention.
[0097] "Transient transformation" in the context of a polynucleotide, polypeptide and / or ribonucleoprotein means that the polynucleotide, polypeptide and / or ribonucleoprotein is introduced into a cell and does not integrate into the genome of the cell.
[0098] "Stably introduced" or "stably introduced" in the context of a polynucleotide being introduced into a cell means that the introduced polynucleotide is stably incorporated into the genome of the cell, and thus the cell is stably transformed with the polynucleotide.
[0099] As used herein, "stable transformation" or "stably transformed" refers to a nucleic acid molecule that is introduced into a cell and integrated into the cell's genome. Thus, the integrated nucleic acid molecule is capable of being inherited by its progeny, more specifically, by successive generations. As used herein, "genome" includes both nuclear and plastid genomes, and thus includes integration of a nucleic acid into, for example, a chloroplast or mitochondrial genome. As used herein, stable transformation may also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or plasmid.
[0100] Transient transformation can be detected by, for example, enzyme-linked immunosorbent assay (ELISA) or Western blotting, which can detect the presence of peptides or polypeptides encoded by one or more transgenics introduced into an organism. Stable transformation of a cell can be detected by, for example, Southern blot hybridization assays of the genomic DNA of the cell and nucleic acid sequences, which specifically hybridize to the nucleotide sequences of the transgenics introduced into an organism (e.g., mammals, plants, etc.). Stable transformation of a cell can be detected by, for example, Northern blot hybridization assays of the RNA of the cell and nucleic acid sequences, which specifically hybridize to the nucleotide sequences of the transgenics introduced into the host organism. Stable transformation of a cell can also be detected by, for example, polymerase chain reaction (PCR) or other amplification reactions well known in the art, which employ specific primer sequences that hybridize to the target sequence of the transgenic, resulting in the amplification of the transgenic sequence, thereby detecting the transgenic sequence according to standard methods. Conversion can also be detected by direct sequencing and / or hybridization protocols well known in the art.
[0101] Thus, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs and / or expression cassettes of the present invention can be transiently expressed and / or they can be stably incorporated into the genome of a host organism. Thus, in some embodiments, the nucleic acid constructs of the present invention can be transiently introduced into a cell together with a guide nucleic acid, and thus, no DNA is maintained in the cell.
[0102] Nucleic acid construct of the present invention, polypeptide and / or ribonucleoprotein can be introduced into cell by any method well known by persons skilled in the art.In certain embodiments, conversion method includes but is not limited to: via bacterial-mediated nucleic acid delivery (for example, via agrobacterium), virus-mediated nucleic acid delivery, silicon carbide and / or nucleic acid whisker-mediated nucleic acid delivery, the conversion of liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated conversion, cyclodextrin-mediated conversion, electroporation, nanoparticle-mediated conversion, supersound process, infiltration, PEG-mediated nucleic acid uptake, and causing nucleic acid to be introduced into any other electricity, chemistry, physics (mechanics) and / or biological mechanism in cell (for example, plant cell or animal cell), including any combination thereof.In some embodiments of the invention, the conversion of cell includes nuclear transformation.In certain embodiments, the conversion of cell includes plastid transformation (for example, chloroplast transformation).In certain embodiments, recombinant nucleic acid construct of the present invention can be introduced into cell via conventional breeding technology.
[0103] Procedures for transforming eukaryotic and prokaryotic organisms are well known and routine in the art and are described throughout the literature (see, for example, Jiang et al., 2013. Nat. Biotechnol. 31: 233-239; Ran et al., Nature Protocols 8: 2281-2308 (2013)). General guidance on various plant transformation methods known in the art includes Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, Eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell. Mol. Biol. Lett. 7: 849-858 (2002)).
[0104] In some embodiments, the present invention provides a method for producing a nucleotide sequence, polypeptide and / or ribonucleoprotein that is incorporated into a host organism or its cell in a variety of ways well known in the art. Method of the present invention does not rely on the ad hoc approach for introducing one or more nucleotide sequences, polypeptide and / or ribonucleoprotein into an organism, but only relies on the inside of at least one cell of the organism to enter. When introducing more than one nucleotide sequence, polypeptide and / or ribonucleoprotein, they can be assembled into a part for a single nucleic acid construct, or be assembled into an independent nucleic acid construct, and can be positioned on an identical or different nucleic acid construct. Therefore, nucleotide sequence, polypeptide and / or ribonucleoprotein can be introduced into interested cell in a single transformation event and / or in an independent transformation event, or alternatively, in the relevant case, nucleotide sequence can for example be incorporated into a plant as a part for a breeding scheme. In certain embodiments, cell is a eukaryotic cell (for example, a plant cell or a mammal such as a human cell).
[0105] In some embodiments, the nucleic acid construct of the present invention (e.g., a polynucleotide encoding a fusion protein of the present invention, a polynucleotide encoding a deaminase, and / or a guide nucleic acid, and / or an expression cassette and / or a vector comprising the same) can be operably linked to at least one regulatory sequence, optionally wherein the at least one regulatory sequence can be codon-optimized for expression in a plant. In some embodiments, the at least one regulatory sequence can be, for example, a promoter, an operator, a terminator, or an enhancer. In some embodiments, the at least one regulatory sequence can be a promoter. In some embodiments, the regulatory sequence can be an intron. In some embodiments, the at least one regulatory sequence can be, for example, a promoter operably associated with an intron or a promoter region comprising an intron. In some embodiments, the at least one regulatory sequence can be, for example, a ubiquitin promoter and its associated introns (e.g., Medicago truncatula and / or corn and its associated introns). In some embodiments, the at least one regulatory sequence can be a terminator nucleotide sequence and / or an enhancer nucleotide sequence.
[0106] In some embodiments, the nucleic acid constructs of the invention can be operably associated with a promoter region, wherein the promoter region comprises an intron, optionally wherein the promoter region can be a ubiquitin promoter and intron (e.g., an alfalfa or maize ubiquitin promoter and intron, e.g., SEQ ID NO: 36 or SEQ ID NO: 37). In some embodiments, the nucleic acid constructs of the invention operably associated with a promoter region comprising an intron can be codon-optimized for expression in plants.
[0107] In some embodiments, the nucleic acid construct of the present invention can encode one or more (e.g., 1, 2, 3, 4 or more) polypeptides of interest. The one or more polypeptides of interest can be codon-optimized for expression in eukaryotic organisms (e.g., humans or plants). In some embodiments, the fusion protein can comprise one or more (e.g., 1, 2, 3, 4 or more) polypeptides of interest.
[0108] The polypeptides of interest that can be used in the present invention may include, but are not limited to, polypeptides or protein domains having the following: deaminase activity, nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil-DNA glycosylase inhibitor (UGI)), reverse transcriptase, peptide tag (e.g., GCN4 peptide tag), demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity In some embodiments, the polypeptide of interest is a Fok1 nuclease or uracil-DNA glycosylase inhibitor. When encoded in a nucleic acid (polynucleotide, expression cassette and / or vector), the encoded polypeptide or protein domain can be codon-optimized for expression in an organism. In some embodiments, the polypeptide of interest can be linked to a fusion protein of the present invention or to a CRISPR-Cas effector protein domain to provide a CRISPR-Cas fusion protein. In some embodiments, a CRISPR-Cas fusion protein comprising a CRISPR-Cas effector protein domain linked to a recruitment motif (e.g., a peptide tag) may also be linked to a polypeptide of interest (e.g., a CRISPR-Cas effector protein domain may be linked, for example, to a recruitment motif (e.g., a peptide tag or affinity polypeptide) and, for example, a polypeptide of interest simultaneously).
[0109] In some embodiments, the editing system of the present invention includes CRISPR-Cas effector proteins. As used herein, "CRISPR-Cas effector proteins" are proteins or polypeptides that cut, cut or nick nucleic acids; bind nucleic acids (e.g., target nucleic acids and / or guide nucleic acids); and / or identify, recognize or bind guide nucleic acids as defined herein. In some embodiments, CRISPR-Cas effector proteins can be enzymes (e.g., nucleases, endonucleases, nickases, etc.) and / or can act as enzymes. In some embodiments, CRISPR-Cas effector proteins refer to CRISPR-Cas nucleases. In some embodiments, CRISPR-Cas effector proteins include nuclease activity and / or nickase activity, include nuclease domains whose nuclease activity and / or nickase activity has been reduced or eliminated, include single-stranded DNA cleavage activity (ss DNAse activity) or it has ss DNAse activity that has been reduced or eliminated, and / or include self-processing RNAse activity or it has self-processing RNAse activity that has been reduced or eliminated. CRISPR-Cas effector proteins can bind to target nucleic acids. The CRISPR-Cas effector protein can be a type I, type II, type III, type IV, type V or type VI CRISPR-Cas effector protein. In some embodiments, the CRISPR-Cas effector protein can be from a type I CRISPR-Cas system, a type II CRISPR-Cas system, a type III CRISPR-Cas system, a type IV CRISPR-Cas system, a type V CRISPR-Cas system or a type VI CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein of the present invention can be from a type II CRISPR-Cas system or a type V CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein can be a type II CRISPR-Cas effector protein, for example, a Cas9 effector protein. In some embodiments, the CRISPR-Cas effector protein can be a type V CRISPR-Cas effector protein, for example, a Cas12 effector protein. In some embodiments, the CRISPR-Cas effector protein can be Cas12a, and optionally can have the amino acid sequence of any one of SEQ ID NOs: 38-60 or 192-195 and / or the nucleotide sequence of any one of SEQ ID NOs: 61-63. In some embodiments, the CRISPR-Cas effector protein may be an active Cas12a, and optionally may have the amino acid sequence of SEQ ID NO: 46 or 55. In some embodiments, the CRISPR-Cas effector protein may be an inactive (i.e., dead) Cas12a, and optionally may have the amino acid sequence of SEQ ID NO: 38.
[0110] Exemplary CRISPR-Cas effector proteins include, but are not limited to, Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3 , Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG) and / or Csf5 nuclease, optionally wherein the CRISPR-Cas effector protein can be a Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b and / or Cas14c effector protein.
[0111] In some embodiments, the CRISPR-Cas effector proteins useful in the present invention may comprise mutations in their nuclease active sites and / or nuclease domains (e.g., RuvC, HNH, e.g., RuvC sites of Cas12a nuclease domains; e.g., RuvC sites and / or HNH sites of Cas9 nuclease domains). CRISPR-Cas effector proteins that have mutations in their nuclease active sites and / or nuclease domains and therefore no longer contain nuclease activity are often referred to as “inactive” or “dead,” e.g., dCas12a. In some embodiments, CRISPR-Cas effector proteins that have mutations in their nuclease active sites and / or nuclease domains may have impaired activity or reduced activity (e.g., nickase activity) compared to the same CRISPR-Cas effector proteins without mutations.
[0112] The V-type CRISPR-Cas effector protein that can be used in embodiments of the present invention can be Cas12a. The CRISPR-Cas effector protein can be a V-type clustered regularly interspaced short palindromic repeats (CRISPR)-Cas nuclease. Cas12a is different from the more well-known type II CRISPR Cas9 nuclease in several aspects. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) located at 3' of its guide RNA (gRNA, sgRNA, crRNA, crDNA, CRISPR array) binding site (protospacer, target nucleic acid, target DNA), while Cas12a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) located at 5' of the target nucleic acid. In fact, the orientation of Cas9 and Cas12a in binding to their guide RNAs is almost opposite relative to their N and C termini. In addition, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA), rather than the dual guide RNA (sgRNA (e.g., crRNA and tracrRNA)) found in the natural Cas9 system, and Cas12a processes its own gRNA. In addition, the Cas12a nuclease activity produces staggered DNA double-strand breaks, rather than the blunt ends produced by the Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cut two DNA chains, while Cas9 utilizes the HNH domain and the RuvC domain to cut.
[0113] The CRISPR Cas12a effector protein that can be used in the present invention can be any known or later identified Cas12a (formerly known as Cpf1) (see, for example, U.S. Patent No. 9,790,490, the disclosure of which about the Cpf1 (Cas12a) sequence is incorporated by reference). The term "Cas12a" refers to an RNA-guided protein that can have nuclease activity, the protein comprising a guide nucleic acid binding domain and an active, inactive or partially active DNA cleavage domain, whereby the RNA-guided nuclease activity of Cas12a can be active, inactive or partially active, respectively. In some embodiments, the Cas12a that can be used in the present invention may comprise a mutation in a nuclease active site (e.g., a RuvC site of a Cas12a domain). Cas12a that has a mutation in its nuclease domain and / or nuclease active site and therefore no longer comprises nuclease activity is generally referred to as dead Cas12a (e.g., dCas12a). In some embodiments, Cas12a having a mutation in its nuclease domain and / or nuclease active site may have impaired activity, for example, may have reduced nickase activity.
[0114] In some embodiments, the CRISPR-Cas effector protein (e.g., Cas12a) can be optimized for expression in an organism, such as an animal (e.g., a mammal, such as a human), a plant, a fungus, an archaea, or a bacterium. In some embodiments, the CRISPR-Cas effector protein (e.g., Cas12a) can be optimized for expression in a plant.
[0115] Any deaminase domain / polypeptide that can be used for base editing can be used in the present invention. As used herein, "cytosine deaminase" and "cytidine deaminase" refer to a polypeptide or its domain that catalyzes or can catalyze the deamination of cytosine, because the polypeptide or domain catalyzes or can catalyze the removal of an amine group from a cytosine base. Therefore, cytosine deaminase can cause cytosine to be converted into thymidine (via a uracil intermediate), thereby causing C to be converted to T or G to A in the complementary chain in the genome. Therefore, in some embodiments, the cytosine deaminase encoded by the polynucleotide of the present invention produces C→T conversion in the sense (e.g., "+"; template) chain of the target nucleic acid or produces G→A conversion in the antisense (e.g., "-", complementary) chain of the target nucleic acid. In some embodiments, the cytosine deaminase encoded by the polynucleotide of the present invention produces C to T, G or A conversion in the complementary chain in the genome.
[0116] The cytosine deaminase that can be used in the present invention can be any known or later identified cytosine deaminase from any organism (see, for example, U.S. Patent No. 10,167,457 and Thuronyi et al., Nat. Biotechnol. 37: 1070-1079 (2019), each of which is incorporated herein by reference for the cytosine deaminase disclosed therein). The cytosine deaminase can catalyze the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. Therefore, in some embodiments, the deaminase or deaminase domain that can be used in the present invention can be a cytidine deaminase domain that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the cytosine deaminase can be a variant of a naturally occurring cytosine deaminase, including but not limited to primates (e.g., humans, monkeys, chimpanzees, gorillas), dogs, cows, rats, or mice. Thus, in some embodiments, a cytosine deaminase useful in the present invention can be about 70% to about 100% identical to a wild-type cytosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a naturally occurring cytosine deaminase, and any range or value therein).
[0117] In some embodiments, the cytosine deaminase useful in the present invention may be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase may be an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, an APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570), and evolved versions thereof. Evolved deaminases are disclosed in, for example, U.S. Patent No. 10,113,163, Gaudelli et al. (Nature 551(7681):464-471 (2017)), and Thuronyi et al. (Nature Biotechnology 37:1070-1079 (2019)), each of which is incorporated herein by reference for its disclosure regarding deaminases and evolved deaminases. In some embodiments, the cytosine deaminase may be an APOBEC1 deaminase having an amino acid sequence of SEQ ID NO: 64. In some embodiments, the cytosine deaminase may be an APOBEC3A deaminase having an amino acid sequence of SEQ ID NO: 65. In some embodiments, the cytosine deaminase may be a CDA1 deaminase, optionally a CDA1 having an amino acid sequence of SEQ ID NO: 66. In some embodiments, the cytosine deaminase may be a FERNY deaminase, optionally a FERNY having an amino acid sequence of SEQ ID NO: 67. In some embodiments, the cytosine deaminase may be a rAPOBEC1 deaminase, optionally a rAPOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 68. In some embodiments, the cytosine deaminase may be a hAID deaminase, optionally a hAID having the amino acid sequence of SEQ ID NO: 69 or SEQ ID NO: 70.In some embodiments, the cytosine deaminases useful in the present invention can be about 70% to about 100% identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identical) to the amino acid sequence of a naturally occurring cytosine deaminase (e.g., an "evolved deaminase") (see, e.g., SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73). In some embodiments, a cytosine deaminase useful in the present invention can be about 70% to about 99.5% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical) to the amino acid sequence of any one of SEQ ID NOs: 64-73 (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical). In some embodiments, a polynucleotide encoding a cytosine deaminase can be codon-optimized for expression in plants, and the codon-optimized polypeptide can be about 70% to 99.5% identical to a reference polynucleotide.
[0118] As used herein, "adenine deaminase" and "adenosine deaminase" refer to polypeptides or domains thereof that catalyze or are capable of catalyzing the hydrolytic deamination of adenine or adenosine (e.g., removing an amine group from adenine). In some embodiments, adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminase can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention can produce an A→G conversion in the sense (e.g., "+"; template) strand of the target nucleic acid or a T→C conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid. Adenine deaminase useful in the present invention can be any known or later identified adenine deaminase from any organism (see, e.g., U.S. Patent No. 10,113,163, which incorporates herein by reference the adenine deaminases disclosed therein).
[0119] In some embodiments, the adenosine deaminase can be a variant of a naturally occurring adenine deaminase. Thus, in some embodiments, the adenosine deaminase can be about 70% to 100% identical to a wild-type adenine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identical to a naturally occurring adenine deaminase, and any ranges or values therein). In some embodiments, the deaminase or deaminases do not exist in nature and can be referred to as engineered, mutated or evolved adenosine deaminase. Thus, for example, an engineered, mutant, or evolved adenine deaminase polypeptide or adenine deaminase domain can be about 70% to 99.9% identical to a naturally occurring adenine deaminase polypeptide / domain (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 9 ... In some embodiments, the adenine deaminase polypeptide / domain may be codon-optimized for expression in plants.
[0120] In some embodiments, the adenine deaminase domain can be a wild-type tRNA-specific adenosine deaminase domain, e.g., a tRNA-specific adenosine deaminase (TadA) and / or a mutant / evolved adenosine deaminase domain, e.g., a mutant / evolved tRNA-specific adenosine deaminase domain (TadA*). In some embodiments, the TadA domain can be from Escherichia coli. In some embodiments, TadA can be modified, e.g., truncated, such that one or more N-terminal and / or C-terminal amino acids are deleted relative to full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues may be deleted relative to full-length TadA). In some embodiments, the TadA polypeptide or TadA domain does not comprise an N-terminal methionine. In some embodiments, wild-type E. coli TadA comprises the amino acid sequence of SEQ ID NO: 74. In some embodiments, the mutant / evolved E. coli TadA* comprises the amino acid sequence of any one of SEQ ID NOs: 75-78. In some embodiments, the polynucleotide encoding TadA / TadA* can be codon-optimized for expression in plants. In some embodiments, the adenine deaminase can comprise all or a portion of the amino acid sequence of any one of SEQ ID NOs: 79-84. In some embodiments, the adenine deaminase can comprise the amino acid sequence of SEQ ID NOs: All or part of the amino acid sequence of any one of NOs: 74-84.
[0121] In some embodiments, the nucleic acid constructs of the present invention may further encode a glycosylase inhibitor (e.g., a uracil glycosylase inhibitor (UGI), such as a uracil-DNA glycosylase inhibitor). In some embodiments, the present invention provides a fusion protein comprising a UGI and / or one or more polynucleotides encoding the same, optionally wherein the one or more polynucleotides may be codon-optimized for expression in plants.
[0122] "Uracil glycosylase inhibitors" useful in the present invention may be any protein or polypeptide capable of inhibiting uracil-DNA glycosylase base excision repair enzymes. In some embodiments, the UGI domain comprises wild-type UGI or a fragment thereof. In some embodiments, the UGI domain useful in the present invention may be about 70% to about 100% identical to the amino acid sequence of a naturally occurring UGI domain (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identical and any range or value therein). In some embodiments, the UGI domain can comprise the amino acid sequence of SEQ ID NO:85, or a polypeptide having about 70% to about 99.5% identity to the amino acid sequence of SEQ ID NO:85 (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of SEQ ID NO:85). For example, in some embodiments, the UGI domain may comprise a fragment of the amino acid sequence of SEQ ID NO: 85 that is 100% identical to a portion of contiguous nucleotides (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides; e.g., about 10, 15, 20, 25, 30, 35, 40, 45 to about 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides) of the amino acid sequence of SEQ ID NO: 85. In some embodiments, the UGI domain can be a variant of a known UGI (e.g., SEQ ID NO: 85) having about 70% to about 99.5% identity (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% identity, and any range or value therein) to the known UGI. In some embodiments, the polynucleotide encoding the UGI can be codon-optimized for expression in a plant (e.g., a plant), and the codon-optimized polypeptide can be about 70% to about 99.5% identical to the reference polynucleotide.
[0123] The fusion protein of the present invention can be used in combination with a guide nucleic acid (e.g., a guide RNA (gRNA), a CRISPR array, a CRISPR RNA, a crRNA), which is designed to function together with the fusion protein to modify the target nucleic acid. The guide nucleic acid that can be used in the present invention may comprise at least one spacer sequence and at least one repeat sequence. The guide nucleic acid can form a complex with the fusion protein (e.g., with the nuclease domain of the fusion protein), and the spacer sequence can hybridize with the target nucleic acid, thereby guiding the complex to the target nucleic acid, wherein the target nucleic acid can be modified (e.g., cut or edited) and / or regulated (e.g., regulated transcription) by a deaminase (e.g., cytosine deaminase and / or adenine deaminase) or a reverse transcriptase that is optionally present in the complex and / or recruited to the complex.
[0124] As used herein, "guide nucleic acid", "guide RNA", "gRNA", "CRISPR RNA / DNA", "crRNA" or "crDNA" refers to a nucleic acid comprising at least one spacer sequence and at least one repetitive sequence (e.g., a repetitive sequence or a fragment or portion thereof of a V-type Cas12a CRISPR-Cas system), wherein the spacer sequence is complementary to (and hybridizes with) a target nucleic acid (e.g., a target DNA and / or a protospacer), wherein the repetitive sequence can be connected to the 5' end and / or the 3' end of the spacer sequence. In some embodiments, the guide nucleic acid comprises DNA. In some embodiments, the guide nucleic acid comprises RNA (e.g., is a guide RNA). The design of the gRNA of the present invention can be based on a type I, type II, type III, type IV, type V or type VI CRISPR-Cas system. In some embodiments, the Cas12a gRNA may comprise a repetitive sequence (full length or a portion thereof ("handle"); e.g., a pseudoknot-like structure) and a spacer sequence from 5' to 3'.
[0125] In some embodiments, the guide nucleic acid can comprise more than one repeat-spacer sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). The guide nucleic acids of the present invention are synthetic, artificial, and do not exist in nature. The gRNA can be very long and can be used as an aptamer (e.g., in the MS2 recruitment strategy) or other RNA structures with hanging spacers.
[0126] As used herein, "repetitive sequence" refers to any repetitive sequence of, for example, a wild-type CRISPR Cas locus (e.g., Cas9 locus, Cas12a locus, C2c1 locus, etc.), or a repetitive sequence of a synthetic repetitive sequence (e.g., synthetic crRNA) that plays a role together with the CRISPR-Cas effector protein encoded by the nucleic acid construct of the present invention. The repetitive sequence that can be used for the present invention can be any known or later identified repetitive sequence (e.g., type I, type II, type III, type IV, type V, or type VI) of the CRISPR-Cas locus, or it can be a synthetic repetitive sequence designed to play a role in an I, II, III, IV, V, or VI type CRISPR-Cas system. The repetitive sequence may include a hairpin structure and / or a stem-loop structure. In certain embodiments, the repetitive sequence may form a pseudoknot structure (i.e., "handle") at its 5' end. Thus, in some embodiments, the repetitive sequence may be identical or substantially identical to a repetitive sequence from a wild-type type I CRISPR-Cas locus, a type II CRISPR-Cas locus, a type III CRISPR-Cas locus, a type IV CRISPR-Cas locus, a type V CRISPR-Cas locus, and / or a type VI CRISPR-Cas locus. The repetitive sequence from the wild-type CRISPR-Cas locus can be determined by an established algorithm, for example, using the CRISPRfinder provided by CRISPRdb (see, Grissa et al., Nucleic Acids Res. 35 (web server album): W52-7). In some embodiments, the repetitive sequence or a portion thereof is linked to the 5' end of the spacer sequence at its 3' end, thereby forming a repeat-spacer sequence (e.g., a guide nucleic acid, a guide RNA / DNA, crRNA, crDNA).
[0127] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides, depending on the particular repeat sequence and whether the guide nucleic acid comprising the repeat sequence is processed or unprocessed (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 to 100 or more nucleotides, or any range or value therein; e.g., about). In some embodiments, the repetitive sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100, or more nucleotides.
[0128] The repeat sequence linked to the 5' end of the spacer sequence can comprise a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive nucleotides of a wild-type repeat sequence). In some embodiments, the portion of the repeat sequence linked to the 5' end of the spacer sequence can be about five to about ten consecutive nucleotides (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) in length and has at least 90% sequence identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) to the same region (e.g., 5' end) of the wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, a portion of a repeat sequence may comprise a pseudoknot-like structure (eg, a "handle") at its 5' end.
[0129] As used herein, "spacer sequence" is a nucleotide sequence that is complementary to a target nucleic acid (e.g., target DNA) (e.g., protospacer). The spacer sequence can be fully complementary or substantially complementary to the target nucleic acid (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more)). Therefore, in some embodiments, as compared to the target nucleic acid, the spacer sequence can have one, two, three, four or five mispairings, which can be continuous or discontinuous. In some embodiments, the spacer sequence can have 70% complementarity with the target nucleic acid. In other embodiments, the spacer nucleotide sequence can have 80% complementarity with the target nucleic acid. In still other embodiments, the spacer nucleotide sequence can have 85%, 90%, 95%, 96%, 97%, 98%, 99% or 99.5% complementarity with the target nucleic acid (protospacer). In certain embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence can have a length of about 15 nucleotides to about 30 nucleotides (for example, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides, or any range or value therein). Therefore, in some embodiments, the spacer sequence can have complete complementarity or basic complementarity in the region where the length of the target nucleic acid (for example, protospacer) is at least about 15 nucleotides to about 30 nucleotides. In certain embodiments, the length of the spacer is about 20 nucleotides. In certain embodiments, the length of the spacer is about 21, 22 or 23 nucleotides.
[0130] In some embodiments, the 5' region of the spacer sequence of the guide nucleic acid may be fully complementary to the target nucleic acid, while the 3' region of the spacer may be substantially complementary to the target nucleic acid (e.g., for a spacer of a type V CRISPR-Cas system), or the 3' region of the spacer sequence of the guide nucleic acid may be fully complementary to the target nucleic acid, while the 5' region of the spacer may be substantially complementary to the target nucleic acid (e.g., for a spacer of a type II CRISPR-Cas system), and thus, the overall complementarity of the spacer sequence to the target nucleic acid may be less than 100%. Thus, for example, in a guide nucleic acid of a type V CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 5' region (i.e., the seed region) of a 20-nucleotide spacer sequence may be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence may be substantially complementary to the target nucleic acid (e.g., at least about 70% complementary). In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides, and any range therein) of the 5' end of the spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more)).
[0131] As another example, in a guide nucleic acid of a type II CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 3' region (i.e., seed region) of, for example, a 20-nucleotide spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target nucleic acid. In some embodiments, the first 1 to 10 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, and any ranges therein) of the 3' end of the spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 50% complementary (e.g., at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, or any range or value therein)).
[0132] In some embodiments, the seed region of a spacer can be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.
[0133] In some embodiments, the editing system of the present invention comprises an extended guide nucleic acid, a fusion protein of the present invention, and optionally a reverse transcriptase. In some embodiments, the fusion protein of the present invention comprises all or a portion of a reverse transcriptase. In some embodiments, the fusion protein of the present invention, the extended guide nucleic acid, and optionally a reverse transcriptase can form or be contained in a complex that is capable of interacting with the target nucleic acid.
[0134] In certain embodiments, the guide nucleic acid further comprises a reverse transcriptase template and can be referred to as the guide nucleic acid of extension. As used herein, "the guide nucleic acid of extension" is a guide nucleic acid as described herein, which further comprises a reverse transcriptase template (RTT) and / or a primer binding site (PBS). In certain embodiments, the guide nucleic acid of extension is a lead editing guide RNA (pegRNA) of engineering approaches. The guide nucleic acid of extension can be a targeting allele guide RNA (tagRNA) or a stable targeting allele guide RNA (stagRNA). As used herein, "tagRNA" refers to a guide nucleic acid comprising a PBS and RTT and having an extension of target strand complementarity. As used herein, "stagRNA" refers to a tagRNA comprising a stabilization motif. The stabilization motif can be present at the 3' and / or 5' ends of the tagRNA. In certain embodiments, the stabilization motif is present at the 3' end of the tagRNA. Exemplary stabilization motifs include but are not limited to raising motifs, RNA hairpins, pseudoknot sequences and / or PP7 motifs (e.g., PP7RNA hairpin sequences). In certain embodiments, stagRNA is a tagRNA comprising a PP7 RNA hairpin sequence. In some embodiments, the CRISPR-Cas effector protein (e.g., a type II or type V CRISPR-Cas effector protein), the reverse transcriptase, and the extended guide nucleic acid can form or be contained in a complex.
[0135] In some embodiments, the extended guide nucleic acid comprises an extension portion comprising a primer binding site and a reverse transcriptase template, wherein the reverse transcriptase template comprises a modification (e.g., an edit) to be incorporated into the target nucleic acid. In some embodiments, the extended guide nucleic acid comprises a primer binding site and a modification (e.g., an edit) to be incorporated into the target nucleic acid (e.g., a reverse transcriptase template) at its 3' end. In some embodiments, the extended guide nucleic acid comprises: (1) a sequence that interacts (e.g., recruits and / or binds) with a CRISPR-Cas effector protein (e.g., a CRISPR-Cas nuclease), (2) a spacer that is substantially complementary to a first site on a target nucleic acid (e.g., a CRISPR RNA (crRNA) (first crRNA) and / or tracrRNA+crRNA (sgRNA)), and (3) a nucleic acid-encoded repair template (e.g., an RNA-encoded repair template) comprising a primer binding site and an RNA template (e.g., which encodes the modification to be incorporated into the target nucleic acid). In some embodiments, the extended guide nucleic acid (e.g., extended guide RNA) may comprise a spacer sequence, a repetitive sequence, and an extension from 5' to 3', wherein the extension comprises a reverse transcriptase template and a primer binding site from 5' to 3'. In some embodiments, the extended guide nucleic acid may comprise a spacer sequence, a repetitive sequence, and an extension from 5' to 3', wherein the extension comprises a primer binding site and a reverse transcriptase template from 5' to 3'. In some embodiments, the extended guide nucleic acid may comprise an extension, a spacer sequence, and a repetitive sequence from 5' to 3', wherein the extension comprises a reverse transcriptase template and a primer binding site from 5' to 3'. In some embodiments, the extended guide nucleic acid may comprise an extension, a spacer sequence, and a repetitive sequence from 5' to 3', wherein the extension comprises a primer binding site and a reverse transcriptase template from 5' to 3'.
[0136] According to some embodiments, the guide nucleic acid (for example, pegRNA) of extension may have such as Anzalone et al., Nature, December 2019; 576 (7785): Structure described in 149-157 and / or designed as described therein.In certain embodiments, the guide nucleic acid of extension includes a primer binding site (PBS) optionally with a sequence of 1, 2, 3, 4 or 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 nucleotides and a reverse transcriptase template (RT template) sequence optionally with a sequence of 65 or more nucleotides.In certain embodiments, the PBS of the guide nucleic acid of extension has a sequence less than 15 nucleotides and has a sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14 nucleotides (for example, a sequence of 5 or 6 nucleotides in length). RT template sequence can be after the PBS sequence in 5' to 3' directions. In some embodiments, the RT template sequence of the extended guide nucleic acid has a length greater than 65 nucleotides and may include about 50 or more nucleotides that are heterologous to the target site (e.g., target nucleic acid), followed by about 15 or more nucleotides that are homologous to the target site. In some embodiments, the RT template sequence of the extended guide nucleic acid is after the PBS sequence, and the RT template sequence has a length greater than 65 nucleotides, wherein the sequence includes more than 50 nucleotides that are heterologous to the target site, followed by more than 15 nucleotides that are homologous to the target site. Therefore, in some embodiments, when the extended guide nucleic acid is reverse transcribed, the resulting newly transcribed sequence can hybridize with the unnicked strand of the target site and / or be configured to hybridize with the unnicked strand of the target site, which can thereby produce heteroduplex DNA with a large insertion in the newly synthesized chain. After repairing this mismatched DNA, the resulting repaired DNA can contain a large insertion (e.g., greater than 50 nucleotides) of a DNA sequence. In some embodiments, the method can provide a large deletion (e.g., greater than 50 nucleotides) of a DNA sequence. In some embodiments, the PBS and 15 or more nucleotides homologous to the target site may comprise homology arms that can be used to insert heterologous DNA into the target site, optionally using homology-directed repair. The inserted DNA can correspond to any functional DNA sequence, such as, but not limited to: a functional transgene; a DNA fragment inserted into a gene in a manner that, when transcribed, produces a hairpin RNA sufficient to silence the homologous gene by RNAi; and / or one or more functional site-specific recombination sites, such as lox, frt, which can then be used for subsequent Cre- or Flp-mediated site-specific recombination processes. In some embodiments, the extended guide nucleic acid may be too large to be generated in vivo using a PolIII promoter.In some embodiments, the extended guide nucleic acid can be operably associated with a PolII promoter and / or produced using a PolII promoter. In some embodiments, a DNA binding polypeptide (e.g., a DNA binding domain) and / or a DNA endonuclease may have a structure as described in Anzalone et al., Nature, December 2019; 576(7785): 149-157 and / or be designed as described therein. In some embodiments, the DNA binding domain and / or the DNA endonuclease is a CRISPR Cas polypeptide, such as a Cas9 nickase, a nick variant of another CRISPR Cas polypeptide, or Cas12a.
[0137] In some embodiments, two extended guide nucleic acids (e.g., pegRNA) can be used (e.g., an editing system can include two extended guide nucleic acids). One or both of the two extended guide nucleic acids can have a structure as described in Anzalone et al., Nature, December 2019; 576(7785): 149-157 and / or be designed as described therein. The two extended guide nucleic acids can include a primer binding site (PBS) optionally having a sequence of 1, 2, 3, 4 or 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 nucleotides and a reverse transcriptase template (RT template) sequence optionally having a sequence of 50 or more nucleotides. The RT template sequences of the two extended guide nucleic acids can be complementary to each other, and therefore the polynucleotides reverse transcribed from each RT template will be complementary to each other and will be able to hybridize with each other. This can allow the intermediate produced by this system and / or method to connect two DNA segments together, and the two DNA segments were originally separated by more than 50 nucleotides, such as within a chromosome, or located on two separate DNA fragments, such as on two different chromosomes. After repairing the intermediate, according to the design of the RT template, the resulting products can produce large segment deletions, large segment inversions or interchromosomal recombination. Since all these products are produced by homology-directed repair, these products may be predictable, accurate and / or reproducible. In certain embodiments, DNA binding polypeptides (e.g., DNA binding domains) and / or DNA endonucleases may have structures such as Anzalone et al., Nature, December 2019; 576 (7785): 149-157 and / or are designed as described therein. In certain embodiments, DNA binding polypeptides and / or DNA endonucleases are CRISPR Cas polypeptides, such as Cas9 nickases, similar nick variants of another CRISPR Cas polypeptide, or Cas12a. In some embodiments, the DNA binding polypeptide and / or DNA endonuclease is a Cas9 nuclease, a similar nuclease from another CRISPR Cas polypeptide, or Cas12a. The use of a nuclease (rather than a nickase) can promote intrachromosomal or interchromosomal recombination processes by single-stranded annealing of 3' overhangs of more than 50 nucleotides, which will be generated at each of the two target sites corresponding to the two pegRNA target nucleic acids. In some embodiments, the editing system comprises an extended guide nucleic acid and a guide nucleic acid that does not contain a reverse transcriptase template and / or primer binding site.
[0138] The extended guide nucleic acid can comprise a CRISPR nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid; and (b) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the RT template encodes the modification to be incorporated into the target nucleic acid. The CRISPR nucleic acid can be a type II or type V CRISPR nucleic acid, and / or the tracr nucleic acid can be any tracr corresponding to an appropriate type II or type V CRISPR nucleic acid. In some embodiments, the extended guide nucleic acid comprises: (i) a type V CRISPR nucleic acid or a type II CRISPR nucleic acid (e.g., a type II or V CRISPR RNA, a type II or V CRISPR DNA, a type II or V crRNA, or a type II or V crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid (e.g., a type II or V tracrRNA, a type II or V tracrDNA); and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the type V CRISPR nucleic acid or type II CRISPR nucleic acid comprises a spacer that binds to a first strand (e.g., target strand) of a target nucleic acid (e.g., the spacer is complementary to a portion of contiguous nucleotides in the first strand of the target nucleic acid) and the primer binding site is bound to the first strand (e.g., target strand). In some embodiments, the extension portion can be fused to the 5' end or 3' end of the CRISPR nucleic acid (e.g., from 5' to 3': repeat sequence-spacer-extension portion or extension portion-repeat sequence-spacer) and / or fused to the 5' end or 3' end of the tracr nucleic acid. In some embodiments, the extension portion of the extended guide nucleic acid comprises an RT template (RTT) and a primer binding site (PBS) from 5' to 3' (e.g., 5'-crRNA-spacer-RTT (editing code)-PBS-3'), or comprises a PBS and RTT from 5' to 3' (e.g., 5'-crRNA-spacer-PBS-RTT (editing code)-3'), depending on the position of the extension portion relative to the CRISPR nucleic acid of the extended guide nucleic acid. For example, in some embodiments, the extension portion of the extended guide nucleic acid may comprise an RT template and a primer binding site from 5' to 3' (when the extended guide sequence is attached to the 3' end of the CRISPR nucleic acid). In some embodiments, the extended portion of the extended guide sequence from 5' to 3' may comprise a primer binding site and an RT template (when the extended guide sequence is ligated to the 5' end of the CRISPR nucleic acid).
[0139] In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the second strand (e.g., the non-target, top strand) of the target nucleic acid. In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the first strand of the target nucleic acid (e.g., binds to the target strand, optionally the same strand that recruits the CRISPR-Cas effector protein, the bottom strand). In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the second strand (e.g., the non-target strand, optionally the strand opposite to the strand that recruits the CRISPR-Cas effector protein) of the target nucleic acid. In some embodiments, a reverse transcriptase (RT) can be added to the target strand of the target nucleic acid (e.g., the strand that is complementary to the spacer of the CRISPR nucleic acid of the extended guide nucleic acid and recruits the CRISPR-Cas effector protein). In some embodiments, a reverse transcriptase (RT) is added to the non-target strand of the target nucleic acid (e.g., the strand that is complementary to the spacer of the same CRISPR nucleic acid and recruits the CRISPR-Cas effector protein). Exemplary methods and editing systems are described in International Patent Publication No. WO 2021 / 092130, International Patent Publication No. WO 2022 / 098993, and U.S. Patent Application Publication Nos. 2021 / 0147862, 2021 / 0130835, 2021 / 0147862, and 2022 / 0145334, each of which is incorporated herein by reference in its entirety.
[0140] The RT template of the extended guide nucleic acid can encode one or more modifications to be incorporated into the target nucleic acid (e.g., edits). The one or more modifications can be located at any position within the RT template (e.g., where the positional positioning can be relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid). In some embodiments, the RT template has a modification at one or more positions of -1 to 23 (e.g., -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23) relative to the position of the protospacer adjacent motif (PAM) (e.g., TTTG) in the target nucleic acid. In some embodiments, the RT template may include a modification at nucleotide position -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. In some embodiments, the RT template may comprise a modification located at nucleotide position 4 to nucleotide position 17 (e.g., position 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the PAM of the target nucleic acid. In some embodiments, the RT template may comprise a modification located at nucleotide position 10 to nucleotide position 17 (e.g., position 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the PAM of the target nucleic acid. In some embodiments, the RT template may comprise a modification located at nucleotide position 12 to nucleotide position 15 (e.g., position 12, 13, 14, or 15) of the RT template relative to the position of the PAM of the target nucleic acid.
[0141] In some embodiments, the extension portion of the extended guide nucleic acid from 5' to 3' may comprise an RT template and a primer binding site (e.g., when the extension portion is attached to the 3' end of a CRISPR nucleic acid). In some embodiments, the extension portion of the extended guide nucleic acid from 5' to 3' may comprise a primer binding site and an RT template (RTT) (e.g., when the extension portion is attached to the 5' end of a CRISPR nucleic acid). In some embodiments, the length of the RT template can be from about 1 nucleotide to about 100 nucleotides (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, and any range or value therein), for example, a length of about 1 nucleotide to about 10 nucleotides, about 1 nucleotide to about 15 nucleotides, about 1 nucleotide to about 20 nucleotides, about 1 nucleotide to about 25 nucleotides, about 1 nucleotide to about about 30 nucleotides, about 1 nucleotide to about 35, 36, 37, 38, 39 or 40 nucleotides, about 1 nucleotide to about 50 nucleotides, about 5 nucleotides to about 15 nucleotides, about 5 nucleotides to about 20 nucleotides, about 5 nucleotides to about 25 nucleotides, about 5 nucleotides to about 30 nucleotides, about 5 nucleotides to about 35, 36, 37, 38, 39 or 40 nucleotides, about 5 nucleotides to about 50 nucleotides, about 8 nucleotides to about 15 nucleotides, about 8 nucleotides to about 20 nucleotides, about 8 nucleotides to about 25 nucleotides, about 8 nucleotides to about nucleotides, about 10 nucleotides to about 30 nucleotides, about 10 nucleotides to about 36 nucleotides, about 10 nucleotides to about 40 nucleotides, about 10 nucleotides to about 50 nucleotides, about 10 nucleotides to about 100 nucleotides, about 10 nucleotides to about 15 nucleotides, about 10 nucleotides to about 20 nucleotides, about 10 nucleotides to about 25 nucleotides, about 10 nucleotides to about 30 nucleotides, about 10 nucleotides to about 36 nucleotides, about 10 nucleotides to about 40 nucleotides, about 10 nucleotides to about 50 nucleotides, about 10 nucleotides to about 100 nucleotides in length, and any range or value therein.In some embodiments, the RT template can be at least 8 nucleotides in length, optionally from about 8 nucleotides to about 100 nucleotides in length. In some embodiments, the RT template is 36, 37, 38, 39, or 40 nucleotides in length or less (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length), or any value or range thereof (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length to about 16, 17, 18, 19, 20 47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110 In some embodiments, the RT template may be about 36, 40, 44, 47, 50, 52, 55, 63, 72, or 74 nucleotides in length. One or more modifications may be present within the length of the RTT. The one or more modifications may be located at any position within the RTT, wherein the position of the modification may be described relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid. In some embodiments, the RT template may comprise a nucleotide sequence at position -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, In some embodiments, the RT template may comprise a modification at nucleotide position 4 to nucleotide position 17 (e.g., position 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the protospacer-adjacent motif (PAM) of the target nucleic acid. In some embodiments, the RT template may comprise a modification at nucleotide position 10 to nucleotide position 17 (e.g., position 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the protospacer-adjacent motif (PAM) of the target nucleic acid.In some embodiments, the RT template can comprise a modification at nucleotide position 12 to nucleotide position 15 (e.g., position 12, 13, 14, or 15) of the RT template relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid.
[0142] As used herein, the "primer binding site" (PBS) of the extended portion of the guide nucleic acid (e.g., tagRNA) of extension refers to a region or "primer" that can be combined with a target nucleic acid, such as a continuous nucleotide sequence complementary to a target nucleic acid primer. As an example, a CRISPR Cas effector protein (e.g., type II or type V, such as Cas 9 or Cas12a) can nick / cut DNA, and the 3' end of the cut DNA serves as a primer for the PBS portion of the extended guide nucleic acid. The PBS can be complementary to the 3' end of the chain of the target nucleic acid, and can be combined with the target chain or non-target chain and / or can be configured to be combined with the target chain or non-target chain. The primer binding site can be fully complementary to the primer, or it can be substantially complementary (e.g., at least 70% complementary (e.g., 70% or about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or more)) to the primer of the target nucleic acid. 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110 nucleotides, about 45 nucleotides to about 80 nucleotides, about 45 nucleotides to about 80 nucleotides, about 45 nucleotides to about 80 nucleotides, about 45 nucleotides to about 80 nucleotides, or about 45 nucleotides to about 60 nucleotides, or any range or value therein).72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides in length, or any range or value thereof. In some embodiments, the PBS may be at least 30 nucleotides in length, optionally from about 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length to about 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides in length.
[0143] In some embodiments, the RTT may be about 35 nucleotides to about 75 nucleotides in length and the PBS may be about 30 nucleotides to about 80 nucleotides in length, optionally wherein the PBS may be about 8, 16, 24, 32, 40, 48, 56, 64, 72, or 80 nucleotides in length and the RTT may be about 36, 40, 44, 47, 50, 52, 55, 63, 72, or 74 nucleotides in length, or any combination of RTT length and / or PBS length.
[0144] In some embodiments, the extension portion of the extended guide nucleic acid can be fused to the 5' or 3' end of a type II or type V CRISPR nucleic acid (e.g., from 5' to 3': repeat sequence-spacer-extension portion or extension portion-repeat sequence-spacer) and / or to the 5' or 3' end of a tracr nucleic acid. In some embodiments, when the extension portion is located 5' of the crRNA, the type V CRISPR-Cas effector protein is modified to reduce (or eliminate) self-processing RNAse activity.
[0145] In some embodiments, the extended portion of the extended guide nucleic acid can be connected to a type II or type V CRISPR nucleic acid and / or a type II or type V tracrRNA via a linker. In some embodiments, the linker is about 1 to about 100 or more nucleotides in length (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, and any range therein (e.g., a length of about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, about 40 to about 100, about 50 to about 100, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 nucleotides to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 98, 99, 100, 110, 115, 120, 130, 140, 150 or more nucleotides in length.
[0146] The guide nucleic acid and / or the guide nucleic acid of extension may include one or more recruitment motifs as described herein, which may be connected to the 5' end and / or 3' end of the guide nucleic acid and / or it may be inserted into the guide nucleic acid (e.g., in the hairpin loop of the guide nucleic acid). In certain embodiments, the guide nucleic acid of extension may be connected to an RNA recruitment motif. The guide nucleic acid of extension and / or the guide nucleic acid may be connected to one or two or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs; e.g., at least 10 to about 25 motifs), optionally wherein two or more RNA recruitment motifs may be identical RNA recruitment motifs or different RNA recruitment motifs. In certain embodiments, the RNA recruitment motif may be located at the 3' end of the extension portion of the guide nucleic acid of extension (e.g., from 5' to 3', repeat sequence-spacer-extension portion (RT template-primer binding site)-RNA recruitment motif). In certain embodiments, the RNA recruitment motif may be embedded in the extension portion of the guide nucleic acid of extension.
[0147] In some embodiments, the editing system comprises an extended guide nucleic acid linked to an RNA recruitment motif and a reverse transcriptase as a reverse transcriptase fusion protein, wherein the reverse transcriptase fusion protein comprises a reverse transcriptase polypeptide fused to an affinity polypeptide that binds to the RNA recruitment motif, wherein the extended guide nucleic acid binds to the target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the reverse transcriptase fusion protein to the extended guide nucleic acid and contacting the target nucleic acid with the reverse transcriptase. In some embodiments, two or more reverse transcriptase fusion proteins can be recruited to the extended guide nucleic acid, thereby contacting the target nucleic acid with two or more reverse transcriptase fusion proteins.
[0148] "Target nucleic acid," "target DNA," "target nucleotide sequence," "target region," and "target region in a genome" are used interchangeably herein and refer to a region in the genome of an organism (e.g., a plant) that comprises a sequence that is fully complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) to a spacer sequence in a guide nucleic acid as defined herein. The target nucleic acid is targeted by an editing system (or component thereof) as described herein. The target region that can be used for the CRISPR-Cas system can be positioned immediately 3' (e.g., type V CRISPR-Cas system) or immediately 5' (e.g., type II CRISPR-Cas system) of the PAM sequence in the genome of an organism (e.g., a plant genome or a mammalian (e.g., human) genome). The target region can be selected from any region of at least 15 consecutive nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) positioned immediately adjacent to the PAM sequence.
[0149] As used herein, a "protospacer sequence" or "protospacer" refers to a sequence that is completely or substantially complementary to (and can hybridize with) a spacer sequence of a guide nucleic acid. In some embodiments, a protospacer is all or a portion of a target nucleic acid as defined herein that is completely or substantially complementary to (and hybridizes with) a spacer sequence of a CRISPR repeat-spacer sequence (e.g., a guide nucleic acid, a CRISPR array, a crRNA).
[0150] In the case of type V CRISPR-Cas (e.g., Cas12a) systems and type II CRISPR-Cas (Cas9) systems, the protospacer sequence is flanked by (e.g., immediately adjacent to) a protospacer adjacent motif (PAM). For type IV CRISPR-Cas systems, the PAM is located at the 5' end of the non-target strand and the 3' end of the target strand (see below, as an example).
[0151]
[0152] In the case of type II CRISPR-Cas (e.g., Cas9) systems, the PAM is positioned immediately adjacent to the 3' location of the target region. The PAM of the type I CRISPR-Cas system is located at the 5' of the target strand. There is no known PAM for the type III CRISPR-Cas system. Makarova et al. describe the nomenclature of all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). Guide structures and PAMs are described in R. Barrangou (Genome Biol. 16:247 (2015)).
[0153] Typical Cas12a PAM is rich in T. In some embodiments, a typical Cas12a PAM sequence can be 5'-TTN, 5'-TTTN or 5'-TTTV. In some embodiments, a typical Cas9 (e.g., Streptococcus pyogenes (S. pyogenes)) PAM can be 5'-NGG-3'. In some embodiments, atypical PAMs can be used, but the efficiency may be lower.
[0154] Those skilled in the art can determine additional PAM sequences by established experiments and calculation methods.Therefore, for example, experimental methods include targeting the sequence of all possible nucleotide sequences of the side joint and identifying sequence members that do not experience targeting, such as by conversion of target plasmid DNA (Esvelt et al., 2013.Nat.Methods 10:1116-1121; Jiang et al., 2013.Nat.Biotechnol.31:233-239). In some aspects, calculation methods may include carrying out BLAST search to identify the original target DNA sequence in phage or plasmid to natural spacers, and aligning these sequences to determine conserved sequences adjacent to the target sequence (Briner and Barrangou, 2014.Appl.Environ.Microbiol.80:994-1001; Mojica et al., 2009.Microbiology 155:733-740).
[0155] In some embodiments, the invention provides expression cassettes and / or vectors comprising nucleic acid constructs of the invention (e.g., one or more components of an editing system of the invention). In some embodiments, expression cassettes and / or vectors comprising nucleic acid constructs of the invention and / or one or more guide nucleic acids may be provided. In some embodiments, nucleic acid constructs of the invention encode fusion proteins and / or deaminases, and each may be contained on an expression cassette or vector that is the same as or separate from an expression cassette or vector that comprises one or more guide nucleic acids. When a nucleic acid construct encoding a component of a fusion protein or editing system is contained on an expression cassette or vector that is separate from an expression cassette or vector that comprises a guide nucleic acid, the target nucleic acid and the expression cassette or vector encoding the component of the fusion protein or editing system can be contacted with each other and the guide nucleic acid (e.g., provided together) in any order, such as before, simultaneously with, or after providing an expression cassette comprising a guide nucleic acid (e.g., contacted with the target nucleic acid).
[0156] Methods for recruiting one or more components of an editing system to each other and / or a target nucleic acid are known in the art and may include the use of peptide tags or affinity polypeptides that interact with peptide tags. In some embodiments, a guide nucleic acid can be linked to an RNA recruitment motif, and a deaminase can be linked to an affinity polypeptide that can interact with the RNA recruitment motif, thereby recruiting the deaminase to the target nucleic acid. Alternatively, a chemical interaction can be used to recruit a polypeptide (e.g., a deaminase) to a target nucleic acid.
[0157] As used herein, "raising motif" refers to one half of a binding pair, which can be used to recruit the compound to which the raising motif is bound to another compound (i.e., "corresponding motif") comprising the other half of the binding pair. The raising motif and the corresponding motif can be non-covalently bound. In certain embodiments, the raising motif is an RNA raising motif (e.g., an RNA raising motif capable of binding to an affinity polypeptide and / or configured to bind to an affinity polypeptide), an affinity polypeptide (e.g., an RNA raising motif and / or a peptide tag capable of binding to an affinity polypeptide and / or configured to bind to an RNA raising motif and / or a peptide tag) or a peptide tag (e.g., a peptide tag capable of binding to an affinity polypeptide and / or configured to bind to an affinity polypeptide). For example, when the raising motif is an RNA raising motif, the corresponding motif of the RNA raising motif can be an affinity polypeptide in conjunction with the RNA raising motif. Another example is that when the raising motif is a peptide tag, the corresponding motif of the peptide tag can be an affinity polypeptide in conjunction with the peptide tag. Thus, a compound comprising a recruitment motif (eg, an affinity polypeptide) can be recruited to another compound (eg, a guide nucleic acid) comprising a corresponding motif of the recruitment motif (eg, an RNA recruitment motif).
[0158] Peptide tags (e.g., epitopes) that can be used in the present invention may include, but are not limited to, GCN4 peptide tags (e.g., Sun-Tag), c-Myc affinity tags, HA affinity tags, His affinity tags, S affinity tags, methionine-His affinity tags, RGD-His affinity tags, FLAG octapeptide, strep tags or strep tags II, V5 tags, and / or VSV-G epitopes. Any epitope that can be linked to a polypeptide and for which there is a corresponding affinity polypeptide that can be linked to another polypeptide can be used as a peptide tag in the present invention. In some embodiments, the peptide tag may comprise 1 or 2 or more copies (e.g., repeating units, multimeric epitopes (e.g., tandem repeats)) of the peptide tag (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more repeating units). In some embodiments, the affinity polypeptide that interacts / binds with the peptide tag may be an antibody. In some embodiments, the antibody may be a scFv antibody. In some embodiments, the affinity polypeptide bound to the peptide tag may be synthetic (e.g., evolved for affinity interactions), including but not limited to affibodies, anticalins, monobodies, and / or DARPins (see, e.g., Sha et al., Protein Sci. 26(5):910-924 (2017)); Gilbreth (Curr Opin Struc Biol 22(4):413-420 (2013)), U.S. Patent No. 9,982,053, each of which is incorporated by reference in its entirety for its teachings on affibodies, anticalins, monobodies, and / or DARPins.
[0159] In some embodiments, the guide nucleic acid can be linked to an RNA recruitment motif, and the polypeptide to be recruited (e.g., a deaminase) can be fused to an affinity polypeptide that binds to the RNA recruitment motif, wherein the guide sequence binds to the target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the polypeptide to the guide sequence and contacting the target nucleic acid with the polypeptide (e.g., a deaminase). In some embodiments, two or more polypeptides can be recruited to the guide nucleic acid, thereby contacting the target nucleic acid with two or more polypeptides (e.g., a deaminase).
[0160] In some embodiments of the invention, a guide RNA can be linked to one or two or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs; such as at least 10 to about 25 motifs), optionally wherein the two or more RNA recruitment motifs can be the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide may include, but are not limited to, the telomerase Ku binding motif (e.g., Ku binding hairpin) and the corresponding affinity polypeptide Ku (e.g., Ku heterodimer), the telomerase Sm7 binding motif and the corresponding affinity polypeptide Sm7, the MS2 phage operator stem-loop and the corresponding affinity polypeptide MS2 coat protein (MCP), the PP7 phage operator stem-loop and the corresponding affinity polypeptide PP7 coat protein (PCP), the SfMu phage Com stem-loop and the corresponding affinity polypeptide Com RNA binding protein, the PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF), and / or a synthetic RNA aptamer and aptamer ligand as the corresponding affinity polypeptide. In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide may be the MS2 phage operator stem-loop and the affinity polypeptide MS2 coat protein (MCP). In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide can be a PUF binding site (PBS) and an affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF). Exemplary RNA recruitment motifs and corresponding affinity polypeptides that can be used in the present invention can include, but are not limited to, SEQ ID NOs: 86-96.
[0161] In some embodiments, components for recruiting polypeptides and nucleic acids may include components that function through chemical interactions, which may include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin; SNAP tag; Halo tag; CLIP tag; compound-induced DmrA-DmrC heterodimers; bifunctional ligands (e.g., chemically induced dimerization).
[0162] As described herein, "peptide tags" can be used to recruit one or more polypeptides. A peptide tag can be any polypeptide that can be bound by a corresponding motif (e.g., an affinity polypeptide). A peptide tag may also be referred to as an "epitope," and when provided in multiple copies, is referred to as a "multimerization epitope." Exemplary peptide tags may include, but are not limited to, a GCN4 peptide tag (e.g., Sun-Tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or a strep tag II, a V5 tag, and / or a VSV-G epitope. In some embodiments, the peptide tag may also include a phosphorylated tyrosine in a specific sequence context recognized by an SH2 domain, a characteristic consensus sequence containing phosphoserine recognized by a 14-3-3 protein, a proline-rich peptide motif recognized by an SH3 domain, a PDZ protein interaction domain, or a PDZ signal sequence, and an AGO hook motif from a plant. Peptide tags are disclosed in WO 2018 / 136783 and U.S. Patent Application Publication No. 2017 / 0219596, the disclosures of which regarding peptide tags are incorporated by reference. Peptide tags that can be used in the present invention may include, but are not limited to, SEQ ID NO: 97 and SEQ ID NO: 98. Affinity polypeptides that can be used for peptide tags include, but are not limited to, SEQ ID NO: 99.
[0163] The peptide tag may comprise or be present in one copy or two or more copies of a peptide tag (e.g., a multimeric peptide tag or multimeric epitope) (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, or 25 or more peptide tags). When multimerized, the peptide tags may be directly fused to each other, or they may be linked to each other via one or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more amino acids, optionally about 3 to about 10, about 4 to about 10, about 5 to about 10, about 5 to about 15, or about 5 to about 20 amino acids, etc., and any value or range therein). Thus, in some embodiments, the CRISPR-Cas effector protein of the present invention may comprise a CRISPR-Cas effector protein fused to one peptide tag or to two or more peptide tags, optionally wherein the two or more peptide tags are fused to each other via one or more amino acid residues. In some embodiments, the peptide tag useful in the present invention may be a single copy of a GCN4 peptide tag or epitope, or may be a multimerized GCN4 epitope comprising about 2 to about 25 or more copies of a peptide tag (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more copies of a GCN4 epitope, or any range thereof).
[0164] In some embodiments, the peptide tag can be fused to a CRISPR-Cas polypeptide or domain. In some embodiments, the peptide tag can be fused or connected to the C-terminus of the CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, the peptide tag can be fused or connected to the N-terminus of the CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, the peptide tag can be fused within the CRISPR-Cas effector protein (for example, the peptide tag can be in the loop region of the CRISPR-Cas effector protein). In some embodiments, the peptide tag can be fused to a cytosine deaminase and / or adenine deaminase.
[0165] "Affinity polypeptide" (e.g., "recruiting polypeptide") refers to any polypeptide that can bind to its corresponding peptide tag, peptide tag, or RNA recruitment motif. The affinity polypeptide of the peptide tag can be, for example, an antibody and / or single-chain antibody that specifically binds to the peptide tag, respectively. In some embodiments, the antibody of the peptide tag can be, but is not limited to, a scFv antibody. In some embodiments, the affinity polypeptide can be fused or linked to the N-terminus of a deaminase (e.g., cytosine deaminase or adenine deaminase). In some embodiments, the affinity polypeptide is stable under reducing conditions of a cell or cell extract.
[0166] The nucleic acid constructs of the present invention and / or guide nucleic acids can be contained in one or more expression cassettes as described herein. In certain embodiments, the nucleic acid constructs of the present invention can be contained in an expression cassette or vector that is identical or separate from an expression cassette or vector that contains a guide nucleic acid and / or an extended guide nucleic acid.
[0167] In some embodiments, a nucleic acid construct, expression cassette, or vector of the invention that is optimized for expression in an organism (e.g., a human or a plant) may be about 70% to 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to a nucleic acid construct, expression cassette, or vector comprising the same polynucleotide but that has not been codon-optimized for expression in the organism.
[0168] When used in combination with a guide nucleic acid, nucleic acid construct of the present invention (and expression cassette and / or vector comprising it) can be used to modify target nucleic acid and / or its expression. Before, simultaneously or after contacting the target nucleic acid with a guide nucleic acid / raising guide nucleic acid (and / or expression cassette and vector comprising it), the target nucleic acid can be contacted with nucleic acid construct of the present invention and / or expression cassette and / or vector comprising it.
[0169] According to embodiments of the present invention, provided herein are fusion proteins (e.g., engineered proteins) including intein polypeptides. In certain embodiments, the fusion protein of the present invention includes Cas12a polypeptides and intein polypeptides. In certain embodiments, the fusion protein of the present invention includes a polypeptide of interest and an intein polypeptide. In certain embodiments, the fusion protein of the present invention includes a reverse transcriptase polypeptide and an intein polypeptide. As used herein, "engineered protein" is a polypeptide or protein not naturally found in nature. The fusion protein of the present invention may include a Cas12a polypeptide and / or a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to an intein polypeptide (e.g., connected (link) and / or connected (attach)). Cas12a polypeptides and intein polypeptides can be directly fused (e.g., there is no amino acid residue or linker between the two polypeptides) or indirectly fused (e.g., there is a linker (e.g., amino acid or peptide) or another polypeptide between the two polypeptides). Similarly, a polypeptide of interest (e.g., a reverse transcriptase polypeptide) and an intein polypeptide can be directly fused or indirectly fused. In certain embodiments, the Cas12a polypeptide is directly fused to an intein polypeptide (e.g., via a peptide bond). In some embodiments, the Cas12a polypeptide is indirectly fused to the intein polypeptide (e.g., via a peptide linker). The Cas12a polypeptide and the intein polypeptide can be fused in any orientation. For example, in some embodiments, the N-terminus of the Cas12a polypeptide is fused to the C-terminus of the intein polypeptide or to the N-terminus of the intein polypeptide. In some embodiments, the C-terminus of the Cas12a polypeptide is fused to the C-terminus of the intein polypeptide or to the N-terminus of the intein polypeptide. The polypeptide of interest (e.g., a reverse transcriptase polypeptide) and the intein polypeptide can be fused in any orientation. In some embodiments, the fusion protein of the present invention has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NO: 100-109 or 187-188. According to some embodiments, a nucleic acid molecule is provided that encodes a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to one or more of SEQ ID NOs: 100-109 or 187-188. In some embodiments, the nucleic acid molecule comprises a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to one or more of SEQ ID NOs: 185-186.
[0170] In some embodiments, the intein polypeptide is a portion (e.g., a fragment, such as an N-terminal intein fragment or a C-terminal intein fragment) of an intein, such as a portion of a molecular scaffold formed by two corresponding portions (e.g., two corresponding fragments or a pair of intein polypeptides) that together can or are configured to catalyze the cleavage and formation of peptide bonds. As used herein, an "intein" refers to a catalytically active complex formed by two polypeptides (e.g., a pair of intein polypeptides) that are associated with each other, wherein the complex can or is configured to excise itself (e.g., two polypeptides) from a larger precursor polypeptide, and can or is configured to ligate the ends of the two polypeptides with peptide bonds, optionally wherein the excision and ligation occur simultaneously. In some embodiments, the intein polypeptide is a portion of a split intein, such as a trans-splicing splint intein. As used herein, a "split intein" can undergo protein trans-splicing, wherein two fragments of an intein (e.g., two intein polypeptides or a pair of intein polypeptides) associate (e.g., non-covalently bond) to form a catalytically competent complex or molecular scaffold that catalyzes the excision of the two intein fragments and the ligation of their flanking sequences. In some embodiments, the intein polypeptide is an autocatalytic polypeptide that, together with a corresponding intein polypeptide, forms an intein (e.g., a split intein), capable of excising the intein polypeptide from a larger precursor protein (e.g., a fusion protein of the invention) and enabling the flanking polypeptide sequences (e.g., sequences adjacent to the excised intein polypeptide) to be joined by forming new peptide bonds. A fusion protein of the invention may include an intein polypeptide as one of two total parts, such that the intein polypeptide, together with another intein polypeptide (e.g., the second part), forms an intein, such as a trans-splicing split intein. In some embodiments, the split intein and / or its intein polypeptide may function (e.g., perform protein trans-splicing) without any assistance and / or conditions other than the two parts of the split intein (e.g., the two intein polypeptides that together provide the split intein). For example, two intein polypeptides can spontaneously associate to form an intein and can spontaneously catalyze their own excision and the joining of their flanking sequences without the need for assistance (e.g., external conditions and / or cofactors). In some embodiments, inteins can be used in which protein trans-splicing is controlled (e.g., the intein undergoes conditional trans-splicing). For example, certain conditions (e.g., light and / or cofactors) may be required for the intein to function.In some embodiments, the split intein and / or intein polypeptide is a light-inducible intein (e.g., as described in Wong S et al. (2015) An Engineered Split Intein for Photoactivated Protein Trans-Splicing. PLoS ONE 10(8):e0135965), which uses light to control the association of two intein polypeptides that together provide an intein (e.g., a catalytically active complex). In some embodiments, a cofactor (e.g., a small molecule) and / or an activator is used to bring two intein polypeptides together that together provide an intein to control protein trans-splicing, such as described in Gramespacher, Josef A. et al., J Am Chem Soc. 2019 Sep 4;141(35):13708-13712. In some embodiments, an intein polypeptide of a fusion protein of the invention can be configured to be removed (e.g., excised) from the fusion protein and fused in situ and / or in vivo with another intein polypeptide (e.g., an intein polypeptide that is part of a different fusion protein of the invention).
[0171] In some embodiments, the intein of the present invention is an intein present in the DNA polymerase III gene (DnaE) in cyanobacteria and / or an intein as described in Pinto, F., Thornton, EL and Wang, B. An expanded library of orthogonal split inteins enables modular multi-peptide assemblies. Nat Commun 11, 1529 (2020). Additional exemplary inteins include, but are not limited to, Nostoc punctiforme (Npu) intein and mutants thereof (e.g., NpuGEP, i.e., a mutant containing three amino acid residue mutations). In some embodiments, the intein polypeptide is a portion of the Nostoc punctiforme (Npu) intein and / or a portion of a mutant Npu intein (e.g., a portion of NpuGEP). The intein polypeptide can be the N-terminal portion of an intein, wherein the intein polypeptide comprises the N-terminus of the full-length intein and / or the active complex. In some embodiments, the intein polypeptide can be the C-terminal portion of an intein, wherein the intein polypeptide comprises the C-terminus of a full-length intein and / or an active complex. The intein polypeptide that is the N-terminal portion of an intein and the intein polypeptide that comprises the remainder of the intein (e.g., the C-terminal portion of the intein) together are an intein pair and form an intein and / or active complex. The intein polypeptides of the invention can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 110-112.
[0172] The Cas12a polypeptide of the present invention can be a part of a Cas12a protein, optionally a part of a Cas12a fusion protein. In some embodiments, the Cas12a protein and / or the Cas12a fusion protein can be the protein described in U.S. Patent Application Publication No. 2022 / 0112473, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the Cas12a polypeptide is a part of a sequence of SEQ ID NO: 38-60, 113-149, 192-195 or 196-259. In some embodiments, the Cas12a polypeptide is about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or more consecutive amino acids of a sequence of SEQ ID NO: 38-60, 113-149, 192-195 or 196-259. In some embodiments, the Cas12a polypeptide is a part of the two total parts that form the Cas12a protein together. For example, the first Cas12a polypeptide present in the first fusion protein of the present invention forms the Cas12a protein together with the second Cas12a polypeptide present in the second fusion protein of the present invention. In some embodiments, the Cas12a polypeptide is the N-terminal portion of the Cas12a protein, wherein the Cas12a polypeptide includes the N-terminal of the full-length Cas12a protein. In some embodiments, the Cas12a polypeptide is the C-terminal portion of the Cas12a protein, wherein the Cas12a polypeptide includes the C-terminal of the full-length Cas12a protein. After the two Cas12a polypeptides respectively present in two different fusion proteins of the present invention are fused together, the two Cas12a polypeptides can provide a Cas12a protein as a part of an editing system described herein, such as a CRISPR-Cas editing system. The editing system can be used to modify the target nucleic acid. In some embodiments, the fusion protein of the present invention comprises a Cas12a polypeptide as the N-terminal portion of the Cas12a protein and an intein polypeptide as the N-terminal portion of the intein, wherein the intein polypeptide is located at the C-terminus of the fusion protein and / or the C-terminus of the Cas-12a polypeptide. In some embodiments, the fusion protein of the present invention comprises a Cas12a polypeptide as the C-terminal portion of the Cas12a protein and an intein polypeptide as the C-terminal portion of the intein, wherein the intein polypeptide is located at the N-terminus of the fusion protein and / or the N-terminus of the Cas-12a polypeptide.
[0173] In some embodiments, the Cas12a protein is divided into two parts, and the fusion protein of the present invention comprises one part. For example, the Cas12a protein can be divided into two parts between amino acid residues 173 and 174, 174 and 175, 175 and 176, 309 and 310, 310 and 311, 405 and 406, 406 and 407, 440 and 441, 441 and 442, 549 and 550, or 550 and 551, such that the Cas12a polypeptide includes 173, 174, 175, 309, 310, 405, 406, 440, 441, 549 or 550 consecutive amino acids of the Cas12a protein and the other Cas12a polypeptide includes the remaining portion of the Cas12a protein (e.g., from amino acid residues 174, 175, 176, 310, 311, 406, 407, 441, 442, 549 or 551 to the end of the protein). In some embodiments, the Cas12a polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 150-159 and 175-184.
[0174] In certain embodiments, the fusion protein of the present invention includes all or part of a polypeptide of interest. In certain embodiments, the polypeptide of interest is fused to the Cas12a polypeptide and / or intein polypeptide (directly or via a connexon). In certain embodiments, the polypeptide of interest is fused to the Cas12a polypeptide and intein polypeptide (directly or via a connexon) so that the polypeptide of interest is between the Cas12a polypeptide and the intein polypeptide. In certain embodiments, the polypeptide of interest is fused to the Cas12a polypeptide and intein polypeptide (directly or via a connexon) so that the Cas12a polypeptide is between the polypeptide of interest and the intein polypeptide. In certain embodiments, the fusion protein of the present invention includes the polypeptide of interest fused to the intein polypeptide (directly or via a connexon), and the fusion protein does not contain Cas12a polypeptide. In certain embodiments, the polypeptide of interest is fused to the N-terminus of the intein polypeptide (directly or via a connexon). In certain embodiments, the polypeptide of interest is fused to the C-terminus of the intein polypeptide (directly or via a connexon).
[0175] The fusion protein of the present invention may include a reverse transcriptase. The reverse transcriptase polypeptide of the present invention can be all or part of a reverse transcriptase. In some embodiments, the reverse transcriptase has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NO:160-171. In some embodiments, the reverse transcriptase is fused to the Cas12a polypeptide (directly or via a linker) so that the Cas12a polypeptide is between the reverse transcriptase and the intein polypeptide. In some embodiments, the reverse transcriptase is fused to the Cas12a polypeptide and the intein polypeptide (directly or via a linker) so that the reverse transcriptase is between the Cas12a polypeptide and the intein polypeptide. In some embodiments, the reverse transcriptase is fused to the N-terminal of the Cas12a polypeptide (directly or via a linker). In some embodiments, the reverse transcriptase is fused to the C-terminal of the Cas12a polypeptide (directly or via a linker). In some embodiments, the reverse transcriptase polypeptide is fused to the intein polypeptide (directly or via a linker) to provide a fusion protein. In some embodiments, the fusion protein comprising a reverse transcriptase polypeptide and an intein polypeptide does not contain a Cas12a polypeptide.
[0176] The fusion protein of the present invention may comprise a nuclear localization signal. In some embodiments, the nuclear localization signal has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 172-174.
[0177] In some embodiments, a nucleic acid molecule is provided that encodes a fusion protein of the present invention. The nucleic acid molecule can be operably associated with a promoter. In some embodiments, an expression cassette or vector is provided that comprises a nucleic acid molecule encoding a fusion protein of the present invention. In some embodiments, an AAV vector is provided that comprises a nucleic acid molecule encoding a fusion protein of the present invention.
[0178] According to embodiments of the present invention, a complex may be provided, comprising Cas12a protein, a guide nucleic acid (e.g., a guide RNA and / or an extended guide nucleic acid), an optional reverse transcriptase, and an optional deaminase. In certain embodiments, the complex comprises Cas12a protein, an extended guide nucleic acid, and a reverse transcriptase. In certain embodiments, the complex comprises Cas12a protein, a guide nucleic acid, and a deaminase. The Cas12a protein of the complex of the present invention can be prepared by the first fusion protein of the present invention and the second fusion protein of the present invention, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide. After the first fusion protein is contacted with the second fusion protein (e.g., the first and second fusion proteins are provided together (e.g., in the same composition or cell) under conditions suitable for excision of the first and second intein polypeptides, association of the first and second intein polypeptides, and fusion of the first and second Cas12a polypeptides), the first and second intein polypeptides can associate to form an intein (e.g., an active complex), and the intein can excise the intein (e.g., the first and second intein polypeptides), and optionally the first and second Cas12a polypeptides are fused together with a linker (e.g., a peptide linker) between the first Cas12a polypeptide and the second Cas12a polypeptide.
[0179] In some embodiments, the complex of the present invention comprises an engineered protein (e.g., a fusion protein, a base editor, a templated editor, etc.) and a guide nucleic acid (e.g., a guide RNA). The base editor may comprise a CRISPR-Cas effector protein (e.g., Cas12a) and a deaminase. In some embodiments, the templated editor may comprise a CRISPR-Cas effector protein (e.g., Cas12a) and a reverse transcriptase. In some embodiments, the templated editor may be referred to as a REDRAW editor. In some embodiments, the complex of the present invention comprises: an engineered protein, which is prepared by the first fusion protein of the present invention and the second fusion protein of the present invention, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA). In some embodiments, the first fusion protein of the present invention and the second fusion protein of the present invention provide and / or form an engineered protein together, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide. In some embodiments, the fusion proteins of the invention optionally comprise all or a portion of a polypeptide of interest (e.g., a reverse transcriptase polypeptide), a linker, and an intein polypeptide in the N to C direction, optionally wherein the linker comprises the sequence of SEQ ID NO: 189. In some embodiments, the fusion proteins of the invention optionally comprise all or a portion of an intein polypeptide, a linker, and Cas12a in the N to C direction, optionally wherein the linker comprises the sequence of SEQ ID NO: 190.
[0180] In some embodiments, a composition is provided, comprising: a first fusion protein of the present invention, comprising a first Cas12a polypeptide fused to a first intein polypeptide; and a second fusion protein of the present invention, comprising a second Cas12a polypeptide fused to a second intein polypeptide. The first fusion protein and the second fusion protein may be different from each other. In some embodiments, the first intein polypeptide of the first fusion protein and the second intein polypeptide of the second fusion protein together form an intein and / or constitute two parts of a full-length intein, and / or the first Cas12a polypeptide of the first fusion protein and the second Cas12a polypeptide of the second fusion protein together form a Cas12a protein and / or constitute two parts of a full-length Cas12a protein. In some embodiments, the first and second fusion proteins are present in the same cell and can be optionally delivered to the cell using separate compositions or a composition comprising the first and second fusion proteins at the same time.
[0181] In some embodiments, the composition of the present invention includes: a first fusion protein of the present invention, which includes a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second fusion protein of the present invention, which includes a Cas12a polypeptide fused to a second intein polypeptide. The first fusion protein and the second fusion protein may be different from each other. In some embodiments, the first intein polypeptide of the first fusion protein and the second intein polypeptide of the second fusion protein together form an intein and / or constitute two parts of a full-length intein, and / or the polypeptide of interest of the first fusion protein and the Cas12a polypeptide of the second fusion protein together form a fusion protein (e.g., an engineered protein and / or a templated editor) and / or constitute two parts of a full-length fusion protein (e.g., an engineered protein and / or a templated editor). In some embodiments, the first and second fusion proteins are present in the same cell and can optionally be delivered to the cell using separate compositions or compositions comprising the first and second fusion proteins simultaneously.
[0182] In some embodiments, the composition of the present invention comprises: a first nucleic acid molecule encoding a first fusion protein, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein, wherein the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide. The first nucleic acid molecule can encode the fusion protein of the present invention, and the second nucleic acid molecule can encode the fusion protein of the present invention. In some embodiments, the first nucleic acid molecule is present in a first expression cassette and / or vector, and the second nucleic acid molecule is present in a second expression cassette and / or vector, wherein the first and second expression cassettes and / or vectors are separate and / or different from each other.
[0183] In some embodiments, the composition of the present invention comprises: a first nucleic acid molecule encoding a first fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein, wherein the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide. The first nucleic acid molecule can encode the fusion protein of the present invention, and the second nucleic acid molecule can encode the fusion protein of the present invention. In some embodiments, the first nucleic acid molecule is present in a first expression cassette and / or vector, and the second nucleic acid molecule is present in a second expression cassette and / or vector, wherein the first and second expression cassettes and / or vectors are separate and / or different from each other.
[0184] According to some embodiments of the present invention, a test kit may be provided. The test kit of the present invention may include: a first nucleic acid molecule encoding a first fusion protein, wherein the first fusion protein includes a first Cas12a polypeptide fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein, wherein the second fusion protein includes a second Cas12a polypeptide fused to a second intein polypeptide. In some embodiments, the test kit of the present invention includes: a first nucleic acid molecule encoding a first fusion protein, wherein the first fusion protein includes a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and a second nucleic acid molecule encoding a second fusion protein, wherein the second fusion protein includes a Cas12a polypeptide fused to a second intein polypeptide. In some embodiments, the first nucleic acid molecule of the present invention is present in a first expression cassette and / or vector, and the second nucleic acid molecule of the present invention is present in a second expression cassette and / or vector, wherein the first and second expression cassettes and / or vectors are separate and / or different from each other.
[0185] The complex and / or method of the present invention can use and / or include the Cas12a protein provided (for example, prepared) by two different fusion proteins of the present invention.For example, when two different fusion proteins contact (wherein each fusion protein includes Cas12a polypeptide), the intein polypeptides of the two fusion proteins can associate to form an intein (for example, an active complex), and the intein can excise the intein (for example, the first and second intein polypeptides), and optionally with the first Cas12a polypeptide and the second Cas12a polypeptide between the linker (for example, peptide linker) the first and second Cas12a polypeptides are fused together to form Cas12a protein. Two different fusion proteins can be provided by the composition and / or kit of the present invention and / or be present in the composition and / or kit of the present invention. In certain embodiments, the method of the present invention uses an editing system (for example, a CRISPR-Cas editing system), wherein the Cas12a protein of the present invention (for example, the Cas12a protein formed by the composition and / or method of the present invention) is a part of the editing system, and is provided (for example, prepared) by two different fusion proteins of the present invention. The editing system can be used for modifying target nucleic acid.
[0186] In some embodiments, the complex and / or method of the present invention can use and / or include a fusion protein (e.g., engineered protein) provided (e.g., prepared) by two different fusion proteins of the present invention. For example, when a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) and a first intein polypeptide contacts a second fusion protein comprising a Cas12a polypeptide and a second intein polypeptide, the intein polypeptides of the two fusion proteins can associate to form an intein (e.g., an active complex), and the intein can excise the intein (e.g., the first and second intein polypeptides), and optionally the polypeptide of interest and the Cas12a polypeptide are fused together to form a fusion protein (e.g., engineered protein) with a linker (e.g., a peptide linker) between the polypeptide of interest and the Cas12a polypeptide. Two different fusion proteins can be provided by the composition and / or kit of the present invention and / or be present in the composition and / or kit of the present invention. In some embodiments, the methods of the present invention use an editing system (e.g., a CRISPR-Cas editing system) in which a fusion protein of the present invention (e.g., an engineered protein (e.g., a templated editor) formed by the compositions and / or methods of the present invention) is part of the editing system and is provided (e.g., prepared) by two different fusion proteins of the present invention. The editing system can be used to modify a target nucleic acid.
[0187] According to some embodiments, a method for modifying a target nucleic acid is provided, the method comprising contacting the target nucleic acid with: a Cas12a protein prepared by a first fusion protein of the present invention and a second fusion protein of the present invention, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA). The Cas12a protein and the guide nucleic acid may form a complex or may be contained in a complex. The target nucleic acid may be present in a cell (e.g., a eukaryotic cell). In some embodiments, the target nucleic acid is present in a plant cell or a human cell. A method for modifying a target nucleic acid may comprise introducing a first nucleic acid molecule encoding the first fusion protein into the cell and introducing a second nucleic acid molecule encoding the second fusion protein into the cell, and expressing the first fusion protein and the second fusion protein in the cell. The first and second nucleic acid molecules may be present in the same composition so that the first and second nucleic acid molecules can be introduced together. In some embodiments, the first and second nucleic acid molecules are in different compositions so that the first and second nucleic acid molecules are introduced together in different, separate compositions or sequentially in any order. In certain embodiments, the first nucleic acid molecule and / or the second nucleic acid molecule are present in an expression cassette and / or a vector. In certain embodiments, the expression cassette and / or the vector are AAV vectors. In certain embodiments, the target nucleic acid is present in a cell, optionally in a cell of an organism (e.g., a plant, a human, etc.). In certain embodiments, the method for modifying the target nucleic acid is performed in vitro, in vivo, or in vitro.
[0188] In some embodiments, the method for modifying a target nucleic acid of the present invention comprises contacting the target nucleic acid with an engineered protein (e.g., a templated editor) prepared by a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA). The engineered protein and the guide nucleic acid may form a complex or may be contained in a complex. The target nucleic acid may be present in a cell (e.g., a eukaryotic cell). In some embodiments, the target nucleic acid is present in a plant cell or a human cell. A method for modifying a target nucleic acid may comprise introducing a first nucleic acid molecule encoding the first fusion protein into the cell and introducing a second nucleic acid molecule encoding the second fusion protein into the cell, and expressing the first fusion protein and the second fusion protein in the cell. The first and second nucleic acid molecules may be present in the same composition such that the first and second nucleic acid molecules can be introduced together. In some embodiments, the first and second nucleic acid molecules are in different compositions such that the first and second nucleic acid molecules are introduced together in different, separate compositions or sequentially in any order. In certain embodiments, the first nucleic acid molecule and / or the second nucleic acid molecule are present in an expression cassette and / or a vector. In certain embodiments, the expression cassette and / or the vector are AAV vectors. In certain embodiments, the target nucleic acid is present in a cell, optionally in a cell of an organism (e.g., a plant, a human, etc.). In certain embodiments, the method for modifying the target nucleic acid is performed in vitro, in vivo, or in vitro.
[0189] According to some embodiments, a method for modifying a target nucleic acid is provided, the method comprising: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell, wherein the first nucleic acid molecule encodes a first fusion protein, the first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein, the second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a protein comprising at least a portion of the first Cas12a polypeptide and at least a portion of the second Cas12a polypeptide and a guide nucleic acid (e.g., a guide RNA and / or an extended guide nucleic acid). In some embodiments, the guide nucleic acid and the protein comprising at least a portion of the first Cas12a polypeptide and at least a portion of the second Cas12a polypeptide form a complex or are contained in a complex. The method may include expressing the first fusion protein and the second fusion protein in the cell. In some embodiments, after the introduction step, the method comprises cutting (e.g., excising) the first intein polypeptide from the first fusion protein and cutting (e.g., excising) the second intein polypeptide from the second fusion protein. The method may also include before, during and / or after cutting, causing the first intein polypeptide to associate with the second intein polypeptide to form an intein. The cutting step may cut the first Cas12a polypeptide from the first fusion protein and the second Cas12a polypeptide from the second fusion protein, and / or the method may further include cutting the first Cas12a polypeptide from the first fusion protein and cutting the second Cas12a polypeptide from the second fusion protein. Before, during and / or after cutting the first and second intein polypeptides, the method may include fusing the first and second Cas12a polypeptides together (e.g., via a peptide bond between the first Cas12a polypeptide and the second Cas12a polypeptide) to form a Cas12a protein, wherein the protein in contact with the target nucleic acid in the cell is the Cas12a protein. In some embodiments, the intein fuses the first and second Cas12a polypeptides together. In some embodiments, cutting the first intein polypeptide from the first fusion protein and cutting the second intein polypeptide from the second fusion protein occur simultaneously with fusing the first and second Cas12a polypeptides together. The introducing step may comprise introducing into the cell a first expression cassette and / or vector comprising the first nucleic acid molecule and introducing into the cell a second expression cassette and / or vector comprising the second nucleic acid molecule. The first and / or second expression cassette and / or vector may comprise the guide nucleic acid, or the method may comprise introducing into the cell a third expression cassette and / or vector comprising the guide nucleic acid.In some embodiments, the first, second, and / or third expression cassette and / or vector is an AAV vector.
[0190] In some embodiments, the method of modifying a target nucleic acid of the present invention comprises: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell, wherein the first nucleic acid molecule encodes a first fusion protein, the first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein, the second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a protein (e.g., a templated editor) comprising at least a portion of the polypeptide of interest and at least a portion of the Cas12a polypeptide and a guide nucleic acid (e.g., a guide RNA). In some embodiments, the guide nucleic acid and the protein comprising at least a portion of the polypeptide of interest and at least a portion of the Cas12a polypeptide form a complex or are contained in a complex. The method may comprise expressing the first fusion protein and the second fusion protein in the cell. In some embodiments, after the introduction step, the method comprises cutting (e.g., excising) the first intein polypeptide from the first fusion protein and cutting (e.g., excising) the second intein polypeptide from the second fusion protein. The method may further comprise, before, during, and / or after the cutting, associating the first intein polypeptide with the second intein polypeptide to form an intein. The cutting step may cut the polypeptide of interest from the first fusion protein and cut the Cas12a polypeptide from the second fusion protein, and / or the method may further comprise cutting the polypeptide of interest from the first fusion protein and cutting the Cas12a polypeptide from the second fusion protein. Before, during and / or after cutting the first and second intein polypeptides, the method may comprise fusing the polypeptide of interest and the Cas12a polypeptide together (e.g., via a peptide bond between the polypeptide of interest and the Cas12a polypeptide) to form a fusion protein (optionally wherein the fusion protein is a templated editor), wherein the fusion protein contacts the target nucleic acid in the cell. In some embodiments, the intein fuses the polypeptide of interest and the Cas12a polypeptide together. In some embodiments, cutting the first intein polypeptide from the first fusion protein and cutting the second intein polypeptide from the second fusion protein occur simultaneously with fusing the polypeptide of interest and the Cas12a polypeptide together. The introducing step may comprise introducing a first expression cassette and / or a vector comprising the first nucleic acid molecule into the cell and introducing a second expression cassette and / or a vector comprising the second nucleic acid molecule into the cell. The first and / or second expression cassette and / or vector may comprise the guide nucleic acid, or the method may comprise introducing a third expression cassette and / or vector comprising the guide nucleic acid into the cell. In some embodiments, the first, second, and / or third expression cassette and / or vector is an AAV vector.
[0191] In some embodiments, the efficiency of the target nucleic acid modification method of the present invention is improved compared to the efficiency of the control method. An exemplary control method includes a method for contacting the target nucleic acid with a wild-type CRISPR-Cas effector protein, which is not fused together by protein splicing (e.g., from two different proteins) optionally using an intein. Another exemplary control method includes a method for contacting the target nucleic acid with a fusion protein (e.g., an engineered protein and / or a templated editor), which is not fused together by protein splicing (e.g., from two different proteins) optionally using an intein. Compared to the control method, the present method can produce increased inversion-deletions and / or increased modification (e.g., precise modification) levels. In some embodiments, the editing system used in the methods of the present invention is a Redraw editing system, such as described in U.S. Patent Application Publication No. 2021 / 0130835 and / or U.S. Patent Application Publication No. 2022 / 0145334, the contents of each of which are incorporated herein by reference in their entirety, but optionally wherein the CRISPR-Cas effector protein is a Cas12a protein fused together by two fusion proteins of the present invention via protein splicing, and / or wherein the templated editor is a protein fused together by two fusion proteins of the present invention via protein splicing.
[0192] According to an embodiment of the present invention, the Cas12a protein can be divided into two parts (e.g., two Cas12a polypeptides), and each part can be fused with a portion (e.g., a fragment) of a trans-splicing split intein (e.g., an intein polypeptide) to provide two different fusion proteins. Each of the two fusion proteins can be packaged separately in an AAV vector. For example, a nucleic acid molecule encoding a fusion portion comprising a Cas12a polypeptide and an intein polypeptide can be provided in an AAV vector. In certain embodiments, two fusion proteins (respective Cas12a polypeptides form full-length Cas12a proteins and respective intein polypeptides form full-length inteins) are packaged separately in separate AAV vectors, and when two AAV vectors are introduced (e.g., infected) into the same cell, both fusion proteins can be expressed, and the two Cas12a polypeptides can be fused (e.g., spliced) together in situ.
[0193] In some embodiments, the engineered protein (e.g., a base editor, a templated editor, etc.) can be divided into two parts (e.g., a part comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) and a second part comprising a Cas12a polypeptide), and each part can be fused to a portion (e.g., a fragment) of a trans-splicing split intein (e.g., an intein polypeptide) to provide two different fusion proteins. Each of the two fusion proteins can be packaged separately in an AAV vector. For example, a nucleic acid molecule encoding a fusion portion comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) and an intein polypeptide can be provided in an AAV vector. In some embodiments, two fusion proteins (which together form an engineered protein, and each intein polypeptide forms a full-length intein) are packaged separately in separate AAV vectors, and when the two AAV vectors are introduced (e.g., infected) into the same cell, both fusion proteins can be expressed, and the polypeptide of interest and the Cas12a polypeptide can be fused (e.g., spliced) together in situ.
[0194] In some embodiments, the editing system of the present invention utilizes the Redraw editing system. More details about the Redraw editing system can be found in U.S. Patent Application Publication No. 2021 / 0130835 and / or U.S. Patent Application Publication No. 2022 / 0145334, the contents of each of which are incorporated herein by reference in their entirety.
[0195] As described herein, the fusion proteins, nucleic acids, expression cassettes and / or vectors of the present invention can be codon optimized for expression in an organism. The organisms useful in the present invention can be any organism or cell thereof for which nucleic acid modification can be used. The organism can include, but is not limited to, any animal (e.g., a mammal), any plant, any fungus, any archaea, or any bacteria. In some embodiments, the organism can be a plant or a cell thereof. In some embodiments, the organism is an animal, such as a mammal (e.g., a human).
[0196] Target nucleic acid can be the genomic sequence from any organism (e.g., eukaryotic organisms, such as mammals or plants). In certain embodiments, target nucleic acid is the genomic sequence from a model organism, such as but not limited to Escherichia coli, immortalized human cell lines (e.g., HEK293, HeLa, etc.), Caenorhabditis elegans, Arabidopsis thaliana, and / or Drosophila melanogaster. In certain embodiments, target nucleic acid is the genomic sequence from a non-model organism. Exemplary non-model organisms include but are not limited to crop plants (e.g., fruit crop plants, vegetable crop plants, and / or field crop plants) and / or animals, such as humans, primates, and / or mice. In certain embodiments, non-model organisms are crop plants, such as corn, soybeans, wheat, or mustard. In certain embodiments, non-model organisms are animals for testing and / or using human therapeutic agents.
[0197] Nucleic acid constructs of the present invention can be used to modify the target nucleic acid of any plant or plant part. Fusion proteins of the present invention can be used to modify any plant (or plant grouping, for example, genus or higher classification), including angiosperms, gymnosperms, monocots, dicots, C3, C4, CAM plants, bryophytes, ferns and / or pseudoferns, microalgae and / or macroalgae. Plants and / or plant parts that can be used for the present invention can be plants and / or plant parts of any plant species / variety / cultivar. As used herein, term "plant part" includes but is not limited to embryo, pollen, ovule, seed, leaf, stem, bud, flower, branch, fruit, grain, ear, cob, shell, stalk, root, root tip, anther, plant cell (including complete plant cell in plant and / or plant part), plant protoplast, plant tissue, plant cell tissue culture, plant callus, plant clump etc. As used herein, "bud" refers to the above-ground part, including leaf and stem. In addition, as used herein, "plant cell" refers to the structural and physiological unit of a plant, which includes a cell wall and may also refer to a protoplast. A plant cell may be in the form of an isolated single cell, or may be a cultured cell, or may be a part of a higher-order organizational unit such as a plant tissue or a plant organ.
[0198] Non-limiting examples of plants that can be used in the present invention include lawn grasses (e.g., bluegrass, bentgrass, ryegrass, fescue), feather reed grass, bunch grass, miscanthus, reed, switchgrass, vegetable crops including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce (e.g., cabbage, leaf lettuce, romaine lettuce), yellow taro, melons (e.g., cantaloupe, watermelon, Crenshaw, honeydew, cantaloupe), Brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, kale, Chinese cabbage, bok choy), cardoons, carrots, sauerkraut, okra, onions, celery, parsley, chickpeas, parsnips, endive, peppers, potatoes, cucurbits (e.g., zucchini, cucumber, zucchini, pumpkin, honeydew, watermelon, cantaloupe), radish, dry bulb onions (e.g., zucchini, cucumber, zucchini, pumpkin, honeydew, watermelon, cantaloupe), onion), rutabaga, eggplant, salsify, endive, shallot, endive, garlic, spinach, shallots, pumpkin, leafy greens, beets (sugar beets and fodder beets), sweet potatoes, chard, horseradish, tomatoes, carrots, and spices; fruit crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quinces, figs, nuts (e.g., chestnuts, pecans, pistachios, hazelnuts, pistachios, peanuts, walnuts, macadamia nuts, almonds, etc.), citrus (e.g., clementines, kumquats, oranges, grapefruits, tangerines, mandarins, lemons, limes, etc.), blueberries, black raspberries, boysenberries, cranberries, currants, gooseberries, loganberries, raspberries, strawberries, blackberries, grapes (wine grapes and table grapes), avocados, bananas, Kiwi, persimmon, pomegranate, pineapple, tropical fruits, pome fruit, cantaloupe, mango, papaya and lychee, field crop plants such as clover, alfalfa, timothy, evening primrose, meadowsweet, corn / maize (fodder corn, sweet corn, popcorn corn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, triticale, sorghum, tobacco, kapok, legumes (beans (e.g., green beans and dry beans), lentils, peas, soybeans), oilseed plants (rapeseed, canola, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa bean, peanut, oil palm), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), Cannabis (e.g., hemp, sativa), Cannabis indica and Cannabis ruderalis), plants of the Lauraceae family (cinnamon, camphor) or plants such as coffee, sugar cane, tea and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants (e.g. roses, tulips, violets), and trees, such as forest trees (broadleaf trees and evergreen trees, such as conifers;For example, elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow) as well as shrubs and other seedlings. In some embodiments, the nucleic acid constructs of the present invention and / or expression cassettes and / or vectors encoding the same can be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry and / or cherry.
[0199] In some embodiments, the present invention provides cells (eg, plant cells, animal cells, bacterial cells, archaeal cells, etc.) comprising a polypeptide, polynucleotide, nucleic acid construct, expression cassette, or vector of the present invention.
[0200] The present invention further comprises one or more kits for carrying out the methods of the present invention. The kits of the present invention may comprise reagents, buffers and equipment for mixing, measuring, sorting, labeling, etc., as well as instructions for modifying target nucleic acids.
[0201] In certain embodiments, the present invention provides a kind of test kit, it comprises one or more fusion proteins of the present invention described herein, nucleic acid construct of the present invention and / or comprises its expression cassette and / or vector and / or cell and optional its instruction manual.In certain embodiments, test kit may further comprise CRISPR-Cas guide nucleic acid (corresponding to Cas12a albumen provided herein, it can be by polynucleotide encoding of the present invention) and / or comprises its expression cassette and / or vector and or cell.In certain embodiments, guide nucleic acid can be provided on the same expression cassette and / or vector with one or more nucleic acid constructs of the present invention.In certain embodiments, guide nucleic acid can be provided on expression cassette or vector separated from expression cassette or vector comprising one or more nucleic acid constructs of the present invention.
[0202] Thus, in some embodiments, a kit is provided comprising a nucleic acid construct comprising (a) one or more polynucleotides as provided herein, and (b) a promoter that drives expression of the one or more polynucleotides of (a). In some embodiments, the kit may further comprise a nucleic acid construct encoding a guide nucleic acid, wherein the construct comprises a cloning site for cloning a nucleic acid sequence identical or complementary to the target nucleic acid sequence into the backbone of the guide nucleic acid.
[0203] In some embodiments, the nucleic acid construct of the present invention can be an mRNA, which can encode one or more introns within the encoded polynucleotide. In some embodiments, the nucleic acid construct of the present invention and / or an expression cassette and / or vector comprising the same can further encode one or more selectable markers that can be used to identify transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.).
[0204] The polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems and / or cells of the invention may comprise all or a portion of the sequence of one or more of SEQ ID NOs: 1- 259. In some embodiments, the polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems and / or cells of the invention may comprise at least about 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more contiguous amino acids of the sequence of one or more of SEQ ID NOs: 1-259.
[0205] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claims of the present invention, but are intended to serve as examples of certain embodiments. Any variation of the exemplary methods that a skilled person would like to consider is intended to fall within the scope of the present invention.
[0206] Examples
[0207] Example 1: Validation of Split Intein Protein Reconstitution Using mCherry
[0208] The trans-splicing activities of wild-type Npu (SEQ ID NOs: 110 and 111, N-terminal and C-terminal portions, respectively) and NpuGEP (a mutant containing three residue mutations; SEQ ID NOs: 110 and 112, N-terminal and C-terminal portions, respectively) were evaluated in HEK293T cells. For each of wild-type (WT) Npu and NpuGEP, the N-terminal half of the mCherry protein was fused to the N-terminal portion of the Npu intein (NpuN), and the C-terminal half of mCherry was fused to the C-terminal portion of the Npu intein (NpuC). Figure 1 ). After transfection of plasmids encoding the two halves of mCherry, their splicing and remodeling were measured by flow cytometry 3 days later ( Figure 2 Robust fluorescence close to that of native mCherry was detected from the dual-plasmid setup ( Figure 2 Cells given only half of the mCherry fragment showed no fluorescence, demonstrating that both halves of the protein must be present to function. Figure 2 ). This demonstrates that the split intein is functional in human cells and allows for rapid protein splicing by transient plasmid transfection of each component.
[0209] Example 2: Split-intein Protein Reconstruction Using the Redraw Editor
[0210] Fusion proteins were prepared using the Redraw editor (RE2; SEQ ID NO: 113) and the NpuGEP system. Several sites within RE2 were used to introduce a cleavage site ( Figure 3 ). Specifically, RE2 includes a reverse transcriptase (RT) and a Cas12a sequence, and a fusion protein is prepared, including a portion of RE2, wherein the split occurs in the Cas12a sequence. Therefore, some fusion proteins include the N-terminal portion of the Cas12a sequence, which includes amino acid residues 1-175, 1-310, 1-406, 1-441, or 1-550 of the Cas12 sequence, and other fusion proteins include the remaining C-terminal portion of the Cas-12a sequence (e.g., amino acid residues 176, 311, 407, 442, or 551 to the end of Cas12a). Fusion proteins with sequences of SEQ ID NOs: 100-109 are prepared. In some cases, two amino acids (cysteine and alanine (CA) or cysteine and phenylalanine (CF)) are inserted between the C-terminal portion of the Npu intein (NpuC) and the C-terminal portion of RE2 (e.g., the C-terminal portion of Cas12a), which will leave a "CA" or "CF" scar within the reconstructed RE2 protein. The insertion of two amino acid "scar" residues (here "CA" or "CF") was confirmed in the mature mCherry protein, and it was confirmed that the scar does not affect the ability of the protein to fluoresce ( Figure 4 ).
[0211] Example 3: Intein-mediated Redraw Editor Reconstruction in HEK293T Cells
[0212] HEK293T cells were transfected with a plasmid encoding a fusion protein with the N-terminal component of RE2, a plasmid encoding a fusion protein with the C-terminal component of RE2, and a plasmid encoding a stagRNA or crRNA targeting an endogenous site as described in Example 2. After 3 days, high-throughput amplicon sequencing was performed to quantify Redraw activity at the target site. It was observed that several split RE2 forms were able to splice together in co-transfected HEK293T cells, thereby achieving Redraw activity. Robust inversion-deletion activity was observed ( Figures 5 to 6 ) and precise editing activity comparable to that of single polypeptide RE2 constructs ( Figure 7 ). In addition, it was observed that both halves of RE2 had to be expressed in cells to achieve inversion-deletion and Redraw activity.
[0213] These results indicate that the Redraw editor can be split into two halves (e.g., one half comprising a reverse transcriptase polypeptide (e.g., RT(5M)-NpuN) and the other half comprising a Cas12a polypeptide (e.g., NpuC-Cas12a-Brex27), or, for example, one half comprising a first Cas12a polypeptide and the other half comprising a second Cas12a polypeptide (e.g., RE2 split at 175)) and efficiently reconstituted in cells. This may allow the Redraw editor to be separately packaged into a viral delivery vehicle (e.g., an adeno-associated virus (AAV) vector) and co-infected to perform Redraw editing in cells targeted by the viral delivery vehicle.
[0214] The foregoing is illustrative of the present invention and should not be construed as limiting the present invention. The present invention is defined by the following claims, with equivalents of the claims to be included therein.
Claims
1. A fusion protein comprising Cas12a polypeptide and intein polypeptide fusion.
2. The fusion protein of claim 1, wherein the intein polypeptide is a first portion of a trans-spliced split intein, optionally wherein the intein polypeptide is one portion of two portions that together form the trans-spliced split intein.
3. The fusion protein according to claim 1 or 2, wherein the intein polypeptide is a portion of a Nostoc punctata (Npu) intein and / or a portion of a mutant Npu intein.
4. The fusion protein of any one of the preceding claims, wherein the intein polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 110-112.
5. The fusion protein of any one of the preceding claims, wherein the intein polypeptide is fused to the N-terminus of the Cas12a polypeptide, optionally wherein a linker (e.g., 1, 2, 3, 4 or more amino acid residues) is present between the intein polypeptide and the Cas12a polypeptide.
6. The fusion protein of any one of claims 1 to 4, wherein the intein polypeptide is fused to the C-terminus of the Cas12a polypeptide, optionally wherein a linker (e.g., 1, 2, 3, 4 or more amino acid residues) is present between the intein polypeptide and the Cas12a polypeptide.
7. The fusion protein of any one of the preceding claims, wherein the intein polypeptide is configured to be removed from the fusion protein and fused with another intein polypeptide in situ and / or in vivo.
8. The fusion protein of any one of the preceding claims, wherein the Cas12a polypeptide is a first portion of a Cas12a protein, optionally wherein the Cas12a polypeptide is one portion of two portions that together form the Cas12a protein.
9. The fusion protein of any one of the preceding claims, wherein the Cas12a polypeptide is the N-terminal portion of the Cas12a protein, and the intein polypeptide is the N-terminal portion of an intein, optionally wherein the intein polypeptide is fused to the C-terminus of the Cas12a polypeptide.
10. The fusion protein of any one of claims 1 to 8, wherein the Cas12a polypeptide is the C-terminal portion of the Cas12a protein, and the intein polypeptide is the C-terminal portion of an intein, optionally wherein the intein polypeptide is fused to the N-terminus of the Cas12a polypeptide.
11. The fusion protein of any one of the preceding claims, further comprising a reverse transcriptase fused to the Cas12a polypeptide and the intein polypeptide, optionally wherein the reverse transcriptase is fused to the N-terminus of the Cas12a polypeptide and the intein is fused to the C-terminus of the Cas12a polypeptide.
12. The fusion protein of any one of the preceding claims, wherein the Cas12a polypeptide comprises approximately 175, 310, 406, 441 or 550 consecutive amino acids of a Cas12a protein.
13. The fusion protein of any one of the preceding claims, wherein the Cas12a polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 150-159 and 175-184.
14. The fusion protein of any one of the preceding claims, wherein the fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 100-109 or 188.
15. A nucleic acid molecule encoding the fusion protein according to any one of claims 1 to 14.
16. An expression cassette or vector comprising the nucleic acid molecule according to claim 15.
17. A composite comprising: A Cas12a protein prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide; a guide nucleic acid (e.g., a guide RNA); and Optional deaminase.
18. A method for modifying a target nucleic acid, the method comprising: The target nucleic acid is modified by contacting the target nucleic acid with: A Cas12a protein prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a first Cas12a polypeptide fused to a first intein polypeptide, and the second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA), Optionally wherein the Cas12a protein and the guide nucleic acid form a complex or are contained in a complex.
19. The method according to claim 18, wherein the first fusion protein and / or the second fusion protein is a fusion protein according to any one of claims 1 to 14.
20. The method of claim 18 or 19, wherein the target nucleic acid is present in a cell (eg, a eukaryotic cell), optionally wherein the target nucleic acid is present in a plant cell or a human cell.
21. The method of claim 20, further comprising introducing a first nucleic acid molecule encoding the first fusion protein into the cell and introducing a second nucleic acid molecule encoding the second fusion protein into the cell, and producing the first fusion protein and the second fusion protein in the cell.
22. The method of claim 21, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is present in an expression cassette and / or vector, optionally wherein the expression cassette and / or vector is an adeno-associated viral vector.
23. A composition comprising: A first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and A second fusion protein comprises a second Cas12a polypeptide fused to a second intein polypeptide.
24. The composition of claim 23, wherein the first fusion protein and / or the second fusion protein is a fusion protein according to any one of claims 1 to 14, and wherein the first fusion protein is different from the second fusion protein.
25. A composition comprising: a first nucleic acid molecule encoding a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and A second nucleic acid molecule encodes a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide.
26. A kit comprising: a first nucleic acid molecule encoding a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide; and A second nucleic acid molecule encodes a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide.
27. The composition of claim 25 or the kit of claim 26, wherein a first expression cassette and / or vector comprises the first nucleic acid molecule and a second expression cassette and / or vector comprises the second nucleic acid molecule, optionally wherein the first expression cassette and / or vector is different from the second expression cassette and / or vector.
28. The composition or kit of any one of claims 25 to 27, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is the nucleic acid molecule of claim 15, and wherein the first nucleic acid molecule is different from the second nucleic acid molecule.
29. A method for modifying a target nucleic acid, the method comprising: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell comprising the target nucleic acid, wherein the first nucleic acid molecule encodes a first fusion protein comprising a first Cas12a polypeptide fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein comprising a second Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a guide nucleic acid (e.g., a guide RNA) and a protein comprising at least a portion of the first Cas12a polypeptide and at least a portion of the second Cas12a polypeptide, thereby modifying the target nucleic acid, Optionally wherein the protein and the guide nucleic acid form a complex or are comprised in a complex.
30. The method of claim 29, further comprising producing the first fusion protein and the second fusion protein in the cell.
31. The method of claim 29 or 30, wherein the first fusion protein and / or the second fusion protein is a fusion protein according to any one of claims 1 to 14, and the first fusion protein is different from the second fusion protein.
32. The method of any one of claims 29 to 31, further comprising cleaving the first intein polypeptide from the first fusion protein and cleaving the second intein polypeptide from the second fusion protein, optionally wherein before, during, and / or after cleavage, the first intein polypeptide and the second intein polypeptide associate to form an intein.
33. The method of claim 32, wherein the intein is a Nostoc punctata (Npu) intein, and / or a portion of a mutant Npu intein.
34. The method of claim 32 or 33, wherein the first intein polypeptide and / or the second polypeptide have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 110-112.
35. The method of any one of claims 29 to 35, further comprising cleaving the first Cas12a polypeptide from the first fusion protein and cleaving the second Cas12a polypeptide from the second fusion protein, optionally wherein before, during, and / or after cleavage, the first Cas12a polypeptide and the second Cas12a polypeptide are fused to form a Cas12a protein, wherein the protein that contacts the target nucleic acid in the cell is the Cas12a protein.
36. The method of claim 35, wherein the protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 38-60, 113-149, 192-195, and 196-259.
37. The method of any one of claims 29 to 36, wherein the first fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one of SEQ ID NOs: 100, 102, 104, 106, and 108, respectively, and the second fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one of SEQ ID NOs: 101, 103, 105, 107, and 109, respectively, wherein the first fusion protein is different from the second fusion protein.
38. The method of any one of claims 29 to 37, wherein introducing the first nucleic acid molecule and the second nucleic acid molecule into the cell comprises: introducing a first expression cassette and / or vector comprising the first nucleic acid molecule into the cell, and introducing a second expression cassette and / or vector comprising the second nucleic acid molecule into the cell, optionally wherein the first expression cassette and / or vector, and / or the second expression cassette and / or vector is an expression cassette and / or vector according to claim 16.
39. The method of any one of claims 29 to 38, further comprising introducing a third expression cassette and / or vector comprising the guide nucleic acid.
40. The method of claim 38 or 39, wherein the first expression cassette and / or vector, the second expression cassette and / or vector, and / or the third expression cassette and / or vector is an adeno-associated viral vector.
41. The method of any one of claims 18, 19, and 29 to 40, wherein the efficiency of the method in modifying the target nucleic acid is improved compared to the efficiency of a control method (e.g., a method comprising contacting the target nucleic acid with a wild-type CRISPR-Cas effector protein that is not fused together by protein splicing, optionally using an intein).
42. An engineered protein comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 100-109 and 187-188.
43. A nucleic acid molecule encoding the engineered protein according to claim 42.
44. A nucleic acid molecule comprising a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 187-188.
45. A fusion protein comprising a polypeptide of interest (eg, a reverse transcriptase polypeptide) fused to an intein polypeptide.
46. The fusion protein of claim 45, wherein the intein polypeptide is a first portion of a trans-spliced split intein, optionally wherein the intein polypeptide is one portion of two portions that together form the trans-spliced split intein.
47. The fusion protein of claim 45 or 46, wherein the intein polypeptide is a portion of a Nostoc punctata (Npu) intein and / or a portion of a mutant Npu intein.
48. The fusion protein of any one of claims 45 to 47, wherein the intein polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 110-112.
49. The fusion protein of any one of claims 45 to 48, wherein the intein polypeptide is fused to the N-terminus of the polypeptide of interest, optionally wherein a linker (e.g., 1, 2, 3, 4 or more amino acid residues) is present between the intein polypeptide and the polypeptide of interest.
50. The fusion protein of any one of claims 45 to 48, wherein the intein polypeptide is fused to the C-terminus of the polypeptide of interest, optionally wherein a linker (e.g., 1, 2, 3, 4 or more amino acid residues) is present between the intein polypeptide and the polypeptide of interest.
51. The fusion protein of any one of claims 45 to 50, wherein the intein polypeptide is configured to be removed from the fusion protein and fused with another intein polypeptide in situ and / or in vivo.
52. The fusion protein of any one of claims 45 to 51, wherein the polypeptide of interest has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 160-171.
53. The fusion protein of any one of claims 45 to 52, wherein the intein polypeptide is the N-terminal portion of an intein, optionally wherein the intein polypeptide is fused to the C-terminus of the polypeptide of interest.
54. The fusion protein of any one of claims 45 to 53, further comprising a linker, optionally wherein the linker is located between the polypeptide of interest and the intein polypeptide.
55. The fusion protein of any one of claims 1 to 14 or 45 to 54, further comprising a nuclear localization signal, optionally wherein the nuclear localization signal has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 172-174.
56. A nucleic acid molecule encoding the fusion protein according to any one of claims 45 to 55.
57. An expression cassette or vector comprising the nucleic acid molecule of claim 56.
58. A composite comprising: A fusion protein (e.g., a templated editor) prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and A guide nucleic acid (eg, a guide RNA).
59. A method for modifying a target nucleic acid, the method comprising: The target nucleic acid is modified by contacting the target nucleic acid with: A fusion protein prepared from a first fusion protein and a second fusion protein, wherein the first fusion protein comprises a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide; and a guide nucleic acid (e.g., a guide RNA), Optionally wherein the fusion protein and the guide nucleic acid form a complex or are comprised in a complex.
60. The method of claim 59, wherein the first fusion protein is the fusion protein of any one of claims 45 to 55, and / or the second fusion protein is the fusion protein of any one of claims 1 to 14.
61. The method of claim 59 or 60, wherein the target nucleic acid is present in a cell (eg, a eukaryotic cell), optionally wherein the target nucleic acid is present in a plant cell or a human cell.
62. The method of claim 61, further comprising introducing a first nucleic acid molecule encoding the first fusion protein into the cell and introducing a second nucleic acid molecule encoding the second fusion protein into the cell, and producing the first fusion protein and the second fusion protein in the cell.
63. The method of claim 62, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is present in an expression cassette and / or vector, optionally wherein the expression cassette and / or vector is an adeno-associated viral vector.
64. A composition comprising: a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and A second fusion protein comprises a Cas12a polypeptide fused to a second intein polypeptide.
65. The composition of claim 64, wherein the first fusion protein is a fusion protein according to any one of claims 45 to 55, and / or the second fusion protein is a fusion protein according to any one of claims 1 to 14.
66. A composition comprising: a first nucleic acid molecule encoding a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and A second nucleic acid molecule encodes a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide.
67. A kit comprising: a first nucleic acid molecule encoding a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide; and A second nucleic acid molecule encodes a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide.
68. The composition of claim 66 or the kit of claim 67, wherein a first expression cassette and / or vector comprises the first nucleic acid molecule and a second expression cassette and / or vector comprises the second nucleic acid molecule, optionally wherein the first expression cassette and / or vector is different from the second expression cassette and / or vector.
69. The composition or kit of any one of claims 66 to 68, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is a nucleic acid molecule according to claim 43 or 44.
70. A method for modifying a target nucleic acid, the method comprising: introducing a first nucleic acid molecule and a second nucleic acid molecule into a cell comprising the target nucleic acid, wherein the first nucleic acid molecule encodes a first fusion protein comprising a polypeptide of interest (e.g., a reverse transcriptase polypeptide) fused to a first intein polypeptide, and the second nucleic acid molecule encodes a second fusion protein comprising a Cas12a polypeptide fused to a second intein polypeptide; contacting the target nucleic acid in the cell with a guide nucleic acid (e.g., a guide RNA) and a protein comprising at least a portion of the polypeptide of interest and at least a portion of the Cas12a polypeptide, thereby modifying the target nucleic acid, Optionally wherein the protein and the guide nucleic acid form a complex or are comprised in a complex.
71. The method of claim 70, further comprising producing the first fusion protein and the second fusion protein in the cell.
72. The method of claim 70 or 71, wherein the first fusion protein is a fusion protein according to any one of claims 45 to 55, and / or the second fusion protein is a fusion protein according to any one of claims 1 to 14.
73. The method of any one of claims 70 to 72, further comprising cleaving the first intein polypeptide from the first fusion protein and cleaving the second intein polypeptide from the second fusion protein, optionally wherein before, during, and / or after cleavage, the first intein polypeptide and the second intein polypeptide associate to form an intein.
74. The method of claim 73, wherein the intein is a Nostoc punctata (Npu) intein, and / or a portion of a mutant Npu intein.
75. The method of claim 73 or 74, wherein the first intein polypeptide and / or the second polypeptide have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 110-112.
76. The method of any one of claims 70 to 85, further comprising cleaving the polypeptide of interest from the first fusion protein and cleaving the Cas12a polypeptide from the second fusion protein, optionally wherein before, during, and / or after cleavage, the polypeptide of interest and the Cas12a polypeptide are fused to form a fusion protein, wherein the protein that contacts the target nucleic acid in the cell is the fusion protein.
77. The method of claim 76, wherein the fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of SEQ ID NOs: 38-60, 113-149, 192-195, and 196-259.
78. The method of any one of claims 70 to 77, wherein the first fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 187, and the second fusion protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:
188.
79. The method of any one of claims 70 to 78, wherein introducing the first nucleic acid molecule and the second nucleic acid molecule into the cell comprises: introducing a first expression cassette and / or vector comprising the first nucleic acid molecule into the cell, and introducing a second expression cassette and / or vector comprising the second nucleic acid molecule into the cell.
80. The method of any one of claims 70 to 79, further comprising introducing a third expression cassette and / or vector comprising the guide nucleic acid.
81. The method of claim 79 or 89, wherein the first expression cassette and / or vector, the second expression cassette and / or vector, and / or the third expression cassette and / or vector is an adeno-associated viral vector.
82. The method of any one of claims 59 to 63 and 70 to 81, wherein the efficiency of the method in modifying the target nucleic acid is improved compared to the efficiency of a control method (e.g., a method comprising contacting the target nucleic acid with a templated editor that is not fused via protein splicing, optionally using an intein).
Citation Information
Patent Citations
Seed specific transcriptional regulation
EP0255378A2
Tissue-preferential promoters
EP0452269A2
Adenosine nucleobase editors and uses thereof
US10113163B2
Nucleobase editors and uses thereof
US10167457B2
Synthetic chloroplast transit peptides
US10421972B2
Cited By
AsCas12f segmented expression mediated gene editing method
CN121022930A