Engineered proteins and methods of use thereof
By developing a complex of polypeptide with RNA-dependent DNA polymerase activity, CRISPR-Cas effector protein and guide nucleic acid, the problem of reverse transcriptase generating cDNA at high temperatures was solved, and efficient nucleic acid editing effect was achieved.
Patent Information
- Application Number
- CN202380085367.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-22
AI Technical Summary
Existing reverse transcriptases are difficult to process mRNA with complex secondary structures when generating cDNA, and lack the ability to continuously synthesize under high temperature conditions, which cannot meet the needs of template editing applications such as REDRAW and PRIME.
A polypeptide is developed that is similar to a polypeptide within the range of specific sequence identity, with RNA-dependent DNA polymerase activity, capable of producing DNA from RNA at high temperatures, and forms a complex with the CRISPR-Cas effector protein and an extended guide nucleic acid for modifying the nucleic acid of interest.
It realizes efficient cDNA generation under high temperature conditions, improves the accuracy and efficiency of nucleic acid editing, and is suitable for the editing application of complex RNA templates.
Smart Images

Figure CN120359295A_ABST
Abstract
Description
[0001] Statement Regarding Electronic File of Sequence Listing
[0002] The disclosure of the XML-formatted sequence listing named 1499-109_ST26.xml, sized 668,784 bytes, generated on December 20, 2023, and submitted herewith is hereby incorporated by reference in its entirety.
[0003] Priority Claim
[0004] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 476,393, filed on December 21, 2022, under 35 U.S.C.§119(e), the entire content of which is incorporated herein by reference. Technical Field
[0005] The present invention relates to engineered proteins (e.g., engineered enzymes) and methods of using these proteins. The present invention also relates to compositions and systems for modifying or editing target nucleic acids. Background Art
[0006] Reverse transcriptase (RT) (also known as RNA-dependent DNA polymerase) is an enzyme that generates DNA (e.g., cDNA) from RNA. They typically polymerize from the free 3' end of a primer annealed to the RNA. Depending on the system, these primers can be DNA or RNA. For example, in retroviruses, the first-strand primer is usually transfer RNA. However, DNA is typically used as the primer for RT during second-strand synthesis in vitro and in lentiviruses.
[0007] RT is widely used in laboratories interested in studying mRNA to generate cDNA in vitro. In these cases, poly dT DNA is used as the reverse-strand primer, and random deoxyhexamers are typically used as the forward-strand primer. The most commonly used RTs for this function are Moloney murine leukemia virus reverse transcriptase (MMLV-RT or M-MuLV-RT), avian myeloblastosis virus reverse transcriptase (AMV-RT), and occasionally human immunodeficiency virus reverse transcriptase (HIV-RT). The most commonly used of these three is MMLV-RT.
[0008] Most RT engineering has focused on RTs with long continuous synthesis capabilities and high-temperature tolerance, such as MMLV-RT. The reasons for high-temperature optimization of RT are related to the process of generating cDNA. Large RNA secondary structures can block reverse transcription. High temperatures can be used to denature the RNA portion for which they wish to prepare cDNA. Since many mRNA molecules have more than a thousand nucleotides, continuous synthesis ability is also crucial for generating a good cDNA library. However, new RTs are needed, especially for template editing applications such as REDRAW and PRIME. Summary of the Invention
[0009] A first aspect of the invention is directed to a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96 - 133. In some embodiments, the polypeptide generates DNA (e.g., cDNA) from RNA. The polypeptide can generate (e.g., polymerize) the DNA from one end (e.g., the 3' end) of a DNA and / or RNA primer. In some embodiments, the polypeptide has activity as an RNA-dependent DNA polymerase.
[0010] Another aspect of the invention is directed to a nucleic acid molecule encoding a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96 - 133. In some embodiments, the nucleic acid molecule of the invention has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 134 - 171.
[0011] Another aspect of the invention is directed to a complex comprising: a CRISPR-Cas effector protein; a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96 - 133; and an extended guide nucleic acid.
[0012] Another aspect of the present invention is an expression cassette that is codon-optimized for expression in an organism, the expression cassette comprising: a polynucleotide encoding a promoter sequence; and a polynucleotide encoding a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96-133, wherein the polynucleotide encoding the polypeptide can be codon-optimized for expression in the organism. In some embodiments, the expression cassette of the present invention comprises: a polynucleotide encoding a promoter sequence; and a polynucleotide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 134-171, wherein the polynucleotide can be codon-optimized for expression in an organism.
[0013] Another aspect of the present invention is a method for modifying a target nucleic acid in a cell, the method comprising introducing the expression cassette and / or vector of the present invention into the cell so as to modify the target nucleic acid in the cell.
[0014] Another aspect of the present invention is a method for producing the polypeptide of the present invention, the method comprising: culturing a cell or cell population that has been transformed with a nucleic acid encoding the polypeptide; and isolating the polypeptide so as to produce the polypeptide.
[0015] Another aspect of the present invention is a method for performing reverse transcription, the method comprising contacting a target nucleic acid with a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96-133, wherein the polypeptide reverse transcribes the target nucleic acid to provide DNA (such as cDNA).
[0016] Another aspect of the present invention is a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with the following to modify the target nucleic acid: a CRISPR-Cas effector protein; a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96-133; and an extended guide nucleic acid.
[0017] It should be noted that aspects of the present invention described with respect to one embodiment may be incorporated into different embodiments, although not specifically described thereof. That is, all embodiments and / or features of any embodiment may be combined in any manner and / or combination. The applicant reserves the right to change any originally filed claims and / or to file any new claims accordingly, including the right to modify any originally filed claim to depend on and / or incorporate any feature of any other claim, although the claims were not originally presented in this manner. These and other objects and / or aspects of the present invention will be explained in detail in the specification set forth below. Those of ordinary skill in the art will understand further features, advantages and details of the present invention from reading the accompanying drawings and the detailed description of the preferred embodiments that follow, such description being illustrative only of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a graph showing the precise editing frequencies of putative reverse transcriptase proteins targeting three sites in the DMNT1-001 (PWsp143 and PWsp137) and DMNT1-002 (PWsp139) loci.
[0019] Figure 2 is a graph showing the inversion-deletion (indel) frequencies of putative reverse transcriptase proteins targeting three sites in the DMNT1-001 (PWsp143 and PWsp137) and DMNT1-002 (PWsp139) loci as measured using the stagRNA guide sequence.
[0020] Figure 3 is a graph showing the inversion-deletion frequencies of four different CRISPR RNA (crRNA) guide sequences targeting the DMNT1-001 (PWsp143 and PWsp137) or HEK2 (PWsp450 and PWsp451) loci.
[0021] Figure 4 is a graph showing the precise editing frequencies of putative reverse transcriptase proteins targeting the DMNT1-001 (PWsp143 and PWsp137) or RNF2 (PWsp453 and PWsp454) loci compared to a control reverse transcriptase.
[0022] Figure 5 is a graph showing the inversion-deletion frequencies of putative reverse transcriptase proteins targeting the DMNT1-001 (PWsp143 and PWsp137) or RNF2 (PWsp453 and PWsp454) loci as measured using the stagRNA guide sequence compared to a control reverse transcriptase. DETAILED DESCRIPTION
[0023] The present invention will now be described below with reference to the accompanying drawings and examples, in which embodiments of the present invention are shown. This specific embodiment is not intended to be a detailed catalog of all different ways of implementing the present invention or all features that can be added to the present invention. For example, features shown with respect to one embodiment can be incorporated into other embodiments, and features shown with respect to a particular embodiment can be deleted from that embodiment. Thus, the present invention contemplates that in some embodiments of the present invention, any feature or combination of features set forth herein can be excluded or omitted. Additionally, many variations and additions to the various embodiments proposed herein will be apparent to those skilled in the art in light of the present disclosure, and such many variations and additions do not depart from the present invention. Accordingly, the following description is intended to illustrate some particular embodiments of the present invention, rather than to exhaustively set forth all of its permutations, combinations, and variations.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The terms used in the description of the present invention are for the purpose of describing particular embodiments only and are not intended to limit the present invention.
[0025] All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentences and / or paragraphs in which the references are presented.
[0026] Unless the context otherwise indicates, the various features of the present invention specifically intended to be described herein can be used in any combination. Additionally, the present invention contemplates that in some embodiments of the present invention, any feature or combination of features set forth herein can be excluded or omitted. By way of illustration, if the specification states that a composition contains components A, B, and C, it is specifically intended that any one of A, B, or C, or any combination thereof, can be omitted and disclaimed, either singly or in any combination.
[0027] As used in the description of the present invention and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise.
[0028] Also as used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the absence of a combination when interpreted in the alternative form “or.”
[0029] As used herein, when the term "about" refers to a measurable value such as an amount or concentration, etc., it is intended to cover variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value as well as the specified value. For example, "about X", where X is a measurable value, is intended to include X and variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. The ranges of measurable values provided herein may include any other ranges and / or individual values therein.
[0030] As used herein, phrases such as "between X and Y" and "between about X and Y" shall be interpreted to include X and Y. As used herein, the phrase "between about X and Y" means "between about X and about Y", and the phrase "from about X to Y" means "from about X to about Y".
[0031] Unless otherwise indicated herein, the recitation of numerical ranges herein is merely intended to be a shorthand method for separately referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were recited individually herein. For example, if a range of 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.
[0032] As used herein, the terms "comprising", "including", and "having" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0033] As used herein, the transitional phrase "consisting essentially of" means that the scope of the claim should be interpreted to cover the specified materials or steps recited in the claim and materials or steps that do not materially affect the basic and novel characteristics of the claimed invention. Thus, when used in the claims of the present invention, the term "consisting essentially of" is not intended to be interpreted as equivalent to "comprising".
[0034] As used herein, the terms "increase (increase, increasing)", "enhance (enhance, enhancing)", and "improve (improve, improving)" (and their grammatical variants) describe an elevation of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 150%, 200%, 300%, 400%, 500% or more as compared to another measurable property or quantity (e.g., a control value).
[0035] As used herein, the terms "reduce", "reduced", "reducing", "reduction", "diminish", and "decrease" (and their grammatical variants) describe a reduction of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% compared to another measurable property or quantity (e.g., a control value). In some embodiments, the reduction may result in no or substantially no (i.e., a negligible amount, e.g., less than about 10% or even 5%) detectable activity or amount.
[0036] A "heterologous nucleotide sequence" or "recombinant nucleotide sequence" is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.
[0037] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence that occurs in nature or is endogenous. Thus, for example, a "native nucleic acid" is a nucleic acid that occurs naturally in or is endogenous to a reference organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.
[0038] As used herein, the terms "nucleic acid", "nucleic acid molecule", "nucleotide sequence", and "polynucleotide" refer to linear or branched, single-stranded or double-stranded RNA or DNA, or hybrids thereof. The terms also encompass RNA / DNA hybrids. When synthesizing dsRNA, less common bases such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, etc. may also be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind RNA with high affinity and are potent antisense inhibitors of gene expression. Other modifications may also be made, such as modifications to the phosphodiester backbone or the 2'-hydroxyl group in the RNA ribose group.
[0039] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or a sequence of these nucleotides from the 5' to 3' end of a nucleic acid molecule, and includes DNA or RNA molecules, including cDNA, DNA fragments or parts, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA and antisense RNA, any of which can be single-stranded or double-stranded. The terms "nucleotide sequence", "nucleic acid", "nucleic acid molecule", "nucleic acid construct", "recombinant nucleic acid", "oligonucleotide" and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. Nucleic acid molecules and / or nucleotide sequences provided herein are presented in the 5' to 3' direction from left to right herein, and are represented by the standard code for representing nucleotide characters specified in the U.S. sequence rules 37CFR§§1.821-1.825 and the World Intellectual Property Organization (WIPO) standard ST.25. As used herein, "5' district" can represent the polynucleotide district closest to the 5' end of a polynucleotide. Therefore, for example, the elements in the 5' district of a polynucleotide can be located at any position from the first nucleotide located at the 5' end of a polynucleotide to the nucleotide located in the middle of a polynucleotide. As used herein, "3' region" may refer to the region of a polynucleotide closest to the 3' end of a polynucleotide. Thus, for example, elements in the 3' region of a polynucleotide may be located anywhere from the first nucleotide at the 3' end of the polynucleotide to a nucleotide in the middle of the polynucleotide.
[0040] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotides (AMOs), etc. A gene may or may not be able to produce a functional protein or gene product. A gene may include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions).
[0041] A polynucleotide, gene or polypeptide can be "isolated", which means that the nucleic acid or polypeptide is substantially or essentially free of components that normally accompany the nucleic acid or polypeptide in nature. In some embodiments, these components include other cellular materials, culture medium from recombinant production and / or various chemicals used to chemically synthesize nucleic acids or polypeptides.
[0042] The term "mutation" refers to a point mutation (e.g., missense or nonsense, or insertion or deletion of a single base pair causing a frameshift), insertion, deletion and / or truncation. When the mutation is a substitution of one residue within an amino acid sequence by another residue, or a deletion or insertion of one or more residues within the sequence, the mutation is typically described by identifying the original residue, then identifying the position of the residue within the sequence, and the identity of the newly substituted residue.
[0043] As used herein, "non-natural mutation" refers to a mutation that is generated through human intervention and is different from a naturally occurring mutation that exists in the same gene or polypeptide (e.g., naturally occurring and not the result of a modification performed by a human).
[0044] As used herein, the terms "complementary" or "complementarity" refer to the natural binding of polynucleotides through base pairing under permissive salt and temperature conditions. For example, the sequence "A-G-T" (5' to 3') binds to the complementary sequence "T-C-A" (3' to 5'). Complementarity between two single-stranded molecules can be "partial", where only some of the nucleotides bind, or can be complete when there is full complementarity between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between the nucleic acid strands.
[0045] As used herein, "complementary" can mean 100% complementary to a reference nucleotide sequence, or can mean less than 100% complementary (e.g., "substantially complementary", e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% complementary, etc.).
[0046] A "portion" or "fragment" of a nucleotide sequence or polypeptide (including a domain) should be understood to refer, respectively, to a nucleotide sequence or polypeptide that is reduced in length (e.g., reduced by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more residues (e.g., nucleotides or peptides)) relative to a reference nucleotide sequence or polypeptide, and that consists of, consists essentially of, and / or is composed of a nucleotide sequence or polypeptide of contiguous residues that is the same as or nearly the same as (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) to the reference nucleotide sequence or polypeptide. In some embodiments, a portion of a reference nucleotide sequence or polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or more of the full-length reference nucleotide sequence or polypeptide. Such nucleic acid fragments or portions according to the invention may, where appropriate, be included in the larger polynucleotides of which they are a component. As an example, the repeat sequences of the guide nucleic acids of the invention may comprise a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type type V CRISPR Cas repeat sequence, e.g., a repeat sequence from a CRISPR Cas system, including but not limited to Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c, etc.). Similarly, a portion of a polypeptide may be included in the larger polypeptide of which it is a component.
[0047] Different nucleic acids or proteins having homology are referred to herein as "homologs". The term homolog includes homologous sequences from the same species and other species, as well as orthologous sequences from the same species and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences, expressed as a percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Thus, the compositions and methods of the present invention further comprise homologs of the nucleotide sequences and polypeptides of the present invention. As used herein, "orthologous" and "ortholog" refer to homologous nucleotide sequences and / or amino acid sequences in different species that are generated from a common ancestral gene during speciation. Homologs or orthologs of the nucleotide sequences of the present invention have significant sequence identity with the nucleotide sequences of the present invention (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100%).
[0048] As used herein, "sequence identity" refers to the degree of identity of two optimally aligned polynucleotide or polypeptide sequences over the entire component (e.g., nucleotide or amino acid) alignment window. "Identity" can be readily calculated by known methods, including but not limited to the methods described in Computational Molecular Biology (Lesk, A.M. ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D.W. ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A.M. and Griffin, H.G. eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G. ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J. eds.) Stockton Press, New York (1991).
[0049] As used herein, the term "percent sequence identity" or "identity percent" refers to the percentage of identical nucleotides in the linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "identity percent" may refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.
[0050] As used herein, in the context of two nucleic acid molecules, nucleotide sequences or protein sequences, the phrase "substantially identical" or "substantial identity" refers to two or more sequences or subsequences having at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% nucleotide or amino acid residue identity when measured using one of the following sequence comparison algorithms or when compared and aligned by visual inspection to obtain maximal correspondence. In some embodiments of the invention, substantial identity exists within a continuous nucleotide region of a nucleotide sequence of the invention, the length of which region is from about 10 nucleotides to about 20 nucleotides, from about 10 nucleotides to about 25 nucleotides, from about 10 nucleotides to about 30 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 30 nucleotides to about 40 nucleotides, from about 50 nucleotides to about 60 nucleotides, from about 70 nucleotides to about 80 nucleotides, from about 90 nucleotides to about 100 nucleotides, or more nucleotides, and any range therein, up to the full length of the sequence. In some embodiments, the nucleotide sequences may be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, substantially identical nucleotide or protein sequences perform substantially the same function as the nucleotide (or encoded protein sequence) to which they are substantially identical.
[0051] For sequence comparison, typically one sequence acts as a reference sequence to which one or more test sequences are compared. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, subsequence coordinates are specified if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence relative to the reference sequence based on the designated program parameters.
[0052] Optimal alignment of sequences for comparison windows is well known to those skilled in the art and can be performed using tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, etc., and optionally through computerized implementations of these algorithms, such as GAP, BESTFIT, FASTA, and TFASTA provided as part of Wisconsin (Accelrys Inc., San Diego, CA), as well as web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee. In some embodiments, the "optimal alignment" of two sequences (e.g., two polypeptide sequences) is the alignment with the highest score, optionally from alignments performed using tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, as Wisconsin GAP, BESTFIT, FASTA and TFASTA, Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSSWater, LALIGN, GGSEARCH2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE and / or T-Coffee, provided in part by (Accelrys Inc., San Diego, CA). In some embodiments, a "best alignment" of two sequences (e.g., two polypeptide sequences) is an alignment that provides the highest percentage of sequence identity, optionally allowing one or more gaps to be introduced into one or both sequences. The "identity score" of an alignment segment of a test sequence and a reference sequence is the number of identical components shared by the two aligned sequences divided by the total number of components in the reference sequence segment (e.g., the entire reference sequence or a smaller defined portion of the reference sequence). The percentage of sequence identity is expressed as the identity score multiplied by 100. The comparison of one or more polynucleotide sequences can be a comparison with a full-length polynucleotide sequence or a portion thereof, or a comparison with a longer polynucleotide sequence. For purposes of the present invention, the "percentage of identity" and / or the best alignment can be determined using the Basic Local Alignment Search Tool (BLAST) provided by the National Center for Biotechnology Information, e.g., BLASTX for translated nucleotide sequences, BLASTN for polynucleotide sequences, and BLASTP for polypeptide sequences.
[0053] Two nucleotide sequences are also considered to be substantially complementary when they hybridize to each other under stringent conditions. In some representative embodiments, two nucleotide sequences that are considered to be substantially complementary hybridize to each other under highly stringent conditions.
[0054] In the context of nucleic acid hybridization experiments (e.g., Southern and Northern hybridizations), "stringent hybridization conditions" and "stringent hybridization wash conditions" are sequence-dependent and vary under different environmental parameters. Extensive guidelines for nucleic acid hybridization are available in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). Generally, highly stringent hybridization and wash conditions are selected to be about 5 °C lower than the thermal melting temperature (T m ) of a particular sequence at a defined ionic strength and pH.
[0055] T m is the temperature (at a defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are selected to be equal to the T m。In Southern or Northern blotting, examples of stringent hybridization conditions for hybridizing complementary nucleotide sequences having more than 100 complementary residues to a filter membrane are hybridization overnight at 42 °C with 50% formamide and 1 mg heparin. Examples of highly stringent washing conditions are washing for about 15 minutes at 72 °C with 0.15 M NaCl. An example of stringent washing conditions is washing for 15 minutes at 65 °C with 0.2x SSC (for a description of SSC buffer, see Sambrook below). Typically, low stringency washing is performed prior to high stringency washing to remove background probe signal. An example of medium stringency washing for a duplex of more than 100 nucleotides is washing for 15 minutes at 45 °C with 1x SSC. An example of low stringency washing for a duplex of more than 100 nucleotides is washing for 15 minutes at 40 °C with 4 - 6x SSC. For short probes (e.g., about 10 to 50 nucleotides), stringent conditions generally involve a salt concentration of less than about 1.0 M Na ions, usually about 0.01 to 1.0 M Na ion concentration (or other salts) at pH 7.0 to 8.3, and a temperature of at least about 30 °C. The addition of destabilizing agents such as formamide can also achieve stringent conditions. Generally, in a particular hybridization assay, a signal-to-noise ratio that is 2-fold (or higher) than that observed for an unrelated probe indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are still substantially the same if the proteins encoded by the nucleotide sequences are substantially the same. This occurs, for example, when copies of a nucleotide sequence are made using the maximum codon degeneracy permitted by the genetic code.
[0056] The polynucleotides and / or recombinant nucleic acid constructs of the present invention can be codon-optimized for expression. In some embodiments, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the present invention (e.g., those that comprise / encode a polypeptide of the present invention (e.g., an engineered protein), a nucleic acid-binding polypeptide (e.g., a DNA-binding polypeptide, such as a sequence-specific DNA-binding domain from a polynucleotide-guided endonuclease, zinc finger nuclease, transcription activator-like effector nuclease (TALEN), Argonaute protein, and / or CRISPR-Cas effector protein), a guide nucleic acid, and / or a reverse transcriptase) can be codon-optimized for expression in an organism (e.g., an animal (e.g., a human), a plant, a fungus, an archaeon, or a bacterium). In some embodiments, the codon-optimized nucleic acid constructs, polynucleotides, expression cassettes, and / or vectors of the present invention have about 70% to about 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%) identity or higher identity to a reference nucleic acid construct, polynucleotide, expression cassette, and / or vector that has not been codon-optimized.
[0057] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the present invention can be operably associated with a variety of promoters and / or other regulatory elements for expression in an organism or its cells (e.g., a mammal and / or mammalian cells, a plant and / or plant cells, etc.). Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the present invention can further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter can be operably associated with an intron (e.g., the Ubi1 promoter and intron). In some embodiments, a promoter associated with an intron can be referred to as a "promoter region" (e.g., the Ubi1 promoter and intron).
[0058] As used herein, "operably linked" or "operably associated" in reference to polynucleotides means that the indicated elements are functionally related to each other and are generally also physically related. Thus, as used herein, the terms "operably linked" or "operably associated" refer to nucleotide sequences that are functionally associated on a single nucleic acid molecule. Accordingly, a first nucleotide sequence that is operably linked to a second nucleotide sequence refers to a situation where the first nucleotide sequence is in a functional relationship with the second nucleotide sequence. For example, if a promoter affects the transcription or expression of a nucleotide sequence, the promoter is operably associated with the nucleotide sequence. Those skilled in the art will understand that a control sequence (e.g., a promoter) need not be adjacent to the nucleotide sequence with which it is operably associated, so long as the control sequence functions to direct its expression. Thus, for example, there may be intervening untranslated but transcribed nucleic acid sequences between a promoter and a nucleotide sequence, and the promoter may still be considered to be "operably linked" to the nucleotide sequence.
[0059] As used herein, the terms "linked" or "fused" in reference to polypeptides refer to the covalent linking of one polypeptide to another polypeptide. A polypeptide may be linked or fused (e.g., at the N-terminus or C-terminus) to another polypeptide directly (e.g., via a peptide bond) or via a linker (e.g., a peptide linker). The direct fusion of two polypeptides (e.g., direct linking) means that an amino acid residue of the first polypeptide among the two polypeptides is covalently linked to an amino acid residue of the second polypeptide among the two polypeptides, without an intervening element being inserted between the two amino acid residues. By way of example, the first polypeptide and the second polypeptide may be directly linked via a peptide bond between the first polypeptide and the second polypeptide, without an intervening element (e.g., a linker) being inserted between the first polypeptide and the second polypeptide. The indirect fusion of two polypeptides (e.g., indirect linking) means that there is an intervening element (e.g., a linker, e.g., a peptide linker) between the two polypeptides, and the intervening element is covalently linked to each polypeptide, and optionally the intervening element may link one end of the first polypeptide among the two polypeptides to one end of the second polypeptide among the two polypeptides.
[0060] As used herein, "fusion protein" refers to two or more polypeptides that are covalently linked (e.g., directly or indirectly) such that they are transcribed and translated as a single unit, resulting in a single polypeptide that comprises the two or more polypeptides. In some embodiments, the two or more polypeptides may be naturally encoded by separate genes but are encoded by a single gene in the form of a fusion protein. The term "linker" is well recognized in the art and refers to a chemical group or molecule that links two molecules or moieties, such as two polypeptides or domains of a fusion protein, e.g., a CRISPR-Cas effector protein and a peptide tag and / or a polypeptide of interest. A linker may comprise a single linking molecule (e.g., a single amino acid) or may comprise more than one linking molecule. In some embodiments, a linker may be an organic molecule, group, polymer, or chemical moiety, such as a divalent organic moiety. In some embodiments, a linker may be an amino acid or may be a peptide. In some embodiments, the linker is a peptide (e.g., a peptide linker).
[0061] In some embodiments, the length of the peptide linker useful in the present invention can be from about 2 to about 100 or more amino acids, such as a length of about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids (e.g., a length of about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids (e.g., a length of about 105, 110, 115, 120, 130, 140, 150 or more amino acids)). In some embodiments, the peptide linker can be a GS linker. In some embodiments, the peptide linker is a GS linker having 2, 3 or 4 amino acid residues, optionally having 2 or 4 amino acid residues. In some embodiments, the peptide linker has one of the amino acid sequences of SEQ ID NO: 1-35 or 222. In some embodiments, the peptide linker may comprise CA, CF, (GGS) n, GS, SG, GSSG (SEQ ID NO:31), GSSGSS (SEQ ID NO:32), GSSGSSGS (SEQ ID NO:33), (GSS) n (SEQ ID NO:34), (GSS) n GS (SEQ ID NO:35), S(GGS) n (SEQ ID NO:25), SGGS (SEQ ID NO:26), (GSS) n The amino acid sequence of G (SEQ ID NO:222) or (GGGGGS)n (SEQ ID NO:27), where n is an integer from 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the peptide linker may comprise the following amino acid sequence: SGGSGGSGGS (SEQ ID NO:28). In some embodiments, the peptide linker may comprise the following amino acid sequence: SGSETPGTSESATPES (SEQ ID NO:29), also known as the XTEN linker. In some embodiments, the peptide linker may comprise the following amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO:30), also known as the GS-XTEN-GS linker.
[0062] As used herein, the terms "ligation" or "fusion" with respect to polynucleotides refer to the covalent joining of one polynucleotide to another polynucleotide. In some embodiments, two or more polynucleotide molecules can be joined by a linker, which can be an organic molecule, group, polymer, or chemical moiety, such as a divalent organic moiety. Polynucleotides can be joined or fused (at the 5' end or 3' end) to another polynucleotide either via a direct covalent linkage or via one or more linking nucleotides. In some embodiments, a polynucleotide motif of a certain structure can be inserted within another polynucleotide sequence (e.g., an extension of a hairpin structure in a guide RNA). In some embodiments, the linking nucleotides can be naturally occurring nucleotides. In some embodiments, the linking nucleotides can be non-naturally occurring nucleotides. The direct fusion of two polynucleotides (e.g., direct ligation) refers to the covalent joining of a nucleotide of a first polynucleotide among two polynucleotides to a nucleotide of a second polynucleotide among the two polynucleotides, without an intervening element between the two polynucleotides. By way of example, a first polynucleotide and a second polynucleotide can be directly joined via a phosphodiester bond between the first polynucleotide and the second polynucleotide, without an intervening element (e.g., a linker) between the first polynucleotide and the second polynucleotide. The indirect fusion of two polynucleotides (e.g., indirect ligation) refers to the presence of an intervening element (e.g., a linker, such as a polynucleotide linker) between two polynucleotides, and the intervening element is covalently joined to each polynucleotide, optionally joining one end of a first polynucleotide among the two polynucleotides to one end of a second polynucleotide among the two polynucleotides.
[0063] A "promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) operably associated with the promoter. The coding sequence controlled or regulated by the promoter can encode a polypeptide and / or functional RNA. Generally, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. Generally, a promoter is located 5' or upstream of the start point of the coding region of the corresponding coding sequence. A promoter can contain other elements that act as regulators of gene expression; for example, a promoter region. These include the TATA box consensus sequence and usually also the CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box can be replaced by the AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region can contain at least one intron (e.g., SEQ ID NO:36 or SEQ ID NO:37).
[0064] Promoters useful in the present invention can include, for example, constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred and / or tissue-specific promoters for preparing recombinant nucleic acid molecules, e.g., "synthetic nucleic acid constructs" or "protein-RNA complexes". These different types of promoters are known in the art.
[0065] The choice of promoter can vary depending on the temporal and spatial requirements of expression and also on the host cell to be transformed. Promoters from many different organisms are well known in the art. Based on the extensive knowledge available in the art, an appropriate promoter can be selected for a particular host organism of interest. Thus, for example, much is known about the promoters upstream of genes that are highly constitutively expressed in model organisms, and this knowledge can be readily obtained and implemented in other systems as appropriate.
[0066] In some embodiments, promoters functional in plants can be used with the constructs of the present invention. Non-limiting examples of promoters that can be used to drive expression in plants include the promoter of the RubisCo small subunit gene 1 (PrbcS1), the promoter of the actin gene (Pactin), the promoter of the nitrate reductase gene (Pnr), and the promoter of the repetitive carbonic anhydrase gene 1 (Pdca1) (see Walker et al., Plant Cell Rep. 23:727-735 (2005); Li et al., Gene 403:132-142 (2007); Li et al., Mol Biol Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, and Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al., Gene 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al., Mol Biol Rep. 37:1143-1154 (2010)). In some embodiments, the promoters useful in the present invention are RNA polymerase II (Pol II) promoters. In some embodiments, the U6 promoter or 7SL promoter from Zea mays can be used in the constructs of the present invention. In some embodiments, the U6c promoter and / or 7SL promoter from maize can be used to drive the expression of the guide nucleic acid. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from Glycine max can be used in the constructs of the present invention. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean can be used to drive the expression of the guide nucleic acid.
[0067] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Patent No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and U.S. Patent No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci. USA 84:6624-6629), the sucrose synthase promoter (Yang and Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and ubiquitin promoters. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants, e.g., sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Molec. Biol. 12:619-632), and Arabidopsis (Norris et al., 1993. Plant Molec. Biol. 21:895-906). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocot systems, and its sequence and vectors constructed for monocot transformation are disclosed in European Patent Publication EP0342926. The ubiquitin promoter is suitable for expressing the nucleotide sequences of the present invention in transgenic plants, particularly monocots. In addition, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231:150-160 (1991)) can be readily modified for expressing the nucleotide sequences of the present invention and is particularly suitable for monocot hosts.
[0068] In some embodiments, tissue-specific / tissue-preferred promoters can be used to express heterologous polynucleotides in plant cells. Tissue-specific or preferred expression patterns include, but are not limited to, green tissue-specific or preferred, root-specific or preferred, stem-specific or preferred, flower-specific or preferred, or pollen-specific or preferred. Promoters suitable for expression in green tissue include many promoters that regulate genes involved in photosynthesis, many of which have been cloned from monocotyledonous and dicotyledonous plants. In one embodiment, a promoter useful in the present invention is the maize PEPC promoter from the phosphoenolpyruvate carboxylase gene (Hudspeth and Grula, Plant Molec. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those associated with genes encoding seed storage proteins (such as β-conglycinin, cruciferin, napin, and phaseolin), zein, or oleosin (such as oleosin) or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad 2-1)) and other nucleic acids expressed during embryo development (such as Bce4, see, for example, Kridl et al. (1991) Seed Sci. Res. 1:209-219; and EP patent No. 255378). Tissue-specific or tissue-preferred promoters that can be used to express the nucleotide sequences of the present invention in plants, particularly maize, include, but are not limited to, those that direct expression in roots, pith, leaves, or pollen. These promoters are disclosed, for example, in WO 93 / 07278, the disclosure of which regarding promoters is incorporated herein by reference.Other non-limiting examples of tissue-specific or tissue-preferred promoters useful in the present invention are the cotton rubisco promoter disclosed in U.S. Patent No. 6,040,504; the rice sucrose synthase promoter disclosed in U.S. Patent No. 5,604,121; the root-specific promoter described by de Framond (FEBS 290:103-106 (1991); European Patent EP 0452269 to Ciba-Geigy); the stem-specific promoter described in U.S. Patent No. 5,625,136 (to Ciba-Geigy) that drives the expression of the maize trpA gene; the Cestrum yellow leaf curling virus promoter disclosed in WO 01 / 73087; and pollen-specific or pollen-preferred promoters, including but not limited to ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al., Plant Biotechnol. Reports 9(5):297-306 (2015)), ZmSTK2_USP from maize (Wang et al., Genome 60(6):485-495 (2017)), LAT52 and LAT59 from tomato (Twell et al., Development 109(3):705-713 (1990)), Zm13 (U.S. Patent No. 10,421,972), the PLA2-δ promoter from Arabidopsis thaliana (U.S. Patent No. 7,141,424) and / or the ZmC5 promoter from maize (International PCT Publication No. WO 1999 / 042587).
[0069] Additional examples of plant tissue-specific / tissue-preferred promoters include but are not limited to the root hair-specific cis-element (RHE) (K IMet al., The Plant Cell 18:2958-2970 (2006)), the root-specific promoter RCc3 (Jeong et al., Plant Physiol. 153:185-197 (2010)) and RB7 (U.S. Patent No. 5,459,252), the lectin promoter (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), the maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology, 37(8):1108-1115), the maize light-harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), the maize heat shock protein promoter (O'Dell et al. (1985) EMBO J. 5:451-458; and Rochester et al. (1986) EMBO J. 5:451-458), the pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase" pp. 29-39, in: Genetic Engineering of Plants (Hollaender ed., Plenum Press 1983; and Poulsen et al. (1986) Mol. Gen. Genet. 205:193-200), the Ti plasmid mannosine synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), the Ti plasmid nopaline synthase promoter (Langridge et al. (1989), supra), the petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7:1257-1263), the legume glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev. 3:1639-1646), the truncated CaMV 35S promoter (O'Dell et al. (1985) Nature 313:810-812), the potato glycoprotein promoter (Wenzler et al. (1989) Plant Mol.Biol. 13: 347-354), root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res. 18: 7449), zein promoter (Kriz et al. (1987) Mol. Gen. Genet. 207: 90-98; Langridge et al. (1983) Cell 34: 1015-1022; Reina et al. (1990) Nucleic Acids Res. 18: 6425; Reina et al. (1990) Nucleic Acids Res. 18: 7449; and Wandelt et al. (1989) Nucleic Acids Res. 17: 2354), globulin-1 promoter (Belanger et al. (1991) Genetics 129: 863-872), α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215: 431-440), PEPCase promoter (Hudspeth and Grula (1989) Plant Mol. Biol. 12: 579-589), R gene complex-related promoter (Chandler et al. (1989) Plant Cell 1: 1175-1183), and chalcone synthase promoter (Franken et al. (1991) EMBO J. 10: 2605-2612).
[0070] Useful for seed-specific expression is the legumin promoter (Czako et al. (1992) Mol. Gen. Genet. 235: 33-40; and the seed-specific promoter disclosed in U.S. Patent No. 5,625,136). Promoters useful for expression in mature leaves are those that switch at the onset of senescence, such as the SAG promoter from Arabidopsis (Gan et al. (1995) Science 270: 1986-1988).
[0071] In addition, promoters that are functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful for the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).
[0072] Additional regulatory elements useful for the present invention include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.
[0073] The introns useful in the present invention can be introns that are identified in plants, isolated from plants, and then inserted into an expression cassette for plant transformation. As understood by those skilled in the art, an intron can contain the sequences required for self-excision and be incorporated in-frame into a nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein-coding sequences in a nucleic acid construct, or an intron can be used within a protein-coding sequence, for example, to stabilize mRNA. If they are used within a protein-coding sequence, they are inserted "in-frame" and include the excision site. Introns can also be associated with a promoter to improve or alter expression. As an example, promoter / intron combinations useful in the present invention include, but are not limited to, the promoter / intron combination of the maize Ubi1 promoter and an intron.
[0074] Non-limiting examples of introns useful in the present invention include introns from the following genes: the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), ubiquitin gene (Ubi1), RuBisCO small subunit (rbcS) gene, RuBisCO large subunit (rbcL) gene, actin gene (e.g., actin-1 intron), pyruvate dehydrogenase kinase gene (pdk), nitrate reductase gene (nr), repetitive carbonic anhydrase gene 1 (Tdca1), psbA gene, atpA gene, or any combination thereof.
[0075] As used herein, an "editing system" refers to any site-specific (e.g., sequence-specific) nucleic acid editing system known now or developed later, which can introduce modifications (e.g., mutations) into nucleic acids in a target-specific manner. For example, an editing system (e.g., a site-specific and / or sequence-specific editing system) can include, but is not limited to, a CRISPR-Cas editing system, a meganuclease editing system, a zinc finger nuclease (ZFN) editing system, a transcription activator-like effector nuclease (TALEN) editing system, a base editing system, and / or a prime editing system, each of which can contain one or more polypeptides and / or one or more polynucleotides, which can modify a target nucleic acid (e.g., mutate the target nucleic acid) in a sequence-specific manner when they are present and / or expressed together (e.g., as a system) in a composition and / or a cell. In some embodiments, an editing system (e.g., a site-specific and / or sequence-specific editing system) can contain one or more polynucleotides and / or one or more polypeptides, including but not limited to a nucleic acid-binding polypeptide (e.g., a DNA-binding domain), a nuclease, another polypeptide, and / or a polynucleotide. In some embodiments, a CRISPR-Cas editing system containing the polypeptide of the present invention is provided and / or used.
[0076] In some embodiments, the editing system comprises one or more sequence-specific nucleic acid-binding polypeptides (e.g., DNA-binding domains), which can be from, for example, polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, transcription activator-like effector nucleases (TALENs), and / or Argonaute proteins. In some embodiments, the editing system comprises one or more cleavage polypeptides (e.g., nucleases), including but not limited to endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, and / or transcription activator-like effector nucleases (TALENs).
[0077] As used herein, "nucleic acid-binding polypeptide" refers to a polypeptide or domain that binds and / or is capable of binding to a nucleic acid (e.g., a target nucleic acid). A DNA-binding domain is an exemplary nucleic acid-binding polypeptide and can be a site- and / or sequence-specific nucleic acid-binding domain. In some embodiments, the nucleic acid-binding polypeptide can be a sequence-specific nucleic acid-binding polypeptide, such as but not limited to sequence-specific binding domains from, for example, polynucleotide-guided endonucleases, CRISPR-Cas effector proteins (e.g., CRISPR-Cas endonucleases), zinc finger nucleases, transcription activator-like effector nucleases (TALENs), and / or Argonaute proteins. In some embodiments, the nucleic acid-binding polypeptide comprises a cleavage domain (e.g., a nuclease domain), such as but not limited to endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases, zinc finger nucleases, and / or transcription activator-like effector nucleases (TALENs). In some embodiments, the nucleic acid-binding polypeptide associates with and / or is capable of associating with one or more nucleic acid molecules (e.g., forming a complex) (e.g., forming a complex with a guide nucleic acid as described herein), which can direct and / or guide the nucleic acid-binding polypeptide to a specific target nucleotide sequence (e.g., a locus in the genome) complementary to the one or more nucleic acid molecules (or a portion or region thereof), such that the nucleic acid-binding polypeptide binds to the nucleotide sequence at the specific target site. In some embodiments, the nucleic acid-binding polypeptide is a CRISPR-Cas effector protein as described herein.
[0078] In some embodiments, the editing system comprises or is a ribonucleoprotein, such as an assembled ribonucleoprotein complex (e.g., a ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and optionally a reverse transcriptase). In some embodiments, the ribonucleoproteins of the editing system can be assembled together (e.g., a pre-assembled ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and an optional reverse transcriptase), such as when contacting a target nucleic acid or when introduced into a cell (e.g., a mammalian cell or a plant cell) (e.g., when contacting the components of the ribonucleoprotein with the target nucleic acid and / or when introducing the components of the ribonucleoprotein into the cell). In some embodiments, the ribonucleoproteins of the editing system can assemble into a complex (e.g., a non-covalently bound complex) when a portion of the ribonucleoprotein contacts the target nucleic acid and / or can assemble after and / or during introduction into a plant cell. In some embodiments, the ribonucleoproteins of the editing system can contact the target nucleic acid and / or can be introduced into a plant cell. In some embodiments, the editing system can assemble (e.g., assemble into a non-covalently bound complex) when introduced into a plant cell. In some embodiments, the ribonucleoprotein can comprise a polypeptide of the invention, a guide nucleic acid, and an optional reverse transcriptase.
[0079] In some embodiments, the editing system of the invention comprises a reverse transcriptase (which can be a polypeptide of the invention), an extended guide nucleic acid, and a CRISPR-Cas effector protein (e.g., a type II CRISPR-Cas effector protein or a type V CRISPR-Cas effector protein). In some embodiments, the type V CRISPR-Cas effector protein or type II CRISPR-Cas effector protein, reverse transcriptase, and extended guide nucleic acid can form a complex or can be included in a complex capable of interacting with a target nucleic acid.
[0080] In some embodiments, the guide nucleic acid further comprises a reverse transcriptase template and may be referred to as an extended guide nucleic acid. As used herein, an "extended guide nucleic acid" is a guide nucleic acid as described herein that further comprises a reverse transcriptase template (RTT) and / or a primer binding site (PBS). In some embodiments, the extended guide nucleic acid is an engineered prime editing guide RNA (pegRNA). The extended guide nucleic acid can be a targeted allele guide RNA (tagRNA) or a stable targeted allele guide RNA (stagRNA). As used herein, "tagRNA" refers to an extended guide nucleic acid that comprises a PBS and an RTT and has target strand complementarity. As used herein, "stagRNA" refers to a tagRNA that comprises a stabilizing motif. The stabilizing motif can be present at the 3' and / or 5' end of the tagRNA. In some embodiments, the stabilizing motif is present at the 3' end of the tagRNA. Exemplary stabilizing motifs include, but are not limited to, recruitment motifs, RNA hairpins, pseudoknot sequences, and / or PP7 motifs (e.g., PP7 RNA hairpin sequences). In some embodiments, the stagRNA is a tagRNA that comprises a PP7 RNA hairpin sequence. In some embodiments, a CRISPR-Cas effector protein (e.g., a type II or type V CRISPR-Cas effector protein), a reverse transcriptase, and an extended guide nucleic acid can form a complex or be included in a complex.
[0081] In some embodiments, the extended guide nucleic acid comprises an extension portion that includes a primer binding site and a reverse transcriptase template, wherein the reverse transcriptase template comprises a modification (e.g., an edit) to be incorporated into the target nucleic acid. In some embodiments, the extended guide nucleic acid comprises a primer binding site and a modification (e.g., an edit) to be incorporated into the target nucleic acid (e.g., the reverse transcriptase template) at its 3' end. In some embodiments, the extended guide nucleic acid comprises: (1) a sequence that interacts (e.g., recruits and / or binds) with a CRISPR-Cas effector protein (e.g., a CRISPR-Cas nuclease), (2) a spacer that is substantially complementary to a first site on the target nucleic acid (e.g., a CRISPR RNA (crRNA) (first crRNA) and / or a tracrRNA+crRNA (sgRNA)), and (3) a nucleic acid-encoded repair template (e.g., an RNA-encoded repair template) that includes a primer binding site and an RNA template (e.g., that encodes a modification to be incorporated into the target nucleic acid). In some embodiments, the extended guide nucleic acid (e.g., the extended guide RNA) can comprise, from 5' to 3', a spacer sequence, a repeat sequence, and an extension portion that comprises, from 5' to 3', a reverse transcriptase template and a primer binding site. In some embodiments, the extended guide nucleic acid can comprise, from 5' to 3', a spacer sequence, a repeat sequence, and an extension portion that comprises, from 5' to 3', a primer binding site and a reverse transcriptase template. In some embodiments, the extended guide nucleic acid can comprise, from 5' to 3', an extension portion, a spacer sequence, and a repeat sequence, wherein the extension portion comprises, from 5' to 3', a reverse transcriptase template and a primer binding site. In some embodiments, the extended guide nucleic acid can comprise, from 5' to 3', an extension portion, a spacer sequence, and a repeat sequence, wherein the extension portion comprises, from 5' to 3', a primer binding site and a reverse transcriptase template.
[0082] According to some embodiments, an extended guide nucleic acid (e.g., pegRNA) can have a structure as described in Anzalone et al., Nature, December 2019; 576(7785):149-157 and / or be designed as described therein. In some embodiments, the extended guide nucleic acid comprises a primer binding site (PBS) optionally having a sequence of 1, 2, 3, 4, or 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides and a reverse transcriptase template (RT template) sequence optionally having a sequence of 65 or more nucleotides. In some embodiments, the PBS of the extended guide nucleic acid has a sequence of less than 15 nucleotides and has a sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 nucleotides (e.g., a sequence of length 5 or 6 nucleotides). The RT template sequence can be after the PBS sequence in the 5' to 3' direction. In some embodiments, the RT template sequence of the extended guide nucleic acid has a length greater than 65 nucleotides and can comprise about 50 or more nucleotides heterologous to the target site (e.g., the target nucleic acid), followed by about 15 or more nucleotides homologous to the target site. In some embodiments, the RT template sequence of the extended guide nucleic acid is after the PBS sequence and the RT template sequence has a length greater than 65 nucleotides, wherein the sequence comprises more than 50 nucleotides heterologous to the target site, followed by more than 15 nucleotides homologous to the target site. Thus, in some embodiments, when reverse transcribing the extended guide nucleic acid, the resulting newly transcribed sequence can hybridize to and / or be configured to hybridize to the non-nicked strand of the target site, which can thereby result in a heteroduplex DNA with a large insertion in the newly synthesized strand. After repairing this mismatched DNA, the resulting repaired DNA can contain a large insertion (e.g., greater than 50 nucleotides) of a DNA sequence. In some embodiments, the method can provide a large deletion (e.g., greater than 50 nucleotides) of a DNA sequence. In some embodiments, the PBS and 15 or more nucleotides homologous to the target site can comprise a homology arm, which can be used to optionally insert the heterology into the target site using homology-directed repair. The inserted DNA can correspond to any functional DNA sequence, such as but not limited to: a functional transgene; a DNA fragment inserted into a gene in such a way that when the gene is transcribed, it produces a hairpin RNA sufficient to silence a homologous gene by RNAi; and / or one or more functional site-specific recombination sites, such as lox, frt, which can then be used in subsequent Cre- or Flp-mediated site-specific recombination processes. In some embodiments, the extended guide nucleic acid may be too large to be produced in vivo using a PolIII promoter.In some embodiments, the extended guide nucleic acid can be operably associated with and / or generated using a PolII promoter. In some embodiments, the DNA-binding polypeptide (e.g., DNA-binding domain) and / or DNA endonuclease can have a structure as described in Anzalone et al., Nature, December 2019; 576(7785):149-157 and / or be designed as described therein. In some embodiments, the DNA-binding domain and / or DNA endonuclease is a CRISPR Cas polypeptide, such as Cas9 nickase, a nickase variant of another CRISPR Cas polypeptide, or Cas12a.
[0083] In some embodiments, two extended guide nucleic acids (e.g., pegRNAs) can be used (e.g., the editing system can comprise two extended guide nucleic acids). One or both of the two extended guide nucleic acids can have the structure as described in Anzalone et al., Nature, December 2019; 576(7785):149-157 and / or be designed as described therein. The two extended guide nucleic acids can comprise a primer binding site (PBS) optionally having a sequence of 1, 2, 3, 4, or 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides and a reverse transcriptase template (RT template) sequence optionally having a sequence of 50 or more nucleotides. The RT template sequences of the two extended guide nucleic acids can be complementary to each other, and thus the polynucleotides reverse transcribed from each RT template will be complementary to each other and will be able to hybridize to each other. This can allow the intermediates generated by this system and / or method to ligate two DNA segments together that were originally separated by more than 50 nucleotides, e.g., within one chromosome, or located on two separate DNA fragments, e.g., on two different chromosomes. After repair of the intermediate, depending on the design of the RT template, the resulting product can generate large deletions, large inversions, or interchromosomal recombinations. Since all of these products are generated by homology-directed repair, these products can be predictably precise and / or reproducible. In some embodiments, the DNA-binding polypeptide (e.g., DNA-binding domain) and / or DNA endonuclease can have the structure as described in Anzalone et al., Nature, December 2019; 576(7785):149-157 and / or be designed as described therein. In some embodiments, the DNA-binding polypeptide and / or DNA endonuclease is a CRISPR Cas polypeptide, such as Cas9 nickase, a similar nickase variant of another CRISPR Cas polypeptide, or Cas12a. In some embodiments, the DNA-binding polypeptide and / or DNA endonuclease is a Cas9 nuclease, a similar nuclease of another CRISPR Cas polypeptide, or Cas12a. Using a nuclease (instead of a nickase) can facilitate the intrachromosomal or interchromosomal recombination process by single-strand annealing of 3' overhangs of more than 50 nucleotides that will be generated at each of the two target sites corresponding to the two pegRNA target nucleic acids. In some embodiments, the editing system comprises one extended guide nucleic acid and one guide nucleic acid that does not contain a reverse transcriptase template and / or a primer binding site.
[0084] The extended guide nucleic acid can comprise a CRISPR nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid; and (b) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the RT template encodes a modification to be incorporated into the target nucleic acid. The CRISPR nucleic acid can be a type II or type V CRISPR nucleic acid, and / or the tracr nucleic acid can be any tracr corresponding to the appropriate type II or type V CRISPR nucleic acid. In some embodiments, the extended guide nucleic acid comprises: (i) a type V CRISPR nucleic acid or a type II CRISPR nucleic acid (e.g., type II or type V CRISPR RNA, type II or type V CRISPR DNA, type II or type V crRNA, or type II or type V crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid (e.g., type II or type V tracrRNA, type II or type V tracrDNA); and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the type V CRISPR nucleic acid or type II CRISPR nucleic acid comprises a spacer that binds to the first strand of the target nucleic acid (e.g., the target strand) (e.g., the spacer is complementary to a portion of contiguous nucleotides in the first strand of the target nucleic acid) and the primer binding site binds to the first strand (e.g., the target strand). In some embodiments, the extension portion can be fused to the 5' end or the 3' end of the CRISPR nucleic acid (e.g., from 5' to 3': repeat-spacer-extension portion or extension portion-repeat-spacer) and / or fused to the 5' end or the 3' end of the tracr nucleic acid. In some embodiments, the extension portion of the extended guide nucleic acid comprises an RT template (RTT) and a primer binding site (PBS) from 5' to 3' (e.g., 5'-crRNA-spacer-RTT (edit-encoding)-PBS-3'), or comprises a PBS and an RTT from 5' to 3' (e.g., 5'-crRNA-spacer-PBS-RTT (edit-encoding)-3'), depending on the position of the extension portion relative to the CRISPR nucleic acid of the extended guide nucleic acid. For example, in some embodiments, the extension portion of the extended guide nucleic acid can comprise an RT template and a primer binding site from 5' to 3' (when the extended guide sequence is linked to the 3' end of the CRISPR nucleic acid). In some embodiments, the extension portion of the extended guide sequence can comprise a primer binding site and an RT template from 5' to 3' (when the extended guide sequence is linked to the 5' end of the CRISPR nucleic acid).
[0085] In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the second strand of the target nucleic acid (e.g., the non-target, top strand). In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the first strand of the target nucleic acid (e.g., binds to the target strand, optionally the same strand that recruits the CRISPR-Cas effector protein, the bottom strand). In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the second strand of the target nucleic acid (e.g., the non-target strand, optionally the strand opposite the strand that recruits the CRISPR-Cas effector protein). In some embodiments, reverse transcriptase (RT) can be added to the target strand of the target nucleic acid (e.g., the strand complementary to the spacer of the CRISPR nucleic acid of the extended guide nucleic acid and that recruits the CRISPR-Cas effector protein). In some embodiments, reverse transcriptase (RT) is added to the non-target strand of the target nucleic acid (e.g., the strand complementary to the strand complementary to the spacer of the CRISPR nucleic acid and that recruits the CRISPR-Cas effector protein). Exemplary methods and editing systems are described in International Patent Publication No. WO 2021 / 092130, International Patent Publication No. WO 2022 / 098993, and U.S. Patent Application Publication Nos. 2021 / 0147862, 2021 / 0130835, 2021 / 0147862, and 2022 / 0145334, each of which is incorporated herein by reference in its entirety.
[0086] The RT template of the extended guide nucleic acid can encode one or more modifications (e.g., edits) to be incorporated into the target nucleic acid. The one or more modifications can be located anywhere within the RT template (e.g., where the location can be relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid). In some embodiments, the RT template has a modification at one or more positions from -1 to 23 (e.g., -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23) relative to the position of the protospacer adjacent motif (PAM) (e.g., TTTG) in the target nucleic acid. In some embodiments, the RT template can comprise a modification at nucleotide positions -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. In some embodiments, the RT template can comprise a modification at nucleotide positions 4 to nucleotide position 17 (e.g., positions 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the PAM of the target nucleic acid. In some embodiments, the RT template can comprise a modification at nucleotide positions 10 to nucleotide position 17 (e.g., positions 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the PAM of the target nucleic acid. In some embodiments, the RT template can comprise a modification at nucleotide positions 12 to nucleotide position 15 (e.g., positions 12, 13, 14, or 15) of the RT template relative to the position of the PAM of the target nucleic acid.
[0087] In some embodiments, the extension portion of the extended guide nucleic acid may comprise an RT template and a primer binding site from 5' to 3' (e.g., when the extension portion is linked to the 3' end of the CRISPR nucleic acid). In some embodiments, the extension portion of the extended guide nucleic acid may comprise a primer binding site and an RT template (RTT) from 5' to 3' (e.g., when the extension portion is linked to the 5' end of the CRISPR nucleic acid). In some embodiments, the length of the RT template may be from about 1 nucleotide to about 100 nucleotides (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 nucleotides or more, and any range or value therein), e.g., a length of from about 1 nucleotide to about 10 nucleotides, from about 1 nucleotide to about 15 nucleotides, from about 1 nucleotide to about 20 nucleotides, from about 1 nucleotide to about 25 nucleotides, from about 1 nucleotide to about 30 nucleotides, from about 1 nucleotide to about 35, 36, 37, 38, 39 or 40 nucleotides, from about 1 nucleotide to about 50 nucleotides, from about 5 nucleotides to about 15 nucleotides, from about 5 nucleotides to about 20 nucleotides, from about 5 nucleotides to about 25 nucleotides, from about 5 nucleotides to about 30 nucleotides, from about 5 nucleotides to about 35, 36, 37, 38, 39 or 40 nucleotides, from about 5 nucleotides to about 50 nucleotides, from about 8 nucleotides to about 15 nucleotides, from about 8 nucleotides to about 20 nucleotides, from about 8 nucleotides to about 25 nucleotides, from about 8 nucleotides to about 30 nucleotides, from about 8 nucleotides to about 35, 36, 37, 38, 39 or 40 nucleotides, from about 8 nucleotides to about 50 nucleotides, a length of from about 8 nucleotides to about 100 nucleotides, from about 10 nucleotides to about 15 nucleotides, from about 10 nucleotides to about 20 nucleotides, from about 10 nucleotides to about 25 nucleotides, from about 10 nucleotides to about 30 nucleotides, from about 10 nucleotides to about 36 nucleotides, from about 10 nucleotides to about 40 nucleotides, from about 10 nucleotides to about 50 nucleotides, from about 10 nucleotides to about 100 nucleotides, and any range or value therein.In some embodiments, the length of the RT template can be at least 8 nucleotides, optionally from about 8 nucleotides to about 100 nucleotides. In some embodiments, the length of the RT template is 36, 37, 38, 39, or 40 nucleotides or less (e.g., a length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides, or any value or range therein (e.g., a length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides to a length of about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides). In some embodiments, the length of the RT template can be at least 30 nucleotides, optionally from a length of about 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides to a length of about 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides, or any range or value therein. In some embodiments, the length of the RT template can be about 36, 40, 44, 47, 50, 52, 55, 63, 72, or 74 nucleotides. One or more modifications can be present within the length of the RTT. The one or more modifications can be located at any position within the RTT, where the position of the modification can be described relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid. In some embodiments, the RT template can contain a modification at nucleotide positions -1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. In some embodiments, the RT template can contain a modification at nucleotide positions 4 to nucleotide position 17 (e.g., positions 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid. In some embodiments, the RT template can contain a modification at nucleotide positions 10 to nucleotide position 17 (e.g., positions 10, 11, 12, 13, 14, 15, 16, or 17) of the RT template relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid.In some embodiments, the RT template may contain a modification located at nucleotide positions 12 to nucleotide position 15 (e.g., position 12, 13, 14, or 15) of the RT template relative to the position of the protospacer adjacent motif (PAM) of the target nucleic acid.
[0088] As used herein, the "primer binding site" (PBS) of the extension portion of an extended guide nucleic acid (e.g., tagRNA) refers to a continuous nucleotide sequence that can bind to a region or "primer" on a target nucleic acid, e.g., complementary to a target nucleic acid primer. As an example, a CRISPR Cas effector protein (e.g., type II or type V, e.g., Cas 9 or Cas12a) can nick / cut DNA, and the 3' end of the cut DNA serves as a primer for the PBS portion of the extended guide nucleic acid. The PBS can be complementary to the 3' end of the strand of the target nucleic acid and can bind to the target strand or the non-target strand and / or can be configured to bind to the target strand or the non-target strand. The primer binding site can be fully complementary to the primer, or it can be substantially complementary to the primer of the target nucleic acid (e.g., at least 70% complementary (e.g., 70% or about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or higher)). In some embodiments, the length of the primer binding site of the extension portion can be from about 1 nucleotide to about 100 nucleotides in length (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, or any value or range therein), or from about 4 nucleotides to about 85 nucleotides, from about 10 nucleotides to about 80 nucleotides, from about 20 nucleotides to about 80 nucleotides, from about 25 nucleotides to about 80 nucleotides, from about 30 nucleotides to about 80 nucleotides, from about 40 nucleotides to about 80 nucleotides, from about 45 nucleotides to about 80 nucleotides, from about 45 nucleotides to about 75 nucleotides, or from about 45 nucleotides to about 60 nucleotides, or any range or value therein.In some embodiments, the length of the PBS can be at least 30 nucleotides, optionally from about 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides to about 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides, or any range or value therebetween. In some embodiments, the length of the PBS can be about 8, 16, 24, 32, 40, 48, 56, 64, 72, or 80 nucleotides.
[0089] In some embodiments, the length of the RTT can be from about 35 nucleotides to about 75 nucleotides and the length of the PBS can be from about 30 nucleotides to about 80 nucleotides, optionally wherein the length of the PBS can be about 8, 16, 24, 32, 40, 48, 56, 64, 72, or 80 nucleotides and the length of the RTT can be about 36, 40, 44, 47, 50, 52, 55, 63, 72, or 74 nucleotides, or any combination of RTT length and / or PBS length.
[0090] In some embodiments, the extended portion of the extended guide nucleic acid can be fused to the 5'-end or 3'-end of a type II or type V CRISPR nucleic acid (e.g., 5' to 3': repeat-spacer-extension or extension-repeat-spacer) and / or fused to the 5'-end or 3'-end of the tracr nucleic acid. In some embodiments, when the extended portion is at the 5' of the crRNA, the type V CRISPR-Cas effector protein is modified to reduce (or eliminate) self-processing RNase activity.
[0091] In some embodiments, the extension portion of the extended guide nucleic acid can be linked to a type II or type V CRISPR nucleic acid and / or a type II or type V tracrRNA via a linker. In some embodiments, the linker has a length of from about 1 to about 100 or more nucleotides (e.g., a length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, and any range therein (e.g., a length of from about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, about 40 to about 100, about 50 to about 100, or from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 nucleotides to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides (e.g., a length of about 105, 110, 115, 120, 130, 140, 150 or more nucleotides).
[0092] A guide nucleic acid and / or an extended guide nucleic acid can comprise one or more recruitment motifs as described herein, which can be linked to the 5' end and / or the 3' end of the guide nucleic acid and / or which can be inserted into the guide nucleic acid (e.g., within a hairpin loop of the guide nucleic acid). In some embodiments, the extended guide nucleic acid can be linked to an RNA recruitment motif. The extended guide nucleic acid and / or the guide nucleic acid can be linked to one or two or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs; e.g., at least 10 to about 25 motifs), optionally wherein two or more of the RNA recruitment motifs can be the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, the RNA recruitment motif can be located at the 3' end of the extended portion of the extended guide nucleic acid (e.g., from 5' to 3', repeat-spacer-extension (RT template - primer binding site) - RNA recruitment motif). In some embodiments, the RNA recruitment motif can be embedded within the extended portion of the extended guide nucleic acid.
[0093] In some embodiments, the editing system comprises an extended guide nucleic acid linked to an RNA recruitment motif and a reverse transcriptase as a reverse transcriptase fusion protein, wherein the reverse transcriptase fusion protein comprises a reverse transcriptase polypeptide fused to an affinity polypeptide that binds the RNA recruitment motif, wherein the extended guide nucleic acid binds to a target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the reverse transcriptase fusion protein to the extended guide nucleic acid and bringing the target nucleic acid into contact with the reverse transcriptase. In some embodiments, two or more reverse transcriptase fusion proteins can be recruited to the extended guide nucleic acid, thereby bringing the target nucleic acid into contact with two or more reverse transcriptase fusion proteins.
[0094] As used herein, the terms "transgenic" or "transgenized" refer to at least one nucleic acid sequence that is obtained or synthetically produced from the genome of one organism and then introduced into a host cell (e.g., a plant cell) or an organism or tissue of interest and subsequently integrated into the genome of the host by "stable" transformation or transfection methods. In contrast, the terms "transient" transformation or transfection or introduction refer to the manner of introducing molecular tools, which include at least one nucleic acid (DNA, RNA, single-stranded or double-stranded or a mixture thereof) and / or at least one amino acid sequence, optionally containing suitable chemical or biological agents, to effect transfer into at least one compartment of interest in the cell, including but not limited to the cytoplasm, organelles (including the nucleus, mitochondria, vacuoles, chloroplasts) or membranes, thereby resulting in transcription and / or translation and / or association and / or activity of at least one introduced molecule without achieving stable integration or incorporation into the genome, and thus the corresponding at least one molecule introduced into the cell genome is not inherited. The term "transgene-free" refers to the state in which no transgenic is present or found in the genome of a host cell or tissue or organism of interest.
[0095] In some embodiments, the polynucleotides and / or nucleic acid constructs of the present invention may be an "expression cassette" or may be contained within an expression cassette. As used herein, an "expression cassette" refers to a recombinant nucleic acid molecule that contains, for example, a nucleic acid construct of the present invention (e.g., a polynucleotide encoding a polypeptide of the present invention (e.g., an engineered protein), a polynucleotide encoding a nuclease, a polynucleotide encoding a reverse transcriptase, a polynucleotide encoding a reverse transcriptase fusion protein, a polynucleotide encoding a peptide tag, a polynucleotide encoding an affinity polypeptide, a polynucleotide encoding a glycosylase, and / or a polynucleotide containing a guide nucleic acid), wherein the nucleic acid construct is operably associated with at least a control sequence (e.g., a promoter). Thus, some embodiments of the present invention provide expression cassettes that are designed to express, for example, the nucleic acid constructs of the present invention. When an expression cassette contains more than one polynucleotide, the polynucleotides may be operably linked to a single promoter that drives the expression of all the polynucleotides, or the polynucleotides may be operably linked to one or more separate promoters (e.g., three polynucleotides may be driven by one, two, or three promoters in any combination). Thus, for example, the polynucleotide encoding a polypeptide of the present invention, the polynucleotide encoding a CRISPR-Cas effector protein, and the polynucleotide containing a guide nucleic acid contained in the expression cassette may each be operably associated with a single promoter, or one or more of the polynucleotides may be operably associated with separate promoters (e.g., two or three promoters) in any combination, and the promoters may be the same or different from each other.
[0096] In some embodiments, the expression cassette containing the polynucleotide / nucleic acid construct of the present invention may be optimized for expression in an organism (e.g., an animal, a plant, a bacterium, etc.).
[0097] An expression cassette containing the nucleic acid construct of the present invention can be chimeric, meaning that at least one of its components is heterologous with respect to at least one of its other components (e.g., a promoter from a host organism is operably linked to a polynucleotide of interest to be expressed in the host organism, where the polynucleotide of interest is from an organism different from the host, or does not normally co-occur with the promoter). The expression cassette can also be naturally occurring, but has been obtained in a recombinant form useful for heterologous expression.
[0098] The expression cassette can optionally include transcriptional and / or translational termination regions (i.e., termination regions) and / or enhancer regions functional in the selected host cell. A variety of transcriptional terminators and enhancers are known in the art and can be used in the expression cassette. Transcriptional terminators are responsible for terminating transcription and correct mRNA polyadenylation. The termination region and / or enhancer region can be native to the transcriptional initiation region, can be native to a gene encoding a CRISPR-Cas effector protein or a gene encoding a polypeptide of the present invention, can be native to the host cell, or can be native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the CRISPR-Cas effector protein, the host cell, or any combination thereof).
[0099] The expression cassette of the present invention can also include a polynucleotide encoding a selectable marker, which can be used to select transformed host cells. As used herein, "selectable marker" refers to a polynucleotide sequence that, when expressed, confers a unique phenotype on the host cell expressing the marker, thereby allowing such transformed cells to be distinguished from cells without the marker. Such a polynucleotide sequence can encode a selectable or screenable marker, depending on whether the marker confers a trait that can be selected by chemical means (e.g., by using a selection agent (e.g., an antibiotic, etc.)), or whether the marker is a trait that can be identified simply by observation or testing (e.g., by screening (e.g., fluorescence)). Many examples of suitable selectable markers are known in the art and can be used in the expression cassette described herein.
[0100] The expression cassettes, nucleic acid molecules / constructs, and polynucleotide sequences described herein can be used in conjunction with vectors. The term "vector" refers to a composition for transferring, delivering, or introducing one or more nucleic acids into a cell. A vector can comprise a nucleic acid construct that contains one or more nucleotide sequences to be transferred, delivered, or introduced into the cell. Vectors for transforming host organisms are well known in the art. Non-limiting examples of general classes of vectors include viral vectors (e.g., adeno-associated virus (AAV) vectors), plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, phages, artificial chromosomes, minicircles, or Agrobacterium binary vectors in double-stranded or single-stranded linear or circular form, which may or may not be self-propagating or mobile. In some embodiments, viral vectors can include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated viral, or herpes simplex viral vectors. A vector as defined herein can transform prokaryotic or eukaryotic hosts by integration into the cell genome or by being episomal (e.g., an autonomously replicating plasmid with an origin of replication). Also included are shuttle vectors, which are DNA vectors capable of replicating natively or intentionally in two different host organisms, which can be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammals, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of an appropriate promoter or other regulatory element and is operably linked to an appropriate promoter or other regulatory element for transcription in the host cell. The vector can be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, this can contain its own promoter and / or other regulatory elements, while in the case of cDNA, this can be under the control of an appropriate promoter and / or other regulatory elements for expression in the host cell. Thus, the nucleic acid constructs and / or expression cassettes of the present invention can be contained in vectors described herein and known in the art.
[0101] As used herein, "contact", "contacting", "contacted" and their grammatical variants refer to bringing together the components of a desired reaction under conditions suitable for the desired reaction (e.g., transformation, transcriptional control, genome editing, creating nicks and / or cleavage). Thus, for example, a target nucleic acid can be contacted with a nucleic acid construct of the invention encoding, e.g., a nucleic acid-binding polypeptide (e.g., a DNA-binding domain, e.g., a sequence-specific DNA-binding protein (e.g., a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein (e.g., a CRISPR-Cas endonuclease), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein)), an extended guide nucleic acid, a polynucleotide encoding a polypeptide of the invention, and optionally a guide nucleic acid, under conditions for expressing the nucleic acid-binding polypeptide (e.g., a CRISPR-Cas effector protein), and the nucleic acid-binding polypeptide forms a complex with the extended guide nucleic acid, the complex hybridizes to the target nucleic acid, and optionally the polypeptide of the invention can be recruited to the nucleic acid-binding polypeptide (and thus to the target nucleic acid), or the polypeptide of the invention is fused to the nucleic acid-binding polypeptide, thereby modifying the target nucleic acid. In some embodiments, the polypeptide of the invention and the nucleic acid-binding polypeptide are optionally localized to the target nucleic acid by covalent and / or non-covalent interactions. Other methods for recruiting reverse transcriptase that utilize other protein-protein interactions, RNA-protein interactions, and / or chemical interactions can be used.
[0102] In some embodiments, a target nucleic acid can be contacted with a nucleic acid construct of the invention encoding a polypeptide of the invention, a CRISPR-Cas effector protein, and a guide nucleic acid under conditions for expressing the polypeptide, or a target nucleic acid can be contacted with a polypeptide of the invention, a CRISPR-Cas effector protein, and a guide nucleic acid. The CRISPR-Cas effector protein can form a complex with the guide nucleic acid, and the complex can hybridize to the target nucleic acid, and / or the polypeptide of the invention is recruited to the CRISPR-Cas effector protein (and thus to the target nucleic acid), and / or the polypeptide of the invention is fused to the CRISPR-Cas effector protein, thereby modifying the target nucleic acid. The polypeptide of the invention can optionally be localized at the target nucleic acid by covalent and / or non-covalent interactions.
[0103] As used herein, "modifying" or "modification" of a target nucleic acid includes editing (e.g., mutating), covalently modifying, exchanging / substituting nucleic acid / nucleobases, deleting, cleaving, and / or nicking the target nucleic acid to thereby provide a modified nucleic acid and / or altering the transcriptional control of the target nucleic acid to thereby provide a modified nucleic acid. In some embodiments, the modification can include insertions and / or deletions of any size and / or any type of single-base change (SNP). In some embodiments, the modification includes SNPs. In some embodiments, the modification includes exchanging and / or substituting one or more (e.g., 1, 2, 3, 4, 5 or more) nucleotides. In some embodiments, the length of the insertion or deletion can be from about 1 base to about 30,000 bases or longer (e.g., lengths of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500,4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, 20,500, 21,000, 21,500, 22,000, 22,500, 23,000, 23,500, 24,000, 24,500, 25,000, 25,500, 26,000, 26,500, 27,000, 27,500, 28,000, 28,500, 29,000, 29,500, 30,000 bases or longer, or any value or range therein). Thus, in some embodiments, the length of the insertion or deletion can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900,910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 bases, or any range or value therein; lengths of about 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bases to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases or longer, or any value or range therein; lengths of about 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases to about 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000,9,500 or 10,000 bases or longer, or any value or range therebetween; or a length of about 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690 or 700 bases to about 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500 or 5000 bases or longer, or any value or range therebetween. In some embodiments, the length of the insertion or deletion can be about 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500 or 10,000 bases to about 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, 20,500, 21,000, 21,500, 22,000, 22,500, 23,000, 23,500, 24,000, 24,500, 25,000, 25,500, 26,000, 26,500, 27,000, 27,500, 28,000, 28,500, 29,000, 29,500 or 30,000 bases or longer, or any value or range therebetween.
[0104] As used herein, "recruit", "recruiting", or "recruitment" refers to the attraction of one or more polypeptides or polynucleotides to another polypeptide or polynucleotide (e.g., a specific location in the genome) using protein-protein interactions, nucleic acid-protein interactions (e.g., RNA-protein interactions), and / or chemical interactions. Protein-protein interactions can include, but are not limited to, peptide tags (epitopes, polymerization epitopes) and corresponding affinity polypeptides, RNA recruitment motifs and corresponding affinity polypeptides, and / or chemical interactions. Exemplary chemical interactions that can be used with polypeptides and polynucleotides for recruitment purposes can include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin interaction; SNAP-tag (Hussain et al., Curr Pharm Des. 19(30):5437-42 (2013)); Halo-tag (Los et al., ACS Chem Biol. 3(6):373-82 (2008)); CLIP-tag (Gautier et al., Chemistry&Biology 15:128-136 (2008)); compound-induced DmrA-DmrC heterodimer (Tak et al., Nat Methods 14(12):1163-1166 (2017)); and / or bifunctional ligand approaches such as chemical-induced dimerization (Voβ et al., Curr Opin Chemical Biology 28:194-201 (2015)) (e.g., dihydrofolate reductase (DHFR) (Kopyteck et al., Cell Cehm Biol 7(5):313-321 (2000)). In some embodiments, the recruitment methods and / or systems of the present invention use protein-protein interactions and / or nucleic acid-protein interactions (e.g., RNA-protein interactions) to attract polypeptides or polynucleotides to another polypeptide or polynucleotide (e.g., to a specific location in the genome).
[0105] In the context of a polynucleotide or editing system of interest, "introducing", "introduce", "introduced" (and their grammatical variants) refer to presenting a nucleotide sequence of interest (e.g., a polynucleotide, nucleic acid construct, and / or guide nucleic acid) and / or an editing system (e.g., a polynucleotide, polypeptide, and / or ribonucleoprotein) to a host organism or a cell of the organism (e.g., a host cell; e.g., a plant cell) in a manner such that the nucleotide sequence and / or the editing system gain entry into the interior of the cell. Thus, for example, a nucleic acid construct of the invention encoding a polypeptide of the invention, a CRISPR-Cas effector protein of the invention, and / or a guide nucleic acid can be introduced into a cell of an organism to transform the cell with the polypeptide, CRISPR-Cas effector protein, guide nucleic acid, and reverse transcriptase. In some embodiments, a polypeptide and / or a guide nucleic acid of the invention can be introduced into a cell of an organism, optionally wherein the polypeptide and the guide nucleic acid can be comprised in a complex (e.g., a ribonucleoprotein). In some embodiments, the organism is a eukaryote (e.g., a mammal, e.g., a human).
[0106] As used herein, the term "transformation" refers to the introduction of a nucleic acid, polypeptide, and / or ribonucleoprotein (e.g., a heterologous nucleic acid, polypeptide, and / or ribonucleoprotein) into a cell. Transformation of a cell can be stable or transient. Thus, in some embodiments, a host cell or host organism can be stably transformed with a polynucleotide / nucleic acid molecule of the invention. In some embodiments, a host cell or host organism can be transiently transformed with a nucleic acid construct, polypeptide, and / or ribonucleoprotein of the invention.
[0107] In the context of a polynucleotide, polypeptide, and / or ribonucleoprotein, "transient transformation" means that the polynucleotide, polypeptide, and / or ribonucleoprotein are introduced into a cell and do not integrate into the genome of the cell.
[0108] In the context of a polynucleotide being introduced into a cell, "stable introduction" or "is stably introduced" means that the introduced polynucleotide is stably incorporated into the genome of the cell, such that the cell is stably transformed with the polynucleotide.
[0109] As used herein, "stable transformation" or "is stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the genome of the cell. Thus, the integrated nucleic acid molecule can be inherited by its progeny, and more specifically, can be inherited by successive generations of progeny. As used herein, "genome" includes the nuclear genome and the plastid genome, and thus includes the integration of a nucleic acid into, for example, a chloroplast or mitochondrial genome. As used herein, stable transformation can also refer to a transgene maintained episomally, e.g., as a minichromosome or a plasmid.
[0110] Transient transformation can be detected by, for example, enzyme-linked immunosorbent assay (ELISA) or Western blotting, which can detect the presence of peptides or polypeptides encoded by one or more transgenes introduced into an organism. Stable transformation of cells can be detected by, for example, Southern blot hybridization assay of the genomic DNA of the cells with a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of the transgene introduced into an organism (such as a mammal, a plant, etc.). Stable transformation of cells can also be detected by, for example, Northern blot hybridization assay of the RNA of the cells with a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of the transgene introduced into the host organism. Stable transformation of cells can also be detected by, for example, polymerase chain reaction (PCR) or other amplification reactions known in the art, which employ specific primer sequences that hybridize to the target sequence of the transgene, resulting in the amplification of the transgene sequence, and thus the transgene sequence can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.
[0111] Thus, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the present invention can be transiently expressed and / or they can be stably incorporated into the genome of a host organism. Thus, in some embodiments, the nucleic acid construct of the present invention can be transiently introduced into a cell together with a guide nucleic acid, and thus, no DNA is maintained in the cell.
[0112] The nucleic acid constructs, polypeptides, and / or ribonucleoproteins of the present invention can be introduced into cells by any method known to those skilled in the art. In some embodiments, the transformation methods include, but are not limited to: transformation via bacteria-mediated nucleic acid delivery (e.g., via Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide- and / or whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, particle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, osmosis, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanism that results in the introduction of nucleic acids into cells (such as plant cells or animal cells), including any combination thereof. In some embodiments of the present invention, the transformation of cells includes nuclear transformation. In some embodiments, the transformation of cells includes plastid transformation (e.g., chloroplast transformation). In some embodiments, the recombinant nucleic acid constructs of the present invention can be introduced into cells via conventional breeding techniques.
[0113] Procedures for transforming eukaryotes and prokaryotes are well-known and routine in the art and are described throughout the literature (see, e.g., Jiang et al., 2013. Nat. Biotechnol. 31:233-239; Ran et al., Nature Protocols 8:2281-2308 (2013)). General guidelines for various plant transformation methods known in the art include Miki et al. (“Procedures for Introducing Foreign DNA into Plants” in Methods in Plant Molecular Biology and Biotechnology, Glick, B.R. and Thompson, J.E., Eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell. Mol. Biol. Lett. 7:849-858 (2002)).
[0114] Thus, nucleotide sequences, polypeptides, and / or ribonucleoproteins can be introduced into a host organism or its cells in a variety of ways well-known in the art. The methods of the present invention do not depend on a particular method for introducing one or more nucleotide sequences, polypeptides, and / or ribonucleoproteins into an organism, but rather only on their entry into the interior of at least one cell of the organism. In cases where more than one nucleotide sequence, polypeptide, and / or ribonucleoprotein is to be introduced, they can be assembled as part of a single nucleic acid construct or into separate nucleic acid constructs and can be located on the same or different nucleic acid constructs. Thus, nucleotide sequences, polypeptides, and / or ribonucleoproteins can be introduced into the cells of interest in a single transformation event and / or in separate transformation events, or alternatively, in relevant cases, the nucleotide sequences can be incorporated into a plant, for example, as part of a breeding program. In some embodiments, the cells are eukaryotic cells (e.g., plant cells or mammalian cells such as human cells).
[0115] In some embodiments, the nucleic acid constructs of the present invention (e.g., polynucleotides encoding CRISPR-Cas effector proteins, polynucleotides encoding the polypeptides of the present invention, and / or guide nucleic acids, and / or expression cassettes and / or vectors comprising the same) may be operably linked to at least one regulatory sequence, optionally wherein the at least one regulatory sequence may be codon-optimized for expression in plants. In some embodiments, the at least one regulatory sequence may be, for example, a promoter, an operator, a terminator, or an enhancer. In some embodiments, the at least one regulatory sequence may be a promoter. In some embodiments, the regulatory sequence may be an intron. In some embodiments, the at least one regulatory sequence may be, for example, a promoter operably associated with an intron or a promoter region comprising an intron. In some embodiments, the at least one regulatory sequence may be, for example, the ubiquitin promoter and its associated intron (e.g., Medicago truncatula and / or maize and its associated intron). In some embodiments, the at least one regulatory sequence may be a terminator nucleotide sequence and / or an enhancer nucleotide sequence.
[0116] In some embodiments, the nucleic acid constructs of the present invention may be operably associated with a promoter region that comprises an intron, optionally wherein the promoter region may be the ubiquitin promoter and an intron (e.g., the alfalfa or maize ubiquitin promoter and intron, such as SEQ ID NO:36 or SEQ ID NO:37). In some embodiments, the nucleic acid constructs of the present invention operably associated with a promoter region comprising an intron may be codon-optimized for expression in plants.
[0117] In some embodiments, the nucleic acid constructs of the present invention may encode one or more (e.g., 1, 2, 3, 4 or more) polypeptides of interest. The one or more polypeptides of interest may be codon-optimized for expression in eukaryotes (e.g., humans or plants). In some embodiments, the polypeptides of the present invention may comprise one or more (e.g., 1, 2, 3, 4 or more) polypeptides of interest.
[0118] Polypeptides of interest useful in the present invention may include, but are not limited to, polypeptides or protein domains having the following: deaminase activity, nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil-DNA glycosylase inhibitor (UGI)), reverse transcriptase, peptide tags (e.g., GCN4 peptide tag), demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fok1), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, nuclear localization sequence or activity, affinity polypeptide, peptide tag, and / or photolyase activity. In some embodiments, the polypeptide of interest is a Fok1 nuclease or a uracil-DNA glycosylase inhibitor. When encoded in a nucleic acid (polynucleotide, expression cassette, and / or vector), the encoded polypeptide or protein domain may be codon optimized for expression in an organism. In some embodiments, the polypeptide of interest may be linked to a CRISPR-Cas effector protein domain to provide a CRISPR-Cas fusion protein. In some embodiments, a CRISPR-Cas fusion protein comprising a CRISPR-Cas effector protein domain linked to a peptide tag may also be linked to a polypeptide of interest (e.g., the CRISPR-Cas effector protein domain may be linked, for example, simultaneously to a peptide tag (or affinity polypeptide) and, for example, a polypeptide of interest).
[0119] In some embodiments, the editing system of the present invention comprises a CRISPR-Cas effector protein. As used herein, a "CRISPR-Cas effector protein" is a protein or polypeptide that cleaves, incises, or nicks nucleic acids; binds nucleic acids (e.g., target nucleic acids and / or guide nucleic acids); and / or identifies, recognizes, or binds to a guide nucleic acid as defined herein. In some embodiments, the CRISPR-Cas effector protein can be an enzyme (e.g., nuclease, endonuclease, nickase, etc.) and / or can act as an enzyme. In some embodiments, the CRISPR-Cas effector protein refers to a CRISPR-Cas nuclease. In some embodiments, the CRISPR-Cas effector protein comprises nuclease activity and / or nickase activity, comprises a nuclease domain whose nuclease activity and / or nickase activity has been reduced or eliminated, comprises single-stranded DNA cleavage activity (ssDNAse activity) or has ssDNAse activity that has been reduced or eliminated, and / or comprises self-processing RNase activity or has self-processing RNase activity that has been reduced or eliminated. The CRISPR-Cas effector protein can bind to a target nucleic acid. The CRISPR-Cas effector protein can be a type I, type II, type III, type IV, type V, or type VI CRISPR-Cas effector protein. In some embodiments, the CRISPR-Cas effector protein can be from a type I CRISPR-Cas system, a type II CRISPR-Cas system, a type III CRISPR-Cas system, a type IV CRISPR-Cas system, a type V CRISPR-Cas system, or a type VI CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein can be from a type II CRISPR-Cas system or a type V CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein can be a type II CRISPR-Cas effector protein, e.g., a Cas9 effector protein. In some embodiments, the CRISPR-Cas effector protein can be a type V CRISPR-Cas effector protein, e.g., a Cas12 effector protein. In some embodiments, the CRISPR-Cas effector protein can be Cas12a and optionally can have the amino acid sequence of any one of SEQ ID NOs: 38-60, 183, and 212-221 and / or the nucleotide sequence of any one of SEQ ID NOs: 61-63. In some embodiments, the CRISPR-Cas effector protein can be active Cas12a and optionally can have the amino acid sequence of SEQ ID NO: 46 or 55. In some embodiments, the CRISPR-Cas effector protein can be inactive (i.e., dead) Cas12a and optionally can have the amino acid sequence of SEQ ID NO: 38.In some embodiments, the CRISPR-Cas effector protein can be Cas12b and optionally can have the amino acid sequence of SEQ ID NO:64.
[0120] Exemplary CRISPR-Cas effector proteins include, but are not limited to, Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4 (dinG), and / or Csf5 nucleases, optionally wherein the CRISPR-Cas effector protein can be a Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c effector protein.
[0121] In some embodiments, the CRISPR-Cas effector protein useful in the present invention can contain mutations in its nuclease active site and / or nuclease domain (e.g., RuvC, HNH, such as the RuvC site of the Cas12a nuclease domain; such as the RuvC site and / or HNH site of the Cas9 nuclease domain). A CRISPR-Cas effector protein having a mutation in its nuclease active site and / or nuclease domain that results in the protein not having nuclease activity is generally referred to as "inactive" or "dead", such as dCas9. In some embodiments, a CRISPR-Cas effector protein having a mutation in its nuclease active site and / or nuclease domain may have impaired activity or reduced activity (e.g., nickase activity) compared to the same CRISPR-Cas effector protein without the mutation.
[0122] The CRISPR Cas9 effector protein or Cas9 useful in the present invention can be any known or later identified Cas9 nuclease. In some embodiments, Cas9 can be a protein from, for example, Streptococcus spp. (e.g., S. pyogenes, S. thermophilus), Lactobacillus spp., Bifidobacterium spp., Kandleria spp., Leuconostoc spp., Oenococcus spp., Pediococcus spp., Weissella spp., and / or Olsenella spp. In some embodiments, the CRISPR-Cas effector protein can be Cas9 and optionally can have a nucleotide sequence of any one of SEQ ID NOs: 65-79 and / or an amino acid sequence of any one of SEQ ID NOs: 80-81.
[0123] In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus pyogenes and / or can recognize the PAM sequence motifs NGG, NAG, NGA (Mali et al., Science 2013; 339(6121):823-826). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus thermophilus and / or can recognize the PAM sequence motifs NGGNG and / or NNAGAAW (W = A or T) (see, e.g., Horvath et al., Science, 2010; 327(5962):167-170, and Deveau et al., J Bacteriol 2008; 190(4):1390-1400). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus mutans and / or can recognize the PAM sequence motifs NGG and / or NAAR (R = A or G) (see, e.g., Deveau et al., JBACTERIOL 2008; 190(4):1390-1400). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Staphylococcus aureus and / or can recognize the PAM sequence motif NNGRR (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Staphylococcus aureus and / or can recognize the PAM sequence motif N GRRT (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Staphylococcus aureus and / or can recognize the PAM sequence motif N GRRV (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Neisseria meningitidis and / or can recognize the PAM sequence motifs NGATT or N GCTT (R = A or G, V = A, G or C) (see, e.g., Hou et al., PNAS 2013, 1-6). In the foregoing embodiments in this paragraph, N in the PAM sequence motif can be any nucleotide residue, e.g., any one of A, G, C or T. In some embodiments, the CRISPR-Cas effector protein can be Cas13a derived from Leptotrichia shahii and / or can recognize a protospacer flanking sequence (PFS) (or RNA PAM (rPAM)) sequence motif of a single 3'A, U or C, which can be located within the target nucleic acid).
[0124] The type V CRISPR-Cas effector proteins useful in embodiments of the present invention can be any type V CRISPR-Cas nuclease. Exemplary type V CRISPR-Cas effector proteins include, but are not limited to, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c nucleases. In some embodiments, the type V CRISPR-Cas effector protein can be Cas12a. In some embodiments, the type V CRISPR-Cas effector protein can be a nickase, optionally, a Cas12a nickase. In some embodiments, the type V CRISPR-Cas effector protein can be Cas12b (e.g., SEQ ID NO: 64).
[0125] In some embodiments, the CRISPR-Cas effector protein can be a type V clustered regularly interspaced short palindromic repeats (CRISPR)-Cas nuclease. Cas12a differs from the better-known type II CRISPR Cas9 nuclease in several respects. For example, Cas9 recognizes a guanine-rich protospacer adjacent motif (PAM) (3'-NGG) located 3' of its guide RNA (gRNA, sgRNA, crRNA, crDNA, CRISPR array) binding site (protospacer, target nucleic acid, target DNA), while Cas12a recognizes a thymine-rich PAM (5'-TTN, 5'-TTTN) located 5' of the target nucleic acid. In fact, the orientation of Cas9 and Cas12a binding their guide RNAs is almost opposite with respect to their N and C termini. In addition, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA), rather than the dual guide RNAs (sgRNA (e.g., crRNA and tracrRNA)) found in the native Cas9 system, and Cas12a processes its own gRNA. Further, Cas12a nuclease activity produces staggered DNA double-strand breaks, rather than blunt ends produced by Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cleave both DNA strands, while Cas9 utilizes an HNH domain and an RuvC domain to cleave.
[0126] The CRISPR Cas12a effector protein useful in the present invention can be any known or later identified Cas12a (previously known as Cpf1) (see, for example, U.S. Patent No. 9,790,490, the disclosure of which regarding Cpf1 (Cas12a) sequences is incorporated by reference). The term "Cas12a" refers to an RNA-guided protein that may have nuclease activity, the protein comprising a guide nucleic acid-binding domain and an active, inactive, or partially active DNA cleavage domain, whereby the RNA-guided nuclease activity of Cas12a can be active, inactive, or partially active, respectively. In some embodiments, the Cas12a useful in the present invention may comprise a mutation in a nuclease active site (e.g., the RuvC site of the Cas12a domain). Cas12a having a mutation in its nuclease domain and / or nuclease active site and thus no longer comprising nuclease activity is generally referred to as dead Cas12a (e.g., dCas12a). In some embodiments, Cas12a having a mutation in its nuclease domain and / or nuclease active site may have impaired activity, e.g., may have reduced nickase activity. In some embodiments, Cas12a can have an amino acid sequence of any one of SEQ ID NOs: 38-59, 183, or 212-221.
[0127] In some embodiments, the CRISPR-Cas effector protein can be optimized for expression in an organism, such as an animal (e.g., a mammal, such as a human), a plant, a fungus, an archaeon, or a bacterium. In some embodiments, the CRISPR-Cas effector protein (e.g., a Cas12a polypeptide / domain or a Cas9 polypeptide / domain) can be optimized for expression in a plant.
[0128] In some embodiments, the CRISPR-Cas effector protein (e.g., an engineered CRISPR-Cas effector protein) is a target-strand nickase and / or a non-target-strand nickase. In some embodiments, the CRISPR-Cas effector protein (e.g., an engineered CRISPR-Cas effector protein) is a non-target-strand nickase. In some embodiments, the CRISPR-Cas effector protein is an engineered protein having an amino acid sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to one or more of SEQ ID NOs: 223-314.
[0129] The polypeptides of the present invention can be used in combination with guide nucleic acids (e.g., guide RNA (gRNA), CRISPR array, CRISPR RNA, crRNA, or extended guide nucleic acid), which are designed to act together with a CRISPR-Cas effector protein to modify a target nucleic acid. The guide nucleic acids useful in the present invention can comprise at least one spacer sequence and at least one repeat sequence. The guide nucleic acid can be capable of forming a complex with a CRISPR-Cas effector protein (e.g., with the nuclease domain of the protein), and the spacer sequence can hybridize to the target nucleic acid, thereby guiding the complex to the target nucleic acid, wherein the target nucleic acid can be modified (e.g., cleaved or edited) and / or regulated (e.g., transcriptional regulation) by the polypeptides of the present invention optionally present in and / or recruited to the complex.
[0130] In some embodiments, a CRISPR-Cas effector protein comprising a Cas9 domain (or a nucleic acid construct encoding the same) can be used in combination with a Cas9 guide nucleic acid to modify a target nucleic acid and can be in or form a complex.
[0131] Similarly, the CRISPR-Cas effector protein can comprise a Cas12a domain (or other selected CRISPR-Cas nucleases, such as C2c1, C2c3, Cas12b, Cas12c, Cas12d, Cas12f, Cas12i, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4 (dinG) and / or Csf5), which can be used in combination with a Cas12a guide nucleic acid (or a guide nucleic acid of another selected CRISPR-Cas nuclease) to modify a target nucleic acid, thereby editing the target nucleic acid.
[0132] As used herein, "guide nucleic acid", "guide RNA", "gRNA", "CRISPR RNA / DNA", "crRNA", or "crDNA" refers to a nucleic acid that comprises at least one spacer sequence that is complementary (and hybridizes) to a target nucleic acid (e.g., a target DNA and / or protospacer) and at least one repeat sequence (e.g., a repeat sequence of a type V Cas12a CRISPR-Cas system, or a fragment or portion thereof; a repeat sequence of a type II Cas9 CRISPR-Cas system, or a fragment thereof; a repeat sequence of a type V C2c1 CRISPR-Cas system, or a fragment thereof; e.g., a repeat sequence of a CRISPR-Cas system of C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12i, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG), and / or Csf5, or a fragment thereof), wherein the repeat sequence may be linked to the 5'-end and / or 3'-end of the spacer sequence. In some embodiments, the guide nucleic acid comprises DNA. In some embodiments, the guide nucleic acid comprises RNA (e.g., is a guide RNA). The design of the gRNA of the present invention may be based on a type I, type II, type III, type IV, type V, or type VI CRISPR-Cas system.
[0133] In some embodiments, a Cas12a gRNA may comprise, from 5' to 3', a repeat sequence (full-length or a portion thereof ("handle"); e.g., a pseudoknot structure) and a spacer sequence.
[0134] In some embodiments, the guide nucleic acid can comprise more than one repeat-spacer sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). The guide nucleic acids of the present invention are synthetic, artificial and do not exist in nature. The gRNA can be very long and can be used as an aptamer (as in the MS2 recruitment strategy) or other RNA structures with hanging spacers.
[0135] As used herein, "repeat sequence" refers to any repeat sequence of, for example, a wild-type CRISPR Cas locus (e.g., Cas9 locus, Cas12a locus, C2c1 locus, etc.), or the repeat sequence of a synthetic crRNA that functions with a CRISPR-Cas effector protein encoded by a nucleic acid construct of the present invention. The repeat sequences that can be used in the present invention can be any known or later identified repeat sequences of CRISPR-Cas loci (e.g., type I, type II, type III, type IV, type V or type VI), or it can be a synthetic repeat sequence designed to function in type I, II, III, IV, V or VI CRISPR-Cas systems. The repeat sequence can comprise a hairpin structure and / or a stem-loop structure. In some embodiments, the repeat sequence can form a pseudoknot-like structure (i.e., "handle") at its 5' end. Thus, in some embodiments, the repeat sequence can be the same as or substantially the same as the repeat sequences from wild-type type I CRISPR-Cas loci, type II CRISPR-Cas loci, type III CRISPR-Cas loci, type IV CRISPR-Cas loci, type V CRISPR-Cas loci and / or type VI CRISPR-Cas loci. The repeat sequences from wild-type CRISPR-Cas loci can be determined by established algorithms, such as using CRISPRfinder provided by CRISPRdb (see, Grissa et al., Nucleic Acids Res. 35 (Web Server issue): W52-7). In some embodiments, the repeat sequence or a portion thereof is linked to the 5' end of the spacer sequence at its 3' end, thereby forming a repeat-spacer sequence (e.g., guide nucleic acid, guide RNA / DNA, crRNA, crDNA).
[0136] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides, depending on the particular repeat sequence and whether the guide nucleic acid comprising the repeat sequence is processed or unprocessed (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 to 100 or more nucleotides, or any range or value therein; e.g., about). In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100 or more nucleotides.
[0137] The repeat sequence linked to the 5' end of the spacer sequence may comprise a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive nucleotides of the wild-type repeat sequence). In some embodiments, the length of a portion of the repeat sequence linked to the 5' end of the spacer sequence may be about five to about ten consecutive nucleotides (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and have at least 90% sequence identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher) with the same region (e.g., the 5' end) of the wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, a portion of the repeat sequence may comprise a pseudoknot structure (e.g., a "handle") at its 5' end.
[0138] As used herein, an "spacer sequence" is a nucleotide sequence that is complementary to a target nucleic acid (e.g., target DNA) (e.g., a protospacer). The spacer sequence can be fully complementary or substantially complementary to the target nucleic acid (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher)). Thus, in some embodiments, the spacer sequence can have one, two, three, four, or five mismatches compared to the target nucleic acid, which can be contiguous or non - contiguous. In some embodiments, the spacer sequence can have 70% complementarity with the target nucleic acid. In other embodiments, the spacer nucleotide sequence can have 80% complementarity with the target nucleic acid. In still other embodiments, the spacer nucleotide sequence can have 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity with the target nucleic acid (protospacer), etc. In some embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence can have a length of about 15 nucleotides to about 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, the spacer sequence can have full or substantial complementarity over a region of the target nucleic acid (e.g., protospacer) that is at least about 15 nucleotides to about 30 nucleotides in length. In some embodiments, the length of the spacer is about 20 nucleotides. In some embodiments, the length of the spacer is about 21, 22, or 23 nucleotides.
[0139] In some embodiments, the 5' region of the spacer sequence of the guide nucleic acid can be fully complementary to the target nucleic acid, while the 3' region of the spacer can be substantially complementary to the target nucleic acid (e.g., for the spacer of type V CRISPR-Cas systems), or the 3' region of the spacer sequence of the guide nucleic acid can be fully complementary to the target nucleic acid, while the 5' region of the spacer can be substantially complementary to the target nucleic acid (e.g., for the spacer of type II CRISPR-Cas systems), and thus, the overall complementarity of the spacer sequence to the target nucleic acid can be less than 100%. Thus, for example, in the guide nucleic acid of a type V CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 5' region (i.e., the seed region) of a spacer sequence of, for example, 20 nucleotides can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 70% complementary). In some embodiments, the first 1 to 8 nucleotides at the 5' end of the spacer sequence (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides and any range therein) can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher)).
[0140] As another example, in the guide nucleic acid of a type II CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 3' region (i.e., the seed region) of a 20-nucleotide spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 70% complementary). In some embodiments, the first 1 to 10 nucleotides at the 3' end of the spacer sequence (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides and any range therein) can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target nucleic acid (e.g., at least about 50% complementary (e.g., at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher, or any range or value therein)). The recruitment guide RNA further comprises one or more recruitment motifs as described herein, which can be linked to the 5' or 3' end of the guide, or which can be inserted into the recruitment guide nucleic acid (e.g., within a hairpin loop).
[0141] In some embodiments, the length of the seed region of the spacer can be about 8 to about 10 nucleotides, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.
[0142] In some embodiments, the guide nucleic acid further comprises a reverse transcriptase template and can be referred to as an extended guide nucleic acid.
[0143] The guide nucleic acid and / or the extended guide nucleic acid can comprise one or more recruitment motifs as described herein, which can be linked to the 5' end and / or 3' end of the guide nucleic acid and / or which can be inserted into the guide nucleic acid (e.g., within a hairpin loop of the guide nucleic acid).
[0144] "Target nucleic acid", "target DNA", "target nucleotide sequence", "target region", and "target region in the genome" are used interchangeably herein and refer to a region in the genome of an organism (e.g., a plant) that contains a sequence that is perfectly complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher)) to the spacer sequence in the guide nucleic acid as defined herein. The target nucleic acid is targeted by an editing system (or its components) as described herein. The target region for use in a CRISPR-Cas system can be located adjacent to the 3' (e.g., for type V CRISPR-Cas systems) or adjacent to the 5' (e.g., for type II CRISPR-Cas systems) of a PAM sequence in the genome of an organism (e.g., a plant genome or a mammalian (e.g., human) genome). The target region can be selected from any region of at least 15 contiguous nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) located adjacent to a PAM sequence.
[0145] As used herein, "protospacer sequence" or "protospacer" refers to a sequence that is perfectly or substantially complementary (and can hybridize) to the spacer sequence of the guide nucleic acid. In some embodiments, the protospacer is all or a portion of the target nucleic acid as defined herein that is perfectly or substantially complementary (and hybridizes) to the spacer sequence of a CRISPR repeat-spacer sequence (e.g., a guide nucleic acid, a CRISPR array, a crRNA).
[0146] In the case of type V CRISPR-Cas (e.g., Cas12a) systems and type II CRISPR-Cas (Cas9) systems, the protospacer sequence is flanked by (e.g., adjacent to) a protospacer adjacent motif (PAM). For type IV CRISPR-Cas systems, the PAM is located at the 5' end of the non-target strand and the 3' end of the target strand (see below, as an example).
[0147]
[0148] In the case of type II CRISPR-Cas (e.g., Cas9) systems, the PAM is located adjacent to the 3' end of the target region. For type I CRISPR-Cas systems, the PAM is located at the 5' end of the target strand. There is no known PAM for type III CRISPR-Cas systems. Makarova et al. described the nomenclature for all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). R. Barrangou (Genome Biol. 16:247 (2015)) described the guide architecture and PAM.
[0149] Typical Cas12a PAMs are T-rich. In some embodiments, a typical Cas12a PAM sequence can be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, a typical Cas9 (e.g., Streptococcus pyogenes) PAM can be 5'-NGG-3'. In some embodiments, non-typical PAMs can be used, but with potentially lower efficiency.
[0150] One of ordinary skill in the art can determine additional PAM sequences through established experimental and computational methods. Thus, for example, experimental methods include targeting sequences flanking all possible nucleotide sequences and identifying sequence members that do not undergo targeting, e.g., by transformation of target plasmid DNA (Esvelt et al., 2013. Nat. Methods 10:1116-1121; Jiang et al., 2013. Nat. Biotechnol. 31:233-239). In some aspects, computational methods can include performing a BLAST search of native spacers to identify the original target DNA sequences in phages or plasmids and aligning these sequences to determine conserved sequences adjacent to the target sequences (Briner and Barrangou, 2014. Appl. Environ. Microbiol. 80:994-1001; Mojica et al., 2009. Microbiology 155:733-740).
[0151] In some embodiments, the present invention provides an expression cassette and / or vector comprising a nucleic acid construct of the present invention (e.g., one or more components of the editing system of the present invention). In some embodiments, an expression cassette and / or vector comprising a nucleic acid construct of the present invention and / or one or more guide nucleic acids may be provided. In some embodiments, the nucleic acid construct of the present invention encodes a polypeptide of the present invention and / or a CRISPR-Cas effector protein, and each may be contained on the same or a separate expression cassette or vector as an expression cassette or vector comprising one or more guide nucleic acids. When the nucleic acid construct encoding a polypeptide of the present invention or a component of the editing system is contained on an expression cassette or vector separate from the expression cassette or vector comprising the guide nucleic acid, the target nucleic acid and the expression cassette or vector encoding a polypeptide of the present invention or a component of the editing system may contact each other and the guide nucleic acid in any order (e.g., provided together), such as before, simultaneously, or after providing the expression cassette comprising the guide nucleic acid (e.g., contacting the target nucleic acid).
[0152] Methods for recruiting one or more components of an editing system to each other and / or to a target nucleic acid are known in the art and may include using peptide tags or affinity polypeptides that interact with peptide tags. In some embodiments, a guide nucleic acid may be linked to an RNA recruitment motif, and a polypeptide of the present invention may be linked to an affinity polypeptide capable of interacting with the RNA recruitment motif, thereby recruiting the polypeptide of the present invention to the target nucleic acid. Alternatively, chemical interactions may be used to recruit a polypeptide (e.g., a polypeptide of the present invention) to the target nucleic acid.
[0153] As used herein, a "recruitment motif" refers to one half of a binding pair that can be used to recruit a compound to which the recruitment motif binds to another compound (i.e., the "corresponding motif") that includes the other half of the binding pair. The recruitment motif and the corresponding motif may bind non-covalently. In some embodiments, the recruitment motif is an RNA recruitment motif (e.g., an RNA recruitment motif capable of binding to and / or configured to bind to an affinity polypeptide), an affinity polypeptide (e.g., an affinity polypeptide capable of binding to and / or configured to bind to an RNA recruitment motif and / or a peptide tag), or a peptide tag (e.g., a peptide tag capable of binding to and / or configured to bind to an affinity polypeptide). By way of example, when the recruitment motif is an RNA recruitment motif, the corresponding motif of the RNA recruitment motif may be an affinity polypeptide that binds the RNA recruitment motif. Another example is when the recruitment motif is a peptide tag, the corresponding motif of the peptide tag may be an affinity polypeptide that binds the peptide tag. Thus, a compound comprising a recruitment motif (e.g., an affinity polypeptide) may be recruited to another compound (e.g., a guide nucleic acid) comprising the corresponding motif of the recruitment motif (e.g., an RNA recruitment motif).
[0154] Peptide tags (e.g., epitopes) useful in the present invention may include, but are not limited to, the GCN4 peptide tag (e.g., Sun-Tag), c-Myc affinity tag, HA affinity tag, His affinity tag, S affinity tag, methionine-His affinity tag, RGD-His affinity tag, FLAG octapeptide, strep tag or strep tag II, V5 tag, and / or VSV-G epitope. Any epitope that can be linked to a polypeptide and for which there is a corresponding affinity polypeptide that can be linked to another polypeptide can be used as a peptide tag in the present invention. In some embodiments, the peptide tag may comprise 1 or 2 or more copies of the peptide tag (e.g., repeat units, multimerized epitopes (e.g., tandem repeats)) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more repeat units). In some embodiments, the affinity polypeptide that interacts / binds to the peptide tag may be an antibody. In some embodiments, the antibody may be a scFv antibody. In some embodiments, the affinity polypeptide that binds to the peptide tag may be synthetic (e.g., evolved for affinity interaction), including but not limited to affibody, anticalin, monobody, and / or DARPin (see, e.g., Sha et al., Protein Sci. 26(5):910-924 (2017)); Gilbreth (Curr Opin Struc Biol 22(4):413-420 (2013)), U.S. Patent No. 9,982,053, the teachings of which regarding affibody, anticalin, monobody, and / or DARPin are incorporated by reference in their entireties.
[0155] In some embodiments, a guide nucleic acid may be linked to an RNA recruitment motif, and the polypeptide to be recruited (e.g., a polypeptide of the present invention) may be fused to an affinity polypeptide that binds to the RNA recruitment motif, wherein the guide sequence binds to the target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the polypeptide to the guide sequence and bringing the target nucleic acid into contact with the polypeptide (e.g., a polypeptide of the present invention). In some embodiments, two or more polypeptides may be recruited to the guide nucleic acid, thereby bringing the target nucleic acid into contact with two or more polypeptides (e.g., one or more polypeptides of the present invention).
[0156] In some embodiments of the present invention, the guide RNA can be linked to one or two or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs; e.g., at least 10 to about 25 motifs), optionally where two or more RNA recruitment motifs can be the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide can include, but are not limited to, the telomerase Ku-binding motif (e.g., Ku-binding hairpin) and the corresponding affinity polypeptide Ku (e.g., Ku heterodimer), the telomerase Sm7-binding motif and the corresponding affinity polypeptide Sm7, the MS2 phage operator stem-loop and the corresponding affinity polypeptide MS2 coat protein (MCP), the PP7 phage operator stem-loop and the corresponding affinity polypeptide PP7 coat protein (PCP), the SfMu phage Com stem-loop and the corresponding affinity polypeptide Com RNA-binding protein, the PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA aptamer and aptamer ligand as the corresponding affinity polypeptide. In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide can be the MS2 phage operator stem-loop and the affinity polypeptide MS2 coat protein (MCP). In some embodiments, the RNA recruitment motif and the corresponding affinity polypeptide can be the PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA-binding factor (PUF). Exemplary RNA recruitment motifs and corresponding affinity polypeptides useful in the present invention can include, but are not limited to, SEQ ID NOs: 82-92.
[0157] In some embodiments, the components for recruiting polypeptides and nucleic acids can include components that act through chemical interactions, which can include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin; SNAP tag; Halo tag; CLIP tag; compound-induced DmrA-DmrC heterodimerization; bifunctional ligands (e.g., chemical-induced dimerization).
[0158] As described herein, a "peptide tag" can be used to recruit one or more polypeptides. A peptide tag can be any polypeptide that can be bound by a corresponding motif (e.g., an affinity polypeptide). A peptide tag can also be referred to as an "epitope" and, when provided in multiple copies, as a "multimerized epitope". Exemplary peptide tags can include, but are not limited to, the GCN4 peptide tag (e.g., Sun-Tag), the c-Myc affinity tag, the HA affinity tag, the His affinity tag, the S affinity tag, the methionine-His affinity tag, the RGD-His affinity tag, the FLAG octapeptide, the strep tag or strep tag II, the V5 tag, and / or the VSV-G epitope. In some embodiments, the peptide tag can also include a phosphorylated tyrosine in a specific sequence context recognized by an SH2 domain, a characteristic consensus sequence containing a phosphoserine recognized by a 14-3-3 protein, a proline-rich peptide motif recognized by an SH3 domain, a PDZ protein interaction domain, or a PDZ signaling sequence, and an AGO hook motif from plants. Peptide tags are disclosed in WO2018 / 136783 and U.S. Patent Application Publication No. 2017 / 0219596, the disclosures of which are incorporated by reference with respect to the peptide tags. Peptide tags useful in the present invention can include, but are not limited to, SEQ ID NO:93 and SEQ ID NO:94. Affinity polypeptides useful for peptide tags include, but are not limited to, SEQ ID NO:95.
[0159] The peptide tag can be included or present in one copy or two or more copies of the peptide tag (e.g., a multimerized peptide tag or multimerized epitope) (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, or 25 or more peptide tags). When multimerized, the peptide tags can be fused directly to each other, or they can be linked to each other via one or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more amino acids, optionally about 3 to about 10, about 4 to about 10, about 5 to about 10, about 5 to about 15, or about 5 to about 20 amino acids, etc., and any value or range therein). Thus, in some embodiments, the CRISPR-Cas effector protein and / or polypeptide of the invention can be fused to one peptide tag or to two or more peptide tags, optionally wherein the two or more peptide tags are fused to each other via one or more amino acid residues. In some embodiments, the peptide tag useful in the invention can be a single copy of a GCN4 peptide tag or epitope, or can be a multimerized GCN4 epitope, which comprises about 2 to about 25 or more copies of the peptide tag (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more copies of the GCN4 epitope, or any range therein).
[0160] In some embodiments, the peptide tag can be fused to a CRISPR-Cas polypeptide or domain. In some embodiments, the peptide tag can be fused or linked to the C-terminus of a CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, the peptide tag can be fused or linked to the N-terminus of a CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, the peptide tag can be fused within a CRISPR-Cas effector protein (e.g., the peptide tag can be in a loop region of the CRISPR-Cas effector protein). In some embodiments, the peptide tag can be fused to or a polypeptide of the invention.
[0161] An "affinity polypeptide" (e.g., "recruitment polypeptide") refers to any polypeptide that is capable of binding to its corresponding peptide tag, peptide tag, or RNA recruitment motif. The affinity polypeptide of a peptide tag can be, for example, an antibody and / or single-chain antibody that specifically binds to the peptide tag, respectively. In some embodiments, the antibody of the peptide tag can be, but is not limited to, an scFv antibody. In some embodiments, the affinity polypeptide can be fused or linked to the N-terminus of a polypeptide of the invention. In some embodiments, the affinity polypeptide is stable under reducing conditions in a cell or cell extract.
[0162] The nucleic acid constructs and / or guide nucleic acids of the present invention may be included in one or more expression cassettes as described herein. In some embodiments, the nucleic acid constructs of the present invention may be included in an expression cassette or vector that is the same as or separate from the expression cassette or vector containing the guide nucleic acid and / or extended guide nucleic acid.
[0163] In some embodiments, the nucleic acid constructs, expression cassettes or vectors of the present invention optimized for expression in an organism (e.g., a human or a plant) may be about 70% to 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100%) to nucleic acid constructs, expression cassettes or vectors containing the same polynucleotide but not codon-optimized for expression in the organism.
[0164] When used in combination with a guide nucleic acid, the nucleic acid constructs of the present invention (and expression cassettes and / or vectors containing the same) can be used to modify a target nucleic acid and / or its expression. The target nucleic acid can be contacted with the nucleic acid constructs of the present invention and / or expression cassettes and / or vectors containing the same before, simultaneously with, or after contacting the target nucleic acid with the guide nucleic acid / recruiting guide nucleic acid (and / or expression cassettes and vectors containing the same).
[0165] According to embodiments of the present invention, polypeptides (e.g., engineered proteins) are provided herein. As used herein, an "engineered protein" refers to a polypeptide or protein that is not naturally found in nature. Engineered proteins may be referred to as mutant proteins. In some embodiments, the mutant protein comprises a non-natural mutation. In some embodiments, the polypeptides of the present invention have at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96 - 133. The polypeptides of the present invention can generate DNA from RNA (e.g., cDNA). In some embodiments, the polypeptides of the present invention have the activity of an RNA-dependent DNA polymerase, and / or the polypeptide generates DNA by polymerizing DNA from one end (e.g., the 3' end) of a DNA primer and / or an RNA primer. In some embodiments, the polypeptides of the present invention can be monomers or homodimers, and / or can generate DNA from RNA as monomers or as homodimers. In some embodiments, the polypeptides of the present invention have reverse transcriptase activity. The activity of the polypeptides of the present invention can be measured by polymerase activity using methods known in the art. Exemplary assays for measuring and / or determining polymerase activity include, but are not limited to, measuring the incorporation of labeled nucleotides, such as radioactively labeled and / or colorimetric (e.g., fluorescent) labeled nucleotides. In some embodiments, radioactively labeled nucleotides are used to measure and / or determine polymerase activity (e.g., measuring the amount of radioactively labeled nucleotides included in the polymerized nucleic acid). In some embodiments, a primer extension assay is used to measure and / or determine polymerase activity, optionally wherein the primer extension assay uses a labeled primer (e.g., a fluorescently labeled primer, such as a Cy3 or Cy5 labeled primer) to determine the synthesis rate. In some embodiments, a cleavage assay is used to measure and / or determine polymerase activity (e.g., RNase activity).
[0166] In some embodiments, the activity of the polypeptide of the present invention is measured by the number of nucleotides generated (e.g., polymerized) during one cell division and / or within about 20 minutes. A person skilled in the art can easily determine cell division and / or the time period of one cell division. In some embodiments, the cell division is the cell division of bacterial cells, human cells, or plant cells. In some embodiments, the time of one cell division is about 18 to about 22 minutes, about 19 to about 21 minutes, or about 20 minutes. In some embodiments, the polypeptide of the present invention generates (e.g., polymerizes) at least 18, 19, 20, 21, 22, 23, 24, 25, or more nucleotides during one cell division and / or within about 20 minutes, optionally at a temperature of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C. The polypeptide activity can be measured by the rate of nucleotide generation (e.g., polymerization) over a period of time. In some embodiments, the polypeptide of the present invention polymerizes at least 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides within 20 minutes. In some embodiments, the polypeptide of the present invention polymerizes at least 0.9, 0.95, 1.0, 1.05, 1.1, 1.15, 1.20, 1.25, or more nucleotides per minute at a physiologically relevant temperature. In some embodiments, the polypeptide of the present invention generates nucleotides at a rate of at least 0.9, 0.95, 1.0, 1.05, 1.1, 1.15, 1.20, or 1.25 nucleotides per minute at a temperature of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C. In some embodiments, the polypeptide of the present invention generates (e.g., polymerizes) at least 23 nucleotides during one cell division and / or within about 20 minutes, optionally at a temperature of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C.
[0167] In some embodiments, the polypeptide of the present invention generates DNA from RNA at a temperature in the range of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C or 50°C to about 55°C, 60°C, 65°C, 70°C, 75°C or 80°C. In some embodiments, the polypeptide of the present invention generates DNA from RNA at a temperature of about 50°C or lower, for example, at a temperature of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C or 45°C. In some embodiments, the polypeptide of the present invention generates DNA from RNA at a temperature in the range of about 10°C to about 40°C, about 15°C to about 35°C, about 15°C to about 40°C, about 15°C to about 50°C, about 18°C to about 30°C, about 18°C to about 25°C, about 20°C to about 25°C, about 20°C to about 80°C, about 20°C to about 50°C, about 20°C to about 45°C, about 30°C to about 50°C, about 30°C to about 45°C, about 30°C to about 40°C, about 20°C to about 22°C, at a temperature of about 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C or 25°C or about room temperature.
[0168] As used herein, "processive synthesis ability" refers to the number of nucleotides generated (e.g., synthesized) in a single binding event of the polypeptide of the present invention. The processive synthesis ability can be measured in vitro and / or in vivo using methods known in the art. In some embodiments, the processive synthesis ability can be measured by the number of bases generated per unit of enzyme (e.g., the polypeptide of the present invention) over a period of time, where 1 unit of enzyme is the amount of enzyme that incorporates 1 nmol of dTTP into acid-insoluble material in a total reaction volume of 50 μl at 37°C for 10 minutes using poly(rA)·oligo(dT) as a template primer and 50 mM Tris-HCl (pH 8.3), 6 mM MgCl2, 10 mM dithiothreitol, 0.5 mM [3H]-dTTP, and 0.4 mM poly(rA)·oligo(dT)12-18. In some embodiments, the polypeptide of the present invention has a processive synthesis ability of at least about 100, 250, 500, 1000, 1500, or 2000 or more nucleotides, such as about 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700, 1750, 1800, 1850, 1900, 1950, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 or more nucleotides. In some embodiments, the polypeptide of the present invention has a processive synthesis ability of about 100 to about 600 nucleotides, about 200 to about 500 nucleotides, or at least about 300 nucleotides.
[0169] Compared to the processive synthesis ability of a control reverse transcriptase, the polypeptides of the present invention may have an improved processive synthesis ability. As used herein, "control reverse transcriptase" or "control RT" refers to a naturally occurring reverse transcriptase or a commercially available reverse transcriptase. Exemplary control RTs include, but are not limited to, Moloney murine leukemia virus reverse transcriptase (MMLV-RT or M-MuLV-RT), mutant MMLV (e.g., 5M-MMLV), avian myeloblastosis virus reverse transcriptase (AMV-RT), human immunodeficiency virus reverse transcriptase (HIV-RT). In some embodiments, the control RT can be MMLV and / or 5M-MMLV. In some embodiments, the control RT can have the sequence of one of SEQ ID NOs: 172-182. The processive synthesis ability of the polypeptides of the present invention and the control RT can be measured by using one or more sequences under the same reaction conditions (e.g., the same time, temperature, concentration, sequence identity, modification, etc.). In some embodiments, the polypeptides of the present invention can have a processive synthesis ability that is reduced compared to the processive synthesis ability of the control RT (e.g., lower than the processive synthesis ability of the control RT) and / or can have a processive synthesis ability that is faster than the ability of a DNA repair enzyme to correct modifications, such that the polypeptide can outperform the DNA repair enzyme. In some embodiments, the processive synthesis ability of the polypeptides of the present invention at a first temperature is equal to or superior to the processive synthesis ability of a control reverse transcriptase at a second temperature, wherein the first temperature is lower than the second temperature. For example, the processive synthesis ability of the polypeptides of the present invention at room temperature can be equal to or superior to the processive synthesis ability of a control reverse transcriptase (e.g., 5M-MMLV) at a temperature higher than room temperature (e.g., about 42 °C or about 55 °C). In some embodiments, the processive synthesis ability of the polypeptides of the present invention is within ± about 5%, about 10%, about 15%, about 20%, or about 25% of the processive synthesis ability of a DNA polymerase in a cell. In some embodiments, the processive synthesis ability of the polypeptides of the present invention is within about 5%, about 10%, about 15%, about 20%, or about 25% of the rate of one or more replicative polymerases of a cell, one or more repair polymerases of a cell, or a combination thereof.
[0170] In some embodiments, the polypeptide of the invention lacks at least a portion of the RNaseH domain. The RNaseH domain is conserved in viral RTs and is structurally similar to the RNaseH domains in Escherichia coli, Bacillus halodurans, and human RNase H1. See, e.g., Champoux et al., FEBS Journal, 276:6, 1506-1516 (2009). The enzymatic activity is conferred by the DEDD sequence motif, which is a conserved sequence composed of aspartic acid and glutamic acid residues, and which are Asp443, Glu478, Asp498, and Asp549 in HIV-1 RNase H, and the corresponding active site amino acids in the M-MLV enzyme are Asp524, Glu562, Asp583, and Asp653.
[0171] In some embodiments, the polypeptide of the invention has reduced (e.g., about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 15%, 10%, 5% or less) ribonuclease (RNase) activity or no RNase activity. In some embodiments, the polypeptide of the invention has reduced (e.g., about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 15%, 10%, 5% or less) ribonuclease (RNase) activity or no RNase activity compared to a control RT. In some embodiments, the control reverse transcriptase can be an RT with high RNase activity, such as avian myeloblastosis virus (AMV), or an RT with moderate RNasH activity, such as MMLV-RT and / or 5M-MMLV.
[0172] In some embodiments, the polypeptides of the present invention are capable of and / or can generate DNA from RNA in two or more different species (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more different species). In some embodiments, the polypeptides of the present invention generate DNA from RNA in two or more different species and optionally have a continuous synthesis ability of at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, or 2,000 or more nucleotides at a temperature in the range of about 25°C, 30°C, 35°C, 40°C, 45°C, or 50°C to about 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C. In some embodiments, the polypeptides of the present invention generate DNA from RNA in two or more different species and, for each of the two or more different species, at a temperature in the range of about 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, or 50°C to about 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C. In some embodiments, the polypeptides of the present invention generate DNA from RNA in two or more different species and optionally have a continuous synthesis ability of at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, or 2,000 or more nucleotides at a temperature below 50°C (e.g., about 20°C, 25°C, 30°C, 35°C, 40°C, or 45°C) or in the range of about 20°C, 25°C, or 30°C to about 35°C, 40°C, or 45°C.
[0173] The polypeptides of the present invention can generate DNA from RNA in prokaryotes or eukaryotes. In some embodiments, the polypeptides of the present invention generate DNA from RNA in plants and / or animals. In some embodiments, the polypeptides of the present invention generate DNA from RNA in: maize, soybean, canola, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, blackberry, raspberry, black raspberry, and / or Brassica spp.
[0174] The polypeptides of the present invention can be fusion proteins. In some embodiments, the polypeptides of the present invention are fused to a CRISPR-Cas effector protein or a portion thereof (e.g., directly or indirectly, e.g., via a linker).
[0175] The polypeptide of the present invention (optionally in the complex of the present invention) may have the same or higher (e.g., increased compared thereto) editing efficiency (e.g., inversion-deletion percentage) as a control reverse transcriptase (optionally in a complex having components similar to the complex having the polypeptide of the present invention). In some embodiments, the polypeptide of the present invention has the same or higher editing efficiency as the control RT when performed at a temperature below 50 °C, such as about 45 °C, about 42 °C, about 37 °C, about 30 °C, about 25 °C, or about room temperature.
[0176] In some embodiments, a polypeptide of the present invention having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to one or more of SEQ ID NOs: 96 - 133 has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or higher sequence identity to one or more of SEQ ID NOs: 187 - 195. In some embodiments, a polypeptide of the present invention (e.g., having at least about 70% sequence identity to one or more of SEQ ID NOs: 96 - 133) has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or higher sequence identity to one or more of SEQ ID NOs: 187 - 195. In some embodiments, a polypeptide of the present invention (e.g., having at least about 70% sequence identity to one or more of SEQ ID NOs: 96 - 133) has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or higher sequence identity to one or more of SEQ ID NOs: 187 - 195 and has 1, 2, 3, 4, 5, 6, 7, 8, or more mutations (e.g., one or more non-natural mutations) relative to the sequences of SEQ ID NOs: 187 - 195, optionally when optimally aligned therewith.
[0177] According to some embodiments of the present invention, there is provided a complex comprising: a CRISPR-Cas effector protein; a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96-133; and an extended guide nucleic acid. In some embodiments, the CRISPR-Cas effector protein is a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein. The CRISPR-Cas protein may be a fusion protein. In some embodiments, the CRISPR-Cas effector protein may comprise a CRISPR-Cas effector polypeptide fused to a peptide tag. In some embodiments, the CRISPR-Cas effector protein may comprise a CRISPR-Cas effector polypeptide fused to an affinity polypeptide capable of binding a peptide tag. In some embodiments, the CRISPR-Cas effector protein may comprise a CRISPR-Cas effector polypeptide fused to an affinity polypeptide capable of binding an RNA recruitment motif. In some embodiments, when the CRISPR-Cas protein comprises a peptide tag, the polypeptide of the present invention comprises an affinity polypeptide capable of binding the peptide tag. In some embodiments, when the CRISPR-Cas protein comprises an affinity polypeptide, the polypeptide of the present invention comprises a peptide tag capable of binding the affinity polypeptide. In some embodiments, the polypeptide of the present invention may be fused to an affinity polypeptide capable of binding an RNA recruitment polypeptide. The complex of the present invention may further comprise a guide nucleic acid, which optionally does not contain a reverse transcriptase template. In some embodiments, the guide nucleic acid is directed against a target nucleic acid different from the extended guide nucleic acid.
[0178] Nucleic acid molecules can encode the polypeptides of the present invention, and such nucleic acid molecules can be present in expression cassettes and / or vectors. The polynucleotides and / or recombinant nucleic acid constructs of the present invention can be codon-optimized for expression. In some embodiments, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the present invention (e.g., those comprising / encoding the polypeptides of the present invention, nucleic acid-binding polypeptides (e.g., DNA-binding domains, such as sequence-specific DNA-binding domains from polynucleotide-guided endonucleases such as CRISPR-Cas effector proteins) and / or extended guide nucleic acids) can be codon-optimized for expression in an organism (e.g., an animal, a plant, a fungus, an archaeon, or a bacterium). In some embodiments, the expression cassette and / or vector of the present invention comprises a polynucleotide encoding a promoter sequence and a polynucleotide encoding a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96 - 133 and 183 - 195. The polynucleotide encoding the polypeptide of the present invention can be codon-optimized for expression in a specific organism (e.g., a human or a plant). In some embodiments, the expression cassette and / or vector of the present invention comprises a polynucleotide encoding a promoter sequence and a polynucleotide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 134 - 171 and 196 - 204. In some embodiments, the expression cassette and / or vector further comprises a polynucleotide encoding a CRISPR-Cas effector protein, which is codon-optimized for expression in the organism. In some embodiments, the organism is an animal (e.g., a human), a plant, a fungus, an archaeon, or a bacterium.
[0179] Methods for modifying a target nucleic acid in a cell are provided, and the methods can comprise introducing the expression cassette and / or vector of the present invention into the cell to provide a modified target nucleic acid. In some embodiments, the cell is a plant cell, and the method further comprises regenerating the plant cell comprising the modified target nucleic acid to produce a plant comprising the modified target nucleic acid. In some embodiments, the introduction of the expression cassette is carried out at a temperature of about 20°C to about 42°C.
[0180] Methods for producing the polypeptides of the present invention are provided herein, and the methods can comprise: culturing a cell or cells that have been transformed with a nucleic acid encoding the polypeptide of the present invention; and isolating the polypeptide of the present invention, thereby producing the polypeptide. In some embodiments, the isolation step is carried out by dialysis, centrifugation, column purification, etc.
[0181] The present invention provides methods for performing reverse transcription, and the methods may include contacting a target nucleic acid with a polypeptide of the present invention, which has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96-133. In some embodiments, the methods of the present invention may include comparing the activity (e.g., reverse transcriptase activity), percentage of precise editing, and / or percentage of inversion-deletion after contacting the polypeptide of the present invention with the target nucleic acid with the same activity, percentage of precise editing, and / or percentage of inversion-deletion of a protein after contacting the protein with the same target nucleic acid, wherein the polypeptide of the present invention has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96-133, and the protein has the sequence of one of SEQ ID NOs: 183-195. The polypeptide of the present invention can reverse transcribe the target nucleic acid to provide DNA (e.g., cDNA). In some embodiments, the target nucleic acid is in a plant and / or in a plant cell including a cell wall. In some embodiments, the target nucleic acid is in a mammal and / or in a mammalian cell.
[0182] The present invention provides methods for modifying a target nucleic acid, and the methods may include contacting the target nucleic acid with a CRISPR-Cas effector protein, a polypeptide of the present invention, and an extended guide nucleic acid, wherein the polypeptide of the present invention has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96-133. In some embodiments, the target nucleic acid is in a plant and / or in a plant cell including a cell wall. In some embodiments, the method for modifying the target nucleic acid is a templated editing method.
[0183] The method of the present invention may comprise contacting the target nucleic acid with an extended guide nucleic acid. In some embodiments, the extended guide nucleic acid comprises a primer binding site, and the target nucleic acid is double-stranded and comprises a first strand and a second strand. In some embodiments, the primer binding site of the extended guide nucleic acid binds to the first strand or the second strand of the target nucleic acid. In some embodiments, the second strand is the non-target strand of the target nucleic acid. In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site binds to the first strand of the target nucleic acid. The first strand may be the target strand of the target nucleic acid, and / or the CRISPR-Cas effector protein is recruited to the first strand. In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site of the extended guide nucleic acid binds to the second strand of the target nucleic acid. In some embodiments, the second strand is the non-target strand of the target nucleic acid, and / or the CRISPR-Cas effector protein is recruited to the second strand. In some embodiments, the CRISPR-Cas effector protein is a double-stranded nuclease that cleaves the first strand and the second strand of the target nucleic acid, resulting in a double-strand break. In some embodiments, the CRISPR-Cas effector protein, the polypeptide of the present invention, and the extended guide nucleic acid form a complex or are included in a complex.
[0184] The method of the present invention may comprise contacting a target nucleic acid with an extended guide nucleic acid and / or introducing the extended guide nucleic acid into a cell, the extended guide nucleic acid comprising: (i) a CRISPR nucleic acid and / or a CRISPR nucleic acid and a tracr nucleic acid; and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template). In some embodiments, the extension portion of the extended guide nucleic acid is fused to the 5' end or the 3' end of the CRISPR nucleic acid (e.g., from 5' to 3': repeat sequence - spacer - extension portion or extension portion - repeat sequence - spacer) and / or fused to the 5' end or the 3' end of the tracr nucleic acid. In some embodiments, the extension portion of the extended guide nucleic acid comprises an RT template and a primer binding site from 5' to 3'. In some embodiments, the extension portion of the extended guide nucleic acid is located 5' of the crRNA.
[0185] In some embodiments, the length of the primer binding site is from about 1 nucleotide to about 100 nucleotides. In some embodiments, the length of the primer binding site is at least 45 nucleotides, or the length is from about 45 nucleotides to about 100 nucleotides, such as 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleotides.
[0186] In some embodiments, the length of the RT template is from about 1 to about 100 nucleotides, or the length of the RT template can be about 40 nucleotides or less, such as 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide.
[0187] In some embodiments, the extension portion of the extended guide nucleic acid is linked to the CRISPR nucleic acid and / or the tracrRNA via a linker. The length of the linker can be from about 1 to about 100 nucleotides, such as a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleotides.
[0188] In some embodiments, in the methods of the present invention, the CRISPR-Cas effector protein is a fusion protein and / or the polypeptide of the present invention is a fusion protein. In some embodiments, the CRISPR-Cas effector protein, the polypeptide of the present invention, and / or the extended guide nucleic acid are fused with one or more components that recruit the polypeptide to the CRISPR-Cas effector protein. In some embodiments, the CRISPR-Cas effector protein is a type V CRISPR-Cas effector fusion protein comprising a type V CRISPR-Cas effector polypeptide fused (e.g., linked) to a peptide tag (e.g., an epitope or a multimerization epitope), and the polypeptide of the present invention is a reverse transcriptase fusion protein comprising the polypeptide fused (e.g., linked) to an affinity polypeptide that binds the peptide tag. In some embodiments, the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins. In some embodiments, the CRISPR-Cas effector protein is a type II CRISPR-Cas effector fusion protein comprising a type II CRISPR-Cas effector polypeptide fused (e.g., linked) to a peptide tag (e.g., an epitope or a multimerization epitope), and the polypeptide of the present invention is a reverse transcriptase fusion protein comprising the polypeptide fused (linked) to an affinity polypeptide that binds the peptide tag. In some embodiments, the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins.
[0189] In some embodiments, the extended guide nucleic acid is linked to an RNA recruitment motif, and the polypeptide of the present invention is a reverse transcriptase fusion protein comprising the polypeptide fused (e.g., linked) to an affinity polypeptide that binds the RNA recruitment motif. In some embodiments, the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins. In some embodiments, the extended guide nucleic acid (e.g., extended guide RNA) is linked to two or more RNA recruitment motifs, optionally wherein the two or more RNA recruitment motifs are the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, at least one of the two or more RNA recruitment motifs is located at the 3' end of the extended portion of the extended guide nucleic acid or is embedded in the extended portion.
[0190] The method of the present invention may comprise contacting the target nucleic acid with a Dna2 polypeptide and / or a 5'-flap endonuclease (FEN). In some embodiments, the FEN and / or Dna2 polypeptide is overexpressed (e.g., overexpressed in the presence of the target nucleic acid). In some embodiments, the FEN is a fusion protein comprising a FEN domain fused to a CRISPR-Cas effector protein, and / or wherein the Dna2 polypeptide is a fusion protein comprising a Dna2 domain fused to a CRISPR-Cas effector protein. In some embodiments, the CRISPR-Cas effector protein is a first CRISPR-Cas effector protein, and the method includes the step of contacting the target nucleic acid with a second CRISPR-Cas effector protein. In some embodiments, the first CRISPR-Cas effector protein nicks or cuts a first site on the first strand of a double-stranded target nucleic acid. In some embodiments, the nick or cut site is located upstream or downstream (e.g., 5' or 3') of a second site on the second strand of the target nucleic acid that has been nicked by a second CRISPR-Cas effector protein, at about 10 to about 125 base pairs.
[0191] In some embodiments, in terms of modifying the target nucleic acid, the efficiency of the method of the present invention is higher than the efficiency of a control method (e.g., a method using a control RT and performed under the same conditions). Compared with the control method, the method of the present invention can produce an increased level of inversion-deletion and / or an increased level of modification (e.g., precise modification).
[0192] In some embodiments, the complex and / or method of the present invention may be the complex and / or method described in U.S. Patent Application Publication No. 2021 / 0130835 and / or U.S. Patent Application Publication No. 2022 / 0145334, the contents of each of which are incorporated herein by reference in their entireties, but wherein the reverse transcriptase is the polypeptide of the present invention.
[0193] In some embodiments, the editing system of the present invention is used for prime editing. As used herein, "prime editing" and its grammatical variants refer to a nucleic acid editing technique that uses a Cas9 nickase domain fused to a reverse transcriptase to modify a target nucleic acid without a double-strand break or a donor DNA template. In prime editing, the Cas9 nickase domain cuts the non-complementary strand of DNA upstream of the PAM site, thereby providing a 3'-flap that is extended with a modified extension. More details regarding prime editing can be found in Anzalone et al. (2019) Nature 576, 149-157 and / or U.S. Patent Application Publication No. 2021 / 0147862, the contents of each of which are incorporated herein by reference in their entireties.
[0194] In some embodiments, the editing system of the present invention utilizes a Redraw editing system. More details regarding the Redraw editing system can be found in U.S. Patent Application Publication No. 2021 / 0130835 and / or U.S. Patent Application Publication No. 2022 / 0145334, the contents of each of which are incorporated herein by reference in their entireties.
[0195] As described herein, the polypeptides, nucleic acids, expression cassettes, and / or vectors of the present invention can be codon-optimized for expression in an organism. The organism that can be used in the present invention can be any organism or its cells for which nucleic acid modification can be used. The organism can include, but is not limited to, any animal (e.g., a mammal), any plant, any fungus, any archaea, or any bacterium. In some embodiments, the organism can be a plant or its cells. In some embodiments, the organism is an animal, such as a mammal (e.g., a human).
[0196] The target nucleic acid can be a genomic sequence from any organism (e.g., a eukaryote, such as a mammal or a plant). In some embodiments, the target nucleic acid is a genomic sequence from a model organism, such as, but not limited to, Escherichia coli, an immortalized human cell line (e.g., HEK293, HeLa, etc.), Caenorhabditis elegans, and / or Drosophila Melanogaster. In some embodiments, the target nucleic acid is a genomic sequence from a non-model organism. Exemplary non-model organisms include, but are not limited to, crop plants (e.g., fruit crop plants, vegetable crop plants, and / or field crop plants) and / or animals, such as humans, primates, and / or mice. In some embodiments, the non-model organism is a crop plant, such as corn, soybean, wheat, or canola. In some embodiments, the non-model organism is an animal used for testing and / or using human therapeutics.
[0197] The target nucleic acid of any plant or plant part can be modified using the nucleic acid constructs of the present invention. The polypeptides of the present invention can be used to modify any plant (or plant grouping, such as a genus or higher taxon), including angiosperms, gymnosperms, monocots, dicots, C3, C4, CAM plants, bryophytes, ferns and / or fern allies, microalgae and / or macroalgae. The plants and / or plant parts useful in the present invention can be plants and / or plant parts of any plant species / variety / cultivar. As used herein, the term "plant part" includes, but is not limited to, embryo, pollen, ovule, seed, leaf, stem, bud, flower, branch, fruit, grain, spike, rachis, husk, straw, root, root tip, anther, plant cell (including intact plant cells in plants and / or plant parts), plant protoplast, plant tissue, plant cell tissue culture, plant callus, plant clump, etc. As used herein, "bud" refers to the above-ground part, including leaves and stems. In addition, as used herein, "plant cell" refers to the structural and physiological unit of a plant, which contains a cell wall and can also refer to a protoplast. A plant cell can be in the form of an isolated single cell, or can be a cultured cell, or can be part of a higher tissue unit such as a plant tissue or plant organ.
[0198] Non-limiting examples of plants useful in the present invention include turfgrasses (e.g., Poa, Agrostis, Lolium, Festuca), Calamagrostis acutiflora, Deschampsia cespitosa, Miscanthus, Arundo donax, Panicum virgatum, vegetable crops, including artichoke, kohlrabi, arugula, leek, asparagus, lettuce (e.g., iceberg lettuce, leaf lettuce, romaine lettuce), taro, melons (e.g., muskmelon, watermelon, crenshaw, honeydew, cantaloupe), Brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, collard greens, napa cabbage, bok choy), cardoon, carrot, bok choy, okra, onion, celery, parsley, chickpea, parsnip, chicory, pepper, potato, cucurbit plants (e.g., zucchini, cucumber, Italian cucumber, pumpkin, squash, honeydew, watermelon, cantaloupe), radish, dry bulb onion, rutabaga, eggplant, salsify, broadleaf chicory, scallion, endive, garlic, spinach, green onion, squash, leafy greens, beets (sugar beet and fodder beet), sweet potato, chard, horseradish, tomato, radish, and spices; fruit crops, such as apple, apricot, cherry, nectarine, peach, pear, plum, prune, cherry, quince, fig, nuts (e.g., chestnut, pecan, pistachio, hazelnut, pistachio, peanut, walnut, macadamia, almond, etc.), citrus (e.g., clementine, kumquat, orange, grapefruit, tangerine, mandarin, lemon, lime, etc.), blueberry, black raspberry, boysenberry, cranberry, currant, gooseberry, loganberry, raspberry, strawberry, blackberry, grape (wine grape and table grape), avocado, banana, kiwi, persimmon, pomegranate, pineapple, tropical fruit, pome fruit, melon, mango, papaya, and lychee, field crop plants such as clover, alfalfa, timothy, evening primrose, camelina, corn / maize (forage corn, sweet corn, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, triticale, sorghum, tobacco, kapok, leguminous plants (beans (e.g., green beans and dry beans), lentils, peas, soybeans), oil plants (rape, canola, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa bean, peanut, oil palm), duckweed, Arabidopsis thaliana, fiber plants (cotton, flax, hemp, jute), Cannabis (e.g., Cannabis sativa, Cannabis indica, and Cannabis ruderalis), Lauraceae plants (cinnamon, camphor) or plants such as coffee tree, sugarcane, tea, and natural rubber plants; and / or flower bed plants, such as flowering plants, cacti, succulents, and / or ornamental plants (e.g., rose, tulip, violet), and trees, such as forest trees (broadleaf trees and evergreens, such as conifers;For example, elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow), as well as shrubs and other seedlings. In some embodiments, the nucleic acid constructs and / or expression cassettes and / or vectors encoding the same of the present invention can be used to modify maize, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry, and / or cherry.;
[0199] In some embodiments, the present invention provides cells (e.g., plant cells, animal cells, bacterial cells, archaeal cells, etc.) comprising the polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, or vectors of the present invention.
[0200] The present invention further encompasses one or more kits for practicing the methods of the present invention. The kits of the present invention can include reagents, buffers, and equipment for mixing, measuring, sorting, labeling, etc., as well as instructions suitable for modifying target nucleic acids.
[0201] In some embodiments, the present invention provides a kit comprising one or more polypeptides of the present invention described herein, nucleic acid constructs of the present invention, and / or expression cassettes and / or vectors and / or cells comprising the same, and optionally instructions for its use. In some embodiments, the kit can further comprise a CRISPR-Cas guide nucleic acid (corresponding to the CRISPR-Cas effector protein provided herein, which can be encoded by a polynucleotide) and / or an expression cassette and / or vector and / or cell comprising the same. In some embodiments, the guide nucleic acid can be provided on the same expression cassette and / or vector as one or more nucleic acid constructs of the present invention. In some embodiments, the guide nucleic acid can be provided on an expression cassette or vector separate from the expression cassette or vector comprising one or more nucleic acid constructs of the present invention.
[0202] Thus, in some embodiments, a kit is provided that comprises a nucleic acid construct comprising (a) one or more polynucleotides as provided herein, and (b) a promoter driving the expression of the one or more polynucleotides of (a). In some embodiments, the kit can further comprise a nucleic acid construct encoding a guide nucleic acid, wherein the construct comprises a cloning site for cloning a nucleic acid sequence identical or complementary to a target nucleic acid sequence into the backbone of the guide nucleic acid.
[0203] In some embodiments, the nucleic acid constructs of the present invention can be mRNA, which can encode one or more introns within the encoded polynucleotide. In some embodiments, the nucleic acid constructs and / or expression cassettes and / or vectors comprising the same of the present invention can further encode one or more selectable markers (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.) that can be used to identify transformants.
[0204] The polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems, and / or cells of the present invention may comprise all or a portion of the sequences of one or more of SEQ ID NOs: 1-314. In some embodiments, the polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems, and / or cells of the present invention may comprise at least about 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more consecutive amino acids of the sequences of one or more of SEQ ID NOs: 1-314.
[0205] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claims of the present invention, but are intended as examples of certain embodiments. Any variations of the exemplary methods contemplated by those skilled in the art are intended to fall within the scope of the present invention.
[0206] Examples
[0207] Example 1: Alternative reverse transcriptase (RT) activity in REDRAW editing
[0208] The activities and fidelities of 47 putative RT proteins were tested in a trans configuration in the REDRAW editing system and compared to MMLV-RT(5M) (RT(5M); SEQ ID NO: 186) as a control. In the absence of antibiotics, HEK293T cells were seeded into 48-well collagen-coated plates (Corning) using DMEM medium. At 70-80% confluence, cells were transfected with 500 ng of control or putative RT protein plasmid and 500 ng of guide RNA plasmid using 1.5 μL of LTX (ThermoFisher Scientific) according to the manufacturer's protocol. Three days later, cells were lysed using a crude extraction method with SDS buffer. Using the stagRNA guide sequence, each control or putative RT was scored based on the precise base pair editing and the percentage of inversion-deletion placement at three target sites on the DMNT1-001 and DMNT1-002 loci. As Figure 1 shown, certain putative RTs exhibited a certain degree of activity in the REDRAW system. The inversion-deletion frequencies of all tested putative RTs were low ( Figure 2 ).
[0209] Example 2: Testing RT activity in the REDRAW editing system
[0210] Select the RTs listed in Table 1 that were tested in Example 1 for further examination and / or modification to provide additional putative RTs for testing. The putative RTs tested in Example 2 are listed in Table 2 by their SEQ ID NO and the corresponding vectors encoding the putative RTs. Table 2 also provides a description of each putative RT. For example, as shown in Table 2, the putative RT identified as RT49 has the sequence of SEQ ID NO:96 and is the putative RT identified as RT1 of the sequence with SEQ ID NO:187 having D200N, T306K, and L603W mutations, and the putative RT identified as RT75 has the sequence of SEQ ID NO:97 and is the putative RT identified as RT3 of the sequence with SEQ ID NO:188 having D198N, T304K, W311F, E328P, and L602W mutations.
[0211] Table 1: RTs from Example 1.
[0212]
[0213]
[0214] Table 2: Putative RTs tested in Example 2.
[0215]
[0216]
[0217]
[0218] Table 3: Controls tested in Example 2.
[0219] vector description of the content encoded by the vector SEQ ID NO of the encoded content pWISE121 LbCas12a nuclease SEQ ID NO:183 pWISE6099 RE2 SEQ ID NO:184 pWISE6806 RE4 SEQ ID NO:185
[0220] As described in Example 1 above, the activities of the additional putative RTs listed in Table 2 and the controls listed in Table 3 were tested in HEK293T cells, except that the putative RTs were in an orientation fused to the CRISPR-Cas effector protein to provide fusion proteins comprising the putative RTs, rather than in the trans configuration as in Example 1. The fusion proteins comprising the putative RTs have the same structure as Control 3 (SEQ ID NO:185), except that the reverse transcriptase of Control 3 (which has the sequence of SEQ ID NO:186) was replaced by the SEQ ID NOs listed in Table 2. Thus, for example, pWISE7583 encodes a fusion protein having the same structure as pWISE6806, except that SEQ ID NO:97 was used instead of SEQ ID NO:186.
[0221] The guide RNAs tested (as crRNAs) for the inversion-deletion frequencies targeting the target DNMT-1 and HEK2 loci were tested as controls ( Figure 3 ). Precise base pair editing ( Figure 4 ) and the percentage of inversion-deletion placement ( Figure 5 ) in the DMNT1-001 and RNF2 loci were tested using four different stagRNA guide sequences. Fusion proteins including SEQ ID NO: 96, 125, 98, 127, 99, 100, 129, 101 or 128 were the sequences with the best performance in precise editing compared to the Cas12a control (SEQ ID NO: 183) and the REDRAW editors RE2 and RE4 (SEQ ID NO: 184 and 185, respectively). The inversion-deletion frequency of RT was assumed to be comparable to that of the control.
[0222] The foregoing is a description of the invention and should not be construed as a limitation thereof. The invention is defined by the following claims, and equivalents of the claims are included therein.
Claims
1. A polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NO: 96 - 133, optionally wherein the polypeptide generates DNA (e.g., cDNA) from RNA.
2. The polypeptide according to claim 1, wherein the polypeptide has the activity of an RNA - dependent DNA polymerase, and / or wherein the polypeptide generates the DNA by polymerizing the DNA from one end (e.g., the 3' end) of a DNA and / or RNA primer.
3. The polypeptide according to claim 1 or 2, wherein the polypeptide generates the DNA from the RNA at a temperature in the range of about 25°C, 30°C, 35°C, 40°C, 45°C or 50°C to about 55°C, 60°C, 65°C, 70°C, 75°C or 80°C.
4. The polypeptide according to any one of the preceding claims, wherein the polypeptide generates the DNA from the RNA at a temperature of about 40°C or lower.
5. The polypeptide according to any one of the preceding claims, wherein the polypeptide has a continuous synthesis ability of at least about 100, 250, 500, 1000, 1500 or 2000 nucleotides.
6. The polypeptide according to any one of the preceding claims, wherein the continuous synthesis ability of the polypeptide is improved compared to that of a control reverse transcriptase.
7. The polypeptide according to any one of the preceding claims, wherein the continuous synthesis ability of the polypeptide at a first temperature is equal to or superior to the continuous synthesis ability of a control reverse transcriptase at a second temperature, wherein the first temperature is lower than the second temperature.
8. The polypeptide according to any one of the preceding claims, wherein the polypeptide generates (e.g., polymerizes) at least 18, 19, 20, 21, 22, 23, 24, 25 or more nucleotides during one cell division and / or within about 20 minutes, optionally at a temperature of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C or 80°C.
9. The polypeptide according to any one of the preceding claims, wherein the polypeptide lacks at least a portion of the RNaseH domain.
10. The polypeptide according to any one of the preceding claims, wherein the polypeptide generates the DNA from the RNA in two or more different species.
11. The polypeptide according to any one of the preceding claims, wherein the polypeptide has a continuous synthesis ability of at least about 100, 500, 1,000 or 2,000 nucleotides at a temperature in the range of about 25°C, 30°C, 35°C, 40°C, 45°C or 50°C to about 55°C, 60°C, 65°C, 70°C, 75°C or 80°C in two or more different species.
12. The polypeptide according to any one of the preceding claims, wherein the polypeptide generates the DNA from the RNA as a monomer or as a homodimer.
13. The polypeptide according to any one of the preceding claims, wherein the cDNA synthesis rate of the polypeptide is equal to or greater than that of a control reverse transcriptase.
14. The polypeptide according to any one of the preceding claims, wherein optionally compared to a control reverse transcriptase, the polypeptide has reduced or no ribonuclease (RNase) activity.
15. The polypeptide according to any one of the preceding claims, wherein the polypeptide generates the DNA from the RNA in: maize, soybean, canola, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, blackberry, raspberry, black raspberry, and / or Brassica spp.
16. The polypeptide according to any one of the preceding claims, which further comprises a CRISPR-Cas effector polypeptide fused to the polypeptide.
17. The polypeptide according to any one of the preceding claims, wherein the polypeptide comprises at least a portion of an RNAseH domain.
18. The polypeptide according to any one of the preceding claims, wherein the polypeptide has the same or higher editing efficiency (e.g., percentage of inversion-deletion) as a control reverse transcriptase.
19. A nucleic acid encoding the polypeptide according to any one of the preceding claims.
20. A complex comprising: a CRISPR-Cas effector protein; a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96-133; and an extended guide nucleic acid.
21. The complex according to claim 20, wherein the CRISPR-Cas effector protein is a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein.
22. The complex according to claim 20 or 21, wherein the CRISPR-Cas effector protein is a fusion protein comprising a CRISPR-Cas effector polypeptide fused to a peptide tag.
23. The complex according to claim 20 or 21, wherein the CRISPR-Cas effector protein is a fusion protein comprising a CRISPR-Cas effector polypeptide fused to an affinity polypeptide capable of binding to a peptide tag.
24. The complex according to claim 20 or 21, wherein the CRISPR-Cas effector protein is a fusion protein, the fusion protein comprising a CRISPR-Cas effector polypeptide fused to an affinity polypeptide capable of binding to an RNA recruitment motif.
25. The complex according to any one of claims 20 to 24, wherein the polypeptide is fused to a peptide tag.
26. The complex according to any one of claims 20 to 25, wherein the polypeptide is fused to an affinity polypeptide capable of binding to a peptide tag.
27. The complex according to any one of claims 20 to 26, wherein the polypeptide is fused to an affinity polypeptide capable of binding to an RNA recruitment polypeptide.
28. The complex according to any one of claims 20 to 27, further comprising a guide nucleic acid.
29. The complex according to any one of claims 20 to 28, which is contained in an expression cassette, optionally wherein the expression cassette is contained in a vector.
30. An expression cassette that is codon-optimized for expression in an organism, the expression cassette comprising: a polynucleotide encoding a promoter sequence, and a polynucleotide encoding a polypeptide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 96-133, and / or a polynucleotide having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to one or more of SEQ ID NOs: 134-171, optionally wherein the polynucleotide encoding the polypeptide is codon-optimized for expression in the organism.
31. The expression cassette according to claim 30, further comprising a polynucleotide encoding a CRISPR-Cas effector protein, the polynucleotide being codon-optimized for expression in the organism.
32. The expression cassette according to claim 30 or 31, wherein the organism is an animal, a plant, a fungus, an archaeon or a bacterium.
33. A method of modifying a target nucleic acid in a cell, the method comprising: introducing the expression cassette according to any one of claims 30 to 32 into the cell, thereby modifying the target nucleic acid in the cell.
34. The method according to claim 33, wherein the cell is a plant cell, and the method further comprises regenerating the plant cell comprising the modified target nucleic acid to produce a plant comprising the modified target nucleic acid.
35. The method according to any one of claims 30 to 34, wherein the introducing is carried out at a temperature of about 20°C to about 42°C.
36. A method for producing the polypeptide according to any one of claims 1 to 18, the method comprising: culturing a cell or cell population that has been transformed with a nucleic acid encoding the polypeptide; and Isolate the polypeptide (e.g., by dialysis, centrifugation, column purification, etc.) to thereby produce the polypeptide.
37. A method for performing reverse transcription, the method comprising: Contacting a target nucleic acid with a polypeptide that has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96 - 133, wherein the polypeptide reverse transcribes the target nucleic acid to provide DNA (e.g., cDNA).
38. A method for modifying a target nucleic acid, the method comprising: Contacting the target nucleic acid with the following to thereby modify the target nucleic acid a CRISPR - Cas effector protein; a polypeptide that has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NOs: 96 - 133; and an extended guide nucleic acid.
39. The method according to claim 37 or 38, wherein the polypeptide is the polypeptide according to any one of claims 1 to 18.
40. The method according to any one of claims 37 to 39, wherein the extended guide nucleic acid comprises a primer binding site, and wherein the target nucleic acid is double - stranded and comprises a first strand and a second strand, optionally wherein the primer binding site binds to the first strand or the second strand of the target nucleic acid, further optionally wherein the second strand is the non - target strand of the target nucleic acid.
41. The method according to any one of claims 37 to 40, wherein the target nucleic acid is double - stranded and comprises a first strand and a second strand, and the primer binding site binds to the first strand of the target nucleic acid, optionally wherein the first strand is the target strand of the target nucleic acid and / or wherein the CRISPR - Cas effector protein is recruited to the first strand.
42. The method according to any one of claims 37 to 41, wherein the target nucleic acid is double - stranded and comprises a first strand and a second strand, and the primer binding site binds to the second strand of the target nucleic acid, optionally wherein the second strand is the non - target strand of the target nucleic acid and / or wherein the CRISPR - Cas effector protein is recruited to the second strand.
43. The method according to any one of claims 37 to 42, wherein the CRISPR - Cas effector protein is a double - strand nuclease that cleaves the first strand and the second strand of the target nucleic acid to result in a double - strand break.
44. The method according to any one of claims 37 to 43, wherein the CRISPR - Cas effector protein, the polypeptide, and the extended guide nucleic acid form a complex or are included in a complex.
45. The method according to any one of claims 37 to 44, wherein the extended guide nucleic acid comprises: (i) A CRISPR nucleic acid, and / or a CRISPR nucleic acid and a tracr nucleic acid; and (ii) An extension portion comprising a primer binding site and a reverse transcriptase template (RT template).
46. The method according to claim 45, wherein the extension portion is fused to the 5'-end or 3'-end of the CRISPR nucleic acid (e.g., 5' to 3': repeat-spacer-extension portion, or extension portion-repeat-spacer) and / or to the 5'-end or 3'-end of the tracr nucleic acid.
47. The method according to claim 45 or 46, wherein the extension portion of the extended guide nucleic acid comprises an RT template and a primer binding site from 5' to 3'.
48. The method according to any one of claims 45 to 47, wherein the length of the primer binding site is from about one nucleotide to about 100 nucleotides, optionally wherein the length of the primer binding site is at least 45 nucleotides or from about 45 nucleotides to about 100 nucleotides.
49. The method according to any one of claims 45 to 48, wherein the length of the RT template is from about 1 to about 100 nucleotides, optionally wherein the length of the RT template is about 40 nucleotides or less.
50. The method according to any one of claims 45 to 49, wherein the extension portion of the extended guide nucleic acid is linked to the CRISPR nucleic acid and / or the tracrRNA via a linker, optionally wherein the length of the linker is from about 1 to about 100 nucleotides.
51. The method according to any one of claims 45 to 50, wherein when the extension portion is located at the 5' of the crRNA, the CRISPR-Cas effector protein is modified to reduce (or eliminate) self-processing RNase activity.
52. The method according to any one of claims 37 to 51, wherein the CRISPR-Cas effector protein is a fusion protein and / or the polypeptide is a fusion protein, optionally wherein the CRISPR-Cas effector protein, the polypeptide and / or the extended guide nucleic acid are fused to one or more components that recruit the polypeptide to the CRISPR-Cas effector protein.
53. The method according to any one of claims 37 to 52, wherein the CRISPR-Cas effector protein is a type V CRISPR-Cas effector fusion protein, the effector fusion protein comprising a type V CRISPR-Cas effector polypeptide fused (linked) to a peptide tag (e.g., an epitope or a multimerization epitope), and the polypeptide is a reverse transcriptase fusion protein, the reverse transcriptase fusion protein comprising the polypeptide fused (linked) to an affinity polypeptide that binds to the peptide tag, optionally wherein the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins.
54. The method according to any one of claims 37 to 53, wherein the CRISPR-Cas effector protein is a type II CRISPR-Cas effector fusion protein, the effector fusion protein comprises a type II CRISPR-Cas effector polypeptide fused (linked) to a peptide tag (e.g., an epitope or a multimerization epitope), and the polypeptide is a reverse transcriptase fusion protein, the reverse transcriptase fusion protein comprises the polypeptide fused (linked) to an affinity polypeptide that binds to the peptide tag, optionally wherein the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins.
55. The method according to any one of claims 37 to 54, wherein the extended guide nucleic acid is linked to an RNA recruitment motif, and the polypeptide is a reverse transcriptase fusion protein, the reverse transcriptase fusion protein comprises the polypeptide fused (linked) to an affinity polypeptide that binds to the RNA recruitment motif, optionally wherein the target nucleic acid is contacted with two or more reverse transcriptase fusion proteins.
56. The method according to any one of claims 37 to 55, wherein the extended guide RNA is linked to two or more RNA recruitment motifs, optionally wherein the two or more RNA recruitment motifs are the same RNA recruitment motif or different RNA recruitment motifs.
57. The method according to claim 56, wherein at least one of the two or more RNA recruitment motifs is located at the 3' end of the extended portion of the extended guide nucleic acid or embedded in the extended portion.
58. The method according to any one of claims 37 to 57, further comprising contacting the target nucleic acid with a Dna2 polypeptide and / or a 5'-flap endonuclease (FEN), optionally wherein the FEN and / or Dna2 polypeptide is overexpressed (e.g., overexpressed in the presence of the target nucleic acid), and / or optionally wherein the FEN is a fusion protein comprising a FEN domain fused to the CRISPR-Cas effector protein, and / or wherein the Dna2 polypeptide is a fusion protein comprising a Dna2 domain fused to the CRISPR-Cas effector protein.
59. The method according to any one of claims 37 to 58, wherein the CRISPR-Cas effector protein is a first CRISPR-Cas effector protein, and the method further comprises contacting the target nucleic acid with a second CRISPR-Cas effector protein.
60. The method according to claim 59, wherein the first CRISPR-Cas effector protein nicks or cuts a first site on the first strand of the target nucleic acid, the first site being located about 10 to about 125 base pairs (5' or 3') from a second site on the second strand that has been nicked by the second CRISPR-Cas effector protein.
61. The method according to claim 59, wherein the first CRISPR-Cas effector protein nicks or cuts a first site on the second strand of the target nucleic acid, the first site being located about 10 to about 125 base pairs (5' or 3') from a second site on the first strand that has been nicked by the second CRISPR-Cas effector protein.
Citation Information
Patent Citations
Seed specific transcriptional regulation
EP0255378A2
Tissue-preferential promoters
EP0452269A2
Synthetic chloroplast transit peptides
US10421972B2
Improvement in safety-valves
US160167A
A protein tagging system for in vivo single molecule imaging and control of gene transcription
US20170219596A1