Fusion protein, recombinant nucleic acid, DNA construct, vector, cell, nucleic acid editing method, and plant genome.
Fusion proteins with a site-directed nuclease and recruiter domain enhance the precision and efficiency of genome editing by tethering donor polynucleotides to target sites, addressing the limitations of existing SDNs.
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-06-23
- Publication Date
- 2026-07-07
AI Technical Summary
Existing site-directed nucleases (SDNs) for genome editing, such as CRISPR/Cas systems, lack precision and efficiency in site-specific integration, often resulting in off-target edits and low frequency of desired insertion events.
Fusion proteins are developed, combining a site-directed nuclease with a recruiter domain comprising a site-specific DNA binding domain, such as a Cro repressor family protein, to enhance the specificity and frequency of genome editing by tethering a donor polynucleotide to the target site using homologous recombination.
The fusion proteins significantly increase the efficiency of site-specific integration and reduce off-target effects, improving the precision and frequency of desired edits in genomic DNA.
Smart Images

Figure 00000122_0000 
Figure 00000123_0000 
Figure 00000124_0000
Abstract
Description
1 / 116 “FUSION PROTEIN, RECOMBINANT NUCLEIC ACID, DNA CONSTRUCTOR, VECTOR, CELL, NUCLEIC ACID EDITING METHOD AND PLANT GENOME” CROSS-REFERENCE TO RELATED REQUESTS
[001] This request claims priority over the PCT Request No. PCT / CN2023 / 080827, filed on March 10, 2023, which is incorporated by reference. AREA
[002] The present invention relates to methods for increasing site-specific integration. The methods presented here are applicable to both non-homologous end joining (NHEJ) and homology-dependent repair (HDR) mechanisms. SEQUENCE LISTING
[003] This request is accompanied by a sequence listing titled 82448SL.xml, created on March 6, 2023, which is approximately 62.5 kilobytes in size. This sequence listing is incorporated herein by reference in its entirety. BACKGROUND
[004] Site-directed nucleases (SDNs) (e.g., zinc finger nucleases, transcription activator-like effector nucleases, CRISPR-associated nucleases) have gained increasing popularity in the gene editing space. These SDNs act as endonucleases and generally create double-strand breaks (DSBs) in specific DNA sequences, activating intrinsic cell repair mechanisms (e.g., homologous recombination). During the repair process, site-directed modification of the specific DNA sequence can be achieved. The CRISPR (Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated) system evolved in bacteria and Petition 870260060590, dated 06 / 22 / 2026, pp. 133 / 261 2 / 116 Archaebacteria as an adaptive immune system to defend themselves against viral attack. In recent years, the CRISPR / Cas system has attracted particular interest as a tool for genome editing. CRISPR / Cas systems that generate site-specific double-strand breaks (DSBs) can be used to edit DNA in eukaryotic cells, for example, by producing deletions, insertions, and / or changes in the nucleotide sequence.
[005] Site-directed modifications induced by SDNs often lack precision (e.g., off-target edits can occur), and they frequently occur at a low frequency. For example, when a CRISPR / Cas system is configured to cause site-specific integration using a donor template, the specificity of DSB targeting can vary, and the frequency of desired insertion events can be low. As such, there is a need for methods to increase the efficiency of targeted genome editing using SDNs. BRIEF SUMMARY
[006] The Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify critical or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter.
[007] In one aspect, this description provides a fusion protein comprising a site-directed nuclease fused to a recruiting domain comprising a site-specific DNA binding domain.
[008] In some embodiments, the site-directed nuclease comprises a CRISPR-associated nuclease. In some embodiments, the CRISPR-associated nuclease is selected from the group consisting of Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Petition 870260060590, dated 06 / 22 / 2026, pp. 134 / 261 3 / 116 Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, Cas14 and nicase or deactivated versions thereof. In some embodiments, the CRISPR-associated nuclease is a Cas9 enzyme. In some embodiments, the CRISPR-associated nuclease is a Cas12a enzyme.
[009] In some embodiments, the recruiter domain is a Cro repressor family protein. In some embodiments, the Cro repressor family protein comprises Cro of N15, Cro of lambda, Cro of P22, Cro of 434, or a combination of any of them. In some embodiments, the recruiter domain comprises an amino acid sequence having at least 90% identity with any of the SEQ IDs NOs: 1-4. In some embodiments, the recruiter domain comprises a dimerization domain.
[0010] In some embodiments, the fusion protein comprises a linker located between the site-directed nuclease and the recruiting domain. In some embodiments, the linker comprises any of the SEQ ID NO: 6, 7, 15, or 16. In some embodiments, the fusion protein comprises a nuclear localization signal.
[0011] In some embodiments, the fusion protein comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 11 or 13.
[0012] In another aspect, a recombinant nucleic acid encoding the fusion protein is provided in this description, the fusion protein comprising a site-directed nuclease fused to a recruiter domain comprising a site-specific DNA binding domain.
[0013] In another aspect, this description provides a DNA construct comprising a promoter operationally linked to recombinant nucleic acid. In some embodiments, the Petition 870260060590, dated 06 / 22 / 2026, pp. 135 / 261 4 / 116 promoter comprises at least one of an inducible promoter, a constitutive promoter, an egg cell-specific promoter, a pollen-specific promoter, or an apical meristem tissue-specific promoter. In some embodiments, the promoter is a ubiquitin 4 promoter, an actin promoter, a tubulin promoter, a MADS box promoter, or a plant virus promoter.
[0014] In another aspect, a vector comprising recombinant nucleic acid or DNA construct is provided in this description.
[0015] In another aspect, this description provides a cell comprising recombinant nucleic acid, DNA construct or vector. In some embodiments, the cell is a plant cell. In some embodiments, the plant cell is a corn plant cell, a soybean plant cell, a rice plant cell, a wheat plant cell and / or a sunflower plant cell.
[0016] In another aspect, a method for editing a nucleic acid is provided in this description, the method comprising: (a) providing at least one fusion protein as described in this description; (b) providing the nucleic acid, wherein the nucleic acid comprises a first binding site and a target region comprising a portion of the nucleic acid, wherein the first binding site is within or adjacent to the target region; (c) providing a donor polynucleotide comprising a donor nucleotide region and at least one recruiter sequence that is specifically bound by the recruiter domain of the at least one fusion protein; and (d) contacting the nucleic acid and the donor polynucleotide with the at least one fusion protein, wherein the at least one fusion protein specifically binds to the first binding site of the nucleic acid and to the recruiter sequence of the polynucleotide. Petition 870260060590, dated 06 / 22 / 2026, pp. 136 / 261 5 / 116 donor, thus resulting in an edit in the target region of the nucleic acid.
[0017] In some embodiments, the first binding site is adjacent to a 5' end or a 3' end of the target region. In some embodiments, the edit in the nucleic acid target region is a substitution of at least one portion of the target region by at least one portion of the donor polynucleotide.
[0018] In some embodiments, the nucleic acid additionally comprises a second binding site, wherein the second binding site is within or adjacent to the target region, and wherein at least one fusion protein binds specifically to the first binding site and the second binding site of the nucleic acid. In some embodiments, the second binding site is adjacent to a 5' end or a 3' end of the target region.
[0019] In some embodiments, the recruiting domain of the fusion protein comprises a Cro repressor family protein and at least one recruiting sequence comprises a Cro OR3 operon sequence. In some embodiments, the Cro OR3 operon sequence comprises an N15 OR3 operon sequence (optionally, SEQ ID NO:18), a lambda OR3 operon sequence, a P22 OR3 operon sequence, a 434 OR3 operon sequence, or a combination thereof.
[0020] In some embodiments, the donor polynucleotide comprises at least one homology arm, wherein the at least one homology arm comprises a nucleotide sequence having complementarity with a portion of the target region of the nucleic acid. In some embodiments, the donor polynucleotide comprises at least two recruiting sequences. In some embodiments, the donor polynucleotide comprises a first recruiting sequence adjacent to a 5' end of the nucleotide region. Petition 870260060590, dated 06 / 22 / 2026, pp. 137 / 261 6 / 116 donor and a second recruiting sequence adjacent to a 3' end of the donor nucleotide region. In some embodiments, at least the two recruiting sequences are not within the donor nucleotide region.
[0021] In some embodiments, the site-directed nuclease of at least one fusion protein comprises a CRISPR-associated nuclease and the method further comprises providing at least one guide RNA, wherein the at least one guide RNA comprises a nucleotide sequence having complementarity with the first binding site and / or the second nucleic acid binding site. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] This application includes the following figures. The figures are intended to illustrate specific modalities and / or characteristics of the compositions and methods and to complement any description(s) of the compositions and methods. The figures do not limit the scope of the compositions and methods unless the written description expressly indicates otherwise.
[0023] FIG. 1 shows schematic illustrations of various types of donor polynucleotides described herein, in accordance with aspects of this description. Donor DNA polynucleotides (“Donor DNA”) with a portion designed for insertion into a target site (“Insertion / Substitution Sequence”), as described herein, are illustrated. Left and right homology arms (“LHA” and “RHA”, respectively) and Cro OR3 recruiting sequences (“Cro OR3”) are also shown.
[0024] FIG. 2 shows a schematic illustration of one embodiment of the methods provided herein, in accordance with aspects of this description. A nucleic acid (“Genomic DNA”) with a target region and a donor DNA polynucleotide (“Donor DNA”) with a portion designed for insertion into the target region (“Sequence”) are illustrated. Petition 870260060590, dated 06 / 22 / 2026, pp. 138 / 261 7 / 116 insertion / substitution”), as described herein. Left and right homology arms (“LHA” and “RHA”, respectively) and Cro OR3 recruiting sequences (“Cro OR3”) are also shown. Potential cleavage sites that are more or less useful (“On-target cuts” and “Off-target cuts”, respectively) in particular embodiments provided herein are indicated on the genomic DNA locus.
[0025] FIG. 3 shows schematic illustrations of two fusion proteins provided herein located at a target site on a genomic DNA sequence and tethered to a donor DNA sequence, in accordance with aspects of this description.
[0026] FIG. 4 shows a schematic illustration of an N15 Cas9-Cro fusion protein, in accordance with aspects of this description.
[0027] FIG. 5 shows a schematic illustration of an N15 Cas12a-Cro fusion protein, in accordance with aspects of this description. DETAILED DESCRIPTION
[0028] The following description sets forth various aspects and embodiments of the present compositions and methods. No particular embodiment is intended to define the scope of compositions and methods. Instead, the embodiments merely provide non-limiting examples of various compositions and methods that are at least included within the scope of the disclosed compositions and methods. The description should be read from the perspective of a person skilled in the art; therefore, information well known to the person skilled in the art is not necessarily included. I. Terminology
[0029] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning. Petition 870260060590, dated 06 / 22 / 2026, pp. 139 / 261 8 / 116 meaning as commonly understood by someone skilled in the art. References to techniques used herein are intended to refer to techniques as commonly understood in the art, including variations of those techniques and / or substitutions for equivalent techniques that would be apparent to someone skilled in the art. Although it is believed that the following terms are well understood by someone skilled in the art, the following definitions are provided to facilitate the explanation of the subject matter presently disclosed.
[0030] As used herein, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, the reference to “an enzyme” optionally includes a combination of two or more such and similar molecules.
[0031] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0032] The term “about” as used here refers to the usual error range for the respective value readily known to the expert in the art in this technical area, for example ± 20%, ± 10% or ± 5% are within the intended meaning of the recited value.
[0033] As used herein, the term “comprising” or “comprises” is open-ended. When used in connection with a nucleic acid (or amino acid sequence) in question, it refers to a nucleic acid sequence (or an amino acid sequence) that includes the sequence in question as a part or as its entire sequence.
[0034] As used herein, the transitional phrase “consisting essentially of” means that the scope of a claim should be interpreted as encompassing the materials or steps specified in the claim and those that do not materially affect the(s) Petition 870260060590, dated 06 / 22 / 2026, pp. 140 / 261 9 / 116 basic and novel feature(s) of the claimed matter. Thus, the term “essentially consisting of” when used in a claim of this description is not intended to be interpreted as being equivalent to “comprising”.
[0035] The term “plurality” refers to more than one entity. Thus, a “plurality of individuals” refers to at least two individuals. In some modalities, the term plurality refers to more than half of the whole. For example, in some modalities, a “plurality of a population” refers to more than half of the members of that population.
[0036] The term “plant” as used herein refers to any plant at any stage of development, particularly a seed plant. The term “plant cell” as used herein refers to a structural and physiological unit of a plant, comprising a protoplast and a cell wall. The plant cell may be in the form of a single isolated cell or a cultured cell or as part of a higher-organized unit such as, for example, plant tissue, a plant organ or a whole plant. The plant cell may be derived from or part of an angiosperm or gymnosperm.A plant cell can be a monocotyledonous plant cell (for example, a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a grass cell, or an ornamental grass cell) or a dicotyledonous plant cell (for example, a tobacco cell, a pepper cell, an eggplant cell, a sunflower cell, a cruciferous cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar beet cell, or a rapeseed cell). The term “plant cell culture” as used here refers to cultures of plant units such as, for example, Petition 870260060590, dated 06 / 22 / 2026, pp. 141 / 261 10 / 116 protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at various stages of development. The term “plant tissue” as used herein refers to a group of plant cells organized into a structural and functional unit. Any tissue of a plant in planta or in culture is included. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue culture, and any group of plant cells organized into structural and / or functional units. The use of this term in conjunction with, or in the absence of, any specific type of plant tissue as listed above or otherwise covered by this definition is not intended to be exclusive of any other type of plant tissue.The term “plant part” as used herein refers to a part of a plant, including single cells and cellular tissues, such as plant cells that are intact in plants, cell clumps, and tissue cultures from which plants can be regenerated. Examples of plant parts include, but are not limited to, single cells and tissues of pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, buds, cuttings, and seeds; as well as pollen, ovules, ovule cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, buds, cuttings, shoots, rhizomes, seeds, protoplasts, calluses, and the like.
[0037] The terms “polypeptide”, “peptide”, and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.
[0038] The terms “nucleic acid” and “polynucleotide” are used Petition 870260060590, dated 06 / 22 / 2026, pp. 142 / 261 11 / 116 interchangeably and as used herein refer to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in single or double-stranded form, as well as to both sense and antisense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers thereof. In higher plants, DNA is the genetic material while RNA is involved in the transfer of information contained within DNA into proteins. A “genome” is the entire body of genetic material contained in each cell of an organism. It is understood that when an RNA is described, its corresponding cDNA is also described, where uridine is represented as thymidine. In particular embodiments, a nucleotide refers to a ribonucleotide, deoxynucleotide, or a modified form of any type of nucleotide and combinations thereof.Additionally, a polynucleotide disclosed herein may include one or both naturally occurring and modified nucleotides linked together by naturally occurring and / or unnaturally occurring nucleotide bonds. Nucleic acid molecules may be chemically or biochemically modified or may contain unnaturally occurring or derivatized nucleotide bases, as will be readily understood by those skilled in the art.Such modifications include, for example, tagging, methylation, substitution of one or more of the naturally occurring nucleotides by an analog, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates and the like), charged linkages (e.g., phosphorothioates, phosphorodithioates and the like), pendant fragments (e.g., polypeptides), intercalators (e.g., acridine, psoralen and the like), chelating agents, alkylating agents and modified linkages (e.g., anomeric alpha nucleic acids and the like). The above term is also intended to include any topological conformation. Petition 870260060590, dated 06 / 22 / 2026, pp. 143 / 261 12 / 116 including single-stranded, double-stranded, partially duplexed, triplex, hairpin, circular, and blocked conformations. A reference to a nucleic acid sequence includes its complement unless otherwise specified. Thus, a reference to a nucleic acid molecule having a particular sequence should be understood as including its complementary strand, with its complementary sequence. Nucleotide sequences are “complementary” when they specifically hybridize in solution (e.g., according to Watson-Crick base-pairing rules). The term also includes codon-optimized nucleic acids encoding the same polypeptide sequence. It is also understood that nucleic acids may be unpurified, purified, or attached, for example, to a synthetic material such as a spherule or column array.
[0039] The term “corresponding to” in the context of nucleic acid sequences means that, when the amino acid sequences of certain sequences are aligned with each other, the nucleic acids that “correspond to” certain positions enumerated in the present invention are those that align with these positions in a reference sequence, but are not necessarily in these exact numerical positions relative to a particular amino acid sequence of the invention. Optimal sequence alignment for comparison can be conducted by computerized implementations of known algorithms or by visual analysis. Readily available sequence comparison and multiple sequence alignment algorithms are, respectively, the Basic Local Alignment Search Tool (BLAST) and the ClustalW / ClustalW2 / Clustal Omega programs available on the internet (e.g., the EMBL-EBI website).Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and... Petition 870260060590, dated 06 / 22 / 2026, pp. 144 / 261 13 / 116 FASTA, which are part of the Accelrys GCG package available from Accelrys, Inc. of San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988; and Sambrook & Russell, 2001.
[0040] Unless otherwise indicated, a particular nucleic acid sequence also implicitly includes conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the explicitly indicated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected codons (or all) is replaced by mixed base and / or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19: 5081 (1991); Ohtsuka et al., J. Biol. Chem. 260: 2605-2608 (1985); and Rossolini et al., Mol. Cell. Sondas 8: 91-98 (1994).
[0041] The terms “identity” or “substantial identity,” as used in the context of a polynucleotide or polypeptide sequence described herein, refer to a sequence that has at least 60% sequence identity with a reference sequence. Alternatively, the percentage of identity may be any integer from 60% to 100%. Exemplary embodiments include at least: 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, compared to a reference sequence using the programs described herein; preferably BLAST using standard parameters, as described below. An expert will recognize that these values can be appropriately adjusted to determine the corresponding identity of the proteins encoded by two nucleotide sequences, taking into account codon degeneracy, the Petition 870260060590, dated 06 / 22 / 2026, pp. 145 / 261 14 / 116 amino acid similarity, reading frame positioning and similar.
[0042] For sequence comparison, typically one sequence acts as a reference sequence against which test sequences are compared. When using a sequence comparison algorithm, the test and reference sequences are entered into a computer, subsequence coordinates are assigned if necessary, and the sequence algorithm program parameters are assigned. Standard program parameters can be used, or alternative parameters can be assigned. The sequence comparison algorithm then calculates the percentage of sequence identities for the test sequences with respect to the reference sequence, based on the program parameters.
[0043] A “comparison window”, as used herein, includes reference to a segment of any one of a number of contiguous positions selected from the group consisting of 20 to 600, usually about 50 to about 200, more usually about 100 to about 150, in which a sequence can be compared with a reference sequence of the same number of contiguous positions after the two sequences are ideally aligned. Sequence alignment methods for comparison are well known in the art. Ideal sequence alignment for comparison can be conducted by the local homology algorithm of Smith and Waterman Add. APL. Math. 2: 482 (1981), by the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48: 443 (1970), by the similarity search method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85: 2444 (1988), by computerized implementations of these algorithms (e.g., BLAST) or by manual alignment and visual inspection. Petition 870260060590, dated 06 / 22 / 2026, pp. 146 / 261 15 / 116
[0044] The algorithms that are suitable for determining sequence identity percentage and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res. 25: 3389-3402, respectively. The software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (NCBI) website. The algorithm involves first identifying high-score sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or satisfy some positive value threshold score T when aligned with a word of the same length in a sequence in databases. T is referred to as a neighborhood word threshold score (Altschul et al, supra).These initial neighborhood word matches act as seeds to initiate searches to find longer HSPs containing the same words. The word matches are then extended in both directions along each sequence as much as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a matching residue pair; always >0) and N (penalty score for non-matching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score.The extension of word hits in each direction is interrupted when: the cumulative alignment score falls by an amount X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more alignments of negative score residues; or the end of any sequence is reached. The W, T, and X parameters of the BLAST algorithm determine the sensitivity and... Petition 870260060590, dated 06 / 22 / 2026, pp. 147 / 261 16 / 116 the alignment speed. The BLASTN program (for nucleotide sequences) uses as standards a word size (W) of 28, an expectation (E) of 10, M = 1, N = -2 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as standards a word size (W) of 3, an expectation (E) of 10 and the BLOSUM62 scoring matrix. See Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1989).
[0045] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences. See, for example, Karlin & Altschul, Proc. Natl. Acad. Sci. USA 90: 5873-5787 (1993). One similarity measure provided by the BLAST algorithm is the least-sum probability (P(N)), which gives an indication of the probability that a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the least-sum probability in a comparison of the test nucleic acid with the reference nucleic acid is less than about 0.01, more preferably less than about 10⁻⁵, and most preferably less than about 10⁻²⁰.
[0046] “Recombination” is the exchange of DNA strands to produce new arrangements of nucleotide sequences. The term can refer to the homologous recombination process that occurs in the repair of double-strand DNA breaks, where one polynucleotide is used as a template to repair a homologous polynucleotide. The term can also refer to the exchange of information between two homologous chromosomes during meiosis. The frequency of double recombination is the product of the frequencies of single recombinants. For example, a recombinant in an area of 10 cM can be found with a frequency of 10%, and double recombinants are found with a frequency of 10% x 10% = 1% (1 centimorgan is defined as progeny). Petition 870260060590, dated 06 / 22 / 2026, pp. 148 / 261 17 / 116 recombinant at 1% in a crossover test).
[0047] A “gene” is a defined region located within a genome that, in addition to the aforementioned coding nucleic acid sequence, comprises other nucleic acid sequences, primarily regulatory, responsible for controlling the expression, i.e., the transcription and translation of the coding portion. Genes may include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein, including regulatory sequences. Genes may or may not be capable of being used to produce a functional protein. In some embodiments, a gene refers only to the coding region. The term “native gene” refers to a gene as found in nature.The term “chimeric gene” refers to any gene that contains 1) DNA sequences, including regulatory and coding sequences, that are not found together in nature or 2) sequences encoding protein parts that are not naturally joined or 3) promoter parts that are not naturally joined. Accordingly, a chimeric gene may comprise regulatory and coding sequences that are derived from different sources or comprise regulatory and coding sequences derived from the same source but arranged in a manner different from that found in nature. A gene may be “isolated,” which means a nucleic acid molecule that is substantially or essentially devoid of components normally found in association with the nucleic acid molecule in its natural state.These components include other cellular material, recombinant production culture medium, and / or various chemicals used in the chemical synthesis of the nucleic acid molecule. Petition 870260060590, dated 06 / 22 / 2026, pp. 149 / 261 18 / 116
[0048] A “gene of interest” or “nucleotide sequence of interest” refers to any gene that, when transferred to a plant, confers upon the plant a desirable characteristic such as antibiotic resistance, virus resistance, insect resistance, disease resistance or resistance to other pests, herbicide tolerance, improved nutritional value, improved performance in an industrial process or altered reproductive capacity. The “gene of interest” may also be one that is transferred to plants for the production of commercially valuable enzymes or metabolites in the plant.
[0049] An “isolated” nucleic acid molecule or nucleotide sequence or an “isolated” polypeptide is a nucleic acid molecule, nucleotide sequence, or polypeptide that, by human intervention, exists in addition to its native environment and / or has a function that is different, modified, modulated, and / or altered compared to its function in its native environment and is therefore not a product of nature. An isolated nucleic acid molecule or isolated polypeptide may exist in a purified form or may exist in a non-native environment such as, for example, a recombinant host cell. Thus, for example, with regard to polynucleotides, the term isolated means that it is separated from the chromosome and / or cell in which it naturally occurs.A polynucleotide is also considered isolated if it is separated from the chromosome and / or cell in which it naturally occurs and is then inserted into a genetic context, a chromosome, a chromosome location, and / or a cell in which it does not naturally occur. The recombinant nucleic acid molecules and nucleotide sequences of the invention may be considered to be “isolated” as defined above.
[0050] Thus, an “isolated nucleic acid molecule” or “isolated nucleotide sequence” is a nucleic acid molecule or Petition 870260060590, dated 06 / 22 / 2026, pp. 150 / 261 19 / 116 A nucleotide sequence that is not immediately contiguous with the nucleotide sequences with which it is immediately contiguous (one at the 5' end and one at the 3' end) in the naturally occurring genome of the organism from which it is derived. Accordingly, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding sequences (e.g., promoters) that are immediately contiguous to a coding sequence. The term therefore includes, for example, a recombinant nucleic acid that is incorporated into a vector, a plasmid, or a self-replicating virus, or into the genomic DNA of a prokaryote or eukaryote, or that exists as a separate molecule (e.g., a cDNA or a fragment of genomic DNA produced by PCR or treatment with restriction endonucleases), independently of other sequences.It also includes a recombinant nucleic acid that is part of a hybrid nucleic acid encoding an additional polypeptide or peptide sequence. An “isolated nucleic acid molecule” or “isolated nucleotide sequence” may also include a nucleotide sequence derived from and inserted into the same original, natural cell type, but present in an unnatural state, for example, present in a different number of copies and / or under the control of regulatory sequences different from those found in the native state of the nucleic acid molecule.
[0051] The term “isolate” may additionally refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide, or fragment that is substantially free of cellular material, viral material, and / or culture medium (e.g., when produced by recombinant DNA techniques) or chemical precursors or other chemicals (e.g., when chemically synthesized). Furthermore, an “isolated fragment” is a fragment of a nucleic acid molecule, nucleotide sequence, or polypeptide that Petition 870260060590, dated 06 / 22 / 2026, pp. 151 / 261 20 / 116 does not occur naturally as a fragment and would not be found as such in its natural state. “Isolated” does not necessarily mean that the preparation is technically pure (homogeneous), but it is pure enough to provide the polypeptide or nucleic acid in a form in which it can be used for its intended purpose.
[0052] “Homology-dependent repair” or “homology-directed repair” or “HDR” refers to a mechanism for repairing ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by the cell when an HDR template exists with a sequence that has significant homology to the lesion site. The term “perfect HDR” refers to a situation in which the genomic homology junctions in the substituted allele have undergone complete HDR, and “imperfect HDR” refers to a situation in which the genomic homology junctions in the substituted allele have undergone partial or incomplete HDR. In some embodiments, a donor polynucleotide molecule with homology to the cleaved target DNA sequence is used as a template for repairing the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. As such, new nucleic acid material can be inserted / copied at the site.In some cases, a target DNA is contacted with a donor molecule, for example a donor polynucleotide molecule. In some cases, a donor polynucleotide molecule is introduced into a cell. In some cases, at least one segment of a donor polynucleotide molecule integrates into the cell's genome.
[0053] “Microhomology-mediated end joining” or “MMEJ” or “alternative non-homologous end joining” (AltNHEJ) refers to a way of repairing double-strand breaks in DNA. This repair mechanism uses microhomologous sequences to align the broken strands. “Non-homologous end joining” or Petition 870260060590, dated 06 / 22 / 2026, pp. 152 / 261 21 / 116 “NHEJ” refers to a method of repairing double-strand breaks in DNA. Double-strand breaks are repaired by directly ligating the broken ends together. Generally, in the absence of a donor polynucleotide, no new nucleic acid material is inserted at the site, although some nucleic acid material may be lost or added, resulting in a small deletion or a small insertion. In some embodiments, a donor polynucleotide molecule can be provided (e.g., introduced into a cell) and a portion of the donor polynucleotide can be inserted into the genome via MMEJ or NHEJ. Some embodiments of the methods provided here increase the probability of donor polynucleotide insertion by tethering the donor polynucleotide to a target site, as described below. II. Introduction
[0054] Fusion proteins and associated recombinant nucleic acids, systems, and methods are provided here for increasing the efficiency of genome editing using SDNs and donor polynucleotide tethering methods. The description is based in part on the inventors' discovery that 1) fusing an SDN to a recruiter domain comprising a site-specific DNA binding domain and 2) using a homologous donor polynucleotide repair template comprising a binding site for the recruiter domain results in increased HDR frequency, as demonstrated in the Examples herein. Without being limited by any particular theory, it is possible that the recruiter domain binds to the binding site on the donor polynucleotide template and tethers the donor polynucleotide to the cleavage site (i.e., through fusion of the recruiter domain to the SDN that forms the cleavage). Additionally, it is possible that this tethering increases the probability of cleavage repair. Petition 870260060590, dated 06 / 22 / 2026, pp. 153 / 261 22 / 116 mediated by HDR (e.g., by promoting spatial proximity between the cleavage site and the donor polynucleotide template). III. Fusion proteins
[0055] In one aspect, fusion proteins comprising a site-directed nuclease linked to a recruiter domain comprising a site-specific DNA-binding domain are provided here. As used throughout the text, a “fusion protein” is a protein comprising two different polypeptide sequences, namely, a site-directed nuclease polypeptide sequence and a recruiter domain polypeptide sequence, which are joined or linked to form a single polypeptide. In some embodiments, the two amino acid sequences are encoded by separate nucleic acid sequences that have been joined so that they are transcribed and translated to produce a single polypeptide. The site-directed nuclease and the recruiter domain may be linked in any order and orientation relative to each other. For example, the C'-terminal end of the site-directed nuclease may be linked to the N'-terminal end or to the C'-terminal end of the recruiter domain.The site-directed nuclease and recruiter domain can also be separated by one or more additional fusion protein domains, as described below. A. Site-directed nuclease
[0056] The fusion proteins provided herein comprise a site-directed modifying polypeptide (e.g., a site-directed nuclease). A “site-directed modifying polypeptide” modifies target DNA (e.g., through cleavage or methylation of target DNA) and / or a polypeptide associated with target DNA (e.g., methylation or acetylation of a histone tail). In some embodiments, a site-directed modifying polypeptide interacts with a guide RNA, which is a single RNA molecule or an RNA duplex of at least Petition 870260060590, dated 06 / 22 / 2026, pp. 154 / 261 23 / 116 two RNA molecules, and is guided to a DNA sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, e.g., an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by virtue of its association with guide RNA. In some embodiments, the site-directed polypeptide is a site-directed nuclease, which is capable of cleaving one or both strands of DNA at a specified target sequence.
[0057] The term “cleavage” or “to cleave” refers to the breaking of the covalent phosphodiester bond in the ribosylphosphodiester backbone of a polynucleotide and encompasses both single-strand and double-strand breaks. Double-strand cleavage can occur as a result of two distinct single-strand cleavage events. Cleavage can result in the production of blunt ends or staggered ends (also known as cohesive ends). A “nuclease cleavage site” or “genomic nuclease cleavage site” is a region of nucleotides within which a site-directed nuclease cleaves (e.g., when bound to a proximal binding site). When the polynucleotide is DNA (e.g., genomic DNA), one or both strands may be cleaved at the nuclease cleavage site. This cleavage by the nuclease enzyme initiates DNA repair mechanisms within the cell, establishing an environment for homologous recombination to occur.
[0058] Several site-directed nucleases can be used in the fusion proteins, systems, and methods disclosed herein. Suitable nucleases include, but are not limited to, CRISPR-associated proteins (Cas) or Cas nucleases; zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); meganucleases; RNA-binding proteins (RBPs); CRISPR-associated RNA-binding proteins; recombinases; flippases; Petition 870260060590, dated 06 / 22 / 2026, pp. 155 / 261 24 / 116 transposases; Argonaut (Ago) proteins (e.g., prokaryotic Argonaut (pAgo), archaebacterial Argonaut (aAgo), eukaryotic Argonaut (eAgo), and Natronobacterium gregoryi Argonaut (NgAgo); RNA-acting adenosine deaminases (ADAR); CRISPR-Cas-inspired RNA targeting system (CIRT); Pumilio / fem-3 binding factor (PUF); homing endonuclease or any functional fragment thereof, any derivative thereof; any variant thereof; and any fragment thereof.
[0059] In some embodiments, the site-directed nuclease is a naturally occurring site-directed nuclease. Exemplary naturally occurring site-directed nucleases are known in the art (see, for example, Makarova et al., 2017, Cell 168: 328-328.e1, and Shmakov et al., 2017, Nat Rev Microbiol 15 (3): 169-182, both incorporated herein by reference). In some embodiments, a site-directed nuclease binds to a targeting polynucleotide DNA (e.g., a guide RNA) and is thereby directed to a specific sequence within a target DNA and cleaves the target DNA.
[0060] In some embodiments, the site-directed nuclease is modified from its natural sequence (e.g., through mutation of one or more amino acid residues) to change its function. For example, the site-directed nuclease can be modified to be enzymatically inactive. The term “enzymatically inactive” can refer to a site-directed nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner but cannot cleave a target polynucleotide. An enzymatically inactive site-directed polypeptide may comprise an enzymatically inactive domain (e.g., nuclease domain). Enzymatically inactive can refer to none Petition 870260060590, dated 06 / 22 / 2026, pp. 156 / 261 25 / 116 activity. Enzymatically inactive may refer to substantially no activity. Enzymatically inactive may refer to essentially no activity. Enzymatically inactive may refer to activity not greater than 1%, not greater than 2%, not greater than 3%, not greater than 4%, not greater than 5%, not greater than 6%, not greater than 7%, not greater than 8%, not greater than 9%, or not greater than 10% of activity compared to an exemplary wild-type activity (e.g., nucleic acid cleavage activity, wild-type Cas9 activity).
[0061] In some embodiments, the site-directed nuclease (e.g., an enzymatically inactive site-directed nuclease) is fused to one or more transcriptional repressor domains, activator domains, epigenetic domains, recombinase domains, transposase domains, flippase domains, nicase domains, cleavage domains, or any combination thereof. The activator domain may include one or more tandem activation domains located at the carboxyl terminus of the enzyme. In other cases, the actuator moiety includes one or more tandem repressor domains located at the carboxyl terminus of the protein. Non-limiting exemplary activation domains include GAL4, herpes simplex activation domain VP16, VP64 (a tetramer of the herpes simplex activation domain VP16), NF-KB p65 subunit, Epstein-Barr virus transactivator R (Rta) and are described in Chavez et al., Nat Methods, 2015, 12 (4): 326-328 and in U.S. Patent Application No. Publ. 20140068797.Exemplary, non-limiting repression domains include Koxl's KRAB (Kruppel-associated box) domain, the Mad mSIN3 interaction domain (SID), the ERF repressor domain (ERD), and are described in Chavez et al., Nat Methods, 2015, 12(4): 326-328 and in U.S. Patent Application No. Publ. 20140068797. A nuclease. Petition 870260060590, dated 06 / 22 / 2026, pp. 157 / 261 26 / 116 can also be fused to a heterologous polypeptide, providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the C-terminus, C-terminal, or internally within the nuclease. CRISPR / Cas nucleases
[0062] In some embodiments, the site-directed nuclease comprises a CRISPR-associated protein (Cas) or a Cas nuclease that functions in a CRISPR (Regularly Interspaced Short Palindromic Repeats) / Cas system. In bacteria, this system can provide adaptive immunity against foreign DNA (Barrangou, R., et al., “CRISPR provides acquired resistance against viruses in prokaryotes”, Science (2007) 315: 1709-1712; Makarova, KS, et al., “Evolution and classification of the CRISPR-Cas systems”, Nat Rev Microbiol (2011) 9: 467-477; Garneau, JE, et al., “The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA”, Nature (2010) 468: 67-71; Sapranauskas, R., et al., “The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli”, Nucleic Acids Res (2011) 39: 9275-9282). In a wide variety of organisms, including various mammals, animals, plants, microbes, and yeasts, a CRISPR / Cas system (e.g., modified and / or unmodified) can be used as a genome manipulation tool. A CRISPR / Cas system may comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing.An RNA-guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease) can specifically bind to a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave DNA (Gasiunas, G., et al., “Cas9-crRNA ribonucleoprotein complex mediates. Petition 870260060590, dated 06 / 22 / 2026, pp. 158 / 261 27 / 116 specific DNA cleavage for adaptive immunity in bacteria”, Proc Natl Acad Sci USA (2012) 109: E2579-E2 86; Jinek, M., et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity”, Science (2012) 337: 816-821; Sternberg, SH, et al., “DNA interrogation by CRISPR Cas9 RNA-directed endonuclease”, Nature (2014) 507: 62; Deltcheva, E., et al., “CRISPR RNA maturation by small RNA-encoded trans- and host factor RNase III”, Nature (2011) 471: 602-607). of DNA (e.g., double-strand breaks) can result in DNA break repair that allows the introduction of gene modification(s) (e.g., (nucleic acid editing). DNA break repair can occur through non-homologous end joining (NHEJ), micro-homology-mediated end joining (MMEJ), or homology-directed repair (HDR).In some embodiments, donor polynucleotides are used to promote HDR, as detailed below in the “Systems” section. CRISPR-Cas systems have been widely used for programmable genome editing in a variety of organisms and model systems (Cong, L., et al., “Multiplex genome engineering using CRISPR Cas systems”, Science (2013) 339: 819-823; Jiang, W., et al., “RNA-guided editing of bacterial genomes using CRISPR-Cas systems”, Nat. Biotechnol. (2013) 31: 233-239; Sander, JD & Joung, J. K, “CRISPR-Cas systems for editing, regulating and targeting genomes”, Nature Biotechnol. (2014) 32: 347-355).
[0063] In some embodiments, the site-directed nuclease described herein comprises a Cas protein that forms a complex with a guide nucleic acid, such as a guide RNA (described further in the “Systems” section). In some embodiments, the site-directed nuclease comprises a Cas protein that forms a complex with a single guide nucleic acid, such as a single guide RNA (sgRNA). In some embodiments, the site-directed nuclease comprises a Petition 870260060590, dated 06 / 22 / 2026, pp. 159 / 261 28 / 116 RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), which is capable of forming a complex with a Cas protein. In some cases, RNA-guided Cas proteins recognize DNA targets that are complementary to a portion of the gRNA known as a CRISPR RNA (crRNA) sequence. The target sequence is often referred to as a protospacer, and the portion of the crRNA sequence that is complementary to the protospacer is often referred to as a spacer. In order to function (e.g., to cleave DNA), many Cas nucleases also require a specific adjacent protospacer motif (PAM), usually a DNA sequence 2 to 6 base pairs immediately following the protospacer sequence.
[0064] Various site-directed Cas nucleases (e.g., proteins Cas proteins from different species may be useful in fusion proteins, systems, and methods provided here based on the various enzymatic characteristics of different Cas proteins (e.g., different protospacer adjacent motif (PAM) sequence preferences; increased or decreased enzymatic activity; increased or decreased level of cellular toxicity; propensity to result in one or more NHEJ, homology-directed repair, single-strand breaks, double-strand breaks, etc.). Cas proteins from various species (e.g., those disclosed in Shmakov et al., 2017, or polypeptides derived from them) may require different PAM sequences in the target DNA. Thus, for a particular Cas enzyme of choice, the PAM sequence requirement may be different from the 5'-NGG-3' sequence (where N is an A, T, C, or G) known to be required for Cas activity.Many Cas9 orthologs from a wide variety of species have been identified, and the proteins share only a few identical amino acids. All of them... Petition 870260060590, dated 06 / 22 / 2026, pp. 160 / 261 29 / 116 identified Cas9 orthologs share the same domain architecture with a central HNH endonuclease domain and a split RuvC / RNaseH domain. Cas9 proteins share 4 critical motifs with a conserved architecture; Motifs 1, 2, and 4 are RuvC-like motifs, while motif 3 is an HNH motif. In contrast, Cas12a proteins from various species may have different PAM sequence requirements compared to the canonical LbCas12a PAM from TTTV.
[0065] Any suitable CRISPR / Cas system can be used. A CRISPR / Cas system can be referred to using a variety of nomenclature systems. Exemplary nomenclature systems are provided in Makarova, KS et al., “An updated evolutionary classification of CRISPR-Cas systems”, Nat Rev Microbiol (2015) 13: 722-736 and Shmakov, S. et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems”, Mol Cell (2015) 60: 1-13. A CRISPR / Cas system can be a type I, type II, type III, type IV, type V, type VI, or any other suitable CRISPR / Cas system. A CRISPR / Cas system as used here can be a Class 1, Class 2, or any other appropriately classified CRISPR / Cas system. The determination of Class 1 or Class 2 can be based on the genes encoding the effector module.Class 1 systems generally have a multi-subunit crRNA-effector complex, while Class 2 systems generally have a single protein, such as Cas9, Cpfl (also referred to as Cas12a), C2c1, C2c2, C2c3, or a crRNA-effector complex. A Class 1 CRISPR / Cas system may use a complex of multiple Cas proteins to effect regulation. A Class 1 CRISPR / Cas system may comprise, for example, CRISPR / Cas type I (e.g., I, IA, IB, IC, ID, IE, IF, IU), type III (e.g., III, IIIA, IIIB, IIIC, IIID), and type IV (e.g., IV, IVA, IVB). Petition 870260060590, dated 06 / 22 / 2026, pp. 161 / 261 30 / 116 Class II CRISPR / Cas systems can use a single large Cas protein to effect regulation. A Class II CRISPR / Cas system may comprise, for example, type II CRISPR / Cas (e.g., II, IIA, IIB) and type V. CRISPR systems may be complementary to each other and / or may borrow trans functional units to facilitate CRISPR locus targeting.
[0066] A Cas protein can be from any suitable organism. Non-limiting examples include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas nap hthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicellulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans , Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp.Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some aspects, the organism is Streptococcus pyogenes. Petition 870260060590, dated 06 / 22 / 2026, pp. 162 / 261 31 / 116 (S. pyogenes). In some respects, the organism is Staphylococcus aureus (S. aureus). In some respects, the organism is Streptococcus thermophilus (S. thermophilus).
[0067] A Cas protein can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp.Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsoginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella. Petition 870260060590, dated 06 / 22 / 2026, pp. 163 / 261 32 / 116 pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.
[0068] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl (also referred to as Cas12a), Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4, and Cul966 and homologs or modified versions thereof. In some embodiments, the site-directed nuclease of the fusion proteins provided herein comprises a CRISPR-associated nuclease, wherein the CRISPR-associated nuclease is Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, or Cas14. In some embodiments, the CRISPR-associated nuclease is a Cas9 enzyme. In some embodiments, the CRISPR-associated nuclease is a Cas12a enzyme.In some embodiments, the CRISPR-associated nuclease is a nicase or a deactivated version of a CRISPR-associated nuclease.
[0069] Lachnospiraceae bacterium Cpf1 (LbCpf1) is one of many Cpfl proteins in a large group. The terms “Cpfl” and “Cas12a” are used interchangeably throughout this description. Cpf1 is a Cas protein. In some embodiments, the site-directed nuclease is a catalytically active Cas12a from Lachnospiraceae bacterium (“LbCas12a”) or Moraxella bovoculi AAX08_00205 (“Mb2Cas12a”). In some embodiments, the site-directed nuclease domain of the fusion protein is a Cas12a protein from any of the Lachnospiraceae bacterium, Acidaminococcus sp., Moraxella bovoculi, Thiomicrospira Petition 870260060590, dated 06 / 22 / 2026, pp. 164 / 261 33 / 116 sp., Moraxella lacunata, Methanomethylophilus alvus, Btyrivibrio sp., or Bacteroidetesoral sp.
[0070] A Cas protein may comprise one or more domains. Non-limiting examples of domains include guide nucleic acid recognition and / or binding domains, nuclease domains (e.g., DNase or RNase, RuvC, HNH domains), DNA binding domains, RNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. A guide nucleic acid recognition and / or binding domain may interact with a guide nucleic acid. A nuclease domain may comprise catalytic activity for nucleic acid cleavage. A nuclease domain may lack catalytic activity to prevent nucleic acid cleavage. A Cas protein may be a chimeric Cas protein that is fused to other proteins or polypeptides. A Cas protein may be a chimera of several Cas proteins, for example, comprising domains from different Cas proteins.
[0071] A Cas protein used herein may be an active variant, inactive variant, or fragment of a wild-type or modified Cas protein. A Cas protein may comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof, relative to a wild-type version of the Cas protein. A Cas protein may be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to an exemplary wild-type Cas protein. A Cas protein can be a polypeptide with at most approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity or sequence similarity. Petition 870260060590, dated 06 / 22 / 2026, pp. 165 / 261 34 / 116 with an exemplary wild-type Cas protein. The variants or fragments may comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity with a wild-type or modified Cas protein or a portion thereof. The variants or fragments may be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.
[0072] In some embodiments, a modified Cas protein has diminished function compared to the unmodified form. In some embodiments, a modified Cas protein is deficient in a function of the unmodified form. For example, a nuclease-deficient Cas protein retains the ability to bind to DNA but lacks or has reduced nucleic acid cleavage activity. A Cas nuclease (e.g., retaining wild-type nuclease activity, having reduced nuclease activity, and / or lacking nuclease activity) can function in a CRISPR / Cas system to regulate the level and / or activity of a target gene or protein (e.g., decrease, increase, or elimination). The Cas protein can bind to a target polynucleotide and prevent transcription by physical obstruction or edit a nucleic acid sequence to produce non-functional gene products.In some embodiments, the modified Cas protein has no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the function (e.g., nuclease activity) of the wild-type Cas protein (e.g., Cas9 from S. pyogenes). In some embodiments, the modified Cas protein has no substantial function of the wild-type Cas protein. Petition 870260060590, dated 06 / 22 / 2026, pp. 166 / 261 35 / 116 wild type. When a Cas protein is a modified form that lacks substantial nucleic acid cleavage activity, it may be referred to as enzymatically inactive and / or “dead” (abbreviated by “d”). A dead Cas protein (e.g., dCas, dCas9) may bind to a target polynucleotide but may not cleave the target polynucleotide. In some respects, a dead Cas protein is either a dead Cas9 protein or a dead Cas12a protein.
[0073] In some embodiments, a modified Cas protein can be a modified Cas “base editor.” Base editing allows direct, irreversible conversion of a target DNA base to another in a programmable manner, without requiring DNA cleavage or a donor polynucleotide molecule. For example, Komor et al. (2016, Nature 533: 420-424) teach a Cas9-cytidine deaminase fusion, where Cas9 was also manipulated to be inactivated and not induce double-strand DNA breaks. Additionally, Gaudelli et al. (2017, Nature, doi:10.1038 / nature24644) teach a catalytically defective Cas9 fused to an adenosine deaminase tRNA, which can mediate the conversion of an A / T to G / C in a target DNA sequence. Another class of manipulated Cas9 nucleases that can be used as the site-directed nuclease in the fusion proteins of this description are variants that can recognize a wide range of PAM sequences, including NG, GAA, and GAT (Hu et al., 2018, Nature, doi: 10.1038 / nature26155).
[0074] A Cas protein can be modified to optimize the regulation of gene expression. A Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of Petition 870260060590, dated 06 / 22 / 2026, pp. 167 / 261 36 / 116 Cas proteins can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the protein's function or to optimize (e.g., enhance or reduce) the Cas protein's activity to regulate gene expression.
[0075] One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein may be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. For example, in a Cas protein comprising at least two nuclease domains (e.g., Cas9), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, known as a nicase, may generate a single-strand break in a CRISPR RNA recognition sequence (crRNA) within a double-stranded DNA but not a double-strand break. Such a nicase may cleave the complementary strand or the non-complementary strand, but may not cleave both. In some embodiments, the targeting specificity of double-strand breaks is improved by targeting a nicase to opposite strands at two nearby loci. If a nicase cleaves the single strand at both loci, a double-strand break is formed and can be repaired via HR as described herein.If all the nuclease domains of a Cas protein (for example, RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are deleted or mutated, the resulting Cas protein may have a reduced or no ability to cleave both strands of double-stranded DNA. Zinc Finger Nucleases
[0076] In some embodiments, a site-directed nuclease suitable for use in the fusion proteins or methods described herein is a “zinc finger nuclease” or “ZFN”. ZFNs refer to a Petition 870260060590, dated 06 / 22 / 2026, pp. 168 / 261 37 / 116 Fusion between a cleavage domain, such as a Fokl cleavage domain, and at least one zinc finger motif (e.g., at least 2, 3, 4, or 5 zinc finger motifs) that can bind to polynucleotides such as DNA and RNA. Heterodimerization at certain positions on a polynucleotide of two individual ZFNs in a certain orientation and spacing can lead to cleavage of the polynucleotide. For example, ZFN binding to DNA can induce a double-strand break in the DNA. In order to allow two cleavage domains to dimerize and cleave DNA, two individual ZFNs can bind to opposite strands of DNA with their C-terminus a certain distance apart. In some cases, linking sequences between the zinc finger domain and the cleavage domain may require that the 5' edge of each binding site be separated by about 5-7 base pairs. In some cases, a cleavage domain is fused to the C-terminal of each zinc finger domain.Exemplary ZFNs include, but are not limited to, those described in Urnov et al., Nature Reviews Genetics, 2010, 11: 636-646; Gaj et al., Nat Methods, 2012, 9 (8): 8057; U.S. Patents Nos. 6,534,261; 6,607,882; 6,746,838; 6,794,136; 6,824,978; 6,866,997; 6,933,113; 6,979,539; 7,013,219; 7,030,215;. 7,220,719; 7,241,573; 7,241,574; 7,585,849; 7,595,376; 6,903,185; 6,479,626; and U.S. Publications Nos. 2003 / 0232410 and 2009 / 0203140.
[0077] In some embodiments, a nuclease comprising a ZFN can generate a double-strand break in a target polynucleotide, such as DNA. A double-strand break in DNA can result in DNA break repair that allows the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a polynucleotide repair template Petition 870260060590, dated 06 / 22 / 2026, pp. 169 / 261 38 / 116 A donor or template polynucleotide containing homology arms flanking target DNA sites can be provided. In some embodiments, a ZFN is a zinc finger nicase that induces site-specific single-strand DNA breaks or cuts, thus resulting in HR. Descriptions of zinc finger nicases are found, for example, in Ramirez et al., Nucl Acids Res, 2012, 40 (12): 5560-8; Kim et al., Genome Res, 2012, 22 (7): 1327-33. In some embodiments, a ZFN binds to a polynucleotide (e.g., DNA and / or RNA) but is unable to cleave the polynucleotide.
[0078] In some embodiments, the cleavage domain of a nuclease comprising a ZFN comprises a modified form of a wild-type cleavage domain. The modified form of the cleavage domain may comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid cleavage activity of the cleavage domain. For example, the modified form of the cleavage domain may have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid cleavage activity of the wild-type cleavage domain. The modified form of the cleavage domain may have no substantial nucleic acid cleavage activity. In some forms, the cleavage domain is enzymatically inactive. TAL-Effector Nucleases
[0079] In some embodiments, a site-directed nuclease suitable for use in the fusion proteins, systems, or methods described herein is a “TALEN” or “effector TAL nuclease.” TALENs refer to manipulated transcription activator-like effector nucleases that generally contain a central repeat domain. Petition 870260060590, dated 06 / 22 / 2026, pp. 170 / 261 39 / 116 in tandem DNA binding and a cleavage domain. TALENs can be produced by fusing an effector DNA-binding domain (TAL) to a DNA cleavage domain. In some cases, a tandem DNA-binding repeat comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. A transcription activator-like effector protein (TALE) can be fused to a nuclease such as a wild-type or mutated Fok1 endonuclease or to the catalytic domain of Fok1. Several mutations to Fok1 have been made for use in TALENs, which, for example, improve specificity or cleavage activity. Such TALENs can be manipulated to bind to any desired DNA sequence.TALENs can be used to generate gene modifications (e.g., nucleic acid sequence editing) by creating a double-strand break in a target DNA sequence, which in turn undergoes NHEJ or HR. A double-strand break in DNA can result in DNA break repair that allows the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur through non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor polynucleotide repair template or template polynucleotide containing homology arms flanking target DNA sites can be provided. In some cases, a single-stranded donor polynucleotide repair template is provided to promote HR. Detailed descriptions of TALENs and their uses for gene editing are found, for example, in U.S. Pat. Nos. 8,440,431; 8,440,432; 8,450,471; 8,586,363; and. 8,697,853; Scharenberg et al., Curr Gene Ther, 2013, 13 (4): 291-303; Gaj et al., Nat Methods, 2012, 9 (8): 805-7; Beurdeley et al., Nat Petition 870260060590, dated 06 / 22 / 2026, pp. 171 / 261 40 / 116 Commun, 2013, 4: 1762; and Joung and Sander, Nat Rev Mol Cell Biol, 2013, 14 (1): 49-55.
[0080] In some embodiments, a TALEN is manipulated for reduced nuclease activity. In some embodiments, the nuclease domain of a TALEN comprises a modified form of a wild-type nuclease domain. The modified form of the nuclease domain may comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, not more than 20%, not more than 10%, not more than 5%, or not more than 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may have no substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is enzymatically inactive.
[0081] In some embodiments, the transcription activator-like effector protein (TALE) is fused to a domain that can modulate transcription and does not comprise a nuclease. In some embodiments, the transcription activator-like effector protein (TALE) is designed to function as a transcriptional activator. In some embodiments, the transcription activator-like effector protein (TALE) is designed to function as a transcriptional repressor. For example, the DNA-binding domain of the transcription activator-like effector protein (TALE) may be fused (e.g., linked) to one or more transcriptional activation domains or to one or more transcriptional repression domains. Non-limiting examples of a transcriptional activation domain include a domain Petition 870260060590, dated 06 / 22 / 2026, pp. 172 / 261 41 / 116 activation of VP16 in herpes simplex and a tetrameric repeat of the VP16 activation domain, for example, a VP64 activation domain. A non-limiting example of a transcriptional repression domain includes a Kruppel-associated box domain. Meganucleases
[0082] In some embodiments, a site-directed nuclease suitable for use in the fusion proteins, systems, or methods described herein is a meganuclease. Meganucleases generally refer to rare-cutting endonucleases or homing endonucleases that can be highly specific. Meganucleases can recognize target DNA sites ranging from at least 12 base pairs in length, for example, 12 to 40 base pairs, 12 to 50 base pairs, or 12 to 60 base pairs in length. Meganucleases can be modular DNA-binding nucleases such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA- or protein-binding domain specifying a target nucleic acid sequence. The DNA-binding domain may contain at least one motif that recognizes single-stranded or double-stranded DNA. A meganuclease can generate a double-strand break.A double-strand break in DNA can result in DNA break repair that allows the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur through non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor polynucleotide template containing homology arms flanking target DNA sites can be provided. The meganuclease can be monomeric or dimeric. In some embodiments, the meganuclease is naturally occurring (found in nature) or wild-type, and in other cases, the meganuclease is non-natural, artificial, or engineered. Petition 870260060590, dated 06 / 22 / 2026, pp. 173 / 261 42 / 116 synthetic, rationally designed or man-made. In some embodiments, the meganuclease of the present description includes an I-CreI meganuclease, I-CeuI meganuclease, I-Msol meganuclease, I-SceI meganuclease, their variants, their derivatives and their fragments. Detailed descriptions of useful meganucleases and their application in gene editing are found, for example, in Silva et al., Curr Gene Ther, 2011, 11 (1): 11-27; Zaslavoskiy et al., BMC Bioinformatics, 2014, 15: 191; Takeuchi et al., Proc Natl Acad Sci USA, 2014, 111 (11): 4061-4066, and US Patent Nos. 7,842,489; 7,897,372; 8,021,867; 8,163,514; 8,133,697; 8,021,867; 8,119,361; 8,119,381; 8,124,36; and 8,129,134.
[0083] In some embodiments, the nuclease domain of a meganuclease comprises a modified form of a wild-type nuclease domain. The modified form of the nuclease domain may comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may have no substantial nucleic acid cleavage activity. In some forms, the nuclease domain is enzymatically inactive.In some forms, a meganuclease can bind to DNA but cannot cleave the DNA. B. Recruiter domain
[0084] The fusion proteins provided herein comprise a recruiting domain comprising a DNA-binding domain. Petition 870260060590, dated 06 / 22 / 2026, pp. 174 / 261 43 / 116 site-specific. A recruiting domain of fusion proteins here can comprise any polypeptide with a unique recognition motif.
[0085] In some embodiments, the recruiting domain is a Cro repressor family protein. Cro repressor family proteins are bacteriophage transcription factors that function as homodimers. See, for example, MS Dubrava, et al., N15 Cro and λ Cro: Orthologous DNA-binding domains with completely different but equally effective homodimer interfaces, Protein Science 17: 803-812 (2008). In some embodiments, the recruiting domain comprises an N15 Cro protein, a lambda Cro protein, a P22 Cro protein (see, for example, AR Poteete, et al., Bacteriophage P22 Cro Protein: Sequence, Purification, and Properties, Biochemistry 25: 251-256 (1986)), a 434 Cro protein (see, for example, C. Wolberger, et al., Structure of a phage 434 Cro / DNA complex, Nature 335: 789-795 (1988)), or a combination thereof. Additional information on Cro family proteins of N15, lambda, P22, and 434 phages can be found in BM Hall, et al., Extreme divergence between one-to-one orthologs: the structure of N15 Cro bound to operator DNA and its relationship to the λ Cro complex, Nucleic Acids Research, 47 (13): 7118-7129 (2019).
[0086] In some embodiments, the recruiting domain of a fusion protein provided herein comprises all or most of the polypeptide sequence of a naturally occurring DNA-binding protein. In some embodiments, the recruiting domain comprises only the DNA-binding domain of a naturally occurring DNA-binding protein. In some embodiments, the recruiting domain comprises one or more modifications (e.g., as described in the “Variants” section below) relative to the protein from which they are derived. In some embodiments, the domain Petition 870260060590, dated 06 / 22 / 2026, pp. 175 / 261 44 / 116 recruiter comprises a synthetic DNA-binding polypeptide sequence.
[0087] In some embodiments of the fusion proteins provided herein, the recruiting domain functions (i.e., binds to a specific DNA sequence) as an oligomer. In some embodiments, the recruiting domain functions as a homo-oligomer. In some embodiments, the recruiting domain functions as a homodimer, homotrimer, homotetramer, or a higher-order homo-oligomer. In such embodiments, a fusion protein provided herein may comprise more than one monomer (e.g., two, three, four, or more monomers) of the recruiting domain. In some embodiments, the monomers are separated by a linker (e.g., any of the linkers described herein). For example, Cro repressor family proteins generally function as homodimers. In some embodiments, the fusion proteins provided herein comprise two or more monomers of a recruiting domain sequence of Cro repressor family proteins.In some embodiments, the two or more monomers are connected by flexible linkers, allowing the monomers to interact and form homo-oligomers. Exemplary fusion proteins comprising two monomers of Cro repressor family proteins are described in the Examples herein and illustrated in FIG. 4 and FIG. 5.
[0088] In some embodiments, the recruiting domain functions as a hetero-oligomer. In such embodiments, a fusion protein provided herein may comprise at least one monomer of two or more recruiting domain proteins.
[0089] In some embodiments, the recruiting domain comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, Petition 870260060590, dated 06 / 22 / 2026, pp. 176 / 261 45 / 116 at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity with any of the SEQ ID Nos: 1-4. C. Additional fusion protein domains
[0090] In some embodiments, the fusion proteins provided herein comprise one or more linkers. Linkers, also referred to as spacers, as used herein, are flexible molecules or a flexible stretch of molecules that joins or connects two portions (e.g., domains) of a fusion protein or a modified protein as provided herein. In some embodiments, the linker is a polypeptide. Proteins with domains linked by polypeptide linkers are referred to as fusion proteins. In some embodiments, the linker is a non-peptide linker. Proteins with domains linked by polypeptide linkers are referred to as modified proteins. It will be understood that where fusion proteins are discussed throughout the present description, modified proteins are also generally contemplated, where feasible.
[0091] The linker can increase the range of orientations that can be adopted by the domains of the fusion protein or modified protein. The linker can be optimized to produce desired effects on the fusion protein or modified protein. Aspects of linker design and considerations are described, for example, in Chen, X. et al., Adv Drug Deliv Rev. October 15, 2013; 65(10):1357-1369, and Klein, JS et al. 2014 Protein Eng. Des. Sel. 27(10):325-330. In some embodiments, the proteins provided herein comprise a peptide linker. In some embodiments, the proteins provided herein comprise a non-peptide linker. In some embodiments, the proteins provided herein comprise Petition 870260060590, dated 06 / 22 / 2026, pp. 177 / 261 46 / 116 a non-peptidic ligand and a non-peptidic ligand. The proteins provided herein may also comprise a plurality of ligands, including at least one peptide ligand, at least one non-peptidic ligand, or at least one peptide ligand and at least one non-peptidic ligand.
[0092] Binders can be short or long, flexible or rigid. See, for example, PCT / US2020 / 051383 incorporated herein by reference in its entirety, and WO 2020 / 168102, incorporated herein by reference in its entirety, and US 2021 / 0017506, incorporated herein by reference in its entirety.
[0093] In some embodiments, the length of a ligand can affect one or more functions of the fusion protein. The selection of ligands to achieve the desired length is within the capabilities of a person skilled in the art. In some embodiments, a peptide ligand may have, for example, 5 to 100 or more amino acids in length (e.g., 5 aa, 10 aa, 15 aa, 20 aa, 25 aa, 30 aa, 35 aa, 40 aa, 45 aa, 50 aa, 55 aa, 60 aa, 65 aa, 70 aa, 75 aa, 80 aa, 85 aa, 90 aa, 95 aa, or 100 aa).
[0094] Depending on its length, the linker sequence can have various conformations in its secondary structure, such as helical, β-strand, spiral / curve, and turns. In some cases, a linker sequence may have an extended conformation and function as an independent domain that does not interact with adjacent protein domains. Linker sequences can be flexible or rigid. Flexible linkers provide a degree of movement or interaction between polypeptide domains and are generally rich in small or polar amino acids such as Gly and Ser (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or all amino acid residues of the linker are Gly or Ser). A rigid linker can be used to maintain a fixed distance between domains and to Petition 870260060590, dated 06 / 22 / 2026, pp. 178 / 261 47 / 116 help maintain their independent functions. The attachment of the ligand can be through an amide bond (e.g., a peptide bond) or other functionalities as further discussed below.
[0095] In some embodiments, a peptide linker described herein comprises an amino acid sequence with at least 90% sequence identity with any of the SEQ ID NOs: 6, 7, 15, or 16. In some embodiments, the linker comprises an XTEN linker sequence. See, for example, X. Li, et al., Base editing with a Cpf1cytidine deaminase fusion, Nature Biotechnology 36: 324-327 (2018); Y. Zong, et al., Precise base editing in rice, wheat, and maize with a Cas9-cytidine deaminase fusion, Nature Biotechnology 35: 438-440 (2017); and V. Schellenberger, et al., A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a more finely tuned way, Nature Biotechnology 27: 1189-1190 (2009). In some modalities, the binder comprises one or more repetitions (e.g., 2 repetitions, 3 repetitions, 4 repetitions, 5 repetitions, 6 repetitions or more) of SGGS (SEQ ID NO: 22), GGGS (SEQ ID NO: 23), GGGGS (SEQ ID NO: 24) and / or one or more repetitions of GSSGSS (SEQ ID NO: 25).Additional exemplary peptide ligands include, but are not limited to, peptide ligands comprising SGSETPGTSESATPE (“XTEN03”, SEQ ID NO: 27), SGSETPGTSESATPES (“XTEN-01” SEQ ID NO: 28), SGSETPGTSESATPELK (“XTEN-02,” SEQ ID NO: 29), (GGGGS)3 (SEQ ID NO: 30), (GGGGS)5 (SEQ ID NO: 31), (GGGGGS)10 (SEQ ID NO: 32), GGGGGGGG (SEQ ID NO: 33), GSAGSAAGSGEF (SEQ ID NO: 34), A(EAAAK)3A (SEQ ID NO: 35), or A(EAAAK)10A (SEQ ID NO: 36). PCT / US2020 / 051383, Chen et al., Adv. Drug Deliv. Petition 870260060590, dated 06 / 22 / 2026, pp. 179 / 261 48 / 116 both of which are incorporated herein by reference.
[0096] In some embodiments, a non-peptidic linker may comprise any of a number of known chemical linkers. Exemplary chemical linkers may include one or more beta-alanine units, 4-aminobutyric acid (GABA), (2-aminoethoxy)acetic acid (AEA), 5-aminohexanoic acid (Ahx), PEG multimers, and trioxatridecanosuccinamic acid (Ttds). In some embodiments, the non-peptidic linker comprises one or more polyethylene glycol (PEG) units, which is commonly used as a linker for conjugation of polypeptide domains due to its water solubility, lack of toxicity, low immunogenicity, and well-defined chain lengths. See, for example, RamirezPaz, J., et al., PLoS One 13 (7): e0197643 (2018). The number of PEG binding units can be selected based on the desired binder length.
[0097] Modified proteins comprising a non-peptidic linker can be produced in a variety of ways. For example, a site-directed nuclease and a recruiting domain can be produced separately (e.g., in vitro or by expression and purification of host cells) and chemically linked in vitro. In some embodiments, a site-directed nuclease, a recruiting domain, and a linker can each be produced separately and chemically linked in vitro. Various chemical linkers can be used to crosslink two amino acid residues.
[0098] Also contemplated here are modalities in which the site-directed nuclease and recruiting domain as described above are used separately (e.g., introduced into cells separately or applied to target nucleic acids separately) and brought into proximity to form a complex without using ligands as described above. Several methods of complex formation between Petition 870260060590, dated 06 / 22 / 2026, pp. 180 / 261 49 / 116 Two or more polypeptides are known in the art and include, but are not limited to, using protein-protein interaction strategies (e.g., SunTag, coiled spiral, etc.), using RNA aptamers and associated binding proteins (e.g., MS2, N22, etc.), and Tag:Catcher strategies. For example, a site-directed nuclease of the present description may comprise an MS2 RNA aptamer, which would facilitate interaction with a recruiting domain comprising an MS2 coat protein.
[0099] In some embodiments, the fusion proteins provided herein comprise a targeting sequence that mediates the localization (or retention) of a protein at a subcellular location, for example, plasma membrane or membrane of a given organelle, nucleus, cytosol, mitochondria, endoplasmic reticulum (ER), Golgi, chloroplast, apoplast, peroxisome, or other organelle. For example, a targeting sequence may direct a protein (e.g., a nuclease) to a nucleus using a nuclear localization signal (NLS); out of a cell nucleus, for example to the cytoplasm, using a nuclear export signal (NES); mitochondria using a mitochondrial targeting signal; the endoplasmic reticulum (ER) using an ER retention signal; a peroxisome using a peroxisomal targeting signal; plasma membrane using a membrane localization signal; or combinations thereof.In some embodiments, the fusion protein comprises a nuclear localization signal. Non-limiting examples of NLSs include an NLS sequence derived from: the SV40 virus large T antigen NLS, having the amino acid sequence PKKKRKV (SEQ ID NO: 37); the nucleoplasmin NLS (e.g., the split nucleoplasmin NLS with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 38)); the c-myc NLS having the amino acid sequence PAAKRVKLD. Petition 870260060590, dated 06 / 22 / 2026, pp. 181 / 261 50 / 116 (SEQ ID NO: 39) or RQRRNELKRSP (SEQ ID NO: 40); the NLS of hRNPAI M9 having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 41); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 42) of the importin-alpha IBB domain; the VSRKRPRP (SEQ ID NO: 43) and PPKKARED (SEQ ID NO: 44) sequences of the myoma T protein; the PQPKKKPL (SEQ ID NO: 45) sequence of human p53; the SALIKKKKKMAP (SEQ ID NO: 46) sequence of mouse c-abl IV; the DRLRR (SEQ ID NO: 47) and PKQKKRK (SEQ ID NO: 48) sequences of influenza virus NS1; the RKLKKKIKKL (SEQ ID NO: 49) sequence of hepatitis virus delta antigen; the REKKKFLKRR (SEQ ID NO: 50) sequence of mouse Mx1 protein; the KRKGDEVDGVDEVAKKKSKK sequence (SEQ ID NO: 51) of human poly(ADP-ribose) polymerase; the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 52) of glucocorticoid (human) steroid hormone receptors; and the KRPRDRHDGELGGRKRAR (SEQ ID NO: 53) sequence of the Agrobacterium VirD2 protein.
[00100] In some embodiments, the fusion protein provided herein comprises an amino acid sequence having at least 70% (for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identity with SEQ ID NOS: 11 or 13.
[00101] Any of the polypeptides and fusion proteins described herein may additionally comprise a detectable fraction, for example, a fluorescent protein or fragment thereof. Examples of fluorescent proteins include, but are not limited to Petition 870260060590, dated 06 / 22 / 2026, pp. 182 / 261 51 / 116 limited to yellow fluorescent protein (YFP, e.g., Venus), green fluorescent protein (GFP), and red fluorescent protein (RFP), as well as derivatives, e.g., mutant derivatives, of these proteins. See, for example, Chudakov et al. “Fluorescent Proteins and Their Applications in Imaging Living Cells and Tissues”, Physiological Reviews 90 (3): 1103-1163 (2010); and Specht et al., “A Critical and Comparative Review of Fluorescent Tools for Live-Cell Imaging”, Annual Review of Physiology 79: 93-117 (2017)).
[00102] Any of the polypeptides described herein may additionally comprise an affinity label, for example, a polyhistidine label (e.g., (His)6 (SEQ ID NO: 54)), an HA label (e.g., YPYDVPDYA (SEQ ID NO: 55)), albumin-binding protein, alkaline phosphatase, an AU1 epitope, an AU5 epitope, a biotin-carboxy carrier protein (BCCP), a FLAG epitope (e.g., DYKDDDDK (SEQ ID NO: 56)) or a MYC epitope (e.g., EQKLISEEDL (SEQ ID NO: 57)), to name a few. See Kimple et al. “Overview of Affinity Tags for Protein Purification”, Curr. Protoc. Protein Sci. 73: Unit-9.9 (2013). D. Variants
[00103] Variants of the polypeptides described herein are also provided. The polypeptide variants retain their respective biological activity unless explicitly noted otherwise. For example, variants of a site-directed nuclease polypeptide retain the biological function of the full-length, native-sequence site-directed nuclease. In another example, variants of the recruiter domain retain the biological function of the full-length, native-sequence recruiter domain.
[00104] Modifications to any of the polypeptides or proteins provided here are made by known methods. For example, modifications are made by site mutagenesis. Petition 870260060590, dated 06 / 22 / 2026, pp. 183 / 261 52 / 116 specific nucleotides in a nucleic acid encoding the polypeptide, thereby producing a DNA encoding the modification and subsequently expressing the DNA in recombinant cell culture to produce the encoded polypeptide. Techniques for making substitution mutations at predetermined locations in DNA having a known sequence are well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations. Any of the nucleic acid sequences provided here can have codon optimization to alter, for example, to maximize expression, in a host cell or organism.
[00105] The amino acids in the polypeptides described herein may be any of the 20 naturally occurring amino acids, destereoisomers of naturally occurring amino acids, non-natural amino acids, and chemically modified amino acids. Non-natural amino acids (i.e., those not found naturally in proteins) are also known in the art, as presented in, for example, Zhang et al. “Protein engineering with non-natural amino acids”, Curr. Opin. Struct. Biol. 23 (4): 581-587 (2013); Xie et al. “Adding amino acids to the genetic repertoire”, 9 (6): 548-54 (2005)); and all references cited therein. β and γ amino acids are known in the art and are also contemplated herein as non-natural amino acids.
[00106] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, a side chain may be modified to comprise a signaling moiety, such as a fluorophore or a radio tag. A side chain may also be modified to comprise a new functional group, such as a thiol, carboxylic acid, or amino group. Post-translationally modified amino acids Petition 870260060590, dated 06 / 22 / 2026, pp. 184 / 261 53 / 116 are also included in the definition of chemically modified amino acids.
[00107] Conservative amino acid substitutions are also contemplated. By way of example, conservative amino acid substitutions can be made at one or more amino acid residues, for example, at one or more lysine residues of any of the polypeptides provided herein. A person skilled in the art would know that a conservative substitution is the substitution of an amino acid residue for another that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions with each other: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), Methionine (M).
[00108] By way of example, when an arginine for serine is mentioned, a conservative substitution for serine (e.g., threonine) is also contemplated. Non-conservative substitutions, for example, replacing a lysine with an asparagine, are also contemplated. IV. Recombinant nucleic acids, constructs, vectors and host cells
[00109] Recombinant nucleic acids encoding any of the polypeptides described herein are also provided herein. For example, a recombinant nucleic acid encoding a polypeptide that has at least 70% (e.g., at least 75%), Petition 870260060590, dated 06 / 22 / 2026, pp. 185 / 261 54 / 116 at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity with SEQ ID NOs: 12 or 14 is also provided. Recombinant nucleic acids having at least 70% identity with either SEQ ID NOs: 12 or 14 are also provided.
[00110] A DNA construct comprising a promoter operationally linked to a recombinant nucleic acid encoding a fusion protein or domains thereof as described herein is also provided. A nucleic acid is “operationally linked” when it is placed in a functional relationship with another nucleic acid sequence. Numerous promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the transcription start site that is involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.
[00111] The term “promoter” as used herein refers to a nucleotide sequence, usually upstream (5') of its coding sequence, that controls the expression of the coding sequence by providing recognition of RNA polymerase and other factors required for proper transcription. “Promoter regulatory sequences” consist of proximal and more distal upstream elements. Promoter regulatory sequences influence the transcription, processing, or stability of RNA or the translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. They include natural and synthetic sequences as well as sequences that may be a Petition 870260060590, dated 06 / 22 / 2026, pp. 186 / 261 55 / 116 combination of synthetic and natural sequences. An “enhancer” is a DNA sequence that can stimulate promoter activity and can be an innate promoter element or a heterologous element inserted to increase the level or tissue specificity of a promoter. It is capable of operating in both orientations (normal or inverted) and is able to function even when moved upstream or downstream of the promoter. The meaning of the term “promoter” includes “promoter regulatory sequences”.
[00112] The choice of promoters to be included depends on several factors, including, but not limited to, efficiency, selectivity, inducibility, desired expression level, and preferential expression in cells or tissues. It is a routine matter for an expert in the technique to modulate the expression of a sequence by appropriate selection and positioning of promoters and other regulatory regions relative to that sequence.
[00113] It has been shown that certain promoters are capable of directing RNA synthesis at a higher rate than others. These are called “strong promoters”. Certain other promoters have been shown to direct RNA synthesis to higher levels only in particular cell or tissue types and are often referred to as “tissue-specific promoters” or “tissue-preferred promoters” if the promoters preferentially direct RNA synthesis in certain tissues (RNA synthesis may occur in other tissues at reduced levels). Since the expression patterns of a chimeric gene (or genes) introduced into a plant are controlled using promoters, there is continued interest in isolating innovative promoters that are capable of controlling the expression of a chimeric gene (or genes) at certain levels in specific tissue types or at specific plant developmental stages. Petition 870260060590, dated 06 / 22 / 2026, pp. 187 / 261 56 / 116
[00114] Certain promoters are capable of directing RNA synthesis at relatively similar levels in all tissues of a plant. These are called “constitutive promoters” or “tissue-independent” promoters. Constitutive promoters can be divided into strong, moderate, and weak categories according to their effectiveness in directing RNA synthesis. Since it is often necessary to simultaneously express a chimeric gene (or genes) in different plant tissues to obtain the desired functions of the gene (or genes), constitutive promoters are especially useful in this respect.Although many constitutive promoters have been discovered from plants and plant viruses and characterized, there is still ongoing interest in isolating more innovative constitutive promoters, synthetic or native, that are capable of controlling the expression of a chimeric gene (or genes) at different levels and the expression of multiple genes in the same transgenic plant for gene stacking.
[00115] Among the most commonly used promoters are the nopaline synthase (NOS) promoter (Ebert et al., Proc. Natl. Acad. Sci. USA 84: 5745-5749 (1987)); the octapine synthase (OCS) promoter; kaulimovirus promoters such as the 19S promoter of cauliflower mosaic virus (CaMV) (Lawton et al., Plant Mol. Biol. 9: 315-324 (1987)); the light-inducible promoter of the small rubisco subunit (Pellegrineschi et al., Biochem. Soc. Trans. 23 (2): 247-250 (1995)); the Adh promoter (Walker et al., Proc. Natl. Acad. Sci. USA 84: 6624-66280 (1987)); the promoter of sucrose synthase (Yang et al., Proc. Natl. Acad. Sci. USA 87: 414-44148 (1990)); the promoter of the R gene complex (Chandler et al., Plant Cell 1: 1175-1183 (1989)); the promoter of the chlorophyll a / b binding protein gene; and similar ones.
[00116] Furthermore, it is contemplated that promoters combining elements from more than one promoter can be useful. Petition 870260060590, dated 06 / 22 / 2026, pp. 188 / 261 57 / 116 For example, U.S. Patent No. 5,491,288 discloses combining a cauliflower mosaic virus promoter with a histone promoter. Thus, the promoter elements disclosed here can be combined with elements from other promoters. Promoters that are useful for plant transgene expression include those that are inducible, viral, synthetic, constitutive (Odell Nature 313: 810-812 (1985)), temporally regulated, spatially regulated, tissue-specific, and spatially and temporally regulated. Using the regulatory elements described here, numerous agronomic genes can be expressed in transformed plants. More particularly, plants can be genetically manipulated to express various phenotypes of agronomic interest.
[00117] In some embodiments of the DNA constructs provided herein, the promoter may be a eukaryotic or prokaryotic promoter. In some embodiments, the promoter is an inducible promoter, a native inducible promoter (e.g., drought-inducible Rab17), a synthetic inducible promoter (e.g., auxin-inducible DR5, estradiol-inducible XVE / pLex, dexamethasone-inducible GVG / Gal4), a constitutive promoter (e.g., ZmUbq1, OsAct1, OsTub3, EF), an egg cell-specific promoter (e.g., EC1, EC2, EC3, EC4, EC5), a pollen-specific promoter, an apical meristem tissue-specific promoter, or a promoter with enriched expression in the zygote. In some embodiments, the promoter is a floral mosaic promoter (e.g., ZmBde1, OsAP1). In some embodiments, the promoter is a ubiquitin 4 promoter, an actin promoter, a tubulin promoter, a MADS box promoter, or a plant virus promoter.Suitable promoters are disclosed, for example, in U.S. Pat. No. 10,519,456, the entire contents of which are incorporated herein. Petition 870260060590, dated 06 / 22 / 2026, pp. 189 / 261 58 / 116 reference, and PCT / US2022 / 020690, incorporated herein by reference.
[00118] The recombinant nucleic acids provided herein may be included in expression cassettes for expression in a host cell or organism of interest. The cassette will include 5' and 3' regulatory sequences operationally linked to a recombinant nucleic acid provided herein that allows expression of a fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed in the cell or organism. Where additional genes or elements are included, the components are operationally linked. Alternatively, the additional gene(s) or element(s) may be provided in multiple expression cassettes. Such an expression cassette is provided with a plurality of restriction sites and / or recombination sites so that the insertion of polynucleotides is under the transcriptional regulation of the regulatory regions.The expression cassette may additionally contain a selectable marker gene. The expression cassette will include in the 5' to 3' direction of transcription: a transcriptional and translational initiation region (i.e., a promoter), a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region) that is functional in the cell or organism of interest. The promoters of the invention are capable of directing or driving the expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is subsequently translated to produce a protein) in a host cell. The regulatory regions (i.e., promoters, transcriptional regulatory regions, and translational termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, “heterologous” in. Petition 870260060590, dated 06 / 22 / 2026, pp. 190 / 261 59 / 116 reference to a sequence is a sequence that originates from an external species or, if from the same species, is substantially modified relative to its native form in genomic composition and / or locus by deliberate human intervention.
[00119] Additional regulatory signals include, but are not limited to, transcriptional initiation start sites, operators, activators, enhancers, other regulatory elements, ribosomal binding sites, an initiation codon, termination signals, and the like. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Handbook, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and the references cited therein.
[00120] The expression cassette may comprise a selectable marker gene for the selection of transformed cells. Marker genes include genes that confer antibiotic resistance, such as those conferring resistance to hygromycin, ampicillin, gentamicin, and neomycin, to name a few. Additional selectable markers are known and any of them may be used.
[00121] In preparing the expression cassette, the various DNA fragments can be manipulated to provide the DNA sequences in the appropriate orientation and, as appropriate, in the appropriate reading frame. For this purpose, adapters or linkers may be used to join the DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove superfluous DNA, remove restriction sites, or the like. For this purpose, in vitro mutagenesis, primer repair, restriction, pairing, re-substitutions, for example, transitions and Petition 870260060590, dated 06 / 22 / 2026, pp. 191 / 261 60 / 116 cross-sections.
[00122] In preparing the expression cassette, the various DNA fragments can be manipulated to provide the DNA sequences in the appropriate orientation and, as appropriate, in the appropriate reading frame. For this purpose, adapters or linkers can be used to join the DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove superfluous DNA, remove restriction sites, or the like. For this purpose, in vitro mutagenesis, primer repair, restriction, pairing, re-substitutions, for example, transitions and transversions, can be used.
[00123] A vector comprising a recombinant nucleic acid or DNA construct presented herein is further provided. The vector is contemplated to have the necessary functional elements that direct and regulate the transcription of the inserted nucleic acid. These functional elements include, but are not limited to, a promoter, upstream or downstream regions of the promoter such as enhancers that can regulate the transcriptional activity of the promoter, an origin of replication, restriction sites appropriate to facilitate the cloning of inserts adjacent to the promoter, antibiotic resistance genes or other markers that can serve to select cells containing the vector or the vector containing the insert, RNA splice junctions, a transcription termination region or any other region that can serve to facilitate the expression of the inserted gene or hybrid gene. See, generally, Sambrook et al. Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector, for example, could be a plasmid.
[00124] There are numerous E. coli expression vectors known to a person skilled in the art that are useful for the expression of Petition 870260060590, dated 06 / 22 / 2026, pp. 192 / 261 61 / 116 a nucleic acid. Other suitable microbial hosts for use include bacilli, such as Bacillus subtilis, and other Enterobacteriaceae, such as Salmonella, Senatia, and various Pseudomonas species. In these prokaryotic hosts, it is also possible to produce expression vectors, which will typically contain expression control sequences compatible with the host cell (e.g., an origin of replication). Additionally, any number of a variety of well-known promoters will be present, such as the lactose promoter system, a tryptophan (Trp) promoter system, a beta-lactamase promoter system, or a lambda phage promoter system. Additionally, yeast expression can be used. A nucleic acid encoding a polypeptide of the present invention is provided herein, wherein the nucleic acid can be expressed by a yeast cell. More specifically, the nucleic acid can be expressed by Pichia pastoris or S. cerevisiae.
[00125] Mammalian cells also allow the expression of proteins in an environment that favors important post-translational modifications such as cysteine folding and pairing, addition of complex carbohydrate structures, and secretion of active protein. Useful vectors for the expression of active proteins in mammalian cells are known in the art and may contain genes conferring resistance to hygromycin, resistance to geneticin or G418, or other genes or phenotypes suitable for use as selectable markers, or resistance to methotrexate for gene amplification. A number of suitable host cell lines capable of secreting intact human proteins have been developed in the art and include CHO cells, HeLa cells, HEK-293 cells, HEK-293T cells, U2OS cells, or any other primary or transformed cell line. Other suitable host cell lines include COS-7 cells, myeloma cell lines, Jurkat cells, Petition 870260060590, dated 06 / 22 / 2026, pp. 193 / 261 62 / 116 etc. Expression vectors for these cells may include expression control sequences, such as an origin of replication, a promoter, an enhancer, and necessary information processing sites, such as ribosome binding sites, RNA splice sites, polyadenylation sites, and transcriptional terminator sequences. Preferred expression control sequences are promoters derived from immunoglobulin genes, SV40, Adenovirus, Bovine Papillomavirus, etc.
[00126] The expression vectors described herein may also include nucleic acids as described herein under the control of an inducible promoter such as a tetracycline-inducible promoter or a glucocorticoid-inducible promoter. The nucleic acids of the present invention may also be under the control of a tissue-specific promoter to promote nucleic acid expression in specific cells, tissues, or organs. Any tunable promoter, such as a metallothionein promoter, a heat shock promoter, and other tunable promoters, of which many examples are well known in the art, are also contemplated. Furthermore, a Cre-loxP inducible system may also be used, as well as an Flp recombinase-inducible promoter system, both of which are known in the art.
[00127] Insect cells also allow the expression of polypeptides. Recombinant proteins produced in insect cells with baculovirus vectors undergo post-translational modifications similar to those of wild-type mammalian proteins.
[00128] Host cells comprising recombinant nucleic acids, DNA constructs and / or vectors described herein, as well as methods for producing such cells, are also provided herein. In some embodiments, the cell is a plant cell. In Petition 870260060590, dated 06 / 22 / 2026, pp. 194 / 261 63 / 116 In some forms, the plant cell is a corn plant cell, a soybean plant cell, a rice plant cell, a wheat plant cell and / or a sunflower plant cell.
[00129] A host cell comprising a nucleic acid or vector described herein is provided. The host cell may be an in vitro, ex vivo, or in vivo host cell. The host cells as provided herein are capable of expressing the fusion protein. Cell populations of any of the host cells described herein are also provided. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a recombinant nucleic acid encoding the fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a DNA construct encoding the fusion protein as described herein.In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a nucleic acid or a DNA construct encoding the fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a plurality of any of the host cells described herein. In some embodiments, a plurality of cells from any of the cell populations described herein expresses a fusion protein as described herein.
[00130] In some embodiments, the proportioned cells express the fusion protein stably or transiently. Stable expression of the fusion protein in a cell refers to the integration of any of the nucleic acids, DNA constructs, or vectors described herein into the cell's genome, thereby enabling the cell to express the fusion protein. Transient expression refers to Petition 870260060590, dated 06 / 22 / 2026, pp. 195 / 261 64 / 116 expression of the fusion protein directly from any of the nucleic acids, DNA constructs and / or vectors after introduction into the cell (i.e., the gene encoding the fusion protein is not integrated into the cell's genome).
[00131] In some embodiments, the proportioned cells express the fusion protein constitutively or inducibly. Constitutive expression refers to the continued, continuous expression of a gene (i.e., a protein), whereas inducible expression refers to gene (protein) expression that is responsive to a stimulus. Inducible expression is generally regulated through an inducible promoter, the description of which is included above.
[00132] A cell culture comprising one or more host cells described herein is also provided. Methods for the culture and production of many cells, including cells of bacterial (e.g., E. coli and other bacterial strains), animal (especially mammalian), and archaebacterial origin are available in the art. See, for example, Sambrook, supra; Ausubel, ed. (1995) Current Protocols in Molecular Biology, John Wiley & Sons, as well as Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3rd ed., Wiley-Liss, New York and the references cited therein; Doyle and Griffiths (1997) Mammalian Cell Culture: Essential Techniques, John Wiley and Sons, NY; Humason (1979) Animal Tissue Techniques, 4th ed., W.H. Freeman and Company; and Ricciardelli, et al., (1989) In vitro Cell Dev. Biol. 25: 1016-1024.
[00133] The host cell may be a prokaryotic cell, including, for example, a bacterial cell. Alternatively, the cell may be a eukaryotic cell, for example, a plant cell, a yeast cell, an insect cell, an avian cell, or a mammalian cell. The plant cell may be a field-cultured or greenhouse-cultured plant cell, including but not limited to Petition 870260060590, dated 06 / 22 / 2026, pp. 196 / 261 65 / 116 large-acre crop plants, fruits and vegetables, perennial tree plants, and ornamental plants. In some embodiments, the plant cell is a plant cell of sugarcane, pumpkin, corn (maize), wheat, rice, cassava, soybeans, hay, potatoes, cotton, tomatoes, alfalfa, and green algae. In some embodiments, the plant cell is a plant cell of any legume and vegetable, such as cabbage, turnip, carrot, parsnip, beetroot, lettuce, beans, broad beans, peas, potato, eggplant, tomato, cucumber, pumpkin, zucchini, onion, garlic, leek, pepper, spinach, yam, sweet potato, and cassava. In some embodiments, the plant cell is a plant cell of corn, soybeans, sunflower, tomato, rice, or wheat. In some embodiments, the mammalian cell may be a HEK-293T cell, a HEK293 cell, a Chinese hamster ovary (CHO) cell, a U2OS cell, a COS-7 cell, a HELA cell, or any other primary or transformed cell.A number of other suitable host cell lines have been developed and include myeloma cell lines, fibroblast cell lines, and a variety of tumor cell lines such as melanoma cell lines. Vectors containing the nucleic acid segments of interest can be transferred or introduced into the host cell by well-known methods, which vary depending on the type of host cell.
[00134] As used herein, the phrase “introduce” in the context of introducing a nucleic acid into a cell (e.g., a prokaryotic cell, a bacterial cell, a eukaryotic cell, a plant cell) refers to the translocation of the nucleic acid sequence from outside a cell to inside the cell. In some cases, introducing refers to the translocation of the nucleic acid from outside the cell to inside the cell nucleus. Where more than one nucleic acid molecule is to be introduced, these nucleic acid molecules may be Petition 870260060590, dated 06 / 22 / 2026, pp. 197 / 261 66 / 116 assembled as part of a single polynucleotide or nucleic acid construct, or as separate polynucleotide or nucleic acid constructs, and may be located in the same nucleic acid construct or in different nucleic acid constructs. Accordingly, such polynucleotides may be introduced into cells (e.g., plant cells) in a single transformation event, in separate transformation events, or, for example, as part of a breeding protocol.Several methods of introducing a nucleic acid into a cell are contemplated, including, but not limited to, electroporation, nanoparticle delivery, biolistic transformation, viral delivery, contact with nanowires or nanotubes, receptor-mediated internalization, translocation via cell-penetrating peptides, liposome-mediated translocation, DEAE dextran, lipofectamine, calcium phosphate, or any method now known or identified in the future for introducing nucleic acids into prokaryotic or eukaryotic cell hosts. A targeted nuclease system (e.g., an RNA-guided nuclease, a transcription activator-like effector nuclease (TALEN), a zinc finger nuclease (ZFN), or a megaTAL (MT)) can also be used to introduce a nucleic acid, for example, a nucleic acid encoding a fusion protein described herein, into a host cell. See Li et al. Signal Transduction and Targeted Therapy 5, Article No. 1 (2020).
[00135] The transformation of a cell can be stable or transient. Thus, a cell, plant cell, plant and / or transgenic plant part of the invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in genetically stable heritability. In some embodiments, the introduction into a plant, part Petition 870260060590, dated 06 / 22 / 2026, pp. 198 / 261 67 / 116 of plant and / or plant cell is through bacterial-mediated transformation, particle bombardment transformation, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, liposome-mediated transformation, nanoparticle-mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, silicon carbide fiber-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation or any other electrical, chemical, physical and / or biological mechanism that results in the introduction of nucleic acid into the plant, plant part and / or its cell or any combination thereof.
[00136] Plant transformation procedures are well known and routine in the art and are described throughout the literature. Non-limiting examples of plant transformation methods include transformation via bacterial-mediated nucleic acid delivery (e.g., via bacteria of the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid capillary crystal-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, as well as any other electrical, chemical, physical (mechanical) and / or biological mechanism that results in the introduction of nucleic acid into the plant cell, including any combination thereof.General guidelines for various known plant transformation methods include Miki et al. (“Procedures for Introducing Foreign DNA into Plants” in Methods in Plant Molecular Biology and Biotechnology, Glick, BR. Petition 870260060590, dated 06 / 22 / 2026, pp. 199 / 261 68 / 116 and Thompson, JE, Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7: 849-858 (2002)).
[00137] Agrobacterium-mediated transformation is a commonly used method for transforming plants due to its high transformation efficiency and its wide utility with many different species. Agrobacterium-mediated transformation typically involves transferring the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain that may depend on the complement of genes carried by the host Agrobacterium strain in a co-resident Ti plasmid or chromosomally (Uknes et al. 1993, Plant Cell 5: 159-169). Transfer of the recombinant binary vector to Agrobacterium can be achieved by a triparental mating procedure using Escherichia coli carrying the recombinant binary vector, a strain of E.A helper E. coli that carries a plasmid capable of mobilizing the recombinant binary vector to the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation (Hofgen and Willmitzer 1988, Nucleic Acids Res 16: 9877).
[00138] Transformation of a plant by recombinant Agrobacterium usually involves co-culturing Agrobacterium with plant explants and follows well-known methods in the art. The transformed tissue is typically regenerated in selection medium carrying a marker of antibiotic or herbicide resistance between the T-DNA boundaries of the binary plasmid.
[00139] Another method for transforming plants, plant parts, and plant cells involves propelling inert or biologically active particles into plant tissues and cells. See, for example, U.S. Patents Nos. 4,945,050; 5,036,006 and 5,100,792. Generally, this method involves propelling inert particles or Petition 870260060590, dated 06 / 22 / 2026, pages 200 / 261 69 / 116 biologically active particles for plant cells under conditions effective in penetrating the outer surface of the cell and resulting in incorporation within the cell. When inert particles are used, the vector can be introduced into the cell by coating the particles with the vector containing the nucleic acid of interest. Alternatively, a cell or cells can be surrounded by the vector such that the vector is transported into the cell by the particle's trail. Biologically active particles (e.g., dried yeast cells, dried bacteria, or a bacteriophage, each containing one or more nucleic acids to be introduced) can also be propelled into plant tissue.As used herein, the phrase “biological transformation” refers to a method of introducing RNA or DNA into cells (e.g., plant cells) directly, in which RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., plant cell) using high-velocity pressure to allow the RNA or DNA to penetrate the cell (e.g., penetrate the plant cell wall).
[00140] The CRISPR / Cas system can also be used to edit the genome of a host cell or organism. As detailed above, the “CRISPR / Cas” system refers to a widespread class of bacterial systems for defense against foreign nucleic acid. Any of the components of the CRISPR / Cas system described herein can be used to introduce fusion proteins, recombinant nucleic acids, or systems into the genome of a host cell or organism. Methods for genome editing mediated by the CRISPR / Cas system are known in the art. It will be understood that the use of a CRISPR / Cas system for introducing fusion proteins, recombinant nucleic acids, or systems described herein into the genome of a host cell or organism is different from the methods and systems described herein. Petition 870260060590, dated 06 / 22 / 2026, pages 201 / 261 70 / 116 private individuals provided here.
[00141] Any of the fusion proteins described herein can be purified or isolated from a host cell or population of host cells. For example, a recombinant nucleic acid encoding any of the fusion proteins described herein can be introduced into a host cell under conditions that permit expression of the fusion protein. In some embodiments, the recombinant nucleic acid has codon optimization for expression. After expression in the host cell, the fusion protein can be isolated or purified using purification methods known in the art. V. Systems
[00142] In another aspect, useful systems for editing one or more nucleic acids are provided herein. The systems comprise one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, the systems additionally comprise one or more additional elements that are useful for editing one or more nucleic acids. For example, a system provided herein may further comprise a donor polynucleotide. As another example, a system comprising a fusion protein comprising a Cas nuclease may further comprise one or more guide nucleic acids and / or one or more donor polynucleotide sequences. Donor polynucleotides and guide nucleic acids are detailed below. The systems provided herein are useful for carrying out the methods described in Section VI of this description. A. Donor polynucleotides
[00143] The systems and methods of the present description may comprise a donor polynucleotide. A “donor polynucleotide”, “donor molecule” or “donor template” is a polymer or oligomer of Petition 870260060590, dated 06 / 22 / 2026, pages 202 / 261 71 / 116 nucleotides intended for insertion into a target polynucleotide, typically a target genomic site. The donor sequence may be one or more transgenes, expression cassettes, or nucleotide sequences of interest. A donor molecule may be a donor DNA molecule, either single-stranded, partially double-stranded, or double-stranded. The donor polynucleotide may be a natural or modified polynucleotide, an RNA-DNA chimera, or a DNA fragment, either single-stranded or at least partially double-stranded, or a fully double-stranded DNA molecule, or a PGR-amplified ssDNA or at least partially double-stranded dsDNA fragment. In some embodiments, the donor DNA molecule is part of a circularized DNA molecule. In some cases, a fully double-stranded donor DNA may provide increased stability because dsDNA fragments are generally more resistant than ssDNA to degradation by nucleases.
[00144] In some embodiments, the donor polynucleotide comprises at least one recruiting sequence to which the recruiting domain of at least one fusion protein provided herein binds. In some embodiments, the donor polynucleotide comprises two, three, four, five, six, seven, eight, or more recruiting sequences. In some embodiments, the two or more recruiting sequences are the same sequence. In some embodiments, the two or more recruiting sequences are different sequences. In some embodiments, the recruiting sequence comprises at least 10 contiguous nucleotides (e.g., at least 12, at least 14, at least 16, at least 18, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50 or more) that are at least 70% identical to a recognition motif that is specifically bound by the recruiting domain. In some embodiments, the recruiting domain Petition 870260060590, dated 06 / 22 / 2026, pp. 203 / 261 72 / 116 comprises a naturally occurring sequence that is specifically linked by a DNA-binding protein. In some embodiments, the recruiter domain comprises one or more modifications with respect to the naturally occurring sequence. In some embodiments, the recruiter domain comprises a synthetic sequence that is specifically linked by a recruiter domain described herein.
[00145] In some embodiments, the recruiting sequence comprises a protein operon sequence from the Cro repressor family (a “Cro OR3” sequence). In some embodiments, the recruiting sequence comprises an N15 OR3 operon sequence, a lambda OR3 operon sequence, a P22 OR3 operon sequence, a 434 OR3 operon sequence, or a combination thereof. In some embodiments, the recruiting sequence comprises a nucleotide sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with any of the SEQ ID Nos: 5-10.
[00146] The donor molecule may comprise at least 10 contiguous nucleotides (often referred to as a homology arm), wherein the nucleic acid molecule is at least 70% identical to a genomic nucleotide sequence, such that these contiguous nucleotides are sufficient for homologous recombination of the donor polynucleotide molecule into the cell genome at the targeted genomic DNA sequence after cleavage, for example, by a site-directed nuclease. In some embodiments, the donor polynucleotide molecule may comprise at least approximately 10, 20, 30, 50, 70, 80, 100, 150, 200, 250, 300, 250, 400, 450, 500, 600, Petition 870260060590, dated 06 / 22 / 2026, pages 204 / 261 73 / 116 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 7500, 10000, 15,000 or 20,000 nucleotides, including any value within that range not explicitly stated herein, wherein the donor polynucleotide molecule is at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a genomic nucleic acid sequence. In some embodiments, the donor molecule comprises at least one homology arm. In some embodiments, the donor molecule comprises two homology arms (which may be referred to as a left homology arm and a right homology arm).
[00147] In some embodiments, the donor polynucleotide molecule may be substantially complementary to a genomic nucleic acid sequence. In some embodiments, the donor polynucleotide molecule comprises a heterologous nucleic acid sequence. In some embodiments, the donor polynucleotide molecule comprises at least one expression cassette. In some embodiments, the donor polynucleotide molecule may comprise a transgene, which comprises at least one expression cassette. In some embodiments, the donor polynucleotide molecule comprises an allelic modification of a gene that is native to the target genome. The allelic modification may comprise at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution. In some embodiments, the allelic modification may comprise a small insertion or deletion.
[00148] The donor polynucleotide can be any suitable nucleic acid. In some embodiments, the donor polynucleotide is a portion of a donor template. In some embodiments, the donor template is part of a plasmid or linear nucleic acid. In some embodiments, the donor polynucleotide is a portion of a Petition 870260060590, dated 06 / 22 / 2026, pages 205 / 261 74 / 116 chromosome.
[00149] The various donor polynucleotide sequences described herein may be arranged in any suitable configuration. Exemplary embodiments of donor polynucleotides comprising Cro OR3 recruiting sequences and left and right homology arms flanking an insertion sequence designed for integration into a target site (e.g., according to the methods described herein) are illustrated in FIG. 1. In some embodiments, the donor polynucleotides comprise two homology arms, in which the homology arms flank an insertion sequence. In embodiments of donor polynucleotides comprising one or more recruiting sequences, the recruiting sequences may be upstream and / or downstream of the insertion sequence. In some embodiments, the recruiting sequences are between the homology arms and the insertion sequence. In some embodiments, the recruiting sequences are outside the homology arms.In some embodiments, the donor polynucleotide comprises more than one recruiting sequence comprising the same sequence in tandem series.
[00150] In some embodiments, the donor polynucleotide comprises a nucleotide sequence having at least 70% (for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identity with SEQ ID NO:20 or 21. B. Guide nucleic acids
[00151] In some cases, the systems and methods described herein comprise at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein comprise Petition 870260060590, dated 06 / 22 / 2026, pages 206 / 261 75 / 116 a plurality of guide nucleic acids. In some embodiments, the polynucleotide may be deoxyribonucleic acid (DNA). In some cases, the DNA sequence may be single-stranded or double-stranded. In some embodiments, at least one guide nucleic acid polynucleotide may be ribonucleic acid (guide RNA).
[00152] In some embodiments, the nuclease may be complexed with at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide may comprise a nucleic acid-targeting region comprising a sequence complementary to a nucleic acid sequence in the targeted polynucleotide such as targeted genomic loci or genes to confer sequence specificity of nuclease targeting. In some embodiments, the at least one guide RNA polynucleotide may comprise two separate nucleic acid molecules, which may be referred to as a double guide nucleic acid, or a single nucleic acid molecule, which may be referred to as a single guide nucleic acid (e.g., single guide RNA or sgRNA). In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a fused CRISPR RNA (crRNA) and a transactivator crRNA (tracrRNA).In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA. In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA but lacking a tracrRNA. In some embodiments, the guide nucleic acid is a double guide nucleic acid comprising unfused crRNA and tracrRNA. An exemplary double guide nucleic acid might comprise a crRNA-like molecule and a tracrRNA-like molecule. An exemplary single guide nucleic acid might comprise a crRNA-like molecule. An exemplary single guide nucleic acid might comprise a fused crRNA-like molecule and a tracrRNA-like molecule. Petition 870260060590, dated 06 / 22 / 2026, pp. 207 / 261 76 / 116
[00153] A crRNA may comprise the nucleic acid-targeting segment (e.g., spacer region) of the guide nucleic acid and a nucleotide stretch that may form half of a double-stranded duplex of the Cas protein-binding segment of the guide nucleic acid.
[00154] A tracrRNA may comprise a nucleotide stretch that forms the other half of the double-stranded duplex of the Cas protein-binding segment of the gRNA. A nucleotide stretch of a crRNA may be complementary to and hybridize with a nucleotide stretch of a tracrRNA to form the double-stranded duplex of the Cas protein-binding domain of the guide nucleic acid.
[00155] CrRNA and tracrRNA can hybridize to form a guide nucleic acid. The crRNA can also provide a segment targeting single-stranded nucleic acids (e.g., a spacer region) that hybridizes with a target nucleic acid recognition sequence (e.g., protospacer). The sequence of a crRNA, including the spacer region, or tracrRNA molecule can be designed to be specific to the species in which the guide nucleic acid is to be used.
[00156] Whether a nuclease requires only a crRNA molecule or whether it requires both a crRNA molecule and a tracrRNA molecule (covalently linked or not) depends on the CRISPR-associated nuclease used.
[00157] In some embodiments, the nucleic acid-targeting region of a guide nucleic acid may be between 18 and 72 nucleotides in length. The nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) may be about 12 nucleotides long to about 100 nucleotides long. For example, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) may be about 12 nucleotides (nt) long to about 80 nt, or about 12 nt long to about 50 nt, Petition 870260060590, dated 06 / 22 / 2026, pp. 208 / 261 77 / 116 from about 12 nt to about 40 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 12 nt to about 18 nt, from about 12 nt to about 17 nt, from about 12 nt to about 16 nt or from about 12 nt to about 15 nt.Alternatively, the segment targeting DNA may have a length of approximately 18 nt to approximately 20 nt, approximately 18 nt to approximately 25 nt, approximately 18 nt to approximately 30 nt, approximately 18 nt to approximately 35 nt, approximately 18 nt to approximately 40 nt, approximately 18 nt to approximately 45 nt, approximately 18 nt to approximately 50 nt, approximately 18 nt to approximately 60 nt, approximately 18 nt to approximately 70 nt, approximately 18 nt to approximately 80 nt, approximately 18 nt to approximately 90 nt, approximately 18 nt to approximately 100 nt, approximately 20 nt to approximately 25 nt, approximately 20 nt to approximately 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, from about 20 nt to about 60 nt, from about 20 nt to about 70 nt, from about 20 nt to about 80 nt nt, from about 20 nt to about 90 nt or from about 20 nt to about 100 nt.The length of the region targeting nucleic acids can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The length of the region targeting nucleic acids (e.g., spacer region) can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides.
[00158] In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 17 nucleotides in length. Petition 870260060590, dated 06 / 22 / 2026, pp. 209 / 261 78 / 116 length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 16 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 22 nucleotides in length.
[00159] The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid may have a length of, for example, at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 45 nt, about 12 nt to about 40 nt, about 12 nt to about 35 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, about 12 nt to about 19 nt, about 19 nt to about 20 nt, about 19 nt to about 25 nt,from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt.,
[00160] A protospacer sequence of a polynucleotide Petition 870260060590, dated 06 / 22 / 2026, pp. 210 / 261 The targeted 79 / 116 can be identified by identifying an adjacent protospacer motif (PAM) within a region of interest and selecting a region of a desired size upstream or downstream of the PAM as the protospacer. A corresponding spacer sequence can be projected by determining the complementary sequence of the protospacer region.
[00161] A spacer sequence can be identified using a computer program (e.g., machine-readable code). The computer program can use variables such as predicted melting temperature, secondary structure formation and predicted pairing temperature, sequence identity, genomic context, chromatin accessibility, % GC, genomic occurrence frequency, methylation status, presence of SNPs, and the like.
[00162] The percentage of complementarity between the sequence targeting nucleic acids (for example, a spacer sequence of at least one guide polynucleotide as disclosed here) and the target nucleic acid (for example, a protospacer sequence of one or more target loci as disclosed here) may be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percentage of complementarity between the sequence targeting nucleic acids and the target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over approximately 20 contiguous nucleotides.
[00163] The Cas-binding segment of a guide nucleic acid may comprise two nucleotide stretches (e.g., crRNA and tracrRNA) that are complementary to each other. The two stretches Petition 870260060590, dated 06 / 22 / 2026, pages 211 / 261 80 / 116 nucleotides (e.g., crRNA and tracrRNA) that are complementary to each other can be covalently linked by intervening nucleotides (e.g., a ligand in the case of a single guide nucleic acid). The two nucleotide stretches (e.g., crRNA and tracrRNA) that are complementary to each other can hybridize to form a double-stranded RNA duplex or hairpin segment for Cas protein binding, thus resulting in a stem-loop structure. The crRNA and tracrRNA can be covalently linked through the 3' end of the crRNA and the 5' end of the tracrRNA. Alternatively, the tracrRNA and crRNA can be covalently linked through the 5' end of the tracrRNA and the 3' end of the crRNA.
[00164] The Cas-binding segment of a guide nucleic acid can have a length of about 10 nucleotides to about 100 nucleotides, for example, from about 10 nucleotides (nt) to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt or from about 90 nt to about 100 nt. For example, the Cas-binding segment of a guide nucleic acid may have a length of about 15 nucleotides (nt) to about 80 nt, about 15 nt to about 50 nt, about 15 nt to about 40 nt, about 15 nt to about 30 nt, or about 15 nt to about 25 nt.
[00165] The duplex dsRNA of the Cas protein-binding segment of the guide nucleic acid can have a length of about 6 base pairs (bp) to about 50 bp. For example, the duplex dsRNA of the protein-binding segment can have a length of about 6 bp to about 40 bp, about 6 bp to about 30 bp, about 6 bp to about 25 bp, about 6 bp to about 20 bp, about 25 bp. Petition 870260060590, dated 06 / 22 / 2026, pp. 212 / 261 81 / 116 bp to about 15 bp, from about 8 bp to about 40 bp, from about 8 bp to about 30 bp, from about 8 bp to about 25 bp, from about 8 bp to about 20 bp or from about 8 bp to about 15 bp. For example, the dsRNA duplex of the Cas protein-binding segment can have a length of about 8 bp to about 10 bp, about 10 bp to about 15 bp, about 15 bp to about 18 bp, about 18 bp to about 20 bp, about 20 bp to about 25 bp, about 25 bp to about 30 bp, about 30 bp to about 35 bp, about 35 bp to about 40 bp, or about 40 bp to about 50 bp.
[00166] In some embodiments, the Cas protein-binding segment duplex dsRNA may have a length of 36 base pairs. The complementarity percentage between the nucleotide sequences that hybridize to form the protein-binding segment duplex dsRNA may be at least about 60%. For example, the complementarity percentage between the nucleotide sequences that hybridize to form the protein-binding segment dsRNA duplex may be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some cases, the percentage of complementarity between the nucleotide sequences that hybridize to form the duplex dsRNA of the protein-binding segment is 100%.
[00167] The linker (for example, the sequence that links a crRNA and a tracrRNA into a single guide nucleic acid) can have a length of about 3 nucleotides to about 100 nucleotides. For example, the linker can have a length of about 3 nucleotides (nt) to about 90 nt, about 3 nucleotides (nt) to about 80 nt, about 3 nucleotides (nt) to about 70 nt, about 3 nucleotides (nt) to about 60 nt, about 3 nucleotides (nt) to about 50 nt, Petition 870260060590, dated 06 / 22 / 2026, pp. 213 / 261 82 / 116 approximately 3 nucleotides (nt) to approximately 40 nt, approximately 3 nucleotides (nt) to approximately 30 nt, approximately 3 nucleotides (nt) to approximately 20 nt or approximately 3 nucleotides (nt) to approximately 10 nt. For example, the binder may have a length of about 3 nt to about 5 nt, about 5 nt to about 10 nt, about 10 nt to about 15 nt, about 15 nt to about 20 nt, about 20 nt to about 25 nt, about 25 nt to about 30 nt, about 30 nt to about 35 nt, about 35 nt to about 40 nt, about 40 nt to about 50 nt, about 50 nt to about 60 nt, about 60 nt to about 70 nt, about 70 nt to about 80 nt, about 80 nt to about 90 nt, or about 90 nt to about 100 nt. In some modalities, the RNA ligand targeting DNA is 4 nt.
[00168] The guide nucleic acids in the description systems may include modifications or sequences that provide additional desirable features (e.g., modified or regulated stability; subcellular targeting; tracking with a fluorescent tag; a binding site for a protein or protein complex; and the like).Examples of such modifications include, for example, a 5' cap (a 7-methylguanylate (m7G) cap); a 3' polyadenylated cap (a 3' poly(A) cap); a riboswitch sequence (e.g., to allow regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (a hairpin)); a modification or sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows fluorescent detection, and so on); a modification or sequence that provides a tracking site. Petition 870260060590, dated 06 / 22 / 2026, pp. 214 / 261 83 / 116 binding to proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).
[00169] A guide nucleic acid may comprise one or more modifications (e.g., a base modification, a backbone modification) to provide the nucleic acid with a new or enhanced characteristic (e.g., improved stability). A guide nucleic acid may comprise a nucleic acid affinity label. A nucleoside may be a base-sugar combination. The base portion of the nucleotide may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. Nucleotides may be nucleosides that additionally include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In the formation of guide nucleic acids, phosphate groups covalently link adjacent nucleosides to each other to form a linear polymeric compound.In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound; however, linear compounds may be suitable. Additionally, linear compounds may have internal complementarity of nucleotide bases and may therefore fold in a manner to produce a fully or partially double-stranded compound. Furthermore, within guide nucleic acids, phosphate groups may be commonly referred to as forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid may be a 3' to 5' phosphodiester bond. Petition 870260060590, dated 06 / 22 / 2026, pages 215 / 261 84 / 116
[00170] A guide nucleic acid may comprise a modified backbone and / or modified internucleoside linkages. Modified backbones may include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.
[00171] Suitable modified guide nucleic acid skeletons containing a phosphorus atom therein may include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl phosphonates and others such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates including 3'-amino phosphoramidate and aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates and boranephosphates having normal 3'-5' linkages, 2'-5' linked analogs and those having inverted polarity in which one or more internucleotide linkages are a 3' to 3', 5' to 5' or 2' linkage for 2'.Guide nucleic acids having suitable inverted polarity may comprise a single 3' to 3' linkage in the internucleotide plus 3' bond (such as a single inverted nucleoside residue in which the nucleobase is absent or has a hydroxyl group in its place). Various salts (e.g., potassium chloride or sodium chloride), mixed salts, and free acid forms may also be included.
[00172] A guide nucleic acid may comprise one or more phosphorothioate and / or heteroatom internucleoside linkages, in particular -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (a methylene (methylimino) or MMI backbone), -CH2-ON(CH3)-CH2-, -CH2-N(CH3)N(CH3)-CH2- and -ON(CH3)-CH2-CH2- (where the native phosphodiester internucleoside linkage is represented as -OP(=O)(OH)-O-CH2-).
[00173] A guide nucleic acid may comprise a morpholino scaffold structure. For example, a nucleic acid may Petition 870260060590, dated 06 / 22 / 2026, pp. 216 / 261 85 / 116 comprise a 6-membered morpholino ring in place of a ribose ring. In some of these embodiments, a phosphorodiamidate or other non-phosphodiester internucleosidic bond replaces a phosphodiester bond.
[00174] A guide nucleic acid may comprise polynucleotide backbones that are formed by short-chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short-chain heteroatomic or heterocyclic internucleoside linkages. These may include those having morpholino linkages (formed in part from the sugar portion of a nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; riboacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and others having component parts of N, O, S and CH2 mixed together.
[00175] A guide nucleic acid may comprise a nucleic acid mimetic. The term “mimetic” may refer to polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced by groups other than furanose; the replacement of only the furanose ring may also be referred to as a sugar substitute. The heterocyclic base moiety or a modified heterocyclic base moiety may be retained for hybridization with an appropriate target nucleic acid. Such a nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide may be replaced by an amide-containing backbone, in particular an aminoethylglycine backbone. Nucleotides may be retained and are linked directly or indirectly to the aza nitrogen atoms of the moiety. Petition 870260060590, dated 06 / 22 / 2026, pp. 217 / 261 86 / 116 of the amide backbone. The backbone in PNA compounds may comprise two or more linked aminoethylglycine units that give the PNA an amide-containing backbone. The heterocyclic base moieties may be linked directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone.
[00176] A guide nucleic acid may comprise linked morpholino units (morpholino nucleic acid) having heterocyclic bases attached to the morpholino ring. Linking groups may link the monomeric morpholino units into a morpholino nucleic acid. Non-ionic morpholino-based oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides may be non-ionic mimetics of guide nucleic acids. A variety of compounds within the morpholino class may be joined using different linking groups. An additional class of polynucleotide mimetic may be referred to as cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in a nucleic acid molecule may be replaced by a cyclohexenyl ring. DMT-protected CeNA phosphoramidite monomers may be prepared and used for the synthesis of oligomeric compounds using phosphoramidite chemistry.The incorporation of CeNA monomers into a nucleic acid chain can increase the stability of a DNA / RNA hybrid. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to the native complexes. A further modification can include Blocked Nucleic Acids (LNAs) in which the 2'hydroxyl group is attached to the 4' carbon atom of the sugar ring, thus forming a 2'-C,4'-C-oxymethylene linkage, resulting in a bicyclic sugar moiety. The linkage can be a methylene (-CH2) group, bridging the 2' oxygen atom and the 4' carbon atom. Petition 870260060590, dated 06 / 22 / 2026, pages 218 / 261 87 / 116 carbon 4' where n is 1 or 2. LNA and LNA analogs can exhibit very high thermal duplex stabilities with complementary nucleic acid (Tm = +3 to +10 °C), stability towards 3' exonucleolytic degradation, and good solubility properties.
[00177] A guide nucleic acid may comprise one or more substituted sugar moieties. Suitable polynucleotides may comprise a sugar substituent group selected from: OH; F; O-, S- or N-alkyl; O-, S- or N-alkenyl; O-, S- or N-alkynyl; or O-alkyl-O-alkyl, wherein alkyl, alkenyl and alkynyl may be C1 to C10 alkyl or substituted or unsubstituted C2 to C10 alkenyl and alkynyl. O((CH2)nO)mCH3, O(CH2)nOCH3, O(CH2)nNH2, O(CH2)nCH3, O(CH2)nONH2 and O(CH2)nON((CH2)nCH3)2 are particularly suitable, wherein neither are from 1 to about 10.A sugar substituent group may be selected from: C1 to C10 lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, an RNA cleavage group, a reporter group, an intercalator, a group to improve the pharmacokinetic properties of a guide nucleic acid or a group to improve the pharmacodynamic properties of a guide nucleic acid and other substituents having similar properties. A suitable modification may include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE, an alkoxyalkoxy group).A suitable additional modification may include 2'-dimethylamino-oxyethoxy, (an O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE) and 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-aminoethoxy-ethyl or 2'-DMAEOE), 2'-O-CH2-O-CH2-N(CH3)2.
[00178] Other suitable sugar substitute groups may Petition 870260060590, dated 06 / 22 / 2026, pp. 219 / 261 88 / 116 include methoxy (-O-CH3), aminopropoxy (--OCH2CH2NH2), allyl (-CH2CH=CH2), -O-allyl (--O--CH2—CH=CH2), and fluorine (F). The substituent groups of the 2'-sugar can be in the arabino (upward) or ribo (downward) position. A suitable 2'-arabino modification is 2'-F. Similar modifications can also be made at other positions in the oligomeric compound, particularly at the 3' position of the sugar in the 3'-terminal nucleoside or in 2'-5' linked nucleotides and at the 5' position of the 5'-terminal nucleotide. Oligomeric compounds may also have sugar mimetics such as cyclobutyl moieties in place of the pentofuranosyl sugar.
[00179] A guide nucleic acid may also include nucleobase (or “base”) modifications or substitutions. As used herein, “unmodified” or “natural” nucleobases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)).Modified nucleobases may include other synthetic and natural nucleobases such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other adenines and guanines. 8-substituted, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other uracils and 5-substituted cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-aza-adenine, 7-desazaguanine and 7-desaza-adenine and 3-desaza-adenine.The modified nucleobases may include tricyclic pyrimidines such as phenoxazine cytidine (1Hpyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1HPetition 870260060590, 22 / 06 / 2026, page 220 / 261. 89 / 116 pirimido(5,4-b)(1,4)benzothiazine-2(3H)-one), G-clamps such as a substituted phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-Hpirimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2Hpirimido(4,5-b)indol-2-one), pyridoindol cytidine (Hpirido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).
[00180] Heterocyclic base moieties may include those in which the purine or pyrimidine base is replaced by other heterocycles, for example 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine and 2-pyridone. Nucleobases may be useful for increasing the binding affinity of a polynucleotide compound. These may include 5-substituted pyrimidines, 6-azapyrimidines and N2, N-6 and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil and 5-propynylcytosine. 5-Methylcytosine substitutions may increase nucleic acid duplex stability at 0.6-1.2 °C and may be suitable base substitutions (e.g., when combined with 2'-O-methoxyethyl sugar modifications).
[00181] A modification of a guide nucleic acid may comprise chemically attaching to the guide nucleic acid one or more fragments or conjugates that can enhance the activity, cellular distribution, or cellular uptake of the guide nucleic acid. These fragments or conjugates may include groups covalently conjugated to functional groups such as primary or secondary hydroxyl groups. Conjugated groups may include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of oligomers, and groups that can enhance the pharmacokinetic properties of oligomers. Conjugated groups may include, but are not limited to, cholesterols, lipids, phospholipids, biotin, phenazine, folate, phenanthridine, anthraquinone, acridine, fluoresceins, rhodamines, coumarins, and dyes. Petition 870260060590, dated 06 / 22 / 2026, pages 221 / 26190 / 116 groups that enhance pharmacodynamic properties include groups that improve uptake, enhance resistance to degradation, and / or strengthen specific hybridization of the sequence with the target nucleic acid. Groups that may enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of a nucleic acid. Conjugated fractions may include, but are not limited to, lipid fractions such as a cholesterol fraction, cholic acid, a thioether (e.g., hexyl-S-tritylthiol), a thiocholesterol, an aliphatic chain (e.g., dodecandiol or undecyl residues), a phospholipid (e.g., dihexadecyl-rac-glycerol or 1,2-di-O-hexadecyl-rac-glycero-3H-triethylammonium phosphonate), a polyamine, or a polyethylene glycol or adamantaneacetic acid chain, a palmityl fraction, or an octadecylamine or hexylamino-carbonyl-oxycholesterol fraction.
[00182] In some embodiments, at least one guide RNA polynucleotide of a system or method provided herein may bind to at least one portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). In some cases, at least one guide RNA polynucleotide is capable of forming a complex with a site-directed nuclease to direct the site-directed nuclease to target the portion of a target nucleic acid (e.g., a location in a genome or a gene).
[00183] In some embodiments, the systems described herein comprise at least one guide RNA polynucleotide that is capable of forming a complex with a site-directed nuclease portion of a system fusion protein. In some embodiments, the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides that are capable of forming a complex with a site-directed nuclease portion of a Petition 870260060590, dated 06 / 22 / 2026, pp. 222 / 261 91 / 116 system fusion protein.
[00184] In some embodiments, the guide nucleic acid comprises a nucleotide sequence having at least 70% (for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identity with SEQ IDs NOS: 17 or 19.
[00185] Kits that include the components of the systems described in this description are also provided here. In some embodiments, the kits include one or more of the fusion proteins and / or polynucleotides described herein. VI. Methods
[00186] In another aspect, methods are provided here for editing one or more nucleic acids using the fusion proteins and / or systems described herein. In some embodiments, the methods comprise contacting a nucleic acid comprising fusion protein binding sites (i.e., the nucleic acid to be edited) with at least one fusion protein as described herein, wherein contacting the nucleic acid with at least one fusion protein results in an edit to the nucleic acid. The nucleic acid (i.e., the nucleic acid to be edited) may be any suitable nucleic acid. In some embodiments, the nucleic acid is a portion of a chromosome. In some embodiments, the nucleic acid is a portion of a genome (e.g., a plant genome).
[00187] As described here and demonstrated in the Examples below, the methods provided here can result in increased frequency of one or more desired nucleic acid editing outcomes (e.g., fragment replacement via HDR).
[00188] In some modes, the nucleic acid to be edited Petition 870260060590, dated 06 / 22 / 2026, pp. 223 / 261 92 / 116 by the method comprises a target region. As used herein, “target region” refers to a portion of a nucleic acid that is targeted for editing. For example, a target region may be a portion of a gene that is to be edited. In some embodiments, at least one portion of the target region is replaced by at least one portion of a donor polynucleotide. In some embodiments, the target region comprises at least one (e.g., two, three, four, five, six, seven, eight, nine, 10 or more) nuclease cleavage site. In some embodiments, the nucleic acid comprises at least one binding site. In some embodiments, the target region is flanked by nuclease cleavage sites. In some embodiments, the nucleic acid comprises a first binding site that is adjacent to the 5' end of the target region and a second binding site that is adjacent to the 3' end of the target region.
[00189] In some embodiments, as detailed below, the nucleic acid to be edited comprises a first binding site and a second binding site. In some embodiments, the first binding site and the second binding site are different sequences, and the method comprises providing two different fusion proteins, one to bind to the first binding site and one to bind to the second binding site. In some embodiments, the first binding site and the second binding site are the same sequence, and the method comprises providing a fusion protein that can bind to both the first binding site and the second binding site.
[00190] In some embodiments, the methods herein comprise providing a donor polynucleotide. In some embodiments, the donor polynucleotide comprises a left homology arm (i.e., a homology arm that is complementary to a sequence upstream of the target region) and a right homology arm (i.e., Petition 870260060590, dated 06 / 22 / 2026, pp. 224 / 261 93 / 116 a homology arm that is complementary to a sequence downstream of the target region). In such embodiments, the nucleic acid target region to be edited is flanked by homology arms, and the target region comprises at least one fusion protein binding site (i.e., at least one cleavage site). An exemplary embodiment is illustrated in FIG. 2.
[00191] In some embodiments of the methods provided herein, the recruiting domain of the fusion protein binds specifically to a recruiting sequence in the donor polynucleotide. Without being constrained by any particular theory, a fusion protein that binds both to a binding site in the target region of nucleic acid (i.e., via site-directed nuclease binding) and to a recruiting sequence in the donor polynucleotide (i.e., via recruiter domain-specific binding) ties the donor polynucleotide in close proximity to the cleavage site formed by the SDN. Such close proximity, in some embodiments, increases the likelihood that the cleaved nucleic acid will be repaired in a manner that leads to the incorporation of at least a portion of the donor polynucleotide into the target region. In some embodiments, the donor polynucleotide comprises at least one homology arm, and repair occurs via HDR.In some modalities, repair occurs through NHEJ or MMEJ, in which the ends of the donor polynucleotide are joined to the cleaved nucleic acid ends in the target region.
[00192] In some embodiments of the methods provided herein, for example, in the Examples herein, the site-directed nuclease of at least one fusion protein comprises a CRISPR-associated nuclease. In such embodiments, the method may further comprise providing guide RNAs to direct fusion proteins to binding sites. In some embodiments, the method comprises Petition 870260060590, dated 06 / 22 / 2026, pp. 225 / 261 94 / 116 provide at least one first guide RNA and at least one second guide RNA. In some embodiments, the at least first guide RNA comprises a nucleotide sequence with complementarity to the first nucleic acid binding site to be edited. In some embodiments, the at least second guide RNA comprises a nucleotide sequence with complementarity to the second nucleic acid binding site to be edited.
[00193] The methods herein comprise providing at least one fusion protein, a nucleic acid to be edited, and a donor polynucleotide, and may also comprise providing at least one guide RNA. These various components may be provided using any suitable technique. For example, providing a fusion protein may comprise introducing the fusion protein into a cell or introducing a recombinant nucleic acid, construct, or vector encoding the fusion protein into a cell. Similarly, a gRNA may be provided by introducing the gRNA itself or a nucleic acid sequence encoding the gRNA. In some embodiments, a fusion protein and a gRNA may be encoded by the same DNA construct or vector. EXAMPLES Example 1. The Cro-Cas9 fusion induces allele substitution at the ZmALS2 locus.
[00194] This example demonstrates that a CroCas9 fusion protein improves the efficiency of homology-directed repair (HDR)-mediated allele replacement using donor DNA containing Cro-binding sequences. In this example, as shown in FIG. 3 (top panel), the donor DNA contains a Cro operator site (OR3). The Cro protein, functioning as a homodimer, binds to the operator site. The Cro protein is tethered to Cas9 via a linker. Petition 870260060590, dated 06 / 22 / 2026, pages 226 / 261 95 / 116 XTEN.
[00195] The schematic design of the Cro-Cas9 fusion is shown in FIG. 4. A single-chain dimer (two monomers linked by a 15-amino acid linker: SEQ ID NO: 15) of the Cro protein from bacteriophage N15 was fused to the N-terminal of Streptococcus pyogenes Cas9 (SpCas9) via a 32-amino acid linker (SEQ ID NO: 16). The fusion protein is labeled with a 3X FLAG peptide followed by a nuclear localization signal (NLS) from SV40 at the N-terminal and another NLS at the C-terminal.
[00196] Svitashev et al. (2015, Plant Physiol.) reported Cas9-induced allele substitution at the ZmALS2 locus (the maize acetolactate synthase homolog on chromosome 5). In this example, the same target site (gRNA is represented by SEQ ID NO: 17) was chosen to evaluate the effectiveness of Cro-Cas9 fusion in improving allele substitution efficiency.
[00197] A DNA construct was built to express Cro-Cas9 fusion proteins and gRNA in maize cells. The maize codon-optimized coding sequence of the fusion protein was operationally ligated to a sugarcane ubiquitin 4 promoter and an Agrobacterium tumefaciens nopalin synthase terminator. The single guide RNA (sgRNA) targeting the ZmALS2 locus was operationally ligated to an Oryza sativa U3 promoter and a stretch of nine consecutive thymine bases followed by the sgRNA spacer sequence to terminate transcription. A control construct lacking Cro domains was created but otherwise identical to the Cro-Cas9 expressing vector.
[00198] The donor DNA serving as a repair template for HDR was based on a 657 bp homology fragment taken from the coding region of ZmALS2, starting from the fourth codon (GCT for alanine) and ending at the 222nd codon (CGG for arginine). A Petition 870260060590, dated 06 / 22 / 2026, pp. 227 / 261 The 96 / 116 donor DNA sequence is represented by SEQ ID NO: 20. A total of eight single nucleotide substitutions were introduced into the homology fragment around the Cas9 cleavage site. Among them, only two conferred amino acid changes. The G-for-C substitution at codon 111 changed methionine to isoleucine while destroying the TGG PAM of the gRNA, protecting the donor DNA from cleavage by Cas9. The C-for-T substitution at codon 119 changed proline to serine, making the gene product resistant to the herbicide chlorsulfurone. The nucleotide substitutions also created three restriction sites as molecular barcodes that facilitate the detection of the edited allele by HDR. A 20 bp Or3 operon sequence from bacteriophage N15 (SEQ ID NO: 18) specifically recognized and bound by Cro proteins was added to each flank of the homology fragment containing substitutions to prepare the donor DNA.The donor DNA was cloned into a high-copy cloning vector flanked by PmeI and ASCI restriction sites for excision.
[00199] The linearized vector (0.2 pmol) was biolistically co-delivered with double-stranded donor DNA (2 pmol) into maize embryonic callus cells. Transgenic calli were selected in PMI selection medium and subsequently regenerated into transgenic plants via standard tissue culture procedures. Regenerated plants were sampled for DNA extraction, and a TaqMan assay was designed to distinguish the HDR-substituted resistant allele from the wild-type allele. A primer pair outside the homology region was designed to amplify a 1.5 kb fragment spanning the entire homology by PCR, and amplicons were subjected to restriction fragment length polymorphism (RFLP) analysis and Sanger sequencing for sequence validation.
[00200] The results are summarized in Table 1. In each one Petition 870260060590, dated 06 / 22 / 2026, pages 228 / 261 Of the three plants with substituted alleles generated by the Cro-Cas9 fusion construct, 97 / 116 of the two alleles underwent HDR-mediated allele substitution with all eight nucleotide substitutions installed, while the other allele harbored an indel mutation at the Cas9 cleavage site. The control vector generated only one plant in which one of the two alleles underwent partial substitution with a subset of the eight installed substitutions. Table 1. Efficiency of allele substitution at the ZmALS2 locus. Number of Number of Number of Efficiency Construct plants replacement global replacement transgenic total perfect partial (perfect + partial) Fusion Cro- 240 3 0 1.25% Cas9 Cas9 of 160 0 1 0.63% control Example 2. The Cro-Cas12a fusion induces allele substitution at the ZmGL2 locus.
[00201] This Example demonstrates that a CroCas12a fusion protein improves the efficiency of homology-directed repair (HDR)-mediated allele replacement using donor DNA containing Cro-binding sequences. In this Example, as shown in FIG. 3 (bottom panel), the donor DNA contains a Cro operator site (Or3). The Cro protein, functioning as a homodimer, binds to the operator site. The Cro protein is tethered to Cas12a via an XTEN linker.
[00202] The schematic design of the Cro-Cas12a fusion is shown in FIG. 5. The same single-chain N15 Cro dimer as described in Example 1 was fused to the N-terminal of the D156R variant of Cas12a from Lachnospiraceae bacterium ND2006 (LbCas12a) via a 30 amino acid linker (GGGGS)e (SEQ ID NO:26). The protein Petition 870260060590, dated 06 / 22 / 2026, pp. 229 / 261 The 98 / 116 fusion is flanked by one SV40 NLS at the N-terminal and two SV40 NLS at the C-terminal.
[00203] K. Lee, et al., Activities and specificities of CRISPR / Cas9 and Cas12a nucleases for targeted mutagenesis in maize, Plant Biotech. J., 17: 362-372 (2019) reported that LbCas12a could efficiently induce double-strand DNA breaks at the maize Glossy2 locus (ZmGL2). In this example, the same target site (the gRNA is SEQ ID NO: 19) was chosen to evaluate the effectiveness of Cro-LbCas12a fusion in improving allele replacement efficiency.
[00204] A DNA construct was built to express Cro-LbCas12a fusion proteins and gRNA in maize cells. The maize codon-optimized coding sequence of the fusion protein was operationally ligated to a sugarcane ubiquitin 4 promoter and an Agrobacterium tumefaciens nopalin synthase terminator. Mature crRNA targeting the ZmGL2 locus, flanked with hammerhead (HH) ribozymes and hepatitis delta virus (HDV) for processing, was operationally ligated to another copy of the sugarcane ubiquitin 4 promoter and another copy of the nopalin synthase terminator. An identical construct lacking Cro domains was built as a control.
[00205] The donor DNA serving as the repair template for HDR was based on an 859 bp homology fragment taken from the coding region of ZmGL2. The donor DNA sequence is represented by SEQ ID NO: 21. The left (upstream) cleavage site homology resided in the first intron, and the right (downstream) cleavage site homology resided in the second exon. A consecutive eight-nucleotide substitution (ACAAACTT to TAGTGACC) was introduced into the homology fragment in the middle of the gRNA protospacer, which created two consecutive premature stop codons and protected the donor from cleavage by Cas12a. The same sequences Petition 870260060590, dated 06 / 22 / 2026, pp. 230 / 261 99 / 116 N15 Or3, as described above, were added to each flank of the homology fragment containing substitutions to prepare the donor DNA. The donor DNA was cloned into a cloning vector and could be prepared in linear form by PCR using the vector as a template.
[00206] A mixture of 0.15 pmol linearized vector and 2.25 pmol double-stranded donor DNA was biolistically co-delivered into maize embryonic callus cells. Transgenic calli were selected in PMI selection medium and subsequently regenerated into transgenic plants via tissue culture procedures. Regenerated plants were sampled for DNA extraction, and a TaqMan assay was designed to distinguish the HDR-substituted functional silencing allele from the wild-type allele. A primer pair outside the homology region was designed to amplify a 1.4 kb fragment spanning the entire homology by PCR, and the amplicons were subjected to Sanger sequencing to validate expected mutations and genomic homology joins.
[00207] The results are summarized in Table 2. In all plants, the 8-nucleotide mutation was installed at the ZmGL2 locus. Among these plants, six generated by the Cro-LbCas12a fusion and three generated by the LbCas12a control had both genomic homology junctions in the substituted allele undergoing perfect HDR. In contrast, two plants generated by the Cro-LbCas12a fusion and seven generated by the LbCas12a control had at least one of the genomic homology junctions in the substituted allele failing to undergo perfect HDR. Notably, one plant, generated by CroLbCas12a, had both alleles undergoing perfect HDR substitution. Table 2. Efficiency of allele substitution at the ZmGL2 locus. Construction No. of plants No. of Efficiency No. Petition 870260060590, dated 06 / 22 / 2026, pages 231 / 261 100 / 116 Total transgenic replacement by perfect HDR replacement imperfect replacement of perfect HDR Fusion Cro- LbCas12a 160 6* 2 3.8% Control LbCas12a 198 3 7 1.3% * One plant has homozygous perfect HDR replacement. LIST OF REFERENCED SEQUENCES SEQ ID NO:1 Cro monomer amino acid sequence of N15 MKPEELVRHFGDVEKAAVGVGVTPGAVYQWLQAGEIPPLRQSDIEV RTAYKLKSDFTSQRMGKEGHNSGTK SEQ ID NO:2 Amino acid sequence of Cro lambda monomer MEQRITLKDYAMRFGQTKTAKDLGVYQSAINKAIHAGRKIFLTINADG SVYAEEVKPFPSNKKTTA SEQ ID NO:3 Cro monomer amino acid sequence of P22 MYKKDVIDHFGTQRAVAKALGISDAAVSQWKEVIPEKDAYRLEIVTAG ALKYQENAYRQAA SEQ ID NO:4 Cro monomer amino acid sequence of 434 MQTLSERLKKRRIALKMTQTELATKAGVKQQSIQLIEAGVTKRPRFLF EIAMALNCDPVWLQYGTKRGKAA SEQ ID NO:5 Or3 de Cro de N15 TTATAGCTGGCTATAA SEQ ID NO:6 Or3 de Cro de Lamba (native) TATCACCGCAAGGGATA Petition 870260060590, dated 06 / 22 / 2026, pp. 232 / 261 101 / 116 SEQ ID NO:7 Lamba Cro Or3 (synthetic) TATCACCGGGGTGATA SEQ ID NO:8 Or3 de Cro de P22 AGTTAAGTCTTAAAT SEQ ID NO:9 Cro Or3 of 434 (native) ACAAGAAAAACTGT SEQ ID NO:10 434 Cro Or3 (synthetic) ACAATATATATTGT SEQ ID NO:11 Cro-Cas9 amino acid sequence MDYKDHDGDYKDHDIDYKDDDDKGAPKKKRKVGGGGSGKPEELVR HFGDVEKAAVGVGVTPGAVYQWLQAGEIPPLRQSDIEVRTAYKLKSD FTSQRMGKEGHNSGTKGGGSGGGSGGGSGGGKPEELVRHFGDVE KAAVGVGVTPGAVYQWLQAGEIPPLRQSDIEVRTAYKLKSDFTSQRM GKEGHNSGTKSGGSSGGSSGSETPGTSESATPESSGGSSGGSMDK KYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALL FDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFH RLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTD KADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLF EENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALS LGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAA KNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQ LPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLV KLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKI EKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGAS AQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGM Petição 870260060590, de 22 / 06 / 2026, pág. 233 / 261 102 / 116 RKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGV EDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIE ERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTIL DFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANL AGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQK NSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDM YVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNV PSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFI KRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDF RKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFKTEITLANGEIRKRPLI ETNGETGEIVWDKGRDFATVRLKVLSMPQVNIVKKTEVQTGGFSKESI LPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKL KSVKELVGITIMERSSFEKNPVDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPI REqaENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSI TGLYETRIDLSQLGGDSSPPKKKRKVSWKDASGWSRM SEQ ID NO:12 Cro-Cas9 nucleotide sequence ATGGATTACAAGGACCACGACGGGGACTACAAGGACCACGACATT GACTACAAGGAACGACGATGATAAGGGGGCTCCGAAGAAGAAGAG GAAGGTCGGCGGCGGCGGAGCGCAAGCCCGAGGAGCTGGTG TACCCCGGGGGCCGTCTACCAGTGGCTGCAGGCGGGCGAGATC CCTCCGCTGCGCCAGAGCGACATCGAGGTTCGGAGCGTACAA GCTCAAGAGCGATTTCACCAGCCAGAGGATGGGCAAGGAGGGGC GGGCGGCTCCGGGGGCGGCAAGCCAGAGGAGCTGGTTAGGCAC TTCGGCGACGTCGAGAAGGCTGCCGTGGGGGTTGGCGTGACTCC Petition 870260060590, of 22 / 06 / 2026, p. 234 / 261 103 / 116 AGGGGCCGTGTACCAGTGGCTCCAGGCCGGCGAGATTCCGCCAC TGCGGCAGTCCGACATCGAGGTGCGCACCGCTTACAAGCTGAAG TCCGACTTCACCTCACAGAGGATGGGGAAGGAGGGCCATAACAG CGGCACAAAGTCAGGGGGCTCCTCAGGCGGGTCCAGCGGGTCG GAAACCCCAGGGACCAGCGAGTCGGCCACCCCAGAGTCGTCGG GCGGCAGCAGCGGCGGGAGCATGGACAAGAAGTACAGCATCGG CCTGGACATCGGCACCAACAGCGTGGGCTGGGCCGTGATCACCG ACGAGTACAAGGTGCCGAGCAAGAAGTTCAAGGTGCTGGGCAAC ACCGACAGGCACAGCATCAAGAAGAACCTGATCGGCGCCCTGCT GTTCGACAGCGGCGAGACCGCCGAGGCCACCAGGCTGAAGAGG ACCGCCAGGAGGAGGTACACCAGGAGGAAGAACAGGATCTGCTA CCTGCAGGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACA GCTTCTTCCACAGGCTGGAGGAGAGCTTCCTGGTGGAGGAGGAC AAGAAGCACGAGAGGCACCCGATCTTCGGCAACATCGTGGACGA GGTGGCCTACCACGAGAAGTACCCGACCATCTACCACCTGAGGA AGAAGCTGGTGGACAGCACCGACAAGGCCGACCTGAGGCTGATC TACCTGGCCCTGGCCCACATGATCAAGTTCAGGGGCCACTTCCTG ATCGAGGGCGACCTGAACCCGGACAACAGCGACGTGGACAAGCT GTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGAGGAGAA CCCGATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGAGCG CCAGGCTGAGCAAGAGCAGGAGGCTGGAGAACCTGATCGCCCAG CTGCCGGGCGAGAAGAAGAACGGCCTGTTCGGCAACCTGATCGCCCTGAGCCTGGGCCTGACCCCGAACTTCAAGAGCAACTTCGACCT GGCCGAGGACGCCAAGCTGCAGCTGAGCAAGGACACCTACGACG ACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCC GACCTGTTCCTGGCCGCCAAGAACCTGAGCGACGCCATCCTGCT GAGCGACATCCTGAGGGTGAACACCGAGATCACCAAGGCCCCGC TGACGCGCCAGCATGATCAAGAGGTACGACGAGCACCACCAGGAC CTGACCCTGCTGAAGGCCCTGGTGAGGCAGCAGCTGCCGGAGAA GTACAAGGAGATCTTCTTCGACCAGAGCAAGAACGGCTACGCCG Petition 870260060590, dated 06 / 22 / 2026, pages 235 / 261 104 / 116 GCTACATCGACGGCGGCGCCAGCCAGGAGGAGTTCTACAAGTTC ATCAAGCCGATCCTGGAGAAGATGGACGGCACCGAGGAGCTGCT GGTGAAGCTGAACAGGGAGGACCTGCTGAGGAAGCAGAGGACCT TCGACAACGGCAGCATCCCGCACCAGATCCACCTGGGCGAGCTG CACGCCATCCTGAGGAGGCAGGAGGACTTCTACCCGTTCCTGAA GGACAACAGGGAGAAGATCGAGAAGATCCTGACCTTCCGCATCC CGTACTACGTGGGCCCGCTGGCCAGGGGCAACAGCAGGTTCGCC TGGATGACCAGGAAGAGCGAGGAGACCATCACCCCGTGGAACTT CGAGGAGGTGGTGGACAAGGGCGCCAGCGCCCAGAGCTTCATC GAGAGGATGACCAACTTCGACAAGAACCTGCCGAACGAGAAGGT GCTGCCGAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTACAA CGAGCTGACCAAGGTGAAGTACGTGACCGAGGGCATGAGGAAGC CGGCCTTCCTGAGCGGCGAGCAGAAGAAGGCCATCGTGGACCTG CTGTTCAAGACCAACAGGAAGGTGACCGTGAAGCAGCTGAAGGA GGACTACTTCAAGAAGATCGAGTGCTTCGACAGCGTGGAGATCAG CGGCGTGGAGGACAGGTTCAACGCCAGCCTGGGCACCTACCACG ACCTGCTGAAGATCATCAAGGACAAGGACTTCCTGGACAACGAGG AGAACGAGGACATCCTGGAGGACATCGTGCTGACCCTGACCCTG TTCGAGGACAGGGAGATGATCGAGGAGAGGCTGAAGACCTACGC CCACCTGTTCGACGACAAGGTGATGAAGCAGCTGAAGAGGAGGA GGTACACCGGCTGGGGCAGGCTGAGCAGGAAGCTGATCAACGG CATCAGGGACAAGCAGAGCGGCAAGACCATCCTGGACTTCCTGAAGAGCGACGGCTTCGCCAACAGGAACTTCATGCAGCTGATCCAC GACGACAGCCTGACCTTCAAGGAGGACATCCAGAAGGCCCAGGT GAGCGGCCAGGGCGACAGCCTGCACGAGCACATCGCCAACCTG GCCGGCAGCCCGGCCATCAAGAAGGGCATCCTGCAGACCGTGAA GGTGGTGGACGAGCTGGTGAAGGTGATGGGCAGGCACAAGCCG GAGAACATCGTGATCGAGATGGCCAGGGAGAACCAGACCACCCA GAAGGGCCAGAAGAACAGCAGGGAGAGGATGAAGAGGATCGAG GAGGGCATCAAGGAGCTGGGCAGCCAGATCCTGAAGGAGCACCC Petition 870260060590, dated 06 / 22 / 2026, pages 236 / 261 105 / 116 GGTGGAGAACACCCAGCTGCAGAACGAGAAGCTGTACCTGTACT ACCTGCAGAACGGCAGGGACATGTACGTGGACCAGGAGCTGGAC ATCAACAGGCTGAGCGACTACGACGTGGACCACATCGTGCCGCA GAGCTTCCTGAAGGACGACAGCATCGACAACAAGGTGCTGACCA GGAGCGACAAGAACAGGGGGCAAGAGCGACAACGTGCCGAGCGA GGAGGTGGTGAAGAATGAAAAAACTACTGGAGGCAGCTGCTGA ACGCCAAGCTGATCACCCAGGAAGTTCGACAACCTGACCAAG GCCGAGAGGGGCGGCCTGAGCGAGCTGGACAAGGCCGGCTTCA TTAAAAGGCAGCTGGTGGAGACCAGGCAGATCACCAAGCACGTG GCCCAGATCCTGGACAGCAGGATGAACACCAAGTACGACGAGAA CGACAAGCTGATCAGGGAGGTGAAGGTGATCACCCTGAAGAGCA AGCTGGTGAGCGACTTCAGGAAGGACTTCCAGTTCTACAAGGTGA GGGAGATCAATAATTACCACCACGCCCACGCCTACCTGAACG CCGTGGTGGGCACCGCCTGAAAAAGTACCCGAAGCTGGAG AGCGAGTTCGTGTACGGCGCACTACAAGGTGTACGACGTGAGGAA GATGATCGCCAAGAGCGAGCAGGAGATCGGCAAGGCCACGCCA AGTACTTCTTCTACAGCAACATCATGAACTTCTTCAAGACCGAGAT CACCCTGGCCAACGGCGAGATCAGGAAGGGCCGCTGATCGATGGAGA CCAACGGCGAGACCGGCGAGATCGTGTGGGGACAAGGGCAGGA CTTCGCCACCGTGAGGAAGGTGCTGCCATGCCGCAGGTGAACA TCGTGAAGAAGACCGAGGTGCAGACCGGCGGCTTCAGCAAGGAG AGCATCCTGCCGAAGAGGAACGCGACAAGCTGATCGCCAGGAGAAGGACTGGGATCCGAAGAAGTACGGCGGCTTCGACAGCCCGA CCGTGGCCTACAGCGTGCTGGTGGTGGCCAAGGTGGAGAAGGG CAAGGCAAGAAGCTGAGAGCGTGAGGAGCTGGTGGGCATCA CCATCATGGAGGAGCCTTCGTGAGAAGTCAACCCA CTGGAGGCCAAGGGCTACAAGGAGGTGAAGAAGGACCTGATCAT TAAACTGCCGAAGTACAGCCTGTTCGAGCTGGAGAACGGCAGGA AGAGGATGCTGGCCAGCCGGCGAGCTGCAGAAGGGCAACGA GCTGGCCCTGCCGAGCAAGTACGTGAACTTCCA Petition 870260060590, dated 6 / 22 / 2026, p. 237 / 261 106 / 116 GCCACTACGAGAAGCTGAGGCAGCCCGGAGGACAACGAGCA GAAGCAGCTGTTCGTGGAGCAGCAAGCACTACCTGGACGAGA TCATCGAGCAGATCAGCGAGTTCAGAGGGTGATCCTGGCC GACGCCAACCTGGACAAGGTGCTGAGCGCCTACAACAAGCCAAG GGACAAGCCGATCAGGGAGCAGGCCGAACATCATCCACCTGT TCACCCTGACCAACCTGGCGCCCCGGCCGCCTTCAAGTACTTC GACACCACCATCGACAGGAAGAGGTACACCAGCACCAAGGAGGT GCTGGACCCCACTGATCCACCAGAGTCACCGGCCTGTACG AGACCAGGATCGACCTGAGCCAGCTGGGCGGCGACAGCAGCCC GCCGAAAGAAGAGGAAGGTGAGCTGGAAGGACGCCAGCGGC TGGAGCAGGATGTGA SEQ ID NO:13 Amino acid sequence of Cro-Cas12a MPKKKRKVSGGSSGGSKPEELVRHFGDVEKAAVGVGVTPGAVYQW LQAGEIPPLRQSDIEVRTAYKLKSDFTSQRMGKEGHNSGTKGGGGSG GGSGGGSGGGGKPEELVRHFGDVEKAAVGVGVTPGAVYQWLQAGEI PPLRQSDIEVRTAYKLKSDFTSQRMGKEGHNSGTKGGGGSGGGS GGGGSGGGSGGGSGGGSGGGGSMSKLEKFTNCYSLSKTLRFKAIPVG KTQENIDNKRLLVEDEKRAEDYKGVKKLLDRYLSFINDVLHSIKLKNL NNYISLFRKKTTRTEKENKELENLEINLRKEIAKAFKGNEGYKSLFKKDII ETILPEFLDDKDEIALVNSFNGFTTAFTGFRNRENMFSEAKSTSIAF RCINENLTRYISNMDIFEKVDAIFDKHEVQEIKEKILNSDYDVEDFFEG EFFNFVLTQEGIDVYNAIIGGFVTESGEKIKGLNEYINLYNQKTKQKLP KFKPLYKQVLSDRESLSFYGEGYTSDEEVLEVFRNTLNKNSEIFFSSIK KLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEYD DIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLK EIIIQKVDEIYKVYGSSEKLFDADFVLEKSLKNDAVVAIMKDLLDSVKS FENYIKAFFGEGKETNRDESFYGDFVLAYDILLKVDHIYDAIRNYVTQK PYSKDKFKLYFQNPQFMGGWDKDKETDYRATILRYGSKYYLAIMDKK YAKCLQKIDKDDVNGNYEKINYKLLPGPNKMLPKVFFSKKWMAYYNP Petition 870260060590, de 22 / 06 / 2026, pág. 238 / 261 107 / 116 SEDIQKIYKNGTFKKGDMFNLNDCHKLIDFFKDSISRYPKWSNAYDFN FSETEKYKDIAGFYREVEEQGYKVSFESASKKEVDKLVEEGKLYMFQ IYNKDFSDKSHGTPNLHTMYFKLLFDENNHGQIRLSGGAELFMRRAS LKKEELVVHPANSPIANKNPDNPKKTTTLSYDVYKDKRFSEDQYELHI PIAINKCPKNIFKINTEVRVLLKHDDNPYVIGIDRGERNLLYIVVVDGKG NIVEQYSLNEIINNFGIRIKTDYHSLLDKKEKERFEARQNWTSIENIKE LKAGYISQVVHKICELVEKYDAVIALEDLNSGFKNSRVKVEKQVYQKF EKMLIDKLNYMVDKKSNPCATGGALKGYQITNKFESFKSMSTQNGFI FYIPAWLTSKIDPSTGFVNLLKTKYTSIADSKKFISSFDRIMYVPEEDLF EFALDYKNFSRTDADYIKKWKLYSYGNRIRIFRNPKKNNVFDWEEVC LTSAYKELFNKYGINYQQGDIRALLCEQSDKAFYSSFMALMSLMLQM RNSITGRTDVDFLISPVKNSDGIFYDSRNYEAQENAILPKNADANGAY NIARKVLWAIGQFKKAEDEKLDKVKIAISNKEWLEYAQTSVKHGSPKK KRKVSGGSSGGSPKKKRKV SEQ ID NO:14 Sequência nucleotídica de CRO-Cas12a ATGCCGAAGAAGAAGCGCAAGGTCTCCGGCGGCAGCTCCGGCG GCAGCAAGCCCGAGGAGCTGGTGCGGCACTTCGGCGATGTGGA GAAGGCTGCTGTTGGCGTGGGGGTTACCCCGGGGGCCGTCTACC AGTGGCTGCAGGCGGGCGAGATCCCTCCGCTGCGCCAGAGCGA CATCGAGGTTCGGACAGCGTACAAGCTCAAGAGCGATTTCACCAG CCAGAGGATGGGCAAGGAGGGGCATAATTCAGGCACCAAGGGCG GCGGGTCAGGGGGCGGCAGCGGGGGCGGCTCCGGGGGCGGCA AGCCAGAGGAGCTGGTTAGGCACTTCGGCGACGTCGAGAAGGCT GCCGTGGGGGTTGGCGTGACTCCAGGGGCCGTGTACCAGTGGC TCCAGGCCGGCGAGATTCCGCCACTGCGGCAGTCCGACATCGAG GTGCGCACCGCTTACAAGCTGAAGTCCGACTTCACCTCACAGAGG ATGGGGAAGGAGGGCCATAACAGCGGCACAAAGGGGGGCGGGG GCTCAGGCGGGGGCGGGAGCGGCGGCGGGGGCTCTGGGGGCG GCGGCAGCGGCGGGGGCGGCAGCGGGGGCGGCGGGTCGATGA Petição 870260060590, de 22 / 06 / 2026, pág. 239 / 261 108 / 116 GCAAGCTGGAGAAGTTCACGAACTGCTACTCCCTCAGCAAGACCC TGAGGTTCAAGGCGATCCCGGTCGGCAAGACCCAGGAGAACATC GACAACAAGCGGCTGCTGGTGGAGGACGAGAAGAGGGCTGAGG ACTACAAGGGCGTGAAGAAGCTCCTGGACCGCTACTACCTGTCCT TCATCAACGACGTGCTCCACAGCATCAAGCTCAAGAACCTGAACA ACTACATCAGCCTCTTCAGGAAGAAGACGCGCACCGAGAAGGAG AACAAGGAGCTCGAGAACCTGGAGATCAACCTGAGGAAGGAGAT CGCCAAGGCGTTCAAGGGCAACGAGGGCTACAAGTCCCTCTTCA AGAAGGACATCATCGAGACGATCCTCCCGGAGTTCCTGGACGAC AAGGACGAGATCGCCCTGGTCAACTCCTTCAACGGCTTCACCACG GCGTTCACCGGCTTCTTCAGGAACCGCGAGAACATGTTCAGCGA GGAGGCCAAGTCCACGAGCATCGCGTTCAGGTGCATCAACGAGA ACCTCACCCGCTACATCTCCAACATGGACATCTTCGAGAAGGTCG ACGCGATCTTCGACAAGCACGAGGTGCAGGAGATCAAGGAGAAG ATCCTGAACAGCGACTACGACGTCGAGGACTTCTTCGAGGGCGA GTTCTTCAACTTCGTCCTCACGCAGGAGGGCATCGACGTGTACAA CGCCATCATCGGTGGCTTCGTGACCGAGTCCGGCGAGAAGATCA AGGGCCTGAACGAGTACATCAACCTCTACAACCAGAAGACCAAGC AGAAGCTGCCGAAGTTCAAGCCCCTGTACAAGCAGGTGCTCTCC GACAGGGAGTCCCTCAGCTTCTACGGCGAGGGCTACACGAGCGA CGAGGAGGTCCTGGAGGTGTTCCGCAACACCCTCAACAAGAACA GCGAGATCTTCTCCAGCATCAAGAAGCTCGAGAAGCTGTTCAAGAACTTCGACGAGTACTCCAGCGCCGGCATCTTCGTCAAGAACGGC CCGGCGATCTCCACGATCAGCAAGGACATCTTCGGCGAGTGGAA CGTGATCCGCGACAAGTGGAACGCCGAGTACGACGACATCCACC TCAAGAAGAAGGCGGTGGTCACCGAGAAGTACGAGGACGACAGG CGCAAGTCCTTCAAGAAGATCGGCTCCTTCAGCCTCGAGCAGCTG CAGGAGTACGCCGACGCGGACCTGAGGCGTGGTCGAGAAGCTCAA GGAGATCATCATCCAGAAGGTCGACGAGATCTACAAGGTGTACGG CTCCAGCGAGAAGCTCTTCGACGCGGACTTCGTCCTCGAGAAGT Petition 870260060590, dated 06 / 22 / 2026, pages 240 / 261 109 / 116 CCCTGAAGAAGAACGACGCCGTGGTCGCGATCATGAAGGACCTC CTGGACTCCGTGAAGAGCTTCGAGAATTACATCAAGGCCTTCTTC GGCGAGGGCAAGGAGACGAACAGGGACGAGTCCTTCTACGGCG ACTTCGTCCTGGCCTACGACATCCTCCTGAAGGTGGACCACATCT ACGACGCGATCCGCAACTACGTGACCCAGAAGCCGTACAGCAAG GACAAGTTCAAGCTCTACTTCCAGAACCCCCAGTTCATGGGCGGC TGGGACAAGGACAAGGAGACGGACTACAGGGCGACCATCCTGCG CTACGGCAGCAAGTACTACCTCGCCATCATGGACAAGAAGTACGC GAAGTGCCTGCAGAAGATCGACAAGGACGACGTCAACGGCAACT ACGAGAAGATCAACTACAAGCTCCTGCCGGGCCCCAACAAGATG CTCCCGAAGGTGTTCTTCTCCAAGAAGTGGATGGCCTACTACAAC CCCAGCGAGGACATCCAGAAGATCTACAAGAACGGCACGTTCAA GAAGGGCGACATGTTCAACCTGAACGACTGCCACAAGCTCATCGA CTTCTTCAAGGACTCCATCAGCCGCTACCCGAAGTGGTCCAACGC CTACGACTTCAACTTCAGCGAGACCGAGAAGTACAAGGACATCGC GGGCTTCTACCGCGAGGTCGAGGAGCAGGGCTACAAGGTGTCCT TCGAGTCCGCCAGCAAGAAGGAGGTCGACAAGCTGGTGGAGGAG GGCAAGCTCTACATGTTCCAGATCTACAACAAGGACTTCTCCGAC AAGAGCCACGGCACGCCCAACCTGCACACCATGTACTTCAAGCTC CTGTTCGACGAGAACAACCACGGCCAGATCAGGCTGTCCGGCGG CGCCGAGCTCTTCATGAGGAGGGCGAGCCTGAAGAAGGAGGAGC TGGTGGTCCACCCCGCTAACAGCCCAATCGCGAACAAGAACCCGGACAACCCCAAGAAGACCACGACCCTGTCCTACGACGTGTACAAG GACAAGAGGTTCAGCGAGGACCAGTACGAGCTCCACATCCCGAT CGCGATCAACAAGTGCCCCAAGAACATCTTCAAGATCAACACCGA GGTCCGCGTGCTCCTGAAGCACGACGACAACCCCTACGTGATCG GCATCGACAGGGGCGAGAGGAACCTCCTGTACATCGTGGTCGTG GACGGCAAGGGCAACATCGTGGAGCAGTACTCCCTCAACGAGAT CATCAACAACTTCAACGGCATCAGGATCAAGACGGACTACCACAG CCTCCTGGACAAGAAGGAGAAGGAGAGGTTCGAGGCCCGCCAGA Petição 870260060590, de 22 / 06 / 2026, pág. 241 / 261 110 / 116 ACTGGACCTCCATCGAGAACATCAAGGAGCTGAAGGCGGGCTAC ATCAGCCAGGTCGTGCACAAGATCTGCGAGCTCGTCGAGAAGTA CGACGCCGTGATCGCCCTCGAGGACCTGAACTCCGGCTTCAAGA ACAGCCGCGTCAAGGTGGAGAAGCAGGTCTACCAGAAGTTCGAG AAGATGCTCATCGACAAGCTGAACTACATGGTGGACAAGAAGTCC AACCCCTGCGCTACGGGCGGCGCGCTGAAGGGCTACCAGATCAC CAACAAGTTCGAGAGCTTCAAGTCCATGAGCACTCAGAACGGCTT CATCTTCTACATCCCGGCGTGGCTCACGTCCAAGATCGACCCCAG CACCGGCTTCGTCAACCTCCTGAAGACGAAGTACACCTCCATCGC CGACAGCAAGAAGTTCATCTCCAGCTTCGACCGCATCATGTATGT GCCGGAGGAGGACCTGTTCGAGTTCGCCCTCGACTACAAGAACT TCTCCCGCACGGACGCGGACTACATCAAGAAGTGGAAGCTGTAC AGCTACGGCAACCGCATCCGCATCTTCAGGAACCCCAAGAAGAAC AACGTCTTCGACTGGGAGGAGGTGTGCCTGACCTCCGCGTACAA GGAGCTCTTCAACAAGTACGGCATCAACTACCAGCAGGGCGACAT CAGGGCTCTCCTGTGCGAGCAGAGCGACAAGGCCTTCTACTCCA GCTTCATGGCGCTGATGTCCCTCATGCTGCAGATGAGGAACTCGA TCACCGGCAGGACGGACGTGGACTTCCTCATCTCCCCGGTGAAG AACAGCGACGGCATCTTCTACGACTCCAGGAACTACGAGGCCCA GGAGAACGCGATCCTCCCAAAGAACGCGGACGCCAACGGCGCCT ACAACATCGCCAGGAAGGTCCTCTGGGCTATCGGCCAGTTCAAGA AGGCGGAGGACGAGAAGCTGGACAAGGTGAAGATCGCCATCAGCAACAAGGAGTGGCTCGAGTACGCCCAGACCTCGGTCAAGCACGG CAGCCCGAAGAAGAAGCGCAAGGTGTCCGGCGGCAGCTCCGGC GGCAGCCCGAAGAAGAAGCGCAAAGTGTGA SEQ ID NO: 15 linker of 15 amino acids. GGGSGGGSGGGSGGG SEQ ID NO: 16 linker of 32 amino acids. Petition 870260060590, dated 06 / 22 / 2026, pp. 242 / 261 111 / 116 SGGSSGGSSGSETPGTSESATPESSGGSSGGS SEQ ID NO: 17 ZmALS2 gRNA (maize acetolactate synthase homolog) GCTGCTCGATTCCGTCCCCA SEQ ID NO: 18 N15 bacteriophage OR3 operon of 20 bp. CTTTATAGCTGGCTATAATT SEQ ID NO: 19 gRNA from Glossy2 corn (ZmGL2) GTCACAGATCACAAACTTCAAATG SEQ ID NO: 20 Donor DNA serving as a repair template for HDR based on a 657 bp homology fragment taken from the coding region of ZmALS2, starting from the fourth codon (GCT for alanine) and ending at the 222nd codon (CGG for arginine). GCTCCCCCGGCCACCCCGCTCCGGCCGTGGGGCCCCACCGATC CCCGCAAGGGCGCCGACATCCTCGTCGAGTCCCTCGAGCGCTGC GGCGTCCGCGACGTCTTCGCCTACCCCGGCGGCGCGTCCATGGA GATCCACCAGGCACTCACCCGCTCCCCCCGTCATCGCCAACCACC TCTTCCGCCACGAGCAAGGGGAGGCCTTTGCGGCCTCCGGCTAC GCGCGCTCCTCGGGCCGCGTCGGCGTCTGCATCGCCACCTCCG GCCCCGGCGCCACCAACCTTGTCTCCGCGCTCGCCGACGCGCTG CTCGATTCTGTGCCGATCGTCGCTATCACCGGTCAGGTGTCGCGA CGCATGATTGGCACCGACGCCTTCCAGGAGACGCCCATCGTCGA GGTCACCCGCTCCATCACCAAGCACAACTACCTGGTCCTCGACGT CGACGACATCCCCCGCGTCGTGCAGGAGGCTTTCTTCCCTCGCCT CCTCTGGTCGACCGGGGCCGGTGCTTGTCGACATCCCCAAGGAC ATCCAGCAGCAGATGGCGGTGCCTGTCTGGGACAAGCCCATGAG TCTGCCTGGGTACATTGCGCGCCTTCCCAAGCCCCCTGCGACTG AGTTGCTTGAGCAGGTGCTGCGTCTTGTTGGTGAATCCCGG Petition 870260060590, dated 06 / 22 / 2026, pages 243 / 261 112 / 116 SEQ ID NO: 21 Donor DNA serving as the repair template for HDR based on an 859 bp homology fragment taken from the coding region of ZmGL2. TACCCACATAGCACGTACGGAGTATGAAGTCAATCAAAGGTGCAA ATCACGGCTTTGTATACTAGCTAGACTAGTCTATTACAGTACAGAC AATTACTGCAAACTAATTGTGCTGGTCGATTACTTCTGTTACTAAG GGCATGTACAACTTAGATACAACTGCACGGTACTCCAAGTATAAG ACACAACTAAAACACAATATAATACAGTGGTCATGTCTAAAACATG TGTCTTACCATATTTATTGTACCAATCAGGGCATTCAATAAATTAAA GTGACCAATCAGATAGTCTCATGTCTCGAATATAGAGCTAAGACAC CGTGTCTTCGTCAAAATACATGTCTTGAGATTTTTTACATTCACCCT CCTAGACACACTCTAAGACACAACTTAAGACACCCCACGGTACAT GCCCTAACTACGTACTCCCTCCGTCCTTTTTTATTTATCGTTTCTTT GGTCACAGATCTAGTGACCCAAATGCGGTGGGCTGGCGCTGGGG TTCAGCTGGGCGCACCTCATCGGCGACATCCCGTCGGCCGCCAC CTGCTTCAACAAGTGGGCGCAGATCCTCAGCGGCAAGAAGCCGG AAGCCACCGTCCTCACCCCGCCGAACCAGCCGCTGCAGGGCCAG TCCCCCGCGGCGCCGCGCTCCGTCAAGCAGGTCGGGCCCATGG AGGACCTCTGGCTGGTCCCCGCGGGCCGCGACATGGCGTGCTAC TCCTTCCACGTCAGCGACGCGGTGCTCAAGAAGCTCCACCAGCA GCAGAATGGGCGCCAGGACGCCGCCGCTGGCACCTTCGAGCTC GTGTCGGCGCTGGTGTGGCAGGCGGTGGCCAAGATCAGGGGCG ACGTGG SGGS (SEQ ID NO:22) GGGS (SEQ ID NO:23) GGGGS (SEQ ID NO:24) GSSGSS (SEQ ID NO:25) (GGGGS)6 (SEQ ID NO:26) SGSETPGTSESATPE (SEQ ID NO:27) Petition 870260060590, on 06 / 22 / 2026, page. 244 / 261 113 / 116 SSGSETPGTSESATPES (SEQ ID NO:28) SSGETPGTSESATPELK (SEQ ID NO:29) (GGGGS)3 (SEQ ID NO:30) (GGGGS)5 (SEQ ID NO:31) (GGGGS)io (SEQ ID NO:32) GGGGGGGG (SEQ ID NO:33) GSAGSAAGSGEF (SEQ ID NO:34) A(EAAAK)3A (SEQ ID NO:35) A(EAAAK)ioA (SEQ ID NO:36) NLS does antígeno Télécharger sur SV40 PKKKRKV (SEQ ID NO:37) NLS de nucleoplasmina KRPAATKKAGQAKKKK (SEQ ID NO:38) NLS de C-myc PAAKRVKLD (SEQ ID NO:39) NLS de C-myc RQRRNELKRSP (SEQ ID NO:40) NLS de hRNPA1 M9 NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:41) Domino IBB importina alfa RMRIZFKNKGKDTAELRRRRVEVSVELKAKKDEQILKRRNV (SEQ ID NO:42) Fibroid T protein NLS VSRKRPRP (SEQ ID NO:43) Fibroid T protein NLS PPKKARED (SEQ ID NO:44) NLS of human p53 PQPKKKPL (SEQ ID NO:45) NLS c-abl IV of camundongo Petition 870260060590, of 22 / 06 / 2026, p. 245 / 261 114 / 116 SALIKKKKMAP (SEQ ID NO:46) NLS of influenza virus NS1 DRLRR (SEQ ID NO:47) NLS of influenza virus NS1 PKQKKRK (SEQ ID NO:48) NLS of hepatitis virus delta antigen RKLKKKKKKL (SEQ ID NO:49) NLS of camundongo protein Mx1 RECKFLKRR (SEQ ID NO:50) NLS of human poly(ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:51) Glycocorticoid NLS of steroid hormone receptors (human) RCCLQAGMNLEARCTKK (SEQ ID NO:52) NLS of Agrobacterium VirD2 protein KRPRDRHDGELGGRKRAR (SEQ ID NO:53) (His)6 HHHHHH (SEQ ID NO:54) Ha Label YPYDVPDYA (SEQ ID NO:55) Epitope FLAG DYKDDDDK (SEQ ID NO:56) MYC Epitope EQKLISEEDL (SEQ ID NO:57)
[00208] All patents, patent publications, patent applications, scientific journal articles, books, technical references and the like discussed in this description are incorporated herein by reference in their entirety for all purposes.
[00209] It should be understood that the figures and descriptions in the description have been simplified to illustrate elements that are relevant to a Petition 870260060590, dated 06 / 22 / 2026, pp. 246 / 261 115 / 116 clear understanding of the description. It should be appreciated that the figures are presented for illustrative purposes and not as construction drawings. Omitted details and modifications or alternative embodiments are within the scope of experts in the art.
[00210] It may be appreciated that, in certain aspects of the description, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to provide an element or structure or to perform a given function or functions. Except where such substitution is not operational for practicing certain embodiments of the description, such substitution is considered within the scope of the description.
[00211] The examples presented here are intended to illustrate potential and specific implementations of the description. It may be appreciated that the examples are primarily intended for purposes of illustrating the description for those skilled in the art. Variations to these diagrams or to the operations described herein may exist without departing from the spirit of the description. For example, in certain cases, the steps or operations of the method may be performed or executed in a different order, or operations may be added, deleted, or modified.
[00212] Where a range of values is provided, it is understood that every intermediate value, down to the smallest fraction of a unit of the lower limit, unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Any narrower range between any stated values or unstated intermediate values in a stated range and any other stated or intermediate value in that stated range is encompassed. The upper and lower limits of such smaller ranges may be independently included or excluded from the range, and each range where one, neither, or both limits are included Petition 870260060590, dated 06 / 22 / 2026, pp. 247 / 261 116 / 116 in the lower bands is also encompassed within the technology, subject to any limit specifically excluded in the indicated band. Where the indicated band includes one or both of the limits, the bands excluding one or both of these included limits are also included.
[00213] In the preceding description, numerous specific details are presented to provide a more complete understanding of the present invention. However, it will be evident to a person skilled in the art that the invention described in this description can be practiced without one or more of these specific details. In other cases, features and procedures well known to those skilled in the art have not been described so as to avoid obscuring the invention. The embodiments described are for illustrative and not restrictive purposes. Although the present invention is described primarily with reference to specific embodiments, it is also anticipated that other embodiments will become apparent to those skilled in the art by reading this description, and it is intended that such embodiments are contained within the present inventive methods.Accordingly, the present description is not limited to the embodiments described above or illustrated in the drawings, and various embodiments and modifications may be made without departing from the scope of the claims below. Petition 870260060590, dated 06 / 22 / 2026, pp. 248 / 261
Claims
1 / 5 CLAIMS 1. Fusion protein, characterized in that it comprises a site-directed nuclease fused to a recruiter domain comprising a site-specific DNA binding domain, wherein the recruiter domain comprises a protein of the Cro repressor family.
2. Fusion protein, according to claim 1, characterized in that the site-directed nuclease comprises a CRISPR-associated nuclease.
3. Fusion protein according to claim 2, characterized in that the CRISPR-associated nuclease is selected from the group consisting of Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, Cas14 and nicase or deactivated versions thereof, optionally wherein: (a) the CRISPR-associated nuclease is a Cas9 enzyme; or (b) the CRISPR-associated nuclease is a Cas12a enzyme.
4. Fusion protein, according to any one of claims 1 to 3, characterized in that the recruiting domain: (a) comprises Cro of N15, Cro of P22, Cro of 434 or a combination of any of the same. (b) comprises an amino acid sequence having at least 90% identity with any of the SEQ ID NOs: 1, 3 or 4; and / or (c) comprises a dimerization domain.
5. Fusion protein, according to any one of claims 1 to 4, characterized in that the fusion protein comprises a linker located between the site-directed nuclease and the recruiting domain; optionally wherein the linker comprises Petition 870260060590, dated 06 / 22 / 2026, p. 249 / 261 2 / 5 any one of the SEQ ID Nos: 6, 7, 15 or 16.
6. Fusion protein, according to any one of claims 1 to 5, characterized in that the fusion protein comprises: (a) a nuclear localization signal; and / or (b) an amino acid sequence having at least 90% identity with SEQ ID NO: 11 or 13.
7. Recombinant nucleic acid, characterized in that it encodes the fusion protein, as defined in any one of claims 1 to 6.
8. DNA construct, characterized in that it comprises a promoter operatively linked to recombinant nucleic acid, as defined in claim 7.
9. DNA construct according to claim 8, characterized in that the promoter: (a) comprises at least one of an inducible promoter, a constitutive promoter, an egg cell-specific promoter, a pollen-specific promoter or an apical meristem tissue-specific promoter; and / or (b) is a ubiquitin 4 promoter, an actin promoter, a tubulin promoter, a MADS box promoter or a plant virus promoter.
10. Vector, characterized in that it comprises recombinant nucleic acid, as defined in claim 7, or DNA construct, as defined in claim 8 or 9.
11. Cell, characterized in that it comprises recombinant nucleic acid, as defined in claim 7, the DNA construct, as defined in claim 8 or 9, or the vector, as defined in claim 10.
12. Cell, according to claim 11, characterized in Petition 870260060590, dated 06 / 22 / 2026, pp. 250 / 261 3 / 5 by the fact that the cell is a plant cell; optionally wherein the plant cell is a corn plant cell, a soybean plant cell, a rice plant cell, a wheat plant cell or a sunflower plant cell.
13. A method for editing a nucleic acid, characterized in that it comprises: a. providing at least one fusion protein, as defined in any one of claims 1 to 6; b. providing the nucleic acid, wherein the nucleic acid comprises a first binding site and a target region comprising a portion of the nucleic acid, wherein the first binding site is within or adjacent to the target region; c. providing a donor polynucleotide comprising a donor nucleotide region and at least one recruiting sequence that is specifically bound by the recruiting domain of the at least one fusion protein; and d. contacting the nucleic acid and the donor polynucleotide with the at least one fusion protein, wherein the at least one fusion protein specifically binds to the first binding site of the nucleic acid and to the recruiting sequence of the donor polynucleotide, thereby resulting in an edit in the target region of the nucleic acid.
14. Method according to claim 13, characterized in that: (a) the first binding site is adjacent to a 5' end or a 3' end of the target region; (b) the nucleic acid further comprises a second binding site, wherein the second binding site is within or adjacent to the target region, and wherein at least one fusion protein binds specifically to the first binding site and the second nucleic acid binding site; optionally where the second binding site is adjacent to a 5' end or a 3' end of the target region; and / or (c) at least one recruiting sequence comprises a Cro OR3 operon sequence; optionally where the Cro OR3 operon sequence comprises an N15 OR3 operon sequence (optionally, SEQ ID NO:18), a P22 OR3 operon sequence, a 434 OR3 operon sequence, or a combination thereof.
15. Method according to claim 13 or 14, characterized in that the donor polynucleotide comprises: (a) at least one homology arm, wherein the at least one homology arm comprises a nucleotide sequence having complementarity with a portion of the target region of the nucleic acid; (b) at least two recruiting sequences; optionally wherein: (i) the donor polynucleotide comprises a first recruiting sequence adjacent to a 5' end of the donor nucleotide region and a second recruiting sequence adjacent to a 3' end of the donor nucleotide region; and / or (ii) the at least two recruiting sequences are not within the donor nucleotide region;(c) the site-directed nuclease of at least one fusion protein comprises a CRISPR-associated nuclease and the method further comprises providing at least one guide RNA, wherein the at least one guide RNA comprises a nucleotide sequence having complementarity with the first binding site and / or the second binding site of the nucleic acid; and / or (d) the edit in the target region of the nucleic acid is a substitution of at least a portion of the target region by at least one portion of the donor polynucleotide.
16. Cell genome, characterized in that it comprises recombinant nucleic acid, as defined in claim 7, the DNA construct, as defined in claim 8 or 9, or the vector, as defined in claim 10. Petition 870260060590, dated 06 / 22 / 2026, pp. 253 / 261