Cas polypeptide with modified PAM recognition

By modifying the PAM specificity of Cas-alpha polypeptides through amino acid incorporation, the targeting range and efficiency of genome editing are enhanced, addressing the limitations of existing CRISPR techniques.

JP2026510998APending Publication Date: 2026-04-10PIONEER HI BREED INTERNATIONAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing genome editing techniques using CRISPR-related Cas polypeptides are limited by their protospacer-adjacent motif (PAM) requirement, restricting targeting range and necessitating costly and time-consuming redesign for each target site.

Method used

Modifying the PAM specificity of Cas-alpha polypeptides by comparing and incorporating amino acids from orthologous Cas-alpha polypeptides into structurally similar positions to enhance targeting flexibility and efficiency.

Benefits of technology

Enhances the targeting range and specificity of Cas-alpha polypeptides, reducing the need for costly redesign and improving genome editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510998000091
    Figure 2026510998000091
  • Figure 2026510998000092
    Figure 2026510998000092
  • Figure 2026510998000093
    Figure 2026510998000093
Patent Text Reader

Abstract

This disclosure relates to methods for modifying the PAM specificity of Cas-alpha polypeptides. This disclosure also relates to Cas-alpha-10 polypeptides having modified PAM specificity, as well as methods and compositions for their use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 491,118, filed on March 20, 2023, and U.S. Provisional Patent Application No. 63 / 585,659, filed on September 27, 2023, which are hereby incorporated by reference in their entirety.

[0002] Reference to Electronically Submitted Sequence Listing The official copy of the sequence listing was electronically submitted as a sequence listing in XML format with the file name "108587 - WO - SEC - 1_Sequence_Listing_ST26", created on March 18, 2024, having a size of 119 kilobytes, and was submitted simultaneously with this specification. The sequence listing contained in this XML - formatted document is part of this specification and is hereby incorporated by reference in its entirety.

[0003] The present disclosure relates to the field of molecular biology, and more particularly, to compositions of novel polypeptide - derived Cas polypeptides, as well as compositions and methods for editing or modifying the genome of cells.

Background Art

[0004] Recombinant DNA technology has made it possible to insert DNA sequences at target genomic locations and / or modify specific endogenous chromosomal sequences. Site - specific integration techniques using site - specific recombination systems, like other types of recombination techniques, are used for the generation of targeted insertions of target genes in various organisms. Genome - editing techniques such as designer zinc - finger nucleases (ZFNs), transcription activator - like effector nucleases (TALENs), or homing meganucleases are available for the generation of target genome perturbations, but these systems have low specificity and require redesign for each target site, thus tending to use engineered nucleases that are costly and time - consuming to produce.

[0005] A newer technique has been identified that utilizes the adaptive immune system of archaea or bacteria, called CRISPR (Clustered Regular Arrangement Short Palindromic Sequence Repeat), which involves various domains of effector proteins encompassing diverse activities (DNA recognition, binding, and selective cleavage).

[0006] The protospacer-adjacent motif (PAM) requirement of CRISPR-related (Cas) polypeptides limits their targeting range (Shmakov et al, 2015; Zetsche et al, 2015; Burstein et al, 2017; Karvelis et al, 2020; Pausch et al, 2020). This is particularly evident in genome editing applications where the outcome depends on proximity to the desired editing cleavage site (e.g., template-free editing and homologous recombination repair) or in approaches that impose additional sequence requirements on target selection (e.g., base editing; Anzalone et al, 2020).

[0007] This specification discloses methods for modifying the PAM specificity of Cas-alpha polypeptides. Also disclosed are Cas-alpha 10 polypeptides having modified PAM specificity, as well as methods and compositions for their use. [Overview of the Initiative]

[0008] In a first aspect, the Disclosure provides a method for modifying the protospacer adjacent motif (PAM) specificity of a target Cas-alpha polypeptide, comprising: (a) comparing the PAM interaction (PI) domain of an ortholog Cas-alpha polypeptide with the PI domain of a target Cas-alpha polypeptide, wherein the ortholog Cas-alpha polypeptide has different PAM specificity from that of the target Cas-alpha polypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the ortholog Cas-alpha polypeptide; (c) incorporating the one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the ortholog Cas-alpha polypeptide into one or more structurally similar positions of the target Cas-alpha polypeptide to obtain a modified target Cas-alpha polypeptide; and (d) determining the PAM recognition of the modified target Cas-alpha polypeptide.

[0009] In some examples of methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the orthologous Cas-alpha polypeptide has different PAM recognition than the target Cas-alpha polypeptide.

[0010] In some examples of methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the target Cas-alpha polypeptide is Cas-alpha-10 polypeptide.

[0011] In some examples of methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the Cas-alpha-10 polypeptide contains an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, and the PI domain contains amino acids S63 to I196.

[0012] In some examples of methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the orthologous Cas-alpha polypeptide is Cas-alpha 1, Cas-alpha 2, Cas-alpha 3, Cas-alpha 4, Cas-alpha 5, Cas-alpha 6, Cas-alpha 7, Cas-alpha 8, or Cas-alpha 11.

[0013] In a second aspect, the disclosure provides synthetic or non-natural Cas-alpha-10 polypeptides comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87. In some examples of this second aspect, the synthetic or non-natural Cas-alpha-10 polypeptide has endonuclease activity. In some examples of this second aspect, the synthetic or non-natural Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of this second aspect, the synthetic or non-natural Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

[0014] In a third aspect, the Disclosure provides synthetic or non-natural Cas-alpha-10 polypeptides comprising a PAM interaction (PI) domain, wherein the PI domain is 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3' It recognizes PAM sequences including 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C). In some examples of this third embodiment, synthetic or unnatural Cas-alpha-10 polypeptides have endonuclease activity. In some examples of this third embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of this third embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In further examples of this third embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide has nickase activity. In yet another example of this third embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide is complexed with or operably associated with a reverse transcriptase.

[0015] In a fourth aspect, the Disclosure provides a synthetic or non-natural Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence comprises mutations to the amino acid positions of SEQ ID NO: 2, the mutations being: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y7 2D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or one or more combinations of the following mutations: combination of K85Q mutation and N92L mutation; combination of N88H mutation and Q89G mutation; Combinations of N88K and Q89G mutations; combinations of N88Q and Q89G mutations; combinations of K85S and N92L mutations; combinations of K85S and N92Q mutations; combinations of K85S and N92C mutations; combinations of K85S and N92H mutations; combinations of K85S and N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; K85N and N 92M mutation combination; K85N mutation and N92Q mutation combination; K85N mutation and N92I mutation combination; K85S mutation, N88D mutation and Q89G mutation combination; K85S mutation, N88H mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation combination; K85S mutation, N88D mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation combination; K85S mutation, N92L mutation and Q125R mutation combination;The combinations include: a combination of the Y72S mutation, the K85D mutation, the Q125R mutation, and the N127R mutation; a combination of the Y72A mutation, the K85S mutation, the Q89D mutation, the N92L mutation, and the Q125R mutation; a combination of the Y72C mutation, the N88H mutation, the Q89G mutation, and the Q125R mutation; a combination of the Y72C mutation, the N88D mutation, the Q89D mutation, the N92W mutation, and the Q125R mutation; or a combination of the K85Q mutation and the N92W mutation. In certain examples of this fourth embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes the K85S mutation. In other examples of this fourth embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes a combination of the K85N mutation and the N92L mutation. In some examples of this fourth embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide has endonuclease activity. In some examples of this fourth embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of this fourth embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In further examples of this fourth embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide has nickase activity. In yet another example of this fourth embodiment, the synthetic or unnatural Cas-alpha-10 polypeptide is complexed with or operably associated with a reverse transcriptase.

[0016] In a fifth aspect, the disclosure provides a synthetic composition comprising (a) a Cas-alpha-10 polypeptide having DNA-binding activity and comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, the complex of which binds to the target polynucleotide. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide has endonuclease activity. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this fifth embodiment, the Cas-alpha-10 polypeptide has nickase activity. In yet another example of this fifth embodiment, the Cas-alpha-10 polypeptide is complexed with or operably associated with reverse transcriptase.

[0017] In a sixth aspect of the disclosure, (a) a Cas-alpha-10 polypeptide having DNA-binding activity comprising a PAM interaction (PI) domain, the PI domain recognizes a PAM sequence on a target polynucleotide, the PAM sequences being 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'TTTY-3', 5'TTYTY-3', 5'TYTY-3', 5'TTNN-3', 5 The present invention provides a synthetic composition comprising (b) a Cas-alpha-10 polypeptide comprising '-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C); and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide forms a complex with the at least one guide polynucleotide, the complex of which binds to the target polynucleotide, and the guide polynucleotide. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide has endonuclease activity. In some examples of synthetic compositions, the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of synthetic compositions, the Cas-alpha-10 polypeptide is complexed with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity.In a further example of this sixth embodiment, the Cas-alpha-10 polypeptide has nickase activity. In yet another example of this sixth embodiment, the Cas-alpha-10 polypeptide is complexed with or operably associated with reverse transcriptase.

[0018] In a seventh aspect, the disclosure provides (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence comprises mutations to the amino acid positions of SEQ ID NO: 2, the mutations being: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation Mutations; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or combinations of the following mutations: combination of K85Q mutation and N92L mutation; N88H mutation and Q89G mutation Mutation combinations; N88K mutation and Q89G mutation combination; N88Q mutation and Q89G mutation combination; K85S mutation and N92L mutation combination; K85S mutation and N92Q mutation combination; K85S mutation and N92C mutation combination; K85S mutation and N92H mutation combination; K85S mutation and N92A mutation combination; K85S mutation and N92M mutation combination; K85N mutation and N92L mutation combination; K85N mutation and N92H mutation combination; K85N mutation and N92A mutation combination; K85N mutation and N92C mutation combination; K85N Combinations of mutations and N92M mutations; combinations of K85N mutations and N92Q mutations; combinations of K85N mutations and N92I mutations; combinations of K85S mutations, N88D mutations and Q89G mutations; combinations of K85S mutations, N88H mutations, Q89G mutations and N92L mutations; combinations of Y72A mutations, N88D mutations, Q89D mutations and Q125R mutations; combinations of K85S mutations, N88D mutations, Q89G mutations and N92L mutations; combinations of Y72A mutations, N88D mutations, Q89G mutations and Q125R mutations; combinations of K85S mutations, N92L mutations and Q125R mutations;(b) A Cas-alpha-10 polypeptide comprising one or more of the following: a combination of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; a combination of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; a combination of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; a combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or a combination of K85Q and N92W; (b) a guide polynucleotide comprising at least one guide polynucleotide having a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with at least one guide polynucleotide, the complex of which binds to the target polynucleotide;

[0019] In certain examples of this seventh embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes the K85S mutation. In other examples of this seventh embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes a combination of the K85N mutation and the N92L mutation. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide has endonuclease activity. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase. In some examples of the synthetic composition, the Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In further examples of this seventh embodiment, the Cas-alpha-10 polypeptide has nickase activity. In yet another example of this seventh embodiment, the Cas-alpha-10 polypeptide is either complexed with or operably associated with the reverse transcriptase.

[0020] In the eighth aspect, the Disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing a cell with a Cas-alpha-10 polypeptide having at least 90% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0021] In some examples of the eighth aspect, the PAM sequences recognized by the Cas-alpha-10 polypeptide are 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT Includes -3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C).

[0022] In some examples of the eighth aspect, the method further comprises providing a donor DNA molecule or a polynucleotide modification template to a cell.

[0023] In some examples of the eighth aspect, the cell is derived from or obtained from an animal, a fungus or a plant. In some examples of the eighth aspect, the plant is a monocotyledon or a dicotyledon. In some aspects, the plant is corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, triticale, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, peanut, potato, tobacco, Arabidopsis, safflower or tomato.

[0024] In some examples of the eighth aspect, the Cas-alpha10 polypeptide has endonuclease activity.

[0025] In some examples of the eighth aspect, at least one guide polynucleotide comprises a plurality of guide polynucleotides, and the Cas-alpha10 polypeptide is an inactivated Cas-alpha10 endonuclease complexed with a deaminase.

[0026] In some examples of the eighth aspect, the Cas-alpha10 polypeptide complexes with a heterologous protein domain via a linker, where the heterologous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this eighth aspect, the Cas-alpha10 polypeptide has nickase activity. In yet another example of this eighth aspect, the Cas-alpha10 polypeptide is complexed with or operably associated with a reverse transcriptase.

[0027] In a ninth aspect, the disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing a Cas-alpha-10 polypeptide comprising a PAM interaction (PI) domain to a cell, wherein the PI domain recognizes a PAM sequence on the target polynucleotide, and the PAM sequence is 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTY C-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC- 3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3 (b) comprising ', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C); (b) providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0028] In some examples of the ninth aspect, the method further comprises providing a donor DNA molecule or a polynucleotide modification template to a cell.

[0029] In some examples of the ninth aspect, the cell is derived from or obtained from an animal, fungus or plant. In some examples, the plant is a monocotyledon or dicotyledon. In some examples, the plant is maize, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, triticale, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, peanut, potato, tobacco, Arabidopsis, safflower or tomato.

[0030] In some examples of the ninth aspect, the Cas-alpha10 polypeptide has endonuclease activity.

[0031] In some examples of the ninth aspect, at least one guide polynucleotide comprises a plurality of guide polynucleotides, and the Cas-alpha10 polypeptide is an inactivated Cas-alpha10 endonuclease complexed with a deaminase.

[0032] In some examples of the ninth aspect, the Cas-alpha10 polypeptide is complexed with a heterologous protein domain via a linker, wherein the heterologous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this ninth aspect, the Cas-alpha10 polypeptide has nickase activity. In yet another example of this ninth aspect, the Cas-alpha10 polypeptide is complexed with or operably associated with a reverse transcriptase.

[0033] In a tenth aspect, the Disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing a cell with a Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence comprises mutations to the amino acid positions of SEQ ID NO: 2, the mutations being: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or the following combination of mutations: K85Q mutation and N9 2L mutation combination; N88H mutation and Q89G mutation combination; N88K mutation and Q89G mutation combination; N88Q mutation and Q89G mutation combination; K85S mutation and N92L mutation combination; K85S mutation and N92Q mutation combination; K85S mutation and N92C mutation combination; K85S mutation and N92H mutation combination; K85S mutation and N92A mutation combination; K85S mutation and N92M mutation combination; K85N mutation and N92L mutation combination; K85N mutation and N92H combination; K85N mutation and N92A mutation combination ;K85N mutation and N92C mutation combination;K85N mutation and N92M mutation combination;K85N mutation and N92Q mutation combination;K85N mutation and N92I mutation combination;K85S mutation, N88D mutation and Q89G mutation combination;K85S mutation, N88H mutation, Q89G mutation and N92L mutation combination;Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation combination;K85S mutation, N88D mutation, Q89G mutation and N92L mutation combination;Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation combination;The Cas-alpha-10 polypeptide is one or more of the following combinations: a combination of K85S mutation, N92L mutation, and Q125R mutation; a combination of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; a combination of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; a combination of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; a combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or a combination of K85Q and N92W mutation. (b) Recognizing a PAM sequence on a target polynucleotide; (c) Providing the cell with at least one guide polynucleotide containing a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (d) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0034] In some examples of this tenth embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes the K85S mutation. In other examples of this tenth embodiment, the amino acid sequence of the Cas-alpha-10 polypeptide includes a combination of the K85N mutation and the N92L mutation.

[0035] In some examples of the 10th embodiment, the PAM sequences recognized by the Cas-alpha-10 polypeptide are 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTT Includes T-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C).

[0036] In some examples of the tenth embodiment, the method further comprises providing a donor DNA molecule or a polynucleotide modification template to a cell.

[0037] In some examples of the tenth embodiment, the cells are derived from or obtained from animals, fungi, or plants. In some embodiments, the plants are monocots or dicots. In some examples, the plants are maize, soybeans, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, peanut, potato, tobacco, Arabidopsis, safflower, or tomato.

[0038] In some examples of the tenth embodiment, the Cas-alpha-10 polypeptide has endonuclease activity.

[0039] In some examples of the tenth embodiment, at least one guide polynucleotide comprises multiple guide polynucleotides, and the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with a deaminase.

[0040] In some examples of the tenth embodiment, the Cas-alpha-10 polypeptide complexes with a heterogeneous protein domain via a linker, where the heterogeneous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In further examples of this tenth embodiment, the synthetic or non-natural Cas-alpha-10 polypeptide has nickase activity. In yet another example of this tenth embodiment, the synthetic or non-natural Cas-alpha-10 polypeptide is complexed with or operably associated with reverse transcriptase.

[0041] In an eleventh aspect, the Disclosure provides animal or fungal cells comprising either a synthetic or non-natural Cas-alpha-10 polypeptide as described herein.

[0042] In a twelfth aspect, the Disclosure provides plant cells, plant parts, microplants, or plants comprising any of the synthetic or non-natural Cas-alpha-10 polypeptides described herein.

[0043] In a thirteenth aspect, the Disclosure provides a method for modifying the protospacer-adjacent motif (PAM) specificity of a target Cas-alpha-10 polypeptide, comprising: (a) comparing the PAM interaction (PI) domain of an ortholog Cas-alpha polypeptide with the PI domain of a target Cas-alpha-10 polypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the ortholog Cas-alpha polypeptide; (c) incorporating the one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the ortholog Cas-alpha polypeptide into one or more structurally similar positions of the target Cas-alpha-10 polypeptide to obtain a modified target Cas-alpha-10 polypeptide; and (d) determining the PAM recognition of the modified target Cas-alpha-10 polypeptide. In some examples of this thirteenth aspect, the ortholog Cas-alpha polypeptide has different PAM recognition than the target Cas-alpha-10 polypeptide. In some examples of this thirteenth embodiment, the Cas-alpha-10 polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2. In further examples of this thirteenth embodiment, the orthologous Cas-alpha polypeptide is Cas-alpha-1, Cas-alpha-2, Cas-alpha-3, Cas-alpha-4, Cas-alpha-5, Cas-alpha-6, Cas-alpha-7, Cas-alpha-8, or Cas-alpha-11.

[0044] Brief description of the drawings and sequence list This disclosure can be better understood from the following detailed description, which forms part of this application, as well as from the drawings and sequence listings accompanying this specification. [Brief explanation of the drawing]

[0045] [Figure 1]The phylogenetic relationships between several Cas-alpha orthollogs are shown. Three supergroups were identified (I, II, and III). Group I included clade 1 (Candidate Archaea and the phylum Aureabacteria (where the gene locus typically encodes Cas1, Cas2, and Cas4)). Group II consists of clade 2 (Aquificae (Sulfurihydrogenibium and Hydrogenivirga genera) and Deltaproteobacteria (Desulfovibrio genus)), clade 3 (Candidate Archaea (typically encoding Cas1, Cas2, and Cas4 at the locus)), clade 4 (Bacteroidetes (Prevotella and Bacteroides genera)), and clade 5 (Candidate Levibacterium). Group III included clade 6 (Clostridia (Dorea, Ruminococcus, Clostridium, Clostridioides, Peptocolstridium, Cellulosilyticym, Eubacterium)) and clade 7 (Bacilli (Bacillus, Acidi This included the genera Acidibacillus, Aneurinibacillus, Brevibacillus, Parageobacillus, and Alicyclobacillus, as well as clade 8 (Negativicutes (Phascolarctobacterium)) and clade 9 (Flavobacteriia (Flavobacterium)).The diamond symbol represents Cas-alpha-endonucleases 1-11, and their PAM recognition is described in U.S. Patent No. 10934536. [Figure 2] This specification shows an expression cassette expressing Cas-alpha-10 or a synthetic or unnatural Cas-alpha-10 polypeptide. [Figure 3] This document describes a method for modifying the PAM specificity of Cas-alpha polypeptides. [Figure 4] This document presents another method for modifying PAM specificity in Cas-alpha polypeptides. [Figure 5] Here is yet another method for modifying PAM specificity in Cas-alpha polypeptides. [Figure 6] This graph shows the number of target sites in the maize and human genomes targeted by several PAM variants in Example 2.

[0046] Sequence ID 1 is the PRT sequence of the Cas-alpha 10 polypeptide derived from Syntrophomonas palmitatica.

[0047] Sequence ID 2 is the PRT sequence of the first exemplary synthetic or non-natural Cas-alpha-10 polypeptide.

[0048] Sequence ID 3 is the PRT sequence of a Cas-alpha 4 polypeptide derived from an uncultured archaeon.

[0049] Sequence ID 4 is the PRT sequence of the Cas-alpha 8 polypeptide derived from Acidibacillus sulfuroxidans.

[0050] Sequence ID 5 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85A mutation compared to Sequence ID 2.

[0051] Sequence ID 6 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S mutation relative to Sequence ID 2.

[0052] Sequence ID 7 is the PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92R mutation relative to Sequence ID 2.

[0053] Sequence ID 8 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92K mutation relative to Sequence ID 2.

[0054] Sequence ID 9 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Q125K mutation relative to Sequence ID 2.

[0055] Sequence ID 10 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88H mutation relative to Sequence ID 2.

[0056] Sequence ID 11 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88H and Q89G mutations compared to Sequence ID 2.

[0057] Sequence ID 12 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88K mutation relative to Sequence ID 2.

[0058] Sequence ID 13 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88K and Q89G mutations compared to Sequence ID 2.

[0059] Sequence ID 14 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88Q mutation relative to Sequence ID 2.

[0060] Sequence ID 15 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88Q and Q89G mutations relative to Sequence ID 2.

[0061] Sequence ID 16 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Q125F mutation relative to Sequence ID 2.

[0062] Sequence ID 17 is a PRT sequence Cas-alpha-1 polypeptide derived from Candidatus Micrarchaeota archaeon.

[0063] Sequence ID 18 is a PRT sequence Cas-alpha-2 polypeptide derived from Candidatus Micrarchaeota archaeon.

[0064] Sequence ID 19 is a PRT sequence Cas-alpha-3 polypeptide derived from Candidatus Aureabacteria bacterium.

[0065] Sequence ID 20 is a PRT sequence Cas-alpha 5 polypeptide derived from Candidatus Micrarchaeota archaeon.

[0066] Sequence ID 21 is a PRT sequence Cas-alpha-6 polypeptide derived from uncultured archaea.

[0067] Sequence ID 22 is a PRT sequence Cas-alpha 7 polypeptide derived from Parageobacillus thermoglucosidasius.

[0068] Sequence ID 23 is a PRT sequence Cas-alpha-9 polypeptide derived from the genus Ruminococcus.

[0069] Sequence ID 24 is a PRT sequence Cas-alpha-11 polypeptide derived from Clostridium novyi.

[0070] Sequence ID 25 is a PRT sequence Cas-alpha-13 polypeptide derived from Clostridium paraputrificum.

[0071] Sequence ID 26 is a PRT sequence Cas-alpha-24 polypeptide derived from Bacillus toyonensis.

[0072] Sequence ID 27 is a PRT sequence Cas-alpha-29 polypeptide derived from the genus Peptoclostridium sp.

[0073] Sequence ID 28 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72V mutation compared to Sequence ID 2.

[0074] Sequence ID 29 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72E mutation relative to Sequence ID 2.

[0075] Sequence ID 30 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72Q mutation compared to Sequence ID 2.

[0076] Sequence ID 31 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72T mutation relative to Sequence ID 2.

[0077] Sequence ID 32 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72C mutation compared to Sequence ID 2.

[0078] Sequence ID 33 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72A mutation compared to Sequence ID 2.

[0079] Sequence ID 34 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72S mutation relative to Sequence ID 2.

[0080] Sequence ID 35 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72P mutation relative to Sequence ID 2.

[0081] Sequence ID 36 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72G mutation relative to Sequence ID 2.

[0082] Sequence ID 37 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72D mutation relative to Sequence ID 2.

[0083] Sequence ID 38 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72L mutation compared to Sequence ID 2.

[0084] Sequence ID 39 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85G mutation relative to Sequence ID 2.

[0085] Sequence ID 40 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85D mutation relative to Sequence ID 2.

[0086] Sequence ID 41 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85N mutation relative to Sequence ID 2.

[0087] Sequence ID 42 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N88D mutation relative to Sequence ID 2.

[0088] Sequence ID 43 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Q89D mutation relative to Sequence ID 2.

[0089] Sequence ID 44 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92D mutation relative to Sequence ID 2.

[0090] Sequence ID 45 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92C mutation relative to Sequence ID 2.

[0091] Sequence ID 46 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92A mutation relative to Sequence ID 2.

[0092] Sequence ID 47 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92G mutation relative to Sequence ID 2.

[0093] Sequence ID 48 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92F mutation relative to Sequence ID 2.

[0094] Sequence ID 49 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92E mutation relative to Sequence ID 2.

[0095] Sequence ID 50 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92H mutation relative to Sequence ID 2.

[0096] Sequence ID 51 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92I mutation relative to Sequence ID 2.

[0097] Sequence ID 52 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92L mutation relative to Sequence ID 2.

[0098] Sequence ID 53 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92V mutation relative to Sequence ID 2.

[0099] Sequence ID 54 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92W mutation relative to Sequence ID 2.

[0100] Sequence ID 55 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92Q mutation relative to Sequence ID 2.

[0101] Sequence ID 56 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92S mutation relative to Sequence ID 2.

[0102] Sequence ID 57 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92P mutation relative to Sequence ID 2.

[0103] Sequence ID 58 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92Y mutation relative to Sequence ID 2.

[0104] Sequence ID 59 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92T mutation relative to Sequence ID 2.

[0105] Sequence ID 60 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the N92M mutation relative to Sequence ID 2.

[0106] Sequence ID 61 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Q125R mutation relative to Sequence ID 2.

[0107] Sequence ID 62 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Q125P mutation relative to Sequence ID 2.

[0108] Sequence ID 63 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92L mutations relative to Sequence ID 2.

[0109] Sequence ID 64 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92Q mutations compared to Sequence ID 2.

[0110] Sequence ID 65 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92C mutations compared to Sequence ID 2.

[0111] Sequence ID 66 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92H mutations compared to Sequence ID 2.

[0112] Sequence ID 67 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92A mutations compared to Sequence ID 2.

[0113] Sequence ID 68 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85S and N92M mutations relative to Sequence ID 2.

[0114] Sequence ID 69 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92L mutations compared to Sequence ID 2.

[0115] Sequence ID 70 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92H mutations compared to Sequence ID 2.

[0116] Sequence ID 71 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92A mutations compared to Sequence ID 2.

[0117] Sequence ID 72 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92C mutations compared to Sequence ID 2.

[0118] Sequence ID 73 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92M mutations compared to Sequence ID 2.

[0119] Sequence ID 74 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92Q mutations compared to Sequence ID 2.

[0120] Sequence ID 75 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85N and N92I mutations compared to Sequence ID 2.

[0121] Sequence ID 76 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85S, N88D, and Q89G mutations compared to Sequence ID 2.

[0122] Sequence ID 77 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85S, N88H, Q89G, and N92L mutations compared to Sequence ID 2.

[0123] Sequence ID 78 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having Y72A, N88D, Q89D, and Q125R mutations compared to Sequence ID 2.

[0124] Sequence ID 79 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85S, N88D, Q89G, and N92L mutations compared to Sequence ID 2.

[0125] Sequence ID 80 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having Y72A, N88D, Q89G, and Q125R mutations compared to Sequence ID 2.

[0126] Sequence ID 81 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having K85S, N92L, and Q125R mutations compared to Sequence ID 2.

[0127] Sequence ID 82 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having Y72S, K85D, Q125R, and N127R mutations compared to Sequence ID 2.

[0128] Sequence ID 83 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the Y72A, K85S, Q89D, N92L, and Q125R mutations compared to Sequence ID 2.

[0129] Sequence ID 84 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having Y72C, N88H, Q89G, and Q125R mutations compared to Sequence ID 2.

[0130] Sequence ID 85 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having Y72C, N88D, Q89D, N92W, and Q125R mutations compared to Sequence ID 2.

[0131] Sequence ID 86 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85Q and N92L mutations compared to Sequence ID 2.

[0132] Sequence ID 87 is a PRT sequence of a synthetic or non-natural Cas-alpha-10 polypeptide having the K85Q and N92W mutations relative to Sequence ID 2.

[0133] Sequence ID 88 is the PRT sequence of a second exemplary synthetic or non-natural Cas-alpha-10 polypeptide.

[0134] Sequence ID 89 is the PRT sequence of a third exemplary synthetic or non-natural Cas-alpha-10 polypeptide. [Modes for carrying out the invention]

[0135] Novel CRISPR effector systems and elements comprising such systems are provided, for example, but not limited to, novel guide polynucleotide / endonuclease complexes, guide polynucleotides, guide RNA elements, Cas polypeptides and endonucleases, and compositions and methods for proteins comprising endonuclease functionality (domains). Also provided are compositions and methods for the direct delivery of endonucleases, cleavage-ready complexes, guide RNAs, and guide RNA / Cas polypeptide complexes. The disclosure further includes compositions and methods for genomic modification, gene editing, and insertion of a target polynucleotide into the cellular genome of a target sequence.

[0136] The terms used in the claims and specification are defined below unless otherwise specified. It should be noted that, as used herein and in the appended claims, the singular forms “a,” “an,” and “it” refer to multiple subjects unless the context clearly indicates otherwise.

[0137] definition As used herein, “nucleic acid” means polynucleotide, and includes single-stranded or double-stranded polymers of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may also include fragments and modified nucleotides. Accordingly, the terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” and “nucleic acid fragment” are used interchangeably to refer to single-stranded or double-stranded RNA, and / or DNA, and / or RNA-DNA polymers, and optionally include synthetic nucleotide bases, non-natural nucleotide bases, or modified nucleotide bases. Nucleotides (usually found in 5'-monophosphate form) are designated by single letters as follows: "A" for adenosine or deoxyadenosine (respectively RNA or DNA, respectively), "C" for cytosine or deoxycytosine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A, C, or T, "I" for inosine, and "N" for any nucleotide.

[0138] When the term "genome" is applied to prokaryotic and eukaryotic cells or biological cells, it encompasses not only chromosomal DNA found in the nucleus but also organelle DNA found in the cellular components of the cell (e.g., mitochondria or plastids).

[0139] "Open Reading Frame" is abbreviated as ORF.

[0140] The term "selective hybridization" refers to the fact that, under stringent hybridization conditions, a nucleic acid sequence hybridizes to a specific nucleic acid target sequence to a detectable degree (e.g., at least twice the background) than it hybridizes to a non-target nucleic acid sequence, substantially excluding the non-target nucleic acid. Sequences that selectively hybridize typically have at least about 80% or 90% sequence identity with each other, and up to 100% sequence identity (i.e., fully complementary).

[0141] The terms "stringent conditions" or "stringent hybridization conditions" refer to the conditions under which a probe will selectively hybridize to its target sequence in an in vitro hybridization assay. Stringent conditions will vary depending on the sequence and the environment. By controlling the stringency of the hybridization and / or washing conditions, it is possible to identify a target sequence that is 100% complementary to the probe (homologous probing). Alternatively, stringency conditions can be adjusted to tolerate sequence mismatch, allowing for the detection of lower similarity (non-homologous probing). Generally, probes are less than approximately 1000 nucleotides in length, and sometimes less than 500 nucleotides. Typically, stringent conditions are a pH of 7.0–8.3, at least about 30°C for short probes (e.g., 10–50 nucleotides) and at least about 60°C for long probes (e.g., more than 50 nucleotides), with a salt concentration of about 1.5 M Na ion concentration, typically about 0.01–1.0 M Na ion concentration. Stringent conditions can also be created by adding destabilizing agents such as formamide. Examples of low stringency conditions include hybridization at 37°C with a buffer consisting of 30–35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), and washing at 50–55°C with 1 × 2 × SSC (20 × SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Examples of moderate stringency conditions include hybridization at 37°C with 40-45% formamide, 1M NaCl, and 1% SDS, and washing with 0.5×~1× SSC at 55-60°C. Examples of high stringency conditions include hybridization at 37°C with 50% formamide, 1M NaCl, and 1% SDS, and washing with 0.1× SSC at 60-65°C.

[0142] "Homologousity" refers to similar DNA sequences. For example, a "homologous region to a genomic region" found on donor DNA is a region of DNA that has a sequence similar to a given "genomic region" in the genome of a cell or organism. A homologous region can be of any length sufficient to promote homologous recombination at the target site to be cleaved. For example, a homologous region may be at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5 These can include lengths of ~700, 5~800, 5~900, 5~1000, 5~1100, 5~1200, 5~1300, 5~1400, 5~1500, 5~1600, 5~1700, 5~1800, 5~1900, 5~2000, 5~2100, 5~2200, 5~2300, 5~2400, 5~2500, 5~2600, 5~2700, 5~2800, 5~2900, 5~3000, 5~3100 bases or longer. "Sufficient homology" means that two polynucleotide sequences have sufficient structural similarity to act as substrates in a homologous recombination reaction. This structural similarity includes the total length of each polynucleotide fragment and the sequence similarity of the polynucleotides. Sequence similarity can be described by the percentage of sequence identity over the entire length of the sequence and / or the percentage of sequence identity over a portion of the sequence length, including conserved regions with localized similarities such as consecutive nucleotides having 100% sequence identity.

[0143] As used herein, “genomic region” is a portion of a chromosome within the genome of a cell, located on either side of the target site, or including part of the target site. This genomic region includes at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, It may contain 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases, thereby giving this genomic region sufficient homology to undergo homologous recombination with the corresponding homologous region.

[0144] As used herein, "homologous recombination" (HR) involves the exchange of DNA fragments between two DNA molecules at a homologous site. The frequency of homologous recombination is influenced by many factors. The amount of homologous recombination and the relative ratio of homologous to non-homologous recombination vary among different organisms. Generally, the length of the homologous region affects the frequency of homologous recombination events; the longer the homologous region, the higher the frequency. The length of the homologous region required to observe homologous recombination also varies by species. While homology of at least 5kb has often been used, homologous recombination has been observed using homology as small as 25-50bp. For example, Singer et al., (1982) Cell 31:25-33; Shen and Huang, (1986) Genetics 112:441-57; Watt et al., (1985) Proc. Natl. Acad. Sci. USA 82:4768-72, Sugawara and Haber, (1992) Mol Cell Biol. 12:563-75; Rubnitz and Subramani, (1984) Mol Cell Biol 4:2253-8; Ayares et al., (1986) Proc. Natl. Acad. Sci. USA 83:5199-203; Liskay et al., (1987) Genetics 115:161-7.

[0145] In the context of nucleic acid sequences or polypeptide sequences, "sequence identity" or "identity" refers to the same nucleic acid base or amino acid residues in two sequences when aligned to obtain the greatest possible correspondence across a given comparison window.

[0146] The term "percentage of sequence identity" refers to a numerical value determined by comparing two optimally aligned sequences across the entire comparison window, where portions of the polynucleotide or polypeptide sequence within this comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where identical nucleic acid bases or amino acid residues occur in both sequences, obtaining the number of matched positions, dividing this number of matched positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Useful examples of sequence identity percentages include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage between 50% and 100%. These identities can be determined using any of the programs described herein.

[0147] Sequence alignment and the calculation of identity or similarity percentages may be determined using various comparison methods designed to detect homologous sequences, such as, but not limited to, the MegAlign® program of the LASERGENE Bioinformatics Computing Suite (DNASTAR Inc., Madison, WI). In connection with this application, if sequence analysis software is used for analysis, it will be understood that, unless otherwise specified, the analysis results will be based on the “default values” of the program mentioned. As used herein, “default values” means any set of values ​​or parameters that are initially loaded into the software when it is first initialized.

[0148] The "Clustal V method for alignment" corresponds to the alignment method found in the MegAlign® program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI) and is also referred to as Clustal V (described in Higgins and Sharp, (1989) CABIOS 5:151-153 and Higgins et al., (1992) Comput Appl Biosci 8:189-191). For multiple alignments, the default values ​​correspond to GAP PENALTY=10 and GAP LENGTH PENALTY=10. The default parameters for pairwise alignment of protein sequences using the Clustal method and for calculating identity percentage are KTUPLE=1, GAP PENALTY=3, WINDOW=5, and DIAGONALS SAVED=5. For nucleic acids, these parameters are KTUPLE=2, GAP PENALTY=5, WINDOW=4, and DIAGONALS SAVED=4. After sequence alignment using the Clustal V program, the "identity percentage" can be obtained by examining the "sequence distance" table in the same program. The "Clustal W method of alignment" is abbreviated as Clustal W (described by Higgins and Sharp, (1989) CABIOS 5:151-153, Higgins et al., (1992) Comput Appl Biosci 8:189-191) and corresponds to the alignment method found in the MegAlign® v6.1 program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, WI).Default parameters for multiple alignments (GAP PENALTY=10, GAP LENGTH PENALTY=0.2, Delay Divergen Seqs(%)=30, DNA Transition Weight=0.5, Protein Weight Matrix=Gonnet Series, DNA Weight Matrix=IUB). After aligning sequences using the Clustal W program, the "identity percentage" can be obtained by examining the "sequence distance" table in the same program. Unless otherwise specified, the sequence identity / similarity values ​​shown herein refer to values ​​obtained using GAP version 10 (GCG, Accelrys, San Diego, CA) with the following parameters: nucleotide sequence identity % and similarity % are obtained using a gap generation penalty weight of 50, a gap lengthening penalty weight of 3, and the nwsgapdna.cmp scoring matrix; amino acid sequence identity % and similarity % are obtained using a gap generation penalty weight of 8, a gap lengthening penalty of 2, and the BLOSUM62 scoring matrix (Henikoff and Henikoff, (1989) Proc. Natl. Acad. Sci. USA 89:10915). GAP uses the algorithm of Needleman and Wunsch, (1970) J Mol Biol 48:443-53 to find the overall alignment of two sequences that maximizes the number of matches and minimizes the number of gaps. GAP considers all possible alignments and gap locations and uses gap generation and gap elongation penalties per matched base to create alignments with the maximum number of matched bases and the minimum gaps. "BLAST" is a search algorithm provided by the National Center for Biotechnology Information (NCBI) used to find similar regions between biological sequences.This program compares nucleotide or protein sequences with a sequence database, calculates the statistical significance of the match, and identifies sequences sufficiently similar to the query sequence so that the similarity is not expected to have occurred randomly. BLAST reports the identified sequences and their local alignments to the query sequence. It will be well understood by those skilled in the art that many levels of sequence identity are useful for identifying polypeptides from other species or modified naturally or synthetically (such polypeptides have the same or similar function or activity). Useful examples of identity percentages include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage between 50% and 100%. In fact, any amino acid identity of 50% to 100%, such as 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity may be useful in describing this disclosure.

[0149] Polynucleotide sequences and polypeptide sequences, their variants, and the structural relationships between these sequences may be described herein by the terms “homology,” “homologous,” “substantially identical,” “substantially similar,” and “substantially corresponding,” which are used interchangeably. These refer to polypeptide sequences or nucleic acid sequences in which a change in one or more amino acids or nucleotide bases does not affect the function of the molecule (e.g., the ability to mediate gene expression or the ability to produce a particular phenotype). These terms also refer to modifications of nucleic acid sequences in which the functional properties of the resulting nucleic acid are substantially unchanged compared to the original, unmodified nucleic acid. These modifications include the deletion, substitution, and / or insertion of one or more nucleotides in the nucleic acid fragment. Substantially similar nucleic acid sequences can be defined by their ability to hybridize with the sequences exemplified herein (under moderate stringency conditions, e.g., 0.5 × SSC, 0.1% SDS, 60°C) or with any portion of the nucleotide sequences disclosed herein, and they are functionally equivalent to any of the nucleic acid sequences disclosed herein. By adjusting the stringency conditions, fragments with moderate similarity, such as homologous sequences from distantly related organisms, can be screened against fragments with high similarity, such as genes that replicate functional enzymes from closely related organisms. Washing after hybridization determines the stringency conditions.

[0150] A "centimorgan" (cM) or "map unit" is the distance between any pair of two polynucleotide sequences, binding genes, markers, target sites, loci, or any pair thereof, where 1% of the results of meiosis are recombinations. Therefore, a centimorgan is equivalent to the distance between any pair of two binding genes, markers, target sites, loci, or any pair thereof equal to the mean recombination frequency of 1%.

[0151] "Isolated" or "purified" nucleic acid molecules, polynucleotides, polypeptides, or proteins, or their biologically active portions, are substantially or essentially free from components that typically co-occur or interact with the polynucleotide or protein as they do in their naturally occurring environment. Therefore, isolated or purified polynucleotides, polypeptides, or proteins, if produced by recombinant technology, are substantially free from other cellular material or culture media, or, if chemically synthesized, substantially free from chemical precursors or other chemicals. Optimally, an "isolated" polynucleotide (optimally, a protein-coding sequence) does not contain sequences naturally adjacent to it in the genomic DNA of the organism from which it originates (i.e., sequences located at the 5' and 3' ends of the polynucleotide). For example, in various embodiments, an isolated polynucleotide may contain approximately 5kb, 4kb, 3kb, 2kb, 1kb, 0.5kb, or less than 0.1kb of nucleotide sequences naturally adjacent to it in the genomic DNA of the cell from which it originates. Isolated polynucleotides can be purified from cells in which they naturally occur. Isolated polynucleotides can also be obtained using conventional nucleic acid purification methods known to those skilled in the art. This term also includes recombinant polynucleotides and chemically synthesized polynucleotides.

[0152] The term "fragment" refers to a contiguous set of nucleotides or amino acids. In one embodiment, a fragment is a contiguous sequence of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 nucleotides. In one embodiment, a fragment is a contiguous sequence of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 amino acids. A fragment may or may not exhibit the function of a sequence that shares some percentage of identity over the entire length of the fragment.

[0153] The terms “functionally equivalent fragment” and “functionally equivalent fragment” are used interchangeably herein. These terms refer to an isolated nucleic acid fragment or part or subsequence of a polypeptide that exhibits the same activity or function as the longer sequence from which it originates. For example, a fragment retains the ability to alter gene expression or produce a particular phenotype, regardless of whether the fragment encodes an active protein. For instance, a fragment can be used to design a gene that produces a desired phenotype in a modified plant. A gene may be designed to be used repressively by conjugating the nucleic acid fragment, whether it encodes an active enzyme or not, to a plant promoter sequence in sense or antisense orientation.

[0154] A "gene" is a nucleic acid fragment that expresses a functional molecule, and includes, for example, certain proteins that include regulatory sequences preceding (5' non-coding) and succeeding (3' non-coding) a coding sequence. A "native gene" refers to a gene found in a naturally occurring location along with its own regulatory sequence.

[0155] The term "endogenous" refers to sequences or other molecules that are naturally present in cells or organisms. In some aspects, endogenous polynucleotides are typically found in the genome of a cell, i.e., they are not heterologous.

[0156] An "allele" is one of several alternative forms of a gene that occupies a given locus on a chromosome. If all alleles present at a given locus on a chromosome are the same, the plant is homozygous at that locus. If the alleles present at a given locus on a chromosome are different, the plant is heterozygous at that locus.

[0157] A "coding sequence" refers to a polynucleotide sequence that codes for a specific amino acid sequence. A "regulatory sequence" refers to a nucleotide sequence located upstream (5' non-coding sequence), within the coding sequence, or downstream (3' non-coding sequence) of the coding sequence, which affects the transcription, RNA processing, stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translational leader sequences, 5' non-coding sequences, 3' non-coding sequences, introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0158] A “mutant gene” is a gene that has been altered by human intervention. Such a “mutant gene” has a sequence that differs from the sequence of the corresponding non-mutant gene due to the addition, deletion, or substitution of at least one nucleotide. In certain aspects of this disclosure, the mutant gene includes alterations resulting from the guide polynucleotide / Cas polypeptide system disclosed herein. A mutant plant is a plant containing the mutant gene.

[0159] As used herein, “targeted mutation” means a mutation in a gene (referred to as a target gene), including innate genes, which is produced by altering a target sequence within a target gene using any method known to those skilled in the art, including a method comprising the inducible Cas polypeptide system disclosed herein.

[0160] The terms “knockout,” “gene knockout,” and “genetic knockout” are used interchangeably herein. Knockout refers to a cellular DNA sequence that has been partially or completely disabled by targeting with a Cas polypeptide, for example, the pre-knockout DNA sequence may have encoded an amino acid sequence and may have had a regulatory function (e.g., a promoter).

[0161] The terms “knock-in,” “gene knock-in,” “gene insertion,” and “genetic knock-in” are used interchangeably herein. Knock-in means replacing or inserting a DNA sequence at a specific DNA sequence in a cell by targeting using a Cas polypeptide (for example, by homologous recombination (HR), also using a suitable donor DNA polynucleotide). Examples of knock-ins include the specific insertion of a heterologous amino acid coding sequence in the coding region of a gene, or the specific insertion of a transcriptional regulatory element at a locus.

[0162] A "domain" refers to a continuous sequence of nucleotides (which may be RNA, DNA, and / or RNA-DNA combination sequences) or amino acids.

[0163] The terms "conserved domain" or "motif" refer to a set of polynucleotides or amino acids that are conserved at specific locations along the aligned sequence of evolutionarily related proteins. While amino acids at other locations may vary among homologous proteins, those highly conserved at specific locations are those essential to the structure, stability, or activity of the protein. Because they are identified by the high degree of conservation of the aligned sequence of their protein homolog family, they can be used as identifiers or "signatures" to determine whether a protein with a newly determined sequence belongs to a previously identified protein family.

[0164] A "codon-modifying gene," "codon-preferential gene," or "codon-optimized gene" is a gene with codon usage frequencies designed to mimic the preferred codon usage frequencies of a host cell.

[0165] An "optimized" polynucleotide is a sequence that has been optimized to enhance expression in a specific heterologous host cell.

[0166] A "plant-optimized nucleotide sequence" is a nucleotide sequence optimized for expression in plants, particularly for increased expression in plants. Plant-optimized nucleotide sequences include codon-optimized genes. Plant-optimized nucleotide sequences can be synthesized by modifying a nucleotide sequence encoding a protein, such as the Cas polypeptide disclosed herein, using one or more plant-preferred codons to enhance expression. For example, see Campbell and Gowri (1990) Plant Physiol. 92:1-11 for a discussion of the use of host-preferred codons.

[0167] A "promoter" is a region of DNA involved in the recognition and binding of RNA polymerase and other proteins that initiate transcription. A promoter sequence consists of proximal and more distal upstream elements, the latter of which are often referred to as enhancers. Enhancers are DNA sequences that can stimulate promoter activity and may be promoter-specific elements or heterologous elements inserted to enhance the promoter's level or tissue specificity. A promoter may be entirely derived from a native gene, or it may consist of different elements derived from different naturally occurring promoters and / or may include synthetic DNA segments. It is understood by those skilled in the art that various promoters can induce gene expression in various tissue or cell types, at various developmental stages, or under various environmental conditions. Furthermore, since the precise boundaries of regulatory sequences are often not fully defined, it is recognized that several variations of DNA fragments may possess identical promoter activity.

[0168] The promoter that causes the most gene expression in most cell types is generally called a “constitutive promoter.” The term “inducible promoter” refers to a promoter that selectively expresses coding sequences or functional RNA in response to the presence of endogenous or exogenous stimuli, or in response to the environment, hormones, chemical and / or growth signals, for example, by chemical compounds (chemical inducers). Examples of inducible or regulatory promoters include those induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals such as ethanol, abscisic acid (ABA), jasmonic acid, salicylic acid, or toxicity mitigators.

[0169] A "translation leader sequence" refers to a polynucleotide sequence located between the promoter sequence and the coding sequence of a gene. Translation leader sequences are located upstream of the translation initiation sequence of mRNA. Translation leader sequences can influence the processing of the primary transcript into mRNA, mRNA stability, or translation efficiency. Examples of translation leader sequences have been reported (e.g., Turner and Foster, (1995) Mol Biotechnol 3:225-236).

[0170] The terms "3' non-coding sequence," "transcription terminator," or "termination sequence" refer to DNA sequences located downstream of coding sequences, including polyadenylation recognition sequences and other sequences that encode regulatory signals that may affect mRNA processing or gene expression. Polyadenylation signals are typically characterized by their influence on the addition of a polyadenylate region to the 3' end of the mRNA precursor. The use of different 3' non-coding sequences is exemplified by Ingelbrecht et al., (1989) Plant Cell 1:671-680.

[0171] An "RNA transcript" refers to the product of RNA polymerase-catalyzed transcription of a DNA sequence. When an RNA transcript is a complete complementary copy of a DNA sequence, it is called a primary transcript or pre-mRNA. When an RNA transcript is an RNA sequence obtained by post-transcriptional processing of a primary transcript or preRNA, it is called mature RNA or mRNA. "Messenger RNA" or "mRNA" refers to RNA that does not contain introns and can be translated into protein by cells. "cDNA" refers to DNA that is complementary to the mRNA template and is synthesized from the mRNA template using reverse transcriptase. cDNA can be single-stranded or can be converted into a double-stranded form using the Klenow fragment of DNA polymerase I. "Sense" RNA refers to an RNA transcript containing mRNA that can be translated into protein intracellularly or in vitro. "Antisense RNA" means an RNA transcript that is complementary to all or part of a target primary transcript or mRNA and blocks the expression of the target gene (see, for example, U.S. Patent No. 5,107,065). The complementarity of antisense RNA may be complementary to any part of a particular gene transcript, i.e., a 5' non-coding sequence, a 3' non-coding sequence, an intron, or a coding sequence. "Functional RNA" refers to antisense RNA, ribozyme RNA, or other RNA that cannot be translated into polypeptides but nevertheless affects intracellular processes. The terms "complementary" and "reverse complementary" are used interchangeably herein with respect to mRNA transcripts and are intended to define antisense RNA for a message.

[0172] The term "genome" refers to the entire set of genetic material (genes and non-coding sequences) present in each cell, virus, or organelle of an organism; and / or a set of chromosomes inherited from one parent as a haploid unit.

[0173] The term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one is regulated by the other. For example, a promoter is operably linked to a coding sequence if it can control the expression of that coding sequence (i.e., the coding sequence is under the transcriptional control of the promoter). A coding sequence can be operably linked to a regulatory sequence in either a sense or antisense direction. In another example, a complementary RNA region can be operably linked directly or indirectly to the 5' end or 3' end of a target mRNA, or within the target mRNA, or the first complementary region is the 5' end of the target mRNA, and its complement is the 3' end of the target mRNA.

[0174] Generally, “host” refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, “host cell” refers to a single-celled organism cultured in vivo or in vitro eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell), or cell derived from a multicellular organism (e.g., cell line) into which a heterologous polynucleotide or polypeptide has been introduced. In some embodiments, the cell is selected from the group consisting of: archaeal cells, bacterial cells, eukaryotic cells, eukaryotic single-celled organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, avian cells, insect cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In some examples, the cell is in vitro. In some examples, the cell is in vivo.

[0175] The term "recombination" refers to the artificial combination of two otherwise separate sequence segments, for example, by chemical synthesis or by manipulating isolated nucleic acid segments using genetic engineering techniques.

[0176] The terms “plasmid,” “vector,” and “cassette” refer to linear or circular additional chromosomal elements that often carry genes that are not part of the cell’s central metabolism, and typically take the form of double-stranded DNA. Such elements can be linear or circular, single-stranded or double-stranded DNA or RNA, of any origin, self-replicating sequences, genomic integration sequences, phages, or nucleotide sequences, in which many nucleotide sequences are bound to or recombinant into specific constructs that can introduce the desired polynucleotide into the cell. A “transformation cassette” refers to a specific vector that contains a gene and, in addition to this gene, elements that promote the transformation of a particular host cell. An “expression cassette” refers to a specific vector that contains a gene and, in addition to this gene, elements that promote the expression of this gene in the host.

[0177] The terms “recombinant DNA molecule,” “recombinant DNA construct,” “expression construct,” “construct,” and “recombinant construct” are used interchangeably herein. A recombinant DNA construct includes an artificial combination of nucleic acid sequences (e.g., regulatory and coding sequences) that are not necessarily found together in nature. For example, a recombinant DNA construct may include regulatory and coding sequences from different origins, or regulatory and coding sequences from the same origin but sequenced in a way different from that found in nature. Such constructs may be used on their own or in conjunction with a vector. When a vector is used, the choice of vector depends on the method used to introduce the vector into host cells, as is well known to those skilled in the art. For example, plasmid vectors may be used. Those skilled in the art are familiar with the genetic elements that must be present on a vector in order to transform, select, and grow host cells without issue. Those skilled in the art will recognize that different, independent transformation events can result in different levels and patterns of expression (Jones et al., (1985) EMBO J 4:2411-2418; De Almeida et al., (1989) Mol Gen Genetics 218:78-86), and therefore, multiple events are usually screened to obtain a strain exhibiting the desired level and pattern of expression. Such screening can be performed by standard molecular biological assays, biochemical assays, and other assays, including Southern blot analysis of DNA, Northern blot analysis of mRNA expression, PCR, real-time quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting analysis of protein expression, enzyme or activity analysis, and / or phenotypic analysis.

[0178] The term “heterogeneous” refers to a difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. Non-limiting examples include: differences in taxonomic origin (e.g., a polynucleotide sequence obtained from Zea mays is heterogeneous if inserted into the genome of an Oryza sativa plant or a different subspecies or variety of Zea mays; or a polynucleotide obtained from bacteria introduced into plant cells); or differences in sequence (e.g., a polynucleotide sequence obtained from Zea mays, isolated, modified, and reintroduced into a maize plant). As used herein, “heterogeneous” in relation to a sequence may refer to a sequence originating from a different species, subspecies, or alien species, or, if from the same species, a sequence whose composition and / or genomic locus have been substantially altered from its natural form by intentional human intervention. For example, a promoter operably ligated to a heterologous polynucleotide may be from a different species than the one from which the polynucleotide originates, or, if from the same / similar species, one or both may be substantially modified from their original form and / or genomic locus, or the promoter may not be the native promoter of the operably ligated polynucleotide. Alternatively, one or more regulatory regions and / or polynucleotides shown herein may be entirely synthetic. In another example, the target polynucleotide for cleavage by a Cas polypeptide may be from a different organism than the Cas polypeptide. In yet another example, the Cas polypeptide and guide RNA may be introduced into the target polynucleotide along with an additional polynucleotide that acts as a template or donor for insertion into the target polynucleotide, where the additional polynucleotide is heterologous to the target polynucleotide and / or Cas polypeptide.

[0179] As used herein, the term "expression" refers to the production of a functional end product (e.g., mRNA, guide RNA, or protein) in its precursor or mature form.

[0180] A "mature" protein refers to a polypeptide that has been processed after translation (i.e., a polypeptide from which any prepeptides or propeptides present in the primary translation product have been removed).

[0181] A "precursor" protein refers to the primary product of mRNA translation (i.e., prepeptides and propeptides are still present). Prepeptides and propeptides can, but are not limited to, intracellular localization signals.

[0182] A "CRISPR" (Clustered Regular Arrangement Short Palindromic Sequence Repeat) locus refers to a specific gene locus that codes for components of the DNA cleavage system used, for example, by bacterial and archaeal cells to disrupt foreign DNA (Horvath and Barrangou, 2010, Science 327:167-170; International Publication No. 2007025097, published March 1, 2007). A CRISPR locus can consist of a CRISPR array containing short direct repeats (CRISPR repeats) separated by short variable DNA sequences (called spacers), which may be flanked by a variety of Cas (CRISPR-related) genes.

[0183] As used herein, “effector” or “effector protein” is a protein that encompasses activity including recognizing, binding to, and / or cleaving or nicking a polynucleotide target. An effector or effector protein may also be an endonuclease. The “effector complex” of the CRISPR system contains a Cas polypeptide involved in the recognition and binding of crRNA and the target. A portion of the Cas polypeptide component may further contain a domain involved in the cleavage of the target polynucleotide.

[0184] The term “Cas polypeptide” refers to polypeptides encoded by Cas (CRISPR-related) genes. Cas polypeptides include proteins encoded by genes at the Cas locus and also include adaptation molecules and interference molecules. Interference molecules in bacterial adaptive immune complexes include endonucleases. Cas endonucleases described herein include one or more nuclease domains. Examples of Cas endonucleases include, but are not limited to, the novel Cas-alpha polypeptides disclosed herein, Cas9 protein, Cpf1 (Cas12) protein, C2c1 protein, C2c2 protein, C2c3 protein, Cas3, Cas3-HD, Cas5, Cas7, Cas8, Cas10, or combinations or complexes thereof. When complexed with a suitable polynucleotide component, a Cas polypeptide may be a “Cas endonuclease” or “Cas effector protein” capable of recognizing, binding to, and optionally nicking or cleaving all or part of a specific polynucleotide target sequence. The Cas-alpha endonucleases of this disclosure include those having one or more RuvC nuclease domains.Cas polypeptide is further defined as a functional fragment or functional variant of a natural Cas polypeptide, or a sequence of at least 50, 50-100, at least 100, 100-150, at least 150, 150-200, at least 200, 200-250, at least 250, 250-300, at least 300, 300-350, at least 350, 350-400, at least 400, 400-450, at least 500, or more than 500 amino acids of a natural Cas polypeptide, and at least 50%, 50%-55%, at least 55%, 55%-60% A protein is further defined as one that shares sequence identity of %, at least 60%, 60%-65%, at least 65%, 65%-70%, at least 70%, 70%-75%, at least 75%, 75%-80%, at least 80%, 80%-85%, at least 85%, 85%-90%, at least 90%, 90%-95%, at least 95%, 95%-96%, at least 96%, 96%-97%, at least 97%, 97%-98%, at least 98%, 98%-99%, at least 99%, 99%-100%, or 100%, and retains the activity of at least a portion of the native sequence.

[0185] A “functional fragment” of a Cas polypeptide refers to a portion or sub-sequence of a Cas polypeptide of this disclosure that retains the ability to recognize and bind to a target site and to selectively unwind, nick, or cleave (introduce a single-strand or double-strand break into the target site). This portion or sub-sequence of a Cas polypeptide may include a complete or partial (functional) peptide of any one of its domains, for example, but not limited to, the entire functional portion of a Cas3 HD domain, the entire functional portion of a Cas3 helicase domain, or the entire functional portion of a protein (for example, Cas5, Cas5d, Cas7, and Cas8b1).

[0186] The term “functional variant” of a Cas polypeptide or Cas effector protein refers to a variant of a Cas effector protein disclosed herein that retains the ability to recognize, bind to, and selectively unwind, nick, or cleave all or part of a target sequence.

[0187] Cas endonucleases may also include multifunctional Cas endonucleases. The terms “multifunctional Cas endonuclease” and “multifunctional Cas endonuclease polypeptide” are used interchangeably herein and include references to a single polypeptide having Cas endonucleas function (including at least one protein domain capable of acting as a Cas endonuclease) and at least one other function, such as but not limited to complexing function (including at least a second protein domain capable of complexing with other proteins). In some embodiments, a multifunctional Cas endonuclease includes at least one additional protein domain (internal, upstream (5'), downstream (3'), or both internal 5' and 3', or any combination thereof) relative to the domain typical of a Cas endonuclease.

[0188] The terms “cascade” and “cascade complex” are used interchangeably herein and include references to multi-subunit protein complexes that can assemble with polynucleotides to form polynucleotide-protein complexes (PNPs). A cascade is a PNP that is polynucleotide-dependent in terms of complex assembly and stability, as well as the identification of target nucleic acid sequences. A cascade functions as a surveillance complex that finds and optionally binds to a target nucleic acid complementary to the variable targeting domain of a guide polynucleotide.

[0189] The terms “5'-cap” and “7-methylguanylate (m7G) cap” are used interchangeably herein. The 7-methylguanylate residue is located at the 5' end of messenger RNA (mRNA) in eukaryotic cells. RNA polymerase II (Pol II) transcribes mRNA in eukaryotic cells. Messenger RNA capping generally occurs as follows: the terminal 5' phosphate group of the mRNA transcript is removed by RNA terminal phosphatase, leaving two terminal phosphates. Guanosine monophosphate (GMP) is added to the terminal phosphate of the transcript by guanylyltransferase, leaving a 5'-5' triphosphate-linked guanine at the end of the transcript. Finally, the 7-nitrogen of this terminal guanine is methylated by methyltransferase.

[0190] In this specification, the term “5'-cap-less” is used to refer to RNA that has, for example, a 5'-hydroxyl group instead of a 5'-cap. Such RNA may be referred to, for example, “decapped RNA.” Since 5'-capped RNA is transported out of the nucleus, decapped RNA can accumulate more abundantly in the nucleus after transcription. One or more RNA components in this specification are decapped.

[0191] As used herein, the term “guide polynucleotide” refers to a polynucleotide sequence that can form a complex with a Cas polypeptide (e.g., the Cas polypeptides described herein) and that enables the Cas polypeptide to recognize, selectively bind to, and selectively cleave DNA target sites. The guide polynucleotide sequence may be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence).

[0192] The terms “functional fragment” of guide RNA, crRNA, or tracrRNA are used interchangeably herein and refer to a portion or sub-sequence of a guide RNA, crRNA, or tracrRNA of this disclosure that retains the ability to function as a guide RNA, crRNA, or tracrRNA, respectively.

[0193] The terms “functional variant” of guide RNA, crRNA, or tracrRNA (each) are used interchangeably herein and refer to a variant of the guide RNA, crRNA, or tracrRNA of this disclosure that retains the ability to function as guide RNA, crRNA, or tracrRNA, respectively.

[0194] The terms “single guide RNA” and “sgRNA” are used interchangeably herein and relate to the synthetic fusion of two RNA molecules, where a crRNA (CRISPR RNA) containing a variable targeting domain fused to a tracrRNA (trans-activated CRISPR RNA) (bound to a tracr mate sequence that hybridizes to the tracrRNA). A single guide RNA may comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of a type II CRISPR / Cas system capable of forming a complex with a type II Cas polypeptide, the guide RNA / Cas polypeptide complex capable of guiding the Cas polypeptide to a DNA target site, allowing the Cas polypeptide to recognize that DNA target site, optionally bind thereto, and optionally introduce a nick or cleave (introduce a single-strand or double-strand break).

[0195] The terms "variable targeting domain" or "VT domain" are used interchangeably herein and include nucleotide sequences that can hybridize (are complementary to) one strand (nucleotide sequence) of a double-stranded DNA target site. The complementarity rate between the first nucleotide sequence domain (VT domain) and the target sequence may be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. The length of the variable targeting domain may be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the variable targeting domain contains a contiguous sequence of 12 to 30 nucleotides. The variable targeting domain may consist of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.

[0196] The terms “Cas endonuclease recognition domain” or “CER domain” (of a guide polynucleotide) are used interchangeably herein and include a nucleotide sequence that interacts with the Cas polypeptide. The CER domain includes a (trans-acting) tracr nucleotide mate sequence followed by a tracr nucleotide sequence. The CER domain may consist of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence (see, for example, U.S. Patent Application Publication No. 20150059010A1, published February 26, 2015), or any combination thereof.

[0197] As used herein, the terms “guide polynucleotide / Cas polypeptide complex,” “guide polynucleotide / Cas polypeptide system,” “guide polynucleotide / Cas complex,” “guide polynucleotide / Cas system,” “inducible Cas system,” “polynucleotide-inducible endonuclease,” and “PGEN” are used interchangeably herein and refer to at least one guide polynucleotide and at least one Cas polypeptide that can form a complex, wherein the guide polynucleotide / Cas polypeptide complex can guide the Cas polypeptide to a DNA target site, enabling the Cas polypeptide to recognize the DNA target site, bind thereto, and optionally introduce or cleave a nick (introduce a single-strand or double-strand break). The guide polynucleotide / Cas polypeptide complex described herein may comprise a Cas polypeptide from any of the known CRISPR systems (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15; Zetsche et al., 2015, Cell 163,1-13; Shmakov et al., 2015, Molecular Cell 60,1-13) and a suitable polynucleotide component.

[0198] The terms “guide RNA / Cas polypeptide complex,” “guide RNA / Cas polypeptide system,” “guide RNA / Cas complex,” “guide RNA / Cas system,” “gRNA / Cas complex,” “gRNA / Cas system,” “RNA-inducible endonuclease,” and “RGEN” are used interchangeably herein and refer to at least one RNA component and at least one Cas polypeptide that can form a complex, wherein the guide RNA / Cas polypeptide complex can induce a Cas polypeptide to a DNA target site, enabling the Cas polypeptide to recognize, bind to, and optionally introduce or cleave a nick (introduce a single-strand or double-strand break) at the DNA target site.

[0199] The terms “target site,” “target sequence,” “target site sequence,” “target DNA,” “target locus,” “genomic target site,” “genomic target sequence,” “genomic target locus,” and “protospacer” are used interchangeably herein and refer to polynucleotide sequences, such as nucleotide sequences of any other DNA molecules in a cell, chromosome, episome, locus, or genome (e.g., chromosomal DNA, chloroplast DNA, mitochondrial DNA, plasmid DNA), that a guide polynucleotide / Cas polypeptide complex can recognize, bind to, and optionally nick or cleave. A target site may be an endogenous site in the cell’s genome, or a target site may be heterologous to the cell and therefore not naturally occurring in the cell’s genome, or a target site may be found at a genomic location heterologous to where it naturally occurs. Where used herein, the terms “endogenous target sequence” and “natural target sequence” are used interchangeably herein and refer to a target sequence that is endogenous or natural in the cell’s genome and is located at an endogenous or natural location within the cell’s genome. "Artificial target site" or "artificial target sequence" is used interchangeably herein and refers to a target sequence introduced into the genome of a cell. Such an artificial target sequence may be identical in sequence to an endogenous or natural target sequence in the cell's genome, but may be located at a different location in the cell's genome (i.e., a non-endogenous or non-developmental location).

[0200] As used herein, “protospacer adjacent motif” (PAM) refers to a short nucleotide sequence adjacent to a target sequence (protospacer) recognized (targeted) by the guide polynucleotide / Cas polypeptide system described herein. If a Cas polypeptide does not follow the target DNA sequence, it may not be able to properly recognize the target DNA sequence. The sequence and length of the PAM as used herein may vary depending on the Cas polypeptide or Cas polypeptide complex used. While the PAM sequence can be of any length, it is generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long.

[0201] "Modified target site," "modified target sequence," "altered target site," and "altered target sequence" are used interchangeably herein and refer to a target sequence disclosed herein that includes at least one modification compared to the unmodified target sequence. Such "modifications" include, for example, (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, (iv) chemical modification of at least one nucleotide, or (v) any combination of (i) to (iv).

[0202] A “modified nucleotide” or “edited nucleotide” refers to a target nucleotide sequence that includes at least one modification compared to its unmodified nucleotide sequence. Such “modifications” include, for example, (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, (iv) chemical modification of at least one nucleotide, or (v) any combination of (i) to (iv).

[0203] The terms "modifying a target site" and "altering a target site" are used interchangeably in this specification and refer to methods for generating a modified target site.

[0204] As used herein, “donor DNA” is a DNA construct containing the polynucleotide of interest to be inserted into a genomic target site by homologous recombination repair.

[0205] The term "polynucleotide modification template" includes a polynucleotide containing at least one nucleotide modification when compared to the nucleotide sequence to be edited. The nucleotide modification may be the substitution, addition, or deletion of at least one nucleotide. Optionally, the polynucleotide modification template may further include homologous nucleotide sequences adjacent to at least one nucleotide modification, which provide sufficient homology to the desired nucleotide sequence to be edited.

[0206] As used herein, the term "plant-optimized Cas polypeptide" refers to a Cas polypeptide (e.g., a multifunctional Cas polypeptide) encoded by a nucleotide sequence optimized for expression in plant cells or within plants.

[0207] "Plant-optimized nucleotide sequence encoding a Cas polypeptide," "plant-optimized construct encoding a Cas polypeptide," and "plant-optimized polynucleotide encoding a Cas polypeptide" are used interchangeably herein and refer to a nucleotide sequence encoding a Cas polypeptide, or a variant or functional fragment thereof, that is optimized for expression in plant cells or plants. A plant containing a plant-optimized Cas polypeptide includes a plant containing a nucleotide sequence encoding a Cas sequence, and / or a plant containing a Cas polypeptide. In some embodiments, the plant-optimized Cas polypeptide nucleotide sequence is a maize-optimized, rice-optimized, wheat-optimized, soybean-optimized, cotton-optimized, or canola-optimized Cas polypeptide.

[0208] The term “plant” generally includes whole plants, plant organs, plant tissues, seeds, plant cells, and their offspring. Plants are monocots or dicots. Plant cells include, but are not limited to, cells derived from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen, and microspores. “Plant element” is intended to refer to a whole plant or plant component, which may include differentiated and / or undifferentiated tissues (e.g., plant tissues, parts, and cell types, but not limited to these). In one embodiment, a plant element is one of the following: whole plants, seedlings, meristematic tissues, basic tissues, vascular tissues, epidermal tissues, seeds, leaves, roots, shoots, stems, flowers, fruits, stolons, bulbs, tubers, corms, keikis, shoots, buds, tumor tissues, and cells and cultures in various forms (e.g., single cells, protoplasts, embryos, callus tissue). It should be noted that because protoplasts lack a cell wall, they are not technically "incomplete" plant cells (like those found in nature with all their components intact). The term "plant organ" refers to a plant tissue or group of tissues that constitute a morphologically and functionally independent part of a plant. As used herein, "plant element" is synonymous with "part" of a plant and refers to any part of a plant, which may include distinct tissues and / or organs, and may be used interchangeably with the term "tissue" throughout. Similarly, "plant reproductive element" is intended to refer to any part of a plant that can initiate other plants by sexual or asexual reproduction of that plant, such as, but not limited to, seeds, seedlings, roots, shoots, cuttings, scions, grafts, stolons, bulbs, tubers, corms, keikis, or buds. Plant elements may be present in plants, or in plant organs, tissue cultures, or cell cultures.

[0209] "Offspring" includes any subsequent generations of a plant.

[0210] As used herein, the term “plant part” means plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant masses, and intact plant cells in plants or plant parts (e.g., embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, grains, spikes, rachis, pods, stalks, roots, root tips, anthers, etc., and the parts themselves). “Grain” is intended to mean mature seeds produced by growers for purposes other than cultivating or propagating seeds. Offspring, variants, and mutants of regenerated plants are also included in this disclosure if any of them contain transgenes.

[0211] The term “monocotyledonous” or “monocot” refers to a subclass of angiosperms, also known as the “monocotyledoneae,” whose seeds typically contain only one initial leaf or cotyledon. The term includes references to the whole plant, plant elements, plant organs (e.g., leaves, stems, roots), seeds, plant cells, and their offspring.

[0212] The term “dicotyledonous” or “dicot” refers to a subclass of angiosperms, also known as the “dicotyledoneae,” whose seeds typically contain only two initial leaves or cotyledons. The term includes references to the whole plant, plant elements, plant organs (e.g., leaves, stems, roots), seeds, plant cells, and their offspring.

[0213] The term “non-conventional yeast” as used herein refers to any yeast species that is neither a Saccharomyces species (e.g., S. cerevisiae) nor a Schizosaccharomyces species. (See “Non-Conventional Yeasts in Genetics, Biochemistry and Biotechnology: Practical Protocols”, K. Wolf, KDBreunig, G. Barth, Eds., Springer-Verlag, Berlin, Germany, 2003).

[0214] In relation to this disclosure, the terms “hybrid,” “cross-pollinated,” or “hybridized” mean the fusion of gametes by pollination to produce offspring (i.e., cells, seeds, or plants). The terms encompass both sexual pollination (pollination between one plant species and another) and self-pollination (i.e., pollen and ovules (or microspores and megaspores) originating from the same plant or genetically identical plants).

[0215] The term "gene transfer" refers to the transfer of a desired allele at a gene locus from one genetic background to another. For example, the transfer of a desired allele at a specific locus can be transmitted to at least one offspring plant through sexual mating between two parent plants, where at least one of the parent plants has the desired allele in its genome. Alternatively, for example, allele transfer may occur by recombination between two donor genomes in a fusion protoplast, where at least one of the donor protoplasts has the desired allele in its genome. The desired allele may be, for example, a transgene, a modified (mutated or edited) native allele, or a selected allele of a marker or QTL.

[0216] "Introduction" is intended to mean providing a polynucleotide, polypeptide, or polynucleotide-protein complex to a target such as a cell or organism in a manner that allows the component to enter the interior of the organism's cells or the cells themselves.

[0217] A “polynucleotide of interest” includes any nucleotide sequence encoding a protein or polypeptide that improves the desirability of a crop, i.e., a trait of agroeconomic value. Examples of polynucleotides of interest include, but are not limited to, polynucleotides encoding agroeconomically important traits, herbicide resistance, insecticide resistance, disease resistance, nematode resistance, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, cereal characteristics, marketability, phenotypic markers, or any other trait of agricultural or commercial importance. Polynucleotides of interest may also be available in sense-oriented or antisense orientation. Furthermore, multiple polynucleotides of interest may be used together or “stacked” to provide further benefits.

[0218] A "complex trait locus" includes genomic loci that contain multiple genetically linked transgenes.

[0219] The terms “reduced,” “less,” “slower,” and “increased,” “faster,” “enhanced,” and “greater,” as used herein, refer to a reduction or increase in the properties of a modified plant element or the resulting plant compared to an unmodified plant element or the resulting plant. For example, the reduction in properties may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5%-10%, at least 10%, 10%-20%, at least 15%, at least 20%, 20%-30%, at least 25%, at least 30%, 30%-40%, at least 35%, at least 40%, 40%-50%, at least 45%, at least 50%, 50%-60%, at least about 60%, 60%-70%, 70%-80%, at least 75%, at least about 80%, 80%-90%, at least about 90%, 90%-100%, at least 100%, 100%-200%, at least 200%, at least about 300%, at least about 400%, or even lower than the untreated control. The increase in properties may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% to 20%, at least 15%, at least 20%, 20% to 30%, at least 25%, at least 30%, 30% to 40%, at least 35%, at least 40%, 40% to 50%, at least 45%, at least 50%, 50% to 60%, at least about 60%, 60% to 70%, 70% to 80%, at least 75%, at least about 80%, 80% to 90%, at least about 90%, 90% to 100%, at least 100%, 100% to 200%, at least 200%, at least about 300%, at least about 400%, or more than that of the untreated control.

[0220] As used herein, the term “before” refers to the presence of one sequence upstream or 5' side of another sequence, with respect to sequence position.

[0221] The meanings of the abbreviations are as follows: "sec" means seconds, "min" means minutes, "h" means hours, "d" means days, "μL" means microliters, "mL" means milliliters, "L" means liters, "μM" means "micromoles," "mM" means "millimolees," "M" means "moles," "mmol" means millimoles, "μmole" or "umole" means micromoles, "g" means grams, "μg" or "ug" means micrograms, "ng" means nanograms, "U" means units, "bp" means base pairs, and "kb" means kilobases.

[0222] Classification of CRISPR-Cas systems CRISPR-Cas systems are classified according to the sequence and structural analysis of their components. Multiple CRISPR / Cas systems are described, including Class 1 systems with multi-subunit effector complexes (including types I, III, and IV) and Class 2 systems with single protein effectors (including types II, V, and VI) (Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15; Zetsche et al., 2015, Cell 163,1-13; Shmakov et al., 2015, Molecular Cell 60,1-13; Haft et al., 2005, Computational Biology, PLoS Comput Biol 1(6):e60; and Koonin et al. 2017, Curr Opinion Microbiology 37:67-78).

[0223] The CRISPR-Cas system comprises at least a CRISPR RNA (crRNA) molecule and at least one CRISPR-related (Cas) protein, forming a crRNA-ribonucleoprotein (crRNP) effector complex. The CRISPR-Cas locus contains an array of identical repeats interspersed with DNA targeting spacers encoding the crRNA component, and an operon-like unit of the cas gene encoding the Cas polypeptide component. The resulting ribonucleoprotein complex recognizes polynucleotides in a sequence-specific manner (Jore et al., Nature Structural & Molecular Biology 18, 529-536 (2011)). The crRNA functions as a guide RNA that sequence-specifically binds the effector (protein or complex) to a double-stranded DNA sequence by forming a so-called R-loop by moving the non-complementary strand and forming base pairs with the complementary DNA strand (Jore et al., 2011, Nature Structural & Molecular Biology 18, 529-536).

[0224] RNA transcripts at CRISPR loci (pre-crRNAs) are specifically cleaved at repeat sequences by CRISPR-associated (Cas) endoribonucleases in type I and type III systems, or by RNase III in type II systems. The number of CRISPR-associated genes at a given CRISPR locus can vary by species.

[0225] Different Cas genes encoding proteins with different domains exist in different CRISPR systems. A Cas operon contains genes encoding one or more effector endonucleases and other Cas polypeptides. Protein subunits include those described in Makarova et al. 2011, Nat Rev Microbiol. 2011 9(6):467-477; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; and Koonin et al. 2017, Current Opinion Microbiology 37:67-78. Domain types include those involved in expression (pre-crRNA processing, e.g., Cas 6 or RNase III), those involved in buffering (including effector modules for crRNA and target binding, and domains for target cleavage), those involved in adaptation (spacer insertion, e.g., Cas1 or Cas2), and those involved in support (regulation, or helper, or unknown function). Some domains may function for multiple purposes; in particular, Cas9, for example, includes a domain for endonuclease function and a domain for targeted cleavage.

[0226] Cas polypeptides recognize DNA target sites adjacent to protospacer fringe motifs (PAMs) via direct RNA-DNA base pairing, induced by a single CRISPR RNA (crRNA) (Jore, MM et al., 2011, Nat. Struct. Mol. Biol. 18:529-536, Westra, ER et al., 2012, Molecular Cell 46:595-605, and Sinkunas, T. et al., 2013, EMBO J. 32:385-394).

[0227] Class I CRISPR-Cas system Class I CRISPR-Cas systems include types I, III, and IV. A distinctive feature of class I systems is the presence of an effector endonuclease complex instead of a single protein. The cascade complex contains a nucleic acid-binding domain, which is the core fold of the RNA recognition motif (RRM) and various RAMP (repeat-associated mysterious protein) protein superfamilies (Makarova et al. 2013, Biochem Soc Trans 41, 1392-1400; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13: 1-15). The RAMP protein subunits include Cas5 and Cas7 (containing the crRNA-effector complex backbone) (where the Cas5 subunit binds to the 5' handle of crRNA and interacts with the larger subunit), and often Cas6, which loosely associates with the effector complex and typically functions as a repeat-specific RNase in pre-crRNA processing (Charpentier et al., FEMS Microbiol Rev 2015, 39:428-441; Niewoehner et al., RNA 2016, 22:318-329).

[0228] The type I CRISPR-Cas system includes a complex of effector proteins called a cascade (a CRISPR-associated complex for antiviral defense), which contains at least Cas5 and Cas7. The effector complex works together with a single CRISPR RNA (crRNA) and Cas3 to defend against invading viral DNA (Brouns, SJJ et al. Science 321:960-964; Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15). The type I CRISPR-Cas locus contains the signature gene cas3 (or variant cas3' or cas3") which encodes a metal-dependent nuclease with a single-stranded DNA (ssDNA)-stimulated superfamily 2 helicase that has been demonstrated to unwind double-stranded DNA (dsDNA) and RNA-DNA double strands (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13: 1-15). Following target recognition, the Cas3 endonuclease is recruited to the cascade-crRNA-target DNA complex to cleave and degrade the DNA target (Westra, E et al. (2012) Molecular Cell 46: 595-605, Sinkunas, T. et al. (2011) EMBO J. 30: 1335-1342, and Sinkunas, T. et al. (2013) EMBO J. J.32:385-394). In some type I systems, Cas6 may be an active endonuclease involved in crRNA processing, and Cas5 and Cas7 function as non-catalyzed RNA-binding proteins, but in type I-C systems, crRNA processing can be catalyzed by Cas5 (Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15). Type I systems can be divided into seven subtypes (Makarova et al. 2011, Nat Rev Microbiol. 2011 9(6):467-477; Koonin et al. 2017, Curr Opinion Microbiology 37:67-78).A modified type I CRISPR-associated complex (cascade) for adaptive antiviral defense, comprising at least the protein subunits Cas7, Cas5, and Cas6, wherein one of these subunits is synthetically fused to a Cas3 endonuclease or a modified restriction endonuclease FokI, is described (International Publication No. 2013098244, published July 4, 2013).

[0229] The type III CRISPR-Cas system, containing multiple Cas7 genes, targets ssRNA or ssDNA and functions as either an RNase or a target RNA-activated DNA nuclease (Tamulaitis et al., Trends in Microbiology 25(10)49-61, 2017). The Csm (type III-A) and Cmr (type III-B) complexes function as RNA-activated single-stranded (ss)DNases that link target RNA binding / cleavage with ssDNA degradation. Upon infection with foreign DNA, CRISPR RNA (crRNA)-induced binding of the Csm or Cmr complex to a new transcript recruits Cas10 DNase to actively transcribed phage DNA, resulting in degradation of both the transcript and phage DNA, but not the host DNA. The Cas10 HD domain is involved in ssDNase activity, and the Csm3 / Cmr4 subunit is involved in the endoribonuclease activity of the Csm / Cmr complex. The 3' adjacent sequence of the target RNA is important for Csm / Cmr ssDNase activity; base pairing with the 5' handle of the crRNA protects host DNA from degradation.

[0230] Type IV systems include typical Type I cas5 and cas7 domains in addition to the cas8-like domain, but may lack the CRISPR array that is characteristic of most other CRISPR-Cas systems.

[0231] Class II CRISPR-Cas system Class II CRISPR-Cas systems include types II, V, and VI. A distinctive feature of class II systems is the presence of a single Cas effector protein instead of an effector complex. Type II and V Cas polypeptides contain a RuvC endonuclease domain that employs an RNase H fold.

[0232] The Type II CRISPR / Cas system uses crRNA and tracrRNA (trans-activated CRISPR RNA) to guide Cas polypeptides to their DNA targets. The crRNA contains a spacer region complementary to one strand of the double-stranded DNA target, and a region that base-pairs with tracrRNA to form an RNA double-strand that cleaves the DNA target to the Cas polypeptide, leaving a blunt end. The spacer is obtained through a not-fully understood process involving the Cas1 and Cas2 proteins. Type II CRISPR / Cas loci typically include the cas9 gene in addition to the cas1 and cas2 genes (Chylinski et al., 2013, RNA Biology 10:726-737; Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15). Type II CRISPR-Cas loci can encode tracrRNAs that are partially complementary to the repeats within each CRISPR array and may include other proteins such as Csn1 and Csn2. The presence of cas9 near the cas1 and cas2 genes is characteristic of type II loci (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13: 1-15).

[0233] The V-type CRISPR / Cas system contains a single Cas polypeptide including Cpf1 (Cas12) (Koonin et al., Curr Opinion Microbiology 37:67-78, 2017), and unlike Cas9, it is an active RNA-induced endonuclease that does not necessarily require additional trans-activated CRISPR(tracr)RNA for target cleavage.

[0234] The type VI CRISPR-Cas system contains the cas13 gene, which encodes a nuclease independent of tracrRNA activity and possesses two HEPN (higher eukaryote and prokaryote nucleotide-binding) domains but lacks both an HNH domain and a RuvC domain. The majority of the HEPN domains contain conserved motifs that constitute the metal-independent endo-RNase active site (Anantharam et al., Biol Direct 8:15, 2013). Because of this characteristic, the type VI system is thought to act on RNA targets instead of the DNA targets common to other CRISPR-Cas systems.

[0235] In a first aspect, the disclosure provides a method for modifying the protospacer-adjacent motif (PAM) specificity of a target Cas-alpha polypeptide. As used herein, “protospacer-adjacent motif” (PAM) refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that can be recognized (targeted) by a guide polynucleotide / Cas polypeptide system. If a Cas polypeptide does not have a PAM sequence following the target DNA sequence, it may not be able to recognize the target DNA sequence properly. The sequence and length of the PAM as used herein may vary depending on the Cas polypeptide or Cas polypeptide complex used. The PAM sequence may be of any length, but is generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long.

[0236] In some embodiments, a method for modifying the PAM specificity of a target Cas-alpha polypeptide includes: (a) comparing the PAM interaction (PI) domain of a heterologous orthologous Cas-alpha polypeptide with the PI domain of the target Cas-alpha polypeptide, wherein the orthologous Cas-alpha polypeptide has different PAM recognition than the target Cas-alpha polypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-alpha polypeptide; (c) incorporating the one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the orthologous Cas-alpha polypeptide into one or more structurally similar positions of the target Cas-alpha polypeptide to obtain a modified target Cas-alpha polypeptide; and (d) determining the PAM recognition in the modified target Cas-alpha polypeptide.

[0237] As used herein, "orthologous Cas-alpha polypeptide" or "Cas-alpha ortholog" refers to a Cas-alpha polypeptide comprising a CRISPR-Cas polypeptide containing at least one zinc finger-like domain, at least one bridge helix-like domain, and a tri-split RuvC domain (including a discontinuous RuvC-I domain, a RuvC-II domain, and a RuvC- domain) in the size range of approximately 327 to 777 amino acids, and GxxxG, ExL, and Cx n C or Cx n It contains a (C,H) motif (where x represents any amino acid, and n = one or more amino acids).

[0238] Cas-alpha orthologues that can be used in the methods disclosed herein include Cas-alpha 1, Cas-alpha 2, Cas-alpha 3, Cas-alpha 4, Cas-alpha 5, Cas-alpha 6, Cas-alpha 7, Cas-alpha 8, Cas-alpha 10, Cas-alpha 11, Cas-alpha 13, Cas-alpha 24, and Cas-alpha 29.

[0239] Figure 1 shows the phylogenetic relationships among some Cas-alpha orthollogs, which are divided into three supergroups (I, II, and III). Group I includes clade 1 (Candidate Archaea and the phylum Aureabacteria (where the gene locus typically encodes Cas1, Cas2, and Cas4)). Group II consists of clade 2 (Aquificae (Sulfurihydrogenibium and Hydrogenivirga genera) and Deltaproteobacteria (Desulfovibrio genus)), clade 3 (Candidate Archaea (typically encoding Cas1, Cas2, and Cas4 at the locus)), clade 4 (Bacteroidetes (Prevotella and Bacteroides genera)), and clade 5 (Candidate Levibacterium). Group III includes clade 6 (Clostridia (Dorea, Ruminococcus, Clostridium, Clostridioides, Peptocolstridium, Cellulosilyticym, Eubacterium)), and clade 6 (Clostridia (Dorea, Ruminococcus, Clostridium, Clostridioides, Peptocolstridium, Cellulosilyticym, Eubacterium)). Group III includes clade 7 (Bacilli (Bacillus, Acidi This includes the genera Acidibacillus, Aneurinibacillus, Brevibacillus, Parageobacillus, and Alicyclobacillus, as well as clade 8 (Negativicutes (Phascolarctobacterium)) and clade 9 (Flavobacteriia (Flavobacterium)).The diamond symbol indicates orthologous Cas-alpha 1-11 endonucleases.

[0240] As used herein, “structurally similar position” refers to a coordinate in a polypeptide that occupies a similar three-dimensional position when its predicted (e.g., by a neural network-based information model such as AlphaFold) or determined (e.g., using cryo-electron microscopy (Cryo-EM), X-ray crystallography, and NMR spectroscopy) structure is aligned with or superimposed on the predicted or determined orthologous structure.

[0241] Methods for comparing the PI domains of target Cas-alpha polypeptides and orthologous Cas-alpha polypeptides include, but are not limited to, log-expected multiple sequence comparison (MUSCLE), multiple sequence comparison using Clustal Omega, root mean square distance (RMSD), distance matrix alignment (DALI), structural homology by environment-based alignment (SHEBA), combinatorial extension (CE), homologous structural alignment database (HOMSTRAD), protein structure classification (SCOP), FatCat, and PhyreStorm.

[0242] In some examples of the disclosed methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the step of selecting one or more amino acids from the PI domain of an ortholog Cas-alpha polypeptide includes comparing the PI domain sequence and / or structure of the ortholog Cas-alpha polypeptide with the PI domain sequence and / or structure of a target Cas-alpha polypeptide whose PAM specificity is to be altered, and substituting one or more amino acids from the ortholog Cas-alpha polypeptide into the Cas-alpha polypeptide target.

[0243] In some examples of the disclosed methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the step of selecting one or more polypeptide chains from the PI domain of an ortholog Cas-alpha polypeptide includes comparing the PI domain sequence and / or structure of the ortholog Cas-alpha polypeptide with the PI domain sequence and / or structure of a target Cas-alpha polypeptide whose PAM specificity is to be altered, and substituting one or more polypeptide chains from the ortholog Cas-alpha polypeptide with the Cas-alpha polypeptide target.

[0244] Methods for determining PAM recognition of a modified target Cas-alpha polypeptide include, but are not limited to, transcription and translation of the modified Cas-alpha polypeptide in cells or cell-free mixtures, complexing the modified Cas-alpha polypeptide with a guide RNA to form a ribonucleoprotein (RNP), incubating the RNP with a DNA species containing a fixed guide RNA target and a set of PAM sequences different from the target, capturing DNA molecules supporting target cleavage, sequencing the PAM region from the DNA species supporting target cleavage, and calculating a consensus PAM using a position-frequency matrix to summarize changes in PAM specificity.

[0245] In some examples of the disclosed methods for modifying the PAM specificity of a target Cas-alpha polypeptide, the target Cas-alpha polypeptide is a Cas-alpha-10 polypeptide. The target Cas-alpha polypeptide may be a Cas-alpha-10 polypeptide having an amino acid sequence that has at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with SEQ ID NO: 2.

[0246] In another aspect, the disclosure provides synthetic or non-natural Cas-alpha-10 polypeptides having modified PAM specificity. More specifically, the Specified Discloses Cas-alpha-10 polypeptides comprising a modified PAM interaction (PI) domain such that the resulting Cas-alpha-10 polypeptide recognizes a PAM sequence in a target polynucleotide other than 5'-TTC-3'.

[0247] In another aspect, the Disclosure provides a modified PAM-specific synthetic or non-natural Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with any one of SEQ ID NOs.

[0248] In this embodiment, a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity may contain an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NOs.

[0249] In another example of this embodiment, a synthetic or unnatural Cas-alpha-10 polypeptide having modified PAM specificity may comprise the amino acid sequences of SEQ ID NOs. 5-16 or 28-87.

[0250] In another embodiment, the present disclosure provides synthetic or non-natural Cas-alpha-10 polypeptides having modified PAM specificity, wherein the Cas-alpha-10 polypeptides are 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5 It includes a PI domain that recognizes PAM sequences containing '-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C).

[0251] In another aspect, the Disclosure provides a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity, wherein the Cas-alpha-10 polypeptide comprises an amino acid sequence having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with SEQ ID NO: 2, and the synthetic or non-natural Cas-alpha-10 polypeptide sequence has modified PAM specificity with respect to the amino acid positions of SEQ ID NO: 2. This includes K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K Includes 85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0252] In this embodiment, a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity may comprise an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, where the Cas-alpha-10 polypeptide sequence comprises K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N8 Includes 8Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0253] In another example of this embodiment, a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity may comprise the amino acid sequence of SEQ ID NO: 2, where the Cas-alpha-10 polypeptide sequence comprises the following mutations relative to the amino acid positions of SEQ ID NO: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation Mutations include: Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0254] In yet another aspect, the Disclosure provides a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity, wherein the Cas-alpha-10 polypeptide is at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or It may contain an amino acid sequence having at least 99% or 100% sequence identity, and the Cas-alpha-10 polypeptide sequence may be in the following combinations with respect to the amino acid position of SEQ ID NO: 2; a combination of K85Q mutation and N92L mutation; a combination of N88H mutation and Q89G mutation; a combination of N88K mutation and Q89G mutation; a combination of N88Q mutation and Q89G mutation; a combination of K85S mutation and N92L mutation; a combination of K85S mutation and N92Q mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92H mutation; Combinations of K85S and N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; combinations of K85S, N88D, and Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; Y72 Combinations of A mutation, N88D mutation, Q89D mutation, and Q125R mutation; combinations of K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; combinations of K85S mutation, N92L mutation, and Q125R mutation; combinations of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; combinations of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; combinations of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation;A combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or one or more combinations of K85Q and N92W mutation.

[0255] In this embodiment, a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity may comprise an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, wherein the Cas-alpha-10 polypeptide sequence comprises the following combinations with respect to the amino acid position of SEQ ID NO: K85Q mutation and N92L mutation Different combinations; combination of N88H mutation and Q89G mutation; combination of N88K mutation and Q89G mutation; combination of N88Q mutation and Q89G mutation; combination of K85S mutation and N92L mutation; combination of K85S mutation and N92Q mutation; combination of K85S mutation and N92C mutation; combination of K85S mutation and N92H mutation; combination of K85S mutation and N92A mutation; combination of K85S mutation and N92M mutation; combination of K85N mutation and N92L mutation; combination of K85N mutation and N92H mutation; combination of K85N mutation and N92 A mutation combination; K85N mutation and N92C mutation combination; K85N mutation and N92M mutation combination; K85N mutation and N92Q mutation combination; K85N mutation and N92I mutation combination; K85S mutation, N88D mutation and Q89G mutation combination; K85S mutation, N88H mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation combination; K85S mutation, N88D mutation and Q89G mutation and N92L mutation combination; Y72A mutation and N88D mutation and A combination of Q89G and Q125R mutations; a combination of K85S, N92L, and Q125R mutations; a combination of Y72S, K85D, Q125R, and N127R mutations; a combination of Y72A, K85S, Q89D, N92L, and Q125R mutations; a combination of Y72C, N88H, Q89G, and Q125R mutations; a combination of Y72C, N88D, Q89D, N92W, and Q125R mutations; or one or more combinations of K85Q and N92W mutations.

[0256] In another example of this embodiment, a synthetic or non-natural Cas-alpha-10 polypeptide having modified PAM specificity may comprise the amino acid sequence of SEQ ID NO: 2, where the Cas-alpha-10 polypeptide sequence comprises the following combinations with respect to the amino acid positions of SEQ ID NO: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92 L combination; K85S mutation and N92Q mutation combination; K85S mutation and N92C mutation combination; K85S mutation and N92H mutation combination; K85S mutation and N92A mutation combination; K85S mutation and N92M mutation combination; K85N mutation and N92L mutation combination; K85N mutation and N92H mutation combination; K85N mutation and N92A mutation combination; K85N mutation and N92C mutation combination; K85N mutation and N92M mutation combination; K8 Combinations of 5N mutation and N92Q mutation; combination of K85N mutation and N92I mutation; combination of K85S mutation, N88D mutation and Q89G mutation; combination of K85S mutation, N88H mutation, Q89G mutation and N92L mutation; combination of Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation; combination of K85S mutation, N88D mutation, Q89G mutation and N92L mutation; combination of Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation; K85S mutation A combination of different mutations, N92L mutation, and Q125R mutation; a combination of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; a combination of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; a combination of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; a combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or one or more combinations of K85Q and N92W mutation.

[0257] In further embodiments, the Disclosure provides a synthetic composition comprising (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising an amino acid sequence having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide, and forms a complex with the at least one guide polynucleotide, the complex of which binds to the target polynucleotide.

[0258] In one embodiment of this invention, the synthetic composition may include (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, the complex may include a guide polynucleotide that binds to the target polynucleotide.

[0259] In another embodiment of this model, the synthetic composition may include (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising one amino acid sequence of SEQ ID NOs. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, the complex of which may include the guide polynucleotide that binds to the target polynucleotide.

[0260] In another aspect, the Disclosure provides (a) a Cas-alpha-10 polypeptide having DNA-binding activity comprising a PAM interaction (PI) domain, the PI domain recognizing a PAM sequence on a target polynucleotide, the PAM sequences being 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5 The present invention provides a synthetic composition comprising (b) a Cas-alpha-10 polypeptide comprising '-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C); and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide forms a complex with the at least one guide polynucleotide, and the complex comprises the guide polynucleotide which binds to the target polynucleotide.

[0261] In yet another aspect, the Disclosure relates to (a) a Cas-alpha-10 polypeptide having DNA-binding activity, wherein it contains at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 9% of SEQ ID NO: 2. The Cas-alpha-10 polypeptide contains an amino acid sequence having 8%, or at least 99%, or 100% sequence identity, where the Cas-alpha-10 polypeptide has the following mutations relative to the amino acid position of SEQ ID NO: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation Cas-alpha-10 polypeptide containing the following mutations: Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation. (b) a synthetic composition comprising a guide polynucleotide having a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide.

[0262] In one example of this embodiment, the synthetic composition is (a) a Cas-alpha-10 polypeptide having DNA binding activity, comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas-alpha-10 polypeptide comprises K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation (b) a Cas-alpha-10 polypeptide comprising the following mutations: N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with at least one guide polynucleotide, the complex may include a guide polynucleotide that binds to the target polynucleotide.

[0263] In another example of this embodiment, the synthetic composition is (a) a Cas-alpha-10 polypeptide having DNA-binding activity, wherein the amino acid sequence of SEQ ID NO: 2, and K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation (b) a Cas-alpha-10 polypeptide comprising heterologous;N92L mutation;N92V mutation;N92W mutation;N92Q mutation;N92S mutation;N92P mutation;N92Y mutation;N92T mutation;N92M mutation;Q125R mutation; or Q125P mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with at least one guide polynucleotide, the complex may include a guide polynucleotide that binds to the target polynucleotide.

[0264] In yet another aspect, the Disclosure provides (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising an amino acid sequence having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with SEQ ID NO: 2 For amino acid position number 2, the following combinations are possible: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; K85S mutation and N92C mutation; K85S mutation and N92H mutation; K85S mutation and N92A mutation; K85S mutation and N92M mutation; K85N mutation and N92L mutation; K85N mutation and N9 2H combination; K85N mutation and N92A mutation combination; K85N mutation and N92C mutation combination; K85N mutation and N92M mutation combination; K85N mutation and N92Q mutation combination; K85N mutation and N92I mutation combination; K85S mutation, N88D mutation and Q89G mutation combination; K85S mutation, N88H mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation combination; K85S mutation, N88D mutation, Q89G mutation and N92L mutation combination; Y72A mutation and N88D Cas-alpha-10 polypeptide containing one or more of the following: mutations, Q89G mutation and Q125R mutation; K85S mutation, N92L mutation and Q125R mutation; Y72S mutation, K85D mutation and Q125R mutation and N127R mutation; Y72A mutation and K85S mutation and Q89D mutation and N92L mutation and Q125R mutation; Y72C mutation and N88H mutation and Q89G mutation and Q125R mutation; Y72C mutation and N88D mutation and Q89D mutation and N92W mutation and Q125R mutation; or K85Q and N92W combination;(b) a synthetic composition comprising a guide polynucleotide having a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide.

[0265] In one example of this embodiment, the synthetic composition is (a) a Cas-alpha-10 polypeptide having DNA-binding activity, comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, and comprising the following combinations with respect to the amino acid positions of SEQ ID NO: a combination of the K85Q mutation and the N92L mutation; a combination of the N88H mutation and the Q89G mutation. Combinations; combination of N88K mutation and Q89G mutation; combination of N88Q mutation and Q89G mutation; combination of K85S mutation and N92L mutation; combination of K85S mutation and N92Q mutation; combination of K85S mutation and N92C mutation; combination of K85S mutation and N92H mutation; combination of K85S mutation and N92A mutation; combination of K85S mutation and N92M mutation; combination of K85N mutation and N92L mutation; combination of K85N mutation and N92H mutation; combination of K85N mutation and N92A mutation; combination of K85N mutation and N92 C mutation combinations; K85N mutation and N92M mutation combination; K85N mutation and N92Q mutation combination; K85N mutation and N92I mutation combination; K85S mutation, N88D mutation and Q89G mutation combination; K85S mutation, N88H mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation combination; K85S mutation, N88D mutation, Q89G mutation and N92L mutation combination; Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation combination Cas-alpha-10 polypeptide containing one or more of the following combinations: K85S mutation, N92L mutation, and Q125R mutation; Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or K85Q and N92W combinations.and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with at least one guide polynucleotide, the complex may include the guide polynucleotide that binds to the target polynucleotide.

[0266] In another example of this embodiment, a synthetic or non-natural composition is (a) a Cas-alpha-10 polypeptide having DNA-binding activity, wherein the amino acid sequence of SEQ ID NO: 2 and the following combinations: combination of K85Q mutation and N92L mutation; combination of N88H mutation and Q89G mutation; combination of N88K mutation and Q89G mutation; combination of N88Q mutation and Q89G mutation; combination of K85S mutation and N92L mutation; combination of K85S mutation and N92Q mutation; K Combinations of 85S mutation and N92C mutation; combination of K85S mutation and N92H mutation; combination of K85S mutation and N92A mutation; combination of K85S mutation and N92M mutation; combination of K85N mutation and N92L mutation; combination of K85N mutation and N92H mutation; combination of K85N mutation and N92A mutation; combination of K85N mutation and N92C mutation; combination of K85N mutation and N92M mutation; combination of K85N mutation and N92Q mutation; K85N mutation A combination of the N92I mutation; a combination of the K85S mutation, N88D mutation, and Q89G mutation; a combination of the K85S mutation, N88H mutation, Q89G mutation, and N92L mutation; a combination of the Y72A mutation, N88D mutation, Q89D mutation, and Q125R mutation; a combination of the K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; a combination of the Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; a combination of the K85S mutation, N92L mutation, and Q125R mutation Cas-alpha-10 polypeptide containing one or more of the following combinations: Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or K85Q and N92W combination.(b) and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide, the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with at least one guide polynucleotide, the complex may include the guide polynucleotide that binds to the target polynucleotide.

[0267] In a further embodiment, the present disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) a Cas-alpha-10 polynucleotide having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87. (b) providing a lipeptide to a cell, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on a target polynucleotide; (b) providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0268] In one example of this embodiment, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-alpha-10 polypeptide having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any one of SEQ ID NOs. 5-16 or 28-87, wherein the Cas-alpha-10 polypeptide recognizes the PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0269] In another embodiment of this, a method for editing a target polynucleotide in a cell may include: (a) providing a cell with one of the Cas-alpha-10 polypeptides of sequence numbers 5-16 or 28-87, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0270] In some examples of the aforementioned methods for editing target polynucleotides in cells, the PAM sequences recognized by the Cas-alpha-10 polypeptide are 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY- Includes 3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C).

[0271] In some examples of the disclosed methods for editing a target polynucleotide in a cell, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0272] In a further embodiment, the Disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing a Cas-alpha-10 polypeptide comprising a PAM interaction (PI) domain to a cell, wherein the PI domain recognizes a PAM sequence on the target polynucleotide, and the PAM sequence is 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY- (b) comprising 3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C); (b) providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0273] In some embodiments of a method for editing a target polynucleotide in a cell, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0274] In another aspect, the present disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 9 The present invention provides cells with a Cas-alpha-10 polypeptide containing an amino acid sequence having 9% or 100% sequence identity, wherein the Cas-alpha-10 polypeptide sequence has the following amino acid positions relative to SEQ ID NO: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation. The Cas-alpha-10 polypeptide includes mutations such as Y72D mutation, Y72L mutation, K85G mutation, K85D mutation, K85N mutation, N88D mutation, Q89D mutation, N92D mutation, N92C mutation, N92A mutation, N92G mutation, N92F mutation, N92E mutation, N92H mutation, N92I mutation, N92L mutation, N92V mutation, N92W mutation, N92Q mutation, N92S mutation, N92P mutation, N92Y mutation, N92T mutation, N92M mutation, Q125R mutation, or Q125P mutation, and is located on the target polynucleotide. (b) Recognizing the PAM sequence of the Cas-alpha-10 polypeptide; (c) Providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (d) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0275] In one example of this embodiment, a method for editing a target polynucleotide in a cell may include: (a) providing a cell with a Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein Cas-alpha The polypeptide sequence is as follows, relative to the amino acid position of SEQ ID NO: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; (b) providing the cell with at least one guide polynucleotide comprising a Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on a target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0276] In another example of this embodiment, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-alpha-10 polypeptide containing the amino acids of SEQ ID NO: 2, wherein the Cas-alpha-10 polypeptide has K85S mutations; K85A mutations; N92R mutations; N92K mutations; Q125K mutations; N88H mutations; N88K mutations at the amino acid positions of SEQ ID NO: 2 ;N88Q mutation;Q125F mutation;Y72V mutation;Y72E mutation;Y72Q mutation;Y72T mutation;Y72C mutation;Y72A mutation;Y72S mutation;Y72P mutation;Y72G mutation;Y72D mutation;Y72L mutation;K85G mutation;K85D mutation;K85N mutation;N88D mutation;Q89D mutation;N92D mutation;N92C mutation;N92A mutation;N92G mutation;N92F mutation;N92E mutation;N92 (b) providing the cell with at least one guide polynucleotide comprising an H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on a target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex of which binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0277] In another aspect, the present disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) an nucleotide having at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or 100% sequence identity with SEQ ID NO: 2 The present invention provides cells with Cas-alpha-10 polypeptide containing a no-acid sequence, wherein the Cas-alpha-10 polypeptide sequence is a combination of the following amino acid positions in SEQ ID NO: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; K85S mutation and N92C mutation; K85S mutation and N92H mutation. se; combination of K85S mutation and N92A mutation; combination of K85S mutation and N92M mutation; combination of K85N mutation and N92L mutation; combination of K85N mutation and N92H mutation; combination of K85N mutation and N92A mutation; combination of K85N mutation and N92C mutation; combination of K85N mutation and N92M mutation; combination of K85N mutation and N92Q mutation; combination of K85N mutation and N92I mutation; combination of K85S mutation, N88D mutation and Q89G mutation; combination of K85S mutation, N88H mutation, Q89G mutation and N92L mutation; Y7 Combinations of 2A mutation, N88D mutation, Q89D mutation, and Q125R mutation; combinations of K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; combinations of K85S mutation, N92L mutation, and Q125R mutation; combinations of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; combinations of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; combinations of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation;(b) comprising one or more combinations of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or K85Q and N92W mutation, wherein the Cas-alpha-10 polypeptide recognizes the PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide containing a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0278] In one example of this embodiment, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas-alpha-10 polypeptide sequence comprises the following combinations with respect to the amino acid position of SEQ ID NO: K85Q mutation and N92L mutation Combinations of; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; K85S mutation and N92C mutation; K85S mutation and N92H mutation; K85S mutation and N92A mutation; K85S mutation and N92M mutation; K85N mutation and N92L mutation; K85N mutation and N92H mutation; K85N mutation and N92A mutation; K85N mutation A combination of K85N mutation and N92C mutation; a combination of K85N mutation and N92M mutation; a combination of K85N mutation and N92Q mutation; a combination of K85N mutation and N92I mutation; a combination of K85S mutation, N88D mutation and Q89G mutation; a combination of K85S mutation, N88H mutation, Q89G mutation and N92L mutation; a combination of Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation; a combination of K85S mutation, N88D mutation, Q89G mutation and N92L mutation; a combination of Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation; a combination of K85S mutation and N92 The Cas-alpha-10 polypeptide includes one or more of the following combinations: L mutation and Q125R mutation; Y72S mutation, K85D mutation, Q125R mutation and N127R mutation; Y72A mutation, K85S mutation, Q89D mutation, N92L mutation and Q125R mutation; Y72C mutation, N88H mutation, Q89G mutation and Q125R mutation; Y72C mutation, N88D mutation, Q89D mutation, N92W mutation and Q125R mutation; or K85Q and N92W mutation, wherein the Cas-alpha-10 polypeptide recognizes the PAM sequence on the target polynucleotide;(b) providing the cell with at least one guide polynucleotide containing a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0279] In another example of this embodiment, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-alpha-10 polypeptide containing the amino acids of SEQ ID NO: 2, wherein the Cas-alpha-10 polypeptide sequence has the following combinations with respect to the amino acid positions of SEQ ID NO: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation 2L combination; K85S mutation and N92Q mutation combination; K85S mutation and N92C mutation combination; K85S mutation and N92H mutation combination; K85S mutation and N92A mutation combination; K85S mutation and N92M mutation combination; K85N mutation and N92L mutation combination; K85N mutation and N92H combination; K85N mutation and N92A mutation combination; K85N mutation and N92C mutation combination; K85N mutation and N92M mutation combination; K85N mutation and N92Q mutation combination Combination; K85N mutation and N92I mutation; K85S mutation, N88D mutation and Q89G mutation; K85S mutation, N88H mutation, Q89G mutation and N92L mutation; Y72A mutation, N88D mutation, Q89D mutation and Q125R mutation; K85S mutation, N88D mutation, Q89G mutation and N92L mutation; Y72A mutation, N88D mutation, Q89G mutation and Q125R mutation; K85S mutation, N92L mutation and Q125R mutation; Y72S mutation and K The Cas-alpha-10 polypeptide includes one or more of the following combinations: 85D mutation, Q125R mutation, and N127R mutation; Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or K85Q mutation and N92W mutation, wherein the Cas-alpha-10 polypeptide recognizes the PAM sequence on the target polynucleotide;(b) providing the cell with at least one guide polynucleotide containing a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide.

[0280] In some embodiments of methods for editing target polynucleotides in cells, the PAM sequences recognized by the Cas-alpha-10 polypeptide are 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3 Includes ', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N=A, C, G, or T, Y=T or C, D=G, A, or T, W=A or T, H=A, T, or C).

[0281] In some embodiments of a method for editing a target polynucleotide in a cell, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0282] CRISPR-Cas system components Caspylene Numerous proteins, including those involved in adaptation (spacer insertion), interference (effector module target binding, target nicking, or cleavage—e.g., endonuclease activity), expression (pre-crRNA processing), regulation, or other functions, can be encoded by the CRISPR cas operon.

[0283] The two proteins Cas1 and Cas2 are conserved across many CRISPR systems (as described, e.g., Koonin et al., Curr Opinion Microbiology 37:67-78, 2017). Cas1 is a metal-dependent DNA-specific endonuclease that generates double-stranded DNA fragments. In some systems, Cas1 forms a stable complex with Cas2, which is essential for spacer acquisition and insertion for CRISPR systems (Nunez et al., Nature Str Mol Biol 21:528-534, 2014).

[0284] Many other proteins, including Cas4 (which may have similarities to the RecB nuclease), have been identified in various systems and are thought to play a role in capturing novel viral DNA sequences for incorporation into CRISPR arrays (Zhang et al., PLOS One 7(10):e47232, 2012).

[0285] Some proteins can encompass multiple functions. For example, Cas9, a signature protein of the Class II system, has been shown to be involved in pre-crRNA processing, target binding, and target cleavage.

[0286] Cas Endonuclease and Effects An endonuclease is an enzyme that cleaves phosphodiester bonds within a polynucleotide chain and includes restriction endonucleases that cleave DNA at specific sites without damaging the bases. Examples of endonucleases include restriction endonucleases, meganucleases, TAL effector nucleases (TALENs), zinc finger nucleases, and Cas (CRISPR-associated) effector endonucleases.

[0287] A Cas endonuclease, either as a single effector protein or forming an effector complex with other components, unwinds a DNA double strand at a target sequence and optionally cleaves at least one DNA strand when mediated by recognition of the target sequence by a polynucleotide (e.g., but not limited to, crRNA or guide RNA) that forms a complex with the Cas effector protein. Such recognition and cleavage of the target sequence by a Cas endonuclease typically occurs when an exact protospacer adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. Alternatively, a Cas endonuclease herein may lack DNA cleavage activity or nicking activity but can still specifically bind to a DNA target sequence when complexed with a suitable RNA component (see also U.S. Patent Application Publication No. 20150082478, published March 19, 2015, and U.S. Patent Application Publication No. 20150059010, published February 26, 2015).

[0288] A Cas endonuclease can occur as part of an individual effector (class 2 CRISPR system) or a larger effector complex (class I CRISPR system).

[0289] Examples of the described Cas endonucleases include, for example, Cas9, Cas12f (Cas-alpha, Cas14), Cas12l (Cas-beta), Cas12a (Cpf1), Cas12b (C2c1 protein), Cas13 (C2c2 protein), Cas12c (C2c3 protein), Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas3, Cas3-HD, Cas5, Cas6, Cas7, Cas8, Cas10, or combinations or complexes thereof. The Cas endonucleases and effector proteins can be used for targeted genome editing (by single and multiple double-strand breaks and nicks) and targeted genome regulation (by attaching epigenetic effector domains to Cas polypeptides or sgRNAs). The Cas endonucleases can also be engineered to function as RNA-guided recombinases and can function as scaffolds for the assembly of multi-protein and nucleic acid complexes by RNA tethers (Mali et al., 2013, Nature Methods Vol.10:957-963).

[0290] Cas-alpha endonuclease Cas-alpha endonuclease (e.g., also known as Cas-alpha 10, Cas12f) is defined as a functional RNA-guided PAM-dependent dsDNA cleavage protein consisting of less than 800 amino acids, divided into three subdomains and further comprising a C-terminal RuvC catalytic domain containing a bridging helix and one or more zinc finger motifs; and an N-terminal Rec subunit having a helical bundle, a WED wedge (or "oligonucleotide binding domain", OBD) domain, and optionally a zinc finger motif. Some exemplary Cas-alpha endonucleases are described, for example, in U.S. Patent No. 10,934,536 and International Publication No. 2022 / 082179 pamphlet.

[0291] The RuvC domain has been shown in the literature to encompass the functionality of endonucleases. Cas-alpha endonucleases can be isolated or identified from a locus comprising a cas-alpha gene encoding an effector protein and an array containing multiple repeats. In some embodiments, the cas-alpha locus may further comprise part or all of the cas1 gene, the cas2 gene, and / or the cas4 gene.

[0292] A zinc finger motif is a domain that stabilizes a zinc fold by coordinating one or more zinc ions, typically with cysteine ​​and histidine side chains. Zinc fingers are named according to the pattern of cysteine ​​and histidine residues that coordinate to the zinc ion (for example, C4 means that four cysteine ​​residues coordinate to the zinc ion, and C3H means that three cysteine ​​residues and one histidine residue coordinate to the zinc ion).

[0293] Cas-alpha polypeptides contain one or more zinc finger (ZFN) coordination motifs that can form a zinc-binding domain. These zinc finger-like motifs can assist in the separation of target and non-target strands, as well as the loading of guide RNA to the DNA target. Cas-alpha polypeptides containing one or more zinc finger motifs may provide further stability to ribonucleoprotein complexes on target polynucleotides. Cas-alpha polypeptides contain a C4 or C3H zinc-binding domain.

[0294] Cas-alpha-endonuclease is an RNA-induced endonuclease that can bind to and cleave a double-stranded DNA target containing (1) a sequence homologous to the nucleotide sequence of a guide RNA and (2) a PAM sequence. In some embodiments, the PAM is T-rich. In some embodiments, the PAM is C-rich.

[0295] Cas-alpha endonucleases function as double-strand break inducers and may also be nickases or single-strand break inducers. In some embodiments, catalytically inactive Cas-alpha endonucleases may be used to target or guide a target DNA sequence, but without inducing cleavage. In some embodiments, catalytically inactive Cas-alpha polypeptides may be used in conjunction with functional endonucleases that cleave the target sequence. In some embodiments, catalytically inactive Cas-alpha polypeptides may be combined with base-editing molecules such as deaminases. A "deaminase" is an enzyme that catalyzes deamination reactions. For example, deamination of adenine by adenine deaminase results in the formation of inosine. Inosine selectively base-pairs with cytosine instead of thymine. This results in post-replication transfer mutations, which convert the original AT base pair to a GC base pair. In another example, cytosine deamination leads to uracil formation, which can be repaired by cellular repair mechanisms and revert to CT base pairs or TA, GC, or AT base pairs. This heterogeneity in repair can be suppressed by introducing uracil glycosylase inhibitors so that DNA repair or replication converts the original CT base pairs to TA base pairs (Burnett et al. (2022) Frontiers in Genome Editing. 4, 923718). In the case of both adenine and cytosine deaminases, the introduction of nicks promotes the change in their respective base pairs (Burnett et al., 2022). In some embodiments, the deaminase may be cytidine deaminase. In some embodiments, the deaminase may be adenine deaminase. In some embodiments, the deaminase may be ADAR-2.

[0296] The "functional fragments" of Cas-alpha-endonuclease possess the ability to recognize, bind to, or cleave a single strand of a double-stranded polynucleotide, or the ability to cleave both strands of a double-stranded polynucleotide, or any combination thereof.

[0297] Cas polypeptides, effector proteins, or functional fragments thereof for use in the methods of this disclosure may be isolated from natural sources or from recombinant sources in which genetically modified host cells have been modified to express the nucleic acid sequence encoding this protein. Alternatively, Cas polypeptides may be produced using cell-free protein expression systems or synthetically. Effector Cas nucleases may be isolated and introduced into heterologous cells or modified from their natural form to exhibit a different type or size of activity than that of their natural source. Such modifications include, but are not limited to, fragments, variants, substitutions, deletions, and insertions.

[0298] Cas polypeptides and Cas effector protein fragments and variants can be obtained by methods such as site-directed mutagenesis and synthetic construction. Methods for measuring endonuclease activity are well known in the art and include, but are not limited to, International Publication No. 2013166113, published November 7, 2013, International Publication No. 2016186953, published November 24, 2016, and International Publication No. 2016186946, published November 24, 2016.

[0299] The Cas polypeptides disclosed herein can be modified. Modifications of Cas polypeptides include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the naturally occurring nuclease activity of the Cas polypeptide. For example, in some cases, modified Cas polypeptides have less than 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nuclease activity of the corresponding wild-type Cas polypeptide (U.S. Patent Application Publication No. 20140068797, published March 6, 2014). In some cases, modified Cas polypeptides have substantially no nuclease activity and are referred to as catalytically inactivated Cas or inactivated Cas (dCas). Inactivated Cas / inactivated Cas includes inactivated Cas endonuclease (dCas). Catalytically inactive Cas effector proteins can be fused to heterologous sequences to induce or modify their activity.

[0300] The Cas polypeptide disclosed herein may be part of a fusion protein comprising one or more heterologous protein domains (e.g., one, two, three or more domains in addition to the Cas polypeptide). Such a fusion protein may comprise any further protein sequence and, optionally, any linker sequence between any two domains (e.g., between Cas and a first heterologous domain). Examples of protein domains that can be fused to the Cas polypeptides herein include, but are not limited to, epitope tags (e.g., histidine [His], V5, FLAG, influenza hemagglutinin [HA], myc, VSV-G, thioredoxin [Trx]), reporters (e.g., glutathione-5-transferase [GST], horseradish peroxidase [HRP], chloramphenicol acetyltransferase [CAT], beta-galactosidase, beta-glucuronidase [GUS], luciferase, green fluorescent protein [GFP], HcRed, DsRed, cyan fluorescent protein [CFP], yellow fluorescent protein [YFP], blue fluorescent protein [BFP]), and domains having one or more of the following activities: methyltransferase activity, demethyltransferase activity, transcriptional activating activity (e.g., VP16 or VP64), transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. The disclosed Cas polypeptide can also fuse with DNA molecules or other molecules, such as maltose-binding proteins (MBPs), S-tags, Lex A DNA-binding domains (DBDs), GAL4A DNA-binding domains, and proteins that bind to herpes simplex virus (HSV) VP16.

[0301] Catalytically active and / or inactive Cas polypeptides can be fused to heterologous sequences (U.S. Patent Application Publication No. 20140068797, published March 6, 2014). Suitable fusion partners include, but are not limited to, polypeptides that provide activity to indirectly increase transcription by directly acting on target DNA or polypeptides associated with target DNA (e.g., histones or other DNA-binding proteins). Further suitable fusion partners include, but are not limited to, polypeptides that provide methyltransferase activity, demethylation activity, acetyltransferase activity, deacetylation activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity. Other suitable fusion partners include, but are not limited to, polypeptides that directly induce increased transcription of the target nucleic acid (e.g., transcription activators or their fragments, proteins or their fragments that induce transcription activators, small molecule / drug-responsive transcription regulators, etc.). Partially active or catalytically inactive Cas-alpha endonucleases can also fuse to other proteins or domains (e.g., Clo51 nuclease or FokI nuclease) to generate double-strand breaks (Guilinger et al. Nature biotechnology, volume 32, number 6, June 2014).

[0302] Catalytically active or inactive Cas polypeptides, such as the Cas-alpha polypeptide disclosed herein, may also be fused with molecules that direct the editing of one or more bases in a polynucleotide sequence, such as site-directed deaminases that can change nucleotide identity, for example, from C·G to T·A or A·T to G·C (Gaudelli et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al., "Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems." Science 353(6305)(2016); Komor et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage." Nature) 533(7603)(2016):420-4). Base editing fusion proteins may include, for example, active (double-strand breaking), partially active (nickase), or inactivated (catalytically inactive) Cas-alpha endonucleases and deaminases (e.g., cytidine deaminase, adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABE, etc., but not limited to these). Base editing repair inhibitors and glycosylase inhibitors (e.g., uracil glycosylase inhibitors (which prevent the removal of uracil)) are intended in some embodiments as other components of the base editing system.

[0303] Any of the Cas polypeptides disclosed herein (e.g., catalytically inactive Cas polypeptides) can also be complexed with base-editing molecules via RNA aptamer systems, such as those described in International Publication No. 2021 / 055459, published on March 25, 2021. RNA aptamer systems may include, for example, (A) RNA motifs, e.g., (1) telomerase Ku-binding motifs, (2) telomerase SM7-binding motifs, (3) MS2 phage operator stem-loops, (6) PP7 phage operator stem-loops, (5) SfMu phage COM stem-loops, (4) chemically modified versions of such aptamers, or (7) non-natural RNA aptamers and (B) their corresponding aptamer ligands or RNA-binding sections. See also U.S. Patent No. 11,479,793, published on July 12, 2018, and International Publication No. 2018 / 129129.

[0304] The Cas polypeptides and endonucleases described herein can be expressed and purified by methods known in the art, for example, by the method described in International Publication No. 2016 / 186953, published on November 24, 2016.

[0305] To date, many Cas polypeptides and endonucleases capable of recognizing specific PAM sequences (International Publication No. 2016186953, published November 24, 2016; International Publication No. 2016186946, published November 24, 2016; and Zetsche B et al. 2015. Cell 163, 1013) and cleaving target DNA at specific locations have been described. Based on the methods and embodiments described herein using novel inducible Cas systems, those skilled in the art will understand that these methods can be adapted so that they can be used with any inducible endonuclease system.

[0306] The Cas effector protein may contain a heterologous nuclear localization sequence (NLS). The heterologous NLS amino acid sequences described herein may have sufficient strength to drive, for example, the accumulation of a detectable amount of Cas polypeptide in the nucleus of a yeast cell as described herein. The NLS may contain a short sequence (e.g., 2 to 20 residues) of one (single-segmented) or more (e.g., bi-segmented) basic, positively charged residues (e.g., lysine and / or arginine), and may be positioned at any site in the Cas amino acid sequence, but exposed to the surface of the protein. The NLS may be operably ligated to the N-terminus or C-terminus of the Cas polypeptide as described herein. For example, two or more NLS sequences may be ligated to the Cas polypeptide, for example, to both the N-terminus and C-terminus of the Cas polypeptide. The Cas polypeptide gene can be operably ligated upstream of the SV40 nuclear targeting signal in the Cas codon region and downstream of the bifid VirD2 nuclear stereotactic signal in the Cas codon region (Tinland et al. (1992) Proc. Natl. Acad. Sci. USA 89:7442-6). Non-limiting examples of preferred NLS sequences as used herein are disclosed in U.S. Patent Nos. 6,660,830 and 7,309,576.

[0307] Guide polynucleotides The guide polynucleotide enables target recognition, binding to the target, and selective cleavage of the target by the Cas polypeptide, and may be a single or double molecule. The guide polynucleotide sequence may be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence). Optionally, the guide polynucleotide may include at least one nucleotide, a phosphodiester bond, or a binding modification, including, but not limited to, locked nucleic acid (LNA), 5-methyl dC, 2,6-diaminopurine, 2'-fluoro A, 2'-fluoro U, 2'-O-methyl RNA, a phosphorothioate bond, linking to a cholesterol molecule, linking to a polyethylene glycol molecule, linking to a spacer 18 (hexaethylene glycol chain) molecule, or a 5'-to-3' covalent link resulting in cyclization. Guide polynucleotides containing only ribonucleic acid are also called “guide RNA” or “gRNA” (U.S. Patent Application Publication No. 20150082478, published March 19, 2015, and U.S. Patent Application Publication No. 20150059010, published February 26, 2015). Guide polynucleotides may be manipulated or synthesized.

[0308] Guide polynucleotides include chimeric non-natural guide RNAs that contain regions not found together in nature (i.e., these are heterogeneous). For example, in a chimeric non-natural guide RNA containing a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in the target DNA, ligated to a second nucleotide sequence capable of recognizing a Cas polypeptide, the first and second nucleotide sequences are not found ligated together in nature.

[0309] Guide polynucleotides can be double molecules (also called double-stranded guide polynucleotides) containing a cr nucleotide sequence (e.g., crRNA) and a tracr nucleotide sequence (e.g., tracrRNA). In some examples, there are linker polynucleotides that ligate crRNA and tracrRNA to form a single guide (e.g., sgRNA).

[0310] In some embodiments, a cr nucleotide comprises a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in target DNA, and a second nucleotide sequence (also called a tracr mate sequence) that is part of a Cas polypeptide recognition (CER) domain. The tracr mate sequence can hybridize to the tracr nucleotide along a complementary region, together forming the Cas polypeptide recognition domain, i.e., the CER domain. The CER domain can interact with the Cas polypeptide. The cr nucleotides and tracr nucleotides of a double-stranded guide polynucleotide can be RNA, DNA, and / or RNA-DNA combination sequences. In some embodiments, the cr nucleotide molecule of a double-stranded guide polynucleotide is called "crDNA" (when composed of a continuous sequence of DNA nucleotides), "crRNA" (when composed of a continuous sequence of RNA nucleotides), or "crDNA-RNA" (when composed of a combination of DNA nucleotides and RNA nucleotides). The cr nucleotide may include fragments of crRNA naturally occurring in bacteria and archaea. The size of the fragments of naturally occurring crRNA in bacteria and archaea that may be present in the cr nucleotides disclosed herein may vary from, but are not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides.

[0311] In some embodiments, the tracr nucleotide is referred to as "tracrRNA" (when consisting of a continuous stretch of RNA nucleotides), "tracrDNA" (when consisting of a continuous stretch of DNA nucleotides), and "tracrDNA-RNA" (when consisting of a combination of DNA and RNA nucleotides). In one embodiment, the RNA that guides the RNA Cas9 endonuclease complex is a double-stranded RNA comprising a double-stranded crRNA-tracrRNA. The tracrRNA (trans-activating CRISPR RNA) includes, in the 5' to 3' direction, (i) a sequence that anneals to the repeat region of a CRISPR type II crRNA, and (ii) a stem-loop-containing portion (Deltcheva et al., Nature 471:602-607). The double-stranded guide polynucleotide can form a complex with a Cas polypeptide, and the guide polynucleotide / Cas polypeptide complex (also referred to as a guide polynucleotide / Cas polypeptide system) can direct the Cas polypeptide to a genomic target site, allowing the Cas polypeptide to recognize, bind to, and optionally nick or cleave (introduce a single-stranded or double-stranded break) the target site. (U.S. Patent Application Publication No. 20150082478, published on March 19, 20 as well as U.S. Patent Application Publication No. 20150059010, published on February 26, 2015).

[0312] In some embodiments, the guide polynucleotide is a guide polynucleotide that can form a PGEN as described herein, and the guide polynucleotide includes a first nucleotide sequence domain complementary to the nucleotide sequence of the target DNA and a second nucleotide sequence domain that interacts with the Cas polypeptide.

[0313] In some embodiments, the guide polynucleotide is a guide polynucleotide as described herein, and each of the first nucleotide sequence and the second nucleotide sequence domain is selected from the group consisting of a DNA sequence, an RNA sequence, and combinations thereof.

[0314] In some embodiments, the guide polynucleotide is a guide polynucleotide as described herein and further comprises a second nucleotide sequence domain selected from the group consisting of stability-enhancing RNA backbone modifications, stability-enhancing DNA backbone modifications, and combinations thereof (see Kanasty et al., 2013, Common RNA-backbone modifications, Nature Materials 12:976-977; U.S. Patent Application Publication No. 20150082478, published March 19, 2015, and U.S. Patent Application Publication No. 20150059010, published February 26, 2015).

[0315] The guide RNA may contain a bimolecule comprising a chimeric non-natural crRNA ligated to at least one tracrRNA. The chimeric non-natural crRNA may include a crRNA containing regions that are not found together in nature (i.e., they are heterogeneous). For example, a crRNA containing a first nucleotide sequence domain (called a variable targeting domain or VT domain) ligated to a second nucleotide sequence (also called a tracrmate sequence) that can hybridize to a nucleotide sequence in the target DNA (therefore, the first and second sequences are never found ligated together in nature).

[0316] A guide polynucleotide can also be a single molecule (also called a single guide polynucleotide) containing a cr nucleotide sequence linked to a tracr nucleotide sequence. A single guide polynucleotide contains a first nucleotide sequence domain (called a variable target domain or VT domain) that can hybridize to a nucleotide sequence in target DNA, and a Cas endonuclease recognition domain (CER domain) that interacts with the Cas polypeptide.

[0317] The VT domain and / or CER domain of a single guide polynucleotide may contain an RNA sequence, a DNA sequence, or an RNA-DNA combination sequence. A single guide polynucleotide composed of sequences derived from cr nucleotides and tracr nucleotides may be referred to as a "single guide RNA" (when composed of a continuous sequence of RNA nucleotides), a "single guide DNA" (when composed of a continuous sequence of DNA nucleotides), or a "single guide RNA-DNA" (when composed of a combination of RNA nucleotides and DNA nucleotides). A single guide polynucleotide may form a complex with a Cas polypeptide, and the guide polynucleotide / Cas polypeptide complex (also called a guide polynucleotide / Cas polypeptide system) may guide the Cas polypeptide to a genomic target site, allowing the Cas polypeptide to recognize this target site, bind to it, and selectively nick or cleave (introduce single-strand or double-strand breaks) at this target site. (U.S. Patent Application Publication No. 20150082478, published March 19, 2015, and U.S. Patent Application Publication No. 20150059010, published February 26, 2015).

[0318] Chimeric non-natural single guide RNAs (sgRNAs) include sgRNAs that contain regions not found together in nature (i.e., these are heterogeneous). For example, an sgRNA containing a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize to a nucleotide sequence in the target DNA, ligated to a second nucleotide sequence (also called a tracr-mate sequence) (these are not found ligated together in nature).

[0319] The nucleotide sequence linking the cr nucleotide and tracr nucleotide of a single guide polynucleotide may include an RNA sequence, a DNA sequence, or an RNA-DNA combination sequence. In one embodiment, the nucleotide sequence linking the cr nucleotide and tracr nucleotide of a single guide polynucleotide (also called a "loop") includes at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43 The lengths may be 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides. In another embodiment, the nucleotide sequence linking the cr nucleotide and tracr nucleotide of the single-stranded guide polynucleotide may include, but is not limited to, a tetraloop sequence such as a GAAA tetraloop sequence.

[0320] Guide polynucleotides can be generated by any method known in the art, including, but not limited to, the chemical synthesis of guide polynucleotides (e.g., Hendel et al. 2015, Nature Biotechnology 33, 985-989), the in vitro generation of guide polynucleotides, and / or the self-splicing of guide RNA (e.g., but not limited to, Xie et al., 2015, PNAS 112:3570-3575).

[0321] Guide polynucleotide / Cas polypeptide complex The guide polynucleotide / Cas polypeptide complexes described herein can recognize, bind to, and optionally nick, unnick, or cleave all or part of a target sequence (e.g., a Cas endonuclease or Cas polypeptide having nicking or cleavage activity, or a Cas polypeptide having nuclease or endonuclease activity).

[0322] A guide polynucleotide / Cas polypeptide complex capable of cleaving both strands of a DNA target sequence typically comprises a Cas polypeptide having all of its endonuclease domains in a functional state (e.g., a wild-type endonuclease domain, or a variant thereof that retains some or all of the activity of each endonuclease domain). Therefore, a wild-type Cas polypeptide (e.g., the Cas polypeptides disclosed herein) or a variant thereof that retains some or all of the activity of each endonuclease domain is a preferred example of a Cas endonuclease capable of cleaving both strands of a DNA target sequence.

[0323] A guide polynucleotide / Cas endonuclease complex capable of cleaving a single strand of a DNA target sequence may be characterized herein as having nickase activity (e.g., partial cleavage ability). A Cas nickase typically comprises one functional endonuclease domain that allows Cas to cleave only a single strand of the DNA target sequence (i.e., nick it). For example, a Cas9 nickase may comprise (i) a mutated, dysfunctional RuvC domain and (ii) a functional HNH domain (e.g., a wild-type HNH domain). Another example of a Cas9 nickase may comprise (i) a functional RuvC domain (e.g., a wild-type RuvC domain) and (ii) a mutated, dysfunctional HNH domain. Non-limiting examples of Cas9 nickases suitable for use herein are disclosed in U.S. Patent Application Publication No. 20140189896, published July 3, 2014. A pair of Cas nickases can be used to increase the specificity of DNA targeting. Generally, this can be done by providing two Cas nickases that target a DNA sequence on the reverse strand within a region for desired targeting, by binding to an RNA component with different guide sequences, and making a nick near it. Such a break near each DNA strand causes a double-strand break (i.e., a DSB with a single-strand overhang), which is then recognized as a substrate for non-homologous end-joining NHEJ (prone to incomplete repair leading to mutation) or homologous recombination HR. Each nick in these embodiments may be spaced apart from each other by, for example, at least about 5, 5-10, at least 10, 10-15, at least 15, 15-20, at least 20, 20-30, at least 30, 30-40, at least 40, 40-50, at least 50, 50-60, at least 60, 60-70, at least 70, 70-80, at least 80, 80-90, at least 90, 90-100, or 100 or more (or any integer between 5 and 100) bases. One or two Cas nickase proteins in this specification may be used in a Cas nickase pair.For example, a Cas9 nickase having a mutated RuvC domain but a functional HNH domain (i.e., Cas9 HNH+ / RuvC-) can be used (e.g., Streptococcus pyogenes Cas9 HNH+ / RuvC-). Each Cas9 nickase (e.g., Cas9 HNH+ / RuvC-) can be directed to specific DNA sites that are close to each other (up to 100 base pairs apart) by using a suitable RNA component herein that has a guide RNA sequence that directs each nickase to target its respective specific DNA site.

[0324] In certain embodiments, guide polynucleotide / Cas polypeptide complexes can bind to a DNA target site sequence but do not cleave any strands at the target site sequence. Such complexes may contain a Cas polypeptide in which all of its nuclease domains are mutated and dysfunctional. For example, the Cas9 protein, which can bind to a DNA target site sequence but does not cleave any strands at the target site sequence, may contain both a mutated dysfunctional RuvC domain and a mutated dysfunctional HNH domain. Gene expression can be regulated using the Cas polypeptides herein that bind to a target DNA sequence but do not cleave it, for example, in which case the Cas polypeptide may fuse with a transcription factor (or part thereof) (e.g., a repressor or activator, e.g., either of those disclosed herein).

[0325] In some embodiments, the guide polynucleotide / Cas endonuclease complex (PGEN) described herein is a PGEN in which the Cas endonuclease is optionally covalently or noncovalently linked to or associated with at least one protein subunit or functional fragment thereof.

[0326] In some embodiments, the guide polynucleotide / Cas endonuclease complex is a guide polynucleotide / Cas endonuclease complex (PGEN) comprising at least one guide polynucleotide and at least one Cas endonuclease polypeptide, wherein the Cas endonuclease polypeptide comprises at least one protein subunit or a functional fragment thereof, the guide polynucleotide is a chimeric non-natural guide polynucleotide, and the guide polynucleotide / Cas endonuclease complex can recognize, bind to, and optionally nick, unnick, or cleave all or part of a target sequence.

[0327] The Cas effector protein may be one of the Cas-alpha effector proteins disclosed herein.

[0328] In some embodiments, the guide polynucleotide / Cas effector complex is a guide polynucleotide / Cas effector protein complex (PGEN) comprising at least one guide polynucleotide and a Cas-alpha effector protein, wherein the guide polynucleotide / Cas effector protein complex can recognize, bind to, and optionally nicking, unraveling, or cleaving all or part of a target sequence.

[0329] PGEN can be a guide polynucleotide / Cas effector protein complex, where the Cas effector protein further comprises at least one protein subunit or one or more copies of a functional fragment thereof. In some embodiments, the protein subunit is selected from the group consisting of Cas1 protein subunit, Cas2 protein subunit, Cas4 protein subunit, and any combination thereof. PGEN can be a guide polynucleotide / Cas effector protein complex, where the Cas effector protein further comprises at least two different protein subunits selected from the group consisting of Cas1, Cas2, and Cas4.

[0330] PGEN may be a guide polynucleotide / Cas effector protein complex, where the Cas effector protein further comprises at least three different protein subunits or functional fragments thereof selected from the group consisting of Cas1, Cas2, and optionally one further Cas polypeptide including Cas4.

[0331] In some embodiments, the guide polynucleotide / Cas effector protein complex described herein is a PGEN, wherein the Cas effector protein is covalently or acovalently linked to at least one protein subunit or functional fragment thereof. A PGEN may be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein polypeptide is covalently or acovalently linked to or associated with one or more copies of at least one protein subunit or functional fragment thereof selected from the group consisting of a Cas1 protein subunit, a Cas2 protein subunit, a further Cas polypeptide optionally comprising a Cas4 protein subunit, and any combination thereof. A PGEN may be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein is covalently or acovalently linked to or associated with at least two different protein subunits selected from the group consisting of a further Cas polypeptide optionally comprising Cas1, Cas2, and Cas4. PGEN can be a guide polynucleotide / Cas effector protein complex, where the Cas effector protein is covalently or noncovalently linked to at least three different protein subunits or functional fragments thereof selected from the group consisting of Cas1, Cas2, and optionally Cas4, and any combination thereof.

[0332] Any component of the guide polynucleotide / Cas effector protein complex, the guide polynucleotide / Cas effector protein complex itself, and the polynucleotide modification template and / or donor DNA can be introduced into heterologous cells or organisms by any method known in the art.

[0333] Recombinant constructs for cell transformation Any combination thereof, including the guide polynucleotide, Cas polypeptide, polynucleotide modification template, donor DNA, the guide polynucleotide / Cas polypeptide system disclosed herein, and optionally one or more polynucleotides of interest, can be introduced into cells. Examples of cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, unconventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein.

[0334] The standard recombinant DNA and molecular cloning techniques used herein are known in the art and are described more fully in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory: Cold Spring Harbor, NY (1989). Transformation methods are known to those skilled in the art and are described below.

[0335] The vectors and constructs include circular plasmids and linear polynucleotides, which include the polynucleotide of interest and other optional components (e.g., linkers, adapters, regulatory elements, or analytical elements). In some examples, the recognition site and / or target site may be contained within introns, coding sequences, 5'UTR, 3'UTR, and / or regulatory regions.

[0336] NHEJ and HDR In some embodiments, the Cas polypeptides described herein may comprise one or more guide polynucleotides and optionally donor DNA, and editing a target polynucleotide sequence may be part of a genome editing system further comprising non-homologous end joining (NHEJ) or homologous recombination (HR) after a Cas polypeptide-mediated double-strand break. When a double-strand break is introduced into DNA, the cell's DNA repair mechanisms are activated to repair the break. The most common repair mechanism for joining broken ends is the non-homologous end joining pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). Chromosomal structural integrity is typically maintained by repair, but deletions, insertions, or other rearrangements are also possible (Siebert and Puchta, (2002) Plant Cell 14:1121-31; Pacher et al., (2007) Genetics 175:21-9). Alternatively, double-strand breaks can be repaired by homologous recombination between homologous DNA sequences. If the sequence near a double-strand break is altered, for example, by the activity of exonucleases involved in the maturation of double-strand breaks, the gene conversion pathway can be repaired to the original structure if homologous sequences, such as homologous chromosomes in non-dividing cells or sister chromatids after DNA replication, are available (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenetic DNA sequences can also function as DNA repair templates for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0337] As used herein, “donor DNA” is a DNA construct containing a polynucleotide of interest to be inserted into a genomic target site, where the insertion is mediated by a Cas polypeptide. When a double-strand break is introduced into the target site by an endonuclease, the homologous first and second regions of the donor DNA undergo homologous recombination with their corresponding homologous genomic regions, thereby resulting in an exchange of DNA between the donor and the target genome. Thus, the provided method incorporates the polynucleotide of interest from the donor DNA into a double-strand break at a target site in the plant genome, thereby modifying the original target site and generating a modified genomic target site.

[0338] Base editing In some embodiments, the Cas polypeptide described herein may be part of a genome editing system further comprising a base editing agent and a plurality of guide polynucleotides, wherein editing a target polynucleotide sequence involves introducing a plurality of nucleic acid base edits into the target polynucleotide sequence resulting in a variant nucleotide sequence.

[0339] In some cases, one or more nucleic acid bases of a target polynucleotide can be chemically altered to change a base from one type to another, for example, from cytosine to thymine, or from adenine to guanine. In some embodiments, plants having multiple altered bases can be produced by modifying or changing multiple bases, for example, two or more, five or more, ten or more, twenty or more, 30 or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, or even several thousand bases.

[0340] Using any base editing complex, such as a base editing agent associated with an RNA guide protein, it is possible to target and bind to a desired gene locus in the genome of an organism and chemically modify one or more components of the target polynucleotide.

[0341] Site-directed base editing can be achieved by manipulating one or more nucleotide changes to create one or more edits in the genome. These include, for example, site-directed base editing mediated by base-editing deaminase enzymes, such as C·G to T·A or A·T to G·C (Gaudelli et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al., "Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems." Science 353(6305)(2016); Komor et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage." Nature) 533(7603)(2016):420-4). Catalytically “dead” or inactive Cas9 (dCas9) fused to cytidine deaminase or adenine deaminase proteins, such as the catalytically inactive “dead” versions of Cas endonuclease disclosed herein, become specific base-editing agents capable of altering DNA bases without inducing DNA cleavage. A base editor that converts C->T (or G->A on the opposite strand) or adenine base editor that converts adenine to inosine results in an A->G change within the editing window identified by gRNA. Any molecule that results in a change of nucleic acid bases is a “base-editing agent”.

[0342] For many target traits, the generation of a single double-strand break and subsequent repair by HDR or NHEJ is not ideal for quantitative traits. The observed phenotype includes both genotype and environmental effects. Genotype effects further include additive, dominant, and superiority effects. Depending on the method used and the selected site, the probability of no effect per any single edit can be greater than zero, and the phenotypic effect of any single edit can be small. Double-strand break repair can also be noisy and have low reproducibility.

[0343] One approach to improve the probability of no effect or small phenotypic effect per edit is to perform multiple genome modifications so that multiple target sites are altered. Single base substitutions can be made possible by methods of altering the genome sequence that do not introduce double-strand breaks. Combining these approaches, multiple base editing is useful for purifying numerous genotype edits that can produce observable phenotypic alterations. In some cases, tens, hundreds, or even thousands of sites can be edited within one or a few generations of an organism.

[0344] Multiplexing approaches to base editing in organisms have the potential to generate multiple significant phenotypic variations in one or several generations, with a positive directional bias in the effects. In some embodiments, the organism is a plant. A plant or population of plants with multiple edits can be crossbred to produce offspring plants, some of which include multiple edits from the parent plant. In this way, accelerated breeding of desired traits can be achieved in parallel in one or several generations, replacing conventional time-consuming sequential crossbreeding and breeding over multiple generations.

[0345] Base-editing deaminases, such as cytidine deaminase or adenine deaminase, can be fused to RNA-induced endonucleases that are either inactivated (e.g., "DCAS" such as inactivated Cas9) or partially active (e.g., "nCas" such as Cas9 nickase) so as not to cleave the target site where they are induced. The dCas forms a functional complex with a guide polynucleotide that shares homology to the polynucleotide sequence of the target site, and further complexes with the deaminase molecule. The induced Cas polypeptide recognizes and binds to the double-stranded target sequence, opening the double strand and exposing the individual bases. In the case of cytidine deaminase, the deaminase deaminates the cytosine base to produce uracil. Uracil glycosylase inhibitors (UGIs) are provided to prevent the conversion of U back to C. The DNA replication or repair mechanism then converts uracil to thymine (U to T), and then repairs the opposite base (the previous G in the original GC pair) to adenine to produce a TA pair. For example, see Komor et al. Nature Volume 533, Pages 420-424, May 19, 2016. Base editing performance can be enhanced using RNA aptamer systems, such as those discussed in International Publication No. 2021 / 055459, published March 25, 2021; U.S. Patent No. 11,479,793; and International Publication No. 2018 / 129129, published July 12, 2018.

[0346] Prime Edit In some embodiments, the Cas polypeptides described herein may be part of a genome editing system further comprising a prime editing agent and a guide polynucleotide, wherein editing a target nucleotide sequence involves introducing one or more insertions, deletions, or nucleotide base exchanges into the target nucleotide sequence without generating double-strand DNA breaks. See, for example, Anzalone, AV, Randolph, PB, Davis, J.Ret al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019).

[0347] In some embodiments, the prime editor is a Cas polypeptide fused to a reverse transcriptase, where the Cas polypeptide is modified to nick the DNA rather than produce double-strand breaks. This Cas-polypeptide-reverse transcriptase fusion may also be called a "prime editor" or "PE". In some embodiments, the guide polynucleotide comprises a prime editing guide polynucleotide (pegRNA) that is larger than a standard sgRNA commonly used in CRISPR gene editing (e.g., more than 100 nucleic acid bases). The pegRNA comprises a primer-binding sequence (PBS) and a template containing the desired or target RNA sequence at its 3' end.

[0348] During prime editing, the PE:pegRNA complex binds to the target DNA sequence, and the modified Cas polypeptide nicks one target DNA strand to create a flap. PBS on the pegRNA binds to the DNA flap, and the target RNA sequence is reverse transcribed using reverse transcriptase. The edited strand is incorporated into the target DNA at the nicked end, and the target DNA sequence is repaired with the newly reverse-transcribed DNA.

[0349] Other reverse transcriptase-based genome modification systems Another option for genome modification using the CRISPR-Cas system appears to rely on reverse transcriptase-based methods that reverse transcribe the desired genome edit onto the complement of the PAM-containing target strand DNA (i.e., the target strand) compared to the non-target strand DNA described in prime editing. One aspect of this RNA-coding DNA substitution of an allele by CRISPR (called REDRAW) is disclosed, for example, in Kim et al., bioRxiv 2022.12.13.520319 (2022) and U.S. Patent Application Publication No. 20210130835A1, which are incorporated herein by reference to the extent necessary for use with the CRISPR-Cas polypeptides disclosed herein.

[0350] Components for the expression and utilization of novel CRISPR-Cas systems in prokaryotic and eukaryotic cells This disclosure further provides expression constructs for expressing a guide RNA / Cas system in prokaryotic cells / organisms or eukaryotic cells / organisms that recognizes all or part of a target sequence, binds to it, and optionally nicks, unnicks, or cleaves it.

[0351] In some embodiments, the expression construct of the Disclosure comprises a promoter operably ligated to a nucleotide sequence encoding a Cas gene (or an optimized plant comprising the Cas polypeptide gene described herein), and a promoter operably ligated to the guide RNA of the Disclosure. The promoters can drive the expression of the operably ligated nucleotide sequence in prokaryotic cells / organisms or eukaryotic cells / organisms.

[0352] Nucleotide sequence modifications of the guide polynucleotide, VT domain, and / or CER domain may be selected from, but are not limited to, the group consisting of a 5' cap, a 3' polyadenylated tail, a riboswitch sequence, a stable regulatory sequence, a sequence forming a dsRNA double helix, a modification or sequence that targets the guide polynucleotide to an intracellular location, a modification or sequence that provides tracking, a modification or sequence that provides a binding site for a protein, locked nucleic acid (LNA), 5-methyldC nucleotide, 2,6-diaminopurine nucleotide, 2'-fluoroA nucleotide, 2'-fluoroU nucleotide; 2'-O-methylRNA nucleotide, phosphorothioate linkage, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 molecule, a 5' to 3' covalent linkage, or any combination thereof. These modifications may result in at least one additional advantageous feature, which may be selected from the following: modified or regulated stability, intracellular targeting, tracking, fluorescent labeling, binding sites for proteins or protein complexes, modified binding affinity to complementary target sequences, modified resistance to cytodegradation, and increased cellular permeability.

[0353] A method for expressing RNA components such as gRNA in eukaryotic cells for Cas9-mediated DNA targeting involves using the RNA polymerase III (Pol III) promoter, which enables transcription of RNA with precisely defined, unmodified 5' and 3' ends (DiCarlo et al., Nucleic Acids Res. 41:4336-4343; Ma et al., Mol.Ther.Nucleic Acids 3:e161). This strategy has been successfully applied to cells of several different species, including maize and soybeans (US Patent Application Publication No. 20150082478, published March 19, 2015). A method for expressing RNA components without a 5' cap is described (International Publication No. 2016 / 025131, published February 18, 2016).

[0354] Various methods and compositions can be used to obtain cells or organisms having a target polynucleotide to be inserted into a target site for a Cas polypeptide. Such methods may involve homologous recombination (HR) to incorporate the target polynucleotide into the target site. In one method described herein, the target polynucleotide is introduced into an organismal cell by a donor DNA construct.

[0355] The donor DNA construct further includes first and second homologous regions adjacent to the target polynucleotide. The first and second homologous regions of the donor DNA are homologous to first and second genomic regions located in or adjacent to the target site in the genome of a cell or organism, respectively.

[0356] Donor DNA can be bound to guide polynucleotides. Tethered donor DNA can enable co-localization of target and donor DNA, which is useful in genome editing, gene insertion, and target genome regulation, and may also be useful in targeting mitotic terminal cells in which the function of the endogenous HR mechanism is expected to be significantly impaired (Mali et al., 2013, Nature Methods Vol.10:957-963).

[0357] The amount of homology or sequence identity shared by target and donor polynucleotides can vary, ranging from approximately 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, and 450-90 bp. The ranges include 0 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5-3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5-7 kb, 4-8 kb, 5-10 kb, or the total length and / or region having unit integer values ​​within the ranges less than or equal to the total length of the target site. These ranges include all integers within this range; for example, the range 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 bp. The amount of homology can also be explained by the percentage of sequence identity over the entire fully aligned length of the two polynucleotides, which includes at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%-99%, 99%, 99%-100%, or 100% of sequence identity. Sufficient homology includes any combination of polynucleotide length, overall sequence identity percentage, and any optional conserved region of consecutive nucleotides or local sequence identity percentage. For example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity with respect to the region of the target locus.Sufficient homology can also be explained by the predictive ability of two polynucleotides to specifically hybridize under high stringency conditions; see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology--Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0358] The structural similarity between a given genomic region and the corresponding homologous region found on donor DNA can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of homology or sequence identity shared by the “homologous region” of donor DNA and the “genomic region” of the biological genome may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, such that those sequences undergo homologous recombination.

[0359] Homologous regions on donor DNA may have homology to any sequence adjacent to the target site. In some cases, homologous regions share considerable sequence homology to genomic sequences directly adjacent to the target site, but it is recognized that these homologous regions can be designed to have sufficient homology to a further 5' or 3' region relative to the target site. Homologous regions may also have homology to fragments of the target site in addition to downstream genomic regions.

[0360] In one embodiment, the first homologous region further includes a first fragment of the target site, and the second homologous region includes a second fragment of the target site, wherein the first and second fragments are distinct.

[0361] Target polynucleotide This specification further describes the polynucleotides of interest, including those that reflect the interests of commercial markets and those related to crop development. The crops and markets of interest will change, and as developing countries expand their global markets, newer crops and technologies will also emerge. Furthermore, as our understanding of agricultural traits and characteristics such as yield and hybrid vigor will improve, the selection of genes for genetic modification will change accordingly.

[0362] General categories of target polynucleotides include, for example, target genes involved in information transfer, such as zinc fingers; target genes involved in signal transduction, such as kinases; and target genes involved in housekeeping, such as heat shock proteins. More specific target polynucleotides include, but are not limited to, the following: genes involved in traits of agricultural benefit, including, but not limited to, those affecting crop yield, grain quality, crop nutrient content, starch and carbohydrate quality and quantity, as well as grain size, sucrose load, protein quality and quantity, nitrogen fixation and / or utilization, fatty acids and oil composition; genes encoding proteins that confer resistance to abiotic stress (e.g., drought, nitrogen, temperature, salinity, toxic metals or trace elements, or toxins such as insecticides and herbicides); and genes encoding proteins that confer resistance to biological stress (e.g., attacks by fungi, viruses, bacteria, insects, and nematodes, and the progression of diseases associated with these organisms).

[0363] Agriculturally important traits, such as oil, starch, and protein content, can be genetically modified in addition to using conventional breeding methods. Further modifications include increasing the content of oleic acid, saturated and unsaturated oils, increasing lysine and sulfur concentrations, providing essential amino acids, and modifying starch. Modifications of horidothionine protein are described in U.S. Patents No. 5,703,049, No. 5,885,801, No. 5,885,802, and No. 5,990,389.

[0364] The target polynucleotide sequence may encode proteins involved in the introduction of disease resistance or pest resistance. "Disease resistance" or "pest resistance" is intended to enable plants to avoid harmful symptoms resulting from interactions between plants and pathogens. Pest resistance genes may encode resistance to high-yield-inhibiting pests such as cutworms, armyworms, and European corn borers. Examples of useful gene products include disease and insect resistance genes such as lysozyme or cecropin to protect antimicrobial properties, proteins such as defensins, glucanases, or chitinases to protect antifungal properties, or Bacillus thuringiensis endotoxin, protease inhibitors, collagenases, lectins, or glucosidases to control nematodes or insects. Genes encoding disease resistance traits include detoxification genes for fumonisin (U.S. Patent No. 5,792,931), non-pathogenicity (avr) genes, and disease resistance (R) genes (Jones et al. (1994) Science 266:789; Martin et al. (1993) Science 262:1432; and Mindrinos et al. (1994) Cell 78:1089), and similar genes. Insect resistance genes may encode resistance to pests that cause yield drag, such as cutworms, armyworms, European corn borers, and similar species. Examples of such genes include the Bacillus thuringiensis toxic protein gene (US Patent Nos. 5,366,892; 5,747,450; 5,736,514; 5,723,756; 5,593,881; and Geiser et al. (1986) Gene 48:109), and similar genes.

[0365] Proteins obtained by the expression of "herbicide-resistant proteins" or "nucleic acid molecules encoding herbicide resistance" include proteins that confer upon cells the ability to withstand higher concentrations of herbicides than cells that do not express this protein, or the ability to withstand a specific concentration of herbicide for a longer period than cells that do not express this protein. Herbicide resistance traits can be introduced into plants by genes encoding resistance to herbicides that act by inhibiting the activity of acetolactate synthase (ALS, also known as acetohydroxy acid synthase, AHAS), particularly sulfonylurea (UK: sulfonylurea (sulphonylurea)) type herbicides, genes that act by inhibiting the activity of glutamine synthase, for example, genes encoding resistance to phosphinothricin or basta (e.g., the bar gene), genes encoding resistance to glyphosate (e.g., the EPSP synthase gene and the GAT gene), genes encoding resistance to HPPD inhibitors (e.g., the HPPD gene), or other such genes known in the art. For example, see U.S. Patent Nos. 7,626,077, 5,310,667, 5,866,775, 6,225,114, 6,248,876, 7,169,970, 6,867,293, and 9,187,762. The bar gene encodes resistance to the herbicide basta, the nptII gene encodes resistance to the antibiotics kanamycin and genethecin, and the ALS gene variant encodes resistance to the herbicide chlorosulfuron.

[0366] Furthermore, the polynucleotide of interest may also include an antisense sequence that is complementary to at least a portion of the messenger RNA (mRNA) for the target gene sequence of interest. The antisense nucleotide is constructed to hybridize with the corresponding mRNA. Modifications of the antisense sequence can be made insofar as this sequence hybridizes with the corresponding mRNA and interferes with the expression of the corresponding mRNA. In this way, antisense constructs having 70%, 80%, or 85% sequence identity to the corresponding antisense sequence can be used. In addition, portions of the antisense nucleotide may be used to disrupt the expression of the target gene. Generally, sequences of at least 50 nucleotides, 100 nucleotides, 200 nucleotides, or more can be used.

[0367] In addition, the target polynucleotide can also be used in sense orientation to repress the expression of endogenous genes in plants. Methods for repressing gene expression in plants using sense-oriented polynucleotides are known in the art. These methods generally involve transforming the plant with a DNA construct that includes a promoter that drives expression in the plant, which is operably ligated to at least a portion of the nucleotide sequence corresponding to the transcription of the endogenous gene. Typically, such nucleotide sequences have considerable sequence identity to the transcription sequence of the endogenous gene, generally more than about 65%, more than about 85%, or more than about 95%. See U.S. Patents No. 5,283,184 and No. 5,034,323.

[0368] The polynucleotide of interest may also be a phenotypic marker. A phenotypic marker is a screening or selectable marker, including visual markers and selectable markers, whether positive or negative selectable markers. Any phenotypic marker may be used. Specifically, a selectable marker or screening marker often includes a DNA segment that, under certain conditions, allows for the identification of a molecule or cells containing this molecule, or for selection favorably or unfavorably for this molecule or cells. These markers may encode activity such as the production of RNA, peptides, or proteins, or they may provide binding sites for RNA, peptides, proteins, inorganic and organic compounds or compositions, and the like.

[0369] Examples of selection markers include, but are not limited to, the following: DNA segments containing restriction enzyme sites; DNA segments encoding products that provide resistance to otherwise toxic compounds, including antibiotics such as spectinomycin, ampicillin, kanamycin, tetracycline, Basta, neomycin phosphotransferase II (NEO), and hygromycin phosphotransferase (HPT); DNA segments encoding products that are otherwise deleted in recipient cells (e.g., tRNA genes, nutritional requirement markers); DNA segments encoding products that can be easily identified (e.g., phenotypic markers such as β-galactosidase and GUS; fluorescent proteins such as green fluorescent protein (GFP), cyan (CFP), yellow (YFP), red (RFP), and cell surface proteins); the generation of novel primer sites for PCR (e.g., paralleling of two previously unjointed DNA sequences), the inclusion of DNA sequences that have not been or have been acted upon by restriction endonucleases or other DNA-modifying enzymes, chemicals, etc., and the inclusion of DNA sequences necessary for specific modifications (e.g., methylation) that enable their identification.

[0370] Further selection markers include genes conferring resistance to herbicides such as sulfonylurea, glufosinate ammonium, bromoxynyl, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D). See, for example, acetolactate synthase (ALS) for resistance to sulfonylurea, imidazolinone, triazolopyrimidine sulfonamide, pyrimidinyl salicylate, and sulfonylaminocarbonyltriazolinone (Shaner and Singh, 1997, Herbicide Activity: Toxicol Biochem Mol Biol 69-110); and glyphosate-resistant 5-enolpyruvirsicite-3-phosphate (EPSPS) (Saroha et al. 1998, J. Plant Biochemistry & Biotechnology Vol 7: 65-72).

[0371] The polynucleotide of interest includes genes that can be stacked or used in combination with other traits (e.g., herbicide resistance, or any other traits described herein). The polynucleotide of interest and / or traits can be stacked together in composite trait loci as described in U.S. Patent Application Publication No. 20130263324, published October 3, 2013, and International Publication No. 2013 / 112686, published August 1, 2013.

[0372] The target polypeptide includes any protein or polypeptide encoded by the target polynucleotide described herein.

[0373] Furthermore, a method is provided for identifying at least one plant cell species containing a target polynucleotide in its genome, integrated into a target site. Various methods can be used to identify plant cells in which an insertion has been made into or near the target site of the genome. Such methods can be considered as direct analysis of the target sequence to detect changes in the target sequence, and include, but are not limited to, PCR, sequencing, nuclease digestion, Southern blotting, and any combination thereof. See, for example, U.S. Patent Application Publication No. 20090133152, published May 21, 2009. This method also includes recovering the plant from the plant cell containing the target polynucleotide integrated into its genome. The plant may be sterile or fertile. Any target polynucleotide is provided and confirmed to be able to be integrated into a target site in the plant genome and expressed in the plant.

[0374] Sequence optimization for expression in plants Methods for synthesizing plant-optimized genes are available in the art. See, for example, U.S. Patent Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498). Further sequence modifications are known to enhance gene expression in plant hosts. These modifications include, for example, the removal of one or more sequences encoding pseudo-polyadenylation signals, the removal of one or more exon-intron splice site signals, the removal of one or more transposon-like repeats, and the removal of other well-characterized sequences that may be detrimental to gene expression. The GC content of a sequence may be adjusted to the average level of a given plant host (calculated by referring to known genes expressed in this host plant cell). Where possible, the sequence is modified to avoid the predicted hairpin secondary structure of one or more mRNAs. Thus, the “plant-optimized nucleotide sequences” of this disclosure include one or more such sequence modifications.

[0375] Expression element Any polynucleotide encoding a Cas polypeptide or other CRISPR system component disclosed herein may be functionally linked to a heterologous expression element to facilitate transcription or regulation in a host cell. Such expression elements include, but are not limited to, promoters, leaders, introns, and terminators. An expression element may be “minimal,” meaning a short sequence derived from a natural source that still functions as an expression regulator or modifier. Alternatively, an expression element may be “optimized,” meaning its polynucleotide sequence has been modified from its natural state to function with more desirable characteristics in a particular host cell (for example, a bacterial promoter may be “maize-optimized” to improve expression in maize plants). Alternatively, an expression element may be “synthetic,” meaning the expression element is designed in silico and synthesized for use in a host cell. Synthetic expression elements may be entirely synthetic or partially synthetic (containing fragments of naturally occurring polynucleotide sequences).

[0376] Certain promoters have been shown to induce RNA synthesis at a higher rate than others. These are called "strong promoters." If a particular promoter is shown to induce high levels of RNA synthesis only in specific cell or tissue types, and the promoter favorably induces RNA synthesis in certain tissues but at lower levels in others, it is often called a "tissue-specific promoter" or "tissue-dominant promoter."

[0377] Plant promoters include promoters that can initiate transcription in plant cells. For a description of plant promoters, see Potenza et al., 2004, In vitro Cell Dev Biol 40:1-22; Porto et al., 2014, Molecular Biotechnology (2014), 56(1), 38-49.

[0378] Examples of constitutive promoters include the core CaMV 35S promoter (Odell et al., (1985) Nature 313:810-2); rice actin (McElroy et al., (1990) Plant Cell 2:163-71); ubiquitin (Christensen et al., (1989) Plant Mol Biol 12:619-32; ALS promoter (U.S. Patent No. 5,659,026) and similar ones).

[0379] Tissue-preferred promoters can be used to enhance expression within specific plant tissues. As an organization-priority promoter, for example, the international publication No. 2013103367, published on July 11, 2013, includes the following publications: Kawamata et al., (1997) Plant Cell Physiol 38:792-803; Hansen et al., (1997) Mol Gen Genet 254:337-43; Russell et al., (1997) Transgenic Res 6:157-68; Rinehart et al., (1996) Plant Physiol 112:1331-41; Van Camp et al., (1996) Plant Physiol 112:525-35; Canevascini et al., (1996) Plant Physiol 112:513-524; Lam, (1994) Results Probl Cell Differ References include 20:181-96 and Guevara-Garcia et al., (1993) Plant J 4:495-505. Examples of leaf-preferential promoters include Yamamoto et al., (1997) Plant J 12:255-65; Kwon et al., (1994) Plant Physiol 105:357-67; Yamamoto et al., (1994) Plant Cell Physiol 35:773-8; Gotor et al., (1993) Plant J 3:509-18; Orozco et al., (1993) Plant Mol Biol 23:1129-38; Matsuoka et al., (1993) Proc.Natl.Acad.Sci.USA 90:9586-90; Simpson et al., (1958) EMBO J 4:2723-9; and Timko et al., (1988) Nature 318:57-8.Examples of root-preferential promoters include: Hire et al., (1992) Plant Mol Biol 20:207-18 (soybean root-specific glutamine synthase gene); Miao et al., (1991) Plant Cell 3:11-22 (cytosolic glutamine synthase (GS)); Keller and Baumgartner, (1991) Plant Cell 3:1051-61 (root-specific regulatory element in the green bean GRP 1.8 gene); Sanger et al., (1990) Plant Mol Biol 14:433-43 (root-specific promoter of A. tumefaciens mannopine synthase (MAS)); Bogusz et al., (1990) Plant Cell 2:633-41 (Parasponia andersonii Root-specific promoters isolated from *A. andersonii* and *Trema tomentosa*); Leach and Aoyagi, (1991) Plant Sci 79:69-76 (*A. rhizogenes* rolC and rolD root induction genes); Teeri et al., (1989) EMBO J 8:343-50 (*Agrobacterium* wound induction TR1' and TR2' genes); VfENOD-GRP3 gene promoter (Kuster et al., (1995) Plant Mol Biol 29:759-72); and rolB promoter (Capana et al., (1994) Plant Mol Biol 25:681-91; phaseolin gene (Murai et al., (1983) Science 23:476-82; Sengopta-Gopalen et See also U.S. Patent Nos. 5,837,876, 5,750,386, 5,633,363, 5,459,252, 5,401,836, 5,110,732, and 5,023,179.

[0380] Seed-preferred promoters include both seed-specific promoters active during seed development and seed germination promoters active during seed germination. See Thompson et al., (1989) BioEssays 10:108. Examples of seed-preferred promoters include, but are not limited to, Cim1 (cytokinin-inducible message); cZ19B1 (maize 19kDa zein); and milps (myo-inositol-1-phosphate synthase); as well as those disclosed, for example, in International Publication No. 2000011177, published March 2, 2000, and U.S. Patent No. 6,225,529. Examples of seed-preferred promoters for dicotyledonous plants include, but are not limited to, bean-β-phaseolin, napin, β-conglycinin, soybean lectin, clusiferin, and similar products. Examples of suitable seed promoters for monocotyledonous plants include, but are not limited to, maize 15kDa zein, 22kDa zein, 27kDa gamma zein, waxy, shrunken 1, shrunken 2, globulin 1, oleosin, and nuc1. See also International Publication No. 2000012733, published on March 9, 2000, which discloses seed-preferential promoters derived from the END1 and END2 genes.

[0381] Chemically inducible (regulatory) promoters can be used to regulate gene expression in prokaryotic and eukaryotic cells or organisms by the application of exogenous chemical modifiers. These promoters may be chemically inducible promoters, where the application of a chemical induces gene expression, or chemically repressive promoters, where the application of a chemical suppresses gene expression. Examples of chemically inducible promoters include, but are not limited to, the maize In2-2 promoter activated by benzenesulfonamide herbicide toxicity mitigators (De Veylder et al., (1997) Plant Cell Physiol 38:568-77), the maize GST promoter activated by hydrophobic electrophilic compounds used as pre-germination herbicides (GST-II-27, International Publication No. 1993001294, published January 21, 1993), and the tobacco PR-1a promoter activated by salicylic acid (Ono et al., (2004) Biosci Biotechnol Biochem 68:803-7). Other chemomotor promoters include steroid-responsive promoters (e.g., glucocorticoid-inducible promoters (see Schena et al., (1991) Proc. Natl. Acad. Sci. USA 88:10421-5; McNellis et al., (1998) Plant J 14:247-257); and tetracycline-inducible and tetracycline-inhibiting promoters (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Patent Nos. 5,814,618 and 5,789,156)).

[0382] Examples of pathogen-inducible promoters induced after pathogen infection include, but are not limited to, those that regulate the expression of PR proteins, SAR proteins, beta-1,3-glucanase, chitinase, etc.

[0383] An example of a stress-inducible promoter is the RD29A promoter (Kasuga et al. (1999) Nature Biotechnol. 17:287-91). Those skilled in the art are familiar with protocols for simulating stress conditions such as drought, osmotic stress, salt stress, and temperature stress, as well as protocols for evaluating the stress tolerance of plants subjected to simulated or naturally occurring stress conditions.

[0384] Another example of an inducible promoter useful for plant cells is the ZmCAS1 promoter described in U.S. Patent Application Publication No. 20130312137, published on November 21, 2013.

[0385] Various new promoters useful for plant cells are constantly being discovered, and many examples can be found in The Biochemistry of Plants, Vol. 115, Stumpf and Conn, eds (New York, NY: Academic Press), pp. 1-82, edited by Okamuro and Goldberg (1989).

[0386] Genetic targeting The guide polynucleotide / Cas system described herein may be used for gene targeting.

[0387] Generally, DNA targeting can be performed by cleaving one or both strands at a specific polynucleotide sequence in a cell using a Cas polypeptide associated with a suitable polynucleotide component. When single-strand or double-strand breaks are induced in the DNA, the cell's DNA repair mechanisms are activated, and the breaks are repaired by non-homologous end joining (NHEJ) or homologous recombination repair (HDR) processes, thereby potentially modifying the target site.

[0388] The length of the DNA sequence at the target site can vary, for example, including target sites with a length of at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more than 30 nucleotides. The target site may be palindromic, meaning that the sequence on one strand can be read in the opposite direction on the complementary strand. The cleavage site may be located within the target sequence, or it may be located outside the target sequence. In another variation, the cleavage may occur at nucleotide positions directly opposite each other, resulting in a blunt-end cleavage, or in other cases, the cleavage may be offset, resulting in a single-stranded overhang (which may be a 5' or 3' overhang), also known as a "sticky end." Active variants of the genomic target site may also be used. These active variants may contain at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with respect to a given target site, and the active variants may retain biological activity and therefore be recognized and cleaved by Cas polypeptides.

[0389] Assays for measuring single-strand or double-strand breaks at target sites by endonucleases are known in the art and generally measure the overall activity and specificity of the agent acting on the DNA substrate containing the recognition site.

[0390] The targeting methods described herein may be carried out, for example, so that two or more DNA target sites are targeted in this manner. Such methods may be optionally characterized as multiplexing. In certain embodiments, two, three, four, five, six, seven, eight, nine, ten or more target sites may be targeted simultaneously. Multiplexing is typically carried out by the targeting methods described herein, which provide multiple different RNA components (each designed to guide a guide polynucleotide / Cas polypeptide complex to a specific DNA target site).

[0391] gene editing A genome sequence editing process combining a DSB and a modification template generally involves introducing into a host cell a DSB inducer or nucleic acid encoding a DSB inducer that recognizes a target sequence in a chromosomal sequence and can introduce a DSB into the genome sequence, and at least one polynucleotide modification template that includes a change of at least one nucleotide compared to the nucleotide sequence to be edited. The polynucleotide modification template may further include at least one nucleotide change and an adjacent nucleotide sequence, the adjacent sequence being substantially homologous to a chromosomal region adjacent to the DSB. Genome editing using DSB inducers (e.g., Cas-gRNA complexes) is described, for example, in U.S. Patent Application Publication No. 20150082478, published March 19, 2015; International Publication No. 2015026886, published February 26, 2015; International Publication No. 2016007347, published January 14, 2016; and International Publication No. 2016 / 025131, published February 18, 2016.

[0392] Several uses of guide RNA / Cas polypeptide systems have been described (see, for example, U.S. Patent Application Publication No. 20150082478A1, published March 19, 2015; International Publication No. 2015026886, published February 26, 2015; and U.S. Patent Application Publication No. 20150059010, published February 26, 2015), including, but not limited to, modification or substitution of a target nucleotide sequence (e.g., regulatory elements), insertion of a target polynucleotide, gene knockout, gene knock-in, modification of a splicing site, and / or introduction of an alternative splicing site, modification of a nucleotide sequence encoding a target protein, fusion of amino acids and / or proteins, and gene silencing by expressing a reverse repeat in a target gene.

[0393] Proteins can be modified in various ways, including amino acid substitutions, deletions, truncations, and insertions. Methods for such operations are generally known. For example, amino acid sequence variants of proteins can be prepared by mutations in DNA. Examples of mutagenesis and nucleotide sequence modification include Kunkel, (1985) Proc. Natl. Acad. Sci. USA 82:488-92; Kunkel et al., (1987) Meth Enzymol 154:367-82; U.S. Patent No. 4,873,192; Walker and Gaastra, eds. (1983) Techniques in Molecular Biology (MacMillan Publishing Company, New York), and the literature cited therein. Guidance on amino acid substitutions that are unlikely to affect the biological activity of proteins can be found, for example, in the model of Dayhoff et al. (1978) Atlas of Protein Sequence and Structure (Natl Biomed Res Found, Washington, DC). Conservative substitutions, such as replacing one amino acid with another amino acid having similar properties, are sometimes preferred. Conservative deletions, insertions, and amino acid substitutions are expected not to cause radical changes in the protein's properties, and the effects of substitutions, deletions, insertions, or combinations thereof can be evaluated by routine screening analyses. Assays for double-strand break-inducing activity are known and generally measure the overall activity and specificity of the active substance on a DNA substrate containing the target site.

[0394] This specification describes a genome editing method using Cas polypeptides and a complex comprising Cas polypeptides and guide polynucleotides. Following the characterization of guide RNA and PAM sequences, components of endonucleases and associated CRISPR RNA (crRNA) can be used to modify chromosomal DNA in organisms including plants. To facilitate optimal expression and nuclear localization (in eukaryotic cells), the gene comprising this complex can be optimized as described in International Publication No. 2016186953, published November 24, 2016, and then delivered to cells as a DNA expression cassette by methods known in the art. The components necessary for the active complex can be delivered as RNA with or without modifications to protect the RNA from degradation, or as capped or uncapped mRNA (Zhang, Y. et al., 2016, Nat.Commun. 7:12617), or as a Cas polypeptide-guided polynucleotide complex (International Publication No. 2017070032, published April 27, 2017), or as any combination thereof. In addition, some or more parts of the complex and crRNA can be expressed from a DNA construct, while the other components can be delivered as RNA with or without modifications to protect the RNA from degradation, i.e., as capped or uncapped mRNA (Zhang et al. 2016 Nat.Commun. 7:12617), or as a Cas polypeptide-guided polynucleotide complex (International Publication No. 2017 / 070032, published April 27, 2017), or as any combination thereof. To generate crRNA in vivo, for example, as described in International Publication No. 2017105991, published on June 22, 2017, endogenous RNAse may also be recruited to cleave the crRNA transcript into a mature form that can induce a complex at its DNA target site, using tRNA-derived elements. The nickase complex may be used individually or in coordination to generate one or more DNA nicks on one or both of the DNA strands.Furthermore, the cleavage activity of Cas endonuclease can be inactivated by modifying key catalytic residues in its cleavage domain (Sinkunas, T et al., 2013, EMBO J.32:385-394), resulting in an RNA-induced helicase that can be used to enhance homologous recombination repair, induce transcriptional activation, or reconstruct DNA local structures. Moreover, the activity of both the Cas cleavage domain and the helicase domain can be knocked out, and these can be used in combination with other DNA cleavage, DNA nicking, DNA binding, transcriptional activation, transcriptional repression, DNA reconstruction, DNA deamination, DNA unwinding, DNA recombination enhancement, DNA integration, DNA inversion, and DNA repair agents.

[0395] The transcriptional directions of the tracrRNA (if present) and other components of the CRISPR-Cas system (e.g., variable targeting domain, crRNA repeats, loops, anti-repeats) can be estimated as described in International Publication No. 2016186946 and International Publication No. 2016186953, both published on November 24, 2016.

[0396] Once appropriate guide RNA requirements are established as described herein, PAM priorities for each of the novel systems disclosed herein can be investigated. If the cleavage complex results in degradation of the randomized PAM library, the complex can be converted to a nickase by mutagenerating key residues or by assembling the reaction in the absence of ATP as previously described (Sinkunas, T. et al., 2013, EMBO J.32:385-394) by inactivating ATPase-dependent helicase activity. Two regions of PAM randomization separated by two protospacer targets can be used to generate double-strand DNA breaks, which can be captured and sequenced to investigate the PAM sequences supporting the breaks by each complex.

[0397] In some embodiments, a method for modifying a target site in the genome of a cell comprises introducing at least one PGEN as described herein into a cell and identifying at least one cell having the modification at the target, wherein the modification at the target site is selected from the group consisting of (i) replacement of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, chemical modification of at least one nucleotide, and (v) any combination of (i) to (iv).

[0398] The nucleotide to be edited may be located within or outside the target site that is recognized and cleaved by the Cas endonuclease. In one embodiment, at least one nucleotide modification is not a modification at the target site recognized and cleaved by the Cas endonuclease. In another embodiment, there are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 900, or 1000 nucleotides between the at least one nucleotide to be edited and the genomic target site.

[0399] Knockout can be caused by indels (insertion or deletion of nucleotide bases in the target DNA sequence by NHEJ) or by the specific removal of a sequence that reduces or completely eliminates the function of a sequence at or near the target site.

[0400] Guide polynucleotide / Cas endonuclease-induced targeted mutations can occur within nucleotide sequences located inside or outside the genomic target site that is recognized and cleaved by Cas endonuclease.

[0401] In one embodiment, the present disclosure describes a method for modifying a target site in the genome of a cell, comprising introducing at least one PGEN as described herein and at least one donor DNA into the cell, wherein the donor DNA comprises a polynucleotide of interest, and further comprising optionally identifying at least one cell into which the polynucleotide of interest is incorporated at or near the target site.

[0402] In some embodiments, the methods disclosed herein may utilize homologous recombination (HR) to provide the integration of the desired polynucleotide at a target site.

[0403] Various methods and compositions can be used to generate cells or organisms having a target polynucleotide inserted into a target site by the activity of the CRISPR-Cas system components described herein. In one method described herein, the target polynucleotide is introduced into a biological cell by a donor DNA construct. As used herein, “donor DNA” is a DNA construct containing the target polynucleotide to be inserted into the genomic target site of the Cas polypeptide. The donor DNA construct further comprises first and second homologous regions adjacent to the target polynucleotide. The first and second homologous regions of the donor DNA are homologous to first and second genomic regions present in or adjacent to the target site in the genome of the cell or organism, respectively.

[0404] Donor DNA can be bound to guide polynucleotides. Tethered donor DNA can enable co-localization of target and donor DNA, which is useful in genome editing, gene insertion, and target genome regulation, and may also be useful in targeting mitotic terminal cells in which the function of the endogenous HR mechanism is expected to be significantly impaired (Mali et al., 2013, Nature Methods Vol.10:957-963).

[0405] The amount of homology or sequence identity shared by target and donor polynucleotides can vary, ranging from approximately 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, and 450-90 bp. The ranges include 0 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5-3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5-7 kb, 4-8 kb, 5-10 kb, or the total length and / or region having unit integer values ​​within the ranges less than or equal to the total length of the target site. These ranges include all integers within this range; for example, the range 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 bp. The amount of homology can also be explained by the percentage of sequence identity over the entire fully aligned length of the two polynucleotides, which includes at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of sequence identity. Sufficient homology includes any combination of polynucleotide length, overall sequence identity percentage, and any optional conserved region of consecutive nucleotides or local sequence identity percentage. For example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity with respect to the region of the target locus.Sufficient homology can also be explained by the predictive ability of two polynucleotides to specifically hybridize under high stringency conditions; see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology--Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0406] Episomal DNA molecules can also be ligated into double-strand breaks, for example, to incorporate T-DNA into chromosomal double-strand breaks (Chilton and Que, (2003) Plant Physiol 133:956-65; Salomon and Puchta, (1998) EMBO J 17:6086-95). If the sequence surrounding a double-strand break is altered, for example, by exonuclease activity involved in the maturation of the double-strand break, the gene conversion pathway can revert to its original structure if homologous sequences, such as homologous chromosomes in non-cleavageable somatic cells or sister chromatids after DNA replication, are available (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenetic DNA sequences can also function as DNA repair templates for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0407] In one embodiment, the present disclosure is a method for editing a nucleotide sequence in the genome of a cell, comprising introducing at least one PGEN as described herein and a polynucleotide modification template into the cell, wherein the polynucleotide modification template comprises at least one nucleotide modification of the nucleotide sequence, and further comprising optionally selecting at least one cell containing the edited nucleotide sequence.

[0408] This guide polynucleotide / Cas endonuclease system can be used in combination with at least one polynucleotide modification template to enable the editing (modification) of a target genomic nucleotide sequence. (See also U.S. Patent Application Publication No. 20150082478, published March 19, 2015, and International Publication No. 2015026886, published February 26, 2015).

[0409] The target polynucleotides and / or traits can be stacked together in a complex trait locus, as described in International Publication No. 2012129373, published September 27, 2012, and International Publication No. 2013112686, published August 1, 2013. The guide polynucleotide / Cas9 endonuclease system described herein provides a system for efficiently generating double-strand breaks, enabling the stacking of traits in a complex trait locus.

[0410] The guide polynucleotide / Cas system described herein, which mediates gene targeting, may be used in a manner similar to that disclosed in International Publication No. 2012129373, published September 27, 2012, in a method for inducing heterogeneous gene insertion and / or generating complex trait loci containing multiple heterogenes, in which the guide polynucleotide / Cas system disclosed herein is used instead of the use of a double-strand break inducer to introduce the gene of interest. The transgenes can be propagated as a single locus by inserting independent transgenes within 0.1, 0.2, 0.3, 0.4, 0.5, 1.0, 2, or 5 centimorgans (cM) of each other (see, for example, U.S. Patent Application Publication No. 20130263324, published October 3, 2013, or International Publication No. 2012129373, published March 14, 2013). After selecting plants containing the introduced gene, the plants containing (at least) one introduced gene can be crossbred to form F1 plants containing both introduced genes. Of these F1 (F2 or BC1) offspring, 1 / 500 will have two different introduced genes rearranged on the same chromosome. This composite locus can then be propagated as a single locus along with the traits of both introduced genes. By repeating this process, any desired traits can be accumulated.

[0411] Further use of guide RNA / Cas polypeptide systems is described (e.g., U.S. Patent Application Publication No. 20150082478, published March 19, 2015; International Publication No. 2015026886, published February 26, 2015; U.S. Patent Application Publication No. 20150059010, published February 26, 2015; International Publication No. 2016007347, published January 14, 2016; and PCT Application International Publication No. 20160251, published February 18, 2016). (See Pamphlet No. 31) Examples of such methods include, but are not limited to, modification or substitution of a target nucleotide sequence (e.g., regulatory elements), insertion of a target polynucleotide, gene knockout, gene knock-in, modification of a splicing site and / or introduction of an alternative splicing site, modification of a nucleotide sequence encoding a target protein, fusion of amino acids and / or proteins, and gene silencing by expressing a reverse repeat in a target gene.

[0412] The characteristics obtained from the gene editing compositions and methods described herein can be evaluated. Chromosome intervals that correlate with the desired phenotype or trait can be identified. Various methods well known in the art can be used to identify chromosome intervals. The boundaries of such chromosome intervals are drawn to encompass markers linked to genes controlling the desired trait. In other words, chromosome intervals are defined such that any marker present within the interval (including terminal markers defining the interval boundaries) can be used as markers for a particular trait. In one embodiment, a chromosome interval may contain at least one QTL, and may even contain multiple QTLs. If multiple QTLs are very close together within the same interval, the association of a particular marker with a particular QTL becomes unclear, as one marker may be linked to one or more QTLs. Conversely, for example, if two very close markers cosegregate with the desired phenotypic trait, it is often unclear whether each of these markers identifies the same QTL or two different QTLs. The term “quantitative trait locus” or “QTL” refers to a region of DNA associated with differences in the expression of a quantitative phenotypic trait within at least one genetic background, e.g., at least one breeding population. A QTL region contains or is closely associated with a gene that affects the trait in question. A “QTL allele” can contain multiple genes or other genetic factors within a contiguous genomic region or linkage group, such as a haplotype. A QTL allele can mean a haplotype within a particular window, which is a contiguous genomic region defined and traceable by a set of one or more polymorphic markers. A haplotype can be defined by the allele’s unique fingerprint at each marker position within a particular window.

[0413] Introduction of CRISPR-Cas system components into cells The methods and compositions described herein do not depend on any specific method for introducing a sequence into an organism or cell, but merely involve introducing a polynucleotide or polypeptide into the interior of at least one cell of that organism. Introduction includes references to the integration of nucleic acids into eukaryotic or prokaryotic cells, where the nucleic acid can be incorporated into the cell's genome, and also includes references to the transient (direct) delivery of nucleic acids, proteins, or polynucleotide-protein complexes (PGEN, RGEN) to cells.

[0414] Methods for introducing polynucleotides or polypeptides, or polynucleotide-protein complexes, into cells or organisms are known in the art and include, but are not limited to, microinjection, electroporation, stable transformation, transient transformation, ballistic particle acceleration (particle impact), whisker-mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, virus-mediated transfer, transfection, transduction, cell-permeable peptides, mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, topical application, sexual hybridization, sexual reproduction, and any combination thereof.

[0415] For example, guide polynucleotides (guide RNA, crnucleotide + tracr nucleotide, guide DNA, and / or guide RNA-DNA molecules) can be directly (transiently) introduced into cells as single-stranded or double-stranded polynucleotide molecules. Guide RNA (or crRNA + tracrRNA) can also be introduced indirectly into cells by introducing a recombinant DNA molecule containing a heterologous nucleic acid fragment encoding the guide RNA (or crRNA + tracrRNA), which is operably ligated to a specific promoter capable of transcribing the guide RNA (crRNA + tracrRNA molecule) within the cell. Certain promoters may be RNA polymerase III promoters that enable transcription of RNA having strictly defined and unmodified 5' and 3' ends, though not limited to these specific promoters (Ma et al., 2014, Mol.Ther.Nucleic Acids 3:e161; DiCarlo et al., 2013, Nucleic Acids Res.41:4336-4343; International Publication No. 2015026887, published February 26, 2015). Any promoter capable of transcribing guide RNA within the cell may be used, including heat shock / heat-inducible promoters operably ligated to the nucleotide sequence encoding the guide RNA.

[0416] Unlike animal cells (e.g., human cells), fungal cells (e.g., yeast cells), and protoplasts, plant cells, for example, contain a plant cell wall that can act as a barrier to the delivery of components.

[0417] Delivery of Cas polypeptides and / or guide RNAs and / or ribonucleoprotein complexes and / or polynucleotides encoding any one or more of the aforementioned into plant cells can be achieved by methods known in the art, including, but not limited to, the following: Rhizobiales-mediated transformation (e.g., Agrobacterium, Ochrobactrum), particle-mediated delivery (particle impact), polyethylene glycol (PEG)-mediated transfection (e.g., into protoplasts), electroporation, cell-permeable peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery.

[0418] Cas polypeptides, such as those described herein, can be introduced into cells by direct delivery of the Cas polypeptide itself (referred to as direct delivery of the Cas polypeptide), the mRNA encoding the Cas polypeptide, and / or the guide polynucleotide / Cas polypeptide complex itself, using any method known in the art. Cas polypeptides can also be introduced indirectly into cells by introducing a recombinant DNA molecule encoding the Cas polypeptide. Endonucleases can be transiently introduced into cells or incorporated into the host cell genome using any method known in the art. The uptake of endonucleases and / or derived polynucleotides into cells can be facilitated using cell-permeable peptides (CPPs), as described in International Publication No. 2016073433, published May 12, 2016. Any promoter capable of expressing Cas polypeptides in cells can be used, including heat shock / heat-inducible promoters operably ligated to the nucleotide sequence encoding the Cas polypeptide.

[0419] Direct delivery of polynucleotide-modified templates to plant cells can be achieved by particle-mediated delivery, and any other direct delivery method, including but not limited to polyethylene glycol (PEG)-mediated transfection of protoplasts, whisker-mediated transformation, electroporation, particle impaction, cell-permeable peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, can be used without problems for the delivery of polynucleotide-modified templates in eukaryotic cells such as plant cells.

[0420] Donor DNA can be introduced by any means known in the art. Donor DNA can be provided by any transformation method known in the art, including, for example, Agrobacterium-mediated transformation or bioristic particle impaction. Donor DNA may exist transiently within cells or may be introduced via viral replicons. In the presence of Cas endonuclease and a target site, donor DNA is inserted into the transformed plant genome.

[0421] Direct delivery of any one of the inducible Cas system components may be accompanied by the direct delivery (co-delivery) of other mRNAs that may facilitate the enrichment and / or visualization of cells receiving the guide polynucleotide / Cas polypeptide complex component. For example, direct co-delivery of the guide polynucleotide / Cas polypeptide component (and / or the guide polynucleotide / Cas polypeptide complex itself) together with mRNA encoding a phenotypic marker (e.g., but not limited to CRC (Bruce et al. 2000 The Plant Cell 12:65-79) or other transcription activators may enable cell selection and enrichment without the use of exogenous selection markers by restoring the function of non-functional gene products, as described in International Publication No. 2017070032, published on April 27, 2017.

[0422] The introduction of the guide RNA / Cas polypeptide complex (representing the cleavage-ready complex as described herein) into cells includes introducing the individual components of the complex separately or together into cells, and introducing them directly (delivered directly as RNA for the guide and protein and protein subunits for the Cas polypeptide, or functional fragments thereof) or by recombinant constructs expressing the components (guide RNA, Cas polypeptide, protein subunits, or functional fragments thereof). The introduction of the guide RNA / Cas endonuclease complex (RGEN) into cells includes introducing the guide RNA / Cas polypeptide complex into cells as a ribonucleotide protein. This ribonucleotide protein may be assembled before introduction into cells as described herein. The components constituting the guide RNA / Cas endonuclease ribonucleotide protein (at least one Cas endonuclease, at least one guide RNA, and at least one protein subunit) may be assembled in vitro or before introduction into cells (which are targeted for genome modification as described herein) by any means known in the art.

[0423] Direct delivery of RGEN ribonucleoprotein enables genome editing at target sites in the cellular genome, with the complex subsequently degrading rapidly and its presence in the cell becoming transient. This transient presence of the RGEN complex can lead to a reduction in off-target effects. In contrast, delivery of RGEN components (guide RNA, Cas9 endonuclease) via plasmid DNA sequences can result in sustained expression of RGEN from these plasmids, which can increase off-target effects (Cradick, T.Jet et al. (2013) Nucleic Acids Res 41:9584-9592; Fu, Y. et al. (2014) Nat. Biotechnol. 31:822-826).

[0424] Direct delivery can be achieved by combining any one component of a guide RNA / Cas endonuclease complex (RGEN) (representing the cleavage-ready complex described herein) (e.g., at least one guide RNA, at least one Cas polypeptide, and one optional additional protein) with a delivery matrix containing microparticles (e.g., but not limited to gold particles, tungsten particles, and silicon carbide whisker particles) (see also International Publication No. 2017070032, published April 27, 2017). This delivery matrix may contain any one of the components, e.g., a Cas endonuclease, which is attached to a solid matrix (e.g., impact particles).

[0425] In some embodiments, the guide polynucleotide / Cas polypeptide complex is a complex in which the guide RNA and Cas polypeptide that form the guide RNA / Cas polypeptide complex are introduced into the cell as RNA and protein, respectively.

[0426] In some embodiments, the guide polynucleotide / Cas polypeptide complex is a complex in which the guide RNA and Cas polypeptide protein that form the guide RNA / Cas polypeptide complex, as well as at least one protein subunit of the complex, are introduced into the cell as RNA and protein, respectively.

[0427] In some embodiments, the guide polynucleotide / Cas endonuclease complex is a complex in which the guide RNA and Cas endonuclease protein, which form the guide RNA and Cas endonuclease complex (ready-to-cleave complex), and at least one protein subunit of the complex are pre-assembled in vitro and introduced into cells as a ribonucleotide-protein complex.

[0428] Known protocols for introducing polynucleotides, polypeptides, or polynucleotide-protein complexes (PGENs, RGENs) into eukaryotic cells such as plants or plant cells include: microinjection (Crossway et al., (1986) Biotechniques 4:320-34 and U.S. Patent No. 6,300,543), meristematic tissue transformation (U.S. Patent No. 5,736,369), electroporation (Riggs et al., (1986) Proc. Natl. Acad. Sci. USA 83:5602-6, Agrobacterium-mediated transformation (U.S. Patent Nos. 5,563,055 and 5,981,840), and whisker-mediated transformation (Ainley et al. 2013, Plant Biotechnology Journal 11:1126-1134; Shaheen A. and M. Arshad 2011 Properties and Applications of Silicon Carbide (2011), 345-358 Editor(s): Gerhardt, Rosario. Publisher: InTech, Rijeka, Croatia. CODEN: 69PQBP; ISBN: 978-953-307-201-2), direct gene transfer (Paszkowski et al., (1984) EMBO J 3:2717-22), and ballistic particle acceleration (US Patent Nos. 4,945,050; 5,879,918; 5,886,244; 5,932,782; Tomes et al., (1995) “Direct DNA Transfer into Intact Plant Cells via Microprojectile Bombardment” in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, ed. Gamborg & Phillips (Springer-Verlag, Berlin); McCabe et al., (1988) Biotechnology 6:923-6; Weissinger et al.,(1988)Ann Rev Genet 22:421-77;Sanford et al.,(1987)Particulate Science and Technology 5:27-37(Onion);Christou et al.,(1988)Plant Physiol 87:671-4(Soybean);Finer and McMullen,(1991)In vitro Cell Dev Biol 27P:175-82(Soybean);Singh et al.,(1998)Theor Appl Genet 96:319-24(Soybean);Datta et al.,(1990)Biotechnology 8:736-40(Rice);Klein et al.,(1988)Proc.Natl.Acad.Sci.USA 85:4305-9(Maize);Klein et al.,(1988)Biotechnology 6:559-63 (Maize); U.S. Patent No. 5,240,855; U.S. Patent No. 5,322,783, and U.S. Patent No. 5,324,646; Klein et al., (1988) Plant Physiol 91:440-4 (Maize); Fromm et al., (1990) Biotechnology 8:833-9 (Maize); Hooykaas-Van Slogteren et al., (1984) Nature 311:763-4; U.S. Patent No. 5,736,369 (Grain Grasses); Bytebier et al., (1987) Proc. Natl. Acad. Sci. USA 84:5345-9 (Liliaceae); De Wet et al, (1985) in The Experimental Manipulation of Ovule Tissues, ed. Chapman et al., (Longman, New York), pp. 197-209 (pollen); Kaeppler et al., (1990) Plant Cell Rep 9:415-8), and Kaeppler et al., (1992) Theor Appl Genet 84:560-6 (whisker-borne transformation); D'Halluin et al.(1992) Plant Cell 4:1495-505 (Electroporation); Li et al., (1993) Plant Cell Rep 12:250-5; Christou and Ford (1995) Annals Botany 75:407-13 (Rice), and Osjoda et al., (1996) Nat Biotechnol 14:745-50 (Maize via Agrobacterium tumefaciens).

[0429] Alternatively, polynucleotides may be introduced into plants or plant cells by contacting cells or organisms with viruses or viral nucleic acids. Generally, such methods involve incorporating polynucleotides into viral DNA or RNA molecules. In some examples, the polypeptide of interest may first be synthesized as part of a viral polyprotein, which is then subjected to in vivo or in vitro proteolytic treatment to produce the desired recombinant protein. Methods for introducing polynucleotides, including viral DNA or RNA molecules, into plants and expressing the proteins encoded therein are known; see, for example, U.S. Patents 5,889,191, 5,889,190, 5,866,785, 5,589,367, and 5,316,931.

[0430] Polynucleotides or recombinant DNA constructs can be provided to or introduced into prokaryotic and eukaryotic cells or organisms using various transient transformation methods. Such transient transformation methods include, but are not limited to, the direct introduction of polynucleotide constructs into plants.

[0431] Nucleic acids and proteins may be delivered to cells by any method, including using molecules (e.g., cell-permeable peptides and nanocarriers) that facilitate the uptake of any or all components (proteins and / or nucleic acids) of the inducible Cas system. See also U.S. Patent Application Publication No. 20110035836, published February 10, 2011, and European Patent Application Publication No. 2821486A1, published January 7, 2015.

[0432] Other methods for introducing polynucleotides into prokaryotic and eukaryotic cells or parts of organisms or plants (e.g., plastid transformation and methods for introducing polynucleotides into tissues from seedlings or mature seeds) may be used.

[0433] Stable transformation is intended to mean that a nucleotide construct introduced into an organism is integrated into the organism's genome and can be passed on to its offspring. Transient transformation is intended to mean that a polynucleotide is introduced into an organism but is not integrated into the organism's genome, or that a polypeptide is introduced into the organism. Transient transformation indicates that the introduced composition is expressed or present only temporarily within the organism.

[0434] Various methods can be used to identify cells with altered genomes at or near a target site without using screening marker phenotypes. Such methods can be considered as processes that directly analyze the target sequence to detect any changes within the target sequence, and include, but are not limited to, PCR, sequencing, nuclease digestion, Southern blotting, and any combination thereof.

[0435] Cells and plants The polynucleotides and polypeptides of this disclosure can be introduced into cells. Examples of cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, fungal, insect, yeast, unconventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein. Monocotyledonous and dicotyledonous plants, and any plant containing plant elements, can be used in conjunction with the compositions and methods described herein.

[0436] Examples of usable monocots include, but are not limited to, the following: maize (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), barnyard millet (e.g., pearl millet (Pennisetum glaucum)), foxtail millet (Panicum miliaceum), proso millet (Setaria italica), finger millet (Eleusine coracana), and wheat (species of the genus Triticum, e.g., Triticum aestibum). aestivum), Triticum monococcum, sugarcane (Saccharum spp.), wild oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), pineapple (Ananas comosus), banana (Musa spp.), palms, ornamental plants, turfgrass, and other grasses.

[0437] Examples of usable dicotyledonous plants include, but are not limited to, soybeans (Glycine max), Brassica species (e.g., rapeseed or canola, but not limited to) (Brassica napus, B. campestris, Brassica rapa, Brassica juncea), alfalfa (Medicago sativa), tobacco (Nicotiana tabacum), Arabidopsis (Arabidopsis thaliana), sunflower (Helianthus annuus), and cotton (Gossypium arboreum). arboreum, Gossypium barbadense, peanuts (Arachis hypogaea), tomatoes (Solanum lycopersicum), and potatoes (Solanum tuberosum).

[0438] Further usable plants include: safflower (Carthamus tinctorius), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea species), coconut (Cocos nucifera), citrus (Citrus species), cocoa (Theobroma cacao), tea plant (Camellia sinensis), banana (Musa species), avocado (Persea americana), and fig (Ficus cacica). casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beet (Beta vulgaris), vegetables, ornamental plants, and conifers.

[0439] Suitable vegetables include: tomatoes (Lycopersicon esculentum), lettuce (e.g., lettuce (Lactuca sativa)), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus species), and members of the Cucumis genus such as cucumbers (C. sativus), cantaloupes (C. cantalupensis) and muskmelons (C. melo). Ornamental plants include: azaleas (Rhododendron species), hydrangeas (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa species), tulips (Tulipa species), trumpet daffodils (Narcissus species), petunias (Petunia hybrida), carnations (Dianthus caryophyllus), poinsettias (Euphorbia pulcherrima), and chrysanthemums.

[0440] Suitable conifers include: pine, e.g., loblolly pine (Pinus taeda), slash pine (Pinus elliotii), Ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesii); American hemlock (Tsuga canadensis); Western spruce (Picea glauca); American redwood (Sequoia sempervirens); Sakhalin fir, e.g., European fir (Abies amabilis) amabilis)) and balsam fir (Abies balsamea); and Himalayan cedars, such as Western red cedar (Thuja plicata) and Alaskan yellow cedar (Chamaecyparis nootkatensis).

[0441] In some aspects of this disclosure, a fertile plant is a plant that produces viable male and female gametes and is self-fertile. Such a self-fertile plant can produce offspring plants without the contribution of gametes and the genetic material contained therein from other plants. Other aspects of this disclosure may involve the use of a non-self-fertile plant, such that the plant does not produce viable or otherwise fertilizable male or female gametes, or both.

[0442] This disclosure finds use in plant breeding of plants that include one or more introduced traits or edited genomes.

[0443] A non-limiting example of how two traits can be stacked in the genome, for example, at a genetic distance of 5 cM from each other, is described as follows: A first plant containing a first transgenic target site integrated into a first DSB target site within the genome window, but lacking a first target genomic locus, is crossed with a second transgenic plant containing the target genomic locus at a different genomic insertion site within the genome window, the second plant lacking the first transgenic target site. Approximately 5% of the plant offspring from this cross possess both the first transgenic target site integrated into the first DSB target site and the first target genomic locus integrated at a different genomic insertion site within the genome window. Progeny plants having both sites within the defined genome window may be further crossed with a third transgenic plant that contains a second transgenic target site incorporated into a second DSB target site and / or a second target genomic locus within the defined genome window, but lacks the first transgenic target site and the first target genomic locus. After this, progeny having the first transgenic target site, the first target genomic locus, and the second target genomic locus incorporated into a different genomic insertion site within the genome window are selected. Using such methods, transgenic plants can be created that contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 or more complex trait loci having transgenic target sites incorporated into DSB target sites and / or desired genomic loci incorporated into different sites within the genome window. In this way, a variety of complex trait loci can be generated.

[0444] Cells and animals The polynucleotides and polypeptides of this disclosure can be introduced into animal cells. Animal cells include, but are not limited to, organisms of the phylums including chordates, arthropods, mollusks, annelids, cnidarians, or echinoderms; or organisms of the classes including mammals, insects, birds, amphibians, reptiles, or fish. In some embodiments, the animals are humans, mice, C. elegans, rats, fruit flies (Drosophila species), zebrafish, chickens, dogs, cats, guinea pigs, hamsters, medaka, sea lampreys, pufferfish, tree frogs (e.g., Xenopus species), monkeys, or chimpanzees. Specific cell types that may be considered include: monocytes, diploid cells, germ cells, neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, germ cells, hematopoietic cells, osteocytes, germ cells, somatic cells, stem cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some embodiments, multiple cells of biological origin may be used.

[0445] The Cas polypeptides disclosed herein may be used to edit the genome of animal cells in various ways. In some embodiments, deletion of one or more nucleotides may be desired. In other embodiments, insertion of one or more nucleotides may be desired. In some embodiments, substitution of one or more nucleotides may be desired. In other embodiments, modification of one or more nucleotides may be desired by covalent or non-covalent interactions with other atoms or molecules.

[0446] Genome modification using Cas polypeptides can be used to alter the genotype and / or phenotype of a target organism. Such alterations are preferably related to the improvement of a desired phenotype or physiologically important trait, the correction of an endogenous defect, or the expression of certain types of expression markers. In some embodiments, the desired phenotype or physiologically important trait is related to the overall health, adaptability, or reproductive capacity of the animal, the ecological adaptability of the animal, or the relationship or interaction of this animal with other organisms in its environment. In some embodiments, the target phenotype or physiologically important characteristic is selected from the following groups: improvement of overall health, improvement of disease, improvement of disease, stabilization of disease, prevention of disease, treatment of parasitic infections, treatment of viral infections, treatment of retroviral infections, treatment of bacterial infections, treatment of neurological disorders (e.g., multiple sclerosis, but not limited to these), correction of endogenous gene defects (e.g., metabolic disorders, chondrodysplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Barth syndrome, breast cancer, Charcot-Marie-Tooth disease, colorectal cancer, cat cry syndrome, Crohn's disease, cystic fibrosis, Darkham's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombotic diathesis, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease) , hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland syndrome, porphyria, progeria, prostate cancer, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, skin cancer, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatinofacial syndrome, WAGR syndrome, and Wilson's disease), treatment of congenital immunodeficiencies (e.g., immunoglobulin subclass deficiencies, but not limited to these), treatment of acquired immunodeficiencies (e.g., AIDS and other HIV-related disorders, but not limited to these), treatment of cancer, and treatment of diseases including rare or "orphan" conditions for which there are no other effective treatment options.

[0447] Cells genetically modified using the compositions or methods disclosed herein may be transplanted into subjects for purposes such as gene therapy, for example, to treat diseases, or as antiviral, antipathogen, or anticancer agents, for the production of genetically modified organisms in agriculture, or for biological research.

[0448] In vitro detection, binding, and modification of polynucleotides The compositions disclosed herein may, in some embodiments involving isolated polynucleotide sequences, be further used as compositions for use in in vitro methods. The isolated polynucleotide sequence may comprise one or more target sequences for modification. In some embodiments, the isolated polynucleotide sequence may be genomic DNA, a PCR product, or a synthesized oligonucleotide.

[0449] composition Modification of the target sequence may take the form of nucleotide insertion, nucleotide deletion, nucleotide substitution, addition of atomic molecules to existing nucleotides, nucleotide alteration, or conjugation of heterologous polynucleotides or polypeptides to the target sequence. Insertion of one or more nucleotides may be achieved by including a donor polynucleotide in the reaction mixture, the donor polynucleotide being inserted into the double-strand break region generated by the Cas-alpha ortholog polypeptide. This insertion may be carried out by non-homologous end joining or homologous recombination.

[0450] In some embodiments, the sequence of the target polynucleotide is known before modification and is compared with the sequence of the polynucleotide obtained from treatment with a Cas-alpha ortholog. In some embodiments, the sequence of the target polynucleotide is unknown before modification and treatment with a Cas-alpha ortholog is used as part of a method for determining the sequence of the target polynucleotide.

[0451] In some embodiments, the Cas-alpha-ortholog may be selected from the group consisting of unmodified wild-type Cas-alpha-orthologs, functional Cas-alpha-ortholog variants, functional Cas-alpha-ortholog fragments, fusion proteins containing active or inactivated Cas-alpha-orthologs, Cas-alpha-orthologs further containing one or more nuclear localization sequences (NLS) at the C-terminus or N-terminus, or both the N-terminus and C-terminus, biotinylated Cas-alpha-orthologs, Cas-alpha-ortholog niccasases, Cas-alpha-ortholog endonucleases, Cas-alpha-orthologs further containing histidine tags, and any two or more mixtures thereof.

[0452] In some embodiments, the Cas-alpha ortholog is a fusion protein further comprising a nuclease domain, a transcriptional activator domain, a transcriptional repressor domain, an epigenetic modification domain, a cleavage domain, a nuclear localization signal, a cell permeability domain, a translocation domain, a marker, or a target polynucleotide sequence or a transgene heterologous to the cell from which the target polynucleotide sequence is obtained or induced.

[0453] In some embodiments, multiple Cas-alpha-orthologues may be desired. In some embodiments, the multiple Cas-alpha-orthologues may include Cas-alpha-orthologues derived from different biosources or from different gene loci within the same organism. In some embodiments, the multiple Cas-alpha-orthologues may include Cas-alpha-orthologues having different binding specificities to target polynucleotides. In some embodiments, the multiple Cas-alpha-orthologues may include Cas-alpha-orthologues having different cleavage efficiencies. In some embodiments, the multiple Cas-alpha-orthologues may include Cas-alpha-orthologues having different PAM properties. In some embodiments, the multiple Cas-alpha-orthologues may include orthologues with different molecular compositions, i.e., polynucleotide Cas-alpha-orthologues and polypeptide Cas-alpha-orthologues.

[0454] Guide polynucleotides may be provided as a single guide RNA (sgRNA), a chimeric molecule containing tracrRNA, a chimeric molecule containing crRNA, a chimeric RNA-DNA molecule, a DNA molecule, or a polynucleotide containing one or more chemically modified nucleotides.

[0455] Storage conditions for Cas-alpha-orthologs and / or guide polynucleotides include parameters relating to temperature, state of matter, and time. In some embodiments, Cas-alpha-orthologs and / or guide polynucleotides are stored at approximately -80°C, approximately -20°C, approximately 4°C, approximately 20-25°C, or approximately 37°C. In some embodiments, Cas-alpha-orthologs and / or guide polynucleotides are stored as liquid, cryogenic liquid, or lyophilized powder. In some embodiments, Cas-alpha-orthologs and / or guide polynucleotides are stable for at least one day, at least one week, at least one month, at least one year, or for more than one year.

[0456] Any or all of the reactable polynucleotide components (e.g., guide polynucleotides, donor polynucleotides, and optionally Cas-alpha polynucleotides) may be supplied as part of a vector, construct, linearized or cyclic plasmid, or as part of a chimeric molecule. Each component may be supplied separately or together to the reaction mixture. In some embodiments, one or more polynucleotide components are operably linked to heterogeneous non-coding regulatory elements that modulate their expression.

[0457] A method for modifying a target polynucleotide comprises combining minimal elements to form a reaction mixture comprising a Cas-alpha ortholog (or a variant, fragment, or other related molecule as described above), a guide polynucleotide containing a sequence substantially complementary to or selectively hybridizing with the target polynucleotide sequence of the target polynucleotide, and the target polynucleotide for modification. In some embodiments, the Cas-alpha ortholog is provided as a polypeptide. In some embodiments, the Cas-alpha ortholog is provided as a Cas-alpha ortholog polynucleotide. In some embodiments, the guide polynucleotide is provided as a polynucleotide molecule containing an RNA molecule, a DNA molecule, an RNA:DNA hybrid, or a chemically modified nucleotide.

[0458] Any one of the components of the storage buffer or reaction mixture may be optimized for stability, efficacy, or other parameters. Further components of the storage buffer or reaction mixture may include buffer compositions, Tris, EDTA, dithiothreitol (DTT), phosphate-buffered saline (PBS), sodium chloride, magnesium chloride, HEPES, glycerol, BSA, salts, emulsifiers, detergents, chelating agents, redox agents, antibodies, nuclease-free water, proteinases, and / or viscous agents. In some embodiments, the storage buffer or reaction mixture further includes a buffer comprising at least one of the following components: HEPES, MgCl2, NaCl, EDTA, proteinase, proteinase K, glycerol, and nuclease-free water.

[0459] Incubation conditions vary according to the desired results. The temperature is preferably at least 10°C, 10–15°C, at least 15°C, 15–17°C, at least 17°C, 17–20°C, at least 20°C, 20–22°C, at least 22°C, 22–25°C, at least 25°C, 25–27°C, at least 27°C, 27–30°C, at least 30°C, 30–32°C, at least 32°C, 32–35°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, or greater than 40°C. The incubation time is at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, or greater than 10 minutes.

[0460] The sequences of polynucleotides in the reaction mixture before, during, or after incubation can be determined by any method known in the art. In some embodiments, modifications of the target polynucleotide can be confirmed by comparing the sequence of a polynucleotide purified from the reaction mixture with the sequence of the target polynucleotide before combining it with a Cas-alpha ortholog.

[0461] One or more of the compositions disclosed herein, useful for in vitro or in vivo detection, binding, and / or modification of polynucleotides, may be included in the kit. The kit comprises a Cas-alpha-ortholog or a polynucleotide Cas-alpha-ortholog encoding such, optionally further comprising a buffer component enabling efficient storage, and one or more further compositions enabling the introduction of the Cas-alpha-ortholog or polynucleotide Cas-alpha-ortholog into a heterogeneous polynucleotide, wherein the Cas-alpha-ortholog or polynucleotide Cas-alpha-ortholog may result in modification, addition, deletion, or substitution of at least one nucleotide of the heterogeneous polynucleotide. In a further embodiment, the Cas-alpha-orthologs disclosed herein may be used for enrichment of one or more polynucleotide target sequences from a mixed pool. In a further embodiment, the Cas-alpha-orthologs disclosed herein may be immobilized on a matrix for use in in vitro detection, binding, and / or modifi...

Claims

1. A non-natural Cas-alpha-10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence comprises mutations relative to SEQ ID NO: 2, the mutations being: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72 Non-natural Cas-alpha-10 polypeptides containing the following mutations: G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

2. The aforementioned amino acid sequences are combinations of the following mutations relative to Sequence ID No. 2: a combination of K85Q mutation and N92L mutation; a combination of N88H mutation and Q89G mutation; a combination of N88K mutation and Q89G mutation; a combination of N88Q mutation and Q89G mutation; a combination of K85S mutation and N92L mutation; a combination of K85S mutation and N92Q mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92H mutation; K85 Combinations of S mutation and N92A mutation; combination of K85S mutation and N92M mutation; combination of K85N mutation and N92L mutation; combination of K85N mutation and N92H mutation; combination of K85N mutation and N92A mutation; combination of K85N mutation and N92C mutation; combination of K85N mutation and N92M mutation; combination of K85N mutation and N92Q mutation; combination of K85N mutation and N92I mutation; combination of K85S mutation, N88D mutation and Q89G mutation Mutation combinations; combinations of K85S mutation, N88H mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89D mutation, and Q125R mutation; combinations of K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; combinations of K85S mutation, N92L mutation, and Q125R mutation; Y72S mutation, K85D mutation, and Q125 The non-natural Cas-alpha-10 polypeptide according to claim 1, comprising one or more of the following combinations: R mutation and N127R mutation; Y72A mutation, K85S mutation, Q89D mutation, N92L mutation and Q125R mutation; Y72C mutation, N88H mutation, Q89G mutation and Q125R mutation; Y72C mutation, N88D mutation, Q89D mutation, N92W mutation and Q125R mutation; or K85Q mutation and N92W mutation.

3. The non-natural Cas-alpha 10 polypeptide according to claim 1 or claim 2, wherein the amino acid sequence comprises a PAM interaction (PI) domain containing amino acids from S63 to I196.

4. The aforementioned amino acid sequence comprises the K85S mutation, as described in any one of claims 1 to 3, for the non-natural Cas-alpha 10 polypeptide.

5. The non-natural Cas-alpha 10 polypeptide according to any one of claims 1 to 3, wherein the amino acid sequence includes a combination of the K85Q mutation and the N92L mutation.

6. The non-natural Cas-alpha 10 polypeptide according to any one of claims 1 to 3, wherein the amino acid sequence includes a combination of the K85N mutation and the N92L mutation.

7. The Cas-alpha-10 polypeptide is a non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 6, wherein the Cas-alpha-10 polypeptide has endonuclease activity.

8. The non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 6, wherein the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with deaminase.

9. The Cas-alpha-10 polypeptide is a non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 6, wherein the Cas-alpha-10 polypeptide has nickase activity.

10. The non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 6 or 9, wherein the Cas-alpha-10 polypeptide is complexed with a reverse transcriptase or operably associated with a reverse transcriptase.

11. The non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 6, wherein the Cas-alpha-10 polypeptide complexes with a heterologous protein domain via a linker, and the heterologous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity.

12. A non-natural Cas-alpha 10 polypeptide according to any one of claims 1 to 11, comprising an amino acid sequence having at least 90% sequence identity with any one of sequence numbers 6, 5, 7-16, or 28-87.

13. A non-natural Cas-alpha 10 polypeptide containing a PAM interaction (PI) domain, wherein the PI domain is 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC- A non-natural Cas-alpha-10 polypeptide that recognizes PAM sequences including 3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3' (wherein N = A, C, G, or T, Y = T or C, D = G, A, or T, W = A or T, H = A, T, or C).

14. The non-natural Cas-alpha-10 polypeptide according to claim 13, wherein the Cas-alpha-10 polypeptide has endonuclease activity.

15. The non-natural Cas-alpha-10 polypeptide according to claim 13, wherein the Cas-alpha-10 polypeptide is an inactivated Cas-alpha-10 endonuclease complexed with deaminase.

16. The Cas-alpha-10 polypeptide is a non-natural Cas-alpha-10 polypeptide according to claim 13, wherein the Cas-alpha-10 polypeptide has nicasse activity.

17. The non-natural Cas-alpha-10 polypeptide according to claim 13 or claim 16, wherein the Cas-alpha-10 polypeptide is complexed with or operably associated with a reverse transcriptase.

18. The non-natural Cas-alpha-10 polypeptide according to claim 13, wherein the Cas-alpha-10 polypeptide complexes with a heterologous protein domain via a linker, and the heterologous protein domain has methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity.

19. A synthetic composition, (a) A Cas-alpha-10 polypeptide having DNA-binding activity, the non-natural Cas-alpha-10 polypeptide according to any one of claims 1 to 18; and (b) At least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterogeneous to the source of the Cas-alpha-10 polypeptide. A synthetic composition comprising the following: the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, the complex of which binds to the target polynucleotide.

20. A method for editing targeted polynucleotides in cells, (a) Providing a Cas-alpha-10 polypeptide according to any one of claims 1 to 18 to a cell, wherein the Cas-alpha-10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) Providing a cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-alpha-10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) Introducing at least one nucleotide modification into a target polynucleotide via a complex, wherein the target polynucleotide is heterogeneous to the Cas-alpha-10 polypeptide. Methods that include...

21. The method according to claim 20, further comprising providing a donor DNA molecule or a polynucleotide modification template to cells.

22. The method according to claim 20, wherein the cells are derived from or obtained from animals, fungi, or plants.

23. The method according to claim 22, wherein the plant is a monocotyledonous plant or a dicotyledonous plant.

24. The method according to claim 22, wherein the plant is corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, or tomato.

25. An animal, fungus, or cell thereof comprising a non-natural Cas-alpha-10 polypeptide as described in any one of claims 1 to 18.

26. A plant or plant cell comprising a non-natural Cas-alpha-10 polypeptide as described in any one of claims 1 to 18.

27. A method for modifying the protospacer adjacent motif (PAM) specificity of a target Cas-alpha 10 polypeptide, (a) Comparing the PAM interaction (PI) domain of the ortholog Cas-alpha polypeptide with the PI domain of the target Cas-alpha 10 polypeptide; (b) Selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the ortholog Cas-alpha polypeptide; (c) incorporating one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the ortholog Cas-alpha polypeptide into one or more structurally similar positions of the target Cas-alpha 10 polypeptide to obtain a modified target Cas-alpha 10 polypeptide; and (d) Determine the PAM recognition of the modified target Cas-alpha 10 polypeptide. Methods that include...

28. The method according to claim 27, wherein the ortholog Cas-alpha polypeptide has PAM recognition different from that of the target Cas-alpha 10 polypeptide.

29. The method according to claim 27, wherein the Cas-alpha-10 polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

2.

30. The method according to any one of claims 27 to 29, wherein the ortholog Cas-alpha polypeptide is Cas-alpha 1, Cas-alpha 2, Cas-alpha 3, Cas-alpha 4, Cas-alpha 5, Cas-alpha 6, Cas-alpha 7, Cas-alpha 8, or Cas-alpha 11.