Cas polypeptides with altered PAM recognition

By modifying the PAM interaction domain of the Cas-α peptide to change its specificity, a Cas-α peptide with multiple active functions was prepared, which solved the problems of limited targeting range and high cost of the CRISPR-Cas system in genome editing, and realized flexible and efficient genome editing.

CN121152874APending Publication Date: 2025-12-16PIONEER HI BREED INTERNATIONAL INC
View PDF 89 Cites 0 Cited by

Patent Information

Application Number
CN202480033412.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-27
Filing Date
2024-03-19
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems are limited by PAM specificity in genome editing, resulting in limited targeting range and high preparation costs, making it difficult to meet diverse editing needs.

Method used

By modifying the PAM interaction domain of Cas-α peptides to change their specificity, Cas-α peptides with different PAM recognition capabilities, including Cas-α 10 peptides and their combinations, can be prepared and bind to heterologous protein domains to achieve a variety of active functions.

Benefits of technology

This improves the flexibility and efficiency of targeted editing of Cas-α peptides, reduces preparation costs, and expands the application scope of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure relates to methods for altering the PAM specificity of Cas-alpha polypeptides. The present disclosure also relates to Cas-alpha 10 polypeptides having altered PAM specificity, and methods and compositions of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 491,118, filed March 20, 2023, and U.S. Provisional Application No. 63 / 585,659, filed September 27, 2023, which are incorporated herein by reference in their entirety. References to sequence lists submitted electronically

[0002] An official copy of the sequence list is submitted electronically as an XML sequence list, named "108587-WO-SEC-1_Sequence_Listing_ST26", created on March 18, 2024, and is 119 kilobytes in size, and is submitted with this specification. The sequence list contained in this XML file is part of this specification and is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to the field of molecular biology, and more particularly to compositions of novel polynucleotide-guided Cas peptides, and compositions and methods for editing or modifying cellular genomes. Background Technology

[0004] Recombinant DNA technology enables the insertion of DNA sequences and / or modification of specific endogenous chromosomal sequences at target genomic locations. Site-specific integration techniques employing site-specific recombination systems, along with other types of recombination techniques, have been used to generate targeted insertions of target genes in various organisms. Genome editing techniques such as engineered zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or homing broad-spectrum nucleases can be used to generate targeted genomic interference; however, these systems tend to have low specificity and require the use of engineered nucleases that need to be redesigned for each target site, making their preparation costly and time-consuming.

[0005] A newer technique utilizing archaea or the adaptive immune system of bacteria, known as CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats), has been identified. C lustered R egularly I nterspaced S hort P alindromic R epeats contain different domains of effector proteins, which have a variety of activities (DNA recognition, binding, and selective cleavage).

[0006] CRISPR-associated (Cas) peptides' prespacer neighbor motif (PAM) requirement limits their targeting range (Shmakov et al., 2015; Zetsche et al., 2015; Burstein et al., 2017; Karvelis et al., 2020; Pausch et al., 2020). This becomes particularly evident in genome editing applications, where the outcome depends on the proximity of the desired edit to the cleavage site (e.g., template-free editing and homology-directed repair), or on methods that impose additional sequence requirements on target selection (e.g., base editing; Anzalone et al., 2020).

[0007] This document discloses a method for altering the PAM specificity of Cas-α peptides. It also discloses Cas-α 10 peptides with altered PAM specificity, methods of use therein, and compositions thereof. Summary of the Invention

[0008] In a first aspect, this disclosure provides a method for altering the prespacer neighbor motif (PAM) specificity of a target Cas-α polypeptide, the method comprising: (a) comparing a PAM interaction (PI) domain of an orthologous Cas-α polypeptide with a PI domain of the target Cas-α polypeptide, wherein the orthologous Cas-α polypeptide has a different PAM specificity than the target Cas-α polypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-α polypeptide; (c) incorporating one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the orthologous Cas-α polypeptide into one or more structurally similar positions of the target Cas-α polypeptide to produce a modified target Cas-α polypeptide; and (d) determining the PAM recognition of the modified target Cas-α polypeptide.

[0009] In some instances of methods used to alter the PAM specificity of a target Cas-α peptide, the orthologous Cas-α peptide has a different PAM recognition than the target Cas-α peptide.

[0010] In some instances of methods used to alter the PAM specificity of a target Cas-α peptide, the target Cas-α peptide is a Cas-α 10 peptide.

[0011] In some examples of methods for altering the PAM specificity of the target Cas-α polypeptide, the Cas-α 10 polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, and the PI domain comprises amino acids from S63 to I196.

[0012] In some examples of methods used to alter the PAM specificity of target Cas-α peptides, the orthologous Cas-α peptides are Cas-α 1, Cas-α 2, Cas-α 3, Cas-α 4, Cas-α 5, Cas-α 6, Cas-α 7, Cas-α 8, or Cas-α 11.

[0013] In a second aspect, this disclosure provides synthetic or non-naturally occurring Cas-α10 polypeptides comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID Nos. 5-16 or 28-87. In some examples of this second aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide has endonuclease activity. In some examples of this second aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase. In some examples of this second aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain has methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

[0014] In a third aspect, this disclosure provides synthetic or non-naturally occurring Cas-α10 peptides comprising a PAM interaction (PI) domain, wherein the PI domain recognizes PAM sequences comprising: 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'- GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3', or 5'-NNCN-3', where N = A, C, G, or T, Y = T or C, D = G, A, or T, W = A or T, and H = A, T, or C. In some examples of this third aspect, the synthetic or non-naturally occurring Cas-α 10 polypeptides possess endonuclease activity. In some examples of this third aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase. In some examples of this third aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this third aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide possesses nickase activity. In yet another further example of this third aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide complexes with or operatively associates with a reverse transcriptase.

[0015] In a fourth aspect, this disclosure provides synthetic or non-naturally occurring Cas-α10 polypeptides comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence contains mutations at amino acid positions relative to SEQ ID NO: 2, wherein the mutations include K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92 F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or one or more of the following mutation combinations: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; K85S mutation and N92Q mutation. Combinations of K85S and N92H mutations; combinations of K85S and N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; combinations of K85S, N88D, and Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations. Combinations of the following mutations: Y72A, N88D, Q89D, and Q125R; K85S, N88D, Q89G, and N92L; Y72A, N88D, Q89G, and Q125R; K85S, N92L, and Q125R; Y72S, K85D, Q125R, and N127R; Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R.The mutations are: Y72C, N88D, Q89D, N92W, and Q125R; or K85Q and N92W. In a specific example of this fourth aspect, the amino acid sequence of the Cas-α10 polypeptide contains the K85S mutation. In other examples of this fourth aspect, the amino acid sequence of the Cas-α10 polypeptide contains a combination of the K85N and N92L mutations. In some examples of this fourth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide possesses endonuclease activity. In some examples of this fourth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase. In some examples of this fourth aspect, a synthetic or non-naturally occurring Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this fourth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide possesses nickase activity. In yet another further example of this fourth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide complexes with or operatively associates with reverse transcriptase.

[0016] In a fifth aspect, this disclosure provides synthetic compositions comprising: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID Nos. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide. In some examples of the synthetic compositions, the Cas-α10 polypeptide has endonuclease activity. In some examples of the synthetic compositions, the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase. In some examples of the synthesized compositions, the Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this fifth aspect, the Cas-α10 polypeptide possesses nickase activity. In yet another further example of this fifth aspect, the Cas-α10 polypeptide complexes with or operably associates with reverse transcriptase.

[0017] In a sixth aspect, this disclosure provides a synthetic composition comprising: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising a PAM interaction (PI) domain, wherein the PI domain recognizes a PAM sequence on a target polynucleotide, wherein the PAM sequence comprises 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3'. ', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G, or T, Y = T or C, D = G, A, or T, W = A or T, and H = A, T, or C; and (b) at least one guiding polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the source of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide forms a complex with the at least one guiding polynucleotide, and wherein the complex binds to the target polynucleotide. In some examples of the synthesized compositions, the Cas-α10 polypeptide has endonuclease activity. In some examples of the synthesized compositions, the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase. In some examples of the synthesized compositions, the Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, and wherein the heterologous protein domain has methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this sixth aspect, the Cas-α10 polypeptide possesses nicking enzyme activity. In yet another further example of this sixth aspect, the Cas-α10 polypeptide is complexed with or operatively associated with reverse transcriptase.

[0018] In a seventh aspect, this disclosure provides a synthetic composition comprising: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence is relative to SEQ ID NO: The amino acid position of 2 contains mutations, including the following mutations: K85S; K85A; N92R; N92K; Q125K; N88H; N88K; N88Q; Q125F; Y72V; Y72E; Y72Q; Y72T; Y72C; Y72A; Y72S; Y72P; Y72G; Y72D; Y72L; K85G; K85D; K85N; N88D; Q89D; N92D; N92C; N92A; N9 2G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or one or more of the following mutation combinations: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; N92Q mutation and N92Q mutation. Combinations of 2Q mutations; combinations of K85S and N92C mutations; combinations of K85S and N92H mutations; combinations of K85S and N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; combinations of K85S, N88D, and Q89G mutations; K85 Combinations of S mutation, N88H mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89D mutation, and Q125R mutation; combinations of K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; combinations of K85S mutation, N92L mutation, and Q125R mutation; combinations of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; combinations of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation;A combination of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; a combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or a combination of K85Q and N92W mutations; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0019] In a specific example of this seventh aspect, the amino acid sequence of the Cas-α 10 polypeptide contains a K85S mutation. In other examples of this seventh aspect, the amino acid sequence of the Cas-α 10 polypeptide contains a combination of K85N and N92L mutations. In some examples of the synthetic compositions, the Cas-α 10 polypeptide possesses endonuclease activity. In some examples of the synthetic compositions, the Cas-α 10 polypeptide is an inactivated Cas-α 10 endonuclease complexed with a deaminase. In some examples of the synthetic compositions, the Cas-α 10 polypeptide complexes with a heterologous protein domain via a linker, and wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of this seventh aspect, the Cas-α 10 polypeptide possesses nickase activity. In a further example of this seventh aspect, the Cas-α10 polypeptide is complexed with or operatively associated with reverse transcriptase.

[0020] In an eighth aspect, this disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing the cell with a Cas-α10 polypeptide having at least 90% sequence identity with any one of SEQ ID No. 5-16 or 28-87, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0021] In some instances of the eighth aspect, the PAM sequences recognized by the Cas-α10 peptide include 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', and 5'-DTTY. -3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

[0022] In some instances of the eighth aspect, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0023] In some instances of aspect eight, the cells are derived from or obtained from animals, fungi, or plants. In some instances of aspect eight, the plants are dicotyledons or monocotyledons. In some aspects, the plants are corn, soybeans, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanuts, potatoes, Arabidopsis, safflower, or tomatoes.

[0024] In some examples of the eighth aspect, the Cas-α10 polypeptide exhibits endonuclease activity.

[0025] In some instances of the eighth aspect, at least one guiding polynucleotide comprises multiple guiding polynucleotides, and the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase.

[0026] In some examples of the eighth aspect, the Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of the eighth aspect, the Cas-α10 polypeptide possesses nickase activity. In yet another further example of the eighth aspect, the Cas-α10 polypeptide complexes with or operatively associates with reverse transcriptase.

[0027] In a ninth aspect, this disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing the cell with a Cas-α10 polypeptide comprising a PAM interaction (PI) domain, wherein the PI domain recognizes a PAM sequence on the target polynucleotide, and wherein the PAM sequence comprises 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-G TTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'- TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0028] In some instances of the ninth aspect, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0029] In some instances of the ninth aspect, the cells are derived from or obtained from animals, fungi, or plants. In some instances, the plants are dicotyledonous or monocotyledonous plants. In some instances, the plants are corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, or tomato.

[0030] In some examples of the ninth aspect, the Cas-α10 polypeptide exhibits endonuclease activity.

[0031] In some instances of the ninth aspect, at least one guiding polynucleotide comprises multiple guiding polynucleotides, and the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase.

[0032] In some examples of the ninth aspect, the Cas-α 10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of the ninth aspect, the Cas-α 10 polypeptide possesses nickase activity. In yet another further example of the ninth aspect, the Cas-α 10 polypeptide complexes with or operatively associates with reverse transcriptase.

[0033] In a tenth aspect, this disclosure provides a method for editing target polynucleotides in cells, the method comprising: (a) providing the cell with a Cas-α10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence contains a mutation at an amino acid position relative to SEQ ID NO: 2, wherein the mutation comprises K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N9 2G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; or one or more of the following mutation combinations: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; N92Q mutation and N92Q mutation. Combinations of 2Q mutations; combinations of K85S and N92C mutations; combinations of K85S and N92H mutations; combinations of K85S and N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; combinations of K85S, N88D, and Q89G mutations; K85 Combinations of S mutation, N88H mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89D mutation, and Q125R mutation; combinations of K85S mutation, N88D mutation, Q89G mutation, and N92L mutation; combinations of Y72A mutation, N88D mutation, Q89G mutation, and Q125R mutation; combinations of K85S mutation, N92L mutation, and Q125R mutation; combinations of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; combinations of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation;A combination of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; a combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or a combination of K85Q and N92W mutations, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide.

[0034] In some examples of this tenth aspect, the amino acid sequence of the Cas-α 10 polypeptide contains a K85S mutation. In other examples of this tenth aspect, the amino acid sequence of the Cas-α 10 polypeptide contains a combination of K85N and N92L mutations.

[0035] In some instances of the tenth aspect, the PAM sequence recognized by the Cas-α 10 peptide includes 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', and 5'-DTTY. -3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

[0036] In some instances of the tenth aspect, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0037] In some instances of aspect ten, the cells are derived from or obtained from animals, fungi, or plants. In some aspects, the plants are dicotyledons or monocotyledons. In some instances, the plants are corn, soybeans, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanuts, potatoes, Arabidopsis, safflower, or tomatoes.

[0038] In some examples of the tenth aspect, the Cas-α10 polypeptide exhibits endonuclease activity.

[0039] In some instances of the tenth aspect, at least one guiding polynucleotide comprises multiple guiding polynucleotides, and the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase.

[0040] In some examples of the tenth aspect, the Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, wherein the heterologous protein domain possesses methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity. In a further example of the tenth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide possesses nickase activity. In yet another further example of the tenth aspect, the synthetic or non-naturally occurring Cas-α10 polypeptide complexes with or operatively associates with a reverse transcriptase.

[0041] In an eleventh aspect, this disclosure provides animal cells or fungal cells containing any one of the synthetic or non-naturally occurring Cas-α10 polypeptides as described herein.

[0042] In a twelfth aspect, this disclosure provides plant cells, plant parts, plantlets or plants comprising any of the synthetic or non-naturally occurring Cas-α 10 polypeptides as described herein.

[0043] In a thirteenth aspect, this disclosure provides a method for altering the prespacer neighbor motif (PAM) specificity of a target Cas-α 10 polypeptide, the method comprising: (a) comparing a PAM interaction (PI) domain of an orthologous Cas-α polypeptide with a PI domain of the target Cas-α 10 polypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-α polypeptide; (c) incorporating one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the orthologous Cas-α polypeptide into one or more structurally similar positions of the target Cas-α 10 polypeptide, thereby producing a modified target Cas-α 10 polypeptide; and (d) determining the PAM recognition of the modified target Cas-α 10 polypeptide. In some examples of this thirteenth aspect, the orthologous Cas-α polypeptide has a different PAM recognition than the target Cas-α 10 polypeptide. In some examples of this thirteenth aspect, the Cas-α 10 polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2. In a further example of this thirteenth example, the orthologous Cas-α polypeptide is Cas-α 1, Cas-α 2, Cas-α 3, Cas-α 4, Cas-α 5, Cas-α 6, Cas-α 7, Cas-α 8 or Cas-α 11. Description of the attached figures and sequence listing

[0044] This disclosure will be more fully understood from the following detailed description, accompanying drawings, and sequence listing, which form part of this application.

[0045] Figure 1Phylogenetic relationships among some Cas-α orthologs were elucidated. Three supergroups (I, II, and III) were identified. Group I includes clade 1 (candidate archaea and Aureabacteria (typically encoding Cas1, Cas2, and Cas4 at the locus)). Group II includes clade 2 (Aquagenic Bacteria (Sulfurihydrogenibium and Hydrogenivirga) and Deltaproteobacteria (Desulfovibrio)), clade 3 (candidate archaea (typically encoding Cas1, Cas2, and Cas4 at the locus)), clade 4 (Bacteroidetes (Prevotella and Bacteroides))), clade 5 (candidate Levybacterium), and clade 6 (Clostridia (Dorea, Ruminococcus, Clostridium, Clostridioides, Peptocolstridium, Cellulosilyticym, Eubacteria)). Group III includes clade 7 (Bacilli, Acidibacillus, Aneurinibacillus, Brevibacillus, Parageobacillus, Alicyclobacillus), clade 8 (Negativicutes, Phascolarctobacterium), and clade 9 (Flavobacteriia, Flavobacterium). The diamond symbol represents Cas-α endonucleases 1-11, whose PAM recognition is described in US10934536.

[0046] Figure 2 This describes expression cassettes for expressing Cas-α10 as described herein, or synthetic or non-naturally occurring Cas-α10 peptides.

[0047] Figure 3 The method for altering the PAM specificity in Cas-α peptides is described.

[0048] Figure 4This describes another method for altering the PAM specificity in Cas-α peptides.

[0049] Figure 5 This illustrates yet another method for altering the PAM specificity in Cas-α peptides.

[0050] Figure 6 The figure illustrates the number of target sites in the corn and human genomes targeted by several PAM variants of Example 2.

[0051] SEQ ID NO: 1 is the PRT sequence of the Cas-α10 polypeptide from Syntrophomonas palmitatica.

[0052] SEQ ID NO: 2 is the PRT sequence of a first exemplary synthetic or non-naturally occurring Cas-α 10 polypeptide.

[0053] SEQ ID NO: 3 is the PRT sequence of the Cas-α4 polypeptide from uncultured archaea.

[0054] SEQ ID NO: 4 is the PRT sequence of the Cas-α 8 polypeptide from Acidibacillus sulfuroxidans.

[0055] SEQ ID NO: 5 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the K85A mutation relative to SEQ ID NO: 2.

[0056] SEQ ID NO: 6 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having a K85S mutation relative to SEQ ID NO: 2.

[0057] SEQ ID NO: 7 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92R mutation relative to SEQ ID NO: 2.

[0058] SEQ ID NO: 8 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having an N92K mutation relative to SEQ ID NO: 2.

[0059] SEQ ID NO: 9 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having a Q125K mutation relative to SEQ ID NO: 2.

[0060] SEQ ID NO: 10 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N88H mutation relative to SEQ ID NO: 2.

[0061] SEQ ID NO: 11 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having N88H and Q89G mutations relative to SEQ ID NO: 2.

[0062] SEQ ID NO: 12 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having an N88K mutation relative to SEQ ID NO: 2.

[0063] SEQ ID NO: 13 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having N88K and Q89G mutations relative to SEQ ID NO: 2.

[0064] SEQ ID NO: 14 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N88Q mutation relative to SEQ ID NO: 2.

[0065] SEQ ID NO: 15 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having N88Q and Q89G mutations relative to SEQ ID NO: 2.

[0066] SEQ ID NO: 16 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Q125F mutation relative to SEQ ID NO: 2.

[0067] SEQ ID NO: 17 is a PRT sequence Cas-α1 polypeptide from the archaea Candidatus Micrarchaeota.

[0068] SEQ ID NO: 18 is a PRT sequence Cas-α2 polypeptide from the archaea Candidatus Micrarchaeota.

[0069] SEQ ID NO: 19 is a PRT sequence Cas-α3 polypeptide from the bacterium Candidatus Aureabacteria.

[0070] SEQ ID NO: 20 is a PRT sequence Cas-α5 polypeptide from the archaea Candidatus Micrarchaeota.

[0071] SEQ ID NO: 21 is a PRT sequence of Cas-α6 polypeptide from an uncultured archaea.

[0072] SEQ ID NO: 22 is a PRT sequence Cas-α7 polypeptide from Parageobacillus thermoglucosidasius.

[0073] SEQ ID NO: 23 is a PRT sequence Cas-α 9 polypeptide from a species of Ruminococcus.

[0074] SEQ ID NO: 24 is a PRT sequence Cas-α 11 polypeptide from Clostridium novyi.

[0075] SEQ ID NO: 25 is a PRT sequence Cas-α 13 polypeptide from Clostridium paraputrificum.

[0076] SEQ ID NO: 26 is a PRT sequence Cas-α 24 polypeptide from Bacillus toyonensis.

[0077] SEQ ID NO: 27 is a PRT sequence Cas-α 29 polypeptide from a species of the genus Peptoclostridium.

[0078] SEQ ID NO: 28 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72V mutation relative to SEQ ID NO: 2.

[0079] SEQ ID NO: 29 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72E mutation relative to SEQ ID NO: 2.

[0080] SEQ ID NO: 30 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72Q mutation relative to SEQ ID NO: 2.

[0081] SEQ ID NO: 31 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72T mutation relative to SEQ ID NO: 2.

[0082] SEQ ID NO: 32 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72C mutation relative to SEQ ID NO: 2.

[0083] SEQ ID NO: 33 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72A mutation relative to SEQ ID NO: 2.

[0084] SEQ ID NO: 34 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72S mutation relative to SEQ ID NO: 2.

[0085] SEQ ID NO: 35 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72P mutation relative to SEQ ID NO: 2.

[0086] SEQ ID NO: 36 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72G mutation relative to SEQ ID NO: 2.

[0087] SEQ ID NO: 37 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72D mutation relative to SEQ ID NO: 2.

[0088] SEQ ID NO: 38 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Y72L mutation relative to SEQ ID NO: 2.

[0089] SEQ ID NO: 39 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the K85G mutation relative to SEQ ID NO: 2.

[0090] SEQ ID NO: 40 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the K85D mutation relative to SEQ ID NO: 2.

[0091] SEQ ID NO: 41 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the K85N mutation relative to SEQ ID NO: 2.

[0092] SEQ ID NO: 42 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N88D mutation relative to SEQ ID NO: 2.

[0093] SEQ ID NO: 43 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Q89D mutation relative to SEQ ID NO: 2.

[0094] SEQ ID NO: 44 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92D mutation relative to SEQ ID NO: 2.

[0095] SEQ ID NO: 45 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92C mutation relative to SEQ ID NO: 2.

[0096] SEQ ID NO: 46 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92A mutation relative to SEQ ID NO: 2.

[0097] SEQ ID NO: 47 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92G mutation relative to SEQ ID NO: 2.

[0098] SEQ ID NO: 48 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92F mutation relative to SEQ ID NO: 2.

[0099] SEQ ID NO: 49 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92E mutation relative to SEQ ID NO: 2.

[0100] SEQ ID NO: 50 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92H mutation relative to SEQ ID NO: 2.

[0101] SEQ ID NO: 51 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92I mutation relative to SEQ ID NO: 2.

[0102] SEQ ID NO: 52 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92L mutation relative to SEQ ID NO: 2.

[0103] SEQ ID NO: 53 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92V mutation relative to SEQ ID NO: 2.

[0104] SEQ ID NO: 54 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92W mutation relative to SEQ ID NO: 2.

[0105] SEQ ID NO: 55 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92Q mutation relative to SEQ ID NO: 2.

[0106] SEQ ID NO: 56 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92S mutation relative to SEQ ID NO: 2.

[0107] SEQ ID NO: 57 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92P mutation relative to SEQ ID NO: 2.

[0108] SEQ ID NO: 58 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92Y mutation relative to SEQ ID NO: 2.

[0109] SEQ ID NO: 59 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92T mutation relative to SEQ ID NO: 2.

[0110] SEQ ID NO: 60 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the N92M mutation relative to SEQ ID NO: 2.

[0111] SEQ ID NO: 61 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Q125R mutation relative to SEQ ID NO: 2.

[0112] SEQ ID NO: 62 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having the Q125P mutation relative to SEQ ID NO: 2.

[0113] SEQ ID NO: 63 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92L mutations relative to SEQ ID NO: 2.

[0114] SEQ ID NO: 64 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92Q mutations relative to SEQ ID NO: 2.

[0115] SEQ ID NO: 65 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92C mutations relative to SEQ ID NO: 2.

[0116] SEQ ID NO: 66 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92H mutations relative to SEQ ID NO: 2.

[0117] SEQ ID NO: 67 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92A mutations relative to SEQ ID NO: 2.

[0118] SEQ ID NO: 68 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S and N92M mutations relative to SEQ ID NO: 2.

[0119] SEQ ID NO: 69 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92L mutations relative to SEQ ID NO: 2.

[0120] SEQ ID NO: 70 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92H mutations relative to SEQ ID NO: 2.

[0121] SEQ ID NO: 71 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92A mutations relative to SEQ ID NO: 2.

[0122] SEQ ID NO: 72 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92C mutations relative to SEQ ID NO: 2.

[0123] SEQ ID NO: 73 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92M mutations relative to SEQ ID NO: 2.

[0124] SEQ ID NO: 74 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92Q mutations relative to SEQ ID NO: 2.

[0125] SEQ ID NO: 75 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85N and N92I mutations relative to SEQ ID NO: 2.

[0126] SEQ ID NO: 76 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S, N88D and Q89G mutations relative to SEQ ID NO: 2.

[0127] SEQ ID NO: 77 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S, N88H, Q89G and N92L mutations relative to SEQ ID NO: 2.

[0128] SEQ ID NO: 78 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72A, N88D, Q89D, and Q125R relative to SEQ ID NO: 2.

[0129] SEQ ID NO: 79 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S, N88D, Q89G and N92L mutations relative to SEQ ID NO: 2.

[0130] SEQ ID NO: 80 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72A, N88D, Q89G, and Q125R relative to SEQ ID NO: 2.

[0131] SEQ ID NO: 81 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85S, N92L and Q125R mutations relative to SEQ ID NO: 2.

[0132] SEQ ID NO: 82 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72S, K85D, Q125R and N127R relative to SEQ ID NO: 2.

[0133] SEQ ID NO: 83 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72A, K85S, Q89D, N92L and Q125R relative to SEQ ID NO: 2.

[0134] SEQ ID NO: 84 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72C, N88H, Q89G, and Q125R relative to SEQ ID NO: 2.

[0135] SEQ ID NO: 85 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having mutations of Y72C, N88D, Q89D, N92W, and Q125R relative to SEQ ID NO: 2.

[0136] SEQ ID NO: 86 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85Q and N92L mutations relative to SEQ ID NO: 2.

[0137] SEQ ID NO: 87 is the PRT sequence of a synthetic or non-naturally occurring Cas-α 10 polypeptide having K85Q and N92W mutations relative to SEQ ID NO: 2.

[0138] SEQ ID NO: 88 is the PRT sequence of a second exemplary synthetic or non-naturally occurring Cas-α 10 polypeptide.

[0139] SEQ ID NO: 89 is the PRT sequence of a third exemplary synthetic or non-naturally occurring Cas-α 10 polypeptide. Detailed Implementation

[0140] Compositions and methods are provided for novel CRISPR effector systems and elements comprising such systems, including, but not limited to, novel guiding polynucleotide / endonuclease complexes, guiding polynucleotides, guiding RNA elements, Cas peptides, and endonucleases, as well as proteins comprising endonuclease functional domains. Compositions and methods for directly delivering endonucleases, cleavage-ready complexes, guiding RNA, and guiding RNA / Cas peptide complexes are also provided. This disclosure further includes compositions and methods for genomic modification of target sequences in the cellular genome, for gene editing, and for inserting target polynucleotides into the cellular genome.

[0141] Unless otherwise specified, the terms used in the claims and description are defined as set forth below. It should be noted that, unless the context clearly indicates otherwise, the singular forms “a / an” and “the” as used in this specification and the appended claims include plural indicators.

[0142] definition

[0143] As used herein, “nucleic acid” means polynucleotide and includes single-stranded or double-stranded polymers comprising deoxyribonucleotide or ribonucleotide bases. Nucleic acids may also include fragments and modified nucleotides. Therefore, the terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” and “nucleic acid fragment” are used interchangeably to refer to single-stranded or double-stranded RNA and / or DNA and / or RNA-DNA polymers, optionally including synthetic, non-natural, or altered nucleotide bases. Nucleotides (usually found in their 5'-monophosphate form) are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (for RNA or DNA, respectively), “C” for cytosine or deoxycytosine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.

[0144] The term "genome," when applied to prokaryotic or eukaryotic cells or somatic cells, encompasses not only chromosomal DNA found in the cell nucleus but also organelle DNA found in subcellular components of the cell, such as mitochondria or plastids.

[0145] The abbreviation for "readable box" is ORF.

[0146] The term "selective hybridization" refers to the hybridization of a nucleic acid sequence to a specific target nucleic acid sequence under stringent hybridization conditions, where the hybridization is detectably greater (e.g., at least twice the background) than the hybridization of the same nucleic acid sequence to a non-target nucleic acid sequence, and substantially excludes non-target nucleic acids. Selectively hybridized sequences typically have at least 80% sequence identity, or 90% sequence identity, up to and including 100% sequence identity (i.e., complete complementarity) with each other.

[0147] The term "strict conditions" or "strict hybridization conditions" refers to conditions under which a probe will selectively hybridize with its target sequence in an in vitro hybridization assay. Strict conditions are sequence-dependent and will vary under different conditions. By controlling the strictness of hybridization and / or washing conditions, target sequences that are 100% complementary to the probe can be identified (homologous detection). Alternatively, strict conditions can be tuned to allow some mismatches in the sequence for detection of a lower degree of similarity (heterologous detection). Typically, the probe length is less than about 1000 nucleotides, optionally less than 500 nucleotides. Typically, strict conditions will be the following: a salt concentration of less than about 1.5 M Na ions, typically about 0.01 M to 1.0 M Na ion concentration (or one or more other salts), at pH 7.0 to 8.3, and at least about 30°C for short probes (e.g., 10 to 50 nucleotides) and at least about 60°C for long probes (e.g., more than 50 nucleotides). Strict conditions can also be achieved by adding a destabilizing agent such as formamide. Exemplary low-toughness conditions include hybridization at 37°C with a buffer solution of 30% to 35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), followed by washing at 50°C to 55°C in 1X to 2X SSC (20X SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Exemplary medium-toughness conditions include hybridization at 37°C with 40% to 45% formamide, 1 M NaCl, and 1% SDS, followed by washing at 55°C to 60°C in 0.5X to 1X SSC. Exemplary high-toughness conditions include hybridization at 37°C with 50% formamide, 1 M NaCl, and 1% SDS, followed by washing at 60°C to 65°C in 0.1X SSC.

[0148] "Homologous" means that the DNA sequences are similar. For example, a "region homologous to a genomic region" found on donor DNA is a region of DNA that has a similar sequence to a given "genomic sequence" in the genome of a cell or organism. Homologous regions can have any length sufficient to promote homologous recombination at the cleavage target site. For example, the length of a homologous region can include at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5- 1200, 5-1300, 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases are allowed to ensure that homologous regions have sufficient homology to undergo homologous recombination with corresponding genomic regions. "Sufficient homology" means that two polynucleotide sequences have enough structural similarity to serve as substrates for homologous recombination reactions. Structural similarity includes the total length of each polynucleotide fragment and the sequence similarity of the polynucleotides. Sequence similarity can be described by percentage sequence identity over the entire length of the sequence and / or by conserved regions containing local similarity (e.g., consecutive nucleotides with 100% sequence identity) and percentage sequence identity over a portion of the sequence length.

[0149] As used herein, a "genomic region" is a chromosomal segment of the cell's genome that is located on either side of the target site, or alternatively, also contains a portion of the target site. A genomic region may contain at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, or 5-120 chromosomal regions. 0, 5-1300, 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases, thus ensuring sufficient homology for genomic regions to undergo homologous recombination with corresponding homologous regions.

[0150] As used in this article, “homologous recombination (HR)” refers to the exchange of DNA fragments between two DNA molecules at homologous sites. The frequency of homologous recombination is influenced by several factors. Different organisms vary in the amount of homologous recombination and the relative ratio of homologous to non-homologous recombination. Generally, the length of the homologous region affects the frequency of homologous recombination events: the longer the homologous region, the higher the frequency. The length of the homologous region required to observe homologous recombination also varies from species to species. In many cases, homology of at least 5 kb has been utilized, but homologous recombination with homology as low as 25–50 bp has been observed. See, for example, Singer et al., (1982) Cell 31:25-33; Shen and Huang, (1986) Genetics 112:441-57; Watt et al., (1985) Proc. Natl. Acad. Sci. USA 82:4768-72; Sugawara and Haber, (1992) Mol Cell Biol 12:563-75; Rubnitz and Subramani, (1984) Mol Cell Biol 4:2253-8; Ayares et al., (1986) Proc. Natl. Acad. Sci. USA 83:5199-203; Liskay et al., (1987) Genetics 115:161-7.

[0151] In the context of nucleic acid or polypeptide sequences, "sequence identity" or "identity" means that the nucleic acid bases or amino acid residues in two sequences are identical when compared against the maximum correspondence in a specified comparison window.

[0152] The term "percentage of sequence identity" refers to a value determined by comparing two best-aligned sequences within a comparison window, where the portion of a polynucleotide or polypeptide sequence within the comparison window may contain additions or deletions (i.e., vacancies) compared to a reference sequence (which contains no additions or deletions). This percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue appears to produce a number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and then multiplying these results by 100 to produce the percentage of sequence identity. Useful examples of percentage sequence identity include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage from 50% to 100%. These identities can be determined using any of the procedures described herein.

[0153] Sequence alignment and percentage identity or similarity calculations can be determined using a variety of comparison methods designed for detecting homologous sequences, including but not limited to the MegAlign™ program of the LASERGENE bioinformatics computing package (DNASTAR Inc., Madison, Wisconsin). In the context of this application, it should be understood that when using sequence analysis software for analysis, the results will be based on the “default values” of the referenced program, unless otherwise stated. As used herein, “default values” will mean any set of values ​​or parameters initially loaded when the software is first initialized.

[0154] The “Clustal V method for alignment” corresponds to the alignment method labeled Clustal V (described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992) Comput Appl Biosci [Computer Applications in Biological Sciences] 8:189-191), and is found in the MegAlign™ program of the LASERGENE Bioinformatics Computing Package (DNASTAR, Madison, Wisconsin). For multiple alignments, the default values ​​correspond to gap penalty = 10 and gap length penalty = 10. The default parameters for performing side-by-side alignments and calculating percentage identity of protein sequences using the Clustal method are KTUPLE = 1, gap penalty = 3, window = 5, and stored diagonals = 5. For nucleic acids, these parameters are KTUPLE=2, void penalty=5, window=4, and storage diagonal=4. After aligning sequences using the Clustal V program, the percentage identity can be obtained by looking at the "Sequence Distance" table in the same program. "Alignment by Clustal W method" corresponds to the alignment method labeled Clustal W (described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992) Comput Appl Biosci [Computer Applications in Biological Sciences] 8:189-191) and is found in the MegAlign™ v6.1 program of the LASERGENE Bioinformatics Computing Package (DNASTAR, Madison, Wisconsin). The default parameters for multiple alignments are: gap penalty = 10, gap length penalty = 0.2, delay divergence sequences (%) = 30, DNA conversion weight = 0.5, protein weight matrix = Gonnet series, DNA weight matrix = IUB. After aligning sequences using the Clustal W program, the percentage identity can be obtained by checking the "Sequence Distance" table in the same program.Unless otherwise stated, the sequence identity / similarity values ​​provided herein refer to values ​​obtained using GAP version 10 (GCG, Accelrys, San Diego, CA) with the following parameters: % identity and % similarity of nucleotide sequences weighted by a 50-void penalty and a 3-void length extension penalty, and an nwsgapdna.cmp scoring matrix; % identity and % similarity of amino acid sequences weighted by an 8-void penalty and a 2-void length extension penalty, and a BLOSUM62 scoring matrix (Henikoff and Henikoff, (1989) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 89:10915). GAP uses the algorithm of Needleman and Wunsch (1970) J Mol Biol [Journal of Molecular Biology] 48:443-53 to find alignments of two complete sequences that maximize the number of matches and minimize the number of voids. GAP considers all possible alignments and vacancy positions, and uses penalties and vacancy extension penalties in units of matching bases to produce alignments with the maximum number of matching bases and the minimum number of vacancy positions. BLAST is a search algorithm provided by the National Center for Biotechnology Information (NCBI) for finding regions of similarity between biological sequences. This program compares nucleotide or protein sequences to a sequence database and calculates the statistical significance of matches to identify sequences with sufficient similarity to the query sequence so that the similarity is not predicted as having occurred randomly. BLAST reports the identified sequences and their local alignments to the query sequence. Those skilled in the art will readily understand that many levels of sequence identity are useful in identifying peptides or modified natural or synthetic peptides from other species that have the same or similar functions or activities. Useful examples of percentage identity include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage from 50% to 100%.In fact, in describing this disclosure, any amino acid identity from 50% to 100% may be useful, such as 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0155] Polynucleotide and polypeptide sequences, their variants, and the structural relationships of these sequences can be described using the terms “homology,” “homologous,” “substantially identical,” “substantially similar,” and “substantially corresponding,” which are used interchangeably herein. These refer to polypeptide or nucleic acid sequences in which changes at one or more amino acids or nucleotide bases do not affect the function of the molecule, such as its ability to mediate gene expression or produce a particular phenotype. These terms also refer to one or more modifications of a nucleic acid sequence that substantially do not alter the functional properties of the resulting nucleic acid relative to the initially unmodified nucleic acid. These modifications include the deletion, substitution, and / or insertion of one or more nucleotides in the nucleic acid fragment. Substantially similar nucleic acid sequences covered can be defined by their ability to hybridize with sequences exemplified herein, or with any portion of a nucleotide sequence disclosed herein and functionally equivalent to any nucleic acid sequence disclosed herein (under moderately stringent conditions, e.g., 0.5X SSC, 0.1% SDS, 60°C). Stringent conditions can be adjusted to screen for moderately similar fragments (such as homologous sequences from distantly related organisms) to highly similar fragments (such as genes replicating functional enzymes from closely related organisms). The washing process after hybridization dictates strict conditions.

[0156] A centimeter (cM) or map distance unit is the distance between two polynucleotide sequences, linked genes, markers, target sites, loci, or any pair thereof, where 1% of the meiotic product is recombination. Therefore, one centimeter is equivalent to the distance equal to 1% of the average recombination frequency between two linked genes, markers, target sites, loci, or any pair thereof.

[0157] "Isolated" or "purified" nucleic acid molecules, polynucleotides, polypeptides, or proteins, or their biologically active portions, are substantially or essentially free of components that normally accompany or interact with polynucleotides or proteins as found in their natural environment. Therefore, isolated or purified polynucleotides, polypeptides, or proteins are substantially free of other cellular material or culture medium when produced by recombinant technology, or substantially free of chemical precursors or other chemicals when chemically synthesized. Preferably, "isolated" polynucleotides do not contain sequences naturally flanking the polynucleotide in the genomic DNA of the organism from which the polynucleotide is derived (i.e., sequences located at the 5' and 3' ends of the polynucleotide) (preferably protein-coding sequences). For example, in various respects, isolated polynucleotides may contain nucleotide sequences smaller than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb, which are naturally located flanking the polynucleotide in the genomic DNA of the cell from which the polynucleotide is derived. Isolated polynucleotides can be purified from the cells in which they are naturally present. Conventional nucleic acid purification methods known to those skilled in the art can be used to obtain isolated polynucleotides. The term also covers recombinant polynucleotides and chemically synthesized polynucleotides.

[0158] The term "fragment" refers to a continuous set of nucleotides or amino acids. On one hand, a fragment is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more consecutive nucleotides. On the other hand, a fragment is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more consecutive amino acids. A fragment may or may not exhibit the function of a sequence that has a certain percentage of identity in length.

[0159] The terms “functionally equivalent fragment” and “functionally equivalent fragment” are used interchangeably herein. These terms refer to a portion or subsequence of an isolated nucleic acid fragment or polypeptide that exhibits the same activity or function as the longer sequence from which it is derived. In one instance, regardless of whether the fragment encodes an active protein, the fragment retains the ability to alter gene expression or produce a certain phenotype. For example, the fragment can be used to design genes to produce a desired phenotype in modified plants. Genes can be designed for use in repression by linking nucleic acid fragments with a sense or antisense orientation relative to the plant promoter sequence, regardless of whether they encode an active enzyme.

[0160] A “gene” includes a segment of nucleic acid that expresses a functional molecule (such as, but not limited to, a specific protein), which includes a regulatory sequence preceding the coding sequence (5' non-coding sequence) and a regulatory sequence following it (3' non-coding sequence). A “natural gene” is a gene that has its own regulatory sequence found in its natural endogenous location.

[0161] The term "endogenous" refers to sequences or other molecules that are naturally present in cells or organisms. In some respects, endogenous polynucleotides are typically found in the genome of a cell; that is, they are not heterologous.

[0162] An "allele" is one of several alternative forms of a gene that occupies a given locus on a chromosome. A plant is homozygous at a given locus when all alleles present at that locus are identical. A plant is heterozygous at a given locus if the alleles present at that locus are different.

[0163] A “coding sequence” refers to a polynucleotide sequence that encodes a specific amino acid sequence. A “regulatory sequence” refers to a nucleotide sequence located upstream (5' non-coding), inside, or downstream (3' non-coding) of a coding sequence, and that affects the transcription, RNA processing or stability, or translation of the related coding sequence. Regulatory sequences include, but are not limited to: promoters, pretranslational sequences, 5' non-translated sequences, 3' non-translated sequences, introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0164] A "mutated gene" is a gene that has been altered through human intervention. Such a "mutated gene" has a sequence that differs from the sequence of a corresponding non-mutated gene by at least one nucleotide addition, deletion, or substitution. In some aspects of this disclosure, the mutated gene contains alterations caused by a guiding polynucleotide / Cas polypeptide system as disclosed herein. A mutated plant is a plant that contains a mutated gene.

[0165] As used herein, a “targeted mutation” is a mutation in a gene (called a target gene) (including natural genes) that is produced by altering a target sequence within a target gene using any method known to those skilled in the art, including methods involving a guided Cas peptide system as disclosed herein.

[0166] The terms “knockout,” “gene knock-out,” and “genetic knock-out” are used interchangeably in this document. Knockout means that the DNA sequence of a cell has been partially or completely invalidated by targeting with a Cas peptide; for example, the DNA sequence may have encoded an amino acid sequence or may have had a regulatory function (e.g., a promoter) before the knockout.

[0167] The terms “knock-in,” “gene knock-in,” “gene insertion,” and “geneticknock-in” are used interchangeably in this document. Knock-in refers to the substitution or insertion of a DNA sequence at a specific DNA sequence in a cell through targeted use of a Cas peptide (e.g., via homologous recombination (HR), in which a suitable donor DNA polynucleotide is also used). Examples of knock-in are the specific insertion of a heterologous amino acid coding sequence into a gene coding region, or the specific insertion of a transcriptional regulatory element into a genetic locus.

[0168] "Domain" refers to a continuous extension of a nucleotide (which can be RNA, DNA, and / or RNA-DNA combination sequences) or amino acid.

[0169] The term "conserved domain" or "motif" refers to a group of polynucleotides or amino acids that are conserved at a specific position along the aligned sequence of an evolutionarily related protein. While amino acids can vary at other positions between homologous proteins, highly conserved amino acids at specific positions indicate that they are essential for the protein's structure, stability, or activity. Because they are identified by their high conservation in the aligned sequences of protein homologues, they can be used as identifiers or "characteristics" to determine whether a protein with a newly identified sequence belongs to a previously identified protein family.

[0170] "Codon-modified genes," "codon-preferred genes," or "codon-optimized genes" are genes whose codon usage frequencies are designed to mimic the frequency of codon usage preferences of the host cell.

[0171] "Optimized" polynucleotides are sequences that have been optimized to improve expression in specific heterologous host cells.

[0172] “Plant-optimized nucleotide sequences” are nucleotide sequences optimized for expression in plants (specifically, for increased expression in plants). Plant-optimized nucleotide sequences include codon-optimized genes. Plant-optimized nucleotide sequences can be synthesized by modifying the nucleotide sequence encoding a protein (such as the Cas polypeptide disclosed herein) using one or more plant-preferred codons to improve expression. For a discussion of the use of host-preferred codons, see, for example, Campbell and Gowri, (1990) Plant Physiol. [Plant Physiology] 92:1-11.

[0173] A promoter is a DNA region involved in the recognition and binding of RNA polymerases and other proteins to initiate transcription. A promoter sequence consists of a proximal upstream element and a distal upstream element, the latter often called an enhancer. An enhancer is a DNA sequence that can stimulate promoter activity and can be an intrinsic element of the promoter or a heterologous element inserted to enhance the promoter's level or tissue specificity. A promoter may be derived entirely from a natural gene, or may consist of different elements derived from different promoters existing in nature, and / or contain synthetic DNA segments. Those skilled in the art will understand that different promoters may guide gene expression in different tissues or cell types, at different developmental stages, or in response to different environmental conditions. It is further recognized that, since the exact boundaries of regulatory sequences are not fully defined in most cases, some variant DNA fragments may possess the same promoter activity.

[0174] In most cases, promoters that induce gene expression in most cell types are generally called "constitutive promoters." The term "inducible promoter" refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of endogenous or exogenous stimuli (e.g., by chemical compounds (chemical inducers)), or in response to environmental, hormone, chemical, and / or developmental signals. Inducible or regulatory promoters include promoters that are induced or regulated, for example, by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals (such as ethanol, abscisic acid (ABA), jasmonic acid esters, salicylic acid, or safeners).

[0175] A "leader sequence" is a polynucleotide sequence located between the promoter and coding sequences of a gene. Leader sequences are located upstream of the translation initiation sequence in mRNA. Leader sequences can influence the processing of primary transcripts into mRNA, mRNA stability, or translation efficiency. Examples of leader sequences have been described (e.g., Turner and Foster, (1995) Mol Biotechnol [Molecular Biotechnology] 3:225-236).

[0176] "3' non-coding sequence," "transcription terminator," or "termination sequence" refers to a DNA sequence located downstream of a coding sequence and includes polyadenylation recognition sequences and other sequences encoding regulatory signals that can affect mRNA processing or gene expression. Polyadenylation signals are typically characterized by influencing the addition of polyadenylates to the 3' end of mRNA precursors. The uses of different 3' non-coding sequences are illustrated by Ingelbrecht et al. (1989) Plant Cell 1:671-680.

[0177] “RNA transcript” refers to the product of transcription catalyzed by RNA polymerase of a DNA sequence. When the RNA transcript is a completely complementary copy of the DNA sequence, it is called primary transcript or pre-mRNA. When the RNA transcript is a post-transcriptionally processed RNA sequence derived from primary transcript pre-mRNA, it is called mature RNA or mRNA. “Merchant RNA” or “mRNA” refers to RNA that does not contain introns and can be translated into protein by a cell. “cDNA” refers to DNA that is complementary to an mRNA template and synthesized from the mRNA template using reverse transcriptase. cDNA can be single-stranded or can be converted into double-stranded form using the Klenow fragment of DNA polymerase I. “Sense” RNA refers to RNA transcript that includes mRNA and can be translated into protein in cells or in vitro. “Antisense RNA” refers to RNA transcript that is wholly or partially complementary to a target primary transcript or mRNA and blocks the expression of the target gene (see, for example, U.S. Patent No. 5,107,065). Antisense RNA can be complementary to any part of a specific gene transcript, namely the 5' non-coding sequence, 3' non-coding sequence, intron, or coding sequence. "Functional RNA" refers to antisense RNA, ribozyme RNA, or other RNA that may not be translated but still plays a role in cellular processes. The terms "complementary sequence" and "reverse complementary sequence" are used interchangeably in this document with respect to mRNA transcripts and are intended to define antisense RNA for messengers.

[0178] The term “genome” refers to the complete complementary sequence of genetic material (genes and non-coding sequences) present in every cell of an organism, virus, or organelle; and / or the complete set of chromosomes inherited as (haploid) units from a parent.

[0179] The term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one nucleic acid sequence is regulated by the other. For example, a promoter is operably linked to a coding sequence when it can regulate the expression of that sequence (i.e., the coding sequence is under the transcriptional control of the promoter). The coding sequence can be operably linked to the regulatory sequence in either a sense or antisense orientation. In another instance, complementary RNA regions can be operably linked directly or indirectly at the 5' or 3' of the target mRNA, or within the target mRNA, or the first complementary region is at the 5' of the target mRNA and its complementary sequence is at the 3' of the target mRNA.

[0180] Generally, "host" refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, "host cell" refers to a eukaryotic cell, prokaryotic cell (e.g., bacterial or archaea cell), or cell (e.g., cell line) derived from a multicellular organism cultured as a single-celled entity into which a heterologous polynucleotide or polypeptide has been introduced. In some aspects, the cell is selected from the group consisting of: primitive cells, bacterial cells, eukaryotic cells, eukaryotic single-celled organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, avian cells, insect cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In some cases, the cell is an in vitro cell. In some cases, the cell is an in vivo cell.

[0181] The term "recombination" refers to the artificial combination of two originally separate sequence segments, for example, through chemical synthesis or by manipulating isolated nucleic acid segments using genetic engineering techniques.

[0182] The terms "plasmid," "vector," and "cassette" refer to linear or circular extrachromosomal elements that typically carry a gene that is not part of the cell's central metabolism and are usually in the form of double-stranded DNA. Such elements can be autonomously replicating sequences, genome-integrated sequences, bacteriophages, or nucleotide sequences derived from any source, single-stranded or double-stranded DNA or RNA, in straight or circular form, where many nucleotide sequences have been linked or reassembled into a unique structure capable of introducing a target polynucleotide into the cell. A "transformation cassette" is a specific vector containing a gene and other elements besides the gene that promotes transformation of a specific host cell. An "expression cassette" is a specific vector containing a gene and other elements besides the gene that allows expression of that gene in the host.

[0183] The terms “recombinant DNA molecule,” “recombinant DNA construct,” “expression construct,” “construct,” and “recombinant construct” are used interchangeably herein. A recombinant DNA construct contains nucleic acid sequences, such as an artificial combination of regulatory and coding sequences that are not found together in nature. For example, a recombinant DNA construct may contain regulatory and coding sequences derived from different sources, or it may contain regulatory and coding sequences derived from the same source but arranged in a manner different from that of natural occurrence. Such constructs may be used alone or in combination with a vector. If a vector is used, the choice of vector depends on the methods known to those skilled in the art for introducing the vector into host cells. For example, a plasmid vector may be used. Those skilled in the art fully understand that genetic elements must be present on the vector for successful transformation, selection, and propagation of host cells. Those skilled in the art will also recognize that different independent transformation events can lead to different expression levels and patterns (Jones et al., (1985) EMBO J [Journal of the European Society for Molecular Biology] 4:2411-2418; De Almeida et al., (1989) Mol Gen Genetics [Molecular and General Genetics] 218:78-86), and therefore typically multiple events are screened to obtain strains exhibiting the desired expression levels and patterns. Such screening can be performed using standard molecular biological assays, biochemical assays, and other assays, including Western blot analysis of DNA, Northern blotting analysis of mRNA expression, PCR, real-time quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting analysis of protein expression, enzyme assays or activity assays, and / or phenotypic analysis.

[0184] The term "heterologous" refers to a difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. Non-limiting examples include taxonomic derivation (e.g., a polynucleotide sequence obtained from Zeamays would be heterologous if it were inserted into the genome of a rice plant or a different variety or cultivar of Zeamays; or a polynucleotide obtained from bacteria introduced into plant cells) or sequence differences (e.g., a polynucleotide sequence obtained from Zeamays, isolated, modified, and reintroduced into Zeamays). As used herein, "heterologous" with respect to a sequence can mean that the sequence originates from a different species, variety, or alien species, or, if originating from the same species, has been substantially modified from its natural form in the composition and / or genomic locus through deliberate human intervention. For example, the promoter operatively linked to the heterologous polynucleotide may originate from a species different from the species from which the polynucleotide is derived, or, if from the same / similar species, one or both may be modified substantially from their original form and / or genomic loci, or the promoter may not be a natural promoter of the polynucleotide operatively linked to it. Alternatively, one or more regulatory regions and / or polynucleotides provided herein may be synthesized integrally. In another instance, the target polynucleotide for cleavage by the Cas polypeptide may belong to an organism different from the Cas polypeptide. In yet another instance, the Cas polypeptide and guide RNA may be introduced into the target polynucleotide along with another polynucleotide serving as a template or donor for insertion into the target polynucleotide, wherein the other polynucleotide is heterologous to the target polynucleotide and / or the Cas polypeptide.

[0185] As used herein, the term “expression” refers to the production of a functional end product (e.g., mRNA, guide RNA, or protein) in its precursor or mature form.

[0186] "Mature" proteins are post-translational processed polypeptides (i.e., polypeptides from which any pre-peptides or pro-peptides present in the primary translation product have been removed).

[0187] "Precursor" proteins refer to the primary products of mRNA translation (i.e., still containing propeptides or propeptides). Propeptides and propeptides can be, but are not limited to, intracellular localization signals.

[0188] "CRISPR" (clustered, regularly spaced short palindromic repeating sequences) C lustered R egularly I nterspaced S hort P alindromicR CRISPR loci are genetic loci that encode components of a DNA cutting system, such as those used by bacterial and archaea cells to destroy foreign DNA (Horvath and Barrangou, 2010, Science 327:167-170; WO 2007025097, published March 1, 2007). A CRISPR locus can consist of a CRISPR array whose flanking structures can be different Cas (CRISPR-associated) genes, containing short, homologous repeat sequences (CRISPR repeats) separated by short, variable DNA sequences called 'spacers'.

[0189] As used herein, an "effecton" or "effecton protein" is a protein that has activities including recognizing, binding to, and / or cleaving or nicking polynucleotide targets. An effector or effecton protein can also be an endonuclease. The "effecton complex" of the CRISPR system includes a Cas polypeptide involved in the recognition and binding of crRNA and targets. Some component Cas polypeptides may additionally contain domains involved in the cleavage of target polynucleotides.

[0190] The term "Cas polypeptide" refers to a polypeptide composed of Cas ( CCas polypeptides are polypeptides encoded by genes associated with RISPR. Cas polypeptides include proteins encoded by genes in the Cas locus and include adaptive molecules as well as interfering molecules. Interfering molecules of bacterial adaptive immune complexes include endonucleases. The Cas endonucleases described herein comprise one or more nuclease domains. Cas endonucleases include, but are not limited to: novel Cas-α polypeptides disclosed herein, Cas9 proteins, Cpf1 (Cas12) proteins, C2c1 proteins, C2c2 proteins, C2c3 proteins, Cas3, Cas3-HD, Cas5, Cas7, Cas8, Cas10, or combinations or complexes thereof. When complexed with suitable polynucleotide components, Cas polypeptides can be “Cas endonucleases” or “Cas effector proteins” capable of recognizing, binding to, and optionally cleaving or cutting all or part of a specific polynucleotide target sequence. The Cas-α endonucleases disclosed herein may include those endonucleases having one or more RuvC nuclease domains. Cas peptides are further defined as functional fragments or functional variants of natural Cas peptides, or having at least 50, between 50 and 100, at least 100, between 100 and 150, at least 150, between 150 and 200, at least 200, between 200 and 250, at least 250, between 250 and 300, at least 300, between 300 and 350, at least 350, between 350 and 400, at least 400, between 400 and 450, or at least 500 or more consecutive amino acids of natural Cas peptides, having at least 50%, between 50% and 55%, or at least 55% of the total number of amino acids. Proteins with 55% and 60%, at least 60%, 60% and 65%, at least 65%, 65% and 70%, at least 70%, 70% and 75%, at least 75%, 75% and 80%, at least 80%, 80% and 85%, at least 85%, 85% and 90%, at least 90%, 90% and 95%, at least 95%, 95% and 96%, at least 96%, 96% and 97%, at least 97%, 97% and 98%, at least 98%, 98% and 99%, at least 99%, 99% and 100%, or 100% sequence identity and retaining at least partial activity of the native sequence.

[0191] A “functional fragment” of a Cas polypeptide refers to a portion or subsequence of the Cas polypeptide disclosed herein that retains the ability to recognize, bind to, and optionally unwind, cleave, or cut (introduce single- or double-strand breaks) target sites. A portion or subsequence of a Cas polypeptide may contain a complete peptide or a partial (functional) peptide of any of its domains, such as, but not limited to, a complete functional portion of the Cas3 HD domain, a complete functional portion of the Cas3 helicase domain, or a complete functional portion of a protein (e.g., but not limited to, Cas5, Cas5d, Cas7, and Cas8b1).

[0192] The term “functional variant” of Cas peptide or Cas effector protein refers to a variant of the Cas effector protein disclosed herein that retains all or part of the ability to recognize, bind to, and optionally unwind, cleave, or cleave a target sequence.

[0193] Cas endonucleases may also include multifunctional Cas endonucleases. The terms “multifunctional Cas endonuclease” and “multifunctional Cas endonuclease polypeptide” are used interchangeably herein and include reference to a single polypeptide having Cas endonuclease function (containing at least one protein domain that can function as a Cas endonuclease) and at least one other function, such as, but not limited to, the function of forming a complex (which contains at least a second protein domain that can form a complex with other proteins). In some aspects, multifunctional Cas endonucleases contain at least one additional protein domain relative to those typical domains of Cas endonucleases (inside, upstream (5'), or downstream (3'), or in both 5' and 3' inside, or any combination thereof).

[0194] The terms “cascade” and “cascade complex” are used interchangeably herein and include reference to multi-subunit protein complexes that can assemble with polynucleotides to form polynucleotide-protein complexes (PNPs). A cascade is a polynucleotide-dependent PNP that enables complex assembly and stability, as well as the identification of target nucleic acid sequences. A cascade functions as a surveillance complex that detects and optionally binds to target nucleic acids complementary to variable targeting domains that guide the polynucleotide.

[0195] The terms “5’-cap” and “7-methylguanosine (m7G) cap” are used interchangeably in this document. A 7-methylguanosine residue is located at the 5’ end of messenger RNA (mRNA) in eukaryotes. In eukaryotes, RNA polymerase II (Pol II) transcribes mRNA. Messenger RNA capping typically occurs as follows: the terminal 5’ phosphate group of the mRNA transcript is removed by an RNA terminal phosphatase, leaving two terminal phosphates. Guanylan monophosphate (GMP) is added to the terminal phosphates of the transcript by a guanylate transferase, leaving a 5′-5′ triphosphate-linked guanine at the end of the transcript. Finally, the 7-nitrogen of this terminal guanine is methylated by a methyltransferase.

[0196] The term "without a 5' cap" in this article refers to RNA that has, for example, a 5'-hydroxyl group instead of a 5'-cap. Such RNA, for example, could be called "uncapped RNA." Because 5'-capped RNA has a tendency to export to the nucleus, uncapped RNA can accumulate more effectively in the nucleus after transcription. One or more RNA components described in this article are uncapped.

[0197] As used herein, the term "guide polynucleotide" refers to a polynucleotide sequence that can form a complex with a Cas polypeptide (including the Cas polypeptide described herein) and enable the Cas polypeptide to recognize, optionally bind to, and optionally cleave a DNA target site. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combined sequence).

[0198] The terms “functional fragment” of guide RNA, crRNA or tracrRNA are used interchangeably herein and refer respectively to a portion or subsequence of the guide RNA, crRNA or tracrRNA of this disclosure, which retains the ability to function as guide RNA, crRNA or tracrRNA.

[0199] The terms “functional variant” of guide RNA, crRNA or tracrRNA (respectively) are used interchangeably herein and refer respectively to variants of the guide RNA, crRNA or tracrRNA of this disclosure that retain the ability to function as guide RNA, crRNA or tracrRNA.

[0200] The terms “single guide RNA” and “sgRNA” are used interchangeably in this document and refer to the synthetic fusion of two RNA molecules, comprising a crRNA (CRISPRRNA) containing a variable targeting domain (linked to a tracr-pairing sequence that hybridizes with tracrRNA) and a tracrRNA (trans-activating RNA). CRISPR RNA) fusion. The single guide RNA may contain a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of a type II CRISPR / Cas system that can form a complex with a type II Cas polypeptide, wherein the guide RNA / Cas polypeptide complex can guide the Cas polypeptide to a DNA target site, enabling the Cas polypeptide to recognize, optionally bind to, and optionally cleave or cut (introduce single-strand or double-strand breaks) the DNA target site.

[0201] The terms “variable targeting domain” or “VT domain” are used interchangeably herein and include a nucleotide sequence that can hybridize (complement) with one strand (nucleotide sequence) of a double-stranded DNA target site. The percentage of complementarity between the first nucleotide sequence domain (VT domain) and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. The variable targeting domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some respects, the variable targeting domain comprises a continuous extension of 12 to 30 nucleotides. The variable targeting domain can consist of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.

[0202] The terms “Cas endonuclease recognition domain” or “CER domain” (for guiding polynucleotides) are used interchangeably herein and include the nucleotide sequence that interacts with the Cas polypeptide. The CER domain contains a (trans-acting) tracr nucleotide chaperone sequence, followed by a tracr nucleotide sequence. The CER domain may consist of a DNA sequence, an RNA sequence, a modified DNA sequence, a modified RNA sequence (see, for example, US20150059010A1, published February 26, 2015), or any combination thereof.

[0203] As used herein, the terms “guided polynucleotide / Cas peptide complex,” “guided polynucleotide / Cas peptide system,” “guided polynucleotide / Cas complex,” “guided polynucleotide / Cas system,” “guided Cas system,” “polynucleotide-guided endonuclease,” and “PGEN” are used interchangeably and refer to at least one guide polynucleotide and at least one Cas peptide capable of forming a complex, wherein the guide polynucleotide / Cas peptide complex guides the Cas peptide to a DNA target site, enabling the Cas peptide to recognize, bind to, and optionally cleave or cut (introduce single-strand or double-strand breaks) the DNA target site. The guiding polynucleotide / Cas peptide complexes described herein may comprise one or more Cas peptides and one or more suitable polynucleotide components from any of the known CRISPR systems (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al., 2015, Nature Reviews Microbiology, Vol. 13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13).

[0204] The terms “guide RNA / Cas polypeptide complex,” “guide RNA / Cas polypeptide system,” “guide RNA / Cas complex,” “guide RNA / Cas system,” “gRNA / Cas complex,” “gRNA / Cas system,” “RNA-guided endonuclease,” and “RGEN” are used interchangeably herein and refer to at least one RNA component and at least one Cas polypeptide capable of forming a complex, wherein the guide RNA / Cas polypeptide complex guides the Cas polypeptide to a DNA target site, enabling the Cas polypeptide to recognize, bind to, and optionally cleave or cut (introduce single-strand or double-strand breaks) the DNA target site.

[0205] The terms “target site,” “target sequence,” “target site sequence,” “target DNA,” “target locus,” “genomic target site,” “genomic target sequence,” “genomic target locus,” and “anterior spacer” are used interchangeably herein and refer to a polynucleotide sequence, such as, but not limited to, a nucleotide sequence on any other DNA molecule (including chromosomal DNA, chloroplast DNA, mitochondrial DNA, plasmid DNA) in the chromosome, episome, locus, or genome of a cell, at which a polynucleotide / Cas polypeptide complex can recognize, bind, and optionally cleave or cut. A target site may be an endogenous site in the cell’s genome, or alternatively, a target site may be heterologous to the cell and thus not naturally present in the cell’s genome, or a target site may be found in a heterologous genomic location compared to its location in nature. As used herein, the terms “endogenous target sequence” and “natural target sequence” are used interchangeably herein and refer to a target sequence that is endogenous or natural to the cell’s genome and is located at an endogenous or natural location in the cell’s genome. “Artificial target site” or “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the cell’s genome. Such artificial target sequences can be identical in sequence to endogenous or natural target sequences in the cell's genome, but located at a different position in the cell's genome (i.e., a non-endogenous or non-natural position).

[0206] In this document, “pre-spacer adjacent motif” (PAM) refers to a short nucleotide sequence adjacent to the (targeted) target sequence (pre-spacer) recognized by the polynucleotide / Cas peptide system described herein. If the target DNA sequence is not followed by a PAM sequence, the Cas peptide may not be able to successfully recognize the target DNA sequence. The sequence and length of the PAM in this document can vary depending on the Cas peptide or Cas peptide complex used. The PAM sequence can be of any length, but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0207] The terms “altered target site,” “altered target sequence,” “modified target site,” and “modified target sequence” are used interchangeably herein and refer to a target sequence as disclosed herein that contains at least one alteration when compared to an unaltered target sequence. Such an “alteration” includes, for example: (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, (iv) chemical change of at least one nucleotide, or (v) any combination of (i)–(iv).

[0208] "Modified nucleotide" or "edited nucleotide" means a target nucleotide sequence that contains at least one alteration when compared to its unmodified nucleotide sequence. Such "alteration" includes, for example: (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, (iv) chemical change of at least one nucleotide, or (v) any combination of (i)-(iv).

[0209] The terms “modify target site” and “change target site” are used interchangeably in this article and refer to the methods used to produce the changed target site.

[0210] As used in this article, “donor DNA” is a DNA construct containing a target polynucleotide to be inserted into a genomic target site through homologous targeted repair.

[0211] The term "polynucleotide modification template" includes a polynucleotide that contains at least one nucleotide modification when compared to the nucleotide sequence to be edited. The nucleotide modification can be a substitution, addition, or deletion of at least one nucleotide. Optionally, the polynucleotide modification template may further include homologous nucleotide sequences flanking the at least one nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.

[0212] The term "plant-optimized Cas peptide" in this article refers to a Cas peptide encoded by a nucleotide sequence that has been optimized for expression in plant cells or plants, including multifunctional Cas peptides.

[0213] The terms “plant-optimized nucleotide sequence encoding Cas peptide,” “plant-optimized construct encoding Cas peptide,” and “plant-optimized polynucleotide encoding Cas peptide” are used interchangeably herein and refer to a nucleotide sequence encoding a Cas peptide, or a variant or functional fragment thereof, which has been optimized for expression in plant cells or plants. Plants containing plant-optimized Cas peptides include: plants containing nucleotide sequences encoding Cas sequences, and / or plants containing Cas peptides. In some respects, plant-optimized Cas peptide nucleotide sequences are Cas peptides optimized for maize, rice, wheat, soybean, cotton, or canola.

[0214] The term "plant" generally includes the whole plant, plant organs, plant tissues, seeds, plant cells, and the offspring of a plant. A plant can be monocotyledonous or dicotyledonous. Plant cells include, but are not limited to, cells derived from: seeds, suspension cultures, embryos, meristematic regions, callus, leaves, roots, buds, gametophytes, sporophytes, pollen, and microspores. "Plant element" is intended to refer to the whole plant or plant component, which may include differentiated and / or undifferentiated tissues, such as, but not limited to, plant tissues, parts, and cell types. In one aspect, a plant element is one of the following: the whole plant, seedling, meristematic tissue, ground tissue, vascular tissue, cortical tissue, seeds, leaves, roots, buds, stems, flowers, fruits, stolons, bulbs, tubers, corms, asexual terminal shoots, buds, young shoots, tumor tissue, and various forms of cells and cultures (e.g., single cells, protoplasts, embryos, and callus). It should be noted that protoplasts are not technically “complete” plant cells (all components are naturally present) because they lack cell walls. The term “plant organ” refers to a plant tissue or group of tissues that constitutes a morphologically and functionally distinct part of a plant. As used herein, “plant element” is a synonym for “part” of a plant, referring to any part of a plant and may include different tissues and / or organs, and may be used interchangeably with the term “tissue” throughout the text. Similarly, “plant reproductive element” is intended generally to refer to any plant part capable of creating another plant through the sexual or asexual reproduction of that plant, such as, but not limited to: seeds, seedlings, roots, buds, cuttings, scions, grafted seedlings, stolons, bulbs, tubers, corms, asexual terminal shoots, or buds. Plant elements can exist in plants or in plant organs, tissue cultures, or cell cultures.

[0215] "Offspring" includes any subsequent generations of a plant.

[0216] As used herein, the term "plant part" means plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant masses, and intact plant cells in a plant or plant part (such as embryo, pollen, ovule, seed, leaf, flower, branch, fruit, kernel, spike, rachis, husk, stem, root, root tip, anther, etc.), together with these parts themselves. "Grain" means mature seed produced by commercial growers for purposes other than cultivation or propagation of a species. Progeny, variants, and mutants of regenerated plants are also included within the scope of this disclosure, provided that these parts contain introduced polynucleotides.

[0217] The term "monocotyledonous" or "monocotyledonous" refers to a subclass of angiosperms, also known as the "monocotyledonous class," whose seeds typically contain only one embryonic leaf or cotyledon. The term encompasses references to the whole plant, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and their offspring.

[0218] The term "dicotyledonous" or "dicotyledonous" refers to a subclass of angiosperms, also known as the "dicotyledonous plant class," whose seeds typically contain two embryonic leaves or cotyledons. The term encompasses references to the whole plant, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and their offspring.

[0219] The term “unconventional yeast” in this article refers to any yeast species that is not a species of the genus *Saccharomyces* (e.g., *Saccharomyces cerevisiae*) or *Schizosaccharomyces*. (See “Non-Conventional Yeasts in Genetics, Biochemistry and Biotechnology: Practical Protocols”, K. Wolf, KDBreunig, G. Barth, eds., Springer-Verlag, Berlin, Germany, 2003).

[0220] In the context of this disclosure, the term "hybrid" or "cross" (or "crossing") refers to the fusion of gametes via pollination to produce offspring (i.e., cells, seeds, or plants). This term encompasses sexual hybridization (one plant being pollinated by another) and self-pollination (self-pollination, i.e., when pollen and ovules (or microspores and megaspores) are from the same plant or plants with identical genes).

[0221] The term "introgression" refers to the phenomenon of a desired allele at a genetic locus being transferred from one genetic background to another. For example, introgression of a desired allele at a designated locus can be transferred to at least one progeny plant via sexual hybridization between two parent plants, where at least one of the parent plants carries the desired allele in its genome. Alternatively, allele transfer can occur, for example, via recombination between two donor genomes, such as in fused protoplasts, where at least one of the donor protoplasts carries the desired allele in its genome. The desired allele can be, for example, a transgenic, modified (mutated or edited) natural allele, or a selected allele of a marker or QTL.

[0222] "Introduction" is intended to mean providing a polynucleotide or polypeptide or polynucleotide-protein complex to a target, such as a cell or organism, in such a way that one or more components enter the interior of the cell or reach the cell itself.

[0223] "Target polynucleotide" includes any nucleotide sequence encoding a protein or polypeptide that improves crop desirability (i.e., agronomically significant traits). Target polynucleotides include, but are not limited to, polynucleotides encoding important traits such as agronomical, herbicide-resistance, insecticide-resistance, disease resistance, nematode resistance, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, grain characteristics, commercial products, phenotypic markers, or any other trait of agronomical or commercial importance. Target polynucleotides may also be utilized in a sense or antisense orientation. Furthermore, more than one target polynucleotide may be utilized together or "stacked" to provide additional benefits.

[0224] "Complex trait loci" include genomic loci of multiple transgenes that are genetically linked to each other.

[0225] As used herein, the terms “reduced,” “fewer,” “slower,” and “increased,” “faster,” “enhanced,” and “greater” refer to a reduction or increase in a characteristic aspect of a modified plant element or produced plant compared to an unmodified plant element or produced plant. For example, a reduction in a characteristic aspect could be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% to 20%, at least 15%, at least 20%, 20% to 30%, at least 25%, at least 30%, 30% to 40%, at least 35%, at least 40%, 40% to 50%, at least 45%, at least 50%, 50% to 60%, at least about 60%, 60% to 70%, 70% to 80%, at least 75%, at least about 80%, 80% to 90%, at least about 90%, 90% to 100%, at least 100%, 100% and 200%, at least 200%, at least about 300%, at least about 400%, or More, and the increase can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% to 20%, at least 15%, at least 20%, 20% to 30%, at least 25%, at least 30%, 30% to 40%, at least 35%, at least 40%, 40% to 50%, at least 45%, at least 50%, 50% to 60%, at least about 60%, 60% to 70%, 70% to 80%, at least 75%, at least about 80%, 80% to 90%, at least about 90%, 90% to 100%, at least 100%, 100% and 200%, at least 200%, at least about 300%, at least about 400% or more higher than the untreated control.

[0226] As used in this article, when referring to sequence position, the term "previous" means that one sequence appears upstream or 5' upstream of another sequence.

[0227] The abbreviations have the following meanings: "sec" means second, "min" means minute, "h" means hour, "d" means day, "µL" means microliter, "mL" means milliliter, "L" means liter, "µM" means micromolar, "mM" means millimolecular concentration, "M" means mole, "mmol" means millimole, "µmole" or "umole" means micromolar, "g" means gram, "µg" or "ug" means microgram, "ng" means nanogram, "U" means unit, "bp" means base pair, and "kb" means kilobase.

[0228] Classification of CRISPR-Cas systems

[0229] CRISPR-Cas systems have been classified based on the sequence and structural analysis of their components. Multiple CRISPR / Cas systems have been described, including Class 1 systems with multi-subunit effector complexes (including types I, III, and IV) and Class 2 systems with single protein effectors (including types II, V, and VI) (Makarova et al. 2015, Nature Reviews Microbiology, Vol. 13: 1-15; Zetsche et al. 2015, Cell, 163, 1-13; Shmakov et al. 2015, Molecular Cell, 60, 1-13; Haft et al. 2005, Computational Biology, PLoS Comput Biol, 1(6):e60; and Koonin et al. 2017, Curr Opinion Microbiology, 37: 67-78).

[0230] The CRISPR-Cas system comprises at least a CRISPR RNA (crRNA) molecule and at least one CRISPR-associated (Cas) protein to form a crRNA-ribonucleoprotein (crRNP) effector complex. The CRISPR-Cas locus contains a series of identical repetitive sequences interspersed with DNA-targeting spacers encoding crRNA components and operon-like units of the cas gene encoding Cas polypeptide components. The resulting ribonucleoprotein complex recognizes polynucleotides in a sequence-specific manner (Jore et al., Nature Structural & Molecular Biology 18, 529-536 (2011)). The crRNA acts as a guide RNA for the sequence-specific binding of effectors (proteins or complexes) to double-stranded DNA sequences by forming base pairs with the complementary DNA strand while simultaneously substituting non-complementary strands to form so-called R loops (Jore et al., 2011. Nature Structural & Molecular Biology 18, 529-536).

[0231] The RNA transcript (precrRNA) at a CRISPR locus is specifically cleaved in the repetitive sequence by a CRISPR-associated (Cas) ribonuclease in type I and III systems, or by RNase III in type II systems. The number of CRISPR-associated genes at a given CRISPR locus can vary between species.

[0232] Different CRISPR systems contain different cas genes that encode proteins with different domains. The cas operon contains genes encoding one or more effector endonucleases and other cas polypeptides. Protein subunits include those described in: Makarova et al. 2011, Nat Rev Microbiol. 2011 9(6):467-477; Makarova et al. 2015, Nature Reviews Microbiology 13:1-15; and Koonin et al. 2017, Current Opinion Microbiology 37:67-78. Domain types include those involved in expression (precrRNA processing, such as Cas 6 or RNase III), interference (including effector modules for crRNA and target binding, and one or more domains for target cleavage), adaptation (spacer insertion, such as Cas1 or Cas2), and auxiliary (regulatory, auxiliary, or unknown functions). Some domains can serve more than one function. For example, Cas9 includes domains for endonuclease function and domains for target cleavage.

[0233] The Cas peptide is directed by a single CRISPR RNA (crRNA) that recognizes DNA target sites near the preseptal neighbor motif (PAM) via direct RNA-DNA base pairing (Jore, MM et al., 2011, Nat. Struct. Mol. Biol. [Nature Structural & Molecular Biology] 18:529-536; Westra, ER et al., 2012, Molecular Cell [Molecular Cell] 46:595-605; and Sinkunas, T. et al., 2013, EMBO J. [Journal of the European Society for Molecular Biology] 32:385-394).

[0234] Class I CRISPR-Cas systems

[0235] Class I CRISPR-Cas systems include types I, III, and IV. Class I systems are characterized by the presence of an effector endonuclease complex rather than a single protein. The Cascade complex comprises an RNA recognition motif (RRM) and a nucleic acid-binding domain, which is the core fold of various RAMP (repetitive sequence-associated mystery protein) superfamilies (Makarova et al. 2013, Biochem Soc Trans. 41, 1392-1400; Makarova et al. 2015, Nature Reviews Microbiology, Vol. 13, 1-15). RAMP protein subunits include Cas5 and Cas7 (which contain the backbone of the crRNA–effector complex), with the Cas5 subunit binding to the 5' stalk of the crRNA and interacting with the large subunit, and typically including Cas6, which loosely associates with the effector complex and typically acts as a repeat sequence-specific RNase in precrRNA processing (Charpentier et al., FEMS Microbiol Rev [FEMS Microbiology Review] 2015, 39:428-441; Niewoehner et al., RNA 2016, 22:318-329).

[0236] The type I CRISPR-Cas system contains an effector protein complex called Cascade (a CRISPR-associated complex for antiviral defense), which contains at least Cas5 and Cas7. The effector complex functions together with a single CRISPR RNA (crRNA) and Cas3 to defend against invading viral DNA (Brouns, SJJ et al., Science 321:960-964; Makarova et al. 2015, Nature Reviews Microbiology, Vol. 13:1-15). The type I CRISPR-Cas locus contains the characteristic gene cas3 (or variants cas3' or cas3''), which encodes a metal-dependent nuclease that is a superfamily 2 helicase stimulated by single-stranded DNA (ssDNA) and capable of unwinding double-stranded DNA (dsDNA) and RNA-DNA duplexes (Makarova et al. 2015, Nature Reviews; Microbiology, Vol. 13:1-15). Following target recognition, the Cas3 endonuclease is recruited into the Cascade-crRNA-target DNA complex to cleave and degrade the DNA target (Westra, ER et al. (2012) Molecular Cell 46:595-605, Sinkunas, T. et al. (2011) EMBO J. 30:1335-1342, and Sinkunas, T. et al. (2013) EMBO J. 32:385-394). In some type I systems, Cas6 is the active endonuclease responsible for crRNA processing, while Cas5 and Cas7 function as non-catalytic RNA-binding proteins; although in type I systems, crRNA processing can be catalyzed by Cas5 (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). Type I system is divided into seven subtypes (Makarova et al. 2011, Nat Rev Microbiol. [Nature Review Microbiology] 2011 9(6):467-477; Koonin et al. 2017, Curr Opinion Microbiology [New Insights in Microbiology] 37:67-78).A modified type I CRISPR-associated complex (Cascade) for adaptive antiviral defense has been described, comprising at least the protein subunits Cas7, Cas5 and Cas6, wherein one of these subunits is synthetically fused with either the Cas3 endonuclease or the modified restriction endonuclease FokI (WO 2013098244, published July 4, 2013).

[0237] The type III CRISPR-Cas system (including multiple cas7 genes) targets ssRNA or ssDNA and functions as an RNase and a DNA nuclease activated by RNA (Tamulaitis et al., Trends in Microbiology 25(10)49-61, 2017). The Csm (type III-A) and Cmr (type III-B) complexes function as RNA-activated single-stranded (ss) DNases (coupling target RNA binding / cleavage with ssDNA degradation). Upon infection with foreign DNA, the CRISPR RNA (crRNA)-directed Csm or Cmr complexes bind to newly generated transcripts, recruiting Cas10 DNases to actively transcribed phage DNA, leading to degradation of the transcripts and phage DNA, rather than host DNA. The Cas10 HD-domain is responsible for ssDNase activity, and the Csm3 / Cmr4 subunits are responsible for the endonuclease activity of the Csm / Cmr complex. The 3' flanking sequence of the target RNA is crucial for the ssDNA enzyme activity of Csm / Cmr: base pairing with the 5'-stalk of the crRNA protects the host DNA from degradation.

[0238] Type IV systems, despite including typical Type I Cas5 and Cas7 domains as well as Cas8-like domains, may lack a CRISPR array, which is a feature of most other CRISPR-Cas systems.

[0239] Type II CRISPR-Cas systems

[0240] Class II CRISPR-Cas systems include types II, V, and VI. Type II systems are characterized by the presence of a single Cas effector protein, rather than an effector complex. Type II and V Cas peptides contain a RuvC endonuclease domain that employs RNase H folding.

[0241] Type II CRISPR / Cas systems use crRNA and tracrRNA (trans-activating CRISPR RNA) to guide the Cas peptide to its DNA target. The crRNA contains a spacer region complementary to one strand of the double-stranded DNA target and a region that pairs with the tracrRNA to form an RNA duplex, which guides the Cas peptide to cleave the DNA target, leaving blunt ends. The spacer is obtained through a process involving the Cas1 and Cas2 proteins that is not fully understood. Type II CRISPR / Cas loci typically include the cas1 and cas2 genes, as well as the cas9 gene (Chylinski et al., 2013, RNA Biology 10:726-737; Makarova et al., 2015, Nature Reviews Microbiology Vol. 13:1-15). Type II CRISPR-Cas loci can encode tracrRNA, which is partially complementary to a repetitive sequence within the corresponding CRISPR array, and may contain other proteins such as Csn1 and Csn2. The presence of cas9 near cas1 and cas2 genes is a marker of type II loci (Makarova et al. 2015, Nature Reviews Microbiology, Vol. 13: 1-15).

[0242] The type V CRISPR / Cas system contains a single Cas polypeptide, including Cpf1 (Cas12) (Koonin et al., CurrOpinion Microbiology 37:67-78, 2017), which is an active RNA-directed endonuclease that does not necessarily require additional trans-activated CRISPR (tracr) RNA for target cleavage, unlike Cas9.

[0243] The type VI CRISPR-Cas system contains the cas13 gene, which encodes a nuclease with two HEPN (higher eukaryotic and prokaryotic nucleotide-binding) domains but lacking the HNH or RuvC domains, and is independent of tracrRNA activity. Most HEPN domains contain conserved motifs that constitute the metal-independent internal RNase active site (Anantharam et al., Biol Direct 8:15, 2013). Due to this characteristic, the type VI system is thought to act on RNA targets, rather than the DNA targets common in other CRISPR-Cas systems.

[0244] In a first aspect, this disclosure provides a method for altering the specificity of the preseptal neighbor motif (PAM) of a target Cas-α polypeptide. As used herein, a “preseptal neighbor motif” (PAM) refers to a short nucleotide sequence adjacent to a (targeted) target sequence (preseptum) that can be recognized by a guiding polynucleotide / Cas polypeptide system. If the target DNA sequence is not followed by a PAM sequence, the Cas polypeptide may fail to recognize the target DNA sequence. The sequence and length of the PAM as used herein can vary depending on the Cas polypeptide or Cas polypeptide complex used. The PAM sequence can be of any length, but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0245] In some aspects, methods for altering the PAM specificity of a target Cas-α peptide include: (a) comparing a PAM interaction (PI) domain of a heterologous, orthologous Cas-α peptide with a PI domain of the target Cas-α peptide, wherein the orthologous Cas-α peptide has a different PAM recognition than the target Cas-α peptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-α peptide; (c) incorporating one or more amino acids and / or one or more polypeptide chains selected from the PI domain of the orthologous Cas-α peptide into one or more structurally similar positions of the target Cas-α peptide to produce a modified target Cas-α peptide; and (d) determining the PAM recognition in the modified target Cas-α peptide.

[0246] As used herein, "orthologous Cas-α polypeptide" or "Cas-α ortholog" refers to a Cas-α polypeptide containing a CRISPR-Cas polypeptide that comprises at least one zinc finger-like domain, at least one bridge-helix-like domain, a triplet RuvC domain (comprising discontinuous RuvC-I, RuvC-II, and RuvC-III domains) ranging in size from approximately 327 to 777 amino acids, and contains GxxxG, ExL, and Cx. n C or Cx n (C,H) motif (where x represents any amino acid and n = one or more amino acids).

[0247] The Cas-α orthologs that can be used in the methods disclosed herein include Cas-α 1, Cas-α 2, Cas-α 3, Cas-α 4, Cas-α 5, Cas-α 6, Cas-α 7, Cas-α 8, Cas-α 10, Cas-α 11, Cas-α 13, Cas-α 24 and Cas-α 29.

[0248] Figure 1 This study elucidates the phylogenetic relationships among some Cas-α orthologs divided into three supergroups (I, II, and III). Group I includes clade 1 (candidate archaea and Aureabacteria (typically encoding Cas1, Cas2, and Cas4 at the locus)). Group II includes clade 2 (Aquagenic Bacteria (*Sulfurihydrogenibium* and *Hydrogenivirga*) and Deltaproteobacteria (*Desulfovibrio*), clade 3 (candidate archaea (typically encoding Cas1, Cas2, and Cas4 at the locus)), clade 4 (Bacteroidetes (*Prevotella* and *Bacteroidetes*)), clade 5 (candidate Levybacterium), and clade 6 (Clostridium (*Dorea*, *Ruminococcus*, *Clostridium*, *Krosterus*, *Peptocolstridium*, *Cellulosilyticym*, *Eubacteria*, *Mutobacteria*)). Group III includes... Clade 7 (Bacilli (Bacillus, Acidibacillus, Aneurinibacillus, Brachybacillus, Parageobacillus, Alicyclobacillus)), clade 8 (Negativicutes (Phascolarctobacterium))), and clade 9 (Flavobacteriia (Flavobacterium))). The diamond symbol represents the orthologous Cas-α1-11 endonuclease.

[0249] As used in this paper, “structurally similar position” refers to the coordinates of a similar three-dimensional position in a polypeptide when the structure of the polypeptide (predicted (e.g., informational models based on neural networks, such as AlphaFold)) or determined (e.g., using cryo-electron microscopy (Cryo-EM), X-ray crystallography, and NMR spectroscopy)) is aligned or superimposed with a predicted or determined orthologous structure.

[0250] Methods for comparing the PI domains of target Cas-α peptides and orthologous Cas-α peptides include, but are not limited to, multiple sequence comparisons by log-expectation (MUSCLE), multiple sequence comparisons using Clustal Omega, root mean square distance (RMSD), distance matrix alignment (DALI), structural homology by environment-based alignment (SHEBA), combinatorial extension (CE), homology alignment database (HOMSTRAD), protein structure classification (SCOP), FatCat, and PhyreStorm.

[0251] In some examples of the disclosed methods for altering the PAM specificity of a target Cas-α peptide, the step of selecting one or more amino acids from the PI domain of an orthologous Cas-α peptide includes comparing the PI domain sequence and / or structure of the orthologous Cas-α peptide with the PI domain sequence and / or structure of the target Cas-α peptide from which its PAM specificity is to be altered, and replacing one or more amino acids from the orthologous Cas-α peptide into the Cas-α peptide target.

[0252] In some examples of the disclosed methods for altering the PAM specificity of a target Cas-α peptide, the step of selecting one or more polypeptide chains from the PI domain of an orthologous Cas-α peptide includes comparing the PI domain sequence and / or structure of the orthologous Cas-α peptide with the PI domain sequence and / or structure of the target Cas-α peptide from which its PAM specificity is to be altered, and replacing one or more polypeptide chains from the orthologous Cas-α peptide into the Cas-α peptide target.

[0253] Methods for determining PAM recognition of modified target Cas-α peptides include, but are not limited to, transcription and translation of the modified Cas-α peptide in a cellular or cell-free mixture, complexing the modified Cas-α peptide with a guide RNA to form a ribonucleoprotein (RNP), incubating the RNP with a DNA species containing a fixed guide RNA target and a collection of different PAM sequences, capturing DNA molecules that support target cleavage, sequencing the PAM regions of the DNA species that support target cleavage, and using a position frequency matrix to calculate common PAMs to summarize changes specific to PAM.

[0254] In some examples of the disclosed methods for altering the PAM specificity of a target Cas-α polypeptide, the target Cas-α polypeptide is a Cas-α 10 polypeptide. The target Cas-α polypeptide may be a Cas-α 10 polypeptide having an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2.

[0255] On the other hand, this disclosure provides synthetic or non-naturally occurring Cas-α10 peptides with modified PAM specificity. More specifically, this disclosure discloses Cas-α10 peptides comprising a modified PAM interaction (PI) domain, such that the resulting Cas-α10 peptide recognizes PAM sequences in target polynucleotides other than 5'-TTC-3'.

[0256] On the other hand, this disclosure provides synthetic or non-naturally occurring Cas-α10 polypeptides with altered PAM specificity, comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with any of SEQ ID NO: 5-16 or 28-87.

[0257] In examples of this, synthetic or non-naturally occurring Cas-α 10 polypeptides with modified PAM specificity may contain an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 5-16 or 28-87.

[0258] In another example of this, synthetic or non-naturally occurring Cas-α10 polypeptides with altered PAM specificity may contain the amino acid sequence of SEQ ID NO: 5-16 or 28-87.

[0259] On the other hand, this disclosure provides synthetic or non-naturally occurring Cas-α10 peptides with modified PAM specificity, wherein the Cas-α10 peptide contains a PI domain that recognizes a PAM sequence comprising 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', and 5'-GTTY-3'. 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

[0260] On the other hand, this disclosure provides synthetic or non-naturally occurring Cas-α10 polypeptides with modified PAM specificity, wherein the Cas-α10 polypeptide comprises an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with respect to the amino acid position of SEQ ID NO: 2, and wherein the synthetic or non-naturally occurring Cas-α10 polypeptide contains an amino acid sequence with at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with respect to the amino acid position of SEQ ID NO: 2. The 10 polypeptide sequences contain the following mutations: K85S, K85A, N92R, N92K, Q125K, N88H, N88K, N88Q, Q125F, Y72V, Y72E, Y72Q, Y72T, Y72C, Y72A, Y72S, Y72P, Y72G, Y72D, Y72L, K85G, and K85. D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0261] In examples of this, synthetic or non-naturally occurring Cas-α 10 polypeptides with altered PAM specificity may comprise an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, wherein the Cas-α 10 polypeptide sequence contains, relative to the amino acid positions of SEQ ID NO: 2, the following mutations: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85 D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0262] In another example of this, a synthetic or non-naturally occurring Cas-α10 polypeptide with altered PAM specificity may comprise the amino acid sequence of SEQ ID NO: 2, wherein the Cas-α10 polypeptide sequence contains, relative to the amino acid positions of SEQ ID NO: 2, the following mutations: K85S, K85A, N92R, N92K, Q125K, N88H, N88K, N88Q, Q125F, Y72V, Y72E, Y72Q, Y72T, Y72C, Y72A, Y72S, Y72P, Y72G, Y72D, Y72L, K85G, K85 D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

[0263] In another aspect, this disclosure provides synthetic or non-naturally occurring Cas-α 10 polypeptides with modified PAM specificity, wherein the Cas-α 10 polypeptide may comprise an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2, and wherein, relative to the amino acid position of SEQ ID NO: 2, the Cas-α The 10 polypeptide sequence contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S and N88D mutations; and... Combinations of Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations. The following combinations of mutations are possible: Y72S, K85D, Q125R, and N127R; Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or K85Q and N92W.

[0264] In examples of this, synthetic or non-naturally occurring Cas-α 10 polypeptides with modified PAM specificity may comprise an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, and wherein the Cas-α 10 polypeptide sequence is relative to SEQ ID NO: The amino acid position of 2 contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and... Combinations of Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations. The following combinations of mutations are possible: Y72S, K85D, Q125R, and N127R; Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or K85Q and N92W.

[0265] In another example of this, a synthetic or non-naturally occurring Cas-α10 polypeptide with altered PAM specificity may comprise the amino acid sequence of SEQ ID NO: 2, wherein the Cas-α10 polypeptide sequence is relative to SEQ ID NO: The amino acid position of 2 contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and... Combinations of Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations. The following combinations of mutations are possible: Y72S, K85D, Q125R, and N127R; Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or K85Q and N92W.

[0266] In a further aspect, this disclosure provides a synthetic composition comprising: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with any one of SEQ ID Nos. 5-16 or 28-87; and (b) at least one guiding polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide, wherein the Cas-α The 10 polypeptide recognizes the PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, wherein the complex binds to the target polynucleotide.

[0267] In an example of this, the synthesized composition may comprise: (a) a CAS-α 10 polypeptide having DNA-binding activity, the CAS-α 10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of SEQ ID NO. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0268] In another example of this aspect, the synthesized composition may comprise: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising the amino acid sequence of any one of SEQ ID No. 5-16 or 28-87; and (b) at least one guide polynucleotide comprising a region complementary to a target polynucleotide, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0269] In another aspect, this disclosure provides synthetic compositions comprising: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising a PAM interaction (PI) domain, wherein the PI domain recognizes a PAM sequence on a target polynucleotide, wherein the PAM sequence comprises 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', etc. ', 5'-GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C; and (b) at least one guiding polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide forms a complex with the at least one guiding polynucleotide, and wherein the complex binds to the target polynucleotide.

[0270] In another aspect, this disclosure provides a synthetic composition comprising: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2, and wherein the Cas-α 10 polypeptide is relative to SEQ ID NO: The amino acid position of 2 contains the following mutations: K85S, K85A, N92R, N92K, Q125K, N88H, N88K, N88Q, Q125F, Y72V, Y72E, Y72Q, Y72T, Y72C, Y72A, Y72S, Y72P, Y72G, Y72D, Y72L, K85G, and K85D. (a) mutations; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0271] In an example of this, the synthesized composition may comprise: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, and wherein the Cas-α 10 polypeptide is relative to SEQ ID NO: The amino acid positions of 2 include K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D. (a) mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0272] In another example of this aspect, the synthesized composition may comprise: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising the amino acid sequence of SEQ ID NO: 2 and the following mutations: K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation. (a) K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0273] In another aspect, this disclosure provides a synthetic composition comprising: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2, and wherein the Cas-α 10 polypeptide is relative to SEQ ID NO: The amino acid position of 2 contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and... Combinations of Q89G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations. Combinations of Y72S, K85D, Q125R, and N127R mutations; combinations of Y72A, K85S, Q89D, N92L, and Q125R mutations; combinations of Y72C, N88H, Q89G, and Q125R mutations; combinations of Y72C, N88D, Q89D, N92W, and Q125R mutations; or combinations of K85Q and N92W mutations.(b) includes at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0274] In an example of this, the synthesized composition may comprise: (a) a Cas-α 10 polypeptide having DNA-binding activity, the Cas-α 10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, and wherein the Cas-α 10 polypeptide is relative to SEQ ID NO: The amino acid position of 2 contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and Q8... Combinations of 9G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Y7 The combination of 2S mutation, K85D mutation, Q125R mutation and N127R mutation; the combination of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation and Q125R mutation; the combination of Y72C mutation, N88H mutation, Q89G mutation and Q125R mutation; the combination of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation and Q125R mutation; or the combination of K85Q and N92W mutation; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0275] In another example of this aspect, the synthetic or non-naturally occurring composition may comprise: (a) a Cas-α10 polypeptide having DNA-binding activity, the Cas-α10 polypeptide comprising SEQ ID NO: The amino acid sequence of 2 and one or more of the following combinations: K85Q mutation and N92L mutation; N88H mutation and Q89G mutation; N88K mutation and Q89G mutation; N88Q mutation and Q89G mutation; K85S mutation and N92L mutation; K85S mutation and N92Q mutation; K85S mutation and N92C mutation; K85S mutation and N92H mutation; K85S mutation and N92A mutation; K85S mutation and N92M mutation; K85N mutation and N92L mutation; K85N mutation and N92H mutation; K85N mutation and N92A mutation; K85N mutation and N92C mutation; K85N mutation and N92M mutation; K85N mutation and N92Q mutation; K85N mutation and N92I mutation; K85S mutation, N88D mutation and Q89 Combinations of G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Y72 Combinations of S mutations, K85D mutations, Q125R mutations, and N127R mutations; combinations of Y72A mutations, K85S mutations, Q89D mutations, N92L mutations, and Q125R mutations; combinations of Y72C mutations, N88H mutations, Q89G mutations, and Q125R mutations; combinations of Y72C mutations, N88D mutations, Q89D mutations, N92W mutations, and Q125R mutations; or combinations of K85Q and N92W mutations; and (b) at least one guide polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the origin of the Cas-α10 polypeptide, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

[0276] In a further aspect, this disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing the cell with a Cas-α 10 polypeptide having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with any of SEQ ID Nos. 5-16 or 28-87, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) At least one nucleotide modification is introduced into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0277] In an example of this aspect, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-α10 polypeptide having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of SEQ ID Nos. 5-16 or 28-87, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0278] In another example of this aspect, a method for editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-α 10 polypeptide of any one of SEQ ID Nos. 5-16 or 28-87, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide.

[0279] In some examples of the aforementioned methods for editing target polynucleotides in cells, the PAM sequence recognized by the Cas-α10 peptide includes 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', and 5'-DTTY. -3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

[0280] In some instances, the disclosed methods for editing target polynucleotides in cells further include providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0281] In a further aspect, this disclosure provides a method for editing a target polynucleotide in a cell, the method comprising: (a) providing the cell with a Cas-α10 polypeptide comprising a PAM interaction (PI) domain, wherein the PI domain recognizes a PAM sequence on the target polynucleotide, and wherein the PAM sequence comprises 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-G TTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'- TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0282] In some aspects of methods for editing target polynucleotides in cells, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0283] On the other hand, this disclosure provides a method for editing target polynucleotides in cells, the method comprising: (a) providing the cell with a Cas-α 10 polypeptide comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2, wherein the Cas-α 10 polypeptide sequence is relative to SEQ ID NO: The amino acid positions of 2 include K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide.

[0284] In an example of this, a method of editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-α10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas-α10 polypeptide sequence is relative to SEQ ID NO: The amino acid positions of 2 include K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide.

[0285] In another example of this aspect, a method for editing target polynucleotides in a cell may include: (a) providing the cell with a Cas-α10 polypeptide comprising the amino acids of SEQ ID NO: 2, wherein the Cas-α10 polypeptide comprises, relative to the amino acid position of SEQ ID NO: 2, a K85S mutation; a K85A mutation; an N92R mutation; an N92K mutation; a Q125K mutation; an N88H mutation; an N88K mutation; an N88Q mutation; a Q125F mutation; a Y72V mutation; a Y72E mutation; a Y72Q mutation; a Y72T mutation; a Y72C mutation; a Y72A mutation; a Y72S mutation; a Y72P mutation; a Y72G mutation; a Y72D mutation; a Y72L mutation; a K85G mutation; a K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation, wherein the Cas-α 10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α 10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α 10 polypeptide.

[0286] On the other hand, this disclosure provides a method for editing target polynucleotides in cells, the method comprising: (a) providing the cell with a Cas-α 10 polypeptide comprising an amino acid sequence having at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90%, alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, or alternatively at least 100% sequence identity with SEQ ID NO: 2, wherein the Cas-α 10 polypeptide sequence is relative to SEQ ID NO: The amino acid position of 2 contains one or more of the following combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and Q89G mutations. Combinations of mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Y72S mutations... The following combinations of mutations are possible: Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or K85Q and N92W mutations, wherein the Cas-α 10 polypeptide recognizes the PAM sequence on the target polynucleotide.(b) Providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0287] In an example of this aspect, a method of editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-α 10 polypeptide comprising an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas-α 10 polypeptide sequence, relative to the amino acid position of SEQ ID NO: 2, comprises one or more of the following combinations: a combination of K85Q mutation and N92L mutation; a combination of N88H mutation and Q89G mutation; a combination of N88K mutation and Q89G mutation; a combination of N88Q mutation and Q89G mutation; a combination of K85S mutation and N92L mutation; a combination of K85S mutation and N92Q mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92H ...C mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92C mutation; Combinations of N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; K85S, N88D, and Q89G mutations. Combinations of mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Y72S mutations... (a) A combination of mutations: Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or a combination of mutations: K85Q and N92W, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) Providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0288] In another example of this aspect, a method of editing a target polynucleotide in a cell may include: (a) providing the cell with a Cas-α10 polypeptide comprising the amino acid of SEQ ID NO: 2, wherein the Cas-α10 polypeptide sequence comprises one or more of the following combinations relative to the amino acid position of SEQ ID NO: 2: a combination of K85Q mutation and N92L mutation; a combination of N88H mutation and Q89G mutation; a combination of N88K mutation and Q89G mutation; a combination of N88Q mutation and Q89G mutation; a combination of K85S mutation and N92L mutation; a combination of K85S mutation and N92Q mutation; a combination of K85S mutation and N92C mutation; a combination of K85S mutation and N92H ... Combinations of N92A mutations; combinations of K85S and N92M mutations; combinations of K85N and N92L mutations; combinations of K85N and N92H mutations; combinations of K85N and N92A mutations; combinations of K85N and N92C mutations; combinations of K85N and N92M mutations; combinations of K85N and N92Q mutations; combinations of K85N and N92I mutations; K85S, N88D, and Q89G mutations. Combinations of mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Y72S mutations... (a) A combination of mutations: Y72A, K85S, Q89D, N92L, and Q125R; Y72C, N88H, Q89G, and Q125R; Y72C, N88D, Q89D, N92W, and Q125R; or a combination of mutations: K85Q and N92W, wherein the Cas-α10 polypeptide recognizes a PAM sequence on the target polynucleotide; (b) Providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; and (c) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

[0289] In some aspects of methods for editing target polynucleotides in cells, the PAM sequence recognized by the Cas-α10 peptide includes 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5'-GTTY-3', and 5'-DTTY. -3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

[0290] In some aspects of methods for editing target polynucleotides in cells, the method further includes providing the cell with a donor DNA molecule or a polynucleotide modification template.

[0291] CRISPR-Cas system components

[0292] Cas polypeptide

[0293] Many proteins can be encoded in CRISPR cas operons, including those involved in adaptation (spacer insertion), interference (target binding, target cleavage, or cutting – such as endonuclease activity), expression (precrRNA processing), regulation, or others.

[0294] Cas1 and Cas2 are conserved in many CRISPR systems (e.g., Koonin et al., Curr Opinion Microbiology 37:67-78, 2017). Cas1 is a metal-dependent DNA-specific endonuclease that produces double-stranded DNA fragments. In some systems, Cas1 and Cas2 form a stable complex, which is crucial for spacer acquisition and insertion in CRISPR systems (Nuñez et al., Nature Structural Molecular Biology 21:528-534, 2014).

[0295] Many other proteins have been identified in different systems, including Cas4 (which may be similar to RecB nuclease) and are thought to play a role in capturing novel viral DNA sequences for integration into CRISPR arrays (Zhang et al., PLOS One 7(10):e47232, 2012).

[0296] Some proteins may have multiple functions. For example, Cas9, a characteristic protein of the type II system, has been shown to be involved in precrRNA processing, target binding, and target cleavage.

[0297] Cas endonucleases and effectors

[0298] Nucleotide endonucleases are enzymes that cleave phosphodiester bonds within polynucleotide chains, and include restriction endonucleases that cleave DNA at specific sites without damaging bases. Examples of endonucleases include restriction endonucleases, macronucleases, TAL effector nucleases (TALENs), zinc finger nucleases, and Cas (…). C RISPR- as (sociated) effector endonucleases.

[0299] Cas endonucleases (either as a single effector protein or as an effector complex with other components) unwind the DNA double helix at a target sequence and optionally cleave at least one DNA strand, as mediated by recognition of the target sequence by a polynucleotide (e.g., but not limited to crRNA or guide RNA) complexed with a Cas effector protein. Such recognition and cleavage of the target sequence by Cas endonucleases typically occurs if the correct prespacer adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. Alternatively, the Cas endonucleases described herein may lack DNA cleavage or nicking activity, but may still specifically bind to the DNA target sequence when complexed with a suitable RNA component. (See also U.S. Patent Application US20150082478, published March 19, 2015, and US20150059010, published February 26, 2015).

[0300] Cas endonucleases can appear as single effectors (type 2 CRISPR systems) or as part of larger effector complexes (type 1 CRISPR systems).

[0301] The Cas endonucleases described include, but are not limited to, Cas9, Cas12f (Cas-α, Cas14), Cas12l (Cas-β), Cas12a (Cpf1), Cas12b (C2c1 protein), Cas13 (C2c2 protein), Cas12c (C2c3 protein), Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas3, Cas3-HD, Cas5, Cas6, Cas7, Cas8, Cas10, or combinations or complexes thereof. Cas endonucleases and effector proteins can be used for targeted genome editing (via single and multiple double-strand breaks and nicks) and targeted genome regulation (via linking epigenetic effector domains to Cas peptides or sgRNAs). Cas endonucleases can also be engineered to function as RNA-directed recombinases and can act as scaffolds for assembling multiprotein and nucleic acid complexes via RNA lineages (Mali et al., 2013 Nature Methods, Vol. 10: 957-963).

[0302] Cas-α endonuclease

[0303] Cas-α endonucleases (e.g., Cas-α 10, also known as Cas12f) are defined as functional RNA-directed PAM-dependent dsDNA cleaving proteins having fewer than 800 amino acids and comprising: a C-terminal RuvC catalytic domain (split into three subdomains) and further comprising a bridge-helix and one or more zinc finger motifs; and an N-terminal Rec subunit with a helical bundle, a WED wedge-like (or “oligonucleotide-binding domain”, OBD) domain, and optional zinc finger motifs. For example, some exemplary Cas-α endonucleases are described in US10934536 and WO 2022082179.

[0304] Existing literature has demonstrated that the RuvC domain contains endonuclease function. Cas-α endonucleases can be isolated or identified from loci containing cas-α genes encoding effector proteins and from arrays containing multiple repetitive sequences. In some respects, cas-α loci may also contain partial or complete cas1, cas2, and / or cas4 genes.

[0305] Zinc finger motifs are domains that typically coordinate one or more zinc ions via cysteine ​​and histidine side chains to stabilize their folding. Zinc fingers are named according to the pattern of cysteine ​​and histidine residues coordinated with the zinc ion (e.g., C4 indicates a zinc ion coordinated with four cysteine ​​residues; C3H indicates a zinc ion coordinated with three cysteine ​​residues and one histidine residue).

[0306] Cas-α peptides contain one or more zinc finger (ZFN) coordination motifs that can form zinc-binding domains. Zinc finger-like motifs help separate target and non-target strands and load guide RNA into DNA targets. Cas-α peptides containing one or more zinc finger motifs can provide additional stability to ribonucleoprotein complexes on target polynucleotides. Cas-α peptides contain a C4 or C3H zinc-binding domain.

[0307] Cas-α endonucleases are RNA-directed endonucleases that bind to and cleave double-stranded DNA targets comprising: (1) a sequence homologous to the nucleotide sequence of the guiding RNA, and (2) a PAM sequence. In some respects, PAM is T-rich. In some respects, PAM is C-rich.

[0308] Cas-α endonucleases function as double-strand break inducers, but can also be cleavage enzymes or single-strand break inducers. In some respects, catalytically inactive Cas-α endonucleases can be used to target or recruit to target DNA sequences without inducing cleavage. In some respects, catalytically inactivated Cas-α peptides can be used in conjunction with functional endonucleases to cleave target sequences. In some respects, catalytically inactivated Cas-α peptides can be combined with base-editing molecules, such as deaminases. A "deaminase" is an enzyme that catalyzes deamination reactions. For example, deamination of adenine with adenine deaminase results in the formation of inosine. Inosine pairs selectively with cytosine, but not thymine. This leads to a post-replication conversion mutation, converting the original AT base pair to a GC base pair. In another instance, cytosine deamination leads to the formation of uracil, which can be repaired back to a CT base pair or a TA, GC, or AT base pair via cellular repair mechanisms. This heterogeneity in repair can be suppressed by introducing uracil glycosylase inhibitors, causing DNA repair or replication to convert the original CT base pairs to TA base pairs (Burnett et al. (2022) Frontiers in Genome Editing. 4, 923718). In the case of adenine and cytosine deaminases, the introduction of a nick promotes the corresponding base pair changes (Burnett et al., 2022). In some respects, the deaminase can be acytidine deaminase. In some respects, the deaminase can be adenine deaminase. In some respects, the deaminase can be ADAR-2.

[0309] The "functional fragment" of the Cas-α endonuclease retains the ability to recognize, bind to, or cleave a single strand of a double-stranded polynucleotide, or cleave both strands of a double-stranded polynucleotide, or any combination thereof.

[0310] The Cas peptides, effector proteins, or functional fragments thereof used in the methods of this disclosure can be isolated from natural or recombinant sources, in which genetically modified host cells are modified to express the nucleic acid sequence encoding the protein. Alternatively, Cas peptides can be generated using a cell-free protein expression system or synthesized. Effector Cas nucleases can be isolated and introduced into heterologous cells, or modified from their natural form to exhibit a different type or magnitude of activity than that from their natural source. Such modifications include, but are not limited to, fragmentation, variants, substitution, deletion, and insertion.

[0311] Fragments and variants of Cas peptides and Cas effector proteins can be obtained via methods such as site-directed mutagenesis and synthetic construction. Methods for measuring endonuclease activity are well known in the art, for example, but not limited to, WO 2013166113 disclosed on November 7, 2013, WO 2016186953 disclosed on November 24, 2016, and WO2016186946 disclosed on November 24, 2016.

[0312] The Cas peptides disclosed herein can be modified. Modified forms of Cas peptides may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nuclease activity of the naturally occurring Cas peptide. For example, in some cases, the modified forms of Cas peptides have less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the corresponding wild-type Cas peptide (US20140068797, disclosed March 6, 2014). In some cases, the modified forms of Cas peptides have no substantial nuclease activity and are referred to as catalytically “inactivated Cas” or “dCas”. Inactivated Cas / dCas includes inactivated Cas endonuclease (dCas). Catalytically inactivated Cas effector proteins can be fused with heterologous sequences to induce or modify activity.

[0313] The Cas peptide disclosed herein may be part of a fusion protein comprising one or more heterologous protein domains (e.g., 1, 2, 3 or more domains other than the Cas peptide). Such a fusion protein may comprise any additional protein sequence, as well as an optional linker sequence between any two domains (e.g., between the Cas peptide and the first heterologous domain). Examples of protein domains that can be fused with the Cas peptides described herein include, but are not limited to, epitope tags (e.g., histidine [His], V5, FLAG, influenza hemagglutinin [HA], myc, VSV-G, thioredoxin [Trx]); reporter molecules (e.g., glutathione-5-transferase [GST], horseradish peroxidase [HRP], chloramphenicol acetyltransferase [CAT], β-galactosidase, β-glucuronidase [GUS], luciferase, green fluorescent protein [GFP], HcRed, DsRed, cyan fluorescent protein [CFP], yellow fluorescent protein [YFP], blue fluorescent protein [BFP]); and domains containing one or more of the following activities: methyltransferase activity, demethyltransferase activity, transcriptional activation activity (e.g., VP16 or VP64), transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. The disclosed Cas peptide can also be fused with proteins that bind to DNA molecules or other molecules, such as maltose-binding protein (MBP), S-tag, Lex A DNA-binding domain (DBD), GAL4A DNA-binding domain, and herpes simplex virus (HSV) VP16.

[0314] Catalytically active and / or catalytically inactive Cas peptides can be fused to a heterologous sequence (US20140068797, published March 6, 2014). Suitable fusion couplers include, but are not limited to, peptides that provide activity that indirectly increases transcription by acting directly on target DNA or on peptides associated with that target DNA (e.g., histones or other DNA-binding proteins). Other suitable fusion couplers include, but are not limited to, peptides that provide methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinase activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristylation activity, or demyristylation activity. Further suitable fusion partners include, but are not limited to, polypeptides that directly provide increased transcription of the target nucleic acid (e.g., transcription activators or fragments thereof, proteins or fragments thereof that recruit transcription activators, small molecule / drug-responsive transcription regulators, etc.). Partially active or catalytically inactivated Cas-α endonucleases may also fuse with another protein or domain, such as Clo51 or FokI nucleases, to produce double-strand breaks (Guilinger et al., Nature Biotechnology, Vol. 32, No. 6, June 2014).

[0315] Catalytically active or catalytically inactive Cas peptides, such as the Cas-α peptide disclosed herein, can also be fused with molecules that guide the editing of single or multiple bases in a polynucleotide sequence. These molecules are, for example, site-specific deaminases that can alter the identity of nucleotides, for example, from C•G to T•A or from A•T to G•C (Gaudelli et al., Programmable baseediting of A•T to G•C in genomic DNA without DNA cleavage). *Nature* (2017); Nishida et al., “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” *Science* 353 (6305) (2016); Komor et al., “Programmable editing of a targetbase in genomic DNA without double-stranded DNA cleavage.” [Programmable editing of target bases in genomic DNA without double-strand DNA cleavage]"Nature [Nature] 533 (7603) (2016):420-4. Base editing fusion proteins may contain, for example, active (generating double-strand breaks), partially active (cleavage enzymes), or inactivated (catalytically inactivated) Cas-α endonucleases and deaminases (e.g., but not limited to cytidine deaminase, adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABE, etc.). In some respects, base editing repair inhibitors and glycosylation enzyme inhibitors (e.g., uracil glycosylation enzyme inhibitors (to prevent uracil removal)) are considered other components of base editing systems.

[0316] Any of the Cas peptides disclosed herein (e.g., catalytically inactivated Cas peptides) may also be complexed with base-editing molecules via an RNA-aptamer system (such as any of those described in WO 2021 / 055459, published March 25, 2021). The RNA-aptamer system may include, for example, (A) RNA motifs such as (1) a telomerase Ku-binding motif, (2) a telomerase Sm7-binding motif, (3) an MS2 phage operator stem-loop, (4) a PP7 phage operator stem-loop, (5) an SfMu phage Com stem-loop, (6) a chemically modified version of such an aptamer, or (7) a non-natural RNA aptamer, and (B) the corresponding aptamer ligand or its RNA-binding segment. See also U.S. Patent 11,479,793 and WO 2018 / 129129, published July 12, 2018.

[0317] The Cas peptides and endonucleases described herein can be expressed and purified using methods known in the art, such as those described in WO / 2016 / 186953, published on November 24, 2016.

[0318] To date, numerous Cas peptides and endonucleases have been described that recognize specific PAM sequences (WO 2016186953, WO 2016186946, and Zetsche B et al., 2015. Cell [Cell] 163, 1013) and cleave target DNA at specific sites. It should be understood that, based on the methods and aspects of using novel guided Cas systems described herein, those skilled in the art can now tailor these methods to utilize any guided endonuclease system.

[0319] Cas effector proteins may contain heterologous nuclear localization sequences (NLS). For example, the heterologous NLS amino acid sequence described herein may have sufficient strength to drive the accumulation of detectable amounts of Cas peptides in the nuclei of the yeast cells described herein. NLS may contain one (monotype) or more (e.g., ditype) short sequences (e.g., 2 to 20 residues) of basic, positively charged residues (e.g., lysine and / or arginine) and may be located anywhere in the Cas amino acid sequence, but such that it is exposed on the protein surface. For example, NLS may be operatively linked to the N-terminus or C-terminus of the Cas peptide described herein. Two or more NLS sequences may be linked to the Cas peptide, for example, at both the N-terminus and C-terminus of the Cas peptide. Cas peptide genes may be operatively linked to the SV40 nuclear targeting signal upstream of the Cas codon region and the ditype VirD2 nuclear localization signal downstream of the Cas codon region (Tinland et al. (1992) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 89:7442-6). Non-limiting examples of suitable NLS sequences herein include those disclosed in U.S. Patent Nos. 6,660,830 and 7,309,576.

[0320] Guided polynucleotides

[0321] Guide polynucleotides enable Cas peptides to target recognition, binding, and optionally cleavage, and can be single-molecule or bimolecule-molecule. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combinatorial sequence). Optionally, the guide polynucleotide may contain at least one nucleotide, a phosphodiester bond, or a linker modification, such as, but not limited to, locked nucleic acid (LNA), 5-methyldC, 2,6-diaminopurine, 2'-fluoroA, 2'-fluoroU, 2'-O-methylRNA, a thiophosphate bond, a linker to a cholesterol molecule, a linker to a polyethylene glycol molecule, a linker to a spacer 18 (hexaethylene glycol chain) molecule, or a 5' to 3' covalent linker leading to cyclization. Guide polynucleotides containing only ribonucleic acid are also called "guide RNA" or "gRNA" (US20150082478, published March 19, 2015, and US20150059010, published February 26, 2015). Guide polynucleotides can be engineered or synthesized.

[0322] Guide polynucleotides include chimeric, non-naturally occurring guide RNAs that contain regions not found together in nature (i.e., they are heterologous to each other). For example, a chimeric, non-naturally occurring guide RNA contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize with a nucleotide sequence in the target DNA, which is linked to a second nucleotide sequence that recognizes the Cas polypeptide, such that the first and second nucleotide sequences are not found linked together in nature.

[0323] Guide polynucleotides can be bimolecules (also known as double-stranded guide polynucleotides) containing cr nucleotide sequences (e.g., crRNA) and tracr nucleotide sequences (e.g., tracrRNA). In some cases, there are linker polynucleotides that connect crRNA and tracrRNA to form a single guide (e.g., sgRNA).

[0324] In some respects, the cr nucleotide includes a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize with a nucleotide sequence in the target DNA and serves as a... Cas polypeptide The second nucleotide sequence (also known as the tracr pairing sequence) is part of the recognition (CER) domain. The tracr pairing sequence can hybridize with the tracr nucleotide along the complementary region and together form Cas polypeptide The recognition domain or CER domain is used. The CER domain is capable of interacting with the Cas polypeptide. The cr and tracr nucleotides of the double-stranded polynucleotide can be RNA, DNA, and / or RNA-DNA combination sequences. In some respects, the cr nucleotide molecule of the double-stranded polynucleotide is referred to as “crDNA” (when it is composed of successive extensions of DNA nucleotides) or “crRNA” (when it is composed of successive extensions of RNA nucleotides) or “crDNA-RNA” (when it is composed of a combination of DNA and RNA nucleotides). The cr nucleotide can be a fragment of crRNA naturally occurring in bacteria and archaea. The size of the crRNA fragment naturally occurring in bacteria and archaea that can be present in the cr nucleotides disclosed herein can be, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides.

[0325] In some respects, tracr nucleotides are referred to as “tracrRNA” (when composed of consecutive extensions of RNA nucleotides) or “tracrDNA” (when composed of consecutive extensions of DNA nucleotides) or “tracrDNA-RNA” (when composed of a combination of DNA and RNA nucleotides). In one respect, the RNA guiding the RNA / Cas9 endonuclease complex is a double-stranded RNA comprising double-stranded crRNA-tracrRNA. In the 5'-to-3' direction, tracrRNA (trans-activating CRISPR RNA) comprises (i) a sequence annealed to the repeat region of CRISPR type II crRNA and (ii) a stem-loop portion (Deltcheva et al., Nature [Nature] 471:602-607). The double-stranded guiding polynucleotide can form a complex with a Cas polypeptide, wherein the guiding polynucleotide / Cas polypeptide complex (also referred to as the guiding polynucleotide / Cas polypeptide system) guides the Cas polypeptide to a genomic target site, enabling the Cas polypeptide to recognize, bind to, and optionally cleave or cut (introduce single-strand or double-strand breaks) the target site. (US20150082478, published on March 19, 2015, and US20150059010, published on February 26, 2015).

[0326] In some respects, the guiding polynucleotide is a guiding polynucleotide capable of forming a PGEN as described herein, wherein the guiding polynucleotide comprises a first nucleotide sequence domain complementary to a nucleotide sequence in the target DNA and a second nucleotide sequence domain interacting with the Cas polypeptide.

[0327] In some respects, the guiding polynucleotide is the guiding polynucleotide as described herein, wherein each of the first nucleotide sequence domain and the second nucleotide sequence domain is selected from the group consisting of DNA sequences, RNA sequences, and combinations thereof.

[0328] In some respects, the guiding polynucleotide is the guiding polynucleotide described herein, which further comprises a second nucleotide sequence domain selected from the group consisting of: RNA backbone modifications that enhance stability, DNA backbone modifications that enhance stability, and combinations thereof (see Kanasty et al., 2013, Common RNA-backbone modifications, Nature Materials 12:976-977; US20150082478 published March 19, 2015 and US20150059010 published February 26, 2015).

[0329] The guide RNA can comprise a bimolecule containing a chimeric, non-naturally occurring crRNA linked to at least one tracrRNA. Chimeric, non-naturally occurring crRNAs include crRNAs containing regions that are not found together in nature (i.e., they are heterologous to each other). For example, the crRNA contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) capable of hybridizing with a nucleotide sequence in the target DNA, which is linked to a second nucleotide sequence (also called a tracr pairing sequence) such that the first and second sequences are not found linked together in nature.

[0330] Guided polynucleotides can also be single molecules containing a cr nucleotide sequence linked to a tracr nucleotide sequence (also known as a single-guided polynucleotide). A single-guided polynucleotide contains a first nucleotide sequence domain (called a variable-targeting domain) that can hybridize with a nucleotide sequence in the target DNA. V ariable T The Cas endonuclease recognition domain (or VT domain) and the Cas endonuclease recognition domain (CER domain) that interacts with the Cas polypeptide.

[0331] The VT domain and / or CER domain of a single-direction polynucleotide can contain an RNA sequence, a DNA sequence, or an RNA-DNA combination sequence. A single-direction polynucleotide composed of sequences from cr and tracr nucleotides can be referred to as a “single-direction RNA” (when composed of consecutive extensions of RNA nucleotides), a “single-direction DNA” (when composed of consecutive extensions of DNA nucleotides), or a “single-direction RNA-DNA” (when composed of a combination of RNA and DNA nucleotides). The single-direction polynucleotide can form a complex with a Cas polypeptide, wherein the directing polynucleotide / Cas polypeptide complex (also referred to as a directing polynucleotide / Cas polypeptide system) directs the Cas polypeptide to a genomic target site, enabling the Cas polypeptide to recognize, bind to, and optionally cleave or cut (introduce single-strand or double-strand breaks) the target site. (US20150082478, published March 19, 2015, and US20150059010, published February 26, 2015).

[0332] Chimeric, non-naturally occurring single guide RNAs (sgRNAs) include sgRNAs containing regions that are not found together in nature (i.e., they are heterologous to each other). For example, an sgRNA contains a first nucleotide sequence domain (called a variable targeting domain or VT domain) that can hybridize with a nucleotide sequence in a target DNA, which is linked to a second nucleotide sequence (also called a tracr pairing sequence) that is not found to be linked together in nature.

[0333] The nucleotide sequence linking the single-guided polynucleotide (cr) and tracr nucleotide can comprise an RNA sequence, a DNA sequence, or an RNA-DNA combination sequence. On one hand, the nucleotide sequence linking the single-guided polynucleotide (cr) and tracr nucleotide (also called a "loop") can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46. The length can be 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides. On the other hand, the nucleotide sequence linking the cr and tracr nucleotides of a single-direction polynucleotide can contain a tetranucleotide ring sequence, such as, but not limited to, the GAAA tetranucleotide ring sequence.

[0334] Guide polynucleotides can be produced by any method known in the art, including chemically synthesized guide polynucleotides (such as, but not limited to, Hendel et al. 2015, Nature Biotechnology 33, 985-989), in vitro generated guide polynucleotides, and / or self-splicing guide RNAs (such as, but not limited to, Xie et al. 2015, PNAS 112:3570-3575).

[0335] Guided polynucleotide / Cas polypeptide complex

[0336] The guiding polynucleotide / Cas polypeptide complex described herein can recognize, bind to, and optionally cleave, unwind, or cleave all or part of a target sequence (e.g., a Cas endonuclease or a Cas polypeptide with cleavage or cleavage activity or a Cas polypeptide with nuclease or endonuclease activity).

[0337] A guide polynucleotide / Cas polypeptide complex capable of cleaving both strands of a DNA target sequence typically comprises a Cas polypeptide having all of its endonuclease domains in a functional state (e.g., wild-type endonuclease domains or variants thereof retaining some or all of the activity in each endonuclease domain). Therefore, a wild-type Cas polypeptide (e.g., the Cas polypeptide disclosed herein) or a variant thereof retaining some or all of the activity in each endonuclease domain of the Cas polypeptide is a suitable example of a Cas endonuclease capable of cleaving both strands of a DNA target sequence.

[0338] A guide polynucleotide / Cas endonuclease complex capable of cleaving only one strand of a DNA target sequence can be characterized herein as having nicking activity (e.g., partial cleavage capability). Cas nickases typically contain a functional endonuclease domain that allows Cas to cleave only one strand of the DNA target sequence (i.e., form a nick). For example, a Cas9 nickase may contain (i) a mutated, dysfunctional RuvC domain and (ii) a functional HNH domain (e.g., a wild-type HNH domain). As another example, a Cas9 nickase may contain (i) a functional RuvC domain (e.g., a wild-type RuvC domain) and (ii) a mutated, dysfunctional HNH domain. Non-limiting examples of Cas9 nickases applicable herein are disclosed in US20140189896, published July 3, 2014. A pair of Cas nickases can be used to increase the specificity of DNA targeting. Generally, this can be achieved by providing two Cas nickases that target and nick DNA sequences in the desired target region on opposite strands by associating with RNA components that have different guide sequences. Such a nearby cleavage of each DNA strand produces a double-strand break (i.e., a DSB with a single-stranded overhang), which is then recognized as a substrate for non-homologous end joining (NHEJ) (which tends to produce incomplete repair leading to mutations) or homologous recombination (HR). Each nick in these aspects may be spaced apart from each other, for example, by at least about 5, between 5 and 10, at least 10, between 10 and 15, at least 15, between 15 and 20, at least 20, between 20 and 30, at least 30, between 30 and 40, at least 40, between 40 and 50, at least 50, between 50 and 60, at least 60, between 60 and 70, at least 70, between 70 and 80, at least 80, between 80 and 90, at least 90, between 90 and 100, or 100 or more (or any integer between 5 and 100). One or both of the Cas nickase proteins described herein may be used for Cas nickase pairs. For example, a Cas9 nickase with a mutated RuvC domain but a functional HNH domain (i.e., Cas9 HNH+ / RuvC-) (e.g., Streptococcus pyogenes Cas9 HNH+ / RuvC-) may be used. By using the appropriate RNA components described herein, which have guide RNA sequences that target each cas9 nickase to each specific DNA site, each cas9 nickase (e.g., cas9 HNH+ / RuvC-) is directed to a specific DNA site adjacent to each other (separated by up to 100 base pairs).

[0339] In some respects, a Cas polypeptide complex can bind to a DNA target sequence but not cleave any strand at the target sequence. Such a complex can contain a Cas polypeptide in which all nuclease domains are mutated or dysfunctional. For example, a Cas9 protein that can bind to a DNA target sequence but does not cleave any strand at the target sequence can contain a mutated, dysfunctional RuvC domain and a mutated, dysfunctional HNH domain. Cas polypeptides of this document that bind to but do not cleave target DNA sequences can be used to regulate gene expression, for example, in which case the Cas polypeptide can be fused with a transcription factor (or a portion thereof) (e.g., a repressor or activator, such as any of those disclosed herein).

[0340] In some respects, the guiding polynucleotide / Cas endonuclease complex (PGEN) described herein is a PGEN, wherein the Cas endonuclease is optionally covalently or non-covalently linked to or assembled to at least one protein subunit or a functional fragment thereof.

[0341] In some aspects, a guide polynucleotide / Cas endonuclease complex is a guide polynucleotide / Cas endonuclease complex (PGEN) comprising at least one guide polynucleotide and at least one Cas endonuclease polypeptide, wherein the Cas endonuclease polypeptide comprises at least one protein subunit or a functional fragment thereof, wherein the guide polynucleotide is a chimeric, non-naturally occurring guide polynucleotide, and wherein the guide polynucleotide / Cas endonuclease complex is capable of recognizing, binding to, and optionally cleaving, unwinding, or cutting all or part of a target sequence.

[0342] Cas effector proteins can be Cas-α effector proteins as disclosed herein.

[0343] In some aspects, a guiding polynucleotide / Cas effector complex is a guiding polynucleotide / Cas effector protein complex (PGEN) comprising at least one guiding polynucleotide and a Cas-α effector protein, wherein the guiding polynucleotide / Cas effector protein complex is capable of recognizing, binding to, and optionally cleaving, unwinding or cutting all or part of a target sequence.

[0344] PGEN can be a guiding polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises one or more copies of at least one protein subunit or a functional fragment thereof. In some aspects, the protein subunit is selected from the group consisting of: Cas1 protein subunit, Cas2 protein subunit, Cas4 protein subunit, and any combination thereof. PGEN can be a guiding polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises at least two distinct protein subunits selected from the group consisting of Cas1, Cas2, and Cas4.

[0345] PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein further comprises at least three different protein subunits or functional fragments thereof, the protein subunits being selected from the group consisting of Cas1, Cas2 and optionally an additional Cas polypeptide containing Cas4.

[0346] In some aspects, the guiding polynucleotide / Cas effector protein complex described herein is a PGEN, wherein the Cas effector protein is covalently or non-covalently linked to at least one protein subunit or a functional fragment thereof. A PGEN can be a guiding polynucleotide / Cas effector protein complex wherein the Cas effector protein polypeptide is covalently or non-covalently linked to or assembled to one or more copies of at least one protein subunit or a functional fragment thereof, the at least one protein subunit being selected from the group consisting of: a Cas1 protein subunit, a Cas2 protein subunit, an additional Cas polypeptide optionally containing a Cas4 protein subunit, and any combination thereof. A PGEN can be a guiding polynucleotide / Cas effector protein complex wherein the Cas effector protein is covalently or non-covalently linked to or assembled to at least two different protein subunits, the at least two different protein subunits being selected from the group consisting of: Cas1, Cas2, and an additional Cas polypeptide optionally containing Cas4. PGEN can be a guide polynucleotide / Cas effector protein complex, wherein the Cas effector protein is covalently or non-covalently linked to at least three different protein subunits or functional fragments thereof, the at least three different protein subunits being selected from the group consisting of Cas1, Cas2, and an additional Cas polypeptide optionally containing Cas4, and any combination thereof.

[0347] Any component of the directing polynucleotide / Cas effector protein complex, the directing polynucleotide / Cas effector protein complex itself, and one or more polynucleotide modification templates and / or one or more donor DNAs may be introduced into a heterologous cell or organism by any method known in the art.

[0348] Recombinant constructs for cell transformation

[0349] The disclosed guide polynucleotides, Cas peptides, polynucleotide modification templates, donor DNA, the guide polynucleotide / Cas peptide systems disclosed herein, and any combination thereof (optionally further comprising one or more target polynucleotides) can be introduced into cells. Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, unconventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein.

[0350] The standard recombinant DNA and molecular cloning techniques used in this paper are well known in the art and are described more fully in Sambrook et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory, Cold Spring Harbor, New York (1989). The transformation methods are well known to those skilled in the art and are described below.

[0351] Vectors and constructs include circular plasmids and linear polynucleotides containing the target polynucleotide and optionally other components, including linkers, adaptors, regulatory components, or analytical components. In some instances, recognition sites and / or target sites may be contained within introns, coding sequences, 5' UTRs, 3' UTRs, and / or regulatory regions.

[0352] NHEJ and HDR

[0353] In some respects, the Cas peptide described herein can be part of a genome editing system that further comprises one or more guide polynucleotides and optionally donor DNA, and the editable target polynucleotide sequence includes non-homologous end joining (NHEJ) or homologous recombination (HR) following a Cas peptide-mediated double-strand break. Once a double-strand break is induced in DNA, the cell's DNA repair mechanisms are activated to repair the break. The most common repair mechanism that brings the broken ends together is the non-homologous end joining pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). The structural integrity of chromosomes is typically preserved through repair, but deletions, insertions, or other rearrangements are possible (Siebert and Puchta, (2002) Plant Cell 14:1121-31; Pacher et al., (2007) Genetics 175:21-9). Alternatively, the double-strand break can be repaired by homologous recombination between homologous DNA sequences. Once the sequence surrounding a double-strand break is altered, for example through the activity of a mature exonuclease involved in the double-strand break, gene conversion pathways can restore the original structure if homologous sequences are present, such as homologous chromosomes in non-dividing somatic cells, or sister chromatids after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenetic DNA sequences can also serve as DNA repair templates for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0354] As used herein, “donor DNA” is a DNA construct containing a target polynucleotide to be inserted into a genomic target site, wherein the insertion is mediated by a Cas polypeptide. Once a double-strand break is introduced at the target site via a nuclease, the first and second homologous regions of the donor DNA can undergo homologous recombination with their corresponding homologous genomic regions, resulting in an exchange of DNA between the donor and target genomes. Therefore, the method provided leads to the integration of the target polynucleotide of the donor DNA into a double-strand break at the target site in the plant genome, thereby altering the original target site and producing an altered genomic target site.

[0355] Base editing

[0356] In some respects, the Cas peptide described herein can be part of a genome editing system that further includes a base editing agent and multiple guide polynucleotides, and editing the target polynucleotide sequence involves introducing multiple nucleobase edits into the target polynucleotide sequence to produce a variant nucleotide sequence.

[0357] One or more nucleobases of a target polynucleotide can be chemically altered, in some cases changing the base from one type to another, such as from cytosine to thymine or from adenine to guanine. In some respects, multiple bases, such as 2 or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or even more than 100, 200 or more, up to thousands of bases, can be modified or altered to produce plants with multiple modified bases.

[0358] Any base editing complex, such as a base editing agent associated with an RNA-directing protein, can be used to target and bind to a desired locus in an organism's genome and chemically modify one or more components of the target polynucleotide.

[0359] It can achieve site-specific base conversions to engineer one or more nucleotide changes, thereby creating one or more edits in the genome. These include, for example, site-specific base editing mediated by C•G to T•A or A•T to G•C base editing deaminases (Gaudelli et al., Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage, Nature [Nature] (2017); Nishida et al., “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems, Science [Science] 353 (6305) (2016); Komor et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage, Nature [Nature] 533 (7603)). (2016):420-4). Catalytic “death” or inactivation of Cas9 (dCas9) fused with cytidine deaminase or adenine deaminase proteins (e.g., the catalytically inactivated “death” form of the Cas endonucleases disclosed herein) becomes a specific base editor that can alter DNA bases without inducing DNA breaks. Base editors convert C->T (or G->A on the opposite strand) or adenine base editors convert adenine to inosine, thereby causing an A->G change within the editing window specified by the gRNA. Any molecule that affects nucleobase changes is a “base editor”.

[0360] For many target traits, the creation of single- and double-strand breaks and subsequent repair via HDR or NHEJ are not ideal for quantifying the trait. Observed phenotypes include genotypic and environmental effects. Genotypic effects further include additive, dominant, and epistatic effects. The probability that any single edit has no effect may be greater than zero, and any single phenotypic effect may be small, depending on the method used and the site chosen. Double-strand break repair can also be "noisy" and have low reproducibility.

[0361] One approach to improve the probability of no effect or small phenotypic effect per edit is multiple genomic editing, which modifies multiple target sites. Methods that modify genomic sequences without introducing double-strand breaks will allow for single-base substitutions. Combining these methods, multiple base editing facilitates the creation of a large number of genotype edits that can produce observable phenotypic modifications. In some cases, dozens, hundreds, or thousands of sites can be edited within one or several generations of an organism.

[0362] Multiple approaches to base editing in organisms have the potential to create multiple significant phenotypic variations across one or several generations, with a positive bias towards the effect. In some respects, this organism is a plant. Plants or plant populations with multiple edits can be hybridized to produce progeny plants, some of which will contain multiple edits from the parent lines. In this way, accelerated breeding of desired traits can be achieved in parallel across one or several generations, replacing time-consuming traditional sequential hybridization and breeding across multiple generations.

[0363] Base-editing deaminases, such as cytidine deaminases or adenine deaminases, can fuse with RNA-directed endonucleases that can be inactivated (“dCas”, such as inactivated Cas9) or partially activated (“nCas”, such as Cas9 nickases), preventing them from cleaving the target site to which they are directed. dCas forms a functional complex at the target site with a guide polynucleotide sharing homology with the polynucleotide sequence and further complexes with the deaminase molecule. The guide Cas polypeptide recognizes and binds to the double-stranded target sequence, thereby opening the double strand to expose individual bases. In the case of cytidine deaminase, the deaminase deaminates the cytosine base and produces uracil. Uracil glycosylation inhibitors (UGIs) are provided to prevent U from converting back to C. DNA replication or repair mechanisms then convert uracil to thymine (U to T) and subsequently repair the opposite base (previously G in the original GC pair) to adenine, resulting in a TA pair. For example, see Komor et al., Nature, Vol. 533, pp. 420-424, May 19, 2016. Base editing capabilities can be enhanced using RNA aptamer systems, such as those discussed below: WO 2021 / 055459, published March 25, 2021; U.S. Patent 11,479,793; and WO 2018 / 129129, published July 12, 2018.

[0364] Preview Editor

[0365] In some respects, the Cas peptide described herein can be part of a genome editing system that further comprises a lead editor and a guiding polynucleotide, and editing the target nucleotide sequence involves introducing one or more insertions, deletions, or nucleobase exchanges into the target nucleotide sequence without generating double-strand DNA breaks. See, for example, Anzalone, AV, Randolph, PB, Davis, JR, et al. Search-and-replace genomeediting without double-strand breaks or donor DNA. Nature 576, 149–157 (2019).

[0366] In some respects, the lead editor is a Cas polypeptide fused with reverse transcriptase, wherein the Cas polypeptide is modified to nick DNA rather than generate double-strand breaks. This Cas-polypeptide-reverse transcriptase fusion may also be referred to as a "lead editor" or "PE". In some respects, the guiding polynucleotide contains a lead editing guiding polynucleotide (pegRNA) and is larger than the standard sgRNA typically used for CRISPR gene editing (e.g., >100 nucleotides). The pegRNA contains a primer-binding sequence (PBS) and a template containing the desired or target RNA sequence at its 3' end.

[0367] During lead editing, the PE:pegRNA complex binds to the target DNA sequence, and a modified Cas polypeptide cleaves one strand of the target DNA, creating a single-stranded overhang. PBS on the pegRNA binds to the DNA single-stranded overhang, and the target RNA sequence is reverse transcribed using reverse transcriptase. The edited strand is incorporated into the target DNA at the end of the cleaved single-stranded overhang, and the target DNA sequence is repaired with the new reverse-transcribed DNA.

[0368] Other reverse transcriptase-based genome modification systems

[0369] Another option for genome modification via the CRISPR-Cas system appears to rely on a reverse transcriptase-based approach that, compared to the non-target strand DNA described in the lead editing, reverse transcribes the desired genome edit into the complementary strand of a PAM-containing target strand DNA (i.e., the target strand). Aspects of RNA-encoded allele DNA substitution using CRISPR (designated REDRAW) are disclosed, for example, in Kim et al., bioRxiv [Biology Preprint Database] 2022.12.13.520319(2022) and US20210130835A1, the disclosures of which are incorporated herein by reference to the extent necessary for use with the CRISPR-Cas peptides disclosed herein.

[0370] Components for the expression and utilization of novel CRISPR-Cas systems in prokaryotic and eukaryotic cells

[0371] This disclosure also provides expression constructs for expressing a guide RNA / Cas system in prokaryotic or eukaryotic cells / organisms, the guide RNA / Cas system being able to recognize, bind to, and optionally cleave, unwind, or cleave all or part of a target sequence.

[0372] In some aspects, the expression constructs of this disclosure include a promoter operatively linked to a nucleotide sequence encoding a Cas gene (or a plant-optimized Cas polypeptide gene, including those described herein) and a promoter operatively linked to a guide RNA of this disclosure. The promoters are capable of driving the expression of the operatively linked nucleotide sequence in prokaryotic or eukaryotic cells / organisms.

[0373] Nucleotide sequence modifications guiding polynucleotides, VT domains, and / or CER domains may be selected from, but are not limited to, the group consisting of: 5' caps, 3' polyadenylated tails, riboswitch sequences, stability control sequences, sequences forming dsRNA duplexes, modifications or sequences that guide polynucleotide targeting to subcellular locations, modifications or sequences that provide tracking, modifications or sequences that provide protein binding sites, locked nucleic acids (LNAs), 5-methyldC nucleotides, 2,6-diaminopurine nucleotides, 2'-fluoroA nucleotides, 2'-fluoroU nucleotides; 2'-O-methylRNA nucleotides, phosphate thioester bonds, linkages to cholesterol molecules, linkages to polyethylene glycol molecules, linkages to spacer 18 molecules, 5' to 3' covalent linkages, or any combination thereof. These modifications may produce at least one additional beneficial feature, wherein the additional beneficial feature is selected from the group consisting of: modified or regulated stability, subcellular targeting, tracking, fluorescent labeling, binding sites for proteins or protein complexes, modified binding affinity to complementary target sequences, modified cellular degradation resistance, and increased cell permeability.

[0374] Methods for expressing RNA components (e.g., gRNA) in eukaryotic cells for Cas9-mediated DNA targeting have utilized the RNA polymerase III (Pol III) promoter, which allows transcription of RNA with precisely defined unmodified 5'- and 3'-termini (DiCarlo et al., Nucleic Acids Res. 41: 4336-4343; Ma et al., Mol. Ther. Nucleic Acids 3:e161). This strategy has been successfully applied in cells of several different species, including maize and soybean (US20150082478, published March 19, 2015). Methods for expressing RNA components that do not have a 5' cap have been described (WO2016 / 025131, published February 18, 2016).

[0375] Various methods and compositions can be used to obtain cells or organisms having a target polynucleotide inserted into a target site for a Cas polypeptide. Such methods may employ homologous recombination (HR) to provide integration of the target polynucleotide at the target site. In one method described herein, the target polynucleotide is introduced into an organism cell via a donor DNA construct.

[0376] The donor DNA construct further includes a first homologous region and a second homologous region located flanking the target polynucleotide. The first and second homologous regions of the donor DNA are homologous to the first and second genomic regions present in or flanking the target site in a cell or organism's genome, respectively.

[0377] Donor DNA can be ligated with guide polynucleotides. Ligated donor DNA can allow colocalization of the target and donor DNA, and can be used for genome editing, gene insertion and targeted genome regulation, and can also be used to target cells in late mitosis, in which the function of endogenous HR mechanisms is expected to be greatly reduced (Mali et al., 2013 Nature Methods, Vol. 10: 957-963).

[0378] The amount of homology or sequence identity between the target and donor polynucleotides can vary and includes the total length and / or the sequence length, ranging from approximately 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5–3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5-7 kb, 4-8 kb, etc. Regions of single integer values ​​within the range of kb, 5-10 kb, or up to and including the total length of the target site. These ranges include each integer within the range; for example, a range of 1-20 bp includes 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, and 20 bp. The amount of homology can also be described by percentage sequence identity over the complete alignment length of two polynucleotides, which includes at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98% to 99%, 99%, 99% to 100% or 100% percentage sequence identity. Sufficient homology includes any combination of polynucleotide length, overall percentage sequence identity, and optionally conserved regions or local percentage sequence identity of continuous nucleotides. For example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity with a region of the target locus.Sufficient homology can also be described by the predicted ability of two polynucleotides to hybridize specifically under highly stringent conditions, see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., eds. (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, (Elsevier, NY).

[0379] The structural similarity between a given genomic region and a corresponding homologous region found on the donor DNA can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of homology or sequence identity between a “homologous region” of the donor DNA and a “genomic region” of the organism’s genome can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, allowing the sequence to undergo homologous recombination.

[0380] Homologous regions on the donor DNA can be homologous to any sequence flanking the target site. While in some cases homologous regions show significant sequence homology to genomic sequences immediately flanking the target site, it should be recognized that homologous regions can be engineered to have sufficient homology to regions that may be 5' or 3' further away from the target site. Homologous regions can also be homologous to fragments of the target site and downstream genomic regions.

[0381] In one aspect, the first homologous region further includes a first fragment in the target site, and the second homologous region includes a second fragment in the target site, wherein the first fragment and the second fragment are different.

[0382] Target polynucleotide

[0383] This paper further describes the target polynucleotides, including those reflecting the commercial markets and interests involved in crop development. Target crops and markets are changing, and new crops and technologies will emerge as developing countries open up international markets. Furthermore, as our understanding of agronomic traits and characteristics, such as increased yield and heterosis, deepens, the selection of genes for genetic engineering will change accordingly.

[0384] General categories of target polynucleotides include, for example, those target genes involved in information (e.g., zinc fingers), those involved in communication (e.g., kinases), and those involved in housekeeping (e.g., heat shock proteins). More specific target polynucleotides include, but are not limited to, genes involved in traits of agronomic importance, such as, but not limited to: genes for crop yield, grain quality, crop nutrient composition, starch and carbohydrate quality and quantity; those genes that affect grain size, sucrose load, protein quality and quantity, nitrogen fixation and / or nitrogen use, fatty acid and oil composition; genes encoding proteins that confer resistance to abiotic stresses (e.g., drought, nitrogen, temperature, salinity, toxic metals, or trace elements), or those that confer resistance to toxins (e.g., pesticides and herbicides); and genes encoding proteins that confer resistance to biotic stresses (e.g., attacks by fungi, viruses, bacteria, insects, and nematodes, and the development of diseases associated with these organisms).

[0385] In addition to traditional breeding methods, agronomically important traits (such as oil, starch, and protein content) can be altered genetically. Modifications include increasing oleic acid, saturated and unsaturated oil content, increasing lysine and sulfur levels, providing essential amino acids, and modifying starch. Protein modifications of hordothionin are described in U.S. Patent Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389.

[0386] The target polynucleotide sequence can encode proteins involved in providing resistance to disease or pests. "Disease resistance" or "pest resistance" is intended to refer to the plant's avoidance of harmful symptoms as a consequence of plant-pathogen interactions. Pest resistance genes can encode resistance to pests that severely impact yield, such as rootworms, root-cutting moths, and European corn borers. Examples of useful gene products include disease resistance genes and insect resistance genes, such as lysozymes or silkworm-killing peptides for antibacterial protection, or proteins for antifungal protection, such as defensins, glucanases, or chitinases, or Bacillus thuringiensis endotoxins, protease inhibitors, collagenases, lectins, or glycosidases for controlling nematodes or insects. Genes encoding disease resistance traits include detoxification genes, such as those against fumonisin (US Patent No. 5,792,931); avirulence (avr) and disease resistance (R) genes (Jones et al. (1994) Science 266:789; Martin et al. (1993) Science 262:1432; and Mindrinos et al. (1994) Cell 78:1089); etc. Insect resistance genes can encode resistance to pests that severely impact yields, such as rootworms, root-cutting moths, and the European corn borer. Such genes include, for example, the Bacillus thuringiensis toxic protein gene (US Patent Nos. 5,366,892; 5,747,450; 5,736,514; 5,723,756; 5,593,881; and Geiser et al. (1986) Gene [gene] 48:109), etc.

[0387] "Herbicide resistance proteins" or proteins expressed from "herbicide resistance-encoding nucleic acid molecules" include proteins that confer tolerance to higher concentrations of herbicides compared to cells that do not express the protein, or that confer tolerance to a certain concentration of herbicide for a longer period of time compared to cells that do not express the protein. Herbicide resistance traits can be introduced into plants through genes encoding resistance to herbicides that inhibit acetolactate synthase (ALS, also known as acetylhydroxy acid synthase, AHAS) (especially sulfonylurea herbicides), genes encoding resistance to herbicides that inhibit glutamine synthase (e.g., glufosinate or basta) (e.g., the bar gene), genes encoding resistance to glyphosate (e.g., EPSP synthase genes and GAT genes), genes encoding resistance to HPPD inhibitors (e.g., the HPPD gene), or other such genes known in the art. See, for example, U.S. Patent Nos. 7,626,077, 5,310,667, 5,866,775, 6,225,114, 6,248,876, 7,169,970, 6,867,293, and 9,187,762. The bar gene encodes resistance to the herbicide basta, the nptII gene encodes resistance to the antibiotics kanamycin and genimycin, and the ALS- gene mutant encodes resistance to the herbicide chlorsulfuron.

[0388] Furthermore, it is recognized that the target polynucleotide may also include an antisense sequence complementary to at least a portion of the messenger RNA (mRNA) of the targeted gene sequence. The antisense nucleotide is constructed to hybridize with the corresponding mRNA. The antisense sequence can be modified as long as it hybridizes with the corresponding mRNA and interferes with the expression of the corresponding mRNA. In this manner, antisense constructs having 70%, 80%, or 85% sequence identity with the corresponding antisense sequence can be used. Furthermore, a portion of the antisense nucleotide can be used to disrupt the expression of the target gene. Typically, sequences of at least 50 nucleotides, 100 nucleotides, 200 nucleotides, or more nucleotides can be used.

[0389] Furthermore, the target polynucleotide can also be used in a sense-oriented manner to suppress the expression of endogenous genes in plants. Methods for using polynucleotides in a sense-oriented manner to suppress gene expression in plants are known in the art. These methods typically involve transforming plants with a DNA construct containing a promoter operatively linked to at least a portion of the nucleotide sequence of a transcript corresponding to an endogenous gene, driving expression in the plant. Typically, such a nucleotide sequence has substantial sequence identity with the sequence of the transcript of the endogenous gene, typically greater than about 65%, about 85%, or greater than about 95%. See U.S. Patent Nos. 5,283,184 and 5,034,323.

[0390] The target polynucleotide can also be a phenotypic biomarker. A phenotypic biomarker is a screenable or selectable biomarker, which includes both visual and selectable biomarkers, whether positively or negatively selectable. Any phenotypic biomarker can be used. In particular, selectable or screenable biomarkers contain a DNA segment that allows for the identification of molecules or cells containing the biomarker under specific conditions, or for selection of molecules or cells containing the biomarker. These biomarkers may encode activities, such as, but not limited to, the production of RNA, peptides, or proteins, or may provide binding sites for RNA, peptides, proteins, inorganic and organic compounds or compositions thereof.

[0391] Examples of optional markers include, but are not limited to, DNA segments containing restriction endonuclease sites; DNA segments encoding products that provide resistance to other toxic compounds, including antibiotics such as spectinomycin, ampicillin, kanamycin, tetracycline, Basta, neomycin phosphotransferase II (NEO), and hygromycin phosphotransferase (HPT); DNA segments encoding products that are not present in the receiving somatic cells themselves (e.g., tRNA genes, auxotrophic markers); DNA segments encoding easily identifiable products (e.g., phenotypic markers such as β-galactosidase, GUS; fluorescent proteins such as green fluorescent protein (GFP), cyan (CFP), yellow (YFP), red (RFP), and cell surface proteins); generating new primer sites for PCR (e.g., juxtaposition of two previously non-juxtaposed DNA sequences); DNA sequences that are inactive or active by restriction endonucleases or other DNA-modifying enzymes, chemicals, etc.; and DNA sequences that contain specific modifications (e.g., methylation) that allow for their identification.

[0392] Other alternative markers include genes conferring resistance to herbicide compounds such as sulfonylureas, glufosinate, bromosulfuron, imidazolinones, and 2,4-dichlorophenoxyacetate (2,4-D). See, for example, genes for resistance to sulfonylureas, imidazolinones, triazolopyrimidine sulfonamides, pyrimidine salicylic acid, and sulfonylaminocarbonyl-triazolinones (Shaner and Singh, 1997, Herbicide Activity: Toxicol Biochem Mol Biol [Herbicide Activity: Toxicology, Biochemistry, Molecular Biology] 69-110); glyphosate-resistant 5-enolacetone-shikimic acid-3-phosphate (EPSPS) (Saroha et al., 1998, J. Plant Biochemistry & Biotechnology [Journal of Plant Biochemistry & Biotechnology] Vol. 7: 65-72); acetolactate synthase (ALS).

[0393] The target polynucleotide includes genes stacked or used in combination with other traits, such as, but not limited to, herbicide resistance or any other trait described herein. Target polynucleotides and / or traits may be stacked together in complex trait loci, as described in US20130263324, published October 3, 2013, and WO / 2013 / 112686, published August 1, 2013.

[0394] The target polypeptide includes any protein or polypeptide encoded by the target polynucleotide described herein.

[0395] Further, methods are provided for identifying at least one plant cell containing a target polynucleotide integrated at a target site in its genome. A variety of methods can be used to identify those plant cells having an insertion at or near a target site in the genome. Such methods can be considered as directly analyzing the target sequence to detect any changes in the target sequence, including but not limited to PCR methods, sequencing methods, nuclease digestion, DNA blotting, and any combination thereof. See, for example, US20090133152, published May 21, 2009. The method also includes reproducing the plant from a plant cell containing a target polynucleotide integrated into its genome. The plant can be sterile or fertile. It should be appreciated that any target polynucleotide can be provided, integrated into a target site in the plant genome, and expressed in the plant.

[0396] Optimization of sequences for expression in plants

[0397] Methods for synthesizing plant-preferred genes are available in the art. See, for example, U.S. Patent Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. [Nucleic Acid Research] 17:477-498. Other sequence modifications are known to enhance gene expression in plant hosts. These include, for example, the elimination of: one or more sequences encoding pseudopolyadenylation signals, one or more exon-intron splicing site signals, one or more transposon-like repeats, and other well-characterized sequences that may be detrimental to gene expression. The GC content of a sequence can be adjusted to the average level for a given plant host, as calculated by referencing known genes expressed in host plant cells. When possible, sequences are modified to avoid one or more predicted hairpin secondary mRNA structures. Therefore, the “plant-optimized nucleotide sequences” of this disclosure include one or more such sequence modifications.

[0398] Expression elements

[0399] Any polynucleotide encoding the Cas peptide or other CRISPR system components disclosed herein can be functionally linked to a heterologous expression element to promote transcription or regulation in a host cell. Such expression elements include, but are not limited to, promoters, precursors, introns, and terminators. Expression elements can be “minimal” – meaning a shorter sequence derived from a natural source that still functions as an expression regulator or modifier. Alternatively, expression elements can be “optimized” – meaning their polynucleotide sequence has been altered from its natural state to function with more desired characteristics in a particular host cell (e.g., but not limited to, “corn-optimized” bacterial promoters to improve their expression in maize plants). Alternatively, expression elements can be “synthetic” – meaning they are computer-designed and synthesized for use in a host cell. Synthetic expression elements can be fully synthetic or partially synthetic (comprising fragments of naturally occurring polynucleotide sequences).

[0400] It has been shown that some promoters can direct RNA synthesis at a higher rate than others. These are called "strong promoters." Some other promoters have been shown to direct RNA synthesis at a higher level only in specific types of cells or tissues, and if a promoter preferentially directs RNA synthesis in some tissues but also directs RNA synthesis at a lower level in others, it is generally called a "tissue-specific promoter" or "tissue-biased promoter."

[0401] Plant promoters include promoters that can initiate transcription in plant cells. For a review of plant promoters, see Potenza et al., 2004 In vitro Cell Dev Biol [In vitro cell and developmental biology] 40:1-22; Porto et al., 2014, Molecular Biotechnology [Molecular Biotechnology] (2014), 56(1), 38-49.

[0402] Constitutive promoters include, for example, the core CaMV 35S promoter (Odell et al., (1985) Nature 313:810-2); rice actin (McElroy et al., (1990) Plant Cell 2:163-71); ubiquitin (Christensen et al., (1989) Plant Mol Biol 12:619-32); and the ALS promoter (US Patent No. 5,659,026), etc.

[0403] Tissue-biased promoters can be used to target enhanced expression in specific plant tissues. Organization-preferred promoters include, for example, WO 2013103367 published on July 11, 2013; Kawamata et al., (1997) Plant Cell Physiol [Plant Cell Physiology] 38:792-803; Hansen et al., (1997) Mol Gen Genet [Molecular and General Genetics] 254:337-43; Russell et al., (1997) Transgenic Research [Transgenic Research] 6:157-68; Rinehart et al., (1996) Plant Physiol [Plant Physiology] 112:1331-41; Van Camp et al., (1996) Plant Physiol. [Plant Physiology] 112:525-35; Canevascini et al., (1996) Plant Physiol. [Plant Physiology] 112:513-524; Lam, (1994) Results Probl Cell Differ 20:181-96; and Guevara-Garcia et al., (1993) Plant J 4:495-505. Leaf-preferred promoters include, for example, Yamamoto et al., (1997) Plant J [Plant Journal] 12:255-65; Kwon et al., (1994) Plant Physiol [Plant Physiology] 105:357-67; Yamamoto et al., (1994) Plant Cell Physiol [Plant Cell Physiology] 35:773-8; Gotor et al., (1993) Plant J [Plant Journal] 3:509-18; Orozco et al., (1993) Plant Mol Biol [Plant Molecular Biology] 23:1129-38; Matsuoka et al., (1993) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 90:9586-90; Simpson et al., (1958) EMBO J [Journal of the European Society for Molecular Biology] 4:2723-9; Timko et al., (1988) Nature 318:57-8.Root-specific promoters include, for example, Hire et al., (1992) Plant Mol Biol 20:207-18 (soybean root-specific glutamine synthase gene); Miao et al., (1991) Plant Cell 3:11-22 (cytoplasmic glutamine synthase (GS)); Keller and Baumgartner, (1991) Plant Cell 3:1051-61 (root-specific control element in the GRP 1.8 gene of common bean); Sanger et al., (1990) Plant Mol Biol 14:433-43 (root-specific promoter of mannosine synthase (MAS) of Agrobacterium tumefaciens); Bogusz et al., (1990) Plant Cell 2:633-41 (Root-specific promoters isolated from Parasponia and Trema tomentosa, Ulmaceae); Leach and Aoyagi, (1991) Plant Sci 79:69-76 (Root-inducible genes of Agrobacterium rhizogenes rolC and rolD); Teeri et al., (1989) EMBO J 8:343-50 (Agrobacterium wound-induced TR1' and TR2' genes); VfENOD-GRP3 gene promoter (Kuster et al., (1995) Plant Mol Biol 29:759-72); and rolB promoter (Capana et al., (1994) Plant Mol Biol 25:681-91); bean globulin gene (Murai et al., (1983) Science 23:476-82; Sengopta-Gopalen et al., (1988) Proc. Natl. Acad. Sci. USA 82:3320-4). See also U.S. Patent Nos. 5,837,876; 5,750,386; 5,633,363; 5,459,252; 5,401,836; 5,110,732 and 5,023,179.

[0404] Seed-preferred promoters include both seed-specific promoters active during seed development and seed-germinating promoters active during seed germination. See Thompson et al., (1989) BioEssays [Biological Analysis] 10:108. Seed-preferred promoters include, but are not limited to, Cim1 (cytokinin-induced messenger); cZ19B1 (19 kDa corn gliadin); and milps (inositol-1-phosphate synthase); and, for example, those disclosed in WO2000011177 (published March 2, 2000) and U.S. Patent 6,225,529. For dicotyledonous plants, seed-preferred promoters include, but are not limited to, beta-coumarin, rapeseed protein, β-conglycinin, soybean lectin, cruciferous proteins, etc. For monocotyledons, seed-preferred promoters include, but are not limited to, 15 kDa zeatin, 22 kDa zeatin, 27 kDa γ-zeatin, waxes, thymol 1, thymol 2, globulin 1, olein, and nuc1. See also WO 2000012733, published March 9, 2000, which discloses seed-preferred promoters from the END1 and END2 genes.

[0405] Chemically inducible (regulatory) promoters can be used to regulate gene expression in prokaryotic and eukaryotic cells or organisms by applying exogenous chemical regulators. The promoter can be a chemically inducible promoter when gene expression is induced by chemicals, or a chemically repressive promoter when gene expression is repressed by chemicals. Chemically induced promoters include, but are not limited to: the corn In2-2 promoter activated by benzenesulfonamide herbicide safener (De Veylder et al., (1997) Plant Cell Physiol 38:568-77), the corn GST promoter activated by a hydrophobic electrophilic compound used as a pre-emergence herbicide (GST-II-27, WO 1993001294 published on January 21, 1993), and the tobacco PR-1a promoter activated by salicylic acid (Ono et al., (2004) Biosci Biotechnol Biochem 68:803-7). Other chemical-regulated promoters include steroid-responsive promoters (see, for example, glucocorticoid-inducible promoters (Schena et al., (1991) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 88:10421-5; McNellis et al., (1998) Plant J [Journal of Plant Science] 14:247-257); tetracycline-inducible promoters and tetracycline-repressive promoters (Gatz et al., (1991) Mol Gen Genet [Molecular and General Genetics] 227:229-37; US Patent Nos. 5,814,618 and 5,789,156)).

[0406] Pathogen-inducible promoters induced after infection by pathogens include, but are not limited to, promoters that regulate the expression of PR protein, SAR protein, β-1,3-glucanase, chitinase, etc.

[0407] Stress-induced promoters include the RD29A promoter (Kasuga et al. (1999) Nature Biotechnol. 17:287-91). Those skilled in the art are familiar with procedures for simulating stress conditions (such as drought, osmotic stress, salt stress, and temperature stress) and evaluating the stress tolerance of plants that have been subjected to simulated or naturally occurring stress conditions.

[0408] Another example of an inducible promoter useful in plant cells is the ZmCAS1 promoter, described in US20130312137, published on November 21, 2013.

[0409] New promoters of different types that are useful in plant cells are constantly being discovered; many examples can be found in the compilation of Okamuro and Goldberg, (1989) The Biochemistry of Plants, Vol. 115, edited by Stumpf and Conn (New York, NY: Academic Press), pp. 1-82.

[0410] Gene Targeting

[0411] The guided polynucleotide / Cas system described in this article can be used for gene targeting.

[0412] Generally, DNA targeting can be achieved by cleaving one or two strands at specific polynucleotide sequences in cells that have Cas peptides associated with suitable polynucleotide components. Once a single-strand or double-strand break is induced in the DNA, the cell's DNA repair mechanisms are activated to repair the break via non-homologous end joining (NHEJ) or homologous directed repair (HDR) processes that lead to modifications at the target site.

[0413] The length of the DNA sequence at the target site can vary and includes target sites of, for example, at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length. It is also possible for the target site to be palindromic, i.e., the sequence on one strand is identical when read in the opposite direction on the complementary strand. The nick / cut site can be inside or outside the target sequence. In another variation, the cut can occur at exactly opposite nucleotide positions to produce a blunt-end cut, or in other cases, the cuts can be staggered to produce single-stranded overhangs, also known as “sticky ends,” which can be 5’ or 3’ overhangs. Active variants of genomic target sites can also be used. Such active variants may contain at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with a given target site, wherein these active variants retain biological activity and are therefore able to be recognized and cleaved by Cas peptides.

[0414] Measurements of single-strand or double-strand breaks at target sites caused by endonucleases are known in the art and are typically used to measure the overall activity and specificity of a drug on a DNA substrate containing a recognition site.

[0415] The targeting method described herein can be performed, for example, by targeting two or more DNA target sites. This method can optionally be characterized as a multiplexing method. In some respects, two, three, four, five, six, seven, eight, nine, ten or more target sites can be targeted simultaneously. Multiplexing methods are typically performed through the targeting method described herein, in which multiple different RNA components are provided, each designed to guide the guide polynucleotide / Cas polypeptide complex to a unique DNA target site.

[0416] Gene editing

[0417] The process of editing a genome sequence by combining a DSB and a modification template typically involves: introducing a DSB inducer or a nucleic acid encoding a DSB inducer (which recognizes a target sequence in the chromosomal sequence and is capable of inducing a DSB in the genome sequence) into a host cell, and at least one polynucleotide modification template containing at least one nucleotide change compared to the nucleotide sequence to be edited. The polynucleotide modification template may further contain a nucleotide sequence flanking the at least one nucleotide change, wherein the flanking sequence is substantially homologous to the chromosomal region flanking the DSB. Genome editing using DSB inducers (such as Cas-gRNA complexes) has been described, for example, in: US20150082478, published March 19, 2015; WO 2015026886, published February 26, 2015; WO 2016007347, published January 14, 2016; and WO / 2016 / 025131, published February 18, 2016.

[0418] Some uses of the RNA / Cas peptide system have been described (see, for example, US20150082478 A1, published March 19, 2015; WO 2015026886, published February 26, 2015; and US20150059010, published February 26, 2015) and include, but are not limited to, modification or substitution of target nucleotide sequences (such as regulatory elements), insertion of target polynucleotides, gene knockout, gene knock-in, modification of splice sites and / or introduction of alternative splice sites, modification of nucleotide sequences encoding target proteins, amino acid and / or protein fusion, and gene silencing induced by expression of inverted repeat sequences in target genes.

[0419] Proteins can be altered in various ways, including by amino acid substitution, deletion, truncation, and insertion. Methods for such manipulations are generally known. For example, amino acid sequence variants of one or more proteins can be prepared by mutations in the DNA. Methods for mutagenesis and nucleotide sequence alteration include, for example, Kunkel, (1985) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 82:488-92; Kunkel et al., (1987) Meth Enzymol [Enzymological Methods] 154:367-82; U.S. Patent No. 4,873,192; Walker and Gaastra, ed. (1983) Techniques in Molecular Biology [Techniques in Molecular Biology] (MacMillan Publishing Company, New York), and the references cited therein. Guidelines have been found regarding amino acid substitutions that are unlikely to affect the biological activity of proteins, for example, in the model of Dayhoff et al., (1978) Atlas of Protein Sequence and Structure (Natl Biomed Res Found, Washington, DC). Conserved substitutions, such as exchanging one amino acid with another that has similar properties, are preferred. Conserved deletions, insertions, and amino acid substitutions are expected not to produce fundamental changes in protein characteristics, and the effects of any substitution, deletion, insertion, or combination thereof can be evaluated by routine screening assays. Assays of double-strand break-inducible activity are known and are typically used to measure the overall activity and specificity of an agent against a DNA substrate containing a target site.

[0420] This article describes a method for genome editing using Cas peptides and complexes of Cas peptides and guide polynucleotides. After characterization of the guide RNA and PAM sequences, chromosomal DNA in other organisms, including plants, can be modified using components such as endonucleases and associated CRISPR RNA (crRNA). To facilitate optimal expression and nuclear localization (for eukaryotic cells), the gene-containing complex can be optimized as described in WO 2016186953, published November 24, 2016, and then delivered to cells as a DNA expression cassette using methods known in the art. Alternatively, the components necessary for the active complex can be delivered as RNA (with or without modifications to protect RNA from degradation) or as capped or uncapped mRNA (Zhang, Y. et al., 2016, Nat. Commun. [Nature Communications] 7:12617) or a Cas peptide-guided polynucleotide complex (published April 27, 2017, WO 2017070032), or any combination thereof. Furthermore, one or more portions of the complex and crRNA can be expressed from the DNA construct, while other components are delivered as RNA (with or without modifications to protect RNA from degradation) or as capped or uncapped mRNA (Zhang et al. 2016 Nat. Commun. [Nature Communications] 7:12617) or Cas peptide-guided polynucleotide complexes (disclosed in WO 2017070032, April 27, 2017), or any combination thereof. For in vivo production of crRNA, tRNA-derived elements can also be used to recruit endogenous RNases to cleave the crRNA transcript into a mature form capable of guiding the complex to its DNA target site, for example, as described in WO 2017105991, June 22, 2017. The nicking enzyme complex can be used alone or in combination to generate single or multiple DNA nicks on one or two DNA strands. Furthermore, the cleavage activity of Cas endonucleases can be inactivated by altering key catalytic residues in the cleavage domain (Sinkunas, T. et al., 2013, EMBO J. [Journal of the European Society for Molecular Biology] 32:385-394), thereby generating RNA-directed helicases that can be used to enhance homology-directed repair, induce transcriptional activation, or remodel local DNA structures. Additionally, the activity of both the Cas cleavage and helicase domains can be knocked out and used in combination with other DNA cleavage, DNA nicking, DNA binding, transcriptional activation, transcriptional repression, DNA remodeling, DNA deamination, DNA unwinding, DNA recombination enhancement, DNA integration, DNA inversion, and DNA repair agents.

[0421] The transcriptional direction of tracrRNA for the CRISPR-Cas system (if present) and other components of the CRISPR-Cas system (e.g., variable target domains, crRNA repeat sequences, loops, anti-repetitive sequences) can be deduced as described in WO 2016186946 and WO2016186953 published on November 24, 2016.

[0422] As described herein, once appropriate guide RNA requirements are established, the PAM preference of each novel system disclosed herein can be examined. If cleavage of the complex leads to the degradation of the randomized PAM library, the complex can be converted into a cleavage enzyme by mutagenesis of key residues or by inactivating ATPase-dependent helicase activity through an assembly reaction in the absence of ATP, as previously described (Sinkunas, T. et al., 2013, EMBO J. [Journal of the European Society for Molecular Biology] 32:385-394). Double-stranded DNA breaks can be generated using two regions of PAM randomization separated by two anterior spacer targets, which can be captured and sequenced to examine the PAM sequences supporting cleavage of their respective complexes.

[0423] In some aspects, methods for modifying target sites in the genome of a cell include introducing at least one PGEN as described herein into a cell and identifying at least one cell having a modification at the target site, wherein the modification at the target site is selected from the group consisting of: (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, chemical alteration of at least one nucleotide, and (v) any combination of (i)–(iv).

[0424] The nucleotide to be edited can be located inside or outside the target site recognized and cleaved by the Cas endonuclease. On the one hand, at least one nucleotide modification is not a modification at the target site recognized and cleaved by the Cas endonuclease. On the other hand, there are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 900, or 1000 nucleotides between the at least one nucleotide to be edited and the genomic target site.

[0425] Knockouts can be generated by insertion or deletion (by inserting or deleting nucleotide bases in the target DNA sequence via NHEJ) or by specifically removing sequences that reduce or completely disrupt sequence function at or near the target site.

[0426] Guided polynucleotide / Cas endonuclease-induced targeted mutations can occur in nucleotide sequences located inside or outside the genomic target site recognized and cleaved by Cas endonucleases.

[0427] In one aspect, this disclosure describes a method for modifying target sites in the genome of a cell, the method comprising introducing at least one PGEN and at least one donor DNA as described herein into the cell, wherein the donor DNA comprises a target polynucleotide, and optionally, the method further comprises identifying at least one cell in which the target polynucleotide is integrated into or near the target site.

[0428] In some respects, the methods disclosed herein can employ homologous recombination (HR) to provide integration of the target polynucleotide at the target site.

[0429] Various methods and compositions can be used to generate cells or organisms having a target polynucleotide with an active insertion target site via the CRISPR-Cas system components described herein. In one method described herein, the target polynucleotide is introduced into an organism cell via a donor DNA construct. As used herein, “donor DNA” is a DNA construct containing the target polynucleotide at a genomic target site to be inserted into a Cas polypeptide. The donor DNA construct further includes a first homologous region and a second homologous region flanking the target polynucleotide. The first and second homologous regions of the donor DNA are homologous to the first and second genomic regions present in or flanking the target site of the cell or organism's genome, respectively.

[0430] Donor DNA can be ligated with guide polynucleotides. Ligated donor DNA can allow colocalization of the target and donor DNA, and can be used for genome editing, gene insertion and targeted genome regulation, and can also be used to target cells in late mitosis, in which the function of endogenous HR mechanisms is expected to be greatly reduced (Mali et al., 2013 Nature Methods, Vol. 10: 957-963).

[0431] The amount of homology or sequence identity between the target and donor polynucleotides can vary and includes the total length and / or the sequence length, ranging from approximately 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200-400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5–3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5-7 kb, 4-8 kb, etc. Regions of single integer values ​​within the range of kb, 5-10 kb, or up to and including the total length of the target site. These ranges include each integer within the range; for example, a range of 1-20 bp includes 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, and 20 bp. The amount of homology can also be described by percentage sequence identity over the complete alignment length of two polynucleotides, which includes approximately 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% percentage sequence identity. Sufficient homology includes any combination of polynucleotide length, overall percentage sequence identity, and optionally conserved regions or local percentage sequence identity of consecutive nucleotides. For example, sufficient homology can be described as a 75-150 bp region having at least 80% sequence identity with a region of the target locus.Sufficient homology can also be described by the predicted ability of two polynucleotides to hybridize specifically under highly stringent conditions, see, for example, Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., eds. (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, (Elsevier, NY).

[0432] Additional DNA molecules can also be ligated to double-strand breaks, for example, by integrating T-DNA into a chromosomal double-strand break (Chilton and Que, (2003) Plant Physiol 133:956-65; Salomon and Puchta, (1998) EMBO J. 17:6086-95). Genetic conversion pathways can restore the original structure once the sequence surrounding the double-strand break is altered, for example by changes in the activity of mature exonucleases involved in the break, if homologous sequences are present, such as homologous chromosomes in non-dividing somatic cells, or sister chromatids after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenetic DNA sequences can also serve as templates for DNA repair in homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0433] In one aspect, this disclosure includes a method for editing nucleotide sequences in the genome of a cell, the method comprising introducing at least one PGEN as described herein and a polynucleotide modification template, wherein the polynucleotide modification template comprises at least one nucleotide modification of the nucleotide sequence, and the method optionally further comprising selecting at least one cell containing the edited nucleotide sequence.

[0434] The guided polynucleotide / Cas endonuclease system can be used in combination with at least one polynucleotide modification template to allow editing (modification) of the nucleotide sequence of a target genome. (See also US20150082478, published March 19, 2015, and WO 2015026886, published February 26, 2015).

[0435] Target polynucleotides and / or traits can be stacked together in complex trait loci, as described in WO 2012129373, published September 27, 2012, and WO 2013112686, published August 1, 2013. The guided polynucleotide / Cas9 endonuclease system described herein provides an efficient system for generating double-strand breaks and allowing traits to be stacked in complex trait loci.

[0436] The directed polynucleotide / Cas system for gene targeting as described herein can be used in methods for guiding heterologous gene insertion and / or generating complex trait loci containing multiple heterologous genes in a manner similar to that disclosed in WO 2012129373, published September 27, 2012, wherein the directed polynucleotide / Cas system as disclosed herein is used instead of a double-strand break inducer to introduce the target gene. These transgenes can be bred as single genetic loci by inserting independent transgenes into each other at 0.1, 0.2, 0.3, 0.4, 0.5, 1.0, 2, or even 5 centimoles (cM) (e.g., see US20130263324, published October 3, 2013, or WO2012129373, published March 14, 2013). After selecting plants containing transgenes, plants containing (at least) one transgene can be crossed to form F1 containing two transgenes. Of the offspring from these F1 (F2 or BC1) generations, 1 in 500 will have two distinct transgenes recombined on the same chromosome. The complex locus can then be propagated into a single locus with both transgene traits. This process can be repeated to stack as many traits as possible.

[0437] Further uses of the RNA / Cas peptide system have been described (see, for example, US 20150082478, published March 19, 2015; WO 2015026886, published February 26, 2015; US20150059010, published February 26, 2015; WO 2016007347, published January 14, 2016; and PCT application WO2016025131, published February 18, 2016) including but not limited to modification or substitution of target nucleotide sequences (such as regulatory elements), insertion of target polynucleotides, gene knockout, gene knock-in, modification of splice sites and / or introduction of alternative splice sites, modification of nucleotide sequences encoding target proteins, amino acids and / or protein fusions, and gene silencing induced by expression of inverted repeat sequences in target genes.

[0438] The characteristics produced by the gene-editing compositions and methods described herein can be evaluated. Chromosomal regions associated with a target phenotype or trait can be identified. Various methods well known in the art can be used to identify chromosomal regions. The boundaries of such chromosomal regions are extended to encompass markers that will be linked to genes controlling the target trait. In other words, chromosomal regions are extended such that any marker located within the region (including terminal markers defining the boundaries of the region) can be used as a marker for the specific trait. On one hand, a chromosomal region contains at least one QTL, and in addition, it can indeed contain more than one QTL. Multiple QTLs that are very close together in the same region can confuse the association of a specific marker with a specific QTL, because a marker may show linkage with more than one QTL. Conversely, for example, if two very close markers show co-segregation with the desired phenotypic trait, it is sometimes difficult to distinguish whether each of those markers identifies the same QTL or two different QTLs. The term “quantitative trait locus” or “QTL” refers to a region of DNA associated with differential expression of a quantitative phenotypic trait in at least one genetic context (e.g., in at least one breeding population). The region of a QTL encompasses or is closely linked to one or more genes that influence the trait under consideration. An “allele of a QTL” can comprise multiple genes or other genetic factors, such as haplotypes, within a contiguous genomic region or linkage group. A QTL allele can represent a haplotype within a specified window, where the window is a contiguous genomic region that can be defined and tracked using a set of one or more polymorphic markers. A haplotype can be specified as a unique fingerprint defined by the allele at each marker within the window.

[0439] Introducing CRISPR-Cas system components into cells

[0440] The methods described herein are not dependent on any specific method for introducing a sequence into an organism or cell, as long as the polynucleotide or polypeptide enters the interior of at least one cell of the organism. Introduction includes reference to incorporating nucleic acids into eukaryotic or prokaryotic cells, wherein the nucleic acids may be incorporated into the cell's genome, and includes reference to the transient (direct) delivery of nucleic acids, proteins, or polynucleotide-protein complexes (PGEN, RGEN) into the cell.

[0441] Methods for introducing polynucleotides or peptides or polynucleotide-protein complexes into cells or organisms are known in the art and include, but are not limited to, microinjection, electroporation, stable transformation methods, transient transformation methods, ballistic particle acceleration (particle bombardment), whisker-mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, virus-mediated introduction, transfection, transduction, cell-penetrating peptides, mesoporous silica nanoparticles (MSN)-mediated direct protein delivery, topical application, sexual hybridization, sexual breeding, and any combination thereof.

[0442] For example, guide polynucleotides (guide RNA, crRNA + tracrRNA, guide DNA, and / or guide RNA-DNA molecules) can be introduced directly into cells (transiently) as single-stranded or double-stranded polynucleotide molecules. Guide RNA (or crRNA + tracrRNA) can also be indirectly introduced into cells by introducing a recombinant DNA molecule containing a heterologous nucleic acid fragment encoding a guide RNA (or crRNA + tracrRNA), the heterologous nucleic acid fragment being operatively linked to a specific promoter capable of transcribing the guide RNA (crRNA + tracrRNA molecule) in the cells. Specific promoters can be, but are not limited to, RNA polymerase III promoters, which allow transcription of RNA with precisely defined unmodified 5'- and 3'-termini (Ma et al., 2014, Mol. Ther. Nucleic Acids [Molecular Therapy-Nucleic Acids] 3:e161; DiCarlo et al., 2013, Nucleic Acids Res. [Nucleic Acids Research] 41:4336-4343; WO 2015026887, published February 26, 2015). Any promoter capable of transcribing guide RNA in the cell can be used, and these promoters include heat shock / heat inducible promoters operatively linked to the nucleotide sequence encoding the guide RNA.

[0443] Plant cells differ from animal cells (such as human cells), fungal cells (such as yeast cells), and protoplasts, including, for example, plant cells containing plant cell walls (which can act as a barrier for component delivery).

[0444] Delivery of Cas peptides, and / or guide RNA, and / or ribonucleoprotein complexes, and / or polynucleotides encoding any one or more of the foregoing into plant cells can be achieved by methods known in the art, such as, but not limited to: rhizobium-mediated transformation (e.g., Agrobacterium, Ochrobactrum), particle-mediated delivery (particle bombardment), polyethylene glycol (PEG)-mediated transfection (e.g., protoplasts), electroporation, cell-penetrating peptides, or direct protein delivery mediated by mesoporous silica nanoparticles (MSN).

[0445] Cas peptides, such as those described herein, can be introduced into cells by any method known in the art, either directly into the Cas peptide itself (referred to as direct delivery of the Cas peptide), into mRNA encoding the Cas protein, and / or into the guiding polynucleotide / Cas peptide complex itself. Cas peptides can also be introduced into cells indirectly by introducing a recombinant DNA molecule encoding the Cas peptide. Endonucleases can be introduced into cells transiently using any method known in the art, or endonucleases can be incorporated into the genome of a host cell. Cell-penetrating peptides (CPPs), as described in WO 2016073433 published May 12, 2016, can facilitate the uptake of endonucleases and / or guided polynucleotides into cells. Any promoter capable of expressing the Cas peptide in cells can be used, and these promoters include heat shock / heat-inducible promoters operatively linked to a nucleotide sequence encoding the Cas peptide.

[0446] Direct delivery of polynucleotide modified templates into plant cells can be achieved through particle-mediated delivery, and any other direct delivery method, such as, but not limited to, polyethylene glycol (PEG)-mediated protoplast transfection, whisker-mediated transformation, electroporation, particle bombardment, cell-penetrating peptides, or mesoporous silica nanoparticles (MSN)-mediated direct protein delivery, can be successfully used to deliver polynucleotide modified templates in eukaryotic cells, such as plant cells.

[0447] Donor DNA can be introduced by any means known in the art. Donor DNA can be provided by any transformation method known in the art, including, for example, Agrobacterium-mediated transformation or bio-projectile particle bombardment. Donor DNA can be transiently present in the cell or can be introduced via viral replicons. In the presence of Cas endonucleases and target sites, the donor DNA is inserted into the genome of the transformed plant.

[0448] Direct delivery of any of the guided Cas system components may be accompanied by direct delivery (co-delivery) of other mRNAs that can promote enrichment and / or visualization in cells receiving the guided polynucleotide / Cas polypeptide complex component. For example, direct co-delivery of the guided polynucleotide / Cas polypeptide component (and / or the guided polynucleotide / Cas polypeptide complex itself) with mRNAs encoding phenotypic markers (such as, but not limited to, transcriptional activators such as CRC (Bruce et al. 2000 The Plant Cell 12:65-79)) can enable cell selection and enrichment by restoring the function of nonfunctional gene products without the use of exogenous selectable markers, as described in WO 2017070032, published April 27, 2017.

[0449] Introducing the guide RNA / Cas polypeptide complex (representing the cleavable complex as described herein) into cells includes introducing the components of the complex individually or in combination into the cells, either directly (as RNA (for the guide) and protein (for the Cas polypeptide and protein subunits or functional fragments thereof)) or via a recombinant construct expressing these components (guide RNA, Cas polypeptide, protein subunits or functional fragments thereof). Introducing the guide RNA / Cas endonuclease complex (RGEN) into cells includes introducing the guide RNA / Cas polypeptide complex into the cells as a ribonucleotide-protein. This ribonucleotide-protein can be assembled prior to introduction into the cells as described herein. Components comprising the guide RNA / Cas endonuclease ribonucleotide protein (at least one Cas endonuclease, at least one guide RNA, and at least one protein subunit) can be assembled in vitro or by any method known in the art prior to introduction into cells (which are targeted for genomic modifications as described herein).

[0450] Direct delivery of the RGEN ribonucleoprotein allows for genome editing at target sites within the cell's genome, followed by rapid degradation of the complex, which is then allowed only transient presence within the cell. This transient presence of the RGEN complex may reduce off-target effects. In contrast, delivery of RGEN components (guide RNA, Cas9 endonuclease) via plasmid DNA sequences can result in constant expression of RGEN from these plasmids, which can enhance off-target effects (Cradick, TJ et al. (2013) Nucleic Acids Res [Nucleic Acids Research] 41:9584-9592; Fu, Y et al. (2014) Nat. Biotechnol. [Nature Biotechnology] 31:822-826).

[0451] Direct delivery can be achieved by combining any component of the guide RNA / Cas endonuclease complex (RGEN) (representing the cleavage-ready complex as described herein) (e.g., at least one guide RNA, at least one Cas polypeptide, and optionally at least one additional protein) with a delivery matrix containing microparticles (e.g., but not limited to gold particles, tungsten particles, and silicon carbide whisker particles) (see also WO 2017070032, published April 27, 2017). The delivery matrix may contain any of these components, such as a Cas endonuclease, attached to a solid matrix (e.g., particles for bombardment).

[0452] In some respects, the guide polynucleotide / Cas polypeptide complex is a complex in which the guide RNA and Cas polypeptide protein that form the guide RNA / Cas polypeptide complex are introduced into the cell as RNA and protein, respectively.

[0453] In some respects, the guide polynucleotide / Cas polypeptide complex is a complex in which the guide RNA and Cas polypeptide protein that form the guide RNA / Cas polypeptide complex and at least one protein subunit of the complex are introduced into the cell as RNA and protein, respectively.

[0454] In some respects, the guide polynucleotide / Cas endonuclease complex is a complex in which the guide RNA and Cas endonuclease protein that form the guide RNA / Cas endonuclease complex (cutting-ready complex) and at least one protein subunit of the complex are pre-assembled in vitro and introduced into the cell as a ribonucleotide-protein complex.

[0455] Protocols for introducing polynucleotides, polypeptides, or polynucleotide-protein complexes (PGEN, RGEN) into eukaryotic cells (such as plants or plant cells) are known and include microinjection (Crossway et al., (1986) Biotechniques 4:320-34 and U.S. Patent No. 6,300,543), meristematic transformation (U.S. Patent No. 5,736,369), electroporation (Riggs et al., (1986) Proc. Natl. Acad. Sci. USA 83:5602-6), Agrobacterium-mediated transformation (U.S. Patent Nos. 5,563,055 and 5,981,840), whisker-mediated transformation (Ainley et al. 2013, Plant Biotechnology Journal 11:1126-1134; Shaheen A. and M. Arshad 2011 Properties and Applications of SiliconCarbide). [Properties and Applications of Silicon Carbide] (2011), 345-358. Edited by Gerhardt and Rosario. Published by InTech, Rijeka, Croatia.CODEN:69PQBP; ISBN:978-953-307-201-2), direct gene transfer (Paszkowski et al., (1984) EMBO J [Journal of the European Society for Molecular Biology] 3:2717-22), and ballistic particle acceleration (US Patent Nos. 4,945,050; 5,879,918; 5,886,244; 5,932,782; Tomes et al., (1995) “Direct DNA Transfer into Intact Plant Cells via Microprojectile Bombardment” in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, edited by Gamborg and Phillips (Springer-Verlag, Berlin); McCabe et al., (1988) Biotechnology. 6:923-6; Weissinger et al., (1988) Ann Rev Genet [Genetics Annals] 22:421-77; Sanford et al., (1987) Particulate Science and Technology [Particulate Science and Technology] 5:27-37 (Onion); Christou et al., (1988) Plant Physiology [Plant Physiology] 87:671-4 (Soybean); Finer and McMullen, (1991) In vitro Cell Dev Biol [In vitro Cell Biology and Developmental Biology] 27P:175-82 (Soybean); Singh et al., (1998) Theor Appl Genet [Theoretical and Applied Genetics] 96:319-24 (Soybean); Datta et al., (1990) Biotechnology [Biotechnology] 8:736-40 (Rice); Klein et al., (1988) Proc. Natl. Acad. Sci.USA [Proceedings of the National Academy of Sciences] 85:4305-9 (Corn); Klein et al., (1988) Biotechnology 6:559-63 (Corn); US Patent Nos. 5,240,855; 5,322,783 and 5,324,646; Klein et al., (1988) Plant Physiol 91:440-4 (Corn); Fromm et al., (1990) Biotechnology 8:833-9 (Corn); Hooykaas-Van Slogteren et al., (1984) Nature 311:763-4; US Patent No. 5,736,369 (Cereals); Bytebier et al., (1987) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 84:5345-9 (Liliaceae); De Wet et al., (1985) The Experimental Manipulation of Ovule Tissues, edited by Chapman et al., (Longman, New York), pp. 197-209 (pollen); Kaeppler et al., (1990) Plant Cell Reports 9:415-8 and Kaeppler et al., (1992) Theor Appl Genet 84:560-6 (whisker-mediated transformation); D'Halluin et al., (1992) Plant Cell 4:1495-505 (electroporation); Li et al., (1993) Plant Cell Reports 12:250-5; Christou and Ford (1995) Annals Botany [Annals of Botany] 75:407-13 (rice) and Osjoda et al., (1996) Nat Biotechnol 14:745-50 (maize transformed by Agrobacterium tumefaciens).

[0456] Alternatively, polynucleotides can be introduced into plants or plant cells by contacting cells or organisms with viruses or viral nucleic acids. Typically, such methods involve incorporating the polynucleotide into viral DNA or RNA molecules. In some instances, the target polypeptide can be initially synthesized as part of a viral polyprotein, which is then processed in vivo or in vitro via proteolytic hydrolysis to produce the desired recombinant protein. Methods for introducing polynucleotides into plants and expressing proteins encoded therein (involving viral DNA or RNA molecules) are known, see, for example, U.S. Patent Nos. 5,889,191, 5,889,190, 5,866,785, 5,589,367, and 5,316,931.

[0457] A variety of transient transformation methods can be used to provide or introduce polynucleotides or recombinant DNA constructs into prokaryotic and eukaryotic cells or organisms. These transient transformation methods include, but are not limited to, the direct introduction of polynucleotide constructs into plants.

[0458] Nucleic acids and proteins can be delivered to cells by any method, including methods that use molecules to facilitate the uptake of any or all components of a guided Cas system (proteins and / or nucleic acids), such as cell-penetrating peptides and nanocarriers. See also US20110035836, published February 10, 2011, and EP2821486A1, published January 7, 2015.

[0459] Other methods for introducing polynucleotides into prokaryotic and eukaryotic cells or organisms or plant parts can be used, including plastid transformation methods and methods for introducing polynucleotides into tissues from seedlings or mature seeds.

[0460] "Stable transformation" refers to the integration of a nucleotide construct introduced into an organism into its genome and its ability to be inherited by its offspring. "Transient transformation" refers to the introduction of a polynucleotide into an organism that does not integrate into its genome, or the introduction of a polypeptide into the organism. Transient transformation indicates that the introduced composition is only temporarily expressed or present in the organism.

[0461] A variety of methods can be used to identify cells with altered genomes at or near the target site without using selectable biomarker phenotypes. Such methods can be considered as directly analyzing the target sequence to detect any changes in the target sequence, including but not limited to PCR methods, sequencing methods, nuclease digestion, DNA blotting, and any combination thereof.

[0462] Cells and plants

[0463] The polynucleotides and polypeptides disclosed herein can be introduced into cells. Cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, fungal, insect, yeast, unconventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein. Any plant (including monocots and dicots, and plant elements) can be used with the compositions and methods described herein.

[0464] Examples of monocotyledonous plants that can be used include, but are not limited to, maize (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet, Pennisetum glaucum), foxtail millet (Panicum miliaceum), foxtail millet (Setaria italica), wheat (Eleusine coracana), and wheat (species of the genus Triticum, such as Triticum aestivum, Triticum). Monococcum), sugarcane (species of the genus Saccharum spp.), oats (species of the genus Avena), barley (species of the genus Hordeum), switchgrass (species of the genus Panicum virgatum), pineapple (species of the genus Ananas comosus), banana (species of the genus Musa spp.), palms, ornamental plants, turfgrass, and other grasses.

[0465] Examples of dicotyledonous plants that may be used include, but are not limited to, soybean (Glycine max), Brassica species (e.g., but not limited to: rapeseed or canola rapeseed) (European rapeseed (Brassica napus) and Chinese rapeseed (B. campestris), turnip (Brassica rapa), mustard (Brassica juncea)), alfalfa (Medicago sativa)), tobacco (Nicotiana tabacum)), Arabidopsis (Arabidopsis), sunflower (Helianthus annuus)), cotton (Gossypium arboreum, Gossypium barbadense)), and peanut (Arachis hypogaea)), tomato (Solanum lycopersicum)), potato (Solanum tuberosum) etc.

[0466] Other plants that can be used include safflower (Carthamus tinctorius), sweet potato (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), citrus (Citrus spp.), cocoa (Theobroma cacao), tea (Camelliasinensis), banana (Musa spp.), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), and olive (Olea oleifera). Europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beets (Beta vulgaris), vegetables, ornamental plants and conifers.

[0467] Vegetables that can be used include tomatoes, lettuce (e.g., Lactucasativa), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus spp.), and members of the Cucumber genus such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and musk melon (C. melo). Ornamental plants include azaleas (Rhododendron spp.), hydrangeas (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnations (Dianthus caryophyllus), poinsettias (Euphorbia pulcherrima), and chrysanthemums.

[0468] Suitable coniferous trees include pine trees such as loblolly pine (Pinustaeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsugamenziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); redwood (Sequoia sempervirens); and true fir trees, such as silver fir (Abies spp.). amabilis) and balsam fir (Abies balsamea); as well as cedar, such as western red cedar (Thuja plicata) and Alaskan yellow cedar (Chamaecyparis nootkatensis).

[0469] In some aspects of this disclosure, fertile plants are those that produce viable male and female gametes and are self-fertile. Such self-fertile plants can produce offspring without the contribution of gametes from any other plant or the genetic material contained therein. Other aspects of this disclosure may involve the use of non-self-fertile plants, since these plants do not produce viable or otherwise fertile male or female gametes, or both.

[0470] This disclosure can be used for breeding plants that contain one or more introduced traits or edited genomes.

[0471] The following describes a non-restrictive example of how two traits can be stacked in the genome at a genetic distance of, for example, 5 cM from each other: A first plant containing a first transgenic target site integrated into a first DSB target site within a genomic window and lacking a first target genomic locus is crossed with a second transgenic plant containing a target genomic locus at a different genomic insertion site within the genomic window, and the second plant does not contain the first transgenic target site. Approximately 5% of the progeny from this cross will have the first transgenic target site integrated into the first DSB target site and the first target genomic locus integrated at a different genomic insertion site within the genomic window. Progeny plants with two sites within the defined genomic window can be further crossed with a third transgenic plant containing a second transgenic target site integrated into a second DSB target site and / or a second target genomic locus within the defined genomic window and lacking the first transgenic target site and the first target genomic locus. Progeny with the first transgenic target site, the first target genomic locus, and the second target genomic locus integrated at different genomic insertion sites within the genomic window are then selected. This method can be used to generate transgenic plants containing complex trait loci that have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 or more transgenic target sites integrated into DSB target sites and / or target genomic loci integrated into different sites within a genomic window. In this way, various complex trait loci can be generated.

[0472] Cells and Animals

[0473] The polynucleotides and polypeptides disclosed herein can be introduced into animal cells. Animal cells can include, but are not limited to, organisms belonging to the following phyla: Chordata, Arthropoda, Molluscs, Annelids, Cnidaria, or Echinodermata; and organisms belonging to the following classes: mammals, insects, birds, amphibians, reptiles, or fish. In some respects, the animal is a human, mouse, *C. elegans*, rat, fruit fly (*Drosophila* species), zebrafish, chicken, dog, cat, guinea pig, hamster, Japanese rice fish, lamprey, pufferfish, tree frog (e.g., *Xenopus* species), monkey, or chimpanzee. The specific cell types anticipated include haploid cells, diploid cells, germ cells, neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, embryonic cells, hematopoietic cells, bone cells, germ cells, somatic cells, stem cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some respects, multiple cells from an organism may be used.

[0474] The Cas peptides disclosed herein can be used to edit the genome of animal cells in various ways. In some respects, it may be desirable to delete one or more nucleotides. In other respects, it may be desirable to insert one or more nucleotides. In some respects, it may be desirable to substitute one or more nucleotides. In other respects, it may be desirable to modify one or more nucleotides through covalent or non-covalent interactions with another atom or molecule.

[0475] Genomic modifications via Cas peptides can be used to achieve genotypic and / or phenotypic alterations in target organisms. Such alterations are preferably associated with improvements in traits of a target phenotype or physiological importance, correction of endogenous defects, or expression of some type of expression marker. In some respects, traits of a target phenotype or physiological importance are related to: the animal's overall health, adaptability or fertility, the animal's ecological adaptability, or the animal's relationships or interactions with other organisms in the environment.In some respects, phenotypic or physiologically important features are selected from the following groups: improved general health, disease reversal, disease modification, disease stabilization, disease prevention, treatment of parasitic infections, treatment of viral infections, treatment of retroviral infections, treatment of bacterial infections, treatment of neurological disorders (e.g., but not limited to: multiple sclerosis), and endogenous genetic defects (e.g., but not limited to: metabolic disorders, rickets, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Barth syndrome, breast cancer, Charcot-Marie-Tooth disease, colon cancer, Cri-du-chat syndrome, Crohn's disease, cystic fibrosis, Dercum disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombosis). Leiden Thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, Fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, prostate cancer, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, skin cancer, spinal muscular atrophy, Tay-Sachs amaurosis, thalassemia, trimethylamineuria, Turner syndrome, Velocardiofacial syndrome, WAGR syndrome, and Wilson's disease. Treatments for diseases such as the correction of congenital immune disorders (e.g., but not limited to: immunoglobulin subclass deficiencies), the treatment of acquired immune disorders (e.g., but not limited to: AIDS and other HIV-related disorders), cancer, and diseases including rare or “orphan” conditions for which no other effective treatment options have been found.

[0476] Genetically modified cells using the compositions or methods disclosed herein can be transplanted into subjects for purposes such as gene therapy, for example, to treat diseases or as antiviral, antipathogenic or anticancer therapeutic agents, for the production of genetically modified organisms in agriculture or for biological research.

[0477] In vitro polynucleotide detection, binding and modification

[0478] In some aspects, the compositions disclosed herein can be further used (in some aspects together with one or more isolated polynucleotide sequences) as compositions for use in in vitro methods. The one or more isolated polynucleotide sequences may contain one or more target sequences for modification. In some aspects, the one or more isolated polynucleotide sequences may be genomic DNA, PCR products, or synthetic oligonucleotides.

[0479] Composition

[0480] Modification of the target sequence can take the following forms: nucleotide insertion, nucleotide deletion, nucleotide substitution, addition of atoms or molecules to an existing nucleotide, nucleotide modification, or binding of a heterologous polynucleotide or polypeptide to the target sequence. Insertion of one or more nucleotides can be accomplished by including a donor polynucleotide in the reaction mixture: the donor polynucleotide is inserted into a double-strand break generated by the Cas-α orthologous polypeptide. Insertion can be performed via non-homologous end joining or via homologous recombination.

[0481] In some respects, the sequence of the target polynucleotide is known prior to modification and is compared with one or more sequences of one or more polynucleotides produced by Cas-α ortholog treatment. In other respects, the sequence of the target polynucleotide is unknown prior to modification, and Cas-α ortholog treatment is used as part of a method for determining the sequence of said target polynucleotide.

[0482] In some respects, Cas-α orthologs may be selected from the group consisting of: unmodified wild-type Cas-α orthologs; functional Cas-α ortholog variants; functional Cas-α ortholog fragments; fusion proteins containing active or inactivated Cas-α orthologs; Cas-α orthologs that further contain one or more nuclear localization sequences (NLS) at the C-terminus, the N-terminus, or both the N-terminus and the C-terminus; biotinylated Cas-α orthologs; Cas-α ortholog nickases; Cas-α ortholog endonucleases; Cas-α orthologs that further contain histidine tags; and mixtures of any two or more of the above.

[0483] In some respects, Cas-α orthologs are fusion proteins that further include nuclease domains, transcription activator domains, transcription repressor domains, epigenetic modification domains, cleavage domains, nuclear localization signals, cell penetration domains, translocation domains, markers, or cellular heterologous transgenes of the target polynucleotide sequence or from which the target polynucleotide sequence is obtained or derived.

[0484] In some aspects, multiple Cas-α orthologs are desirable. In some aspects, the multiple may comprise Cas-α orthologs derived from different organisms or from different loci within the same organism. In some aspects, the multiple may comprise Cas-α orthologs with different binding specificities to target polynucleotides. In some aspects, the multiple may comprise Cas-α orthologs with different cleavage efficiencies. In some aspects, the multiple may comprise Cas-α orthologs with different PAM specificities. In some aspects, the multiple may comprise orthologs with different molecular compositions (i.e., polynucleotide Cas-α orthologs and polypeptide Cas-α orthologs).

[0485] Guide polynucleotides can be provided as a single guide RNA (sgRNA), a chimeric molecule containing tracrRNA, a chimeric molecule containing crRNA, a chimeric RNA-DNA molecule, a DNA molecule, or a polynucleotide containing one or more chemically modified nucleotides.

[0486] Storage conditions for Cas-α orthologs and / or guide polynucleotides include parameters related to temperature, state of matter, and time. In some respects, Cas-α orthologs and / or guide polynucleotides are stored at approximately -80°C, approximately -20°C, approximately 4°C, approximately 20–25°C, or approximately 37°C. In some respects, Cas-α orthologs and / or guide polynucleotides are stored in liquid, frozen liquid, or lyophilized powder form. In some respects, Cas-α orthologs and / or guide polynucleotides are stable for at least one day, at least one week, at least one month, at least one year, or even longer than one year.

[0487] Any or all possible polynucleotide components of the reaction (e.g., guide polynucleotides, donor polynucleotides, optionally Cas-α polynucleotides) may be provided as part of a vector, construct, linearized or circularized plasmid, or as part of a chimeric molecule. Each component may be provided to the reaction mixture individually or together. In some aspects, one or more polynucleotide components may be operatively linked to a heterologous noncoding regulatory element that regulates their expression.

[0488] Methods for modifying target polynucleotides involve assembling a minimal set of elements into a reaction mixture comprising: a Cas-α ortholog (or a variant, fragment, or other relevant molecule as described above), a guiding polynucleotide (containing a sequence substantially complementary to or selectively hybridizing with the target polynucleotide sequence of the target polynucleotide), and the target polynucleotide for modification. In some aspects, the Cas-α ortholog is provided as a polypeptide. In some aspects, the Cas-α ortholog is provided as a Cas-α ortholog polynucleotide. In some aspects, the guiding polynucleotide is provided as an RNA molecule, a DNA molecule, an RNA:DNA hybrid, or a polynucleotide molecule containing a chemically modified nucleotide.

[0489] Storage buffers or reaction mixtures of any of their components can be optimized for stability, efficacy, or other parameters. Additional components of the storage buffer or reaction mixture may include a buffer composition, Tris, EDTA, dithiothreitol (DTT), phosphate-buffered saline (PBS), sodium chloride, magnesium chloride, HEPES, glycerol, BSA, salts, emulsifiers, detergents, chelating agents, redox agents, antibodies, nuclease-free water, proteases, and / or viscosity agents. In some aspects, the storage buffer or reaction mixture further comprises a buffer solution having at least one of the following components: HEPES, MgCl2, NaCl, EDTA, protease, proteinase K, glycerol, and nuclease-free water.

[0490] Incubation conditions will vary depending on the desired results. The preferred temperatures are at least 10°C, between 10°C and 15°C, at least 15°C, between 15°C and 17°C, at least 17°C, between 17°C and 20°C, at least 20°C, between 20°C and 22°C, at least 22°C, between 22°C and 25°C, at least 25°C, between 25°C and 27°C, at least 27°C, between 27°C and 30°C, at least 30°C, between 30°C and 32°C, at least 32°C, between 32°C and 35°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, or even greater than 40°C. The incubation time is at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, or even more than 10 minutes.

[0491] Before, during, or after incubation, one or more sequences of one or more polynucleotides in the reaction mixture can be determined by any method known in the art. In some aspects, the modification of the target polynucleotide can be determined by comparing one or more sequences of one or more polynucleotides purified from the reaction mixture with the sequence of the target polynucleotide prior to binding to the Cas-α ortholog.

[0492] The kit may contain any one or more of the compositions disclosed herein that can be used for in vitro or in vivo polynucleotide detection, binding, and / or modification. The kit contains a Cas-α ortholog or encoding such a Cas-α ortholog or polynucleotide Cas-α ortholog, and optionally further contains a buffer component capable of efficient storage, and one or more additional compositions capable of introducing the Cas-α ortholog or Cas-α ortholog into a heteropolynucleotide, wherein the Cas-α ortholog or Cas-α ortholog enables modification, addition, deletion, or substitution of at least one nucleotide of the heteropolynucleotide. In another aspect, the Cas-α orthologs disclosed herein can be used to enrich one or more polynucleotide target sequences from a mixing pool. In yet another aspect, the Cas-α orthologs disclosed herein can be immobilized on a matrix for in vitro target polynucleotide detection, binding, and / or modification.

[0493] For storage, purification, and / or characterization purposes, Cas-α endonucleases can be attached, bound, or affixed to a solid matrix. Examples of solid matrices include, but are not limited to, filters, chromatographic resins, assay plates, test tubes, and cryogenic vials. Cas-α endonucleases can be thoroughly purified and stored in suitable buffer solutions or lyophilized.

[0494] Detection methods

[0495] Methods for detecting Cas-α endonuclease-guided polynucleotide complexes bound to target polynucleotides may include any methods known in the art, including but not limited to microscopy, chromatographic separation, electrophoresis, immunoprecipitation, filtration, nanopore separation, microarrays, and those described below.

[0496] Electrophoretic mobility shift assay (EMSA): This technique studies proteins that bind to known DNA oligonucleotide probes and assesses the specificity of the interaction. It is based on the principle that protein-DNA complexes migrate more slowly than free DNA molecules during polyacrylamide or agarose gel electrophoresis. Because DNA migration is impeded upon protein binding, this assay is also known as a gel retardation assay. Adding a protein-specific antibody to the binding component produces a larger complex (antibody-protein-DNA) that migrates even more slowly during electrophoresis; this is called hypervariable and can be used to confirm protein identity.

[0497] DNA pull-down assays use DNA probes labeled with high-affinity tags (such as biotin), which allow for probe recovery or immobilization. The DNA probes can be complexed with proteins from cell lysates obtained in a similar reaction used in EMSA and then purified using agarose or magnetic beads. The proteins are then eluted from the DNA and identified by Western blotting or mass spectrometry. Alternatively, the proteins can be labeled with affinity tags, or the DNA-protein complexes can be separated using antibodies against the target protein (similar to ultravariable assays). In this case, the unknown DNA sequence bound to the protein is detected by Western blotting or PCR analysis.

[0498] Reporter assays provide real-time in vivo readouts of the translational activity of a target promoter. A reporter gene is a fusion of a target promoter DNA sequence and a reporter gene DNA sequence (which is custom-designed by the researcher and encodes a protein with detectable properties, such as firefly / Raeneria luciferase or alkaline phosphatase). These genes produce the enzyme only when the target promoter is activated. The enzyme then catalyzes the substrate to produce a light or color change that can be detected by spectroscopic instruments. The signal from the reporter gene serves as an indirect determinant of the translation of endogenous proteins driven by the same promoter.

[0499] Microplate capture and detection assays use immobilized DNA probes to capture specific protein-DNA interactions and confirm protein identity and relative abundance with target-specific antibodies. Typically, DNA probes are immobilized on the surface of 96- or 384-well microplates coated with streptavidin. Cell extracts are prepared and added to bind proteins to oligonucleotides. The extracts are then removed, and each well is washed several times to remove non-specifically bound proteins. Finally, the protein is detected using a labeled, specific antibody. This method is highly sensitive, capable of detecting less than 0.2 pg of target protein per well. This method can also be used for oligonucleotides labeled with other tags, such as primary amines that can be immobilized on microplates coated with amine-reactive surface chemicals.

[0500] DNA footprinting is one of the most widely used methods for obtaining detailed information about individual nucleotides within protein-DNA complexes and even the interior of living cells. In such experiments, chemicals or enzymes are used to modify or digest DNA molecules. When sequence-specific proteins bind to DNA, they can protect the binding sites from modification or digestion. This is then visualized by denaturing gel electrophoresis, where the unprotected DNA is more or less randomly cut. Therefore, it appears as a “ladder” of bands, and the protein-protected sites do not have corresponding bands and look like footprints in the banding pattern. The footprints are left here by identifying specific nucleotides at the protein-DNA binding sites.

[0501] Microscopy techniques include optical, fluorescence, electron, and atomic force microscopy (AFM).

[0502] Chromatin immunoprecipitation (ChIP) allows proteins to covalently...

Claims

1. A non-naturally occurring Cas-α10 polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2, wherein the amino acid sequence comprises mutations relative to SEQ ID NO: 2, wherein the mutations comprise K85S mutation; K85A mutation; N92R mutation; N92K mutation; Q125K mutation; N88H mutation; N88K mutation; N88Q mutation; Q125F mutation; Y72V mutation; Y72E mutation; Y72Q mutation; Y72T mutation; Y72C mutation; Y72A mutation; Y72S mutation; Y72P mutation; Y72G mutation; Y72D mutation; Y72L mutation; K85G mutation; K85D mutation; K85N mutation; N88D mutation; Q89D mutation; N92D mutation; N92C mutation; N92A mutation; N92G mutation; N92F mutation; N92E mutation; N92H mutation; N92I mutation; N92L mutation; N92V mutation; N92W mutation; N92Q mutation; N92S mutation; N92P mutation; N92Y mutation; N92T mutation; N92M mutation; Q125R mutation; or Q125P mutation.

2. The non-naturally occurring Cas-α10 polypeptide of claim 1, wherein the amino acid sequence corresponds to SEQ ID NO: 2 includes one or more of the following mutation combinations: K85Q and N92L mutations; N88H and Q89G mutations; N88K and Q89G mutations; N88Q and Q89G mutations; K85S and N92L mutations; K85S and N92Q mutations; K85S and N92C mutations; K85S and N92H mutations; K85S and N92A mutations; K85S and N92M mutations; K85N and N92L mutations; K85N and N92H mutations; K85N and N92A mutations; K85N and N92C mutations; K85N and N92M mutations; K85N and N92Q mutations; K85N and N92I mutations; K85S, N88D, and Q8... Combinations of 9G mutations; combinations of K85S, N88H, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89D, and Q125R mutations; combinations of K85S, N88D, Q89G, and N92L mutations; combinations of Y72A, N88D, Q89G, and Q125R mutations; combinations of K85S, N92L, and Q125R mutations; Combinations of Y72S mutation, K85D mutation, Q125R mutation, and N127R mutation; combinations of Y72A mutation, K85S mutation, Q89D mutation, N92L mutation, and Q125R mutation; combinations of Y72C mutation, N88H mutation, Q89G mutation, and Q125R mutation; combinations of Y72C mutation, N88D mutation, Q89D mutation, N92W mutation, and Q125R mutation; or combinations of K85Q and N92W mutations.

3. The non-naturally occurring Cas-α10 polypeptide of claim 1 or claim 2, wherein the amino acid sequence comprises a PAM interaction (PI) domain, the PAM interaction (PI) domain comprising amino acids from S63 to I196.

4. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-3, wherein the amino acid sequence comprises the K85S mutation.

5. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-3, wherein the amino acid sequence comprises a combination of the K85Q mutation and the N92L mutation.

6. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-3, wherein the amino acid sequence comprises a combination of the K85N mutation and the N92L mutation.

7. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-6, wherein the Cas-α10 polypeptide has endonuclease activity.

8. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-6, wherein the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase.

9. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-6, wherein the Cas-α10 polypeptide has cleavage enzyme activity.

10. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-6 or 9, wherein the Cas-α10 polypeptide is complexed with or operatively associated with reverse transcriptase.

11. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-6, wherein the Cas-α10 polypeptide is complexed with a heterologous protein domain via a linker, the heterologous protein domain having methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity.

12. The non-naturally occurring Cas-α10 polypeptide according to any one of claims 1-11, comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID No. 6, 5, 7-16 or 28-87.

13. A non-naturally occurring Cas-α10 polypeptide comprising a PAM interaction (PI) domain, wherein the PI domain recognizes a PAM sequence comprising: 5'-YTC-3', 5'-TTTC-3', 5'-GTTC-3', 5'-GTYC-3', 5'-GTWC-3', 5'-DTHC-3', 5'-GTNC-3', 5'-NNNC-3', 5'-GNNC-3', 5'-HCTC-3', 5'-NTCC-3', 5'-TNC-3', 5' -GTTY-3', 5'-DTTY-3', 5'-GTTT-3', 5'-GTHC-3', 5'-NYTY-3', 5'-TTTY-3', 5'-TTYTY-3', 5'-TYTY-3', 5'-TTNN-3', 5'-NNCC-3', 5'-NCCC-3', 5'-NTNN-3', 5'-NCTY-3', 5'-NTHC-3', 5'-NNNN-3' or 5'-NNCN-3', where N = A, C, G or T, Y = T or C, D = G, A or T, W = A or T, and H = A, T or C.

14. The non-naturally occurring Cas-α10 polypeptide of claim 13, wherein the Cas-α10 polypeptide has endonuclease activity.

15. The non-naturally occurring Cas-α10 polypeptide of claim 13, wherein the Cas-α10 polypeptide is an inactivated Cas-α10 endonuclease complexed with a deaminase.

16. The non-naturally occurring Cas-α10 polypeptide of claim 13, wherein the Cas-α10 polypeptide has cleavage enzyme activity.

17. The non-naturally occurring Cas-α10 polypeptide of claim 13 or claim 16, wherein the Cas-α10 polypeptide is complexed with or operatively associated with reverse transcriptase.

18. The non-naturally occurring Cas-α10 polypeptide of claim 13, wherein the Cas-α10 polypeptide complexes with a heterologous protein domain via a linker, and wherein the heterologous protein domain has methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, reverse transcriptase activity, RNA aptamer binding activity, or nucleic acid binding activity.

19. A synthetic composition comprising: (a) a Cas-α10 polypeptide with DNA-binding activity, the non-naturally occurring Cas-α10 polypeptide as described in any one of claims 1-18; and (b) at least one guiding polynucleotide comprising a region complementary to the target polynucleotide, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide. The Cas-α10 polypeptide recognizes the PAM sequence on the target polynucleotide and forms a complex with the at least one guide polynucleotide, and the complex binds to the target polynucleotide.

20. A method for editing target polynucleotides in cells, the method comprising: (a) Providing the cell with the Cas-α10 polypeptide as described in any one of claims 1-18, wherein the Cas-α10 polypeptide recognizes the PAM sequence on the target polynucleotide; (b) Providing the cell with at least one guide polynucleotide comprising a region complementary to the target polypeptide, wherein the Cas-α10 polypeptide forms a complex with the guide polynucleotide, and the complex binds to the target polynucleotide; as well as (c) Introducing at least one nucleotide modification into the target polynucleotide via the complex, wherein the target polynucleotide is heterologous to the Cas-α10 polypeptide.

21. The method of claim 20, further comprising providing the cell with a donor DNA molecule or a polynucleotide modification template.

22. The method of claim 20, wherein the cell is derived from or obtained from an animal, fungus, or plant.

23. The method of claim 22, wherein the plant is a dicotyledonous plant or a monocotyledonous plant.

24. The method of claim 22, wherein the plant is corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, or tomato.

25. An animal, fungus, or cell thereof, wherein the animal, fungus, or cell thereof contains a non-naturally occurring Cas-α10 polypeptide as claimed in any one of claims 1-18.

26. A plant or plant cell comprising a non-naturally occurring Cas-α10 polypeptide as described in any one of claims 1-18.

27. A method for altering the specificity of the prespacer neighbor motif (PAM) of a target Cas-α10 polypeptide, the method comprising: (a) Comparing the PAM interaction (PI) domain of the orthologous Cas-α polypeptide with the PI domain of the target Cas-α 10 polypeptide; (b) Select one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-α polypeptide; (c) Incorporating one or more amino acids selected from the PI domain of the orthologous Cas-α polypeptide and / or one or more polypeptide chains into one or more structurally similar positions of the target Cas-α 10 polypeptide to produce a modified target Cas-α 10 polypeptide. as well as (d) Determine the PAM recognition of the modified target Cas-α10 peptide.

28. The method of claim 27, wherein the orthologous Cas-α polypeptide has a different PAM recognition than the target Cas-α 10 polypeptide.

29. The method of claim 27, wherein the Cas-α 10 polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

2.

30. The method of any one of claims 27-29, wherein the orthologous Cas-α polypeptide is Cas-α 1, Cas-α 2, Cas-α 3, Cas-α 4, Cas-α 5, Cas-α 6, Cas-α 7, Cas-α 8 or Cas-α 11.

Citation Information

Patent Citations

  • Method of introducing nucleic acid into plant cells

    EP2821486A1

  • CRISPR-CAS systems for genome editing

    US10934536B2

  • Honeycomb formed body and method for producing honeycomb structure

    US11261134B2

  • Nuclease-independent targeted gene editing platform and uses thereof

    US11479793B2

  • Methods for altering the genome of a monocot plant cell

    US20090133152A1