Engineered crispr-CAS12 systems for genome editing

Engineered Cas12l polypeptides with altered sequences and guide RNAs address the specificity and cost issues of existing CRISPR-Cas systems, enabling efficient and cost-effective genome editing in eukaryotic organisms.

WO2025166125A1PCT designated stage Publication Date: 2025-08-07PIONEER HI BREED INTERNATIONAL INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013977
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2025-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing genome-editing technologies, such as CRISPR-Cas systems, face challenges with low specificity and the need for redesigning nucleases for each target site, leading to high costs and inefficiencies, particularly in eukaryotic organisms like animals and plants.

Method used

Development of engineered Cas12l polypeptides with altered amino acid sequences and guide RNAs that enhance DNA recognition, binding, and cleavage activities, allowing for targeted genome editing with improved specificity and flexibility.

Benefits of technology

The engineered Cas12l polypeptides and guide RNAs enable precise and efficient genome editing in eukaryotic cells, reducing the need for redesign and lowering costs by improving interaction with DNA substrates and enhancing editing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000017_0001
    Figure IMGF000017_0001
  • Figure IMGF000023_0001
    Figure IMGF000023_0001
  • Figure IMGF000023_0002
    Figure IMGF000023_0002
Patent Text Reader

Abstract

Provided herein are engineered Cas12l (Cas-beta) polypeptides, interacting guide RNAs, as well as methods and compositions for use thereof; also provided is a method for selectively enriching for Cas polypeptide-guide RNA ribonucleoproteins (RNPs) having double stranded polynucleotide cleavage activity.
Need to check novelty before this filing date? Find Prior Art

Description

ENGINEERED CRISPR-CAS12 SYSTEMS FOR GENOME EDITING CROSS-REFRENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional No.63 / 549,039, filed February 2, 2024, the disclosure of which is incorporated by reference herein in its entirety. REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The official copy of the sequence listing is submitted electronically via Patent Center as an XML formatted sequence listing with a file named 211821-WO-SEC-1_ST26 created on January 30, 2025, and having a size of 475,123 bytes, and is filed concurrently with the specification. The sequence listing comprised in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety. FIELD

[0003] The disclosure relates to the field of molecular biology, in particular to compositions of novel, improved RNA-guided Cas endonuclease systems, and compositions and methods for editing or modifying the genome of a cell. BACKGROUND

[0004] Recombinant DNA technology has made it possible to insert DNA sequences at targeted genomic locations and / or modify specific endogenous chromosomal sequences. Site- specific integration techniques, which employ site-specific recombination systems, as well as other types of recombination technologies, have been used to generate targeted insertions of genes of interest in a variety of organism. Genome-editing techniques such as designer zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or homing meganucleases, are available for producing targeted genome perturbations, but these systems tend to have low specificity and employ designed nucleases that need to be redesigned for each target site, which renders them costly and time-consuming to prepare.

[0005] Newer technologies utilizing archaeal or bacterial adaptive immunity systems have been identified, called CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats), which comprise different domains of effector proteins that encompass a variety of activities (DNA recognition, binding, and optionally cleavage).

[0006] The protospacer adjacent motif (PAM) requirement of CRISPR-associated (Cas) polypeptides restricts their targeting range (Shmakov et al, 2015; Zetsche et al, 2015; Burstein et al, 2017; Karvelis et al, 2020; Pausch et al, 2020). This becomes particularly apparent ingenome editing applications where the outcome is dependent on the proximity of the desired edit to the cut site (e.g., template-free editing and homology-directed repair) or for approaches that impose additional sequence requirements on target selection (e.g., base editing; Anzalone et al, 2020).

[0007] There remains a need for engineering novel Cas effectors and systems, e.g., to improve their activity in eukaryotes (particularly animals and plants), to improve their interactions with RNA and DNA substrates, to improve their DNA recognition, binding, and optionally cleavage activity, and to improve their ability to edit endogenous (or previously introduced heterologous) polynucleotides.

[0008] Herein are described are novel engineered Cas12l polypeptides as well as novel interacting guide RNAs, and methods and compositions for use thereof. The terms Cas12l and Cas-beta are used herein interchangeably, such that a Cas 12-beta polypeptide, endonuclease or effector refers to a Cas12l polypeptide, endonuclease, or effector respectively. SUMMARY

[0009] Disclosed herein are novel engineered Cas 12l, i.e., Cas-beta, polypeptides, compositions, and methods of use thereof. These Cas polypeptides can be guided by a guide polynucleotide to target double-stranded DNA in a PAM-dependent fashion. In some examples, the engineered Cas polypeptides are active endonucleases capable of introducing a break at the target site of the target double-stranded DNA. In other examples, the Cas polypeptide comprises one or more mutations that render it incapable of double-strand cutting, but permit single-strand cutting. In certain examples, the Cas polypeptide comprises one or more mutations that render it incapable of cleaving either or both strands of a double-stranded polynucleotide, but it retains the ability to bind to a target polynucleotide sequence.

[0010] In all aspects, the novel engineered Cas-beta polypeptides disclosed herein are engineered to include different amino acid sequence from the native Cas12l effectors obtained or derived from Armatimonadetes bacteria (Urbaitis et al. 2022 EMBO Reports, 23(12): e55481).

[0011] Also disclosed herein are novel sgRNA variants, compositions, and methods of use thereof. These sgRNA variants are capable of interacting and forming a complex with Cas- beta (Cas12l) polypeptides (including the novel engineered Cas-beta polypeptides disclosed herein) and a target polynucleotide.

[0012] In a first aspect, provided herein is an engineered Cas-beta polypeptide, or a polynucleotide encoding the engineered Cas-beta polypeptide, wherein the Cas-betapolypeptide comprises one or more altered amino at the acids corresponding to the residue(s) at position 17, 21, 68, 131, 142, 153, 244, 253, 297, 301, 342, 452, 572, 576, 607, 615, 762, 766, or 771 of SEQ ID NO:1. As used herein, a reference to an engineered Cas-beta polypeptide that comprises an amino acid or multiple amino acids corresponding to the residue(s) at position(s) [x] of SEQ ID NO:[y] means that, when the Cas-beta is aligned to SEQ ID NO:[y], the engineered Cas-beta comprises an altered amino acid at the Cas-beta position that is aligned with position(s) [x] of SEQ ID NO:[y].

[0013] In one example, the first aspect provides an engineered Cas-beta polypeptide, or a polynucleotide encoding the engineered Cas-beta polypeptide, wherein the Cas-beta polypeptide comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:1 and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residue(s) at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of of SEQ ID NO:1. Thus, the altered amino acids can be one or more of the altered amino acids of in any one of SEQ ID NOs:32-73 and SEQ ID NOs:180-201. For example, the engineered Cas-beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in any one of SEQ ID NOs:32-73 and SEQ ID NOs:180-201. In particular examples, the engineered Cas- beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in SEQ ID NO:68 (correspond to positions 21 and 615 of SEQ ID NO:1), SEQ ID NO:72 (correspond to positions 297, 301, 342, 615, 762, and 766 of SEQ ID NO:1), or SEQ ID NO:189 (correspond to positions 297, 301, 342, 572, 607, 615, 762, and 766 of SEQ ID NO:1.

[0014] In certain examples, the provided engineered Cas-beta polypeptide comprises the sequence of any one of SEQ ID NOs: 32-73 and SEQ ID NOs:180-201 such as, e.g., SEQ ID NO:189. Also provided is a polynucleotide comprising coding sequence for any one of SEQ ID NOs: 32-73 and SEQ ID NOs:180-201, including a mammalian codon-optimized or plant codon-optimized sequence encoding any one of SEQ ID NOs: 32-73 and SEQ ID NOs:180- 201, e.g., provided is a polynucleotide comprising any one of SEQ ID NOs:134-173. For example, provided is a polynucleotide comprising coding sequence for SEQ ID NO:189, which can be a codon-optimized for expression in mammals or plants.

[0015] In another example of the first aspect, provided is an engineered Cas-beta polypeptide, or a polynucleotide encoding the engineered Cas-beta polypeptide, wherein the Cas-beta polypeptide comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:129, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0016] In yet another example of the first aspect, provided is an engineered Cas-beta polypeptide, or a polynucleotide encoding the engineered Cas-beta polypeptide, wherein the Cas-beta polypeptide comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:130, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0017] In still another example of the first aspect, provided is an engineered Cas-beta polypeptide, or a polynucleotide encoding the engineered Cas-beta polypeptide, wherein the Cas-beta polypeptide comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:131, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0018] In particular examples, each of the foregoing examples of an engineered Cas-beta polypeptide can have endonuclease activity or, alternatively, be deactivated to lack endonuclease activity. When deactivated, each of the foregoing examples of an engineered Cas-beta polypeptide can be complexed to a base editor, such as a nucleoside deaminase.

[0019] In particular examples, each of the foregoing examples of an engineered Cas-beta polypeptide can be complexed to a heterologous protein domain through a linker, and wherein the heterologous protein domain has transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

[0020] In a second aspect, provided is a synthetic composition comprising the engineered Cas12l polypeptide disclosed herein and at least one guide polynucleotide (e.g., gRNA) that comprises a region of complementarity to a target polynucleotide, wherein the engineered Cas12l polypeptide forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide. Optionally, the synthetic composition is in a complex that comprises the at least one guide polynucleotide and the target polynucleotide. The target polynucleotide can be from any organism, e.g., a mammal, fungus, plant, bacteria, protozoa, or archaebacteria cell. In certain preferred examples, the target polynucleotide is from a human cell.

[0021] In one example, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises one or more altered amino at the acids corresponding to the residue(s) at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0022] In another example, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises (i) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:1 and (ii) one or more altered amino acids corresponding to the residue(s) at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. Thus, the altered amino acids can be one or more of the altered amino acids of in any one or more of SEQ ID NOs:32-73 and 180-201. For example, synthetic composition can comprise an engineered Cas-beta polypeptide comprising all of the altered amino acids (relative to SEQ ID NO:1) in any one of SEQ ID NOs:32-73 and SEQ ID NOs:180-201. In particular examples, the engineered Cas-beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in SEQ ID NO:68 (correspond to positions 21 and 615 of SEQ ID NO:1), SEQ ID NO:72 (correspond to positions 297, 301, 342, 615, 762, and 766 of SEQ ID NO:1), or SEQ ID NO:189 (correspond to positions 297, 301, 342, 572, 607, 615, 762, and 766 of SEQ ID NO:1.

[0023] In certain examples, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises the sequence of any one of SEQ ID NOs: 32- 73 and 180-201.

[0024] In yet another example, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:129, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0025] In still another example, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:130, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0026] In another example, the synthetic composition of the second aspect comprises an engineered Cas-beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:131, and wherein the Cas-beta polypeptide comprises one or more altered amino acidscorresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0027] In particular examples, each of the foregoing examples of a synthetic composition of the second aspect, the engineered Cas-beta polypeptide can have endonuclease activity or, alternatively, be deactivated to lack endonuclease activity. When deactivated, the engineered Cas-beta polypeptide can be complexed to a base editor, such as a nucleoside deaminase. In additional particular examples of a synthetic composition of the second aspect, the engineered Cas-beta polypeptide can be complexed to a heterologous protein domain through a linker, and wherein the heterologous protein domain has transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

[0028] In a third aspect, provided herein is a Cas-beta variant single guide RNA (sgRNA) capable of directing a Cas-beta polypeptide to a target DNA, wherein the sgRNA variant comprises an altered stem-loop hairpin relative to an unaltered Cas12-beta single guide, wherein the stem-loop hairpin is the first hairpin closest to 5’ terminus of the unaltered sgRNA. The alteration can be a deletion of one or more nucleotides of the stem-loop hairpin relative to the unaltered Cas12l sgRNA. For example, the alteration can be a deletion of three or more nucleotides of the stem-loop hairpin relative to the unaltered Cas12l sgRNA. In a different example, the alteration can be an inserted, additional hairpin sequence, which is inserted at the stem-loop hairpin relative to the unaltered Cas-beta sgRNA. In particular examples of this third aspect, the variant sgRNA comprises any one of SEQ ID NOs:74-121, 123-128, 202-216, 283, 284, and 287-316.

[0029] In a fourth aspect, provided herein is a construct comprising nucleotide sequence encoding any example of the engineered Cas-beta polypeptide provided by the first aspect disclosed herein. Such a construct can be an expression construct, which can be used in a method that includes expressing the encoded sequence to generate any one (or more) of the engineered Cas-beta polypeptide provided by the first aspect disclosed herein.

[0030] In a fifth aspect, provided herein is a construct comprising nucleotide sequence encoding any example of the Cas-beta variant single guide RNA (sgRNA) disclosed herein. Also provided is a construct comprising nucleotide sequence encoding (i) an engineered Cas-beta polypeptide disclosed herein and (ii) a Cas-beta variant sgRNA disclosed herein. Such a construct can be an expression construct, which can be used in a method that includes expressing the encoded sequence to generate any one (or more) of the Cas-beta variant sgRNA provided by the second aspect disclosed herein.

[0031] In a sixth aspect, provided herein is a method of generating an engineered Cas-beta, said method comprising producing any one (or more) of the engineered Cas-beta polypeptide provided by the first aspect disclosed herein. These can be produced by expressing the encoded Cas-beta polypeptide from a construct of the fourth aspect disclosed herein.

[0032] In a seventh aspect, provided herein is a method of generating an engineered Cas- beta, said method comprising providing sequence of an unaltered Cas-beta polypeptide, and altering one or more amino acids in the Cas-beta polypeptide corresponding to one or more of the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1, thereby producing the engineered Cas-beta polypeptide. In one example of the seventh aspect, the method includes (a) providing coding sequence encoding an unaltered Cas- beta polypeptide; (b) altering the coding sequence encoding one or more amino acids in the Cas-beta polypeptide corresponding to one or more of the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1; and (c) expressing the altered coding sequence thereby producing the engineered Cas-beta polypeptide.

[0033] In a particular example of the seventh aspect, the method includes (a) providing coding sequence encoding an unaltered Cas-beta polypeptide; (b) aligning the amino acid sequence of the unaltered Cas12l polypeptide to SEQ ID NO:1; (c) using the alignment to determine the one or more amino acids in the unaltered Cas12l polypeptide corresponding to one or more of the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1, respectively; and (d) altering the coding sequence encoding one or more amino acids in the unaltered Cas-beta polypeptide corresponding to one or more of the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1; and (c) expressing the altered coding sequence and thereby producing the engineered Cas-beta polypeptide. For example, the engineered Cas-beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in any one of SEQ ID NOs:32-73 and SEQ ID NOs:180-201. In particular examples, the engineered Cas-beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in SEQ ID NO:68 (correspond to positions 21 and 615 of SEQ ID NO:1), SEQ ID NO:72 (correspond to positions 297, 301, 342, 615, 762, and 766 of SEQ ID NO:1), or SEQ ID NO:189 (correspond to positions 297, 301, 342, 572, 607, 615, 762, and 766 of SEQ ID NO:1.

[0034] In an eight aspect, provided herein is a method of generating an Cas-beta variant sgRNA, said method comprising synthesizing any one (or more) of the examples of Cas-beta variant sgRNA provided by the second aspect disclosed herein. These can be produced by expressing the encoded Cas-beta polypeptide from a construct of the fifth aspect disclosed herein.

[0035] In a particular example of the eighth aspect, the method includes (a) providing the sequence of an unaltered Cas-beta sgRNA sequence comprising a stem-loop hairpin, wherein the stem-loop hairpin is the first hairpin closest to 5’ terminus of the sgRNA, and (b) synthesizing an Cas-beta variant sgRNA in which the stem-loop hairpin is altered relative to the unaltered Cas-beta sgRNA, thereby generating a Cas-beta variant sgRNA of the second aspect.

[0036] In a ninth aspect, provided is a method of editing a target polynucleotide in a cell (, the method comprising: (a) providing to the cell the engineered Cas-beta polypeptide of the first aspect (e.g., via a construct of the fifth aspect), (b) providing the cell with at least one guide polynucleotide comprising a region of complementarity to the target polypeptide, wherein the engineered Cas12l polypeptide forms a complex with the guide polynucleotide and the complex binds the target polynucleotide; and (c) introducing at least one nucleotide modification in the target polynucleotide via the complex.

[0037] In one example of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises one or more altered amino at the acids corresponding to the residue(s) at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0038] In another example of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises (i) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:1 and (ii) one or more altered amino acids corresponding to the residue(s) at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. Thus, the altered amino acids can be one or more of the altered amino acids of in any one or more of SEQ ID NOs:32-73 and SEQ ID NOs:180-201. For example, the engineered Cas-beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in any oneof SEQ ID NOs:32-73 and SEQ ID NOs:180-201. In particular examples, the engineered Cas- beta polypeptide can comprise all of the altered amino acids (relative to SEQ ID NO:1) in SEQ ID NO:68 (correspond to positions 21 and 615 of SEQ ID NO:1), SEQ ID NO:72 (correspond to positions 297, 301, 342, 615, 762, and 766 of SEQ ID NO:1), or SEQ ID NO:189 (correspond to positions 297, 301, 342, 572, 607, 615, 762, and 766 of SEQ ID NO:1.

[0039] In certain examples of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises the sequence of any one of SEQ ID NOs: 32-73 and SEQ ID NOs:180-201.

[0040] In yet another example of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises an engineered Cas- beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:129, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0041] In still another example of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:130, and wherein the Cas-beta polypeptide comprises one or more altered amino acids corresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0042] In another example of the ninth aspect, step (a) of the method comprises providing to the cell an engineered Cas-beta polypeptide that comprises a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO:131, and wherein the Cas-beta polypeptide comprises one or more altered amino acidscorresponding to the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1. The one or more altered amino acids can be one or more substitutions (e.g. one or more substitutions with a positively charged amino acid), deletions, or insertions at residue(s) corresponding to position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:1.

[0043] In some cases of each of the foregoing examples of the ninth aspect, the engineered Cas-beta polypeptide provided to the cell in step (a) can have endonuclease activity or, alternatively, be deactivated to lack endonuclease activity. When deactivated, the engineered Cas-beta polypeptide can be complexed to a base editor, such as a nucleoside deaminase. In additional particular examples of a synthetic composition of the second aspect, the engineered Cas-beta polypeptide can be complexed to a heterologous protein domain through a linker, and wherein the heterologous protein domain has transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

[0044] In each of the foregoing examples of the ninth aspect, the cell can be a human, mammalian, fungus, plant, bacteria, protozoa, or archaebacteria cell.

[0045] In a tenth aspect, provided herein is a method of A method for selectively enriching for one or more Cas polypeptide-guide RNA ribonucleoproteins (RNPs) having double stranded polynucleotide cleavage activity, the method comprising combining the following in a reaction volume: Cas endonucleases, single guide RNAs (sgRNAs) that each comprises a region of complementarity (ROC) to a target sequence, and double stranded target polynucleotides that comprise the target sequence and a cleavage site, wherein each target polynucleotide comprises two termini and each target polynucleotide comprise a first tag at the terminus that is more proximal to the cleavage site relative to the ROC target site and a different second tag at the polynucleotide terminus that is more proximal to the ROC target site than the cleavage site. The reaction volume can include (i) Cas endonuclease variants combined with one kind of sgRNA, (ii) one kind of Cas endonuclease combined with sgRNA variants, or (iii) Cas endonuclease variants combined with sgRNA variants. The method further comprises forming RNPs from the Cas endonucleases, sgRNAs, and target polynucleotides and then allowing for Cas cleavage of the target polynucleotide. A first substrate that preferentially binds to the terminus comprising the first tag is then added; and the first substrate, including materials bound to the first substrate, are separated from the remaining reaction volume. This separation removes materials bound to the first substrate such as target polynucleotides that didnot form RNP (may indicate non-functional CRISPR system), RNPs bound to intact target polynucleotides (RNPs that failed to cleave the target polynucleotide), and / or cleaved portions of target polynucleotides. Then a second substrate that preferentially binds to the terminus comprising the second tag is added to the remaining reaction volume, which allows for the enrichment for the second substrate and materials bound to the second substrate, thereby enriching for RNPs comprising cleaved target polynucleotides.

[0046] In one example of the tenth aspect, the method selectively enriches for sgRNA variants that provide improved double stranded polynucleotide cleavage activity relative to an original sgRNA. In this example, the reaction volume comprises (at least predominantly) one kind of Cas endonuclease combined with sgRNA variants (and optionally the original sgRNA); different RNPs are formed that comprise the same Cas endonuclease, different sgRNA variants, and the double stranded target polynucleotides; and the step of enriching for the second substrate and materials bound to the second substrate thereby enriches for RNPs comprising sgRNA variants that provide improved double stranded polynucleotide cleavage activity.

[0047] In another example of the tenth aspect, the method selectively enriches for Cas endonuclease variants that provide improved double stranded polynucleotide cleavage activity relative to an original Cas endonuclease. In this example, the reaction volume comprises an Cas endonuclease variants (and optionally the original Cas endonuclease) combined with (at least predominantly) one kind of sgRNA; different RNPs are formed that comprise the different Cas endonuclease variants, the same kind of sgRNA, and the double stranded target polynucleotides; and the step of enriching for the second substrate and materials bound to the second substrate thereby enriches for SNPs comprising Cas endonuclease variants that provide improved double stranded polynucleotide cleavage activity.

[0048] In still another example of the tenth aspect, the method selectively enriches different combinations of Cas endonuclease variants and sgRNA variants that provide improved double stranded polynucleotide cleavage activity relative to an original combination of a Cas endonuclease and sgRNA. In this example, the reaction volume comprises Cas endonuclease variants (and optionally the Cas endonuclease from the original combination) combined with sgRNA variants (and optionally the sgRNA from the original combination); different RNPs are formed that comprise different combinations of the Cas endonuclease variants, sgRNA variants, and the double stranded target polynucleotides; and the step of enriching for the second substrate and materials bound to the second substrate thereby enriches for SNPs comprising combinations of Cas endonuclease variants and sgRNA variants that provide improved double stranded polynucleotide cleavage activity.

[0049] In each the foregoing examples of the tenth aspect, the first and second tag can each be selected from the group consisting of biotin, azide, folate, polyhistidine, FLAG, myc, maltose binding protein, metal chelating peptides, histidine-tryptophan modules, a protein A domain, and any suitable tag for enrichment.

[0050] In certain examples of the tenth aspect, after separating the first substrate from the remaining reaction volume, the method includes treating the remaining reaction volume with a a reagent that modifies the second tag to introduce a third tag and the second substrate preferentially binds to the second tag via the third tag. For example, in a method described herein the first tag is biotin, the second tag is azide and after separating material bound to the first substrate, the remaining reaction volume is treated with a reagent that modifies the second tag to introduce a biotin; and then the second substrate (which can be the same as the first substrate) can be used to enrich for the SNPs comprising biotin-modified second tag.

[0051] In each the foregoing examples of the tenth aspect, suitable substrates can be used to separate and enrich for materials comprising the first and second tags. Such substrates can comprise (i) magnetic beads or particles, (ii) beads or particles of agarose, cellulose, glass, polystyrene, or polyacrylamide, (iii) a microtiter plate, (iv) a tube, or (v) a membrane. BRIEF DESCRIPTION OF THE DRAWINGS AND THE SEQUENCE LISTING

[0052] The disclosure can be more fully understood from the following detailed description and the accompanying drawings and Sequence Listing, which form a part of this application.

[0053] FIG.1 is a schematic diagram of a Cas-beta expression cassette suitable for expressing in human cells the wild type or novel engineered Cas-beta polypeptides disclosed herein.

[0054] FIG. 2 is a schematic diagram of a Cas-beta sgRNA expression cassette suitable for expressing in human cells the unaltered or altered sgRNA variants disclosed herein.

[0055] FIG. 3 is a bar graph showing genome editing efficiency at RunXI target site (R1); WTAP target site (W6); RunXI target site (R3); and WTAP target site (W1) in human cells transfected with construct encoding Cas-beta2 endonuclease system; efficiency in negative control, non-transfected human cells indicated by “NC”.

[0056] FIG. 4 is a protein ribbon diagram based on Cryo-electron microscopy (“Cryo-EM”) macromolecular reconstruction of Cas-beta2-sgRNA-DNA; positions of amino acids selected for mutagenesis are depicted in black.

[0057] FIG.5 provides Cas-beta2 amino acid sequence (SEQ ID NO:1); amino acids selected for mutagenesis or alteration to generate variants disclosed herein are underlined.

[0058] FIG. 6 is a bar graph showing genome editing efficiency in human cells of Cas-beta2 variants generated by mutagenesis; dashed line indicates average editing efficiency of wild- type Cas-beta2.

[0059] FIG.7 is a bar graph showing genome editing efficiency at R1, W6, R3, and W1 target sites in human cells transfected with the indicated single, double, and triple mutation containing Cas-beta variants - as well as editing efficiency of control cells transfected with wild type (WT) Cas-beta or non-transfected cells (NC).

[0060] FIG.8 is a bar graph showing genome editing efficiency at R1, W6, R3, and W1 target sites in human cells transfected with indicated Cas-beta2 variants – as well as editing efficiency of control cells transfected with wild type (WT) Cas-beta or non-transfected cells (NC).

[0061] FIG. 9 is a schematic diagram of a Cas polypeptide-guide RNA ribonucleoprotein (RNP) and target DNA complex pull-down assay that was used to detect binding activity of sgRNA variants disclosed herein.

[0062] FIG. 10 is a bar graph showing the relative fold-change (y-axis) of sgRNA variants enriched by pull-down assay, as compared to unaltered sgRNA (fold change =1, indicated by dashed line).

[0063] FIG. 11 is a schematic diagram of Cas-beta2 sgRNA scaffold secondary structure; regions expendable for RNP formation as determined by the RNP pulldown assay are highlighted and labeled (D1, D9, D17, and D16).

[0064] FIG.12 is a schematic diagram of Cas-beta2 sgRNA scaffold secondary structure which shows the portion of stem-loop hairpin that was deleted in scaffold region of certain Cas-beta2 sgRNA variants disclosed herein.

[0065] FIG. 13 is a bar graph showing genome editing efficiency in human cells transfected with Cas-beta2 polypeptide and either wild type (WT-sgRNA) or sgRNA variants containing the deletion of the scaffold stem-loop hairpin shown in Fig.12, which variants were directed to R1, W6, R3, or W1 target sites.

[0066] FIG.14 is a schematic diagram of Cas-beta2 sgRNA scaffold secondary structure which shows an additional hairpin that was inserted in the scaffold region of certain Cas-beta2 sgRNA variants disclosed herein.

[0067] FIG. 15 is a bar graph showing genome editing efficiency in human cells transfected with the indicated Cas-beta2 polypeptide and either wild type (WT-sgRNA) or sgRNA variants containing the hairpin insertion (SS-HP sgRNA) shown in Fig.14, which variants were directed to R1, W6, R3, or W1 target sites.

[0068] FIG.16 provides amino acid sequence alignment of Cas-beta2 (SEQ ID NO:1) and Cas- beta 1 (SEQ ID NO:129); amino acids selected for mutagenesis or alteration to create variants are boxed.

[0069] FIG.17 provides amino acid sequence alignment of Cas-beta2 (SEQ ID NO:1) and Cas- beta 3 (SEQ ID NO:130); amino acids selected for mutagenesis or alteration to create variants are boxed.

[0070] FIG.18 provides amino acid sequence alignment of Cas-beta2 (SEQ ID NO:1) and Cas- beta 4 (SEQ ID NO:131); amino acids selected for mutagenesis or alteration to create variants are boxed.

[0071] FIG. 19 is a set of four bar graphs showing genome editing efficiency in human cells of M67 Cas-beta2 as compared to the indicated further M67 Cas-beta2 variants (each expressed from M67 vector); NC indicates no Cas endonuclease, i.e, empty vector control.

[0072] FIG. 20 is a set of four bar graphs showing indel formation frequency of Cas-beta variants M43 and M67, as well as versions of M43 and M67 that were further altered to contain Q572R / F607S substitutions.

[0073] FIG.21 is a schematic diagram of optimized plasmid vector that can express Cas-beta2 nucease and Cas-beta2 variants in mammalian cells and generate genome edits therein.

[0074] FIG. 22 is a set of four bar graphs showing indel formation frequency of Cas-beta variants M43, M67, M43 (Q572R / F607S), and M67 (Q572R / F607S) when Cas-beta variants are delivered to mammalian cells by mRNA transfection.

[0075] FIG. 23 is a bar graph that shows mean percentage of GFP-positive cells for each indicated nuclease variant as well as control reactions where the HDR template encoding the EGFP gene or single guide RNA were omitted.

[0076] FIG. 24 is a set of three bar graphs showing DNA editing activity (indel formation) in mammalian cells of engineered guide RNA variants and wild type guide RNA.

[0077] FIG. 25 is a schematic diagram showing steps of an RNP pulldown assay for eliminating CRISPR-Cas binding events that do not result in substrate cleavage.

[0078] FIG.26 is a bar graph that shows HDR genome editing efficiency of an engineered Cas- beta2 variant disclosed herein when used with indicated concentrations of single stranded oligodeoxynucleotide (ODN) template.

[0079] Nucleic acid sequences listed in the accompanying sequence listing and referenced herein are shown using standard letter abbreviations for nucleotide bases. While only one strand of each nucleic acid sequence is shown, the complementary strand is understood to beincluded in any reference to the displayed strand. Sequence listings are described in the following Table 1. TABLE 1 Sequence Description SEQ ID NO:1 Wild type Cas-beta2SEQ ID NO:2 SV40 nuclear localization amino acid se uenceSEQ ID NO:47 S383R Cas-beta2 variant protein SEQ ID NO:48 V499R Cas-beta2 variant protein SEQ ID NO:49 N478R Cas-beta2 variant proteinSEQ ID NO:100 Cas-beta2 sgRNA deletion variant D27 SEQ ID NO:101 Cas-beta2 sgRNA deletion variant D28 SEQ ID NO:102 Cas-beta2 sgRNA deletion variant D29SEQ ID NO:190 Cas-beta2 M67(Q572R / F607D) variant protein SEQ ID NO:191 Cas-beta2 M67(Q572R / F607Q) variant proteinSEQ ID NO:233 Reverse primer for RunX1 region amplification for T7 Endonuclease I assay SEQ ID NO:234 Forward primer for WTAP region amplification for T7 Endonuclease I assaySEQ ID NO:276 Reverse primer 1 for RunX1 region amplification for Illumina SEQ SEQ ID NO:277 Reverse primer 2 for RunX1 region amplification for Illumina SEQ r or s 'SEQ ID NO:317 DNA sequence of ssODN HDR templateDETAILED

[0080] Compositions and methods are provided for novel CRISPR effector systems and elements comprising such systems, including, but not limiting to, novel guide polynucleotide / endonuclease complexes, guide polynucleotides, guide RNA elements, Cas polypeptides, and endonucleases, as well as proteins comprising an endonuclease functionality (domain). Compositions and methods are also provided for direct delivery of endonucleases, cleavage ready complexes, guide RNAs, and guide RNA / Cas endonuclease complexes. The present disclosure further includes compositions and methods for genome modification of a target sequence in the genome of a cell, for gene editing, and for inserting a polynucleotide of interest into the genome of a cell.

[0081] Terms used in the claims and specification are defined as set forth below unless otherwise specified. It must be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless the context clearly dictates otherwise. Definitions

[0082] As used herein, “nucleic acid” means a polynucleotide and includes a single or a double-stranded polymer of deoxyribonucleotide or ribonucleotide bases. Nucleic acids may also include fragments and modified nucleotides. Thus, the terms “polynucleotide”, “nucleic acid sequence”, “nucleotide sequence” and “nucleic acid fragment” are used interchangeably to denote a polymer of RNA and / or DNA and / or RNA-DNA that is single- or double-stranded, optionally comprising synthetic, non-natural, or altered nucleotide bases. Nucleotides (usually found in their 5’-monophosphate form) are referred to by their single letter designation as follows: “A” for adenosine or deoxyadenosine (for RNA or DNA, respectively), “C” for cytosine or deoxycytosine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purines (A or G), “Y” for pyrimidines (C or T), “K” for G or T, “H” for A or C or T, “I” for inosine, and “N” for any nucleotide.

[0083] The term “genome” as it applies to a prokaryotic and eukaryotic cell or organism cells encompasses not only chromosomal DNA found within the nucleus, but organelle DNA found within subcellular components (e.g., mitochondria, or plastid) of the cell.

[0084] “Open reading frame” is abbreviated ORF.

[0085] By “homology” is meant DNA sequences that are similar. For example, a “region of homology to a genomic region” that is found on the donor DNA is a region of DNA that has a similar sequence to a given “genomic region” in the cell or organism genome. A region of homology can be of any length that is sufficient to promote homologous recombination at the cleaved target site. For example, the region of homology can comprise at least 5-10, 5-15, 5- 20, 5-25, 5-30, 5-35, 5-40, 5-45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5-85, 5-90, 5-95, 5- 100, 5-200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5- 2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases in length such that the region of homology has sufficient homology to undergo homologous recombination with the corresponding genomic region. “Sufficient homology” indicates that two polynucleotide sequences have sufficient structural similarity to act as substrates for a homologous recombination reaction. The structural similarity includes overall length of each polynucleotide fragment, as well as the sequence similarity of the polynucleotides. Sequence similarity can be described by the percent sequence identity over the whole length of the sequences, and / or by conserved regions comprising localized similarities such as contiguous nucleotides having 100% sequence identity, and percent sequence identity over a portion of the length of the sequences.

[0086] As used herein, a “genomic region” is a segment of a chromosome in the genome of a cell that is present on either side of the target site or, alternatively, also comprises a portion of the target site. The genomic region can comprise at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5- 40, 5-45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5- 400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5- 1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5- 2700, 5-2800. 5-2900, 5-3000, 5-3100 or more bases such that the genomic region has sufficient homology to undergo homologous recombination with the corresponding region of homology.

[0087] As used herein, “homologous recombination” (HR) includes the exchange of DNA fragments between two DNA molecules at the sites of homology. The frequency of homologous recombination is influenced by a number of factors. Different organisms vary with respect to the amount of homologous recombination and the relative proportion of homologous to non-homologous recombination. Generally, the length of the region of homology affects the frequency of homologous recombination events: the longer the region of homology, the greater the frequency. The length of the homology region needed to observe homologousrecombination is also species-variable. In many cases, at least 5 kb of homology has been utilized, but homologous recombination has been observed with as little as 25-50 bp of homology. See, for example, Singer et al., (1982) Cell 31:25-33; Shen and Huang, (1986) Genetics 112:441-57; Watt et al., (1985) Proc. Natl. Acad. Sci. USA 82:4768-72, Sugawara and Haber, (1992) Mol Cell Biol 12:563-75, Rubnitz and Subramani, (1984) Mol Cell Biol 4:2253-8; Ayares et al., (1986) Proc. Natl. Acad. Sci. USA 83:5199-203; Liskay et al., (1987) Genetics 115:161-7.

[0088] “Sequence identity” or “identity” in the context of nucleic acid or polypeptide sequences refers to the nucleic acid bases or amino acid residues in two sequences that are the same when aligned for maximum correspondence over a specified comparison window.

[0089] The term “percentage of sequence identity” refers to the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the results by 100 to yield the percentage of sequence identity. Useful examples of percent sequence identities include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage from 50% to 100%. These identities can be determined using any of the programs described herein.

[0090] Sequence alignments and percent identity or similarity calculations may be determined using a variety of comparison methods designed to detect homologous sequences including,but not limited to, the MegAlign program of the LASERGENE bioinformatics computingsuite (DNASTAR Inc., Madison, WI). Within the context of this application it will be understood that where sequence analysis software is used for analysis, that the results of the analysis will be based on the “default values” of the program referenced, unless otherwise specified. As used herein “default values” will mean any set of values or parameters that originally load with the software when first initialized.

[0091] The “Clustal V method of alignment” corresponds to the alignment method labeled Clustal V (described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992)Comput Appl Biosci 8:189-191) and found in the MegAlign program of the LASERGENEbioinformatics computing suite (DNASTAR Inc., Madison, WI). For multiple alignments, the default values correspond to GAP PENALTY=10 and GAP LENGTH PENALTY=10. Default parameters for pairwise alignments and calculation of percent identity of protein sequences using the Clustal method are KTUPLE=1, GAP PENALTY=3, WINDOW=5 and DIAGONALS SAVED=5. For nucleic acids these parameters are KTUPLE=2, GAP PENALTY=5, WINDOW=4 and DIAGONALS SAVED=4. After alignment of the sequences using the Clustal V program, it is possible to obtain a “percent identity” by viewing the “sequence distances” Table in the same program. The “Clustal W method of alignment” corresponds to the alignment method labeled Clustal W (described by Higgins and Sharp, (1989) CABIOS 5:151-153; Higgins et al., (1992) Comput Appl Biosci 8:189-191) and foundin the MegAlign v6.1 program of the LASERGENE bioinformatics computing suite(DNASTAR Inc., Madison, WI). Default parameters for multiple alignment (GAP PENALTY=10, GAP LENGTH PENALTY=0.2, Delay Divergen Seqs (%)=30, DNA Transition Weight=0.5, Protein Weight Matrix=Gonnet Series, DNA Weight Matrix=IUB). After alignment of the sequences using the Clustal W program, it is possible to obtain a “percent identity” by viewing the “sequence distances” Table in the same program. Unless otherwise stated, sequence identity / similarity values provided herein refer to the value obtained using GAP Version 10 (GCG, Accelrys, San Diego, CA) using the following parameters:% identity and % similarity for a nucleotide sequence using a gap creation penalty weight of 50 and a gap length extension penalty weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for an amino acid sequence using a GAP creation penalty weight of 8 and a gap length extension penalty of 2, and the BLOSUM62 scoring matrix (Henikoff and Henikoff, (1989) Proc. Natl. Acad. Sci. USA 89:10915). GAP uses the algorithm of Needleman and Wunsch, (1970) J Mol Biol 48:443-53, to find an alignment of two complete sequences that maximizes the number of matches and minimizes the number of gaps. GAP considers all possible alignments and gap positions and creates the alignment with the largest number of matched bases and the fewest gaps, using a gap creation penalty and a gap extension penalty in units of matched bases. “BLAST” is a searching algorithm provided by the National Center for Biotechnology Information (NCBI) used to find regions of similarity between biological sequences. The program compares nucleotide or protein sequences to sequence databases and calculates the statistical significance of matches to identify sequences having sufficient similarity to a query sequence such that the similarity would not be predicted to have occurredrandomly. BLAST reports the identified sequences and their local alignment to the query sequence.

[0092] An "isolated" or "purified" nucleic acid molecule, polynucleotide, polypeptide, or protein, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polynucleotide or protein as found in its naturally occurring environment. Thus, an isolated or purified polynucleotide or polypeptide or protein is substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Optimally, an "isolated" polynucleotide is free of sequences (optimally protein encoding sequences) that naturally flank the polynucleotide (i.e., sequences located at the 5' and 3' ends of the polynucleotide) in the genomic DNA of the organism from which the polynucleotide is derived. For example, in various aspects, the isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequence that naturally flank the polynucleotide in genomic DNA of the cell from which the polynucleotide is derived. Isolated polynucleotides may be purified from a cell in which they naturally occur. Conventional nucleic acid purification methods known to skilled artisans may be used to obtain isolated polynucleotides. The term also embraces recombinant polynucleotides and chemically synthesized polynucleotides.

[0093] The term “fragment” refers to a contiguous set of nucleotides or amino acids. In one aspect, a fragment is 2, 3, 4, 5, 6, 78, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or greater than 20 contiguous nucleotides. In one aspect, a fragment is 2, 3, 4, 5, 6, 78, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or greater than 20 contiguous amino acids. A fragment may or may not exhibit the function of a sequence sharing some percent identity over the length of said fragment.

[0094] The terms “fragment that is functionally equivalent” and “functionally equivalent fragment” are used interchangeably herein. These terms refer to a portion or subsequence of an isolated nucleic acid fragment or polypeptide that displays the same activity or function as the longer sequence from which it derives. In one example, the fragment retains the ability to alter gene expression or produce a certain phenotype whether or not the fragment encodes an active protein. For example, the fragment can be used in the design of genes to produce the desired phenotype in a modified plant. Genes can be designed for use in suppression by linking a nucleic acid fragment, whether or not it encodes an active enzyme, in the sense or antisense orientation relative to a plant promoter sequence.

[0095] “Gene” includes a nucleic acid fragment that expresses a functional molecule such as, but not limited to, a specific protein, including regulatory sequences preceding (5’ non-coding sequences) and following (3’ non-coding sequences) the coding sequence. “Native gene” refers to a gene as found in its natural endogenous location with its own regulatory sequences.

[0096] By the term “endogenous” it is meant a sequence or other molecule that naturally occurs in a cell or organism. In some aspects, an endogenous polynucleotide is normally found in the genome of a cell; that is, the endogenous polynucleotide is not heterologous.

[0097] An “allele” is one of several alternative forms of a gene occupying a given locus on a chromosome.

[0098] “Coding sequence” refers to a polynucleotide sequence which codes for a specific amino acid sequence. “Regulatory sequences” refer to nucleotide sequences located upstream (5’ non-coding sequences), within, or downstream (3’ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leader sequences, 5’ untranslated sequences, 3’ untranslated sequences, introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0099] A “mutated gene” is a gene that has been altered through human intervention. Such a “mutated gene” has a sequence that differs from the sequence of the corresponding non- mutated gene by at least one nucleotide addition, deletion, or substitution. In certain aspects of the disclosure, the mutated gene comprises an alteration that results from a guide polynucleotide / Cas endonuclease system as disclosed herein. A mutated plant is a plant comprising a mutated gene.

[0100] As used herein, a “targeted mutation” is a mutation in a gene (referred to as the target gene), including a native gene, that was made by altering a target sequence within the target gene using any method known to one skilled in the art, including a method involving a guided Cas endonuclease system as disclosed herein.

[0101] The terms “knock-out”, “gene knock-out” and “genetic knock-out” are used interchangeably herein. A knock-out represents a DNA sequence of a cell that has been rendered partially or completely inoperative by targeting with a Cas polypeptide; for example, a DNA sequence prior to knock-out could have encoded an amino acid sequence, or could have had a regulatory function (e.g., promoter).

[0102] The terms “knock-in”, “gene knock-in, “gene insertion” and “genetic knock-in” are used interchangeably herein. A knock-in represents the replacement or insertion of a DNAsequence at a specific DNA sequence in cell by targeting with a Cas polypeptide (for example by homologous recombination (HR), wherein a suitable donor DNA polynucleotide is also used). examples of knock-ins are a specific insertion of a heterologous amino acid coding sequence in a coding region of a gene, or a specific insertion of a transcriptional regulatory element in a genetic locus.

[0103] By “domain” it is meant a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA-DNA-combination sequence) or amino acids.

[0104] A “codon-modified gene” or “codon-preferred gene” or “codon-optimized gene” is a gene having its frequency of codon usage designed to mimic the frequency of preferred codon usage of the host cell.

[0105] An “optimized” polynucleotide is a sequence that has been optimized for improved expression in a particular heterologous host cell.

[0106] A “promoter” is a region of DNA involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. The promoter sequence consists of proximal and more distal upstream elements, the latter elements often referred to as enhancers. An “enhancer” is a DNA sequence that can stimulate promoter activity, and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue- specificity of a promoter. Promoters may be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, and / or comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental conditions. It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of some variation may have identical promoter activity.

[0107] Promoters that cause a gene to be expressed in most cell types at most times are commonly referred to as “constitutive promoters”. The term “inducible promoter” refers to a promoter that selectively express a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, for example by chemical compounds (chemical inducers) or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulated promoters include, for example, promoters induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, phytohormones, wounding, or chemicals such as ethanol, abscisic acid (ABA), jasmonate, salicylic acid, or safeners.

[0108] “Translation leader sequence” refers to a polynucleotide sequence located between the promoter sequence of a gene and the coding sequence. The translation leader sequence is present in the mRNA upstream of the translation start sequence. The translation leader sequence may affect processing of the primary transcript to mRNA, mRNA stability or translation efficiency. Examples of translation leader sequences have been described (e.g., Turner and Foster, (1995) Mol Biotechnol 3:225-236).

[0109] “3’ non-coding sequences”, “transcription terminator” or “termination sequences” refer to DNA sequences located downstream of a coding sequence and include polyadenylation recognition sequences and other sequences encoding regulatory signals capable of affecting mRNA processing or gene expression. The polyadenylation signal is usually characterized by affecting the addition of polyadenylic acid tracts to the 3’ end of the mRNA precursor. The use of different 3’ non-coding sequences is exemplified by Ingelbrecht et al., (1989) Plant Cell 1:671-680.

[0110] “RNA transcript” refers to the product resulting from RNA polymerase-catalyzed transcription of a DNA sequence. When the RNA transcript is a perfect complimentary copy of the DNA sequence, it is referred to as the primary transcript or pre-mRNA. A RNA transcript is referred to as the mature RNA or mRNA when it is a RNA sequence derived from post- transcriptional processing of the primary transcript pre-mRNA. “Messenger RNA” or “mRNA” refers to the RNA that is without introns and that can be translated into protein by the cell. “cDNA” refers to a DNA that is complementary to, and synthesized from, an mRNA template using the enzyme reverse transcriptase. The cDNA can be single-stranded or converted into double-stranded form using the Klenow fragment of DNA polymerase I. “Sense” RNA refers to RNA transcript that includes the mRNA and can be translated into protein within a cell or in vitro. “Antisense RNA” refers to an RNA transcript that is complementary to all or part of a target primary transcript or mRNA, and that blocks the expression of a target gene (see, e.g., U.S. Patent No. 5,107,065). The complementarity of an antisense RNA may be with any part of the specific gene transcript, i.e., at the 5’ non-coding sequence, 3’ non-coding sequence, introns, or the coding sequence. “Functional RNA” refers to antisense RNA, ribozyme RNA, or other RNA that may not be translated but yet has an effect on cellular processes. The terms “complement” and “reverse complement” are used interchangeably herein with respect to mRNA transcripts, and are meant to define the antisense RNA of the message.

[0111] The term "genome" refers to the entire complement of genetic material (genes and non- coding sequences) that is present in each cell of an organism, or virus or organelle; and / or a complete set of chromosomes inherited as a (haploid) unit from one parent.

[0112] The term “operably linked” refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is regulated by the other. For example, a promoter is operably linked with a coding sequence when it is capable of regulating the expression of that coding sequence (i.e., the coding sequence is under the transcriptional control of the promoter). Coding sequences can be operably linked to regulatory sequences in a sense or antisense orientation. In another example, the complementary RNA regions can be operably linked, either directly or indirectly, 5’ to the target mRNA, or 3’ to the target mRNA, or within the target mRNA, or a first complementary region is 5’ and its complement is 3’ to the target mRNA.

[0113] Generally, “host” refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, a "host cell" refers to an in vivo or in vitro eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell), or cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, into which a heterologous polynucleotide or polypeptide has been introduced. In some aspects, the cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-cell organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, in invertebrate cell, a vertebrate cell, a fish cell, a frog cell, a bird cell, an insect cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some cases, the cell is in vitro. In some cases, the cell is in vivo.

[0114] The term “recombinant” refers to an artificial combination of two otherwise separated segments of sequence, e.g., by chemical synthesis, or manipulation of isolated segments of nucleic acids by genetic engineering techniques.

[0115] The terms “plasmid”, “vector” and “cassette” refer to a linear or circular extra chromosomal element often carrying genes that are not part of the central metabolism of the cell, and usually in the form of double-stranded DNA. Such elements may be autonomously replicating sequences, genome integrating sequences, phage, or nucleotide sequences, in linear or circular form, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a polynucleotide of interest into a cell. “Transformation cassette” refers to a specific vector comprising a gene and having elements inaddition to the gene that facilitates transformation of a particular host cell. “Expression cassette” refers to a specific vector comprising a gene and having elements in addition to the gene that allow for expression of that gene in a host.

[0116] The terms “recombinant DNA molecule”, “recombinant DNA construct”, “expression construct”, “construct”, and “recombinant construct” are used interchangeably herein. A recombinant DNA construct comprises an artificial combination of nucleic acid sequences, e.g., regulatory and coding sequences that are not all found together in nature. For example, a recombinant DNA construct may comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source, but arranged in a manner different than that found in nature. Such a construct may be used by itself or may be used in conjunction with a vector. If a vector is used, then the choice of vector is dependent upon the method that will be used to introduce the vector into the host cells as is well known to those skilled in the art. For example, a plasmid vector can be used. The skilled artisan is well aware of the genetic elements that must be present on the vector in order to successfully transform, select and propagate host cells. The skilled artisan will also recognize that different independent transformation events may result in different levels and patterns of expression (Jones et al., (1985) EMBO J 4:2411-2418; De Almeida et al., (1989) Mol Gen Genetics 218:78-86), and thus that multiple events are typically screened in order to obtain lines displaying the desired expression level and pattern. Such screening may be accomplished standard molecular biological, biochemical, and other assays including Southern analysis of DNA, Northern analysis of mRNA expression, PCR, real time quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting analysis of protein expression, enzyme or activity assays, and / or phenotypic analysis.

[0117] The term “heterologous” refers to the difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. Non-limiting examples include differences in taxonomic derivation (e.g., a polynucleotide sequence obtained from E.coli would be heterologous if inserted into the genome of a mammal, or of a different bacterium (other than E.coli); or a polynucleotide obtained from a bacterium was introduced into a cell of a plant), or sequence (e.g., a polynucleotide sequence obtained from Zea mays, isolated, modified, and re-introduced into a maize plant). As used herein, “heterologous” in reference to a sequence can refer to a sequence that originates from a different species, variety, foreign species, or, if from the same species, is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, a promoter operably linked toa heterologous polynucleotide is from a species different from the species from which the polynucleotide was derived, or, if from the same / analogous species, one or both are substantially modified from their original form and / or genomic locus, or the promoter is not the native promoter for the operably linked polynucleotide. Alternatively, one or more regulatory region(s) and / or a polynucleotide provided herein may be entirely synthetic. In another example, a target polynucleotide for cleavage by a Cas endonuclease may be of a different organism than that of the Cas endonuclease. In another example, a Cas endonuclease and guide RNA may be introduced to a target polynucleotide with an additional polynucleotide that acts as a template or donor for insertion into the target polynucleotide, wherein the additional polynucleotide is heterologous to the target polynucleotide and / or the Cas endonuclease.

[0118] The term “expression”, as used herein, refers to the production of a functional end- product (e.g., an mRNA, guide RNA, or a protein) in either precursor or mature form.

[0119] A “mature” protein refers to a post-translationally processed polypeptide (i.e., one from which any pre- or propeptides present in the primary translation product have been removed).

[0120] “Precursor” protein refers to the primary product of translation of mRNA (i.e., with pre- and propeptides still present). Pre- and propeptides may be but are not limited to intracellular localization signals.

[0121] “CRISPR” (Clustered Regularly Interspaced Short Palindromic Repeats) loci refers to certain genetic loci encoding components of DNA cleavage systems, for example, used by bacterial and archaeal cells to destroy foreign DNA (Horvath and Barrangou, 2010, Science 327:167-170; WO2007025097, published 01 March 2007). A CRISPR locus can consist of a CRISPR array, comprising short direct repeats (CRISPR repeats) separated by short variable DNA sequences (called spacers), which can be flanked by diverse Cas (CRISPR-associated) genes.

[0122] As used herein, an “effector” or “effector protein” is a protein that encompasses an activity including recognizing, binding to, and / or cleaving or nicking a polynucleotide target. An effector, or effector protein, may also be an endonuclease. The “effector complex” of a CRISPR system includes Cas polypeptides involved in crRNA and target recognition and binding. Some of the component Cas polypeptides may additionally comprise domains involved in target polynucleotide cleavage.

[0123] The term “Cas polypeptide” refers to a polypeptide encoded by a Cas (CRISPR- associated) gene. A Cas polypeptide includes proteins encoded by a gene in a cas locus, and include adaptation molecules as well as interference molecules. An interference molecule of abacterial adaptive immunity complex includes endonucleases. A Cas endonuclease described herein can comprise one or more nuclease domains. A Cas polypeptide may be a “Cas endonuclease” or “Cas effector protein” that, when in complex with a suitable polynucleotide component, is capable of recognizing, binding to, and optionally nicking or cleaving all or part of a specific polynucleotide target sequence. A Cas polypeptide is further defined as a functional fragment or functional variant of a native Cas polypeptide, or a protein that shares at least 50%, between 50% and 55%, at least 55%, between 55% and 60%, at least 60%, between 60% and 65%, at least 65%, between 65% and 70%, at least 70%, between 70% and 75%, at least 75%, between 75% and 80%, at least 80%, between 80% and 85%, at least 85%, between 85% and 90%, at least 90%, between 90% and 95%, at least 95%, between 95% and 96%, at least 96%, between 96% and 97%, at least 97%, between 97% and 98%, at least 98%, between 98% and 99%, at least 99%, between 99% and 100%, or 100% sequence identity with at least 50, between 50 and 100, at least 100, between 100 and 150, at least 150, between 150 and 200, at least 200, between 200 and 250, at least 250, between 250 and 300, at least 300, between 300 and 350, at least 350, between 350 and 400, at least 400, between 400 and 450, at least 500, or greater than 500 contiguous amino acids of a native Cas polypeptide, and retains at least partial activity of the native sequence.

[0124] A “functional fragment”, “fragment that is functionally equivalent” and “functionally equivalent fragment” of a Cas endonuclease are used interchangeably herein, and refer to a portion or subsequence of the Cas endonuclease of the present disclosure in which the ability to recognize, bind to, and optionally unwind, nick or cleave (introduce a single or double-strand break in) the target site is retained.

[0125] The terms “functional variant”, “variant that is functionally equivalent” and “functionally equivalent variant” of a Cas endonuclease or Cas effector protein, including Cas- beta variant described herein, are used interchangeably herein, and refer to a variant of the Cas- beta variant effector protein disclosed herein in which the ability to recognize, bind to, and optionally unwind, nick or cleave all or part of a target sequence is retained.

[0126] A Cas endonuclease may also include a multifunctional Cas endonuclease. The term “multifunctional Cas endonuclease” and “multifunctional Cas endonuclease polypeptide” are used interchangeably herein and includes reference to a single polypeptide that has Cas endonuclease functionality (comprising at least one protein domain that can act as a Cas endonuclease) and at least one other functionality, such as but not limited to, the functionality to form a complex (comprises at least a second protein domain that can form a complex with other proteins). In some aspects, the multifunctional Cas endonuclease comprises at least oneadditional protein domain relative (either internally, upstream (5’), downstream (3’), or both internally 5’ and 3’, or any combination thereof) to those domains typical of a Cas endonuclease.

[0127] The terms “cascade” and “cascade complex” are used interchangeably herein and include reference to a multi-subunit protein complex that can assemble with a polynucleotide forming a polynucleotide-protein complex (PNP). Cascade is a PNP that relies on the polynucleotide for complex assembly and stability, and for the identification of target nucleic acid sequences. Cascade functions as a surveillance complex that finds and optionally binds target nucleic acids that are complementary to a variable targeting domain of the guide polynucleotide.

[0128] The terms ”cleavage-ready Cascade”, “crCascade”, ”cleavage-ready Cascade complex”, “crCascade complex”, ”cleavage-ready Cascade system”, “CRC” and “crCascade system”, are used interchangeably herein and include reference to a multi-subunit protein complex that can assemble with a polynucleotide forming a polynucleotide-protein complex (PNP), wherein one of the cascade proteins is a Cas endonuclease capable of recognizing, binding to, and optionally unwinding, nicking, or cleaving all or part of a target sequence.

[0129] The terms “5’-cap” and “7-methylguanylate (m7G) cap” are used interchangeably herein. A 7-methylguanylate residue is located on the 5′ terminus of messenger RNA (mRNA) in eukaryotes. RNA polymerase II (Pol II) transcribes mRNA in eukaryotes. Messenger RNA capping occurs generally as follows: the most terminal 5’ phosphate group of the mRNA transcript is removed by RNA terminal phosphatase, leaving two terminal phosphates. A guanosine monophosphate (GMP) is added to the terminal phosphate of the transcript by a guanylyl transferase, leaving a 5′-5′ triphosphate-linked guanine at the transcript terminus. Finally, the 7-nitrogen of this terminal guanine is methylated by a methyl transferase.

[0130] The terminology “not having a 5’-cap” herein is used to refer to RNA having, for example, a 5’-hydroxyl group instead of a 5’-cap. Such RNA can be referred to as “uncapped RNA”, for example. Uncapped RNA can better accumulate in the nucleus following transcription, since 5’-capped RNA is subject to nuclear export. One or more RNA components herein are uncapped.

[0131] As used herein, the term “guide polynucleotide”, relates to a polynucleotide sequence that can form a complex with a Cas endonuclease, including the Cas endonuclease described herein, and enables the Cas endonuclease to recognize, optionally bind to, and optionally cleave a DNA target site. The guide polynucleotide sequence can be a RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence).

[0132] The terms “functional fragment”, “fragment that is functionally equivalent” and “functionally equivalent fragment” of a guide RNA, crRNA or tracrRNA are used interchangeably herein, and refer to a portion or subsequence of the guide RNA, crRNA or tracrRNA, respectively, of the present disclosure in which the ability to function as a guide RNA, crRNA or tracrRNA, respectively, is retained.

[0133] The terms “functional variant”, “variant that is functionally equivalent” and “functionally equivalent variant” of a guide RNA, crRNA or tracrRNA (respectively) are used interchangeably herein, and refer to a variant of the guide RNA, crRNA or tracrRNA, respectively, of the present disclosure in which the ability to function as a guide RNA, crRNA or tracrRNA, respectively, is retained.

[0134] The terms “single guide RNA” and “sgRNA” are used interchangeably herein and relate to a synthetic fusion of two RNA molecules, a crRNA (CRISPR RNA) comprising a variable targeting domain (linked to a tracr mate sequence that hybridizes to a tracrRNA), fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA can comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of the type V CRISPR / Cas system that can form a complex with a type V Cas endonuclease, wherein said guide RNA / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, optionally bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site.

[0135] The term “variable targeting domain” or “VT domain” is used interchangeably herein and includes a nucleotide sequence that can hybridize (is complementary) to one strand (nucleotide sequence) of a double strand DNA target site. The percent complementation between the first nucleotide sequence domain (VT domain) and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. The variable targeting domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. In some aspects, the variable targeting domain comprises a contiguous stretch of 12 to 30 nucleotides. The variable targeting domain can be composed of a DNA sequence, a RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.

[0136] The term “Cas endonuclease recognition domain” or “CER domain” (of a guide polynucleotide) is used interchangeably herein and includes a nucleotide sequence that interacts with a Cas endonuclease polypeptide. A CER domain comprises a (trans-acting)tracrNucleotide mate sequence followed by a tracrNucleotide sequence. The CER domain can be composed of a DNA sequence, a RNA sequence, a modified DNA sequence, a modified RNA sequence (see for example US20150059010A1, published 26 February 2015), or any combination thereof.

[0137] As used herein, the terms “guide polynucleotide / Cas endonuclease complex”, “guide polynucleotide / Cas endonuclease system”, “ guide polynucleotide / Cas complex”, “guide polynucleotide / Cas system” and “guided Cas system” “Polynucleotide-guided endonuclease” , “PGEN” are used interchangeably herein and refer to at least one guide polynucleotide and at least one Cas endonuclease, that are capable of forming a complex, wherein said guide polynucleotide / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site. A guide polynucleotide / Cas endonuclease complex herein can comprise Cas polypeptide(s) and suitable polynucleotide component(s) of any of the known CRISPR systems (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13).

[0138] The terms “guide RNA / Cas polypeptide complex”, “guide RNA / Cas polypeptide system”, “guide RNA / Cas complex”, “guide RNA / Cas system”, “gRNA / Cas complex”, “gRNA / Cas system”, “RNA-guided polypeptide” , “RGEN” are used interchangeably herein and refer to at least one RNA component and at least one Cas polypeptide that are capable of forming a complex. When said guide RNA / Cas polypeptide includes a Cas endonuclease, the complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site.

[0139] The terms “target site”, “target sequence”, “target site sequence, ”target DNA”, “target locus”, “genomic target site”, “genomic target sequence”, “genomic target locus” and “protospacer”, are used interchangeably herein and refer to a polynucleotide sequence such as, but not limited to, a nucleotide sequence on a chromosome, episome, a locus, or any other DNA molecule in the genome (including chromosomal, chloroplastic, mitochondrial DNA, plasmid DNA) of a cell, at which a guide polynucleotide / Cas polypeptide complex can recognize, bind to, and optionally nick or cleave . The target site can be an endogenous site in the genome of a cell, or alternatively, the target site can be heterologous to the cell and thereby not be naturally occurring in the genome of the cell, or the target site can be found in a heterologous genomic location compared to where it occurs in nature. As used herein, terms “endogenous targetsequence” and “native target sequence” are used interchangeable herein to refer to a target sequence that is endogenous or native to the genome of a cell and is at the endogenous or native position of that target sequence in the genome of the cell. An “artificial target site” or “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence can be identical in sequence to an endogenous or native target sequence in the genome of a cell but be located in a different position (i.e., a non-endogenous or non-native position) in the genome of a cell.

[0140] A “protospacer adjacent motif” (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas polypeptide system described herein. The Cas polypeptide may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM herein can differ depending on the Cas polypeptide or Cas polypeptide complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.

[0141] An “altered target site”, “altered target sequence”, “modified target site”, “modified target sequence” are used interchangeably herein and refer to a target sequence as disclosed herein that comprises at least one alteration when compared to non-altered target sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, (iv) a chemical alteration of at least one nucleotide, or (v) any combination of (i) – (iv).

[0142] A “modified nucleotide” or “edited nucleotide” refers to a nucleotide sequence of interest that comprises at least one alteration when compared to its non-modified nucleotide sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, (iv) a chemical alteration of at least one nucleotide, or (v) any combination of (i) – (iv).

[0143] Methods for “modifying a target site” and “altering a target site” are used interchangeably herein and refer to methods for producing an altered target site.

[0144] As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of a Cas endonuclease.

[0145] The term “polynucleotide modification template” includes a polynucleotide that comprises at least one nucleotide modification when compared to the nucleotide sequence to be edited. A nucleotide modification can be at least one nucleotide substitution, addition or deletion. Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the at least one nucleotide modification, whereinthe flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.

[0146] The term “codon-optimized Cas endonuclease” herein refers to a Cas polypeptide, including a multifunctional Cas polypeptide, encoded by a nucleotide sequence that has been optimized for expression in a host cell, e.g., a mammalian or plant host cell.

[0147] A “codon-optimized nucleotide sequence encoding a Cas endonuclease”, “codon- optimized construct encoding a Cas endonuclease” and a “codon-optimized polynucleotide encoding a Cas endonuclease” are used interchangeably herein and refer to a nucleotide sequence encoding a Cas polypeptide, or a variant or functional fragment thereof, that has been optimized for expression in a host cell. In some aspects, the codon-optimized Cas polypeptide nucleotide sequence is a mammalian-optimized, human-optimized, maize-optimized, rice- optimized, wheat-optimized, soybean-optimized, cotton-optimized, or canola-optimized Cas polypeptide.

[0148] “Progeny” comprises any subsequent generation of an organism.

[0149] As used herein, the term “plant part” refers to plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant calli, plant clumps, and plant cells that are intact in plants or parts of plants such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, stalks, roots, root tips, anthers, and the like, as well as the parts themselves. Grain is intended to mean the mature seed produced by commercial growers for purposes other than growing or reproducing the species. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the disclosure, provided that these parts comprise the introduced polynucleotides.

[0150] The term “non-conventional yeast” herein refers to any yeast that is not a Saccharomyces (e.g., S. cerevisiae) or Schizosaccharomyces yeast species. (see “Non- Conventional Yeasts in Genetics, Biochemistry and Biotechnology:Practical Protocols”, K. Wolf, K.D. Breunig, G. Barth, Eds., Springer-Verlag, Berlin, Germany, 2003).

[0151] The term “crossed” or “cross” or “crossing” in the context of this disclosure means the fusion of gametes via pollination to produce progeny (i.e., cells, seeds, or plants). The term encompasses both sexual crosses (the pollination of one plant by another) and selfing (self- pollination, i.e., when the pollen and ovule (or microspores and megaspores) are from the same plant or genetically identical plants).

[0152] The term “introgression” refers to the transmission of a desired allele of a genetic locus from one genetic background to another. For example, introgression of a desired allele at a specified locus can be transmitted to at least one progeny plant via a sexual cross between twoparent plants, where at least one of the parent plants has the desired allele within its genome. Alternatively, for example, transmission of an allele can occur by recombination between two donor genomes, e.g., in a fused protoplast, where at least one of the donor protoplasts has the desired allele in its genome. The desired allele can be, e.g., a transgene, a modified (mutated or edited) native allele, or a selected allele of a marker or QTL.

[0153] The term “isoline” is a comparative term, and references organisms that are genetically identical, but differ in treatment. In one example, two genetically identical mammalian cell cultures or maize plant embryos may be separated into two different groups, one receiving a treatment (such as the introduction of a CRISPR-Cas effector endonuclease) and one control that does not receive such treatment. Any phenotypic differences between the two groups may thus be attributed solely to the treatment and not to any inherency of the plant's endogenous genetic makeup.

[0154] "Introducing" is intended to mean presenting to a target, such as a cell or organism, a polynucleotide or polypeptide or polynucleotide-protein complex, in such a manner that the component(s) gains access to the interior of a cell of the organism or to the cell itself.

[0155] The terms "decreased," "fewer," "slower" and "increased" "faster" "enhanced" "greater" as used herein refers to a decrease or increase in a characteristic of the modified organism or cell as compared to an unmodified organism or cell. For example, a decrease in a characteristic may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400%) or more lower than the untreated control and an increase may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400% or more higher than the untreated control.

[0156] As used herein, the term “before”, in reference to a sequence position, refers to an occurrence of one sequence upstream, or 5’, to another sequence.

[0157] The meaning of abbreviations is as follows: “sec” means second(s), “min” means minute(s), “h” means hour(s), “d” means day(s), “μL” means microliter(s), “mL” means milliliter(s), “L” means liter(s), “μM” means micromolar, “mM” means millimolar, “M” means molar, “mmol” means millimole(s), “μmole” or “umole” mean micromole(s), “g” means gram(s), “μg” or “ug” means microgram(s), “ng” means nanogram(s), “U” means unit(s), “bp” means base pair(s) and “kb” means kilobase(s). Classification of CRISPR-Cas Systems

[0158] CRISPR-Cas systems have been classified according to sequence and structural analysis of components. Multiple CRISPR / Cas systems have been described including Class 1 systems, with multisubunit effector complexes (comprising type I, type III, and type IV), and Class 2 systems, with single protein effectors (comprising type II, type V, and type VI) (Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13; Haft et al., 2005, Computational Biology, PLoS Comput Biol 1(6):e60; and Koonin et al. 2017, Curr Opinion Microbiology 37:67-78).

[0159] A CRISPR-Cas system comprises, at a minimum, a CRISPR RNA (crRNA) molecule and at least one CRISPR-associated (Cas) protein to form crRNA ribonucleoprotein (crRNP) effector complexes. CRISPR-Cas loci comprise an array of identical repeats interspersed with DNA-targeting spacers that encode the crRNA components and an operon-like unit of cas genes encoding the Cas polypeptide components. The resulting ribonucleoprotein complex recognizes a polynucleotide in a sequence-specific manner (Jore et al., Nature Structural & Molecular Biology 18, 529–536 (2011)). The crRNA serves as a guide RNA for sequence specific binding of the effector (protein or complex) to double strand DNA sequences, by forming base pairs with the complementary DNA strand while displacing the noncomplementary strand to form a so-called R-loop. (Jore et al., 2011. Nature Structural & Molecular Biology 18, 529–536).

[0160] RNA transcripts of CRISPR loci (pre-crRNA) are cleaved specifically in the repeat sequences by CRISPR associated (Cas) endoribonucleases in type I and type III systems or by RNase III in type II systems. The number of CRISPR-associated genes at a given CRISPR locus can vary between species.

[0161] Different cas genes that encode proteins with different domains are present in different CRISPR systems. The cas operon comprises genes that encode for one or more effector endonucleases, as well as other Cas polypeptides. Protein subunits include those described inMakarova et al. 2011, Nat Rev Microbiol. 20119(6):467–477; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; and Koonin et al. 2017, Current Opinion Microbiology 37:67-78). The types of domains include those involved in Expression (pre-crRNA processing, for example Cas 6 or RNaseIII), Interference (including an effector module for crRNA and target binding, as well as domain(s) for target cleavage), Adaptation (spacer insertion, for example Cas1 or Cas2), and Ancillary (regulation or helper or unknown function). Some domains may serve more than one purpose, for example Cas9 comprises domains for endonuclease functionality as well as for target cleavage, among others.

[0162] The Cas endonuclease is guided by a single CRISPR RNA (crRNA) through direct RNA-DNA base-pairing to recognize a DNA target site that is in close vicinity to a protospacer adjacent motif (PAM) (Jore, M.M. et al., 2011, Nat. Struct. Mol. Biol. 18:529-536, Westra, E.R. et al., 2012, Molecular Cell 46:595-605, and Sinkunas, T. et al., 2013, EMBO J.32:385- 394). Class I CRISPR-Cas Systems

[0163] Class I CRISPR-Cas systems comprise Types I, III, and IV. A characteristic feature of Class I systems is the presence of an effector endonuclease complex instead of a single protein. A Cascade complex comprises a RNA recognition motif (RRM) and a nucleic acid-binding domain that is the core fold of the diverse RAMP (Repeat-Associated Mysterious Proteins) protein superfamily (Makarova et al. 2013, Biochem Soc Trans 41, 1392-1400; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15). RAMP protein subunits include Cas5 and Cas7 (which comprise the skeleton of the crRNA–effector complex), wherein the Cas5 subunit binds the 5’ handle of the crRNA and interacts with the large subunit, and often includes Cas6 which is loosely associated with the effector complex and typically functions as the repeat-specific RNase in the pre-crRNA processing (Charpentier et al., FEMS Microbiol Rev 2015, 39:428-441; Niewoehner et al., RNA 2016, 22:318-329). Class II CRISPR-Cas Systems

[0164] Class II CRISPR-Cas systems comprise Types II, V, and VI. A characteristic feature of Class II systems is the presence of a single Cas effector protein instead of an effector complex. Types II and V Cas polypeptides comprise an RuvC endonuclease domain that adopts the RNase H fold.

[0165] Type II CRISPR / Cas systems employ a crRNA and tracrRNA (trans-activating CRISPR RNA) to guide the Cas endonuclease to its DNA target. The crRNA comprises a spacer region complementary to one strand of the double strand DNA target and a region thatbase pairs with the tracrRNA (trans-activating CRISPR RNA) forming a RNA duplex that directs the Cas endonuclease to cleave the DNA target, leaving a blunt end. Spacers are acquired through a not fully understood process involving Cas1 and Cas2 proteins. Type II CRISPR / Cas loci typically comprise cas1 and cas2 genes in addition to the cas9 gene (Chylinski et al., 2013, RNA Biology 10:726-737; Makarova et al. 2015, Nature Reviews Microbiology Vol.13:1-15). Type II CRISR-Cas loci can encode a tracrRNA, which is partially complementary to the repeats within the respective CRISPR array, and can comprise other proteins such as Csn1 and Csn2. The presence of cas9 in the vicinity of cas1 and cas2 genes is the hallmark of type II loci (Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1- 15).

[0166] Type V CRISPR / Cas systems comprise a single Cas endonuclease, including Cpf1 (Cas12) (Koonin et al., Curr Opinion Microbiology 37:67-78, 2017), that is an active RNA- guided endonuclease that does not necessarily require the additional trans-activating CRISPR (tracr) RNA for target cleavage, unlike Cas9.

[0167] Type VI CRISPR-Cas systems comprise a cas13 gene that encodes a nuclease with two HEPN (Higher Eukaryotes and Prokaryotes Nucleotide-binding) domains but no HNH or RuvC domains, and are not dependent upon tracrRNA activity. The majority of HEPN domains comprise conserved motifs that constitute a metal-independent endoRNase active site (Anantharam et al., Biol Direct 8:15, 2013). Because of this feature, it is thought that type VI systems act on RNA targets instead of the DNA targets that are common to other CRISPR-Cas systems.

[0168] In a first aspect, the disclosure provides methods for altering protospacer adjacent motif (PAM) specificity of a target Cas-beta polypeptide. As used herein, a “protospacer adjacent motif” (PAM) refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that can be recognized (targeted) by a guide polynucleotide / Cas polypeptide system. The Cas polypeptide may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM herein can differ depending on the Cas polypeptide or Cas polypeptide complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.

[0169] In some aspects, a method for altering PAM specificity of a target Cas-beta polypeptide comprises: (a) comparing the PAM interacting (PI) domain of a heterologous, orthologous Cas- beta polypeptide with the PI domain of the target Cas-beta polypeptide, wherein the orthologous Cas-beta polypeptide has different PAM recognition than the target Cas-betapolypeptide; (b) selecting one or more amino acids and / or one or more polypeptide chains from the PI domain of the orthologous Cas-beta polypeptide; (c) incorporating the one or more amino acids and / or the one or more polypeptide chains selected from the PI domain of the orthologous Cas-beta polypeptide into one or more structurally similar positions of the target Cas-beta polypeptide resulting in a modified target Cas-beta polypeptide; and (d) determining PAM recognition in the modified target Cas-beta polypeptide.

[0170] As used herein the target Cas-beta polypeptide can be cas-beta2 or any one of the Cas- beta2 variants disclosed herein. As used herein, an “orthologous Cas-beta polypeptide” or “Cas-beta ortholog” refers to a Cas-beta polypeptides: Cas-beta1, Cas-beta3, Cas-beta4 or any other member of the Cas12l family (See Urbaitis et al.2022 EMBO Reports, 23(12): e55481).

[0171] As used herein, a “structurally similar position” refers to a coordinate in a polypeptide that occupies a similar three-dimensional location when its structure, either predicted (e.g. neural network-based informatic models such as AlphaFold) or determined (e.g. using cryogenic electron microscopy (Cryo-EM), X-ray crystallography, and NMR spectroscopy), is aligned or superimposed with an orthologous structure either predicted or determined.

[0172] Methods for comparing the PI domain of a target Cas-beta polypeptide and an orthologous Cas-beta polypeptide include, but are not limited to, multiple sequence comparison by log-expectation (MUSCLE), multiple sequence comparison using Clustal Omega, root mean square distance (RMSD), distance matrix alignment (DALI), structural homology by environment-based alignment (SHEBA), combinatorial extension (CE), homologous structure alignment database (HOMSTRAD), structural classification of proteins (SCOP), FatCat, and PhyreStorm.

[0173] In some aspects, selecting one or more amino acids from the PI domain of an orthologous Cas-beta polypeptide comprises comparing the PI domain sequence and / or structure of the orthologous Cas-beta polypeptide with that from the target Cas-beta polypeptide whose PAM specificity is to be altered and substituting one or more amino acids from the orthologous Cas-beta polypeptide into the Cas-beta polypeptide target.

[0174] In some aspects, selecting one or more polypeptide chains from the PI domain of an orthologous Cas-beta polypeptide comprises comparing the PI domain sequence and / or structure of the orthologous Cas-beta polypeptide with that from the target Cas-beta polypeptide whose PAM specificity is to be altered and substituting one or more polypeptide chains from the orthologous Cas-beta polypeptide into the Cas-beta polypeptide target.

[0175] Methods for determining the PAM recognition of a modified target Cas-beta polypeptide include, but are not limited to, transcription and translation of the modified Cas-beta polypeptide in a cell or in a cell-free mixture, complexing the modified Cas-beta polypeptide with a guide RNA to form a ribonucleoprotein (RNP), incubating the RNP with DNA species that contain a fixed guide RNA target and a collection of different PAM sequences, capturing DNA molecules that support target cleavage, sequencing the PAM region from DNA species that support target cleavage, and calculating a consensus PAM using a position frequency matrix to summarize alterations to PAM specificity.

[0176] In some aspects of the method, the target Cas-beta polypeptide has an amino acid sequence with at least 50%, alternatively at least 55%, alternatively at least 60%, alternatively at least 65%, alternatively at least 70%, alternatively at least 75%, alternatively at least 80%, alternatively at least 85%, alternatively at least 90% alternatively at least 95%, alternatively at least 96%, alternatively at least 97%, alternatively at least 98%, alternatively at least 99%, alternatively 100% sequence identity to SEQ ID NO:1, SEQ ID NO:129, SEQ ID NO:130, or SEQ ID NO:131.

[0177] In another aspect, the disclosure provides synthetic Cas-beta polypeptides having altered PAM specificity. More specifically, disclosed herein are Cas-beta polypeptides comprising modified PAM interacting (PI) domains such that the resulting Cas-beta polypeptides recognize a PAM sequence in a target polynucleotide other than 5’-CCY-3’. CRISPR-Cas System Components Cas polypeptides

[0178] A number of proteins may be encoded in the CRISPR cas operon, including those involved in adaptation (spacer insertion), interference (effector module target binding, target nicking or cleavage – e.g. endonuclease activity), expression (pre-crRNA processing), regulation, or other.

[0179] Two proteins, Cas1 and Cas2, are conserved among many CRISPR systems (for example, as described in Koonin et al., Curr Opinion Microbiology 37:67-78, 2017). Cas1 is a metal-dependent DNA-specific endonuclease that produces double-stranded DNA fragments. In some systems Cas1 forms a stable complex with Cas2, which is essential to spacer acquisition and insertion for CRISPR systems (Nuñez et al., Nature Str Mol Biol 21:528-534, 2014).

[0180] A number of other proteins have been identified across different systems, including Cas4 (which may have similarity to a RecB nuclease) and is thought to play a role in the capture of new viral DNA sequences for incorporation into the CRISPR array (Zhang et al., PLOS One 7(10):e47232, 2012).

[0181] Some proteins may encompass a plurality of functions. For example, Cas9, the signature protein of Class 2 type II systems, has been demonstrated to be involved in pre- crRNA processing, target binding, as well as target cleavage. Cas Endonucleases and Effectors

[0182] Endonucleases are enzymes that cleave the phosphodiester bond within a polynucleotide chain, and include restriction endonucleases that cleave DNA at specific sites without damaging the bases. Examples of endonucleases include restriction endonucleases, meganucleases, TAL effector nucleases (TALENs), zinc finger nucleases, and Cas (CRISPR- associated) effector endonucleases.

[0183] Cas endonucleases, either as single effector proteins or in an effector complex with other components, unwind the DNA duplex at the target sequence and optionally cleave at least one DNA strand, as mediated by recognition of the target sequence by a polynucleotide (such as, but not limited to, a crRNA or guide RNA) that is in complex with the Cas effector protein. Such recognition and cutting of a target sequence by a Cas endonuclease typically occurs if the correct protospacer-adjacent motif (PAM) is located at or adjacent to the 3' end of the DNA target sequence. Alternatively, a Cas endonuclease herein may lack DNA cleavage or nicking activity, but can still specifically bind to a DNA target sequence when complexed with a suitable RNA component. (See also U.S. Patent Application US20150082478 published 19 March 2015 and US20150059010 published 26 February 2015).

[0184] Cas endonucleases may occur as individual effectors (Class 2 CRISPR systems) or as part of larger effector complexes (Class I CRISPR systems).

[0185] Cas endonucleases that have been described include, but are not limited to, for example: Cas9, Cas12f (Cas-alpha, Cas14), Cas12l (Cas-beta), Cas12a (Cpf1), Cas12b (a C2c1 protein), Cas13 (a C2c2 protein), Cas12c (a C2c3 protein), Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas3, Cas3-HD, Cas 5, Cas6, Cas7, Cas8, Cas10, or combinations or complexes of these. Cas endonucleases and effector proteins can be used for targeted genome editing (via simplex and multiplex double-strand breaks and nicks) and targeted genome regulation (via tethering of epigenetic effector domains to either the Cas polypeptide or sgRNA. A Cas endonuclease can also be engineered to function as an RNA-guided recombinase, and via RNA tethers could serve as a scaffold for the assembly of multiprotein and nucleic acid complexes (Mali et al., 2013, Nature Methods Vol.10:957-963).Cas-beta Endonucleases

[0186] A Cas-beta endonuclease (also known as Cas12l) has been characterized as a functional RNA-guided, PAM-dependent dsDNA cleavage protein of fewer than 900 amino acids, comprising: a C-terminal region comprising a RuvC catalytic domain split into three subdomains and an N-terminal region comprising an “Oligonucleotide Binding Domain” (OBD) domain split by two helical regions containing a bridge-helix (BH)-like motif and a helix-turn-helix (HTH) DNA binding domain. Urbaitis et al. 2022 EMBO Reports, 23(12): e55481.

[0187] Cas-beta endonucleases are RNA-guided endonucleases capable of binding to, and cleaving, a double-strand DNA target that comprises: (1) a sequence sharing homology with a nucleotide sequence of the guide RNA, and (2) a PAM sequence. In some aspects, the PAM is C-rich, e.g., comprising 5’-CCY-3’.

[0188] In some aspects, a catalytically inactive Cas-beta polypeptide may be used with a functional endonuclease, to cleave a target sequence. In some aspects, a catalytically inactive Cas-beta polypeptide may be combined with a base editing molecule, such as a deaminase. A “deaminase” is an enzyme that catalyzes a deamination reaction. For example, deamination of adenine with an adenine deaminase results in the formation of inosine. Inosine selectively base pairs with cytosine instead of thymine. This results in a post-replicative transition mutation, such that the original A – T base pair transforms into a G – C base pair. In another example, cytosine deamination results in the formation of uracil, which can be repaired by cellular repair mechanisms back to a C – T base pair or to a T – A, G – C, or A – T base pair. This heterogeneity in repair can be suppressed by the introduction of a uracil glycosylase inhibitor, such that DNA repair or replication transforms the original C – T base pair into a T – A base pair (Burnett et al. (2022) Frontiers in Genome Editing. 4, 923718). In the case of both adenine and cytosine deaminases, the introduction of a nick promotes the respective base pair change (Burnett et al., 2022). In some aspects, a deaminase may be a cytidine deaminase. In some aspects, a deaminase may be an adenine deaminase. In some aspects, a deaminase may be ADAR-2.

[0189] A “functional fragment” of a Cas-beta endonuclease retains the ability to recognize, or bind, or nick a single strand of a double-stranded polynucleotide, or cleave both strands of a double-stranded polynucleotide, or any combination of the preceding.

[0190] A Cas endonuclease, effector protein, or functional fragment thereof, for use in the disclosed methods, can be isolated from a native source, or from, a recombinant source where the genetically modified host cell is modified to express the nucleic acid sequence encoding the protein. Alternatively, the Cas polypeptide can be produced using cell free proteinexpression systems, or be synthetically produced. Effector Cas nucleases may be isolated and introduced into a heterologous cell, or may be modified from its native form to exhibit a different type or magnitude of activity than what it would exhibit in its native source. Such modifications include but are not limited to: fragments, variants, substitutions, deletions, and insertions.

[0191] Fragments and variants of Cas endonucleases and Cas effector proteins can be obtained via methods such as site-directed mutagenesis and synthetic construction. Methods for measuring endonuclease activity are well known in the art such as, but not limiting to, WO2013166113 published 07 November 2013, WO2016186953 published 24 November 2016, and WO2016186946 published 24 November 2016.

[0192] The Cas endonuclease can comprise a modified form of the Cas polypeptide. The modified form of the Cas polypeptide can include an amino acid change (e.g., deletion, insertion, or substitution) that reduces the naturally-occurring nuclease activity of the Cas polypeptide. For example, in some instances, the modified form of the Cas polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas polypeptide (US20140068797 published 06 March 2014). In some cases, the modified form of the Cas polypeptide has no substantial nuclease activity and is referred to as catalytically “inactivated Cas” or “deactivated Cas (dCas).” An inactivated Cas / deactivated Cas includes a deactivated Cas endonuclease (dCas). A catalytically inactive Cas effector protein can be fused to a heterologous sequence to induce or modify activity.

[0193] A Cas endonuclease can be part of a fusion protein comprising one or more heterologous protein domains (e.g., 1, 2, 3, or more domains in addition to the Cas polypeptide). Such a fusion protein may comprise any additional protein sequence, and optionally a linker sequence between any two domains, such as between Cas and a first heterologous domain. Examples of protein domains that may be fused to a Cas polypeptide herein include, without limitation, epitope tags (e.g., histidine [His], V5, FLAG, influenza hemagglutinin [HA], myc, VSV-G, thioredoxin [Trx]), reporters (e.g., glutathione-5- transferase [GST], horseradish peroxidase [HRP], chloramphenicol acetyltransferase [CAT], beta-galactosidase, beta-glucuronidase [GUS], luciferase, green fluorescent protein [GFP], HcRed, DsRed, cyan fluorescent protein [CFP], yellow fluorescent protein [YFP], blue fluorescent protein [BFP]), and domains having one or more of the following activities: transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity (e.g., VP16 or VP64), transcription repression activity, transcription releasefactor activity, histone modification activity, RNA cleavage activity and nucleic acid binding activity. A Cas polypeptide can also be in fusion with a protein that binds DNA molecules or other molecules, such as maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD), GAL4A DNA binding domain, and herpes simplex virus (HSV) VP16.

[0194] A catalytically active and / or inactive Cas endonuclease can be fused to a heterologous sequence (US20140068797 published 06 March 2014). Suitable fusion partners, e.g., for a catalytically inactive Cas polypeptide can include, but are not limited to, a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA. Additional suitable fusion partners include, but are not limited to, a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity. Other suitable fusion partners for catalytically inactive Cas polypeptide can include a polypeptide that provides (i) transposase activity to enable targeted transposition (insertion of DNA payload or transposon) to a target site or (ii) integrase activity to enable targeted integration of DNA at a target site. Further suitable fusion partners include, but are not limited to, a polypeptide that directly provides for increased transcription of the target nucleic acid (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription regulator, etc.). A partially active or catalytically inactive Cas-alpha endonuclease can also be fused to another protein or domain, for example Clo51 or FokI nuclease, to generate double- strand breaks (Guilinger et al. Nature Biotechnology, volume 32, number 6, June 2014).

[0195] A catalytically active or inactive Cas polypeptide, such as the Cas-alpha polypeptide described herein, can also be in fusion with a base editor molecule that directs editing of single or multiple bases in a polynucleotide sequence, for example a site-specific deaminase that can change the identity of a nucleotide, for example from C•G to T•A or an A•T to G•C (Gaudelli et al., Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” Science 353 (6305) (2016); Komor et al. “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage.” Nature 533 (7603) (2016):420-4. A base editing fusion protein may comprise, for example, an active (double strand break creating), partially active (nickase) or deactivated(catalytically inactive) Cas-beta endonuclease and a deaminase (such as, but not limited to, a cytidine deaminase, an adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABEs, or the like). Base edit repair inhibitors and glycosylase inhibitors (e.g., uracil glycosylase inhibitor (to prevent uracil removal)) are contemplated as other components of a base editing system, in some aspects.

[0196] The Cas endonucleases described herein can be expressed and purified by methods known in the art, for example as described in WO / 2016 / 186953 published 24 November 2016.

[0197] Many Cas endonucleases have been described to date that can recognize specific PAM sequences (WO2016186953 published 24 November 2016, WO2016186946 published 24 November 2016, and Zetsche B et al. 2015. Cell 163, 1013) and cleave the target DNA at a specific position. It is understood that based on the methods and aspects described herein utilizing a novel guided Cas system one skilled in the art can now tailor these methods such that they can utilize any guided endonuclease system.

[0198] A Cas effector protein can comprise a heterologous nuclear localization sequence (NLS). A heterologous NLS amino acid sequence herein may be of sufficient strength to drive accumulation of a Cas polypeptide in a detectable amount in the nucleus of a yeast cell herein, for example. An NLS may comprise one (monopartite) or more (e.g., bipartite) short sequences (e.g., 2 to 20 residues) of basic, positively charged residues (e.g., lysine and / or arginine), and can be located anywhere in a Cas amino acid sequence but such that it is exposed on the protein surface. An NLS may be operably linked to the N-terminus or C-terminus of a Cas polypeptide herein, for example. Two or more NLS sequences can be linked to a Cas polypeptide, for example, such as on both the N- and C-termini of a Cas polypeptide. The Cas endonuclease gene can be operably linked to a SV40 nuclear targeting signal upstream of the Cas codon region and a bipartite VirD2 nuclear localization signal (Tinland et al. (1992) Proc. Natl. Acad. Sci. USA 89:7442-6) downstream of the Cas codon region. Non-limiting examples of suitable NLS sequences herein include those disclosed in U.S. Patent Nos.6,660,830 and 7,309,576. Guide Polynucleotides

[0199] The guide polynucleotide enables target recognition, binding, and optionally cleavage by the Cas polypeptide, and can be a single molecule or a double molecule. The guide polynucleotide sequence can be a RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence). Optionally, the guide polynucleotide can comprise at least one nucleotide, phosphodiester bond or linkage modification such as, but not limited, to Locked Nucleic Acid (LNA), 5-methyl dC, 2,6-Diaminopurine, 2’-Fluoro A, 2’-Fluoro U, 2'-O-Methyl RNA, phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 (hexaethylene glycol chain) molecule, or 5’ to 3’ covalent linkage resulting in circularization. A guide polynucleotide that solely comprises ribonucleic acids is also referred to as a “guide RNA” or “gRNA” (US20150082478 published 19 March 2015 and US20150059010 published 26 February 2015). A guide polynucleotide may be engineered or synthetic.

[0200] A guide polynucleotide includes a chimeric non-naturally occurring guide RNA comprising regions that are not found together in nature (i.e., they are heterologous with each other). For example, a chimeric non-naturally occurring guide RNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA, linked to a second nucleotide sequence that can recognize the Cas endonuclease, such that the first and second nucleotide sequence are not found linked together in nature.

[0201] A guide polynucleotide can be a double molecule (also referred to as duplex guide polynucleotide) comprising a crNucleotide sequence (such as a crRNA) and a tracrNucleotide (such as a tracrRNA) sequence. In some cases, there is a linker polynucleotide that connects the crRNA and tracrRNA to form a single guide, for example an sgRNA.

[0202] In some aspects, the crNucleotide includes a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA and a second nucleotide sequence (also referred to as a tracr mate sequence) that is part of a Cas endonuclease recognition (CER) domain. The tracr mate sequence can hybridized to a tracrNucleotide along a region of complementarity and together form the Cas endonuclease recognition domain or CER domain. The CER domain is capable of interacting with a Cas endonuclease polypeptide. The crNucleotide and the tracrNucleotide of the duplex guide polynucleotide can be RNA, DNA, and / or RNA-DNA- combination sequences. In some aspects, the crNucleotide molecule of the duplex guide polynucleotide is referred to as “crDNA” (when composed of a contiguous stretch of DNA nucleotides) or “crRNA” (when composed of a contiguous stretch of RNA nucleotides), or “crDNA-RNA” (when composed of a combination of DNA and RNA nucleotides). The crNucleotide can comprise a fragment of the crRNA naturally occurring in Bacteria and Archaea. The size of the fragment of the crRNA naturally occurring in Bacteria and Archaea that can be present in a crNucleotide disclosed herein can range from, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9,10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides.

[0203] In some aspects, the tracrNucleotide is referred to as “tracrRNA” (when composed of a contiguous stretch of RNA nucleotides) or “tracrDNA” (when composed of a contiguous stretch of DNA nucleotides) or “tracrDNA-RNA” (when composed of a combination of DNA and RNA nucleotides. In one aspect, the RNA that guides the RNA / Cas polypeptide complex is a duplexed RNA comprising a duplex crRNA-tracrRNA. The tracrRNA (trans-activating CRISPR RNA) comprises, in the 5’-to-3’ direction, (i) a sequence that anneals with the repeat region of CRISPR type II crRNA and (ii) a stem loop-comprising portion (Deltcheva et al., Nature 471:602-607). The duplex guide polynucleotide can form a complex with a Cas endonuclease, wherein said guide polynucleotide / Cas polypeptide complex (also referred to as a guide polynucleotide / Cas polypeptide system) can direct the Cas endonuclease to a genomic target site, enabling the Cas polypeptide to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) into the target site. (US20150082478 published 19 March 2015 and US20150059010 published 26 February 2015).

[0204] In some aspects, the guide polynucleotide is a guide polynucleotide capable of forming a PGEN as described herein, wherein said guide polynucleotide comprises a first nucleotide sequence domain that is complementary to a nucleotide sequence in a target DNA, and a second nucleotide sequence domain that interacts with said Cas polypeptide.

[0205] In some aspects, the guide polynucleotide is a guide polynucleotide described herein, wherein the first nucleotide sequence and the second nucleotide sequence domain is selected from the group consisting of a DNA sequence, a RNA sequence, and a combination thereof.

[0206] In some aspects,, the guide polynucleotide is a guide polynucleotide described herein, wherein the first nucleotide sequence and the second nucleotide sequence domain is selected from the group consisting of RNA backbone modifications that enhance stability, DNA backbone modifications that enhance stability, and a combination thereof (see Kanasty et al., 2013, Common RNA-backbone modifications, Nature Materials 12:976-977; US20150082478 published 19 March 2015 and US20150059010 published 26 February 2015)

[0207] The guide RNA includes a dual molecule comprising a chimeric non-naturally occurring crRNA linked to at least one tracrRNA. A chimeric non-naturally occurring crRNA includes a crRNA that comprises regions that are not found together in nature (i.e., they are heterologous with each other. For example, a crRNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA, linked to a second nucleotide sequence (also referred to as a tracr mate sequence) such that the first and second sequence are not found linked together in nature.

[0208] The guide polynucleotide can also be a single molecule (also referred to as single guide polynucleotide) comprising a crNucleotide sequence linked to a tracrNucleotide sequence. The single guide polynucleotide comprises a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA and a Cas endonuclease recognition domain (CER domain), that interacts with a Cas endonuclease polypeptide.

[0209] The VT domain and / or the CER domain of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA-combination sequence. The single guide polynucleotide being comprised of sequences from the crNucleotide and the tracrNucleotide may be referred to as “single guide RNA” (when composed of a contiguous stretch of RNA nucleotides) or “single guide DNA” (when composed of a contiguous stretch of DNA nucleotides) or “single guide RNA-DNA” (when composed of a combination of RNA and DNA nucleotides). The single guide polynucleotide can form a complex with a Cas endonuclease, wherein said guide polynucleotide / Cas endonuclease complex (also referred to as a guide polynucleotide / Cas endonuclease system) can direct the Cas endonuclease to a genomic target site, enabling the Cas endonuclease to recognize, bind to, and optionally nick or cleave (introduce a single or double-strand break) the target site. (US20150082478 published 19 March 2015 and US20150059010 published 26 February 2015).

[0210] A chimeric non-naturally occurring single guide RNA (sgRNA) includes a sgRNA that comprises regions that are not found together in nature (i.e., they are heterologous with each other. For example, a sgRNA comprising a first nucleotide sequence domain (referred to as Variable Targeting domain or VT domain) that can hybridize to a nucleotide sequence in a target DNA linked to a second nucleotide sequence (also referred to as a tracr mate sequence) that are not found linked together in nature.

[0211] The nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA combination sequence. In one aspect, the nucleotide sequence linking the crNucleotide and the tracrNucleotide of a single guide polynucleotide (also referred to as “loop”) can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleotides in length. In another aspect, the nucleotide sequence linking the crNucleotide and thetracrNucleotide of a single guide polynucleotide can comprise a tetraloop sequence, such as, but not limiting to a GAAA tetraloop sequence.

[0212] The guide polynucleotide can be produced by any method known in the art, including chemically synthesizing guide polynucleotides (such as but not limiting to Hendel et al.2015, Nature Biotechnology 33, 985–989), in vitro generated guide polynucleotides, and / or self- splicing guide RNAs (such as but not limited to Xie et al.2015, PNAS 112:3570-3575). Guide Polynucleotide / Cas Polypeptide Complexes

[0213] A guide polynucleotide / Cas polypeptide complex described herein is capable of recognizing, binding to, and optionally nicking, unwinding, or cleaving all or part of a target sequence.

[0214] A guide polynucleotide / Cas polypeptide complex that can cleave both strands of a DNA target sequence typically comprises a Cas polypeptide that has all of its endonuclease domains in a functional state (e.g., wild type endonuclease domains or variants thereof retaining some or all activity in each endonuclease domain). Thus, a wild type Cas polypeptide (e.g., a Cas- beta2 polypeptide disclosed herein), or a variant thereof retaining some or all activity in each endonuclease domain of the Cas polypeptide, is a suitable example of a Cas endonuclease that can cleave both strands of a DNA target sequence.

[0215] A guide polynucleotide / Cas endonuclease complex in certain aspects can bind to a DNA target site sequence, but does not cleave any strand at the target site sequence. Such a complex may comprise a Cas polypeptide in which all of its nuclease domains are mutant, dysfunctional. For example, a Cas9 protein that can bind to a DNA target site sequence, but does not cleave any strand at the target site sequence, may comprise both a mutant, dysfunctional RuvC domain and a mutant, dysfunctional HNH domain. A Cas polypeptide herein that binds, but does not cleave, a target DNA sequence can be used to modulate gene expression, for example, in which case the Cas polypeptide could be fused with a transcription factor (or portion thereof) (e.g., a repressor or activator, such as any of those disclosed herein).

[0216] In some aspects, the guide polynucleotide / Cas polypeptide complex (PGEN) described herein is a PGEN, wherein said Cas polypeptide is optionally covalently or non-covalently linked, or assembled to at least one protein subunit, or functional fragment thereof.

[0217] In some aspects, the guide polynucleotide / Cas polypeptide complex is a guide polynucleotide / Cas endonuclease complex (PGEN) comprising at least one guide polynucleotide and at least one Cas polypeptide, wherein said Cas polypeptide comprises at least one protein subunit, or a functional fragment thereof, wherein said guide polynucleotideis a chimeric non-naturally occurring guide polynucleotide, wherein said guide polynucleotide / Cas endonuclease complex is capable of recognizing, binding to, and optionally nicking, unwinding, or cleaving all or part of a target sequence.

[0218] In some aspects, the guide polynucleotide / Cas effector complex is a guide polynucleotide / Cas effector protein complex (PGEN) comprising at least one guide polynucleotide and a Cas-beta effector protein, wherein said guide polynucleotide / Cas effector protein complex is capable of recognizing, binding to, and optionally nicking, unwinding, or cleaving all or part of a target sequence.

[0219] The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein further comprises one copy or multiple copies of at least one protein subunit, or a functional fragment thereof. In some aspects, said protein subunit is selected from the group consisting of a Cas1 protein subunit, a Cas2 protein subunit, a Cas4 protein subunit, and any combination thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein further comprises at least two different protein subunits of selected from the group consisting of a Cas1, Cas2, and Cas4.

[0220] The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein further comprises at least three different protein subunits, or functional fragments thereof, selected from the group consisting of Cas1, Cas2, and one additional Cas polypeptide, optionally comprising Cas4.

[0221] In some aspects, the guide polynucleotide / Cas effector protein complex described herein is a PGEN, wherein said Cas effector protein is covalently or non-covalently linked to at least one protein subunit, or functional fragment thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein polypeptide is covalently or non-covalently linked, or assembled to one copy or multiple copies of at least one protein subunit, or a functional fragment thereof, selected from the group consisting of a Cas1 protein subunit, a Cas2 protein subunit, a one additional Cas polypeptide optionally comprising Cas4 protein subunit, and any combination thereof. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein is covalently or non-covalently linked or assembled to at least two different protein subunits selected from the group consisting of a Cas1, a Cas2, and one additional Cas polypeptide, optionally comprising Cas4. The PGEN can be a guide polynucleotide / Cas effector protein complex, wherein said Cas effector protein is covalently or non-covalently linked to at least three different protein subunits, or functional fragments thereof, selected from the group consistingof a Cas1, a Cas2, and one additional Cas polypeptide, optionally comprising Cas4, and any combination thereof.

[0222] Any component of the guide polynucleotide / Cas effector protein complex, the guide polynucleotide / Cas effector protein complex itself, as well as the polynucleotide modification template(s) and / or donor DNA(s), can be introduced into a heterologous cell or organism by any method known in the art. Recombinant Constructs for Transformation of Cells

[0223] The disclosed guide polynucleotides, Cas endonucleases, polynucleotide modification templates, donor DNAs, guide polynucleotide / Cas endonuclease systems disclosed herein, and any one combination thereof, optionally further comprising one or more polynucleotide(s) of interest, can be introduced into a cell. Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells as well as plants and seeds produced by the methods described herein.

[0224] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described more fully in Sambrook et al., Molecular Cloning:A Laboratory Manual; Cold Spring Harbor Laboratory:Cold Spring Harbor, NY (1989). Transformation methods are well known to those skilled in the art and are described infra.

[0225] Vectors and constructs include circular plasmids, and linear polynucleotides, comprising a polynucleotide of interest and optionally other components including linkers, adapters, regulatory or analysis. In some examples a recognition site and / or target site can be comprised within an intron, coding sequence, 5' UTRs, 3' UTRs, and / or regulatory regions.

[0226] NHEJ and HDR. In some aspects, the Cas polypeptide disclosed herein can be part of a genome editing system further comprising one or more guide polynucleotides and optionally donor DNA, and editing a target polynucleotide sequence comprises nonhomologous end- joining (NHEJ) or homologous recombination (HR) following a Cas endonuclease-mediated double-strand break. Once a double-strand break is induced in the DNA, the cell's DNA repair mechanism is activated to repair the break. The most common repair mechanism to bring the broken ends together is the nonhomologous end-joining pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). The structural integrity of chromosomes is typically preserved by the repair, but deletions, insertions, or other rearrangements are possible (Siebert and Puchta, (2002) Plant Cell 14:1121-31; Pacher et al., (2007) Genetics 175:21-9). Alternatively, the double-strand break can be repaired by homologous recombination between homologous DNA sequences. Once the sequence around the double-strand break is altered, for example, by exonucleaseactivities involved in the maturation of double-strand breaks, gene conversion pathways can restore the original structure if a homologous sequence is available, such as a homologous chromosome in non-dividing somatic cells, or a sister chromatid after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenic DNA sequences may also serve as a DNA repair template for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0227] As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of a Cas endonuclease. Once a double-strand break is introduced in the target site by the endonuclease, the first and second regions of homology of the donor DNA can undergo homologous recombination with their corresponding genomic regions of homology resulting in exchange of DNA between the donor and the target genome. As such, the provided methods result in the integration of the polynucleotide of interest of the donor DNA into the double-strand break in the target site in the plant genome, thereby altering the original target site and producing an altered genomic target site.

[0228] Base Editing. In some aspects, the Cas polypeptide disclosed herein can be part of a genome editing system further comprising a base editor and a plurality of guide polynucleotides and editing a target polynucleotide sequence comprises introducing a plurality of nucleobase edits in the target polynucleotide sequence resulting in a variant nucleotide sequence.

[0229] One or more nucleobases of a target polynucleotide can be chemically altered, in some cases to change the base from one type to another, for example from a Cytosine to a Thymine, or an Adenine to a Guanine. In some aspects, a plurality of bases, for example 2 or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more 90 or more, 100 or more, or even greater than 100, 200 or more, up to thousands of bases may be modified or altered, to produce a plant with a plurality of modified bases.

[0230] Any base editing complex, such as a base editor associated with an RNA-guided protein, may be used to target and bind to a desired locus in the genome of an organism and chemically modify one or more components of a target polynucleotide.

[0231] Site-specific base conversions can be achieved to engineer one or more nucleotide changes to create one or more edits into the genome. These include for example, a site-specific base edit mediated by an C•G to T•A or an A•T to G•C base editing deaminase enzymes (Gaudelli et al., Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” Science 353 (6305) (2016); Komor et al. “Programmable editing of a target base in genomic DNA without double-stranded DNAcleavage.” Nature 533 (7603) (2016):420-4. A catalytically “dead” or inactive Cas9 (dCas9), for example a catalytically inactive “dead” version of a Cas endonuclease disclosed herein, fused to a cytidine deaminase or an adenine deaminase protein becomes a specific base editor that can alter DNA bases without inducing a DNA break. Base editors convert C->T (or G->A on the opposite strand) or an adenine base editor that would convert adenine to inosine, resulting in an A->G change within an editing window specified by the gRNA. Any molecule that effects a change in a nucleobase is a “base editor” or “base editing agent”.

[0232] For many traits of interest, the creation of single double-strand breaks and the subsequent repair via HDR or NHEJ is not ideal for quantitative traits. An observed phenotype includes both genotype effects and environmental effects. The genotype effects further comprise additive effects, dominance effects, and epistatic effects. The probability of no effect per any single edit can be greater than zero, and any single phenotypic effect can be small, depending on the method used and site selected. Double-stranded break repair can additionally be “noisy” and have low repeatability.

[0233] One approach to ameliorate the probability of no effect per edit or small phenotypic effect outcome is to multiplex genome modification, such that a plurality of target sites are modified. Methods to modify a genomic sequence that do not introduce double-strand breaks would allow for single base substitutions. Combining these approaches, multiplexed base editing is beneficial for creating large numbers of genotype edits that can produce observable phenotype modifications. In some cases, dozens or hundreds or thousands of sites can be edited within one or a few generations of an organism.

[0234] A multiplexed approach to base editing in an organism, has the potential to create a plurality of significant phenotypic variations in one or a few generations, with a positive directional bias to the effects. In some aspects, the organism is a plant. A plant or a population of plants with a plurality of edits can be cross-bred to produce progeny plants, some of which will comprise multiple pluralities of edits from the parental lines. In this way, accelerated breeding of desired traits can be accomplished in parallel in one or a few generations, replacing time-consuming traditional sequential crossing and breeding across multiple generations.

[0235] A base editing deaminase, such as a cytidine deaminase or an adenine deaminase, may be fused to an RNA-guided endonuclease that can be deactivated (“dCas”, such as a deactivated Cas9) or partially active (“nCas”, such as a Cas9 nickase) so that it does not cleave a target site to which it is guided. The dCas forms a functional complex with a guide polynucleotide that shares homology with a polynucleotide sequence at the target site, and is further complexed with the deaminase molecule. The guided Cas endonuclease recognizes and binds to a double-stranded target sequence, opening the double-strand to expose individual bases. In the case of a cytidine deaminase, the deaminase deaminates the cytosine base and creates a uracil. Uracil glycosylase inhibitor (UGI) is provided to prevent the conversion of U back to C. DNA replication or repair mechanisms then convert the Uracil to a thymine (U to T), and subsequent repair of the opposing base (formerly G in the original G-C pair) to an Adenine, creating a T- A pair. For example, see Komor et al. Nature Volume 533, Pages 420-424, 19 May 2016.

[0236] Prime Editing. In some aspects, the Cas polypeptide described herein can be part of a genome editing system further comprising a prime editing agent and a guide polynucleotide and editing a target nucleotide sequence comprises introducing one or more insertions, deletions, or nucleobase swaps in a target nucleotide sequence without generating a double- stranded DNA break.

[0237] In some aspects, the prime editing agent is a Cas polypeptide disclosed herein fused to a reverse transcriptase, wherein the Cas polypeptide is modified to nick DNA rather than generating double-strand break. This Cas-polypeptide-reverse transcriptase fusion can also be referred to as a “prime editor” or “PE”. In some aspects, the guide polynucleotide comprises a prime editing guide polynucleotide (pegRNA), and is larger than standard sgRNAs commonly used for CRISPR gene editing (e.g., >100 nucleobases). The pegRNA comprises a primer binding sequence (PBS) and a template containing the desired or target RNA sequence at its 3’ end.

[0238] During prime editing, the PE:pegRNA complex binds to a target DNA sequence and the modified Cas polypeptide nicks one target DNA strand resulting in a flap. The PBS on the pegRNA binds to the DNA flap and the target RNA sequence is reverse transcribed using the reverse transcriptase. The edited strand is incorporated into the target DNA at the end of the nicked flap, and the target DNA sequence is repaired with the new reverse transcribed DNA. Components for Expression and Utilization of Novel CRISPR-Cas Systems in Prokaryotic and Eukaryotic cells.

[0239] The disclosure further provides expression constructs for expressing in a prokaryotic or eukaryotic cell / organism a guide RNA / Cas polypeptide system that is capable of recognizing, binding to, and optionally nicking, unwinding, or cleaving all or part of a target sequence.

[0240] In some aspects, the expression constructs of the disclosure comprise a promoter operably linked to a nucleotide sequence encoding a Cas polypeptide (e.g., a mammal- optimized coding sequence for a Cas-beta variant described herein) and a promoter operably linked to a guide RNA (e.g., sgRNA) of the present disclosure. The promoter is capable ofdriving expression of an operably linked nucleotide sequence in a prokaryotic or eukaryotic cell / organism.

[0241] Nucleotide sequence modification of the guide polynucleotide, VT domain and / or CER domain can be selected from, but not limited to , the group consisting of a 5' cap, a 3' polyadenylated tail, a riboswitch sequence, a stability control sequence, a sequence that forms a dsRNA duplex, a modification or sequence that targets the guide poly nucleotide to a subcellular location, a modification or sequence that provides for tracking , a modification or sequence that provides a binding site for proteins , a Locked Nucleic Acid (LNA), a 5-methyl dC nucleotide, a 2,6-Diaminopurine nucleotide, a 2’-Fluoro A nucleotide, a 2’-Fluoro U nucleotide; a 2'-O-Methyl RNA nucleotide, a phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 molecule, a 5’ to 3’ covalent linkage, or any combination thereof. These modifications can result in at least one additional beneficial feature, wherein the additional beneficial feature is selected from the group of a modified or regulated stability, a subcellular targeting, tracking, a fluorescent label, a binding site for a protein or protein complex, modified binding affinity to complementary target sequence, modified resistance to cellular degradation, and increased cellular permeability.

[0242] A method of expressing RNA components such as gRNA in eukaryotic cells for performing Cas9-mediated DNA targeting has been to use RNA polymerase III (Pol III) promoters, which allow for transcription of RNA with precisely defined, unmodified, 5’- and 3’-ends (DiCarlo et al., Nucleic Acids Res.41:4336-4343; Ma et al., Mol. Ther. Nucleic Acids 3:e161). This strategy has been successfully applied in cells of several different species including maize and soybean (US20150082478 published 19 March 2015). Methods for expressing RNA components that do not have a 5’ cap have been described (WO2016 / 025131 published 18 February 2016).

[0243] Various methods and compositions can be employed to obtain a cell or organism having a polynucleotide of interest inserted in a target site for a Cas endonuclease. Such methods can employ homologous recombination (HR) to provide integration of the polynucleotide of interest at the target site. In one method described herein, a polynucleotide of interest is introduced into the organism cell via a donor DNA construct.

[0244] The donor DNA construct further comprises a first and a second region of homology that flank the polynucleotide of interest. The first and second regions of homology of the donor DNA share homology to a first and a second genomic region, respectively, present in or flanking the target site of the cell or organism genome.

[0245] The donor DNA can be tethered to the guide polynucleotide. Tethered donor DNAs can allow for co-localizing target and donor DNA, useful in genome editing, gene insertion, and targeted genome regulation, and can also be useful in targeting post-mitotic cells where function of endogenous HR machinery is expected to be highly diminished (Mali et al., 2013, Nature Methods Vol.10:957-963).

[0246] The amount of homology or sequence identity shared by a target and a donor polynucleotide can vary and includes total lengths and / or regions having unit integral values in the ranges of about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200- 400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5–3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5- 7 kb, 4-8 kb, 5-10 kb, or up to and including the total length of the target site. These ranges include every integer within the range, for example, the range of 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 and 20 bps. The amount of homology can also be described by percent sequence identity over the full aligned length of the two polynucleotides which includes percent sequence identity at least of about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, between 98% and 99%, 99%, between 99% and 100%, or 100%. Sufficient homology includes any combination of polynucleotide length, global percent sequence identity, and optionally conserved regions of contiguous nucleotides or local percent sequence identity, for example sufficient homology can be described as a region of 75-150 bp having at least 80% sequence identity to a region of the target locus. Sufficient homology can also be described by the predicted ability of two polynucleotides to specifically hybridize under high stringency conditions, see, for example, Sambrook et al., (1989) Molecular Cloning:A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.); and, Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology--Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0247] The structural similarity between a given genomic region and the corresponding region of homology found on the donor DNA can be any degree of sequence identity that allows for homologous recombination to occur. For example, the amount of homology or sequence identity shared by the “region of homology” of the donor DNA and the “genomic region” of the organism genome can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%,84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity, such that the sequences undergo homologous recombination

[0248] The region of homology on the donor DNA can have homology to any sequence flanking the target site. While in some instances the regions of homology share significant sequence homology to the genomic sequence immediately flanking the target site, it is recognized that the regions of homology can be designed to have sufficient homology to regions that may be further 5' or 3' to the target site. The regions of homology can also have homology with a fragment of the target site along with downstream genomic regions

[0249] In one aspect, the first region of homology further comprises a first fragment of the target site and the second region of homology comprises a second fragment of the target site, wherein the first and second fragments are dissimilar. Polynucleotides of Interest

[0250] Polynucleotides of interest are further described herein and include polynucleotides reflective of the commercial markets and interests of those involved in medicine, pharmaceuticals, diagnostics, or agriculture.

[0251] General categories of polynucleotides of interest include, for example, genes of interest involved in information or signaling, such as zinc fingers, those involved in communication, such as kinases, and those involved in housekeeping, such as heat shock proteins. Polynucleotides of interest include, but are not limited to, those involved in disease.

[0252] Furthermore, it is recognized that the polynucleotide of interest may also comprise antisense sequences complementary to at least a portion of the messenger RNA (mRNA) for a targeted gene sequence of interest. Antisense nucleotides are constructed to hybridize with the corresponding mRNA. Modifications of the antisense sequences may be made as long as the sequences hybridize to and interfere with expression of the corresponding mRNA. In this manner, antisense constructions having 70%, 80%, or 85% sequence identity to the corresponding antisense sequences may be used. Furthermore, portions of the antisense nucleotides may be used to disrupt the expression of the target gene. Generally, sequences of at least 50 nucleotides, 100 nucleotides, 200 nucleotides, or greater may be used.

[0253] In addition, the polynucleotide of interest may also be used in the sense orientation to suppress the expression of endogenous genes in mammals or plants. Methods for suppressing gene expression in plants using polynucleotides in the sense orientation are known in the art. The methods generally involve transforming plants with a DNA construct comprising a promoter that drives expression in a plant operably linked to at least a portion of a nucleotidesequence that corresponds to the transcript of the endogenous gene. Typically, such a nucleotide sequence has substantial sequence identity to the sequence of the transcript of the endogenous gene, generally greater than about 65% sequence identity, about 85% sequence identity, or greater than about 95% sequence identity. See U.S. Patent Nos. 5,283,184 and 5,034,323.

[0254] The polynucleotide of interest can also be a disease or phenotypic marker. A phenotypic marker is screenable or a selectable marker that includes visual markers and selectable markers whether it is a positive or negative selectable marker. Any disease or phenotypic marker can be used.

[0255] Examples of selectable markers include, but are not limited to, DNA segments that comprise restriction enzyme sites; DNA segments that encode products which provide resistance against otherwise toxic compounds including antibiotics, such as, spectinomycin, ampicillin, kanamycin, tetracycline, Basta, neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT)); DNA segments that encode products which are otherwise lacking in the recipient cell (e.g., tRNA genes, auxotrophic markers); DNA segments that encode products which can be readily identified (e.g., phenotypic markers such as β- galactosidase, GUS; fluorescent proteins such as green fluorescent protein (GFP), cyan (CFP), yellow (YFP), red (RFP), and cell surface proteins); the generation of new primer sites for PCR (e.g., the juxtaposition of two DNA sequence not previously juxtaposed), the inclusion of DNA sequences not acted upon or acted upon by a restriction endonuclease or other DNA modifying enzyme, chemical, etc.; and, the inclusion of a DNA sequences required for a specific modification (e.g., methylation) that allows its identification.

[0256] Additional selectable markers include genes that confer resistance to herbicidal compounds, such as sulphonylureas, glufosinate ammonium, bromoxynil, imidazolinones, and 2,4-dichlorophenoxyacetate (2,4-D). See for example, Acetolactase synthase (ALS) for resistance to sulfonylureas, imidazolinones, triazolopyrimidine sulfonamides, pyrimidinylsalicylates and sulphonylaminocarbonyl-triazolinones (Shaner and Singh, 1997, Herbicide Activity:Toxicol Biochem Mol Biol 69-110); glyphosate resistant 5- enolpyruvylshikimate-3-phosphate (EPSPS) (Saroha et al. 1998, J. Plant Biochemistry & Biotechnology Vol 7:65-72);

[0257] A polypeptide of interest includes any protein or polypeptide that is encoded by a polynucleotide of interest described herein.Optimization of Sequences for Expression in Eukaryotes, Mammals, and Plants

[0258] Methods are available in the art for synthesizing coding sequences which are optimized for expression in eukaryotic cells. Eukaryotic cells may be those of or derived from a particular organism, such as a mammal, including but not limited to human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to a process of modifying a nucleic acid sequence for enhanced expression in the host cells of interest by replacing at least one codon (e.g. about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with codons that are more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Various species exhibit particular bias for certain codons of a particular amino acid. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is in turn believed to be dependent on, among other things, the properties of the codons being translated and the availability of particular transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell is generally a reflection of the codons used most frequently in peptide synthesis. Accordingly, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the “Codon Usage Database”, and these tables can be adapted in a number of ways. See Nakamura, Y., et al. “Codon usage tabulated from the international DNA sequence databases: status for the year 2000” Nucl. Acids Res. 28:292 (2000). Yeast codon usage is described in the publicly available Yeast Genome database as well as in Bennetzen and Hall, J Biol Chem. 1982. Computer algorithms for codon optimizing a particular sequence for expression in a particular host cell are also available, such as Gene Forge (Aptagen; Jacobus, Pa.), are also available. Cas9 human codon optimized sequence is disclosed in PCT Application Publication WO 2014 / 093622.

[0259] Methods are available in the art for synthesizing plant-preferred genes. See, for example, U.S. Patent Nos. 5,380,831, and 5,436,391, and Murray et al. (1989) Nucleic Acids Res.17:477-498. Additional sequence modifications are known to enhance gene expression in a plant host. These include, for example, elimination of: one or more sequences encoding spurious polyadenylation signals, one or more exon-intron splice site signals, one or more transposon-like repeats, and other such well-characterized sequences that may be deleterious to gene expression. The G-C content of the sequence may be adjusted to levels average for a given plant host, as calculated by reference to known genes expressed in the host plant cell. When possible, the sequence is modified to avoid one or more predicted hairpin secondarymRNA structures. Thus, "a plant-optimized nucleotide sequence" of the present disclosure comprises one or more of such sequence modifications. Expression Elements

[0260] Any polynucleotide encoding a Cas polypeptide or other CRISPR system component disclosed herein may be functionally linked to a heterologous expression element, to facilitate transcription or regulation in a host cell. Such expression elements include but are not limited to:promoter, leader, intron, and terminator. Expression elements may be “minimal” – meaning a shorter sequence derived from a native source, that still functions as an expression regulator or modifier. Alternatively, an expression element may be “optimized” – meaning that its polynucleotide sequence has been altered from its native state in order to function with a more desirable characteristic in a particular host cell (for example, but not limited to, a bacterial promoter may be “maize-optimized” to improve its expression in corn plants). Alternatively, an expression element may be “synthetic” – meaning that it is designed in silico and synthesized for use in a host cell. Synthetic expression elements may be entirely synthetic, or partially synthetic (comprising a fragment of a naturally-occurring polynucleotide sequence).

[0261] It has been shown that certain promoters are able to direct RNA synthesis at a higher rate than others. These are called “strong promoters”. Certain other promoters have been shown to direct RNA synthesis at higher levels only in particular types of cells or tissues and are often referred to as “tissue specific promoters”, or “tissue-preferred promoters” if the promoters direct RNA synthesis preferably in certain tissues but also in other tissues at reduced levels.

[0262] A plant promoter includes a promoter capable of initiating transcription in a plant cell. For a review of plant promoters, see, Potenza et al., 2004, In vitro Cell Dev Biol 40:1-22; Porto et al., 2014, Molecular Biotechnology (2014), 56(1), 38-49.

[0263] Constitutive promoters are active in cells under most or all circumstances and they can include, for example, SV40, CMV, UBC, EF1A, PGK and CAGG for mammalian systems (Qin et al., 2010 PLosOne 5(5):e10611); and core CaMV 35S promoter; rice actin ubiquitin and ALS promoters in plant systems. (U.S. Patent No.5,659,026) and the like.

[0264] Tissue-preferred or regulated promoters can be utilized to target enhanced expression within a particular tissue or type of cell. Tissue-preferred promoters include, for example, PDGF, PDGFRb, CX3CR1, TRPA1, Krt5, actb, aMHC, 1a1kin in mammalian sytems and root-preferred or seed-preferred promoters in plants.

[0265] Chemical inducible (regulated) promoters can be used to modulate the expression of a gene in a prokaryotic and eukaryotic cell or organism through the application of an exogenouschemical regulator. The promoter may be a chemical-inducible promoter, where application of the chemical induces gene expression, or a chemical-repressible promoter, where application of the chemical represses gene expression. Chemical-inducible promoters include steroid- responsive promoters, tetracycline-inducible and tetracycline-repressible promoters. Gene Targeting

[0266] The guide polynucleotide / Cas systems described herein can be used for gene targeting.

[0267] In general, DNA targeting can be performed by cleaving one or both strands at a specific polynucleotide sequence in a cell with a Cas polypeptide associated with a suitable polynucleotide component. Once a single or double-strand break is induced in the DNA, the cell’s DNA repair mechanism is activated to repair the break via nonhomologous end-joining (NHEJ) or Homology-Directed Repair (HDR) processes which can lead to modifications at the target site.

[0268] The length of the DNA sequence at the target site can vary, and includes, for example, target sites that are at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more than 30 nucleotides in length. It is further possible that the target site can be palindromic, that is, the sequence on one strand reads the same in the opposite direction on the complementary strand. The nick / cleavage site can be within the target sequence or the nick / cleavage site could be outside of the target sequence. In another variation, the cleavage could occur at nucleotide positions immediately opposite each other to produce a blunt end cut or, in other cases, the incisions could be staggered to produce single-stranded overhangs, also called “sticky ends”, which can be either 5' overhangs, or 3' overhangs. Active variants of genomic target sites can also be used. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the given target site, wherein the active variants retain biological activity and hence are capable of being recognized and cleaved by an Cas endonuclease.

[0269] Assays to measure the single or double-strand break of a target site by an endonuclease are known in the art and generally measure the overall activity and specificity of the agent on DNA substrates comprising recognition sites.

[0270] A targeting method herein can be performed in such a way that two or more DNA target sites are targeted in the method, for example. Such a method can optionally be characterized as a multiplex method. Two, three, four, five, six, seven, eight, nine, ten, or more target sites can be targeted at the same time in certain aspects. A multiplex method is typically performed by a targeting method herein in which multiple different RNA components are provided, eachdesigned to guide a guide polynucleotide / Cas endonuclease complex to a unique DNA target site. Gene Editing

[0271] The process for editing a genomic sequence combining DSB and modification templates generally comprises: introducing into a host cell a DSB-inducing agent, or a nucleic acid encoding a DSB-inducing agent, that recognizes a target sequence in the chromosomal sequence and is able to induce a DSB in the genomic sequence, and at least one polynucleotide modification template comprising at least one nucleotide alteration when compared to the nucleotide sequence to be edited. The polynucleotide modification template can further comprise nucleotide sequences flanking the at least one nucleotide alteration, in which the flanking sequences are substantially homologous to the chromosomal region flanking the DSB. Genome editing using DSB-inducing agents, such as Cas-gRNA complexes, has been described, for example in US20150082478 published on 19 March 2015, WO2015026886 published on 26 February 2015, WO2016007347 published 14 January 2016, and WO / 2016 / 025131 published on 18 February 2016.

[0272] Some uses for guide RNA / Cas endonuclease systems have been described (see for example:US20150082478 A1 published 19 March 2015, WO2015026886 published 26 February 2015, and US20150059010 published 26 February 2015) and include but are not limited to modifying or replacing nucleotide sequences of interest (such as a regulatory elements), insertion of polynucleotides of interest, gene knock-out, gene-knock in, modification of splicing sites and / or introducing alternate splicing sites, modifications of nucleotide sequences encoding a protein of interest, amino acid and / or protein fusions, and gene silencing by expressing an inverted repeat into a gene of interest.

[0273] Proteins may be altered in various ways including amino acid substitutions, deletions, truncations, and insertions. Methods for such manipulations are generally known. For example, amino acid sequence variants of the protein(s) can be prepared by mutations in the DNA. Methods for mutagenesis and nucleotide sequence alterations include, for example, Kunkel, (1985) Proc. Natl. Acad. Sci. USA 82:488-92; Kunkel et al., (1987) Meth Enzymol 154:367- 82; U.S. Patent No. 4,873,192; Walker and Gaastra, eds. (1983) Techniques in Molecular Biology (MacMillan Publishing Company, New York) and the references cited therein. Guidance regarding amino acid substitutions not likely to affect biological activity of the protein is found, for example, in the model of Dayhoff et al., (1978) Atlas of Protein Sequence and Structure (Natl Biomed Res Found, Washington, D.C.). Conservative substitutions, suchas exchanging one amino acid with another having similar properties, may be preferable. Conservative deletions, insertions, and amino acid substitutions are not expected to produce radical changes in the characteristics of the protein, and the effect of any substitution, deletion, insertion, or combination thereof can be evaluated by routine screening assays. Assays for double-strand-break-inducing activity are known and generally measure the overall activity and specificity of the agent on DNA substrates comprising target sites.

[0274] Described herein are methods for genome editing with a Cas endonuclease and complexes with a Cas endonuclease and a guide polynucleotide. Following characterization of the guide RNA and PAM sequence, components of the endonuclease and associated CRISPR RNA (crRNA) may be utilized to modify chromosomal DNA in other organisms including plants. To facilitate optimal expression and nuclear localization (for eukaryotic cells), the genes comprising the complex may be optimized as described in WO2016186953 published 24 November 2016, and then delivered into cells as DNA expression cassettes by methods known in the art. The components necessary to comprise an active complex may also be delivered as RNA with or without modifications that protect the RNA from degradation or as mRNA capped or uncapped (Zhang, Y. et al., 2016, Nat. Commun. 7:12617) or Cas polypeptide guide polynucleotide complexes (WO2017070032 published 27 April 2017), or any combination thereof. Additionally, a part or part(s) of the complex and crRNA may be expressed from a DNA construct while other components are delivered as RNA with or without modifications that protect the RNA from degradation or as mRNA capped or uncapped (Zhang et al. 2016 Nat. Commun.7:12617) or Cas polypeptide guide polynucleotide complexes (WO2017070032 published 27 April 2017) or any combination thereof. To produce crRNAs in-vivo, tRNA derived elements may also be used to recruit endogenous RNAses to cleave crRNA transcripts into mature forms capable of guiding the complex to its DNA target site, as described, for example, in WO2017105991 published 22 June 2017. Furthermore, the cleavage activity of the Cas endonuclease may be deactivated by altering key catalytic residues in its cleavage domain (Sinkunas, T. et al., 2013, EMBO J.32:385-394) resulting in a RNA guided helicase that may be used to enhance homology directed repair, induce transcriptional activation, or remodel local DNA structures. Moreover, the activity of the Cas cleavage and helicase domains may both be knocked-out and used in combination with other DNA cutting, DNA nicking, DNA binding, transcriptional activation, transcriptional repression, DNA remodeling, DNA deamination, DNA unwinding, DNA recombination enhancing, DNA integration, DNA inversion, and DNA repair agents.

[0275] The transcriptional direction of the tracrRNA for the CRISPR-Cas system (if present) and other components of the CRISPR-Cas system (such as variable targeting domain, crRNA repeat, loop, anti-repeat) can be deduced as described in WO2016186946 published 24 November 2016, and WO2016186953 published 24 November 2016.

[0276] As described herein, once the appropriate guide RNA requirement is established, the PAM preferences for each new system disclosed herein may be examined. Two regions of PAM randomization separated by two protospacer targets may be utilized to generate a double- stranded DNA break which may be captured and sequenced to examine the PAM sequences that support cleavage by the respective complex.

[0277] In some aspects, a method for modifying a target site in the genome of a cell comprises introducing into a cell at least one PGEN described herein, and identifying at least one cell that has a modification at said target, wherein the modification at said target site is selected from the group consisting of (i) a replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, the chemical alteration of at least one nucleotide, and (v) any combination of (i) – (iv).

[0278] The nucleotide to be edited can be located within or outside a target site recognized and cleaved by a Cas endonuclease. In one aspect, the at least one nucleotide modification is not a modification at a target site recognized and cleaved by a Cas endonuclease. In another aspect, there are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 900 or 1000 nucleotides between the at least one nucleotide to be edited and the genomic target site.

[0279] A knock-out may be produced by an indel (insertion or deletion of nucleotide bases in a target DNA sequence through NHEJ), or by specific removal of sequence that reduces or completely destroys the function of sequence at or near the targeting site.

[0280] A guide polynucleotide / Cas endonuclease induced targeted mutation can occur in a nucleotide sequence that is located within or outside a genomic target site that is recognized and cleaved by the Cas endonuclease.

[0281] In one aspect, the disclosure describes a method for modifying a target site in the genome of a cell, the method comprising introducing into a cell at least one PGEN described herein and at least one donor DNA, wherein said donor DNA comprises a polynucleotide of interest, and optionally, further comprising identifying at least one cell that said polynucleotide of interest integrated in or near said target site.

[0282] In some aspects, the methods disclosed herein may employ homologous recombination (HR) to provide integration of the polynucleotide of interest at the target site.

[0283] Various methods and compositions can be employed to produce a cell or organism having a polynucleotide of interest inserted in a target site via activity of a CRISPR-Cas system component described herein. In one method described herein, a polynucleotide of interest is introduced into the organism cell via a donor DNA construct. As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of a Cas endonuclease. The donor DNA construct further comprises a first and a second region of homology that flank the polynucleotide of interest. The first and second regions of homology of the donor DNA share homology to a first and a second genomic region, respectively, present in or flanking the target site of the cell or organism genome.

[0284] The donor DNA can be tethered to the guide polynucleotide. Tethered donor DNAs can allow for co-localizing target and donor DNA, useful in genome editing, gene insertion, and targeted genome regulation, and can also be useful in targeting post-mitotic cells where function of endogenous HR machinery is expected to be highly diminished (Mali et al., 2013, Nature Methods Vol.10:957-963).

[0285] The amount of homology or sequence identity shared by a target and a donor polynucleotide can vary and includes total lengths and / or regions having unit integral values in the ranges of about 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100-250 bp, 150-300 bp, 200- 400 bp, 250-500 bp, 300-600 bp, 350-750 bp, 400-800 bp, 450-900 bp, 500-1000 bp, 600-1250 bp, 700-1500 bp, 800-1750 bp, 900-2000 bp, 1-2.5 kb, 1.5–3 kb, 2-4 kb, 2.5-5 kb, 3-6 kb, 3.5- 7 kb, 4-8 kb, 5-10 kb, or up to and including the total length of the target site. These ranges include every integer within the range, for example, the range of 1-20 bp includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 and 20 bps. The amount of homology can also be described by percent sequence identity over the full aligned length of the two polynucleotides which includes percent sequence identity of about at least 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. Sufficient homology includes any combination of polynucleotide length, global percent sequence identity, and optionally conserved regions of contiguous nucleotides or local percent sequence identity, for example sufficient homology can be described as a region of 75-150 bp having at least 80% sequence identity to a region of the target locus. Sufficient homology can also be described by the predicted ability of two polynucleotides to specifically hybridize under high stringency conditions, see, for example, Sambrook et al., (1989) Molecular Cloning:A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY); Current Protocols in Molecular Biology, Ausubel et al., Eds (1994) Current Protocols, (Greene PublishingAssociates, Inc. and John Wiley & Sons, Inc.); and, Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology--Hybridization with Nucleic Acid Probes, (Elsevier, New York).

[0286] Episomal DNA molecules can also be ligated into the double-strand break, for example, integration of T-DNAs into chromosomal double-strand breaks (Chilton and Que, (2003) Plant Physiol 133:956-65; Salomon and Puchta, (1998) EMBO J. 17:6086-95). Once the sequence around the double-strand breaks is altered, for example, by exonuclease activities involved in the maturation of double-strand breaks, gene conversion pathways can restore the original structure if a homologous sequence is available, such as a homologous chromosome in non- dividing somatic cells, or a sister chromatid after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenic DNA sequences may also serve as a DNA repair template for homologous recombination (Puchta, (1999) Genetics 152:1173-81).

[0287] In one aspect, the disclosure comprises a method for editing a nucleotide sequence in the genome of a cell, the method comprising introducing into at least one PGEN described herein, and a polynucleotide modification template, wherein said polynucleotide modification template comprises at least one nucleotide modification of said nucleotide sequence, and optionally further comprising selecting at least one cell that comprises the edited nucleotide sequence.

[0288] The guide polynucleotide / Cas endonuclease system can be used in combination with at least one polynucleotide modification template to allow for editing (modification) of a genomic nucleotide sequence of interest. (See also US20150082478, published 19 March 2015 and WO2015026886 published 26 February 2015).

[0289] Polynucleotides of interest and / or traits can be stacked together in a complex trait locus as described in PCT Application Publications WO2012 / 129373 and WO 2013 / 112686or in a modified locus as described in PCT Application Publication WO 2022 / 040134. The guide polynucleotide / Cas-beta endonuclease system described herein provides for an efficient system to generate double-strand breaks and allows for traits to be stacked in a complex trait locus.

[0290] Further uses for guide RNA / Cas endonuclease systems have been described (See for example:US20150082478 published 19 March 2015, WO2015026886 published 26 February 2015, US20150059010 published 26 February 2015, WO2016007347 published 14 January 2016, and PCT application WO2016025131 published 18 February 2016) and include but are not limited to modifying or replacing nucleotide sequences of interest (such as a regulatory elements), insertion of polynucleotides of interest, gene knock-out, gene-knock in, modification of splicing sites and / or introducing alternate splicing sites, modifications ofnucleotide sequences encoding a protein of interest, amino acid and / or protein fusions, and gene silencing by expressing an inverted repeat into a gene of interest.

[0291] Resulting characteristics from the gene editing compositions and methods described herein may be evaluated. Chromosomal intervals that correlate with a phenotype or trait of interest can be identified. A variety of methods well known in the art are available for identifying chromosomal intervals. The boundaries of such chromosomal intervals are drawn to encompass markers that will be linked to the gene controlling the trait of interest. In other words, the chromosomal interval is drawn such that any marker that lies within that interval (including the terminal markers that define the boundaries of the interval) can be used as a marker for a particular trait. In one aspect, the chromosomal interval comprises at least one QTL, and furthermore, may indeed comprise more than one QTL. Close proximity of multiple QTLs in the same interval may obfuscate the correlation of a particular marker with a particular QTL, as one marker may demonstrate linkage to more than one QTL. Conversely, e.g., if two markers in close proximity show co-segregation with the desired phenotypic trait, it is sometimes unclear if each of those markers identifies the same QTL or two different QTL. The term “quantitative trait locus” or “QTL” refers to a region of DNA that is associated with the differential expression of a quantitative phenotypic trait in at least one genetic background, e.g., in at least one breeding population. The region of the QTL encompasses or is closely linked to the gene or genes that affect the trait in question. An “allele of a QTL” can comprise multiple genes or other genetic factors within a contiguous genomic region or linkage group, such as a haplotype. An allele of a QTL can denote a haplotype within a specified window wherein said window is a contiguous genomic region that can be defined, and tracked, with a set of one or more polymorphic markers. A haplotype can be defined by the unique fingerprint of alleles at each marker within the specified window. Introduction of CRISPR-Cas System Components into a Cell

[0292] The methods and compositions described herein do not depend on a particular method for introducing a sequence into an organism or cell, only that the polynucleotide or polypeptide gains access to the interior of at least one cell of the organism. Introducing includes reference to the incorporation of a nucleic acid into a eukaryotic or prokaryotic cell where the nucleic acid may be incorporated into the genome of the cell, and includes reference to the transient (direct) provision of a nucleic acid, protein or polynucleotide-protein complex (PGEN, RGEN) to the cell.

[0293] Methods for introducing polynucleotides or polypeptides or a polynucleotide-protein complex into cells or organisms are known in the art including, but not limited to, microinjection, electroporation, stable transformation methods, transient transformation methods, ballistic particle acceleration (particle bombardment), whiskers mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, viral-mediated introduction, transfection, transduction, cell-penetrating peptides, mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, topical applications, sexual crossing , sexual breeding, and any combination thereof.

[0294] For example, the guide polynucleotide (guide RNA, crNucleotide + tracrNucleotide, guide DNA and / or guide RNA-DNA molecule) can be introduced into a cell directly (transiently) as a single stranded or double stranded polynucleotide molecule. The guide RNA (or crRNA + tracrRNA) can also be introduced into a cell indirectly by introducing a recombinant DNA molecule comprising a heterologous nucleic acid fragment encoding the guide RNA (or crRNA + tracrRNA), operably linked to a specific promoter that is capable of transcribing the guide RNA (crRNA+tracrRNA molecules) in said cell. The specific promoter can be, but is not limited to, a RNA polymerase III promoter, which allow for transcription of RNA with precisely defined, unmodified, 5’- and 3’-ends (Ma et al., 2014, Mol. Ther. Nucleic Acids 3:e161; DiCarlo et al., 2013, Nucleic Acids Res. 41:4336-4343; WO2015026887, published 26 February 2015). Any promoter capable of transcribing the guide RNA in a cell can be used and includes a heat shock / heat inducible promoter operably linked to a nucleotide sequence encoding the guide RNA.

[0295] Plant cells differ from animal cells (such as human cells), fungal cells (such as yeast cells) and protoplasts, including for example plant cells comprise a plant cell wall which may act as a barrier to the delivery of components.

[0296] Delivery of the Cas polypeptide, and / or the guide RNA, and / or a ribonucleoprotein complex, and / or a polynucleotide encoding any one or more of the preceding, into mammalian cells can be achieved through methods known in the art, for example but not limited to the following. In particular the relatively smaller size of Cas-beta variants described herein facilitates their in vivo or ex vivo delivery into mammalian cells by the following methods. Viral vectors such as retroviral (e.g. murine leukemia, HIV, or lentiviral) or DNA viruses (e.g. adenovirus, herpes simplex, and adeno-associated) or non-viral vectors can be used. Transfection methods are non-viral delivery methods that can be used and can include contacting the cell with DEAE-Dextran, calcium phosphate, liposomes or electroporation of a plasmid into a cell. Additional methods of non-viral delivery include electroporation,lipofection, microinjection, biolistics, virosomes, liposomes, immuno liposomes, polycation or lipid: nucleic acid conjugates, naked DNA, naked RNA, artificial virions, and agent-enhanced uptake of DNA. Sonoporation using, e.g., the Sonitron 2000 system (Rich-Mar) can also be used for delivery of nucleic acids. In some embodiments, one or more nucleic acids are delivered as mRNA. In some embodiments, capped mRNAs are used to increase translational efficiency and / or mRNA stability. In some embodiments, ARCA (anti-reverse cap analog) caps or variants thereof are used. See U.S. Pat. Nos.7,074,596 and 8,153,773.

[0297] Delivery of the Cas polypeptide, and / or the guide RNA, and / or a ribonucleoprotein complex, and / or a polynucleotide encoding any one or more of the preceding, into plant cells can be achieved through methods known in the art, for example but not limited to: Rhizobiales- mediated transformation (e.g., Agrobacterium, Ochrobactrum), particle mediated delivery (particle bombardment), polyethylene glycol (PEG)-mediated transfection (for example to protoplasts), electroporation, cell-penetrating peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery.

[0298] Cas polypeptides, such as the Cas polypeptide described herein, can be introduced into a cell by directly introducing the Cas polypeptide itself (referred to as direct delivery of Cas endonuclease), the mRNA encoding the Cas polypeptide, and / or the guide polynucleotide / Cas endonuclease complex itself, using any method known in the art. The Cas endonuclease can also be introduced into a cell indirectly by introducing a recombinant DNA molecule that encodes the Cas endonuclease. The endonuclease can be introduced into a cell transiently or can be incorporated into the genome of the host cell using any method known in the art. Uptake of the endonuclease and / or the guided polynucleotide into the cell can be facilitated with a Cell Penetrating Peptide (CPP) as described in WO2016073433 published 12 May 2016. Any promoter capable of expressing the Cas endonuclease in a cell can be used and includes a heat shock / heat inducible promoter operably linked to a nucleotide sequence encoding the Cas endonuclease.

[0299] Direct delivery of a polynucleotide modification template into plant cells can be achieved through particle mediated delivery, and any other direct method of delivery, such as but not limiting to, polyethylene glycol (PEG)-mediated transfection to protoplasts, whiskers mediated transformation, electroporation, particle bombardment, cell-penetrating peptides, or mesoporous silica nanoparticle (MSN)-mediated direct protein delivery can be successfully used for delivering a polynucleotide modification template in eukaryotic cells, such as plant cells.

[0300] The donor DNA can be introduced by any means known in the art. The donor DNA may be provided by any transformation method known in the art including, for example, Agrobacterium-mediated transformation or biolistic particle bombardment. The donor DNA may be present transiently in the cell or it could be introduced via a viral replicon. In the presence of the Cas endonuclease and the target site, the donor DNA is inserted into the transformed plant’s genome.

[0301] Direct delivery of any one of the guided Cas system components can be accompanied by direct delivery (co-delivery) of other mRNAs that can promote the enrichment and / or visualization of cells receiving the guide polynucleotide / Cas polypeptide complex components. For example, direct co-delivery of the guide polynucleotide / Cas polypeptide components (and / or guide polynucleotide / Cas endonuclease complex itself) together with mRNA encoding phenotypic markers (such as but not limiting to transcriptional activators such as CRC (Bruce et al. 2000 The Plant Cell 12:65-79) can enable the selection and enrichment of cells without the use of an exogenous selectable marker by restoring function to a non-functional gene product as described in WO2017070032 published 27 April 2017.

[0302] Introducing a guide RNA / Cas polypeptide complex described herein, (representing the cleavage ready complex described herein) into a cell includes introducing the individual components of said complex either separately or combined into the cell, and either directly (direct delivery as RNA for the guide and protein for the Cas polypeptide and protein subunits, or functional fragments thereof) or via recombination constructs expressing the components (guide RNA, Cas endonuclease, protein subunits, or functional fragments thereof). Introducing a guide RNA / Cas endonuclease complex (RGEN) into a cell includes introducing the guide RNA / Cas endonuclease complex as a ribonucleotide-protein into the cell. The ribonucleotide- protein can be assembled prior to being introduced into the cell as described herein. The components comprising the guide RNA / Cas endonuclease ribonucleotide protein (at least one Cas endonuclease, at least one guide RNA, at least one protein subunit) can be assembled in vitro or assembled by any means known in the art prior to being introduced into a cell (targeted for genome modification as described herein).

[0303] Direct delivery of the RGEN ribonucleoprotein, allows for genome editing at a target site in the genome of a cell which can be followed by rapid degradation of the complex, and only a transient presence of the complex in the cell. This transient presence of the RGEN complex may lead to reduced off-target effects. In contrast, delivery of RGEN components (guide RNA, Cas9 endonuclease) via plasmid DNA sequences can result in constant expressionof RGENs from these plasmids which can intensify off target effects (Cradick, T. J. et al. (2013) Nucleic Acids Res 41:9584-9592; Fu, Y et al. (2014) Nat. Biotechnol.31:822-826).

[0304] Direct delivery can be achieved by combining any one component of the guide RNA / Cas endonuclease complex (RGEN), representing the cleavage ready complex described herein, (such as at least one guide RNA, at least one Cas polypeptide, and optionally one additional protein), with a delivery matrix comprising a microparticle (such as but not limited to of a gold particle, tungsten particle, and silicon carbide whisker particle) (see also WO2017070032 published 27 April 2017). The delivery matrix may comprise any one of the components, such as the Cas endonuclease, which is attached to a solid matrix (e.g., a particle for bombardment).

[0305] In some aspects, the guide polynucleotide / Cas endonuclease complex, is a complex wherein the guide RNA and Cas endonuclease forming the guide RNA / Cas endonuclease complex are introduced into the cell as RNA and protein, respectively.

[0306] In some aspects, the guide polynucleotide / Cas endonuclease complex, is a complex wherein the guide RNA and Cas endonuclease protein and the at least one protein subunit of a complex forming the guide RNA / Cas endonuclease complex are introduced into the cell as RNA and proteins, respectively.

[0307] In some aspects, the guide polynucleotide / Cas endonuclease complex, is a complex wherein the guide RNA and Cas endonuclease protein and the at least one protein subunit of a complex forming the guide RNA / Cas endonuclease complex (cleavage ready complex) are preassembled in vitro and introduced into the cell as a ribonucleotide-protein complex.

[0308] Protocols for introducing polynucleotides, polypeptides or polynucleotide-protein complexes (PGEN, RGEN) into eukaryotic cells, such as plants or plant cells are known and include microinjection (Crossway et al., (1986) Biotechniques 4:320-34 and U.S. Patent No. 6,300,543), meristem transformation (U.S. Patent No.5,736,369), electroporation (Riggs et al., (1986) Proc. Natl. Acad. Sci. USA 83:5602-6, Agrobacterium-mediated transformation (U.S. Patent Nos. 5,563,055 and 5,981,840), whiskers mediated transformation (Ainley et al. 2013, Plant Biotechnology Journal 11:1126-1134; Shaheen A. and M. Arshad 2011 Properties and Applications of Silicon Carbide (2011), 345-358 Editor(s):Gerhardt, Rosario. Publisher:InTech, Rijeka, Croatia. CODEN:69PQBP; ISBN:978-953-307-201-2), direct gene transfer (Paszkowski et al., (1984) EMBO J 3:2717-22), and ballistic particle acceleration (U.S. Patent Nos. 4,945,050; 5,879,918; 5,886,244; 5,932,782; Tomes et al., (1995) "Direct DNA Transfer into Intact Plant Cells via Microprojectile Bombardment" in Plant Cell, Tissue, and Organ Culture:Fundamental Methods, ed. Gamborg & Phillips (Springer-Verlag, Berlin);McCabe et al., (1988) Biotechnology 6:923-6; Weissinger et al., (1988) Ann Rev Genet 22:421- 77; Sanford et al., (1987) Particulate Science and Technology 5:27-37 (onion); Christou et al., (1988) Plant Physiol 87:671-4 (soybean); Finer and McMullen, (1991) In vitro Cell Dev Biol 27P:175-82 (soybean); Singh et al., (1998) Theor Appl Genet 96:319-24 (soybean); Datta et al., (1990) Biotechnology 8:736-40 (rice); Klein et al., (1988) Proc. Natl. Acad. Sci. USA 85:4305-9 (maize); Klein et al., (1988) Biotechnology 6:559-63 (maize); U.S. Patent Nos. 5,240,855; 5,322,783 and 5,324,646; Klein et al., (1988) Plant Physiol 91:440-4 (maize); Fromm et al., (1990) Biotechnology 8:833-9 (maize); Hooykaas-Van Slogteren et al., (1984) Nature 311:763-4; U.S. Patent No. 5,736,369 (cereals); Bytebier et al., (1987) Proc. Natl. Acad. Sci. USA 84:5345-9 (Liliaceae); De Wet et al., (1985) in The Experimental Manipulation of Ovule Tissues, ed. Chapman et al., (Longman, New York), pp. 197-209 (pollen); Kaeppler et al., (1990) Plant Cell Rep 9:415-8) and Kaeppler et al., (1992) Theor Appl Genet 84:560-6 (whisker-mediated transformation); D'Halluin et al., (1992) Plant Cell 4:1495-505 (electroporation); Li et al., (1993) Plant Cell Rep 12:250-5; Christou and Ford (1995) Annals Botany 75:407-13 (rice) and Osjoda et al., (1996) Nat Biotechnol 14:745-50 (maize via Agrobacterium tumefaciens).

[0309] Alternatively, polynucleotides may be introduced into plant or plant cells by contacting cells or organisms with a virus or viral nucleic acids. Generally, such methods involve incorporating a polynucleotide within a viral DNA or RNA molecule. In some examples a polypeptide of interest may be initially synthesized as part of a viral polyprotein, which is later processed by proteolysis in vivo or in vitro to produce the desired recombinant protein. Methods for introducing polynucleotides into plants and expressing a protein encoded therein, involving viral DNA or RNA molecules, are known, see, for example, U.S. Patent Nos. 5,889,191, 5,889,190, 5,866,785, 5,589,367 and 5,316,931.

[0310] The polynucleotide or recombinant DNA construct can be provided to or introduced into a prokaryotic and eukaryotic cell or organism using a variety of transient transformation methods. Such transient transformation methods include, but are not limited to, the introduction of the polynucleotide construct directly into the plant.

[0311] Nucleic acids and proteins can be provided to a cell by any method including methods using molecules to facilitate the uptake of anyone or all components of a guided Cas system (protein and / or nucleic acids), such as cell-penetrating peptides and nanocarriers. See also US20110035836 published 10 February 2011, and EP2821486A1 published 07 January 2015.

[0312] Other methods of introducing polynucleotides into a prokaryotic and eukaryotic cell or organism or plant part can be used, including plastid transformation methods, and the methods for introducing polynucleotides into tissues from seedlings or mature seeds.

[0313] Stable transformation is intended to mean that the nucleotide construct introduced into an organism integrates into a genome of the organism and is capable of being inherited by the progeny thereof. Transient transformation is intended to mean that a polynucleotide is introduced into the organism and does not integrate into a genome of the organism or a polypeptide is introduced into an organism. Transient transformation indicates that the introduced composition is only temporarily expressed or present in the organism.

[0314] A variety of methods are available to identify those cells having an altered genome at or near a target site without using a screenable marker phenotype. Such methods can be viewed as directly analyzing a target sequence to detect any change in the target sequence, including but not limited to PCR methods, sequencing methods, nuclease digestion, Southern blots, and any combination thereof. Cells and Animals

[0315] The presently disclosed polynucleotides and polypeptides can be introduced into an animal cell. Animal cells can include, but are not limited to: an organism of a phylum including chordates, arthropods, mollusks, annelids, cnidarians, or echinoderms; or an organism of a class including mammals, insects, birds, amphibians, reptiles, or fishes. In some aspects, the animal is human, mouse, C. elegans, rat, fruit fly (Drosophila spp.), zebrafish, chicken, dog, cat, guinea pig, hamster, chicken, Japanese ricefish, sea lamprey, pufferfish, tree frog (e.g., Xenopus spp.), monkey, or chimpanzee. Particular cell types that are contemplated include haploid cells, diploid cells, reproductive cells, neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, embryonic cells, hematopoietic cells, bone cells, germ cells, somatic cells, stem cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some aspects, a plurality of cells from an organism may be used.

[0316] The Cas endonucleases disclosed herein may be used to edit the genome of an animal cell in various ways. In some aspects, it may be desirable to delete one or more nucleotides. In another aspect, it may be desirable to insert one or more nucleotides. In some aspects, it may be desirable to replace one or more nucleotides. In another aspect, it may be desirable to modify one or more nucleotides via a covalent or non-covalent interaction with another atom or molecule.

[0317] Genome modification via a Cas endonuclease may be used to effect a genotypic and / or phenotypic change on the target organism. Such a change is preferably related to an improved phenotype of interest or a physiologically-important characteristic, the correction of an endogenous defect, or the expression of some type of expression marker. In some aspects, the phenotype of interest or physiologically-important characteristic is related to the overall health, fitness, or fertility of the animal, the ecological fitness of the animal, or the relationship or interaction of the animal with other organisms in its environment. In some aspects, the phenotype of interest or physiologically-important characteristic is selected from the group consisting of: improved general health, disease reversal, disease modification, disease stabilization, disease prevention, treatment of parasitic infections, treatment of viral infections, treatment of retroviral infections, treatment of bacterial infections, treatment of neurological disorders (for example but not limited to: multiple sclerosis), correction of endogenous genetic defects (for example but not limited to: metabolic disorders, Achondroplasia, Alpha-1 Antitrypsin Deficiency, Antiphospholipid Syndrome, Autism, Autosomal Dominant Polycystic Kidney Disease, Barth syndrome, Breast cancer, Charcot-Marie-Tooth, Colon cancer, Cri du chat, Crohn's Disease, Cystic fibrosis, Dercum Disease, Down Syndrome, Duane Syndrome, Duchenne Muscular Dystrophy, Factor V Leiden Thrombophilia, Familial Hypercholesterolemia, Familial Mediterranean Fever, Fragile X Syndrome, Gaucher Disease, Hemochromatosis, Hemophilia, Holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, Myotonic Dystrophy, Neurofibromatosis, Noonan Syndrome, Osteogenesis Imperfecta, Parkinson's disease, Phenylketonuria, Poland Anomaly, Porphyria, Progeria, Prostate Cancer, Retinitis Pigmentosa, Severe Combined Immunodeficiency (SCID), Sickle cell disease, Skin Cancer, Spinal Muscular Atrophy, Tay-Sachs, Thalassemia, Trimethylaminuria, Turner Syndrome, Velocardiofacial Syndrome, WAGR Syndrome, and Wilson Disease), treatment of innate immune disorders (for example but not limited to: immunoglobulin subclass deficiencies), treatment of acquired immune disorders (for example but not limited to: AIDS and other HIV-related disorders), treatment of cancer, as well as treatment of diseases, including rare or “orphan” conditions, that have eluded effective treatment options with other methods.

[0318] Cells that have been genetically modified using the compositions or methods disclosed herein may be transplanted to a subject for purposes such as gene therapy, e.g. to treat a disease, or as an antiviral, antipathogenic, or anticancer therapeutic, for the production of genetically modified organisms in agriculture, or for biological research.Cells and Plants

[0319] The presently disclosed polynucleotides and polypeptides can be introduced into a cell. Cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells as well as plants and seeds produced by the methods described herein. Any plant can be used with the compositions and methods described herein, including monocot and dicot plants, and plant elements.

[0320] Examples of monocot plants that can be used include, but are not limited to, corn (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet (Pennisetum glaucum), proso millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana)), wheat (Triticum species, for example Triticum aestivum, Triticum monococcum), sugarcane (Saccharum spp.), oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), pineapple (Ananas comosus), banana (Musa spp.), palm, ornamentals, turfgrasses, and other grasses.

[0321] Examples of dicot plants that can be used include, but are not limited to, soybean (Glycine max), Brassica species (for example but not limited to:oilseed rape or Canola) (Brassica napus, B. campestris, Brassica rapa, Brassica. juncea), alfalfa (Medicago sativa),), tobacco (Nicotiana tabacum), Arabidopsis (Arabidopsis thaliana), sunflower (Helianthus annuus), cotton (Gossypium arboreum, Gossypium barbadense), and peanut (Arachis hypogaea), tomato (Solanum lycopersicum), potato (Solanum tuberosum.

[0322] Additional plants that can be used include safflower (Carthamus tinctorius), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), citrus trees (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musa spp.), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beets (Beta vulgaris), vegetables, ornamentals, and conifers.

[0323] Vegetables that can be used include tomatoes (Lycopersicon esculentum), lettuce (e.g., Lactuca sativa), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus spp.), and members of the genus Cucumis such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and musk melon (C. melo). Ornamentals include azalea (Rhododendron spp.), hydrangea (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnation (Dianthus caryophyllus), poinsettia (Euphorbia pulcherrima), and chrysanthemum.

[0324] Conifers that may be used include pines such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); redwood (Sequoia sempervirens); true firs such as silver fir (Abies amabilis) and balsam fir (Abies balsamea); and cedars such as Western red cedar (Thuja plicata) and Alaska yellow cedar (Chamaecyparis nootkatensis).

[0325] In some aspects of the disclosure, a fertile plant is a plant that produces viable male and female gametes and is self-fertile. Such a self-fertile plant can produce a progeny plant without the contribution from any other plant of a gamete and the genetic material comprised therein. Other aspects of the disclosure can involve the use of a plant that is not self-fertile because the plant does not produce male gametes, or female gametes, or both, that are viable or otherwise capable of fertilization.

[0326] The present disclosure finds use in the breeding of plants comprising one or more introduced traits, or edited genomes.

[0327] A non-limiting example of how two traits can be stacked into the genome at a genetic distance of, for example, 5 cM from each other is described as follows:A first plant comprising a first transgenic target site integrated into a first DSB target site within the genomic window and not having the first genomic locus of interest is crossed to a second transgenic plant, comprising a genomic locus of interest at a different genomic insertion site within the genomic window and the second plant does not comprise the first transgenic target site. About 5% of the plant progeny from this cross will have both the first transgenic target site integrated into a first DSB target site and the first genomic locus of interest integrated at different genomic insertion sites within the genomic window. Progeny plants having both sites in the defined genomic window can be further crossed with a third transgenic plant comprising a second transgenic target site integrated into a second DSB target site and / or a second genomic locus of interest within the defined genomic window and lacking the first transgenic target site and the first genomic locus of interest. Progeny are then selected having the first transgenic target site, the first genomic locus of interest and the second genomic locus of interest integrated at different genomic insertion sites within the genomic window. Such methods can be used to produce a transgenic plant comprising a complex trait locus having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14,15, 16, 17, 19, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 or more transgenic target sites integrated into DSB target sites and / or genomic loci of interest integrated at different sites within the genomic window. In such a manner, various complex trait loci can be generated.In vitro Polynucleotide Detection, Binding, and Modification

[0328] The compositions disclosed herein may further be used as compositions for use in in vitro methods, in some aspects with isolated polynucleotide sequence(s). Said isolated polynucleotide sequence(s) may comprise one or more target sequence(s) for modification. In some aspects, said isolated polynucleotide sequence(s) may be genomic DNA, a PCR product, or a synthesized oligonucleotide. Compositions

[0329] Modification of a target sequence may be in the form of a nucleotide insertion, a nucleotide deletion, a nucleotide substitution, the addition of an atom molecule to an existing nucleotide, a nucleotide modification, or the binding of a heterologous polynucleotide or polypeptide to said target sequence. The insertion of one or more nucleotides may be accomplished by the inclusion of a donor polynucleotide in the reaction mixture: said donor polynucleotide is inserted into a double-strand break created by a Cas-beta variant disclosed herein. The insertion may be via non-homologous end joining or via homologous recombination.

[0330] In some aspects, the sequence of the target polynucleotide is known prior to modification, and compared to the sequence(s) of polynucleotide(s) that result from treatment with the Cas-beta variant. In some aspects, the sequence of the target polynucleotide is not known prior to modification, and the treatment with the Cas-beta variant is used as part of a method to determine the sequence of said target polynucleotide.

[0331] In some aspects, the Cas-beta variant may be selected from the group consisting of: an unmodified wild type Cas-beta ortholog, a functional Cas-beta variant, a functional Cas-beta fragment, a fusion protein comprising an active Cas-beta variant or a de-activated Cas-beta variant that lacks endonuclease activity, a Cas-beta variant further comprising one or more nuclear localization sequences (NLS) on the C-terminus or on the N-terminus or on both the N- and C-termini, a biotinylated Cas-beta variant, a Cas-beta variant endonuclease, a Cas-beta variant further comprising a Histidine tag, and a mixture of any two or more of the preceding.

[0332] In some aspects, the Cas-beta variant disclosed herein is a fusion protein further comprising a nuclease domain, a transcriptional activator domain, a transcriptional repressor domain, an epigenetic modification domain, a cleavage domain, a nuclear localization signal, a cell-penetrating domain, a translocation domain, a marker, or a transgene that is heterologous to the target polynucleotide sequence or to the cell from which said target polynucleotide sequence is obtained or derived.

[0333] In some aspects, a plurality of Cas-beta variants may be desired. In some aspects, said plurality may comprise Cas-beta variants derived from different source organisms or from different loci within the same organism. In some aspects, said plurality may comprise Cas-beta variants with different binding specificities to the target polynucleotide. In some aspects, said plurality may comprise Cas-beta variants with different cleavage efficiencies. In some aspects, said plurality may comprise Cas-beta variants with different PAM specificities.

[0334] The guide polynucleotide may be provided as a single guide RNA (sgRNA), a chimeric molecule comprising a tracrRNA, a chimeric molecule comprising a crRNA, a chimeric RNA- DNA molecule, a DNA molecule, or a polynucleotide comprising one or more chemically modified nucleotides.

[0335] The storage conditions of the Cas-beta variant and / or the guide polynucleotide disclosed herein include parameters for temperature, state of matter, and time. In some aspects, the Cas-beta variant and / or the guide polynucleotide is stored at about -80 degrees Celsius, at about -20 degrees Celsius, at about 4 degrees Celsius, at about 20-25 degrees Celsius, or at about 37 degrees Celsius. In some aspects, the Cas-beta ortholog and / or the guide polynucleotide is stored as a liquid, a frozen liquid, or as a lyophilized powder. In some aspects, the Cas-beta ortholog and / or the guide polynucleotide is stable for at least one day, at least one week, at least one month, at least one year, or even greater than one year.

[0336] Any or all of the possible polynucleotide components of the reaction (e.g., guide polynucleotide, donor polynucleotide, optionally a Cas-beta polynucleotide) may be provided as part of a vector, a construct, a linearized or circularized plasmid, or as part of a chimeric molecule. Each component may be provided to the reaction mixture separately or together. In some aspects, one or more of the polynucleotide components are operably linked to a heterologous noncoding regulatory element that regulates its expression.

[0337] The method for modification of a target polynucleotide comprises combining the minimal elements into a reaction mixture comprising: a Cas-beta variant (fragment, or other related molecule as described above), a guide polynucleotide comprising a sequence that is substantially complementary to, or selectively hybridizes to, the target polynucleotide sequence of the target polynucleotide, and a target polynucleotide for modification. In some aspects, the Cas-beta ortholog is provided as a polypeptide. In some aspects, the Cas-alpha ortholog is provided as a Cas-beta ortholog polynucleotide. In some aspects, the guide polynucleotide is provided as an RNA molecule, a DNA molecule, an RNA:DNA hybrid, or a polynucleotide molecule comprising a chemically-modified nucleotide.

[0338] The storage buffer of any one of the components, or the reaction mixture, may be optimized for stability, efficacy, or other parameters. Additional components of the storage buffer or the reaction mixture may include a buffer composition, Tris, EDTA, dithiothreitol (DTT), phosphate-buffered saline (PBS), sodium chloride, magnesium chloride, HEPES, glycerol, BSA, a salt, an emulsifier, a detergent, a chelating agent, a redox reagent, an antibody, nuclease-free water, a proteinase, and / or a viscosity agent. In some aspects, the storage buffer or reaction mixture further comprises a buffer solution with at least one of the following components: HEPES, MgCl2, NaCl, EDTA, a proteinase, Proteinase K, glycerol, nuclease-free water.

[0339] Incubation conditions will vary according to desired outcome. The temperature is preferably at least 10 degrees Celsius, between 10 and 15, at least 15, between 15 and 17, at least 17, between 17 and 20, at least 20, between 20 and 22, at least 22, between 22 and 25, at least 25, between 25 and 27, at least 27, between 27 and 30, at least 30, between 30 and 32, at least 32, between 32 and 35, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, or even greater than 40 degrees Celsius. The time of incubation is at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, or even greater than 10 minutes.

[0340] The sequence(s) of the polynucleotide(s) in the reaction mixture prior to, during, or after incubation may be determined by any method known in the art. In some aspects, modification of a target polynucleotide may be ascertained by comparing the sequence(s) of the polynucleotide(s) purified from the reaction mixture to the sequence of the target polynucleotide prior to combining with the Cas-beta variant.

[0341] Any one or more of the compositions disclosed herein, useful for in vitro or in vivo polynucleotide detection, binding, and / or modification, may be comprised within a kit. A kit comprises a Cas-beta ortholog or a polynucleotide Cas-beta variant encoding such, optionally further comprising buffer components to enable efficient storage, and one or more additional compositions that enable the introduction of said Cas-beta variant or Cas-beta variant to a heterologous polynucleotide, wherein said Cas-alpha ortholog or Cas-beta variant is capable of effecting a modification, addition, deletion, or substitution of at least one nucleotide of said heterologous polynucleotide. In an additional aspect, a Cas-beta variant disclosed herein may be used for the enrichment of one or more polynucleotide target sequences from a mixed pool. In an additional aspect, a Cas-beta variant disclosed herein may be immobilized on a matrix for use in in vitro target polynucleotide detection, binding, and / or modification.

[0342] A Cas-beta variant may be attached, associated with, or affixed to a solid matrix for the purposes of storage, purification, and / or characterization. Examples of a solid matrix include, but are not limited to: a filter, a chromatography resin, an assay plate, a test tube, a cryogenic vial, etc. A Cas-beta variant may be substantially purified and stored in an appropriate buffer solution, or lyophilized. Methods of Detection

[0343] Methods of detecting the Cas-beta variant-guide polynucleotide complex bound to the target polynucleotide may include any known in the art, including but not limited to microscopy, chromatographic separation, electrophoresis, immunoprecipitation, filtration, nanopore separation, microarrays, as well as those described below.

[0344] A DNA Electrophoretic Mobility Shift Assay (EMSA): studies proteins binding to known DNA oligonucleotide probes and assesses the specificity of the interaction. The technique is based on the principle that protein-DNA complexes migrate more slowly than free DNA molecules when subjected to polyacrylamide or agarose gel electrophoresis. Because the rate of DNA migration is retarded upon protein binding, the assay is also called a gel retardation assay. Adding a protein-specific antibody to the binding components creates an even larger complex (antibody-protein-DNA) which migrates even slower during electrophoresis, this is known as a supershift and can be used to confirm protein identities.

[0345] DNA Pull-down Assays use a DNA probe labelled with a high affinity tag, such as biotin, which allows the probe to be recovered or immobilized. A DNA probe can be complexed with a protein from a cell lysate in a reaction similar to that used in the EMSA and then used to purify the complex using agarose or magnetic beads. The proteins are then eluted from the DNA and detected by Western blot or identified by mass spectrometry. Alternatively, the protein may be labelled with an affinity tag or the DNA-protein complex may be isolated using an antibody against the protein of interest (similar to a supershift assay). In this case, the unknown DNA sequence bound by the protein is detected by Southern blotting or through PCR analysis.

[0346] Reporter assays provide a real-time in vivo read-out of translational activity for a promoter of interest. Reporter genes are fusions of a target promoter DNA sequence and a reporter gene DNA sequence which is customized by the researcher and the DNA sequence codes for a protein with detectable properties like firefly / Renilla luciferase or alkaline phosphatase. These genes produce enzymes only when the promoter of interest is activated. The enzyme, in turn, catalyses a substrate to produce either light or a colour change that can bedetected by spectroscopic instrumentation. The signal from the reporter gene is used as an indirect determinant for the translation of endogenous proteins driven from the same promoter.

[0347] Microplate Capture and Detection Assays use immobilized DNA probes to capture specific protein-DNA interactions and confirm protein identities and relative amounts with target specific antibodies. Typically, a DNA probe is immobilized on the surface of 96- or 384- well microplates coated with streptavidin. A cellular extract is prepared and added to allow the binding protein to bind to the oligonucleotide. The extract is then removed and each well is washed several times to remove non-specifically bound proteins. Finally, the protein is detected using a specific antibody labelled for detection. This method can be extremely sensitive , detecting less than 0.2pg of the target protein per well. This method may also be utilized for oligonucleotides labelled with other tags, such as primary amines that can be immobilized on microplates coated with an amine-reactive surface chemistry.

[0348] DNA Footprinting is one of the most widely used methods for obtaining detailed information on the individual nucleotides in protein–DNA complexes, even inside living cells. In such an experiment, chemicals or enzymes are used to modify or digest the DNA molecules. • When sequence specific proteins bind to DNA, they can protect the binding sites from modification or digestion. This can subsequently be visualized by denaturing gel electrophoresis, where unprotected DNA is cleaved more or less at random. Therefore it appears as a ‘ladder’ of bands and the sites protected by proteins have no corresponding bands and look like foot prints in the pattern of bands. The foot prints there by identify specific nucleosides at the protein–DNA binding sites.

[0349] Microscopic techniques include optical, fluorescence, electron, and atomic force microscopy (AFM).

[0350] Chromatin immunoprecipitation analysis (ChIP) causes proteins to bind covalently to their DNA targets, after which they are unlinked and characterized separately.

[0351] Systematic Evolution of Ligands by EXponential enrichment (SELEX) exposes target proteins to a random library of oligonucleotides. Those genes that bind are separated and amplified by PCR.

[0352] The following are Examples of specific aspects of the disclosure. The Examples are offered for illustrative purposes only, and are not intended to limit the scope of the disclosure in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should, of course, be allowed for.EXAMPLE 1

[0353] The following example describes compositions and methods to evaluate Cas-beta endonuclease-mediated gene editing efficiency in mammalian cells.

[0354] Mammalian expression vector pcDNA3.1 containing a human codon optimized Cas- beta2 gene (SEQ ID NO:1) was synthesized to order by GenScript (Piscataway, NJ USA). The expression construct harbored sequence encoding a SV40 nuclear localization signal (NLS) (SEQ ID NO:2) at the N-terminus and sequence encoding a nucleoplasmin NLS (SEQ ID NO:3) at the C-terminus of the Cas-beta2 protein as shown in Fig.1. To express guide RNAs (gRNAs) that direct Cas-beta2 to target sites, plasmids were synthesized that encode U6 polymerase III promoter sequence, an individual Cas-beta2 gRNA sequences SEQ ID NO:4-7, HDV ribozyme sequence, and a U6 terminator. At the 3’ end of each gRNA sequence (see Fig. 2), each of these plasmids includes a variable spacer sequence that targets a site in the RunXI gene or WTAP gene of Human embryonic kidney (HEK) cell line 293T. Plasmids encoding different spacer sequences were constructed by restriction cloning. Sequences of genomic target sites are shown in Table 2. TABLE2 Target Gene Site name Sequence Listing Target site sequence 5’ – 3’ RunX1 “R1” SEQ ID NO:8 GCCTTCAGAAGAGGGTGCATTTTC T C C [03ression plasmids as follows. HEK293T cells were cultured in Dulbecco’s modified Eagle’s Medium (DMEM) with GlutaMAX (Thermo Fisher Scientific, Waltham, MA, USA), supplemented with 10 % fetal bovine serum and 10,000 units / mL penicillin, and 10,000 μg / mL streptomycin at 37ºC in 5 % CO2. HEK293T cells were seeded into 96-well plates one day before transfection at a density of 2x105 / ml, 90 μl per well. Cells were transfected using transfection reagent FuGENE®HD (Promega Corporation, Madison, WI, USA) according to the manufacturer’s protocol. For each sample, 450 ng of DNA containing 64 fmol of plasmid encoding Cas-beta2 and 64 fmol of plasmid encoding U6-gRNA-HDV was used. A 3:1 ratio of FuGENE®HD Transfection Reagent to DNA was used (0.3μl reagent:100ng DNA per well). Cells were incubated at 37ºC with 5% CO2 for 4 days post transfection before genomic DNA extraction. Cells were washed twice with 200 μl pre-warmed to 37ºC 1X DPBS (Thermo Fisher Scientific) and resuspended in 25 μl cell lysis solution containing 50 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, pH 7.6 and 0.2 mg / ml Proteinase K (New England Biolabs, Ipswich,MA USA). Resuspended cells were incubated at 55ºC for 1 hour and 95ºC for 15 minutes. Cell lysates were prepared from transfected HEK293T and negative control (untrasfected) HEK293T for use in T7 Endonuclease I assay and deep sequencing.

[0356] T7 Endonuclease I Assay. Cell lysates were used in PCR amplification of genomic region comprising each of the Cas-beta2 target sequences. PCR amplification was performed using Q5 Hot Start High-Fidelity 2X Master Mix (New England Biolabs) following the manufacturer’s recommended instructions. PCR reactions were performed using 1μl of the cell lysate and 0.5 μM of each primer in a final reaction volume of 25 μl. PCR primers are shown in Table 3. TABLE 3 Target Primer Sequence Listing Primer sequence 5’ – 3’ C G [0357using the following T7 Endonuclease I assay.25 μL of each completed PCR reaction was combined with 3 μL NEBuffer 2 (New England Biolabs) and 7 μL of water before denaturation at 95ºC for 5 minutes and re-annealing by temperature ramping from 95-85ºC at -2ºC / second followed by ramping from 85-25ºC at -0.1ºC / second. After hybridization reaction 1 μL of T7 Endonuclease I (New England Biolabs) was added to each re-annealed sample and cleavage reactions were incubated at 37ºC for 20 minutes. Fragmented PCR products were analyzed by performing gel electrophoresis using E-Gel Precast Agarose Gel Electrophoresis system (Invitrogen).2 % gel with Ethidium bromide dye was used.8 μL of each sample was combined with 7 μL of E-Gel Sample Loading Buffer (Invitrogen), whole volume (15 μl) was loaded to the well of E-Gel. Genomic modification percentage was calculated by densitometric analysis using ImageJ software.

[0358] Deep Sequencing. Cell lysates of transfected HEK293T cells were used to amplify genomic target regions and fragments were extended by Illumina sequencing that included a unique index for each sample through two rounds of PCR. For primary PCR, Q5 HotStart 2x MasterMix (New England Biolabs) was used. Reactions were set up using 4 μl of cell lysate as template and 0.2 μM of each primer in a final volume of 25μl. Cycling conditions used were:98°C for 2 minutes 30 seconds, 25 cycles of 98°C for 30 seconds, 56.5°C for 30 seconds, 72°C for 25 seconds, and final extension at 72°C for 2minutes. Custom primers were used that are complementary to sequences surrounding genomic targets and include non-complementary ‘tails’ with Illumina adapter sequences. Custom primers used for primary PCR are shown in Table 4 (underline indicates sequence that is complementary to genomic DNA). TABLE 4 SEQ ID NO: Primer Sequence 5’-3’ Description31 GTGACTGGAGTTCAGACGTGTGCTCTTCCGA Indels Primary PCR WTAP TCTCCCTGTACATTTACACTTGA target 1 reverse primer 2 wEngland Biolabs) and used in secondary PCR reactions.

[0360] Secondary PCR was performed using PCR Add-on Kit for Illumina from Lexogen GmbH (Vienna, Austria) with primers encoding Illumina sequences and a 6 nucleotide i7 index (Lexogen i76 nt Index Set (7001-7096)). Cycling conditions for secondary PCR were 98°C for 30 seconds, 9 cycles at 98°C for 10 seconds, 65°C for 20 seconds and 72°C for 30 seconds, followed by final extension at 72°C for 1 minute. Secondary PCR products were purified using Monarch PCR & DNA Cleanup Kit (New England Biolabs), their quantity and quality checked using NanoPhotometer® NP80 spectrophotometer from Implen (München, Germany) and QUBIT 1X dsDNA HS Assay and QUBIT 4 fluorometry system from Invitrogen.

[0361] Purified samples were pooled in an equimolar ratio. The resulting library was analyzed using Bioanalyzer from Agilent (Santa Clara, CA, USA) and prepared for deep sequencing according to Illumina specifications. Sequencing was performed using the MISEQ Reagent Kit v2 (300-cycles) on the MISEQ System with 7% PHIX v.3 all from Illumina (San Diego, CA, USA). Sequencing data was analyzed using Geneious Prime 2023.1.1 software from Biomatters (Auckland, NZ). Briefly, reads were filtered and trimmed using BBDuk (Joint Genome Institute, Berkeley, CA, USA) and a 50bp editing window around the cleavage site was checked for the formation of indels. Table 5 shows editing outcomes relative to WTAP W6 reference sequence: CCCGTGGGAGAGTCTACACT*T*TCATA*C*CCCGCACTGAGTTGATTTACGTAAC (SEQ ID NO:132), wherein PAM sequence CCC (bold nucleotides) is followed by target sequence (underlined nucleotides) and asterisks indicate expected endonuclease cut sites. Table 6 shows editing outcomes relative to RunX1 R1 reference sequence: CCCGCCTTCAGAAGAGGGTG*C*ATTTT*C*AGGAGGAAGCGATGGCTTCAGACA G (SEQ ID NO:133). In foregoing W6 and R1 reference sequences, PAM sequence CCC (bold nucleotides) is followed by target sequence (underlined nucleotides) and asterisks indicate expected endonuclease cut sites. In following Table 5 and Table 6, dashes represent nucleotide deletions in outcome sequence.TABLE 5 Outcome Outcome SEQ ID Outcome Sequence Frequency Description NO C C C C C COutcome Outcome SEQ Sequence Frequency Description ID G G G4 nt 0.64% Deletion (1551) pos.17-20 133 CCCGCCTTCAGAAGAG- - - - CATTTTCAGGAGGAAGCGATGGCTTCAGACAG Gig. 3. Editing efficiency indicates the fraction of DNA that was edited, relative to total DNA, at the target locus of transfected mammalian cells as determined by Illumina NGS. EXAMPLE 2

[0363] The following example describes methods of engineering improved Cas-beta polypeptides and provides engineered Cas-beta polypeptides with improved binding and / or editing efficiency in mammalian cells.

[0364] Cryo-EM-derived structure of the Cas-beta2-sgRNA-DNA complex was analyzed to identify Cas-beta amino acids in contact with the guide RNA or DNA substrate (Fig. 4). Residues were considered to be in contact they contained atoms with van der Waals radii overlapping by more than 0.4 Å, as determined by ChimeraX software (Pettersen 2020). 34 neutral or acidic amino acids were selected and separately mutated to arginines (Fig.4 and Fig. 5) to determine if any improve interactions between Cas-beta2 and its guide RNA and DNA substrate. DNA fragments coding for the respective mutation were synthesized (Twist Biosciences) and cloned into the Cas-beta2 expression vector. HEK293T cells were transfected with plasmids encoding Cas-beta variants (SEQ ID NOs:32-65) along with plasmids encoding guide RNAs targeting RunXI (Target R1) or WTAP (Target W6) target sites as described in Example 1.

[0365] Each of the engineered variants Cas-beta variants (SEQ ID NOs:32-65) was evaluated for editing efficiency using the T7 Endonuclease I assay described in Example 1. The observed editing efficiency for each variant is shown in Fig. 6, which indicates editing efficiencies for each variant at the two targets (results from two independent, repeated experiments).

[0366] Single amino acid mutations that conferred increased activity were combined to produce double amino acid Cas-beta variants (SEQ ID NOs:66-69) and triple amino acid Cas- beta variants (SEQ ID NOs:70-71). Double Cas-beta variants include S17R / E771R (SEQ ID NO:66); V21R / D342R (SEQ ID NO:67) V21R / E615R (SEQ ID NO:68) and V21R / E766R (SEQ ID NO:69). Triple Cas-beta variants include V21R / D342R / E766R (SEQ ID NO:70) andV21R / D342R / E771R (SQ ID NO:71). HEK293T cells were transfected with plasmids encoding the foregoing single, double and triple Cas-beta2 mutants along with plasmids encoding guide RNA targeting RunXI (Target R1 and R3) or WTAP (Target W1 and W6) sites. Editing efficiency was evaluated by Illumina SEQ. Fig. 7 shows the observed editing efficiencies for each of the indicated Cas-beta variant.

[0367] Further variants were produced, including S297R / T301R / D342R / E615R / E762R / E766R (SEQ ID NO:72) and S17R / V21R / D142R / L253R / E615R / E762R / E766R (SEQ ID NO:73). HEK293T cells were transfected with empty plasmids, wild type Cas-beta2 and mutant Cas-beta variants along with plasmids encoding guide RNA targeting RunXI (Target R1 and R3) or WTAP (Target W1 and W6) sites. Editing efficiency was evaluated by T7 Endonuclease I assay as described above. Fig. 8 shows the observed DNA editing efficiencies for the indicated Cas-beta variants as compared against wild-type (WT) enzyme and empty plasmid (NC).

[0368] The foregoing results demonstrate that the engineered Cas-beta2 variants disclosed herein provided improved NHEJ-mediated gene editing activity on genomic targets in mammalian cells. EXAMPLE 3

[0369] The following example describes the identification of guide RNA residues and regions of residues involved in Cas polypeptide-guide RNA ribonucleoprotein (RNP) effector complexes and provides engineered guide RNA variants with improved binding and / or editing efficiency in mammalian cells.

[0370] To systematically probe which regions of Cas-beta2 guide RNA are either required for or not significantly involved in RNP formation, an RNP pulldown assay was designed. A schematic illustration of the pulldown assay is shown in Fig. 9. First, a library of Casbeta2 sgRNAs containing 3 nt stepwise deletions spanning the entire sgRNA scaffold sequence was designed and synthesized: SEQ ID NOs:74-121. Then, 2 μM of the sgRNA library was incubated with 1 μM of Casbeta2 and 100 nM of a DNA oligoduplex substrate (SEQ ID NO:122) containing a Biotin moiety at the 5’ end of the non-target strand at 37°C for 30 minutes in CutSmart buffer (New England Biolabs). The reaction mixture was then incubated with 10 μl of Dynabeads™ M-280 Streptavidin (Invitrogen) at room temperature for 15 minutes. The beads were then washed to remove unbound components. Finally, the beads were incubated at 75°C for 15 minutes to denature the Cas-beta RNPs and the resulting mixturepurified using Monarch® RNA Cleanup Kit (New England Biolabs) to collect the sgRNAs that facilitated RNP formation and substrate cleavage.

[0371] Guide RNAs from control library and pulldown experiments were prepared for sequencing. Collibri™ Stranded RNA Library Prep Kit for Illumina™ Systems (Thermo Fisher Scientific) was used according to the manufacturer‘s instructions but with omission of fragmentation step in the protocol. Sample libraries were pooled and sequenced using the MiSeq Reagent Kit v2 (300-cycles) (Illumina) on the MiSeq System (Illumina) with 5% PhiX v.3 (Illumina). Reads matching references in the guide RNA library with deletions were counted. Control sample counts were used for normalization and normalized frequencies were used to determine which deletions were enriched for in the sample pool by calculating the fold change of guide RNAs with specific deletions. Formula used for normalization: Normalized Frequency = [Read With Deletion Frequency] / [Control Frequency / Average Control Frequency].

[0372] Overall enrichment or depletion of particular sgRNA variants is shown in Fig. 10. In Fig.10, the fold-change (y-axis) indicated the change in the fraction of the indicated sgRNA variant (x-axis) in the combined sgRNA pool that was obtained from the pull-down assay versus the fraction of that sgRNA variant in the initial sgRNA pool. Several variants showed a higher than 3-fold sequencing read frequency change in the pulldown samples compared to the initial sgRNA pool and the nucleotide deletions that these variants encoded are shown in Fig. 11. These experiments reveal positions of sgRNA which could be modified to improve RNP binding and substrate cleavage activity.

[0373] Based on results showing that deletions of sgRNA regions D16 and D17 (see Fig. 11) were particularly enriched in the RNP-pulldown screen (Fig. 10), sgRNA was modified by deletion of the region in the first (closest to 5’ terminus) stem-loop hairpin formed by bases 33- 45, such that deletion encompass regions D16 and D17. Versions of this deletion variant sgRNA(Δ33-45) were generated that target R1, R3, W1, and W6 target sites respectively (SEQ ID NO:123-126) (Fig. 12A). HEK293T cells were transfected with plasmids encoding wild- type Casbeta-2 along with plasmids encoding one of the wild-type (SEQ ID NO:4-7) or modified (SEQ ID NO:123-126) sgRNAs. Fig. 12B shows the observed DNA editing efficiency as determined by T7 Endonuclease I assay.

[0374] Although not bound by theory or any particular mechanism, it was hypothesized the stabilization of RNA secondary structure of the Cas-beta2 sgRNA as observed in the cryo-EM derived Cas-beta2-sgRNA-DNA complex might improve the efficiency of complex formation and DNA editing activity. Hairpin insertion variants “SS-HP sgRNAs” (SEQ ID NOs: 127 and128) were generated by inserting sequence coding for RNA hairpin-forming sequence (5’- GGACUUCGGUCC-3’, positions 39-50 of SEQ ID NOs: 127 and 128) after the 38thbase in the Cas-beta2 sgRNA molecule (Fig. 14A). See, e.g., Riesenberg et al. 2022, Nature Communications 13(489):1-8. HEK293T cells were transfected with plasmids encoding wild- type Casbeta-2 as well as V21R / E615R and S297R / T301R / D342R / E615R / E762R / E766R Casbeta variants along with plasmids encoding wild-type (SEQ ID NO:4-7) and SS-HP sgRNAs (SEQ ID NOs: 127-128). Fig. 14B shows the observed DNA editing efficiency as determined by T7 Endonuclease I assay.

[0375] The foregoing results demonstrate that modified sgRNAs variants disclosed herein provided increased NHEJ-mediated gene editing activity on genomic targets in mammalian cells, as compared to unmodified sgRNA counterpart. EXAMPLE 4

[0376] The following provides additional examples of engineered Cas-beta polypeptides designed to provide improved binding and / or editing efficiency in mammalian cells.

[0377] Additional Cas-beta polypeptides (SEQ ID NOs: 129-131) were aligned to the Cas- beta2 amino acid sequence using EMBOSS Needle (Madeira, F et al (2022) Nucleic Acids Res 50(W1):W276-W279) (Fig. 13-15). For certain amino acids that conferred improved activity in Cas-beta2 polypeptide as demonstrated above, the corresponding amino acids were identified in the other Cas-beta polypeptides (SEQ ID NOs: 129-131). In particular, amino acids were identified corresponding to each of the following engineered (mutated) amino acids of SEQ ID NO:1: serine at position 17, valine at position 21, aspartate at position 142, aspartate at position 244, serine at position 297, threonine at position 301, aspartate at position 342, glutamate at position 615, glutamate at position 762, glutamate at position 766, and glutamate at position 771.

[0378] Provided is a Cas-beta1 polypeptide comprising one or more altered amino acid residue, such as a deletion, insertion or substitution, at any one or more of the corresponding Cas beta- 1 positions shown in Table 7. Thus for example provided is a Cas beta-1 polypeptide comprising a substitution (e.g. arginine) at position 20, position 145, position 235, position 288, position 333, or position 749 (or any combination of the foregoing substitutions) of SEQ ID NO:129, which preferably provides improved binding and / or editing activity. TABLE 7 Cas beta-2 position Corresponding Cas beta-1 positionSEQ ID NO:1 pos 142 SEQ ID NO:129 pos 145 SEQ ID NO:1 pos 244 SEQ ID NO:129 pos 235ue, such as a deletion, insertion or substitution, at any one or more of the corresponding Cas beta- 3 positions shown in Table 8. Thus for example provided is a Cas beta-3 polypeptide comprising a substitution (e.g. arginine) at position 15, position 19, position 136, position 234, position 287, position 346, or position 759 (or any combination of the foregoing substitutions) of SEQ ID NO:130, which preferably provides improved binding and / or editing activity. TABLE 8 Cas beta-2 position Corresponding Cas beta-3 position SEQ ID NO:1 os 17 SEQ ID NO:130 os 15

[0380] Provided is a Cas-beta4 polypeptide comprising one or more altered amino acid residue, such as a deletion, insertion or substitution, at any one or more of the corresponding Cas beta- 3 positions shown in Table 9. Thus for example provided is a Cas beta-3 polypeptide comprising a substitution (e.g. arginine) at position 7, position 11, position 132, position 234, position 242, position 287, position 291, position 598, position 749, or position (or any combination of the foregoing substitutions) of SEQ ID NO:131, which preferably provides improved binding and / or editing activity. TABLE 9 Cas beta-2 position Corresponding Cas beta-4 positionSEQ ID NO:1 pos 142 SEQ ID NO:131 pos 132 SEQ ID NO:1 pos 244 SEQ ID NO:131 pos 234

[0381] The following provides examples of engineered Cas-beta polypeptides and the use of large language model prediction in conjunction with experimentation to identify variants of Cas-beta that provide improved binding and / or editing efficiency in mammalian cells.

[0382] An open-source protein language model ESM-2 was employed to predict variant effects of all mutations on the function of Cas-beta2 (Lin et al 2022 bioRxiv 2022.07.20.500902; doi.org / 10.1101 / 2022.07.20.500902; Rives et al 2021 Proc. Nat’l Acad Sci 118(15): e2016239118). This was measured via masked marginal scoring function which gives the log odds ratio of the mutated position, allowing to infer the effect of the mutation on protein (Meier et al 2021 bioRxiv 2021.07.09.450648; doi.org / 10.1101 / 2021.07.09.450648). According to the model, the higher the score, the higher probability of the mutation to yield a positive effect. For our case, we used ESM-2 model with 15B parameters to predict mutational landscape of the S297R / T301R / D342R / E615R / E762R / E766R (SEQ ID NO:72) variant Cas-beta2 (referred to herafter as “M67”). Out of all the predictions, we selected 9 unique alterations in the M67 background that scored the highest (see Table 10), constructed expression vectors for the resulting Cas-beta2 M67 variant vectors, and tested the variants for DNA editing activity in HEK293T cells as described below. Table 10 shows the further alterations of wild type residues that were made in M67 (SEQ ID NO:72) variant background. Table 10 Variants of M67 Substition Wild type Substituted Score 2 1 9 9 2W560L (SEQ ID NO:185) 560 W L 4,237652 L576K (SEQ ID NO:186) 576 L K 4,020621 Q572 R (SEQ ID NO 187) 572 Q R 3709371 9 [0383 get sites,plasmids encoding a U6 polymerase III promoter sequence, the Cas-beta2 single guide RNA sequence (SEQ ID NOs: 202-212), a HDV ribozyme sequence and a U6 terminator sequence were synthesized. At the 3’ end of the single guide RNA sequence was encoded a variable spacer sequence targeting sites in human embryonic kidney (HEK) cell line 293T RunX1 and WTAP genes (SEQ ID NOs: 217-227). Plasmids encoding different spacer sequences were constructed by restriction cloning.

[0384] To evaluate Cas-beta2-mediated gene editing efficiency, HEK cell line 293T (ATCC- CRL-3216) was transfected with Cas-beta2 variants and their guide RNA expression plasmids. HEK293T cells were cultured in Dulbecco’s modified Eagle’s Medium (DMEM) with GlutaMAX (Thermo Fisher Scientific), supplemented with 10 % fetal bovine serum (Thermo Fisher Scientific) and 10,000 units / mL penicillin, and 10,000 μg / mL streptomycin (Thermo Fisher Scientific) at 37ºC in 5 % CO2. HEK293T cells were seeded into 96-well plates (Thermo Fisher Scientific) one day before transfection at a density of 2x105 / ml, 90 μl per well. Cells were transfected using transfection reagent FuGENE HD (Promega Corporation) according to the manufacturer’s protocol. For each sample 300 ng of DNA containing 64 fmol of plasmid encoding Cas-beta2 variant and 64 fmol of plasmid encoding U6-sgRNA-HDV was used. For experiments FuGENE® HD Transfection Reagent:DNA ratio 3:1 used (0.3μl reagent:100ng DNA per well).

[0385] Cells were incubated at 37ºC with 5% CO2 for 4 days post-transfection before genomic DNA extraction. Cells were washed twice with 200 μl pre-warmed to 37ºC 1X DPBS (Thermo Fisher Scientific) and resuspended in 25 μl cell lysis solution containing 50 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, pH 7.6 (Sigma Aldrich) and 0.2 mg / ml Proteinase K (New England Biolabs). Resuspended cells were incubated at 55ºC for 1 hour and 95ºC for 15 minutes. Genomic DNA was extracted by cell lysis and genomic region surrounding each CasB2 target site was PCR amplified (SEQ ID NOs:232-235). PCR amplification was performed using Q5 Hot Start High-Fidelity 2X Master Mix (New England Biolabs) according to the manufacturer’s instructions. The reaction was set up using 1μl of the cell lysate and 0.5 μM of each primer in a final reaction volume of 25 μl.

[0386] Genome editing frequencies were estimated using T7 Endonuclease I assays.20 μL of each PCR reaction was combined with 3 μL NEBuffer 2 (New England Biolabs) and 7 μL of water before denaturation at 95ºC for 5 minutes and re-annealing by temperature ramping from 95-85ºC at -2ºC / s followed by ramping from 85-25ºC at -0.1ºC / s. 1 μL of T7 Endonuclease I (New England Biolabs) was added to each re-annealed sample and cleavage reactions were incubated at 37ºC for 20 min. Fragments were analyzed by performing gel electrophoresis using E-Gel Precast Agarose Gel Electrophoresis system (Invitrogen). 8 μL of each sample was mixed with 7 μL of E-Gel Sample Loading Buffer (Invitrogen), then the whole volume was loaded to the well. Target cleavage percentage was calculated using ImageJ software.

[0387] Table 11 shows NHEJ DNA editing efficiency of M67 Cas-beta2 as compared to each further altered M67 variants listed in Table 10. Editing efficiency was measured at WTAP target sites W3 and W6 and RunXI target sites R1 and R5 in human cells. The two M67 variants containing a Q572R or F607S substitution, respectively, exhibited comparable or higher mean editing activity than original M67. Table 11 SEQ ID Cas-Editing Efficiency NO:beta

[0388] These, t containing both substitutions (SEQ ID NO:189), were chosen for more thorough analysis on a larger set of genomic targets. The Q572R / F607S variant (SEQ ID NO:189) exhibited higher mean rate of indel formation on most targets than the parent version Cas-beta2 M67 (Figure 2).

[0389] To generate results shown in Figure 19, indel formation rate was evaluated by Illumina NGS. Lysates of HEK293T cells were prepared for deep sequencing to determine the activity of Cas-beta2 variants in eukaryotic cells by studying the rates of NHEJ outcomes and resulting indels in the genomic target sites of treated cells. Briefly, the genomic target regions were amplified and fragments extended with Illumina sequences including a unique index for eachsample through two rounds of PCR. 4 μL of each lysate sample was used as template in the primary PCR reaction. For the primary PCR custom primers were used that were complementary to the sequences surrounding the genomic targets and had non-complementary ‘tails’ with Illumina adapter sequences (SEQ ID NOs:244-279). Q5 HotStart 2x MasterMix (New England Biolabs) was used for the primary PCR and the reaction set up using 4 μL of cell lysate as template and 0.2 μM of each primer in a final volume of 25 μL. The cycling conditions used were: 98°C for 2min 30s, 24 cycles of  98°C for 30s, 56.5°C for 30s, 72°C for 25s, and final extension at 72°C for 2min. The primary PCR product was purified using Ampure XP (Beckman Coulter) magnetic beads and used for the secondary PCR.

[0390] Secondary PCR was performed using either NEBNext Ultra II Q5 Master Mix (New England Biolabs) or 2X Q5 High Fidelity Hot Start PCR Master Mix (New England Biolabs) with either custom primers (synthesis ordered from Integrated DNA Technologies) containing Illumina sequences and i7 and i5 index (on reverse and forward primer, respectively) or NEBNext Multiplex oligos for Illumina (Dual Index Primers Sets 1 or 2). Cycling conditions for secondary PCR were as follows: 98°C for 2 min, 10 cycles at 98°C for 10s, 60°C for 30s and 72°C for 2 min, followed by final extension at 72°C for 2 min, when using 2X Q5 High Fidelity Hot Start PCR Master Mix (New England Biolabs) and custom primers, or 98°C for 30s, 10 cycles at 98°C for 10s, 65°C for 1min 15s, followed by final extension at 65°C for 5min. The secondary PCR products were purified using Monarch PCR & DNA Cleanup Kit (New England Biolabs), their quantity and quality checked using spectrophotometry (NanoPhotometer® NP80, IMPLEN) and fluorimetry (Qubit 1X dsDNA HS Assay, Qubit 1X dsDNA BR Assay and Qubit 4, Thermo Fisher Scientific). The purified samples were pooled in an equimolar ratio and size selection performed using Ampure XP (Beckman Coulter) magnetic beads. The resulting library quantified using NEBNext Library Quant Kit for Illumina (New England Biolabs). Final library pool was prepared for deep sequencing according to Illumina’s specifications. Paired end sequencing was performed using the MiSeq Reagent Kit v2 (300-cycles) (Illumina) on the MiSeq System (Illumina) with 7% PhiX v.3 (Illumina). All sequencing data analysis was done using Geneious Prime 2025.0. Reads were trimmed and filtered using BBDuk and mapped to the reference sequence with a 50 nt quantification window set around the target site. Reads that had differences from the reference sequence in this window were included in genome editing efficiency calculations.EXAMPLE 6

[0391] The following example describes the delivery to cells of Cas-beta2 variants further altered to include Q572R / F607S substitutions along with synthetic guides as RNPs. The results confirm the improved DNA editing efficiency of the Q572R / F607S variants.

[0392] HEK293T cells were electroporated with (i) RNPs of the following purified Cas-beta2 variants: V21R / E615R (SEQ ID NO:68) (referred to hereafter as “M43”) (SEQ ID NO:68), M67 (SEQ ID NO:72), M43 further altered to include Q572R / F607S substitutions (SEQ ID NO:201) or M67(Q572R / F607S) (SEQ ID NO:189) nucleases and (ii) synthetic sgRNAs targeting W6, CD34, CD151 and AAVS1 (SEQ ID NOs: 203; 213-215). HEK293T cells were maintained in Dulbecco’s modified Eagle’s Medium (DMEM) with GlutaMAX (Thermo Fisher Scientific), supplemented with 10 % fetal bovine serum (Thermo Fisher Scientific) and 10,000 units / mL penicillin, and 10,000 μg / mL streptomycin (Thermo Fisher Scientific) at 37ºC with 5 % CO2 incubation. Two days before electroporation, cells were split and maintained as described above. On the day of the experiment, cells were washed using Dulbecco’s Phosphate Saline Buffer (DPBS) (Thermo Fisher Scientific) and trypsinized using TrypLE Express Enzyme (1X) (Thermo Fisher Scientific). Collected cells resuspended in DPBS were diluted using SF Cell Line Nucleofector Solution (Lonza) supplemented with Supplement 1 (Lonza). RNPs were assembled in Nucleofector Solution supplemented with Supplement 1 using 126 pmol of each purified nuclease and 160 pmol of sgRNA. RNP assembly reactions were incubated at room temperature for 20 minutes, then kept at 4°C until nucleofection.2×105cells per sample were added into the RNP assembly reaction and whole volume was transferred to 16-well Nucleocuvette Strips (Lonza). Cells were electroporated using preprogrammed protocol CM-130 using 4D-Nucleofector (Lonza). After electroporation, contents of each well were split and seeded into 4 wells containing pre-warmed Dulbecco’s modified Eagle’s Medium (DMEM) with GlutaMAX (Thermo Fisher Scientific), supplemented with 10 % fetal bovine serum (Thermo Fisher Scientific) in 96-well plate (Thermo Fisher Scientific). Cells were then grown for 48 h at 37ºC in 5 % CO2 before genomic DNA extraction. The cells were washed twice with 200 μl 1X DPBS (Thermo Fisher Scientific) and resuspended in 25 μl 50 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, pH 7.6 (Sigma Aldrich) and 0.2 mg / ml Proteinase K (New England Biolabs) lysis buffer. Resuspended cells were lysed by incubating at 55ºC for 60 minutes and 95ºC for 15 minutes. Target loci were PCR amplified (SEQ ID NOs: 234-235; 236-241) and T7 Endonuclease I assay performed to assess NHEJ rate. Afterwards, amplified genomic DNA was submitted for Illumina NGS to confirm indel formation rate. The results of these experiments shown in Figure 20 demonstrate that variantscontaining Q572R / F607S mutations exhibited higher indel formation efficiency than their respective parent nucleases M43 or M67. EXAMPLE 7

[0393] The following example describes the delivery of Cas-beta variants by mRNA transfection of mammalian cells and demonstrates that for most of the targets tested Cas-beta variants containing Q572R / F607S substitutions induced indel formation at higher frequencies, as compared to equivalent Cas-beta variants lacking the Q572R / F607S substituions.

[0394] Cas-beta2 and mutant variant mRNA was synthesized in vitro for use in transfection experiments. To produce in vitro transcription templates, Cas-beta2 and mutant sequences were cloned to a plasmid vector optimized for mRNA synthesis via the process described hereafter. This optimized plasmid vector contains 5’ and 3’ human hemoglobin beta subunit (HBB) UTR sequences (SEQ ID NOs: 280-281), T7 RNA polymerase promoter, and kanamicyn resistance gene (Figure 21). To clone target genes into this optimized vector, wild type Cas-beta2 and Cas-beta2 M67 variant sequences were amplified by PCR during which overlap sequences from the destination vector were added; vector was also amplified by PCR during which overlap sequences of the target genes were added to the 5’ and 3’ ends of a linear PCR product. Products were analyzed on agarose gel and purified from the reaction using New England Biolabs Monarch PCR and DNA Cleanup Kit (5 μg). Assembly reactions were set up using New England Biolabs NEBuilder HiFi DNA Assembly Master Mix and incubated for 1hour at 50°C, then transformed to E. coli DH5 cells (New England Biolabs). Plasmid DNAwas purified from several clones of each of the constructs transformed and sequences were verified by whole plasmid sequencing.

[0395] After sequence verification, correct plasmids were used as PCR templates to produce in vitro transcription templates for mRNA synthesis. During this PCR step, target sequences were amplified using universal primers. Reverse primer is elongated and used for 120 nt poly(A) tail addition. After PCR, products were confirmed by gel electrophoresis and purified from PCR reactions using New England Biolabs MONARCH PCR and DNA Cleanup Kit (5 μg), concentrations measured using spectrophotometry (NanoPhotometer® NP80, IMPLEN). Approximately 1 μg of each purified PCR product was then used as a template for in vitro transcription reaction performed using HiScribe® T7 mRNA Kit with CleanCap® Reagent AG (New England Biolabs); kit-provided UTP was not used, instead, N1-Methylspeudouridine was added to a final concentration of 5 mM. In vitro transcription reaction was run for 1 hour at 37°C after which DNaseI and nuclease-free water was added to the reaction and incubatedfor 30 min at 37°C to remove template DNA. After this step, IVT reaction product was purified using the MONARCH RNA Cleanup Kit (500 μg) (New England Biolabs). Part of the purified samples were set aside to assess their quality. Sample concentrations were measured using spectrophotometry (NanoPhotometer® NP80, IMPLEN) and fluorimetry (Qubit RNA Broad Range Assay Kit and Qubit 4, Thermo Fisher Scientific). Integrity of mRNA was assessed by performing capillary electrophoresis using 4150 TapeStation System together with RNA Screen Tape reagents (Agilent). Synthesized mRNA was aliquoted and kept at -80°C until use. HEK293T (ATCC-CRL-3216), cells were transfected with mRNA of each Cas-beta2 nuclease and sgRNA of one of four genomic targets (CD34, CD151, VEGFA and W6) (SEQ ID NOs: 213, 214, 216, and 203).

[0396] HEK293T cells were maintained in Dulbecco’s modified Eagle’s Medium (DMEM) with GlutaMAX (Thermo Fisher Scientific), supplemented with 10 % fetal bovine serum (Thermo Fisher Scientific) and 10,000 units / mL penicillin, and 10,000 μg / mL streptomycin (Thermo Fisher Scientific) at 37ºC with 5 % CO2 incubation. HEK293T cells were seeded into 96-well plates (Thermo Fisher Scientific) one day prior to transfection at a density of 18,000 cells per well. Cells were transfected using Lipofectamine MessengerMAX Transfection Reagent (Invitrogen) following the manufacturer’s recommended protocol. For each well of a 96-well plate a total amount of 100 ng RNA was used; in some experiments, different total RNA amounts were used (200 ng and / or 300 ng). This total amount consists of mRNA and sgRNA in a ratio of 1:4. Cells were incubated at 37ºC for 72 hours post transfection in 5 % CO2 before genomic DNA extraction. The cells were washed twice with 200 μl 1X DPBS (Thermo Fisher Scientific) and resuspended in 25 μl 50 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, pH 7.6 (Sigma Aldrich) and 0.2 mg / ml Proteinase K (New England Biolabs) lysis buffer. Resuspended cells were incubated at 55ºC for 60 minutes and 95ºC for 15 minutes. Target loci were PCR amplified (SEQ ID NOs: 236-243) and T7 Endonuclease I assay performed to assess NHEJ rate. Afterwards, amplified genomic DNA was submitted for Illumina NGS to confirm indel formation rate. Results shown in Figure 22 demonstrate that for 3 of the 4 tested targets, variants containing Q572R / F607S mutations exhibited higher indel formation efficiency than their parent nucleases M43 or M67. EXAMPLE 8

[0397] The following example provides engineered Cas-beta variants and describes the use of a second round of large language model prediction in conjunction with experimentation toidentify variants of Cas-beta that provide improved binding and / or editing efficiency in mammalian cells.

[0398] The Cas-beta2 M67(Q572R / F607S) variant was iterated again through ESM-2 model. Predicted scores were compared with its parent version Cas-beta2 M67. From this, two strategies for further selection were applied. One strategy involved selecting the 3 best variants that had both the positive score in general and greatest score difference between 1st and 2nd iterations, namely variants containing V610H, K619T, K557V. Second strategy involved selecting the 3 best variants that already had a high likelihood score (see Table 10) and whose score difference slightly increased (>0.1) in second iteration, namely variants containing A603N, C442F, E683V. Table 12 shows the scores and score differences used to selected new substation variants). Additionally, since we observed that mutation at position 607 (variant F607S) caused an increase in activity with some targets (Figure 20; Figure 22), we designed 5 variants that contained a variable amino acid, namely F607D, F607Q, F607A, F607R, F607Y (SEQ ID NOs: 189-194). This was done to evaluate whether the specific amino acid or its property would cause an even greater increase in activity. Table 12 shows the position of each substitution Table 12 PositionWild typeSubstituted Score difference to Variant aa aaScoreIteration 1 SEQ ID NO. [iting activity in HEK293T cells and compared against Cas-beta2 M67 and Cas-beta2 M67(Q572R / F607S). The editing efficiency of each of these variants at each of genomic targets W2, W3, W4, W5, R3, R4, and R5 is shown in Table 13, which demonstrates that substitution of F607S to F607R or incorporation of E683V yielded increased mean DNA editing efficiency on some targets. Table 13 SEQ Editing Efficiency193 M67 Q572R / F607R 35.1 3.8 2.2 36.5 23.6 5.5 9.9 194 M67 Q572R / F607Y 22.6 3.7 12.0 31.0 25.7 8.6 6.6 190 M67 Q572R / F607D 263 90 47 277 208 61 146

[0400] The following example describes the use of Cas-beta variants disclosed herein for SDN- 3 insertion of a gene of interest at a desired target site.

[0401] Wild type Cas-beta2 (SEQ ID NO:1), Cas-beta2 M67 and Cas-beta2 M67(Q572R / F607S) (SEQ ID NO:189) were evaluated for homology directed repair (HDR) efficiency and compared against SpCas9. HEK293T cells were electroporated with individual nuclease–sgRNA RNP complexes, targeting the AAVS1 gene, along with a dsDNA HDR template encoding the EGFP gene and ~500 nt homology arms for AAVS1.

[0402] Electroporation was performed using 4D-Nucleofector System (Lonza) and SF Cell Line 4D-Nucleofector X Kit S (Lonza) according to manufacturer’s protocol for HEK293 cells. 126 pmol of each purified Cas-beta2 nuclease variant was mixed with 160 pmol of synthetic sgRNA (SEQ ID NO:282) . In the case of SpCas9, 50 pmol of EnGen® Spy Cas9 NLS (New England BioLabs) with 100 pmol sgRNA (SEQ ID NO:316) The mixtures were incubated for 15-20 min at room temperature. Meanwhile, HEK293T cells were washed with 1X DPBS (Thermo Fisher Scientific) and harvested. 2 x 105cells / sample were resuspended in 10,5 μl / sample supplemented Nucleofector solution.9,6 μl of prepared ribonucleoprotein complex, 0,5 μg of dsDNA template and 10,5 μl ofHEK293T cells were mixed (final volume 25 μl) and entire volume of mixture was transferred into electroporation cuvette. Cells were electroporated using the program CM-130. After the electroporation, 75 μl of prewarmed Dulbecco‘s modified Eagle's medium (DMEM) supplemented with 10 % Fetal Bovine Serum (FBS) (without antibiotics) was added to the cells in the cuvette and 25 μl of cells were transferred into 96 well plate filled with 175 μl prewarmed media without antibiotics. Cells were harvested five days after electroporation and washed twice with 1X DPBS before analysis. HDR efficiency was determined by MACSQuant® Analyzer 16 Flow Cytometer (Miltenyi Biotec). Flow cytometry results were analyzed with FlowJo™ Software v10.10.0 (BD Biosciences).

[0403] Figure 23 shows the mean percentage of GFP-positive cells for each nuclease variant as well as control reactions where the HDR template encoding the EGFP gene or guide RNAwere omitted. Cas-beta2 variant M67(Q572R / F607S) exhibited the highest rate of EGFP gene insertion. EXAMPLE 10

[0404] The following example provides variant sgRNAs and describes a method of disrupting a 5’-UUU-3’ tract in tracrRNA to provide sgRNAs variants that improve Cas-beta editing activity.

[0405] To improve Cas-beta2 mediated DNA editing efficiency via engineering of the single guide RNA scaffold (single guide RNA without the spacer sequence), a 5’-UUU-3’ tract in the tracrRNA of Cas-beta2 was disrupted to evaluate whether such alteration reduces the likelihood of premature transcription termination. In one case, base pair of U88 : A135 was flipped to A88 : U135 (MOD21_sgRNA (SEQ ID:283). In another case, the U88 : A135 base pair bases were altered to be C88 : G135 (MOD25_sgRNA (SEQ ID:284). These guide RNA variants were compared against wild type guide RNA in facilitating DNA editing activity using Cas- beta2 M67 in HEK293T cells. Results shown in Figure 24 indicated that at least MOD25_sgRNA improved Cas-beta genome editing activity. EXAMPLE 11

[0406] The following example provides altered Cas-beta2 guide RNAs and demonstates that in certain cases alterations or subtitutions of only a few nucleotides improved Type V CRISPR- Cas system DNA editing efficiency. The following example also describes a method of selectively enriching for sgRNA variants that facilitate Cas endonuclease DNA cleavage, which method can be used, e.g., to select sgRNA modifications that improve DNA cleavage activity and / or shorter sgRNA molecule that retain DNA cleavage activity. To systematically probe which regions of Cas-beta2 guide RNA are expendable for RNP formation and evaluate potential RNA mutations beneficial to target cleavage a modified RNP pulldown assay that could eliminate binding events that did not result in substrate cleavage was designed. The method involved the use of a DNA double stranded substrate, the 5’ end of each strand was coupled to a different affinity tag to enable two-step enrichment: the first step isolated and selectively removed unbound DNA substrate and RNPs bound to intact (uncleaved) substrate; the second step isolated and selected for RNPs bound to cleaved substrate as illustrated in Figure 25.

[0407] A library of altered Casbeta2 sgRNAs containing 1 and 3 nt stepwise deletions, 4 nt GAAA insertions and 1 nt substitutions spanning the whole sgRNA scaffold sequence wasdesigned and synthesized. Then, 2 μM of the sgRNA library was incubated with 1 μM of Casbeta2 and 100 nM of a DNA oligoduplex substrate (the duplex non-target strand SEQ ID NO:2855’ terminus was linked to azide tag and the target strand SEQ ID NO: 2865’ terminus was linked to biotin tag) at 37°C for 30 min in CutSmart buffer (New England Biolabs). The reaction mixture was then incubated with 5 μl of Dynabeads™ M-280 Streptavidin (Invitrogen) at room temperature for 15 min. Beads were then washed to remove unbound DNA substrate and RNPs bound to intact substrate. Afterwards, 1 nmol of DBCO-PEG-biotin conjugate was added to introduce a biotin moiety via click-chemistry to the non-target strand of the cleaved oligoduplex that remains bound to the RNP and reaction incubated for 30 min at 37°C. Samples were diluted and dialyzed using 30 kDa filter spin columns (Thermo Fisher Scientific) to remove free conjugate. 5 μl of Dynabeads™ M-280 Streptavidin to the volume that was retained on the filter membrane and incubated at room temperature for 15 min and vortexed. Beads were washed and resuspended in 50 μl 1x CutSmart buffer. RNA was released from RNPs by adding 1 μl of Proteinase K (Thermo Fisher Scientific) and 1 μl 0.5 M EDTA and incubating at 55°C for 10 min and 75°C for 10 min. The resulting mixture was purified using MONARCH RNA Cleanup Kit (New England Biolabs) to collect the sgRNAs that facilitated RNP formation and substrate cleavage. Guide RNAs from control library and pulldown experiments were prepared for sequencing. Collibri™ Stranded RNA Library Prep Kit for Illumina™ Systems (Thermo Fisher Scientific) was used according to the manufacturer‘s instructions but without the fragmentation step in the protocol. Sample libraries were pooled and sequenced using the MiSeq Reagent Kit v2 (300-cycles) (Illumina) on the MiSeq System (Illumina) with 5% PhiX v.3 (Illumina). Reads matching references in the guide RNA library with deletions were counted. Control sample counts were used for normalization and normalized frequencies were used to determine which deletions were enriched for in the sample pool by calculating the fold change of guide RNAs with specific deletions. The formula used for normalization was as follows: Normalized Frequency = /

[0408] of the ratio of sgRNA variant frequency in the treated sample versus the frequency in the untreated library. The frequency change of the 20 of the most highly enriched sgRNA variants (SEQ ID NOs:287-306) are shown in Table 14. Based on prior experiments and this sequencing data regions observed to be dispensable to DNA binding and cleavage in vitro were removed to yield altered sgRNA variants provided in Table 15, wherein underline indicates bases addedrelative to wild type sgRNA, dashes indicate bases deleted relative to wild-type, and uridine nucleosides are shown as T in accordance with sequence listing standards. Table 14 SEQ ID sgRNA Log2Frequency SEQ ID sgRNA Log2Frequency NO: Variant Change NO: Variant Change 87 i 2 2229848 297 G5C 1186315SEQ ID NOsgRNA Variant SequencesA A T C T C A A C G C G C G C G C GTTAGGCGTTCCGTCTCGACTATGCCGTACCACTAGACCGAGCCTACACGGCACGCGGTCATAGC GTTAACCAAGGCGTGGTGACAAGCCTCTTTCAGGCGTCGGACACTTAAGAGCGTTGAAAAACG G135A CTCTTAGAGAATGAAAG C G C G C G A A C A C G C G C G C G C Gned sgRNA secondary structure, truncated sgRNA variants containing nucleotide deletions of varying length at several positions have been designed (Table 16). They were tested for DNA editing activity in HEK293T cells with Cas-beta2 M67(Q572R / F607S) (SEQ ID NO:189). One variant (trunc2 (SEQ ID NO:308) exhibits DNA editing activity comparable to wild type sgRNA (SEQ ID NO:204) as determined by T7 Endonuclease I assay (Figure 26). Table 16 SEQ ID sgRNA sgRNA sequenceT T308 trunc2 GTTCCGTCTCGACTATGCCGTACCACTAGACCGAGCCTACACGGCACGCGGTCAT AGCGTTAACCAAGGCGTGGTGACAAGCCTCTTTCAGGCGTCGGACACTTAAGGA AACTTAGGGAATGAAAGGCCTTCAGAAGAGGGTGCATTTTC C T C G A T G TEXAMPLE 12

[0410] The following example demonstrates the use of an engineered Cas-beta2 variant disclosed herein to introduce a single nucleotide change via HDR.Cas-beta2 M67(Q572R / F607S) was also evaluated for HDR efficiency using a single-stranded oligodeoxynucleotide (ssODN) DNA template. HEK293T cells were electroporated with nuclease–sgRNA RNP complex, targeting the AAVS1 gene, along with varying amounts of ssODN template encoding a 1 nt substitution which introduced an Alw26I restriction site, a 51 nt left homology arm (LHA) and 50 nt right homology arm (RHA) (SEQ ID NO:317). HEK293T cells were subcultured 2-3 days before the electroporation in Dulbecco‘s modified Eagle's medium (DMEM) (Thermo Fisher Scientific) supplemented with 10 % Fetal Bovine Serum (FBS) (Thermo Fisher Scientific) and 1 % Streptomycin / Penicillin (Thermo Fisher Scientific) and incubated in humidified incubator with 5 % CO2 at 37 °C.

[0411] Electroporation was performed using 4D-Nucleofector System (Lonza) and SF Cell Line 4D-Nucleofector X Kit S (Lonza) according to manufacturer’s protocol for HEK293 cells. 126 pmol of purified Cas-beta2 M67(Q572R / F607S) nuclease was mixed with 160 pmol of synthetic sgRNA (SEQ ID NO:282). The mixtures were incubated for 15-20 min at roomtemperature. Meanwhile, HEK293T cells were washed with 1X DPBS (Thermo Fisher Scientific) and harvested. 2 x 105cells / sample were resuspended in 10.5 μl / sample supplemented Nucleofector solution. 9.6 μl of prepared ribonucleoprotein complex, 1.6, 3.2, 16, 25, 37.5, 50, 62.5, and 100 pmol of ssODN HDR template and 10.5 μl of prepared HEK293T cells were mixed (final volume 25 μl) and entire volume of mixture was transferred into electroporation cuvette. Cells were electroporated using the program CM-130. After the electroporation, 75 μl of prewarmed Dulbecco‘s modified Eagle's medium (DMEM) supplemented with 10 % Fetal Bovine Serum (FBS) (without antibiotics) was added to the cells in the cuvette and 25 μl of cells were transferred into 96 well plate filled with 175 μl prewarmed media without antibiotics. Cells were incubated at 37ºC with 5% CO2 for 2 days post electroporation before genomic DNA extraction. The cells were washed twice with 200 μl pre- warmed to 37ºC 1X DPBS (Thermo Fisher Scientific) and resuspended in 25 μl cell lysis solution containing 50 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, pH 7.6 (Sigma Aldrich) and 0.2 mg / ml Proteinase K (New England Biolabs). Resuspended cells were incubated at 55ºC for 1 hour and 95ºC for 15 minutes.

[0412] Genomic region surrounding HDR-mediated nucleotide substitution, which introduces new Alw26I restriction site, was amplified by PCR using Q5 Hot Start High-Fidelity 2X Master Mix (New England Biolabs) according to the manufacturer’s instructions. The reaction was set up using 1μl of the cell lysate and 0.5 μM of each primer (SEQ ID NOs:240-241) in a final reaction volume of 25 μl. The PCR products were purified using Monarch PCR & DNA Cleanup Kit (New England Biolabs), theirquality and quantity checked using spectrophotometry (NanoPhotometer® NP80, IMPLEN) and fluorimetry (Qubit 1X dsDNA HS Assay, Qubit 1X dsDNA BR Assay and Qubit 4, Thermo Fisher Scientific), respectively. 200 ng of the product was used in a restriction digestion reaction with Alw26I (Thermo Fisher Scientific) according to the manufacturer’s instructions. 10 ng of restriction products were analyzed by 4150 TapeStation System (Agilent) using D1000 ScreenTape with D1000 Reagents (Agilent). HDR efficiency was defined and calculated as the fraction of the integrated area of the restriction endonuclease cleavage product peaks to the sum integrated area of cleaved and non-cleaved DNA peaks. HDR efficiency results are shown in Figure 27.

Claims

CLAIMS What is claimed:

1. An engineered Cas12l polypeptide comprising an amino acid amino acid sequence having at least 90% sequence identity to any one of SEQ ID Nos.1, 129, 130, or 131, wherein the Cas12l polypeptide comprises one or more altered amino acids corresponding to the residues at positions 615, 17, 21, 68, 131, 142, 153, 244, 253, 297, 301, 342, 452, 572, 576, 607, 762, 766, or 771 of SEQ ID NO:

1.

2. The engineered Cas12l polypeptide of claim 1, wherein the Cas12l polypeptide has endonuclease activity.

3. The engineered Cas12l polypeptide of claim 1, wherein the Cas12l polypeptide is a deactivated Cas (dCas) or a dCas complexed to a heterologous protein domain that has base editing activity, nucleoside deaminase activity, transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

4. The engineered Cas12l polypeptide of any one of claims 1-3, wherein each of the one or more altered amino acids comprises (i) a substitution at each altered amino acid corresponding to position 615, 17, 21, 68, 131, 142, 153, 244, 253, 297, 301, 342, 452, 572, 576, 607, 762, 766, or 771 of SEQ ID NO:1 or (ii) a substitution with a positively charged amino acid at each altered amino acid corresponding to position 615, 17, 21, 131, 142, 244, 253, 297, 301, 342, 762, 766, or 771 of SEQ ID NO:

1.

5. The engineered Cas12l polypeptide of any one of claims 1-4 comprising (i) one or more of the altered amino acids in one or more of SEQ ID NOs:32-73 or 180-201, (ii) any one of SEQ ID NOs:32-73, (iii) any one of SEQ ID NOs:180-201, or (iv) SEQ ID NO:

189.

6. A synthetic composition comprising: (a) the engineered Cas12l polypeptide of any one of claims 1-5; and(b) at least one guide polynucleotide comprising a region of complementarity to a target polynucleotide, wherein the Cas12l polypeptide forms a complex with the at least one guide polynucleotide, and wherein the complex binds to the target polynucleotide.

7. The synthetic composition of claim 6, wherein the composition further comprises the target polynucleotide.

8. The synthetic composition of claim 6 or 7, wherein the target polynucleotide comprises target sequence from a mammal, fungus, plant, bacteria, protozoa, or archaebacteria cell.

9. The synthetic composition of claim 8, wherein the target polynucleotide comprises target sequence from a mammalian cell or a human cell.

10. A variant Cas12l single guide RNA (sgRNA) capable of directing Cas12l polypeptide to a target DNA, wherein the sgRNA variant (i) comprises an altered stem-loop hairpin relative to an unaltered Cas12l single guide, wherein the stem-loop hairpin is the first hairpin closest to 5’ terminus of the unaltered sgRNA or (ii) is missing one or more nucleotides of the stem-loop hairpin relative to the unaltered Cas12l single guide.

11. The variant Cas12l sgRNA of claim 10, wherein the sgRNA variant is (i) missing three or more nucleotides of the stem-loop hairpin relative to the unaltered Cas12l single guide, (ii) is missing three to six nucleotides of the stem-loop hairpin relative to the unaltered Cas12l single guide, or (iii) includes additional hairpin sequence inserted at the stem-loop hairpin relative to the unaltered Cas12l single guide.

12. A Cas12l single guide RNA (sgRNA) variant comprising any one of SEQ ID NOs:74- 121, 123-128, 202-216, 283, 284, and 287-316.

13. A polynucleotide or construct comprising sequence encoding the synthetic Cas12l polypeptide of any one of claims 1-5.

14. The polynucleotide or construct of claim 13, further comprising sequence encoding the variant Cas12l sgRNA of claim 11 or 12.

15. A polynucleotide or construct comprising sequence encoding the variant Cas12l sgRNA of any one of claim 11 or 12.

16. A method of generating an engineered Cas12l, said method comprising producing the engineered Cas12l polypeptide of any one of claims 1-5, or expressing the construct of claim 13 or 14.

17. The method of claim 16, said method comprising (a) providing sequence of an unaltered Cas12 polypeptide; and (b) altering one or more amino acids in the Cas12 polypeptide corresponding to one or more of the residues at position 615, 17, 21, 68, 131, 142, 153, 244, 253, 297, 301, 342, 452, 572, 576, 607, 762, 766, or 771 of SEQ ID NO:1, thereby producing the engineered Cas12l polypeptide.

18. The method of claim 24, wherein (a) the unaltered Cas12 polypeptide is an unaltered Cas12l polypeptide and the method further comprises (c) expressing the altered coding sequence thereby producing the engineered Cas12l polypeptide.

19. The method of claim 18 or 19, said method further comprising producing an alignment of the amino acid sequence of the unaltered Cas12l polypeptide and SEQ ID NO:1 and using the alignment to determine the one or more amino acids in the unaltered Cas12l polypeptide corresponding to one or more of the residues at position 17, 21, 131, 142, 244, 253, 297, 301, 342, 615, 762, 766, or 771 of SEQ ID NO:

1.

20. A method of generating a variant single guide RNA (sgRNA), said method comprising synthesizing the variant Cas12l sgRNA of any one of claims 10-12 or expressing the construct of claims 13 or 14.

21. The method of claim 20, said method comprising(a) providing the sequence of an unaltered Cas12l sgRNA sequence comprising a stem-loop hairpin, wherein the stem-loop hairpin is the first hairpin closest to 5’ terminus of the sgRNA, and (b) synthesizing an altered Cas12l sgRNA which comprises an altered stem-loop hairpin relative to the unaltered Cas12l single guide, thereby generating the variant Cas12l sgRNA of any one of claims 10-12.

22. The method of claim 21, wherein the method comprises (a’) providing coding sequence encoding the unaltered Cas12l sgRNA comprising a stem-loop hairpin; (b’) altering the coding sequence encoding the comprising a stem-loop hairpin; and (c) expressing the altered coding sequence thereby producing the variant Cas12l sgRNA.

23. A method of editing a target polynucleotide in a cell, the method comprising: (a) providing to the cell the engineered Cas12l polypeptide of any one of claims 1-5 or expressing the construct of any one of claims 13-15 in the cell. (b) providing the cell with at least one guide polynucleotide comprising a region of complementarity to the target polypeptide, wherein the engineered Cas12l polypeptide forms a complex with the guide polynucleotide and the complex binds the target polynucleotide; and (c) introducing at least one nucleotide modification in the target polynucleotide via the complex.

24. The method of claim 23, further comprising providing the cell with a donor DNA molecular or a polynucleotide modification template.

25. The method of claim 23 or 24, wherein the cell is derived or obtained from (i) a mammal, fungus, plant, bacteria, protozoa, or archaebacteria or (ii) a human.

26. The method of any one of claims 23-25, wherein the engineered Cas12l polypeptide has endonuclease activity.

27. The method of any one of claims 23-25, wherein the engineered Cas12l polypeptide is a deactivated Cas (dCas), a dCas complexed to a heterologous protein domain that has base editing activity, nucleoside deaminase activity, transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity 28. The method of any one of claims 23-25, wherein the at least one guide polynucleotide comprises a plurality of guide polynucleotides and the engineered Cas12l polypeptide is a dCas complexed to a base editor or deaminase.

29. A method of editing a target polynucleotide in a cell, the method comprising: (a) providing to the cell the variant Cas12l sgRNA of any one of claims 14-19 or expressing the construct of claim 22 in the cell; (b) providing the cell with at least one guide polynucleotide comprising a region of complementarity to the target polypeptide, wherein the engineered Cas12l polypeptide forms a complex with the guide polynucleotide and the complex binds the target polynucleotide; and (c) introducing at least one nucleotide modification in the target polynucleotide via the complex.

30. The method of claim 29, further comprising providing the cell with a donor DNA molecular or a polynucleotide modification template.

31. The method of claim 29 or 30, wherein the cell is derived or obtained from (i) a mammal, fungus, plant, bacteria, protozoa, or archaebacteria or (ii) a human.

32. The method of any one of claims 29-31, wherein the engineered Cas12l polypeptide has endonuclease activity.

33. The method of any one of claims 29-31, wherein the engineered Cas12l polypeptide is a deactivated Cas (dCas), a dCas complexed to a heterologous protein domain that has base editing activity, nucleoside deaminase activity, transposase activity, integrase activity, methylase activity, demethylase activity, transcription activation activity,transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, or nucleic acid binding activity.

34. The method of any one of claims 29-31, wherein the at least one guide polynucleotide comprises a plurality of guide polynucleotides and the engineered Cas12l polypeptide is a dCas complexed to a base editor or deaminase.

35. A method for selectively enriching for one or more Cas polypeptide-guide RNA ribonucleoproteins (RNPs) having double stranded polynucleotide cleavage activity, the method comprising (a) combining in a reaction volume: Cas endonucleases, single guide RNAs (sgRNAs) that each comprises a region of complementarity (ROC) to a target sequence, and double stranded target polynucleotides that comprise the target sequence and a cleavage site, wherein each target polynucleotide comprises a first tag at the terminus that is more proximal to the cleavage site relative to the ROC target site and a different second tag at the polynucleotide terminus that is more proximal to the ROC target site than the cleavage site and wherein (i) the Cas endonucleases comprise a Cas endonuclease or variants thereof, (ii) the sgRNAs comprise an sgRNA and variants thereof, or both (i) and (ii); (b) assembling RNPs comprising the Cas endonucleases, sgRNAs, and target polynucleotides and allowing for Cas cleavage of the target polynucleotide; (c) adding to the reaction volume a first substrate that preferentially binds to the terminus comprising the first tag; (d) separating the first substrate, including materials bound to the first substrate from the remaining reaction volume, wherein the materials bound to the first substrate may include target polynucleotides that did not form RNP, RNPs bound to intact target polynucleotides, and cleaved portions of target polynucleotides; (e) adding to the remaining reaction volume a second substrate that preferentially binds to the terminus comprising the second tag; and (f) enriching for the second substrate and materials bound to the second substrate, thereby enriching for RNPs comprising cleaved target polynucleotides.

36. The method of claim 35, wherein the method selectively enriches for sgRNA variants that provide improved double stranded polynucleotide cleavage activity relative to an original sgRNA, wherein: (a’) the Cas endonucleases comprise one kind of Cas endonuclease and the sgRNAs comprise an original sgRNA; and (f’) enriching for the second substrate and materials bound to the second substrate, thereby enriches for RNPs comprising sgRNA variants that provide improved double stranded polynucleotide cleavage activity.

37. The method of claim 35, wherein the method selectively enriches for Cas endonuclease variants that provide improved double stranded polynucleotide cleavage activity relative to an original Cas endonuclease, wherein: (a’’) the Cas endonucleases comprise an original Cas endonuclease and variants thereof, the sgRNAs comprise predominantly one kind of sgRNA; and (f’’) enriching for the second substrate and materials bound to the second substrate, thereby enriches for RNPs comprising Cas endonuclease variants that provide improved double stranded polynucleotide cleavage activity.

38. The method of any one of claims 35-37, wherein the first and second tag are each selected from the group consisting of biotin, azide, folate, polyhistidine, FLAG, myc, maltose binding protein, metal chelating peptides, histidine-tryptophan modules, and a protein A domain.

39. The method of any one of claims 35-37, wherein the remaining reaction volume in step (e) is treated with a reagent that modifies the second tag to introduce a third tag and the second substrate in step (e) preferentially binds to the second tag via the third tag.

40. The method of claim 39, wherein the first tag is biotin, the second tag is azide, the second tag is modified to introduce a biotin, and the second substrate is the same as the first substrate.

1. The method of any on of claim 35-40, wherein the substrate comprises (i) magnetic beads or particles, (ii) beads or particles of agarose, cellulose, glass, polystyrene, or polyacrylamide, (iii) a microtiter plate, (iv) a tube, or (v) a membrane.

Citation Information

Patent Citations

  • Crispr-cas effector polypeptides and methods of use thereof

    US20210254038A1

  • Crispr-cas effector polypeptides and methods of use thereof

    US20230028178A1

  • Novel crispr-CAS systems for genome editing

    US20230084762A1