Deamination polyproteins and methods of use thereof
The engineered dimeric deamination polyprotein addresses the inefficiencies of existing genome-editing technologies by introducing targeted single nucleotide polymorphisms in eukaryotic cells, enabling precise modulation of gene expression and disruption.
Patent Information
- Application Number
- PCT/US2025/033535
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-27
- Filing Date
- 2025-06-13
- Publication Date
- 2025-12-18
AI Technical Summary
Existing genome-editing technologies, such as CRISPR systems, require redesign for each target site, making them costly and time-consuming, and there is a need for novel effectors and systems that can efficiently edit endogenous and heterologous polynucleotides in eukaryotes like animals and plants.
Development of an engineered, non-natural dimeric deamination polyprotein comprising nucleotide deaminases, Cas-alpha 10 polypeptides, and uracil glycosylase inhibitors that form dimers upon binding to guide polynucleotides, introducing single nucleotide polymorphisms in target and non-target strands of dsDNA to modulate gene expression.
The dimeric deamination polyprotein efficiently introduces targeted single nucleotide polymorphisms and mutations in eukaryotic cells, including plants, allowing precise modulation of gene expression and disruption of gene function.
Smart Images

Figure US2025033535_18122025_PF_FP_ABST
Abstract
Description
DEAMINATION POLYPROTEINSAND METHODS OF USE THEREOFCROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to US Provisional Application No. 63 / 660,013, filed June 14, 2024, and US Provisional Application No. 63 / 700,108, filed September 27, 2024, which are incorporated by reference herein in their entireties.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0002] The official copy of the sequence listing is submitted electronically as an XML formatted sequence listing with a file named “212381-WO-SEC-l_Sequence_Listing_ST26” created on May 29, 2025 having a size of 138 kilobytes, and is filed concurrently with the specification. The sequence listing comprised in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety.BACKGROUND
[0003] Recombinant DNA technology has made it possible to insert DNA sequences at targeted genomic locations and / or modify specific endogenous chromosomal sequences. Site-specific integration techniques, which employ site-specific recombination systems, as well as other types of recombination technologies, have been used to generate targeted insertions of genes of interest in a variety of organisms. Genome-editing techniques such as designer zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or homing meganucleases, are available for producing targeted genome perturbations, but these systems employ designed nucleases that need to be redesigned for each target site, which renders them costly and timeconsuming to prepare.
[0004] Newer technologies utilizing archaeal or bacterial adaptive immunity systems have been identified, called CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats), which comprise different domains of effector proteins that encompass a variety of activities (DNA recognition, binding, and optionally cleavage).
[0005] Despite the identification and characterization of some of these systems, there remains a need for identifying novel effectors and systems, as well as demonstrating activity in eukaryotes,particularly animals and plants, to effect editing of endogenous and previously-introduced heterologous polynucleotides.SUMMARY
[0006] In a first aspect, the disclosure provides an engineered, non-natural dimeric, multi-domain deamination polyprotein, the deamination polyprotein comprising: (i) a first nucleotide deaminase; (ii) a second nucleotide deaminase; (iii) a first Cas-alpha 10 polypeptide; (iv) a second Cas-alpha 10 polypeptide; (v) a first uracil glycosylase inhibitor; and (vi) a second uracil glycosylase inhibitor.
[0007] In an example of this first aspect, the first nucleotide deaminase is a cytosine deaminase and the second nucleotide deaminase is a cytosine deaminase, and the deamination polyprotein deaminates a cytosine polynucleotide on a target or a non-target dsDNA target, resulting in the transition of a cytosine-guanine base pair to a thymine-adenine base pair or a guanine-cytosine base pair to an adenine-thymine base pair.
[0008] In an example of this first aspect, the first and / or the second nucleotide deaminase can have an amino acid sequence having at least 90% sequence identity to one of SEQ ID Nos: 1-4. The first and / or the second uracil glycosylase inhibitor can have an amino acid sequence having at least 90% sequence identity to one of SEQ ID Nos: 7-10.
[0009] In an example of this first aspect, the first Cas-alpha 10 polypeptide is a first dead Cas- alpha 10 (dCas-alpha 10) polypeptide and the second Cas-alpha 10 polypeptide is a second dead Cas-alpha 10 polypeptide. The first and / or second dCas-alpha 10 polypeptide can have an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 5 or SEQ ID NO: 6.
[0010] The deamination polyprotein of this first aspect can comprise two monomers, the first monomer comprising the first nucleotide deaminase, the first dCas-alpha 10 polypeptide, and the first uracil glycosylase inhibitor and the second monomer comprising the second nucleotide deaminase, the second dCas-alpha 10 polypeptide, and the second uracil glycosylase inhibitor, wherein the first and second monomers dimerize upon binding to a guide polynucleotide. The first nucleotide deaminase and the second nucleotide deaminase can be the same and the polyprotein is a homodimer. Alternatively, the first nucleotide deaminase and the second nucleotide deaminase can be different and the polyprotein is a heterodimer.
[0011] In an example of this first aspect, the first monomer comprises the C-terminus of the first nucleotide deaminase linked to the N-terminus of the first dCas-alpha 10 polypeptide, and the C- terminus of the first dCas-alpha 10 polypeptide linked to the N-terminus of the first uracil glycosylase inhibitor; and the second monomer comprises the C-terminus of the second nucleotide deaminase linked to the N-terminus of the second dCas-alpha 10 polypeptide, and the C-terminus of the second dCas-alpha 10 polypeptide linked to the N-terminus of the second uracil glycosylase inhibitor.
[0012] In another example of this first aspect, the first Cas-alpha 10 polypeptide is a first Cas- alpha 10 polypeptide and the second Cas-alpha 10 polypeptide is a second Cas-alpha lOCas-alpha 10 polypeptide. The first and / or the second Cas-alpha 10 polypeptide can have an amino acid sequence having at least 90% sequence identity to SEQ ID NOs: 79 and 90-92. The deamination polyprotein can comprise two monomers, the first monomer comprising the first nucleotide deaminase, the first Cas-alpha 10 polypeptide, and the first uracil glycosylase inhibitor and the second monomer comprising the second nucleotide deaminase, the second Cas-alpha 10 polypeptide, and the second uracil glycosylase inhibitor, wherein the first and second monomers dimerize upon binding to a guide polynucleotide. In an example, first monomer comprises the C- terminus of the first nucleotide deaminase linked to the N-terminus of the first Cas-alpha 10 polypeptide, and the C-terminus of the first Cas-alpha 10 polypeptide linked to the N-terminus of the first uracil glycosylase inhibitor; and the second monomer comprises the C-terminus of the second nucleotide deaminase linked to the N-terminus of the second Cas-alpha 10 polypeptide, and the C-terminus of the second Cas-alpha 10 polypeptide linked to the N-terminus of the second uracil glycosylase inhibitor.
[0013] In an example of this first aspect, the first nucleotide deaminase is an adenine deaminase and the second nucleotide deaminase is an adenine deaminase. In another example, the first nucleotide deaminase is an adenine deaminase and the second nucleotide deaminase is a cytosine deaminase. In yet another example, the first nucleotide deaminase is a cytosine deaminase and the second nucleotide deaminase is a cytosine deaminase.
[0014] In a second aspect, the disclosure provides a composition comprising a deamination polyprotein according to any of the examples described above and herein and a guide polynucleotide comprising a sequence that shares homology with a dsDNA target site in a cell, wherein the deamination polyprotein introduces single nucleotide polymorphism in both a targetstrand and a non-target strand of a dsDNA target site. In an example of this second aspect, a variable targeting domain or spacer of the guide polynucleotide is between 12-16 nucleotides in length.
[0015] In a third aspect, the disclosure provides a cell comprising a deamination polyprotein according to any of the examples described above and herein. The cell can be a eukaryotic cell, such as a plant cell.
[0016] In a fourth aspect, the disclosure provides a method of modifying a dsDNA target site in a cell, the method comprising: providing to the cell: c (ii) a guide polynucleotide comprising a sequence that shares homology with the dsDNA target site in the cell, wherein the dimeric deamination polyprotein and the guide polynucleotide form a complex that recognizes and binds to the dsDNA target site, and introducing single nucleotide polymorphisms in a target strand and a non-target strand of the dsDNA target site.
[0017] In a fifth aspect, the disclosure provides a multiplexed method of modifying a plurality of dsDNA target sites in a cell, the method comprising: providing to the cell: (i) a deamination polyprotein according to any of the examples described above and herein; and (ii) a plurality of guide polynucleotides, each guide polynucleotide comprising a sequence that shares homology with a dsDNA target site of the plurality of dsDNA target sites, wherein the dimeric deamination polyprotein and each guide polynucleotide form a complex that recognizes and binds to each dsDNA target site, and introducing single nucleotide polymorphisms in a target strand and a non- target strand of each dsDNA target site.
[0018] Examples of the fourth and / or fifth aspects, modifying a target site in a cell result in disruption of gene expression, increased or decreased gene expression, or novel variation in a polypeptide encoded by a gene. The methods of the fourth and / or fifth aspects can also result in: (a) introducing one or more premature stop codons within the open reading frame of a gene; (b) introducing an in-frame stop codon at an exon -intron splice site; (c) introducing a C-G polymorphism at a TATA box, initiator element (Inr), downstream promoter element (DPE), TFIIB recognition element (BRE), downstream core element (DCE), motif ten element (MTE), X core promoter element (XCPE), intron, 5’ untranslated region (UTR), and / or 3’ UTR); and / or (d) introducing into a gene a new codon encoding a novel amino acid.
[0019] In the methods of the fourth and / or fifth aspects the cell can be a eukaryotic cell. In an example, the cell is a plant cell, and method can further comprise obtaining progeny from the plant cell, wherein the progeny comprises at least one single nucleotide polymorphism.
[0020] In an example of the fourth and / or fifth aspects, introducing single nucleotide polymorphisms comprises: deaminating a cytosine polynucleotide on the target or the non-target dsDNA target, resulting in the transition of a cytosine-guanine base pair to a thymine-adenine base pair or a guanine-cytosine base pair to an adenine-thymine base pair.
[0021] In an example of the fourth and / or fifth aspects, the guide polynucleotide is a gRNA comprises a spacer of between 12-18 nucleotides, alternatively between 12-16 nucleotides.
[0022] In a sixth aspect, the disclosure provides a method of modulating gene expression in a cell, the method comprising: providing to the cell: (i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and a second nuclease-active Cas-alpha polypeptide (i.e., having double-strand break activity); and (ii) a first guide polynucleotide that shares homology with a first dsDNA target site in the cell, the first guide polynucleotide comprising a spacer that is 10-16 nucleotides in length; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the first dsDNA target site, wherein the dimeric polyprotein and the first guide polynucleotide form a first complex that recognizes a PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and optionally nicks the first dsDNA target site; wherein introducing the one or more single nucleotide polymorphisms (SNPs) modulates gene expression or alters a resulting polypeptide expressed by a gene in the cell by: (a) disrupting gene expression by introducing a premature stop codon within an open reading frame of the gene; (b) altering a polypeptide expressed from a gene by removing an exon-intron splice site; (c) increasing or decreasing gene expression by introducing the one or more SNPs in a gene expression regulatory element; and / or (d) altering a polypeptide expressed from a gene by introducing a codon encoding a novel amino acid into the gene.
[0023] In a seventh aspect, the disclosure provides a method of modulating gene expression in a cell, the method comprising: providing to the cell: (i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and a secondnuclease-active Cas-alpha polypeptide (i.e., having double-strand break activity); and (ii) a first guide polynucleotide that shares homology with a first dsDNA target site in the cell, the first guide polynucleotide comprising a spacer that is 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pairs with a non-target strand of the first dsDNA target site; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the first dsDNA target site, wherein the dimeric polyprotein and the first guide polynucleotide form a first complex that recognizes a PAM sequence on the first dsDNA target site, binds to the first dsDNA target site, and optionally nicks the first dsDNA target site; wherein introducing the one or more single nucleotide polymorphisms (SNPs) modulates gene expression or alters a resulting polypeptide expressed by a gene in the cell by: (a) disrupting gene expression by introducing a premature stop codon within an open reading frame of the gene; (b) altering a polypeptide expressed from a gene by removing an exon-intron splice site; (c) increasing or decreasing gene expression by introducing the one or more SNPs in a gene expression regulatory element; and / or (d) altering a polypeptide expressed from a gene by introducing a codon encoding a novel amino acid into the gene.
[0024] In an example of the sixth or seventh aspects, the method further comprises providing to the cell a second guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length. The method can further comprise: providing to the cell the second guide polynucleotide that shares homology with a second dsDNA target site in the cell; inducing a double-strand break in the second dsDNA target site, wherein the dimeric polyprotein and the second guide polynucleotide form a second complex that recognizes a PAM sequence on the second dsDNA target site, binds to the second dsDNA target site, and cleaves the second dsDNA target site; and introducing an insertion, deletion, inversion, or translocation mutation at the second dsDNA target site via non -homologous end joining or homology-directed repair. The method can further comprise providing to the cell a polynucleotide modification template or a donor DNA.
[0025] In an example of the sixth or seventh aspects, the gene expression regulatory element is a TATA box, an initiator element, a downstream promoter element, a TFIIB recognition element, a downstream core element, a motif ten element, an X core promoter element, an intron, a 5’ untranslated region (UTR), or 3’ UTR.
[0026] In an example of the sixth or seventh aspects, the first nucleotide deaminase is a first cytosine deaminase and the dimeric polyprotein further comprises a first uracil glycosylaseinhibitor; and the second nucleotide deaminase is a second cytosine deaminase and the dimeric polyprotein further comprises a second uracil glycosylase inhibitor. The dimeric polyprotein can comprise two monomers: the first monomer comprising the first cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and the first uracil glycosylase inhibitor; and the second monomer comprising the second cytosine deaminase, the second nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and the second uracil glycosylase inhibitor; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide. Optionally, in the foregoing examples of the sixth and seventh aspects, the method further comprises providing to the cell additional guide polynucleotides (e.g., a third guide polynucleotide and / or a fourth guide polynucleotide and / or more guide polynucleotides) each comprising a spacer.
[0027] In an example of the sixth or seventh aspects, the first nucleotide deaminase is a cytosine deaminase and the dimeric polyprotein further comprises a uracil glycosylase inhibitor; and the second nucleotide deaminase is an adenine deaminase. The dimeric polyprotein can comprise two monomers: the first monomer comprising the cytosine deaminase, the first Cas-alpha 10 polypeptide, and the uracil glycosylase inhibitor; and the second monomer comprising the adenine deaminase and the second Cas-alpha 10 polypeptide; wherein the first and second monomers dimerize upon binding to the first, second, third, fourth, and / or additional guide polynucleotide(s).
[0028] In an example of the sixth or seventh aspects, the first nucleotide deaminase is a first adenine deaminase and the second nucleotide deaminase is a second adenine deaminase. The dimeric polyprotein comprises two monomers: the first monomer comprising the first adenine deaminase and the first Cas-alpha 10 polypeptide; and the second monomer comprising the second adenine deaminase and the second Cas-alpha 10 polypeptide; wherein the first and second monomers dimerize upon binding to the first, second, third, fourth, and / or additional guide polynucleotides.
[0029] In another example of the sixth aspect, the spacer of the first guide polynucleotide is 10-12 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site and binds to the first dsDNA target site without nicking or cleaving the first dsDNA target site.
[0030] In yet another example of the sixth aspect, the spacer of the first guide polynucleotide is 12-16 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and nicks the first dsDNA target site.
[0031] In an eighth aspect, the disclosure provides a non-natural, multi-domain, dimeric polyprotein comprising: (i) a first nucleotide deaminase; (ii) a second nucleotide deaminase; (iii) a first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity); and (iv) a second nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity).
[0032] In an example of the eighth aspect, the first nucleotide deaminase is a first cytosine deaminase and the dimeric polyprotein further comprises a first uracil glycosylase inhibitor; and the second nucleotide deaminase is a second cytosine deaminase and the dimeric polyprotein further comprises a second uracil glycosylase inhibitor. The dimeric polyprotein can comprise two monomers: the first monomer comprising the first cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and the first uracil glycosylase inhibitor; and the second monomer comprising the second cytosine deaminase, the second nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and the second uracil glycosylase inhibitor; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
[0033] In an example of the eighth aspect, the first nucleotide deaminase is a cytosine deaminase and the dimeric polyprotein further comprises a uracil glycosylase inhibitor; and the second nucleotide deaminase is an adenine deaminase. The dimeric polyprotein can comprise two monomers: the first monomer comprising the cytosine deaminase, the first Cas-alpha 10 polypeptide, and the uracil glycosylase inhibitor; and the second monomer comprising the adenine deaminase and the second Cas-alpha 10 polypeptide; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
[0034] In an example of the eighth aspect, the first nucleotide deaminase is a first adenine deaminase and the second nucleotide deaminase is a second adenine deaminase. The dimeric polyprotein can comprise two monomers: the first monomer comprising the first adenine deaminase and the first Cas-alpha 10 polypeptide; and the second monomer comprising the second adenine deaminase and the second Cas-alpha 10 polypeptide; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
[0035] In a ninth aspect, the disclosure provides a composition comprising: (i) a dimeric polyprotein according to any example of the eighth aspect disclosed herein; and (ii) at least one guide polynucleotide that shares homology with a dsDNA target site in a cell, wherein the at least one guide polynucleotide comprises: (a) a guide polynucleotide comprising a spacer that is about17-23 nucleotides in length; (b) a guide polynucleotide comprising a spacer that is 12-16 nucleotides in length; (c) a guide polynucleotide comprising a spacer that is 10-12 nucleotides in length; (d) a guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site; or (e) any combination of (a) - (d), wherein the dimeric polyprotein is capable of introducing single nucleotide polymorphism resulting from deamination of a target strand and / or a non-target strand of a dsDNA target site and / or an insertion, deletion, inversion, or translocation mutation in the dsDNA target site.
[0036] In an example of the ninth aspect, the composition can comprise the guide polynucleotide of (a), the guide polynucleotide of (b), the guide polynucleotide of (c), and the guide polynucleotide of (d), wherein each guide polynucleotide shares homology with a unique dsDNA target site.
[0037] In a tenth aspect, the disclosure provides a cell comprising any dimeric polyprotein described herein. The cell can be a eukaryotic cell; such has a plant cell.
[0038] In an eleventh aspect, the disclosure provides a method of modifying a dsDNA target site in a cell, the method comprising: providing to the cell: (i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide (i.e., having double-strand break activity), and a second nuclease-active Cas-alpha polypeptide (i.e., having double-strand break activity); and (ii) one or more guide polynucleotides, each guide polynucleotide sharing homology with a unique dsDNA target site in the cell, wherein the one or more guide polynucleotides comprise: (a) a first guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length; (b) a second guide polynucleotide comprising a spacer that is 12-16 nucleotides in length; (c) a third guide polynucleotide comprising a spacer that is 10-12 nucleotides in length; or (d) a fourth guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the dsDNA target site, wherein the dimeric polyprotein and the guide polynucleotide form a complex that recognizes a PAM sequence of the dsDNA target site, binds to the dsDNA target site, and optionally nicks the first dsDNA target site and / or introducing an insertion, deletion, inversion, or translocation mutation at the dsDNA target site byinducing a double-strand break in the dsDNA target site, wherein the dimeric polyprotein and the guide polynucleotide form a complex that recognizes a PAM sequence on the dsDNA target site, binds to the dsDNA target site, and cleaves the dsDNA target site.
[0039] In a twelfth aspect, the disclosure provides a method of modulating gene expression in a cell, the method comprising: providing to the cell: (i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide, and a second nuclease-active Cas-alpha 10 polypeptide; and (ii) a first guide polynucleotide that shares homology with a first dsDNA target site in the cell; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the first dsDNA target site, wherein the dimeric polyprotein and the first guide polynucleotide form a first complex that recognizes a PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and optionally nicks the first dsDNA target site; wherein introducing the one or more single nucleotide polymorphisms (SNPs) modulates gene expression or alters a resulting polypeptide expressed by a gene in the cell by: disrupting gene expression by introducing a premature stop codon within an open reading frame of the gene; altering a polypeptide expressed from a gene by removing an exon-intron splice site; increasing or decreasing gene expression by introducing the one or more SNPs in a gene expression regulatory element; and / or altering a polypeptide expressed from a gene by introducing a codon encoding a novel amino acid into the gene.
[0040] In an example of the twelfth aspect, the first guide polynucleotide comprises a spacer that is 12-16 nucleotides in length, a spacer that is 10-12 nucleotides in length, or a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site.
[0041] In an example of the twelfth aspect, the method further comprises providing to the cell a second guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length.
[0042] In an example of the twelfth aspect, the method further comprises, providing to the cell the second guide polynucleotide that shares homology with a second dsDNA target site in the cell; inducing a double-strand break in the second dsDNA target site, wherein the dimeric polyprotein and the second guide polynucleotide form a second complex that recognizes a PAM sequence on the second dsDNA target site, binds to the second dsDNA target site, and cleaves the second dsDNA target site; and introducing an insertion, deletion, inversion, or translocation mutation atthe second dsDNA target site via non-homologous end joining or homology-directed repair. The method can further comprise providing to the cell a polynucleotide modification template or a donor DNA.
[0043] In an example of the twelfth aspect, the gene expression regulatory element is a TATAbox, an initiator element, a downstream promoter element, a TFIIB recognition element, a downstream core element, a motif ten element, an X core promoter element, an intron, a 5’ untranslated region (UTR), or 3’ UTR.
[0044] In an example of the twelfth aspect, the first nucleotide deaminase is a first cytosine deaminase and the dimeric polyprotein further comprises a first uracil glycosylase inhibitor; and the second nucleotide deaminase is a second cytosine deaminase and the dimeric polyprotein further comprises a second uracil glycosylase inhibitor.
[0045] In an example of the twelfth aspect, the dimeric polyprotein comprises two monomers: the first monomer comprising the first cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the first uracil glycosylase inhibitor; and the second monomer comprising the second cytosine deaminase, the second nuclease-active Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the second uracil glycosylase inhibitor; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
[0046] In an example of the twelfth aspect, the first nucleotide deaminase is a cytosine deaminase and the dimeric polyprotein further comprises a uracil glycosylase inhibitor; and the second nucleotide deaminase is an adenine deaminase.
[0047] In an example of the twelfth aspect, the dimeric polyprotein comprises two monomers: the first monomer comprising the cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the uracil glycosylase inhibitor; and the second monomer comprising the adenine deaminase and the second nuclease-active Cas-alpha 10 polypeptide, wherein the second nucleaseactive Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
[0048] In an example of the twelfth aspect, the first nucleotide deaminase is a first adenine deaminase and the second nucleotide deaminase is a second adenine deaminase.
[0049] In an example of the twelfth aspect, the dimeric polyprotein comprises two monomers: the first monomer comprising the first adenine deaminase and the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; and the second monomer comprising the second adenine deaminase and the second nuclease-active Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
[0050] In an example of the twelfth aspect, the spacer of the first guide polynucleotide is 10-12 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site and binds to the first dsDNA target site without nicking or cleaving the first dsDNA target site.
[0051] In an example of the twelfth aspect, the spacer of the first guide polynucleotide is 12-16 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and nicks the first dsDNA target site.BRIEF DESCRIPTION OF THE DRAWINGS AND SEQUENCE LISTING
[0052] FIG. 1A illustrates a homodimeric, nuclease-inactivated Cas-alpha 10 polypeptide, multidomain deamination polyprotein complexed with a guide RNA (gRNA), the polyprotein comprising cytosine deaminases. The variable targeting domain or spacer of the gRNA is indicated with long dashes. CD = cytosine deaminase; UGI = uracil glycosylase inhibitor.
[0053] FIG. IB illustrates a homodimeric, nuclease-inactivated Cas-alpha 10 polypeptide, multidomain deamination polyprotein complexed with a gRNA, the polyprotein comprising adenine deaminases. The variable targeting domain or spacer of the gRNA is indicated with long dashes. AD = adenine deaminase.
[0054] FIG. 1C illustrates a heterodimeric, nuclease-inactivated Cas-alpha 10 polypeptide, multidomain deamination polyprotein complexed with a gRNA, the polyprotein comprising an adenine deaminase and a cytosine deaminase. The variable targeting domain or spacer of the gRNA is indicated with long dashes. CD = cytosine deaminase; UGI = uracil glycosylase inhibitor; AD = adenine deaminase.
[0055] FIG. 2A is a graph showing the frequency of cytosine to thymine transition polymorphisms as a function of a homodimeric, multi-domain deamination polyprotein having a rAPOBECl, a hAPOBEC3A, a StsDaOl, or a 1APOBEC1 CD domain.
[0056] FIG. 2B shows the eight dsDNA target sites of FIG. 2A.
[0057] FIG. 2C is a graph showing the frequency of cytosine to thymine transition polymorphisms as a function of a homodimeric, multi-domain deamination polyproteins having a sAPOBEC3G CD domain fused to the amino terminal of Cas-alpha 10 dALT114.
[0058] FIG. 2D shows the four dsDNA target sites of FIG. 2C.
[0059] FIG. 3A is a graph showing the percentage of single nucleotide polymorphisms (SNPs) recovered in human cells at 10 different dsDNA target sites when using a hAPOBEC3 A Cas-alpha 10 dALT114 polyprotein-gRNA complex.
[0060] FIG. 3B is a graph showing the percentage of insertion / deletion mutations (indels) recovered in human cells at 10 different dsDNA target sites when using a hAPOBEC3 A Cas-alpha 10 dALT114 polyprotein-gRNA complex.
[0061] FIG. 3C is a graph showing the percentage of SNPs recovered in human cells at four different dsDNA target sites when using a hAPOBEC3A or sAPOBEC3G Cas-alpha 10 dALT135- F54 polyprotein-gRNA complex.
[0062] FIG. 3D is a graph showing the viability of human cells after delivery of plasmid DNA expressing a hAPOBEC3A or sAPOBEC3G Cas-alpha 10 dALT135-F54 polyprotein-gRNA complex.
[0063] FIGS. 4A- 4H show frequencies of next-generation sequence (NGS) reads and observed targeted modifications from hAPOBEC3A Cas-alpha 10 dALT114 polyprotein editing experiments in human cells at the DNMT2 (FIG. 4A), FANCFT1 (FIG. 4B), VEGFA2 (FIG. 4C), WTAP4n (FIG. 4D), WTAPT3 (FIG. 4E), DNMT1 (FIG. 4F), HBBln (FIG. 4G), and WTAPT5 (FIG. 4H) target sites. FIGS. 41 - 4L show frequences of NGS reads and observed target modifications from hAPOBEC3A Cas-alpha 10 dALT135-F54 polyprotein or sAPOBEC3G Cas- alpha 10 dALT135-F54 polyprotein editing experiments in human cells at the WTAPT3 (FIG. 41), VEGFA7 (FIG. 4J), RUNX2n (FIG. 4K), and FANCFT5 (FIG. 4L) target sites. Transition polynucleotide polymorphisms are in lowercase and underlined.
[0064] FIG. 5A is a diagram illustrating a mechanism for cytosine (C) to thymine (T) transition mutations by a nuclease-inactive homodimeric, multi-domain deamination polyprotein-gRNA complex. Cytosines on the non-target strand of a dsDNA target are deaminated to uracil (U) and used as a template during DNA replication or repair resulting in C-G to T-A single nucleotide polymorphisms.
[0065] FIG. 5B is a diagram illustrating a mechanism for guanine (G) to adenine (A) transition mutations by a nuclease-inactive homodimeric, multi-domain deamination polyprotein-gRNA complex. Cytosines on the target strand of a dsDNA target are deaminated to uracil (U) and used as a template during DNA replication or repair resulting in G-C to A-T single nucleotide polymorphisms.
[0066] FIG. 6 is a graph showing average cytosine (C) to thymine (T) and guanine (G) to adenine (A) editing relative to the Cas-alpha 10 target site.
[0067] FIG. 7 is a graph showing the average frequency of cytosine (C) to thymine (T) and guanine (G) to adenine (A) single nucleotide polymorphisms (SNPs) recovered using hAPOBEC3A Cas-alpha 10 dALT114 and hAPOBEC3A Cas-alpha 10 dALT135-F54 polyproteins.
[0068] FIG. 8A illustrates a homodimeric, nuclease-active Cas-alpha 10 polypeptide (i.e., having endonuclease or double-strand break activity), multi-domain deamination polyprotein complexed with gRNAs, the polyprotein comprising cytosine deaminases. Each gRNAhas a different variable targeting domain or spacer: (A) a gRNA having a variable targeting domain or spacer of around 20 nucleotides, (B) a gRNA having a variable targeting domain or spacer of 12-16 nucleotides, (C) a gRNA having 10-12 nucleotides, and (D) a gRNA having additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand. The variable targeting domain or spacer of the gRNA is indicated with long dashes. CD = cytosine deaminase; UGI = uracil glycosylase inhibitor.
[0069] FIG. 8B illustrates a homodimeric, nuclease-active Cas-alpha 10 polypeptide (i.e., having endonuclease or double-strand break activity), multi-domain deamination polyprotein complexed with gRNAs, the polyprotein comprising adenine deaminases. Each gRNAhas a different variable targeting domain or spacer: (A) a gRNA having a variable targeting domain or spacer of around 20 nucleotides, (B) a gRNA having a variable targeting domain or spacer of 12-16 nucleotides, (C) a gRNA having 10-12 nucleotides, and (D) a gRNA having additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand. The variable targeting domain or spacer of the gRNA is indicated with long dashes. AD = adenine deaminase.
[0070] FIG. 8C illustrates a heterodimeric, nuclease-activate Cas-alpha 10 polypeptide (i.e., having endonuclease or double-strand break activity), multi-domain deamination polyprotein complexed with gRNAs, the polyprotein comprising an adenine deaminase and a cytosinedeaminase. Each gRNA has a different variable targeting domain or spacer: (A) a gRNA having a variable targeting domain or spacer of around 20 nucleotides, (B) a gRNA having a variable targeting domain or spacer of 12-16 nucleotides, (C) a gRNA having 10-12 nucleotides, and (D) a gRNA having additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand. The variable targeting domain or spacer of the gRNA is indicated with long dashes. CD = cytosine deaminase; UGI = uracil glycosylase inhibitor; AD = adenine deaminase.
[0071] FIG. 9A depicts the effect of reducing the length of the variable targeting domain or spacer in a gRNA from 20 to 12-16 nucleotides when complexed with a nuclease-active polyprotein. Lightning bolts indicate the positions of DNA phosphodiester backbone cleavage.
[0072] FIG. 9B depicts the effect of reducing the length of the variable targeting domain or spacer in a gRNA from 20 to 10-12 nucleotides when complexed with a nuclease-active polyprotein. Lightning bolts indicate the positions of DNA phosphodiester backbone cleavage.
[0073] FIG. 9C depicts the effect of adding nucleotides at the 3’ end of the variable targeting domain or spacer in a gRNA capable of base pairing with the non-target strand of a double- stranded DNA target site when complexed with a nuclease-active polyprotein. Lightning bolts indicate the positions of DNA phosphodiester backbone cleavage.
[0074] FIG. 9D depicts another effect of adding nucleotides at the 3’ end of the variable targeting domain or spacer in a gRNA capable of base pairing with the non-target strand of a double-stranded DNA target site when complexed with a nuclease-active polyprotein. Lightning bolts indicate the positions of DNA phosphodiester backbone cleavage.
[0075] FIG. 10A is a diagram illustrating a mechanism by which the introduction of single nucleotide polymorphisms (SNPs) can be biased towards guanine (G) to adenine (A) transition mutations by nicking the non-target strand of a dsDNA target site and using the target strand as a DNA repair template. Lightning bolt indicates the position of DNA nick.
[0076] FIG. 10B is a diagram illustrating a mechanism by which the introduction of SNPS can be biased towards cytosine (C) to thymine (T) transition mutations by nicking the target strand of a dsDNA target site and using the non-target strand as a DNA repair template. Lightning bolt indicates the position of DNA nick.
[0077] FIGS. 11A - HE are graphs demonstrating the effect of gRNA variable targeting domain or spacer length on non-target and target strand cleavage for the TTR-sg7 dsDNA target site.
[0078] FIGS. 12A - 12E are graphs demonstrating the effect of gRNA variable targeting domain or spacer length on non-target and target strand cleavage for the TTR-sg8 dsDNA target site.
[0079] FIGS. 13A- 13E are graphs demonstrating the effect of gRNA variable targeting domain or spacer length on non-target and target strand cleavage for the RUNX1 dsDNA target site.
[0080] FIGS. 14A- 14E are graphs demonstrating the effect of gRNA variable targeting domain or spacer length on non-target and target strand cleavage for the RUNXl-pn2 dsDNA target site.
[0081] FIG. 15 is a graph showing double-stranded (ds) DNA target binding of nuclease-active Cas-alpha 10 ALT114 in a Saccharomyces cerevisiae cell as the length of the variable targeting domain (aka spacer) in the guide RNA (gRNA) is shortened.
[0082] FIG. 16A is a diagram illustrating a mechanism for cytosine (C) to thymine (T) and adenine (A) to guanine (G) transition mutations by a heterodimeric, multi-domain deamination polyprotein-gRNA complex. Cytosines and adenines on the non-target strand of a dsDNA target are deaminated to uracil (U) and inosine (I), respectively, and used as a template during DNA replication or repair resulting in C-G to T-A and A-T to G-C single nucleotide polymorphisms.
[0083] FIG. 16B is a diagram illustrating a mechanism for guanine (G) to adenine (A) and thymine (T) to cytosine (C) transition mutations by a heterodimeric, multi-domain deamination polyprotein-gRNA complex. Cytosines and adenines on the target strand of a dsDNA target are deaminated to uracil (U) and inosine (I), respectively, and used as a template during DNA replication or repair resulting in G-C to A-T and T-A to C-G single nucleotide polymorphisms.
[0084] FIG. 16C is a diagram illustrating a mechanism by which the introduction of single nucleotide polymorphisms (SNPs) can be biased towards guanine (G) to adenine (A) and thymine (T) to cytosine (C) transition mutations by nicking the non-target strand of a dsDNA target site.
[0085] FIGS. 17A - 17F are diagrams illustrating applications for the introduction of polymorphisms using the deamination polyproteins described herein. The size of the arrow towards the beginning of the gene represents the relative strength of expression such that a longer arrow represents an increase in gene transcription (FIG. 17D) and a shorter one a decrease in gene transcription (FIG. 17E). Transition polynucleotide polymorphisms are in lowercase and underlined. UTR = untranslated region; Pro = promoter; Term = terminator. Changes in FIGS.17B - 17E are relative to FIG. 17A.DETAILED DESCRIPTION
[0086] DNA deaminases in conjunction with CRISPR-Cas tools offer to provide a step change in our ability to introduce precise DNA modifications into plant genomes. First, they don't require the introduction of a DNA double-strand break or repair template to precisely modify a chromosomal DNA target, ultimately, enabling multi-site editing approaches. Next, the frequencies of desirable base conversion far exceed those offered by homology directed repair approaches in plants.
[0087] The introduction of targeted variation in multiplex offers the potential to engineer quantitative traits in plants like yield. Described herein is a Cas-alpha 10 (also known as SpaCasl2fl) base-editing platform (described herein as “high yielding precision randomization” (HYPeR)) that produces single nucleotide polymorphisms stemming from the deamination of one or more nucleotides in both target and non-target strands of a double-stranded DNA (dsDNA) target site.
[0088] Terms used in the claims and specification are defined as set forth below unless otherwise specified. It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise.
[0089] The meaning of abbreviations is as follows: “sec” means second(s), “min” means minute(s), “h” means hour(s), “d” means day(s), “microliters” means microliter(s), “mL” means milliliter(s), “L” means liter(s), “uM” means micromolar, “mM” means millimolar, “M” means molar, “mmol” means millimole(s), “umole” mean micromole(s), “g” means gram(s), “micrograms” or “ug” means microgram(s), “ng” means nanogram(s), “U” means unit(s), “bp” means base pair(s) and “kb” means kilobase(s).
[0090] An “altered target site”, “altered target sequence”, “modified target site”, “modified target sequence” are used interchangeably herein and refer to a target sequence as disclosed herein that comprises at least one alteration or modification when compared to a non-altered target sequence. Such alterations or modifications include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) - (iii).
[0091] As used herein, the term “before”, in reference to a sequence position, refers to an occurrence of one sequence upstream, or 5’, to another sequence.
[0092] The compositions and methods herein can provide for an improved “agronomic trait” or “trait of agronomic importance” or “trait of agronomic interest” to a plant, which can include, butnot be limited to, the following: disease resistance, drought tolerance, heat tolerance, cold tolerance, salinity tolerance, metal tolerance, herbicide tolerance, improved water use efficiency, improved nitrogen utilization, improved nitrogen fixation, pest resistance, herbivore resistance, pathogen resistance, yield improvement, health enhancement, vigor improvement, growth improvement, photosynthetic capability improvement, nutrition enhancement, altered protein content, altered oil content, increased biomass, increased shoot length, increased root length, improved root architecture, modulation of a metabolite, modulation of the proteome, increased seed weight, altered seed carbohydrate composition, altered seed oil composition, altered seed protein composition, altered seed nutrient composition, as compared to an isoline plant not comprising a modification derived from the methods or compositions herein.
[0093] “Agronomic trait potential” is intended to mean a capability of a plant element for exhibiting a phenotype, preferably an improved agronomic trait, at some point during its life cycle, or conveying said phenotype to another plant element with which it is associated in the same plant.
[0094] An “allele” is one of several alternative forms of a gene occupying a given locus on a chromosome. When all the alleles present at a given locus on a chromosome are the same, that plant is homozygous at that locus. If the alleles present at a given locus on a chromosome differ, that plant is heterozygous at that locus.
[0095] “Coding sequence” refers to a polynucleotide sequence that codes for a specific amino acid sequence.
[0096] A “codon-modified gene”, “codon-preferred gene”, or “codon-optimized gene” is a gene having its frequency of codon usage designed to mimic the frequency of preferred codon usage of a host cell.
[0097] A “complex trait locus” includes a genomic locus that has multiple transgenes genetically linked to each other.
[0098] As used herein, “crossed” or “cross” or “crossing” refers to the fusion of gametes via pollination to produce progeny (i.e., cells, seeds, or plants). The term encompasses both sexual crosses (the pollination of one plant by another) and selfing (self-pollination, i.e., when the pollen and ovule (or microspores and megaspores) are from the same plant or genetically identical plants).
[0099] A “deaminase” is an enzyme that catalyzes a deamination reaction. For example, deamination of adenine with an adenine deaminase results in the formation of inosine. Inosine selectively base pairs with cytosine instead of thymine. This results in a post-replicative transitionmutation, such that the original A - T base pair transforms into a G - C base pair. In another example, cytosine deamination results in the formation of uracil, which can be repaired by cellular repair mechanisms back to a C - T base pair or to a T - A, G - C, or A - T base pair. This heterogeneity in repair can be suppressed by the introduction of a uracil glycosylase inhibitor, such that DNA repair or replication transforms the original C - T base pair into a T - A base pair (Burnett et al. (2022) Frontiers in Genome Editing. 4, 923718). In the case of both adenine and cytosine deaminases, the introduction of a nick promotes the respective base pair change (Burnett et al., 2022).
[0100] The terms “decreased”, “fewer”, “reduced”, “slower” and “increased”, “faster”, “enhanced”, “greater” as used herein refer to a decrease or increase in a characteristic of a modified plant element or resulting plant compared to an unmodified plant element or resulting plant. For example, a decrease in a characteristic can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400%) or more lower than the untreated control and an increase can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400% or more higher than an untreated or unmodified control.
[0101] The terms “dicotyledonous” and “dicot” refer to the subclass of angiosperm plants also knows as “dicotyledoneae”, whose seeds typically comprise two embryonic leaves, or cotyledons. The term includes references to whole plants, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and progeny of the same.
[0102] As used herein, “domain” refers to a contiguous stretch of nucleotides (that can be RNA, DNA, and / or an RNA-DNA-combination sequence) or amino acids. The terms “conserveddomain” or “motif’ refer to a set of polynucleotides or amino acids conserved at specific positions along an aligned sequence of evolutionarily related proteins. While amino acids at other positions can vary between homologous proteins, amino acids that are highly conserved at specific positions indicate amino acids that are essential to the structure, the stability, or the activity of a protein. Because they are identified by their high degree of conservation in aligned sequences of a family of protein homologues, they can be used as identifiers, or “signatures”, to determine if a protein with a newly determined sequence belongs to a previously identified protein family.
[0103] As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into a genomic target site.
[0104] The term “endogenous” refers to a sequence or other molecule that naturally occurs in a cell or organism. For example, an endogenous polynucleotide is normally found in the genome of a cell (i.e., is not heterologous).
[0105] “Expression” as used herein refers to the production of a functional end-product (e g., an mRNA, guide RNA, or a protein) in either precursor or mature form.
[0106] The term “fragment” refers to a contiguous set of nucleotides or amino acids. For example, a fragment is 2, 3, 4, 5, 6, 7 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or greater than 20 contiguous nucleotides. A fragment may or may not exhibit the function of a sequence sharing some percent identity over the length of said fragment.
[0107] A “fragment that is functionally equivalent” and “functionally equivalent fragment” are used interchangeably herein. These terms refer to a portion or sub-sequence of an isolated nucleic acid fragment or polypeptide that displays the same activity or function as the longer sequence from which it is derived. For example, the fragment retains the ability to alter gene expression or produce a certain phenotype whether or not the fragment encodes an active protein. For example, the fragment can be used in the design of genes to produce the desired phenotype in a modified plant. Genes can be designed for use in suppression by linking a nucleic acid fragment, whether or not it encodes an active enzyme, in the sense or antisense orientation relative to a plant promoter sequence.
[0108] As used herein, “gene” refers to a nucleic acid fragment that expresses a functional molecule such as, but not limited to, a specific protein, including regulatory sequences preceding (5’ non-coding sequences) and following (3’ non-coding sequences) the coding sequence. “Nativegene” refers to a gene as found in its natural endogenous location with its own regulatory sequences.
[0109] As used herein “genome” refers to the entire complement of genetic material (genes and non-coding sequences) that is present in each cell of an organism, or virus or organelle; and / or a complete set of chromosomes inherited as a (haploid) unit from one parent. The term “genome” as it applies to a prokaryotic and eukaryotic cell or organism cells encompasses not only chromosomal DNA found within the nucleus, but organelle DNA found within subcellular components (e.g., mitochondria, or plastid) of the cell.
[0110] As used herein, a “genomic region” is a segment of a chromosome in the genome of a cell that is present on either side of a target site or, alternatively, also comprises a portion of a target site. A genomic region can comprise at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500, 5-600, 5- 700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5-1600, 5-1700, 5-1800, 5- 1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800. 5-2900, 5-3000, 5-3100 or more bases such that the genomic region has sufficient homology to undergo homologous recombination with the corresponding region of homology.
[0111] The term “heterologous” refers to the difference between the original environment, location, or composition of a particular polynucleotide or polypeptide sequence and its current environment, location, or composition. Non-limiting examples include differences in taxonomic derivation (e g., a polynucleotide sequence obtained from Zea mays would be heterologous if inserted into the genome of an Oryza sativa plant, or of a different variety or cultivar of Zea mays; or a polynucleotide obtained from a bacterium was introduced into a cell of a plant), or sequence (e.g., a polynucleotide sequence obtained from Zea mays, isolated, modified, and re-introduced into a maize plant). As used herein, “heterologous” in reference to a sequence can refer to a sequence that originates from a different species, variety, foreign species, or, if from the same species, is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, a promoter operably linked to a heterologous polynucleotide is from a species different from the species from which the polynucleotide was derived, or, if from the same / analogous species, one or both are substantially modified from their original form and / or genomic locus, or the promoter is not the native promoter for the operablylinked polynucleotide. Alternatively, one or more regulatory region(s) and / or a polynucleotide provided herein can be entirely synthetic.
[0112] As used herein, “homology” is meant describe DNA sequences that are similar. For example, a “region of homology to a genomic region” that is found on a donor DNA is a region of DNA that has a similar sequence to a given “genomic region” in the genome of a cell or organism. A region of homology can be of any length that is sufficient to promote homologous recombination at a cleaved target site. For example, a region of homology can comprise at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5- 50, 5-55, 5-60, 5-65, 5- 70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5- 200, 5-300, 5-400, 5-500, 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5-1200, 5-1300, 5-1400, 5- 1500, 5-1600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-2300, 5-2400, 5-2500, 5-2600, 5-2700, 5-2800, 5-2900, 5-3000, 5-3100 or more bases in length such that the region of homology has sufficient homology to undergo homologous recombination with the corresponding genomic region. “Sufficient homology” indicates that two polynucleotide sequences have structural similarity such that they are capable of acting as substrates for homologous recombination. The structural similarity includes overall length of each polynucleotide fragment, as well as the sequence similarity of the polynucleotides. Sequence similarity can be described by the percent sequence identity over the whole length of the sequences, and / or by conserved regions comprising localized similarities such as contiguous nucleotides having 100% sequence identity, and percent sequence identity over a portion of the length of the sequences.
[0113] As used herein, “homologous recombination” (HR) includes the exchange of DNA fragments between two DNA molecules at the sites of homology. The frequency of homologous recombination is influenced by a number of factors. Different organisms vary with respect to the amount of homologous recombination and the relative proportion of homologous to non- homologous recombination. Generally, the length of a region of homology affects the frequency of homologous recombination events: the longer the region of homology, the greater the frequency. The length of a homology region needed to observe homologous recombination is also speciesvariable. In many cases, at least 5 kb of homology has been utilized, but homologous recombination has been observed with as little as 25-50 bp of homology.
[0114] As used herein, “host” refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, a "host cell" refers to an in vivo or in vitro eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell),or cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, into which a heterologous polynucleotide or polypeptide has been introduced. In some embodiments, the cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-cell organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, in invertebrate cell, a vertebrate cell, a fish cell, a frog cell, a bird cell, an insect cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some cases, the cell is in vitro. In some cases, the cell is in vivo.
[0115] As used herein, “introducing” and “providing” are intended to mean presenting a subject molecule to a target, such as a cell or organism, a polynucleotide or polypeptide or polynucleotide- protein complex, in such a manner that the subject gains access to the target, such as the interior of a cell of the organism or to the cell itself, or in the case of a target polynucleotide, presented to the polynucleotide in such a way that the subject has capability of physical or chemical contact with the polynucleotide. An “isolated” or “purified” nucleic acid molecule, polynucleotide, polypeptide, or protein, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polynucleotide or protein as found in its naturally occurring environment. Thus, an isolated or purified polynucleotide, polypeptide, or protein is substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Generally, an “isolated” polynucleotide is free of sequences (e g., protein encoding sequences) that naturally flank the polynucleotide (i.e., sequences located at the 5' and 3' ends of the polynucleotide) in the genomic DNA of the organism from which the polynucleotide is derived. For example, in various embodiments, the isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequence that naturally flank the polynucleotide in genomic DNA of the cell from which the polynucleotide is derived. Isolated polynucleotides can be purified from a cell in which they naturally occur. The term also embraces recombinant polynucleotides and chemically synthesized polynucleotides.
[0116] The term “isoline” is a comparative term, and references organisms that are genetically identical, but differ in treatment. In one example, two genetically identical maize plant embryos can be separated into two different groups, one receiving a treatment (such as the introduction of a Cas polypeptide) and one control that does not receive such treatment. Any phenotypicdifferences between the two groups can thus be attributed solely to the treatment and not to any inherent of the plant's endogenous genetic makeup.
[0117] A “mature” protein refers to a post-translationally processed polypeptide (i.e., one from which any pre- or propeptides present in the primary translation product have been removed).
[0118] A “modified nucleotide” or “edited nucleotide” refers to a nucleotide sequence of interest that comprises at least one alteration when compared to its non-modified nucleotide sequence. Such “alterations” include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) - (iii).
[0119] The terms “monocotyledonous” and “monocot” refer to the subclass of angiosperm plants also known as “monocotyledoneae”, whose seeds typically comprise only one embryonic leaf, or cotyledon. The term includes references to whole plants, plant elements, plant organs (e g., leaves, stems, roots, etc.), seeds, plant cells, and progeny of the same.
[0120] A “mutated gene” is a gene that has been altered through human intervention. Such a “mutated gene” has a sequence that differs from the sequence of the corresponding non-mutated gene by at least one nucleotide addition, deletion, or substitution. A mutated plant is a plant comprising a mutated polynucleotide sequence or gene.
[0121] As used herein, “nucleic acid” generally refers to a polynucleotide and includes a single or a double-stranded polymer of deoxyribonucleotide or ribonucleotide bases. Nucleic acids can also include fragments and modified nucleotides. Thus, the terms “polynucleotide”, “nucleic acid sequence”, “nucleotide sequence” and “nucleic acid fragment” are used interchangeably to denote a polymer of RNA and / or DNA and / or RNA-DNA that is single- or double-stranded, optionally comprising synthetic, non-natural, or altered nucleotide bases. Nucleotides (usually found in their 5 ’-monophosphate form) are referred to by their single letter designation as follows: “A” for adenosine or deoxyadenosine (for RNA or DNA, respectively), “C” for cytosine or deoxy cytosine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purines (A or G), “Y” for pyrimidines (C or T), “K” for G or T, “H” for A or C or T, “I” for inosine, and “N” for any nucleotide.
[0122] An “optimized” polynucleotide is a sequence that has been optimized for improved expression in a particular heterologous host cell.
[0123] An “optimized nucleotide sequence" is a nucleotide sequence that has been optimized for expression in a particular organism. A plant-optimized nucleotide sequence includes a codon- optimized gene. A plant-optimized nucleotide sequence can be synthesized by modifying a nucleotide sequence encoding a protein such as, for example, a Cas polypeptide as disclosed herein, using one or more plant-preferred codons for improved expression. See, for example, Campbell and Gowri (1990) Plant Physiol. 92: 1-11 for a discussion of host-preferred codon usage.
[0124] As used herein, “open reading frame” is abbreviated ORF.
[0125] The term “operably linked” refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is regulated by the other. For example, a promoter is operably linked with a coding sequence when it is capable of regulating the expression of that coding sequence (i.e., the coding sequence is under the transcriptional control of the promoter). Coding sequences can be operably linked to regulatory sequences in a sense or antisense orientation. In another example, the complementary RNA regions can be operably linked, either directly or indirectly, 5 ’ to the target mRNA, or 3 ’ to the target mRNA, or within the target mRNA, or a first complementary region is 5’ and its complement is 3’ to the target mRNA.
[0126] The terms “plasmid”, “vector”, and “cassette” refer to a linear or circular extra chromosomal element often carrying genes that are not part of the central metabolism of the cell, and usually in the form of double-stranded DNA. Such elements can be autonomously replicating sequences, genome integrating sequences, phage, or nucleotide sequences, in linear or circular form, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a polynucleotide of interest into a cell. “Transformation cassette” refers to a specific vector comprising a gene and having elements in addition to the gene that facilitates transformation of a particular host cell. “Expression cassette” refers to a specific vector comprising a gene and having elements in addition to the gene that allow for expression of that gene in a host.
[0127] The term “plant” generically includes whole plants, plant organs, plant tissues, seeds, plant cells, seeds and progeny of the same. Plant cells include, without limitation, cells from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen and microspores. Plant cells comprise a plant cell wall, and as such are distinct, with different biochemical characteristics, from protoplasts that lack a cell wall.
[0128] A "plant element" or “plant part” is intended to reference either a whole plant or a plant component, which can comprise differentiated and / or undifferentiated tissues, for example but not limited to plant tissues, parts, and cell types. In one embodiment, a plant element is one of the following: whole plant, seedling, meristematic tissue, ground tissue, vascular tissue, dermal tissue, seed, leaf, root, shoot, stem, flower, fruit, stolon, bulb, tuber, corm, keiki, shoot, bud, tumor tissue, and various forms of cells and culture (e.g., single cells, protoplasts, embryos, callus tissue), plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant calli, plant clumps, and plant cells that are intact in plants or parts of plants such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, stalks, roots, root tips, anthers, and the like, as well as the parts themselves. Grain is intended to mean the mature seed produced by commercial growers for purposes other than growing or reproducing the species. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the invention, provided that these parts comprise the introduced polynucleotides. The term "plant organ" refers to plant tissue or a group of tissues that constitute a morphologically and functionally distinct part of a plant. As used herein, a "plant element" is synonymous to a "portion" or “part” of a plant, and refers to any part of the plant, and can include distinct tissues and / or organs, and can be used interchangeably with the term "tissue" throughout. Similarly, a "plant reproductive element" is intended to generically reference any part of a plant that is able to initiate other plants via either sexual or asexual reproduction of that plant, for example but not limited to: seed, seedling, root, shoot, cutting, scion, graft, stolon, bulb, tuber, corm, keiki, or bud. The plant element can be in plant or in a plant organ, tissue culture, or cell culture.
[0129] As used herein, a “polynucleotide of interest” encodes a protein or polypeptide that is “of interest” for a particular purpose, e.g. a selectable marker. A trait or polynucleotide “of interest” can be one that improves a desirable phenotype of a plant, particularly a crop plant, i.e. a trait of agronomic interest. Polynucleotides of interest: include, but are not limited to, polynucleotides encoding important traits for agronomics, herbicide-resistance, insecticidal resistance, disease resistance, nematode resistance, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, grain characteristics, commercial products, phenotypic marker, or any other trait of agronomic or commercial importance. A polynucleotide of interest can additionally be utilized in either the sense or anti-sense orientation. Further, more than one polynucleotide of interest can be utilized together, or “stacked”, to provide additional benefit. A“polynucleotide of interest” can encode a gene expression regulatory element, for example a promoter, intron, terminator, 5’UTR, 3’UTR, or other noncoding sequence. A “polynucleotide of interest” can comprise a DNA sequences that encodes for an RNA molecule, for example a functional RNA, siRNA, miRNA, or a guide RNA that is capable of interacting with a Cas polypeptide to bind to a target polynucleotide sequence.
[0130] A “population” of plants refers to a plurality of individual plants that share temporal and spatial location, and can further share one or more characteristic(s), such as a common genotype.
[0131] A “precursor” protein refers to the primary product of translation of mRNA (i.e., with pre- and propeptides still present). Pre- and pro-peptides can be but are not limited to intracellular localization signals.
[0132] “Progeny” comprises any subsequent generation of a plant.
[0133] A “promoter” is a region of DNA involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. The promoter sequence consists of proximal and more distal upstream elements, the latter elements often referred to as enhancers. An “enhancer” is a DNA sequence that can stimulate promoter activity, and can be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue-specificity of a promoter. Promoters can be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, and / or comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters can direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental conditions. It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of some variation can have identical promoter activity.
[0134] Promoters that cause a gene to be expressed in most cell types at most times are commonly referred to as “constitutive promoters”. The term “inducible promoter” refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, for example by chemical compounds (chemical inducers) or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulated promoters include, for example, promoters induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, phytohormones, wounding, or chemicals such as ethanol, abscisic acid (ABA), j asm onate, salicylic acid, or safeners.
[0135] A “protospacer adjacent motif’ (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide / Cas polypeptide complex described herein. The Cas polypeptide may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM herein can differ depending on the Cas polypeptide or Cas polypeptide complex used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides long.
[0136] The term “recombinant” refers to an artificial combination of two otherwise separated segments of sequence, e.g., by chemical synthesis, or manipulation of isolated segments of nucleic acids by genetic engineering techniques.
[0137] The terms “recombinant DNA molecule”, “recombinant DNA construct”, “expression construct”, “construct”, and “recombinant construct” are used interchangeably herein. A recombinant DNA construct comprises an artificial combination of nucleic acid sequences, e.g., regulatory and coding sequences that are not all found together in nature. For example, a recombinant DNA construct can comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source, but arranged in a manner different than that found in nature. Such a construct can be used by itself or can be used in conjunction with a vector. If a vector is used, then the choice of vector is dependent upon the method that will be used to introduce the vector into the host cells as is well known to those skilled in the art.
[0138] “Regulatory sequences” refer to nucleotide sequences located upstream (5’ non-coding sequences), within, or downstream (3’ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leader sequences, 5’ untranslated sequences, 3’ untranslated sequences, introns, polyadenylation target sequences, RNA processing sites, effector binding sites, and stem-loop structures.
[0139] “3’ non-coding sequences”, “transcription terminator”, and “termination sequences” refer to DNA sequences located downstream of a coding sequence and include polyadenylation recognition sequences and other sequences encoding regulatory signals capable of affecting mRNA processing or gene expression. The polyadenylation signal is usually characterized by affecting the addition of polyadenylic acid tracts to the 3’ end of the mRNA precursor. The use ofdifferent 3’ non-coding sequences is exemplified by Ingelbrecht et al., (1989) Plant Cell 1 :671 - 680.
[0140] “RNA transcript” refers to the product resulting from RNA polymerase-catalyzed transcription of a DNA sequence. When the RNA transcript is a perfect complimentary copy of the DNA sequence, it is referred to as the primary transcript or pre-mRNA. A RNA transcript is referred to as the mature RNA or mRNA when it is a RNA sequence derived from post- transcriptional processing of the primary transcript pre-mRNA. “Messenger RNA” or “mRNA” refers to the RNAthat is without introns and that can be translated into protein by the cell. “cDNA” refers to a DNA that is complementary to, and synthesized from, an mRNA template using the enzyme reverse transcriptase.
[0141] “Sequence identity” or “identity” in the context of nucleic acid or polypeptide sequences refers to nucleic acid bases or amino acid residues in two sequences that are the same when aligned for maximum correspondence over a specified comparison window. As used herein, “percentage of sequence identity” refers to the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window can comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the results by 100 to yield the percentage of sequence identity. Useful examples of percent sequence identities include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, or any percentage from 50% to 100%.
[0142] Polynucleotide and polypeptide sequences, variants thereof, and the structural relationships of these sequences can be described by the terms “homology”, “homologous”, “substantially identical”, “substantially similar”, and “corresponding substantially” which are used interchangeably herein. These refer to polypeptide or nucleic acid sequences wherein changes in one or more amino acids or nucleotide bases do not affect the function of the molecule, such as the ability to mediate gene expression or to produce a certain phenotype. These terms also refer to modification(s) of nucleic acid sequences that do not substantially alter the functional properties of the resulting nucleic acid relative to the initial, unmodified nucleic acid. These modificationsinclude deletion, substitution, and / or insertion of one or more nucleotides in the nucleic acid fragment. Substantially similar nucleic acid sequences encompassed can be defined by their ability to hybridize (under moderately stringent conditions, e.g., 0.5X SSC, 0.1% SDS, 60°C) with the sequences exemplified herein, or to any portion of the nucleotide sequences disclosed herein and which are functionally equivalent to any of the nucleic acid sequences disclosed herein. Stringency conditions can be adjusted to screen for moderately similar fragments, such as homologous sequences from distantly related organisms, to highly similar fragments, such as genes that duplicate functional enzymes from closely related organisms. Post-hybridization washes determine stringency conditions.
[0143] As used herein, a “targeted mutation” is a mutation in a gene (referred to as the target gene), including a native gene, made by altering a target sequence within the target gene using any method known to one skilled in the art, including a method involving a guided deamination polyprotein as disclosed herein.
[0144] The terms “target site”, “target sequence”, “target site sequence, ’’target DNA”, “target locus”, “genomic target site”, “genomic target sequence”, “genomic target locus”, “target polynucleotide”, and “protospacer”, are used interchangeably herein and refer to a polynucleotide sequence such as, but not limited to, a nucleotide sequence on a chromosome, episome, a locus, or any other DNA molecule in the genome (including chromosomal, chloroplastic, mitochondrial DNA, plasmid DNA) of a cell, at which a guide polynucleotide / dimeric deamination polyprotein complex can recognize, bind to, and optionally nick or cleave. The target site can be an endogenous site in the genome of a cell, or alternatively, the target site can be heterologous to the cell and thereby not be naturally occurring in the genome of the cell, or the target site can be found in a heterologous genomic location compared to where it occurs in nature. As used herein, terms “endogenous target sequence” and “native target sequence” are used interchangeable herein to refer to a target sequence that is endogenous or native to the genome of a cell and is at the endogenous or native position of that target sequence in the genome of the cell. An “artificial target site” or “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence can be identical in sequence to an endogenous or native target sequence in the genome of a cell but be located in a different position (i.e., a non-endogenous or non-native position) in the genome of acell. Methods for “modifying a target site” and “altering a target site” are used interchangeably herein and refer to methods for producing an altered target site.
[0145] “Translation leader sequence” refers to a polynucleotide sequence located between the promoter sequence of a gene and the coding sequence. The translation leader sequence is present in the mRNA upstream of the translation start sequence. The translation leader sequence can affect processing of the primary transcript to mRNA, mRNA stability or translation efficiency. Examples of translation leader sequences have been described (e.g., Turner and Foster, (1995) Mol Biotechnol 3:225-236).
[0146] In a first aspect, the disclosure provides an engineered, non-natural dimeric, multi-domain deamination polyprotein (referred to herein as a “deamination polyprotein”) comprising two nucleotide deaminases (a first nucleotide deaminase and a second nucleotide deaminase), two nuclease-inactive or “dead” Cas-alpha 10 polypeptides (a first dead Cas-alpha 10 polypeptide and a second dead Cas-alpha 10 polypeptide), and, if cytosine deaminases are used, two uracil glycosylase inhibitors (a first uracil glycosylase inhibitor and a second uracil glycosylase inhibitor).
[0147] In an example of this first aspect, the first and second nucleotide deaminases are cytosine deaminases. Turning to FIG. 1A, each monomer of the deamination polyprotein includes a cytosine deaminase, a nuclease-inactive or “dead” Cas-alpha 10 polypeptide (dCas-alpha 10
[0148] ), and a uracil glycosylase inhibitors. The Cas-alpha 10 domain in the dCas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. Each monomer of the deamination polyprotein can be constructed by linking the C- terminus of the cytosine deaminase (CD) to the N-terminus of the dCas-alpha 10 polypeptide and the C-terminus of the dCas-alpha 10 polypeptide to the N-terminus of the uracil glycosylase inhibitor (UGI) using suitable linkers. For example, the linker between the C-terminus of the CD and the N-terminus of the dCas-alpha 10 polypeptide can be a XTEN linker (mXTEN) (SEQ ID NO: 11) while the linker between the C-terminus of the dCas-alpha 10 polypeptide and the N- terminus of the UGI can be a flexible GGS linker (SEQ ID NO: 12).
[0149] In another example, as shown in FIG. IB, each monomer of the deamination polyprotein includes an adenine deaminase (AD) and a dCas-alpha 10 polypeptide. The Cas-alpha 10 domain in the dCas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. Optionally, the adenine deaminase monomer can be complexed to ainosine glycosylase, such as by linking the C-terminus of the dCas-alpha 10 polypeptide to the N- terminus of the inosine glycosylase.
[0150] In yet another example, as shown in FIG. 1C, a first monomer of the deamination polyprotein includes an adenine deaminase and a dCas-alpha 10 polypeptide and a second monomer of the deamination polyprotein includes a cytosine deaminase, a dCas-alpha 10 polypeptide, and a uracil glycosylase inhibitor. The Cas-alpha 10 domain in the dCas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. Optionally, the adenine deaminase monomer can be complexed to a inosine glycosylase, such as by linking the C-terminus of the dCas-alpha 10 polypeptide to the N-terminus of the inosine glycosylase.
[0151] In a second aspect, the disclosure provides a method of modifying a dsDNA target site in a cell by introducing single nucleotide polymorphisms in both the target and non -target strand of the dsDNA target site using a nuclease-inactive deamination polyprotein as described herein.
[0152] As shown in FIG. 5A, when a nuclease-inactive deamination polyprotein-gRNA complex deaminates cytosines on the non-target strand of a dsDNA target site to uracil (U) (top of FIG. 5A), the non-target strand serves as the template during DNA replication and repair resulting in cytosine-guanine (C-T) to thymine-adenine (T-A) single nucleotide polymorphisms (also known as transition mutations) (bottom of FIG. 5A).
[0153] Turning to FIG. 5B, the nuclease-inactive deamination polyprotein-gRNA complex deaminates cytosines on the target strand of a dsDNA target site to uracil (top of FIG. 5B), and the target strand serves as the template during DNA replication and repair resulting in G-C to A-T single nucleotide polymorphisms (bottom of FIG. 5B).
[0154] In athird aspect, the disclosure provides an engineered, non-natural dimeric, multi-domain deamination polyprotein comprising two nucleotide deaminases, two Cas-alpha 10 polypeptides, each having endonuclease or double-strand break inducing activity, and, if cytosine deaminases are used, two UGIs (a first UGI and a second UGI). The deamination polyprotein (also described herein as a nuclease-active deamination polyprotein) can then be paired with one or guide polynucleotides having various spacer lengths to achieve different genome-editing outcomes.
[0155] In an example of this third aspect, the first and second nucleotide deaminases are cytosine deaminases. Turning to FIG. 8A, each monomer of the nuclease-active deamination polyprotein includes a cytosine deaminase, a Cas-alpha 10 polypeptide having double-strand break activity,and a UGI. Each monomer of the deamination polyprotein can be constructed by linking the C- terminus of the cytosine deaminase (CD) to the N-terminus of the Cas-alpha 10 polypeptide and the C-terminus of the Cas-alpha 10 polypeptide to the N-terminus of the uracil glycosylase inhibitor (UGI) using suitable linkers. For example, the linker between the C-terminus of the CD and the N-terminus of the Cas-alpha 10 polypeptide can be a XTEN linker (mXTEN) (SEQ ID NO: 11) while the linker between the C-terminus of the Cas-alpha 10 polypeptide and the N- terminus of the UGI can be a flexible GGS linker (SEQ ID NO: 12). The Cas-alpha 10 domain in the Cas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. As further shown in FIG. 8A, a nuclease-active deamination polyprotein can be paired with various guide polynucleotides, each having a different variable targeting domain or spacer length, to achieve different genome-editing outcomes. For example, guide polynucleotide “A” in FIG. 8A, having a spacer length of 17-23 nucleotides can be used to introduce a double-strand break in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “B” in FIG. 8A, having a spacer length of 12-16 nucleotides, can be used to introduce a nick (single- strand break) in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “C” in FIG. 8A, having a spacer length of 10- 12 nucleotides, can be used to bind a DNA target site without introducing a double- or singlestrand break when complexed with the deamination polyprotein. Guide polynucleotide “D” in FIG. 8A, having a spacer length of 17-23 nucleotides and additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand (blocking nuclease activity on the non-target strand), can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein.
[0156] In another example, as shown in FIG. 8B, each monomer of the nuclease-active deamination polyprotein includes an adenine deaminase (AD) and a Cas-alpha 10 polypeptide having double-strand break activity. The Cas-alpha 10 domain in the Cas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. Optionally, the adenine deaminase monomer can be complexed to a inosine glycosylase, such as by linking the C-terminus of the Cas-alpha 10 polypeptide to the N-terminus of the inosine glycosylase. As further shown in FIG. 8B, the nuclease-active deamination polyprotein can be paired with various guide polynucleotides, each having a different variable targeting domain or spacer length, to achieve different genome-editing outcomes. For example, guide polynucleotide“A” in FIG. 8B, having a spacer length of 17-23 nucleotides can be used to introduce a doublestrand break in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “B” in FIG. 8B, having a spacer length of 12-16 nucleotides, can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “C” in FIG. 8B, having a spacer length of 10-12 nucleotides, can be used to bind a DNA target site without introducing a double- or single-strand break when complexed with the deamination polyprotein. Guide polynucleotide “D” in FIG. 8B, having a spacer length of 17-23 nucleotides and additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand (blocking nuclease activity on the non-target strand), can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein.
[0157] In yet another example, as shown in FIG. 8C, a first monomer of the nuclease-active deamination polyprotein includes an adenine deaminase and a Cas-alpha 10 polypeptide having double-strand break activity and a second monomer of the nuclease-active deamination polyprotein includes a cytosine deaminase, a Cas-alpha 10 polypeptide having double-strand break activity, and a uracil glycosylase inhibitor. The Cas-alpha 10 domain in the Cas-alpha 10 polypeptide oligomerizes upon binding to a guide polynucleotide, resulting in a dimerized polyprotein. Optionally, the adenine deaminase monomer can be complexed to a inosine glycosylase, such as by linking the C-terminus of the Cas-alpha 10 polypeptide to the N-terminus of the inosine glycosylase. As further shown in FIG. 8C, the nuclease-active deamination polyprotein can be paired with various guide polynucleotides, each having a different variable targeting domain or spacer length, to achieve different genome-editing outcomes. For example, guide polynucleotide “A” in FIG. 8C, having a spacer length of 17-23 nucleotides can be used to introduce a double-strand break in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “B” in FIG. 8C, having a spacer length of 12-16 nucleotides, can be used to introduce a nick (single- strand break) in a DNA target site when complexed with the deamination polyprotein. Guide polynucleotide “C” in FIG. 8C, having a spacer length of 10- 12 nucleotides, can be used to bind a DNA target site without introducing a double- or singlestrand break when complexed with the deamination polyprotein. Guide polynucleotide “D” in FIG. 8C, having a spacer length of 17-23 nucleotides and additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand (blocking nuclease activity onthe non-target strand), can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein.
[0158] In a fourth aspect, the disclosure provides a method of modifying a dsDNA target site in a cell by introducing single nucleotide polymorphisms in both the target and non-target strand of the dsDNA target site using a nuclease-active deamination polyprotein as described herein. Using a nuclease-active deamination polyprotein biases the introduction of transition polymorphisms to either C to T or G to A.
[0159] This strand-specific editing outcome is a result of which DNA strand in the dsDNA target site is utilized as template for cellular DNA repair, with the strand opposite the one nicked serving as the repair template. By introducing a nick in the non-target strand of a dsDNA target site, the target strand is utilized as a template during DNA repair, which introduces G-C to A-T polymorphisms in the repaired target site. Alternatively, when the target strand of a dsDNA target site is nicked, the non-target strand is used as a template during DNA repair, which introduces C- G to T-A polymorphisms in the repaired target site.
[0160] In a fifth aspect, the disclosure provides methods of using a nuclease-inactive deamination polyprotein and / or a nuclease-active deamination polyprotein as described herein for introducing polymorphisms to disrupt gene expression, increase or decrease gene expression, and / or introduce novel variation in a polypeptide encoded by a gene.
[0161] In a first example of this fifth aspect, the expression of a polypeptide encoded by a gene is disrupted (FIGS. 17B and 17C) by introducing one or more premature stop codon(s) within the open reading frame (ORF) of a gene (FIG. 17B; changes relative to FIG. 17A). For example, the introduction of C to T transition polymorphisms at CAA or CAG codons encoding glutamine, or a CGA codon encoding arginine produces a TAA, TAG, or TGA stop codon, respectively. Similarly, as shown in FIG. 17C (changes relative to FIG. 17A), a G to A SNP can be introduced, such that a TGG codon encoding tryptophan results in a TAG, TGA, or TAA stop codon. Alternatively or additionally, G to A polymorphisms can be introduced at exon-intron splice sites resulting in translation of an intron containing an in-frame stop codon or incorporation of the intron’s ORF into the polypeptide that results in a protein with reduced or abolished functionality (as shown in FIG. 17C, the canonical plant GT and AG splice sites are modified to AT and AA, respectively).
[0162] In a second example of this fifth aspect, targeted C and G transition mutations can be introduced into one or more regions capable of modulating gene expression (for example but notlimited to a TATA box, initiator element (Inr), downstream promoter element (DPE), TFTIB recognition element (BRE), downstream core element (DCE), motif ten element (MTE), X core promoter element (XCPE), intron, 5’ untranslated region (UTR), and / or 3’ UTR) (FIG. 17D; changes relative to FIG. 17A).
[0163] In a third example of this fifth aspect, new codons encoding novel amino acids can be introduced into a gene and the resulting variation screened for changes in protein function (FIG. 17E; changes relative to FIG. 17A). In this case, protein functionality can be reduced or lost, a dominant-negative mutation introduced, or a gain-of-function acquired.
[0164] In a sixth aspect, the disclosure provides methods of modulating gene expression in a cell by introducing a nuclease-active deamination polyprotein described herein with guide polynucleotides having different spacer lengths to achieve various genome-editing outcomes. The method comprises providing to a cell a nuclease-active deamination polyprotein as described herein along with one or more guide polynucleotides and introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of a dsDNA target site. Introducing the one or more single nucleotide polymorphisms modulates gene expression or alters a resulting polypeptide expressed by a gene in the cell by: disrupting gene expression by introducing a premature stop codon within an open reading frame of the gene; altering a polypeptide expressed from a gene by removing an exon-intron splice site; increasing or decreasing gene expression by introducing the one or more SNPs in a gene expression regulatory element; and / or altering a polypeptide expressed from a gene by introducing a codon encoding a novel amino acid into the gene.
[0165] In an example of this sixth aspect, a nuclease-active deamination polyprotein can be paired with various guide polynucleotides, each having a different variable targeting domain or spacer length, to achieve different genome-editing outcomes. For example, a guide polynucleotide having a spacer length of 17-23 nucleotides can be used to introduce a double-strand break in the DNA target site when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 12-16 nucleotides can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 10-12 nucleotides can be used to bind a DNA target site without introducing a double- or single-strand break when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 17-23 nucleotides and additional nucleotides added to the3’ end of the spacer capable of base pairing with the non-target strand (blocking nuclease activity on the non-target strand) can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein.
[0166] In a seventh aspect, the disclosure provides methods of modifying a dsDNA target site in a cell by introducing a nuclease-active deamination polyprotein described herein with guide polynucleotides having different spacer lengths to achieve various genome-editing outcomes. The method comprises providing to a cell a nuclease-active deamination polyprotein as described herein along with one or more guide polynucleotides and introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of a dsDNA target site and / or introducing an insertion, deletion, inversion, or translocation mutation at the dsDNA target site by inducing a double-strand break in the dsDNA target site.
[0167] In an example of this seventh aspect, the deamination polyprotein can be paired with various guide polynucleotides, each having a different variable targeting domain or spacer length, to achieve different genome-editing outcomes. For example, a guide polynucleotide having a spacer length of 17-23 nucleotides can be used to introduce a double-strand break in the DNA target site when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 12-16 nucleotides can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 10-12 nucleotides can be used to bind a DNA target site without introducing a double- or single-strand break when complexed with the deamination polyprotein. A guide polynucleotide having a spacer length of 17-23 nucleotides and additional nucleotides added to the 3’ end of the spacer capable of base pairing with the non-target strand (blocking nuclease activity on the non-target strand) can be used to introduce a nick (single-strand break) in a DNA target site when complexed with the deamination polyprotein.
[0168] NHEJ andHDR
[0169] In some aspects of the methods and compositions described herein, a genome editing system comprises a deamination polyprotein, one or more guide polynucleotides, and optionally donor DNA, and editing a target polynucleotide sequence comprises nonhomologous end-joining (NHEJ) or homologous recombination (HR) following a Cas endonuclease-mediated doublestrand break. Once a double-strand break is induced in the DNA, the cell's DNA repair mechanism is activated to repair the break. The most common repair mechanism to bring the broken endstogether is the nonhomologous end-joining pathway (Bleuyard et al., (2006) DNA Repair 5: 1 -12). The structural integrity of chromosomes is typically preserved by repair, but deletions, insertions, or other rearrangements are possible (Siebert and Puchta, (2002) Plant Cell 14: 1121-31; Pacher et al., (2007) Genetics 175:21-9). Alternatively, the double-strand break can be repaired by homologous recombination between homologous DNA sequences. Once the sequence around the double-strand break is altered, for example, by exonuclease activities involved in the maturation of double-strand breaks, gene conversion pathways can restore the original structure if a homologous sequence is available, such as a homologous chromosome in non-dividing somatic cells, or a sister chromatid after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenic DNA sequences may also serve as a DNA repair template for homologous recombination (Puchta, (1999) Genetics 152: 1173-81).
[0170] In some aspects of the methods and compositions described herein, the genome editing system comprises a deamination polyprotein, one or more guide polynucleotides, and a donor DNA. As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of a Cas endonuclease. Once a double-strand break is introduced in the target site by the endonuclease, the first and second regions of homology of the donor DNA can undergo homologous recombination with their corresponding genomic regions of homology resulting in exchange of DNA between the donor and the target genome. As such, the provided methods result in the integration of the polynucleotide of interest of the donor DNA into the double-strand break in the target site in the plant genome, thereby altering the original target site and producing an altered genomic target site.
[0001] In some aspects of the methods and compositions described herein, the deamination polyprotein utilizes a polynucleotide modification template. The term “polynucleotide modification template” includes a polynucleotide that comprises at least one nucleotide modification when compared to a nucleotide sequence to be edited. A nucleotide modification can be at least one nucleotide substitution, addition, or deletion. Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the at least one nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.
[0171] As used herein, a “polynucleotide of interest” encodes a protein or polypeptide that is “of interest” for a particular purpose, e.g. a selectable marker. In some aspects, a trait or polynucleotide“of interest” is one that improves a desirable phenotype of a plant, particularly a crop plant, i.e. a trait of agronomic interest. Polynucleotides of interest: include, but are not limited to, polynucleotides encoding important traits for agronomics, herbicide-resistance, insecticidal resistance, disease resistance, nematode resistance, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, grain characteristics, commercial products, phenotypic marker, or any other trait of agronomic or commercial importance. A polynucleotide of interest may additionally be utilized in either the sense or anti-sense orientation. Further, more than one polynucleotide of interest may be utilized together, or “stacked”, to provide additional benefit. In some aspects, a “polynucleotide of interest” may encode a gene expression regulatory element, for example a promoter, intron, terminator, 5’UTR, 3’UTR, or other noncoding sequence. In some aspects, a “polynucleotide of interest” may comprise a DNA sequences that encodes for an RNA molecule, for example a functional RNA, siRNA, miRNA, or a guide RNAthat is capable of interacting with a Cas endonuclease to bind to a target polynucleotide sequence.
[0172] Cas Polypeptide Base Editing
[0173] Cas polypeptides of the present disclosure refer to a polypeptide encoded by a Cas (CRISPR-associated) gene. The deamination polyproteins and methods described herein utilize a Cas-alpha polypeptide for base editing. A Cas polypeptide is further defined as a functional fragment or functional variant of a native Cas polypeptide, or a protein that shares at least 50%, between 50% and 55%, at least 55%, between 55% and 60%, at least 60%, between 60% and 65%, at least 65%, between 65% and 70%, at least 70%, between 70% and 75%, at least 75%, between 75% and 80%, at least 80%, between 80% and 85%, at least 85%, between 85% and 90%, at least 90%, between 90% and 95%, at least 95%, between 95% and 96%, at least 96%, between 96% and 97%, at least 97%, between 97% and 98%, at least 98%, between 98% and 99%, at least 99%, between 99% and 100%, or 100% sequence identity with at least 50, between 50 and 100, at least 100, between 100 and 150, at least 150, between 150 and 200, at least 200, between 200 and 250, at least 250, between 250 and 300, at least 300, between 300 and 350, at least 350, between 350 and 400, at least 400, between 400 and 450, at least 500, or greater than 500 contiguous amino acids of a native Cas polypeptide, and retains at least partial activity.
[0174] A Cas-alpha polypeptide is a functional RNA-guided, PAM-dependent protein of fewer than 800 amino acids, comprising: a C-terminal RuvC catalytic domain split into three subdomains and further comprising bridge-helix and one or more Zinc finger motif(s); and an N-terminal Recsubunit with a helical bundle, WED wedge-like (or “Oligonucleotide Binding Domain”, OBD) domain, and, optionally, a Zinc finger motif.
[0175] Cas-alpha polypeptides comprise one or more Zinc Finger (ZFN) coordination motif(s) that can form aZinc binding domain. Zinc Finger-like motifs can aid in target and non-target strand separation and loading of the guide polynucleotide into the DNA target. Cas-alpha polypeptides comprising one or more Zinc Finger motifs can provide additional stability to a ribonucleoprotein complex on a target polynucleotide. Cas-alpha polypeptides can comprise C4 or C3H zinc binding domains.
[0176] Cas-alpha polypeptides can function as double-strand-break-inducing agent or singlestrand-break inducing agent (nickases). In an example of the deamination polyproteins and methods described herein, the Cas-alpha polypeptide is a catalytically “dead” or inactive Cas-alpha polypeptide that can alter DNA bases without inducing a double- or single-strand break in a dsDNA target site. In another example of the deamination polyproteins and methods described herein, the Cas-alpha polypeptide has nicking active (i.e., is a Cas-alpha polypeptide nickase) that introduces a single-strand break in a dsDNA target site. In another example of the deamination polyproteins and methods described herein, the Cas-alpha polypeptide is catalytically active, also known as “nuclease active”, and has double-strand break inducing activity.
[0177] For many traits of interest, the creation of single double-strand breaks and the subsequent repair via HDR or NHEJ is not ideal for quantitative traits. An observed phenotype includes both genotype effects and environmental effects. The genotype effects further comprise additive effects, dominance effects, and epistatic effects. The probability of no effect per any single edit can be greater than zero, and any single phenotypic effect can be small, depending on the method used and site selected. Double-stranded break repair can additionally be “noisy” and have low repeatability.
[0178] One approach to ameliorate the probability of no effect per edit or small phenotypic effect outcome is to multiplex genome modification, such that a plurality of target sites are modified. Methods to modify a genomic sequence that do not introduce double-strand breaks would allow for single base substitutions. Combining these approaches, multiplexed base editing is beneficial for creating large numbers of genotype edits that can produce observable phenotype modifications. In some cases, dozens or hundreds or thousands of sites can be edited within one or a few generations of an organism.
[0179] A multiplexed approach to base editing in an organism, has the potential to create a plurality of significant phenotypic variations in one or a few generations, with a positive directional bias to the effects. In some aspects, the organism is a plant. A plant or a population of plants with a plurality of edits can be cross-bred to produce progeny plants, some of which will comprise multiple pluralities of edits from the parental lines. In this way, accelerated breeding of desired traits can be accomplished in parallel in one or a few generations, replacing time-consuming traditional sequential crossing and breeding across multiple generations.
[0180] The compositions (e.g., deamination polyproteins) and methods described herein can introduce a plurality of nucleobase edits in a dsDNA target polynucleotide resulting in a variant nucleotide sequence. One or more nucleobases of a target polynucleotide can be chemically altered to change the base from one type to another, for example from a Cytosine to a Thymine, or an Adenine to a Guanine. In some aspects, a plurality of bases, for example 2 or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more 90 or more, 100 or more, or even greater than 100, 200 or more, up to thousands of bases can be modified or altered, to produce a cell or organism with a plurality of modified bases.
[0181] The deamination polyproteins described herein can comprise abase editing deaminase and a polynucleotide-guided (e.g., gRNA) Cas-alpha polypeptide that form a functional complex with a guide polynucleotide that shares homology with a polynucleotide sequence at the dsDNA target site. The guided Cas-alpha polypeptide, of the dimeric deamination polyprotein, recognizes and binds to a double- stranded target sequence, opening the double-strand to expose individual bases. In the case of a cytidine deaminase, the deaminase deaminates the cytosine base and creates a uracil. Uracil glycosylase inhibitor (UGI) is provided to prevent the conversion of U back to C. DNA replication or repair mechanisms then convert the Uracil to a thymine (U to T), and subsequent repair of the opposing base (formerly G in the original G-C pair) to an Adenine, creating a T-A pair. (Komor et al. Nature Volume 533, Pages 420-424, 19 May 2016).
[0182] As used herein, “guide polynucleotide” refers to a polynucleotide sequence that can form a complex with a Cas polypeptide, including the Cas polypeptides described herein, and enables the Cas polypeptide to recognize, optionally bind to, and optionally cleave or nick a DNA target site. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence). The guide polynucleotide can contain modified or substitute bases. The terms “functional fragment”, “fragment that is functionallyequivalent” and “functionally equivalent fragment” of a guide RNA, crRNA or tracrRNA are used interchangeably herein, and refer to a portion or subsequence of the guide RNA, crRNA or tracrRNA, respectively, of the present disclosure in which the ability to function as a guide RNA, crRNA or tracrRNA, respectively, is retained. The terms “functional variant”, “variant that is functionally equivalent” and “functionally equivalent variant” of a guide RNA, crRNA or tracrRNA (respectively) are used interchangeably herein, and refer to a variant of the guide RNA, crRNA or tracrRNA, respectively, of the present disclosure in which the ability to function as a guide RNA, crRNA or tracrRNA, respectively, is retained. The terms “single guide RNA” and “sgRNA” are used interchangeably herein and relate to a synthetic fusion of two RNA molecules, a crRNA (CRISPR RNA) comprising a variable targeting domain (linked to a tracr mate sequence that hybridizes to a tracrRNA), fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA can comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment that can form a complex with a Cas polypeptide, wherein the guide RNA / Cas polypeptide complex can direct the Cas polypeptide (and the dimeric deamination polyprotein) to a DNA target site.
[0183] The terms “variable targeting domain”, “VT domain”, and “spacer” are used interchangeably herein and refer to a nucleotide sequence of a guide RNA that can hybridize (is complementary) to one strand (nucleotide sequence) of a dsDNA target site. The percent complementation between the first nucleotide sequence domain (VT domain) and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. The variable targeting domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. In an example of the compositions and methods described herein, the gRNAfor a deamination polyprotein comprises a spacer of between 12-18 nucleotides. In another example of the compositions and methods described herein, the gRNA for a deamination polyprotein comprises a spacer of between 12-16 nucleotides.
[0184] Cytosine Deaminase
[0185] Suitable CDs for use in the deamination polyprotein as described herein include the rat apolipoprotein B mRNA editing enzyme catalytic subunit 1 (rAPOBECl) (SEQ ID NO: 1), the human APOBEC subunit 3A (hAPOBEC3A) (SEQ ID NO: 2), a carboxyl (C) terminal deaminasemotif from a LamG-like jellyroll fold domain-containing protein from Streptomyces (SEQ ID NO: 3), the lizard APOBECl-like deaminase (1APOBEC1) (SEQ ID NO: 4), and a C-terminal deaminase domain from the APOBEC3G-like protein from Sorex araneus, the common shrew (SEQ ID NO: 106). The CDs of the deamination polyprotein can have amino acid sequences having at least 75% identity, alternatively at least 80% identity, alternatively at least 85% identity, alternatively at least 90% identity, alternatively at least 95% identity, alternatively 100% identity to SEQ ID Nos: 1-4 or 106.
[0186] Additional CDs for use in the deamination polyprotein as described herein include human activation induced cytidine deaminase (AID), Petromyzon marinus cytidine deaminase 1 (PmCDAl), human APOBEC subunit 3G (hAPOBEC3G), mouse APOBEC subunit 3G (mAPOBEC3G), double-strand DNA-specific deaminase toxin A (DddA). For review of CDs, see Xiang et al. 2023 Biophysics Reports 9(6):325-337.
[0187] Adenine Deaminase
[0188] Adenine deaminases (ADs) suitable for use in the deamination polyprotein as described herein include modified versions of A. coli TadA, such as TadA-7.10, TadA-7.10- F148A, V106W, TadA-minABEmax, TadA-8e, and TadA-9. Other ADs that can be used in the deamination polyprotein as described herein include ADAR1, ADAR2, and proteins containing their adenine deaminase domains e.g., ADAR2DD and modified versions of ADAR2DD. For review of ADs see Xiang et al. Ibid.
[0189] Uracil Glycosylase Inhibitor
[0190] The UGIs of the present disclosure are capable of introducing cytosine (C) to thymine (T) polymorphisms within a Cas-alpha 10 dsDNA target site. Examples of UGIs for use in the deamination polyprotein include, but are not limited to, amino acid sequences having at least 75% identity, alternatively at least 80% identity, alternatively at least 85% identity, alternatively at least 90% identity, alternatively at least 95% identity, alternatively 100% identity to SEQ ID Nos: 7-10.
[0191] dCas-alpha 10
[0192] The dCas-alpha 10 polypeptide of the deamination polyprotein comprises a different amino acid sequence from a native effector protein obtained or derived from Syntrophomonas palmitatica. Some exemplary Cas-alpha polynucleotides and dCas-alpha 10 polypeptides are described in, for example, US10934536, WO2022082179, and WO2023244992. For example, such a polypeptide can comprise a variant sequence (such as alanine, deletion, insertion, or anothernon-native amino acid) at a position corresponding to position 228, 327, or 434 of wild type Cas- alpha 10. In a particular example, the monomers of the deamination polyprotein include engineered dCas-alpha 10 polypeptides having an amino acid sequence at least 75% identical, alternatively at least 80% identical, alternatively at least 85% identical, alternatively at least 90% identical, alternatively at least 95% identical, alternatively 100% identical to SEQ ID NO: 5 or SEQ ID NO: 6.
[0193] Cas-alpha 10
[0194] The Cas-alpha 10 polypeptide of the deamination polyprotein comprises a different amino acid sequence from a native effector protein obtained or derived from Syntrophomonas palmitatica. Some exemplary Cas-alpha polynucleotides are described in, for example, US10934536, WO2022082179, and WO2023244992. In a particular example, the monomers of the deamination polyprotein include an engineered Cas-alpha 10 polypeptide having an amino acid sequence at least 75% identical, alternatively at least 80% identical, alternatively at least 85% identical, alternatively at least 90% identical, alternatively at least 95% identical, alternatively 100% identical to SEQ ID NOs: 79 or 90-92.
[0195] Evaluation and Selection of Target Sites
[0196] The length of the dsDNA sequence at the target site can vary, and includes, for example, target sites that are at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more than 30 nucleotides in length. It is further possible that the target site can be palindromic, that is, the sequence on one strand reads the same in the opposite direction on the complementary strand.
[0197] Any cell genome can be evaluated for potential target sites for multiplexed base editing. In one non-limiting example, an elite inbred line of a particular plant is selected, wherein the elite inbred line displays one or more desirable phenotypes. However, the elite inbred line can further comprise some alleles that can be optimized. Candidate editing sites are selected via one of several methods that identifies such optimizable alleles, that can be neutral or deleterious in the original genome but positive, beneficial, or desirable after editing, such as with a deaminase base editor.
[0198] Any approach or combination of approaches can be used to find candidate sites in a genome for multiple site editing, for example but not limited to: evolutionary conservation, computational functional prediction, or candidate QTNs identified from experimentation or the literature. Because single edit effects are likely to be small, multiplexing of the edits is an important feature.In some cases, individual phenotypes affected cannot be obvious, but fitness association would likely impact yield components of a plant.
[0199] In one approach, calculations of evolutionary conservation of individual nucleotides are performed and evaluated, such as Genomic Evolutionary Rate Profiling (GERP), which is a method for producing position-specific estimates of evolutionary constraint using maximum likelihood evolutionary rate estimation. Several "constrained elements" where multiple positions combine to give a signal that is indicative of a putative functional element are identified; this track shows the position-specific scores only, not the element predictions. Constraint intensity at each individual alignment position is quantified in terms of a "rejected substitutions" (RS) score, defined as the number of substitutions expected under neutrality minus the number of substitutions "observed" at the position.
[0200] Genomic sites for potential editing are scored independently. Positive scores represent a substitution deficit (i.e., fewer substitutions than the average neutral site) and thus indicate that a site can be under evolutionary constraint. Negative scores indicate that a site is probably evolving neutrally; negative scores should not be interpreted as evidence of accelerated rates of evolution because of too many strong confounders, such as alignment uncertainty or rate variance. Positive scores scale with the level of constraint, such that the greater the score, the greater the level of evolutionary constraint inferred to be acting on that site.
[0201] Potential sites for selection for editing, including base editing, are made of alleles at evolutionarily conserved loci that display nucleotide positions that are different than the consensus across multiple genomes. In plants, incomplete dominance of deleterious alleles contributes to trait variation and heterosis in maize.
[0202] In another method, candidate edits can be selected based on attention-based predictor network algorithms. Sites with a mutation increasing expression levels of a gene are more conserved, while sites with a mutation decreasing expression levels of a gene are less conserved. For example, conserved sites with a particular allele that has predicted regulatory effects could be one target for base editing.
[0203] Identification of candidate alleles for editing can be, for example but not limited to, nonconserved bases in a particular cell line at a conserved (polymorphic) site, bases producing nonsense codons, and / or predicted rare nonsynonymous high impact substitutions.
[0204] Recombinant Constructs for Transformation of Cells
[0205] The disclosed deamination polyproteins, and monomers thereof, and guide polynucleotides can be introduced into a cell. Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, non-conventional yeast, and plant cells as well as plants and seeds produced by the methods described herein.
[0206] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described more fully in Sambrook et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory: Cold Spring Harbor, NY (1989). Transformation methods are well known to those skilled in the art and are described infra.
[0207] Vectors and constructs include circular plasmids, and linear polynucleotides, comprising a polynucleotide of interest and optionally other components including linkers, adapters, regulatory or analysis. In some examples a recognition site and / or target site can be comprised within an intron, coding sequence, 5' UTRs, 3' UTRs, and / or regulatory regions.
[0208] Components for Expression and Utilization of CRISPR-Cas Systems in Prokaryotic and Eukaryotic Cells
[0209] The disclosure further provides expression constructs for expressing in a prokaryotic or eukaryotic cell or organism a guide polynucleotide and deamination polyprotein as described herein
[0210] The expression constructs of the disclosure can comprise a promoter operably linked to a nucleotide sequence encoding a Cas gene (or optimized sequence, including a Cas polypeptide gene described herein) and a promoter operably linked to a guide RNA of the present disclosure. The promoter is capable of driving expression of an operably linked nucleotide sequence in a prokaryotic or eukaryotic cell / organism.
[0211] Nucleotide sequence modification of the guide polynucleotide can be selected from, but not limited to, the group consisting of a 5' cap, a 3' polyadenylated tail, a riboswitch sequence, a stability control sequence, a sequence that forms a dsRNA duplex, a modification or sequence that targets the guide poly nucleotide to a subcellular location, a modification or sequence that provides for tracking , a modification or sequence that provides a binding site for proteins , a Locked Nucleic Acid (LNA), a 5-methyl dC nucleotide, a 2,6-Diaminopurine nucleotide, a 2’-Fluoro A nucleotide, a 2’ -Fluoro U nucleotide; a 2'-O-Methyl RNA nucleotide, a phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 molecule, a 5’ to 3’ covalent linkage, or any combination thereof. These modifications can result in at leastone additional beneficial feature, wherein the additional beneficial feature is selected from the group of a modified or regulated stability, a subcellular targeting, tracking, a fluorescent label, a binding site for a protein or protein complex, modified binding affinity to complementary target sequence, modified resistance to cellular degradation, and increased cellular permeability.
[0212] Expression Elements
[0213] Any polynucleotide encoding a component disclosed herein can be functionally linked to a heterologous expression element, to facilitate transcription or regulation in a host cell. Such expression elements include but are not limited to: promoter, leader, intron, and terminator. Expression elements can be “minimal” - meaning a shorter sequence derived from a native source, that still functions as an expression regulator or modifier. Alternatively, an expression element can be “optimized” - meaning that its polynucleotide sequence has been altered from its native state in order to function with a more desirable characteristic in a particular host cell. Alternatively, an expression element can be “synthetic” - meaning that it is designed in silico and synthesized for use in a host cell. Synthetic expression elements can be entirely synthetic, or partially synthetic (comprising a fragment of a naturally-occurring polynucleotide sequence).
[0214] Amethod of expressing RNAcomponents such as gRNAin eukaryotic cells for performing Cas9-mediated DNA targeting has been to use RNA polymerase III (Pol III) promoters, which allow for transcription of RNA with precisely defined, unmodified, 5’- and 3 ’-ends (DiCarlo et al., Nucleic Acids Res. 41: 4336-4343; Ma et al., Mol. Ther. Nucleic Acids 3:el61). This strategy has been successfully applied in cells of several different species including maize and soybean (US20150082478 published 19 March 2015). Methods for expressing RNA components that do not have a 5’ cap have been described (W02016 / 025131 published 18 February 2016).
[0215] Optimization of Sequences for Expression in Plants
[0216] Additional sequence modifications are known to enhance gene expression in a plant host. These include, for example, elimination of one or more sequences encoding spurious polyadenylation signals, one or more exon-intron splice site signals, one or more transposon-like repeats, and other such well-characterized sequences that can be deleterious to gene expression. The G-C content of the sequence can be adjusted to levels average for a given plant host, as calculated by reference to known genes expressed in the host plant cell. When possible, the sequence is modified to avoid one or more predicted hairpin secondary mRNA structures. Thus,"a plant-optimized nucleotide sequence" of the present disclosure comprises one or more of such sequence modifications.
[0217] Polynucleotides of Interest
[0218] Polynucleotides of interest or DNA target sites can be endogenous to the organism being edited, or can be provided as heterologous molecules to the organism.
[0219] General categories of polynucleotides of interest (DNA target sites) include, for example, genes of interest involved in information, such as zinc fingers, those involved in communication, such as kinases, and those involved in housekeeping, such as heat shock proteins. More specific polynucleotides of interest include, but are not limited to, genes involved in crop yield, grain quality, crop nutrient content, starch and carbohydrate quality and quantity as well as those affecting kernel size, sucrose loading, protein quality and quantity, nitrogen fixation and / or utilization, fatty acid and oil composition, genes encoding proteins conferring resistance to abiotic stress (such as drought, nitrogen, temperature, salinity, toxic metals or trace elements, or those conferring resistance to molecules such as pesticides or herbicides), genes encoding proteins conferring resistance to biotic stress (such as attacks by fungi, viruses, bacteria, insects, or nematodes, and development of diseases associated with these organisms).
[0220] Agronomically important traits such as oil, starch, and protein content can be genetically altered in addition to using traditional breeding methods. Modifications include increasing content of oleic acid, saturated and unsaturated oils, increasing levels of lysine and sulfur, providing essential amino acids, and also modification of starch.
[0221] Polynucleotide sequences of interest can encode proteins involved in providing disease or pest resistance. By "disease resistance" or "pest resistance" is intended that the plants avoid the harmful symptoms that are the outcome of the plant-pathogen interactions.
[0222] An "herbicide resistance protein" or a protein resulting from expression of an "herbicide resistance-encoding nucleic acid molecule" includes proteins that confer upon a cell the ability to tolerate a higher concentration of an herbicide than cells that do not express the protein, or to tolerate a certain concentration of an herbicide for a longer period of time than cells that do not express the protein. Herbicide resistance traits can be introduced into plants by genes coding for resistance to herbicides that act to inhibit the action of acetolactate synthase (ALS, also referred to as acetohydroxyacid synthase, AHAS), in particular the sulfonylurea (UK: sulphonylurea) type herbicides, genes coding for resistance to herbicides that act to inhibit the action of glutaminesynthase, such as phosphinothricin orbasta (e g., the bar gene), glyphosate (e g., the EPSP synthase gene and the GAT gene), HPPD inhibitors (e.g., the HPPD gene) or other such genes known in the art. See, for example, US Patent Nos. 7,626,077, 5,310,667, 5,866,775, 6,225,114, 6,248,876, 7,169,970, 6,867,293, and 9,187,762. The bar gene encodes resistance to the herbicide basta, the nptll gene encodes resistance to the antibiotics kanamycin and geneticin, and the ALS-gene mutants encode resistance to the herbicide chlorsulfuron.
[0223] Furthermore, it is recognized that the polynucleotide of interest can also comprise antisense sequences complementary to at least a portion of the messenger RNA (mRNA) for a targeted gene sequence of interest. Antisense nucleotides are constructed to hybridize with the corresponding mRNA. Modifications of the antisense sequences can be made as long as the sequences hybridize to and interfere with expression of the corresponding mRNA. In this manner, antisense constructions having 70%, 80%, or 85% sequence identity to the corresponding antisense sequences can be used. Furthermore, portions of the antisense nucleotides can be used to disrupt the expression of the target gene. Generally, sequences of at least 50 nucleotides, 100 nucleotides, 200 nucleotides, or greater can be used.
[0224] In addition, the polynucleotide of interest can also be used in the sense orientation to suppress the expression of endogenous genes in plants. Methods for suppressing gene expression in plants using polynucleotides in the sense orientation are known in the art. The methods generally involve transforming plants with a DNA construct comprising a promoter that drives expression in a plant operably linked to at least a portion of a nucleotide sequence that corresponds to the transcript of the endogenous gene. Typically, such a nucleotide sequence has substantial sequence identity to the sequence of the transcript of the endogenous gene, generally greater than about 65% sequence identity, about 85% sequence identity, or greater than about 95% sequence identity.
[0225] The polynucleotide of interest can also be a phenotypic marker. A phenotypic marker is screenable or a selectable marker that includes visual markers and selectable markers whether it is a positive or negative selectable marker. Any phenotypic marker can be used. Specifically, a selectable or screenable marker comprises a DNA segment that allows one to identify, or select for or against a molecule or a cell that comprises it, often under particular conditions. These markers can encode an activity, such as, but not limited to, production of RNA, peptide, or protein, or can provide a binding site for RNA, peptides, proteins, inorganic and organic compounds or compositions and the like.
[0226] Introduction of Deamination Polyprotein / Guide Polynucleotide Components into a Cell
[0227] The methods and compositions described herein do not depend on a particular method for introducing a sequence into an organism or cell, only that the polynucleotide or polypeptide gains access to the interior of at least one cell of the organism. Introducing includes reference to the incorporation of a nucleic acid into a eukaryotic or prokaryotic cell where the nucleic acid can be incorporated into the genome of the cell, and includes reference to the transient (direct) provision of a nucleic acid, protein or polynucleotide-protein complex (PGEN, RGEN) to the cell.
[0228] Methods for introducing polynucleotides or polypeptides or a polynucleotide-protein complex into cells or organisms are known in the art including, but not limited to, microinjection, electroporation, stable transformation methods, transient transformation methods, ballistic particle acceleration (particle bombardment), whiskers mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, viral -mediated introduction, transfection, transduction, cellpenetrating peptides, mesoporous silica nanoparticle (MSN)-mediated direct protein delivery, topical applications, sexual crossing , sexual breeding, and any combination thereof.
[0229] Protocols for introducing polynucleotides, polypeptides or polynucleotide-protein complexes (PGEN, RGEN) into eukaryotic cells, such as plants or plant cells are known.
[0230] Alternatively, polynucleotides can be introduced into plant or plant cells by contacting cells or organisms with a virus or viral nucleic acids. Generally, such methods involve incorporating a polynucleotide within a viral DNA or RNA molecule. In some examples a polypeptide of interest can be initially synthesized as part of a viral polyprotein, which is later processed by proteolysis in vivo or in vitro to produce the desired recombinant protein. Methods for introducing polynucleotides into plants and expressing a protein encoded therein, involving viral DNA or RNA molecules, are known, see, for example, U.S. Patent Nos. 5,889,191, 5,889,190, 5,866,785, 5,589,367 and 5,316,931.
[0231] The polynucleotide or recombinant DNA construct can be provided to or introduced into a prokaryotic and eukaryotic cell or organism using a variety of transient transformation methods. Such transient transformation methods include, but are not limited to, the introduction of the polynucleotide construct directly into the plant.
[0232] Nucleic acids and proteins can be provided to a cell by any method including methods using molecules to facilitate the uptake of anyone or all components of a guided Cas system(protein and / or nucleic acids), such as cell-penetrating peptides and nanocarriers. See also US20110035836 published 10 February 2011, and EP2821486A1 published 07 January 2015.
[0233] Other methods of introducing polynucleotides into a prokaryotic and eukaryotic cell or organism or plant part can be used, including plastid transformation methods, and the methods for introducing polynucleotides into tissues from seedlings or mature seeds.
[0234] Stable transformation is intended to mean that the nucleotide construct introduced into an organism integrates into a genome of the organism and is capable of being inherited by the progeny thereof. Transient transformation is intended to mean that a polynucleotide is introduced into the organism and does not integrate into a genome of the organism or a polypeptide is introduced into an organism. Transient transformation indicates that the introduced composition is only temporarily expressed or present in the organism.
[0235] A variety of methods are available to identify those cells having an altered genome at or near a target site without using a screenable marker phenotype. Such methods can be viewed as directly analyzing a target sequence to detect any change in the target sequence, including but not limited to PCR methods, sequencing methods, nuclease digestion, Southern blots, and any combination thereof.
[0236] The presently disclosed polynucleotides and polypeptides can be introduced into a cell. Cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, protist, fungal, insect, yeast, non-conventional yeast, and plant cells, as well as plants and seeds produced by the methods described herein. The cell of the organism can be a reproductive cell, a somatic cell, a meiotic cell, a mitotic cell, a stem cell, or a pluripotent stem cell.
[0237] Cells and Plants
[0238] The presently disclosed polynucleotides and polypeptides can be introduced into a plant cell. Plant cells include, well as plants and seeds produced by the methods described herein. Any plant can be used with the compositions and methods described herein, including monocot and dicot plants, and plant elements.
[0239] Genome modification via the compositions and methods described herein can be used to effect a genotypic and / or phenotypic change on the target organism. Such a change is preferably related to an improved trait of interest or an agronomically-important characteristic, the correction of an endogenous defect, or the expression of some type of expression marker. For example, the trait of interest or agronomically-important characteristic can be related to the overall health,fitness, or fertility of the plant, the yield of a plant product, the ecological fitness of the plant, or the environmental stability of the plant. In some aspects, the trait of interest or agronomically- important characteristic is selected from the group consisting of: agronomics, herbicide resistance, insecticide resistance, disease resistance, nematode resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, grain characteristics, commercial product production. In some aspects, the trait of interest or agronomically-important characteristic is selected from the group consisting of disease resistance, drought tolerance, heat tolerance, cold tolerance, salinity tolerance, metal tolerance, herbicide tolerance, improved water use efficiency, improved nitrogen utilization, improved nitrogen fixation, pest resistance, herbivore resistance, pathogen resistance, yield improvement, health enhancement, vigor improvement, growth improvement, photosynthetic capability improvement, nutrition enhancement, altered protein content, altered starch content, altered carbohydrate content, altered sugar content, altered fiber content, altered oil content, increased biomass, increased shoot length, increased root length, improved root architecture, modulation of a metabolite, modulation of the proteome, increased seed weight, altered seed carbohydrate composition, altered seed oil composition, altered seed protein composition, altered seed nutrient composition, as compared to an isoline plant not comprising a modification derived from the methods or compositions herein.
[0240] Examples of monocot plants that can be used include, but are not limited to, com (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e g., pearl millet (Pennisetum glaucum), proso millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana)), wheat (Triticum species, for example Triticum aestivum, Triticum monococcum), sugarcane (Saccharum spp.), oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), pineapple (Ananas comosus), banana (Musa spp.), palm, ornamentals, turfgrasses, and other grasses.
[0241] Examples of dicot plants that can be used include, but are not limited to, soybean (Glycine max), Brassica species (for example but not limited to: oilseed rape or Canola) (Brassica napus, B. campestris, Brassica rapa, Brassica juncea), alfalfa (Medicago sativa), tobacco (Nicotiana tabacum), Arabidopsis (Arabidopsis thaliana), sunflower (Helianthus annuus), cotton (Gossypium arboreum, Gossypium barbadense), and peanut (Arachis hypogaea), tomato (Solanum lycopersicum), potato (Solanum tuberosum).
[0242] Additional plants that can be used include safflower (Carthamus tinctorius), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), citrus trees (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musa spp.), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beets (Beta vulgaris), vegetables, ornamentals, and conifers.
[0243] Vegetables that can be used include tomatoes (Lycopersicon esculentum), lettuce (e.g., Lactuca sativa), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus spp.), and members of the genus Cucumis such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and musk melon (C. melo). Ornamentals include azalea (Rhododendron spp.), hydrangea (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnation (Dianthus caryophyllus), poinsettia (Euphorbia pulcherrima), and chrysanthemum.
[0244] Conifers that can be used include pines such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); redwood (Sequoia sempervirens); true firs such as silver fir (Abies amabilis) and balsam fir (Abies balsamea); and cedars such as Western red cedar (Thuja plicata) and Alaska yellow cedar (Chamaecyparis nootkatensis).
[0245] In certain embodiments of the disclosure, a fertile plant is a plant that produces viable male and female gametes and is self-fertile. Such a self-fertile plant can produce a progeny plant without the contribution from any other plant of a gamete and the genetic material comprised therein. Other embodiments of the disclosure can involve the use of a plant that is not self-fertile because the plant does not produce male gametes, or female gametes, or both, that are viable or otherwise capable of fertilization.
[0246] The present disclosure finds use in the breeding of plants comprising one or more edited alleles created by the methods or compositions disclosed herein. In some aspects, the edited alleles influence the phenotypic expression of one or more traits, such as plant health, growth, or yield. In some aspects, two plants can be crossed via sexual reproduction to create progeny plant(s) that comprise some or all of the edits from both parental plants.
[0247] Cells and Animals
[0248] The presently disclosed polynucleotides and polypeptides can be introduced into an animal cell. Animal cells can include, but are not limited to: an organism of a phylum including chordates, arthropods, mollusks, annelids, cnidarians, or echinoderms; or an organism of a class including mammals, insects, birds, amphibians, reptiles, or fishes. In some aspects, the animal is human, mouse, C. elegans, rat, fruit fly (Drosophila spp.), zebrafish, chicken, dog, cat, guinea pig, hamster, chicken, Japanese ricefish, sea lamprey, pufferfish, tree frog (e.g., Xenopus spp.), monkey, or chimpanzee. Particular cell types that are contemplated include haploid cells, diploid cells, reproductive cells, neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, embryonic cells, hematopoietic cells, bone cells, germ cells, somatic cells, stem cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some aspects, a plurality of cells from an organism can be used.
[0249] Genome modification via the compositions and methods described herein can be used to effect a genotypic and / or phenotypic change on the target organism. Such a change is preferably related to an improved phenotype of interest or a physiologically-important characteristic, the correction of an endogenous defect, or the expression of some type of expression marker. In some aspects, the phenotype of interest or physiologically-important characteristic is related to the overall health, fitness, or fertility of the animal, the ecological fitness of the animal, or the relationship or interaction of the animal with other organisms in its environment. In some aspects, the phenotype of interest or physiologically-important characteristic is selected from the group consisting of: improved general health, disease reversal, disease modification, disease stabilization, disease prevention, treatment of parasitic infections, treatment of viral infections, treatment of retroviral infections, treatment of bacterial infections, treatment of neurological disorders (for example but not limited to: multiple sclerosis), correction of endogenous genetic defects (for example but not limited to: metabolic disorders, Achondroplasia, Alpha-1 Antitrypsin Deficiency, Antiphospholipid Syndrome, Autism, Autosomal Dominant Polycystic Kidney Disease, Barth syndrome, Breast cancer, Charcot-Marie-Tooth, Colon cancer, Cri du chat, Crohn's Disease, Cystic fibrosis, Dercum Disease, Down Syndrome, Duane Syndrome, Duchenne Muscular Dystrophy, Factor V Leiden Thrombophilia, Familial Hypercholesterolemia, Familial Mediterranean Fever, Fragile X Syndrome, Gaucher Disease, Hemochromatosis, Hemophilia, Holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, MyotonicDystrophy, Neurofibromatosis, Noonan Syndrome, Osteogenesis Imperfecta, Parkinson's disease, Phenylketonuria, Poland Anomaly, Porphyria, Progeria, Prostate Cancer, Retinitis Pigmentosa, Severe Combined Immunodeficiency (SCID), Sickle cell disease, Skin Cancer, Spinal Muscular Atrophy, Tay-Sachs, Thalassemia, Trimethylaminuria, Turner Syndrome, Velocardiofacial Syndrome, WAGR Syndrome, and Wilson Disease), treatment of innate immune disorders (for example but not limited to: immunoglobulin subclass deficiencies), treatment of acquired immune disorders (for example but not limited to: AIDS and other HIV-related disorders), treatment of cancer, as well as treatment of diseases, including rare or “orphan” conditions, that have eluded effective treatment options with other methods.
[0250] Cells that have been genetically modified using the compositions or methods disclosed herein can be transplanted to a subject for purposes such as gene therapy, e.g. to treat a disease, or as an antiviral, antipathogenic, or anticancer therapeutic, for the production of genetically modified organisms in agriculture, or for biological research.
[0251] While the invention has been particularly shown and described with reference to a preferred embodiment and various alternate embodiments, it will be understood by persons skilled in the relevant art that various changes in form and details can be made therein without departing from the spirit and scope of the invention. For instance, while the particular examples below can illustrate the methods and embodiments described herein using a specific plant, the principles in these examples can be applied to any plant. Therefore, it will be appreciated that the scope of this invention is encompassed by the embodiments of the inventions recited herein and in the specification rather than the specific examples that are exemplified below. All cited patents and publications referred to in this application are herein incorporated by reference in their entirety, for all purposes, to the same extent as if each were individually and specifically incorporated by reference.EXAMPLESExample 1: Engineering Enzymes for Target and Non-target DNA Strand Deamination
[0252] In this Example, methods of engineering enzymes capable of deaminating one or more polynucleotides on one or both polynucleotide strands of a cellular double-stranded (ds) DNA target are described.
[0253] In one method, a nuclease-inactive or dead (d) homodimeric multi-domain heterologous polyprotein was engineered, the polyprotein comprised two cytosine deaminases (one or more of SEQ ID NOs: 1-4, 106), two nuclease-inactive or “dead” Cas-alpha 10 proteins (dCas-alpha 10) engineered to have enhanced RNA-guided DNA target recognition (one or both of SEQ ID Nos: 5-6), and two uracil glycosylase inhibitors (one or more of SEQ ID Nos: 7-10). A monomer of the polyprotein was initially constructed by linking the carboxyl terminus (C-terminus) of a cytosine deaminase to the amino terminus (N-terminus) of a Cas-alpha 10 protein using a 32 aa long modified XTEN linker (mXTEN) (SEQ ID NO: 11) (Schellenberger et al. (2009) Nature Biotechnology. 27: 1186-1190) and the C-terminus of the Cas-alpha 10 protein was attached to the N-terminus of a uracil glycosylase inhibitor using a flexible GGS linker (SEQ ID NO: 12) (Rosmalen et al. (2017) Biochemistry. 56:6565-6574). Since the Cas-alpha 10 domain in the polyprotein oligomerizes into two copies upon binding to a guide polynucleotide (Bigelyte et al. (2021) Nature Communications. 12:6191), it served as the basis for dimerization in the polyprotein.
[0254] Next, the frequency and location of polymorphism(s) in a Cas-alpha dsDNA target site was evaluated using different cytosine deaminase in the polyprotein were used. For this, a dual plasmid DNA A. coli assay was developed. The first plasmid containing a pBR322 origin of replication (ORI) and Kanamycin resistance gene expressed the polyprotein and Cas-alpha guide RNA while the second plasmid containing a pl 5a ORI and carbenicillin marker harbored the dsDNA target. In this way, the dual plasmid design allowed for interchangeability of the polyprotein and the dsDNA target.
[0255] In the first plasmid, the sequence encoding the polyprotein was E. coli codon-optimized. The resulting codon-optimized sequence was then synthesized in three fragments (GenScript, USA), and seamlessly inserted in an expression plasmid between an isopropyl [3-d- 1- thiogalactopyranoside (IPTG) inducible lac promoter (Donovan et al. (1996) Journal of Industrial Microbiology. 16:145-154) (SEQ ID NO: 13) and a terminator derived from the histidine biosynthetic operon (Di Nocera et al. (1978) Proceedings of the National Academy of Science, USA. 75:4276-4280) (SEQ ID NO: 14) using a NEBuilder HiFi DNA Assembly kit per the manufacturer’s instruction (NEB, USA). Once assembled, an expression cassette for a gRNA capable of complexing with the Cas-alpha 10 domain in the polyprotein and directing it to a DNA target site was synthesized (GenScript, USA) and introduced by seamless assembly into a sitedownstream of the histidine terminator. The gRNA expression cassette included a J23119 constitutive promoter (SEQ ID NO: 15) (iGEM Registry of Standard Biological Parts, www.partsregistry.org), a sequence encoding the gRNA (SEQ ID NO: 16), a sequence encoding the Hepatitis Delta Virus (HDV) ribozyme (SEQ ID NO: 17) (Kuo et al. (1988) Journal of Virology. 62:4439-4444), and a terminator from the T1 region of the rrnB E. coli gene (SEQ ID NO: 18) (Orosz et al. (1991) Europea Journal of Biochemistry . 201 :653-659).
[0256] The second plasmid containing the Cas-alpha 10 dsDNA target site was built by synthesizing complementary oligonucleotides with 28 nt overhangs (IDT, USA) followed by PCR amplification of the backbone and seamless assembly to insert the resulting duplexed DNA target into the plasmid. In a 5’ to 3’ direction, the target site included a 5’-TTC-3’ protospacer adjacent motif (PAM) followed immediately by a 20 nt protospacer or region capable of base pairing with the variable targeting domain or spacer in the gRNA. To accurately characterize the activity of the polyprotein, multiple target plasmids were built containing cytosines at various positions in the target site and the sequence encoding the gRNA adjusted accordingly. The dsDNA target site plasmid was next transformed into E. coli cells (E. cloni 10G cells, Lucigen Corporation, USA) per the manufacturer’s instruction and plated on LB agar containing Carbenicillin (100 pg / mL). Colonies were then selected, grown overnight at 37°C in liquid Terrific Broth (TB) supplemented with Carbenicillin (100 pg / mL), diluted 500-fold in fresh TB containing Carbenicillin, grown for 4-6 hrs at 37°C, and electrocompetent cells prepared (iGEM protocol). Next, the plasmid encoding the polyprotein and gRNA was transformed into the dsDNA target-containing competent cells and plated on LB agar containing Carbenicillin (100 pg / mL) and Kanamycin (50 pg / mL). Single colonies from each treatment were then grown overnight in liquid TB containing Carbenicillin (100 pg / mL) and Kanamycin (50 pg / mL). Cultures were then diluted 1000-fold into fresh TB supplemented with Carbenicillin and Kanamycin, IPTG was added at a final concentration of 0.05 mM, and cells grown overnight at 37°C. Plasmid DNA was then purified using a Qiagen plasmid miniprep kit (Qiagen, Netherlands), the dsDNA target plasmids evaluated for the presence of polymorphisms within or near the Cas-alpha dsDNA target site using targeted amplicon deep sequencing (Illumina, USA), and sequence alterations quantified using CRISPResso2 (Clement et al. (2019) Nature Biotechnology. 37:224-226). Treatments that used a gRNA with five or more mismatches between its variable targeting domain and the dsDNA target site served as a negative control.
[0257] Five different DNA cytosine deaminase (CD) domains were initially tested in combination with a dCas-alpha 10 (referred to herein as “dALTl 14” (SEQ ID NO: 5)) and the uracil glycosylase inhibitor from Bacillus subtilis bacteriophage PBS1 (SEQ ID NO: 7) (Mol et al. (1995) Cell. 82:701-708). The first cytosine deaminase domain was the rat apolipoprotein B mRNA editing enzyme catalytic subunit 1 (rAPOBECl) (SEQ ID NO: 1) (Petersen-Mahrt et al. (2003) The Journal of Biological Chemistry. 278: 19583-19586). The second cytosine deaminase domain was the human APOBEC subunit 3A (hAPOBEC3A) (SEQ ID NO: 2) (Chen et al. (2006) Current Biology. 16:480-485). The third cytosine deaminase domain was a carboxyl (C) terminal deaminase motif (SEQ ID NO: 3) from a LamG-like jellyroll fold domain-containing protein from Streptomyces (StsDaOl) (Thompson et al. (1999) Structure. 7:169-177). The fourth cytosine deaminase domain was the lizard APOBEC 1 -like deaminase (1APOBEC1) (SEQ ID NO: 4) (Severi et al. (2010) Molecular Biology and Evolution. 28: 1125-1129). The fifth cytosine deaminase was a C-terminal deaminase domain from the APOBEC3G-like protein from Sorex araneus, the common shrew (SEQ ID NO: 106).
[0258] FIG. 2A demonstrates the frequencies of DNA single nucleotide polymorphisms, and more specifically the frequency of cytosine (C) to thymine (T) editing as a function of the rAPOBECl, hAPOBEC3A, StsDaOl, and 1APOBEC1 CD domains fused to the amino terminal end of Cas- alpha 10 dALT114. Evaluations were carried-out in E. colixismg eight dsDNA plasmid target sites, DNA Target lA (“Tla”; SEQ ID NO: 80), DNA Target 1 (“Tl”; SEQ ID NO: 81), DNA Target 2A (“T2a”; SEQ ID NO: 82), DNA Target 2 (“T2”; SEQ ID NO: 83), DNA Target 3A (“T3a”; SEQ ID NO: 84), DNA Target 3 (“T3”; SEQ ID NO: 85), DNA Target 4A (“T4a”; SEQ ID NO: 86), and DNA Target 4 (“T4”; SEQ ID NO: 87) (FIG. 2B). Rates of C to T transition polymorphisms for dsDNA target sites 100% matching the variable targeting domain or spacer of the gRNA as well as sites containing mismatches between the spacer and DNA target site are shown.
[0259] As averaged across the eight dsDNA target sites, no or minimal sequence polymorphisms were identified for rAPOBECl relative to the negative control. In contrast, the hAPOBEC3A, StsDaOl, and 1APOBEC1 CD domains produced measurable (>5% of sequence reads) rates of cytosine to thymine transition mutations at positions 2-13 of the target site. 1APOBEC1 exhibited the highest rates of editing (greater than 90% of sequence reads with cytosine to thymine modification) although targets with a mismatch outside of positions 1-6 in the gRNA targeting portion of the DNA target site also showed editing indicating that targeted deamination occursrapidly even during transient DNA target binding and R-loop formation. Moreover, the cytosine to thymine polymorphisms observed in these cases with 1APOBEC1 were restricted to positions 1-6 that demonstrated 100% base pairing complementation with the gRNA targeting sequence, ultimately, reducing the window of editing in the DNA target site.
[0260] FIG. 2C demonstrates the frequencies of DNA single nucleotide polymorphisms in E. coli, and more specifically the frequency of cytosine (C) to thymine (T) editing when the sAPOBEC3G CD domain was fused to the amino terminal end of Cas-alpha 10 dALT114. Experiments were performed using four dsDNA plasmid target sites to examine sequence preferences of SAPOBEC3G (e g., TC, AC, CC, and GC), DNA Target 5 (“T5”; SEQ ID NO: 107), DNA Target 6 (“T6”; SEQ ID NO: 108), DNA Target 7 (“T7”; SEQ ID NO: 109), and DNA Target 8 (“T8”; SEQ ID NO: 110) (FIG. 2D).
[0261] In E. coli, sAPOBEC3G showed similar rates of C to T editing as observed with hAPOBEC3A. Additionally, editing occurred when preceded by a T, A, C, and G, although, rates of C to T transition SNPs were slightly higher in the context of a TC.
[0262] Next, the ability of the hAPOBEC3A polyprotein to induce RNA-guided C to T transition mutations in human HEK293T cells was tested. This was first accomplished by building a human- optimized polyprotein and gRNA plasmid DNA expression cassettes. To enable active transport into the nucleus, the respective polyprotein genes were modified by appending a sequence encoding an N-terminus bipartite nuclear localization signal (NLS) (SEQ ID NO: 19) onto the 5’ end and a MYC proto-oncogene NLS (SEQ ID NO:20) onto the 3’ end (Xu etal. (2021) Molecular Cell. 81 :4333-4345). The sequence encoding the resulting NLS tagged polyprotein was next codon-optimized for expression in human cells using methods known in the art (Koblan et al. (2018) Nature Biotechnology . 36:843-846), synthesized (GenScript, USA), and operably subcloned into a human cell expression plasmid DNA between a human betaherpesvirus promoter and enhancer (SEQ ID Nos: 21 and 22) and human-optimized sequences encoding a ribosome skipping T2A peptide (SEQ ID NO: 23) (Ahier et al. (2014) Genetics.196:605-613), green fluorescent protein (GFP) (SEQ ID NO: 24), and a terminator derived from the human gammaherpesvirus (SEQ ID NO: 25) (GenScript).
[0263] Human-optimized gRNA expression plasmids were next constructed. Briefly, a human U6 polymerase III promoter (SEQ ID NO: 26) was operably linked to a sequence encoding a Cas- alpha 10 gRNA (SEQ ID NO: 16) containing a 20 nt region at its 3’ end capable of base pairingwith one strand of a dsDNAtarget site. To ensure accurate maturation, a sequence encoding a HDV ribozyme (SEQ ID NO: 17) was next appended to the 3’ of the resulting sequence. This was then followed by the addition of a U6 terminator. The resulting gRNA expression cassette was then synthesized (GenScript, USA) and cloned into a pucl9 plasmid DNA backbone over EcoRI and Kpnl restriction sites (GenScript, USA).
[0264] Polyprotein and gRNA expression cassettes were next transfected into HEK293T cells and dsDNAtarget sites evaluated for the presence of sequence modifications. Transformation was first performed by culturing HEK293T cells (ATCC catalog umber CRL-3216) in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS) and penicillin (100 U / ml) / streptomycin (100 ug / ml) at 37°C in 5% CO2. Cells were then seeded in 96-well plates at a 1.8x 104 cells / well density. After approximately one day growth, cells were transfected with plasmids encoding the polyprotein and U6-guide RNA expression cassettes using FuGENE HD (Promega) per manufacturer’s instruction. Briefly, on the day of transfection, 0.450 pg of plasmid DNA (0.353 pg of polyprotein and 0.099 pg of gRNA expression cassettes) was diluted in a final volume of 20 pl of room temperature Opti-MEM (Gibco), mixed, and 1.35 pl of FuGENE HD was added directly to the solution (not allowing undiluted transfection reagent to contact the sides or bottom of the tube or plate) to achieve a 3: 1 ratio of transfection reagent to plasmid DNA. The resulting mixture was then incubated for 10 mins at room temperature, 20 pl added to each HEK293T seeded well, and mixed gently by pipetting or by using a shaker. Transfected cells were then grown at 37°C with 5% CO2 for 4 days and GFP expression monitored by fluorescence microscope. Cells were then collected by trypsinization and genomic DNA isolated using QuickExtract (Lucigen, USA). Changes to the DNA target site were then assayed using targeted amplicon deep sequencing (Illumina, USA) (Bigelyte et al. (2021) Nature Communications. 12:6191) and the resulting sequence reads analyzed with CRISPResso2 (Clement et al. (2019) Nature Biotechnology. 37:224-226). Changes within and immediately adjacent to the gRNA target site (+ / -10 bp) were validated as true alterations by referencing the results from negative control experiments (transfections where the polyprotein and / or gRNA plasmid DNA expression cassettes were omitted).
[0265] For the hAPOBEC3A polyprotein, eight of the ten sites targeted for editing in HEK293T demonstrated frequencies of DNA single nucleotide polymorphisms (SNPs) (FIG. 3A) that weresignificantly greater than the negative controls with low rates of insertion or deletion (indel) (FIG. 3B) mutations. Frequencies of SNPs ranged from four to fifteen percent.
[0266] FIGS. 4A-4H illustrate next-generation sequence (NGS) reads and observed targeted modifications in HEK293T cells from hAPOBEC3A Cas-alpha 10 dALT114 polyprotein editing experiments at the DNMT2 (FIG. 4A), FANCFT1 (FIG. 4B), VEGFA2 (FIG. 4C), WTAP4n (FIG. 4D), WTAPT3 (FIG. 4E), DNMT1 (FIG. 4F), HBB ln (FIG. 4G), and WTAPT5 (FIG. 4H) target sites. The observed targeted modifications included both cytosine (C) to thymine (T) and guanine (G) to adenine (A) transition mutations indicating deamination and repair of cytosines on both the non-target and target strands in the three nucleic acid structure formed by Cas-alpha 10 RNA-guided target recognition (FIGS. 5A & 5B). In general, the presence of C to T or G to A targeted polymorphisms in the eight target sites were mutually exclusive (FIGS. 4A-4H) suggesting that incorporation of the transition mutation was dependent on which DNA strand (nontarget or target) was used as template during DNA repair or replication. Additionally, for the eight modified targets, single base substitutions were identified across a 28 bp window starting at the - 1 position and extending out to +27 with the highest rate of modification found between +3 and +17 (FIG. 6). Data averaged across eight HEK293T targets with at least two sites in the collection having either a C or G present at each position in the non-target strand. Exceptions include two positions designated with open circles and represent the TT in the TTC PAM for Cas-alpha 10.
[0267] To enhance the observed editing efficiency in HEK293T cells, a different variant of the dCas-alpha 10 protein was used in the hAPOBEC3A polyprotein, dALT135-F54 (SEQ ID NO: 6). On average, it resulted in a 1.77-fold gain in the percentage of reads with a C to T or G to A polymorphisms (FIG. 7; data collected and averaged across eight target sites). The pattern of non- target and target strand cytosine editing, generally mutually exclusive C to T or G to A editing, and the window of editing remained similar to the initial design that used dALT114 (SEQ ID NO: 5).
[0268] To confirm the observed editing patterns, a second CD polyprotein, sAPOBEC3G, was evaluated. FIGS. 4I-4L show NGS reads and observed targeted modifications in HEK293T cells when either a hAPOBEC3A or a sAPOBEC3G Cas-alpha 10 dALT135-F54 polyprotein was used. In these side-by-side editing experiments (setup in triplicate), the frequency and occurrence of both C to T and G to A polymorphisms were similar between both CD polyproteins (FIGS. 3C, 4I-4L). It was also observed that the use of sAPOBEC3G resulted in a higher proportion of viableHEK293T cells compared to hAPOBEC3A as assayed using CyQUANT XTT (ThermoFisher Scientific, USA) when delivered on a plasmid DNA expression cassette (FIG. 3D).
[0269] In a second method, a nuclease-active homodimeric polyprotein gRNA complex capable of binding to and / or nicking a dsDNA target site was engineered. For this, a Cas-alpha 10 polypeptide (SEQ ID NO: 79) was utilized in the polyprotein and the gRNA variable targeting domain or spacer was modified (FIGS. 9A-9D). In a first experiment, the spacer of the gRNA was shortened to 12-16 nts resulting in nicking of only one DNA phosphodiester backbone in the dsDNA target site (FIG. 9A). This can be advantageous as the editing outcome observed above with a nuclease-inactive polyprotein can be shifted from a mixture of C to T or G to A SNPs to G to A polymorphisms (FIG. 10A) and likewise, a nick in the target strand can bias repair towards C to T transition mutations (FIG. 10B). In a second experiment, a gRNA spacer length of 10-12 nts are used to abolish any dsDNA target cleavage or nicking resulting in a gRNA capable of directing dsDNA target binding (FIG. 9B). It is expected that editing outcomes will be similar to those produced with a nuclease-inactive polyprotein resulting a mixture of C to T or G to A transition mutations (FIGS. 5A and 5B). In a third experiment, dsDNA target binding or nicking is engineered by appending a sequence onto the 3’ end of the spacer that is capable of base pairing with the non-target strand (NTS) (FIG. 9C and 9D). In this way, the RuvC nuclease domain(s) of Cas-alpha 10 are unable to access their single-stranded (ss) DNA substrate resulting in either dsDNA target nicking or binding.
[0270] The effect of gRNA variable targeting domain or spacer length on dsDNA target nicking was evaluated biochemically. Briefly, reactions were assembled in vitro with purified Cas-alpha 10 ALT135-F54 protein (having double-strand break / nuclease activity), gRNA, and linear dsDNA target where the non-target and target strands were 5’ labeled with FAM and ROX, respectively. Denaturing capillary electrophoresis (Stebbins, M. et al. (1997) J. Chromatogn, B, Biomed. Sci. Appl. 697:181-188 and Huai et al. (2017) Nat. Commun. 8: 1375) was then used to monitor rates of non-target and target strand cleavage 15 and 60 minutes after initiating the reaction. The effect of gRNA spacer length on non-target and target DNA strand cleavage is shown at TTR-sg7 (FIGS. 11A-11E), TTR-sg8 (FIGS. 12A-12E), RUNX1 (FIGS. 13A-13E), and RUNXl-pn2 (FIGS. 14A-14E). These data demonstrate that a variable targeting domain (aka spacer) of between 12-16 nucleotides (nts) produced DNA target nicking.
[0271] In FIGS. 11A-11E, these data demonstrate that a variable targeting domain (also known as a spacer) of between 14-16 nucleotides resulted in dsDNA target nicking. More specifically, a variable target domain of 18, 16, and 14 nucleotides produced DNAtarget nicking in the non-target strand while a variable target domain of 18 and 16 nucleotides produced nicking in the target strand.
[0272] In FIGS. 12A-12E, these data demonstrate that a variable targeting domain of between 12- 16 nucleotides resulted in dsDNA target nicking. More specifically, a variable target domain of 18, 16, 14, and 12 nucleotides resulted produced DNA target nicking in the non-target strand while a variable target domain of 18, 16, and 14 nucleotides produced nicking in the target strand.
[0273] In FIGS. 13A-13E, these data demonstrate that a variable targeting domain of between 14- 16 nucleotides resulted in dsDNA target nicking. More specifically, a variable target domain of 18, 16, and 14 nucleotides produced DNAtarget nicking in the non-target strand while a variable target domain of 18 and 16 nucleotides produced nicking in the target strand.
[0274] In FIGS. 14A-14E, these data demonstrate that a variable targeting domain of between 12- 16 nucleotides resulted in dsDNA target nicking. More specifically, a variable target domain of 18, 16, 14, and 12 nucleotides resulted produced DNA target nicking in the non-target strand while a variable target domain of 18, 16, and 14 nucleotides produced nicking in the target strand.
[0275] The effect of gRNA variable targeting domain or spacer length on dsDNA target binding was examined using a Saccharomyces cerevisiae colorimetric CRISPR activation (CRISPRa) assay. Briefly, sequences encoding nuclease-active ALT114 (SEQ ID NO: 92) or nuclease-active ALT135-F54 (SEQ ID NO: 79) (each having double-strand break / endonuclease activity) were codon optimized and linked in-frame to the 5’ end of a yeast optimized sequence encoding a flexible VP16 transcriptional activation domain linker (SEQ ID NO: 93) and a VPR transcriptional activation domain (SEQ ID NO: 94) (Chavez et al. (2015) Nature Methods. 12:326-328). Sequences encoding a simian virus 40 (SV40) NLS (SEQ ID NO: 95) were next added onto the 5’ and 3’ ends. The resulting NLS tagged gene was then synthesized and operably cloned between a ROX3 promoter (SEQ ID NO: 96) (Generoso et al. (2016) Journal of Microbiological Methods . 127:203-205) and CYC1 terminator (SEQ ID NO: 97) on a low copy (2-5 copies per cell) CEN6 replicating plasmid (Karim et al. (2013) FEMS Yeast Research. 13: 107-116) (GenScript, USA). To transcribe the gRNA necessary for directing Cas-alpha 10 dsDNA target binding, a DNA sequence encoding the HDV ribozyme (SEQ ID NO: 17) was first appended to the 3’ end of aDNA sequence encoding a Cas-alpha 10 single gRNA (SEQ ID NO: 16). Next, the SNR52 promoter (SEQ ID NO: 98) and SUP4 terminator (SEQ ID NO: 99) were operably linked to the ends of the resulting sequence incorporating a G bp onto the 3’ end of the SNR52 to promote transcription. DNA fragments were then synthesized and cloned into the CEN6 vector containing the cas-alpha 10 gene (GenScript). A reporter cassette comprising a weak pRevl promoter (SEQ ID NO: 100) (Lee el al. (2015) ACS Synthetic Biology. 4:975-986) operably linked to a sequence encoding a Cas-alpha 8 nuclease (SEQ ID NO: 101) followed by an alcohol dehydrogenase 1 (ADH1) terminator (SEQ ID NO: 102) (Lee el al. (2015) ACS Synthetic Biology. 4:975-986) was next inserted into the S. cerevisiae genome in safe harbor X-4 (Mikkelsen el al. (2012) Metabolic Engineering. 14: 104-111) using standard yeast recombineering techniques (Rothstein (1991) Methods Enzymology. 194:281-301). Next, the nALT114-VPR or nALT135-F54-VPR was directed to a single dsDNA target site in the pRevl promoter using gRNAs containing spacer lengths varying between 10 and 20 nts. Upon binding to the pRevl promoter, nALT114-VPR or nALT135-F54-VPR stimulated Cas-alpha 8 expression triggering the recognition and cleavage of a target site in the ADE2 gene. Note the Cas-alpha 8 gRNA (SEQ ID NO: 103) was expressed from a second SNR52-SUP4 expression cassette on the CEN6 plasmid. Upon disruption of ADE2, the color of the yeast cells shifted from white to red (pink) and this phenotypic change was used to read-out nALT114 or nALT135-F54 dsDNA binding. It was quantified by first capturing yeast colony images using a Nikon Digital Sight Ds-Fil camera (Nikon, Japan) and NIS-Elements BR software (version 4.00.07) (Nikon, Japan). Next, the total yeast area (as pixels) was measured, and the total percentage of red calculated using custom scripts. The effect of gRNA variable targeting domain or spacer length on nuclease-active Cas-alpha 10 (nATL114-VPR) dsDNA target binding is shown in FIG. 15. Values in the graph were calculated by dividing the percentage of red yeast for a given treatment by the percentage of red yeast from experiments using a nuclease-inactive or “dead” (d) Cas-alpha 10 ALT 114 and a gRNA containing a 20 nt length spacer (treatment / dALT114+20 nt spacer baseline). In this way, values less than one reflect reduced dsDNA target binding, those around one equivalent dsDNA target binding, and those greater than one improved dsDNA target binding. A variable targeting domain length of around 11 nts supports equivalent or enhanced dsDNA target binding compared with a nuclease-inactive Cas-alpha 10 protein and gRNA with a 20 nt length spacer.
[0276] In a third method, a heterodimeric polyprotein gRNA complex is engineered. For this, two different polyprotein expression cassettes are built. The first expression cassette encodes a polyprotein having a CD (e.g., SEQ ID Nos: 1-4) fused to the N-terminus of a Cas-alpha 10 polypeptide having endonuclease (double-strand break) activity (e.g., SEQ ID Nos: 79, 90-92) followed by a UGI domain (e.g., SEQ ID Nos: 7-10). The second expression cassette encodes a polyprotein having a monomeric or dimeric adenine deaminase (AD) (e.g., SEQ ID Nos: 104 and 105) linked in-frame to the N-terminus of a Cas-alpha 10 polypeptide having endonuclease (double-strand break) activity (e.g., SEQ ID Nos: 79, 90-92) with an optional C-terminus fusion of a UGI (e.g., SEQ ID Nos: 7-10). Next, upon expression in a cell, a mixture of homodimeric and heterodimeric polyproteins are formed around the respective gRNA capable of directing dsDNA target binding or nicking. It is expected if a binding gRNA is used, C to T, G to A, A to G, and T to C transition polymorphisms can all be recovered at the dsDNA target site (FIG. 16A and 16B). It is expected if a gRNA that results in NTS nicking is used, editing outcomes are limited to G to A and T to C SNPs (FIG. 16C).Example 2: High Yielding Precision Randomization (HYPeR)
[0277] In this Example, methods of introducing one or more transition polymorphisms stemming from the deamination of the non-target strand and / or target strand within or in the vicinity of a dsDNA target site of a homodimeric, heterologous multi-domain polyprotein are described. Termed “High yielding precision randomization (HYPeR)”, the method can be used to disrupt gene expression, increase or decrease gene expression, and / or introduce novel variation in a polypeptide encoded by a gene. FIGS. 17B-17E illustrate various methods of modulating a dsDNA target site (FIG. 17A) using a homodimeric multi-domain heterologous polyprotein.
[0278] In a first method, the expression of a polypeptide encoded by a gene is disrupted (FIGS. 17B and 17C). This is accomplished by introducing one or more premature stop codon(s) within the open reading frame (ORF) of a gene (FIG. 17B; changes relative to FIG. 17A). In this case, the introduction of C to T transition polymorphisms at CAA or CAG codons encoding glutamine, or a CGA codon encoding arginine produces a TAA, TAG, or TGA stop codon, respectively. Similarly, as shown in FIG. 17C (changes relative to FIG. 17A), a G to A SNP is introduced, and a TGG codon encoding tryptophan results in a TAG, TGA, or TAA stop codon. Alternatively, or additionally, G to A polymorphisms are introduced at exon-intron splice sites resulting in translation of an intron containing an in-frame stop codon or incorporation of the intron’s ORFinto the polypeptide that results in a protein with reduced or abolished functionality (as shown in FIG. 17C, the canonical plant GT and AG splice sites are modified to AT and AA, respectively).
[0279] In a second method, premature stop codons are removed from genes. For this, repair of deaminated adenines as described above is used to convert TAA, TAG, or TGA stop codons to TGG, CAA, CAG, or CGAthat encode tryptophan, glutamine, and arginine, respectively.
[0280] In a third method, targeted C, G, and A transition mutations can be introduced into one or more regions capable of modulating gene expression (for example but not limited to a TATA box, initiator element (Inr), downstream promoter element (DPE), TFIIB recognition element (BRE), downstream core element (DCE), motif ten element (MTE), X core promoter element (XCPE), intron, 5’ untranslated region (UTR), and / or 3’ UTR) (FIG. 17D and 17E; changes relative to FIG. 17A). For this, a neural network trained for prediction of gene expression in plants was used to select homodimeric heterologous multi-domain polypeptide targets. Initially, genes known to impact Zea mays plant height were selected through prior publications and other sources including mutant analyses, GWAS, or enrichment within expression datasets. The upstream 2 kb promoters of these genes were then collected. Within each promoter, sites were chosen by identifying the Cas-alpha 10 PAM, 5’-DTTC-3’, and then selecting 20 nts 3’ of the PAM for use as the gRNA target (on both the positive and negative DNA strands). Within each target site, the rate of transition polymorphisms was simulated by randomly sampling from a population of target sequences that contained C to T and G to A SNPs with a normal distribution, a mean of 3.5, and a sigma of 1.9. To ensure that a SNP was made when less than one mutation was sampled, one polymorphism was simulated. Similarly, if more than 10 mutations were sampled, only 10 alterations were introduced. For a given number of mutations, positions were randomly sampled from the DNA target site with 1-3 and 19-27 having a 35% chance of being modified while sites 4-18 had a probability of 65%. Once sites were selected, single base pair substitutions would occur if a C or G was present, substituting the base to T or A, respectively. The permutated sequence within the 2 kb promoter was then input into a transformer-based deep neural network trained for prediction of multi -tissue expression. Here, the predicted effect of substitutions within a single tissue were estimated relative to an unmutated sequence. To ensure each site was properly sampled, permutations within each target sequence was performed 500 times and the mean and standard deviation of the predicted effects were calculated for each target site within each tissue. The median standard deviation was then calculated across eight key tissues for each guide. Target sites were further filtered to ensurethat there was a lack of four T bases, the GC content was between 0.35-0.65, and that the target was specific (for example number of sites with no mismatch < 1, number of sites with 1 mismatch < 1, number of sites with 2 mismatches = 0, and number of sites with 3 mismatches < 10). Targets that passed these filters and had the largest median standard deviation, and thus effects on expression across key tissues, were selected for editing.
[0281] In a fourth method, new codons encoding novel amino acids are introduced into a gene and the resulting variation screened for changes in protein function (FIG. 17F; changes relative to FIG. 17A). In this case, protein functionality can be reduced or lost, a dominant-negative mutation introduced, or a gain-of-function acquired (Blackwell and Marsh (2022) Annual Review of Genomics and Human Genetics. 23:475-498). To increase the number of potential DNA target sites within or near these key motifs, a Cas-alpha 10 mutant, K85S, previously shown to shift PAM recognition from 5’-DTTC-3’ to 5’-YTC-3’ (9516-WO-PCT) was used (SEQ ID Nos: 88-91). To discriminate between a closely related homologs, target sites were selected to have at least one mismatch in positions 1-6 3’ of the PAM in the DNA target site.
[0282] In a fifth method, a combination of the aforementioned methods is performed concurrently (FIGS. 17B-17F). When using a nuclease-active polyprotein, one or more gRNA(s) with a 20 nt length spacer can also be used to simultaneously introduce one or more double-strand break(s) (DSBs) at one or more DNA target site(s) in a cell. In this way, genomic modifications resulting from DSB repair using either the non-homologous end-joining (NHEJ) or homologous recombination (HR) pathways (Hsu et al. (2014) Cell. 157: 1262-1278) can be introduced in a target or gRNA specific fashion in addition to alterations resulting from cytosine and / or adenine deamination.
Claims
CLAIMS1. A method of modulating gene expression in a cell, the method comprising: providing to the cell:(i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide, and a second nuclease-active Cas-alpha 10 polypeptide; and(ii) a first guide polynucleotide that shares homology with a first dsDNA target site in the cell; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the first dsDNA target site, wherein the dimeric polyprotein and the first guide polynucleotide form a first complex that recognizes a PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and optionally nicks the first dsDNA target site; wherein introducing the one or more single nucleotide polymorphisms (SNPs) modulates gene expression or alters a resulting polypeptide expressed by a gene in the cell by:(a) disrupting gene expression by introducing a premature stop codon within an open reading frame of the gene;(b) altering a polypeptide expressed from a gene by removing an exon-intron splice site;(c) increasing or decreasing gene expression by introducing the one or more SNPs in a gene expression regulatory element; and / or(d) altering a polypeptide expressed from a gene by introducing a codon encoding a novel amino acid into the gene.
2. The method of claim 1, wherein the first guide polynucleotide comprises a spacer that is 12-16 nucleotides in length, a spacer that is 10-12 nucleotides in length, or a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site.
3. The method of claim 1 or claim 2, further comprising providing to the cell a second guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length.
4. The method of claim 3, further comprising: providing to the cell the second guide polynucleotide that shares homology with a second dsDNA target site in the cell; inducing a double-strand break in the second dsDNA target site, wherein the dimeric polyprotein and the second guide polynucleotide form a second complex that recognizes a PAM sequence on the second dsDNA target site, binds to the second dsDNA target site, and cleaves the second dsDNA target site; and introducing an insertion, deletion, inversion, or translocation mutation at the second dsDNA target site via non-homologous end joining or homology-directed repair.
5. The method of claim 4, further comprising providing to the cell a polynucleotide modification template or a donor DNA.
6. The method of any one of claims 1-5, wherein the gene expression regulatory element is a TATA box, an initiator element, a downstream promoter element, a TFIIB recognition element, a downstream core element, a motif ten element, an X core promoter element, an intron, a 5’ untranslated region (UTR), or 3’ UTR.
7. The method of any one of claims 1-6, wherein: the first nucleotide deaminase is a first cytosine deaminase and the dimeric polyprotein further comprises a first uracil glycosylase inhibitor; and the second nucleotide deaminase is a second cytosine deaminase and the dimeric polyprotein further comprises a second uracil glycosylase inhibitor.
8. The method of claim 7, wherein the dimeric polyprotein comprises two monomers: the first monomer comprising the first cytosine deaminase, the first nuclease-active Cas- alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has doublestrand break activity, and the first uracil glycosylase inhibitor; andthe second monomer comprising the second cytosine deaminase, the second nucleaseactive Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the second uracil glycosylase inhibitor; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
9. The method of any one of claims 1-6, wherein: the first nucleotide deaminase is a cytosine deaminase and the dimeric polyprotein further comprises a uracil glycosylase inhibitor; and the second nucleotide deaminase is an adenine deaminase.
10. The dimeric polyprotein of claim 9, wherein the dimeric polyprotein comprises two monomers: the first monomer comprising the cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the uracil glycosylase inhibitor; and the second monomer comprising the adenine deaminase and the second nuclease-active Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
11. The method of any one of claims 1-6, wherein the first nucleotide deaminase is a first adenine deaminase and the second nucleotide deaminase is a second adenine deaminase.
12. The polyprotein of claim 11, wherein the dimeric polyprotein comprises two monomers: the first monomer comprising the first adenine deaminase and the first nuclease-activeCas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; andthe second monomer comprising the second adenine deaminase and the second nucleaseactive Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to the first and / or second guide polynucleotide.
13. The method of claim 1, wherein the spacer of the first guide polynucleotide is 10-12 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site and binds to the first dsDNA target site without nicking or cleaving the first dsDNA target site.
14. The method of claim 1, wherein the spacer of the first guide polynucleotide is 12-16 nucleotides in length and the first complex recognizes the PAM sequence of the first dsDNA target site, binds to the first dsDNA target site, and nicks the first dsDNA target site.
15. A non-natural, multi-domain, dimeric polyprotein comprising:(i) a first nucleotide deaminase;(ii) a second nucleotide deaminase;(iii)a first nuclease-active Cas-alpha 10 polypeptide; and(iv)a second nuclease-active Cas-alpha 10 polypeptide.
16. The dimeric polyprotein of claim 15, wherein: the first nucleotide deaminase is a first cytosine deaminase and the dimeric polyprotein further comprises a first uracil glycosylase inhibitor; and the second nucleotide deaminase is a second cytosine deaminase and the dimeric polyprotein further comprises a second uracil glycosylase inhibitor.
17. The dimeric polyprotein of claim 16, wherein the dimeric polyprotein comprises two monomers:the first monomer comprising the first cytosine deaminase, the first nuclease-active Cas- alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has doublestrand break activity, and the first uracil glycosylase inhibitor; and the second monomer comprising the second cytosine deaminase, the second nucleaseactive Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the second uracil glycosylase inhibitor; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
18. The dimeric polyprotein of claim 15, wherein: the first nucleotide deaminase is a cytosine deaminase and the dimeric polyprotein further comprises a uracil glycosylase inhibitor; and the second nucleotide deaminase is an adenine deaminase.
19. The dimeric polyprotein of claim 18, wherein the dimeric polyprotein comprises two monomers: the first monomer comprising the cytosine deaminase, the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity, and the uracil glycosylase inhibitor; and the second monomer comprising the adenine deaminase and the second nuclease-active Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
20. The dimeric polyprotein of claim 15, wherein the first nucleotide deaminase is a first adenine deaminase and the second nucleotide deaminase is a second adenine deaminase.
21. The polyprotein of claim 20, wherein the dimeric polyprotein comprises two monomers: the first monomer comprising the first adenine deaminase and the first nuclease-active Cas-alpha 10 polypeptide, wherein the first nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; andthe second monomer comprising the second adenine deaminase and the second nucleaseactive Cas-alpha 10 polypeptide, wherein the second nuclease-active Cas-alpha 10 polypeptide has double-strand break activity; wherein the first and second monomers dimerize upon binding to a guide polynucleotide.
22. A non-natural composition comprising:(i) a dimeric polyprotein according to any one of claims 15-21; and(ii) at least one guide polynucleotide that shares homology with a dsDNA target site in a cell, wherein the at least one guide polynucleotide comprises:(a) a guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length;(b) a guide polynucleotide comprising a spacer that is 12-16 nucleotides in length;(c) a guide polynucleotide comprising a spacer that is 10-12 nucleotides in length;(d) a guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site;(e) any combination of (a) - (d), wherein the dimeric polyprotein is capable of introducing single nucleotide polymorphism resulting from deamination of a target strand and / or a non-target strand of a dsDNA target site and / or an insertion, deletion, inversion, or translocation mutation in the dsDNA target site.
23. A cell comprising the dimeric polyprotein of any one of claims 15-21.
24. The cell of claim 23, wherein the cell is an eukaryotic cell.
25. The cell of claim 24, wherein the cell is a plant cell.
26. The composition of claim 22, comprising the guide polynucleotide of (a), the guide polynucleotide of (b), the guide polynucleotide of (c), and the guide polynucleotide of (d), wherein each guide polynucleotide shares homology with a unique dsDNA target site.
27. A method of modifying a dsDNA target site in a cell, the method comprising: providing to the cell:(i) a non-natural, multi-domain, dimeric polyprotein comprising a first nucleotide deaminase, a second nucleotide deaminase, a first nuclease-active Cas-alpha 10 polypeptide, and a second nuclease-active Cas-alpha 10 polypeptide; and(ii) one or more guide polynucleotides, each guide polynucleotide sharing homology with a unique dsDNA target site in the cell, wherein the one or more guide polynucleotides comprise:(a) a first guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length;(b) a second guide polynucleotide comprising a spacer that is 12-16 nucleotides in length;(c) a third guide polynucleotide comprising a spacer that is 10-12 nucleotides in length; or(d) a fourth guide polynucleotide comprising a spacer that is about 17-23 nucleotides in length with additional nucleotides at the 3’ end of the spacer that base pair with a non-target strand of the dsDNA target site; introducing one or more single nucleotide polymorphisms resulting from deamination of a target strand and / or a non-target strand of the dsDNA target site, wherein the dimeric polyprotein and the guide polynucleotide form a complex that recognizes a PAM sequence of the dsDNA target site, binds to the dsDNA target site, and optionally nicks the first dsDNA target site and / or introducing an insertion, deletion, inversion, or translocation mutation at the dsDNA target site by inducing a double-strand break in the dsDNA target site, wherein the dimeric polyprotein and the guide polynucleotide form a complex that recognizes a PAM sequence on the dsDNA target site, binds to the dsDNA target site, and cleaves the dsDNA target site.
Citation Information
Patent Citations
Method of introducing nucleic acid into plant cells
EP2821486A1
CRISPR-CAS systems for genome editing
US10934536B2
Nanocarrier based plant transfection and transduction
US20110035836A1
Plant genome modification using guide RNA / CAS endonuclease systems and methods of use
US20150082478A1
Glyphosate-tolerant 5-enolpyruvyl-3-phosphoshikimate synthases
US5310667A