Engineered CAS9 nuclease
Engineered SPCas9 polypeptides with specific amino acid substitutions address the issue of off-target effects in RNA-guided nucleases, achieving improved specificity and reduced regulatory concerns in genome editing applications.
Patent Information
- Application Number
- JP2023165873
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-01
- Filing Date
- 2023-09-27
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2038-06-08
AI Technical Summary
Existing RNA-guided nucleases, such as wild-type Cas9, exhibit off-target binding, cleavage, and editing, which raises regulatory concerns and reduces the precision of genome editing applications.
Engineered SPCas9 polypeptides with specific amino acid substitutions, such as D23A, Y128V, T67L, and D1251G, are developed to enhance target specificity, reducing off-target effects while maintaining or improving on-target activity.
The engineered SPCas9 polypeptides demonstrate significantly reduced off-target editing rates, typically by 5% to 95% compared to wild-type SPCas9, while maintaining or enhancing on-target binding and cleavage efficiency.
Smart Images

Figure 0007695315000031 
Figure 0007695315000032 
Figure 0007695315000033
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 517,811, filed on June 9, 2017, and U.S. Provisional Patent Application No. 62 / 665,388, filed on May 1, 2018. The entire contents of both are hereby incorporated by reference into this specification.
[0002] Sequence Listing This specification refers to a sequence listing (submitted electronically on June 8, 2018, as a.txt file named "2011271 - 0077_SL.txt"). The.txt file was created on June 4, 2018, and is 41,513 bytes in size. The sequence listing is hereby incorporated by reference in its entirety into this specification.
[0003] The present disclosure relates to CRISPR / Cas - related methods and components for editing a target nucleic acid sequence or regulating the expression of a target nucleic acid sequence, and related uses thereof. More particularly, the present disclosure relates to engineered Cas9 nucleases with altered and improved target specificity.
Background Art
[0004] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat) has evolved in bacteria and archaea as an adaptive immune system for defense against viral attack. Immediately after exposure to a virus, short segments of viral DNA are integrated into the CRISPR locus. RNA is transcribed from a portion of the CRISPR locus that contains the viral sequence. RNA containing a sequence complementary to the viral genome mediates the targeting of an RNA - guided nuclease protein, such as Cas9 or Cpf1, to a target sequence within the viral genome. Then, the RNA - guided nuclease silences the viral target by cleaving it.
[0005] The CRISPR system has been adapted for genome editing in eukaryotic cells. The system typically includes a protein component (RNA-guided nuclease) and a nucleic acid component (usually referred to as guide RNA or "gRNA"). These two components interact with a specific target DNA sequence recognized or complementary to the two components of the system, and in some cases, form a complex that edits or modifies the target sequence, for example, by site-specific DNA cleavage.
[0006] The value of nucleases such as these as a means for treating genetic diseases is widely recognized. For example, the U.S. Food and Drug Administration (FDA) held a Science Board Meeting on November 15, 2016, to address the use of such systems and the potential legal issues arising therefrom. In this committee, the FDA pointed out that the Cas9 / guide RNA (gRNA) ribonucleoprotein (RNP) complex can be customized to effect precise editing at the target locus, but this complex can also interact with and cleave other "off-target" positions. The possibility of off-target cleavage ("off-target") ultimately gives rise to the potential for legal issues regarding the approval of therapies utilizing these nucleases, at least. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0007] The present disclosure, in part, addresses potential regulatory considerations by providing engineered RNA-guided nucleases that exhibit improved specificity for targeting DNA sequences, for example, as compared to wild-type nucleases. The improved specificity can be, for example, (i) an increase in on-target binding, cleavage, and / or editing of DNA, and / or (ii) a decrease in off-target binding, cleavage, and / or editing of DNA as compared to, for example, wild-type RNA-guided nucleases and / or another mutant nuclease.
[0008] In one aspect, the present disclosure provides an isolated SPCas9 polypeptide comprising an amino acid substitution at one or more of the following positions: D23, D1251, Y128, T67, N497, R661, Q695, and / or Q926 as compared to wild-type Staphylococcus pyogenes Cas9 (SPCas9). In some embodiments, the isolated SPCas9 polypeptide comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of SEQ ID NO: 13, and the polypeptide comprises an amino acid substitution at one or more of the following positions of SEQ ID NO: 13: D23, D1251, Y128, T67, N497, R661, Q695, and / or Q926.
[0009] In some embodiments, the isolated polypeptide comprises one or more of the following amino acid substitutions: D23A, Y128V, T67L, N497A, D1251G, R661A, Q695A, and / or Q926A. In some embodiments, the isolated polypeptide comprises the following amino acid substitutions: D23A / Y128V / D1251G / T67L.
[0010] In another aspect, the present disclosure provides a fusion protein comprising an isolated polypeptide described herein fused to a heterologous functional domain, optionally via a linker that does not interfere with the activity of the fusion protein. In some embodiments, the heterologous functional domain is selected from the group consisting of: VP64, NF-kappaB p65, Krueppel-associated box (KRAB) domain, ERF repressor domain (ERD), mSin3A interaction domain (SID), heterochromatin protein 1 (HP1), DNA methyltransferase (DNMT), TET protein, histone acetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferase (HMT), or histone demethylase (HDM), MS2, Csy4, lambda N protein, and FokI.
[0011] In another aspect, the present disclosure features a genome editing system comprising an isolated polypeptide described herein.
[0012] In another aspect, the present disclosure features a nucleic acid encoding an isolated polypeptide described herein. In another aspect, the present disclosure features a vector comprising the nucleic acid.
[0013] In another aspect, the present disclosure features a composition comprising an isolated polypeptide described herein, a genome editing system described herein, a nucleic acid described herein, and / or a vector described herein, and optionally a pharmaceutically acceptable carrier.
[0014] In another aspect, the present disclosure features a method of altering a cell, the method comprising contacting the cell with such a composition. In another aspect, the present disclosure features a method of treating a patient, the method comprising administering such a composition to the patient.
[0015] In another aspect, the present disclosure features a polypeptide comprising an amino acid sequence that is at least about 80% identical (e.g., at least about 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% identical) to SEQ ID NO: 13 and having an amino acid substitution at one or more of positions D23, T67, Y128, and D1251 of SEQ ID NO: 13. In some embodiments, the polypeptide comprises an amino acid substitution at D23. In some embodiments, the polypeptide comprises an amino acid substitution at D23 and at least one amino acid substitution at T67, Y128, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at D23 and at least two substitutions at T67, Y128, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at T67. In some embodiments, the polypeptide comprises an amino acid substitution at T67 and at least one amino acid substitution at D23, Y128, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at T67 and at least two substitutions at D23, Y128, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at Y128. In some embodiments, the polypeptide comprises an amino acid substitution at Y128 and at least one amino acid substitution at D23, T67, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at Y128 and at least two substitutions at D23, T67, or D1251. In some embodiments, the polypeptide comprises an amino acid substitution at D1251. In some embodiments, the polypeptide comprises a substitution at D1251 and at least one amino acid substitution at D23, T67, or Y128. In some embodiments, the polypeptide comprises an amino acid substitution at D1251 and at least two substitutions at D23, T67, or Y128. In some embodiments, the polypeptide comprises amino acid substitutions at D23, T67, Y128, and D1251. In some embodiments, the polypeptide further comprises at least one additional amino acid substitution described herein.
[0016] In some embodiments, when the polypeptide contacts the target double-stranded DNA (dsDNA), the off-target editing rate is lower than the observed off-target editing rate of the target by wild-type SPCas9. In some embodiments, the off-target editing rate by the polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% lower than that by wild-type SPCas9. In some embodiments, the off-target editing rate is measured by evaluating the level (e.g., fraction or percentage) of indels at the off-target site.
[0017] In another aspect, the present disclosure features a polypeptide and a fusion protein comprising one or more of a nuclear localization sequence, a cell membrane permeable peptide sequence, and / or an affinity tag. In another aspect, the present disclosure features a fusion protein comprising a polypeptide fused to a heterologous functional domain, optionally via a linker that does not interfere with the activity of the fusion protein.
[0018] In some embodiments, the heterologous functional domain is a transcriptional transactivation domain. In some embodiments, the transcriptional transactivation domain is derived from VP64 or NFk-B p65. In some embodiments, the heterologous functional domain is a transcriptional silencer or transcriptional repression domain. In some embodiments, the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, an ERF repressor domain (ERD), or an mSin3A interaction domain (SID). In some embodiments, the transcriptional silencer is heterochromatin protein 1 (HP1). In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA (e.g., DNA methyltransferase (DNMT) or TET protein). In some embodiments, the TET protein is TETI. In some embodiments, the heterologous functional domain is an enzyme that modifies histone subunits (e.g., histone acetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferase (HMT), or histone demethylase). In some embodiments, the heterologous functional domain is a biological tether (e.g., MS2, Csy4, or lambda N protein). In some embodiments, the heterologous functional domain is Fokl.
[0019] In another aspect, the disclosure features an isolated nucleic acid encoding a polypeptide described herein. In another aspect, the disclosure features a vector comprising such an isolated nucleic acid. In another aspect, the disclosure features a host cell comprising such a vector.
[0020] In another aspect, the present disclosure features a polypeptide that comprises an amino acid sequence that is at least about 80% identical (e.g., at least about 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% identical) to SEQ ID NO: 13 and that comprises one or more of the following amino acid substitutions: D23A, T67L, Y128V, and D1251G. In some embodiments, the polypeptide comprises D23A; and / or the polypeptide comprises D23A and at least one of T67L, Y128V, and D1251G; and / or the polypeptide comprises D23A and at least two of T67L, Y128V, and D1251G; and / or the polypeptide comprises T67L; and / or the polypeptide comprises T67L and at least one of D23A, Y128V, and D1251G; and / or the polypeptide comprises T67L and at least two of D23A, Y128V, and D1251G; and / or the polypeptide comprises Y128V; and / or the polypeptide comprises Y128V and at least one of D23A, T67L, and D1251G; and / or the polypeptide comprises Y128V and at least two of D23A, T67L, and D1251G; and / or the polypeptide comprises D1251G; and / or the polypeptide comprises D1251G and at least one of D23A, T67L, and Y128V; and / or the polypeptide comprises D1251G and at least two of D23A, T67L, and Y128V; and / or the polypeptide comprises D23A, T67L, Y128V, and D1251G.
[0021] In some embodiments, when the polypeptide contacts a double-stranded DNA (dsDNA) target, the off-target editing rate is lower than the observed off-target editing rate of the target by wild-type SPCas9. In some embodiments, the off-target editing rate by the polypeptide is about 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% lower than that by wild-type SPCas9. In some embodiments, the off-target editing rate is measured by evaluating the level (e.g., ratio or percentage) of indels at the off-target site.
[0022] In another aspect, the present disclosure features a polypeptide and a fusion protein comprising one or more of a nuclear localization sequence, a cell membrane permeable peptide sequence, and / or an affinity tag.
[0023] In another aspect, the present disclosure features a fusion protein comprising a polypeptide that fuses to a heterologous functional domain, optionally via a linker, wherein the linker does not interfere with the activity of the fusion protein. In some embodiments, the heterologous functional domain is a transcriptional transactivation domain (e.g., a transactivation domain derived from VP64 or NFk-B p65). In some embodiments, the heterologous functional domain is a transcriptional silencer or transcriptional repressor domain (e.g., a Krueppel-associated box (KRAB) domain, an ERF repressor domain (ERD), or an mSin3A interaction domain (SID)). In some embodiments, the transcriptional silencer is heterochromatin protein 1 (HP1). In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA (e.g., a DNA methyltransferase (DNMT) or a TET protein). In some embodiments, the TET protein is TETI. In some embodiments, the heterologous functional domain is an enzyme that modifies histone subunits (e.g., a histone acetyltransferase (HAT), a histone deacetylase (HDAC), a histone methyltransferase (HMT), or a histone demethylase). In some embodiments, the heterologous functional domain is a biological tether (e.g., MS2, Csy4, or lambda N protein). In some embodiments, the heterologous functional domain is Fokl.
[0024] In another aspect, the present disclosure features an isolated nucleic acid encoding such a polypeptide. In another aspect, the present disclosure features a vector comprising such an isolated nucleic acid. In another aspect, the present disclosure features a host cell comprising such a vector.
[0025] In another aspect, the present disclosure is a method of genetically engineering a population of cells, comprising altering the genomes of at least a plurality of cells by expressing or contacting the cells with a polypeptide of the present disclosure (e.g., a mutant nuclease described herein) and a guide nucleic acid having a region complementary to a target sequence on a target nucleic acid of the genome of the cells.
[0026] In some embodiments, the off-target editing rate by the polypeptide is lower than the observed off-target editing rate of the target sequence by wild-type SPCas9. In some embodiments, the off-target editing rate by the polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% lower than that by wild-type SPCas9. In some embodiments, the off-target editing rate is measured by assessing the level (e.g., ratio or percentage) of indels at the off-target site.
[0027] In some embodiments, the polypeptide and the guide nucleic acid are administered as a ribonucleoprotein (RNP). In some embodiments, the RNP is administered at a dose of 1×10 -4 μM to 1 μM RNP.
[0028] In another aspect, the present disclosure is a method of editing a population of double-stranded DNA (dsDNA) molecules, comprising editing a plurality of dsDNA molecules by contacting the dsDNA molecules with a polypeptide of the present disclosure (e.g., a mutant nuclease described herein) and a guide nucleic acid having a region complementary to a target sequence of the dsDNA molecules.
[0029] In some embodiments, the off-target editing rate by the polypeptide is lower than the observed off-target editing rate of the target sequence by wild-type SPCas9. In some embodiments, the off-target editing rate by the polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% lower than that by wild-type SPCas9. In some embodiments, the off-target editing rate is measured by evaluating the level (e.g., ratio or percentage) of indels at the off-target site.
[0030] In some embodiments, the polypeptide and the guide nucleic acid are administered as a ribonucleoprotein (RNP). In some embodiments, the RNP is administered at a dose of 1×10 -4 μM to 1 μM RNP. BRIEF DESCRIPTION OF THE DRAWINGS
[0031]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 14C
Figure 15
Figure 16
Figure 17
Figure 18
[0032] Definitions Throughout this specification, several terms defined in the following paragraphs are used. Other definitions are also referred to within the body of this specification.
[0033] As used herein, the terms “about” and “approximately” with respect to a number are used herein to include numbers within 20%, 10%, 5%, or 1% in either direction (greater than or less than) of that number, unless otherwise stated in this specification or otherwise apparent from the context (except where such numbers would exceed 100% of the possible values).
[0034] As used herein, the term “cleavage” refers to the breaking of the covalent backbone of a DNA molecule. “Cleavage,” as used herein, refers to the breaking of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two separate single-strand cleavage events. DNA cleavage can result in the production of either blunt ends or sticky ends.
[0035] As used herein, “conservative substitution” refers to an amino acid substitution that occurs between amino acids within the following groups: i) methionine, isoleucine, leucine, valine, ii) phenylalanine, tyrosine, tryptophan, iii) lysine, arginine, histidine, iv) alanine, glycine, v) serine, threonine, vi) glutamine, asparagine, and vii) glutamic acid, aspartic acid. In some embodiments, a conservative amino acid substitution refers to an amino acid substitution that does not change the relative charge or size characteristics of the protein in which the amino acid substitution occurred.
[0036] As used herein, the term "fusion protein" refers to a protein resulting from the binding of two or more originally distinct proteins or portions thereof. In some embodiments, a linker or spacer will be present between each protein.
[0037] As used herein, the term "heterologous" with respect to a polypeptide domain refers to the fact that the polypeptide domains do not occur together naturally (e.g., within the same polypeptide). For example, within a fusion protein created by human hand, a polypeptide domain from one polypeptide may be fused to a polypeptide domain from a different polypeptide. Since the two polypeptide domains do not occur together naturally, they will be considered "heterologous" to each other.
[0038] As used herein, the term "host cell" is a cell that is engineered according to the present invention, e.g., into which a nucleic acid is introduced. A "transformed host cell" is a cell that has undergone transformation to take up an exogenous substance such as exogenous genetic material, e.g., an exogenous nucleic acid.
[0039] As used herein, the term "identity" refers to the overall correlation between polymer molecules, such as between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polymer molecules are considered to be "substantially identical" to each other if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. The calculation of percent identity between two nucleic acid sequences or polypeptide sequences can be performed, for example, by aligning the two sequences so as to optimize the comparison (e.g., gaps may be introduced into one or both of the first and second sequences so as to optimize the alignment, and non-identical sequences may be ignored for the comparison). In certain embodiments, the length of the sequences aligned for comparison is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100% of the length of the reference sequence. Next, the nucleotides at the corresponding positions are compared. The comparison of sequences and the determination of percent identity between two sequences can be accomplished using mathematical algorithms. As is well known in the art, amino acid sequences or nucleic acid sequences can be compared using a variety of algorithms, including those available in commercial computer programs such as BLASTN and BLASTP for nucleotide sequences, gapped BLAST, and PSI-BLAST for amino acid sequences.Such exemplary programs are described in Altschul, et al., Basic local alignment search tool, J. Mol. Biol., 215(3):403-410, 1990; Altschul, et al., Methods in Enzymology; Altschul et al., Nucleic Acids Res. 25:3389-3402, 1997; Baxevanis et al., Bioinformatics: A Practical Guide to the Analysis of Genes and Proteins, Wiley, 1998; and Misener, et al., (eds.), Bioinformatics Methods and Protocols (Methods in Molecular Biology, Vol. 132), Humana Press, 1999.
[0040] As used herein in the context of polynucleotides, the term "library" refers to a population of two or more different polynucleotides. In some embodiments, the library comprises at least two polynucleotides comprising different sequences encoding nucleases, and / or at least two polynucleotides comprising different sequences encoding guide RNAs. In some embodiments, the library is at least 10 1 at least 10 2 at least 10 3 at least 10 4 at least 10 5 at least 10 6 at least 10 7 at least 10 8 at least 10 9 at least 10 10 at least 10 11 at least 10 12 at least 10 13 at least 10 14 or at least 10 15comprises a plurality of different polynucleotides. In some embodiments, the members of the library may comprise a randomized sequence, e.g., a completely or partially randomized sequence. In some embodiments, the library comprises nucleic acids comprising polynucleotides that are unrelated to each other, e.g., a completely randomized sequence. In other embodiments, at least some members of the library may be related, e.g., variants or derivatives of a particular sequence.
[0041] As used herein, the term "operably linked" refers to juxtaposition where the components so described are in a relationship that allows them to function as intended. A regulatory element "operably linked" to a functional element is related such that the expression and / or activity of the functional element is achieved under conditions compatible with the regulatory element. In some embodiments, an "operably linked" regulatory element is contiguous with the coding element of interest (e.g., covalently linked); in some embodiments, the regulatory element acts in trans to or at / from the functional element of interest.
[0042] As used herein, the term "nuclease" refers to a polypeptide capable of cleaving the phosphodiester bond between nucleotide subunits of a nucleic acid; the term "endonuclease" refers to a polypeptide capable of cleaving a phosphodiester bond within a polynucleotide chain.
[0043] As used herein, the terms "nucleic acid", "nucleic acid molecule", or "polynucleotide" are used interchangeably herein. These refer to polymers of deoxyribonucleotides or ribonucleotides in single-stranded or double-stranded form, and include known analogs of natural nucleotides that can function in a manner similar to natural nucleotides unless otherwise described. The term includes nucleic acid-like structures having synthetic backbones, as well as amplification products. DNA and RNA are both polynucleotides. This polymer can include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0044] As used herein, the term "oligonucleotide" refers to a chain of nucleotides or nucleotide analogs. Oligonucleotides can be obtained by several methods, including, for example, chemical synthesis, restriction enzyme digestion, or PCR. As recognized by those skilled in the art, the length of an oligonucleotide (i.e., the number of nucleotides) can vary widely depending on the intended function or use of the oligonucleotide. Throughout this specification, whenever an oligonucleotide is represented by a sequence of letters (selected from the four-letter alphabet: A, C, G, and T, which represent adenosine, cytidine, guanosine, and thymidine, respectively), the nucleotides are shown in the 5' to 3' order from left to right. In certain embodiments, the sequence of an oligonucleotide includes one or more degenerate residues as described herein.
[0045] As used herein, the term "off-target" refers to the unintended or unexpected binding, cleavage, and / or editing of DNA by an RNA-guided nuclease. In some embodiments, a region of DNA is an off-target region if it is different from the region of DNA that is intended or expected to be bound, cleaved, and / or edited by 1, 2, 3, 4, 5, 6, 7, or more nucleotides.
[0046] As used herein, the term "on-target" refers to the intended or expected binding, cleavage, and / or editing of DNA by an RNA-guided nuclease.
[0047] As used herein, the term "polypeptide" generally has its recognized meaning in the art as a polymer of amino acids. The term is also used to refer to specific functional classes of polypeptides, such as nucleases, antibodies, etc.
[0048] As used herein, the term "regulatory element" refers to a DNA sequence that controls or affects one or more aspects of gene expression. In some embodiments, the regulatory element is, or includes, a promoter, enhancer, silencer, and / or termination signal. In some embodiments, the regulatory element controls or affects inducible expression.
[0049] As used herein, the term "target site" refers to a nucleic acid sequence that defines a portion of a nucleic acid to which a binding molecule will bind, provided that conditions are sufficient for binding to exist. In some embodiments, the target site is a nucleic acid sequence to which the nucleases described herein bind and / or are cleaved by such nucleases. In some embodiments, the target site is a nucleic acid sequence to which the guide RNA described herein binds. The target site can be single-stranded or double-stranded. In the context of a dimerizing nuclease, such as a nuclease comprising the Fokl DNA cleavage domain, the target site typically includes a left half-site (to which one monomer of the nuclease binds), a right half-site (to which the second monomer of the nuclease binds), and a spacer sequence between the half-sites at which cleavage occurs. In some embodiments, the left half-site and / or the right half-site are 10 to 18 nucleotides in length. In some embodiments, either or both of the half-sites are shorter or longer. In some embodiments, the left half-site and the right half-site contain different nucleic acid sequences. In the context of zinc finger nucleases, the target site can, in some embodiments, include two half-sites, each 6 to 18 nucleotides in length, flanking an unspecified spacer region that is 4 to 8 bp in length. In the context of TALENs, the target site can, in some embodiments, include two half-sites, each 10 to 23 bp in length, flanking an unspecified spacer region that is 10 to 30 bp in length. In the context of an RNA-guided (e.g., RNA-programmable) nuclease, the target site typically includes a nucleotide sequence complementary to the guide RNA of the RNA-programmable nuclease, and a protospacer adjacent motif (PAM) at the 3' or 5' end flanking the guide RNA complementary sequence. For the RNA-guided nuclease Cas9, the target site can, in some embodiments, be 16 to 24 base pairs + 3 to 6 base pairs (PAM) (e.g., NNN, where N corresponds to any nucleotide).Exemplary target sites for RNA-guided nucleases such as Cas9 are known to those of skill in the art and include, without limitation, NNG, NGN, NAG, NGA, NGG, NGAG, and NGCG (where N corresponds to any nucleotide). Further, Cas9 nucleases from different species (e.g., Streptococcus thermophilus instead of Streptococcus pyogenes) recognize a PAM containing the sequence NGGNG. Additional PAM sequences are known, including, without limitation, NNAGAAW and NAAR (see, e.g., Esvelt and Wang, Molecular Systems Biology, 9:641 (2013), the entire contents of which are incorporated herein by reference). For example, the target site for an RNA-guided nuclease such as Cas9 can include the structure [Nz]-[PAM] (where each N is independently any nucleotide and z is an integer from 1 to 50). In some embodiments, z is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50. In some embodiments, z is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50. In some embodiments, Z is 20.
[0050] As used herein, the term "variant" refers to an entity that exhibits significant structural identity with a reference entity (e.g., a wild-type sequence), but whose structure differs from the reference entity in the presence or level of one or more chemical moieties. In many embodiments, the variant also differs functionally from its reference entity. Generally, whether a particular entity is appropriately considered a "variant" of a reference entity is based on the degree of structural identity with the reference entity. As will be understood by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. By definition, a variant is a different chemical entity that shares one or more of such characteristic structural elements. To give a few examples, a polypeptide may contain a plurality of amino acids whose positions are specified relative to each other in linear or three-dimensional space and / or may have characteristic sequence elements that contribute to a particular biological function; a nucleic acid may have characteristic sequence elements that contain a plurality of nucleotide residues whose positions are specified relative to each other in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide as a result of one or more differences in the amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids) covalently attached to the polypeptide backbone. In some embodiments, the variant polypeptide exhibits an overall sequence identity with the reference polypeptide (e.g., a nuclease described herein) of at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. Alternatively, or in addition, in some embodiments, the variant polypeptide does not share at least one characteristic sequence element with the reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, the variant polypeptide shares one or more biological activities of the reference polypeptide, such as nuclease activity. In some embodiments, the variant polypeptide lacks one or more of the biological activities of the reference polypeptide.In some embodiments, the mutant polypeptide exhibits a decrease in the level of one or more biological activities (e.g., nuclease activity, e.g., off-target nuclease activity) as compared to a reference polypeptide. In some embodiments, if a polypeptide of interest has an amino acid sequence that is identical to the parental amino acid sequence but has a few sequence changes at specific positions, the polypeptide of interest is considered a “variant” of the parental polypeptide or the reference polypeptide. Typically, less than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% of the residues are substituted in the variant as compared to the parent. In some embodiments, the variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 residue substituted as compared to the parent. Often, the variant has a very small number (e.g., 5, 4, 3, 2, or less than 1) of substituted functional residues (i.e., residues involved in a particular biological activity). In some embodiments, the variant has an addition or deletion that is 5, 4, 3, 2, or 1 or less as compared to the parent, and often there is no addition or deletion. Further, any addition or deletion is typically less than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6 residues, and generally less than about 5, about 4, about 3, or about 2 residues. In some embodiments, the parental polypeptide or the reference polypeptide is one found in nature.
[0051] Overview The present disclosure, in part, encompasses the discovery of RNA-guided nucleases that target DNA sequences and exhibit improved specificity, e.g., over wild-type nucleases. Provided herein are such RNA-guided nuclease variants, compositions, and systems that include such nuclease mutants, and methods of making and using such nuclease mutants, e.g., to edit one or more target nucleic acids.
[0052] RNA-guided nucleases RNA-guided nucleases according to the present disclosure include, but are not limited to, class 2 CRISPR nucleases of natural origin, such as Cas9 and Cpf1, and other nucleases derived from or obtained therefrom. For example, other nucleases derived from or obtained therefrom include mutant nucleases. In some embodiments, the mutant nuclease comprises an alteration of one or more enzymatic properties (e.g., an alteration of nuclease activity or helicase activity) as compared to a naturally occurring or other reference nuclease molecule (including a nuclease molecule that has already been engineered or modified). In some embodiments, the mutant nuclease can have nickase activity or non-cleavage activity (opposite to double-stranded nuclease activity). In another embodiment, the mutant nuclease has an alteration that changes its size, such as a deletion of an amino acid sequence that decreases its size, but the effect on one or more, or any, nuclease activity may or may not be significant. In another embodiment, the mutant nuclease can recognize a different PAM sequence. In some embodiments, the different PAM sequence is a PAM sequence other than that recognized by the endogenous wild-type PI domain of the reference nuclease, such as a non-canonical sequence.
[0053] RNA-guided nucleases, in terms of function, are defined as nucleases that (a) interact (e.g., form a complex) with a gRNA; and (b) together with the gRNA, bind to, and optionally cleave or modify, a target region of DNA that includes (i) a sequence complementary to the targeting domain of the gRNA and optionally (ii) an additional sequence called a "protospacer adjacent motif" or "PAM" (which is described in more detail below). RNA-guided nucleases can be broadly defined by their PAM specificity and cleavage activity, although there may be differences between individual RNA-guided nucleases having the same PAM specificity or cleavage activity. Those skilled in the art will recognize that some aspects of the present disclosure relate to systems, methods, and compositions that can be practiced using any suitable RNA-guided nuclease having a certain PAM specificity and / or cleavage activity. For this reason, unless specified otherwise, the term RNA-guided nuclease should be understood as an inclusive term and is not limited to any particular type (e.g., Cas9 versus Cpf1), species (e.g., Streptococcus pyogenes versus Staphylococcus aureus), or difference (e.g., full length versus truncated or fragmented; native origin PAM specificity versus engineered PAM specificity, etc.) of RNA-guided nucleases.
[0054] The PAM sequence is named for its contiguous relationship to the "protospacer" sequence that is complementary to the gRNA targeting domain (or "spacer"). Together with the protospacer sequence, the PAM sequence defines the target region or sequence for a particular RNA-guided nuclease / gRNA combination.
[0055] Various RNA-guided nucleases may require different contiguous relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence that is 3' of the protospacer visualized relative to the guide RNA targeting domain.
[0056] On the other hand, Cpf1 usually recognizes the PAM sequence that is 5' of the protospacer.
[0057] In addition to recognizing specific continuous directions of the PAM and protospacer, RNA-guided nucleases can also recognize specific PAM sequences. For example, Staphylococcus aureus (S. aureus) Cas9 recognizes PAM sequences of NNGRRT or NNGRRV (where the N residue is 3' adjacent to the region recognized by the gRNA targeting domain). Streptococcus pyogenes (S. pyogenes) Cas9 recognizes the NGG PAM sequence. Also, Francisella novicida (F. novicida) Cpf1 recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-guided nucleases, and strategies for identifying novel PAM sequences are described by Shmakov et al., 2015, Molecular Cell 60, 385-397, November 5, 2015. It should also be noted that engineered RNA-guided nucleases may have PAM specificities different from those of reference molecules (e.g., in the case of an engineered RNA-guided nuclease, the reference molecule can be a naturally occurring variant from which the RNA-guided nuclease is derived, or a naturally occurring variant with the highest amino acid sequence homology to the engineered RNA-guided nuclease).
[0058] In addition to its PAM specificity, an RNA-guided nuclease can be characterized by its DNA cleavage activity: naturally occurring RNA-guided nucleases typically form DSBs in the target nucleic acid, but result in only SSBs (see Ran & Hsu, et al., Cell 154(6), 1380-1389, September 12, 2013 ("Ran")), which is incorporated herein by reference, or engineered variants that do not cleave at all have been produced.
[0059] Cas9 The crystal structures of Streptococcus pyogenes (S. pyogenes) Cas9 (Jinek et al., Science 343(6176), 1247997, 2014 (“Jinek 2014”)) and Staphylococcus aureus (S. aureus) Cas9 in complex with single-molecule guide RNA and target DNA (Nishimasu 2014; Anders et al., Nature. 2014 Sep 25;513(7519):569-73 (“Anders 2014”); and Nishimasu 2015) have been determined.
[0060] Naturally occurring Cas9 proteins contain two lobes: a recognition (REC) lobe and a nuclease (NUC) lobe; each of these contains specific structural and / or functional domains. The REC lobe contains an arginine-rich bridging helix (BH) domain and at least one REC domain (e.g., the REC1 domain and, optionally, the REC2 domain). The REC lobe has no structural similarity to other known proteins, suggesting that this is a unique functional domain. Without wishing to be bound by any theory, mutational analysis suggests specific functional roles for the BH and REC domains: the BH domain appears to play a role in gRNA:DNA recognition, whereas the REC domain is thought to interact with the repeat:anti-repeat duplex of the gRNA and mediate the formation of the Cas9 / gRNA complex.
[0061] The NUC lobe contains an RuvC domain, an HNH domain, and a PAM interacting (PI) domain. The RuvC domain has structural similarity to members of the retroviral integrase superfamily and cleaves the non-complementary (i.e., bottom) strand of the target nucleic acid. This may be formed from two or more split RuvC motifs (such as RuvC I, RuvCII, and RuvCIII in Streptococcus pyogenes and Staphylococcus aureus). On the other hand, the HNH domain is structurally similar to the HNN endonuclease motif and cleaves the complementary (i.e., top) strand of the target nucleic acid. The PI domain contributes to PAM specificity, as its name suggests.
[0062] Certain functions of Cas9 are related to the specific domains described above (although not necessarily fully determined thereby), but these functions and other functions may be mediated or affected by other Cas9 domains or by multiple domains on either lobe. For example, in Streptococcus pyogenes Cas9, as described in Nishimasu 2014, the gRNA repeat:anti-repeat duplex enters the groove between the REC and NUC lobes, and the nucleotides within this duplex interact with amino acids in the BH, PI, and REC domains. Some nucleotides within the first stem-loop structure also interact with amino acids in multiple domains (PI, BH, and REC1), and some nucleotides within the second and third stem-loops are similar (RuvC and PI domains).
[0063] Mutant Cas9 nuclease The present disclosure includes, for example, mutant RNA-guided nucleases with an increased level of specificity for a target relative to a wild-type nuclease. For example, the mutant RNA-guided nucleases of the present disclosure exhibit an increased level of on-target binding, editing, and / or cleavage activity compared to a wild-type nuclease. Additionally, or alternatively, the mutant RNA-guided nucleases of the present disclosure exhibit a decreased level of off-target binding, editing, and / or cleavage activity compared to a wild-type nuclease.
[0064] The mutant nucleases described herein include variants of Streptococcus pyogenes (S. pyogenes) Cas9 and Neisseria meningitidis (N. meningitidis) (SEQ ID NO: 14). The amino acid sequence of wild-type Streptococcus pyogenes (S. pyogenes) Cas9 is described as SEQ ID NO: 13. The mutant nuclease may include amino acid substitutions at a single position or multiple positions relative to the wild-type nuclease, for example, at 2, 3, 4, 5, 6, 7, 8, 9, 10, or more positions. In some embodiments, the mutant nuclease includes an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the wild-type nuclease.
[0065] One or more wild-type amino acids may be substituted with alanine. Additionally, or alternatively, one or more wild-type amino acids may be substituted with a conservative mutant amino acid. Additionally, or alternatively, one or more wild-type amino acids may be substituted with a non-conservative mutant amino acid.
[0066] For example, the mutant nuclease described in this specification may include substitutions at all of positions 1, 2, 3, 4, 5, 6, 7, or 8 relative to the wild-type nuclease (e.g., SEQ ID NO: 13): D23, D1251, Y128, T67, N497, R661, Q695, and / or Q926 (e.g., conservative and / or non-conservative substitutions of alanine at one or all of these positions). Exemplary mutant nucleases may include all of the following substitutions at positions 1, 2, 3, 4, 5, 6, 7, or 8 relative to the wild-type nuclease: D23A, D1251G, Y128V, T67L, N497A, R661A, Q695A, and / or Q926A. A specific nuclease variant of the present disclosure includes the following substitutions relative to the wild-type nuclease: D23A, D1251G, Y128V, and T67L.
[0067] In some embodiments, the mutant nuclease includes an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13 and includes all of the following substitutions at positions 1, 2, 3, 4, 5, 6, 7, or 8: D23A, D1251G, Y128V, T67L, N497A, R661A, Q695A, and / or Q926A. For example, the mutant nuclease may include an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13 and includes the following substitutions: D23A, D1251G, Y128V, and T67L.
[0068] In addition to substitutions at 1, 2, 3, 4, 5, 6, 7, or all 8 of D23, D1251, Y128, T67, N497, R661, Q695, and / or Q926 (e.g., D23A, D1251G, Y128V, and T67L), the Streptococcus pyogenes Cas9 variant may also include substitutions at one or more of the following positions: L169; Y450; M495; W659; M694; H698; A728; E1108; V1015; R71; Y72; R78; R165; R403; T404; F405; K1107; S1109; R1114; S1116; K1118; D1135; S1136; K1200; S1216; E1219; R1333; R1335; T1337; Y72; R75; K76; L101; S104; F105; R115; H116; I135; H160; K163; Y325; H328; R340; F351; D364; Q402; R403; I1110; K1113; R1122; Y1131; R63; R66; R70; R71; R74; R78; R403; T404; N407; R447; I448; Y450; K510; Y515; R661; V1009; Y1013; K30; K33; N46; R40; K44; E57; T62; R69; N77; L455; S460; R467; T472; I473; H721; K742; K1097; V1100; T1102; F1105; K1123; K1124; E1225; Q1272; H1349; S1351; and / or Y1356, e.g., the substitutions described in U.S. Patent No. 9,512,446.
[0069] In some embodiments, the Streptococcus pyogenes variant may include substitutions at one or more of the following positions: N692, K810, K1003, R1060, and G1218. In some embodiments, the Streptococcus pyogenes variant includes one or more of the following substitutions: N692A, K810A, K1003A, R1060A, and G1218R.
[0070] Table 1 lists exemplary Streptococcus pyogenes (S. pyogenes) Cas9 mutants containing 3 to 5 substitutions according to certain embodiments of the present disclosure. For clarity, the present disclosure encompasses Cas9 mutant proteins having mutations at 1, 2, 3, 4, 5, or more of the sites shown hereinbefore in the present disclosure and in other sections. Exemplary triple mutants, quadruple mutants, and quintuple mutants are shown in Table 1 and are described, for example, in Chen et al., Nature 550:407-410 (2017); Slaymaker et al., Science 351:84-88 (2015); Kleinstiver et al., Nature 529:490-495; Kleinstiver et al., Nature 523:481-485 (2015); Kleinstiver et al., Nature Biotechnology 33:1293-1298 (2015).
[0071] [Table 1]
[0072] Examples of Streptococcus pyogenes (S. pyogenes) Cas9 variants that reduce or abolish the nuclease activity of Cas9 include one or more amino acid substitutions at: D10, E762, D839, H983, or D986, and H840 or N863. For example, Streptococcus pyogenes (S. pyogenes) Cas9 may include the amino acid substitutions D10A / D10N and H840A / H840N / H840Y that catalytically inactivate the nuclease portion of the protein. Substitutions at these positions may be alanine or other residues, such as glutamine, asparagine, tyrosine, serine, or aspartate, for example, E762Q, H983N, H983Y, D986N, N863D, N863S, or N863H (Nishimasu et al., Cell 156, 935-949 (2014); WO 2014 / 152432). In some embodiments, the variant includes a single amino acid substitution of D10A or H840A, which results in a single-stranded nickase enzyme. In some embodiments, the mutant polypeptide includes the amino acid substitutions of D10A and H840A, which inactivate the nuclease activity (e.g., known as dead Cas9 or dCas9). Also included as mutant nucleases described herein are variants of Neisseria meningitidis (N. meningitidis) (Hou et al., PNAS Early Edition 2013, 1-6; incorporated herein by reference). The amino acid sequence of wild-type Neisseria meningitidis (N. meningitidis) Cas9 is set forth as SEQ ID NO: 14. Comparison of the Cas9 sequences of Neisseria meningitidis (N. meningitidis) and Streptococcus pyogenes (S. pyogenes) shows that certain regions are conserved (see WO 2015 / 161276). Thus, the disclosure includes Neisseria meningitidis (N. meningitidis) Cas9 variants that include one or more of the substitutions described herein in the context of Streptococcus pyogenes (S. pyogenes) Cas9, for example, at one or more corresponding amino acid positions of Neisseria meningitidis (N. meningitidis) Cas9.For example, substitutions at amino acid positions D29, D983, L101, S66, Q421, E459, and Y671 of Neisseria meningitidis (N. meningitidis) correspond to amino acid positions D23, D1251, Y128, T67, R661, Q695, and / or Q926 of Streptococcus pyogenes (S. pyogenes), respectively (Figure 17).
[0073] The mutant Neisseria meningitidis (N. meningitidis) nuclease may contain amino acid substitutions at a single position or at multiple positions, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more positions, relative to the wild-type nuclease. In some embodiments, the mutant Neisseria meningitidis (N. meningitidis) nuclease comprises an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the wild-type nuclease.
[0074] The mutant nuclease retains one or more functional activities of the wild-type nuclease, such as the ability to cleave double-stranded DNA, the ability to cleave single-stranded DNA (e.g., a nickase), the ability to target DNA but not cleave it (e.g., a dead nuclease), and / or the ability to interact with a guide nucleic acid. In some embodiments, the mutant nuclease has the same or approximately the same level of on-target activity as the wild-type nuclease. In some embodiments, the mutant nuclease has one or more functional activities that are improved compared to the wild-type nuclease. For example, the mutant nucleases described herein can exhibit an increase in the level of on-target activity (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 125%, 150%, 175%, 200%, or more relative to the wild-type) and / or a decrease in the level of off-target activity (e.g., 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the wild-type activity). In some embodiments, when the mutant nucleases described herein contact double-stranded DNA (dsDNA) (e.g., target dsDNA), the off-target editing (e.g., off-target editing rate) is lower than the observed or measured off-target editing rate of the target dsDNA by the wild-type nuclease. For example, the off-target editing rate by the mutant nuclease can be about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% lower than that by the wild-type nuclease.
[0075] The activity of the mutant nuclease (e.g., on-target activity and / or off-target activity) can be evaluated using any method known in the art, such as GUIDE-seq (see, e.g., Tsai et al., Nat. Biotechnol. 33:187-197 (2015)); CIRCLE-seq (see, e.g., Tsai et al., Nature Methods 14:607-614 (2017)); Digenome-seq (see, e.g., Kim et al., Nature Methods 12:237-243 (2015)); or ChIP-seq (see, e.g., O’Geen et al., Nucleic Acids Res. 43:3389-3404 (2015)). In some embodiments, the off-target editing rate is evaluated by determining the indel% at the off-target site.
[0076] As is well known to those skilled in the art, there are various methods for introducing substitutions into the amino acid sequence of a polypeptide. Nucleic acids encoding mutant nucleases can be introduced into viral or non-viral vectors for expression in host cells (e.g., human cells, animal cells, bacterial cells, yeast cells, insect cells). In some embodiments, the nucleic acid encoding the mutant nuclease is operably linked to one or more regulatory domains for expression of the nuclease. As will be understood by those skilled in the art, suitable bacterial and eukaryotic promoters are well known in the art and are described, for example, in Sambrook et al., Molecular Cloning, A Laboratory Manual (3d ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel et al., eds., 2010). Bacterial expression systems for expressing engineered proteins are available, for example, in Escherichia coli (E. coli), Bacillus species, and Salmonella (Paiva et al., 1983, Gene 22:229-235).
[0077] Cpf1 The crystal structure of the complex of crRNA with double-stranded (ds) DNA target (including the TTTN PAM sequence) of Acidaminococcus sp. Cpf1 has been elucidated by Yamano et al. (Cell. 2016 May 5;165(4):949-962(“Yamano”), which is incorporated herein by reference). Cpf1, like Cas9, has two lobes: the REC (recognition) lobe and the NUC (nuclease) lobe. The REC lobe contains the REC1 and REC2 domains that have no similarity to any known protein structure. On the other hand, the NUC lobe contains three RuvC domains (RuvC-I, -II, and -III), as well as the BH domain. However, in contrast to Cas9, the Cpf1 REC lobe lacks the HNH domain and also contains other domains that have no similarity to any known protein structure: the structurally unique PI domain, three Wedge (WED) domains (WED-I, -II, and -III), and the nuclease (Nuc) domain.
[0078] Although Cas9 and Cpf1 have similarities in structure and function, it should be recognized that certain Cpf1 activities are mediated by structural domains that are not similar to any Cas9 domain. For example, cleavage of the complementary strand of the target DNA appears to be mediated by the Nuc domain, which is both sequence- and spatially different from the HNH domain of Cas9. Furthermore, the non-targeting portion (handle) of the Cpf1 gRNA adopts a pseudoknot structure rather than the stem-loop structure formed by the repeat:anti-repeat duplex in the Cas9 gRNA.
[0079] A nucleic acid encoding an RNA-guided nuclease Nucleic acids encoding RNA-guided nucleases, such as Cas9, Cpf1, or functional fragments thereof, are provided herein. Exemplary nucleic acids encoding RNA-guided nucleases have been described previously (see, e.g., Cong et al., Science. 2013 Feb 15;339(6121):819-23 (“Cong 2013”); Wang et al., PLoS One. 2013 Dec 31;8(12):e85650 (“Wang 2013”); Mali 2013; Jinek 2012).
[0080] In some cases, the nucleic acid encoding the RNA-guided nuclease can be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule can be chemically modified. In certain embodiments, the mRNA encoding the RNA-guided nuclease will have one or more (e.g., all) of the following properties: it can be capped; it can be polyadenylated; and it can be substituted with 5-methylcytidine and / or pseudouridine.
[0081] The synthetic nucleic acid sequence can also be codon-optimized. For example, at least one non-common or less common codon has been replaced by a common codon. For example, the synthetic nucleic acid can direct the synthesis of an optimized messenger mRNA, optimized for expression, e.g., in a mammalian expression system as described herein. Examples of codon-optimized Cas9 coding sequences are shown in WO 2016 / 073990 pamphlet (“Cotta-Ramusino”).
[0082] Additionally or alternatively, the nucleic acid encoding the RNA-guided nuclease can include a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.
[0083] Guide RNA (gRNA) molecule The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates the specific binding (or "targeting") of an RNA-guided nuclease, such as Cas9 or Cpf1, to a target sequence, such as an intracellular genomic or episomal sequence. The gRNA can be single-molecule (comprising a single RNA molecule, or also referred to as chimeric) or modular (comprising two or more, typically two separate RNA molecules, such as crRNA and tracrRNA, which are usually bound to each other, for example, by duplex formation). The gRNA and its components are described throughout the literature, for example, in Briner et al. (Molecular Cell 56(2), 333-339, October 23, 2014 ("Briner")), which is incorporated by reference, and also in Cotta-Ramusino.
[0084] In bacteria and archaea, type II CRISPR systems generally include an RNA-guided nuclease protein such as Cas9, a CRISPR RNA (crRNA) containing a 5' region complementary to the foreign sequence, and a trans-activating crRNA (tracrRNA) containing a 5' region complementary to and forming a duplex with the 3' region of the crRNA. Without intending to be bound by any theory, it is thought that this duplex facilitates the formation of the Cas9 / gRNA complex - which is required for the activity of the Cas9 / gRNA complex. Since type II CRISPR systems have been adapted for use in gene editing, in one non-limiting example, it has been discovered that the crRNA and tracrRNA can be linked using a four-nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that connects the complementary regions of the crRNA (its 3' end) and tracrRNA (its 5' end) to form a single single-molecule or chimeric guide RNA. (Mali et al. Science. 2013 Feb 15;339(6121):823-826 ("Mali 2013"); Jiang et al. Nat Biotechnol. 2013 Mar;31(3):233-239 ("Jiang"); and Jinek et al., 2012 Science Aug. 17;337(6096):816-821 ("Jinek 2012") (all of which are incorporated herein by reference)).
[0085] The guide RNA, whether single molecule or modular, includes a "targeting domain" that is fully or partially complementary to a target domain within a target sequence, such as a DNA sequence within the genome of a cell in which editing is desired. The targeting domain is referred to by various names in the literature, including, but not limited to, "guide sequence" (Hsu et al., Nat Biotechnol. 2013 Sep;31(9):827-832 ("Hsu"), incorporated herein by reference), "complementary region" (Cotta-Ramusino), "spacer" (Briner), and the generic "crRNA" (Jiang). Regardless of the name given, the targeting domain typically is 10 to 30 nucleotides in length, and in certain embodiments is 16 to 24 nucleotides in length (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides in length), and is located at or near the 5' end in the case of Cas9 gRNA or at or near the 3' end in the case of Cpf1 gRNA.
[0086] In addition to the targeting domain, the gRNA typically (but not necessarily as discussed below) contains multiple domains that can affect the formation or activity of the gRNA / Cas9 complex. For example, as mentioned above, the double-stranded structure formed by the first and second complementary domains of the gRNA (also called the repeat:anti-repeat duplex) can interact with the recognition (REC) lobe of Cas9 and mediate the formation of the Cas9 / gRNA complex. (Both Nishimasu et al., Cell 156, 935-949, February 27, 2014 (“Nishimasu 2014”) and Nishimasu et al., Cell 162, 1113-1126, August 27, 2015 (“Nishimasu 2015”), which are incorporated herein by reference.) It should be noted that the first and / or second complementary domains may contain one or more poly-A tracts that may be recognized as termination signals by RNA polymerase. Thus, the sequences of the first and second complementary domains may, in some cases, be modified to remove these tracts and promote complete in vitro transcription of the gRNA, for example, through the use of A-G exchanges or A-U exchanges as described by Briner. These and other similar modifications to the first and second complementary domains are within the scope of the present disclosure.
[0087] Together with the first and second complementary domains, the Cas9 gRNA typically contains two or more additional double-stranded regions that are involved in nuclease activity in vivo, although not necessarily in vitro (Nishimasu 2015). The first stem-loop proximal to the 3’ portion of the second complementary domain is variously called the “proximal domain” (Cotta-Ramusino), “stem-loop 1” (Nishimasu 2014 and 2015), and “nexus” (Briner). Near the 3’ end of the gRNA, one or more additional stem-loop structures are generally present, and the number varies by species: the S. pyogenes gRNA typically contains two 3’ stem-loops (for a total of four stem-loop structures including the repeat:anti-repeat duplex), whereas S. aureus and other species have only one (for a total of three stem-loop structures). An account of the conserved stem-loop structures (and, more generally, gRNA structure) organized by species is provided by Briner.
[0088] The foregoing description has focused on gRNAs for use with Cas9, but it should be recognized that other RNA-guided nucleases that utilize gRNAs that differ in some respects from those described to date have been discovered or invented (or may be in the future). For example, Cpf1 (“CRISPR from Prevotella and Franciscella 1”) is a recently discovered RNA-guided nuclease that does not require a tracrRNA to function. (Zetsche et al., 2015, Cell 163, 759-771 October 22, 2015(“Zetsche I”), which is incorporated herein by reference). gRNAs for use in the Cpf1 genome editing system generally include a targeting domain and a complementary domain (also sometimes referred to as the “handle”). It should also be noted that in gRNAs for use with Cpf1, the targeting domain is typically present at or near the 3’ end rather than the 5’ end as described above for Cas9 gRNAs (the handle is at or near the 5’ end of the Cpf1 gRNA).
[0089] However, one of ordinary skill in the art will recognize that while there may be structural differences between gRNAs from different prokaryotic species or between Cpf1 gRNAs and Cas9 gRNAs, the general principles by which the gRNAs operate are consistent. From this invariance in operation, gRNAs can be broadly defined by their targeting domain sequences, and one of ordinary skill in the art will recognize that a given targeting domain sequence can be incorporated into any suitable gRNA, including single molecule or chimeric gRNAs, or gRNAs that include one or more chemical and / or sequence modifications (substitutions, additional nucleotides, truncations, etc.). Thus, for the sake of brevity of presentation in this disclosure, gRNAs can be described only with respect to their targeting domain sequences.
[0090] More generally, one of ordinary skill in the art will recognize that some aspects of the present disclosure relate to systems, methods, and compositions that can be carried out using multiple RNA-guided nucleases. For this reason, unless specified otherwise, the term gRNA should be understood to encompass not only gRNAs that are compatible with a particular species of Cas9 or Cpf1, but also any suitable gRNA that can be used with any RNA-guided nuclease. By way of example, in some embodiments, the term gRNA can include any RNA-guided nuclease present in a class 2 CRISPR system, such as a type II or type V or CRISPR system, or an RNA-guided nuclease derived from or adapted therefrom, for use with such RNA-guided nucleases.
[0091] Selection method The present disclosure also provides a competition-based selection strategy that will select the variant with the greatest fitness (e.g., between libraries of variants / mutants) under a set of conditions. The selection methods of the present invention are useful, for example, in directed evolution strategies, such as strategies that include selection following one or more rounds of mutagenesis. In certain embodiments, the methods of the present disclosure enable higher throughput directed evolution strategies than are typically observed in current polypeptide and / or polypeptide evolution strategies.
[0092] Selection based on binding to a DNA target site within a phagemid In one aspect, the present disclosure provides a method of selecting a version of a polypeptide or polynucleotide of interest based on binding to a DNA target site. The method generally comprises: (a) providing a library of polynucleotides, wherein the various polynucleotides in the library encode various versions of the polypeptide of interest or function as templates for various versions of the polynucleotide of interest; (b) introducing the library of polynucleotides into host cells, whereby each transformed host cell contains a polynucleotide that encodes a version of the polypeptide of interest or functions as a template for a version of the polynucleotide of interest; (c) providing a plurality of bacteriophages comprising a phagemid that encodes a first selection agent and contains the DNA target site; (d) incubating the transformed host cells from step (b) together with the plurality of bacteriophages under culture conditions such that the plurality of bacteriophages infect the transformed host cells, wherein expression of the first selection agent confers a survival benefit or detriment to the infected host cells; and (e) selecting the host cells that exhibit a survival benefit (e.g., the survival benefit described herein) in (d). For example, survival in step (d) is based on various schemes, as outlined below.
[0093] Binding at the DNA target site reduces the expression of a selection agent encoded by or contained on the phagemid (further discussed herein). As further discussed herein, the selection agent can confer a survival disadvantage or a survival benefit. Reduction in the expression of the selection agent can occur by transcriptional repression of the selection agent mediated by binding at the DNA target site. In some embodiments, reduction in the expression of the selection agent occurs by cleavage of the phagemid at or near the DNA target site. For example, the polypeptide or polynucleotide of interest binds to DNA and cleaves the DNA. Some classes of enzymes bind to specific DNA recognition sites and cleave the DNA at or near the binding site. Thus, the DNA cleavage site may be the same as or different from the DNA binding site. In some embodiments where the DNA cleavage site is different from the DNA binding site, the two sites are close to each other (e.g., within 20, 15, 10, or 5 base pairs). Additionally, or alternatively, the two sites are not within 20, 15, 10, or 5 base pairs of each other.
[0094] When cleavage is involved, the cleavage can be cleavage of one strand of the phagemid (referred to as "nicking") that replicates as a double-stranded plasmid inside the host cell, or cleavage of both strands, as further discussed below.
[0095] In certain embodiments, binding at the DNA target site increases the expression of a selection agent encoded by the phagemid or of a selection agent whose template is on the phagemid. The increase in the expression of the selection agent can occur, for example, by transcriptional activation of the selection agent mediated by binding at the DNA target site.
[0096] A version of a polypeptide or polynucleotide that binds to a DNA target site can be selected such that a host cell transformed with such a version exhibits maintenance of cell growth kinetics, an increase in cell growth kinetics (e.g., an increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics); while a host cell transformed with a version that does not bind to the DNA target site does not exhibit an increase in cell growth kinetics (e.g., exhibits a decrease in cell growth kinetics and / or cell division), and / or dies. In certain embodiments, a version of a polypeptide or polynucleotide that binds to a DNA target site is selected such that a host cell transformed with such a version survives, while a host cell transformed with a version that does not bind to the DNA target site does not survive.
[0097] A version of a polypeptide or polynucleotide that does not bind to a DNA target site can be selected such that a host cell transformed with such a version exhibits maintenance of cell growth kinetics, an increase in cell growth kinetics (e.g., an increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics); while a host cell transformed with a version that binds to the DNA target site does not exhibit an increase in cell growth kinetics (e.g., exhibits a decrease in cell growth kinetics and / or cell division), and / or dies. In certain embodiments, a version of a polypeptide or polynucleotide that does not bind to a DNA target site is selected such that a host cell transformed with such a version survives, while a host cell transformed with a version that binds to the DNA target site does not survive.
[0098] In alternative embodiments, a Cas9 protein, a gRNA targeting a target sequence, and a phagemid encoding a phage origin F1 element (e.g., pEvol_CAS) can be constructed (Figure 16). In some embodiments, the phagemid constitutively expresses beta-lactamase, which confers resistance to ampicillin (AmpR) or a similar antibiotic, such as carbenicillin. And it expresses an inducible arabinose promoter (Ara) that controls the expression of Cas9. In some embodiments, pEvol_CAS can be packaged into a helper bacteriophage for introduction into transformed host cells. Plasmids, such as pSelect_MUT and pSelect_WT, each containing a potential target site, can also be constructed. Alternatively, or in addition, these plasmids may contain a constitutively expressed chloramphenicol resistance gene (CmR) and a bacterial toxin under the control of the lac promoter, which allows for induction of toxin expression, for example, by IPTG (isopropyl β-D-1-thiogalactopyranoside).
[0099] To engineer allele specificity, for example, a phagemid library of Cas9 mutants can be generated using the pEvol_CAS phagemid as the first template for mutagenesis and an inclusive and unbiased mutagenesis method that targets all codons and allows for adjustment of the mutation rate.
[0100] Alternatively, or in addition, in some embodiments, each round of evolution involves exposing a phagemid library of pEvol_CAS mutants to positive selection for cleavage against Escherichia coli (E. coli) containing, for example, pSelect_MUT or pSelect_WT in a competition culture. For example, bacteria containing pSelect_MUT or pSelect_WT can be infected with phage packaging the pEvol_CAS mutants, and the bacteria can be cultured in liquid culture in ampicillin.
[0101] In some embodiments, after the first incubation and infection, the stringency of positive selection with the toxin can be evaluated, for example, by adding IPTG to induce toxin expression. In some embodiments, the expression of Cas9 and guide RNA can be induced by the addition of arabinose. During positive selection, the bacteria are continuously infected by the phages present in the liquid culture, so they can exhibit a continuous challenge to cleave the target.
[0102] In some embodiments, after the first incubation and infection, the stringency of negative selection with an antibiotic, such as chloramphenicol. In some embodiments, the expression of Cas9 and guide RNA can be induced by adding arabinose. During negative selection, the bacteria are continuously infected by the phages present in the liquid culture, so they can exhibit a continuous challenge not to cleave the target.
[0103] In another aspect, these methods typically include: (a) preparing a polynucleotide that encodes a first selectable agent and includes a DNA target site; (b) introducing the polynucleotide into a host cell such that each transformed host cell includes the polynucleotide encoding the first selectable agent and the DNA target site; (c) preparing a plurality of bacteriophages that include phagemids encoding a library of polynucleotides, wherein the various polynucleotides within the library encode various versions of a polypeptide of interest or function as templates for various versions of a polynucleotide of interest; (d) incubating the transformed host cells from step (b) together with the plurality of bacteriophages under culture conditions such that the plurality of bacteriophages infect the transformed host cells, wherein expression of the first selectable agent confers a survival benefit or detriment to the infected host cells; and (e) selecting host cells that exhibit a survival benefit (e.g., a survival benefit as described herein) in (d). For example, survival in step (d) is based on various schemes, as outlined previously.
[0104] Selection based on binding at the DNA target site when expression of a detrimental selectable agent reduces binding In certain embodiments, the selection agent confers a survival disadvantage to the host cell (e.g., a decrease in cell growth kinetics (e.g., growth delay) and / or inhibition of cell division) and / or kills the host cell, and binding at the DNA target site decreases the expression of the selection agent. Survival then is based on binding at the DNA target site: Host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of a polynucleotide that binds to the DNA target site exhibit maintenance of cell growth kinetics, an increase in cell growth kinetics (e.g., an increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics). In contrast, host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of a polynucleotide that does not bind to the DNA target site exhibit a survival disadvantage (e.g., a decrease in cell growth kinetics (e.g., growth delay) and / or inhibition of cell division), fail to survive, and / or die.
[0105] In some embodiments, expression of the selection agent is decreased by cleaving the phagemid at or near the DNA target site, e.g., by cleaving one strand of the phagemid ("nicking") or both strands.
[0106] The polynucleotides within the library may, for example, contain an antibiotic resistance gene, and the selection agent (encoded by the phagemid or with the template on the phagemid) inhibits the product of the antibiotic resistance gene. Such culture conditions during selection may, for example, include exposure to the antibiotic to which the antibiotic resistance gene confers resistance. In one example, the antibiotic resistance gene encodes beta-lactamase, the antibiotic is ampicillin or penicillin, or another beta-lactam antibiotic, and the selection agent is a beta-lactamase inhibitor protein (BLIP).
[0107] Selection based on lack of binding at a DNA target site when expression of a favorable selection agent reduces binding In certain embodiments, the selection agent confers upon a host cell a survival benefit (e.g., maintenance of cell growth kinetics, increase in cell growth kinetics (e.g., increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics)), and binding at the DNA target site reduces expression of the selection agent. Survival is based on lack of binding at the DNA target site: Host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide that does not bind to the DNA target site exhibit maintenance of cell growth kinetics, increase in cell growth kinetics (e.g., increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics). In contrast, host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide that binds to the DNA target site exhibit a survival disadvantage (e.g., decrease in cell growth kinetics (e.g., growth delay), inhibition of cell division), do not survive, and / or die.
[0108] In some embodiments, expression of the selection agent is reduced by cleaving the phagemid at or near the DNA target site, e.g., by cleaving one strand of the phagemid (“nicking”) or both strands.
[0109] Selection based on binding at a DNA target site when expression of a favorable selection agent increases binding In certain embodiments, the selection agent confers a survival benefit on the host cell (e.g., maintenance of cell growth kinetics, increase in cell growth kinetics (e.g., increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics)), and binding at the DNA target site increases the expression of the selection agent. Survival is based on binding at the DNA target site: host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of a polynucleotide that binds to the DNA target site exhibit maintenance of cell growth kinetics, increase in cell growth kinetics (e.g., increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics). In contrast, host cells transformed with a polynucleotide from the library that encodes a version of the polypeptide of interest or that functions as a template for a version of a polynucleotide that does not bind to the DNA target site exhibit a survival disadvantage (e.g., decrease in cell growth kinetics (e.g., growth delay), inhibition of cell division), do not survive, and / or die.
[0110] Selection based on lack of binding at the DNA target site when binding increases the expression of an adverse selection agent In certain embodiments, the selection agent confers a survival disadvantage (e.g., a decrease in cell growth kinetics (e.g., growth delay) and / or inhibition of cell division) on the host cell and / or kills the host cell, and binding at the DNA target site increases expression of the selection agent. Survival is based on the absence of binding at the DNA target site: Host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or functions as a template for a version of the polynucleotide that does not bind to the DNA target site exhibit maintenance of cell growth kinetics, an increase in cell growth kinetics (e.g., an increase in cell division), and / or reversal of a decrease in cell growth kinetics (e.g., at least partial rescue from a decrease in cell growth kinetics). In contrast, host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or functions as a template for a version of the polynucleotide that binds to the DNA target site exhibit a survival disadvantage (e.g., a decrease in cell growth kinetics (e.g., growth delay), inhibition of cell division), do not survive, and / or die.
[0111] Selection based on binding to a DNA target site in the host cell genome in the presence of a selection agent In one aspect, the disclosure provides a method of selecting for a version of a polypeptide or polynucleotide of interest based on whether it binds to a DNA target site in the presence of a selection agent.
[0112] As further discussed herein, these methods typically: (a) preparing a library of polynucleotides, wherein the various polynucleotides in the library encode various versions of a polypeptide or polynucleotide of interest; (b) introducing the library of polynucleotides into host cells such that each transformed host cell contains a polynucleotide encoding a version of the polypeptide or polynucleotide of interest, wherein the host cell genome contains a DNA target site; (c) preparing a plurality of bacteriophages comprising a phagemid encoding a first selection agent, wherein the first selection agent is a first selection polynucleotide; (d) incubating the transformed host cells from step (b) with the plurality of bacteriophages under culture conditions such that the plurality of bacteriophages infect the transformed host cells, wherein binding of the DNA target site in the presence of the first selection agent confers a survival benefit or detriment to the infected host cell; and (e) selecting the host cells that survived step (d).
[0113] Survival in step (d) is based on various schemes outlined below.
[0114] The DNA target site may be, for example, within the host cell genome. Additionally or alternatively, the DNA target site may be within an essential survival gene of the host cell. Additionally or alternatively, the DNA target site may be within a gene whose product interferes with the expression of a survival gene.
[0115] In some embodiments, the selection polynucleotide is a guide RNA for a CRISPR-associated (Cas) nuclease.
[0116] Survival based on lack of binding at the DNA target site when binding is detrimental In certain embodiments, binding at a DNA target site within a host cell in the presence of a selection agent (which is a polynucleotide) is disadvantageous (e.g., because binding at the DNA target site results in disruption of an essential survival gene within the host cell). Survival is based on the absence of binding at the DNA target site: Host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide of interest that does not bind to the DNA target site survive. In contrast, host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide of interest that binds to the DNA target site do not survive.
[0117] Survival binding at a DNA target site when binding is advantageous In certain embodiments, binding at a DNA target site within a host cell in the presence of a selection agent (which is a polynucleotide) is advantageous (e.g., because binding at the DNA target site results in expression of a survival gene within the host cell). Survival is based on binding at the DNA target site: Host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide of interest that binds to the DNA target site survive. In contrast, host cells transformed with a polynucleotide from a library that encodes a version of the polypeptide of interest or that functions as a template for a version of the polynucleotide of interest that does not bind to the DNA target site do not survive.
[0118] Selection based on induction of polypeptide expression In one aspect, the present disclosure provides a method of selecting a version of a polypeptide or polynucleotide of interest based on modulation or control of the expression of the polypeptide.
[0119] Furthermore, as contemplated herein, these methods typically include: (a) preparing a library of polynucleotides, wherein the various polynucleotides within the library encode various versions of the polypeptide of interest or function as templates for various versions of the polynucleotide of interest; (b) introducing the library of polynucleotides into host cells such that each transformed host cell contains a polynucleotide encoding a version of the polypeptide of interest or functions as a template for a version of the polynucleotide of interest; (c) inducing expression of the polypeptide so as to control the amount of polypeptide present in the culture; (d) preparing a plurality of bacteriophages comprising a phagemid encoding a first selection agent and containing a DNA target site; (e) incubating the transformed host cells from step (b) together with the plurality of bacteriophages under culture conditions such that the plurality of bacteriophages infect the transformed host cells, wherein expression of the first selection agent confers a survival benefit or detriment to the infected host cells; and (f) selecting the host cells that survived step (d). Survival in step (d) is based on various schemes described herein.
[0120] The polynucleotide may include an inducible promoter, such as an inducible promoter described herein, and expression is induced by contacting the polynucleotide with one or more inducing agents described herein. For example, the polynucleotide may include an arabinose promoter, and expression from the polynucleotide can be induced by contacting the polynucleotide with arabinose. In another example, the polynucleotide may include a tac promoter, and expression from the polynucleotide can be induced by contacting the polynucleotide with IPTG. In yet another example, the polynucleotide may include an rhaBAD promoter, and expression from the polynucleotide can be induced by contacting the polynucleotide with rhamnose.
[0121] Library of polynucleotides The methods of the present disclosure can begin, for example, with the step of providing a library of polynucleotides (e.g., a plasmid library), where the various polynucleotides within the library encode various versions of the polypeptide of interest (or, in the case of the polynucleotide of interest, the library comprises various versions of the polynucleotide of interest, and / or various versions of the polynucleotide of interest that function as templates for various versions of the polynucleotide of interest).
[0122] The libraries described herein may, for example, comprise polynucleotides operably linked to an inducible promoter. For example, induction of the promoter may induce the expression of the polypeptide encoded by the polynucleotide. In some embodiments, induction of the promoter for inducing the expression of the polypeptide encoded by the polynucleotide affects the efficiency of the selection method. For example, the efficiency of the selection method may be improved and / or increased as compared to the efficiency of a selection method that does not use an inducible promoter.
[0123] Such libraries can be obtained, for example, by using existing libraries, such as those available through commercially available and / or publicly available collections, or by purchasing them. Alternatively, or in addition, the library can be obtained from mutagenesis methods. For example, the library can be obtained by a random mutagenesis method or a comprehensive mutagenesis method, such as a method that randomly targets polynucleotides throughout a predefined target region for mutagenesis.
[0124] The library can also be obtained by a targeted mutagenesis method. For example, a polynucleotide of interest, or a sub-region of a polypeptide of interest, may be targeted for mutagenesis. Additionally, or alternatively, the entire polynucleotide of interest, or the entire polypeptide of interest, may be targeted for mutagenesis.
[0125] The polypeptide or polynucleotide of interest typically has DNA-binding ability, but it is expected that not all versions of the polypeptide or polynucleotide of interest encoded by different polynucleotides in the library will necessarily be able to bind to DNA. Furthermore, it is expected that the ability to bind to DNA may vary among versions of the polypeptide or polynucleotide of interest encoded by different polynucleotides in the library. Indeed, the selection methods of the present disclosure involve distinguishing between versions of the polypeptide or polynucleotide of interest that can and cannot bind to a DNA target site. In certain embodiments, many or most versions of the polypeptide or polynucleotide of interest do not bind to DNA.
[0126] Similarly, in embodiments where the polypeptide or polynucleotide of interest can cleave DNA, not all versions of the polypeptide or polynucleotide of interest will necessarily be able to cleave DNA.
[0127] host cell The methods of the present disclosure may include, after the step of preparing a library of polynucleotides, introducing the library of polynucleotides into host cells such that each transformed host cell contains a polynucleotide encoding a version of the polypeptide of interest or functions as a template for a version of the polynucleotide of interest.
[0128] A host cell generally refers to a cell that can take up exogenous substances, such as nucleic acids (e.g., DNA and RNA), polypeptides, or ribonucleoproteins. The host cell can be, for example, a unicellular organism, such as a microorganism, or a eukaryotic cell, such as a yeast cell, a mammalian cell (e.g., in culture).
[0129] In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell, such as Escherichia coli (E. coli). The bacterial cell can be Gram-negative or Gram-positive and can belong to the Bacteria (formerly called Eubacteria) domain or the Archaea (formerly called Archaebacteria) domain. Any of these types of bacteria can grow in a laboratory environment and can be suitable as a host cell as long as it can take up exogenous substances.
[0130] The host cell can be a competent or made competent bacterial cell in that it can take up exogenous substances, such as genetic material.
[0131] There are various mechanisms by which exogenous substances, such as genetic material, can be introduced into a host cell. For example, in bacteria, there are three general mechanisms, classified as transformation (uptake and incorporation of extracellular nucleic acids such as DNA), transduction (transfer of genetic material from one cell to another, e.g., by a plasmid or by a virus that infects cells such as a bacteriophage), and conjugation (direct transfer of nucleic acids between two cells that are temporarily joined). A host cell into which genetic material has been introduced by transformation is usually referred to as a "transformed host cell".
[0132] In some embodiments, a library of polynucleotides is introduced into a host cell by transformation. Protocols for transforming host cells are known in the art. For bacterial cells, for example, there are methods based on electroporation, methods based on lipofection, methods based on heat shock, methods based on agitation with glass beads, methods based on chemical transformation, and methods based on bombardment with particles coated with an exogenous substance (e.g., DNA or RNA). One of ordinary skill in the art will be able to select a method based on the art and / or the protocol provided by the manufacturer of the host cell.
[0133] The transformed host cells may each contain, for example, a version of a polypeptide of interest or a polynucleotide that functions as a template for a version of a polynucleotide of interest.
[0134] The library can be introduced into a population of host cells such that the population of transformed host cells collectively contains all members of the library of polynucleotides. That is, for all versions of the polynucleotides within the library, at least one host cell within the population contains a version of the polynucleotide such that all versions of the polynucleotides within the library are represented within the population of transformed host cells.
[0135] Bacteriophage The methods of the present disclosure may include, after the step of introducing a library of polynucleotides into a host cell, the step of preparing a plurality of bacteriophages comprising a phagemid that encodes a first selection agent and contains a DNA target site.
[0136] A bacteriophage is a virus that infects bacteria and injects its genome (and / or any phagemid packaged within the bacteriophage) into the cytoplasm of the bacterium. Typically, bacteriophages replicate within bacteria. However, replication-defective bacteriophages also exist.
[0137] In some embodiments, a plurality of bacteriophages comprising the phagemid described herein are incubated with the transformed host cell under conditions under which the bacteriophage can infect the transformed host cell. The bacteriophage can be replication-competent; for example, the bacteriophage replicates within the transformed host cell, and the replicated viral particles are released as virions into the medium, enabling reinfection of other host cells by the bacteriophage.
[0138] Virions can be released from the host cell without lysing the host cell.
[0139] In some embodiments, the plurality of bacteriophages present a continuous sensitization to the host cell by continuously infecting (infecting and reinfecting) the transformed host cell.
[0140] The bacteriophage can be a "helper phage" in that it selectively packages the phagemid relative to the phage DNA. For example, the bacteriophage selectively packages the phagemid relative to the phage DNA at a ratio of at least 3:1, at least 4:1, at least 5:1, at least 6:1, at least 7:1, at least 8:1, at least 9:1, or at least 10:1.
[0141] In some embodiments, the bacteriophage does not normally lyse the host cell. For example, the bacteriophage does not lyse the host cell under conditions under which the transformed host cell is incubated with the plurality of bacteriophages.
[0142] The bacteriophage can be a filamentous bacteriophage. Filamentous bacteriophages typically infect Gram-negative bacteria (such as, among others, Escherichia coli (E. coli), Pseudomonas aeruginosa (P. aeruginosa), Neisseria gonorrhoeae (N. gonorrhoeae), and Yersinia pestis (Y. pestis)) and have a single-stranded DNA genome.
[0143] For example, the filamentous bacteriophage may be an Ff phage, which infects Escherichia coli (E. coli) having an F episome. Examples of such phages include, but are not limited to, M13 bacteriophage, f1 phage, fd phage, and their derivatives and variants.
[0144] In addition, or alternatively, the bacteriophage may be an M13 bacteriophage or a derivative or variant thereof. For example, the bacteriophage may be M13KO7 (a derivative of M13 having a kanamycin resistance marker and a p15A origin of replication). M13KO7 is characterized as functioning as a useful helper phage due to its relatively high phagemid-to-phage packaging ratio of approximately 10:1.
[0145] In addition, or alternatively, the bacteriophage may be VCSM13 (a derivative of M13KO7).
[0146] The bacteriophage may also be an f1 bacteriophage or a derivative or variant thereof. For example, the bacteriophage may be R408 (a derivative of f1 having no antibiotic selection marker).
[0147] In addition, or alternatively, the bacteriophage may be CM13 (a derivative of M13KO7 that has been reported to more reliably produce virions than M13KO7).
[0148] A pool of bacteriophages containing various phagemids may be used in the methods of the present disclosure. For example, as further discussed herein, if it is desired to select for binding to multiple off-target sites and / or cleavage of multiple off-target sites, different off-site targets may be presented on different phagemids contained within the same pool of bacteriophages.
[0149] Phagemid Since phagemids are circular plasmids having an f1 origin of replication from the f1 phage, they can replicate as plasmids and can be packaged as single-stranded DNA by bacteriophages. Phagemids also contain an origin of replication for double-stranded replication (e.g., while inside the host cell).
[0150] Phagemids suitable for use in the present invention typically encode a selection agent or function as a template for a selection agent and contain a DNA target site. Thus, phagemids used in the present invention typically contain a regulatory element that drives expression of the selection agent, operably linked to the selection agent, a gene element that encodes the selection agent or functions as a template for the selection agent.
[0151] As noted above, the DNA target site can be included anywhere within the phagemid. For example, the DNA target site can be positioned within the regulatory element, within the gene element, outside and distal to both the regulatory element and the gene element, or outside both elements but proximal to at least one of the elements.
[0152] The location of the DNA target site can depend on the embodiment. For example, the DNA target site can be positioned within the regulatory element. This positioning can be appropriate, for example, in embodiments where the polypeptide of interest is a transcription factor, such as a transcriptional activator or repressor, and the selection is based on whether the transcription factor binds to the DNA target site.
[0153] There is no limitation on where the DNA target site can be positioned in terms of the binding of the polypeptide of interest anywhere within the phagemid increasing or decreasing the expression of the selection agent. For example, binding of the polypeptide of interest at the DNA target site can result in cleavage of the phagemid at or near the DNA target site. Cleavage of the phagemid anywhere within the phagemid results in linearization of the phagemid, whereby the phagemid is no longer replicated within the host cell and thus expression of the selection agent ends.
[0154] Phagemids can be packaged into bacteriophages using methods known in the art, such as those provided by bacteriophage manufacturers. For example, commonly used protocols involve forming a double-stranded plasmid version of the desired phagemid construct, transforming the double-stranded plasmid into a host cell such as a bacterium, and then inoculating a culture of such transformed host cells with a helper bacteriophage capable of packaging the double-stranded plasmid as a single-stranded phagemid.
[0155] Culture conditions The method of the present disclosure may include, after preparing a plurality of bacteriophages containing a phagemid encoding a first selection agent and containing a DNA target site, incubating the transformed host cell (into which a library of polynucleotides has been introduced) together with the plurality of bacteriophages under culture conditions such that the plurality of bacteriophages infect the transformed host cell. Usually, the conditions are such that the expression of the first selection agent confers a survival disadvantage or a survival benefit, depending on the embodiment.
[0156] In certain embodiments, the culture conditions are competitive culture conditions. "Competitive culture conditions" refer to conditions under which a population of organisms (such as host cells) are grown together and must compete for the same limited resources, such as nutrients, oxygen.
[0157] The host cells may be incubated in an environment with no or little fresh nutrient input. For example, the host cells may be incubated in an environment with no or little fresh oxygen input, such as in a sealed container, such as a flask.
[0158] In addition, or alternatively, the host cells may be incubated in a medium that is well mixed throughout the incubation period, for example, in a liquid shake culture. Typically, under such well-mixed conditions, the host cells will compete for nutrients and / or oxygen (in the case of aerobic organisms) as their nutritional requirements are similar and the nutrients and / or oxygen are depleted by the growing population.
[0159] In addition, or alternatively, the host cells may be incubated at a substantially constant temperature, for example, at a temperature most suitable for the type of host cell. For certain bacterial species such as Escherichia coli (E. coli), the host cells are typically incubated at a temperature of approximately 37°C. In some embodiments, the host cells are incubated within 5°C, 4°C, 3°C, 2°C, or 1°C of 37°C, for example, at approximately 37°C.
[0160] The host cells may be incubated in a liquid shake culture. This shaking is typically vigorous enough to prevent non-uniform distribution of nutrients and / or precipitation of some host cells at the bottom of the culture. For example, the host cells may be shaken at at least 100 revolutions per minute (rpm), at least 125 rpm, at least 150 rpm, at least 175 rpm, at least 200 rpm, at least 225 rpm, at least 250 rpm, at least 275 rpm, or at least 300 rpm. In some embodiments, the host cells are shaken at 100 rpm to 400 rpm, for example, 200 to 350 rpm, for example, at approximately 300 rpm.
[0161] The host cells may be incubated for a period of time before a plurality of bacteriophages are introduced into the culture. This period allows, for example, the host cell population to recover from a storage state and / or reach a particular desired density prior to the introduction of the plurality of bacteriophages. A selection pressure may or may not be applied during this period before the plurality of bacteriophages are introduced.
[0162] The culture conditions may include continuous incubation of the host cells with the bacteriophage for a period of time, for example, at least 4 hours, at least 8 hours, at least 12 hours, or at least 16 hours. Additionally, or alternatively, the culture conditions may include continuous incubation of the host cells with the bacteriophage until the growth of the host cells saturates.
[0163] The culture conditions may allow for continuous infection of the host cells by the bacteriophage. That is, the host cells become infected and, if they survive, are continuously reinfected during the incubation period.
[0164] Additionally, or alternatively, a selection pressure is introduced into the culture. For example, particularly when the host cells are transformed with exogenous DNA (such as a plasmid), a selection pressure can be introduced to preferentially select host cells that maintain the exogenous DNA. As a commonly used scheme, the use of one or more antibiotics as a selection pressure and the corresponding antibiotic resistance gene within the exogenous DNA to be maintained can be mentioned. This selection pressure may be the same as or different from that related to the selection agent discussed herein, and in some embodiments, both are used, for example, sequentially and / or simultaneously.
[0165] In some embodiments, at least during the period in which the transformed host cells are incubated with bacteriophage, the culture conditions include exposure to one or more antibiotics, and some host cells may be resistant to the antibiotics by virtue of an antibiotic resistance gene present on a phagemid, a polynucleotide in the library, or both. For example, both the phagemid and the polynucleotide in the library may have an antibiotic resistance gene, and for example, the antibiotic resistance genes may be the same or different. If the phagemid contains an antibiotic resistance gene (a "first antibiotic resistance gene" that confers resistance to a "first antibiotic") and the polynucleotide contains a different antibiotic resistance gene (a "second antibiotic resistance gene" that confers resistance to a "second antibiotic"), the culture conditions can include any of a variety of schemes. Non-limiting examples are that the conditions are: 1) simultaneous exposure to both the first antibiotic and the second antibiotic; 2) continuous exposure to the second antibiotic for a period of time (e.g., during the period in which the host cells are incubated before the bacteriophage is introduced into the culture), followed by i) exposure to the first antibiotic, or ii) exposure to both the first antibiotic and the second antibiotic (e.g., during the period in which the host cells are incubated with the bacteriophage), or; 3) exposure to only one of the relevant antibiotics (e.g., the first antibiotic) during the course of the incubation.
[0166] Selective agent The methods described herein may include providing a plurality of bacteriophages that include a phagemid that encodes a selective agent or functions as a template for a selective agent. Depending on the embodiment, the selective agent may confer a survival benefit or a survival detriment to the host cell under conditions in which the host cell is incubated with the bacteriophage. The selective agent may confer an increase in cell growth kinetics, or a decrease in cell growth kinetics to the host cell under conditions in which the host cell is incubated with the bacteriophage.
[0167] The selective agent may be, for example, a polypeptide and / or a polynucleotide.
[0168] In some embodiments, the selection agent confers a survival benefit on the host cell.
[0169] In some embodiments, the selection agent is encoded by a gene essential for the survival of the host cell. Examples of such essential survival genes include, but are not limited to, genes involved in fatty acid biosynthesis; genes involved in amino acid biosynthesis; genes involved in cell division; genes involved in global regulatory functions; genes involved in protein translation and / or modification; genes involved in transcription; genes involved in proteolysis; genes encoding heat shock proteins; genes involved in ATP transport; genes involved in peptidoglycan synthesis; genes involved in DNA replication, repair, and / or modification; genes involved in tRNA modification and / or synthesis; and genes encoding ribosomal components and / or involved in ribosome synthesis.For example, in Escherichia coli, several essential survival genes are known in the art and include, but are not limited to, accD (acetyl-CoA carboxylase, carboxyltransferase component, beta subunit), acpS (CoA:apo-[acyl-carrier-protein] pantetheine phosphotransferase), asd (aspartate semialdehyde dehydrogenase), dapE (N-succinyl-diaminopimelate desacylase), dnaJ (chaperone of DnaK; heat shock protein), dnaK (chaperone Hsp70), era (GTP-binding protein), frr (ribosome release factor), ftsI (septum formation; penicillin-binding protein 3; peptidoglycan synthase), ftsL (cell division protein; inner growth of the septum wall); ftsN (essential cell division protein); ftsZ (cell division; forms a peripheral ring; tubulin-like GTP-binding protein and GTPase), gcpE, grpE (phage lambda replication; host DNA synthesis; heat shock protein; protein repair), hflB (degrades sigma32, integral membrane peptidase, cell division protein), infA (protein chain initiation factor IF-1), lgt (phosphatidylglycerol prolipoprotein diacylglyceryl transferase; major membrane phospholipid), lpxC (UDP-3-O-acyl-N-acetylglucosamine deacetylase; lipid A biosynthesis), map (methionine aminopeptidase), mopA (GroEL, chaperone Hsp60, peptide-dependent ATPase, heat shock protein), mopB (GroES, 10 Kd chaperone binds to Hsp60 (in press).MgATP, which inhibits its ATPase activity), msbA ATP-binding transport protein; multicopy suppressor of htrB), murA (the first step in murein biosynthesis; UDP-N-acetylglucosamine 1-carboxyvinyltransferase), murI (glutamate racemase, required for the biosynthesis of D-glutamate and peptidoglycan), nadE (NAD synthetase, which selects NH3 over glutamine), nusG (a component of transcriptional antitermination), parC (DNA topoisomerase IV subunit A), ppa (inorganic pyrophosphatase), proS (proline tRNA synthetase), pyrB (aspartate carbamoyltransferase, catalytic subunit), rpsB (30S ribosomal subunit protein S2), trmA (tRNA(uracil-5-)-methyltransferase), ycaH, ycfB, yfiL, ygjD (putative O-sialoglycoprotein endopeptidase), yhbZ (putative GTP-binding factor), yihA, and yjeQ. Additional essential genes in Escherichia coli (E. coli) include those listed in "Experimental Determination and System-Level Analysis of Essential Genes in E. coli MG1655" by Gerdes 2003 (e.g., in Supplementary Tables 1, 2, and 6).
[0170] In some embodiments, the selection agent is encoded by an antibiotic resistance gene, as further discussed below. For example, the culture conditions may include exposure to an antibiotic to which the antibiotic resistance gene confers resistance.
[0171] In some embodiments, the selection agent inhibits a gene product that confers a survival disadvantage.
[0172] In certain embodiments, the selection agent confers a survival disadvantage on the host cell. The selection agent can be toxic to the host cell. For example, the selection agent can be a toxin, many of which are known in the art and many of which have been identified in various bacterial species. Examples of such toxins include, but are not limited to, ccdB, FlmA, fst, HicA, Hok, Ibs, Kid, LdrD, MazF, ParE, SymE, Tisb, TxpA / BrnT, XCV2162, yafO, Zeta, and tse2. For example, the selection agent can be ccdB, which is found in Escherichia coli (E. coli). In other examples, the selection agent is tse2.
[0173] The selection agent can be toxic because it produces a toxic substance. For example, the production of the toxic substance can occur only in the presence of another agent, the presence of which can or cannot be controlled externally.
[0174] In addition or alternatively, the selection agent can inhibit a gene product that confers a survival benefit. By way of non-limiting example, the selection agent can be a beta-lactamase inhibitor protein (BLIP), which inhibits, among other things, beta-lactamases such as ampicillin and penicillin.
[0175] Inducing agent The methods described herein can include the step of preparing a library of polynucleotides, wherein the various polynucleotides within the library encode various versions of the polypeptide of interest (or, in the case of the polynucleotide of interest, function as templates for various versions of the polynucleotide of interest). The polynucleotide can include, for example, regulatory elements such as a promoter, which can control the expression of the polypeptide. The regulatory element can be an inducible promoter and the expression can be induced by an inducing agent. Such an inducing agent and / or the induced expression can increase or improve the efficiency of selection.
[0176] The inducer may be a polypeptide and / or a polynucleotide. The inducer may also be a small molecule, light, temperature, or an intracellular metabolite.
[0177] In some embodiments, the inducer is arabinose, anhydrotetracycline, lactose, IPTG, propionate, blue light (470 nm), red light (650 nm), green light (532 nm), or L-rhamnose. For example, the inducer may be arabinose.
[0178] A library of polynucleotides encoding various versions of the Cas9 molecule In some embodiments, the methods and compositions of the invention may be used with a library of polynucleotides encoding various versions of the Cas9 molecule or Cas9 polypeptide (e.g., an inclusive and unbiased library of Cas9 mutants spanning all or part of the Cas9 molecule or Cas9 polypeptide). In certain embodiments, the methods and compositions of the invention can be used to select one or more members of the library based on specific properties. In typical embodiments, the Cas9 molecule or Cas9 polypeptide has the ability to interact with the gRNA molecule and, in cooperation with the gRNA molecule, localize to an internal nucleic acid site. Other activities, such as PAM specificity, cleavage activity, or helicase activity, may vary more widely in the Cas9 molecule and Cas9 polypeptide.
[0179] In some embodiments, the methods and compositions of the invention can be used to select one or more versions of a Cas9 molecule or Cas9 polypeptide that include an alteration in enzymatic properties (e.g., an alteration in nuclease activity or helicase activity) compared to other reference Cas9 molecules (including naturally occurring Cas9 molecules or Cas9 molecules that have already been engineered or modified). As discussed herein, mutant versions of a reference Cas9 molecule or Cas9 polypeptide can have nickase activity or non-cleavage activity (opposite to double-stranded nuclease activity). In one embodiment, the methods and compositions of the invention can be used to select one or more versions of a Cas9 molecule or Cas9 polypeptide that have an alteration that changes its size (e.g., a deletion of an amino acid sequence that decreases its size), where the effect on one or more, or any, Cas9 activities may or may not be significant. In one embodiment, the methods and compositions of the invention can be used to select one or more versions of a Cas9 molecule or Cas9 polypeptide that recognize different PAM sequences (e.g., a version of a Cas9 molecule can be selected to recognize a PAM sequence other than that recognized by the endogenous wild-type PI domain of a reference Cas9 molecule).
[0180] A library having various versions of a Cas9 molecule or a Cas9 polypeptide can be prepared by any method, for example, by modifying a parental, for example, naturally occurring, Cas9 molecule or Cas9 polypeptide to provide a library of modified Cas9 molecules or Cas9 polypeptides. For example, one or more mutations or differences can be introduced compared to a parental Cas9 molecule, for example, a naturally occurring or engineered Cas9 molecule. Such mutations and differences include: substitutions (e.g., non-conservative substitutions or substitutions of essential amino acids); insertions; or deletions. In one embodiment, the Cas9 molecule or Cas9 polypeptide within the library of the present invention has one or more mutations or differences, for example, compared to a reference Cas9 molecule, for example, a parental Cas9 molecule, at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, or 50 mutations, and less than 200, 100, or 80 mutations.
[0181] Library of guide RNA molecules In some embodiments, the methods and compositions of the present disclosure may be used with a library of guide RNA molecules and / or polynucleotides encoding guide RNA molecules. For example, a library can be provided and / or generated that includes DNA molecules encoding guide RNAs having (i) different targeting domains described herein; (ii) different first and / or second complementary domains described herein; and / or (iii) different stem-loops described herein. As described herein, the library can be introduced into a host cell. In some embodiments, a nucleic acid encoding an RNA-guided nuclease, for example, a Cas9 molecule or a Cas9 polypeptide, is also introduced into the host cell.
[0182] In certain embodiments, the methods and compositions of the disclosure can be used to select one or more members of a guide RNA library based on specific properties, e.g., localizing to a site within a nucleic acid, and / or interacting with a Cas9 molecule or Cas9 polypeptide, and / or localizing a Cas9 molecule or Cas9 polypeptide to a site within a nucleic acid.
[0183] Libraries having various versions of guide RNAs can be prepared using any method that provides a library of modified guide RNAs, e.g., by modification of a parental, e.g., naturally occurring, guide RNA. For example, one or more mutations or differences can be introduced compared to a parental guide RNA, e.g., a naturally occurring or engineered guide RNA. Such mutations and differences include: substitutions; insertions; or deletions. In some embodiments, the guide RNAs within the libraries of the disclosure can include at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, or 50 mutations, and less than 200, 100, or 80 mutations, compared to a reference guide RNA, e.g., a parental guide RNA.
[0184] DNA target site Generally, the DNA target site for a particular method of the invention can depend on the physical location of the DNA target site (e.g., in some embodiments, the DNA target site can be positioned on a phagemid while in other embodiments, the DNA target site can be positioned within the host cell genome), the nature of the polypeptide or polynucleotide of interest, the nature of the selection process, and / or the desired outcome of the selection process. The DNA target site can be positioned within various types of nucleotide sequences. For example, in some embodiments, the DNA target site can be positioned within a non-transcribed element, within an element that encodes a polypeptide or functions as a template for a polynucleotide (e.g., non-coding RNA), within a regulatory element that controls the expression of a polypeptide, among others.
[0185] As described herein, in some embodiments, the DNA target site may be positioned on a phagemid. In some embodiments, the DNA target site may be positioned on a plasmid. In situations where the selection process depends on cleavage (or non-cleavage) of the phagemid or plasmid, the DNA target site may be positioned anywhere on the phagemid or plasmid. This is because selection can be derived from cleavage at any position on the phagemid or plasmid, depending on linearization (and subsequent destruction) of the phagemid or plasmid. In situations where the selection process depends on suppression (or activation) of the expression of a selection agent, the DNA target site may be positioned within a regulatory element that drives the expression of the selection agent. In some embodiments, the regulatory element may be an inducible regulatory element.
[0186] As described herein, in some embodiments, the DNA target site may be positioned within the host cell genome. In situations where the selection process depends on cleavage of an endogenous gene (an "essential gene") that is essential for the survival of the host cell, the DNA target site may be positioned, for example, within the coding or regulatory elements of the essential gene. In situations where the selection process depends on suppression of the essential gene, the DNA target site may be at any position within the host cell genome that leads to suppression of the essential gene when bound by the polypeptide of interest (e.g., within the regulatory elements of the essential gene, such as between the promoter and coding region of the essential gene).
[0187] The specific nucleotide sequence of the DNA target site (i.e., independent and distinct from whether it is positioned on a phagemid or within the host cell genome) will typically depend on the nature of the polypeptide of interest, the nature of the selection process, and the desired outcome of the selection process. As an example, if the polypeptide of interest is a reference nuclease that recognizes a first nucleotide sequence (e.g., a meganuclease, TALEN, or zinc finger nuclease), and the method of the invention is used to select one or more modified versions of the reference nuclease that selectively bind to a second nucleotide sequence that is different from the first nucleotide sequence (e.g., 1, 2, 3 bases, etc.), the method of the invention may involve the use of DNA target sites corresponding to the second nucleotide sequence in the positive selection step and DNA target sites corresponding to the first nucleotide sequence in the negative selection step (i.e., for selecting versions of the reference nuclease that bind to the second nucleotide sequence but not to the first nucleotide sequence).
[0188] In the case of a Cas molecule (e.g., a Cas9 molecule), the DNA target site will be determined, in part, based on the PAM of the Cas molecule and the sequence of the targeting domain of the gRNA used to localize the Cas molecule to the DNA target site. As an example, if the polypeptide of interest is a reference Cas9 molecule that recognizes a first PAM sequence, and the method of the invention is used to select one or more modified versions of the reference Cas9 molecule that selectively recognize a second PAM sequence that is different from the first PAM sequence (e.g., 1, 2, 3 bases, etc.), the method of the invention may involve the use of DNA target sites containing the second PAM sequence in the positive selection step and DNA target sites containing the first PAM sequence in the negative selection step (i.e., for selecting versions of the reference Cas9 molecule that recognize the second PAM sequence but not the first PAM sequence). In both cases, the DNA target site will also include a sequence that is complementary to the sequence of the targeting domain of the gRNA used to localize the Cas9 molecule to the DNA target site.
[0189] In some embodiments, the methods provided herein can be used to evaluate the ability of PAM variants to direct cleavage of a target site by an RNA-guided nuclease, such as mutated Streptococcus pyogenes Cas9.
[0190] In some embodiments, the library comprises a plurality of nucleic acid templates further comprising nucleotide sequences containing PAM variants adjacent to the target site. In some embodiments, the PAM sequence comprises the sequence NGA, NGAG, NGCG, NNGRRT, NNGRRA, or NCCRRC.
[0191] Some of the methods provided herein enable the simultaneous evaluation of multiple PAM variants for any given target site and, in some embodiments, in combination with mutated Streptococcus pyogenes Cas9. Thus, data obtained from such methods can be used to compile a list of PAM variants that mediate cleavage of a particular target site, in combination with wild-type Streptococcus pyogenes Cas9 or mutated Streptococcus pyogenes Cas9. In some embodiments, a sequencing method is used to generate quantitative sequencing data and determine the relative amount of cleavage of a particular target site mediated by a particular PAM variant.
[0192] Antibiotic resistance gene In certain embodiments, the plasmids within the library and / or phagemids contain an antibiotic resistance gene. In some embodiments, the antibiotic resistance gene confers resistance to an antibiotic that kills or inhibits the growth of bacteria, such as Escherichia coli (E. coli). Non-limiting examples of such antibiotics include ampicillin, bleomycin, carbenicillin, chloramphenicol, erythromycin, kanamycin, penicillin, polymyxin B, spectinomycin, streptomycin, and tetracycline. Various antibiotic resistance gene cassettes are known and are available and / or commercially available in the art, for example, as elements within plasmids. For example, there are some commercially available plasmids or genetic elements having ampR (ampicillin resistance), bleR (bleomycin resistance), carR (carbenicillin resistance), cmR (chloramphenicol resistance), kanR (kanamycin resistance), and / or tetR (tetracycline resistance). An additional example of an antibiotic resistance gene is beta-lactamase.
[0193] In some embodiments, the phagemid contains a first antibiotic resistance gene and the plasmid within the library contains a second antibiotic resistance gene. In some embodiments, the first antibiotic resistance gene is different from the second antibiotic resistance gene. For example, in some embodiments, the first antibiotic resistance gene is the cmR (chloramphenicol resistance) gene and the second antibiotic resistance gene is the ampR (ampicillin resistance) gene.
[0194] Regulatory elements In certain embodiments, a genetic element (e.g., one encoding a selectable agent, an antibiotic resistance gene, a polypeptide, a polynucleotide) is operably linked to a regulatory element to enable the expression of one or more other elements, such as a selectable agent, an antibiotic resistance gene, a polypeptide, a polynucleotide.
[0195] In some embodiments, the phagemid contains, on the phagemid, one or more genetic elements, such as regulatory elements that drive the expression of a selection agent.
[0196] In some embodiments, the polynucleotide within the library contains, on the polynucleotide, one or more genetic elements, such as a polypeptide or polynucleotide of interest, and, if present, regulatory elements that drive the expression of a genetic element encoding a selection agent, such as an antibiotic resistance gene.
[0197] There are a wide variety of gene regulatory elements. The type of regulatory element used may depend, for example, on the host cell, the type of gene intended to be expressed, and other factors such as the transcription factors used.
[0198] Gene regulatory elements include, but are not limited to, enhancers, promoters, operators, terminators, and combinations thereof. As a non-limiting example, the regulatory element may include both a promoter and an operator.
[0199] In some embodiments, the regulatory element is constitutive in that it is active intracellularly in all situations. For example, constitutive elements such as constitutive promoters can be used to express gene products without the need for additional regulation.
[0200] In some embodiments, the regulatory element is inducible, i.e., it is active only in response to a specific stimulus.
[0201] For example, the lac operator is inducible in that it can be activated in the presence of IPTG (isopropyl β-D-1-thiogalactopyranoside). Another example is the arabinose promoter that is activated in the presence of arabinose.
[0202] In some embodiments, the regulatory element is bidirectional in that it can drive the expression of genes located on opposite sides within an array. Thus, in some embodiments, the expression of at least two gene elements can be driven by the same gene element.
[0203] Gene segments that function as regulatory elements are readily available in the art and many are commercially available from vendors. For example, expression plasmids or other vectors that already contain one or more regulatory elements for expressing a gene segment of interest are readily available.
[0204] Analysis of Selected Versions of Polypeptides and / or Polynucleotides After one or more rounds of selection, selected versions of the polypeptide and / or polynucleotide can be recovered and analyzed from the host cells that survived the selection. In the context of using multiple cycles of evolution (mutation followed by one or more rounds of selection), the analysis can be performed at the end of each cycle or only within some of the cycles.
[0205] Examples of types of analysis include, but are not limited to, sequencing, binding, and / or cleavage assays (including in vitro assays), and confirmation of the activity of selected versions in cell types other than the host cell type.
[0206] As a non-limiting example, next-generation sequencing (also known as high-throughput sequencing) may be performed to sequence all or most of the selected mutants.
[0207] In some embodiments, deep sequencing is performed, meaning that each nucleotide is read several times during the sequencing process at a depth that is, for example, greater than at least 7, at least 10, at least 15, at least 20, or even greater. Here, the depth (D) is D = N × L / G (Equation 1) (where N is the number of reads, L is the length of the actual genome, and G is the length of the polynucleotide to be sequenced).
[0208] In some embodiments, Sanger sequencing is used to analyze at least some of the selected versions.
[0209] Analysis of the sequences may be used, for example, to check for enriched amino acid residues or nucleotides indicative of the selected versions.
[0210] Alternatively, or in addition, samples of the selected versions may be sequenced, for example, from individual host cell colonies (e.g., bacterial colonies).
[0211] Binding and / or cleavage assays are known in the art. Some of these assays are performed in vitro using, for example, components of cells or isolated molecules (e.g., polypeptides, polynucleotides, or ribonucleoproteins) rather than whole cells.
[0212] In some embodiments, in vitro assays for the binding and / or cleavage of DNA substrates are performed. In some embodiments, the assay tests the activity of lysates extracted from host cells that survived one or more rounds of selection. In some embodiments, the assay tests the activity of polypeptides, polynucleotides, and / or ribonucleoproteins, or complexes thereof, extracted from host cells that survived one or more rounds of selection.
[0213] In some embodiments, the analysis includes performing one or more assays to test one or more functions of the products of selected versions of polynucleotides in the library (e.g., polypeptides encoded by the selected versions, or polynucleotides whose template is the selected version).
[0214] Use In some embodiments, the selection method of the present invention is used in conjunction with a mutagenesis method for generating a library of plasmids. Any mutagenesis method can be used in conjunction with the selection method of the present invention.
[0215] In some embodiments, one or more rounds of selection following one round of mutagenesis are used. This cycle may be performed once, for example, as part of a directed evolution strategy, or may be repeated one or more times. The versions of the polypeptides and / or polynucleotides of interest selected within one cycle are mutated in the mutagenesis round of the next cycle. For example, the cycle may be repeated many times until the selected versions of the polypeptides and / or polynucleotides of interest obtained meet a certain criterion and / or until a desired number of selected polypeptides and / or polynucleotides that meet a certain criterion are obtained.
[0216] In some embodiments, within one cycle, following one round of mutagenesis, one round of positive selection (e.g., for versions of the polypeptides and / or polynucleotides of interest that cleave and / or bind to a DNA target site) follows.
[0217] In some embodiments, within one cycle, following one round of mutagenesis, one round of positive selection follows, and then one round of negative selection (e.g., for versions of the polypeptides and / or polynucleotides of interest that do not cleave and / or do not bind to a DNA target site) follows.
[0218] In some embodiments, within one cycle, following one round of mutagenesis, one round of negative selection follows, and then one round of positive selection follows.
[0219] In embodiments where multiple cycles are performed, the cycles do not need to be generally the same with respect to rounds of mutagenesis and selection. Additionally, other details do not need to be the same between cycles. For example, the method of mutagenesis does not need to be the same from one cycle to the next, and the exact conditions or generalities of the selection rounds also do not need to be the same.
[0220] Thus, the selection methods of the present disclosure can be used to select polypeptides and / or polynucleotides of interest that have specificity for a desired binding and / or cleavage site.
[0221] For example, the selection method can be used to select polypeptides and / or polynucleotides of interest that bind to one allele but not to another. For example, the ability to distinguish between a disease allele and a wild-type allele can be used, for example, to develop therapies based on techniques of gene editing, gene suppression, and / or gene activation. In some embodiments, for example, a positive selection is performed to select a polypeptide or polynucleotide of interest that recognizes one allele (e.g., a disease allele), and then a negative selection is performed to select a polypeptide or polynucleotide of interest that recognizes another allele (e.g., a wild-type allele). In some embodiments, a negative selection is performed to select a polypeptide or polynucleotide of interest that recognizes one allele (e.g., a wild-type allele), and then a positive selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes another allele (e.g., a disease allele).
[0222] As shown in the examples, the selection method of the present invention was used in an evolution scheme to evolve polypeptides that have the ability to distinguish between alleles that differ by only one base change.
[0223] As another example, for instance, a selection method may be used to select a polypeptide and / or polynucleotide of interest whose binding selection has been altered as compared to a naturally occurring polypeptide and / or polynucleotide of interest. For example, certain DNA-binding proteins (including enzymes) have very limited binding specificities, thus limiting their uses. Selection and / or evolution of a site-specific DNA-binding domain or protein whose binding specificity has been altered (e.g., as compared to a naturally occurring polypeptide and / or polynucleotide of interest) may increase the scope of its uses.
[0224] In some embodiments, positive selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes a certain DNA target site (e.g., a desired new target site), and then negative selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes another DNA target site (e.g., a natural target site).
[0225] In some embodiments, positive selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes a certain DNA target site (e.g., a desired new target site), and negative selection is not performed.
[0226] In some embodiments, negative selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes a certain DNA target site (e.g., a natural target site), and then positive selection is performed to select a polypeptide and / or polynucleotide of interest that recognizes another DNA target site (e.g., a desired new target site).
[0227] As another example, a selection method may be used to select a polypeptide or polynucleotide of interest with reduced off-target activity. Certain DNA-binding proteins are classified as specific for a particular recognition sequence, but some may exhibit promiscuity in that they bind to one or more off-target sites to some extent.
[0228] In some embodiments, for example, a negative selection is performed to select a polypeptide or polynucleotide of interest that does not recognize one or more off-target sites. If it is desired to select multiple off-target sites, in some embodiments, a pool of bacteriophages containing various phagemids is used, and each of the various phagemids contains a DNA target site corresponding to one of the off-target sites. Since the host cell can be repeatedly infected by various bacteriophages during the incubation step, it is possible to select for binding to or cleavage at multiple off-targets in a single round of negative selection.
[0229] In some embodiments, for example, a positive selection is performed to select a polypeptide or polynucleotide of interest that recognizes a specific recognition sequence, and then a negative selection is performed to select a polypeptide or polynucleotide of interest that recognizes one or more off-target sites. In some embodiments, a negative selection is performed to select a polypeptide or polynucleotide of interest that recognizes one or more off-target sites, and then a positive selection is performed to select a polypeptide or polynucleotide of interest that recognizes a specific recognition sequence.
[0230] In some embodiments where multiple rounds of selection (e.g., a positive selection round followed by a negative selection round) are used within one cycle, the method includes the step of pelleting the host cells (e.g., by centrifugation) between rounds of selection. Such a pelleting step can, for example, remove agents (e.g., antibiotics, inducers of gene expression such as IPTG) used during previous selection rounds.
[0231] The mutant nucleases identified using the methods described herein can be used to genetically engineer a population of cells. To alter or engineer a population of cells, the cells may be contacted with a mutant nuclease described herein, or a vector capable of expressing the mutant nuclease, and a guide nucleic acid. As is known in the art, the guide nucleic acid will have a region complementary to a target sequence on a target nucleic acid of the cell's genome. In some embodiments, the mutant nuclease and the guide nucleic acid are administered as a ribonucleoprotein (RNP). In some embodiments, the RNP is administered at a dose of 1×10 -4 μM to 1 μM of RNP. In some embodiments, less than 1%, 5%, 10%, 15%, or 20% of the alterations constitute alterations of off-target sequences within the population of cells. In some embodiments, greater than 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% of the alterations constitute alterations of on-target sequences within the population of cells.
[0232] The mutant nucleases identified using the methods described herein can be used to edit a population of double-stranded DNA (dsDNA) molecules. To edit a population of dsDNA molecules, the molecules may be contacted with a mutant nuclease and a guide nucleic acid described herein. As is known in the art, the guide nucleic acid will have a region complementary to the target sequence of the dsDNA. In some embodiments, the mutant nuclease and the guide nucleic acid are administered as a ribonucleoprotein (RNP). In some embodiments, the RNP is administered at a dose of 1×10 -4 μM to 1 μM of RNP. In some embodiments, less than 1%, 5%, 10%, 15%, or 20% of the editing constitutes editing of off-target sequences within the population of dsDNA molecules. In some embodiments, greater than 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% of the editing constitutes editing of on-target sequences within the population of dsDNA molecules.
[0233] Implementation of Genome Editing Systems: Delivery, Formulation, and Routes of Administration As discussed above, the genome editing systems of the present disclosure can be implemented in any suitable manner, which means that the components of such systems, including but not limited to RNA-guided nucleases (e.g., RNA-guided nuclease variants described herein), gRNAs, and optional donor template nucleic acids, can be delivered, formulated, or administered in any suitable form or combination of forms that result in the transduction, expression, or introduction of the genome editing system and / or cause the desired therapeutic outcome (e.g., ex vivo and / or in vivo) in cells, tissues, or subjects. Tables 2 and 3 show some non-limiting examples of the implementation of genome editing systems. However, those skilled in the art will understand that these lists are not exhaustive and that other implementations are possible. In some embodiments, one or more of the components described herein are delivered / administered in vivo. In some embodiments, one or more of the components described herein are delivered / administered ex vivo. For example, in some embodiments, an RNA (e.g., mRNA) encoding an RNA-guided nuclease variant described herein is delivered / administered to cells in vivo or ex vivo. In some embodiments, an RNA-guided nuclease variant described herein is delivered / administered to cells in vivo or ex vivo as a ribonucleoprotein (RNP) complex with a gRNA or without a gRNA. With particular reference to Table 2, the table lists some exemplary implementations of genome editing systems that include a single gRNA and an optional donor template. However, genome editing systems according to the present disclosure may incorporate multiple gRNAs, multiple RNA-guided nucleases (e.g., multiple RNA-guided nuclease variants described herein), and other components, such as proteins, and various implementations will be apparent to those skilled in the art based on the principles shown in Table 2. In Table 2, "[N / A]" indicates that the genome editing system does not include the indicated component.
[0234] [Table 2]
[0235]
Table 3
[0236] Table 3 summarizes various delivery methods for the components of the genome editing system, as described herein. Again, the list is intended to be illustrative rather than limiting.
[0237]
Table 4
[0238]
Table 5
[0239] Nucleic acid-based delivery of the genome editing system Nucleic acids encoding various elements of the genome editing system according to the present disclosure can be administered to a subject or delivered into cells by methods known in the art or as described herein. For example, DNA encoding an RNA-guided nuclease (e.g., an RNA-guided nuclease variant described herein) and / or encoding a gRNA, as well as donor template nucleic acids, can be delivered, for example, by a vector (e.g., a viral vector or a non-viral vector), a non-vector-based method (e.g., using naked DNA or DNA complexes), or a combination thereof.
[0240] Nucleic acids encoding the genome editing system or its components can be delivered directly into cells as naked DNA or RNA (e.g., mRNA), for example, by transfection or electroporation, or may be conjugated to a molecule (e.g., N-acetylgalactosamine) that promotes uptake by target cells (e.g., erythrocytes, HSCs). Nucleic acid vectors, such as those summarized in Table 3, may be used.
[0241] The nucleic acid vector may contain one or more sequences encoding genome editing system components, such as RNA-guided nucleases (e.g., RNA-guided nuclease variants described herein), gRNAs, and / or donor templates. The vector may also contain a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization) associated with (e.g., inserted into, fused to) a protein-encoding sequence. As an example, the nucleic acid vector may contain a Cas9-coding sequence containing one or more nuclear localization sequences (e.g., derived from SV40).
[0242] The nucleic acid vector may also contain any suitable number of regulatory / control elements, such as promoters, enhancers, introns, polyadenylation signals, Kozak consensus sequences, or internal ribosome entry sites (IRES). Such elements are well known in the art and are described in Cotta-Ramusino.
[0243] The nucleic acid vectors according to the present disclosure include recombinant viral vectors. Exemplary viral vectors are shown in Table 3, and additional suitable viral vectors, as well as their use and production, are described in Cotta-Ramusino. Other viral vectors known in the art may also be used. Also, viral particles may be used to deliver genome editing system components in nucleic acid and / or peptide form. For example, "empty" viral particles can be assembled to contain any suitable cargo. Viral vectors and viral particles can also be engineered to incorporate targeting ligands to alter target tissue specificity.
[0244] In addition to viral vectors, non-viral vectors can be used to deliver nucleic acids encoding a genome editing system according to the present disclosure. One important category of non-viral nucleic acid vectors are nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art and are summarized in Cotta-Ramusino. Any suitable nanoparticle design can be used to deliver genome editing system components or nucleic acids encoding such components. For example, organic (e.g., lipid and / or polymer) nanoparticles may be suitable for use as delivery vehicles in certain embodiments of the present disclosure. Exemplary lipids used in nanoparticle formulations and / or gene transfer are shown in Table 4, and Table 5 lists exemplary polymers used in gene transfer and / or nanoparticle formulations.
[0245]
Table 6
[0246]
Table 7
[0247]
Table 8
[0248] Non-viral vectors may, in some cases, include targeting modifications that improve uptake and / or selectively target specific cell types. Such targeting modifications can include, for example, cell-specific antigens, monoclonal antibodies, single-chain antibodies, aptamers, polymers, sugars (such as N-acetylgalactosamine (GalNAc)), and cell-penetrating peptides. Such vectors may also, in some cases, use fusogenic and endosome-destabilizing peptides / polymers to undergo acid-triggered conformational changes (e.g., to facilitate endosomal escape of the cargo) and / or incorporate stimuli-cleavable polymers, for example, for release within cellular compartments. For example, disulfide-based cationic polymers that are cleaved within a reducing cellular environment may be used.
[0249] In certain embodiments, one or more nucleic acid molecules (e.g., DNA molecules) other than the components of the genome editing system, such as the RNA-guided nuclease components and / or gRNA components described herein, are delivered. In one embodiment, the nucleic acid molecule is delivered simultaneously with the delivery of one or more components of the genome editing system. In one embodiment, the nucleic acid molecule is delivered before or after (e.g., within about 30 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 9 hours, 12 hours, 1 day, 2 days, 3 days, 1 week, 2 weeks, or less than 4 weeks) the delivery of one or more components of the genome editing system. In one embodiment, the nucleic acid molecule is delivered by a means different from that by which one or more components of the genome editing system, such as the RNA-guided nuclease component and / or gRNA component, are delivered. The nucleic acid molecule may be delivered by any of the delivery methods described herein. For example, the nucleic acid molecule may be delivered by a viral vector, such as an integration-deficient lentivirus, and the RNA-guided nuclease molecular component and / or gRNA component may be delivered by electroporation, for example, so as to be able to reduce the toxicity caused by the nucleic acid (e.g., DNA). In one embodiment, the nucleic acid molecule encodes a therapeutic protein, such as a protein described herein. In one embodiment, the nucleic acid molecule encodes an RNA molecule, such as an RNA molecule described herein.
[0250] Delivery of RNP and / or RNA Encoding Genome Editing System Components RNP (a complex of gRNA and an RNA-guided nuclease (e.g., an RNA-guided nuclease variant described herein)), and / or an RNA-guided nuclease (e.g., an RNA-guided nuclease variant described herein) and / or an RNA (e.g., mRNA) encoding gRNA can be delivered into cells or administered to a subject by methods known in the art, some of which are described in Cotta-Ramusino. In vitro, an RNA-guided nuclease (e.g., an RNA-guided nuclease variant described herein) and / or an RNA (e.g., mRNA) encoding gRNA can be delivered, for example, by microinjection, electroporation, transient cell compression or squeezing (see, e.g., Lee 2012). Also, lipid-mediated transfection, peptide-mediated delivery, GalNAc- or other conjugate-mediated delivery, and combinations thereof may be used for delivery in vitro and in vivo.
[0251] Delivery via electroporation in vitro is defined as mixing cells, with or without an RNA-guided nuclease and / or an RNA (e.g., mRNA) encoding gRNA, and a donor template nucleic acid molecule in a cartridge, chamber, or cuvette, and applying one or more electrical impulses of defined time and amplitude. Systems and protocols for electroporation are known in the art, and any suitable electroporation tool and / or protocol may be used in connection with various embodiments of the present disclosure.
[0252] Route of administration A genome editing system, or a cell modified or manipulated using such a system, can be administered to a subject by any suitable mode or route, whether locally or systemically. Systemic administration modes include oral and parenteral routes. Examples of parenteral routes include intravenous, intramedullary, intraarterial, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. Components to be systemically administered may be modified or formulated to target, for example, HSCs, hematopoietic stem / progenitor cells, or erythroid precursors or progenitor cells.
[0253] Examples of local administration modes include intramedullary injection into the vertebral column or intramuscular injection into the thigh into the marrow cavity, and injection into the portal vein. In one embodiment, a significantly smaller amount of the component (compared to a systemic approach) may be effective when administered locally (e.g., directly into the bone marrow) compared to when administered systemically (e.g., intravenously). The local administration mode can reduce or eliminate the incidence of potentially toxic side effects that may occur when a therapeutically effective amount of the component is administered systemically.
[0254] Administration may be achieved as a periodic bolus (e.g., intravenously) or as a continuous infusion from an internal reservoir or an external reservoir (e.g., an intravenous bag or an implantable pump). Components may be administered locally, for example, by continuous release from a sustained release drug delivery device.
[0255] In addition, the component may be formulated to enable long-term release. Examples of release systems may include biodegradable substances or matrices of substances that release incorporated components by diffusion. The components can be distributed homogeneously or heterogeneously within the release system. Various release systems may be useful. However, the selection of an appropriate system will depend on the rate of release required for a particular application. Both non-degradable and degradable release systems can be used. Suitable release systems include polymers and polymer matrices, non-polymer matrices, or inorganic and organic excipients and diluents, such as, but not limited to, calcium carbonate and sugars (e.g., trehalose). The release system may be natural or synthetic. However, synthetic release systems are preferred because they generally provide a more reliable, reproducible, and well-defined release profile. The release system material can be selected such that components of different molecular weights are released by diffusion of the substance or through degradation.
[0256] Representative synthetic, biodegradable polymers include, for example: polyamides such as poly(amino acids) and poly(peptides); polyesters such as poly(lactic acid), poly(glycolic acid), poly(lactic-co-glycolic acid), and poly(caprolactone); poly(anhydrides); polyorthoesters; polycarbonates; and their chemical derivatives (substituents, chemical groups such as addition of alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof. Representative synthetic, non-degradable polymers include, for example: polyethers such as poly(ethylene oxide), poly(ethylene glycol), and poly(tetramethylene oxide); vinyl polymers - polyacrylates and polymethacrylates such as methyl, ethyl, other alkyl, hydroxyethyl methacrylate, acrylic acid and methacrylic acid, and others such as poly(vinyl alcohol), poly(vinyl pyrrolidone), and poly(vinyl acetate); poly(urethanes); cellulose and its derivatives such as alkyl, hydroxyalkyl, ether, ester, nitrocellulose, and various cellulose acetates; polysiloxanes; and any of their chemical derivatives (substituents, chemical groups such as addition of alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof.
[0257] Poly(lactide-co-glycolide) microspheres can also be used. Typically, the microspheres are composed of polymers of lactic acid and glycolic acid, which are configured to form hollow spheres. The spheres can have a diameter of approximately 15 - 30 microns and can load the components described herein.
[0258] Multimodal or differential delivery of components One skilled in the art will understand that the various components of the genome editing system may be delivered together or separately, and simultaneously or non-simultaneously. Separate and / or asynchronous delivery of the genome editing system components may be desired, particularly to achieve temporal or spatial control over the function of the genome editing system and to limit the specific effects caused by its activity.
[0259] As used herein, different modes or differential modes refer to delivery modes that confer various pharmacodynamic or pharmacokinetic properties to the target component molecules, such as RNA-guided nuclease molecules, gRNAs, template nucleic acids, or payloads. For example, the delivery mode can result in different tissue distributions, different half-lives, or different temporal distributions, for example, in a selected compartment, tissue, or organ.
[0260] Some modes of delivery, such as delivery into cells or by nucleic acid vectors that persist in the progeny of cells, for example, by autonomous replication or insertion into cellular nucleic acids, result in more persistent expression and presence of the components. Examples include delivery of viruses such as AAV or lentivirus.
[0261] As an example, the components of the genome editing system, such as RNA-guided nucleases and gRNAs, may be delivered by different modes with respect to the resulting half-life or persistence of the delivered components in the body or in a specific compartment, tissue, or organ. In one embodiment, the gRNA may be delivered by such a mode. The RNA-guided nuclease molecule components may be delivered by a mode with less persistence in the body or in a specific compartment, tissue, or organ, or less exposure to them.
[0262] More generally, in one embodiment, a first delivery mode is used to deliver a first component, and a second delivery mode is used to deliver a second component. The first delivery mode confers a first pharmacodynamic or pharmacokinetic property. The first pharmacodynamic property may be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component within the body, compartment, tissue, or organ. The second delivery mode confers a second pharmacodynamic or pharmacokinetic property. The second pharmacodynamic property may be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component within the body, compartment, tissue, or organ.
[0263] In certain embodiments, the first pharmacodynamic or pharmacokinetic property, such as distribution, persistence, or exposure, is more restricted than the second pharmacodynamic or pharmacokinetic property.
[0264] In certain embodiments, the first delivery mode is selected to optimize, for example minimize, a pharmacodynamic or pharmacokinetic property, such as distribution, persistence, or exposure.
[0265] In certain embodiments, the second delivery mode is selected to optimize, for example maximize, a pharmacodynamic or pharmacokinetic property, such as distribution, persistence, or exposure.
[0266] In certain embodiments, the first delivery mode includes the use of relatively persistent elements, such as nucleic acids, such as plasmids or viral vectors, such as AAV or lentivirus. Since such vectors are relatively persistent, the products transcribed from the vectors will be relatively persistent.
[0267] In certain embodiments, the second delivery mode includes relatively transient elements, such as RNA or protein.
[0268] In certain embodiments, the first component comprises the gRNA and the delivery mode is relatively persistent. For example, the gRNA is transcribed from a plasmid or viral vector, such as AAV or lentivirus. Transcription of these genes would have little physiological consequence, since the genes do not encode a protein product and the gRNA cannot function in isolation. The second component, the RNA-guided nuclease molecule, is delivered in a transient manner, for example, as mRNA encoding a protein or as a protein. This ensures that the complete RNA-guided nuclease molecule / gRNA complex exists and is active for only a short period of time.
[0269] Furthermore, the components can be delivered in various molecular forms or by different delivery vectors that complement each other to enhance safety and tissue specificity.
[0270] The use of differential delivery modes can enhance performance, safety, and / or efficacy. For example, the likelihood of final off-target modification can be reduced. Delivery of immunogenic components, such as the Cas9 molecule, by a less persistent mode can reduce immunogenicity, since peptides of the bacterially derived Cas enzyme are presented on the cell surface by MHC molecules. A bipartite delivery system can mitigate these drawbacks.
[0271] The differential delivery mode can be used to deliver components to different but overlapping target regions. The formation of the active complex is minimized outside the overlapping portion of the target region. Thus, in one embodiment, a first component, e.g., a gRNA, is delivered by a first delivery mode to result in a first spatial, e.g., tissue, distribution. A second component, e.g., an RNA-guided nuclease molecule, is delivered by a second delivery mode to result in a second spatial, e.g., tissue, distribution. In one embodiment, the first mode comprises a first element selected from liposomes, nanoparticles, e.g., polymeric nanoparticles, and nucleic acids, e.g., viral vectors. The second mode comprises a second element selected from the group. In one embodiment, the first delivery mode comprises a first targeting element, e.g., a cell-specific receptor or an antibody, and the second delivery mode does not comprise such an element. In certain embodiments, the second delivery mode comprises a second targeting element, e.g., a second cell-specific receptor or a second antibody.
[0272] When an RNA-guided nuclease molecule is incorporated into and delivered by a viral delivery vector, liposome, or polymeric nanoparticle, there is a potential for delivery to multiple tissues and therapeutic activity within multiple tissues, if it is desired to target only a single tissue. A bipartite delivery system can solve this problem and enhance tissue specificity. If the gRNA and the RNA-guided nuclease molecule are packaged within separate delivery vehicles that have different but overlapping tissue tropisms, a fully functional complex is formed only within the tissue targeted by both vectors.
[0273] All publications, patent applications, patents, and other references mentioned in this specification are hereby incorporated by reference in their entirety. Further, the materials, methods, and examples are illustrative only and not intended to be limiting. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described herein.
[0274] This disclosure is further illustrated by the following examples. These examples are provided for illustrative purposes only and are not to be construed as limiting the scope or content of this disclosure in any way.
Examples
[0275] Example 1: Evolution of Allele-Specific Cas9 to a Single Base Pair Mutation Conferring Cone-Rod Dystrophy 6 (CORD6) This example demonstrates that the selection method of the present invention can be used in an evolutionary strategy to evolve a site-specific nuclease that is specific for a disease allele that differs only by a point mutation (a single base change) compared to the wild-type non-disease allele.
[0276] When Cas9 or other targeted nucleases are used for allele-specific cleavage of a heterozygous sequence, this use is hampered by erratic activity, especially when the alleles differ by a single base. The inventors sought to engineer a Cas9 mutant that could selectively cleave only one allele, specifically an allele having an R838S mutation within the retinal guanylate cyclase (GUCY2D) protein that confers the CORD6 disease phenotype. The inventors constructed plasmid pEvol_CORD6. This encodes the Cas9 protein and a gRNA targeting the CORD6 sequence TAACCTGGAGGATCTGATCC (SEQ ID NO: 1). pEvol_CORD6 also constitutively expresses beta-lactamase, which confers resistance to ampicillin. Two phagemids (plasmids containing the phage origin f1 element), pSelect_CORD6 and pSelect_GUCY2DWT, were also constructed. These contained the potential target sites
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
[0277] pSelect_CORD6 and pSelect_GUCY2DWT also each contain a constitutively expressed chloramphenicol resistance gene and ccdB (a bacterial toxin) under the control of the lac promoter, which allows for induction of ccdB expression by IPTG (isopropyl β-D-1-thiogalactopyranoside).
[0278] pSelect_CORD6 and pSelect_GUCY2DWT were separately packaged into helper bacteriophages.
[0279] To engineer allele specificity, two Escherichia coli (E. coli) bacterial libraries of Cas9 mutants were generated using the pEvol_CORD6 plasmid as the initial template for mutagenesis, targeting all codons and using an inclusive and unbiased mutagenesis method that allows for adjustment of the mutation rate. One library was adjusted to have a median of three amino acid mutations per Cas9 polypeptide ("low" mutation rate). The other had a median of five amino acid mutations per Cas9 polypeptide ("high" mutation rate).
[0280] In each round of evolution, the inventors exposed each bacterial library of pEvol_CORD6 mutants, first, to positive selection for cleavage of phage-containing pSelect_CORD6 in competitive cultures by serial sensitization induction by phage, and then, to negative selection for cleavage of pSelect_GUCY2DWT, as follows:
[0281] To infect the bacteria, phages packaging the appropriate pSelect plasmid were added to a saturated culture of bacteria containing a library of pEvol_CORD6 mutants, and the bacterial library was liquid cultured in ampicillin. For each library, the entire library was cultured in the same liquid culture.
[0282] After this initial incubation and infection, positive selection was performed by adding 1 mM IPTG, which induces ccdB. Next, the cultures were grown overnight, for example, for at least 12 hours. The cells were then pelleted, removing some of the IPTG. Negative selection was then performed by growing the bacteria in the presence of 50 μg / ml chloramphenicol (constitutively expressed by both pSelect plasmids) and in the absence of IPTG during a second overnight culture. Between both positive and negative selections, the bacteria were continuously infected by phages present in the liquid culture, presenting a continuous sensitization induction to either cleavage (in the case of positive selection) or non-cleavage (in the case of negative selection).
[0283] Pooled plasmid DNA from all selected library members after negative selection was used as a template for the next mutagenesis reaction. The inventors thus repeated three rounds of mutagenesis (generating a library), positive selection, and negative selection. By applying a dual selection pressure to each library, stringent selection was performed for Cas9 mutants containing a PAM specific to the CORD6 allele.
[0284] PacBio next-generation sequencing of plasmid DNA isolated from pooled selected library members was performed after each round of negative selection and after each evolutionary cycle. Only after the first evolutionary cycle, the inventors found that a particular mutant accounted for approximately 20% of the population and exhibited a high selection intensity. The inventors advanced testing of the cleavage activity of this mutant using an E. coli cell lysate containing the mutant protein on amplicons containing either the wild-type or mutated GUCY2D sequence. The inventors observed cleavage only on the CORD6 amplicon (Figure 1). Further analysis of the PAM selection of this mutant also showed two-fold higher specificity for A rather than position 6 C. The activity of this highly selected mutant confirms the designed selection pressure and demonstrates the success of the unbiased mutagenesis method and the manipulation of allele-specific Cas9 mutants by the competitive selection strategy.
[0285] Example 2: Evolution of Cas9 with reduced off-target activity using known off-targets This example describes how the selection method of the present invention can be used in an evolutionary strategy to reduce the off-target activity of site-specific DNA binding enzymes.
[0286] Off-target cleavage is a common side effect of Cas9 target DNA cleavage. To mitigate this effect, the selection of on-target cleavage ("positive selection") can be coupled with the selection of known or potential off-target sequences ("negative selection") in the inventors' system. Off-targets discovered by GUIDE-SEQ or other methods can be counter-selected based on information. Additionally, a library of potential off-targets such as single base pair mismatches can also be selected. Thus, by combining specific guides with Cas9 evolved to reduce off-target cleavage, it can be tuned to selectively cleave at the appropriate site.
[0287] In this case, the evolution proceeds by first selecting on-target cleavage in positive selection and then negative selecting a mixed phage population of designated off-targets, followed by optional deep sequencing confirmation (Figure 2). This evolutionary algorithm may be repeated over several rounds.
[0288] Example 3: Evolution of allele-specific Cas9 This example demonstrates that the selection method of the present invention can be used in an evolutionary strategy to evolve a site-specific nuclease that is specific for a disease allele that differs only by a point mutation (a single base change) compared to the wild-type non-disease allele. However, using the selection method of the present invention, a site-specific nuclease that is specific for an allele that differs more than a single base change (e.g., a mutant or disease allele) compared to another allele (e.g., a wild-type or non-disease allele) can also be evolved.
[0289] When Cas9 or other target nucleases are used for allele-specific cleavage of a heterozygous sequence, this use is hampered by erratic activity, especially when the alleles differ by a single base. The inventors sought to engineer a Cas9 mutant that selectively cleaves only one allele (e.g., allele 1, the mutant allele) and not alleles that differ by a single base (e.g., allele 2, the wild-type allele). The inventors also sought to improve the efficiency of a method for selecting Cas9 mutants to achieve the greatest discrimination for cleavage of one allele. Selection of the most discriminatory Cas9 mutants can be achieved, for example, by controlling the amount of Cas9 present within the selection process and / or by improving the efficiency of positive and negative selection. The amount of Cas9 in the selection system can be controlled, for example, by use of a plasmid expressing Cas9 at a low copy number. In some embodiments, the amount of Cas9 is controlled by placing the expression of Cas9 under the control of an inducible promoter. In some embodiments, the inducible promoter is an arabinose promoter (Figure 3).
[0290] The processes of positive and negative selection can rely on the inducible expression of a toxin molecule and / or the expression of resistance to a drug such as an antibiotic. For example, if the expression of the toxin is plasmid-derived, only cells containing a Cas9 mutant that recognizes and cleaves an appropriate target (e.g., allele 1, the mutant allele) within the plasmid will survive. Cells containing Cas9 that does not recognize and cleave the appropriate target will die by the toxin. This positive selection step selects for all Cas9 molecules that can recognize and cleave an appropriate target (e.g., allele 1, the mutant allele).
[0291] In another embodiment, when cells are treated with an antibiotic, only the cells containing Cas9 that cleaves an inappropriate target within a plasmid conferring resistance to the antibiotic will die. Cells containing Cas9 that do not recognize the inappropriate target (e.g., allele 2, wild-type allele) will survive, maintaining resistance to the antibiotic. This negative selection step selects for Cas9 molecules that can recognize and cleave inappropriate targets (e.g., allele 2, wild-type allele).
[0292] The usefulness of the positive and negative selection steps for the identification of highly selective Cas9 molecules depends, at least in part, on a high degree of discrimination in cell death. Comparison of the cell growth kinetics during selection can characterize the selection efficiency of the optimal Cas9 molecule.
[0293] Selection efficiency using tse2 A plasmid, pEvol_CAS, encoding Cas9 protein and a gRNA targeting a target sequence was constructed. A plasmid, pEvol_NONTARGETING, encoding Cas9 protein and a non-targeting gRNA was also constructed. Both plasmids constitutively express beta-lactamase conferring resistance to ampicillin (AmpR) and an inducible arabinose promoter (Ara) controlling the expression of Cas9. Phagemids (plasmids containing the phage origin f1 element), pSelect_MUT and pSelect_WT, were also constructed. These each contain potential target sites. The phagemids also contain a constitutively expressed chloramphenicol resistance gene (CmR) and tse2 (bacterial toxin) under the control of the lac promoter, which allows induction of tse2 expression by IPTG (isopropyl β-D-1-thiogalactopyranoside). pSelect_MUT and pSelect_WT were each separately packaged into helper bacteriophage.
[0294] To engineer allele specificity, two Escherichia coli (E. coli) bacterial libraries of Cas9 mutants were generated using the pEvol_CAS plasmid as the first template for mutagenesis, with a comprehensive and unbiased mutagenesis method that targets all codons and allows for the adjustment of the mutation rate. One library was adjusted to have a median of three amino acid mutations per Cas9 polypeptide (“low” mutation rate). The other had a median of five amino acid mutations per Cas9 polypeptide (“high” mutation rate).
[0295] In each round of evolution, the inventors exposed each bacterial library of pEvol_CAS mutants to positive selection for cleavage of phage-containing pSelect_MUT in competitive cultures by successive sensitization induction with phage as follows:
[0296] To infect the bacteria, phage packaging the pSelect_MUT plasmid was added to bacteria containing the pEvol_CAS mutant or pEvol_NONTARGETING library, and the bacterial library was liquid cultured in ampicillin. For each library, the entire library was cultured in the same liquid culture.
[0297] After this first incubation and infection, the stringency of positive selection with tse2 was evaluated by adding 1 mM IPTG, which induces tse2 expression, to subsets of the pEvol_CAS cultures and to subsets of the pEvol_NONTARGETING cultures. Expression of Cas9 and guide RNA was induced by addition of arabinose. Expression of Cas9 and guide RNA was not induced in subsets of the pEvol_CAS cultures treated with IPTG. Next, the cultures were grown overnight, for example for at least 12 hours. During positive selection, the bacteria were continuously infected with phage present in the liquid culture to present successive sensitization induction to target cleavage.
[0298] As shown in Fig. 4, cultures that did not express Cas9 but expressed tse2 (-Cas + tse2), or expressed a non-targeting guide RNA (+Nontargeting Cas + tse2), showed significant growth delays due to the induction of tse2. In comparison, cultures induced to express tse2 but also express Cas9 and a targeting guide RNA (+Cas + tse2) (expected to cleave the target and suppress the expression of tse2) demonstrated rapid growth over approximately 7 hours. Cultures that expressed either Cas9 and either a targeting or non-targeting guide RNA but were not induced to express tse2 demonstrated rapid cell growth over the first 6 hours. These data demonstrate that in the absence of Cas9 or when the guide RNA does not recognize the target, tse2 has a significant cell-killing effect. These data also demonstrate that appropriately targeted Cas9 and guide RNA offset the effect of the induction of tse2 expression.
[0299] Efficiency of selection by modulating Cas9 expression Using the pEvol_WTCAS plasmid as the first template for mutagenesis and an inclusive and unbiased mutagenesis method that targets all codons and allows adjustment of the mutation rate, a plasmid library, pEvol_CASLIBRARY, was generated. The plasmid encodes the Cas9 protein and a gRNA that targets the target sequence. A plasmid pEvol_WTCAS encoding the wild-type Cas9 protein and the targeting gRNA was also constructed. Both plasmids constitutively express beta-lactamase, which confers resistance to ampicillin (AmpR), and an inducible arabinose promoter (Ara) that controls the expression of Cas9. As described above, phagemids (plasmids containing the phage origin f1 element), pSelect_MUT and pSelect_WT, were also constructed. These contain potential target sites.
[0300] To infect bacteria, the phage packaging the pSelect_MUT plasmid was added to saturated bacteria containing a library of pEvol_CASLIBRARY mutants or pEvol_WTCAS, and the bacterial library was liquid cultured in ampicillin. For each library, the entire library was cultured in the same liquid culture.
[0301] After this first incubation and infection, the stringency of positive selection using tse2 and wild-type Cas or a Cas library was evaluated by adding 1 mM IPTG, which induces tse2 expression, to subsets of the pEvol_CASLIBRARY cultures and to subsets of the pEvol_WTCAS cultures. Expression of Cas9 and gRNA was induced by the addition of arabinose. Expression of Cas9 and gRNA was not induced in subsets of the pEvol_CASLIBRARY cultures and pEvol_WT CAS cultures treated with IPTG. Next, the cultures were grown overnight, for example for at least 12 hours. During positive selection, the bacteria were continuously infected by the phage present in the liquid culture to present continuous sensitization induction to target cleavage.
[0302] As shown in Fig. 5, cultures expressing tse2 but not wild-type Cas9 (-WTCas + tse2) or not expressing the mutant Cas9 library (-Cas Library + tse2) showed significant growth delays due to the induction of tse2. However, wild-type Cas9 cultures showed a greater growth delay than mutant Cas9 library cultures, which showed leaky expression of Cas9 mutants even in the absence of arabinose. In comparison, cultures induced to express tse2 as well as Cas9 and the targeting guide RNA (+WTcas + tse2 or +Cas Library + tse2), which were expected to cleave the target and suppress the expression of tse2, demonstrated rapid growth over approximately 7 hours. Cultures expressing either wild-type Cas9 or a Cas9 library mutant but not induced to express tse2 demonstrated rapid cell growth over the first 6 hours. The difference in cell growth between cultures expressing wild-type Cas9 with or without tse2 was smaller than the difference in cell growth between cultures expressing the Cas9 library mutant with or without tse2. These data suggest that the Cas9 library mutants exhibit greater cleavage activity than wild-type Cas9. These data also confirmed that in the absence of Cas9, tse2 has a significant cell-killing effect.
[0303] Negative selection was also performed by growing the bacteria in the presence of 50 μg / ml chloramphenicol (resistance to chloramphenicol is constitutively expressed by the pSelect_MUT phagemid and the pSelect_WT phagemid) during an overnight culture. Control cultures were not treated with chloramphenicol. During negative selection, the bacteria were continuously infected with the phages present in the liquid culture to present continuous sensitization to appropriate target cleavage (allele 1, mutant allele) and inappropriate target non-cleavage (allele 2, wild-type allele). Both wild-type Cas9 (WTCas+Cm) and the mutant Cas9 library (Library+Cm) showed a significant growth delay due to the removal of chloramphenicol resistance by off-target cleavage (Figure 6). However, the mutant Cas9 library mutants demonstrated the recovery of growth resulting from the selection of Cas9 mutants that did not show off-target cleavage and maintained chloramphenicol resistance.
[0304] Library evolution Successive rounds of library evolution generated Cas9 mutants with a high level of selectivity for cleaving the target. This was demonstrated by the successive decrease in the growth delay when the culture was induced to express tse2. Figure 7 shows a significant growth delay and a significant negative delta for the wild-type Cas9 culture (WTcas+tse2) when tse2 was induced, compared to the growth of a culture expressing wild-type Cas9 without induction of tse2 (WTcas-tse2). The delta of cell growth decreased after one round of mutagenesis (e.g., Round 1+tse2 vs. Round 1-tse2). After three rounds of mutagenesis, the cell growth curves were nearly identical for the cultures induced to express tse2 and those not induced to express tse2. These data indicate that the use of tse2 and the selective rounds of mutagenesis can generate mutant Cas9 that is highly selective for on-target cleavage.
[0305] Example 4: Evolution of Streptococcus pyogenes Cas9 to Reduce Off-Target Cleavage Using the systems and methods of the present disclosure, various nuclease properties can be selected, as further illustrated by the following examples. Using the selection methods disclosed herein, Streptococcus pyogenes cas9 mutants were identified from a mutagenized library that maintained on-target cleavage efficiency but reduced cleavage at off-target loci, providing a means to potentially rescue promiscuous guides for therapeutic use. A mutagenized Streptococcus pyogenes Cas9 library was generated using random target scanning mutagenesis (SMART). The library was then transformed into Escherichia coli and phage-induced for three rounds of both positive and negative selection. The positive selection step utilized a positive selection plasmid containing a cleavage cassette with an on-target sequence (SEQ ID NO: 4) for a guide RNA directed to a human genomic locus with multiple known off-targets as determined by GUIDE-Seq, as shown in Table 6:
[0306] [Table 9]
[0307] A single negative selection step utilized four pooled constructs, each containing a unique off-target that differed from the on-target sequence by four residues (SEQ ID NOs: 5-7). After the positive and negative selection steps, clones were selected, sequenced by next-generation sequencing (NGS), reads were aligned, and the most commonly mutated amino acid residues were identified by comparison to wild-type Streptococcus pyogenes Cas9 (SEQ ID NO: 13). In vitro cleavage assays using on-target and off-target substrates demonstrated that the clones exhibited on-target cleavage efficiency comparable to WT Cas9 (Figure 8A), while a few clones showed reduced off-target cleavage compared to WT, and other clones showed substantially the same off-target cleavage efficiency in vitro (Figure 8B). Also, on-target and off-target analyses were performed on on-target and off-target genomic loci in human T cells as shown in Figures 9A and 9B. Wild-type and mutant Cas9 / guide RNA ribonucleoprotein complexes (RNPs) were delivered at various concentrations over a >2-log range, genomic DNA was harvested, on-target and off-target sites were amplified, and sequenced by next-generation sequencing. As shown in Figure 10, some mutant clones showed slightly reduced on-target cleavage activity compared to WT Cas9, while also showing substantially lower off-target cleavage than WT. Together, these data confirm that the phage-selection method described herein can be successfully applied to reduce off-target cleavage observed with specific gRNAs by selecting compensating Cas9 mutant proteins.
[0308] Table 2 shows the selected amino acid residues that are mutated within the clones identified in this screen, and the residues that can be substituted at each position to generate mutants with reduced off-target activity:
[0309]
Table 10
[0310] Table 8 shows exemplary single, double, and triple Streptococcus pyogenes (S. pyogenes) Cas9 mutants according to certain embodiments of the present disclosure. For clarity, only single, double, and triple mutants are listed in the table to streamline the presentation, but the present disclosure encompasses Cas9 mutant proteins having mutations at 1, 2, 3, 4, 5, or more of the sites shown in Table 8.
[0311] [Table 11]
[0312] [Table 12]
[0313] [Table 13]
[0314] [Table 14]
[0315] [Table 15]
[0316] Without limiting the foregoing, the present disclosure encompasses the following mutants:
[0317] [Table 16]
[0318] The present disclosure also encompasses a genome editing system comprising a mutated Streptococcus pyogenes (S. pyogenes) Cas9 described herein.
[0319] The isolated SpCas9 mutant proteins described herein are, in certain embodiments of the present disclosure, fused to a heterologous functional domain, optionally via a linker that does not interfere with the activity of the fusion protein. In some embodiments, the heterologous functional domain is a transcriptional activation domain. In some embodiments, the transcriptional activation domain is derived from VP64 or NF-kappaB p65. In some embodiments, the heterologous functional domain is a transcriptional silencer or transcriptional repressor domain. In some embodiments, the transcriptional repressor domain is a Krueppel-associated box (KRAB) domain, an ERF repressor domain (ERD), or an mSin3A interaction domain (SID). In some embodiments, the transcriptional silencer is a heterochromatin protein 1 (HP1), such as HP1 alpha, or HP1 beta. In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein. In some embodiments, the TET protein is TET1. In some embodiments, the heterologous functional domain is an enzyme that modifies a histone subunit. In some embodiments, the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), a histone deacetylase (HDAC), a histone methyltransferase (HMT), or a histone demethylase. In some embodiments, the heterologous functional domain is a biological tether. In some embodiments, the biological tether is MS2, Csy4, or lambda N protein. In some embodiments, the heterologous functional domain is FokI.
[0320] In addition to the inclusion of the isolated nucleic acids encoding the mutant SpCas9 proteins described herein, the present disclosure encompasses both viral vectors and non-viral vectors containing such isolated nucleic acids. These are, in some cases, operably linked to one or more regulatory domains for expressing the mutant SpCas9 proteins described herein. The present disclosure also includes host cells, such as mammalian host cells, containing the nucleic acids described herein and, in some cases, expressing one or more of the mutant SpCas9 proteins described herein.
[0321] The mutant SpCas9 proteins described herein can be used to alter a cell's genome, for example, intracellularly, by expressing an isolated mutant SaCas9 or SpCas9 protein described herein and a guide RNA having a region complementary to a selected portion of the cell's genome. Alternatively, or in addition, the present disclosure further encompasses methods of altering, for example, selectively altering, a cell's genome by contacting the cell with a protein variant described herein and a guide RNA having a region complementary to a selected portion of the cell's genome. In some embodiments, the cell is a stem cell, such as an embryonic stem cell, a mesenchymal stem cell, or an induced pluripotent stem cell; is within a living animal; or is within an embryo, such as a mammalian, insect, or fish (e.g., zebrafish) embryo or embryonic cell.
[0322] In some embodiments, the isolated protein or fusion protein comprises one or more of a nuclear localization sequence, a cell membrane permeable peptide sequence, and / or an affinity tag.
[0323] Furthermore, the present disclosure encompasses methods of altering double-stranded DNA (dsDNA) molecules intracellularly, such as in vitro methods, ex vivo methods, and in vivo methods. The methods include contacting the dsDNA molecule with one or more of the mutant proteins described herein and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.
[0324] Example 5: Evolution of Streptococcus pyogenes Cas9 to Reduce Off-Target Cleavage Using the systems and methods of the present disclosure, various nuclease properties can be selected, as further illustrated by the following examples. Using the selection methods disclosed herein, Streptococcus pyogenes cas9 mutants were identified from a mutagenized library that maintained on-target cleavage efficiency but reduced cleavage at off-target loci, providing a means to potentially rescue promiscuous guides for therapeutic use. A mutagenized Streptococcus pyogenes Cas9 library was generated using random target scanning mutagenesis (SMART). The library was then transformed into Escherichia coli (E. coli) and phage-sensitized for three rounds of both positive and negative selection (Figures 1 and 2). After the positive and negative selection steps, clones were selected, sequenced by next-generation sequencing (NGS), the reads were aligned, and the most commonly mutated amino acid residues were identified compared to wild-type Streptococcus pyogenes Cas9 (SEQ ID NO: 13). The frequency of the identified mutations was determined by codon position and according to amino acid substitution (Figure 11). Mutations were identified within the RuvC domain (e.g., D23A), REC domain (e.g., T67L, Y128V), and PAM interaction domain (PI) (e.g., D1251G). Expression constructs containing combinations of four different mutations (D23A, T67L, Y128V, and D1251G) were prepared and tested in vitro for on-target and off-target editing efficiency in human T cells. In this example, the construct containing the mutations D23A, Y128V, D1251G, and T67L was called "Mut6" or "SpartaCas".
[0325] An in vitro dose-response study was performed for on-target and off-target loci of the genome in T cells using Mut6 and wild-type Streptococcus pyogenes Cas9, as shown in Figure 12. Wild-type or mutant Cas9 / guide RNA ribonucleoprotein complexes (RNPs) were delivered at various concentrations over a >2 log range. As shown in Figure 12, the Mut6 construct showed on-target cleavage activity comparable to wild-type Streptococcus pyogenes Cas9, while also showing substantially lower off-target cleavage than wild-type.
[0326] On-target editing at six different loci (SiteA, SiteB, SiteC, SiteD, SiteE, and SiteF) was evaluated using wild-type Streptococcus pyogenes Cas9 (WT SPCas9), "SpartaCas" (Mut 6), and two known mutant Streptococcus pyogenes Cas9 proteins, eCas (Slaymaker et al. Science (2015) 351:84-88) and HF1 Cas9 (Kleinstiver et al. Nature (2016) 529:490-495). Either 1 μm (loci 1 and 2) or 5 μm (loci 3-6) of either wild-type or mutant Cas9 / guide RNA RNP was delivered to human T cells. The editing efficiency was locus-dependent. The editing efficiency of SpartaCas was higher than that of HF1 Cas9 at all loci and higher than that of eCas at four out of six loci (Figure 13).
[0327] Furthermore, on-target editing dose-response studies were performed in human T cells using wild-type Streptococcus pyogenes Cas9 ("Spy"), wild-type Staphylococcus aureus Cas9 ("Sau"), wild-type Acidaminococcus sp. Cpf1 ("AsCpf1"), SpartaCas ("Mut6" or "S6"), and alternative mutant Streptococcus pyogenes Cas9 proteins (HF1 Cas9, eCas9, and Alt-R® Cas9 ("AltR Cas9") (www.idtdna.com)). On-target editing was target-dependent, and SpartaCas demonstrated high efficiency of on-target editing (Figures 14A-14C).
[0328] SpartaCas (Mut6) was further evaluated for on-target and off-target cleavage efficiency at RNP doses ranging from 0.03125 μM to 4 μM. On-target cleavage was comparable to wild-type Streptococcus pyogenes Cas9, while off-target cleavage was significantly reduced (Figure 15).
[0329] Example 6: Evaluation of off-target cleavage by Streptococcus pyogenes Cas9 mutants using GUIDE-seq Off-target cleavage by wild-type Cas9, eCas, and SpartaCas was evaluated using GUIDE-Seq using the following method.
[0330] RNP complex formation The RNP was complexed with a two-part gRNA synthesized by Integrated DNA Technologies. All guides were annealed to a final concentration of 200 μM. The ratio of crRNA to tracrRNA was 1:1. The RNP was complexed to achieve an enzyme-to-guide ratio of 1:2. A 1:1 volume ratio of 100 μM enzyme and 200 μM gRNA was used to achieve a final RNP concentration of 50 μM. The RNP was complexed for 30 minutes at room temperature. Next, the RNP was serially diluted 2-fold across 8 concentrations. The RNP was frozen at -80 °C until nucleofection.
[0331] Culturing of T cells T cells were cultured in Lonza X-Vivo 15 medium. The cells were thawed and cultured with Dynabeads Human T-Activator CD3 / CD28 for T Cell Expansion and Activation. On the second day after thawing, the cells were removed from the beads. The cells were cultured until the fourth day, on which they were spin-down for nucleofection. On the second day after nucleofection, the cell volume was split in half into a new plate and allowed to continue growing at room temperature.
[0332] T cell nucleofection Cells were counted using a BioRad T-20 cell counter. Cells were mixed 1:1 with trypan blue and counted. The total amount of cells required (sufficient for 500k cells per well) was aliquoted into separate tubes and then spun down at 1500 RPM for 5 minutes. Next, the cells were resuspended in Lonza P2 nucleofection solution. Next, the cells were placed into a Lonza 96-well nucleofection plate at 20 μL per well. Next, the cells and the RNP plate were transferred to a BioMek FX robot. Using a 96-well head, 2 μL of each RNP was transferred into the nucleofection plate and mixed. Next, the nucleofection plate was immediately transferred to a Lonza shuttle system where it was nucleofected with a DS-130 pulse code. Next, the cells were immediately transferred back to the BioMek FX where they were transferred and mixed into a pre-warmed 96-well untreated media plate. Next, the cell plate was placed at 37 °C for incubation.
[0333] gDNA extraction On day 4, the cells were spun down in the plate at 2000 RPM for 5 minutes. Next, the media was decanted. Next, the cell pellet was resuspended in Agencourt DNAdvance lysis solution. gDNA was extracted using the DNAdvance protocol for the BioMek FX.
[0334] GUIDE-seq GUIDE-seq was performed based on the protocol of Tsai et al. (Nat. Biotechnol. 33:187-197 (2015)) and adapted to T cells as follows. 10 μL of 4.4 μM RNP was combined with 4 μL of 100 μM dsODN and 6 μL of 1×H150 buffer to a total volume of 20 μL. The RNP was kept on ice until nucleofection. T cells were counted using a BioRad T-20 cell counter. The cells were mixed 1:1 with trypan blue and counted. The total amount of cells necessary (sufficient for 2 million cells per cuvette) was aliquoted into separate tubes and then spun down at 1500 RPM for 5 minutes. Next, the cells were resuspended in 80 μL of Lonza P2 solution. Next, the cells were pipetted into each cuvette and 20 μL of RNP / dsODN was added to each cuvette and the entire solution was gently mixed. Next, the cells were nucleofected using the CA-137 pulse code. Next, the cells were immediately pipetted into pre-warmed non-coated media plates. Next, the cell plates were placed at 37 °C for incubation. Genomic DNA was extracted and analyzed using only bidirectional reads with the protocol of Tsai et al. (Nat. Biotechnol. 33:187-197 (2015)).
[0335] Results Figures 18A and 18B show off-target cleavage for "SiteG" and "SiteH", respectively. As shown in Figures 18A and 18B, both SpartaCas and eCas decreased the total number of off-targets compared to wild-type Cas9. Furthermore, most of the remaining off-targets decreased the read count (shown in cyan).
[0336] Equivalents The present invention has been described in detail along with its detailed description, but it will be understood that the foregoing description is intended to illustrate the scope of the present invention as defined by the scope of the appended claims and not to limit the scope of the present invention. Other aspects, advantages, and modifications are within the scope of the following claims.
[0337] Array list An exemplary codon-optimized nucleic acid sequence (SEQ ID NO: 9) encoding the Cas9 molecule of Streptococcus pyogenes (S. pyogenes).
Chem.
Chem.
[0338] An exemplary codon-optimized nucleic acid sequence (SEQ ID NO: 10) encoding the Cas9 molecule of Staphylococcus aureus (S. aureus).
Chem.
[0339] An exemplary codon-optimized nucleic acid sequence (SEQ ID NO: 11) encoding the Cas9 molecule of Staphylococcus aureus (S. aureus).
Chem.
[0340] An exemplary codon-optimized nucleic acid sequence (SEQ ID NO: 12) encoding the Cas9 molecule of Staphylococcus aureus (S. aureus).
Chem.
[0341] An exemplary Streptococcus pyogenes (S. pyogenes) Cas9 amino acid sequence (SEQ ID NO: 13).
Chem.
[0342] An exemplary Neisseria meningitidis Cas9 amino acid sequence (SEQ ID NO: 14).
Chem.
Claims
1. An in vitro or ex vivo method of modifying a cell, comprising: administering to the cell: (i) a polypeptide having nuclease activity and comprising an amino acid sequence having at least 90% identity to SEQ ID NO: 13, wherein the polypeptide has the following amino acid substitutions relative to SEQ ID NO: 13: D23A and Y128V, and the polypeptide has reduced off-target cleavage activity compared to the polypeptide of SEQ ID NO: 13; and (ii) a guide nucleic acid in a step. Method.
2. The method of claim 1, wherein the guide nucleic acid is guide RNA. Method.
3. The method of claim 1, wherein the cell comprises a target sequence and an off-target sequence, and the off-target editing rate of the off-target sequence by the polypeptide is lower than the off-target editing rate of the off-target sequence observed when using a control polypeptide comprising SEQ ID NO:
13. Method.
4. The method of claim 3, wherein the off-target editing rate by the polypeptide is at least 5% lower than the off-target editing rate of the control polypeptide. Method.
5. The method of claim 3, wherein the off-target editing rate is measured by evaluating the level of indels in the off-target sequence. Method.
6. The method of claim 1, wherein the cell is modified using a ribonucleoprotein (RNP) complex comprising the polypeptide and the guide nucleic acid. Method.
7. The method of claim 1, wherein the polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 13, Method.
8. The method of claim 1, wherein the polypeptide comprises an amino acid sequence having at least 99% identity to SEQ ID NO: 13, Method.
9. The method of claim 1, wherein the polypeptide further has an amino acid substitution at one or more of the following positions of SEQ ID NO: 13: T67 and D1251, Method.
10. The method of claim 1, wherein the polypeptide further has an amino acid substitution at the following positions of SEQ ID NO: 13: T67 and D1251, Method.
11. The method of claim 1, wherein the polypeptide further has one or more of the following amino acid substitutions relative to SEQ ID NO: 13: T67L and D1251G, Method.
12. The method of claim 1, wherein the polypeptide further has the following amino acid substitutions relative to SEQ ID NO: 13: T67L and D1251G, Method.
13. An in vitro or ex vivo method of editing a population of double-stranded DNA (dsDNA) molecules, wherein the method comprises the step of subjecting the dsDNA molecules to (i) A polypeptide comprising an amino acid sequence having at least 90% identity to SEQ ID NO: 13, wherein said polypeptide has the following amino acid substitutions relative to SEQ ID NO: 13: D23A and Y128V, and said polypeptide has reduced off-target cleavage activity compared to the polypeptide of SEQ ID NO: 13; and, (ii) A guide nucleic acid comprising a region complementary to the target sequence of said dsDNA molecule comprising contacting, whereby a plurality of said dsDNA molecules are edited. Method.
14. The method of claim 13, wherein the off-target editing rate of non-target sequences by said polypeptide is lower than the off-target editing rate of said non-target sequences observed when using a polypeptide comprising the amino acid sequence of SEQ ID NO:
13. Method.
15. The method of claim 14, wherein the off-target editing rate by said polypeptide is at least 5% lower than the off-target editing rate of the polypeptide comprising the amino acid sequence of SEQ ID NO:
13. Method.
16. The method of claim 14, wherein the off-target editing rate is measured by evaluating the level of indels in said non-target sequence. Method.
17. The method of claim 13, wherein the step of contacting said dsDNA molecule with an RNP complex comprising said polypeptide and a guide nucleic acid is included. Method.
18. The method of claim 17, wherein the step of contacting said dsDNA molecule with said RNP complex is included, and said RNP complex is at a dose of 1×10 -4 μM to 1 μM RNP. Method.
19. The method according to claim 13, wherein the polypeptide further has an amino acid substitution at one or more of the following positions of SEQ ID NO: 13: T67 and D1251, Method.
20. The method according to claim 13, wherein the polypeptide further has an amino acid substitution at the following positions of SEQ ID NO: 13: T67 and D1251, Method.
21. The method according to claim 13, wherein the polypeptide further has one or more of the following amino acid substitutions with respect to SEQ ID NO: 13: T67L and D1251G, Method.
22. The method according to claim 13, wherein the polypeptide further has the following amino acid substitutions with respect to SEQ ID NO: 13: T67L and D1251G, Method.
Citation Information
Patent Citations
Engineered crispr-cas9 nucleases
WO2017040348A1