High fidelity spcas9 nuclease for genome modification
Patent Information
- Application Number
- CN202180020084.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-11
- Filing Date
- 2021-03-11
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2041-03-11
AI Technical Summary
然而,这些变体中的大多数通过以质粒形式筛选得到鉴定,并且它们借助于核糖核蛋白(RNP)递送转换为用于基因组修饰的重组蛋白质经常导致低活性
[0013] Other purposes and features will be apparent in part and indicated in part below.
Smart Images

Figure CN115244177B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 988,279, filed March 11, 2020, the entire contents of which are incorporated herein by reference.
[0003] sequence list
[0004] This application contains a sequence list, which has been electronically submitted in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy created on March 11, 2021, is named P20-035_WO-PCT_SL.txt and is 49,120 bytes in size. Technical Field
[0005] This disclosure relates to modified Cas9 protein variants and systems, nucleic acids encoding said protein variants and systems, and methods for preparing and using said protein variants and systems for genome modification. Background Technology
[0006] CRISPR Cas9 (SpCas9) has been widely used as a genome-editing endonuclease in many cell types and organisms in Streptococcus pyogenes. However, wild-type nucleases tend to induce mutations at unintended genomic sites carrying sequences similar to the target site. Several SpCas9 variants with improved specificity have been developed to mitigate this drawback. These include eSpCas9 1.0 (K810A, K1003A, R1060A), eSpCas9 1.1 (K848A, K1003A, R1060A), SpCas9-HF1 (N497A, R661A, Q695A, Q926A), HypaCas9 (N692A, M694A, Q695A, H698A), EvoCas9 (M495V, Y515N, K526E, R661L), SniperCas9 (F539S, M763I, K890N), HiFi Cas9 V3 (R691A), Opti-SpCas9 (R661A and K1003H) and OptiHF-SpCas9 (Q695A, K848A, E293M, T924V and Q926A). (Slaymaker et al., Science 351, 84-88; Kleinstiver et al., Nature 523, 490-495; Chen et al., Nature 550, 407-410; Casini et al., Nature Biotechnology 36, 265-271; Lee et al., Nature Communications 9, 3048; Vakulskas et al., Nature Medicine 24, 1216-1224; Choi et al., Nature Methods 16, 722-730). However, most of these variants were identified by screening in plasmid form, and their conversion into recombinant proteins for genome modification via ribonucleoprotein (RNP) delivery often resulted in low activity.
[0007] Due to the greatly increased demand for SpCas9 recombinant proteins in genome modification, there is a need for nucleases in the form of recombinant proteins that can function across different genomic sites with improved specificity and sustained activity. Summary of the Invention
[0008] Various aspects of the present invention provide modified Cas9 protein variants and systems comprising them.
[0009] Therefore, in summary, this disclosure relates to modified Streptococcus pyogenes Cas9 (SpCas9) protein variants comprising modifications at one, two, or more of the amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060, wherein lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine (R) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q). For example, in one exemplary embodiment, the modified SpCas9 protein variant comprises the K855L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 1003, and 1060 (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another exemplary embodiment, the modified SpCas9 protein variant comprises the R661L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 652, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In a particular embodiment, the mutation is selected from the group consisting of: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; and K652Q-R661L-K855Q (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0010] Another aspect of this disclosure relates to a modified Cas9 system comprising a modified Cas9 protein variant disclosed herein and at least one modified guide RNA, wherein each modified guide RNA is designed to be complexed with the modified Cas9 protein variant.
[0011] Another aspect of this disclosure relates to nucleic acids encoding modified Cas9 protein variants and systems comprising them. Vectors containing nucleic acids are also provided.
[0012] Another aspect of this disclosure relates to methods for preparing and using the modified Cas9 protein variants and systems described herein.
[0013] Other purposes and features will be apparent in part and indicated in part below. Attached Figure Description
[0014] Figure 1AThe figure illustrates the intermediate target activity of five different substitutions at the K855 residue of the HEKSite4 target site in human U-2 OS cells. K855E and K855A resulted in reduced intermediate target activity (Example 1). This figure discloses SEQ ID NO: 73.
[0015] Figure 1B The off-target activities of five different substitutions at the K855 residue of HEKSite4 off-target site are shown in human U-2 OS cells (Example 1). This figure discloses SEQ ID NO: 74.
[0016] Figure 2 This figure illustrates the mid-target activity at the HEKSite4 target site with different substitutions of residues R661, N692, or Q695 in human U-2 OS cells (Example 2). SEQ ID NO: 73 is disclosed in this figure.
[0017] Figure 3A The figure illustrates the mid-target activity of triple and quadruple mutant proteins at the FANCF02 target site in human K562 cells (Example 3). SEQ ID NO: 75 is disclosed in this figure.
[0018] Figure 3B The off-target activity of triple and quadruple mutant proteins at a single mismatch off-target site of FANCF02 in human K562 cells is shown (Example 3). This figure discloses SEQ ID NO: 76.
[0019] Figure 3C The figure illustrates the mid-target activity of triple and quadruple mutant proteins at the HBB03 target site in human K562 cells (Example 3). SEQ ID NO: 77 is disclosed in this figure.
[0020] Figure 3D The off-target activity of triple and quadruple mutant proteins at a single mismatch off-target site of HBB03 in human K562 cells is shown (Example 3). This figure discloses SEQ ID NO: 78.
[0021] Figure 4 The figure illustrates the mid-target activity of a group of selected mutant proteins at five different genomic loci in K562 cells (Example 4). SEQ ID NOs 79-83 are disclosed in the order of appearance. Detailed Implementation
[0022] Due to the significantly increased demand for recombinant SpCas9 proteins in genome modification, there is a need for nucleases in recombinant protein form that can function across different genomic sites with improved specificity and sustained activity. Using a recombinant protein-based screening approach, at least two distinct groups of SpCas9 variants with varying levels of specificity and activity have been identified. One group exhibits extremely high specificity relative to other SpCas9 variants, but with highly variable activity across different genomic sites. The other group offers a balanced combination of specificity and activity, outperforming the well-established eSpCas9 1.1 in activity and the recently developed HiFi Cas9 V3 in specificity. This group of nucleases holds great potential for widespread application in genome modification in eukaryotic cells.
[0023] Previous attempts to develop high-fidelity SpCas9 variants have largely relied on some form of plasmid expression-based selection scheme. These variants often exhibit low activity when used as recombinant proteins. Without being bound by any particular theory, it is hypothesized that plasmid overexpression in mammalian cells can mask the diminished activity of these variants caused by mutations, thereby increasing specificity. To avoid this confounding effect of plasmid overexpression, recombinant protein-based screening methods have been employed to improve the nuclease. Furthermore, in contrast to previous attempts that invariably used alanine substitutions to mutate key residues, optimal amino acid substitutions were used to maintain mid-target activity while improving specificity. These differences methodologically distinguish the proteins disclosed herein from those derived from previous SpCas9 protein engineering efforts.
[0024] Mutations and their combinations contain key residues that are presumed to be involved in different mechanisms of Cas9 DNA substrate binding stability, whereas previous attempts limited mutation combinations to key residues that were presumed to be involved in a single mechanism. For example, eSpCas9 was developed by mutating conserved positively charged amino acid residues that interact with the negatively charged phosphate backbone of the non-target strand, based on the assumption that positively charged residues stabilize the non-target strand during strand separation; and subsequently stabilizing the formation of the RNA-target DNA heteroduplex (Slaymaker et al., Science 351, 84-88). In contrast, SpCas9-HF1 was developed by reducing hydrogen bonding or charge interactions with the phosphate backbone of the target strand (Kleinstiver et al., Nature 523, 490-495). On the other hand, HypaCas9 is derived by mutating a cluster of conserved residues (N692, M694, Q695, and H698) in the REC3 domain to alanine, which is presumed to sense RNA-DNA interactions and transmit such signals to trigger a conformational change in the HNH nuclease domain (Chen et al., Nature 550, 407-410).
[0025] By employing a unique screening method based on recombinant proteins and extending rational designs to different combinations of mechanisms as disclosed herein, this disclosure has identified at least three distinct groups of SpCas9 variants with varying levels of specificity and activity.
[0026] (I) Modified Cas9 protein
[0027] One aspect of this disclosure relates to modified Cas proteins. Modified Cas proteins comprise at least one, at least two, or at least three amino acid substitutions, insertions, or deletions relative to their wild-type counterparts; that is, compared to wild-type Cas proteins, modified Cas9 proteins comprise modifications or mutations to the amino acid sequence. Among various Cas proteins, for example, Cas9 protein is a single-effect protein in type II CRISPR systems present in various bacteria.
[0028] In one embodiment, the modified Cas9 protein disclosed herein is derived from Streptococcus spp. ( Streptococcus ) species. In another embodiment, for example, the modified Cas9 protein variant is derived from Streptococcus pyogenes (SpCas9). Thus, in some embodiments, the modified Cas9 protein described herein is a homolog of SpCas9.
[0029] Wild-type Cas9 protein contains two nuclease domains, namely the RuvC and HNH domains, each of which cleaves one strand of the double-stranded sequence. Cas9 protein also contains a REC domain that interacts with guide RNA (e.g., REC1, REC2) or RNA / DNA heteroduplexes (e.g., REC3), and a domain that interacts with the prespacer adjacent motif (PAM) (i.e., the PAM interaction domain).
[0030] As noted herein, the Cas9 proteins of this disclosure have been modified to include one or more modifications (i.e., substitution of at least one amino acid, deletion of at least one amino acid, or insertion of at least one amino acid) to give the Cas9 proteins altered activity, specificity, and / or stability. These modified Cas9 proteins are not naturally occurring.
[0031] Generally, known and / or commercially available Cas9 mutants are concentrated on point mutations in specific regions of the protein, independent of other regions and combinations of mutations in different regions. It has been advantageously found that, relative to known Cas9 mutants, combinations of mutations in different regions of the Cas9 protein can lead to improved specificity, activity (e.g., on-target or off-target activity), and / or other beneficial properties.
[0032] For example, the Cas9 protein disclosed herein has at least one mutation in a structural region of a protein including non-target DNA strand contact residues, and / or at least one mutation in a structural region of a protein including target DNA / guide RNA heteroduplex contact residues, and / or at least one mutation in a structural region of a protein including α-spiral leaf residues. For the purposes of this disclosure, non-target DNA strand contact residues include, for example, amino acids R780, K810, K848, K855, K1003, and R1060; target DNA / guide RNA heteroduplex contact residues include, for example, amino acids R661 and R691; and α-spiral leaf residues include, for example, amino acids K526, K562, and K652 (refer to the numbering system for Streptococcus pyogenes Cas9—SpCas9). Therefore, in various embodiments, the Cas9 protein disclosed herein has at least one mutation in a structural region of a protein including non-target DNA strand contact residues, and at least one mutation in a structural region of a protein including target DNA / guide RNA heteroduplex contact residues. In other embodiments, the Cas9 protein disclosed herein has at least one mutation in a structural region of a protein including non-target DNA strand contact residues and at least one mutation in a structural region of a protein including an α-helical leaf. In still other embodiments, the Cas9 protein disclosed herein has at least one mutation in a structural region of a protein including target DNA / guide RNA heteroduplex contact residues and at least one mutation in a structural region of a protein including an α-helical leaf. In still other embodiments, the Cas9 protein disclosed herein has at least one mutation in a structural region of a protein including non-target DNA strand contact residues, at least one mutation in a structural region of a protein including target DNA / guide RNA heteroduplex contact residues, and at least one mutation in a structural region of a protein including an α-helical leaf.
[0033] The Cas9 protein variant disclosed herein has a modified amino acid sequence, which is identified by referring to the amino acid number at the corresponding position in the unmodified mature (wild-type) Streptococcus pyogenes Cas9 (SEQ ID NO: 1). The Cas9 protein variant disclosed herein preferably has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity with SEQ ID NO: 1.
[0034]
[0035] For easy reference, the symbols and their single-letter codes for the 20 essential amino acids are shown in Table A below.
[0036] Table A: Amino Acids
[0037] amino acids Single-letter codes alanine A Arginine R Asparagine N Aspartic acid D Cysteine C glutamine Q glutamic acid E glycine G Histidine H Isoleucine I Leucine L Lysine K Methionine M Phenylalanine F proline P Serine S threonine T Tryptophan W Tyrosine Y Valine V
[0038] It should be understood that the amino acid modifications described herein utilize a nomenclature that begins with a letter (single-letter code) referring to the affected amino acid and ends with a letter (single-letter code) specifying the change, with the amino acid residue position between the two letters. For example, a hypothetical protein might have an alanine residue at hypothetical amino acid position 100 and is designated A100. As a further example, a modification at hypothetical amino acid position 100 from alanine to valine is designated A100V. Modifications selected from two or more options can be designated using " / ", for example, a modification at hypothetical amino acid position 100 from alanine to valine or serine is designated A100V / S.
[0039] In one embodiment, the modified Cas9 protein variant includes mutations at one or more of the following amino acid positions: 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, the modified Cas9 protein variant includes mutations at two or more of the following amino acid positions: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at two or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at three or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at four or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at five or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at six or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at seven or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at eight or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060.As another example, a modified Cas9 protein may include mutations at nine or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at ten or more of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. As another example, a modified Cas9 protein may include mutations at each of the following locations: K526, K562, K652, R661, R691, R780, K810, K848, K855, K1003, and R1060. In some of these different embodiments, for example, lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine (R) at one or more of the aforementioned amino acid positions (R) is changed to leucine (L) or glutamine (Q).
[0040] In one embodiment, for example, the modified SpCas9 variant includes the K526L / Q mutation, and at least one other mutation at amino acid positions 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the K562L / Q mutation, and at least one other mutation at amino acid positions 526, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the K652L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the R661L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 691, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the R691L / Q mutation, and at least one other mutation at amino acid positions 562, 661, 780, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the R780L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 810, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the K810L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 848, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the K848L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 810, 855, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9).In another embodiment, for example, the modified SpCas9 variant includes the K855L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 810, 848, 1003, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the K1003L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 810, 848, 855, and 1060 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another embodiment, for example, the modified SpCas9 variant includes the R1060L / Q mutation, and at least one other mutation at amino acid positions 526, 562, 661, 691, 780, 810, 848, 855, and 1003 (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0041] Therefore, in some embodiments, for example, the modified SpCas9 protein variants include two mutations at two different amino acid positions selected from K526L / Q, K562L / Q, K652L / Q, K810L / Q, K848L / Q, K855L / Q, R661L / Q, R691L / Q, R780L / Q, K1003L / Q, and R1060L / Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9). In other embodiments, for example, the modified SpCas9 protein variants include three mutations at three different amino acid positions selected from K526L / Q, K562L / Q, K652L / Q, K810L / Q, K848L / Q, K855L / Q, R661L / Q, R691L / Q, R780L / Q, K1003L / Q, and R1060L / Q (referencing the numbering system for Streptococcus pyogenes Cas9-SpCas9). It should be understood that other implementation schemes are provided in which four, five, six, seven, eight, nine, ten, or eleven mutations may be present at amino acid positions selected from K526L / Q, K562L / Q, K652L / Q, K810L / Q, K848L / Q, K855L / Q, R661L / Q, R691L / Q, R780L / Q, K1003L / Q, and R1060L / Q (refer to the numbering system of Streptococcus pyogenes Cas9-SpCas9).
[0042] It should be understood that, in the foregoing paragraphs, mutants known in the art (if any) within the scope of a particular implementation or instance are excluded by way of condition.
[0043] In another specific embodiment, the modified SpCas9 protein variant includes at least one of the following mutations: K526L, K526Q, K562L, K562Q, K652L, K652Q, K810L, K810Q, K848L, K848Q, K855L, K855Q, R661L, R661Q, R691L, R691Q, R780L, R780Q, K1003L, K1003Q, R1060L, and R1060Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0044] In another specific embodiment, the modified SpCas9 protein variant includes at least two of the following mutations: K526L, K526Q, K562L, K562Q, K652L, K652Q, K810L, K810Q, K848L, K848Q, K855L, K855Q, R661L, R661Q, R691L, R691Q, R780L, R780Q, K1003L, K1003Q, R1060L, and R1060Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0045] In yet another specific implementation, the modified SpCas9 protein variant includes at least three of the following mutations: K526L, K526Q, K562L, K562Q, K652L, K652Q, K810L, K810Q, K848L, K848Q, K855L, K855Q, R661L, R661Q, R691L, R691Q, R780L, R780Q, K1003L, K1003Q, R1060L, and R1060Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0046] In yet another implementation, the modified SpCas9 protein variant includes at least four, five, six, seven, eight, nine, ten, or eleven of the following mutations: K526L, K526Q, K562L, K562Q, K652L, K652Q, K810L, K810Q, K848L, K848Q, K855L, K855Q, R661L, R661Q, R691L, R691Q, R780L, R780Q, K1003L, K1003Q, R1060L, and R1060Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0047] In one particular embodiment, the modified SpCas9 protein is selected from one of the following variant groups: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; K652Q-R661L-K855Q; R661L-K855Q-K1003Q; and R661L-K855Q-R1060Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9). In another particular embodiment, the modified SpCas9 protein is selected from one of the following variant groups: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; and K652Q-R661L-K855Q. Therefore, for example, the modified SpCas9 protein variant could be K562L-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be K562Q-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be K652L-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be K652Q-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K855Q-K1003Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K855Q-R1060Q. Members of this group of variants have a relatively balanced specificity and activity, superior in activity to the well-established eSpCas9 1.1 and superior in specificity to the recently developed HiFi Cas9V3.
[0048] In another specific embodiment, the modified SpCas9 protein is selected from one of the following variant groups: K526L-R661L-K855Q; R661L-R691L-K855Q; R661L-R780L-K855Q; R661L-R780Q-K855Q; R661L-K810L-K855Q; and R661L-K848L-K855Q (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9). Thus, for example, the modified SpCas9 protein variant could be K526L-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-R691L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-R780L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-R780Q-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K810L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K848L-K855Q. Members of this group of variants have extremely high levels of specificity but exhibit highly diverse activities across target sites.
[0049] In another specific embodiment, the modified SpCas9 protein is selected from one of the following variant groups: K526Q-R661L-K855Q; R661L-K810Q-K855Q; R661L-K855Q-K1003L; and R661L-K855Q-R1060L (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9). Thus, for example, the modified SpCas9 protein variant could be K526Q-R661L-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K810Q-K855Q. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K855Q-K1003L. Alternatively, for example, the modified SpCas9 protein variant could be R661L-K855Q-R1060L. Members of this group of variants are similar to eSpCas9 1.1 in both specificity and activity levels; however, they differ from eSpCas9 1.1 in terms of mutation profile.
[0050] In addition to the mutations discussed above, the Cas9 protein can also be modified by one or more mutations and / or deletions to inactivate one or both of its nuclease domains. Inactivation of one nuclease domain generates a Cas9 protein that cleaves one strand of a double-stranded sequence (i.e., the Cas9 cleavage enzyme). The RuvC domain can be inactivated by mutations such as D10A, D8A, E762A, and / or D986A, and the HNH domain can be inactivated by mutations such as H840A, H559A, N854A, N856A, and / or N863A (refer to the numbering system for Streptococcus pyogenes Cas9—SpCas9). Inactivation of both nuclease domains generates a Cas9 protein without cleavage activity (i.e., catalytically inactivated or dead Cas9).
[0051] In addition to the various mutations discussed above, the Cas9 protein can also be modified by one or more amino acid substitutions, deletions, and / or insertions to achieve improved target specificity, improved fidelity, altered PAM specificity, reduced off-target effects, and / or increased stability. Non-limiting examples of one or more mutations that improve target specificity, improve fidelity, and / or reduce off-target effects include N497A, R661A, Q695A, K810A, K848A, K855A, Q926A, K1003A, R1060A, and / or D1135E (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9).
[0052] In addition to the modifications discussed above, the Cas9 protein can also be engineered to include at least one heterodomain, i.e., Cas9 is fused to one or more heterodomains. In cases where two or more heterodomains are fused to Cas9, the two or more heterodomains can be identical or they can be different. One or more heterodomains can be fused to the N-terminus, C-terminus, internal location, or a combination thereof. Fusion can be direct via chemical bonds, or the bonding can be indirect via one or more linkers. In various embodiments, the heterodomain is selected from nuclear localization signals, cell-penetrating domains, detection-enhancing markers or reporter domains (fluorescent or enzymatic reporter proteins), chromatin modification domains, epigenetic modification domains (e.g., cytidine deaminase domains, histone acetyltransferase domains, etc.), transcriptional regulatory domains, DNA or RNA deaminase domains, uracil-DNA glycosylase domains, reverse transcriptase domains, recombinase domains, RNA aptamer-binding domains, or non-Cas9 nuclease domains.
[0053] (a) Nuclear localization signal
[0054] In some implementations, one or more heterogeneous structural domains may be nuclear localization signals (NLS). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 2), PKKKRRV (SEQ ID NO: 3), KRPAATKKAGQAKKKK (SEQ ID NO: 4), YGRKKRRQRRR (SEQ ID NO: 5), RKKRRQRRR (SEQ ID NO: 6), PAAKRVKLD (SEQ ID NO: 7), RQRRNELKRSP (SEQ ID NO: 8), VSRKRPRP (SEQ ID NO: 9), PPKKARED (SEQ ID NO: 10), PQPKKKPL (SEQ ID NO: 11), SALIKKKKKMAP (SEQ ID NO: 12), PKQKKRK (SEQ ID NO: 13), RKLKKKIKKL (SEQ ID NO: 14), REKKKFLKRR (SEQ ID NO: 15), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 16), and KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 17). 16), RKCLQAGMNLEARKTKK (SEQ ID NO: 17), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 18) and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 19).
[0055] (b) Cell penetration domain
[0056] In other implementations, one or more heterologous domains may be cell-penetrating domains. Examples of suitable cell-penetrating domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 20), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 21), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 22), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 23), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 24), YARAAARQARA (SEQ ID NO: 25), THRLPRRRRRR (SEQ ID NO: 26), GGRRARRRRRR (SEQ ID NO: 27), RRQRRTSKLMKR (SEQ ID NO: 28), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 29), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 20). 30) and RQIKIWFQNRRMKWKK (SEQ ID NO: 31).
[0057] (c) Marker domain
[0058] In an alternative implementation, one or more heterologous domains may be marker domains. Marker domains include fluorescent proteins and purification tags or epitope tags. Suitable fluorescent proteins include, but are not limited to, green fluorescent proteins (e.g., GFP, eGFP, GFP-2, tagGFP, turboGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., BFP, EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), and red fluorescent proteins (e.g., mKa...). The labeling domain may contain tandem repeats of one or more fluorescent proteins (e.g., mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, MonomericKusabira-Orange, mTangerine, tdTomato), or combinations thereof. The labeling domain may contain tandem repeats of one or more fluorescent proteins (e.g., Suntag). Non-limiting examples of suitable purification tags or epitope tags include 6xHis (SEQ ID NO: 32), FLAG, etc. ® HA, GST, Myc, SAM, etc. Non-limiting examples of heterologous fusions that promote the detection or enrichment of CRISPR complexes include streptavidin (Kipriyanov et al., Human Antibodies, 1995, 6(3):93-101), avidin (Airenne et al., Biomolecular Engineering, 1999, 16(1-4):87-92), monomeric forms of avidin (Laitinen et al., Journal of Biological Chemistry, 2003, 278(6):4010-4014), and peptide tags that promote biotinylation during recombinant production (Cull et al., Methods in Enzymology, 2000, 326:430-440).
[0059] (d) Chromatin regulatory motifs
[0060] In other embodiments, one or more heterologous domains may be chromatin regulatory motifs (CMMs). Non-limiting examples of CMMs include nucleosome-interacting peptides derived from high mobility family (HMG) proteins (e.g., HMGB1, HMGB2, HMGB3, HMGN1, HMGN2, HMGN3a, HMGN3b, HMGN4, and HMGN5 proteins), central globular domains of histone H1 variants (e.g., histones H1.0, H1.1, H1.2, H1.3, H1.4, H1.5, H1.6, H1.7, H1.8, H1.9, and H1.10), or DNA-binding domains of chromatin remodeling complexes (e.g., SWI / SNF (conversion / sucrose non-fermentation), ISWI (simulated conversion / sucrose non-fermentation)). I mitation SWI The CMM comprises tch), CHD (chromatin domain-helicase-DNA binding), Mi-2 / NuRD (nucleosome remodeling and deacetylation enzyme), INO80, SWR1, and RSC complexes. In other embodiments, the CMM may also be derived from topoisomerases, helicases, or viral proteins. The source of the CMM can and will vary. The CMM can be derived from humans, animals (i.e., vertebrates and invertebrates), plants, algae, or yeast. Non-limiting examples of specific CMMs are listed in Table B below. Those skilled in the art can readily identify homologs and / or related fusion motifs in other species.
[0061] Table B: Chromatin Regulatory Motifs
[0062] protein Login ID Fusion motif HMGN1 P05114 full length HMGN2 P05204 full length HMGN3a Q15651 full length HMGN3b Q15651-2 full length HMGN4 O00479 full length HMGN5 P82970 Nucleosome binding motif HMGB1 P09429 Box A Human histone H1.0 P07305 globular motif Human histone H1.2 P16403 globular motif Human CHD1 O14646 DNA binding motif Yeast CHD1 P32657 DNA binding motif Yeast ISWI P38144 DNA binding motif TOP1 P11387 DNA binding motif Human herpesvirus 8 (LANA) J9QSF0 Nucleosome binding motif Human CMV IE1 P13202 chromatin tethering motif Mycobacterium leprae DNA helicase P40832 HhH binding motif
[0063] (e) Epigenetic modification domains
[0064] In other embodiments, one or more heterologous domains may be epigenetic modification domains. Non-limiting examples of suitable epigenetic modification domains include those having the following functions: DNA deamination (e.g., cytidine deaminase, adenosine deaminase, guanine deaminase), DNA methyltransferase activity (e.g., cytosine methyltransferase), DNA demethylase activity, DNA amination, DNA oxidation activity, DNA helicase activity, histone acetyltransferase (HAT) activity (e.g., a HAT domain derived from E1A-binding protein p300), histone deacetylase activity, and histone methylation activity. Histone transmethylase activity, histone demethylase activity, histone kinase activity, histone phosphatase activity, histone ubiquitin ligase activity, histone deubiquitination activity, histone adenylation activity, histone deadenylation activity, histone SUMOylation activity, histone deSUMOylation activity, histone ribosylation activity, histone deribosylation activity, histone myristoylation activity, histone demyristoylation activity, histone citrullination activity, histone alkylation activity, histone dealkylation activity, or histone oxidation activity. In specific embodiments, the epigenetic modification domain may include cytidine deaminase activity, adenosine deaminase activity, histone acetyltransferase activity, or DNA methyltransferase activity.
[0065] (f) Transcriptional regulatory domains
[0066] In other embodiments, one or more heterologous domains may be transcriptional regulatory domains (i.e., transcriptional activation domains or transcriptional repression domains). Suitable transcriptional activation domains include, but are not limited to, the herpes simplex virus VP16 domain, VP64 (i.e., four tandem copies of VP16), VP160 (i.e., ten tandem copies of VP16), NFκB p65 activation domain (p65), EBV R transactivator (Rta) domain, VPR (i.e., VP64+p65+Rta), p300-dependent transcriptional activation domain, p53 activation domains 1 and 2, heat shock factor 1 (HSF1) activation domain, Smad4 activation domain (SAD), cAMP response element binding protein (CREB) activation domain, E2A activation domain, activated T cell nuclear factor (NFAT) activation domain, or combinations thereof. Non-restricted examples of suitable transcriptional repression domains include the Kruppel-associated box (KRAB) repression domain, the Mxi repression domain, the inducible early cAMP repression (ICER) domain, the YYl glycine-rich repression domain, Sp1-like repressors, E(spl) repressors, IκB repressors, Sin3 repressors, methyl-CpG-binding protein 2 (MeCP2) repressors, or combinations thereof. Transcriptional activation or repression domains can be genetically fused to the Cas9 protein or bound via non-covalent protein-protein, protein-RNA, or protein-DNA interactions.
[0067] (g) RNA aptamer binding domain
[0068] In a further implementation, one or more heterologous domains may be RNA aptamer-binding domains (Konermann et al., Nature, 2015, 517(7536):583-588; Zalatan et al., Cell, 2015, 160(1-2):339-50). Examples of suitable RNA aptamer protein domains include MS2 capsid protein (MCP), PP7 bacterial phage capsid protein (PCP), μ bacterial phage Com protein, λ bacterial phage N22 protein, stem-loop binding protein (SLBP), Fragile X-associated intellectual disability syndrome-associated protein 1 (FXR1), and proteins derived from bacterial phages such as AP205, BZ13, f1, f2, fd, fr, ID2, JP34 / GA, JP501, JP34, JP500, KU1, M11, M12, MX1, NL95, PP7, ϕCb5, ϕCb8r, ϕCb12r, ϕCb23r, Qβ, R17, SP-β, TW18, TW19, and VK, fragments thereof, or derivatives thereof.
[0069] (h) Non-Cas9 nuclease domains
[0070] In other embodiments, one or more heterologous domains may be non-Cas9 nuclease domains. Suitable nuclease domains can be obtained from any endonuclease or exonuclease. Non-limiting examples of nuclease domains derived from them include, but are not limited to, restriction endonucleases and homing endonucleases. In some embodiments, the nuclease domain may be derived from type II-S restriction endonucleases. Type II-S endonucleases cleave DNA at sites typically a few base pairs away from the recognition / binding site and therefore have separable binding and cleaving domains. These enzymes are generally monomers that transiently associate to form dimers to cleave each DNA strand at staggered locations. Non-limiting examples of suitable type II-S endonucleases include BfiI, BpmI, BsaI, BsgI, BsmBI, BsmI, BspMI, FokI, MboII, and SapI. In some embodiments, the nuclease domain may be a FokI nuclease domain or a derivative thereof. The type II-S nuclease domain can be modified to promote the dimerization of two different nuclease domains. For example, the cleavage domain of FokI can be modified by mutating certain amino acid residues. As a non-limiting example, amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of the FokI nuclease domain are targets for modification. In a specific embodiment, the FokI nuclease domain may include a first FokI hemidomain containing mutations of Q486E, I499L, and / or N496D, and a second FokI hemidomain containing mutations of E490K, I538K, and / or H537R.
[0071] (i) Nucleotide modifying enzymes
[0072] The modified Cas9 variants described in this article may also contain nucleobase-modifying enzymes or their catalytic domains.
[0073] Various nucleobase-modifying enzymes are suitable for the systems disclosed herein. Nucleobase-modifying enzymes can be DNA base editors. In some embodiments, the DNA base editor can be a cytidine deaminase that converts cytidine to uridine, which is then read as thymine by a polymerase. Non-limiting examples of cytidine deaminases include cytidine deaminase 1 (CDA1), cytidine deaminase 2 (CDA2), activation-induced cytidine deaminase (AICDA), apolipoprotein B mRNA editing complex (APOBEC) family cytidine deaminases (e.g., APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4), APOBEC1 complement / APOBEC1 stimulatory factor (ACF1 / ASF) cytidine deaminases, RNA-acting cytosine deaminases (CDAR), and bacterial long isotype cytidine deaminases (CDD). L The DNA base editor can be an adenosine deaminase that converts adenosine to inosine, which is then read as guanosine by a polymerase. Non-limiting examples of adenosine deaminases include tRNA adenine deaminase, adenosine deaminase, adenosine deaminase acting on RNA (ADAR), and adenosine deaminase acting on tRNA (ADAT).
[0074] Nucleobase modifying enzymes (base editors) can be wild-type or fragments thereof, modified forms (e.g., non-essential domains may be missing), or altered forms. Nucleobase modifying enzymes (base editors) can have eukaryotic, bacterial, or archaea origins.
[0075] In some implementations, the nucleobase modifying enzyme (base editor) can be a cytidine deaminase or its catalytic domain. Cytidine deaminase can have human, mouse, lamprey, abalone, or *Escherichia coli* (…) E. coli Origin. In embodiments where the nucleobase modifying enzyme is a cytidine deaminase, the RNA-guided nucleobase modification system may further include at least one uracil glycosylation inhibitor (UGI) domain. UGI inhibits the removal of uracil from DNA, a result of cytosine deamination. Suitable UGI domains are known in the art.
[0076] In some implementations, systems employing cytidine deaminase and UGI may have negative effects if these components are overexpressed. To prevent overexpression, degradation tags can be added. Degradation tags indicate that the protein has been degraded through a protein recycling system. These degradation tags result in different protein half-lives. Examples of non-restrictive degradation tags are LVA, AAV, ASV, and LAA.
[0077] (j) Reverse transcriptase
[0078] In some implementations, the domain fused with the modified SpCas9 variant described herein is a reverse transcriptase. Examples of reverse transcriptases include avian myeloblastoma virus (AMV) reverse transcriptase and Moloney murine leukemia virus (MMLV) reverse transcriptase.
[0079] (k) Recombinase / Integrase
[0080] In some embodiments, the domain fused with the modified SpCas9 variant described herein is a recombinase or integrase. Non-limiting examples of suitable recombinases include Cre recombinase, FLP recombinase, Gin recombinase, Bacteroides intN2 tyrosine integrase (encoded by the NBU2 gene), Streptomyces phage phiC31 (φC31) recombinase, Escherichia coli phage P4 recombinase, Escherichia coli phage λ integrase, Listeria A118 phage recombinase, lentiviral or HIV integrase, and Actinomyces phage R4 Sre recombinase. The recombinase / integrase mediates recombination between two sequence-specific recognition (or attachment) sites (e.g., attP and attB sites or two Cre / loxP sites), or it can randomly insert DNA, as with HIV integrase.
[0081] (l) Connector
[0082] One or more heterodomains can be directly attached to the Cas9 protein via one or more chemical bonds (e.g., covalent bonds), or one or more heterodomains can be indirectly attached to the Cas9 protein via one or more linkers.
[0083] A linker is a chemical group that connects one or more other chemical groups via at least one covalent bond. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). A linker may include one or more spacer groups, including but not limited to alkylene, alkenylene, ynylene, alkyl, alkenyl, ynyl, alkoxy, aryl, heteroaryl, aralkyl, aryl-alkenyl, arylynyl, etc. A linker may be neutral or carry a positive or negative charge. Additionally, a linker may be cleavable, such that the covalent bond connecting the linker to another chemical group can be broken or cleaved under certain conditions, including pH, temperature, salt concentration, light, catalysts, or enzymes. In some embodiments, the linker may be a peptide linker. A peptide linker may be a flexible amino acid linker (e.g., comprising small, nonpolar or polar amino acids). Non-limiting examples of flexible joints include LEGGGS (SEQ ID NO: 33), TGSG (SEQ ID NO: 34), GGSGGGSG (SEQ ID NO: 35), and (GGGGS). 1-4 (SEQ ID NO: 36) and (Gly) 6-8 (SEQ ID NO:37). Alternatively, the peptide linker can be a rigid amino acid linker. Such linkers include (EAAAK). 1-4 (SEQ ID NO:38), A (EAAAK) 2-5 A (SEQ ID NO: 39), PAPAP (SEQ ID NO: 40), and (AP) 6-8 (SEQ ID NO:41). Other examples of suitable connectors are well known in the art, and procedures for designing connectors are readily available (e.g., Crasto et al., Protein Eng., 2000, 13(5):309-312).
[0084] (m) Generation of modified Cas9 protein
[0085] In some embodiments, the modified Cas9 protein can be recombinantly produced in cell-free systems, bacterial cells, or eukaryotic cells and purified using conventional purification methods. In other embodiments, the modified Cas9 protein is produced in vivo in target eukaryotic cells from nucleic acids encoding the modified Cas9 protein (see section (III) below and incorporated herein by reference in section (I)).
[0086] In embodiments where the modified Cas9 protein contains nuclease or nickase activity, the modified Cas9 protein may further include at least one nuclear localization signal, a cell penetration domain and / or a marker domain, and at least one chromatin disruption domain. In embodiments where the modified Cas9 protein is linked to an epigenetic modification domain, the modified Cas9 protein may further include at least one nuclear localization signal, a cell penetration domain and / or a marker domain, and at least one chromatin disruption domain. Furthermore, in embodiments where the modified Cas9 protein is linked to a transcriptional regulatory domain, the modified Cas9 protein may further include at least one nuclear localization signal, a cell penetration domain and / or a marker domain, and at least one chromatin disruption domain and / or at least one RNA aptamer binding domain.
[0087] (II) Modified Cas9 system
[0088] Another aspect of this disclosure provides a modified Cas9 system comprising a modified Cas9 protein variant as discussed in the preceding section (I) incorporated herein by reference (e.g., a modified Cas9 protein variant including modifications at one or more (e.g., two or three) of amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9), wherein lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine at one or more of the aforementioned amino acid positions (R) is changed to leucine (L) or glutamine (Q)) and a modified guide RNA, wherein each modified guide RNA is designed to complex with a particular modified Cas9 protein. Each modified guide RNA contains a 5' guide sequence designed to hybridize with a target sequence in a double-stranded sequence, wherein the target sequence is the 5' of a prespacer adjacent motif (PAM).
[0089] (a) Modified guide RNA
[0090] The modified guide RNA is designed to complex with a specific modified Cas9 protein. The guide RNA comprises (i) a CRISPR RNA (crRNA) containing a guide sequence at its 5' end for hybridization with the target sequence, and (ii) a trans-acting crRNA (tracrRNA) sequence for recruiting the Cas9 protein. The crRNA guide sequence is different for each guide RNA (i.e., sequence-specific). The tracrRNA sequence is generally identical within the guide RNA, which is designed to complex with a Cas9 protein from a specific bacterial species.
[0091] The crRNA guide sequence is designed to hybridize with the target sequence (i.e., the prespacer sequence) in the double-stranded sequence. Generally, the complementarity between the crRNA and the target sequence is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. In specific embodiments, the complementarity is perfect (i.e., 100%). In various embodiments, the length of the crRNA guide sequence can range from about 15 nucleotides to about 25 nucleotides. For example, the length of the crRNA guide sequence can be about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides. In specific embodiments, the crRNA is about 19, 20, or 21 nucleotides long. In one embodiment, the crRNA guide sequence has a length of 20 nucleotides.
[0092] The guide RNA contains a repetitive sequence forming at least one stem-loop structure that interacts with the Cas9 protein, and a 3' sequence that maintains a single strand. The length of each loop and stem can vary. For example, the length of a loop can range from about 3 to about 10 nucleotides, while the length of a stem can range from about 6 to about 20 base pairs. The stem may contain one or more protrusions of about 1 to about 10 nucleotides. The length of the single-stranded 3' region can vary. The tracrRNA sequence in the modified guide RNA is generally based on the coding sequence of the wild-type tracrRNA in the target bacterial species. The wild-type sequence can be modified to promote secondary structure formation, increase secondary structure stability, promote expression in eukaryotic cells, etc. For example, one or more nucleotide changes can be introduced into the coding sequence of the guide RNA (see Example 3 below). The length of the tracrRNA sequence can range from about 50 nucleotides to about 300 nucleotides. In various embodiments, the length of tracrRNA can range from about 50 to about 90 nucleotides, about 90 to about 110 nucleotides, about 110 to about 130 nucleotides, about 130 to about 150 nucleotides, about 150 to about 170 nucleotides, about 170 to about 200 nucleotides, about 200 to about 250 nucleotides, or about 250 to about 300 nucleotides.
[0093] Generally, the modified guide RNA is a single molecule (i.e., a chimeric single guide RNA or sgRNA) in which a crRNA sequence is linked to a tracrRNA sequence. However, in some embodiments, the modified guide RNA can be two separate molecules (e.g., a bimolecular guide RNA). For example, the guide RNA may include a first molecule (or region) containing crRNA and a second molecule (or region) containing tracrRNA, wherein the crRNA contains a 3' sequence (containing about 6 to about 20 nucleotides) capable of pairing with the 5' base of the second molecule, and the tracrRNA contains a 5' sequence (containing about 6 to about 20 nucleotides) capable of pairing with the 3' base of the first molecule (or region).
[0094] In some implementations, the tracrRNA sequence of the modified guide RNA may be modified to include one or more aptamer sequences (Konermann et al., Nature, 2015, 517(7536):583-588; Zalatan et al., Cell, 2015, 160(1-2):339-50). Suitable aptamer sequences include those that bind to adaptor proteins selected from: MCP, PCP, Com, SLBP, FXR1, AP205, BZ13, f1, f2, fd, fr, ID2, JP34 / GA, JP501, JP34, JP500, KU1, M11, M12, MX1, NL95, PP7, ϕCb5, ϕCb8r, ϕCb12r, ϕCb23r, Qβ, R17, SP-β, TW18, TW19, VK, fragments thereof, or derivatives thereof. Those skilled in the art will understand that the length of the aptamer sequence can vary.
[0095] In other embodiments, the guide RNA may further comprise at least one detectable marker. The detectable marker may be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags, or a suitable fluorescent dye), a detection tag (e.g., biotin, digitoxin, etc.), a quantum dot, or a gold particle.
[0096] Guide RNA may contain standard ribonucleotides and / or modified ribonucleotides. In some embodiments, guide RNA may contain standard or modified deoxyribonucleotides. In embodiments in which guide RNA is enzymatically synthesized (i.e., in vivo or in vitro), guide RNA generally contains standard ribonucleotides. In embodiments in which guide RNA is chemically synthesized, guide RNA may contain standard or modified ribonucleotides and / or deoxyribonucleotides. Modified ribonucleotides and / or deoxyribonucleotides include base modifications (e.g., pseudouridine, 2-thiouridine, N6-methyladenosine, etc.) and / or sugar modifications (e.g., 2'-O-methyl, 2'-fluorine, 2'-amino, locked nucleic acid (LNA), etc.). The backbone of guide RNA may also be modified to include phosphate thioester bonds, borane phosphate bonds, or peptide nucleic acids.
[0097] (b) PAM sequence
[0098] The modified Cas9 system detailed above targets a specific sequence in double-stranded DNA upstream of the PAM sequence. The PAM sequence can include canonical 5'-NGG-3' PAM or non-canonical PAMs, such as 5'-NAG-3' PAM. In some embodiments, the modified Cas9 system detailed above can be modified to recognize alternative PAMs, such as 5'–NGAN–3', 5'–NGNG–3', and 5'–NGCG–3' PAM.
[0099] (III) Nucleic Acids
[0100] A further aspect of this disclosure provides nucleic acids encoding modified Cas9 protein variants and systems described in preceding sections (I) and (II) incorporated herein by reference (III) (e.g., modified Cas9 protein variants including modifications at one or more (e.g., two or three) of amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9), wherein lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine (R) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q)). Proteins and systems may be encoded by a single nucleic acid or multiple nucleic acids. Nucleic acids may be DNA or RNA, linear or circular, single-stranded or double-stranded. RNA or DNA may be codon-optimized for efficient translation into proteins in target eukaryotic cells. Codon optimization programs are available as free software or from commercial sources.
[0101] In some embodiments, the nucleic acid encoding the modified Cas9 protein can be RNA. The RNA can be synthesized enzymatically in vitro. For this purpose, the DNA encoding the modified Cas9 protein can be operatively linked to a promoter sequence recognized by a bacteriophage RNA polymerase for in vitro RNA synthesis. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variant of the T7, T3, or SP6 promoter sequence. The DNA encoding the modified protein can be a portion of a vector, as detailed below. In such embodiments, the RNA transcribed in vitro can be purified, capped, and / or polyadenylated. In other embodiments, the RNA encoding the modified Cas9 protein can be a portion of self-replicating RNA (Yoshioka et al., Cell Stem Cell, 2013, 13:246-254). Self-replicating RNA can be derived from non-infectious, self-replicating Venezuelan equine encephalitis (VEE) virus RNA replicons, which are positive single-stranded RNAs capable of self-replicating a limited number of cell divisions and can be modified to encode target proteins (Yoshioka et al., Cell Stem Cell, 2013, 13:246-254).
[0102] In other embodiments, the nucleic acid encoding the modified Cas9 protein can be DNA. The DNA coding sequence can be operatively linked to at least one promoter control sequence for expression in target cells. In some embodiments, the DNA coding sequence can be operatively linked to a promoter sequence for expressing the modified Cas9 protein in bacterial (e.g., *E. coli*) or eukaryotic (e.g., yeast, insect, or mammalian) cells. Suitable bacterial promoters include, but are not limited to, the T7 promoter. lac Operator promoter, trp promoter, tac promoter (which is) trp and lacHybrids of promoters), any of the foregoing variants, and any combination thereof. Non-limiting examples of suitable eukaryotic promoters include constitutive, regulatory, or cell or tissue-specific promoters. Suitable constitutive eukaryotic promoter control sequences include, but are not limited to, cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglycerate kinase (PGK) promoter, elongation factor (ED1)-α promoter, ubiquitin promoter, actin promoter, microtubule promoter, immunoglobulin promoter, fragments thereof, or any combination thereof. Examples of suitable regulatory eukaryotic promoter control sequences include, but are not limited to, those regulated by heat shock, metals, steroids, antibiotics, or alcohols. Non-limiting examples of tissue-specific promoters include the B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase-1 promoter, endothelial glycoprotein promoter, fibronectin promoter, Flt-1 promoter, GFAP promoter, GPIIb promoter, ICAM-2 promoter, INF-β promoter, Mb promoter, NphsI promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter. The promoter sequence can be wild-type, or it can be modified for more efficient or effective expression. In some embodiments, the DNA coding sequence may also be linked to a polyadenylation signal (e.g., SV40 polyA signal, bovine growth hormone (BGH) polyA signal, etc.) and / or at least one transcription termination sequence. In some cases, the modified Cas9 protein can be purified from bacterial or eukaryotic cells.
[0103] In other embodiments, the modified guide RNA may be encoded by DNA. In some cases, the DNA encoding the modified guide RNA may be operatively ligated to a promoter sequence recognized by bacteriophage RNA polymerase for in vitro RNA synthesis. For example, the promoter sequence may be a T7, T3, or SP6 promoter sequence, or a variant of the T7, T3, or SP6 promoter sequence. In other cases, the DNA encoding the modified guide RNA may be operatively ligated to a promoter sequence recognized by RNA polymerase III (PolIII) for expression in target eukaryotic cells. Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters.
[0104] In various embodiments, the nucleic acid encoding the modified Cas9 protein may be present in the vector. In some embodiments, the vector may further contain nucleic acid encoding a modified guide RNA. Suitable vectors include plasmid vectors, viral vectors, and self-replicating RNA (Yoshioka et al., Cell Stem Cell, 2013, 13:246-254). In some embodiments, the nucleic acid encoding the complex or fusion protein may be present in the plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and variants thereof. In other embodiments, the nucleic acid encoding the complex or fusion protein may be part of a viral vector (e.g., lentiviral vector, adeno-associated virus vector, adenovirus vector, etc.). The plasmid or viral vector may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), optional marker sequences (e.g., antibiotic resistance genes), origin of replication, etc. Further information about the vector and its uses can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001.
[0105] (IV) Eukaryotic cells
[0106] Another aspect of this disclosure includes eukaryotic cells comprising at least one modified Cas9 protein variant as detailed in section (I) above, incorporated herein by reference (IV) (e.g., a modified Cas9 protein variant including modifications at one or more (e.g., two or three) of amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (refer to the numbering system of Streptococcus pyogenes Cas9-SpCas9), wherein lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine (R) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q)), and / or at least one nucleic acid encoding a modified Cas9 protein and / or system and / or modified guide RNA as detailed in sections (I), (II), and (III) above (each incorporated herein by reference (IV)).
[0107] Eukaryotic cells can be human cells, non-human mammalian cells, non-mammal vertebrate cells, invertebrate cells, plant cells, or single-celled eukaryotes. Examples of suitable eukaryotic cells are detailed in section (V)(c) below. Eukaryotic cells can be in vitro, ex vivo, or in vivo.
[0108] (V) Methods for modifying sequences
[0109] A further aspect of this disclosure includes methods for modifying chromosome sequences in eukaryotic cells. Generally, the method includes introducing at least one modified Cas9 system as detailed in section (II) above into a target eukaryotic cell, said modified Cas9 system further comprising a modified Cas9 protein variant as detailed in section (I) above, wherein each of sections (I) and (II) is incorporated herein by reference (V) (e.g., including a modified Cas9 protein variant modified at one or more (e.g., two or three) of amino acid positions 526, 562, 652, 661, 691, 780, 810, 848, 855, 1003, and 1060 (refer to the numbering system for Streptococcus pyogenes Cas9-SpCas9), wherein lysine (K) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q), and / or arginine (R) at one or more of the aforementioned amino acid positions is changed to leucine (L) or glutamine (Q)); and / or encoding as described in sections (I), (II), and (III) above. (each of which is incorporated herein by reference in this section (V)) at least one nucleic acid of the modified Cas9 protein and / or system and / or modified guide RNA, as detailed therein, is introduced into the target eukaryotic cell.
[0110] In embodiments where the modified Cas9 protein contains nuclease or nickase activity, the chromosomal sequence modification may include the substitution, deletion, or insertion of at least one nucleotide. In some iterations, the method includes introducing one modified Cas9 system containing nuclease activity or two modified Cas9 systems containing nickase activity but without a donor polynucleotide into a eukaryotic cell, causing one or more modified Cas9 systems to introduce double-strand breaks at target sites in the chromosomal sequence, and inactivating the chromosomal sequence (i.e., insertion / deletion) through double-strand break repair via a cellular DNA repair process. In other iterations, the method includes introducing one modified Cas9 system containing nuclease activity or two modified Cas9 systems containing nickase activity and a donor polynucleotide into a eukaryotic cell, causing one or more modified Cas9 systems to introduce double-strand breaks at target sites in the chromosomal sequence, and inactivating the chromosomal sequence (i.e., gene knock-in) through double-strand break repair via a cellular DNA repair process.
[0111] In embodiments in which the modified Cas9 protein contains epigenetic modification activity or transcriptional regulatory activity, chromosomal sequence modification may include the conversion of at least one nucleotide on or near a target site, modification of at least one nucleotide on or near a target site, modification of at least one histone on or near a target site, and / or transcriptional changes on or near a target site.
[0112] Furthermore, it should be understood that the modified Cas9 variants described herein can also be used to modify genomes other than those of eukaryotic cells, such as microbial genomes.
[0113] (a) Introduced into cells
[0114] As mentioned above, the method involves introducing at least one modified Cas9 system and / or a nucleic acid (and optionally a donor polynucleotide) encoding said system into a eukaryotic cell. At least one system and / or nucleic acid / donor polynucleotide can be introduced into the target cell by various means.
[0115] In some embodiments, cells can be transfected with appropriate molecules (i.e., proteins, DNA, and / or RNA). Suitable transfection methods include nuclear transfection (or electroporation), calcium phosphate-mediated transfection, cationic polymer transfection (e.g., DEAE-glucan or polyethyleneimine), viral transduction, virion transfection, viral particle transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipid transfection, dendritic molecule transfection, heat shock transfection, magnetic transfection, lipid transfection, gene gun delivery, puncture transfection, sonoporosis, phototransfection, and nucleic acid uptake enhanced with proprietary reagents. Transfection methods are well known in the art (see, for example, “Current Protocols in Molecular Biology”, Ausubel et al., John Wiley & Sons, New York, 2003, or “Molecular Cloning: A Laboratory Manual”, Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001). In other embodiments, molecules can be introduced into cells via microinjection. For example, molecules can be injected into the cytoplasm or nucleus of target cells. The amount of each molecule introduced into the cell can vary, but those skilled in the art are familiar with the means used to determine the appropriate amount.
[0116] Various molecules can be introduced into the cell simultaneously or sequentially. For example, a modified Cas9 system (or its encoded nucleic acid) and a donor polynucleotide can be introduced simultaneously. Alternatively, one can be introduced first, followed by another.
[0117] Generally, cells are maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art and are described, for example, in Santiago et al., Proc. Natl. Acad. Sci. USA, 2008, 105:5809-5814; Moehle et al., Proc. Natl. Acad. Sci. USA, 2007, 104:3055-3060; Urnov et al., Nature, 2005, 435:646-651; and Lombardo et al., Nat. Biotechnol., 2007, 25:1298-1306. Those skilled in the art will understand that methods used for culturing cells are known in the art and can and will vary depending on the cell type. In all cases, routine optimization can be used to determine the best technique for a particular cell type.
[0118] (b) Optional donor polynucleotides
[0119] In embodiments where the modified Cas9 protein contains nuclease or nickase activity, the method may further include introducing at least one donor polynucleotide into the cell. The donor polynucleotide may be single-stranded or double-stranded, linear or circular, and / or RNA or DNA. In some embodiments, the donor polynucleotide may be a vector, such as a plasmid vector.
[0120] The donor polynucleotide contains at least one donor sequence. In some respects, the donor sequence of the donor polynucleotide can be a modified form of an endogenous or natural chromosomal sequence. For example, the donor sequence can be substantially equivalent to a portion of a chromosomal sequence at or near the sequence targeted by a modified Cas9 system, but it contains at least one nucleotide variation. Thus, after integration or exchange with the natural sequence, the sequence at the target chromosomal location contains at least one nucleotide variation. For example, the variation can be the insertion of one or more nucleotides, the deletion of one or more nucleotides, the substitution of one or more nucleotides, or a combination thereof. As a result of the integration of the modified sequence as a form of “genetic correction,” the cell can produce a modified gene product from the target chromosomal sequence.
[0121] In other respects, the donor sequence of the donor polynucleotide can be a foreign sequence. As used herein, a "foreign" sequence refers to a sequence that is not native to the cell or whose native location is different from that in the cell's genome. For example, a foreign sequence can contain a protein-coding sequence that can be operatively linked to a foreign promoter control sequence, such that upon integration into the genome, the cell can express the protein encoded by the integrated sequence. Alternatively, the foreign sequence can be integrated into a chromosomal sequence such that its expression is regulated by an endogenous promoter control sequence. In other iterations, the foreign sequence can be a transcriptional control sequence, another expression control sequence, an RNA-coding sequence, etc. As noted above, the integration of a foreign sequence into a chromosomal sequence is referred to as "knock-in".
[0122] As those skilled in the art will understand, the length of the donor sequence can and will vary. For example, the length of the donor sequence can range from a few nucleotides to hundreds of nucleotides to hundreds of thousands of nucleotides.
[0123] Typically, the donor sequence in a donor polynucleotide is flanked by upstream and downstream sequences, which share fundamental sequence identity with the sequences located upstream and downstream of the sequence targeted by the modified Cas9 system, respectively. Due to this sequence similarity, the upstream and downstream sequences of the donor polynucleotide allow homologous recombination between the donor polynucleotide and the target chromosomal sequence, enabling the donor sequence to integrate into (or exchange with) the chromosomal sequence.
[0124] As used herein, an upstream sequence refers to a nucleic acid sequence that shares substantially the same sequence identity as a chromosomal sequence upstream of a sequence targeted by a modified Cas9 system. Similarly, a downstream sequence refers to a nucleic acid sequence that shares substantially the same sequence identity as a chromosomal sequence downstream of a sequence targeted by a modified Cas9 system. As used herein, the phrase "substantially the same sequence identity" means a sequence having at least about 75% sequence identity. Therefore, the upstream and downstream sequences in the donor polynucleotide can have about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the sequence upstream or downstream of the target sequence. In one exemplary embodiment, the upstream and downstream sequences in the donor polynucleotide can have about 95% or 100% sequence identity with the chromosomal sequence upstream or downstream of the sequence targeted by a modified Cas9 system.
[0125] In some embodiments, the upstream sequence shares basic sequence identity with a chromosome sequence located directly upstream of the sequence targeted by the modified Cas9 system. In other embodiments, the upstream sequence shares basic sequence identity with a chromosome sequence located within approximately one hundred (100) nucleotides upstream of the target sequence. Thus, for example, the upstream sequence may share basic sequence identity with a chromosome sequence located within approximately 1 to approximately 20, approximately 21 to approximately 40, approximately 41 to approximately 60, approximately 61 to approximately 80, or approximately 81 to approximately 100 nucleotides upstream of the target sequence. In some embodiments, the downstream sequence shares basic sequence identity with a chromosome sequence located directly downstream of the sequence targeted by the modified Cas9 system. In other embodiments, the downstream sequence shares basic sequence identity with a chromosome sequence located within approximately one hundred (100) nucleotides downstream of the target sequence. Therefore, for example, the downstream sequence can share basic sequence identity with a chromosome sequence located about 1 to about 20, about 21 to about 40, about 41 to about 60, about 61 to about 80, or about 81 to about 100 nucleotides downstream of the target sequence.
[0126] Each upstream or downstream sequence can range in length from about 20 nucleotides to about 5000 nucleotides. In some embodiments, the upstream and downstream sequences can comprise about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, or 5000 nucleotides. In a specific embodiment, the upstream and downstream sequences can range in length from about 50 to about 1500 nucleotides.
[0127] (c) Cell type
[0128] Various cells are suitable for use in the methods disclosed herein, including prokaryotic cells (e.g., bacteria) and eukaryotic cells (e.g., animal cells, insect cells, and plant cells). For example, the cells can be human cells, non-human mammalian cells, non-mammal vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotes. In some embodiments, the cells can be single-celled embryos. For example, non-human mammalian embryos include rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cattle, horse, and primate embryos. In other embodiments, the cells can be stem cells, such as embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, etc. In one embodiment, the stem cells are not human embryonic stem cells. Furthermore, stem cells can include those prepared using the techniques disclosed in WO2003 / 046141, which is incorporated herein in its entirety, or Chung et al. (Cell Stem Cell, 2008, 2:113-117). Cells can be in vitro (i.e., in culture), ex vivo (i.e., within a tissue isolated from an organism), or in vivo (i.e., within an organism). In exemplary embodiments, the cells are mammalian cells or mammalian cell lines. In particular embodiments, the cells are human cells or human cell lines.
[0129] For example, in some implementations, the eukaryotic cells or population of eukaryotic cells are T cells, CD8 cells, etc. + T cells, CD8 + Naïve T cells, central memory T cells, effector memory T cells, CD4 +T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, pancreatic progenitor cells, endocrine progenitor cells, exocrine progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocyte erythroid progenitor cells, monocyte precursor cells, endocrine precursor cells, exocrine cells, fibroblasts, hepatocytes, myoblasts, macrophages, pancreatic β cells, cardiomyocytes, blood cells, ductal cells, acinar cells, α cells, β cells, δ cells, PP cells, bile duct cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells. Cells, alveolar epithelial cells, lung epithelial progenitor cells, skeletal muscle cells, cardiomyocytes, muscle satellite cells, myocytes, neurons, neuronal stem cells, mesenchymal stem cells, induced pluripotent stem (iPS) cells, embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, proto-B cells, memory B cells, plasma B cells), gastrointestinal epithelial cells, bile duct epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, hepatic macrophages, osteoblasts, osteoclasts, adipocytes (e.g., brown adipocytes or white adipocytes), preadipocytes, pancreatic progenitor cells, islet cells, pancreatic β cells, pancreatic α cells, pancreatic δ cells, pancreatic exocrine cells, Schwann cells or oligodendrocytes, or such cell populations.Non-limiting examples of suitable mammalian cells or cell lines include human induced pluripotent stem cells (hiPSCs), human T cells (autologous or allogeneic), human B cells, human macrophages, human hematopoietic stem cells (hHSCs), human hepatocytes, human retinal cells, pancreatic islets, human embryonic kidney cells (HEK293, HEK293T); human cervical cancer cells (HELA); human lung cells (W138); and human hepatocytes (Hep). G2); human U2-OS osteosarcoma cells, human A549 cells, human A-431 cells, and human K562 cells; Chinese hamster ovary (CHO) cells and young hamster kidney (BHK) cells; mouse myeloma NS0 cells, mouse embryonic fibroblast 3T3 cells (NIH3T3), and mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse myeloma SP2 / 0 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse cancer CT26 cells, small Mouse prostate DuCuP cells; mouse mammary EMT6 cells; mouse hepatocellular carcinoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelial MTD-1A cells; mouse cardiac MyEnd cells; mouse kidney RenCa cells; mouse pancreatic RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells; rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatocellular carcinoma cells (HTC); buffalo rat liver BRL 3A cells; canine kidney cells (MDCK); canine mammary gland (CMT) cells; rat osteosarcoma D17 cells; rat monocyte / macrophage DH82 cells; monkey kidney SV-40 transformed fibroblasts (COS7) cells; monkey kidney CVI-76 cells; African green monkey kidney (VERO-76) cells. A comprehensive list of mammalian cell lines can be found in the American Center for Type Culture Collections (ATCC, Manassas, VA).
[0130] Other aspects of this disclosure include animals modified to encode the nucleic acids or vectors described above, or animals permanently modified by a SpCas9 variant modified by this disclosure. For example, the animal could be a model organism (i.e., *Drosophila melanogaster*). Drosophila melanogaster Animals can be (e.g., mice, mosquitoes, rats), or farm animals, farmed fish, or pets. As another example, an animal can be a carrier of at least one disease. As yet another example, an organism can be a carrier of a human disease (i.e., mosquitoes, ticks, birds).
[0131] Other aspects of this disclosure include plants modified using the nucleic acids or vectors described above, or plants transiently or permanently modified by SpCas9 variants modified by this disclosure. For example, the plants can be crops (i.e., rice, soybeans, wheat, tobacco, cotton, alfalfa, low-erucic acid rapeseed, corn, sugar beets, etc.).
[0132] (VI) Application
[0133] The compositions and methods disclosed herein can be used in a variety of therapeutic, diagnostic, industrial, and research applications. In some embodiments, the disclosure can be used to modify any target chromosomal sequence in cells, animals, or plants to construct gene function models and / or study gene function, investigate target genetic or epigenetic conditions, or study biochemical pathways involved in various diseases or conditions. For example, transgenic organisms can be generated to construct models of diseases or conditions in which the expression of one or more nucleic acid sequences associated with the disease or condition is altered. Disease models can be used to study the effects of mutations on organisms, study the development and / or progression of diseases, study the effects of pharmaceutically active compounds on diseases, and / or evaluate the efficacy of potential gene therapy strategies.
[0134] In other embodiments, the compositions and methods can be used to perform efficient and cost-effective functional genomic screening, which can be used to study the function of genes involved in specific biological processes and how any alterations in gene expression can affect biological processes, or to perform saturation or depth scan mutagenesis of genomic loci linked to cell phenotypes. For example, saturation or depth scan mutagenesis can be used to determine the critical minimum characteristics and discrete vulnerabilities of functional elements required for gene expression, drug resistance, and disease reversal.
[0135] In further embodiments, the compositions and methods disclosed herein can be used for diagnostic testing to determine the presence of a disease or condition and / or to determine treatment options. Examples of suitable diagnostic tests include detecting specific mutations in cancer cells (e.g., specific mutations in EGFR, HER2, etc.), detecting specific mutations associated with specific diseases (e.g., trinucleotide repeats, mutations in β-globin associated with sickle cell disease, specific SNPs, etc.), detecting hepatitis, detecting viruses (e.g., Zika virus), and so on.
[0136] In other embodiments, the compositions and methods disclosed herein can be used to correct genetic mutations associated with specific diseases or conditions, such as correcting mutations in globin genes associated with sickle cell disease or thalassemia, correcting mutations in adenosine deaminase genes associated with severe combined immunodeficiency (SCID), reducing the expression of the causative gene HTT in Huntington's disease, or correcting mutations in rhodopsin genes for the treatment of retinitis pigmentosa. Such modifications can be performed in vitro in cells.
[0137] In other embodiments, the compositions and methods disclosed herein can be used to generate crop plants with improved traits or increased resistance to environmental stresses. This disclosure can also be used to generate farm animals or production animals with improved traits. For example, pigs possess many characteristics that make them attractive as biomedical models, particularly in regenerative medicine or xenotransplantation.
[0138] For example, this disclosure provides nucleotide sequences or nucleic acids or vectors as described above for use as agents for gene therapy. This disclosure also provides pharmaceutical compositions comprising nucleotide sequences or nucleic acids or vectors as described above, and at least one pharmaceutically acceptable excipient. The invention further provides pharmaceutical compositions comprising a recombinant Cas9 polypeptide containing the above-described mutation and at least one pharmaceutically acceptable excipient. Pharmaceutically acceptable excipients typically include inactive ingredients used as mediators (e.g., water, capsule shells, etc.), diluents, or components constituting a dosage form or pharmaceutical composition containing a drug such as a therapeutic agent. Pharmaceutically acceptable excipients also typically include inactive ingredients that impart adhesive (i.e., adhesive), disintegrant (i.e., disintegrant), lubricant (lubricant), and / or other functions (i.e., solvent, surfactant, etc.) to the composition. Furthermore, this disclosure provides in vitro uses of nucleotide sequences or nucleic acids or vectors as described above for genome engineering, cell engineering, protein expression, or other biotechnological applications. Furthermore, this disclosure provides in vitro uses of recombinant Cas9 peptides containing the above-described mutations, together with guide RNA (e.g., single-molecule (i.e., chimeric) guide RNA or bimolecule (i.e., two-part) guide RNA) for genome engineering, cell engineering, protein expression, or other biotechnological applications.
[0139] Other aspects of this disclosure relate to kits that include various components described herein, such as Cas9 protein variants, guide RNA, vectors, primers, etc., including instructions for their use in genome engineering, cell engineering, protein expression, or other biotechnological applications.
[0140] definition
[0141] The following definitions and methods are provided to better define the invention and to guide those skilled in the art in its practice. Unless otherwise stated, the terminology should be understood by those skilled in the art based on its conventional usage.
[0142] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. The following references provide general definitions for many of the terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, unless otherwise stated, the following terms have their respective meanings.
[0143] When describing elements of this disclosure or its preferred embodiments, the articles “a,” “an,” “the,” and “described” are intended to mean the presence of one or more elements. The terms “comprising,” “including,” and “having” are intended to be included and mean that there may be additional elements besides those listed.
[0144] When used in relation to a numerical value x, the term “about” means, for example, x ± 5%.
[0145] As used herein, the terms “complementary” or “complementarity” refer to the association of double-stranded nucleic acids via base pairing through specific hydrogen bonds. Base pairing can be standard Watson-Crick base pairing (e.g., 5'-AGT C-3' paired with the complementary sequence 3'-TCA G-5'). Base pairing can also be Hoogsteen or reverse Hoogsteen hydrogen bonds. Complementarity is typically measured relative to the double-stranded region and therefore, for example, excludes overhangs. If only some (e.g., 70%) bases are complementary, the complementarity between the two strands of the double-stranded region can be partial and expressed as a percentage (e.g., 70%). Bases that are not complementary are “mismatched”. If all bases in the double-stranded region are complementary, the complementarity can also be complete (i.e., 100%).
[0146] As used herein, the term “CRISPR / Cas system” or “Cas9 system” refers to a complex containing the Cas9 protein (i.e., nuclease, nickase, or catalytic death protein) and guide RNA.
[0147] As used in this article, the term "endogenous sequence" refers to the chromosome sequence native to a cell.
[0148] As used in this article, the term "exogenous" refers to a sequence that is not natural to the cell, or a chromosomal sequence whose natural location in the cell's genome is at a different chromosomal position.
[0149] As used herein, “gene” refers to the DNA region (including exons and introns) that encodes a gene product, as well as all DNA regions that regulate the production of gene products, regardless of whether such regulatory sequences are adjacent to coding and / or transcriptional sequences. Accordingly, a gene includes, but is not limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions.
[0150] The term "heterogeneous" refers to an entity that is not endogenous or natural for the target cell. For example, a heterologous protein is a protein derived from or originally derived from a foreign source (e.g., a foreign-introduced nucleic acid sequence). In some cases, heterologous proteins are not typically produced by the target cell.
[0151] The term "nicking enzyme" refers to an enzyme that cuts one strand of a double-stranded nucleic acid sequence (i.e., creates a nick in the double-stranded sequence). For example, nucleases with double-strand cleaving activity can be modified by mutation and / or deletion to act as nicking enzymes and cut only one strand of a double-stranded sequence.
[0152] As used in this article, the term "nuclease" refers to an enzyme that cuts both strands of a double-stranded nucleic acid sequence.
[0153] The terms “nucleic acid” and “polynucleotide” refer to polymers of deoxyribonucleotides or ribonucleotides in linear or cyclic conformations and in single- or double-stranded form. For the purposes of this disclosure, these terms should not be construed as limitations on polymer length. These terms may include known analogs of natural nucleotides, as well as nucleotides modified in the base, sugar, and / or phosphate ester moieties (e.g., phosphate thioester backbone). Generally, analogs of a particular nucleotide have the same base-pairing specificity; that is, an analog of A will pair with a T base.
[0154] The term "nucleotide" refers to deoxyribonucleotides or ribonucleotides. Nucleotides can be standard nucleotides (i.e., adenosine, guanosine, cytidine, thymidine, and uridine), nucleotide isomers, or nucleotide analogs. Nucleotide analogs refer to nucleotides having modified purine or pyrimidine bases or modified ribose moieties. Nucleotide analogs can be naturally occurring nucleotides (e.g., inosine, pseudouridine, etc.) or non-naturally occurring nucleotides. Non-limiting examples of modifications to the sugar or base moieties of nucleotides include the addition (or removal) of acetyl, amino, carboxyl, carboxymethyl, hydroxy, methyl, phosphoryl, and thiol groups, as well as the substitution of the carbon and nitrogen atoms of the base by other atoms (e.g., 7-denitropurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methylnucleotides, locked nucleic acids (LNAs), peptide nucleic acids (PNAs), and morpholinonucleotides.
[0155] The terms “polypeptide” and “protein” are used interchangeably to refer to polymers of amino acid residues.
[0156] The terms “target sequence,” “target chromosome sequence,” and “target site” are used interchangeably to refer to a specific sequence in the chromosomal DNA targeted by the modified Cas9 system, and the site where the modified Cas9 system modifies the DNA or DNA-related proteins.
[0157] Techniques for determining the identity of nucleic acid and amino acid sequences are known in the art. Typically, such techniques involve determining the nucleotide sequence of a gene's mRNA and / or the amino acid sequence it encodes, and comparing these sequences with second nucleotide or amino acid sequences. Genomic sequences can also be determined and compared in this manner. Generally, identity refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence between two polynucleotide or polypeptide sequences, respectively. Two or more sequences (polynucleotides or amino acids) can be compared by determining their percentage identity. The percentage identity of two sequences (whether nucleic acid or amino acid sequences) is the number of exact matches between the two aligned sequences divided by the length of the shorter sequence and multiplied by 100. Approximate alignments of nucleic acid sequences are provided by the local homology algorithm in Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm can be applied to amino acid sequences using a scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff (ed.), 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, DC, USA, and normalized by Gribskov, Nucl. Acids Res. 14 (6):6745-6763 (1986). An exemplary implementation of this algorithm for determining the percentage identity of sequences is provided by Genetics Computer Group (Madison, Wis.) in the "BestFit" utility. Other suitable procedures for calculating the percentage identity or similarity between sequences are generally known in the art; for example, another alignment procedure is BLAST used with default parameters. For example, the following default parameters can be used to use BLASTN and BLASTP: Genetic code = standard; Filter = none; Strains = two; Truncation = 60; Expectation = 10; Matrix = BLOSUM62; Description = 50 sequences; Sort = high score; Database = non-redundant, GENBANK+EMBL+DDBJ+PDB+GENBANK CDS translation + Swiss protein + Spupdate + PIR. Details of these procedures can be found on the GENBANK NIH Genetic Sequence Database website.
[0158] The invention has been described in detail, and it will be apparent that modifications and variations are possible without departing from the scope of the invention as defined in the appended claims. Furthermore, it should be understood that all examples in this disclosure are provided as non-limiting examples. Example
[0159] The following non-limiting embodiments are provided to further illustrate the invention. Those skilled in the art will understand that the techniques disclosed in the following embodiments represent methods that the inventors have found to function well in the practice of the invention, and can therefore be considered as examples constituting their mode of practice. However, in view of this disclosure, those skilled in the art will understand that many changes can be made to the specific embodiments disclosed, and the same or similar results can still be obtained without departing from the spirit and scope of the invention.
[0160] Example 1: Different amino acid substitutions on K855 exhibit different mid-target activities.
[0161] Wild-type SpCas9 was mutated at the K855 residue to alanine, glutamic acid, isoleucine, methionine, or glutamine, and the recombinant proteins were purified from *E. coli* to greater than 95% homogeneity. The amino acid sequences of the K855Q mutant proteins are listed in Table 1. All K855 mutant proteins, except for the single K855 mutation, share the same polypeptide sequence. Wild-type SpCas9 protein was purchased from MilliporeSigma as a control. A chemically synthesized HEKSite4 single guide RNA (sgRNA) with a guide sequence of 5'-GGCACUGCGGCUGGAGGUGG-3' (SEQ ID NO: 42) was also purchased from MilliporeSigma. Each protein was tested in triplicate.
[0162] The ribonucleoprotein (RNP) complex was prepared by adding buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5), 150 pmol sgRNA, and 8 µg Cas9 protein to a 10 µL total reaction volume in a 1.5 mL microcentrifuge tube. The sgRNA / Cas9 protein molar ratio was approximately 3:1. The complex was incubated at room temperature for 15 min and then kept on ice until transfection. Human U-2OS cells, confluent at approximately 80%, were detached with trypsin solution and washed twice with Hank's balanced salt solution. The cells were then incubated at approximately 0.25 x 10⁻⁶ cells per 100 µL. 6Cells were resuspended in Nucleofector Solution V (Lonza). Nuclear transfection was performed by transferring 100 µL of cells into the RNP complex and immediately mixing by gentle up-and-down aspiration without introducing air bubbles, followed by transfer to cuvettes for electroporation using the Amaxa X-001 procedure. Cells were immediately transferred to 6-well plates containing 2 mL of preheated medium per well and grown at 37°C and 5% CO2 for 3 days, after which they were harvested for genomic modification assays.
[0163] Genomic DNA extracts from transfected cells were prepared using QuickExtract solution. PCR amplification of the target genomic region was performed using the KAPA HiFi HotStart ReadyMix PCR Kit (Roche) with next-generation sequencing (NGS) primers under the following cycling conditions: 95℃ / 3m; 98℃ / 20s, 68℃ / 30s, and 72℃ / 45s for 34 cycles; 72℃ / 5m. The NGS primers for the HEKSite4 target site were: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNNNNNNGGAACCCAGGTAGCCAGAGA-3' (forward) (SEQ ID NO: 43) and 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNNNNNGGGGTGGGGTCAGACGT-3' (reverse) (SEQ ID NO: 44). The NGS primers for the HEKSite4 off-target site are: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNNNNNNCTAGAGCAAACCTTGGCATTGTCC-3' (forward) (SEQ ID NO: 45) and 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNNNNNACCCTCTACCCTCCCTGATG-3' (reverse) (SEQ ID NO: 46). The primary PCR product was then re-amplified using the JumpStart™ TaqReadyMix™ for Quantitative PCR Kit (MilliporeSigma) with Illumina index primers under the following cycling conditions: 95℃ / 3m; 95℃ / 30s, 55℃ / 30s, and 72℃ / 30s for 8 cycles; 72℃ / 5m. Indexed PCR products were purified using the Select-a-Size DNA Clean & Concentrator kit (Zymo) and quantified using PicoGreen (ThermoFisher). The PCR products were then normalized and pooled to prepare NGS libraries. NGS was performed using an Illumina MiSeq instrument and a 2 x 300 bp kit. The FASTQ file for each sample was analyzed for genome editing frequencies using the NGS analysis pipeline.
[0164] The results are presented in Figure 1A and 1BThe results showed that different K855 mutant proteins exhibited varying levels of mid-target activity, and all five K855 mutant proteins reduced off-target effects to essentially similar levels. The results also indicated that glutamate and alanine were not optimal substitutes for K855 in maintaining mid-target activity.
[0165] Table 1. SpCas9 K855Q amino acid sequence
[0166]
[0167] Example 2: Dual mutant variants with optimal amino acid substitutions maintain intermediate target activity
[0168] Different amino acid substitutions at R661, N692, or Q695 were introduced into the K855M and K855Q mutant background to generate double mutants. The recombinant proteins were purified from *E. coli* to greater than 95% homogeneity. Except for the specified mutations, all double mutants shared the same polypeptide sequence as the K855Q mutant listed in Table 1. Each protein was tested for the same HEKSite4 target site in U2-OS cells in three biological replicates. RNP complex preparation, cell transfection, and NGS analysis were performed as described in Example 1.
[0169] The results are presented in Figure 2 The results showed that different amino acid substitutions at R661, N692, or Q695 resulted in varying levels of mid-target activity. At the R661 residue, isoleucine substitution led to a significant reduction in activity, while leucine, asparagine, or glutamine substitutions maintained the same activity level as WT Cas9. Substitutions of the two uncharged residues were less predictable.
[0170] Example 3: A triple mutant characterized by a balance of specificity and activity
[0171] Leucine or glutamine substitutions at K526, K562, K652, R691, R780, K810, K848, K1003, or R1060 were introduced into the R661L-K855Q background to generate 18 triple mutants and one quadruple mutant (R661L-K855Q-K1003Q-R1060Q). Except for the specified mutations, all triple and quadruple mutants shared the same polypeptide sequence as the K855Q mutant listed in Table 1. Recombinant proteins were purified from *E. coli* to greater than 95% homogeneity. Synthetic sgRNAs targeting human FANCF02 and HBB03 were purchased from Millipore Sigma. The guide sequences of these sgRNAs are listed in Table 2. eSpCas9 1.1 protein was purchased from Millipore Sigma, and HiFi Cas9 V3 protein was purchased from Integrated DNA Technologies. Each protein was tested in three biological replicates.
[0172] The RNP complex was prepared as described in Example 1. Human K562 cells were injected with 0.25 x 10⁻⁶ cells one day before transfection. 6 Seed at 100 cells / mL, and at transfection at approximately 0.5 x 10⁻⁶ cells / mL. 6 Cells / mL. Cells were washed twice with Hank's balanced salt solution, then diluted at approximately 0.35 x 10⁻⁶ cells / mL. 6 Cells were resuspended in Nucleofector Solution V (Lonza). Nuclear transfection was performed by transferring 100 µL of cells into the RNP complex and immediately mixing by gently aspirating without introducing air bubbles, then transferring the mixture into cuvettes for electroporation using the Amaxa program T-016. Cells were immediately transferred to 6-well plates containing 2 mL of preheated medium per well and grown at 37°C and 5% CO2 for 3 days, then harvested for genomic modification assays. Genomic DNA extracts from transfected cells were prepared using QuickExtract solution. PCR amplification of the target genomic region was performed using the JumpStart™ Taq ReadyMix™ for Quantitative PCR Kit (MilliporeSigma) with NGS primers under the following cycling conditions: 98°C / 2 min; 98°C / 15 s, 62°C / 30 s, and 72°C / 45 s for 34 cycles; 72°C / 5 min. NGS primer sequences are listed in Table 2. NGS library preparation, sequencing, and data analysis were performed as described in Example 1.
[0173] The results are presented in Figure 3A , 3B In 3C and 3D. Figure 3A and 3B The results showed that all proteins were highly active at the FANCF02 target site with only minor variations. However, there was a wide range of variations in the off-target mutation frequency at the single mismatch off-target site of FANCF02. Six triple mutant proteins outperformed eSpCas9 1.1 in reducing off-target effects. These included K526L-R661L-K855Q, R661L-R691L-K855Q, R661L-R780L-K855Q, R661L-R780Q-K855Q, R661L-K810L-K855Q, and R661L-K848L-K855Q. Compared to WTCas9, except for the aberrant mutant R661L-R691Q-K855Q, the remaining mutant proteins are comparable to eSpCas9 1.1 or fall between eSpCas9 1.1 and HiFi Cas9 V3 in terms of reducing off-target mutation frequency. Figure 3C and 3D The results further differentiated the levels of intermediate-target activity and specificity in these proteins. For example, the highly specific mutant protein identified at the FANCF02 site almost completely lost all intermediate-target activity at the HBB03 site. However, six triple mutant proteins had off-target mutation frequencies similar to eSpCas9 1.1, but at the HBB03 site, they exhibited substantially higher levels of intermediate-target activity than eSpCas9 1.1. Based on the combined results, this group of mutant proteins was identified as having a balanced specificity and activity. The selected mutant proteins included K562L-R661L-K855Q, K562Q-R661L-K855Q, K652L-R661L-K855Q, K652Q-R661L-K855Q, R661L-K855Q-K1003Q, and R661L-K855Q-R1060Q. Based on the combined results, four eSpCas9 1.1-like triple mutant proteins were also identified, including K526Q-R661L-K855Q, R661L-K810Q-K855Q, R661L-K855Q-K1003L, and R661L-K855Q-R1060L.
[0174] Table 2. sgRNA guide sequence and NGS primers
[0175]
[0176] Example 4: Specifically improved SpCas9 nuclease-mediated efficient editing across different genomic sites
[0177] sgRNAs targeting five human genomic loci were purchased from Millipore Sigma. The guide sequences of these sgRNAs are listed in Table 3. The RNP complex was prepared as described in Example 1. Human k562 cells were cultured at 0.25 x 10⁻⁶ cells per cell line one day prior to transfection. 6 Seed at 100 cells / mL, and at transfection at approximately 0.5 x 10⁻⁶ cells / mL. 6 Cells / mL. Cells were washed twice with Hank's balanced salt solution, then diluted at approximately 0.35 x 10⁻⁶ cells / mL. 6 Cells were resuspended in Nucleofector Solution V (Lonza). Nuclear transfection was performed by transferring 100 µL of cells into the RNP complex and immediately mixing by gentle up-and-down aspiration without introducing air bubbles, followed by transfer to cuvettes for electroporation using Amaxa program T-016. Cells were immediately transferred to 6-well plates containing 2 mL of preheated medium per well and grown at 37°C and 5% CO2 for 3 days, after which they were harvested for genomic modification assays.
[0178] Genomic DNA extracts from transfected cells were prepared using QuickExtract solution. PCR amplification of the targeted genomic region was performed using the JumpStart™ Taq ReadyMix™ for Quantitative PCR Kit (MilliporeSigma) with NGS primers under the following cycling conditions: 98°C / 2m; 98°C / 15s, 62°C / 30s, and 72°C / 45s for 34 cycles; 72°C / 5m. The NGS primer sequences are listed in Table 3. NGS library preparation, sequencing, and data analysis were performed as described in Example 1. Results are presented in… Figure 4 The results showed that the four triple mutant proteins identified, with balanced specificity and activity, had substantially higher editing efficiency than eSpCas9 1.1 and were comparable to WTCas9 across all five genomic target sites.
[0179] Table 3. sgRNA guide sequence and NGS primers
[0180] . <110> Sigma-Aldrich LLC <120> High-fidelity SpCas9 nuclease for genome modification <130> P20-035 WO-PCT <140> <141> <150> 62 / 988,279 <151> 2020-03-11 <160> 83 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> Streptococcus pyogenes <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915,920,925 Lys Tyr Asp Ser Arg Met Asn Thr Lys Tyr Asp 930,935,940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 2 <211> 7 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 2 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 3 <211> 7 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 3 Pro Lys Lys Lys Arg Arg Val 1 5 <210> 4 <211> 16 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 4 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 5 <211> 11 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 5 Tyr Gly Arg Lys Lys Arg Arg Gln Arg Arg Arg 1 5 10 <210> 6 <211> 9 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 6 Arg Lys Lys Arg Arg Gln Arg Arg Arg 1 5 <210> 7 <211> 9 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 7 Pro Ala Ala Lys Arg Val Lys Leu Asp 1 5 <210> 8 <211> 11 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 8 Arg Gln Arg Arg Asn Glu Leu Lys Arg Ser Pro 1 5 10 <210> 9 <211> 8 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 9 Val Ser Arg Lys Arg Pro Arg Pro 1 5 <210> 10 <211> 8 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 10 Pro Pro Lys Lys Ala Arg Glu Asp 1 5 <210> 11 <211> 8 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 11 Pro Gln Pro Lys Lys Lys Pro Leu 1 5 <210> 12 <211> 12 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 12 Ser Ala Leu Ile Lys Lys Lys Lys Lys Lys Met Ala Pro 1 5 10 <210> 13 <211> 7 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 13 Pro Lys Gln Lys Lys Arg Lys 1 5 <210> 14 <211> 10 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 14 Arg Lys Leu Lys Lys Lys Ile Lys Lys Leu 1 5 10 <210> 15 <211> 10 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 15 Arg Glu Lys Lys Lys Phe Leu Lys Arg Arg 1 5 10 <210> 16 <211> 20 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 16 Lys Arg Lys Gly Asp Glu Val Asp Gly Val Asp Glu Val Ala Lys Lys 1 5 10 15 Lys Ser Lys Lys 20 <210> 17 <211> 17 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 17 Arg Lys Cys Leu Gln Ala Gly Met Asn Leu Glu Ala Arg Lys Thr Lys 1 5 10 15 Lys <210> 18 <211> 38 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic polypeptide" <400> 18 Asn Gln Ser Ser Asn Phe Gly Pro Met Lys Gly Gly Asn Phe Gly Gly 1 5 10 15 Arg Ser Ser Gly Pro Tyr Gly Gly Gly Gly Gln Tyr Phe Ala Lys Pro 20 25 30 Arg Asn Gln Gly Gly Tyr 35 <210> 19 <211> 42 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic polypeptide" <400> 19 Arg Met Arg Ile Glx Phe Lys Asn Lys Gly Lys Asp Thr Ala Glu Leu 1 5 10 15 Arg Arg Arg Arg Val Glu Val Ser Val Glu Leu Arg Lys Ala Lys Lys 20 25 30 Asp Glu Gln Ile Leu Lys Arg Arg Asn Val 35 40 <210> 20 <211> 20 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 20 Gly Arg Lys Lys Arg Arg Gln Arg Arg Arg Pro Pro Gln Pro Lys Lys 1 5 10 15 Lys Arg Lys Val 20 <210> twenty one <211> 19 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> twenty one Pro Leu Ser Ser Ile Phe Ser Arg Ile Gly Asp Pro Pro Lys Lys Lys 1 5 10 15 Arg Lys Val <210> twenty two <211> twenty four <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> twenty two Gly Ala Leu Phe Leu Gly Trp Leu Gly Ala Ala Gly Ser Thr Met Gly 1 5 10 15 Ala Pro Lys Lys Lys Arg Lys Val 20 <210> twenty three <211> 27 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> twenty three Gly Ala Leu Phe Leu Gly Phe Leu Gly Ala Ala Gly Ser Thr Met Gly 1 5 10 15 Ala Trp Ser Gln Pro Lys Lys Lys Arg Lys Val 20 25 <210> twenty four <211> twenty one <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> twenty four Lys Glu Thr Trp Trp Glu Thr Trp Trp Thr Glu Trp Ser Gln Pro Lys 1 5 10 15 Lys Lys Arg Lys Val 20 <210> 25 <211> 11 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 25 Tyr Ala Arg Ala Ala Ala Arg Gln Ala Arg Ala 1 5 10 <210> 26 <211> 11 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 26 Thr His Arg Leu Pro Arg Arg Arg Arg Arg Arg 1 5 10 <210> 27 <211> 11 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 27 Gly Gly Arg Arg Ala Arg Arg Arg Arg Arg Arg 1 5 10 <210> 28 <211> 12 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 28 Arg Arg Gln Arg Arg Thr Ser Lys Leu Met Lys Arg 1 5 10 <210> 29 <211> 27 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 29 Gly Trp Thr Leu Asn Ser Ala Gly Tyr Leu Leu Gly Lys Ile Asn Leu 1 5 10 15 Lys Ala Leu Ala Ala Leu Ala Lys Lys Ile Leu 20 25 <210> 30 <211> 33 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic polypeptide" <400> 30 Lys Ala Leu Ala Trp Glu Ala Lys Leu Ala Lys Ala Leu Ala Lys Ala 1 5 10 15 Leu Ala Lys His Leu Ala Lys Ala Leu Ala Lys Ala Leu Lys Cys Glu 20 25 30 Ala <210> 31 <211> 16 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 31 Arg Gln Ile Lys Ile Trp Phe Gln Asn Arg Arg Met Lys Trp Lys Lys 1 5 10 15 <210> 32 <211> 6 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic 6xHis tag" <400> 32 His His His His His His 1 5 <210> 33 <211> 6 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 33 Leu Glu Gly Gly Gly Ser 1 5 <210> 34 <211> 4 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 34 Thr Gly Ser Gly 1 <210> 35 <211> 8 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 35 Gly Gly Ser Gly Gly Gly Ser Gly 1 5 <210> 36 <211> 20 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <220> <221> site <222> (1)..(20) <223> / Note: This sequence can contain 1-4 repeating units of 'Gly Gly Gly Gly Ser'. <400> 36 Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly 1 5 10 15 Gly Gly Gly Ser 20 <210> 37 <211> 8 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <220> <221> site <222> (1)..(8) <223> / Note: This sequence can contain 6-8 residues.* <400> 37 Gly Gly Gly Gly Gly Gly Gly Gly 1 5 <210> 38 <211> 20 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <220> <221> site <222> (1)..(20) <223> / Note: This region can contain 1-4 'Glu Ala Ala Ala Lys' repeating units. <400> 38 Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Glu 1 5 10 15 Ala Ala Ala Lys 20 <210> 39 <211> 27 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <220> <221> site <222> (2)..(26) <223> / Note: This sequence may contain 2-5 'Glu Ala Ala Ala Lys' repeating units. <400> 39 Ala Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys 1 5 10 15 Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Ala 20 25 <210> 40 <211> 5 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <400> 40 Pro Ala Pro Ala Pro 1 5 <210> 41 <211> 16 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic peptide" <220> <221> site <222> (1)..(16) <223> / Note: This sequence can contain 6-8 'Ala Pro' repeating units. / <400> 41 Ala Pro Ala Pro Ala Pro Ala Pro Ala Pro Ala Pro Ala Pro Ala Pro 1 5 10 15 <210> 42 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 42 ggcacugcgg cuggaggugg 20 <210> 43 <211> 59 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 43 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnng gaacccaggt agccagaga 59 <210> 44 <211> 57 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 44 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn ggggtggggt cagacgt 57 <210> 45 <211> 63 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 45 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnc tagagcaaac cttggcattg 60 tcc 63 <210> 46 <211> 60 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 46 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn accctctacc ctccctgatg 60 <210> 47 <211> 1404 <212> PRT <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic polypeptide" <400> 47 Pro Ala Ala Lys Arg Val Lys Leu Asp Gly Gly Gly Gly Ser Thr Gly 1 5 10 15 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 20 25 30 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 35 40 45 Lys Val Leu Gly Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 50 55 60 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 65 70 75 80 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr 85 90 95 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 100 105 110 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 115 120 125 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 130 135 140 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 145 150 155 160 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 165 170 175 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 180 185 190 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 195 200 205 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 210 215 220 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 225 230 235 240 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 245 250 255 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 260 265 270 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 275 280 285 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 290 295 300 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 305 310 315 320 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 325 330 335 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 340 345 350 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 355 360 365 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 370 375 380 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 385 390 395 400 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 405 410 415 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 420 425 430 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 435 440 445 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 450 455 460 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 465 470 475 480 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 485 490 495 Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 500 505 510 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 515 520 525 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 530 535 540 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 545 550 555 560 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 565 570 575 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser 580 585 590 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 595 600 605 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 610 615 620 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 625 630 635 640 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 645 650 655 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 660 665 670 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 675 680 685 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 690 695 700 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 705 710 715 720 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 725 730 735 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 740 745 750 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 755 760 765 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 770 775 780 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 785 790 795 800 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 805 810 815 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 820 825 830 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 835 840 845 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 850 855 860 Asp Ser Ile Asp Asn Gln Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 865 870 875 880 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 885 890 895 Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 900 905 910 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys 915 920 925 Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys 930,935,940 His Gln Ile Leu Asp Ser Arg With Asn Thr Lys Tyr Asp Glu 945 950 955 960 Asn Asp Lys With Arg Glu Val Val Lys With Thr Lys Ser Ser Lys 965,970,975 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 980,985,990 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val Val 995 1000 1005 Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 1010 1015 1020 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1025 1030 1035 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1040 1045 1050 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1055 1060 1065 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1070 1075 1080 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1085 1090 1095 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1100 1105 1110 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1115 1120 1125 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1130 1135 1140 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1145 1150 1155 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1160 1165 1170 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1175 1180 1185 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1190 1195 1200 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1205 1210 1215 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1220 1225 1230 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1235 1240 1245 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1250 1255 1260 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1265 1270 1275 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1280 1285 1290 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1295 1300 1305 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1310 1315 1320 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1325 1330 1335 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1340 1345 1350 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1355 1360 1365 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1370 1375 1380 Glu Phe Pro Lys Lys Lys Arg Lys Val Gly Gly Gly Gly Ser Pro 1385 1390 1395 Lys Lys Lys Arg Lys Val 1400 <210> 48 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 48 gcugcagaag ggauuccaug 20 <210> 49 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 49 cacguucacc uugccccaca 20 <210> 50 <211> 59 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 50 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnna atggggccat gccgaccaa 59 <210> 51 <211> 63 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 51 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn agttgcccag agtcaaggaa 60 cac 63 <210> 52 <211> 60 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 52 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnt ctttccctca ctctggctcg 60 <210> 53 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 53 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn tggaatgaat ggggtgggag 60 g 61 <210> 54 <211> 63 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 54 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnt aggcagagag agtcagtgcc 60 tat 63 <210> 55 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 55 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn ccaatctact cccaggagca 60 g 61 <210> 56 <211> 59 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 56 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnt cactggagca gggaggaca 59 <210> 57 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 57 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn gggtaggaaa acagcccaag 60 g 61 <210> 58 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 58 cucccuccca ggauccucuc 20 <210> 59 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 59 gccaguagcc agccccgucc 20 <210> 60 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 60 caggcagucu ucauccccgu 20 <210> 61 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 61 gaagcgugau gacaaagagg 20 <210> 62 <211> 20 <212> RNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: Synthetic oligonucleotide" <400> 62 auucugguca acguguccuu 20 <210> 63 <211> 62 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 63 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnc ttgggaagtg taaggaagct 60 gc 62 <210> 64 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 64 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn gcctcctcct tcctagtctc 60 c 61 <210> 65 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 65 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnng ctgcagcttc cttacacttc 60 c 61 <210> 66 <211> 64 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 66 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn gaggaatatg tcccagatag 60 cact 64 <210> 67 <211> 64 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 67 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnc tgtgggattt catggaagtt 60 cagc 64 <210> 68 <211> 59 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 68 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn atcctggctg gcaaggtgg 59 <210> 69 <211> 60 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 69 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnnt tccttggcct ctgactgttg 60 <210> 70 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 70 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn ttcctgccca ccatctactc 60 c 61 <210> 71 <211> 61 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (34) (39) <223> a, c, t, g, unknown or other <400> 71 tcgtcggcag cgtcagatgt gtataagaga cagnnnnnng atgggcctca gtaccacatt 60 g 61 <210> 72 <211> 62 <212> DNA <213> Artificial sequence <220> <221> source <223> / Note="Description of artificial sequence: synthetic primers" <220> <221> Modified bases <222> (35) (40) <223> a, c, t, g, unknown or other <400> 72 gtctcgtggg ctcggagatg tgtataagag acagnnnnnn caacctttgc cttcccctaa 60 cc 62 <210> 73 <211> twenty three <212> DNA <213> Homo sapiens <400> 73 ggcactgcgg ctggaggtgg ggg 23 <210> 74 <211> twenty three <212> DNA <213> Homo sapiens <400> 74 ggcacgacgg ctggaggtgg ggg 23 <210> 75 <211> twenty three <212> DNA <213> Homo sapiens <400> 75 gctgcagaag ggattccatg agg 23 <210> 76 <211> 23 <212> DNA <213> Homo sapiens <400> 76 gctgcagaag ggattccaag ggg 23 <210> 77 <211> 23 <212> DNA <213> Homo sapiens <400> 77 cacgttcacc ttgccccaca ggg 23 <210> 78 <211> 23 <212> DNA <213> Homo sapiens <400> 78 cacgttcact ttgccccaca ggg 23 <210> 79 <211> 23 <212> DNA <213> Homo sapiens <400> 79 ctccctccca ggatcctctc tgg 23 <210> 80 <211> 23 <212> DNA <213> Homo sapiens <400> 80 gccagtagcc agccccgtcc tgg 23 <210> 81 <211> 23 <212> DNA <213> Homo sapiens <400> 81 caggcagtct tcatccccgt agg 23 <210> 82 <211> 23 <212> DNA <213> Homo sapiens <400> 82 gaagcgtgat gacaaagagg agg 23 <210> 83 <211> 23 <212> DNA <213> Homo sapiens <400> 83 attctggtca acgtgtcctt cgg 23
Claims
1. A modified Streptococcus pyogenes Cas9 (SpCas9) protein variant, wherein the variant is represented by an amino acid sequence differing from SEQ ID NO. 1 by the following amino acid mutations: a K855Q mutation and two other mutations, said two other mutations being two mutations selected from two different amino acid positions of K526L / Q, K562L / Q, K652L / Q, K810L / Q, K848L / Q, R661L / Q, R691L / Q, R780L / Q, K1003L / Q, and R1060L / Q, referring to the SpCas9 numbering system of SEQ ID NO. 1, wherein said mutations are selected from the group consisting of: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; K652Q-R661L-K855Q; R661L-K855Q-K1003Q; R661L-K855Q-R1060Q; K526L-R661L-K855Q; R661L-R691L-K855Q; R661L-R780L-K855Q; R661L-R780Q-K855Q; R661L-K810L-K855Q; R661L-K848L-K855Q; K526Q-R661L-K855Q; R661L-K810Q-K855Q; R661L-K855Q-K1003L; and R661L-K855Q-R1060L.
2. The modified SpCas9 protein variant of claim 1, wherein the mutation is selected from the group consisting of: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; K652Q-R661L-K855Q; R661L-K855Q-K1003Q; and R661L-K855Q-R1060Q.
3. The modified SpCas9 protein variant of claim 1, wherein the mutation is selected from the group consisting of: K562L-R661L-K855Q; K562Q-R661L-K855Q; K652L-R661L-K855Q; and K652Q-R661L-K855Q.
4. The modified SpCas9 protein variant of claim 1, wherein the mutation is selected from the group consisting of: K526L-R661L-K855Q; R661L-R691L-K855Q; R661L-R780L-K855Q; R661L-R780Q-K855Q; R661L-K810L-K855Q, and R661L-K848L-K855Q.
5. The modified SpCas9 protein variant of claim 1, wherein the mutation is selected from the group consisting of: K526Q-R661L-K855Q; R661L-K810Q-K855Q; R661L-K855Q-K1003L; and R661L-K855Q-R1060L.
6. A modified SpCas9 protein variant of any one of claims 1-5, further comprising one or more heterologous domains fused to the N-terminus, C-terminus, internal location, or a combination thereof.
7. The modified SpCas9 protein variant of claim 6, wherein the heterologous domain is selected from nuclear localization signal, cell penetration domain, marker or reporter domain that facilitates detection, chromatin modification domain, epigenetic modification domain, transcriptional regulation domain, DNA or RNA deaminase domain, uracil-DNA glycosylase domain, reverse transcriptase domain, recombinase domain, RNA aptamer binding domain, and non-Cas9 nuclease domain.
8. The modified SpCas9 protein variant of any one of claims 1-5, further comprising at least one nuclear localization signal fused to the N-terminus, C-terminus, internal location, or a combination thereof.
9. A modified SpCas9 protein variant of any one of claims 1-5, further comprising at least one mutation in the RuvC domain and / or at least one mutation in the HNH domain.
10. The modified SpCas9 protein variant of claim 9, wherein at least one mutation in the RuvC domain comprises at least one mutation selected from D10A, D8A, E762A, and D986A.
11. The modified SpCas9 protein variant of claim 9, wherein at least one mutation in the HNH domain comprises at least one mutation selected from H840A, H559A, N854A, N856A, and N863A.
12. A modified Cas9 system comprising a modified SpCas9 protein variant of any one of claims 1-11 and at least one modified guide RNA, wherein the at least one modified guide RNA is designed to be complexed with the modified SpCas9 protein variant.
13. A variety of nucleic acids encoding a modified SpCas9 protein variant of any one of claims 1-11.
14. Multiple nucleic acids encoding the modified SpCas9 system of claim 12.
15. The plurality of nucleic acids of claim 14, said plurality of nucleic acids comprising at least one nucleic acid encoding the modified SpCas9 protein variant and at least one nucleic acid encoding the modified guide RNA.
16. The plurality of nucleic acids according to any one of claims 13-15, wherein at least one nucleic acid is RNA.
17. The plurality of nucleic acids according to any one of claims 13-15, wherein at least one nucleic acid is DNA.
18. A plurality of nucleic acids according to any one of claims 13-15, wherein at least one nucleic acid encoding a modified SpCas9 protein variant is codon-optimized for expression in eukaryotic cells.
19. The various nucleic acids of claim 18, wherein the eukaryotic cell is a human cell, a non-human mammalian cell, a non-mammal vertebrate cell, an invertebrate cell, a plant cell, or a single-celled eukaryotic organism.
20. The plurality of nucleic acids of claim 12, wherein at least one nucleic acid encoding the modified guide RNA is DNA.
21. The plurality of nucleic acids of claim 12, wherein at least one nucleic acid encoding a modified SpCas9 protein variant is operatively linked to a phage promoter sequence for in vitro RNA synthesis or protein expression in bacterial cells, and at least one nucleic acid encoding a modified guide RNA is operatively linked to a phage promoter sequence for in vitro RNA synthesis.
22. The plurality of nucleic acids of claim 8, wherein at least one nucleic acid encoding a modified Cas9 protein variant is operatively linked to a eukaryotic promoter sequence for expression in eukaryotic cells, and at least one nucleic acid encoding a modified guide RNA is operatively linked to a eukaryotic promoter sequence for expression in eukaryotic cells.
23. At least one vector comprising a plurality of nucleic acids according to any one of claims 13-22.
24. At least one vector of claim 23, which is a plasmid vector, a viral vector or a self-replicating viral RNA replicon.
25. A eukaryotic cell comprising a modified Cas9 system of claim 12 or a plurality of nucleic acids of claim 14 or 15, wherein the cell is a human cell, a non-human mammalian cell, a non-mammal vertebrate cell, an invertebrate cell, or a single-celled eukaryote.
26. The eukaryotic cell of claim 25, wherein it is a human cell, a non-human mammalian cell, a non-mammal vertebrate cell, or an invertebrate cell.
27. The eukaryotic cell of claim 26, whether in vivo or in vitro.
28. A modified SpCas9 protein variant of any one of claims 1-5, which is a Cas9 homolog.
29. A ribonucleoprotein (RNP) complex comprising a modified SpCas9 protein variant of any one of claims 1-9.
30. A fusion protein comprising a modified SpCas9 protein variant of any one of claims 1-9.
31. A pharmaceutical composition comprising a modified SpCas9 protein variant of any one of claims 1-9 and at least one pharmaceutically acceptable excipient.
32. A method for modifying the chromosome sequence of a eukaryotic cell, the method comprising expressing, in the eukaryotic cell, a modified SpCas9 protein variant of any one of claims 1-9 together with guide RNA.
Citation Information
Patent Citations
Methods for making and using reprogrammed human somatic cell nuclei and autologous and isogenic human stem cells
WO2003046141A2
CRISPR enzyme mutations reducing off-target effects
CN108290933A
Programmable cas9-recombinase fusion proteins and uses thereof
CN109804066A