Improved nuclease compositions and methods
By introducing engineered KFERQ motifs or KFERQ-like motifs and amino acid modifications, the degradation rate and target specificity of recombinant Cas9 proteins are improved, and off-target modification and cytotoxicity problems in the CRISPR/Cas9 system are solved, achieving safer and more efficient gene editing effects.
Patent Information
- Application Number
- CN201980058359.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-07
- Filing Date
- 2019-09-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-09-06
AI Technical Summary
The existing CRISPR/Cas9 system has off-target modification problems in clinical applications, resulting in unpredictable and undesirable results. At the same time, the cytotoxicity of Cas9 has also attracted attention.
The recombinant Cas9 protein with an engineered KFERQ motif or KFERQ-like motif is enhanced, and the chaperone-mediated autophagy target motif is introduced through amino acid modification to reduce off-target modification.
Reduced off-target modifications were achieved, reducing the cytotoxicity of Cas9, while maintaining target-on-target efficiency similar to wild-type Cas9.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present disclosure provides a recombinant Cas9 protein with a faster degradation rate than wild-type Cas9. The present disclosure also provides a recombinant Cas9 protein and a CRISPR-Cas system with reduced off-target modifications. Also provided herein is a method for site-specific modification with reduced off-target modifications using the recombinant Cas9 protein of the present disclosure. Background Art
[0002] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) systems are prokaryotic immune systems first discovered by Ishino in Escherichia coli (Ishino et al., Journal of Bacteriology 169(12):5429-5433 (1987)). The prokaryotic immune system provides immunity to viruses and plasmids by targeting their nucleic acids in a sequence-specific manner. See also Soret et al., Nature Reviews Microbiology 6(3):181-186 (2008). The CRISPR immune response involves two main phases: the first is acquisition and the second is interference. The acquisition phase involves cutting the genomes of invading viruses and plasmids and integrating segments of the invading viral and plasmid genomes into the CRISPR locus of the organism. The segments integrated into the CRISPR locus of the organism are called protospacer sequences and help protect the organism from subsequent attacks by the same virus or plasmid. The second phase involves attacking the invading virus or plasmid. In the second stage, the protospacer sequence is transcribed into RNA, which, after some processing, hybridizes with a complementary sequence in the DNA of the invading virus or plasmid and also associates with a protein or protein complex that efficiently cleaves the DNA.
[0003] Depending on the bacterial species, the process of CRISPR RNA processing is different. For example, in the type II system originally described in the bacterium Streptococcus pyogenes, the transcribed RNA is paired with the trans-activating RNA (tracrRNA), and then cleaved by RNase III to form a single CRISPR-RNA (crRNA). After being bound by the Cas9 nuclease, the crRNA is further processed to produce mature crRNA. The crRNA / Cas9 complex then binds to the DNA containing a sequence complementary to the capture region (called the protospacer sequence). Then, the Cas9 protein cleaves the two chains of DNA in a site-specific manner to form a double-strand break (DSB). This provides a DNA-based "memory", which causes the virus or plasmid DNA to degrade rapidly after repeated exposure and / or infection. There has been a comprehensive review of the native CRISPR system (see, for example, Barrangou et al., Cell [Cell] 54 (2): 234-244 (2014)).
[0004] Since its initial discovery, multiple research groups have conducted extensive research on the potential applications of CRISPR systems in genetic engineering, including gene editing (Jinek et al., Science 337(6096):816-821 (2012); Cong et al., Science 339(6121):819-823 (2013); and Mali et al., Science 339(6121):823-826 (2013)). The CRISPR-Cas9 gene editing system has been successfully used in a wide range of organisms and cell lines. In addition to genome editing, the CRISPR system has many other applications, including regulating gene expression, gene circuit construction, and functional genomics (reviewed in Sander et al., Nature Biotechnology 32:347-355 (2014)).
[0005] The suitability of CRISPR / Cas9 for therapeutic applications is a topic of intense concern. However, off-target modifications of the target genome (i.e., double-stranded DNA breaks at the locus of the unintended target sequence) may lead to unpredictable and undesirable results, causing concern about the use of CRISPR systems in clinical applications. See, for example, Hsu et al., Nature Biotechnology [Natural Biotechnology] 31 (9): 827-834 (2013); Hsu et al., Cell [Cell] 157 (6): 1262-1278 (2014); and Schaefer et al., Nature Methods [Natural Methods] 14 (6): 547-548 (2017).
[0006] The cytotoxicity of the CRISPR / Cas9 system has also received attention. Studies have shown that when the efficiency of the Cas9 nuclease is increased, a tp53-dependent toxic response is triggered in cells (Ihry et al., bioRxiv (2017), doi: 10.1101 / 168443).
[0007] Efforts have been made to reduce Cas9 off-target modifications. Fu et al. described methods using truncated guide RNAs with short target complementary regions to reduce the off-target effects of Cas9 by reducing the length of the guide RNA-target DNA interface (Nature Biotechnology 32(3):279-284(2014)). Kleinstiver et al. described engineered Cas9 variants with reduced contact with the target DNA sequence to minimize off-target binding (Nature 529(7587):490-495(2016)). However, despite the reduction in Cas9 off-target activity, two studies also showed a corresponding reduction in on-target nuclease efficiency.
[0008] Therefore, there remains a need in the art for improved CRISPR / Cas9 systems with reduced off-target activity that maintain on-target efficiency. Summary of the invention
[0009] In some embodiments, the present disclosure provides a recombinant Cas9 protein having a faster degradation rate than wild-type Cas9. In some embodiments, the present disclosure also provides a recombinant Cas9 protein and CRISPR-Cas system with reduced off-target modifications. In some embodiments, the present disclosure provides a method for site-specific modification with reduced off-target modifications using the recombinant Cas9 protein described herein.
[0010] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif.
[0011] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is selected from KFERQ (SEQ ID NO:24), RKVEQ (SEQ ID NO:25), QDLKF (SEQ ID NO:26), QRFFE (SEQ ID NO:27), NRVVD (SEQ ID NO:28), QRDKV (SEQ ID NO:29), QKILD (SEQ ID NO:30), QKKEL (SEQ ID NO:31), QFREL (SEQ ID NO:32), IKLDQ (SEQ ID NO:33), DVVRQ (SEQ ID NO:34), QRIVE (SEQ ID NO:35), VKELQ (SEQ ID NO:36), QKVFD (SEQ ID NO:37), QELLR (SEQ ID NO:38), VDKLN (SEQ ID NO:39), RIKEN (SEQ ID NO:40), NKKFE (SEQ ID NO:41), and combinations thereof. In some embodiments, the engineered KFERQ-like motif is VDKLN (SEQ ID NO: 39).
[0012] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the REC lobe of the recombinant Cas9 protein. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the Rec2 domain of the REC lobe. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the HNH domain, RuvC domain, or PI domain of the recombinant Cas9 protein.
[0013] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in a surface exposed region of a recombinant Cas9 protein. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is at the N-terminus or C-terminus of a recombinant Cas9 protein.
[0014] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the modifications introduce a chaperone-mediated autophagy (CMA) target motif or an endosomal microautophagy (eMI) target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the recombinant Cas9 protein is degraded in vivo at least 50% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the recombinant Cas9 protein is degraded in vivo at least 80% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif.
[0015] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the recombinant Cas9 protein comprises a CMA target motif or an eMI target motif.
[0016] In some embodiments, the CMA target motif or eMI target motif is selected from KFERQ (SEQ ID NO:24), RKVEQ (SEQ ID NO:25), QDLKF (SEQ ID NO:26), QRFFE (SEQ ID NO:27), NRVVD (SEQ ID NO:28), QRDKV (SEQ ID NO:29), QKILD (SEQ ID NO:30), QKKEL (SEQ ID NO:31), QFREL (SEQ ID NO:32), IKLDQ (SEQ ID NO:33), DVVRQ (SEQ ID NO:34), QRIVE (SEQ ID NO:35), VKELQ (SEQ ID NO:36), QKVFD (SEQ ID NO:37), QELLR (SEQ ID NO:38), VDKLN (SEQ ID NO:39), RIKEN (SEQ ID NO:40), NKKFE (SEQ ID NO:41), and combinations thereof. In some embodiments, the CMA target motif or eMI target motif is VDKLN (SEQ ID NO: 39). In some embodiments, the one or more amino acid substitutions are in a surface exposed region of the recombinant Cas9 protein.
[0017] In some embodiments, the present disclosure provides a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, which recombinant Cas9 protein comprises an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO:1.
[0018] In some embodiments, the amino acid modification comprises one or more of the following mutations: F185N; A547E / I548L; T560E / V561Q; D829L / I830R; L1087E / S1088Q; or P1199D / K1200Q. In some embodiments, the amino acid modification is a mutation at F185. In some embodiments, the mutation is F185N. In some embodiments, the amino acid modification results in a CMA target motif or an eMI target motif.
[0019] In some embodiments, the recombinant Cas9 protein of the present disclosure is at least 90% identical to SEQ ID NO:1.
[0020] In some embodiments, the present disclosure provides a recombinant Cas9 protein capable of binding to heat shock cognate protein of 70 kD (HSC70).
[0021] In some embodiments, the present disclosure provides a recombinant protein (SpCas9) isolated from Streptococcus pyogenes, comprising an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO: 1. In some embodiments, the KFERQ-like motif is VDKLN (SEQ ID NO: 39).
[0022] In some embodiments, the recombinant Cas9 protein of the present disclosure further comprises a mutation at position D10, H840, or a combination thereof in SEQ ID NO: 1. In some embodiments, the mutation is selected from D10A or D10N; H840A, H840N, or H840Y; and combinations thereof. In some embodiments, the recombinant Cas9 protein of the present disclosure produces a sticky end.
[0023] In some embodiments, the recombinant Cas9 protein of the present disclosure further comprises one or more nuclear localization signals.
[0024] In some embodiments, the disclosure provides polynucleotide sequences encoding the recombinant Cas9 of the disclosure. In some embodiments, the polynucleotide sequences are codon optimized for expression in eukaryotic cells.
[0025] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, comprising: a recombinant Cas9 protein of the present disclosure; and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0026] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, comprising: a polynucleotide sequence encoding a recombinant Cas9 protein of the present disclosure; and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0027] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, comprising: a regulatory element operably linked to a polynucleotide sequence encoding a recombinant Cas9 protein of the present disclosure; and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and includes a guide sequence.
[0028] In some embodiments of the CRISPR-Cas system, the guide sequence is linked to the direct repeat sequence.
[0029] In some embodiments of the CRISPR-Cas system, the guide polynucleotide comprises a tracrRNA sequence. In some embodiments, the CRISPR-Cas system comprises a separate polynucleotide comprising a tracrRNA sequence.
[0030] In some embodiments of the CRISPR-Cas system, the polynucleotide sequence encoding the recombinant Cas9 protein and the guide polynucleotide are on a single vector. In some embodiments of the CRISPR-Cas system, the polynucleotide sequence encoding the recombinant Cas9 protein, the guide polynucleotide and the tracrRNA sequence are on a single vector.
[0031] In some embodiments, the delivery particle comprises a CRISPR-Cas system of the present disclosure. In some embodiments, the vesicle comprises a CRISPR-Cas system of the present disclosure. In some embodiments, the vesicle is an exosome or a liposome.
[0032] In some embodiments, the viral vector comprises a CRISPR-Cas system of the present disclosure. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus vector.
[0033] In some embodiments, the present disclosure provides a method of providing a site-specific modification at a target sequence in the genome of a cell, the method comprising introducing a CRISPR-Cas system of the present disclosure into the cell.
[0034] In some embodiments of the method, the modification comprises a deletion of at least a portion of the target sequence. In some embodiments of the method, the modification comprises a mutation of the target sequence. In some embodiments of the method, the modification comprises an insertion of a sequence of interest (SoI) at the target sequence.
[0035] In some embodiments of the method, the off-target modification in the genome of the cell is less than about 5% of the modification in the genome produced by the recombinant Cas9. In some embodiments of the method, the off-target modification in the genome of the cell is less than about 2% of the modification in the genome produced by the recombinant Cas9. In some embodiments of the method, the off-target modification in the genome of the cell is less than about 1% of the modification in the genome produced by the recombinant Cas9. In some embodiments of the method, the off-target modification in the genome of the cell is reduced by at least about 50% relative to wild-type CRISPR-Cas9 or Cas9 that does not include a KFERQ motif or a KFERQ-like motif.
[0036] In some embodiments of the method, the cell is a bacterial cell, a mammalian cell, or a plant cell. In some embodiments of the method, the cell is a human cell. In some embodiments of the method, the cell is a pluripotent stem cell. In some embodiments of the method, the cell is an induced pluripotent stem cell.
[0037] In some embodiments of the method, the guide sequence of the guide polynucleotide is capable of hybridizing to a target sequence in the cell genome. In some embodiments of the method, the CRISPR-Cas system is introduced into the cell via a delivery particle, vesicle, or viral vector. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figures 1 and 2 relate to the experiments described in Example 4.
[0039] Figure 1A showed that after stable transfection of Cas9 into human urothelial cells (SVHUC-1), cell numbers were reduced. Figure 1B showed that the body weight of mice expressing Cas9 decreased after induction of Cas9 expression.
[0040] Figure 2 The left panel is a microscope image of induced pluripotent stem cells (iPSCs). Figure 2 The right panel is a microscope image of iPSCs expressing Cas9.
[0041] Figure 3 Schematic diagram showing the Cas9 protein from Streptococcus pyogenes (SpCas9).
[0042] Figure 4 The crystal structure of SpCas9 bound to guide RNA (Sg RNA) and DNA is shown.
[0043] 5 to 12 relate to the experiments described in Example 1.
[0044] Figure 5A Schematic diagram showing plasmids containing wild-type Cas9 and Cas9 including a KFERQ motif. Figure 5B Schematic diagram showing plasmids containing wild-type Cas9 and FaDe-Cas9 tagged with FLAG tags, respectively.
[0045] Fig. 6A and 6B Western blots (immunoblots) are shown to detect the presence of Cas9 or FaDe-Cas9. Fig. 6A In the Cas9-specific antibody, wild-type Cas9 was detected, but KFERQ-Cas9 was not detected. Figure 6B In the assay, FLAG-tagged wild-type Cas9 could be detected by antibodies specific for FLAG and Cas9, but none of the antibodies could detect FLAG-tagged KFERQ-Cas9.
[0046] Fig. 7A and 7B The expression of Cas9 or FaDe-Cas9 over time is shown. Fig. 7A Western blot in shows that wild-type Cas9 levels increased over time, but FaDe-Cas9 was not detected by Cas9-specific antibodies. Low exp: low exposure; high exp: high exposure; ctr: control (no Cas9). Figure 7B It was shown that the mRNA transcript levels of Cas9 and FaDe-Cas9 were comparable at the same time points.
[0047] Figure 8 A schematic diagram showing a dual reporter vector comprising one promoter for the expression of Cas9 fused to GFP and a second promoter for the expression of mCherry.
[0048] Fig. 9A and 9B are fluorescence microscopy images showing cells expressing Cas9-GFP and FaDe-Cas9-GFP, respectively.
[0049] Fig. 10A Shown are the transfection efficiencies of Cas9 and FaDe-Cas9 measured by mCherry fluorescence. Fig. 10B The expression levels of Cas9 and FaDe-Cas9 measured by GFP fluorescence are shown.
[0050] Fig.11A and 11B Fluorescence microscopy images of Cas9-GFP and mCherry are shown, respectively. Fig. 11Cshow Fig.11A and 11B A merge of GFP and mCherry was shown in Figure 2, which shows that the same cells expressing GFP also express mCherry.
[0051] Fig.12 Shown are western blots indicating the expression levels of Cas9 and FaDe-Cas9 in different cell types over time as detected by Cas9-specific antibodies.
[0052] 13 to 16 relate to the experiments described in Example 2.
[0053] Fig.13A Blot showing co-immunoprecipitation of Cas9 or FaDe-Cas9 with HSC70 detected by HSC70-specific antibody. Fig. 13B Immunofluorescence images of Cas9 or FaDe-Cas9 and Lamp-2A are shown.
[0054] Fig.14 Schematic showing two plasmids: the first plasmid expresses Lamp-2A fused to dsRed, and the second plasmid expresses FaDe-Cas9 fused to GFP.
[0055] Fig.15 Fluorescence microscopy images showing Lamp-2A-dsRed, FaDe-Cas9-GFP, and a merged image showing co-localization of Lamp-2A and FaDe-Cas9.
[0056] Fig.16 Western blots are shown, which indicate the localization of Cas9 and FaDe-Cas9 in the cytosol or nucleus.
[0057] 17 to 19 relate to the experiments described in Example 3.
[0058] Fig.17A Shown are the results of the Surveyor nuclease assay (Cellular Assay) testing Cas9 and FaDe-Cas9 nuclease activities in HEK cells. Fig. 17B Next generation sequencing results demonstrating the nuclease efficiency of Cas9 and FaDe-Cas9 are shown.
[0059] Fig.18 Assay results testing Cas9 and FaDe-Cas9 nuclease activity in hiPScs are shown. RNP: ribonucleoprotein; pl: plasmid.
[0060] Fig.19The results of the analysis of off-target modifications by Cas9 and FaDe-Cas9 at the EMX and FANCF loci are shown. The left, middle, and right panels compare the on-target efficiency, off-target efficiency, and normalized on-target efficiency between Cas9 and FaDe-Cas9, respectively.
[0061] Fig. 20A Presents data showing that FaDe-Cas9 has comparable on-target efficiency and reduced off-target activity compared to Cas9. Fig. 20B It was shown that cells transfected with Cas9 had a reduced proliferation rate compared to FaDe-Cas9 and untransfected cells.
[0062] Fig.21 showed that cells edited with FaDe-Cas9 resulted in reduced chromosomal translocations compared to Cas9.
[0063] Fig.22A Schematic diagram showing experiments testing cells' tolerance to Cas9 and FaDe-Cas9. Fig. 22B and 22C showed that cells can tolerate more copies of FaDe-Cas9 than Cas9.
[0064] Fig.23A and 23B Quantification of Western blots showing Cas9 in cells transfected with Cas9 or FaDe-Cas9 at time points from 0 to 100 hours ( Fig.23A ), and close-ups of time points from 0 to 24 hours ( Fig. 23B ).
[0065] Figures 24A-24C Shown are Cas9 and FaDe-Cas9 intracellular protein levels measured by ELISA assay.
[0066] Fig.25A and 25B Shown are the results of analyzing KD lamp2a efficiency and protein accumulation by GFP signal.
[0067] Fig.26 Immunoprecipitation results are shown, which show that FaDe-Cas9 binds with high affinity to the master regulator of chaperone-mediated autophagy, HSC70.
[0068] Fig.27A and 27B Displays in vivo on-target / off-target analysis plots.
[0069] Fig.28A showed that mouse livers exposed to FaDe-Cas9 displayed higher gene editing due to its rapid degradation in liver tissue, which in turn led to increased survival of edited cells. Fig.28B and 28C We show that the rapid in vivo turnover of FaDe-Cas9 results in lower hepatotoxicity compared with Cas9, as evidenced by unchanged hepatic glycogen, barely detectable mitotic marker Ki67, less amount of infiltrates, and minimal cell necrosis. Fig.28D showed that virus-mediated expression of FaDe-Cas9 in vivo resulted in lower cytotoxic T lymphocyte immune responses. DETAILED DESCRIPTION
[0070] Described herein are components of the CRISPR-Cas system, which can be used for genome editing, genome engineering, and altering the expression of genes and / or genetic elements. The CRISPR-Cas system can be used for a variety of therapeutic applications, including the treatment of genetic diseases. Also described herein are fast-degrading variants of the Cas9 protein (sometimes referred to as "FaDe-Cas9"), which help reduce the off-target activity of the CRISPR-Cas9 system. Other advantages of rapidly degradable Cas9 proteins are described herein, including, but not limited to, on-target efficiency comparable to wild-type Cas9 and / or reduced toxicity compared to wild-type Cas9.
[0071] definition
[0072] As used herein, "a" or "an" may mean one or more. As used in the specification and one or more claims herein, when used in conjunction with the word "comprising", the word "a" may mean one or more than one. As used herein, "another" may mean at least a second or more.
[0073] Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error of the method / device employed to determine the value, or the variation between study subjects. Typically, the term is meant to cover approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20% variation, depending on the specific circumstances.
[0074] The term "or" used in the claims is intended to mean "and / or" unless explicitly indicated to refer to only alternatives or the alternatives are mutually exclusive, although the present disclosure supports a definition referring to only alternatives and "and / or."
[0075] As used in this specification and in one or more claims, the words "comprising" (and any forms of comprising, such as "comprising" and "including"), "having" (and any forms of having, such as "having" and "having"), "including" (and any forms of including, such as "including" and "including"), or "containing" (and any forms of containing, such as "containing" and "containing") are inclusive or open-ended, and do not exclude additional unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method, system, host cell, expression vector and / or composition of the present disclosure. In addition, the compositions, systems, host cells and / or vectors of the present disclosure can be used to implement the methods and proteins of the present disclosure.
[0076] The use of the term "for example" and its corresponding short form "such as" (whether italicized or not) means that the specific terms listed are representative examples and embodiments of the present disclosure, and is not intended to be limited to the specific examples cited or listed, unless expressly stated otherwise.
[0077] "Nucleic acid", "nucleic acid molecule", "nucleotide", "nucleotide sequence", "oligonucleotide" or "polynucleotide" means a polymeric compound comprising covalently linked nucleotides. The term "nucleic acid" includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA), both of which can be single-stranded or double-stranded. DNA includes, but is not limited to, complementary DNA (cDNA), genomic DNA, plasmid or vector DNA, and synthetic DNA. In some embodiments, the present disclosure provides polynucleotides encoding any of the polypeptides disclosed herein, for example, the present disclosure relates to polynucleotides encoding Cas proteins or variants thereof.
[0078] "Gene" refers to an assembly of nucleotides that encodes a polypeptide, and includes cDNA and genomic DNA nucleic acid molecules."Gene" also refers to a nucleic acid fragment that may serve as regulatory sequences preceding (5' non-coding sequences) and following (3' non-coding sequences) the coding sequence.
[0079] A nucleic acid molecule is "hybridizable" or "hybridizes" with another nucleic acid molecule (such as a cDNA, genomic DNA, or RNA) when the single-stranded form of the nucleic acid molecule can anneal to the other nucleic acid molecule under suitable conditions of temperature and solution ionic strength. Hybridization and washing conditions are known and are exemplified in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 thereof. Temperature and ionic strength conditions determine the "stringency" of the hybridization. Stringent conditions can be adjusted to screen for moderately similar fragments (such as homologous sequences from distantly related organisms) to highly similar fragments (such as genes that duplicate functional enzymes from closely related organisms). For preliminary screening of homologous nucleic acids, a T corresponding to 55°C can be used. m Low stringency hybridization conditions, for example, 5X SSC, 0.1% SDS, 0.25% milk and no formamide; or 30% formamide, 5X SSC, 0.5% SDS. Moderate stringency hybridization conditions correspond to higher T m , for example, 40% formamide and 5X or 6X SCC. High stringency hybridization conditions correspond to the highest Tm, for example, 50% formamide, 5X or 6X SCC. Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases may exist depending on the stringency of the hybridization.
[0080] The term "complementary" is used to describe the relationship between nucleotide bases that are capable of hybridizing to each other. For example, for DNA, adenosine is complementary to thymine, and cytosine is complementary to guanine. Therefore, the present disclosure also includes isolated nucleic acid fragments that are complementary to the complete sequences disclosed or used herein, as well as those substantially similar nucleic acid sequences.
[0081] DNA "coding sequence" is a double-stranded DNA sequence, when placed under the control of an appropriate regulatory sequence, the double-stranded DNA sequence is transcribed and translated into a polypeptide in vitro or in vivo. "Suitable regulatory sequence" refers to a nucleotide sequence located upstream (5' non-coding sequence), internal or downstream (3' non-coding sequence) of the coding sequence, and the nucleotide sequence affects transcription, RNA processing or stability or the translation of the relevant coding sequence. Regulatory sequences may include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites and stem-loop structures. The boundaries of the coding sequence are determined by the start codon at the 5' (amino) end and the translation stop codon at the 3' (carboxyl) end. The coding sequence may include but is not limited to prokaryotic sequences, cDNA from mRNA, genomic DNA sequences, and even synthetic DNA sequences. If the coding sequence is intended to be expressed in eukaryotic cells, polyadenylation signals and transcription termination sequences are usually present at the 3' end of the coding sequence.
[0082] The abbreviation of “open reading frame” is ORF, which means a nucleic acid sequence (DNA, cDNA or RNA) that includes a translation initiation signal or start codon (such as ATG or AUG) and a stop codon and that may be translated into a polypeptide sequence.
[0083] The term "homologous recombination" refers to the insertion of a foreign DNA sequence into another DNA molecule, for example, a vector is inserted into a chromosome. In some cases, the vector targets a specific chromosome site for homologous recombination. For specific homologous recombination, the vector generally contains a sufficiently long region with homology to the chromosome sequence to allow complementary binding of the vector to the chromosome and incorporation of the vector into the chromosome. Longer homology regions and greater sequence similarity can improve the efficiency of homologous recombination.
[0084] According to the disclosure herein, polynucleotides can be amplified using methods known in the art. Once a suitable host system and growth conditions are established, recombinant expression vectors can be amplified and prepared in large quantities. As described herein, expression vectors that can be used include, but are not limited to, the following vectors or their derivatives: human or animal viruses, such as vaccinia virus or adenovirus; insect viruses, such as baculovirus; yeast vectors; phage vectors (e.g., λ), and plasmid and cosmid DNA vectors.
[0085] As used herein, "operably linked" means that a polynucleotide of interest, such as a polynucleotide encoding a Cas9 protein, is linked to a regulatory element in a manner that allows expression of the polynucleotide sequence. In some embodiments, the regulatory element is a promoter. In some embodiments, the polynucleotide of interest is operably linked to a promoter on an expression vector.
[0086] As used herein, "promoter", "promoter sequence" or "promoter region" refers to a DNA regulatory region / sequence that is capable of binding RNA polymerase and is involved in initiating transcription of downstream coding or non-coding sequences. In some examples of the present disclosure, the promoter sequence includes a transcription start site and extends upstream to include the minimum number of bases or elements used to initiate transcription at a level detectable above background. In some embodiments, the promoter sequence includes a transcription start site, and a protein binding domain responsible for RNA polymerase binding. Eukaryotic promoters typically, but not always, contain multiple "TATA" boxes and "CAT" boxes. Various promoters, including inducible promoters, can be used to drive various vectors of the present disclosure.
[0087] "Vector" is any tool for cloning and / or transferring nucleic acid into a host cell. A vector may be a replicon that may be attached to another DNA segment so as to produce replication of the attached segment. "Replicon" is any genetic factor (e.g., plasmid, phage, cosmid, chromosome, virus) that acts as an automatic unit for DNA replication in vivo, i.e., can replicate under its self-control. In some embodiments of the present disclosure, the vector is an additional vector that is removed / lost from a cell population by, for example, asymmetric distribution after many cell generations. The term "vector" includes viral and non-viral tools for introducing the nucleic acid into cells in vitro, in vitro or in vivo. A large number of vectors known in the art can be used to manipulate nucleic acids, integrate response elements and promoters into genes, etc. Possible vectors include, for example, plasmids or modified viruses, including, for example, phages such as lambda derivatives, or plasmids such as pBR322 or pUC plasmid derivatives, or Bluescript vectors. For example, inserting a DNA fragment corresponding to a response element and a promoter into a suitable vector can be accompanied by connecting the suitable DNA fragment to a selected vector with a complementary binding end. Alternatively, the ends of the DNA molecules can be modified by enzyme catalysis or any site is produced by connecting a nucleotide sequence (joint) to the DNA ends. Such vectors can be engineered to include a selectable marker gene that provides for selection of cells, which incorporate the marker into the cell genome. Such markers allow identification and / or selection of host cells that incorporate and express the protein encoded by the marker.
[0088] Viral vectors, particularly retroviral vectors, have been used in many gene delivery applications in cells and living animals. Available viral vectors include, but are not limited to, retrovirus, adeno-associated virus, poxvirus, baculovirus, vaccinia, herpes simplex, Epstein-Barr virus, adenovirus, geminivirus and cauliflower mosaic virus vectors. Non-viral vectors include, but are not limited to, plasmids, liposomes, charged lipids (cytofectins), DNA-protein complexes and biopolymers. In addition to nucleic acids, vectors may also include one or more regulatory regions and / or selective markers for selecting, measuring and monitoring nucleic acid transfer results (transferred to which tissue, expression duration, etc.).
[0089] The vector can be introduced into the desired host cell by known methods, including but not limited to transfection, transduction, cell fusion and lipofection. The vector may include various regulatory elements, including promoters. In some embodiments, the vector design can be based on multiple constructs designed by Mali et al. "Cas9 as a versatile tool for engineering biology", Nature Methods [Natural Methods] 10: 957-63 (2013). In some embodiments, the present disclosure provides an expression vector comprising any polynucleotide described herein, for example, an expression vector comprising a polynucleotide encoding a Cas protein or a variant thereof. In some embodiments, the present disclosure provides an expression vector comprising a polynucleotide encoding a Cas9 protein or a variant thereof.
[0090] The term "plasmid" refers to an extra chromosomal element that usually carries genes that are not involved in the central metabolism of the cell and is usually in the form of a circular double-stranded DNA molecule. Such elements can be linear, circular or supercoiled self-replicating sequences of single-stranded or double-stranded DNA or RNA derived from any source, genome integration sequences, bacteriophages or nucleotide sequences in which many nucleotide sequences have been linked or recombined into a unique structure that is capable of introducing promoter fragments and DNA sequences for selected gene products together with appropriate 3' untranslated sequences into the cell.
[0091] As used herein, "transfection" means introducing exogenous nucleic acid molecules (including vectors) into cells. "Transfected" cells include exogenous nucleic acid molecules inside the cell, and "converted" cells are cells in which the exogenous nucleic acid molecules in the cell induce phenotypic changes in the cell. The transfected nucleic acid molecules can be integrated into the genomic DNA of the host cell and / or can be temporarily or for a long time maintained outside the chromosome by the cell. Host cells or organisms expressing exogenous nucleic acid molecules or fragments are referred to as "recombinant", "converted" or "transgenic" organisms. In some embodiments, the present disclosure provides host cells including any expression vector described herein (e.g., an expression vector including a polynucleotide encoding a Cas protein or a variant thereof). In some embodiments, the present disclosure provides host cells including an expression vector, the expression vector including a polynucleotide encoding a Cas9 protein or a variant thereof.
[0092] The term "host cell" refers to a cell into which a recombinant expression vector has been introduced. The term "host cell" refers not only to the cell into which the expression vector has been introduced (the "parent" cell), but also to the progeny of such a cell. Because modifications may occur in the progeny, for example due to mutations or environmental influences, the progeny may be different from the parent cell, but are still included within the scope of the term "host cell".
[0093] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein to refer to polymeric forms of amino acids of any length, which may include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones.
[0094] The starting point of a protein or polypeptide is called the "N-terminus" (or amino-terminus, NH2-terminus, N-terminus or amine-terminus), which refers to the free amine (-NH2) group of the first amino acid residue of the protein or polypeptide. The end of a protein or polypeptide is called the "C-terminus" (or carboxyl-terminus, carboxyl-terminus, C-terminus or COOH-terminus), which refers to the free carboxyl group (-COOH) of the last amino acid residue of the protein or peptide.
[0095] As used herein, "amino acid" refers to a compound that includes a carboxyl group (-COOH) and an amino group (-NH2). "Amino acid" refers to both natural and non-natural (i.e., synthetic) amino acids. Natural amino acids and their three-letter and one-letter abbreviations include: alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V).
[0096] "Amino acid substitution" refers to a polypeptide or protein comprising one or more wild-type or naturally occurring amino acids replaced by an amino acid different from the wild-type or naturally occurring amino acid at the amino acid residue. The substituted amino acid can be a synthetic or naturally occurring amino acid. In certain embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y and V. Substitution mutants can be described using an abbreviation system. For example, a substitution mutation in which the fifth (5th) amino acid residue is substituted can be abbreviated as "X5Y", wherein "X" is the wild-type or naturally occurring amino acid replaced, "5" is the position of the amino acid residue in the amino acid sequence of the protein or polypeptide, and "Y" is a substituted or non-wild-type or non-naturally occurring amino acid.
[0097] An "isolated" polypeptide, protein, peptide or nucleic acid is a molecule that has been removed from its natural environment. It is also understood that an "isolated" polypeptide, protein, peptide or nucleic acid can be formulated with an excipient (such as a diluent) or adjuvant and still be considered isolated.
[0098] The term "recombinant" when used to refer to nucleic acid molecules, peptides, polypeptides or proteins means a new combination of genetic material not known to exist in nature or produced therefrom. Recombinant molecules can be produced by any well-known technique currently available in the field of recombinant technology, including but not limited to polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid phase synthesis of nucleic acid molecules, peptides or proteins.
[0099] When used to refer to a polypeptide or protein, the term "domain" means a unique functional and / or structural unit in a protein. A domain is sometimes responsible for a specific function or interaction, contributing to the overall effect of the protein. Domains can exist in a variety of biological contexts. Similar domains can be found in proteins with different functions. Alternatively, domains with low sequence identity (i.e., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5%, or less than about 1% sequence identity) may have the same function. In some embodiments, the Cas9 domain is a RuvC domain. In some embodiments, the Cas9 domain is a HNH domain. In some embodiments, the Cas9 domain is a Rec domain.
[0100] When used for polypeptide or protein, the term "motif" generally refers to a group of typically shorter than 20 amino acid conserved amino acid residues, which may be important for protein function. Specific sequence motifs can mediate common functions in multiple proteins, such as protein binding or targeting specific subcellular locations. Examples of motifs include but are not limited to nuclear localization signals, microbody targeting motifs, motifs that prevent or promote secretion, and motifs that promote protein recognition and binding. Motif databases and / or motif search tools are well known to those skilled in the art, and include for example PROSITE (expasy.ch / sprot / prosite.html), Pfam (pfam.wu stl.edu), PRINTS (biochem.ucl.ac.uk / bsm / dbbrowser / PRINTS / PRINT S.html) and Minimotif Miner (cse-mnm.engr.uconn.edu:8080 / MNM / SMS SearchServlet).
[0101] As used herein, an "engineered" protein refers to a protein that includes one or more modifications in the protein to obtain a desired property. Exemplary modifications include, but are not limited to, insertion, deletion, substitution, or fusion with another domain or protein. The engineered proteins of the present disclosure include engineered Cas9 proteins.
[0102] In some embodiments, the engineered protein is produced by a wild-type protein. As used herein, a "wild-type" protein or nucleic acid is a naturally occurring unmodified protein or nucleic acid. For example, a wild-type Cas9 protein can be isolated from the organism Streptococcus pyogenes and can include the amino acid sequence of SEQ ID NO: 1. Wild-type is contrasted with a "mutant", which includes one or more modifications in the amino acid and / or nucleotide sequence of a protein or nucleic acid. For example, a mutant variant of Streptococcus pyogenes Cas9 can include the amino acid sequence of SEQ ID NO: 2, which has a single amino acid substitution relative to wild-type Streptococcus pyogenes Cas9 (SEQ ID NO: 1).
[0103] When used for polypeptides or proteins, the term "degrade" or "degradation" generally refers to the breakdown of proteins into smaller peptide fragments or single amino acids via a process commonly referred to as proteolysis. Intracellular degradation of proteins can be achieved in lysosomes or proteasomes. Lysosomal degradation is generally a non-selective process, except for pathways such as the selective chaperone-mediated autophagy pathway described herein. In lysosomal degradation, cytoplasmic proteins are endocytosed into lysosomes for degradation. Proteasomal degradation is generally selective, where the protein to be degraded is tagged with ubiquitin. For a review of the proteasomal protein degradation pathway, see, e.g., Ciechanover, Cell [Cell] 79 (1): 13-21 (1994); Hasselgren et al., Ann Surg [Annals of Surgery] 225 (3): 307-316 (1997); Collins et al., Cell [Cell] 169 (5): 792-806 (2017). Generally, the degradation rate of a protein is related to its function and biochemical characteristics in the cell. For example, proteins with stretches rich in proline, glutamate, serine, and threonine (sometimes referred to as PEST proteins) have short half-lives (see, e.g., Voet & Voet, Biochemistry, 2nd Edition, John Wiley & Sons, pp. 1010-1014 (1995), which is incorporated by reference in its entirety). Other factors that affect the rate of protein degradation include: the deamination rates of glutamine and asparagine; the oxidation rates of cysteine, histidine, and methionine; the absence of stabilizing ligands; the presence of attached carbohydrates or phosphate groups; the presence of free α-amino groups; the charge of the protein; and the flexibility and stability of the protein (see, e.g., Creighton "Chapter 10—Degradation" in Proteins: Structures and Molecular Properties 2 ndEd. [Proteins: Structure and Molecular Properties 2nd Edition "Chapter 10 - Degradation"] WH Freeman and Company, pp. 463-473 (1993), which is incorporated by reference in its entirety). Methods for measuring protein degradation rates include, for example, amino acid isotope pulse-chase (e.g., stable isotope labeling with amino acids in cell culture or SILAC), post-synthesis radiolabeling, or reporter-dependent methods, such as global protein stability profiling (GPSP), which utilizes, for example, GFP as a reporter protein (see, e.g., Yewdell et al., Cell Biol Int [International Cell Biology] 35(5): 457-462 (2011)). Another method for measuring protein degradation rates is by quantifying the amount of protein in cells at different time points using, for example, a densitometry analysis of immunoblots, plotting protein levels over time, and determining degradation rates from protein level versus time plots. A method for determining protein degradation rates can be selected by one skilled in the art.
[0104] As used herein, the term "sequence similarity" or "% similarity" refers to the degree of identity or correspondence between nucleic acid sequences or amino acid sequences. As used herein, "sequence similarity" refers to nucleic acid sequences in which changes in one or more nucleotide bases result in the substitution of one or more amino acids, but do not affect the functional properties of the protein encoded by the DNA sequence. "Sequence similarity" also refers to modifications of nucleic acids, such as the deletion or insertion of one or more nucleotide bases that do not substantially affect the functional properties of the resulting transcript. Therefore, it should be understood that the present disclosure does not only cover specific exemplary sequences. Methods for making nucleotide base substitutions and methods for determining the retention of biological activity of the encoded products are known.
[0105] In addition, the skilled person recognizes that the similar sequences encompassed by the present disclosure are also defined by their ability to hybridize with the sequences exemplified herein under stringent conditions. The similar nucleic acid sequences disclosed herein are those nucleic acids whose DNA sequences are at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% identical to the DNA sequences of the nucleic acids disclosed herein. The similar nucleic acid sequences disclosed herein are those nucleic acids whose DNA sequences are about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 99%, at least about 99%, or about 100% identical to the DNA sequences of the nucleic acids disclosed herein.
[0106] As used herein, "sequence similarity" refers to two or more amino acid sequences in which greater than about 40% of the amino acids are identical, or greater than about 60% of the amino acids are functionally identical. Functionally identical or functionally similar amino acids have chemically similar side chains. For example, amino acids can be grouped according to functional similarity in the following manner:
[0107] Positively charged side chains: Arg, His, Lys;
[0108] Negatively charged side chains: Asn, Glu;
[0109] Polar, uncharged side chains: Ser, Thr, Asn, Gln;
[0110] Hydrophobic side chains: Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp;
[0111] Others: Cys, Gly, Pro.
[0112] In some embodiments, similar amino acid sequences of the disclosure have at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 99% identical amino acids.
[0113] In some embodiments, the similar amino acid sequences disclosed herein have at least 60%, at least 70%, at least 80%, at least 90% or at least 95% functionally identical amino acids. In some embodiments, the similar amino acid sequences disclosed herein have about 40%, at least about 40%, about 45%, at least about 45%, about 50%, at least about 50%, about 55%, at least about 55%, about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% identical amino acids.
[0114] In some embodiments, the similar amino acid sequences disclosed herein have about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% functionally identical amino acids.
[0115] As used herein, the term "identical protein" refers to a protein having a substantially similar structure or amino acid sequence to a reference protein, which performs the same biochemical function as the reference protein, and may include a protein that differs from the reference protein by substitution or deletion of one or more amino acids at one or more sites in the amino acid sequence, i.e., at least about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% identical amino acids. In one aspect, "identical protein" refers to a protein having the same amino acid sequence as a reference protein.
[0116] Sequence similarity can be determined by sequence alignment using conventional methods in the art, such as BLAST, MUSCLE, Clustal (including ClustalW and ClustalX) and T-Coffee (including variants such as M-Coffee, R-Coffee and Expresso).
[0117] In the context of nucleic acid sequences or amino acid sequences, the term "sequence identity" or "percent identity %" refers to the percentage of identical residues in the sequences compared when the sequences are compared over a specified comparison window. In some embodiments, only specific portions of two or more sequences are compared to determine sequence identity. In some embodiments, only specific domains of two or more sequences are compared to determine sequence similarity. The comparison window can be a segment of at least 10 to more than 1000 residues, at least 20 to about 1000 residues, or at least 50 to 500 residues, in which these sequences can be compared and compared. Alignment methods for determining sequence identity are well known and can be performed using publicly available databases such as BLAST. When referring to amino acid sequences, "percent identity" or "percent identity %" can be determined by methods known in the art. For example, in some embodiments, the "percent identity" of two amino acid sequences is determined using the algorithm of Karlin and Altschul, Proc Nat Acad Sci USA 87:2264-2268 (1990), as modified by Karlin and Altschul, Proc Nat Acad Sci USA 90:5873-5877 (1993). This algorithm is incorporated into BLAST programs, such as BLAST+ or NBLAST and XBLAST programs described in Altschul et al., Journal of Molecular Biology, 215:403-410 (1990). BLAST protein searches can be performed using programs such as the XBLAST program (score = 50, wordlength = 3) to obtain amino acid sequences homologous to the protein molecules of the present disclosure. In the case where there is a gap between two sequences, the gap BLAST program described in, for example, Altschul et al., Nucleic Acids Research [Nucleic Acids Research] 25 (17): 3389-3402 (1997) can be used. When utilizing BLAST programs and gap BLAST programs, the default parameters of the corresponding programs (e.g., XBLAST and NBLAST) can be used.
[0118] In some embodiments, the polypeptide or nucleic acid molecule has 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, or at least 99%, or 100% sequence identity to a reference polypeptide or nucleic acid molecule (or a fragment of a reference polypeptide or nucleic acid molecule), respectively. In some embodiments, the polypeptide or nucleic acid molecule has about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% sequence identity to a reference polypeptide or nucleic acid molecule (or a fragment of a reference polypeptide or nucleic acid molecule), respectively. Overview of CRISPR-Cas Systems
[0119] CRISPR-associated protein 9 (Cas9) is an RNA-guided nuclease of the type II CRISPR adaptive immune system found in bacteria, including but not limited to bacteria such as Streptococcus pyogenes, Streptococcus thermophilus, Staphylococcus aureus and Neisseria meningitidis. For an overview of the CRISPR-Cas9 system, see, for example, Sander et al., Nature Biotechnology [Natural Biotechnology] 32: 347-355 (2014). Generally, CRISPR or CRISPR-Cas systems are characterized by elements that promote the formation of CRISPR complexes at the site of the target sequence, and the CRISPR complexes include guiding polynucleotides and Cas9 nucleases (interchangeably referred to herein as "Cas9 proteins" or "Cas9 nucleases"). In naturally occurring CRISPR-Cas systems, exogenous DNA is introduced into the CRISPR array, and then crRNA (CRISPR-RNA) with a "protospacer sequence" region complementary to the exogenous DNA site is produced. CrRNA hybridizes with tracrRNA (also encoded by the CRISPR system), and the pair of RNAs associates with the Cas9 nuclease. The crRNA / tracrRNA / Cas9 complex recognizes and cleaves exogenous DNA with the protospacer sequence.
[0120] In some embodiments, the present disclosure provides an engineered CRISPR-Cas system. In some embodiments, the engineered CRISPR-Cas system includes an engineered Cas9 protein comprising one or more modifications relative to wild-type Cas9. In some embodiments, the engineered Cas9 protein includes one or more motifs not present in wild-type Cas9. One or more motifs introduced into wild-type Cas9 can be referred to as "engineered" motifs. In some embodiments, one or more engineered motifs in the Cas9 protein are chaperone-mediated autophagy (CMA) motifs.
[0121] In some embodiments, the engineered CRISPR-Cas system includes an engineered guide polynucleotide, which includes one or more modifications relative to wild-type crRNA and / or tracrRNA. In some embodiments, the engineered CRISPR-Cas system utilizes a portion of crRNA and tracrRNA sequences (i.e., a single guide polynucleotide) fusion. Therefore, in this case, a complex is formed between Cas9 and a single guide polynucleotide. A single guide polynucleotide forms a complex with Cas9 to mediate the cleavage of a target sequence, which is complementary to the first (5') 20 nucleotides of the guide polynucleotide (i.e., the guide sequence portion of the guide polynucleotide), and is adjacent to the protospacer sequence adjacent to the motif (PAM) sequence. In other embodiments, the engineered CRISPR-Cas system includes a separate polynucleotide comprising a tracrRNA sequence, i.e., tracrRNA is not a part of the guide polynucleotide comprising a guide sequence. In this case, a complex is formed between Cas9, a guide polynucleotide and tracrRNA. In some embodiments, the tracrRNA component of the guide polynucleotide activates the Cas9 protein. In some embodiments, the activation of the Cas9 protein activates or increases the nuclease activity of Cas9. In some embodiments, the Cas9 protein is not active until it forms a complex with the crRNA and tracrRNA.
[0122] Cas9 endonuclease produces double-stranded DNA breaks at the target sequence upstream of the protospacer sequence adjacent to the motif (PAM). The repair of double-stranded breaks may result in insertion or deletion at the double-stranded break site. In certain embodiments, the endogenous DNA repair pathway of the cell is used to insert the target sequence into the target sequence. Endogenous DNA repair pathways include non-homologous end joining (NHEJ) pathways, microhomology-mediated end joining (MMEJ) pathways, and homology-directed repair (HDR) pathways. NHEJ, MMEJ, and HDR pathways can repair double-stranded DNA breaks. In NHEJ, the break in repair DNA does not require a homologous template. NHEJ repairs may be error-prone, but when DNA breaks include compatible overhangs, errors can be reduced. NHEJ and MMEJ are mechanistically distinct DNA repair pathways, each of which involves a different subset of DNA repair enzymes. Unlike NHEJ, which may be accurate or error-prone in some cases, MMEJ is always error-prone and can result in deletions and insertions at the repair site. MMEJ-related deletions are attributed to microhomologies (2-10 base pairs) on both sides of the double-strand break. In contrast, HDR requires a homologous template to directly repair, but HDR repairs typically have high fidelity and are less error-prone. In some embodiments, the error-prone nature of NHEJ and MMEJ repairs is utilized to introduce nonspecific nucleotide substitutions in the target sequence.
[0123] As described herein, some CRISPR-Cas systems may have undesirable off-target activity or off-target genome editing. "Off-target" used in the context of genome editing refers to non-specific and unexpected genetic modifications, which is contrary to "on target", which refers to modifications at the expected locus. When, for example, the Cas9 nuclease does not bind to its expected target sequence (i.e., a genomic sequence complementary to the guide sequence on the guide polynucleotide), off-target modifications may result, which may be caused by homologous sequences and / or mismatch tolerance. Off-target modifications may include, but are not limited to, unexpected point mutations, deletions, insertions, inversions, and translocations. In some embodiments, compared to wild-type Cas9 protein, the engineered Cas9 protein of the present disclosure has reduced off-target activity. In some embodiments, compared to wild-type Cas9 protein, the off-target activity of the engineered Cas9 protein of the present disclosure is reduced by at least about 50%. In some embodiments, compared to wild-type Cas9, the off-target activity of the engineered Cas9 protein of the present disclosure is reduced by at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 100%. Off-target modifications can use, for example, targeted sequencing, exome sequencing, whole genome sequencing, BLESS (direct in situ break tagging, streptavidin enrichment and next generation sequencing), GUIDE-seq (whole genome, unbiased identification of DSBs by sequencing), LAM-HTGTS (linear amplification-mediated high-throughput whole genome translocation sequencing) and Digenome-seq (whole genome sequencing of in vitro Cas9 digestion). Off-target modification detection and quantitative methods are described in, for example, Zhang et al., Mol Ther Nucleic Acids [molecular therapeutic nucleic acids] 4: e264 (2014); and Zischewski et al., Biotechnol Adv [biotechnological progress] 35: 95-104 (2017).
[0124] Cas9 protein
[0125] In some embodiments, the Cas9 protein is derived from the following species: Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus dysgalactiae, Streptococcus mutans, Listeria innocua, Staphylococcus aureus or Klebsiella pneumoniae. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 1) of Streptococcus pyogenes Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 17) of Streptococcus thermophilus Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 18) of Streptococcus dysgalactiae Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 19) of Streptococcus mutans Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 20) of Listeria innocua Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence (SEQ ID NO: 21) of Staphylococcus aureus Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Klebsiella pneumoniae Cas9 protein (SEQ ID NO: 22).
[0126] In some embodiments, the term Cas9 refers to a polypeptide comprising SEQ ID NO: 1. In some embodiments, the Cas9 protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 1. In some embodiments, Cas9 is a polypeptide encoded by a polynucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to SEQ ID NO: 3.
[0127] In some embodiments, the term Cas9 refers to Cas9 that can produce sticky ends. As used herein, the term "cohesive end", "staggered end" or "sticky end" refers to a nucleic acid fragment with chains of unequal lengths. In contrast to "flat ends", sticky ends are produced by staggered cutting of nucleic acids (typically DNA). Sticky or sticky ends have protruding single-stranded chains (these chains have unpaired nucleotides) or overhangs, for example, 3' or 5' overhangs. Each overhang can be annealed with another complementary overhang to form a base pair. Two complementary sticky ends can be annealed together via interactions such as hydrogen bonding. The stability of the annealed sticky ends depends on the melting temperature of the paired overhangs. Two complementary sticky ends can be connected together by chemical or enzymatic connection (for example, by DNA ligase).
[0128] In certain embodiments, the term Cas9 refers to a Cas9 variant with a changed function, such as a Cas9 hybrid protein. For example, the binding domain of Cas9 or a Cas9 protein with an inactive DNA cleavage domain can be used as a binding domain specifically bound to a desired target sequence via a guidance polynucleotide. The binding domain (i.e., inactive Cas9) can be fused or conjugated to a cleavage domain (e.g., a cleavage domain of an endonuclease FokI) to produce an engineered hybrid nuclease. Cas9-FokI hybrid proteins are further described in, for example, U.S. Patent Publication No. 2015 / 0071899 and Guilinger et al., Nature Biotechnology [Natural Biotechnology] 32: 577-582 (2014). Other examples of engineered hybrid nucleases are described in, e.g., Wah et al., Proc Nat Acad Sci [Proceedings of the National Academy of Sciences of the United States of America] 95:10564-10569 (1996); Li et al., Nucl Acids Res [Nucleic Acids Research] 39(1):359-372 (2011); and Kim et al., Proc Nat Acad Sci [Proceedings of the National Academy of Sciences of the United States of America] 93:1156-1160 (1996).
[0129] Cpf1 (centromere and promoter factor 1) is also an RNA-guided nuclease of the type II CRISPR system. Cpf1 produces sticky ends. The CRISPR / Cpf1 system is similar to the CRISPR / Cas9 system. However, there are some differences between Cas9 and Cpf1. Unlike Cas9, Cpf1 does not use tracrRNA. The Cpf1 protein recognizes a PAM sequence different from that of Cas9, and Cpf1 cleaves at a different site from that of Cas9. Cas9 cleaves at a sequence adjacent to the PAM, while Cpf1 cleaves at a sequence away from the PAM. The Cpf1 protein is further described in, for example, foreign patent disclosures GB 1506509.7, U.S. Patent No. 9,580,701, U.S. Patent Publication 2016 / 0208243, and Zetsche et al., Cell [Cell] 163 (3): 759-771 (2015). According to the present disclosure, enzymes that are functionally similar to Cpf1 can be used. Thus, in some embodiments, the present disclosure provides recombinant Cpf1 proteins comprising the amino acid modifications described herein.
[0130] Some wild-type or naturally occurring Cas9 proteins (e.g., Cas9 proteins from Streptococcus pyogenes) have six domains: Rec1, Rec2, Bridge Helix (BH), PAM interaction (PI), HNH, and RuvC. The Rec1 domain is responsible for binding to the guide polynucleotide. The BH domain is responsible for initiating cleavage activity when bound to the target sequence. The PI domain confers PAM specificity and is responsible for initiating binding to the target sequence. The HNH and RuvC domains are nuclease domains that cut DNA. Structural studies of the Cas9 protein have revealed that the protein has a recognition lobe ("REC lobe"), which includes the BH, Rec1, and Rec2 domains; and a nuclease lobe ("NUC lobe"), which includes RuvC (divided into RuvC I, RuvC II, and RuvC III subdomains), HNH, and PI domains. See Figure 3 and 4Protein domains can be identified using domain structure prediction tools based on protein amino acid sequences, such as SMART (Letunic et al., Nucleic Acids Research (2017), doi: 10.1093 / nar / gkx922), PANDA (Wang et al., Scientific Reports 8:3484 (2018)), or InterPro (Finn et al., Nucleic Acids Research (2017), doi: 10.1093 / nar / gkw1107). Protein domains can also be identified based on protein structure (e.g., by visual inspection) or by using algorithms such as PUU (Holm et al., Proteins 19(3):256-268 (1994)), RigidFinder (Abyzov et al., Proteins 78(2):309-324 (2010)), or PiSQRD (Aleksiev et al., Bioinformatics 25(20):2743-2744 (2009)). Identification of Cas9 domains based on structural characterization is described, for example, in Jinek et al., Science 337:816-821 (2012); Nishimasu et al., Cell 156(5):935-949 (2014); Anders et al., Nature 513:569-573 (2014); and Sternberg et al., Nature 507(7490):62-67 (2014).
[0131] In some embodiments, the Cas9 proteins of the present disclosure include a REC lobe that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the Cas9 proteins of the present disclosure include a NUC lobe that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 6-7.
[0132] In some embodiments, the Cas9 protein of the present disclosure comprises a BH domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 8. In some embodiments, the Cas9 protein of the present disclosure comprises a Rec1 domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 9-10. In some embodiments, the Cas9 protein of the present disclosure comprises a Rec2 domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 11. In some embodiments, the Cas9 protein of the present disclosure comprises a RuvC domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 12-14. In some embodiments, the Cas9 protein of the present disclosure includes an HNH domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 15. In some embodiments, the Cas9 protein of the present disclosure includes a PI domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 16.
[0133] Structural studies (e.g., crystal structures) of Cas9 proteins can reveal regions of surface-exposed proteins. As used herein, "surface-exposed regions" refer to regions of proteins that are accessible to the surrounding environment, i.e., regions on the outer "surface" of the protein. Similarly, "surface-exposed residues" include amino acid residues of proteins that are in surface-exposed regions. Surface-exposed residues are opposite to "buried" residues, which face inward toward the center of the protein and form a "buried zone" that cannot be approached by the surrounding environment. Surface-exposed residues on proteins may play an important role in interactions with other molecules (e.g., other proteins or cell structures). Therefore, in some embodiments, certain residues on a protein (e.g., Cas9 protein) are surface-exposed in a conformational state, and are not surface-exposed in different conformational states. For example, the Cas9 protein can undergo a conformational change upon binding to a guide RNA such that a previously unexposed region of the Cas9 protein becomes surface exposed upon guide RNA binding, or vice versa (see, e.g., Fagerlund et al., Proc Nat Acad Sci 114(26):E5211-E5128 (2017)).
[0134] Surface exposed residues may also determine the physical properties of proteins and limit the folding structure of proteins. When viewing protein crystal structures in programs such as PyMOL (pymol.org) or Swiss PDB Viewer (spdbv.vital-it.ch), surface exposed residues can be determined. Surface exposed residues can also be calculated using programs such as NACCESS (bioinf.manchester.ac.uk / naccess). The surface exposed residues of proteins can also be determined by computational prediction, for example, when the crystal structure is not available. The computational prediction tools for surface exposed residues in protein sequences include, for example, SARpred (Garg et al., Proteins [protein] 61: 318-24 (2005)), PSA / TEM (Mizuguchi et al., Bioinformatics [bioinformatics] 14: 617-623 (1998)) and RSARF (caps.ncbs.res.in / download / pugal / RSARF) in the JOY program package.
[0135] Overview of chaperone-mediated autophagy
[0136] "Chaperone-mediated autophagy" or CMA refers to a selective protein degradation process that involves chaperone-dependent selection of cytosolic proteins, which are then targeted to lysosomes and translocated across the lysosomal membrane for degradation. An exemplary chaperone protein for CMA is heat shock cognate protein of 70 kD, or HSC70. "Endosomal microautophagy" or eMI refers to a protein degradation process similar to CMA, except that eMI selectively targets proteins that include a KFERQ motif or a KFERQ-like motif to late endosomes rather than lysosomes for degradation. Like CMA, HSC70 is also a chaperone protein for eMI. See, e.g., Kaushik et al., Trends Cell Biol 22(8):407-417 (2012); Tekirdag et al., J Biol Chem 293:5414-5424 (2018); and Pereira et al., Int J Cell Biol 2012(4):931956 (2012).
[0137] The "KFERQ motif" referred to herein is a pentapeptide sequence: Lys-Phe-Glu-Arg-Gln (SEQ ID NO: 24). The "KFERQ-like motif" referred to herein is a motif that is biochemically similar or biochemically related to KFERQ. As described herein, a biochemically similar or biochemically related motif may include functionally equivalent amino acid residues. Thus, the KFERQ motif may be any pentapeptide having the following parameters: one or two positively charged residues (e.g., Lys or Arg); one or two bulky hydrophobic residues (e.g., Phe, Ile, Leu, or Val); a negatively charged residue (e.g., Asp or Glu); and Gln or Asn located on either side of the pentapeptide. See, for example, Dice et al., Trends Biochem Sci 15(8):305-309 (1990); and Kaushik et al., Trends Cell Biol 22(8):407-417 (2012). Examples of KFERQ-like motifs include, but are not limited to, those listed in Table 1.
[0138] Table 1. KFERQ-like motifs
[0139]
[0140]
[0141] Proteins comprising at least one KFERQ motif or KFERQ-like motif can be recognized by components of CMA or eMI. Therefore, in some embodiments, the KFERQ motif or KFERQ-like motif is a chaperone-mediated autophagy (CMA) target motif. In some embodiments, the KFERQ motif or KFERQ-like motif is an endosomal microautophagy (eMI) target motif. Without being bound by a particular theory, for the purpose of illustrating the present disclosure, CMA and eMI are described herein, and it should be understood that the KFERQ motif or KFERQ-like motif can be used as a target for other protein degradation pathways, and other consensus sequences or motifs (different from the KFERQ motif or KFERQ-like motif described herein) can be CMA or eMI target motifs.
[0142] HSC70 recognizes and binds to CMA or eMI target motifs on proteins, such as KFERQ motifs or KFERQ-like motifs, to form a chaperone protein complex. The chaperone protein complex then binds to the 2A type lysosomal associated membrane protein (LAMP-2A) receptor. The protein unfolds, which triggers the multimerization of LAMP-2A. Subsequently, the unfolded protein is translocated across the lysosomal membrane via LAMP-2A, and the transported protein is finally degraded. See, for example, Kaushik et al., Trends Cell Biol [Cell Biology Trends] 22 (8): 407-417 (2012).
[0143] Recombinant Cas9 protein
[0144] Compared to wild-type Cas9, the recombinant Cas9 protein of the present disclosure is a functional Cas9 nuclease and has reduced off-target modifications. "Functional Cas9 nuclease" means that the recombinant Cas9 protein has at least about the same level of nuclease activity as the wild-type Cas9 protein, as measured by a Cas9 activity assay. "Functional Cas9 nuclease" also means that the recombinant Cas9 has about the same level of on-target modification (i.e., genome editing efficiency) as the wild-type Cas9 protein, as measured by a Cas9 efficiency assay.
[0145] In some embodiments, the recombinant Cas9 protein of the present disclosure has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 100% of the nuclease activity of the wild-type Cas9 protein, as measured by a Cas9 activity assay. In some embodiments, the recombinant Cas9 protein of the present disclosure has higher nuclease activity than the wild-type Cas9 protein, as measured by a Cas9 activity assay. Non-limiting examples of Cas9 activity assays include T7 endonuclease I assays and SURVEYOR assays (reviewed in Vouillot et al., G3 (Bethesda) 5 (3): 407-415 (2015)). In some embodiments, the recombinant Cas9 protein of the present disclosure has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 100% of the on-target modification of the wild-type Cas9 protein, as measured by a Cas9 efficiency assay. In some embodiments, the recombinant Cas9 protein of the present disclosure has higher on-target modification than wild-type Cas9 protein, as measured by Cas9 efficiency assay. Non-limiting examples of Cas9 efficiency assays include mismatch detection assays and sequencing-based assays (reviewed in Zischewski et al., Biotechnol Adv [Biological Technology Progress] 35: 95-104 (2017)).
[0146] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif.
[0147] As described herein, the KFERQ motif or KFERQ-like motif is recognized by a component of a CMA or eMI. Thus, in some embodiments, a Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif is recognized by a component of a CMA or eMI. In some embodiments, the KFERQ motif or KFERQ-like motif is any one of SEQ ID NOs: 24-41. Thus, in some embodiments, the KFERQ motif or KFERQ-like motif is KFERQ (SEQ ID NO:24), RKVEQ (SEQ ID NO:25), QDLKF (SEQ ID NO:26), QRFFE (SEQ ID NO:27), NRVVD (SEQ ID NO:28), QRDKV (SEQ ID NO:29), QKILD (SEQ ID NO:30), QKKEL (SEQ ID NO:31), QFREL (SEQ ID NO:32), IKLDQ (SEQ ID NO:33), DVVRQ (SEQ ID NO:34), QRIVE (SEQ ID NO:35), VKELQ (SEQ ID NO:36), QKVFD (SEQ ID NO:37), QELLR (SEQ ID NO:38), VDKLN (SEQ ID NO:39), RIKEN (SEQ ID NO:40), or NKKFE (SEQ ID NO:41). In some embodiments, the engineered KFERQ motif or KFERQ-like motif is VDKLN (SEQ ID NO: 39).
[0148] In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence KFERQ (SEQ ID NO: 24). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence RKVEQ (SEQ ID NO: 25). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QDLKF (SEQ ID NO: 26). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QRFFE (SEQ ID NO: 27). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence NRVVD (SEQ ID NO: 28). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QRDKV (SEQ ID NO: 29).
[0149] In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QKILD (SEQ ID NO: 30). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QKKEL (SEQ ID NO: 31). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QFREL (SEQ ID NO: 32). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence IKLDQ (SEQ ID NO: 33). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence DVVRQ (SEQ ID NO: 34). In some embodiments, the recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QRIVE (SEQ ID NO: 35). In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence VKELQ (SEQ ID NO: 36).
[0150] In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QKVFD (SEQ ID NO: 37). In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence QELLR (SEQ ID NO: 38). In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence VDKLN (SEQ ID NO: 39). In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence RIKEN (SEQ ID NO: 40). In some embodiments, the recombinant Cas9 protein comprises an engineered KFERQ motif or KFERQ-like motif having the amino acid sequence NKKFE (SEQ ID NO: 41).
[0151] In some embodiments, the engineered KFERQ motif or KFERQ-like motif precedes the first amino acid residue of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 1 to 100 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 100 to 300 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 300 to 700 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 700 to 900 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 900 to 1100 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is between amino acid residues 1100 to 1300 of SEQ ID NO: 1. In some embodiments, the engineered KFERQ motif or KFERQ-like motif follows the last amino acid residue of SEQ ID NO:1.
[0152] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the REC lobe of the Cas9 protein. In some embodiments, the REC lobe of the Cas9 protein includes a BH domain, a Rec1 domain, and a Rec2 domain. In some embodiments, the REC lobe has an amino acid sequence of SEQ ID NO:5. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the Rec1 domain of the REC lobe. In some embodiments, the Rec1 domain has an amino acid sequence of SEQ ID NO:9-10. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the Rec2 domain of the REC lobe. In some embodiments, the Rec2 domain has an amino acid sequence of SEQ ID NO:11. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the BH domain of the REC lobe. In some embodiments, the BH domain has an amino acid sequence of SEQ ID NO:8.
[0153] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the NUC lobe of the Cas9 protein. In some embodiments, the NUC lobe of the Cas9 protein includes a RuvC domain, an HNH domain, and a PI domain. In some embodiments, the NUC lobe has an amino acid sequence of SEQ ID NO: 6-7. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the RuvC domain, the HNH domain, and the PI domain of the Cas9 protein. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the RuvC domain. In some embodiments, the RuvC domain has an amino acid sequence of SEQ ID NO: 12-14. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the HNH domain. In some embodiments, the HNH domain has an amino acid sequence of SEQ ID NO: 15. In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in the PI domain. In some embodiments, the PI domain has an amino acid sequence of SEQ ID NO: 16.
[0154] In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif comprises a lobe that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif comprises a lobe that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 6-7.
[0155] In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif comprises a BH domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 8. In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif comprises a Rec1 domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 9-10. In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif comprises a Rec2 domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO:11.
[0156] In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif comprises a RuvC domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 12-14. In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif comprises a HNH domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 15. In some embodiments, the Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif comprises a PI domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identity to the amino acid sequence of SEQ ID NO:16.
[0157] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is in a surface-exposed region of the recombinant Cas9 protein. As described herein, the surface-exposed region refers to a region of the Cas9 protein that is accessible to the surrounding environment, for example, a component of a protein degradation pathway that is accessible. In some embodiments, the surface-exposed region of the recombinant Cas9 protein is in the REC lobe of the Cas9 protein. In some embodiments, the surface-exposed region of the recombinant Cas9 protein is in the NUC lobe of the Cas9 protein. In some embodiments, the surface-exposed region of the recombinant Cas9 protein is in the Rec1 domain, Rec2 domain, BH domain, RuvC domain, HNH domain, or PI domain of the Cas9 protein. In some embodiments, the surface-exposed region of the recombinant Cas9 protein is between amino acid residues 150 and 250 of the Cas9 protein.
[0158] In some embodiments, the engineered KFERQ motif or KFERQ-like motif is located at the N-terminus or C-terminus of the recombinant Cas9 protein. As described herein, the N-terminus is the "start" of a protein or polypeptide, and the C-terminus is the "end" of a protein or polypeptide. Therefore, in some embodiments, the KFERQ motif or KFERQ-like motif is located at the "start" of the N-terminus of the Cas9 protein. In some embodiments, the KFERQ motif or KFERQ-like motif is located at the "end" of the C-terminus of the Cas9 protein. In some embodiments, adding an engineered motif to the N-terminus or C-terminus of the protein does not affect the folding, structure or dynamics of the protein. In some embodiments, the N-terminus of Cas9 is surface exposed. In some embodiments, the C-terminus of Cas9 is surface exposed.
[0159] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the modifications introduce a chaperone-mediated autophagy (CMA) target motif or an endosomal microautophagy (eMI) target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the recombinant Cas9 protein is degraded in vivo at least 50% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the recombinant Cas9 protein is degraded in vivo at least 80% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif.
[0160] In some embodiments, the present disclosure provides a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the recombinant Cas9 protein comprises a CMA target motif or an eMI target motif.
[0161] As described herein, proteins containing a CMA motif or an eMI motif are targets of the CMA or eMI protein degradation pathway. Thus, in some embodiments, a recombinant Cas9 protein of the present disclosure comprising one or more amino acid modifications that introduce a CMA or eMI target motif is targeted for protein degradation via CMA or eMI. Likewise, in some embodiments, a recombinant Cas9 protein of the present disclosure comprising a CMA target motif or an eMI target motif is targeted for protein degradation via CMA or eMI.
[0162] In some embodiments, the recombinant Cas9 protein comprising a CMA or eMI target motif is degraded in vivo at least 20% faster, at least 30% faster, at least 40% faster, at least 50% faster, at least 60% faster, at least 70% faster, at least 80% faster, at least 90% faster, at least 100% faster, at least 150% faster, at least 200% faster, at least 500% faster than a wild-type Cas9 protein or a Cas9 protein that does not comprise a CMA or eMI target motif, as measured by immunoblotting or a GFP reporter assay. In some embodiments, if the same cell expresses: (a) a recombinant Cas9 comprising one or more amino acid modifications that introduce a CMA or eMI target motif, and (b) a wild-type Cas9, the recombinant Cas9 is completely degraded, while at least 50% of the wild-type Cas9 remains within the cell. Similarly, in some embodiments, if the same cell expresses: (a) a recombinant Cas9 comprising one or more amino acid modifications that introduce a CMA or eMI target motif, and (b) a Cas9 protein that does not include a CMA or eMI target motif, the recombinant Cas9 is completely degraded, while at least 50% of the Cas9 protein that does not include a CMA or eMI target motif remains in the cell. In some embodiments, the recombinant Cas9 is completely degraded, while at least 60%, at least 70%, at least 80%, at least 90%, or at least 100% of the wild-type Cas9 or the Cas9 protein that does not include a CMA or eMI target motif remains in the cell. In some embodiments, the recombinant Cas9 is completely degraded within 12 hours, within 24 hours, within 36 hours, within 48 hours, or within 72 hours of introduction into the cell. As used in the embodiments herein, "complete degradation" refers to a protein that is below the detection level of a GFP reporter gene assay or immunoblotting. As described herein, methods for measuring protein degradation rates include, for example, amino acid isotope pulse-chase (e.g., stable isotope labeling with amino acids in cell culture or SILAC), post-synthesis radiolabeling, or reporter-dependent methods, such as global protein stability profiling (GPSP), which utilizes, for example, GFP as a reporter protein (see, e.g., Yewdell et al., Cell Biol Int [International Cell Biology] 35(5): 457-462 (2011)). In some embodiments, the degradation rate of the Cas9 protein is measured by quantifying the amount of Cas9 protein in the cell at different time points using, for example, a densitometry analysis of an immunoblot, plotting the protein level over time, and determining the degradation rate from the Cas9 protein level versus time plot.
[0163] In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions F185 of SEQ ID NO: 1. In some embodiments, the mutation is F185N. In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions A547 and I548 of SEQ ID NO: 1. In some embodiments, the mutations are A547E and I548L. In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions T560 and V561 of SEQ ID NO: 1. In some embodiments, the mutations are T560E and V561Q. In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions D829 and I830 of SEQ ID NO: 1. In some embodiments, the mutations are D829L and I830R. In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions L1087 and S1088 of SEQ ID NO: 1. In some embodiments, the mutations are L1087E and S1088Q. In some embodiments, the one or more amino acid modifications in the recombinant Cas9 include mutations at positions P1199 and K1200 of SEQ ID NO: 1. In some embodiments, the mutations are P1199D and K1200Q.
[0164] In some embodiments, one or more amino acid modifications in the recombinant Cas9 include a combination of any mutations described herein. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from F185N, A547E, I548L, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from F185N, T560E, V561Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from F185N, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from F185N, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from F185N, P1199D, K1200Q, and combinations thereof.
[0165] In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from A547E, I548L, T560E, V561Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from A547E, I548L, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from A547E, I548L, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from A547E, I548L, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from T560E, V561Q, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from T560E, V561Q, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from T560E, V561Q, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from D829L, I830R, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from D829L, I830R, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant Cas9 include mutations selected from L1087E, S1088Q, P1199D, K1200Q, and combinations thereof. In some embodiments, as described herein, one or more amino acid modifications in the recombinant Cas9 protein result in one or more CMA target motifs or eMI target motifs.
[0166] In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 50% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 60% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 70% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 80% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 90% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to SEQ ID NO: 1, and includes one or more amino acid modifications described herein.
[0167] In some embodiments, the present disclosure provides a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, which recombinant Cas9 protein comprises an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO:1.
[0168] In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include a mutation at position F185 of SEQ ID NO: 1. In some embodiments, the mutation is F185N. In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include mutations at positions A547 and I548 of SEQ ID NO: 1. In some embodiments, the mutations are A547E and I548L. In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include mutations at positions T560 and V561 of SEQ ID NO: 1. In some embodiments, the mutations are T560E and V561Q. In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include mutations at positions D829 and I830 of SEQ ID NO: 1. In some embodiments, the mutations are D829L and I830R. In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include mutations at positions L1087 and S1088 of SEQ ID NO: 1. In some embodiments, the mutations are L1087E and S1088Q. In some embodiments, the one or more amino acid modifications in the recombinant SpCas9 include mutations at positions P1199 and K1200 of SEQ ID NO: 1. In some embodiments, the mutations are P1199D and K1200Q.
[0169] In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include a combination of any mutations described herein. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from F185N, A547E, I548L, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from F185N, T560E, V561Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from F185N, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from F185N, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from F185N, P1199D, K1200Q, and combinations thereof.
[0170] In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from A547E, I548L, T560E, V561Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from A547E, I548L, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from A547E, I548L, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from A547E, I548L, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from T560E, V561Q, D829L, I830R, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from T560E, V561Q, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from T560E, V561Q, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from D829L, I830R, L1087E, S1088Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from D829L, I830R, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 include mutations selected from L1087E, S1088Q, P1199D, K1200Q, and combinations thereof. In some embodiments, one or more amino acid modifications in the recombinant SpCas9 protein result in one or more CMA target motifs or eMI target motifs. CMA target motifs and eMI target motifs are as described herein.
[0171] In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 50% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein. In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 60% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein. In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 70% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein. In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 80% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein. In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 90% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein. In some embodiments, the recombinant SpCas9 protein has an amino acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 1 and includes one or more amino acid modifications described herein.
[0172] In some embodiments, the present disclosure provides a recombinant Cas9 protein capable of binding to heat shock cognate protein of 70 kD (HSC70).
[0173] As described herein, HSC70 is a molecular chaperone protein in the degradation pathway of CMA and eMI proteins. HSC70 binds to proteins targeted for degradation and transports the proteins to lysosomes (in the case of CMA) or late endosomes (in the case of eMI) for degradation. Therefore, in some embodiments, proteins with higher binding affinity to HSC70 are degraded faster than proteins with lower binding affinity to HSC70. In some embodiments, the binding ability and / or affinity of a protein to HSC70 is determined by the presence of a CMA or eMI target motif, such as a KFERQ motif or a KFERQ-like motif, on the protein.
[0174] In some embodiments, the recombinant Cas9 protein of the present disclosure is capable of binding to HSC70. As described and exemplified herein, wild-type Cas9 or Cas9 proteins that do not include a KFERQ motif or a KFERQ-like motif do not bind to HSC70. In some embodiments, the recombinant Cas9 protein of the present disclosure is capable of binding to HSC70 with an affinity that is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100% higher than that of the wild-type Cas9 protein or the Cas9 protein that does not include a KFERQ motif or a KFERQ-like motif. Methods for determining binding affinity between proteins are known in the art and include, for example, biochemical methods, such as: co-immunoprecipitation, bimolecular fluorescence complementation, affinity electrophoresis, pull-down assays, phage display, in vivo cross-linking, tandem affinity purification, mass spectrometry after cross-linking, and proximity ligation assays; biophysical methods, such as: biolayer interferometry, dynamic light scattering, surface plasmon resonance, fluorescence resonance energy transfer, and isothermal titration calorimetry; and / or genetic methods, such as: yeast two-hybrid screening and bacterial two-hybrid screening. For an overview of methods for measuring binding affinity and detecting protein interactions, see, e.g., Meyerkord and Fu, Protein-Protein Interactions: Methods and Applications 2 nd Ed. [Protein-Protein Interactions: Methods and Applications 2nd Edition] 2015, Humana Press. In some embodiments, the recombinant Cas9 disclosed herein can be detected by HSC70 antibody after incubation with HSC70 for a period of time, while the wild-type Cas9 or the Cas9 protein that does not include the KFERQ motif or the KFERQ-like motif is not detected by the HSC70 antibody after incubation for the same period of time.
[0175] In some embodiments, the binding affinity between HSC70 and recombinant Cas9 is at least 2 times higher, at least 3 times higher, at least 4 times higher, at least 5 times higher, at least 6 times higher, at least 7 times higher, at least 8 times higher, at least 9 times higher, at least 10 times higher, at least 20 times higher, at least 30 times higher, at least 40 times higher, at least 50 times higher, at least 60 times higher, at least 70 times higher, at least 80 times higher, at least 90 times higher, at least 100 times higher, at least 500 times higher, or at least 1000 times higher than the binding affinity between HSC70 and wild-type Cas9 or a Cas9 protein that does not include a KFERQ motif or a KFERQ-like motif.
[0176] In some embodiments, the present disclosure provides a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, which recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO:1.
[0177] As described herein, SEQ ID NO: 1 includes the amino acid sequence of a wild-type Cas9 protein (SpCas9) from Streptococcus pyogenes. Amino acid position 185 of SEQ ID NO: 1 is in the region corresponding to the Rec2 domain of SpCas9. In some embodiments, the amino acid residue at position 185 of SEQ ID NO: 1 is modified to produce a KFERQ motif or a KFERQ-like motif. In some embodiments, the KFERQ-like motif at position 185 of SEQ ID NO: 1 is VDKLN. In some embodiments, the recombinant SpCas9 protein includes a mutation at position 185 of SEQ ID NO: 1. In some embodiments, the mutation is F185N.
[0178] In some embodiments, the recombinant Cas9 of the present disclosure further includes mutations at positions D10 and / or H840 of SEQ ID NO:1. Mutations at positions D10 and / or H840 of wild-type Cas9 produce Cas9 with nickase activity, also referred to herein as "Cas9 nickase". Cas9 nickase is only capable of cleaving one strand of double-stranded DNA (i.e., "nicking" the DNA). Cas9 nickase is described in, for example, Cho et al., Genome Res [Genome Research] 24: 132-141 (2013). In some embodiments, the recombinant Cas9 protein of the present disclosure further includes mutations at amino acid position D10 of SEQ ID NO:1. In some embodiments, the recombinant Cas9 protein of the present disclosure further includes mutations at amino acid position H840 of SEQ ID NO:1. In some embodiments, the recombinant Cas9 protein of the present disclosure further includes mutations at amino acid position D10 of SEQ ID NO:1 and mutations at amino acid position H840. In some embodiments, the mutation at position D10 is D10A. In some embodiments, the mutation at position D10 is D10N. In some embodiments, the mutation at position H840 is H840A. In some embodiments, the mutation at position H840 is H840N. In some embodiments, the mutation at position H840 is H840Y. In some embodiments, the recombinant Cas9 protein has a F185N mutation and a D10A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation and a D10N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation and a H840N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation and a H840Y mutation.
[0179] In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10A mutation, and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10A mutation, and a H840N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840Y mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10A mutation, and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10A mutation, and a H840N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10A mutation, and a H840Y mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840A mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840N mutation. In some embodiments, the recombinant Cas9 protein has a F185N mutation, a D10N mutation, and a H840Y mutation.
[0180] In some embodiments, the recombinant Cas9 protein disclosed herein produces sticky ends. As described herein, sticky ends refer to nucleic acid fragments with chains of unequal lengths. In some embodiments, the recombinant Cas9 protein that produces sticky ends is a recombinant Cas9-FokI hybrid protein. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and a D10A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and a D10N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and an H840A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and an H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and an H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has an F185N mutation and an H840Y mutation in Cas9.
[0181] In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840Y mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840Y mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10A mutation, and a H840Y mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840A mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840N mutation in Cas9. In some embodiments, the recombinant Cas9-FokI hybrid protein has a F185N mutation, a D10N mutation, and a H840Y mutation in Cas9.
[0182] In some embodiments, the recombinant Cas9 protein that produces sticky ends includes an engineered KFERQ motif or KFERQ-like motif on a wild-type Cas9 protein that can produce sticky ends. In some embodiments, the wild-type Cas9 protein that can produce sticky ends is isolated from Francisella novicida (FnCas9) (SEQ ID NO: 23). In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least about 90% sequence identity to SEQ ID NO: 23, and includes an engineered KFERQ motif or KFERQ-like motif as described herein. In some embodiments, the recombinant Cas9 protein has an amino acid sequence with at least about 90% sequence identity to SEQ ID NO: 23, and includes a CMA target motif or an eMI target motif. In some embodiments, the recombinant Cas9 has an amino acid sequence with at least about 90% sequence identity to SEQ ID NO: 23, and is capable of binding to HSC70 with a higher affinity than wild-type Cas9. In some embodiments, the recombinant Cas9 has an amino acid sequence with at least about 90% sequence identity to SEQ ID NO:23, and degrades faster than wild-type Cas9.
[0183] In some embodiments, the recombinant Cas9 of the present disclosure includes one or more nuclear localization signals. "Nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein to be imported into the nucleus by nuclear transport, i.e., a protein with NLS is transported to the nucleus. Typically, NLS includes positively charged Lys or Arg residues exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLSs from the following: SV40 large T antigen, nucleoplasmic protein, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS includes the sequence PKKKRKV (SEQ ID NO: 42). In some embodiments, the NLS includes the sequence AVKRPAATKKAGQAKKKKLD (SEQ ID NO: 43). In some embodiments, the NLS includes the sequence PAAKRVKLD (SEQ ID NO: 44). In some embodiments, the NLS includes the sequence MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 45). In some embodiments, the NLS includes the sequence KLKIKRPVK (SEQ ID NO: 46). Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK (SEQ ID NO: 47) in the yeast transcription repressor Matα2, and PY-NLS.
[0184] Nucleotides
[0185] In some embodiments, the present disclosure provides polynucleotide sequences encoding recombinant Cas9 proteins described herein. In some embodiments, the present disclosure provides polynucleotide sequences encoding recombinant Cas9 comprising an engineered KFERQ motif or a KFERQ-like motif. In some embodiments, the present disclosure provides polynucleotide sequences encoding recombinant Cas9 proteins, the recombinant Cas9 proteins comprising one or more amino acid modifications of wild-type Cas9 proteins, the modifications introducing CMA target motifs or eMI target motifs into the Cas9 proteins, wherein the recombinant Cas9 proteins are degraded in vivo at least 20% faster than wild-type Cas9 proteins. In some embodiments, the present disclosure provides polynucleotide sequences encoding recombinant Cas9 proteins, the recombinant Cas9 proteins comprising one or more amino acid modifications of wild-type Cas9 proteins, wherein the recombinant Cas9 proteins comprise CMA target motifs or eMI target motifs. In some embodiments, the present disclosure provides a polynucleotide sequence encoding a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, the recombinant Cas9 protein comprising an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO: 1. In some embodiments, the present disclosure provides a polynucleotide sequence encoding a recombinant Cas9 protein capable of binding to HSC70. In some embodiments, the present disclosure provides a polynucleotide sequence encoding a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, the recombinant Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO: 1.
[0186] In some embodiments, the polynucleotide sequence has at least 50% sequence identity to SEQ ID NO: 3. In some embodiments, the polynucleotide sequence has at least 60% sequence identity to SEQ ID NO: 3. In some embodiments, the polynucleotide sequence has at least 70% sequence identity to SEQ ID NO: 3. In some embodiments, the polynucleotide sequence has at least 80% sequence identity to SEQ ID NO: 3. In some embodiments, the polynucleotide sequence has at least 90% sequence identity to SEQ ID NO: 3. In some embodiments, the polynucleotide sequence has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3.
[0187] In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon optimized for expression in eukaryotic cells. In some embodiments, the polynucleotide sequence encoding stiCas9 is codon optimized for expression in animal cells. In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon optimized for expression in human cells. In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon optimized for expression in plant cells. Codon optimization is to adjust the codons to match the tRNA abundance of the expression host to improve the yield and efficiency of recombinant or heterologous protein expression. Codon optimization methods are conventional methods in the art, and software programs can be used to perform, such as integrated DNA technology company (Integrated DNA Technologies) codon optimization tools, Entelechon codon usage table analysis tools, GENEMAKER Blue Heron software, Aptagen's Gene Forge software, DNABuilder software, universal codon usage analysis software, publicly available OPTIMIZER software and Kingsray's OptimumGene algorithm.
[0188] CRISPR-Cas system
[0189] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, which includes: a recombinant Cas9 protein provided herein, and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and includes a guide sequence.
[0190] In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes an engineered KFERQ motif or a KFERQ-like motif. In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes one or more amino acid modifications of the wild-type Cas9 protein, which introduces a CMA target motif or an eMI target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes one or more amino acid modifications of the wild-type Cas9 protein, wherein the recombinant Cas9 protein includes a CMA target motif or an eMI target motif. In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is isolated from Streptococcus pyogenes (SpCas9) and includes an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO: 1. In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is capable of binding to HSC70. In some embodiments, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is isolated from Streptococcus pyogenes (SpCas9) and includes an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO: 1.
[0191] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, which includes: a polynucleotide sequence encoding a recombinant Cas9 protein provided herein, and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and includes a guide sequence.
[0192] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system, comprising: a regulatory element operably linked to a polynucleotide sequence encoding a recombinant Cas9 protein provided herein, and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0193] In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 comprising an engineered KFERQ motif or a KFERQ-like motif. In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, the modification introducing a CMA target motif or an eMI target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the recombinant Cas9 protein comprises a CMA target motif or an eMI target motif. In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, the recombinant Cas9 protein comprising an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof, of SEQ ID NO: 1. In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 protein capable of binding to HSC70. In some embodiments, the polynucleotides of the non-naturally occurring CRISPR-Cas system encode a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, the recombinant Cas9 protein comprising an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO: 1.
[0194] In some embodiments, the regulatory element connected to the polynucleotide sequence encoding the recombinant Cas9 protein is a promoter. In some embodiments, the regulatory element is a bacterial promoter. In some embodiments, the regulatory element is a viral promoter. In some embodiments, the regulatory element is a eukaryotic regulatory element, i.e., a eukaryotic promoter. In some embodiments, the eukaryotic regulatory element is a mammalian promoter.
[0195] In some embodiments, the guiding polynucleotide of the non-naturally occurring CRISPR-Cas system is an RNA molecule. The RNA molecule that binds to the CRISPR-Cas component and targets it to a specific position in the target DNA is referred to herein as a "guide RNA", "gRNA" or "small guide RNA", and may also be referred to herein as "DNA-targeting RNA". A guiding polynucleotide, such as a guiding RNA, includes at least two nucleotide segments: at least one "DNA binding segment" and at least one "polypeptide binding segment". A "segment" refers to a portion, section, or region of a molecule, for example, a continuous stretch of nucleotides of a guiding polynucleotide molecule. Unless otherwise explicitly defined, the definition of a "segment" is not limited to a specific number of total base pairs.
[0196] In some embodiments, the DNA binding segment (or "DNA targeting sequence") of the guide polynucleotide hybridizes to a target sequence in a cell. In some embodiments, the DNA binding segment of the guide polynucleotide (e.g., guide RNA) includes a polynucleotide sequence that is complementary to a specific sequence within the target DNA.
[0197] In some embodiments, the guidance polynucleotides of the present disclosure have a guidance sequence that hybridizes with a target sequence in a bacterial cell. In some embodiments of the method, the target sequence is in a bacterial cell. In some embodiments, the bacterial cell is a laboratory strain. Examples of such cells include, but are not limited to, Escherichia coli, Staphylococcus aureus, Vibrio cholerae, Streptococcus pneumoniae, Bacillus subtilis, Caulobacter crescentus, Mycoplasma genitalium, Aspergillus freudenreich, Synechocystis, Pseudomonas fluorescens, Azotobacter vinelandii, Streptomyces coelicolor. In some embodiments, the bacterial cell is a bacterium for preparing food and / or beverages. Non-limiting exemplary genera of such cells include, but are not limited to, Acetobacter, Arthrobacter, Bacillus, Bifidobacterium, Brevibacterium, Brevibacterium, Carnobacterium, Corynebacterium, Enterococcus, Gluconacetobacter, Hafnia, Halomonas, Coxsella, Lactobacillus (including acid-fasting Lactobacillus, acidophilus Lactobacillus, peptic Lactobacillus, Lactobacillus brevis, Lactobacillus buchnera, Lactobacillus casei, Lactobacillus curvatus, Lactobacillus fermentum, Lactobacillus hilgardii, Lactobacillus jensenii, Lactobacillus kimchii, Lactobacillus lactis, Lactobacillus paracasei, Lactobacillus plantarum, and Lactobacillus sakei), Leuconostoc, Microbacterium, Pediococcus, Propionibacterium, Weissella, and Zymomonas.
[0198] In some embodiments, the guidance polynucleotides disclosed herein have a guidance sequence that hybridizes with a target sequence in a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is a human or rodent or bovine cell line or cell strain. Examples of such cells / cell lines or cell strains include, but are not limited to, mouse myeloma (NSO) cell lines, Chinese hamster ovary (CHO) cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK (baby hamster kidney cells), VERO, SP2 / 0, YB2 / 0, Y0, C127, L cells, COS (e.g., COS1 and COS7), QC1-3, HEK-293, VERO, PER.C6, HeLA, EB1, EB2, EB3, oncolytic or hybridoma cell lines. In some embodiments, the eukaryotic cell is a CHO cell line. In some embodiments, the eukaryotic cell is a CHO cell. In some embodiments, the cell is a CHO-K1 cell, a CHO-K1 SV cell, a DG44 CHO cell, a DUXB11 CHO cell, a CHOS, a CHO GS knockout cell, a CHO FUT8 Gs knockout cell, a CHOZN or a CHO derived cell. CHO GS knockout cells (e.g., GSKO cells) are, for example, CHO-K1 SV GS knockout cells. CHO FUT8 knockout cells are, for example, POTELLIGENT CHOK1 SV (Lonza Biologics, Inc.). Eukaryotic cells can also be avian cells, cell lines or cell strains, such as EBX cells, EB14, EB24, EB26, EB66 or EBv13.
[0199] In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a stem cell. Stem cells can be, for example, pluripotent stem cells, including embryonic stem cells (ESC), adult stem cells, induced pluripotent stem cells (iPSC), tissue-specific stem cells (e.g., hematopoietic stem cells) and mesenchymal stem cells (MSC). In some embodiments, the human cell is a differentiated form of any cell described herein. In some embodiments, the eukaryotic cell is a cell derived from any primary cell in culture.
[0200] In some embodiments, the eukaryotic cell is a hepatocyte, such as a human hepatocyte, an animal hepatocyte, or a non-parenchymal cell. For example, the eukaryotic cell can be a culturable metabolically qualified human hepatocyte, a culturable induction qualified human hepatocyte, a culturable human hepatocyte, a suspension qualified human hepatocyte (including 10-donor and 20-donor pooled hepatocytes), a human liver Kupffer cell, a human hepatic stellate cell, a dog hepatocyte (including single and pooled beagle hepatocytes), a mouse hepatocyte (including CD-1 and C57BI / 6 hepatocytes), a rat hepatocyte (including Sprague-Dawley, Wistar Han and Wistar hepatocytes), a monkey hepatocyte (including cynomolgus monkey or rhesus monkey hepatocytes), a cat hepatocyte (including domestic shorthair cat hepatocytes) and a rabbit hepatocyte (including New Zealand white rabbit hepatocytes).
[0201] In some embodiments, the eukaryotic cell is a plant cell. For example, the plant cell can be a cell of a crop such as cassava, corn, sorghum, wheat or rice. The plant cell can be a cell of algae, trees or vegetables. The plant cell can be a cell of a monocot or a dicot, or can be a cell of a crop or a cereal plant, a production plant, a fruit or a vegetable. For example, the plant cell can be a cell of a tree, and the tree is, for example, a citrus fruit tree, such as an orange tree, a grapefruit tree or a lemon tree; a peach tree or a nectarine tree; an apple tree or a pear tree; a nut tree, such as an almond tree or a walnut tree or a pistachio tree; a nightshade, for example, a potato, a Brassica plant, a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0202] In some embodiments, the guide sequence of the guide polynucleotide is about 5 to about 50 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 6 to about 45 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 7 to about 40 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 8 to about 35 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 9 to about 30 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 10 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 12 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 14 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 16 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 18 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 5 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 6 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 7 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 8 to about 10 nucleotides. The length of the guide sequence can be determined by a skilled artisan using a guide sequence design tool, such as the CRISPR design tool (Hsu et al., Nat Biotechnol [Nature Biotechnology] 31(9):827-832 (2013)), ampliCan (Labun et al., bioRxiv 2018, doi: 10.1101 / 249474), CasFinder (Alach et al., bioRxiv 2014, doi: 10.1101 / 005074), CHOPCHOP (Labun et al., Nucleic Acids Res [Nucleic Acids Research] 2016, doi: 10.1093 / nar / gkw398), etc.
[0203] In some embodiments, the guidance polynucleotide (e.g., guide RNA) disclosed herein includes a polypeptide binding sequence / segment. The polypeptide binding segment (or "protein binding sequence") of the guidance polynucleotide (e.g., guide RNA) interacts with the polynucleotide binding domain of the Cas protein disclosed herein. Such polypeptide binding segments or sequences are known to those skilled in the art, for example, those disclosed in U.S. Patents 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899 and 2015 / 0071906, and the disclosures disclosed herein are incorporated herein in their entirety. In some embodiments, the polypeptide binding segment of the guidance polynucleotide binds to Cas9. In some embodiments, the polypeptide binding segment of a guide polynucleotide binds to a recombinant Cas9 protein provided herein.
[0204] In some embodiments, the guide polynucleotide is at least about 10, 15, 20, 25 or 30 nucleotides and at most about 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 nucleotides. In some embodiments, the guide polynucleotide is about 10 to about 150 nucleotides. In some embodiments, the guide polynucleotide is about 20 to about 120 nucleotides. In some embodiments, the guide polynucleotide is about 30 to about 100 nucleotides. In some embodiments, the guide polynucleotide is about 40 to about 80 nucleotides. In some embodiments, the guide polynucleotide is about 50 to about 60 nucleotides. In some embodiments, the guide polynucleotide is about 10 to about 35 nucleotides. In some embodiments, the guide polynucleotide is about 15 to about 30 nucleotides. In some embodiments, the guide polynucleotide is about 20 to about 25 nucleotides.
[0205] A guide polynucleotide (e.g., a guide RNA) can be introduced into a target cell as a separate molecule (e.g., an RNA molecule) or introduced into the cell using an expression vector comprising DNA encoding the guide polynucleotide (e.g., a guide RNA).
[0206] In some embodiments, the guiding polynucleotide of the CRISPR-Cas system is connected to the same direction repeat sequence. The same direction repeat sequence or DR sequence is a repetitive sequence array in the CRISPR locus, separated by a non-repetitive sequence (inter-region sequence) of a short stretch. The inter-region sequence targets the pre-inter-region sequence adjacent to the motif (PAM) on the target sequence. When the non-coding portion of the CRISPR locus (i.e., guiding polynucleotides and tracrRNA) is transcribed, the transcript is cleaved into multiple short crRNAs on the DR sequence, and these crRNAs include a single inter-region sequence, which guides the Cas9 nuclease to PAM. In some embodiments, the DR sequence is RNA. In some embodiments, the DR sequence is encoded by nucleic acid. In some embodiments, the DR sequence is connected to the guiding polynucleotide. In some embodiments, the DR sequence is connected to the guiding sequence of the guiding polynucleotide. In some embodiments, the DR sequence includes a secondary structure. In some embodiments, the DR sequence includes a stem-loop structure. In some embodiments, the DR sequence is 10 to 20 nucleotides. In some embodiments, the DR sequence is at least 16 nucleotides. In some embodiments, the DR sequence is at least 16 nucleotides and includes a single stem-loop. In some embodiments, the DR sequence comprises an RNA aptamer. In some embodiments, the secondary structure or stem loop in the DR is recognized by a nuclease for cleavage. In some embodiments, the nuclease is a ribonuclease. In some embodiments, the nuclease is an RNase III.
[0207] In some embodiments, the CRISPR-Cas system of the present disclosure further includes tracrRNA. "TracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with a precursor crRNA or a precursor CRISPR-RNA, which is then cleaved by an RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide RNA includes a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide RNA activates the Cas9 protein. In some embodiments, the guiding polynucleotide of the CRISPR-cas system includes a tracrRNA sequence. In some embodiments, the CRISPR-Cas system includes a separate polynucleotide including a tracrRNA sequence.
[0208] In some embodiments, the polynucleotide encoding recombinant Cas9 and the guide polynucleotide are on a single vector. In some embodiments, the polynucleotide encoding recombinant Cas9, the guide polynucleotide (or the nucleotide that can be transcribed into the guide polynucleotide) and the tracrRNA are on a single vector. In some embodiments, the polynucleotide encoding recombinant Cas9, the guide polynucleotide (or the nucleotide that can be transcribed into the guide polynucleotide), the tracrRNA and the same direction repeat sequence are on a single vector. In some embodiments, the vector is an expression vector. In some embodiments, the vector is a mammalian expression vector. In some embodiments, the vector is a human expression vector. In some embodiments, the vector is a plant expression vector.
[0209] In some embodiments, the polynucleotide encoding the recombinant Cas9 and the guide polynucleotide is a single nucleic acid molecule. In some embodiments, the polynucleotide encoding the recombinant Cas9, the guide polynucleotide and the tracrRNA is a single nucleic acid molecule. In some embodiments, the polynucleotide encoding the recombinant Cas9, the guide polynucleotide, the tracrRNA and the direct repeat sequence is a single nucleic acid molecule. In some embodiments, the single nucleic acid molecule is an expression vector. In some embodiments, the single nucleic acid molecule is a mammalian expression vector. In some embodiments, the single nucleic acid molecule is a human expression vector. In some embodiments, the single nucleic acid molecule is a plant expression vector.
[0210] In some embodiments, the recombinant Cas9 and the guide polynucleotide are capable of forming a complex. In some embodiments, the complex of the recombinant Cas9 and the guide polynucleotide does not exist in nature.
[0211] Various methods for delivering CRISPR-Cas systems are known in the art. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered by delivery particles. Delivery particles are biological delivery systems or formulations comprising particles. As defined herein, a "particle" is an entity with a maximum diameter of about 100 microns (μm). In some embodiments, the maximum diameter of the particle is about 10 μm. In some embodiments, the maximum diameter of the particle is about 2000 nanometers (nm). In some embodiments, the maximum diameter of the particle is about 1000nm. In some embodiments, the maximum diameter of the particle is about 900nm, about 800nm, about 700nm, about 600nm, about 500nm, about 400nm, about 300nm, about 200nm or about 100nm. In some embodiments, the diameter of the particle is about 25nm to about 200nm. In some embodiments, the diameter of the particle is about 50nm to about 150nm. In some embodiments, the diameter of the particle is about 75nm to about 100nm.
[0212] The delivery particles can be provided in any form, including but not limited to: solid, semi-solid, emulsion or colloidal particles. In some embodiments, the delivery particles are lipid-based systems, liposomes, micelles, microvesicles, exosomes or gene guns. In some embodiments, the delivery particles include a CRISPR-Cas system. In some embodiments, the delivery particles include a CRISPR-Cas system, which includes a recombinant Cas9 and a guide polynucleotide. In some embodiments, the delivery particles include a CRISPR-Cas system, which includes a recombinant Cas9 and a guide polynucleotide, wherein the recombinant Cas9 and the guide polynucleotide exist as a complex. In some embodiments, the delivery particles include a CRISPR-Cas system, which includes a recombinant Cas9, a guide polynucleotide and a polynucleotide containing tracrRNA. In some embodiments, the delivery particles include a CRISPR-Cas system, which includes a recombinant Cas9, a guide polynucleotide and a tracrRNA.
[0213] In some embodiments, the delivery particle further comprises a lipid, a sugar, a metal or a protein. In some embodiments, the delivery particle is a lipid envelope. For example, Su et al., Molecular Pharmacology [Molecular Pharmacology] 8 (3): 774-784 (2011) describes mRNA delivery using a lipid envelope or a delivery particle comprising lipids. In some embodiments, the delivery particle is a sugar-based particle, for example, GalNAc. Sugar-based particles are described in WO 2014 / 118272 and Nair et al., J Am Chem Soc [Journal of the American Chemical Society] 136 (49): 16958-16961 (2014).
[0214] In some embodiments, the delivery particle is a nanoparticle. The nanoparticles covered by the present disclosure can be provided in different forms, for example, as solid nanoparticles (e.g., metals, such as silver, gold, iron, titanium), non-metals, lipid-based solids, polymers, nanoparticles or combinations thereof. Metals, dielectrics and semiconductor nanoparticles and hybrid structures (e.g., core-shell nanoparticles) can be prepared. If the nanoparticles made of semiconductor materials are small enough (typically less than 10nm), the electronic energy levels can be quantified, and then they can also be labeled as quantum dots. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications, and can be adjusted to be suitable for similar uses in the present disclosure.
[0215] Preparation of delivery particles is further described in U.S. Patent Publication Nos. 2011 / 0293703, 2012 / 0251560, and 2013 / 0302401; and U.S. Patent Nos. 5,543,158, 5,855,913, 5,895,309, 6,007,845, and 8,709,843.
[0216] In some embodiments, the vesicle includes the CRISPR-Cas system of the present disclosure. A "vesicle" is a small structure with a fluid surrounded by a lipid bilayer in a cell. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered by a vesicle. In some embodiments, the vesicle includes recombinant Cas9 and a guide polynucleotide. In some embodiments, the vesicle includes recombinant Cas9 and a guide polynucleotide, wherein the recombinant Cas9 and the guide polynucleotide exist as a complex. In some embodiments, the vesicle includes a CRISPR-Cas system, and the CRISPR-Cas system includes recombinant Cas9, a guide polynucleotide, and a polynucleotide containing tracrRNA. In some embodiments, the vesicle includes a CRISPR-Cas system, and the CRISPR-Cas system includes recombinant Cas9, a guide polynucleotide, and tracrRNA.
[0217] In some embodiments, the vesicles comprising recombinant Cas9 and guide polynucleotides are exosomes or liposomes. In some embodiments, the vesicles are exosomes. In some embodiments, the exosomes are used to deliver the CRISPR-Cas system disclosed herein. Exosomes are endogenous nanovesicles (i.e., having a diameter of about 30 nm to about 100 nm) that can transport RNA and proteins, and can deliver RNA to the brain and other target organs. For example, Alvarez-Erviti et al., Nature Biotechnology [Natural Biology] 29: 341 (2011), El-Andaloussi et al., Nature Protocols [Natural Experiment Manual] 7: 2112-2116 (2012), and Wahlgren et al., Nucleic Acids Research [Nucleic Acids Research] 40 (17): e130 (2012) describe engineered exosomes for delivering endogenous biological materials to target organs.
[0218] In some embodiments, the vesicles including stiCas9 and guide polynucleotides are liposomes. In some embodiments, the liposomes are used to deliver the CRISPR-Cas system disclosed herein. Liposomes are spherical vesicle structures having at least one lipid bilayer and can be used as vehicles for nutrient and drug administration. Liposomes are generally composed of phospholipids (particularly phosphatidylcholine) and other lipids (such as egg phosphatidylethanolamine). The types of liposomes include, but are not limited to, multilamellar vesicles, small unilamellar vesicles, large unilamellar vesicles, and cochlear vesicles. See, for example, Spuch and Navarro, Journal of Drug Delivery, article number 469679 (2011). For example, Morrissey et al., Nature Biotechnology 23(8):1002-1007 (2005), Zimmerman et al., Nature Letters 441:111-114 (2006), and Li et al., Gene Therapy 19:775-780 (2012) describe liposomes for delivering biological materials such as CRISPR-Cas components.
[0219] In some embodiments, the viral vector includes the CRISPR-Cas system of the present disclosure. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered by a viral vector. In some embodiments, the viral vector includes recombinant Cas9 and a guide polynucleotide. In some embodiments, the viral vector includes recombinant Cas9 and a guide polynucleotide, wherein the recombinant Cas9 and the guide polynucleotide exist as a complex. In some embodiments, the viral vector includes a CRISPR-Cas system, which includes recombinant Cas9, a guide polynucleotide, and a polynucleotide containing tracrRNA. In some embodiments, the viral vector includes a CRISPR-Cas system, which includes recombinant Cas9, a guide polynucleotide, and tracrRNA. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus vector. Examples of viral vectors are provided herein.
[0220] In some embodiments, adeno-associated virus (AAV) and / or lentiviral vectors can be used as viral vectors comprising elements of the CRISPR-Cas system described herein. In some embodiments of the present disclosure, the Cas protein is expressed intracellularly by cells transduced by a viral vector.
[0221] In some embodiments, the Cas proteins and methods disclosed herein are used for ex vivo gene editing, such as CAR-T type therapies. These embodiments may involve modification of cells from human donors. In these cases, viral vectors may also be used; however, there are other options for directly transfecting Cas9 proteins (along with in vitro transcribed guide RNA and donor DNA) into cultured cells.
[0222] In some embodiments, the recombinant Cas9 protein of the present disclosure is a part of a fusion protein including one or more heterologous protein domains (e.g., about or at least about 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or more domains in addition to the recombinant Cas9 protein). The Cas9 fusion protein may include any other protein sequence, and optionally a linker sequence between any two domains. Examples of protein domains that can be fused to the recombinant Cas9 protein include, but are not limited to, epitope tags, reporter gene sequences, and protein domains having one or more of the following activities: methylase activity, demethylase activity, transcription activation activity, transcription inhibition activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include: histidine (His) tags, V5 tags, FLAG tags, influenza virus hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), autofluorescent protein (including blue fluorescent protein (BFP)) and mCherry. In some embodiments, the recombinant Cas9 protein is fused to a protein or protein fragment that binds to a DNA molecule or binds to other cellular molecules, including but not limited to: maltose binding protein (MBP), S tag, Lex A DNA binding domain (DBD), GAL4 DNA binding domain and herpes simplex virus (HSV) BP16 protein. Other domains that may form part of a fusion protein including a Cas9 protein are described in U.S. Patent Publication 2011 / 0059502. In some embodiments, the tagged recombinant Cas9 protein is used to identify the location of the target sequence.
[0223] In some embodiments, the recombinant Cas9 protein can form a component of an inducible system. The inducible nature of the system allows the use of a certain form of energy to control gene editing or gene expression in time and space. This form of energy may include, but is not limited to, electromagnetic radiation, acoustic energy, chemical energy, and thermal energy. Non-limiting examples of inducible systems include: tetracycline inducible promoters (Tet-On or Tet-Off), small molecule double hybrid transcription activation systems (FKBP, ABA, etc.), or light-induced systems (phytochrome, LOV domains, or cryptochromes). In some embodiments, the Cas9 protein is part of a light-induced transcription effector (LITE) that guides changes in transcriptional activity in a sequence-specific manner. The components of light may include Cas9 protein, light-responsive cytochrome heterodimers (e.g., from Arabidopsis thaliana) and transcriptional activation / repression domains. Other examples of inducible DNA binding proteins and methods of using the same are provided in International Application Publication Nos. WO 2014 / 018423 and WO 2014 / 093635; U.S. Patent Nos. 8,889,418 and 8,895,308; and U.S. Patent Publication Nos. 2014 / 0186919, 2014 / 0242700, 2014 / 0273234, and 2014 / 0335620;
[0224] Site-specific modification methods
[0225] In some embodiments, the present disclosure provides a method of providing a site-specific modification at a target sequence in the genome of a cell, the method comprising introducing a non-naturally occurring CRISPR-Cas system described herein into the cell.
[0226] In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes an engineered KFERQ motif or a KFERQ-like motif. In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes one or more amino acid modifications of the wild-type Cas9 protein, which introduces a CMA target motif or an eMI target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif. In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system includes one or more amino acid modifications of the wild-type Cas9 protein, wherein the recombinant Cas9 protein includes a CMA target motif or an eMI target motif. In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is isolated from Streptococcus pyogenes (SpCas9) and includes an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO:1. In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is capable of binding to HSC70. In some embodiments of the method, the recombinant Cas9 protein of the non-naturally occurring CRISPR-Cas system is isolated from Streptococcus pyogenes (SpCas9) and includes an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO:1.
[0227] Modifications of the target sequence encompass single nucleotide substitutions, multiple nucleotide substitutions, insertions (ie, knock-ins) and deletions (ie, knock-outs) of nucleic acids, frameshift mutations, and other nucleic acid modifications.
[0228] In some embodiments of the method, the modification is a deletion of at least a portion of the target sequence. The target sequence can be cleaved at two different sites and produce complementary sticky ends, and these complementary sticky ends can be reconnected to remove the sequence portion between the two sites.
[0229] In some embodiments of the method, the modification is a mutation of the target sequence. Site-specific mutagenesis can be achieved by using a site-specific nuclease that promotes homologous recombination of an exogenous polynucleotide template (also referred to as a "donor polynucleotide" or "donor vector") containing the desired mutation. In some embodiments, the sequence of interest (SoI) includes the desired mutation.
[0230] In some embodiments of the method, the modification is the insertion of a sequence of interest (SoI) into the target sequence. The SoI can be introduced as an exogenous polynucleotide template. In some embodiments, the exogenous polynucleotide comprises a blunt end. In some embodiments, the exogenous polynucleotide template comprises a sticky end. In some embodiments, the exogenous polynucleotide template comprises a sticky end that is complementary to a sticky end in the target sequence.
[0231] The exogenous polynucleotide template can have any suitable length, such as about or at least about 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 500 or 1000 or more nucleotides in length. In some embodiments, the exogenous polynucleotide template is complementary to a portion of a polynucleotide comprising a target sequence. When optimally aligned, the exogenous polynucleotide template overlaps one or more nucleotides (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 or more nucleotides) of the target sequence. In some embodiments, when the exogenous polynucleotide template and the polynucleotide comprising the target sequence are optimally aligned, the closest nucleotide of the exogenous polynucleotide template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides from the target sequence.
[0232] In some embodiments, the exogenous polynucleotide is DNA, e.g., a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear fragment of single-stranded or double-stranded DNA, an oligonucleotide, a PCR fragment, a naked nucleic acid, or a nucleic acid complexed with a delivery vehicle such as a liposome.
[0233] In some embodiments, the exogenous polynucleotide is inserted into the target sequence using the cell's endogenous DNA repair pathways. Endogenous DNA repair pathways include NHEJ, MMEJ, and HDR, each of which is described herein. During the repair process, an exogenous polynucleotide template including a Sol can be introduced into the target sequence. In some embodiments, an exogenous polynucleotide template including a Sol flanked by an upstream sequence and a downstream sequence is introduced into the cell, wherein the upstream and downstream sequences have sequence similarity to either side of the integration site in the target sequence. In some embodiments, the exogenous polynucleotide including the Sol includes, for example, a mutant gene. In some embodiments, the exogenous polynucleotide includes a sequence that is endogenous or exogenous to the cell. In some embodiments, the Sol includes a polynucleotide encoding a protein, or a non-coding sequence, such as, for example, a microRNA. In some embodiments, the Sol is operably linked to a regulatory element. In some embodiments, the Sol is a regulatory element. In some embodiments, the Sol includes a resistance cassette, for example, a gene that confers resistance to an antibiotic. In some embodiments, the Sol includes a mutation of a wild-type target sequence. In some embodiments, the Sol disrupts or corrects the target sequence by generating a frameshift mutation or a nucleotide substitution. In some embodiments, the Sol includes a marker. Introducing a marker into the target sequence can facilitate screening for targeted integration. In some embodiments, the marker is a restriction site, a fluorescent protein, or a selectable marker. In some embodiments, the SoI is introduced as a vector comprising the SoI.
[0234] The upstream and downstream sequences in the exogenous polynucleotide template are selected to promote homologous recombination between the target sequence and the exogenous polynucleotide. The upstream sequence is a nucleic acid sequence having sequence similarity to the upstream sequence of the targeted site for integration (target sequence). Similarly, the downstream sequence is a nucleic acid sequence having sequence similarity to the downstream sequence of the target site for integration. Thus, in some embodiments, the exogenous polynucleotide template comprising the SoI is inserted into the target sequence by homologous recombination at the upstream and downstream sequences. In some embodiments, the upstream and downstream sequences in the exogenous polynucleotide template have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the upstream and downstream sequences of the targeted genomic sequence, respectively. In some embodiments, the upstream or downstream sequence has at least about 20, 50, 100, 150, 200, 250, 300, 350, 400 or 500 base pairs and at most about 600, 750, 1000, 1250, 1500, 1750 or 2000 base pairs. In some embodiments, the upstream or downstream sequence has about 20 to 2000 base pairs, or about 50 to 1750 base pairs, or about 100 to 1500 base pairs, or about 200 to 1250 base pairs, or about 300 to 1000 base pairs, or about 400 to about 750 base pairs, or about 500 to 600 base pairs. In some embodiments, the upstream or downstream sequence has about 50, about 100, about 250, about 500, about 100, about 1250, about 1500, about 1750, about 2000, about 2250, or about 2500 base pairs.
[0235] In some embodiments of the method, the modification in the target sequence is the inactivation of expression of the target sequence in the cell. For example, after the CRISPR-Cas complex binds to the target sequence, the target sequence is inactivated so that the sequence is not transcribed, the encoded protein is not produced, and / or the sequence does not function like the wild-type sequence. For example, a protein or microRNA coding sequence may be inactivated so that no protein is produced.
[0236] In some embodiments, the regulatory sequence may be inactivated so that it no longer functions as a regulatory sequence. Examples of regulatory sequences include promoters, transcription terminators, enhancers, and other regulatory elements described herein. The inactivated target sequence may include a deletion mutation (i.e., a deletion of one or more nucleotides), an insertion mutation (i.e., an insertion of one or more nucleotides), or a nonsense mutation (i.e., replacing a single nucleotide with another nucleotide to introduce a stop codon). In some embodiments, the inactivation of the target sequence results in a "knockout" of the target sequence.
[0237] In some embodiments of the method comprising the recombinant Cas9 provided herein, the off-target modification in the cell genome is reduced by at least about 50% relative to wild-type Cas9 or Cas9 that does not include a KFERQ motif or a KFERQ-like motif. As described herein, off-target modifications are non-specific and unexpected genetic modifications, such as unexpected point mutations, deletions, insertions, inversions, and translocations. In some embodiments, the recombinant Cas9 protein disclosed herein has reduced off-targets in cells due to a faster degradation rate. In some embodiments, the recombinant Cas9 protein disclosed herein has reduced off-targets in cells due to the lower availability of cells to Cas9. In some embodiments, the recombinant Cas9 protein disclosed herein has reduced off-targets in cells due to the shorter exposure time of cells to Cas9.
[0238] In some embodiments of the method comprising a recombinant Cas9 provided herein, off-target modifications are reduced relative to wild-type Cas9, and on-target modifications are at least approximately the same level. In some embodiments, the on-target modification of the recombinant Cas9 is at least about 20%, at least about 15%, at least about 10%, at least about 5%, at least about 4%, at least about 3%, at least about 2%, at least about 1%, at least about 0.5% of the on-target modification of the wild-type Cas9. In some embodiments comprising a recombinant Cas9 provided herein, off-target modifications are reduced relative to wild-type Cas9, and on-target modifications are increased. In some embodiments, the on-target modification of the recombinant Cas9 is at least about 5%, at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about 15%, at least about 16%, at least about 17%, at least about 18%, at least about 19%, or at least about 20% higher than that of the wild-type Cas9.
[0239] In some embodiments, off-target modifications in the genome of the cell are reduced by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 150%, or at least about 200% relative to wild-type Cas9 or Cas9 that does not include a KFERQ motif or a KFERQ-like motif.
[0240] In some embodiments of the methods comprising a recombinant Cas9 provided herein, off-target modifications in the genome of the cell are reduced by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold, at least about 500-fold, or at least about 1000-fold relative to wild-type Cas9 or Cas9 that does not comprise a KFERQ motif or a KFERQ-like motif.
[0241] In some embodiments of the method, the off-target modification in the cell genome is less than about 5% of all modifications in the genome produced by the recombinant Cas9 with a KFERQ motif or a KFERQ-like motif. In some embodiments of the method, the off-target modification in the cell genome is less than about 2% of all modifications in the genome produced by the recombinant Cas9 with a KFERQ motif or a KFERQ-like motif. In some embodiments of the method, the off-target modification in the cell genome is less than about 1% of all modifications in the genome produced by the recombinant Cas9 with a KFERQ motif or a KFERQ-like motif. As described herein, the off-target modification of wild-type Cas9 can be at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% of all modifications in the genome of wild-type Cas9. In some embodiments, the off-target modification of the recombinant Cas9 with a KFERQ motif or a KFERQ-like motif is less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1.5%, less than about 1%, less than about 0.5%, or less than about 0.1% of all modifications in the genome produced by the recombinant Cas9 with a KFERQ motif or a KFERQ-like motif. The amount of off-target modification can vary depending on the sequence of the guide polynucleotide and the target genomic locus. Typically, when the same guide polynucleotide is used, the recombinant Cas9 protein with a KFERQ motif or a KFERQ-like motif has reduced off-target modifications compared to wild-type Cas9.
[0242] In some embodiments of the method, the target sequence is in a bacterial cell. In some embodiments, the bacterial cell is a laboratory strain. Examples of such cells include, but are not limited to, Escherichia coli, Staphylococcus aureus, Vibrio cholerae, Streptococcus pneumoniae, Bacillus subtilis, Caulobacter crescentus, Mycoplasma genitalium, Aspergillus freundii, Synechocystis, Pseudomonas fluorescens, Azotobacter vinelandii, Streptomyces coelicolor. In some embodiments, the bacterial cell is a bacterium used to prepare food and / or beverages. Non-limiting exemplary genera of such cells include, but are not limited to, Acetobacter, Arthrobacter, Bacillus, Bifidobacterium, Brevibacterium, Brevibacterium, Carnobacterium, Corynebacterium, Enterococcus, Gluconacetobacter, Hafnia, Halomonas, Coxsella, Lactobacillus (including acid-fasting Lactobacillus, acidophilus Lactobacillus, peptic Lactobacillus, brevis, buchnera Lactobacillus, casei Lactobacillus, curvature Lactobacillus, fermentum Lactobacillus, hilarix, jensenii, kingellus, Lactobacillus lactis, Lactobacillus paracasei, Lactobacillus plantarum, and Lactobacillus sakazakii), Leuconostoc, Microbacterium, Pediococcus, Propionibacterium, Weissella, and Zymomonas.
[0243] In some embodiments of the method, the target cell is in a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is a human or rodent or cattle cell line or cell strain. Examples of such cells / cell lines or cell strains include, but are not limited to, mouse myeloma (NSO) cell lines, Chinese hamster ovary (CHO) cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK (baby hamster kidney cells), VERO, SP2 / 0, YB2 / 0, Y0, C127, L cells, COS (e.g., COS1 and COS7), QC1-3, HEK-293, VERO, PER.C6, HeLA, EB1, EB2, EB3, oncolytic or hybridoma cell lines. In some embodiments, the eukaryotic cell is a CHO cell line. In some embodiments, the eukaryotic cell is a CHO cell. In some embodiments, the cell is a CHO-K1 cell, a CHO-K1 SV cell, a DG44 CHO cell, a DUXB11 CHO cell, a CHOS, a CHO GS knockout cell, a CHO FUT8 Gs knockout cell, a CHOZN or a CHO derived cell. CHO GS knockout cells (e.g., GSKO cells) are, for example, CHO-K1 SV GS knockout cells. CHOFUT8 knockout cells are, for example, POTELLIGENT CHOK1 SV (Lonza Biotech). Eukaryotic cells can also be avian cells, cell lines or cell strains, such as EBX cells, EB14, EB24, EB26, EB66 or EBv13.
[0244] In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a stem cell. Stem cells can be, for example, pluripotent stem cells, including embryonic stem cells (ESC), adult stem cells, induced pluripotent stem cells (iPSC), tissue-specific stem cells (e.g., hematopoietic stem cells) and mesenchymal stem cells (MSC). In some embodiments, the cell is a pluripotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell. In some embodiments, the human cell is a differentiated form of any cell described herein. In some embodiments, the eukaryotic cell is a cell derived from any primary cell in culture.
[0245] In some embodiments, the eukaryotic cell is a hepatocyte, such as a human hepatocyte, an animal hepatocyte, or a non-parenchymal cell. For example, the eukaryotic cell can be a culturable metabolically qualified human hepatocyte, a culturable induction qualified human hepatocyte, a culturable human hepatocyte, a suspension qualified human hepatocyte (including 10-donor and 20-donor pooled hepatocytes), a human liver Kupffer cell, a human hepatic stellate cell, a dog hepatocyte (including single and pooled beagle hepatocytes), a mouse hepatocyte (including CD-1 and C57BI / 6 hepatocytes), a rat hepatocyte (including Sprague-Dawley, Wistar Han and Wistar hepatocytes), a monkey hepatocyte (including cynomolgus monkey or rhesus monkey hepatocytes), a cat hepatocyte (including domestic shorthair cat hepatocytes) and a rabbit hepatocyte (including New Zealand white rabbit hepatocytes).
[0246] In some embodiments, the eukaryotic cell is a plant cell. For example, the plant cell can be a cell of a crop such as cassava, corn, sorghum, wheat or rice. The plant cell can be a cell of algae, trees or vegetables. The plant cell can be a cell of a monocot or a dicot, or can be a cell of a crop or a cereal plant, a production plant, a fruit or a vegetable. For example, the plant cell can be a cell of a tree, and the tree is, for example, a citrus fruit tree, such as an orange tree, a grapefruit tree or a lemon tree; a peach tree or a nectarine tree; an apple tree or a pear tree; a nut tree, such as an almond tree or a walnut tree or a pistachio tree; a nightshade, for example, a potato, a Brassica plant, a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0247] In some embodiments of the method, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in the cell. In some embodiments, the DNA binding segment of the instructing polynucleotide hybridizes with the target sequence in the cell. In some embodiments, the DNA binding segment of the instructing polynucleotide (e.g., guide RNA) includes a polynucleotide sequence complementary to the specific sequence in the target DNA. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in bacterial cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in eukaryotic cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in mammalian cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in human cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in pluripotent stem cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in induced pluripotent stem cells. In some embodiments, the guide sequence of the instructing polynucleotide can hybridize with the target sequence in plant cells.
[0248] In some embodiments, the guide sequence of the guide polynucleotide is at least about 5, 6, 7, 8, 9, 10, 12, 14, 16, 18 or 20 nucleotides and at most about 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 5 to about 50 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 6 to about 45 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 7 to about 40 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 8 to about 35 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 9 to about 30 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 10 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 12 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 14 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 16 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 18 to about 20 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 5 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 6 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 7 to about 10 nucleotides. In some embodiments, the guide sequence of the guide polynucleotide is about 8 to about 10 nucleotides. The length of the guide sequence can be determined by one of skill in the art using the guide sequence design tools described herein.
[0249] In some embodiments of the method, the CRISPR-Cas system is introduced into the cell via delivery particles, vesicles or viral vectors. In some embodiments of the method, the CRISPR-Cas system including recombinant Cas9 and instructing polynucleotides is introduced into the cell via delivery particles. In some embodiments of the method, the CRISPR-Cas system including recombinant Cas9 and instructing polynucleotides is introduced into the cell via vesicles. In some embodiments of the method, the CRISPR-Cas system including recombinant Cas9 and instructing polynucleotides is introduced into the cell via a vector. In some embodiments of the method, the CRISPR-Cas system including recombinant Cas9 and instructing polynucleotides is introduced into the cell via a viral vector. In some embodiments of the method, the polynucleotides encoding the components of the complex including recombinant Cas9 and instructing polynucleotides are introduced into one or more vectors. Examples of delivery particles, vesicles, vectors, viral vectors and methods (such as transfection of vectors) delivered to cells are provided herein.
[0250] All references cited herein, including patents, patent applications, articles, textbooks, etc. and references cited therein (to the same extent as if they had not already been cited), are hereby incorporated by reference in their entirety.
[0251] Other Exemplary Embodiments
[0252] Example 1 is a recombinant Cas9 protein comprising an engineered KFERQ motif or a KFERQ-like motif.
[0253] Embodiment 2 includes the recombinant Cas9 protein of embodiment 1, wherein the engineered KFERQ motif or KFERQ-like motif is selected from KFERQ (SEQ ID NO:24), RKVEQ (SEQ ID NO:25), QDLKF (SEQ ID NO:26), QRFFE (SEQ ID NO:27), NRVVD (SEQ ID NO:28), QRDKV (SEQ ID NO:29), QKILD (SEQ ID NO:30), QKKEL (SEQ ID NO:31), QFREL (SEQ ID NO:32), IKLDQ (SEQ ID NO:33), DVVRQ (SEQ ID NO:34), QRIVE (SEQ ID NO:35), VKELQ (SEQ ID NO:36), QKVFD (SEQ ID NO:37), QELLR (SEQ ID NO:38), VDKLN (SEQ ID NO:39), RIKEN (SEQ ID NO:40), NKKFE (SEQ ID NO:41), and combinations thereof.
[0254] Embodiment 3 includes the recombinant Cas9 protein of embodiment 1 or 2, wherein the engineered KFERQ-like motif is VDKLN (SEQ ID NO: 39).
[0255] Embodiment 4 includes the recombinant Cas9 protein of embodiment 1, wherein the engineered KFERQ motif or KFERQ-like motif is in the REC lobe of the Cas9 protein.
[0256] Embodiment 5 includes the recombinant Cas9 protein of embodiment 2, wherein the engineered KFERQ motif or KFERQ-like motif is in the Rec2 domain of the REC lobe.
[0257] Embodiment 6 includes the recombinant Cas9 protein of embodiment 1, wherein the engineered KFERQ motif or KFERQ-like motif is in the HNH domain, RuvC domain or PI domain of the recombinant Cas9 protein.
[0258] Embodiment 7 includes the recombinant Cas9 protein of any one of embodiments 1 to 4, wherein the engineered KFERQ motif or KFERQ-like motif is in a surface-exposed region of the recombinant Cas9 protein.
[0259] Embodiment 8 includes the recombinant Cas9 protein of any one of embodiments 1 to 6, wherein the engineered KFERQ motif or KFERQ-like motif is at the N-terminus or C-terminus of the recombinant Cas9 protein.
[0260] Embodiment 9 is a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, which modifications introduce a chaperone-mediated autophagy (CMA) target motif or an endosomal microautophagy (eMI) target motif into the Cas9 protein, wherein the recombinant Cas9 protein is degraded in vivo by at least 20% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif.
[0261] Embodiment 10 includes the recombinant Cas9 protein of embodiment 9, wherein the recombinant Cas9 protein is degraded in vivo at least 50% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif.
[0262] Embodiment 11 includes the recombinant Cas9 protein of embodiment 9 or 10, wherein the recombinant Cas9 protein is degraded in vivo at least 80% faster than the wild-type Cas9 protein or the Cas9 protein that does not include the CMA or eMI target motif.
[0263] Embodiment 12 is a recombinant Cas9 protein comprising one or more amino acid modifications of a wild-type Cas9 protein, wherein the recombinant Cas9 protein comprises a CMA target motif or an eMI target motif.
[0264] Embodiment 13 includes the recombinant Cas9 protein of any one of embodiments 9 to 12, wherein the CMA target motif or the eMI target motif is selected from KFERQ (SEQ ID NO:24), RKVEQ (SEQ ID NO:25), QDLKF (SEQ ID NO:26), QRFFE (SEQ ID NO:27), NRVVD (SEQ ID NO:28), QRDKV (SEQ ID NO:29), QKILD (SEQ ID NO:30), QKKEL (SEQ ID NO:31), QFREL (SEQ ID NO:32), IKLDQ (SEQ ID NO:33), DVVRQ (SEQ ID NO:34), QRIVE (SEQ ID NO:35), VKELQ (SEQ ID NO:36), QKVFD (SEQ ID NO:37), QELLR (SEQ ID NO:38), VDKLN (SEQ ID NO:39), RIKEN (SEQ ID NO:40), NKKFE (SEQ ID NO:41), and combinations thereof.
[0265] Embodiment 14 includes the recombinant Cas9 protein of embodiment 13, wherein the CMA target motif or eMI target motif is VDKLN (SEQ ID NO: 39).
[0266] Embodiment 15 includes the recombinant Cas9 protein of any one of embodiments 9 to 14, wherein the one or more amino acid substitutions are in a surface-exposed region of the recombinant Cas9 protein.
[0267] Embodiment 16 comprises a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, which recombinant Cas9 protein comprises an amino acid modification at one or more of positions F185, A547, I548, T560, V561, D829, I830, L1087, S1088, P1199, K1200, or a combination thereof of SEQ ID NO:1.
[0268] Embodiment 17 includes the recombinant Cas9 protein of any one of embodiments 9 to 16, wherein the amino acid modification includes one or more of the following mutations: F185N; A547E / I548L; T560E / V561Q; D829L / I830R; L1087E / S1088Q; or P1199D / K1200Q.
[0269] Embodiment 18 includes the recombinant Cas9 protein of any one of embodiments 9 to 17, wherein the amino acid modification is a mutation at F185.
[0270] Embodiment 19 includes the recombinant Cas9 protein of embodiment 18, wherein the mutation is F185N.
[0271] Embodiment 20 includes the recombinant Cas9 protein of any one of embodiments 16 to 19, wherein the amino acid modification results in a CMA target motif or an eMI target motif.
[0272] Embodiment 21 includes the recombinant Cas9 protein of any one of embodiments 9 to 20, wherein the recombinant Cas9 protein has at least 90% identity with SEQ ID NO:1.
[0273] Example 22 is a recombinant Cas9 protein capable of binding to 70kD heat shock cognate protein (HSC70).
[0274] Example 23 is a recombinant Cas9 protein (SpCas9) isolated from Streptococcus pyogenes, which recombinant Cas9 protein includes an engineered KFERQ motif or KFERQ-like motif at amino acid position 185 of SEQ ID NO:1.
[0275] Embodiment 24 includes the recombinant Cas9 protein of embodiment 23, wherein the KFERQ-like motif is VDKLN (SEQ ID NO: 39).
[0276] Embodiment 25 includes the recombinant Cas9 protein of any one of embodiments 1 to 24, further comprising a mutation at position D10, H840, or a combination thereof in SEQ ID NO: 1.
[0277] Embodiment 26 includes the recombinant Cas9 protein of embodiment 25, wherein the mutation is selected from D10A or D10N; H840A, H840N or H840Y; and combinations thereof.
[0278] Embodiment 27 includes the recombinant Cas9 protein of any one of embodiments 1 to 26, wherein the recombinant Cas9 protein produces sticky ends.
[0279] Embodiment 28 includes the recombinant Cas9 protein of any one of embodiments 1 to 27, further comprising one or more nuclear localization signals.
[0280] Embodiment 29 is a polynucleotide sequence encoding the recombinant Cas9 protein of any one of embodiments 1 to 28.
[0281] Embodiment 30 comprises the polynucleotide sequence of embodiment 29, wherein the polynucleotide sequence is codon optimized for expression in a eukaryotic cell.
[0282] Embodiment 31 is a non-naturally occurring CRISPR-Cas system, comprising: the recombinant Cas9 protein of any one of embodiments 1 to 28; and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0283] Embodiment 32 is a non-naturally occurring CRISPR-Cas system, comprising: the polynucleotide sequence of embodiment 29 or 30; and a nucleotide sequence encoding a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0284] Embodiment 33 is a non-naturally occurring CRISPR-Cas system, comprising: a regulatory element operably linked to the polynucleotide sequence of embodiment 29 or 30; and a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
[0285] Embodiment 34 includes the system of any one of embodiments 31 to 33, wherein the guide sequence is linked to a direct repeat sequence.
[0286] Embodiment 35 includes the system of any one of embodiments 31 to 34, wherein the guide polynucleotide comprises a tracrRNA sequence.
[0287] Embodiment 36 includes the system of any one of embodiments 31 to 34, further comprising a separate polynucleotide comprising a tracrRNA sequence.
[0288] Embodiment 37 includes the system of any one of embodiments 31 to 35, wherein the polynucleotide sequence encoding the recombinant Cas9 protein and the guide polynucleotide are on a single vector.
[0289] Embodiment 38 includes the system of embodiment 36, wherein the polynucleotide sequence encoding the recombinant Cas9 protein, the guide polynucleotide and the tracrRNA sequence are on a single vector.
[0290] Embodiment 39 is a delivery particle of the system comprising any one of embodiments 31 to 38.
[0291] Embodiment 40 is a vesicle comprising the system of any one of embodiments 31 to 38.
[0292] Embodiment 41 includes the vesicle of embodiment 40, wherein the vesicle is an exosome or a liposome.
[0293] Embodiment 42 is a viral vector comprising the system according to any one of embodiments 31 to 38.
[0294] Embodiment 43 includes the viral vector of embodiment 42, wherein the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus vector.
[0295] Embodiment 44 is a method of providing site-specific modification at a target sequence in the genome of a cell, the method comprising introducing the CRISPR-Cas system of any one of embodiments 31 to 38 into the cell.
[0296] Embodiment 45 includes the method of embodiment 44, wherein the modification comprises deletion of at least a portion of the target sequence.
[0297] Embodiment 46 includes the method of embodiment 44, wherein the modification comprises mutation of the target sequence.
[0298] Embodiment 47 includes the method of embodiment 44, wherein the modification comprises inserting a sequence of interest (SoI) at the target sequence.
[0299] Embodiment 48 includes the method of any one of embodiments 44 to 47, wherein the off-target modifications in the genome of the cell are less than about 5% of the modifications in the genome produced by recombinant Cas9.
[0300] Embodiment 49 includes the method of any one of embodiments 44 to 48, wherein the off-target modifications in the genome of the cell are less than about 2% of the modifications in the genome produced by recombinant Cas9.
[0301] Embodiment 50 includes the method of any one of embodiments 44 to 49, wherein the off-target modifications in the genome of the cell are less than about 1% of the modifications in the genome produced by the recombinant Cas9.
[0302] Embodiment 51 includes the method of any one of embodiments 44 to 50, wherein off-target modifications in the genome of the cell are reduced by at least about 50% relative to wild-type CRISPR-Cas9 or Cas9 that does not include a KFERQ motif or a KFERQ-like motif.
[0303] Embodiment 52 includes the method of any one of embodiments 44 to 51, wherein the cell is a bacterial cell, a mammalian cell, or a plant cell.
[0304] Embodiment 53 includes the method of embodiment 52, wherein the cell is a human cell.
[0305] Embodiment 54 includes the method of embodiment 53, wherein the cell is a pluripotent stem cell.
[0306] Embodiment 55 includes the method of embodiment 54, wherein the cell is an induced pluripotent stem cell.
[0307] Embodiment 56 includes the method of any one of embodiments 44 to 55, wherein the guide sequence of the guide polynucleotide is capable of hybridizing to a target sequence in the genome of the cell.
[0308] Embodiment 57 includes the method of any one of embodiments 44 to 56, wherein the CRISPR-Cas system is introduced into the cell via a delivery particle, vesicle, or viral vector.
[0309] sequence
[0310] Table 4 below lists the sequences provided herein.
[0311] Table 4. Sequence Listing
[0312]
[0313]
[0314] SpCas9 - Cas9 from Streptococcus pyogenes (SEQ ID NO: 1)
[0315]
[0316] FaDe-SpCas9—SpCas9 with F185N mutation (SEQ ID NO: 2)
[0317]
[0318] SpCas9 nucleotide sequence (SEQ ID NO: 3)
[0319]
[0320]
[0321] FaDe-SpCas9 nucleotide sequence (SEQ ID NO:4)
[0322]
[0323]
[0324] SpCas9 REC lobe: amino acids 61-718 of SpCas9 (SEQ ID NO: 5)
[0325]
[0326] SpCas9 NUC leaf 1: amino acids 1-60 of SpCas9 (SEQ ID NO: 6)
[0327]
[0328] SpCas9 NUC leaf 2: amino acids 719-1368 of SpCas9 (SEQ ID NO: 7)
[0329]
[0330] SpCas9 BH domain: amino acids 61-94 of SpCas9 (SEQ ID NO: 8)
[0331]
[0332] SpCas9 Rec1 domain 1: amino acids 95-180 of SpCas9 (SEQ ID NO: 9)
[0333]
[0334] SpCas9 Rec1 domain 2: amino acids 309-718 of SpCas9 (SEQ ID NO: 10)
[0335]
[0336] SpCas9 Rec2 domain: amino acids 181-308 of SpCas9 (SEQ ID NO: 11)
[0337]
[0338] SpCas9 RuvC I domain: amino acids 1-59 of SpCas9 (SEQ ID NO: 12)
[0339]
[0340] SpCas9 RuvC II domain: amino acids 718-774 of SpCas9 (SEQ ID NO: 13)
[0341]
[0342] SpCas9 RuvC III domain: amino acids 909-1098 of SpCas9 (SEQ ID NO: 14)
[0343]
[0344] SpCas9 HNH domain: amino acids 775-908 of SpCas9 (SEQ ID NO: 15)
[0345]
[0346] SpCas9 PI domain: amino acids 1099-1368 of SpCas9 (SEQ ID NO: 16)
[0347]
[0348] Cas9 from Streptococcus thermophilus (SEQ ID NO: 17)
[0349]
[0350] Cas9 from Streptococcus dysgalactiae (SEQ ID NO: 18)
[0351]
[0352] Cas9 from Streptococcus mutans (SEQ ID NO: 19)
[0353]
[0354] Cas9 from Listeria innocua (SEQ ID NO: 20)
[0355]
[0356] Cas9 from Staphylococcus aureus (SEQ ID NO: 21)
[0357]
[0358] Cas9 from Klebsiella pneumoniae (SEQ ID NO: 22)
[0359]
[0360] FnCas9 - Cas9 from Francisella novicida (SEQ ID NO: 23)
[0361]
[0362] Table 5. KFERQ or KFERQ-like motifs
[0363] SEQ ID NO:24 QUR SEQ ID NO:25 RKVEQ SEQ ID NO:26 QDQ SEQ ID NO:27 QRFFE SEQ ID NO:28 NRVVD SEQ ID NO:29 QUR SEQ ID NO:30 QKILD SEQ ID NO:31 QKKEL SEQ ID NO:32 QFREL SEQ ID NO:33 IKDJ SEQ ID NO:34 DVQ SEQ ID NO:35 QRIVE SEQ ID NO:36 VKELQ SEQ ID NO:37 QVFD SEQ ID NO:38 QELLR SEQ ID NO:39 VDK SEQ ID NO:40 RIKEN SEQ ID NO:41 NKKFE
[0364] Table 6. Nuclear localization signals
[0365] SEQ ID NO:42 PKKKRKV SEQ ID NO:43 AVKRPAATKKAGQAKKKKLD SEQ ID NO:44 PAAKRVKLD SEQ ID NO:45 MSRRRKANPTKLSENAKKLAKEVEN SEQ ID NO:46 KLKIKRPVK SEQ ID NO:47 KIPIK
[0366] Table 7. Primers
[0367] SEQ ID NO. Gene sequence SEQ ID NO:48 EMX1-T forward TTCCAGAACCGGAGGACAAAG SEQ ID NO:49 EMX1-T Reverse CCACCCTAGTCATTGGAGGT SEQ ID NO:50 EMX1-OT1 forward TTTATTATCTGCACATGTATG SEQ ID NO:51 EMX1-OT1 reverse CTACCTGTACATCTGCACAAG SEQ ID NO:52 EMX1-OT2 forward ATGTGCTTCAACCCATCACG SEQ ID NO:53 EMX1-OT2 reverse GTTGGCTTTCACAAGGATGC SEQ ID NO:54 FANCF-T forward CACGGATAAAGACGCTGGGA SEQ ID NO:55 FANCF-T Reverse TCCCAGGTGCTGACGTAGG SEQ ID NO:56 FANCF-OT1 forward TAGCACTGGGTGCTTAATCCG SEQ ID NO:57 FANCF-OT2 Reverse GGGTTTGGTTGGCTGCTCAT SEQ ID NO:58 AAVS1T2 forward ACCGGGGCCACTAGGGACAGGAT SEQ ID NO:59 AAVS1T2 Reverse AAACATCCTGTCCCTAGTGGCCC SEQ ID NO:60 Cas9 forward GATAAAGCAGACCTGCGGCTGATCTATC SEQ ID NO:61 Cas9 reverse CTGGCAGCTGAGCGATCAGGTTCTC SEQ ID NO:62 A1AT forward GATGCCCACCTTCCCCTCTC SEQ ID NO:63 A1AT Reverse AGTGGTGGCCTCATTCTGGA SEQ ID NO:64 ABCB1 forward GGCTTCACGAGAAAAGTTGATG SEQ ID NO:65 ABCB1 reverse GGATTCACAGGCTTCACCTAC
[0368] Table 8. Guide RNA sequences
[0369]
[0370] Examples
[0371] Materials and methods
[0372] The following materials and methods were used in the experiments described in the Examples. Oligonucleotides and guide RNA synthesis were provided by Sigma and Synthego. Unless otherwise stated, reagents and kits were purchased from ThermoFisher.
[0373] List of reagents and kits
[0374] Subcellular fractionation kit;
[0375] EDTA-free protease inhibitor cocktail (Sigma);
[0376] NOVEX NUPAGE protein gel, Bis-Tris 4%-12%, 1.5 mm thick, 10 wells;
[0377] 4X Laemmli buffer (BioRad);
[0378] NUPAGE MOPS20X SDS running buffer;
[0379] NUPAGE 20X transfer buffer;
[0380] NOVEX Sharp prestained protein standard molecular weight marker;
[0381] Nitrocellulose pre-cut blotting membrane, pore size 0.45 μm;
[0382] Primary antibody: Cas9 mouse monoclonal (Abcam);
[0383] monoclonal anti-FLAG M2 antibody (Sigma);
[0384] monoclonal anti-α-tubulin (Sigma);
[0385] Secondary antibody: IRDYE 680RD donkey anti-mouse IgG (H+L), 0.1 mg (Li-COR);
[0386] HSC70 mouse monoclonal (Santa Cruz);
[0387] Protein A agarose (Abcam);
[0388] REVERTAID RT Kit;
[0389] DNase I, no RNase;
[0390] Phusion Flash High-Fidelity PCR Master Mix;
[0391] Q5 Hot Start Hi-Fi 2X Master Mix;
[0392] Gentra Puregene kit (Qiagen);
[0393] Lipofectamine LTX Plus (Invitrogen);
[0394] FUGENE HD (Promega)
[0395] Cas9 (CRISPR-associated protein 9) ELISA kit (Cell Biolabs)
[0396] Cycloheximide CHX (Sigma)
[0397] Leupeptin (Sigma)
[0398] Anti-CRISPR-Cas9 antibody [EPR19799] (ab203933)
[0399] Anti-Ki67 antibody (ab15580)
[0400] Anti-cleaved caspase 3 antibody (ab2302)
[0401] Anti-γH2A.X (phospho S139) antibody [9F3] (ab26350)
[0402] Anti-CD8α antibody [144B] (ab17147)
[0403] Anti-CD4 Antibody [EPR19514] - Low Endotoxin, Azide Free (ab221775)
[0404] The primers used in the experiments described herein are listed in Table 2.
[0405] Table 2. Primers
[0406]
[0407] The guide RNAs used in the experiments described herein are listed in Table 3.
[0408] Table 3. Guide RNA sequences
[0409]
[0410] Experimental Procedure
[0411] Cell culture. SV-HUC-1 cells were cultured in F-12K medium (Ham's Kaighn modification) supplemented with 10% (v / v) fetal bovine serum (Gibco) and 1% antibiotics (100 U / mL penicillin and 100 mg / L streptomycin sulfate).
[0412] HEK293 cells were cultured in Dulbecco's modified Eagle's medium (Invitrogen) supplemented with 10% (v / v) fetal bovine serum (Gibco) and 1% antibiotics (100 U / mL penicillin and 100 mg / L streptomycin sulfate). All cells were grown at 37° C. in a humidified atmosphere with 5% CO 2 .
[0413] Neuro-2a cells were cultured in Eagle's Minimum Essential Medium (EMEM) supplemented with 10% (v / v) fetal bovine serum (Gibco) and 1% antibiotics (100 U / mL penicillin and 100 mg / L streptomycin sulfate).
[0414] HCT cells were cultured in RPMI-1640 medium (ATCC 30-2001) supplemented with 10% (v / v) fetal bovine serum (Gibco) and 1% antibiotics (100 U / mL penicillin and 100 mg / L streptomycin sulfate).
[0415] Human iPSCs were cultured in CELLARTIS DEF-CS Basal Medium supplemented with GF-1 Supplement (diluted 1:333), GF-2 Supplement (diluted 1:1000), and GF-3 Supplement (diluted 1:1000).
[0416] Plasmid mutagenesis. Mutagenesis and synthesis of G-Blocks gene fragments.
[0417] Example 1. Cas9 vs. FaDe-Cas9: protein turnover analysis
[0418] Plasmid transfection. Reverse transfection was performed with plasmids encoding Cas9 or FaDe-Cas9 (FaDe-Cas9: wild-type Cas9 with F185N mutation; Figure 5A Immortalized human or mouse cell lines were transfected with plasmids (shown as KFERQ-Cas9 in Figure 1). 3 μg of plasmid (Cas9 or FaDe-Cas9) was mixed with transfection agent (Lipofectamine LTX) in OPTIMEM medium and incubated in 6-well plates for 25 minutes. After incubation, cells were detached and counted (50×10 4), and resuspended in 2 mL of complete medium and added to the wells containing the transfection reagent mixture. 24 hours after transfection, cells were analyzed by GFP expression to assess transfection efficiency, and cells were harvested at different time points for Western blot analysis.
[0419] Cell lysis and protein extraction for Western blot. Harvest transfected cells with trypsin and centrifuge at 2000rpm for 5 minutes. After washing with cold PBS1X, the cell pellet is suspended in cold "2-step lysis buffer" and protease inhibitors (2-step lysis buffer: 10mM KCl, 20mM Tris HCl pH 7.4, 10mM MgCl2, 20mM EDTA, 10% glycerol, 0.8% TRITON), vortexed, sonicated and incubated on ice for 20 minutes. The ultrasonic program includes two treatments: peak power 20.0-duty cycle 40.0-pulse cycle 50-duration 15 seconds. Then the cell lysate is incubated with 420nM NaCl for 5 minutes to separate protein from nucleic acid, and clarified by centrifugation at 15,000rpm for 30 minutes at 4°C. After centrifugation, the precipitate containing DNA and membrane is discarded, and the protein concentration of the supernatant is measured by NANODROP.
[0420] Immunoblotting. The clarified lysates from cells transfected as described above were mixed with loading buffer (10% β-mercaptoethanol) and boiled for 8 minutes. The samples were loaded into NUPAGE Bis-Tris SDS protein gels and separated by running at 200V (MOPS1X) for 40 minutes. The protein gel was then transferred to a nitrocellulose membrane (tank transfer) at 35V using NUPAGE transfer buffer plus 20% methanol. The membrane was blocked in 1% BSA for 1 hour and incubated with the primary antibody at 4°C. After washing 3 times with PBS Tween 0.2%, the membrane was incubated with the secondary antibody (1:10000 in 1% BSA) at room temperature for 1 hour and washed 3 times with PBS Tween 0.2%. Detection was performed using the ODYSSEY imaging system.
[0421] The results are Fig. 6A , 6B and 7. Analysis of Cas9 and FaDe-Cas9 showed that the expression of FaDe-Cas9 was below the detection level of Western blot ( Fig. 6A ).like Figure 6B As shown, the mutation in Cas9 that generates FaDe-Cas9 does not impair Cas9 antibody specificity. Fig. 7AAs shown, FaDe-Cas9 showed very low expression in a short time window compared to Cas9. While Cas9 was detected by anti-Cas9 at increasing levels from 8 to 72 hours after transfection, FaDe-Cas9 was not detected by anti-Cas9 even at 8 hours after transfection.
[0422] Protein turnover was analyzed by GFP fusion protein expression. 3 HEK293 cells seeded into 96-well plates at a cell density were transfected with 100 ng of dual promoter-driven reporter vectors encoding Cas9-GFP-Fused or FaDe-Cas9-GFP-Fused, and mCherry expression under its own promoter. Transfection was performed using FUGENE HD transfection reagent. The plates were then placed in INCUCYTE to monitor the in vivo levels of Cas9 and FaDe-Cas9 proteins by measuring the GFP fluorescence signal over time. mCherry fluorescence was analyzed to assess transfection efficiency. Figure 8 Schematic diagram of the dual reporter vector transfected into cells for transfection efficiency assessment.
[0423] The results are shown in Figures 9 to 11. Although the transfection efficiency was comparable (as measured by mCherry fluorescence, Fig. 10A ), but FaDe-Cas9-GFP-Fused showed lower expression levels (as measured by GFP fluorescence, Fig. 10B ). Fig.12 The results of Cas9 and FaDe-Cas9 protein turnover experiments were shown to be similar in HEK, HCT, hIPSc, and Neuro-2a cells, indicating that the high turnover of FaDe-Cas9 is independent of the cell type.
[0424] mRNA stability was assessed by semi-quantitative RT-PCR. RNA was isolated from SVHUC1 cells transfected with plasmids encoding Cas9 or FaDe-Cas9. The cell pellets harvested at different time points were suspended in 1 ml TRIZOL. After 5 minutes at room temperature, 200 μL of chloroform was added and then incubated at room temperature for 3 minutes. The samples were centrifuged at 15,000 rpm for 15 minutes at 4 ° C, and the aqueous phase (containing RNA) was collected in a separate test tube. RNA was then precipitated by adding 500 μL of isopropanol to the aqueous phase, incubating at room temperature for 10 minutes and centrifuging at 15,000 rpm. RNA was collected at the bottom as a gel-like precipitate, washed with 70% ethanol and dissolved in RNase-free water. The RNA sample was treated with DNase in 10 μL of a reaction containing 1 μg RNA, 1X DNase buffer and 1 μL DNase plus ultrapure water. The mixture was incubated at 37 ° C for 30 minutes, and the RNA concentration was measured with NANODROP.
[0425] cDNA synthesis. The reaction was carried out in a 20 μL reaction with 500 ng RNA, 1 μL random hexamer primers, 4 μL 5X reaction buffer, 1 μL RIBOLOCK RNase inhibitor (20 U / μL), 2 μL 10 mM dNTP Mix, 1 μL REVERTAID RT (200 U / μL) and nuclease-free water. The reaction was incubated at 42°C for 60 minutes and terminated by heating at 70°C for 5 minutes.
[0426] PCR. PCR was performed in a 20 μL reaction volume (0.5 μL forward and reverse primers (10 μM), 1 μL cDNA and water) using 1 μL Q5 Hot Start High Fidelity 2X Master Mix. Each reaction was repeated three times. PCR was performed as follows: denaturation at 95°C for 1 minute, followed by 24 cycles of denaturation at 95°C for 30 seconds, annealing at 58°C for 30 seconds, extension at 72°C for 30 seconds, and a final extension of 2 minutes.
[0427] The results are shown in Figure 7B , indicating similar mRNA transcription levels from Cas9 and FaDe-Cas9. These data confirm that the transfection efficiency of Cas9 and FaDe-Cas9 is comparable.
[0428] Example 2. Evaluation of the effect of CMA on protein turnover and subcellular localization of FaDe-Cas9Cas9-HSC70 co-immunoprecipitation. SV-HUC-1 cells inoculated with 70% confluence were transfected with plasmids encoding Cas9 or FaDe-Cas9. Cells were harvested 48 hours after transfection and lysed with CO-IP lysis buffer (140mM KCl, 3mM MgCl2, 0.5% NONIDET P-40, 20mM HEPES pH 7.4, 1mM EDTA, 1.5mM EGTA, protease inhibitors (EDTA-free protease inhibitor cocktail)). The cell pellet was suspended in freshly prepared cold lysis buffer and passed through a 25-gauge needle 5-6 times using a 1mL syringe. The lysate was incubated on ice for 30 minutes and centrifuged at 15,000rpm for 20 minutes at 4°C. The clarified lysate was collected in a new test tube and analyzed with NANODROP to measure protein concentration. Cas9 immunoprecipitation was performed on 800 μg of clarified lysate, diluted in lysis buffer, and incubated with Cas9 primary antibody (1:100) overnight at 4°C. The next day, the immune complex was immobilized on 50 μL Protein GSEPHAROSE beads on a rotating drum for 4 hours at 4°C. The beads were then washed with CO-IP lysis buffer (3X), resuspended in 30 μL SDS sample buffer, boiled for 5 minutes, and subjected to western blotting with HSC70 primary antibody (1:1000).
[0429] The results are shown in Fig.13A Middle. Co-immunoprecipitation showed that FaDe-Cas9, but not Cas9, interacted with the CMA chaperone HSC70. As shown in the anti-HSC70 blot, FaDe-Cas9, but not Cas9, was detected using the anti-HSC70 antibody.
[0430] Colocalization of FaDe-Lamp2A. 3 HEK293 cells seeded in 96-well plates were co-transfected with 50 ng Cas9-GFP-Fused or FaDe-Cas9-GFP-Fused and 30 ng Lamp2A-dsRed-Fused. 24 hours after transfection, cells were analyzed by INCUCYTE zoom. Fig.14 Schematic diagram of the plasmids transfected into cells.
[0431] The results are shown in Fig. 13B and 15 Middle. GFP-tagged Cas9 or FaDe-Cas9 showed green fluorescence, while mCherry-tagged Lamp2A showed red fluorescence. Visualization of immunofluorescence signals showed that FaDe-Cas9 colocalized with lysosomal proteins and the CMA regulator Lamp2A in the cytosol. Colocalization with Lamp2A indicated active degradation.
[0432] Subcellular protein fractionation. Harvest cells with trypsin-EDTA and centrifuge at 500×g for 5 minutes. Wash the cell pellet with ice-cold PBS, and dry and remove the supernatant and discard. Add ice-cold cell extraction buffer (CEB) containing protease inhibitors to the cell pellet for cytoplasmic extraction and incubate at 4°C for 10 minutes. Centrifuge the lysate at 500×g for 5 minutes. Then transfer the supernatant (cytoplasmic extract) to a clean pre-cooled tube on ice. Add ice-cold membrane extraction buffer (MEB) containing protease inhibitors to the pellet to extract the membrane, incubate at 4°C for 10 minutes, and then centrifuge at 3000×g for 5 minutes. For nuclear extraction, add ice-cold nuclear extraction buffer (NEB) containing protease inhibitors plus 5μL 100mM CaCl2 and 3μL micrococcal nuclease (300 units) to the pellet per 100μL, incubate at 4°C for 10 minutes, and then incubate at room temperature for 5 minutes. The lysates were then clarified by centrifugation at 15,000 rpm for 10 min at 4° C. Protein extracts from each cell compartment were measured to quantify protein concentration and separated on NUPAGE Bis-Tris SDS protein gels.
[0433] 48 hours after transfection, the subcellular localization results of Cas9 and FaDe-Cas9 were shown in Fig.16 Cas9 and FaDe-Cas9 were found in the same subcellular location, suggesting that the F185N mutation in FaDe-Cas9 does not affect compartmentalization.
[0434] Example 3. Analysis of enzyme activity, on-target and off-target indels
[0435] Genomic DNA extraction. Genomic DNA extraction was performed using the Gentra Puregene kit. The cell pellet was suspended in 300 μL of lysis buffer and incubated on ice for 5 minutes after adding protein precipitation buffer (100 μL). The lysate was then centrifuged at 14,000 rpm for 10 minutes; the supernatant was mixed with 300 μL of isopropanol to precipitate the DNA, and then centrifuged at 14,000 rpm for 10 minutes. The DNA pellet was washed with 100 μL of 70% ethanol and suspended in 30 μL of water. DNA concentration was measured using NANODROP.
[0436] Surveyor nuclease assay. PCR: PCR was performed in a total volume of 20 μL with 100 ng gDNA and 10 μL 2X master mix (PHUSION Flash High Fidelity PCR) plus 1 μL forward and reverse primers (10 μM). PCR was performed as follows: denaturation at 95°C for 3 minutes, followed by 35 cycles of denaturation at 95°C for 5 seconds, annealing at 58°C for 30 seconds, extension at 72°C (1 min / 1,000 bps), and a final extension of 5 minutes.
[0437] Digestion: PCR products were denatured by heating at 99°C for 5 minutes and then reannealed to form heteroduplexes by cooling to 65°C for 30 minutes and to 23°C for 30 minutes.
[0438] The hybridized heteroduplexes or homoduplexes were treated with Surveyor nuclease (also known as CEL nuclease), which recognizes mismatches present in heteroduplex DNA and cleaves both strands at the 3' side of the mismatch distortion. Therefore, 20 μL of unpurified PCR product (about 250 ng) in a 50 μL reaction volume was digested with 1 μL surveyor enzyme + 1 μL enhancer at 42°C for 20 minutes. 10 μL of treated DNA was then separated by electrophoresis on a 10% acrylamide gel for 40 minutes.
[0439] The results of the assay to determine the nuclease efficiency of Cas9 and FaDe-Cas9 in HEK cells are shown in Fig.17A Middle. Cas9 and FaDe-Cas9 nuclease activities are comparable as shown by the gel band showing “cleaved” DNA. Fig. 17B Next generation sequencing analysis shown confirms comparable nuclease efficiency of Cas9 (12.8%) and FaDe-Cas9 (16.5%). Fig.18 It was shown that in hiPSCs, FaDe-Cas9 had a level of nuclease activity comparable to that of Cas9.
[0440] EMX1 / FANCF1 off-target analysis. HEK293 cells were cultured at 60×10 4The density of 100 μg / mL was inoculated into 12-well plates and co-transfected with 800 ng of plasmid encoding Cas9 or FaDe-Cas9 and 200 ng of plasmid encoding gRNA (EMX1 or FANCF1). FUGENE HD transfection reagent was used for transfection. Cells were harvested 72 hours after transfection and lysed for genomic DNA extraction. 100 ng gDNA, 1 μl PHUSION FLASH high-fidelity PCR master mix, and forward and reverse primers fused to the adapters designed according to the manufacturer's recommendations from Illumina were used for PCR amplification of the on-target and off-target regions. PCR was performed as follows: denaturation at 95 ° C for 1 minute, followed by denaturation at 95 ° C for 30 seconds, annealing at 58 ° C for 30 seconds, and extension at 72 ° C for 30 seconds, 30 cycles, and finally extension for 2 minutes. Purify the PCR product and analyze it by next-generation sequencing.
[0441] The results are shown in Fig.19 Although FaDe-Cas9 has slightly lower on-target modification than Cas9, the off-target modification of FaDe-Cas9 is significantly reduced compared to Cas9. Therefore, FaDe-Cas9 shows an increased normalized on-target modification efficiency than Cas9.
[0442] Electroporation of ribonucleoprotein (RNP). Pre-complexation of Cas9 / RNP: In a 1.5 mL tube, 100 pmol of Cas9 or FaDe-Cas9 was mixed with 120 pmol of synthetic dual gRNA in 10 μL 1X Cas9 buffer and incubated at room temperature for 20 minutes. 4 HEK293 cells of 100 μL density were suspended in 20 μL electroporation buffer SF and incubated with RNP complexes for 2 minutes. Cells were electroporated using a 4D nuclear transfection instrument (Lonza). After nuclear transfection, cells were seeded in 12-well plates with 1 mL of complete medium (DMEM) and harvested at different time points (3 hours; 6 hours; 10 hours; 24 hours).
[0443] RNP transfection. Human iPSCs were cultured at 20×10 4 The cells were seeded at a density of 100 μM and transfected with increasing concentrations of Cas9 or FaDe-Cas9 and 3 μM dual gRNA. Transfection was performed using Lipofectamine CRISPRMAX transfection reagent. 48 hours after infection, cells were lysed to extract genomic DNA, and deletions were confirmed by PCR.
[0444] Example 4. Effects of Cas9 expression on cells
[0445] Example 4.1
[0446] Human urothelial cells (SVHUC-1) were transfected with wild-type Cas9. Cell clones containing Cas9 were confirmed by ddPCR ( Figure 1A , upper panel). Four weeks after the generation of the Cas9 stable cell line, cells proliferated and were counted at 24, 48, and 72 hours. Cell counts from wild-type cells (no Cas9), clone 1 (heterozygous Cas9 integration), and clone 3 (homozygous Cas9 integration) showed a decreasing trend ( Figure 1A , small picture below).
[0447] Example 4.2
[0448] Mice were transfected with wild-type Cas9 on a doxycycline-inducible promoter. The body weight of mice was measured after induction ("iCas dox") and compared with wild-type mice ("WT water"), wild-type mice with doxycycline ("WT dox"), and mice transfected but not expressing Cas9 ("iCas water"). Figure 1B The results in the study showed that mice transfected with and expressing wild-type Cas9 exhibited weight loss compared to mice not expressing Cas9.
[0449] Example 4.3
[0450] Human induced pluripotent stem cells (hiPSCs) were transiently transfected with wild-type Cas9. Microscope images of cells were taken 5 weeks after transient expression of Cas9. Figure 2 As shown, hiPSCs lose their undifferentiated phenotype when Cas9 is transiently expressed.
[0451] Example 5. Protein turnover analysis of Cas9 relative to FaDe-Cas9
[0452] Example 5.1 Cas9 and FaDe-Cas9 intracellular protein levels measured by ELISA assay
[0453] 20x10 4HEK293 cells at a density of 10 μl were suspended in 20 μl electroporation buffer SF and incubated with the RNP complex for 2 minutes. The cells were electroporated using a 4D nuclear transfection instrument (4D nuclear transfection instrument core unit: Lonza, AAF-1002B; 4D nuclear transfection instrument X unit: AAF-1002X; Lonza). After nuclear transfection, the cells were seeded in a 12-well plate with 1 ml of complete medium (DMEM) and harvested at 24 hours. The cells were lysed and the protein level was analyzed by commercial kit ELISA assay (Cell Biology Laboratory) according to the instruction manual: the cell or tissue lysate was sonicated or homogenized in a lysis buffer such as RIPA buffer (25 mM Tris·HCl pH7.6, 150 mM NaCl, 1% NP-40, 1% sodium deoxycholate, 0.1% SDS) and centrifuged at 10,000 x g for 10 minutes at 4°C before the assay.
[0454] The results are as follows Fig.24A As shown. When RNP was electroporated at 7.5ug / 10^5 cells, about 5% of intracellular Cas9 was recovered at 24h compared to <0.1% FaDe-Cas9. At 24h after electroporation, the abundance of FaDe-Cas9 in cells was reduced by >97%.
[0455] Example 5.2 Measurement of protein turnover and degradation
[0456] Use GFP-fused Cas9 or FaDe-Cas9 at 30x10 4 HEK293 cells were transfected at a density of 1:1 and 1:1 and GFP expression was analyzed over time by incucyte.
[0457] At 12 hours after transfection, cells were treated with CHX (10 ug / ml) to inhibit protein synthesis, and the degradation protein of Cas9 and FaDe-Cas9 was measured following the GFP signal compared to untreated cells.
[0458] FaDe-Cas9 was less abundant and never reached the intracellular protein levels observed with Cas9 via GFP expression ( Fig. 24B Cells exposed to the protein translation inhibitor cycloheximide (CHX) showed a decrease in FaDe-Cas9 levels over time, whereas Cas9 protein levels remained constant over time ( Fig.24C ).
[0459] Example 6. Effect of CMA on high protein turnover and protein subcellular localization of FaDe-Cas9
[0460] Example 6.1. Lamp2A knockdown
[0461] Using RNAiMAX transfection reagent and OPTIMEM, 7x10 cells were co-transfected with GFP fusion expression vector Cas9 or FaDe-Cas9 plus Ds-Redlamp2a vector (2:1), gRNA (3:1) and increasing doses of siRNA (10-20-40-60-90-100ng). 3 HEK293 cells were co-transfected with scrambled siRNA as a control. Transfected cells were monitored over time under incucyte zoom to analyze KD lamp2a efficiency and protein accumulation by GFP signal.
[0462] To inhibit CMA, the expression of the lysosomal receptor Lampa2a was reduced by siRNA. siRNA transfection resulted in a dose-dependent reduction in the expression of Lamp2a receptor ( Fig.25A / B), but the protein level of Cas9 was unchanged ( Fig.25A ), FaDe-Cas9 showed dose-dependent accumulation, indicating that its degradation was dependent on CMA ( Fig.25B ).
[0463] Example 6.2 Cas9-HSC70-co-immunoprecipitation.
[0464] Using MaxCyte, HEK293s cells (50 million) were electroporated with Cas9 fused with 20ug of the marker tag of FaDe-Cas9. 24 hours after electroporation, cells were treated with 100uM leupeptin (cathepsin B inhibitor) to temporarily inhibit lysosomal degradation. Cells were harvested 48 hours after electroporation and lysed with CO-IP lysis buffer (KCl 140mM, 3mMMgCl2, 0.5% Nonidet P-40, 20mM Hepes pH 7.4, 1mM EDTA, 1.5mM EGTA, protease inhibitors (EDTA-free protease inhibitor cocktail, Sigma)). The cell pellet was suspended in cold and freshly prepared lysis buffer and passed through a 25-gauge needle 5-6 times using a 1ml syringe. The lysate was incubated on ice for 30 minutes and centrifuged at 15,000 for 20 minutes at 4°C. The clarified lysate was collected in a new tube and analyzed with Nanodrop to measure protein concentration. Cas9 immunoprecipitation was performed on 800ug clarified lysate, which was diluted in lysis buffer and incubated overnight at 4°C with a labeled primary antibody (anti-FLAG F7425 SIGMA 1:500). The next day, the immune complex was fixed on protein G-agarose beads (50ul) on a drum at 4°C for 4 hours. The beads were washed with CO-IP lysis buffer (3X), resuspended in SDS sample buffer (30ul), boiled for 5 minutes, and western blotted with Cas9 and HSC70 primary antibodies (1:1000). 30ug lysate was used as INPUT to confirm the total amount of protein.
[0465] Immunoprecipitation revealed that FaDe-Cas9 bound with high affinity to the master regulator of chaperone-mediated autophagy, HSC70. Fig.26 middle.
[0466] Example 7. Cas9 versus FaDe-Cas9 in vivo mouse model
[0467] Example 7.1 Adenovirus constructs
[0468] Adenoviruses expressing Cas9 or FaDe-Cas9 and gRNA (Ad-Cas9-gMH and Ad-Cas9-gP and Ad-FaDe-Cas9-gMH and Ad-FaDe-Cas9-gP) were generated by Vector Biolabs (Malvern). Adv Cas9 / FaDe and gRNA were expressed from chicken β-actin hybrid (CBh) and U6 promoters, respectively, in a replication-deficient adenovirus serotype 5 (dE1 / E3) backbone. Negative control adenoviruses expressing Cas9 and GFP from CBh and CMV promoters, respectively, but without gRNA, were also generated (Ad-Cas9-GFP and Ad-FaDe-Cas9-GFP).
[0469] Example 7.2 In vivo on-target / off-target analysis
[0470] For in vivo off-target editing analysis, (gP), a promiscuous guide RNA targeting the mouse PCSK9 locus, was chosen due to its high potential to induce multiple off-target mutations in the mouse genome ( Fig.27A ).
[0471] Male mice aged 9 to 11 weeks were injected intravenously with 1 × 10 9 Infectious units (IFU) of adenovirus (Ad-Cas9-gp or Ad-FaDe-Cas9-gp, or Ad-Cas9-GFP or Ad-FaDe-Cas9-GFP) in 200 μl of phosphate-buffered saline diluent. Peripheral blood was collected before virus administration (baseline), one week after virus administration, and at termination (four days or three weeks after virus administration).
[0472] Genomic DNA was extracted from liver tissue of adenovirus-injected mice on day 7 after treatment for indel analysis. Off-targets were identified by targeted deep sequencing with CIRCLE-seq by selecting sites with read counts above 50% of the on-target and various lower-ranked sites (containing up to 6 mismatches relative to the on-target). Evaluation of the gene editing frequency of on-target + 9 off-target sites of gP in the mouse model showed that 4 of the 9 different off-target sites showed significantly reduced FaDe-Cas9 gene editing compared to Cas9 ( Fig.27B ).
[0473] Example 7.3 Next Generation Sequencing (NGS)
[0474] PCR products were purified using magnetic beads, quantified using the QuantiFlor dsDNA System kit (Promega), normalized to 10 ng / μl per amplicon, and pooled. The pooled samples were end-repaired and A-tailed using the End Prep enzyme mix and the reaction buffer in the NEBNextUltra II DNA library preparation kit from Enomina, and ligated to the Illumina TruSeq adapter using the ligation master mix and ligation enhancer in the same kit. The library samples were then purified using magnetic beads, size-selected using PEG / NaCl SPRI solution (KAPA Biosystems), quantified using droplet digital PCR (BioRad), and loaded onto the Illumina MiSeq for deep sequencing.
[0475] Example 7.4 In vivo genome editing and viability assessment
[0476] For in vivo Pcsk9 gene editing, 9- to 11-week-old humanized PCSK9 mice (PCSK9KIKO) carrying a single allele of mouse or human PSCK9 were injected with a tail vein injection of 1 × 109 infectious units (IFU) of adenovirus (Ad-Cas9-gMH or Ad-FaDe-Cas9-gMH and Ad-Cas9-GFP or Ad-FaDe-Cas9-GFP) in 200 μl of phosphate-buffered saline diluent. Genomic DNA was extracted from liver tissues of adenovirus-injected mice on day 7 after treatment for insertion-deletion analysis of the human and mouse PCSK9 loci, respectively, by NGS. Liver lobes were included in paraffin blocks and stained with hematoxylin and eosin (H&E), mitotic markers Ki6, Cas9, cleaved caspase 3, p-H2AX, and CD4 / CD8.
[0477] The livers of mice exposed to FaDe-Cas9 showed higher gene editing due to its rapid degradation in liver tissue, which in turn led to increased survival of edited cells, such as Fig.28A Compared with Cas9, the rapid in vivo turnover of FaDe-Cas9 resulted in lower hepatotoxicity, as evidenced by unchanged hepatic glycogen, barely detectable mitotic marker Ki67, less infiltrates, and minimal cell necrosis. Fig.28B / C).
[0478] Furthermore, we found that viral-mediated expression of FaDe-Cas9 in vivo resulted in lower cytotoxic T lymphocyte immune responses ( Fig.28D). One week after AdV delivery, Cas9 was still highly expressed in hepatocytes, while FaDe-Cas9 was undetectable (A, E). The shorter incubation period of FaDe-Cas9 resulted in lower levels of apoptosis (cleaved caspase 3 IHC; B, F) and DNA double-strand breaks (phospho-H2AX IHC; C, G) in hepatocytes. Although the immune response to Cas9 delivery in murine livers consisted of moderate numbers of CD4+ (memory) lymphocytes and many CD8+ (cytotoxic) lymphocytes, the latter were significantly reduced in FaDe-Cas9 AdV-infected livers (CD4-CD8 IHC; D, H).
Claims
1. A recombinant Cas9 protein comprising an engineered KFERQ motif, wherein the engineered KFERQ motif is in a surface-exposed region of the recombinant Cas9 protein, and wherein the engineered KFERQ motif is VDKLN (SEQ ID NO: 39), and the recombinant Cas9 protein is obtained by introducing a F185N mutation at amino acid position 185 of SEQ ID NO:
1.
2. A recombinant Cas9 protein, which, compared with the recombinant Cas9 protein as described in claim 1, further comprises a mutation at position D10, H840 or a combination thereof in SEQ ID NO: 1, wherein the mutation is selected from D10A or D10N; H840A, H840N or H840Y; and a combination thereof.
3. The recombinant Cas9 protein of claim 1 or 2, wherein the recombinant Cas9 protein produces sticky ends.
4. A recombinant Cas9 protein, which, compared with the recombinant Cas9 protein as described in claim 1 or 2, further comprises one or more nuclear localization signals.
5. A polynucleotide molecule encoding the recombinant Cas9 protein according to any one of claims 1 to 4.
6. The polynucleotide molecule of claim 5, wherein the polynucleotide molecule is codon optimized for expression in eukaryotic cells.
7. A non-naturally occurring CRISPR-Cas system comprising: (a) the recombinant Cas9 protein of any one of claims 1 to 4; and (b) a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
8. A non-naturally occurring CRISPR-Cas system comprising: (a) a polynucleotide molecule as claimed in claim 5 or 6; and (b) a nucleotide sequence encoding a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
9. A non-naturally occurring CRISPR-Cas system comprising: (a) a regulatory element operably linked to a polynucleotide molecule as claimed in claim 5 or 6; and (b) a guide polynucleotide that forms a complex with the recombinant Cas9 protein and comprises a guide sequence.
10. The system of any one of claims 7 to 9, wherein the guide sequence is linked to a direct repeat sequence.
11. The system of any one of claims 7 to 9, wherein the guide polynucleotide comprises a tracrRNA sequence.
12. A delivery particle comprising the system of any one of claims 7 to 11.
13. A vesicle comprising the system of any one of claims 7 to 11. The vesicle of claim 13 , wherein the vesicle is an exosome or a liposome.
15. A viral vector comprising the system of any one of claims 7 to 11.
16. An in vitro method for providing site-specific modification at a target sequence in the genome of a cell, the method comprising introducing into the cell a CRISPR-Cas system as claimed in any one of claims 7 to 11.
Citation Information
Patent Citations
Multiple domain proteins
US20110059502A1
Aminoalcohol lipidoids and uses thereof
US20110293703A1
Conjugated lipomers and uses thereof
US20120251560A1
Poly(beta-amino alcohols), their preparation, and uses thereof
US20130302401A1
Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription
US20140068797A1