Methods for modulating VEGF and uses thereof
Patent Information
- Application Number
- JP2024542251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-14
- Filing Date
- 2023-01-10
- Publication Date
- 2026-01-20
AI Technical Summary
The prior art is difficult to effectively regulate and silen endogenous genes, especially VEGF genes, in vivo, and there are problems such as safety, toxicity, immunogenicity and non-targeting effects.
CRISPR/Cas9 gene editing technology is used to design sgRNA to target the transcription start site of VEGF gene, bind DNA binding proteins such as dCas9-KRAB to form fusion molecules, and regulate VEGF gene expression through DNA methyltransferases, etc.
Efficient targeted regulation and silencing of VEGF genes are achieved, limiting viral vector delivery is avoided, and safety and specificity are improved.
Smart Images

Figure 00000077_0000 
Figure 00000077_0001 
Figure 00000077_0002
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to the fields of molecular biology, immunology, and medicine. More specifically, the present disclosure relates to CRISPR / Cas9-based fusion molecules and methods of use thereof for use in the targeted reduction or elimination of VEGF gene products in vivo.
[0002] This application includes a Sequence Listing which has been submitted in electronic format herewith and is hereby incorporated by reference herein. [Background technology]
[0003] Engineered DNA-binding proteins that can be customized to target any gene in mammalian cells have enabled rapid advances in biomedical research and are a promising platform for gene therapy. The RNA-guided CRISPR-Cas9 system has emerged as a promising platform for programmable targeted gene control. Fusion of catalytically inactive "dead" Cas9 (dCas9) to a Kruppel-associated box (KRAB) domain generates synthetic repressors capable of highly specific and potent regulation or silencing of target genes in cell culture experiments.
[0004] However, the sustained regulation and silencing of endogenous genes using synthetic dCas9-KRAB fusion proteins has been difficult for use in in vivo therapy. Synthetic repressors exceed the size packaging limits of viral vector delivery methods. Safety, toxicity, immunogenicity, and off-target effects are other difficulties that limit the use of synthetic repressors in vivo. Summary of the Invention [Problem to be solved by the invention]
[0005] There is a need in the art for alternative approaches to generate engineered synthetic gene repressors and for in vivo delivery of synthetic gene repressors for use as therapeutics. The present disclosure addresses this unmet need in the art. [Means for solving the problem]
[0006] In one aspect, the disclosure provides an sgRNA comprising a sequence complementary to a target DNA sequence located within 500 bp upstream to 500 bp downstream of the transcription start site of the vascular endothelial growth factor (VEGF) gene.
[0007] In some embodiments, the sgRNA comprises the nucleic acid sequence of any one of SEQ ID NOs: 29-58 and 60-84.
[0008] In some embodiments, the VEGF gene is a VEGF-A gene from a mammal, such as human, monkey, mouse, rat, and rabbit.
[0009] In one aspect, the disclosure provides DNA sequences encoding the sgRNAs disclosed herein.
[0010] In one aspect, the present disclosure provides a method for producing a method for manufacturing a semiconductor device comprising: (a) a fusion molecule, or a nucleic acid sequence encoding the fusion molecule, comprising at least one DNA binding protein and at least one modulator of gene expression; and (b) a guide molecule comprising an sgRNA as disclosed herein and a protein-binding sequence capable of binding to at least one DNA-binding protein, or a nucleic acid sequence encoding the guide molecule. providing a composition comprising At least one modulator of gene expression provides for the modification of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element.
[0011] In some embodiments, at least one modulator of gene expression described herein provides an alteration of at least one nucleotide within 1,000 bp upstream to 1,000 bp downstream of the transcription start site of the VEGF gene.
[0012] In one aspect, the disclosure provides a fusion molecule comprising at least one DNA binding protein and at least one modulator of gene expression, or a composition comprising a nucleic acid sequence encoding the fusion molecule, wherein the fusion molecule targets a genomic region near the VEGF gene and / or within a VEGF regulatory element, and wherein the at least one modulator of gene expression provides an alteration of at least one nucleotide near the VEGF gene and / or within the VEGF regulatory element.
[0013] In some embodiments, the at least one modulator of gene expression comprises a DNA methyltransferase (DNMT), a DNA demethylase, a histone methyltransferase, a histone demethylase or a portion thereof, or a zinc finger protein based transcription factor or a portion thereof, or a combination thereof.
[0014] In some embodiments, the at least one DNA binding protein is Cas9, dCas9, Cpf1, a zinc finger nuclease (ZNF), a transcription activator-like effector nuclease (TALEN), a homing endonuclease, a dCas9-FokI nuclease, or a MegaTal nuclease.
[0015] In some embodiments, the VEGF gene is a VEGF-A, VEGF-B, VEGF-C, VEGF-D, or VEGF-E gene.
[0016] In some embodiments, the VEGF regulatory element is a transcription initiation site, a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, or a locus control region.
[0017] In some embodiments, the modification of at least one nucleotide near the VEGF gene and / or within the VEGF regulatory element is located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream or downstream of the transcription start site of the VEGF gene.
[0018] In some embodiments, the modification of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element is located within 500 bp upstream or downstream of the transcription start site of the VEGF gene.
[0019] In some embodiments, the modification of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element is located within 300 bp upstream or downstream of the transcription start site of the VEGF gene.
[0020] In some embodiments, the modification of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element is located within 1000 bp upstream of the transcription start site to within 300 bp downstream of the transcription start site of the VEGF gene.
[0021] In some embodiments, the modification of at least one nucleotide is DNA methylation.
[0022] In some embodiments, the at least one modulator of gene expression comprises one or more selected from DNA methyltransferases (DNMTs), zinc finger protein-based transcription factors, portions thereof, and any combination thereof. The DNA methyltransferases may be DNMT3A, DNMT3B, DNMT3L, DNMT1, and / or DNMT2.
[0023] In some embodiments, DNMT3A comprises the amino acid sequence of SEQ ID NO:23 and / or DNMT3L comprises the amino acid sequence of SEQ ID NO:24.
[0024] In some embodiments, the zinc finger protein transcription factor is a Kruppel-associated repression box (KRAB). Specifically, KRAB may comprise the amino acid sequence of SEQ ID NO:22.
[0025] In some embodiments, the at least one modulator of gene expression comprises a DNA methyltransferase or a portion thereof and a zinc finger protein transcription factor or a portion thereof. Specifically, the DNA methyltransferase may be selected from DNMT3A and DNMT3L and combinations thereof, and the zinc finger protein transcription factor may be KRAB.
[0026] In some embodiments, at least one DNA binding protein is Cas9, dCas9, Cpf1, zinc finger nuclease (ZNF), transcription activator-like effector nuclease (TALEN), homing endonuclease, dCas9-FokI nuclease or MegaTal nuclease.For example, at least one DNA binding protein is dCas9.
[0027] In some embodiments, the dCas9 is selected from the group consisting of Staphylococcus aureus dCas9, Streptococcus pyogenes dCas9, Campylobacter jejuni dCas9, Corynebacterium diphtheria dCas9, Eubacterium ventriosum dCas9, Streptococcus pasteurianus dCas9, Lactobacillus farciminis dCas9, Sphaerochaete globus dCas9, and / or other dCas9 inhibitors. globus dCas9, Azospirillum (e.g. strain B510) dCas9, Gluconacetobacter diazotrophicus dCas9, Neisseria cinerea dCas9, Roseburia intestinalis dCas9, Parvibaculum lavamentivorans dCas9, Nitratifractor salsuginis (e.g. strain DSM 16511) dCas9, Campylobacter lari (e.g. strain CF89-12) dCas9, Streptococcus thermophilus dCas9, thermophilus (e.g., strain LMD-9) dCas9.
[0028] In some embodiments, the dCas9 comprises the amino acid sequence of SEQ ID NO:1.
[0029] In some embodiments, the fusion molecules disclosed herein comprise at least one modulator of gene expression fused to the C-terminus, the N-terminus, or both, of at least one DNA binding protein.
[0030] In some embodiments, at least one modulator of gene expression is fused directly to at least one DNA binding protein.
[0031] In some embodiments, at least one modulator of gene expression is indirectly fused to at least one DNA binding protein via a non-modulator, a second modulator, or a linker.
[0032] In some embodiments, the fusion molecules disclosed herein comprise dCas9 fused to KRAB at its C-terminus and DNMT3A and DNMT3L at its N-terminus.
[0033] In some embodiments, the fusion molecule comprises the amino acid sequence of SEQ ID NO:28.
[0034] In some embodiments, the fusion molecule further comprises at least one nuclear localization sequence, which may be directly or indirectly fused to the C-terminus, N-terminus, or both, of the at least one DNA binding protein.
[0035] In some embodiments, the nucleic acid sequence encoding the fusion molecule is deoxyribonucleic acid (DNA) or messenger ribonucleic acid (mRNA).
[0036] In some embodiments, the compositions disclosed herein further comprise at least one single guide RNA (sgRNA) that is complementary to a target DNA sequence near the VEGF gene and / or within a VEGF regulatory element.
[0037] In some embodiments, the target DNA sequence is located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream or downstream of the transcription start site of the VEGF gene.
[0038] In some embodiments, the sgRNA comprises the nucleic acid sequence of SEQ ID NOs: 29-58 and 60-84.
[0039] In some embodiments, the fusogenic molecule is packaged in a liposome or lipid nanoparticle.
[0040] In some embodiments, the fusion molecule and the sgRNA are packaged in a liposome or lipid nanoparticle. The fusion molecule and the sgRNA can be packaged in the same liposome or lipid nanoparticle, or in different liposomes or lipid nanoparticles.
[0041] In some embodiments, the liposomes or lipid nanoparticles are comprised of an ionizable lipid (20%-70% molar ratio), a PEGylated lipid (0%-30% molar ratio), a supporting lipid (30%-50% molar ratio), and cholesterol (10%-50% molar ratio).
[0042] In some embodiments, the ionizable lipid is selected from the group consisting of a pH-responsive ionizable lipid, a thermally responsive ionizable lipid, and a light-responsive ionizable lipid.
[0043] In some embodiments, the fusion molecule is packaged in an AAV vector.
[0044] In some embodiments, the fusion molecule and the sgRNA are packaged in an AAV vector. The fusion molecule and the sgRNA may be packaged in the same AAV vector or in different AAV vectors.
[0045] In some embodiments, the compositions disclosed herein are pharmaceutical compositions that include a pharma- ceutically acceptable carrier.
[0046] In one aspect, the disclosure provides a method for modulating (e.g., reducing or eliminating) expression of a VEGF gene product in a cell, the method comprising introducing into a cell a fusion molecule, or a nucleic acid sequence encoding the fusion molecule, comprising at least one DNA binding protein and at least one modulator of gene expression, wherein the modulator of gene expression provides an alteration of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element, thereby modulating (e.g., reducing or eliminating) expression of a VEGF gene product in the cell.
[0047] In one aspect, the disclosure provides an in vivo method of modulating (e.g., reducing or eliminating) expression of a VEGF gene product in a subject, the method comprising introducing into a cell of the subject a fusion molecule, or a nucleic acid sequence encoding the fusion molecule, comprising at least one DNA binding protein and at least one modulator of gene expression, wherein the modulator of gene expression provides an alteration of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element, thereby modulating (e.g., reducing or eliminating) expression of a VEGF gene product in the subject.
[0048] In one aspect, the disclosure provides a method for treating or alleviating a symptom of a VEGF-related disorder in a subject, the method comprising introducing into a cell of the subject a fusion molecule, or a nucleic acid sequence encoding the fusion molecule, comprising at least one DNA binding protein and at least one modulator of gene expression, wherein the modulator of gene expression provides an alteration of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element, thereby treating or alleviating a symptom of a VEGF-related disorder in the subject.
[0049] In some embodiments, the VEGF regulatory element is a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, or a locus control region.
[0050] In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within about 100bp, about 200bp, about 300bp, about 400bp, about 500bp, about 600bp, about 700bp, about 800bp, about 900bp, about 1000bp, about 1100bp, about 1200bp, about 1300bp, about 1400bp, or about 1500bp upstream of the transcription start site of the VEGF gene. In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within 1000bp upstream of the transcription start site of the VEGF gene. In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within 300bp upstream of the transcription start site of the VEGF gene.
[0051] In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within about 100bp, about 200bp, about 300bp, about 400bp, about 500bp, about 600bp, about 700bp, about 800bp, about 900bp, about 1000bp, about 1100bp, about 1200bp, about 1300bp, about 1400bp, or about 1500bp downstream of the transcription start site of the VEGF gene. In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within about 300bp downstream of the transcription start site of the VEGF gene. In some embodiments, the modification of at least one nucleotide in the vicinity of the VEGF gene and / or in the VEGF regulatory element is located within 1000bp upstream of the transcription start site of the VEGF gene and within 300bp downstream of the transcription start site.
[0052] In some embodiments, the modification of at least one nucleotide is DNA methylation.
[0053] In some embodiments, the at least one modulator of gene expression comprises a DNA methyltransferase (DNMT), a DNA demethylase, a histone methyltransferase, a histone demethylase, or a portion thereof.
[0054] In some embodiments, the at least one modulator of gene expression comprises a DNA methyltransferase (DNMT) or a portion thereof. In some embodiments, the DNA methyltransferase is DNMT3A, DNMT3B, DNMT3L, DNMT1, or DNMT2. In some embodiments, the DNMT3A comprises the amino acid sequence of SEQ ID NO: 23. In some embodiments, the DNMT3L comprises the amino acid sequence of SEQ ID NO: 24.
[0055] In some embodiments, at least one modulator of gene expression comprises a zinc finger protein transcription factor or a part thereof. In some embodiments, the zinc finger protein transcription factor is Kruppel-associated repression box (KRAB). In some embodiments, KRAB comprises the amino acid sequence of SEQ ID NO: 22.
[0056] In some embodiments, the at least one modulator of gene expression comprises a DNA methyltransferase or a portion thereof and a zinc finger protein transcription factor or a portion thereof. In some embodiments, the DNA methyltransferase is selected from DNMT3A and DNMT3L and combinations thereof, and the zinc finger protein transcription factor is KRAB.
[0057] In some embodiments, at least one DNA binding protein is Cas9, dCas9, Cpf1, zinc finger nuclease (ZNF), transcription activator-like effector nuclease (TALEN), homing endonuclease, dCas9-FokI nuclease, or MegaTal nuclease.In some embodiments, at least one DNA binding protein is dCas9. In some embodiments, the dCas9 is selected from the group consisting of Staphylococcus aureus dCas9, Streptococcus pyogenes dCas9, Campylobacter jejuni dCas9, Corynebacterium diphtheriae dCas9, Eubacterium ventriosum dCas9, Streptococcus pasteurianus dCas9, Lactobacillus farciminis dCas9, Spirochaete globus dCas9, Azospirillum (e.g., strain B510) dCas9, Gluconacetobacter diazotrophicus Cas9, Neisseria cinerea dCas9, Roseburia intestinalis dCas9, Parvibaculum labamentivorans dCas9, Nitratifructator sarsuginis (e.g., DSM In some embodiments, the dCas9 comprises the amino acid sequence of SEQ ID NO:1.
[0058] In some embodiments, fusion molecules comprise at least one modulator of gene expression fused to the C-terminus, the N-terminus, or both, of at least one DNA binding protein.
[0059] In some embodiments, at least one modulator of gene expression is directly fused to at least one DNA binding protein.In some embodiments, at least one modulator of gene expression is indirectly fused to at least one DNA binding protein via a non-modulator, a second modulator, or a linker.In some embodiments, the fusion molecule comprises dCas9 fused to KRAB at C-terminus and DNMT3A and DNMT3L at N-terminus.In some embodiments, the fusion molecule comprises the amino acid sequence of SEQ ID NO:28.
[0060] In some embodiments, the fusion molecule further comprises at least one nuclear localization sequence.In some embodiments, at least one nuclear localization sequence is directly fused to the C-terminus, N-terminus, or both of at least one DNA binding protein.In some embodiments, at least one nuclear localization sequence is indirectly fused to the C-terminus, N-terminus, or both of at least one DNA binding protein via a linker.
[0061] In some embodiments, the nucleic acid sequence encoding the fusion molecule is deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid sequence encoding the fusion molecule is messenger ribonucleic acid (mRNA).
[0062] In some embodiments, the method further comprises introducing at least one single guide RNA (sgRNA) that is complementary to a DNA sequence near the VEGF gene and / or within a VEGF regulatory element, thereby targeting the fusion molecule to the VEGF gene or VEGF regulatory element, or introducing DNA encoding the sgRNA. In some embodiments, the sgRNA comprises the nucleic acid sequence of SEQ ID NOs: 29-58 and 60-84.
[0063] In some embodiments, the fusion molecule is formulated in a liposome or lipid nanoparticle. In some embodiments, the fusion molecule and sgRNA are formulated in a liposome or lipid nanoparticle. In some embodiments, the fusion molecule and sgRNA are formulated in the same liposome or lipid nanoparticle. In some embodiments, the fusion molecule and sgRNA are formulated in different liposomes or lipid nanoparticles.
[0064] In some embodiments, the liposome or lipid nanoparticle is comprised of an ionizable lipid (20%-70% molar ratio), a PEGylated lipid (0%-30% molar ratio), a supporting lipid (30%-50% molar ratio), and cholesterol (10%-50% molar ratio). In some embodiments, the ionizable lipid is selected from the group consisting of a pH-responsive ionizable lipid, a thermally responsive ionizable lipid, and a light-responsive ionizable lipid.
[0065] In some embodiments, the fusion molecule is formulated in an AAV vector.In some embodiments, the fusion molecule and sgRNA are formulated in an AAV vector.In some embodiments, the fusion molecule and sgRNA are formulated in the same AAV vector.In some embodiments, the fusion molecule and sgRNA are formulated in different AAV vectors.
[0066] In some embodiments, the fusion molecule is delivered to the cell by local injection, systemic injection, or a combination thereof, hi some embodiments, the fusion molecule is delivered to the subject's eye by intraocular or intravitreal injection.
[0067] In some embodiments, the subject is a mammal, such as a human, monkey, mouse, rat, rabbit, pig, horse, cat, and dog.
[0068] In some embodiments, the VEGF-related disorder is associated with angiogenesis.In some embodiments, the VEGF-related disorder is neovascular disorder, for example, ophthalmic neovascular disorder, including age-related macular degeneration (AMD).In some embodiments, the VEGF-related disorder is wet AMD or dry AMD.
[0069] The present disclosure provides an sgRNA comprising any one of the nucleic acid sequences of SEQ ID NOs: 29 to 58 and 60 to 84. The present disclosure provides a DNA sequence encoding any one of the sgRNAs disclosed herein.
[0070] In one aspect, provided herein is a composition as described above for use in treating or alleviating a symptom of a VEGF-related disorder in a subject.
[0071] In some embodiments, the VEGF-related disorder is a neovascular disorder, such as an ophthalmic neovascular disorder, including age-related macular degeneration (AMD).
[0072] In one aspect, provided herein is the use of the above-described composition in the manufacture of a medicament for treating or alleviating a symptom of a VEGF-related disorder in a subject.
[0073] In one aspect, provided herein is a kit comprising a container containing the above-described composition or fusion molecule. [Brief description of the drawings]
[0074] [Figure 1]FIG. 1A is a schematic diagram showing the "EPICAS" dual plasmid system and sgRNA tiling screen design targeting mouse VEGF-A expression. The first plasmid (the "catalytic protein" plasmid or "fusion molecule" plasmid) encodes DNMT3A-DNMT3L (3A3L)-dCas9-KRAB under the control of a CAG promoter, and a GFP marker separated by a 2A element. The second plasmid (the "sgRNA" plasmid) has an sgRNA scaffold under the control of a U6 promoter, and an mCherry marker under the control of a CMV promoter. The tiling screen sgRNA targets the transcription start site (TSS) +250 bp upstream of the mouse VEGF-A protein coding sequence (CDS). FIG. 1B is a bar graph showing the relative mRNA expression 48 hours after transfection of the mouse N2A cell line with catalytic protein plasmids and various single VEGF sgRNA plasmids. Figure 1C is a bar graph showing relative mRNA expression one week after transfection of the mouse N2A cell line with catalytic protein plasmid and various single VEGF sgRNA plasmids or mixtures of VEGF sgRNA plasmids. Figure 1D is a schematic showing the results of bisulfite PCR analysis of the VEGF-A locus. Each row represents one single clone and each column represents one specific genomic location. Black dots represent sites of successful methylation. [Diagram 2]Figure 2A is a schematic diagram showing EPICAS mRNA plasmid design. EPICAS ORF contains DNMT3A-DNMT3L-dCas9-KRAB cassette. The plasmid can be digested with XbaI and BpiI restriction sites to form a linearized plasmid. Figure 2B is an electrophoretic diagram of mRNA expressed and purified from EPICAS mRNA plasmid. Figure 2C is a schematic diagram showing that EPICAS mRNA can successfully knock down endogenous VEGFA gene in primary mouse hepatocytes. Left panel: photomicrograph of primary mouse hepatocytes; middle panel: flow cytometry graph showing expression of GFP 72 hours after transfection with and without GFP-P2A-Casoff mRNA and sgRNA treatment; right panel: relative VEGFA mRNA expression in GFP positive cells between control and GFP-P2A-Casoff mRNA and sgRNA treated groups. [Diagram 3] FIG. 3A is a schematic diagram of lipid nanoparticle (LNP) design. Extragenic CRISPR / Cas elements and sgRNA elements can be encapsulated by LNPs. FIG. 3B is a transmission electron microscopy image showing LNPs containing EPICAS. FIG. 3C is a graph showing the particle size distribution of LNPs. FIG. 3D is a series of photographs showing in vivo fluorescence imaging of luciferase mRNA delivered by intravitreal injection of lipid nanoparticles into the eyes of Ai9 mice. FIG. 3E is a schematic diagram of the in vivo experimental design for delivery of LNPs containing EPICAS mRNA and sgRNA targeting mouse VEGF. LNPs were administered to Ai9 mice by injection into the posterior region of the eye, and VEGFA gene expression in the retina and choroid was analyzed 5 days after injection. [Figure 4]Figure 4A is a schematic diagram of the rabbit VEGFA sgRNA tiling screen design in 293T reporter cell line. The tiling screen sgRNAs target 500 bp upstream and downstream of the transcription start site (TSS) of the rabbit VEGFA gene. Figure 4B is a series of graphs showing the fluorescence intensity of reporter cells 72 hours after transfection with EPICAS plasmid and sgRNAs targeting rabbit VEGFA. Figure 4C is a series of graphs showing VEGFA mRNA expression in rabbit RK-13 cells after transfection of six sgRNAs targeting the endogenous gene VEGFA in rabbit cells, which had good knockdown effect together with EPICAS plasmid. [Diagram 5] Figure 5A is a schematic diagram showing the experimental design of a sgRNA tiling screen targeting a conserved region of human and mouse VEGF-A, specifically within 300 bp upstream to 300 bp downstream of the transcription start site (TSS) of the human VEGFA gene. Figure 5B is a series of graphs showing VEGFA mRNA expression 48 hours after transfection with the EPICAS dual plasmid system using various sgRNAs targeting VEGF. Figure 5C is a series of graphs showing VEGFA mRNA expression 96 hours after transfection with the EPICAS dual plasmid system using various sgRNAs targeting VEGF. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0075] The present disclosure overcomes the problems associated with current technology by providing engineered fusion molecules (e.g., DNMT3A-DNMT3L (3A3L)-dCas9-KRAB fusion molecules) for targeted reduction or elimination of gene products (e.g., VEGF) in cells for use in in vivo gene therapy. The engineered fusion molecules of the present disclosure are useful for treating genetic diseases, including, for example, liver diseases, diseases associated with high cholesterol, and diseases associated with cholesterol dysregulation (e.g., low-density lipoprotein (LDL) cholesterol). Thus, methods for making engineered fusion molecules and pharmaceutical formulations thereof (e.g., lipid nanoparticle formulations) for use in in vivo delivery are also provided.
[0076] I. Definition As used herein, the term "coding sequence" or "encoding nucleic acid" refers to a nucleic acid (RNA or DNA molecule) that comprises a nucleotide sequence that codes for a protein. The coding sequence may further comprise initiation and termination signals operably linked to control elements, including a promoter and polyadenylation signal, that are capable of directing expression in cells of an individual or mammal to which the nucleic acid is administered. The coding sequence may be codon-optimized.
[0077] The term "complement" or "complementary" as used herein in reference to nucleic acids can refer to Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to the property shared between two nucleic acid sequences, such as when they are aligned antiparallel to each other, such that the nucleotide bases at each position are complementary.
[0078] The terms "correct", "genome edit" and "restore" refer to alterations that alter a mutant gene that encodes a mutant or truncated protein or no protein at all, resulting in expression of a full-length or partially full-length functional protein. Correcting or restoring a mutant gene can include replacing a region of the gene that has a mutation, or replacing the entire mutant gene with a copy of the gene that does not have the mutation by a repair mechanism such as homology-directed repair (HDR). Correcting or restoring a mutant gene can also include repairing a frameshift mutation that causes a premature stop codon, an aberrant splice acceptor site, or an aberrant splice donor site by generating a double-strand break in the gene and then repairing it using non-homologous end joining (NHEJ). NHEJ can add or delete at least one base pair during repair to restore the correct reading frame and eliminate the premature stop codon. The correction or restoration of mutant genes also includes disrupting the abnormal splice acceptor site or splice donor sequence. The correction or restoration of mutant genes also includes deleting non-essential gene segments by the simultaneous action of two nucleases on the same DNA strand to remove the DNA between the two nuclease target sites and restore the correct reading frame by repairing the DNA break by NHEJ.
[0079] As used herein, the terms "donor DNA," "donor template," and "repair template" refer to a double-stranded DNA fragment or molecule that contains at least a portion of a gene of interest. The donor DNA can encode a fully functional protein or a partially functional protein.
[0080] As used herein, the terms "frameshift" or "frameshift mutation" are used interchangeably and refer to a type of genetic mutation in which the addition or deletion of one or more nucleotides causes a shift in the reading frame of a codon in an mRNA. A shift in the reading frame can result in a change in the amino acid sequence, such as a missense mutation or a premature stop codon during protein translation.
[0081] As used herein, the terms "functional" and "fully functional" describe a protein that has biological activity. A "functional gene" refers to a gene that is transcribed into mRNA that is translated into a functional protein.
[0082] As used herein, the term "fusion protein" refers to a chimeric protein created directly or indirectly by the covalent or non-covalent association of two or more genes that originally encoded separate proteins. In some embodiments, translation of the fusion gene results in a single polypeptide having functional properties derived from each of the original proteins.
[0083] As used herein, the term "genetic construct" refers to a DNA or RNA molecule that contains a nucleotide sequence that encodes a protein. The coding sequence includes start and stop signals operably linked to control elements, including a promoter and polyadenylation signal, that are capable of directing expression in a cell.
[0084] The term "homologous recombination repair" or "HDR" used interchangeably herein refers to the machinery in cells that repairs double-stranded DNA damage when homologous pieces of DNA are present in the nucleus, mainly in the G2 and S phases of the cell cycle. HDR uses donor DNA templates to guide repair and create specific sequence changes to genomes, including targeted addition of entire genes. If donor templates are provided with site-specific nucleases, such as CRISPR / Cas9-based systems, the cellular machinery will repair the break by homologous recombination, which is enhanced by several orders of magnitude in the presence of DNA breaks. In the absence of homologous DNA pieces, non-homologous end joining can occur instead.
[0085] As used herein, the term "genome editing" refers to modifying a gene. Genome editing can include correcting or restoring a mutant gene. Genome editing can include knocking out a gene, such as a mutant gene or a normal gene. Genome editing is used to treat a disease by modifying a gene of interest.
[0086] The term "identical" or "identity" as used herein in relation to two or more nucleic acid or polypeptide sequences means that the sequences have a specified percentage of residues that are the same over a specified region. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over a specified region, determining the number of positions where identical residues occur in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the specified region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences are of different length, or if the alignment results in one or more sticky ends such that the identified comparison region contains only a single sequence, the residues of the single sequence are included in the denominator of the calculation but not in the numerator. When comparing DNA and RNA, thymine (T) and uracil (U) may be considered equivalent. Identity may be calculated manually or using a computer sequence algorithm such as BLAST or BLAST 2.0. The identity of related peptides can be readily calculated by known methods.Such methods include, but are not limited to, those described in Computational Molecular Biology, edited by Lesk, AM, Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, edited by Smith, DW, Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part 1, edited by Griffin, AM and Griffin, HG, Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Sequence Analysis Primer, edited by Gribskov, M. and Devereux, J., M. Stockton Press, New York, 1991; and Carillo et al., SIAM J. Applied Math. 48, 1073 (1988), the entireties of which are incorporated herein by reference.
[0087] As used herein, the terms "mutated gene" or "mutated gene", used interchangeably herein, refer to a gene that has undergone a detectable mutation. A mutant gene has undergone an alteration, such as a loss, gain, or exchange of genetic material that affects the normal transmission and expression of the gene. As used herein, a "disrupted gene" refers to a mutant gene that has a mutation that causes a premature stop codon. A disrupted gene product is truncated relative to the non-disrupted full-length gene product.
[0088] As used herein, the term "modulator of extragenic modification" refers to an agent that targets gene expression through extragenic modification (e.g., through histone acetylation or methylation, or DNA methylation at the target gene's control elements, such as promoters, enhancers, or transcription start sites). Chromatin remodeling and DNA methylation are two major mechanisms that control gene transcription. Specific extragenic marks (e.g., DNA methylation) structurally or biochemically direct gene transcription or gene silencing / suppression. For example, DNA methylation in regions that control transcriptional activity alters gene expression without changing the underlying DNA sequence. Transcriptional control using extragenic modification (e.g., DNA methylation) allows targeted regulation of gene expression without affecting the expression of other gene products.
[0089] As used herein, the term "non-homologous end joining (NHEJ) pathway" refers to a pathway that repairs double-stranded breaks in DNA by directly ligating the broken ends without the need for a homologous template. Template-independent religation of DNA ends by NHEJ is a stochastic and error-prone repair process that introduces random microinsertions and microdeletions (indels) at the breaks in DNA. This method is used to intentionally disrupt, delete, or change the reading frame of a targeted gene sequence. NHEJ typically uses short homologous DNA sequences, called microhomologies, to guide the repair. These microhomologies are often present in single-stranded overhangs at the ends of the double-stranded break. When the overhangs are perfectly matched, NHEJ usually repairs the break accurately, but incomplete repairs resulting in loss of nucleotides can also occur. This is, however, more common when the overhangs are not matched.
[0090] As used herein, the term "normal gene" refers to a gene that has not undergone alteration, such as the loss, gain, or exchange of genetic material. A normal gene undergoes normal gene transmission and gene expression.
[0091] As used herein, the term "nuclease-mediated NHEJ" refers to NHEJ that is initiated after a nuclease, such as cas9, breaks double-stranded DNA.
[0092] As used herein, the term "nucleic acid" or "oligonucleotide" or "polynucleotide" refers to at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary strand. That is, the nucleic acid also encompasses the complementary strand of the depicted single strand. Many variants of a nucleic acid may be used for the same purpose as a given nucleic acid. That is, the nucleic acid also encompasses substantially identical nucleic acids and their complements. A single strand provides a probe that can hybridize to a target sequence under stringent hybridization conditions. That is, the nucleic acid also encompasses a probe that hybridizes under stringent hybridization conditions. A nucleic acid may be single-stranded or double-stranded, and may include portions of both double-stranded and single-stranded sequences. The nucleic acid may be DNA, RNA, or hybrid, both genomic and cDNA, and may contain combinations of deoxyribonucleotides and ribonucleotides, as well as combinations that include uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. The nucleic acid may be obtained by chemical synthesis or recombinant methods.
[0093] As used herein, the term "operably linked" refers to the juxtaposition of two or more biological sequences of interest, with or without spacers or linkers, in a relationship that allows them to function in an intended manner. When used in reference to a polypeptide, it is intended to mean that the polypeptide sequences are linked in a manner that allows the linked product to have the intended biological function. When used in reference to a polynucleotide, as an example, when a polynucleotide encoding a polypeptide is operably linked to a control sequence (e.g., a promoter, enhancer, silencer sequence, etc.), it is intended to mean that the polynucleotide sequence is linked in a manner that allows for controlled expression of the polypeptide from the polynucleotide. In some embodiments, the expression of a gene is under the control of a promoter to which it is spatially connected. The promoter may be 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene may be approximately the same as the distance between the promoter and the gene it controls in the gene from which the promoter is derived. As known in the art, variations in this distance may be accommodated without loss of promoter function.
[0094] As used herein, the term "partially functional" describes a protein encoded by a mutant gene that has less biological activity than a functional protein, but more than a non-functional protein. In one embodiment, a partially functional protein exhibits less than 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, or 30% of the biological activity of the corresponding functional protein.
[0095] The terms "premature stop codon" or "out-of-frame stop codon," as used interchangeably herein, refer to a nonsense mutation in the sequence of DNA that results in a stop codon at a position not normally found in the wild-type gene. A premature stop codon can cause a truncated or shortened protein compared to the full-length version of the protein.
[0096] The term "promoter" as used herein refers to a synthetic or naturally derived molecule that can confer, activate, or enhance expression of a nucleic acid in a cell. A promoter may contain one or more specific transcriptional control sequences that further enhance expression and / or alter the spatial and / or temporal expression of the nucleic acid. A promoter may also contain distal enhancer or repressor elements, which may be located as far as several thousand base pairs away from the start site of transcription. Promoters may be derived from sources including viruses, bacteria, fungi, plants, insects, and animals. A promoter may control expression of genetic components constitutively, or differentially with respect to the cell, tissue, or organ in which expression occurs, or with respect to the developmental stage in which expression occurs, or in response to an external stimulus such as a physiological stress, a pathogen, a metal ion, or an inducer. Representative examples of promoters include the bacteriophage T7 promoter, the bacteriophage T3 promoter, the SP6 promoter, the lac operator promoter, the tac promoter, the SV40 late promoter, the SV40 early promoter, the RSV-LTR promoter, and the CMV IE promoter.
[0097] As used herein, the term "target gene" refers to any nucleotide sequence that encodes a known or hypothesized gene product. A target gene may be a mutated gene involved in a genetic disease or disorder.
[0098] As used herein, the term "target region" refers to the region of a target gene to which a site-specific nuclease is designed to bind.
[0099] As used herein, the term "transgene" refers to a gene or genetic material that contains a genetic sequence that is isolated from one organism and introduced into another organism. Alternatively, the term "transgene" refers to a gene or genetic material that is chemically synthesized and introduced into an organism. This non-natural segment of DNA may retain the ability to produce RNA or protein in the transgenic organism, or may change the normal function of the genetic code of the transgenic organism. The introduction of a transgene has the potential to change the phenotype of the organism.
[0100] As used herein, the term "variant" when used in reference to a nucleic acid means (i) a portion or fragment of a referenced nucleotide sequence, (ii) a complement of a referenced nucleotide sequence or a portion thereof, (iii) a nucleic acid that is substantially identical to a referenced nucleic acid or its complement, or (iv) a nucleic acid that hybridizes under stringent conditions to a referenced nucleic acid, its complement, or a sequence substantially identical thereto. "Variants" in reference to peptides or polypeptides that differ in amino acid sequence by amino acid insertions, deletions, or conservative substitutions but retain at least one biological activity.
[0101] A variant may also refer to a protein having an amino acid sequence substantially identical to a referenced protein, with the amino acid sequence retaining at least one biological activity. Conservative substitutions of amino acids, i.e., replacing an amino acid with a different amino acid having similar properties (e.g., hydrophilicity, degree and distribution of charged regions), are recognized in the art to typically include minor changes. These minor changes may be identified in part by considering the hydropathic index of an amino acid, as understood in the art. Kyte et al., J. Mol. Biol. 157:105-132 (1982), incorporated herein by reference in its entirety. The hydropathic index of an amino acid is based on a consideration of its hydrophobicity and charge. It is known in the art that amino acids with similar hydropathic indexes can be substituted and still retain protein function. In one embodiment, amino acids with hydropathic indexes of ±2 are substituted. The hydrophilicity of an amino acid is also used to identify substitutions that result in a protein that retains biological function. Considering the hydrophilicity of an amino acid in the context of a peptide allows for the calculation of the maximum local average hydrophilicity of that peptide. Substitutions may be made for amino acids having hydrophilicity values within ±2 of each other. Both the hydrophobicity index and hydrophilicity value of an amino acid are influenced by the particular side chain of that amino acid. Consistent with that observation, it is understood that amino acid substitutions that are compatible with biological function depend on the relative similarity of amino acids, particularly their side chains, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.
[0102] As used herein, the term "vector" refers to a nucleic acid sequence that includes an origin of replication.The vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome.The vector can be a DNA or RNA vector.The vector can be a self-replicating extrachromosomal vector, such as a DNA plasmid.
[0103] As used herein, the terms "gene transfer," "gene delivery," and "gene transduction" refer to a method or system for reliably inserting a specific nucleotide sequence (e.g., DNA or RNA), a fusion protein, a polypeptide, etc., into a target cell.
[0104] As used herein, the terms "adenovirus-associated virus (AAV) vector," "AAV gene therapy vector," and "gene therapy vector" refer to vectors with functional or partially functional ITR sequences and transgenes. As used herein, the term "ITR" refers to inverted terminal repeats (ITRs). ITR sequences can be derived from adeno-associated virus serotypes, including, without limitation, AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, and AAV-6. However, ITRs do not have to be wild-type nucleotide sequences and can be altered (e.g., by nucleotide insertion, deletion, or substitution) so long as the sequences retain the ability to provide functional rescue, replication, and packaging. AAV vectors can be deleted in whole or in part of one or more of the AAV wild-type genes, preferably the rep and / or cap genes, but retain adjacent functional ITR sequences. Functional ITR sequences function, for example, to rescue, replicate, and package AAV virions or particles. Thus, an "AAV vector" is defined herein to include at least the sequences necessary to insert a transgene into a cell of a subject, and may include sequences required in cis for viral replication and packaging (e.g., functional ITRs).
[0105] As used herein, the term "gene therapy" refers to a method of treating a patient in which a polypeptide or nucleic acid sequence is transferred into the patient's cells, thereby modulating the activity and / or expression of a particular gene.In certain embodiments, the expression of a gene is suppressed.In certain embodiments, the expression of a gene is enhanced.In certain embodiments, the temporal or spatial pattern of the expression of a gene is modulated.
[0106] A "transgene" may include a transgenic sequence or a native or wild-type DNA sequence. A transgene may become part of the genome of a primate subject. A transgenic sequence may be partially or wholly heterologous. That is, the transgenic sequence or a portion thereof may originate from a different species than the cell into which it is introduced.
[0107] As used herein, the term "stably maintained" refers to the characteristic of a transgenic subject (e.g., a human or non-human primate) that has maintained at least one of its transgenic elements (i.e., the desired element) through many generations of cells. For example, the term is intended to encompass many cell division cycles of the originally transfected cell. The term "stable transfection" or "stably transfected" refers to the introduction and integration of foreign DNA into the genome of a cell. The term "stable transfectant" refers to a cell that has foreign DNA stably integrated into its genomic DNA.
[0108] As used herein, the terms "transgene coding", "nucleic acid molecule coding", "DNA sequence coding" and "DNA coding" refer to the order or sequence of deoxyribonucleotides along a strand of deoxyribonucleic acid. The order of these deoxyribonucleotides can determine the order of amino acids along a chain of, for example, a polypeptide (protein). That is, a DNA sequence can code for an amino acid sequence.
[0109] As used herein, the term "wild type" (wt) refers to a gene or gene product that has the characteristics of that gene or gene product when isolated from a naturally occurring source. A wild type gene is the gene that is most frequently observed in a population, and thus the "normal" or "wild type" form of the gene is arbitrarily designated. In contrast, the term "modified" or "mutant" refers to a gene or gene product that exhibits modified sequence and / or functional properties (i.e., altered characteristics) when compared to the wild type gene or gene product. It is noted that naturally occurring mutants can be isolated that are identified by the acquisition of altered characteristics when compared to the wild type gene or gene product.
[0110] As used herein, the term "transfection" refers to the uptake of foreign nucleic acid (e.g., DNA or RNA) by a cell. A cell is "transfected" when an exogenous nucleic acid (DNA or RNA) is introduced inside the cell membrane. Several transfection techniques are generally known in the art (see, e.g., Graham et al., Virol., 52:456 (1973); Sambrook et al., Molecular Cloning, a Laboratory Manual, Cold Spring Harbor Laboratories, New York (1989); Davis et al., Basic Methods in Molecular Biology, Elsevier, (1986); and Chu et al., Gene 13:197 (1981), all of which are incorporated herein by reference in their entirety). Such techniques may be used to introduce one or more exogenous DNA moieties, such as gene transfer vectors and other nucleic acid molecules, into suitable recipient cells.
[0111] As used herein, the terms "stable transfection" and "stably transfected" refer to the introduction and integration of foreign DNA into the genome of a transfected cell. The term "stable transfectant" refers to a cell that has foreign DNA stably integrated into its genomic DNA.
[0112] As used herein, the term "transient transfection" or "transiently transfected" refers to the introduction of foreign DNA into a cell, where the foreign DNA is not integrated into the genome of the transfected cell but is maintained as an episome. During this time, the foreign DNA is subject to the regulatory controls that govern the expression of endogenous genes in chromosomes. The term "transient transfectant" refers to a cell that has taken up foreign DNA but has not been able to integrate this DNA. As used herein, the term "transduction" refers to the delivery of DNA molecules to recipient cells in vivo or in vitro via a replication-defective viral vector, e.g., via a recombinant AAV virion.
[0113] As used herein, the term "recipient cell" refers to a cell that has been or can be transfected or transduced by a nucleic acid construct or vector carrying a selected nucleotide sequence of interest. The term includes the progeny of a parent cell, whether or not identical in morphology or genetic make-up to the original parent, so long as the selected nucleotide sequence is present. Recipient cells may be cells of a subject to which gene therapy particles and / or gene therapy vectors have been administered.
[0114] As used herein, the term "recombinant DNA molecule" refers to a DNA molecule that is comprised of segments of DNA joined together by means of molecular biological techniques.
[0115] As used herein, the term "control element" refers to a genetic element that can control the expression of a nucleic acid sequence. For example, a promoter is a control element that facilitates the initiation of transcription of an operably linked coding region. Other control elements are splicing signals, polyadenylation signals, termination signals, etc.
[0116] The term DNA "control sequence" collectively refers to control elements such as promoter sequences, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites ("IRES"), enhancers, etc., which collectively provide for the replication, transcription, and translation of a coding sequence in a recipient cell. Not all of these control sequences need be present.
[0117] Transcriptional control signals in eukaryotes generally include "promoter" and "enhancer" elements. Promoters and enhancers consist of short arrays of DNA sequences that specifically interact with cellular proteins involved in transcription (Maniatis et al., Science 236:1237 (1987), which is incorporated herein by reference in its entirety). Promoter and enhancer elements have been isolated from a variety of eukaryotic sources, including genes in yeast, insect, and mammalian cells, as well as viruses (analogous control sequences, i.e., promoters, are also found in prokaryotes). The selection of a particular promoter and enhancer depends on the type of recipient cell. Some eukaryotic promoters and enhancers have a broad host range, while others are functional in a limited subset of cell types (see, e.g., Voss et al., Trends Biochem. Sci., 11:287 (1986), which is incorporated herein by reference in its entirety; and Maniatis et al., supra, for review). For example, the SV40 early gene enhancer is highly active in a wide variety of cell types from many mammalian species and has been used to express proteins in a wide range of mammalian cells [Dijkema et al., EMBO J. 4:761 (1985), incorporated herein by reference in its entirety]. Promoter and enhancer elements derived from the human elongation factor 1-alpha gene (Uetsuki et al., J. Biol. Chem., 264:5791 (1989); Kim et al., Gene 91:217 (1990); and Mizushima and Nagata, Nucl. Acids. Res., 18:5322 (1990)], the long terminal repeat of Rous sarcoma virus (Gorman et al., Proc. Natl. Acad. Sci. USA 79:6777 (1982)], and human cytomegalovirus (Boshart et al., Cell 41:521 (1985)] are also useful for expression of proteins in a variety of mammalian cell types and are incorporated herein by reference in their entirety. Promoters and enhancers can be found naturally, either alone or together.For example, retroviral long terminal repeats contain both promoter and enhancer elements. Generally, promoters and enhancers are independent of the gene being transcribed or translated. That is, the enhancers and promoters used may be "endogenous," "exogenous," or "heterologous" with respect to the gene to which they are operably linked. An "endogenous" enhancer / promoter is one that is naturally linked to a given gene in the genome. An "exogenous" or "heterologous" enhancer or promoter is one that has been juxtaposed to a gene by genetic engineering (i.e., molecular biology techniques) such that transcription of that gene is directed by the linked enhancer / promoter.
[0118] As used herein, the term "tissue-specific" refers to control elements or sequences, such as promoters, enhancers, etc., where expression of a nucleic acid sequence is substantially greater in a particular cell type or tissue.
[0119] The presence of "splicing signals" on expression vectors often results in high levels of recombinant transcript expression. Splicing signals mediate the removal of introns from the primary RNA transcript and consist of splice donor and acceptor sites [Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, New York (1989), pp. 16.7-16.8, the entirety of which is incorporated herein by reference]. A commonly used splice donor and acceptor site is the splice junction from the 16S RNA of SV40.
[0120] Transcription termination signals are generally found downstream of polyadenylation signals and are several hundred nucleotides in length. As used herein, the term "poly A site" or "poly A sequence" refers to a DNA sequence that directs both the termination and polyadenylation of nascent RNA transcripts. Efficient polyadenylation of recombinant transcripts is desirable, since transcripts lacking a poly A tail are unstable and rapidly degraded. The poly A signal utilized in an expression vector may be "heterologous" or "endogenous." An endogenous poly A signal is one that is naturally found at the 3' end of the coding region of a given gene in the genome. A heterologous poly A signal is one that is isolated from one gene and operably linked to the 3' end of another gene. A commonly used heterologous poly A signal is the SV40 poly A signal. The SV40 poly A signal is contained on a 237 bp BamHI / BclI restriction fragment and directs both termination and polyadenylation (Sambrook et al., supra, pp. 16.6-16.7, incorporated herein by reference in its entirety).
[0121] As used herein, the terms "subject" and "patient" are used interchangeably herein and refer to both human and non-human animals. The term "non-human animals" of the present disclosure includes all vertebrates, e.g., mammals and non-mammals, e.g., non-human primates, sheep, dogs, cats, horses, cows, chickens, amphibians, reptiles, etc.
[0122] As defined herein, a "therapeutically effective amount" or "therapeutically effective dose" is an amount or dose of a fusion protein, polypeptide, nucleic acid, lipid nanoparticle, liposome, AAV particle, or virion that can produce a sufficient amount of a desired protein to modulate the activity of the protein in a desired manner, thereby providing a palliative tool for clinical intervention. In some embodiments, a therapeutically effective amount or dose of a transfected fusion protein, polypeptide, nucleic acid, AAV particle, or virion described herein is sufficient to confer suppression of the gene targeted by the fusion protein / gene therapy construct.
[0123] As used herein, the term "treating," e.g., a disorder, means that a subject (e.g., a human) having, at risk of having, and / or experiencing symptoms of the disorder, in one embodiment, when administered a fusion molecule or a nucleic acid encoding a fusion molecule, and / or a gRNA or a nucleic acid encoding a gRNA, as described herein, has less severe symptoms and / or recovers more quickly than if the fusion molecule or a nucleic acid encoding a fusion molecule, and / or a gRNA or a nucleic acid encoding a gRNA, had never been administered.
[0124] II. DNA-binding proteins In certain embodiments of the method and composition defined herein by the present disclosure, DNA binding protein (e.g., DNA targeting agent) comprises (DNA) nuclease, for example, nuclease that can target DNA in sequence-specific manner or can be directed or instructed to target DNA in sequence-specific manner, such as CRISPR-Cas system, zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), or meganuclease.In some embodiments, DNA binding protein is the DNA nuclease derived from CRISPR-Cas system.
[0125] Transcription Activator-Like Effector Nuclease (TALEN) System In certain embodiments, the nucleic acid binding protein is a (modified) transcription activator-like effector nuclease (TALEN) system. Transcription activator-like effectors (TALEs) can be engineered to bind to virtually any desired DNA sequence. Exemplary methods of genome editing using the TALEN system can be found, for example, in Cermak T. Doyle EL. Christian M. Wang L. Zhang Y. Schmidt C et al., Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic Acids Res. 2011; 39:e82; Zhang F. Cong L. Lodato S. Kosuri S. Church GM. Arlotta P Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nat Biotechnol. 2011; 29:149-153, and U.S. Patent Nos. 8,450,471, 8,440,431, and 8,440,432, each of which is incorporated by reference in its entirety.
[0126] By way of further guidance and without limitation, naturally occurring TALEs or "wild-type TALEs" are nucleic acid binding proteins secreted by numerous Proteobacterial species. TALE polypeptides contain a nucleic acid binding domain consisting of tandem repeats of highly conserved monomeric polypeptides that are primarily 33, 34, or 35 amino acids in length and differ from each other primarily at amino acid positions 12 and 13. In some embodiments, the nucleic acid is DNA.
[0127] As used herein, the term "polypeptide monomer" or "TALE monomer" will be used to refer to the highly conserved repeated polypeptide sequence within a TALE nucleic acid binding domain, and the term "repeat variable dinucleotide" or "RVD" will be used to refer to the highly variable amino acids at positions 12 and 13 of the polypeptide monomer.
[0128] As provided throughout this disclosure, the amino acid residues of the RVD are represented using the IUPAC single letter code for amino acids. A common designation for TALE monomers contained within a DNA binding domain is X1-11-(X12X13)-X14-33 or 34 or 35, where the subscripts indicate the amino acid position and X represents any amino acid. X12X13 represents the RVD. In some polypeptide monomers, the variable amino acid at position 13 is missing or absent, and in such polypeptide monomers the RVD consists of a single amino acid. In such cases, the RVD is instead represented by X. * where X represents X12, ( *) indicates the absence of X13. The DNA binding domain comprises several repeats of TALE monomers, represented as [X1-ll-(X12X13)-X14-33 or 34 or 35]z, where in advantageous embodiments z is at least 5-40. In further advantageous embodiments z is at least 10-26. A TALE monomer has a nucleotide binding affinity determined by the identity of the amino acids in its RVD. For example, a polypeptide monomer with an RVD of NI preferentially binds adenine (A), a polypeptide monomer with an RVD of NG preferentially binds thymine (T), a polypeptide monomer with an RVD of HD preferentially binds cytosine (C), and a polypeptide monomer with an RVD of NN preferentially binds both adenine (A) and guanine (G). In yet another embodiment of the present disclosure, a polypeptide monomer with an RVD of IG preferentially binds T. That is, the number and order of polypeptide monomer repeats in the nucleic acid binding domain of a TALE determines its nucleic acid target specificity. In further embodiments of the present disclosure, the polypeptide monomer with the NS RVD can recognize all four base pairs and bind to A, T, G, or C. The structure and function of TALEs are further described, for example, in Moscou et al., Science 326:1501 (2009); Boch et al., Science 326:1509-1512 (2009); and Zhang et al., Nature Biotechnology 29:149-153 (2011), each of which is incorporated by reference in its entirety. In certain embodiments, targeting is achieved by a polynucleic acid-binding TALEN fragment. In certain embodiments, the targeting domain comprises or consists of a catalytically inactive TALEN or a nucleic acid-binding fragment thereof.
[0129] Zinc finger nuclease (ZFN) system In certain embodiments, the targeting domain comprises or consists of a (modified) zinc finger nuclease (ZFN) system, which uses an artificial restriction enzyme generated by fusing a zinc finger DNA binding domain to a DNA cleavage domain that can be engineered to target a desired DNA sequence. Exemplary methods of genome editing using ZFNs can be found, for example, in U.S. Patent Nos. 6,534,261, 6,607,882, 6,746,838, 6,794,136, 6,824,978, 6,866,997, 6,933,113, 6,979,539, 7,013,219, 7,030,215, 7,220,719, 7,241,573, 7,241,574, 7,585,849, 7,595,376, 6,903,185, and 6,479,626, each of which is incorporated by reference in its entirety. By way of further guidance and without limitation, artificial zinc finger (ZF) technology involves the creation of arrays of ZF modules that target novel DNA binding sites in the genome. Each finger module in the ZF array targets three DNA bases. Customized arrays of individual zinc finger domains are assembled into ZF proteins (ZFPs). ZFPs may contain functional domains. The first synthetic zinc finger nucleases (ZFNs) were developed by fusing ZF proteins to the catalytic domain of the type IIS restriction enzyme FokI (Kim, YG et al., 1994, Chimeric restriction endonuclease, Proc. Natl. Acad. Sci. USA 91, 883-887; Kim, YG et al., 1996, Hybrid restriction enzymes: zinc finger fusions to FokI cleavage domain. Proc. Natl. Acad. Sci. USA 93, 1156-1160). Increased cleavage specificity can be achieved along with reduced off-target activity by using pairs of ZFN heterodimers, each targeting a different nucleotide sequence separated by a short spacer.(Doyon, Y. et al., 2011, Enhancing zinc-finger-nuclease activity with improved obligate heterodimeric architectures. Nat. Methods 8, pp. 74-79). ZFPs can also be designed as transcriptional activators and repressors and have been used to target many genes in many organisms. In certain embodiments, the targeting domain comprises or consists of a nucleic acid-binding zinc finger nuclease or a nucleic acid-binding fragment thereof. In certain embodiments, the nucleic acid that binds to the zinc finger nuclease (fragment) is catalytically inactive.
[0130] Meganuclease In certain embodiments, the targeting domain comprises a (modified) meganuclease, which is an endonuclease characterized by a large recognition site (double-stranded DNA sequence of 12-40 base pairs). Exemplary methods using meganucleases can be found in U.S. Patent Nos. 8,163,514, 8,133,697, 8,021,867, 8,119,361, 8,119,381, 8,124,369, and 8,129,134, each of which is incorporated by reference in its entirety. In certain embodiments, targeting is achieved by a polynucleic acid-binding meganuclease fragment. In certain embodiments, targeting is achieved by a polynucleic acid-binding catalytically inactive meganuclease (fragment). Thus, in certain embodiments, the targeting domain comprises or consists of a nucleic acid-binding meganuclease or a nucleic acid-binding fragment thereof.
[0131] CRISPR-Cas system In some embodiments, the DNA binding protein and single guide RNA sequence of the present disclosure are derived from a CRISPR-Cas system. The present disclosure provides a CRISPR / Cas9-based engineered system for use in genome editing and treating genetic diseases. The CRISPR / Cas9-based engineered system may be designed to target any gene, including genes involved in genetic diseases, liver diseases, and dysregulation of cholesterol, such as LDL, (e.g., VEGF). The present disclosure provides a CRISPR-Cas system that includes engineered Cas proteins and / or guide RNAs with desired specificity and activity (e.g., reducing or eliminating expression of the VEGF gene product). The CRISPR / Cas9-based system may include a Cas9 protein, a mutated Cas9 protein, or a Cas9 fusion protein (e.g., a DNMT3A-DNMT3L (3A3L)-dCas9-KRAB fusion molecule) and at least one sgRNA (e.g., a VEGF sgRNA). A Cas9 fusion protein may contain a domain (e.g., DNMT3A, DNMT3L, or KRAB) that has an activity different from that endogenous to Cas9.
[0132] Generally, a Cas protein (used interchangeably herein as CRISPR protein, CRISPR enzyme, CRISPR-Cas protein, CRISPR-Cas enzyme, Cas, CRISPR effector, or Cas effector protein) and / or a guide sequence are components of a CRISPR-Cas system. CRISPR-Cas system or CRISPR system collectively refers to transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (transcriptionally activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (including "direct repeats" and tracrRNA-processed partial direct repeats in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or "RNAs" as that term is used herein (e.g., RNAs that guide Cas, e.g., CRISPR RNA and transcriptionally activating (tracr) RNA, or single guide RNAs (also known as sgRNAs, chimeric RNAs) or other sequences and transcripts from a CRISPR locus.
[0133] In general, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at the site of target sequence (also referred to as protospacer in the context of endogenous CRISPR systems). In the engineered system of the present disclosure, direct repeats can include naturally occurring or non-naturally occurring sequences. The direct repeats of the present disclosure are not limited to naturally occurring lengths and sequences. In addition, the direct repeats of the present disclosure can include insertions of nucleotides such as sequences that bind to aptamers or adaptor proteins (for association with functional domains). In certain embodiments, one end of the direct repeat that includes insertions is approximately the first half of the short DR, and the end is approximately the second half of the short DR.
[0134] In the context of forming a CRISPR complex, "target sequence" or "target polynucleotide" refers to the sequence to which guide sequence is designed to have complementarity, and hybridization between the target sequence and guide sequence promotes the formation of a CRISPR complex.Target sequence may comprise any polynucleotide, for example, DNA or RNA polynucleotide.In some embodiments, target sequence is located in the nucleus or cytoplasm of a cell.
[0135] In general, a guide sequence (or spacer sequence) can be any polynucleotide sequence that has sufficient complementarity (e.g., perfect complementarity) with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more when optimally aligned using a suitable alignment algorithm.
[0136] In certain embodiments, the adjustment of cleavage efficiency can be utilized by introducing one or more mismatches, such as one or two mismatches between the spacer sequence and the target sequence, including the position of the mismatch along the spacer / target. The more central (i.e., not 3' or 5') the double mismatch is, the more the cleavage efficiency is affected. Thus, the cleavage efficiency can be adjusted by selecting the position of the mismatch along the spacer. As an example, when less than 100% cleavage of the target is desired (e.g., in a cell population), one or more, such as preferably two, mismatches between the spacer and the target sequence may be introduced into the spacer sequence. The more central the position of the mismatch is along the spacer, the lower the cleavage percentage.
[0137] The CRISPR-Cas system or its components may be used to introduce one or more mutations into a target locus or nucleic acid sequence. The mutations may include the introduction, deletion, or substitution of one or more nucleotides via a guide RNA or sgRNA in the respective target sequence of the cell. The mutations may include the introduction, deletion, or substitution of 1-75 nucleotides via a guide RNA in the respective target sequence of the cell.
[0138] Typically, in the context of an endogenous CRISPR-Cas system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of the target sequence), which may depend, for example, on secondary structure, particularly in the case of RNA targets. In some cases, in the context of an endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands (if applicable) in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of the target sequence).
[0139] In some embodiments, a guide RNA (capable of guiding Cas to a target locus) may comprise (1) a guide sequence capable of hybridizing to a target locus (a polynucleotide target locus, e.g., an RNA target locus) in a eukaryotic cell, and (2) direct repeat (DR) sequences, which are present in a single RNA, i.e., sgRNA (arranged in a 5' to 3' orientation) or crRNA.
[0140] For general information regarding CRISPR-Cas systems, their components, and delivery of those components, including methods, materials, delivery vehicles, vectors, particles, AAVs, and their production and use, including amounts and formulations, all useful in the practice of the present disclosure, see U.S. Pat. Nos. 8,999,641, 8,993,223, and 8,993,232, each of which is incorporated herein by reference in its entirety. 33, No. 8,945,839, No. 8,932,814, No. 8,906,616, No. 8,895,308, No. 8,889,418, No. 8,889,356, No. 8,871,445, No. 8,865,406, No. 8,795,965, No. 8,771,945, and No. 8,697,359, U.S. Patent Publication Nos. 2014-0310830 and 2014-0287938 A1, 2014-0273234 A1, 2014-0273232 A1, 2014-0273231 A1, 2014-0256046 A1, 2014-0248702 A1, 2014-0242700 A1, 2014-0242699 A1, 2014-0242664 A1, 2014-0234972 A1, 2014-0227787 A1, 2014-0189896 A1, 2014-0186958, 2014-0186919 A1, 2014-0186843 A1, 2014-0179770 A1, and 2014-0179006 A1, 2014-0170753, European Patents 2784162 B1 and 2771468 B1, European Patent Applications 2771468, 2764103, and 2784162, and PCT Patent Publications WO 2021 / 183807A1 (PCT / US2021 / 021973), WO 2014 / 093661 (PCT / US2013 / 074743), WO 2014 / 093694 (PCT / US2013 / 074790), WO 2014 / 093595 (PCT / US2013 / 074611), WO 2014 / 093718 (PCT / US2013 / 074825), WO 2014 / 093709(PCT / US2013 / 074812), WO 2014 / 093622(PCT / US2013 / 074667), WO 2014 / 093635(PCT / US2013 / 074691), WO 2014 / 093655(PCT / US2013 / 074736), WO2014 / 093712(PCT / US2013 / 074819), WO 2014 / 093701(PCT / US2013 / 074800), WO 2014 / 018423(PCT / US2013 / 051418), WO 2014 / 204723(PCT / US2014 / 041790), WO 2014 / 204724(PCT / US2014 / 041800), WO2014 / 204725(PCT / US2014 / 041803), WO 2014 / 204726(PCT / US2014 / 041804), WO See PCT Application No. 2014 / 204727 (PCT / US2014 / 041806), WO 2014 / 204728 (PCT / US2014 / 041808), and WO 2014 / 204729 (PCT / US2014 / 041809).
[0141] Cas proteins A Cas protein (e.g., an engineered Cas protein) can have substantially the same (e.g., 80%-100%, 90%-100%, 95%-100%, 98%-100%, 99%-100%, 99.9%-100%, or about 100%) nuclease activity as the wild-type counterpart Cas protein. In certain cases, the engineered Cas protein has a higher (e.g., at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% higher) nuclease activity than the wild-type counterpart Cas protein.
[0142] Alternatively or additionally, a Cas protein (e.g., an engineered Cas protein) may have at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% higher specificity than a wild-type counterpart Cas protein. In certain examples, a Cas protein (e.g., an engineered Cas protein) may have at least 30% higher specificity than a wild-type counterpart Cas protein. As used herein, the term "specificity" of Cas may correspond to the number or percentage of on-target polynucleotide cleavage events relative to the number or percentage of all polynucleotide cleavage events, including on-target and off-target events. The activity and specificity of the Cas proteins are consistent with that described in Hsu PD et al., DNA targeting specificity of RNA-guided Cas9 nucleases, Nat Biotechnol. 2013 Sep;31(9):827-832; and Slaymaker IM et al., Rationally engineered Cas9 nucleases with improved specificity, Science. 2016 Jan l;351(6268):84-88, which also describe exemplary methods for detecting Cas protein activity and specificity, and are incorporated herein by reference in their entirety, and described in detail elsewhere herein.
[0143] In some embodiments, the Cas protein (e.g., its RuvC domain) slides one base upstream (with respect to the PAM) to produce a sticky break, which can be filled in to result in a single base duplication (i.e., +1 insertion). Examples of +1 insertion positions are described in Zuo, Z. and Liu, J. (2016). Cas9-catalyzed DNA Cleavage Generates Staggered Ends: Evidence from Molecular Dynamics Simulations. Scientific Reports 6, 37584 pages. In some embodiments, the engineered Cas protein has a different +1 insertion frequency than the wild-type counterpart Cas protein. For example, the +1 insertion frequency when a guanine is at the -2 position with respect to the PAM is higher than the +1 insertion frequency when a thymidine, cytidine, or adenine is at the -2 position with respect to the PAM. In some cases, the +1 insertion is due to the host machinery in human cells. In some cases, the Cas protein can generate a sticky break. The sticky break can be a 1 bp or 1 nucleotide 5' overhang. The staggered cleavage may be at a 1 bp or 1 nucleotide 3' overhang.
[0144] The nucleic acid molecule encoding Cas may be codon-optimized. An example of a codon-optimized sequence is a sequence optimized for expression in a eukaryote, such as a human (i.e., optimized for expression in a human), or another eukaryote, animal, or mammal discussed herein. See, for example, the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (PCT / US2013 / 074667). While this is preferred, it is recognized that other examples are possible, and codon optimization for host species other than humans, or for specific organs, is known. In some embodiments, the enzyme coding sequence encoding Cas is codon-optimized for expression in a particular cell, such as a eukaryotic cell. The eukaryotic cell may be or be derived from a cell of a particular organism, such as a human or non-human eukaryote, or an animal or mammal, such as a mouse, rat, rabbit, dog, livestock, or a mammal, including but not limited to a non-human mammal or primate, discussed herein. In some embodiments, processes for modifying the genetic identity of the germline of humans and / or processes for modifying the genetic identity of animals that may cause suffering without any substantial medical benefit to humans or animals, as well as animals resulting from these processes, may be excluded. In general, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with a codon that is more or most frequently used in the genes of that host cell, while maintaining the native amino acid sequence. Various species exhibit a particular bias toward certain codons for certain amino acids. Codon bias (the difference in codon usage between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which is believed to depend, among other things, on the properties of the codon being translated and the availability of certain transfer RNA (tRNA) molecules. The predominance of selected tRNAs within a cell is generally a reflection of the codons most frequently used in peptide synthesis.Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the "Codon Usage Database" available at www.kazusa.orjp / codon / , and these tables can be adapted in several ways. See Nakamura, Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, for example, Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the Cas-encoding sequence correspond to the most frequently used codon for a particular amino acid.
[0145] In some embodiments, Cas protein can have nucleic acid cleavage activity.Cas protein can have RNA binding and DNA cleavage functions.In some embodiments, Cas can direct the cleavage of one or more nucleic acid strands at or near the position of target sequence, for example, in the target sequence and / or in the complement of the target sequence, or in the sequence associated with the target sequence, for example, within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the Cas protein may direct two or more cuts (e.g., one, two, three, four, five, or more cuts) of one or two strands within the target sequence, and / or within the complement of the target sequence or a sequence associated with the target sequence, and / or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the cuts may be blunt, i.e., generating blunt ends. In some embodiments, the cuts may be staggered, i.e., generating sticky ends.
[0146] In some embodiments, the vector encodes a nucleic acid-targeting Cas protein that may be mutated relative to the corresponding wild-type enzyme, whereby the mutated nucleic acid-targeting Cas protein lacks the ability to cleave one or two strands of the target polynucleotide containing the target sequence. For example, the HNH domain is altered or mutated to produce a mutated Cas that is substantially devoid of all DNA cleavage activity, for example, the DNA cleavage activity of the mutated enzyme is no more than about 25%, 10%, 5%, 1%, 0.1%, 0.01% or less than the nucleic acid cleavage activity of the non-mutated form of the enzyme. For example, the nucleic acid cleavage activity of the mutant form may be absent or negligible compared to the non-mutated form. As used herein, the term "derived" with respect to an enzyme means that the derived enzyme is based primarily on the wild-type enzyme in the sense that it has a high degree of sequence homology, but is mutated (modified) by methods known in the art or described herein.
[0147] Typically, in the context of an endogenous nucleic acid targeting system, formation of a nucleic acid targeting complex (comprising a guide RNA or crRNA hybridized to a target sequence and complexed with one or more effector proteins that target the nucleic acid) results in DNA strand cleavage in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of the target sequence). As used herein, the term "sequence associated with a target locus of interest" refers to a sequence near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of the target sequence, where the target sequence is contained within the target locus of interest).
[0148] It will be appreciated that effector proteins are based on or derived from enzymes, and thus in some embodiments the term "effector protein" does include "enzymes." However, it will also be appreciated that effector proteins may have DNA or RNA binding activity but not necessarily cleavage or damage activity, including dead Cas protein function, as required in some embodiments.
[0149] In some embodiments, the Cas protein may form a component of an inducible system. The inducible nature of the system will allow for spatiotemporal control of gene editing or gene expression using forms of energy. Forms of energy include, but are not limited to, electromagnetic radiation, acoustic energy, chemical energy, and thermal energy. Examples of inducible systems include tetracycline-inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcription activation systems (FKBP, ABA, etc.), or light-inducible systems (Phytochrome, LOV domain, or Cryptochrome). In one embodiment, the CRISPR effector protein may be part of a light-induced transcription effector (LITE) that directs alteration of transcription activity in a sequence-specific manner. The light component may include the CRISPR effector protein, a light-responsive cytochrome heterodimer (e.g., from Arabidopsis thaliana), and a transcription activation / repression domain. Further examples of inducible DNA binding proteins and methods for their use are provided in U.S. 61 / 736465 and 61 / 721,283, and WO 2014018423 A2, the entireties of which are incorporated by reference herein.
[0150] In some embodiments, the mutated Cas may have one or more mutations that result in reduced off-target effects, such as improved CRISPR enzymes for use in complexing with guide RNA to achieve modification to the target locus, but reduce or eliminate activity toward off-targets, as well as improved CRISPR enzymes for use in complexing with guide RNA to increase activity of the CRISPR enzyme, such as when complexed with guide RNA. It should be understood that the mutated enzymes described herein below may be used in any of the methods according to the present disclosure described elsewhere herein. Any of the methods, products, compositions, and uses described elsewhere herein are equally applicable with mutated CRISPR enzymes, as described in more detail below.
[0151] The methods and mutations that can be employed in various combinations to increase or decrease on-target and off-target activity and / or specificity, or to increase or decrease on-target and off-target binding and / or specificity, can be used to compensate for or enhance mutations or modifications made to promote other effects. Such mutations or modifications made to promote other effects include mutations or modifications to Cas and / or mutations or modifications made to guide RNA. The methods and mutations of the present disclosure are used to modulate the activity of Cas nuclease and / or binding to chemically modified guide RNA.
[0152] In certain embodiments, the catalytic activity of the Cas protein of the present disclosure is altered or modified. A mutated Cas is understood to have altered or modified catalytic activity if the catalytic activity is different from that of a corresponding wild-type Cas protein (e.g., a non-mutated Cas protein). Catalytic activity can be determined by means known in the art. As an example and without limitation, catalytic activity can be determined in vitro or in vivo by determining the percentage of indels (e.g., after a given time or at a given dose). In certain embodiments, catalytic activity is increased. In certain embodiments, catalytic activity is increased by at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In certain embodiments, catalytic activity is decreased. In certain embodiments, catalytic activity is reduced by at least 5%, preferably at least 10%, more preferably at least 20%, such as at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or (substantially) 100%. One or more mutations herein may inactivate catalytic activity, thereby reducing substantially all catalytic activity, reducing activity below detectable levels, or reducing to no measurable catalytic activity.
[0153] One or more characteristics of the engineered Cas protein may differ from the corresponding wild-type Cas protein. Examples of such characteristics include catalytic activity, gRNA binding, specificity of the Cas protein (e.g., specificity of editing a defined target), stability of the Cas protein, off-target binding, target binding, protease activity, nickase activity, PFS recognition. In some examples, the engineered Cas protein may include one or more mutations of the corresponding wild-type Cas protein. In some embodiments, the catalytic activity of the engineered Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the catalytic activity of the engineered Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the gRNA binding of the engineered Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the gRNA binding of the engineered Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the specificity of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the specificity of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the stability of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the stability of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the engineered Cas protein further comprises one or more mutations that inactivate catalytic activity. In some embodiments, the off-target binding of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the off-target binding of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the target binding of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the target binding of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the engineered Cas protein has increased protease activity or polynucleotide binding capacity compared to the corresponding wild-type Cas protein.In some embodiments, the PFS recognition is altered compared to the corresponding wild-type Cas protein.
[0154] Examples of Cas proteins Examples of Cas proteins include class I (e.g., type I, III, and IV) and class 2 (e.g., type II, V, and VI) Cas proteins, such as Cas9, Cas12 (e.g., Cas12a, Cas12b, Cas12c, Cas12d), Cas13 (e.g., Cas13a, Cas13b, Cas13c, Cas13d), CasX, CasY, Cas14, variants thereof (e.g., mutants, truncations), homologs thereof, and orthologs thereof. The terms "ortholog" and "homolog" are well known in the art. As further guidance, a "homolog" of a protein as used herein is a protein of the same species that performs the same or similar function as the protein to which it is a homolog. Homolog proteins may be structurally related, but may not be related, or may only be partially structurally related. As used herein, an "ortholog" of a protein is a protein in a different species that performs the same or similar function as the protein to which it is an ortholog. Orthologous proteins may be structurally related, but may also be unrelated, or only partially structurally related.
[0155] Class 2 Cas proteins In some embodiments, the Cas protein is a class 2 Cas protein, i.e., a Cas protein of a class 2 CRISPR-Cas system. The class 2 CRISPR-Cas system can be a subtype, such as type II-A, type II-B, type II-C, type VA, type VB, type VC, or type VU. In some embodiments, the Cas protein is Cas9, Cas12a, Cas12b, Cas12c, or Cas12d. In some embodiments, Cas9 can be SpCas9, SaCas9, StCas9, and other Cas9 orthologues. Cas12 can be Cas12a, Cas12b, and Cas12c, including FnCas12a, or a homolog or orthologue thereof. Definitions and exemplary members of CRISPR-Cas systems include those described in Kira S. Makarova and Eugene V. Koonin, Annotation and Classification of CRISPR-Cas systems, Methods Mol Biol. 2015;1311:47-75; and Sergey Shmakov et al., Diversity and evolution of class 2 CRISPR-Cas systems, Nat Rev Microbial. 2017 Mar;15(3):169-182.
[0156] Cas protein linker In some examples, the Cas protein comprises at least one RuvC domain and at least one HNH domain. The Cas protein may further comprise a first and a second linker domain connecting the RuvC domain and the HNH domain. The first linker (L1) and the second linker (L2) connecting the HNH domain and the RuvC domain in Cas9 are described in Nishimasu, H. et al., "Crystal structure of Cas9 in complex with guide RNA and target RNA," Cell 156 (Feb. 27, 2014): 935-949, and Ribeiro, L. et al., (2018) "Protein engineering strategies to expand CRISPR-Cas9 applications," International Journal of Genomics, Vol. 2018, Article ID 1652567 (doi.org / 10.1155 / 2018 / 1652567). Figure 1 of Ribeiro shows the overall organization, structure, and function of Cas9 and is specifically incorporated herein by reference. Specifically, Figure 1A shows a schematic representation of the domain organization of SpCas9 showing the genetic organization of the HNH and RuvC domains, including linkers L1 (spanning amino acids 765-780) and L2 (spanning amino acids 906-918), as described herein.
[0157] Similarly, the domain organization of Staphylococcus aureus Cas9 (SaCas9) can be used to reference the first and second linker domains. In one aspect, the linker 1 domain region spans residues 481-519 and connects the RuvC-II domain to the HNH domain in SaCas9. In some embodiments, the linker 2 region spans residues 629-649 and connects the RuvC-III and HNH domains of SaCas9. Thus, the first and / or second linker domains may be mutated in a Cas9 ortholog and may reference amino acid residues that correspond to amino acids in wild-type SaCas9. See Nishimasu, Cell. 2015 Aug 27;162(5):1113-1126; doi:10.1016 / j.cell.2015.08.007, incorporated herein by reference. In particular, Figures 1, S1-S3 of Nishimasu detail the domain organization of the Cas9 protein and are specifically incorporated by reference herein for their teachings.
[0158] The first and second linkers may comprise about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or more amino acids. The first and second linkers may correspond to wild-type linkers. In some embodiments, the first and second linkers may comprise one or more mutations in the first and / or second linker. In one embodiment, the first and / or second linker comprises one or more mutations that improve the specificity of the Cas9 protein.
[0159] In some embodiments, the linkers L1 and L2 connecting the HNH and RuvC domains of Cas9 comprise wild-type amino acid sequences. In some embodiments, the linkers connecting the HNH and RuvC domains comprise mutations at one or more amino acids. In an exemplary embodiment, the first linker (L1) comprises a mutation corresponding to the amino acid T769I of SpCas9, and / or the second linker (L2) comprises a mutation corresponding to the amino acid G915M of SpCas9. In an exemplary embodiment, the one or more linker mutations, e.g., T769I and G915M, confer improved specificity to the Cas9 protein.
[0160] In one embodiment, one or more mutations in the first and second linkers may be combined with one or more mutations in other parts of the Cas9 protein, as described herein, to further improve specificity and / or retain activity substantially equivalent to the wild-type Cas9 protein. In one embodiment, mutations in the linkers and / or further mutations in the Cas protein can be identified using the methods detailed herein that enhance / improve specificity and substantially retain wild-type activity to wild-type Cas9.
[0161] Class 2, type II Cas proteins (e.g., Cas9) In some embodiments, the Cas protein may be a class 2, type II Cas protein of a CRISPR-Cas system (a type II Cas protein). In some embodiments, the Cas protein may be a class 2, type II Cas protein, such as Cas9. In some embodiments, the CRISPR / Cas9-based system may include a Cas9 protein or fragment thereof, a Cas9 fusion protein, a nucleic acid encoding a Cas9 protein or fragment thereof, or a nucleic acid encoding a Cas9 fusion protein. "Cas9 (CRISPR associated protein 9)" refers to a polypeptide or fragment thereof having at least about 85% amino acid identity with NCBI Accession No. NP_269215 and having RNA binding activity, DNA binding activity, and / or DNA cleavage activity (e.g., endonuclease or nickase activity). "Cas9 function" may be defined by any of several assays, including, but not limited to, a fluorescence polarization-based nucleic acid binding assay, a fluorescence polarization-based strand invasion assay, a transcription assay, an EGFP disruption assay, a DNA cleavage assay, and / or a surveyor assay, for example as described herein. "Cas9 nucleic acid molecule" refers to a polynucleotide encoding a Cas9 polypeptide or a fragment thereof. An exemplary Cas9 nucleic acid molecule sequence is provided in genomic sequence number NC_002737. In some embodiments, inhibitors of Cas9, such as naturally occurring S. pyogenes Cas9 (SpCas9) or S. aureus Cas9 (SaCas9) or variants thereof, are disclosed herein. Cas9 recognizes foreign DNA using a protospacer adjacent motif (PAM) sequence and base pairing of the target DNA with a guide RNA (gRNA). The relative ease of inducing targeted strand breaks at any genomic locus by Cas9 has enabled efficient genome editing in many cell types and organisms. Cas9 derivatives can also be used as transcriptional activators / repressors.
[0162] In some cases, the CRISPR-Cas protein is Cas9 or a variant thereof. In some cases, Cas9 can be wild-type Cas9, including any naturally occurring bacterial Cas9. Cas9 orthologs typically share a common organization of 3-4 RuvC domains and an HNH domain. The 5'-most RuvC domain cleaves the non-complementary strand, and the HNH domain cleaves the complementary strand. All designations are in reference to the guide sequence. The catalytic residues of the 5' RuvC domain are identified by homology comparison of the Cas9 of interest with other Cas9 orthologs [from S. pyogenes type II CRISPR locus, S. thermophilus CRISPR locus 1, S. thermophilus CRISPR locus 3, and Franciscilla novicida type II CRISPR locus], and the conserved Asp residue (D10) is mutated to alanine, converting Cas9 into a complementary strand cleavage enzyme. Thus, the Cas enzyme may be a wild-type Cas9, including any naturally occurring bacterial Cas9. The CRISPR, Cas, or Cas9 enzyme may be codon-optimized or may be a modified version, including any chimera, mutant, homolog, or ortholog. In further aspects of the present disclosure, the Cas9 enzyme may include one or more mutations and may be used as a general DNA-binding protein, with or without fusion with a functional domain.
[0163] The mutation may be an artificially introduced mutation, or a gain-of-function mutation, or a loss-of-function mutation. In some embodiments, the transcription activation domain may be VP64. In some embodiments, the transcription repressor domain may be KRAB or SID4X. Other aspects of the present disclosure relate to mutated Cas9 enzymes fused with domains including, but not limited to, nucleases, transcription activators, repressors, recombinases, transposases, histone remodelers, demethylases, DNA methyltransferases, cryptochromes, light-inducible / regulatory domains, or chemically-inducible / regulatory domains. The present disclosure may include sgRNAs or tracrRNAs, or guides or chimeric guide sequences that allow for enhanced performance of these RNAs in cells. The type II CRISPR enzyme may be any Cas enzyme. In some cases, the Cas9 enzyme is or is derived from SpCas9 or SaCas9. As used herein, the term "derived" in reference to an enzyme means that the derived enzyme is based primarily on a wild-type enzyme in the sense of having a high degree of sequence identity, but is mutated (modified) in some manner known in the art or described herein. In one example, the mutations may include one or more mutations in the first linker domain, the second linker domain, and / or other portions of the protein. The high degree of sequence identity may include at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more relative to the wild-type enzyme.
[0164] The Cas enzyme may be an identified Cas, since it may refer to a general class of enzymes that share homology with the largest nucleases with multiple nuclease domains from type II CRISPR systems. In some cases, the Cas9 enzyme is derived from or derived from SpCas9 (S. pyogenes Cas9) or saCas9 (S. aureus Cas9). "StCas9" refers to wild-type Cas9 from S. thermophilus (UniProt ID: G3ECR1). Similarly, "SpCas9" refers to wild-type Cas9 from S. pyogenes (UniProt ID: Q99ZW2). As used herein, the term "derived" in relation to an enzyme means that the derived enzyme is based primarily on the wild-type enzyme in the sense of having a high degree of sequence identity, but is mutated (modified) in some manner known in the art or described herein. It will be appreciated that the terms Cas and CRISPR enzyme are generally used interchangeably herein unless otherwise clear. As noted above, much of the residue numbering used herein refers to the Cas9 enzyme from the Type II CRISPR locus in Streptococcus pyogenes.
[0165] In certain embodiments, the effector protein is selected from the group consisting of Streptococcus, Campylobacter, Nitratiphracter, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Spirochete, Lactobacillus, Eubacterium, Corynebacterium, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus us, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, Acidaminococcus, Streptococcus, Campylobacter, Nitratifractor, Staphylococcus Bacteroides, Parvibacterium, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Spirochete, Lactobacillus, Eubacterium, Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium,The Cas9 effector protein is derived from or originates from an organism from a genera including Spirochete, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratiphracter, Mycoplasma, or Campylobacter.
[0166] In some embodiments, the Cas9 protein is selected from the group consisting of S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, S. pneumonia, C. jejuni, C. coli, N. salsuginis, N. tergarcus, S. auricularis, S. carnosus, N. meningitides, N. gonorrhoeae, L. monocytogenes, L. ivanovii, C. botulinum, C. botulinum, C. difficile, C. tetani, or C. sordellii, Francisella tularensis 1, Francisella tularensis subsp.novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011 GWA2_33_10, Parcubacteria bacterium GW2011 GWC2_44_17, Smithella subsp. SCADC, Acidaminococcus subsp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis3, Prevotella disiens, and Porphyromonas macacae. In some embodiments, the Cas9 effector protein is derived from or originates from an organism selected from Streptococcus pyogenes, Staphylococcus aureus, or Streptococcus thermophilus Cas9 derived organisms.
[0167] In a more preferred embodiment, the Cas9 protein is derived from a bacterial species selected from Streptococcus pyogenes, Staphylococcus aureus, or Streptococcus thermophilus Cas9. In certain embodiments, the Cas9 is derived from Francisella tularensis 1, Prevotella arvensis, Lachnospiraceae bacterium MC20171, Butyrivibrio proteoclasticus, Pellegrinibacterium bacterium GW2011 GWA2 33 JO, Parvum bacterium b ... GWC2_44_17, Smicella subsp. SCADC, Acidaminococcus subsp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma thermitum, Eubacterium eligens, Morachesella bobocri237, Leptospira inadae, Lachnospiraceae bacterium ND2006, Porphyromonas clevioricanis3, Prevotella digens, and Porphyromonas macacae. In certain embodiments, the Cas9 protein is derived from a bacterial species selected from Acidaminococcus subsp. BV3L6, Lachnospiraceae bacterium MA2020. In certain embodiments, the effector protein is derived from a subspecies of Francisella tularensis 1, including but not limited to Francisella tularensis subspecies novicida.
[0168] The Cas9 enzymes were identified in S. pyogenes serovar M1 (UniProt ID: Q99ZW2), S. aureus Cas9 (UniProt ID: J7RUA5), Eubacterium ventriosum Cas9 (UniProt ID: A5Z395), Azospirillum (strain B510) Cas9 (UniProt ID: D3NT09), Gluconacetobacter diazotrophicus (strain ATCC 49037) Cas9 (UnitProt ID: A9HKP2), Neisseria cinerea Cas9 (UniProt ID: D0W2Z9), Roseburia intestinalis Cas9 (UniProt ID: C7G697), Parvibaculum labamentivorans (strain DS-1) Cas9 (UniProt ID: A7HP89), Nitratifructator sarsuginis (DSM) Cas9 (UniProt ID: A7HP89), and 16511) Cas9 (UniProt ID: E6WZS9), Campylobacter lari Cas9 (UniProt ID: G1UFN3).
[0169] Enzymatic action by Cas9 derived from Streptococcus pyogenes or any closely related Cas9 generates a double-stranded break at the target site sequence that hybridizes with 20 nucleotides of the guide sequence and has a protospacer adjacent motif (PAM) sequence (examples include NGG / NRG or a PAM that can be determined as described herein) after 20 nucleotides of the target sequence. CRISPR activity by Cas9 for site-specific DNA recognition and cleavage is defined by the guide sequence, the tracr sequence that partially hybridizes with the guide sequence, and the PAM sequence. Further aspects of the CRISPR system are described in Karginov and Hannon, The CRISPR system: small RNA-guided defense in bacteria and archaea, Mole Cell 2010, January 15;37(1):7. The type II CRISPR locus from Streptococcus pyogenes SF370 contains a cluster of four genes, Cas9, Cas1, Cas2, and Csnl, as well as two non-coding RNA elements, tracrRNA, and a characteristic array of repeat sequences (direct repeats) flanked by short stretches of non-repetitive sequences (spacers, each ~30 bp). In this system, targeted DNA double-strand breaks (DSBs) are generated in four sequential steps. First, two non-coding RNAs, the pre-crRNA array and tracrRNA, are transcribed from the CRISPR locus. Second, tracrRNA hybridizes to the direct repeats of the pre-crRNA, which are then processed into mature crRNAs containing distinct spacer sequences. Third, the mature crRNA:tracrRNA complex directs Cas9 to a DNA target consisting of a protospacer and a corresponding PAM via heteroduplex formation between the spacer region of the crRNA and the protospacer DNA. Finally, Cas9 mediates cleavage of the target DNA upstream of the PAM to create a DSB in the protospacer. Pre-crRNA arrays consisting of a single spacer flanked by two direct repeats (DRs) are also encompassed by the term "tracr mate sequence."In certain embodiments, Cas9 can be constitutively present, inducibly present, conditionally present, or administered or delivered.Optimization of Cas9 can be used to enhance function or develop new function.Chimeric Cas9 protein can be generated, and Cas9 can be used as general DNA binding protein.The structural information provided about Cas9 can be used to further manipulate and optimize CRISPR-Cas system, and can be extended to explore the structure-function relationship in other CRISPR enzyme systems, particularly in other type II CRISPR enzymes or Cas9 orthologs. The crystal structure information (described in U.S. Provisional Application Nos. 61 / 915,251, filed December 12, 2013, 61 / 930,214, filed January 22, 2014, and 61 / 980,012, filed April 15, 2014, and Nishimasu et al., "Crystal Structure of Cas9 in Complex with Guide RNA and Target DNA," Cell 156(5):935-949, DOI: http: / / dx.doi.org / 10.1016 / j.cell.2014.02.001 (2014), each and every one of which is incorporated herein by reference) provides structural information for truncating and creating modular or multipart CRISPR enzymes that can be incorporated into inducible CRISPR-Cas systems. In particular, structural information is provided for S. pyogenes Cas9 (SpCas9), which may be extended to other Cas9 orthologues or other type II CRISPR enzymes. The Cas9 gene is found in several diverse bacterial genomes, typically at the same locus as the casl, cas2, and cas4 genes, as well as CRISPR cassettes. Furthermore, the Cas9 protein contains a readily identifiable C-terminal region that is homologous to the transposon ORF-B, and contains an active RuvC-like nuclease, an arginine-rich region.
[0170] dCas9 Cas9 protein can be mutated so that its nuclease activity is not inactivated. Recently, inactivated Cas9 protein from S. pyogenes (iCas9, also referred to as "dCas9"), which does not have endonuclease activity, has been targeted to genes in bacteria, yeast, and human cells by gRNA to silence gene expression by steric hindrance. As used herein, "dCas molecule" may refer to dCas protein or a fragment thereof. As used herein, "dCas molecule" may refer to dCas protein or a fragment thereof. As used herein, the terms "iCas" and "dCas" are used interchangeably and refer to catalytically inactive CRISPR-associated proteins. In one embodiment, the dCas molecule comprises one or more mutations in the DNA cleavage domain. In one embodiment, the dCas molecule comprises one or more mutations in the RuvC or domain. In one embodiment, the dCas molecule comprises one or more mutations in both the RuvC and HNH domains. In one embodiment, the dCas molecule is a fragment of a wild-type Cas molecule. In one embodiment, the dCas molecule comprises a functional domain from a wild-type Cas molecule, the functional domain being selected from a Reel domain, a bridge helix domain, or a PAM interaction domain. In one embodiment, the nuclease activity of the dCas molecule is reduced by at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% compared to the activity of the corresponding wild-type Cas molecule.
[0171] A suitable dCas molecule can be derived from a wild-type Cas molecule. The Cas molecule can be from a type I, type II, or type III CRISPR-Cas system. In one embodiment, a suitable dCas molecule can be derived from a Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, or Cas10 molecule. In one embodiment, the dCas molecule is derived from a Cas9 molecule. The dCas9 molecule can be obtained by introducing point mutations (e.g., substitutions, deletions, or additions) into the Cas9 molecule, for example in the DNA cleavage domain, for example in the nuclease domain, for example in the RuvC and / or HNH domain. See, for example, Jinek et al., Science (2012) 337:816-21, the entirety of which is incorporated herein by reference. For example, by introducing two point mutations in RuvC and HNH domain, Cas9 nuclease activity is reduced while retaining Cas9 sgRNA and DNA binding activity.In one embodiment, the two point mutations in RuvC and HNH active sites are D10A and H840A mutations in S. pyogenes Cas9 molecule.Alternatively, D10 and H840 can be deleted to disable Cas9 nuclease activity while retaining S. pyogenes Cas9 molecule sgRNA and DNA binding activity.In one embodiment, the two point mutations in RuvC and HNH active sites are D10A and N580A mutations in S. pyogenes Cas9 molecule.
[0172] In various embodiments, the disclosure includes variants or mutants of any of the dCas molecules or variants thereof. All variants and mutants of dCas9 are also known as SpCas9 (Cas9 isolated from Streptococcus pyogenes), SaCas9 (Cas9 isolated from Staphylococcus aureus), StCas9 (Cas9 isolated from Streptococcus thermophilus), NmCas9 (Cas9 isolated from Neisseria meningitides), FnCas9 (Cas9 isolated from Francisella novicida), CjCas9 (Cas9 isolated from Campylobacter jejuni), ScCas9 [Cas9 isolated from Streptococcus canis], as well as any variants and mutant forms of the above Cas9, such as high-fidelity Cas9 (Kleinstiver et al., Nature. 2016 Jan 28) and enhanced SpCas9 (Slaymaker et al., Sciences. 2016 Any of the above-listed compounds may be used in the methods, compositions, or kits disclosed herein, including, but not limited to, those derived from the method, composition, or kit disclosed herein (Jan 01, 2001). This list is only intended to provide some exemplary options and is not intended to be exhaustive.
[0173] In one embodiment, the dCas molecule is a Streptococcus pyogenes dCas9 molecule comprising mutations at D10 and / or H840, numbered according to SEQ ID NO: 1. In one embodiment, the dCas molecule is a Streptococcus pyogenes dCas9 molecule comprising mutations at D10A and / or H840A, numbered according to SEQ ID NO: 1.
[0174] Streptococcus pyogenes dCas9
[0175] In one embodiment, the dCas9 molecule is a Staphylococcus aureus dCas9 molecule comprising the amino acid sequence of SEQ ID NO:2 or 3, a sequence substantially identical to SEQ ID NO:2 or 3 (e.g., at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more in sequence identity), or a sequence having one, two, three, four, five or more alterations, e.g., amino acid substitutions, insertions, or deletions, relative to SEQ ID NO:2 or 3, or any fragment thereof.
[0176] Similar mutations can also be applied to any other naturally occurring Cas9 (e.g., Cas9 from other species) or engineered Cas9 molecules. In certain embodiments, the dCas9 molecule is a Streptococcus pyogenes dCas9 molecule, a Staphylococcus aureus dCas9 molecule, a Campylobacter jejuni dCas9 molecule, a Corynebacterium diphtheriae dCas9 molecule, a Eubacterium ventriosum dCas9 molecule, a Streptococcus pasteurianus dCas9 molecule, a Lactobacillus farciminis dCas9 molecule, a Spirochaete globus dCas9 molecule, a Azospirillum (B 510 strain) dCas9 molecule, a Gluconacetobacter diazotrophicus dCas9 molecule, a Neisseria cinerea dCas9 molecule, a Roseburia intestinalis dCas9 molecule, a Parvibaculum labamentivorans dCas9 molecule, a Nitratifructator sarusginis (DSM 1651 The dCas9 molecule includes a Campylobacter lari (strain CF89-12) dCas9 molecule, a Streptococcus thermophilus (strain LMD-9) dCas9 molecule, or a fragment thereof.
[0177] In certain embodiments, the present disclosure provides a dCas9 molecule for a Streptococcus pyogenes dCas9 molecule, a Staphylococcus aureus dCas9 molecule, a Campylobacter jejuni dCas9 molecule, a Corynebacterium diphtheriae dCas9 molecule, a Eubacterium ventriosum dCas9 molecule, a Streptococcus pasteurianus dCas9 molecule, a Lactobacillus farciminis dCas9 molecule, a Spirochaete globus dCas9 molecule, a Azospirillum (B 510 strain) dCas9 molecule, a Gluconacetobacter diazotrophicus dCas9 molecule, a Neisseria cinerea dCas9 molecule, a Roseburia intestinalis dCas9 molecule, a Parvibaculum labamentivorans dCas9 molecule, a Nitratifructator sarusginis (DSM 1651 The present invention provides a vector comprising a nucleotide sequence encoding a Campylobacter lari (strain CF89-12) dCas9 molecule, a Streptococcus thermophilus (strain LMD-9) dCas9 molecule, or a fragment thereof.
[0178] Exemplary dCas9 proteins include, but are not limited to, those listed in Table 1.
[0179] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8]
[0180] Cas9 fusion protein The CRISPR / Cas9 system may comprise a fusion molecule [e.g., DNMT3A-DNMT3L (3A3L)-dCas9-KRAB]. The fusion molecule may comprise at least one DNA binding protein (e.g., dCas9) and at least one modulator of gene expression (e.g., KRAB, DNMT3A, DNMT3L, DNMT3A-DNMT3L fusion peptide). In some embodiments, the modulator of gene expression is selected from a repressor of gene expression (e.g., KRAB), an activator of gene expression, or a modulator of extragenic modification (e.g., DNMT3A, DNMT3L, DNMT3A-DNMT3L fusion protein), or any combination thereof. Various modulators of gene expression are known in the art. See, for example, Thakore et al., Nat Methods. 2016;13:127-37, which is incorporated herein by reference in its entirety.
[0181] Repressors of gene expression In some embodiments, the modulator of gene expression comprises a repressor of gene expression. The repressor can be any known repressor of gene expression, such as a repressor selected from Kruppel-associated box (KRAB) domain, mSin3 interaction domain (SID), MAX-interacting protein 1 (MXI1), chromoshadow domain, EAR repression domain (SRDX), eukaryotic release factor 1 (ERFl), eukaryotic release factor 3 (ERF3), tetracycline repressor, lad repressor, Catharanthus roseus G-box binding factor 1 and 2, Drosophila Groucho, tripartite motif-containing 28 (TRTM28), nuclear receptor co-repressor 1, nuclear receptor co-repressor 2, or fragments or fusions thereof.
[0182] Kruppel related box (KRAB) The KRAB domain is a type of transcriptional repression domain present in the N-terminal portion of many zinc finger protein transcription factors. When tethered to target DNA by the DNA binding domain, the KRAB domain functions as a transcriptional repressor. The KRAB domain is enriched in charged amino acids and may be divided into subdomains A and B. The KRAB A and B subdomains may be separated by a variable spacer segment, and many KRAB proteins contain only the A subdomain. A 45-amino acid sequence in the KRAB A subdomain has been shown to be important for transcriptional repression. The B subdomain does not repress transcription by itself, but enhances the repression exerted by the KRAB A subdomain. The KRAB domain recruits the co-repressors KAP1 (KRAB-associated protein 1, also known as transcriptional intermediary 1 beta, KRAB-A interacting protein, and tripartite motif protein 28) and heterochromatin protein 1 (HP1), as well as other chromatin regulatory proteins, resulting in transcriptional repression through the formation of heterochromatin. In one embodiment, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to a KRAB domain or a fragment thereof. In one embodiment, the KRAB domain or a fragment thereof is fused to the N-terminus of the dCas9 molecule. In one embodiment, the KRAB domain or a fragment thereof is fused to the C-terminus of the dCas9 molecule. In one embodiment, the KRAB domain or a fragment thereof is fused to both the N-terminus and the C-terminus of the dCas9 molecule. In one embodiment, the fusion molecule comprises a sequence of SEQ ID NO:22, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO:22, or a sequence having one, two, three, four, five or more changes, such as amino acid substitutions, insertions, or deletions, relative to SEQ ID NO:22, or any fragment thereof.
[0183] Exemplary KRAB domain sequence: RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEP (SEQ ID NO:22)
[0184] mSin3 interaction domain (SID) The mSin3 interaction domain (SID) is an interaction domain present on several transcriptional repressor proteins. It interacts with the paired amphipathic alpha helix 2 (PAH2) domain of mSin3, a transcriptional repressor domain bound to transcriptional repressor proteins such as mSin3 A co-repressor. In one embodiment, the methods and compositions disclosed herein include fusion molecules comprising a dCas9 molecule fused to a mSin3 interaction domain or a fragment thereof. In one embodiment, the methods and compositions disclosed herein include fusion molecules comprising a dCas9 molecule fused to four linked mSin3 interaction domains (SID4X). In one embodiment, the four linked mSin3 interaction domains (SID4X) are fused to the C-terminus of the dCas9 molecule.
[0185] MAX-interacting protein 1 (MXI1) Mxi1 is a repressor of MYC. Mxi1 antagonizes the transcriptional activity of MYC, possibly by competing for binding to MYC-associated factor X (MAX), which is required for MYC to bind and function. In one embodiment, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused with Mxi1 or a fragment thereof. In one embodiment, Mxi1 is fused to the C-terminus of the dCas9 molecule.
[0186] Activators of gene expression In some embodiments, the modulator of gene expression comprises an activator of gene expression. The activator can be any known activator of gene expression, such as a VP16 activation domain, a VP64 activation domain, a p65 activation domain, an Epstein-Barr virus R transactivator Rta molecule, or a fragment thereof. Activators that can be used with dCas9 molecules are known in the art. See, for example, Chavez et al., Nat Methods. (2016) 13:563-67, the entirety of which is incorporated herein by reference.
[0187] VP16, VP64, VP160 VP16 is a 16 amino acid viral protein sequence that recruits transcriptional activators to promoters and enhancers. VP64 is a transcriptional activator that contains four copies of VP16, such as a molecule that contains four tandem copies of VP16 connected by a Gly-Ser linker. VP160 is a transcriptional activator that contains ten copies of VP16. In one embodiment, the methods and compositions disclosed herein include fusion molecules that contain a dCas9 molecule fused to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies of VP16. In one embodiment, the methods and compositions disclosed herein include fusion molecules that contain a dCas9 molecule fused to VP64. In one embodiment, the methods and compositions disclosed herein include fusion molecules that contain a dCas9 molecule fused to VP160. In one embodiment, VP64 is fused to the C-terminus, N-terminus, or both the N-terminus and C-terminus of the dCas9 molecule.
[0188] p65 activation domain (p65AD) p65AD is the main transactivation domain of the 65 kDa polypeptide of the nuclear form of F-κB transcription factor.An exemplary sequence of human transcription factor p65 is available in the Uniprot database under accession number Q04206.In one embodiment, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused with p65 or a fragment thereof, such as p65AD.
[0189] Epstein-Barr virus (EBV) R transactivator (Rta) EBV immediate early protein Rta is a transcriptional activator that induces lytic gene expression and induces viral reactivation. In one embodiment, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to Rta or a fragment thereof.
[0190] VP64, p65, Rta fusion In one embodiment, the methods and compositions disclosed herein include fusion molecules comprising dCas9 molecules fused with VP64, p65, Rta, or any combination thereof. In this, the tripartite activator VP64-p65-Rta (also known as VPR), in which three transcription activation domains are fused with a short amino acid linker, can effectively upregulate target gene expression when fused with dCas9 molecules. In one embodiment, the methods and compositions disclosed herein include fusion molecules comprising dCas9 molecules fused with VPR.
[0191] Synergistic Activation Mediator (SAM) In one embodiment, the methods and compositions disclosed herein include a CRISPR-Cas system that includes three components: (1) a dCas9-VP64 fusion, (2) a gRNA incorporating two MS2 RNA aptamers in the tetraloop and stem loop, and (3) an MS2-P65-HSF1 activation helper protein. This system, named synergistic activation mediator (SAM), brings together three activator domains VP64, P65, and HSFl, and is described in Konermann et al., Nature. 2015;517:583-8, which is incorporated herein by reference in its entirety.
[0192] Ldbl self-association domain In one embodiment, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to an Ldbl self-association domain, which recruits enhancer-associated endogenous Ldbl.
[0193] Modulators of extragenic modifications In one embodiment, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused with a modulator of gene expression. In some embodiments, the modulator of gene expression comprises a modulator of extragenic modification. In one embodiment, the fusion molecule regulates the expression of a target gene through extragenic modification, for example, through histone acetylation or methylation, or DNA methylation, at the control element of the target gene, for example, promoter, enhancer, or transcription start site. The modulator can be any known modulator of extragenic modification, for example, histone acetyltransferase (e.g., p300 catalytic domain), histone deacetylase, histone methyltransferase [e.g., SUV39H1 or G9a (EHMT2)], histone demethylase (e.g., LSD1), DNA methyltransferase (e.g., DNMT3a or DNMT3a-DNMT3L), DNA demethylase (e.g., TET1 catalytic domain or TDG), or a fragment thereof.
[0194] Histone modifying activity In some embodiments, a modulator of an extragenic modification may have histone modifying activity, which may include, but is not limited to, histone deacetylase, histone acetyltransferase, histone demethylase, or histone methyltransferase activity.
[0195] In some embodiments, the modulator of extragenic modification may have histone acetyltransferase activity. Histone acetyltransferase may be p300 or CREB binding protein (CBP) protein or fragment thereof. In some embodiments, the methods and compositions disclosed herein comprise fusion molecules comprising dCas9 molecules fused with acetyltransferase p300 or fragment thereof, such as the catalytic core of p300. In some embodiments, the methods and compositions disclosed herein comprise fusion molecules comprising dCas9 molecules fused with CREB binding protein (CBP) protein or fragment thereof.
[0196] In some embodiments, the modulator of extragenic modification may comprise histone demethylase activity. For example, the modulator of extragenic modification may comprise an enzyme that removes methyl (CH3-) groups from nucleic acids or proteins (e.g., histones). In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule that comprises a dCas9 molecule fused with Lys-specific histone demethylase 1 (LSD1) or a fragment thereof.
[0197] In some embodiments, the modulator of extragenic modification may have histone methyltransferase activity.In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused with SUV39H1 or a fragment thereof.In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused with G9a (EHMT2) or a fragment thereof.
[0198] DNA demethylase activity In some embodiments, the modulator of extragenic modification may have DNA demethylase activity. For example, the modulator of extragenic modification may covert methyl groups to hydroxymethylcytosine as a mechanism for demethylating DNA. In some embodiments, the methods and compositions disclosed herein comprise fusion molecules comprising dCas9 molecules fused with 10-11 translocation methylcytosine dioxygenase 1 (TET1) or fragments thereof. In some embodiments, the methods and compositions disclosed herein comprise fusion molecules comprising dCas9 molecules fused with thymine DNA glycosylase (TDG) or fragments thereof.
[0199] DNA methylase activity In some embodiments, the modulator of the extragenic modification may have DNA methylase activity. For example, the modulator of the extragenic modification may have methylase activity involved in the transfer of methyl groups to DNA, RNA, proteins, small molecules, cytosine, or adenine. In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to DNMT3A or a fragment thereof. In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to DNMT3L or a fragment thereof. In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to DNMT3L and DNMT3L or a fragment thereof. In some embodiments, the methods and compositions disclosed herein comprise a fusion molecule comprising a dCas9 molecule fused to DNMT3A-DNMT3L fusion peptide.
[0200] DNMT3A: MNHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQVDRYIASEVCEDSITVGMVRHQGKIMYVGDVRSVTQKHIQEWGPFDLVIGGSPCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPFFWLFENVVAMGVSDKRDISRFLESNPVMIDAKEVSAAHRARYFWGNLPGMNRPLASTVNDKLELQECLEHGRIAKFSKVRTITTRSNSIKQGKDQHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRSWSVPVIRHLFAPLKEYFACV (SEQ ID NO: 23)
[0201] DNMT3L: MGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGTLKYVEDVTNVVRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFHRILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQDVRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVRSRSKLDAPKVDLLVKNCLLPLREYFKYFSQ NSLPL (SEQ ID NO: 24)
[0202] DNMT3A-DNMT3L fusion peptide (SEQ ID NO:27)
[0203] In one embodiment, the Cas9 fusion protein also comprises a nuclear localization sequence (NLS), e.g., an LS fused to the N-terminus and / or C-terminus of Cas9.
[0204] Nuclear localization sequences are known in the art. In one embodiment, the NLS comprises the amino acid sequence of SEQ ID NO:25 or 26, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO:25 or 26, or a sequence having one, two, three, four, five or more alterations, such as amino acid substitutions, insertions, or deletions, relative to SEQ ID NO:25 or 26, or any fragment thereof.
[0205] SEQ ID NO:25 (exemplary nuclear localization sequence): APKKKRKVGIHGVPAA
[0206] SEQ ID NO:26 (an exemplary nuclear localization sequence): KRPAATKKAGQAKKKK
[0207] In some embodiments, the CRISPR / Cas9 system may include a dCas9 molecule and a modulator of gene expression, or a nucleic acid encoding the dCas9 molecule and the modulator of gene expression. In one embodiment, the dCas9 molecule and the modulator of gene expression are covalently linked. In one embodiment, the modulator of gene expression is directly covalently fused to the dCas9 molecule. In one embodiment, the modulator of gene expression is indirectly covalently fused to the dCas9 molecule, either through a non-modulator or a linker, or through a second modulator. In one embodiment, the modulator of gene expression is at the N-terminus and / or C-terminus of the dCas9 molecule. In one embodiment, the dCas9 molecule and the modulator of gene expression are non-covalently linked. Exemplary sequences include, but are not limited to, those listed in Table 2. In some embodiments, the linker between the dCas9 and at least one modulator of gene expression comprises an amino acid sequence corresponding to a linker listed in Table 2.
[0208] [Table 2]
[0209] In one embodiment, the dCas9 molecule is fused to a first tag, e.g., a first peptide tag. In one embodiment, the modulator of gene expression is fused to a second tag, e.g., a second peptide tag. In one embodiment, the first and second tags, e.g., the first peptide tag and the second peptide tag, interact with each other non-covalently, thereby bringing the dCas9 molecule and the modulator of gene expression into close proximity.
[0210] In one embodiment, the CRISPR / Cas9-based system comprises a fusion molecule or a nucleic acid encoding the fusion molecule. In one embodiment, the fusion molecule comprises a sequence comprising dCas9 fused to a modulator of gene expression. In one embodiment, the dCas9 molecule is a Streptococcus pyogenes dCas9 molecule, a Staphylococcus aureus dCas9 molecule, a Campylobacter jejuni dCas9 molecule, a Corynebacterium diphtheriae dCas9 molecule, a Eubacterium ventriosum dCas9 molecule, a Streptococcus pasteurianus dCas9 molecule, a Lactobacillus farciminis dCas9 molecule, a Spirochaete globus dCas9 molecule, a Azospirillum (B510 strain) dCas9 molecule, a Gluconacetobacter diazotrophicus dCas9 molecule, a Neisseria cinerea dCas9 molecule, a Roseburia intestinalis dCas9 molecule, a Parvibaculum labamentivorans dCas9 molecule, a Nitratifructator sarsuginis (DSM) dCas9 molecule, or a combination thereof. The dCas9 molecule includes a Campylobacter lari (strain CF89-12) dCas9 molecule, a Streptococcus thermophilus (strain LMD-9) dCas9 molecule, or a fragment thereof.
[0211] In one embodiment, the fusion molecule is a DNMT3A-DNMT3L(3A3L)-dCas9-KRAB fusion molecule comprising a DNMT3A-DNMT3L fusion peptide (3A3L), a dCas9 peptide, and a KRAB peptide domain fused directly or indirectly (e.g., via a linker) from the N-terminus to the C-terminus.
[0212] In one embodiment, the fusion molecule is a DNMT3A-DNMT3L(3A3L)-dCas9-KRAB fusion molecule comprising a DNMT3A-DNMT3L fusion peptide (3A3L), a dCas9 peptide, and a KRAB peptide domain fused directly or indirectly (e.g., via a linker) from the N-terminus to the C-terminus.
[0213] In one embodiment, the fusion molecule comprises the amino acid sequence of SEQ ID NO:28, a sequence substantially identical to SEQ ID NO:28 (e.g., having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity), or a sequence having one, two, three, four, five or more alterations, e.g., substitutions, insertions, or deletions, relative to SEQ ID NO:28, or any fragment thereof.
[0214] DNMT3A-DNMT3L(3A3L)-dCas9-KRAB [ka]
[0215] gRNA As used herein, the term "guide sequence" in reference to the CRISPR-Cas system includes any polynucleotide sequence that hybridizes with a target nucleic acid sequence and has sufficient complementarity with the target nucleic acid sequence to direct sequence-specific binding of a nucleic acid target complex to the target nucleic acid sequence. The guide sequence may form a duplex with the target sequence. The duplex may be a DNA duplex, an RNA duplex, or an RNA / DNA duplex. The terms "guide molecule" and "guide RNA" and "single guide RNA" are used interchangeably herein to refer to an RNA-based molecule that includes a guide sequence that can form a complex with a CRISPR-Cas protein and has sufficient complementarity with the target nucleic acid sequence to hybridize with the target nucleic acid sequence to direct sequence-specific binding of the complex to the target nucleic acid sequence. Guide molecules or guide RNAs, as described herein, specifically encompass RNA-based molecules with one or more chemical modifications, such as chemical linking of two ribonucleotides or replacement of one or more ribonucleotides with one or more deoxyribonucleotides.
[0216] A guide molecule or guide RNA of a CRISPR-Cas protein may comprise a tracr mate sequence (including a "direct repeat" in the context of an endogenous CRISPR system) and a guide sequence (also referred to as a "spacer" in the context of an endogenous CRISPR system). In some embodiments, the CRISPR-Cas system or complex described herein does not comprise and / or rely on the presence of a tracr sequence. In certain embodiments, a guide molecule may comprise, consist essentially of, or consist of a direct repeat sequence fused or linked to a guide sequence or spacer sequence. In some embodiments, a guide molecule or sgRNA comprises a tracr sequence as set forth in SEQ ID NO:59.
[0217] An exemplary tracr sequence: gtttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc (SEQ ID NO: 59)
[0218] Generally, CRISPR-Cas systems are characterized by elements that promote the formation of CRISPR complexes at the site of target sequences. In the context of the formation of CRISPR complexes, "target sequence" refers to the sequence to which guide sequence is designed to have complementarity, and hybridization of the target DNA sequence with the guide sequence promotes the formation of CRISPR complexes.
[0219] In certain embodiments, the length of the guide sequence or spacer of the guide molecule is 15-50 nucleotides in length. In certain embodiments, the length of the spacer of the guide RNA is at least 15 nucleotides in length. In certain embodiments, the length of the spacer is 15-17 nucleotides in length, 17-20 nucleotides in length, 20-24 nucleotides in length, 23-25 nucleotides in length, 24-27 nucleotides in length, 27-30 nucleotides in length, 30-35 nucleotides in length, or more than 35 nucleotides in length.
[0220] In some embodiments, the guide sequence is of length 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0221] In some embodiments, the sequence of the guide molecule (direct repeat and / or spacer) is selected to reduce the degree of secondary structure within the guide molecule. In some embodiments, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of the guide RNA targeting a nucleic acid contribute to self-complementary base pairing when optimally folded. Optimal folding may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculation of minimum Gibbs free energy. One example of such an algorithm is mFold, described by Zuker and Stiegler [Nucleic Acids Res. 9 (1981), pp. 133-148]. Another example of a folding algorithm is the online web server RNAfold, which uses a centroid structure prediction algorithm developed at the Institute for Theoretical Chemistry at the University of Vienna [see, e.g., A. R. Gruber et al., 2008, Cell 106(1):23-24; and P. A. Carr and G. M. Church, 2009, Nature Biotechnology 27(12):1151-62].
[0222] As mentioned above, the CRISPR / Cas9 system utilizes a gRNA that provides targeting for the CRISPR / Cas9-based system. The gRNA is a fusion of two non-coding RNAs, the crRNA and the tracrRNA. The sgRNA can target any desired DNA sequence by exchanging a sequence encoding a 20 bp protospacer that confers target specificity through complementary base pairing with the desired DNA target. The gRNA mimics the naturally occurring crRNA:tracrRNA duplex involved in type II effector systems. This duplex may include, for example, a 42 nucleotide crRNA and a 75 nucleotide tracrRNA, which acts as a guide for Cas9 to cleave the target nucleic acid.
[0223] The terms "target region," "target sequence," or "protospacer," as used interchangeably herein, refer to a region of a target gene targeted by a CRISPR / Cas9-based system. A CRISPR / Cas9-based system may include at least one gRNA, where the gRNAs target different DNA sequences. The target DNA sequences may overlap. The target sequence or protospacer is followed by a PAM sequence at the 3' end of the protospacer. Different type II systems require different PAMs. For example, the S. pyogenes type II system uses a "NGG" sequence, where "N" can be any nucleotide.
[0224] In some embodiments, the number of gRNAs administered to a cell may be at least 1 gRNA, at least 2 different gRNAs, at least 3 different gRNAs, at least 4 different gRNAs, at least 5 different gRNAs, at least 6 different gRNAs, at least 7 different gRNAs, at least 8 different gRNAs, at least 9 different gRNAs, at least 10 different gRNAs, at least 11 different gRNAs, at least 12 different gRNAs, at least 13 different gRNAs, at least 14 different gRNAs, at least 15 different gRNAs, at least 16 different gRNAs, at least 17 different gRNAs, at least 18 different gRNAs, at least 19 different gRNAs, at least 20 different gRNAs, at least 25 different gRNAs, at least 30 different gRNAs, at least 35 different gRNAs, at least 40 different gRNAs, at least 45 different gRNAs, or at least 50 different gRNAs.
[0225] In some embodiments, the number of gRNAs administered to a cell is at least 1 gRNA to at least 50 different gRNAs, at least 1 gRNA to at least 45 different gRNAs, at least 1 gRNA to at least 40 different gRNAs, at least 1 gRNA to at least 35 different gRNAs, at least 1 gRNA to at least 30 different gRNAs, at least 1 gRNA to at least 25 different gRNAs, at least 1 gRNA to at least 20 different gRNAs, at least 1 gRNA to at least 16 different gRNAs, at least 1 gRNA to at least 12 different gRNAs, at least 1 gRNA to at least 8 different gRNAs, at least 1 gRNA to at least 4 different gRNAs, at least 4 gRNA to at least 50 different gRNAs, at least 4 different gRNAs to at least 45 different gRNAs, at least 4 different gRNAs to at least 40 different gRNAs, at least 4 different gRNAs to at least 35 different gRNAs, at least Also, between 4 different gRNAs to at least 30 different gRNAs, at least 4 different gRNAs to at least 25 different gRNAs, at least 4 different gRNAs to at least 20 different gRNAs, at least 4 different gRNAs to at least 16 different gRNAs, at least 4 different gRNAs to at least 12 different gRNAs, at least 4 different gRNAs to at least 8 different gRNAs, at least 8 different gRNAs to at least 50 different gRNAs, at least 8 different gRNAs to at least 45 different gRNAs, at least 8 different gRNAs to at least 40 different gRNAs, at least 8 different gRNAs to at least 35 different gRNAs, at least 8 different gRNAs to at least 30 different gRNAs, at least 8 different gRNAs to at least 25 different gRNAs, 8 different gRNAs to at least 20 different gRNAs, at least 8 different gRNAs to at least 16 different gRNAs, or 8 different gRNAs to at least 12 different gRNAs.
[0226] In some embodiments, the gRNA is selected to increase or decrease transcription of the target gene. In some embodiments, the gRNA targets a region upstream of the transcription start site (TSS) of the target gene (e.g., VEGF), for example, 0-1000 bp upstream of the transcription start site of the target gene. In some embodiments, the gRNA targets a region 0-50 bp, 0-100 bp, 0-150 bp, 0-200 bp, 0-250 bp, 0-300 bp, 0-350 bp, 0-400 bp, 0-450 bp, 0-500 bp, 0-550 bp, 0-600 bp, 0-650 bp, 0-700 bp, 0-750 bp, 0-800 bp, 0-850 bp, 0-900 bp, 0-950 bp, or 0-1000 bp upstream of the transcription start site of the target gene. In some embodiments, the gRNA targets a region within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the transcription start site of the target gene. In one embodiment, the gRNA targets a region 0-300 bp upstream of the TSS of the target gene.
[0227] In some embodiments, the gRNA targets a region downstream of the transcription start site of the target gene, for example, 0-1000 bp downstream of the transcription start site of the target gene. In some embodiments, the gRNA targets a region 0-50 bp, 0-100 bp, 0-150 bp, 0-200 bp, 0-250 bp, 0-300 bp, 0-350 bp, 0-400 bp, 0-450 bp, 0-500 bp, 0-550 bp, 0-600 bp, 0-650 bp, 0-700 bp, 0-750 bp, 0-800 bp, 0-850 bp, 0-900 bp, 0-950 bp, or 0-1000 bp downstream of the transcription start site of the target gene. In some embodiments, the gRNA targets a region within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp downstream of the transcription start site of the target gene. In one embodiment, the gRNA targets a region 0-300 bp downstream of the TSS of the target gene.
[0228] VEGF As used herein, the term "VEGF" refers to vascular endothelial growth factor. The VEGF pathway is involved in many aspects of blood vessel development and includes a family of proteins that act as angiogenic activators, including VEGF-A, VEGF-B, VEGF-C, VEGF-E, and their respective receptors. VEGF-A, also called VEGF or vascular permeability factor (VPF), is the target of antiangiogenic therapy. VEGF-A exists in five isoforms resulting from alternative splicing of the mRNA of a single VEGF gene, namely VEGFm, VEGF45, VEGFies, VEGF189, and VEGF206.
[0229] Human VEGFA has a cytogenetic location of 6p21.1 with genomic coordinates at positions 43,770,209-43,786,487 on the forward strand of chromosome 6. Examples of human VEGF sequences can be found at NCBI Gene ID 7422 and Ensembl Gene ID ENSG00000112715. VEGFA induces proliferation and migration of vascular endothelial cells and is important for both physiological and pathological angiogenesis. Disruption of this gene in mice resulted in abnormalities in embryonic vascular formation.
[0230] In some embodiments, the method shows a significant reduction in the expression level of VEGFA after transfection with the CRISPR / Cas9 system disclosed herein. In some embodiments, the reduction in VEGF expression level is about 80% or more, about 85% or more, about 90% or more, or about 95% or more compared to a control (e.g., transfection with a sgRNA that does not target VEGFA).
[0231] In some embodiments, the reduction in VEGFA expression levels is maintained for at least about 96 hours, at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 6 weeks, at least about 2 months, or longer following transfection with the CRISPR / Cas9 system disclosed herein. In some embodiments, the reduction in VEGF levels is maintained for about 1 week to about 4 weeks, about 1 week to about 3 weeks, about 1 week to about 2 weeks, about 2 weeks to about 4 weeks, about 2 weeks to about 3 weeks, or about 3 weeks to about 4 weeks.
[0232] In some embodiments, the reduction in the expression level of VEGFA is based on a comparison of a baseline or a predetermined level of VEGFA. In some embodiments, the reduction in the level of VEGFA may be based on a comparison of a first level of VEGF from a first sample from a subject and a second level of VEGF from a second sample from a subject.
[0233] The present disclosure provides sgRNA sequences that target mouse or rabbit VEGFA target genes.Exemplary sgRNAs include, but are not limited to, those listed in Table 3a and Table 3b.The present disclosure also provides sgRNA sequences that target human VEGFA (which also target the homologous region in monkey VEGFA).Exemplary sgRNAs include, but are not limited to, those listed in Table 4.
[0234] [Table 3]
[0235] [Table 4]
[0236] [Table 5]
[0237] In some embodiments, the gRNA targets the promoter region of the target gene. In one embodiment, the gRNA targets the enhancer region of the target gene. The gRNA can be divided into a target binding region, a Cas9 binding region, and a transcription termination region. The target binding region hybridizes with a target region in the target gene. Methods for designing such target binding regions are known in the art. See, for example, Doench et al., Nat Biotechnol. (2014) 32:1262-7; and Doench et al., Nat Biotechnol. (2016) 34:184-91, the entire contents of which are incorporated herein by reference. Design tools are available, for example, Feng Zhang lab's target Finder, Michael Boutros lab's Target Finder (E-CRISP), RGEN Tools (Cas-OF Finder), CasFinder, and CRISPR Optimal Target Finder. In certain embodiments, the target binding region may be between about 15 and 50 nucleotides in length (about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or about 50 nucleotides in length). In certain embodiments, the target binding region may be between about 19 and about 21 nucleotides in length. In one embodiment, the target binding region is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.
[0238] In one embodiment, the target binding region is complementary, for example completely complementary, to the target region in the target gene.In one embodiment, the target binding region is substantially complementary to the target region in the target gene.In one embodiment, the target binding region comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or less nucleotides that are not complementary to the target region in the target gene.
[0239] In one embodiment, the target binding region is engineered to improve stability or extend half-life, for example by incorporating non-natural or modified nucleotides into the target binding region, by removing or modifying RNA destabilizing sequence elements, by adding RNA stabilizing sequence elements, or by increasing the stability of the Cas9 / gRNA complex. In one embodiment, the target binding region is engineered to enhance its transcription. In one embodiment, the target binding region is engineered to reduce secondary structure formation. In one embodiment, the Cas9 binding region of the gRNA is modified to enhance transcription of the gRNA. In one embodiment, the Cas9 binding region of the gRNA is modified to improve the stability or assembly of the Cas9 / gRNA complex.
[0240] Delivery System The present disclosure also provides delivery systems for introducing the components of the systems and compositions herein into a cell, tissue, organ, or organism. The delivery systems may include one or more delivery vehicles and / or cargo.
[0241] cargo The delivery system may include one or more cargoes. The cargoes may include one or more components of the systems and compositions herein. The cargoes may include one or more of the following: i) a plasmid encoding one or more Cas proteins, ii) a plasmid encoding one or more guide RNAs, iii) an mRNA of one or more Cas proteins, iv) one or more guide RNAs, v) one or more Cas proteins, vi) any combination thereof. In some examples, the cargoes may include a plasmid encoding one or more Cas proteins and one or more (e.g., multiple) guide RNAs. In some embodiments, the cargoes may include an mRNA encoding one or more Cas proteins and one or more guide RNAs.
[0242] In some examples, the cargo may include one or more Cas proteins and one or more guide RNAs, for example, in the form of a ribonucleoprotein complex (RNP). The ribonucleoprotein complex may be delivered by the methods and systems herein. In some cases, the ribonucleoprotein may be delivered by a polypeptide-based shuttle agent. In one example, the ribonucleoprotein may be delivered to a histidine-rich domain and a CPD using a synthetic peptide that includes an endosomal leakage domain (ELD) operably linked to a cell-penetrating domain (CPD), for example, as described in WO2016161516.
[0243] Physical Delivery In some embodiments, the cargo may be introduced into the cell by a physical delivery method. Examples of physical methods include microinjection, electroporation, and hydrodynamic delivery.
[0244] Microinjection Direct microinjection of cargo into cells can achieve high efficiencies, for example, greater than 90% or about 100%. In some embodiments, microinjection may be performed using a microscope and a needle (e.g., 0.5-5.0 μm diameter) to puncture the cell membrane and deliver the cargo directly to the target site within the cell. Microinjection may be used for in vitro and ex vivo delivery.
[0245] Plasmids containing sequences encoding Cas protein and / or guide RNA, mRNA, and / or guide RNA may be microinjected. In some cases, microinjection may be used to i) deliver DNA directly to cell nucleus, and / or ii) deliver mRNA (e.g., in vitro transcribed) to cell nucleus or cytoplasm. In certain examples, microinjection may be used to deliver sgRNA directly to nucleus and deliver mRNA encoding Cas to cytoplasm, for example, to facilitate translation and shuttling of Cas to nucleus.
[0246] Microinjection can be used to generate genetically modified offspring animals. For example, gene editing cargo can be injected into zygotes to allow efficient germline modification. Such an approach can obtain normal embryos and full-term offspring with desired modifications. Microinjection can also be used to transiently up-regulate or down-regulate specific genes in the genome of cells, for example, using CRISPR and CRISPRi.
[0247] Electroporation In some embodiments, cargo and / or delivery vehicle can be delivered by electroporation.Electroporation uses high-voltage pulse current to transiently open nanometer-sized pores in the cell membrane of cells suspended in buffer solution, allowing components with hydrodynamic diameters of tens of nanometers to enter the cells.In some cases, electroporation is used in various cell types to efficiently transfer cargo into cells.Electroporation can be used in in vitro and ex vivo delivery.
[0248] Electroporation may be used to deliver cargo to the nucleus of mammalian cells by applying specific voltages and agents, for example by nucleofection. Such approaches include those described in Wu Y, et al. (2015). Cell Res 25:67-79; Ye L, et al. (2014). Proc Natl Acad Sci USA 111:9591-6; Choi PS, Meyerson M. (2014). Nat Commun 5:3728; Wang J, Quake SR. (2014). Proc Natl Acad Sci 111:13157-62. Electroporation may be used to deliver cargo in vivo, for example by the method described in Zuckermann M, et al. (2015). Nat Commun 6:7391.
[0249] Hydraulic Delivery Hydrodynamic delivery may also be used to deliver cargo, e.g., for in vivo delivery. In some examples, hydrodynamic delivery may be performed by rapidly pushing a large volume (e.g., 8-10% of body weight) of solution containing the gene editing cargo into the bloodstream of a subject (e.g., animal or human), e.g., through the tail vein for mice. Because blood is incompressible, a large volume of liquid injected at once results in an increase in hydrodynamic pressure, which temporarily enhances the permeability to endothelial and parenchymal cells, allowing the passage of cargo into cells that would normally not be able to cross the cell membrane. This approach may be used to deliver naked DNA plasmids and proteins. The delivered cargo is enriched in the liver, kidneys, lungs, muscle, and / or heart.
[0250] Transfection Cargo, e.g., nucleic acid, may be introduced into cells by a transfection method for introducing nucleic acid into cells. Examples of transfection methods include calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, imparefection, optical transfection, and proprietary agent-enhanced nucleic acid uptake.
[0251] Delivery Vehicle The delivery system may include one or more delivery vehicles. The delivery vehicle may deliver the cargo to a cell, tissue, organ, or organism (e.g., an animal or plant). The cargo may be packaged, carried, or otherwise associated with the delivery vehicle. The delivery vehicle is selected based on the type of cargo to be delivered and / or the delivery is in vitro and / or in vivo. Examples of delivery vehicles include vectors, viruses, non-viral vehicles, and other delivery agents described herein.
[0252] A delivery vehicle according to the present disclosure may have a maximum dimension (e.g., diameter) of less than 100 microns (μm). In some embodiments, the delivery vehicle has a maximum dimension of less than 10 μm. In some embodiments, the delivery vehicle may have a maximum dimension of less than 2000 nanometers (nm). In some embodiments, the delivery vehicle may have a maximum dimension of less than 1000 nanometers (nm). In some embodiments, the delivery vehicle may have a maximum dimension (e.g., diameter) of less than 900 nm, less than 800 nm, less than 700 nm, less than 600 nm, less than 500 nm, less than 400 nm, less than 300 nm, less than 200 nm, less than 150 nm, or less than 100 nm, less than 50 nm. In some embodiments, the delivery vehicle may have a maximum dimension in the range of 25 nm to 200 nm.
[0253] In some embodiments, the delivery vehicle may be or include particles. For example, the delivery vehicle may be or include nanoparticles (e.g., particles having a maximum dimension (e.g., diameter) of 1000 nm or less). The particles may be provided in various forms, such as solid particles (e.g., metals such as silver, gold, iron, titanium, etc., non-metals, liquid-based solids, polymers), suspensions of particles, or combinations thereof. Metallic, dielectric, and semiconductor particles may be prepared, as well as hybrid structures (e.g., core-shell particles).
[0254] vector The system, composition, and / or delivery system may include one or more vectors. The present disclosure also includes vector systems. The vector system may include one or more vectors. In some embodiments, a vector refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include single-stranded, double-stranded, or partially double-stranded nucleic acid molecules, nucleic acid molecules with one or more free ends or no free ends (e.g., circular), nucleic acid molecules comprising DNA, RNA, or both, and various other polynucleotides known in the art. A vector may be a plasmid, e.g., a circular double-stranded DNA loop into which additional DNA segments can be inserted by standard molecular cloning techniques. Certain vectors may be capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors and episomal mammalian vectors having a bacterial origin of replication). Some vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell, and are thereby replicated along with the host genome. In certain instances, the vector may be an expression vector capable of, e.g., directing the expression of a gene to which it is operably linked. In some cases, the expression vector may be for expression in eukaryotic cells. Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
[0255] Examples of vectors include pGEX, pMAL, pRIT5, E. coli expression vectors (e.g., pTrc, pET l ld), yeast expression vectors (e.g., pYepSecl, pMFa, pJRY88, pYES2, and picZ), baculovirus vectors (e.g., for expression in insect cells such as SF9 cells) (e.g., the pAc series and pVL series), mammalian expression vectors (e.g., pCDM8 and pMT2PC).
[0256] A vector may comprise i) a Cas coding sequence, and / or ii) a single or at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 32, at least 48, at least 50 guide RNA coding sequences. There may be a promoter for each RNA coding sequence in a single vector. Alternatively or additionally, there may be promoters controlling (e.g. driving transcription and / or expression of) multiple RNA coding sequences in a single vector.
[0257] Control Elements The vector may comprise one or more control elements. The control elements may be operably linked to the coding sequence of a Cas protein, an accessory protein, a guide RNA (e.g., a single guide RNA, a crRNA, and / or a tracrRNA), or a combination thereof. The term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the control element in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system, or in a host cell when the vector is introduced into the host cell). In certain examples, the vector may comprise a first control element operably linked to a nucleotide sequence encoding a Cas protein, and a second control element operably linked to a nucleotide sequence encoding a guide RNA.
[0258] Examples of control elements include promoters, enhancers, internal ribosomal entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Such control elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). Control elements include control elements that direct constitutive expression of a nucleotide sequence in many types of host cells, and control elements that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific control sequences). Tissue-specific promoters may direct expression primarily in a desired tissue of interest, such as muscle, nerve, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). Control elements may direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific.
[0259] Exemplary promoters include one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or combinations thereof. Exemplary pol III promoters include, but are not limited to, U6 and HI promoters. Exemplary pol II promoters include, but are not limited to, retroviral Rous sarcoma virus promoter (RSV) LTR promoter (which may include an RSV enhancer), cytomegalovirus (CMV) promoter (which may include a CMV enhancer), SV40 promoter, dihydrofolate reductase promoter, actin promoter, phosphoglycerol kinase (PGK) promoter, and EF1a promoter.
[0260] Viral Vectors Cargo may be delivered by virus. In some embodiments, viral vectors are used. Viral vectors may contain virus-derived DNA or RNA sequences for packaging into viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by viruses for transfection into host cells. Viruses and viral vectors may be used for in vitro, ex vivo, and / or in vivo delivery.
[0261] Adeno-associated virus (AAV) The system and composition herein may be delivered by adeno-associated virus (AAV). AAV vectors may be used for such delivery. AAV of the genus Dependovirus and family Parvoviridae is a single-stranded DNA virus. In some embodiments, the genomic material delivered by AAV can exist indefinitely in cells, for example as exogenous DNA, or can directly integrate into host DNA with some modification, so that AAV can provide a continuous source of donated DNA. In some embodiments, AAV does not cause disease or is not associated with disease in humans. The virus itself can efficiently infect cells with little or no elicitation of innate or adaptive immune responses or associated toxicity.
[0262] Examples of AAV that can be used herein include AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-8, and AAV-9. The type of AAV may be selected with respect to the cells to be targeted, for example, AAV serotypes 1, 2, 5, or hybrid capsids AAV1, AAV2, AAV5, or any combination thereof, can be selected to target brain or neuronal cells, and AAV4 can be selected to target cardiac tissue. AAV8 is useful for delivery to the liver. AAV-2 based vectors were originally proposed for delivery of CFTR to CF airways. Other serotypes such as AAV-1, AAV-5, AAV-6, and AAV-9 exhibit improved gene transfer efficiency in various models of lung epithelial tissue. Examples of cell types targeted by AAV are described in Grimm, D. et al., J. Virol. 82:5887-5911 (2008) and WO 2021 / 183807A1, the entireties of which are incorporated by reference herein.
[0263] CRISPR-Cas AAV particles may be created in HEK 293T cells. Once particles with specific tropism are created, they can be used to infect target cell lines in the same way as natural virus particles. This allows for the sustained presence of CRISPR-Cas components in the infected cell type, making this version of delivery particularly suitable for cases where long-term expression is desired. Examples of AAV dosages and formulations that can be used include those described in U.S. Patent Nos. 8,454,972 and 8,404,658.
[0264] Various strategies are used for delivery by the system and composition herein using AAV.In some cases, the coding sequences of Cas and gRNA are directly packaged in one DNA plasmid vector and delivered via one AAV particle.In some cases, AAV can be used to deliver gRNA to cells that have already been engineered to express Cas.In some cases, the coding sequences of Cas and gRNA are produced in two separate AAV particles, which are used for co-transfection of target cells.In some cases, markers, tags, and other sequences can be packaged in the same AAV particle as the coding sequences of Cas and / or gRNA.
[0265] Lentivirus The system and composition of the present invention can be delivered by lentivirus.Lentivirus vector can be used for such delivery.Lentivirus is a complex retrovirus that has the ability to infect and express genes in both mitotic and postmitotic cells.
[0266] Examples of lentivirus include human immunodeficiency virus (HIV), which can target a wide range of cell types using envelope glycoproteins from other viruses, and the minimal non-primate lentivirus vector of the equine infectious anemia virus (EIAV) system used in ophthalmology treatment.In certain embodiments, self-inactivating lentivirus vectors carrying siRNA targeting the common exon shared by HIV tat / rev, nucleolus-localized TAR decoy, and anti-CCR5 specific hammerhead ribozyme (see, for example, DiGiusto et al., (2010) Sci Transl Med 2:36ra43) can be used and / or adapted to the nucleic acid targeting system herein.
[0267] Lentivirus can be pseudotyped with other viral proteins, such as the G protein of vesicular stomatitis virus.When doing this, the cellular tropism of lentivirus can be changed to be broad or narrow as desired.In some cases, the second and third generation lentivirus system can split essential genes into three plasmids to improve stability, which can reduce the possibility of accidental reconstitution of live viral particles in cells.
[0268] In some instances, taking advantage of their integration capabilities, lentiviruses may be used to create libraries of cells containing different genetic modifications, for example, for screening and / or studying genes and signaling pathways.
[0269] Adenovirus The system and composition herein can be delivered by adenovirus.Adenovirus vector can be used for such delivery.Adenovirus includes non-enveloped virus with icosahedral nucleocapsid containing double-stranded DNA genome.Adenovirus can infect dividing and non-dividing cells.In some embodiments, adenovirus is not integrated into the genome of host cell, and can be used to limit the off-target effect of CRISPR-Cas system in gene editing application.
[0270] Non-viral media The delivery vehicle may include a non-viral vehicle. In general, any method and vehicle capable of delivering nucleic acid and / or protein may be used to deliver the system composition herein. Examples of non-viral vehicles include lipid nanoparticles, cell penetrating peptides (CPPs), DNA nanoclues, gold nanoparticles, streptolysin 0, multifunctional enveloped nanodevices (MENDs), lipid-coated mesoporous silica particles, and other inorganic nanoparticles.
[0271] lipid particles Delivery vehicles may include lipid particles, such as lipid nanoparticles (LNPs) and liposomes.
[0272] Lipid Nanoparticles (LNPs) LNP encapsulates nucleic acid in cationic lipid particles (e.g., liposomes) and can be relatively easily delivered to cells. In some cases, lipid nanoparticles do not contain any viral components, thereby helping to minimize safety and immunogenicity concerns. Lipid particles can be used for in vitro, ex vivo, and in vivo delivery. Lipid particles can be used for cell populations at various scales.
[0273] LNPs can be easily prepared by various methods known in the art, such as mixing an organic phase with an aqueous phase. Mixing of the two phases can be achieved by microfluidic devices and impinging flow reactors. Better mixing of the organic phase with the aqueous phase results in better embedding rate and particle size distribution of LNPs. Preferably, the particle size of LNPs can be adjusted by changing the mixing rate of the organic phase with the aqueous phase. Faster mixing rate prepares LNPs with smaller particle size. Embedding efficiency can be optimized by controlling the N / P ratio of the LNP system. In some preferred embodiments, the N / P ratio is 1:1 to 9:1.
[0274] In some examples, LNPs may be used to deliver DNA molecules (e.g., DNA molecules containing coding sequences for Cas and / or gRNA) and / or RNA molecules (e.g., Cas mRNA, gRNA). In certain cases, LNPs may be used to deliver RNP complexes of Cas / gRNA.
[0275] In some embodiments, LNPs are used to deliver mRNA and gRNA (e.g., an mRNA fusion molecule comprising DNMT3A-DNMT3L (3A-3L)-dCas9-KRAB and at least one sgRNA targeting VEGF).
[0276] Components of LNPs may include cationic lipids, i.e., 1,2-dilineoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-. In some embodiments, LNPs may include ionizable lipids. In some embodiments, ionizable lipids include, but are not limited to, pH-responsive ionizable lipids, thermally responsive ionizable lipids, and light-responsive ions. In some embodiments, ionizable lipids include cationic lipids and anionic lipids that are ionized under certain conditions, such as, but not limited to, pH, temperature, or light. In some embodiments, the molar ratio of ionizable lipids in the LNP is between 20% and 70% (about 20% and 70%, about 20% and 65%, about 20% and 60%, about 20% and 55%, about 20% and 50%, about 20% and 45%, about 20% and 40%, about 20% and 35%, about 20% and 30%, about 20% and 25%, about 30% and 70%, about 30% and 65%, about 30% and 60 ... % to 55%, about 30% to 50%, about 30% to 45%, about 30% to 40%, about 30% to 35%, about 40% to 70%, about 40% to 65%, about 40% to 60%, about 40% to 55%, about 40% to 50%, about 40% to 45%, about 50% to 70%, about 50% to 65%, about 50% to 60%, about 50% to 55%, about 60% to 70%, or about 60% to 65%).
[0277] In some embodiments, the LNP may include a PEGylated lipid. In some embodiments, the molar ratio of the PEGylated lipid in the LNP is 0% to about 30% (e.g., about 0% to about 30%, about 0% to about 25%, about 0% to about 20%, about 0% to about 15%, about 0% to about 10%, about 10% to about 30%, about 10% to about 25%, about 10% to about 20%, about 10% to about 15%, about 20% to about 30%, or about 20% to about 25%).
[0278] In some embodiments, the LNP may include a supporting lipid. In some embodiments, the molar ratio of the supporting lipid in the LNP is 30% to about 50% (e.g., about 30% to about 50%, about 30% to about 45%, about 30% to about 40%, about 30% to about 35%, about 40% to about 50%, or about 40% to about 45%).
[0279] In some embodiments, the LNP may contain cholesterol. In some embodiments, the molar ratio of cholesterol in the LNP is 10% to about 50% (e.g., about 10% to about 50%, about 10% to about 45%, about 10% to about 40%, about 10% to about 35%, about 10% to about 30%, about 10% to about 25%, about 10% to about 20%, about 10% to about 15%, about 20% to about 50%, about 20% to about 45%, about 20% to about 40%, about 20% to about 35%, about 20% to about 30%, about 20% to about 25%, about 30% to about 50%, about 30% to about 45%, about 30% to about 40%, about 30% to about 35%, about 40% to about 50%, or about 40% to about 45%).
[0280] In some embodiments, the LNPs may comprise a mixture of ionizable lipids (20%-70% molar ratio), PEGylated lipids (0%-30% molar ratio), supporting lipids (30%-50% molar ratio), and cholesterol (10%-50% molar ratio).
[0281] Liposomes In some embodiments, the lipid particle may be a liposome. Liposomes are spherical vesicular structures consisting of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. In some embodiments, liposomes are biocompatible, non-toxic, and can deliver both hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their load through biological membranes and the blood-brain barrier (BBB).
[0282] Liposomes can be made from several different types of lipids, such as phospholipids. Liposomes may include natural phospholipids and lipids, such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, monosialoganglioside, or any combination thereof.
[0283] Several other additives may be added to the liposomes to modify their structure and properties. For example, the liposomes may further contain cholesterol, sphingomyelin, and / or 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), for example, to increase stability and / or prevent leakage of the liposome's internal cargo.
[0284] Stable Nucleic Acid-Lipid Particles (SNALPs) In some embodiments, the lipid particle may be a stable nucleic acid lipid particle (SNALP). The SNALP may comprise ionizable lipids (DLinDMA) (e.g., cationic at low pH), neutral helper lipids, cholesterol, diffusible polyethylene glycol (PEG)-lipids, or any combination thereof. In some examples, the SNALP may comprise synthetic cholesterol, dipalmitoyl phosphatidylcholine, 3-N- or other lipids.
[0285] The lipid particles may also contain one or more other types of lipids, for example cationic lipids, such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-lipoplexes and / or polyplexes.
[0286] In some embodiments, the delivery vehicle comprises lipoplexes and / or polyplexes. Lipoplexes bind to negatively charged cell membranes and induce endocytosis into cells. Examples of lipoplexes can be complexes containing lipids and non-lipid components. Examples of lipoplexes and polyplexes include FuGENE-6 reagent, non-liposomal solutions containing lipids and other components, zwitterionic amino lipids (ZAL), Ca2p (e.g., forms DNA / Ca2+ microcomplexes), polyetheneimine (PEI) (e.g., branched PEI), and poly(L-lysine) (PLL).
[0287] Cell-penetrating peptides In some embodiments, the delivery vehicle comprises a cell-penetrating peptide (CPP), which is a short peptide that facilitates the cellular uptake of a variety of molecular cargoes, from nano-sized particles to small chemical molecules and large fragments of DNA.
[0288] CPPs can be of various sizes, amino acid sequences, and charges.In some cases, CPPs can translocate plasma membranes to facilitate the delivery of various molecular cargoes to cytoplasm or organelles.CPPs can be introduced into cells through various mechanisms, such as direct penetration in membranes, entry mediated by endocytosis, and translocation by forming transient structures.
[0289] CPPs may have an amino acid composition that contains a high relative abundance of positively charged amino acids such as lysine or arginine, or sequences that contain alternating patterns of polar / charged and non-polar, hydrophobic amino acids. These two types of structures are referred to as polycationic or amphipathic, respectively. A third class of CPPs are hydrophobic peptides that have low net charge and contain only non-polar residues, or contain hydrophobic amino acid groups that are important for cellular uptake. Another type of CPP is the transactivating transcription activator (Tat) from human immunodeficiency virus I (HIV-I). Examples of CPPs include penetratin, Tat(48-60), transportan, and (R-AhX-R4) (Ahx refers to aminohexanoyl). Examples of CPPs and related applications also include those described in U.S. Pat. No. 8,372,951.
[0290] CPPs can be used very easily for in vitro and ex vivo work, and extensive optimization for each cargo and cell type is usually required. In some cases, CPPs can be directly covalently linked to Cas proteins, which are then complexed with gRNA and delivered to cells. In some cases, separate delivery of CPP-Cas and CPP-gRNA to multiple cells can be performed. CPPs can also be used to deliver RNPs.
[0291] DNA Nano Crew In some embodiments, the delivery vehicle comprises a DNA nanocleu. A DNA nanocleu refers to a DNA in a spherical structure (e.g., having the shape of a yam ball). The nanocleu may be synthesized by rolling circle amplification with a palindromic sequence that aids in the self-assembly of the structure. The sphere may then be loaded with a payload. Examples of DNA nanocleu are described in Sun W et al., J Am Chem Soc. 2014 Oct 22; 136(42): 14722-5; and Sun Wet et al., Angew Chem Int Ed Engl. 2015 Oct 5; 54(41): 12029-33. The DNA nanocleu may have a palindromic sequence that should be partially complementary to the gRNA in the Cas:gRNA ribonucleoprotein complex. The DNA nanocleu may be coated, for example, with PEI, to induce endosomal escape.
[0292] Gold Nanoparticles In some embodiments, the delivery vehicle comprises gold nanoparticles (also referred to as AuNPs or colloidal gold). The gold nanoparticles may be complexed with a cargo, such as Cas:gRNA RNP. The gold nanoparticles may be coated, such as with silicate and endosome-disrupting polymer PAsp(DET). Examples of gold nanoparticles include AuraSense Therapeutics' Spherical Nucleic Acid [SNA™] constructs, and those described in Mout R, et al. (2017). ACS Nano 11:2452-8; Lee K, et al. (2017). Nat Biomed Eng 1:889-901.
[0293] iTOP In some embodiments, the delivery vehicle comprises iTOP. iTOP refers to a combination of small molecules that drive highly efficient delivery of native proteins into cells, independent of any transduction peptide. iTOP is used to induce osmocytosis and transduction with propane betaine, and induces macropinocytotic uptake of extracellular macromolecules into cells by the combination of NaCl-mediated hyperosmolarity and the transduction compound (propane betaine). Examples of iTOP methods and reagents include those described in D'Astolfo DS, Pagliero RJ, Pras A, et al. (2015) Cell 161:674-690.
[0294] Polymer-Based Particles In some embodiments, the delivery vehicle may comprise a polymer-based particle (e.g., nanoparticle). In some embodiments, the polymer-based particle may mimic the viral mechanism of membrane fusion. The polymer-based particle may be a synthetic copy of the influenza virus mechanism, forming a transfection complex with various types of nucleic acids (siRNA, miRNA, plasmid DNA, or shRNA, mRNA) that are taken up into cells via the endocytic pathway, a process that involves the formation of acidic compartments. The low pH of the late endosome acts as a chemical switch that makes the particle surface hydrophobic and facilitates crossing the membrane. Once in the cytosol, the particle releases its payload for cellular action. This active endosomal escape technique uses a natural uptake pathway, so it is safe and maximizes transfection efficiency. In some embodiments, the polymer-based particle may comprise alkylated and carboxyalkylated branched polyethyleneimines. In some examples, the polymer-based particle is VIROMER, such as VIROMER RNAi, VIROMER RED, VIROMER mRNA, VIROMER CRISPR. Examples of methods of delivering the systems and compositions herein include those described in Bawage SS et al., Synthetic mRNA expressed Casl3a mitigates RNA viral infections, www.biorxiv.org / content / l0.l l01 / 370460v1.full doi:doi.org / 10.1101 / 370460; Viromer® RED, a powerful tool for transfection of keratinocytes. doi:10.13140 / RG.2.2.16993.61281; Viromer® Transfection - Factbook 2018: technology, product overview, users' data., doi:10.13140 / RG.2.2.23912.16642.
[0295] Streptolysin O (SLO) The delivery vehicle may be streptolysin O (SLO). SLO is a toxin produced by group A streptococci that acts by creating pores in mammalian cell membranes. SLO acts in a reversible manner, allowing the delivery of proteins (e.g., up to 100 kDa) into the cytosol of cells without sacrificing overall viability. Examples of SLO include those described in Sierig G et al., (2003) Infect Immun 71:446-55; Walev I et al., (2001) Proc Natl Acad Sci USA 98:3185-90; Teng KW et al., (2017) Elife 6:e25460.
[0296] Multifunctional Envelope Nanodevices (MEND) The delivery vehicle may include a multifunctional enveloped nanodevice (MEND). The MEND may include a condensed plasmid DNA, a PLL core, and a lipid film shell. The MEND may further include a cell-penetrating peptide (e.g., stearyl octaarginine). The cell-penetrating peptide may be present in the lipid shell. The lipid envelope may be modified with one or more functional components, such as polyethylene glycol (e.g., to increase vascular circulation time), a ligand for targeting specific tissues / cells, an additional cell-penetrating peptide (e.g., for greater cellular delivery), lipids for enhancing endosomal escape, and a nuclear delivery tag. In some examples, the MEND may be a tetralamellar MEND (T-MEND), which may target the cell nucleus and mitochondria. In certain examples, the MEND may be a PEG-peptide-DOPE conjugate MEND (PPD-MEND), which may target bladder cancer cells. Examples of MENDs include those described in Kogure K et al. (2004). J Control Release 98:317-23; Nakamura T et al. (2012). Ace Chem Res 45:1113-21.
[0297] Lipid-coated mesoporous silica particles The delivery vehicle may include lipid-coated mesoporous silica particles. The lipid-coated mesoporous silica particles may include a mesoporous silica nanoparticle core and a lipid membrane shell. The silica core has a large internal surface area, which may result in high cargo loading capacity. In some embodiments, the pore size, pore chemistry, and overall particle size may be modified to load various types of cargo. The lipid coating of the particles may be modified to maximize cargo loading, increase circulation time, and provide precise targeting and cargo release. Examples of lipid-coated mesoporous silica particles include those described in Du X et al. (2014). Biomaterials 35:5580-90; Durfee PN et al. (2016). ACS Nano 10:8325-45.
[0298] Inorganic Nanoparticles The delivery vehicle may comprise inorganic nanoparticles. Examples of inorganic nanoparticles include carbon nanotubes (CNTs) [e.g., as described in Bates Kand Kostarelos K. (2013). Adv Drug Deliv Rev 65:2023-33], bare mesoporous silica nanoparticles (MSNPs) [e.g., as described in Luo GF et al. (2014). Sci Rep 4:6064], and dense silica nanoparticles (SiNPs) [e.g., as described in Luo D and Saltzman WM. (2000). Nat Biotechnol 18:893-5].
[0299] How to use The compositions and systems herein may be used for a variety of applications, including the modification of non-animal organisms, such as plants and fungi, and the modification of animals, the treatment and diagnosis of diseases in plants, animals, and humans. Generally, the compositions and systems may be introduced into a cell, tissue, organ, or organism where the expression and / or activity of one or more genes is modified.
[0300] Cells and organisms The present disclosure provides cells, tissues, and organisms comprising engineered Cas proteins, CRISPR-Cas systems, polynucleotides encoding one or more components of the CRISPR-Cas system, and / or vectors comprising the polynucleotides. The present disclosure also provides nucleotide sequences encoding effector proteins that are codon-optimized for expression in eukaryotic organisms or eukaryotic cells in any of the methods or compositions described herein. In one embodiment of the present disclosure, the codon-optimized effector protein is any Cas protein discussed herein, codon-optimized for operability in eukaryotic cells or organisms, such as cells or organisms described elsewhere herein, such as, but not limited to, yeast cells or mammalian cells or organisms, including mouse cells, rat cells, and human cells, or non-human eukaryotic organisms, such as plants.
[0301] In certain embodiments, modification of a target locus of interest may result in a eukaryotic cell that comprises altered expression of at least one gene product, a eukaryotic cell that comprises altered expression of at least one gene product that has increased expression of at least one gene product, a eukaryotic cell that comprises altered expression of at least one gene product that has decreased expression of at least one gene product, or a eukaryotic cell that comprises an edited genome.
[0302] In certain embodiments, the eukaryotic cell may be a mammalian cell or a human cell.
[0303] In further embodiments, the non-naturally occurring or engineered compositions, vector systems, or delivery systems described herein may be used for site-specific gene knockout, site-specific genome editing, RNA sequence-specific interference, or multiplex genome engineering.
[0304] Also provided are gene products from the cells, cell lines, or organisms described herein. In certain embodiments, the amount of expressed gene product may be greater or less than the amount of gene product from a cell that does not have altered expression or an edited genome. In certain embodiments, the gene product may be altered compared to the gene product from a cell that does not have altered expression or an edited genome.
[0305] Exemplary Therapies The present disclosure provides the use of CRISPR-Cas system for treatment in various diseases and disorders. In some embodiments, the present disclosure described herein relates to a method for treatment, in which cells are edited ex vivo by CRISPR or base editor to regulate VEGF (e.g., VEGFA) gene, and then the edited cells are administered to a patient in need thereof. In some embodiments, the editing comprises knocking in, knocking out, or knocking down the expression of VEGF (e.g., VEGFA) gene in cells.
[0306] In some embodiments, the VEGFA-targeted CRISPR-Cas systems described herein are useful for inhibiting cellular processes mediated by VEGFA and have indications for the prevention or treatment of disorders associated with aberrant angiogenesis and / or lymphangiogenesis stimulated by the action of VEGFA or a VEGFA-associated receptor (e.g., various ophthalmic disorders and cancers).
[0307] The VEGFA-targeting CRISPR-Cas systems described herein, comprising fusion molecules comprising at least one DNA binding protein and at least one modulator of gene expression, or nucleic acid sequences encoding the fusion molecules, and sgRNAs designed to target DNA sequences near the VEGF gene and / or within VEGF regulatory elements, are therapeutically useful for treating or preventing any disease whose condition is improved, ameliorated, inhibited, or prevented by the removal, inhibition, or reduction of VEGF-A.
[0308] A non-exhaustive list of specific conditions ameliorated by inhibition or reduction of VEGFA includes clinical conditions characterized by excessive proliferation of vascular endothelial cells, such as cerebral edema associated with trauma, stroke, or tumors, vascular permeability, edema, edema associated with inflammatory disorders such as psoriasis or arthritis, including rheumatoid arthritis, asthma, generalized edema associated with burns, ascites and pleural effusions associated with tumors, inflammation, or trauma, chronic airway inflammation, capillary leak syndrome, sepsis, kidney disease associated with increased protein leakage, and ophthalmic disorders such as age-related macular degeneration and diabetic retinopathy.
[0309] "Neovascular disorders" are disorders or disease states characterized by altered, dysregulated, or unregulated angiogenesis. Examples of neovascular disorders include neovascular transformation (e.g., cancer) and ophthalmic neovascular disorders, including diabetic retinopathy and age-related macular degeneration.
[0310] An "ophthalmic neovascular disorder" is a disorder characterized by altered, dysregulated, or unregulated angiogenesis in the eye of a patient. Such disorders include optic disc neovascularization, iris neovascularization, retinal neovascularization, choroidal neovascularization, corneal neovascularization, vitreous neovascularization, glaucoma, pannus, pterygium, macular edema, diabetic retinopathy, diabetic macular edema, vascular retinopathy, retinal degeneration, uveitis, retinal inflammatory disease, and proliferative vitreoretinopathy.
[0311] In some embodiments, the disease to be treated by the compositions and methods disclosed herein involves angiogenesis (the formation of blood vessels). In some embodiments, the disease is a neovascular disorder, such as an ophthalmic neovascular disorder, including age-related macular degeneration (AMD), including dry AMD and wet AMD. EXAMPLES
[0312] [Example 1] Fusion molecule plasmid construction and knockdown efficiency
[0313] Two plasmids were constructed to form the "EPICAS" system (Figure 1A). The "fusion molecule" or "catalytic protein" plasmid encodes dCas9, DNMT3A, DNMT3L, and KRAB peptides. The fused DNMT3A and DNMT3L (3A3L) peptides are present at the N-terminus of dCas9, and KRAB is present at the C-terminus of dCas9. That is, the fusion molecule has 3A3L-dCas9-KRAB from the N-terminus to the C-terminus. The "sgRNA" plasmid encodes an sgRNA sequence that targets the VEGFa gene. The "scaffold" is the sequence of the gene or promoter of the VEGFA gene. A number of sgRNAs were designed to target regions within 250 bp upstream and downstream of the transcription start site (TSS) of the mouse VEGFa gene. Specifically, seven sgRNAs (SEQ ID NOs: 29-35) were designed and the corresponding sgRNA plasmids were generated for subsequent transfection.
[0314] Individual sgRNA plasmids were co-transfected with catalytic protein plasmids into mouse N2A cell line (National collection of Authenticated Cell cultures). After 48 hours, the top 10% GFP+ and mCherry+ cells were sorted by FACS. RT-QPCR experiments were performed to evaluate the mRNA expression level of VEGFa. All seven tested sgRNAs showed significantly downregulated VEGF expression in N2A cells, with knockdown efficiency of approximately 80%. Among them, cells transfected with sgRNA3, sgRNA5, and sgRNA7 showed the best knockdown effect, reaching approximately 84% (Figure 1B).
[0315] Next, sgRNA3, sgRNA4, and sgRNA5 were selected to further test the retention of the knockdown effect over a long period of time. The sgRNA plasmids were co-transfected with the catalytic protein plasmid into mouse N2A cell line. After 48 hours, the top 10% GFP+ and mCherry+ cells were sorted by FACS and continued to be cultured. After one week, the cells were collected and RT-QPCR experiments were performed to evaluate the mRNA expression level of VEGFa. The results showed that the gene silencing effect was retained and even improved compared to 48 hours after transfection, reaching more than 90% (Figure 1C). VEGFa mRNA has a relatively long half-life in cells, and it is speculated that both the degradation of existing VEGFa mRNA and the impediment of the synthesis of new VEGFa mRNA contributed to the low level of VEGFa mRNA after one week.
[0316] In addition, to determine whether a combination of two or more sgRNAs can further reduce the gene expression level of VEGFa in N2A cells, a combination of sgRNA3, sgRNA4, and sgRNA5 (sgMix) was tested. The results showed that sgMix significantly knocked down the expression level of VEGFa (Figure 1C). [Example 2] In vitro transcription of mRNA encoding the fusion molecule
[0317] In vitro transcription and purification were used to produce mRNAs corresponding to the fusion molecules or catalytic proteins of the EPICAS system. First, a plasmid was constructed containing all of the fusion molecule elements, including the cassette of 5'UTR-DNMT3A-DNMT3L-dCas9-KRAB-3'UTR-polyA. The plasmid sequence was linearized by XbaI and BpiI restriction enzyme digestion (Figure 2A). An in vitro transcription reaction containing the linearized DNA template, T7 RNA polymerase, NTPs, and a cap analog was performed to produce mRNAs containing N1-methylpseudouridine. After digestion of the DNA template by DNase I, the mRNA product was purified and buffer exchanged, and the purity of the final mRNA product was assessed by capillary gel electrophoresis (Figure 2B). 100-mer sgRNAs were chemically synthesized with minimal terminal modifications under solid-phase synthesis conditions by a commercial supplier (Genewiz).
[0318] To test the function of the in vitro transcribed mRNA, we constructed a Snrpn-GFP reporter system in HEK293T cells. The reporter system uses a synthetic methylation-sensing promoter (a conserved sequence element derived from the promoter of the imprinted gene Snrpn) to control the expression of GFP. By inserting this reporter construct into a genomic locus, the methylation status of nearby sequences was indicated. The above in vitro transcribed mRNA (GFP-P2A-Casoff mRNA) was co-transfected with sgRNA targeting the Snrpn gene into primary mouse hepatocytes using Lipofectamine Messenger MAX (Figure 2C, left panel). FACS analysis showed that the percentage of GFP+ cells was significantly increased 72 hours after transfection with mRNA and sgRNA, suggesting that the reporter system was successfully established in mouse hepatocytes (Figure 2C, center panel). Moreover, the expression level of VEGFA mRNA was obviously reduced 72 hours after transfection (Figure 2C, right panel).
[0319] Taken together, these results indicate that EPICAS mRNA can silence VEGFA expression for a long period of time using transient transfection. [Example 3] Lipid nanoparticle encapsulation of mRNA and sgRNA encoding fusion molecules
[0320] Using standard methods known in the art, LNPs were formulated for delivery of fusion molecule mRNA and sgRNA to mouse choroid. The LNPs contained fusion molecule mRNA and sgRNA targeting VEGFA gene in a 1:1 weight ratio (Figure 3A). A mixture of ionizable lipids (49.5% molar ratio), PEGylated lipids (2.5% molar ratio), supporting lipid (DPPC) (9.9% molar ratio), and cholesterol (38.1% molar ratio), i.e., lipid nanoparticles (LNPs), was formulated using a properly designed impinging flow reactor or microfluidic device. By changing the percentage of ionizable lipid, the release rate of sgRNA and mRNA can be modified. By increasing the percentage of ionizable lipid (>55% molar ratio), sgRNA was released faster than mRNA. Transmission electron microscope (TEM) images showed that the LNPs were spherical and nano-sized particles (Figure 3B). The LNPs had a uniform size (78.2 ± 5.2 nm, PDI < 0.10) using dynamic light scattering (NanoSZ, Malvern) (Figure 3C).
[0321] To test whether LNPs can successfully deliver mRNA to the posterior retina and choroidal regions in vivo, LNPs containing luciferase mRNA were produced and administered by intravitreal injection into the eyes of Ai9 mice (Jax Labs). AAV8 virus expressing EF1A-Cre-GFP was used as a positive control, and PBS was used as a negative control. Fluorescence in vivo imaging results showed that LNPs could deliver mRNA to the retinal pigment epithelium (RPE) and choroid in a highly efficient manner (Figure 3D). The positive control AAV8 could efficiently infect the RPE and choroid. [Example 4] Silencing of the VEGF gene in mice using LNP delivery of mRNA and sgRNA encoding a fusion molecule
[0322] To test whether the mRNA and sgRNA delivered by LNPs could successfully knock down the levels of VEGFa mRNA, LNPs containing EPICAS mRNA and sgRNA3 were produced and administered by intravitreal injection into the eyes of Ai9 mice. Five days after injection, the mice were euthanized, and the retinas and choroids were obtained and processed for purification of mRNA. RT-QPCR experiments were performed to evaluate the mRNA expression levels of VEGFa.
[0323] The results showed significantly downregulated expression of VEGFa in the choroid compared to the control PBS group (Figure 3E), indicating the effectiveness of the EPICAS system in silencing VEGF gene expression in vivo. On the other hand, the reduction in VEGFa expression in the retina was not significant compared to the control group, indicating a potential preference for the choroid in reducing gene expression. [Example 5] Silencing of the rabbit VEGF gene in rabbit cells
[0324] A reporter system was established to screen rabbit VEGFa sgRNAs (SEQ ID NOs: 60-84). An artificial sequence was constructed that included 500 bp upstream of the transcription start site (TSS) of the rabbit VEGFa gene to the first exon and GFP at the C-terminus of the first exon (Figure 4A). The artificial sequence was integrated into the genome of 293T cells using the piggyback transposon system. A cell line stably transfected with GFP was obtained and used to screen sgRNAs. In this experiment, reporter cells were transfected with EPICAS plasmid and sgRNA, and the fluorescence intensity of reporter cells was detected after 72 hours. The results of flow cytometry analysis showed that most of the sgRNAs significantly reduced the intensity of GFP (Figure 4B). Four of the six sgRNAs that had good knockdown effects were transfected into rabbit RK-13 cells together with EPICAS plasmid to target the endogenous gene VEGFA in rabbit cells. QPCR results showed that these sgRNAs significantly reduced VEGFA mRNA expression in RK-13 cells (Figure 4C). [Example 6] Silencing of the human VEGF gene in human cell lines
[0325] To test the efficacy of VEGF gene silencing in human cell lines, a reporter cell line was constructed. A plasmid was constructed to carry a cassette driven by the CMV promoter, where the cassette had the following elements in the 5' to 3' direction: 5'-pCMV-300bp-TSS-+300bp-VEGF exon 1-2A-GFP-3'. In this reporter system, the CMV promoter drives the expression of VEGF and GFP fluorescence. If VEGF is silenced, the transcription of GFP is terminated. The reporter plasmid was transfected into HEK293T cells together with the piggyBac transposase (PBase) plasmid. Cells with successful integration of the reporter cassette were FACS sorted by expression of GFP fluorescence.
[0326] A number of sgRNAs were designed to target homologous regions within 300 bp upstream and downstream of the transcription start site (TSS) of monkey and human VEGF genes (Figure 5A). Twenty-three sgRNAs (sequence numbers 36-58) were selected for the construction of plasmids encoding each one of the sgRNAs. Individual sgRNA plasmids were co-transfected with catalytic protein (DNMT3A-DNMT3L-dCas9-KRAB) plasmids into HEK293T cells. After 48 or 96 h, GFP+ and mCherry+ double positive cells were sorted by FACS. RT-QPCR experiments were performed to evaluate the mRNA expression levels of VEGFa.
[0327] Most sgRNAs tested showed significantly downregulated VEGFa expression in 293T cells (Figure 5B). Cells transfected with sgRNA10, sgRNA19, sgRNA20, sgRNA21, sgRNA22, and sgRNA23 resulted in more than 50% downregulation of VEGFa after 48 hours. After 96 hours, VEGFA expression levels were lower, with sgRNA19, sgRNA20, and sgRNA22 reaching more than 80% downregulation.
[0328] Taken together, these results show that the EPICAS system successfully silenced VEGFA expression with high efficiency and persistence in both mouse and human cells, supporting the silencing of VEGF gene expression by exogenous editing. LNPs were used to successfully deliver the EPICAS system in vivo. Thus, LNP formulations of the EPICAS system can be used to treat VEGF-related diseases, such as AMD.
Claims
1. A composition comprising a fusion molecule comprising at least one DNA binding protein and at least one modulator of gene expression, or a nucleic acid sequence encoding said fusion molecule, wherein said fusion molecule targets a genomic region near a VEGF gene and / or within a VEGF regulatory element, and wherein said at least one modulator of gene expression provides an alteration of at least one nucleotide near the VEGF gene and / or within a VEGF regulatory element; the at least one modulator of gene expression comprises a DNA methyltransferase (DNMT), a DNA demethylase, a histone methyltransferase, a histone demethylase or a portion thereof, or a zinc finger protein-based transcription factor or a portion thereof, or a combination thereof; The composition, wherein the at least one DNA binding protein is Cas9, dCas9, Cpf1, a zinc finger nuclease (ZNF), a transcription activator-like effector nuclease (TALEN), a homing endonuclease, a dCas9-FokI nuclease, or a MegaTal nuclease.
2. The composition of claim 1, wherein the VEGF gene is a VEGF-A gene.
3. The composition of claim 1 , wherein the VEGF regulatory element is a transcription initiation site, a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, or a locus control region.
4. The modification of at least one nucleotide near the VEGF gene and / or within the VEGF regulatory element is (a) located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the transcription start site of the VEGF gene; (b) located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp downstream of the transcription start site of the VEGF gene; (c) located within 500 bp upstream to 500 bp downstream of the transcription start site of the VEGF gene; optionally, located within 300 bp upstream to 300 bp downstream of the transcription start site of the VEGF gene; or (d) located within 1000 bp upstream of the transcription start site of the VEGF gene and within 300 bp downstream of the transcription start site; The composition of claim 1.
5. The composition of claim 1 , wherein the at least one nucleotide modification is DNA methylation.
6. 2. The composition of claim 1, wherein the at least one modulator of gene expression comprises one or more selected from DNA methyltransferase (DNMT), zinc finger protein-based transcription factor, a portion thereof, and any combination thereof, optionally comprising DNA methyltransferase or a portion thereof, and zinc finger protein-based transcription factor or a portion thereof.
7. The composition of claim 6, wherein the DNA methyltransferase is DNMT3A, DNMT3B, DNMT3L, DNMT1, or DNMT2, and optionally, the DNMT3A comprises the amino acid sequence of SEQ ID NO: 23 and / or the DNMT3L comprises the amino acid sequence of SEQ ID NO:
24.
8. The composition of claim 6, wherein the zinc finger protein transcription factor is a Kruppel-associated repression box (KRAB), and optionally, the KRAB comprises the amino acid sequence of SEQ ID NO:
22.
9. The composition of claim 6, wherein the DNA methyltransferase is selected from DNMT3A and DNMT3L and combinations thereof, and the zinc finger protein transcription factor is KRAB.
10. 2. The composition of claim 1, wherein the at least one DNA binding protein is Cas9, dCas9, Cpf1, a zinc finger nuclease (ZNF), a transcription activator-like effector nuclease (TALEN), a homing endonuclease, a dCas9-FokI nuclease, or a MegaTal nuclease.
11. The at least one DNA-binding protein is dCas9, and optionally, the dCas9 is selected from the group consisting of Staphylococcus aureus dCas9, Streptococcus pyogenes dCas9, Campylobacter jejuni dCas9, Corynebacterium diphtheriae dCas9, Eubacterium ventriosum dCas9, and Streptococcus pasteurianus dCas9. , Lactobacillus farciminis dCas9, Spirochaete globus dCas9, Azospirillum (e.g., strain B510) dCas9, Gluconacetobacter diazotrophicus dCas9, Neisseria cinerea dCas9, Roseburia intestinalis dCas9, Parvibaculum lavamentivorans dCas9, Nitratifractor sarusuginis (e.g., strain DSM 16511) dCas9, Campylobacter lari (e.g., strain CF89-12) dCas9, or Streptococcus thermophilus (e.g., strain LMD-9) dCas9, optionally wherein the dCas9 comprises the amino acid sequence of SEQ ID NO:
1.
12. 2. The composition of claim 1, wherein the fusion molecule comprises the at least one modulator of gene expression fused to the C-terminus, N-terminus, or both, of the at least one DNA binding protein, and optionally (a) the at least one modulator of gene expression is fused directly to the at least one DNA binding protein, or (b) the at least one modulator of gene expression is fused indirectly to the at least one DNA binding protein via a non-modulator, a second modulator, or a linker.
13. The composition of claim 12, wherein the fusion molecule comprises dCas9 fused to KRAB at its C-terminus and DNMT3A and DNMT3L at its N-terminus, and optionally, the fusion molecule comprises the amino acid sequence of SEQ ID NO:
28.
14. The composition of claim 1, wherein the fusion molecule further comprises at least one nuclear localization sequence, and optionally, the at least one nuclear localization sequence is fused directly or indirectly to the C-terminus, N-terminus, or both, of the at least one DNA binding protein.
15. The composition of claim 1 , wherein the nucleic acid sequence encoding the fusion molecule is deoxyribonucleic acid (DNA) or messenger ribonucleic acid (mRNA).
16. 2. The composition of claim 1, further comprising at least one single guide RNA (sgRNA) complementary to a target DNA sequence near the VEGF gene and / or within a VEGF regulatory element, optionally wherein the target DNA sequence is located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream or downstream of the transcription start site of the VEGF (e.g., VEGF-A) gene.
17. 17. The composition of claim 16, wherein the sgRNA comprises a nucleic acid sequence of any one of SEQ ID NOs: 29-58 and 60-84.
18. the fusion molecule is packaged in a liposome or lipid nanoparticle; Optionally, the fusion molecule and sgRNA are packaged in a liposome or lipid nanoparticle; Optionally, the fusion molecule and the sgRNA are packaged in the same liposome or lipid nanoparticle or in different liposomes or lipid nanoparticles; Optionally, the liposome or lipid nanoparticle is composed of an ionizable lipid (20% to 70% by mole), a PEGylated lipid (0% to 30% by mole), a supporting lipid (30% to 50% by mole), and cholesterol (10% to 50% by mole); Optionally, the ionizable lipid is selected from the group consisting of a pH-responsive ionizable lipid, a thermally responsive ionizable lipid, and a light-responsive ionizable lipid; Optionally, the fusion molecule is packaged in an AAV vector; Optionally, the fusion molecule and sgRNA are packaged in an AAV vector; or Optionally, the fusion molecule and the sgRNA are packaged in the same AAV vector or in different AAV vectors. The composition of claim 1.
19. A pharmaceutical composition comprising the composition of any one of claims 1 to 18 and a pharmaceutically acceptable carrier.
20. An sgRNA comprising a sequence complementary to a target DNA sequence located near a VEGF gene and / or within a VEGF regulatory element, and optionally located within 500 bp upstream to 500 bp downstream of the transcription start site of the VEGF gene; Optionally, the sgRNA, wherein the VEGF gene is a VEGF-A gene derived from a mammal, such as a human, monkey, mouse, rat, or rabbit.
21. comprising the nucleic acid sequence of any one of SEQ ID NOs: 29-58 and 60-84, and optionally further comprising the tracr sequence set forth in SEQ ID NO: 59; Optionally, the VEGF gene is a VEGF-A gene derived from a mammal, such as a human, monkey, mouse, rat, or rabbit. The sgRNA of claim 20.
22. A nucleic acid molecule encoding the sgRNA of claim 20 or 21.
23. (a) a fusion molecule comprising at least one DNA binding protein and at least one modulator of gene expression, or a nucleic acid sequence encoding said fusion molecule; and (b) a guide molecule comprising the sgRNA of claim 20 and a protein-binding sequence capable of binding to the at least one DNA-binding protein, or a nucleic acid sequence encoding the guide molecule. A composition comprising: A composition wherein said at least one modulator of gene expression provides at least one nucleotide modification near said VEGF gene and / or within a VEGF regulatory element.
24. 24. A composition according to any one of claims 1 to 18 and 23 for use in a method for reducing or eliminating expression of a VEGF gene product in a cell, the method comprising introducing the composition into the cell, thereby reducing or eliminating expression of the VEGF gene product in the cell.
25. 24. A composition described in any one of claims 1 to 18 and 23 for use in an in vivo method of reducing or eliminating expression of a VEGF gene product in a subject, the method comprising the step of introducing the composition into cells of the subject, thereby reducing or eliminating expression of the VEGF gene product in the subject.
26. 26. The composition of claim 25, wherein the subject is a mammal, such as a human, monkey, mouse, rat, rabbit, pig, horse, cat, and dog.
27. 25. The composition of claim 24, wherein the cell is a retinal cell, a retinal pigment epithelial (RPE) cell, or a choroidal cell.
28. 26. The composition of claim 25, wherein the fusion molecule is delivered to the subject by local injection, such as intraocular injection or intravitreal injection.
29. 24. The composition of any one of claims 1 to 18 and 23 for use in treating or alleviating symptoms of a VEGF-associated disorder in a subject.
30. 30. The composition of claim 29, wherein the VEGF-related disorder is a disease involving angiogenesis, optionally a neovascular disorder, such as an ophthalmic neovascular disorder, including age-related macular degeneration (AMD).
31. 30. The composition of claim 29, wherein the subject is a mammal, such as a human, monkey, mouse, rat, rabbit, pig, horse, cat, and dog.
32. 24. A kit comprising a container containing the composition of any one of claims 1 to 18 and 23.