Fusion protein and application thereof
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2026-03-19
AI Technical Summary
The existing CRISPR-Cas9 system has shortcomings in DNA cleavage specificity and editing efficiency, especially in terms of high specificity cleavage and low risk of off-targeting.
A fusion protein was designed to form a dCas9-FokI consortium by combining non-catalytically active Cas9 and FokI nucleases, and to increase the specificity and efficiency of DNA cleavage by replacing the irregularly curled domain of dCas9 into the FokI domain.
High specific cleavage, high editing efficiency, and low off-target risk, and in some cases surpass the performance of wild-type SpCas9.
Abstract
Description
A fusion protein and its application
[0001] This application claims the benefit of Chinese Patent Application No. 2023112129154, filed September 20, 2023. This application incorporates the entirety of the aforementioned Chinese Patent Application. Technical Field
[0002] The present disclosure belongs to the field of biotechnology, and specifically relates to a fusion protein and its application. Background Art
[0003] CRISPR-Cas9 is an adaptive immune defense system developed by bacteria and archaea over a long period of evolution, used to combat invading viruses and foreign DNA. SpCas9, derived from the CRISPR / Cas9 system of Streptococcus pyogenes, is widely used in genetic engineering due to its ease of use and high efficiency. The CRISPR-Cas9 system consists of Cas9, crRNA, and tracrRNA. The Cas9 protein is a DNA endonuclease expressed by many bacteria. The Cas9 endonuclease must be guided by a guide RNA (a complex of crRNA and tracrRNA) to cleave DNA. Cas9, crRNA, and tracrRNA first combine to form a complex that recognizes and binds to the crRNA's complementary target sequence. Subsequently, the Cas9 protein's HNH nuclease domain active site cleaves the crRNA's complementary DNA strand, while the Ruvc nuclease domain active site cleaves the non-complementary strand, creating a double-strand break (DSB) in the DNA. This allows for specific DNA editing through NHEJ or HDR.
[0004] Summary of the Invention
[0005] The present disclosure provides a fusion protein and its application. The fusion protein disclosed herein is used for gene editing and can achieve high specific cleavage, high editing efficiency and / or low off-target risk.
[0006] To improve DNA cleavage specificity, we generated a fusion of the catalytically inactive Cas9 and the FokI nuclease. We designed, screened, and tested the position of FokI within the fusion protein, ultimately obtaining a new fusion protein that, in some instances, exhibited superior editing activity to currently known Cas9 terminally fused to FokI.
[0007] In some cases, the indel efficiency of the dCas9-FokI disclosed herein combined with a pair of gRNAs is even higher than that of wild-type SpCas9.
[0008] It should be understood by those skilled in the art that the dCas9-FokI disclosed herein, in combination with a pair of gRNAs, generates two single-strand breaks on complementary DNA duplexes, thereby forming a double-strand break. Compared to the direct introduction of double-strand breaks by wild-type SpCas9, the dCas9-FokI disclosed herein is expected to reduce off-target activity, for example, see Maryam Saifaldeen, et al., CRISPR FokIDead Cas9 System: Principles and Applications in Genome Engineering, Cells, 2020, 9, 2518.
[0009] A first aspect of the present disclosure provides a fusion protein obtained by replacing consecutive amino acid residues of the random coil structure of dCas9 with a functional domain, wherein the dCas9 is a Cas9 protein without nuclease activity;
[0010] Optionally, the functional domain is selected from a deaminase domain (including but not limited to a cytosine deaminase domain, an adenine deaminase domain), a UGI domain, a UDG domain, a methylation domain, a demethylation domain, a subcellular localization signal (including but not limited to a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, a chloroplast localization signal), a transcription activation domain, a transcription repression domain, a nuclease domain (including but not limited to a FokI domain), a histone deacetylation domain, and a DNA ligation domain; further, the functional domain is a nuclease domain; further, the functional domain is a FokI domain;
[0011] Optionally, the random coil structure is located in the REC1-A and / or REC1-B domain of dCas9;
[0012] Optionally, the random coil structure is located in the RuvC, loop L2 domain and / or HNH domain of dCas9; or the random coil structure is located in the RuvC domain of dCas9; or the random coil structure is located in the RuvC-Ⅲ domain of dCas9.
[0013] A second aspect of the present disclosure provides a fusion protein comprising dCas9 and a functional domain, wherein the consecutive amino acid residues in the random coil in the dCas9 are replaced with the functional domain; the dCas9 is a Cas9 protein without nuclease activity;
[0014] Optionally, the functional domain is selected from a deaminase domain (including but not limited to a cytosine deaminase domain, an adenine deaminase domain), a UGI domain, a UDG domain, a methylation domain, a demethylation domain, a subcellular localization signal (including but not limited to a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, a chloroplast localization signal), a transcription activation domain, a transcription repression domain, a nuclease domain (including but not limited to a FokI domain), a histone deacetylation domain, and a DNA ligation domain; further, the functional domain is a nuclease domain; further, the functional domain is a FokI domain;
[0015] Optionally, the random coil structure is located in the REC1-A and / or REC1-B domain of dCas9;
[0016] Optionally, the random coil structure is located in the RuvC, loop L2 domain and / or HNH domain of dCas9; or the random coil structure is located in the RuvC domain of dCas9; or the random coil structure is located in the RuvC-Ⅲ domain of dCas9.
[0017] In some embodiments of the present disclosure, the fusion protein is capable of specifically binding to the target nucleic acid under the guidance of a guide polynucleotide.
[0018] In some embodiments of the present disclosure, the fusion protein is capable of sequence-specifically binding to a target nucleic acid under the guidance of a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence, and the guide sequence hybridizes to the target nucleic acid.
[0019] In some embodiments of the present disclosure, the fusion protein is capable of forming a complex with a guide polynucleotide comprising a guide sequence engineered to direct sequence-specific binding of the complex to a target nucleic acid.
[0020] In some specific embodiments of the present disclosure, the functional domain is a FokI domain, and the fusion protein is capable of recognizing a double-stranded target nucleic acid and nicking the target nucleic acid under the guidance of a guide polynucleotide; or the fusion protein is capable of forming a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the complex sequence to specifically bind to a double-stranded target nucleic acid and nick the target nucleic acid.
[0021] In some preferred embodiments of the present disclosure, the guide polynucleotide further comprises a backbone sequence; further, the backbone sequence interacts with the dCas9 portion of the fusion protein, or the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein.
[0022] In some embodiments of the present disclosure, the Cas9 protein is selected from SpCas9, SaCas9, NmeCas9, StCas9 and CjCas9.
[0023] In some preferred embodiments of the present disclosure, the Cas9 protein is SpCas9 (Streptococcus pyogenes Cas9) or SaCas9 (Staphylococcus aureus Cas9).
[0024] In some specific embodiments of the present disclosure, the amino acid sequence of the SpCas9 is shown in SEQ ID NO: 1.
[0025] In some embodiments of the present disclosure, the amino acid sequence of the Cas9 protein has ≥80%, ≥85%, ≥90%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99% or ≥99.5% sequence identity with the amino acid sequence shown in SEQ ID NO:1.
[0026] In some specific embodiments of the present disclosure, the amino acid sequence of the dCas9 is as shown in SEQ ID NO: 2 or has a sequence identity of ≥80%, ≥85%, ≥90%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99% or ≥99.5% to the amino acid sequence shown in SEQ ID NO: 2.
[0027] In some embodiments of the present disclosure, the dCas9 is a Cas9 protein in which a mutation, such as a point mutation, occurs in the RuvC domain and / or the HNH domain, rendering the Cas9 protein nuclease-inactive.
[0028] In some preferred embodiments of the present disclosure, the dCas9 is a Cas9 in which the positions corresponding to D10 and N863 of SpCas9 are mutated to A; or the dCas9 is a Cas9 in which the positions corresponding to D10 and H840 of SpCas9 are mutated to A. The corresponding positions can be determined by sequence alignment of the Cas9 and SpCas9.
[0029] In some preferred embodiments of the present disclosure, the dCas9 is SpCas9 with D10A and N863A mutations, or SpCas9 with D10A and H840A mutations.
[0030] In some preferred embodiments of the present disclosure, the dCas9 is SaCas9 with D10A and H557A mutations.
[0031] In some embodiments of the present disclosure, the functional domain is fused to dCas9 directly or through a linker. In some embodiments of the present disclosure, the functional domain is directly fused to dCas9 (i.e., the N-terminal residue of the functional domain is directly covalently linked to the N-terminal residue of the dCas9 split by a peptide bond, and the C-terminal residue of the functional domain is directly covalently linked to the C-terminal residue of the dCas9 split by a peptide bond), or the functional domain is fused to dCas9 through an amino acid sequence. In some embodiments of the present disclosure, the functional domain is directly fused to dCas9, or the functional domain is fused to dCas9 through a linker sequence of 1-50, 1-40, 1-30, 1-20, 5-16, 1-10, 10-20, 3-8 or 15-20 amino acid residues in length.
[0032] In some specific embodiments of the present disclosure, the linker is an amino acid residue sequence GGSGS (SEQ ID NO: 50) or SGSETPGTSESATPES (SEQ ID NO: 51).
[0033] In some embodiments of the present disclosure, the fusion protein is obtained by replacing 1-20 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain.
[0034] In some embodiments of the present disclosure, the fusion protein is obtained by replacing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain.
[0035] In some embodiments of the present disclosure, the fusion protein is obtained by replacing 1-15, 3-11, 3-10 or 9-11 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain.
[0036] In some embodiments of the present disclosure, the length of the replaced random coil structure is 1-15 amino acid residues.
[0037] In some specific embodiments of the present disclosure, the length of the replaced random coil structure is 3-11 amino acid residues.
[0038] In other specific embodiments of the present disclosure, the length of the replaced random coil structure is 3-10 amino acid residues.
[0039] In some other specific embodiments of the present disclosure, the length of the replaced random coil structure is 9-11 amino acid residues.
[0040] In some other specific embodiments of the present disclosure, the length of the replaced random coil structure is 3 amino acid residues.
[0041] In some embodiments of the present disclosure, the random coil structure corresponds to positions 37-41, 107-113, 147-149, 169-172, 308-314, 792-799, 864-871, 906-908, 943-952 and / or 1015-1017 of the sequence shown in SEQ ID NO: 2.
[0042] In some preferred embodiments of the present disclosure, when the amino acid sequence of the dCas9 is as shown in SEQ ID NO: 2, the substitution occurs at positions 37-41, 107-113, 147-149, 169-172, 308-314, 792-799, 864-871, 906-908, 943-952 or 1015-1017.
[0043] In the present disclosure, the position of the random coil structure can be determined by sequence alignment with the sequence shown in SEQ ID NO: 2.
[0044] In some embodiments of the present disclosure, the functional domain is fused to dCas9 directly or through a linker after replacing the replaced sequence of dCas9.
[0045] In some embodiments of the present disclosure, the length of the replaced sequence is 1-15 amino acid residues.
[0046] In other specific embodiments of the present disclosure, the length of the replaced sequence is 3-11 amino acid residues.
[0047] In other specific embodiments of the present disclosure, the length of the replaced sequence is 3-10 amino acid residues.
[0048] In other specific embodiments of the present disclosure, the length of the replaced sequence is 9-11 amino acid residues.
[0049] In other specific embodiments of the present disclosure, the length of the replaced sequence is 3 amino acid residues.
[0050] In some specific embodiments of the present disclosure:
[0051] When the replaced sequence is located at positions 37-41 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 5 amino acid residues;
[0052] When the replaced sequence is located at positions 107-113 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 7 amino acid residues;
[0053] When the replaced sequence is located at positions 147-149 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues;
[0054] When the replaced sequence is located at positions 169-172 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 4 amino acid residues;
[0055] When the replaced sequence is located at positions 308-314 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 7 amino acid residues;
[0056] When the replaced sequence is located at positions 792-799 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 8 amino acid residues;
[0057] When the replaced sequence is located at positions 864-871 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 8 amino acid residues;
[0058] When the replaced sequence is located at positions 906-908 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues;
[0059] When the replaced sequence is located at positions 943-952 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 10 amino acid residues;
[0060] When the replaced sequence is located at positions 1015-1017 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues.
[0061] In some embodiments of the present disclosure, the FokI domain includes the DNA binding domain and the cleavage domain of the FokI endonuclease.
[0062] In some embodiments of the present disclosure, the FokI domain includes or is the DNA cleavage domain of FokI endonuclease.
[0063] In the present disclosure, the FokI domain can exist in the form of a monomer or a dimer. For example, the FokI domain can exist in the form of a monomer in solution.
[0064] In some embodiments of the present disclosure, the FokI domain has nuclease activity.
[0065] In one example of the present disclosure, the first fusion protein can bind to one single strand of a double-stranded target nucleic acid together with the first guide polynucleotide, and the second fusion protein can bind to the other single strand of the target nucleic acid together with the second guide polynucleotide;
[0066] The FokI domains of the first fusion protein and the second fusion protein have a spacer sequence on the double-stranded target nucleic acid and can form a dimer, wherein the length of the spacer sequence is 1-100bp, and the dimer can cut the double-stranded target nucleic acid when the first fusion protein, the first guide polynucleotide, the second fusion protein and the second guide polynucleotide bind to the target nucleic acid.
[0067] In one example of the present disclosure, the FokI domain can dimerize after dCas9 binds to the double-stranded target DNA to achieve cutting of the double-stranded target DNA through its nuclease activity. When the FokI domain dimerizes, the two FokI domains have a spacer sequence on the double-stranded target DNA to cut the double-stranded target DNA. The spacer sequence here does not refer to the spacer sequence of the CRISPR array in the CRISPR system.
[0068] In the present disclosure, the length of the spacer sequence may be 1-100bp, 1-90bp, 1-80bp, 1-70bp, 1-60bp, 1-50bp, 1-40bp, 1-30bp, 1-20bp, 1-10bp, 1-5bp, 2-100bp, 2-90bp, 2-80bp, 2-70bp, 2-60bp, 2-50bp, 2-40bp, 2-30bp, 2-2 0bp, 2-10bp, 2-5bp, 3-100bp, 3-90bp, 3-80bp, 3-70bp, 3-60bp, 3-50bp, 3-40bp, 3-30bp, 3-2 0bp, 3-10bp, 3-5bp, 4-100bp, 4-90bp, 4-80bp, 4-70bp, 4-60bp, 4-50bp, 4-40bp, 4-30bp, 4-2 0bp, 4-10bp, 4-5bp, 5-100bp, 5-90bp, 5-80bp, 5-70bp, 5-60bp, 5-50bp, 5-40bp, 5-30bp, 5- 20bp, 5-10bp, 10-100bp, 10-90bp, 10-80bp, 10-70bp, 10-60bp, 10-50bp, 10-40bp, 10-30bp, In some embodiments, the spacer sequence has a length of 4-30bp. In some embodiments, the spacer sequence has a length of 6-20bp. In some embodiments of the present disclosure, the length of the spacer sequence is 1bp, 2bp, 3bp, 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 11bp, 12bp, 13bp, 14bp, 15bp, 16bp, 17bp, 18bp, 19bp, 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp or 40bp.The length of the spacer sequence is, for example, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, 20 bp, 21 bp, 22 bp, 23 bp, 24 bp, 25 bp, 26 bp, 27 bp, 28 bp, 29 bp or 30 bp.
[0069] In some embodiments of the present disclosure, the amino acid sequence of the FokI domain has a sequence identity of ≥50%, ≥60%, ≥70%, ≥80%, ≥90%, ≥95%, ≥98% or ≥99% to the amino acid residue sequence at positions 6 to 201 of the sequence shown in SEQ ID NO:4.
[0070] In some embodiments of the present disclosure, the FokI domain comprises the amino acid residue sequence from position 6 to position 201 of the sequence shown in SEQ ID NO:4.
[0071] In some embodiments of the present disclosure, the FokI domain consists of the amino acid residue sequence from position 6 to position 201 of the sequence shown in SEQ ID NO:4.
[0072] In some embodiments of the present disclosure, the amino acid sequence of the FokI domain is as shown in SEQ ID NO:52 or has a sequence identity of ≥50%, ≥60%, ≥70%, ≥80%, ≥90%, ≥95%, ≥98% or ≥99% to the amino acid sequence shown in SEQ ID NO:52.
[0073] In some embodiments of the present disclosure, the guide polynucleotide is a single molecule guide polynucleotide or a bimolecule guide polynucleotide. For example, the guide polynucleotide is a single molecule guide polynucleotide (sgRNA) formed by chimerizing crRNA and tracrRNA in the CRISPR-Cas9 system. For another example, the guide polynucleotide is a group consisting of two separate molecules of crRNA and tracrRNA in the CRISPR-Cas9 system, which is the bimolecule.
[0074] In some embodiments of the present disclosure, the fusion protein further comprises any one or more of the following: a subcellular localization signal, a DNA binding domain, a transcription activation domain, a transcription repression domain, a nuclease domain, a deaminase domain, a UDG domain, a UGI domain, a methylase, a demethylase, a histone deacetylase, a DNA ligase, an epitope tag, and a reporter protein;
[0075] Optionally, the subcellular localization signal is selected from: a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, and a chloroplast localization signal.
[0076] In the present disclosure, the fusion protein is preferably a fusion protein obtained by replacing amino acid residues 943-952 of a nuclease activity-inactivated mutant of the Cas9 protein with a functional domain as shown in SEQ ID NO: 1;
[0077] Optionally, the inactivation mutant has D10A and N863A mutations, or D10A and H840A mutations relative to SEQ ID NO: 1;
[0078] Optionally, the functional domain is a FokI domain; further, the FokI domain is a DNA cleavage domain of a FokI endonuclease; further, the amino acid sequence of the FokI domain is shown in SEQ ID NO: 52;
[0079] Optionally, the functional domain is connected to the N-terminus and C-terminus of the inactivated Cas9 mutant via a linker sequence; further, the functional domain is connected to the N-terminus and C-terminus of the inactivated Cas9 mutant via a linker with an amino acid sequence as shown in SEQ ID NO: 50 and SEQ ID NO: 51, respectively;
[0080] Optionally, the N-terminus and / or C-terminus of the fusion protein is fused with at least one nuclear localization sequence;
[0081] Optionally, the fusion protein is capable of forming a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the complex sequence to specifically bind to the double-stranded target DNA and nick the target DNA; further, the guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide; further, the guide polynucleotide is a single-molecule guide polynucleotide formed by chimerizing crRNA and tracrRNA in the CRISPR-Cas9 system, or the guide polynucleotide is a group consisting of two separate molecules, crRNA and tracrRNA in the CRISPR-Cas9 system, and the crRNA consists of a guide sequence and a direct repeat sequence.
[0082] In some embodiments of the present disclosure, the fusion protein comprises the amino acid sequence shown in SEQ ID NO:3.
[0083] In some specific embodiments of the present disclosure, the amino acid sequence of the fusion protein is shown in SEQ ID NO: 3.
[0084] A second aspect of the present disclosure provides a ribonucleoprotein complex (RNP complex), comprising a guide polynucleotide and a fusion protein as described in the first aspect.
[0085] In some embodiments of the present disclosure, the guide polynucleotide guides the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid.
[0086] In some embodiments of the present disclosure, the guide polynucleotide comprises a guide sequence that is engineered to direct the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid.
[0087] In some preferred embodiments of the present disclosure, the guide polynucleotide further comprises a backbone sequence; the backbone sequence interacts with the dCas9 portion of the fusion protein, or the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein.
[0088] In some embodiments of the present disclosure, the guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide.
[0089] A third aspect of the present disclosure provides a gene editing system, comprising:
[0090] (1) the fusion protein according to the first aspect, or a nucleic acid encoding the same; and / or,
[0091] (2) A guiding polynucleotide, or a nucleic acid encoding the same.
[0092] In some embodiments of the present disclosure, the guide polynucleotide guides the fusion protein sequence to specifically bind to the target nucleic acid.
[0093] In some embodiments of the present disclosure, the guide polynucleotide comprises a guide sequence that is engineered to direct the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid.
[0094] In some preferred embodiments of the present disclosure, the guide polynucleotide further comprises a backbone sequence; the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein.
[0095] In some embodiments of the present disclosure, the gene editing system comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different fusion proteins. In some embodiments of the present disclosure, the gene editing system comprises 1 fusion protein.
[0096] In some embodiments of the present disclosure, the gene editing system comprises at least two different guide polynucleotides.
[0097] In some embodiments of the present disclosure, the gene editing system comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different guide polynucleotides. In some embodiments of the present disclosure, the gene editing system comprises 2 different guide polynucleotides.
[0098] In some embodiments of the present disclosure, the guide sequences are different between the different guide polynucleotides, and the backbone sequences are the same. In some embodiments of the present disclosure, the guide sequences are different between the different guide polynucleotides, and the backbone sequences are different.
[0099] In some embodiments of the present disclosure, the guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide. In some embodiments of the present disclosure, the guide polynucleotide is a single-molecule guide polynucleotide obtained by chimerizing crRNA and tracrRNA. In some embodiments of the present disclosure, the guide polynucleotide is a single-molecule guide polynucleotide obtained by chimerizing crRNA and tracrRNA through nucleotide sequences. In some embodiments of the present disclosure, the guide polynucleotide is a double-molecule guide polynucleotide composed of a crRNA molecule and a tracrRNA molecule.
[0100] The fourth aspect of the present disclosure provides a polynucleotide encoding the fusion protein as described in the first aspect or the ribonucleoprotein complex as described in the second aspect.
[0101] In some embodiments of the present disclosure, in the polynucleotide encoding the ribonucleoprotein complex, the polynucleotide encoding the guide polynucleotide and the polynucleotide encoding the fusion protein are on the same or different nucleic acid chains.
[0102] The fifth aspect of the present disclosure provides an expression cassette, comprising a promoter and the polynucleotide according to the fourth aspect; the promoter sequence regulates the expression of the coding sequence.
[0103] The sixth aspect of the present disclosure provides a recombinant expression vector comprising the polynucleotide as described in the fourth aspect, and optionally a cis-acting element and / or a trans-acting factor encoding gene.
[0104] In some embodiments of the present disclosure, the cis-acting element is selected from a promoter, an enhancer, and a silencer.
[0105] In some embodiments of the present disclosure, the trans-acting factor encoding gene is selected from nucleotides encoding polymerases, transcription factors and / or transcription regulatory factors.
[0106] The seventh aspect of the present disclosure provides a composition comprising a delivery vector, and the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect and / or the recombinant expression vector as described in the sixth aspect.
[0107] The eighth aspect of the present disclosure provides a recombinant cell, which comprises the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect and / or the composition as described in the seventh aspect.
[0108] The ninth aspect of the present disclosure provides a pharmaceutical composition, comprising the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect, the composition as described in the seventh aspect and / or the recombinant cell as described in the eighth aspect, and optionally a pharmaceutically acceptable carrier and / or excipient.
[0109] The tenth aspect of the present disclosure provides a kit, which comprises the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect, the composition as described in the seventh aspect, the recombinant cell as described in the eighth aspect and / or the pharmaceutical composition as described in the ninth aspect.
[0110] The eleventh aspect of the present disclosure provides a fusion protein as described in the first aspect, a ribonucleoprotein complex as described in the second aspect, a gene editing system as described in the third aspect, a polynucleotide as described in the fourth aspect, an expression cassette as described in the fifth aspect, a recombinant expression vector as described in the sixth aspect, a composition as described in the seventh aspect, a recombinant cell as described in the eighth aspect, a pharmaceutical composition as described in the ninth aspect and / or a kit as described in the tenth aspect for the preparation of a drug for diagnosing, preventing and / or treating a disease or condition associated with a target nucleic acid.
[0111] The twelfth aspect of the present disclosure provides a method for diagnosing, preventing and / or treating a disease or condition associated with a target nucleic acid, the method comprising administering to a patient in need thereof the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect, the composition as described in the seventh aspect, the recombinant cell as described in the eighth aspect, the pharmaceutical composition as described in the ninth aspect and / or the kit as described in the tenth aspect.
[0112] The thirteenth aspect of the present disclosure provides a fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect, the composition as described in the seventh aspect, the recombinant cell as described in the eighth aspect, the pharmaceutical composition as described in the ninth aspect and / or the kit as described in the tenth aspect, wherein the fusion protein, ribonucleoprotein complex, gene editing system, polynucleotide, expression cassette, recombinant expression vector, composition, recombinant cell, pharmaceutical composition or kit is used for diagnosing, preventing and / or treating diseases or conditions associated with the target nucleic acid.
[0113] The fourteenth aspect of the present disclosure provides a method for editing a target nucleic acid in vitro, ex vivo or in vivo, the method comprising the step of contacting the fusion protein as described in the first aspect, the ribonucleoprotein complex as described in the second aspect, the gene editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the expression cassette as described in the fifth aspect, the recombinant expression vector as described in the sixth aspect, the composition as described in the seventh aspect, the recombinant cell as described in the eighth aspect, the pharmaceutical composition as described in the ninth aspect and / or the kit as described in the tenth aspect with the target nucleic acid, so that the sequence of the target nucleic acid changes or the expression level of the target nucleic acid is changed.
[0114] In one example of the present disclosure, the method includes the steps of contacting the first fusion protein, the second fusion protein, the first guide polynucleotide, and the second guide polynucleotide as described herein with a double-stranded target nucleic acid, so that the sequence of the target nucleic acid changes or the expression level of the target nucleic acid is changed;
[0115] The first fusion protein can bind to one single strand of the double-stranded target nucleic acid together with the first guide polynucleotide, and the second fusion protein can bind to the other single strand of the target nucleic acid together with the second guide polynucleotide;
[0116] The FokI domains of the first fusion protein and the second fusion protein have a spacer sequence on the double-stranded target nucleic acid and can form a dimer, wherein the length of the spacer sequence is 1-100bp, and the dimer can cut the double-stranded target nucleic acid when the first fusion protein, the first guide polynucleotide, the second fusion protein and the second guide polynucleotide bind to the target nucleic acid.
[0117] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure.
[0118] The positive progress of this disclosure is:
[0119] The gene editing system containing the fusion protein disclosed herein has high editing efficiency and / or low off-target risk. BRIEF DESCRIPTION OF THE DRAWINGS
[0120] Figure 1 is a schematic diagram of replacing the FokI domain with single-stranded nucleic acid cleavage capability within the dCas9 protein.
[0121] Figure 2 is a schematic diagram of dual gRNA guiding dCas9-FokI to target and cleave two single strands, thereby forming a double-strand break;
[0122] The nucleotide sequence in the figure is shown as SEQ ID NO: 32, and the amino acid sequence is shown as SEQ ID NO: 33.
[0123] Figure 3 is a schematic diagram of the flow cytometry detection results after dCas9-FokI fusion protein editing.
[0124] Figure 4 is a schematic diagram of the editing efficiency of different recombinant protein targeting reporter systems.
[0125] Figure 5 is a schematic diagram of the editing efficiency of different recombinant proteins targeting endogenous genes. DETAILED DESCRIPTION
[0126] The present disclosure is further illustrated by way of examples below, but the present disclosure is not limited to the scope of the examples. Experimental methods in the following examples without specifying specific conditions were performed according to conventional methods and conditions, or selected according to the product specifications.
[0127] the term
[0128] In this disclosure, unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of this disclosure, definitions and explanations of relevant terms are provided below.
[0129] In the present disclosure, the letters in the amino acid sequence represent the single-letter abbreviations of amino acids well known in the art, such as those described in J. Biol. Chem, 243, p3558 (1968): Alanine: Ala-A, Arginine: Arg-R, Aspartic acid: Asp-D, Cysteine: Cys-C, Glutamine: Gln-Q, Glutamic acid: Glu-E, Histidine: His-H, Glycine: Gly-G, Asparagine: Asn-N, Tyrosine: Tyr-Y, Proline: Pro-P, Serine: Ser-S, Methionine: Met-M, Lysine: Lys-K, Valine: Val-V, Isoleucine: Ile-I, Phenylalanine: Phe-F, Leucine: Leu-L, Tryptophan: Trp-W, Threonine: Thr-T.
[0130] In this disclosure, taking the SpCas9 protein as an example, the numbers represented by the sites refer to the positions of the amino acid residues corresponding to the wild-type SpCas9 protein amino acid sequence SEQ ID NO: 1. In this disclosure, the letters before the site represent the original amino acid residue, and the letters after the site represent the substituted amino acid residue.
[0131] In the present disclosure, the term "identity" (identity or percent identity) is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. When a certain position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (for example, a certain position in each of the two DNA molecules is occupied by adenine, or a certain position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent sequence identity" (percent identity) between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100%. For example, if 6 out of 10 positions of the two sequences match, then the two sequences have 60% sequence identity. Typically, two sequences are compared when they are aligned to produce maximum sequence identity. Such an alignment can be performed using published and commercially available alignment algorithms and programs, such as, but not limited to, ClustalΩ, MAFFT, Probcons, T-Coffee, Probalign, BLAST, which can be reasonably selected for use by one of ordinary skill in the art. Those skilled in the art can determine appropriate parameters for aligning sequences, including, for example, any algorithms needed to achieve better alignment or optimal comparison over the entire length of the sequences being compared, as well as any algorithms needed to achieve better alignment or optimal comparison over a portion of the sequences being compared.
[0132] In the present disclosure, the comparison of sequences and the determination of the percent identity between two sequences can be achieved using a mathematical algorithm. In some cases, the percent identity between two amino acid sequences is determined using the Needleman and Wunsch ((1970) J. Mol. Biol. 48:444-453) algorithm, using the Blossum 62 scoring matrix, with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5, which has been incorporated into the GAP program in the GCG software package.
[0133] Cas9 protein and gene editing
[0134] Cas9 is a common Cas protein in CRISPR II systems. Some wild-type Cas9s have been shown to have nuclease activity, capable of cleaving nucleic acids or nucleic acid chains. Cas9 proteins typically form RNP complexes with guide polynucleotides and are drawn to the cleavage site by the guide polynucleotide.
[0135] Cas9 molecules of many species can be used for the fusion proteins described herein. Although Streptococcus pyogenes and Staphylococcus aureus Cas9 molecules are the subject of many disclosures herein, Cas9 molecules of other species listed herein, Cas9 molecules derived from or based on the Cas9 proteins of the species may likewise be used. In other words, although many descriptions herein use Streptococcus pyogenes and Staphylococcus aureus Cas9 molecules, Cas9 molecules from other species may replace them.
[0136] Gene editing is the editing of a single base or nucleic acid chain based on a variety of gene editing tools, thereby changing the original nucleic acid at the molecular level. Therefore, gene editing includes single base editing as well as editing of nucleic acid chains. As used herein, the term "single base" refers to one and only one nucleotide in a nucleic acid sequence. When used in the context of single base editing, it refers to replacing a base at a specific position in a nucleic acid sequence with a different base. Such replacement can occur through many mechanisms, including but not limited to substitution or modification.
[0137] As used herein, the term "nicking" refers to breaking the interstrand or internucleotide bonds in a nucleic acid chain, thereby creating a gap in a single-stranded or double-stranded nucleic acid. The term "cleavage" refers to breaking at least the interstrand bonds in a double-stranded nucleic acid, or at least the internucleotide bonds in a single-stranded nucleic acid, thereby completely separating the nucleic acid chains upstream and downstream of the cleavage site.
[0138] FokI domain
[0139] FokI is a type II restriction endonuclease comprising a DNA recognition domain and a catalytic (endonuclease) domain. The fusion proteins described herein can include the entire FokI, or only the nuclease domain of FokI, e.g., amino acids 388-583 or 408-583 of GenBank Accession No. AAA24927.1, e.g., as described in Li et al., Nucleic Acids Res. 39(1):359–372 (2011); Cathomen and Joung, Mol. Ther. 16:1200–1207 (2008), or as described in Miller et al., Nat Biotechnol 25:778–785 (2007); Szczepek et al., Nat Biotechnol 25:786–793 (2007); or a mutant form of FokI as described in Bitinaite et al., Proc. Natl. Acad. Sci. USA. 95:10570–10575 (1998).
[0140] The FokI domain can be the DNA cleavage domain of a wild-type FokI endonuclease; the FokI domain can be a mutant of the DNA cleavage domain of a wild-type FokI endonuclease, including but not limited to single-point mutants and multiple-point mutants. In some embodiments, two FokI domains can form a dimer. In some embodiments, two FokI domains cannot form a dimer. In some embodiments, the two FokI domains dimerize to obtain nuclease activity. In some embodiments, the two FokI domains dimerize to obtain target nucleic acid cleavage. In some embodiments, the FokI domains can cleave target nucleic acid without dimerization. In some embodiments, the FokI domains do not dimerize to obtain nuclease activity.
[0141] Guide polynucleotide
[0142] As used herein, the term "guide polynucleotide" is used to refer to a molecule in the CRISPR-Cas system that forms a CRISPR complex with the Cas protein and guides the CRISPR complex to the target nucleic acid. Typically, the guide polynucleotide comprises a backbone sequence connected to a guide sequence that can hybridize with the target nucleic acid sequence. The backbone sequence typically comprises a direct repeat sequence and may also comprise a tracrRNA sequence. The direct repeat sequence and the guide sequence constitute crRNA.
[0143] In some embodiments, the guidance polynucleotide is a guide RNA. In some embodiments, the guidance polynucleotide is a guide DNA. In some embodiments, at least one base of the guidance polynucleotide is a DNA base. In some embodiments, the guidance polynucleotide is a chemically modified guidance polynucleotide. In some embodiments, the guidance polynucleotide comprises at least one chemically modified nucleotide.
[0144] In some embodiments, the guide polynucleotide comprises at least one guide sequence (also known as a spacer sequence) linked to at least one direct repeat (DR). In some embodiments, the guide sequence is located at the 3' end of the backbone sequence. In some embodiments, the guide sequence is located at the 5' end of the backbone sequence.
[0145] In some embodiments, the same direction repeat sequence comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55 or at least 60 nucleotides. In some embodiments, the same direction repeat sequence comprises no more than 70 nucleotides, no more than 60 nucleotides, no more than 55 nucleotides, no more than 50 nucleotides, no more than 45 nucleotides, no more than 40 nucleotides, no more than 35 nucleotides, no more than 30 nucleotides, no more than 25 nucleotides, no more than 20 nucleotides or no more than 15 nucleotides. In some embodiments, the same direction repeat sequence comprises 5-70 nucleotides, 5-50 nucleotides, 5-30 nucleotides, 5-20 nucleotides, 5-15 nucleotides, 10-40 nucleotides, 10-30 nucleotides, 10-20 nucleotides or 10-15 nucleotides.
[0146] In some embodiments, the tracrRNA sequence comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, or at least 80 nucleotides. In some embodiments, the tracrRNA sequence comprises no more than 120 nucleotides, no more than 110 nucleotides, no more than 100 nucleotides, no more than 90 nucleotides, no more than 80 nucleotides, no more than 70 nucleotides, no more than 60 nucleotides, no more than 55 nucleotides, no more than 50 nucleotides, no more than 45 nucleotides, no more than 40 nucleotides, no more than 35 nucleotides, no more than 30 nucleotides, no more than 25 nucleotides, no more than 20 nucleotides, or no more than 15 nucleotides. In some embodiments, the direct repeat sequence comprises 10-120 nucleotides, 20-100 nucleotides, 30-90 nucleotides, 40-80 nucleotides, 40-70 nucleotides, 50-70 nucleotides, or 60-70 nucleotides.
[0147] In some embodiments, the guide sequence comprises at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, or at least 30 nucleotides. In some embodiments, the guide sequence comprises no more than 60 nucleotides, no more than 55 nucleotides, no more than 50 nucleotides, no more than 45 nucleotides, no more than 40 nucleotides, no more than 35 nucleotides, no more than 30 nucleotides, no more than 25 nucleotides, or no more than 20 nucleotides. In some embodiments, the guide sequence comprises 15-60 nucleotides, 15-40 nucleotides, 15-30 nucleotides, 15-25 nucleotides, or 17-25 nucleotides.
[0148] In some embodiments, the guide sequence has sufficient complementarity to the target nucleic acid to hybridize to the target nucleic acid and guide sequence-specific binding of the fusion protein of the present disclosure to the target nucleic acid. In some embodiments, the guide sequence has 100% complementarity to the target nucleic acid, but the guide sequence can have less than 100% complementarity to the target nucleic acid DNA, for example, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementarity.
[0149] In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than two nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than one nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with or without mismatches.
[0150] The guiding polynucleotide can be a single molecule or a double molecule. The guiding polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence). Optionally, the guiding polynucleotide can comprise at least one nucleotide, a phosphodiester bond, or a linkage modification, such as, but not limited to, locked nucleic acid (LNA), 5-methyl dC, 2,6-diaminopurine, 2'-fluoro A, 2'-fluoro U, 2'-O-methyl RNA, a phosphorothioate bond, a connection to a cholesterol molecule, a connection to a polyethylene glycol molecule, a connection to a spacer 18 (hexaethylene glycol chain) molecule, or a 5' to 3' covalent connection resulting in cyclization.
[0151] The guide polynucleotide can be a single molecule guide polynucleotide or a bimolecule guide polynucleotide. Non-limiting examples include, for example, crRNA and tracrRNA in the CRISPR-Cas9 system are chimeric (e.g., connected by a GAAA sequence) to form a single molecule guide polynucleotide (sgRNA). For another example, the guide polynucleotide is a group consisting of two separate molecules of crRNA and tracrRNA in the CRISPR-Cas9 system, i.e., the bimolecule.
[0152] The ability of the guide polynucleotide to guide the sequence-specific binding of the ribonucleoprotein complex to the target DNA can be evaluated by any suitable assay. For example, the components of the CRISPR system sufficient to form the ribonucleoprotein complex, including the guide polynucleotide to be tested, can be provided to a host cell with the corresponding target nucleic acid DNA molecule, such as by transfection of a vector encoding the components of the ribonucleoprotein complex, and then the preferential cutting within the target sequence is evaluated. Similarly, the cutting of the target nucleic acid DNA sequence can be evaluated in a test tube by providing the target nucleic acid DNA, the components of the ribonucleoprotein complex, including the guide polynucleotide to be tested and the control guide polynucleotide different from the test guide polynucleotide, and comparing the ability to bind to the target nucleic acid DNA or the rate of cutting the target DNA between the guide polynucleotide to be tested and the control guide polynucleotide. The ability of the ribonucleoprotein complex to cut the target nucleic acid or target DNA can also be evaluated by the above-mentioned assay.
[0153] Gene editing system
[0154] The gene editing systems disclosed herein include, but are not limited to, systems for the following uses: for cutting a target nucleic acid, for introducing a base transition on a target nucleic acid, for inhibiting or reducing the expression of a specific gene on a target nucleic acid, for activating or increasing the expression of a specific gene on a target nucleic acid, for visualizing or detecting a target nucleic acid, for masking a specific gene on a target nucleic acid, and for silencing a specific gene on a target nucleic acid.
[0155] Vector system
[0156] Another aspect of the present disclosure relates to a vector system comprising the fusion protein described herein, the vector system comprising one or more vectors (expression vectors), the vector comprising a polynucleotide sequence encoding the fusion protein and a polynucleotide sequence encoding the guide polynucleotide. Alternatively, the vector comprises a polynucleotide sequence encoding the fusion protein.
[0157] In some embodiments, the vector system comprises at least one plasmid or viral vector (e.g., retrovirus, lentivirus, adenovirus, adeno-associated virus, or herpes simplex virus). In some embodiments, the polynucleotide sequence encoding the fusion protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same vector. In some embodiments, the polynucleotide sequence encoding the fusion protein and the polynucleotide sequence encoding the guide polynucleotide are located on two or more vectors.
[0158] In some embodiments, the polynucleotide sequence encoding the fusion protein and / or the polynucleotide sequence encoding the guidance polynucleotide are operably connected to a regulatory element. Regulatory elements include promoters, enhancers, internal ribosome entry sites (IRES) and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences). Regulatory elements include regulatory elements that make the nucleotide sequence constitutively expressed in many types of host cells, and regulatory elements (e.g., tissue-specific regulatory sequences) that make the nucleotide sequence expressed only in certain host cells. Tissue-specific promoters can be directly expressed mainly in desired tissues of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas) or specific cell types (e.g., lymphocytes). Regulatory elements can also guide expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type-specific. In some embodiments, the regulatory element is an enhancer element, such as WPRE, the CMV enhancer, the R-U5 segment in the LTR of HTLV-1, the SV40 enhancer, or the intronic sequence between exons 2 and 3 of rabbit β-globin.
[0159] In some embodiments, the vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, a phosphoglycerol kinase (PGK) promoter, or an EF1α promoter), or a pol III promoter and a pol II promoter.
[0160] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).
[0161] In some embodiments, the promoter is a tissue-specific promoter, which can be used to drive the tissue-specific expression of the fusion protein. Suitable muscle-specific promoters include but are not limited to CK8, MHCK7, myoglobin promoter (Mb), desmin (Desmin) promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoters. Suitable immune cell-specific promoters include but are not limited to B29 promoter (B cell), CD14 promoter (monocyte), CD43 promoter (leukocyte and platelet), CD68 (macrophage) and SV40 / CD43 promoter (leukocyte and platelet). Suitable blood cell-specific promoters include but are not limited to CD43 promoter (leukocyte and platelet), CD45 promoter (hematopoietic cell), INF-β (hematopoietic cell), WASP promoter (hematopoietic cell), SV40 / CD43 promoter (leukocyte and platelet), and SV40 / CD45 promoter (hematopoietic cell). Suitable pancreas-specific promoters include but are not limited to elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astroglial cells), the SYN1 promoter (neurons), and the NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.
[0162] AAV vectors
[0163] Another aspect of the present disclosure relates to an adeno-associated viral (AAV) vector comprising a fusion protein or gene editing system described herein, wherein the adeno-associated viral (AAV) vector comprises DNA encoding the fusion protein and guide polynucleotide described herein.
[0164] Delivery of the CRISPR-Cas system via an AAV vector is described in Maeder et al., Nature Medicine 25:229-233 (2019), which has been clinically demonstrated to be safe and effective for subretinal delivery of AAV. Local delivery via subretinal injection, the natural tropism of AAV5 for photoreceptor cells, and the use of the photoreceptor-specific GRK1 promoter are all used to limit the expression of the CRISPR / Cas system to therapeutic target tissues and cell types, which is incorporated herein by reference in its entirety. In some embodiments, the AAV vector comprises an ssDNA genome comprising a coding sequence for a fusion protein and a guide polynucleotide flanked by ITRs.
[0165] In some embodiments, the fusion protein or gene editing system described herein is packaged in an AAV vector, such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAVrh74. In some embodiments, the fusion protein or gene editing system described herein is packaged in an AAV vector comprising an engineered capsid with tissue tropism, such as an engineered muscle tropism capsid. Tabebordbar et al., Cell 184:4919-4938 (2021) describes engineering of AAV capsids with tissue tropism by directed evolution, identifies a class of capsids containing RGD motifs, and systemic injection of MyoAAV can efficiently transduce primate muscles, which is incorporated herein by reference in its entirety.
[0166] lipid nanoparticles
[0167] Another aspect of the present disclosure relates to lipid nanoparticles (LNPs) comprising a gene editing system described herein, wherein the LNP comprises a guide polynucleotide described herein and an mRNA encoding a fusion protein described herein.
[0168] Gillmore et al., N.Engl.J.Med., 385:493-502 (2021) describes the LNP delivery of the CRISPR-Cas system, lipid nanoparticles (LNPs) are composed of four lipids, including a proprietary ionizable lipid LP000001; DSPC; cholesterol and DMG-PEG2k, and the LNP suspension is formulated in an aqueous buffer of Tris, NaCl and sucrose, pH 7.4, which is incorporated herein by reference in its entirety. In some embodiments, in addition to the RNA payload (mRNA and guide polynucleotides), the lipid nanoparticle (LNP) also includes four components: a cationic or ionizable lipid, cholesterol, a helper lipid, and a PEG-lipid. In some embodiments, the cation or ionizable lipids include cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102, and Intellia LP01. In some embodiments, the PEG-lipids include PEG-2000-C-DMG, PEG-2000-DMG, or ALC-0159. In some embodiments, the helper lipids include DSPC. The components of LNPs are described in Paunovska et al., Nature Reviews Genetics 23: 265-280 (2022), and FDA-approved LNPs contain variants of four basic ingredients: cation or ionizable lipids, cholesterol, helper lipids, and polyethylene glycol (PEG) lipids, which are incorporated herein by reference in their entirety.
[0169] Lentiviral vectors
[0170] Another aspect of the present disclosure relates to a lentiviral vector comprising a gene editing system as described herein, wherein the lentiviral vector comprises a guide polynucleotide as described herein and an mRNA encoding a fusion protein as described herein. In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein, such as VSV-G. In some embodiments, the mRNA encoding the fusion protein is linked to an aptamer sequence.
[0171] RNP complex
[0172] Another aspect of the present disclosure relates to a ribonucleoprotein complex comprising the ribonucleoprotein complex described herein, wherein the ribonucleoprotein complex is formed by a guidance polynucleotide and a fusion protein as described herein. In some embodiments, the ribonucleoprotein complex can be delivered to eukaryotic cells, mammalian cells, or human cells by microinjection or electroporation. In some embodiments, the ribonucleoprotein complex can be packaged in virus-like particles and delivered to mammals or human subjects in vivo.
[0173] virus-like particles
[0174] Another aspect of the present disclosure relates to a virus-like particle (VLP) comprising the gene editing system described herein, wherein the virus-like particle comprises the guide polynucleotide and fusion protein described herein or a ribonucleoprotein complex consisting of the guide polynucleotide and fusion protein.
[0175] Banskota et al. Cell 185(2):250-265(2022) reported the development and application of DNA-free virus-like particles (eVLPs) for efficient packaging and delivery of base editors or Cas9 ribonucleoprotein; Mangeot et al., Nature Communications 10(1):1-15(2019) used engineered mouse leukemia virus-like particles (Nanoblades) loaded with Cas9-sgRNA ribonucleoprotein to induce efficient genome editing in cell lines and primary cells (including human induced pluripotent stem cells, human hematopoietic stem cells and mouse bone marrow cells); Campbell, et al., Molecular Therapy 27:151-163 (2019) utilizes a kind of specialized extracellular vesicle called " gesicle ", effectively but transiently delivers Cas9 targeting HIV long terminal repeat sequence (LTR) in the form of ribonucleoprotein, Gesicles is produced by expressing vesicular stomatitis virus glycoprotein and packaging protein (as its cargo), so there is no need for transgenic delivery, so that Cas9 expression and Mangeot et al.Molecular Therapy, 19 (9): 1656-1666 (2011) reported that the overexpression of the spike glycoprotein of vesicular stomatitis virus (VSV-G) in human cells induced the release of fusogenic vesicles called gesicles, biochemical and functional studies have shown that glial cells bind proteins from production cells and can transport them to receptor cells, and this protein transduction method allows direct transport of cytoplasmic, nuclear or surface proteins in target cells. These documents all describe engineered VLPs, which are incorporated herein by reference in their entirety.
[0176] In some embodiments, engineered virus-like particles (VLPs) are pseudotyped with homologous or heterologous envelope proteins, such as VSV-G. In some embodiments, the fusion protein is fused to a gag protein (e.g., MLVgag) via a cleavable linker, wherein cleavage of the linker in the target cell exposes an NLS between the linker and the fusion protein. In some embodiments, the fusion protein or conjugate comprises (e.g., from 5' to 3') a gag protein (e.g., MLVgag), one or more NESs, a cleavable linker, one or more NLSs, and Cas12, as described in Banskota et al. Cell 185(2):250-265(2022).
[0177] In some embodiments, the fusion protein is fused to a first dimerization domain that is capable of dimerizing or heterodimerizing with a second dimerization domain fused to a membrane protein, wherein the presence of a ligand promotes the dimerization and enriches the Cas12 protein or fusion protein or conjugate into the VLP as described in Campbell, et al., Molecular Therapy 27: 151-163 (2019).
[0178] cell
[0179] Another aspect of the present disclosure relates to cells comprising fusion proteins as described herein and gene editing systems containing them.Cells (for example, which can be used to produce cell-free systems) can be eukaryotic or prokaryotic.Examples of such cells include, but are not limited to, bacteria, archaebacteria, plants, fungi, yeasts, insects, and mammalian cells, such as lactobacilli, lactococci, bacillus (for example, bacillus subtilis), Escherichia (for example, Escherichia coli), Clostridium, Saccharomyces or Pichia (such as saccharomyces cerevisiae or Pichia pastoris), Kluyveromyces lactis, Salmonella typhimurium, Drosophila cells, Caenorhabditis elegans cells, African clawed frog cells, SF9 cells, C129 cells, 293 cells, Neurospora and immortalized mammalian cell lines (for example, Hela cells, bone marrow cell lines, and lymphoid cell lines).
[0180] In some embodiments, the cell is a prokaryotic cell, such as a bacterial cell, such as escherichia coli. In some embodiments, the cell is a eukaryotic cell, such as a mammalian cell or a human cell. In some embodiments, the cell is a primary eukaryotic cell, a stem cell, a tumor / cancer cell, a circulating tumor cell (CTC), a blood cell (for example, T cell, B cell, NK cell, Tregs etc.), a hematopoietic stem cell, a specialized immune cell (such as tumor infiltrating lymphocytes or tumor suppressor lymphocytes), a stromal cell (such as cancer associated fibroblasts etc.) in a tumor microenvironment. In some embodiments, the cell is the brain or neuronal cell (for example, neuron, astrocyte, microglia, retinal ganglion cell, rod / cone cell etc.) of a central or peripheral nervous system.
[0181] Target nucleic acid or target DNA
[0182] The fusion proteins described herein and the gene editing systems containing them can be used to target one or more target nucleic acids, such as target DNA molecules, such as target DNA molecules present in biological samples, environmental samples (such as soil, air or water samples), etc.
[0183] In some embodiments, the target nucleic acid is a disease-related gene or a signaling biochemical pathway-related gene, or the target nucleic acid is a reporter gene. Non-limiting examples of such target nucleic acids include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed on December 12, 2012 and January 2, 2013, respectively, and International Application No. PCT / US2013 / 074667, filed on December 12, 2013, all of which are incorporated herein by reference.
[0184] Therapeutic applications
[0185] The present disclosure discloses compositions and pharmaceutical compositions comprising the fusion protein or gene editing system. In some embodiments, the pharmaceutical composition is delivered to a human subject in vivo. The pharmaceutical composition can be delivered by any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, intracerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intracameral injection, intratympanic injection, intranasal administration, and inhalation.
[0186] As used herein, the term "effective amount" refers to a specific amount of a pharmaceutical composition comprising a therapeutic agent that achieves a clinically beneficial outcome (i.e., for example, alleviating symptoms). The toxicity and therapeutic efficacy of such a composition can be determined by standard pharmaceutical procedures in cell culture or experimental animals, such as for determining LD50 (the dose that kills 50% of a population) and ED50 (the dose that is therapeutically effective in 50% of a population). The dose ratio between toxicity and therapeutic effect is the therapeutic index, and it can be expressed as the ratio LD50 / ED50. Compounds that exhibit a high therapeutic index are preferred. The data obtained from these cell culture assays and other animal studies can be used to formulate a series of doses for human use. The dosage of such a compound is preferably within the range of a circulating concentration of an ED50 that has little or no toxicity. Depending on the dosage form employed, the sensitivity of the patient, and the route of administration, the dosage varies within this range.
[0187] As used herein, the term "disease" or "disorder" refers to any disruption of the normal state of an organism or plant, or one of its parts, or an impairment that alters the performance of a vital function. It typically manifests as obvious signs and symptoms, which are usually a response to: i) environmental factors (such as malnutrition, industrial hazards, or climate); ii) specific infectious agents (such as helminths, bacteria, or viruses); iii) an inherent defect of the organism (such as a genetic abnormality); and / or iv) a combination of these factors.
[0188] As used herein, the term "administering" or "administering" refers to any method of providing a composition to a patient so that the composition has the intended effect on the patient. Exemplary methods of administration are by direct mechanisms such as local tissue administration (i.e., extravascular placement), oral, transdermal patch, topical, inhalation, suppository, etc.
[0189] As used herein, the term "patient" or "subject" is a human or animal and need not be hospitalized. For example, an outpatient, a person in a nursing home is a "patient." A patient can include humans or non-human animals of any age, and thus includes adults and minors (i.e., children). The term "patient" does not imply a need for medical treatment, and therefore, a patient can voluntarily or involuntarily participate in an experiment, whether clinical or in support of basic scientific research.
[0190] Example 1: Design of different dCas9-FokI fusion proteins and vector plasmid construction
[0191] In order to insert the FokI nuclease domain into the dCas9 protein, bioinformatics analysis and AI methods were used to predict and simulate different possible insertion sites of the dCas9 protein. Under the premise that it was predicted that it would not affect the ability of dCas9 to bind to DNA, the FokI domain with single-stranded nucleic acid cleavage ability was inserted or replaced. Specifically, the random coil of SpCas9 was marked on PyMOL, and the positions where the marked random coil region was greater than 3aa were counted, and then the corresponding domains were marked. The random coil sequence of dSpCas9 was replaced with the FokI domain. The first priority was to replace the amino acid sequence with the FokI domain in the RuvC region; the second priority was to consider that the longer the length of the random coil, the smaller the potential impact of the inserted FokI domain on the overall structure of SpCas9, so the priority was higher. Based on this design, the dCas9-FokI-S1-S10 fusion protein was obtained.
[0192] As shown in Figure 1 and Table 1, molecular biology methods were used to construct vector plasmids of different dCas9-FokI fusion proteins.
[0193] In this example, the amino acid sequence of dCas9 was determined based on SpCas9 (SEQ ID NO: 1), as shown in SEQ ID NO: 2.
[0194] Table 1. Different designed dCas9-FokI fusion proteins
[0195] * indicates that the FokI domain with linker (SEQ ID NO: 4) was substituted at the indicated position.
[0196] The amino acid sequence from position 6 to position 201 of the sequence shown is the FokI domain (SEQ ID NO: 52).
[0197] **Domain division of SpCas9 (SEQ ID NO: 1) refers to the article Zhu, X., et al. Cryo-EM structures reveal coordinated domain motions that govern DNA cleavage by Cas9. Nat Struct Mol Biol 26, 679–685 (2019). https: / / doi.org / 10.1038 / s41594-019-0258-2.
[0198] Table 2. List of primers corresponding to different dCas9-FokI fusion proteins
[0199] The expression vector pCDNA3.1(+) (Miaoling Biotechnology, P0157) was double-digested with Acc65I and EcoRI, and the linearized vector was recovered by agarose gel electrophoresis.
[0200] Using FokI-dCas9 (Miaoling Bio, P5109) as a template, PCR amplification was performed using the FokI-PF1 and FokI-PR1 primer pairs in Table 2 to obtain the FokI coding sequence (SEQ ID NO: 5), which encodes the FokI domain with linker.
[0201] Using pCDNA3.1-dCas9 (Miaoling Biotechnology, P15213) as a template, PCR amplification was performed using the amplification primer combinations of different fusion protein coding sequence fragments in Table 3 (primer sequences are shown in Table 2) to obtain different fragments.
[0202] Table 3. Primer combinations for amplification of different dCas9-FokI fusion protein coding sequence fragments
[0203] Different fusion protein coding sequences, FokI coding sequences and linearized vectors were cloned using Gibson Homologous recombination was performed using Master Mix (NEB) to integrate the fusion protein coding sequence into the cloning region of the vector pCDNA3.1(+) to construct recombinant vectors encoding the different structural fusion proteins (dCas9-FokI-S1 to dCas9-FokI-S10) listed in Table 1. For example, the dCas9-FokI-S1 vector was constructed by homologous recombination of the linearized pCDNA3.1(+) vector + fragment S1-F2 + FokI coding sequence + fragment S1-F1.
[0204] The reaction solution was transformed into Stbl3 competent cells, coated on ampicillin-resistant LB plates, and cultured overnight at 37°C. Clones were picked and sequenced to obtain plasmid vectors of fusion proteins dCas9-FokI-S1 to dCas9-FokI-S10, respectively.
[0205] The amino acid sequence of the dCas9-FokI-S9 fusion protein is:
[0206] The amino acid sequences of dCas9-FokI-S1 to dCas9-FokI-S8, and dCas9-FokI-S10 are simply FokI domains with linkers substituted into different positions shown in Table 1, and thus can be derived by analogy.
[0207] Example 2: Construction of pCDH-CMV-EGFP-Reporter3-EF1a-Puro cell line
[0208] a. Construction of the lentiviral expression plasmid pCDH-CMV-EGFP-Reporter3-EF1a-Puro for the GFP reporter system
[0209] The EGFP coding sequence of the detection system was synthesized and the synthesized fragment was digested with XbaI+NotI and then ligated with the vector pCDH-CMV-MCS-EF1-Puro (Ubao Bio, VT1480) digested with XbaI+NotI using T4 DNA ligase (Thermo Scientific TM , EL0012) were ligated and transformed into Stbl3, and plates containing ampicillin were coated. After overnight culture at 37°C, clones were picked and sequenced to obtain the pCDH-CMV-EGFP-Reporter3-EF1a-Puro vector (SEQ ID NO: 30).
[0210] The synthetic EGFP coding sequence is as follows:
[0211] The underlined parts are the targeting sites of the two gRNAs.
[0212] b. Lentiviral packaging
[0213] The sequence-identified pCDH-CMV-EGFP-Reporter3-EF1a-Puro plasmid was mixed with the viral packaging helper plasmids pMD2.G (Miaoling Biotechnology, P0262) and psPAX2 (Miaoling Biotechnology, P0261) at a molar ratio of 1:1:1 and transfected into 293T cells using polyethylene glycol (PEI). After 48 hours of transfection, the culture supernatant was collected and filtered through a 0.45 μm filter to obtain the crude pCDH-CMV-EGFP-Reporter3-EF1a-Puro virus.
[0214] c. Virus infection of 293T cells to construct a test cell line
[0215] 293T cells were infected with 1 / 4 volume of crude pCDH-CMV-EGFP-Reporter3-EF1a-Puro virus in the culture medium. After 48 hours of infection, the medium was changed and 2 μg / ml of Puromycin was added for selection. The selected cells were screened for single clones using limiting dilution. The selected single clones were used as the cell line for testing.
[0216] Example 3: Editing efficiency test of dCas9-FokI fusion proteins with different structures
[0217] Because FokI can only cleave single-stranded DNA, the inventors used dual gRNAs to direct the dCas9-FokI fusion protein to cleave double-stranded DNA. The cleaved DNA undergoes NHEJ repair within the cell, causing the GFP gene in some cells to produce a normal reading frame (as shown in Figure 2), thereby emitting fluorescence. Flow cytometry was used to measure the proportion of cells emitting green fluorescence to characterize the editing efficiency of different dCas9-FokI fusion proteins.
[0218] gRNA vector construction
[0219] The guide sequence coding sequence was introduced by enzyme digestion-ligase ligation to obtain two gRNA expression vectors, gRNA1-SpCas9-pUC57Kan (SEQ ID NO: 34) and gRNA2-SpCas9-pUC57Kan (SEQ ID NO: 35), which can be driven by the U6 promoter to express gRNA1 or gRNA2. The gRNA contains a guide sequence targeting EGFP and a gRNA backbone sequence corresponding to SpCas9.
[0220] Take 250ng of each of the dCas9-FokI-S1 to S9 recombinant plasmid vectors from Example 1, mix them with 125ng of the gRNA1-SpCas9-pUC57Kan plasmid and 125ng of the gRNA2-SpCas9-pUC57Kan plasmid, respectively, and 100μl of opti-MEM (Gibco). Then add 1.5μl of PEI (YEASEN) and mix evenly. Let it stand for 20min. Then, add it to the test cell line constructed in Example 2 for cell transfection. After 72h of transfection, the cells were collected and analyzed by flow cytometry.
[0221] The negative control (NC) group was the test cell line of Example 2, which was not transfected with the CRISPR system-related plasmid.
[0222] The results are shown in Figure 3 and Table 4.
[0223] Table 4. Editing efficiency
[0224] Flow cytometry results showed that the editing efficiency of dCas9-FokI-S9 was much higher than that of dCas9-FokI-S1 to dCas9-FokI-S8, and dCas9-FokI-S10.
[0225] Example 4: Construction of dCas9-FokI-S9 control plasmid and PAMless plasmid
[0226] A similar method to Example 1 was used to construct the dCas9-FokI-953-01 plasmid (SEQ ID NO: 36) and the dCas9-FOKI-953-02 plasmid (SEQ ID NO: 37), both of which were used to express a fusion protein fused to FokI within dSpCas9. The difference between the two is that the recombinant protein expressed by the dCas9-FOKI-953-01 plasmid is directly inserted into the FokI domain between 953aa-954aa; the recombinant protein expressed by the dCas9-FOKI-953-02 plasmid is inserted into the FokI domain between 953aa-954aa and the FokI domain ends with a linker sequence for connection to the Cas9 sequence.
[0227] A dCas9PAMless-FokI-S9 plasmid (SEQ ID NO: 38) was constructed using a method similar to that of Example 1 for expressing a recombinant protein with the FokI domain inserted into the dead PAMless Cas9 (SpRY variant).
[0228] Example 5: Editing reporter systems with different dCas9-FokI fusion proteins
[0229] Cas9 was fused to FokI to generate different fusion proteins. Following the protocol described in Example 3, the editing efficacy was initially verified using a reporter system. Combining different fusion proteins with gRNAs edited the EGFP gene, generating indels and restoring fluorescent EGFP expression. The editing efficacy of the fusion proteins was assessed by measuring the eGFP ratio.
[0230] The dCas9-FokI positive control group was transfected with 250 ng of dCas9-FokI plasmid (Miaoling Bio P5109, expressing dCas9 with N-terminal fusion FokI), and 125 ng each of gRNA1 plasmid and gRNA2 plasmid.
[0231] The Cas9 nickase positive control group was transfected with 250 ng of Cas9 nickase plasmid (SEQ ID NO: 39, expressing Cas9 nickase) and 125 ng each of gRNA1 plasmid and gRNA2 plasmid.
[0232] The dCas9-FokI-S9 group, dCas9PAMless-FokI-S9 experimental group, dCas9-FokI-953-01 group, and dCas9-FokI-953-02 group were transfected with 250 ng of the corresponding recombinant protein plasmid, and 125 ng each of the gRNA1 plasmid and gRNA2 plasmid.
[0233] The SpCas9-gRNA1+2 control group was transfected with 250 ng of pX459 V2.0 plasmid (expressing wild-type SpCas9), and 125 ng each of gRNA1 plasmid and gRNA2 plasmid.
[0234] The SpCas9-gRNA1 control group was transfected with 250 ng of pX459 V2.0 plasmid (expressing wild-type SpCas9) and 250 ng of gRNA1 plasmid.
[0235] The SpCas9-gRNA2 control group was transfected with 250 ng of pX459 V2.0 plasmid (expressing wild-type SpCas9) and 250 ng of gRNA2 plasmid.
[0236] The negative control (NC) group was the test cell line of Example 2, which was not transfected with the CRISPR system-related plasmid.
[0237] Three transfection reagents, PEI, EndoFectin and Lipo2000, were used for testing.
[0238] The specific results are shown in Table 5 and Figure 4. It can be seen that the editing efficiency of the dCas9-FokI-S9 fusion protein for this reporter system is comparable to that of Cas9 nickase, even better than that of SpCas9, much higher than that of the fusion protein with FokI inserted or replaced at the 953aa position, and better than the fusion protein with the FokI domain fused to the end of dCas9.
[0239] Table 5. Editing efficiency of different recombinant proteins
[0240] Example 6: dCas9-FokI-S9 endogenous gene editing efficiency test
[0241] Plasmids expressing gRNA targeting the HBB gene and VEGFA gene (containing the gRNA backbone sequence corresponding to SpCas9, i.e., SEQ ID NO: 40) were constructed respectively, and the U6 promoter drove the expression.
[0242] The guide sequences of gRNAs are shown in Table 6.
[0243] Table 6. Endogenous gene editing efficiency detection
[0244] Take 250 ng of the recombinant protein plasmid of the previous example and 250 ng of the gRNA plasmid of this example (250 ng when one gRNA plasmid is used, and about 125 ng each when two gRNA plasmids are used), mix them, add 1.5 μl PEI, and transfect 293T cells.
[0245] 72 hours after transfection, cells were washed once with PBS, digested with trypsin, harvested by centrifugation at 300 rcf for 5 minutes, and resuspended in 500 μl of PBS. A 50 μl aliquot of the resuspended cells was treated with 48 μl of DirectPCR Lysis Reagent (Cell) and 2 μl of Proteinase K. 1 μl of the treated lysate was used for PCR amplification. Amplification primer sequences are shown in Table 7.
[0246] Table 7. Primers for amplifying target sequences HBB and VEGFA
[0247] The amplified products were subjected to Sanger sequencing, and the sequencing results were subjected to TIDE analysis to confirm the editing efficiency. The results are shown in Table 8 and Figure 5.
[0248] Table 8. Endogenous gene editing efficiency
[0249] TIDE analysis results showed that dCas9-FokI-S9 had a good editing effect on endogenous genes, with editing efficiency comparable to SpCas9 (slightly better in some cases), significantly superior to dCas9-FokI with FokI fused to the terminal, dCas9-FokI-953-01 and dCas9-FokI-953-02 with FokI inserted or replaced at position 953aa, and PAMless Cas9, and superior to Cas9 nickase.
[0250] Although specific embodiments of the present disclosure have been described above, those skilled in the art will appreciate that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present disclosure. Therefore, the scope of protection of the present disclosure is defined by the appended claims.
Claims
1. A fusion protein, characterized in that The fusion protein is obtained by replacing the continuous amino acid residues of the random coil structure of dCas9 with the functional domain, and the dCas9 is a Cas9 protein without nuclease activity; Optionally, the functional domain is selected from a deaminase domain (e.g., a cytosine deaminase domain or an adenine deaminase domain), a UGI domain, a UDG domain, a methylation domain, a demethylation domain, a subcellular localization signal (e.g., a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, or a chloroplast localization signal), a transcription activation domain, a transcription repression domain, a nuclease domain (e.g., a FokI domain), a histone deacetylation domain, and a DNA ligation domain; further, the functional domain is a nuclease domain; further, the functional domain is a FokI domain; Optionally, the random coil structure is located in the REC1-A and / or REC1-B domain of dCas9; Optionally, the random coil structure is located in RuvC, loop L2 domain and / or HNH domain of dCas9; or the random coil structure is located in the RuvC domain of dCas9; or the random coil structure is located in the RuvC-Ⅲ domain of dCas9.
2. A fusion protein, characterized in that: The fusion protein includes dCas9 and a functional domain, wherein the continuous amino acid residues in the random coil in the dCas9 are replaced by the functional domain; the dCas9 is a Cas9 protein without nuclease activity; Optionally, the functional domain is selected from a deaminase domain (e.g., a cytosine deaminase domain or an adenine deaminase domain), a UGI domain, a UDG domain, a methylation domain, a demethylation domain, a subcellular localization signal (e.g., a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, or a chloroplast localization signal), a transcription activation domain, a transcription repression domain, a nuclease domain (e.g., a FokI domain), a histone deacetylation domain, and a DNA ligation domain; further, the functional domain is a nuclease domain; further, the functional domain is a FokI domain; Optionally, the random coil structure is located in the REC1-A and / or REC1-B domain of dCas9; Optionally, the random coil structure is located in RuvC, loop L2 domain and / or HNH domain of dCas9; or the random coil structure is located in the RuvC domain of dCas9; or the random coil structure is located in the RuvC-Ⅲ domain of dCas9.
3. The fusion protein according to claim 1 or 2, characterized in that The fusion protein is capable of specifically binding to a target nucleic acid under the guidance of a guide polynucleotide; Optionally, the fusion protein is capable of sequence-specific binding to a target nucleic acid under the guidance of a guide polynucleotide, the guide polynucleotide comprising a guide sequence, the guide sequence hybridizing to the target nucleic acid; Alternatively, the fusion protein is capable of forming a complex with a guide polynucleotide comprising a guide sequence engineered to direct sequence-specific binding of the complex to a target nucleic acid.
4. The fusion protein according to any one of claims 1 to 3, characterized in that The functional domain is a FokI domain, and the fusion protein is capable of recognizing a double-stranded target nucleic acid and nicking the target nucleic acid under the guidance of a guide polynucleotide; or the fusion protein is capable of forming a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the complex sequence to specifically bind to a double-stranded target nucleic acid and nick the target nucleic acid.
5. The fusion protein according to any one of claims 1 to 4, characterized in that The guide polynucleotide also comprises a backbone sequence; further, the backbone sequence interacts with the dCas9 portion of the fusion protein, or the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein.
6. The fusion protein according to any one of claims 1 to 5, characterized in that The Cas9 protein is selected from SpCas9, SaCas9, NmeCas9, StCas9 and CjCas9, preferably SpCas9 (Streptococcus pyogenes Cas9) or SaCas9 (Staphylococcus aureus Cas9); Optionally, the amino acid sequence of the SpCas9 is as shown in SEQ ID NO: 1; Optionally, the amino acid sequence of the Cas9 protein has ≥80%, ≥85%, ≥90%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99% or ≥99.5% sequence identity with the amino acid sequence shown in SEQ ID NO: 1; Optionally, the amino acid sequence of the dCas9 is as shown in SEQ ID NO:2 or has a sequence identity of ≥80%, ≥85%, ≥90%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99% or ≥99.5% to the amino acid sequence shown in SEQ ID NO:
2.
7. The fusion protein according to any one of claims 1 to 6, characterized in that The dCas9 is a Cas9 protein in which a mutation, such as a point mutation, occurs in the RuvC domain and / or the HNH domain, so that the Cas9 protein has no nuclease activity; Optionally, the dCas9 is a Cas9 in which the positions corresponding to D10 and N863 of SpCas9 are mutated to A; or the dCas9 is a Cas9 in which the positions corresponding to D10 and H840 of SpCas9 are mutated to A; the corresponding positions can be determined by sequence alignment of the Cas9 and SpCas9; Optionally, the dCas9 is SpCas9 with D10A and N863A mutations, or SpCas9 with D10A and H840A mutations; Optionally, the dCas9 is SaCas9 with D10A and H557A mutations.
8. The fusion protein according to any one of claims 1 to 7, characterized in that The functional domain is fused to dCas9 directly or through a linker; Optionally, the functional domain is directly fused to dCas9, or the functional domain is fused to dCas9 via an amino acid sequence; Optionally, the functional domain is fused to dCas9 via a linker sequence having a length of 1-50, 1-40, 1-30, 1-20, 5-16, 1-10, 10-20, 3-8 or 15-20 amino acid residues; Preferably, the linker is an amino acid residue sequence as shown in SEQ ID NO: 50 or 51.
9. The fusion protein according to any one of claims 1 to 8, characterized in that The fusion protein is obtained by replacing 1-20 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain; Optionally, the fusion protein is obtained by replacing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain; Optionally, the fusion protein is obtained by replacing 1-15, 3-11, 3-10 or 9-11 consecutive amino acid residues of the random coil structure of dCas9 with a functional domain; Optionally, the length of the replaced random coil structure is 1-15, 3-11, 3-10, 9-11 or 3 amino acid residues; Optionally, the random coil structure corresponds to positions 37-41, 107-113, 147-149, 169-172, 308-314, 792-799, 864-871, 906-908, 943-952 and / or 1015-1017 of the sequence shown in SEQ ID NO:
2.
10. The fusion protein according to any one of claims 1 to 9, characterized in that When the amino acid sequence of the dCas9 is as shown in SEQ ID NO: 2, the substitution occurs at positions 37-41, 107-113, 147-149, 169-172, 308-314, 792-799, 864-871, 906-908, 943-952, or 1015-1017; Optionally, the functional domain is fused to dCas9 directly or through a linker after replacing the replaced sequence of the dCas9; Optionally, the length of the replaced sequence is 1-15, 3-11, 3-10, 9-11 or 3 amino acid residues; Optionally: When the replaced sequence is located at position 37-41 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 5 amino acid residues; When the replaced sequence is located at positions 107-113 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 7 amino acid residues; When the replaced sequence is located at positions 147-149 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues; When the replaced sequence is located at positions 169-172 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 4 amino acid residues; When the replaced sequence is located at positions 308-314 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 7 amino acid residues; When the replaced sequence is located at positions 792-799 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 8 amino acid residues; When the replaced sequence is located at positions 864-871 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 8 amino acid residues; When the replaced sequence is located at positions 906-908 of the amino acid sequence as shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues; When the replaced sequence is located at positions 943-952 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 10 amino acid residues; When the replaced sequence is located at positions 1015-1017 of the amino acid sequence shown in SEQ ID NO: 2, the length of the replaced sequence is 3 amino acid residues.
11. The fusion protein according to any one of claims 1 to 10, characterized in that The FokI domain includes the DNA binding domain and the cleavage domain of the FokI endonuclease; Optionally, the FokI domain comprises or is the DNA cleavage domain of a FokI endonuclease; Optionally, the FokI domain may exist in a monomeric or dimer form; for example, the FokI domain exists in a monomeric form in solution; Optionally, the FokI domain has nuclease activity.
12. The fusion protein according to any one of claims 1 to 11, characterized in that The first fusion protein can bind to one single strand of the double-stranded target nucleic acid together with the first guide polynucleotide, and the second fusion protein can bind to the other single strand of the target nucleic acid together with the second guide polynucleotide; The FokI domains of the first fusion protein and the second fusion protein have a spacer sequence on the double-stranded target nucleic acid and can form a dimer, wherein the length of the spacer sequence is 1-100 bp, and the dimer can cut the double-stranded target nucleic acid when the first fusion protein, the first guiding polynucleotide, the second fusion protein and the second guiding polynucleotide bind to the target nucleic acid; Optionally, the FokI domain can dimerize after dCas9 binds to the double-stranded target DNA to achieve cutting of the double-stranded target DNA through its nuclease activity; when the FokI domain dimerizes, the two FokI domains have a spacer sequence on the double-stranded target DNA to cut the double-stranded target DNA; Preferably, the length of the spacer sequence is 1-100bp, 1-90bp, 1-80bp, 1-70bp, 1-60bp, 1-50bp, 1-40bp, 1-30bp, 1-20bp, 1-10bp, 1-5bp, 2-100bp, 2-90bp, 2-80bp, 2-70bp, 2-60bp, 2-50bp, 2-40bp, 2-30bp, 2-20bp , 2-10bp, 2-5bp, 3-100bp, 3-90bp, 3-80bp, 3-70bp, 3-60bp, 3-50bp, 3-40bp, 3-30bp, 3-20bp , 3-10bp, 3-5bp, 4-100bp, 4-90bp, 4-80bp, 4-70bp, 4-60bp, 4-50bp, 4-40bp, 4-30bp, 4-20bp, 4-10bp, 4-5bp, 5-100bp, 5-90bp, 5-80bp, 5-70bp, 5-60bp, 5-50bp, 5-40bp, 5-30bp, 5-20bp, 5-10bp, 6-20bp, 10-100bp, 10-90bp, 10-80bp, 10-70bp, 10-60bp, 10-50bp, 10-40bp, 10-30bp , 10-20bp, 15-100bp, 15-90bp, 15-80bp, 15-70bp, 15-60bp, 15-50bp, 15-40bp, 15-30bp, 15-20bp, 20-100bp, 20-90bp, 20-80bp, 20-70bp, 20-60bp, 20-50bp, 20-40bp, 20-30bp or 20-25bp.
13. The fusion protein according to any one of claims 1 to 12, characterized in that The amino acid sequence of the FokI domain has a sequence identity of ≥50%, ≥60%, ≥70%, ≥80%, ≥90%, ≥95%, ≥98% or ≥99% to the amino acid residue sequence of positions 6 to 201 of the sequence shown in SEQ ID NO:4; Optionally, the FokI domain comprises or consists of the amino acid residue sequence from position 6 to position 201 of the sequence shown in SEQ ID NO:4; Optionally, the amino acid sequence of the FokI domain is as shown in SEQ ID NO:52 or has a sequence identity of ≥50%, ≥60%, ≥70%, ≥80%, ≥90%, ≥95%, ≥98% or ≥99% to the amino acid sequence as shown in SEQ ID NO:
52.
14. The fusion protein according to any one of claims 1 to 13, characterized in that The guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide; Optionally, the guide polynucleotide is formed by the chimera of crRNA and tracrRNA in the CRISPR-Cas9 system to form a single-molecule guide polynucleotide (sgRNA); optionally, the guide polynucleotide is composed of crRNA and tracrRNA in the CRISPR-Cas9 system, that is, the double molecule.
15. The fusion protein according to any one of claims 1 to 14, characterized in that The fusion protein further comprises any one or more of the following: a subcellular localization signal, a DNA binding domain, a transcription activation domain, a transcription repression domain, a nuclease domain, a deaminase domain, a UDG domain, a UGI domain, a methylase, a demethylase, a histone deacetylase, a DNA ligase, an epitope tag, and a reporter protein; Optionally, the subcellular localization signal is selected from: a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, and a chloroplast localization signal.
16. The fusion protein according to any one of claims 1 to 15, characterized in that The fusion protein is a fusion protein obtained by replacing the amino acid residues 943-952 of the nuclease activity inactivation mutant of the Cas9 protein with the amino acid sequence as shown in SEQ ID NO: 1 with the functional domain; Optionally, the inactivation mutant has D10A and N863A mutations, or D10A and H840A mutations relative to SEQ ID NO: 1; Optionally, the functional domain is a FokI domain; further, the FokI domain is a DNA cleavage domain of a FokI endonuclease; further, the amino acid sequence of the FokI domain is as shown in SEQ ID NO: 52; Optionally, the functional domain is connected to the N-terminus and C-terminus of the Cas9 inactivation mutant respectively through a linker sequence; further, the functional domain is connected to the N-terminus and C-terminus of the Cas9 inactivation mutant respectively through a linker with an amino acid sequence as shown in SEQ ID NO: 50 and SEQ ID NO: 51; Optionally, at least one nuclear localization sequence is fused to the N-terminus and / or C-terminus of the fusion protein; Optionally, the fusion protein is capable of forming a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the complex sequence to specifically bind to a double-stranded target DNA and nick the target DNA; further, the guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide; further, the guide polynucleotide is a chimera of crRNA and tracrRNA in the CRISPR-Cas9 system to form a single-molecule guide polynucleotide or the guide polynucleotide is a group consisting of two separate molecules of crRNA and tracrRNA in the CRISPR-Cas9 system, and the crRNA consists of a guide sequence and a direct repeat sequence; Optionally, the fusion protein comprises or consists of the amino acid sequence shown in SEQ ID NO:
3.
17. A ribonucleoprotein complex (RNP complex), characterized in that The ribonucleoprotein complex (RNP complex) comprises a guide polynucleotide and a fusion protein according to any one of claims 1 to 16; Optionally, the guide polynucleotide guides the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid; Optionally, the guide polynucleotide comprises a guide sequence engineered to direct the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid; Optionally, the guide polynucleotide further comprises a backbone sequence; the backbone sequence interacts with the dCas9 portion of the fusion protein, or the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein; Optionally, the guide polynucleotide is a single-molecule guide polynucleotide or a double-molecule guide polynucleotide.
18. A gene editing system, characterized in that: The gene editing system comprises: (1) the fusion protein according to any one of claims 1 to 16, or a nucleic acid encoding the fusion protein; and / or, (2) a guide polynucleotide, or a nucleic acid encoding the same; Optionally, the guide polynucleotide guides the fusion protein sequence to specifically bind to a target nucleic acid; Optionally, the guide polynucleotide comprises a guide sequence engineered to direct the ribonucleoprotein complex to sequence-specifically bind to a target nucleic acid; Optionally, the guide polynucleotide further comprises a backbone sequence; the backbone sequence enables the guide polynucleotide to form a complex with the fusion protein.
19. The gene editing system according to claim 18, wherein: The gene editing system comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different fusion proteins; optionally, the gene editing system comprises 1 fusion protein; Optionally, the gene editing system comprises at least 2 different guide polynucleotides; Optionally, the gene editing system comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different guide polynucleotides; the gene editing system preferably comprises 2 different guide polynucleotides; Optionally, the different guide polynucleotides have different guide sequences and the same backbone sequences; or, the different guide polynucleotides have different guide sequences and the same backbone sequences.
20. A polynucleotide, characterized in that The polynucleotide encodes the fusion protein according to any one of claims 1 to 16 or the ribonucleoprotein complex according to claim 17; Optionally, in the polynucleotide encoding the ribonucleoprotein complex, the polynucleotide encoding the guide polynucleotide and the polynucleotide encoding the fusion protein are on the same or different nucleic acid chains.
21. An expression cassette, characterized in that The expression cassette comprises a promoter and the polynucleotide of claim 20; the promoter sequence regulates the expression of the coding sequence.
22. A recombinant expression vector, characterized in that: The recombinant expression vector comprises the polynucleotide according to claim 20, and optionally a cis-acting element and / or a trans-acting factor encoding gene; Optionally, the cis-acting element is selected from a promoter, an enhancer and a silencer; Optionally, the trans-acting factor encoding gene is selected from nucleotides encoding polymerases, transcription factors and / or transcription regulatory factors.
23. A composition, characterized in that The composition comprises a delivery vector, and a fusion protein as described in any one of claims 1 to 16, a ribonucleoprotein complex as described in claim 17, a gene editing system as described in claim 18 or 19, a polynucleotide as described in claim 20, an expression cassette as described in claim 21 and / or a recombinant expression vector as described in claim 22.
24. A recombinant cell, characterized in that The recombinant cell comprises the fusion protein of any one of claims 1 to 16, the ribonucleoprotein complex of claim 17, the gene editing system of claim 18 or 19, the polynucleotide of claim 20, the expression cassette of claim 21, the recombinant expression vector of claim 22 and / or the composition of claim 23.
25. A pharmaceutical composition, characterized in that The pharmaceutical composition comprises a fusion protein as described in any one of claims 1 to 16, a ribonucleoprotein complex as described in claim 17, a gene editing system as described in claim 18 or 19, a polynucleotide as described in claim 20, an expression cassette as described in claim 21, a recombinant expression vector as described in claim 22, a composition as described in claim 23 and / or a recombinant cell as described in claim 24, and optionally a pharmaceutically acceptable carrier and / or excipient.
26. A kit, characterized in that The kit comprises the fusion protein according to any one of claims 1 to 16, the ribonucleoprotein complex according to claim 17, the gene editing system according to claim 18 or 19, the polynucleotide according to claim 20, the expression cassette according to claim 21, the recombinant expression vector according to claim 22, the composition according to claim 23, the recombinant cell according to claim 24 and / or the pharmaceutical composition according to claim 25.
27. Use of a fusion protein as described in any one of claims 1 to 16, a ribonucleoprotein complex as described in claim 17, a gene editing system as described in claim 18 or 19, a polynucleotide as described in claim 20, an expression cassette as described in claim 21, a recombinant expression vector as described in claim 22, a composition as described in claim 23, a recombinant cell as described in claim 24, a pharmaceutical composition as described in claim 25, and / or a kit as described in claim 26 in the preparation of a drug for diagnosing, preventing and / or treating a disease or condition associated with a target nucleic acid.
28. A method for diagnosing, preventing and / or treating a disease or condition associated with a target nucleic acid, characterized in that: The method comprises administering to a patient in need thereof a fusion protein as described in any one of claims 1 to 16, a ribonucleoprotein complex as described in claim 17, a gene editing system as described in claim 18 or 19, a polynucleotide as described in claim 20, an expression cassette as described in claim 21, a recombinant expression vector as described in claim 22, a composition as described in claim 23, a recombinant cell as described in claim 24, a pharmaceutical composition as described in claim 25, and / or a kit as described in claim 26.
29. A fusion protein as described in any one of claims 1-16, a ribonucleoprotein complex as described in claim 17, a gene editing system as described in claim 18 or 19, a polynucleotide as described in claim 20, an expression cassette as described in claim 21, a recombinant expression vector as described in claim 22, a composition as described in claim 23, a recombinant cell as described in claim 24, a pharmaceutical composition as described in claim 25, and / or a kit as described in claim 26, wherein the fusion protein, ribonucleoprotein complex, gene editing system, polynucleotide, expression cassette, recombinant expression vector, composition, recombinant cell, pharmaceutical composition or kit is used to diagnose, prevent and / or treat a disease or condition associated with a target nucleic acid.
30. A method for editing a target nucleic acid in vitro, ex vivo or in vivo, characterized in that: The method comprises the step of contacting the fusion protein according to any one of claims 1 to 16, the ribonucleoprotein complex according to claim 17, the gene editing system according to claim 18 or 19, the polynucleotide according to claim 20, the expression cassette according to claim 21, the recombinant expression vector according to claim 22, the composition according to claim 23, the recombinant cell according to claim 24, the pharmaceutical composition according to claim 25, and / or the kit according to claim 26 with the target nucleic acid, so that the sequence of the target nucleic acid changes or the expression level of the target nucleic acid changes.