Connexon and application thereof
By developing new linkers with specific amino acid sequences or homology polypeptides, the problem of insufficient linkers in the construction of fusion proteins in the prior art is solved, and efficient protein folding and biological activity are achieved.
Patent Information
- Application Number
- CN202311682917.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art lacks effective linkers when constructing fusion proteins, resulting in misfolding of proteins, low yields or impaired biological activity.
A new linking peptide has been developed to ensure correct folding and high biological activity of the fusion protein by providing linkers with specific amino acid sequences or homology polypeptides.
By using these new linkers, the yield and biological activity of the fusion protein is significantly improved, meeting the needs of building efficient base editors.
Smart Images

Figure BDA0004596785320000141 
Figure BDA0004596785320000151 
Figure BDA0004596785320000231
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biology, and in particular, to a linker and its application. Background Art
[0002] When constructing fusion proteins, it is often necessary to select a suitable linker to connect polypeptide domains. Direct fusion of functional domains (without using a linker) may cause many problems, such as misfolding of the fusion protein, low protein yield, or impaired biological activity.
[0003] The activity of the fusion protein may be related to the linker. For example, when constructing base editor fusion proteins, the fusion of nuclease and deaminase domains is mostly through a linker, and the choice of linker is related to the editing efficiency of the base editor.
[0004] Therefore, there is an urgent need in the art to develop new linkers to meet the requirements for constructing fusion proteins. Summary of the Invention
[0005] The purpose of the present invention is to develop new linkers to meet the requirements for constructing fusion proteins.
[0006] In the first aspect of the present invention, a connecting peptide is provided, and the connecting peptide is selected from the following group:
[0007] (a) a polypeptide having an amino acid sequence shown in any one of SEQ ID NO: 1-36;
[0008] (b) a polypeptide having a homology (or identity) of ≥30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with the amino acid sequence shown in any one of SEQ ID NO: 1-36, and the polypeptide has the biological function of SEQ ID NO: 1-36;
[0009] (c) a derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acid residues in the amino acid sequence shown in any one of SEQ ID NO: 1-36, and retaining the biological function of SEQ ID NO: 1-36.
[0010] In another preferred embodiment, the connecting peptide is synthetic.
[0011] In another preferred example, compared with the amino acid sequence shown in any one of SEQ ID NO:1-36, the linking peptide has a sequence with substitution, deletion or addition of one or more bases, and substantially retains the biological function of the sequence from which it is derived.
[0012] In another preferred example, the linking peptide has a polypeptide with a homology (or identity) of ≥30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with the amino acid sequence shown in any one of SEQ ID NO:2, 18, 21-25, 29-30, and the polypeptide has the biological function of SEQ ID NO:2, 18, 21-25, 29-30.
[0013] In another preferred example, the linking peptide has a polypeptide with a homology (or identity) of ≥30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with the amino acid sequence shown in any one of SEQ ID NO:3-11, 16, 19, 26, and the polypeptide has the biological function of SEQ ID NO:3-11, 16, 19, 26.
[0014] In another preferred example, the linking peptide has a polypeptide with a homology (or identity) of ≥40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with the amino acid sequence shown in any one of SEQ ID NO:17, 31-33, and the polypeptide has the biological function of SEQ ID NO:17, 31-33.
[0015] In another preferred example, the linker peptide has a polypeptide with a homology (or identity) of ≥40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with any of the amino acid sequences shown in SEQ ID NO: 27, 34, 36, and the polypeptide has the biological functions of SEQ ID NO: 15, 20, 27, 34, 36.
[0016] In another preferred example, the linker peptide has a polypeptide with a homology (or identity) of ≥70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with any of the amino acid sequences shown in SEQ ID NO: 12 - 14, and the polypeptide has the biological functions of SEQ ID NO: 12 - 14.
[0017] In another preferred example, the linker peptide is selected from polypeptides having the amino acid sequences shown in SEQ ID NO: 2, 18, 21 - 25, 29 - 30.
[0018] In another preferred example, the linker peptide is selected from polypeptides having the amino acid sequences shown in SEQ ID NO: 3 - 11, 16, 19, 26.
[0019] In another preferred example, the linker peptide is selected from polypeptides having the amino acid sequences shown in SEQ ID NO: 17, 31 - 33.
[0020] In another preferred example, the linker peptide is selected from polypeptides having the amino acid sequences shown in SEQ ID NO: 27, 34, 36.
[0021] In another preferred example, the linker peptide is selected from polypeptides having the amino acid sequences shown in SEQ ID NO: 12 - 14.
[0022] In the second aspect of the present invention, a fusion protein is provided. The fusion protein includes the linker peptide described in the first aspect of the present invention and one or more functional domains, or includes the linker peptide described in the first aspect of the present invention and at least one nuclease.
[0023] In another preferred example, the functional domain is selected from the group consisting of: a Cas protein targeting moiety, a DNA binding domain, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a lytic activity polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, and a base excision repair inhibitor (such as uracil-DNA glycosylase inhibitor (UGI)), or a combination thereof.
[0024] In one embodiment, the fusion protein comprises a nuclease, the linking peptide described in the first aspect of the present invention, and one or more deamination domains.
[0025] In some specific embodiments, the deamination domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.
[0026] In some specific embodiments, the adenosine deaminase catalytic domain is adenosine deaminase.
[0027] In some specific embodiments, the adenine deaminase (or adenosine deaminase) is any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).
[0028] In some specific embodiments, adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively.
[0029] In some specific embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in DNA.
[0030] In some specific embodiments, the adenine deaminase can be any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).
[0031] In some specific embodiments, the adenosine deaminase has an amino acid sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100% sequence identity to any one of SEQ ID NOs: 96 - 98.
[0032] In some specific embodiments, the deamination domain is a cytidine deaminase catalytic domain.
[0033] In some specific embodiments, the cytidine deaminase catalytic domain comprises an APOBEC deaminase.
[0034] In some specific embodiments, the cytidine deaminase can be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570), and / or variant versions thereof.
[0035] In some specific embodiments, the cytidine deaminase is APOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 99.
[0036] In some specific embodiments, the cytidine deaminase can be APOBEC3A deaminase having the amino acid sequence of SEQ ID NO: 100.
[0037] In some specific embodiments, the cytidine deaminase can have about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identity to the amino acid sequence of a naturally occurring cytidine deaminase.
[0038] In some specific embodiments, the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 100% identity to the amino acid sequence shown in SEQ ID NO: 99 or SEQ ID NO: 100.
[0039] In some embodiments, the polynucleotide encoding the cytidine deaminase can be codon-optimized for expression in an organism, and the codon-optimized polypeptide can have about 70% to 99.5% identity to a reference polynucleotide.
[0040] In another preferred embodiment, the nuclease is selected from nucleic acid programmable nucleotide-binding proteins (napDNAbp) or homologs thereof.
[0041] In another preferred embodiment, the nuclease is selected from RNA-guided nucleic acid programmable nucleotide-binding proteins, homing endonucleases (Meganuclease), zinc finger fusion proteins (ZFN), or TALEN, or homologs thereof.
[0042] In another preferred embodiment, the nuclease is selected from type II CRISPR-Cas polypeptides, type I CRISPR-Cas polypeptides, type III CRISPR-Cas polypeptides, type IV CRISPR-Cas polypeptides, type V CRISPR-Cas polypeptides, type VI CRISPR-Cas polypeptides, type VII CRISPR-Cas polypeptides, IscB polypeptides, TnpB polypeptides, IsrB polypeptides, or homologs thereof.
[0043] In another preferred embodiment, the nuclease is selected from Cas9, CasX, CasY, Cas12a (Cpf1), Cas 12b (C2cl), Cas13a (C2c2), Cas12c (C2c3), Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, Argonaute (Ago), or homologs thereof.
[0044] In another preferred embodiment, the nuclease has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with the amino acids shown in any one of SEQ ID NOs: 92-95.
[0045] In another preferred embodiment, the fusion protein comprises the following elements fused together:
[0046] (a) A first polypeptide domain;
[0047] (b) The linker peptide described in the first aspect of the present invention;
[0048] (c) A second polypeptide domain.
[0049] In some specific embodiments, the first polypeptide domain is linked to the linker peptide described in the first aspect of the present invention at its C-terminus.
[0050] In some specific embodiments, the first polypeptide domain is linked to the linker peptide described in the first aspect of the present invention at its N-terminus.
[0051] In some specific embodiments, the first polypeptide domain is selected from nucleic acid programmable nucleotide-binding proteins (napDNAbp) or homologs thereof.
[0052] In some specific embodiments, the first polypeptide domain is selected from RNA-guided nucleic acid programmable nucleotide-binding proteins, homing endonucleases (Meganuclease), zinc finger fusion proteins (ZFN), or TALEN, or homologs thereof.
[0053] In some specific embodiments, the first polypeptide domain is selected from type II CRISPR-Cas polypeptides, type I CRISPR-Cas polypeptides, type III CRISPR-Cas polypeptides, type IV CRISPR-Cas polypeptides, type V CRISPR-Cas polypeptides, type VI CRISPR-Cas polypeptides, type VII CRISPR-Cas polypeptides, IscB polypeptides, TnpB polypeptides, IsrB polypeptides, or homologs thereof.
[0054] In some specific embodiments, the first polypeptide domain is selected from Cas9, CasX, CasY, Cas 12a (Cpf1), Cas 12b (C2cl), Cas 13a (C2c2), Cas 12c (C2c3), Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, Argonaute (Ago), or homologs thereof.
[0055] In some specific embodiments, the first polypeptide domain has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to the amino acids shown in any one of SEQ ID NOs: 92-95.
[0056] In some specific embodiments, the second polypeptide domain is selected from a localization signal, a reporter protein, a CRISPR-Cas effector protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage activity polypeptide, a ligase.
[0057] In another preferred example, the second polypeptide domain comprises a deamination domain.
[0058] In some specific embodiments, the deamination domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.
[0059] In some specific embodiments, the adenosine deaminase catalytic domain is adenosine deaminase.
[0060] In some specific embodiments, the adenine deaminase (or adenosine deaminase) is any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).
[0061] In some specific embodiments, adenine deaminase can catalyze the hydrolysis deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively.
[0062] In some specific embodiments, the adenosine deaminase catalyzes the hydrolysis deamination of adenine or adenosine in DNA.
[0063] In some specific embodiments, the adenine deaminase can be any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).
[0064] In some specific embodiments, the adenosine deaminase has an amino acid sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100% sequence identity to any one of SEQ ID NOs: 96 - 98.
[0065] In some specific embodiments, the deamination domain is a cytidine deaminase catalytic domain.
[0066] In some specific embodiments, the cytidine deaminase catalytic domain includes an APOBEC deaminase.
[0067] In some specific embodiments, the cytidine deaminase can be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570), and / or variant versions thereof.
[0068] In some specific embodiments, the cytidine deaminase is APOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 99.
[0069] In some specific embodiments, the cytidine deaminase is APOBEC3A deaminase having the amino acid sequence of SEQ ID NO: 100.
[0070] In some specific embodiments, the amino acid sequence of the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity with the amino acid sequence of the wild-type cytidine deaminase (such as POBEC1 deaminase, APOBEC3A deaminase).
[0071] In some specific embodiments, the amino acid sequence of the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 100% identity with the amino acid sequence shown in SEQ ID NO: 99 or SEQ ID NO: 100.
[0072] In some embodiments, the polynucleotide encoding the cytidine deaminase can be codon-optimized for expression in an organism, and the codon-optimized polypeptide can have about 70% to 99.5% identity with the reference polynucleotide.
[0073] In another preferred example, the fusion protein has the structure shown in Formula I or II below:
[0074] X-Z-Y (I)
[0075] Y-Z-X (II);
[0076] Wherein,
[0077] X is the first polypeptide domain;
[0078] Y is the second polypeptide domain;
[0079] Z is the linking peptide described in the first aspect of the present invention;
[0080] "-" represents a peptide bond or a peptide linker connecting the above elements.
[0081] In another preferred embodiment, any two of X, Y, and Z are connected in a head-to-head, head-to-tail, tail-to-head, or tail-to-tail manner.
[0082] In another preferred embodiment, the "head" refers to the N-terminus of the polypeptide domain.
[0083] In another preferred embodiment, the "tail" refers to the C-terminus of the polypeptide domain.
[0084] In another preferred embodiment, the length of the peptide linker is 0-20 amino acids, preferably 0-10 amino acids.
[0085] The third aspect of the present invention provides an isolated polynucleotide encoding the linking peptide described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention.
[0086] In certain embodiments, the isolated polynucleotide includes a sequence optimized by humanization.
[0087] In certain embodiments, the polynucleotide further contains an auxiliary element selected from the group consisting of a signal peptide, a secretion peptide, a tag sequence (such as 6His), or a combination thereof, flanking the ORF of the fusion protein.
[0088] In certain embodiments, the polynucleotide is selected from the group consisting of a genomic sequence, a cDNA sequence, an RNA sequence, or a combination thereof.
[0089] In certain embodiments, the polynucleotide further contains a promoter operably linked to the ORF sequence of the fusion protein.
[0090] In certain embodiments, the promoter is selected from the group consisting of a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter.
[0091] In certain embodiments, the polynucleotide is a polynucleotide optimized according to the codon preference of the host cell.
[0092] In certain embodiments, the host cell comprises a prokaryotic cell or a eukaryotic cell.
[0093] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell or a mammalian cell (including human and non-human mammals).
[0094] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0095] In certain embodiments, the yeast cell is selected from yeast of one or more sources in the following group: Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell comprises: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.
[0096] In certain embodiments, the host cell is selected from the following group: Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.
[0097] The fourth aspect of the present invention provides a vector, which contains the polynucleotide described in the third aspect of the present invention.
[0098] In certain embodiments, the vector comprises one or more promoters, which are operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, replication origin, selection marker, nucleic acid restriction site, and / or homologous recombination site.
[0099] In certain embodiments, the vector includes a plasmid, a viral vector.
[0100] In certain embodiments, the viral vector is selected from the following group: adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpes virus, SV40, poxvirus, or a combination thereof.
[0101] In certain embodiments, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, a multifunctional vector.
[0102] The fifth aspect of the present invention provides a CRISPR-Cas complex, which comprises:
[0103] (i) a protein component, selected from the following group: the fusion protein described in the second aspect of the present invention; and
[0104] (ii) a nucleic acid component, selected from the following group: guide RNA, nucleic acid encoding the guide RNA, precursor RNA of the guide RNA, nucleic acid encoding the precursor RNA of the guide RNA, or a combination thereof;
[0105] Wherein, the protein component and the nucleic acid component are combined with each other to form a complex.
[0106] The sixth aspect of the present invention provides a CRISPR-Cas composition, comprising:
[0107] (i) A first component selected from the group consisting of: the fusion protein described in the second aspect of the present invention, the nucleotide sequence encoding the fusion protein described in the second aspect of the present invention, and any combination thereof; and
[0108] (ii) A second component, which is one or more guide RNAs, or the nucleotide sequence encoding the guide RNA;
[0109] The guide RNA can form a complex with the fusion protein described in (i).
[0110] In certain embodiments, the guide RNA comprises a direct repeat sequence and a spacer sequence from the 5' to 3' direction, and the spacer sequence can hybridize with a target sequence.
[0111] In certain embodiments, the composition further comprises a pharmaceutically acceptable carrier.
[0112] In certain embodiments, the composition comprises a pharmaceutical composition.
[0113] In certain embodiments, the dosage form of the composition is selected from the group consisting of: freeze-dried preparation, liquid preparation, or a combination thereof.
[0114] In certain embodiments, the dosage form of the composition is a liquid preparation.
[0115] In certain embodiments, the dosage form of the composition is an injection dosage form.
[0116] In certain embodiments, the composition is a cell preparation.
[0117] The seventh aspect of the present invention provides a CRISPR-Cas system, comprising one or more vectors, and the one or more vectors comprise:
[0118] (i) A first nucleic acid, which is the nucleotide sequence encoding the fusion protein described in the second aspect of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and
[0119] (ii) A second nucleic acid, which is the nucleotide sequence encoding the guide RNA; optionally, the second nucleic acid is operably linked to a second regulatory element;
[0120] Wherein:
[0121] The first nucleic acid and the second nucleic acid are present on the same or different vectors;
[0122] The guide RNA is capable of forming a complex with the fusion protein described in (i).
[0123] In certain embodiments, the vector includes a plasmid, a viral vector.
[0124] In certain embodiments, the guide RNA includes a spacer sequence capable of hybridizing with a target sequence; and a direct repeat (DR) sequence that is linked to the spacer sequence and capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.
[0125] In certain embodiments, the guide RNA includes unmodified and modified guide RNAs.
[0126] In certain embodiments, the modified guide RNA includes chemical modifications of bases.
[0127] In certain embodiments, the chemical modification includes methylation modification, methoxy modification, fluorination modification or thiolation modification.
[0128] In certain embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.
[0129] In certain embodiments, at least one component in the composition is non-naturally occurring or modified.
[0130] In certain embodiments, the spacer sequence is linked to the 3' end of the direct repeat (DR) sequence.
[0131] In certain embodiments, the spacer sequence contains a complementary sequence of the target sequence.
[0132] In certain embodiments, the target sequence is DNA from a prokaryotic cell or a eukaryotic cell or DNA formed by reverse transcription based on RNA; or, the target sequence is non-naturally occurring DNA or DNA formed by reverse transcription based on RNA.
[0133] In certain embodiments, the target sequence includes a cDNA sequence.
[0134] In certain embodiments, the target sequence includes a single-stranded DNA, double-stranded DNA sequence.
[0135] In certain embodiments, the target sequence is present inside the cell.
[0136] In certain embodiments, the target sequence is present within the cell nucleus or within the cytoplasm (e.g., an organelle).
[0137] In certain embodiments, the cell is a eukaryotic cell.
[0138] In certain embodiments, the cell is a prokaryotic cell.
[0139] In certain embodiments, the target sequence is present outside the cell.
[0140] In certain embodiments, the fusion protein comprises one or more NLS sequences.
[0141] In certain embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the fusion protein described in the second aspect of the present invention.
[0142] In certain embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the fusion protein described in the second aspect of the present invention.
[0143] The eighth aspect of the present invention provides a kit comprising one or more components selected from the following: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.
[0144] In certain embodiments, the kit further comprises a label or instructions.
[0145] In certain embodiments, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, cleaving a target gene or a non-target gene.
[0146] The ninth aspect of the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.
[0147] In certain embodiments, the delivery vector is a particle.
[0148] In certain embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).
[0149] The tenth aspect of the present invention provides a host cell, comprising the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.
[0150] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).
[0151] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0152] In certain embodiments, the yeast cell is selected from yeast of one or more sources in the following group: Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.
[0153] In certain embodiments, the host cell is selected from the following group: Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.
[0154] The eleventh aspect of the present invention provides an enzyme preparation, which includes the fusion protein described in the second aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.
[0155] In certain embodiments, the enzyme preparation includes an injection and / or a freeze-dried preparation.
[0156] The twelfth aspect of the present invention provides a kit, comprising:
[0157] A first container, and in the first container, the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or a drug containing the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.
[0158] In certain embodiments, the drug in the first container is a single-agent preparation containing the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.
[0159] In some embodiments, the dosage form of the drug is selected from the group consisting of: freeze-dried preparations, liquid preparations, or combinations thereof.
[0160] In some embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.
[0161] In some embodiments, the kit further contains an instruction manual.
[0162] The thirteenth aspect of the present invention provides a kit, comprising:
[0163] (a1) A first container, and the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector, or a drug containing the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector, located in the first container;
[0164] (b1) Optionally, a second container, and the guide RNA or its expression vector, or a drug containing the guide RNA or its expression vector, located in the second container.
[0165] In some embodiments, the first container and the second container are different containers.
[0166] In some embodiments, the drug in the first container is a single-agent preparation containing the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector.
[0167] In some embodiments, the drug in the second container is a single-agent preparation containing the guide RNA or its expression vector.
[0168] In some embodiments, the dosage form of the drug is selected from the group consisting of: freeze-dried preparations, liquid preparations, or combinations thereof.
[0169] In some embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.
[0170] In some embodiments, the kit further contains an instruction manual.
[0171] The fourteenth aspect of the present invention provides a method for targeting and editing a target gene, comprising: contacting the fusion protein described in the second aspect of the present invention, or the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention, or the enzyme preparation described in the eleventh aspect of the present invention, or the kit described in the twelfth or thirteenth aspect of the present invention with the target gene, or delivering it into a cell containing the target gene, wherein the target sequence is present in the target gene.
[0172] In some embodiments, the target gene is present inside the cell.
[0173] In some embodiments, the cell is a prokaryotic cell.
[0174] In some embodiments, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.
[0175] In some embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro.
[0176] In some embodiments, editing the target gene includes site-directed mutagenesis of the target sequence.
[0177] In some embodiments, the site-directed mutagenesis includes base changes of A→G or C→T.
[0178] In some embodiments, the target gene includes DNA, RNA.
[0179] In some embodiments, the DNA includes single-stranded DNA, double-stranded DNA.
[0180] In some embodiments, the RNA includes single-stranded RNA, double-stranded RNA.
[0181] The fifteenth aspect of the present invention provides a cell or its progeny obtained by the method described in the fourteenth aspect of the present invention, wherein the cell contains a modification that does not exist in its wild type.
[0182] In another preferred example, the wild-type cell contains an abnormality, and the abnormality of the cell modified by the method has been resolved or corrected.
[0183] The sixteenth aspect of the present invention provides a cell product of the cell or its progeny described in the fifteenth aspect of the present invention.
[0184] The seventeenth aspect of the present invention provides a cell or cell line or their progeny in vitro, ex vivo or in vivo, wherein the cell or cell line or their progeny contains: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, or the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.
[0185] In some embodiments, the cell is a prokaryotic cell.
[0186] In some embodiments, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.
[0187] In some embodiments, the cell is a stem cell or a stem cell line.
[0188] The eighteenth aspect of the present invention provides a cell preparation, comprising the host cell described in the tenth aspect of the present invention, or the cell described in the fifteenth aspect of the present invention or its progeny, or the cell product of the cell described in the sixteenth aspect of the present invention or its progeny, or the cell or cell line described in the seventeenth aspect of the present invention or their progeny.
[0189] In another preferred example, the cell preparation further comprises a pharmaceutically acceptable carrier or excipient.
[0190] In another preferred example, the cell preparation comprises an injection and / or a freeze-dried preparation.
[0191] The nineteenth aspect of the present invention provides the use of the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention or the system described in the seventh aspect of the present invention, the kit described in the eighth aspect of the present invention, the delivery composition described in the ninth aspect of the present invention, the enzyme preparation described in the eleventh aspect of the present invention, the kit described in the twelfth aspect or the thirteenth aspect of the present invention, for preparing a drug or preparation for nucleic acid editing (e.g., gene or genome editing).
[0192] In certain embodiments, the gene or genome editing includes site-directed mutagenesis.
[0193] In certain embodiments, the site-directed mutagenesis includes base changes of A→G or C→T.
[0194] The twentieth aspect of the present invention provides the use of the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention or the system described in the seventh aspect of the present invention, the kit described in the eighth aspect of the present invention, the delivery composition described in the ninth aspect of the present invention, the enzyme preparation described in the eleventh aspect of the present invention, the kit described in the twelfth aspect or the thirteenth aspect of the present invention, for preparing a drug or preparation for one or more of the following selected from the group consisting of:
[0195] (i) ex vivo gene or genome editing;
[0196] (ii) editing a target sequence in a target locus to modify a biological or non-human organism;
[0197] (iii) treating a disorder caused by a defect in a target sequence in a target locus;
[0198] (iv) treating a disorder or disease of a subject in need.
[0199] In certain embodiments, the disorder or disease is selected from Angelman syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, Pompe disease, myoclonic dystrophy, Huntington's disease (HTT), fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, hemoglobinopathy (e.g., β-thalassemia, sickle cell disease), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, hypercholesterolemia, heart disease (e.g., hypertrophic cardiomyopathy (HCM)), and cancer.
[0200] In certain embodiments, the disorder or disease includes cancer, infectious disease, metabolic disease, and neurological disease.
[0201] In certain embodiments, the disorder or disease includes a disease caused by a single-base mutation.
[0202] In certain embodiments, the disease or disorder includes one or more of hypercholesterolemia, transthyretin amyloidosis, and β-hemoglobinopathy.
[0203] In certain embodiments, the disorder or disease is caused by a pathogenic point mutation.
[0204] It should be understood that within the scope of the present invention, the above-described technical features of the present invention and the technical features specifically described hereinafter (such as in the examples) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS
[0205] Figure 1 Shows the plasmid map of the dual-luciferase reporter system.
[0206] FIG. 2 shows the editing efficiency of base editors with different linkers detected using the dual-luciferase reporter system, where Figure 2A describes the editing activity of base editors containing Linker1-22 and GS-XTEN-GS, Figure 2BThe editing activities of base editors containing Linker1, 23-31, 35, GS-XTEN-GS are described.
[0207] Figure 3 shows the editing efficiencies of base editors with different linkers at positions A5 and A7 of site1. Among them, Figure 3A The editing activities of base editors containing Linker2, 4, 6-8, 11-12, 14, 18-19, 21-22, GS-XTEN-GS at site1 are described. Figure 3B The editing activities of base editors containing Linker23-31, 35, GS-XTEN-GS at site1 are described.
[0208] Figure 4 shows the base editing efficiencies of base editors with different linkers in the PCSK9 gene. Among them Figure 4A The editing activities of base editors containing Linker23-36, GS-XTEN-GS in the PCSK9 gene are described. Figure 4B The editing activities of base editors containing Linker2, 4, 6-8, 11-12, 14, 18-19, 21-22, GS-XTEN-GS in the PCSK9 gene are described. Detailed implementation mode
[0209] Through extensive and in-depth research and a large number of screenings, the present invention has for the first time discovered new linkers and provided fusion proteins, CRISPR-Cas systems and complexes containing such linkers. Base editors containing the linkers are also provided, such as fusing DNA editing proteins such as deaminases (such as adenosine deaminase or cytidine deaminase) with nucleases (such as Cas9, Cas 12, Cas13, etc.) into fusion proteins for base modification (such as A-to-G) editing of target DNA or target RNA. The present invention also provides the applications of the fusion proteins, CRISPR-Cas systems or complexes containing the fusion proteins and the methods for editing or modifying target nucleic acids, such as for gene editing of disease targets, so as to achieve the prevention and treatment of diseases. The linkers of the present invention can significantly improve the editing efficiency. On this basis, the inventors of the present invention have completed the present invention.
[0210] Terms
[0211] The following examples are only used to describe the present invention and do not limit the present invention. Unless otherwise specified, the experiments and methods described in the examples are basically carried out according to the conventional methods well-known in the art and described in various references.
[0212] In addition, for those conditions not specified in the embodiments, they are carried out according to conventional conditions or the conditions recommended by the manufacturer. For reagents or instruments whose manufacturers are not specified, they are all conventional products that can be obtained through commercial purchase. Those skilled in the art know that the embodiments describe the present invention by way of example and are not intended to limit the scope claimed by the present invention. All the published cases and other reference materials mentioned herein are incorporated herein by reference in their entirety.
[0213] To make the present disclosure easier to understand, certain terms are first defined. As used in this application, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.
[0214] The term "about" can refer to a value or a composition within an acceptable error range of a specific value or composition determined by a person of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0215] As used herein, the terms "comprising" or "including (containing)" can be open-ended, semi-closed, and closed. In other words, the terms also include "consisting essentially of...", or "consisting of...".
[0216] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions where identical residues occur. Generally, this is expressed as a percentage. Methods for measuring sequence identity of nucleotide sequences are well known to those skilled in the art.
[0217] General definitions
[0218] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention pertains. The terms used in the description of the present invention are only for describing specific embodiments and are not intended to limit the present invention.
[0219] Unless otherwise indicated by the context, the various features described in the present invention can be used in any combination. In addition, the present invention also contemplates that in some embodiments of the present invention, any feature or combination of features set forth herein can be excluded or omitted. By way of example, if the specification states that the composition contains components 1, 2, and 3, then any one of components 1, 2, 3 or any combination thereof can be omitted and waived individually or in any combination.
[0220] As used in the description of the present invention and the appended claims, the articles "a / an" and "the" are used in the present invention to refer to one or more than one (i.e., at least one) of the grammatical objects of the said articles. For example, "element" means one element or more than one element, unless the context clearly indicates otherwise.
[0221] The use of alternatives (such as "or") should be understood to mean any one, both, or any combination of the alternatives. The term "and / or" should be understood to mean either or both of the alternatives.
[0222] As used in the present invention, when referring to measurable values such as amounts or concentrations, etc., the term "about" means covering variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value as well as the specified value. For example, "about X" (where X is a measurable value) means including X, as well as variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. The ranges provided for measurable values in the present invention can include any other ranges and / or individual values therein.
[0223] As used in the present invention, the terms "comprising", "containing", and "including" clearly state the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0224] As used in the present invention, the transitional phrase "consisting essentially of" means that the scope of the claim will be interpreted to cover the specified materials or steps recited in the claim, as well as those that do not materially affect the basic and novel features of the present invention claimed. Thus, when used in the claims of the present invention, the term "consisting essentially of" is not intended to be interpreted as equivalent to "comprising".
[0225] As used in the present invention, the terms "increase" and "enhance" (and their grammatical variations) describe an improvement of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to a control.
[0226] As used in the present invention, the terms "reduce", "decrease" and "lessen" (and their grammatical variations) describe, for example, a reduction of at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% compared to a control. In particular embodiments, the reduction may result in no or substantially no (i.e., a non-significant amount, such as less than about 10% or even 5%) detectable activity or amount.
[0227] Throughout this specification, reference to "one embodiment", "an embodiment", "a particular embodiment", "a related embodiment", "some specific embodiments", "a certain embodiment", "another embodiment" or "a further embodiment" or combinations thereof means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Thus, the appearances of the foregoing phrases throughout the specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner in one or more embodiments.
[0228] "Sequence identity" or "sequence similarity" between two polypeptide or nucleic acid sequences refers to the percentage of the number of identical residues between the sequences out of the total number of residues, and the calculation of the total number of residues is determined based on the type of mutations. Types of mutations include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / replacements of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence. Taking polypeptides as an example (similarly for nucleotides), if the type of mutation is one or more of the following: substitutions / replacements of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence, then the total number of residues is calculated as the larger of the two sequences being compared. If the type of mutation also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, then the number of amino acids inserted or deleted at either or both ends (e.g., the number of insertions or deletions at both ends is less than 20) is not counted in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps in the alignment (if any) are resolved by a specific algorithm.
[0229] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.
[0230] A "natural" or "wild-type" nucleic acid, nucleotide sequence, polypeptide or amino acid sequence refers to a nucleic acid, nucleotide sequence, polypeptide or amino acid sequence that occurs naturally or is endogenous. Thus, for example, a "wild-type mRNA" is an mRNA that occurs naturally in an organism or is endogenous to the organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.
[0231] Those skilled in the art will appreciate that the structure of a protein can be altered without adversely affecting its activity and functionality. For example, one or more conservative amino acid substitutions can be introduced into the amino acid sequence of the protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule.
[0232] Those skilled in the art are aware of examples and embodiments of conservative amino acid substitutions. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group as the site to be replaced, i.e., a nonpolar amino acid residue can be replaced with another nonpolar amino acid residue, a polar uncharged amino acid residue can be replaced with another polar uncharged amino acid residue, a basic amino acid residue can be replaced with another basic amino acid residue, and an acidic amino acid residue can be replaced with another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, conservative substitutions in which one amino acid is replaced by another amino acid belonging to the same group fall within the scope of the present invention. Thus, the proteins of the present invention can contain one or more conservative substitutions in the amino acid sequence, and these conservative substitutions are preferably made according to Table A. In addition, the present invention also encompasses proteins that further contain one or more other non-conservative substitutions, as long as the non-conservative substitutions do not significantly affect the desired functions and biological activities of the proteins of the present invention.
[0233] Conservative amino acid substitutions can be made at one or more predicted non-essential amino acid residues. A "non-essential" amino acid residue is an amino acid residue that can be altered (deleted, substituted or replaced) without changing the biological activity, while an "essential" amino acid residue is required for biological activity. A "conservative amino acid substitution" is a substitution in which an amino acid residue is replaced by an amino acid residue having a similar side chain. Amino acid substitutions can be made in non-conserved regions of the Cas enzyme. Generally, such substitutions are not made on conserved amino acid residues or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants can have fewer conservative or non-conservative changes in conserved regions.
[0234] Table A
[0235]
[0236]
[0237] Those skilled in the art are aware that changing (substituting, deleting, truncating or inserting) one or more amino acid residues from the N- and / or C-terminus of a protein can still retain its functional activity. Therefore, proteins in which one or more amino acid residues have been changed from the N- and / or C-terminus of the Cas protein of the present invention while retaining its desired functional activity are also within the scope of the present invention. These changes can include those introduced by modern molecular methods such as PCR, which include PCR amplification that changes or extends the protein coding sequence by means of including amino acid coding sequences among the oligonucleotides used in the PCR amplification.
[0238] It should be recognized that proteins can be modified in various ways, including amino acid substitutions, deletions, truncations and insertions, and the methods for such operations are generally known to those skilled in the art.
[0239] Linker / Adapter / Linking Peptide
[0240] The terms "linker", "adapter", "linking peptide" and "adapter polypeptide" are used interchangeably and all refer to a polypeptide amino acid sequence of two or more amino acids that covalently links a first protein or polypeptide to a second protein polypeptide.
[0241] The present invention provides a variety of new linkers. For example, the linker can be Linker1-linker36, and the sequences are shown in Table 1.
[0242] In some specific embodiments, the linker has an amino acid sequence having at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence shown in any one of SEQ ID NO:1-36.
[0243] In some embodiments, the linker has a sequence with one or more base substitutions, deletions or additions compared to the amino acid sequence shown in any one of SEQ ID NO:1-36, and substantially retains the biological function of the sequence from which it is derived.
[0244] Cas Protein
[0245] In the present invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably. Cas protein is taken in its broadest sense and includes wild-type Cas protein, its derivatives or variants, analogs, and its functional fragments such as oligonucleotide-binding fragments.
[0246] In some embodiments, the Cas protein comprises an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence shown in any one of SEQ ID NOs: 92-95.
[0247] In some embodiments, the Cas protein comprises the amino acid sequence shown in any one of SEQ ID NOs: 92-95.
[0248] Orthologue, ortholog
[0249] As used herein, the term "orthologue, ortholog" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein as described herein refers to a protein from a different species that performs the same or similar function as the protein of which it is an orthologue.
[0250] Nucleic acid cleavage of the present invention includes: DNA or RNA breaks in the target nucleic acid generated by the Cas protein (Cis cleavage), and DNA or RNA breaks in a collateral nucleic acid substrate (single-stranded nucleic acid substrate) caused by the trans-cleavage activity of the Cas protein (i.e., non-specific or non-targeted, Trans cleavage). In some specific embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0251] Trans cleavage refers to the fact that in certain environments, the activated Cas12 family protein remains active after binding to the target sequence and continues to non-specifically cleave non-target oligonucleotides. This trans-cleavage activity can be used to detect the presence of specific target oligonucleotides using the Cas system. For example, the Cas12i system has been engineered to non-specifically cleave ssDNA or transcripts. The trans-cleavage activity is used in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, which can be used for many clinical diagnoses (Gootenberg, J.S. et al., Nucleic acid detection with CRISPR-Cas13a / C2c2. Science 356, 438-442 (2017)).
[0252] Fusion protein
[0253] As used in the present invention, the term "fusion protein" refers to a hybrid (e.g., chimeric, recombinant) polypeptide that contains protein domains from at least two different proteins. One protein domain can be located in the amino-terminal (N-terminal) portion of the fusion protein and will contain the free N-terminus of the fusion protein (e.g., the amino (NH2) group), and this protein domain of the fusion protein can be referred to as the "amino-terminal fusion protein" or "amino-terminal fusion protein domain". Similarly, one protein domain can be located in the carboxyl-terminal (C-terminal) portion of the fusion protein and will contain the free C-terminus of the fusion protein (e.g., the carboxyl (COOH) group), and this protein domain of the fusion protein can be referred to as the "carboxyl-terminal fusion protein" or "carboxyl-terminal fusion protein domain".
[0254] In one embodiment, the fusion protein comprises the amino acid sequence shown in SEQ ID NO:73.
[0255] The present invention provides a fusion protein comprising Linker1-36 (SEQ ID NO:1-36).
[0256] In some embodiments, the linker GS-XTEN-GS (SEQ ID NO.75) in the amino acid sequence shown in SEQ ID NO:73 is replaced with Linker1-36 (SEQ ID NO:1-36) to obtain a fusion protein comprising Linker1-36 (SEQ ID NO:1-36).
[0257] In some embodiments, the fusion protein comprises an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence of the fusion protein comprising Linker1-36 (SEQ ID NO:1-36).
[0258] CRISPR system
[0259] The term "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" is used interchangeably and has the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of guiding the activity of the Cas genes.
[0260] CRISPR / Cas complex
[0261] The term "CRISPR / Cas complex" refers to a complex formed by the binding of a gRNA (guide RNA) or a mature crRNA (or guide RNA) to a Cas protein or a fusion protein containing the same, which contains a direct repeat sequence hybridized to the guide sequence of the target sequence and bound to the Cas protein, and the complex is capable of recognizing and cleaving a target nucleotide that can hybridize to the guide RNA or mature crRNA.
[0262] gRNA (guide RNA, guideRNA)
[0263] The terms "guide RNA (guide RNA, gRNA)", "mature crRNA", "guide sequence", "guide RNA" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may contain a direct repeat (DR) sequence and a guide sequence, or consist essentially of or consist of a direct repeat (DR) sequence and a guide sequence.
[0264] In some cases, the guide sequence is any polynucleotide sequence that has sufficient complementarity to the target sequence to hybridize to the target sequence and direct the specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. The guide sequence contains a sequence (such as a spacer (DR) sequence) that has sufficient complementarity to the target nucleic acid sequence to hybridize to the target nucleic acid sequence and direct the sequence-specific binding of the complex to the target nucleic acid sequence.
[0265] Target nucleic acid
[0266] In the present invention, the terms "target nucleic acid", "target sequence", "target nucleic acid sequence" are used interchangeably and refer to a specific nucleic acid that contains a nucleic acid sequence that is fully or partially complementary to the guide sequence in the gRNA. The "target sequence" refers to a polynucleotide targeted by the guide sequence in the gRNA, such as a sequence complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including a Cas protein and a gRNA). Perfect complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. In some embodiments, the target nucleic acid contains a non-coding region (e.g., a promoter or a terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded.
[0267] The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located intracellularly or extracellularly. In some cases, the target sequence is located within the cell nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts).
[0268] The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be related to a protospacer adjacent motif (PAM).
[0269] Donor template
[0270] In the present invention, the donor template nucleic acid or donor template can be used interchangeably and refers to a nucleic acid molecule that, after the Cas protein described herein has altered the target nucleic acid, one or more cellular proteins can use to alter the structure of the target nucleic acid.
[0271] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid or a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template nucleic acid is an exogenous nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination can be achieved using the donor template, and the recombination is homologous recombination.
[0272] Cleavage
[0273] Cleavage refers to a DNA break in the target nucleic acid generated by the Cas protein described herein. In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break.
[0274] In the present invention, the meanings of cleaving the target nucleic acid or modifying the target nucleic acid can overlap. Modifying the target nucleic acid includes not only modification of a single nucleotide, such as site-directed mutagenesis, but also insertion or deletion of nucleic acid fragments.
[0275] Functional domain
[0276] In the present invention, the functional domain is taken in its broadest sense and includes a protein such as an enzyme or a factor itself or a fragment / domain thereof having a specific function. A Cas protein (e.g., a dCas protein) is linked / associated with one or more functional domains selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a lytic activity polypeptide, a ligase. When more than one functional domain is included, the functional domains can be the same or different.
[0277] Deamination domain
[0278] In the present invention, the deamination domain includes a deaminase (such as adenosine deaminase or cytidine deaminase) catalytic domain. As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to inosine (or the inosine portion of a molecule), as shown below. In some specific embodiments, the adenine-containing molecule is adenosine (A), and the inosine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0279] Adenosine deaminases include, but are not limited to, members of the enzyme family called adenosine deaminases acting on RNA (ADAR), members of the enzyme family called adenosine deaminases acting on tRNA (ADAT), and other family members containing an adenosine deaminase domain (ADAD). According to the present disclosure, adenosine deaminases are capable of targeting adenine in RNA / DNA and RNA duplexes. In certain embodiments, adenosine deaminases have been modified to increase their ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.
[0280] In some embodiments, the deaminase is a cytidine deaminase. The term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts cytidine (or the cytidine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytidine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytidine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0281] Cytidine deaminases include, but are not limited to, members of the enzyme family called apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-induced deaminase (AID), or cytidine deaminase 1 (CDA1). In certain embodiments, APOBEC family deaminases.
[0282] Vector
[0283] A vector is a nucleic acid molecule capable of transporting another nucleic acid molecule linked thereto.
[0284] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or no free ends (such as circular ones); nucleic acid molecules including DNA, RNA, or both; and other diverse polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection, enabling the genetic material elements it carries to be expressed in the host cell. A vector can be introduced into a host cell and thereby produce transcripts, proteins, or peptides, including those from proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts such as nucleic acid transcripts, proteins, or enzymes). A vector can contain various elements controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. A vector can also contain an origin of replication.
[0285] Vectors include plasmids and viral vectors. A plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. A viral vector, in which virus-derived DNA or RNA sequences are present in the vector used for packaging the virus. Viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. A viral vector also contains the polynucleotide carried by the virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0286] Other vectors (e.g., non-episomal mammalian vectors) integrate into the genome of the host cell after being introduced into the host cell and thereby replicate together with the host genome. Moreover, certain vectors are capable of directing the expression of genes operably linked to them. Such vectors are called "expression vectors".
[0287] Regulatory elements
[0288] "Regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression control elements (e.g., transcriptional termination signals such as polyadenylation signals, poly-U sequences), a detailed description of which can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can predominantly direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or particular cell types (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell type-specific.
[0289] "Promoter" refers to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as in response to a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulatable promoters include, for example, promoters induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, phytohormones, wounding, or chemicals (such as ethanol, abscisic acid (ABA), jasmonates, salicylic acid, or safeners).
[0290] Host cell
[0291] "Host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism cultured as a single cell entity (e.g., a cell line), which is used as a recipient for a nucleic acid (e.g., an expression vector), and includes the progeny of the original cell that has been genetically modified by the nucleic acid.
[0292] It should be understood that the progeny of a single cell may, due to natural, accidental or deliberate mutations, not necessarily have exactly the same morphology or genome as the original parental cell. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.
[0293] Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of host cell to be transformed, the desired level of expression, etc.
[0294] NLS
[0295] NLS refers to the "nuclear localization sequence" or "nuclear localization signal", which is an amino acid sequence that promotes the entry of a protein into the cell nucleus. Nuclear localization sequences are known in the art (for example, as described in International PCT Application PCT / EP2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001), which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequences: KRTADGSEFESPKKKRKV, AVKRPAATKKAGQAKKKKLD, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0296] Operably linked
[0297] "Operably linked" means that a target nucleotide sequence is linked to a regulatory element in a manner that permits the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can also be selected to target specific types of cells.
[0298] Complementary
[0299] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the percentage of complementarity is 50%, 60%, 70%, 80%, 90%, and 100%). "Fully complementary" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.
[0300] The term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid that is complementary to a target sequence hybridizes predominantly to the target sequence and substantially not to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence hybridizes specifically to its target sequence.
[0301] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding of the bases between these nucleotide residues. The complex can comprise two strands that form a duplex, three or more strands that form a multi-stranded complex, a single self-hybridizing strand, or any combination of these. The hybridization reaction can constitute a step in a broader process (such as the start of PCR, or the cleavage of a polynucleotide by an enzyme). A sequence that can hybridize to a given sequence is called the "complement" of the given sequence.
[0302] Hybridization of the target sequence with the gRNA means that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or it represents that at least 12, 15, 16, 17, 18, 19, 20, or more bases of the nucleic acid sequences of the target sequence and the gRNA can be complementary paired to hybridize and form a complex.
[0303] Expression
[0304] Nucleic acid expression includes one or more of generating an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5′ capping, and / or 3′ end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.
[0305] Delivery
[0306] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein-RNA. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.
[0307] Kit
[0308] In one aspect, the present invention provides a kit, which includes the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned vector, the aforementioned CRISPR-Cas composition, the use of the aforementioned system in the preparation of a kit, and the components of the kit are in the same or different containers.
[0309] In one aspect, the present invention further provides a container that contains the aforementioned kit.
[0310] In one embodiment, the container includes a sterile container;
[0311] In one embodiment, the container includes a syringe.
[0312] In some embodiments, the kit further includes instructions for using the kit, such as instructions in more than one language. The kit can also contain one or more reagents for use in the process using the aforementioned one or more components. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction or storage buffers. The aforementioned reagents can be provided in a form that requires the addition of one or more other components before use (e.g., in a concentrated or lyophilized form); the buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. The buffer can have a suitable pH value (pH), for example, it can be alkaline. In some embodiments, the pH of the buffer is between about 7 and 10.
[0313] Treatment
[0314] "Treatment" means treating or curing a subject's disorder, delaying the onset of symptoms of the disorder, and / or delaying the severity of the disorder. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals such as bovines, equines, ovines, porcines, canines, felines, lagomorphs (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder caused by a disease-related gene defect). "Plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developmental stage.
[0315] In one aspect, the present invention also provides the use of the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, complex, system, kit, delivery composition, enzyme preparation, host cell in the preparation of a medicament for treating a disorder or disease of a subject in need thereof.
[0316] In one embodiment, the use includes administering the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, complex, system, kit, delivery composition, enzyme preparation to the subject or an isolated cell of the subject.
[0317] In one embodiment, the disorder or disease includes a disease caused by a single-base mutation.
[0318] In one embodiment, the disorder or disease is selected from Angelman syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, Pompe disease, myoclonic dystrophy, Huntington's disease (HTT), fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, hemoglobinopathy (e.g., β-thalassemia, sickle cell disease), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, hypercholesterolemia, heart disease (e.g., hypertrophic cardiomyopathy (HCM)), and cancer.
[0319] In one embodiment, the clinical variant database is obtained from the NCBI ClinVar database available on the NCBI ClinVar website. Pathogenic single nucleotide polymorphisms (SNPs) are identified from this list. Using genomic locus information, CRISPR targets in the regions overlapping with and surrounding each SNP are identified. The selection of SNPs that can be corrected by using base editing in combination with a Cas protein or its variant to target causal mutations is listed in the table below, with only one alias for each disease listed in the table. "RS#" corresponds to the RS accession number in the SNP database on the NCBI website. "AlleleID" corresponds to the causal allele accession number. The "Name" column contains the locus identifier of the gene, the gene name, the position of the mutation in the gene, and the change caused by the mutation.
[0320] Disease targets of base editing
[0321]
[0322]
[0323]
[0324]
[0325]
[0326]
[0327]
[0328]
[0329]
[0330]
[0331]
[0332]
[0333]
[0334]
[0335]
[0336]
[0337]
[0338]
[0339]
[0340]
[0341]
[0342] In one embodiment, the disorder or disease includes cancer, infectious disease, metabolic disease, and neurological disease.
[0343] In one embodiment, the cancer includes one or more of the diseases or disorders including hypercholesterolemia, transthyretin amyloidosis, and β-hemoglobinopathy.
[0344] The main advantages of the present invention include:
[0345] (1) The linker mined and designed in the present invention has a large difference from the linkers used in the prior art. For example, compared with the amino acid sequence shown in GS-XTEN-GS (SEQ ID NO. 75), its similarity is very low (for example, the amino acid sequence homology of Linker12 and GS-XTEN-GS is only less than 19%).
[0346] (2) Compared with the widely used linker, such as GS-XTEN-GS (SEQ ID NO. 75), the base editing efficiency of the base editors constructed with various Linkers of the present invention is equivalent or higher. For example, as can be seen from Figure 2, the editing efficiency of Linker22-5V17.2-nCas9 is significantly higher than that of GS-XTEN-GS-5V17.2-nCas9.
[0347] (3) The base editors (such as adenosine base editors or cytidine base editors) composed of the linkers designed in the present invention can be used for the treatment of diseases caused by single-base mutations.
[0348] (4) The linkers designed in the present invention can be widely used for the connection between various molecules or parts, and have broad application prospects.
[0349] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are generally carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight percentage and weight parts.
[0350] Unless otherwise specified, the reagents and materials in the embodiments of the present invention are all commercially available products.
[0351] Example 1. Construction of base editors with different linkers
[0352] 1. A variety of linkers were designed, and the sequences are shown in the following table:
[0353] Table 1 Amino acid sequences of different linkers and nucleotide coding sequences after codon optimization
[0354]
[0355]
[0356] 2. Construction of an adenosine base editor expression vector with linker Linker1-36
[0357] Taking the base editor 5V17.2-nCas9 in the prior art as a basis, its amino acid sequence (SEQ ID NO.73) is shown as follows:
[0358]
[0359] Among them, the double-underlined sequence is the nCas9 amino acid sequence, the single-underlined sequence is the deaminase 5V17.2 amino acid sequence, the bold italic sequence is the linker GS-XTEN-GS (SEQ ID NO.75), which is located between the 976th and 1071st positions of the expression vector, the sequences without labels at both ends are NLS sequences, and the asterisk at the C-terminus represents the position of the stop codon.
[0360] The corresponding nucleotide coding sequence of the above base editor is shown as SEQ ID NO.74, and the nucleotide coding sequence of GS-XTEN-GS is shown as SEQ ID NO.76.
[0361] Using the base editor 5V17.2-nCas9 expression vector as a template, upstream and downstream primers for the vector backbone were designed based on the position of the linker GS-XTEN-GS (SEQ ID NO.75) on the expression vector, and PCR amplification was performed to obtain the backbone vector PCR product.
[0362] The upstream and downstream primers for the vector backbone were vector backbone-F (SEQ ID NO.77) and vector backbone-R (SEQ ID NO.78) respectively. The base editor 5V17.2-nCas9 expression vector template was amplified using the 2X Phanta Flash master mix (Novizan, P520-01). The PCR amplification reaction system and procedure were as follows:
[0363] Reaction system: 1 μl of vector backbone-F (10 μM), 1 μl of vector backbone-R (10 μM), 10 μl of 2X Phanta Flash master mix (Novizan, P520-01), 20 ng of template plasmid, and sterilized water was added to make up to 20 μl.
[0364] Reaction procedure: Pre-denaturation at 98°C for 30 s; denaturation at 98°C for 10 s, annealing at 60°C for 5 s, extension at 72°C for 45 s; 30 cycles; final extension at 72°C for 1 min.
[0365] After that, the universal DNA purification and recovery kit (Tiangen, DP214) was used to purify and recover the backbone vector PCR product. The amplified backbone sequence is shown in SEQ ID NO.79.
[0366] Each linker nucleotide coding sequence fragment in Table 1 of Example 1 was synthesized, and PCR primers corresponding to each linker nucleotide coding sequence were designed. The 5' ends of the upstream and downstream primers for the linker each carried a homologous arm sequence of the vector backbone. The designed primers are shown in Table 2. Each linker nucleotide coding sequence fragment and primer were synthesized by Biosciences (Shanghai) Co., Ltd.
[0367] Table 2 Upstream and downstream primer sequences for each linker
[0368]
[0369]
[0370] After that, each linker sequence fragment was amplified by PCR using the corresponding upstream and downstream primers in Table 4, and the linker PCR product was obtained. The amplification reaction system and reaction procedure were as follows:
[0371] Reaction system: 1 μl of Linker-F (10 μM), 1 μl of Linker-R (10 μM), 10 μl of 2X Phanta Flash mastermix (Novizan, P520-01), and make up to 20 μl with sterilized water.
[0372] The reaction procedure is as follows:
[0373] Pre-denaturation at 98°C for 30 s; denaturation at 98°C for 10 s, annealing at 60°C for 5 s, extension at 72°C for 2 s; 20 cycles; final extension at 72°C for 1 min.
[0374] Purify and recover the target fragment using a universal DNA purification and recovery kit (Tiangen, catalog number DP214).
[0375] 3. Using the ClonExpress MultiS One Step Cloning Kit (Novizan Biotech, catalog number C113-01), a non-ligase-dependent multi-fragment one-step cloning kit, recombine and ligate each linker fragment with the backbone vector fragment. Immediately place it on ice after reacting at 37°C for 30 min to obtain the recombinant base editor expression plasmid. Then transform the recombinant base editor expression plasmid into DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.). The specific method is as follows:
[0376] Take out the DH5α competent cells from the refrigerator and thaw them on ice. Add 5 μl of the recombinant base editor expression plasmid to 50 μl of the competent cells, mix gently, place on ice for 20 min, heat shock in a 42°C water bath for 1 min, quickly transfer to ice after taking out, and let it stand for 2 min. Add 500 μl of LB medium and shake the bacteria in a shaker at 220 rpm for 30 min. Take 100 μl of the bacterial solution and spread it on a plate. Invert the plate and place it in a 37°C incubator. After overnight culture, pick single colonies, shake the bacteria, and send the bacterial solution to a sequencing company for Sanger sequencing. Use the Novizan plasmid miniprep kit (catalog number DC203-01) to extract the plasmid according to the instructions for the plasmid bacteria solution with correct sequencing.
[0377] Obtain 5V17.2-nCas9 recombinant vectors containing Linker1-36, and name them Linker1-5V17.2-nCas9 recombinant vector, Linker2-5V17.2-nCas9 recombinant vector, Linker3-5V17.2-nCas9 recombinant vector, Linker4-5V17.2-nCas9 recombinant vector, Linker5-5V17.2-nCas9 recombinant vector, Linker6-5V17.2-nCas9 recombinant vector, Linker7-5V17.2-nCas9 recombinant vector, Linker8-5V17.2-nCas9 recombinant vector, Linker9-5V17.2-nCas9 recombinant vector, Linker10-5V17.2-nCas9 recombinant vector, Linker11-5V17.2-nCas9 recombinant vector, Linker12-5V17.2-nCas9 recombinant vector, Linker13-5V17.2-nCas9 recombinant vector, Linker14-5V17.2-nCas9 recombinant vector, Linker15-5V17.2-nCas9 recombinant vector, Linker16-5V17.2-nCas9 recombinant vector, Linker17-5V17.2-nCas9 recombinant vector, Linker18-5V17.2-nCas9 recombinant vector, Linker19-5V17.2-nCas9 recombinant vector, Linker20-5V17.2-nCas9 recombinant vector, Linker21-5V17.2-nCas9 recombinant vector, Linker22-5V17.2-nCas9 recombinant vector, Linker23-5V17.2-nCas9 recombinant vector, Linker24-5V17.2-nCas9 recombinant vector, Linker25-5V17.2-nCas9 recombinant vector, Linker26-5V17.2-nCas9 recombinant vector, Linker27-5V17.2-nCas9 recombinant vector, Linker28-5V17.2-nCas9 recombinant vector, Linker29-5V17.2-nCas9 recombinant vector, Linker30-5V17.2-nCas9 recombinant vector, Linker31-5V17.2-nCas9 recombinant vector, Linker32-5V17.2-nCas9 recombinant vector, Linker33-5V17.2-nCas9 recombinant vector, Linker34-5V17.2-nCas9 recombinant vector, Linker35-5V17.2-nCas9 recombinant vector, Linker36-5V17.2-nCas9 recombinant vector.
[0378] Example 2. Comparing the editing efficiency of base editors with different linkers using a dual-luciferase reporter system
[0379] To efficiently and sensitively detect the editing activities of various base editors, a dual-luciferase reporter system vector (SEQ ID NO. 81, see the plasmid map in Figure 1 ) was constructed, which contains the Fluc luciferase (Firefly luciferase) coding sequence (SEQ ID NO. 82) and the NanoLuc luciferase (NanoLuc) coding sequence. A target sequence (SEQ ID NO. 80) and a PAM site 5'-NGG located at one end of the target sequence were designed in the front section of the NLuc sequence. Since the complementary sequence of the target sequence contains the stop codon TGA, under normal circumstances, the NanoLuc coding sequence is not expressed. When the target sequence is effectively edited by the adenosine base editor, the stop codon TGA will become CGA, enabling the normal expression of NanoLuc and emitting fluorescence. Therefore, taking Fluc as an internal reference, the base editing efficiency of the editor can be obtained by calculating the luminescence reading value of NanoLuc.
[0380] The construction method is as follows:
[0381] The pmEGFP-C1 plasmid (Addgene, Plasmid, #36412) was digested with the restriction endonucleases AseI (Thermo, ER0911), NheI (Thermo, FD0974), and XmaI (NEB) to obtain digestion products 1 (582 bp), product 2 (795 bp), and product 3 (3353 bp). Products 1 and 3 were recovered.
[0382] The F-Luc sequence, N-Luc sequence, N-Luc' sequence, EF1a promoter, and BGH-polyA sequence were synthesized respectively, and the synthesis method is as follows:
[0383]
[0384] The PCR reaction system is as follows: 33 μl of sterile water, 5 μl of 10×KOD Buffer, 5 μl of dNTP, 3 μl of MgSO4, 1 μl of each upstream and downstream primer, 1 μl of PCR template (about 10 ng), and 1 μl of KOD enzyme. After preparation, vortex and mix well, and centrifuge briefly at high speed, then place it in the PCR instrument. The PCR reaction program is as follows: pre-reaction at 98°C for 1 min; thermal denaturation at 98°C for 15 s, annealing at 60°C for 30 s, extension at 72°C for 30 s / kb, and after 30 cycles, extension at 72°C for 10 min. The PCR products were stored at 12°C, and then the DNA fragments were recovered by cutting the gel after agarose gel electrophoresis, and the EF1a promoter, N-Luc, N-Luc’, and BGH-polyA fragments were amplified.
[0385] The EF1a promoter, N-Luc, and BGH-polyA fragments were subjected to PCR splicing to obtain the first splicing fragment. The EF1a promoter, N-Luc’, and BGH-polyA fragments were subjected to PCR splicing to obtain the first splicing fragment’. The first splicing fragment and the first splicing fragment’ were respectively spliced with the obtained product 1, product 3, and the F-Luc sequence. The splicing was carried out using the Gibson Master Mix kit (NEB, E2611L), and the operation was carried out according to the kit instructions to obtain the reporter plasmid and the reference plasmid.
[0386]
[0387] The spacer sequence was designed according to the target sequence (SEQ ID NO.80) and constructed onto the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain the sgRNA recombinant expression vector, which can transcribe the sgRNA sequence (SEQ ID NO.83).
[0388] The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vector including each linker were co-transfected into HEK293T cells (purchased from ATCC) with the sgRNA recombinant expression vector and the dual-luciferase reporter system vector (50 ng each of the base editor expression vector, sgRNA expression vector, and dual-luciferase reporter system vector) by the PEI transfection method. After 48 h of transfection, the detection of dual-luciferase expression was carried out. The detection reagent was the Dual-Luciferase Reporter Assay Kit (Beyotime, product number RG028), and the Fluc fluorescence value was used as the internal reference. The editing efficiency was as Figure 2A 、 2B shown. Through analysis, it was found that the editing activity of Linker13-5V17.2-nCas9 was comparable to that of GS-XTEN-GS-5V17.2-nCas9. The editing activities of the base editors including linkers Linker1-12, 14-31, and 35 were significantly higher than those of GS-XTEN-GS-5V17.2-nCas9. In particular, the editing activity of Linker22-5V17.2-nCas9 was approximately twice that of GS-XTEN-GS-5V17.2-nCas9.
[0389] Example 3. Editing Activity of Base Editors with Different Linkers in HEK293T Cells
[0390] To compare the editing activities of various base editors on endogenous genes in mammalian cells, the site1 locus (see Gaudelli NM, Komor AC, Rees HA, Packer MS, Badran AH, Bryson DI, Liu DR. Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage. Nature. 2017 Nov 23;551(7681):464-471. doi:10.1038 / nature24644. Epub 2017 Oct 25. Erratum in: Nature. 2018 May 2;: PMID:29160308; PMCID:PMC5726555.) was selected as the target.
[0391] According to the site1 target sequence (SEQ ID NO.84), a spacer sequence was designed and constructed onto the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain the site1-sgRNA expression vector, which can transcribe the sgRNA sequence (SEQ ID NO.85) targeting the site1 locus. The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vector including each linker were co-transfected into HEK293T cells with the lenti U6-sgRNA / EF1a-mCherry recombinant expression vector (75 ng of each of the base editor expression vector plasmid and the sgRNA expression vector) by the PEI transfection method. After continued culture for 24 h, the genomic DNA of the transfected positive cells was extracted, and the upstream primer site1-F (SEQ ID NO.86) and the downstream primer site1-R (SEQ ID NO.87) were designed and synthesized to perform PCR amplification on the site1 target sequence locus. The amplified product was subjected to Sanger sequencing (Bosoom Biotech) and analyzed for editing efficiency using the online tool https: / / hanlab.cc / beat / Figure 3A 、 3B) At the A5 site of site1, the editing efficiency of the base editor containing Linker11, 22-24, 26-31, 35 is significantly higher than or comparable to that of GS-XTEN-GS-5V17.2-nCas9. At the A7 site of site1, the editing efficiency of the base editor containing Linker2, 4, 6, 11-12, 18-19, 22-24, 26-31, 35 is significantly higher than or comparable to that of GS-XTEN-GS-5V17.2-nCas9.
[0392] The site of the target sequence (SEQ ID NO.88) in the PCSK9 locus was selected for testing. According to the target sequence (SEQ ID NO.88), a spacer sequence was constructed and built into the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain a recombinant expression vector, which can transcribe the sgRNA sequence (SEQ ID NO.89) targeting the PCSK9 target sequence site.
[0393] The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vectors including each linker were co-transfected into HEK293T cells with the lenti U6-sgRNA / EF1a-mCherry recombinant expression vector by the PEI transfection method. After continued culture for 24 h, the genomes of the transfected positive cells were extracted. The upstream primer PCSK9-F (SEQ ID NO.90) and the downstream primer PCSK9-F (SEQ ID NO.91) were designed and synthesized to perform PCR amplification on the PCSK9 locus. The amplified products were subjected to fastNGS sequencing (Qingke Biotechnology). It was found through testing that at the PCSK9 locus, each base editor had a significant editing efficiency. In particular, the editing activities of the base editors of Linker23, Linker29, Linker35-5V17.2-nCas9 were significantly higher than those of GS-XTEN-GS-5V17.2-nCas9, and the editing activity of Linker24-5V17.2-nCas9 at this locus was comparable to that of GS-XTEN-GS-5V17.2-nCas9. The base editors containing Linker25-28, 30-34, 36 all had obvious editing activities ( Figure 4A ) The base editors including Linker2, 4, 6-8, 11-12, 14, 18-19, 21-22 had significant editing activities at the PCSK9 locus, and the editing activity of Linker7-5V17.2-nCas9 was comparable to that of GS-XTEN-GS-5V17.2-nCas9 ( Figure 4B)。
[0394] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0395] Sequence information
[0396]
[0397]
[0398]
[0399]
[0400]
[0401]
[0402]
[0403]
[0404]
[0405]
[0406]
[0407]
[0408]
[0409]
[0410]
[0411]
[0412]
[0413]
[0414]
[0415]
[0416]
[0417]
[0418]
[0419]
[0420]
[0421]
[0422]
[0423]
[0424]
[0425] All documents mentioned in the present invention are cited herein by reference as if each individual document was cited separately. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of the present application.
Claims
1. A linking peptide, characterized in that, the linking peptide is selected from the following group: (a) a polypeptide having any one of the amino acid sequences shown in SEQ ID NO: 1-36; (b) a polypeptide having a homology (or identity) of ≥30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with any one of the amino acid sequences shown in SEQ ID NO: 1-36, and the polypeptide has the biological function of SEQ ID NO: 1-36; (c) a derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acid residues in any one of the amino acid sequences shown in SEQ ID NO: 1-36, and retaining the biological function of SEQ ID NO: 1-36.
2. A fusion protein, characterized in that, the fusion protein comprises the linking peptide according to claim 1 and one or more functional domains, or comprises the linking peptide according to claim 1 and at least one nuclease.
3. An isolated polynucleotide encoding the linking peptide according to claim 1 or the fusion protein according to claim 2.
4. A vector, characterized in that, the vector contains the polynucleotide according to claim 3.
5. A CRISPR-Cas complex, characterized in that, the complex comprises: (i) a protein component, selected from the following group: the fusion protein according to claim 2; and (ii) a nucleic acid component, selected from the following group: guide RNA, a nucleic acid encoding the guide RNA, a precursor RNA of the guide RNA, a nucleic acid encoding the precursor RNA of the guide RNA, or a combination thereof; wherein the protein component and the nucleic acid component bind to each other to form a complex.
6. A CRISPR-Cas composition, characterized in that, it comprises: (i) a first component, selected from the following group: the fusion protein according to claim 2, a nucleotide sequence encoding the fusion protein according to claim 2, and any combination thereof; and (ii) a second component, the second component being one or more guide RNAs, or a nucleotide sequence encoding the guide RNA; the guide RNA can form a complex with the fusion protein in (i).
7. A CRISPR-Cas system, characterized in that, it comprises one or more vectors, and the one or more vectors comprise: (i) a first nucleic acid, which is a nucleotide sequence encoding the fusion protein according to claim 2; optionally the first nucleic acid is operably linked to a first regulatory element; and (ii) a second nucleic acid, which is a nucleotide sequence encoding the guide RNA; optionally the second nucleic acid is operably linked to a second regulatory element; wherein: the first nucleic acid and the second nucleic acid are present on the same or different vectors; The guide RNA is capable of forming a complex with the fusion protein described in (i).
8. A kit, characterized in that, it comprises one or more components selected from the following: the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7.
9. A delivery composition, characterized in that, it contains a delivery vector and one or more selected from the following: the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7.
10. A host cell, characterized in that, it contains the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9.
11. An enzyme preparation, characterized in that, the enzyme preparation comprises the fusion protein according to claim 2, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9.
12. A medicine box, characterized in that, it comprises: a first container, and the complex according to claim 5, or the CRISPR-Cas composition according to claim 6, or the system according to claim 7 located in the first container, or a drug containing the complex according to claim 5, or the CRISPR-Cas composition according to claim 6, or the system according to claim 7.
13. A medicine box, characterized in that, it comprises: (a1) a first container, and the fusion protein according to claim 2, or its coding gene or its expression vector located in the first container, or a drug containing the fusion protein according to claim 2, or its coding gene or its expression vector; (b1) optionally a second container, and the guide RNA or its expression vector located in the second container, or a drug containing the guide RNA or its expression vector.
14. A method for targeting and editing a target gene, characterized in that, it comprises: contacting the fusion protein according to claim 2, or the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9, or the enzyme preparation according to claim 11, or the medicine box according to claim 12 or 13 with the target gene, or delivering it into a cell containing the target gene, wherein the target sequence is present in the target gene.
15. A cell or its progeny obtained by the method according to claim 14, wherein the cell contains a modification that does not exist in its wild type.
16. The cell product of the cell or its progeny according to claim 15.
17. An in vitro, ex vivo or in vivo cell or cell line or their progeny, wherein, the cell or cell line or their progeny comprises: the fusion protein according to claim 2, the polynucleotide according to claim 3, or the complex according to claim 5, the CRISPR-Cas composition according to claim 6 or the system according to claim 7 or the delivery composition according to claim 9.
18. A cell preparation, wherein, it comprises the host cell according to claim 10 or the cell according to claim 15 or their progeny, or the cell product of the cell according to claim 16 or their progeny, or the cell or cell line according to claim 17 or their progeny.
19. Use of the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6 or the system according to claim 7, the kit according to claim 8, the delivery composition according to claim 9, the enzyme preparation according to claim 11, the kit according to claim 12 or 13, wherein, for the preparation of a drug or preparation for nucleic acid editing (e.g., gene or genome editing).
20. Use of the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6 or the system according to claim 7, the kit according to claim 8, the delivery composition according to claim 9, the enzyme preparation according to claim 11, the kit according to claim 12 or 13, wherein, for the preparation of a drug or preparation for use in one or more of the following selected from the group consisting of: (i) ex vivo gene or genome editing; (ii) editing a target sequence in a target locus to modify a biological or non-human organism; (iii) treating a disorder caused by a defect in a target sequence in a target locus; (iv) treating a disorder or disease in a subject in need thereof.
Citation Information
Patent Citations
Adenosine nucleobase editors and uses thereof
US10113163B2
Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells
WO2001038547A2