Connecting peptide and application thereof

By developing new linkers with specific amino acid sequences or homology polypeptides, the difficulty in selecting linkers in fusion protein construction is solved, and the correct folding and high biological activity of fusion proteins are achieved.

CN120040556APending Publication Date: 2025-05-27YOLTECH THERAPEUTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311594041.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to select suitable linkers when constructing fusion proteins, resulting in misfolding of fusion proteins, low protein yield or impaired biological activity.

Method used

A new linking peptide was developed to ensure correct folding and high biological activity of the fusion protein by providing linkers with specific amino acid sequences or homology polypeptides.

Benefits of technology

By using these new linkers, the activity and yield of the fusion protein is significantly improved, and the problem of difficulty in selecting linkers in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004572286760000141
    Figure BDA0004572286760000141
  • Figure BDA0004572286760000231
    Figure BDA0004572286760000231
  • Figure BDA0004572286760000241
    Figure BDA0004572286760000241
Patent Text Reader

Abstract

The invention provides a connecting peptide and application thereof, particularly provides the connecting peptide and a fusion protein containing the connecting peptide, and further provides a base editor containing the connecting peptide and application of the base editor. The linker peptide provided by the invention can significantly improve the editing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biology, and in particular, to a linker peptide and its application. Background Art

[0002] When constructing fusion proteins, it is often necessary to select a suitable linker to connect polypeptide domains. Direct fusion of functional domains (without using a linker) may cause many problems, such as misfolding of the fusion protein, low protein yield, or impaired biological activity.

[0003] The activity of the fusion protein may be related to the linker. For example, when constructing base editor fusion proteins, the fusion of nuclease and deaminase domains is mostly through a linker, and the choice of linker is related to the editing efficiency of the base editor.

[0004] Therefore, there is an urgent need in the art to develop new linkers to meet the needs of constructing fusion proteins. Summary of the Invention

[0005] The purpose of the present invention is to develop new linkers to meet the needs of constructing fusion proteins.

[0006] In the first aspect of the present invention, a linker peptide is provided, and the linker peptide is selected from the following group:

[0007] (a) a polypeptide having an amino acid sequence shown in any one of SEQ ID NO: 1-9;

[0008] (b) a polypeptide having a homology (or identity) of ≥20%, 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with the amino acid sequence shown in any one of SEQ ID NO: 1-9, and the polypeptide has the biological function of SEQ ID NO: 1-9;

[0009] (c) a derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acid residues in the amino acid sequence shown in any one of SEQ ID NO: 1-9, and retaining the biological function of SEQ ID NO: 1-9.

[0010] In another preferred embodiment, the linker peptide is synthetic.

[0011] In another preferred example, compared with the amino acid sequence shown in any of SEQ ID NO: 1-9, the linker peptide has a sequence with one or more base substitutions, deletions or additions, and substantially retains the biological function of the sequence from which it is derived.

[0012] In another preferred example, the linker peptide has a polypeptide with a homology (or identity) of ≥40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with the amino acid sequence shown in any of SEQ ID NO: 1-2, 8-9, and the polypeptide has the biological function of SEQ ID NO: 1-2, 8-9.

[0013] In another preferred example, the linker peptide has a polypeptide with a homology (or identity) of ≥60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% or 100% with the amino acid sequence shown in any of SEQ ID NO: 4-6, and the polypeptide has the biological function of SEQ ID NO: 4-6.

[0014] In another preferred example, the linker peptide is selected from the polypeptides with the amino acid sequence shown in SEQ ID NO: 1-2.

[0015] In another preferred example, the linker peptide is selected from the polypeptides with the amino acid sequence shown in SEQ ID NO: 8-9.

[0016] The second aspect of the present invention provides a fusion protein, which includes the linker peptide described in the first aspect of the present invention and one or more functional domains, or includes the linker peptide described in the first aspect of the present invention and at least one nuclease.

[0017] In another preferred example, the functional domain is selected from the following group: Cas protein targeting part, DNA binding domain, transcriptional activation domain, transcriptional repression domain, nuclease, deamination domain, methylase, demethylase, transcriptional release factor, HDAC, cleavage active polypeptide, ligase, integrase, transposase, recombinase, polymerase and base excision repair inhibitor (such as uracil-DNA glycosylase inhibitor (UGI)), or a combination thereof.

[0018] In one embodiment, the fusion protein includes a nuclease, the linker peptide described in the first aspect of the present invention and one or more deamination domains.

[0019] In some specific embodiments, the deamination domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0020] In some specific embodiments, the adenosine deaminase catalytic domain is adenosine deaminase.

[0021] In some specific embodiments, the adenine deaminase (or adenosine deaminase) is any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).

[0022] In some specific embodiments, adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively.

[0023] In some specific embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in DNA.

[0024] In some specific embodiments, the adenine deaminase can be any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).

[0025] In some specific embodiments, the adenosine deaminase has an amino acid sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100% sequence identity to any one of SEQ ID NOs: 60 - 62.

[0026] In some specific embodiments, the deamination domain is a cytidine deaminase catalytic domain.

[0027] In some specific embodiments, the cytidine deaminase catalytic domain includes an APOBEC deaminase.

[0028] In some specific embodiments, the cytidine deaminase can be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570), and / or variant versions thereof.

[0029] In some specific embodiments, the cytidine deaminase is APOBEC1 deaminase having the amino acid sequence of SEQ ID NO:63.

[0030] In some specific embodiments, the cytidine deaminase can be APOBEC3A deaminase having the amino acid sequence of SEQ ID NO:64.

[0031] In some specific embodiments, the cytidine deaminase can have about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence of a naturally occurring cytidine deaminase.

[0032] In some specific embodiments, the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 100% identity to the amino acid sequence shown in SEQ ID NO:63 or SEQ ID NO:64.

[0033] In some embodiments, the polynucleotide encoding the cytidine deaminase can be codon-optimized for expression in an organism, and the codon-optimized polypeptide can have about 70% to 99.5% identity to a reference polynucleotide.

[0034] In another preferred example, the nuclease is selected from a nucleic acid programmable nucleotide-binding protein (napDNAbp) or its homolog.

[0035] In another preferred example, the nuclease is selected from RNA-guided nucleic acid programmable nucleotide-binding proteins, homing endonucleases (Meganuclease), zinc finger fusion proteins (ZFN), or TALEN, or homologs thereof.

[0036] In another preferred example, the nuclease is selected from type II CRISPR-Cas polypeptides, type I CRISPR-Cas polypeptides, type III CRISPR-Cas polypeptides, type IV CRISPR-Cas polypeptides, type V CRISPR-Cas polypeptides, type VI CRISPR-Cas polypeptides, type VII CRISPR-Cas polypeptides, IscB polypeptides, TnpB polypeptides, IsrB polypeptides, or homologs thereof.

[0037] In another preferred example, the nuclease is selected from Cas9, CasX, CasY, Cas12a (Cpf1), Cas12b (C2cl), Cas13a (C2c2), Cas12c (C2c3), Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Argonaute (Ago), or homologs thereof.

[0038] In another preferred example, the nuclease has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with the amino acids shown in any one of SEQ ID NOs: 56-59.

[0039] In another preferred example, the fusion protein comprises the following elements fused together:

[0040] (a) A first polypeptide domain;

[0041] (b) The linker peptide as claimed in claim 1;

[0042] (c) A second polypeptide domain.

[0043] In some specific embodiments, the first polypeptide domain is linked to the linker peptide described in the first aspect of the present invention at its C-terminus.

[0044] In some specific embodiments, the first polypeptide domain is linked to the linker peptide described in the first aspect of the present invention at its N-terminus.

[0045] In some specific embodiments, the first polypeptide domain is selected from nucleic acid programmable nucleotide-binding proteins (napDNAbp) or homologs thereof.

[0046] In some specific embodiments, the first polypeptide domain is selected from an RNA-guided nucleic acid programmable nucleotide-binding protein, a homing endonuclease (Meganuclease), a zinc finger fusion protein (ZFN), or a TALEN, or a homolog thereof.

[0047] In some specific embodiments, the first polypeptide domain is selected from a type II CRISPR-Cas polypeptide, a type I CRISPR-Cas polypeptide, a type III CRISPR-Cas polypeptide, a type IV CRISPR-Cas polypeptide, a type V CRISPR-Cas polypeptide, a type VI CRISPR-Cas polypeptide, a type VII CRISPR-Cas polypeptide, an IscB polypeptide, a TnpB polypeptide, an IsrB polypeptide, or a homolog thereof.

[0048] In some specific embodiments, the first polypeptide domain is selected from Cas9, CasX, CasY, Cas 12a (Cpf1), Cas 12b (C2cl), Cas 13a (C2c2), Cas 12c (C2c3), Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Argonaute (Ago), or a homolog thereof.

[0049] In some specific embodiments, the first polypeptide domain has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with the amino acid shown in any one of SEQ ID NOs: 56-59.

[0050] In some specific embodiments, the second polypeptide domain is selected from a localization signal, a reporter protein, a CRISPR-Cas effector protein targeting moiety, a DNA-binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage activity polypeptide, a ligase.

[0051] In another preferred example, the second polypeptide domain includes a deamination domain.

[0052] In some specific embodiments, the deamination domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0053] In some specific embodiments, the adenosine deaminase catalytic domain is adenosine deaminase.

[0054] In some specific embodiments, the adenine deaminase (or adenosine deaminase) is any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).

[0055] In some specific embodiments, the adenine deaminase can catalyze the hydrolysis and deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively.

[0056] In some specific embodiments, the adenosine deaminase catalyzes the hydrolysis and deamination of adenine or adenosine in DNA.

[0057] In some specific embodiments, the adenine deaminase can be any known or later identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure regarding adenine deaminase).

[0058] In some specific embodiments, the adenosine deaminase has an amino acid sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100% sequence identity to any one of SEQ ID NOs: 60 - 62.

[0059] In some specific embodiments, the deamination domain is a cytidine deaminase catalytic domain.

[0060] In some specific embodiments, the cytidine deaminase catalytic domain includes an APOBEC deaminase.

[0061] In some specific embodiments, the cytidine deaminase can be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570) and / or variant versions thereof.

[0062] In some specific embodiments, the cytidine deaminase is APOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 63.

[0063] In some specific embodiments, the cytidine deaminase is the APOBEC3A deaminase having the amino acid sequence of SEQ ID NO: 64.

[0064] In some specific embodiments, the amino acid sequence of the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity with the amino acid sequence of the wild-type cytidine deaminase (such as POBEC1 deaminase, APOBEC3A deaminase).

[0065] In some specific embodiments, the amino acid sequence of the cytidine deaminase has about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 100% identity with the amino acid sequence shown in SEQ ID NO: 63 or SEQ ID NO: 64.

[0066] In some embodiments, the polynucleotide encoding the cytidine deaminase can be codon-optimized for expression in an organism, and the codon-optimized polypeptide can have about 70% to 99.5% identity with the reference polynucleotide.

[0067] In another preferred example, the fusion protein has the structure shown in Formula I or II below:

[0068] X-Z-Y (I)

[0069] Y-Z-X (II);

[0070] In the formula,

[0071] X is the first polypeptide domain;

[0072] Y is the second polypeptide domain;

[0073] Z is the linker peptide as claimed in claim 1;

[0074] "-" represents a peptide bond or a peptide linker connecting the above elements.

[0075] In another preferred example, any two of X, Y, and Z are connected in a head-to-head, head-to-tail, tail-to-head, or tail-to-tail manner.

[0076] In another preferred embodiment, the "head" refers to the N-terminus of the polypeptide domain.

[0077] In another preferred embodiment, the "tail" refers to the C-terminus of the polypeptide domain.

[0078] In another preferred embodiment, the length of the peptide linker is 0 - 20 amino acids, preferably 0 - 10 amino acids.

[0079] In another preferred embodiment, the fusion protein is selected from the group consisting of:

[0080] (A) a polypeptide having the amino acid sequence shown in any one of SEQ ID NOs: 77 - 85;

[0081] (B) a polypeptide having at least 80% homology (preferably at least 90% homology; more preferably at least 95% homology; most preferably at least 97% homology, such as above 98%, above 99%) with the amino acid sequence shown in any one of SEQ ID NOs: 77 - 85, and the polypeptide has gene editing activity;

[0082] (C) a derivative polypeptide formed by substituting, deleting or adding 1 - 5 amino acid residues to the amino acid sequence shown in any one of SEQ ID NOs: 77 - 85, and retaining gene editing activity.

[0083] The third aspect of the present invention provides an isolated polynucleotide encoding the linker peptide described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention.

[0084] In certain embodiments, the isolated polynucleotide includes a sequence optimized by humanization.

[0085] In certain embodiments, the polynucleotide further contains an auxiliary element selected from the group consisting of: a signal peptide, a secretion peptide, a tag sequence (such as 6His), or a combination thereof, flanking the ORF of the fusion protein.

[0086] In certain embodiments, the polynucleotide is selected from the group consisting of: a genomic sequence, a cDNA sequence, an RNA sequence, or a combination thereof.

[0087] In certain embodiments, the polynucleotide further contains a promoter operably linked to the ORF sequence of the fusion protein.

[0088] In certain embodiments, the promoter is selected from the group consisting of: a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter.

[0089] In certain embodiments, the polynucleotide is a polynucleotide codon-optimized according to the codon preference of the host cell.

[0090] In certain embodiments, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0091] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).

[0092] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.

[0093] In certain embodiments, the yeast cell is selected from yeast of one or more sources in the following group: Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.

[0094] In certain embodiments, the host cell is selected from the following group: Escherichia coli, wheat germ cell, insect cell, SF9, Hela, HEK293, CHO, yeast cell, or a combination thereof.

[0095] The fourth aspect of the present invention provides a vector, and the vector contains the polynucleotide described in the third aspect of the present invention.

[0096] In certain embodiments, the vector contains one or more promoters, and the promoter is operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, replication origin, selection marker, nucleic acid restriction site, and / or homologous recombination site.

[0097] In certain embodiments, the vector includes a plasmid, a viral vector.

[0098] In certain embodiments, the viral vector is selected from the following group: adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or a combination thereof.

[0099] In certain embodiments, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, a multifunctional vector.

[0100] The fifth aspect of the present invention provides a CRISPR-Cas complex, and the complex includes:

[0101] (i) a protein component, selected from the following group: the fusion protein described in the second aspect of the present invention; and

[0102] (ii) A nucleic acid component selected from the group consisting of: guide RNA, a nucleic acid encoding the guide RNA, a precursor RNA of the guide RNA, a nucleic acid of the precursor RNA encoding the guide RNA, or a combination thereof;

[0103] Wherein, the protein component and the nucleic acid component are combined with each other to form a complex.

[0104] The sixth aspect of the present invention provides a CRISPR-Cas composition, comprising:

[0105] (i) A first component selected from the group consisting of: the fusion protein described in the second aspect of the present invention, a nucleotide sequence encoding the fusion protein described in the second aspect of the present invention, and any combination thereof; and

[0106] (ii) A second component, which is one or more guide RNAs, or a nucleotide sequence encoding the guide RNA;

[0107] The guide RNA is capable of forming a complex with the fusion protein described in (i).

[0108] In certain embodiments, the guide RNA comprises a direct repeat sequence and a spacer sequence from the 5' to 3' direction, and the spacer sequence is capable of hybridizing with a target sequence.

[0109] In certain embodiments, the composition further comprises a pharmaceutically acceptable carrier.

[0110] In certain embodiments, the composition comprises a pharmaceutical composition.

[0111] In certain embodiments, the dosage form of the composition is selected from the group consisting of: freeze-dried preparation, liquid preparation, or a combination thereof.

[0112] In certain embodiments, the dosage form of the composition is a liquid preparation.

[0113] In certain embodiments, the dosage form of the composition is an injection dosage form.

[0114] In certain embodiments, the composition is a cell preparation.

[0115] The seventh aspect of the present invention provides a CRISPR-Cas system, comprising one or more vectors, and the one or more vectors comprise:

[0116] (i) A first nucleic acid, which is a nucleotide sequence encoding the fusion protein described in the second aspect of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0117] (ii) A second nucleic acid, which is a nucleotide sequence encoding a guide RNA; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0118] Wherein:

[0119] The first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0120] The guide RNA is capable of forming a complex with the fusion protein described in (i).

[0121] In certain embodiments, the vector includes a plasmid, a viral vector.

[0122] In certain embodiments, the guide RNA includes a spacer sequence capable of hybridizing with a target sequence; and a direct repeat (DR) sequence that is linked to the spacer sequence and is capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.

[0123] In certain embodiments, the guide RNA includes unmodified and modified guide RNAs.

[0124] In certain embodiments, the modified guide RNA includes chemical modifications of bases.

[0125] In certain embodiments, the chemical modification includes methylation modification, methoxy modification, fluorination modification or thiolation modification.

[0126] In certain embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0127] In certain embodiments, at least one component in the composition is non-naturally occurring or modified.

[0128] In certain embodiments, the spacer sequence is linked to the 3' end of the direct repeat (DR) sequence.

[0129] In certain embodiments, the spacer sequence contains a complementary sequence of the target sequence.

[0130] In certain embodiments, the target sequence is a DNA from a prokaryotic cell or a eukaryotic cell or a DNA sequence formed by reverse transcription based on RNA; alternatively, the target sequence is a non-naturally occurring DNA or a DNA sequence formed by reverse transcription based on RNA.

[0131] In certain embodiments, the target sequence includes a cDNA sequence.

[0132] In some embodiments, the target sequence comprises single-stranded DNA, double-stranded DNA sequences.

[0133] In some embodiments, the target sequence is present within a cell.

[0134] In some embodiments, the target sequence is present within the cell nucleus or within the cytoplasm (e.g., an organelle).

[0135] In some embodiments, the cell is a eukaryotic cell.

[0136] In some embodiments, the cell is a prokaryotic cell.

[0137] In some embodiments, the target sequence is present outside the cell.

[0138] In some embodiments, the fusion protein comprises one or more NLS sequences.

[0139] In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the fusion protein described in the second aspect of the present invention.

[0140] In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the fusion protein described in the second aspect of the present invention.

[0141] The eighth aspect of the present invention provides a kit comprising one or more components selected from the following: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.

[0142] In some embodiments, the kit further comprises a label or instructions.

[0143] In some embodiments, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, cleaving a target gene or a non-target gene.

[0144] The ninth aspect of the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.

[0145] In some embodiments, the delivery vector is a particle.

[0146] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microbubbles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0147] The tenth aspect of the present invention provides a host cell comprising the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.

[0148] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).

[0149] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.

[0150] In certain embodiments, the yeast cell is selected from one or more sources of yeast in the following group: Pichia pastoris, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.

[0151] In certain embodiments, the host cell is selected from the following group: Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

[0152] The eleventh aspect of the present invention provides an enzyme preparation, which includes the fusion protein described in the second aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.

[0153] In certain embodiments, the enzyme preparation includes an injection and / or a freeze-dried preparation.

[0154] The twelfth aspect of the present invention provides a kit, comprising:

[0155] A first container, and in the first container, the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or a drug containing the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.

[0156] In certain embodiments, the drug in the first container is a single-agent formulation containing the complex described in the fifth aspect of the present invention, or the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention.

[0157] In certain embodiments, the dosage form of the drug is selected from the group consisting of: freeze-dried preparations, liquid preparations, or combinations thereof.

[0158] In certain embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0159] In certain embodiments, the kit further contains instructions.

[0160] The thirteenth aspect of the present invention provides a kit, comprising:

[0161] (a1) a first container, and the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector, or a drug containing the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector, located in the first container;

[0162] (b1) an optional second container, and the guide RNA or its expression vector, or a drug containing the guide RNA or its expression vector, located in the second container.

[0163] In certain embodiments, the first container and the second container are different containers.

[0164] In certain embodiments, the drug in the first container is a single-agent formulation containing the fusion protein described in the second aspect of the present invention, or its coding gene or its expression vector.

[0165] In certain embodiments, the drug in the second container is a single-agent formulation containing the guide RNA or its expression vector.

[0166] In certain embodiments, the dosage form of the drug is selected from the group consisting of: freeze-dried preparations, liquid preparations, or combinations thereof.

[0167] In certain embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0168] In certain embodiments, the kit further contains instructions.

[0169] The fourteenth aspect of the present invention provides a method for targeting and editing a target gene, comprising: contacting the fusion protein described in the second aspect of the present invention, or the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention, or the enzyme preparation described in the eleventh aspect of the present invention, or the kit described in the twelfth aspect or the thirteenth aspect of the present invention with the target gene, or delivering it into a cell containing the target gene, wherein the target sequence is present in the target gene.

[0170] In certain embodiments, the target gene is present within a cell.

[0171] In certain embodiments, the cell is a prokaryotic cell.

[0172] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.

[0173] In certain embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro.

[0174] In certain embodiments, editing the target gene includes site-directed mutagenesis of the target sequence.

[0175] In certain embodiments, the site-directed mutagenesis includes base changes of A→G or C→T.

[0176] In certain embodiments, the target gene includes DNA, RNA.

[0177] In certain embodiments, the DNA includes single-stranded DNA, double-stranded DNA.

[0178] In certain embodiments, the RNA includes single-stranded RNA, double-stranded RNA.

[0179] The fifteenth aspect of the present invention provides a cell or its progeny obtained by the method described in the fourteenth aspect of the present invention, wherein the cell contains a modification that does not exist in its wild type.

[0180] In another preferred example, the wild-type cell contains an abnormality, and the abnormality of the cell modified by the method has been resolved or corrected.

[0181] The sixteenth aspect of the present invention provides a cell product of the cell or its progeny described in the fifteenth aspect of the present invention.

[0182] A seventeenth aspect of the present invention provides an in vitro, ex vivo or in vivo cell or cell line or their progeny, said cell or cell line or their progeny comprising: the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, or the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, or the delivery composition described in the ninth aspect of the present invention.

[0183] In certain embodiments, the cell is a prokaryotic cell.

[0184] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.

[0185] In certain embodiments, the cell is a stem cell or a stem cell line.

[0186] An eighteenth aspect of the present invention provides a cell preparation, comprising the host cell described in the tenth aspect of the present invention, or the cell or its progeny described in the fifteenth aspect of the present invention, or the cell product of the cell or its progeny described in the sixteenth aspect of the present invention, or the cell or cell line or their progeny described in the seventeenth aspect of the present invention.

[0187] In another preferred example, the cell preparation further comprises a pharmaceutically acceptable carrier or excipient.

[0188] In another preferred example, the cell preparation comprises an injection and / or a freeze-dried preparation.

[0189] A nineteenth aspect of the present invention provides the use of the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention, or the system described in the seventh aspect of the present invention, the kit described in the eighth aspect of the present invention, the delivery composition described in the ninth aspect of the present invention, the enzyme preparation described in the eleventh aspect of the present invention, or the kit described in the twelfth or thirteenth aspect of the present invention, for preparing a drug or preparation for nucleic acid editing (e.g., gene or genome editing).

[0190] In certain embodiments, the gene or genome editing includes site-directed mutagenesis.

[0191] In certain embodiments, the site-directed mutagenesis includes a base change of A→G or C→T.

[0192] The twentieth aspect of the present invention provides the use of the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the complex described in the fifth aspect of the present invention, the CRISPR-Cas composition described in the sixth aspect of the present invention or the system described in the seventh aspect of the present invention, the kit described in the eighth aspect of the present invention, the delivery composition described in the ninth aspect of the present invention, the enzyme preparation described in the eleventh aspect of the present invention, the medicine box described in the twelfth aspect or the thirteenth aspect of the present invention, for preparing a drug or a preparation, and the drug or the preparation is used for one or more selected from the following groups:

[0193] (i) Ex vivo gene or genome editing;

[0194] (ii) Editing a target sequence in a target locus to modify a biological or non-human organism;

[0195] (iii) Treating a disease caused by a defect in a target sequence in a target locus;

[0196] (iv) Treating a disease or disorder of a subject in need thereof.

[0197] In certain embodiments, the disease or disorder is selected from the group consisting of Angelman syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, Pompe disease, myoclonic muscular dystrophy, Huntington's disease (HTT), fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, hemoglobinopathy (e.g., β-thalassemia, sickle cell disease), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, hypercholesterolemia, heart disease (e.g., hypertrophic cardiomyopathy (HCM)) and cancer.

[0198] In certain embodiments, the disease or disorder includes cancer, infectious disease, metabolic disease, neurological disease.

[0199] In certain embodiments, the disease or disorder includes a disease caused by a single-base mutation.

[0200] In certain embodiments, the disease or disorder includes one or more of hypercholesterolemia, transthyretin amyloidosis, and β-hemoglobinopathy.

[0201] In certain embodiments, the disorder or disease is caused by a pathogenic point mutation.

[0202] It should be understood that within the scope of the present invention, the above-described technical features of the present invention and the technical features specifically described hereinafter (such as in the examples) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS

[0203] Figure 1 Shows the plasmid map of the dual-luciferase reporter system.

[0204] Figure 2 Shows the detection and comparison of the editing efficiency of base editors with different linkers using the dual-luciferase reporter system.

[0205] Figure 3 Shows the editing efficiency of base editors with different linkers at positions A5 and A7 of site1.

[0206] Figure 4 Shows the base editing efficiency of base editors with different linkers in the PCSK9 gene. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0207] Through extensive and in-depth research and a large number of screenings, the present invention has for the first time discovered a new linker and provided a fusion protein, a CRISPR-Cas system, and a complex containing the linker. Also provided are base editors containing the linker, such as fusing DNA editing proteins such as deaminases (e.g., adenosine deaminase or cytidine deaminase) with nucleases (e.g., Cas9, Cas12, Cas13, etc.) to form a fusion protein for base modification (e.g., A-to-G) editing of target DNA or target RNA. The present invention also provides the applications of the fusion protein, the CRISPR-Cas system or complex containing the fusion protein, and the method for editing or modifying target nucleic acids, such as for gene editing of disease targets to achieve the prevention and treatment of diseases. The linker of the present invention can significantly improve the editing efficiency. On this basis, the inventors have completed the present invention.

[0208] TERMS

[0209] The following examples are only used to describe the present invention and do not limit the present invention. Unless otherwise specified, the experiments and methods described in the examples are basically carried out according to the conventional methods well-known in the art and described in various references.

[0210] In addition, for those conditions not specified in the examples, they shall be carried out according to conventional conditions or conditions recommended by the manufacturer. For reagents or instruments whose manufacturers are not specified, they are all conventional products that can be obtained through commercial purchase. Those skilled in the art know that the examples describe the present invention by way of illustration and are not intended to limit the scope of the present invention claimed. All the published cases and other reference materials mentioned herein are incorporated herein by reference in their entirety.

[0211] To make the present disclosure more easily understood, certain terms are first defined. As used in this application, unless otherwise clearly specified herein, each of the following terms shall have the meanings given below. Other definitions are set forth throughout the application.

[0212] The term "about" can refer to a value or a composition within an acceptable error range of a specific value or composition determined by a person of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0213] As used herein, the terms "comprising" or "including (containing)" can be open-ended, semi-closed, and closed. In other words, the terms also include "consisting essentially of..." or "consisting of...".

[0214] Sequence identity (or homology) is determined by comparing two aligned sequences along a predefined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions where identical residues occur. Generally, this is expressed as a percentage. Methods for measuring the sequence identity of nucleotide sequences are well known to those skilled in the art.

[0215] General definitions

[0216] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by a person of ordinary skill in the art to which the present invention pertains. The terms used in the description of the present invention are only for describing specific embodiments and are not intended to limit the present invention.

[0217] Unless otherwise indicated by the context, the various features described in the present invention can be used in any combination. In addition, the present invention also contemplates that in some embodiments of the present invention, any feature or combination of features set forth herein can be excluded or omitted. For purposes of illustration, if the specification states that the composition contains components 1, 2, and 3, then any one of components 1, 2, 3 or any combination thereof can be omitted and waived individually or in any combination.

[0218] As used in the description of the present invention and the appended claims, the articles "a / an" and "the" are used in the present invention to refer to one or more than one (i.e., at least one) of the grammatical objects of the said articles. For example, "an element" means one element or more than one element, unless the context clearly indicates otherwise.

[0219] The use of alternatives (such as "or") should be understood to mean any one, both, or any combination of the alternatives. The term "and / or" should be understood to mean either or both of the alternatives.

[0220] As used in the present invention, when referring to measurable values such as amounts or concentrations, etc., the term "about" means encompassing variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value, as well as the specified value. For example, "about X" (where X is a measurable value) means including X, as well as variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. The ranges provided for measurable values in the present invention may include any other ranges and / or individual values therein.

[0221] As used in the present invention, the terms "comprising", "containing", and "including" specifically state the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0222] As used in the present invention, the transitional phrase "consisting essentially of" means that the scope of the claim will be interpreted to cover the specified materials or steps recited in the claim, as well as those that do not substantially affect the basic and novel features of the present invention being claimed. Thus, when used in the claims of the present invention, the term "consisting essentially of" is not intended to be interpreted as equivalent to "comprising".

[0223] As used in the present invention, the terms "increase" and "enhance" (and their grammatical variations) describe an improvement of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to a control.

[0224] As used herein, the terms "reduce", "decrease", and "lessen" (and grammatical variations thereof) describe, for example, at least about a 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% decrease compared to a control. In particular embodiments, the reduction may result in no or substantially no (i.e., a non-significant amount, such as less than about 10% or even 5%) detectable activity or amount.

[0225] Throughout this specification, reference to "one embodiment", "an embodiment", "a particular embodiment", "a related embodiment", "some embodiments", "a certain embodiment", "an additional embodiment", "a further embodiment", or combinations thereof means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Thus, the appearances of the foregoing phrases throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0226] "Sequence identity" and "sequence similarity" between two polypeptide or nucleic acid sequences represent the percentage of the number of identical residues between the sequences out of the total number of residues, and the calculation of the total number of residues is determined based on the type of mutation. Types of mutations include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / replacements of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence. Taking polypeptides as an example (nucleotides are the same), if the type of mutation is one or more of the following: substitutions / replacements of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence, then the total number of residues is calculated as the larger of the molecules being compared. If the type of mutation also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, then the number of amino acids inserted or deleted at either or both ends (e.g., the number inserted or deleted at both ends is less than 20) is not counted in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps in the alignment (if any) are resolved by a specific algorithm.

[0227] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.

[0228] "Natural" or "wild-type" nucleic acids, nucleotide sequences, polypeptides or amino acid sequences refer to nucleic acids, nucleotide sequences, polypeptides or amino acid sequences that occur naturally or are endogenous. Thus, for example, "wild-type mRNA" is mRNA that occurs naturally in an organism or is endogenous to the organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.

[0229] Those skilled in the art will appreciate that the structure of a protein can be altered without adversely affecting its activity and functionality, for example by introducing one or more conservative amino acid substitutions in the amino acid sequence of the protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule.

[0230] Those skilled in the art are aware of examples and embodiments of conservative amino acid substitutions. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group as the site to be replaced, i.e., a nonpolar amino acid residue is replaced with another nonpolar amino acid residue, a polar uncharged amino acid residue is replaced with another polar uncharged amino acid residue, a basic amino acid residue is replaced with another basic amino acid residue, and an acidic amino acid residue is replaced with another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, conservative substitutions in which one amino acid is replaced with another amino acid belonging to the same group fall within the scope of the present invention. Thus, the proteins of the present invention may contain one or more conservative substitutions in the amino acid sequence, and these conservative substitutions are preferably made according to Table A. In addition, the present invention also encompasses proteins that further contain one or more other non-conservative substitutions, as long as such non-conservative substitutions do not significantly affect the desired functions and biological activities of the proteins of the present invention.

[0231] Conservative amino acid substitutions can be made at one or more predicted non-essential amino acid residues. A "non-essential" amino acid residue is an amino acid residue that can be changed (deleted, substituted or replaced) without changing biological activity, while an "essential" amino acid residue is required for biological activity. A "conservative amino acid substitution" is a substitution in which an amino acid residue is replaced with an amino acid residue having a similar side chain. Amino acid substitutions can be made in non-conserved regions of the Cas enzyme. Generally, such substitutions are not made on conserved amino acid residues or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will appreciate that functional variants may have fewer conservative or non-conservative changes in conserved regions.

[0232] Table A

[0233]

[0234] Those skilled in the art are aware that changing (substituting, deleting, truncating or inserting) one or more amino acid residues from the N and / or C termini of a protein can still retain its functional activity. Therefore, proteins in which one or more amino acid residues have been changed from the N and / or C termini of the Cas proteins of the present invention while retaining their required functional activity are also within the scope of the present invention. These changes can include those introduced by modern molecular methods such as PCR, which include PCR amplification that changes or extends the protein coding sequence by means of including amino acid coding sequences among the oligonucleotides used in the PCR amplification.

[0235] It should be recognized that proteins can be modified in various ways, including amino acid substitution, deletion, truncation and insertion, and the methods for such operations are generally known to those skilled in the art.

[0236] Linker / Adapter / Linking Peptide

[0237] The terms "linker", "adapter", "linking peptide" and "adapter polypeptide" are used interchangeably and all refer to a polypeptide amino acid sequence of two or more amino acids that covalently links a first protein or polypeptide to a second protein polypeptide.

[0238] The present invention provides a variety of new linkers. For example, the linker can be Linker1 (SEQ ID NO:1), Linker2 (SEQ ID NO:2), Linker3 (SEQ ID NO:3), Linker4 (SEQ ID NO:4), Linker5 (SEQ ID NO:5), Linker6 (SEQ ID NO:6), Linker7 (SEQ ID NO:7), Linker8 (SEQ ID NO:8), Linker9 (SEQ ID NO:9).

[0239] In some specific embodiments, the linker has an amino acid sequence having at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence shown in any one of SEQ ID NOs: 1-9.

[0240] In some embodiments, the linker has a sequence with one or more base substitutions, deletions or additions compared to the amino acid sequence shown in any one of SEQ ID NOs: 1-9, and substantially retains the biological function of the sequence from which it is derived.

[0241] Cas Protein

[0242] In the present invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably. Cas protein is taken in its broadest sense and includes wild-type Cas protein, its derivatives or variants, analogs, and its functional fragments such as oligonucleotide-binding fragments.

[0243] In some embodiments, the Cas protein comprises an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence shown in any one of SEQ ID NO: 56-59.

[0244] In some embodiments, the Cas protein comprises the amino acid sequence shown in any one of SEQ ID NO: 56-59.

[0245] Orthologue

[0246] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or similar function as the protein of which it is an orthologue.

[0247] The nucleic acid cleavage of the present invention includes: DNA or RNA cleavage in the target nucleic acid generated by the Cas protein (Cis cleavage), and DNA or RNA cleavage in the collateral nucleic acid substrate (single-stranded nucleic acid substrate) caused by the collateral activity of the Cas protein (i.e., non-specific or non-targeted, Trans cleavage). In some specific embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0248] Trans cleavage refers to the situation where, in certain environments, the activated Cas12 family proteins remain active after binding to the target sequence and continue to non-specifically cleave non-target oligonucleotides. This collateral cleavage activity can be used to detect the presence of specific target oligonucleotides using the Cas system. For example, the Cas12i system has been engineered to non-specifically cleave ssDNA or transcripts. The collateral cleavage activity is used in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, which can be used for many clinical diagnoses (Gootenberg, J.S. et al., Nucleic acid detection with CRISPR-Cas13a / C2c2. Science 356, 438-442 (2017)).

[0249] Fusion protein

[0250] As used in the present invention, the term "fusion protein" refers to a hybrid (e.g., chimeric, recombinant) polypeptide that contains protein domains from at least two different proteins. One protein domain can be located in the amino-terminal (N-terminal) portion of the fusion protein and will contain the free N-terminus of the fusion protein (e.g., the amino (NH2) group), and this protein domain of the fusion protein can be referred to as the "amino-terminal fusion protein" or "amino-terminal fusion protein domain". Similarly, one protein domain can be located in the carboxyl-terminal (C-terminal) portion of the fusion protein and will contain the free C-terminus of the fusion protein (e.g., the carboxyl (COOH) group), and this protein domain of the fusion protein can be referred to as the "carboxyl-terminal fusion protein" or "carboxyl-terminal fusion protein domain".

[0251] The present invention provides a fusion protein comprising Linker1-9 (SEQ ID NO: 1-9).

[0252] In some embodiments, the fusion protein comprises an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identity to the amino acid sequence shown in any one of SEQ ID NO: 77-85.

[0253] In some embodiments, the fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 77-85.

[0254] CRISPR system

[0255] The term "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" is used interchangeably and has the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of guiding the activity of said Cas genes.

[0256] CRISPR / Cas complex

[0257] The term "CRISPR / Cas complex" refers to a complex formed by the binding of a gRNA (guide RNA) or a mature crRNA (or guide RNA) to a Cas protein or a fusion protein containing it, which contains a direct repeat sequence hybridized to the guide sequence of the target sequence and bound to the Cas protein, and the complex is capable of recognizing and cleaving a target nucleotide that can hybridize to the guide RNA or mature crRNA.

[0258] gRNA (guide RNA)

[0259] The terms "guide RNA (gRNA)", "mature crRNA", "guide sequence", "guide RNA" are used interchangeably and have the meaning commonly understood by those skilled in the art. Generally, the guide RNA can include a direct repeat (DR) sequence and a guide sequence, or consist essentially of or consist of a direct repeat (DR) sequence and a guide sequence.

[0260] In some cases, the guide sequence is any polynucleotide sequence that has sufficient complementarity to the target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. The guide sequence contains a sequence (such as a spacer (DR) sequence) that has sufficient complementarity to the target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide the sequence-specific binding of the complex to the target nucleic acid sequence.

[0261] Target nucleic acid

[0262] In the present invention, "target nucleic acid", "target sequence", and "target nucleic acid sequence" are used interchangeably and refer to a specific nucleic acid that contains a nucleic acid sequence that is fully or partially complementary to the guide sequence in the gRNA. The "target sequence" refers to the polynucleotide targeted by the guide sequence in the gRNA, such as a sequence that is complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including the Cas protein and the gRNA). Perfect complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. In some embodiments, the target nucleic acid contains a non-coding region (e.g., a promoter or a terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded.

[0263] The target sequence can comprise any polynucleotide, such as DNA or RNA. In certain cases, the target sequence is located intracellularly or extracellularly. In certain cases, the target sequence is located within the nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts) of a cell.

[0264] The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In certain cases, the target sequence should be related to the protospacer adjacent motif (PAM).

[0265] Donor template

[0266] In the present invention, the donor template nucleic acid or donor template is used interchangeably and refers to a nucleic acid molecule that one or more cellular proteins can use to alter the structure of a target nucleic acid after the Cas protein described herein has altered the target nucleic acid.

[0267] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid or a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template nucleic acid is an exogenous nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination can be achieved using the donor template, and the recombination is homologous recombination.

[0268] Cleavage

[0269] Cleavage refers to a DNA break in the target nucleic acid generated by the Cas protein described herein. In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break.

[0270] In the present invention, the meanings of cleaving the target nucleic acid or modifying the target nucleic acid can overlap. Modifying the target nucleic acid includes not only the modification of a single nucleotide, such as site-directed mutagenesis, but also the insertion or deletion of nucleic acid fragments.

[0271] Functional domain

[0272] In the present invention, the functional domain is taken in its broadest sense and includes a protein such as an enzyme or a factor itself or a fragment / domain thereof having a specific function. A Cas protein (such as a dCas protein) is linked / associated with one or more functional domains selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage active polypeptide, a ligase, or a combination thereof. When more than one functional domain is included, the functional domains may be the same or different.

[0273] Deamination domain

[0274] In the present invention, the deamination domain includes a catalytic domain of a deaminase (such as adenosine deaminase or cytidine deaminase). As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts adenine (or the adenine moiety of a molecule) to hypoxanthine (or the hypoxanthine moiety of a molecule), as shown below. In some specific embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0275] Adenosine deaminases include, but are not limited to, members of the enzyme family called adenosine deaminases acting on RNA (ADAR), members of the enzyme family called adenosine deaminases acting on tRNA (ADAT), and other family members containing an adenosine deaminase domain (ADAD). According to the present disclosure, adenosine deaminases are capable of targeting adenine in RNA / DNA and RNA duplexes. In certain embodiments, adenosine deaminases have been modified to increase their ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0276] In some embodiments, the deaminase is a cytidine deaminase. The term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts cytidine (or the cytidine moiety of a molecule) to uracil (or the uracil moiety of a molecule). In some embodiments, the cytidine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytidine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0277] Cytidine deaminases include, but are not limited to, members of the family of enzymes known as the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases, activation-induced deaminase (AID), or cytidine deaminase 1 (CDA1). In certain embodiments, APOBEC family deaminases.

[0278] vector

[0279] A vector is a nucleic acid molecule capable of transporting another nucleic acid molecule linked thereto.

[0280] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, no free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and other diverse polynucleotides known in the art. Vectors can be introduced into host cells by transformation, transduction, or transfection such that the genetic material elements carried by them are expressed in the host cells. A vector can be introduced into a host cell and thereby produce transcripts, proteins, or peptides, including those from proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts such as nucleic acid transcripts, proteins, or enzymes). A vector can contain various elements for controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Vectors can also contain origins of replication.

[0281] Vectors include plasmids, viral vectors. A plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Viral vectors, in which virus-derived DNA or RNA sequences are present in the vector for packaging the virus, and viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by the virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.

[0282] Other vectors (e.g., non-episomal mammalian vectors) integrate into the genome of the host cell after being introduced into the host cell and are thereby replicated together with the host genome. Moreover, certain vectors are capable of directing the expression of genes operably linked to them. Such vectors are referred to as "expression vectors".

[0283] regulatory element

[0284] "Regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals, poly-U sequences), the detailed description of which can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell type-specific.

[0285] "Promoter" refers to a non-coding nucleotide sequence located upstream of a gene that can initiate the expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as in response to a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulatable promoters include, for example, promoters that are induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, phytohormones, wounding, or chemicals (such as ethanol, abscisic acid (ABA), jasmonates, salicylic acid, or safeners).

[0286] Host cell

[0287] "Host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism (e.g., a cell line) cultured as a single cell entity, which is used as a recipient for a nucleic acid (e.g., an expression vector) and includes the progeny of the original cell that has been genetically modified by the nucleic acid.

[0288] It should be understood that the progeny of a single cell may, due to natural, accidental or deliberate mutations, not necessarily have exactly the same morphology or genome as the original parental cell. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.

[0289] Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of host cell to be transformed, the desired level of expression, etc.

[0290] NLS

[0291] NLS refers to a "nuclear localization sequence" or "nuclear localization signal", which is an amino acid sequence that promotes the entry of a protein into the cell nucleus. Nuclear localization sequences are known in the art (for example, described in International PCT Application PCT / EP2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001), and this patent is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequences: KRTADGSEFESPKKKRKV, AVKRPAATKKAGQAKKKKLD, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0292] Operably linked

[0293] "Operably linked" means that a target nucleotide sequence is linked to a regulatory element in a manner that permits the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the types of these vectors can also be selected to target specific types of cells.

[0294] Complementary

[0295] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the percentage of complementarity is 50%, 60%, 70%, 80%, 90%, and 100%). "Fully complementary" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0296] The term "stringent conditions" related to hybridization refers to conditions under which a nucleic acid that is complementary to a target sequence hybridizes primarily to the target sequence and substantially not to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence hybridizes specifically to its target sequence.

[0297] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding of the bases between these nucleotide residues. The complex can include two strands that form a duplex, three or more strands that form a multi-stranded complex, a single self-hybridizing strand, or any combination of these. The hybridization reaction can constitute a step in a broader process (such as the start of PCR, or the cleavage of a polynucleotide by an enzyme). A sequence that can hybridize to a given sequence is called the "complement" of the given sequence.

[0298] Hybridization of the target sequence with the gRNA means that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or it represents that at least 12, 15, 16, 17, 18, 19, 20, or more bases of the nucleic acid sequences of the target sequence and the gRNA can be complementary paired to hybridize and form a complex.

[0299] Expression

[0300] Nucleic acid expression includes one or more of generating an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5′ capping, and / or 3′ end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.

[0301] Delivery

[0302] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein-RNA. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.

[0303] Kit

[0304] In one aspect, the present invention provides a kit, which includes the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned vector, the aforementioned CRISPR-Cas composition, the use of the aforementioned system in the preparation of a kit, and the components of the kit are in the same or different containers.

[0305] In one aspect, the present invention further provides a container that contains the aforementioned kit.

[0306] In one embodiment, the container includes a sterile container;

[0307] In one embodiment, the container includes a syringe.

[0308] In some embodiments, the kit further includes instructions for using the kit, such as instructions in more than one language. The kit may also contain one or more reagents for use in the process using one or more of the above components. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction or storage buffers. The above reagents can be provided in a form that requires the addition of one or more other components before use (e.g., in a concentrated or lyophilized form); the buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. The buffer can have a suitable pH value (pH), for example, it can be alkaline. In some embodiments, the pH of the buffer is between about 7 and 10.

[0309] Treatment

[0310] "Treatment" refers to treating or curing a subject's disorder, delaying the onset of symptoms of the disorder, and / or delaying the severity of the disorder. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals such as bovines, equines, ovines, porcines, canines, felines, lagomorphs (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder caused by a disease-related gene defect). A "plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developmental stage.

[0311] In one aspect, the present invention also provides the use of the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, complex, system, kit, delivery composition, enzyme preparation, host cell in the preparation of a medicament for treating a disorder or disease of a subject in need thereof.

[0312] In one embodiment, the use includes administering the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, complex, system, kit, delivery composition, enzyme preparation to the subject or an isolated cell of the subject.

[0313] In one embodiment, the disorder or disease includes a disease caused by a single-base mutation.

[0314] In one embodiment, the disorder or disease is selected from Angelman syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, Pompe disease, myoclonic dystrophy, Huntington's disease (HTT), fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, hemoglobinopathy (e.g., β-thalassemia, sickle cell disease), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, hypercholesterolemia, heart disease (e.g., hypertrophic cardiomyopathy (HCM)), and cancer.

[0315] In one embodiment, the clinical variant database is obtained from the NCBI ClinVar database available on the NCBI ClinVar website. Pathogenic single nucleotide polymorphisms (SNPs) are identified from this list. Using genomic locus information, CRISPR targets in the regions overlapping and surrounding each SNP are identified. The selection of SNPs that can be corrected using base editing in combination with a Cas protein or its variant to target the causal mutation is listed in the table below, with only one alias for each disease listed in the table. "RS#" corresponds to the RS accession number in the SNP database on the NCBI website. "AlleleID" corresponds to the causal allele accession number. The "Name" column contains the locus identifier of the gene, the gene name, the position of the mutation in the gene, and the change caused by the mutation.

[0316] Disease targets of base editing

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328]

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335]

[0336]

[0337]

[0338] In one embodiment, the disorder or disease includes cancer, infectious disease, metabolic disease, and neurological disease.

[0339] In one embodiment, the cancer includes one or more of the diseases or disorders including hypercholesterolemia, transthyretin amyloidosis, and β-hemoglobinopathy.

[0340] The main advantages of the present invention include:

[0341] (1) The linker mined and designed in the present invention has a large difference from the linkers used in the prior art. For example, compared with the amino acid sequence shown in GS-XTEN-GS (SEQ ID NO: 21), its similarity is very low.

[0342] (2) Compared with the widely used linker, such as GS-XTEN-GS (SEQ ID NO: 21), the base editing efficiency of the CRISPR-Cas system constructed with the base editors containing various Linkers of the present invention is equivalent or higher. For example, it can be seen from Figure 2 that the base editing efficiency of the base editor using Linker9 is greatly improved.

[0343] (3) The base editors (such as adenosine base editor or cytidine base editor) composed of the linkers designed in the present invention can be used for the treatment of diseases caused by single-base mutations.

[0344] (4) The linkers designed in the present invention can be widely used for the connection between various molecules or parts, and have broad application prospects.

[0345] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions in the following embodiments are usually carried out under conventional conditions, such as the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and weight parts.

[0346] Unless otherwise specified, the reagents and materials in the embodiments of the present invention are all commercially available products.

[0347] Example 1. Construction of Base Editors with Different Linkers

[0348] 1. A variety of linkers were designed, and the sequences are shown in the following table:

[0349] Table 1 Amino Acid Sequences of Different Linkers and Nucleotide Coding Sequences after Codon Optimization

[0350] Linker Name Linker Amino Acid Sequence Linker Nucleotide Coding Sequence Linker1 SEQ ID NO:1 SEQ ID NO:10 Linker2 SEQ ID NO:2 SEQ ID NO:11 Linker3 SEQ ID NO:3 SEQ ID NO:12 Linker4 SEQ ID NO:4 SEQ ID NO:13 Linker5 SEQ ID NO:5 SEQ ID NO:14 Linker6 SEQ ID NO:6 SEQ ID NO:15 Linker7 SEQ ID NO:7 SEQ ID NO:16 Linker8 SEQ ID NO:8 SEQ ID NO:17 Linker9 SEQ ID NO:9 SEQ ID NO:18

[0351] 2. Construct adenosine base editor expression vectors with Linker1 - 10

[0352] Taking the base editor 5V17.2 - nCas9 in the prior art as a basis, its amino acid sequence (SEQ ID NO:19) is shown as follows:

[0353]

[0354]

[0355] KKRKV*(SEQ ID NO:19). Among them, the double - underlined sequence is the nCas9 amino acid sequence, the single - underlined sequence is the deaminase 5V17.2 amino acid sequence, the bold and italic sequence is the linker GS - XTEN - GS (SEQ ID NO:21), which is located between the 976th and 1071st positions of the expression vector, the sequences at both ends are not marked as NLS sequences, and the asterisk at the C - terminus represents the position of the stop codon.

[0356] The corresponding nucleotide coding sequence of the above - mentioned base editor is as shown in SEQ ID NO:20, and the nucleotide coding sequence of GS - XTEN - GS is as shown in SEQ ID NO:22.

[0357] Using the base editor 5V17.2 - nCas9 expression vector as a template, upstream and downstream primers of the vector backbone were designed according to the position of the linker GS - XTEN - GS (SEQ ID NO:22) on the expression vector, and PCR amplification was carried out to obtain the backbone vector PCR product.

[0358] The upstream and downstream primers for the vector backbone are vector backbone-F (SEQ ID NO: 23) and vector backbone-R (SEQ ID NO: 24) respectively. The template of the base editor 5V17.2-nCas9 expression vector was amplified using the kit 2X Phanta Flash master mix (Novizan, P520-01). The PCR amplification reaction system and procedure are as follows:

[0359] Reaction system: 1 μl of vector backbone-F (10 μM), 1 μl of vector backbone-R (10 μM), 10 μl of 2X Phanta Flash master mix (Novizan, P520-01), 20 ng of template plasmid, and made up to 20 μl with sterile water.

[0360] The reaction procedure is as follows: pre-denaturation at 98°C for 30 s; denaturation at 98°C for 10 s, annealing at 60°C for 5 s, extension at 72°C for 45 s; 30 cycles; final extension at 72°C for 1 min.

[0361] After that, the universal DNA purification and recovery kit (Tiangen, DP214) was used to purify and recover the PCR product of the backbone vector. The obtained backbone sequence by amplification is shown in SEQ ID NO: 25.

[0362] Synthesize each linker nucleotide coding sequence fragment in Table 1 of Example 1, and design PCR primers corresponding to each linker nucleotide coding sequence. The 5' ends of the upstream and downstream primers of the linker each carry a homologous arm sequence of the vector backbone. The designed primers are shown in Table 2. Each linker nucleotide coding sequence fragment and primer were synthesized by Platinum Biotechnology (Shanghai) Co., Ltd.

[0363] Table 2 Upstream and downstream primer sequences of each linker

[0364]

[0365]

[0366] After that, each linker sequence fragment was amplified by PCR using the corresponding upstream and downstream primers in Table 2, and the linker PCR product was obtained. The amplification reaction system and reaction procedure are as follows:

[0367] Reaction system: 1 μl of Linker-F (10 μM), 1 μl of Linker-R (10 μM), 10 μl of 2X Phanta Flash master mix, and made up to 20 μl with sterile water.

[0368] The reaction procedure is as follows:

[0369] Pre-denaturation at 98°C for 30 s; denaturation at 98°C for 10 s, annealing at 60°C for 5 s, extension at 72°C for 2 s; 20 cycles; final extension at 72°C for 1 min.

[0370] The target fragment was purified and recovered using a universal DNA purification and recovery kit (Tiangen, product number DP214).

[0371] 3. Using the non-ligase-dependent multi-fragment one-step cloning kit ClonExpress MultiS One Step Cloning Kit (Vazyme Biotech, product number C113-01), each linker fragment was recombinantly ligated with the backbone vector fragment. After reacting at 37°C for 30 min, it was immediately placed on ice to obtain the recombinant base editor expression plasmid. Then, the recombinant base editor expression plasmid was transformed into DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.). The specific method is as follows:

[0372] Take out the DH5α competent cells from the refrigerator and thaw them on ice. Add 5 μl of the recombinant base editor expression plasmid to 50 μl of the competent cells, gently mix, place on ice for 20 min, heat shock in a 42°C water bath for 1 min, quickly transfer to ice after taking out, and let it stand for 2 min. Add 500 μl of LB medium and shake the bacteria in a shaker at 220 rpm for 30 min. Take 100 μl of the bacterial solution and spread it on a plate. Invert the plate and place it in a 37°C incubator. After overnight culture, pick single colonies, shake the bacteria, and send the bacterial solution to a sequencing company for Sanger sequencing. The plasmid bacterial solution with correct sequencing was used to extract the plasmid according to the instructions of the Vazyme plasmid mini-prep kit (product number DC203-01).

[0373] Obtain the 5V17.2-nCas9 recombinant vectors including Linker1-9, and name them Linker1-5V17.2-nCas9 recombinant vector, Linker2-5V17.2-nCas9 recombinant vector, Linker3-5V17.2-nCas9 recombinant vector, Linker4-5V17.2-nCas9 recombinant vector, Linker5-5V17.2-nCas9 recombinant vector, Linker7-5V17.2-nCas9 recombinant vector, Linker8-5V17.2-nCas9 recombinant vector, and Linker9-5V17.2-nCas9 recombinant vector respectively.

[0374] Example 2. Comparing the editing efficiency of base editors with different linkers using the dual-luciferase reporter system

[0375] To efficiently and sensitively detect the editing activities of various base editors, a dual-luciferase reporter system vector (SEQ ID NO: 45, see plasmid map in Figure 1 ) was constructed. It contains the coding sequence of Fluc luciferase (Firefly luciferase) (SEQ ID NO: 46) and the coding sequence of NanoLuc luciferase (NanoLuc) (SEQ ID NO: 86). A target sequence (SEQ ID NO: 44) and a PAM site 5'-NGG located at one end of the target sequence were designed in the front section of the NLuc sequence. Since the complementary sequence of the target sequence contains the stop codon TGA, under normal circumstances, the NanoLuc coding sequence is not expressed. When the target sequence is effectively edited by an adenosine base editor, the stop codon TGA will become CGA, enabling the normal expression of NanoLuc and emitting fluorescence. Therefore, taking Fluc as an internal reference, by calculating the luminescence reading value of NanoLuc, the base editing efficiency of the editor can be obtained.

[0376] The construction method is as follows:

[0377] The pmEGFP-C1 plasmid (Addgene, Plasmid, #36412) was digested with restriction endonucleases AseI (Thermo, ER0911), NheI (Thermo, FD0974), and XmaI (NEB) to obtain digestion product 1 (582 bp), product 2 (795 bp), and product 3 (3353 bp). Products 1 and 3 were recovered.

[0378] The F-Luc sequence, N-Luc sequence, N-Luc' sequence, EF1a promoter, and BGH-polyA sequence were synthesized respectively. The synthesis method is as follows:

[0379]

[0380] The PCR reaction system is as follows: 33 μl of sterile water, 5 μl of 10×KOD Buffer, 5 μl of dNTP, 3 μl of MgSO4, 1 μl each of upstream and downstream primers, 1 μl of PCR template (about 10 ng), and 1 μl of KOD enzyme. After preparation, it was vortexed and mixed evenly, centrifuged briefly at high speed, and then placed in a PCR instrument. The PCR reaction program is as follows: pre-reaction at 98°C for 1 min; thermal denaturation at 98°C for 15 s, annealing at 60°C for 30 s, extension at 72°C for 30 s / kb, 30 cycles, and then 72°C for 10 min. The PCR products were stored at 12°C, and then the DNA fragments were recovered by cutting and gel electrophoresis after agarose gel electrophoresis, and the EF1a promoter, N-Luc, N-Luc’, and BGH-polyA fragments were amplified.

[0381] The EF1a promoter, N-Luc, and BGH-polyA fragments were subjected to PCR splicing to obtain the first splicing fragment. The EF1a promoter, N-Luc’, and BGH-polyA fragments were subjected to PCR splicing to obtain the first splicing fragment’. The first splicing fragment and the first splicing fragment’ were respectively spliced with the obtained product 1, product 3, and the F-Luc sequence. The splicing was carried out using the Gibson Master Mix kit (NEB, E2611L), and the operation was carried out according to the kit instructions to obtain the reporter plasmid and the reference plasmid.

[0382]

[0383] The spacer sequence was designed according to the target sequence (SEQ ID NO:44), and the spacer sequence was constructed onto the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain the sgRNA expression vector, and this recombinant expression vector can transcribe the sgRNA sequence (SEQ ID NO:47).

[0384] The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vector including Linker1-9 were co-transfected into HEK293T cells (purchased from ATCC) with the lenti U6-sgRNA / EF1a-mCherry recombinant expression vector and the dual luciferase reporter system vector (50 ng each of the base editor expression vector, sgRNA expression vector, and dual luciferase reporter system vector) by the PEI transfection method. After 48 h of transfection, the detection of dual luciferase expression was carried out. The detection reagent was the Dual-Luciferase Reporter Assay Kit (Beyotime, catalog number RG028), and the Fluc fluorescence value was used as the internal reference. The editing efficiency was as Figure 2 shown. Through analysis, it was found that 5V17.2-nCas9 with Linker1-9 all had significant base editing activities. Among them, the editing activities of the base editors with Linker1, 3-9 for 5V17.2-nCas9 were all significantly higher than those of GS-XTEN-GS-5V17.2-nCas9.

[0385] Example 3. Editing activities of base editors with different linkers in HEK293T cells

[0386] To compare the editing activities of various base editors on endogenous genes in mammalian cells, site1 locus (see Gaudelli NM, Komor AC, Rees HA, Packer MS, Badran AH, Bryson DI, Liu DR. Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage. Nature. 2017 Nov 23;551(7681):464-471. doi:10.1038 / nature24644. Epub 2017 Oct 25. Erratum in: Nature. 2018 May 2;: PMID:29160308; PMCID:PMC5726555.) was selected as the target.

[0387] According to the site1 target sequence (SEQ ID NO:48), a spacer sequence was designed and constructed onto the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain the site1-sgRNA expression vector, which can transcribe the sgRNA sequence (SEQ ID NO:49) targeting the site1 locus. The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vector including Linker1-9 were co-transfected into HEK293T cells with the lenti U6-sgRNA / EF1a-mCherry recombinant expression vector (75 ng each of the base editor expression vector plasmid and the sgRNA expression vector) using the PEI transfection method. After continued culture for 24 h, the genomic DNA of the transfected positive cells was extracted, and the upstream primer site1-F (SEQ ID NO:50) and the downstream primer site1-R (SEQ ID NO:51) were designed and synthesized to perform PCR amplification on the site1 target sequence locus. The amplified product was subjected to sanger sequencing (Bosoom Biotech) and the editing efficiency was analyzed using the online tool https: / / hanlab.cc / beat / Figure 3) At the A5 site of site1, the editing efficiencies of base editors Linker2-5V17.2-nCas9 and Linker5-5V17.2-nCas9 are comparable to that of 5V17.2-nCas9. At the A7 site of site1, the editing efficiency of base editor Linker5-5V17.2-nCas9 is higher than that of GS-XTEN-GS-5V17.2-nCas9. The editing efficiencies of Linker1-5V17.2-nCas9, Linker2-5V17.2-nCas9, Linker6-5V17.2-nCas9, and Linker8-5V17.2-nCas9 are comparable to that of GS-XTEN-GS-5V17.2-nCas9. 5V17.2-nCas9 with other linkers also has significant editing activity at the site1 locus.

[0388] The site of the target sequence (SEQ ID NO:52) in the PCSK9 locus was selected for testing. According to the targeting sequence (SEQ ID NO:52), a spacer sequence was constructed and cloned into the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) to obtain a recombinant expression vector, which can transcribe the sgRNA sequence (SEQ ID NO:53) targeting the PCSK9 targeting sequence site.

[0389] The base editor 5V17.2-nCas9 expression vector and the 5V17.2-nCas9 recombinant vector containing linker Linker1-9 were co-transfected into HEK293T cells with the lenti U6-sgRNA / EF1a-mCherry recombinant expression vector by the PEI transfection method. After continuous culture for 24 h, the genomes of the transfected positive cells were extracted. The upstream primer PCSK9-F (SEQ ID NO:54) and the downstream primer PCSK9-R (SEQ ID NO:55) were designed and synthesized to perform PCR amplification on the PCSK9 locus. The amplification products were subjected to fastNGS sequencing (Tsingke Biological). It was found through testing that at the PCSK9 locus, each base editor had significant editing efficiency. In particular, the editing activities of base editors with Linker2, 3, and 5 were higher than that of GS-XTEN-GS-5V17.2-nCas9. 5V17.2-nCas9 with other linkers also had significant activity at this locus ( Figure 4 ).

[0390] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0391] Sequence information:

[0392]

[0393]

[0394]

[0395]

[0396]

[0397]

[0398]

[0399]

[0400]

[0401]

[0402]

[0403]

[0404]

[0405]

[0406]

[0407]

[0408]

[0409]

[0410]

[0411]

[0412]

[0413]

[0414]

[0415]

[0416]

[0417] All documents mentioned in this invention are cited herein for reference as if each individual document was cited for reference. In addition, it should be understood that after reading the above teachings of this invention, those skilled in the art can make various changes or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

Claims

1. A linking peptide, characterized in that, the linking peptide is selected from the following group: (a) a polypeptide having any one of the amino acid sequences shown in SEQ ID NO: 1-9; (b) a polypeptide having a homology (or identity) of ≥20%, 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with any one of the amino acid sequences shown in SEQ ID NO: 1-9, and the polypeptide has the biological function of SEQ ID NO: 1-9; (c) a derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, and even more preferably 1-5) amino acid residues in any one of the amino acid sequences shown in SEQ ID NO: 1-9, and retaining the biological function of SEQ ID NO: 1-9.

2. A fusion protein, characterized in that, the fusion protein comprises the linking peptide according to claim 1 and one or more functional domains, or comprises the linking peptide according to claim 1 and at least one nuclease. Preferably, the fusion protein comprises a nuclease, the linking peptide according to claim 1 and one or more deaminase domains.

3. An isolated polynucleotide, characterized in that, the polynucleotide encodes the linking peptide according to claim 1 or the fusion protein according to claim 2.

4. A vector, characterized in that, the vector contains the polynucleotide according to claim 3.

5. A CRISPR-Cas complex, characterized in that, the complex comprises: (i) a protein component, selected from the following group: the fusion protein according to claim 2; and (ii) a nucleic acid component, selected from the following group: a guide RNA, a nucleic acid encoding the guide RNA, a precursor RNA of the guide RNA, a nucleic acid encoding the precursor RNA of the guide RNA, or a combination thereof; wherein the protein component and the nucleic acid component bind to each other to form a complex.

6. A CRISPR-Cas composition, characterized in that, it comprises: (i) a first component, selected from the following group: the fusion protein according to claim 2, a nucleotide sequence encoding the fusion protein according to claim 2, and any combination thereof; and (ii) a second component, the second component being one or more guide RNAs, or a nucleotide sequence encoding the guide RNA; the guide RNA can form a complex with the fusion protein described in (i).

7. A CRISPR-Cas system, characterized in that, it comprises one or more vectors, and the one or more vectors comprise: (i) a first nucleic acid, which is a nucleotide sequence encoding the fusion protein according to claim 2; optionally the first nucleic acid is operably linked to a first regulatory element; and (ii) a second nucleic acid, which is a nucleotide sequence encoding a guide RNA; optionally the second nucleic acid is operably linked to a second regulatory element; wherein: The first nucleic acid and the second nucleic acid are present on the same or different vectors; The guide RNA is capable of forming a complex with the fusion protein described in (i).

8. A kit comprising one or more components selected from the following: the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7.

9. A delivery composition, characterized in that it comprises a delivery vector and one or more selected from the following: the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7.

10. A host cell, characterized in that it comprises the fusion protein according to claim 2, the polynucleotide according to claim 3, the vector according to claim 4, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9.

11. An enzyme preparation, characterized in that the enzyme preparation comprises the fusion protein according to claim 2, the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9.

12. A medicine box, characterized in that it comprises: a first container, and the complex according to claim 5, or the CRISPR-Cas composition according to claim 6, or the system according to claim 7 located in the first container, or a drug containing the complex according to claim 5, or the CRISPR-Cas composition according to claim 6, or the system according to claim 7.

13. A medicine box, characterized in that it comprises: (a1) A first container, and the fusion protein according to claim 2, or its coding gene, or its expression vector located in the first container, or a drug containing the fusion protein according to claim 2, or its coding gene, or its expression vector; (b1) Optionally, a second container, and the guide RNA or its expression vector located in the second container, or a drug containing the guide RNA or its expression vector.

14. A method for targeting and editing a target gene, characterized in that it comprises: contacting the fusion protein according to claim 2, or the complex according to claim 5, the CRISPR-Cas composition according to claim 6, or the system according to claim 7, or the delivery composition according to claim 9, or the enzyme preparation according to claim 11, or the medicine box according to claim 12 or 13 with the target gene, or delivering it into a cell containing the target gene, and a target sequence is present in the target gene.

15. A cell or its progeny obtained by the method according to claim 14, wherein the cell contains a modification that does not exist in its wild type.

16. A cell product of the cell or its progeny according to claim 15.

17. An in vitro, ex vivo or in vivo cell or cell line or their progeny, said cell or cell line or their progeny comprising: the fusion protein of claim 2, the polynucleotide of claim 3, or the complex of claim 5, the CRISPR-Cas composition of claim 6 or the system of claim 7 or the delivery composition of claim 9.

18. A cell preparation, wherein, it comprises the host cell of claim 10 or the cell of claim 15 or their progeny, or the cell product of the cell of claim 16 or their progeny, or the cell or cell line of claim 17 or their progeny.

19. Use of the fusion protein of claim 2, the polynucleotide of claim 3, the vector of claim 4, the complex of claim 5, the CRISPR-Cas composition of claim 6 or the system of claim 7, the kit of claim 8, the delivery composition of claim 9, the enzyme preparation of claim 11, the kit of claim 12 or claim 13, wherein, for preparing a drug or preparation for nucleic acid editing (e.g., gene or genome editing).

20. Use of the fusion protein of claim 2, the polynucleotide of claim 3, the vector of claim 4, the complex of claim 5, the CRISPR-Cas composition of claim 6 or the system of claim 7, the kit of claim 8, the delivery composition of claim 9, the enzyme preparation of claim 11, the kit of claim 12 or claim 13, wherein, for preparing a drug or preparation for one or more selected from the group consisting of: (i) ex vivo gene or genome editing; (ii) editing a target sequence in a target locus to modify a biological or non-human organism; (iii) treating a disorder caused by a defect in a target sequence in a target locus; (iv) treating a disorder or disease in a subject in need thereof.

Citation Information

Patent Citations

  • Adenosine nucleobase editors and uses thereof

    US10113163B2

  • Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells

    WO2001038547A2