Stacas9 protein or mutant thereof, gene editing system containing stacas9 protein or mutant thereof and application
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI PUDONG HOSPITAL
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-07
AI Technical Summary
但是这些Cas9蛋白或者容易脱靶(即非靶向位点切割),或者PAM序列复杂,或者编辑活性低,难以进行广泛应用
[0018]本发明人开发了可在真核细胞环境高效进行基因编辑的CRISPR/StaCas9编辑工具。该StaCas9蛋白具有较少数量的氨基酸,特别是具有目前可用于真核基因编辑器中的较少数量的氨基酸,因此可有效地包装到表达载体例如腺相关病毒载体中。并且,该蛋白具有特异性高、PAM简单的特性,而且蛋白分子量小可轻易被腺相关病毒等载体工具包装,非常适合后期作为基因治疗工具的开发。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, specifically to the StaCas9 protein or its mutants, gene editing systems containing the StaCas9 protein or its mutants, and their applications. Background Technology
[0002] The CRISPR / Cas system is an acquired immune system evolved by bacteria and archaea to defend against invasion by exogenous viruses or plasmids. The CRISPR / Cas9 system contains tracrRNA (trans-activating crRNA) and crRNA (CRISPR-derived RNA), which, together with Cas9, form a complex to function. tracrRNA and crRNA can be fused together via a linker sequence to form single-stranded guide RNA (sgRNA). When DNA breaks occur, repair is handled by two main DNA damage repair mechanisms within the cell: non-homologous end-joining (NHEJ) and homologous recombination (HR). NHEJ repair results in base deletions or insertions, which can be used for gene knockout; HR repair, when a homologous template is provided, can be used for site-specific gene insertion and precise base substitution.
[0003] Beyond basic research, the CRISPR / Cas9 gene editing system holds broad clinical application prospects. It can be used for gene therapy by introducing Cas enzymes and single-stranded guide RNA into the body. Currently, the most effective expression vector used in gene therapy is adeno-associated virus (AAV). However, AAV viral packaging has limitations on DNA length, generally not exceeding 4.5 kb. Therefore, although SpCas9 has been widely used in basic research due to its simple PAM sequence (recognizing NGG) and high activity, the 1368 amino acids in the SpCas9 protein, combined with the sgRNA and promoter required for gene editing systems, make it difficult to effectively package into AAV viruses, limiting its clinical application. To overcome this problem, several small-molecule Cas9 proteins have been developed, such as SaCas9 (PAM sequence NNGRRT), St1Cas9 (PAM sequence NNAGAW), NmCas9 (PAM sequence NNNNGATT), Nme2Cas9 (PAM sequence NNNNCC), and CjCas9 (PAM sequence NNNNRYAC). However, these Cas9 proteins are either prone to off-target cleavage (i.e., non-target site cleavage), have complex PAM sequences, or have low editing activity, making them difficult to widely apply.
[0004] Therefore, finding a small CRISPR / Cas9 system with high editing activity, high specificity, and simple PAM sequences is the hope for solving the above problems. Summary of the Invention
[0005] To address the aforementioned problems, the inventors, through repeated research, discovered a Cas9 protein (StaCas9, derived from...) Staphylococcus The invention was completed by obtaining the wild-type StaCas9 protein (sp. 17KM0847) and further obtaining the corresponding mutant and designing the corresponding single-stranded guide RNA. The wild-type StaCas9 protein or its mutant and the above-mentioned single-stranded guide RNA can constitute a CRISPA / StaCas9 gene editing system for effective gene editing.
[0006] Therefore, in a first aspect, the present invention provides a StaCas9 mutant protein, said StaCas9 mutant protein comprising mutations at E887 and / or S998 relative to the amino acid sequence of the wild-type StaCas9 protein as shown in SEQ ID NO: 1.
[0007] In a second aspect, the present invention provides a conjugate comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. b) Modified parts; and c) An optional adapter for connecting the StaCas9 protein or its homolog to the modified portion.
[0008] In a third aspect, the present invention provides a fusion protein comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. b) Other proteins or peptides; and c) Optional adapters for connecting the StaCas9 protein or its homologs to the other protein or polypeptide.
[0009] In a fourth aspect, the present invention provides a single-stranded guide RNA comprising a CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence that shares at least 90% sequence identity with the nucleic acid sequence shown in SEQ ID NO: 5 and retains the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence that retains the biological activity of SEQ ID NO: 5, obtained by modifying the nucleic acid sequence described in SEQ ID NO: 5.
[0010] In a fifth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a StaCas9 mutant protein or a homolog of the StaCas9 mutant protein of the first aspect of the present invention, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention.
[0011] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0012] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a StaCas9 mutant protein or a homolog of the StaCas9 mutant protein of the first aspect of the present invention, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention.
[0013] In an eighth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0014] In a ninth aspect, the present invention provides a CRISPR / StaCas9 gene editing system comprising: 1) Protein components, which include StaCas9 protein or its homologs, conjugates or fusion proteins, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified moiety or the other protein or polypeptide; and 2) Nucleic acid components, which include: The fourth aspect of this invention involves a single-stranded guide RNA; Furthermore, the protein component and the nucleic acid component combine with each other to form a complex.
[0015] In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules of the fifth and / or sixth aspects of the present invention, or a vector of the seventh and / or eighth aspects of the present invention.
[0016] In an eleventh aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (4) with the target sequence in the intracellular or in vitro environment: (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (3) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; or (4) The CRISPR / StaCas9 gene editing system of the ninth aspect of the present invention; Upon contact with the target sequence, the StaCas9 protein or its homologs, conjugates, or fusion proteins recognize the PAM sequence 5'-NNG located at the 3' end of the target sequence.
[0017] In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: i. Choose any one of the following (1) to (6): (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The fifth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or an isolated nucleic acid molecule containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the sixth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (3) The fifth aspect of the present invention comprises an isolated nucleic acid molecule encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention and comprising a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or the sixth aspect of the present invention comprises a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention and comprising a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (4) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (5) The seventh aspect of the present invention comprises a nucleic acid sequence encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or the eighth aspect of the present invention comprises a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention and a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (6) The CRISPR / StaCas9 gene editing system of the ninth aspect of the present invention; and ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.
[0018] The inventors have developed a CRISPR / StaCas9 editing tool capable of efficient gene editing in a eukaryotic environment. This StaCas9 protein has a relatively small number of amino acids, particularly a small number currently available for eukaryotic gene editors, allowing for efficient packaging into expression vectors such as adeno-associated virus vectors. Furthermore, the protein exhibits high specificity, simple PAM (Physical Amino Acid) structure, and its small molecular weight allows for easy packaging into vectors such as adeno-associated viruses, making it highly suitable for future development as a gene therapy tool.
[0019] Furthermore, the PAM of the StaCas9 protein of this invention is NNG, a very simple sequence, thus enabling the CRISPR / StaCas9 editing system to have a wide editing range. Moreover, experiments conducted by the inventors have demonstrated that the StaCas9 mutant protein of this invention exhibits a significant advantage in editing efficiency at random sites compared to the wild-type StaCas9 protein, and demonstrates strong gene editing capabilities in a eukaryotic environment. Compared to other Cas9 proteins in the same series, the mutant protein enStaCas9 possesses extremely significant editing advantages, making it more suitable for the development and application research of gene editing.
[0020] The StaCas9 protein of this invention has high activity and high specificity, and has a relatively simple PAM sequence, which expands the field of StaCas9 protein and increases its application range. Attached Figure Description
[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below.
[0022] Figure 1 A PAM logo diagram is shown for recognizing PAM sequences by the CRISPR / StaCas9 gene editing system according to an embodiment of the present invention.
[0023] Figure 2 The results of gene editing efficiency of the CRISPR / StaCas9 gene editing system according to an embodiment of the present invention after gene editing at twelve target sites are shown.
[0024] Figure 3 The results show the specificity detection of the CRISPR / StaCas9 gene editing system according to an embodiment of the present invention in the GFP reporter system HEK293T cell line.
[0025] Figure 4 The results of gene editing efficiency of the CRISPR / enStaCas9 gene editing system according to an embodiment of the present invention after gene editing at six target sites are shown. Detailed Implementation
[0026] Unless otherwise stated, the scientific and technical terms used in this application have the meanings commonly understood by those skilled in the art. Definitions and explanations of relevant terms are provided below for a better understanding of the invention.
[0027] The terms “StaCas9 protein,” “StaCas9,” “StaCas,” and “Cas” used herein are interchangeable and refer to RNA-guided nucleases, including the StaCas9 protein or its functionally active fragments. Therefore, in this document, the term “StaCas9 protein” may refer to the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1, a mutant of the wild-type StaCas9 protein, or both, as the case may be. The StaCas9 protein is a protein component of the CRISPR / StaCas9 genome editing system that, guided by single-stranded guide RNA (sgRNA), targets and cleaves DNA target sequences, forming DNA double-strand breaks (DSBs). DNA double-strand breaks can activate the cell’s inherent repair mechanisms non-homologous end-joining (NHEJ) and homologous recombination (HR), thereby repairing DNA damage in the cell. During the repair process, the specific DNA sequence is edited at specific sites.
[0028] The terms “single-stranded guide RNA” and “sgRNA” as used herein are interchangeable and have the meanings commonly understood by those skilled in the art. Generally, a single-stranded guide RNA may comprise a scaffold sequence and a guide sequence, which is also referred to herein as guide RNA (or gRNA). In the context of an endogenous CRISPR system, the guide sequence is also referred to as a spacer sequence. In some cases, the guide sequence is any polynucleotide sequence that is sufficiently similar to a target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / StaCas9 complex to said target sequence. In some embodiments, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining optimal alignment is within the capabilities of those skilled in the art. For example, there are publicly available and commercially available comparison algorithms and programs, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.
[0029] As used herein, the term "CRISPR / StaCas9 complex" refers to a complex formed by the binding of a single-stranded guide RNA or mature crRNA:tracrRNA hybrid to a StaCas9 protein (e.g., wild-type StaCas9 protein or a mutant thereof), which contains a guide sequence that hybridizes to a target sequence and thereby enables the StaCas9 protein to bind to said target sequence. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the single-stranded guide RNA or mature crRNA.
[0030] Therefore, in the formation of the CRISPR / StaCas9 complex, the "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, where hybridization between the target sequence and the guide sequence will promote StaCas9's activity, such as cleavage of the target sequence. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote StaCas9's activity. The target sequence can include any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In other cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast.
[0031] As used herein, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas protein used, but the PAM is typically a 2-5 base sequence adjacent to the protospacer sequence (target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas protein.
[0032] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” used herein are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.
[0033] The terms “polypeptide,” “peptide,” and “protein” as used herein are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, and to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.
[0034] The terms “sequence identity” or “homology” as used herein have their generally accepted meaning in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of that molecule. (See, e.g., Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; SequenceAnalysis in Molecular Biology, von Heinje, G., Academic Press, 1987; andSequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While many methods exist for measuring the identity between two polynucleotides or polypeptides, the term "identity" is known to those skilled in the art to refer to conserved amino acid substitutions in peptides or proteins that can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide does not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0035] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. A vector is called an expression vector when it enables the expression of a protein encoded by the inserted polynucleotide, or when it enables transcription of the inserted polynucleotide (e.g., to generate mRNA or functional RNA). Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to, plasmid vectors and viral vectors. Vectors may also contain various regulatory sequences that regulate expression. The terms "regulatory sequence" and "regulatory element" are used interchangeably herein, referring to a nucleotide sequence located upstream (5' non-coding sequence), midway, or downstream (3' non-coding sequence) of a coding sequence that affects transcription, RNA processing, stability, or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. These regulatory sequences may originate from different sources or from the same source but arranged in a manner different from what is typically found naturally. Additionally, vectors may contain a replication initiation site.
[0036] As used herein, the term "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.
[0037] As used in this article, the term "constitutive promoter" refers to a promoter that generally causes gene expression in most cell types and under most conditions. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to promoters that are primarily, but not necessarily, expressed specifically in one tissue or organ, and may also be expressed in a specific cell type. "Developmental regulatory promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses a manipulated DNA sequence in response to endogenous or exogenous stimuli (environment, hormones, chemical signals, etc.).
[0038] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.
[0039] As used in this article, the term "stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its subsequent generations.
[0040] The term "transient transformation" as used in this article refers to the introduction of nucleic acid molecules or proteins into cells to perform their functions without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.
[0041] The terms “identity,” “consistency,” or “homology” used in this article have the generally accepted meanings in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule (see, for example, Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While many methods exist for measuring the identity between two polynucleotides or peptides, the term "identity" is known to those skilled in the art to refer to conserved amino acid substitutions in peptides or proteins that can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a peptide does not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0042] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. Complementarity percentage indicates the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% if 5, 6, 7, 8, 9, or 10 out of 10 are complementary). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. As used in this article, "substantially complementary" means a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or it means two nucleic acids that hybridize under strict conditions.
[0043] As used in this paper, the term "strict condition" in relation to hybridization refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of a strict condition is described in Tijssen (1993), *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes*, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assay", Elsevier, New York.
[0044] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. Hydrogen bonds can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.
[0045] Proteins, their conjugates and fusion proteins As described above, the inventors unexpectedly discovered that wild-type and mutant StaCas9 proteins, when combined with sgRNA, can effectively form a CRISPR / StaCas9 gene editing system. StaCas9 proteins have a relatively small number of amino acids, high editing efficiency, and recognize very simple PAM sequences (NNG). Therefore, the CRISPR / StaCas9 gene editing system of this invention has a wide editing range and can be effectively used for site-specific editing of target sequences. This completes the present invention.
[0046] Therefore, in a first aspect, the present invention provides a StaCas9 mutant protein, said StaCas9 mutant protein comprising mutations at E887 and / or S998 relative to the amino acid sequence of the wild-type StaCas9 protein as shown in SEQ ID NO: 1.
[0047] In one implementation, the mutation includes E887R, S998K, or E887R and S998K.
[0048] Experiments have demonstrated that the StaCas9 mutant protein of this invention exhibits a significant advantage in editing efficiency at random sites compared to the wild-type StaCas9 protein. Furthermore, compared to other StaCas9 proteins in the same series, the double-point mutant proteins enStaCas9 (E887R and S998K) possess extremely significant editing advantages, making them more suitable for the development and application of gene editing technologies.
[0049] Furthermore, the StaCas9 protein can be derivatized, for example, by linking it to other molecules (e.g., other proteins or peptides, detectable markers). Typically, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to single-stranded guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, in this invention, the StaCas9 protein can be functionally linked (through chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular moieties, such as other proteins or peptides, detectable markers, pharmaceutical reagents, etc.
[0050] Specifically, the StaCas9 protein can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the protein's ability to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the StaCas9 protein targeted. For example, it can be linked to a detectable tag to facilitate the detection of the StaCas9 protein. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the StaCas9 protein.
[0051] Therefore, in a second aspect, the present invention provides a conjugate comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; b) Modified parts; and c) An optional adapter for connecting the StaCas9 protein or its homolog to the modified portion.
[0052] In this invention, the “biological activity” of the StaCas9 protein refers to its activity in binding to single-stranded guide RNA, its endonuclease activity (including single-stranded and double-stranded cleavage activity), and / or its activity in binding to and cleaving a specific site of a target sequence under the guidance of guide RNA (gRNA), but is not limited thereto.
[0053] It is understandable that, in addition to the StaCas9 protein itself, the StaCas9 protein can also be combined with other substances, such as other proteins or tagged objects, to give it other functions.
[0054] Therefore, in one embodiment, the modified portion may be another protein or peptide, a detectable marker, or a combination thereof.
[0055] In a further embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0056] Epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).
[0057] Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.
[0058] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0059] The StaCas9 protein of the present invention can be coupled, conjugated, or fused to the modified moiety via a linker, or it can be directly linked to the modified moiety without a linker. Linkers are well known in the art, and examples of them may include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0060] In a third aspect, the present invention provides a fusion protein comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; b) Other proteins or peptides; and c) Optional adapters for connecting the StaCas9 protein or its homologs to the other protein or polypeptide.
[0061] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0062] Single-stranded guide RNA In a fourth aspect, the present invention provides a single-stranded guide RNA comprising a CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identical to the nucleic acid sequence shown in SEQ ID NO: 5 and retains the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence obtained by modifying the nucleic acid sequence described in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5.
[0063] In one embodiment, the modification may be one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening.
[0064] In a further embodiment, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine, or ten bases relative to the base sequence.
[0065] In yet another embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 3' end of the CRISPR scaffold sequence, the CRISPR spacer sequence being a sequence of 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides in length and capable of complementary pairing with the target sequence.
[0066] In a preferred embodiment, the CRISPR spacer sequence is a 20-nucleotide sequence that is complementary to the target sequence.
[0067] In a further embodiment, the single-stranded guide RNA further includes a terminator at the 3' end of the spacer sequence. As an example, the terminator may be a plurality of terminators, such as at least six (e.g., seven or eight) U.
[0068] The single-stranded guide RNA can bind to the StaCas9 protein, conjugate, or fusion protein mentioned above to form a complex. This complex can recognize the corresponding PAM and thereby bind to the target sequence, thus achieving the cleavage of the target sequence or gene editing.
[0069] Nucleic acid encoding and vectors In a fifth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a StaCas9 mutant protein or a homolog of the StaCas9 mutant protein, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention, wherein the homolog of the StaCas9 mutant protein has an amino acid sequence that ... Homologous to the StaCas9 mutant protein with at least 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95%, 99.99%, or at least 99.999% sequence identity and retaining the biological activity of the StaCas9 mutant protein.
[0070] In one embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10.
[0071] In a further embodiment, the isolated nucleic acid molecule also comprises a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0072] As an example, the isolated nucleic acid molecule contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and also contains the nucleic acid sequence shown in SEQ ID NO: 6.
[0073] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0074] In one embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 6.
[0075] In a preferred embodiment, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a CRISPR spacer sequence.
[0076] In one embodiment, the isolated nucleic acid molecule further comprises encoding such as SEQ ID NO: The nucleic acid sequence of wild-type StaCas9 protein or a homolog of wild-type StaCas9 protein shown in 1, wherein the homolog of wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of wild-type StaCas9 protein and retains the biological activity of wild-type StaCas9 protein.
[0077] In a preferred embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 6 and also comprises the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
[0078] After the isolated nucleic acid molecules of the present invention are transfected into corresponding cells using certain tools known in the art, such as expression vectors, the isolated nucleic acid molecules of the present invention can express the StaCas9 mutant protein or its homologs described in the first aspect of the present invention, the conjugates of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, the wild-type StaCas9 protein or its homologs as shown in SEQ ID NO: 1, and / or the single-stranded guide RNA described above, and perform corresponding functions therein, such as gene editing.
[0079] In addition, the isolated nucleic acid molecules of the present invention can express the StaCas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention, the wild-type StaCas9 protein or its homolog as shown in SEQ ID NO: 1, and single-stranded guide RNA, individually or separately, or can express the expression products in one go. The choice of expression method depends on the specific circumstances.
[0080] Furthermore, the expressed product has the corresponding effects and / or functions described above, which will not be repeated here for the sake of brevity.
[0081] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a StaCas9 mutant protein or a homolog of the StaCas9 mutant protein according to the first aspect of the present invention, a conjugate according to the second aspect of the present invention, or a fusion protein according to the third aspect of the present invention, wherein the homolog of the StaCas9 mutant protein has an amino acid sequence having an amino acid sequence that has ... Homologous to the StaCas9 mutant protein with at least 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95%, 99.99%, or at least 99.999% sequence identity and retaining the biological activity of the StaCas9 mutant protein.
[0082] In one embodiment, the vector comprises the nucleic acid sequence or a degenerate sequence thereof shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10.
[0083] In one embodiment, the vector may be an expression vector, such as a plasmid vector like pUC19 vector, an applicator vector, a pAAV2_ITR vector, a retroviral vector, a lentiviral vector, an adenovirus vector, or an adeno-associated virus vector.
[0084] In yet another embodiment, the vector further comprises a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0085] As an example, the vector contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and also contains the nucleic acid sequence shown in SEQ ID NO: 6.
[0086] In an eighth aspect, the present invention provides a vector comprising a nucleic acid molecule encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0087] In one embodiment, the vector comprises the nucleic acid sequence shown in SEQ ID NO: 6 or a degenerate sequence thereof.
[0088] In a preferred embodiment, the vector further comprises a nucleic acid sequence encoding a CRISPR spacer sequence.
[0089] In one embodiment, the vector further comprises a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, wherein the homolog of the wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type StaCas9 protein and retains the biological activity of the wild-type StaCas9 protein.
[0090] As an example, the vector contains the nucleic acid sequence shown in SEQ ID NO: 6 and also contains the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
[0091] As described above, after the vector of the present invention is transfected into cells, the nucleic acid sequence cloned in the vector can be expressed as the StaCas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, the wild-type StaCas9 protein or its homolog as shown in SEQ ID NO: 1, and / or the single-stranded guide RNA described above, and perform corresponding functions, such as gene editing.
[0092] Alternatively, multiple vectors, such as two vectors, can be transfected into cells. One vector expresses the StaCas9 mutant protein or its homologs according to the first aspect of this invention, the conjugate of the second aspect of this invention, or the fusion protein of the third aspect of this invention, or the wild-type StaCas9 protein or its homologs as shown in SEQ ID NO: 1, while the other vector expresses single-stranded guide RNA. Subsequently, the expressed StaCas9 mutant protein or its homologs according to the first aspect of this invention, the conjugate of the second aspect of this invention, or the fusion protein of the third aspect of this invention, or the wild-type StaCas9 protein or its homologs as shown in SEQ ID NO: 1, are then combined with the expressed single-stranded guide RNA to form a complex, which performs the corresponding function, such as gene editing.
[0093] Of course, the nucleic acid sequence encoding the StaCas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, or the wild-type StaCas9 protein or its homolog as shown in SEQ ID NO: 1, and the nucleic acid sequence encoding the single-stranded guide RNA can also be cloned into a vector, such that after the vector is transfected into cells, it expresses both the StaCas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, or the wild-type StaCas9 protein or its homolog as shown in SEQ ID NO: 1, and the single-stranded guide RNA, and performs the corresponding functions, such as gene editing.
[0094] CRISPR / StaCas9 gene editing system In a ninth aspect, the present invention provides a CRISPR / StaCas9 gene editing system comprising: 1) Protein components, comprising StaCas9 protein or its homologs, conjugates, or fusion proteins, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence shown by the StaCas9 protein and retains the biological activity of the StaCas9 protein; The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified moiety or the other protein or polypeptide; and 2) Nucleic acid components, comprising: single-stranded guide RNA of the fourth aspect of the present invention; Furthermore, the protein component and the nucleic acid component combine with each other to form a complex.
[0095] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0096] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0097] As an example, the protein component comprises: wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and the nucleic acid component comprises the nucleic acid sequence shown in SEQ ID NO: 6.
[0098] The CRISPR / StaCas9 gene editing system of the present invention can be directly constructed from the StaCas9 protein or its homologs, conjugates or fusion proteins described herein and the single-stranded guide RNA described herein, or it can be constructed from the expression product obtained by the vector described herein.
[0099] The CRISPR / StaCas9 gene editing system of the present invention achieves the recognition, localization, cleavage and gene editing of target sequences through the combined action of the StaCas9 protein or its homologs, conjugates or fusion proteins contained therein and single-stranded guide RNA.
[0100] The CRISPR / StaCas9 gene editing system of this invention can precisely locate target sequences. "Precisely located" has two meanings: first, the CRISPR / StaCas9 gene editing system itself can recognize and bind to the target sequence; second, the CRISPR / StaCas9 gene editing system can bring other proteins fused with the StaCas9 protein or proteins that specifically recognize the sgRNA to the location of the target sequence.
[0101] The CRISPR / StaCas9 gene editing system of the present invention has low tolerance for non-target sequences. In this document, "low tolerance" means that the CRISPR / StaCas9 gene editing system of the present invention is substantially or completely unable to recognize and bind to non-target sequences, or substantially or completely unable to bring other proteins fused with the StaCas9 protein or proteins that specifically recognize the sgRNA to the location of the non-target sequence.
[0102] The CRISPR / StaCas9 gene editing system of the present invention can target more DNA sequences in the genome because the PAM sequence on the target sequence recognized by the StaCas9 protein contained therein is simpler.
[0103] cell In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules of the fifth and / or sixth aspects of the present invention, or a vector of the seventh and / or eighth aspects of the present invention.
[0104] As an example, the cell can be a prokaryotic cell or a eukaryotic cell. For the eukaryotic cell, as an example, it can be a plant cell or an animal cell. For the animal cell, as an example, it can be a mammalian cell.
[0105] As an example, the cell could also be a human cell.
[0106] It should be noted that the term "human cells" as used in this article does not include human germ cells or fertilized eggs, but may include human embryonic stem cells that have not yet developed in the body and are within 14 days of fertilization. Furthermore, the term "animal cells" as used in this article does not include animal embryonic stem cells.
[0107] method In an eleventh aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising contacting any one of the following (1) to (4) with the target sequence in the intracellular or in vitro environment: (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (3) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; or (4) The CRISPR / StaCas9 gene editing system of the ninth aspect of the present invention; Upon contact with the target sequence, the StaCas9 protein or its homologs, conjugates, or fusion proteins recognize the PAM sequence 5'-NNG located at the 3' end of the target sequence.
[0108] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0109] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0110] As an example, for item (1) above, it can be: wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5.
[0111] As an example, for item (2) above, it can be: a vector containing the nucleic acid sequence or its degenerate sequence shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:9 or SEQ ID NO:10, and a vector containing the nucleic acid sequence shown in SEQ ID NO:6.
[0112] In one embodiment, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell being, for example, a plant cell or an animal cell, the animal cell being, for example, a mammalian cell such as a human cell. Similarly, the term "human cell" as used herein does not include human germ cells or fertilized eggs, but may include human embryonic stem cells that have not yet developed in vivo and are within 14 days of fertilization, and the term "animal cell" as used herein does not include animal embryonic stem cells.
[0113] In one embodiment, the gene editing includes one or more of the following: gene knockout of a target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking.
[0114] Furthermore, in one embodiment, the single base conversion includes adenine to guanine, cytosine to thymine, or cytosine to uracil.
[0115] In one embodiment, in the method, the CRISPR spacer sequence of the single-stranded guide RNA forms a fully complementary base pairing structure with the target sequence, and an incompletely complementary base pairing structure with a non-target sequence.
[0116] In this document, the incomplete base pairing structure refers to a structure that includes a portion of base pairing and a portion of non-base pairing, wherein the non-base pairing includes, for example, base mismatch and / or base bulge.
[0117] In one embodiment, the incomplete base complementary pairing structure includes one or more, for example, two or more base mismatches.
[0118] Therefore, the StaCas9 protein or its homologs, conjugates, or fusion proteins of the present invention can cleave target sites on the target sequence, and the cleavage action of the StaCas9 protein or its homologs, conjugates, or fusion proteins causes double-strand breaks in the target sequence. Furthermore, when the method is performed intracellularly, the cleaved target sequence can be repaired through intracellular non-homologous end joining repair or homologous recombination repair pathways, thereby achieving gene editing of the target sequence.
[0119] The CRISPR / StaCas9 gene editing system and gene editing method of the present invention have been experimentally shown to have an editing efficiency of 13%-79% for wild-type StaCas9. Furthermore, the mutant enStaCas9 exhibits even higher editing efficiency compared to wild-type StaCas9. For example, at sites where the PAM recognition sequence is NNGC, the mutant enStaCas9 demonstrates an editing efficiency of 13%-36%, while the wild-type StaCas9 achieves only 2%-34% efficiency at these sites. Additionally, the CRISPR / StaCas9 gene editing system exhibits a very low mismatch rate in the first 21 bp of the guide RNA. Therefore, this gene editing system can specifically edit target genes, exhibiting high editing efficiency and low off-target rate, and can be widely applied to gene editing in cells or in vitro environments.
[0120] Reagent test kit In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: i. Choose any one of the following (1) to (6): (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein of the first aspect of the present invention. The homolog is an amino acid sequence that has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The fifth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or an isolated nucleic acid molecule containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the sixth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (3) The fifth aspect of the present invention comprises a nucleic acid sequence encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, a conjugate of the second aspect of the present invention or a fusion protein of the third aspect of the present invention and comprising a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or the sixth aspect of the present invention comprises a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention and comprising a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (4) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein; and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (5) The seventh aspect of the present invention comprises a nucleic acid sequence encoding the StaCas9 mutant protein of the first aspect of the present invention or a homolog of the StaCas9 mutant protein, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or the eighth aspect of the present invention comprises a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention and a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein. (6) The CRISPR / StaCas9 gene editing system of the ninth aspect of the present invention; and ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.
[0121] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0122] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0123] As an example, for item (1) above, it can be: wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5.
[0124] As an example, for item (2) above, it can be: an isolated nucleic acid molecule containing the nucleic acid sequence or its degenerate sequence shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and an isolated nucleic acid molecule containing the nucleic acid sequence shown in SEQ ID NO: 6.
[0125] As an example, for item (4) above, it can be: a vector containing the nucleic acid sequence or its degenerate sequence shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and a vector containing the nucleic acid sequence shown in SEQ ID NO: 6.
[0126] Of course, those skilled in the art will understand that the kit of the present invention may also contain other reagents that facilitate gene editing.
[0127] A brief description of the sequence involved in this invention. SEQ ID NO: 1: Wild-type StaCas9 protein sequence SEQ ID NO: 2: enStaCas9 mutant protein sequence SEQ ID NO: 3: The coding sequence of wild-type StaCas9 protein SEQ ID NO: 4: Coding sequence of the enStaCas9 mutant protein SEQ ID NO: 5: Single-stranded guide RNA scaffold sequence SEQ ID NO: 6: DNA sequence of a single-stranded guide RNA scaffold. SEQ ID NO: 7: StaCas9-E887R mutant protein sequence SEQ ID NO: 8: StaCas9-S998K mutant protein sequence SEQ ID NO: 9: Coding sequence of the StaCas9-E887R mutant protein SEQ ID NO: 10: Coding sequence of the StaCas9-S998K mutant protein Example The invention will now be described with reference to the following embodiments, which are intended to be illustrative and not limiting. Those skilled in the art will understand that the embodiments provided herein are for the purpose of describing the invention in detail only and are not intended to limit the scope of protection claimed by the invention.
[0128] Unless otherwise specified, the experiments and methods described in the examples were generally performed according to conventional methods well known in the art and described in the various references. Furthermore, for conditions not specifically specified in the examples, conventional conditions or conditions recommended by the manufacturer were followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0129] Example 1 (1) Constructing plasmid pAAV-CMV-StaCas9-Sa sgRNA-puro The amino acid sequence of the wild-type StaCas9 protein (SEQ ID NO: 1, gene search number WP_180810311.1) was downloaded from NCBI.
[0130] The nucleic acid sequence encoding the wild-type StaCas9 protein was codon-optimized to obtain the gene sequence of the wild-type StaCas9 protein that is highly expressed in human cells, as shown in SEQ ID NO: 3.
[0131] The gene sequence shown in SEQ ID NO: 3 obtained above was used for gene synthesis and constructed into the pAAV-CMV-SauriCas9-puro backbone plasmid (Addgene platform, catalog# 135965) to obtain the plasmid pAAV-CMV-StaCas9-Sa sgRNA-puro. The specific steps are as follows.
[0132] For the pAAV-CMV-SauriCas9-puro plasmid and the synthesized StaCas9 egg gene fragment, the restriction sites are AgeI and BamHI.
[0133] The pAAV-CMV-SauriCas9-puro plasmid was digested for 4 h at 37°C using the following enzyme digestion reaction system.
[0134]
[0135] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min. DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions, and finally eluted with ultrapure water.
[0136] The digested plasmid backbone and wild-type StaCas9 DNA fragment were ligated using Instant Sticky-end Ligase Master Mix. The ligation reaction system is as follows:
[0137] The connection reaction conditions are: room temperature reaction for 5-10 min.
[0138] The above ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and revive Escherichia coli DH5α competent cells.
[0139] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0140] The correctly ligated E. coli DH5α clone was subjected to culture and then the plasmid was extracted to obtain the plasmid pAAV-CMV-StaCas9-Sa sgRNA-puro.
[0141] (2) Constructing plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro The DNA sequence of the single-stranded guide RNA scaffold sequence for use with StaCas9 protein was synthesized (SEQ ID NO: 6), namely Sta sgRNA. The Sa sgRNA sequence of the pAAV-CMV-StaCas9-Sa sgRNA-puro plasmid in (1) was replaced with the StasgRNA sequence, and the primers and Sta sgRNA sequence were constructed as follows: Sta sgRNA: GTTTTAGTACTCTGGAAACAGAATCTACTAAAACAAGACATTATGTCGTGTTTATCCCATCATTCTTGATGGGA Primer-F: CCATCATTCTTGATGGGATTTTTGCGGCCGCAGGAACC Primer-R: TCCAGAGTACTAAAACTGAGACCTGCCGTGGTCT The reaction system is as follows:
[0142] The PCR procedure is as follows:
[0143] Following the manufacturer's instructions, the PCR products were purified into DNA fragments of the target size using a gel extraction kit. The purified target DNA fragments were then subjected to recombination reactions using the NEBuilder HiFi DNA Assembly Cloning Kit. The recombination system is as follows:
[0144] The reaction system was allowed to react at 50°C for 30 minutes.
[0145] 10 μL of the recombinant product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 h to activate and reactivate Escherichia coli DH5α competent cells.
[0146] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0147] Sequencing was used to verify the correct ligation of the E. coli DH5α clone, and the plasmid was extracted. The plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro was prepared for later use.
[0148] (3) Preparation of linearized plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro The plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro prepared in (2) was digested with BsaI restriction endonuclease. The digestion system is as follows:
[0149] The enzyme digestion system was reacted at 37°C for 1 hour.
[0150] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min. DNA fragments were excised from the agarose gel and recovered using a gel extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions, and finally eluted with ultrapure water. The DNA fragment is the linearized plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro, containing the encoding genes of the wild-type StaCas9 protein and sgRNAscaffold, with a size of 9152 bp.
[0151] The DNA concentration of the recovered linearized plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.
[0152] (4) Preparation of plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA The gRNA sequence was designed, and sticky end sequences (indicated by uppercase letters) corresponding to the linearized plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro were added to the sense and antisense strands, respectively. The two oligonucleotide single-stranded DNAs were then synthesized, and their specific sequences are shown below: gRNA: gctcggagatcatcattgcg Oligo-F: CACCgctcggagatcatcattgcg Oligo-R:AAACcgcaatgatgatctccgagc.
[0153] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained product was ligated to the linearized pAAV-CMV-StaCas9-Sta sgRNA-puro plasmid obtained in step (3) using DNA ligase (purchased from NEB).
[0154] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0155] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0156] The correctly ligated E. coli DH5α clone was subjected to culture and then the plasmid was extracted to obtain the plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA containing the target sgRNA sequence, which was then used for later use.
[0157] (5) Transfection of HEK293T cell line library containing target sequence of GFP reporter system with plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA The HEK293T cell line library containing the target sequence of the GFP reporter system was obtained as follows: a 25 bp protospacer (as the target sequence) and a 7 bp random sequence (as the PAM sequence) were inserted between the start codon ATG and the GFP coding sequence, resulting in a frameshift mutation that prevents GFP expression. This GFP gene containing the inserted fragment was started using the CMV promoter and constructed into a lentiviral expression vector. This sequence was then randomly inserted into the genome of HEK293T cells via lentiviral mediation, creating a stable GFP reporter cell line library. After the gene editing system successfully cleaved the target sequence, the cells' self-repair system caused some cells to recover the GFP reading frame, producing green fluorescence. Flow cytometry analysis was used to statistically analyze the percentage of GFP-positive cells to assess the editing capability and specificity of the gene editing system.
[0158] The transfection process includes the following steps: On day 0, the HEK293T cell line library containing the target sequence of the GFP reporter system was plated in 10 cm dishes as required for transfection, with the cell density controlled at 30%.
[0159] The HEK293T cell line library containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-target site-PAM-GFP, where the PAM sequence is a 7 bp random sequence and the target site sequence is GACGGCTCGGAGATCATCATTGCG.
[0160] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.
[0161] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (PEI) (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0162] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium of the HEK293T cell line library containing the target sequence of the GFP reporter system. The mixture was then placed in a 37°C, 5% CO2 incubator for further culture.
[0163] Five days after culture, the editing of target genes in the HEK293T cell line library by the CRISPR / StaCas9 system was observed under a fluorescence microscope. Under the microscope, cells transfected with the CRISPR / StaCas9 system produced green fluorescence, indicating that the CRISPR / StaCas9 gene editing system of this invention successfully edited the target genes in HEK293T cells. Subsequently, cells expressing GFP fluorescence were sorted using flow cytometry and further enriched for growth.
[0164] (6) Preparation of next-generation sequencing libraries The enriched HEK293T cell library was collected, and genomic DNA was extracted using a genomic DNA extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the kit.
[0165] The first round of PCR for library preparation was performed using Q5. ® PCR was performed using High-Fidelity 2× Master Mix. The PCR primers are shown below: F1 primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNgcgagaaaagccttgttt; R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNctgaacttgtggccgtttac.
[0166] The PCR reaction system is as follows:
[0167] The PCR procedure is as follows:
[0168] Perform sequencing, library preparation, and second-round PCR using Q5. ® PCR was performed using High-Fidelity 2× Master Mix. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0169] The PCR reaction system is as follows:
[0170] The PCR procedure is as follows:
[0171] The second-round PCR products were purified into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions, and the next-generation sequencing library was prepared.
[0172] (7) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqX Ten (illumina).
[0173] Based on the results obtained from next-generation sequencing, the PAM sequences of the HEK293T cell library were analyzed, and a PAM logo diagram was drawn, as shown below. Figure 1 As shown in the figure, the wild-type StaCas9 protein recognizes the PAM sequence NNG. This PAM sequence is very simple, indicating that the wild-type StaCas9 protein can be used for cell gene editing and shows good application potential.
[0174] Example 2 (1)-(3): Same as steps (1)-(3) of Example 1, to prepare linearized plasmid pAAV-CMV-StaCas9-StasgRNA-puro.
[0175] (4) Preparation of plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA Each gRNA was designed, and its sequence is shown in Table 1 below. The corresponding sticky end sequences (shown in bold) on both sides of the linearized plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro were added to the sense and antisense strands, respectively. These oligonucleotide single-stranded DNAs were then synthesized, and their specific sequences are also shown in Table 1 below.
[0176] Table 1. Sequences of gRNA and oligonucleotide single-stranded DNA
[0177] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained product was ligated to the linearized pAAV-CMV-StaCas9-Sta sgRNA-puro plasmid obtained in step (3) using DNA ligase (purchased from NEB).
[0178] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0179] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0180] The correctly ligated E. coli DH5α clone was subjected to culture and then the plasmid was extracted to obtain the plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA containing the target sgRNA sequence, which was then used for later use.
[0181] (5) Transfection of HEK293T cell line with plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA On day 0, HEK293T cells containing the target sequence were seeded in 6-well plates as needed for transfection, with a cell density of approximately 30%.
[0182] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-sgRNA to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.
[0183] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI, purchased from Polysciences). Add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0184] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 3 days.
[0185] (6) Preparation of next-generation sequencing libraries HEK293T cells were collected three days after editing, and genomic DNA was extracted using a genomic DNA extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the kit.
[0186] The first round of PCR for library preparation was performed using Q5. ® PCR was performed using High-Fidelity 2× Master Mix. The PCR primers are shown in Table 2 below. Table 2. List of primers for the first round of PCR in next-generation sequencing
[0187] The PCR reaction system is as follows:
[0188] The PCR procedure is as follows:
[0189] Perform sequencing, library preparation, and second-round PCR using Q5. ® PCR was performed using High-Fidelity 2× Master Mix. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0190] The PCR reaction system is as follows:
[0191] The PCR procedure is as follows:
[0192] The second-round PCR products were purified into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions, and the next-generation sequencing library was prepared.
[0193] (7) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqX Ten (illumina).
[0194] The editing efficiency of the CRISPR / Cas9 gene editing system of this invention at 12 target sites, as calculated by next-generation sequencing, is as follows: Figure 2 As shown in the figure, the X-axis represents the target site, and the Y-axis represents the editing efficiency (Indels%). This figure demonstrates that the gene editing system of this invention has an editing efficiency of at least 12%, and a maximum of at least 80%, indicating that the gene editing system of this invention can be effectively used for cell gene editing.
[0195] Example 3 (1)-(3): Same as steps (1)-(3) of Example 1, to prepare linearized plasmid pAAV-CMV-StaCas9-StasgRNA-puro.
[0196] (4) Preparation of plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-on-target-sgRNA or pAAV-CMV-StaCas9-Sta sgRNA-puro-mismatch-sgRNA The sequences of each on-target gRNA and mismatch gRNA were designed, and their corresponding oligonucleotide single-stranded DNA are shown in Table 3 below. The mismatch bases are shown as bold bases with underlined lines in the sequence listing.
[0197] Table 3. Oligonucleotide single-stranded DNA corresponding to on-target gRNA and mismatch gRNA
[0198] The oligonucleotide single-stranded DNA corresponding to the obtained on-target gRNA and the oligonucleotide single-stranded DNA corresponding to different mismatch gRNAs were annealed separately. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained products were ligated into the obtained linearized pAAV-CMV-StaCas9-Sta sgRNA-puro plasmid using DNA ligase (purchased from NEB).
[0199] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were then cultured at 37℃ for 1 h to activate and revitalize Escherichia coli DH5α competent cells.
[0200] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0201] The correctly ligated *E. coli* DH5α clones were subjected to culture in a shaking process, and then plasmids were extracted to obtain plasmids expressing the above-target sgRNA sequence pAAV-CMV-StaCas9-Sta sgRNA-puro-on-target-sgRNA and plasmids expressing different mismatch sgRNA sequences pAAV-CMV-StaCas9-Sta sgRNA-puro-mismatch-sgRNA, respectively, for later use.
[0202] (5) Transfection of HEK293T cell line with plasmids pAAV-CMV-StaCas9-Sta sgRNA-puro-on-target-sgRNA and pAAV-CMV-StaCas9-Sta sgRNA-puro-mismatch-sgRNA The obtained plasmids pAAV-CMV-StaCas9-Sta sgRNA-puro-on-target-sgRNA and pAAV-CMV-StaCas9-Sta sgRNA-puro-mismatch-sgRNA were transfected into the HEK293T cell line containing the target sequence using liposomes. The transfection process included the following steps: On day 0, the HEK293T cell line containing the target sequence of the GFP reporter system was seeded in 6-well plates as required for transfection, with the cell density controlled at 30%.
[0203] The HEK293T cell line containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-target-site-PAM-GFP, where the PAM sequence is CTGG and the target site sequence is GGCTCGGAGATCATCATTGCG (see [link to relevant documentation]). Figure 3 ).
[0204] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid containing only the StaCas9 coding sequence, or 2 μg of the plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro-on-target-sgRNA or pAAV-CMV-StaCas9-Sta sgRNA-puro-mismatch-sgRNA, and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.
[0205] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0206] The diluted plasmid and diluted transfection reagent were mixed and gently blown to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium of the HEK293T cell line containing the target sequence of the GFP reporter system. The cell line was then placed in a 37°C, 5% CO2 incubator for further culture.
[0207] The efficiency and off-target rate of the CRISPR gene editing system of this invention in editing target sequences were analyzed using flow cytometry.
[0208] Specifically, HEK293T cell lines cultured in a CO2 incubator for 3 days were collected, and their specificity was detected using a flow cytometer (BDBiosciences FACSCalibur). The GFP positivity rate was analyzed and plotted using FlowJo analysis software.
[0209] The specificity detection results of the CRISPR / StaCas9 gene editing system of the present invention in the HEK293T cell line containing the target sequence of the GFP reporter system are shown in the figure. Figure 3 The top bar shows a schematic diagram of the GFP reporter system, where a specific target sequence and PAM sequence are inserted between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. After the gene editing system successfully cuts the target sequence, the cell's own repair system allows some cells to recover the GFP reading frame, producing green fluorescence. Figure 3 In the bar chart below, the Y-axis represents the percentage of GFP-positive cells (%), and the X-axis represents the oligonucleotide single-stranded DNA sequences corresponding to on-target gRNA and mismatch gRNA. Figure 3 As can be seen, the CRISPR / StaCas9 gene editing system of this invention edited all target sites in the HEK293T cell line of the GFP reporter system. However, the proportion of gene editing mediated by mismatch gRNA was significantly lower than that mediated by on-target gRNA, indicating that the CRISPR / StaCas9 gene editing system of this invention has high editing activity, low off-target rate, and high specificity. Furthermore, in the study results of the CRISPR / StaCas9 gene editing system, mismatches only occurred in the M10 and M16 dibase mismatches, indicating that the CRISPR / StaCas9 gene editing system of this invention has extremely high requirements for perfect pairing between gRNA and target sequence, exhibiting low error tolerance and high safety in practical applications.
[0210] Example 4 (1) Preparation of mutant plasmid pAAV-CMV-enStaCas9-Sta sgRNA-puro The plasmid pAAV-CMV-StaCas9-Sta sgRNA-puro was constructed by introducing E887R and S998K mutations (StaCas9 with double mutations of E887R and S998K is denoted as enStaCas9). Specifically, the mutant bases corresponding to S998K were designed and synthesized at the 5' end of primers, and reverse PCR was performed. The primer sequences are as follows: Primer F-E887R: GCACCTGGATATCACCCA; Primer R-E887R: GTGATATCCAGGTGCCTTCCTAGCTTCTTGGAGTAG; Primer F-S998K: CCAGAAACAATATGCTAAGC; Primer R-S998K: GCATATTGTTTCTGGCGGATGTGCACATTGTCCACC.
[0211] The PCR reaction system is as follows:
[0212] The PCR procedure is as follows:
[0213] Purify the PCR products into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions. Then, use the NEBuilder HiFi DNA Assembly Cloning Kit to perform recombination reactions on the purified target DNA fragments. The recombination system consisted of: 0.1 μg DNA fragment, 5 μL NEBuilder® HiFi DNA Assembly Master Mix, and water to a final volume of 10 μL. Incubate the reaction at 50°C for 30 minutes.
[0214] 10 μL of the recombinant product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0215] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0216] The correctly ligated E. coli DH5α clone was subjected to culture and then the plasmid was extracted to obtain the plasmid pAAV-CMV-enStaCas9-Sta sgRNA-puro containing the E887R and S998K mutations, which was then used for later use.
[0217] (2) Preparation of mutant plasmid PAAV-CMV-enStaCas9-Sta sgRNA-puro-sgRNA The primers for constructing each gRNA and its corresponding plasmid are designed and their sequences are shown in Table 4. The plasmid PAAV-CMV-enStaCas9-Sta sgRNA-puro prepared in (1) is used as a template for PCR reaction. The gRNA and primer sequences are as follows: Table 4. gRNA and primer sequences
[0218] The PCR reaction system is as follows:
[0219] The PCR procedure is as follows:
[0220] Purify the PCR products into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions. Then, use the NEBuilder HiFi DNA Assembly Cloning Kit to perform recombination reactions on the purified target DNA fragments. The recombination system consisted of: 0.1 μg DNA fragment, 5 μL NEBuilder® HiFi DNA Assembly Master Mix, and water to a final volume of 10 μL. Incubate the reaction at 50°C for 30 minutes.
[0221] 10 μL of the recombinant product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0222] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0223] The correctly ligated E. coli DH5α clone was subjected to culture by shaking, and then the plasmid was extracted to obtain the plasmid pAAV-CMV-enStaCas9-Sta sgRNA-puro-sgRNA, which was then used for later use.
[0224] (3) Transfection of HEK293T cells with plasmid pAAV-CMV-enStaCas9-Sta sgRNA-puro-sgRNA The plasmids pAAV-CMV-enStaCas9-Sta sgRNA-puro-sgRNA prepared in (2) were transfected into HEK293T cells to evaluate the editing capability of the CRISPR / Cas9 gene editing system containing the enStaCas9 protein. The transfection steps are as follows: On day 0, HEK293T cells were seeded in 48-well plates as needed for transfection, with the cell density controlled at 30%.
[0225] Day 1, transfection was performed. The transfection process is as follows: Take 0.5 μg of the plasmid pAAV-CMV-enStaCas9-Sta sgRNA-puro-sgRNA to be transfected and add it to 12.5 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.
[0226] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 1 μL of Lipofectamine® 2000 or PEI to 12.5 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0227] The diluted plasmid and diluted transfection reagent were mixed and gently blown to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then dropped into the culture medium of HEK293T cells. The cells were then placed in a 37°C, 5% CO2 incubator for further culture.
[0228] 48 hours after transfection, the medium was replaced with one containing 1 μg / mL puromycin (Puro) and cultured for another 48 hours. The medium was then replaced with the normal medium and cultured for 2 days before the cells were harvested.
[0229] (4) Preparation of next-generation sequencing libraries Collect enriched HEK293T cells, centrifuge to remove supernatant, resuspend the cell pellet in 20 μL QuickExtract DNAExtraction Solution (purchased from Lucigen) and lyse. Lysis program: 65 ℃ for 6 min, 98 ℃ for 2 min, and store at 4 ℃.
[0230] The first round of PCR for library preparation was performed using Q5® High-Fidelity 2× Master Mix. The PCR primers are shown in Table 5. Table 5. List of primers for the first round of PCR in next-generation sequencing
[0231] The PCR reaction system is as follows:
[0232] The PCR procedure is as follows:
[0233] For the second round of PCR for sequencing library preparation, Q5® High-Fidelity 2× Master Mix was used. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNN NNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNN NNGTGACTGGAGTTCAGACGTGTG.
[0234] The PCR reaction system is as follows:
[0235] The PCR procedure is as follows:
[0236] The second-round PCR products were purified into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions, and the next-generation sequencing library was prepared.
[0237] (5) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqX Ten (illumina).
[0238] Based on the results obtained from next-generation sequencing, the indels of enStaCas9 at endogenous sites in HEK293T cells were analyzed, and an editing efficiency histogram was plotted, such as... Figure 4As shown in the figure, the CRISPR / Cas9 gene editing system containing single mutations of E887R and S998K, as well as double mutations of E887R and S998K, can all effectively perform gene editing. The enStaCas9 system, containing the double mutations of E887R and S998K, exhibits superior editing efficiency. Specifically, the editing efficiency of the E887R mutant is 4%-31%, with an average efficiency of 12%, which is 1.6 times that of the wild-type protein; the editing efficiency of the S998K mutant is 5%-46%, with an average efficiency of 16%, which is 2.1 times that of the wild-type protein; and the editing efficiency of enStaCas9, containing the double mutants of E887R and S998K, is 13%-36%, with an average efficiency of 28%, which is 3.7 times that of the wild-type protein.
[0239] The sequences used in this invention: SEQ ID NO: 1: Wild-type StaCas9 protein sequence SEQ ID NO: 2: enStaCas9 mutant protein sequence SEQ ID NO: 3: The coding sequence of wild-type StaCas9 protein SEQ ID NO: 4: Coding sequence of the enStaCas9 mutant protein SEQ ID NO: 5: Single-stranded guide RNA scaffold sequence GUUUUAGUACUCUGGAAACAGAAUCUACUAAAACAAGACAUUAUGUCGUGUUUAUCCCAUCAUUCUUGAUGGGA SEQ ID NO: 6: DNA sequence of a single-stranded guide RNA scaffold. GTTTTAGTACTCTGGAAACAGAATCTACTAAAACAAGACATTATGTCGTGTTTATCCCATCATTCTTGATGGGA SEQ ID NO: 7: StaCas9-E887R mutant protein sequence SEQ ID NO: 8: StaCas9-S998K mutant protein sequence SEQ ID NO: 9: Coding sequence of the StaCas9-E887R mutant protein SEQ ID NO: 10: Coding sequence of the StaCas9-S998K mutant protein
Claims
1. A StaCas9 mutant protein, said StaCas9 mutant protein comprising mutations at E887 and / or S998 relative to the amino acid sequence of the wild-type StaCas9 protein as shown in SEQ ID NO:1; Preferably, the mutation includes E887R, S998K, or E887R and S998K.
2. A conjugate, said conjugate comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein as described in claim 1. The homolog is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; b) Modified parts; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as c) An optional adapter for connecting the StaCas9 protein or its homolog to the modified portion.
3. A fusion protein, said fusion protein comprising: a) StaCas9 protein or its homologs, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein as described in claim 1. The homolog is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; b) Other proteins or peptides; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as c) Optional adapters for connecting the StaCas9 protein or its homologs to the other protein or polypeptide.
4. A single-stranded guide RNA comprising a CRISPR scaffold sequence, said CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identical to the nucleic acid sequence shown in SEQ ID NO: 5 and retains the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence modified based on the nucleic acid sequence described in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO:
5. For example, the modification may be one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening. For example, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the base sequence.
5. The single-stranded guide RNA according to claim 4, wherein, The single-stranded guide RNA further includes a CRISPR spacer sequence at the 3' end of the CRISPR scaffold sequence. The CRISPR spacer sequence is a sequence of 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides (preferably 20 nucleotides) in length and is complementary to the target sequence.
6. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the StaCas9 mutant protein of claim 1 or a homolog of the StaCas9 mutant protein, the conjugate of claim 2, or the fusion protein of claim 3, wherein the homolog of the StaCas9 mutant protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the StaCas9 mutant protein and retains the biological activity of the StaCas9 mutant protein; For example, the isolated nucleic acid molecule contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO:
10.
7. The isolated nucleic acid molecule according to claim 6, wherein the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding the single-stranded guide RNA according to any one of claims 4 to 5; For example, the isolated nucleic acid molecule contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and also contains the nucleic acid sequence shown in SEQ ID NO:
6.
8. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA as described in any one of claims 4 to 5; For example, the isolated nucleic acid molecule contains the nucleic acid sequence shown in SEQ ID NO: 6, and preferably also contains a nucleic acid sequence encoding a CRISPR spacer sequence.
9. The isolated nucleic acid molecule according to claim 8, wherein the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, wherein the homolog of the wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type StaCas9 protein and retains the biological activity of the wild-type StaCas9 protein; For example, the isolated nucleic acid molecule contains the nucleic acid sequence shown in SEQ ID NO: 6 and also contains the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
10. A vector comprising a nucleic acid sequence encoding the StaCas9 mutant protein of claim 1 or a homolog of the StaCas9 mutant protein, the conjugate of claim 2, or the fusion protein of claim 3, wherein the homolog of the StaCas9 mutant protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 mutant protein and retains the biological activity of the StaCas9 mutant protein; For example, the vector contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10; For example, the vector is a plasmid vector such as pUC19 vector, attachor vector, pAAV2_ITR vector, retroviral vector, lentiviral vector, adenovirus vector, or adeno-associated virus vector.
11. The carrier according to claim 10, wherein, The vector further comprises a nucleic acid sequence encoding the single-stranded guide RNA as described in any one of claims 4 to 5; For example, the vector contains the nucleic acid sequence or degenerate sequence shown in SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and also contains the nucleic acid sequence shown in SEQ ID NO:
6.
12. A vector comprising a nucleic acid sequence encoding a single-stranded guide RNA as described in any one of claims 4 to 5; For example, the vector contains the nucleic acid sequence shown in SEQ ID NO: 6, and preferably also contains a nucleic acid sequence encoding a CRISPR spacer sequence.
13. The vector of claim 12, wherein the vector further comprises a nucleic acid sequence encoding a wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, wherein the homolog of the wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type StaCas9 protein and retains the biological activity of the wild-type StaCas9 protein; For example, the vector contains the nucleic acid sequence shown in SEQ ID NO: 6 and also contains the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
14. A CRISPR / StaCas9 gene editing system, comprising: 1) Protein components, which include StaCas9 protein or its homologs, conjugates or fusion proteins, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein as described in claim 1; The homolog is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence shown by the StaCas9 protein and retains the biological activity of the StaCas9 protein; The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as 2) A nucleic acid component comprising: the single-stranded guide RNA as described in any one of claims 4 to 5; Furthermore, the protein component and the nucleic acid component bind together to form a complex; For example, the protein component comprises: wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and the nucleic acid component comprises the nucleic acid sequence shown in SEQ ID NO:
5.
15. A cell comprising: an isolated nucleic acid molecule according to any one of claims 6 to 8, or a vector according to any one of claims 9 to 11; For example, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell being, for example, a plant cell or an animal cell, the animal cell being, for example, a mammalian cell such as a human cell.
16. A method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: Contact any of the following (1) to (4) with the target sequence in the intracellular or in vitro environment: (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA according to any one of claims 4 to 5, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein as described in claim 1. The homolog is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein; The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; For example, wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5; (2) The vector according to claim 10 or a vector comprising a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, and the vector according to claim 12; Wherein, the homolog of the wild-type StaCas9 protein is a homolog that has an amino acid sequence identity of at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% and retains the biological activity of the wild-type StaCas9 protein; For example, vectors containing the nucleic acid sequences or degenerate sequences shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:9 or SEQ ID NO:10, and vectors containing the nucleic acid sequence shown in SEQ ID NO:6; (3) The carrier according to claim 11 or 13; or (4) The CRISPR / StaCas9 gene editing system according to claim 14; Upon contact with the target sequence, the StaCas9 protein or its homologs, conjugates or fusion proteins recognize the PAM sequence 5'-NNG located at the 3' end of the target sequence. For example, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell is, for example, a plant cell or an animal cell, the animal cell is, for example, a mammalian cell such as a human cell; For example, the gene editing includes one or more of the following: gene knockout of the target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking. For example, the single base conversion includes the conversion of adenine to guanine, cytosine to thymine, or cytosine to uracil.
17. The method according to claim 16, wherein, The CRISPR spacer sequence of the single-stranded guide RNA forms a completely complementary base pairing structure with the target sequence, while forming an incomplete complementary base pairing structure with the non-target sequence. For example, the incomplete base complementary pairing structure includes one or more structures, such as two or more base mismatches.
18. A kit for gene editing of a target sequence in an intracellular or in vitro environment, comprising: i. Choose any one of the following (1) to (6): (1) StaCas9 protein or its homologs, conjugates or fusion proteins, and the single-stranded guide RNA according to any one of claims 4 to 5, wherein: The StaCas9 protein is the wild-type StaCas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the StaCas9 mutant protein as described in claim 1. The homolog is an amino acid sequence that has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the StaCas9 protein and retains the biological activity of the StaCas9 protein. The conjugate comprises: the StaCas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the StaCas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the StaCas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; For example, wild-type StaCas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or StaCas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7 and SEQ ID NO: 8, and single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5; (2) The isolated nucleic acid molecule according to claim 6 or the isolated nucleic acid molecule containing a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, and the isolated nucleic acid molecule according to claim 8, wherein the homolog of the wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type StaCas9 protein and retains the biological activity of the wild-type StaCas9 protein; For example, isolated nucleic acid molecules containing the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and isolated nucleic acid molecules containing the nucleic acid sequence shown in SEQ ID NO: 6; (3) The isolated nucleic acid molecule according to claim 7 or 9; (4) The vector according to claim 10 or a vector comprising a nucleic acid sequence encoding the wild-type StaCas9 protein as shown in SEQ ID NO: 1 or a homolog of the wild-type StaCas9 protein, and the vector according to claim 12, wherein the homolog of the wild-type StaCas9 protein is a homolog whose amino acid sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type StaCas9 protein and retains the biological activity of the wild-type StaCas9 protein; For example, vectors containing the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 9 or SEQ ID NO: 10, and vectors containing the nucleic acid sequence shown in SEQ ID NO: 6; (5) The carrier according to claim 11 or 13; or (6) The CRISPR / StaCas9 gene editing system according to claim 14; as well as ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.