Hsp8Cas9 or mutant thereof, gene editing system containing Hsp8Cas9 or mutant thereof and application
By developing the Hsp8Cas9 protein and its mutants, the problem of packaging the SpCas9 protein in AAV viruses has been solved, enabling efficient and highly specific gene editing. It is suitable for adeno-associated virus vectors and has broad potential for gene editing applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-10
AI Technical Summary
In existing CRISPR/Cas9 systems, the SpCas9 protein is difficult to package effectively into AAV viruses due to its complex PAM sequence or low editing activity, which limits its clinical application.
Develop an Hsp8Cas9 protein and its mutants, which have a small number of amino acids and high editing activity, and can form an effective CRISPR/Cas9 gene editing system with single-stranded guide RNA, suitable for adeno-associated virus vectors.
It achieves efficient and highly specific gene editing, reduces off-target rates, is suitable for gene editing needs in different environments, and has broad clinical application prospects.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, specifically to an Hsp8Cas9 or a mutant thereof, a gene editing system containing Hsp8Cas9 or a mutant thereof, and related applications. Background Technology
[0002] CRISPR / Cas9 is an acquired immune system evolved by bacteria and archaea to defend against invasion by exogenous viruses or plasmids. In the CRISPR / Cas9 system, when crRNA (CRISPR-derived RNA), tracrRNA (trans-activating RNA), and Cas9 protein form a complex, they can recognize the PAM (Protospacer Adjacent Motif) sequence at the target site. The crRNA forms a complementary structure with the target DNA sequence, while the Cas9 protein performs the function of cutting the DNA, causing DNA breakage. TracrRNA and crRNA can fuse into single-stranded guide RNA (sgRNA) through a linker sequence. After DNA breakage, repair is handled by two main DNA damage repair mechanisms within the cell: non-homologous end-joining (NHEJ) and homologous recombination (HR). NHEJ repair results in base deletions or insertions, which can be used for gene knockout; when a homologous template is provided, HR repair can be used for site-specific gene insertion and precise base substitution.
[0003] Beyond basic research, CRISPR / Cas9 also holds broad clinical application prospects. The CRISPR / Cas9 system can be used for gene therapy, which involves introducing Cas9 and sgRNA into the body. Currently, the most effective expression vector used in gene therapy is adeno-associated virus (AAV). However, AAV viral packaging has requirements on DNA length, generally not exceeding 4.5 kb. Therefore, although SpCas9 has been widely used in basic research due to its simple PAM sequence (recognizing NGG) and high activity, the 1368 amino acids in the SpCas9 protein, along with the sgRNA and promoter required for gene editing systems, make it difficult to effectively package into AAV viruses, limiting its clinical application. To overcome this problem, several smaller Cas9 proteins have been developed, such as SaCas9 (PAM sequence NNGRRT), St1Cas9 (PAM sequence NNAGAW), NmCas9 (PAM sequence NNNNGATT), Nme2Cas9 (PAM sequence NNNNCC), and CjCas9 (PAM sequence NNNNRYAC). However, these Cas9 proteins are either prone to off-target cleavage (i.e., non-target site cleavage), have complex PAM sequences, or have low editing activity, making them difficult to widely apply.
[0004] Therefore, finding a small CRISPR / Cas9 system with high editing activity, high specificity, and simple PAM sequences is the hope for solving the above problems. Summary of the Invention
[0005] To address the aforementioned problems, the inventors, through repeated research, discovered an Hsp8Cas9 protein and its corresponding single-stranded guide RNA, which together constitute an effective CRISPA / Cas9 gene editing system, thus completing this invention.
[0006] This application aims to provide Hsp8Cas9 or its mutants, the nucleotide sequence encoding the protein or its mutants, and the expression vector thereof. The Hsp8Cas9 or its mutants consist of only 1028 amino acids, which is fewer than existing Cas9 proteins, and can be efficiently packaged into expression vectors such as adeno-associated virus vectors.
[0007] Therefore, in a first aspect, the present invention provides an Hsp8Cas9 mutant protein, wherein the Hsp8Cas9 mutant protein contains one or more mutations at V895, V941, Y942, and H585 relative to the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1.
[0008] In a second aspect, the present invention provides a conjugate comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein. b) Modified parts; and c) An optional adapter for connecting the Cas9 protein or its homolog to the modified portion.
[0009] In a third aspect, the present invention provides a fusion protein comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein. b) Other proteins or peptides; and c) Optional adapters for connecting the Cas9 protein or its homologs to the other protein or polypeptide.
[0010] In a fourth aspect, the present invention provides a single-stranded guide RNA comprising a CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence having at least 90% sequence identity with the nucleic acid sequence shown in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence obtained by modifying the nucleic acid sequence described in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5.
[0011] In a fifth aspect, the present invention provides an isolated nucleic acid molecule that encodes the nucleic acid sequence of the Hsp8Cas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention.
[0012] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0013] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding an Hsp8Cas9 mutant protein or homolog thereof of the first aspect of the present invention, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention.
[0014] In an eighth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0015] In a ninth aspect, the present invention provides a CRISPR / Cas9 gene editing system, the CRISPR / Cas9 gene editing system comprising: 1) Protein components, comprising the nucleic acid sequences of the Cas9 protein, its homologs, conjugates, or fusion proteins, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein. The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified moiety or the other protein or polypeptide; and 2) Nucleic acid components, comprising: single-stranded guide RNA of the fourth aspect of the present invention; Furthermore, the protein component and the nucleic acid component combine with each other to form a complex.
[0016] In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules of the fifth and / or sixth aspects of the present invention, or a vector of the seventh and / or eighth aspects of the present invention.
[0017] In an eleventh aspect, the present invention provides a method for gene editing in an intracellular or in vitro environment, the method comprising contacting any one of the following (1) to (4) with a target sequence in an intracellular or in vitro environment: (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein. The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or contains a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (3) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; or (4) The CRISPR / Cas9 gene editing system of the ninth aspect of the present invention; Upon contact with the target sequence, the Cas9 protein or its homologs, conjugates, or fusion proteins recognize the PAM sequence NNNNRYA located at the 3' end of the target sequence.
[0018] In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: i. Choose any one of the following (1) to (6): (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein. The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The fifth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or an isolated nucleic acid molecule containing a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the sixth aspect of the present invention does not contain an isolated nucleic acid molecule encoding the nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (3) The fifth aspect of the present invention comprises a nucleic acid molecule that encodes the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention and comprises a nucleic acid sequence that encodes the single-stranded guide RNA of the fourth aspect of the present invention; or the sixth aspect of the present invention comprises a nucleic acid molecule that encodes the single-stranded guide RNA of the fourth aspect of the present invention and comprises a nucleic acid sequence that encodes the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; (4) The seventh aspect of the present invention does not contain a vector encoding a nucleic acid sequence of a single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the eighth aspect of the present invention does not contain a vector encoding a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (5) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; or (6) The CRISPR / Cas9 gene editing system of the ninth aspect of the present invention; as well as ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.
[0019] The beneficial technical effects of the present invention include at least one or more of the following: The wild-type or mutant Hsp8Cas9 protein of the present invention, or its homologs, conjugates or fusion proteins, have only 1028 amino acids compared to the existing Cas9 protein in the prior art. The fewer amino acids, the more effectively they can be packaged into vectors such as adeno-associated virus vectors, thus solving the problem in the prior art that the Cas9 protein is too large to be packaged into adeno-associated virus along with sgRNA. Furthermore, the CRISPR / Cas9 gene editing system of the present invention has high editing efficiency and can edit target DNA with high specificity and low off-target rate; Furthermore, the CRISPR / Cas9 gene editing system of this application can be designed with sgRNA that can perform base complementary pairing with the target DNA sequence to be edited according to the needs of the DNA sequence to be edited, and the sgRNA can be modified to a certain extent in a manner known in the art. Therefore, it can meet the needs of different gene editing in different environments and has broad application prospects in the field of gene editing. Attached Figure Description
[0020] Figure 1 A shows a micrograph of edited GFP reporter cells. Figure 1 B shows a fluorescence image of a GFP reporter cell line 5 days after editing using the CRISPR / Hsp8Cas9 gene editing system. Figure 1 C shows a schematic diagram of the CRISPR / Hsp8Cas9 gene editing system recognizing PAM sequences.
[0021] Figure 2This diagram illustrates the efficiency results of the CRISPR / Hsp8Cas9 gene editing system and 23 modified CRISPR / Hsp8Cas9 gene editing systems after gene editing at a single target site.
[0022] Figure 3 This diagram illustrates the editing efficiency results of the CRISPR / Hsp8Cas9 gene editing system and the CRISPR / Hsp8Cas9 gene editing system based on 7 single mutants and combined mutants after gene editing at two target sites.
[0023] Figure 4 This diagram illustrates a comparison of the editing efficiency of CRISPR / Hsp8Cas9, CRISPR / enHsp8Cas9, and CRISPR / SpCas9 gene editing systems after performing gene editing on eight target sites.
[0024] Figure 5 This diagram illustrates the results of specific detection of double base mismatches at a single target site using the CRISPR / Hsp8Cas9 and CRISPR / enHsp8Cas9 gene editing systems.
[0025] Figure 6 This diagram illustrates a comparison of detection results between a guided editing system based on enHsp8Cas9 and a guided editing system based on SpCas9-NG, showing precise editing at three target sites. Detailed Implementation
[0026] Unless otherwise stated, the scientific and technical terms used in this application have the meanings commonly understood by those skilled in the art. Definitions and explanations of relevant terms are provided below for a better understanding of the invention.
[0027] The specific technical solution of this application will be described in detail below.
[0028] The terms “Cas9 protein,” “Hsp8Cas9,” and “Cas” used herein are interchangeable and refer to RNA-guided nucleases, including the Cas9 protein or its functionally active fragments. Therefore, in this document, the Cas9 protein may refer to the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1, a mutant of the wild-type Hsp8Cas9 protein (such as enHsp8Cas9), or both, as the case may be. The Cas9 protein is a protein component of the CRISPR / Cas9 genome editing system that, guided by single-stranded guide RNA (sgRNA), targets and cleaves DNA target sequences, forming DNA double-strand breaks (DSBs). DNA double-strand breaks can activate the cell’s inherent repair mechanisms of non-homologous end-joining (NHEJ) and homologous recombination (HR), thereby repairing DNA damage in the cell. During the repair process, the specific DNA sequence is edited at specific sites.
[0029] The terms “single-stranded guide RNA” and “sgRNA” (single guided RNA) used herein are interchangeable and have the meanings commonly understood by those skilled in the art. Generally, a single-stranded guide RNA may comprise a scaffold sequence and a guide sequence, which is also referred to herein as guide RNA (or gRNA). In the context of an endogenous CRISPR system, the guide sequence is also referred to as a spacer sequence. In some cases, the guide sequence is any polynucleotide sequence that is sufficiently similar to a target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas9 complex to said target sequence. In some embodiments, when optimally aligned, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of those skilled in the art. For example, there are publicly available and commercially available comparison algorithms and programs, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.
[0030] As used herein, the term "CRISPR / Cas9 complex" refers to a complex formed by the binding of a single-stranded guide RNA or mature crRNA:tracrRNA hybrid to a Cas9 protein (e.g., wild-type Hsp8Cas9 protein or a mutant thereof), which contains a guide sequence that hybridizes to a target sequence and thereby enables the Cas9 protein to bind to said target sequence. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the single-stranded guide RNA or mature crRNA.
[0031] Therefore, in the formation of a CRISPR / Cas9 complex, the "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, where hybridization between the target sequence and the guide sequence will promote Cas9's activity, such as cleavage of the target sequence. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote Cas9's activity. The target sequence can include any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In other cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast.
[0032] As used herein, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas protein used, but the PAM is typically a 2-5 base sequence adjacent to the protospacer sequence (target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas protein.
[0033] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” used herein are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.
[0034] The terms “polypeptide,” “peptide,” and “protein” as used herein are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, and to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.
[0035] The terms “sequence identity” or “homology” as used herein have their generally accepted meaning in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of that molecule. (See, e.g., Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; SequenceAnalysis in Molecular Biology, von Heinje, G., Academic Press, 1987; andSequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While many methods exist for measuring the identity between two polynucleotides or peptides, the term "identity" is known to those skilled in the art to refer to conserved amino acid substitutions in peptides or proteins that can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a peptide does not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0036] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. A vector is called an expression vector when it enables the expression of a protein encoded by the inserted polynucleotide, or when it enables transcription of the inserted polynucleotide (e.g., to generate mRNA or functional RNA). Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to, plasmid vectors and viral vectors. Vectors may also contain various regulatory sequences that regulate expression. The terms "regulatory sequence" and "regulatory element" are used interchangeably herein, referring to a nucleotide sequence located upstream (5' non-coding sequence), midway, or downstream (3' non-coding sequence) of a coding sequence that affects transcription, RNA processing, stability, or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. These regulatory sequences may originate from different sources or from the same source but arranged in a manner different from what is typically found naturally. Additionally, vectors may contain a replication initiation site.
[0037] As used herein, the term "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.
[0038] As used in this article, the term "constitutive promoter" refers to a promoter that generally causes gene expression in most cell types and under most conditions. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to promoters that are primarily, but not necessarily, expressed specifically in one tissue or organ, and may also be expressed in a specific cell type. "Developmental regulatory promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses a manipulated DNA sequence in response to endogenous or exogenous stimuli (environment, hormones, chemical signals, etc.).
[0039] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.
[0040] As used in this article, the term "stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its subsequent generations.
[0041] The term "transient transformation" as used in this article refers to the introduction of nucleic acid molecules or proteins into cells to perform their functions without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.
[0042] The terms “identity,” “consistency,” or “homology” used in this article have the generally accepted meanings in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule (see, for example, Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While many methods exist for measuring the identity between two polynucleotides or peptides, the term "identity" is known to those skilled in the art to refer to conserved amino acid substitutions in peptides or proteins that can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a peptide does not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0043] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The complementarity percentage indicates the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementarity out of 10). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0044] As used in this paper, the term "strict condition" in relation to hybridization refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of a strict condition is described in Tijssen (1993), *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes*, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assay", Elsevier, New York.
[0045] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. Hydrogen bonds can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.
[0046] Proteins, their conjugates and fusion proteins As described above, the inventors unexpectedly discovered that wild-type and mutant Hsp8Cas9 proteins, when combined with sgRNA, can effectively form a CRISPR / Hsp8Cas9 gene editing system. The Hsp8Cas9 protein has a relatively small number of amino acids, high editing efficiency, and recognizes a relatively simple PAM sequence NNNNRYA. Therefore, the CRISPR / Hsp8Cas9 gene editing system of this invention has a wide editing range and can be effectively used for site-specific editing of target sequences. This completes the present invention.
[0047] Therefore, in a first aspect, the present invention provides an Hsp8Cas9 mutant protein, wherein the Hsp8Cas9 mutant protein contains one or more mutations at V895, V941, Y942, and H585 relative to the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1.
[0048] The wild-type Hsp8Cas9 protein of this invention belongs to the genus Helicobacter ( ). Helicobacter sp. Its UniProt access number is A0A2A2GRY8.
[0049] Wild-type Hsp8Cas9 protein is derived from Helicobacter sp. TUL The Cas9 protein of this invention has a sequence of 1028 amino acids as shown in SEQ ID NO:1. The inventors have verified that the wild-type Hsp8Cas9 protein of this invention has a simpler PAM (NNNNRYA) compared to the commonly used CjCas9 in the prior art, and that its editing efficiency is significantly improved after mutation.
[0050] In one implementation, the mutation includes one or more of V895R, V941R, Y942K, and H585A.
[0051] It is understood that wild-type Hsp8Cas9 has double-strand cleavage activity, and the Hsp8Cas9 mutant protein may include Cas9 proteins with no cleavage activity, or only single-strand cleavage activity, or double-strand cleavage activity.
[0052] In a preferred embodiment, the mutations include V941R and Y942K, and this combined mutant is also referred to herein as enHsp8Cas9 or the enHsp8Cas9 protein.
[0053] In a more preferred embodiment, the mutations include V941R, Y942K, and H585A, and this combined mutant possesses only single-strand cleavage activity, also referred to herein as enHsp8Cas9-H585A. A guide editor based on this combined mutant can precisely edit target sites in the HEK293T cell line, and at certain sites, its editing efficiency is significantly higher than that of the SpCas9-NG guide editor, indicating that this tool has broad application prospects in gene therapy.
[0054] It is understood that the Cas9 protein can be derivatized, for example, by linking it to other molecules (e.g., other proteins or peptides, detectable markers). Generally, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to single-stranded guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, in this invention, the Cas9 protein can be functionally linked (through chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular moieties, such as other proteins or peptides, detectable markers, pharmaceutical reagents, etc.
[0055] Specifically, the Cas9 protein can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the protein's ability to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the Cas9 protein targeted. For example, it can be linked to a detectable tag to facilitate the detection of the Cas9 protein. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the Cas9 protein.
[0056] Therefore, in a second aspect, the present invention provides a conjugate comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; b) Modified parts; and c) An optional adapter for connecting the Cas9 protein or its homolog to the modified portion.
[0057] In this invention, the so-called "biological activity" of the Cas9 protein refers to its activity in binding to single-stranded guide RNA, its endonuclease activity (including single-stranded cleavage activity and double-stranded cleavage activity), and / or its activity in binding to and cleaving a specific site of a target sequence under the guidance of guide RNA (gRNA), but is not limited thereto.
[0058] It is understandable that, in addition to the Cas9 protein itself, the Cas9 protein can also be combined with other substances, such as other proteins or tagged objects, to give it other functions.
[0059] Therefore, in one embodiment, the modified portion may be another protein or peptide, a detectable marker, or a combination thereof.
[0060] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0061] Epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).
[0062] Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.
[0063] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0064] The Cas9 protein of the present invention can be coupled, conjugated, or fused to the modified portion via a linker, or it can be directly linked to the modified portion without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0065] In a third aspect, the present invention provides a fusion protein comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; b) Other proteins or peptides; and c) Optional adapters for connecting the Cas9 protein or its homologs to the other protein or polypeptide.
[0066] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0067] Single-stranded guide RNA In a fourth aspect, the present invention provides a single-stranded guide RNA comprising a CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity with the nucleic acid sequence shown in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence obtained by modifying the nucleic acid sequence described in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5.
[0068] In one embodiment, the modification may be one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening.
[0069] In a further embodiment, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine, or ten bases relative to the base sequence.
[0070] In yet another embodiment, the single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the CRISPR scaffold sequence, the CRISPR spacer sequence being a sequence of 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides in length and capable of complementary pairing with the target sequence.
[0071] In a preferred embodiment, the CRISPR spacer sequence is a 22-nucleotide sequence that is complementary to the target sequence.
[0072] The single-stranded guide RNA can bind to the Cas9 protein, conjugate, or fusion protein mentioned above to form a complex. This complex can recognize the corresponding PAM and thereby bind to the target sequence, thus achieving the cleavage of the target sequence or gene editing.
[0073] Nucleic acid encoding and vectors In a fifth aspect, the present invention provides an isolated nucleic acid molecule that encodes the nucleic acid sequence of the Hsp8Cas9 mutant protein or its homolog of the first aspect of the present invention, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the Hsp8Cas9 mutant protein and retains the biological activity of the Hsp8Cas9 mutant protein.
[0074] In one embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence or degenerate sequence shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22.
[0075] In one embodiment, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0076] As an example, the isolated nucleic acid molecule comprises the nucleic acid sequence or degenerate sequence thereof shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and the isolated nucleic acid molecule further comprises the nucleic acid sequence shown in SEQ ID NO: 6.
[0077] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0078] In one embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 6.
[0079] In a preferred embodiment, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a CRISPR spacer sequence.
[0080] In one embodiment, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein.
[0081] In a preferred embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 6 and the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
[0082] After the isolated nucleic acid molecules of the present invention are transfected into corresponding cells using certain tools known in the art, such as expression vectors, the isolated nucleic acid molecules of the present invention can express the Hsp8Cas9 mutant protein or its homologs described above, the conjugates of the second aspect of the present invention or the fusion protein of the third aspect of the present invention, encoding the wild-type Hsp8Cas9 protein or its homologs as shown in SEQ ID NO: 1, and / or the single-stranded guide RNA described above, and perform corresponding functions, such as gene editing.
[0083] In addition, the isolated nucleic acid molecules of the present invention can express Hsp8Cas9 mutant protein or its homologs, conjugates of the second aspect of the present invention or fusion proteins of the third aspect of the present invention, and single-stranded guide RNA individually or separately, or can express the above expression products in one go. The choice of expression method depends on the specific circumstances.
[0084] Furthermore, the aforementioned expressive products have the corresponding effects and / or functions described above, which will not be repeated here for the sake of brevity.
[0085] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding an Hsp8Cas9 mutant protein or a homolog thereof of the first aspect of the present invention, a conjugate of the second aspect of the present invention, or a fusion protein of the third aspect of the present invention, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the Hsp8Cas9 mutant protein and retains the biological activity of the Hsp8Cas9 mutant protein.
[0086] In one embodiment, the vector comprises a nucleic acid sequence or a degenerate sequence thereof shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22.
[0087] In a preferred embodiment, the vector is a plasmid vector, a retroviral vector, an adenovirus vector, or an adeno-associated virus vector.
[0088] In a more preferred embodiment, the plasmid vector is the pAAV2_ITR plasmid vector.
[0089] In a further embodiment, the vector further comprises a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0090] As an example, the vector comprises the nucleic acid sequence or degenerate sequence thereof shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and the vector further comprises the nucleic acid sequence shown in SEQ ID NO: 6.
[0091] In an eighth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect of the present invention.
[0092] In one embodiment, the vector comprises the nucleic acid sequence shown in SEQ ID NO: 6 or a degenerate sequence thereof.
[0093] In a preferred embodiment, the vector further comprises a nucleic acid sequence encoding a CRISPR spacer sequence.
[0094] In one embodiment, the vector further comprises a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or a homolog as shown in SEQ ID NO: 1, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein.
[0095] In a preferred embodiment, the vector comprises the nucleic acid sequence shown in SEQ ID NO: 6 and the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
[0096] As described above, after the vector of the present invention is transfected into cells, the nucleic acid sequence cloned in the vector can be expressed as an Hsp8Cas9 mutant protein or its homolog, a conjugate of the second aspect of the present invention or a fusion protein of the third aspect of the present invention, a wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and / or the single-stranded guide RNA described above, and perform corresponding functions, such as gene editing.
[0097] Alternatively, multiple vectors, such as two vectors, can be transfected into cells. One vector expresses the Hsp8Cas9 mutant protein or its homologs, the conjugate of the second aspect of the invention, the fusion protein of the third aspect of the invention, or the wild-type Hsp8Cas9 protein or its homologs as shown in SEQ ID NO: 1, while the other vector expresses single-stranded guide RNA. Subsequently, the expressed Hsp8Cas9 mutant protein or its homologs, the conjugate of the second aspect of the invention, the fusion protein of the third aspect of the invention, or the wild-type Hsp8Cas9 protein or its homologs as shown in SEQ ID NO: 1, complexes with the expressed single-stranded guide RNA to form a complex, which performs the corresponding function, such as gene editing.
[0098] Alternatively, the nucleic acid sequence encoding the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, the fusion protein of the third aspect of the present invention, or the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the nucleic acid sequence encoding the single-stranded guide RNA can be cloned into a vector, such that after the vector is transfected into cells, it expresses both the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, the fusion protein of the third aspect of the present invention, or the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the single-stranded guide RNA, and performs the corresponding functions, such as gene editing.
[0099] CRISPR / Cas9 gene editing system In a ninth aspect, the present invention provides a CRISPR / Cas9 gene editing system, the CRISPR / Cas9 gene editing system comprising: 1) Protein components, comprising the nucleic acid sequences of the Cas9 protein, its homologs, conjugates, or fusion proteins, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified moiety or the other protein or polypeptide; and 2) Nucleic acid components, comprising: single-stranded guide RNA of the fourth aspect of the present invention; Furthermore, the protein component and the nucleic acid component combine with each other to form a complex.
[0100] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0101] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0102] As an example, the protein component comprises: a wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or an Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19 or SEQ ID NO: 21, and the nucleic acid component comprises a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5.
[0103] The CRISPR / Cas9 gene editing system of the present invention can be directly constructed from the Cas9 protein or its homologs, conjugates or fusion proteins described herein and the single-stranded guide RNA described herein, or it can be constructed from the expression products obtained by the vectors described herein.
[0104] In one embodiment, the wild-type Hsp8Cas9 protein or the Hsp8Cas9 mutant protein and the single-stranded guide RNA are obtained by expression using the vector described herein.
[0105] The CRISPR / Cas9 gene editing system of the present invention achieves the recognition, localization, cleavage and gene editing of target sequences through the combined action of the Cas9 protein or its homologs, conjugates or fusion proteins contained therein and single-stranded guide RNA.
[0106] The CRISPR / Cas9 gene editing system of this invention can precisely locate target sequences. "Precise location" has two meanings: first, the CRISPR / Cas9 gene editing system itself can recognize and bind to the target sequence; second, the CRISPR / Cas9 gene editing system can bring other proteins fused with the Cas9 protein or proteins that specifically recognize the sgRNA to the location of the target sequence.
[0107] When the CRISPR / Cas9 gene editing system comes into contact with target DNA in the intracellular or in vitro environment, the Cas9 protein or its homologs, conjugates, or fusion proteins recognize the PAM sequence NNNNRYA on the target DNA, and the 5' 22 bp sequence of the sgRNA forms a complementary base pair with the target DNA. Thus, the Cas9 protein or its homologs, conjugates, or fusion proteins can cleave the target site on the target DNA, and double-strand breaks occur in the target DNA under the cleavage action of the Cas9 protein or its homologs, conjugates, or fusion proteins. Furthermore, when the CRISPR / Cas9 gene editing system is in a cell, the cleaved DNA can be repaired through intracellular non-homologous end ligation repair or homologous recombination repair pathways, thereby enabling gene editing of the target DNA.
[0108] cell In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules of the fifth and / or sixth aspects of the present invention, or a vector of the seventh and / or eighth aspects of the present invention.
[0109] As an example, the cell can be a prokaryotic cell or a eukaryotic cell. For the eukaryotic cell, as an example, it can be a plant cell or an animal cell. For the animal cell, as an example, it can be a mammalian cell, such as a human cell.
[0110] method In an eleventh aspect, the present invention provides a method for gene editing in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (4) with a target sequence in an intracellular or in vitro environment: (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The seventh aspect of the present invention does not contain a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or contains a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the eighth aspect of the present invention does not contain a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (3) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; or (4) The CRISPR / Cas9 gene editing system of the ninth aspect of the present invention; Upon contact with the target sequence, the Cas9 protein or its homologs, conjugates, or fusion proteins recognize the PAM sequence NNNNRYA located at the 3' end of the target sequence.
[0111] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0112] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0113] As an example, for item (1) above, it can be: wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19 or SEQ ID NO: 21, and a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5.
[0114] As an example, for item (2) above, it can be: a vector containing any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20 or SEQ ID NO:22, and a vector containing the nucleic acid sequence shown in SEQ ID NO:6.
[0115] In a preferred embodiment, the gene editing includes gene knockout of target DNA, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, or chromatin imaging tracking.
[0116] In a more preferred embodiment, the single base conversion includes adenine to guanine, cytosine to thymine, or cytosine to uracil.
[0117] In one embodiment, the cell is a eukaryotic cell or a prokaryotic cell.
[0118] In a preferred embodiment, the eukaryotic cell includes mammalian cells or plant cells; In a more preferred embodiment, the mammalian cells include Chinese hamster ovary cells, young hamster kidney cells, mouse Sertoli cells, mouse mammary tumor cells, buffalo rat hepatocytes, rat hepatoma cells, monkey kidney CVI line transformed with SV40, monkey kidney cells, canine kidney cells, human cervical cancer cells, human lung cells, human hepatocytes, HIH / 3T3 cells, human U2-OS osteosarcoma cells, human A549 cells, human K562 cells, human HEK293 cells, human HEK293T cells, human HCT116 cells, or human MCF-7 cells.
[0119] In one embodiment, the CRISPR spacer sequence of the single-stranded guide RNA forms a fully complementary base pairing structure with the target sequence, and an incompletely complementary base pairing structure with the non-target sequence.
[0120] In this document, the incomplete base pairing structure refers to a structure that includes a portion of base pairing and a portion of non-base pairing, wherein the non-base pairing includes, for example, base mismatch and / or base bulge.
[0121] In a preferred embodiment, the incomplete base complementarity pairing includes one or more, for example, two or more base mismatches.
[0122] Therefore, the Cas9 protein or its homologs, conjugates, or fusion proteins of the present invention can cleave target sites on the target sequence, and the cleavage action of the Cas9 protein or its homologs, conjugates, or fusion proteins causes double-strand breaks in the target sequence. Furthermore, when the method is performed intracellularly, the cleaved target sequence can be repaired through intracellular non-homologous end joining repair or homologous recombination repair pathways, thereby achieving gene editing of the target sequence.
[0123] The CRISPR / Cas9 gene editing system and gene editing method using this invention have been experimentally shown to have an editing efficiency of up to 88.57% and an extremely low off-target rate. Therefore, this gene editing system can specifically edit target genes, exhibiting high editing efficiency and low off-target rate, and can be widely applied to gene editing in cells.
[0124] Reagent test kit In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: i. Choose any one of the following (1) to (6): (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA of the fourth aspect of the present invention, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein of the first aspect of the present invention. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; (2) The fifth aspect of the present invention does not contain an isolated nucleic acid molecule encoding a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention or an isolated nucleic acid molecule containing a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the sixth aspect of the present invention does not contain an isolated nucleic acid molecule encoding the nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (3) The fifth aspect of the present invention comprises a nucleic acid molecule that encodes the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention or the fusion protein of the third aspect of the present invention and comprises a nucleic acid sequence that encodes the single-stranded guide RNA of the fourth aspect of the present invention; or the sixth aspect of the present invention comprises a nucleic acid molecule that encodes the single-stranded guide RNA of the fourth aspect of the present invention and comprises a nucleic acid sequence that encodes the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; (4) The seventh aspect of the present invention does not contain a vector encoding a nucleic acid sequence of a single-stranded guide RNA of the fourth aspect of the present invention or a vector containing a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the eighth aspect of the present invention does not contain a vector encoding a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1. (5) A vector comprising, in the seventh aspect of the present invention, a nucleic acid sequence encoding the Hsp8Cas9 mutant protein or its homolog, the conjugate of the second aspect of the present invention, or the fusion protein of the third aspect of the present invention, and a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention; or a vector comprising, in the eighth aspect of the present invention, a nucleic acid sequence encoding the single-stranded guide RNA of the fourth aspect of the present invention, and a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1; or (6) The CRISPR / Cas9 gene editing system of the ninth aspect of the present invention; and ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.
[0125] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof.
[0126] In one embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0127] As an example, for item (1) above, it can be: wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1 or Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19 or SEQ ID NO: 21, and a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5.
[0128] As an example, for item (2) above, it can be: an isolated nucleic acid molecule containing any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 1, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and an isolated nucleic acid molecule containing the nucleic acid sequence shown in SEQ ID NO: 6.
[0129] As an example, for item (4) above, it can be: a vector containing the nucleic acid sequence or its degenerate sequence shown in any one of SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and a vector containing the nucleic acid sequence shown in SEQ ID NO: 6.
[0130] Of course, those skilled in the art will understand that the kit of the present invention may also contain other reagents that facilitate gene editing.
[0131] A brief description of the sequence involved in this invention. SEQ ID NO: 1: Amino acid sequence of Hsp8Cas9 protein SEQ ID NO: 2: Amino acid sequence of enHsp8Cas9 protein SEQ ID NO: 3: Coding sequence of Hsp8Cas9 protein SEQ ID NO: 4: Coding sequence of the enHsp8Cas9 protein SEQ ID NO: 5: Single-stranded guide RNA scaffold sequence SEQ ID NO: 6: DNA sequence of a single-stranded guide RNA scaffold. SEQ ID NO: 7: Amino acid sequence of Hsp8Cas9-V895R protein SEQ ID NO: 8: Coding sequence of Hsp8Cas9-V895R protein SEQ ID NO: 9: Amino acid sequence of Hsp8Cas9-V941R protein SEQ ID NO: 10: Coding sequence of Hsp8Cas9-V941R protein SEQ ID NO: 11: Amino acid sequence of Hsp8Cas9-Y942K protein SEQ ID NO: 12: Coding sequence of the Hsp8Cas9-Y942K protein SEQ ID NO: 13: Amino acid sequence of Hsp8Cas9-V895R&Y942K protein SEQ ID NO: 14: Coding sequence of the Hsp8Cas9-V895R&Y942K protein SEQ ID NO: 15: Amino acid sequence of Hsp8Cas9-V895R&V941R protein SEQ ID NO: 16: Coding sequence of Hsp8Cas9-V895R&V941R protein SEQ ID NO: 17: Amino acid sequence of Hsp8Cas9-V895R&Y942K&V941R protein SEQ ID NO: 18: Coding sequence of Hsp8Cas9-V895R&Y942K&V941R protein SEQ ID NO: 19: Amino acid sequence of Hsp8Cas9-H585A protein SEQ ID NO: 20: Coding sequence of Hsp8Cas9-H585A protein SEQ ID NO: 21: Amino acid sequence of Hsp8Cas9-V941R&Y942K&H585A protein SEQ ID NO: 22: The coding sequence of the Hsp8Cas9-V941R&Y942K&H585A protein. Example The invention will now be described with reference to the following embodiments, which are intended to be illustrative and not limiting. Those skilled in the art will understand that the embodiments provided herein are for the purpose of describing the invention in detail only and are not intended to limit the scope of protection claimed by the invention.
[0132] Unless otherwise specified, the experiments and methods described in the examples were generally performed according to conventional methods well known in the art and described in the various references. Furthermore, for conditions not specifically specified in the examples, conventional conditions or conditions recommended by the manufacturer were followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0133] Example 1 (1) Constructing plasmid pAAV2_Hsp8Cas9_ITR ① Download the amino acid sequence of the Hsp8Cas9 gene based on the search number A0A2A2GRY8 on UniProt. The corresponding amino acid sequence of the Hsp8Cas9 protein is shown in SEQ ID NO:1.
[0134] ② The amino acid sequence of the Hsp8Cas9 protein obtained above was codon optimized to obtain the gene sequence of Hsp8Cas9 protein that is highly expressed in human cells, as shown in SEQ ID NO:3.
[0135] ③ The Hsp8Cas9 protein highly expressed gene sequence obtained above, as shown in SEQ ID NO:3, was synthesized and constructed into the pAAV2_ITR backbone plasmid to obtain plasmid pAAV2_Hsp8Cas9_ITR.
[0136] (2) Constructing the plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold The pSKB plasmid (available commercially from the Addgene platform, catalog #62540) was digested with ClaI and XhoI restriction endonucleases. The digestion system consisted of 1 μg pSKB plasmid, 5 μL 10×rCutSmart buffer (from NEB), 1 μL ClaI and 1 μL XhoI restriction endonuclease (from NEB), and water to a final volume of 50 μL. The digestion was incubated at 37°C for 1 hour.
[0137] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0138] The target size DNA fragment was cut from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, it was eluted with ultrapure water.
[0139] Genes were synthesized based on the sgRNA scaffold sequence (SEQ ID NO: 5) and constructed on a linearized pSKB backbone to obtain the plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold.
[0140] (3) Preparation of linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold The plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold prepared in (2) was digested with BbsI restriction endonuclease. The digestion system consisted of 1 μg plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold, 5 μL 10×CutSmart buffer (purchased from NEB), 1 μL BbsI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion system was incubated at 37°C for 1 hour.
[0141] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0142] DNA fragments were excised from the agarose gel and recovered using a gel extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. The fragments were then eluted with ultrapure water. The resulting DNA fragment was a linearized plasmid containing the gene encoding the sgRNA scaffold, with a size of 4740 bp.
[0143] The DNA concentration of the recovered linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.
[0144] (4) Preparation of plasmid Hsp8Cas9-PSK-mU6-PAM2sgRNA-scaffold The gRNA sequence was designed, and sticky end sequences (indicated by uppercase letters) corresponding to the flanking sides of the linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold were added to its sense and antisense strands, respectively. The two oligonucleotide single-stranded DNAs were then synthesized, and their specific sequences are shown below: gRNA: ggtagaggcggccacgacctg Oligo-F:TTTGgagtagaggcggccacgacctgGT Oligo-R: TAAAACcaggtcgtggccgcctctactc.
[0145] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained product was ligated to the linearized Hsp8Cas9-PSK-mU6-sgRNA-scaffold plasmid obtained in step (3) using DNA ligase (purchased from NEB).
[0146] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0147] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0148] The correctly ligated E. coli DH5α clone was subjected to culture by shaking, and then the plasmid was extracted to obtain the plasmid Hsp8Cas9-PSK-mU6-PAM2sgRNA-scaffold containing the target sgRNA sequence, which was then used for later use.
[0149] (5) Transfection of HEK293T cell line library containing target sequence of GFP reporter system with plasmids expressing Cas protein and sgRNA The HEK293T cell line library containing the target sequence of the GFP reporter system was obtained as follows: a 24 bp protospacer (as the target sequence) and an 8 bp random sequence (as the PAM sequence) were inserted between the start codon ATG and the GFP coding sequence, resulting in a frameshift mutation that prevents GFP expression. This GFP gene containing the inserted fragment was started using a CMV promoter and constructed into a lentiviral expression vector. This sequence was then randomly inserted into the genome of HEK293T cells via lentiviral mediation, creating a stable GFP reporter cell line library. When the target sequence was cut using a gene editing system, the cells' self-repair system caused some cells to recover the GFP reading frame, producing green fluorescence. Flow cytometry analysis was used to statistically analyze the percentage of GFP-positive cells, which allowed for the assessment of the editing capability and specificity of the gene editing system.
[0150] The transfection process includes the following steps: ① On day 0, according to the transfection requirements, the HEK293T cell line library containing the target sequence of the GFP reporter system was plated in a 10cm dish, and the cell density was controlled at 30%.
[0151] The HEK293T cell line library containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-target site-PAM-GFP, where the PAM sequence is an 8 bp random sequence and the target site sequence is GAGAGTAGAGGCGGCCACGACCTG.
[0152] ②On day 1, transfection was performed. The transfection process is as follows: Add 10 μg of pAAV2_Hsp8Cas9_ITR plasmid and 5 μg of Hsp8Cas9-PSK-mU6-PAM2sgRNA-scaffold to 100 μL of Opti-MEM medium (purchased from Gibco) and gently pipette to mix.
[0153] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0154] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium of the HEK293T cell line library containing the target sequence of the GFP reporter system. The mixture was then placed in a 37°C, 5% CO2 incubator for further culture.
[0155] Five days after culture, the editing of target genes in the HEK293T cell line library by the CRISPR / Hsp8Cas9 system was observed under a fluorescence microscope. Under the microscope, cells transfected with the CRISPR / Hsp8Cas9 system produced green fluorescence, indicating successful editing of the target genes in the cells. Cell fluorescence images are shown below. Figure 1 As shown in B. Subsequently, cells expressing GFP fluorescence were sorted using a flow cytometry system and then further enriched and grown.
[0156] (6) Preparation of next-generation sequencing libraries ① Collect the enriched HEK293T cell library and extract genomic DNA using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.
[0157] ② For the first round of PCR library construction, use a 2×Q5 Mastermix for the PCR reaction. The PCR primers are shown below: F1 primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNgcgagaaaagccttgttt; R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNctgaacttgtggccgtttac.
[0158] The PCR reaction system is as follows:
[0159] The PCR procedure is as follows:
[0160] ③ For the second round of PCR for sequencing library preparation, use a 2×Q5 Mastermix for the PCR reaction. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0161] The PCR reaction system is as follows:
[0162] The PCR procedure is as follows:
[0163] ④ Use a gel extraction kit to purify the DNA fragments of the target size from the second round of PCR products, following the steps provided by the manufacturer. The next-generation sequencing library is now ready.
[0164] (7) Analysis of second-generation sequencing results.
[0165] The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0166] Based on the results obtained from next-generation sequencing, the PAM sequences of the HEK293T cell library were analyzed, and a PAM logo diagram was drawn, as shown below. Figure 1 As shown in C, the PAM sequence recognized by the Hsp8Cas9 protein is NNNNRYA, which is similar to the PAM of CjCas9 (NNNNRYAC), but simpler.
[0167] Example 2 (1) Construction of plasmid pAAV2_Hsp8Cas9_ITR and 23 modified pAAV2_vHsp8Cas9_ITR plasmids ①-③ The plasmid pAAV2_Hsp8Cas9_ITR was prepared using the same method as in Example 1 (1) ①-③.
[0168] ④ Based on plasmid pAAV2_Hsp8Cas9_ITR, the following 23 amino acid mutations were introduced via circular PCR: L57K, Y204K, Y204R, G257R, A289N, H343K, D472K, N546K, N608R, N608K, S665R, R668K, Y669R, M860K, M860R, V895R, V941R, Y942K, V945K, V945R, T947K, S949K, and Q986K. The primer sequences used are shown in Table 1 below. Table 1. Primer sequences for pAAV2_Hsp8Cas9_ITR circular PCR
[0169] The PCR reaction system is as follows:
[0170] The PCR procedure is as follows:
[0171] The PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel recovery kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.
[0172] The 23 recovered circular PCR gel products were transformed. 5 μl of each product was added to *E. coli* DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, then 900 μL of LB medium was added, and the cells were incubated at 37℃ for 1 hour to activate and reactivate the *E. coli* DH5α competent cells.
[0173] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance, incubated upside down in a 37°C incubator, and the obtained Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0174] The E. coli DH5α clones that were correctly sequenced and introduced into the mutation site were cultured by shaking, and then plasmids were extracted to obtain 23 modified pAAV2_vHsp8Cas9_ITR plasmids for later use.
[0175] (2) Construction of plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using a method similar to that in Example 1 (2), the plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was prepared using the sgRNA scaffold sequence (SEQ ID NO: 5).
[0176] (3) Preparation of linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using the same method as in Example 1 (3), the linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was prepared, and the DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). It was then stored for later use or at -20°C for long-term preservation.
[0177] (4) Preparation of plasmid Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold ① Design a gRNA sequence and add sticky end sequences (indicated by uppercase letters) to the sense and antisense strands of the linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold, respectively. Then synthesize these two oligonucleotide single-stranded DNA sequences, as shown below: A3-gRNA: cggggaaggaggatgccccagg Oligo-F:TTTGcggggaaggaggatgccccaggGT Oligo-R: TAAAACcctggggcatcctccttccccg.
[0178] ② Anneal the oligonucleotide single-stranded DNA corresponding to the sgRNA. After annealing, the resulting products are ligated to the linearized Hsp8Cas9-PSK-mU6-sgRNA-scaffold plasmid obtained in step (3) using DNA ligase (purchased from NEB) to obtain the corresponding sgRNA ligation products.
[0179] The annealing reaction described above included: mixing 1 μL of 100 μM oligo-F, 1 μL of 100 μM oligo-R, and 28 μL of water by shaking, then placing the mixture in a PCR instrument and running the annealing program. The annealing program was as follows: 95℃ for 5 min, 85℃ for 1 min, 75℃ for 1 min, 65℃ for 1 min, 55℃ for 1 min, 45℃ for 1 min, 35℃ for 1 min, 25℃ for 1 min, and storing at 4℃ with a cooling rate of 0.3℃ / s.
[0180] The above ligation process was performed according to the instructions for DNA ligase (purchased from NEB).
[0181] ③ Take 1 μL of the ligation product of the corresponding sgRNA obtained in step ② and add it to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.). Incubate on ice for 30 min, heat shock at 42℃ for 1 min, incubate on ice for 2 min, add 900 μL of LB medium, and incubate at 37℃ for 1 h to activate and revive Escherichia coli DH5α competent cells.
[0182] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0183] ④ Perform a shake culture on the correctly ligated E. coli DH5α clone verified by sequencing, and then extract the plasmid to obtain the plasmid Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold containing the above sgRNA sequence, for later use.
[0184] (5) pAAV2_Hsp8Cas9_ITR and 23 modified pAAV2_vHsp8Cas9_ITR plasmids were used in combination with plasmid Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold to transfect HEK293T cell lines containing target DNA. The above transfection process includes the following steps: ① On day 0, HEK293T cells containing the target DNA were seeded in 48-well plates as needed for transfection, with a cell density of about 30%.
[0185] ②On day 1, transfection was performed. The transfection process is as follows: 350 ng each of the pAAV2_Hsp8Cas9_ITR and 23 modified pAAV2_vHsp8Cas9_ITR plasmids were added to 12.5 μL of Opti-MEM medium (Gibco) along with 150 ng of Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold plasmid. The mixture was gently pipetted to obtain dilution A. Lipofectamine® 2000 (Invitrogen) or polyethyleneimine (PEI) (Polysciences) were then gently mixed. 0.5 μL of either Lipofectamine® 2000 or PEI was added to 12.5 μL of Opti-MEM medium (Gibco), gently mixed, and incubated at room temperature for 5 min to obtain dilution B.
[0186] Mix the dilution A and dilution B obtained above, gently pipette to mix, and obtain a mixture C containing the transfection reagent and the plasmid to be transfected. Let it stand at room temperature for 20 min, and then add the mixture C to the HEK293T cell culture medium containing the target DNA in ①. Then place it in a 37℃, 5% CO2 incubator and continue to culture for 3 days.
[0187] (6) Preparation of next-generation sequencing libraries ① Collect HEK293T cells containing target DNA that have been successfully transfected in step (5), and extract genomic DNA from HEK293T cells containing target DNA using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) and following the steps provided by the DNA kit.
[0188] ② The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers are shown below: F1 primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNGACTCCACCAACGCCGAC R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNTGGTCAGCACAGAGTGGCTA The PCR reaction system is as follows:
[0189] The PCR procedure is as follows:
[0190] ③ Perform the second round of PCR for library construction using a 2×Q5 Mastermix. The PCR primers are shown below: F1 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0191] The PCR reaction system is as follows:
[0192] The PCR procedure is as follows:
[0193] ④ Use a gel extraction kit to purify the DNA fragment of about 400 bp from the second round of PCR products, following the steps provided by the manufacturer. This completes the preparation of the next-generation sequencing library.
[0194] (7) Analysis of second-generation sequencing results The second-generation sequencing library obtained above was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0195] The editing efficiency of Hsp8Cas9 and 23 modified Hsp8Cas9 at the same endogenous site, calculated by next-generation sequencing, is as follows: Figure 2As shown, the X-axis represents mutation sites, and the Y-axis represents editing efficiency (Indels%). From Figure 2 As can be seen, compared with the unmodified Hsp8Cas9, the mutants V895R, V941R, and Y942K all exhibited significantly improved editing efficiency. This indicates that, in this invention, by rationally modifying key amino acid sites of Hsp8Cas9, gene editing performance can be significantly improved, thus making it effective for cellular gene editing.
[0196] Next, we will perform pairwise and three-site combinations of the three amino acid mutations V895R, V941R, and Y942K to further test the editing efficiency of the mutants.
[0197] Example 3 (1) Construction of pAAV2_Hsp8Cas9_ITR plasmid and modified pAAV2_v1Hsp8Cas9_ITR and pAAV2_v2Hsp8Cas9_ITR plasmids Genes were synthesized using the gene sequence shown in SEQ ID NO:3 and the gene sequences of 3 single mutants (V895R, V941R, Y942K) and 4 combined mutants (V895R&V941R, V895R&Y942K, V941R&Y942K, V895R&V941R&Y942K), and constructed into the pAAV2_ITR backbone plasmid to obtain the pAAV2_Hsp8Cas9_ITR plasmid, the 3 single mutant pAAV2_v1Hsp8Cas9_ITR plasmid, and the 4 combined mutant pAAV2_v2Hsp8Cas9_ITR plasmid.
[0198] (2) Construction of plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using the same method as in Example 1 (2), plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was prepared using the sgRNA scaffold sequence (SEQ ID NO: 5).
[0199] (3) Preparation of linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using the same method as in Example 1 (3), the linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was prepared, and the DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). It was then stored for later use or at -20°C for long-term preservation.
[0200] (4) Preparation of plasmids Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold and Hsp8Cas9-PSK-mU6-A4sgRNA-scaffold Two target sites, A3 and A4, were selected, and gRNA sequences were designed. Sticky end sequences (indicated by uppercase letters) corresponding to the linearization plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold were added to the sense and antisense strands, respectively. Oligonucleotide single-stranded DNA was then synthesized, and its specific sequence is shown below: A3-gRNA: cggggaaggaggatgccccagg Oligo-F:TTTGcggggaaggaggatgccccaggGT Oligo-R: TAAAACcctggggcatcctccttccccg A4-gRNA:ggccatcctaagaaacgagaga Oligo-F:TTTGggccatcctaagaaacgagagaGT Oligo-R: TAAAACtctctcgtttcttaggatggcc Then, using a method similar to that in Example 2 (4), plasmids Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold and Hsp8Cas9-PSK-mU6-A4sgRNA-scaffold containing the above-mentioned sgRNA sequences were prepared for later use.
[0201] (5) pAAV2_Hsp8Cas9_ITR, three single mutant pAAV2_v1Hsp8Cas9_ITR plasmids and four combined mutant pAAV2_v2Hsp8Cas9_ITR plasmids were combined with plasmids Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold and Hsp8Cas9-PSK-mU6-A4sgRNA-scaffold to transfect HEK293T cell lines containing target DNA. Using a method similar to that in Example 2 (5), the above-mentioned pAAV2_Hsp8Cas9_ITR, three single mutant pAAV2_v1Hsp8Cas9_ITR plasmids, and four combined mutant pAAV2_v2Hsp8Cas9_ITR plasmids were transfected into the HEK293T cell line in combination with plasmids Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold and Hsp8Cas9-PSK-mU6-A4sgRNA-scaffold, respectively. The above transfection process includes the following steps: ① On day 0, HEK293T cells containing the target DNA were seeded in 48-well plates as needed for transfection, with a cell density of about 30%.
[0202] ②On day 1, transfection was performed. The transfection process is as follows: Take 350 ng of each of the pAAV2_Hsp8Cas9_ITR, 3 single mutant pAAV2_v1Hsp8Cas9_ITR plasmids, and 4 combined mutant pAAV2_v2Hsp8Cas9_ITR plasmids, and add them together with 150 ng of Hsp8Cas9-PSK-mU6-A3sgRNA-scaffold plasmid to 12.5 μL of Opti-MEM medium (purchased from Gibco). Gently pipette and mix well to obtain dilution A. Similarly, 350 ng each of the pAAV2_Hsp8Cas9_ITR plasmid, the three single mutant pAAV2_v1Hsp8Cas9_ITR plasmid, and the four combined mutant pAAV2_v2Hsp8Cas9_ITR plasmid were added to 12.5 μL of Opti-MEM medium (Gibco) and gently pipetted to obtain dilution A. Lipofectamine® 2000 transfection reagent (Invitrogen) or polyethyleneimine (PEI) (Polysciences) was gently mixed, and 0.5 μL of Lipofectamine® 2000 or PEI was added to 12.5 μL of Opti-MEM medium (Gibco), gently mixed, and incubated at room temperature for 5 min to obtain dilution B.
[0203] Mix the dilution A and dilution B obtained above, gently pipette to mix, and obtain a mixture C containing the transfection reagent and the plasmid to be transfected. Let it stand at room temperature for 20 min, and then add the mixture C to the HEK293T cell culture medium containing the target DNA in ①. Then place it in a 37℃, 5% CO2 incubator and continue to culture for 3 days.
[0204] (6) Preparation of next-generation sequencing libraries ① Collect HEK293T cells containing target DNA that have been successfully transfected in step (5), and extract genomic DNA from HEK293T cells containing target DNA using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) and following the steps provided by the DNA kit.
[0205] ② The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers targeting the A3 site are shown below: F1 primer: ACACTCTTTCCCTACACGACCGCTTCCGATCTNNNNgactccaccaacgccgac R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNtggtcagcacagagtggcta The PCR primers targeting A4 are shown below: F1 primer: ACACTCTTTCCCTACACGACCGCTTCCGATCTNNNNcatctctcctccctcaccca R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNggagacggggtactttgg The PCR reaction system is as follows:
[0206] The PCR procedure is as follows:
[0207] ③ Perform the second round of PCR for library construction using a 2×Q5 Mastermix. The first round product is used as a template. The forward primer (P5-NGS-F) sequence is: “AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGAC”; the reverse primer (P7-NGS-R) sequence is: “CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTG”. Here, “N” represents an 8-bit index (Illumina Novaseq V1.5).
[0208] The PCR reaction system is as follows:
[0209] The PCR procedure is as follows:
[0210] ④ Use a gel extraction kit to purify the DNA fragment of about 400 bp from the second round of PCR products, following the steps provided by the manufacturer. This completes the preparation of the next-generation sequencing library.
[0211] (7) Analysis of second-generation sequencing results The second-generation sequencing library obtained above was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0212] The editing efficiency of Hsp8Cas9 at the same endogenous site, as calculated by next-generation sequencing, and that of modified single mutants and combined mutants of Hsp8Cas9 are as follows: Figure 3 As shown, the X-axis represents the target site, and the Y-axis represents the editing efficiency (Indels%). From Figure 3 As can be seen, compared with the unmodified Hsp8Cas9 and the single-point mutant, the editing efficiency of the combined mutant is further improved. Among them, the V941R&Y942K combination is slightly better than the other combinations. We finally chose the V941R&Y942K combined mutant (referred to as enHsp8Cas9) for further testing.
[0213] Example 4 (1) Construction of plasmid pAAV2_enHsp8Cas9_ITR Genes of Hsp8Cas9 protein overexpression (as shown in SEQ ID NO:3) and enHsp8Cas9 were synthesized and constructed into the pSpCas9(BB)-2A-Puro (PX459) backbone plasmid V2.0 (Addgene platform, catalog#62988) to obtain plasmids pHsp8Cas9-2A-Puro and penHsp8Cas9-2A-Puro.
[0214] (2) Construction of plasmids pHsp8Cas9-sgRNA-2A-Puro and penHsp8Cas9-sgRNA-2A-Puro vectors The pHsp8Cas9-2A-Puro and penHsp8Cas9-2A-Puro plasmids expressing wild-type and mutant Hsp8Cas9 proteins, respectively, were linearized using PCR and recombinant adapters were added. The PCR primer sequences used were: pCMV-F: CCTTTAGACTGACACCTCttttttgttttagagctagaaatagc pCMV-R: CCCTTTTCAGGGACTAAAACaggtcttctcgaagacccgg The PCR reaction system is as follows:
[0215] The PCR procedure is as follows:
[0216] The PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.
[0217] Genes were synthesized based on the DNA sequence of the sgRNA scaffold (SEQ ID NO: 6) and constructed on the linearized pHsp8Cas9-2A-Puro and penHsp8Cas9-2A-Puro backbones to obtain plasmids pHsp8Cas9-sgRNA-2A-Puro and penHsp8Cas9-sgRNA-2A-Puro.
[0218] (3) Preparation of linearized plasmids pHsp8Cas9-sgRNA-2A-Puro, penHsp8Cas9-sgRNA-2A-Puro and pSpCas9(BB)-2A-Puro (PX459) V2.0 ① In the enzyme digestion system, plasmids pHsp8Cas9-sgRNA-2A-Puro, penHsp8Cas9-sgRNA-2A-Puro, and pSpCas9(BB)-2A-Puro (PX459) V2.0 (purchased from Addgene, catalog number: 62988) were digested with BbsI restriction endonuclease to obtain linearized plasmid fragments.
[0219] The temperature was controlled at 37℃ and the time was 1 hour during the above-mentioned enzyme digestion linearization reaction.
[0220] In the above enzyme digestion system, the amount of each substance used in the linearization reaction is calculated based on a 50 μL enzyme digestion system, which contains 1 μg plasmid, 5 μL 10 xCutSmart buffer (purchased from NEB), 1 μL BbsI restriction endonuclease (purchased from NEB), and the remainder water.
[0221] ② The linearized plasmid fragments obtained above were electrophoresed on an agarose gel with a volume percentage of 1%. The electrophoresis process was controlled at a voltage of 120 V and a time of 30 min.
[0222] ③Then DNA fragments of 8129 bp and 9171 bp were excised and recovered using a gel extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water to obtain linearized plasmids pHsp8Cas9-sgRNA-2A-Puro, penHsp8Cas9-sgRNA-2A-Puro, and pSpCas9(BB)-2A-Puro (PX459) V2.0.
[0223] ④ The DNA concentration of the recovered linearized plasmids pHsp8Cas9-sgRNA-2A-Puro, penHsp8Cas9-sgRNA-2A-Puro, and pSpCas9(BB)-2A-Puro (PX459) V2.0 was determined using a NanoDrop™ Lite spectrophotometer (ThermoScientific) for later use or for long-term storage at -20℃.
[0224] (4) Preparation of plasmids pHsp8Cas9-target-sgRNA-2A-Puro, penHsp8Cas9-target-sgRNA-2A-Puro and pSpCas9(BB)-target-sgRNA-2A-Puro (PX459) V2.0 targeting different endogenous sites Each gRNA was designed, and its sequence is shown in Table 2. Sticky end sequences (indicated by uppercase letters) corresponding to the linearized plasmids pHsp8Cas9-sgRNA-2A-Puro, penHsp8Cas9-sgRNA-2A-Puro, and pSpCas9(BB)-2A-Puro(PX459) V2.0 were added to the sense and antisense strands, respectively. These oligonucleotide single-stranded DNAs were then synthesized, and their specific sequences are also shown in Table 2 below.
[0225] Table 2. Sequences of gRNA and oligonucleotide single-stranded DNA
[0226] Then, using a method similar to that in Example 1 (4), plasmids pHsp8Cas9-target-sgRNA-2A-Puro, penHsp8Cas9-target-sgRNA-2A-Puro, and pSpCas9(BB)-target-sgRNA-2A-Puro (PX459) V2.0 containing the target sgRNA sequence targeting different endogenous sites were prepared for later use.
[0227] (5) Transfection of HEK293T cell line with plasmids pHsp8Cas9-target-sgRNA-2A-Puro, penHsp8Cas9-target-sgRNA-2A-Puro and pSpCas9(BB)-target-sgRNA-2A-Puro (PX459) V2.0 targeting different endogenous sites ①On day 0, HEK293T cells containing the target sequence were seeded in 6-well plates as needed for transfection, with a cell density of about 30%.
[0228] ②On day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid to be transfected, pHsp8Cas9-target-sgRNA-2A-Puro, penHsp8Cas9-target-sgRNA-2A-Puro, or pSpCas9(BB)-target-sgRNA-2A-Puro (PX459) V2.0, and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.
[0229] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI, purchased from Polysciences). Add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0230] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 5 days.
[0231] (6) Preparation of next-generation sequencing libraries ① Collect HEK293T cells edited for 5 days in step (5), and extract the cell genome using QuickExtract DNA Extraction Solution according to the instructions. The extracted genome is then used for library construction through two rounds of PCR.
[0232] ② The first round of PCR for library construction was performed using a 2×Q5 Mastermix. The PCR primers are shown in Table 3. Table 3. List of primers for the first round of PCR in next-generation sequencing
[0233] The PCR reaction system is as follows:
[0234] The PCR procedure is as follows:
[0235] ③ For the second round of PCR library construction, use 2×Q5 Mastermix for the PCR reaction. The first round product is used as a template for the second round of PCR. The forward primer (P5-NGS-F) sequence is: “AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGAC”; the reverse primer (P7-NGS-R) sequence is: “CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTG”. Here, “N” represents an 8-bit index (Illumina Novaseq V1.5). The reaction system is shown below:
[0236] The PCR procedure is as follows:
[0237] ④ Use a gel extraction kit to purify the DNA fragments from the second round of PCR products, following the manufacturer's instructions. This completes the preparation of the next-generation sequencing library.
[0238] (7) Analysis of second-generation sequencing results The second-generation sequencing library obtained above was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0239] The editing efficiency of Hsp8Cas9, enHsp8Cas9, and SpCas9 on target sites obtained from next-generation sequencing calculations is as follows: Figure 4As shown, the X-axis represents the target site, and the Y-axis represents the editing efficiency (Indels%). From Figure 4 As can be seen, the editing efficiency of enHsp8Cas9 reached over 75% at nine different endogenous sites, significantly higher than that of wild-type Hsp8Cas9, and also higher than the most widely used SpCas9 at most endogenous sites. This result fully demonstrates that the CRISPR / enHsp8Cas9 gene editing system constructed in this invention can be effectively used for cellular gene editing. Compared with existing CRISPR / Cas9 gene editing systems, enHsp8Cas9, modified from Hsp8Cas9, significantly improves editing performance while maintaining a small molecular weight for easy delivery, demonstrating its great potential in gene therapy, especially in in vivo applications.
[0240] Example 5 (1) Construction of plasmids pAAV2_Hsp8Cas9_ITR and pAAV2_enHsp8Cas9_ITR Genes of Hsp8Cas9 protein overexpression (as shown in SEQ ID NO:3) and enHsp8Cas9 (as shown in SEQ ID NO:4) were synthesized and constructed into the pAAV2_ITR backbone plasmid to obtain plasmids pAAV2_Hsp8Cas9_ITR and pAAV2_enHsp8Cas9_ITR.
[0241] (2) Construction of plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using the same method as in Example 1 (2), the coding sequence (SEQ ID NO: 6) of the sgRNA scaffold sequence (SEQ ID NO: 5) was synthesized and constructed on a linearized pSKB backbone to obtain the plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold.
[0242] (3) Preparation of linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold Using the same method as in Example 1 (3), the linearized plasmid Hsp8Cas9-PSK-mU6-sgRNA-scaffold was prepared, and the DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). It was then stored for later use or at -20°C for long-term preservation.
[0243] (4) Preparation of plasmids Hsp8Cas9-PSK-mU6-On target-sgRNA-scaffold and Hsp8Cas9-PSK-mU6-Mismatch-sgRNA-scaffold The oligonucleotide single-stranded DNA corresponding to the on-target sgRNA and the mismatch sgRNA was designed and synthesized. The mismatch bases are shown as bold underlined bases in the sequence listing. The specific sequences are shown in Table 4.
[0244] Table 4. Oligonucleotide single-stranded DNA corresponding to on target gRNA and mismatch gRNA
[0245] Then, using a method similar to that in Example 2 (4), the above-mentioned Hsp8Cas9-PSK-mU6-Ontarget-sgRNA-scaffold and 21 Hsp8Cas9-PSK-mU6-Mismatch-sgRNA-scaffold plasmids were prepared for later use.
[0246] (5) The plasmid Hsp8Cas9-PSK-mU6-On target-sgRNA-scaffold expressing the on targets gRNA sequence and 21 plasmids Hsp8Cas9-PSK-mU6-Mismatch-sgRNA-scaffold expressing different mismatch gRNA sequences were used in combination with plasmids expressing Hsp8Cas9 and enHsp8Cas9 proteins to transfect HEK293T cell lines. After a gene editing system cuts the target DNA, the editing capability and specificity of the gene editing system can be assessed by analyzing the indel ratio through next-generation sequencing.
[0247] The above transfection process includes the following steps: ① On day 0, HEK293T cells were seeded in 24-well plates as needed for transfection, with the cell density controlled at 30%.
[0248] ②On day 1, transfection was performed. The transfection process is as follows: Take 500 ng of the pAAV2_Hsp8Cas9_ITR or pAAV2_Hsp8Cas9_ITR plasmid expressing the protein and add 300 ng of the Hsp8Cas9-PSK-mU6-On target-sgRNA-scaffold or Hsp8Cas9-PSK-mU6-Mismatch-sgRNA-scaffold plasmid expressing sgRNA to 25 μL of Opti-MEM medium (purchased from Gibco). Mix gently by pipetting. This is called dilution A.
[0249] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 0.8 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min to obtain dilution B.
[0250] Mix the dilution A and dilution B obtained above, gently blow them to mix, let the mixture stand at room temperature for 20 minutes, and then add the mixture C to the culture medium of the HEK293T cell line in ①.
[0251] ③ The HEK293T cell line obtained in ② after adding mixture C was placed in a 37℃, 5% CO2 incubator for further culture.
[0252] (6) Preparation of next-generation sequencing libraries ① Collect HEK293T cells containing target DNA that have been successfully transfected in step (5), and extract the cell genome using QuickExtract DNA Extraction Solution according to the instructions. The extracted genome is then used for library construction through two rounds of PCR.
[0253] ② For the first round of PCR library construction, use a 2×Q5 Mastermix for the PCR reaction. The PCR primers are shown below. Test-A2A3-F: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNgactccaccaacgccgac Test-A2A3-R:ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNtggtcagcacagagtggcta The PCR reaction system is as follows:
[0254] The PCR procedure is as follows:
[0255] ③ For the second round of PCR library construction, use 2×Q5 Mastermix for the PCR reaction. The first round product is used as a template for the second round of PCR. The forward primer (P5-NGS-F) sequence is: “AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGAC”; the reverse primer (P7-NGS-R) sequence is: “CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTG”. Here, “N” represents an 8-bit index (Illumina Novaseq V1.5). The reaction system is shown below:
[0256] The PCR procedure is as follows:
[0257] ④ Use a gel extraction kit to purify the DNA fragments from the second round of PCR products, following the manufacturer's instructions. This completes the preparation of the next-generation sequencing library.
[0258] (7) Analysis of second-generation sequencing results The second-generation sequencing library obtained above was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0259] The editing efficiency of target sites obtained by second-generation sequencing calculations is as follows: Figure 5 As shown, the Y-axis represents editing efficiency (Indels%), the X-axis represents the oligonucleotide single-stranded DNA sequences corresponding to the on-target sgRNA and mismatch sgRNA, ON represents the oligonucleotide single-stranded DNA sequence corresponding to the on-target sgRNA, M1M2-M21M22 are the oligonucleotide single-stranded DNA sequences corresponding to the mismatch sgRNA, and NC represents the transfection group without sgRNA. Figure 5 As can be seen, the CRISPR / Hsp8Cas9 gene editing system of this invention exhibits low detectable off-target activity when double-base mismatches occur in the sgRNA sequence, especially when mismatches occur in the sgRNA sequence near the PAM end, where off-target effects are almost invisible, demonstrating high specificity. For the modified enHsp8Cas9, when mismatches occur in the sgRNA sequence far from the PAM end, activity comparable to that on the target sequence is detected, with no obvious off-target editing observed at some locations, maintaining high specificity.
[0260] Example 6 (1) Construction of plasmid pCMV-enHsp8Cas9-PEmax-P2A-GFP The gene sequence of the mutant enHsp8Cas9-H585A, based on the modified enHsp8Cas9, was synthesized and constructed onto the pCMV-PEmax-P2A-GFP (Addgene platform, commercially available, catalog #180020) backbone. That is, the sequence of SpCas9-H840A was replaced by the mutant enHsp8Cas9-H585A gene sequence to obtain pCMV-enHsp8Cas9-PEmax-P2A-GFP. The gene sequences of SpCas9-NG-H840A (L1111R, G1218R, E1219F, A1322R, R1335V, T1337R) were synthesized and constructed onto the pCMV-PEmax-P2A-GFP (Addgene platform, commercially available, catalog #180020) backbone to obtain pCMV-PEmax-NG-P2A-GFP.
[0261] (2) Construction of Hsp8Cas9-PSK-mU6-sgRNA-scaffold plasmid Using a method similar to that in Example 1 (2), plasmids Hsp8Cas9-PSK-mU6-sgRNA-scaffold and SpCas9-PSK-mU6-sgRNA-scaffold were prepared using the sgRNA scaffold sequence (SEQ ID NO:5) and the sgRNA scaffold sequence of SpCas9 (detailed sequence reference: pSpCas9(BB)-2A-Puro (PX459) V2.0 (Addgene platform, catalog#62988)).
[0262] (3) Preparation of expression pegRNA plasmids for E4, PR, and A14 The PSK-mU6 backbone and corresponding sgRNA-scaffold in (2) were amplified by PCR, and the sgRNA target sequence and repair template were introduced by primers. The 3' extension of the PegRNA contained the PBS sequence and the RT template sequence. The information sequences contained on the PegRNA for each target site are shown in Table 5 below, so as to complete the precise editing of different types of targets (insertion, mutation, deletion, etc.).
[0263] Table 5. RTT and PBS sequences of PegRNA and specific editing types
[0264] The primer sequences for backbone amplification are as follows, and the amplified fragment is 4597 bp: MS2-F: acatgaggatcacccatgtTTTTTTACTAGTCTCGAGGGGGG pegmU6-R:AAGGCTTTTCTCCAAGGGAT; The sgRNA-scaffold sequence was amplified using Hsp8Cas9-PSK-mU6-sgRNA-scaffold and SpCas9-PSK-mU6-sgRNA-scaffold plasmids as templates, respectively. The primer sequences are shown in Table 6 below. Table 6. Primer sequences for pegRNA plasmid construction
[0265] The PCR reaction system is as follows:
[0266] The PCR procedure is as follows:
[0267] The PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel recovery kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.
[0268] The PSK-mU6 backbone fragment was homologously recombinated with the enHsp8Cas9-mU6-sgRNA-scaffold fragment and the SpCas9-mU6-sgRNA-scaffold fragment respectively in the proportions required by the manufacturer's instructions. The homologous recombinase used was NEBuilder® High Fidelity DNA Assembly Premix (NEB). The reaction system is as follows:
[0269] The reaction conditions are as follows:
[0270] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0271] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance, incubated upside down in a 37°C incubator, and the obtained Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0272] The correctly ligated E. coli DH5α clones were subjected to culture in a shaking process, and then plasmids were extracted to obtain expression pegRNA plasmids targeting E4, PR, and A14, namely enHsp8Cas9-E4-pegRNA, enHsp8Cas9-PR-pegRNA, enHsp8Cas9-A14-pegRNA, and SpCas9-E4-pegRNA, SpCas9-PR-pegRNA, SpCas9-A14-pegRNA.
[0273] (4) Preparation of nick sgRNA plasmids for the expression of E4, PR, and A14 ① Linearized plasmids Hsp8Cas9-PSK-mU6-sgRNA-scaffold and SpCas9-PSK-mU6-sgRNA-scaffold were prepared using a method similar to that in Example 1 (3), and the DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). They were then used for future reference or stored at -20°C for long-term preservation.
[0274] ②Design and synthesize oligonucleotide single-stranded DNA corresponding to nick sgRNA, the specific sequence of which is shown in Table 7.
[0275] Table 7. Oligonucleotide single-stranded DNA corresponding to nick sgRNA
[0276] ③ The oligonucleotide single-stranded DNA corresponding to the nick sgRNA obtained in step ② was annealed. After annealing, the resulting products were ligated to the linearized Hsp8Cas9-PSK-mU6-sgRNA-scaffold and SpCas9-PSK-mU6-sgRNA-scaffold plasmids obtained in step ① using DNA ligase (purchased from NEB) to obtain the ligation products of the corresponding nick sgRNAs.
[0277] The annealing reaction described above included: mixing 1 μL of 100 μM oligo-F, 1 μL of 100 μM oligo-R, and 28 μL of water by shaking, then placing the mixture in a PCR instrument and running the annealing program. The annealing program was as follows: 95℃ for 5 min, 85℃ for 1 min, 75℃ for 1 min, 65℃ for 1 min, 55℃ for 1 min, 45℃ for 1 min, 35℃ for 1 min, 25℃ for 1 min, and storing at 4℃ with a cooling rate of 0.3℃ / s.
[0278] The above ligation process was performed according to the instructions for DNA ligase (purchased from NEB).
[0279] ④ Take 1 μL of the ligation product of the corresponding nick sgRNA obtained in step ③ and add it to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.). Incubate on ice for 30 min, heat shock at 42℃ for 1 min, incubate on ice for 2 min, add 900 μL of LB medium, and incubate at 37℃ for 1 h to activate and revive Escherichia coli DH5α competent cells.
[0280] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The obtained Escherichia coli DH5α monoclonal cells were then verified by Sanger sequencing.
[0281] ⑤ Perform culture on the correctly ligated E. coli DH5α clones verified by sequencing, and then extract plasmids to obtain expression nick sgRNA plasmids targeting E4, PR, and A14, namely enHsp8Cas9-E4-nicksgRNA, enHsp8Cas9-PR-nicksgRNA, enHsp8Cas9-A14-nicksgRNA, and SpCas9-E4-nicksgRNA, SpCas9-PR-nicksgRNA, SpCas9-A14-nicksgRNA, respectively, for later use.
[0282] (5) PE system plasmids targeting the endogenous sites of E4, PR, and A14, expressing enHsp8Cas9 and SpCas9, were transfected into HEK293T cell lines. The pCMV-enHsp8Cas9-PEmax-P2A-GFP and the corresponding pegRNA and nick sgRNA expression plasmids for E4, PR, and A14 obtained in step (1) were co-transfected into the HEK293T cell line using liposomes. The pCMV-PEmax-NG-P2A-GFP plasmid (available commercially on the Addgene platform, catalog #180020) and the corresponding pegRNA and nick sgRNA expression plasmids for E4, PR, and A14 expression plasmids for SpCas9 were co-transfected into the HEK293T cell line using liposomes.
[0283] After the gene editing system cuts the target DNA and repairs it according to the reverse transcription template, the editing capability of the PE gene editing system can be evaluated by analyzing the correct editing ratio through next-generation sequencing.
[0284] The above transfection process includes the following steps: ① On day 0, HEK293T cells were seeded in 24-well plates as needed for transfection, with the cell density controlled at 30%.
[0285] ②On day 1, transfection was performed. The transfection process is as follows: For transfection of the enHsp8Cas9 PE system, 450 ng of pCMV-enHsp8Cas9-PEmax-P2A-GFP, 250 ng of pegRNA, and 100 ng of nick sgRNA plasmid were added to 25 μL of Opti-MEM medium (purchased from Gibco). Similarly, for transfection of the SpCas9 PE system, 450 ng of pCMV-PEmax-NG-P2A-GFP, 250 ng of pegRNA, and 100 ng of nick sgRNA plasmid were added to 25 μL of Opti-MEM medium (purchased from Gibco), and the mixture was gently pipetted to mix. This was recorded as dilution A.
[0286] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 0.8 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min to obtain dilution B.
[0287] Mix the dilution A and dilution B obtained above, gently blow them to mix, let the mixture stand at room temperature for 20 minutes, and then add the mixture C to the culture medium of the HEK293T cell line in ①.
[0288] ③ The HEK293T cell line obtained in ② after adding mixture C was placed in a 37℃, 5% CO2 incubator for further culture.
[0289] (6) Preparation of next-generation sequencing libraries ① Collect HEK293T cells containing target DNA that have been successfully transfected in step (5), and extract the cell genome using QuickExtract DNA Extraction Solution according to the instructions. The extracted genome is then used for library construction through two rounds of PCR.
[0290] ② For the first round of PCR library construction, use 2×Q5 Mastermix for PCR reaction. The PCR primers are shown in Table 8 below.
[0291] Table 8. First-round primer sequences for PE editing site detection
[0292] The PCR reaction system is as follows:
[0293] The PCR procedure is as follows:
[0294] ③ For the second round of PCR library construction, use 2×Q5 Mastermix for the PCR reaction. The first round product is used as a template for the second round of PCR. The forward primer (P5-NGS-F) sequence is: “AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGAC”; the reverse primer (P7-NGS-R) sequence is: “CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTG”. Here, “N” represents an 8-bit index (Illumina Novaseq V1.5). The reaction system is shown below:
[0295] The PCR procedure is as follows:
[0296] ④ Use a gel extraction kit to purify the DNA fragments from the second round of PCR products, following the manufacturer's instructions. This completes the preparation of the next-generation sequencing library.
[0297] (7) Analysis of second-generation sequencing results The second-generation sequencing library obtained above was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).
[0298] The editing efficiency of target sites obtained by second-generation sequencing calculations is as follows: Figure 6 As shown, the Y-axis represents the percentage of sequence reads with specified edits, and the X-axis represents different target sites in the cell genome. From Figure 6 As can be seen, the guide editor derived from the CRISPR / enHsp8Cas9 gene editing system of this invention can precisely edit target sites in the HEK293T cell line, and its editing efficiency at certain sites is significantly higher than that of the SpCas9-NG guide editor, indicating that this tool has broad application prospects in gene therapy.
[0299] The above specific embodiments are merely detailed explanations of the technical solutions of this application. This application is not limited to the above embodiments. Those skilled in the art should understand that any improvements or substitutions based on the above principles and spirit should be within the protection scope of this application.
[0300] The sequences used in this invention: SEQ ID NO: 1: Amino acid sequence of Hsp8Cas9 protein SEQ ID NO:2: Amino acid sequence of enHsp8Cas9 protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNKNIKQDQDKKDKEEGKILSALSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEIQRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENEYRAPKYSLSAVEFVTISKVINILASCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSLKALGEILPFMRNGIRYDQACKNANLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDLPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYIARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKKFYAVPIYTMDIALGVLPNKAVVGGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYF RK FGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:3: Coding sequence of Hsp8Cas9 protein SEQ ID NO:4: Coding sequence of the enHsp8Cas9 protein CGGAAA TTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATC CAGCGGTTAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO: 5: Single-stranded guide RNA scaffold sequence GUUUUAGUCCCUGAAAAGGGACUAAAAUAAGCUUGUCAGUCGGCGCAGGGUGAGGCUCUAUGAGUGCUCGAAAAUCCCUUUAGACUGACACCUC SEQ ID NO: 6: DNA sequence of a single-stranded guide RNA scaffold. GTTTTAGTCCCTGAAAAGGGACTAAAATAAGCTTGTCAGTCGGCGCAGGGTGAGGCTCTATGAGTGCTCGAAAATCCCTTTAGACTGACACCTC SEQ ID NO:7: Amino acid sequence of Hsp8Cas9-V895R protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNNKNIKQDQDKKDKEEGKILSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEI QRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENYRAPKYSLSAVEFVTISVINILASSCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSKALGEILPFMRNGIRYDQACKNA NLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYI ARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKFYAVPIYTMDIALGVLPNKAV RGGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYFVYFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:8: Coding sequence of Hsp8Cas9-V895R protein CGG GGGGGAAAGTCCAATGGCGTGATCAAAGACTGGATCGAGATGGACGAGAACTACGAGTTCTGTTTCAGCCTGTTTAAAGGAGATGTGGTTCTGGTGCAGAAATCTAGTATGGAAAAACCTGAGTTCGCGTACTTCGTGTACTTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO:9: Amino acid sequence of Hsp8Cas9-V941R protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNNKNIKQDQDKKDKEEGKILSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEIQRGFGFAMTQEL CDEVLHIAFYQRPLKDFSHLVGKCTFYENEYRAPKYSLSAVEFVTISVINILASCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSKALGEILPFMRNGIRYDQNANLVTRHNDKKMKFLPPNESMYA QDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDLPNLENNILKLRLFREQDEVCAYSGKKITLDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYIARLIASYTREYLLACPLDCNEDTMLVAGEKGSKIH VETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKFYAVPIYTMDIALGVLPNKAVVGGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEAYF RYFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:10: Coding sequence of Hsp8Cas9-V941R protein CGG TACTTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO:11: Amino acid sequence of Hsp8Cas9-Y942K protein SEQ ID NO:12: Coding sequence of Hsp8Cas9-V942K protein AAA TTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO:13: Amino acid sequence of Hsp8Cas9-V895R&Y942K protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNNKNIKQDQDKKDKEEGKILSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEI QRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENYRAPKYSLSAVEFVTISVINILASSCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSKALGEILPFMRNGIRYDQACKNA NLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYI ARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKFYAVPIYTMDIALGVLPNKAV R GGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYFV KFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:14: Coding sequence of the Hsp8Cas9-V895R&Y942K protein CGG GGGGGAAAGTCCAATGGCGTGATCAAAGACTGGATCGAGATGGACGAGAACTACGAGTTCTGTTTCAGCCTGTTTAAAGGAGATGTGGTTCTGGTGCAGAAATCTAGTATGGAAAAACCTGAGTTCGCGTACTTCGTG AAA TTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO:15: Amino acid sequence of Hsp8Cas9-V895R&V941R protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNNKNIKQDQDKKDKEEGKILSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEI QRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENYRAPKYSLSAVEFVTISVINILASSCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSKALGEILPFMRNGIRYDQACKNA NLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYI ARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKFYAVPIYTMDIALGVLPNKAV R GGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYF RYFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:16: Coding sequence of Hsp8Cas9-V895R&V941R protein CGG GGGGGAAAGTCCAATGGCGTGATCAAAGACTGGATCGAGATGGACGAGAACTACGAGTTCTGTTTCAGCCTGTTTAAAGGAGATGTGGTTCTGGTGCAGAAATCTAGTATGGAAAAACCTGAGTTCGCGTACTTC CGG TACTTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO:17: Amino acid sequence of Hsp8Cas9-V895R&Y942K&V941R protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNNKNIKQDQDKKDKEEGKILSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEI QRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENYRAPKYSLSAVEFVTISVINILASSCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSKALGEILPFMRNGIRYDQACKNA NLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEIDHILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYI ARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKFYAVPIYTMDIALGVLPNKAV R GGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYF RKFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:18: Coding sequence of Hsp8Cas9-V895R&Y942K&V941R protein CGG GGGGGAAAGTCCAATGGCGTGATCAAAGACTGGATCGAGATGGACGAGAACTACGAGTTCTGTTTCAGCCTGTTTAAAGGAGATGTGGTTCTGGTGCAGAAATCTAGTATGGAAAAACCTGAGTTCGCGTACTTC CGGAAA TTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC SEQ ID NO: 19: Amino acid sequence of Hsp8Cas9-H585A protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNKNIKQDQDKKDKEEGKILSALSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEIQRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENEYRAPKYSLSAVEFVTISKVINILASCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSLKALGEILPFMRNGIRYDQACKNANLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDLPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEID AILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYIARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKKFYAVPIYTMDIALGVLPNKAVVGGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYFVYFGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO:20: Coding sequence of Hsp8Cas9-H585A protein GCC SEQ ID NO:21: Amino acid sequence of the Hsp8Cas9-V941R&Y942K&H585A protein MKIVGIDIGITSIGWAFVENGELKDCGVRIFTGAENPKNGESLALPRRQARSVRRRLARRRGRLESLKKLLTNAWHLSYEDYIATDGELPKAFYGKRFLNPYQLRYEALERLLSKEELMRVILHIAKHRGYGNKNIKQDQDKKDKEEGKILSALSENAKKISHYRTAGEYFYKEFCAVQDVGNAQDKTPMIRVLKPIRNKALSYANCVSQVALQNELSLIFEIQRGFGFAMTQELCDEVLHIAFYQRPLKDFSHLVGKCTFYENEYRAPKYSLSAVEFVTISKVINILASCSKESGEIYTQEQYQKILSSVFEEVCNKGSLSFAQLRKMLNIDESIQFQELKHSNLEKPETKKFVEFKNFKKFCHILGDVKQDRETSNNIARDLTLTKDKEQLKQKLCKYPALTQDQIMQLSEIDFDHHISLSLKALGEILPFMRNGIRYDQACKNANLVTRHNDKKMKFLPPFNESMYAQDLHNPVVLRAISEYRKVLNALIKKYGRFHKIHIELAREVGKSYKDRGRYQKEIDTNYKNREQAKEMCKKIDLPLNENNILKLRLFREQDEVCAYSGKKITLDDLKNAGALEID AILPYSRSFDNGYLNKVLVFTKENQNKGNKTPFEAFGADCIKWSKIVSLAMRNGYPKKKAQNITTTTFATRESGFKARNLSDTRYIARLIASYTREYLACLPLDCNEDTMLVAGEKGSKIHVETINGMLTTTMRHFWGLPVKNRFEHTHHAIDAIIIAYSNASMIKRFSDFVKNHQETLKAELYAKELAQDTFKHQRKFFEPFAGFREHALHKIAQIFVSHSPNRRVRGALHEETFKSFNDATYQKAYGGLAGIQKALELGKIRQIGTKLVANGEMVRVDIFAHKKSKKFYAVPIYTMDIALGVLPNKAVVGGKSNGVIKDWIEMDENYEFCFSLFKGDVVLVQKSSMEKPEFAYF RK FGVSTASIALQTHDNNIKTLTPNQQKLFTSPTQEKVTAESLGIQRLKVFEKYKVSPLGELKPSKYEPRQPIALKTTPKTQKPNKDS SEQ ID NO: : Encoding sequence of Hsp8Cas9-V941R&Y942K&H585A protein GCC CGGAAA TTCGGAGTGTCCACCGCTTCTATCGCCCTGCAGACCCACGACAACAACATCAAGACACTGACACCAAACCAGCAGAAGCTGTTCACCAGCCCCACCCAAGAGAAGGTGACCGCCGAAAGCCTGGGCATCCAGCGGTTAAAAGTGTTCGAAAAATACAAAGTGTCTCCACTGGGCGAGCTGAAGCCCAGCAAGTACGAGCCTAGACAGCCTATCGCTCTGAAAACCACCCCTAAGACCCAGAAGCCTAACAAGGACAGC
Claims
1. An Hsp8Cas9 mutant protein, wherein the Hsp8Cas9 mutant protein contains one or more mutations at V895, V941, Y942, and H585 relative to the amino acid sequence of the wild-type Hsp8Cas9 protein as shown in SEQ ID NO: 1; Preferably, the mutation includes one or more of V895R, V941R, Y942K, and H585A; More preferably, the mutations include V941R and Y942K; More preferably, the mutations include V941R, Y942K, and H585A.
2. A conjugate, said conjugate comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein as described in claim 1. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; b) Modified parts; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as c) An optional adapter for connecting the Cas9 protein or its homolog to the modified portion.
3. A fusion protein, said fusion protein comprising: a) Cas9 protein or its homologs, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein as described in claim 1. The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; b) Other proteins or peptides; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as c) Optional adapters for connecting the Cas9 protein or its homologs to the other protein or polypeptide.
4. A single-stranded guide RNA, said single-stranded guide RNA comprising a CRISPR scaffold sequence, said CRISPR scaffold sequence having: a) The nucleic acid sequence shown in SEQ ID NO: 5; b) A nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity with the nucleic acid sequence shown in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5; or c) A nucleic acid sequence obtained by modifying the nucleic acid sequence described in SEQ ID NO: 5 and retaining the biological activity of SEQ ID NO: 5; For example, the modification is one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening; For example, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the base sequence.
5. The single-stranded guide RNA according to claim 4, wherein, The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the CRISPR scaffold sequence. The CRISPR spacer sequence is a sequence of 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides (preferably 22 nucleotides) in length and is complementary to the target sequence.
6. An isolated nucleic acid molecule, said isolated nucleic acid molecule encoding the nucleic acid sequence of the Hsp8Cas9 mutant protein of claim 1 or its homolog, the conjugate of claim 2, or the fusion protein of claim 3, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the Hsp8Cas9 mutant protein and retains the biological activity of the Hsp8Cas9 mutant protein. Preferably, the isolated nucleic acid molecule comprises any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO:
22.
7. The isolated nucleic acid molecule according to claim 6, wherein the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding the single-stranded guide RNA according to any one of claims 4 to 5; For example, the isolated nucleic acid molecule contains any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and the isolated nucleic acid molecule also contains the nucleic acid sequence shown in SEQ ID NO:
6.
8. An isolated nucleic acid molecule, said isolated nucleic acid molecule encoding the nucleic acid sequence of the single-stranded guide RNA of any one of claims 4 to 5; For example, the isolated nucleic acid molecule contains the nucleic acid sequence shown in SEQ ID NO: 6, and preferably also contains a nucleic acid sequence encoding a CRISPR spacer sequence.
9. The isolated nucleic acid molecule according to claim 8, wherein the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or a homolog thereof as shown in SEQ ID NO: 1, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein; For example, the isolated nucleic acid molecule contains the nucleic acid sequence shown in SEQ ID NO: 6 and the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
10. A vector comprising a nucleic acid sequence encoding the Hsp8Cas9 mutant protein of claim 1 or a homolog thereof, the conjugate of claim 2, or the fusion protein of claim 3, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the Hsp8Cas9 mutant protein and retains the biological activity of the Hsp8Cas9 mutant protein. Preferably, the vector comprises the nucleic acid sequence or degenerate sequence thereof shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22; Preferably, the vector is a plasmid vector, a retroviral vector, an adenovirus vector, or an adeno-associated virus vector; More preferably, the plasmid vector is the pAAV2_ITR plasmid vector.
11. The carrier according to claim 10, characterized in that, The vector further comprises a nucleic acid sequence encoding the single-stranded guide RNA as described in any one of claims 4 to 5. For example, the vector contains the nucleic acid sequence or degenerate sequence thereof shown in any one of SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and the vector further contains the nucleic acid sequence shown in SEQ ID NO:
6.
12. A vector comprising a nucleic acid sequence encoding a single-stranded guide RNA as described in any one of claims 4 to 5; For example, the vector contains the nucleic acid sequence shown in SEQ ID NO: 6, and preferably also contains a nucleic acid sequence encoding a CRISPR spacer sequence.
13. The vector of claim 12, wherein the vector further comprises a nucleic acid sequence encoding a wild-type Hsp8Cas9 protein or a homolog thereof as shown in SEQ ID NO: 1, wherein: The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein; For example, the vector contains the nucleic acid sequence shown in SEQ ID NO: 6 and the nucleic acid sequence shown in SEQ ID NO: 3 or a degenerate sequence thereof.
14. A CRISPR / Cas9 gene editing system, said CRISPR / Cas9 gene editing system comprising: 1) Protein components, comprising the nucleic acid sequences of the Cas9 protein, its homologs, conjugates, or fusion proteins, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein as described in claim 1; The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; as well as 2) A nucleic acid component comprising: the single-stranded guide RNA as described in any one of claims 4 to 5; Furthermore, the protein component and the nucleic acid component bind together to form a complex; Preferably, the protein component comprises: a wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or an Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19 or SEQ ID NO: 21, and the nucleic acid component comprises a single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO:
5.
15. The CRISPR / Cas9 gene editing system of claim 14, wherein the wild-type Hsp8Cas9 or the Hsp8Cas9 mutant protein and the single-stranded guide RNA are obtained by expression of the vector of claim 11 or 13.
16. A cell comprising: an isolated nucleic acid molecule according to any one of claims 6 to 9, or a vector according to any one of claims 10 to 13; For example, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell being, for example, a plant cell or an animal cell, the animal cell being, for example, a mammalian cell such as a human cell.
17. A method for gene editing in cells or in an in vitro environment, characterized in that, The method includes contacting any one of the following (1) to (4) with a target sequence in an intracellular or in vitro environment: (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA according to any one of claims 4 to 5, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein as described in claim 1; The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; For example, wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19 or SEQ ID NO: 21, and single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5; (2) The vector according to claim 10 or a vector comprising a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the vector according to claim 12. Wherein, the homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein; For example, a vector containing any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20 or SEQ ID NO:22, and a vector containing the nucleic acid sequence shown in SEQ ID NO:6; (3) The carrier according to claim 11 or 13; or (4) The CRISPR / Cas9 gene editing system according to claim 14 or 15; Upon contact with the target sequence, the Cas9 protein or its homologs, conjugates or fusion proteins recognize the PAM sequence NNNNRYA located at the 3' end of the target sequence. Preferably, the gene editing includes gene knockout of target DNA, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, or chromatin imaging tracking. More preferably, the single base conversion includes adenine to guanine, cytosine to thymine, or cytosine to uracil.
18. The method of claim 17, wherein the cell is a eukaryotic cell or a prokaryotic cell; Preferably, the eukaryotic cells include mammalian cells or plant cells; More preferably, the mammalian cells include Chinese hamster ovary cells, young hamster kidney cells, mouse Sertoli cells, mouse mammary tumor cells, buffalo rat hepatocytes, rat hepatoma cells, monkey kidney CVI line transformed by SV40, monkey kidney cells, canine kidney cells, human cervical cancer cells, human lung cells, human hepatocytes, HIH / 3T3 cells, human U2-OS osteosarcoma cells, human A549 cells, human K562 cells, human HEK293 cells, human HEK293T cells, human HCT116 cells, or human MCF-7 cells.
19. The method according to any one of claims 17-18, wherein the CRISPR spacer sequence of the single-stranded guide RNA forms a completely complementary base pairing structure with the target sequence, and forms an incomplete complementary base pairing structure with the non-target sequence; Preferably, the incomplete base complementarity pairing includes one or more, for example, two or more base mismatches.
20. A kit for gene editing of a target sequence in an intracellular or in vitro environment, comprising: i. Choose any one of the following (1) to (6): (1) Cas9 protein, its homologs, conjugates or fusion proteins, and the single-stranded guide RNA according to any one of claims 4 to 5, wherein: The Cas9 protein is the wild-type Hsp8Cas9 protein with the amino acid sequence shown in SEQ ID NO: 1 or the Hsp8Cas9 mutant protein as described in claim 1; The homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the amino acid sequence of the Cas9 protein and retains the biological activity of the Cas9 protein; The conjugate comprises: the Cas9 protein or its homolog, a modified moiety, and an optional linker; The fusion protein comprises: the Cas9 protein or its homologs, additional proteins or peptides, and optional linkers; The optional connector is used to connect the Cas9 protein or its homolog to the modified portion or the other protein or polypeptide; For example, the modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), reverse transcriptase, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI; For example, wild-type Hsp8Cas9 protein with an amino acid sequence as shown in SEQ ID NO: 1, or Hsp8Cas9 mutant protein with an amino acid sequence as shown in any one of SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19, or SEQ ID NO: 21, and single-stranded guide RNA containing the CRISPR scaffold sequence shown in SEQ ID NO: 5; (2) The isolated nucleic acid molecule according to claim 6, or the isolated nucleic acid molecule comprising a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the isolated nucleic acid molecule according to claim 8. Wherein, the homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein; For example, isolated nucleic acid molecules containing any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 1, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and isolated nucleic acid molecules containing the nucleic acid sequence shown in SEQ ID NO: 6; (3) The isolated nucleic acid molecule according to claim 7 or 9; (4) The vector according to claim 10 or a vector comprising a nucleic acid sequence encoding the wild-type Hsp8Cas9 protein or its homolog as shown in SEQ ID NO: 1, and the vector according to claim 12. Wherein, the homolog is a homolog whose amino acid sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, or at least 99.999% sequence identity with the wild-type Hsp8Cas9 protein and retains the biological activity of the wild-type Hsp8Cas9 protein; For example, a vector containing any one of the nucleic acid sequences or degenerate sequences shown in SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20 or SEQ ID NO: 22, and a vector containing the nucleic acid sequence shown in SEQ ID NO: 6; (5) The carrier according to claim 11 or 13; or (6) The CRISPR / Cas9 gene editing system according to claim 14 or 15; as well as ii. Instructions on how to perform gene editing on target sequences in intracellular or in vitro environments.