Cas9 Protein, Gene Editing System Containing Cas9 Protein and Application
By modifying Cas9 protein to improve its binding ability to sgRNA and target DNA cleavage specificity, the problem of high off-target rate of CRISPR/Cas9 system in gene therapy is solved, and efficient gene editing effect is achieved, which is suitable for the development of gene therapy tools.
Patent Information
- Application Number
- CN202211134139.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-09-16
AI Technical Summary
The existing CRISPR/Cas9 gene editing system is difficult to accurately distinguish target sites from highly similar sequences in gene therapy, resulting in a high off-target rate, limiting its widespread clinical application.
The highly specific Cas9 proteins SauriCas9-HF, Sha2Cas9-HF and Sa-SlugCas9-HF were developed. By modifying their amino acid sequences, the binding ability of Cas9 protein to single-strand guide RNA and the specificity of targeted DNA cleavage are improved, thereby reducing off-target rate.
These modified Cas9 proteins can form complexes with sgRNA for precise gene editing, with an off-target rate of nearly 0%. They are suitable as gene therapy tools and expand the application range of gene editing.
Smart Images

Figure CN116144629B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of gene editing technology, and specifically to Cas9 protein, a gene editing system containing the Cas9 protein, and related applications thereof. Background Art
[0002] The CRISPR / Cas9 system is an acquired immune system evolved by bacteria and archaea to defend against invading viruses and plasmids. The CRISPR / Cas9 system consists of tracrRNA (trans-activating RNA) and crRNA (CRISPR-derived RNA), which together with Cas9 form a complex to function. The tracrRNA and crRNA can be fused to form a single-stranded guide RNA (sgRNA) via a linker sequence. When DNA breaks and damage occur, two major DNA repair mechanisms within the cell are responsible: non-homologous end-joining (NHEJ) and homologous recombination (HR). NHEJ repair results in base deletions or insertions, enabling gene knockout. HR repair, when provided with a homologous template, allows for site-specific insertions and precise base substitutions.
[0003] In addition to basic scientific research, the CRISPR / Cas9 gene editing system also has broad clinical application prospects. When using the CRISPR / Cas9 gene editing system for gene therapy, Cas nuclease-mediated precise editing is required. If the target site and highly similar sequences cannot be accurately distinguished, off-target effects will occur, which is also one of the urgent problems to be solved in gene therapy. One way to solve this problem is to develop Cas9 mutants, improve its editing specificity, and enhance its ability to distinguish between target site sequences and highly similar sequences. The inventors previously invented SauriCas9, Sha2Cas9, and Sa-SlugCas9 proteins, but their specificity is low, and they can easily cause off-target effects, making them difficult to widely use. Therefore, improving the specificity of these three proteins and reducing the off-target rate are the keys to applying these three proteins to gene therapy. Summary of the Invention
[0004] To address the above issues, the inventors conducted repeated studies and obtained three highly specific Cas9 proteins through modification. They can form a CRISPR / Cas9 gene editing system that effectively performs gene editing with the same single-stranded guide RNA, thus completing the present invention.
[0005] Therefore, in a first aspect, the present invention provides a Cas9 protein, wherein the Cas9 protein is:
[0006] A SauriCas9-HF protein having the amino acid sequence of SEQ ID NO: 1, or a homolog thereof having at least 80% sequence identity to the amino acid sequence of SEQ ID NO: 1 and retaining its biological activity;
[0007] A Sha2Cas9-HF protein having the amino acid sequence of SEQ ID NO: 2, or a homolog thereof having at least 80% sequence identity to the amino acid sequence of SEQ ID NO: 2 and retaining its biological activity; or
[0008] A Sa-SlugCas9-HF protein having the amino acid sequence shown in SEQ ID NO: 3, or a homolog of the amino acid sequence having at least 80% sequence identity with the amino acid sequence shown in SEQ ID NO: 3 and retaining its biological activity.
[0009] In a second aspect, the present invention provides a conjugate comprising:
[0010] a) the Cas9 protein described in the first aspect;
[0011] b) modified parts; and
[0012] c) an optional linker for connecting the Cas9 protein to the modification moiety.
[0013] In a third aspect, the present invention provides a fusion protein comprising:
[0014] a) the Cas9 protein described in the first aspect;
[0015] b) additional proteins and peptides; and
[0016] c) Optional linkers for connecting the Cas9 protein to the additional proteins and polypeptides.
[0017] In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding:
[0018] a) the Cas9 protein described in the first aspect;
[0019] b) the conjugate according to the second aspect; or
[0020] c) The fusion protein according to the third aspect.
[0021] In a fifth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding:
[0022] a) the Cas9 protein described in the first aspect;
[0023] b) the conjugate according to the second aspect; or
[0024] c) The fusion protein according to the third aspect.
[0025] In a sixth aspect, the present invention provides a CRISPR / Cas9 gene editing system comprising:
[0026] a) a protein component comprising the Cas9 protein described in the first aspect, the conjugate described in the second aspect; or the fusion protein described in the third aspect;
[0027] b) a nucleic acid component comprising: a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a scaffold sequence, wherein the scaffold sequence has:
[0028] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0029] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0030] (iii) a nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity;
[0031] Furthermore, the protein component and the nucleic acid component are combined with each other to form a complex.
[0032] In a seventh aspect, the present invention provides a cell comprising: the isolated nucleic acid molecule according to the fourth aspect, or the vector according to the fifth aspect.
[0033] In an eighth aspect, the present invention provides a method for gene editing a target sequence in a cell or in vitro environment, the method comprising: contacting any one of the following (1) to (3) with the target sequence in the cell or in vitro environment:
[0034] (1) The Cas9 protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and a single-stranded guide RNA;
[0035] (2) the carrier described in the fifth aspect; and
[0036] (3) The CRISPR / Cas9 gene editing system described in the sixth aspect;
[0037] wherein, upon contact with a target sequence, the Cas9 protein, the conjugate or the fusion protein recognizes the respective protospacer adjacent sequence (PAM), the PAM being located at the 5' end of the target sequence and having the sequence 5'-NNGG;
[0038] Wherein, the single-stranded guide RNA includes a scaffold sequence; the scaffold sequence has:
[0039] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0040] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0041] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0042] In a ninth aspect, the present invention provides a kit for performing gene editing on a target sequence in a cell or in vitro environment, comprising:
[0043] a) Any one of the following 1) to 4):
[0044] 1) The Cas9 protein of the first aspect, the conjugate of the second aspect, or the fusion protein of the third aspect, and a single-stranded guide RNA;
[0045] 2) the isolated nucleic acid molecule of the fourth aspect;
[0046] 3) The vector described in the fifth aspect; or
[0047] 4) the CRISPR / Cas9 gene editing system described in the sixth aspect; and
[0048] b) Instructions for performing gene editing on the target sequence in a cell or in vitro setting;
[0049] Wherein, the single-stranded guide RNA includes a scaffold sequence, and the scaffold sequence has:
[0050] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0051] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0052] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0053] As previously mentioned, the present invention modifies various Cas9 proteins capable of gene editing in eukaryotic cells. These Cas9 proteins possess high specificity and can form complexes with the same sgRNA for gene editing, with an off-target rate close to 0%, making them well-suited for future development as gene therapy tools. This invention thus expands the scope of gene editing and has broad application prospects in the field of gene editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific implementation or the description of the prior art.
[0055] Figure 1 Schematic diagram showing the editing efficiency results after gene editing of four target sites by the CRISPR / SauriCas9-HF gene editing system;
[0056] Figure 2 Schematic diagram showing the editing efficiency results after gene editing of eight target sites by the CRISPR / Sha2Cas9-HF gene editing system;
[0057] Figure 3 Schematic diagram showing the editing efficiency results after gene editing of eight target sites by the CRISPR / Sa-SlugCas9-HF gene editing system;
[0058] Figure 4 Schematic diagram showing the improved specificity of the CRISPR / SauriCas9-HF gene editing system in the GFP reporter system HEK293T cell line;
[0059] Figure 5 Schematic diagram showing the improved specificity of the CRISPR / Sha2Cas9-HF gene editing system in the GFP reporter system HEK293T cell line.
[0060] Figure 6 Schematic diagram showing the improved specificity detection results of the CRISPR / Sa-SlugCas9-HF gene editing system in the GFP reporter system HEK293T cell line. DETAILED DESCRIPTION
[0061] Hereinafter, the present invention will be described in detail with reference to the accompanying drawings. It should be understood that the following description is merely illustrative of the present invention and is not intended to limit the scope of the present invention. The scope of protection of the present invention shall be subject to the appended claims. Furthermore, those skilled in the art will appreciate that the technical solutions of the present invention may be modified without departing from the spirit and purpose of the present invention. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0062] Where a numerical range is provided, such as a concentration range, a percentage range, or a ratio range, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of the range and any other stated or intervening values in the stated range are encompassed within the subject matter unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and such embodiments are also encompassed within the subject matter, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed within the subject matter.
[0063] In the context of the present invention, many embodiments use the expressions "comprising", "including" or "consisting essentially / mainly of..." The expressions "comprising", "including" or "consisting essentially / mainly of..." can generally be understood as open-ended expressions, indicating that in addition to the elements, components, assemblies, method steps, etc. specifically listed after the expression, other elements, components, assemblies, method steps, etc. are also included. In addition, in this document, the expressions "comprising", "including" or "consisting essentially / mainly of..." can also be understood as closed-ended expressions in some cases, indicating that only the elements, components, assemblies, method steps specifically listed after the expression are included, and no other elements, components, assemblies, method steps are included. In this case, the expression is equivalent to the expression "consisting of..."
[0064] For a better understanding of the present teachings and without limiting the scope of the present teachings, all numbers and other numerical values expressing quantities, percentages or ratios used in the specification and claims should be understood as being modified in all cases by the term "about", unless otherwise indicated. Therefore, unless indicated to the contrary, the numerical parameters set forth in the following specification and the appended claims are approximate values that may vary depending on the desired properties sought to be obtained. At the very least, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0065] definition
[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter of the present invention belongs. Before describing the present invention in detail, the following definitions are provided for a better understanding of the present invention.
[0067] The terms "Cas9 protein," "Cas9," and "Cas" are used interchangeably in this application and refer to RNA-guided nucleases, including the Cas9 protein or its functionally active fragments. The Cas9 protein is a protein component of the CRISPR / Cas9 genome editing system that, under the guidance of a single-stranded guide RNA (gRNA), targets and cuts the target DNA sequence, forming a DNA double-strand break (DSB). DNA double-strand breaks can activate the cell's inherent repair mechanisms, non-homologous end joining (NHEJ) and homologous recombination (HR), thereby repairing DNA damage in the cell. During the repair process, the specific DNA sequence is edited at a specific site.
[0068] As used herein, the terms "single-stranded guide RNA," "gRNA," "sgRNA (single guide RNA)" or "mature crRNA" are used interchangeably in this application and have the meanings commonly understood by those skilled in the art. Generally speaking, a single-stranded guide RNA may comprise a scaffold sequence and a guide sequence, which is also referred to herein as a guide RNA (guide RNA or gRNA). In the context of an endogenous CRISPR system, a guide sequence is also referred to as a spacer. In some cases, a guide sequence is any polynucleotide sequence that has sufficient similarity to a target sequence to hybridize to the target sequence and guide specific binding of the CRISPR / Cas9 complex to the target sequence. In certain embodiments, when optimally aligned, the degree of complementarity between a guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of a person of ordinary skill in the art. For example, there are publicly and commercially available alignment algorithms and programs such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.
[0069] As used herein, the term "CRISPR / Cas9 complex" refers to a complex formed by the binding of a single-stranded guide RNA (single guide RNA) or mature crRNA to the Cas9 protein, which contains a guide sequence that hybridizes to a target sequence and thereby binds the Cas9 protein to the target sequence. This complex is capable of recognizing and cleaving polynucleotides that hybridize to the single-stranded guide RNA or mature crRNA.
[0070] Therefore, in the case of forming a CRISPR / Cas9 complex, a "target sequence" refers to a polynucleotide targeted by a guide sequence designed to have targeting, such as a sequence complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of the CRISPR / Cas9 complex. Complete complementarity is not required, as long as there is enough complementarity to cause hybridization and promote the formation of the CRISPR / Cas complex. The target sequence can include any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast.
[0071] The term "target sequence" or "target polynucleotide" as used herein can be any polynucleotide that is endogenous or exogenous to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas protein used, but the PAM is typically a 2-5 base sequence adjacent to the protospacer sequence (target sequence). Those skilled in the art will be able to identify the PAM sequence for use with a given Cas protein.
[0072] As used herein, the terms "polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and refer to single-stranded or double-stranded RNA or DNA polymers that optionally contain synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single-letter designations as follows: "A" for adenosine or deoxyadenosine (RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for a purine (A or G), "Y" for a pyrimidine (C or T), "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide.
[0073] As used herein, the terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residues is an artificial chemical analog of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" may also include modified forms, including, but not limited to, glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.
[0074] As used herein, the terms sequence "identity" or "homology" have their art-recognized meanings, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the entire length of a polynucleotide or polypeptide or along a region of the molecule (see, e.g., Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are many methods for measuring the identity between two polynucleotides or polypeptides, the term "identity" is well known to those skilled in the art to apply conservative amino acid substitutions in peptides or proteins, and generally can be made without altering the biological activity of the resulting molecule. In general, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially alter biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0075] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which a polynucleotide can be inserted. A vector is called an expression vector when it can express the protein encoded by the inserted polynucleotide, or when it can transcribe the inserted polynucleotide (e.g., to produce mRNA or functional RNA). A vector can be introduced into a host cell through transformation, transduction, or transfection, enabling expression of the genetic material it carries. Vectors are well known to those skilled in the art and include, but are not limited to, plasmid vectors and viral vectors. Vectors may also contain various regulatory sequences that control expression. "Regulatory sequence" and "regulatory element" are used interchangeably herein to refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence that influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. These regulatory sequences may be from different sources or from the same source but arranged in a manner different from that typically found in nature. A vector may also contain an origin of replication.
[0076] As used herein, the term "promoter" refers to a nucleic acid fragment that is capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter that is capable of controlling the transcription of a gene in a cell, whether or not derived from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, or an inducible promoter.
[0077] As used herein, the term "constitutive promoter" refers to a promoter that will generally cause a gene to be expressed in most cell types under most circumstances. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed primarily, but not necessarily exclusively, in one tissue or organ, and may also be expressed in one specific cell or cell type. A "developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. An "inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (environmental, hormonal, chemical signals, etc.).
[0078] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism means transforming the cells of the organism with the nucleic acid or protein so that the nucleic acid or protein can function in the cell. "Transformation," as used herein, includes both stable transformation and transient transformation.
[0079] The term "stable transformation" as used herein refers to the introduction of an exogenous nucleotide sequence into a genome, resulting in the stable inheritance of the exogenous gene. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof.
[0080] As used herein, the term "transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform its function without the exogenous gene being stably inherited. In transient transformation, the exogenous nucleic acid sequence does not integrate into the genome.
[0081] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence, either by traditional Watson-Crick or other non-traditional methods. Percent complementarity refers to the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 out of 10 complementarity represents 50%, 60%, 70%, 80%, 90%, and 100% complementarity). "Complete complementarity" means that all contiguous residues of one nucleic acid sequence form hydrogen bonds with the same number of contiguous residues in the other nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.
[0082] As used herein, the term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and substantially does not hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. Generally speaking, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes to its target sequence. Non-limiting examples of stringent conditions are described in Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes, Section 1, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assay," Elsevier, NY, 1993.
[0083] The term "hybridization" as used herein refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding of the bases between the nucleotide residues. Hydrogen bonding can occur by means of Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex can comprise two chains forming a duplex, three or more chains forming a multi-chain complex, a single self-hybridizing chain, or any combination thereof. A hybridization reaction can constitute a step in a broader process (such as the beginning of PCR or the cutting of polynucleotides via an enzyme). A sequence that can hybridize to a given sequence is referred to as the "complement" of the given sequence.
[0084] Cas9 protein
[0085] In a first aspect, the present invention provides a Cas9 protein, wherein the Cas9 protein is:
[0086] A SauriCas9-HF protein having the amino acid sequence of SEQ ID NO: 1, or a homolog thereof having at least 80% sequence identity to the amino acid sequence of SEQ ID NO: 1 and retaining its biological activity;
[0087] A Sha2Cas9-HF protein having the amino acid sequence of SEQ ID NO: 2, or a homolog thereof having at least 80% sequence identity to the amino acid sequence of SEQ ID NO: 2 and retaining its biological activity; or
[0088] A Sa-SlugCas9-HF protein having the amino acid sequence shown in SEQ ID NO: 3, or a homolog of the amino acid sequence having at least 80% sequence identity with the amino acid sequence shown in SEQ ID NO: 3 and retaining its biological activity.
[0089] The "at least 80% sequence identity" can be at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, at least 99.999%, at least 100%, or any percentage of sequence identity between 80% and 100%.
[0090] In the present invention, the so-called "biological activity" of the Cas9 protein refers to the activity of the protein binding to the single-stranded guide RNA, the endonuclease activity (including single-stranded cleavage activity and double-stranded cleavage activity), and / or the activity of binding to and cleaving a specific site of the target sequence under the guidance of the guide RNA (gRNA), but is not limited thereto.
[0091] Derivatized proteins
[0092] The Cas9 protein can be derivatized, for example, by linking it to another molecule (e.g., another protein or polypeptide). Generally, derivatization of the protein (e.g., labeling) does not adversely affect the desired activity of the protein (e.g., its activity in binding to a single-stranded guide RNA, its endonuclease activity, its activity in binding to and cleaving a specific site of a target sequence under the guidance of a guide RNA). Therefore, the Cas9 protein of the present invention is also intended to include such derivatized forms. For example, the Cas9 protein of the present invention can be functionally linked (by chemical coupling, genetic fusion, non-covalent linkage, or other means) to one or more other molecular moieties, such as another protein or polypeptide, a detectable label, a pharmaceutical agent, and the like.
[0093] In particular, the Cas9 protein can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to provide targeting to the Cas9 protein of the present invention. For example, it can be linked to a detectable label to facilitate detection of the Cas9 protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing, and / or purification of the Cas9 protein of the present invention.
[0094] Thus, in a second aspect, the present invention provides a conjugate comprising:
[0095] a) the Cas9 protein described in the first aspect;
[0096] b) modified parts; and
[0097] c) an optional linker for connecting the Cas9 protein to the modification moiety.
[0098] It is understood that in addition to the Cas9 protein itself, the Cas9 protein can also be combined with other substances such as other proteins or labelable tags to impart other functionalities.
[0099] Thus, in a specific embodiment, the modifying moiety may be another protein or polypeptide, a detectable label, or a combination thereof.
[0100] In a further embodiment, the additional protein or polypeptide is selected from one or more of an epitope tag, a reporter protein or a nuclear localization signal (NLS) sequence, a cytosine deaminase (CBE), adenine deaminase (ABE), a cytosine methylase DNMT3A and MQ1, a cytosine demethylase Tet1, a transcriptional activator protein VP64, p65 and RTA, a transcriptional repressor protein KRAB, a histone acetylase p300, a histone deacetylase LSD1, and an endonuclease FokI.
[0101] Epitope tags are well known to those skilled in the art, and examples thereof include but are not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select a suitable epitope tag according to the desired purpose (e.g., purification, detection or tracing).
[0102] Reporter proteins are well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0103] Detectable labels are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0104] The Cas9 protein of the present invention can be coupled, conjugated, or fused to the modifying moiety via a linker, or can be directly linked to the modifying moiety without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers comprising 1-50 amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0105] In a third aspect, the present invention provides a fusion protein comprising:
[0106] a) the Cas9 protein described in the first aspect;
[0107] b) additional proteins and peptides; and
[0108] c) Optional linkers for connecting the Cas9 protein to the additional proteins and polypeptides.
[0109] As with the second aspect of the present invention, the additional protein or polypeptide may be selected from one or more of an epitope tag, a reporter protein or a nuclear localization signal (NLS) sequence, cytosine deaminase (CBE), adenine deaminase (ABE), cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcriptional activators VP64, p65 and RTA, transcriptional repressor KRAB, histone acetylase p300, histone deacetylase LSD1, and endonuclease FokI.
[0110] Epitope tags are well known to those skilled in the art, and examples thereof include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, and the like. Those skilled in the art will also know how to select an appropriate epitope tag based on the desired purpose (e.g., purification, detection, or tracing). Reporter proteins are well known to those skilled in the art, and examples thereof include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.
[0111] Reporter proteins are well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0112] Detectable labels are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0113] The Cas9 protein of the present invention can be coupled, conjugated, or fused to the additional protein or polypeptide via a linker, or can be directly linked to the additional protein or polypeptide without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers comprising 1-50 amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0114] This invention modifies three Cas9 proteins capable of gene editing in eukaryotic cells—SauriCas9-HF, Sha2Cas9-HF, and Sa-SlugCas9-HF. These modified proteins possess high specificity and can form complexes with the same sgRNA for relatively precise gene editing, making them well-suited for future development as gene therapy tools. This invention expands the scope of gene editing and holds broad application prospects in the field.
[0115] Encoding nucleic acid and vector
[0116] In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding:
[0117] a) the Cas9 protein described in the first aspect;
[0118] b) the conjugate according to the second aspect; or
[0119] c) The fusion protein according to the third aspect.
[0120] In a specific embodiment, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a single-stranded guide RNA; the single-stranded guide RNA comprises a scaffold sequence, wherein the scaffold sequence has:
[0121] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0122] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0123] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0124] The "at least 90% sequence identity" can be at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.9% or at least 100% sequence identity.
[0125] In a specific embodiment, the modification can be one or more of base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening and sequence lengthening.
[0126] In a further embodiment, the shortening of the sequence and the lengthening of the sequence comprise a deletion or addition of 1-10 bases relative to the basic sequence, for example, a deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the basic sequence.
[0127] In another specific embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 5′ end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
[0128] In a preferred embodiment, the CRISPR spacer sequence is a sequence with a length of 21 nucleotides and capable of complementary pairing with the target sequence.
[0129] In a further embodiment, the single-stranded guide RNA further comprises a terminator at the 3′ end of the spacer sequence. As an example, the terminator can be a terminator consisting of a plurality of, for example, at least six (e.g., seven or eight) Us.
[0130] The single-stranded guide RNA can bind to the above-mentioned Cas9 protein, conjugate or fusion protein to form a complex, which can recognize the corresponding PAM and thereby bind to the target sequence, thereby achieving shearing of the target sequence or gene editing.
[0131] In a further embodiment, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 8 or a degenerate sequence thereof, and preferably further comprises a nucleic acid sequence encoding a CRISPR spacer sequence.
[0132] After the isolated nucleic acid molecule of the present invention is transfected into the corresponding cells using certain tools known in the art, such as expression vectors, the isolated nucleic acid molecule of the present invention can express the Cas9 protein, its conjugate or fusion protein, and / or the single-stranded guide RNA described above, and perform corresponding functions therein, such as gene editing.
[0133] In addition, the isolated nucleic acid molecules of the present invention can express Cas9 protein, its conjugate or fusion protein, and single-stranded guide RNA individually / separately, or can express the expression products as a whole. The choice of expression method depends on the specific situation.
[0134] Furthermore, the expression products have the corresponding effects and / or functions described above, which will not be described again for the sake of brevity.
[0135] In a fifth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding:
[0136] a) the Cas9 protein described in the first aspect;
[0137] b) the conjugate according to the second aspect; or
[0138] c) The fusion protein according to the third aspect.
[0139] In a specific embodiment, the vector comprises the nucleic acid sequence shown in any one of SEQ ID NO: 4, SEQ ID NO: 5 and SEQ ID NO: 6 or a degenerate sequence thereof.
[0140] The vector can be an expression vector, for example a plasmid vector such as a pUC19 vector, an episomal vector, a pAAV2_ITR vector, a retroviral vector, a lentiviral vector, an adenoviral vector or an adeno-associated viral vector.
[0141] In another specific embodiment, the vector further comprises a nucleic acid sequence encoding a single-stranded guide RNA. The single-stranded guide RNA comprises a scaffold sequence having:
[0142] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0143] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0144] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0145] The "at least 90% sequence identity" can be at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.9% or at least 100% sequence identity.
[0146] In a specific embodiment, the modification can be one or more of base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening and sequence lengthening.
[0147] In a further embodiment, the shortening of the sequence and the lengthening of the sequence comprise a deletion or addition of 1-10 bases relative to the basic sequence, for example, a deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the basic sequence.
[0148] In another specific embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 5′ end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
[0149] In a preferred embodiment, the CRISPR spacer sequence is a sequence with a length of 21 nucleotides and capable of complementary pairing with the target sequence.
[0150] In a further embodiment, the single-stranded guide RNA further comprises a terminator at the 3′ end of the spacer sequence. As an example, the terminator can be a terminator consisting of a plurality of, for example, at least six (e.g., seven or eight) Us.
[0151] As described above, after the vector of the present invention is transfected into cells, the coding sequence cloned in the vector can be expressed as Cas9 protein, its conjugate or fusion protein, and / or the single-stranded guide RNA described above, and perform corresponding functions therein, such as gene editing.
[0152] Alternatively, multiple vectors, such as two vectors, can be transfected into cells, one of which expresses the Cas9 protein, its conjugate, or fusion protein, while the other expresses the single-stranded guide RNA. The expressed Cas9 protein, its conjugate, or fusion protein then forms a complex with the expressed single-stranded guide RNA, where it can perform its corresponding function, such as gene editing.
[0153] Of course, the nucleic acid sequence encoding the Cas9 protein, its conjugate or fusion protein and the nucleic acid sequence encoding the single-stranded guide RNA can also be cloned into a vector, so that after the vector is transfected into the cell, both the Cas9 protein, its conjugate or fusion protein and the single-stranded guide RNA are expressed and perform corresponding functions, such as gene editing.
[0154] CRISPR / Cas9 gene editing system
[0155] In a sixth aspect, the present invention provides a CRISPR / Cas9 gene editing system comprising:
[0156] a) a protein component comprising the Cas9 protein described in the first aspect, the conjugate described in the second aspect; or the fusion protein described in the third aspect;
[0157] b) a nucleic acid component comprising: a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a scaffold sequence, wherein the scaffold sequence has:
[0158] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0159] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0160] (iii) a nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity;
[0161] Furthermore, the protein component and the nucleic acid component are combined with each other to form a complex.
[0162] The "at least 90% sequence identity" can be at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.9% or at least 100% sequence identity.
[0163] In a specific embodiment, the modification can be one or more of base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening and sequence lengthening.
[0164] In a further embodiment, the shortening of the sequence and the lengthening of the sequence comprise a deletion or addition of 1-10 bases relative to the basic sequence, for example, a deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the basic sequence.
[0165] In another specific embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 5′ end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
[0166] In a preferred embodiment, the CRISPR spacer sequence is a sequence with a length of 21 nucleotides and capable of complementary pairing with the target sequence.
[0167] In a further embodiment, the single-stranded guide RNA further comprises a terminator at the 3′ end of the spacer sequence. As an example, the terminator can be a terminator consisting of a plurality of, for example, at least six (e.g., seven or eight) Us.
[0168] The CRISPR / Cas9 gene editing system of the present invention can be directly composed of the Cas9 protein described herein, its homologs, or their conjugates or fusion proteins and the single-stranded guide RNA described herein, or can be composed of the expression product obtained by expressing the vector described herein. The CRISPR / Cas9 gene editing system achieves recognition, localization, cleavage and gene editing of the target sequence through the combined action of the Cas9 protein and the single-stranded guide RNA contained therein.
[0169] The CRISPR / Cas9 gene editing system of the present invention is capable of precisely localizing the target sequence. The term "precisely localizing" has two meanings: first, the CRISPR / Cas9 gene editing system itself is capable of recognizing and binding to the target sequence; second, it is capable of directing other proteins fused to the Cas9 protein, or proteins that specifically recognize the sgRNA, to the target sequence.
[0170] The CRISPR / Cas9 gene editing system of the present invention has a low tolerance for non-target sequences. Herein, "low tolerance" means that the CRISPR / Cas9 gene editing system of the present invention is substantially unable or completely unable to recognize and bind to non-target sequences, or substantially unable or completely unable to bring other proteins fused to the Cas9 protein or proteins that specifically recognize the sgRNA to the location of non-target sequences.
[0171] cell
[0172] In a seventh aspect, the present invention provides a cell comprising: the isolated nucleic acid molecule according to the fourth aspect, or the vector according to the fifth aspect.
[0173] As an example, the cell can be a prokaryotic cell or a eukaryotic cell. For the eukaryotic cell, as an example, it can be an animal cell. For the animal cell, as an example, it can be a mammalian cell such as a human cell.
[0174] method
[0175] In an eighth aspect, the present invention provides a method for gene editing a target sequence in a cell or in vitro environment, the method comprising: contacting any one of the following (1) to (3) with the target sequence in the cell or in vitro environment:
[0176] (1) The Cas9 protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and a single-stranded guide RNA;
[0177] (2) the carrier described in the fifth aspect; and
[0178] (3) The CRISPR / Cas9 gene editing system described in the sixth aspect;
[0179] wherein, upon contact with a target sequence, the Cas9 protein, the conjugate or the fusion protein recognizes the respective protospacer adjacent sequence (PAM), the PAM being located at the 5' end of the target sequence and having the sequence 5'-NNGG;
[0180] Wherein, the single-stranded guide RNA includes a scaffold sequence; the scaffold sequence has:
[0181] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0182] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0183] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0184] In a specific embodiment, the modification can be one or more of base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening and sequence lengthening.
[0185] In a further specific embodiment, the shortening of the sequence and the lengthening of the sequence comprise a deletion or addition of 1-10 bases relative to the basic sequence, for example, a deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the basic sequence.
[0186] In another specific embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 5′ end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
[0187] In a preferred embodiment, the CRISPR spacer sequence is a sequence with a length of 21 nucleotides and capable of complementary pairing with the target sequence.
[0188] In a further embodiment, the single-stranded guide RNA further comprises a terminator at the 3′ end of the spacer sequence. As an example, the terminator can be a terminator consisting of a plurality of, for example, at least six (e.g., seven or eight) Us.
[0189] In a specific embodiment, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell is, for example, an animal cell, the animal cell is, for example, a mammalian cell such as a human cell.
[0190] In another specific embodiment, the gene editing includes one or more of gene knockout of the target sequence, site-specific base changes, site-specific insertions, regulation of gene transcription levels, DNA methylation regulation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking.
[0191] In a further embodiment, the single base transition comprises a transition of the bases adenine to guanine, cytosine to thymine, or cytosine to uracil.
[0192] In another specific embodiment, in the method, the CRISPR spacer sequence of the single-stranded guide RNA forms a perfect base complementary pairing structure with the target sequence, and forms an incomplete base complementary pairing structure with the non-target sequence.
[0193] Herein, the incomplete base complementary pairing structure refers to a structure including a portion of base complementary pairing and a portion of non-base complementary pairing, wherein the non-base complementary pairing includes, for example, base mismatch and / or base bulge.
[0194] In yet another specific embodiment, the incomplete base complementary pairing structure comprises one or more, for example, two or more base mismatches.
[0195] Thus, the Cas9 protein of the present invention can cleave the target site on the target sequence, and under the cleavage action of the Cas9 protein, the target sequence undergoes a double-strand break. Furthermore, when the method is performed intracellularly, the cleaved target sequence can be repaired through intracellular non-homologous end joining repair or homologous recombination repair pathways, thereby achieving gene editing of the target sequence.
[0196] The CRISPR / Cas9 gene editing system of the present invention and the gene editing method using the gene editing system have been found to be able to form a complex with the same sgRNA for gene editing, and SEQ ID NO: 1, SEQ ID NO: 2 and SEQ ID NO: 3 have an editing efficiency of 5%-25%. In addition, for the SauriCas9-HF protein, Sha2Cas9-HF and Sa-SlugCas9-HF protein gene editing systems, the mismatch guide RNA contained therein has a fault tolerance rate close to 0%. Therefore, these gene editing systems can edit target genes with high specificity, have the characteristics of high editing efficiency and low off-target rate, and can be widely used in gene editing in cells or in vitro environments.
[0197] Reagent test kit
[0198] In a ninth aspect, the present invention provides a kit for performing gene editing on a target sequence in a cell or in vitro environment, comprising:
[0199] a) Any one of the following 1) to 4):
[0200] 1) The Cas9 protein of the first aspect, the conjugate of the second aspect, or the fusion protein of the third aspect, and a single-stranded guide RNA;
[0201] 2) the isolated nucleic acid molecule of the fourth aspect;
[0202] 3) The vector described in the fifth aspect; or
[0203] 4) the CRISPR / Cas9 gene editing system described in the sixth aspect; and
[0204] b) Instructions for performing gene editing on the target sequence in a cell or in vitro setting;
[0205] Wherein, the single-stranded guide RNA includes a scaffold sequence; the scaffold sequence has:
[0206] (i) the nucleic acid sequence shown in SEQ ID NO: 7;
[0207] (ii) a nucleic acid sequence that has at least 90% sequence identity to the nucleic acid sequence shown in SEQ ID NO: 7 and retains its biological activity; or
[0208] (iii) A nucleic acid sequence modified based on the nucleic acid sequence of SEQ ID NO: 7 and retaining its biological activity.
[0209] In a specific embodiment, the modification can be one or more of base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening and sequence lengthening.
[0210] In a further specific embodiment, the shortening of the sequence and the lengthening of the sequence comprise a deletion or addition of 1-10 bases relative to the basic sequence, for example, a deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the basic sequence.
[0211] In another specific embodiment, the single-stranded guide RNA may further include a CRISPR spacer sequence at the 5′ end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
[0212] In a preferred embodiment, the CRISPR spacer sequence is a sequence with a length of 21 nucleotides and capable of complementary pairing with the target sequence.
[0213] In a further embodiment, the single-stranded guide RNA further comprises a terminator at the 3′ end of the spacer sequence. As an example, the terminator can be a terminator consisting of a plurality of, for example, at least six (e.g., seven or eight) Us.
[0214] Of course, those skilled in the art will appreciate that the kit of the present invention may also include other reagents that are helpful for gene editing.
[0215] Brief description of the sequence involved in the present invention
[0216] SEQ ID NO: 1: SauriCas9-HF protein sequence
[0217] SEQ ID NO: 2: Sha2Cas9-HF protein sequence
[0218] SEQ ID NO: 3: Sa-SlugCas9-HF protein sequence
[0219] SEQ ID NO: 4: Coding sequence of SauriCas9-HF protein
[0220] SEQ ID NO: 5: Coding sequence of Sha2Cas9-HF protein
[0221] SEQ ID NO: 6: Coding sequence of Sa-SlugCas9-HF protein
[0222] SEQ ID NO: 7: Scaffold sequence used in conjunction with Cas9 protein
[0223] SEQ ID NO: 8: DNA sequence of the scaffold sequence of the single-stranded guide RNA associated with the Cas9 protein Example
[0224] In the following examples, exemplary CRISPR / Cas9 gene editing systems and related applications of the present invention are shown. Unless otherwise specified, the experimental methods used are conventional methods, and unless otherwise specified, the experimental materials used in the following examples are purchased from conventional reagent stores. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention.
[0225] It should be noted that the terms used in the description of the present invention are intended only to describe specific embodiments and are not intended to limit the present invention. The above summary of the invention and the detailed description below are intended only to illustrate the present invention and are not intended to limit the present invention in any way. Without departing from the spirit and purpose of the present invention, the scope of the present invention is determined by the appended claims. Example 1
[0226] (1) Construction of plasmid pAAV2_Cas9_ITR
[0227] The amino acid sequences of SauriCas9-HF protein, Sha2Cas9-HF and Sa-SlugCas9-HF protein are shown in SEQ ID NO: 1, SEQ ID NO: 2 and SEQ ID NO: 3, respectively.
[0228]
[0229] The amino acid sequences of each of the aforementioned Cas9 proteins were codon-optimized to obtain gene sequences that are highly expressed in human cells. The optimized gene sequences for the SauriCas9-HF protein, Sha2Cas9-HF protein, and Sa-SlugCas9-HF protein are shown in SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6, respectively.
[0230] The highly expressed gene sequences of the Cas9 proteins represented by SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6 obtained above were synthesized and constructed into the slugCas9 backbone plasmid (Addgene platform, catalog #163793) to obtain the plasmid pAAV2_Cas9_ITR.
[0231] (2) Preparation of linearized plasmid hU6-Sa_tracr
[0232] Plasmid hU6-Sa_tracr (Addgene platform, catalog #135973) was digested with BsaI restriction endonuclease. The scaffold sequence in this plasmid is the sequence shown in SEQ ID NO: 8. The digestion system consisted of 1 μg of plasmid hU6-Sa_tracr, 5 μL of 10× CutSmart buffer (purchased from New England Biolabs), 1 μL of BsaI restriction endonuclease (purchased from New England Biolabs), and 50 μL of water. The digestion system was incubated at 37°C overnight.
[0233] Then, the digested products were electrophoresed on 1% agarose gel at 120 V for 30 min.
[0234] DNA fragments were excised from the agarose gel and recovered using a gel extraction kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water. The DNA fragment was the 3088-bp linearized plasmid hU6-Sa_tracr containing the SaCas9 RNA scaffold.
[0235] The recovered linearized plasmid hU6-Sa_tracr was TM The DNA concentration was determined using a Lite spectrophotometer (ThermoScientific) and the DNA was kept for later use or stored at −20°C for long-term storage.
[0236] (3) Preparation of plasmid hU6-Sa_sgRNA
[0237] Each gRNA was designed, and its sequence is shown in Table 2 below. The corresponding sticky end sequences on both sides of the linearized plasmid hU6-Sa_tracr were added to the sense and antisense strands of each designed gRNA sequence pair, and two oligonucleotide single-stranded DNAs were synthesized. The specific sequences of these two oligonucleotide single-stranded DNAs are also shown in Table 2 below.
[0238]
[0239] The oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system was: 1 μL 100 μM oligo-F, 1 μL 100 μM oligo-R, 28 μL water. After the annealing system was shaken and mixed, it was placed in a PCR instrument and the annealing program was run. The annealing program was: 95°C_5min, 85°C_1min, 75°C_1min, 65°C_1min, 55°C_1min, 45°C_1min, 35°C_1min, 25°C_1min, stored at 4°C, and the cooling rate was 0.3°C / s. After annealing, the obtained product was ligated to the linearized hU6-Sa_tracr plasmid obtained in step (2) using DNA ligase (purchased from NEB).
[0240] Take 1 μL of the obtained ligation product and add it to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubate on ice for 30 min, heat shock at 42°C for 1 min, incubate on ice for 2 min, then add it to 900 μL LB medium and culture at 37°C for 1 hour to activate and recover the Escherichia coli DH5α competent cells.
[0241] The revived E. coli DH5α competent cells were spread on LB solid plates containing corresponding resistance and cultured upside down in a 37°C incubator. The obtained E. coli DH5α monoclonal clones were verified by Sanger sequencing.
[0242] The E. coli DH5α clone that was verified to be correctly connected by sequencing was shaken and the plasmid was extracted to obtain the plasmid hU6-Sa_sgRNA containing the target sgRNA sequence for future use.
[0243] (4) Transfection of HEK293T cell lines with plasmid pAAV2_Cas9_ITR expressing Cas protein and plasmid hU6-Sa_sgRNA expressing sgRNA
[0244] On day 0, HEK293T cells containing the target sequence were plated in 24-well plates according to the transfection requirements, with a cell density of about 30%.
[0245] On day 1, transfection was performed. The transfection process was as follows:
[0246] Take 500 ng of plasmid pAAV2_Cas9_ITR and 300 ng of plasmid hU6-Sa_sgRNA, mix them and add them to 25 μL Opti-MEM medium (purchased from Gibco), and gently pipette to mix.
[0247] The transfection reagent liposome (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI, 100 μM) (purchased from Polysciences) was gently flicked to mix, 1.6 μL or 0.8 μL PEI was added to 25 μL Opti-MEM medium (purchased from Gibco), gently mixed, and allowed to stand at room temperature for 5 min.
[0248] Mix the diluted transfection reagent and diluted plasmid, pipette gently to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37°C, 5% CO2 incubator for further culturing for 3 days.
[0249] (5) Preparation of next-generation sequencing library
[0250] Three days after editing, HEK293T cells were collected and genomic DNA was extracted using a DNA kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP304) according to the instructions provided by the DNA kit.
[0251] The first round of PCR library construction was performed using 2×Q5 Master mix (purchased from NEB) for PCR reaction. The PCR primers are as follows:
[0252]
[0253] The reaction system is as follows:
[0254]
[0255] The PCR operation program is as follows:
[0256]
[0257] The second round of PCR for sequencing library construction was performed using 2×Q5 Master mix. The PCR primers are as follows:
[0258] F2 primer: AATGATACGGCCCACCGAGATCTACACTATAGCCTACACTCTTTCCCTACACGAC
[0259] R2 primer: CAAGCAGAAGACGGCATACGAGATTGCTGGGTGTGACTGGAGTTCAGACGTGTG
[0260] The reaction system is as follows:
[0261]
[0262] The PCR operation program is as follows:
[0263]
[0264] The second-round PCR products were purified using a gel extraction kit according to the manufacturer's instructions to obtain DNA fragments of 241 bp, 185 bp, 266 bp, 266 bp, 274 bp, 266 bp, 266 bp, 249 bp, 241 bp, 185 bp, and 385 bp, representing the sizes of the E4, E7, G1, G3, G4, G5, G6, G8, G9, G10, and S3 sites, respectively. Thus, the next-generation sequencing library was prepared.
[0265] (6) Analysis of second-generation sequencing results
[0266] The prepared second-generation sequencing library was subjected to paired-end sequencing on a high-throughput sequencer Hiseq XTen (Illumina).
[0267] The editing efficiency of each target site calculated by the second generation sequencing is as follows Figure 1-3 As shown, the X-axis represents the target site and the Y-axis represents the editing efficiency (Indels%). Figure 1-3 It can be seen that the gene editing systems containing SauriCas9-HF protein, Sha2Cas9-HF and Sa-SlugCas9-HF protein can all be used for cell gene editing. Example 2
[0268] (1) Construction of plasmid pAAV2_Cas9_ITR
[0269] The amino acid sequences of SauriCas9-HF protein, Sha2Cas9-HF protein and Sa-SlugCas9-HF protein are shown in SEQ ID NO: 1 to SEQ ID NO: 3, respectively.
[0270] The amino acid sequence of the Cas9 protein obtained above was codon-optimized to obtain a gene sequence that is highly expressed in human cells. The gene sequences of the SauriCas9-HF protein, Sha2Cas9-HF protein, and Sa-SlugCas9-HF protein are shown in SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6, respectively.
[0271] The gene sequences for highly expressed Cas9 proteins represented by SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6 obtained above were synthesized and constructed into the slugCas9 backbone plasmid (Addgene platform, catalog #163793) to obtain the plasmid pAAV2_Cas9_ITR.
[0272] (2) Preparation of linearized plasmid hU6-Sa_tracr
[0273] Plasmid hU6-Sa_tracr was digested with BsaI restriction endonuclease. The scaffold sequence in this plasmid is the sequence shown in SEQ ID NO: 8. The digestion system consisted of: 1 μg of plasmid hU6-Sa_tracr, 5 μL of 10× CutSmart buffer (purchased from New Brunswick), 1 μL of BsaI restriction endonuclease (purchased from New Brunswick), and 50 μL of water. The digestion system was incubated at 37°C overnight.
[0274] Then, the digested products were electrophoresed on 1% agarose gel at 120 V for 30 min.
[0275] DNA fragments were excised from the agarose gel and recovered using a gel extraction kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water. The DNA fragment was the 3088-bp linearized plasmid hU6-Sa_tracr containing the SaCas9 RNA scaffold.
[0276] The recovered linearized plasmid hU6-Sa_tracr was used to measure the DNA concentration using a NanoDropTM Lite spectrophotometer and then stored for later use or at -20°C for long-term storage.
[0277] (3) Preparation of plasmid hU6-Sa-on target sgRNA or hU6-Sa-mismatch sgRNA
[0278] The sequences of on-target gRNA and mismatch gRNA were designed, and their corresponding oligonucleotide single-stranded DNAs are shown in Table 4 below, where the mismatch bases are shown as underlined bold bases in the sequence table.
[0279] The resulting single-stranded oligonucleotide DNA corresponding to the on-target gRNA and the oligonucleotide single-stranded DNA corresponding to different mismatch gRNAs were annealed separately. The annealing reaction system consisted of 1 μL of 100 μM oligo-F, 1 μL of 100 μM oligo-R, and 28 μL of water. After vortexing, the annealing system was placed in a PCR instrument and the annealing program was run as follows: 95°C for 5 minutes, 85°C for 1 minute, 75°C for 1 minute, 65°C for 1 minute, 55°C for 1 minute, 45°C for 1 minute, 35°C for 1 minute, and 25°C for 1 minute, then stored at 4°C with a cooling rate of 0.3°C / s. After annealing, the resulting products were ligated to the linearized hU6-Sa_tracr plasmid using DNA ligase (purchased from NEB).
[0280] The revived E. coli DH5α competent cells were spread on LB solid plates containing corresponding resistance and cultured upside down in a 37°C incubator. The obtained E. coli DH5α monoclonal clones were verified by Sanger sequencing.
[0281] The Escherichia coli DH5α clones verified to be correctly connected by sequencing were shaken and plasmids were extracted to obtain the plasmid hU6-Sa-on target sgRNA expressing the above-mentioned On target gRNA sequence and the plasmid hU6-Sa-mismatch sgRNA expressing the above-mentioned different mismatch gRNA sequences, respectively, for use.
[0282] (4) The obtained plasmid hU6-Sa-on target sgRNA expressing the on target gRNA sequence and the plasmid hU6-Sa-mismatch sgRNA expressing the above-mentioned different mismatch gRNA sequences were transfected with pAAV2_Cas9_ITR into the GFP reporter system HEK293T cell line containing the target sequence (GGCTCGGAGATCATCATTGCG) by liposome transfection.
[0283]
[0284] The HEK293T cell line containing the target sequence GFP reporter system is obtained by inserting a PAM sequence and a specific target sequence between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. The PAM sequence is then integrated into HEK293T cells via lentiviral infection to obtain a HEK293T cell line containing the target sequence GFP reporter system. After the gene editing system cuts the target sequence, the cell's own repair system causes some cells to restore the GFP reading frame, producing green fluorescence. Flow cytometric analysis of the GFP-positive cell ratio can be used to evaluate the editing ability and specificity of the gene editing system.
[0285] The above transfection process includes the following steps:
[0286] On day 0, HEK293T cells containing the target sequence of the GFP reporter system were plated in 24-well plates according to the transfection requirements, and the cell density was controlled at 30%.
[0287] The GFP reporter system HEK293T cell line containing the target sequence contains the nucleotide sequence of CMV-ATG-PAM-target site-GFP, where the PAM sequence is shown in Figure 4 , the sequence of the target site is the target sequence GGCTCGGAGATCATCATTGCG.
[0288] On day 1, transfection was performed. The transfection process was as follows:
[0289] (1) 500 ng of pAAV2_Cas9_ITR plasmid and 300 ng of hU6-Sa_on target gRNA plasmid, or (2) 500 ng of pAAV2_Cas9_ITR plasmid and 300 ng of hU6-Sa_mismatch gRNA plasmid were mixed and added to 25 μL Opti-MEM medium and gently pipetted to mix. Lipofectamine®2000 (purchased from Invitrogen) or PEI (purchased from Polysciences) was gently flicked to mix, and 1.6 μL of Lipofectamine®2000 or 0.8 μL of PEI was added to 25 μL Opti-MEM medium, gently mixed, and allowed to stand at room temperature for 5 minutes.
[0290] The diluted plasmid and the diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 20 min, then added to the culture medium of the HEK293T cell line containing the GFP reporter system of the target sequence and placed in a 37°C, 5% CO2 incubator for further culture.
[0291] Flow cytometry was used to analyze the editing efficiency and off-target rate of the CRISPR / Cas9 gene editing system of the present invention on the target sequence.
[0292] Specifically, HEK293T cell lines were collected after being cultured in a CO2 incubator for 5 days, and their specificity was detected by flow cytometry (BD Biosciences FACSCalibur). The GFP positive ratio was analyzed and plotted using FlowJo analysis software.
[0293] The specific detection results of the CRISPR / Cas9 gene editing system of the present invention in the GFP reporter system HEK293T cell line containing the target sequence are shown in Figure 4-6 The top horizontal bar shows a schematic diagram of the GFP reporter system. A specific PAM sequence and target sequence are inserted between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. Therefore, when the gene editing system cuts the target sequence, the cell's own repair system restores the GFP reading frame in some cells, producing green fluorescence. Figure 4-6 The Y-axis in the middle and lower bar graph represents the percentage of GFP-positive cells (%), and the X-axis represents the oligonucleotide single-stranded DNA sequences corresponding to the on-target gRNA and mismatch gRNA. Figure 4-6 It can be seen that the CRISPR gene editing system of the present invention edited the target sites in the GFP reporter system HEK293T cell line, and the gene editing ratio mediated by mismatch gRNA was significantly lower than that mediated by on-target gRNA. In addition, in the research results of the gene editing system containing SauriCas9-HF protein, Sha2Cas9-HF protein and Sa-SlugCas9-HF protein, no obvious mismatch was found in all double mismatches, indicating that the gene editing system containing SauriCas9-HF protein, Sha2Cas9-HF protein and Sa-SlugCas9-HF protein has extremely high requirements for complete matching between gRNA and target sequence, and has a low fault tolerance rate and high safety in practical application.
Claims
1. A Cas9 protein, wherein the Cas9 protein is a SauriCas9-HF protein having an amino acid sequence shown in SEQ ID NO:
1.
2. A conjugate comprising: a) the Cas9 protein according to claim 1; b) a modifying moiety; the modifying moiety is selected from another protein or polypeptide, a detectable label or a combination thereof; the additional protein or polypeptide is selected from one or more of an epitope tag, a reporter protein or a nuclear localization signal sequence, a cytosine deaminase, adenine deaminase, a cytosine methylase DNMT3A and MQ1, a cytosine demethylase Tet1, a transcriptional activator protein VP64, p65 and RTA, a transcriptional repressor protein KRAB, a histone acetylase p300, a histone deacetylase LSD1, and an endonuclease FokI; and c) a linker for connecting the Cas9 protein and the modified portion. The conjugate according to claim 2 , wherein the linker is a linker with a length of 1-50 amino acids.
4. A fusion protein comprising: a) the Cas9 protein according to claim 1; b) additional proteins and polypeptides; the additional proteins and polypeptides are selected from one or more of epitope tags, reporter proteins or nuclear localization signal sequences, cytosine deaminase, adenine deaminase, cytosine methylases DNMT3A and MQ1, cytosine demethylase Tet1, transcriptional activators VP64, p65 and RTA, transcriptional repressors KRAB, histone acetylase p300, histone deacetylase LSD1, and endonuclease FokI; and c) a linker for connecting the Cas9 protein to the additional protein and polypeptide. The fusion protein according to claim 4 , wherein the linker is a linker with a length of 1-50 amino acids.
6. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding: a) the Cas9 protein according to claim 1; b) the conjugate according to claim 2 or 3; or c) The fusion protein according to claim 4 or 5.
7. The isolated nucleic acid molecule of claim 6, wherein the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a single-stranded guide RNA; the single-stranded guide RNA comprises a scaffold sequence, the nucleic acid sequence of the scaffold sequence being: (i) the nucleic acid sequence shown in SEQ ID NO: 7; or (ii) a nucleic acid sequence modified based on the nucleic acid sequence shown in SEQ ID NO: 7 and retaining its biological activity; in, The modification is one or more of base phosphorylation, base sulfidation, base methylation, and base hydroxylation.
8. The isolated nucleic acid molecule according to claim 7, wherein The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
9. The isolated nucleic acid molecule according to claim 8, wherein The CRISPR spacer sequence is a sequence with a length of 21 nucleotides that can complementarily pair with the target sequence.
10. A vector comprising a nucleic acid sequence encoding: a) the Cas9 protein according to claim 1; b) the conjugate according to claim 2 or 3; or c) The fusion protein according to claim 4 or 5.
11. The carrier according to claim 10, wherein The nucleic acid sequence of the vector is the nucleic acid sequence shown in SEQ ID NO: 4 or a degenerate sequence thereof.
12. The carrier according to claim 10 or 11, wherein The vector is a plasmid vector, an episome vector, a retroviral vector, a lentiviral vector, an adenoviral vector or an adeno-associated viral vector.
13. The carrier according to claim 12, wherein The vector is a pAAV2_ITR vector.
14. The carrier according to claim 12, wherein The plasmid vector is a pUC19 vector.
15. The carrier according to claim 10 or 11, wherein The vector further comprises a nucleic acid sequence encoding a single-stranded guide RNA; the single-stranded guide RNA comprises a scaffold sequence, and the nucleic acid sequence of the scaffold sequence is: (i) the nucleic acid sequence shown in SEQ ID NO: 7; or (ii) a nucleic acid sequence modified based on the nucleic acid sequence shown in SEQ ID NO: 7 and retaining its biological activity; Wherein, the modification is one or more of base phosphorylation, base sulfidation, base methylation, and base hydroxylation.
16. The carrier according to claim 15, wherein The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the scaffold sequence, wherein the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and is capable of complementary pairing with the target sequence.
17. The carrier according to claim 16, wherein The CRISPR spacer sequence is a sequence with a length of 21 nucleotides that can complementarily pair with the target sequence.
18. A CRISPR / Cas9 gene editing system comprising: a) a protein component comprising the Cas9 protein according to claim 1, the conjugate according to claim 2 or 3; or the fusion protein according to claim 4 or 5; b) a nucleic acid component comprising: a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a scaffold sequence, and the nucleic acid sequence of the scaffold sequence is: (i) the nucleic acid sequence shown in SEQ ID NO: 7; or (ii) a nucleic acid sequence modified based on the nucleic acid sequence shown in SEQ ID NO: 7 and retaining its biological activity; in, The modification is one or more of base phosphorylation, base sulfidation, base methylation, and base hydroxylation; Furthermore, the protein component and the nucleic acid component are combined with each other to form a complex.
19. The CRISPR / Cas9 gene editing system according to claim 18, wherein The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the scaffold sequence; the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and can complementarily pair with the target sequence.
20. The CRISPR / Cas9 gene editing system according to claim 19, wherein The CRISPR spacer sequence is a sequence with a length of 21 nucleotides that can complementarily pair with the target sequence.
21. A cell comprising: the isolated nucleic acid molecule of any one of claims 6 to 9, or the vector of any one of claims 10 to 17.
22. The cell according to claim 21, wherein The cell is a prokaryotic cell or an animal cell.
23. The cell according to claim 22, wherein The animal cells are mammalian cells.
24. A method for non-therapeutic gene editing of a target sequence in a cell or in vitro environment, the method comprising: Contacting any of the following (1) to (3) with a target sequence in a cell or in vitro environment: (1) The Cas9 protein according to claim 1, the conjugate according to claim 2 or 3, or the fusion protein according to claim 4 or 5, and a single-stranded guide RNA; (2) The carrier according to any one of claims 15 to 17; and (3) The CRISPR / Cas9 gene editing system according to any one of claims 18 to 20; wherein, upon contact with a target sequence, the Cas9 protein, the conjugate or the fusion protein recognizes the respective protospacer adjacent sequence (PAM), the PAM being located at the 5' end of the target sequence and having the sequence 5'-NNGG; Wherein, the single-stranded guide RNA includes a scaffold sequence; the nucleic acid sequence of the scaffold sequence is: (i) the nucleic acid sequence shown in SEQ ID NO: 7; or (ii) a nucleic acid sequence modified based on the nucleic acid sequence shown in SEQ ID NO: 7 and retaining its biological activity; Wherein, the modification is one or more of base phosphorylation, base sulfidation, base methylation and base hydroxylation.
25. The method according to claim 24, wherein The cell is a prokaryotic cell or an animal cell.
26. The method according to claim 25, wherein The animal cells are mammalian cells.
27. The method according to claim 24, wherein The gene editing includes one or more of gene knockout of the target sequence, site-specific base changes, site-specific insertions, regulation of gene transcription levels, DNA methylation regulation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking.
28. The method according to claim 27, wherein The single base conversion includes the conversion of the base adenine to guanine, the conversion of cytosine to thymine, or the conversion of cytosine to uracil.
29. The method according to claim 24, wherein The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the scaffold sequence; the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and can complementarily pair with the target sequence.
30. The method according to claim 29, wherein The CRISPR spacer sequence is a sequence with a length of 21 nucleotides that can complementarily pair with the target sequence.
31. The method according to claim 29, wherein The CRISPR spacer sequence forms a perfect base complementary pairing structure with the target sequence, and forms an incomplete base complementary pairing structure with the non-target sequence.
32. The method according to claim 31, wherein The incomplete base complementary pairing structure includes a structure with one or more base mismatches.
33. The method according to claim 32, wherein The incomplete base complementary pairing structure includes a structure with two or more base mismatches.
34. A kit for gene editing a target sequence in a cell or in vitro environment, comprising: a) Any one of the following 1) to 4): 1) The Cas9 protein according to claim 1, the conjugate according to claim 2 or 3, or the fusion protein according to claim 4 or 5, and a single-stranded guide RNA; 2) The isolated nucleic acid molecule according to any one of claims 7 to 9; 3) The vector according to any one of claims 15 to 17; or 4) The CRISPR / Cas9 gene editing system according to any one of claims 18 to 20; and b) Instructions for performing gene editing on the target sequence in a cell or in vitro setting; Wherein, the single-stranded guide RNA includes a scaffold sequence; the nucleic acid sequence of the scaffold sequence is: (i) the nucleic acid sequence shown in SEQ ID NO: 7; or (ii) a nucleic acid sequence modified based on the nucleic acid sequence shown in SEQ ID NO: 7 and retaining its biological activity; Wherein, the modification is one or more of base phosphorylation, base sulfidation, base methylation, and base hydroxylation.
35. The kit according to claim 34, wherein The single-stranded guide RNA further includes a CRISPR spacer sequence at the 5' end of the scaffold sequence; the CRISPR spacer sequence is a sequence with a length of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides and can complementarily pair with the target sequence.
36. The kit according to claim 35, wherein The CRISPR spacer sequence is a sequence with a length of 21 nucleotides that can complementarily pair with the target sequence.
Citation Information
Patent Citations
CRISPR / SauriCas9 gene editing system and application thereof
CN110499335A
Cas9 protein, gene editing system containing Cas9 protein and application
CN113583999A