Cas12b proteins, single guide rnas, gene editing systems comprising the same, and related applications
By mutating the ChCas12b protein and optimizing the sgRNA scaffold sequence, the problem that the CRISPR/Cas system cannot distinguish single-base mutations has been solved, achieving highly specific and efficient gene editing and expanding the application prospects of gene editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUDAN UNIVERSITY
- Filing Date
- 2023-04-24
- Publication Date
- 2026-05-05
AI Technical Summary
Existing CRISPR/Cas systems cannot effectively distinguish single-base mutations, leading to off-target problems, potentially damaging normal genes, and lacking therapeutic effects on dominant mutations.
By mutating amino acid residues R453, D496, and R1137 of the wild-type ChCas12b protein, a Cas12b mutant protein was developed, and the scaffold sequence of the single-stranded guide RNA was optimized to form a CRISPR/Cas12b gene editing system with higher specificity and editing efficiency.
It enables the identification of single-base differences at target sites and the knockout of mutated alleles, improving the specificity and efficiency of gene editing and making it suitable for development as a gene therapy tool.
Smart Images

Figure CN116751762B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, and more specifically to a Cas12b mutant protein, a single-stranded guide RNA used in conjunction with said protein, a CRISPR / Cas12b gene editing system comprising them, and their related applications. Background Technology
[0002] The CRISPR / Cas12b system is an acquired immune system evolved by bacteria and archaea to defend against invasion by exogenous viruses or plasmids. The CRISPR / Cas12b system contains tracrRNA (trans-activating RNA) and crRNA (CRISPR-derived RNA), which together with Cas12b form a complex to function. tracrRNA and crRNA can fuse into single-stranded guide RNA (sgRNA) through a linker sequence. When DNA breaks occur, two main DNA damage repair mechanisms within the cell are responsible for repair: non-homologous end-joining (NHEJ) and homologous recombination (HR). NHEJ repair results in base deletions or insertions, allowing for gene knockout; HR repair, when a homologous template is provided, allows for site-specific gene insertion and precise base substitution.
[0003] Beyond basic scientific research, the CRISPR / Cas12b gene editing system also holds broad clinical application prospects. When using CRISPR / Cas gene editing systems for gene therapy, the biggest concern is off-target effects. Off-target effects can damage normal genes, leading to cancer. Most CRISPR / Cas systems exhibit off-target effects. Therefore, there is an urgent need to improve specificity and reduce off-target effects. Furthermore, many dominant mutations are caused by single-base mutations. If the mutated allele can be knocked out, leaving only the normal allele, a therapeutic effect can be achieved. This requires the CRISPR / Cas system to be able to distinguish single-base mutations. Currently, most CRISPR / Cas systems cannot distinguish single-base mutations. Summary of the Invention
[0004] Through repeated research, the inventors identified mutation sites related to the cleavage specificity of wild-type ChCas12b protein, thereby obtaining a series of Cas12b mutant proteins, all of which can form a CRISPR / Cas12b gene editing system with enhanced cleavage specificity for effective gene editing with single-stranded guide RNA, thus completing this invention.
[0005] In summary, in a first aspect of the present invention, a Cas12b mutant protein is provided, the Cas12b mutant protein comprising a mutation at one or more amino acid residues of R453, D496 and R1137 corresponding to the wild-type ChCas12b protein.
[0006] In a second aspect, the present invention provides a conjugate comprising:
[0007] a) The Cas12b mutant protein described in the first aspect;
[0008] b) Modified parts; and
[0009] c) Optional adapters for connecting the Cas12b mutant protein to the modified portion.
[0010] In a third aspect, the present invention provides a fusion protein comprising:
[0011] a) The Cas12b mutant protein described in the first aspect;
[0012] b) Other proteins and peptides; and
[0013] c) Optional adapters for connecting the Cas12b mutant protein to the other proteins and peptides.
[0014] In a fourth aspect, the present invention provides a single-stranded guide RNA comprising a scaffold sequence and a CRISPR spacer sequence from the 5' end to the 3' end, the scaffold sequence comprising tracrRNA and a repeat sequence from the 5' end to the 3' end, and having a nucleic acid sequence that is truncated or mutated relative to the nucleic acid sequence shown in SEQ ID NO:2.
[0015] In a fifth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding the following:
[0016] a) The Cas12b mutant protein described in the first aspect;
[0017] b) The conjugate described in the second aspect; or
[0018] c) The fusion protein described in the third aspect.
[0019] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect.
[0020] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding the following:
[0021] a) The Cas12b mutant protein described in the first aspect;
[0022] b) The conjugate described in the second aspect; or
[0023] c) The fusion protein described in the third aspect.
[0024] In an eighth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect.
[0025] In a ninth aspect, the present invention provides a CRISPR / Cas12b gene editing system comprising:
[0026] 1) Protein components, which include:
[0027] a) The Cas12b mutant protein described in the first aspect;
[0028] b) The conjugate described in the second aspect; or
[0029] c) The fusion protein described in the third aspect;
[0030] d) Wild-type ChCas12b protein, its conjugates, or fusion proteins;
[0031] 2) The nucleic acid sequence of the single-stranded guide RNA described in the fourth aspect, or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 5;
[0032] The protein component described herein, when it is d) wild-type ChCas12b protein, its conjugate or fusion protein, may only be used in conjunction with the single-stranded guide RNA described in the fourth aspect;
[0033] Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex.
[0034] In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules as described in the fifth or sixth aspect, or a carrier as described in the seventh or eighth aspect.
[0035] In an eleventh aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (7) with the target sequence in the intracellular or in vitro environment:
[0036] (1) The Cas12b mutant protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and single-stranded guide RNA;
[0037] (2) Wild-type ChCas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA described in the fourth aspect;
[0038] (3) The isolated nucleic acid molecules described in the fifth aspect and, as appropriate, isolated nucleic acid molecules containing nucleic acid sequences encoding single-stranded guide RNA;
[0039] (4) An isolated nucleic acid molecule comprising encoding wild-type ChCas12b protein, its conjugate or fusion protein, and the isolated nucleic acid molecule described in the sixth aspect;
[0040] (5) The vectors described in the seventh aspect and, as appropriate, vectors containing nucleic acid sequences encoding single-stranded guide RNA;
[0041] (6) A vector comprising a nucleic acid sequence encoding the wild-type ChCas12b protein, its conjugates, or a fusion protein, and the vector described in aspect eight; and
[0042] (7) The CRISPR / Cas12b gene editing system described in the ninth aspect;
[0043] The Cas12b protein, the conjugate, or the fusion protein recognizes a protospacer adjacent sequence (PAM) located at the 5' end of the target sequence and having a 5'-WTN sequence.
[0044] In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising:
[0045] a) Choose any one of (1) to (7) below:
[0046] (1) The Cas12b mutant protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and single-stranded guide RNA;
[0047] (2) Wild-type ChCas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA described in the fourth aspect;
[0048] (3) The isolated nucleic acid molecules described in the fifth aspect and, as appropriate, isolated nucleic acid molecules containing nucleic acid sequences encoding single-stranded guide RNA;
[0049] (4) An isolated nucleic acid molecule comprising encoding wild-type ChCas12b protein, its conjugate or fusion protein, and the isolated nucleic acid molecule described in the sixth aspect;
[0050] (5) The vectors described in the seventh aspect and, as appropriate, vectors containing nucleic acid sequences encoding single-stranded guide RNA;
[0051] (6) A vector comprising a nucleic acid sequence encoding the wild-type ChCas12b protein, its conjugates, or a fusion protein, and the vector described in aspect eight; and
[0052] (7) The CRISPR / Cas12b gene editing system described in aspect nine; and
[0053] b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment.
[0054] The inventors of this invention have developed a variety of Cas12b mutant proteins based on the wild-type ChCas12b protein (having the protein sequence shown in SEQ ID NO:1). The Cas12b mutant proteins contain mutations at one or more amino acid residues, such as two or three mutations, corresponding to R453, D496 and R1137 of the wild-type ChCas12b protein, thereby obtaining higher specificity.
[0055] Furthermore, the inventors of this invention have developed a variety of sgRNA scaffold sequence variants based on the wild-type ChCas12b-sgRNA scaffold sequence (having the nucleic acid sequence shown in SEQ ID NO:2). These variants are appropriately truncated or have base mutations at the bubble positions relative to the wild-type ChCas12b-sgRNA scaffold sequence, thereby enabling the CRISPR / Cas12b gene editing system containing these sgRNA scaffold sequence variants to achieve higher editing efficiency.
[0056] The Cas12b mutant protein of this invention can form a complex with further modified sgRNA for gene editing. The gene editing tool using the Cas12b mutant protein and sgRNA of this invention can recognize very simple PAMs, i.e., WTNs, and has high editing efficiency and specificity. Furthermore, due to the relatively small number of amino acids and small molecular weight of the protein, it can be easily packaged into vectors such as adeno-associated viruses, making it very suitable for later development as a gene therapy tool. More importantly, the gene editing system using the Cas12b mutant protein and sgRNA of this invention overcomes the limitation of existing CRISPR / Cas systems in distinguishing single-base mutations, enabling the differentiation of single-base differences at target sites. This allows for the knockout of mutated alleles, providing an important foundation for clinical gene therapy using the CRISPR / Cas12b gene editing system. Therefore, this invention further expands the scope of gene editing and has broad application prospects in the field of gene editing. Attached Figure Description
[0057] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below.
[0058] Figure 1 The structure of wild-type ChCas12b-sgRNA is shown, where the blue part is the repeat sequence, the black part is the tracrRNA sequence, and the red part is the spacer sequence.
[0059] Figure 2 Exemplary scaffold sequence variants 1-5 of the present invention are shown, wherein blue represents repetitive sequences, black represents tracrRNA sequences, and red uppercase represents point-mutated nucleotide sequences.
[0060] Figure 3 This diagram illustrates a comparison of the specificity detection results of the CRISPR / ChCas12b-R453A gene editing system and the CRISPR / ChCas12b gene editing system in the HEK293T cell line of the GFP reporter system.
[0061] Figure 4 This diagram illustrates a comparison of the specificity detection results of the CRISPR / ChCas12b-D496A gene editing system and the CRISPR / ChCas12b gene editing system in the HEK293T cell line of the GFP reporter system.
[0062] Figure 5 This diagram illustrates a comparison of the specificity detection results of the CRISPR / ChCas12b-R1137A gene editing system and the CRISPR / ChCas12b gene editing system in the HEK293T cell line of the GFP reporter system.
[0063] Figure 6 The results show the editing efficiency of gene editing systems containing ChCas12b-sgRNA scaffold sequence variants and wild-type ChCas12b-sgRNA in the HEK293T cell line containing the target sequence of the GFP reporter system.
[0064] Figure 7 The results show the editing efficiency of a gene editing system containing wild-type ChCas12b protein and different sgRNA scaffold sequences (wild-type ChCas12b-sgRNA, ChCas12b-sgRNA scaffold sequence variant 2 or 5) for three target sites.
[0065] Figure 8The results show the editing efficiency of gene editing systems containing different Cas12b proteins (wild-type ChCas12b protein, ChCas12b-R453A, ChCas12b-D496A, or ChCas12b-R1137A mutant protein) and ChCas12b-sgRNA scaffold sequence variant 2 for two endogenous target sites, VEGFA and EMX1.
[0066] Figure 9 The results show the editing efficiency of gene editing systems containing different Cas12b proteins (wild-type ChCas12b protein, ChCas12b-R453A, ChCas12b-D496A, or ChCas12b-R1137A mutant protein) and wild-type sgRNA for multiple target sites.
[0067] Figure 10 yes Figure 9 The summary chart of the results also shows the editing efficiency of different Cas12b proteins on multiple target sites.
[0068] Figure 11 The results show the editing efficiency of a gene editing system containing the ChCas12b-D496A mutant protein and the ChCas12b-sgRNA scaffold sequence variant 2 for target sites containing single nucleotide polymorphisms (SNPs). Detailed Implementation
[0069] The present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the following description is merely illustrative and is not intended to limit the scope of the invention; the scope of protection of the invention is defined by the appended claims. Furthermore, those skilled in the art will understand that modifications can be made to the technical solutions of the present invention without departing from its spirit and intent. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0070] definition
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter pertains. Before a detailed description of the invention, the following definitions are provided to better understand it.
[0072] In cases where numerical ranges are provided, such as concentration ranges, percentage ranges, or ratio ranges, it should be understood that, unless the context explicitly specifies otherwise, all intermediate values between the upper and lower limits of the range, up to one-tenth of the lower limit unit, and any other values or intermediate values within the range are included in the subject matter. The upper and lower limits of these smaller ranges may be independently included in the smaller ranges, and such embodiments are also included in the subject matter, limited by any specific excluded limit values within the range. Where the range includes one or two limit values, the range excluding any one or both of those included limit values is also included in the subject matter.
[0073] In the context of this invention, many embodiments use the expressions "comprising," "including," or "basically / mainly composed of." The expressions "comprising," "including," or "basically / mainly composed of" are generally understood as open-ended expressions, indicating that they include not only the elements, components, parts, or method steps specifically listed after the expression, but also other elements, components, parts, or method steps. However, in this document, the expressions "comprising," "including," or "basically / mainly composed of" can also be understood as closed-ended expressions in certain cases, indicating that they only include the elements, components, parts, or method steps specifically listed after the expression, and do not include any other elements, components, parts, or method steps. In this case, the expression is equivalent to the expression "composed of."
[0074] To better understand this teaching and without limiting its scope, all figures and other numerical values used in the specification and claims to express quantities, percentages, or proportions should, in all cases, be understood to be modified by the term "about." Therefore, unless otherwise stated, the numerical parameters set forth in the following specification and appended claims are approximate values that may vary depending on the desired properties sought. At a minimum, each numerical parameter should be interpreted based at least on the reported significant figures and by applying common rounding techniques.
[0075] The terms "Cas12b protein," "Cas12b," and "Cas" used herein are interchangeable and refer to RNA-guided nucleases, including the Cas12b protein or its functionally active fragments. The Cas12b protein is a protein component of the CRISPR / Cas12b genome editing system that, guided by single-stranded guide RNA (gRNA), targets and cleaves DNA target sequences, forming DNA double-strand breaks (DSBs). DNA double-strand breaks can activate the cell's inherent repair mechanisms of non-homologous end-joining (NHEJ) and homologous recombination (HR), thereby repairing DNA damage in the cell. During the repair process, the specific DNA sequence is edited at specific sites.
[0076] The terms “guide RNA,” “gRNA,” “sgRNA,” or “mature crRNA” as used herein are interchangeable and have the meanings commonly understood by those skilled in the art. Generally, a single-stranded guide RNA may comprise a scaffold sequence and a guide sequence, also referred to herein as guide RNA (or gRNA). In the context of an endogenous CRISPR system, the guide sequence is also referred to as a spacer sequence. In some cases, the guide sequence is any polynucleotide sequence that is sufficiently similar to a target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas12 complex to said target sequence. In some embodiments, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining optimal alignment is within the capabilities of those skilled in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan. As used herein, the term "CRISPR / Cas12b complex" refers to a complex formed by the binding of a single-stranded guide RNA or mature crRNA to the Cas12b protein, which contains a guide sequence that hybridizes to a target sequence and thereby enables the Cas12b protein to bind to said target sequence. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the single-stranded guide RNA.
[0077] Therefore, in the formation of the CRISPR / Cas12 complex, the "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, where hybridization between the target sequence and the guide sequence will promote the formation of the CRISPR / Cas12 complex. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of the CRISPR / Cas12 complex. The target sequence can include any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In other cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast.
[0078] As used herein, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide could be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas protein used, but the PAM is typically a 2-5 base sequence adjacent to the protospacer sequence (target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas protein.
[0079] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” used herein are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.
[0080] In the sequences used in this invention, degenerate bases are sometimes used to represent bases at one or more positions. Degenerate bases can be represented by the letters R, Y, M, K, S, W, H, B, V, D, and N, where R represents A / G, Y represents C / T, M represents A / C, K represents G / T, S represents C / G, W represents A / T, H represents A / T / C, B represents G / T / C, V represents G / A / C, D represents G / A / T, and N represents A / T / C / G.
[0081] The terms “polypeptide,” “peptide,” and “protein” as used herein are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, and also to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.
[0082] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. A vector is called an expression vector when it enables the expression of a protein encoded by the inserted polynucleotide, or when it enables transcription of the inserted polynucleotide (e.g., to generate mRNA or functional RNA). Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to, plasmid vectors and viral vectors. Vectors may also contain various regulatory sequences that regulate expression. The terms "regulatory sequence" and "regulatory element" are used interchangeably herein, referring to a nucleotide sequence located upstream (5' non-coding sequence), midway, or downstream (3' non-coding sequence) of a coding sequence that affects transcription, RNA processing, or stability or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. These regulatory sequences may originate from different sources or from the same source but arranged in a manner different from what is typically found naturally. Additionally, vectors may contain a replication initiation site.
[0083] As used herein, the term "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.
[0084] As used in this article, the term "constitutive promoter" refers to a promoter that generally causes gene expression in most cell types and under most conditions. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to promoters that are primarily, but not necessarily, expressed specifically in a single tissue or organ, and may also be expressed in a specific cell type. "Developmental regulatory promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses a manipulated DNA sequence in response to endogenous or exogenous stimuli (environment, hormones, chemical signals, etc.).
[0085] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.
[0086] As used in this article, the term "stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its subsequent generations.
[0087] The term "transient transformation" as used in this article refers to the introduction of nucleic acid molecules or proteins into cells to perform their functions without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.
[0088] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The complementarity percentage indicates the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementarity out of 10). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0089] As used in this paper, the hybridization-related term "strict condition" refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of a strict condition is described in Tijssen, 1993, *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes*, Section I, Chapter II, "Overview of principles of hybridization and the strategy of nucleic acid probe assay", Elsevier, New York.
[0090] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. Hydrogen bonds can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.
[0091] Cas12b protein
[0092] In a first aspect of the invention, a Cas12b mutant protein is provided, the Cas12b mutant protein comprising a mutation at one or more amino acid residues of R453, D496 and R1137 corresponding to the wild-type ChCas12b protein.
[0093] The wild-type ChCas12b protein used in this paper has the amino acid sequence shown in SEQ ID NO:1.
[0094] In a preferred embodiment, the mutation is one or more mutations selected from R453A, D496A and R1137A.
[0095] In a further preferred embodiment, the mutation is R453A, D496A, or R1137A.
[0096] In yet another specific implementation, the mutation is multiple mutations among R453A, D496A, and R1137A, such as two or three mutations.
[0097] In one specific implementation, the Cas12b mutant protein has at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even higher (e.g., 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 99.99%) sequence identity with the wild-type ChCas12b protein.
[0098] Derivatized proteins
[0099] The Cas12b protein can be derivatized, for example, by linking it to other molecules (e.g., other proteins or peptides). Generally, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to single-stranded guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, the Cas12b protein of the present invention is also intended to include such derivatized forms. For example, the Cas12b protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular moieties, such as other proteins or peptides, detectable labels, pharmaceutical reagents, etc.
[0100] In particular, the Cas12b protein can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the Cas12b protein of the present invention targeted. For example, it can be linked to a detectable tag to facilitate the detection of the Cas12b protein of the present invention. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the Cas12b protein of the present invention.
[0101] Therefore, in a second aspect, the present invention provides a conjugate comprising:
[0102] a) The Cas12b mutant protein described in the first aspect;
[0103] b) Modified parts; and
[0104] c) Optional adapters for connecting the Cas12b mutant protein to the modified portion.
[0105] It is understandable that, in addition to the Cas12b mutant protein itself, the Cas12b mutant protein can also be combined with other substances, such as other proteins or tagged substances, to endow it with other functions.
[0106] Therefore, in one specific implementation, the modified portion may be another protein or polypeptide, a detectable label, or a combination thereof.
[0107] In a further embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0108] Epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).
[0109] Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.
[0110] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0111] The Cas12b protein of the present invention can be coupled, conjugated, or fused to the modified portion via a linker, or it can be directly linked to the modified portion without a linker. Linkers are well known in the art, and examples of them may include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0112] In a third aspect, the present invention provides a fusion protein comprising:
[0113] a) The Cas12b mutant protein described in the first aspect;
[0114] b) Other proteins and peptides; and
[0115] c) Optional adapters for connecting the Cas12b mutant protein to the other proteins and peptides.
[0116] Similar to the second aspect of the invention, the additional protein or polypeptide may be selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
[0117] Epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing). Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0118] Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.
[0119] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.
[0120] The Cas12b mutant protein of the present invention can be coupled, conjugated, or fused to the other proteins or peptides via a linker, or can be directly linked to the other proteins or peptides without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0121] The Cas12b mutant protein of this invention has a relatively small number of amino acids, enabling it to form a complex with sgRNA for precise gene editing, and it can perform gene editing in a eukaryotic environment. More specifically, gene editing tools containing the Cas12b mutant protein can recognize very simple PAMs, i.e., WNTs, such as TTG and ATN, and all exhibit high editing efficiency and specificity. Furthermore, due to the small molecular weight of the protein, it can be easily packaged into vector tools such as adeno-associated viruses, making it highly suitable for later development as a gene therapy tool. This invention expands the scope of gene editing and has broad application prospects in the field of gene editing.
[0122] Single-stranded guide RNA
[0123] In a fourth aspect, the present invention provides a single-stranded guide RNA (sgRNA) comprising a scaffold sequence and a CRISPR spacer sequence from the 5' end to the 3' end, the scaffold sequence comprising tracrRNA and a repeat sequence from the 5' end to the 3' end, and having a nucleic acid sequence that is truncated or mutated relative to the nucleic acid sequence shown in SEQ ID NO:2.
[0124] In this paper, the truncated nucleic acid sequences are obtained by truncating the 5' end sequence of the repeat sequence of sgRNA and the 3' end sequence of tracrRNA according to the bubble position, as shown in the figure. Figure 1 As shown, this is caused by base mismatches, including, for example, the mismatch between the 70th "U" (on the tracrRNA sequence) and the 118th "U" (on the repeat sequence) of the sgRNA, the mismatch between the 75th "U" (on the tracrRNA sequence) and the 113th "C" (on the repeat sequence) of the sgRNA, and the mismatch between the 78th "A" (on the tracrRNA sequence) and the 110th "C" (on the repeat sequence) of the sgRNA. Therefore, truncation can occur after the 70th "U" and before the 118th "U" of the sgRNA, i.e., a deletion of 49 nucleotides; it can occur after the 75th "U" and before the 113th "C" of the sgRNA, i.e., a deletion of 39 nucleotides; or it can occur after the 78th "A" and before the 110th "C" of the sgRNA, i.e., a deletion of 33 nucleotides. Of course, any other number of nucleotide deletions are also possible, such as 34, 35, 36, 37, 38, 40, 41, 42, 43, 44, 45, 46, 47 or 48 nucleotide deletions.
[0125] In a preferred embodiment, the truncated nucleic acid sequence is a nucleic acid sequence in which nucleotides 33-49 of positions 70-118 of SEQ ID NO:2 are missing.
[0126] In a preferred embodiment, the truncated nucleic acid sequence is the nucleic acid sequence shown in SEQ ID NO:3 or SEQ ID NO:4.
[0127] In yet another specific embodiment, the mutated nucleic acid sequence is a nucleic acid sequence that mutates one or more nucleotides of the tracrRNA at at least one bubble position formed by the tracrRNA of the nucleic acid sequence shown in SEQ ID NO:2 and the repeating sequence.
[0128] In a preferred embodiment, the mutated nucleotide is complementary to the nucleotide at the corresponding position of the repeating sequence.
[0129] The phrase "one or more nucleotides of the tracrRNA at at least one bubble location formed by the tracrRNA and the repeating sequence" is similar to... Figure 1 As shown, for example, "U" at position 70, "U" at position 75, and "A" at position 78. Therefore, in a more preferred embodiment, the mutated nucleic acid sequence contains one or more mutations corresponding to nucleotides 70, 75, and 78 of the nucleic acid sequence shown in SEQ ID NO:2.
[0130] In a more preferred embodiment, the mutation is one or more mutations selected from sgRNA 70U>A, sgRNA 75U>G, and sgRNA78A>G.
[0131] In a further preferred embodiment, the mutated nucleic acid sequence has at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even higher (e.g., 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 99.99%) sequence identity with SEQ ID NO:2.
[0132] In yet another specific implementation, the CRISPR spacer sequence is a sequence of 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28 nucleotides in length that is complementary to the target sequence.
[0133] In a preferred embodiment, the CRISPR spacer sequence is a 24-nucleotide sequence that is complementary to the target sequence.
[0134] In a further embodiment, the single-stranded guide RNA further includes a terminator at the 5' end of the spacer sequence. As an example, the terminator may be a plurality of terminators, such as at least six (e.g., seven or eight) U.
[0135] The single-stranded guide RNA can bind to wild-type or mutant ChCas12b protein, conjugates, or fusion proteins to form a complex. This complex can recognize the corresponding PAM and thereby bind to the target sequence, thus achieving the cleavage of the target sequence or gene editing.
[0136] Nucleic acid encoding and vectors
[0137] In a fifth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding the following:
[0138] a) The Cas12b mutant protein described in the first aspect;
[0139] b) The conjugate described in the second aspect; or
[0140] c) The fusion protein described in the third aspect.
[0141] In one specific implementation, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a single-stranded guide RNA for the fourth aspect.
[0142] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA of the fourth aspect.
[0143] After the isolated nucleic acid molecules of the present invention are transfected into the corresponding cells using certain tools known in the art, such as expression vectors, the isolated nucleic acid molecules of the present invention can express the Cas12b mutant protein, its conjugates or fusion proteins, and / or the single-stranded guide RNA described above, and perform the corresponding functions, such as gene editing.
[0144] In addition, the isolated nucleic acid molecules of the present invention can express Cas12b protein, its conjugates or fusion proteins, and single-stranded guide RNA individually or separately, or they can express the expression products together, depending on the specific circumstances.
[0145] Furthermore, the expressed product has the corresponding effects and / or functions described above, which will not be repeated here for the sake of brevity.
[0146] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding the following:
[0147] a) The Cas12b mutant protein described in the first aspect;
[0148] b) The conjugate described in the second aspect; or
[0149] c) The fusion protein described in the third aspect.
[0150] In one specific implementation, the vector can be an expression vector, such as a plasmid vector like pUC19 vector, an applicator vector, a pAAV2_ITR vector, a retroviral vector, a lentiviral vector, an adenovirus vector, or an adeno-associated virus vector.
[0151] In yet another specific embodiment, the vector further comprises a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect of the present invention.
[0152] In an eighth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect.
[0153] As described above, after the vector of the present invention is transfected into cells, the coding sequence cloned in the vector can be expressed as the Cas12b protein, its conjugates or fusion proteins, and / or the single-stranded guide RNA described above, and perform corresponding functions therein. For example, gene editing.
[0154] Alternatively, multiple vectors, such as two vectors, can be transfected into cells. One vector expresses the Cas12b protein, its conjugates, or fusion proteins, while the other vector expresses single-stranded guide RNA. Subsequently, the expressed Cas12b protein, its conjugates, or fusion proteins combine with the expressed single-stranded guide RNA to form a complex, which then performs its corresponding function, such as gene editing.
[0155] Alternatively, the nucleic acid sequence encoding the Cas12b protein, its conjugates or fusion proteins, and the nucleic acid sequence encoding the single-stranded guide RNA can be cloned into a vector, so that after the vector is transfected into cells, it expresses both the Cas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA, and performs the corresponding functions, such as gene editing.
[0156] CRISPR / Cas12b gene editing system
[0157] In a ninth aspect, the present invention provides a CRISPR / Cas12b gene editing system comprising:
[0158] 1) Protein components, which include:
[0159] a) The Cas12b mutant protein described in the first aspect;
[0160] b) The conjugate described in the second aspect; or
[0161] c) The fusion protein described in the third aspect;
[0162] d) Wild-type ChCas12b protein, its conjugates, or fusion proteins;
[0163] 2) The single-stranded guide RNA described in the fourth aspect, or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 2;
[0164] The protein component described herein, when it is d) wild-type ChCas12b protein, its conjugate or fusion protein, may only be used in conjunction with the single-stranded guide RNA described in the fourth aspect;
[0165] Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex.
[0166] The CRISPR / Cas12b gene editing system of the present invention can be directly constructed from the Cas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA described herein, or it can be constructed from the expression products obtained by expressing isolated nucleic acid molecules or vectors as described herein. The CRISPR / Cas12b gene editing system of the present invention achieves the recognition, localization, cleavage, and gene editing of target sequences through the combined action of the Cas12b protein and the single-stranded guide RNA contained therein.
[0167] The CRISPR / Cas12b gene editing system of this invention can precisely locate the target sequence. "Precisely located" has two meanings: first, the CRISPR / Cas12b gene editing system itself can recognize and bind to the target sequence; second, the CRISPR / Cas12b gene editing system can bring other proteins fused with the Cas12b protein or proteins that specifically recognize the sgRNA to the location of the target sequence.
[0168] The CRISPR / Cas12b gene editing system of the present invention has low tolerance for non-target sequences. In this document, "low tolerance" means that the CRISPR / Cas12b gene editing system of the present invention is substantially or completely unable to recognize and bind to non-target sequences, or substantially or completely unable to bring other proteins fused with the Cas12b protein or proteins that specifically recognize the sgRNA to the location of the non-target sequence.
[0169] The CRISPR / Cas12b gene editing system of the present invention can target more DNA sequences in the genome because the PAM sequence on the target sequence recognized by the Cas12b protein contained therein is simpler.
[0170] More importantly, as mentioned above, the CRISPR / Cas12b gene editing system of the present invention overcomes the deficiency of existing CRISPR / Cas systems in being unable to distinguish single-base mutations. It can distinguish single-base differences at target sites, thereby enabling the knockout of mutated alleles. This provides an important foundation for the clinical gene therapy using this type of gene editing system.
[0171] cell
[0172] In a tenth aspect, the present invention provides a cell comprising: isolated nucleic acid molecules as described in the fifth or sixth aspect, or a carrier as described in the seventh or eighth aspect.
[0173] As an example, the cell can be a prokaryotic cell or a eukaryotic cell. For the eukaryotic cell, as an example, it can be a plant cell or an animal cell. For the animal cell, as an example, it can be a mammalian cell, such as a human cell.
[0174] method
[0175] In an eleventh aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (7) with the target sequence in the intracellular or in vitro environment:
[0176] (1) The Cas12b mutant protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and single-stranded guide RNA;
[0177] (2) Wild-type ChCas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA described in the fourth aspect;
[0178] (3) The isolated nucleic acid molecules described in the fifth aspect and, as appropriate, isolated nucleic acid molecules containing nucleic acid sequences encoding single-stranded guide RNA;
[0179] (4) An isolated nucleic acid molecule comprising encoding wild-type ChCas12b protein, its conjugate or fusion protein, and the isolated nucleic acid molecule described in the sixth aspect;
[0180] (5) The vectors described in the seventh aspect and, as appropriate, vectors containing nucleic acid sequences encoding single-stranded guide RNA;
[0181] (6) A vector comprising a nucleic acid sequence encoding the wild-type ChCas12b protein, its conjugates, or a fusion protein, and the vector described in aspect eight; and
[0182] (7) The CRISPR / Cas12b gene editing system described in the ninth aspect;
[0183] The Cas12b protein, the conjugate, or the fusion protein recognizes a protospacer adjacent sequence (PAM) located at the 5' end of the target sequence and having a 5'-WTN sequence.
[0184] In one specific implementation, the single-stranded guide RNA of item (1) may be the single-stranded guide RNA described in the fourth aspect or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 2.
[0185] Regarding items (3) and (5) above, as can be understood from the above description, the isolated nucleic acid molecule described in the fifth aspect of the present invention and the vector described in the seventh aspect may, in some cases, contain only a nucleic acid sequence encoding the Cas12b mutant protein or its conjugate or fusion protein, and in other cases, may contain a nucleic acid sequence encoding the Cas12b mutant protein or its conjugate or fusion protein and a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect. Therefore, if the isolated nucleic acid molecule or the vector does not contain a nucleic acid sequence encoding a single-stranded guide RNA, it is necessary to add another isolated nucleic acid molecule or vector containing a nucleic acid sequence encoding a single-stranded guide RNA. Furthermore, it can be understood that the single-stranded guide RNA mentioned in the description "isolated nucleic acid molecule or vector containing a nucleic acid sequence encoding a single-stranded guide RNA" can be either a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO:2 or a single-stranded guide RNA variant of the fourth aspect of the present invention.
[0186] In one specific implementation, the cell is a prokaryotic cell or a eukaryotic cell, the eukaryotic cell being, for example, a plant cell or an animal cell, the animal cell being, for example, a mammalian cell such as a human cell.
[0187] In yet another specific implementation, the gene editing includes one or more of the following: gene knockout of a target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single-base conversion, and chromatin imaging tracking. For example, the single-base conversion includes adenine to guanine, cytosine to thymine, or cytosine to uracil.
[0188] In yet another specific implementation, in the method, the CRISPR spacer sequence of the single-stranded guide RNA forms a fully complementary base pairing structure with the target sequence, and an incomplete complementary base pairing structure with the non-target sequence.
[0189] In this document, the incomplete base pairing structure refers to a structure that includes a portion of base pairing and a portion of non-base pairing, wherein the non-base pairing includes, for example, base mismatch and / or base bulge.
[0190] In a further specific embodiment, the incomplete base complementary pairing structure includes one or more, for example, two or more base mismatches.
[0191] Therefore, the Cas12b mutant protein of the present invention can cleave the target site on the target sequence, and under the cleavage action of the Cas12b mutant protein, a double-strand break occurs in the target sequence. Furthermore, when the method is performed intracellularly, the cleaved target sequence can be repaired through intracellular non-homologous end joining repair or homologous recombination repair pathways, thereby achieving gene editing of the target sequence.
[0192] In the CRISPR / Cas12b gene editing system and gene editing method of the present invention, the Cas12b mutant protein of the present invention has been experimentally found to form a complex with sgRNA for gene editing, and its mismatch-containing guide RNA has a near 0% error tolerance. Therefore, these gene editing systems can edit target genes with high specificity, and have the characteristics of high editing efficiency and low off-target rate, and can be widely used in gene editing in cells or in vitro environments.
[0193] Reagent test kit
[0194] In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising:
[0195] a) Choose any one of (1) to (7) below:
[0196] (1) The Cas12b protein described in the first aspect, the conjugate described in the second aspect, or the fusion protein described in the third aspect, and single-stranded guide RNA;
[0197] (2) Wild-type ChCas12b protein, its conjugates or fusion proteins, and the single-stranded guide RNA described in the fourth aspect;
[0198] (3) The isolated nucleic acid molecules described in the fifth aspect and, as appropriate, isolated nucleic acid molecules containing nucleic acid sequences encoding single-stranded guide RNA;
[0199] (4) An isolated nucleic acid molecule comprising encoding wild-type ChCas12b protein, its conjugate or fusion protein, and the isolated nucleic acid molecule described in the sixth aspect;
[0200] (5) The vectors described in the seventh aspect and, as appropriate, vectors containing nucleic acid sequences encoding single-stranded guide RNA;
[0201] (6) A vector comprising a nucleic acid sequence encoding the wild-type ChCas12b protein, its conjugates, or a fusion protein, and the vector described in aspect eight; and
[0202] (7) The CRISPR / Cas12b gene editing system described in aspect nine; and
[0203] b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment.
[0204] Of course, those skilled in the art will understand that the kit of the present invention may also contain other reagents that facilitate gene editing.
[0205] In one specific implementation, the single-stranded guide RNA of item (1) may be the single-stranded guide RNA described in the fourth aspect or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 2.
[0206] Regarding items (3) and (5) above, as can be understood from the above description, the isolated nucleic acid molecule described in the fifth aspect of the present invention and the vector described in the seventh aspect may, in some cases, contain only a nucleic acid sequence encoding a Cas12b mutant protein or its conjugate or fusion protein, and in other cases, contain a nucleic acid sequence encoding a Cas12a mutant protein or its conjugate or fusion protein and a nucleic acid sequence encoding the single-stranded guide RNA described in the fourth aspect. Therefore, if the isolated nucleic acid molecule or the vector does not contain a nucleic acid sequence encoding a single-stranded guide RNA, it is necessary to add another isolated nucleic acid molecule or vector containing a nucleic acid sequence encoding a single-stranded guide RNA. Furthermore, it can be understood that the single-stranded guide RNA mentioned in the description "isolated nucleic acid molecule or vector containing a nucleic acid sequence encoding a single-stranded guide RNA" can be either a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO:2 or a single-stranded guide RNA variant of the fourth aspect of the present invention.
[0207] Of course, those skilled in the art will understand that the kit of the present invention may also contain other reagents that facilitate gene editing.
[0208] A brief description of the sequence involved in this invention.
[0209] SEQ ID NO:1: Wild-type ChCas12b protein sequence;
[0210] SEQ ID NO:2: Wild-type ChCas12b-sgRNA scaffold sequence;
[0211] SEQ ID NO:3: ChCas12b-sgRNA scaffold sequence variant 1; and
[0212] SEQ ID NO:4: ChCas12b-sgRNA scaffold sequence variant 2.
[0213] Example
[0214] The invention will now be described with reference to the following embodiments, which are intended to be illustrative and not limiting. Those skilled in the art will understand that the embodiments provided herein are for the purpose of describing the invention in detail only and are not intended to limit the scope of protection claimed by the invention.
[0215] Unless otherwise specified, the experiments and methods described in the examples were generally performed according to conventional methods well known in the art and described in the various references. Furthermore, for conditions not specifically specified in the examples, conventional conditions or conditions recommended by the manufacturer were followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0216] Example 1:
[0217] (1) Construction of ChCas12b point mutant plasmid
[0218] Circular PCR was performed using pAAV-CMV-Hsp2Cas9 plasmid (Addgene platform, catalog #192126) as a template. Primer sequences are shown in Table 1 below:
[0219] Table 1: PCR primers for constructing the ChCas12b point mutant
[0220]
[0221] The reaction system is as follows:
[0222]
[0223] The PCR procedure is as follows:
[0224]
[0225] PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel extraction kit following the manufacturer's instructions. DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). T4 PNK and T4 DNA ligase treatments were then performed. The reaction system is as follows:
[0226]
[0227] The reaction conditions are as follows:
[0228]
[0229] Add 1 μL of T4 DNA ligase (NEB) to the reaction system, vortex to mix, and incubate at room temperature for 2 h. Add the ligation product to E. coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubate on ice for 30 min, heat shock at 42℃ for 1 min, incubate on ice for 2 min, add 900 μL of LB medium, and incubate at 37℃ for 1 h to activate and revive E. coli DH5α competent cells.
[0230] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0231] Sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmids were extracted to obtain the point mutants ChCas12b-R453A, ChCas12b-D496A, and ChCas12b-R1137A plasmids, which can be used for later use or stored at -20℃ for long-term storage.
[0232] (2) Plasmid hU6 - OQB30769_tracr - Linearization preparation of BsaI
[0233] The hU6-Sa_tracr plasmid (Addgene platform, catalog #135973) was digested with BsaI and NotI restriction endonucleases. The digestion system consisted of 1 μg of plasmid hU6-Sa_tracr, 5 μL of 10×CutSmart buffer (NEB), 1 μL of BsaI and 1 μL of NotI restriction endonuclease (NEB), with water added to a final volume of 50 μL. The digestion was incubated at 37°C for 3 hours.
[0234] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.
[0235] A 2981bp DNA fragment was excised from an agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragment was eluted with ultrapure water.
[0236] The sgRNA scaffold sequence (SEQ ID NO: 2) was synthesized and constructed on the linearized hU6-Sa_tracr backbone to obtain the plasmid hU6-OQB30769_tracr-BsaI.
[0237] The hU6-OQB30769_tracr-BsaI plasmid was digested with the BsaI restriction endonuclease (NEB). The reaction mixture consisted of 2 μg of plasmid hU6-OQB30769_tracr-BsaI, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of BsaI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.
[0238] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0239] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.
[0240] The DNA concentration of the recovered linearized fragment hU6-OQB30769_tracr-BsaI was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20°C for long-term preservation.
[0241] (3) Plasmid hU6 - OQB30769_tracr - BsaI - Construction of on / off target gRNA
[0242] The sequences of on / off target gRNAs were designed, and their corresponding oligonucleotide single-stranded DNAs are shown in Table 2 below.
[0243] The oligonucleotide single-stranded DNA corresponding to the obtained on / off target gRNA was annealed. The annealing reaction system consisted of 1 μL 100 μM oligo-F, 1 μL 100 μM oligo-R, and 28 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the resulting product was ligated into the obtained linearized hU6-OQB30769_tracr-BsaI plasmid using DNA ligase (purchased from NEB).
[0244] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were cultured at 37℃ for 1 h to activate and revitalize Escherichia coli DH5α competent cells.
[0245] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0246] After sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmid was extracted to obtain the plasmid hU6-OQB30769_tracr-BsaI-on / off target gRNA expressing the above on / off target gRNA sequence, which can be used for later use.
[0247] (4) The plasmids ChCas12b-R453A, ChCas12b-D496A and ChCas12b-R1137A expressing the Cas12b mutant protein were co-transfected with hU6-OQB30769_tracr-BsaI-on / off target gRNA into the HEK293T cell line library containing the target sequence (GGATATGTTGAAGAACACCATGAC) using liposomes.
[0248] The HEK293T cell line containing the target sequence of the GFP reporter system was obtained as follows: a PAM sequence and a specific target sequence were inserted between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. This mutation was then integrated into HEK293T cells via lentiviral infection, resulting in the HEK293T cell line containing the target sequence of the GFP reporter system. After the gene editing system cuts the target sequence, the cells' self-repair system causes some cells to recover the GFP reading frame, producing green fluorescence. Flow cytometry analysis of the GFP-positive cell ratio can be used to assess the editing capability and specificity of the gene editing system.
[0249]
[0250] The above transfection process includes the following steps:
[0251] On day 0, as required for transfection, the HEK293T cell line containing the target sequence of the GFP reporter system was seeded in 48-well plates at a cell density of 30%. This HEK293T cell line containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-PAM-target site-GFP, where the PAM sequence is TTG and the target site sequence is GGATATGTTGAAGAACACCATGAC.
[0252] On day 1, transfection was performed as follows: 0.5 μg of plasmids ChCas12b-R453A, ChCas12b-D496A, and ChCas12b-R1137A expressing the ChCas12b point mutant protein were taken and added together with 0.3 μg of hU6-OQB30769_tracr-BsaI-on / off target gRNA into 17.5 μL of Opti-MEM medium (purchased from Gibco) and gently pipetted to mix.
[0253] Gently mix polyethyleneimine (PEI, purchased from Polysciences), add 0.8 μL of PEI to 17.5 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0254] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 15 min, and then added to the culture medium of the HEK293T cell line containing the target sequence of the GFP reporter system. The cell line was then placed in a 37°C, 5% CO2 incubator for further culture.
[0255] Flow cytometry was used to analyze the editing efficiency of the ChCas12b point mutant on the target sequence. Specifically, HEK293T cells were collected after 5 days of culture in a CO2 incubator, and their specificity was detected using a flow cytometer (BD Biosciences FACSCalibur). The GFP positivity rate was analyzed and plotted using FlowJo analysis software.
[0256] The editing efficiency of the ChCas12b point mutant in the HEK293T cell line containing the target sequence of the GFP reporter system is shown in the figures below. Figure 3-5 When the ChCas12b point mutant cleaves the target sequence, the cell's own repair system allows some cells to recover the GFP reading frame, producing green fluorescence. Figure 3-5 In the graph, the Y-axis represents the percentage of GFP-positive cells (%), the X-axis represents on / off target gRNA, and NC represents the negative control group (no transfected plasmid). Figure 3-5 As can be seen, the ChCas12b point mutant edited all target sites in the HEK293T cell line of the GFP reporter system. The on-target editing efficiency of the ChCas12b point mutant was comparable to that of wild-type ChCas12b, but the off-target editing efficiency was significantly reduced at all sites.
[0257] Example 2
[0258] (1)ChCas12b - sgRNA scaffold - Construction of variant mutant plasmids
[0259] The wild-type ChCas12b-sgRNA scaffold (SEQ ID NO:2) consists of a full-length direct repeat sequence and a full-length tracrRNA from the bacterial genome. Analysis using the RNAfold website yielded a schematic diagram of the sgRNA secondary structure. To further optimize the sgRNA scaffold, we constructed five sgRNA variants (variants 1-5) based on their secondary structure. The truncated variants 1 and 2 are created by combining the 5' end sequence of the direct repeat sequence and the 3' end sequence of the tracrRNA according to the bubble position (e.g., ...). Figure 1 The portions at V1 and V2 shown in the diagram have been appropriately truncated. Variants 3, 4, and 5 are at partial bubble positions (such as...). Figure 1 The bases on the tracrRNA are mutated at positions V3, V4, and V5 (shown in the diagram) to make them complementary to the direct repeat sequence. The nucleic acid sequences of these five sgRNA scaffold-variants (including their respective tracrRNA sequences and direct repeat sequences) are as follows: Figure 2 As shown.
[0260] Using the hU6-OQB30769_tracr-BsaI plasmid as a template, circular PCR was performed to obtain five ChCas12b-sgRNA scaffold-variants. The circular PCR primer sequences are shown in Table 3 below.
[0261] Table 3: PCR primers for constructing the ChCas12b-sgRNA scaffold-variant mutant
[0262]
[0263] The reaction system is as follows:
[0264]
[0265] The PCR procedure is as follows:
[0266]
[0267] PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel extraction kit following the manufacturer's instructions. DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific). T4 PNK and T4 DNA ligase treatments were then performed. The reaction system is as follows:
[0268]
[0269] The reaction conditions are as follows:
[0270]
[0271] Add 1 μL of T4 DNA ligase (NEB) to the reaction system, vortex to mix, and incubate at room temperature for 2 h.
[0272] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.
[0273] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0274] Sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmid was extracted to obtain the mutant ChCas12b-sgRNA scaffold-variant 1-5 plasmid, which can be used for later use or stored at -20℃ for long-term preservation.
[0275] (2) plasmid ChCas12b - sgRNA scaffold - Linearization preparation of variant
[0276] ChCas12b-sgRNA scaffold-variant plasmids 1-5 were digested with BsaI restriction endonuclease (NEB). The reaction system consisted of 2 μg of ChCas12b-sgRNA scaffold-variant plasmid, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of BsaI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.
[0277] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0278] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.
[0279] The DNA concentration of the recovered linearized ChCas12b-sgRNA scaffold-variant was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.
[0280] (3) plasmid ChCas12b - sgRNA scaffold - variant - Construction of target gRNA
[0281] The sequence of the on-target gRNA was designed, and its corresponding oligonucleotide single-stranded DNA is shown in Table 4 below:
[0282] Table 4: Oligonucleotide single-stranded DNA on target gRNA
[0283]
[0284] The obtained oligonucleotide single-stranded DNA corresponding to the on-target gRNA was annealed. The annealing reaction system consisted of 1 μL 100 μM oligo-F, 1 μL 100 μM oligo-R, and 28 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the product was ligated into the obtained linearized ChCas12b-sgRNAscaffold-variant 1-5 plasmid using DNA ligase (purchased from NEB).
[0285] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were cultured at 37℃ for 1 h to activate and reactivate Escherichia coli DH5α competent cells.
[0286] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0287] After sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmid was extracted to obtain the plasmid ChCas12b-sgRNA scaffold-variant-on target gRNA expressing the above-mentioned ontarget gRNA sequence, which was then ready for use.
[0288] The obtained plasmids expressing the crRNA mutant, namely ChCas12b-sgRNA scaffold-variant 1-on target gRNA, ChCas12b-sgRNA scaffold-variant 2-on target gRNA, ChCas12b-sgRNA scaffold-variant 3-on target gRNA, ChCas12b-sgRNA scaffold-variant 4-on target gRNA, and ChCas12b-sgRNA scaffold-variant 5-on target gRNA, were co-transfected with plasmids expressing wild-type ChCas12b protein into a GFP reporter system HEK293T cell line library containing the target sequence (GGATATGTTGAAGAACACCATGAC) using liposomes.
[0289] The HEK293T cell line containing the target sequence of the GFP reporter system was obtained as follows: a PAM sequence and a specific target sequence were inserted between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. This mutation was then integrated into HEK293T cells via lentiviral infection, resulting in the HEK293T cell line containing the target sequence of the GFP reporter system. After the gene editing system cuts the target sequence, the cells' self-repair system causes some cells to recover the GFP reading frame, producing green fluorescence. Flow cytometry analysis of the GFP-positive cell ratio can be used to assess the editing capability and specificity of the gene editing system.
[0290] The above transfection process includes the following steps:
[0291] On day 0, the HEK293T cell line containing the target sequence of the GFP reporter system was seeded in 48-well plates as required for transfection, with the cell density controlled at 30%.
[0292] The HEK293T cell line containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-PAM-target site-GFP, where the PAM sequence is TTG and the target site sequence is GGATATGTTGAAGAACACCATGAC.
[0293] Day 1, transfection was performed. The transfection process is as follows:
[0294] Take 0.3 μg of each of the following ChCas12b-sgRNA scaffold-variant 1-on target gRNA, ChCas12b-sgRNA scaffold-variant 2-on target gRNA, ChCas12b-sgRNA scaffold-variant 3-on target gRNA, ChCas12b-sgRNA scaffold-variant 4-on target gRNA, and ChCas12b-sgRNA scaffold-variant 5-on target gRNA, and add them together with 0.5 μg of plasmid expressing wild-type ChCas12b protein into 17.5 μL of Opti-MEM medium (purchased from Gibco). Mix gently by pipetting.
[0295] Gently mix the PEI (purchased from Polysciences), add 0.8 μL of PEI to 17.5 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.
[0296] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 15 min, and then added to the culture medium of the HEK293T cell line containing the target sequence of the GFP reporter system. The cell line was then placed in a 37°C, 5% CO2 incubator for further culture.
[0297] Flow cytometry was used to analyze the editing efficiency of wild-type ChCas12b protein on target sequences.
[0298] Specifically, HEK293T cell lines cultured in a CO2 incubator for 5 days were collected, and their specificity was detected using a flow cytometer (BDBiosciences FACSCalibur). The GFP positivity rate was analyzed and plotted using FlowJo analysis software.
[0299] The editing efficiency of the ChCas12b-sgRNA scaffold-variant mutant in the HEK293T cell line containing the target sequence of the GFP reporter system is shown in the figure. Figure 6 When wild-type ChCas12b protein and the ChCas12b-sgRNA scaffold-variant mutant cleave the target sequence, the cell's own repair system can restore the GFP reading frame in some cells, producing green fluorescence. Figure 6The Y-axis represents the percentage of GFP-positive cells (%), and the X-axis represents the ChCas12b-sgRNAscaffold-WT and ChCas12b-sgRNAscaffold-variant mutants. From... Figure 6 As can be seen, the ChCas12b-sgRNA scaffold-variant mutant edited all target sites in the HEK293T cell line of the GFP reporter system, and the editing efficiency of ChCas12b-sgRNA scaffold-variant 2 and ChCas12b-sgRNA scaffold-variant 5 was higher than that of wild-type sgRNA.
[0300] The above experiments show that, compared with the gene editing system composed of wild-type ChCas12b protein and wild-type ChCas12b-sgRNA, the gene editing system containing wild-type ChCas12b protein and the sgRNA variant of the present invention exhibits higher editing efficiency. Figure 6 ).
[0301] Example 3:
[0302] (1) Plasmid hU6 - OQB30769_tracr - BsaI, ChCas12b - sgRNA scaffold - variant 2 and ChCas12b - sgRNA scaffold - Linearization preparation of variant 5
[0303] The hU6-OQB30769_tracr-BsaI, ChCas12b-sgRNAscaffold-variant 2, and ChCas12b-sgRNAscaffold-variant 5 plasmids were digested using the BsaI restriction endonuclease (NEB). The reaction system consisted of: 2 μg plasmid (hU6-OQB30769_tracr-BsaI, ChCas12b-sgRNAscaffold-variant 2, or ChCas12b-sgRNAscaffold-variant 5), 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of BsaI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.
[0304] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0305] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.
[0306] The DNA concentrations of the recovered linearized fragments hU6-OQB30769_tracr-BsaI, ChCas12b-sgRNA scaffold-variant 2, and ChCas12b-sgRNA scaffold-variant 5 were determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and were either kept for later use or stored at -20°C for long-term preservation.
[0307] (2) Plasmid hU6 - OQB30769_tracr - BsaI - gRNA, ChCas12b - sgRNA scaffold - variant 2 - gRNA and ChCas12b - sgRNA scaffold - variant 5 - gRNA construction
[0308] The gRNA sequences were designed, and their corresponding oligonucleotide single-stranded DNA sequences are shown in Table 5 below.
[0309] The oligonucleotide single-stranded DNA corresponding to the obtained gRNA was annealed. The annealing reaction system consisted of 1 μL 100 μM Moligo-F, 1 μL 100 μM oligo-R, and 28 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained products were ligated into the linearized hU6-OQB30769_tracr-BsaI, ChCas12b-sgRNA scaffold-variant 2, and ChCas12b-sgRNA scaffold-variant 5 plasmids using DNA ligase (purchased from NEB).
[0310] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were cultured at 37℃ for 1 h to activate and revitalize Escherichia coli DH5α competent cells.
[0311] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0312]
[0313] After sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmids were extracted to obtain plasmids hU6-OQB30769_tracr-BsaI-gRNA, ChCas12b-sgRNA scaffold-variant 2-gRNA, and ChCas12b-sgRNA scaffold-variant 5-gRNA expressing the above gRNA sequences, which were then ready for use.
[0314] (3 - 1) Plasmids expressing wild-type ChCas12b protein and plasmids expressing sgRNA variants ChCas12b - sgRNA scaffold - variant 2 - gRNA and ChCas12b - sgRNA scaffold - variant 5 - gRNA on HEK293T cells Transfection of the system
[0315] On day 0, HEK293T cells containing the target sequence were seeded in 24-well plates as needed for transfection, with a cell density of approximately 30%.
[0316] Day 1, transfection was performed. The transfection process is as follows:
[0317] Take 0.5 μg of plasmid expressing wild-type ChCas12b protein, and add it together with 0.3 μg of plasmid hU6-OQB30769_tracr-BsaI-sgRNA (as a control), ChCas12b-sgRNA scaffold-variant 2-gRNA or ChCas12b-sgRNA scaffold-variant 5-gRNA into 25 μL of Opti-MEM medium (purchased from Gibco), and gently pipette to mix.
[0318] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI) (purchased from Polysciences). Add 1.6 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0319] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 5 days.
[0320] (3 - 2) Plasmids expressing the ChCas12b point mutant protein and plasmids expressing the sgRNA variant ChCas12b - sgRNA scaffold - variant 2 - gRNA transfection of HEK293T cell line
[0321] On day 0, HEK293T cells containing the target sequence were seeded in 24-well plates as needed for transfection, with a cell density of approximately 30%.
[0322] Day 1, transfection was performed. The transfection process is as follows:
[0323] Take 0.5 μg of plasmid expressing wild-type ChCas12b protein, plasmids expressing ChCas12b point mutant protein (ChCas12b-R453A, ChCas12b-D496A, or ChCas12b-R1137A), and add them together with 0.3 μg of plasmid ChCas12b-sgRNAscaffold-variant 2-gRNA to 25 μL of Opti-MEM medium (purchased from Gibco) and gently pipette to mix.
[0324] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI) (purchased from Polysciences). Add 1.6 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0325] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 5 days.
[0326] (3 - 3) Plasmids expressing the ChCas12b point mutant protein and plasmid hU6 expressing wild-type sgRNA. - OQB30769_tracr - BsaI - gRNA transfection of HEK293T cell line
[0327] On day 0, HEK293T cells containing the target sequence were seeded in 24-well plates as needed for transfection, with a cell density of approximately 30%.
[0328] Day 1, transfection was performed. The transfection process is as follows:
[0329] Take 0.5 μg of plasmid expressing wild-type ChCas12b protein (as a control), and plasmids expressing ChCas12b point mutant protein (ChCas12b-R453A, ChCas12b-D496A, or ChCas12b-R1137A), and add them together with 0.3 μg of plasmid hU6-OQB30769_tracr-BsaI-gRNA into 25 μL of Opti-MEM medium (purchased from Gibco), and gently pipette to mix.
[0330] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI) (purchased from Polysciences). Add 1.6 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0331] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 5 days.
[0332] (4) Preparation of next-generation sequencing libraries
[0333] HEK293T cells were collected 5 days after editing, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.
[0334] The first round of PCR for library preparation was performed using 2×Q5 Master mix. The PCR primers are shown below:
[0335] Table 6: List of primers for first-round PCR in next-generation sequencing
[0336]
[0337] The reaction system is as follows:
[0338]
[0339] The PCR procedure is as follows:
[0340]
[0341] For the second round of PCR to prepare the sequencing library, 2xQ5 Master mix was used for the PCR reaction. The PCR primers are shown below:
[0342] F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC;
[0343] R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0344] The reaction system is as follows:
[0345]
[0346] The PCR procedure is as follows:
[0347]
[0348] The second-round PCR products were purified using a gel extraction kit according to the manufacturer's instructions, resulting in a 300-400 bp DNA fragment. This completed the preparation of the next-generation sequencing library.
[0349] (5) Analysis of second-generation sequencing results
[0350] The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (Illumina).
[0351] Next-generation sequencing calculations yielded the editing efficiency results of multiple gene editing systems for multiple target sites obtained in step (3) above, such as... Figures 7 to 9 As shown, the X-axis represents the target site, and the Y-axis represents the editing efficiency (Indels%). Figure 7The results show the editing efficiency of the gene editing system containing wild-type ChCas12b protein and different sgRNA scaffold sequences (wild-type ChCas12b-sgRNA, ChCas12b-sgRNA scaffold sequence variant 2 or 5) obtained in step (3-1) above for three target sites D1, A1 and A3. The results show that for the three target sites D1, A1 and A3, sgRNA scaffold sequence variants 2 and 5 both show superior editing efficiency compared to wild-type sgRNA. Figure 8 The results show the editing efficiency of gene editing systems containing different Cas12b proteins (wild-type ChCas12b protein, ChCas12bR453A, ChCas12bD496A, or ChCas12bR1137A mutant protein) and ChCas12b-sgRNA scaffold sequence variant 2 on two endogenous target sites, VEGFA and EMX1. The results indicate that for the endogenous target sites VEGFA and EMX1, the ChCas12bR453A and ChCas12bD496A mutant proteins both exhibit superior editing efficiency compared to the ChCas12bR1137A mutant protein and the wild-type protein. In particular, for the endogenous target site EMX1, the editing efficiency of ChCas12bR453A and ChCas12bD496A is significantly better than that of the ChCas12bR1137A mutant protein and the wild-type protein. Figure 9 The results show the editing efficiency of gene editing systems containing different Cas12b proteins (wild-type ChCas12b protein, ChCas12bR453A, ChCas12bD496A, or ChCas12bR1137A mutant protein) and wild-type sgRNA at multiple target sites. The results show that for target sites A5, E1, E2, E3, and G1, there is no significant difference in editing efficiency between gene editing systems containing wild-type ChCas12b protein and its three mutant proteins. However, for target sites E4, E5, G2, and G4, some mutant proteins, especially ChCas12bR453A, show significantly better editing efficiency than wild-type proteins. Figure 10 Using each Cas12b protein as the X-axis, further summaries were made. Figure 9 The results for editing efficiency (Indels%) are shown.
[0352] Example 4:
[0353] (1) Plasmid ChCas12b - sgRNA scaffold - Linearization preparation of variant 2
[0354] The ChCas12b-sgRNA scaffold-variant 2 plasmid was digested using the BsaI restriction endonuclease (NEB). The reaction mixture consisted of 2 μg of ChCas12b-sgRNA scaffold-variant 2 plasmid, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of BsaI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.
[0355] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 min.
[0356] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.
[0357] The DNA concentration of the recovered linearized ChCas12b-sgRNA scaffold-variant 2 was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.
[0358] (2) plasmid ChCas12b - sgRNA scaffold - variant 2 - gRNA construction
[0359] The gRNA sequences were designed, and their corresponding oligonucleotide single-stranded DNA sequences are shown in Table 7 below.
[0360] The resulting oligonucleotide single-stranded DNA corresponding to the gRNA was annealed. The annealing reaction mixture consisted of 1 μL 100 μM Moligo-F, 1 μL 100 μM oligo-R, and 28 μL water. After vortexing and mixing the annealing mixture, it was placed in a PCR instrument and the annealing program was run as follows: 95℃ for 5 min, 85℃ for 1 min, 75℃ for 1 min, 65℃ for 1 min, 55℃ for 1 min, 45℃ for 1 min, 35℃ for 1 min, 25℃ for 1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the product was ligated into the linearized ChCas12b-sgRNA scaffold-variant 2 plasmid using DNA ligase (purchased from NEB).
[0361] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were cultured at 37℃ for 1 h to activate and revitalize Escherichia coli DH5α competent cells.
[0362] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.
[0363] After sequencing and verifying the correct ligation of the E. coli DH5α clone, the plasmid was extracted to obtain the plasmid ChCas12b-sgRNA scaffold-variant 2-gRNA expressing the above gRNA sequence, which can be used for later use.
[0364] Table 7: gRNA and its DNA sequence
[0365]
[0366] Note: The lowercase letters (i.e., the first four letters in Oligo-F and Oligo-R) indicate viscous end segments.
[0367] (3) Expression of ChCas12b - Plasmids for D496A protein and sgRNA (ChCas12b) - sgRNA scaffold - variant 2 - gRNA transfection of HEK293T cell line
[0368] On day 0, HEK293T cells containing the target sequence were seeded in 24-well plates as needed for transfection, with a cell density of approximately 30%.
[0369] Day 1, transfection was performed. The transfection process is as follows:
[0370] Take 0.5 μg of plasmid expressing ChCas12b-D496A protein and 0.3 μg of plasmid hU6-OQB30769_scaffold-variant 2-gRNA and add them together to 25 μL of Opti-MEM medium (purchased from Gibco). Gently pipette and mix well.
[0371] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI) (purchased from Polysciences). Add 1.6 μL of Lipofectamine® 2000 or PEI to 25 μL of Opti-MEM medium (purchased from Gibco), mix gently, and let stand at room temperature for 5 min.
[0372] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 5 days.
[0373] (4) Preparation of next-generation sequencing libraries
[0374] HEK293T cells were collected 5 days after editing, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.
[0375] The first round of PCR for library preparation was performed using 2×Q5 Master mix. The PCR primers are shown below:
[0376] Table 8: List of primers for first-round PCR in next-generation sequencing
[0377]
[0378] The reaction system is as follows:
[0379]
[0380] The PCR procedure is as follows:
[0381]
[0382] For the second round of PCR to prepare the sequencing library, 2xQ5 Master mix was used for the PCR reaction. The PCR primers are shown below:
[0383] F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC;
[0384] R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.
[0385] The reaction system is as follows:
[0386]
[0387] The PCR procedure is as follows:
[0388]
[0389] The second-round PCR products were purified using a gel extraction kit according to the manufacturer's instructions, resulting in a 300-400 bp DNA fragment. This completed the preparation of the next-generation sequencing library.
[0390] (5) Analysis of second-generation sequencing results
[0391] The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (Illumina).
[0392] Next-generation sequencing calculations yielded the efficiency of targeted editing applied to single nucleotide polymorphism (SNP) sites, such as... Figure 11 As shown, the X-axis represents the target site (On) and allele sites (Off), and the Y-axis represents the editing efficiency (Indels%). From... Figure 11 As can be seen, the ChCas12b-D496A mutant protein exhibits editing efficiency at the target SNP site, but lacks editing activity at another allele site, and the editing efficiency differs significantly between the two. This indicates that ChCas12b-D496A possesses high specificity and can distinguish single-base differences at the target site. The inventors also verified that other combinations of the Cas12b mutant protein and various sgRNA variants or wild-type sgRNAs obtained similar technical effects, namely, the ability to distinguish single-base differences at the target site.
Claims
1. A Cas12b mutant protein, said Cas12b mutant protein containing only a mutation at one amino acid residue corresponding to R453 and D496 of the wild-type ChCas12b protein as shown in SEQ ID NO:1: R453A or D496A.
2. A conjugate, said conjugate comprising: a) The Cas12b mutant protein of claim 1; b) Modified parts; and c) A connector for linking the Cas12b mutant protein to the modified portion; The modified portion is selected from other proteins or peptides, detectable markers, or combinations thereof; The additional protein or polypeptide mentioned therein is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
3. The conjugate according to claim 2, wherein the linker is a linker with a length of 1-50 amino acids.
4. A fusion protein, said fusion protein comprising: a) The Cas12b mutant protein of claim 1; b) Other proteins and peptides; and c) Connectors for linking the Cas12b protein to the other proteins and peptides; The additional protein or polypeptide mentioned therein is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminase (CBE), adenine deaminase (ABE), cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressor KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, and endonuclease FokI.
5. The fusion protein according to claim 4, wherein the linker is a linker with a length of 1-50 amino acids.
6. A single-stranded guide RNA, wherein the single-stranded guide RNA comprises a scaffold sequence and a CRISPR spacer sequence from the 5' end to the 3' end, the scaffold sequence comprising tracrRNA and a repeat sequence from the 5' end to the 3' end, and the single-stranded guide RNA is a nucleic acid sequence obtained by truncating or mutating based on SEQ ID NO:2; The truncated nucleic acid sequence is the nucleic acid sequence shown in SEQ ID NO:4; The mutation referred to here means that the 78th nucleotide of the nucleic acid sequence shown in SEQ ID NO:2 undergoes a mutation of A>G.
7. The single-stranded guide RNA according to claim 6, wherein the CRISPR spacer sequence is a sequence of 19-28 nucleotides in length that is complementary to the target sequence.
8. The single-stranded guide RNA of claim 6, wherein the CRISPR spacer sequence is a 24-nucleotide sequence that is complementary to the target sequence.
9. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the following: a) The Cas12b mutant protein of claim 1; b) The conjugate according to claim 2 or 3; or c) The fusion protein according to claim 4 or 5.
10. The isolated nucleic acid molecule according to claim 9 further comprises a nucleic acid sequence encoding a single-stranded guide RNA according to any one of claims 6-8 or a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO:
2.
11. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA as described in any one of claims 6-8.
12. A vector comprising a nucleic acid sequence encoding the following: a) The Cas12b mutant protein of claim 1; b) The conjugate according to claim 2 or 3; or c) The fusion protein according to claim 4 or 5.
13. The vector according to claim 12, wherein the vector is a plasmid vector.
14. The vector according to claim 12, wherein the vector is a pUC19 vector, an applicator vector, a pAAV2_ITR vector, a retroviral vector, a lentiviral vector, an adenovirus vector, or an adeno-associated virus vector.
15. The vector according to any one of claims 12 to 14 further comprises a nucleic acid sequence encoding a single-stranded guide RNA according to any one of claims 6 to 8 or a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO:
2.
16. A vector comprising a nucleic acid sequence encoding a single-stranded guide RNA as described in any one of claims 6-8.
17. A CRISPR / Cas12b gene editing system, comprising: 1) Protein components, which include: a) The Cas12b mutant protein of claim 1; b) The conjugate according to claim 2 or 3; or c) The fusion protein according to claim 4 or 5; d) Wild-type ChCas12b protein, its conjugates, or fusion proteins as shown in SEQ ID NO:1; 2) The single-stranded guide RNA according to any one of claims 6-8, or a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO: 2; The protein component described herein, when it is, as in the wild-type ChCas12b protein, its conjugate, or fusion protein as shown in SEQ ID NO:1, may only be used in conjunction with the single-stranded guide RNA of any one of claims 6-8; Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex.
18. A cell comprising the isolated nucleic acid molecule of any one of claims 9-11; or the vector of any one of claims 12-16.
19. The cell according to claim 18, wherein the cell is a prokaryotic cell or a eukaryotic cell, and the eukaryotic cell is an animal cell.
20. The cell of claim 19, wherein the animal cell is a mammalian cell.
21. A method for non-therapeutic gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: Contact any of the following (1) to (9) with the target sequence in the intracellular or in vitro environment: (1) The Cas12b mutant protein of claim 1, the conjugate of claim 2 or 3, or the fusion protein of claim 4 or 5, and the single-stranded guide RNA, wherein the single-stranded guide RNA is the single-stranded guide RNA of any one of claims 6-8 or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 2; (2) The wild-type ChCas12b protein, its conjugates or fusion proteins as shown in SEQ ID NO:1, and the single-stranded guide RNA of any one of claims 6-8; (3) The isolated nucleic acid molecule of claim 9, and the isolated nucleic acid molecule of claim 11, or an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO: 2; (4) An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the wild-type ChCas12b protein, its conjugate or fusion protein as shown in SEQ ID NO:1, and the isolated nucleic acid molecule as described in claim 11; (5) The isolated nucleic acid molecule as described in claim 10; (6) The vector of any one of claims 12 to 14, and the vector of claim 16 or a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO: 2; (7) A vector comprising encoding the wild-type ChCas12b protein as shown in SEQ ID NO:1, its conjugate or fusion protein, and the vector as described in claim 16; (8) The carrier according to claim 15; or (9) The CRISPR / Cas12b gene editing system as described in claim 17; The Cas12b protein, the conjugate, or the fusion protein recognizes a protospacer adjacent sequence (PAM) located at the 5' end of the target sequence and having a 5'-WTN sequence.
22. The method of claim 21, wherein the Cas12b protein, the conjugate, or the fusion protein recognizes a protospacer adjacent sequence (PAM) located at the 5' end of the target sequence and having the sequence 5'-TTG or 5'-ATN.
23. The method according to claim 21, wherein the cell is a prokaryotic cell or a eukaryotic cell, and the eukaryotic cell is an animal cell.
24. The method of claim 23, wherein the animal cell is a mammalian cell.
25. The method of claim 21, wherein the gene editing comprises one or more of the following: gene knockout of a target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking.
26. The method of claim 25, wherein the single base conversion comprises adenine to guanine, cytosine to thymine, or cytosine to uracil.
27. The method according to claim 21, wherein, The CRISPR spacer sequence of the single-stranded guide RNA forms a completely complementary base pairing structure with the target sequence, while forming an incomplete complementary base pairing structure with non-target sequences.
28. The method of claim 27, wherein the incomplete base complementary pairing structure comprises a structure with one or more base mismatches.
29. The method of claim 27, wherein the incomplete base complementary pairing structure comprises a structure with two or more base mismatches.
30. A kit for gene editing of a target sequence in an intracellular or in vitro environment, comprising: a) Choose any one of (1) to (9) below: (1) The Cas12b mutant protein of claim 1, the conjugate of claim 2 or 3, or the fusion protein of claim 4 or 5, and the single-stranded guide RNA, wherein the single-stranded guide RNA is the single-stranded guide RNA of any one of claims 6-8 or a single-stranded guide RNA containing the scaffold sequence shown in SEQ ID NO: 2; (2) The wild-type ChCas12b protein, its conjugates or fusion proteins as shown in SEQ ID NO:1, and the single-stranded guide RNA of any one of claims 6-8; (3) The isolated nucleic acid molecule of claim 9, and the isolated nucleic acid molecule of claim 11, or an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO: 2; (4) An isolated nucleic acid molecule comprising encoding the wild-type ChCas12b protein as shown in SEQ ID NO:1, its conjugate or fusion protein, and the isolated nucleic acid molecule as described in claim 11; (5) The isolated nucleic acid molecule as described in claim 10; (6) The vector of any one of claims 12-14, and the vector of claim 16 or a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA comprising the scaffold sequence shown in SEQ ID NO: 2; (7) A vector comprising encoding the wild-type ChCas12b protein as shown in SEQ ID NO:1, its conjugate or fusion protein, and the vector as described in claim 16; (8) The carrier according to claim 15; or (9) The CRISPR / Cas12b gene editing system as described in claim 17; as well as b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment.
Citation Information
Patent Citations
Cas12 protein, gene editing system containing Cas12 protein and application thereof
CN113373130A
Field deployable crispr-CAS diagnostics and methods of use thereof
WO2021163584A1