Application of truncated HNH structural domain in improving Cas protein editing efficiency

By fusing a truncated HNH domain into the V-type Cas enzyme, the problem of insufficient Cas12a editing activity was solved, its editing efficiency was improved, and its application range was expanded.

CN121825937APending Publication Date: 2026-04-10SHANDONG SHUNFENG BIOTECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing type V CRISPR endonucleases, such as Cas12a, lack the HNH domain, resulting in insufficient editing activity and limiting their application scope.

Method used

Fusing a truncated HNH domain into type V Cas enzymes enhances their editing activity.

Benefits of technology

This enhances the editing efficiency of the Cas protein and expands its application scope.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005767108870000091
    Figure BDA0005767108870000091
  • Figure BDA0005767108870000101
    Figure BDA0005767108870000101
  • Figure BDA0005767108870000111
    Figure BDA0005767108870000111
Patent Text Reader

Abstract

The invention belongs to the field of nucleic acid editing, and particularly relates to the technical field of regular clustering interval short palindromic repeat (CRISPR). Specifically, the invention provides the application of the truncated HNH structural domain in improving the Cas protein editing efficiency, and the truncated HNH structural domain has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing, particularly to the field of regularly clustered short palindromic repeats (CRISPR) technology. Specifically, this invention relates to the application of truncated HNH domains in improving the efficiency of Cas protein editing. Background Technology

[0002] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to create double-strand breaks, using biological non-homologous end joining or homologous recombination for site-specific gene editing.

[0003] Type II CRISPR endonucleases, such as Cas9, include two nuclease domains: the HNH domain and the RuvC domain. Unlike type II CRISPR endonucleases, type V CRISPR endonucleases, such as Cas12a (or cpf1), only have the RuvC domain and lack the HNH domain; as described in Chinese patent CN109207477B, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein.

[0004] This application discovers that a truncated HNH domain can enhance the editing activity of type V Cas enzymes, and has broad application prospects. Summary of the Invention

[0005] The inventors improved the editing activity and expanded the application range of the V-type Cas enzyme by fusing a truncated HNH domain into it.

[0006] On the one hand, the present invention provides the application of truncated HNH domains in improving the editing efficiency of Cas proteins, or in the preparation of Cas proteins with improved editing efficiency.

[0007] In some embodiments, the HNH domain is the HNH domain of Cas9. The amino acid sequence of the HNH domain is shown in SEQ ID No. 3.

[0008] In some embodiments, the HNH domain is selected from any one of the following groups I-III:

[0009] I. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the HNH domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID No. 3, and substantially retains the biological function of the HNH domain;

[0010] II. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the HNH domain, compared with SEQ ID No. 3, has one or more amino acid substitutions, deletions or additions (e.g., substitutions, deletions or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amino acids), and essentially retains the biological function of the HNH domain;

[0011] III. The HNH domain comprises the amino acid sequence shown in SEQ ID No. 3.

[0012] In some embodiments, the amino acid sequence of the truncated HNH domain is missing amino acids corresponding to positions 1-12 of the sequence shown in SEQ ID No. 3, or missing amino acids corresponding to positions 1-31 of the sequence shown in SEQ ID No. 3, or missing amino acids corresponding to positions 1-52 of the sequence shown in SEQ ID No. 3, or missing amino acids corresponding to positions 1-66 of the sequence shown in SEQ ID No. 3, or missing amino acids corresponding to positions 1-80 of the sequence shown in SEQ ID No. 3.

[0013] Preferably, the amino acid sequence of the truncated HNH domain is missing amino acids 1-80 corresponding to the sequence shown in SEQ ID No. 3 compared to SEQ ID No. 3.

[0014] In some embodiments, the amino acid sequence of the truncated HNH domain is shown in SEQ ID No. 4.

[0015] In some embodiments, the truncated HNH domain is selected from any group of the following I-III:

[0016] I. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the truncated HNH domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID No. 4, and substantially retains the biological function of the truncated HNH domain;

[0017] II. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the truncated HNH domain, compared with SEQ ID No. 4, has one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), and substantially retains the biological function of the truncated HNH domain;

[0018] III. The truncated HNH domain comprises the amino acid sequence shown in SEQ ID No. 4.

[0019] On the other hand, the present invention provides a method for improving the efficiency of Cas protein editing, the method comprising the step of fusing the truncated HNH domain with the Cas protein.

[0020] On the other hand, the present invention provides a fusion protein comprising a Cas protein and the aforementioned truncated HNH domain.

[0021] In one embodiment, the Cas protein is a Cas12 family protein.

[0022] In one embodiment, the Cas protein is selected from one or any of Cas12i, Cas12j, Cas12a, Cas12b, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, and Cas-sf0005.

[0023] In one embodiment, the Cas protein is a Cas protein of the Cas12j family, such as Cas12j19. Preferably, the Cas protein is the Cas12j.19 protein described in Chinese Patent CN111770992B, or the Cas12j19 obtained by amino acid mutation described in Chinese patent applications with application numbers 2023100922319, 2023110952452, and CN118613580A.

[0024] In one embodiment, the Cas protein is the wild-type Cas12j19 or the mutant Cas12j19 described above.

[0025] In one embodiment, the Cas protein is the mutated Cas protein enCas-SF02 recorded in CN118613580A, and preferably, the amino acid sequence of the Cas protein is shown in SEQ ID No. 2.

[0026] In one embodiment, the amino acid sequence of the Cas protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with SEQ ID No. 2.

[0027] In one embodiment, the amino acid sequence of the Cas protein, compared with SEQ ID No. 2, has one or more amino acid substitutions, deletions, or additions, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.

[0028] In one embodiment, the Cas protein is a Cas protein of the Cas12i family, such as Cas12i1, Cas12i2, Cas12i3, or Cas12i12.

[0029] In a preferred embodiment, the Cas protein is Cas12i3, for example, the Cas12f.4 protein described in CN111757889B, or the Cas12i3 obtained by amino acid mutation described in Chinese patent applications with application numbers 2022103148077, 2022102697541, 2022106036073, 2022109432359, 2023100884374, 2023100667809, and 2023104503761.

[0030] In one embodiment, the Cas protein is Cas12a, for example, FnCas12a, AsCas12a, LbCas12a, Lb5Cas12a, HkCas12a, OsCas12a, TsCas12a, BbCas12a, BoCas12a, or Lb4Cas12a; preferably, LbCas12a.

[0031] In some embodiments, the Cas protein is a natural wild-type Cas protein; in other embodiments, the Cas protein is an engineered Cas protein, for example, a Cas protein obtained through site-directed amino acid mutation.

[0032] In some embodiments, the truncated HNH domain is placed at the N-terminus or C-terminus of the Cas protein.

[0033] In some embodiments, the truncated HNH domain is directly linked to the N-terminus or C-terminus of the Cas protein.

[0034] In some embodiments, the truncated HNH domain is linked to the N-terminus or C-terminus of the Cas protein via a linker.

[0035] The term "connector," as is well known in the art when referring to peptide linking, refers to a chemical group or molecule that links two molecules or parts. A connector may consist of a single linking molecule (e.g., a single amino acid) or may include more than one linking molecule. In some embodiments, the connector may be an organic molecule, group, polymer, or chemical part, such as a divalent organic part. In some embodiments, the connector may be an amino acid or a peptide.

[0036] The aforementioned linkers are well known in the art and include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4 or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA or Ava), or PEG, etc.

[0037] In some embodiments, the linker may be a GS linker. In some embodiments, the linker may comprise the amino acid sequence (GGS)n, GS, SG, GSSG, S(GGS)n, SGGS, or (GGGGS)n, where n is an integer from 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the linker may comprise the amino acid sequence: SGGSGGSGGS. In some embodiments, the linker may comprise the amino acid sequence: SGSETPGTSESATPES, also known as an XTEN linker. In some embodiments, the linker may comprise the amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS, also known as a GS-XTEN-GS linker.

[0038] In this invention, the amino acid site refers to the site starting from the N-terminus of the amino acid sequence.

[0039] The present invention also provides a fusion protein, which includes the fusion protein as described above and other modified portions.

[0040] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.

[0041] In one embodiment, the modified portion is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminases (e.g., adenosine deaminase or cytidine deaminase), methyltransferase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof. The NLS sequences are well known to those skilled in the art, and examples include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc, and TUS protein.

[0042] In one embodiment, the NLS sequence is located at, near, or close to the end (e.g., N-terminus, C-terminus, or both ends) of the Cas protein of the present invention.

[0043] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection, or tracing).

[0044] The reporter gene sequences are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0045] In one embodiment, the fusion protein of the present invention includes a domain capable of binding to DNA molecules or intracellular molecules, such as maltose-binding protein (MBP), the DNA-binding domain (DBD) of Lex A, the DBD of GAL4, etc.

[0046] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.

[0047] In one embodiment, the fusion protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.

[0048] In one embodiment, the modified portion is directly connected to the N-terminus or C-terminus of the fusion protein or Cas protein of the present invention.

[0049] In one embodiment, the modified portion is attached to the N-terminus or C-terminus of the fusion protein or Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.

[0050] The fusion protein of the present invention is not limited by its production method. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0051] On the other hand, the present invention provides an isolated polynucleotide comprising:

[0052] (a) The polynucleotide sequence encoding the fusion protein of the present invention;

[0053] Alternatively, a polynucleotide complementary to the polynucleotide described in (a).

[0054] In one embodiment, the nucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the nucleotide sequence is codon-optimized for expression in eukaryotic cells.

[0055] In one embodiment, the cell is an animal cell, such as a mammalian cell.

[0056] In one embodiment, the cell is a human cell.

[0057] In one embodiment, the cell is a plant cell, such as the cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.

[0058] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0059] On the other hand, the present invention provides a gRNA comprising a first segment and a second segment; the first segment is also referred to as a "backbone region", "protein binding region", "protein binding sequence", or "direct repeat sequence"; the second segment is also referred to as a "target sequence for targeting nucleic acids", "target segment for targeting nucleic acids", or "guide sequence for targeting target sequences".

[0060] The first segment of the gRNA can interact with the Cas protein of the present invention, thereby enabling the Cas protein and gRNA to form a complex.

[0061] In a preferred embodiment, the first segment is a repeating sequence in the same direction as described above.

[0062] The target sequence or target region of the nucleic acid targeted by this invention comprises a nucleotide sequence complementary to a sequence in the target nucleic acid. In other words, the target sequence or target region of the nucleic acid targeted by this invention interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the target sequence or target region of the nucleic acid can be altered or modified to hybridize with any desired sequence within the target nucleic acid. The nucleic acid is selected from DNA or RNA.

[0063] The percentage of complementarity between the target sequence or target region of the target nucleic acid and the target sequence of the target nucleic acid may be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).

[0064] The "backbone region," "protein-binding region," "protein-binding sequence," or "direct repeat sequence" of the gRNA of this invention can interact with CRISPR proteins (or Cas proteins). The gRNA of this invention guides the interacting Cas protein to a specific nucleotide sequence within the target nucleic acid through the targeting sequence of the target nucleic acid.

[0065] Preferably, the guide RNA comprises a first segment and a second segment in the 5' to 3' direction.

[0066] In this invention, the second segment can also be understood as a guide sequence for hybridization with the target sequence.

[0067] The gRNA of the present invention can form a complex with the Cas protein.

[0068] The present invention also provides a carrier comprising, as described above, a fusion protein, an isolated nucleic acid molecule or a polynucleotide; preferably, it further comprises a regulatory element operatively linked thereto.

[0069] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.

[0070] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.

[0071] In some implementations, the vectors included in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated vectors, and herpes simplex vectors), and may also be plasmids, viruses, granules, bacteriophages, etc., which are well known to those skilled in the art.

[0072] On the other hand, the present invention provides a composition comprising:

[0073] (i) a protein component selected from: the engineered fusion protein described above; and

[0074] (ii) A nucleic acid component comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a unidirectional repeat sequence capable of binding to the Cas protein in the fusion protein of the present invention.

[0075] The protein components and nucleic acid components combine to form a complex.

[0076] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.

[0077] In one embodiment, the complex or composition is non-natural or modified. In one embodiment, at least one component of the complex or composition is non-natural or modified. In one embodiment, the first component is non-natural or modified; and / or, the second component is non-natural or modified.

[0078] On the other hand, the present invention provides an engineered host cell comprising the above-described fusion protein, or the above-described polynucleotide, or the above-described vector, or the above-described composition.

[0079] In some implementations, the cell is a prokaryotic cell.

[0080] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).

[0081] In some implementations, the cell is a stem cell or stem cell line.

[0082] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.

[0083] The present invention also provides the use of the above-mentioned fusion protein, or the above-mentioned polynucleotide, or the above-mentioned vector, or the above-mentioned composition, or the above-mentioned host cell in gene editing; or, in the preparation of reagents or kits for gene editing.

[0084] The present invention also provides a method for editing a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned fusion protein, or the aforementioned polynucleotide, or the aforementioned vector, or the aforementioned composition, or the aforementioned host cell.

[0085] In one embodiment, the method is to edit, target, or cleave target nucleic acids intracellularly or extracellularly.

[0086] The gene editing includes modifying genes, knocking out genes, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides, and gene mutations.

[0087] The editing can be performed in prokaryotic and / or eukaryotic cells.

[0088] On the other hand, the present invention also provides a kit for gene editing, the kit comprising the above-mentioned fusion protein, or the above-mentioned polynucleotide, or the above-mentioned vector, or the above-mentioned composition, or the above-mentioned host cell.

[0089] Beneficial effects of the invention

[0090] This invention fuses a truncated HNH domain with the Cas protein, thereby enhancing the activity of the Cas protein and showing broad application prospects. Attached Figure Description

[0091] Figure 1 It is the editing efficiency of different truncated Cas proteins.

[0092] Figure 2 It is the truncated variant HNH-Δ80-enCas-SF02 that improves multi-target editing efficiency.

[0093] Figure 3 This is a comparison of the multi-target gene editing efficiency of enCas-SF02 and the truncated variant HNH-Δ80-enCas-SF02 in soybean.

[0094] The sequence information involved in this invention is as follows:

[0095]

[0096] Detailed Implementation

[0097] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0098] The following embodiments are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the embodiments are generally performed in accordance with conventional methods well known in the art and described in various references.

[0099] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0100] Example 1. Obtaining the Cas mutant protein

[0101] The known Cas protein HNH-enCas-SF02 (SaHNH-enCas-SF02 in CN118613580A, an engineered Cas protein obtained by inserting the HNH domain from Cas9 into the space between the first and second amino acids from the N-terminus of enCas-SF02) has the amino acid sequence shown in SEQ ID No. 1. The amino acid sequence of enCas-SF02 is shown in SEQ ID No. 2, and the amino acid sequence of the HNH domain is shown in SEQ ID No. 3. Based on the structural information, the applicant truncated the HNH domain, using HNH-enCas-SF02 as a control, and tested the editing efficiency in cells. The truncated HNH domain portions correspond to amino acids 1-12, 1-31, 1-52, 1-66, and 1-80 at the N-terminus of SEQ ID No. 3, respectively. The truncated Cas proteins are designated as HNH-Δ12-enCas-SF02, HNH-Δ31-enCas-SF02, HNH-Δ52-enCas-SF02, HNH-Δ66-enCas-SF02, and HNH-Δ80-enCas-SF02. Construction was achieved using Gibson clone based on PCR amplification, and the final vector was conjugated into the pcDNA3.3-eGFP vector. Fragment amplification kit: TransStart FastPfu DNA Polymerase (containing 2.5 mM dNTPs). See the instruction manual for detailed experimental procedures. Gel recovery kit: For detailed experimental procedures, please refer to the instruction manual for the Gel DNA Extraction Mini Kit. The reagent kit used for vector construction is the pEASY-Basic Seamless Cloning and Assembly Kit (CU201-03). For detailed experimental procedures, please refer to the instruction manual.

[0102] The vector pcDNA3.3 was modified to carry EGFP fluorescent protein and the PuroR resistance gene. The SV40 NLS-Cas-XX fusion protein was inserted via XbaI and PstI restriction sites; the U6 promoter and gRNA sequence were inserted via Mfe1 restriction site, with the target sequence being ATGgcttctcatcgtctgctcct (uppercase letters represent the PAM sequence). The CMV promoter initiated the expression of the SV40 NLS-Cas-XX-NLS-GFP fusion protein. The Cas-XX-NLS protein and the GFP protein were linked using the linker peptide T2A. The EF-1α promoter initiated the expression of the puromycin resistance gene. Plating: 293T cells were plated when the confluence reached 70-80%, with a cell number of 8*10^4 cells / well in 12-well plates. Transfection: Transfection was performed 24 hours after plating, with 6.25 μl of Hieff Transfection Technology added to 100 μl opti-MEM. TM Liposome nucleic acid transfection reagent, mix well; add 2.5 μg plasmid to 100 μl opti-MEM, mix well. Diluted HieffTrans... TM The liposome nucleic acid transfection reagent was mixed thoroughly with the diluted plasmid and incubated at room temperature for 20 min. The incubated mixture was then added to cell-coated culture medium for transfection. 48 h after transfection, cells were digested with trypsin-EDTA (0.05%) and sorted using flow cytometry (FACS) for GFP signaling.

[0103] DNA extraction, PCR amplification of the region near the editing site, and hiTOM sequencing (target sites and amplification primers are shown in the table below): Cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). Genomic DNA was amplified in the region near the target site. PCR products were then sequenced using hiTOM. Sequencing data were analyzed, and the types and proportions of sequences within a 15nt upstream and 10nt downstream of the target site were statistically analyzed. Sequences with a SNV frequency greater than or equal to 1% or a non-SNV mutation frequency greater than or equal to 0.06% were identified to determine the editing efficiency of different Cas proteins at the target site.

[0104]

[0105]

[0106] Statistics on HNH-enCas-SF02 ( Figure 1 HNH-SF02) and its truncated variant HNH-Δ12-enCas-SF02 ( Figure 1 HNH(1-12)-SF02), HNH-Δ31-enCas-SF02 ( Figure 1HNH(1-31)-SF02), HNH-Δ52-enCas-SF02 ( Figure 1 HNH(1-52)-SF02 and HNH-Δ66-enCas-SF02 ( Figure 1 HNH(1-66)-SF02), HNH-Δ80-enCas-SF02 ( Figure 1 The editing efficiency of HNH(1-80)-SF02) at the target site TTR-g15 was as follows: Figure 1 As shown, compared with HNH-enCas-SF02, there was no significant difference in editing efficiency between HNH-Δ12-enCas-SF02, HNH-Δ52-enCas-SF02 and HNH-Δ66-enCas-SF02, HNH-Δ31-enCas-SF02 had reduced editing efficiency, and HNH-Δ80-enCas-SF02 had an improved editing efficiency of about 10%.

[0107] Statistics on HNH-enCas-SF02 ( Figure 2 HNH-SF02) and its truncated variant HNH-Δ80-enCas-SF02 ( Figure 2 The editing efficiency of HNH(1-80)-SF02 in the TTR and TRAC targets (except TTR-g15) in the table above is as follows: Figure 2 As shown, the average editing efficiency of HNH-enCas-SF02 is 56.78%, while the average editing efficiency of HNH-Δ80-enCas-SF02 is improved to 66.26%.

[0108] The above results indicate that the editing efficiency of HNH-Δ80-enCas-SF02 is significantly higher than that of HNH-enCas-SF02, and that shortening the amino acids 1-80 of HNH-enCas-SF02 can greatly improve its editing efficiency in eukaryotic cells.

[0109] Example 2. Validation of the editing efficiency of Cas protein in plants

[0110] The truncated HNH-Δ80-enCas-SF02 protein obtained in Example 1, compared to enCas-SF02 (enCas-SF02 in CN118613580A, a tetraprotic protein of Cas12j19, with an amino acid sequence as shown in SEQ ID No. 2), is a Cas protein obtained by inserting a truncated HNH-Δ80 domain between the first and second amino acids from the N-terminus of enCas-SF02. The amino acid sequence of the original HNH domain is shown in SEQ ID No. 3. The amino acid sequence of the truncated HNH-Δ80 domain, compared to SEQ ID No. 3, is missing amino acids 1-80. The amino acid sequence of the truncated HNH-Δ80 domain is shown in SEQ ID No. 4.

[0111] In this embodiment, the enCas-SF02 and HNH-Δ80-enCas-SF02 proteins were constructed into the vector of Agrobacterium rhizogenes K599, respectively. Sterilized soybean seeds were sown in a sterile, moist substrate of 2:1 mixture of nutrient soil and vermiculite at a depth of 1 cm. The plants were cultured in a greenhouse under conditions of 25°C, 70% relative humidity, and a 12-hour light / 12-hour dark cycle for 7-14 days. A 1 cm oblique incision was made at the hypocotyl of the soybean seedling using a sterile scalpel. Agrobacterium rhizogenes K599 bacterial suspension containing the expression vector was directly applied to the oblique incision surface. The seedlings were then planted in sterile vermiculite pots and irrigated with 10 ml of Agrobacterium rhizogenes K599 bacterial suspension. A high-humidity environment was maintained for 16 days to promote the growth of hairy roots. Finally, genomic DNA was extracted from the hairy roots, and PCR amplification was performed to analyze the editing efficiency. The target sites and primers for soybean hairy roots are shown in the table below.

[0112]

[0113]

[0114] Editing efficiency such as Figure 3 As shown, in 24 hairy root samples, the enCas-SF02 protein ( Figure 3 The average editing efficiency of SF02 at 6 sites was 66.38%; while the HNH-Δ80-enCas-SF02 protein ( Figure 3 The average editing efficiency of HNH(1-80)-SF02 was 91.67%, and the editing efficiency reached 100% at multiple sites.

[0115] This indicates that, compared with enCas-SF02, HNH-Δ80-enCas-SF02 significantly improves the editing efficiency in soybeans, and the truncated HNH-Δ80 domain can enhance the editing efficiency of enCas-SF02 in soybeans.

Claims

1. The application of truncated HNH domains in improving the editing efficiency of Cas proteins, or in the preparation of Cas proteins with improved editing efficiency; characterized in that, The truncated HNH domain is selected from any one of the following groups I-III: I. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the truncated HNH domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID No. 4, and substantially retains the biological function of the truncated HNH domain; II. The HNH domain is the HNH domain of Cas9, and the amino acid sequence of the truncated HNH domain, compared with SEQ ID No. 4, has one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), and substantially retains the biological function of the truncated HNH domain; III. The truncated HNH domain comprises the amino acid sequence shown in SEQ ID No.

4.

2. A method for improving the efficiency of Cas protein editing, the method comprising the step of fusing the truncated HNH domain of claim 1 with the Cas protein.

3. A fusion protein comprising a Cas protein and a truncated HNH domain as described in claim 1, wherein the Cas protein is a Cas12 family protein.

4. An isolated polynucleotide, characterized in that, The polynucleotide is a polynucleotide sequence encoding the fusion protein of claim 3.

5. A carrier, characterized in that, The vector comprises the polynucleotide of claim 4 and a regulatory element operatively linked thereto.

6. A composition, characterized in that, The composition comprises: (i) a protein component selected from the fusion protein of claim 3; (ii) A nucleic acid component, which is gRNA, said gRNA being capable of binding the Cas protein in the fusion protein of claim 3; The protein components and nucleic acid components combine to form a complex.

7. An engineered host cell, characterized in that, The host cell comprises the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the composition of claim 6.

8. The use of the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the composition of claim 6, or the host cell of claim 7 in gene editing; or, in the preparation of a kit for gene editing.

9. A method for editing a target nucleic acid, the method comprising contacting the target nucleic acid with a fusion protein of claim 3, or a polynucleotide of claim 4, or a vector of claim 5, or a composition of claim 6, or a host cell of claim 7.

10. A kit for gene editing, the kit comprising the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the composition of claim 6, or the host cell of claim 7.

Citation Information

Patent Citations

  • CRISPR enzymes and systems

    CN109207477B

  • Novel CRISPR / Cas12f enzymes and systems

    CN111757889B

  • CRISPR-Cas12j enzymes and systems

    CN111770992B

  • Engineered Cas protein and application thereof

    CN118613580A