Engineered nuclease and application thereof
By modifying the Cas protein domains, removing unnecessary domains, and optimizing the PI, REC, and WED domains, a small molecular weight nuclease was developed, solving the problem of the lack of nuclease activity in traditional Cas proteins and realizing its potential application in gene editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-03-13
AI Technical Summary
Traditional V-type Cas proteins, such as the Cas12j family proteins, lack nuclease activity, which limits their potential application in gene editing.
By modifying the Cas protein domains, an engineered nuclease was developed. The RuvC and HNH domains were knocked out or inactivated, the REC and WED domains were retained, and the arrangement of the PI domains was optimized to form a low molecular weight nuclease.
The development of small molecular weight nucleases has been achieved, which have broad application prospects and can exert nuclease activity in gene editing.
Smart Images

Figure CN121653099A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and more particularly to enzyme modification technology. Specifically, this invention provides an engineered nuclease and its applications; by optimizing and modifying the Cas protein domain, this invention obtains a nuclease with a smaller protein molecular weight. Background Technology
[0002] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to create double-strand breaks, using biological non-homologous end joining or homologous recombination for site-specific gene editing.
[0003] Type V Cas proteins, such as the Cas12j family proteins, mainly have an N-terminal recognition domain composed of three domains: PI, REC-I, and WED. Their main functions are to recognize crRNA, unwound DNA and target DNA, and PAM sequences. Traditionally, they are not considered to have nuclease activity and do not participate in the nuclease function of the C-terminal cleavage domain.
[0004] This application, through the modification and optimization of the Cas protein domains, discovered that a truncated protein composed of the PI, REC-I, and WED domains at the N-terminus of the V-type Cas protein can exert nuclease activity, proposing a new path and research direction for the development and application of nucleases. Summary of the Invention
[0005] The inventors developed a novel small-molecule nuclease by modifying the structural domains of the Cas protein.
[0006] engineered nucleases
[0007] On the one hand, the present invention provides an engineered nuclease comprising a REC domain and a WED domain.
[0008] Furthermore, the nuclease does not include a RuvC domain and / or an HNH domain; or, the nuclease includes a RuvC domain and / or an HNH domain that are nuclease activity inactivated.
[0009] In one embodiment, the nuclease consists of a REC domain and a WED domain.
[0010] In one embodiment, the REC structure domain is located at the N-terminus or C-terminus of the WED structure domain; preferably, the REC structure domain is located at the N-terminus of the WED structure domain.
[0011] Furthermore, the engineered nuclease of the present invention also includes a PI domain; preferably, the PI domain is located at the N-terminus of the nuclease.
[0012] In one embodiment, the nuclease comprises a PI domain, a REC domain, and a WED domain.
[0013] In one embodiment, the REC domain is selected from REC I and REC II domains, preferably REC I domain.
[0014] In one embodiment, the engineered nuclease includes a REC domain and a WED domain from the N-terminus to the C-terminus.
[0015] In other embodiments, the engineered nuclease comprises, from the N-terminus to the C-terminus, a PI domain, a REC domain, and a WED domain.
[0016] In one embodiment, the PI domain, the REC domain, and the WED domain are derived from the Cas protein.
[0017] In one embodiment, the Cas protein is selected from one or more of the V-type Cas proteins, such as Cas12a, Cas12i, Cas12j, and Cas-sf0005; preferably, Cas proteins of the Cas12j family.
[0018] In a preferred embodiment, the Cas protein is selected from one or more of Cas-sf0005 (as described in Chinese Patent CN114438055B), Cas12j-2 (as shown in SEQ ID No. 15), and Cas12j.19 (as described in Patent Application WO2020098772A1).
[0019] In one embodiment, the PI domain, the REC domain, and the WED domain are derived from Cas-sf0005, Cas12j-2 (as shown in SEQ ID No. 15), or Cas12j.19.
[0020] In one embodiment, the amino acid sequence of Cas-sf0005 is shown in SEQ ID No. 14; the amino acid sequence of Cas12j-2 is shown in SEQ ID No. 15; and the amino acid sequence of Cas12j.19 is shown in SEQ ID No. 16.
[0021] In one implementation, the REC domain is the REC I domain.
[0022] In one embodiment, the REC I domain is derived from Cas-sf0005, and the amino acid sequence of the REC I domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with SEQ ID No. 1.
[0023] In one embodiment, the REC I domain is derived from Cas12j-2, and the amino acid sequence of the REC I domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with SEQ ID No. 2.
[0024] In one embodiment, the REC I domain is derived from Cas12j19, and the REC I domain contains the sequences shown in SEQ ID No. 3 and SEQ ID No. 4.
[0025] In other embodiments, the amino acid sequence of the REC I domain has one or more amino acid substitutions, deletions, or additions compared to the amino acid sequences described in any of SEQ ID No. 1-4, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.
[0026] In one embodiment, the amino acid sequence of the WED domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the amino acid sequences described in SEQ ID No. 5-10.
[0027] In other embodiments, the amino acid sequence of the WED domain has one or more amino acid substitutions, deletions, or additions compared to any of the amino acid sequences described in SEQ ID No. 5-10, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.
[0028] In one embodiment, the WED domain is derived from Cas-sf0005, and the WED domain contains the sequences shown in SEQ ID No. 5 and SEQ ID No. 6.
[0029] In one embodiment, the WED domain is derived from Cas12j-2, and the WED domain contains the sequences shown in SEQ ID No. 7 and SEQ ID No. 8.
[0030] In one embodiment, the WED domain is derived from Cas12j19, and the WED domain contains the sequences shown in SEQ ID No. 9 and SEQ ID No. 10.
[0031] In one embodiment, the amino acid sequence of the PI domain has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the amino acid sequences described in SEQ ID No. 11-13.
[0032] In other embodiments, the amino acid sequence of the PI domain has one or more amino acid substitutions, deletions, or additions compared to any of the amino acid sequences described in SEQ ID No. 11-13, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.
[0033] In one embodiment, the PI domain, REC domain, and WED domain of the engineered nuclease of the present invention can be directly linked to each other, or they can be linked through a linker or peptide.
[0034] The term "connector," as is well known in the art when referring to peptide linking, refers to a chemical group or molecule that links two molecules or parts. A connector may consist of a single linking molecule (e.g., a single amino acid) or may include more than one linking molecule. In some embodiments, the connector may be an organic molecule, group, polymer, or chemical part, such as a divalent organic part. In some embodiments, the connector may be an amino acid or a peptide.
[0035] The aforementioned linkers are well known in the art and include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4 or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA or Ava), or PEG, etc.
[0036] In some embodiments, the linker may be a GS linker. In some embodiments, the linker may comprise the amino acid sequence (GGS)n, GS, SG, GSSG, S(GGS)n, SGGS, or (GGGGS)n, where n is an integer from 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the linker may comprise the amino acid sequence: SGGSGGSGGS. In some embodiments, the linker may comprise the amino acid sequence: SGSETPGTSESATPES, also known as an XTEN linker. In some embodiments, the linker may comprise the amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS, also known as a GS-XTEN-GS linker.
[0037] The term "REC domain" or "recognition domain" refers to the α-helical domain that is believed to contact the guide RNA. Generally, it refers to the domain that is believed to interact with the repeat:anti-repeat double helix of gRNA and mediate the formation of the Cas endonuclease / gRNA complex.
[0038] The term "WED domain" or "wedge domain" typically refers to a domain that primarily interacts with the repeat:anti-repeat duplex of sgRNA and PAM duplex.
[0039] The term "PI domain" or "PAM interaction domain" generally refers to the domain that interacts with the protospacer neighbor motif (PAM) outside the seed sequence in the region targeted by the Cas protein.
[0040] The RuvC domain or HNH domain is usually the domain in Cas proteins that performs nuclease cleavage activity.
[0041] In one embodiment, the engineered nuclease is selected from any group I-III of the following:
[0042] I. The amino acid sequence of the engineered nuclease is shown in any one of SEQ ID No. 17-20;
[0043] II. Compared with the engineered nuclease described in I, it has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, and has nuclease activity;
[0044] III. Compared with the engineered nuclease described in I, it has one or more amino acid substitutions, deletions, or additions, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids; and has nuclease activity.
[0045] Those skilled in the art will understand that the structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, i.e., a nonpolar amino acid residue can replace another nonpolar amino acid residue, a polar uncharged amino acid residue can replace another polar uncharged amino acid residue, a basic amino acid residue can replace another basic amino acid residue, and an acidic amino acid residue can replace another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this invention, as long as the substitution does not lead to the inactivation of the protein's biological activity. Therefore, the V-type Cas protein engineered by this invention can contain one or more conserved substitutions in its amino acid sequence, which are preferably generated by substitutions according to Table 1. In addition, the present invention also covers proteins that contain one or more other non-conservative substitutions, provided that such non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present invention.
[0046] Conserved amino acid substitutions can be performed at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing biological activity, while “essential” amino acid residues are required for biological activity. A “conserved amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can be performed in the non-conserved regions of the engineered V-type Cas protein described above. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.
[0047] Table 1
[0048]
[0049]
[0050] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C ends of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from their N and / or C ends while retaining their desired functional activity are also within the scope of this invention. These alterations can include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.
[0051] It should be recognized that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of the aforementioned proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0052] Those skilled in the art will understand that these minor amino acid changes in the Cas protein of this invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0053] Those skilled in the art can identify the essential amino acids of the engineered V-type Cas protein of this invention using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of the protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.
[0054] In this invention, amino acid residues can be represented by a single letter or by three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamic acid (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine (Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), and arginine (Arg, R).
[0055] The engineered nuclease of the present invention is not limited by its production method. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0056] Nucleic acid encoding engineered nucleases
[0057] On the other hand, the present invention provides an isolated polynucleotide comprising:
[0058] (a) A multinucleotide sequence encoding the engineered nuclease of the present invention, or a multinucleotide complementary to the multinucleotide described in (a).
[0059] In one embodiment, the nucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the nucleotide sequence is codon-optimized for expression in eukaryotic cells.
[0060] In one embodiment, the cell is an animal cell, such as a mammalian cell.
[0061] In one embodiment, the cell is a human cell.
[0062] In one embodiment, the cell is a plant cell, such as the cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.
[0063] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0064] carrier
[0065] The present invention also provides a carrier comprising an engineered nuclease or isolated polynucleotide as described above; preferably, it further comprises a regulatory element operatively connected thereto.
[0066] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.
[0067] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.
[0068] In some implementations, the vectors included in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated vectors, and herpes simplex vectors), and may also be plasmids, viruses, granules, bacteriophages, etc., which are well known to those skilled in the art.
[0069] Delivery and delivery composition
[0070] The engineered nucleases, nucleic acid molecules, and vectors of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.
[0071] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery carrier, and one or more of the engineered nucleases, nucleic acid molecules, and carriers of the present invention.
[0072] In one embodiment, the delivery carrier is a particle.
[0073] In one embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0074] host cells
[0075] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the engineered nucleases, nucleic acid molecules, and vectors of the present invention.
[0076] In some implementations, the cell is a prokaryotic cell.
[0077] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).
[0078] In some implementations, the cell is a stem cell or stem cell line.
[0079] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.
[0080] Methods and Applications
[0081] The present invention also provides the use of the above-described engineered nucleases, nucleic acid molecules, vectors or host cells in cleaving or splitting target nucleic acids; or in the preparation of reagents for cleaving or splitting target nucleic acids.
[0082] On the other hand, the present invention also provides a method for cleaving or splitting target nucleic acid, the method comprising the step of contacting the engineered nuclease described above with the target nucleic acid.
[0083] In one embodiment, the target nucleic acid is DNA, such as double-stranded DNA or single-stranded DNA.
[0084] In one embodiment, the target nucleic acid is a circular nucleic acid and / or a linear nucleic acid; for example, circular DNA or linear DNA.
[0085] On the other hand, the present invention also provides an enzyme preparation comprising the above-described engineered nuclease.
[0086] In one embodiment, the engineered nuclease of the present invention exhibits endonuclease or exonuclease activity.
[0087] Terminology Definition
[0088] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0089] identity
[0090] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomeric subunit occupies the same location (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum identity. Such comparisons can be made using methods readily available, for example, computer programs such as the Align program (DNAstar, Inc.) Needleman et al. (1970) J. Mol. Biol. 48: 443-453. The percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) integrated into the ALIGN program (version 2.0), which uses a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the Needleman and Wunsch algorithm (J MoIBiol. 48:444-453 (1970)) in the GAP program integrated into the GCG software package (available at www.gcg.com), which uses a Blossum 62 matrix or a PAM250 matrix, along with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.
[0091] carrier
[0092] The term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication initiation site.
[0093] One type of vector is a "plasmid," which is a circular double-stranded DNA loop into which another DNA fragment can be inserted, for example, using standard molecular cloning techniques.
[0094] Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package the virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by the virus used for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0095] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome upon introduction and thereby replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors."
[0096] host cells
[0097] As used herein, the term “host cell” refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0098] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0099] Control element
[0100] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements may also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; CMV enhancers.
[0101] promoter
[0102] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.
[0103] Operable connection
[0104] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0105] Complementarity
[0106] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0107] connector
[0108] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of multiple amino acid residues via peptide bonds. The linkers of this invention can be synthetically produced amino acid sequences or naturally occurring polypeptide sequences, such as polypeptides with hinge region functions. Such linker polypeptides are well known in the art.
[0109] Beneficial effects of the invention
[0110] This invention utilizes the REC and WED domains in the Cas protein through engineering modification to obtain a small-molecule protein with nuclease activity, which has broad application prospects.
[0111] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0112] Figure 1 Schematic diagram of the .Cas-sf0005 structural domain.
[0113] Figure 2 Results of Cas-sf0005 and truncated proteins N1 and N2 nuclease activities in the presence of CrRNA.
[0114] Figure 3 Results of Cas-sf0005 and truncated protein N2(0053) nuclease activity assays in the absence of CrRNA.
[0115] Figure 4 Results of Cas-sf0005 and truncated protein N1 and N2 nuclease activity assays.
[0116] Figure 5 Schematic diagram of the Cas12j-2 structural domain.
[0117] Figure 6 Results of Cas12j-2 truncated N1 nuclease activity assay.
[0118] Figure 7 Schematic diagram of the .Cas12j19 structural domain.
[0119] Figure 8 Cas12j19 truncated N1 nuclease activity test results
[0120] The sequence information involved in this invention is as follows:
[0121]
[0122]
[0123]
[0124] Detailed Implementation
[0125] The following embodiments are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the embodiments are generally performed in accordance with conventional methods well known in the art and described in various references.
[0126] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.
[0127] Example 1. Modification and functional verification of the Cas-sf0005 protein
[0128] In this embodiment, the structural domains of Cas-sf0005 (amino acid sequence as shown in SEQ ID No. 14, also recorded in Chinese patent CN114438055B) are modified.
[0129] like Figure 1 As shown, the N-terminal amino acid region of Cas-sf0005, from amino acid 1 to 356, is composed of three domains: PI, REC-I, and WED. Its main function is to recognize crRNA, unwound DNA, and target DNA PAM sequences. It is traditionally believed to have no nuclease activity and does not participate in the nuclease function of the C-terminal cleavage domain (which is mainly composed of RuvC1-3 and REC-II domains).
[0130] In this embodiment, we obtained the N-terminal proteins N1 (amino acids 1-356 of the N-terminus, including the PI, REC-I, and WED domains, as shown in SEQ ID No. 17) and N2 (amino acids 57-356 of the N-terminus, including the REC-I and WED domains, as shown in SEQ ID No. 18) of Cas-sf0005. Their nuclease activity was then tested using the following methods:
[0131] Digestion Buffer: PBS+2mM DTT+5mM MgCl2
[0132] A. Configuration method:
[0133] 1. First, prepare approximately 10 μL of buffer: PBS + 2 mM DTT + 5 mM MgCl2. Then, add the protein in sequence, mix well, add the nucleic acid, mix well again, and then let it stand at room temperature.
[0134] 2. Add the target protein. If the concentration is 2 mg / ml, add 1 μL and mix well.
[0135] 3. Add nucleic acid DNA (PCR product, adjust the pH to neutral with Tris-HCl or 10X PBS, circular plasmid approximately 4.0 kb). If the concentration is 100 ng / µl, add 1 µl and mix well.
[0136] 4. For the remaining 25ul, add PDB + 2mM DTT + 5mM MgCl2 and let stand at room temperature for 2 hours.
[0137] B. Nucleic acid electrophoresis
[0138] Add 2mM EDTA and boil for 5-10 minutes to inactivate the enzyme and release nucleic acid. Add loading buffer, run nucleic acid electrophoresis gel, take a picture and save it with inverted color.
[0139] After the reaction time is up, add 1 μL of Beyotime proteinase (Proteinase K, 10 mg / ml) and mix well. Incubate at 37°C for 1 hour. Then add loading buffer, run nucleic acid electrophoresis gel, take a picture and save it in reverse color.
[0140] Figure 2 In this model, pDNA is circular DNA, and lDNA is linear DNA; CrRNA is GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGACUCCCCUACGUGCUGCUGA AG; 005-WT is the full-length Cas-sf0005, 005-N1 is the truncated Cas-sf0005 protein N1, 005-N2 is the truncated Cas-sf0005 protein N2, and 005(3A) is a nuclease domain inactivation mutant of Cas-sf0005. Figure 2 It is known that the full-length Cas-sf0005, the cleavage-inactivating mutant 005(3A), and the truncated proteins N1 and N2 can all exhibit nuclease activity against circular DNA and linear DNA, and this activity is independent of the nuclease domain at its C-terminus (because the 005(3A) inactivating mutant also has nuclease activity).
[0141] Furthermore, in the absence of CrRNA, we used the Cas-sf0005 truncating proteins N1 and N2 to cleave DNA, and the results are as follows: Figure 3 As shown.
[0142] Figure 3 In this DNA sequence, pDNA is circular DNA, while LDNA and LpDNA-2 are linear DNA; CrRNA is GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGACUCCCCUACGUGCUGCUGA AG; 005-N1 is a truncated Cas-sf0005 protein N1, composed of... Figure 3 It can be seen that the Cas-sf0005 truncated protein N1 can still exhibit nuclease activity against both circular and linear DNA even in the absence of crRNA.
[0143] Depend on Figures 2-3 It is known that the nuclease activities of Cas-sf0005 truncated proteins N1 and N2 in cleaving circular and linear DNA are not controlled by CrRNA.
[0144] Furthermore, we conducted a series of studies on the active sites of Cas-sf0005 and truncated proteins N1 and N2, ultimately discovering that the E74A and D303A mutations can inactivate the nuclease activity of truncated proteins N1 and N2. Figure 4 As shown.
[0145] like Figure 4As shown, pDNA is circular DNA, LDNA is linear DNA; CrRNA is GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGACUCCCCUACGUGCUGCUGA AG; 005-N1 is a mutant of Cas-sf0005 truncated protein N1 (E74A, D303A).
[0146] Example 2. Modification and functional verification of Cas12j-2 and Cas12j19 proteins
[0147] In this embodiment, following the procedure in Example 1, the homologous proteins Cas12j-2 and Cas12j19 of Cas-sf0005 (Cas12j family) are truncated to obtain the N-terminal protein.
[0148] The full-length amino acid sequence of Cas12j-2 is shown in SEQ ID No. 15, and its truncated protein includes N-terminal P1, REC-1, and WED domains. Using a similar technical approach as in Example 1, its nuclease activity was tested, and the results are as follows. Figure 6 As shown.
[0149] like Figure 6 As shown, pDNA is circular DNA, and lDNA is linear DNA; Cas12j-2-N1 is the N-terminal truncated protein of Cas12j-2 (amino acid sequence shown in SEQ ID No. 19). The results indicate that the N-terminal truncated protein of Cas12j-2 also possesses nuclease activity.
[0150] like Figure 8 As shown, pDNA is circular DNA, and LDNA is linear DNA; Cas12o1-N1 is an N-terminal truncated protein of Cas-12j19 (amino acid sequence as shown in SEQ ID No. 20, including PI, REC-I and WED domains). Figure 8 The results show that the N-terminus of the Cas12j19 protein also has nuclease activity.
[0151] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.
Claims
1. An engineered nuclease comprising a REC domain and a WED domain, wherein the REC domain and the WED domain are derived from a Cas protein, and wherein the nuclease does not include a RuvC domain and / or an HNH domain.
2. The engineered nuclease according to claim 1, characterized in that, The engineered nuclease consists of a REC domain and a WED domain; the REC domain is selected from REC I and REC II domains.
3. The engineered nuclease according to any one of claims 1-2, characterized in that, The engineered nuclease also includes a PI domain.
4. The engineered nuclease according to any one of claims 1-3, characterized in that, The PI domain, the REC domain, and the WED domain are derived from type V Cas protein.
5. The engineered nuclease according to claim 4, characterized in that, The engineered nuclease is selected from any one of the following groups I-III: I. The amino acid sequence of the engineered nuclease is shown in any one of SEQ ID No. 17-20; II. Compared with the engineered nuclease described in I, it has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, and has nuclease activity; III. Compared with the engineered nuclease described in I, it has one or more amino acid substitutions, deletions, or additions, for example, substitutions, deletions, or additions of 1-20 amino acids, or substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids; Furthermore, it possesses nuclease activity.
6. An isolated polynucleotide comprising: (a) Encoding the polynucleotide sequence of the engineered nuclease according to any one of claims 1-5 Alternatively, a polynucleotide complementary to the polynucleotide described in (a).
7. A vector comprising an engineered nuclease according to any one of claims 1-5 or an isolated polynucleotide according to claim 6.
8. The use of the engineered nuclease of any one of claims 1-5, the isolated polynucleotide of claim 6, or the vector of claim 7 in cleaving or splitting target nucleic acids; or, in the preparation of reagents for cleaving or splitting target nucleic acids.
9. A method for cleaving or splitting target nucleic acid, the method comprising the step of contacting the engineered nuclease of any one of claims 1-5 with the target nucleic acid.
10. An enzyme preparation comprising the engineered nuclease according to any one of claims 1-5.
Citation Information
Patent Citations
Novel CRISPR enzymes and systems and their applications
CN114438055B
Crispr-cas12j enzyme and system
WO2020098772A1