Iscb protein and use thereof
By developing a novel RNA-guided endonuclease IscB protein and ωRNA complex, the problems of large size and insufficient diversity of existing CRISPR/Cas systems have been solved, enabling efficient and flexible gene editing and detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG SHUNFENG BIOTECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing CRISPR/Cas systems suffer from large size, require two RNAs for guidance, and lack diversity, which limits the flexibility and ease of delivery in gene editing.
A novel RNA-guided endonuclease, IscB protein, was developed, which has a smaller size and better modification potential. Based on this, corresponding gene editing tools and nucleic acid detection methods were developed, including the design of amino acid sequence variants of the IscB protein and ωRNA to form complexes that target specific nucleic acid sequences for modification.
It provides a more robust gene editing system, improving the flexibility and delivery convenience of gene editing, enabling efficient targeting and modification of specific nucleic acid sequences, and is suitable for gene editing and detection in various cell types.
Smart Images

Figure CN122104640A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202411689377.2, filed on November 25, 2024, entitled "An IscB protein and its application".
[0002] This application claims priority to Chinese patent application CN202311588489.4, filed on November 27, 2023. The entire contents of the aforementioned Chinese patent application are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of gene editing. Specifically, this invention relates to a novel IscB protein and its applications. More specifically, this invention has screened a novel RNA-guided endonuclease IscB protein and developed corresponding gene editing tools and their applications based on this novel IscB protein. Background Technology
[0004] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to create double-strand breaks, using biological non-homologous end joining or homologous recombination for site-specific gene editing.
[0005] IscB protein is a novel RNA-guided endonuclease with a small size. Its homologues include IsrB and IshB proteins. Together with TnpB protein, they can be referred to as the OMEGA (Obligate Mobile Element Guided Activity) system or the Ω system. IscB protein is also a CRISPR-related protein, containing one or more RNA-guided endonuclease domains capable of modifying target nucleic acids; therefore, IscB protein can be called Cas IscB protein. IscB protein contains multiple domains, including split RuvC domains (Ruv-C I, Ruv-C II, and Ruv-C III), HNH domains, PLMP domains, and bridged helix (BH) domains.
[0006] The different CRISPR / Cas proteins currently available each have their own advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two guide RNAs, while Cpf1 only requires one and can be used for multiplex gene editing. CasX is 980 amino acids in size, while common proteins like Cas9, C2c1, CasY, and Cpf1 are typically around 1300 amino acids. In contrast, IscB is smaller than existing Cas proteins, greatly facilitating subsequent delivery and offering greater potential for modification.
[0007] In conclusion, given the limitations of currently available CRISPR / Cas systems, developing a more robust IscB system with superior performance in multiple aspects is of great significance to the development of biotechnology. Summary of the Invention
[0008] Through extensive experimentation and repeated exploration, the inventors of this application unexpectedly discovered a novel RNA-guided endonuclease protein, IscB. Based on this discovery, the inventors developed a new gene editing system, as well as gene editing methods and nucleic acid detection methods based on this system.
[0009] IscB protein
[0010] On the one hand, the present invention provides a novel IscB protein, which is referred to in the present invention as Cas-sf6003, Cas-sf6004, Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429 and Cas-sf6430, and the amino acid sequences of the above endonucleases are shown in any one of SEQ ID No. 1-20.
[0011] In one embodiment, the amino acid sequence of the IscB protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the sequences in SEQ ID No. 1-20, and substantially retains the biological function of the sequence from which it originated. Preferably, the IscB protein is derived from the same species as Cas-sf6003, Cas-sf6004, Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, and Cas-sf6430.
[0012] In one embodiment, the IscB protein is selected from any group I-III below:
[0013] The amino acid sequence of the IscB protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the sequences in SEQ ID No. 1-20, and substantially retains the biological function of the sequences from which it originated;
[0014] II. The amino acid sequence of the IscB protein, compared with any of the sequences in SEQ ID No. 1-20, has one or more amino acid substitutions, deletions or additions, and essentially retains the biological function of the sequence from which it originated.
[0015] III. The IscB protein contains any of the amino acid sequences shown in SEQ ID No. 1-20.
[0016] In one embodiment, the IscB protein is a protein of the OMEGA (Obligate Mobile Element Guided Activity) family or the Ω family, preferably an IscB protein, IsrB protein, IshB protein, or TnpB protein.
[0017] In one embodiment, the IscB protein comprises an HNH domain and a split RuvC domain, the HNH domain being located between the Ruv-CII and RuvC-III subdomains. In other embodiments, the IscB polypeptide comprises a split RuvC domain but does not contain an HNH domain. In still other embodiments, the IscB polypeptide comprises a split RuvC domain but does not contain an HNH domain.
[0018] In some embodiments, the IscB protein may further include an N-terminal PLMP domain and / or a conserved C-terminal domain.
[0019] In several embodiments, the IscB protein comprises about 200 to about 1000 amino acids.
[0020] Those skilled in the art will understand that the structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, i.e., replacing another nonpolar amino acid residue with a nonpolar amino acid residue, replacing another polar uncharged amino acid residue with a polar uncharged amino acid residue, replacing another basic amino acid residue with a basic amino acid residue, and replacing another acidic amino acid residue with an acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this invention, provided that the substitution does not lead to inactivation of the protein's biological activity. Therefore, the proteins of this invention can contain one or more conserved substitutions in their amino acid sequences, preferably generated by substitutions according to Table 1. Furthermore, this invention also covers proteins that also contain one or more other nonconservative substitutions, provided that such nonconservative substitutions do not significantly affect the desired function and biological activity of the proteins of this invention.
[0021] Conserved amino acid substitutions can occur at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing their biological activity, while “essential” amino acid residues are required for biological activity. A “conserved amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can occur in the non-conserved regions of the aforementioned IscB protein. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.
[0022] Table 1
[0023]
[0024] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C ends of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from the N and / or C ends of an IscB protein while retaining their desired functional activity are also within the scope of this invention. These alterations can include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.
[0025] It should be recognized that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of the aforementioned proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0026] Those skilled in the art will understand that these minor amino acid changes in the IscB protein of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0027] Those skilled in the art can identify the essential amino acids of the IscB protein of the present invention using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of the protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.
[0028] In one embodiment, the IscB protein contains any of the amino acid sequences shown in SEQ ID No. 1-20.
[0029] In one embodiment, the IscB protein is any of the amino acid sequences shown in SEQ ID No. 1-20.
[0030] In one embodiment, the IscB protein is a derivative protein with the same biological function as proteins having the sequences shown in any of SEQ ID Nos. 1-20.
[0031] The biological functions of the IscB protein include, but are not limited to, ωRNA binding activity, endonuclease activity, and the ability to bind to and cleave specific sites of a target sequence under ωRNA guidance.
[0032] The present invention also provides a fusion protein comprising the IscB protein as described above and other modified portions.
[0033] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.
[0034] In one embodiment, the modified portion is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), endonuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, cytidine deaminase, adenosine deaminase, methyltransferase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional releasing factor activity, histone modification activity, endonuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof. The NLS sequences are well known to those skilled in the art, and examples include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc, and TUS protein.
[0035] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection or tracing).
[0036] The reporter gene sequences are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0037] In one embodiment, the fusion protein of the present invention includes a domain capable of binding to DNA molecules or intracellular molecules, such as maltose-binding protein (MBP), the DNA-binding domain (DBD) of Lex A, the DBD of GAL4, etc.
[0038] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.
[0039] In one embodiment, the IscB protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.
[0040] In one embodiment, the modified portion is directly connected to the N-terminus or C-terminus of the IscB protein of the present invention.
[0041] In one embodiment, the modified portion is attached to the N-terminus or C-terminus of the IscB protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0042] The IscB protein, protein derivative, or fusion protein of the present invention is not limited by the manner of its production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0043] Nucleic acid of IscB protein
[0044] On the other hand, the present invention provides an isolated polynucleotide comprising:
[0045] (a) A multinucleotide sequence encoding the IscB protein or fusion protein of the present invention;
[0046] Alternatively, a polynucleotide complementary to the polynucleotide described in (a).
[0047] In one embodiment, the nucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the nucleotide sequence is codon-optimized for expression in eukaryotic cells.
[0048] In one embodiment, the cell is an animal cell, such as a mammalian cell.
[0049] In one embodiment, the cell is a human cell.
[0050] In one embodiment, the cell is a plant cell, such as the cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.
[0051] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0052] ωRNA
[0053] On the other hand, the present invention provides an ωRNA comprising a scaffold region and a guide sequence (i.e., a spacer sequence). The scaffold region of the ωRNA is capable of interacting with the IscB protein of the present invention, thereby enabling the IscB protein and the ωRNA to form a complex and guide the complex to bind to a target nucleic acid. For example, the ωRNAs of Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, or Cas-sf6430 comprise a scaffold region and a guide sequence.
[0054] In one implementation, the guide sequence is located at the 5' end of the skeleton region; the guide sequence is also known as the target sequence.
[0055] The guide sequence of the ωRNA described in this invention comprises a nucleotide sequence complementary to a sequence in the target nucleic acid. In other words, the guide sequence of the ωRNA described in this invention interacts with the target nucleic acid in a sequence-specific manner via hybridization (i.e., base pairing). Therefore, the guide sequence of the ωRNA can be altered or modified to hybridize with any desired sequence within the target nucleic acid. The nucleic acid is selected from DNA or RNA.
[0056] The percentage of complementarity between the guide sequence of the ωRNA and the target sequence of the target nucleic acid described in this invention may be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0057] In one implementation, the backbone region and guide sequence (i.e., spacer sequence) of the ωRNA can be designed as two separate molecules that can be hybridized or covalently linked into a single molecule.
[0058] Preferably, the backbone sequence of the ωRNA of the present invention is shown in any one of SEQ ID No. 23-40.
[0059] In one embodiment, the ωRNA includes tracrRNA and crRNA; the tracrRNA is capable of pairing with the pairing region of the crRNA to form a double strand; the crRNA also includes a region that hybridizes to a target sequence (i.e., a target sequence of a target nucleic acid). For example, the ωRNA of the Cas-sf6003 or Cas-sf6004 protein includes tracrRNA and crRNA.
[0060] In one embodiment, the ωRNA includes a tracrRNA sequence and a pairing region sequence between the crRNA and the tracrRNA.
[0061] In one embodiment, the pairing region sequences of the crRNA and tracrRNA are as shown in any of SEQ ID No. 41-42.
[0062] In one embodiment, the tracrRNA sequence is as shown in any of SEQ ID No. 21-22.
[0063] The guide sequence, guide region, or target sequence of the targeted nucleic acid of this invention comprises a nucleotide sequence complementary to a sequence in the target nucleic acid. In other words, the guide sequence, guide region, or target sequence of the targeted nucleic acid of this invention interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the guide sequence, guide region, or target sequence of the targeted nucleic acid can be altered or modified to hybridize with any desired sequence within the target nucleic acid.
[0064] The ωRNA of the present invention can form a complex with the IscB protein.
[0065] The ωRNA of the IscB protein of the present invention contains a targeting sequence that hybridizes with a target nucleic acid, wherein the target nucleic acid includes a sequence located at the 3' end of the adjacent motif (PAM) in the prototype spacer region.
[0066] carrier
[0067] The present invention also provides a carrier comprising the IscB protein as described above, isolated nucleic acid molecules or polynucleotides; preferably, it further comprises a regulatory element operatively linked thereto.
[0068] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.
[0069] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.
[0070] In some implementations, the vectors included in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated vectors, and herpes simplex vectors), and may also be plasmids, viruses, granules, bacteriophages, etc., which are well known to those skilled in the art.
[0071] IscB-based gene editing systems
[0072] This invention provides an engineered, non-naturally occurring vector system, or an IscB-based gene editing system, comprising an IscB protein or a nucleic acid sequence encoding the IscB protein and a nucleic acid encoding one or more ωRNAs.
[0073] The ωRNA can form a complex with the IscB protein.
[0074] In one embodiment, the nucleic acid sequence encoding the IscB protein and the nucleic acid encoding one or more ωRNAs are artificially synthesized.
[0075] In one embodiment, the nucleic acid sequence encoding the IscB protein and the nucleic acid encoding one or more ωRNAs do not coexist naturally.
[0076] The one or more ωRNAs target one or more target sequences in the cell. The one or more target sequences hybridize with the genomic loci of the DNA molecule encoding one or more gene products and guide the IscB protein to the genomic locus of the DNA molecule of the one or more gene products. After reaching the target sequence location, the IscB protein modifies, edits, or cuts the target sequence, thereby altering or modifying the expression of the one or more gene products.
[0077] The cells of this invention include one or more of animals, plants, or microorganisms.
[0078] In some embodiments, the IscB protein is codon-optimized for expression in cells.
[0079] In some embodiments, the IscB protein cleaves one or both strands at the target sequence location.
[0080] The present invention also provides an engineered, non-naturally occurring carrier system, which may include one or more carriers, the one or more carriers comprising:
[0081] a) A first regulatory element, which is operatively linked to ωRNA.
[0082] b) A second regulatory element operatively linked to the IscB protein;
[0083] The ωRNA can form a complex with the IscB protein;
[0084] Components (a) and (b) are located on the same or different carriers in the system.
[0085] The first and second regulatory elements include promoters (e.g., constitutive or inducible promoters), enhancers (e.g., 35S promoters or 35S enhanced promoters), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and polyU sequences).
[0086] In some embodiments, the vector in the system is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, adeno-associated vector, and herpes simplex vector), or it can be a plasmid, virus, granule, bacteriophage, or other type known to those skilled in the art.
[0087] In some embodiments, the system provided herein is a delivery system. In some embodiments, the delivery system is a nanoparticle, liposome, exosome, microbubble, or gene gun.
[0088] In one embodiment, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In another embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0089] In one embodiment, the target sequence is present within the cell. In another embodiment, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0090] In one embodiment, the IscB protein is linked to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.
[0091] On the other hand, the present invention relates to an engineered IscB-based gene editing system, the system comprising the aforementioned IscB protein and one or more ωRNAs, wherein the ωRNA includes a backbone region and a guide sequence, and the IscB protein is capable of binding the ωRNA and targeting a target nucleic acid sequence complementary to the guide sequence.
[0092] Protein-nucleic acid complexes / compositions
[0093] On the other hand, the present invention provides a complex or composition comprising:
[0094] (i) Protein components selected from: the aforementioned IscB protein, derived proteins, or fusion proteins, and any combination thereof; and
[0095] (ii) A nucleic acid component comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a backbone region capable of binding to the IscB protein of the present invention.
[0096] The protein components and nucleic acid components combine to form a complex.
[0097] In one embodiment, the nucleic acid component is ωRNA from the IscB system.
[0098] In one embodiment, the complex or composition is non-natural or modified. In one embodiment, at least one component of the complex or composition is non-natural or modified. In one embodiment, the first component is non-natural or modified; and / or, the second component is non-natural or modified.
[0099] Activated IscB complex
[0100] On the other hand, the present invention also provides an activated IscB complex comprising: (1) a protein component selected from: the IscB protein, derivatized protein, or fusion protein of the present invention, and any combination thereof; (2) ωRNA comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a backbone region capable of binding to the IscB protein of the present invention; and (3) a target sequence bound to the ωRNA. Preferably, the binding is a binding between the target nucleic acid and the target nucleic acid via the target sequence (or guide sequence) on the ωRNA.
[0101] The term "activated IscB complex," "activated complex," or "ternary complex" used in this article refers to the complex formed by the binding or modification of the IscB protein, ωRNA, and target nucleic acid in the IscB system.
[0102] The IscB protein and ωRNA of this invention can form a binary complex, which is activated upon binding to a nucleic acid substrate to form an activated IscB complex. This nucleic acid substrate is complementary to the spacer sequence (or guide sequence for hybridization with the target nucleic acid) in the ωRNA. In some embodiments, the spacer sequence of the ωRNA perfectly matches the target substrate. In other embodiments, the spacer sequence of the ωRNA partially (continuously or discontinuously) matches the target substrate.
[0103] Delivery and delivery composition
[0104] The IscB protein, ωRNA, fusion protein, nucleic acid molecule, vector, system, complex, and composition of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.
[0105] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery carrier, and selected from one or more of the following: the IscB protein, fusion protein, nucleic acid molecule, carrier, system, complex, and composition of the present invention.
[0106] In one embodiment, the delivery carrier is a particle.
[0107] In one embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0108] host cells
[0109] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the IscB protein of the present invention, a fusion protein, a nucleic acid molecule, a protein-nucleic acid complex, an activated IscB complex, a vector, or the delivery composition of the present invention.
[0110] In some implementations, the cell is a prokaryotic cell.
[0111] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (such as rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (such as chickens), fish, or crustaceans (such as clams or shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).
[0112] In some implementations, the cell is a stem cell or stem cell line.
[0113] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.
[0114] Gene editing methods and applications
[0115] The IscB protein, nucleic acid, the above-described composition, the above-described IscB-based gene editing system, the above-described vector system, the above-described delivery composition, or the above-described activated IscB complex or the above-described host cell of the present invention can be used for any or more of the following purposes: targeting and / or editing target nucleic acids; cleaving double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specifically cleaving and / or degrading side-branched nucleic acids; non-specifically cleaving single-stranded nucleic acids; nucleic acid detection; detecting nucleic acids in target samples; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids. In other embodiments, they can also be used to prepare reagents or kits for any or more of the above purposes.
[0116] The present invention also provides the use of the above-mentioned IscB protein, nucleic acid, the above-mentioned composition, the above-mentioned IscB-based gene editing system, the above-mentioned vector system, the above-mentioned delivery composition, or the above-mentioned activated IscB complex in gene editing, gene targeting, or gene cutting; or, in the preparation of reagents or kits for gene editing, gene targeting, or gene cutting.
[0117] In one embodiment, the gene editing, gene targeting, or gene cutting is performed intracellularly and / or extracellularly.
[0118] The present invention also provides a method for editing, targeting, or cleaving a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned IscB protein, nucleic acid, the aforementioned composition, the aforementioned IscB-based gene editing system, the aforementioned vector system, the aforementioned delivery composition, the aforementioned activated IscB complex, or the aforementioned host cell. In one embodiment, the method comprises editing, targeting, or cleaving the target nucleic acid intracellularly or extracellularly.
[0119] The gene editing or editing of target nucleic acids includes modifying genes, knocking out genes, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides, and gene mutations.
[0120] The editing can be performed in prokaryotic and / or eukaryotic cells.
[0121] On the other hand, the present invention also provides a kit for gene editing, gene targeting or gene cutting, the kit comprising the above-mentioned IscB protein, ωRNA, nucleic acid, the above-mentioned composition, the above-mentioned IscB-based gene editing system, the above-mentioned vector system, the above-mentioned delivery composition, the above-mentioned activated IscB complex or the above-mentioned host cell.
[0122] On the other hand, the invention provides the use of the above-mentioned IscB protein, nucleic acid, the above-mentioned composition, the above-mentioned IscB-based gene editing system, the above-mentioned vector system, the above-mentioned delivery composition, the above-mentioned activated IscB complex, or the above-mentioned host cell in the preparation of formulations or kits, wherein the formulations or kits are used for:
[0123] (i) Gene or genome editing;
[0124] (ii) Target nucleic acid detection and / or diagnosis;
[0125] (iii) Editing target sequences in target loci to modify biological or non-human organisms;
[0126] (iv) Treatment of the disease;
[0127] (v) Target gene.
[0128] Preferably, the above-mentioned gene or genome editing is performed intracellularly or extracellularly.
[0129] Preferably, the target nucleic acid detection and / or diagnosis is performed in vitro.
[0130] Preferably, the treatment of the disease is to treat symptoms caused by defects in the target sequence at the target locus.
[0131] Methods for specifically modifying target nucleic acids
[0132] On the other hand, the present invention also provides a method for specifically modifying target nucleic acids, the method comprising: contacting the target nucleic acid with the above-mentioned IscB protein, nucleic acid, the above-mentioned composition, the above-mentioned IscB-based gene editing system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated IscB complex.
[0133] This specific modification can occur in vivo or in vitro.
[0134] This specific modification can occur either inside or outside the cell.
[0135] In some cases, the cells are selected from prokaryotic or eukaryotic cells, such as animal cells, plant cells, or microbial cells.
[0136] In one embodiment, the modification refers to a break in the target sequence, such as a single-strand / double-strand break in DNA or a single-strand break in RNA.
[0137] In some cases, the method further includes contacting the target nucleic acid with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target nucleic acid.
[0138] In one embodiment, the modification further includes inserting an editing template (e.g., a foreign nucleic acid) into the break.
[0139] In one embodiment, the method further includes contacting the editing template with the target nucleic acid or delivering it to a cell containing the target nucleic acid. In this embodiment, the method repairs the broken target gene by homologous recombination with a foreign template polynucleotide; in some embodiments, the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target gene; in other embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence.
[0140] Terminology Definition
[0141] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0142] CRISPR system
[0143] As used herein, the terms “regularly clustered short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements relating to the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes.
[0144] IscB protein
[0145] As used herein, the term "IscB protein" refers to a novel, small-sized RNA-guided endonuclease. Its homologues include IsrB and IshB proteins, and together with TnpB protein, it can be referred to as the OMEGA (Obligate Mobile Element Guided Activity) system or the Ω system. IscB protein is a CRISPR-associated protein containing one or more RNA-guided endonuclease domains capable of modifying target nucleic acids. The domains of IscB protein include multiple domains such as the splitting RuvC domains (RuvC-1, Ruv-C II, and Ruv-C III), HNH domain, PLMP domain, and bridged helix (BH) domain. Therefore, IscB protein can be called Cas IscB protein or Cas protein. Unlike Cas9 protein, the IscB polypeptide contains a PLMP domain.
[0146] In one exemplary embodiment, the IscB protein, moving from the N-terminus to the C-terminus, comprises a PLMP domain, a RuvCI domain, a bridged helical BH domain, a RuvCII domain, an HNH domain, a RuvCIII domain, and a C-terminal domain.
[0147] In one exemplary embodiment, the amino acid sequence of the IscB protein is shown in any of SEQ ID No. 1-20.
[0148] IscB complex
[0149] As used herein, the term "IscB complex" refers to a complex formed by the binding of ωRNA and IscB protein, which includes a guide sequence that hybridizes to the target sequence and a backbone region that binds to the IscB protein. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the ωRNA.
[0150] ωRNA
[0151] As used herein, ωRNAs may contain a scaffold and a guide sequence (or spacer sequence). For example, the ωRNAs of Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, or Cas-sf6430 include a scaffold and a guide sequence.
[0152] ωRNA can include tracrRNA and crRNA; tracrRNA can pair with the pairing region of crRNA to form a double strand; crRNA also includes a region that hybridizes to the target sequence (i.e., the target sequence of the target nucleic acid). For example, the ωRNA of Cas-sf6003 or Cas-sf6004 proteins includes both tracrRNA and crRNA.
[0153] In some cases, the guide sequence or target sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence to hybridize with the target sequence and guide the IscB complex to specifically bind to the target sequence. In one embodiment, when optimal alignment is achieved, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining the optimal alignment is within the capabilities of a person skilled in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.
[0154] In one exemplary embodiment, the backbone region sequence of ωRNA is shown in SEQ ID No. 23-40.
[0155] In one exemplary embodiment, the pairing region sequences of crRNA and tracrRNA are as shown in any of SEQ ID No. 41-42.
[0156] In one exemplary embodiment, the tracrRNA sequence is as shown in any of SEQ ID No. 21-22.
[0157] target sequence
[0158] A "target sequence" refers to a polynucleotide targeted by a guide sequence in ωRNA, such as a sequence complementary to that guide sequence, where hybridization between the target and guide sequences will promote the formation of an IscB complex (including the IscB protein and ωRNA). Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of an IscB complex.
[0159] The target sequence can contain any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. The sequence or template that can be used for recombination into a target locus containing the target sequence is referred to as an "edit template," "edit polynucleotide," or "edit sequence." In one embodiment, the edit template is a foreign nucleic acid. In one embodiment, the recombination is homologous recombination.
[0160] In this invention, the "target sequence," "target polynucleotide," or "target nucleic acid" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).
[0161] wild type
[0162] As used herein, the term “wildtype” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, is separable from its natural source and has not been intentionally modified by humans.
[0163] Derivatization
[0164] As used herein, the term "derivation" refers to the chemical modification of an amino acid, polypeptide, or protein in which one or more substituents are covalently linked to the amino acid, polypeptide, or protein. Substituents may also be referred to as side chains.
[0165] A derivatized protein is a derivative of the original protein. Generally, the derivatization of a protein does not adversely affect its desired activity (e.g., activity to bind to guide RNA, endonuclease activity, activity to bind to and cleave a target sequence at a specific site under the guidance of guide RNA). In other words, the derivative of a protein has the same activity as the original protein.
[0166] Derivatized proteins
[0167] Also known as "protein derivatives," these are modified forms of proteins, where one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted.
[0168] Not naturally occurring
[0169] As used herein, the terms “non-naturally occurring” or “engineered” are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they indicate that the nucleic acid molecule or peptide is at least substantially free from at least one other component bound to it, either naturally occurring or found in nature.
[0170] Orthologue (ortholog)
[0171] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein, as described herein, refers to a protein belonging to a different species that performs the same or similar function as the protein that is its orthologue.
[0172] identity
[0173] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomeric subunit occupies the same location (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum identity. Such comparisons can be made using methods readily available, for example, computer programs such as the Align program (DNAstar, Inc.) Needleman et al. (1970) J. Mol. Biol. 48:443-453. The percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) integrated into the ALIGN program (version 2.0), which uses a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the Needleman and Wunsch algorithm (J MoI Biol. 48:444-453 (1970)) in the GAP program integrated into the GCG software package (available at www.gcg.com), which uses a Blossum 62 matrix or a PAM250 matrix, along with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.
[0174] carrier
[0175] The term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication initiation site.
[0176] One type of vector is a "plasmid," which is a circular double-stranded DNA loop into which another DNA fragment can be inserted, for example, using standard molecular cloning techniques.
[0177] Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package the virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by the virus used for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0178] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome upon introduction and thereby replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors."
[0179] host cells
[0180] As used herein, the term “host cell” refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0181] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0182] Control element
[0183] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed description can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0184] promoter
[0185] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.
[0186] NLS
[0187] A “nuclear localization signal” or “nuclear localization sequence” (NLS) is an amino acid sequence that “tags” a protein to allow it to be transported to the nucleus via nuclear transport; that is, a protein with an NLS is transported to the nucleus. Typically, an NLS contains positively charged Lys or Arg residues exposed on the protein surface. Exemplary nuclear localization sequences include, but are not limited to, NLS from the following: SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS contains the PKKKRKV sequence. In some embodiments, the NLS contains the AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS contains the PAAKRVKLD sequence. In some embodiments, the NLS contains the MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS contains the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the KIPIK sequence in the yeast transcriptional repressor Matα2, and PY-NLS.
[0188] Operable connection
[0189] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0190] Complementarity
[0191] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0192] Strict conditions
[0193] As used herein, “strict conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence.
[0194] Hybridization
[0195] The terms “hybridization” or “complementary” or “substantially complementary” refer to nucleic acids (such as RNA, DNA) containing nucleotide sequences that enable them to bind non-covalently, that is, to form base pairs and / or G / U base pairs with another nucleic acid in a sequence-specific, antiparallel manner (i.e., nucleic acid-specific binding of complementary nucleic acids), also known as “annealing” or “hybridization”.
[0196] Hybridization requires two nucleic acids to contain complementary sequences, although mismatches between bases are possible. Suitable conditions for hybridization between two nucleic acids depend on their length and degree of complementarity, variables well known in the art. Typically, hybridizable nucleic acids are 8 nucleotides or longer (e.g., 10 nucleotides or longer, 12 nucleotides or longer, 15 nucleotides or longer, 20 nucleotides or longer, 22 nucleotides or longer, 25 nucleotides or longer, or 30 nucleotides or longer).
[0197] It should be understood that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid for specific hybridization. The polynucleotide may contain 60% or higher, 65% or higher, 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 98% or higher, 99% or higher, 99.5% or higher, or have 100% sequence complementarity with the target region of the target nucleic acid sequence it hybridizes with.
[0198] Hybridization of the target sequence with ωRNA means that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and ωRNA can hybridize to form a complex; or it means that at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases of the nucleic acid sequences of the target sequence and ωRNA can be complementary and hybridize to form a complex.
[0199] Express
[0200] As used herein, the term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0201] connector
[0202] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of multiple amino acid residues via peptide bonds. The linkers of this invention can be synthetically produced amino acid sequences or naturally occurring polypeptide sequences, such as polypeptides with hinge region functions. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ et al. (1994) Structure2:1121-1123).
[0203] treat
[0204] As used in this article, the term "treatment" means to treat or cure a disease, to delay the onset of symptoms of a disease, and / or to slow the progression of a disease.
[0205] Subjects
[0206] As used herein, the term “subject” includes, but is not limited to, various animals, plants and microorganisms.
[0207] animal
[0208] For example, mammals, such as bovids, equines, sheep, suidae, canines, felines, lagos, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect).
[0209] plant
[0210] The term "plant" should be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, particularly monocotyledonous or dicotyledonous plants, vegetable crops including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce (e.g., head lettuce, leaf lettuce, longleaf lettuce), bok choy, taro, cucurbits (e.g., melons, watermelons, crenshaw, cantaloupes, Roman melons), rapeseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, headless cabbage, Chinese cabbage, bok choy), artichokes, carrots, napa cabbage, okra, onions, celery, parsley, chickpeas, parsnip, chicory, peppers, potatoes, gourds (e.g., zucchini, cucumbers, baby zucchini, squash, pumpkin), radishes, dried artichokes, etc. Onions, turnips, purple eggplant (also known as eggplant), ginseng, lettuce, scallions, chicory, garlic, spinach, green onions, squash, leafy greens, beets (sugar beets and fodder beets), sweet potatoes, romaine lettuce, wasabi, tomatoes, turnips, and spices; fruits and / or vine crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus fruits, blueberries, boysenberry. y), cranberries, currants, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pears, melons, mangoes, papayas, and lychees; field crops such as clover, alfalfa, evening primrose, miscanthus, corn / maize (feed corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain cereals (barley, oats, rye, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans). Oil-bearing plants (rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plants, cocoa beans, peanuts), Arabidopsis, fiber plants (cotton, flax, hemp, jute), Lauraceae (cinnamon, camphor), or a plant such as coffee, sugarcane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broadleaf trees and evergreen trees, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, as well as shrubs and other seedlings.
[0211] Beneficial effects of the invention
[0212] This invention provides a novel IscB protein. Blast results show that the IscB protein of this application has low similarity to previously reported IscB proteins, belonging to a novel Cas protein with broad application prospects.
[0213] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0214] Figure 1 PAM structure of IscB protein. Detailed Implementation
[0215] The following examples are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the examples are generally performed according to conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in this invention can be found in Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd edition (1989); *Current Protocols in Molecular Biology* (edited by FM. Ausubel et al., (1987)); *Methods in Enzymology* series (academic publishing company): *PCR 2: A PRACTICAL*. APPROACH (edited by MJ MacPherson, BD Hames and GR Taylor (1995)), Harlow and Lane (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (edited by R.R. Freshney (1987)).
[0216] Those skilled in the art will appreciate that the embodiments described herein are by way of example only and are not intended to limit the scope of protection claimed herein. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.
[0217] Example 1. Obtaining IscB protein
[0218] The inventors analyzed the metagenomics of uncultured organisms and identified 20 new IscB proteins through redundancy removal and protein clustering analysis. Blast results showed that the IscB protein had low sequence identity with previously reported IscB proteins. In this invention, these proteins were named Cas-sf6003, Cas-sf6004, Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, and Cas-sf6430. The amino acid sequences of these proteins are shown in Table 2 below. Cas-sf6003 and Cas-sf6430... The tracrRNA sequences of the ωRNA of Cas-sf6004 are shown in Table 3. The ωRNA backbone sequences of the proteins Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, Cas-sf6413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, and Cas-sf6430 are shown in Table 4. The pairing sequences of the crRNA and tracrRNA corresponding to the proteins Cas-sf6003 and Cas-sf6004 are shown in Table 5.
[0219] Table 2. Amino acid sequence of IscB protein
[0220]
[0221]
[0222]
[0223]
[0224] Table 3. TracrRNA sequences of ωRNAs corresponding to Cas-sf6003 and Cas-sf6004 proteins
[0225]
[0226] Table 4. ωRNA backbone regions corresponding to IscB protein
[0227]
[0228]
[0229] Table 5. Pairing region sequences of crRNA and tracrRNA corresponding to Cas-sf6003 and Cas-sf6004 proteins
[0230]
[0231] Example 2. PAM identification of IscB protein
[0232] The expression plasmid for the IscB protein (or Cas protein) in Example 1 was constructed as follows: After codon optimization of its nucleic acid sequence using E. coli, the gene was synthesized and ligated into the E. coli expression vector PeT28(a)+. Simultaneously, the JM23119 promoter was added to initiate the transcription of the ωRNA of the Cas protein, forming the vector: PeT28(a)+-Cas-JM23119-ωRNA. The target sequence is GCCCCCAGCGCTTCAGCGTTC, Cas-sf6401, Cas-sf6407, Cas-sf6411, Cas-sf6412, and Cas-sf6406. The backbone sequences of ωRNAs 413, Cas-sf6416, Cas-sf6417, Cas-sf6418, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, and Cas-sf6430 are detailed in Table 4. The ωRNA sequences of Cas-sf6003 and Cas-sf6004 include the paired region sequences in Table 5 and the tracrRNA sequences in Table 3. Construction of the PAM library: Synthetic sequence
[0233] CGTGTTTCGTAAAGTCTGGAAACGCGGAA GCCCCCAGCGCTTCAGCGTTC NNNNNNTCCCCTACGTGCTGCTGAAGTTGCCCGCAA, where N is a random deoxyribonucleotide and the underlined sequence is the target sequence. After being filled with Klenow enzyme, it was ligated into the pcyc184 vector. Following transformation into E. coli, plasmids were extracted to form a PAM library.
[0234] PAM library reduction experiment: The expression vector PeT28(a)+Cas-JM23119-ωRNA was co-transformed with the PAM library plasmid into competent BL21(DE3) cells. The cells were plated on LB agar plates containing kanamycin and chloramphenicol and incubated overnight at 37°C. The bacterial cells were then collected, and the bacterial concentration was adjusted to OD600 of 0.6-0.8. 0.2 mM IPTG was added, and the cells were induced at 37°C for 4 h. Plasmid extraction was performed using the FastPure EndoFree Plasmid Maxi Kit (vazyme) to obtain the reduced PAM library. Primers: PAM-F: GGTCTTCGGTTTCCGTGTT; PAM-R: TGGCGTTGACTCTCAGTCAT. PCR was performed using 30 ng / μL of the plasmid (PAM library) as a template to obtain the control group samples, and PCR was performed using 30 ng / μL of the plasmid (reduced PAM library) as a template to obtain the experimental group samples. Control group and experimental group samples were sent for next-generation sequencing for data analysis. For 4096 PAM sequences, the frequency of occurrence in the experimental and control groups was counted, and the sequences were standardized using the total number of PAM sequences in each group. The PAM sequences were then analyzed and plotted. Figure 1 As shown in Table 6, the PAM preference of IscB protein was found, where R represents A+G; Y represents C+T; M represents A+C; K represents G+T; S represents C+G; W represents A+T; H represents A+C+T; V represents A+C+G; and N represents A+C+G+T.
[0235] Table 6. PAM preference of IscB protein
[0236]
[0237] Example 3. Editing efficiency of IscB protein in animal cells
[0238] The gene editing activity of proteins Cas-sf6004, Cas-sf6412, Cas-sf6417, Cas-sf6419, Cas-sf6420, Cas-sf6422, Cas-sf6423, Cas-sf6425, Cas-sf6426, Cas-sf6427, Cas-sf6428, Cas-sf6429, and Cas-sf6430 was verified in animal cells. The vector pcDNA3.3 was modified to carry EGFP fluorescent protein. An SV40 NLS-Cas-NLS fusion protein was inserted via the BsmB1 restriction site and initiated by the MV promoter. The IscB protein and the GFP protein were linked using the linker peptide T2A. A U6 promoter and ωRNA sequence were inserted via the Mfe1 restriction site. The ωRNA sequence was the same as in Example 2. Target sequence: GCAACTTCAGCAGCACGTAGGGGAThe intracellular editing vector was obtained as: pcDNA3.3 CS2.0-Flag-Cas-NLS-T2A-ECFP-ωRNA.
[0239] Plating: 293T cells were plated when the confluence reached 70-80%, with a cell number of 8*10^4 cells / well in a 12-well plate.
[0240] Transfection: Transfect after 12-24 hours of cell plating. Add 2 µg of plasmid to 100 µL opti-MEM and mix well. Add 4 µL TransIntro® EL Transfection Reagent (TRAN) to the diluted plasmid and incubate at room temperature for 15-20 minutes. Add the incubated mixture to the culture medium containing cells for transfection. Replace with normal culture medium after 24 hours of transfection. Perform flow cytometry analysis 48 hours after transfection. The editing efficiency of IscB protein is shown in Table 7 below.
[0241] Table 7. Editing efficiency of IscB protein in animal cells
[0242]
[0243] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.
Claims
1. An IscB, characterized in that, The IscB protein is any one of the IscB proteins described in I-III below: I. The amino acid sequence of IscB protein is consistent with SEQ ID No.9, SEQ ID No.20, SEQ ID No.12, SEQ ID No.17, SEQ ID No.14, SEQ ID No.6, SEQ ID No.19, SEQ ID No.13, SEQ ID No.15, SEQ ID No.18, SEQ ID No.11, SEQ ID No.16, SEQ ID No.1, SEQ ID No.3, SEQ ID No.4, SEQ IDNo.5, SEQ ID No.7, SEQ ID No.8 or SEQ ID Compared with any sequence of No. 10, it has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity, and substantially retains the biological function of the sequence from which it originated; II. The amino acid sequence of the IscB protein, compared with any one of the sequences SEQ ID No. 9, SEQ ID No. 20, SEQ ID No. 12, SEQ ID No. 17, SEQ ID No. 14, SEQ ID No. 6, SEQ ID No. 19, SEQ ID No. 13, SEQ ID No. 15, SEQ ID No. 18, SEQ ID No. 11, SEQ ID No. 16, SEQ ID No. 1, SEQ ID No. 3, SEQ ID No. 4, SEQ ID No. 5, SEQ ID No. 7, SEQ ID No. 8, or SEQ ID No. 10, has one or more amino acid substitutions, deletions, or additions, and substantially retains the biological function of its derived sequence; III. The IscB protein comprises any of the amino acid sequences shown in SEQ ID No. 9, SEQ ID No. 20, SEQ ID No. 12, SEQ ID No. 17, SEQ ID No. 14, SEQ ID No. 6, SEQ ID No. 19, SEQ ID No. 13, SEQ ID No. 15, SEQ ID No. 18, SEQ ID No. 11, SEQ ID No. 16, SEQ ID No. 1, SEQ ID No. 3, SEQ ID No. 4, SEQ ID No. 5, SEQ ID No. 7, SEQ ID No. 8, or SEQ ID No.
10.
2. A fusion protein comprising the IscB protein of claim 1 and other modified portions.
3. An isolated polynucleotide, characterized in that, The polynucleotide is a polynucleotide sequence encoding the IscB protein of claim 1, or a polynucleotide sequence encoding the fusion protein of claim 2.
4. A carrier, characterized in that, The vector comprises the polynucleotide of claim 3 and a regulatory element operatively linked thereto.
5. A gene editing system based on IscB protein, characterized in that, The system includes the IscB protein of claim 1 and at least one ωRNA capable of binding to the IscB protein, the ωRNA including a region that binds to the IscB protein of claim 1 and a guide sequence for targeting nucleic acids.
6. A composition, characterized in that, The composition comprises: (i) A protein component selected from: the IscB protein of claim 1 or the fusion protein of claim 2; (ii) A nucleic acid component, which is ωRNA, said ωRNA being capable of binding the IscB protein of claim 1; The protein components and nucleic acid components combine to form a complex.
7. An engineered host cell, characterized in that, The host cell comprises the IscB protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the gene editing system based on the IscB protein of claim 5, or the composition of claim 6.
8. The IscB protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the gene editing system based on IscB protein of claim 5, or the composition of claim 6, or the host cell of claim 7, in gene editing, gene targeting, gene cutting, cutting of double-stranded DNA, single-stranded DNA or single-stranded RNA, non-specific cutting and / or degradation of side-branched nucleic acids, non-specific cutting of single-stranded nucleic acids, nucleic acid detection, specific editing of double-stranded nucleic acids, base editing of double-stranded nucleic acids, and base editing of single-stranded nucleic acids; Alternatively, in the preparation of formulations or kits for: gene editing, gene targeting, gene cleavage, cleavage of double-stranded DNA, single-stranded DNA or single-stranded RNA, nonspecific cleavage and / or degradation of side-branched nucleic acids, nonspecific cleavage of single-stranded nucleic acids, nucleic acid detection, specific editing of double-stranded nucleic acids, base editing of double-stranded nucleic acids, and base editing of single-stranded nucleic acids.
9. A method for editing, targeting, or cleaving a target nucleic acid, the method comprising contacting the target nucleic acid with the IscB protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the gene editing system based on the IscB protein of claim 5, or the composition of claim 6, or the host cell of claim 7.
10. A kit for gene editing, gene targeting, or gene cutting, the kit comprising the IscB protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the IscB protein-based gene editing system of claim 5, or the composition of claim 6, or the host cell of claim 7.