A mini crisper-cas system and applications thereof

By developing the VbCasX protein, a small CRISPR-Cas system, the problems of PAM sequence recognition, off-target effects, and immune reactivity in gene editing of CRISPR-Cas systems have been solved, enabling more efficient and safer gene editing.

CN118792281BActive Publication Date: 2025-11-25TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310381197.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-11-25
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems have limitations in gene editing, including limited PAM sequence recognition range, off-target effects, system size limitations on cell delivery, and immunoreactivity, which restrict their application scope and safety.

Method used

A small CRISPR-Cas system was developed. The VbCasX protein, consisting of only 872 amino acids, recognizes the 5'-TTR-3' PAM sequence, is derived from the non-pathogenic bacterium Verrucomicrobia bacterium, has lower immunogenicity, and can be targeted for editing by forming a guide RNA complex.

Benefits of technology

It expands the editing activity window of the CRISPR-Cas system, improves the efficiency and safety of gene editing, reduces the risk of immune response, and adapts to various vector delivery methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118792281B_ABST
    Figure CN118792281B_ABST
Patent Text Reader

Abstract

The application discloses a small CRISPR-Cas system and application thereof. The application provides a small CRISPR-Cas system (i) VbCasX protein and (ii) VbCasX guide RNA, the VbCasX guide RNA forms a complex with the VbCasX protein, and the VbCasX guide RNA comprises a guide sequence hybridized with a target sequence in a target nucleic acid; wherein the protospacer adjacent motif recognized by the VbCasX protein comprises 5'-TTR-3', wherein R is G or A. The Cas protein in the system provided by the application is smaller than the currently known commonly used Cas family protein, and the PAM sequence recognized by the system is different from the currently known commonly used CRISPR-Cas system, thereby expanding the editing activity window of the CRISPR-Cas system in the genome.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of gene editing, and relates to a small CRISPR-Cas system and application thereof. BACKGROUND

[0002] Clustered Regularly Interspaced Short Palindromic Repeats-CRISPR associated proteins (CRISPR-Cas) system is an adaptive immune system existing in bacteria and archaea, which defends against the invasion of foreign nucleic acids through three stages of adaptation, expression and interference. People use this feature to develop CRISPR-Cas system as a powerful gene editing tool, which is successfully applied to gene expression regulation, gene function exploration, nucleic acid detection and other related scientific research. However, there are some limitations when CRISPR-Cas system is used for genome editing, such as the constraint of protospacer adjacent motif (PAM) on target genes, off-target effects, the limitation of system size on cell delivery and the immunoreactivity of the system, which hinder the further expansion of its application range. The specific description is as follows:

[0003] 1.1 PAM constraint of CRISPR-Cas system

[0004] In recent years, CRISPR-Cas system has been widely used in gene editing of human cells, animals, plants and other organisms. In order to realize the editing of CRISPR-Cas system on multiple gene sites, Cas proteins that can recognize different PAM sequences are needed, and PAM sequences need to be recognized by specific Cas proteins to guide and activate CRISPR-Cas system. However, the existing SpyCas9 can only recognize PAM sequence of 5'-NGG-3', and AsCas12a can only recognize PAM sequence of 5'-TTTN-3'. Since the known types of CRISPR-Cas system are limited, other more types of PAM sequences still cannot be recognized, which restricts the application range of CRISPR-Cas system. Therefore, in order to meet the demand of gene editing at different sites, more new Cas proteins that can recognize different PAM sequences need to be excavated, and the recognition range of PAM sequence of CRISPR-Cas system needs to be expanded, so as to alleviate the PAM constraint.

[0005] 1.2 Off-target effect of CRISPR-Cas system

[0006] Off-target effect is one of the side effects of CRISPR-Cas gene editing technology. When editing a gene, unnecessary DNA mutations may occur at non-target locations in the genome, i.e. the system misrecognizes the target gene that needs to be edited. The current CRISPR-Cas gene editing technology has a relatively high probability of off-targeting. Once off-targeting occurs, it may cause unpredictable genetic variations, which may be inherited from generation to generation and cause uncontrollable harmful variations.

[0007] There are mainly two reasons for the off-target effect of the CRISPR-Cas system, i.e. the off-target effect caused by the mismatch between the single guide RNA (sgRNA) and the non-target DNA sequence and the off-target effect independent of sgRNA. Different Cas protein families have different off-target effects caused by mismatches, which have potential safety risks in application, such as causing many side effects including cancer, etc. At present, it is not yet ready to be used as a clinical technology for popularization and application. Therefore, it is an important goal of the CRISPR-Cas system research to study the reduction of the off-target effect of the CRISPR-Cas system and improve the safety.

[0008] 1.3 Limitation of CRISPR-Cas system size on cell delivery

[0009] To complete the gene editing of the CRISPR-Cas system on an organism, an important step is to use AAV, lentivirus or adenovirus as a carrier to introduce the DNA fragments related to the CRISPR-Cas system into cells to edit the target genes in the cells. However, these carriers have certain limitations on the length of DNA fragments (or the molecular weight of proteins), such as the carrying capacity of AAV is about 4.4 kb, and the carrying capacity of lentivirus and adenovirus is about 8 kb. The proteins of the commonly used CRISPR-Cas gene editing tools SpyCas9 and AsCas12a both exceed 1300 amino acids, i.e. the length of the DNA sequence for expressing the protein is more than 3.9 kb. If AAV is used for delivery, the length of the DNA sequence for expressing the protein will be limited, which will affect the delivery efficiency of the CRISPR-Cas system and further affect the efficiency of gene editing. In order to solve this problem, people usually simplify the Cas protein by deleting the redundant domains of the Cas protein, but this protein engineering method often causes a serious decline in the activity of the Cas protein. The use of double AAV packaging transfection is complex and inefficient. Therefore, in order to adapt to various delivery modes of the carrier, it is the key to break the size limitation of the CRISPR-Cas system on cell delivery to excavate compact and small Cas proteins.

[0010] 1.4 Immune reactivity of CRISPR-Cas system

[0011] Recently, several studies have reported that gene editing systems are immunoreactive. CRISPR-Cas systems derived from Staphylococcus aureus and Streptococcus pyogenes, such as SaCas9 and SpCas9, are susceptible to infection of organisms. Therefore, the organism may have produced an adaptive immune response in advance, and when a CRISPR-Cas system derived from the bacteria is used for gene editing, an immune response to the Cas protein may be produced, resulting in an impact on the gene editing ability of the system. Therefore, it is of great significance to discover and screen CRISPR-Cas systems derived from non-pathogenic bacteria.

[0012] In order to break through the above-mentioned limitations, it is an important strategy to mine and identify new CRISPR-Cas systems in order to perfect and develop CRISPR-Cas systems and expand the CRISPR-Cas toolbox. How to discover new, smaller and low-immunogenic Cas homologous proteins on the basis of existing gene editing tools and further study their molecular mechanisms is the key to solving the limitations of clinical treatment and other applications of CRISPR-Cas systems, and has important research significance and value. SUMMARY

[0013] Problems to be solved by the invention

[0014] Based on the above problems existing in the prior art, the present application provides a small CRISPR-Cas system, the Cas protein in the system has only 872 amino acids, which is smaller than the currently known commonly used Cas family proteins; and the PAM sequence recognized by the system is different from the currently known commonly used CRISPR-Cas systems, which expands the editing activity window of the CRISPR-Cas system in the genome; in addition, the system of the present application is derived from non-pathogenic bacteria, and has lower immunogenicity compared with the CRISPR-Cas system derived from human pathogenic bacteria such as Staphylococcus aureus (SaCas9) or Streptococcus pyogenes (SpCas9).

[0015] Solutions to the problems

[0016] The first aspect of the present application provides a small CRISPR-Cas system, which comprises:

[0017] (i) a VbCasX protein; and,

[0018] (ii) a VbCasX guide RNA, which forms a complex with the VbCasX protein, and the VbCasX guide RNA comprises a guide sequence that hybridizes to a target sequence in a target nucleic acid;

[0019] wherein the protospacer adjacent motif recognized by the VbCasX protein comprises 5'-TTR-3', wherein R is G or A.

[0020] In some embodiments, the VbCasX protein is derived from a Verrucomicrobia bacterium.

[0021] In some embodiments, the VbCasX protein comprises one or more of the following sequences:

[0022] 1) the amino acid sequence set forth in SEQ ID NO: 1;

[0023] 2) an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence set forth in SEQ ID NO: 1, and which retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1;

[0024] 3) an amino acid sequence that adds, substitutes, deletes, or inserts one or more amino acid residues in the amino acid sequence set forth in SEQ ID NO: 1, and which retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1;

[0025] 4) an amino acid sequence encoded by a nucleotide sequence that hybridizes to a polynucleotide sequence encoding the amino acid sequence set forth in SEQ ID NO: 1 under stringent conditions, and which retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1, the stringent conditions being medium stringency conditions, medium-high stringency conditions, high stringency conditions, or very high stringency conditions.

[0026] In some embodiments, the VbCasX guide RNA is a dual guide RNA;

[0027] Preferably, the VbCasX guide RNA comprises:

[0028] an activator RNA comprising a nucleotide sequence set forth in SEQ ID NO: 3 or a nucleotide sequence that is 80% or more identical to the sequence set forth in SEQ ID NO: 3, and,

[0029] a targeter RNA comprising a nucleotide sequence as set forth in SEQ ID NO: 4 or a nucleotide sequence having 80% or more identity to the sequence set forth in SEQ ID NO: 4, and the guide sequence.

[0030] In some preferred embodiments, the VbCasX guide RNA is a single guide RNA.

[0031] Preferably, the VbCasX guide RNA comprises a nucleotide sequence as set forth in SEQ ID NO: 5 or a nucleotide sequence having 80% or more identity to the sequence set forth in SEQ ID NO: 5, and the guide sequence.

[0032] In some embodiments, the mini CRISPR-Cas system further comprises: (iii) a donor polynucleotide.

[0033] A second aspect of the application provides a polynucleotide comprising one or more of:

[0034] i) a nucleotide sequence encoding a VbCasX protein in the mini CRISPR-Cas system as described in the first aspect of the application;

[0035] ii) a nucleotide sequence encoding a VbCasX guide RNA in the mini CRISPR-Cas system as described in the first aspect of the application; and,

[0036] iii) a donor polynucleotide sequence.

[0037] A third aspect of the application provides a vector comprising the polynucleotide as described in the second aspect of the application.

[0038] In some preferred embodiments, the vector is an expression vector.

[0039] A fourth aspect of the application provides a cell comprising one or more of:

[0040] (a) the mini CRISPR-Cas system as described in the first aspect of the application;

[0041] (b) the polynucleotide as described in the second aspect of the application; and,

[0042] (c) the vector as described in the third aspect of the application.

[0043] A fifth aspect of the application provides a kit comprising one or more of:

[0044] (A) the mini CRISPR-Cas system as described in the first aspect of the application;

[0045] (B) a polynucleotide as described in the second aspect of the application;

[0046] (C) a vector as described in the third aspect of the application; and,

[0047] (D) a cell as described in the fourth aspect of the application.

[0048] A sixth aspect of the application provides a method of modifying a target nucleic acid, the method comprising the step of contacting the target nucleic acid with a mini CRISPR-Cas system as described in the first aspect of the application.

[0049] A seventh aspect of the application provides use of a mini CRISPR-Cas system as described in the first aspect of the application, a polynucleotide as described in the second aspect of the application, a vector as described in the third aspect of the application, a cell as described in the fourth aspect of the application or a kit as described in the fifth aspect of the application in modifying a target nucleic acid.

[0050] Effects of the invention

[0051] Through implementation of the above technical solutions, the application can achieve the following technical effects:

[0052] The Cas protein in the mini CRISPR-VbCasX system provided by the application is composed of only 872 amino acids, which is smaller than the currently known PlmCasX protein with 986 amino acids, etc. The CRISPR-VbCasX system has stable double-stranded DNA cleavage ability in vitro, can target and edit the genome of mammalian cells, and is a new type of mini CRISPR-Cas system with mammalian cell editing activity. Moreover, the PlmCasX protein in the mini CRISPR-VbCasX system provided by the application can recognize the protospacer adjacent motif containing 5'-TTR-3', which also expands the editing activity window of the CRISPR-Cas system. Through reasonable design or directed evolution and other modification means, it is expected to be optimized into a more efficient gene editing tool, to make up for the shortcomings of the existing CRISPR-Cas system, and to further expand the application potential of the CRISPR-Cas system. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 The CRISPR-Cas gene sequence distribution diagram provided by the application.

[0054] Figure 2 The UV curve and SDS-PAGE result diagram of VbCasX through a heparin affinity column.

[0055] Figure 3 The PAM identification result diagram.

[0056] Figure 4 A schematic diagram of results for VbCasX to produce edits in the presence of a targeting sgRNA in mammalian cells (cells produce green fluorescent protein GFP).

[0057] Figure 5 A second-generation sequencing result for editing efficiency of VbCasX at different sites, T1, T2, T3 are different editing sites. DETAILED DESCRIPTION

[0058] For the purposes of the present application, certain technical and scientific terms are specifically defined below. Unless specifically defined herein, all other technical and scientific terms used in the present application have the meanings that are commonly understood by one of ordinary skill in the art in the field of the present application.

[0059] In the present specification, a numerical range expressed using "numerical value A ~ numerical value B" means a range including the end point numerical values A and B.

[0060] In the present specification, "substantially" or "essentially" means within 5%, preferably 3%, more preferably 1% of the standard deviation from a theoretical model or theoretical data.

[0061] In the present specification, the meaning of "may" includes both the meaning of performing a certain process and the meaning of not performing a certain process.

[0062] In the present specification, "optional" or "optionally" means that the event or circumstance described next can or can not occur, and the description includes the case where the event occurs and the case where the event does not occur.

[0063] In the present specification, "some specific / preferred embodiments", "other specific / preferred embodiments", "embodiments", and the like refer to the specific elements (e.g., features, structures, properties, and / or characteristics) described in relation to the embodiments are included in at least one embodiment described herein, and can or can not be present in other embodiments. In addition, it should be understood that the elements can be combined in various embodiments in any suitable manner.

[0064] In the present specification, the term "CRISPR" refers to Clustered Regularly Interspaced Short Palindromic Repeats, which is from the immune system of microorganisms.

[0065] In the present specification, the term "Cas protein" refers to CRISPR-associated protein, the Cas protein together with CRISPR sequence constitutes a CRISPR-Cas system, the Cas protein has a nuclease-associated functional domain, and cuts the target sequence at a specific position by recognizing PAM (protospacer adjacent motif).

[0066] In the present specification, the terms "polynucleotide" and "nucleic acid" are used interchangeably and refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, this term includes single-, double- or multi- stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.

[0067] In the present specification, "hybridizable" or "complementary" or "substantially complementary" means that a nucleic acid (e.g., RNA, DNA) comprises a nucleotide sequence that enables the nucleic acid to bind, in a sequence-specific, anti-parallel manner (i.e., the nucleic acid specifically binds to a complementary nucleic acid), non-covalently (i.e., forms Watson-Crick base pairs and / or G / U base pairs), "anneal" or "hybridize" to another nucleic acid under conditions of appropriate temperature and solution ionic strength under in vitro and / or in vivo conditions. Standard Watson-Crick base pairing includes: adenine (A) pairs with thymidine (T), adenine (A) pairs with uracil (U), and guanine (G) pairs with cytosine (C). In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization of a DNA molecule to an RNA molecule (e.g., when a DNA target nucleic acid base pairs with a guide RNA): guanine (G) can also pair with uracil (U). For example, in the case of base pairing between a tRNA anticodon and a codon in mRNA, G / U base pairing is at least partially responsible for the degeneracy of the genetic code.

[0068] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E. F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), in particular in Chapter 11 and Table 11.1 of this reference; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). The conditions of temperature and ionic strength determine the "stringency" of the hybridization.

[0069] In the present application, "moderate stringency", "moderate-high stringency", "high stringency" or "very high stringency" describe conditions for nucleic acid hybridization and washing. Guidance for the performance of hybridization reactions can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6, which is incorporated herein by reference. Both aqueous and nonaqueous methods are described in this reference and either can be used. For example, specific hybridization conditions are as follows: (1) low stringency hybridization conditions in 6x sodium chloride / sodium citrate (SSC) at about 45°C, followed by at least one wash in 0.2x SSC, 0.1% SDS at 50°C (for low stringency conditions, the wash temperature can be increased to 55°C); (2) moderate stringency hybridization conditions in 6x SSC at about 45°C, followed by one or more washes in 0.2x SSC, 0.1% SDS at 60°C; (3) high stringency hybridization conditions in 6x SSC at about 45°C, followed by one or more washes in 0.2x SSC, 0.1% SDS at 65°C and preferably; (4) very high stringency hybridization conditions are 0.5M sodium phosphate, 7% SDS at 65°C, followed by one or more washes in 0.2x SSC, 1% SDS at 65°C.

[0070] Hybridization requires that two nucleic acids contain complementary sequences, although mismatches between bases are possible. Conditions suitable for hybridization between two nucleic acids depend on the length and degree of complementarity of the nucleic acids, which are well known variables in the art. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) value of the hybrid of the nucleic acids having those sequences. For hybridization between nucleic acids having short stretches of complementarity (e.g., 35 or fewer, 30 or fewer, 25 or fewer, 22 or fewer, 20 or fewer, or 18 or fewer nucleotides of complementarity), the position of the mismatch can become important (see Sambrook et al., supra, 11.7-11.8). Generally, the length of the hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more). Temperature, wash solution salt concentration, and other conditions can be adjusted as needed, depending on factors such as the length and degree of complementarity of the region of complementarity.

[0071] In the present application, the terms "peptide," "polypeptide," and "protein" are used interchangeably herein and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemicals or biochemical modifications or derivatization of the amino acids, and polypeptides having modified peptide backbones.

[0072] "Bind" (e.g., RNA binding domain, bind to a target nucleic acid, etc.) as used herein refers to a non-covalent interaction between macromolecules (e.g., a non-covalent interaction between a protein and a nucleic acid; a non-covalent interaction between a Cas protein / guide RNA complex and a target nucleic acid; and the like). When in a non-covalent interaction state, the macromolecules are said to be "associated" or "interacting" or "bound" (e.g., when molecule X is said to be interacting with molecule Y, then it means that molecule X is bound to molecule Y in a non-covalent manner). Binding interactions are typically characterized by dissociation constants (Kd) of less than 10 D M, less than 10 -6 M, less than 10 - 7 M, less than 10 -8 M, less than 10 -9 M, less than 10 -10 M, less than 10 -11 M, less than 10 -12 M, less than 10 -13 M, less than 10 -14 M, or less than 10 -15 M. "Affinity" refers to the strength of binding, with increased binding affinity correlating to lower Kd. D M. "Affinity" refers to the strength of binding, with increased binding affinity correlating to lower Kd.

[0073] In the present application, the term "conservative amino acid substitution" refers to the interchangeability of amino acid residues in a protein that have similar side chains. For example, one group of amino acids that have aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids that have aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids that have amide-containing side chains consists of asparagine and glutamine; a group of amino acids that have aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids that have basic side chains consists of lysine, arginine, and histidine; a group of amino acids that have acidic side chains consists of glutamic acid and aspartic acid; and a group of amino acids that have sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine-glycine, and asparagine-glutamine.

[0074] In the present application, a polynucleotide or polypeptide has a certain percentage of "sequence identity" to another polynucleotide or polypeptide, which means that the percentage of bases or amino acids are the same when aligned, and are in the same relative position when comparing the two sequences. Sequence identity can be determined in many different ways. To determine sequence identity, sequences can be aligned using various convenient methods and computer programs (e.g., BLAST, T-COFFEE, MUSCLE, MAFFT T, etc.), which are available on the World Wide Web at websites including ncbi.nlm.nili.gov / BLAST, ebi.ac.uk / Tools / msa / tcoffee / , ebi.ac.uk / Tools / msa / muscle / , mafft.cbrc.jp / alignment / software / . See, e.g., Altschul et al. (1990), J. Mol. Bioi. 215:403-10.

[0075] In the present application, a DNA sequence that "encodes" a particular RNA is a sequence of DNA nucleotides that is transcribed into RNA. A DNA polynucleotide can encode RNA that is translated into a protein (mRNA) (so both the DNA and the mRNA encode the protein), or a DNA polynucleotide can encode RNA that is not translated into a protein (e.g., tRNA, rRNA, microRNA (miRNA), "non-coding" RNA (ncRNA), guide RNA, etc.).

[0076] In the present application, "protein coding sequence" or a sequence encoding a particular protein or polypeptide is a nucleotide sequence that is transcribed into mRNA (in the case of DNA) and translated into a polypeptide (in the case of mRNA) in vitro or in vivo when placed under the control of appropriate regulatory sequences.

[0077] In the present application, the term "naturally occurring" or "unmodified" or "wild type" applied to a nucleic acid, polypeptide, cell or organism refers to a nucleic acid, polypeptide, cell or organism that occurs in nature. For example, a polypeptide or polynucleotide sequence that occurs in an organism and that can be isolated from a source in nature is naturally occurring.

[0078] In the present invention, "recombinant" means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, polymerase chain reaction (PCR), and / or ligation steps that result in a construct having a structural coding or non-coding sequence that is distinguishable from endogenous nucleic acids present in native systems. DNA sequences encoding polypeptides can be assembled from cDNA fragments or from a series of synthetic oligonucleotides to provide synthetic nucleic acids capable of being expressed from recombinant transcriptional units contained in cells or in cell-free transcription and translation systems. Genomic DNA containing the relevant sequences can also be used in the formation of the recombinant gene or transcriptional unit. Sequences of non-translated DNA can be present at the 5' or 3' end of an open reading frame, where such sequences do not interfere with manipulation or expression of the coding region, and can in fact serve a function in regulating production of the desired product by various mechanisms (see "DNA Regulatory Sequences"). Alternatively, DNA sequences encoding untranslated RNA (e.g., guide RNA) can also be considered recombinant. Thus, for example, the term "recombinant" nucleic acid refers to a polynucleotide or nucleic acid that does not occur naturally as a contiguous sequence, made by the artificial combination of two otherwise separate segments of sequence, e.g., by the manipulation of isolated segments of nucleic acids (e.g., by genetic engineering techniques). Such manipulations are typically done to replace a codon with a codon that encodes the same amino acid, a conserved amino acid, or a non-conserved amino acid. Alternatively, such manipulations are done to join together nucleic acid segments that have a desired functional combination. Such manipulations are often done by chemical synthesis means or by the artificial manipulation of separate segments of nucleic acids (e.g., by genetic engineering techniques). When the recombinant polynucleotide encodes a polypeptide, the sequence of the encoded polypeptide can be naturally occurring ("wild type") or can be a variant (e.g., a mutant) of a naturally occurring sequence. One example of this is DNA (recombinant) that encodes a wild type protein, where the DNA sequence is codon optimized to express the protein in a cell (e.g., a eukaryotic cell) in which the protein does not naturally occur (e.g., to express a Cas protein in a eukaryotic cell). Thus, the codon optimized DNA can be recombinant and non-naturally occurring, while the protein encoded by the DNA can have a wild type amino acid sequence.

[0079] Thus, the term "recombinant" polypeptide does not necessarily refer to a polypeptide whose amino acid sequence does not occur in nature. Rather, a "recombinant" polypeptide is encoded by a recombinant non-naturally occurring DNA sequence, but the amino acid sequence of the polypeptide can be naturally occurring ("wild type") or non-naturally occurring (e.g., a variant, a mutant, etc.). Thus, a "recombinant" polypeptide is the result of artificial intervention, but can have a naturally occurring amino acid sequence.

[0080] In the present invention, a "vector" or "expression vector" is a replicon, such as a plasmid, a bacteriophage, a virus, an artificial chromosome, or a cosmid, to which another DNA segment (i.e., an "insert") can be attached so as to bring about the replication of the attached segment in a cell.

[0081] In the present invention, an "expression cassette" comprises a DNA coding sequence operably linked to a promoter. "Operably linked" means colocalized in such a way as to permit them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence (or the coding sequence can also be said to be operably linked to the promoter).

[0082] In the present invention, the term "recombinant expression vector" or "DNA construct" are used interchangeably herein to refer to a DNA molecule comprising a vector and an insert. Recombinant expression vectors are typically generated for the purpose of expressing and / or propagating one or more inserts, or for the purpose of constructing other recombinant nucleotide sequences. The one or more inserts can or can not be operably linked to a promoter sequence, and can or can not be operably linked to DNA regulatory sequences.

[0083] A cell is "genetically modified" or "transformed" or "transfected" with exogenous DNA or exogenous RNA, e.g., a recombinant expression vector, when such DNA is introduced inside the cell. The presence of the exogenous DNA results in a permanent or transient genetic change. The transforming DNA can or can not be integrated (covalently linked) to chromosome(s) of the cell. In, for example, prokaryotic, yeast, and mammalian cells, the transforming DNA can be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transformed DNA is integrated into a chromosome and replicated as part of a chromosome

[0084] Suitable methods of genetic modification (also referred to as "transformation") include, for example, viral or phage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethylenimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep 13. pii: S0169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.23), and the like.

[0085] The choice of genetic modification method generally depends on the type of cell to be transformed and the environment in which the transformation is taking place (e.g., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel et al., Short Protocols in Molecular Biology, 3rded., Wiley & Sons, 1995.

[0086] In the present invention, a "target nucleic acid" is a polynucleotide (e.g., DNA such as genomic DNA) that includes a site ("target site" or "target sequence") targeted by an RNA-guided endonuclease polypeptide (e.g., a Cas protein, etc.). A target sequence is a sequence with which a guide sequence of a guide RNA will hybridize. For example, a target site (or target sequence) within a target nucleic acid 5'-GAGCAUAUC-3' is targeted by (or bound by, or hybridizes to, or complementary to) the sequence 5'-GAUAUGCUC-3'. Suitable hybridization conditions include physiological conditions normally found in a cell. For a double-stranded target nucleic acid, the strand of the target nucleic acid that is complementary to and hybridizes to a guide RNA is referred to as the "complementary strand" or "target strand"; while the strand of the target nucleic acid that is complementary to the "target strand" (and thus not complementary to the guide RNA) is referred to as the "non-target strand" or "non-complementary strand".

[0087] In the present invention, "cleavage" means the breakage of the covalent backbone of a target nucleic acid molecule (e.g., RNA, DNA). Cleavage can be initiated by a variety of methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two distinct single-strand cleavage events.

[0088] In the present invention, "nuclease" and "endonuclease" are used interchangeably herein to mean an enzyme having catalytic activity for nucleic acid cleavage (e.g., ribonuclease activity (ribose nucleic acid cleavage), deoxyribonuclease activity (deoxyribose nucleic acid cleavage), etc.).

[0089] The following describes the technical solutions of the present invention in detail.

[0090] <CRISPR-Cas system>

[0091] The present application provides a novel small CRISPR-Cas system, which comprises a Cas protein and a corresponding guide RNA. In the present application, the Cas protein (e.g., a VbCasX protein) interacts (binds) with the corresponding guide RNA (e.g., a VbCasX guide RNA) to form a ribonucleoprotein (RNP) complex, which targets a specific site in a target nucleic acid through base pairing between the guide RNA and a target sequence within the target nucleic acid. The guide RNA comprises a nucleotide sequence (guide sequence) that is complementary to a sequence (target site / target sequence) of the target nucleic acid. Thus, in the present application, the VbCasX protein forms a complex with the VbCasX guide RNA, and the guide RNA provides sequence specificity to the RNP complex through the guide sequence. The VbCasX protein of the complex provides site-specific activity. In other words, the VbCasX protein is guided (e.g., stably positioned) to a target site within a target nucleic acid sequence (e.g., a chromosomal sequence or an extra-chromosomal sequence, such as a episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) due to its binding with the guide RNA.

[0092] In some embodiments, the present application provides a CRISPR-Cas system comprising: (i) a VbCasX protein (and / or a nucleic acid encoding the VbCasX protein) and (ii) a VbCasX guide RNA (and / or a nucleic acid encoding the VbCasX guide RNA) (e.g., wherein the VbCasX guide RNA can be in the form of a dual guide RNA or a single guide RNA). In other embodiments, the present application provides a CRISPR-Cas system further comprising: (iii) a donor polynucleotide.

[0093] In the present specification, the CRISPR-Cas system provided by the present application is also referred to as a CRISPR-VbCasX system, or a (based on) VbCasX protein system.

[0094] In other embodiments, the present application provides a nucleic acid / protein complex (RNP complex) comprising: (i) a VbCasX protein of the present application; and (ii) a VbCasX guide RNA (e.g., wherein the VbCasX guide RNA can be in the form of a dual guide RNA or a single guide RNA).

[0095] (VbCasX protein)

[0096] A VbCasX polypeptide (used interchangeably with the term “VbCasX protein”) can bind and / or modify (e.g., cleave) a target nucleic acid. In some embodiments, a VbCasX protein is a naturally occurring protein (e.g., naturally occurring in a prokaryotic cell). In other embodiments, a VbCasX protein is not a naturally occurring polypeptide (e.g., the VbCasX protein is a variant VbCasX protein, a chimeric protein, etc.).

[0097] An assay to determine whether a given protein interacts with a VbCasX guide RNA can be any convenient binding assay to test binding between a protein and a nucleic acid. Suitable binding assays (e.g., gel shift assays) are known to one of ordinary skill in the art (e.g., assays that include adding a VbCasX guide RNA and a protein to a target nucleic acid). An assay to determine whether a protein has activity (e.g., to determine whether a protein has nuclease activity to cleave a target nucleic acid) can be any convenient assay (e.g., any convenient nucleic acid cleavage assay to test cleavage of a nucleic acid). Suitable assays (e.g., cleavage assays) are known to one of ordinary skill in the art.

[0098] A VbCasX protein functions as an endonuclease that catalyzes a double-stranded break at a specific sequence targeted in double-stranded DNA (dsDNA). Sequence specificity is provided by an associated guide RNA that hybridizes to a target sequence within the target DNA.

[0099] In some embodiments, the VbCasX protein provided by the present application is (or is derived from) a naturally occurring (wild-type) protein that is derived from a Verrucomicrobia bacterium. Verrucomicrobia is a group of beneficial bacteria that is mainly found in aquatic and soil environments, or in human feces. Verrucomicrobia exists in the inner layer of the intestinal mucosa, and is abundant in healthy individuals. They can decompose polysaccharide substances, help the human body maintain glucose homeostasis, have anti-inflammatory properties, and can further help intestinal health. Therefore, compared to CRISPR-Cas systems derived from human pathogenic bacteria such as Staphylococcus aureus (SaCas9) or Streptococcus pyogenes (SpCas9), the VbCasX protein has lower immunogenicity.

[0100] In some specific embodiments, the VbCasX protein comprises: 1) an amino acid sequence as set forth in SEQ ID NO: 1.

[0101] The amino acid sequence of the VbCasX protein is as follows (SEQ ID NO: 1):

[0102] MNSVRKSLRELLTLGLKASSSASPQTITRTVKLGVEAKYRSDLATALHKHFDAYEEFRR

[0103] SVLDELEQWWNDDPDSFIQMVKCKKAEPYEDKSSCGAWLFSKFLTGKKLPEGLTNKA

[0104] GFALLDSLAGGLKSFITRRANVVKDIKQREKTNTSEWKKELTALAKELSEEMPEEAPEL

[0105] DFDNVDAAAIEGYNEWVALVRSWCNLILVQRHQLNRRDVPAPRYLKGYPGFPGSQRY

[0106] AENLPLKESLPLLREVTLENIKAAKPLFKAATDEQWQSILERFTVELVSGRARTARQTL

[0107] AQRLTYLVKENPGWAKEKIAKEALDGVMRGAKKLEEHLKNKGLTDSRAVIKLANLYN

[0108] VASVFAIEAIRASGDYVSYYETDTPRRVAFGELRGGLHQASDDTAAIQISGFSLSGDNPQ

[0109] YGGLLTYDPDEIGKEKWSLLYTLDNQAMKLVQPSEKAKGRGFLKSDLQGLAKTGRGD

[0110] EAQLLKGEVWLPSDKDKHPLSLPLQMGTRQGREYFWNFDRGLKNSDAWVLNNGRLL

[0111] RVMPAGRPDLAKFYVTITLGRQAPPIGNIKPKAFIGIDRGEAVPAAYAVIDTKGRLLKSG

[0112] LVQDEYREQQRRFNDKKRELQRQSGGYTKWLRSKERNRAKALGGDVSRELLDLAAK

[0113] HEAPLVMEYLSSGLVTRGGKNTMMSSMQYERVLGTLEQRLAEVGLYELVSDPKFRKG

[0114] DNGFIKLVGPAYTSSTCSECGQVFSTDFYEALTDTIVNSKDETWEVTLPTGKQLNLPEE

[0115] YTYWVRGQGEKTRETHDRLTELIKGKPIAKISKTNRRSLTGLLRGCLVPYRPRQAEFHC

[0116] PCCGYEANADEQAALNIARKLLFREELGDKVKEASESARRNTQKLWQDWYQKKLAK

[0117] VWRK

[0118] Nucleotide sequence encoding a VbCasX protein (SEQ ID NO: 2):

[0119] ATGAACTCCGTAAGAAAATCCCTTCGTGAATTGTTAACCCTCGGGCTGAAAGCCAG

[0120] TTCGTCTGCAAGCCCCCAGACAATTACCCGAACGGTCAAGTTGGGCGTAGAAGCAA

[0121] AGTACCGCTCAGATTTGGCTACAGCACTTCACAAACACTTCGACGCGTACGAAGAG

[0122] TTTCGACGCTCAGTATTAGATGAACTGGAGCAGTGGTGGAACGATGACCCAGATTC

[0123] ATTTATCCAGATGGTCAAATGTAAGAAGGCCGAACCGTACGAAGACAAATCAAGTT

[0124] GCGGGGCTTGGTTATTCAGTAAATTTCTTACGGGCAAAAAGCTCCCTGAAGGGTTA

[0125] ACCAACAAGGCAGGGTTTGCTCTCCTTGATAGTTTGGCTGGCGGCTTAAAGAGTTTT

[0126] ATCACGCGGCGAGCCAATGTGGTTAAAGACATTAAGCAGCGTGAAAAGACCAACA

[0127] CCAGCGAATGGAAAAAAGAACTGACCGCACTGGCAAAAGAACTCAGTGAGGAAAT

[0128] GCCGGAGGAAGCCCCAGAACTCGACTTTGACAATGTCGATGCCGCAGCAATTGAA

[0129] GGGTACAATGAATGGGTCGCCTTGGTTCGCTCGTGGTGCAATTTGATATTAGTACAA

[0130] CGTCACCAGCTTAACCGTAGAGATGTCCCTGCGCCCAGATACCTCAAGGGTTATCCC

[0131] GGTTTTCCCGGCTCACAACGTTACGCCGAAAACTTACCACTCAAAGAGTCACTACC

[0132] CCTACTCAGGGAAGTCACACTCGAAAATATCAAGGCAGCCAAACCTTTGTTCAAGG

[0133] CCGCCACAGATGAGCAATGGCAGAGCATACTTGAACGCTTCACTGTAGAACTCGTT

[0134] TCCGGACGTGCTCGTACTGCAAGGCAAACACTCGCGCAAAGGCTAACGTATCTAGT

[0135] CAAAGAAAACCCCGGTTGGGCGAAGGAGAAAATCGCAAAGGAAGCACTCGATGG

[0136] AGTGATGCGAGGTGCAAAGAAACTTGAAGAGCATCTCAAAAATAAAGGGCTTACT

[0137] GACAGTCGCGCAGTAATCAAGTTGGCTAACCTGTACAACGTTGCAAGCGTGTTTGC

[0138] CATTGAGGCAATTCGCGCTTCAGGAGATTACGTTAGTTATTACGAAACGGACACTCC

[0139] TCGCAGGGTCGCTTTCGGTGAATTGCGCGGTGGTTTGCACCAAGCCAGTGATGACA

[0140] CCGCTGCTATTCAAATCTCCGGCTTCTCACTTTCTGGGGATAACCCGCAATATGGCG

[0141] GTTTACTTACTTATGATCCTGATGAAATTGGCAAAGAGAAATGGAGCTTACTTTATAC

[0142] ATTGGATAATCAGGCAATGAAACTGGTGCAACCCAGCGAGAAGGCCAAAGGCAGG

[0143] GGCTTCCTCAAAAGTGACTTACAGGGCCTAGCCAAAACAGGCAGAGGCGACGAAG

[0144] CCCAACTACTGAAGGGAGAAGTTTGGTTGCCCAGCGACAAAGACAAGCATCCTCTT

[0145] TCACTGCCGCTGCAGATGGGGACGAGGCAAGGGCGAGAATATTTTTGGAACTTTGA

[0146] TCGAGGATTAAAAAACTCAGATGCTTGGGTACTGAACAACGGCCGACTGCTCCGAG

[0147] TAATGCCAGCAGGTCGCCCTGATCTCGCAAAGTTCTATGTGACCATCACACTTGGTC

[0148] GACAAGCTCCACCAATAGGGAACATCAAGCCAAAGGCATTCATTGGCATTGATCGA

[0149] GGTGAGGCGGTGCCGGCCGCATACGCAGTCATTGACACCAAAGGCAGGCTGCTAA

[0150] AATCAGGCTTGGTTCAAGATGAATATCGAGAGCAACAACGCAGATTCAACGACAAG

[0151] AAACGCGAGCTACAAAGGCAAAGCGGCGGCTACACCAAGTGGCTTCGCAGCAAAG

[0152] AGAGAAACCGAGCAAAGGCACTAGGGGGAGACGTTTCCCGCGAGCTACTCGACTT

[0153] AGCAGCAAAGCATGAAGCCCCGCTTGTTATGGAGTACCTTTCTAGTGGATTGGTAAC

[0154] TCGCGGAGGCAAGAACACAATGATGAGTTCAATGCAGTATGAGCGAGTGCTCGGCA

[0155] CGCTAGAGCAACGCTTGGCCGAGGTCGGCCTTTATGAACTTGTTTCTGATCCCAAGT

[0156] TTCGCAAGGGCGACAACGGCTTTATCAAGCTGGTTGGCCCGGCGTACACAAGTTCC

[0157] ACCTGTTCAGAGTGTGGACAGGTTTTTTCCACCGATTTTTATGAGGCCCTCACTGAC

[0158] ACTATCGTCAACAGCAAGGATGAAACTTGGGAGGTAACACTGCCAACAGGCAAAC

[0159] AACTTAATCTCCCCGAAGAGTACACCTACTGGGTTCGAGGTCAGGGCGAGAAAACC

[0160] CGAGAGACCCACGATCGGCTAACAGAGTTAATAAAGGGCAAGCCAATCGCAAAAAT

[0161] CAGCAAGACCAACCGCCGAAGTCTAACCGGCTTACTGCGCGGATGCCTCGTGCCTT

[0162] ACAGACCAAGACAGGCTGAGTTTCATTGCCCCTGCTGTGGCTATGAAGCAAACGCT

[0163] GACGAACAGGCCGCTTTGAACATTGCCCGTAAGCTTCTCTTCCGCGAGGAACTTGG

[0164] TGACAAGGTCAAAGAAGCCAGCGAAAGCGCCAGAAGAAACACCCAGAAACTCTG

[0165] GCAAGACTGGTACCAGAAAAAACTAGCCAAGGTTTGGCGCAAATGA

[0166] In other specific embodiments, the VbCasX protein comprises one or more of the following sequences:

[0167] 2) an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence set forth in SEQ ID NO: 1, and that retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1;

[0168] 3) an amino acid sequence that adds, substitutes, deletes, or inserts one or more amino acid residues in the amino acid sequence set forth in SEQ ID NO: 1, and that retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1; or,

[0169] 4) an amino acid sequence encoded by a nucleotide sequence that hybridizes to a polynucleotide sequence encoding the amino acid sequence set forth in SEQ ID NO: 1 under stringent conditions, and that retains the activity of binding a VbCasX guide RNA and / or nuclease activity of the amino acid sequence set forth in SEQ ID NO: 1, the stringent conditions being medium stringency conditions, medium-high stringency conditions, high stringency conditions, or very high stringency conditions.

[0170] It will be appreciated that the VbCasX proteins provided herein can comprise one or more additional features. For example, in some embodiments, the VbCasX proteins can comprise inhibitors, cytoplasmic localization sequences, export sequences, such as nuclear export sequences or other localization sequences, as well as tags that can be used for solubilization, purification, or detection of the fusion. Suitable tags provided herein include, but are not limited to, a biotin carboxylase carrier protein (BCCP) tag, a myc tag, a calmodulin tag, a FLAG tag, a hemagglutinin (HA) tag, a polyhistidine tag, also known as a histidine tag or His-tag, a maltose binding protein (MBP)-tag, a nus-tag, a glutathione-S-transferase (GST)-tag, a green fluorescent protein (GFP)-tag, a thioredoxin-tag, an S-tag, Softags (e.g., Softag 1, Softag 3), a Strep-tag, a biotin ligase tag, a Flash tag, a V5 tag, an SBP tag, a SUMO tag. Other suitable sequences will be apparent to those skilled in the art.

[0171] In other embodiments, the VbCasX proteins can also be fused (conjugated) to a heterologous polypeptide having an activity of interest to form a fusion protein (chimeric VbCasX protein) to provide it with additional functionality, such as modulating transcription of a target DNA, enzymatic activity to modify a target nucleic acid, enzymatic activity to modify a polypeptide associated with a target nucleic acid, and the like.

[0172] In comparison to previously identified CRISPR-Cas endonucleases, the VbCasX proteins provided herein are shorter, and thus provide a relatively short advantage when using the protein as an alternative to provide a nucleotide sequence encoding the protein. For example, in cases where a nucleic acid encoding a VbCasX protein is needed, such as in cases where a viral vector (e.g., an AAV vector) is used, this can be used for delivery to a cell, such as a eukaryotic cell (e.g., a mammalian cell, a human cell, a mouse cell, in vitro, ex vivo, in vivo) for research and / or clinical applications. And the strain from which the VbCasX protein exists is derived from a non-pathogenic bacteria, which has a lower risk of causing an immune response in humans in later applications compared to CRISPR-Cas systems derived from Staphylococcus aureus (SaCas9) or Streptococcus pyogenes (SpCas9).

[0173] (protospacer adjacent motif (PAM))

[0174] A VbCasX protein binds to a target DNA at a target sequence defined by a region of complementarity between the target DNA and an RNA targeting the DNA. As is the case with many Cas proteins, site-specific binding (and / or cleavage) of a double-stranded target DNA occurs at a position determined by both: (i) base-pairing complementarity between the guide RNA and the target DNA; and (ii) a short motif in the target DNA (i.e., a protospacer adjacent motif (PAM)).

[0175] In some embodiments, the PAM for a VbCasX protein is located directly 5’ of the target sequence on the non-complementary strand of the target DNA (the complementary strand hybridizes to the guide sequence of the guide RNA, while the non-complementary strand does not directly hybridize to the guide RNA and is the reverse complement of the non-complementary strand).

[0176] The PAM sequence recognized by the VbCasX protein-based system provided herein comprises 5’-TTR-3’ (where R is G or A), which distinguishes from the PAM sequence of known CRISPR-Cas systems, effectively expanding the editing activity window of CRISPR-Cas systems for use in the genome. In some specific embodiments, the PAM sequence recognized by the VbCasX protein in the system is selected from the group consisting of: TTG or TTA.

[0177] (VbCasX guide RNA)

[0178] A nucleic acid molecule that binds to a VbCasX protein to form a ribonucleoprotein complex (RNP) and targets the complex to a specific location within a target nucleic acid (e.g., a target DNA) is referred to herein as a “VbCasX guide RNA” or simply “guide RNA.” It should be understood that in some embodiments, a hybrid DNA / RNA can be made such that the VbCasX guide RNA comprises DNA bases in addition to RNA bases, but the term “VbCasX guide RNA” is still used to encompass such molecules herein.

[0179] In some embodiments, a VbCasX guide RNA comprises two segments, a targeting segment and a protein-binding segment. The targeting segment of a VbCasX guide RNA comprises a nucleotide sequence (a guide sequence) that is complementary to (and thus hybridizes with) a particular sequence (a target site) within a target nucleic acid (e.g., a target ssRNA, a target ssDNA, a complementary strand of a double-stranded target DNA, etc.). The protein-binding segment (or “protein-binding sequence”) interacts with (binds to) a VbCasX polypeptide. The protein-binding segment of a VbCasX guide RNA comprises two stretches of complementary nucleotides that hybridize to each other to form a double-stranded RNA duplex (a dsRNA duplex). Site-specific binding and / or cleavage of a target nucleic acid (e.g., genomic DNA) can occur at a location determined by base-pairing complementarity between the VbCasX guide RNA (the guide sequence of the VbCasX guide RNA) and the target nucleic acid (e.g., a target sequence of a target locus).

[0180] A VbCasX guide RNA and a VbCasX protein form a complex (e.g., bind through non-covalent interactions). The VbCasX guide RNA provides target specificity to the complex through the targeting segment, which comprises a guide sequence (a nucleotide sequence that is complementary to a target nucleic acid sequence). The VbCasX protein of the complex provides site-specific activity (e.g., cleavage activity provided by the VbCasX protein). In other words, the VbCasX protein is guided to a target nucleic acid sequence (e.g., a target sequence) as a result of its binding to the VbCasX guide RNA.

[0181] A “guide sequence” (also referred to as a “targeting sequence” of a VbCasX guide RNA) can be modified such that a VbCasX guide RNA can target a VbCasX protein (e.g., a naturally-occurring VbCasX protein, etc.) to any desired sequence of any desired target nucleic acid (in addition to the PAM sequence, which can be considered). Thus, for example, a VbCasX guide RNA can have a guide sequence that is complementary to (e.g., can hybridize with) a sequence in a nucleic acid in a eukaryotic cell, e.g., a viral nucleic acid, a eukaryotic nucleic acid (e.g., a eukaryotic chromosome, a chromosomal sequence, a eukaryotic RNA, etc.), etc.

[0182] In some embodiments, it can also be said that a VbCasX guide RNA comprises an “activator” and a “targeter” (e.g., an activator RNA (e.g., a tracrRNA) and a targeter RNA (e.g., a crRNA), respectively). When the “activator” and the “targeter” are two separate molecules, the guide RNA is referred to herein as a “dual guide RNA,” “Dual guide RNA (dgRNA),” “dual-molecule guide RNA,” or “two-molecule guide RNA” (e.g., a “VbCasX dual guide RNA”).

[0183] In the present disclosure, the term "activator" or "activator RNA" is used herein to mean a tracrRNA-like molecule of a VbCasX dual guide RNA (and thus, when the "activator" and "targeting factor" are linked together by, for example, insertion of nucleotides, a VbCasX single guide RNA). Thus, for example, a VbCasX guide RNA (dgRNA or sgRNA) comprises an activator sequence (e.g., a tracrRNA sequence). The tracr molecule (tracrRNA) is a naturally occurring molecule that hybridizes with a CRISPR RNA molecule (crRNA) to form a VbCasX dual guide RNA. The term "activator" is used herein to encompass not only naturally occurring tracrRNAs, but also tracrRNAs with modifications (e.g., truncations, elongations, sequence variations, base modifications, backbone modifications, bond modifications, etc.) in which the activator retains at least one function of a tracrRNA (e.g., contributes to a dsRNA duplex to which a VbCasX protein binds). In some embodiments, the activator provides one or more stem loops that can interact with a VbCasX protein. The activator can be referred to as having a tracr sequence (tracrRNA sequence), and in some embodiments is a tracrRNA, but the term "activator" is not limited to naturally occurring tracrRNAs.

[0184] In the present disclosure, the term “targeting factor” or “targeting factor RNA” is used herein to refer to the crRNA-like molecule of a VbCasX dual guide RNA (and thus, when the “activator” and “targeting factor” are linked together, e.g., by intervening nucleotides, a VbCasX single guide RNA). Thus, for example, a VbCasX guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment (e.g., a duplex-forming segment of a crRNA, which can also be referred to as a crRNA repeat sequence). Because the sequence of the targeting segment of the targeting factor (the segment that hybridizes to the target sequence of the target nucleic acid) is modified by the user to hybridize to the desired target nucleic acid, the sequence of the targeting factor will generally be a non-naturally occurring sequence. However, the duplex-forming segment of the targeting factor that hybridizes to the duplex-forming segment of the activator (described in greater detail herein) can comprise a naturally occurring sequence (e.g., can comprise the sequence of a naturally occurring duplex-forming segment of a crRNA, which can also be referred to as a crRNA repeat sequence). Thus, the term targeting factor is used herein to distinguish from naturally occurring crRNAs, despite the fact that a portion of the targeting factor (e.g., the duplex-forming segment) generally comprises a naturally occurring sequence from a crRNA. However, the term “targeting factor” encompasses naturally occurring crRNAs.

[0185] In some embodiments, the activator and the targeting factor are covalently linked to one another (e.g., by intervening nucleotides), and the guide RNA is referred to herein as a “single guide RNA” (sgRNA), “single molecule guide RNA,” or “one molecule guide RNA” (e.g., a “VbCasX single guide RNA”). Thus, a VbCasX single guide RNA comprises a targeting factor (e.g., a targeting factor RNA) and an activator (e.g., an activator RNA) linked to one another (e.g., by intervening nucleotides), and a duplex-forming segment of the targeting factor and a duplex-forming segment of the activator hybridize to one another to form a double-stranded RNA duplex (dsRNA duplex) of the protein-binding segment of the guide RNA, resulting in a stem-loop structure. Thus, the targeting factor and the activator each have a duplex-forming segment, wherein the duplex-forming segment of the targeting factor and the duplex-forming segment of the activator are complementary to one another and hybridize to one another.

[0186] In some alternative embodiments, the linker of the VbCasX single guide RNA is a stretch of nucleotides. In some embodiments, the targeting factor and the activator of the VbCasX single guide RNA are connected to each other by an intervening stretch of nucleotides, and the linker can have a length of 3 to 20 nucleotides (nt). In some embodiments, the linker of the VbCasX single guide RNA can have a length of 3 to 100 nucleotides (nt). In some embodiments, the linker of the VbCasX single guide RNA can have a length of 3 to 10 nucleotides (nt).

[0187] In some specific embodiments, for the VbCasX protein, the VbCasX guide RNA in the form of a dual guide RNA comprises an activator RNA (e.g., tracrRNA) and a targeting factor RNA (e.g., crRNA), wherein the activator RNA comprises a nucleotide sequence as set forth in SEQ ID NO: 3 or a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the sequence set forth in SEQ ID NO: 3.

[0188] VbCasX activator RNA (e.g., tracrRNA) (SEQ ID NO: 3):

[0189] GUUAUCAAUUUCUUUGAUACUGUGGUUCUUUCAGCCCCACUGAAUUCGAGCCACUCAGCGAACUUCAUUGAUU

[0190] The repeat sequence of the crRNA comprises a nucleotide sequence as set forth in SEQ ID NO: 4 or a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the sequence set forth in SEQ ID NO: 4.

[0191] VbCasX crRNA repeat sequence (SEQ ID NO: 4):

[0192] AUCGAUGAAGUUCGGUGAGAUGGUUUCAAAG

[0193] In some specific embodiments, for a VbCasX protein, the VbCasX guide RNA in the form of a single guide RNA comprises a nucleotide sequence as set forth in SEQ ID NO: 5 or a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the sequence set forth in SEQ ID NO: 5.

[0194] VbCasX sgRNA (SEQ ID NO: 5):

[0195] GUUAUCAAUUUCUUUGAUACUGUGGUUCUUUCAGCCCCACUGAAUUCGAGCCACUCAGCGAACUUCAUUGAUUUCGAUCGAUGAAGUUCGGUGAGAUGGUUUCAAAG

[0196] (guide sequence of VbCasX guide RNA)

[0197] The targeting segment of a VbCasX guide RNA comprises a guide sequence (i.e., a targeting sequence) that is a nucleotide sequence that is complementary to a sequence (target site / target sequence) in a target nucleic acid. In other words, the targeting segment of a VbCasX guide RNA can interact with a target nucleic acid (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), single-stranded RNA (ssRNA), or double-stranded RNA (dsRNA)) in a sequence-specific manner through hybridization (i.e., base pairing). The guide sequence of a VbCasX guide RNA can be modified (e.g., by genetic engineering) / designed to hybridize to any desired target sequence within a target nucleic acid (e.g., a eukaryotic target nucleic acid, such as genomic DNA) (e.g., when considering a PAM, e.g., when targeting a dsDNA target).

[0198] In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100%.

[0199] In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19-25 contiguous nucleotides.

[0200] In some embodiments, the guide sequence has a length in the range of 19-30 nucleotides (nt) (e.g., 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some embodiments, the guide sequence has a length in the range of 19-25 nucleotides (nt) (e.g., 19-22, 19-20, 20-25, 20-25, or 20-22 nt). In some embodiments, the guide sequence has a length of 19 or more nt (e.g., 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some embodiments, the guide sequence has a length of 19 nt, a length of 20 nt, a length of 21 nt, a length of 22 nt, or a length of 23 nt.

[0201] For example, the guide sequence can be AGGGCGACACCCUGGUGAAC (SEQ ID NO: 6), which is a stretch of protein-coding sequence on the green fluorescent reporter gene EGFP. In some specific embodiments, the sequence of an exemplary VbCasX sgRNA comprising this guide sequence is (SEQ ID NO: 7):

[0202] GUUAUCAAUUUCUUUGAUACUGUGGUUCUUUCAGCCCCACUGAAUUCGAGCCAC

[0203] UCAGCGAACUUCAUUGAUUUCGAUCGAUGAAGUUCGGUGAGAUGGUUUCAAAG

[0204] AGGGCGACACCCUGGUGAAC

[0205] where the underlined portion is the guide sequence.

[0206] (donor polynucleotide / donor template)

[0207] Under the guidance of a VbCasX dual guide or single guide RNA, the VbCasX protein generates a site-specific double-strand break (DSB) or single-strand break (SSB) (e.g., when the VbCasX protein is a nickase variant) within a double-stranded DNA (dsDNA) target nucleic acid in some cases, which is repaired by non-homologous end joining (NHEJ) or homology-directed repair (HDR).

[0208] In some embodiments, contacting the target DNA (in contact with the VbCasX protein and the VbCasX guide RNA) occurs under conditions that allow for non-homologous end joining or homology directed repair. Thus, in some embodiments, the target DNA is contacted with a donor polynucleotide (e.g., by introducing the donor polynucleotide into the cell), wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA.

[0209] <Polynucleotide>

[0210] The present invention provides a polynucleotide comprising one or more of:

[0211] i) a nucleotide sequence encoding a VbCasX protein in the above-described CRISPR-Cas system;

[0212] ii) a nucleotide sequence encoding a VbCasX guide RNA in the above-described CRISPR-Cas system; and,

[0213] iii) a donor polynucleotide sequence.

[0214] In some embodiments, when the VbCasX guide RNA is a single guide RNA, the polynucleotide sequence encoding the guide RNA can comprise a single nucleotide sequence. In other embodiments, when the VbCasX guide RNA is a dual guide RNA, the polynucleotide sequence encoding the guide RNA can comprise two separate nucleotide sequences.

[0215] In some embodiments, the nucleotide sequence encoding the VbCasX protein is codon-optimized. This type of optimization can entail mutation of the nucleotide sequence encoding the VbCasX to mimic the codon bias of the intended host organism or cell while encoding the same protein.

[0216] <Vector>

[0217] The present invention provides a vector comprising the above-described polynucleotide, i.e., comprising one or more of:

[0218] (i) a nucleotide sequence encoding a VbCasX protein in the above-described CRISPR-Cas system;

[0219] (ii) a nucleotide sequence encoding a VbCasX guide RNA in the above-described CRISPR-Cas system; and,

[0220] (iii) a donor polynucleotide sequence.

[0221] In some embodiments, (i)-(iii) above can be in the same vector. In other embodiments, (i)-(iii) above can be in different vectors.

[0222] In some embodiments, the vector is an expression vector, more specifically a recombinant expression vector. Suitable expression vectors include viral expression vectors (e.g., viral vectors based on viruses such as vaccinia virus, polio virus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloid leukemia virus, and mammary tumor virus), etc.

[0223] Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements can be used in the expression vector, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc.

[0224] Methods of introducing nucleic acids into host cells are known in the art, and any convenient method can be used to introduce nucleic acids (e.g., expression constructs) into cells. Suitable methods include, for example, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethylenimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.

[0225] <Cells>

[0226] The present disclosure provides a cell comprising one or more of:

[0227] (a) the CRISPR-Cas system described above;

[0228] (b) the polynucleotide described above; and,

[0229] (c) the vector described above.

[0230] The cell can be any of a variety of cells, including, for example, in vitro cells, in vivo cells, ex vivo cells, primary cells, cancer cells, animal cells, plant cells, algal cells, fungal cells, etc.

[0231] In some embodiments, the cell is a recipient of the CRISPR-Cas system provided herein, which can also be referred to as a "host cell" or a "target cell". The host cell or target cell can be a recipient of the CRISPR-Cas system provided herein. The host cell or target cell can be a recipient of the RNP complex provided herein. The host cell or target cell can be a recipient of a single component of the CRISPR-Cas system provided herein.

[0232] In some specific embodiments, non-limiting examples of the cell include: a prokaryotic cell, a eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a unicellular eukaryotic organism, a protozoan cell, a cell from a plant, an algal cell, a fungal cell, an animal cell, a cell from an invertebrate animal, a cell from a vertebrate animal, a cell from a mammal (e.g., an ungulate; a rodent; a non-human primate; a human; a feline; a canine; etc.), and the like. In some cases, the cell is a cell that is not derived from a natural organism (e.g., the cell can be a synthetic cell; also referred to as an artificial cell).

[0233] <Kit>

[0234] The present disclosure provides a cell comprising one or more of:

[0235] (A) the CRISPR-Cas system described above;

[0236] (B) the polynucleotide described above;

[0237] (C) the vector described above; and,

[0238] (D) the cell described above.

[0239] <Methods of modifying a target nucleic acid and uses>

[0240] The present disclosure provides a method of modifying a target nucleic acid, the method comprising the step of contacting the target nucleic acid with the CRISPR-Cas system provided herein, the polynucleotide provided herein, and / or the vector provided herein. In some embodiments, the contacting results in modification of the target nucleic acid by the VbCasX polypeptide. The present disclosure also provides uses of the CRISPR-Cas system provided herein, the polynucleotide provided herein, the vector provided herein, the cell provided herein, and / or the kit provided herein in modifying a target nucleic acid.

[0241] In some embodiments, the CRISPR-Cas system comprises: a VbCasX polypeptide and a VbCasX guide RNA, wherein the VbCasX guide RNA comprises a guide sequence that hybridizes to a target sequence of the target nucleic acid.

[0242] In some specific embodiments, the modification is cleavage of the target nucleic acid. In some specific embodiments, the target nucleic acid is selected from the group consisting of: double stranded DNA, single stranded DNA, RNA, genomic DNA, and extrachromosomal DNA.

[0243] In some specific embodiments, the contacting occurs in vitro or in vivo. In some specific embodiments, the contacting occurs inside a cell or outside a cell.

[0244] In some specific embodiments, the cell is a eukaryotic cell or a prokaryotic cell.

[0245] In some more specific embodiments, the cell is selected from the group consisting of: a plant cell, a fungal cell, a mammalian cell, a reptilian cell, an insect cell, an avian cell, a fish cell, a parasitic cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, and a human cell.

[0246] In some more specific embodiments, the contacting results in genome editing.

[0247] In some embodiments, the contacting comprises introducing the CRISPR-Cas system into a cell. In some embodiments, the contacting further comprises: introducing a DNA donor template into the cell.

[0248] Examples

[0249] The present application is further described in further detail below in connection with specific embodiments, which are given by way of illustration only and are not intended to be limiting of the present application. The following examples are provided as guidance to one of ordinary skill in the art and are not intended to be limiting of the present application in any way.

[0250] The experimental methods in the following examples are routine methods, which are carried out according to the techniques or conditions described in the literature in the art or according to the instructions of the products, unless otherwise specified. The materials, reagents, etc. used in the following examples are commercially available, unless otherwise specified.

[0251] Example 1: Identification of novel CRISPR-VbCasX system using bioinformatics methods

[0252] Using the designed workflow of mining new CRISPR-Cas systems, a small CRISPR-VbCasX system from non-pathogenic bacteria was retrieved from the self-built Bohai macrogenome database through the prediction of the workflow and the layer-by-layer screening of the scoring function. The Cas protein of the system is composed of only 872 amino acids. Through protein proximity sequence analysis of the host genome, it was found that the system has a complete CRISPR array, and the crRNA and tracrRNA sequences were further identified (such as Figure 1 ). Through protein domain prediction, remote homologous protein retrieval and protein tertiary structure prediction, it was found that the VbCasX protein has a typical RuvC nuclease functional domain (such as Figure 1 ), and it is preliminarily believed that the system has nucleic acid cleavage function.

[0253] Example 2: In vitro purification of VbCasX protein and used for PAM identification

[0254] I. Expression and purification of VbCasX protein

[0255] 1. Replace the small DNA fragment between the recognition sites of restriction enzymes BamHI and KpnI of pSUMOH10 vector (pSUMOH10 from the article https: / / doi.org / 10.1093 / nar / gkaa1130) with the human codon-optimized VbCasX gene (as shown in SEQ ID NO: 25), and the other sequences remain unchanged, to obtain the recombinant plasmid pET28b-10XHis-SUMO-VbCasX.

[0256] SEQ ID NO: 25:

[0257] ATGAATAGCGTGAGAAAGTCCCTGAGAGAGCTGCTGACCCTGGGCCTGAAGGCCAGCTCCAGC

[0258] GCCAGCCCACAGACAATCACCAGAACCGTGAAGCTGGGCGTGGAGGCTAAGTATCGGTCCGATC

[0259] TGGCCACAGCCCTGCATAAGCACTTTGACGCCTATGAGGAGTTCAGGAGAAGCGTGCTGGATGA

[0260] GCTGGAGCAGTGGTGGAACGATGATCCCGACTCCTTTATTCAGATGGTGAAGTGTAAGAAAGCC

[0261] GAGCCCTACGAGGATAAGTCCAGCTGCGGCGCCTGGCTGTTCTCTAAGTTTCTGACCGGCAAAA

[0262] AGCTGCCCGAGGGCCTGACAAACAAGGCCGGCTTCGCCCTGCTGGACTCCCTGGCCGGCGGCC

[0263] TGAAGTCCTTTATCACAAGGCGGGCCAACGTGGTGAAGGACATCAAACAGAGAGAAAAGACAA

[0264] ACACCTCCGAGTGGAAGAAGGAGCTGACCGCCCTGGCCAAGGAGCTGAGCGAGGAGATGCCA

[0265] GAGGAGGCCCCCGAGCTGGACTTCGACAATGTGGATGCCGCCGCCATCGAGGGATACAACGAGT

[0266] GGGTGGCCCTGGTGAGGTCCTGGTGTAACCTGATTCTGGTGCAGAGGCACCAGCTGAATAGACG

[0267] GGACGTGCCCGCCCCCAGGTACCTGAAGGGCTACCCCGGCTTCCCCGGCTCTCAGCGGTACGCT

[0268] GAGAATCTCCCCCTGAAGGAAAGCCTCCCACTGCTGCGGGAGGTGACCCTGGAGAACATCAAG

[0269] GCCGCTAAGCCCCTGTTCAAGGCCGCCACCGACGAGCAGTGGCAGAGCATCCTGGAGCGGTTC

[0270] ACCGTGGAGCTCGTGAGCGGGCGCGCCCGGACCGCCCGGCAGACCCTCGCCCAGCGGCTGACC

[0271] TACCTCGTGAAGGAGAATCCTGGCTGGGCCAAGGAGAAAATCGCCAAGGAAGCCCTGGACGGC

[0272] GTGATGAGGGGCGCCAAGAAGCTGGAGGAGCACCTGAAGAATAAGGGCCTGACCGATAGCCGC

[0273] GCCGTGATTAAGCTGGCCAACCTGTATAATGTGGCCTCCGTGTTCGCCATTGAGGCTATCCGGGC

[0274] CAGCGGCGACTACGTGAGCTACTATGAGACCGACACCCCTCGGAGGGTGGCCTTTGGCGAACTG

[0275] CGGGGCGGCCTGCACCAGGCCAGCGACGACACCGCCGCCATCCAGATCAGCGGCTTCAGCCTG

[0276] TCTGGCGATAACCCCCAGTATGGGGGCCTGCTGACTTATGACCCTGACGAGATCGGGAAAGAGA

[0277] AGTGGTCCCTGCTGTACACCCTCGACAACCAGGCCATGAAGCTGGTGCAGCCTAGCGAGAAGGC

[0278] CAAGGGGCGGGGGTTTCTGAAGTCCGACCTGCAGGGCCTGGCCAAAACCGGCAGAGGCGACG

[0279] AGGCCCAGCTGCTGAAGGGCGAGGTGTGGCTCCCTTCCGATAAGGACAAGCACCCCCTGTCACT

[0280] GCCACTGCAGATGGGCACCAGACAGGGAAGAGAGTACTTCTGGAACTTCGATCGGGGCCTGAA

[0281] GAACTCTGACGCCTGGGTTCTGAATAATGGGCGCCTGCTGCGCGTGATGCCCGCCGGCAGACCA

[0282] GACCTGGCCAAATTCTATGTCACAATTACACTGGGCCGCCAGGCCCCCCCCATCGGCAACATCAA

[0283] GCCCAAGGCCTTCATCGGAATCGATAGGGCGAGGCCGTCCCCGCCGCCTATGCCGTGATCGAC

[0284] ACCAAGGGCCGCTGCTGAAAGCGGCCTGGTGCAGGATGAGTACCGCGAGCAGCAGCGGCGG

[0285] TTTAACGACAAGAAGCGCGAGCTGCAGAGGCAGTCAGGCGGCTACACCAAGTGGCTCCGGTCC

[0286] AAGGAGAGAAATAGACCCAAGGCCCTGGGCGGAGATGTGAGCAGGGAACTGCTGGACCTGGCC

[0287] GCTAAGCATGAGGCTCCCCTCGTGATGGAGTACCTGAGCTCTGGACTGGTGACCAGGGCGGCA

[0288] AGAACACCATGATGAGCAGCATGCAGTATGAGAGAGTGCTGGGGCACCCTGGAGCAGAGGCTGG

[0289] CCGAGGTGGGCCTGTACGAGCTGGTGTCTGATCCTAAGTTTAGAAAGGGAGCAATGGCTTTATT

[0290] AAGCTGGTGGGGCCAGCCTACACCAGCTCCACCTGCTCTGAGTGCGGACAGGTGTTCTCCACCG

[0291] ACTTCTATGAGGCCCTGACTGACACCATCGTGAACAGCAAGGATGAAACCTGGGAGGTGACTCT

[0292] GCCTACCGGCAAGCAGCTGAACCTGCCTGAGGAGTACACCCTACTGGGTGAGAGGGCAGGGCGA

[0293] GAAAACACGGGAGACACACGACCGGCTGACCGAGCTGATCAAGGGCAAGCCCATCGCCAAGAT

[0294] CTCCAAGACAAATAGACGCAGCCTGACAGGGCTGCTGAGGGGGTGTCTGGTGCCCTATAGGCCT

[0295] AGACAGGCCGAGTTTCACTGTCCCTGCTGCGGCTATGAGGCCAATGCCGACGAGCAGGCCGCCC

[0296] TGAATATTGCCAGGAAGCTGCTGTTCCGGGAGGAACTGGGGGATAAAGTGAAGGAGGCCTCCG

[0297] AGTCCGCCAGAAGGAATACCCAGAAGCTGTGGCAGGATTGGTACCAGAAGAAGCTGGCCAAGG

[0298] TGTGGAGAAAGTGA

[0299] 2. The recombinant plasmid pSUMOH10-VbCasX constructed in step 1 was introduced into Escherichia coli Rosetta to obtain recombinant Escherichia coli, named Rosetta / VbCasX.

[0300] 3. After completing step 2, inoculate a single Rosetta / VbCasX colony into 100 mL of TB medium and incubate at 37°C with shaking at 220 rpm for 6-8 hours to obtain culture solution 1. Inoculate culture solution 1 into 2 L of TB medium and incubate at 37°C with shaking at 220 rpm until OD (dose retardation) is reached. 600nm The value was set to 1.0-1.2, resulting in culture solution 2. Culture solution 2 was then cooled to 16℃, and IPTG was added to a concentration of 0.4 mM. The mixture was incubated overnight for 20 h to obtain culture solution 3.

[0301] 4. After completing step 3, take 3 of the culture solution, centrifuge at 4000 rpm for 15 min, and collect the precipitate.

[0302] 5. Take the precipitate collected in step 4, add 70 mL of lysis buffer to resuspend it, and then sonicate for 35 min (220 W, 3 s operation, 7 s interval) to obtain cell lysis buffer. Take the cell lysis buffer, centrifuge at 15000 rpm for 60 min, and collect the supernatant.

[0303] 6. Load the supernatant collected in step 5 onto a gravity column, wash the column with 10 ml of lysis buffer, and finally add 10 column volumes of elution buffer containing 300 mM imidazole to obtain the eluted product of the protein with the His-SUMO tag.

[0304] 7. The elution product of the protein with His-SUMO tag obtained in step 6 was mixed with 200 μL Ulp1 protease, and reacted on ice for 30 min to obtain a sample.

[0305] 8. After completion of step 7, the sample was subjected to SDS-PAGE to identify the efficiency of removal of the His-SUMO tag, and then an equal volume of dilution buffer was added, and the sample was loaded into a 5 ml heparin pre-packed column using a peristaltic pump. Subsequently, a salt concentration gradient from 200 mM to 1 M sodium chloride was used to competitively elute the target protein using an AKTA instrument. The eluted components were identified by SDS-PAGE, and the sample with the correct size of the protein band was loaded into a 30 kD concentration tube, and centrifuged at 3800 rpm for a short time to concentrate the sample to 1 mL.

[0306] 9. The concentrated sample obtained in step 8 was loaded onto a Superdex 200 10 / 300 column through a 1 ml loading loop, and eluted using SEC buffer. The column separates and purifies the protein by the size of the particle diameter. The target protein starts to elute at an elution volume of about 11.5 ml, and the ratio of A280 to A260 is close to 2. After the eluted sample was identified by SDS-PAGE, it was concentrated using a 30 kD concentration tube, and finally 10 μL per tube was stored in liquid nitrogen.

[0307] The experimental results are shown in Figure 2 (left side: elution position, right side: SDS-PAGE result). The results show that the VbCasX protein has high expression yield, high purity, and stable quality in E. coli Rosetta.

[0308] Lysis buffer: solutes and their concentrations are 20 mM HEPES-Na pH = 7.5, 600 mM NaCl, 30 mM imidazole, 10% glycerol, and 1 mM TECP.

[0309] Elution buffer: solutes and their concentrations are 20 mM HEPES-Na pH = 7.5, 600 mM NaCl, 300 mM imidazole, 1 mM TECP, and 10% glycerol.

[0310] Dilution buffer: solutes and their concentrations are 20 mM HEPES-Na pH = 7.5, 200 mM NaCl, and 10% glycerol.

[0311] SEC buffer: solutes and their concentrations are 20 mM HEPES-Na pH = 7.5, 150 mM NaCl, 10% glycerol, and 1 mM TECP.

[0312] II. Identification of PAM recognized by VbCasX protein

[0313] 1. Synthetically synthesize two single-stranded DNA molecules with complementary sequences, 5'-gcctgcaggtcgactctagaggatcNNNNNAGGGCGACACCCTGGTGAACg-3' (SEQ ID NO: 8) and 5'-ggccagtgaattcgagctcggtacGTTCACCAGGGTGTCGCC-3' (SEQ ID NO: 9), wherein N is A, T, G or C; then anneal the two single-stranded DNA molecules in 1x Annealing buffer (pH 8.0, 10 mM Tris-HCl buffer containing 25 mM KCl) (annealing program: 95°C for 5 min), slowly cool to room temperature to obtain the annealing product, wherein N is A, T, G or C.

[0314] 2. Take pUC19 vector (from TIANGEN, Addgene number: #50005), and cut it with restriction endonucleases BamHI and Kpnl, and recover the vector skeleton.

[0315] 3. Homologously recombine the annealing product obtained in step 1 and the vector skeleton recovered in step 2, then transform E. coli DH5a to obtain a recombinant bacterium. Extract the plasmid from the recombinant bacterium using an endotoxin-free plasmid extraction kit (Tiangen, DP117) to obtain a plasmid library carrying random PAM.

[0316] III. Analysis of PAM sequence

[0317] 1. Mix the nCas protein obtained in step one, the VbCasX sgRNA (shown in SEQ ID No: 7) and the plasmid library carrying random PAM obtained in step two according to a molar ratio of 10:15:1, then react in cleavage buffer (solvent: pH 7.5, 20 mM Tris-HCl buffer; solute: 300 mM NaCl, 10 mM MgCl2 and 1 mM DTT) at 37°C for 60 min, and add EDTA to terminate the reaction. Use 1.2% agarose gel to recover the linearized plasmid.

[0318] 2. Assemble the end repair system, then react at 11°C for 20 min and at 75°C for 10 min.

[0319] End repair system was 40 μΐ, consisting of 20 μΐ, linearized plasmid, 8 μΐ, 5x T4 polymerase buffer, 1.6 μΐ, (to the final 0.1 mM each) 10 mM dNTP, 10 μΐ, nuclease free water and 0.4 μΐ, T4 DNA polymerase.

[0320] 3. To the system completed in step 2, add 1 μΐ, dATP, 1 μΐ, Dreamtaq polymerase (Thermo, EP0702), 72 °C for 30 min (for dA addition), recover the product.

[0321] 4. The product recovered in step 3 and adapter sequence (32 bp) are ligated with T4 DNA ligase (Biun, D7003) (room temperature reaction for 30 min), and then the product is recovered using beads (Novagen, N411-03).

[0322] Adapter sequence is: 5'-CGCATCGAGCTGAAGGGCATCGACTTCAAGG-3' SEQ ID NO: 10 and 5'-CCTTGAAGTCGATGCCCTTCAGCTCGATGCGT-3' SEQ ID NO: 11.

[0323] 5. The product recovered in step 4 is used as a template for PCR amplification using F: 5'-ATGTTGTGTGGAATTGTGAGCG-3' SEQ ID NO: 12 and R: 5'-CCTTGAAGTCGATGCCCTTCAG-3' SEQ ID NO: 13, and the PCR amplification product is recovered using beads.

[0324] 6. The product recovered in step 5 is used to construct a DNA library using the TIANSeq Rapid DNA Library Construction Kit (Catalog No: NG102). The PAM identification results are as shown in Figure 3 , and the VbCasX recognition sequence on the target upstream is 5'-TTR-3' (R is G or A).

[0325] Example 3, VbCasX protein and its gRNA can target and edit human cell genome

[0326] In this embodiment, the system based on VbCasX protein is used to edit mammalian cells. Briefly, a plasmid with VbCasX protein and gRNA expression system is constructed, and a gRNA plasmid with VbCasX protein and non-target spacer is used as a negative control (Non target, NC), and spCas9 and its corresponding gRNA are used as a positive control (Sun A, Li CP, Chen Z, Zhang S, Li DY, Yang Y, Li LQ, Zhao Y, Wang K, Li Z, Liu J, Liu S, Wang J, Liu JG. The compact Casπ(Cas12l)'bracelet'provides a unique structural platform for DNA manipulation. Cell Res. 2023 Mar;33(3):229-244.), and a HEK293A-EGFP out of frame cell line is transfected with a Celetrix LE+ electroporator. On the 6th day after transfection, the expression of EGFP in cells is observed under a fluorescence microscope (6 possible editing sites are selected for the first attempt). In addition, the genome of the edited cells is extracted, the DNA sequence of the editing site is obtained by PCR and reannealing, and then the edited sequence is cut by T7E1 nuclease, which can detect VbCasX in vivo editing. The principle is: before editing, there is a length of (3n+2) bp from the translation start site to the middle of the CDS sequence of EGFP, which makes EGFP unable to be normally translated and expressed. If the VbCasX protein and gRNA targeting the upstream sequence of EGFP can edit the genome, the upstream sequence will have base insertion or deletion due to NHEJ (non-homologous end joining), which will make EGFP have a certain probability of returning to the normal reading frame and expressing. Then, the genome of these cells is extracted, the sequence of the editing site is obtained by PCR, and the sequence is reannealed to obtain the mismatched DNA double-stranded sequence of the edited sequence and the unedited sequence. T7E1 nuclease can recognize the mismatched site and cut the double-stranded DNA at that position. Therefore, by cutting the annealed DNA fragments with T7E1, and separating the substrate and product by agarose gel, it can be determined whether VbCasX has editing activity in mammalian cells. On this basis, we use the extracted edited cell genome to obtain the editing site DNA library with Index, P5 adapter and P7 adapter sequences by two rounds of PCR, and then detect the editing efficiency of VbCasX in mammalian cells by second-generation sequencing.

[0327] The following will specifically illustrate this embodiment:

[0328] 1. Construction of plasmid for in cell editing

[0329] (1) VbCasX gene (SEQ ID NO: 25) was inserted behind the chicken beta-actin promoter of pBLO62.5 vector (Addgene plasmid #123124) by homologous recombination method, sgRNA and its targeting sequence (i.e. guide sequence, which recognizes the target sequence, and based on the nucleotide sequence of the target sequence, the specific nucleotide sequence of the guide sequence can be determined) were inserted behind the U6 promoter by homologous recombination method, and E. coli DH5a competent cells were transformed. Coated with resistant plates, picked bacteria for sequencing verification and preservation.

[0330] (2) Target sequence and its PAM are as follows:

[0331] Serial number Target sequence PAM T1 TTCATCTCCACCTGAGCAGA (SEQ ID NO: 14) TTG T2 TCGATCTGCTCCCCGAGCTC (SEQ ID NO: 15) TTG T3 GCGATGGCCTCTGCGTTACT (SEQ ID NO: 16) TTG

[0332] 2. In cell EGFP_light_on experiment

[0333] (1) Electroporation of plasmid

[0334] Cell recovery: HEK293A-Myh8(mouse)-out of frame-EGFP cell line (preserved in our laboratory) was used for gene editing. Heat the water bath to 37°C. Take the cryopreservation tube from liquid nitrogen, check if the cap is screwed tightly, quickly place it in the 37°C water bath and constantly stir. Melt the cryopreserved liquid within 1 minute. Immediately add the cryopreserved liquid to a 15 mL centrifuge tube containing more than 10 times the volume of DMEM complete medium (10% FBS, 1% PS), centrifuge at 800 rpm for 3 minutes. Take out the centrifuge tube, discard the supernatant, add 1-2 mL of complete medium, suspend the cells by blowing, and transfer the suspension to a new 10 cm cell culture dish, place it under an inverted microscope to observe the cell density in the culture dish, then place it in a 37°C, 5% CO2 cell incubator, and gently shake the cells to evenly distribute them. The next day, place the cultured cells under an inverted microscope to observe cell recovery and viability.

[0335] Cell passage: Observe the cells, under the microscope, when the cells cover the bottom of the culture dish, the fusion degree is about 80-90%, the cells can be passaged, the culture medium in the culture dish is discarded, 1 mL trypsin is added to the culture dish, the trypsin is slowly shaken to cover the cells, the culture dish is quickly transferred to a 37°C incubator for 1-2 min of digestion, and the supernatant of the culture medium is observed to have floating white cell clusters with the naked eye. Under the microscope, the cell edge gradually becomes clear, and the cell shape changes to round, which changes very quickly, representing the completion of digestion. 1 mL of complete medium is added to terminate the digestion. The cell suspension after termination of digestion is collected and centrifuged at 800 rpm for 3 minutes. The centrifuge tube is taken out, the supernatant is discarded, 10 ml of complete medium is added to resuspend the cells, and the suspension is transferred to a new 10 cm cell culture dish, which is placed on an inverted microscope to observe the cell density in the culture dish, and then placed in a 37°C CO2 cell incubator for continuous culture.

[0336] Cell electroporation: Observe the cell state on the same day, and the cells can be transfected when the cell fusion degree reaches 70-90%. Add 1 mL trypsin to the culture medium, slowly shake to cover the cells, quickly transfer the culture dish to a 37°C incubator for 1-2 min of digestion, and add 1 mL of complete medium to terminate the digestion after digestion is complete. The cell suspension after termination of digestion is collected and centrifuged at 800 rpm for 3 minutes. The centrifuge tube is taken out, the supernatant is discarded, 4 ml of PBS is added to resuspend the cells, and the cells are centrifuged at 800 rpm for 3 minutes. This step is repeated twice. After the cells are washed twice with PBS, the supernatant PBS is discarded, and 1-2 ml of Opti-MEM TM medium is added to resuspend the cells, counted using a hemocytometer, and 5 x 10 5 cells are taken, centrifuged at 800 rpm for 3 minutes. The centrifuge tube is taken out, the supernatant is discarded, 20 μl of Opti-MEM TM medium is added to resuspend the cells to obtain a cell suspension to be transfected, and the plasmids are electroporated according to Table 1:

[0337] Table 1 Electroporation plasmid system

[0338]

[0339] The transfected cells are gently added to the 24-well plate with preheated complete medium, shaken and mixed. The cells are cultured in a 37°C cell incubator. The day of transfection is recorded as day 0.

[0340] (2) Puromycin screening

[0341] Day 1: Add medicine. The puromycin dry powder is suspended in DMSO and mixed, 24 hours after cell transfection, the cells are passaged, and the culture medium is replaced with 1 ml of fresh complete medium containing a final concentration of 1.5 μg / mL puromycin when resuspended, and the same volume of water as the transfection is set as a blank control.

[0342] Day 3: Cell plate change Replace the culture medium in the well with fresh 1 ml complete medium containing a final concentration of 1.5 μg / mL puromycin.

[0343] Day 4: Observe the state of the cells. If the cells in the blank control group are dead, it indicates that the puromycin screening effect is good. At this time, the cells are passaged, and the medium is replaced with 1 ml of complete medium without puromycin when resuspended.

[0344] Day 6: Observe the cells under a microscope and take pictures. After taking pictures, discard the old culture medium of the cells, blow the cells in the well with 1 ml of PBS, collect the cell suspension, centrifuge at 1200g for 3 min, discard the PBS, and perform cell genomic extraction.

[0345] (3) Cell genomic extraction

[0346] The genomic DNA extraction kit (CW2298M) from Kangwei Century was used to extract the cell genome.

[0347] (4) T7E1 detection of mutation rate

[0348] PCR amplification: Using the genome as a template, a sequence of about 600 bp near the target cleavage site was selected for PCR amplification. The PCR product was identified by agarose gel electrophoresis, and the DNA Clean Beads (Vazyme) was used to recover the PCR product.

[0349] Primer sequence:

[0350] Forward MYH8-F (SEQ ID NO: 17):

[0351] ACGGTGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGAT

[0352] Reverse MYH8-R (SEQ ID NO: 18):

[0353] GTCGGGGTAGCGGCTGAAGCACTGCACGCCGTAGGTCAGGGTGGTCAC

[0354] T7E1 detection: Vazyme's T7 Endonuclease I was used to react with 100 ng of gel-recovered PCR product according to the instructions. 2% agarose gel electrophoresis containing Ultra GelRed (Vazyme) was used for analysis.

[0355] (5) Second-generation sequencing detection:

[0356] One round of PCR amplification:

[0357] PCR amplification: 217 bp of sequence near the target cleavage site was selected for PCR amplification with genomic template, and the amplification was set for 25 cycles. R1 and R2 sequences were introduced at the 5' and 3' of the first round PCR product by primers, respectively. The PCR product was identified by agarose gel electrophoresis, and the first round PCR product was recovered using DNA Clean Beads (Vazyme).

[0358] First round PCR primer sequence:

[0359] NGS-1F (SEQ ID NO: 19):

[0360] ACACTCTTTCCCTACACGACGCTCTTCCGATCTACCGTCAGATCGCCTGGAGACGCCA

[0361] NGS-1R (SEQ ID NO: 20):

[0362] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCTTCCGAAGAGCGGCTGCTGTGGC

[0363] NGS-2F (SEQ ID NO: 21):

[0364] ACACTCTTTCCCTACACGACGCTCTTCCGATCTAGCAGCCGCTCTTCGGAAGA

[0365] NGS-2R (SEQ ID NO: 22):

[0366] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGGCACCACCCCGGTGAAC

[0367] Wherein NGS-1F and NGS-1R amplify the DNA fragment near the editing site T1, and NGS-2F and NGS-2R amplify the DNA fragment near Cas9, T2, T3.

[0368] Second round PCR amplification:

[0369] PCR amplification: Using 50ng-100ng of the first-round PCR product as a template, a 217bp sequence near the target cleavage site was selected for PCR amplification. The amplification was set to 15 cycles. The PCR products were identified by different index sequences at 5' and 3' and P5 and P7 adapters in the second-round PCR products by primers. The PCR products were identified by agarose gel electrophoresis. The second-round PCR products were recovered using DNA Clean Beads (Vazyme). The obtained second-round PCR products were sent for next-generation sequencing.

[0370] Second-round PCR primer sequences:

[0371] 2 nd -NGS-1F(SEQ ID NO:23):

[0372] AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAGCTCTTCCGATCT

[0373] 2 nd -NGS-1R(SEQ ID NO:24):

[0374] CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC

[0375] Where NNNNNNNN represents different deoxynucleotide sequences as different indices, as shown in Table 2:

[0376] Table 2 index sequence

[0377]

[0378] like Figures 4-5 As shown, the VbCasX provided by this invention exhibits approximately 5% editing activity at editing site T1 and approximately 1% editing activity at editing sites T2 and T3, demonstrating that the system can achieve site-specific editing in mammalian cells. Based on this, its editing activity in mammalian cells can be improved through rational design and directed evolution optimization, and it can be developed into a new generation of small cell editing tools.

Claims

1. A mini-CRISPR-Cas system, comprising: (i) a VbCasX protein as set forth in SEQ ID NO: 1; and, (ii) a VbCasX guide RNA that forms a complex with the VbCasX protein, and the VbCasX guide RNA comprises a guide sequence that hybridizes to a target sequence in a target nucleic acid; wherein the protospacer adjacent motif recognized by the VbCasX protein comprises 5’-TTR-3’, wherein R is G or A.

2. The mini-CRISPR-Cas system according to claim 1, wherein, The VbCasX protein is derived from Verrucomicrobia bacterium .

3. The mini-CRISPR-Cas system according to claim 1 or 2, wherein, The VbCasX guide RNA is a dual guide RNA.

4. The mini-CRISPR-Cas system of claim 3, wherein, The VbCasX guide RNA comprises: an activator RNA comprising a nucleotide sequence as set forth in SEQ ID NO: 3, and, a targeting factor RNA comprising a nucleotide sequence as set forth in SEQ ID NO: 4, and the guide sequence.

5. The mini-CRISPR-Cas system according to any one of claims 1, 2 and 4, wherein, The VbCasX guide RNA is a single guide RNA.

6. The mini-CRISPR-Cas system according to claim 5, wherein, The VbCasX guide RNA comprises a nucleotide sequence as set forth in SEQ ID NO: 5, and the guide sequence.

7. The mini-CRISPR-Cas system according to any one of claims 1, 2, 4 and 6, wherein, The mini-CRISPR-Cas system further comprises: (iii) a donor polynucleotide.

8. A polynucleotide, comprising: i) a nucleotide sequence encoding the VbCasX protein in the mini-CRISPR-Cas system of any one of claims 1-7; ii) a nucleotide sequence encoding the VbCasX guide RNA in the mini-CRISPR-Cas system of any one of claims 1-7; and, iii) a donor polynucleotide sequence.

9. A vector comprising the polynucleotide of claim 8.

10. The vector of claim 9, wherein, The vector is an expression vector.

11. A cell comprising one or more of: (a) the mini-CRISPR-Cas system of any one of claims 1-7; (b) the polynucleotide of claim 8; and, (c) the vector of claim 9 or 10.

12. A kit comprising one or more of: (A) the mini-CRISPR-Cas system of any one of claims 1-7; (B) the polynucleotide of claim 8; (C) the vector of claim 9 or 10; and, (D) the cell of claim 11.

13. A method for modifying a target nucleic acid for non-therapeutic purposes, the method comprising the step of contacting the target nucleic acid with the mini-CRISPR-Cas system of any one of claims 1-7, the polynucleotide of claim 8, and / or the vector of claim 9 or 10.

14. Use of the mini-CRISPR-Cas system of any one of claims 1-7, the polynucleotide of claim 8, the vector of claim 9 or 10, the cell of claim 11, and / or the kit of claim 12 for modifying a target nucleic acid for non-therapeutic purposes.

Citation Information

Patent Citations

  • CRISPR / Cas system of small editing genome and special CasX protein thereof

    CN114958808A

  • Novel CRISPR enzyme, system and application

    CN115261359A