Novel crisper-cas12n enzymes and systems

By developing novel RNA-guided endonucleases and CRISPR/Cas systems, the shortcomings of existing CRISPR/Cas systems have been addressed, resulting in more efficient and precise gene editing effects and enhancing the system's robustness and multiplex gene editing capabilities.

CN113930412BActive Publication Date: 2025-12-30CHINA AGRI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010605398.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-29
Publication Date
2025-12-30
Estimated Expiration
2040-06-29

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems each have their own advantages and disadvantages in gene editing. There is a need to develop a more robust new CRISPR/Cas system with good performance in many aspects to improve the efficiency and accuracy of gene editing.

Method used

A novel RNA-guided endonuclease was developed, and a new CRISPR/Cas system was constructed based on this enzyme. The system can perform site-specific cleavage by binding to guide RNA, including the design of derivatized proteins, conjugates, and fusion proteins, to improve the robustness and multiplex gene editing capabilities of the system.

Benefits of technology

It enables more efficient and precise gene editing, reduces off-target effects, and improves the system's robustness and ability to perform multiple gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0002560898930000291
    Figure BDA0002560898930000291
  • Figure BDA0002560898930000301
    Figure BDA0002560898930000301
  • Figure BDA0002560898930000311
    Figure BDA0002560898930000311
Patent Text Reader

Abstract

The present invention relates to the field of nucleic acid editing, in particular the field of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) technology. In particular, the present invention relates to Cas effector proteins, fusion proteins comprising such proteins, and nucleic acid molecules encoding them. The present invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing) comprising the proteins or fusion proteins of the present invention, or nucleic acid molecules encoding them. The present invention also relates to methods for nucleic acid editing (e.g., gene or genome editing) using complexes comprising the proteins or fusion proteins of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nucleic acid editing, particularly to the field of regularly clustered short palindromic repeats (CRISPR) technology. Specifically, this invention relates to Cas effector proteins, fusion proteins comprising such proteins, and nucleic acid molecules encoding them. The invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing) comprising the proteins or fusion proteins of this invention, or nucleic acid molecules encoding them. The invention further relates to methods for nucleic acid editing (e.g., gene or genome editing) using proteins or fusion proteins comprising this invention. Background Technology

[0002] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to create double-strand breaks, using biological non-homologous end joining or homologous recombination for site-specific gene editing.

[0003] The CRISPR / Cas9 system is the most commonly used type II CRISPR system. It recognizes the 3'-NGG PAM motif and performs blunt-end cleavage on the target sequence. CRISPR / Cas Type V systems are a relatively new class of CRISPR systems discovered in the last two years. They possess a 5'-TTN motif and perform sticky-end cleavage on the target sequence; examples include Cpf1, C2c1, CasX, and CasY. However, the different CRISPR / Cas systems currently available each have their own advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two guide RNAs, while Cpf1 only requires one and can be used for multiplex gene editing. CasX is 980 amino acids in size, while common systems like Cas9, C2c1, CasY, and Cpf1 are typically around 1300 amino acids. Furthermore, the PAM sequences of Cas9, Cpf1, CasX, and CasY are relatively complex and diverse, while C2c1 recognizes the strict 5'-TTN, making its target site easier to predict than other systems and reducing potential off-target effects.

[0004] In conclusion, given the limitations of currently available CRISPR / Cas systems, developing a more robust new CRISPR / Cas system with superior performance in multiple aspects is of great significance to the development of biotechnology. Summary of the Invention

[0005] Through extensive experimentation and repeated exploration, the inventors of this application unexpectedly discovered a novel RNA-guided endonuclease. Based on this discovery, the inventors developed a new CRISPR / Cas system and a gene editing method based on this system.

[0006] Cas effector protein

[0007] Therefore, in a first aspect, the present invention provides a protein having an amino acid sequence shown in any one of SEQ ID NO: 1, 2, 3 or an ortholog, homolog, variant or functional fragment thereof; wherein the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived.

[0008] In this invention, the biological functions of the above sequences include, but are not limited to, binding activity with guide RNA, endonuclease activity, and binding and cleaving activity with specific sites of the target sequence under the guidance of guide RNA.

[0009] In some embodiments, the orthologs, homologs, and variants have at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence from which they originate.

[0010] In some embodiments, the ortholog, homolog, or variant has at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in any one of SEQ ID NO: 1, 2, or 3, and substantially retains the biological functions of the sequence from which it is derived (e.g., activity to bind to guide RNA, endonuclease activity, and activity to bind to and cleave a target sequence at a specific site guided by guide RNA).

[0011] In some implementations, the protein is an effector protein in a CRISPR / Cas system.

[0012] In some embodiments, the protein of the present invention comprises, or consists of, sequences selected from, the following:

[0013] (i) The sequence shown in any one of SEQ ID NO: 1, 2, or 3;

[0014] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in any one of SEQ ID NO: 1, 2, or 3; or

[0015] (iii) A sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the sequences shown in SEQ ID NO: 1, 2, or 3.

[0016] In some embodiments, the protein of the present invention has the amino acid sequence shown in any one of SEQ ID NO: 1, 2, or 3.

[0017] In some embodiments, the protein of the present invention comprises, or consists of, sequences selected from, the following:

[0018] (i) The sequence shown in SEQ ID NO: 1;

[0019] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID NO: 1; or

[0020] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 1.

[0021] In some embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO: 1.

[0022] In some embodiments, the protein of the present invention comprises, or consists of, sequences selected from, the following:

[0023] (i) The sequence shown in SEQ ID NO: 2;

[0024] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID NO: 2; or

[0025] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 2.

[0026] In some embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO: 2.

[0027] In some embodiments, the protein of the present invention comprises, or consists of, sequences selected from, the following:

[0028] (i) The sequence shown in SEQ ID NO: 3;

[0029] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID NO: 3; or

[0030] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 3.

[0031] In some embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO: 3.

[0032] In some embodiments, the protein of the present invention is able to recognize specific motif sequences.

[0033] Derived protein

[0034] The proteins of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Generally, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, the proteins of the present invention are also intended to include such derivatized forms. For example, the proteins of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical reagent, etc.

[0035] In particular, the protein of the present invention can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting region to make the protein of the present invention targeted. For example, it can be linked to a detectable tag to facilitate the detection of the protein of the present invention. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the protein of the present invention.

[0036] Conjugate

[0037] Therefore, in a second aspect, the present invention provides a conjugate comprising the protein and the modified portion as described above.

[0038] In some embodiments, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.

[0039] In some embodiments, the additional protein or polypeptide is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting moieties, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, methyltransferase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0040] In some embodiments, the conjugates of the present invention comprise one or more NLS sequences, such as the NLS of the SV40 viral large T antigen. In some exemplary embodiments, the NLS sequence is as shown in SEQ ID NO:19. In some embodiments, the NLS sequence is located at, near, or close to the end (e.g., N-terminus or C-terminus) of the protein of the present invention. In some exemplary embodiments, the NLS sequence is located at, near, or close to the C-terminus of the protein of the present invention.

[0041] In some embodiments, the conjugates of the present invention comprise an epitope tag. Such epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).

[0042] In some embodiments, the conjugates of the present invention comprise a reporter gene sequence. Such reporter genes are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0043] In some embodiments, the conjugates of the present invention comprise a domain capable of binding to DNA molecules or intracellular molecules, such as maltose-binding protein (MBP), the DNA-binding domain (DBD) of Lex A, the DBD of GAL4, etc.

[0044] In some embodiments, the conjugates of the present invention contain a detectable marker, such as a fluorescent dye, such as FITC or DAPI.

[0045] In some embodiments, the protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.

[0046] In some embodiments, the modified portion is directly attached to the N-terminus or C-terminus of the protein of the present invention.

[0047] In some embodiments, the modified portion is attached to the N-terminus or C-terminus of the protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.

[0048] Fusion protein

[0049] In a third aspect, the present invention provides a fusion protein comprising the protein of the present invention as well as other proteins or polypeptides.

[0050] In some embodiments, the additional protein or polypeptide is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting moieties, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, methyltransferase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0051] In some embodiments, the fusion protein of the present invention comprises one or more NLS sequences, such as the NLS of the SV40 viral large T antigen. In some embodiments, the NLS sequence is located at, near, or close to the end (e.g., N-terminus or C-terminus) of the protein of the present invention. In some exemplary embodiments, the NLS sequence is located at, near, or close to the C-terminus of the protein of the present invention.

[0052] In some embodiments, the fusion protein of the present invention includes an epitope tag.

[0053] In some embodiments, the fusion protein of the present invention comprises a reporter gene sequence.

[0054] In some embodiments, the fusion protein of the present invention includes a domain capable of binding to DNA molecules or intracellular molecules.

[0055] In some embodiments, the protein of the present invention is optionally fused to the additional protein or polypeptide via a linker.

[0056] In some embodiments, the additional protein or polypeptide is directly linked to the N-terminus or C-terminus of the protein of the present invention.

[0057] In some embodiments, the additional protein or polypeptide is attached to the N-terminus or C-terminus of the protein of the present invention via a linker.

[0058] In some exemplary embodiments, the fusion protein of the present invention has an amino acid sequence selected from the following: SEQ ID NO:20-22.

[0059] The proteins, conjugates, or fusion proteins of the present invention are not limited by their manner of production; for example, they can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0060] Same-direction repeating sequence

[0061] In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising, or composed of, sequences selected from, or composed of sequences selected from:

[0062] (i) The sequence shown in any one of SEQ ID NO: 10, 11, 12, 16, 17, 18;

[0063] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO: 10, 11, 12, 16, 17, 18;

[0064] (iii) A sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the sequences shown in SEQ ID NO: 10, 11, 12, 16, 17, or 18;

[0065] (iv) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iii); or

[0066] The complementary sequence of the sequence described in any of (v)(i)-(iii);

[0067] Furthermore, the sequence described in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived, the biological function of which refers to its activity as a homologous repeat sequence in the CRISPR-Cas system.

[0068] In some implementations, the isolated nucleic acid molecule is a unidirectional repeat sequence in the CRISPR-Cas system.

[0069] In some implementations, the isolated nucleic acid molecule is RNA.

[0070] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0071] (i) The sequence shown in SEQ ID NO: 10 or 16;

[0072] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO: 10 or 16;

[0073] (iii) A sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 10 or 16;

[0074] (iv) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iii); or

[0075] The complementary sequence of the sequence described in any of (v)(i)-(iii).

[0076] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0077] (a) The nucleotide sequence shown in SEQ ID NO: 10 or 16;

[0078] (b) A sequence that hybridizes with the sequence described in (a) under stringent conditions; or

[0079] (c) The complementary sequence of the nucleotide sequence shown in SEQ ID NO: 10 or 16.

[0080] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0081] (i) The sequence shown in SEQ ID NO: 11 or 17;

[0082] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO: 11 or 17;

[0083] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the sequence shown in SEQ ID NO: 11 or 17;

[0084] (iv) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iii); or

[0085] The complementary sequence of the sequence described in any of (v)(i)-(iii).

[0086] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0087] (a) The nucleotide sequence shown in SEQ ID NO: 11 or 17;

[0088] (b) A sequence that hybridizes with the sequence described in (a) under stringent conditions; or

[0089] (c) The complementary sequence of the nucleotide sequence shown in SEQ ID NO: 11 or 17.

[0090] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0091] (i) The sequence shown in SEQ ID NO: 12 or 18;

[0092] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO: 12 or 18;

[0093] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the sequence shown in SEQ ID NO: 12 or 18;

[0094] (iv) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iii); or

[0095] The complementary sequence of the sequence described in any of (v)(i)-(iii).

[0096] In some embodiments, the isolated nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0097] (a) The nucleotide sequence shown in SEQ ID NO: 12 or 18;

[0098] (b) A sequence that hybridizes with the sequence described in (a) under stringent conditions; or

[0099] (c) The complementary sequence of the nucleotide sequence shown in SEQ ID NO: 12 or 18.

[0100] CRISPR / Cas complex

[0101] In a fifth aspect, the present invention provides a composite comprising:

[0102] (i) a protein component selected from: the proteins, conjugates, or fusion proteins of the present invention, and any combination thereof; and

[0103] (ii) A nucleic acid component comprising, from the 5' to the 3' direction, isolated nucleic acid molecules as described above and a guide sequence capable of hybridizing with the target sequence.

[0104] The protein component and the nucleic acid component combine to form a complex.

[0105] In some implementations, the guide sequence is attached to the 3' end of the nucleic acid molecule.

[0106] In some implementations, the guiding sequence comprises a complementary sequence to the target sequence.

[0107] In some implementations, the nucleic acid component is a guide RNA in a CRISPR-Cas system.

[0108] In some implementations, the nucleic acid molecule is RNA.

[0109] In some embodiments, the complex does not contain trans-acting crRNA (tracrRNA).

[0110] In some embodiments, the complex targets a third component, which is a double-stranded polynucleotide containing a target sequence adjacent to the motif sequence.

[0111] In some implementations, the target sequence is located at the 3' end of the motif sequence.

[0112] In some embodiments, the distance between the target sequence and the motif sequence is less than 40, 35, 30, 25, 20, 15, 10, 5, 4, 3, 2, or 1 nucleotide.

[0113] In some embodiments, the guide sequence is at least 5, at least 10, at least 15, at least 20, at least 25, or at least 30 nucleotides in length. In some embodiments, the guide sequence is 10-30, 15-25, 15-22, 19-25, or 19-22 nucleotides in length.

[0114] In some embodiments, the isolated nucleic acid molecule is 55-70 nucleotides in length, for example, 55-65 nucleotides, for example, 60-65 nucleotides, for example, 62-65 nucleotides, for example, 63-64 nucleotides. In some embodiments, the isolated nucleic acid molecule is 15-30 nucleotides in length, for example, 15-25 nucleotides, for example, 20-25 nucleotides, for example, 22-24 nucleotides, for example, 23 nucleotides.

[0115] Encoding nucleic acids, vectors and host cells

[0116] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising:

[0117] (i) The nucleotide sequence encoding the protein or fusion protein of the present invention;

[0118] (ii) a nucleotide sequence encoding the isolated nucleic acid molecule as described in aspect four; or

[0119] (iii) Nucleotide sequences containing (i) and (ii).

[0120] In some embodiments, the nucleotide sequences described in any one of (i)-(iii) are codon-optimized for expression in prokaryotic cells. In some embodiments, the nucleotide sequences described in any one of (i)-(iii) are codon-optimized for expression in eukaryotic cells.

[0121] In a seventh aspect, the present invention also provides a vector comprising the isolated nucleic acid molecule as described in the sixth aspect. The vector of the present invention can be a cloning vector or an expression vector. In some embodiments, the vector of the present invention is, for example, a plasmid, a granulocyte, a bacteriophage, a Cosmid, etc. In some embodiments, the vector is capable of expressing the protein, fusion protein, isolated nucleic acid molecule as described in the fourth aspect, or complex as described in the fifth aspect of the present invention in a subject (e.g., a mammal, such as a human).

[0122] In an eighth aspect, the present invention also provides host cells comprising the isolated nucleic acid molecules or vectors as described above. Such host cells include, but are not limited to, prokaryotic cells such as *Escherichia coli* cells, and eukaryotic cells such as yeast cells, insect cells, plant cells, and animal cells (such as mammalian cells, such as mouse cells, human cells, etc.). The cells of the present invention can also be cell lines, such as 293T cells.

[0123] Composition and carrier composition

[0124] In a ninth aspect, the present invention also provides a composition comprising:

[0125] (i) A first component selected from: proteins, conjugates, fusion proteins, nucleotide sequences encoding said protein or fusion protein, and any combination thereof; and

[0126] (ii) A second component, which is a nucleotide sequence containing guide RNA, or a nucleotide sequence encoding the nucleotide sequence containing guide RNA;

[0127] The guide RNA contains a unidirectional repeat sequence and a guide sequence from 5' to 3', and the guide sequence is capable of hybridizing with the target sequence.

[0128] The guide RNA can form a complex with the protein, conjugate or fusion protein described in (i).

[0129] In some implementations, the unidirectional repeat sequence is a separate nucleic acid molecule as defined in the fourth aspect.

[0130] In some embodiments, the guide sequence is attached to the 3' end of the same-direction repeat sequence. In some embodiments, the guide sequence comprises a complementary sequence to the target sequence.

[0131] In some embodiments, the composition does not contain crRNA (tracrRNA).

[0132] In some embodiments, the composition is non-natural or modified. In some embodiments, at least one component of the composition is non-natural or modified. In some embodiments, the first component is non-natural or modified; and / or, the second component is non-natural or modified.

[0133] In some implementations, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM).

[0134] In some implementations, when the target sequence is RNA, the target sequence does not have a PAM domain restriction.

[0135] In some embodiments, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In some embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0136] In some embodiments, the target sequence is present within the cell. In some embodiments, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a prokaryotic cell.

[0137] In some embodiments, the protein is linked to one or more NLS sequences. In some embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0138] In a tenth aspect, the present invention also provides a composition comprising one or more carriers, said one or more carriers comprising:

[0139] (i) a first nucleic acid comprising a nucleotide sequence encoding the protein or fusion protein of the present invention; optionally, the first nucleic acid is operatively linked to a first regulatory element; and

[0140] (ii) a second nucleic acid comprising a nucleotide sequence encoding a guide RNA; optionally, the second nucleic acid is operatively linked to a second regulatory element;

[0141] in:

[0142] The first nucleic acid and the second nucleic acid may exist on the same or different vectors;

[0143] The guide RNA contains a homologous repeat sequence and a guide sequence from 5' to 3', and the guide sequence is capable of hybridizing with the target sequence;

[0144] The guide RNA can form a complex with the effector protein or fusion protein described in (i).

[0145] In some implementations, the unidirectional repeat sequence is a separate nucleic acid molecule as defined in the fourth aspect.

[0146] In some embodiments, the guide sequence is attached to the 3' end of the same-direction repeat sequence. In some embodiments, the guide sequence comprises a complementary sequence to the target sequence.

[0147] In some embodiments, the composition does not contain trans-acting crRNA (tracrRNA).

[0148] In some embodiments, the composition is non-natural or modified. In some embodiments, at least one component of the composition is non-natural or modified.

[0149] In some implementations, the first regulating element is a promoter, such as an inducible promoter.

[0150] In some implementations, the second regulating element is an promoter, such as an inductive promoter.

[0151] In some implementations, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM).

[0152] In some implementations, when the target sequence is RNA, the target sequence does not have a PAM domain restriction.

[0153] In some embodiments, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In some embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0154] In some embodiments, the target sequence is present within the cell. In some embodiments, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a prokaryotic cell.

[0155] In some embodiments, the protein is linked to one or more NLS sequences. In some embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0156] In some implementations, one type of vector is a plasmid, which refers to a circular double-stranded DNA loop in which additional DNA fragments can be inserted, for example, using standard molecular cloning techniques. Another type of vector is a viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells in which they are introduced. Other vectors (e.g., non-episodic mammalian vectors) integrate into the genome of the host cell after introduction and thereby replicate along with the host genome. Moreover, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors." Common expression vectors used in recombinant DNA technologies are typically in plasmid form.

[0157] Recombinant expression vectors may contain nucleic acid molecules of the present invention in a form suitable for nucleic acid expression in host cells, meaning that these recombinant expression vectors contain one or more regulatory elements selected based on the host cell to be used for expression, the regulatory elements being operatively linked to the nucleic acid sequence to be expressed.

[0158] Delivery and delivery composition

[0159] The proteins, conjugates, fusion proteins, isolated nucleic acid molecules as described in the fourth aspect, complexes of the present invention, isolated nucleic acid molecules as described in the sixth aspect, vectors as described in the seventh aspect, and compositions as described in the ninth and tenth aspects of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.

[0160] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery carrier and one or more selected from the following: proteins, conjugates, fusion proteins, isolated nucleic acid molecules as described in the fourth aspect, complexes of the present invention, isolated nucleic acid molecules as described in the sixth aspect, carriers as described in the seventh aspect, and compositions as described in the ninth and tenth aspects.

[0161] In some implementations, the delivery carrier is a particle.

[0162] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0163] Reagent test kit

[0164] In another aspect, the present invention provides a kit comprising one or more of the components described above. In some embodiments, the kit comprises one or more components selected from: proteins, conjugates, fusion proteins, isolated nucleic acid molecules as described in the fourth aspect, complexes of the present invention, isolated nucleic acid molecules as described in the sixth aspect, vectors as described in the seventh aspect, and compositions as described in the ninth and tenth aspects.

[0165] In some embodiments, the kit of the present invention comprises the composition as described in aspect nine. In some embodiments, the kit further comprises instructions for using the composition.

[0166] In some embodiments, the kit of the present invention comprises the composition as described in aspect ten. In some embodiments, the kit further comprises instructions for using the composition.

[0167] In some embodiments, the components contained in the kit of the present invention can be provided in any suitable container.

[0168] In some embodiments, the kit further comprises one or more buffers. The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH from about 7 to about 10.

[0169] In some embodiments, the kit further includes one or more oligonucleotides corresponding to a guide sequence for insertion into the vector to operatively link the guide sequence and a regulatory element. In some embodiments, the kit includes homologous recombinant template polynucleotides.

[0170] Methods and Applications

[0171] In another aspect, the present invention provides a method for modifying a target gene, comprising: contacting the target gene with a complex as described in the fifth aspect, a composition as described in the ninth aspect, or a composition as described in the tenth aspect, or delivering it to a cell containing the target gene; wherein the target sequence is present in the target gene.

[0172] In some embodiments, the target gene is present within a cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is selected from non-human primate, bovine, pig, or rodent cells. In some embodiments, the cell is a non-mammalian eukaryotic cell, such as a poultry or fish cell. In some embodiments, the cell is a plant cell, such as cells of cultivated plants (e.g., cassava, corn, sorghum, wheat, or rice), algae, trees, or vegetables.

[0173] In some embodiments, the target gene is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0174] In some implementations, the method results in a break in the target sequence (e.g., a double-strand break in DNA or a single-strand break in RNA).

[0175] In some implementations, the break results in a reduction in the transcription of the target gene.

[0176] In some embodiments, the method further includes contacting the target gene with an editing template (e.g., a foreign nucleic acid) or delivering it to a cell containing the target gene. In such embodiments, the method repairs the broken target gene through homologous recombination with the editing template (e.g., a foreign nucleic acid), wherein the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target gene. In some embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence.

[0177] Therefore, in some embodiments, the modification also includes inserting an editing template (e.g., exogenous nucleic acid) into the break.

[0178] In some embodiments, the protein, conjugate, fusion protein, isolated nucleic acid molecule, complex, carrier, or composition is contained in a delivery vector.

[0179] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0180] In some embodiments, the method is used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0181] In another aspect, the present invention provides a method for altering the expression of a gene product, comprising: contacting a complex as described in the fifth aspect, a composition as described in the ninth aspect, or a composition as described in the tenth aspect with a nucleic acid molecule encoding the gene product, or delivering it to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0182] In some embodiments, the nucleic acid molecules are present within a cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is selected from non-human primate, bovine, pig, or rodent cells. In some embodiments, the cell is a non-mammalian eukaryotic cell, such as a poultry or fish cell. In some embodiments, the cell is a plant cell, such as a cell found in cultivated plants (e.g., cassava, corn, sorghum, wheat, or rice), algae, trees, or vegetables.

[0183] In some embodiments, the nucleic acid molecule is present in an in vitro nucleic acid molecule (e.g., a plasmid). In some embodiments, the nucleic acid molecule is present in a plasmid.

[0184] In some embodiments, the expression of the gene product is altered (e.g., enhanced or reduced). In some embodiments, the expression of the gene product is enhanced. In some embodiments, the expression of the gene product is reduced.

[0185] In some implementations, the gene product is a protein.

[0186] In some embodiments, the protein, conjugate, fusion protein, isolated nucleic acid molecule, complex, carrier, or composition is contained in a delivery vector.

[0187] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0188] In some embodiments, the method is used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0189] In another aspect, the present invention relates to proteins as described in the first aspect, conjugates as described in the second aspect, fusion proteins as described in the third aspect, isolated nucleic acid molecules as described in the fourth aspect, complexes as described in the fifth aspect, isolated nucleic acid molecules as described in the sixth aspect, vectors as described in the seventh aspect, compositions as described in the ninth aspect, compositions as described in the tenth aspect, kits or delivery compositions of the present invention, for use in nucleic acid editing.

[0190] In some implementations, the nucleic acid editing includes gene or genome editing, such as modifying genes, knocking out genes, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides.

[0191] In another aspect, the present invention relates to proteins as described in the first aspect, conjugates as described in the second aspect, fusion proteins as described in the third aspect, isolated nucleic acid molecules as described in the fourth aspect, complexes as described in the fifth aspect, isolated nucleic acid molecules as described in the sixth aspect, carriers as described in the seventh aspect, compositions as described in the ninth aspect, compositions as described in the tenth aspect, kits or delivery compositions of the present invention, and their use in the preparation of formulations for:

[0192] (i) In vitro gene or genome editing;

[0193] (ii) Detection of isolated single-stranded DNA;

[0194] (iii) Editing target sequences in target loci to modify biological or non-human organisms;

[0195] (iv) Treating conditions caused by defects in target sequences at target loci.

[0196] Cells and progeny

[0197] In some cases, modifications introduced into cells by the methods of the present invention can alter cells and their progeny to improve the production of their biological products (such as antibodies, starch, ethanol, or other desired cellular outputs). In other cases, modifications introduced into cells by the methods of the present invention can include changes in cells and their progeny that alter the biological products produced.

[0198] Therefore, in another aspect, the present invention also relates to cells or their progeny obtained by the method described above, wherein said cells contain modifications not present in their wild type.

[0199] The present invention also relates to cell products of the cells or their progeny as described above.

[0200] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: a protein as described in the first aspect, a conjugate as described in the second aspect, a fusion protein as described in the third aspect, an isolated nucleic acid molecule as described in the fourth aspect, a complex as described in the fifth aspect, an isolated nucleic acid molecule as described in the sixth aspect, a carrier as described in the seventh aspect, a composition as described in the ninth aspect, a composition as described in the tenth aspect, a kit or delivery composition of the present invention.

[0201] In some implementations, the cell is a prokaryotic cell.

[0202] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).

[0203] In some implementations, the cell is a stem cell or stem cell line.

[0204] Terminology Definition

[0205] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.

[0206] In this invention, the term "Cas12N" refers to a Cas effector protein that the inventors first discovered and identified, which has an amino acid sequence selected from the following:

[0207] (i) The sequences shown in SEQ ID NO: 1-3;

[0208] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID NO: 1-3; or

[0209] (iii) A sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 1-3.

[0210] The Cas12N of this invention is a nuclease that binds to and cleaves a specific site of a target sequence under the guidance of a guide RNA, and has both DNA and RNA endonuclease activities.

[0211] As used herein, the terms “regularly clustered interspaced short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements associated with the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes. Such transcripts or other elements may include sequences encoding Cas effector proteins and guide RNA containing CRISPR RNA (crRNA), as well as trans-acting crRNA (tracrRNA) sequences found in CRISPR-Cas9 systems, or other sequences or transcripts derived from CRISPR loci. In the Cas12N-based CRISPR system described in this invention, tracrRNA sequences are not required.

[0212] As used herein, the terms "Cas effector protein" and "Cas effector enzyme" are used interchangeably and refer to any protein longer than 800 amino acids presented in the CRISPR-Cas system. In some cases, this category of proteins refers to proteins identified from Cas loci.

[0213] As used herein, the terms “guide RNA” and “mature crRNA” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may comprise a direct repeat sequence and a guide sequence, or consist substantially of or composed of a direct repeat sequence and a guide sequence (also referred to as a spacer sequence in the context of an endogenous CRISPR system). In some cases, the guide sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas complex to said target sequence. In some embodiments, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining optimal alignment is within the capabilities of those skilled in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.

[0214] In some cases, the guide sequence has a length of at least 5, at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides. In some cases, the guide sequence has a length of no more than 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides. In some embodiments, the guide sequence has a length of 10-30, or 15-25, or 15-22, or 19-25, or 19-22 nucleotides.

[0215] In some cases, the same-direction repeat sequence has a length of at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, or at least 70 nucleotides. In some cases, the directionally repeating sequence has a length of no more than 70, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides. In some embodiments, the directionally repeating sequence has a length of 55-70 nucleotides, for example 55-65 nucleotides, for example 60-65 nucleotides, for example 62-65 nucleotides, for example 63-64 nucleotides. In some embodiments, the directionally repeating sequence has a length of 15-30 nucleotides, for example 15-25 nucleotides, for example 20-25 nucleotides, for example 22-24 nucleotides, for example 23 nucleotides.

[0216] As used herein, the term "CRISPR / Cas complex" refers to a ribonucleoprotein complex formed by the binding of guide RNA or mature crRNA to the Cas protein, which contains a guide sequence that hybridizes to a target sequence and binds to the Cas protein. This ribonucleoprotein complex is capable of recognizing and cleaving polynucleotides that hybridize with the guide RNA or mature crRNA.

[0217] Therefore, in the formation of a CRISPR / Cas complex, a "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can contain any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. The sequence or template that can be used for recombination into a target locus containing the target sequence is called an "edit template," "edit polynucleotide," or "edit sequence." In some embodiments, the edit template is a foreign nucleic acid. In some embodiments, the recombination is homologous recombination.

[0218] In this invention, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence is believed to be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas effector enzyme used, but the PAM is typically a 2-5 base pair sequence adjacent to the protospacer sequence (i.e., the target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas effector protein. In this document, "specific motif sequence recognized by the Cas protein" or "motif sequence" refers to the PAM sequence.

[0219] In some cases, target sequences or target polynucleotides may include multiple disease-related genes and polynucleotides, as well as genes and polynucleotides related to signal transduction biochemical pathways. Non-limiting examples of such target sequences or target polynucleotides include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed December 12, 2012 and January 2, 2013, respectively, and International Application PCT / US2013 / 074667, filed December 12, 2013, all of which are incorporated herein by reference.

[0220] In some cases, examples of target sequences or target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-related genes or polynucleotides. A “disease-related” gene or polynucleotide refers to any gene or polynucleotide that produces its transcriptional or translational product at an abnormal level or in an abnormal form in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control tissues or cells. In cases where altered expression is associated with the onset and / or progression of disease, it can be a gene expressed at an abnormally high level; or it can be a gene expressed at an abnormally low level. Disease-related genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked to one or more genes responsible for the etiology of the disease in disequilibrium. The transcribed or translated product can be known or unknown and can be at normal or abnormal levels.

[0221] As used herein, the term “wildtype” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, is separable from its natural source and has not been intentionally modified by humans.

[0222] As used herein, the terms “non-naturally occurring” or “engineered” are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they indicate that the nucleic acid molecule or peptide is at least substantially free from at least one other component bound to it, either naturally occurring or found in nature.

[0223] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein, as described herein, refers to a protein belonging to a different species that performs the same or similar function as the protein that is its orthologue.

[0224] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomeric subunit occupies the same location (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum identity. Such comparisons can be made using methods readily available, for example, computer programs such as the Align program (DNAstar, Inc.) Needleman et al. (1970) J. Mol. Biol. 48: 443-453. The percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) integrated into the ALIGN program (version 2.0), which uses a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the Needleman and Wunsch algorithm (J MoIBiol. 48:444-453 (1970)) in the GAP program integrated into the GCG software package (available at www.gcg.com), which uses a Blossum 62 matrix or a PAM250 matrix, along with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.

[0225] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. When a vector enables the expression of a protein encoded by the inserted polynucleotide, it is called an expression vector. Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material elements they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to: plasmids; phage particles; Cos plasmids; artificial chromosomes, such as yeast artificial chromosomes (YAC), bacterial artificial chromosomes (BAC), or P1-derived artificial chromosomes (PAC); bacteriophages such as λ phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retrotranscriptoviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpesviruses (such as herpes simplex virus), poxviruses, baculoviruses, papillomaviruses, and papillomaviruses (such as SV40). A vector may contain multiple elements controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, a vector may contain a replication initiation site.

[0226] As used herein, the term "host cell" refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, fungal cells such as yeast cells or Aspergillus, insect cells such as S2 Drosophila cells or Sf9, or animal cells such as fibroblasts, CHO cells, COS cells, NSO cells, HeLa cells, BHK cells, HEK 293 cells, or human cells.

[0227] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes) as described herein.

[0228] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed descriptions can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0229] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.

[0230] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).

[0231] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.

[0232] As used herein, “strict conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of strict conditions is described in Tijssen (1993), *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes*, Part I, Chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, New York.

[0233] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding between the bases of these nucleotide residues. Hydrogen bonding can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can constitute a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.

[0234] As used herein, the term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.

[0235] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of multiple amino acid residues via peptide bonds. The linkers of this invention can be synthetically produced amino acid sequences or naturally occurring polypeptide sequences, such as polypeptides with hinge region functions. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ et al. (1994) Structure 2:1121-1123).

[0236] As used in this article, the term "treatment" means to treat or cure a disease, to delay the onset of symptoms of a disease, and / or to slow the progression of a disease.

[0237] As used herein, the term "subject" includes, but is not limited to, various animals, such as mammals, including bovines, equines, sheep, suidae, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect).

[0238] Beneficial effects of the invention

[0239] Compared with existing technologies, the Cas protein and system of the present invention have significant advantages. For example, the Cas effector protein of the present invention is smaller in size than Cas9, C2c1, CasY, and Cpf1 proteins, and therefore has superior transfection efficiency. For example, the Cas effector protein of the present invention can cleave DNA in eukaryotes, and compared with the previously reported FnCpf1 with a 5'-TTN PAM domain, the Cas protein of the present invention exhibits significantly stronger cleavage activity in human cell lines. For example, the Cas protein of the present invention has a more stringent PAM recognition method, which can reduce off-target effects.

[0240] The embodiments of the present invention will be described in detail below through examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of preferred embodiments.

[0241] Sequence information

[0242] Information on some of the sequences involved in this invention is provided in Table 1 below.

[0243] Table 1: Sequence Description

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254] Detailed Implementation

[0255] The invention will now be described with reference to the following embodiments, which are intended to illustrate the invention (and not limit it).

[0256] Unless otherwise specified, the experiments and methods described in the examples are generally performed in accordance with conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in this invention can be found in Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd edition (1989); *Current Protocols in Molecular Biology* (edited by FM. Ausubel et al., (1987)); and the *Methods in Enzymology* series (academic publishing company): *PCR 2: A PRACTICAL*. APPROACH (edited by MJ MacPherson, BD Hames and GR Taylor (1995)), Harlow and Lane (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (edited by R.R. Freshney (1987)).

[0257] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0258] The sources of some of the reagents involved in the following examples are as follows:

[0259] LB liquid medium: 10g tryptone, 5g yeast extract, 10g NaCl, bring to a final volume of 1L, and sterilize. If antibiotics are required, add them after the medium has cooled, at a final concentration of 50μg / ml.

[0260] Chloroform / Isoamyl alcohol: Add 10ml of isoamyl alcohol to 240ml of chloroform and mix well.

[0261] RNP buffer: 100mM sodium chloride, 50mM Tris-HCl, 10mM MgCl2, 100μg / ml BSA, pH 7.9.

[0262] The prokaryotic expression vectors pACYC-Duet-1 and pUC19 were purchased from Beijing TransGen Biotech Co., Ltd.

[0263] Escherichia coli competent cells EC100 were purchased from Epicentre.

[0264] Example 1. Obtaining the Cas12N gene and Cas12N guide RNA

[0265] 1. CRISPR and gene annotation: Prodigal was used to annotate all proteins from microbial genome and metagenomic data in the NCBI and JGI databases. Piler-CR was used to annotate CRISPR loci, with all parameters set to default.

[0266] 2. Protein Filtering: Redundancy removal of annotated proteins is performed based on sequence consistency, eliminating proteins with completely identical sequences. Proteins longer than 800 amino acids are classified as macromolecules. Since most effector proteins discovered so far in Class II CRISPR / Cas systems are longer than 900 amino acids, to reduce computational complexity, we only consider macromolecules when mining CRISPR effector proteins.

[0267] 3. Obtaining CRISPR-related macromolecular proteins: Extending 10Kb upstream and downstream of each CRISPR locus, non-redundant macromolecular proteins in the vicinity of the CRISPR locus will be identified.

[0268] 4. Clustering of CRISPR-related macromolecular proteins: BLASTP was used to perform pairwise alignments of non-redundant macromolecular CRISPR-related proteins, outputting alignment results with an Evalue < 1E-10. MCL was then used to perform cluster analysis on the BLASTP outputs to identify CRISPR-related protein families.

[0269] 5. Identification of CRISPR-enriched macromolecular protein families: BLASTP was used to align CRISPR-related protein families to a non-redundant macromolecular protein database after removing CRISPR-related proteins, outputting alignment results with an Evalue < 1E-10. If less than 100% of the homologous proteins were found in a non-CRISPR-related protein database, it indicates that this protein family is enriched in the CRISPR region. This method was used to identify CRISPR-enriched macromolecular protein families.

[0270] 6. Protein Function and Domain Annotation: Using the Pfam database, NR database, and Cas proteins collected from NCBI, CRISPR-enriched macromolecular protein families were annotated to obtain new CRISPR / Cas protein families. Multiple sequence alignment of each CRISPR / Cas family protein was performed using Maft, followed by conserved domain analysis using JPred and HHpred to identify protein families containing the RuvC domain.

[0271] Based on this, the inventors obtained a novel Cas effector protein, namely Cas12N, and its three active homolog sequences, named Cas12N.1 (SEQ ID NO:1), Cas12N.2 (SEQ ID NO:2), and Cas12N.3 (SEQ ID NO:3), respectively. The coding DNA sequences of the three homologs are shown in SEQ ID NO:4, SEQ ID NO:5, and SEQ ID NO:6, respectively. The prototype homologous repeat sequence (the repeat sequence contained in pre-crRNA) corresponding to Cas12N.1 is shown in SEQ ID NO:7, the prototype homologous repeat sequence (the repeat sequence contained in pre-crRNA) corresponding to Cas12N.2 is shown in SEQ ID NO:8, the prototype homologous repeat sequence (the repeat sequence contained in pre-crRNA) corresponding to Cas12N.3 is shown in SEQ ID NO:9, the mature homologous repeat sequence (the repeat sequence contained in mature crRNA) corresponding to Cas12N.1 is shown in SEQ ID NO:13, the mature homologous repeat sequence (the repeat sequence contained in mature crRNA) corresponding to Cas12N.2 is shown in SEQ ID NO:14, and the mature homologous repeat sequence (the repeat sequence contained in mature crRNA) corresponding to Cas12N.2 is shown in SEQ ID NO:15.

[0272] Example 2. Processing of primitive crRNA by the Cas12N gene

[0273] 1. The double-stranded DNA molecule shown in SEQ ID NO:4 was artificially synthesized, and the double-stranded DNA molecule shown in SEQ ID NO:10 was also artificially synthesized.

[0274] 2. The double-stranded DNA molecule synthesized in step 1 was ligated to the prokaryotic expression vector pACYC-Duet-1 to obtain the recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1.

[0275] The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 was sequenced. Sequencing results showed that the recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 contained the sequences shown in SEQ ID NO:4 and SEQ ID NO:10, and expressed the Cas12N.1 protein shown in SEQ ID NO:1 and the prototypical homologous repeat sequence of Cas12N.1 shown in SEQ ID NO:7. The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 was introduced into *E. coli* EC100 to obtain a recombinant bacterium, which was named EC100 / pACYC-Duet-1+CRISPR / Cas12N.1.

[0276] 3. Take a single clone of EC100 / pACYC-Duet-1+CRISPR / Cas12N.1, inoculate it into 100mL LB liquid medium (containing 50μg / mL ampicillin), and culture at 37℃ and 200rpm for 12h to obtain the culture solution.

[0277] 4. Extracting bacterial RNA: Transfer 1.5 mL of bacterial culture to a pre-chilled microcentrifuge tube and centrifuge at 6000 × g for 5 minutes at 4°C. After centrifugation, discard the supernatant and resuspend the cell pellet in 200 μL of Max Bacterial Enhancement Reagent preheated to 95°C. Mix thoroughly by pipetting. Incubate at 95°C for 4 minutes. Add 1 mL of the dissolved product to the solution. Reagent and mix by pipetting, incubate at room temperature for 5 minutes. Add 0.2 mL of cold chloroform, mix by hand for 15 seconds, and incubate at room temperature for 2-3 minutes. Centrifuge at 12,000 × g for 15 minutes at 4°C. Transfer 600 μL of supernatant to a new tube, add 0.5 mL of cold isopropanol to precipitate RNA, mix by inversion, and incubate at room temperature for 10 minutes. Centrifuge at 15,000 × g for 10 minutes at 4°C, discard the supernatant, add 1 mL of 75% ethanol, and vortex to mix. Centrifuge at 7500 × g for 5 minutes at 4°C, discard the supernatant, and air dry. Dissolve the RNA precipitate in 50 μL of RNase-free water and incubate at 60°C for 10 minutes.

[0278] 5. DNA Digestion: Dissolve 20 μg of RNA in 39.5 μL ddH2O and incubate at 65°C for 5 min. Incubate on ice for 5 min, then add 0.5 μL RNAI, 5 μL buffer, and 5 μL DNaseI. Incubate at 37°C for 45 min (50 μL system). Add 50 μL ddH2O and adjust the volume to 100 μL. Centrifuge at 16000g for 30 s in a 2 mL Phase-Lock tube. Add 100 μL of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μL of digested RNA, shake for 15 s, and centrifuge at 16000g for 12 min at 15°C. Transfer the supernatant to a new 1.5 mL centrifuge tube, add an equal volume of isopropanol (1 / 10 NaoAC), and react for 1 h or overnight at -20°C. Centrifuge at 16000g for 30 min at 4°C and discard the supernatant. Wash the precipitate with 350 μL of 75% ethanol, centrifuge at 16000g for 10 min at 4°C, and discard the supernatant. Air dry, add 20 μL of RNase-free water, and dissolve the precipitate at 65°C for 5 min. Measure the concentration using NanoDrop and run a gel electrophoresis.

[0279] 6. 3' Dephosphorylation and 5' Phosphorylation: Add water to 42.5 μL of each of the 20 μg digested RNA and incubate at 90°C for 2 min. Cool on ice for 5 min. Add 5 μL of 10×T4 PNK buffer, 0.5 μL of RNaI, and 2 μL of T4 PNK (50 μL), and incubate at 37°C for 6 h. Add 1 μL of T4 PNK and 1.25 μL (100 mM) of ATP, and incubate at 37°C for 1 h. Add 47.75 μL of ddH2O and adjust the volume to 100 μL. Centrifuge at 16000g for 30 s using a 2 mL Phase-Locktube centrifuge. Add 100 μL of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μL of digested RNA, shake for 15 s, and centrifuge at 16000g for 12 min at 15°C. Transfer the supernatant to a new 1.5 mL centrifuge tube, add an equal volume of isopropanol (total volume 1 / 10 NaoAC), and react for 1 h or overnight at -20°C. Centrifuge at 16000 g for 30 min at 4°C, and discard the supernatant. Wash the precipitate with 350 μL of 75% ethanol, centrifuge at 16000 g for 10 min at 4°C, and discard the supernatant. Air dry, add 21 μL of RNase-free water, and dissolve the precipitate at 65°C for 5 min. Measure the concentration using NanoDrop.

[0280] 7. RNA Monophosphorylation: 20 μL RNA, incubate at 90℃ for 1 min, then cool on ice for 5 min. Add 2 μL RNA 5'Polphosphatase 10× Reaction buffer, 0.5 μL Inhibitor, 1 μL RNA 5'Polphosphatase (20 Units), and RNase-free water to 20 μL. Incubate at 37℃ for 60 min. Add 80 μL ddH2O to adjust the volume to 100 μL. Centrifuge at 16000g for 30 s in a 2 mL Phase-Lock tube. Add 100 μL phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μL digested RNA, shake for 15 s, and centrifuge at 16000g for 12 min at 15℃. Transfer the supernatant to a new 1.5 mL centrifuge tube, add an equal volume of isopropanol (total volume 1 / 10 NaoAC), and react for 1 h or overnight at -20℃. Centrifuge at 16000g for 30 min at 4℃, discard the supernatant, add 350μL of 75% ethanol to wash the precipitate, centrifuge at 16000g for 10 min at 4℃, discard the supernatant. Air dry, add 21μL of RNase-free water, incubate at 65℃ for 5 min to dissolve the precipitate, and measure the concentration using NanoDrop.

[0281] 8. Preparation of cDNA library: 16.5 μL RNase-free water; 5 μL Poly(A) Polymerase 10× Reaction buffer; 5 μL 10mM ATP; 1.5 μL RiboGuard RNase Inhibitor; 20 μL RNA Substrate; 2 μL Poly(A) Polymerase (4 Units); 50 μL total volume. Incubate at 37℃ for 20 min. Add 50 μL dH2O to adjust the volume to 100 μL. Centrifuge at 16000g for 30 s in a 2 mL Phase-Lock tube, then add 100 μL phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μL digested RNA, shake for 15 s, and centrifuge at 16000g for 12 min at 15℃. Transfer the supernatant to a new 1.5 mL centrifuge tube, add an equal volume of isopropanol (total volume 1 / 10 NaoAC), and react for 1 h or overnight at -20 °C. Centrifuge at 16000 g for 30 min at 4 °C, discard the supernatant, air dry, add 11 μL of RNase-free water, and dissolve the precipitate at 65 °C for 5 min. Measure the concentration using NanoDrop.

[0282] 9. After adding sequencing adapters, the cDNA library was sent to Beijing Berry Genomics for sequencing.

[0283] 10. Perform quality filtering on the raw data to remove sequences with an average base quality value below 30. After removing adapters from the sequences, retain RNA sequences of 25 to 50 nt and align them to the reference sequence of the CRISPR array using bowtie.

[0284] 11. Similar to the steps above, recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.2 containing the sequences shown in SEQ ID NO:5 and SEQ ID NO:11, and recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.3 containing the sequences shown in SEQ ID NO:6 and SEQ ID NO:12 were constructed, transformed into E. coli, cDNA libraries were prepared, and data alignment was performed.

[0285] 12. Through comparison, we found that the pre-crRNA of Cas12N.1, Cas12N.2 and Cas12N.3 can be successfully processed into mature crRNA in E. coli. The mature crRNA consists of a repeat sequence and a guide sequence.

[0286] 13. We performed structural prediction and visualization analysis on mature crRNAs using ViennaRNA and VARNA, respectively. We found that the 3' end of the repeat sequence of Cas12N.1, Cas12N.2, and Cas12N.3 crRNAs can form a neck loop.

[0287] Example 3. Identification of the PAM domain of the Cas12N gene

[0288] 1. The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 was constructed and sequenced. Based on the sequencing results, the structure of the recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 is described as follows: The small fragment between the restriction endonuclease Pml I and Kpn I recognition sequences of the vector pACYC-Duet-1 was replaced with the sequence shown in SEQ ID NO:4, and the prototype direct repeat sequence shown in SEQ ID NO:7 and the guide sequence identified by the PAM domain of Cas12N.1 shown in SEQ ID NO:27 were ligated into the vector. The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 expresses the Cas12N.1 protein shown in SEQ ID NO:1.

[0289] 2. The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 contains an expression cassette, the nucleotide sequence of which is shown in SEQ ID NO:23. In the sequence shown in SEQ ID NO:23, positions 1 to 44 from the 5' end are the nucleotide sequence of the pLacZ promoter, positions 45 to 2870 are the nucleotide sequence of the Cas12N.1 gene, and positions 2871 to 2957 are the nucleotide sequence of the terminator (used to terminate transcription). Positions 2958 to 2992 from the 5' end are the nucleotide sequence of the J23119 promoter, positions 2993 to 3044 are the nucleotide sequence of the CRISPR array, and positions 3045 to 3071 are the nucleotide sequence of the rrnB-T1 terminator (used to terminate transcription).

[0290] 3. Obtaining recombinant Escherichia coli: The recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.1 was introduced into Escherichia coli EC100 to obtain recombinant Escherichia coli, named EC100 / pACYC-Duet-1+CRISPR / Cas12N.1. The recombinant plasmid pACYC-Duet-1 was introduced into Escherichia coli EC100 to obtain recombinant Escherichia coli, named EC100 / pACYC-Duet-1.

[0291] 4. Construction of the PAM library: The sequence shown in SEQ ID NO:26 was synthesized and ligated into the pUC19 vector. The sequence shown in SEQ ID NO:26 includes eight random bases at the 5' end and the target sequence. An eight-base sequence preceding the 5' end of the target sequence of the PAM library was designed to construct a plasmid library. The plasmid was transformed into *E. coli* containing the CRISPR / Cas12N.1 locus and *E. coli* without the CRISPR / Cas12N.1 locus, respectively. After incubation at 37°C for 1 hour, the plasmid was extracted, and the PAM region sequence was amplified by PCR and sequenced.

[0292] 5. Similarly, following the steps described above, we ligated the prototype direct repeat sequence shown in SEQ ID NO:8 and the guide sequence identified by the PAM domain shown in SEQ ID NO:27 into the vector to construct the recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.2. We ligated the prototype direct repeat sequence shown in SEQ ID NO:9 and the guide sequence identified by the PAM domain shown in SEQ ID NO:27 into the vector to construct the recombinant plasmid pACYC-Duet-1+CRISPR / Cas12N.3. The expression cassettes of the two recombinant plasmids have the nucleotide sequences shown in SEQ ID NO:24 and SEQ ID NO:25, respectively. In the sequence shown in SEQ ID NO:24, positions 1 to 44 from the 5' end are the nucleotide sequence of the pLacZ promoter, positions 45 to 2240 are the nucleotide sequence of the Cas12N.2 gene, and positions 2241 to 2327 are the nucleotide sequence of the terminator (used to terminate transcription). From the 5' end, positions 2328 to 2362 are the nucleotide sequence of the J23119 promoter, positions 2363 to 2426 are the nucleotide sequence of the CRISPR array, and positions 2427 to 2453 are the nucleotide sequence of the rrnB-T1 terminator (used to terminate transcription). In the sequence shown in SEQ ID NO:25, from the 5' end, positions 1 to 44 are the nucleotide sequence of the pLacZ promoter, positions 45 to 2327 are the nucleotide sequence of the Cas12N.3 gene, and positions 2328 to 2414 are the nucleotide sequence of the terminator (used to terminate transcription). From the 5' end, positions 2415 to 2449 are the nucleotide sequence of the J23119 promoter, positions 2450 to 2515 are the nucleotide sequence of the CRISPR array, and positions 2516 to 2542 are the nucleotide sequence of the rrnB-T1 terminator (used to terminate transcription).

[0293] 6. Obtaining the PAM library structure domains: The occurrence frequency of PAM sequences in the experimental and control groups was counted, and the sequences were standardized using the total number of PAM sequences in each group. For any PAM sequence, if log2 (standardized value of control group / standardized value of experimental group) is greater than 3.5, we consider this PAM sequence to be significantly consumed. We obtained the significantly consumed PAM sequences from all PAM sequences. Furthermore, Weblogo was used to predict the significantly consumed PAM sequences, ultimately obtaining the PAM structure domains of Cas12N.1, Cas12N.2, and Cas12N.3.

[0294] 7. Validation of PAM Library Domains: Through PAM library consumption experiments, we obtained the PAM domains of Cas12N.1, Cas12N.2, and Cas12N.3. To verify the rigor of this domain, we conducted in vivo experiments with 10 groups of PAMs and sequenced the editing activity of Cas12N.1 on these PAMs. First, we integrated a 30 nt target and PAM sequence into the non-conserved position of the kanamycin-resistant gene in the plasmid, and then mixed it with a complex formed by CRISPR / Cas12N.1 and guide RNA and cultured it for 8 hours. By plating and counting colonies, we were able to determine the consumption activity of Cas12N.1 on different PAM sequences. The experimental results show that the CRISPR / Cas12N.1 system can only effectively edit target sequences with specific PAM domains, while it has no editing activity on other target sequences, thus verifying the accuracy of Cas12N.1's PAM domain recognition. The experimental results for Cas12N.2 and Cas12N.3 are similar to those for Cas12N.1.

[0295] Example 4. Identification of DNA cutting mode of CRISPR / Cas12N system

[0296] I. In vitro expression and purification of Cas12N protein

[0297] The specific steps for in vitro expression and purification of Cas12N.1 protein are as follows:

[0298] 1. The nucleotide sequence shown in SEQ ID NO:4 was artificially synthesized.

[0299] 2. The double-stranded DNA molecule synthesized in step 1 was ligated to the prokaryotic expression vector pET-30a(+) to obtain the recombinant plasmid pET-30a-CRISPR / Cas12N.1. The recombinant plasmid pET-30a-CRISPR / Cas12N.1 was sequenced. Sequencing results showed that the recombinant plasmid pET-30a-CRISPR / Cas12N.1 expressed the Cas12N protein with a nuclear localization signal as shown in SEQ ID NO:20.

[0300] 3. The recombinant plasmid pET-30a-CRISPR / Cas12N.1 was introduced into Escherichia coli EC100 to obtain a recombinant bacterium, which was named EC100-CRISPR / Cas12N.1. A single colony of EC100-CRISPR / Cas12N.1 was picked and inoculated into 100 mL of LB liquid medium (containing 50 μg / mL ampicillin), and cultured at 37 °C with shaking at 200 rpm for 12 h to obtain the bacterial culture.

[0301] 4. Take the cultured bacterial suspension and inoculate it into 50 mL of LB liquid medium (containing 50 μg / mL ampicillin) at a volume ratio of 1:100. Incubate at 37°C with shaking at 200 rpm until OD reaches the target value. 600nm The value was 0.6, then IPTG was added to make the concentration 1mM, and the mixture was cultured at 28℃ and 220rpm for 4h with shaking. After centrifugation at 4℃ and 10000rpm for 10min, the bacterial pellet was collected.

[0302] 5. Take the bacterial cell pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend, and then sonicate to disrupt (ultrasonic power 600 W, cycle program: disrupt for 4 s, pause for 6 s, for a total of 20 min). Then centrifuge at 4 °C and 10,000 rpm for 10 min and collect the supernatant A.

[0303] 6. Take the supernatant A, centrifuge at 4℃ and 12000rpm for 10min, and collect the supernatant B.

[0304] 7. The supernatant B was purified using a nickel column manufactured by GE (refer to the nickel column instruction manual for specific purification steps), and then the Cas12N.1 protein was quantified using a protein quantification kit manufactured by Thermo Fisher Scientific.

[0305] 8. Similarly, repeat the above steps to construct recombinant plasmids pET-30a-CRISPR / Cas12N.2 and pET-30a-CRISPR / Cas12N.3. The two recombinant plasmids express the Cas12N protein with nuclear localization signals as shown in SEQ ID NO:21 and SEQ ID NO:22, respectively, and quantify the Cas12N.2 and Cas12N.3 proteins.

[0306] II. Transcription and purification of Cas12N protein guide RNA:

[0307] 1. Templates for guide RNA transcription were designed separately. The structures of the transcription templates were: (1) T7 promoter + mature Cas12N.1 direct repeat sequence (SEQ ID NO:13) + guide sequence (SEQ ID NO:29); (2) T7 promoter + mature Cas12N.2 direct repeat sequence (SEQ ID NO:14) + guide sequence (SEQ ID NO:29); (3) T7 promoter + mature Cas12N.3 direct repeat sequence (SEQ ID NO:15) + guide sequence (SEQ ID NO:29). Primers were designed using Primer 5.0 software to ensure that the forward primer and the reward primer had at least 18 bp of overlapping sequence.

[0308] 2. Prepare the following reaction system, gently mix by pipetting, briefly centrifuge, and then slowly anneal in a PCR instrument. The PCR system is as follows:

[0309]

[0310] 3. Purify the template using the MinElute PCR Purification Kit, following these steps:

[0311] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2ml collection tube, let stand at room temperature for 2 minutes, 12000g / 2min;

[0312] 2) Discard the waste liquid, add 750 μl of Buffer PE (remember to add ethanol before use), 12000 g / 2 min;

[0313] 3) Discard the waste liquid, add 350 μl of Buffer PE, incubate at 12000 g for 2 min, discard the waste liquid, incubate at 12000 g for 2 min;

[0314] 4) Transfer the MinElute column to a new 1.5ml centrifuge tube, open the cap, and incubate at 65℃ for 2 minutes;

[0315] 5) Add 20 μl of preheated EB solution, let stand for 2 min, then centrifuge at 12000 g / 2 min. To improve the recovery rate, the contents of the centrifuge tube can be passed through a MinElute centrifuge column 2-3 times.

[0316] 6) Determine the concentration using Nanodrop and store at -20℃ for later use.

[0317] 4. Purification of guide RNA: Extraction with phenol:chloroform:isoamyl alcohol (25:24:1) to remove DNAseI from the system;

[0318] 1) Add 80 μl of RNA-free H2O to the post-transcriptional reaction system and adjust the volume to 100 μl;

[0319] 2) Take 2 ml of Phase Lock Gel (PLG) Heavy, centrifuge at 15000g for 2 min, add 100 μl of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μl of RNA digested with DNAseI, gently tap the Phase-Lock tube 5-10 times to mix it evenly, and then centrifuge at 15℃ / 16000g for 12 min;

[0320] 3) Take a new RNA-free 1.5ml centrifuge tube, aspirate the supernatant from the previous centrifugation into the centrifuge tube, being careful not to aspirate the gel, add an equal volume of isopropanol and one-tenth volume of sodium acetate solution, mix well with a pipette tip, and place in a -20℃ refrigerator for 1 hour or overnight.

[0321] 4) Centrifuge at 4℃ / 16000g for 30 min, discard the supernatant, add 75% pre-cooled ethanol, mix the precipitate by suction and beating, centrifuge at 4℃ / 16000g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the surface of the RNA, add 100μl of RNAfree H2O, and mix by suction and beating.

[0322] 5. Measure the concentration of purified crRNA using Nanodrop, and uniformly dilute it to 250 ng / μl. Aliquot the diluted 200 μl PCR centrifuge tubes and store at -80℃ for later use.

[0323] 6. Establishment of a double-stranded DNA restriction enzyme digestion system:

[0324] (1) Prepare the following reaction system, gently mix by pipetting, and briefly centrifuge. Incubate at 37°C for 15 min; the DNA cleavage reaction system is shown below:

[0325]

[0326] (2) Add 300 ng of substrate DNA (100 ng / μl), 3 μL, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 hours;

[0327] (3) Add RNase, place at 37°C for 15 min to fully digest RNA impurities in the system;

[0328] (4) Add proteinase K and place at 58°C for 15 min to digest Cas12N.1 and Cas12N.2 proteins respectively;

[0329] (5) Agarose gel running test.

[0330] Gel running results showed that Cas12N.1, Cas12N.2, and Cas12N.3 can effectively cleave double-stranded DNA.

[0331] Example 5. Cleavage of Cas12N in human cell lines

[0332] A eukaryotic expression vector containing the Cas12N.1 gene and a PCR product containing the U6 promoter and guide RNA (containing the prototype direct repeat sequence shown in SEQ ID NO:7 and the eukaryotic editing guide sequence shown in SEQ ID NO:28) were transfected into human HEK293T cells via liposome transfection and cultured at 37°C and 5% CO2 for 72 h. DNA was extracted from all cells, and a 700 bp sequence containing the target site was amplified. The PCR product was ligated into a B-simple vector for first-generation sequencing, performed by Thermo Fisher Scientific. The sequencing results were aligned to the VEGFA gene in the human genome, identifying the cleavage mechanism of Cas12N.1 at the target site and its editing efficiency for VEGFA. Next-generation sequencing libraries were constructed using Tn5 sequencing, and sequencing was performed by Beijing Annoroad Gene Technology Co., Ltd., confirming the editing efficiency of Cas12N for VEGFA. Simultaneously, the cleavage of the DNMT1 gene by Cas12N was also identified.

[0333] The same method was used to detect the cleavage activity of Cas12N.2 and Cas12N.3 on VEGFA. The prototypical repetitive sequence of Cas12N.2 is shown in SEQ ID NO:8, and the eukaryotic editing guide sequence is shown in SEQ ID NO:28. The prototypical repetitive sequence of Cas12N.3 is shown in SEQ ID NO:9, and the eukaryotic editing guide sequence is shown in SEQ ID NO:28. Sequencing identified the cleavage mode of Cas12N.2 and Cas12N.3 on the target site, as well as the editing efficiency of Cas12N.2 and Cas12N.3 on VEGFA and the cleavage of the DNMT1 gene by Cas12N.2 and Cas12N.3.

[0334] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof. SEQUENCE LISTING <110> China Agricultural University <120> Novel CRISPR-Cas12N enzymes and systems <130> IDC200165 <160> 29 <170> PatentIn version 3.5 <210> 1 <211> 941 <212> PRT <213> artificiality <400> 1 Met Lys Ala Lys Tyr Arg Ser Asn Ser Ile His Ala Ser Leu Arg Gln 1 5 10 15 Leu Leu Glu Leu Gly Leu Ser Lys Ser Ser Ser Ala Glu Pro Gln Lys 20 25 30 Ile Thr Arg Thr Ala Lys Phe Lys Ile Asn Thr Asp Ile Arg Pro Asp 35 40 45 Leu Val Pro Val Leu Asn Arg His Phe Asp Ala Phe Glu Lys Phe Arg 50 55 60 Arg Lys Val Leu Asp Glu Leu Glu Ala Arg Trp Asn Lys Asp Gln Lys 65 70 75 80 Ser Phe Gln Ala Met Val Gln Cys Ser Ala Lys Glu Pro Tyr Gln Lys 85 90 95 Lys Ser Ser Cys Tyr Ala Trp Leu Asp Thr His Phe Ile Thr Glu Ala 100 105 110 Lys Ala Val Leu Asp Leu Pro Arg Lys Pro Ala Thr Ser Leu Leu Tyr 115 120 125 Asn Leu Ser Gly Gly Leu Lys Gly Phe Leu Thr Arg Arg Glu Thr Val 130 135 140 Thr Glu Asp Ile Gln Lys Arg Phe Asn Asp Asn Leu Arg Glu Trp Asn 145 150 155 160 Gly Asp Leu Ser Gln Leu Ala Ser Asp Leu Lys Ala Pro Leu Pro Pro 165 170 175 Ala Pro Pro Asn Leu Asp Phe Glu Asn Leu Thr Glu Arg Ala Ile Glu 180 185 190 Arg Tyr Asn Asp Trp Val Gly Arg Thr Arg Ala Trp Cys Asn Leu Ile 195 200 205 Leu Val Gln Gln Lys Lys Val Glu Arg Arg Asp Ala Cys Leu Pro Arg 210 215 220 Tyr Leu Lys Gly Tyr Pro Gly Phe Phe Gly Ser Gln Arg Tyr Ala Thr 225 230 235 240 Thr Ala Ser Leu Ala Glu Asn Leu Lys Lys Leu Glu Gln Glu Ala Arg 245 250 255 Glu Gln Ser Lys Lys Ala Pro Thr Arg Phe Ala Lys Leu Ser Pro Glu 260 265 270 Ile Trp Thr Ala Ile Gln Glu Arg Phe Ser Pro Pro Glu Gly His Glu 275 280 285 Ala Gly Glu Lys Arg Arg Pro Arg Thr Ala His Gln Thr Val Cys Leu 290 295 300 Arg Phe Ala Ala Leu Arg Ala Ala His Pro Glu Trp Thr Pro Ala Gln 305 310 315 320 Leu Ala Gly Glu Ile Leu Ala Gly Ile Phe His Gly Thr Glu Lys Leu 325 330 335 Lys Lys His Leu Ala Ala Asn Gly Phe Glu Asp Arg Arg Ala Val Ile 340 345 350 Lys Leu Ala Asn Leu Tyr Asn Val Val Val Ala Phe Ser Leu Asp Pro 355 360 365 Ile Arg Ala Ala Gly Asn Tyr Ile Ser Phe Tyr Lys Glu Glu Thr Pro 370 375 380 Lys Arg Asn Ala Phe Gly Asp Val Arg Gly Gly Leu His Gln Pro Ser 385 390 395 400 Asp Glu Ser Ala Ala Ile Glu Ile Met Gly Phe Gly Leu Gln Lys Glu 405 410 415 Ser Gly Lys Pro Leu Tyr Asn Gly Leu Leu Val Cys Lys Lys Ser Glu 420 425 430 Lys Glu His Asp Asp Ser Trp Ala Phe Leu Tyr Cys His Thr Glu Gly 435 440 445 Gln Met Phe Glu Leu Ala Asn Glu Lys Ala Lys Leu Arg Gly Lys Leu 450 455 460 Leu Thr Asp Trp Thr Gly Phe Ala Ser Arg Gly Gly Ser Arg Lys Lys 465 470 475 480 Ala Glu Ala Ser Ala Lys Gln Leu Val Arg Gly Arg Val Trp Ile Ser 485 490 495 Glu Lys Thr Pro Pro Thr Val Met Pro Leu Ala Phe Gly Ser Arg Gln 500 505 510 Gly Arg Glu Tyr Leu Trp His Phe Asp Arg Asp Leu Arg Glu Lys Asn 515 520 525 Glu Trp Val Leu Gly Asn Gly Arg Leu Leu Arg Ile Met Pro Pro Gly 530 535 540 Arg Pro Asn Ala Ala Asp Phe Tyr Leu Thr Ile Thr Leu Glu Arg Gln 545 550 555 560 Ala Pro Pro Val Ala Glu Phe Ile Ala Ser Lys Leu Ile Gly Ile Asp 565 570 575 Arg Gly Glu Ala Val Pro Ala Ala Tyr Ala Leu Ile Asp Arg Glu Gly 580 585 590 Arg Trp Leu Val Asn Gln Lys Arg Phe Gln Glu Phe Leu Lys Ala Leu 595 600 605 Gly Thr Trp His Asp Glu Leu Asp Glu Trp Trp Lys Lys Asn Lys Asn 610 615 620 Thr Pro Ile Asn Asp Arg Val Lys Arg Pro Arg Leu Gln Trp Thr Ser 625 630 635 640 Glu Leu Glu Asp Gly Phe Gly Phe Val Ala Ser Glu Tyr Arg Glu Gln 645 650 655 Gln Ala Asp Phe Asn Asp Gln Lys Arg Glu Leu Gln Arg Thr Gln Gly 660 665 670 Gly Tyr Thr Arg Trp Leu Arg Ser Lys Glu Arg Asn Arg Ala Arg Ala 675 680 685 Leu Gly Gly Glu Val Thr Arg Ala Val Leu Ser Leu Ala Ser Glu His 690 695 700 Arg Ser Pro Leu Val Leu Glu Lys Leu Gly Ser Ala Leu Ala Thr Arg 705 710 715 720 Gly Gly Lys Gly Thr Met Met Ser Gln Met Gln Tyr Glu Arg Met Leu 725 730 735 Thr Thr Leu Glu Gln Lys Leu Ala Glu Val Gly Leu Tyr Ala Met Pro 740 745 750 Ser Ala Pro Lys Tyr Arg Lys Leu Thr Asn Gly Phe Ile Asn Phe Ala 755 760 765 Ala Pro His Tyr Thr Ser Ser Thr Cys Ser Ser Cys Gly Gln Val His 770 775 780 Asp Ser Ile Phe Tyr Asn Glu Leu Ala Lys Lys Ile Val Arg Pro Asn 785 790 795 800 Asp Ser Lys Trp Gln Val Thr Leu Pro Ser Gly Gln Thr Arg Thr Leu 805 810 815 Pro Gln Thr Tyr Ser Tyr Trp Leu Lys Gly Lys Gly Glu Gln Thr Lys 820 825 830 Gln Thr Asn Glu Arg Leu Asn Glu Leu Leu Gly Glu Arg Thr Val Thr 835 840 845 Gln Leu Ser Lys Ser Ser Tyr Met Thr Leu Val Ser Leu Leu Lys Asn 850 855 860 Ser Trp Leu Pro Tyr Arg Pro Gln Gln Ala Asp Phe Cys Cys Leu Asn 865 870 875 880 Cys Gly Tyr Thr Thr Asn Ala Asp Val Gln Gly Ala Leu Asn Ile Ala 885 890 895 Arg Lys Phe Leu Phe Arg Ser Glu Arg Gly Lys Lys Thr Asp Asp Ala 900 905 910 Gly Glu Glu Gly Glu Ala Lys Arg Gly Lys Tyr Lys Asp Asp Trp Gln 915 920 925 Thr Trp Tyr Arg Gln Lys Leu Glu Lys Val Trp Cys Lys 930 935 940 <210> 2 <211> 731 <212> PRT <213> artificia <400> 2 Met Ile Thr Lys Leu Thr Tyr Asn Asp Asn Cys Val Phe Val His Pro 1 5 10 15 Asp Asp Lys Ser Leu Lys Thr Val Lys Ala Phe Thr His Gly Ser Ile 20 25 30 Val Phe Asp Val Asn Ala Arg Lys Val Thr Phe Ala Lys Leu Pro Ala 35 40 45 Val Asp Val Pro Glu Gly Leu Lys Ile Glu Glu Gly Ala Thr Leu Asn 50 55 60 Phe Leu Leu Met Ala Pro Asp Gly Thr Leu Arg Lys Gln Ile Arg Asn 65 70 75 80 Gly Val Leu Gln Ala Asn Glu Lys Gln Pro Gly Phe Leu Arg Ile Gly 85 90 95 Gly Met Ala Gln Ser Pro Arg Lys Phe Ser Arg Glu Asp Gly Trp Gly 100 105 110 Asn Met Ile Tyr Lys Tyr Arg Ala Tyr Phe Thr His Pro Gly Leu Thr 115 120 125 Thr Asp Thr Glu Leu Pro Val Trp Leu Lys Asp Ser Ile Lys Arg Gln 130 135 140 Lys Glu Tyr Trp Asn Arg Leu Ala Trp Leu Cys Arg Glu Ala Arg Arg 145 150 155 160 Lys Cys Ser Pro Val Lys Ser Glu Glu Ile Ala Ala Phe Val Lys Thr 165 170 175 Glu Ile Leu Pro Ala Ile Asp Thr Leu Asn Asp Ser Leu Gly Arg Ser 180 185 190 Lys Asp Lys Met Lys His Pro Ala Lys Leu Lys Ile Glu Asp Pro Gly 195 200 205 Val Asp Gly Leu Trp Lys Phe Val Gly Glu Leu Arg His Arg Val Glu 210 215 220 Lys Asp Arg Pro Val Pro Pro Gly Leu Leu Glu Lys Val Ile Ala Phe 225 230 235 240 Ala Glu Gln Phe Lys Ala Asp Tyr Thr Ala Leu Asn Asp Phe Met Asn 245 250 255 Asn Leu Pro Ser Ile Ala Glu Arg Glu Ala Glu Thr Leu His Leu Arg 260 265 270 Arg Phe Glu Val Arg Pro Thr Leu Ala Ser Phe Arg Ala Thr Leu Asp 275 280 285 Arg Arg Lys Thr Thr Lys Ala Ala Trp Ser Glu Gly Trp Pro Leu Ile 290 295 300 Lys Tyr Pro Asp Ser Pro Lys Ala Glu Asp Trp Gly Leu His Tyr Tyr 305 310 315 320 Phe Asn Lys Ala Gly Ile Gly Ser Asp Leu Leu Glu Glu Gly Asp Gly 325 330 335 Val Pro Gly Leu Ser Phe Gly Pro Ala Leu Ser Pro Ala Lys Thr Gly 340 345 350 His Glu Asn Leu Val Gly Gly Ala Thr Lys Arg Arg Leu Arg Glu Ala 355 360 365 Glu Ile Ser Ile Ser Gly Ser Asn Lys Glu Arg Trp Asp Phe His Phe 370 375 380 Gly Val Leu Glu His Arg Pro Leu Pro Lys His Ser His Leu Lys Glu 385 390 395 400 Trp Lys Leu Leu Phe Gln Asp Gly Ala Leu Trp Leu Cys Leu Val Val 405 410 415 Glu Leu Gln Gln Pro Leu Pro Glu Pro Ser Thr Leu Ala Ala Gly Leu 420 425 430 Asp Ile Gly Trp Arg Arg Thr Glu Glu Gly Ile Arg Phe Gly Thr Leu 435 440 445 Tyr Glu Pro Ala Thr Lys Thr Ile His Glu Leu Ile Ile Asn Phe Gln 450 455 460 Arg Ser Pro Lys Asp Pro Lys Gln Arg Val Pro Phe Arg Ile Asp Leu 465 470 475 480 Gly Pro Thr Arg Trp Glu Lys Arg Asn Ile Ile Lys Leu Leu Pro Lys 485 490 495 Trp Lys Pro Gly Asp Gly Ile Pro Asn Ser Leu Glu Ile Lys Ala Ala 500 505 510 Leu Gln Thr Arg Arg Asp Tyr Phe Lys Asp Thr Ala Lys Ile Leu Leu 515 520 525 Arg Lys His Leu Gly Glu Lys Thr Pro Ala Trp Leu Asp Lys Ala Gly 530 535 540 Arg Ser Gly Leu Leu His Leu Lys Glu Glu Phe Lys Asp Asp Ala Asp 545 550 555 560 Val Gln Glu Ile Leu Asn Thr Trp Ala Thr Asn Asp Glu Gln Val Gly 565 570 575 Lys Leu Ala Ser Lys Tyr Ser Ala Arg Val Thr Arg Arg Val Glu Tyr 580 585 590 Gly Gln Gln Gln Val Ala His Asp Val Cys Arg Tyr Leu Gln Gln Lys 595 600 605 Asp Ile Ala Arg Leu Val Val Glu Lys Asn Phe Leu Ala Lys Ile Ala 610 615 620 Gln His Gln Asp Asn Ser Asp Pro Glu Ser Leu Lys Arg Ser Gln Lys 625 630 635 640 Tyr Arg Gln Phe Ala Ala Val Gly Arg Phe Val Ser Lys Leu Lys Asp 645 650 655 Thr Val Val Lys Tyr Gly Gln Val Val Asp Ser Asp Glu Ala Thr Asn 660 665 670 Thr Thr Arg Met Cys Gln Tyr Cys Asn His Leu Asn Pro Ala Thr Glu 675 680 685 Lys Glu Lys Met Asp Cys Glu Gly Cys Gly Arg Glu Ile Lys Leu Asp 690 695 700 Asp Asn Ala Ala Ile Asn Leu Ser Arg Phe Ala Ser Asp Pro Glu Leu 705 710 715 720 Ala Glu Leu Ala Arg Ala Lys Lys Asn Lys Glu 725 730 <210> 3 <211> 760 <212> PRT <213> artificia <400> 3 Met Ile Lys Glu Leu Thr Tyr Lys Asp Asp Cys Val Phe Val His Ala 1 5 10 15 Ala Glu Thr Ser Leu Lys Thr Thr Lys Ala Phe Thr His Gly Arg Ile 20 25 30 Val Phe Asp Glu Ser Thr Arg Leu Ala Thr Phe Ala Gln Leu Pro Ala 35 40 45 Val Ala Ile Pro Ala Glu Leu His Leu Lys Ala Glu Thr Pro Met Asn 50 55 60 Phe Gln Leu Val Ala Pro Asp Gly Glu Leu Arg Pro Gln Lys Arg Ser 65 70 75 80 Gly Val Leu Leu Val Asp Glu Lys Gln Pro Thr Ile Phe His Ile Gly 85 90 95 Gly Met Ser Lys Thr Pro Arg Gln Phe Ser Ala Glu Asp Gly Trp Thr 100 105 110 Asn Asn Ile Tyr Lys Tyr Arg Ala Tyr Phe Thr His Glu Gly Thr Asn 115 120 125 Thr Ser Gly Asp Val Pro Glu Trp Leu Lys Ala Ser Ile Thr Arg Gln 130 135 140 Arg Asp Leu Trp Asn Arg Leu Ala Trp Leu Cys Arg Glu Ala Arg Arg 145 150 155 160 Gln Cys Ser Ser Gly Ser Pro Glu Glu Ile Thr Ala Phe Val Glu Thr 165 170 175 Val Ile Leu Pro Ala Ile Asp Ala Phe Asn Asp Ser Leu Gly Arg Ser 180 185 190 Lys Asp Lys Met Lys His Pro Ala Lys Leu Lys Val Glu Met Pro Gly 195 200 205 Leu Asp Gly Leu Trp His Phe Val Gly Asp Leu Arg Lys Arg Val Asp 210 215 220 Lys Asp Arg Pro Val Pro Glu Gly Leu Leu Glu Lys Val Ile Asp Phe 225 230 235 240 Ala Gln Gln His Lys Thr Asp Tyr Thr Pro Leu Asp Val Phe Val Lys 245 250 255 Asp Phe Gly Ala Ile Ala Asn Lys Glu Ala Lys Asp Leu Lys Leu Arg 260 265 270 Ser Phe Glu Ile Arg Pro Thr Val Thr Ala Phe Lys Ala Thr Leu Asp 275 280 285 Arg Arg Arg Thr Lys Lys Leu Thr Trp Ser Glu Gly Trp Pro Leu Ile 290 295 300 Lys Tyr Asn Asp Ser Pro Lys Ala Gly Asp Trp Gly Leu His Tyr Tyr 305 310 315 320 Phe Asn Lys Ala Gly Val Asn Ser Ser Leu Leu Glu Ser Lys Glu Gly 325 330 335 Val Pro Gly Leu Ser Phe Gly Ala Pro Leu Asp Pro Ala Asn Thr Gly 340 345 350 His Lys Leu Ile His Thr Ala Ser Gly Gly Lys Thr Leu Arg Glu Ala 355 360 365 Glu Ile Ser Ile Arg Gly Asp Asp Asn Gln Gln Trp Arg Phe Arg Phe 370 375 380 Ala Val Ile Gln Trp Arg Pro Leu Pro Pro Asn Ser His Val Lys Glu 385 390 395 400 Trp Lys Leu Leu Leu Gln Asp Gly Lys Leu Trp Leu Cys Leu Val Val 405 410 415 Glu Leu Gln Arg Asp Leu Pro Thr Ser Thr Ala Leu Ala Ala Gly Leu 420 425 430 Asp Val Gly Trp Arg Arg Thr Glu Val Gly Ile Arg Phe Gly Thr Leu 435 440 445 Tyr Glu Pro Thr Ser Trp Thr Val Arg Glu Leu Val Met Asp Phe Glu 450 455 460 Gln Ser Pro Lys Asp His Lys Asp Arg Val Pro Phe Arg Phe Asp Met 465 470 475 480 Gly Pro Thr Arg Trp Glu Lys Arg Asn Ile Thr Arg Leu Phe Pro Asp 485 490 495 Trp Lys Pro Gly Asp Gly Ile Pro Asn Ala Leu Glu Ile Arg Thr Ala 500 505 510 Leu Gln Thr Arg Arg Asp Tyr Phe Lys Asp Thr Ala Lys Ile Leu Leu 515 520 525 Arg Lys His Leu Gly Glu Lys Thr Pro Ala Trp Leu Glu Lys Ala Gly 530 535 540 Arg Ser Gly Leu Leu Arg Leu Lys Glu Glu Phe Lys Asp Asp Val Asp 545 550 555 560 Val Gln Gly Ile Leu Asn Thr Trp Gly Lys Asn Glu Glu Glu Ile Gly 565 570 575 Thr Leu Ala Ala Ala Tyr Ser Lys Lys Val Thr Thr Arg Ile Glu Tyr 580 585 590 Asn Gln Leu Gln Ile Ala His Asp Val Cys Arg Tyr Leu Lys Gln Lys 595 600 605 Gly Ile Thr Arg Leu Ile Leu Glu Lys Asn Phe Leu Ala Lys Ile Ala 610 615 620 Leu His Gln Glu Asn Thr Asp Pro Val Ser Leu Lys Arg Ser Gln Lys 625 630 635 640 Tyr Arg Gln Phe Ala Ala Val Gly Arg Phe Ile Ser Leu Leu Lys Tyr 645 650 655 Thr Ala Val Lys Tyr Gly Ile Val Thr Glu Pro Tyr Glu Ala Met Asn 660 665 670 Thr Thr Arg Met Cys Gln Tyr Cys Asn His Leu Asn Pro Ser Thr Glu 675 680 685 Lys Asp Arg Phe Gln Cys Glu Ala Cys Gly Arg Thr Ile Asp Gln Asp 690 695 700 Tyr Asn Ala Ala Val Asn Leu Ser Arg Phe Ala Cys Asp Pro Lys Leu 705 710 715 720 Ala Lys Leu Ala Leu Gly Asn Gly Asn Glu Ser Glu Glu Asp Ala Gln 725 730 735 Thr Pro Glu Phe Glu Asp Glu Glu Glu Asp Thr Gln Ser Leu Glu Ser 740 745 750 Asp Glu Ser Gly Thr Asn Asp Leu 755 760 <210> 4 <211> 2826 <212> DNA <213> artificia <400> 4 atgaaagcaa agtatcggtc taactcaatt catgcctcgc tccggcagct tttggagctt 60 ggcttaagca agagcagttc ggcagaacca cagaaaatca cgcgcactgc caagttcaag attackccg acatccggcc ggatttggtt ccggtcttga accggcattt tgatgcgttt 180 gaaaagttcc ggcgcaaggt gctggatgag cttgaagcac gatggaacaa agaccagaaa tcgtttcaag caatggtgca atgcagcgcg aaggaaccgt atcagaagaa atccagttgt 360. tatgcttggc tcgacactca tttcatcacc gaagcaaaag cggttcttga cttaccgcgc aaaccagcca ccagccttct ttacaatttg agtggcggat tgaaaggttt tctgacacgc 420 cgcgagacag tcaccgagga cattcaaaaa cgcttcaatg acaatctccg agaatggaac ggcgatttaa gtcagttggc cagcgacctc aaagctccct tgccgccagc acccccgaat 540 ttggattttg aaaacctgac cgaaagagca atcgaaaggt ataacgattg ggttgggcgg acgcgcgcat ggtgcaatct cattcttgtc cagcaaaaga aagttgaacg ccgcgatgcc tgtttgccgc gctatcttaa aggatatccg ggctttttcg gctcacaacg ctatgccaca 720. acagcgagtc tagccgaga cttgaaaaaa ttggacagg aggcacgaga gcaaagcaaa aaagcgccaa ctcgctttgc caaactctca cctgaaattt ggacagccat tcaagaacga 840 ttctcacctc ccgaagggca tgaagccgga gaaaagcgcc gtccgcgcac ggctcaccag 900 actgtctgcc tccgctttgc tgctttgcgg gcagcgcatc cagagtggac tccggcgcaa 960 ctggccggag aaattttagc tggtattttc catggaacgg aaaaacttaa aaagcatctt 1020 gccgctaacg gctttgaaga tcgccgggcg gtcatcaagt tggcaaatct ttacaatgtt 1080 gtggttgcct tctcgctcga cccaatccgc gccgccggaa attacatttc gttttacaag 1140 gaggagacgc ccaagagaaa cgccttcggc gacgtgcggg gtggactgca tcagccgagt 1200 gatgaatcgg ccgcaattga aatcatgggt tttggtttgc agaaagaatc gggcaaaccc 1260 ctctacaacg gcttgctcgt ctgcaaaaaa tcagaaaagg aacatgacga ttcgtgggcg 1320 tttctttatt gccacaccga aggacaaatg tttgaacttg cgaatgaaaa ggcgaagctt 1380 cgcggcaagt tgctaactga ctggactggc tttgccagtc gtggcggctc gcggaaaaag 1440 gctgaagcaa gtgcgaagca attggtgcga ggacgtgtct ggattagcga aaaaactccg 1500 ccgacagtca tgccgctggc gttcgggtca cgtcagggac gtgaatatct ctggcatttt 1560 gaccgggatt tgcgtgagaa aaacgaatgg gttttgggca atggcagact gctccgcatc 1620 atgccgcccg gacgaccgaa cgccgccgac tttatctga caatcactct cgaacgacaa 1680 gcgccacctg ttgctgaatt tatagcgagc aaattaatcg ggattgatcg cggcgaagcc 1740 gttcccgcag cttatgcgct gattgaccgt gaaggaagat ggttggtcaa tcaaaagcga 1800 tttcaagaat tcctgaaagc gctgggcact tggcatgatg agttggatga gtggtggaag 1860 aaaaacaaga acactccaat aaacgaccgg gtaaaacggc cccggttgca atggacaagt 1920 gaactggagg atggatttgg atttgttgcg agcgaatacc gcgaacaaca agccgatttt 1980 aatgaccaaa agcgtgaatt gcagcgaacc caaggcggtt acacgcgctg gttacgaagc 2040 aaagagcgca atcgcgctcg tgctttgggt ggtgaagtca cacgcgcggt gctgtcgctt 2100 gcctccgaac atcgttcgcc tctggttttg gaaaaacttg ggagcgcgtt ggcgacgcgc 2160 ggcggcaaag gaacaatgat gagtcaaatg footgaac gaatgctcac cacgcttgaa 2220 caaaagctcg cagaggtcgg actttatgcg atgccttcag caccaaaata tcgcaagctc 2280 accaatggct tcattaattt cgccgcacca cactatacgt cctcgacatg ctcttcgtgc 2340 gggcaagtgc atgacagcat tttctacaat gagcttgcga aaaaaatcgt gcgcccaaac 2400 gattcaaagt ggcaagtgac gcttccaagt gggcagacac gaacattgcc tcaaacttac 2460 agctattggt tgaaaggaaa aggtgagcag accaaacaaa cgaacgagcg tctcaacgaa 2520 ctgctgggtg aacgaacagt tactcaattg agcaagagca gctatatgac attggtaagt 2580 ttgctcaaaa atagttggtt gccctaccgc ccacaacaag ctgatttttg ttgtctgaat 2640 tgtggctaca ccacgaatgc cgacgtgcaa ggtgcgttaa acattgcccg aaagtttttg 2700 ttccgttccg aacgaggcaa aaagactgac gacgcaggcg aagaaggtga ggcaaagcgt 2760 ggtaaataca aagacgactg gcagacttgg tatcgccaaa agttggaaaa ggtctggtgc 2820 aagtaa 2826 <210> 5 <211> 2196 <212> DNA <213> artifice <400> 5 60 ctaaaacgg ttaaagcctt tacgcatggc agcatcgtat ttgacgtgaa tgcgcgcaaa 120 gtgacctttg caaagttgcc tgcggttgac gttcctgaag ggctcaagat cgaagagggc 180 gcgaccctga actttttgtt gatggcgccc gatggaacct tgcgtaagca gatccgcaac 240 ggcgtgttgc aggccaacga aaaacagcca ggctttctgc gcattggcgg tatggcacaa 300 agcccacgca agttttcgcg cgaagatggc tggggcaaca tgatctacaa gtaccgcgcc 360 tatttcacgc atccggggtt gacgacagac acagaactgc cggtatggct taggactca 420 480 aaatgttctc cagtcaagag cgaagaaatc gcggcctttg taaaaaccga gattctgccg 540 gcgattgaca ctctgaatga ttcgctcggg cgctccaaag aagaatgaa gcaccctgca 600 660 caccgtgttg aaaaggaccg tccggttcca ccggggctac tggaaaaggt gattgccttt 720 gccgagcaat ttaaagcaga ctacaccgcg ctcaacgatt tcatgaacaa tttgccgtcg 780 attgcggaaa gggaagcaga aactcttcat ctgcgccgtt ttgaagtcag gccgacgctt 840 gcttctttca gggcgacgct cgatcgtcgc aagaccacca aagcagcctg gtcggaaggc 900 tggccgctga tcaaatatcc tgacagtccc aaggcagaag actggggtct gcactactac 960 ttcaacaagg cgggtattgg ctctgatctt ctagaggaag gtgatggcgt cccagggttg 1020 agctttggtc ctgcgctgtc gcccgcaaa acaggacatg aaaatctggt tggcggggcg 1080 acaaagcggc gactacggga agcggagatt tctatttctg gcagcaacaa agaacggtgg 1140 gattttcatt ttggagtgtt ggaacaccgg cctctgccga agcactcgca tctgaaggag 1200 tggaaattgc tttttcaaga tggagcactg tggctttgcc tagttgtgga gctacaacag 1260 ccattgccgg agcccagcac gcttgcggcg gggcttgata ttggctggcg cagaacagaa 1320 gaaggcattc ggttcggaac gctgtatgag cccgcaacca agacaatcca cgaactgatc 1380 atcaattttc agcggtcgcc caaagatccc aagcaacgcg tgcctttccg gattgatctt 1440 ggccccacgc gttgggaaaa acgcaacatt atcaagctgc ttcccaaatg gaagccagga 1500 gacggaattc caaattcttt ggaaatcaaa gccgcgctgc agacgcgtcg tgattacttc 1560 aaggacactg caaagatttt gttgcgcaag catctaggag aaaaaacgcc ggcatggctt 1620 gacaaagctg gccgtagtgg cttgctgcat ctgaaggaag aattcaaaga cgatgccgat 1680 gtgcaagaga tcttgaacac gtgggctaca aatgatgagc aggtcggaaa actggcgtca 1740 aaatactccg cacgagttac tagacgcgtt gaatatggcc aacaacaggt tgcccacgat 1800 gtttgccgtt accttcagca aaaggacatt gctcgtttgg tcgtggagaa gaacttcttg 1860 gccaagattg cccagcatca agacaatagc gatccagaaa gtttgaagcg ttcgcagaaa 1920 taccggcaat ttgcggcagt gggtcgtttt gtttcaaagc tgaaggacac tgtcgttaaa 1980 tacggccaag tggttgactc ggatgaggca acaaatacaa cgcgtatgtg ccagtattgc 2040 aaccatctga atccggcaac cgaaaaagaa aagatggatt gcgagggatg cggaagagaa 2100 atcaagctgg acgacaatgc ggcaataaat ctctcgcgct ttgcttccga tccagagctg 2160 gcagaactgg cgcgagcaaa gaaaaataag gaataa 2196 <210> 6 <211> 2283 <212> DNA <213> artifice <400> 6 atgatcaaag aattgaccta caaggacgat tgcgtctttg tgcatgccgc agaaacgtcc 60 ctgaaaacca ccaaggcgtt tacacatggc cgcattgtgt ttgacgagag cacgcgcttg 120 gcgacgtttg cccagctgcc ggcagtggcg attcccgccg agcttcatct caaggcagag 180 acaccaatga atttcagct ggtggcccct gacggcgagc tgcgtccgca gaaacgatcg 240 ggagtgctcc tggtagacga aaagcaacca accatcttcc acattggcgg aatgtcaaaa 300 acgccgcggc agttttcagc agaagatgga tggacgaaca acatctacaa ataccgcgcg 360 tatttcaccc acgaaggcac aaacacgagc ggcgacgtac ctgaatggct caaagcatcg 420 atcacgcggc aaagggatct ttggaaccgg ttggcgtggc tttgccgcga agcacggcgc 480 cagtgctctt ccggttcgcc cgaggagatt accgcttttg tcgaaaccgt tattttgcca 540 gcgatcgatg cttttaacga ctcgcttggc cgttccaaag acaagatgaa gcacccggcc 600 aagctgaagg tcgagatgcc cggcttggat ggcttgtggc attttgttgg cgatctgcgg 660 aagcgcgtcg acaaagatcg tcccgtgcca gaaggcttgc tggagaaggt gattgacttt 720 gcccagcagc acaagacgga ctatacgccg ctcgacgtgt ttgtaaagga ttttggggcg 780 attgccaaca aggaagcaaa ggatctcaag ctacgaagct ttgaaattcg gccgacagtc 840 accgctttca aggcaaccct ggatcgccgc agaaccaaga agttgacttg gtccgagggc 900 tggccactca tcaagtacaa cgacagcccc aaggccggcg attggggcct gcattactac 960 ttcaacaagg ctggcgtcaa ttcatcactg ctggaatcga aggaaggcgt accaggattg 1020 tcgtttggag cgccgctcga tcctgccaat acggggcata agctaatcca tacggcaagc 1080 ggtggcaaga cgcttcggga agcggagatt tccattcgtg gagacgataa ccagcagtgg 1140 cgttttcgtt ttgcagtcat tcagtggcgg ccgttgcccc caaattccca cgtaaaagag 1200 tggaaactgc tattgcagga tggcaaactt tggctttgcc tggtggttga gctgcaacgt 1260 gacctaccaa cttccactgc ccttgcagcc ggtttggatg ttggctggcg gcgcacagaa 1320 gtgggaattc ggtttggcac actgtatgag ccaacgagct ggacggtgcg cgaacttgtg 1380 atggactttg aacaatcgcc caaagatcac aaagatcgtg tgcccttccg cttcgatatg 1440 gggcctacac gctgggagaa acgcaatatc actcggctat tccctgattg gaagccgggc 1500 gacggaattc ccaacgcgtt ggagatcaga accgcgttgc agacccgccg tgactatttc 1560 aaagatacgg ccaagattct gctgcgcaag catcttgggg aaaaaacgcc agcatggctg 1620 gaaaaagccg gtcgtagtgg attgctgcgg ctcaaggaag aattcaaaga cgatgtcgat 1680 gtgcaaggca tcctgaatac gtgggggaaa aacgaggaag aaatcggaac tcttgcggcc 1740 gcatacagca aaaaggttac cacgcgcatt gagtacaacc agctgcaaat cgcccacgat 1800 gtttgccgct atctgaagca aaagggcatt acacggttga tcctggaaaa gaacttcctg 1860 gccaaaattg cgctccacca ggagaacacc gatcccgtga gcttgaaacg ctcgcagaaa 1920 tatcggcaat ttgcggccgt tggccgcttt atctctctgc tcaagtacac tgccgtaaaa 1980 tacggaattg tgactgagcc gtatgaggca atgaacacaa cgcgtatgtg ccaatattgc 2040 aaccatctga atccatcgac ggaaaaggat cgcttccagt gcgaagcgtg tggaagaacg 2100 attgaccagg actacaacgc tgcggtaaat ctgtcacgtt ttgcctgcga tccaaagttg 2160 gcaaagctgg cgctgggaaa tggaaacgaa tctgaagaag acgcccagac gccggaattt 2220 gaagacgaag aagaagacac gcagtcgctg gaaagtgacg aaagcggcac aaacgatctt 2280 tga 2283 <210> 7 <211> 25 <212> RNA <213> crafts <400> 7 cgaugcacuc aaacugcauu ccuac 25 <210> 8 <211> 37 <212> RNA <213> crafts <400> 8 guugcaucua cgcgaugcac ucaaaccacg uugcuac 37 <210> 9 <211> 39 <212> RNA <213> crafts <400> 9 aguagcaacg uaguugcagu gcaucggaua gaugcaaca 39 <210> 10 <211> 25 <212> DNA <213> crafts <400> 10 cgatgcactc aaactgcatt gctac 25 <210> 11 <211> 37 <212> DNA <213> crafts <400> 11 gttgcatcta cgcgatgcac tcaaaccacg ttgctac 37 <210> 12 <211> 39 <212> DNA <213> crafts <400> 12 agtagcaacg tagttgcagt gcatcggata gatgcaaca 39 <210> 13 <211> 1 <212> DNA <213> crafts <400> 13 n 1 <210> 14 <211> 1 <212> DNA <213> crafts <400> 14 n 1 <210> 15 <211> 1 <212> DNA <213> crafts <400> 15 n 1 <210> 16 <211> 1 <212> DNA <213> crafts <400> 16 n 1 <210> 17 <211> 1 <212> DNA <213> crafts <400> 17 n 1 <210> 18 <211> 1 <212> DNA <213> artificiality <400> 18 n 1 <210> 19 <211> 11 <212> PRT <213> artificiality <400> 19 Ser Arg Ala Asp Pro Lys Lys Lys Arg Lys Val 1 5 10 <210> 20 <211> 952 <212> PRT <213> artificiality <400> 20 Met Lys Ala Lys Tyr Arg Ser Asn Ser Ile His Ala Ser Leu Arg Gln 1 5 10 15 Leu Leu Glu Leu Gly Leu Ser Lys Ser Ser Ser Ala Glu Pro Gln Lys 20 25 30 Ile Thr Arg Thr Ala Lys Phe Lys Ile Asn Thr Asp Ile Arg Pro Asp 35 40 45 Leu Val Pro Val Leu Asn Arg His Phe Asp Ala Phe Glu Lys Phe Arg 50 55 60 Arg Lys Val Leu Asp Glu Leu Glu Ala Arg Trp Asn Lys Asp Gln Lys 65 70 75 80 Ser Phe Gln Ala Met Val Gln Cys Ser Ala Lys Glu Pro Tyr Gln Lys 85 90 95 Lys Ser Ser Cys Tyr Ala Trp Leu Asp Thr His Phe Ile Thr Glu Ala 100 105 110 Lys Ala Val Leu Asp Leu Pro Arg Lys Pro Ala Thr Ser Leu Leu Tyr 115 120 125 Asn Leu Ser Gly Gly Leu Lys Gly Phe Leu Thr Arg Arg Glu Thr Val 130 135 140 Thr Glu Asp Ile Gln Lys Arg Phe Asn Asp Asn Leu Arg Glu Trp Asn 145 150 155 160 Gly Asp Leu Ser Gln Leu Ala Ser Asp Leu Lys Ala Pro Leu Pro Pro 165 170 175 Ala Pro Pro Asn Leu Asp Phe Glu Asn Leu Thr Glu Arg Ala Ile Glu 180 185 190 Arg Tyr Asn Asp Trp Val Gly Arg Thr Arg Ala Trp Cys Asn Leu Ile 195 200 205 Leu Val Gln Gln Lys Lys Val Glu Arg Arg Asp Ala Cys Leu Pro Arg 210 215 220 Tyr Leu Lys Gly Tyr Pro Gly Phe Phe Gly Ser Gln Arg Tyr Ala Thr 225 230 235 240 Thr Ala Ser Leu Ala Glu Asn Leu Lys Lys Leu Glu Gln Glu Ala Arg 245 250 255 Glu Gln Ser Lys Lys Ala Pro Thr Arg Phe Ala Lys Leu Ser Pro Glu 260 265 270 Ile Trp Thr Ala Ile Gln Glu Arg Phe Ser Pro Pro Glu Gly His Glu 275 280 285 Ala Gly Glu Lys Arg Arg Pro Arg Thr Ala His Gln Thr Val Cys Leu 290 295 300 Arg Phe Ala Ala Leu Arg Ala Ala His Pro Glu Trp Thr Pro Ala Gln 305 310 315 320 Leu Ala Gly Glu Ile Leu Ala Gly Ile Phe His Gly Thr Glu Lys Leu 325 330 335 Lys Lys His Leu Ala Ala Asn Gly Phe Glu Asp Arg Arg Ala Val Ile 340 345 350 Lys Leu Ala Asn Leu Tyr Asn Val Val Val Ala Phe Ser Leu Asp Pro 355 360 365 Ile Arg Ala Ala Gly Asn Tyr Ile Ser Phe Tyr Lys Glu Glu Thr Pro 370 375 380 Lys Arg Asn Ala Phe Gly Asp Val Arg Gly Gly Leu His Gln Pro Ser 385 390 395 400 Asp Glu Ser Ala Ala Ile Glu Ile Met Gly Phe Gly Leu Gln Lys Glu 405 410 415 Ser Gly Lys Pro Leu Tyr Asn Gly Leu Leu Val Cys Lys Lys Ser Glu 420 425 430 Lys Glu His Asp Asp Ser Trp Ala Phe Leu Tyr Cys His Thr Glu Gly 435 440 445 Gln Met Phe Glu Leu Ala Asn Glu Lys Ala Lys Leu Arg Gly Lys Leu 450 455 460 Leu Thr Asp Trp Thr Gly Phe Ala Ser Arg Gly Gly Ser Arg Lys Lys 465 470 475 480 Ala Glu Ala Ser Ala Lys Gln Leu Val Arg Gly Arg Val Trp Ile Ser 485 490 495 Glu Lys Thr Pro Pro Thr Val Met Pro Leu Ala Phe Gly Ser Arg Gln 500 505 510 Gly Arg Glu Tyr Leu Trp His Phe Asp Arg Asp Leu Arg Glu Lys Asn 515 520 525 Glu Trp Val Leu Gly Asn Gly Arg Leu Leu Arg Ile Met Pro Pro Gly 530 535 540 Arg Pro Asn Ala Ala Asp Phe Tyr Leu Thr Ile Thr Leu Glu Arg Gln 545 550 555 560 Ala Pro Pro Val Ala Glu Phe Ile Ala Ser Lys Leu Ile Gly Ile Asp 565 570 575 Arg Gly Glu Ala Val Pro Ala Ala Tyr Ala Leu Ile Asp Arg Glu Gly 580 585 590 Arg Trp Leu Val Asn Gln Lys Arg Phe Gln Glu Phe Leu Lys Ala Leu 595 600 605 Gly Thr Trp His Asp Glu Leu Asp Glu Trp Trp Lys Lys Asn Lys Asn 610 615 620 Thr Pro Ile Asn Asp Arg Val Lys Arg Pro Arg Leu Gln Trp Thr Ser 625 630 635 640 Glu Leu Glu Asp Gly Phe Gly Phe Val Ala Ser Glu Tyr Arg Glu Gln 645 650 655 Gln Ala Asp Phe Asn Asp Gln Lys Arg Glu Leu Gln Arg Thr Gln Gly 660 665 670 Gly Tyr Thr Arg Trp Leu Arg Ser Lys Glu Arg Asn Arg Ala Arg Ala 675 680 685 Leu Gly Gly Glu Val Thr Arg Ala Val Leu Ser Leu Ala Ser Glu His 690 695 700 Arg Ser Pro Leu Val Leu Glu Lys Leu Gly Ser Ala Leu Ala Thr Arg 705 710 715 720 Gly Gly Lys Gly Thr Met Met Ser Gln Met Gln Tyr Glu Arg Met Leu 725 730 735 Thr Thr Leu Glu Gln Lys Leu Ala Glu Val Gly Leu Tyr Ala Met Pro 740 745 750 Ser Ala Pro Lys Tyr Arg Lys Leu Thr Asn Gly Phe Ile Asn Phe Ala 755 760 765 Ala Pro His Tyr Thr Ser Ser Thr Cys Ser Ser Cys Gly Gln Val His 770 775 780 Asp Ser Ile Phe Tyr Asn Glu Leu Ala Lys Lys Ile Val Arg Pro Asn 785 790 795 800 Asp Ser Lys Trp Gln Val Thr Leu Pro Ser Gly Gln Thr Arg Thr Leu 805 810 815 Pro Gln Thr Tyr Ser Tyr Trp Leu Lys Gly Lys Gly Glu Gln Thr Lys 820 825 830 Gln Thr Asn Glu Arg Leu Asn Glu Leu Leu Gly Glu Arg Thr Val Thr 835 840 845 Gln Leu Ser Lys Ser Ser Tyr Met Thr Leu Val Ser Leu Leu Lys Asn 850 855 860 Ser Trp Leu Pro Tyr Arg Pro Gln Gln Ala Asp Phe Cys Cys Leu Asn 865 870 875 880 Cys Gly Tyr Thr Thr Asn Ala Asp Val Gln Gly Ala Leu Asn Ile Ala 885 890 895 Arg Lys Phe Leu Phe Arg Ser Glu Arg Gly Lys Lys Thr Asp Asp Ala 900 905 910 Gly Glu Glu Gly Glu Ala Lys Arg Gly Lys Tyr Lys Asp Asp Trp Gln 915 920 925 Thr Trp Tyr Arg Gln Lys Leu Glu Lys Val Trp Cys Lys Ser Arg Ala 930 935 940 Asp Pro Lys Lys Lys Arg Lys Val 945 950 <210> 21 <211> 742 <212> PRT <213> artificia <400> 21 Met Ile Thr Lys Leu Thr Tyr Asn Asp Asn Cys Val Phe Val His Pro 1 5 10 15 Asp Asp Lys Ser Leu Lys Thr Val Lys Ala Phe Thr His Gly Ser Ile 20 25 30 Val Phe Asp Val Asn Ala Arg Lys Val Thr Phe Ala Lys Leu Pro Ala 35 40 45 Val Asp Val Pro Glu Gly Leu Lys Ile Glu Glu Gly Ala Thr Leu Asn 50 55 60 Phe Leu Leu Met Ala Pro Asp Gly Thr Leu Arg Lys Gln Ile Arg Asn 65 70 75 80 Gly Val Leu Gln Ala Asn Glu Lys Gln Pro Gly Phe Leu Arg Ile Gly 85 90 95 Gly Met Ala Gln Ser Pro Arg Lys Phe Ser Arg Glu Asp Gly Trp Gly 100 105 110 Asn Met Ile Tyr Lys Tyr Arg Ala Tyr Phe Thr His Pro Gly Leu Thr 115 120 125 Thr Asp Thr Glu Leu Pro Val Trp Leu Lys Asp Ser Ile Lys Arg Gln 130 135 140 Lys Glu Tyr Trp Asn Arg Leu Ala Trp Leu Cys Arg Glu Ala Arg Arg 145 150 155 160 Lys Cys Ser Pro Val Lys Ser Glu Glu Ile Ala Ala Phe Val Lys Thr 165 170 175 Glu Ile Leu Pro Ala Ile Asp Thr Leu Asn Asp Ser Leu Gly Arg Ser 180 185 190 Lys Asp Lys Met Lys His Pro Ala Lys Leu Lys Ile Glu Asp Pro Gly 195 200 205 Val Asp Gly Leu Trp Lys Phe Val Gly Glu Leu Arg His Arg Val Glu 210 215 220 Lys Asp Arg Pro Val Pro Pro Gly Leu Leu Glu Lys Val Ile Ala Phe 225 230 235 240 Ala Glu Gln Phe Lys Ala Asp Tyr Thr Ala Leu Asn Asp Phe Met Asn 245 250 255 Asn Leu Pro Ser Ile Ala Glu Arg Glu Ala Glu Thr Leu His Leu Arg 260 265 270 Arg Phe Glu Val Arg Pro Thr Leu Ala Ser Phe Arg Ala Thr Leu Asp 275 280 285 Arg Arg Lys Thr Thr Lys Ala Ala Trp Ser Glu Gly Trp Pro Leu Ile 290 295 300 Lys Tyr Pro Asp Ser Pro Lys Ala Glu Asp Trp Gly Leu His Tyr Tyr 305 310 315 320 Phe Asn Lys Ala Gly Ile Gly Ser Asp Leu Leu Glu Glu Gly Asp Gly 325 330 335 Val Pro Gly Leu Ser Phe Gly Pro Ala Leu Ser Pro Ala Lys Thr Gly 340 345 350 His Glu Asn Leu Val Gly Gly Ala Thr Lys Arg Arg Leu Arg Glu Ala 355 360 365 Glu Ile Ser Ile Ser Gly Ser Asn Lys Glu Arg Trp Asp Phe His Phe 370 375 380 Gly Val Leu Glu His Arg Pro Leu Pro Lys His Ser His Leu Lys Glu 385 390 395 400 Trp Lys Leu Leu Phe Gln Asp Gly Ala Leu Trp Leu Cys Leu Val Val 405 410 415 Glu Leu Gln Gln Pro Leu Pro Glu Pro Ser Thr Leu Ala Ala Gly Leu 420 425 430 Asp Ile Gly Trp Arg Arg Thr Glu Glu Gly Ile Arg Phe Gly Thr Leu 435 440 445 Tyr Glu Pro Ala Thr Lys Thr Ile His Glu Leu Ile Ile Asn Phe Gln 450 455 460 Arg Ser Pro Lys Asp Pro Lys Gln Arg Val Pro Phe Arg Ile Asp Leu 465 470 475 480 Gly Pro Thr Arg Trp Glu Lys Arg Asn Ile Ile Lys Leu Leu Pro Lys 485 490 495 Trp Lys Pro Gly Asp Gly Ile Pro Asn Ser Leu Glu Ile Lys Ala Ala 500 505 510 Leu Gln Thr Arg Arg Asp Tyr Phe Lys Asp Thr Ala Lys Ile Leu Leu 515 520 525 Arg Lys His Leu Gly Glu Lys Thr Pro Ala Trp Leu Asp Lys Ala Gly 530 535 540 Arg Ser Gly Leu Leu His Leu Lys Glu Glu Phe Lys Asp Asp Ala Asp 545 550 555 560 Val Gln Glu Ile Leu Asn Thr Trp Ala Thr Asn Asp Glu Gln Val Gly 565 570 575 Lys Leu Ala Ser Lys Tyr Ser Ala Arg Val Thr Arg Arg Val Glu Tyr 580 585 590 Gly Gln Gln Gln Val Ala His Asp Val Cys Arg Tyr Leu Gln Gln Lys 595 600 605 Asp Ile Ala Arg Leu Val Val Glu Lys Asn Phe Leu Ala Lys Ile Ala 610 615 620 Gln His Gln Asp Asn Ser Asp Pro Glu Ser Leu Lys Arg Ser Gln Lys 625 630 635 640 Tyr Arg Gln Phe Ala Ala Val Gly Arg Phe Val Ser Lys Leu Lys Asp 645 650 655 Thr Val Val Lys Tyr Gly Gln Val Val Asp Ser Asp Glu Ala Thr Asn 660 665 670 Thr Thr Arg Met Cys Gln Tyr Cys Asn His Leu Asn Pro Ala Thr Glu 675 680 685 Lys Glu Lys Met Asp Cys Glu Gly Cys Gly Arg Glu Ile Lys Leu Asp 690 695 700 Asp Asn Ala Ala Ile Asn Leu Ser Arg Phe Ala Ser Asp Pro Glu Leu 705 710 715 720 Ala Glu Leu Ala Arg Ala Lys Lys Asn Lys Glu Ser Arg Ala Asp Pro 725 730 735 Lys Lys Lys Arg Lys Val 740 <210> 22 <211> 771 <212> PRT <213> artific <400> 22 Met Ile Lys Glu Leu Thr Tyr Lys Asp Asp Cys Val Phe Val His Ala 1 5 10 15 Ala Glu Thr Ser Leu Lys Thr Thr Lys Ala Phe Thr His Gly Arg Ile 20 25 30 Val Phe Asp Glu Ser Thr Arg Leu Ala Thr Phe Ala Gln Leu Pro Ala 35 40 45 Val Ala Ile Pro Ala Glu Leu His Leu Lys Ala Glu Thr Pro Met Asn 50 55 60 Phe Gln Leu Val Ala Pro Asp Gly Glu Leu Arg Pro Gln Lys Arg Ser 65 70 75 80 Gly Val Leu Leu Val Asp Glu Lys Gln Pro Thr Ile Phe His Ile Gly 85 90 95 Gly Met Ser Lys Thr Pro Arg Gln Phe Ser Ala Glu Asp Gly Trp Thr 100 105 110 Asn Asn Ile Tyr Lys Tyr Arg Ala Tyr Phe Thr His Glu Gly Thr Asn 115 120 125 Thr Ser Gly Asp Val Pro Glu Trp Leu Lys Ala Ser Ile Thr Arg Gln 130 135 140 Arg Asp Leu Trp Asn Arg Leu Ala Trp Leu Cys Arg Glu Ala Arg Arg 145 150 155 160 Gln Cys Ser Ser Gly Ser Pro Glu Glu Ile Thr Ala Phe Val Glu Thr 165 170 175 Val Ile Leu Pro Ala Ile Asp Ala Phe Asn Asp Ser Leu Gly Arg Ser 180 185 190 Lys Asp Lys Met Lys His Pro Ala Lys Leu Lys Val Glu Met Pro Gly 195 200 205 Leu Asp Gly Leu Trp His Phe Val Gly Asp Leu Arg Lys Arg Val Asp 210 215 220 Lys Asp Arg Pro Val Pro Glu Gly Leu Leu Glu Lys Val Ile Asp Phe 225 230 235 240 Ala Gln Gln His Lys Thr Asp Tyr Thr Pro Leu Asp Val Phe Val Lys 245 250 255 Asp Phe Gly Ala Ile Ala Asn Lys Glu Ala Lys Asp Leu Lys Leu Arg 260 265 270 Ser Phe Glu Ile Arg Pro Thr Val Thr Ala Phe Lys Ala Thr Leu Asp 275 280 285 Arg Arg Arg Thr Lys Lys Leu Thr Trp Ser Glu Gly Trp Pro Leu Ile 290 295 300 Lys Tyr Asn Asp Ser Pro Lys Ala Gly Asp Trp Gly Leu His Tyr Tyr 305 310 315 320 Phe Asn Lys Ala Gly Val Asn Ser Ser Leu Leu Glu Ser Lys Glu Gly 325 330 335 Val Pro Gly Leu Ser Phe Gly Ala Pro Leu Asp Pro Ala Asn Thr Gly 340 345 350 His Lys Leu Ile His Thr Ala Ser Gly Gly Lys Thr Leu Arg Glu Ala 355 360 365 Glu Ile Ser Ile Arg Gly Asp Asp Asn Gln Gln Trp Arg Phe Arg Phe 370 375 380 Ala Val Ile Gln Trp Arg Pro Leu Pro Pro Asn Ser His Val Lys Glu 385 390 395 400 Trp Lys Leu Leu Leu Gln Asp Gly Lys Leu Trp Leu Cys Leu Val Val 405 410 415 Glu Leu Gln Arg Asp Leu Pro Thr Ser Thr Ala Leu Ala Ala Gly Leu 420 425 430 Asp Val Gly Trp Arg Arg Thr Glu Val Gly Ile Arg Phe Gly Thr Leu 435 440 445 Tyr Glu Pro Thr Ser Trp Thr Val Arg Glu Leu Val Met Asp Phe Glu 450 455 460 Gln Ser Pro Lys Asp His Lys Asp Arg Val Pro Phe Arg Phe Asp Met 465 470 475 480 Gly Pro Thr Arg Trp Glu Lys Arg Asn Ile Thr Arg Leu Phe Pro Asp 485 490 495 Trp Lys Pro Gly Asp Gly Ile Pro Asn Ala Leu Glu Ile Arg Thr Ala 500 505 510 Leu Gln Thr Arg Arg Asp Tyr Phe Lys Asp Thr Ala Lys Ile Leu Leu 515 520 525 Arg Lys His Leu Gly Glu Lys Thr Pro Ala Trp Leu Glu Lys Ala Gly 530 535 540 Arg Ser Gly Leu Leu Arg Leu Lys Glu Glu Phe Lys Asp Asp Val Asp 545 550 555 560 Val Gln Gly Ile Leu Asn Thr Trp Gly Lys Asn Glu Glu Glu Ile Gly 565 570 575 Thr Leu Ala Ala Ala Tyr Ser Lys Lys Val Thr Thr Arg Ile Glu Tyr 580 585 590 Asn Gln Leu Gln Ile Ala His Asp Val Cys Arg Tyr Leu Lys Gln Lys 595 600 605 Gly Ile Thr Arg Leu Ile Leu Glu Lys Asn Phe Leu Ala Lys Ile Ala 610 615 620 Leu His Gln Glu Asn Thr Asp Pro Val Ser Leu Lys Arg Ser Gln Lys 625 630 635 640 Tyr Arg Gln Phe Ala Ala Val Gly Arg Phe Ile Ser Leu Leu Lys Tyr 645 650 655 Thr Ala Val Lys Tyr Gly Ile Val Thr Glu Pro Tyr Glu Ala Met Asn 660 665 670 Thr Thr Arg Met Cys Gln Tyr Cys Asn His Leu Asn Pro Ser Thr Glu 675 680 685 Lys Asp Arg Phe Gln Cys Glu Ala Cys Gly Arg Thr Ile Asp Gln Asp 690 695 700 Tyr Asn Ala Ala Val Asn Leu Ser Arg Phe Ala Cys Asp Pro Lys Leu 705 710 715 720 Ala Lys Leu Ala Leu Gly Asn Gly Asn Glu Ser Glu Glu Asp Ala Gln 725 730 735 Thr Pro Glu Phe Glu Asp Glu Glu Asp Thr Gln Ser Leu Glu Ser 740 745 750 Asp Glu Ser Gly Thr Asn Asp Leu Ser Arg Ala Asp Pro Lys Lys Lys 755 760 765 Arg Light Val 770 <210> 23 <211> 3071 <212> DNA <213> artifice <400> 23 tttacacttt atgcttccgg ctcgtatgtt aggaggtctt tatcatgaaa gcaaagtatc ggtctaactc aattcatgcc tcgctccggc agcttttgga gcttggctta agcaagagca 120 gttcggcaga accacagaaa atcacgcgca ctgccaagtt caagaat accgacatcc ggccggattt ggttccggtc ttgaaccggc attttgatgc gtttgaaaag ttccggcgca 240 aggtgctgga tgagcttga gcacgatgga acaaagacca gaaatcgttt caagcaatgg 360. tgcaatgcag cgcgaagga ccgtatcaga agaaatccag ttgttatgct tggctcgaca ctcatttcat caccgaagca aaagcggttc ttgacttacc gcgcaaacca gccaccagcc 420 ttctttacaa tttgagtggc ggattgaaag gttttctgac acgccgcgag acagtcaccg 480 aggacattca aaaacgcttc aatgacaatc tccgagaatg gaacggcgat ttaagtcagt 540 tggccagcga cctcaaagct cccttgccgc cagcaccccc gaatttggat tttgaaaacc 600 tgaccgaaag agcaatcgaa aggtataacg attgggttgg gcggacgcgc gcatggtgca 660 atctcattct tgtccagcaa aagaaagttg aacgccgcga tgcctgtttg ccgcgctatc 720 ttaaaggata tccgggcttt ttcggctcac aacgctatgc cacaacagcg agtctagccg 780 agaacttgaa aaaattggaa caggaggcac gagagcaaag caaaaaagcg ccaactcgct 840 ttgccaaact ctcacctgaa atttggacag ccattcaaga acgattctca cctcccgaag 900 ggcatgaagc cggagaaaag cgccgtccgc gcacggctca ccagactgtc tgcctccgct 960 ttgctgcttt gcgggcagcg catccagagt ggactccggc gcaactggcc ggagaaattt 1020 tagctggtat tttccatgga acggaaaaac ttaaaaagca tcttgccgct aacggctttg 1080 aagatcgccg ggcggtcatc aagttggcaa atctttacaa tgttgtggtt gccttctcgc 1140 tcgacccaat ccgcgccgcc ggaaattaca tttcgttttta caaggaggag acgcccaaga 1200 gaaacgcctt cggcgacgtg cggggtggac tgcatcagcc gagtgatgaa tcggccgcaa 1260 ttgaaatcat gggttttggt ttgcagaaag aatcgggcaa acccctctac aacggcttgc 1320 tcgtctgcaa aaaatcagaa aaggaacatg acgattcgtg ggcgtttctt tattgccaca 1380 1440 ctgactggac tggctttgcc agtcgtggcg gctcgcggaa aaaggctgaa gcaagtgcga 1500 agcaattggt gcgaggacgt gtctggatta gcgaaaaaac tccgccgaca gtcatgccgc 1560 tggcgttcgg gtcacgtcag ggacgtgaat atctctggca ttttgaccgg gatttgcgtg 1620 agaaaaacga atgggttttg ggcaatggca gactgctccg catcatgccg cccggacgac 1680 cgaacgccgc cgacttttat ctgacaatca ctctcgaacg acaagcgcca cctgttgctg 1740 aatttatagc gagcaaatta atcgggattg atcgcggcga agccgttccc gcagcttatg 1800 cgctgattga ccgtgaagga agatggttgg tcaatcaaaa gcgatttcaa gaattcctga 1860 aagcgctggg cacttggcat gatgagttgg atgagtggtg gaaaaaac aagaacactc 1920 caataaacga ccgggtaaaa cggccccggt tgcaatggac aagtgaactg gaggatggat 1980 ttggatttgt tgcgagcgaa taccgcgaac aacaagccga ttttaatgac caaaagcgtg 2040 aattgcagcg aacccaaggc ggttacacgc gctggttacg aagcaaagag cgcaatcgcg 2100 ctcgtgcttt gggtggtgaa gtcacacgcg cggtgctgtc gcttgcctcc gaacatcgtt 2160 cgcctctggt tttggaaaaa cttgggagcg cgttggcgac gcgcggcggc aaaggaaaaa 2220 tgatgagtca aatgcagtat gaacgaatgc tcaccacgct tgaaaaag ctcgcagagg 2280 tcggacttta tgcgatgcct tcagcaccaa aatatcgcaa gctcaccaat ggcttcatta 2340 atttcgccgc accacactat acgtcctcga catgctcttc gtgcgggcaa gtgcatgaca 2400 gcattttcta caatgagctt gcgaaaaaaa tcgtgcgccc aaacgattca aagtggcaag 2460 tgacgcttcc aagtgggcag acacgaacat tgcctcaaac ttacagctat tggttgaaag 2520 gaaaaggtga gcagaccaaa caaacgaacg agcgtctcaa cgaactgctg ggtgaacgaa 2580 cagttactca attgagcaag agcagctata tgacattggt aagtttgctc aaaatatagtt 2640 ggttgcccta ccgcccacaa caagctgatt tttgttgtct gaattgtggc tacaccacga 2700 atgccgacgt gcaaggtgcg ttaaacattg cccgaaagtt tttgttccgt tccgaacgag 2760 gcaaaaagac tgacgacgca ggcgaagaag gtgaggcaaa gcgtggtaaa tacaaagacg 2820 actggcagac ttggtatcgc caaaagttgg aaaaggtctg gtgcaagtaa caaataaaac 2880 gaaaggctca gtcgaaagac tgggccttc gttttatctg ttgtttgtcg gtgaacgctc 2940 tcctgagtag gacaaatttg acagctagct cagtcctagg tataatgcta gccgatgcac 3000 tcaaactgca ttgctacggt ataacaactt cgacgagctc tacaagaagg ccatcctgac 3060 ggatggcctt t 3071 <210> 24 <211> 2453 <212> DNA <213> fireworks <400> 24 tttacacttt atgcttccgg ctcgtatgtt aggaggtct tatcatgatc acaaagttga 60 cgtacaacga caattgcgtc tttgtacatc cggacgataa gtctctaaaa acggttaaag 120 cctttacgca tggcagcatc gtatttgacg tgaatgcgcg caaagtgacc tttgcaaagt 180 tgcctgcggt tgacgttcct gaagggctca agatcgaaga gggcgcgacc ctgaactttt 240 tgttgatggc gcccgatgga accttgcgta agcagatccg caacggcgtg ttgcaggcca 300 acgaaaaaca gccaggcttt ctgcgcattg gcggtatggc acaaagccca cgcaagtttt 360 cgcgcgaaga tggctggggc aacatgatct acaagtaccg cgcctatttc acgcatccgg 420 ggttgacgac agacacagaa ctgccggtat ggcttaagga ctcaatcaag cgccagaaag 480 agtattggaa tcggttggcg tggctttgcc gcgaagcgcg gcgcaaatgt tctccagtca 540 agagcgaaga aatcgcggcc tttgtaaaaa ccgagattct gccggcgatt gacactctga 600 atgattcgct cgggcgctcc aaagacaaga tgaagcaccc tgcaaaactc aagattgaag 660 atcctggcgt ggacggcctg tggaagttcg ttggtgaact gcgtcaccgt gttgaaaagg 720 accgtccggt tccaccgggg ctactggaaa aggtgattgc ctttgccgag caatttaaag 780 cagactacac cgcgctcaac gatttcatga acaatttgcc gtcgattgcg gaaagggaag 840 cagaaactct tcatctgcgc cgttttgaag tcaggccgac gcttgcttct ttcagggcga 900 cgctcgatcg tcgcaagacc accaaagcag cctggtcgga aggctggccg ctgatcaaat 960 atcctgacag tcccaaggca gaagactggg gtctgcacta ctacttcaac aaggcgggta 1020 ttggctctga tcttctagag gaaggtgatg gcgtcccagg gttgagcttt ggtcctgcgc 1080 tgtcgcccgc aaaaacagga catgaaaatc tggttggcgg ggcgacaaag cggcgactac 1140 gggaagcgga gatttctatt tctggcagca acaaagaacg gtgggatttt cattttggag 1200 tgttggaaca ccggcctctg ccgaagcact cgcatctgaa ggagtggaaa ttgctttttc 1260 aagatggagc actgtggctt tgcctagttg tggagctaca acagccattg ccggagccca 1320 gcacgcttgc ggcggggctt gatattggct ggcgcagaac agaagaaggc attcggttcg 1380 gaacgctgta tgagcccgca accaagacaa tccacgaact gatcatcaat tttcagcggt 1440 cgcccaaaga tcccaagcaa cgcgtgcctt tccggattga tcttggcccc acgcgttggg 1500 aaaaacgcaa cattatcaag ctgcttccca aatggaagcc aggagacgga attccaaatt 1560 ctttggaaat caaagccgcg ctgcagacgc gtcgtgatta cttcaaggac actgcaaaga 1620 ttttgttgcg caagcatcta ggagaaaaaa cgccggcatg gcttgacaaa gctggccgta 1680 gtggcttgct gcatctgaag gaagaattca aagacgatgc cgatgtgcaa gagatcttga 1740 acacgtgggc tacaaatgat gagcaggtcg gaaaactggc gtcaaaatac tccgcacgag 1800 ttactagacg cgttgaatat ggccaacaac aggttgccca cgatgtttgc cgttaccttc 1860 agcaaaagga cattgctcgt ttggtcgtgg agaagaactt cttggccaag attgcccagc 1920 atcaagacaa tagcgatcca gaaagtttga agcgttcgca gaaataccgg caatttgcgg 1980 cagtgggtcg ttttgtttca aagctgaagg acactgtcgt taatacggc caagtggttg 2040 actcggatga ggcaacaaat acaacgcgta tgtgccagta ttgcaaccat ctgaatccgg 2100 caaccgaaaa agaaaagatg gattgcgagg gatgcggaag agaaatcaag ctggacgaca 2160 atgcggcaat aaatctctcg cgctttgctt ccgatccaga gctggcagaa ctggcgcgag 2220 caaagaaaaa tag caataaaac gaaaggctca gtcgaaagac tgggcctttc gttttatctg ttgtttgtcg gtgaacgctc tcctgagtag gacaaatttg acagctagct cagtcctagg fathergcta gcgttgcatc tacgcgatc actcaaacca cgttgctacg gtatacaac ttcgacgagc tctacaaga ggccatcctg acggatggcc ttt <210> 25 <211> 2542 <212> DNA <213> artifice <400> 25 tttacacttt atgcttccgg ctcgtatgtt aggaggtctt tatcatgatc aaagaattga cctacaagga cgattgcgtc tttgtgcatg ccgcagaaac gtccctgaaa accaccaagg cgtttacaca tggccgcatt gtgtttgacg agagcacgcg cttggcgacg tttgcccagc 180 tgccggcagt ggcgattccc gccgagcttc atctcaaggc agagacacca atgaattttc 240 agctggtggc ccctgacggc gagctgcgtc cgcagaaacg atcgggagtg ctcctggtag 300 360. acgaaaagca accaaccatc ttccacattg gcggaatgtc aaaaacgccg cggcagtttt cagcagaga tggatggacg aacaacatct acaaataccg cgcgtatttc acccacgag gcacaaacac gagcggcgac gtacctgaat ggctcaaagc atcgatcacg cggcaaaggg 480 atctttggaa ccggttggcg tggctttgcc gcgaagcacg gcgccagtgc tcttccggtt 540 cgcccgagga gattaccgct tttgtcgaaa ccgttatttt gccagcgatc gatgctttta 600 acgactcgct tggccgttcc aaagacaaga tgaagcaccc ggccaagctg aaggtcgaga 660 tgcccggctt ggatggcttg tggcatttg ttggcgatct gcggaagcgc gtcgacaaag 720 atcgtcccgt gccagaaggc ttgctggaga aggtgattga ctttgcccag cagcacaaga 780 cggactatac gccgctcgac gtgtttgtaa aggattttgg ggcgattgcc aacaaggaag 840 caaaggatct caagctacga agctttgaaa ttcggccgac agtcaccgct ttcaaggcaa 900 ccctggatcg ccgcagaacc aagaagttga cttggtccga gggctggcca ctcatcaagt 960 acaacgacag ccccaaggcc ggcgattggg gcctgcatta ctacttcaac aaggctggcg 1020 tcaattcatc actgctggaa tcgaaggaag gcgtaccagg attgtcgttt ggagcgccgc 1080 tcgatcctgc caatacgggg cataagctaa tccatacggc aagcggtggc aagacgcttc 1140 gggaagcgga gatttccatt cgtggagacg ataaccagca gtggcgtttt cgttttgcag 1200 tcattcagtg gcggccgttg cccccaaatt cccacgtaaa agagtgaaa ctgctattgc 1260 aggatggcaa actttggctt tgcctggtgg ttgagctgca acgtgaccta ccaacttcca 1320 ctgcccttgc agccggtttg gatgttggct ggcggcgcac agaagtggga attcggtttg 1380 gcacactgta tgagccaacg agctggacgg tgcgcgaact tgtgatggac tttgaacaat 1440 cgcccaaaga tcaaaagat cgtgtgccct tccgcttcga tatggggcct acacgctggg 1500 agaaacgcaa tatcactcgg ctattccctg attggaagcc gggcgacgga attcccaacg 1560 cgttggagat cagaaccgcg ttgcagaccc gccgtgacta tttcaaagat acggccaaga 1620 ttctgctgcg caagcatctt ggggaaaaaa cgccagcatg gctggaaaaa gccggtcgta 1680 gtggattgct gcggctcaag gaagaattca aagacgatgt cgatgtgcaa ggcatcctga 1740 atacgtgggg gaaaacgag gaagaaatcg gaactcttgc ggccgcatac agcaaaaagg 1800 ttaccacgcg cattgagtac aaccagctgc aaatcgccca cgatgtttgc cgctatctga 1860 agcaaaaggg cattacacgg ttgatcctgg aaaagaactt cctggccaaa attgcgctcc 1920 1980 ccgttggccg ctttatctct ctgctcaagt acactgccgt aaaatacgga attgtgactg 2040 2100 cgacggaaaa ggatcgcttc cagtgcgaag cgtgtggaag aacgattgac caggactaca 2160 acgctgcggt aaatctgtca cgttttgcct gcgatccaaa gttggcaaag ctggcgctgg 2220 gaaatggaa cgaatctgaa gaacgccc agacgccgga atttgaac gaagaaag 2280 acacgcagtc gctggaaagt gacgaaagcg gcacaaacga tctttgacaa ataaaacgaa 2340 aggctcagtc gaagaactgg gcctttcgtt ttatctgttg tttgtcggtg aacgctctcc 2400 tgagtaggac aaatttgaca gctagctcag tcctaggtat aatgctagca gtagcaacgt 2460 2520 gccatcctga cggatggcct tt 2542 <210> 26 <211> 35 <212> DNA <213> crafts <220> <221> misc_feature <222> (1)..(8) <223> n is a, c, g, or t <400> 26 nnnnnnnngg tataacaact tcgacgagct ctaca 35 <210> 27 <211> 27 <212> RNA <213> crafts <400> 27 gguauaacaa cuucgacgag cucuaca 27 <210> 28 <211> 21 <212> RNA <213> crafts <400> 28 cuaaggaauau ugaagggggg c 21 <210> 29 <211> 24 <212> RNA <213> crafts <400> 29 cuuccaucag agaaccucac ugcg 24

Claims

1. An effector protein in a CRISPR / Cas system, having an amino acid sequence as set forth in SEQ ID NO:

1.

2. A conjugate comprising the protein of claim 1 and a modification moiety; wherein the modification moiety is selected from the group consisting of another protein or polypeptide, a detectable label, and any combination thereof; wherein the another protein or polypeptide is selected from the group consisting of an epitope tag, a reporter gene sequence, a nuclear localization signal sequence, a transcription activation domain, a transcription repression domain, a nuclease domain, and any combination thereof.

3. The conjugate of claim 2, wherein, the modification moiety is linked to the N- or C-terminus of the protein via a linker, or the modification moiety is fused to the N- or C-terminus of the protein.

4. The conjugate of claim 2, having one or more of the following features: (i) the transcription activation domain is VP64; (ii) the transcription repression domain is a KRAB domain or a SID domain; (iii) the nuclease domain is Fokl.

5. The conjugate of claim 2, wherein, the conjugate comprises an epitope tag.

6. The conjugate of claim 2, wherein, the conjugate comprises an NLS sequence.

7. The conjugate of claim 6, wherein, the NLS sequence is as set forth in SEQ ID NO:

19.

8. The conjugate of claim 7, wherein, the NLS sequence is located at the N- or C-terminus of the protein.

9. A fusion protein comprising the protein of claim 1 and a further protein or polypeptide; wherein, the another protein or polypeptide is selected from the group consisting of an epitope tag, a reporter gene sequence, a nuclear localization signal sequence, a transcription activation domain, a transcription repression domain, a nuclease domain, and any combination thereof.

10. The fusion protein of claim 9, wherein, the another protein or polypeptide is optionally linked to the N- or C-terminus of the protein via a linker.

11. The fusion protein of claim 9, having one or more of the following features: (i) the transcription activation domain is VP64; (ii) the transcription repression domain is a KRAB domain or a SID domain; (iii) the nuclease domain is Fokl.

12. The fusion protein of claim 11, wherein, the fusion protein comprises an epitope tag.

13. The fusion protein of claim 11, wherein, the fusion protein comprises an NLS sequence.

14. The fusion protein of claim 13, wherein, the NLS sequence is as set forth in SEQ ID NO:

19.

15. The fusion protein of claim 14, wherein, the NLS sequence is located at the N- or C-terminus of the protein.

16. The fusion protein of claim 9, wherein, the fusion protein has an amino acid sequence as set forth in SEQ ID NO:

20.

17. An isolated nucleic acid molecule, having a sequence as set forth in SEQ ID NO:

10.

18. A complex comprising: (i) a protein component selected from the group consisting of the protein of claim 1, the conjugate of any one of claims 2-8, or the fusion protein of any one of claims 9-16; and (ii) a nucleic acid component comprising, in the 5’ to 3’ direction, the isolated nucleic acid molecule of claim 17 and a guide sequence capable of hybridizing to a target sequence, wherein, the protein component and the nucleic acid component are associated with each other to form the complex.

19. The complex of claim 18, wherein, the guide sequence is linked to the 3’ end of the nucleic acid molecule.

20. The complex of claim 18, wherein, the guide sequence comprises a complement of the target sequence.

21. The complex of claim 18, wherein, the nucleic acid component is a guide RNA in a CRISPR / Cas system.

22. The complex of claim 18, wherein, the nucleic acid molecule is an RNA.

23. The complex of claim 18, wherein, the complex does not comprise a trans-acting crRNA.

24. An isolated nucleic acid molecule, comprising: (i) a nucleotide sequence encoding the protein of claim 1, or the fusion protein of any one of claims 9-16; (ii) a nucleotide sequence of the isolated nucleic acid molecule of claim 17.

25. The isolated nucleic acid molecule of claim 24, wherein, The nucleotide sequence of any one of (i)-(ii) is codon-optimized for expression in a prokaryotic cell or a eukaryotic cell.

26. A vector comprising the isolated nucleic acid molecule of claim 24 or 25.

27. A host cell comprising the isolated nucleic acid molecule of claim 24 or 25, or the vector of claim 26.

28. A composition comprising: (i) a first component selected from the group consisting of: the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, a nucleotide sequence encoding the protein of claim 1 or the fusion protein of any one of claims 9-16; and (ii) a second component which is a nucleotide sequence of a guide RNA; wherein, the guide RNA comprises, in the 5’ to 3’ direction, a direct repeat sequence and a guide sequence, the guide sequence being capable of hybridizing to a target sequence; the guide RNA is capable of forming a complex with the protein, conjugate or fusion protein of (i).

29. The composition of claim 28, wherein, the direct repeat sequence is the isolated nucleic acid molecule as defined in claim 17.

30. The composition of claim 28, wherein, the guide sequence is linked to the 3’ end of the direct repeat sequence.

31. The composition of claim 28, wherein, the guide sequence comprises a complement of the target sequence.

32. The composition of claim 28, wherein, the composition does not comprise a trans-acting crRNA.

33. A composition comprising one or more vectors, the one or more vectors comprising: (i) a first nucleic acid comprising a nucleotide sequence encoding the protein of claim 1 or the fusion protein of any one of claims 9-16; optionally the first nucleic acid is operably linked to a first regulatory element; and (ii) a second nucleic acid comprising a nucleotide sequence encoding a guide RNA; optionally the second nucleic acid is operably linked to a second regulatory element; wherein: the first nucleic acid and second nucleic acid are present on the same or different vectors; the guide RNA comprises, in the 5’ to 3’ direction, a direct repeat sequence and a guide sequence, the guide sequence being capable of hybridizing to a target sequence; the guide RNA is capable of forming a complex with the protein or fusion protein of (i).

34. The composition of claim 33, wherein, the direct repeat sequence is the isolated nucleic acid molecule as defined in claim 17.

35. The composition of claim 33, wherein, the guide sequence is linked to the 3’ end of the direct repeat sequence.

36. The composition of claim 33, wherein, the guide sequence comprises a complement of the target sequence.

37. The composition of claim 33, wherein, the composition does not comprise a trans-acting crRNA.

38. The composition of claim 33, wherein, the first regulatory element and / or second regulatory element is a promoter.

39. The composition of claim 38, wherein, the promoter is an inducible promoter.

40. The composition of any one of claims 28-38, wherein, when the target sequence is DNA, the target sequence is located 3’ to a protospacer adjacent motif (PAM), wherein the target sequence does not have a PAM domain constraint.

41. The composition of claim 40, wherein, the target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell; or, the target sequence is a non-naturally occurring DNA or RNA sequence.

42. The composition of any one of claims 28-39, 41, wherein, The protein is linked to one or more NLS sequences, or the conjugate or fusion protein comprises one or more NLS sequences.

43. A kit comprising one or more components selected from the group consisting of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the host cell of claim 27, the composition of any one of claims 28-42.

44. A delivery composition comprising a delivery vehicle, and one or more selected from the group consisting of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the host cell of claim 27, the composition of any one of claims 28-42.

45. The delivery composition of claim 44, wherein the delivery vehicle is selected from the group consisting of a lipid particle, a metal particle, a protein particle, an exosome, a gene gun, or a viral vector.

46. The delivery composition of claim 44, wherein the delivery vehicle is selected from the group consisting of a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus.

47. A method of modifying a target gene for non-disease diagnostic and therapeutic purposes comprising: contacting the complex of any one of claims 18-23 or the composition of any one of claims 28-42 with the target gene, or delivering into a cell comprising the target gene; the target sequence is present in the target gene.

48. The method of claim 47, wherein the target gene is present in a cell, or the target gene is present in an in vitro nucleic acid molecule.

49. The method of claim 47, wherein the cell is a prokaryotic cell.

50. The method of claim 47, wherein the cell is a eukaryotic cell.

51. The method of claim 47, wherein the cell is selected from the group consisting of a plant cell.

52. The method of claim 47, wherein the method results in a target sequence break.

53. The method of claim 47, wherein the method results in a DNA double strand break.

54. A method of altering expression of a gene product for non-disease diagnostic and therapeutic purposes comprising: contacting the complex of any one of claims 18-23 or the composition of any one of claims 28-42 with a nucleic acid molecule encoding the gene product, or delivering into a cell comprising the nucleic acid molecule; the target sequence is present in the nucleic acid molecule.

55. The method of claim 54, wherein the nucleic acid molecule is present in a cell.

56. The method of claim 54, wherein the cell is a prokaryotic cell.

57. The method of claim 54, wherein the cell is a eukaryotic cell.

58. The method of claim 54, wherein the cell is selected from the group consisting of an animal cell, a plant cell.

59. The method of claim 54, wherein the nucleic acid molecule is present in an in vitro nucleic acid molecule.

60. The method of claim 54, wherein expression of the gene product is altered.

61. The method of claim 54, wherein the gene product is a protein.

62. The method of any one of claims 47-61, wherein the protein, conjugate, fusion protein, complex, vector, or composition is comprised in a delivery vehicle.

63. The method of claim 62, wherein the delivery vehicle is selected from the group consisting of a metal particle, a protein particle, a liposome, an exosome, a viral vector.

64. The method of any one of claims 47-61, for use in modifying a cell, cell line, or organism by altering one or more target sequences in a target gene or nucleic acid molecule encoding a target gene product.

65. An ex vivo or in vivo cell or cell line, or progeny thereof, comprising: the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the composition of any one of claims 28-42.

66. The cell or cell line, or progeny thereof, of claim 65, wherein the cell is a eukaryotic cell.

67. Use of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the composition of any one of claims 28-42, or the kit of claim 43, in the manufacture of a preparation for nucleic acid editing.

68. The use of claim 67, wherein the nucleic acid editing comprises gene editing.

69. The use of claim 68, wherein the gene editing comprises modifying a gene, knocking out a gene, altering expression of a gene product, repairing a mutation, and / or inserting a polynucleotide.

70. Use of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the composition of any one of claims 28-42, or the kit of claim 43, in the manufacture of a preparation for: (i) ex vivo gene or genome editing; (ii) editing a target sequence in a target locus to modify an organism.

Citation Information

Patent Citations

  • Systems methods and compositions for sequence manipulation

    US61736527P0

  • Systems Methods and Compositions for Sequence Manipulation

    US61748427P0

  • Crispr / cas effector protein and system

    WO2019201331A1

  • Crispr-cas12j enzyme and system

    WO2020098772A1

  • Crispr-cas12a enzyme and system

    WO2020098793A1