A gene-editing protein, its corresponding gene-editing system and its applications

By designing a novel CRISPR/Cas system that combines gene-editing proteins and functional domains, the shortcomings of existing systems have been overcome, enabling efficient and robust targeted editing of nucleic acids or polynucleotides, and enhancing the system's diversity and editing precision.

CN116622678BActive Publication Date: 2026-03-06YOLTECH THERAPEUTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310593098.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2026-03-06
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems have limitations in targeting nucleic acid or polynucleotide editing, necessitating the development of more robust and diverse alternative systems.

Method used

A novel CRISPR/Cas system is provided, comprising gene editing proteins and fusion proteins, which combine functional domains such as localization signals, nucleases, and transcription activation domains. By designing specific guide RNAs to bind to target sequences, efficient targeted editing can be achieved.

Benefits of technology

It enables efficient and robust editing of targeted nucleic acids or polynucleotides, enhancing the system's diversity and editing precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0005577327030000131
    Figure GDA0005577327030000131
  • Figure GDA0005577327030000141
    Figure GDA0005577327030000141
  • Figure GDA0005577327030000311
    Figure GDA0005577327030000311
Patent Text Reader

Abstract

This invention provides a gene-editing protein, its corresponding gene-editing system, and its applications. Specifically, the gene-editing protein of this invention exhibits excellent gene-editing activity in vitro, can effectively edit or cut target genes, and can effectively treat the symptoms or diseases of subjects in need.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing, specifically to a gene editing protein, its corresponding gene editing system, and its applications. Background Technology

[0002] The Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) system is a system developed by bacteria and archaea to defend against invading bacteriophage DNA. The CRISPR system comprises two families: family I is further divided into types I, III, and IV; family II is divided into types II, V, and VI. These six system types are further subdivided into 19 subtypes. Many prokaryotes contain multiple CRISPR-Cas systems, indicating that they are compatible and may share components.

[0003] The most common type II system is the CRISPR / Cas9 system. The Cas9 protein, with the assistance of a trans-encoding small RNA (tracrRNA), processes pre-crRNA into mature crRNA that binds to tracrRNA. Later, it was discovered that by artificially constructing single-stranded chimeric guide RNA (gRNA) that mimics the crRNA-tracrRNA complex, the recognition and cleavage of the Cas9 protein at the target site can be effectively mediated. The three bases immediately adjacent to the 3′ end of the target site must be in the form of 5′-NGG-3′, thus forming the PAM (protospacer adjacent motif) structure required for the Cas / crRNA complex to recognize the target site.

[0004] Currently known CRISPR / Cas have their own advantages and disadvantages. For example, Cas9 requires two RNAs as guide RNAs.

[0005] There is an urgent need for alternative and robust editing systems and techniques for targeted nucleic acids or polynucleotides with broad application value.

[0006] Therefore, the development of biotechnology still requires the development of new CRISPR / Cas systems with diverse characteristics. Summary of the Invention

[0007] The main objective of this invention is to provide a new CRISPR / Cas system with diverse features.

[0008] Another objective of this invention is to discover new CRISPR-Cas systems that provide alternative and robust systems and techniques for targeting nucleic acids or polynucleotides, thereby addressing the shortcomings of currently known CRISPR-Cas systems.

[0009] A first aspect of the present invention provides a gene-editing protein, said protein being selected from the group consisting of:

[0010] (a) A polypeptide having the amino acid sequence shown in SEQ ID NO:1;

[0011] (b) A polypeptide having ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% homology (or identity) with the amino acid sequence shown in SEQ ID NO:1, and said polypeptide having the biological function of SEQ ID NO:1;

[0012] (c) A derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, more preferably 1-5) amino acid residues of any of the amino acid sequences shown in SEQ ID NO:1, and retaining the biological function of SEQ ID NO:1.

[0013] In another preferred embodiment, the gene-editing protein is an effector protein in the CRISPR / Cas system.

[0014] A second aspect of the present invention provides a fusion protein comprising the gene editing protein described in the first aspect of the present invention, and one or more functional domains.

[0015] In another preferred embodiment, the functional domain is selected from localization signals, reporter proteins, Cas protein targeting portions, DNA binding domains, epitope tags, transcription activation domains, transcription repression domains, nucleases, deamination domains, methyltransferases, demethylases, transcription release factors, HDACs, cleavage active peptides, ligases, integrases, transposases, recombinases, polymerases, and base excision repair inhibitors (such as uracil-DNA glycosyltransferase inhibitors (UGIs)).

[0016] In another preferred embodiment, the functional domain includes one or more of the following enzyme activities against the target sequence: methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, demyristylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) and deglycosylation activity.

[0017] In another preferred embodiment, the functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain.

[0018] In another preferred embodiment, the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0019] In another preferred embodiment, the functional structural domain is the full length or a functional segment of TadA8e.

[0020] In another preferred embodiment, the positioning signal includes a nuclear positioning signal (NLS) and / or a nuclear output signal (NES).

[0021] In another preferred embodiment, the sequence of the nuclear localization signal is located at, near, or close to the end (e.g., the N-terminus or C-terminus) of the protein described in the first aspect of the invention.

[0022] In another preferred embodiment, the nuclear output signal includes protein tyrosine kinase 2 (such as human protein tyrosine kinase 2).

[0023] In another preferred embodiment, the reporter protein includes glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, and autofluorescent protein.

[0024] In another preferred embodiment, the autofluorescent protein includes green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), and blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire).

[0025] In another preferred embodiment, the DNA binding domain includes methylation-binding proteins, LexADBD, and Gal4DBD.

[0026] In another preferred embodiment, the epitope tag includes a histidine tag, a V5 tag, a FLAG tag, an influenza virus hemagglutinin tag, a Myc tag, a VSV-G tag, a thioredoxin tag, and a streptavidin tag.

[0027] In another preferred embodiment, the transcriptional activation domain includes VP64 and / or VPR.

[0028] In another preferred embodiment, the transcriptional repression domain includes KRAB and / or SID.

[0029] In another preferred embodiment, the nuclease comprises FokI.

[0030] In another preferred embodiment, the cleavage-active polypeptide includes a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity.

[0031] In another preferred embodiment, the ligase comprises DNA ligase and / or RNA ligase.

[0032] In another preferred embodiment, the functional domain is connected to the N-terminus and / or C-terminus of the gene-editing protein.

[0033] In another preferred embodiment, the functional domain is inserted between the N-terminus and C-terminus of the gene-editing protein.

[0034] In another preferred embodiment, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the gene-editing protein via a adapter.

[0035] In another preferred embodiment, the functional domain is inserted between the N-terminus and C-terminus of the gene-editing protein via a connector.

[0036] In another preferred embodiment, the fusion protein has the following structure from the N-terminus to the C-terminus:

[0037] Z1-Z2(I); or

[0038] Z2-Z1(II); or

[0039] Z3-Z1-Z4(III);

[0040] Z1 is either cytosine deaminase or adenosine deaminase;

[0041] Z2 is the gene editing protein described in the first aspect of this invention;

[0042] Z3 is the N-terminal fragment of the gene-editing protein described in the first aspect of this invention;

[0043] Z4 is the C-terminal fragment of the gene-editing protein described in the first aspect of this invention;

[0044] Furthermore, each "-" independently represents a key or connector.

[0045] A third aspect of the present invention provides an isolated polynucleotide encoding the gene-editing protein described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention.

[0046] In another preferred embodiment, the polynucleotide is selected from the group consisting of:

[0047] (a) A polynucleotide with the sequence shown in SEQ ID NO.2;

[0048] (b) A polynucleotide whose nucleotide sequence is ≥70% homology to the sequence shown in SEQ ID NO.2 (preferably ≥80%, more preferably ≥90%, more preferably ≥95%, best ≥99%) and encodes the polypeptide shown in SEQ ID NO.1;

[0049] (c) A polynucleotide complementary to any of the polynucleotides described in (a)-(b).

[0050] In another preferred embodiment, the polynucleotide further comprises, flanking the ORF of the variant, an auxiliary element selected from the group consisting of: signal peptides, secretory peptides, tag sequences (such as 6His), or combinations thereof.

[0051] In another preferred embodiment, the polynucleotide is selected from the group consisting of genomic sequences, cDNA sequences, RNA sequences, or combinations thereof.

[0052] In another preferred embodiment, the polynucleotide also includes a promoter operatively linked to the ORF sequence of the variant.

[0053] In another preferred embodiment, the promoter is selected from the group consisting of: constitutive promoters, tissue-specific promoters, inducible promoters, or strong promoters.

[0054] In another preferred embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0055] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, plant cell, or mammalian cell (including human and non-human mammals).

[0056] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0057] In another preferred embodiment, the yeast cells are selected from one or more sources of yeast from the group consisting of: Pichia pastoris, Kluyveromyces, or combinations thereof; preferably, the yeast cells include: Kluyveromyces, more preferably Kluyveromyces marxi, and / or Kluyveromyces lactis.

[0058] In another preferred embodiment, the host cell is selected from the group consisting of: Escherichia coli, wheat germ cells, insect cells, SF9, HeLa, HEK293, CHO, yeast cells, or combinations thereof.

[0059] A fourth aspect of the present invention provides an isolated nucleic acid molecule comprising, or composed of, sequences selected from, the following:

[0060] (i) The sequence shown in SEQ ID NO: 3;

[0061] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO: 3;

[0062] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the sequence shown in SEQ ID NO: 3;

[0063] (iv) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iii); or

[0064] The complementary sequence of the sequence described in any of (v)(i)-(iii);

[0065] Furthermore, the sequences described in any one of (ii)-(v) substantially retain the biological function of the sequences from which they are derived;

[0066] For example, the isolated nucleic acid molecule is RNA;

[0067] For example, the isolated nucleic acid molecule contains a homologous repeat sequence in the CRISPR / Cas system.

[0068] In another preferred embodiment, the nucleic acid molecule comprises one or more stem-loops or optimized secondary structures;

[0069] For example, the sequence in any one of (ii)-(v) retains the secondary structure of the sequence from which it originates.

[0070] In another preferred embodiment, the nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0071] (a) The nucleotide sequence shown in SEQ ID NO: 3;

[0072] (b) A sequence that hybridizes with the sequence described in (a) under stringent conditions; or

[0073] (c) The complementary sequence of the nucleotide sequence shown in SEQ ID NO: 3.

[0074] The fifth aspect of the present invention provides a guide RNA (gRNA) comprising a direct repeat (DR) sequence capable of binding to the gene editing protein described in the first aspect of the present invention and a spacer sequence capable of targeting a target sequence.

[0075] A sixth aspect of the present invention provides a composite comprising:

[0076] (i) Protein components selected from the group consisting of: gene-editing proteins described in the first aspect of the present invention, fusion proteins described in the second aspect of the present invention, or combinations thereof; and

[0077] (ii) Nucleic acid components selected from the group consisting of: the guide RNA of the fifth aspect of the present invention, nucleic acid encoding the guide RNA of the fifth aspect of the present invention, precursor RNA of the guide RNA of the fifth aspect of the present invention, precursor RNA nucleic acid encoding the guide RNA of the fifth aspect of the present invention, or combinations thereof;

[0078] The protein component and the nucleic acid component combine to form a complex.

[0079] In another preferred embodiment, the direct repeat (DR) sequence in the guide RNA (gRNA) is attached to the 3' or 5' end of the nucleic acid molecule.

[0080] In another preferred embodiment, the spacer sequence in the guide RNA (gRNA) contains a complementary sequence to the target sequence.

[0081] A seventh aspect of the present invention provides a carrier comprising the polynucleotide described in the third aspect of the present invention or the nucleic acid molecule described in the fourth aspect of the present invention.

[0082] In another preferred embodiment, the carrier comprises:

[0083] (1) A first regulatory element, operatively linked to a nucleotide sequence encoding a gene-editing protein according to a first aspect of the invention or a nucleotide sequence encoding a fusion protein according to a second aspect of the invention; and

[0084] (2) A second regulatory element, operatively linked to a nucleotide sequence encoding a guide RNA, the guide RNA comprising:

[0085] (a) Spacer sequences capable of hybridizing with the target sequence, and

[0086] (b) A direct repeat (DR) sequence, which is linked to the spacer sequence, capable of guiding the gene-editing protein of the first aspect of the invention to bind to the guide RNA to form a complex of the sixth aspect of the invention that targets the target sequence.

[0087] In another preferred embodiment, the first control element and the second control element are located on the same or different carriers.

[0088] In another preferred embodiment, the first regulating element and / or the second regulating element is a promoter, such as an induced promoter.

[0089] In another preferred embodiment, the vector comprises one or more promoters operatively linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectivity marker, nucleic acid restriction site, and / or homologous recombination site.

[0090] In another preferred embodiment, the vector includes plasmids and viral vectors.

[0091] In another preferred embodiment, the viral vector is selected from the group consisting of adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or combinations thereof.

[0092] In another preferred embodiment, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, and a multifunctional vector.

[0093] An eighth aspect of the present invention provides a CRISPR-Cas composition comprising:

[0094] (i) The first component is selected from the group consisting of: the gene-editing protein of the first aspect of the present invention, the fusion protein of the second aspect of the present invention, the nucleotide sequence encoding the gene-editing protein of the first aspect of the present invention or the fusion protein of the second aspect of the present invention, and any combination thereof; and

[0095] (ii) A second component comprising one or more guide RNAs as described in the fifth aspect of the present invention, or a nucleotide sequence encoding one or more guide RNAs as described in the fifth aspect of the present invention;

[0096] The guide RNA is capable of forming a complex with the protein, protein variant, or fusion protein described in (i).

[0097] In another preferred embodiment, the guide RNA comprises a unidirectional repeat sequence and a spacer sequence from the 5' to the 3' direction, the spacer sequence being capable of hybridizing with the target sequence.

[0098] In another preferred embodiment, the same repeat sequence is a nucleic acid molecule as defined in the fourth aspect of the present invention.

[0099] In another preferred embodiment, the composition further includes a pharmaceutically acceptable carrier.

[0100] In another preferred embodiment, the composition comprises a pharmaceutical composition.

[0101] In another preferred embodiment, the dosage form of the composition is selected from the group consisting of lyophilized formulations, liquid formulations, or combinations thereof.

[0102] In another preferred embodiment, the dosage form of the composition is a liquid formulation.

[0103] In another preferred embodiment, the composition is in the form of an injection.

[0104] In another preferred embodiment, the composition is a cell preparation.

[0105] A ninth aspect of the present invention provides a CRISPR-Cas system comprising one or more vectors, said one or more vectors comprising:

[0106] (i) a first nucleic acid, which is a nucleotide sequence encoding the gene-editing protein of the first aspect of the present invention or the fusion protein of the second aspect of the present invention; optionally, the first nucleic acid is operatively linked to a first regulatory element; and

[0107] (ii) a second nucleic acid encoding a nucleotide sequence comprising the guide RNA described in the fifth aspect of the present invention; optionally, the second nucleic acid is operatively linked to a second regulatory element;

[0108] in:

[0109] The first nucleic acid and the second nucleic acid may exist on the same or different vectors;

[0110] The guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

[0111] In another preferred embodiment, the vector includes plasmids and viral vectors.

[0112] In another preferred embodiment, the guide RNA includes a spacer sequence capable of hybridizing with a target sequence; and a direct repeat (DR) sequence linked to the spacer sequence and capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.

[0113] In another preferred embodiment, the guide RNA includes both unmodified and modified guide RNAs.

[0114] In another preferred embodiment, the modified guide RNA includes chemical modifications of the bases.

[0115] In another preferred embodiment, the chemical modification includes methylation, methoxylation, fluorination, or thiolation.

[0116] In another preferred embodiment, the same repeat sequence is a nucleic acid molecule as defined in the fourth aspect of the present invention.

[0117] In another preferred embodiment, the first regulating element and / or the second regulating element is a promoter, such as an induced promoter.

[0118] In another preferred embodiment, at least one component of the composition is non-natural or modified.

[0119] In another preferred embodiment, the spacer sequence is connected to the 3' end of the direct repeat (DR) sequence.

[0120] In another preferred embodiment, the spacer sequence comprises a complementary sequence to the target sequence.

[0121] In another preferred embodiment, when the target sequence is DNA, the target sequence is located at the 3' end of the adjacent motif (PAM) of the original spacer sequence, and the PAM has a 5'-PAM sequence as shown in TTTN, where N is A, T, C or G.

[0122] In another preferred embodiment, the target sequence is DNA from prokaryotic or eukaryotic cells or a DNA sequence formed by reverse transcription of RNA; or, the target sequence is non-naturally occurring DNA or a DNA sequence formed by reverse transcription of RNA.

[0123] In another preferred embodiment, the target sequence comprises a cDNA sequence.

[0124] In another preferred embodiment, the target sequence includes single-stranded DNA and double-stranded DNA sequences.

[0125] In another preferred embodiment, the target sequence is present within the cell.

[0126] In another preferred embodiment, the target sequence is present in the cell nucleus or in the cytoplasm (e.g., organelles).

[0127] In another preferred embodiment, the cell is a eukaryotic cell.

[0128] In another preferred embodiment, the cell is a prokaryotic cell.

[0129] In another preferred embodiment, the target sequence is located outside the cell.

[0130] In another preferred embodiment, the gene-editing protein of the first aspect of the invention is linked to one or more NLS sequences, or the fusion protein comprises one or more NLS sequences.

[0131] In another preferred embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the gene-editing protein described in the first aspect of the invention.

[0132] In another preferred embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the gene-editing protein described in the first aspect of the invention.

[0133] The tenth aspect of the present invention provides a kit comprising one or more components selected from the following: the gene editing protein of the first aspect of the present invention, the fusion protein of the second aspect of the present invention, the polynucleotide of the third aspect of the present invention, the complex of the sixth aspect of the present invention, the vector of the seventh aspect of the present invention, the CRISPR-Cas composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention.

[0134] In another preferred embodiment, the kit also includes a label or instructions.

[0135] In another preferred embodiment, the kit is used for gene or genome editing, disease treatment, targeting a gene, cutting a target gene or a non-target gene, or one or more of these.

[0136] The eleventh aspect of the present invention provides a delivery composition comprising a delivery vector and one or more of the following: the gene editing protein of the first aspect of the present invention, the fusion protein of the second aspect of the present invention, the polynucleotide of the third aspect of the present invention, the complex of the sixth aspect of the present invention, the vector of the seventh aspect of the present invention, the CRISPR-Cas composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention.

[0137] In another preferred embodiment, the delivery carrier is a particle.

[0138] In another preferred embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0139] The twelfth aspect of the present invention provides a host cell comprising the gene editing protein of the first aspect of the present invention, the fusion protein of the second aspect of the present invention, the polynucleotide of the third aspect of the present invention, the nucleic acid molecule of the fourth aspect of the present invention, the complex of the sixth aspect of the present invention, the vector of the seventh aspect of the present invention, the CRISPR-Cas composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, or the delivery composition of the eleventh aspect of the present invention.

[0140] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, plant cell, or mammalian cell (including human and non-human mammals).

[0141] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0142] In another preferred embodiment, the yeast cells are selected from one or more sources of yeast from the group consisting of: Pichia pastoris, Kluyveromyces, or combinations thereof; preferably, the yeast cells include: Kluyveromyces, more preferably Kluyveromyces marxi, and / or Kluyveromyces lactis.

[0143] In another preferred embodiment, the host cell is selected from the group consisting of: Escherichia coli, wheat germ cells, insect cells, SF9, HeLa, HEK293, CHO, yeast cells, or combinations thereof.

[0144] The thirteenth aspect of the present invention provides an enzyme preparation comprising the gene editing protein of the first aspect of the present invention, the fusion protein of the second aspect of the present invention, the complex of the sixth aspect of the present invention, the CRISPR-Cas composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, or the delivery composition of the eleventh aspect of the present invention.

[0145] In another preferred embodiment, the enzyme preparation includes an injection and / or a lyophilized preparation.

[0146] The fourteenth aspect of the present invention provides a medicine box, comprising:

[0147] A first container, and a compound of the sixth aspect of the present invention, a composition of the eighth aspect of the present invention, or a system of the ninth aspect of the present invention, or a drug containing the compound of the sixth aspect of the present invention, the composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, located in the first container.

[0148] In another preferred embodiment, the drug in the first container is a single-component formulation containing the complex described in the sixth aspect of the present invention, the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention.

[0149] In another preferred embodiment, the dosage form of the drug is selected from the group consisting of lyophilized preparations, liquid preparations, or combinations thereof.

[0150] In another preferred embodiment, the dosage form of the drug is an oral dosage form or an injectable dosage form.

[0151] In another preferred embodiment, the medicine box also includes an instruction manual.

[0152] The fifteenth aspect of the present invention provides a medicine box, comprising:

[0153] (a1) A first container, and a gene-editing protein of the first aspect of the present invention, or a fusion protein of the second aspect of the present invention, or its encoding gene or its expression vector, or a drug containing the gene-editing protein of the first aspect of the present invention, or a fusion protein of the second aspect of the present invention, or its encoding gene or its expression vector, located in the first container;

[0154] (b1) An optional second container, and the guide RNA or its expression vector according to the fifth aspect of the present invention located in the second container, or a drug containing the guide RNA or its expression vector according to the fifth aspect of the present invention.

[0155] In another preferred embodiment, the first container and the second container are different containers.

[0156] In another preferred embodiment, the drug in the first container is a single-component formulation containing the gene-editing protein described in the first aspect of the present invention, or the fusion protein described in the second aspect of the present invention, or its encoding gene or expression vector.

[0157] In another preferred embodiment, the drug in the second container is a single-ingredient formulation containing the guide RNA or its expression vector as described in the fifth aspect of the present invention.

[0158] In another preferred embodiment, the dosage form of the drug is selected from the group consisting of lyophilized preparations, liquid preparations, or combinations thereof.

[0159] In another preferred embodiment, the dosage form of the drug is an oral dosage form or an injectable dosage form.

[0160] In another preferred embodiment, the medicine box also includes an instruction manual.

[0161] The sixteenth aspect of the present invention provides a method for targeting and editing or cutting a target gene, comprising: contacting the gene editing protein of the first aspect of the present invention, or the fusion protein of the second aspect of the present invention, or the complex of the sixth aspect of the present invention, or the composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, or the delivery composition of the eleventh aspect of the present invention, or the enzyme preparation of the thirteenth aspect of the present invention, or the cassette of the fourteenth or fifteenth aspect of the present invention with the target gene, or delivering it to a cell containing the target gene, wherein the target sequence is present in the target gene.

[0162] In another preferred embodiment, the target gene is present within the cell.

[0163] In another preferred embodiment, the cell is a prokaryotic cell.

[0164] In another preferred embodiment, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.

[0165] In another preferred embodiment, the target gene is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0166] In another preferred embodiment, the editing or cutting of the target gene includes breaking the target sequence, such as a double-strand break in DNA or a single-strand break in RNA, or inserting a foreign nucleic acid into the break.

[0167] In another preferred embodiment, the target gene comprises DNA.

[0168] In another preferred embodiment, the DNA includes single-stranded DNA and double-stranded DNA.

[0169] The seventeenth aspect of the present invention provides a method for inducing changes in cell state, the method comprising contacting a gene-editing protein of the first aspect of the present invention, or a fusion protein of the second aspect of the present invention, or a complex of the sixth aspect of the present invention, or a composition of the eighth aspect of the present invention, or a system of the ninth aspect of the present invention, or a delivery composition of the eleventh aspect of the present invention, or an enzyme preparation of the thirteenth aspect of the present invention, or a kit of the fourteenth or fifteenth aspect of the present invention with a target gene in a cell.

[0170] The eighteenth aspect of the present invention provides a method for altering the expression of a gene product, comprising: contacting a gene-editing protein according to the first aspect of the present invention, or a fusion protein according to the second aspect of the present invention, or a complex according to the sixth aspect of the present invention, or a composition according to the eighth aspect of the present invention, or a system according to the ninth aspect of the present invention, or a delivery composition according to the eleventh aspect of the present invention, or an enzyme preparation according to the thirteenth aspect of the present invention, or a cassette according to the fourteenth or fifteenth aspect of the present invention with a nucleic acid molecule encoding the gene product, or delivering it to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0171] In another preferred embodiment, the nucleic acid molecule is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0172] In another preferred embodiment, the expression of the gene product is altered (e.g., enhanced or reduced).

[0173] In another preferred embodiment, the gene product is a protein.

[0174] In another preferred embodiment, the protein, fusion protein, polynucleotide, isolated nucleic acid molecule, complex, carrier, or composition is contained in a delivery carrier.

[0175] In another preferred embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0176] In another preferred embodiment, one or more target sequences in a nucleic acid molecule encoding a target gene or a target gene product are used to modify a cell, cell line, or organism.

[0177] The nineteenth aspect of the present invention provides a cell or its progeny obtained by the method of any one of the sixteenth to eighteenth aspects of the present invention, wherein the cell contains modifications not present in its wild type.

[0178] The twentieth aspect of the present invention provides cell products of the cells or their progeny as described in the nineteenth aspect of the present invention.

[0179] The twenty-first aspect of the present invention provides an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the gene editing protein of the first aspect of the present invention, or the fusion protein of the second aspect of the present invention, or the polynucleotide of the third aspect of the present invention, or the complex of the sixth aspect of the present invention, or the vector of the seventh aspect of the present invention, or the composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, or the delivery composition of the eleventh aspect of the present invention.

[0180] In another preferred embodiment, the cell is a prokaryotic cell.

[0181] In another preferred embodiment, the cell is a eukaryotic cell, such as a mammalian cell (e.g., a human cell) or a plant cell.

[0182] In another preferred embodiment, the cell is a stem cell or a stem cell line.

[0183] The twenty-second aspect of the present invention provides the use of the gene-editing protein of the first aspect of the present invention, or the fusion protein of the second aspect of the present invention, or the polynucleotide of the third aspect of the present invention, or the nucleic acid molecule of the fourth aspect of the present invention, or the complex of the sixth aspect of the present invention, or the vector of the seventh aspect of the present invention, or the composition of the eighth aspect of the present invention, or the system of the ninth aspect of the present invention, or the kit of the tenth aspect of the present invention, or the delivery composition of the eleventh aspect of the present invention, or the enzyme preparation of the thirteenth aspect of the present invention, or the cassette of the fourteenth or fifteenth aspect of the present invention, for the preparation of a drug or preparation for nucleic acid editing (e.g., gene or genome editing).

[0184] In another preferred embodiment, the gene or genome editing includes modifying a gene, knocking out a gene, altering the expression of a gene product, repairing mutations, and / or inserting polynucleotides.

[0185] The twenty-third aspect of this invention provides the use of the gene-editing protein described in the first aspect of this invention, or the fusion protein described in the second aspect of this invention, or the polynucleotide described in the third aspect of this invention, or the complex described in the sixth aspect of this invention, or the vector described in the seventh aspect of this invention, or the composition described in the eighth aspect of this invention, or the system described in the ninth aspect of this invention, or the kit described in the tenth aspect of this invention, or the delivery composition described in the eleventh aspect of this invention, or the enzyme preparation described in the thirteenth aspect of this invention, or the cassette described in the fourteenth or fifteenth aspect of this invention, for the preparation of a drug or formulation, said drug or formulation for use in one or more of the following groups:

[0186] (i) In vitro gene or genome editing;

[0187] (ii) Detection of isolated single-stranded DNA;

[0188] (iii) Editing target sequences in target loci to modify biological or non-human organisms;

[0189] (iv) Treating conditions caused by defects in target sequences at target loci;

[0190] (v) Treat the symptoms or diseases of the subject in need.

[0191] In another preferred embodiment, the condition or disease includes cancer, infectious diseases, neurological diseases, eye diseases, and hearing diseases.

[0192] In another preferred embodiment, the disease or condition includes cystic fibrosis, progressive pseudohypertrophic muscular dystrophy (Duchenne muscular dystrophy, DMD), Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, Fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, β-thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia, transthyretin amyloidosis, and retinal diseases. Diseases, macular degeneration, Wilms' tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, bile duct cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid carcinoma, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, non-Hodgkin's lymphoma, and urobladder cancer.

[0193] In another preferred embodiment, the symptom or disease is caused by a pathogenic point mutation.

[0194] The twenty-fourth aspect of the present invention provides a method for detecting the presence of target nucleic acid molecules in a sample, the method comprising contacting the sample with a gene-editing protein as described in the first aspect of the present invention, or a fusion protein as described in the second aspect of the present invention, or a complex as described in the sixth aspect of the present invention, or a composition as described in the eighth aspect of the present invention, or a system as described in the ninth aspect of the present invention, or a kit as described in the tenth aspect of the present invention, or a delivery composition as described in the eleventh aspect of the present invention, or an enzyme preparation as described in the thirteenth aspect of the present invention, and contacting a non-target sequence, detecting a detectable signal generated by the cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with guide RNA.

[0195] In another preferred embodiment, if the non-target sequence is cleaved by a protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates that a target nucleic acid molecule is present in the sample; if the non-target sequence is not cleaved by a protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates that a target nucleic acid molecule is not present in the sample.

[0196] In another preferred embodiment, the target nucleic acid molecule is target DNA.

[0197] In another preferred embodiment, the target DNA includes DNA formed based on RNA reverse transcription.

[0198] In another preferred embodiment, the target DNA includes cDNA.

[0199] In another preferred embodiment, the target DNA is selected from the group consisting of single-stranded DNA, double-stranded DNA, or combinations thereof.

[0200] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0201] Figure 1 The image shows an agarose gel electrophoresis diagram of the EcoRI and NotI double-digested CasW1 insert and pET-28a(+) vector.

[0202] Figure 2 The image shows an SDS-PAGE image of the newly purified CasW1 protein stained with Coomassie Brilliant Blue.

[0203] Figure 3 The image shows an agarose gel electrophoresis diagram of the in vitro preparation of dsDNA template.

[0204] Figure 4 The capillary electrophoresis diagrams showing the in vitro cleavage effect of CasW1 are presented. Compared with the in vitro cleavage effect of Cpf1, it can be seen that CasW1 can cleave a 450bp dsDNA template into two dsDNA fragments, with a cleavage activity close to 100%.

[0205] Figure 5 Capillary electrophoresis images of CasW1 in vitro cleavage effect are shown, with Cas12i.16 and S7R-Cas12i3 (M2869) in the prior art as controls.

[0206] Figure 6 The plasmid map of pET-28a(+)-CasW1 is shown.

[0207] Figure 7 The plasmid map of pET-28a(+)-LbCpf1 is shown.

[0208] Figure 8 The plasmid map of pET-28a(+)-HED Cas12i.16 is shown.

[0209] Figure 9 The plasmid map of pET-28a(+)-S7R Cas12i.3 is shown. Detailed Implementation

[0210] Through extensive and in-depth research, the inventors unexpectedly discovered a novel gene-editing protein. This gene-editing protein exhibits excellent gene-editing activity, effectively editing or cutting target genes, and can effectively treat the symptoms or diseases of subjects in need. Based on this, the inventors completed this invention.

[0211] the term

[0212] The following embodiments are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the embodiments are generally performed in accordance with conventional methods well known in the art and described in various references.

[0213] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0214] To facilitate a clearer understanding of this disclosure, certain terms are first defined. As used herein, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.

[0215] The term “about” can refer to a value or composition within an acceptable margin of error for a particular value or composition as determined by a person skilled in the art, depending in part on how the value or composition is measured or determined. For example, as used herein, the expression “about 100” includes all values ​​between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0216] As used herein, the terms “containing” or “including (comprise)” can be open-ended, semi-closed, or closed. In other words, the terms also include “consistently made of” or “composed of”.

[0217] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which may be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of a reference nucleotide sequence or protein) and determining the number of positions where identical residues occur. This is typically expressed as a percentage. The measurement of sequence identity of nucleotide sequences is a method well known to those skilled in the art.

[0218] Gene editing protein

[0219] In this invention, the gene-editing protein is an effector protein in the CRISPR / Cas system.

[0220] In this invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably. Cas protein is used in the broadest sense, including wild-type Cas protein, its derivatives or variants, analogs, and its functional fragments such as oligonucleotide binding fragments.

[0221] The term “wild type” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, gene, or protein, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, which can be isolated from its natural source and has not been intentionally modified by humans.

[0222] The terms “variant,” “derivative,” and “analyte” refer to polypeptides that substantially retain the function or activity of the Cas protein of the present invention.

[0223] Generally, protein derivatization does not adversely affect the protein's desired activity (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA); that is, the protein derivative has the same activity as the original protein. A modified form of "derivative" includes one or more amino acids of the protein that may be deleted, inserted, modified, and / or substituted. The terms "non-natural" or "engineered" are used interchangeably and indicate artificial involvement.

[0224] In one aspect, the present invention provides a gene-editing protein (Cas protein) comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO.1, and substantially retaining the biological function of the sequence from which it is derived;

[0225] In one embodiment, the amino acid sequence of the Cas protein has one or more amino acid substitutions, deletions, or additions compared to the amino acid sequence of SEQ ID NO.1, and substantially retains the biological function of its derived sequence.

[0226] In one embodiment, the Cas protein comprises the amino acid sequence shown in SEQ ID NO.1;

[0227] Or a sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID NO.1; or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with the amino acid sequence shown in SEQ ID NO.1;

[0228] In one embodiment, the Cas protein has the amino acid sequence shown in SEQ ID NO.1.

[0229] Those skilled in the art will understand that the structure of a protein can be altered without adversely affecting its activity and functionality, for example, by introducing one or more conserved amino acid substitutions into the protein's amino acid sequence without adversely affecting the protein molecule's activity and / or three-dimensional structure.

[0230] Those skilled in the art will recognize examples and implementations of conserved amino acid substitutions. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, i.e., replacing another nonpolar amino acid residue with a nonpolar amino acid residue, replacing another polar uncharged amino acid residue with a polar uncharged amino acid residue, replacing another basic amino acid residue with a basic amino acid residue, and replacing another acidic amino acid residue with an acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this invention, provided that the substitution does not lead to the inactivation of the protein's biological activity. Therefore, the proteins of this invention can contain one or more conserved substitutions in their amino acid sequence, preferably generated by substitutions according to Table A. Furthermore, this invention also covers proteins that also contain one or more other nonconservative substitutions, provided that such nonconservative substitutions do not significantly affect the desired function and biological activity of the proteins of this invention.

[0231] Conserved amino acid substitutions can occur at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing biological activity, while “essential” amino acid residues are required for biological activity. A “conserved amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can occur in non-conserved regions of Cas enzymes. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.

[0232] Table A

[0233]

[0234]

[0235] Those skilled in the art will recognize that one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C-terminus of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from the N and / or C-terminus of the Cas protein of this invention while retaining their desired functional activity are also within the scope of this invention. These alterations may include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.

[0236] It should be recognized that proteins can be altered in a variety of ways, including amino acid substitution, deletion, truncation, and insertion, and the methods used for such operations are generally known to those skilled in the art.

[0237] For example, amino acid sequence variants of the Cas protein can be prepared by mutating the DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, by using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform one or more amino acid substitutions; or one or more amino acid deletions and / or one or more amino acid insertions.

[0238] Those skilled in the art will understand that these minute amino acid changes in the Cas protein of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.

[0239] Those skilled in the art can identify the essential amino acids of Cas proteins using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of the protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.

[0240] Orthologue (ortholog)

[0241] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein, as described herein, refers to a protein belonging to a different species that performs the same or similar function as the protein that is its orthologue.

[0242] The nucleic acid cleavage disclosed herein includes: DNA or RNA breaks in target nucleic acids generated by the Cas protein (Cis cleavage), and DNA or RNA breaks in side-branched nucleic acid substrates (single-stranded nucleic acid substrates) caused by the paracleavage activity of the Cas protein (i.e., non-specific or non-targeted, trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0243] Trans-cleavage refers to the phenomenon where, under certain conditions, activated Cas12 family proteins remain active after binding to a target sequence and continue to nonspecifically cleave non-target oligonucleotides. This para-cleavage activity enables the detection of specific target oligonucleotides using Cas systems. For example, the Cas12i system can be engineered to nonspecifically cleave ssDNA or transcripts. Para-cleavage activity has been used in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, which can be used in many clinical diagnostics (Gootenberg, JS et al., Nucleic acid detection with CRISPR-Cas13a / C2c2. Science 356, 438-442 (2017)).

[0244] Fusion protein

[0245] In one aspect, the present invention provides a fusion protein comprising the Cas protein described in any of the preceding claims and one or more functional domains.

[0246] In one embodiment, the functional domain includes one or more of the following: localization signal, reporter protein, Cas protein targeting portion, DNA binding domain, epitope tag, transcription activation domain, transcription repression domain, nuclease, deamination domain, methyltransferase, demethylase, transcription release factor, HDAC, cleavage active peptide, and ligase.

[0247] In one embodiment, "methyltransferase" is exemplarily, such as HhaI DNAm5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3, ZMET2, CMT1, CMT2, etc.

[0248] Demethylases are enzymes that remove methyl (CH3-) groups from nucleic acids, proteins (e.g., histones), and other molecules. Demethylases play a crucial role in epigenetic modification mechanisms. Demethylase proteins alter transcriptional regulation of the genome by controlling the level of methylation occurring on DNA and histones, and consequently regulate the chromatin state at specific loci in organisms, such as TET1 (ten-eleven translocation 1), ten-eleven translocation (TET) dioxygenase 1 (TET1CD), DME, DML1, DML2, and ROS1.

[0249] In another preferred embodiment, the transcriptional releasing factor, exemplarily, is eukaryotic releasing factor 1 (ERF1) activity or eukaryotic releasing factor 3 (ERF3).

[0250] In one embodiment, the functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain.

[0251] In one embodiment, the positioning signal includes a nuclear positioning signal and / or a nuclear output signal;

[0252] Preferably, the nuclear output signal includes human protein tyrosine kinase 2;

[0253] Preferably, the reporter protein includes one or more of glutathione S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or autofluorescent protein;

[0254] Preferably, the autofluorescent protein includes one or more of green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, or blue fluorescent protein;

[0255] Preferably, the DNA binding domain includes one or more of methylation-binding proteins, LexADBD, or Gal4DBD;

[0256] Preferably, the epitope tag includes one or more of the following: histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, or thioredoxin tag;

[0257] Preferably, the transcriptional activation domain includes VP64 and / or VPR;

[0258] Preferably, the transcriptional repression domain includes KRAB and / or SID;

[0259] Preferably, the nuclease comprises FokI;

[0260] Preferably, the deammoniation domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD;

[0261] Preferably, the cleavage-active polypeptide includes a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity, or a polypeptide with double-stranded DNA cleavage activity.

[0262] Preferably, the ligase includes DNA ligase and / or RNA ligase.

[0263] In one implementation, the functional domain is the full length or a functional segment of TadA8e.

[0264] Polynucleotides

[0265] In one aspect, the present invention provides a polynucleotide that is a polynucleotide sequence encoding the gene-editing protein (Cas protein) or a polynucleotide sequence encoding the aforementioned fusion protein.

[0266] In one embodiment, the polynucleotide (DNA molecule) comprises nucleotides having 70% or more, preferably 90% or more, more preferably 95% or more, further preferably 99%, and even more preferably 100% identity with the nucleotide sequence described in SEQ ID NO.2.

[0267] In one embodiment, the polynucleotide is a DNA molecule codon-optimized according to the codon preference of the host cell.

[0268] The optimizations described in this disclosure may require mutations in the nucleotide sequence of the encoded protein (e.g., the Cas protein of this disclosure) to mimic the codon preferences of a intended host organism or cell that simultaneously encodes the same protein. Therefore, the codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, a human codon-optimized nucleotide sequence encoding the protein can be used. As another non-limiting embodiment, if the intended host cell is an animal cell (e.g., mouse cell, insect cell), an animal codon-optimized nucleotide sequence encoding the protein can be generated. As another non-limiting embodiment, if the intended host cell is a plant cell, a plant codon-optimized nucleotide sequence encoding the protein can be generated.

[0269] Lists of codon choices are readily available, for example, in the "Codon Usage Database" at www.kazusa.or.jp / codon. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasY7 or its variants or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in eukaryotic cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasW1 or its variants or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in animal cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasW1 or its variants or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in fungal cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasW1 or its variants or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in plant cells.

[0270] In one implementation, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0271] CRISPR system

[0272] The terms “regularly clustered short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements relating to the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes.

[0273] CRISPR-Cas Composition

[0274] In one aspect, the present invention also provides a CRISPR-Cas composition comprising:

[0275] (1) Protein components: the aforementioned gene-editing protein (Cas protein), or the aforementioned fusion protein; or nucleic acid molecules encoding the aforementioned gene-editing protein (Cas protein) or the aforementioned fusion protein;

[0276] (2) RNA component: guide RNA, or one or more nucleic acids encoding the guide RNA, or precursor RNA of the guide RNA, or nucleic acid encoding the precursor RNA of the guide RNA;

[0277] The protein components and nucleic acid components combine to form a complex.

[0278] In one embodiment, the composition is an activated CRISPR complex, the activated CRISPR complex further comprising a target sequence of a target nucleic acid bound to the guide RNA.

[0279] In one embodiment, the CRISPR-Cas composition includes one or more carriers, said one or more carriers comprising:

[0280] (1) A first regulatory element, operatively linked to a nucleotide sequence encoding the gene-editing protein (Cas protein) or a nucleotide sequence encoding the fusion protein; and

[0281] (2) A second regulatory element, operatively linked to a nucleotide sequence encoding the guide RNA, the guide RNA comprising:

[0282] (a) Spacer sequences capable of hybridizing with the target sequence of the target nucleic acid, and

[0283] (b) A direct repeat (DR) sequence attached to the spacer sequence that guides the gene editing protein (Cas protein) to bind to the guide RNA to form a CRISPR-Cas complex targeting the target sequence;

[0284] The first control element and the second control element are located on the same or different carriers of the CRISPR-Cas carrier system.

[0285] In one embodiment, the first or second regulatory element includes a promoter, which includes one or more of an inductive promoter, a constitutive promoter, or a tissue-specific promoter;

[0286] In one embodiment, the promoter includes one or more of T7, SP6, T3, CMV, EF1a, SV40, PGK1, humanβ-actin, CAG, U6, H1, T7, T7lac, araBAD, trp, lac, or Ptac;

[0287] In one embodiment, the first control element and the second control element are located on the same or different carriers.

[0288] In one embodiment, the vector includes a retroviral vector, a lentiviral vector, an adenovirus vector, an adeno-associated virus vector, a herpes simplex vector, or a phage particle vector.

[0289] In one embodiment, the vector includes a plasmid vector.

[0290] In one embodiment, the target nucleic acid includes DNA derived from eukaryotes or DNA derived from prokaryotes;

[0291] In one embodiment, the eukaryotes include animals or plants;

[0292] In one embodiment, the target nucleic acid includes non-human mammal DNA, human DNA, insect DNA, bird DNA, reptile DNA, amphibian DNA, rodent DNA, fish DNA, worm DNA, nematode DNA, or yeast DNA.

[0293] In one embodiment, the non-human mammalian DNA includes non-human primate DNA.

[0294] CRISPR / Cas complex

[0295] The term "CRISPR / Cas complex" refers to a complex formed by the binding of a guide RNA, gRNA (guide RNA), or mature crRNA (or directing RNA) to a gene-editing protein (Cas protein). This complex contains a guide sequence that hybridizes to the target sequence and binds to the gene-editing protein (Cas protein). The complex can recognize and cleave target nucleotides that hybridize with the guide RNA or mature crRNA.

[0296] Guide RNA (gRNA)

[0297] The terms “guide RNA (gRNA),” “mature crRNA,” “crRNA,” “guide sequence,” and “guide RNA” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, guide RNA may comprise a direct repeat (DR) sequence and a spacer sequence, or consist essentially of or composed of a direct repeat (DR) sequence and a spacer sequence.

[0298] In some cases, the spacer sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR-Cas complex to said target sequence. In one embodiment, the complementarity between the spacer sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. The guide sequence comprises a sequence (e.g., a direct repeat (DR) sequence) that is sufficiently complementary to the target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide the sequence-specific binding of the complex to the target nucleic acid sequence.

[0299] As is known in the art, complete complementarity is not required to function effectively, provided there is sufficient complementarity. Therefore, when necessary, cleavage efficiency can be modulated by introducing mismatches (e.g., one or more mismatches between the spacer sequence and the target nucleic acid, such as mismatches of 1 or 2 nucleotides (including mismatches along the spacer / target sequence)). For example, if a cleavage rate of less than 100% of the target is desired (e.g., in a cell population), one or two mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.

[0300] In one aspect, the present invention provides a guide RNA comprising a direct repeat (DR) sequence capable of binding the Cas protein and a spacer sequence capable of targeting a target sequence.

[0301] In one embodiment, the direct repeat (DR) sequence comprises the sequence shown in SEQ ID NO.3.

[0302] In one embodiment, the 3' end of the same-direction repeat sequence includes a stem-loop structure, and further includes a stem formed by the hybridization of a first stem nucleotide chain and a second stem nucleotide chain, wherein the loop nucleotide chain forms the loop of the stem-loop structure;

[0303] In one embodiment, the repetitive sequence comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence described in SEQ ID NO.3;

[0304] In one embodiment, the repetitive sequence comprises a nucleotide sequence having at least 85%, more preferably 90%, and even more preferably 95% identity with the nucleotide sequence described in SEQ ID NO.3;

[0305] In one embodiment, the same-direction repeat sequence comprises the nucleotide sequence described in SEQ ID NO.3.

[0306] In one embodiment, more than 80% of the spacer sequence is complementary to the target nucleic acid;

[0307] In one embodiment, more than 90%, more than 95%, more preferably more than 99%, and even more preferably 100% of the spacer sequence is complementary to the target nucleic acid;

[0308] In one embodiment, the length of the spacer sequence is 18-41 nt, for example 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 nt, more preferably 18 to 27 nucleotides, more preferably 18 to 24 nucleotides, and most preferably 18 to 22 nucleotides.

[0309] In one implementation, the spacer sequence is 20 nt in length.

[0310] target nucleic acid

[0311] In this invention, the terms "target nucleic acid" and "target sequence" or "target nucleic acid sequence" or "target nucleic acid molecule" are used interchangeably to refer to a specific nucleic acid containing a nucleic acid sequence that is wholly or partially complementary to the spacer sequence in the guide RNA. "Target sequence" refers to a polynucleotide targeted by the spacer sequence in the guide RNA, such as a sequence complementary to that spacer sequence, wherein hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA). Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR-Cas complex. In some embodiments, the target nucleic acid contains a non-coding region (e.g., a promoter or terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded.

[0312] The target sequence can contain any polynucleotide, such as DNA. In some cases, the target sequence is located inside or outside the cell. In other cases, the target sequence is located in the cell nucleus, cytoplasm, or organelles (such as mitochondria or chloroplasts).

[0313] The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).

[0314] Donor template

[0315] In this invention, the donor template nucleic acid or the donor template can be used interchangeably, meaning that after the gene editing protein (Cas protein) described herein alters the target nucleic acid, one or more cellular proteins can use it to change the structure of the target nucleic acid.

[0316] In some embodiments, the donor template nucleic acid is a double-stranded or single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template nucleic acid is a foreign nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination, specifically homologous recombination, can be achieved using the donor template.

[0317] Cutting

[0318] A cut refers to a break in the DNA of a target nucleic acid produced by the gene-editing protein (Cas protein) described herein. In some embodiments, the cut is a double-stranded DNA break. In some embodiments, the cut is a single-stranded DNA break.

[0319] In this invention, the meanings of cleaving target nucleic acids or modifying target nucleic acids can overlap. Modifying target nucleic acids includes not only the modification of single nucleotides, but also the insertion or deletion of nucleic acid fragments.

[0320] Report nucleic acid

[0321] A reporter nucleic acid is a molecule that can be cleaved or otherwise inactivated by an activated CRISPR system protein as described herein. A reporter nucleic acid comprises a nucleic acid element that can be cleaved by a CRISPR protein (e.g., a single-stranded, non-targeting nucleic acid molecule with distinct reporter groups or labeled molecules at both ends). Cleavage of the nucleic acid element produces a detectable signal. Prior to cleavage, or while the reporter nucleic acid is in an “active” state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that in some example embodiments, minimal background signal may be generated in the presence of an active reporter nucleic acid. A positive detectable signal can be any signal detectable using optical, fluorescent, chemiluminescent, electrochemical, or other detection methods known in the art. For example, in some embodiments, a first signal (i.e., a negative detectable signal) may be detected in the presence of a reporter nucleic acid, and then converted to a second signal (e.g., a positive detectable signal) upon detection of a target molecule and upon cleavage or inactivation by an activated CRISPR protein. The reporter nucleic acid can be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.

[0322] The detection method described in this invention can be used for the quantitative detection of target nucleic acids. The quantitative detection index can be determined based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group or the width of the colored band.

[0323] Functional structural domain

[0324] In this article, the term "functional domain" is used in its broadest sense, encompassing proteins such as enzymes or factors themselves, or fragments / domains with specific functions. Gene-editing proteins (e.g., dCas proteins) are linked / associated with one or more functional domains, selected from one or more of the following: localization signals, reporter proteins, Cas protein targeting portions, DNA-binding domains, epitope tags, transcriptional activation domains, transcriptional repression domains, nucleases, deamination domains, methyltransferases, demethylases, transcription release factors, HDACs, cleavage active peptides, and ligases. When more than one functional domain is included, the functional domains may be the same or different.

[0325] Deamination domain

[0326] In this invention, the deamination domain includes a deaminase (e.g., adenosine deaminase or cytidine deaminase) catalytic domain. As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyze the hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) into hypoxanthine (or the hypoxanthine portion of a molecule).

[0327] In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0328] Adenosine deaminases include, but are not limited to, enzyme family members called RNA-acting adenosine deaminases (ADAR), enzyme family members called tRNA-acting adenosine deaminases (ADAT), and other family members containing an adenosine deaminase domain (ADAD). According to this disclosure, adenosine deaminases are capable of targeting adenine in RNA / DNA and RNA duplexes. In certain embodiments, adenosine deaminases have been modified to increase their ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0329] In some embodiments, the deaminase is a cytidine deaminase. The term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyzes a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0330] Cytidine deaminases include, but are not limited to, members of an enzyme family known as the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases, activation-induced deaminase (AID), or cytidine deaminase 1 (CDA1). In certain embodiments, APOBEC family deaminases are included.

[0331] In some embodiments, the cytidine deaminase comprises the wild-type amino acid sequence of cytosine deaminase. In some embodiments, the cytidine deaminase contains one or more mutations in the cytosine deaminase sequence, such that the editing efficiency and / or substrate editing preference of the cytosine deaminase are altered according to specific needs.

[0332] identity

[0333] "Identity" refers to the sequence matching between two polypeptides or two nucleic acids. "Identity" represents the percentage of identical residues in the polypeptide or nucleic acid sequence out of the total number of residues, and is calculated based on mutation type. Mutation types include insertions (extensions) at either end of a sequence, deletions (truncations) at either end of a sequence, substitutions of one or more amino acids / nucleotides, insertions within a sequence, and deletions within a sequence.

[0334] For example, with polypeptide sequences, if the mutation type is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total residue count is calculated based on the larger of the compared molecules. If the mutation type also includes insertions (extensions) or deletions (truncations) at either end of the sequence, the number of amino acids inserted or deleted at either end (e.g., less than 20 at either end) is not included in the total residue count. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between sequences, and gaps in the alignment (if present) are resolved using a specific algorithm. Nucleotide identity is calculated similarly.

[0335] carrier

[0336] The carrier is a nucleic acid molecule that can transport another nucleic acid molecule that is linked to it.

[0337] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. Vectors can be introduced into host cells through transformation, transduction, or transfection, thereby enabling the expression of their carried genetic material elements in the host cells. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Vectors may also contain a replication initiation site.

[0338] Vectors include plasmids and viral vectors. A plasmid is a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. A viral vector contains a virus-derived DNA or RNA sequence within a vector used to package the virus; viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by a virus intended for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.

[0339] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Furthermore, some vectors can direct the expression of genes they are operatively linked to. Such vectors are called "expression vectors."

[0340] In some embodiments, the vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or plasmid) can be delivered to the target tissue via, for example, intramuscular injection, intravenous administration, percutaneous administration, intranasal administration, oral administration, or mucosal administration. The delivery can be performed via a single dose or multiple doses. Those skilled in the art will understand that the actual dose to be delivered herein can vary considerably depending on a variety of factors, including but not limited to the choice of vector, target cells, organism, tissue, general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the manner of administration, and the type of transformation / modification sought.

[0341] Control element

[0342] In this article, "regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences), for detailed description in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In other cases, regulatory elements may also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell type-specific.

[0343] The term "promoter" refers to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of downstream genes. A constitutive promoter is a nucleotide sequence that, when operatively linked to a polynucleotide encoding or defining a gene product, will result in the production of that gene product in the cell under most or all physiological conditions. An inducible promoter is a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of endogenous or exogenous stimuli, such as through a chemical compound (chemical inducer), or in response to environmental, hormone, chemical, and / or developmental signals. Inducible or regulatory promoters include promoters induced or regulated, for example, by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals (such as ethanol, abscisic acid (ABA), jasmonic acid esters, salicylic acid, or safeners).

[0344] host cells

[0345] In this article, "host cell" refers to eukaryotic cells (e.g., animal cells, plant cells, fungal cells, etc.), prokaryotic cells (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or cells derived from multicellular organisms (e.g., cell lines) cultured in the form of single-celled entities, which are used as recipients of nucleic acids (e.g., expression vectors), and includes the offspring of the original cells that have been genetically modified with nucleic acids.

[0346] It should be understood that the offspring of a single cell can be attributed to natural, accidental, or intentional mutations and do not necessarily have the exact same morphology or genome as the original parent cell. A “recombinant host cell” (also known as a “genetically modified host cell”) is a host cell in which a heterologous nucleic acid, such as an expression vector, has been introduced.

[0347] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.

[0348] In another aspect, the present invention also provides a host cell or its progeny comprising the aforementioned gene-editing protein (Cas protein), or the aforementioned fusion protein, or the aforementioned polynucleotide, or the aforementioned vector system, or the aforementioned CRISPR-Cas system, or the aforementioned composition.

[0349] In one embodiment, the host cell includes non-human mammals, humans, insects, birds, reptiles, amphibians, rodents, fish, worms, nematodes, or yeast cells.

[0350] In one aspect, the present invention also provides a multicellular organism comprising the aforementioned cells or their descendants.

[0351] In one embodiment, the multicellular organism is an animal or plant model used for the relevant disease.

[0352] NLS

[0353] NLS stands for “nuclear localization sequence” or “nuclear localization signal,” which is the amino acid sequence that prompts a protein to enter the cell nucleus. Nuclear localization sequences are known in the art (e.g., described in Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001), and are incorporated herein by reference to their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172.

[0354] Operable connection

[0355] "Operationally ligated" refers to the ligation of a target nucleotide sequence to a regulatory element in a manner that allows for the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when a vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can also be selected to target specific cell types.

[0356] Complementary

[0357] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The complementarity percentage represents the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the complementarity percentages are 50%, 60%, 70%, 80%, 90%, and 100%). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under strict conditions.

[0358] The term "strict condition" associated with hybridization refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence.

[0359] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. This complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.

[0360] Hybridization of the target sequence with gRNA indicates that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and gRNA can hybridize to form a complex; or it indicates that at least 12, 15, 16, 17, 18, 19, 20, or more bases of the nucleic acid sequences of the target sequence and gRNA can complement each other to hybridize and form a complex.

[0361] Express

[0362] Nucleic acid expression includes one or more of the following: generating an RNA template from a DNA sequence (e.g., transcription), processing of RNA transcripts (e.g., by splicing, editing, 5′ cap formation and / or 3′ end processing), translating RNA into a polypeptide or protein, or post-translational modifications of a polypeptide or protein.

[0363] deliver

[0364] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as DNA / RNA or RNA / RNA or a combination of protein and RNA. For example, gene-editing proteins (Cas proteins) can be delivered as polynucleotides encoding DNA or RNA, or as proteins.

[0365] In one aspect, the present invention also provides a delivery system comprising the gene-editing protein (Cas protein) or the fusion protein, or the polynucleotide, or the CRISPR-Cas composition.

[0366] In one embodiment, the delivery system further includes a delivery medium, which includes nanoparticles, liposomes, exosomes, microbubbles, gene guns, or electroporation devices.

[0367] Furthermore, when the target is plant cells, delivery methods such as cell-penetrating peptides (CPPs) are employed. For example, in one embodiment, a gene-editing protein (Cas protein) and / or at least one guide RNA is coupled to one or more CPPs, thereby efficiently transporting the CPP coupled with the gene-editing protein (Cas protein) and / or guide RNA into plant cells (e.g., protoplasts). CPPs are short peptides of fewer than 35 amino acids, derived from proteins or chimeric sequences, capable of transporting biomolecules across the cell membrane in a receptor-independent manner. CPPs can be cationic peptides, peptides with hydrophobic sequences, amphiphilic peptides, peptides rich in proline and antimicrobial sequences, and chimeric or dipeptides. CPPs can penetrate biological membranes and thus trigger the transmembrane movement of different biomolecules into the cytoplasm, improving their intracellular pathways and thus promoting biomolecule-target interactions.

[0368] For example, CPP includes Tat (a nuclear transcription activation protein required for viral replication by HIV type 1), penetrin, Kaposi's fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, sweet arrow peptide, etc.

[0369] connector

[0370] In this article, "connector" refers to a chemical group or molecule that connects two molecules or parts, such as the two domains of a fusion protein, or the gene-editing protein (Cas protein) and a deaminase. In some connection methods, the connector is located between or on the flank of two groups, molecules, or other parts and is connected to them by a covalent bond.

[0371] In some embodiments, the linker is a linear polypeptide formed by linking amino acids or multiple amino acid residues together via peptide bonds. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. The length and type of the linker can be designed as needed. In some embodiments, the linker can be selected from artificially synthesized amino acid sequences or naturally occurring polypeptide sequences.

[0372] Detection

[0373] In one aspect, the present invention also provides a method for targeting and editing a target nucleic acid, the method comprising contacting the target nucleic acid with any of the aforementioned CRISPR-Cas systems or compositions.

[0374] In one aspect, the present invention also provides a method for nonspecifically degrading single-stranded DNA after recognizing a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned CRISPR-Cas composition.

[0375] In one aspect, the present invention also provides a method for targeting and creating a nick in a double-stranded target nucleic acid after recognizing a spacer complementary strand of the double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with the aforementioned CRISPR-Cas system or composition.

[0376] In one aspect, the present invention also provides a method for targeting and cleaving a double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with the aforementioned CRISPR-Cas system or composition.

[0377] In one embodiment, the non-spacer sequence complementary strand of the double-stranded target nucleic acid is nicked before the spacer complementary strand of the double-stranded DNA is nicked.

[0378] In one aspect, the present invention also provides a method for specifically editing double-stranded nucleic acids, the method comprising allowing sufficient contact time under adequate conditions,

[0379] (1) the aforementioned gene-editing protein (Cas protein), or fusion protein, another enzyme with sequence-specific nicking activity, and the guide RNA, the guide RNA instructing the gene-editing protein (Cas protein) or the fusion protein to create a nick in the opposite strand relative to the activity of the other sequence-specific nicking enzyme; and (2) the double-stranded nucleic acid; the method resulting in the formation of a double-strand break.

[0380] In one aspect, the present invention also provides a method for editing double-stranded nucleic acids, the method comprising allowing sufficient contact for a sufficient amount of time under adequate conditions:

[0381] (1) the aforementioned gene-editing protein (Cas protein), or fusion protein, and fusion protein of a protein domain having DNA modification activity, and the RNA guide targeting the double-stranded nucleic acid; and (2) the double-stranded nucleic acid;

[0382] The gene-editing protein (Cas protein) of the fusion protein is modified to create a nick in the non-target strand of the double-stranded nucleic acid.

[0383] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at different sites, resulting in staggered cleavage.

[0384] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at the same site, resulting in a flat double-strand break.

[0385] In one aspect, the present invention also provides a method for targeting and cleaving a single-stranded target nucleic acid, the method comprising contacting the target nucleic acid with the CRISPR-Cas composition described in any of the preceding claims.

[0386] In one aspect, the present invention also provides a method for inducing changes in cell state, the method comprising contacting the aforementioned CRISPR-Cas composition with the target nucleic acid in the cell.

[0387] In one embodiment, the cell state includes apoptosis or dormancy;

[0388] In one embodiment, the cells include eukaryotic cells or prokaryotic cells;

[0389] In one embodiment, the cells include mammalian cells or plant disease cells;

[0390] In one embodiment, the cells include cancer cells;

[0391] In one embodiment, the cells include infectious cells or cells infected by an infectious agent;

[0392] In one embodiment, the cells include virus-infected cells and prion-infected cells;

[0393] In one embodiment, the cells include fungal cells, protozoan cells, or parasitic cells.

[0394] In one aspect, the present invention also provides a method for detecting target nucleic acids in a sample, the method comprising contacting the sample with the aforementioned gene-editing protein (Cas protein), guide RNA, and a non-target sequence; detecting a detectable signal generated by the gene-editing protein (Cas protein) cleaving the non-target sequence, thereby detecting the target nucleic acid; wherein the non-target sequence does not hybridize with the guide RNA.

[0395] Reagent test kit

[0396] In one aspect, the present invention provides a kit comprising the aforementioned gene-editing protein (Cas protein), the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned CRISPR-Cas composition, and the aforementioned host cells for use in preparing the kit, wherein the components of the kit are in the same or different containers.

[0397] In one aspect, the present invention also provides a container comprising the aforementioned reagent kit.

[0398] In one embodiment, the container includes a sterile container;

[0399] In one embodiment, the container includes a syringe.

[0400] In some embodiments, the kit also includes instructions for using the kit, such as instructions in more than one language. The kit may also contain one or more reagents for use in the process of utilizing one or more of the components described above. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided prior to use in a form requiring the addition of one or more other components (e.g., in concentrated or lyophilized form); the buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. The buffer may have a suitable pH value, for example, it may be alkaline. In some embodiments, the pH of the buffer is between about 7 and 10.

[0401] treat

[0402] "Treatment" refers to treating or curing a subject's condition, delaying the onset of symptoms, and / or slowing the severity of the condition. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals such as bovines, equines, sheep, suidae, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect). "Plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development.

[0403] In one aspect, the present invention also provides the use of the aforementioned gene-editing protein (Cas protein), the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned CRISPR-Cas composition, and the aforementioned host cell in the preparation of a medicament for treating a condition or disease of a subject in need.

[0404] In one embodiment, the application includes administering the CRISPR-Cas composition to the subject or to ex vivo cells of the subject;

[0405] In one embodiment, the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid associated with the condition or disease, and the Cas protein or the fusion protein cleaves the target nucleic acid;

[0406] In one implementation, the condition or disease includes cancer or an infectious disease;

[0407] In one embodiment, the cancer includes one or more of the following: Wilms' tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid carcinoma, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, non-Hodgkin's lymphoma, or urobladder cancer.

[0408] In one embodiment, the condition or disease includes one or more of the following: cystic fibrosis, progressive pseudohypertrophic muscular dystrophy, Becker's muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber's congenital amaurosis, sickle cell disease, hypercholesterolemia, transthyretin amyloidosis, or β-thalassemia.

[0409] In one embodiment, the infectious agent of the infectious disease includes one or more of human immunodeficiency virus, herpes simplex virus-1, or herpes simplex virus-2.

[0410] The main advantages of this invention include:

[0411] (a) This invention is the first to discover a novel gene-editing protein (Cas protein). The gene-editing protein (Cas protein) of this invention has very good gene-editing activity and can effectively edit or cut target genes, and can effectively treat the symptoms or diseases of subjects in need (e.g., cystic fibrosis, progressive pseudohypertrophic muscular dystrophy, Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, hypercholesterolemia, transthyretin amyloidosis or β-thalassemia, one or more of these).

[0412] (b) This invention has discovered a novel Cas protein with low homology to previously reported Cas enzymes. Compared to existing Cas enzymes, it exhibits superior DNA nuclease activity and has broad application prospects.

[0413] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0414] Unless otherwise specified, the reagents and materials used in the embodiments of this invention are all commercially available products.

[0415] Example 1: Identification of a novel Cas protein

[0416] First, the inventors used computational programs to mine metagenomic data and analyze the metagenomics of uncultured organisms. Through redundancy removal and protein clustering analysis, they identified a new Cas protease, named CasW1, whose amino acid sequence is shown in SEQ ID No. 1 and nucleotide coding sequence is shown in SEQ ID No. 2. Sequence alignment confirmed that CasW1 belongs to the Cas12 family.

[0417] By analyzing and predicting metagenomics, novel CRISPR-Cas system-related proteins and components were obtained. When the CRISPR-Cas effector protein of this invention was compared with existing effector proteins, it was found that it had a low similarity to known Cas proteins.

[0418] Analysis revealed that the PAM corresponding to the Cas protein obtained in this invention is 5'-TTTN, where N represents A / T / G / C. CRISPR loci were annotated on samples containing CasW1 using PILER-CR, and the corresponding DNA encoding the direct repeat (DR) sequence was obtained as shown in SEQ ID No. 3.

[0419] Example 2: In vitro enzyme digestion experiment to verify the cleavage activity of CasW1.

[0420] To verify whether CasW1 is a double-stranded DNA nuclease, the inventors verified its cleavage activity.

[0421] 1. Construct the expression vector for CasW1.

[0422] The nucleotide sequence fragment encoding CasW1 was synthesized and cloned into the prokaryotic protein expression vector pET-28a(+) (BioLab, #QN1060) using restriction endonuclease digestion and T4 DNA ligase ligation. The ligation product was obtained; see the diagram below. Figure 6 .

[0423] The inventors performed PCR on the obtained CasW1 fragment and digested it with EcoRI / NotI. Simultaneously, the prokaryotic protein expression vector pET-28a(+) was also digested with EcoRI / NotI. The CasW1 fragment digestion product and the pET-28a(+) digestion product were then subjected to agarose gel electrophoresis. The results are as follows: Figure 1 As shown, the correctly sized CasW1 nucleotide fragment and the pET-28a(+) vector fragment with the middle 11 bp removed from the EcoRI / NotI restriction site were obtained.

[0424] The inventors then transformed the ligation product into competent E. coli DH5α cells, which were then inoculated onto LB agar plates coated with kanamycin. After overnight incubation at 37°C, single colonies were picked and Sanger sequencing was performed. Plasmid clones with correct sequences were extracted to obtain the pET-28a(+)-CasW1 expression vector.

[0425] 2. In vitro purification of CasW1 protein.

[0426] The experimental procedure is as follows:

[0427] (1) Transformation: Take a tube of competent Escherichia coli BL21(DE3) (Shanghai Weidi Biotechnology Co., Ltd., EC1002) from a -80℃ freezer and place it on ice to dissolve for 5 min. Then, add 1 ng of pET-28a(+)-CasW1 expression vector plasmid to the competent cells, gently tap the bottom of the tube to mix, and let it stand on ice for 25 min. Heat shock in a 45℃ water bath for 45 sec, then quickly return to ice and let it stand for 2 min. Add 700 μL of antibiotic-free LB medium to the centrifuge tube and incubate at 37℃, 220 rpm for 60 min. After incubation, centrifuge at 5000 rpm for 1 min to collect the bacteria, discard 600 μL of supernatant, and then gently mix the remaining liquid with the competent cells and spread it on an LB agar plate using glass beads.

[0428] (2) Induction: At 9:00 AM the following day, single clones were picked from the transformed LB plates and cultured in 3 ml of LB liquid medium containing kanamycin (10 mg / ml). The plates were then placed on a shaker and cultured at 37°C and 220 rpm for 8 hours. At 5:00 PM, the bacteria in 3 ml of LB medium were transferred to 300 ml of 2x YT medium containing kanamycin. 150 μL of IPTG (final concentration 0.5 mM) was added to the bacterial culture, and the culture was incubated overnight at 16°C for 13-14 hours.

[0429] (3) Harvesting the bacteria: Centrifuge the induced bacterial culture at 4500 rpm and 4℃ for 10 min and discard the supernatant. Resuspend the bacteria in a 50 ml centrifuge tube with 15 ml of imidazole (10 mM) (imidazole is added for competitive elution of CasW1 protein), and then add 150 μl of PMSF protease inhibitor (to inhibit protein degradation and improve yield).

[0430] (4) Ultrasound: Fix the 50ml centrifuge tube vertically in a beaker filled with ice water, and adjust the position so that the ultrasound probe is below the surface of the bacterial culture. Ultrasound mode: 3s operation, 12s interval, 200w, 60 cycles. After ultrasound, add 150μl of PMSF protease inhibitor, and then centrifuge at 11000rpm and 4℃ for 25min.

[0431] (5) Beads pretreatment: Pipette 300 μl of Ni-NTA Agarose beads (QIAGEN, #30230) into a 15 ml centrifuge tube, add 10 ml of PBS, rotate at room temperature for 5 min, centrifuge at 1000 g for 2 min at 4 °C, remove the supernatant with a dropper, and wash once more with PBS. Then add 10 ml of imidazole (10 mM), rotate at room temperature for 5 min, centrifuge at 1000 g for 2 min at 4 °C, carefully remove the supernatant with a dropper, and place the 15 ml centrifuge tube containing the washed beads on ice for later use.

[0432] (6) Aspirate all the supernatant of the bacterial culture obtained after centrifugation in step (4) into the centrifuge tube containing Ni-NTAAgarose beads obtained in step (5), and incubate at 4°C for 1 hour by rotation.

[0433] (7) Washing: Ni-NTA agarose beads were eluted twice with 10 ml of imidazole (40 mM). After each addition of imidazole, the beads were rotated at 4°C for 5 min, and then centrifuged at 1000 × g for 2 min at 4°C. The Ni-NTA agarose beads were resuspended with 500 μl of imidazole (250 mM) and then transferred to a pre-chilled affinity chromatography column (MedChemExpress, #HY-K0221). After equilibration for 5 min, the protein fraction eluted with 250 mM imidazole was collected in a 1.5 ml centrifuge tube. The above elution steps were repeated three times.

[0434] (8) After measuring the protein concentration in each tube using a NanoDrop OD280, the protein solutions were collected together and transferred to PBS buffer via a 30kDa ultrafiltration tube. Glycerol was added to a final concentration of 10%, and the solutions were aliquoted, flash-frozen in liquid nitrogen, and stored at -80°C. SDS-PAGE (polyacrylamide gel electrophoresis) was used to identify protein size and purity. Figure 2 As shown, the size of the CasW1 protein is approximately 130 kDa, indicating that a well-purified CasW1 protein has been obtained.

[0435] 3. In vitro cutting verification

[0436] 3.1 Preparation of in vitro dsDNA cleavage template.

[0437] Using the HepG2 cell (ATCC, catalog number HB-8065) ​​genome as a template, forward and reverse primers were prepared based on the hHPRT1 gene (Genebank, NG_012329.2) dsDNA template. The forward and reverse primers are shown below:

[0438] hHPRT1-dsDNA-F: gtagtgtcaactcattgctg (SEQ ID NO.5);

[0439] hHPRT1-dsDNA-R: gtcaagggcatatcctacaa (SEQ ID NO. 6).

[0440] PCR amplification was performed using Taq polymerase, and the reaction system is shown below:

[0441] 1 μl of genomic DNA (as template) (total 100 ng), 10 μl of 2×Taq PCR mix, 0.5 μl each of upstream and downstream primers, and ddH2O to bring the total volume to 20 μl.

[0442] The PCR reaction program was as follows: 95℃ for 5 min; 94℃ for 30 s, 55℃ for 30 s, 72℃ for 20 s, 35 cycles; 72℃ for 10 min; incubate at 12℃.

[0443] The PCR reaction solution was then subjected to agarose gel electrophoresis, and the electrophoresis results are as follows: Figure 3 As shown in the figure. Then, the DNA was recovered from the gel using an agarose gel DNA recovery kit (TIANGEN, DP219-02), and finally eluted with enzyme-free water to obtain the in vitro cleaved dsDNA template.

[0444] 3.2 In vitro enzymatic digestion reaction.

[0445] To test the cleavage activity of CasW1, the inventors designed two sets of comparative experiments to compare its cleavage activity. The first set compared the cleavage activity of CasW1 with LbCpf1, and the second set compared the cleavage activity of CasW1 with HED Cas12i.16 and S7R-Cas12i.3.

[0446] The amino acid sequence of the LbCpf1 protein is shown in SEQ ID NO.12, and the nucleotide coding sequence is shown in SEQ ID NO.11; the amino acid sequence of the HED Cas 12i.16 protein is shown in SEQ ID NO.10, and the nucleotide sequence is shown in SEQ ID NO.9; the amino acid sequence of the S7R-Cas12i.3 is shown in SEQ ID NO.8, and the nucleotide sequence is shown in SEQ ID NO.7.

[0447] The expression vectors LbCpf1, HED Cas 12i.16, and S7R-Cas 12i.3 were constructed using the method described in step 1 of Example 2, respectively. The recombinant expression vector maps are shown below. Figure 7 , 8 9.

[0448] 3.2.1 Comparison of cleavage activities of CasW1 and LbCpf1

[0449] The specific steps are as follows:

[0450] (1) Prepare CasW1-crRNA and LbCpf1-crRNA respectively.

[0451] A target sequence (spacer) was designed based on the hHPRT1 gene and named hHPRT1-spacer: GGTTAAAGATGGTTAAATGAT (SEQ ID NO.4).

[0452] Based on the DR sequences of CasW1 and LbCpf1, crRNA sequences for the above Cas proteins were designed and named CasW1-hHPRT1-crRNA and LbCpf1-hHPRT1-crRNA, respectively, as follows:

[0453] CasW1-hHPRT1-crRNA:

[0454] GTCTAAATGACCTATAAATTTCTACTATGTGTAGAT GGTTAAAGATGGTTAAATGAT(SEQ ID NO.13), wherein the underlined part of the sequence is the DR sequence of CasW1;

[0455] LbCpf1-hHPRT1-crRNA:

[0456] TAATTTCTACTAAGTGTAGAT GGTTAAAGATGGTTAAATGAT(SEQ ID NO.14), where the underlined part of the sequence is the DR sequence of LbCpf1;

[0457] CasW1-hHPRT1-crRNA and LbCpf1-hHPRT1-crRNA sequence fragments were chemically synthesized (by Nanjing Genscript Biotech Co., Ltd.). A mixed solution of CasW1 and CasW1-hHPRT1-crRNA was then prepared, with the following components:

[0458] 20 μl of enzyme-free H2O, 3 μl of NEBuffer r2.1 (10×, NEB, #B6002S), 3 μl of crRNA (CasW1-hHPRT1-crRNA or LbCpf1-hHPRT1-crRNA) (concentration 30 nM), 1 μl of Cas protein (CasW1 or LbCpf1) (concentration 30 nM), and the total volume of the reaction system was 27 μl.

[0459] The same method was used to prepare a mixed solution of LbCpf1 and LbCpf1-hHPRT1-crRNA.

[0460] Then, the CasW1 and CasW1-hHPRT1-crRNA mixture and the LbCpf1 and LbCpf1-hHPRT1-crRNA mixture were placed in a PCR instrument and reacted at 25°C for 10 min.

[0461] (2) Add 3 μL of 60 nM dsDNA solution (final concentration of 6 nM) to the mixed solution of CasW1 and CasW1-hHPRT1-crRNA and the mixed solution of LbCpf1 and LbCpf1-hHPRT1-crRNA respectively. The total reaction system is 30 μL. Mix thoroughly, incubate briefly in a PCR instrument at 37°C for 10 minutes.

[0462] (3) Add 1 uL Proteinase K to the reaction system of the previous step, mix thoroughly, incubate briefly at room temperature for 10 minutes to digest the protein components in the reaction system;

[0463] (4) Use magnetic beads to purify the DNA fragments after enzyme digestion.

[0464] 20 minutes beforehand, remove the DNA sorting magnetic bead solution (Novizan, catalog number N411-02) from the 4°C freezer and allow it to equilibrate to room temperature. Invert the magnetic beads to mix them thoroughly. Add 3 times the volume (approximately 150 μl) of the magnetic bead solution to the dsDNA cleavage product obtained in step (3), and gently pipette 10 times to mix. Incubate at room temperature for 10 minutes to allow the dsDNA cleavage product to bind to the magnetic beads. Place the PCR tube containing the dsDNA cleavage product sample solution on a magnetic rack. After the solution becomes clear, carefully remove the supernatant. Keep the PCR tube on the magnetic rack at all times, add 200 μl of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant. Repeat the rinsing process once more. Open the cap and dry the magnetic beads at room temperature for 5 minutes. Remove the PCR tube from the magnetic rack, add 15 μl of enzyme-free water, vortex or pipette to mix thoroughly, and let stand at room temperature for 2 minutes. Then place the PCR tube on a magnetic rack and let it stand for 5 minutes until the solution is clear. Carefully aspirate the supernatant into a new nuclease-free PCR tube.

[0465] (5) Subsequently, the dsDNA digestion products were analyzed using a portable bioanalyzer (Houzhe Biotechnology, Qsep1). The enzyme digestion effect was detected using the S1 high-resolution clip (Houzhe Biotechnology, C105102) detection protocol. The analysis results are as follows: Figure 4 Analysis using the Smear analysis option built into Qsep1 showed that the cleavage activity of CasW1 was 99.5%, while that of LbCpf1 was 95%, indicating that the cleavage activity of CasW1 was higher than that of LbCpf1.

[0466] 3.2.1 Comparison of cleavage activity of CasW1, HED Cas12i.16, and S7R-Cas12i3

[0467] The specific steps are as follows:

[0468] (1) Prepare crRNAs of CasW1, HED Cas 12i.16 and S7R-Cas 12i3 respectively.

[0469] The hHPRT1-spacer sequence fragment was synthesized. Based on the DR sequences of CasW1, HED Cas12i.16, and S7R-Cas12i3, crRNA sequences of the above Cas proteins were designed and named CasW1-hHPRT1-crRNA, HEDCas12i.16-hHPRT1-crRNA, and S7R-Cas12i3-hHPRT1-crRNA, respectively. The specific sequences are as follows:

[0470] CasW1-hHPRT1-crRNA:

[0471] GTCTAAATGACCTATAAATTTCTACTATGTGTAGAT GGTTAAAGATGGTTAAATGAT(SEQ ID NO.13), wherein the underlined part of the sequence is the DR sequence of CasW1;

[0472] HED Cas12i.16-hHPRT1-crRNA:

[0473] CTAGCAATGACTCAGAAATGTGTCCCCAGTTGACAC GGTTAAAGATGGTTAAATGAT(SEQ ID NO.15), wherein the underlined portion of the sequence is the DR sequence of HED Cas12i.16;

[0474] S7R-Cas12i3-hHPRT1-crRNA:

[0475] AGAGAATGTGTGCATAGTCACAC GGTTAAAGATGGTTAAATGAT(SEQ ID NO.16), wherein the underlined part of the sequence is the DR sequence of S7R-Cas12i3;

[0476] Sequence fragments of CasW1-hHPRT1-crRNA, HEDCas12i.16-hHPRT1-crRNA, and S7R-Cas12i3-hHPRT1-crRNA were chemically synthesized (by Nanjing Genscript Biotech Co., Ltd.) respectively.

[0477] (2) Referring to the cleavage activity comparison experiment in step 3.2.1, in vitro enzyme digestion experiments were performed with CasW1, HED Cas12i.16 and S7R-Cas 12i3 respectively. Finally, the DNA fragments after enzyme digestion were purified by magnetic beads.

[0478] (3) Subsequently, the dsDNA digestion products were analyzed using a portable bioanalyzer (Houzhe Biotechnology, Qsep1). The enzyme digestion effect was detected using the S1 high-resolution clip (Houzhe Biotechnology, C105102) detection protocol. The analysis results are as follows: Figure 5 Analysis using the Smear analysis option built into Qsep1 showed that the cleavage activity of CasW1 was 100%, that of HED Cas12i.16 was 14%, and that of S7R-Cas12i3 was 58%. The cleavage activity of CasW1 was significantly higher than that of HED Cas12i.16 and S7R-Cas12i3.

[0479] Sequence information:

[0480]

[0481]

[0482]

[0483]

[0484]

[0485]

[0486] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A gene editing protein, characterized in that, The protein is a polypeptide of the amino acid sequence set forth in SEQ ID NO:

1.

2. The gene editing protein of claim 1, wherein, The gene editing protein is used as an effector protein in a CRISPR / Cas system.

3. A fusion protein, characterized in that, The gene editing protein of claim 1; and one or more functional domains selected from the group consisting of a localization signal, a reporter protein, an epitope tag.

4. The fusion protein of claim 3, wherein, The functional domain is linked to the N-terminus and / or C-terminus of the gene editing protein.

5. The fusion protein of claim 3, wherein, The functional domain is inserted between the N-terminus and C-terminus of the gene editing protein.

6. The fusion protein of claim 3, wherein, The one or more functional domains are linked to the N-terminus and / or C-terminus of the gene editing protein via a linker.

7. An isolated polynucleotide, comprising: The polynucleotide encodes the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6.

8. The polynucleotide of claim 7, wherein, The polynucleotide is a polynucleotide of the sequence set forth in SEQ ID NO.

2.

9. An isolated nucleic acid molecule, comprising a nucleic acid sequence encoding a polypeptide of claim 1. The nucleic acid molecule is of the sequence set forth in SEQ ID NO:

3.

10. The isolated nucleic acid molecule of claim 9, wherein, The isolated nucleic acid molecule is an RNA.

11. A guide RNA (gRNA) comprising a sequence selected from the group consisting of SEQ ID NOs: 1- 12. The guide RNA comprises a direct repeat sequence set forth in SEQ ID NO: 3 and a spacer sequence linked thereto that targets a target sequence.

12. A composite, characterized in that, Comprising: (i) a protein component selected from the group consisting of the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, or a combination thereof; and (ii) a nucleic acid component selected from the group consisting of the guide RNA of claim 11, a nucleic acid encoding the guide RNA of claim 11, a precursor RNA of the guide RNA of claim 11, a nucleic acid encoding the precursor RNA of the guide RNA of claim 11, or a combination thereof; wherein the protein component and the nucleic acid component are associated with each other to form a complex.

13. A vector, characterized in that, The polynucleotide of claim 7 or 8, or the nucleic acid molecule of claim 9.

14. The vector of claim 13, wherein The vector comprises: (1) a first regulatory element operably linked to a nucleotide sequence encoding the gene editing protein of claim 1 or 2, or a nucleotide sequence encoding the fusion protein of any one of claims 3-6; and (2) a second regulatory element operably linked to a nucleotide sequence encoding a guide RNA, the guide RNA comprising: (a) a spacer sequence capable of hybridizing to a target sequence, and (b) a direct repeat (DR) sequence linked to the spacer sequence, capable of directing the gene editing protein of claim 1 to bind to the guide RNA to form the complex of claim 12 that targets the target sequence.

15. A CRISPR-Cas composition, characterized in that, Comprising: (i) a first component selected from the group consisting of the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, a nucleotide sequence encoding the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6, and any combination thereof; and (ii) a second component, which is one or more guide RNAs of claim 11, or a nucleotide sequence encoding the one or more guide RNAs of claim 11; the guide RNA is capable of forming a complex with the protein or fusion protein of (i).

16. A CRISPR-Cas system, characterized in that, comprising one or more vectors, which comprise: (i) a first nucleic acid, which is a nucleotide sequence encoding the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6; and (ii) a second nucleic acid, which encodes a nucleotide sequence comprising the guide RNA of claim 11; wherein: the first nucleic acid and the second nucleic acid are present on the same or different vectors; the guide RNA is capable of forming a complex with the protein or fusion protein of (i).

17. The system of claim 16, wherein, the first nucleic acid is operably linked to a first regulatory element; the first regulatory element is a promoter.

18. The system of claim 16, wherein, the second nucleic acid is operably linked to a second regulatory element; the second regulatory element is a promoter.

19. A kit comprising, comprising one or more components selected from the group consisting of: the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, the polynucleotide of claim 7 or 8, the complex of claim 12, the vector of claim 13 or 14, the CRISPR-Cas composition of claim 15, or the system of any one of claims 16-18.

20. A delivery composition, characterized in that, comprising a delivery vehicle, and one or more selected from the group consisting of: the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, the polynucleotide of claim 7 or 8, the complex of claim 12, the vector of claim 13, the CRISPR-Cas composition of claim 15, or the system of any one of claims 16-18.

21. A host cell, characterized in that, comprising the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, the polynucleotide of claim 7 or 8, the nucleic acid molecule of claim 9 or 10, the complex of claim 12, the vector of claim 13 or 14, the CRISPR-Cas composition of claim 15, or the system of any one of claims 16-18, or the delivery composition of claim 20.

22. An enzyme preparation, characterized in that, the enzyme preparation comprises the gene editing protein of claim 1 or 2, the fusion protein of any one of claims 3-6, the complex of claim 12, the CRISPR-Cas composition of claim 15, or the system of any one of claims 16-18, or the delivery composition of claim 20.

23. A kit characterized in that, comprising: a first container, and in the first container, the complex of claim 12, or the composition of claim 15, or the system of any one of claims 16-18, or a medicament comprising the complex of claim 12, or the composition of claim 15, or the system of any one of claims 16-18.

24. A kit characterized in that, comprising: (a1) a first container, and located in said first container a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or an isolated polynucleotide of claim 7, or an expression vector comprising an isolated polynucleotide of claim 7, or a medicament containing a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or an isolated polynucleotide of claim 7, or an expression vector comprising an isolated polynucleotide of claim 7.

25. The kit of claim 24, wherein The kit further comprises: (b1) a second container, and located in said second container a guide RNA of claim 11 or an expression vector thereof, or a medicament containing a guide RNA of claim 11 or an expression vector thereof.

26. A method of targeting and editing a target gene or cleaving a target gene that is not diagnostic and not therapeutic, characterized in that, comprising: contacting a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or a complex of claim 12 or a composition of claim 15 or a system of any one of claims 16-18 or a delivery composition of claim 20 or an enzyme preparation of claim 22 or a kit of any one of claims 23-25 with the target gene, or delivering into a cell comprising the target gene, the target sequence being present in the target gene.

27. The method of claim 26, wherein, The target gene is present in a cell.

28. The method of claim 27, wherein, The cell is a prokaryotic cell.

29. The method of claim 27, wherein, The cell is a eukaryotic cell.

30. The method of claim 29, wherein, The cell is a mammalian cell or a plant cell.

31. The method of claim 30, wherein, The cell is a human cell.

32. The method of claim 26, wherein, The target gene is present in a nucleic acid molecule in vitro.

33. A method of genetically editing a cell that is not diagnostic and not therapeutic, comprising, The method comprises contacting a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or a complex of claim 12 or a composition of claim 15 or a system of any one of claims 16-18 or a delivery composition of claim 20 or an enzyme preparation of claim 22 or a kit of any one of claims 23-25 with the target gene in a cell.

34. A method of altering expression of a gene product, other than for diagnosis or therapy, comprising, comprising: contacting a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or a complex of claim 12 or a composition of claim 15 or a system of any one of claims 16-18 or a delivery composition of claim 20 or an enzyme preparation of claim 22 or a kit of any one of claims 23-25 with a nucleic acid molecule encoding the gene product, or delivering into a cell comprising the nucleic acid molecule, the target sequence being present in the nucleic acid molecule.

35. A cell or cell line in vitro, ex vivo, or in vivo, or a progeny thereof, characterized in that, The cell or cell line or their progeny comprises: a gene editing protein of claim 1 or 2, or a fusion protein of any one of claims 3-6, or a polynucleotide of claim 7 or a complex of claim 12 or a vector of claim 13 or 14 or a composition of claim 15 or a system of any one of claims 16-18 or a delivery composition of claim 20.

36. The cell or cell line of claim 35, or a progeny of either of them, wherein, The cell is a prokaryotic cell.

37. The cell or cell line of claim 35, or a progeny of either of them, wherein The cell is a eukaryotic cell.

38. The cell or cell line of claim 37, or a progeny of either of them, wherein, The cell is a mammalian cell or a plant cell.

39. The cell or cell line of claim 37, or a progeny of either of them, wherein The cell is a human cell.

40. The cell or cell line of any of claims 37-39, or a progeny thereof, wherein, The cell is a stem cell or a stem cell line.

41. Use of the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6, or the polynucleotide of claim 7, or the nucleic acid molecule of claim 9 or 10, or the complex of claim 12, or the vector of claim 13 or 14, or the composition of claim 15, or the system of any one of claims 16-18, or the kit of claim 19, or the delivery composition of claim 20, or the enzyme preparation of claim 22, or the kit of any one of claims 23-25, characterized in that, for making a preparation for nucleic acid editing.

42. The use of claim 41, the preparation is for gene or genome editing.

43. Use of the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6, or the polynucleotide of claim 7 or the complex of claim 12 or the vector of claim 13 or 14, or the composition of claim 15 or the system of any one of claims 16-18 or the kit of claim 19 or the delivery composition of claim 20 or the enzyme preparation of claim 22 or the kit of any one of claims 23-25, characterized in that, for making a preparation for one or more selected from the group consisting of: (i) ex vivo gene or genome editing; (ii) detection of single-stranded DNA ex vivo; (iii) treating a disease in a subject in need thereof, the disease comprising cystic fibrosis, Duchenne muscular dystrophy (DMD), alpha-1-antitrypsin deficiency, myotonic dystrophy, Huntington's disease, amyotrophic lateral sclerosis, sickle cell disease, beta thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, cervical cancer, esophageal cancer, stomach cancer, head and neck cancer, ovarian cancer, lymphoma, leukemia.

44. The use of claim 43, wherein the compound is administered in a daily dose of about 0.1 to about 100 mg / kg. The disease comprises hypercholesterolemia, melanoma, acute lymphoblastic leukemia, acute myelogenous leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin's lymphoma.

45. A method of detecting the presence or absence of a target nucleic acid molecule in a sample that is not diagnostic and not therapeutic, comprising, The method comprises contacting a sample with the gene editing protein of claim 1 or 2, or the fusion protein of any one of claims 3-6, or the complex of claim 12, or the composition of claim 15, or the system of any one of claims 16-18, or the kit of claim 19, or the delivery composition of claim 20, or the enzyme preparation of claim 22, and a non-target sequence, detecting a detectable signal produced by cleavage of the non-target sequence by the protein in the complex or CRISPR-Cas composition or system or delivery composition, thereby detecting the target nucleic acid molecule, the non-target sequence not hybridizing to the guide RNA.

46. The method of claim 45, wherein, The non-target sequence is cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, indicating the presence of the target nucleic acid molecule in the sample; and the non-target sequence is not cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, indicating the absence of the target nucleic acid molecule in the sample.

Citation Information

Patent Citations

  • Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells

    WO2001038547A2

  • Cpf1 protein, V-type gene editing system and application

    CN116751763A