Cas protein, crispr-cas system containing cas protein, and use of cas protein

By providing a new Cas protein containing a specific domain and corresponding guide RNA, the CRISPR-Cas system is formed, which solves the problem of lack of diversified application needs in the prior art and achieves efficient gene editing effects.

WO2025119363A1PCT designated stage expired Publication Date: 2025-06-12YOLTECH THERAPEUTICS CO LTD

Patent Information

Application Number
PCT/CN2024/137580
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

There is a lack of novel Cas protein and CRISPR-Cas systems in the prior art to meet diverse application needs.

Method used

A novel Cas protein is provided, including OBD, REC, RuvC, Helical and Nuc domains, and combines guide RNA to form the CRISPR-Cas system.

Benefits of technology

It achieves efficient gene editing activity and specificity, can effectively edit or cleave the target gene, and is suitable for a variety of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024137580-FTAPPB-I100001
    Figure PCTCN2024137580-FTAPPB-I100001
  • Figure PCTCN2024137580-FTAPPB-I100002
    Figure PCTCN2024137580-FTAPPB-I100002
  • Figure PCTCN2024137580-FTAPPB-I100003
    Figure PCTCN2024137580-FTAPPB-I100003
Patent Text Reader

Abstract

The present disclosure provides a Cas protein, a CRISPR-Cas system containing the Cas protein, and a use of the Cas protein. Specifically, the Cas protein of the present disclosure comprises an OBD domain, a REC domain, a RuvC domain, a helical domain, and a Nuc domain, and has a structure as shown in formula I or formula II. The Cas protein of the present disclosure has very good gene editing activity, can effectively edit or cleave a target gene, and can effectively treat disorders or diseases of a subject in need.
Need to check novelty before this filing date? Find Prior Art

Description

A Cas protein, a CRISPR-Cas system containing the same, and applications thereof

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit and priority of the following patent application: Patent application number CN202311664660.5, filed on December 6, 2023, entitled "A Cas protein, a CRISPR-Cas system comprising the same and its application", and the entire contents of the application, including any sequence listings and drawings, are incorporated herein by reference in their entirety.

[0003] References to electronic sequence listings

[0004] This disclosure contains an electronic sequence listing ("P2024-3068xlb.xml," created by "WIPOSequence" software in accordance with WIPO Standard ST.26), which is incorporated herein by reference in its entirety. In accordance with WIPO Standard ST.26, the symbol "t" is used to represent T in DNA and U in RNA. Therefore, in sequence listings prepared in accordance with ST.26, whenever a sequence is RNA, a T in the sequence should be treated as a U. Technical Field

[0005] The present disclosure relates to the field of gene editing, and in particular, to a Cas protein, a CRISPR-Cas system comprising the same, and applications thereof. Background Art

[0006] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as the CRISPR-Cas or CRISPR / Cas system, are currently considered to be the source of bacterial and archaeal immunity against phage infection. The CRISPR-Cas system of prokaryotic adaptive immunity is a diverse set of protein effectors, non-coding elements, and loci that can be engineered for applications such as gene editing, target detection, and disease treatment.

[0007] Currently, there are many Cas proteins and corresponding editing technologies. The field still needs new Cas proteins and CRISPR-Cas systems to meet diverse application needs. Summary of the Invention

[0008] The main purpose of this disclosure is to provide new Cas proteins and CRISPR-Cas systems to meet diverse application needs.

[0009] In one aspect, the present disclosure provides a Cas protein comprising an OBD domain, a REC domain, a RuvC domain, a Helical domain, and a Nuc domain.

[0010] In some embodiments, the RuvC domain comprises RuvC-I, RuvC-II, and RuvC-III domains.

[0011] In some embodiments, the Cas protein does not comprise an HNH domain and a PI domain.

[0012] In some embodiments, the RuvC-III domain is located between the Nuc-I domain and the Nuc-II domain.

[0013] In some embodiments, the OBD domain is a bi-split domain comprising an OBD-I and an OBD-II domain.

[0014] In some embodiments, the OBD-I domain is located at the N-terminus and the Nuc-II domain is located at the C-terminus.

[0015] In some embodiments, the Cas protein performs nucleic acid cleavage function without the aid of tracrRNA.

[0016] In another aspect, the present disclosure provides a fusion protein comprising the Cas protein of the present disclosure; and one or more functional domains.

[0017] In some embodiments, the functional domain is selected from a localization signal, a reporter protein, a Cas protein targeting portion, a DNA binding domain, an epitope tag, a transcription activation domain, a transcription repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcription release factor, an HDAC, a cleavage-active polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, and a base excision repair inhibitor (such as a uracil-DNA glycosylase inhibitor (UGI)).

[0018] In some embodiments, the functional domain comprises one or more of the following enzymatic activities on a target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity.

[0019] In some embodiments, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0020] In one aspect, the present disclosure provides an isolated polynucleotide encoding the Cas protein of the present disclosure or the fusion protein of the present disclosure.

[0021] In one aspect, the present disclosure provides an isolated nucleic acid molecule comprising the structure shown in Formula IV below:

[0022] 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3'(IV),

[0023] wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1) having a plurality (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) of nucleotide pairs in the Cas protein;

[0024] Segments Ba and Bb do not base pair with each other and form a bulge (B);

[0025] Segments R2a and R2b are reverse complementary sequences and form a second stem (R2), which has multiple (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) base pairs; and L is a loop formed at the second stem portion and formed by multiple (3, 4, 5, 6, 7, 8, 9, 10) nucleotides.

[0026] In some embodiments, the nucleic acid molecule comprises the structure shown in Formula IV below:

[0027] 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3'(IV),

[0028] wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1) having 3 or 5 nucleotide pairs in Cas12o;

[0029] Segments Ba and Bb do not exist simultaneously, and a bulge (B) formed by the presence of segment Ba or segment Bb is formed by 2 or 3 nucleotides;

[0030] Segments R2a and R2b are reverse complementary sequences and form a second stem (R2) having 6 or 7 base pairs; and L is a loop formed at the second stem and having 5 or 7 nucleotides.

[0031] In some embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:

[0032] (i) a sequence shown in any one of SEQ ID NOs: 2, 4, and 6;

[0033] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 2, 4, and 6;

[0034] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 2, 4, and 6;

[0035] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0036] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0037] Furthermore, the sequence of any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived;

[0038] For example, the isolated nucleic acid molecule is RNA;

[0039] For example, the isolated nucleic acid molecule comprises a direct repeat (DR) sequence in the CRISPR / Cas system.

[0040] In some embodiments, the nucleic acid molecule comprises one or more stem-loops or optimized secondary structures;

[0041] For example, the sequence of any of (ii)-(v) retains the secondary structure of the sequence from which it is derived.

[0042] In some embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:

[0043] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 2, 4, and 6;

[0044] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0045] (c) A complementary sequence to the nucleotide sequence shown in any one of SEQ ID NOs: 2, 4, and 6.

[0046] In one aspect, the present disclosure provides a guide RNA (gRNA), which includes a direct repeat (DR) sequence capable of binding to the Cas protein of the present disclosure and a spacer sequence capable of targeting a target sequence.

[0047] In one aspect, the present disclosure provides a vector comprising the polynucleotide of the present disclosure and / or the nucleic acid molecule of the present disclosure.

[0048] In one aspect, the present disclosure provides a composite comprising:

[0049] (i) a protein component selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, or a combination thereof; and

[0050] (ii) a nucleic acid component selected from the group consisting of a guide RNA of the present disclosure, a nucleic acid encoding a guide RNA of the present disclosure, a precursor RNA of the guide RNA of the present disclosure, a precursor RNA nucleic acid encoding a guide RNA of the present disclosure, or a combination thereof;

[0051] In one aspect, the present disclosure provides a CRISPR-Cas composition comprising:

[0052] (i) a first component selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, a nucleotide sequence encoding a Cas protein of the present disclosure or a fusion protein of the present disclosure, and any combination thereof; and

[0053] (ii) a second component comprising one or more guide RNAs disclosed herein, or a nucleotide sequence encoding one or more guide RNAs disclosed herein; the guide RNA comprising:

[0054] (iii) a direct repeat (DR) sequence capable of binding to the Cas protein described in the present disclosure; and

[0055] (iv) a spacer sequence capable of targeting a target sequence of a target DNA, wherein the guide RNA is configured to form a complex with the Cas protein;

[0056] In one aspect, the present disclosure provides a CRISPR-Cas system comprising one or more vectors, wherein the one or more vectors comprise:

[0057] (i) a first nucleic acid, which is a nucleotide sequence encoding the Cas protein of the present disclosure or the fusion protein of the present disclosure; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0058] (ii) a second nucleic acid encoding a nucleotide sequence of a guide RNA of the present disclosure; optionally, the second nucleic acid is operably linked to a second regulatory element; the guide RNA comprises:

[0059] (iii) a direct repeat (DR) sequence capable of binding to the Cas protein of the present disclosure; and

[0060] (iv) a spacer sequence capable of targeting a target sequence of a target DNA, wherein the guide RNA is configured to form a complex with the Cas protein;

[0061] in:

[0062] The first nucleic acid and the second nucleic acid are present on the same or different vectors. The guide RNA is capable of forming a complex with the Cas protein or fusion protein described in (i).

[0063] In some embodiments, the vector comprises a plasmid or a viral vector.

[0064] In some embodiments, the guide RNA includes a spacer sequence capable of hybridizing to a target sequence; and a direct repeat (DR) sequence connected to the spacer sequence and capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.

[0065] In some embodiments, the guide RNA includes unmodified and modified guide RNA.

[0066] In some embodiments, the modified guide RNA includes chemical modifications of bases.

[0067] In some embodiments, the chemical modification comprises methylation modification, methoxy modification, fluorination modification, or thiolation modification.

[0068] In some embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0069] In some embodiments, at least one component of the composition is non-naturally occurring or modified.

[0070] In some embodiments, the spacer sequence is linked to the 3' end of the direct repeat (DR) sequence.

[0071] In some embodiments, the spacer sequence comprises a complementary sequence to the target sequence.

[0072] In some embodiments, when the target sequence is DNA, the target sequence is located 3' to a protospacer adjacent motif (PAM), and the PAM is 5'-TN, wherein N is A, T, G, or C.

[0073] In some embodiments, the target sequence is a DNA from a prokaryotic cell or a eukaryotic cell, or a DNA sequence formed based on RNA reverse transcription; alternatively, the target sequence is a non-naturally occurring DNA, or a DNA sequence formed based on RNA reverse transcription.

[0074] In some embodiments, the target sequence comprises a cDNA sequence.

[0075] In some embodiments, the target sequence comprises a single-stranded DNA or a double-stranded DNA sequence.

[0076] In some embodiments, the target sequence is present within a cell.

[0077] In some embodiments, the target sequence is present in the nucleus or in the cytoplasm (eg, an organelle).

[0078] In some embodiments, the cell is a eukaryotic cell.

[0079] In some embodiments, the cell is a prokaryotic cell.

[0080] In some embodiments, the target sequence is present outside the cell.

[0081] In some embodiments, the Cas protein of the present disclosure is linked to one or more NLS sequences, or the fusion protein comprises one or more NLS sequences.

[0082] In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the Cas protein of the present disclosure.

[0083] In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the Cas protein of the present disclosure.

[0084] In one aspect, the present disclosure provides a kit comprising one or more components selected from the following: a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a CRISPR-Cas system of the present disclosure.

[0085] In some embodiments, the kit further comprises a label or instructions.

[0086] In some embodiments, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, and cleaving a target gene or a non-target gene.

[0087] In one aspect, the present disclosure provides a delivery composition comprising a delivery vector or delivery medium, and one or more selected from the following: a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a CRISPR-Cas system of the present disclosure.

[0088] In one aspect, the present disclosure provides a host cell comprising a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, a CRISPR-Cas system of the present disclosure, or a delivery composition of the present disclosure.

[0089] In one aspect, the present disclosure provides an enzyme preparation comprising the Cas protein of the present disclosure, the fusion protein of the present disclosure, the complex of the present disclosure, the CRISPR-Cas composition of the present disclosure, or the CRISPR-Cas system of the present disclosure, or the delivery composition of the present disclosure.

[0090] In another aspect, the present disclosure provides a kit comprising:

[0091] A first container, and a complex of the present disclosure, or a CRISPR-Cas composition of the present disclosure, or a CRISPR-Cas system of the present disclosure, or a drug containing the complex of the present disclosure, or the CRISPR-Cas composition of the present disclosure, or the CRISPR-Cas system of the present disclosure, located in the first container.

[0092] In another aspect, the present disclosure provides a kit comprising:

[0093] (a1) a first container, and a Cas protein of the present disclosure, or a fusion protein of the present disclosure, or a gene encoding the Cas protein of the present disclosure, or an expression vector thereof, or a drug containing the Cas protein of the present disclosure, or a fusion protein of the present disclosure, or a gene encoding the Cas protein of the present disclosure, or an expression vector thereof, located in the first container;

[0094] (b1) an optional second container, and the guide RNA of the present disclosure or its expression vector, or a drug containing the guide RNA of the present disclosure or its expression vector, located in the second container.

[0095] In another aspect, the present disclosure provides a method for targeting and editing a target gene or cleaving a target gene, comprising: contacting the Cas protein of the present disclosure, the fusion protein of the present disclosure, the complex of the present disclosure, the CRISPR-Cas composition of the present disclosure, the CRISPR-Cas system of the present disclosure, the delivery composition of the present disclosure, the enzyme preparation of the present disclosure, or the drug kit of the present disclosure with the target gene, or delivering it to a cell comprising the target gene, wherein the target sequence is present in the target gene.

[0096] In another aspect, the present disclosure provides a method of inducing a change in a cell state, the method comprising contacting a Cas protein of the present disclosure, a fusion protein of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, a CRISPR-Cas system of the present disclosure, a delivery composition of the present disclosure, an enzyme preparation of the present disclosure, or a drug kit of the present disclosure with a target gene in a cell.

[0097] In another aspect, the present disclosure provides a method for altering the expression of a gene product, comprising: contacting the Cas protein of the present disclosure, the fusion protein of the present disclosure, the complex of the present disclosure, the CRISPR-Cas composition of the present disclosure, the CRISPR-Cas system of the present disclosure, the delivery composition of the present disclosure, the enzyme preparation of the present disclosure, or the drug kit of the present disclosure with a nucleic acid molecule encoding the gene product, or delivering it to a cell comprising the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0098] In another aspect, the present disclosure provides a cell or progeny thereof obtained by any of the methods described herein, wherein the cell comprises a modification that is not present in its wild-type form.

[0099] In another aspect, the present disclosure provides a cell product of a cell of the present disclosure or a progeny thereof.

[0100] In another aspect, the present disclosure provides an in vitro, ex vivo or in vivo cell or cell line or progeny thereof, comprising: a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, a CRISPR-Cas system of the present disclosure, or a delivery composition of the present disclosure.

[0101] In another aspect, the present disclosure provides a cell preparation comprising the host cell of the present disclosure, the cell of the present disclosure or its progeny, or a cell product of the cell of the present disclosure or its progeny, or the cell or cell line of the present disclosure or its progeny.

[0102] On the other hand, the present disclosure also provides uses of the Cas protein of the present disclosure, the fusion protein of the present disclosure, the polynucleotide of the present disclosure, the vector of the present disclosure, the complex of the present disclosure, the CRISPR-Cas composition of the present disclosure, the CRISPR-Cas system of the present disclosure, the kit of the present disclosure, the delivery composition of the present disclosure, the enzyme preparation of the present disclosure, or the drug kit of the present disclosure for preparing a medicament or formulation for nucleic acid editing (e.g., gene or genome editing).

[0103] In another aspect, the present disclosure provides uses of the Cas protein of the present disclosure, the fusion protein of the present disclosure, the polynucleotide of the present disclosure, the vector of the present disclosure, the complex of the present disclosure, the CRISPR-Cas composition of the present disclosure, the CRISPR-Cas system of the present disclosure, the kit of the present disclosure, the delivery composition of the present disclosure, the enzyme preparation of the present disclosure, or the kit of the present disclosure for preparing a medicament or formulation for use in one or more selected from the group consisting of:

[0104] (i) ex vivo gene or genome editing;

[0105] (ii) detection of single-stranded DNA in vitro;

[0106] (iii) editing a target sequence in a target locus to modify an organism or non-human organism;

[0107] (iv) treating a disorder caused by a defect in the target sequence in the target locus;

[0108] (v) treating a condition or disease in a subject in need thereof.

[0109] In another aspect, the present disclosure provides a method for detecting the presence of a target nucleic acid molecule in a sample, the method comprising contacting the sample with a Cas protein, a fusion protein, a complex, a CRISPR-Cas composition, a CRISPR-Cas system, a kit, a delivery composition, or an enzyme preparation, and a non-target sequence, and detecting a detectable signal generated by cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

[0110] In another aspect, the present disclosure provides a method of treating a condition or disease in a subject in need thereof, comprising administering to the subject a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, a CRISPR-Cas system of the present disclosure, a kit of the present disclosure, a delivery composition of the present disclosure, an enzyme preparation of the present disclosure, or a pharmaceutical kit of the present disclosure.

[0111] In another aspect, the present disclosure provides a sterile container comprising a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a CRISPR-Cas system of the present disclosure, or a delivery composition of the present disclosure, or an enzyme preparation of the present disclosure.

[0112] In another aspect, the present disclosure provides an implantable device comprising a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, a CRISPR-Cas system of the present disclosure, a delivery composition of the present disclosure, or an enzyme preparation of the present disclosure.

[0113] It should be understood that within the scope of the present disclosure, the above-mentioned technical features of the present disclosure and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 depicts a phylogenetic tree of Cas12o homologs.

[0115] Figure 2 depicts the schematic domain structure (Figures 2A, 2C) and predicted three-dimensional structure (Figure 2B) of Cas12o.

[0116] Figure 3 depicts the Cas12o1 expression vector (Figure 3A) and the LbCpf1 expression vector (Figure 3B)

[0117] FIG4 depicts the Target plasmid carrying the target sequence hTTR1.

[0118] Figure 5 depicts a comparison of the cleavage activities of Cas12o and LbCpf1 against the hTTR1 target sequence in competent cells.

[0119] Figure 6 describes the secondary structure prediction of the DR sequences of Cas12o1, Cas12o2, and Cas12o3.

[0120] FIG7 depicts the experimentally determined PAM preference of Cas12o1. DETAILED DESCRIPTION

[0121] After extensive and in-depth research, the present inventors have discovered a new Cas protein for the first time. The Cas protein of the present disclosure includes an OBD domain, a RuvC domain, a Helical domain, and a Nuc domain, and has a structure shown in Formula I, Formula II, or Formula III. The Cas protein of the present disclosure has excellent gene editing activity and specificity, can effectively edit or cut target genes, and can be used to treat conditions or diseases in subjects in need.

[0122] Many modifications and other embodiments of the disclosure set forth herein will occur to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description. Therefore, it should be understood that the disclosure is not limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, such terms are used in a generic and descriptive sense only and not for purposes of limitation.

[0123] the term

[0124] The following examples are only used to illustrate the present disclosure, rather than to limit the present disclosure. Unless otherwise specified, the experiments and methods described in the examples were basically carried out according to conventional methods well known in the art and described in various references.

[0125] In addition, if specific conditions are not specified in the examples, the results are carried out according to conventional conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments is not specified, they are all conventional products that can be obtained commercially. It is understood by those skilled in the art that the examples describe the present disclosure by way of example and are not intended to limit the scope of the present disclosure. All public cases and other references mentioned herein are incorporated herein by reference in their entirety.

[0126] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.

[0127] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which can be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein) and determining the number of positions at which identical residues occur. Typically, this is expressed as a percentage. The measurement of sequence identity of nucleotide sequences is a method well known to those skilled in the art.

[0128] General Definition

[0129] The articles "a" and "an" are used herein to refer to one or more than one (ie, at least one) of the grammatical object of the article. For example, "a polypeptide" means one or more polypeptides.

[0130] The term "about" or "approximately" refers to a quantity, level, value, quantity, frequency, percentage, dimension, size, amount, weight, or length that varies by up to 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% compared to a reference quantity, level, value, quantity, frequency, percentage, dimension, size, amount, weight, or length. In one embodiment, the term "about" or "approximately" refers to a quantity, level, value, quantity, frequency, percentage, dimension, size, amount, weight, or length that ranges around a reference quantity, level, value, quantity, frequency, percentage, dimension, size, amount, weight, or length by ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1%.

[0131] The term "optional" or "optionally" means that the subsequently described event, circumstance, or alternative may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0132] Throughout this specification, unless the context requires otherwise, the terms "comprises," "including," "contains," and "having" should be understood to imply the inclusion of a stated step or element or group of steps or elements, but not the exclusion of any other step or element or group of steps or elements. In specific embodiments, the terms "comprises," "including," "contains," and "having" are used synonymously.

[0133] The term "heterologous" refers to a nucleotide or polypeptide sequence that is not present in a natural nucleic acid or protein, respectively. For example, relative to Cas12o, a heterologous polypeptide comprises an amino acid sequence from a protein other than the Cas12o protein. In some cases, a portion of a Cas12o protein from one species is fused with a portion of a Cas12o protein from a different species. Therefore, it can be considered that the Cas12o sequences from each species are heterologous to each other. As another example, the Cas12o protein (e.g., dCas12o protein) can be fused with an active domain from a non-Cas12o protein (e.g., deaminase, histone deacetylase), and the sequence of the active domain can be considered to be a heterologous polypeptide (it is heterologous to the Cas12o protein).

[0134] The terms "orthologs" and "homologs" are well known in the art. As further guidance, a "homolog" of a protein as used herein is a protein from the same species that performs the same or similar function as the protein to which it is a homolog. Homologous proteins can be, but need not be, structurally related, or only partially structurally related. An "ortholog" of a protein as used herein is a protein from a different species that performs the same or similar function as the protein to which it is a homolog. Orthologous proteins can be, but need not be, structurally related, or only partially structurally related. In one embodiment, a homolog or ortholog of a nucleic acid-guided nuclease, such as those mentioned herein, has at least 80%, at least 85%, at least 90%, or at least 95% sequence homology or identity with the nucleic acid-guided nuclease. In another embodiment, a homolog or ortholog of a nucleic acid-guided nuclease has at least 80%, at least 85%, at least 90%, or at least 95% sequence identity with a wild-type nucleic acid-guided nuclease.

[0135] Other orthologs of known nucleic acid-guided nucleases can be identified. Some methods for identifying orthologs of nucleic acid-guided nucleases can involve identifying a tracr sequence in the target genome. Identification of the tracr sequence can involve the following steps: searching a database for direct repeat sequences or tracr mate sequences to identify regions containing the nucleic acid-guided nuclease. Searching for homologous sequences in regions flanking the nucleic acid-guided nuclease in both the sense and antisense directions. Searching for transcription terminators and secondary structures. Identifying any sequence that is not a direct repeat sequence or tracr mate sequence but has greater than 50% identity to a direct repeat sequence or tracr mate sequence as a potential tracr sequence. Obtaining a potential tracr sequence and analyzing its associated transcription terminator sequence.

[0136] The chimeric enzyme can comprise a first segment and a second segment, and the segments can be fragments of nucleic acid-guided nuclease orthologs from a genus or species of organism, for example, the fragments are from nucleic acid-guided nuclease orthologs from different species.

[0137] The terms "polynucleotide" and "nucleic acid" are used interchangeably herein to refer to a polymeric form of nucleotides (ribonucleotides or deoxynucleotides) of any length. Thus, the term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derived nucleotide bases. The terms "polynucleotide" and "nucleic acid" should be understood to include single-stranded (such as sense or antisense strands) and double-stranded polynucleotides as applicable to the described embodiments.

[0138] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymeric form of amino acids of any length, which may include genetically encoded and non-genetically encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. The term includes: fusion proteins, including but not limited to fusion proteins with heterologous amino acid sequences, fusions with heterologous and homologous leader sequences, with or without an N-terminal methionine residue; immunolabeled proteins, etc.

[0139] The term "isolated" is intended to describe a polynucleotide, polypeptide or cell that is in an environment that is different from that in which the polynucleotide, polypeptide or cell naturally occurs. An isolated genetically modified host cell can be present in a mixed population of genetically modified host cells.

[0140] The term "exogenous nucleic acid" refers to a nucleic acid that does not normally or naturally occur in nature and / or is not produced by a given bacterium, organism, or cell. As used herein, the term "endogenous nucleic acid" refers to a nucleic acid that normally occurs in nature and / or is produced by a given bacterium, organism, or cell. "Endogenous nucleic acid" is also referred to as a "native nucleic acid" or a nucleic acid that is "native" to a given bacterium, organism, or cell.

[0141] The term "recombinant," specifically nucleic acid (DNA or RNA), is the product of various combinations of cloning, restriction, and / or ligation steps that produce constructs having structural coding sequences or non-coding sequences that can be distinguished from endogenous nucleic acids present in natural systems. In general, the DNA sequence encoding the structural coding sequence can be assembled from cDNA fragments and short oligonucleotide linkers or from a series of synthetic oligonucleotides to provide a synthetic nucleic acid that can be expressed by a recombinant transcription unit contained in a cell or in a cell-free transcription and translation system. Such sequences can be provided in the form of open reading frames that are not interrupted by internal non-translated sequences or introns, which are typically present in eukaryotic genes. Genomic DNA comprising the relevant sequence can also be used in the formation of recombinant genes or transcription units. The sequence of non-translated DNA can be present at the 5' end or 3' end of the open reading frame, where such sequences do not interfere with the manipulation or expression of the coding region and can in fact play a role in regulating the production of the desired product by various mechanisms.

[0142] Therefore, for example, the term "recombinant" polynucleotide or "recombinant" nucleic acid refers to a non-naturally occurring polynucleotide or nucleic acid, such as a polynucleotide or nucleic acid made by the artificial combination of two otherwise separated segments of a sequence through human intervention. This artificial combination is often accomplished by chemical synthesis means or by artificially manipulating the separation segments of nucleic acid (for example, by genetic engineering techniques). This operation is usually performed to replace codons with redundant codons encoding identical or conservative amino acids, while usually introducing or removing sequence recognition sites. Alternatively, nucleic acid segments with desired functions are linked together to produce desired functional combinations. This artificial combination is often accomplished by chemical synthesis means or by artificially manipulating the separation segments of nucleic acid (for example, by genetic engineering techniques).

[0143] Similarly, the term "recombinant" polypeptide refers to a non-naturally occurring polypeptide, e.g., a polypeptide made by the artificial combination of two otherwise separate segments of amino acid sequence through human intervention. Thus, for example, a polypeptide comprising a heterologous amino acid sequence is recombinant.

[0144] The term "operably linked" refers to a juxtaposition in which the components are in a relationship that allows them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. As used herein, the terms "heterologous promoter" and "heterologous control region" refer to promoters and other control regions that are not normally associated with a particular nucleic acid in nature. For example, a "transcription control region heterologous to a coding region" is a transcription control region that is not normally associated with a coding region in nature.

[0145] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. This is a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment can be inserted to achieve replication of the inserted segment. Generally, a vector is capable of replication when combined with appropriate control elements. In some cases, a vector system comprises a single vector. Alternatively, a vector system comprises multiple vectors. The vector can be a viral vector.

[0146] Vectors include, but are not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other polynucleotide variants known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which other DNA segments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which there is a virally derived DNA or RNA sequence for packaging into a virus (e.g., a retrovirus, a replication-defective retrovirus, adenovirus, a replication-defective adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Certain vectors are capable of autonomous replication in the host cells into which they are introduced (e.g., bacterial vectors and episomal mammalian vectors with a bacterial origin of replication). After being introduced into the host cell, other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell, thereby replicating together with the host genome. In addition, certain vectors are capable of guiding the expression of genes operably linked thereto. Such vectors are referred to herein as "expression vectors." Vectors that are expressed in eukaryotic cells and vectors that cause expression in eukaryotic cells may be referred to herein as “eukaryotic expression vectors.” Common expression vectors useful in recombinant DNA techniques are often in the form of plasmids.

[0147] Recombinant expression vector can be suitable for comprising nucleic acid of the present disclosure in the form of expressing nucleic acid in host cell, this means that recombinant expression vector comprises one or more regulatory elements, and described regulatory elements can be selected according to the host cell to be used for expression, and described nucleic acid is operably connected to nucleic acid sequence to be expressed.In recombinant expression vector, " operably connected " is intended to refer to that target nucleotide sequence is connected to regulatory elements in the mode that allows nucleotide sequence expression (for example, in in vitro transcription / translation system or when vector is introduced into host cell in host cell).Advantageous vector includes slow virus and adeno-associated virus, and the type of these vectors can also be selected to target specific type of cell.

[0148] The term "host cell" refers to a cell (e.g., cell line) from a multicellular organism cultured as a eukaryotic cell, a prokaryotic cell, or as a unicellular entity, which can be used as or has been used as a receptor for nucleic acids (e.g., expression vectors), and includes progeny of the original cells genetically modified by nucleic acid. It should be understood that due to natural, accidental, or intentional mutations, the progeny of the unicellular cell may not necessarily be identical in morphology or in genome or total DNA complement sequence to the original parent." recombinant host cell" (also referred to as "genetically modified host cell") is a host cell into which heterologous nucleic acids (e.g., expression vectors) have been introduced. For example, a subject prokaryotic host cell is a prokaryotic host cell (e.g., a bacterium) that has been genetically modified by introducing a heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the prokaryotic host cell (not normally found in nature) or a recombinant nucleic acid that is not normally found in a prokaryotic host cell, into a suitable prokaryotic host cell; and a subject eukaryotic host cell is a eukaryotic host cell that has been genetically modified by introducing a heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the eukaryotic host cell or a recombinant nucleic acid that is not normally found in a eukaryotic host cell, into a suitable eukaryotic host cell.

[0149] The term "amino acid" refers to the twenty common naturally occurring amino acids. Naturally occurring amino acids include alanine (Ala, A), arginine (Arg, R), asparagine (Asn, N), aspartic acid (Asp, D), cysteine ​​(Cys, C), glutamic acid (Glu, E), glutamine (Gln, Q), glycine (Gly, G), histidine (His, H), isoleucine (Ile, I), leucine (Leu, L), lysine (Lys, K), methionine (Met, M), phenylalanine (Phe, F), proline (Pro, P), serine (Ser, S), threonine (Thr, T), tryptophan (Trp, W), tyrosine (Tyr, Y), and valine (Val, V).

[0150] "Sequence identity" between two polypeptides or nucleic acid sequences represents the percentage of the number of identical residues between the sequences to the total number of residues, and the calculation of the total number of residues is determined based on the mutation type. Mutation types include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence. Taking polypeptides as an example (similar to nucleotides), if the mutation type is one or more of the following: substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence, the total number of residues is calculated using the larger of the molecules being compared. If the mutation type also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., the number of insertions or deletions at both ends is less than 20) is not included in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps (if any) in the alignment are resolved by a specific algorithm.

[0151] The term "conservative amino acid substitution" refers to the replacement of an amino acid with a chemically or functionally similar amino acid. Conservative substitution tables providing similar amino acids are well known in the art. For example, in some embodiments, the following groups of amino acids are considered to be conservative substitutions for each other. In some embodiments, the selected groups of amino acids considered to be conservative substitutions for each other are:

[0152] Table A

[0153] In some embodiments, the selected group of amino acids that are considered conservative substitutions for each other are:

[0154] Table B

[0155] In one embodiment, other selected groups of amino acids that are considered conservative substitutions for each other (see, e.g., Creighton, Proteins (1984)):

[0156] Table C

[0157] In some embodiments, other selected groups of amino acids that are considered conservative substitutions for each other are:

[0158] Table D

[0159] The terms "treatment" and "treating" refer to obtaining a desired pharmacological and / or physiological effect. The effect may be prophylactic, in terms of completely or partially preventing a disease or its symptoms, and / or therapeutic, in terms of partially or completely curing a disease and / or side effects attributable to the disease. As used herein, "treatment" covers any treatment of a disease in a mammal (e.g., a human) and includes: (1) preventing the occurrence of a disease in a subject who may be susceptible to the disease but has not yet been diagnosed with the disease; (2) inhibiting the disease, i.e., arresting its development; and (3) relieving the disease, i.e., causing regression of the disease.

[0160] The terms "individual," "subject," "host," and "patient" are used interchangeably herein to refer to an individual organism, such as a mammal, including but not limited to rodents, simians, humans, mammalian farm animals, mammalian sports animals, and mammalian pets.

[0161] Overview

[0162] The present disclosure provides RNA-guided endonuclease polypeptides, referred to herein as "Cas12o" polypeptides (also referred to as "Cas12o proteins"); nucleic acids encoding Cas12o proteins; and modified host cells comprising Cas12o proteins and / or nucleic acids encoding Cas12o proteins. Cas12o proteins can be used in various applications provided, and are smaller in size than other Cas (e.g., Cas9 or Cas12) and are easier to deliver (including by AAV or LNP delivery).

[0163] The present disclosure provides guide RNAs that bind to Cas12o proteins and provide sequence specificity for Cas12o proteins (referred to herein as "Cas12o guide RNAs," "guide RNAs," "crRNAs," or "guide RNAs (gRNAs)"); nucleic acids encoding Cas12o guide RNAs; and modified host cells comprising Cas12o guide RNAs and / or nucleic acids encoding Cas12o guide RNAs. Cas12o guide RNAs can be used in various applications provided.

[0164] Cas12o protein

[0165] The term "Cas12o protein" includes wild-type Cas12o protein, its derivatives or variants, and functional fragments thereof such as oligonucleotide-binding fragments.

[0166] In some embodiments, the Cas12o protein comprises an OBD domain, a REC domain, a RuvC domain, a Helical domain, and a Nuc domain.

[0167] In some embodiments, the RuvC domain includes a RuvC-I domain, a RuvC-II domain, and a RuvC-III domain.

[0168] In some embodiments, the Cas12o protein does not comprise an HNH domain and a PI domain.

[0169] In some embodiments, the Cas protein is a class 2, type V Cas endonuclease.

[0170] In some embodiments, the OBD domain is a bi-split domain, including discontinuous OBD-I domain and OBD-II domain.

[0171] In some exemplary embodiments, the size of the Cas12o protein is between 500 and 1200 amino acids, between 500 and 1100 amino acids, between 700 and 1100 amino acids, and between 900 and 1000 amino acids, and the size variation may depend in part on the specific domain architecture of Cas12o or its homologs.

[0172] The Cas12o protein may be derived from a naturally occurring protein, a modified naturally occurring protein, a functional fragment or truncated version thereof, or a non-naturally occurring protein. In one embodiment, the Cas12o protein may comprise one or more domains derived from other Cas12o protein nucleases, more particularly from different organisms. In one embodiment, the Cas12o protein nuclease may be designed by a computer method. Examples of computer protein design have been described in the art and are therefore known to those skilled in the art. In a specific embodiment, the Cas12o protein locus is not associated with a CRISPR array.

[0173] Cas12o protein may also encompass homologs or orthologs of the Cas12o protein whose sequence is specifically described herein. Orthologous proteins may, but need not be, structurally related or only partially related in structure. In one embodiment, a homolog or ortholog of the Cas12o protein as mentioned herein has at least 80%, at least 85%, at least 90%, at least 95%, at least 99% sequence homology or identity with the Cas12o protein nuclease. In another embodiment, a homolog or ortholog of the Cas12o protein nuclease has at least 80%, at least 85%, at least 90% or at least 95% sequence identity with the wild-type Cas12o protein nuclease.

[0174] In some embodiments, the Cas protein comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1, 3, 5, 7-9. Exemplarily, the Cas12o protein comprises an amino acid sequence as shown in any one of SEQ ID NOs. 1, 3, 5, 7-9, wherein SEQ ID NOs. 7-9 are functional fragments of Cas12o.

[0175] Cas12o variants

[0176] The Cas12o protein may comprise one or more modifications. The term "modified" generally refers to a variant (Cas12o variant) nuclease of a Cas12o protein having one or more modifications or mutations (including point mutations, truncations, insertions, deletions, chimeras, fusion proteins, etc.) compared to its wild-type counterpart. By derived, it is meant that the derivative enzyme is primarily based on the wild-type enzyme in the sense of having a high degree of sequence homology to the wild-type enzyme, but has been mutated (modified) in a manner known in the art or as described herein.

[0177] In some embodiments, the derivative enzyme of the Cas12o protein includes a Cas12o protein that substantially lacks catalytic activity (deadCas12o, dCas12o) or a Cas12o nickase (nicklase Cas12o, nCas12o) with single-stranded cleavage ability.

[0178] Compared to the wild-type counterpart nuclease, dCas12o may have reduced or no nuclease activity (retaining less than 50% (e.g., less than any about 40%, 35%, 30%, 27.5%, 25%, 22.5%, 20%, 17.5%, 15%, 12.5%, 10%, 7.5%, 25 5%, 4%, 3%, 2.5%, 2%, 1% or less) of the corresponding original Cas12o protein (e.g., a novel Cas12o protein comprising an amino acid sequence of any one of SEQ ID NOs: 1, 3, 5, 7-9) or a variant thereof. An example can be when the nucleic acid cleavage activity of the mutant form is zero or negligible compared to the non-mutant form. Cas12o proteins can be identified by reference to the general class of enzymes having homology to the largest nuclease having multiple nuclease domains from type I, type II, type III, type IV, type V or type VI CRISPR systems.

[0179] In some cases, a catalytically inactive or dead nuclease may have nickase (nCas) activity. In some cases, a catalytically inactive or dead nuclease may not have nickase activity. Such a catalytically inactive or dead nuclease may not cause double-stranded or single-stranded breaks on the target polynucleotide, but may still bind to the target polynucleotide or otherwise form a complex.

[0180] As described herein, the Cas12o nickase (nicklase Cas12o, nCas12o) with single-stranded cleavage ability is obtained by modifying the Cas12o protein, and one or more amino acid mutations are introduced into the Cas12o protein to enable it to have a nickase single-stranded DNA cleavage activity that cuts one chain of double-stranded DNA.

[0181] In one embodiment, the modification of Cas12o protein may or may not result in functional changes. For example, modifications that do not result in functional changes include, for example, codon optimization for expression in a specific host, or providing specific markers to the nuclease (for example, for visualization). Modifications that may result in functional changes may also include mutations, including point mutations, insertions, deletions, truncations (including split nucleases), etc., and chimeric nucleases (for example, comprising domains from different orthologs or homologs) or fusion proteins. The chimeric enzyme may include a first fragment and a second fragment, and the fragment may be a fragment of the Cas12o protein nuclease ortholog of a genus or a species of organism, for example, the fragment is from different species of Cas12o protein nuclease orthologs.

[0182] In one embodiment, the nuclease domain of the Cas12o protein is catalytically inactive, or is modified to be catalytically inactive, or is modified to be a nickase. In one embodiment, both nuclease domains are catalytically inactive.

[0183] In one embodiment, the Cas12o protein nuclease may include one or more modifications that result in enhanced activity and / or specificity, such as including mutant residues that stabilize the targeted or non-targeted chain. In one embodiment, the altered or modified activity of the engineered Cas12o protein includes increased targeting efficiency or reduced off-target binding. In one embodiment, the altered activity of the engineered Cas12o protein nuclease includes a modified cleavage activity. In one embodiment, the altered activity includes an increased cleavage activity to the target polynucleotide locus. In one embodiment, the altered activity includes a reduced cleavage activity to the target polynucleotide locus. In one embodiment, the altered activity includes a reduced cleavage activity to the off-target polynucleotide locus. In one embodiment, the altered or modified activity of the modified nuclease includes altered helicase kinetics. In one embodiment, the modified nuclease includes a modification that changes the association of the protein with a nucleic acid molecule comprising RNA, or a chain of a target polynucleotide locus, or a chain of an off-target polynucleotide. In one aspect of the present disclosure, the engineered Cas12o protein nuclease includes a modification that changes the formation of the Cas12o protein nuclease and the associated complex. In one embodiment, the activity of the change includes the cutting activity to the increase of the polynucleotide locus of off-target. Therefore, in one embodiment, compared to the polynucleotide locus of off-target, the specificity of the target polynucleotide locus is increased. In other embodiments, compared to the polynucleotide locus of off-target, the specificity of the target polynucleotide locus is reduced. In one embodiment, mutation causes the reduction of off-target effect (such as cutting or binding properties, activity or kinetics), such as causing the tolerance of the mismatch between target and crRNA to be reduced. Other mutations may cause off-target effect (such as, cutting or binding properties, activity or kinetics) to increase. Other mutations may cause on-target effect (such as, cutting or binding properties, activity or kinetics) to increase or decrease. In one embodiment, mutation causes the change (such as, increase or decrease) helicase activity, association or formation of the functional nuclease complex. In one embodiment, mutation causes PAM recognition to change, i.e., compared to unmodified Cas12o protein nuclease, it is possible (additionally or alternatively) to identify different PAMs.

[0184] In one embodiment, Cas12o protein can be guided to the position of the target sequence or near the target sequence, such as within the target sequence and / or within the complementary sequence of the target sequence or at the cutting of one or two DNA chains at the sequence associated with the target sequence. In one embodiment, Cas12o protein can guide the cutting of one or two DNA chains within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 300, 400, 500 or more base pairs or nucleotides from the first or last nucleotide of the target sequence. In one embodiment, the cutting position of Cas12o is the cutting of about 12-19 nucleotides to two DNA chains from the first nucleotide of the target sequence. In one embodiment, the cutting can be staggered, i.e., producing sticky ends. In one embodiment, the cutting is a staggered cut with a 5' overhang. In one embodiment, the cleavage is a staggered nick with a 5' overhang of 1 to 15 nucleotides, preferably 4 or 9 nucleotides.

[0185] In one embodiment, the cleavage site is distal to the target adjacent motif (TAM), which is used interchangeably with the term "PAM" herein, e.g., cleavage occurs after the nth nucleotide on the non-target strand and after the nucleotide on the targeted strand. In one embodiment, the cleavage site occurs after an identified nucleotide on the non-target strand (calculated from the PAM) and after a further identified nucleotide on the targeted strand (calculated from the PAM). In one embodiment, the vector encodes a nucleic acid-targeting effector protein that can be mutated relative to the corresponding wild-type enzyme such that the mutated nucleic acid-targeting effector protein lacks the ability to cleave one or both DNA and RNA strands of a target polynucleotide containing a target sequence.

[0186] The reference Cas12o protein disclosed herein comprises REC-I domain, REC-II domain, OBD domain, RuvC-I domain, Helical domain, RuvC-II domain, Nuc-I domain, RuvC-III domain and Nuc-II domain (Fig. 2A) in the order of N-terminal-C-terminal. In some cases, the oligonucleotide binding domain (OBD) is a bi-split domain, including discontinuous OBD-I domain and OBD-II domain, in which case OBD-I is located at the N-terminal of Cas12o protein, and Cas12o comprises OBD-I domain, REC-I domain, REC-II domain, OBD-II domain, RuvC-I domain, Helical domain, RuvC-II domain, Nuc-I domain, RuvC-III domain and Nuc-II domain (Fig. 2C) in the order of N-terminal-C-terminal.

[0187] OBD domain (oligonucleotide binding domain)

[0188] The reference Cas12o protein of the present disclosure comprises an oligonucleotide binding domain (OBD). Certain Cas proteins other than Cas12o have domains that can be named in a similar manner. However, in some embodiments, the OBD comprises one or more unique functional features, or comprises a sequence that is unique relative to the Cas12o protein, or a combination thereof.

[0189] In one embodiment, the OBD domain comprises an OBD-I domain and an OBD-II domain, as shown in the OBD domain distribution in Figure 2C. Exemplarily, the OBD-I domain comprises the amino acid sequence at positions 1-15 in SEQ ID NO:5, and the OBD-II domain comprises the amino acid sequence at positions 333-480 in SEQ ID NO:5.

[0190] In one embodiment, the OBD domain comprises only a single domain, exemplified by the OBD domain arrangement shown in Figure 2A. In this case, an exemplary OBD domain is represented by amino acids 359-474 of SEQ ID NO: 1 (OBD-II) or amino acids 337-452 of SEQ ID NO: 3 (OBD-II).

[0191] RuvC domain

[0192] The reference Cas12o protein disclosed herein comprises a RuvC domain, which includes a tri-split RuvC domain (tri-split RuvC domain), including three discontinuous RuvC domains (RuvC-I, RuvC-II and RuvC-III domains). The RuvC domain is the ancestral domain of all 12-type CRISPR proteins. The RuvC domain is derived from a TnpB (transposase B)-like transposase. Similar to other RuvC domains, the Cas12o RuvC domain has a DED catalytic triad responsible for coordinating magnesium (Mg) ions and cleaving DNA.

[0193] Exemplarily, the RuvC-I domain comprises the amino acid sequence of positions 475-563 in SEQ ID NO:1, the amino acid sequence of positions 453-536 in SEQ ID NO:3, and the amino acid sequence of positions 481-560 in SEQ ID NO:5; the RuvC-II domain comprises the amino acid sequence of positions 766-818 in SEQ ID NO:1, the amino acid sequence of positions 732-784 in SEQ ID NO:3, and the amino acid sequence of positions 758-812 in SEQ ID NO:5; and the RuvC-III domain comprises the amino acid sequence of positions 854-869 in SEQ ID NO:1, the amino acid sequence of positions 817-832 in SEQ ID NO:3, and the amino acid sequence of positions 845-860 in SEQ ID NO:5.

[0194] REC domain

[0195] The REC (recognition) domain comprises at least one REC domain (e.g., a REC-I domain and optionally a REC-II domain), which is believed to interact with the repeat: anti-repeat duplex of crRNA and mediate the formation of the Cas protein / crRNA complex.

[0196] The reference Cas12o protein disclosed herein comprises a REC domain comprising a first REC domain (REC-I) and a second REC domain (REC-II) from N-terminus to C-terminus. Exemplarily, the REC-I domain comprises the amino acid sequence of positions 1-165 in SEQ ID NO: 1, the amino acid sequence of positions 1-183 in SEQ ID NO: 3, and the amino acid sequence of positions 16-196 in SEQ ID NO: 5, and the REC-II domain comprises the amino acid sequence of positions 166-358 in SEQ ID NO: 1, the amino acid sequence of positions 184-336 in SEQ ID NO: 3, and the amino acid sequence of positions 197-332 in SEQ ID NO: 5.

[0197] Helical domain

[0198] The reference Cas12o disclosed herein comprises a Helical domain. Exemplarily, the Helical domain comprises the amino acid sequence at positions 564-765 in SEQ ID NO: 1, the amino acid sequence at positions 537-731 in SEQ ID NO: 3, and the amino acid sequence at positions 561-757 in SEQ ID NO: 5.

[0199] Nuc domain

[0200] The Nuc domain is thought to be involved in target strand cleavage (Yamano et al., Cell 2016, 165: 949-962). Other mutational studies in other Cas12 proteins have shown that the Nuc domain contributes to guide and target binding (Swarts et al., Mol Cell 2017, 66: 221-233). The reference Cas12o disclosed herein comprises a Nuc domain (including a Nuc-I domain and a Nuc-II domain), exemplified by a Nuc-I domain comprising an amino acid sequence of positions 819-853 in SEQ ID NO: 1, an amino acid sequence of positions 785-816 in SEQ ID NO: 3, and an amino acid sequence of positions 813-844 in SEQ ID NO: 5, and a Nuc-II domain comprising an amino acid sequence of positions 870-984 in SEQ ID NO: 1, an amino acid sequence of positions 833-954 in SEQ ID NO: 3, and an amino acid sequence of positions 861-966 in SEQ ID NO: 5.

[0201] Exemplarily, the structural domains and amino acid positions of Cas12o disclosed herein are shown in Table 1:

[0202] Table 1

[0203] Guide RNA (crRNA, sgRNA)

[0204] As used herein, the term "guide RNA" is used interchangeably with guide molecules, guide RNA, gRNA, or crRNA, and refers to nucleic acid-based molecules, including but not limited to RNA-based molecules (e.g., direct repeat (DR) sequences) capable of forming a complex with a CRISPR-Cas protein, and comprising a targeting sequence (e.g., a spacer sequence) that is sufficiently complementary to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide sequence-specific binding of the complex to the target nucleic acid sequence.

[0205] In some embodiments, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 100%.

[0206] In some embodiments, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 100% over the seven contiguous nucleotides most 3' to the target site of the target nucleic acid.

[0207] In some embodiments, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 17 or more (e.g., 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more) consecutive nucleotides. In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 17 or more (e.g., 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more) consecutive nucleotides. In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 17 or more (e.g., 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more) consecutive nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 17 or more (e.g., 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more) consecutive nucleotides.

[0208] In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 consecutive nucleotides. In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 99% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 consecutive nucleotides. In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 consecutive nucleotides. In some cases, the percent complementarity between the targeting sequence and the target site of the target nucleic acid is 100% over 19-25 contiguous nucleotides.

[0209] In some cases, the targeting sequence has a length in the range of 19-30 nucleotides (nt) (e.g., 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some cases, the targeting sequence has a length in the range of 19-25 nucleotides (nt) (e.g., 19-22, 19-20, 20-25, 20-25, or 20-22 nt). In some cases, the targeting sequence has a length of 19 or more nt (e.g., 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the targeting sequence has a length of 19 nt. In some cases, the targeting sequence has a length of 20 nt. In some cases, the guide sequence has a length of 21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.

[0210] In some embodiments, the crRNA of Cas12o comprises, or consists essentially of, or consists of: a direct repeat (DR) sequence and a spacer (Spacer) sequence. In some embodiments, the crRNA comprises, or consists essentially of, or consists of: a direct repeat sequence connected to a spacer sequence. In some embodiments, the crRNA comprises a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-Spacer-DR). This is a typical feature of the precursor crRNA (pre-crRNA) configuration. In some embodiments, the crRNA comprises a direct repeat sequence, a spacer sequence, a direct repeat sequence, and a spacer sequence (DR-Spacer-DR-Spacer). In some embodiments, the crRNA comprises two or more direct repeat sequences and two or more spacer sequences. In some embodiments, the crRNA includes a truncated direct repeat sequence, and a spacer sequence. This is a typical feature of a processed or mature crRNA. In some embodiments, the CRISPR-Cas12o effector protein forms a complex with the crRNA, and the spacer sequence guides the complex to sequence-specific binding with the target nucleic acid, and the target nucleic acid is complementary to the spacer sequence.

[0211] Any DR sequence that can mediate the binding of the Cas12o protein described herein to the corresponding crRNA can be used in the present disclosure.

[0212] In one embodiment, the general DR sequence of Cas12o of the present disclosure comprises 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3', wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1), which has a plurality of (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) nucleotide pairs in Cas12o; segment Ba and Bb do not base pair with each other and form a bulge (B); segments R2a and R2b are reverse complementary sequences and form a second stem (R2), which has multiple (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) base pairs; and L is a loop formed at the second stem portion, formed by multiple (3, 4, 5, 6, 7, 8, 9, 10) nucleotides.

[0213] In one embodiment, the DR sequence comprises 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3', wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1), which has 3 or 5 nucleotide pairs in Cas12o; segments Ba and Bb are not present at the same time, and a bulge (B) formed by the existing segment Ba or segment Bb is formed by 2 or 3 nucleotides; segments R2a and R2b are reverse complementary sequences and form a second stem (R2), which has 6 or 7 base pairs; and L is a loop formed at the second stem portion, formed with 5 or 7 nucleotides.

[0214] In one embodiment, the DR sequence is as shown in FIG6A , which comprises 5′-R1a(ACA)-Ba(absent)-R2a(GGUAUCC)-L(UAAAC)-R2b(GGAUGCU)-Bb(GA)-R1b(UGU)-3′.

[0215] In one embodiment, the DR sequence is as shown in FIG6B , which comprises 5′-R1a(UUACA)-Ba(absent)-R2a(ACUAUUC)-L(UUGAAAC)-R2b(GAAUGGU)-Bb(GAU)-R1b(UGUAA)-3′.

[0216] In one embodiment, the DR sequence is as shown in FIG6C , which comprises 5′-R1a(UCAGU)-Ba(GUG)-R2a(GGUCUG)-L(AAACA)-R2b(CAGACC)-Bb(absent)-R1b(AUUGA)-3′.

[0217] In some embodiments, the DR sequence corresponding to the Cas12o protein of the present disclosure is shown in SEQ ID NO. 2, 4, and 6. In some embodiments, the direct repeat comprises a "functional variant" of the sequence shown in SEQ ID NO. 2, 4, or 6, such as a "functional truncated version", "functionally extended version", or "functionally replaced version". For example, the DR variant obtained by truncating or deleting the DR sequence still has DR function, and the DR "functional variant" is a 5' and / or 3' end extension (functionally extended version) or truncation (functionally truncated version) of the reference DR (such as the parent DR), and / or the reference DR sequence is inserted, deleted, and / or replaced (functionally replaced version) One or more nucleotides still have at least 20% (such as at least about any 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or higher) of the reference DR function, that is, the function of mediating the binding of the Cas12o protein to the corresponding crRNA. DR functional variants generally retain a stem-loop-like secondary structure or portion thereof that can be bound by the Cas12o protein. In one embodiment, the stem-loop structure of the DR sequence of Cas12o1 can be as shown in Figure 6. In some embodiments, the stem of the direct repetition contained in the crRNA consists of 10-13 pairs of complementary bases that hybridize with each other, which generally contain 1 RNA bulge, and the loop length is 5-7 nucleotides. In some embodiments, the loop length is 5 nucleotides; in some embodiments, the loop length is 7 nucleotides. In different embodiments, the stem may include at least 10, at least 11, at least 12, or at least 13 base pairs. In some embodiments, the direct repetition includes two nucleotide complementary segments with a total length of about 10-15 nucleotides, and 5-7 nucleotides that constitute the loop. In some embodiments, the stem-loop structure comprises a first stem nucleotide chain having a length of 10-15 nucleotides; a second stem nucleotide chain having a length of 10-15 nucleotides, wherein the first and second stem nucleotide chains can hybridize with each other; and a cyclic nucleotide chain arranged between the first and second stem nucleotide chains, wherein the cyclic nucleotide chain comprises 5, 6, or 7 nucleotides. In some embodiments, the cyclic nucleotide chain comprised by the stem-loop structure comprises at least 3 adenine nucleotides.

[0218] In one embodiment, the DR sequence that can guide Cas12o to the target site has one or more nucleotide changes selected from nucleotide addition, insertion, deletion and substitution that do not cause substantial differences in the secondary structure compared to the DR sequence shown in any one of SEQ ID NO. 2, 4, and 6. Exemplary DR sequences include nucleotide sequences having 80% or higher identity (e.g., 85% or higher, 90% or higher, 93% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100% identity) to the sequence shown in any one of SEQ ID NO. 2, 4, and 6.

[0219] In some embodiments, the length of the spacer sequence is greater than 17 nucleotides, preferably 17 to 100 nucleotides, more preferably 16 to 50 nucleotides (e.g., 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides), more preferably 17 to 50 nucleotides, more preferably 17 to 40 nucleotides, more preferably 18 to 39 nucleotides, and most preferably 18 to 37 nucleotides.

[0220] Fusion protein

[0221] The Cas12o protein (e.g., dCas12o) can have one or more functional domains associated (e.g., via a fusion protein, or a suitable linker), including, for example, one or more domains from the group comprising, consisting essentially of, or consisting of: associated with one or more functional domains selected from a nuclear localization signal (NLS) domain, a nuclear export signal (NES) domain, a translation activation domain, a transcriptional activation domain (e.g., VP64, p65, MyoD1, HSF1, RTA, and SET7 / 9), a translation initiation domain, a transcriptional repression domain (e.g., a KRAB domain, a NuE domain, an NcoR domain, and a SID domain, such as a SID4X domain), a nuclease domain (e.g., FokI), a histone modification domain (e.g., a histone acetyltransferase), a light-inducible / controllable domain, a chemically-inducible / controllable domain, a transposase domain, a homologous recombination machinery domain, a recombinase domain, an integrase domain, a topoisomerase, and combinations thereof.

[0222] In one embodiment, the functional domain comprises a deaminase. In another embodiment, the functional domain is a transposase. In another embodiment, the functional domain is a reverse transcriptase. In some cases, the CRISPR-Cas12o complex as a whole may be associated with two or more functional domains. For example, there may be two or more functional domains associated with the Cas12o protein, or there may be two or more functional domains associated with the guide RNA or crRNA, or there may be one or more functional domains associated with the effector protein targeting RNA and one or more functional domains associated with the guide RNA or crRNA.

[0223] In one embodiment, the Cas12o protein is associated with one or more functional domains, which can be achieved by direct connection of the effector protein to the functional domain, or by association with crRNA. In a non-limiting example, the crRNA comprises an added or inserted sequence that can associate with the target functional domain, including, for example, an aptamer or nucleotide that binds to a nucleic acid binding adapter protein. The functional domain can be a functional heterologous domain.

[0224] In one embodiment, the Cas12o protein is associated with one or more functional domains, and this association can be achieved by direct connection of the effector protein to the functional domain, or by association with crRNA. In a non-limiting example, the crRNA contains an added or inserted sequence that can associate with the target functional domain, including, for example, an aptamer or nucleotide that binds to a nucleic acid binding adapter protein.

[0225] In some embodiments, the functional domains may be functional heterologous domains. At least one or more heterologous functional domains may be located at or near the amino terminus of the effector protein and / or at least one or more heterologous functional domains may be located at or near the carboxyl terminus of the effector protein. The one or more heterologous functional domains may be fused to the effector protein. The one or more heterologous functional domains may be tethered to the effector protein. The one or more heterologous functional domains may be attached to the effector protein via a linker moiety.

[0226] In one embodiment, the one or more functional domains are heterologous functional domains. In some embodiments, the heterologous functional domains have one or more of the following activities: nuclease activity, methylation activity, demethylation activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), deglycosylation activity, transcriptional repression activity, transcriptional activation activity, translational activation activity, translational repression activity, histone modification activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity, chromatin modification or remodeling activity and detectable activity,

[0227] In one embodiment, Cas12o protein or its ortholog or homolog can be used as a universal nucleic acid binding protein fused to or operably connected to a functional domain. Exemplary functional domains may include, but are not limited to, nuclear localization signal (NLS), nuclear export signal (NES), deaminase (e.g., adenosine deaminase or cytidine deaminase) domain, transcriptional activation domain, DNA methylation catalytic domain, histone residue modification domain, nuclease catalytic domain, fluorescent protein, transcriptional modification factor, light-gating factor, chemical inducible factor, chromatin visualization factor, providing a targeting polypeptide for binding to a cell surface portion on a target cell or target cell type, epigenetic modification domain, transposase domain, reverse transcriptase domain, topoisomerase, phosphatase, polymerase.

[0228] In some embodiments, described functional domain is included in some preferred embodiments, functional domain is transcriptional activation domain, such as but not limited to VP64, p65, MyoD1, HSF1, RTA, SET7 / 9 or histone acetyltransferase.In one embodiment, functional domain is transcriptional repression domain, is preferably KRAB.In one embodiment, transcriptional repression domain is the concatemer (such as SID4X) of SID or SID.In one embodiment, functional domain is epigenetic modification domain, thereby provides epigenetic modification enzyme.In one embodiment, functional domain is transcriptional activation domain, it can be P65 activation domain.

[0229] In some embodiments, the nucleic acid-guided nuclease is associated with a ligase or a functional fragment thereof. The ligase can connect the single-strand breaks (nicks) produced by the nucleic acid-guided nuclease. In some cases, the ligase can connect the double-strand breaks produced by the nucleic acid-guided nuclease. In some instances, the nucleic acid-guided nuclease is associated with a reverse transcriptase or a functional fragment thereof.

[0230] Preferably, a transposase domain, a HR (homologous recombination) machinery domain, a recombinase domain and / or an integrase domain are used as functional domains of the present disclosure. In one embodiment, the DNA integration activity comprises a HR machinery domain, an integrase domain, a recombinase domain and / or a transposase domain.

[0231] In one embodiment, the DNA cleavage activity is due to a nuclease. In one embodiment, the nuclease includes a Fok1 nuclease (e.g., see “Dimeric CRISPR RNA-guided FokI nucleases for highly specific genome editing”, Shengdar Q. Tsai, Nicolas Wyvekens, Cyd Khayter, Jennifer A. Foden, Vishal Thapar, Deepak Reyon, Mathew J. Goodwin, Martin J. Aryee, J. Keith Joung Nature Biotechnology 32(6):569--77 (2014)), which relates to a dimeric RNA-guided FokI nuclease that recognizes extended sequences and can efficiently edit endogenous genes in human cells.

[0232] In one embodiment, the Cas12o protein may include one or more heterologous functional domains. As used herein, a heterologous functional domain is a polypeptide that is not derived from the same species as the nucleic acid-guided nuclease. For example, the heterologous functional domain of the nucleic acid-guided nuclease derived from species A is a polypeptide derived from a species different from species A, or an artificial polypeptide. One or more heterologous functional domains may include one or more nuclear localization signal (NLS) domains. One or more heterologous functional domains may include at least two or more NLSs. One or more heterologous functional domains may include one or more transcriptional activation domains. The transcriptional activation domain may include VP64. One or more heterologous functional domains may include one or more transcriptional repression domains. The transcriptional repression domain may include a KRAB domain or a SID domain. One or more heterologous functional domains may include one or more nuclease domains. One or more nuclease domains may include Fok1.

[0233] In one embodiment, one or more functional domains comprise acetyltransferase, preferably histone acetyltransferase. These can be used in the field of epigenomics, for example, in the method for querying epigenome. The method for querying epigenome can include, for example, targeting epigenomic sequence. Targeting epigenomic sequence can include guiding thing to epigenomic target sequence. Epigenomic target sequence can include, in one embodiment, including promoter, silencer or enhancer sequence.

[0234] Examples of acetyltransferases are known, but may in one embodiment include a histone acetyltransferase. In one embodiment, the histone acetyltransferase may comprise the catalytic core of human acetyltransferase p300 (Gerbasch & Reddy, Nature Biotech 2015 April 6).

[0235] Base editing

[0236] In some embodiments, the Cas12o protein (e.g., dCas12o) can be associated with (e.g., fused to) a deaminase (e.g., an adenosine deaminase or a cytidine deaminase) that can change the identity of a nucleotide, for example, from C·G to T·A or from A·T to G·C (Gaudelli et al., Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. "Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems." Science 353(6305) (2016); Komor et al. "Programmable editing of a target base in genomic DNA without double-stranded DNA." "Programmable editing of target bases in genomic DNA without double-stranded DNA cleavage." Nature 533(7603)(2016): 420-4. The base editing fusion protein can comprise, for example, an active (double-strand break-generating), partially active (nickase), or inactive (catalytically inactive) Cas12o nuclease and deaminase. Base editing repair inhibitors and glycosylase inhibitors (e.g., in some embodiments, uracil glycosylase inhibitors (to prevent uracil removal)) are contemplated as additional components of the base editing system. In some embodiments, in certain instances, the nucleotide deaminase is a mutant form of adenosine deaminase. The mutant form of adenosine deaminase can have both adenosine deaminase and cytidine deaminase activity.

[0237] Adenosine deaminase

[0238] The term "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyzes the hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule). In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0239] According to the present disclosure, the adenosine deaminase that can be used in conjunction with the present disclosure includes but is not limited to an enzyme family member (ADAR) called an adenosine deaminase acting on RNA, an enzyme family member (ADAT) called an adenosine deaminase acting on tRNA, and other family members containing adenosine deaminase domain (ADAD). According to the present disclosure, adenosine deaminase can target adenine in RNA / DNA and RNA duplexes. In fact, Zheng et al., (Nucleic Acids Res.2017, 45 (6): 3369-3377) confirmed that ADAR can perform adenosine to inosine editing reaction on RNA / DNA and RNA / RNA duplexes. In a specific embodiment, adenosine deaminase has been modified to increase the ability of the DNA in the RNA / DNA heteroduplex of its editing RNA duplex, as described in detail below.

[0240] In some embodiments, the adenosine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the adenosine deaminase is human, squid, or Drosophila adenosine deaminase.

[0241] In some embodiments, the adenosine deaminase is a human ADAR, including hADAR1, hADAR2, and hADAR3. In some embodiments, the adenosine deaminase is a Caenorhabditis elegans ADAR protein, including ADR-1 and ADR-2. In some embodiments, the adenosine deaminase is a Drosophila ADAR protein, including dAdar. In some embodiments, the adenosine deaminase is a squid (Loligo pealeii) ADAR protein, including sqADAR2a and sqADAR2b. In some embodiments, the adenosine deaminase is a human ADAT protein. In some embodiments, the adenosine deaminase is a Drosophila ADAT protein. In some embodiments, the adenosine deaminase is a human ADAD protein, including TENR (hADAD1) and TENRL (hADAD2).

[0242] In some embodiments, the adenosine deaminase is a TadA protein, such as Escherichia coli TadA. See Kim et al., Biochemistry 45:6407-6416 (2006); Wolf et al., EMBO J. 21:3841-3851 (2002). In some embodiments, the adenosine deaminase is mouse ADA. See Grunebaum et al., Curr. Opin. Allergy Clin. Immunol. 13:630-638 (2013). In some embodiments, the adenosine deaminase is human ADAT2. See Fukui et al., J. Nucleic Acids 2010:260512 (2010). In some embodiments, the deaminase (e.g., an adenosine or cytidine deaminase) is one or more of those described in Cox et al., Science. 2017 Nov 24;358(6366):1019-1027; Komore et al., Nature. 2016 May 19;533(7603):420-4; and Gaudelli et al., Nature. 2017 Nov 23;551(7681):464-471.

[0243] In some embodiments, the adenosine deaminase protein recognizes one or more target adenosine residues in a double-stranded nucleic acid substrate and converts them into inosine residues. In some embodiments, the double-stranded nucleic acid substrate is an RNA-DNA hybrid duplex. In some embodiments, the adenosine deaminase protein recognizes a binding window on the double-stranded substrate. In some embodiments, the binding window comprises at least one target adenosine residue. In some embodiments, the binding window is in the range of about 3 bp to about 100 bp. In some embodiments, the binding window is in the range of about 5 bp to about 50 bp. In some embodiments, the binding window is in the range of about 10 bp to about 30 bp. In some embodiments, the binding window is about 1 bp, 2 bp, 3 bp, 5 bp, 7 bp, 10 bp, 15 bp, 20 bp, 25 bp, 30 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, or 100 bp.

[0244] In some embodiments, the adenosine deaminase protein comprises one or more deaminase domains. Without wishing to be bound by a particular theory, it is expected that the deaminase domain is used to identify one or more target adenosine (A) residues contained in a double-stranded nucleic acid substrate and convert them into inosine (I) residues. In some embodiments, the deaminase domain comprises an active center. In some embodiments, the active center comprises a zinc ion. In some embodiments, during the AI ​​editing process, the base pairing at the target adenosine residue is destroyed, and the target adenosine residue is "flipped" out of the double helix to become accessible to the adenosine deaminase. In some embodiments, the amino acid residue in or near the active center interacts with one or more nucleotides at the 5' end of the target adenosine residue. In some embodiments, the amino acid residue in or near the active center interacts with one or more nucleotides at the 3' end of the target adenosine residue. In some embodiments, the amino acid residue in or near the active center further interacts with nucleotides complementary to the target adenosine residue on the opposite chain. In some embodiments, the amino acid residue forms a hydrogen bond with the 2' hydroxyl group of the nucleotide.

[0245] In some embodiments, the adenosine deaminase comprises a human ADAR2 holoprotein (hADAR2) or a deaminase domain thereof (hADAR2-D). In some embodiments, the adenosine deaminase is an ADAR family member homologous to hADAR2 or hADAR2-D.

[0246] In particular, in some embodiments, the homologous ADAR protein is human ADAR1 (hADAR1) or its deaminase domain (hADAR1-D). In some embodiments, glycine 1007 of hADAR1-D corresponds to glycine 487 hADAR2-D, and glutamate 1008 of hADAR1-D corresponds to glutamate 488 of hADAR2-D.

[0247] In some embodiments, the adenosine deaminase comprises the wild-type amino acid sequence of hADAR2-D. In some embodiments, the adenosine deaminase comprises one or more mutations in the hADAR2-D sequence such that the editing efficiency and / or substrate editing preference of hADAR2-D is altered according to specific needs.

[0248] In some embodiments, the adenosine deaminase catalytic domain comprises an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO:30, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO:30.

[0249] In some embodiments, the catalytic domain of adenosine deaminase includes a mutant of the amino acid sequence shown in SEQ ID NO: 30: E18K+F19S+N20L, named adenosine deaminase 004V14 (see WO2023193536A1).

[0250] In some embodiments, the adenosine deaminase catalytic domain comprises an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 31 (selected from 005V1 deaminase in CN114634923A, in which the amino acid sequence is SEQ ID NO: 2), and it retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 31.

[0251] In some embodiments, the amino acid sequence of the adenosine deaminase catalytic domain comprises amino acid additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO: 30 or 31.

[0252] In some embodiments, the catalytic domain of adenosine deaminase comprises a mutant of the amino acid sequence shown in SEQ ID NO: 31: Q148G+Q149M+P150R, designated as deaminase 005V1-10-3.

[0253] In some embodiments, the functional domain is full-length or a functional fragment of TadA8e.

[0254] In some embodiments, the adenosine deaminase is 004V1 (SEQ ID NO. 30) or 005V1 (SEQ ID NO. 31).

[0255] Cytidine deaminase

[0256] In some embodiments, the deaminase is a cytidine deaminase. As used herein, the term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze the hydrolytic deamination reaction of cytosine (or the cytosine portion of a molecule) into uracil (or the uracil portion of a molecule), as shown below. In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0257] According to the present disclosure, cytidine deaminases that can be used in conjunction with the present disclosure include, but are not limited to, members of the enzyme family known as apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-induced deaminases (AID), or cytidine deaminase 1 (CDA1). In specific embodiments, the deaminases are APOBEC1 deaminases, APOBEC2 deaminases, APOBEC3A deaminases, APOBEC3B deaminases, APOBEC3C deaminases, and APOBEC3D deaminases, APOBEC3E deaminases, APOBEC3F deaminases, APOBEC3G deaminases, APOBEC3H deaminases, or APOBEC4 deaminases.

[0258] In the methods and systems disclosed herein, cytidine deaminase is capable of targeting cytosine in a single strand of DNA. In certain example embodiments, cytidine deaminase can be edited on a single strand present outside a binding component, such as in conjunction with Cas13. In other example embodiments, cytidine deaminase can be edited in a localized bubble, such as a localized bubble formed by a target editing site but a guide sequence mismatch. In certain example embodiments, cytidine deaminase may include mutations that contribute to focused activity, such as those described in Kim et al., Nature Biotechnology (2017) 35(4): 371-377 (doi: 10.1038 / nbt.3803.

[0259] In some embodiments, the cytidine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the cytidine deaminase is a human, primate, cow, dog, rat, or mouse cytidine deaminase.

[0260] In some embodiments, the cytidine deaminase is human APOBEC, including hAPOBEC1 or hAPOBEC3. In some embodiments, the cytidine deaminase is human AID.

[0261] In some embodiments, the cytidine deaminase protein recognizes one or more target cytosine residues in the single-stranded bubble of the RNA duplex and converts them into uracil residues. In some embodiments, the cytidine deaminase protein recognizes a binding window on the single-stranded bubble of the RNA duplex. In some embodiments, the binding window comprises at least one target cytosine residue. In some embodiments, the binding window is in the range of about 3 bp to about 100 bp. In some embodiments, the binding window is in the range of about 5 bp to about 50 bp. In some embodiments, the binding window is in the range of about 10 bp to about 30 bp. In some embodiments, the binding window is about 1 bp, 2 bp, 3 bp, 5 bp, 7 bp, 10 bp, 15 bp, 20 bp, 25 bp, 30 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, or 100 bp.

[0262] In some embodiments, the cytidine deaminase protein comprises one or more deaminase domains. Without wishing to be bound by theory, it is expected that the deaminase domain is used to recognize one or more target cytosine (C) residues contained in the single-stranded bubble of the RNA duplex and convert them into uracil (U) residues. In some embodiments, the deaminase domain comprises an active center. In some embodiments, the active center comprises a zinc ion. In some embodiments, the amino acid residues in or near the active center interact with one or more nucleotides at 5' of the target cytosine residue. In some embodiments, the amino acid residues in or near the active center interact with one or more nucleotides at 3' of the target cytosine residue.

[0263] In some embodiments, the cytidine deaminase comprises the human APOBEC1 full protein (hAPOBEC1) or its deaminase domain (hAPOBEC1-D) or its C-terminal truncated form (hAPOBEC-T). In some embodiments, the cytidine deaminase is an APOBEC family member homologous to hAPOBEC1, hAPOBEC-D, or hAPOBEC-T. In some embodiments, the cytidine deaminase comprises the human AID1 full protein (hAID) or its deaminase domain (hAID-D) or its C-terminal truncated form (hAID-T). In some embodiments, the cytidine deaminase is an AID family member homologous to hAID, hAID-D, or hAID-T. In some embodiments, hAID-T is an hAID with a C-terminal truncation of approximately 20 amino acids.

[0264] In some embodiments, the cytidine deaminase comprises the wild-type amino acid sequence of a cytidine deaminase. In some embodiments, the cytidine deaminase comprises one or more mutations in the cytidine deaminase sequence such that the editing efficiency and / or substrate editing preference of the cytidine deaminase is altered according to specific needs.

[0265] In this article, "association" is taken in its broadest sense, covering the situation where two functional modules directly or indirectly (for example, through a linker) form a fusion protein, and also covering the situation where two functional modules are independent and bonded together by covalent bonds (such as disulfide bonds, etc.) or non-covalent bonds.

[0266] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. This is a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment can be inserted to achieve replication of the inserted segment. Typically, a vector is capable of replication when combined with appropriate control elements.

[0267] In some cases, the vector system comprises a single vector. Alternatively, the vector system comprises multiple vectors. The vector can be a viral vector.

[0268] In some cases, vectors include but are not limited to single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA or both; and other polynucleotide variants known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop, into which other DNA segments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which there is a virally derived DNA or RNA sequence for packaging into a virus (e.g., a retrovirus, a replication-defective retrovirus, adenovirus, a replication-defective adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Certain vectors are capable of autonomous replication in the host cells into which they are introduced (e.g., bacterial vectors and episomal mammalian vectors with a bacterial origin of replication). After being introduced into the host cell, other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell, thereby replicating together with the host genome. In addition, certain vectors can guide the expression of genes operably linked thereto. Such vectors are referred to herein as "expression vectors." Vectors that are expressed in eukaryotic cells and vectors that cause expression in eukaryotic cells may be referred to herein as “eukaryotic expression vectors.” Common expression vectors useful in recombinant DNA techniques are often in the form of plasmids.

[0269] Recombinant expression vector can be suitable for comprising nucleic acid of the present disclosure in the form of expressing nucleic acid in host cell, this means that recombinant expression vector comprises one or more regulatory elements, and described regulatory elements can be selected according to the host cell to be used for expression, and described nucleic acid is operably connected to nucleic acid sequence to be expressed.In recombinant expression vector, " operably connected " is intended to refer to that target nucleotide sequence is connected to regulatory elements in the mode that allows nucleotide sequence expression (for example, in in vitro transcription / translation system or when vector is introduced into host cell in host cell).Advantageous vector includes slow virus and adeno-associated virus, and the type of these vectors can also be selected to target specific type of cell.

[0270] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES) and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences). Such regulatory elements are described in, for example, Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of nucleotide sequences in many types of host cells and those that direct expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in target desired tissues such as muscle, neurons, bones, skin, blood, specific organs (e.g., liver, pancreas) or specific cell types (e.g., lymphocytes). Regulatory elements can also direct expression in a time-dependent manner, such as directing expression in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific. In some embodiments, the vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5 or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5 or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5 or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer) [see, e.g., Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" also encompasses enhancer elements such as WPRE; CMV enhancer; R-U5' segment in the LTR of HTLV-1 (Mol. Cell. Biol., Vol. 8(1), No. 466-472, 1988); SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), No. 1527-31, 1981). It will be understood by those skilled in the art that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the desired expression level, etc.Vectors can be introduced into host cells to produce transcripts, proteins or peptides encoded by nucleic acids described herein, including fusion proteins or peptides (e.g., clustered regularly interspaced short palindromic repeats (CRISPR) transcripts, proteins, enzymes, mutant forms thereof, fusion proteins thereof, etc.). Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of such vectors can also be selected to target specific types of cells. In specific embodiments, bicistronic vectors are used for guide RNA and (optionally modified or mutated) CRISPR enzymes (e.g., Cas12o).

[0271] Vectors can be designed to express CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, CRISPR transcripts can be expressed in bacterial cells such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are further discussed in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990).

[0272] The vector can be introduced and propagated in a prokaryotic organism or a prokaryotic cell. In some embodiments, the prokaryotic organism is used to amplify a copy of the vector to be introduced into a eukaryotic cell, or as an intermediate vector in the production of the vector to be introduced into a eukaryotic cell (e.g., a plasmid is a part of a viral vector packaging system). In some embodiments, the prokaryotic organism is used to amplify a copy of the vector and express one or more nucleic acids, such as providing a source of one or more proteins for delivery to a host cell or host organism. Protein expression in prokaryotes is most commonly carried out in Escherichia coli with a vector containing a constitutive or inducible promoter that guides the expression of a fusion protein or a non-fusion protein. A fusion vector adds many amino acids to the protein encoded therein, such as to the amino terminus of a recombinant protein. Such fusion vectors can provide one or more purposes, such as: (i) increasing the expression of a recombinant protein; (ii) increasing the solubility of a recombinant protein; and (iii) assisting in the purification of a recombinant protein by acting as a ligand in affinity purification. In some embodiments, the vector is a yeast expression vector. Examples of vectors for expression in the yeast Saccharomyces cerivisae include pYepSec1 (Baldari et al., 1987. EMBO J. 6: 229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30: 933-943), pJRY88 (Schultz et al., 1987. Gene 54: 113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.). In some embodiments, the vector drives protein expression in insect cells using a baculovirus expression vector. Baculovirus vectors useful for expressing proteins in cultured insect cells (eg, SF9 cells) include the pAc series (Smith et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).

[0273] In some embodiments, the vector can use a mammalian expression vector to drive the expression of one or more sequences in mammalian cells. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329: 840) and pMT2PC (Kaufman et al., 1987. EMBO J. 6: 187-195). When used in mammalian cells, the control function of the expression vector is generally provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus, cytomegalovirus, simian virus, and other promoters disclosed herein and known in the art. For other suitable expression systems for prokaryotic and eukaryotic cells, see, for example, Sambrook et al., MOLECULAR CLONING: ALABORATORY MANUAL. 2nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, Chapter 16 and Chapter 17.

[0274] In some embodiments, the recombinant mammalian expression vector is capable of preferentially directing expression of the nucleic acid in a specific cell type (eg, tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific; Pinkert et al., 1987. Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43:235-275), in particular promoters of T-cell receptors (Winoto and Baltimore, 1989. EMBO J. 8:729-733) and immunoglobulins (Baneiji et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), neuron-specific promoters (e.g., neurofilament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreas-specific promoters (Edlund et al., 1985. Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoter; U.S. Patent No. 4,873,316 and European Application Publication No. 264,166).

[0275] In some embodiments, one or more vectors driving the expression of one or more elements of the nucleic acid targeting system are introduced into the host cell so that the expression of the elements of the nucleic acid targeting system guides the formation of the nucleic acid targeting complex at one or more target sites. In some embodiments, a single promoter drives the expression of transcripts encoding Cas12o proteins and guide RNAs, and the transcripts are embedded in one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, Cas12o proteins and guide RNAs can be operably linked to the same promoter and expressed from the same promoter. In some embodiments, the vector comprises one or more insertion sites, such as restriction endonuclease recognition sequences (also referred to as "cloning sites"). In some embodiments, one or more insertion sites (e.g., about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. When using multiple different guide sequences, a single expression construct can be used to target multiple different corresponding target sequences within the cell by nucleic acid targeting activity. For example, a single vector may comprise about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences. In some embodiments, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more such vectors containing guide sequences may be provided and optionally delivered to cells. In some embodiments, the vector comprises a regulatory element operably connected to a coding sequence encoding Cas12o protein. Cas12o protein or one or more nucleic acid targeting guide RNAs may be delivered separately; and advantageously, at least one of these is delivered via a particle complex. Cas12o protein mRNA may be delivered before the guide RNA to allow time for expression of the Cas12o protein. Cas12o protein mRNA may be administered 1-12 hours (preferably about 2-6 hours) before administering the guide RNA. Alternatively, Cas12o protein mRNA and guide RNA may be administered together. Advantageously, a second booster dose of guide RNA can be administered 1-12 hours (preferably about 2-6 hours) after the initial administration of Cas12o protein mRNA + guide RNA. Additional administration of Cas12o protein mRNA and / or guide RNA may be useful for achieving the most effective genome modification level.

[0276] In some embodiments, the vector encodes a Cas12o protein comprising one or more nuclear localization sequences (NLSs), such as about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. More particularly, the vector comprises one or more NLSs that are not naturally present in the Cas12o protein. Most particularly, the NLS is present in the vector 5' and / or 3' of the Cas12o protein sequence. In some embodiments, the effector protein targeting RNA comprises about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus, and about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the carboxyl terminus, or a combination of these (e.g., 0 or at least one or more NLSs at the amino terminus and 0 or one or more NLSs at the carboxyl terminus). When more than one NLS is present, each can be selected independently of the other, such that a single NLS can be present in more than one copy and / or in combination with one or more other NLSs in one or more copies. In some embodiments, an NLS is considered to be near the N-terminus or C-terminus when the closest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus or C-terminus.

[0277] Non-limiting examples of NLSs include NLS sequences derived from the following: the NLS of the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV (SEQ ID NO.28); an NLS from a nucleoplasmic protein (e.g., the nucleoplasmic protein bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO.24)); an NLS having the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO.22), AVKRPAATKKAGQAKKKKLD (SEQ ID NO.23), KKTELQTTNAENKTKKL (SEQ ID NO.25), KRGINDRNFWRGENGRKTR (SEQ ID NO.26), RKSGKIAAIVVKRPRK (SEQ ID NO.27), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO.29). In some embodiments, the NLS sequence further comprises a c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO. 32) or RQRRNELKRSP (SEQ ID NO: 33), an hRNPA1M9 NLS comprising the amino acid sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO. 34); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 35) from the IBB domain of importin-α; the sequences VSRKRPRP (SEQ ID NO: 36) and PPKKARED (SEQ ID NO: 37) of myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 38) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 39) of mouse c-abl IV; the sequences DRLRR and PKQKKRK (SEQ ID NO: 31) of influenza virus NS1. NO: 41); the sequence of hepatitis virus delta antigen RKLKKKIKKL (SEQ ID NO: 42); the sequence of mouse Mx1 protein REKKKFLKRR (SEQ ID NO: 43); the sequence of human poly (ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 44); and the sequence of steroid hormone receptor (human) glucocorticoid RKCLQAGMNLEARKTKK (SEQ ID NO: 45). Typically, one or more NLSs have sufficient strength to drive the accumulation of detectable amounts of DNA / RNA-targeted Cas12o proteins in the eukaryotic cell nucleus.Typically, the strength of nuclear localization activity can be derived from the number of NLSs in the nucleic acid targeting effector protein, the specific NLS used, or a combination of these factors. In preferred embodiments of the Cas12o protein complex and system described herein, the codon-optimized Cas12o effector protein comprises an NLS attached to the C-terminus of the protein. In certain embodiments, other localization tags can be fused to the Cas12o protein, such as, but not limited to, localizing the Cas12o protein to a specific site in the cell, such as an organelle, such as mitochondria, plastids, chloroplasts, vesicles, Golgi apparatus, (nuclear or cell) membrane, ribosomes, nucleolus, ER, cytoskeleton, vacuole, centrosome, nucleosome, granule, centriole, etc.

[0278] In one embodiment, the Cas12o protein of the present disclosure comprises one or more NLSs at its N-terminus and / or C-terminus, preferably one NLS each at its N-terminus and C-terminus.

[0279] Codon optimization

[0280] In the case where the effector protein is to be administered as a nucleic acid, the present disclosure contemplates the use of codon-optimized Cas12 proteins, more particularly nucleic acid sequences (and optionally protein sequences) encoding Cas12o. An example of a codon-optimized sequence, in this case, is a sequence optimized for expression in eukaryotes such as humans (i.e., optimized for expression in humans), or a sequence optimized for another eukaryote, animal, or mammal as discussed herein. Although this is preferred, it should be understood that other examples are also possible, and codon optimization for host species other than humans or codon optimization for specific organs is known. In some embodiments, the enzyme coding sequence encoding the Cas protein targeting DNA / RNA is codon-optimized for expression in specific cells such as eukaryotic cells. Eukaryotic cells can be eukaryotic cells of specific organisms such as plants or mammals, or eukaryotic cells derived from specific organisms such as plants or mammals, including but not limited to humans or non-human eukaryotes or animals or mammals as discussed herein, such as mice, rats, rabbits, dogs, livestock, or non-human mammals or primates. In some embodiments, methods for modifying human germline genetic traits and / or methods for modifying the genetic traits of animals that may cause human suffering without any substantial medical benefit to humans or animals, as well as animals produced by such methods, may be excluded. Generally speaking, codon optimization refers to a method of modifying a nucleic acid sequence in a target host cell to enhance expression by replacing at least one codon of the native sequence (e.g., about or greater than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Various species exhibit specific biases for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is generally correlated with the efficiency of translation of messenger RNA (mRNA), which is believed to be dependent, among other things, on the properties of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons that are most frequently used in peptide synthesis. Therefore, genes can be customized based on codon optimization for optimal gene expression in a given organism. Codon usage tables are readily available, for example, in the "Codon Usage Database" at www.kazusa.orjp / codon / , and these tables can be modified in a variety of ways. See Nakamura, Y. et al., "Codon usage tabulated from the international DNA Sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000).Computer algorithms for codon optimization of specific sequences for expression in specific host cells are also available, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more or all codons) in the sequence encoding the Cas protein targeting DNA / RNA correspond to the most commonly used codons for a specific amino acid. For codon usage in yeast, reference is made to the online yeast genome database available at www.yeastgenome.org / community / codon_usage.shtml, or Codon selection in yeast, Bennetzen and Hall, J Biol Chem. 1982 Mar 25; 257(6): 3026-31. For codon usage in plants including algae, reference is made to Codon usage in higher plants, green algae, and cyan obacteria, Campbell and Gowri, Plant Physiol. 1990 January; 92(1): 1-11.; and Codon usage in plant genes, Murray et al., Nucleic Acids Res. 1989 January 25; 17(2): 477-98; or Selection on the codon bias of chloroplast and cyanell egenes in different plant and algal lineages, Morton BR, J Mol Evol. 1998 April; 46(4): 449-59. In some embodiments, the polynucleotide encoding Cas12o has been codon-optimized and expressed in a corresponding host cell (e.g., for mammalian cells, more specifically, human cells, such as HSC or iPSC).

[0281] connector

[0282] The term "linker" refers to a molecule that connects proteins to form a fusion protein. Typically, such molecules have no specific biological activity other than connecting or maintaining a minimum distance or other spatial relationship between proteins. However, in embodiments, the linker can be selected to affect some properties of the linker and / or fusion protein, such as the folding, net charge, or hydrophobicity of the linker.

[0283] Suitable linkers for use in the methods herein include straight or branched carbon linkers, heterocyclic carbon linkers, or peptide linkers. However, as used herein, linkers may also be covalent bonds (carbon-carbon bonds or carbon-heteroatom bonds). In embodiments, linkers may be chemical moieties, which may be monomers, dimers, multimers, or polymers. Preferably, the linker comprises amino acids. Typical amino acids in flexible linkers include Gly, Asn, and Ser. Therefore, in specific embodiments, the linker comprises a combination of one or more of Gly, Asn, and Ser amino acids. Other near-neutral amino acids, such as Thr and Ala, may also be used for linker sequences. Exemplary, GlySer linkers GGS, GGGS, or GSG may be used. GGS, GSG, GGGS, or GGGGS linkers may be repeated multiple times (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or even more, e.g., (GGS)3 (SEQ ID NO: 10), (GGGGS)3 (SEQ ID NO: 15)) to provide a suitable length.

[0284] Prime editing

[0285] In one embodiment, the present disclosure provides compositions and systems, which may include Cas12o protein or its catalytically inactive form, one or more crRNA or guide molecules, and reverse transcriptase. The system can be used to insert a donor polynucleotide into a target polynucleotide. In some embodiments, the composition or system includes a catalytically inactive Cas12o protein, a reverse transcriptase that associates with or can otherwise form a complex with the Cas12o protein, and a crRNA that can form a complex with the Cas12o protein and guide the site-specific binding of the complex to the target sequence of the target polynucleotide, the crRNA further comprising a donor sequence for inserting the target polynucleotide. In some cases, the catalytically inactive Cas12o protein can be a nickase, such as a DNA nickase. In some cases, the Cas12o protein has one or more mutations.

[0286] Cas12o protein can be associated with reverse transcriptase.The reverse transcriptase domain can be a reverse transcriptase or a fragment thereof.In some cases, the reverse transcriptase is human immunodeficiency virus (HIV) RT, avian myoblast virus (AMV) RT, Moloney murine leukemia virus (M-MLV) RT, group II intron RT, group II intron-like RT, or chimeric RT.In one embodiment, RT includes a modified form of these RTs, such as avian myoblast virus (AMV) RT, Moloney murine leukemia virus (M-MLV) RT or an engineered variant of human immunodeficiency virus (HIV) RT (see, e.g., Anzalone et al., Search-and-replace genome editing without double-strand breaks or donor DNA, Nature. December 2019; 576(7785): 149-157).

[0287] In some embodiments, the compositions and systems may include a Cas12o protein or a variant thereof disclosed herein; a reverse transcriptase (RT) polypeptide linked to or otherwise capable of forming a complex with the Cas12o protein or a variant thereof; and a crRNA capable of forming a complex with the Cas12o protein or a variant thereof and comprising: a crRNA or a guide sequence capable of guiding site-specific binding of the Cas12o protein or a variant thereof to a target sequence of a target polynucleotide; a 3' binding site region capable of binding to the upstream cleavage strand of the target polynucleotide; and an RT template sequence encoding an extension sequence, wherein the extension sequence comprises a variant region and a 3' homologous sequence capable of hybridizing to the downstream cleavage strand of the target polynucleotide.

[0288] The reverse transcriptase domain can be a reverse transcriptase or a fragment thereof. Various reverse transcriptases (RT) can be used for alternative embodiments of the present disclosure, including prokaryotic and eukaryotic RT, provided that these RTs work in the host and produce donor polynucleotide sequences from the RNA template. If desired, the nucleotide sequence of natural RT can be modified, for example, using known codon optimization techniques to modify, so as to optimize expression in the desired host. Reverse transcriptase (RT) is an enzyme for producing complementary DNA (cDNA) from an RNA template, a process known as reverse transcription. In one embodiment, the RT domain of a reverse transcriptase is used in the present disclosure. This domain may only include the DNA polymerase activity that depends on RNA. In some instances, the RT domain is non-mutagenic, i.e., it will not cause a mutation in the donor polynucleotide (e.g., in the reverse transcriptase process). In some cases, in some instances, the RT domain may be a non-reverse transcriptase RT, such as a viral RT or human endogenous RT. In some instances, the RT domain may be a reverse transcriptase RT or a DGR RT. In some instances, RT may be less mutagenic than corresponding wild-type RT. In one embodiment, the RT herein is not mutagenic. The reverse transcriptase may be fused to the C-terminus of Cas12o or a variant thereof. Alternatively or additionally, the reverse transcriptase may be fused to the N-terminus of Cas12o or a variant thereof. The fusion may be performed via a linker. In some instances, the reverse transcriptase may be M-MLV reverse transcriptase or a variant thereof. The M-MLV reverse transcriptase variant may comprise one or more mutations.

[0289] Reverse transcriptase domain

[0290] One or more functional domains can be one or more reverse transcriptase domains. In one embodiment, the system comprises an engineered system for modifying a target polynucleotide, the system comprising: a Cas12o protein or a CRISPR-associated Cas12o protein or a variant thereof (e.g., dCas12o); a reverse transcriptase (RT) domain; an RNA template comprising or encoding a donor polynucleotide of a target sequence to be inserted into the target polynucleotide; and crRNA.

[0291] Retrotranscript

[0292] In one embodiment, the donor template for homologous recombination is produced by using the self-starting RNA template for reverse transcription.A non-limiting example of a self-starting reverse transcription system is a reverse transcriptase system.Term " reverse transcriptase " means a kind of genetic element, and the component of its encoding makes it possible to synthesize single-stranded DNA (msDNA) and the reverse transcriptase that branched RNA is connected.In one embodiment, the reverse transcriptase domain is a reverse transcriptase RT domain.In one embodiment, the RNA template encoding is the reverse transcriptase RNA template identified and reversely transcribed by the reverse transcriptase reverse transcriptase domain.Reverse transcriptase is all conservative in many bacterial species, is the relatively unknown efficient reverse transcription system of function.The reverse transcriptase system is made up of reverse transcriptase RT albumen and msr and msd transcript, and msr and msd transcript serve as primer and template sequence respectively. All components of the reverse transcription subsystem are expressed as single transcripts from a single open reading frame, and the transcript includes msr-msd and encodes reverse transcriptase RT protein (Lampson et al., 2005, Retrons, msDNA, and the bacterial genome. Cytogenet Genome Res 110:491-499). The msr element ORF of the reverse transcriptase provides the RNA portion of the msDNA molecule, and the msd element ORF provides the DNA portion of the msDNA molecule. The main transcript from the msr-msd region is considered to serve as the template and primer for producing msDNA. Use its 2'-OH group to start synthesizing msDNA from the internal rG residues of the RNA transcript. Md or msr can also be modified to allow the RNA template of the coding donor polynucleotide to be inserted in msd without changing the function or production of msDNA. The RNA template of the coding donor polynucleotide sequence can be any length, but is preferably less than about 5kb nucleotides, or also less than about 2kb, or also less than 500 bases, condition is to produce the msDNA product.

[0293] Topoisomerase

[0294] Topoisomerase is a class of enzymes that change the topological state of DNA by breaking and reconnecting nucleic acid chains. In some cases, the topoisomerase may be a DNA topoisomerase, which controls and changes the topological state of DNA during transcription and catalyzes the transient breakage and reconnection of single-stranded DNA, thereby allowing the chains to pass through each other, thereby changing the topological structure of DNA. In some embodiments, one or more functional domains may be one or more topoisomerase domains. In one embodiment, the engineered system for modifying a target polynucleotide comprises: a Cas12o protein; a topoisomerase domain; and a nucleic acid template comprising or encoding a donor polynucleotide of a target sequence to be inserted into the target polynucleotide. In some instances, two or more of the following may form a complex: a Cas12o protein; a topoisomerase domain; and a nucleic acid template. In some instances, two or more of the following may be included in a fusion protein: a Cas12o protein; a topoisomerase domain.

[0295] In one embodiment, the topoisomerase domain can connect the donor polynucleotide to the target polynucleotide. Connection can be achieved by sticky end or blunt end connection. In one example, the donor polynucleotide may include an overhang, which includes a sequence complementary to the region of the target polynucleotide. The example of connecting the donor polynucleotide to the target polynucleotide includes the example of TOPO cloning, for example, "The Technology Behind TOPO Cloning," described in www.thermofisher.com / us / en / home / life-science / cloning / topo / topo-resources / the-technology-behind-topo-cloning.html. In one embodiment, the topoisomerase domain can associate with the donor polynucleotide. For example, the topoisomerase domain is covalently linked to the donor polynucleotide.

[0296] Examples of topoisomerases include type I (including type IA and type IB topoisomerases), which cut a single strand of a double-stranded nucleic acid molecule; and type II topoisomerases (e.g., gyrase), which cut both strands of a double-stranded nucleic acid molecule. In some instances, the topoisomerase is a DNA topoisomerase I, such as vaccinia virus topoisomerase I. The topoisomerase can be preloaded with a donor polynucleotide. The vaccinia virus topoisomerase may require a target comprising a 5'-OH group.

[0297] Phosphatase

[0298] The system herein may also comprise a phosphatase domain. A phosphatase is an enzyme that is capable of removing a phosphate group from a molecule, e.g., a nucleic acid such as DNA. Examples of phosphatases include calf intestinal phosphatase, shrimp alkaline phosphatase, antarctic phosphatase, and APEX alkaline phosphatase.

[0299] In some embodiments, the 5'-OH group in the target polynucleotide can be produced by a phosphatase. A topoisomerase compatible with a 5' phosphate target can be used to produce a stably loaded intermediate. In some cases, a Cas12o protein or a CRISPR-related Cas12o protein that leaves a 5'OH after cutting the target polynucleotide can be used. In some cases, the phosphatase domain can be associated with (e.g., fused to) the Cas12o protein. The phosphatase domain may be able to produce an -OH group at the 5' end of the target polynucleotide. The phosphatase can be separated from other components in the system, for example, as a separate protein, delivered on a carrier separated from other components.

[0300] polymerase

[0301] The system herein may also include a polymerase domain. A polymerase refers to an enzyme that synthesizes a nucleic acid chain. The polymerase may be a DNA polymerase or an RNA polymerase.

[0302] In one embodiment, the system includes an engineered system for modifying a target polynucleotide, the engineered system comprising: a Cas12o protein or a CRISPR-associated Cas12o protein; a DNA polymerase domain; and a DNA template comprising a donor polynucleotide of a target sequence to be inserted into the target polynucleotide. In some instances, two or more of the following may form a complex: a Cas12o protein; a DNA polymerase domain; and a DNA template. In some instances, two or more of the following are contained in a fusion protein: a Cas12o protein; a DNA polymerase domain. For example, a Cas12o protein or a CRISPR-associated Cas12o protein and a DNA polymerase domain may be contained in a fusion protein.

[0303] In one embodiment, the system may include a Cas12o protein or a CRISPR-associated Cas12o protein (or a variant thereof, such as a dCas12o protein or a CRISPR-associated Cas12o protein or a CRISPR-associated Cas12o nickase) and a DNA polymerase (e.g., phi29, T4, T7 DNA polymerase). The system may also include a single-stranded DNA or double-stranded DNA template. The DNA template may include i) a first sequence homologous to the target site of the Cas12o protein on the target polynucleotide, and / or ii) a second sequence homologous to another region of the target polynucleotide. In one embodiment, the template can be a synthetic single-stranded or PCR-generated DNA molecule (optionally end-protected by modified nucleotides), or a viral genome (e.g., AAV). In another embodiment, the template is generated using a reverse transcriptase. When the system is delivered into a cell, an endogenous DNA polymerase in the cell can be used. Alternatively or in addition, an exogenous DNA polymerase can be expressed in the cell. The DNA template can be end-protected by one or more modified nucleotides, or comprise a portion of a viral genome.

[0304] Examples of DNA polymerases include Taq, Tne(exo-), Tma(exo-), Pfu(exo-), Pwo(exo-), Thermoanaerobacter thermohydrosulf uricus DNA polymerase, Thermococcus litoralis DNA polymerase I, Escherichia coli DNA polymerase I, Taq DNA polymerase I, Tth DNA polymerase I, Bacillus stearothermophilus (Bst) DNA polymerase I, Escherichia coli DNA polymerase III, bacteriophage T5 DNA polymerase, bacteriophage M2 DNA polymerase, bacteriophage T4 DNA polymerase, bacteriophage T7 DNA polymerase, bacteriophage phi29 DNA polymerase, bacteriophage PRD1 DNA polymerase, bacteriophage phi15 DNA polymerase, bacteriophage phi21 DNA polymerase, bacteriophage PZE DNA polymerase, bacteriophage PZA DNA polymerase, bacteriophage NfDNA polymerase, bacteriophage M2Y DNA polymerase, bacteriophage B103 DNA polymerase, bacteriophage SF5 DNA polymerase, bacteriophage GA-1 DNA polymerase, bacteriophage Cp-5 DNA polymerase, bacteriophage Cp-7 DNA polymerase, bacteriophage PR4 DNA polymerase, bacteriophage PR5 DNA polymerase, bacteriophage PR722 DNA polymerase and bacteriophage L17 DNA polymerase.

[0305] Systems and complexes

[0306] The present disclosure also provides a nucleic acid targeting system. Such systems can be used for targeting, modifying and otherwise manipulating nucleic acids. In one embodiment, the system comprises Cas12o protein or CRISPR-related Cas12o protein and one or more crRNA or guide RNA. Cas12o protein or CRISPR-related Cas12o protein may have nuclease activity, for example, capable of cutting DNA or RNA. Cas12o protein or CRISPR-related Cas12o protein may have nickase activity, for example, capable of producing single-strand breaks on double-stranded nucleic acids such as dsDNA or dsRNA. Cas12o protein or CRISPR-related Cas12o protein may be in a dead form, for example, with nickase activity, or without nuclease or nickase activity. In one embodiment, the system further comprises one or more functional domains, for example, nucleotide deaminase, reverse transcriptase, non-LTR retrotransposon (and encoded protein), polymerase, elements (and encoded protein) for generating diversity. In some instances, the system further comprises one or more donor polynucleotides. The donor polynucleotide can be inserted into the target polynucleotide by the system. The donor polynucleotide can be contained in or encoded by the nucleic acid template. In some embodiments, two or more components in the present system can form a complex. For example, the components are independent molecules that interact directly or indirectly with each other. Some two or more components in the present system can be contained in a fusion protein.

[0307] The term "target sequence" refers to a sequence to which crRNA is designed to be complementary, wherein the hybridization between the target sequence and the spacer sequence promotes the formation of a complex targeting DNA or RNA. Complete complementarity is not necessarily required, as long as there is enough complementarity to cause hybridization and promote the formation of a complex targeting nucleic acid. The target sequence may comprise an RNA polynucleotide. In one embodiment, the target sequence is located in the nucleus or cytoplasm of the cell. In one embodiment, the target sequence may be in an organelle of a eukaryotic cell, for example, a mitochondria or a chloroplast. A sequence or template that can be used to recombine into a targeting locus comprising a target sequence is referred to as an "editing template" or "editing sequence". In various aspects of the present disclosure, an exogenous template may be referred to as an editing template. In one aspect, recombination is homologous recombination.

[0308] In one embodiment, the formation of a nucleic acid-targeting complex (comprising a guide RNA that hybridizes to a target sequence and complexes with one or more nucleic acid-targeting effector proteins) results in the cleavage of one or both nucleic acid strands in or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of) the target sequence. In one embodiment, one or more vectors that drive the expression of one or more elements of a nucleic acid-targeting system are introduced into a host cell such that expression of these elements of the nucleic acid-targeting system directs the formation of a nucleic acid-targeting complex at one or more target sites. For example, a nucleic acid-targeting effector protein and a crRNA or guide RNA can each be operably linked to a separate regulatory element on a separate vector. Alternatively, two or more of these elements, expressed from the same or different regulatory elements, can be combined in a single vector, with one or more additional vectors providing any components of the nucleic acid-targeting system not contained in the first vector. The nucleic acid-targeting system elements combined in a single vector can be arranged in any suitable orientation, such as one element positioned 5' ("upstream") relative to a second element or 3' ("downstream") relative to the second element. The coding sequence of one element can be located on the same strand or the opposite strand of the coding sequence of the second element and oriented in the same or opposite directions. In one embodiment, a single promoter drives expression of a transcript encoding a nucleic acid-targeting effector protein and a guide RNA embedded within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In one embodiment, the nucleic acid-targeting effector protein and guide RNA are operably linked to and expressed from the same promoter.

[0309] Donor polynucleotide

[0310] In one embodiment, the compositions and systems herein may comprise one or more nucleic acid templates. In some cases, the nucleic acid template may comprise one or more polynucleotides. In some cases, the nucleic acid template may comprise the coding sequence of one or more polynucleotides. The nucleic acid template may be an RNA template. The nucleic acid template may be a DNA template.

[0311] Donor polynucleotides can be used to edit target polynucleotides. In some cases, the donor polynucleotides include one or more mutations to be introduced into the target polynucleotides. Examples of such mutations include substitutions, deletions, insertions, or combinations thereof. Mutations can cause open reading frame shifts on target polynucleotides. In some cases, donor polynucleotides change the stop codons in the target polynucleotides. For example, donor polynucleotides can correct premature stop codons. Correction can be achieved by deleting the stop codons or introducing one or more mutations into the stop codons. In other exemplary embodiments, donor polynucleotides solve loss-of-function mutations, deletions, or translocations that may occur, for example, in certain disease situations, by inserting or restoring a functional copy of a gene or its functional fragments or functional regulatory sequences or functional fragments of regulatory sequences. A functional fragment refers to a gene that is less than a complete copy, in a manner that provides enough nucleotide sequence to restore the functionality of a wild-type gene or a non-coding regulatory sequence (e.g., a sequence encoding a long non-coding RNA). In certain exemplary embodiments, the system disclosed herein can be used to replace a single allele of a defective gene or its defective fragment. In another exemplary embodiment, the system disclosed herein can be used to replace two alleles of a defective gene or a defective gene fragment. A "defective gene" or "defective gene fragment" is a gene or gene portion that, when expressed, fails to produce a functional protein or non-coding RNA as the corresponding wild-type gene. In certain exemplary embodiments, these defective genes may be associated with one or more disease phenotypes. In certain exemplary embodiments, the defective gene or gene fragment is not replaced, but the systems described herein are used to insert a donor polynucleotide encoding a gene or gene fragment that compensates for or overrides the expression of the defective gene, thereby eliminating the cellular phenotype associated with the expression of the defective gene or altering it to a different or desired cellular phenotype.

[0312] In one embodiment, the donor polynucleotide may include but is not limited to a gene or gene fragment, a coded protein to be expressed or an RNA transcript, a regulatory element, a repair template, etc. According to the present disclosure, the donor polynucleotide may include left-end and right-end sequence elements that function together with the transposition component that mediates insertion. In some cases, the donor polynucleotide manipulates the splice site on the target polynucleotide. In some instances, the donor polynucleotide destroys the splice site. Destruction can be achieved by inserting the polynucleotide into the splice site and / or introducing one or more mutations into the splice site. In some instances, the donor polynucleotide can restore the splice site. For example, the polynucleotide may include a splice site sequence.

[0313] The size of the donor polynucleotide to be inserted can be from 10 base pairs or nucleotides to 50 kb in length, for example, 50 to 40k, 100 and 30k, 100 to 10000, 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 600 to 1800, 1700 to 1900, 1800 to 2000 base pairs (bp) or nucleotides in length.

[0314] deliver

[0315] The present disclosure also provides a delivery system for introducing the components of the systems and compositions herein into cells, tissues, organs or organisms. The delivery system may include one or more delivery vehicles and / or cargo. In some embodiments, the components of the CRISPR-Cas system can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein RNA. For example, the Cas12o protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein. The guide can be delivered as a DNA encoding polynucleotide or RNA. All possible combinations, including mixed delivery forms, are envisioned.

[0316] In some aspects, the present disclosure provides methods comprising delivering one or more polynucleotides, eg, one or more vectors as described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell.

[0317] In some embodiments, one or more vectors that drive the expression of one or more elements of a nucleic acid-targeting system are introduced into a host cell such that expression of the elements of the nucleic acid-targeting system directs the formation of a nucleic acid-targeting complex at one or more target sites. For example, a nucleic acid-targeting effector enzyme and a nucleic acid-targeting guide RNA can each be operably linked to separate regulatory elements on separate vectors. RNA of the nucleic acid-targeting system can be delivered to an animal or mammal transgenic for a nucleic acid-targeting effector protein, for example, an animal or mammal that constitutively, inducibly, or conditionally expresses a nucleic acid-targeting effector protein; or an animal or mammal that otherwise expresses a nucleic acid-targeting effector protein or has cells containing the nucleic acid-targeting effector protein, for example, by previously administering thereto one or more vectors encoding and expressing the nucleic acid-targeting effector protein in vivo. Alternatively, two or more elements expressed by the same or different regulatory elements can be combined in a single vector, with one or more additional vectors providing any components of the nucleic acid-targeting system not contained in the first vector. The nucleic acid-targeting system elements combined in a single vector can be arranged in any suitable orientation, for example, with one element positioned 5' relative to ("upstream") a second element or 3' relative to ("downstream") a second element. The coding sequence of one element may be located on the same or opposite strand of the coding sequence of the second element and oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding nucleic acid-targeting effector proteins and nucleic acid-targeting guide RNAs, and the transcripts are embedded in one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the nucleic acid-targeting effector protein and the nucleic acid-targeting guide RNA can be operably linked to and expressed from the same promoter. Delivery vehicles, vectors, particles, nanoparticles, formulations, and components thereof for expressing one or more elements of a nucleic acid-targeting system are as used in WO 2014 / 093622 (PCT / US2013 / 074667). In some embodiments, the vector comprises one or more insertion sites, such as restriction endonuclease recognition sequences (also referred to as "cloning sites"). In some embodiments, one or more insertion sites (e.g., about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. When multiple different guide sequences are used, a single expression construct can be used to target the nucleic acid targeting activity to multiple different corresponding target sequences within a cell. For example, a single vector can contain about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences. In some embodiments, about or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more such guide sequence-containing vectors can be provided and optionally delivered to a cell.In some embodiments, the vector comprises a regulatory element operably linked to an enzyme coding sequence encoding a nucleic acid-targeting effector protein. The nucleic acid-targeting effector protein or one or more nucleic acid-targeting guide RNAs can be delivered separately; and advantageously, at least one of these is delivered via a particle complex. The nucleic acid-targeting effector protein mRNA can be delivered before the nucleic acid-targeting guide RNA to allow time for expression of the nucleic acid-targeting effector protein. The nucleic acid-targeting effector protein mRNA can be administered 1-12 hours (preferably about 2-6 hours) before administering the nucleic acid-targeting guide RNA. Alternatively, the nucleic acid-targeting effector protein mRNA and the nucleic acid-targeting guide RNA can be administered together. Advantageously, a second booster dose of the guide RNA can be administered 1-12 hours (preferably about 2-6 hours) after the initial administration of the nucleic acid-targeting effector protein mRNA+guide RNA. Other administrations of nucleic acid-targeting effector protein mRNA and / or guide RNA may be useful for achieving the most effective level of genome modification.

[0318] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding nucleic acid targeting system components to cells in culture or in host organisms. Non-viral vector delivery systems include DNA plasmids, RNA (transcripts of vectors such as those described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles such as liposomes. Viral vector delivery systems include DNA and RNA viruses that have free or integrated genomes after delivery to cells. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel and Felgner, TIBTECH 11:211-217 (1993); Mitani and Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer and Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.) (1995); and Yu et al., Gene Therapy 1: 13-26 (1994).

[0319] Non-viral delivery methods for nucleic acids include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycationic or lipid:nucleic acid conjugates, naked DNA, artificial virions, and agent-enhanced DNA uptake. Lipofection is described in, for example, U.S. Pat. Nos. 5,049,386, 4,946,787; and 4,897,355, and lipofection reagents are sold commercially (e.g., Transfectam TM and Lipofectin TM Cationic lipids and neutral lipids suitable for efficient receptor recognition lipofection of polynucleotides include those of Felgner, WO 91 / 17424; WO 91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration).

[0320] Plasmid delivery involves cloning the guide RNA into a plasmid expressing the CRISPR-Cas protein and transfecting the DNA in cell culture. Plasmid backbones are commercially available and do not require specific equipment. They have the advantage of modularity and can carry CRISPR-Cas coding sequences of different sizes (including sequences encoding larger proteins) as well as selection markers. At the same time, the advantage of plasmids is that they can ensure transient but sustained expression. However, the delivery of plasmids is not direct, making the in vivo efficiency generally low. Sustained expression can also be disadvantageous because it can increase off-target editing. In addition, excessive accumulation of CRISPR-Cas proteins can be toxic to cells. Finally, plasmids always have the risk of random integration of dsDNA in the host genome, more particularly considering the risk of generating double-strand breaks (on-target and off-target).

[0321] Preparation of lipid:nucleic acid complexes, including targeted liposomes, such as immunolipid complexes, is well known to those skilled in the art (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad ... Res. 52:4817-4820 (1992); U.S. Pat. Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787). This will be discussed in more detail below.

[0322] The system based on RNA or DNA virus is used to deliver nucleic acid and utilizes the specific cell in virus targeting body and the highly evolved process that viral payload is transported to nucleus.Viral vector can be directly applied to patient (in vivo), or they can be used for in vitro treatment cell, and the cell of modification can be optionally applied to patient (ex vivo).Conventional system based on virus can comprise retrovirus, slow virus, adenovirus, adeno-associated virus and herpes simplex virus vector, for gene transfer.Can be integrated into host genome with retrovirus, slow virus and adeno-associated virus gene transfer method, this usually causes the long-term expression of the transgenic of insertion.In addition, high transduction efficiency has been observed in many different cell types and target tissues.

[0323] Retroviral tropism can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors are composed of cis-acting long terminal repeats (LTRs) with the capacity to package up to 6-10 kb of foreign sequence. The minimal cis-acting LTRs are sufficient for replication and packaging of the vector, which is then used to integrate the therapeutic gene into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); PCT / US94 / 05700).

[0324] In the application of preferred transient expression, adenovirus-based systems can be used. Adenovirus-based vectors can achieve very high transduction efficiency in many cell types and do not require cell division. Utilizing such vectors, high titer and expression levels have been obtained. The vector can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to transduce cells with target nucleic acids, for example, in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy programs (see, for example, West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors is described in many publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat and Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989).

[0325] The present disclosure provides AAVs comprising or consisting essentially of an exogenous nucleic acid molecule encoding a CRISPR system, e.g., a plurality of cassettes comprising or consisting of a first cassette comprising or consisting essentially of a promoter, a nucleic acid molecule encoding a CRISPR-associated (Cas) protein (a putative nuclease or helicase protein), e.g., Cas12o, and a terminator, and one or more, advantageously up to the packaging size limit of the vector, e.g., a total of five cassettes (including the first cassette), comprising or consisting essentially of a promoter, a nucleic acid molecule encoding a guide RNA (gRNA), and a terminator (e.g., each cassette schematically represented as Promoter-gRNA1-Terminator, Promoter-gRNA2-Terminator...Promoter-gRNA(N)-Terminator). Terminator, where N is the number of the upper limit of the packaging size limit of the vector that can be inserted), or two or more separate rAAVs, each rAAV containing one or more cassettes of a CRISPR system, for example, a first rAAV containing a first cassette comprising or consisting essentially of: a promoter, a nucleic acid molecule encoding Cas, such as Cas and a terminator, and a second rAAV containing one or more cassettes, each comprising or consisting essentially of: a promoter, a nucleic acid molecule encoding a guide RNA (gRNA), and a terminator (e.g., each cassette is schematically represented as promoter-gRNA1-terminator, promoter-gRNA2-terminator...promoter-gRNA(N)-terminator, where N is the number of the upper limit of the packaging size limit of the vector that can be inserted). Alternatively, since Cas12o can process its own crRNA / gRNA, a single crRNA / gRNA array can be used for multiplex gene editing. Thus, instead of comprising multiple boxes to deliver gRNA, rAAV can contain a single box comprising or consisting essentially of a promoter, multiple crRNA / gRNA, and a terminator (e.g., schematically represented as promoter-gRNA1-gRNA2 ... gRNA (N)-terminator, where N is the number of the upper limit of the packaging size limit of the vector that can be inserted). See Zetsche et al., Nature Biotechnology 35, 31-34 (2017), which is incorporated herein by reference in its entirety. Since rAAV is a DNA virus, the nucleic acid molecules discussed herein about AAV or rAAV are advantageously DNA. In some embodiments, the promoter is advantageously the human synaptic protein 1 promoter (hSyn). Other methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, US20030087817, which is incorporated herein by reference.

[0326] In another embodiment, cocal vesiculovirus envelope pseudotyped retroviral vector particles are contemplated (see, for example, U.S. Patent Publication No. 20120164118, assigned to Fred Hutchinson Cancer Research Center). Cocal virus belongs to the genus Vesiculovirus and is the causative agent of vesicular stomatitis in mammals. Cocal virus was originally isolated from mites in Trinidad (Jonkers et al., Am. J. Vet. Res. 25: 236-242 (1964)) and has been identified in Trinidad, Brazil and Argentina from insects, cattle and horses for infection. Many vesiculoviruses that infect mammals have been isolated from naturally infected arthropods, indicating that they are vector-borne. In rural areas where endemic and laboratory-acquired viruses are present, people generally acquire vesiculovirus antibodies; human infection typically results in flu-like symptoms. The Cocal virus envelope glycoprotein shares 71.5% identity with VSV-G Indiana at the amino acid level, and phylogenetic comparisons of the vesiculovirus envelope genes show that Cocal virus is serologically distinct from, but most closely related to, the VSV-G Indiana strain of the vesiculovirus. Jonkers et al., Am. J. Vet. Res. 25:236-242 (1964) and Travassos da Rosa et al., Am. J. Tropical Med. & Hygiene 33:999-1006 (1984). Cocal vesiculovirus envelope pseudotyped retroviral vector particles may include, for example, lentiviral, alpharetroviral, betaretroviral, gammaretroviral, deltaretroviral, and epsilonretroviral vector particles, which may contain retroviral Gag, Pol, and / or one or more accessory proteins and the Cocal vesiculovirus envelope protein. In certain aspects of these embodiments, Gag, Pol, and accessory proteins are lentiviral and / or gammaretroviral.

[0327] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected when they are naturally present in a subject, and optionally reintroduced therein. In some embodiments, the transfected cells are taken from a subject. In some embodiments, the cells are derived from cells taken from a subject, such as a cell line. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Pancl, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV-434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells, K562 cells, Ku812, KCL22, KG1, KYO1, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCK II, MOR / 0.2R, MONO-MAC 6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic varieties thereof. Cell lines can be obtained from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.).

[0328] In certain embodiments, the transient expression and / or presence of one or more components of the AD-functionalized CRISPR system may be of interest, for example, to reduce off-target effects. In some embodiments, cells transfected with one or more vectors as described herein are used to establish new cell lines comprising one or more vector-derived sequences. In some embodiments, cells transiently transfected with components of the AD-functionalized CRISPR system as described herein (e.g., by transient transfection of one or more vectors, or transfection with RNA) and modified by the activity of the CRISPR complex are used to establish new cell lines comprising cells containing the modification but lacking any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors as described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0329] In some embodiments, it is envisioned that RNA and / or protein can be introduced directly into host cells. For example, the CRISPR-Cas protein can be delivered as encoding mRNA together with in vitro transcribed guide RNA. Such methods can reduce the time required to ensure the action of the CRISPR-Cas protein and further prevent long-term expression of CRISPR system components.

[0330] In some embodiments, the RNA molecules of the present disclosure are delivered in the form of liposomes or lipofectin formulations, and can be prepared by methods well known to those skilled in the art. Such methods are described, for example, in U.S. Patents 5,593,972, 5,589,466, and 5,580,859, which are incorporated herein by reference. Delivery systems specifically designed to enhance and improve the delivery of siRNA into mammalian cells have been developed (see, e.g., Shen et al., FEBS Let. 2003, 539: 111-114; Xia et al., Nat. Biotech. 2002, 20: 1006-1010; Reich et al., Mol. Vision. 2003, 9: 210-216; Sorensen et al., J. Mol. Biol. 2003, 327: 761-766; Lewis et al., Nat. Gen. 2002, 32: 107-108; and Simeoni et al., NAR 2003, 31, 11: 2717-2724) and can be applied to the present disclosure. siRNA has recently been successfully used to inhibit gene expression in primates (see, e.g., Tolentino et al., Retina 24(4): 660), which can also be applied to the present disclosure.

[0331] In fact, RNA delivery is a useful method for in vivo delivery. Cas12o, adenosine deaminase and guide RNA can be delivered to cells using liposomes or particles. Therefore, the delivery of CRISPR-Cas proteins (such as Cas12o), the delivery of adenosine deaminase (which can be fused with CRISPR-Cas proteins or adapter proteins) and / or the delivery of RNA of the present disclosure can be in RNA form and via microvesicles, liposomes or particles or nanoparticles. For example, Cas12o mRNA, adenosine deaminase mRNA and guide RNA can be packaged into liposome particles for in vivo delivery. Liposome transfection reagents, such as lipofectamine from Life Technologies and other reagents on the market, can effectively deliver RNA molecules to the liver. In some embodiments, lipid nanoparticles (LNP) are wrapped around Cas12o and its corresponding crRNA at the same time. In some embodiments, the lipid nanoparticles wrapped around Cas12o and / or its corresponding crRNA are given to subjects (such as people) in need by intravenous injection.

[0332] The delivery method of RNA also preferably includes delivery of RNA via particles (Cho, S., Goldberg, M., Son, S., Xu, Q., Yang, F., Mei, Y., Bogatyrev, S., Langer, R. and Anderson, D., Lipid-like nanoparticles for small interfering RNA delivery to endothelial cells, Advanced Functional Materials, 19: 3112-3118, 2010) or exosomes (Schroeder, A., Levins, C., Cortez, C., Langer, R. and Anderson, D., Lipid-based nanotherapeutics for siRNA delivery, Journal of Internal Medicine, 267: 9-21, 2010, PMID: 20059641). In fact, exosomes have been shown to be particularly useful in delivering siRNA, a system somewhat similar to the CRISPR system. For example, El-Andaloussi S et al. ("Exosome-mediated delivery of siRNA in vitro and in vivo." Nat Protoc. 2012 Dec;7(12):2112-26. doi:10.1038 / nprot.2012.131. Epub 2012 Nov 15) describe how exosomes are promising tools for drug delivery across different biological barriers and can be used for delivery of siRNA in vitro and in vivo. Their approach is to generate targeted exosomes by transfecting an expression vector containing an exosomal protein fused to a peptide ligand. The exosomes are then purified and characterized from the transfected cell supernatant, and RNA is then loaded into the exosomes. Delivery or administration according to the present disclosure can be performed using exosomes, particularly but not limited to the brain. Vitamin E (α-tocopherol) can be conjugated to CRISPR Cas and delivered to the brain with high-density lipoprotein (HDL), for example, in a manner similar to that used by Uno et al. (HUMAN GENE THERAPY 22:711-719 (June 2011)) for the delivery of short interfering RNA (siRNA) to the brain. Mice were infused via an Osmotic minipump (Model 1007D; Alzet, Cupertino, CA) filled with phosphate-buffered saline (PBS) or free TocsiBACE or Toc-siBACE / HDL and connected to a brain infusion kit 3 (Alzet).The brain infusion cannula was placed in the midline approximately 0.5 mm behind the anterior bregma for infusion into the dorsal third ventricle. Uno et al. found that as little as 3 nmol of Toc-siRNA containing HDL could induce a considerable reduction in target by the same ICV infusion method. In the present disclosure, similar doses of CRISPR Cas conjugated to α-tocopherol and co-administered with brain-targeted HDL may be considered for humans, for example, about 3 nmol to about 3 μmol of brain-targeted CRISPR Cas may be considered. Zou et al. (HUMAN GENETHERAPY 22:465-475 (April 2011)) described a lentiviral-mediated delivery method of short hairpin RNA targeting PKCγ for in vivo gene silencing in the spinal cord of rats. Zou et al. administered approximately 10 μl of recombinant lentivirus at a titer of 1×10 through an intrathecal catheter. 9 Transducing units (TU) / ml. In the present disclosure, similar doses of CRISPR Cas expressed in brain-targeted lentiviral vectors can be considered for humans, for example, at a titer of 1×10 9 Approximately 10-50 ml of brain-targeted CRISPR Cas in transducing units (TU) / ml of lentivirus.

[0333] In one embodiment, the delivery system can be used to introduce the components of the system and composition into plant cells. For example, electroporation, microinjection, aerosol injection of plant cell protoplasts, biolistic method, DNA particle bombardment and / or Agrobacterium-mediated transformation can be used to deliver the components to plants. Examples of plant methods and delivery systems include those described in Fu et al., Transgenic Res. 2000 February; 9(1): 11-9; Klein RM et al., Biotechnology. 1992; 24: 384-6; Casas AM et al., Proc Natl Acad Sci U SA. 1993 December 1; 90(23): 11212-11216; and U.S. Patent No. 5,563,055, Davey MR et al., Plant Mol Biol. 1989 September; 13(3): 273-85, which are incorporated herein by reference in their entirety.

[0334] The exemplary delivery compositions, systems, and methods described herein related to compositions or Cas12o proteins or CRISPR-associated Cas12o proteins are also applicable to functional domains and other components (e.g., other proteins and polynucleotides associated with Cas12o proteins or CRISPR-associated Cas12o proteins, such as reverse transcriptases, nucleotide deaminases, retrotransposons, donor polynucleotides, etc.).

[0335] cargo

[0336] The delivery system may include one or more goods. The goods may include one or more components of the systems and compositions herein. The goods may include one or more of the following: i) a plasmid encoding one or more protein components such as Cas12o protein or CRISPR-related Cas12o protein and / or functional domains in the composition and system; ii) a plasmid encoding one or more crRNAs, iii) one or more protein components such as Cas12o protein or CRISPR-related Cas12o protein and / or functional domains in the composition and system; iv) one or more guide RNAs; v) one or more protein components such as Cas12o protein or CRISPR-related Cas12o protein and / or functional domains in the composition and system; vi) any combination thereof. The one or more protein components may include nucleic acid-guided nucleases (e.g., Cas), reverse transcriptases, nucleotide deaminases, retrotransposon proteins, other functional domains, or any combination thereof.

[0337] In some examples, goods may include one or more protein components such as Cas12o protein or CRISPR-related Cas12o protein and / or functional domains and one or more (e.g., multiple) guide RNA plasmids in coding compositions and systems. In some cases, plasmids may also encode recombinant templates (e.g., for HDR). In one embodiment, goods may include mRNA encoding one or more protein components and one or more guide RNAs. In some examples, goods may include one or more protein components and one or more crRNA or guide RNAs, for example, in the form of ribonucleoprotein complexes (RNPs). Ribonucleoprotein complexes can be delivered by the methods and systems herein. In some cases, ribonucleoproteins can be delivered by polypeptide-based shuttle agents. In one example, ribonucleoproteins can be delivered using synthetic peptides, comprising an endosome leakage domain (ELD) operably connected to a cell penetrating domain (CPD), an ELD operably connected to a histidine-rich domain and a CPD, for example, as described in WO2016161516. RNPs can also be used to deliver compositions and systems to plant cells, for example, as described in Wu JW et al., Nat Biotechnol. 2015 Nov;33(11):1162-4.

[0338] Physical delivery

[0339] In one embodiment, the cargo can be introduced into the cell by a physical delivery method. Examples of physical methods include microinjection, electroporation, and hydrodynamic delivery. Both nucleic acids and proteins can be delivered using such methods. For example, one or more protein components can be prepared in vitro, separated (if necessary, refolded, purified), and introduced into the cell.

[0340] Microinjection

[0341] Microinjection of cargo directly into cells can achieve high efficiencies, e.g., greater than 90% or about 100%. In one embodiment, microinjection can be performed using a microscope and a needle (e.g., 0.5-5.0 μm in diameter) to pierce the cell membrane and deliver the cargo directly to the target site within the cell. Microinjection can be used for both in vitro and ex vivo delivery.

[0342] The plasmid containing the coding sequence of one or more protein components and / or crRNA, mRNA and / or guide RNA can be microinjected. In some cases, microinjection can be used for i) DNA is delivered directly to the nucleus, and / or ii) mRNA (e.g., in vitro transcribed) is delivered to the nucleus or cytoplasm. In some instances, microinjection can be used for crRNA is delivered directly to the nucleus and mRNA is delivered to the cytoplasm, thereby, for example, promoting the translation of one or more protein components and shuttling to the nucleus.

[0343] Microinjection can be used to generate genetically modified animals. For example, gene-editing cargo can be injected into fertilized eggs to allow for efficient germline modification. This method can produce normal embryos and full-term mouse pups with the desired modification. Microinjection can also be used, for example, to transiently upregulate or downregulate specific genes within the cellular genome using Cas12o proteins or CRISPR-associated Cas12o proteins.

[0344] electroporation

[0345] In one embodiment, cargo and / or delivery vehicles can be delivered by electroporation. Electroporation can use a pulsed high-voltage current to instantaneously open nanometer-sized pores in the cell membrane of cells suspended in a buffer, thereby allowing components with a hydrodynamic diameter of tens of nanometers to flow into the cell. In some cases, electroporation can be used for various cell types and efficiently transfer cargo into cells. Electroporation can be used for both in vitro and ex vivo delivery.

[0346] Electroporation can also be used to deliver cargo to the nucleus of mammalian cells by applying a specific voltage and reagent, for example, by nuclear transfection. Such methods include those described in Wu Y et al. (2015). Cell Res 25: 67-79; Ye L et al. (2014). Proc Natl Acad Sci USA 111: 9591-6; Choi PS, Meyerson M. (2014). Nat Commun 5: 3728; Wang J, Quake SR. (2014). Proc Natl Acad Sci 111: 13157-62. Electroporation can also be used to deliver cargo in vivo, for example, by using the method described in Zuckermann M et al. (2015). Nat Commun 6: 7391.

[0347] hydrodynamic delivery

[0348] Hydrodynamic delivery can also be used to deliver cargo, for example for in vivo delivery. In some instances, hydrodynamic delivery can be performed by rapidly pushing a large volume (8%-10% body weight) solution containing gene editing cargo into the bloodstream of a subject (e.g., an animal or human), for example, in mice, through the tail vein. Because blood is incompressible, large doses of liquid may cause an increase in hydrodynamic pressure, thereby temporarily enhancing the permeability to endothelial cells and parenchymal cells, thereby allowing cargo that normally cannot pass through the cell membrane to enter the cell. This method can be used to deliver naked DNA plasmids and proteins. The delivered cargo can be enriched in the liver, kidneys, lungs, muscles and / or heart.

[0349] Transfection

[0350] Cargo, such as nucleic acids, can be introduced into cells by transfection methods for introducing nucleic acids into cells. Examples of transfection methods include calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, impalefection, optical transfection, and proprietary agent-enhanced nucleic acid uptake.

[0351] Cell-penetrating peptides

[0352] In one embodiment, the delivery vehicle comprises a cell penetrating peptide (CPP).CPPs are short peptides that facilitate cellular uptake of a variety of molecular cargoes (e.g., from nanometer-sized particles to small chemical molecules and large DNA fragments).

[0353] CPPs can have different sizes, amino acid sequences, and charges. In some instances, CPPs can translocate across the plasma membrane and facilitate the delivery of various molecular cargoes to the cytoplasm or organelles. CPPs can be introduced into cells through different mechanisms, such as direct membrane penetration, endocytosis-mediated entry, and translocation through the formation of transient structures.

[0354] The amino acid composition of a CPP can contain a high relative abundance of positively charged amino acids (such as lysine or arginine), or have a sequence containing an alternating pattern of polar / charged amino acids and non-polar hydrophobic amino acids. These two types of structures are referred to as polycationic or amphiphilic structures, respectively. The third type of CPP is a hydrophobic peptide, which contains only non-polar residues, has a low net charge, or has hydrophobic amino acid groups that are critical for cellular uptake. Another type of CPP is the transactivating transcription activator (Tat) from human immunodeficiency virus 1 (HIV-1). Examples of CPPs include penetratin, Tat (48-60), transporter peptides (Transportan) and (R-AhX-R4) (Ahx refers to aminocaproyl), Kaposi's fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, and sweet arrow peptide. Examples of CPPs and related applications also include those described in US Patent No. 8,372,951.

[0355] CPPs can be easily used for in vitro and ex vivo effects and generally require extensive optimization for each cargo and cell type. In some instances, CPPs can be covalently attached directly to Cas12o proteins, which are then complexed with crRNA and delivered to cells. In some instances, CPP-Cas12o and CPP-crRNA can be delivered separately to multiple cells. CPPs can also be used to deliver RNPs.

[0356] CPPs can be used to deliver compositions and systems to plants. In some examples, CPPs can be used to deliver components to plant protoplasts, which are then regenerated into plant cells and further into plants.

[0357] Gold nanoparticles

[0358] In one embodiment, the delivery vehicle includes gold nanoparticles (also referred to as AuNPs or colloidal gold). The gold nanoparticles can form a complex with cargo, such as Cas12o protein:crRNA RNP. The gold nanoparticles can be coated, for example, in silicate and endosome-destroying polymer PAsp (DET). Examples of gold nanoparticles include Spherical Nucleic Acid (SNATM) constructs from AuraSense Therapeutics, and those described in Mout R, et al. (2017). ACS Nano 11: 2452-8; Lee K et al. (2017). Nat Biomed Eng 1: 889-901.

[0359] Genetically modified cells and organisms

[0360] The present disclosure also provides cells comprising one or more components of the compositions and systems herein, such as Cas12o proteins or CRISPR-related Cas12o proteins and / or crRNA. Additionally provided are cells modified by the systems and methods herein, and cell cultures, tissues, organs, and organisms comprising such cells or their progeny. In one embodiment, the present disclosure provides a method for modifying cells or organisms. The cells may be prokaryotic cells or eukaryotic cells. The cells may be mammalian cells. The mammalian cells may be non-human primate, cattle, pig, rodent, or mouse cells. The cells may be non-mammalian eukaryotic cells, such as poultry, fish, or shrimp. The cells may be therapeutic T cells or antibody-producing B cells. The cells may also be plant cells. Plant cells may be cells of crops such as cassava, corn, sorghum, wheat, or rice. Plant cells may also be cells of algae, trees, or vegetables. Modifications introduced into cells by the present disclosure may alter the production of cells and progeny of cells to increase bioproducts (such as antibodies, starch, alcohol, or other desired cell outputs). The modifications introduced into a cell through the present disclosure can be such that the cell and progeny of the cell include alterations that alter the biological product produced.

[0361] In one embodiment, one or more polynucleotide molecules, vectors, or vector systems that drive expression of one or more elements of a composition, system, or delivery system comprising one or more elements of a nucleic acid-targeting system are introduced into a host cell such that expression of these elements of the nucleic acid-targeting system directs the formation of a nucleic acid-targeting complex at one or more target sites. In one embodiment of the present disclosure, the host cell can be a eukaryotic cell, a prokaryotic cell, or a plant cell.

[0362] In one embodiment, the host cell is a cell of a cell line. Cell lines can be obtained from a variety of sources known to those skilled in the art (see, for example, American Type Culture Collection (ATCC) (Manassus, Va.)). In one embodiment, cells transfected with one or more vectors as described herein are used to establish new cell lines comprising sequences derived from one or more vectors. In one embodiment, cells transiently transfected with components of a system as described herein (such as by one or more vectors for transient transfection, or transfected with RNA) and modified by the activity of the complex are used to establish cell lines comprising cells containing modifications but lacking any other exogenous sequences. In one embodiment, cells transiently or non-transiently transfected with one or more vectors as described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0363] Also contemplated are human cells or tissues, plants, or non-human animals comprising one or more of the polynucleotide molecules, vectors, vector systems, or cells of any one of the embodiments herein. In one aspect, host cells and cell lines modified by or comprising a composition, system, or modified enzyme of the present disclosure are provided, including (isolated) stem cells and progeny thereof.

[0364] In one embodiment, a plant or non-human animal comprises at least one of the system components, polynucleotide molecules, vectors, vector systems or cells of any one of the embodiments herein at least one tissue type of the plant or non-human animal. In one embodiment, a non-human animal comprises at least one of the system components, polynucleotide molecules, vectors, vector systems or cells of any one of the embodiments herein in at least one tissue type. In one embodiment, the presence of system components is transient because they degrade over time. In one embodiment, the expression of the components of the system and composition of any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell is limited to certain tissue types or regions in a plant or non-human animal. In one embodiment, the expression of the components of the system and composition of any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell depends on physiological cues. In one embodiment, the expression of the components of the system and composition of any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell can be triggered by exogenous molecules. In one embodiment, the expression of the components of the system and composition of any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell depends on the expression of non-cas molecules in a plant or non-human animal.

[0365] General applications and uses

[0366] The systems, vector systems, vectors and compositions described herein can be used in various nucleic acid targeting applications to alter or modify the synthesis of gene products (such as proteins), nucleic acid cleavage, nucleic acid editing, nucleic acid splicing; transport of target nucleic acids, tracking of target nucleic acids, isolation of target nucleic acids, visualization of target nucleic acids, etc.

[0367] Thus, aspects of the present disclosure also encompass methods and uses of the compositions and systems described herein in genome engineering, e.g., for altering or manipulating expression of one or more genes or one or more gene products in prokaryotic or eukaryotic cells in vitro, in vivo, or ex vivo. In some instances, the target polynucleotide is a target sequence within genomic DNA (including nuclear genomic DNA, mitochondrial DNA, or chloroplast DNA).

[0368] Typically, in the context of a nucleic acid-targeting system, the formation of a nucleic acid-targeting complex (comprising a crRNA or guide RNA hybridized to a target sequence and complexed with one or more nucleic acid-targeting effector proteins) results in cleavage of one or both DNA or RNA strands in or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of) the target sequence. As used herein, the term "one or more sequences associated with a target locus of interest" refers to sequences that are near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs of) a target sequence, wherein the target sequence is contained in the target locus of interest.

[0369] In one embodiment, the present disclosure provides a method for targeting a polynucleotide, the method comprising contacting a sample (such as a cell, cell population, tissue, organ, or organism) comprising a target polynucleotide with a composition, system, polynucleotide, or vector. The contact can result in modification of a gene product or modification of the amount or expression of a gene product. In some instances, the target sequence of the polynucleotide is a disease-associated target sequence.

[0370] In one embodiment, the present disclosure provides a method of modifying a target polynucleotide, the method comprising delivering one or more polynucleotides in a composition, or one or more vectors, to a cell or cell population comprising the target polynucleotide, wherein the complex guides a reverse transcriptase to the target sequence and the reverse transcriptase promotes insertion of a donor sequence from the crRNA into the target polynucleotide.

[0371] Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. "Disease-associated" genes or polynucleotides refer to any genes or polynucleotides that produce transcription or translation products at abnormal levels or in abnormal forms in cells derived from disease-accumulated tissues compared to tissues or cells of non-disease controls. It may be a gene that becomes expressed at an abnormally high level; it may be a gene that becomes expressed at an abnormally low level, and this altered expression is associated with the occurrence and / or progression of the disease. Disease-associated genes also refer to genes with mutations or genetic variations that are directly responsible for or are in linkage disequilibrium with genes that cause the cause of the disease. The products of transcription or translation may be known or unknown and may be at normal or abnormal levels.

[0372] The target polynucleotide of complex can be any polynucleotide endogenous or exogenous to eukaryotic cells.For example, target polynucleotide can be a polynucleotide present in the nucleus of eukaryotic cells.Target polynucleotide can be a sequence encoding a gene product (for example, a protein) or a non-coding sequence (for example, regulatory polynucleotides or junk DNA). Without wishing to be bound by theory, it is believed that the target sequence should be associated with TAM (target adjacent motif) (i.e., a short sequence identified by the complex). The precise sequence and length requirements of TAM are different due to the Cas12o protein used or the related Cas12o protein of CRISPR, but TAM is typically a 2-5 base pair sequence adjacent to the prototype spacer (i.e., target sequence), and technicians will be able to identify other TAM sequences for use with given Cas12o protein or the related Cas12o protein of CRISPR. In addition, the engineering of TAM interaction domains allows TAM specificity to be programmed, improves target site recognition fidelity, and increases the versatility of Cas12o protein genome engineering platforms.

[0373] In some cases, Cas12o proteins or CRISPR-associated Cas12o proteins can be engineered to change their PAM specificity. In some embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM is 5'-TN, wherein N is A, T, G or C.

[0374] Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. "Disease-associated" genes or polynucleotides refer to any genes or polynucleotides that produce transcription or translation products at abnormal levels or in abnormal forms in cells derived from disease-accumulated tissues compared to tissues or cells of non-disease controls. It may be a gene that becomes expressed at an abnormally high level; it may be a gene that becomes expressed at an abnormally low level, and this altered expression is associated with the occurrence and / or progression of the disease. Disease-associated genes also refer to genes with mutations or genetic variations that are directly responsible for or are in linkage disequilibrium with genes that cause the cause of the disease. The products of transcription or translation may be known or unknown and may be at normal or abnormal levels.

[0375] Various aspects of the present disclosure relate to: a method for targeting a polynucleotide, the method comprising contacting a sample comprising the polynucleotide with a composition, system, or Cas12o protein or CRISPR-associated Cas12o protein as described in any embodiment herein; a delivery system comprising a composition, system, or Cas12o protein or CRISPR-associated Cas12o protein nuclease as described in any embodiment herein; a polynucleotide comprising a composition, system, or Cas12o protein or CRISPR-associated Cas12o protein as described in any embodiment herein; a vector comprising a composition, system, or Cas12o protein or CRISPR-associated Cas12o protein as described in any embodiment herein; or a vector system comprising a composition, system, or Cas12o protein or CRISPR-associated Cas12o protein as described in any embodiment herein. In one embodiment, the target polynucleotide is contacted with at least two different compositions, systems, or Cas12o proteins or CRISPR-associated Cas12o proteins. In another embodiment, the two different Cas12o proteins have different target polynucleotide specificities, or degrees of specificity. In one embodiment, the two different Cas12o proteins or CRISPR-associated Cas12o proteins have different TAM specificities.

[0376] Also contemplated are methods of targeting a polynucleotide comprising contacting a sample comprising the polynucleotide with the compositions and systems, vectors, or polynucleotides herein, wherein the contacting results in modification of a gene product or modification of the amount or expression of a gene product. In one embodiment, expression of the targeted gene product is increased by the method. In one embodiment, expression of the targeted gene product is increased by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%. In one embodiment, the expression of the targeted gene product is increased by at least 1.5 fold, at least 2 fold, at least 2.5 fold, at least 3 fold, at least 3.5 fold, at least 3.5 fold, at least 4 fold, at least 4.5 fold, at least 5 fold, at least 10 fold, at least 10 fold, at least 15 fold, at least 20 fold, at least 25 fold, at least 50 fold, at least 100 fold. In one embodiment, the expression of the targeted gene product is decreased by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%. In one embodiment, the expression of the targeted gene product is reduced by at least 1.5 times, at least 2 times, at least 2.5 times, at least 3 times, at least 3.5 times, at least 3.5 times, at least 4 times, at least 4.5 times, at least 5 times, at least 10 times, at least 10 times, at least 15 times, at least 20 times, at least 25 times, at least 50 times, at least 100 times. In an alternative embodiment, the expression of the targeted gene product is reduced by the method described. In other embodiments, the expression of the targeted gene can be completely eliminated, or can be considered eliminated, if the residual expression level of the targeted gene is reduced below the detection limit of the method known in the art for quantifying, detecting or monitoring expression levels.

[0377] In one embodiment, one or more polynucleotide molecules, vectors, or vector systems that drive expression of a nucleic acid-targeting system or one or more elements of a delivery system comprising one or more elements of a nucleic acid-targeting system are introduced into a host cell such that expression of these elements of the nucleic acid-targeting system directs the formation of a nucleic acid-targeting complex at one or more target sites. In one embodiment of the present disclosure, the host cell can be a eukaryotic cell, a prokaryotic cell, or a plant cell.

[0378] Also contemplated are human cells or tissues, plants, or non-human animals comprising one or more of the polynucleotide molecules, vectors, vector systems, or cells of any one of the embodiments herein. In one aspect, host cells and cell lines modified by or comprising a composition, system, or modified enzyme of the present disclosure are provided, including (isolated) stem cells and progeny thereof.

[0379] In one embodiment, a plant or non-human animal comprises at least one of the compositions, polynucleotide molecules, vectors, vector systems or cells described in any one of the embodiments herein at least one tissue type of the plant or non-human animal. In certain embodiments, a non-human animal comprises at least one of the compositions, polynucleotide molecules, vectors, vector systems or cells described in any one of the embodiments herein in at least one tissue type. In one embodiment, the presence of the composition is transient because they degrade over time. In one embodiment, the expression of the composition described in any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell is limited to certain tissue types or regions in a plant or non-human animal. In one embodiment, the expression of the composition described in any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell depends on physiological cues. In one embodiment, the expression of the composition described in any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell can be triggered by an exogenous molecule. In one embodiment, the expression of the composition described in any one of the embodiments contained in a polynucleotide molecule, vector, vector system or cell depends on the expression of non-Cas molecules in plants or non-human animals.

[0380] In one aspect, the present disclosure provides a method for using one or more elements of a nucleic acid targeting system. The target nucleic acid complex disclosed herein provides an efficient means for modifying target DNA or RNA (single-stranded or double-stranded, linear or supercoiled). The target nucleic acid complex disclosed herein has a variety of practicality, including modification (e.g., deletion, insertion, translocation, inactivation, activation) of target DNA or RNA in a variety of cell types. In this way, the target nucleic acid complex disclosed herein has a broad spectrum of applications in, for example, gene therapy, drug screening, disease diagnosis and prognosis. The target nucleic acid complex of the exemplary target nucleic acid includes a target DNA or RNA compounded with a crRNA or guide RNA hybridized to a target sequence in a target locus.

[0381] In some embodiments, the present disclosure provides a method for cutting a target polynucleotide. The method may include modifying the target polynucleotide using a complex of a targeting nucleic acid that is bound to the target polynucleotide and affects the cutting of the target polynucleotide. In one embodiment, the complex of the targeting nucleic acid of the present disclosure can produce a break (e.g., a single-strand or double-strand break) in the polynucleotide sequence when introduced into a cell. In one embodiment, the method may include allowing a composition to bind to a target DNA or RNA to achieve cutting of the target DNA or RNA so as to modify the target DNA or RNA, wherein the complex of the targeting nucleic acid comprises an effector protein of the targeting nucleic acid that is compounded with a guide RNA that hybridizes to a target sequence within the target DNA or RNA. In one aspect, the present disclosure provides a method for modifying the expression of DNA or RNA in a eukaryotic cell. In one embodiment, the method includes allowing a complex of the targeting nucleic acid to bind to DNA or RNA so that the binding causes the expression of the DNA or RNA to increase or decrease; wherein the complex of the targeting nucleic acid comprises an effector protein of the targeting nucleic acid that is compounded with crRNA or guide RNA. Similar considerations and conditions apply to the above-mentioned method for modifying the target DNA or RNA. In fact, these sampling, culturing and reintroduction options are applicable to all aspects of the present disclosure. In one aspect, the present disclosure provides a method for modifying a target DNA or RNA in a eukaryotic cell, which can be performed in vivo, ex vivo, or in vitro. In one embodiment, the method comprises sampling a cell or cell population from a human or non-human animal, and modifying the one or more cells. Cultivation can be performed ex vivo at any stage. One or more cells can even be reintroduced into a non-human animal or plant. For the reintroduced cells, it is particularly preferred that the cells are stem cells. The compositions described in any embodiment herein can be used to detect nucleic acid identifiers.

[0382] In one embodiment, the composition herein induces double-strand breaks in order to achieve the purpose of inducing HDR-mediated correction. In another embodiment, two or more guide RNAs compounded with Cas12o protein or its ortholog or homolog can be used to induce multiple breaks in order to achieve the purpose of inducing HDR-mediated correction. Recombinant template nucleic acid, as the term is used in this article, refers to a nucleic acid sequence that can be used in combination with the composition disclosed herein to change the structure of the target position. In one embodiment, the target nucleic acid is modified to have some or all of the sequences of the recombinant template nucleic acid, typically at or near the cleavage site. In one embodiment, the recombinant template nucleic acid is single-stranded. In an alternative embodiment, the recombinant template nucleic acid is double-stranded. In one embodiment, the recombinant template nucleic acid is DNA, such as double-stranded DNA. In an alternative embodiment, the recombinant template nucleic acid is single-stranded DNA. In one embodiment, a recombinant template is provided for use as a template in homologous recombination, for example, provided in or near a target sequence nicked or cut by an effector protein of a targeting nucleic acid as a part of a complex of targeting nucleic acid. In one embodiment, nuclease-induced non-homologous end joining (NHEJ) can be used for targeted gene-specific knockout.

[0383] Example Applications

[0384] The present disclosure provides a non-naturally occurring or engineered composition, or one or more polynucleotides encoding the components of the composition, or a vector or delivery system comprising one or more polynucleotides encoding the components of the composition, which is used to modify target cells in vivo, in vitro or in vitro, and the modification can be implemented in such a way that the cells are changed so that once modified, the progeny or cell line of the cell modified by the Cas12o protein or CRISPR-related Cas12o protein retains the changed phenotype. The modified cells and progeny can be part of a multicellular organism, such as a plant or animal in which the composition is applied in vitro or in vivo to a desired cell type. The methods herein include therapeutic methods of treatment. The therapeutic methods of treatment may include gene or genome editing, or gene therapy.

[0385] In one embodiment, one or more vectors described herein are used to produce non-human transgenic animals or transgenic plants. In one embodiment, the transgenic animal is a mammal, such as a mouse, rat or rabbit. Methods for producing transgenic animals and plants are known in the art and generally start from a cell transfection method, such as described herein. In one embodiment, the present disclosure provides a non-naturally occurring composition comprising an engineered non-catalytically inactive Cas12o protein or CRISPR-related Cas12o protein as described herein, and this system is used in detection methods such as fluorescence in situ hybridization (FISH).

[0386] Patient-specific screening methods

[0387] Nucleic acid-targeting systems that target DNA (e.g., trinucleotide repeats) can be used to screen patients or patient samples for the presence of such repetitive sequences. The repetitive sequence can be a target for the RNA of the nucleic acid-targeting system, and if the nucleic acid-targeting system binds to it, the binding can be detected, indicating the presence of such a repetitive sequence. Thus, the nucleic acid-targeting system can be used to screen patients or patient samples for the presence of such repetitive sequences. An appropriate compound can then be administered to the patient to address the condition; alternatively, the nucleic acid-targeting system can be administered to bind and cause insertions, deletions, or mutations that alleviate the condition.

[0388] Genome-wide gene knockout screening

[0389] The Cas12o protein and system described herein can be used to perform efficient and cost-effective functional genomic screening. Such screening can utilize a whole genome library based on the Cas12o protein nuclease. Such screening and libraries can be used to determine the function of a gene, the cellular pathway in which the gene is involved, and the way in which any changes in gene expression lead to specific biological processes. One advantage of the present disclosure is that the composition avoids off-target binding and the side effects thereof. This is achieved using a system arranged to have a high degree of sequence specificity for the target DNA. In a preferred embodiment of the present disclosure, the Cas12o protein or the CRISPR-associated Cas12o protein complex is a Cas12o protein complex.

[0390] Functional changes and screening

[0391] In one embodiment, the present disclosure provides a method for evaluating and screening gene function. Compositions are used to accurately deliver functional domains, activate or repress genes, or change epigenetic states by accurately changing methylation sites on specific target loci, which can be applied to single cells or cell groups together with one or more crRNAs or guide RNAs or applied to genomes in cell banks in vitro or in vivo together with libraries, including administering or expressing libraries comprising multiple crRNAs (comprising guide molecules), and wherein screening also includes using Cas12o proteins or CRISPR-related Cas12o proteins, wherein the complex comprising Cas12o proteins or CRISPR-related Cas12o proteins is modified to include heterologous functional domains.

[0392] Modification of cells or organisms

[0393] The present disclosure also provides cells comprising one or more components of the systems herein, such as Cas12o proteins or CRISPR-related Cas12o proteins and / or crRNA. Additionally provided are cells modified by the systems and methods herein, and cell cultures, tissues, organs, and organisms comprising such cells or their progeny. The present disclosure includes, in one embodiment, a method for modifying cells or organisms. The cells may be prokaryotic cells or eukaryotic cells. The cells may be mammalian cells. The mammalian cells may be non-human primate, cattle, pig, rodent, or mouse cells. The cells may be non-mammalian eukaryotic cells, such as poultry, fish, or shrimp. The cells may also be plant cells. The plant cells may be cells of crops such as cassava, corn, sorghum, wheat, or rice. The plant cells may also be cells of algae, trees, or vegetables. The modifications introduced into cells by the present disclosure may result in changes in the cells and the progeny of the cells to increase the production of biological products (such as antibodies, starch, alcohol, or other desired cell outputs). The modifications introduced into cells by the present disclosure may result in changes in the cells and the progeny of the cells including changes in the biological products produced.

[0394] Therapeutic uses and treatments

[0395] Also provided herein are methods for diagnosing, predicting, treating and / or preventing a disease, state or illness in a subject. In general, methods for diagnosing, predicting, treating and / or preventing a disease, state or illness in a subject may include using a composition as described herein, a system or its components to modify a polynucleotide in a subject or its cell, and / or include using a composition as described herein, a system or its components to detect a diseased or healthy polynucleotide in a subject or its cell. In one embodiment, a treatment or prevention method may include using a composition, a system or its components to modify a polynucleotide of an infectious organism (e.g., a bacterium or virus) in a subject or its cell. In one embodiment, a treatment or prevention method may include using a composition, a system or its components to modify a polynucleotide of an infectious organism or a symbiotic organism in a subject. The composition, system and its components may be used to develop a model for a disease, state or illness. The composition, system and its components may be used to detect a disease state or its correction, such as by a treatment or prevention method as described herein. The composition, system and its components may be used to screen and select cells that can be used as, for example, treatment or prevention as described herein. The composition, system and its components may be used to develop a bioactive agent that can be used to modify one or more biological functions or activities in a subject or its cell.

[0396] In one embodiment, the use comprises administering the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned CRISPR-Cas composition, complex, system, kit, delivery composition, enzyme preparation to the subject or the subject's ex vivo cells.

[0397] In some embodiments, the condition or disease includes metabolic disease, cancer, neurological disease, ophthalmic disease, and infectious disease.

[0398] In some embodiments, the condition or disease comprises a genetic disease.

[0399] In some embodiments, the condition or disease is caused by a pathogenic point mutation.

[0400] In one embodiment, the diseases include cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, alpha-1 antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, beta thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia (FH), atherosclerosis (ASCVD), transthyretin amyloidosis (ATTR), alpha-1 antitrypsin deficiency (AATD), retinal Membranous diseases, macular degeneration, Wilms' tumor, Ewing sarcoma, neuroendocrine tumors, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, non-Hodgkin's lymphoma and urinary bladder cancer, primary hyperoxaluria (PH1), hereditary angioedema (HAE) and hepatitis B (HEPATITIS B).

[0401] In some embodiments, the condition or disease comprises hypercholesterolemia (FH), atherosclerosis (ASCVD), transthyretin amyloidosis (ATTR), Alpha-1 antitrypsin deficiency (AATD), primary hyperoxaluria (PH1), hereditary angioedema (HAE), and hepatitis B.

[0402] In one embodiment, the disorder or disease comprises a disease caused by a single base mutation.

[0403] In one embodiment, the clinical variant database is obtained from the NCBI ClinVar database available on the NCBI ClinVar website. Pathogenic single nucleotide polymorphisms (SNPs) are identified from this list. Using genomic locus information, the CRISPR target in the region overlapping with and surrounding each SNP is identified. The selection of SNPs that can be corrected and targeted for causal mutations using base editing in combination with Cas proteins or their variants is listed in the table below, in which only one alias for each disease is listed. "RS#" corresponds to the RS accession number in the SNP database on the NCBI website. "AlleleID" corresponds to the causal allele accession number. The "Name" column contains the locus identifier of the gene, the gene name, the mutation position in the gene, and the changes caused by the mutation.

[0404] In one embodiment, the compositions, systems, and / or components thereof described herein can be used to treat and / or prevent circulatory system diseases. In one embodiment, the compositions, systems, and / or components thereof described herein can be used to treat neurological diseases. In one embodiment, the compositions, systems, and / or components thereof described herein can be used to treat hearing diseases, such as hearing diseases or hearing loss in one or both ears. Deafness is often caused by the loss or damage of hair cells, which prevents them from transmitting signals to auditory neurons. In such cases, cochlear implants can be used to respond to sound and transmit electrical signals to nerve cells. However, because damaged hair cells release fewer growth factors, these neurons often degenerate and retract from the cochlea. In one embodiment, the compositions, systems, and / or components thereof described herein can be used to treat diseases in non-dividing cells; in one embodiment, the gene or transcript to be corrected is located in a non-dividing cell. Exemplary non-dividing cells are muscle cells or neurons. Non-dividing (especially non-dividing, fully differentiated) cell types raise issues regarding gene targeting or genome engineering, for example because homologous recombination (HR) is generally inhibited during the G1 cell cycle phase. In one embodiment, the disease to be treated is a disease affecting the eye. Thus, in one embodiment, the compositions, systems, or components thereof described herein are delivered to one or both eyes. The compositions, systems, and methods described herein can be used to correct eye defects caused by several gene mutations, which are further described in Genetic Diseases of the Eye, 2nd edition, edited by Elias I. Traboulsi, Oxford University Press, 2012.

[0405] In one embodiment, the condition to be treated or targeted is an ocular condition. In one embodiment, the ocular condition may include glaucoma. In one embodiment, the ocular condition includes a retinal degenerative disease. In one embodiment, the retinal degenerative disease is selected from Stargardt disease, Bardet-Biedl syndrome, Best disease, blue cone achromatopsia, choroideremia, cone-rod dystrophy, congenital stationary night blindness, enhanced S cone syndrome, juvenile X-linked retinoschisis, Leber congenital amaurosis, Malattia Leventinesse, Norrie disease or X-linked familial exudative vitreoretinopathy, pattern dystrophy, Sorsby dystrophy, Usher syndrome, retinitis pigmentosa, color blindness or macular dystrophy or degeneration, retinitis pigmentosa, color blindness and age-related macular degeneration. In one embodiment, retinal degenerative disease is Leber congenital amaurosis (LCA) or retinitis pigmentosa. In one embodiment, a lentiviral vector is used for administration to eyes. In one embodiment, the lentiviral vector is an equine infectious anemia virus (EIAV) vector. Other viral vectors may also be used for delivery to eyes, such as AAV vectors, such as those described in: Campochiaro et al., Human Gene Therapy 17: 167-176 (February 2006); Millington-Ward et al. (Molecular Therapy, Vol. 19 No. 4, 642-649 April 2011; Dalkara et al. (Sci Transl Med 5, 189ra76 (2013)), which may be suitable for use with compositions as described herein, systems. In one embodiment, dosage may be in the range of about 106 to 109.5 particle units.

[0406] In one embodiment, compositions, systems can be used to treat and / or prevent muscle disease and related circulatory system or cardiovascular disease or illness. The disclosure also contemplates that compositions as described herein, systems, such as Cas12o protein or CRISPR-related Cas12o protein system, are delivered to the heart. For the heart, myocardial tropism adeno-associated virus (AAVM) is preferred, particularly AAVM41 (see, e.g., Lin-Yanga et al., PNAS, March 10, 2009, Vol. 106, No. 10) showing preferential gene transfer in the heart. Administration can be systemic or local. For systemic administration, the dosage of about 1-10x 1014 individual vector genomes is considered. See also, e.g., Eulalio et al. (2012) Nature 492: 376 and Somasuntharam et al. (2013) Biomaterials 34: 7790, which teach that content can be suitable for and / or applied to compositions as described herein, systems.

[0407] For example, U.S. Patent Publication No. 20110023139, its teaching content can be suitable for and / or applied to compositions as described herein, system, describes the use of zinc finger nucleases to genetically modify cells, animals and proteins associated with cardiovascular disease. Cardiovascular disease generally includes hypertension, heart attack, heart failure and stroke and TIA. Any chromosomal sequence related to cardiovascular disease or the protein encoded by any chromosomal sequence related to cardiovascular disease can be used for the method disclosed herein. Cardiovascular-related proteins are usually selected based on the experimental association of cardiovascular-related proteins with cardiovascular disease development. For example, relative to a colony lacking cardiovascular disease, in a colony suffering from cardiovascular disease, the production rate or circulating concentration of cardiovascular-related proteins can be increased or decreased. Proteomic techniques can be used to assess the difference in protein levels, and the techniques include but are not limited to western blotting, immunohistochemical staining, enzyme-linked immunosorbent assay (ELISA) and mass spectrometry. Alternatively, cardiovascular-related proteins can be identified by using genomic technology to obtain gene expression profiles of genes encoding proteins, and the techniques include but are not limited to DNA microarray analysis, gene expression serial analysis (SAGE) and quantitative real-time polymerase chain reaction (Q-PCR).

[0408] The compositions and systems herein can be used to treat diseases of the muscular system. The present disclosure also contemplates delivering the compositions, systems, and effector protein systems described herein to muscle. In one embodiment, the muscle disease to be treated is a muscular dystrophy, such as DMD. In one embodiment, the compositions and systems described herein (such as systems capable of RNA modification) can be used to achieve exon skipping to achieve correction of diseased genes. In one embodiment, the method includes treating sickle cell-related diseases, such as sickle cell trait, sickle cell disease such as sickle cell anemia, and beta thalassemia. For example, the methods and systems can be used to modify the genome of sickle cells, for example, by correcting one or more mutations in the beta globin gene. In the case of beta thalassemia, sickle cell anemia can be corrected by modifying HSCs with the system. The system allows for specific editing of the genome of a cell by cutting the cell's DNA and then allowing it to repair itself. The Cas12o protein is inserted and guided to the mutation point by an RNA guide, and then the DNA is cut at that point. At the same time, a healthy version of the sequence is inserted. This sequence is used by the cell's own repair system to repair the induced incision. In this way, the Cas12o protein or CRISPR-associated Cas12o protein allows correction of mutations in previously obtained stem cells. These methods and systems can be used to correct sickle cell anemia-related HSCs using a system that targets and corrects mutations (e.g., using a suitable HDR template that delivers a coding sequence for a beta globin, advantageously a non-sickled beta globin); specifically, the guide RNA can target mutations that cause sickle cell anemia, and HDR can provide coding for the correct expression of the beta globin. The ω RNA or guide RNA targeting particles containing mutations and Cas12o proteins is contacted with HSCs carrying mutations. The particles may also contain a suitable HDR template to correct the mutation so as to correctly express the beta globin; or the HSC may be contacted with a second particle or vector containing or delivering an HDR template. The cells so contacted may be administered; and optionally processed / amplified; see Cartier. The HDR template may enable the HSC to express an engineered beta globin gene (e.g., βA-T87Q) or beta globin.

[0409] In one embodiment, the compositions, systems, or components thereof described herein can be used to treat diseases of the kidney or liver. Thus, in one embodiment, the compositions, systems, or components thereof described herein are delivered to the liver or kidney. Delivery strategies that induce cellular uptake of therapeutic nucleic acids include physical forces or carrier systems, such as viral, lipid, or complex-based delivery, or nanocarriers. Based on initial applications with low potential clinical relevance, various gene therapy viral and non-viral vectors have been used to target post-transcriptional events in vivo in different animal kidney disease models when nucleic acids were delivered to kidney cells by systemic hydrodynamic high-pressure injection ((Csaba Révész and Péter Hamar (2011). Delivery Methods to Target RNAs in the Kidney, Gene Therapy Applications, Prof. Chunsheng Kang (ed.), ISBN: 978-953-307-541-9, InTech, available at: www.intechopen.com / books / gene-therapy-applications / delivery-methods-to-target-rnas-inthe-kidney). Delivery methods to the kidney may include those described in Yuan et al. (Am J Physiol Renal Physiol 295:F605-F617,2008). The method of Yuang et al. can be applied to the compositions of the present disclosure, which considers injecting 1-2g of Cas12o protein conjugated to cholesterol subcutaneously to people for delivery to the kidney. In one embodiment, the method of Molitoris et al. (J Am Soc Nephrol 20:1754-1764,2009) can be suitable for compositions, and a cumulative dose of 12-20mg / kg for people can be used for delivery to the proximal tubular cells of the kidney. In one embodiment, the method of Thompson et al. (Nucleic Acid Therapeutics, Vol. 22, No. 4, 2012) can be suitable for compositions, and a dose of up to 25mg / kg can be delivered by intravenous (iv) administration. In one embodiment, Shimizu et al. (J Am Soc Nephrol 21:622-633, 2010) can be adapted to the composition, and a dose of about 10-20 μmol of the composition complexed with the nanocarrier in about 1-2 liters of normal saline for intraperitoneal (ip) administration can be used.

[0410] In one embodiment, the disease treated or prevented by the compositions and systems described herein can be a pulmonary or epithelial disease. The compositions and systems described herein can be used to treat epithelial and / or pulmonary diseases. The present disclosure also contemplates delivering the compositions and systems described herein to one or both lungs.

[0411] In one embodiment, a viral vector can be used to deliver a composition, a system, or a component thereof to the lung. In one embodiment, AAV is AAV-1, AAV-2, AAV-5, AAV-6, and / or AAV-9 for delivery to the lung. (See, e.g., Li et al., Molecular Therapy, Vol. 17, No. 12, 2067-2077 December 2009). In one embodiment, the MOI can be 1 × 103 to 4 × 105 vector genomes / cell. In one embodiment, the delivery vector can be an RSV vector such as in Zamora et al. (Am J Respir Crit Care Med Vol. 183. Pages 531-538, 2011). The method of Zamora et al. can be applied to the nucleic acid targeting system of the present disclosure, and the composition of aerosolization, for example, with a dosage of 0.6 mg / kg, can be considered for use in the present disclosure.

[0412] The compositions and systems described herein can be used to treat skin disorders.The present disclosure also contemplates delivering the compositions and systems described herein to the skin.

[0413] In one embodiment, the composition, system, or components thereof can be delivered to the skin (intradermal delivery) by one or more microneedles or devices containing microneedles. For example, in one embodiment, the device and the method of Hickerson et al. (Molecular Therapy—Nucleic Acids (2013) 2, e129) can be used and / or suitable for, for example, delivering the composition, system as described herein to the skin at a dose of up to 300 μl of a 0.1 mg / ml composition. In one embodiment, the methods and techniques of Leachman et al. (Molecular Therapy, Vol. 18, No. 2, 442-446 February 2010) can be used and / or suitable for delivering the composition as described herein to the skin. In one embodiment, the methods and techniques of Zheng et al. (PNAS, July 24, 2012, Vol. 109, No. 30, 11975-11980) can be used and / or suitable for delivering nanoparticles of the composition as described herein to the skin. In one embodiment, a dose of about 25 nM applied in a single application can achieve gene knockdown in the skin.

[0414] The compositions and systems described herein can be used to treat cancer. The present disclosure also contemplates delivering the compositions and systems described herein to cancer cells. In addition, as described elsewhere herein, the compositions and systems can be used to modify immune cells, such as CAR or CART cells, which can then be used to treat and / or prevent cancer. This is also described in International Patent Publication No. WO 2015 / 161276, the disclosure of which is hereby incorporated by reference and described below.

[0415] The compositions, systems and components thereof described herein can be used to modify cells for adoptive cell therapy. In one aspect of the present disclosure, methods and compositions for editing target nucleic acid sequences or regulating the expression of target nucleic acid sequences and their application in combination with cancer immunotherapy are understood by adapting the compositions and systems disclosed herein. In some instances, the compositions, systems and methods can be used to modify stem cells (e.g., induced pluripotent cells) to derive modified natural killer cells, γδT cells and αβT cells, which can be used for adoptive cell therapy. In certain instances, the compositions, systems and methods can be used to modify modified natural killer cells, γδT cells and αβT cells.

[0416] As used herein, "ACT", "adoptive cell therapy" and "adoptive cell transfer" are used interchangeably. In one embodiment, adoptive cell therapy (ACT) can refer to transferring cells to a patient, with the aim of transferring functionality and characteristics to a new host by the implantation of cells (see, e.g., Mettananda et al., Editing an α-globin enhancerin primary human hematopoietic stem cells as a treatment for β-thalassemia, Nat Commun. September 4, 2017; 8(1): 424). As used herein, the term "engraftment" or "implantation" refers to the process of incorporating cells into a target tissue in vivo by contact with existing cells of the tissue. Adoptive cell therapy (ACT) can refer to transferring cells (most commonly immunogenic cells) back to the same patient or a new recipient host, with the aim of transferring immune functionality and characteristics to a new host. If possible, autologous cells are used to help the recipient by minimizing the problem of GVHD. Autologous tumor infiltrating lymphocytes (TILs) (Zacharakis et al., (2018) Nat Med. 2018 Jun;24(6):724-730; Besser et al., (2010) Clin. Cancer Res 16(9):2646-55; Dudley et al., (2002) Science 298(5594):850-4; and Dudley et al., (2005) Journal of Clinical Oncology 23(10):2346-57.) or genetically redirected peripheral blood mononuclear cells (Johnson et al., (2009) Blood 114(3):535-46; and Morgan et al., (2006) Science 314(5796)126-9) adoptive transfer has been used to successfully treat patients with advanced solid tumors (including melanoma, metastatic breast cancer and colorectal cancer) and patients with CD19-expressing hematological malignancies (Kalos et al., (2011) Science Translational Medicine 3(95):95ra73). In one embodiment, allogeneic immune cells are transferred (see, for example, Ren et al., (2017) Clin Cancer Res 23(9):2255-2266). As further described herein, allogeneic cells can be edited to reduce alloreactivity and prevent graft-versus-host disease. Thus, the use of allogeneic cells allows cells to be obtained from healthy donors and prepared for use in patients, rather than preparing autologous cells from patients after diagnosis.

[0417] In one embodiment, the antigen (such as a tumor antigen) targeted in adoptive cell therapy (such as in particular CAR or TCR T cell therapy) of a disease (such as in particular a tumor or cancer) can be selected from the group consisting of: MR1 (see, e.g., Crowther et al., 2020, Genome-wide CRISPR-Cas9 screening reveal subiquitous T cell cancer targeting via the monomorphic MHC class I-related protein MR1, Nature Immunology, Vol. 21, pp. 178-185); B cell maturation antigen (BCMA) (see, e.g., Friedman et al., Effective Targeting of Multiple BCMA-Expressing Hematological Malignancies by Anti-BCMACAR T Cells, Hum Gene Ther. Mar 8, 2018; Berdeja JG et al. Durable clinical responses in heavily pretreated patients with relapsed / refractory multiple myeloma: updated results from a multicenter study of bb2121 anti-Bcma CAR T cell therapy. Blood. 2017;130:740; and Mouhieddine and Ghobrial, Immunotherapy in Multiple Myeloma: The Era of CAR T Cell Therapy, Hematologist, May-June 2018, Vol. 15, No. 3); PSA (prostate-specific antigen); prostate-specific membrane antigen (PSMA); PSCA (prostate stem cell antigen); tyrosine-protein kinase transmembrane receptor ROR1; fibroblast activation protein (FAP); tumor-associated glycoprotein 72 (TAG72); carcinoembryonic antigen (CEA); epithelial cell adhesion molecule (EPCAM); mesothelin; human epidermal growth factor receptor 2 (ERBB2 (Her2 / neu)); prostatic enzyme; prostatic acid phosphatase (PAP); elongation factor 2 mutant (ELF2M); insulin-like growth factor 1 receptor (IGF-1R); gp100; BCR-ABL (breakpoint cluster region-Abelson); tyrosinase;New York Esophageal Squamous Cell Carcinoma 1 (NY-ESO-1); kappa-light chain, LAGE (L antigen); MAGE (melanoma antigen); melanoma-associated antigen 1 (MAGE-A1); MAGE A3; MAGE A6; legumin; human papillomavirus (HPV) E6; HPV E7; prostein; survivin; PCTA1 (galectin 8); Melan-A / MART-1; Ras mutants; TRP-1 (tyrosinase-related protein 1 or gp75); tyrosinase-related protein 2 (TRP2); TRP-2 / INT2 (TRP-2 / intron 2); RAGE (renal antigen); receptor for advanced glycation end products 1 (RAGE1); renal ubiquitin 1 and 2 (RU1 and RU2); intestinal carboxylesterase (iCE); heat shock protein 70-2 (HSP70-2) mutants; thyroid-stimulating hormone receptor (TSHR); CD123; CD1 71; CD19; CD20; CD22; CD26; CD30; CD33; CD44v7 / 8 (cluster of differentiation 44, intron 7 / 8); CD53; CD92; CD100; CD148; CD150; CD200; CD261; CD262; CD362; CS-1 (CD2 subset 1, CRACC, SLAMF7, CD319, and 19A24); C-type lectin-like molecule 1 (CLL-1); ganglioside GD3 (aNeu5Ac(2-8)aNeu5Ac(2-3)bDGalp(1-4)bDGlcp(1-1)Cer); Tn antigen (Tn Ag); Fms-like tyrosine kinase 3 (FLT3); CD38; CD138; CD44v6; B7H3 (CD276); KIT (CD117); interleukin-13 receptor subunit alpha-2 (IL-13Ra2); interleukin-11 receptor alpha (IL-11Ra); prostate stem cell antigen (PSCA); serine protease 21 (PRSS21); vascular endothelial growth factor receptor 2 (VEGFR2); Lewis (Y) antigen CD24; platelet-derived growth factor receptor β (PDGFR-β); stage-specific embryonic antigen 4 (SSEA-4); cell surface-associated mucin 1 (MUC1); mucin 16 (MUC16); epidermal growth factor receptor (EGFR); epidermal growth factor receptor variant III (EGFRvIII); neural cell adhesion molecule (NCAM); carbonic anhydrase IX (CAIX); proteasome (macropain) β subunit type 9 (LMP2); ephrin type A receptor 2 (EphA2); ephrin B2; fucosyl GM1; sialyl Lewis adhesion molecule (sLe);Ganglioside GM3 (aNeu5Ac(2-3)bDGalp(1-4)bDGlcp(1-1)Cer); TGS5; high molecular weight melanoma-associated antigen (HMWMAA); o-acetyl-GD2 ganglioside (OAcGD2); folate receptor alpha; folate receptor beta; tumor endothelial marker 1 (TEM1 / CD248); tumor endothelial marker 7-related (TEM7R); tight junction protein (claudin) 6 (CLDN6); G protein-coupled receptor class C group 5 member D (GPRC5D); chromosome X open reading frame 61 (CXORF61); CD97; CD179a; anaplastic lymphoma kinase (ALK); polysialic acid; placenta-specific 1 (PLAC1); globoH ceramide hexasaccharide moiety (GloboH); mammary gland differentiation antigen (NY-BR-1); uroplakin 2 (UPK2); hepatitis A virus cellular receptor 1 (HAVCR1); adrenergic receptor beta 3 (ADRB3); pan-nexin 3 (PANX3); G protein-coupled receptor 20 (GPR20); lymphocyte antigen 6 complex locus K9 (LY6K); olfactory receptor 51E2 (OR51E2); TCR gamma alternate reading frame protein (TARP); Wilms tumor protein (WT1); ETS translocation variant gene 6, located on chromosome 12p (ETV6-AML); sperm protein 17 (SPA17); X antigen family member 1A (XAGE1); angiopoietin-binding cell surface receptor 2 (Tie 2); CT (cancer / testis (antigen)); melanoma cancer testis antigen 1 (MAD-CT-1); melanoma cancer testis antigen 2 (MAD-CT-2); Fos-related antigen 1; p53; p53 mutant; human telomerase reverse transcriptase (hTERT); sarcoma translocation breakpoint; melanoma inhibitor of apoptosis (ML-IAP); ERG (transmembrane protease serine 2 (TMPRSS2) ETS fusion gene); N-acetylglucosaminyltransferase V (NA17); paired box protein Pax-3 (PAX3); androgen receptor; cyclin B1; cyclin D1; v-myc avian myelocytic neoplasia viral oncogene neuroblastoma-derived homolog (MYCN); Ras homolog family member C (RhoC); cytochrome P450 1B1 (CYP1B1); CCCTC-binding factor (zinc finger protein)-like (BORIS); squamous cell carcinoma antigen recognized by T cells 1 or 3 (SART1, SART3); paired box protein Pax-5 (PAX5); pre-acrosin binding protein sp32 (OY-TES1); lymphocyte-specific protein tyrosine kinase (LCK); A kinase anchoring protein 4 (AKAP-4); synovial sarcoma X breakpoint 1, 2, 3, or 4 (SSX1, SSX2, SSX3, SSX4); CD79a;CD79b; CD72; leukocyte-associated immunoglobulin-like receptor 1 (LAIR1); Fc fragment of the IgA receptor (FCAR); leukocyte immunoglobulin-like receptor subfamily A, member 2 (LILRA2); CD300 molecule-like family member f (CD300LF); C-type lectin domain family 12, member A (CLEC12A); bone marrow stromal cell antigen 2 (BST2); mucin-like hormone receptor-like 2 with an EGF-like module (EMR2); lymphocyte antigen 75 (LY75); glypican 3 (GPC3); Fc receptor-like 5 (FCRL5); mouse double minute 2 homolog (MDM2); livin; alpha-fetoprotein (AFP); transmembrane activator and CAML interactor (TACI); B-cell activating factor receptor (BAFF-R); V-Ki-ras2 Kirsten rat sarcoma viral oncogene homolog (KRAS); immunoglobulin lambda-like polypeptide 1 (IGLL1); 707-AP (707 alanine proline); ART-4 (adenocarcinoma antigen recognized by T4 cells); BAGE (B antigen; b-catenin / m, b-catenin / mutant); CAMEL (melanoma antigen recognized by CTL); CAP1 (carcinoembryonic antigen peptide 1); CASP-8 (caspase 8); CDC27m (mutant cell division cycle 27); CDK4 / m (mutant cyclin-dependent kinase 4); Cyp-B (cyclophilin B); DAM (differentiation antigen). melanoma); EGP-2 (epithelial glycoprotein 2); EGP-40 (epithelial glycoprotein 40); Erbb2, 3, 4 (erythroblastic leukemia viral oncogene homolog 2, 3, 4); FBP (folate binding protein); fAchR (fetal acetylcholine receptor); G250 (glycoprotein 250); GAGE ​​(G antigen); GnT-V (N-acetylglucosaminyltransferase V); HAGE (helicase antigen); ULA-A (human leukocyte antigen A); HST2 (human signet ring tumor 2); KIAA0205; KDR (kinase insert domain receptor); LDLR / FUT (low-density lipid receptor / GDP receptor) L-fucose: bD-galactosidase 2-aL fucosyltransferase); L1CAM (L1 cell adhesion molecule); MC1R (melanocortin 1 receptor); Myosin / m (mutant myosin); MUM-1, 2, 3 (melanoma ubiquitously mutated proteins 1, 2, 3); NA88-A (NA cDNA clone from patient M88); KG2D (natural killer group 2 member D) ligand; carcinoembryonic antigen (h5T4); p190 small bcr-abl (190KD bcr-abl protein); Pml / RARa (promyelocytic leukemia / retinoic acid receptor alpha); PRAME (melanoma preferentially expressed antigen); SAGE (sarcoma antigen);TEL / AML1 (translocation Ets family leukemia / acute myeloid leukemia 1); TPI / m (mutated triosephosphate isomerase); CD70; and any combination thereof.

[0418] In some embodiments, the compositions, systems, or components of the present disclosure can be used to treat and / or prevent genetic diseases or diseases with genetic and / or epigenetic aspects. The genes and diseases exemplified herein are not exhaustive. In one embodiment, a method for treating and / or preventing a genetic disease may include administering a composition, system, and / or one or more components thereof to a subject, wherein the composition, system, and / or one or more components thereof are capable of modifying one or more copies of one or more genes associated with a genetic disease or a disease with genetic and / or epigenetic aspects in one or more cells of the subject. In one embodiment, modifying one or more copies of one or more genes associated with a genetic disease or a disease with genetic and / or epigenetic aspects in a subject can eliminate the genetic disease or its symptoms in the subject. In one embodiment, modifying one or more copies of one or more genes associated with a genetic disease or a disease with genetic and / or epigenetic aspects in a subject can reduce the severity of the genetic disease or its symptoms in the subject. In one embodiment, a composition, system, or component thereof can modify one or more genes or polynucleotides associated with one or more diseases, including genetic diseases and / or diseases with genetic and / or epigenetic aspects.

[0419] In one embodiment, the composition, system, or components thereof can be used to diagnose, predict, treat, and / or prevent infectious diseases caused by microorganisms such as bacteria, viruses, fungi, parasites, or combinations thereof. In one embodiment, the system or components thereof can target specific microorganisms in a mixed population. Exemplary methods of such technologies are described, for example, in Gomaa AA, Klumpe HE, Luo ML, Selle K, Barrangou R, Beisel CL. 2014. Programmable removal of bacterial strains by use of genome-targeting composition, systems, mBio5: e00928-13; Citorik RJ, Mimee M, LuTK. 2014. Sequence-specific antimicrobials using efficiently delivered RNA-guided nucleases. Nat Biotechnol 32: 1141-1145, the teachings of which can be adapted for use with the compositions, systems, and components thereof described herein. In one embodiment, the composition, system, and / or components thereof can target pathogenic and / or drug-resistant microorganisms, such as bacteria, viruses, parasites, and fungi. In one embodiment, the compositions, systems and / or components thereof are capable of targeting and modifying one or more polynucleotides in a pathogenic microorganism such that the microorganism is rendered less toxic, killed, inhibited, or otherwise rendered incapable of causing disease and / or infection and / or replication in a host cell.

[0420] In one aspect, some of the most challenging mitochondrial disorders are caused by mutations in mitochondrial DNA (mtDNA), a high copy number genome inherited maternally. In one embodiment, the compositions and systems described herein can be used to modify mtDNA mutations. In one embodiment, the mitochondrial disease that can be diagnosed, predicted, treated and / or prevented can be MELAS (mitochondrial myopathy encephalopathy and lactic acidosis and stroke-like episodes), CPEO / PEO (chronic progressive external ophthalmoplegia syndrome / progressive external ophthalmoplegia), KSS (Kearns-Sayre syndrome), MIDD (maternally inherited diabetes mellitus and deafness), MERRF (myoclonic epilepsy with ragged red fibers), NIDDM (non-insulin-dependent diabetes mellitus), LHON (Leber hereditary optic neuropathy), LS (Leigh syndrome), aminoglycoside-induced hearing impairment, NARP (neuropathy, ataxia and retinopathy pigmentosa), extrapyramidal disorders with akinesia-rigidity, psychosis and SNHL, non-syndromic hearing loss, cardiomyopathy, encephalomyopathy, Pearson's syndrome, or a combination thereof.

[0421] In one embodiment, the subject's mtDNA can be modified in vivo or ex vivo. In one embodiment, where mtDNA is modified ex vivo, cells containing the modified mitochondria can be administered back to the subject after modification. In one embodiment, the composition, system, or component thereof is capable of correcting mtDNA mutations or a combination thereof.

[0422] In some embodiments, the compositions, systems, or components thereof disclosed herein can be used for microbiome modification. The microbiome plays an important role in health and disease. For example, the intestinal microbiome can play a role in health by controlling digestion and preventing the growth of pathogenic microorganisms, and is believed to affect mood and emotions. An unbalanced microbiome can promote disease and is believed to lead to weight gain, uncontrolled blood sugar, high cholesterol, cancer, and other conditions. A healthy microbiome has a series of joint features that can be distinguished from an unhealthy individual, so the detection and identification of disease-associated microbiomes can be used to diagnose and detect individual diseases. The compositions, systems, and components thereof can be used to screen microbiome cell populations and to identify disease-associated microbiomes. Cell screening methods using compositions, systems, and components thereof are described elsewhere herein and can be applied to screen a subject's microbiome, such as the intestinal, skin, vaginal, and / or oral microbiome.

[0423] In one embodiment, the compositions, systems, and / or components thereof described herein can be used to modify the microbial population of a subject's microbiome. In one embodiment, the compositions, systems, and / or components thereof can be used to identify and select one or more cell types in a microbiome and remove them from the microbiome population. Exemplary methods for selecting cells using compositions, systems, and / or components thereof are described elsewhere herein. In this way, the composition or microbial characteristics of a microbiome can be changed. In one embodiment, the change causes a change from a diseased microbiome composition to a healthy microbiome composition. In this way, the ratio of one microbial type or species to another can be modified, such as from a diseased ratio to a healthy ratio. In one embodiment, the selected cells are pathogenic microorganisms.

[0424] In one embodiment, the compositions and systems described herein can be used to modify polynucleotides in a microorganism of a subject's microbiome. In one embodiment, the microorganism is a pathogenic microorganism. In one embodiment, the microorganism is a commensal and non-pathogenic microorganism. Methods for modifying polynucleotides in a subject's cells are described elsewhere herein and can be applied to these embodiments.

[0425] In one embodiment, the present disclosure provides a method of modeling a disease associated with a genomic locus in a eukaryotic or non-human organism, the method comprising manipulating a target sequence within a coding, non-coding or regulatory element of the genomic locus, comprising delivering a non-naturally occurring or engineered composition comprising a viral vector system comprising one or more viral vectors operably encoding a composition for expression thereof, wherein the composition comprises a particle delivery system or a delivery system or viral particle as described in any of the above embodiments or a cell as described in any of the above embodiments.

[0426] The present disclosure also contemplates the use of compositions described herein (e.g., Cas12o systems) to provide RNA-guided DNA nucleases suitable for use to provide modified tissues for transplantation. For example, RNA-guided DNA nucleases can be used to knock out, knock down, or destroy selected genes in animals such as transgenic pigs (such as human heme oxygenase 1 transgenic pig lines), for example, by destroying the expression of genes encoding epitopes recognized by the human immune system, i.e., the expression of xenoantigen genes. Candidate pig genes for destruction may, for example, include α(l,3)-galactosyltransferase and cytidine monophosphate-N-acetylneuraminic acid hydroxylase genes (see PCT Patent Publication WO 2014 / 066505). In addition, genes encoding endogenous retroviruses, such as genes encoding all porcine endogenous retroviruses, may be destroyed (see Yang et al., 2015, Genome-wide inactivation of porcine endogenous retroviruses (PERVs), Science 2015 November 27: Vol. 350, No. 6264, pp. 1101-1104). In addition, RNA-guided DNA nucleases can be used to target the integration sites of other genes in xenotransplant donor animals, such as the human CD55 gene, to improve protection against hyperacute rejection.

[0427] Applications in plants and fungi

[0428] The compositions, systems and methods described herein can be used for gene or genome interrogation or editing or manipulation in plants and fungi. For example, applications include investigation and / or selection and / or interrogation and / or comparison and / or manipulation and / or transformation of plant genes or genomes; for example, to create, identify, develop, optimize or confer plant traits or characteristics, or transform plant or fungal genomes. Thus, the yield of plants, new plants with a combination of new traits or characteristics, or new plants with enhanced traits can be increased. The compositions, systems and methods can be used for plants in site-directed integration (SDI) or gene editing (GE) or any near-reverse breeding (NRB) or reverse breeding (RB) techniques.

[0429] The compositions, systems and methods herein can be used to confer desirable traits (e.g., enhanced nutritional quality, enhanced disease resistance and resistance to biotic and abiotic stresses, and increased production of commercially valuable plant products or heterologous compounds) to essentially any plant and fungus, as well as their cells and tissues. The compositions, systems and methods can be used to modify endogenous genes or modify their expression without permanently introducing any foreign genes into the genome.

[0430] Examples of plants

[0431] The compositions, systems, and methods herein can be used to confer desired traits on essentially any plant. A wide variety of plants and plant cell systems can be engineered to obtain desired physiological and agronomic characteristics. Generally, the term "plant" refers to any of various photosynthetic, eukaryotic, unicellular or multicellular organisms in the plant kingdom that are characterized by growth by cell division, contain chloroplasts, and have cell walls composed of cellulose. The term plant encompasses monocots and dicots. In one embodiment, target plants and plant cells for engineering include those monocots and dicots, such as crops including cereal crops (e.g., wheat, corn, rice, millet, barley), fruit crops (e.g., tomato, apple, pear, strawberry, orange), forage crops (e.g., alfalfa), root vegetable crops (e.g., carrot, potato, beet, yam), leafy vegetable crops (e.g., lettuce, spinach); flowering plants (e.g., petunia, rose, chrysanthemum), conifers and pines (e.g., pine fir, spruce); plants used for phytoremediation (e.g., heavy metal accumulating plants); oil crops (e.g., sunflower, rapeseed) and plants used for experimental purposes (e.g., Arabidopsis thaliana). Specifically, plants are intended to include, but are not limited to, angiosperms and gymnosperms such as acacia, alfalfa, amaranth, apple, apricot, artichoke, ash, asparagus, avocado, banana, barley, beans, beets, birch, beech, blackberry, blueberry, broccoli, Brussels sprouts, cabbage, rapeseed, cantaloupe, carrot, cassava, cauliflower, cedar, cereals, celery, chestnut, cherry, Chinese cabbage, citrus, clementine, clover, coffee, corn, cotton, cowpea, cucumber, cypress, eggplant, elm, chicory, eucalyptus, fennel, fig, fir, geranium, grape, grapefruit, peanut, ground cherry, gum hemlock, pecan, kale, kiwi, kohlrabi, larch, lettuce, leek, lemon, Lime, acacia, pine, maidenhair fern, corn, mango, maple, melon, millet, mushroom, mustard, nuts, oak, oats, oil palm, okra, onion, orange, ornamental plants or flowers or trees, papaya, palm, parsley, parsnip, peas, peach, peanut, pear, peat, pepper, persimmon, pigeon pea, pine, pineapple, plantain, plum, pomegranate, potato, pumpkin, endive, radish, rapeseed, raspberry, rice, rye, sorghum, safflower, yellow willow, soybean, spinach, spruce, winter squash, strawberry, beet, sugarcane, sunflower, sweet potato, sweet corn, tangerine, tea, tobacco, tomato, trees, triticale, lawn grass, turnip, vine, walnut, watercress, watermelon, wheat, yam, yew, and zucchini.

[0432] The term plant also encompasses algae, which are primarily photoautotrophs formed primarily due to a lack of roots, leaves, and other organs characteristic of higher plants. The compositions, systems, and methods can be used for a wide range of "algae" or "algae cells." Examples of algae include eukaryotic phyla, including Rhodophyta (red algae), Chlorophyta (green algae), Phaeophyta (brown algae), Bacillariophyta (diatoms), Eustigmatophyta, and dinoflagellates, as well as prokaryotic Cyanobacteria (blue-green algae).Examples of algal species include those of the genera Amphora, Anabaena, Anikstrodesmis, Botryococcus, Chaetoceros, Chlamydomonas, Chlorella, Chlorococcum, Cyclotella, Cylindrotheca, Dunaliella, Emiliana, Euglena, Hematococcus, Isochrysis, Monochrysis, Monoraphidium, Nannochloris, Nannochloropsis, Navicula , Nephrochloris, Nephroselmis, Nitzschia, Nodularia, Nostoc, Oochromonas, Oocystis, Oscillartoria, Pavlova, Phaeodactylum, Playtmonas, Pleurochrysis, Porhyra, Pseudoanabaena, Pyramimonas, Stichococcus, Synechococcus, Synechocystis, Tetraselmis, Thalassiosira, and Trichodesmium.

[0433] In one embodiment, polynucleotides encoding components of the composition and system can be introduced to stably integrate into the genome of the plant cell. In some cases, vectors or expression systems can be used for such integration. The design of the vector or expression system can be adjusted according to the time, place and conditions of expression of the guide RNA and / or Cas12o protein or CRISPR-related Cas12o protein gene. In some cases, the polynucleotides can be integrated into the organelles of the plant, such as plastids, mitochondria or chloroplasts. The elements of the expression system can be located on one or more expression constructs, which are circular, such as plasmids or transformation vectors, or non-circular, such as linear double-stranded DNA.

[0434] In one embodiment, the integration method generally comprises the following steps: selecting a suitable host cell or host tissue, introducing the construct into the host cell or host tissue, and regenerating a plant cell or plant therefrom. In some examples, the expression system for stable integration into the plant cell genome may contain one or more of the following elements: a promoter element, which can be used to express RNA and / or Cas12o protein in plant cells; a 5' untranslated region for enhancing expression; an intron element for further enhancing expression in certain cells (such as monocotyledonous cells); a multiple cloning site for providing convenient restriction sites for inserting guide RNA and / or Cas12o protein gene sequences and other required elements; and a 3' untranslated region for providing efficient termination of expressed transcripts.

[0435] Transient expression in plants

[0436] In one embodiment, the components of the composition and system can be transiently expressed in plant cells. In some examples, the composition and system can modify the target nucleic acid only when the guide RNA and the Cas12o protein or the CRISPR-associated Cas12o protein are present in the cell, so that the genome modification can be further controlled. Since the expression of the Cas12o protein or the CRISPR-associated Cas12o protein is transient, the plants regenerated from such plant cells are generally free of foreign DNA. In certain examples, the Cas12o protein or the CRISPR-associated Cas12o protein is stably expressed and the guide sequence is transiently expressed.

[0437] DNA and / or RNA (e.g., mRNA) can be introduced into plant cells for transient expression. In such cases, sufficient amounts of the introduced nucleic acid can be provided to modify the cell, but the introduced nucleic acid does not persist after a desired period of time or after one or more cell divisions.

[0438] Transient expression can be achieved using a suitable vector. Exemplary vectors that can be used for transient expression include pEAQ vectors (which can be customized for Agrobacterium-mediated transient expression) and cabbage leaf curl virus (CaLCuV), as well as vectors described in Sainsbury F. et al., Plant Biotechnol J. 2009 September; 7(7): 682-93; and Yin K et al., Scientific Reports Vol. 5, Article No. 14926 (2015).

[0439] Exemplary applications in plants

[0440] The composition, system and method can be used to produce genetic variation in target plants (e.g., crops). One or more ω RNAs targeting one or more positions in the genome can be provided, such as a library of ω RNAs, and introduced into plant cells together with the Cas12o protein nuclease. For example, a group of genome-scale point mutations and gene knockouts can be produced. In some instances, the composition, system and method can be used to produce plant parts or plants from the cells so obtained, and screen cells for target traits. The target gene can include coding regions and non-coding regions simultaneously. In some cases, the trait is stress tolerance, and the method is a method for producing stress-tolerant crop varieties.

[0441] In one embodiment, the compositions, systems and methods are used to modify endogenous genes or modify their expression. The expression of the components can be induced by the direct activity of Cas12o protein or CRISPR-associated Cas12o protein and optionally by the introduction of recombinant template DNA, or by modifying the targeted gene to induce targeted modification of the genome. The different strategies described above allow targeted genome editing mediated by Cas12o protein or CRISPR-associated Cas12o protein without requiring the components to be introduced into the plant genome.

[0442] In some cases, modifications can be made without permanently introducing any foreign genes (including those encoding components of the compositions herein) into the plant genome to avoid the presence of foreign DNA in the plant genome. This may be of interest because regulatory requirements for non-transgenic plants are less stringent. Components that are transiently introduced into plant cells are typically removed upon hybridization.

[0443] For example, modification can be performed by transient expression of the components of the compositions and systems. Transient expression can be performed by delivering the components of the compositions and systems using viral vectors, by delivery into protoplasts via particulate molecules such as nanoparticles or CPPs.

[0444] medicine box

[0445] In one aspect, the present disclosure provides a medicine box containing any one or more elements disclosed in the above-mentioned methods and compositions. In one aspect, the present disclosure provides a medicine box including one or more components as described herein. In one embodiment, the medicine box includes the instructions for use of the composition herein and the medicine box. In one embodiment, the medicine box includes instructions for use of a carrier system and the medicine box. In one embodiment, the medicine box includes instructions for use of a delivery system and the medicine box. In one embodiment, the medicine box includes instructions for use of a carrier system and the medicine box. Each element can be provided individually or in combination, and each element can be provided in any suitable container such as a vial, a bottle or a tube. The medicine box may include crRNA as described herein and an optional unbound protection chain. The medicine box may include crRNA, wherein the protection chain is at least partially bound to the reprogrammable spacer portion (i.e., spacer sequence) of the crRNA sequence. In one embodiment, the medicine box includes instructions in one or more languages, such as instructions in more than one language. These instructions may be specific to the applications and methods described herein.

[0446] In one embodiment, the medicine box includes one or more reagents for use in the process utilizing one or more elements as described herein. Reagent can be provided in any suitable container. For example, the medicine box can provide one or more reaction or storage buffers. Reagent can be provided in a form useful in a specific assay, or provided in the form of needing to add one or more other components before use (for example, provided in a concentrate or lyophilized form). Buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and their combination. In one embodiment, buffer is alkaline. In one embodiment, buffer has a pH of about 7 to about 10. In one embodiment, the medicine box includes homologous recombination template polynucleotide. In one embodiment, the medicine box includes one or more vectors and / or one or more polynucleotides as described herein. The medicine box can advantageously allow all elements of the disclosed system to be provided.

[0447] The main advantages of the present disclosure include:

[0448] The present disclosure discloses a novel Cas protein, which comprises an OBD domain, a REC domain, a RuvC domain, a Helical domain, and a Nuc domain, and has a structure represented by Formula I or Formula II. The Cas protein disclosed herein has excellent gene editing activity, can effectively edit or cut target genes, and can be used to treat conditions or diseases in subjects in need.

[0449] Example

[0450] The present disclosure will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present disclosure only and are not intended to limit the scope of the present disclosure. The experimental methods in the following examples, for which specific conditions are not specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight.

[0451] Unless otherwise specified, the reagents and materials in the examples of this disclosure are all commercially available products.

[0452] Example 1 Identification of Cas12o protein

[0453] By analyzing the collected uncultured metagenome sequences, six Cas12o proteins or fragments associated with the discovered CRISPR system were found in the samples and named Cas12o1-Cas12o6, respectively. Their amino acid sequences are shown in SEQ ID NO.1, 3, 5, 7-9, and their nucleotide coding sequences are shown in SEQ ID NO.16, 40, and 46, respectively. The CRISPR loci of samples containing Cas12o1, Cas12o2, and Cas12o3 were annotated by PILERCR, and the corresponding direct repeat (DR) sequences were obtained, as shown in SEQ ID NO.2, 4, and 6, respectively, as shown in the following table:

[0454] RNAfold was used to further analyze the RNA secondary structure of the three DR sequences mentioned above. The results are shown in Figure 6. All DR sequences clearly have a very conserved secondary structure:

[0455] The above three DR sequences contain a 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3' structure, wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1), which has 3 or 5 nucleotide pairs in Cas12o; segments Ba and Bb do not exist at the same time, and a bulge (B) formed by the existing segment Ba or segment Bb is formed by 2 or 3 nucleotides; segments R2a and R2b are reverse complementary sequences and form a second stem (R2), which has 6 or 7 base pairs; and L is a loop formed at the second stem portion, formed with 5 or 7 nucleotides.

[0456] Specifically, the DR sequence of Cas12o1 is shown in Figure 6A, which contains 5'-R1a(ACA)-Ba(absent)-R2a(GGUAUCC)-L(UAAAC)-R2b(GGAUGCU)-Bb(GA)-R1b(UGU)-3'.

[0457] The DR sequence of Cas12o2 is shown in Figure 6B, which contains 5'-R1a(UUACA)-Ba(absent)-R2a(ACUAUUC)-L(UUGAAAC)-R2b(GAAUGGU)-Bb(GAU)-R1b(UGUAA)-3'.

[0458] The DR sequence of Cas12o3 is shown in Figure 6C, which contains 5'-R1a(UCAGU)-Ba(GUG)-R2a(GGUCUG)-L(AAACA)-R2b(CAGACC)-Bb(not present)-R1b(AUUGA)-3'.

[0459] By analyzing the phylogenetic relationship of the RuvC domain, it was found that Cas12o is a new Cas subtype belonging to class2, type V, which is close to Cas12h, Cas12i, and Cas12b subtypes (Figure 1) on the RuvC domain phylogenetic tree. The protein folding software AlphaFold was used to predict the Cas12o three-dimensional conformation (Figure 2B), and it was found that it has conserved domains REC-I, REC-II, OBD, RuvC-I, Helical, RuvC-II, Nuc-I, RuvC-III, and Nuc-II, and its protein conformation is different from other nucleases of the Cas12 family. By annotating Cas12o1 and Cas12o2 proteins, it was found that their representative conserved motifs are distributed as shown in Figure 2A, and by annotating Cas12o3 proteins, it was found that their conserved motifs are distributed as shown in Figure 2C, and the OBD domain is a binary fission domain, which is divided into discontinuous OBD-I and OBD-II domains.

[0460] Through experimental testing, it was found that Cas12o1 showed a strong preference for sequences containing the target 5' end PAM 5'-TN (where N is A, T, G or C)-3' flanking sequence (Figure 7).

[0461] Example 2 Verification of Cas12o protein cleavage activity

[0462] In order to test the cleavage activity of Cas12o1, vectors expressing Cas12o1 and crRNA were constructed as follows:

[0463] The human TTR gene was selected as the cleavage target gene, and Cas12o1-hTTR1-crRNA (SEQ ID NO.14) with a PAM sequence of TN corresponding to the target sequence was designed based on the hTTR1 target sequence (SEQ ID NO.12).

[0464] In order to compare the cleavage activity of Cas12o1 and the control LbCpf1 (amino acid sequence is SEQ ID NO.18, nucleotide coding sequence is SEQ ID NO.19) at the same target sequence, LbCpf1-hTTR1-crRNA (SEQ ID NO.21) with PAM as TTN was designed.

[0465] The T7 promoter was added to the 5' end of the Cas12o1-hTTR1-crRNA sequence and the LbCpf1-hTTR1-crRNA sequence, and the rrnB T2 terminator was added to the 3' end of the two, respectively, to obtain the Cas12o1-hTTR1-crRNA expression sequence and the LbCpf1-hTTR1-crRNA expression sequence, respectively. Among them, the single underline sequence part is the Cas12o1 / LbCpf1 DR sequence, the double underline sequence part is the spacer sequence, the italic sequence part is the T7 promoter, the wavy underline sequence part is the rrnB T2 terminator sequence, the linker sequence is between the spacer sequence and the rrnB T2 terminator sequence, the dotted sequence part is the MfeI restriction site, the bold sequence part is the MluI restriction site, and CACCG is the linker. To protect the integrity of the sequence fragment, the protective base AGC was introduced at the 5' end of the expression sequence, and the protective base ATA was introduced at the 3' end of the expression sequence.

[0466] The nucleotide coding sequences of Cas12o1 and LbCpf1 were synthesized (by Suzhou Hongxun Biotechnology Co., Ltd. and Beijing Qingke Biotechnology Co., Ltd.) and respectively constructed into the ABE8e plasmid (Addgene, Plasmid #138489) at positions 466-5160 to construct the Cas12o1 expression vector (Figure 3A) and the LbCpf1 expression vector (Figure 3B).

[0467] The Cas12o1-hTTR1-crRNA expression sequence fragment and the LbCpf1-hTTR1-crRNA expression sequence fragment (synthesized by Suzhou Hongxun Biotechnology Co., Ltd.) were treated with double enzyme digestion (MfeI / MluI), and then inserted into the Cas12o1 expression vector backbone and the LbCpf1 expression vector backbone obtained by double enzyme digestion (MfeI / MluI) to obtain the Cas12o1-hTTR1-crRNA expression vector and the LbCpf1-hTTR1-crRNA expression vector.

[0468] The araC-pBAD-CCDB fragment (SEQ ID NO. 17) with the target sequence hTTR1 (SEQ ID NO. 12) was designed and synthesized (Suzhou Hongxun Biotechnology Co., Ltd.) and inserted into the 1284-1300 site of the pKESK2 plasmid 2 (Addgene, Plasmid, #64857) to obtain the Target plasmid (SEQ ID NO. NO.11, see Figure 4 for the map). The Target plasmid carries the CCDB gene, which can express the CCDB toxic protein (CCDB toxic protein acts as a DNA gyrase inhibitor, locking the DNA gyrase and broken double-stranded DNA complex, rendering the DNA gyrase unable to function and ultimately leading to cell death). The L-arabinose-induced PBAD promoter can regulate the expression of the CCDB gene. When the target sequence hTTR1 on the Target plasmid is cleaved, the regulatory expression pathway between the PBAD promoter and the CCDB toxic protein is interrupted, and the host cell will not produce the ccdB toxic protein and survive. Conversely, if the hTTR1 target sequence on the Target plasmid is not cleaved, the PBAD promoter regulates the CCDB gene to express the ccDB toxic protein, leading to host cell death. Therefore, the bacterial survival ratio can indicate the cleavage activity.

[0469] The Target plasmid was transfected into DH5a competent cells, and then the Cas12o1-hTTR1-crRNA expression vector and the LbCpf1-hTTR1-crRNA expression vector were transfected into DH5a competent cells carrying the Target plasmid, respectively.

[0470] Through analysis, it was found that both Cas12o1 and LbCpf1 had obvious cleavage activity, and the cleavage activity of Cas12o1 was much higher than that of LbCpf1 (Figure 5).

[0471] Example 3 Cleavage activity of Cas12o1, Cas12o2, and Cas12o3 in mammalian cells

[0472] To verify the cleavage activity of Cas12o1, Cas12o2, and Cas12o3 in mammalian cells, the nucleotide coding sequence of Cas12o1 (SEQ ID NO.16), the nucleotide coding sequence of Cas12o2 (SEQ ID NO.40), and the nucleotide coding sequence of Cas12o3 (SEQ ID NO.46) were respectively constructed into pcDNA3.1 (+) (Invitrogen, V79020) to construct each Cas protein expression vector. hTTR was selected as the target, and the corresponding crRNA containing the hTTR spacer sequence was constructed into the pGL3-U6-sgRNA-EGFP (Plasmid #107721) plasmid to obtain the crRNA expression vector.

[0473] The target sequence and crRNA are shown in the following table:

[0474] The Cas protein expression vector, crRNA expression vector and Target plasmid were transfected into DH5a competent cells for large-scale preparation, and the concentrations were measured and stored for later use.

[0475] Approximately 16 hours before transfection, HEK293T cells were plated into 24-well plates at a density of 2 × 10 cells per well. 5 (500 μL). Cas protein expression vector, crRNA expression vector, and EGFP-C1 plasmid were mixed separately and then washed with 25 μl of Dilute the transfection-specific reduced serum medium (Source Bio, L530KJ) and add 2 μl of Lipofectamine 3000 (Invitrogen, L3000015) reagent, pipette and mix well as reagent A, and let it stand for 5 minutes. At the same time, 2 μl of Lipofectamine 3000 transfection reagent (Invitrogen, L3000015) was added with 25 μl of Dilute and mix the transfection-specific reduced serum medium (Source Biotechnology, L530KJ) as reagent B and let it stand for 5 minutes.

[0476] The above reagents A and B were mixed and pipetted evenly, and allowed to stand for 20 minutes. After standing, the mixed reagents were added dropwise to the 24-well plate cells to be transfected. The plasmid dosages for each well of the 24-well plate were 0.3 μg of Cas protein expression vector, 0.3 μg of crRNA expression vector, and 0.3 ug of EGFP-C1 plasmid. Return to 37 ° C, 5% CO2 incubator for culture. After 6 hours of transfection, the culture medium was replaced with DMEM medium containing 10% FBS. After 48 hours of transfection, EGFP fluorescent protein expression indicated that the cells were successfully transfected, and cells positive for EGFP expression were sorted for editing efficiency detection. The cells were subjected to genomic extraction (using a genomic DNA extraction kit, TIANGEN, DP304-03), and the PCR products after PCR amplification were used for high-throughput deep sequencing (Qingke Biotechnology Co., Ltd.) or Sanger sequencing (Boshang Biotechnology (Shanghai) Co., Ltd.) to identify the editing efficiency.

[0477] The researchers found that the average insertion / deletion percentages of Cas12o1, Cas12o2, and Cas12o3 in the hTTR target were 16.11%, 5.37%, and 11.89%, respectively, demonstrating that Cas12o1, Cas12o2, and Cas12o3 can achieve efficient cleavage in mammalian cells under the guidance of guide RNA.

[0478] As can be seen from Example 1, the cleavage efficiency of Cas12o1 disclosed in the present invention in Escherichia coli is much higher than that of LbCpf1. Example 2 proves that the CRISPR-Cas12o system disclosed in the present invention can achieve multifunctional and efficient genome editing in mammalian cells. And because the Cas12o series (Cas12o1, Cas12o2, Cas12o3) has a smaller size, simple structure, shorter crRNA and self-processing properties, it is suitable for delivery methods including AAV or LNP, and can be used for multiple gene editing applications in vivo or in vitro in the future.

[0479] Sequence information:

[0480] Without departing from the scope and spirit of the present disclosure, various modifications and variations of the methods, pharmaceutical compositions and kits described in the present disclosure will be apparent to those skilled in the art. Although the present disclosure has been described in conjunction with specific embodiments, it will be understood that the present disclosure is capable of further modifications, and the disclosure claimed for protection should not be unduly limited to such specific embodiments. In fact, various variations of the described modes for implementing the present disclosure that are apparent to those skilled in the art are intended to fall within the scope of the present disclosure. This application is intended to cover any variations, uses or changes that are generally consistent with the principles of the present disclosure, including those that do not fall within the scope of the present disclosure but are known and commonly used technical means in the field to which the present disclosure belongs and that can be applied to the essential features set forth above.

Claims

1. A Cas protein, characterized in that Includes OBD domain, REC domain, RuvC domain, Helical domain, and Nuc domain; Optionally, the RuvC domain comprises RuvC-I, RuvC-II and RuvC-III domains; Optionally, the Cas protein does not comprise a HNH domain and a PI domain; Optionally, the RuvC-III domain is located between the Nuc-I domain and the Nuc-II domain; Optionally, the OBD domain is a bi-split domain, including an OBD-I domain and an OBD-II domain, wherein the OBD-I domain is located at the N-terminus, and the Nuc-II domain is located at the C-terminus; Optionally, the Cas protein performs nucleic acid cleavage function without the aid of tracrRNA.

2. The Cas protein according to claim 1, characterized in that The Cas protein has the following domains from N-terminus to C-terminus: A1-A2-A3-A4-Z5 (I), Among them, A1 is the REC domain; A2 is the OBD domain; A3 is the RuvC domain; A4 is the Helical domain; Z5 contains the RuvC domain and the Nuc domain; And, each "-" is independently a bond or a linker; Optionally, the OBD domain comprises an OBD-II domain; Optionally, the REC domain comprises a REC-I domain and a REC-II domain; Optionally, the RuvC domain comprises a RuvC-I domain, a RuvC-II domain, and a RuvC-III domain; Optionally, Z5 has the structure shown in Formula II: Y1-Y2-Y3-Y4(II), Among them, Y1 is the RuvC-II domain; Y2 is the Nuc-I domain; Y3 is the RuvC-III domain; Y4 is the Nuc-II domain; Furthermore, each "-" is independently a bond or a linker.

3. The Cas protein according to claim 1, characterized in that The Cas protein has the following structure from N-terminus to C-terminus: Z1-Z2-Z3-Z4-X-Z5 (III), Among them, Z1 is none or OBD-I domain; Z2 is the REC domain; Z3 is the OBD-II domain; Z4 is the RuvC-I domain; X is the Helical domain; Z5 contains the RuvC domain and the Nuc domain; And, each "-" is independently a bond or a linker; Optionally, the REC domain comprises a REC-I domain and a REC-II domain; Optionally, the RuvC domain is selected from the group consisting of a RuvC-II domain, a RuvC-III domain, or a combination thereof; Optionally, the Nuc domain is selected from the group consisting of a Nuc-I domain, a Nuc-II domain, or a combination thereof; Optionally, Z5 has the structure shown in Formula II: Y1-Y2-Y3-Y4(II); Among them, Y1 is the RuvC-II domain; Y2 is the Nuc-I domain; Y3 is the RuvC-III domain; Y4 is the Nuc-II domain; Furthermore, each "-" is independently a bond or a linker.

4. The Cas protein according to claim 1, characterized in that The OBD domain, REC domain, RuvC domain, Helical domain, Nuc domain has an amino acid sequence that has at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100%) sequence identity with the amino acid sequence of the OBD domain, REC domain, RuvC domain, Helical domain, Nuc domain in Table 1; Optionally, the Cas protein is about 500 to about 1200 amino acids in length, optionally, the Cas protein comprises about 700 and 1100 amino acids, more optionally, the Cas protein comprises about 900 and 1000 amino acids; Optionally, the Cas protein is a Class 2 V-type Cas endonuclease; Optionally, the Cas protein comprises a sequence having at least 70%, at least 75%, at least 80% or at least 90% sequence identity to any one of SEQ ID NOs: 1, 3, 5, 7-9.

5. A fusion protein, characterized in that Comprising the Cas protein of claim 1; and one or more functional domains; Optionally, the functional domain functional domain is selected from a localization signal, a reporter protein, a Cas protein targeting portion, a DNA binding domain, an epitope tag, a transcription activation domain, a transcription repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcription release factor, an HDAC, a cleavage-active polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, and a base excision repair inhibitor (such as a uracil-DNA glycosylase inhibitor (UGI)).

6. The fusion protein according to claim 5, characterized in that The functional domain includes one or more of the following enzymatic activities on the target sequence: Methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity; Optionally, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain; Optionally, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain comprises one or more of ADAR1, ADAR2, APOBEC, AID or TAD; Optionally, the adenosine deaminase catalytic domain comprises an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO:30, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO:30; Optionally, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions and substitutions relative to the amino acid sequence shown in SEQ ID NO: 30; Optionally, the adenosine deaminase catalytic domain comprises a mutant of the amino acid sequence shown in SEQ ID NO: 30: E18K+F19S+N20L, named adenosine deaminase 004V14 (see WO2023193536A1); Optionally, the adenosine deaminase catalytic domain comprises an amino acid sequence that is at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 31 (selected from 005V1 deaminase in CN114634923A, in which the amino acid sequence is SEQ ID NO: 2), and it retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 31; Optionally, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions and substitutions relative to the amino acid sequence shown in SEQ ID NO: 31; Optionally, the adenosine deaminase catalytic domain comprises a mutant of the amino acid sequence shown in SEQ ID NO: 31: Q148G+Q149M+P150R, named deaminase 005V1-10-3; Optionally, the functional domain is the full length or functional fragment of TadA8e; Optionally, the localization signal comprises a nuclear localization signal (NLS) and / or a nuclear export signal (NES); Optionally, the sequence of the nuclear localization signal is shown in any one of SEQ ID NOs: 22-29, 32-39, 41-45; Optionally, the sequence of the nuclear localization signal is located at, near or close to the end (e.g., N-terminus or C-terminus) of the Cas protein of claim 1; Optionally, the nuclear export signal comprises protein tyrosine kinase 2 (such as human protein tyrosine kinase 2); Optionally, the reporter protein comprises glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, autofluorescent protein; Optionally, the autofluorescent protein includes green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent protein (e.g., (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire); Optionally, the DNA binding domain includes methylation binding protein, LexADBD, Gal4DBD; Optionally, the epitope tag comprises a histidine tag, a V5 tag, a FLAG tag, an influenza virus hemagglutinin tag, a Myc tag, a VSV-G tag, a thioredoxin tag, a streptavidin tag; Optionally, the transcriptional activation domain comprises VP64 and / or VPR; Optionally, the transcriptional repression domain comprises KRAB and / or SID; Optionally, the nuclease comprises FokI; Optionally, the cleavage active polypeptide includes a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity; Optionally, the ligase comprises DNA ligase and / or RNA ligase; Optionally, the functional domain is connected to the N-terminus and / or C-terminus of the Cas protein; Optionally, the functional domain is inserted between the N-terminus and the C-terminus of the Cas protein; Optionally, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the Cas protein via a linker; Optionally, the functional domain is inserted between the N-terminus and the C-terminus of the Cas protein via a linker.

7. The fusion protein according to claim 5, characterized in that The fusion protein has the following structure from N-terminus to C-terminus: Z1-Z2 (I'); or Z2-Z1 (II'); or Z3-Z1-Z4 (III'); wherein Z1 is cytosine deaminase or adenosine deaminase; Z2 is the Cas protein according to claim 1; Z3 is the N-terminal fragment of the Cas protein according to claim 1; Z4 is the C-terminal fragment of the Cas protein according to claim 1; Furthermore, each "-" is independently a bond or a linker.

8. An isolated polynucleotide, characterized in that The polynucleotide encodes the Cas protein according to claim 1 or the fusion protein according to claim 5; Optionally, the isolated nucleotide comprises a sequence optimized for humanization; Optionally, the polynucleotide further contains auxiliary elements selected from the following groups on the flank of the ORF of the variant: a signal peptide, a secretory peptide, a tag sequence (such as 6His), or a combination thereof; Optionally, the polynucleotide is selected from the group consisting of a genomic sequence, a cDNA sequence, an RNA sequence, or a combination thereof; Optionally, the polynucleotide further comprises a promoter operably linked to the ORF sequence of the variant; Optionally, the promoter is selected from the group consisting of a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter; Optionally, the polynucleotide has been codon-optimized for expression in eukaryotic cells; Optionally, the polynucleotide is a polynucleotide that is codon-optimized according to the codon preference of the host cell; Optionally, the host cell comprises a prokaryotic cell or a eukaryotic cell; Optionally, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell or a mammalian cell (including human and non-human mammals); Optionally, the host cell is a prokaryotic cell, such as Escherichia coli; Optionally, the yeast cell is selected from one or more sources of yeast selected from the group consisting of Pichia pastoris, Kluyveromyces, or a combination thereof; Optionally, the yeast cell comprises: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis; Optionally, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof; Optionally, the polynucleotide comprises a sequence having at least 70%, at least 75%, at least 80% or at least 90% sequence identity to any one of SEQ ID NO: 16, 40 or 46.

9. A guide RNA (gRNA), characterized in that The guide RNA comprises (i) capable of binding to the direct repeat (DR) sequence of the Cas protein according to claim 1 and (ii) a spacer sequence capable of targeting a target sequence of a target DNA, wherein the guide RNA is configured to form a complex with the Cas protein.

10. A carrier, characterized in that Comprising the polynucleotide of claim 9.

11. A composite, characterized in that Include: (i) a protein component selected from the group consisting of the Cas protein of claim 1, the fusion protein of claim 5, or a combination thereof; and (ii) a nucleic acid component selected from the group consisting of the guide RNA of claim 9, a nucleic acid encoding the guide RNA of claim 9, a precursor RNA of the guide RNA of claim 9, a precursor RNA nucleic acid encoding the guide RNA of claim 9, or a combination thereof; the protein component and the nucleic acid component are combined with each other to form a complex; Wherein, the guide RNA comprises: (iii) a direct repeat (DR) sequence capable of binding to the Cas protein of claim 1 and (iv) a spacer sequence capable of targeting a target sequence of a target DNA.

12. A CRISPR-Cas composition, characterized in that Include: (i) a first component selected from the group consisting of the Cas protein of claim 1, the fusion protein of claim 5, a nucleotide sequence encoding the Cas protein of claim 1 or the fusion protein of claim 5, and any combination thereof; and (ii) a second component, wherein the second component is a nucleotide sequence comprising one or more guide RNAs according to claim 9, or encoding the nucleotide sequence comprising one or more guide RNAs according to claim 9; the guide RNA comprises: (iii) a direct repeat (DR) sequence capable of binding to the Cas protein of claim 1 and (iv) a spacer sequence capable of targeting a target sequence of a target DNA, wherein the guide RNA is configured to form a complex with the Cas protein; Optionally, the composition comprises a pharmaceutical composition; Optionally, the dosage form of the composition is selected from the group consisting of a lyophilized preparation, a liquid preparation, or a combination thereof; Optionally, the composition is in the form of a liquid preparation; Optionally, the composition is in the form of an injection; Optionally, the composition is a cell preparation.

13. A CRISPR-Cas system, characterized in that: Comprising one or more vectors, the one or more vectors comprising: (i) a first nucleic acid, which is a nucleotide sequence encoding the Cas protein of claim 1 or the fusion protein of claim 5; optionally, the first nucleic acid is operably linked to a first regulatory element; as well as (ii) a second nucleic acid encoding the nucleotide sequence of the guide RNA according to claim 9; Optionally, the second nucleic acid is operably linked to a second regulatory element; The guide RNA comprises: (iii) a direct repeat (DR) sequence capable of binding to the Cas protein of claim 1 and (iv) a spacer sequence capable of targeting a target sequence of a target DNA, wherein the guide RNA is configured to form a complex with the Cas protein; The first nucleic acid and the second nucleic acid are present on the same or different vectors.

14. The complex of claim 11, the CRISPR-Cas composition of claim 12, or the CRISPR-Cas system of claim 13, wherein: in, The length of the spacer sequence is greater than 17 nucleotides, preferably, 17 to 100 nucleotides, more preferably 16 to 50 nucleotides (e.g., 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides), more preferably 17 to 50 nucleotides, more preferably 17 to 40 nucleotides, more preferably 18 to 39 nucleotides, most preferably 18 to 37 nucleotides; Optionally, the target DNA is selected from double-stranded DNA, single-stranded DNA, RNA, genomic DNA and extrachromosomal DNA; Optionally, the spacer sequence is connected to the 3' end of the direct repeat (DR) sequence; Optionally, the spacer sequence comprises a complementary sequence to the target sequence; Optionally, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM is 5'-TN, wherein N is A, T, G or C; Optionally, the target sequence is a DNA from a prokaryotic cell or a eukaryotic cell, or a DNA sequence formed based on RNA reverse transcription; or, the target sequence is a non-naturally occurring DNA, or a DNA sequence formed based on RNA reverse transcription; Optionally, the target DNA is located inside or outside the cell; Optionally, the cell is a eukaryotic cell or a prokaryotic cell; Optionally, the cell is selected from animal cells, plant cells, fungal cells; Optionally, the eukaryotic cell is a plant cell, a mammalian cell, an insect cell, an arthropod cell, a fungal cell, a bird cell, a reptile cell, an amphibian cell, an invertebrate cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, or a human cell; Optionally, a DNA donor template is also included, which can be inserted into the locus of interest by homology-directed repair (HDR); Optionally, the donor template nucleic acid has a length of 8-1000 nucleotides; Optionally, the donor template nucleic acid has a length of 25-500 nucleotides; Optionally, the vector comprises a plasmid or a viral vector; Optionally, the guide RNA includes unmodified and modified guide RNA; Optionally, the modified guide RNA includes chemical modifications of bases; Optionally, the chemical modification includes methylation modification, methoxy modification, fluorination modification or thio modification; Optionally, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter, a constitutive promoter, a ubiquitous promoter, a cell type specific promoter or a tissue specific promoter; Optionally, at least one component of the composition is non-naturally occurring or modified; Optionally, the DR sequence comprises a structure as shown in Formula IV below: 5'-R1a-Ba-R2a-L-R2b-Bb-R1b-3' (IV), wherein segments R1a and R1b are reverse complementary sequences and form a first stem (R1) having a plurality (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) of nucleotide pairs in the Cas protein; Segments Ba and Bb do not base pair with each other and form a bulge (B); Segments R2a and R2b are reverse complementary sequences and form a second stem (R2) having a plurality of (2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10) base pairs; and L is a loop formed at the second stem portion and formed by a plurality of (3, 4, 5, 6, 7, 8, 9, 10) nucleotides; Optionally, the DR sequence has a secondary structure substantially identical to the secondary structure of the DR sequence shown in any one of SEQ ID NOs: 2, 4, and 6; Optionally, the DR sequence has nucleotide additions, insertions, deletions or substitutions that do not result in substantial differences in secondary structure compared to the DR sequence shown in any one of SEQ ID NOs: 2, 4, and 6; Optionally, the DR sequence comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence shown in any one of SEQ ID NOs: 2, 4, and 6; (ii) a sequence having one or more base substitutions, deletions or additions (e.g., substitutions, deletions or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases) compared to the sequence shown in any one of SEQ ID NOs: 2, 4, 6; (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to any of SEQ ID NOs: 2, 4, and 6; (iv) a sequence that hybridizes to the sequence described in any one of (i) to (iii) under stringent conditions; or (v) a complementary sequence of the sequence described in any one of (i) to (iii); Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived.

15. A kit, characterized in that: Comprising one or more components selected from the following: the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, or the CRISPR-Cas system of claim 13; Optionally, the kit further comprises a label or instructions; Optionally, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, and cleaving a target gene or a non-target gene.

16. A delivery composition, characterized in that Comprising a delivery vector or a delivery medium, and one or more selected from the following: the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, or the CRISPR-Cas system of claim 13; Optionally, the delivery vehicle is a particle; Optionally, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses); Optionally, the delivery medium comprises nanoparticles, liposomes, exosomes, microvesicles, an electroporation device or a gene gun.

17. A host cell comprising the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, or the delivery composition of claim 16; Optionally, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell or a mammalian cell (including human and non-human mammals); Optionally, the host cell is a prokaryotic cell, such as Escherichia coli; Optionally, the yeast cell is selected from one or more sources of yeast selected from the group consisting of Pichia pastoris, Kluyveromyces, or a combination thereof; Preferably, the yeast cell comprises: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis; Optionally, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

18. An enzyme preparation, characterized in that The enzyme preparation comprises the Cas protein of claim 1, the fusion protein of claim 5, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, or the delivery composition of claim 16; Optionally, the enzyme preparation includes an injection and / or a lyophilized preparation.

19. A medicine box, characterized in that: include: A first container, and the CRISPR-Cas composition of claim 12 or the CRISPR-Cas system of claim 13 located in the first container, or containing the CRISPR-Cas composition of claim 12 or the CRISPR-Cas system of claim 13; Optionally, the first container comprises the CRISPR-Cas composition of claim 12 or the CRISPR-Cas system of claim 13; Optionally, the composition is a pharmaceutical composition; Optionally, the dosage form of the pharmaceutical composition is selected from the group consisting of a lyophilized preparation, a liquid preparation, or a combination thereof; Optionally, the pharmaceutical composition is in an oral dosage form or an injectable dosage form; Optionally, the kit further comprises instructions.

20. A medicine box, characterized in that: include: (a1) a first container, and the Cas protein according to claim 1, or the fusion protein according to claim 5, or a gene encoding the Cas protein, or an expression vector thereof, or a drug containing the Cas protein according to claim 1, or the fusion protein according to claim 5, or a gene encoding the Cas protein, or an expression vector thereof, located in the first container; (b1) an optional second container, and the guide RNA or its expression vector according to claim 9, or a drug containing the guide RNA or its expression vector according to claim 9, located in the second container; Optionally, the first container and the second container are different containers; Optionally, the drug in the first container is a single preparation containing the Cas protein of claim 1, or the fusion protein of claim 5, or its encoding gene or its expression vector; Optionally, the drug in the second container is a single preparation containing the guide RNA or its expression vector as claimed in claim 9; Optionally, the dosage form of the drug is selected from the group consisting of a lyophilized preparation, a liquid preparation, or a combination thereof; Optionally, the drug is in an oral dosage form or an injectable dosage form; Optionally, the kit further comprises instructions.

21. A method for targeting and editing a target gene or cleaving a target gene, characterized in that: include: The Cas protein of claim 1, the fusion protein of claim 5, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the delivery composition of claim 16, the enzyme preparation of claim 18, or the kit of claim 19 is contacted with the target gene, or delivered to a cell comprising the target gene, wherein the target sequence is present in the target gene; Optionally, the target gene is present in a cell; Optionally, the cell is a prokaryotic cell; Optionally, the cell is a eukaryotic cell, such as a mammalian cell (eg, a human cell) or a plant cell; Optionally, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro; Optionally, the editing of the target gene or the cleavage of the target gene comprises a break in the target sequence, such as a double-strand break in DNA or a single-strand break in RNA, or inserting an exogenous nucleic acid into the break; Optionally, the target gene comprises DNA; Optionally, the DNA includes single-stranded DNA and double-stranded DNA; Optionally, the method is a non-diagnostic and non-therapeutic method.

22. A method for inducing a change in cell state, characterized in that: The method comprises contacting the Cas protein of claim 1, or the fusion protein of claim 5, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the delivery composition of claim 16, the enzyme preparation of claim 18, or the drug kit of claim 19 with a target gene in a cell.

23. A method for altering the expression of a gene product, characterized in that include: contacting the Cas protein of claim 1, or the fusion protein of claim 5, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the delivery composition of claim 16, the enzyme preparation of claim 18, or the kit of claim 19 with a nucleic acid molecule encoding the gene product, or delivering it to a cell comprising the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule; Optionally, the nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro; Optionally, expression of the gene product is altered (e.g., increased or decreased); Optionally, the gene product is a protein; Optionally, the protein, fusion protein, polynucleotide, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vehicle; Optionally, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses); Optionally, the cell, cell line or organism is modified by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

24. A cell or progeny thereof obtained by the method of any one of claims 21 to 23, wherein the cell comprises a modification not present in its wild type.

25. A cell product of the cell of claim 24 or its progeny.

26. An in vitro, ex vivo or in vivo cell or cell line or progeny thereof, characterized in that The cell or cell line or their progeny comprises: the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13 or the delivery composition of claim 16; Optionally, the cell is a prokaryotic cell; Optionally, the cell is a eukaryotic cell, such as a mammalian cell (eg, a human cell) or a plant cell; Optionally, the cell is a stem cell or a stem cell line.

27. A cell preparation, characterized in that The host cell of claim 17 or the cell of claim 24 or its progeny, or the cell product of the cell of claim 25 or its progeny, or the cell or cell line of claim 26 or their progeny; Optionally, the cell preparation further comprises a pharmaceutically acceptable carrier or excipient; Optionally, the cell preparation includes an injection and / or a lyophilized preparation.

28. Use of the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the kit of claim 15, the delivery composition of claim 16, the enzyme preparation of claim 18 or the kit of claim 19, characterized in that: For preparing a drug or preparation for nucleic acid editing; Optionally, the nucleic acid editing comprises gene or genome editing; Optionally, the gene or genome editing comprises modifying a gene, knocking out a gene, altering the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide.

29. Use of the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the kit of claim 15, the delivery composition of claim 16, the enzyme preparation of claim 18 or the kit of claim 19, characterized in that: For use in the preparation of a drug or preparation, the drug or preparation is used for one or more selected from the following groups: (i) ex vivo gene or genome editing; (ii) Detection of single-stranded DNA in vitro; (iii) editing a target sequence in a target locus to modify an organism or non-human organism; (iv) treating a disorder caused by a defect in the target sequence in the target locus; (v) treating a condition or disease in a subject in need thereof.

30. A method of treating a condition or disease in a subject in need thereof, characterized in that Comprising administering to the subject the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the kit of claim 15, the delivery composition of claim 16, the enzyme preparation of claim 18, or the drug kit of claim 19.

31. The method of claim 30, wherein the condition or disease comprises a metabolic disease, cancer, a neurological disease, an ophthalmic disease, and an infectious disease; Optionally, the condition or disease comprises a genetic disease; Optionally, the condition or disease is caused by a pathogenic point mutation; Optionally, the condition or disease comprises hypercholesterolemia (FH), atherosclerosis (ASCVD), Transthyretin amyloidosis (ATTR), alpha-1 antitrypsin deficiency (AATD), primary hyperoxaluria (PH1), hereditary angioedema (HAE), and hepatitis B; Optionally, the disease or disorder is a disease caused by a pathogenic point mutation.

32. A method for detecting whether a target nucleic acid molecule exists in a sample, characterized in that: The method comprises contacting the sample with the Cas protein of claim 1, the fusion protein of claim 5, or the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the kit of claim 15, the delivery composition of claim 16, or the enzyme preparation of claim 18, and a non-target sequence, detecting a detectable signal generated by cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA; Optionally, the non-target sequence is cleaved by a protein in the complex or CRISPR-Cas composition or system or delivery composition, indicating the presence of a target nucleic acid molecule in the sample; and the non-target sequence is not cleaved by a protein in the complex or CRISPR-Cas composition or system or delivery composition, indicating the absence of a target nucleic acid molecule in the sample; Optionally, the target nucleic acid molecule is a target DNA; Optionally, the target DNA includes DNA formed based on RNA reverse transcription; Optionally, the target DNA comprises cDNA; Optionally, the target DNA is selected from the group consisting of single-stranded DNA, double-stranded DNA, or a combination thereof.

33. A sterile container, characterized in that: It comprises the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the delivery composition of claim 16 or the enzyme preparation of claim 18; Optionally, the sterile container is a syringe.

34. An implantable device, characterized in that It comprises the Cas protein of claim 1, the fusion protein of claim 5, the polynucleotide of claim 8, the vector of claim 10, the complex of claim 11, the CRISPR-Cas composition of claim 12, the CRISPR-Cas system of claim 13, the delivery composition of claim 16 or the enzyme preparation of claim 18; Optionally, the Cas protein, the fusion protein, the polynucleotide, the complex, the vector, the CRISPR-Cas composition or the system or the delivery composition or the enzyme preparation is stored in a reservoir.

Citation Information

Patent Citations

  • Engineered CAS12B effector proteins and methods of use thereof

    CN116254246A

  • Split CAS12 system and method of use thereof

    CN117120602A

  • Split cas12 systems and methods of use thereof

    US20230323322A1

  • Engineered proteins and methods of use thereof

    US20230383271A1

  • Engineered CAS effector proteins and methods of use thereof

    WO2022120520A1

Cited By

  • Method for detecting typical pathogenic vibrio of laver

    CN121272070A