Modified CPF1 guide RNA

By adding an extended sequence at the 5' end of Cpf1 crRNA, the negative charge density and structured design are enhanced, and the problem of low delivery efficiency of RNA-guided endonuclease is solved, achieving efficient gene editing effect.

CN120290559APending Publication Date: 2025-07-11GENEDIT INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510240123.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-07-12
Filing Date
2018-10-02
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deliver RNA-guided endonuclease into cells, resulting in inefficient genome modification.

Method used

A nucleic acid containing Cpf1 crRNA was designed to enhance negative charge density and structured design by adding an extended sequence to its 5' end to improve delivery efficiency.

Benefits of technology

It significantly improves the delivery efficiency and gene editing effect of Cpf1 crRNA in cells, enhances the frequency of NHEJ and HDR, and is suitable for gene modification in vitro and in vitro.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290559A_ABST
    Figure CN120290559A_ABST
Patent Text Reader

Abstract

The present invention provides a nucleic acid comprising Cpf1 crRNA, a processing sequence of Cpf1 crRNA 5 ', and an extension sequence of processing sequence 5'. The invention also provides compositions comprising a nucleic acid, a vector, and optionally Cpf1. In addition, the present invention provides a method of genetically modifying an eukaryotic target cell comprising contacting the eukaryotic target cell with a nucleic acid or a composition to genetically modify the target nucleic acid in the cell.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention with the application date of October 2, 2018, application number 201880077637.7, and invention title "Modified CPF1 Guide RNA".

[0002] Materials submitted electronically and introduced as a reference

[0003] Cross-reference to related applications

[0004] This patent application claims priority to U.S. Provisional Patent Application No. 62 / 567,123, filed on October 2, 2017; U.S. Provisional Patent Application No. 62 / 617,138, filed on January 12, 2018; and U.S. Provisional Application No. 62 / 697,327, filed on July 12, 2018. The entire disclosure of the above applications is incorporated herein by reference. Background Art

[0005] RNA-guided endonucleases have proven to be effective tools for genome editing in multiple cell types and microorganisms. RNA-guided endonucleases generate site-specific double-stranded DNA breaks or single-stranded DNA breaks within the target nucleic acid. When cleavage of the target nucleic acid occurs intracellularly, the break in the nucleic acid can be repaired by non-homologous end joining (NHEJ) or homology-directed repair (HDR).

[0006] RNA-guided endonucleases and their gene editing components (e.g., guide RNAs) are directly delivered into cells both in vitro and in vivo, and have great potential as a therapeutic strategy for treating genetic diseases. However, currently, it is challenging to directly deliver these components into cells with reasonable efficiency.

[0007] Accordingly, there is a need to identify new compositions and related methods to improve the cellular delivery and other properties of RNA-guided endonucleases that enhance genome editing. The present invention provides such compositions and related methods. Summary of the Invention

[0008] The present invention is a nucleic acid comprising a Cpf1 crRNA having an extended sequence. In one aspect, the nucleic acid comprises a Cpf1 crRNA and an extended sequence at the 5'-end of the Cpf1 crRNA, wherein the extended sequence comprises less than about 60 nucleotides. In another aspect, the nucleic acid comprises a Cpf1 crRNA, a processing sequence 5' of the Cpf1 crRNA, and an extended sequence 5' of the processing sequence. Also provided are compositions comprising the nucleic acid, a vector, and optionally a Cpf1 protein or a vector encoding the protein.

[0009] The present invention also provides methods for genetically modifying eukaryotic target cells. The methods involve contacting the eukaryotic target cells with a nucleic acid comprising a Cpf1 crRNA as described herein.

[0010] These and other aspects of the invention are described in more detail in the following sections. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1A is a graph comparing delivery of unmodified crRNA complexed with Cpf1 using cationic liposomes (lipofectamine) or electroporation (nucleofection).

[0012] Figure 1B and 1C is a graph comparing cationic liposome-mediated delivery of Cpf1-crRNA complexes and NHEJ generation, the Cpf1-crRNA complexes comprising unmodified crRNA (41 nucleotides in length) and extended crRNA.

[0013] Figure 1D is a schematic diagram of the structures of 41-nucleotide (nt) unmodified crRNA and extended crRNA. Arrows represent Cpf1 cleavage sites.

[0014] Figure 1E is a graph showing Figure 1D the NHEJ efficiency for the crRNA constructs shown in

[0015] Figure 1F is a graph depicting cellular delivery of Cpf1 RNPs using cationic liposomes and crRNAs of different lengths labeled with a fluorescent dye.

[0016] Figure 2 provides a graph showing gene editing efficiency of GFP knock-down according to 5'-extended crRNA delivered via electroporation to GFP-HEK cells (right panel), and a schematic diagram of the crRNA (left panel).

[0017] Figure 3A is a schematic diagram of an in vivo study in AI9 mice.

[0018] Figure 3B is a schematic diagram of the gastrocnemius injection site and imaging section.

[0019] Figure 4A is a graph of HDR frequencies for crRNAs of different extensions delivered with donor DNA using electroporation.

[0020] Figure 4BGraph of the percentage of GFP-cells with crRNAs of different lengths delivered together with donor DNA, showing the NHEJ efficiency using electroporation.

[0021] Figure 4C Graph of the percentage of BFP-cells with crRNAs of different lengths delivered together with donor DNA, showing the NHEJ efficiency using electroporation.

[0022] Figure 4D Graph of the percentage of GFP-cells with extended crRNAs delivered using electroporation, with and without single-stranded DNA (ssDNA) that is not homologous to the target sequence.

[0023] Figure 4E Graph showing gene editing as the percentage of GFP-cells with crRNAs using electroporation, where the crRNAs are extended with 100 nt RNA and 9 nt RNA that are not homologous to the target sequence.

[0024] Figure 4F Graph of the percentage of GFP-cells with crRNAs using electroporation, where the crRNAs have and do not have a 4 nt extension and are further modified with a chemical moiety.

[0025] Figure 5A and 5B Provides a schematic diagram of conjugating crRNA and donor DNA.

[0026] Figure 5C Image of gel electrophoresis separation showing the release of donor DNA and crRNA from the conjugated crRNA / DNA molecules after thiol reduction.

[0027] Figure 6 Graph confirming that Cpf1 conjugated to HD-RNA induces NHEJ in GFP-HEK cells after transfection with PAsp(DET) (i.e., cationic polymer).

[0028] Figure 7 Graph confirming that Cpf1 conjugated to HD-RNA induces HDR in GFP-HEK cells after transfection with PAsp(DET) (i.e., cationic polymer).

[0029] Figure 8 Provides the sequence of the Cpf1 protein.

[0030] Figure 9 Provides examples of Cpf1 processing sequences.

[0031] Figure 10AShows a schematic diagram of crRNA conjugated with donor DNA.

[0032] Figure 10B Shows the sequences used in crRNA conjugated with donor DNA.

[0033] Figure 10C Is a graph of the percentage of RFP+ cells after treatment with various crRNAs and Cpf1 using electroporation in primary Ai9 myoblasts.

[0034] Figure 10D Is a graph of the percentage of RFP cells transfected with 100 nt DNA or RNA into primary Ai9 myoblasts.

[0035] Figure 10E Is a graph of the NHEJ efficiency in HepG2 cells transfected with Cpf1 RNP with or without a 9 nt extension of crRNA targeting the Serpina1 gene using electroporation.

[0036] Figure 11A Shows the RNA structures that can be used for crRNA extension.

[0037] Figure 11B Shows the trinucleotide repeats that can be used to provide various RNA structures.

[0038] Figure 11C Shows the intersection of the hybridization extension sequences of crRNA in a kissing loop, which can be used to form crRNA multimers.

[0039] Figure 11D Shows the intersection of the hybridization extension sequences of crRNA to form trimers (inset (i)) or octamers (inset (ii)).

[0040] Figure 12 Shows the editing efficiency (% BPF-) of various Cpf1 crRNAs in HEK293T cells expressing BFP. MS is a 2'-OMe 3'-thiolphosphate modification on the first three nucleotides starting from the 5' end, +9du is a 2'-deoxy modification on the 9th nucleotide starting from the 5' end, and +9S is a thiolphosphate modification on the first 9 nucleotides starting from the 5' end. At 7 days after electroporation, the BFP knockout efficiency was measured by flow cytometry. Mean ± S.E, n = 3. By Student's t-test, all extended crRNAs showed a statistically significant difference from the unmodified crRNA, with a p-value less than 0.05.

[0041] Figure 13A Is a diagram of unmodified Cas9 sgRNA and Cpf1 crRNA; and

[0042] Figure 13B and 13C is a graph showing the relative activities of Cas9 sgRNA and Cpf1 crRNA according to GFP knockdown.

[0043] Figure 14 is a schematic diagram of an extended crRNA modified with biotin and avidin and linked to a targeting molecule containing biotin.

[0044] Figure 15A is a schematic diagram of chemical modifications performed on the extended crRNA.

[0045] Figure 15B is a graph quantifying the amount of remaining crRNA after incubation in serum.

[0046] Figure 15C is a graph of the percentage of GFP-negative cells after delivery of crRNA with a 9-nt extension and chemical modifications using cationic liposomes.

[0047] Figure 15D is a graph comparing the percentage of GFP-negative cells after delivery of chemically modified extended crRNA and unmodified extended crRNA together with Cpf1 using electroporation.

[0048] Figure 16 is a graph comparing cationic polymer-mediated delivery of the Cpf1-crRNA complex and NHEJ generation, where the Cpf1-crRNA complex contains unmodified crRNA (41 nt), 9-base pair extended crRNA (total 50 nt), or 59-base pair extended crRNA (total 100 nt).

[0049] Figure 17 is the corresponding structure of the representative RNA extension sequence (35 nucleotides) listed in Table 1. Detailed Description

[0050] The present invention provides guide nucleic acids for the modification of Cpf1, referred to as "crRNA", which have enhanced properties compared to conventional crRNA molecules. As used herein, crRNA refers to a nucleic acid sequence (e.g., RNA) that binds to an RNA-guided endonuclease Cpf1 and targets the RNA-guided endonuclease to a specific position within a target nucleic acid to be cleaved by Cpf1. Cpf1 is an RNA-guided endonuclease of the type II CRISPR / Cas system, which is involved in type V adaptive immunity. Cpf1 does not require a tracrRNA molecule like other CRISPR enzymes and only requires a single crRNA molecule to function. Cpf1 prefers the "TTN" PAM motif, which is located 5' upstream of its target. Additionally, the cleavage site for Cpf1 is offset by approximately 3-5 bases, which generates "sticky ends" (Kim et al., 2016. "Genome-wide analysis reveals specificities of Cpf1 endonucleases in human cells", published online on June 6, 2016). These sticky ends with 3-5 bp overhangs are thought to facilitate NHEJ-mediated ligation and improve gene editing of DNA fragments with matching ends.

[0051] Those skilled in the art will appreciate that Cpf1 crRNA can be from any species or any synthetic or naturally occurring variant or ortholog derived from or isolated from any source. That is, Cpf1 crRNA can have the required elements of a crRNA (e.g., sequence or structure) that recognize (by binding) any Cpf1 polypeptide or ortholog, or a synthetic variant thereof, from any bacterial species. Examples of Cpf1 crRNA sequences are provided in Figure 9 ; thus, for example, examples of Cpf1 crRNA include those comprising any one of SEQ ID NOs: 21-39). An example of a Cpf1 crRNA sequence for a synthetic variant Cpf1 is the crRNA corresponding to the MAD7 Cpf1 ortholog of Inscripta, Inc. (CO, USA). Another example of a Cpf1 variant is Cpf1 modified to reduce or eliminate ribonuclease activity, e.g., by introducing modifications (such as the H800A, K809A, K860A, F864A, and R790A mutations in Acidaminococcus Cpf1 (AsCpf1), or the H→A mutation at the corresponding position in a different Cpf1 ortholog).

[0052] Generally, a crRNA comprises a targeting domain and a stem-loop domain located 5' of the targeting domain. There is no particular limitation on the overall length of the crRNA, so long as it can direct Cpf1 to a specific location within the target nucleic acid. The stem-loop domain is generally about 19 - 22 nucleotides (nt) in length, while the targeting / guide domain can be anywhere from about 14 - 25 nt (e.g., at least about 14 nt, 15 nt, 16 nt, 17 nt, or 18 nt). In some embodiments, the overall length of the Cpf1 crRNA can be 20 to 100 nt, 20 nt to 90 nt, 20 nt to 80 nt, 20 nt to 70 nt, 20 nt to 60 nt, 20 nt to 55 nt, 20 nt to 50 nt, 20 nt to 45 nt, 20 nt to 40 nt, 20 nt to 35 nt, 20 nt to 30 nt, or 20 nt to 25 nt in length.

[0053] One aspect of the present disclosure provides a nucleic acid comprising a Cpf1 crRNA, an extension sequence 5' of the crRNA, and optionally a processing sequence, which can be located between the crRNA and the extension sequence, within the extension sequence, or 5' of the extension sequence.

[0054] Extension sequence

[0055] The nucleic acid of the invention comprises an extension sequence located 5' of the crRNA. The extension sequence can comprise any combination of nucleic acids (i.e., any sequence). In one embodiment, the extension sequence increases the overall negative charge density of the nucleic acid molecule and improves the delivery of the nucleic acid, including the crRNA.

[0056] In some embodiments, the extension sequence can be cleaved once in a cell. Without being bound by any particular theory or mechanism of action, it is believed that Cpf1 can cleave the extension sequence. However, in certain applications, constructs are desired in which the extension sequence is not cleaved from the Cpf1 crRNA. Thus, in some embodiments, the extension sequence cannot be cleaved by the Cpf1 crRNA. For example, the extension sequence or some portion or region thereof can comprise one or more modified internucleotide linkages (modified "backbone") that are resistant to cleavage by the Cpf1 crRNA (e.g., nuclease resistant). Examples of modified internucleotide linkages include, but are not limited to: phosphorothioates, dithiophosphates, methylphosphonates, aminophosphonates, 2'-O-methyl, 2'-O-methoxyethyl, 2'-fluoro, bridged nucleic acids (BNA), or phosphotriester modified linkages and combinations thereof. The extension sequence or some portion thereof can also comprise synthetic nucleotides, such as xenonucleic acid (XNA) that is nuclease resistant. XNA is a nucleic acid in which the furanose ring of DNA or RNA is replaced with a five- or six-membered modified ribose molecule, such as 1,5-anhydrohexitol nucleic acid (HNA), cyclohexenyl nucleic acid (CeNA), and 2'4'-C-(N-methylaminomethylene) bridged nucleic acid (BNA), 2′-O,4′-C-methylene-β-D-ribonucleic acid or locked nucleic acid (LNA), ANA (arabino nucleic acid), 2'-fluoro-arabino nucleic acid (FANA), and α-L-threofuranosyl nucleic acid (TNA). Additionally, any combination thereof can also be used.

[0057] There is no particular limitation on the length of the extension sequence, as long as the extension sequence increases the overall negative charge density. For example, the extension sequence can have a length of at least about 2 nucleotides (nt) up to about 1000 nt (e.g., at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, or 900, and up to about 1000 nt). In one aspect, the extension sequence has a length of no more than about 100 nucleotides, such as a length of no more than about 80 nucleotides, a length of no more than about 60 nucleotides, or a length of no more than about 40 nucleotides (e.g., a length of no more than about 30 nucleotides or a length of no more than about 20 nucleotides). Any of the foregoing lower and upper limits on length can be expressed as a range. Shorter sequences can also be used (e.g., no more than about 15 nucleotides, or no more than about 10 nucleotides). In some embodiments, the extension sequence comprises at least about 2 nucleotides, such as at least about 4 nucleotides, at least about 6 nucleotides, or even at least about 9 nucleotides. Any of the foregoing can be expressed as a range. Thus, for example, the extension sequence can be about 2-60 nucleotides (e.g., about 2-40 nucleotides, about 2-30 nucleotides, about 2-20 nucleotides, about 2-15 nucleotides, or about 2-10 nucleotides), about 4-60 nucleotides (e.g., about 4-40 nucleotides, about 4-30 nucleotides, about 4-20 nucleotides, about 4-15 nucleotides, or about 4-10 nucleotides); about 6-60 nucleotides (e.g., about 6-40 nucleotides, about 6-30 nucleotides, about 6-20 nucleotides, about 6-15 nucleotides, or about 6-10 nucleotides); or about 9-60 nucleotides (e.g., about 9-40 nucleotides, about 9-30 nucleotides, about 9-20 nucleotides, about 9-15 nucleotides, or about 9-10 nucleotides).

[0058] In some embodiments, the extension sequence has no function other than to confer a greater overall negative charge density to the nucleic acid construct. In this embodiment, for example, the extension sequence is a random or non-coding sequence. In some cases, such as when a processing sequence is used, the sequence can be degraded upon cleavage of the processing sequence and released from the nucleic acid construct.

[0059] In other embodiments, the extension sequence has a function separate and distinct from conferring a greater overall negative charge density on the nucleic acid construct. The extension sequence can have any additional function. For example, the extension sequence can provide a hybridization site for another nucleic acid, such as a donor nucleic acid. Additionally, in some embodiments, the extension sequence can be an aptamer and / or facilitate cell binding. However, sometimes it is not desirable to recruit proteins other than the RNA-guided endonuclease to bind to the guide RNA. Moreover, aptamer sequences typically have complex folding patterns that can be bulky and not compact. Thus, in other embodiments, the extension sequence is not an aptamer sequence.

[0060] In some embodiments, the extension sequence can comprise a sequence encoding a protein whose expression is desired in the target cell to be edited. For example, the extension sequence can comprise a sequence encoding an RNA-guided endonuclease, such as an RNA-guided endonuclease that pairs (i.e., recognizes and is guided by) with the crRNA used in the nucleic acid construct. The extension sequence can comprise, for example, the sequence of the mRNA of an RNA-guided endonuclease.

[0061] In some embodiments, the extension portion self-folds (self-hybridizes) to provide a structured extension portion. There is no limitation on the type of structure provided. The extension portion can have a random coil structure; however, in some embodiments, the extension portion has a structure that is more compact than a random coil structure of the same number of nucleotides, which provides a greater negative charge density. By increasing the overall length of the extension portion, the negative charge of the molecule increases. When a more compact structure is used, the overall negative charge density of the molecule is further increased. The compactness or charge density can be determined based on the mobility in gel electrophoresis. More particularly, if gel electrophoresis is performed on two nucleic acids having the same number of nucleotides that are run together on the same gel, the nucleic acid having the higher mobility (moving the farthest in the gel) is considered to have a more compact structure.

[0062] In another embodiment, the extension sequence comprises at least one semi-stable hairpin structure, stable hairpin structure, pseudoknot structure, G-quadruplex structure, bulge loop structure, internal loop structure, branched loop structure, or a combination thereof. These types of nucleotide structures are known in the art and are described in Figure 11Aare schematically shown. It should be understood that the illustrations are for the purpose of showing only the general structure and are not intended as a detailed illustration of the actual molecular structure. Those skilled in the art will recognize that a hairpin structure, for example, may have scattered non-complementary regions that create "bulges" or other variations in the structure, and other depicted structures may include similar variations. The structure of a given nucleotide sequence can be determined using available algorithms (e.g., "The mfold Web Server" operated by Rensselaer Polytechnic Institute and The RNA Institute, College of Arts and Sciences, State University of New York in Albany; see also M. Zuker, D. H. Mathews & D. H. Turner. Algorithms and Thermodynamics for RNA Secondary Structure Prediction: A Practical Guide In RNA Biochemistry and Biotechnology, 11 - 43, J. Barciszewski and B. F. C. Clark, editors, NATO ASI Series, Kluwer Academic Publishers, Dordrecht, NL, (1999)).

[0063] The types of structures provided can use repetitive trinucleotide motifs (e.g., Figure 11B) Controlled. A repetitive trinucleotide motif is a motif of three nucleotides that is repeated at least twice in a sequence (e.g., repeated two or more times, three or more times, four or more times, five or more times, six or more times, seven or more times, eight or more times, or ten or more times). Thus, the extended sequence can contain a repetitive trinucleotide motif. In one embodiment, the extended sequence contains a repetitive trinucleotide motif of CAA, UUG, AAG, CUU, CCU, CCA, UAA, or a combination thereof, which provides a random coil sequence. In another embodiment, the extended sequence contains a repetitive trinucleotide motif of CAU, CUA, UUA, AUG, UAG, or a combination thereof, which provides a semi-stable hairpin structure. In another embodiment, the extended sequence contains a repetitive CNG trinucleotide motif (e.g., CGG, CAG, CUG, CCG), a repetitive trinucleotide motif of CGA or CGU, or a combination thereof, which provides a stable hairpin structure. In another embodiment, the extended sequence contains a repetitive trinucleotide motif of AGG, UGG, or a combination thereof, which provides a quadruplex (or G-quadruplex) structure. In yet another embodiment, the extended sequence contains a combination of the aforementioned trinucleotide motifs and a combination of the resulting different structures. For example, the extended sequence can have regions containing a random coil structure, regions containing a semi-stable hairpin, regions containing a stable hairpin, and / or regions containing a quadruplex. Thus, each region can contain a repetitive trinucleotide motif associated with the indicated structure. Non-limiting examples of structures are presented in the following table:

[0064] Table 1. Representative RNA extended sequences (35 nucleotides) and their corresponding structures.

[0065]

[0066] The extended sequences can also be used to generate crRNA multimers; thus, in another embodiment, there is provided a crRNA multimer comprising two or more crRNA molecules (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, even 8 or more crRNA molecules), wherein each crRNA comprises an extended sequence as described herein, and the crRNA molecules of the multimer are linked by their extended sequences, e.g., via base pairing or hybridization. Thus, in one embodiment, each crRNA of the multimer comprises an extended sequence that comprises a region that is sufficiently complementary to the extended region of another crRNA of the multimer to facilitate hybridization. The complementary region can have any suitable length that promotes interaction (e.g., 4 nt or more, 6 nt or more, 8 nt or more, 10 nt or more, 15 nt or more, etc.). The crRNA multimer can be used, for example, to simultaneously deliver multiple crRNAs, e.g., when multiple crRNAs are required for a particular therapeutic strategy. An example of such a use is exon skipping, in which a DNA segment is cut by two crRNAs to restore the functional reading frame (e.g., Ousterout DG et al. (2015), Multiplex CRISPR / Cas9-based genome editing for correction of dystrophin mutations that cause Duchenne Muscular Dystrophy. Nat Commun. 6:6244). Exon skipping requires two crRNAs that each target different sites in the nucleus to be targeted (one at the 5' site and the other at the 3' site). Ideally, the ratio of the two crRNAs should be 1:1; however, it is difficult to maintain this ratio. By pairing the crRNAs in the multimer (e.g., each comprising a different targeting sequence) via an appropriate extended structure, delivery at the desired ratio can be facilitated.

[0067] In one embodiment, two or more crRNAs with structured extensions participate in an RNA "kissing" interaction (also known as a loop-loop interaction), which occurs when unpaired nucleotides in one structured extension sequence (e.g., a hairpin loop) base pair with unpaired nucleobases in another structure (e.g., another hairpin loop) on a second crRNA. Examples of this type of interaction are shown in Figure 11C . The formation of kissing loops or other structures multimerizes two or more crRNA molecules. This strategy can be used to link several crRNA molecules.

[0068] Hybridization of complementary sequences in the extensions on each crRNA can also be used to promote multimerization. For example, supramolecular crRNA structures can be constructed via an extension region with self-assembly capabilities. For example, trimers can be formed from three RNA molecules with appropriately placed hybridization regions (e.g., Figure 11D , panel (i); Shu D, Shu Y, Haque F, Abdelmawla S, & Guo P (2011) Thermodynamically stable RNA three-way junctions as platform for constructing multi-functional nanoparticles for delivery of therapeutics. Nat Nanotechnol. 6(10):658-667.). Similarly, RNA octamers can be generated by assembling sixteen RNA molecules ( Figure 11D , panel (ii); Yu J, Liu Z, Jiang W, Wang G, & Mao C (2014) De novo design of an RNA tile that self-assembles into a homo-octameric nanoprism. Nat Commun. 6:5724)).

[0069] Any of the foregoing types of extension portions can be used with or without a processing sequence. In some embodiments, the nucleic acid can include multiple processing sequences and extension sequences. For example, the nucleic acid can further include a second processing sequence 5' of the first extension sequence and a second extension sequence 5' of the second processing sequence. The second processing and extension sequences can be the same as the first processing and extension sequences (e.g., repeated), or either or both of the second processing sequence and the second extension sequence can be different from the first processing sequence and / or the extension sequence. The nucleic acid is not particularly limited to any number of processing and extension sequences and can have 2, 3, 4, 5, etc. processing and / or extension sequences.

[0070] The 5' end of the nucleic acid construct (i.e., the processing sequence or extension sequence at the 5' end, where applicable) can be further modified as needed. For example, the 5' end can be modified with a functional group (e.g., a functional group involved in bioorthogonal or "click" chemical reactions). For example, the 5' end of the nucleic acid can be chemically modified with an azide, a tetrazine, an alkyne, a strained alkene, or a strained alkyne. Such modifications can use appropriately paired functional groups to facilitate the attachment of the desired chemical moiety or molecule to the construct.

[0071] The 5' end of the nucleic acid can optionally be modified via the above-described bioorthogonal or "click" chemistry to incorporate a biofunctional molecule. The biofunctional molecule can be any molecule that enhances the delivery or activity of an RNA-guided endonuclease, or provides some other desired function, such as targeting the nucleic acid to a specific destination (e.g., a portion that targets a specific protein, cell receptor, tissue, etc.), or facilitating the tracking of the construct (e.g., a detectable label such as a fluorescent label, a radioactive label, etc.). Examples of biofunctional molecules include, for example, endolysosomal polymers, donor DNA molecules, amino sugars (e.g., N-acetylgalactosamine (GalNAc) or tri-GalNAc), guide and / or tracer RNAs (e.g., single guide RNAs), and other peptides, nucleic acids, and targeting ligands (e.g., antibodies, ligands, cell receptors, aptamers, galactose, sugars, small molecules). In one embodiment, the crRNA comprises a biotin or avidin (or streptavidin) molecule conjugated to the crRNA extension, allowing, when appropriate, the modified crRNA to bind to another molecule conjugated to avidin / streptavidin or biotin (e.g., a targeting molecule or peptide) (see, for example, Figure 14 ). In another embodiment, the crRNA extension can be covalently linked to an amino sugar in any suitable manner, such as via a linker. As used herein, "amino sugar" is a sugar molecule in which a hydroxyl group has been replaced with an amine group (e.g., galactosamine) and / or a nitrogen of which it is part of a complex functional group (e.g., N-acetylgalactosamine (GalNAc); tri-N-acetylgalactosamine (trifunctional N-acetylgalactosamine)). The amino sugar can be modified to contain an optional spacer. Examples of amino sugars include N-acetylgalactosamine (GalNAc), trivalent GalNAc, or trifunctional N-acetylgalactosamine. An example of an amino sugar group includes the following:

[0072]

[0073] Wherein said linker can be any commonly known in the art, and each linker can be the same or different from each other. Generally, the linker is a saturated or unsaturated aliphatic or heteroaliphatic chain. The aliphatic or heteroaliphatic chain typically contains 1-30 members (e.g., 1-30 carbon, nitrogen, and / or oxygen atoms), and can be substituted by one or more functional groups (e.g., one or more ketone, ether, ester, amide, alcohol, amine, urea, thiourea, sulfoxide, sulfone, sulfonamide, and / or disulfide groups). In some cases, shorter aliphatic or heteroaliphatic chains are used (e.g., about 1-15 members, about 1-10 members, about 1-5 members, about 3-15 members, about 3-10 members, about 5-15 members, or about 5-10 members in the chain). In other cases, longer aliphatic or heteroaliphatic chains are used (e.g., about 5-30 members, about 5-25 members, about 5-20 members, about 10-30 members, about 10-25 members, about 10-20 members, about 15-30 members, about 15-25 members, or about 15-20 members in the chain). Examples of spacers include substituted and unsubstituted alkyl, alkenyl, and polyethylene glycol (e.g., PEG 1-10 or PEG 1-5), or combinations thereof. More specific examples for illustration are provided below:

[0074]

[0075] Before conjugation to the linker, the amino sugar can contain a functional group (e.g., azide, tetrazine, alkyne, strained alkene, or strained alkyne) that allows conjugation to a properly paired functional group (e.g., at the 5' end) attached to the crRNA extension. Thus, for example, before conjugation to the extended crRNA, the amino sugar can contain:

[0076]

[0077] Where A 2 contains an azide, tetrazine, alkyne, strained alkene, or strained alkyne as described herein. More specific examples are as follows:

[0078]

[0079] Where A 2 contains an azide, tetrazine, alkyne, strained alkene, or strained alkyne as described herein, for example:

[0080]

[0081] Processing sequence

[0082] In some embodiments, the crRNA comprises a processing sequence. The processing sequence is a nucleic acid sequence that is self-cleaved by Cpf1 in vitro or in vivo without a guide / targeting sequence. Without wishing to be bound by any particular theory or mechanism of action, it is believed that when present, the processing sequence is cleaved after entering the cell and releases the crRNA from any extension sequence. The processing sequence can be located between the crRNA and the extension sequence. In this configuration, after cleavage of the processing sequence, the crRNA is released from the extension sequence of the nucleic acid construct provided herein.

[0083] The processing sequence can also be located within the extension sequence, at the 5' of the extension sequence, or can serve as the extension sequence. Additionally, multiple processing sequences can be used. For example, a second processing sequence can serve as the extension sequence alone or together with additional nucleotide sequences. However, the extension sequence is generally different from the processing sequence (if present). Furthermore, in one embodiment, the extension sequence does not contain a processing sequence and / or any other complete (intact) crRNA sequence.

[0084] In some embodiments, the processing sequence is located immediately 5' to the crRNA (i.e., directly attached to the crRNA sequence). In other embodiments, a spacer sequence can be present between the crRNA and the processing sequence. The spacer sequence can have any length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nt), provided that it does not prevent Cpf1 cleavage of the processing sequence or the function of the crRNA released after cleavage.

[0085] In one embodiment, the processing sequence comprises a fragment of the direct repeat sequence of the Cpf1 array. The Cpf1 array (sometimes also referred to as pre-crRNA) is a naturally occurring array that contains direct repeat sequences and spacer sequences between each direct repeat. The direct repeat portion of the array includes two parts: a crRNA sequence part and a processing part. Within a given direct repeat, the processing part is located 5' to the crRNA sequence part, often immediately 5' to the processing part. According to this embodiment, the processing sequence of the nucleic acid provided herein comprises at least one fragment of the processing part of the direct repeat sufficient to effect Cpf1 cleavage. For example, the processing sequence can comprise a fragment of at least 5 contiguous nucleotides of the processing part of the direct repeat sequence, such as at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 nt (or the entire processing part of the direct repeat sequence), the length of which depends on the species from which the direct repeat is derived. In some embodiments, the processing sequence comprises the entire processing part of the direct repeat sequence. The direct repeat can be from the Cpf1 array of any microorganism. In Figure 9 direct repeat sequences, as well as examples of the processing parts of the direct repeat sequences, are provided. The processing sequence of the nucleic acid of the present invention can comprise Figure 9a fragment or the entire sequence of any processing sequence (e.g., SEQ ID NO: 2-20).

[0086] donor nucleic acid

[0087] The nucleic acid constructs provided herein may further comprise a donor nucleic acid (also referred to as a donor polynucleotide). A donor polynucleotide is a nucleic acid that is inserted at a cleavage site induced by an RNA-guided endonuclease (e.g., Cpf1). The nucleic acid of the donor polynucleotide can be any type of nucleic acid known in the art. For example, the nucleic acid can be DNA, RNA, a DNA / RNA hybrid, an artificial nucleic acid, or any combination thereof. In one embodiment, the nucleic acid of the donor polynucleotide is DNA, also referred to herein as "donor DNA".

[0088] The donor polynucleotide is typically single-stranded and serves as a template for generating double-stranded DNA containing the desired sequence. The donor polynucleotide has sufficient identity (e.g., 85%, 90%, 95%, or 100% sequence identity) with the genomic sequence flanking the cleavage site to a region of the genomic sequence proximal to the cleavage site (within about 50 bases or less, within about 30 bases or less, within about 15 bases or less, or within about 10 bases or less, within about 5 bases or less, or immediately adjacent to the cleavage site) to support homology-directed repair between the donor sequence and the genomic sequence flanking the cleavage site (with which the donor sequence has sufficient sequence identity). The donor polynucleotide sequence can be of any length, but must have a sufficient number of nucleotides with sequence identity on either side of the cleavage site to facilitate HDR. These regions of the donor polynucleotide are referred to as homology arms. The homology arms can have the same number of bases or different numbers of bases, and each is generally at least 5 nucleotides in length (e.g., 10 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 150 nucleotides or more, or even 200 nucleotides or more). The donor polynucleotide also contains a central region flanked by homology arms that contains the mutation or other DNA sequence of interest. Thus, the overall length of the donor polynucleotide is typically greater than the combined length of the two homology arms (e.g., about 15 nucleotides or more, about 20 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 150 nucleotides or more, or even 200 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more).

[0089] Donor polynucleotide sequences are typically different from the target genomic sequences. In contrast, donor polynucleotide sequences can contain one or more single-base alterations, insertions, deletions, inversions, or rearrangements with respect to the genomic sequence, provided that the homologous arms have sufficient sequence identity to support HDR. The donor polynucleotide sequence can further contain sequences that facilitate the detection of successful insertion of the donor polynucleotide.

[0090] The ends of the donor polynucleotide can be protected by methods known to those skilled in the art (e.g., from exonuclease degradation). For example, one or more dideoxynucleotide residues are added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides are ligated to one or both ends. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, adding terminal amines and using modified internucleotide linkages such as phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues.

[0091] In some embodiments, the donor polynucleotide (e.g., donor DNA) is covalently linked to the 5' end of the Cpf1 crRNA, the 5' end of the processing sequence, or the 5' end of the extension sequence. In a preferred embodiment, the donor polynucleotide is linked to the 5' end of the extension sequence. In some embodiments, the bond between the donor DNA and the nucleic acid is reversible (e.g., a disulfide bond).

[0092] In some embodiments, the donor DNA is covalently linked to the nucleic acid construct. For example, the donor polynucleotide can be linked to the processing sequence and serve as an extension sequence located 5' of the processing sequence. In another embodiment, the donor polynucleotide can be linked at the 5' of the extension sequence.

[0093] The nucleic acid and the donor DNA can be linked or conjugated by any method known in the art. In some embodiments, the 3' end of the donor DNA and the 5' end of the nucleic acid are modified to facilitate bonding. For example, the 5' end of the nucleic acid can be activated with thiopyridine, while the donor DNA can be thiol-capped, allowing the formation of a disulfide bond between the two molecules. In some embodiments, a bridge DNA complementary to both the nucleic acid and the donor DNA hybridizes and brings the two molecules into proximity to facilitate the reaction. Figures 5A - 5C Non-limiting examples of the synthesis of the conjugation of donor DNA to nucleic acids are provided.

[0094] In other embodiments, the nucleic acid and the donor DNA can be conjugated via functional groups, such as those involved in bioorthogonal or "click" chemical reactions. For example, the 5' end of the nucleic acid can be chemically modified with a functional group such as an azide, a tetrazine, an alkyne, a strained alkene, or a strained alkyne, and the 3' end of the donor DNA can be chemically modified with a suitably paired functional group. For example, if the nucleic acid contains an azide, the azide will react with the alkyne group of the donor DNA via an azide-alkyne cycloaddition reaction (copper-catalyzed), or will react with the strained alkyne group of the donor DNA via an azide-strained alkyne cycloaddition reaction (catalyst-free). Similarly, if the nucleic acid contains a tetrazine, it will react with a strained alkene via a tetrazine / alkene cycloaddition reaction. Similarly, the reverse configuration can be used, e.g., if the nucleic acid contains an alkyne, a strained alkyne, or a strained alkene, it will react with the azide or tetrazine group of the donor DNA via the same cycloaddition reaction.

[0095] In some embodiments, the nucleic acid and the donor DNA are conjugated via a linker. For example, the nucleic acid and the donor DNA can be conjugated via a self-immolative linker. As used herein, a "self-immolative linker" is a linker that hydrolyzes under specific conditions (e.g., a specific pH value), which allows the donor DNA to be released from the nucleic acid.

[0096] Linkers for nucleic acid-donor DNA conjugates encompass any linker known in the art that is capable of covalently linking the donor DNA to the nucleic acid. The linker can be attached to the donor DNA and the nucleic acid at either end. However, in some embodiments, the linker is attached to the 5' end of the nucleic acid (e.g., the 5' end of the crRNA, the processing sequence, or the extension sequence) and the 3' end of the donor DNA. The linker can be attached to the nucleic acid and the donor DNA by any method known in the art, such as those described herein with respect to the conjugation of the donor DNA to the nucleic acid.

[0097] In another embodiment, the donor polynucleotide can hybridize to an extension sequence and / or a processing sequence. Thus, for example, the extension sequence can contain a sequence that is sufficiently complementary to the donor polynucleotide to facilitate hybridization.

[0098] When the donor nucleic acid is covalently or non-covalently linked to the extension sequence, it is sometimes desirable for the donor nucleic acid to be linked to the extension sequence or a portion thereof that is not cleaved by the Cpf1 crRNA, such that when the target gene is edited by Cpf1, the donor nucleic acid binds tightly to the crRNA. It is believed that in certain cases, improved gene editing can be achieved through such constructs. As described above, extension sequences that are not cleaved by Cpf1 include, for example, extension sequences that contain one or more modified internucleotide linkages or synthetic nucleotides.

[0099] Compositions and Vectors

[0100] The present invention also includes compositions comprising any of the nucleic acid molecules and vectors described herein. Any suitable vector for nucleic acid delivery can be used. In some embodiments, the vector can comprise a molecule capable of interacting with any of the nucleic acids described herein and facilitating entry of the nucleic acid into a cell.

[0101] In some embodiments, the vector comprises a cationic lipid. A cationic lipid is an amphiphilic molecule having a positively charged polar head group that is attached via an anchor to a non-polar hydrophobic domain generally comprising two alkyl chains. In some embodiments, the cationic lipid forms a liposome (e.g., a lipid vesicle) around the nucleic acid construct and optionally the Cpf1 protein. Thus, in a related aspect, there is provided a liposome comprising a nucleic acid construct and optionally the Cpf1 protein.

[0102] In yet another embodiment, the carrier comprises a cationic polymer. Examples of the cationic polymers of the compositions of the present invention include polyethyleneimine (PEI), poly(arginine), poly(lysine), poly(histidine), poly-[2-{(2-aminoethyl)amino}-ethyl-asparagine] (pAsp(DET)), block copolymers of polyethylene glycol (PEG) and polyarginine, block copolymers of PEG and polylysine, block copolymers of PEG and poly{N-[N-(2-aminoethyl)-2-aminoethyl] asparagine} (PEG-pAsp[DET]), ({2,2-bis[(9Z,12Z)-octadeca-9,12-dien-1-yl]-1,3-dioxolan-5-yl}methyl)dimethylamine, (3aR,5s,6aS)-N,N-dimethyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydro-3aH-cyclopenta[d][1,3]dioxol-5-amine, (3aR,5r,6aS)-N,N-dimethyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydro-3aH-cyclopenta[d][1,3]dioxol-5-amine, (3aR,5R,7aS)-N,N-dimethyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydrobenzo[d][1,3]dioxol-5-amine, (3aS,5R,7aR)-N,N-dimethyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydrobenzo[d][1,3]dioxol-5-amine, (2-{2,2-bis[(9Z,12Z)-octadeca-9,12-dien-1-yl]-1,3-dioxolan-4-yl}ethyl)dimethylamine, (3aR,6aS)-5-methyl-2-((6Z,9Z)-octadeca-6,9-dien-1-yl)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydro-3aH-[1,3]dioxolo[4,5-c]pyrrole, (3aS,7aR)-5-methyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydro-[1,3]dioxolo[4,5-c]pyridine, (3aR,8aS)-6-methyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydro-3aH-[1,3]dioxolo[4,5-d]azepine, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 2-(dimethylamino)acetate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 3-(dimethylamino)propionate, [6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-(dimethylamino)butanoate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 5-(dimethylamino)pentanoate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 6-(dimethylamino)hexanoate, (3-{2,2-bis[(9Z,12Z)-octadeca-9,12-dien-1-yl]-1,3-dioxolan-4-yl}propyl)dimethylamine, 1-((3aR,5r,6aS)-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydro-3aH-cyclopenta[d][1,3]dioxol-5-yl)-N,N-dimethylmethanamine, 1-((3aR,5s,6aS)-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydro-3aH-cyclopenta[d][1,3]dioxol-5-yl)-N,N-dimethylmethanamine, 8-methyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)-1,3-dioxa-8-azaspiro[4.5]decane, 2-(2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)-1,3-dioxolan-4-yl)-N-methyl-N-(pyridin-3-ylmethyl)ethanamine, 1,3-bis((9Z,12Z)-octadeca-9,12-dien-1-yl) 2-[2-(dimethylamino)ethyl]malonate, N,N-dimethyl-1-((3aR,5R,7aS)-2-((8Z,11Z)-octadeca-8,11-dien-1-yl)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydrobenzo[d][1,3]dioxol-5-yl)methanamine, N,N-dimethyl-1-((3aR,5S,7aS)-2-((8Z,11Z)-octadeca-8,11-dien-1-yl)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydrobenzo[d][1,3]dioxol-5-yl)methanamine, (1s,3R,4S)-N,N-dimethyl-3,4-bis((9Z,12Z)-octadeca-9,12-dien-1-yloxy)cyclopentanamine, (1s,3R,4S)-N,N-dimethyl-3,4-bis((9Z,12Z)-octadeca-9,12-dien-1-yloxy)cyclopentanamine, 2-(4,5-bis((8Z,11Z)-heptadeca-8,11-dien-1-yl)-2-methyl-1,3-dioxolan-2-yl)-N,N-dimethylethanamine, 2,3-bis((8Z,11Z)-heptadeca-8,11-dien-1-yl)-N,N-dimethyl-1,4-dioxaspiro[4.5]dec-8-amine, (6Z,9Z,28Z,(6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-(diethylamino)butanoate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-[bis(prop-2-yl)amino]butanoate, N-(4-N,N-dimethylaminobutanoyl)-(6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-amine, (2-{2,2-bis[(9Z,12Z)-octadeca-9,12-dien-1-yl]-1,3-dioxolan-5-yl}ethyl)dimethylamine, (4-{2,2-bis[(9Z,12Z)-octadeca-9,12-dien-1-yl]-1,3-dioxolan-5-yl}butyl)dimethylamine, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl (2-(dimethylamino)ethyl)carbamate, 2-(dimethylamino)ethyl (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-ylcarbamate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 3-(ethylamino)propanoate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-(prop-2-ylamino)butanoate, N1,N1,N2-trimethyl-N2-((11Z,14Z)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)eicos-11,14-dien-1-yl)ethane-1,2-diamine, 3-(dimethylamino)-N-((11Z,14Z)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)eicos-11,14-dien-1-yl)propanamide, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-(methylamino)butanoate, dimethyl({4-[(9Z,12Z)-octadeca-9,12-dien-1-yloxy]-3-{[(9Z,12Z)-octadeca-9,12-dien-1-yloxy]methyl}butyl})amine, 2,3-bis((8Z,11Z)-heptadeca-8,11-dien-1-yl)-8-methyl-1,4-dioxa-8-azaspiro[4.5]decane, 3-(dimethylamino)propyl (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-ylcarbamate, 2-(dimethylamino)ethyl ((11Z,14Z)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)eicos-11,14-dien-1-yl)carbamate, 1-((3aR,4R,6aR)-6-methoxy-2,2-bis((9Z,12Z)-octadeca-9,(12-dien-1-yl)tetrahydrofuro[3,4-d][1,3]dioxol-4-yl)-N,N-dimethylmethanamine, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-[ethyl(methyl)amino]butanoate, (6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl 4-aminobutanoate, 3-(dimethylamino)propyl ((11Z,14Z)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)eicosa-11,14-dien-1-yl)carbamate, 1-((3aR,4R,6aS)-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)tetrahydrofuro[3,4-d][1,3]dioxol-4-yl)-N,N-dimethylmethanamine, (3aR,5R,7aR)-N,N-dimethyl-2,2-bis((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydrobenzo[d][1,3]dioxol-5-amine, (11Z,14Z)-N,N-dimethyl-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)eicosa-11,14-dien-1-amine, (3aS,4S,5R,7R,7aR)-N,N-dimethyl-2-((7Z,10Z)-octadeca-7,10-dien-1-yl)-2-((9Z,12Z)-octadeca-9,12-dien-1-yl)hexahydro-4,7-methanobenzo[d][1,3]dioxol-5-amine, N,N-dimethyl-3,4-bis((9Z,12Z)-octadeca-9,12-dien-1-yloxy)butan-1-amine, and 3-(4,5-bis((8Z,11Z)-heptadeca-8,11-dien-1-yl)-1,3-dioxolan-2-yl)-N,N-dimethylpropan-1-amine. Any combination of the foregoing polymers can also be used.,

[0103] In other embodiments, the carrier comprises polymeric nanoparticles. For example, the compositions of the invention can be administered as nanoparticles as described in International Patent Application No. PCT / US2016 / 052690, the entire disclosure of which is expressly incorporated by reference.,

[0104] Cpf1 polypeptide or nucleic acid encoding the same

[0105] In some embodiments, including the liposome embodiments described above, the composition further comprises a Cpf1 polypeptide or a nucleic acid encoding the same. Any Cpf1 polypeptide can be used in the compositions of the invention, although the Cpf1 selected should be appropriately chosen so as to act in combination with the crRNA of the nucleic acid construct in the composition to cleave the target nucleic acid and / or the processing sequence of the nucleic acid construct when applicable. The Cpf1 of the composition can be a naturally occurring Cpf1 or a variant or mutant Cpf1 polypeptide. In some embodiments, the Cpf1 polypeptide is enzymatically active; for example, when bound to a guide RNA, the Cpf1 polypeptide cleaves the target nucleic acid. In some embodiments, relative to a wild-type Cpf1 polypeptide (e.g., relative to a Cpf1 polypeptide comprising the amino acid sequence (SEQ ID NO: 1) shown in Figure 8 ), the Cpf1 polypeptide exhibits reduced enzymatic activity and retains DNA-binding activity. Mutations that alter the enzymatic activity of Cpf1 are known in the art.

[0106] For example, Cpf1 can be from a bacterium of the genus Acidaminococcus or from the family Lachnospiraceae, or from any genus or species identified in Figure 9 . Figure 8 Examples of Cpf1 protein sequences are provided in Figure 8 . In some embodiments, the amino acid sequence comprised by the Cpf1 polypeptide has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity to the amino acid sequence shown in Figure 8 . In some embodiments, the amino acid sequence comprised by the Cpf1 polypeptide has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity to an adjacent segment of 100 amino acids to 200 amino acids (aa), 200 aa to 400 aa, 400 aa to 600 aa, 600 aa to 800 aa, 800 aa to 1000 aa, 1000 aa to 1100 aa, 1100 aa to 1200 aa, or 1200 aa to 1300 aa of the amino acid sequence shown in

[0107] In some embodiments, the amino acid sequence comprised by the Cpf1 polypeptide has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity to Figure 8The RuvCI domain of the Cpf1 polypeptide of the amino acid sequence shown has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity. In some embodiments, the Cpf1 polypeptide comprises an amino acid sequence that is Figure 8 The RuvCII domain of the Cpf1 polypeptide of the amino acid sequence shown has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity. In some embodiments, the Cpf1 polypeptide comprises an amino acid sequence that is Figure 8 The RuvCIII domain of the Cpf1 polypeptide of the amino acid sequence shown has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity.

[0108] In some embodiments, the Cpf1 polypeptide is FnCpf1, Lb3Cpf1, BpCpf1, PeCpf1, SsCpf1, AsCpf1, Lb2Cpf1, CMtCpf1, EeCpf1, MbCpf1, LiCpf1, LbCpf1, PcCpf1, PdCpf1 or PmCPf1; or a Cpf1 polypeptide that comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90% or 100% amino acid sequence identity thereto.

[0109] In some embodiments, the Cpfl polypeptide comprises an amino acid substitution (e.g., D→A substitution) at the amino acid residue corresponding to position 917 of the amino acid sequence shown in Figure 8 ; and / or comprises an amino acid substitution (e.g., E→A substitution) at the amino acid residue corresponding to position 1006 of the amino acid sequence shown in Figure 8 ; and / or comprises an amino acid substitution (e.g., D→A substitution) at the amino acid residue corresponding to position 1255 of the amino acid sequence shown in Figure 8 .

[0110] The Cpf1 polypeptide can also be RNAse-inactivated Cpf1, such as Cpf1 containing modifications at H800A, K809A, K860A, F864A or R790A of Acidaminococcus Cpf1 (AsCpf1), or at corresponding positions of different Cpf1 orthologs. Examples of mutant Cpf1 proteins include those proteins disclosed in Zetsche et al., “Multiplex Gene Editing by CRISPR-Cpf1 Through Autonomous Processing of a Single crRNA Array,” Nat. Biotechnol. 2017, 35(1), 31-34. The Cpf1 polypeptide can also be a dCpf1 base editor (such as a Cpf1-cytosine deaminase fusion protein). Examples include, for example, those proteins disclosed in Li et al., Nature Biotechnology, 36324-327 (2018), and Mahfouz et al., Biochem J., 475(11), 1955-1964 (2018). An example of a synthetic variant Cpf1 is the MAD7 Cpf1 ortholog by Inscripta, Inc. (CO, USA). Additional examples of Cpf1 proteins include any of those Cpf1 proteins disclosed in International Patent Application No. PCT / US2016 / 052690, including chimeric or mutant proteins, the entire disclosure of which patent application is expressly incorporated herein by reference.

[0111] Other nucleic acids

[0112] In addition to the crRNA, the composition can further comprise other nucleic acids. For example, as described herein, the composition can comprise a donor polynucleotide. Alternatively or additionally, the composition can comprise one or more additional nucleic acids that are not donor polynucleotides (e.g., nucleic acids that do not have significant sequence identity to the target sequence to be edited, or any endogenous nucleic acid sequence of the edited cell (e.g., a level of sequence identity that is not sufficient to permit homologous recombination)). These additional nucleic acids can be RNA or DNA, such as single-stranded RNA or DNA molecules (or hybrid molecules comprising both RNA and DNA, optionally having synthetic nucleic acid residues). The additional nucleic acids can be of any length, such as at least 5 nucleotides in length (e.g., 10 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 150 nucleotides or more, or even 200 nucleotides or more). In some embodiments, the nucleic acid can comprise 500 nucleotides or more, 1000 nucleotides or more, or even 5000 nucleotides or more). However, in most cases, the nucleic acid comprises about 5000 nucleotides or less, such as about 1000 nucleotides or less, or even 500 nucleotides or less (e.g., 200 nucleotides or less).

[0113] The composition can further comprise a nucleic acid encoding a protein for a particular purpose, such as, for example, an RNA-guided endonuclease (e.g., a Cpf1 polypeptide). The RNA-guided endonuclease can be any as described herein with respect to other aspects of the invention.

[0114] divalent metal ion

[0115] In some embodiments, the composition is substantially or completely free of divalent metal ions (e.g., magnesium) that activate the particular Cpf1 protein used, in order to reduce or prevent premature cleavage of the processing sequence prior to delivery. The composition is considered to be substantially free of magnesium at a concentration that does not permit Cpf1 self-processing enzymatic activity. In some embodiments, the composition comprises about 20 mM or less of NaCl and is substantially or completely free of magnesium or other divalent ions that activate the Cpf1 protein.

[0116] Method for genetically modifying a eukaryotic cell

[0117] The present invention also provides a method for genetically modifying a eukaryotic target cell, which comprises contacting the eukaryotic target cell with any nucleic acid or composition described herein (e.g., comprising a Cpf1 crRNA, an extension sequence of the 5' of the crRNA, and optionally a processing sequence between the crRNA and the extension sequence) to genetically modify the target nucleic acid. In some embodiments, the Cpf1 crRNA of the nucleic acid comprises a targeting sequence that hybridizes to a target sequence in the target cell (e.g., the 3' of the stem-loop domain). In some embodiments, the Cpf1 crRNA comprises a processing sequence that is cleaved after entering the cell, thereby releasing the Cpf1 crRNA from the processing sequence and the extension sequence. In other embodiments, the Cpf1 crRNA does not comprise a processing sequence.

[0118] The target nucleic acid is a polynucleotide (e.g., RNA, DNA) to which the targeting sequence of the crRNA binds and induces cleavage by Cpf1. The target nucleic acid comprises a "target site" or "target sequence", which is a sequence present in the target nucleic acid to which the crRNA hybridizes, which in turn guides the endonuclease to the target nucleic acid.

[0119] "Eukaryotic target cell" can be any eukaryotic cell known in the art and includes both in vivo and in vitro cells. In one embodiment, the target cell is a mammalian cell.

[0120] Any route of administration can be used to deliver the composition to a mammal. Indeed, although more than one route can be used to administer the composition, a particular route may provide a more immediate and more effective response than another route. When administered in vitro or ex vivo to cells, the nucleic acid or composition can be contacted with the cells by any suitable method. For example, the nucleic acid can be introduced in liposomes, encapsulated by cationic polymers, and / or by electroporation. When administered to a subject such as a mammal or a human, the composition can be administered by any of a variety of routes. For example, a dose of the composition can also be applied or infused into a body cavity, absorbed through the skin (e.g., via a transdermal patch), inhaled, ingested, topically applied to tissue, or parenterally administered, for example, via intravenous, intraperitoneal, intraoral, intradermal, subcutaneous, or intraarterial administration.

[0121] The composition can be administered in or on a device that permits controlled or sustained release, such as a sponge, a biocompatible mesh, a mechanical reservoir, or a mechanical implant. Implants (see, e.g., U.S. Patent 5,443,505), devices (see, e.g., U.S. Patent 4,863,457), such as implantable devices, such as mechanical reservoirs or implants, or devices composed of a polymeric composition, are particularly useful for the administration of the composition. The composition can also be administered in the form of a sustained release formulation (see, e.g., U.S. Patent 5,378,475), which includes, for example, a gel foam, hyaluronic acid, gelatin, chondroitin sulfate, polyphosphates such as bis-2-hydroxyethyl terephthalate (BHET) and / or polylactic glycolic acid.

[0122] Examples

[0123] The following examples further illustrate the invention, but, of course, should not be construed in any way as limiting its scope. Table 2 below provides the nucleic acid sequences used in these experiments.

[0124] Table 2

[0125] Supplementary Table 1

[0126]

[0127]

[0128]

[0129]

[0130] *: phosphorothioate

[0131] U: 2'-deoxy

[0132] Underline : pre-crRNA sequence

[0133] Example 1

[0134] This example illustrates that unmodified Cpf1 crRNA tends to provide lower gene editing efficiency compared to Cas9 sgRNA.

[0135] Using a green fluorescent protein (GFP) reporter system, SpCas9 and AsCpf1 were compared, both without the extended portion of the guide RNA. A matching protospacer DNA sequence in the GFP gene that could be recognized by both nucleases was selected for a direct comparison of AsCpf1 and SpCas9. This system was in Figure 13AShown in. Using electroporation and cationic lipids, the RNP complex was introduced into HEK293T cells expressing the GFP gene under the control of a doxycycline-inducible promoter (GFP-HEK). Editing activity was determined by measuring the population of GFP-negative cells, where GFP was disrupted by NHEJ-mediated indel mutations.

[0136] AsCpf1 RNP showed lower gene editing than SpCas9 in both electroporated and cationic liposome-treated cells ( Figure 13B and 13C ). 31% of the cells electroporated with AsCpf1 RNP were GFP-negative, while 41% of the cells electroporated with SpCas9 RNP were GFP-negative ( Figure 13B ). Using the cationic liposome and RNAiMax delivery systems, approximately 8% of the cells treated with AsCpf1 were GFP-negative, while approximately 30% of the cells treated with SpCas9 were GFP-negative ( Figure 13C ).

[0137] Example 2

[0138] This example demonstrated that a nucleic acid comprising a Cpf1 crRNA, a processing sequence 5' of the Cpf1 crRNA, and an extended sequence 5' of the processing sequence (reversibly supercharged crRNA) enhanced Cpf1 delivery via cationic lipids in vitro.

[0139] Cationic materials such as cationic liposomes and polycations are the most commonly used delivery vehicles for nucleic acids in cells and live animals. Therefore, the efficiency of cationic lipid cationic liposomes to transfect the Cpf1 / crRNA complex into HEK cells expressing green fluorescent protein (GFP-HEK) was analyzed. Briefly, a crRNA designed to knock down the GFP gene via indels was complexed with Cpf1 and either electroporated (nucleofected) into the cells or transfected with cationic liposomes. Then, the gene editing efficiency was determined via flow cytometry by measuring the number of GFP knockout cells, where cells that no longer expressed GFP indicated that the cell had been transfected with the Cpf1 / crRNA complex.

[0140] The results from these experiments showed that cationic liposomes could not efficiently transfect unmodified Cpf1 crRNA complexes ( Figure 1A ). Specifically, cells treated with cationic liposomes and the Cpf1 / crRNA complex had only 8% NHEJ efficiency, while cells electroporated with the Cpf1 / crRNA complex had 40% NHEJ efficiency, confirming that delivery limitation was the main cause of the low NHEJ efficiency regarding the use of cationic liposomes.

[0141] To examine the effect of sequence extension, crRNAs of 41 nucleotides in length were extended by 9 nucleotides (crRNA + 9), or by 59 nucleotides (crRNA +59 ). The extended portions included self-processing sequences that self-cleaved into active crRNAs. Unmodified crRNA, crRNA + 9, and crRNA +59 were each individually complexed with Cpf1 and transfected into GFP-HEK cells using cationic liposomes. RNPs were formed under low-salt conditions to prevent potential processing of the 5’-end extensions. The NHEJ level, which is an indicator of transfection efficiency, was determined by measuring the percentage of GFP-negative cells.

[0142] As Figure 1B shown, extending the crRNA with self-processing sequences significantly enhanced the transfection efficiency using cationic mediators. Cells treated with unextended crRNA via cationic liposomes showed 8% GFP-negative, while cells treated with 9-base extended crRNA were 18% GFP-, and cells treated with 59-base extended crRNA were 37%, which was an approximately 4-fold increase over the control unmodified crRNA. Additionally, crRNA+9, crRNA+15, and crRNA+25 were tested with cationic liposomes and showed a length-dependent increase in gene editing efficiency ( Figure 1C ).

[0143] To determine whether there was a specific 5'-extension sequence requirement for this enhancement, three different 59-nucleotide 5'-extensions were compared. The first and original 59-nucleotide extended crRNA was described above and contained one AsCpf1 pre-crRNA (crRNA+59), the second 59-base extended cRNA contained four tandem AsCpf1 pre-crRNA sites (crRNA+59-D2), and the third 59-base extended crRNA contained an FnCpf1 pre-crRNA followed by a scrambled DNA sequence with no homology to any sequence in the human genome (crRNA+59-D3) ( Figure 1D ). These crRNAs were delivered using Lipofectamine 2000, similar to the above paragraph. All three 5'-extensions showed equivalent editing activity: 32% GFP-negative for crRNA+59 cells, 30% GFP-negative for crRNA+59_D2 cells, and 27% GFP-negative for crRNA+59_D3 cells ( Figure 1E)。This suggests that there are no strict sequence requirements for 5' extension enhancement, similar to previous findings with crRNAs extended by 9 nucleotides in electroporated cells. Additionally, these results provide evidence in support of the hypothesis above that increasing the negative charge density on the crRNA and thus also on the AsCpf1 RNP complex can enhance the delivery of AsCpf1 to cells via cationic lipids.

[0144] Extended crRNAs were tested in an in vitro DNA cleavage assay to determine whether the extended crRNAs enhanced the intrinsic nuclease activity of Cpf1. No differences in activity were observed among the three crRNAs tested: wild-type crRNA, crRNA +9 and crRNA +59 with incubation times of 15 minutes and 60 minutes. When the incubation time was only 5 minutes, crRNA +59 even had slower DNA cleavage than wild-type crRNA.

[0145] 5'-extended crRNAs were also studied to determine whether there was enhanced gene editing activity in Cpf1 if delivered via plasmid rather than as an RNP. The Cpf1 plasmid was transfected 24 hours prior to electroporation of the crRNA, and the gene editing activity was determined. No improvement in gene editing efficiency was observed when Cpf1 was produced from the plasmid.

[0146] Further, crRNAs were labeled with a fluorescent dye to determine whether the extended crRNAs had enhanced uptake in cells after delivery via electroporation or cationic liposomes. Electroporation of Cpf1 RNP resulted in more than 90% of the cells being positive for dye-crRNA and showed highly efficient delivery regardless of the length of the crRNA. The delivery efficiency of Cpf1 RNP with cationic liposomes depended on the length of the crRNA, and extended crRNAs were delivered more efficiently into HEK 293T cells than wild-type crRNA ( Figure 1F ).

[0147] The results show that the extended crRNAs provided herein enhance delivery and gene editing efficiency.

[0148] Example 3

[0149] This example confirmed that the activity of Cpf1 was enhanced by 5'-terminal extension in HEK cells using electroporation.

[0150] GFP-targeted crRNAs with 5’-end extensions of 4, 9, 15, 25, and 59 nucleotides were introduced into GFP-HEK cells as RNP complexes with AsCpf1 by electroporation. The sequences with 4 to 25 nucleotide extensions were scrambled, and the 59 nucleotide extension consisted of the AsCpf1 pre-crRNA followed by a scrambled RNA sequence with no homology to the human genomic sequence.

[0151] All crRNAs with 4 to 25 nucleotide 5'-extensions showed a sharp increase in gene editing over crRNAs without extensions. Cells electroporated with unextended cRNA were 30% GFP-negative (crRNA), those electroporated with crRNAs with 4 to 25 nucleotide extensions were 55 to 60% GFP-negative, and those electroporated with the 59 nucleotide extension crRNA were 37% GFP-negative ( Figure 2 ). The gene editing levels of crRNAs with 4 to 25 nucleotide 5'-extensions were comparable to those of SpCas9 RNP-electroporated cells.

[0152] The results confirmed that 5’-extensions of Cpf1 crRNAs increased gene editing efficiency.

[0153] Example 4

[0154] This example demonstrated that nucleic acids containing Cpf1 crRNA and 5'-extension sequences enhanced Cpf1 delivery in vivo by cationic lipids and Cpf1 activity in cells.

[0155] Three different chemical modifications were investigated for extended crRNAs: 2'-O-methyl modification, phosphorothioate bonding, and deoxynucleotide ribose groups ( Figure 15A ). The first 3 of 4 nucleotides were extended with: 2'-O-methyl nucleotides and 3'-phosphorothioate bonding (MS), deoxynucleotide at the 9th position of the 5'-extended crRNA of 9 nucleotides (9dU), 3'-phosphorothioate bonding for all 9 nucleotides plus deoxynucleotide at the 9th position of the 5'-extended crRNA of 9 nucleotides (9s).

[0156] Cpf1 RNPs containing extended and chemically modified crRNAs were electroporated into GFP-HEK cells, and gene editing activity was determined by flow cytometry. Extended crRNAs with chemical modifications had similar activity to unmodified extended crRNAs (41% to 46% GFP-negative cells) ( Figure 15D ).

[0157] In addition, these crRNAs were examined using HEK293T cell lines expressing blue fluorescent protein (BFP) (BFP-HEK). The results are presented in Figure 12 In

[0158] Results from this experiment showed that 5' extension increased the gene editing efficiency of AsCpf1, as well as the tolerance of the 5'-end of crRNA to chemical modification. Further, if the 5'-end of the crRNA was extended, 5'-chemical modification of the crRNA was possible without compromising activity.

[0159] A key benefit of using chemically modified crRNAs is that they are more stable to hydrolysis by serum nucleases. Thus, the serum stability of 5'-chemically modified crRNAs was investigated.

[0160] 5'-chemically modified crRNAs were incubated in diluted fetal bovine serum and their degradation was analyzed by gel electrophoresis. Figure 15B Quantification of the remaining crRNAs after a 15-minute incubation in serum was provided. Results showed that unmodified crRNAs degraded rapidly in serum, while crRNAs containing phosphorothioate backbones +9S were significantly more stable to hydrolysis in serum.

[0161] 5'-modified crRNAs were also investigated to determine whether they could enhance the ability of cationic liposomes to transfect Cpf1 RNPs due to their ability to protect crRNAs from nucleases in cells and serum. Cpf1 with crRNA +9S was 40% more efficient in editing genes in cells than crRNA +9 , which suggests that 5'-crRNA chemical modification achieved by 5'-crRNA extension has numerous applications in gene editing ( Figure 15C ).

[0162] The crystal structure of Cpf1 RNP has recently been resolved and it was confirmed that the AsCpf1 protein forms numerous interactions with the phosphodiester backbone of crRNA. Thus, 5'-chemical modification of unextended crRNAs has a high probability of disrupting important interactions between crRNA and Cpf1, resulting in disruption of AsCpf1 gene editing activity.

[0163] In contrast, crRNAs with 5' extensions appear to be tolerant of chemical modifications, as the nucleotides interacting with the AsCpf1 protein are unmodified. These results provide methods for introducing chemical modifications at the 5' end of crRNAs, which can potentially enhance delivery for ex vivo and in vivo therapeutic applications. Such constructs also allow other molecules, such as targeting ligands, endosomal escape moieties, or other functional molecules, to be conjugated to the extended crRNA and retained together with the Cpf1 molecule. For example, biotin or avidin (or streptavidin) can be conjugated to the crRNA extension, allowing the modified crRNA to bind to another molecule conjugated to streptavidin or biotin (e.g., a targeting molecule) when appropriate (e.g., Figure 14 ). Additionally, the crRNA can be conjugated to the Cpf1 mRNA via the extension, allowing Cpf1 to be delivered by translation of the mRNA. Once the Cpf1 mRNA is translated in the cytoplasm to produce the Cpf1 protein, the Cpf1 protein recognizes the crRNA portion of the construct and optionally processes the linker RNA sequence to separate the Cpf1 mRNA and the crRNA.

[0164] Example 5

[0165] Experiments were also performed to determine whether crRNAs with extended sequences could enhance the ability of cationic polymers to transfect Cpf1. Specifically, crRNAs of various lengths (crRNAs without extension, with 9 nt extensions, and with 59 nt extensions) as used in Example 2 were complexed with Cpf1, mixed with the cationic polymer PAsp(DET), and added to GFP-HEK cells. The NHEJ efficiency of the formulations was determined via flow cytometry by measuring the frequency of GFP-negative cells.

[0166] As Figure 16 shown, the 59-base 5' extension enhanced the delivery of the AsCpf1 RNP to cells mediated by PAsp(DET) by 2-fold. The unextended crRNA (crRNA) was 8% GFP-negative, the crRNA with a 9-nucleotide extension (crRNA+9) was 10% GFP-negative, while the crRNA with a 59-nucleotide extension (crRNA+59) was 18% GFP-negative. These results confirm that extended sequences can improve the delivery of crRNAs to cells using cationic polymers.

[0167] Example 6

[0168] This example confirmed that extended crRNAs enhanced Cpf1 delivery in vivo.

[0169] Experiments were performed to determine the extended crRNA (crRNA +59)Whether cationic lipid-mediated delivery of Cpf1 in vivo can be enhanced. The schematic diagram of the experiment is provided in Figure 3A as follows.

[0170] Studies were performed in Ai9 mice using a previously validated spacer. Ai9 mice were given intramuscular injections of cationic lipids or PAsp(DET) in combination with either the AsCpf1-crRNA complex or the AsCpf1-crRNA +59 complex. Two weeks after injection, the expression of tdTomato (red fluorescence) was imaged in 10-μm sections of the gastrocnemius muscle (muscle map in Figure 3B ). Comparison of the images collected for unextended and extended crRNAs showed that extended crRNAs dramatically enhanced the ability of PAsp(DET) to deliver AsCpf1 RNPs in vivo. Additionally, RNPs with extended crRNAs complexed with PAsp(DET) (Cpf1 RNP+59) induced tdTomato expression several millimeters away from the injection site and throughout the gastrocnemius muscle. The high extent of tdTomato expression in muscle is likely due to the unique multinucleated nature of muscle fibers. This allows TdTomato to be expressed along the entire length of the muscle fiber and can thus be observed over a length of several millimeters.

[0171] The ability of 59-nucleotide extended crRNAs to enhance delivery and, by extension, the editing efficiency of AsCpf1 RNPs in vivo supports the value of Cpf1 as a tool for animal studies and as a potential therapeutic agent for treating human diseases, particularly inherited muscular dystrophies.

[0172] Example 7

[0173] 5.1: Extended crRNAs increase HDR and NHEJ rates

[0174] To examine whether 5′-extension can increase HDR rates in addition to NHEJ levels, AsCpf1 RNPs with crRNAs containing various extensions were introduced into GFP-HEK cells together with single-stranded oligonucleotide donors (ssODNs). NHEJ levels were determined by measuring the population of GFP-negative cells (similar to the first section), while HDR rates were quantified using restriction enzyme digestion assays. For both 4- and 9-nucleotide extended crRNAs, a 2-fold improvement in HDR was observed (in Figure 4A , a 17% HDR frequency for crRNA+4 and an 18% HDR frequency for crRNA+9, relative to 9% for unmodified crRNA). For the 59-base extension, a smaller increase in HDR was observed (in Figure 4A , a 13% HDR rate for crRNA+59).

[0175] Interestingly, ssODNs for HDR also dramatically increased the NHEJ efficiency of AsCpf1. The addition of ssODNs increased the percentage of GFP-negative cells from 30% to 46% for unextended crRNA (crRNA), from 55% to 95% for 4-base extended crRNA (crRNA+4), from 58% to 93% for 9-base extended crRNA (crRNA+9), and from 37% to 58% for 59-base extended crRNA (crRNA+59) ( Figure 4B ). The finding that exogenously added DNA enhanced AsCpf1 RNP-mediated editing was further verified in the BFP reporter system. Single-stranded DNA (ssDNA) with no homology to the human genome was electroporated into BFP-HEK cells with AsCpf1 RNP. Similarly, the addition of ssDNA increased AsCpf1 editing activity 2-fold for both unextended and extended crRNAs. The BFP-negative population increased from 31% to 50% for unextended crRNA (crRNA), from 59% to 91% for 4-nucleotide extended crRNA (crRNA+4), and from 60% to 95% for 9-nucleotide extended crRNA (crRNA+9) ( Figure 4C ). Additional experiments were also performed to determine whether exogenously added DNA had to be homologous to the Cpf1 RNP target site in order to enhance gene editing. AsCpf1 RNP was electroporated into cells together with single-stranded DNA (ssDNA) with no homology to the target sequence, and gene editing efficiency was measured. Similarly, the addition of ssDNA with no homology also increased AsCpf1 editing activity to approximately 90% for both extended crRNAs ( Figure 4D ). The results confirmed that ssDNA could enhance editing with AsCpf1. Additionally, the activity enhancement with 5'-end extension synergized with exogenously added ssDNA, and together they induced gene editing approaching 100%. It was observed that if cationic liposomes were used as the delivery method instead of electroporation, the addition of ssDNA did not enhance Cpf1 gene editing efficiency.

[0176] Sometimes it is preferred to use RNA rather than DNA for gene editing because ssRNA cannot integrate into the genome and can be used more safely. To test the effect of ssRNA rather than DNA, GFP-HEK cells were electroporated with Cpf1 RNP and two different ssRNAs (9nt and 100nt), and the resulting gene editing levels were determined. Two 100nt ssRNAs with minor sequence variations both dramatically increased the gene editing efficiency of Cpf1, resulting in a 2-fold improvement, while 9nt ssRNA induced a 10% enhancement in gene editing efficiency ( Figure 4E ).

[0177] These results confirm that single-stranded nucleic acids can be used to enhance Cpf1 editing activity in cells.

[0178] 5.2: Extended crRNA and donor DNA for single molecule.

[0179] Part 2 of this example confirmed that extended crRNA combined with donor DNA in a single molecule can enhance HDR.

[0180] The vast majority of genetic diseases require gene correction rather than knockout, and thus there is great interest in developing Cpf1-based therapeutic agents that can correct gene mutations via HDR. To address the problem of generating nanoparticles that effectively encapsulate both donor DNA and Cpf1-crRNA, experiments were performed to determine whether crRNA and donor DNA can be combined into a single molecule via reversible disulfide bonds. As previously reported, one challenge is that the 5' end of crRNA is quite sensitive to chemical modification. Therefore, chemical modifications at the 5’ end of the self-processing sequence of extended crRNA were tested.

[0181] crRNA was extended with four additional nucleotides at its 5' end, and chemical modifications (i.e., 5'DBCO, 5'thiol, or 5'azide) were added to the ends of the nucleotides. The activities of 4nt-extended crRNAs with and without chemical modifications at the 5' end were tested. Briefly, 5'-modified supercharged crRNAs (designed to knock down the GFP gene via indel formation) were complexed with Cpf1 and electroporated (nucleofected) into cells. Then, gene editing efficiency was determined by flow cytometry by measuring the number of GFP knockout cells, where cells that no longer expressed GFP indicated that the cell had been transfected by the Cpf1 / crRNA complex.

[0182] Figure 4F It was shown that neither chemical modification nor nucleotide extension affected Cpf1-crRNA activity, as all crRNA complexes showed approximately 40% GFP knockout, which is a level similar to that of unmodified control crRNA.

[0183] Since a chemistry can be added to the 5'-end of crRNA without loss of activity, the 5'-end of crRNA was activated with thiopyridine to react with a thiol-terminated donor DNA (see Figures 5A - 5C ). The reaction between the two macromolecules has slow kinetics. Thus, a method using a bridge DNA complementary to both crRNA and donor DNA was used. The bridge hybridizes and brings the two macromolecules closer to facilitate the reaction in order to enhance the conjugation yield between crRNA and donor DNA. The conjugation yield was up to 40%, and the product was purified via gel extraction. The conjugate was named "homologous DNA-crRNA" (HD-RNA). HD-RNA contains a disulfide bond, which should be reduced in the cytoplasm. Thiol-mediated cleavage of HD-RNA was determined by incubating it in DTT for 6 hours and analyzing its molecular weight via gel electrophoresis. Figure 5C Gel comparison in

[0184] showed that DTT reduced HD-RNA and regenerated donor DNA and crRNA.

[0185] Experiments were performed to determine whether HD-RNA enhanced the HDR efficiency of Cpf1 after transfection with a cationic lipid (i.e., cationic liposome) or a cationic polymer (i.e., PAsp(DET)). For these experiments, HD-RNA complexed with Cpf1 was transfected into GFP-HEK cells using a cationic liposome or PAsp(DET), and the levels of HDR and NHEJ were compared. NHEJ was determined by measuring the frequency of GFP-negative cells and confirmed by performing a Surveyor assay with PCR amplicons of the targeted region of the BFP gene. HDR efficiency was determined by isolating cellular DNA and analyzing the presence of a restriction enzyme site embedded in the donor DNA. The results are presented in Figure 6 and 7 .

[0186] Figure 6 and 7It was demonstrated that HD-RNA enhanced both NHEJ and HDR efficiencies of Cpf1 after delivery with PAsp(DET). Specifically, HDR was detected in up to 60% of the cells treated with the HD-RNA / Cpf1 complex delivered with PAsp(DET), which was significantly higher than the HDR rate of the cells treated with Cpf1 / crRNA and donor DNA complexed with PAsp(DET). Additionally, the 60% HDR rate observed with the HD-RNA / Cpf1 complex delivered with PAsp(DET) was even higher than the HDR rate observed with electroporation of the Cpf1 / crRNA complex, and suggested that having donor DNA near the Cpf1 cleavage site could contribute to HDR.

[0187] These results confirmed that the extended crRNA could enhance HDR.

[0188] Example 9

[0189] This example confirmed the use of the extended crRNA in different cell types and the utility of the method for treating genetic disorders.

[0190] HD-RNA has numerous potential applications due to its ability to enhance the generation of HDR by Cpf1 in cells after delivery with cationic lipids. Duchenne muscular dystrophy (DMD) was tested as the initial medical application for HD-RNA. DMD is an early-onset fatal disease caused by mutations in the dystrophin gene; it is the most common congenital myopathy, and approximately 30% of DMD patients have single-base mutations or small deletions, which can potentially be treated with HDR-based therapeutic agents.

[0191] Therefore, HD-RNA designed to target the dystrophin gene was tested for its ability to correct the dystrophin mutation in myoblasts derived from mdx mice via HDR. The HD-RNA was designed to cleave the dystrophin gene and also contained donor DNA designed to correct the C-to-T mutation present in its dystrophin gene (see Figure 10A and 10B ). The HDR rate in mdx myoblasts treated with Cpf1 / HD-RNA + cationic liposomes was determined and compared to mdx myoblasts treated with Cpf1-crRNA, donor DNA, and cationic liposomes. It was found that Cpf1 complexed with HD-RNA was more effective in generating HDR in mdx myoblasts compared to cells treated with Cpf1 RNP and donor DNA. For example, the cells treated with HD-RNA had an HDR rate of 5 - 10%, while the control cells had only a 1% HDR rate.

[0192] In another experiment, primary myoblasts isolated from Ai9 mice, a transgenic mouse strain containing stop codons in all three reading frames coupled to a triple poly(A) signal upstream of the tdTomato reporter gene, were electroporated with AsCpf1 RNP complexed with crRNAs with and without 5’ extensions. The Ai9 mice are a transgenic mouse strain containing the tdTomato reporter gene that has stop codons in all three reading frames coupled to the triple poly(A) signal. The AsCpf1 spacer was designed to introduce multiple breaks into the DNA, which results in the removal of the stop sequence by genomic deletion. Successful gene editing was indicated by the expression of tdTomato (red fluorescent protein, RFP), which could be visualized by fluorescence microscopy and quantified using flow cytometry. Extended crRNAs increased gene editing by more than 40 - 50% over unextended crRNAs. Myoblasts treated with unextended crRNAs were 12% RFP positive; myoblasts treated with 2-nucleotide extended crRNAs were 15% RFP positive; myoblasts treated with 9-nucleotide extended crRNAs were 18% RFP positive, while myoblasts treated with 59-nucleotide extended crRNAs were 16% RFP positive( Figure 10C ). Additionally, the efficiency of gene editing was tested using crRNAs in combination with ssDNA / ssRNA (100 nt) that had no sequence homology to the target DNA primary myoblasts. Both ssDNA and ssRNA enhanced the gene editing efficiency( Figure 10D ).

[0193] Collectively, the deletion of the target sequence in primary myoblasts suggests that extended crRNA-enhanced gene editing can be widely applied across genetic targets and cell types. These results confirm that HD-RNA can be used as a therapeutic agent for genetic diseases.

[0194] Example 10

[0195] The following examples demonstrate that the gene editing effect enhanced by 5' crRNA extension can be widely applied across genetic targets and cell types. Using Serpina1 as a test platform, these crRNAs were tested to see if they could enhance the ability of Cpf1 RNP to edit endogenous genes.

[0196] Using electroporation, primary myoblasts were transfected with crRNAs or crRNAs targeting the Serpina1 gene +9Cpf1 was transfected into HepG2 cells. Serpina1 was chosen for further investigation because mutations in the Serpina1 gene cause α1-antitrypsin deficiency, making it a target for therapeutic gene editing. Droplet digital PCR was performed on genomic DNA from HepG2 cells to quantify NHEJ efficiency.

[0197] like Figure 10E As shown in , compared with wild-type crRNA, +9 Cpf1 RNP has enhanced NHEJ efficiency. These results further indicate that the gene editing effect of 5'crRNA extension can be widely applied across genetic targets.

[0198] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0199] In the context of describing the present invention (especially in the context of the following claims), the use of the terms "one" and "a kind of", as well as "the / said" and "at least one / kind", and similar indicators, should be interpreted as covering both the singular and the plural, unless otherwise specified herein or clearly contradicted by the context. The use of the term "at least one" followed by a list of one or more items (e.g., "at least one of A and B") should be understood to mean an item selected from the listed items (A or B), or any combination of two or more listed items (A and B), unless otherwise specified herein or clearly contradicted by the context. Unless otherwise specified, the terms "comprising", "having", "including" and "containing" should be interpreted as open terms (i.e., meaning "including but not limited to"). Unless otherwise specified herein, the description of the numerical range herein is intended to serve only as a shorthand method of individually referring to each separate value falling within the range, and each separate value is incorporated into the specification as if it were individually described herein. Unless otherwise specified herein or clearly contradicted by the context, all methods described herein can be performed in any suitable order. Unless otherwise required, the use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the present invention and is not intended to limit the scope of the present invention. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the present invention.

[0200] This document describes the preferred embodiments of the present invention, including the best mode known to the inventors for practicing the present invention. After reading the foregoing specification, variations of those preferred embodiments will become apparent to those of ordinary skill in the art. The inventors expect skilled artisans to appropriately employ such variations, and the inventors intend the invention to be practiced otherwise than as specifically described herein. Accordingly, the present invention includes all modifications and equivalents of the subject matter recited in the appended claims as permitted by applicable law. In addition, any combination of the above-described elements in all possible variations thereof is covered by the present invention unless otherwise stated herein or clearly contradicted by context.

Claims

1. A nucleic acid comprising a Cpf1 crRNA, an extension sequence 5' of said crRNA, and optionally a processing sequence between said crRNA and said extension sequence, wherein said processing sequence is a sequence that is self-cleaved by Cpf1.

2. The nucleic acid of claim 1, wherein said nucleic acid comprises a processing sequence, and said processing sequence comprises a fragment of a direct repeat sequence of a Cpf1 array, wherein said direct repeat sequence comprises a crRNA sequence portion and a processing portion located 5' of said crRNA sequence portion, and said fragment comprises at least 5 contiguous nucleotides of the processing portion of said direct repeat sequence.

3. The nucleic acid of claim 2, wherein said processing sequence comprises a fragment of at least 10 nucleotides of the processing portion of said direct repeat sequence.

4. The nucleic acid of claim 2, wherein said processing sequence comprises the entire processing portion of said direct repeat sequence.

5. The nucleic acid of any one of claims 1-4, wherein said extension sequence does not comprise the sequence of said processing sequence or said crRNA.

6. The nucleic acid of any one of claims 1-5, wherein said extension sequence comprises at least 2 nucleotides.

7. The nucleic acid of any one of claims 1-6, wherein said extension sequence comprises 10 to 100 nucleotides.

8. The nucleic acid of any one of claims 1-7, wherein said nucleic acid contains only a single Cpf1 crRNA sequence.

9. The nucleic acid of any one of claims 1 to 8, further comprising a second processing sequence 5' of said extension sequence and a second extension sequence 5' of said second processing sequence.

10. The nucleic acid of any one of claims 1-9, further comprising a donor nucleic acid hybridized or covalently linked thereto.

11. The nucleic acid of claim 10, wherein said donor nucleic acid is covalently linked to the 5′ of said processing sequence or the 5′ of said extension sequence.

12. The nucleic acid of claim 11, wherein said donor nucleic acid is linked to said processing sequence or extension sequence via a linker group.

13. The nucleic acid of claim 10, wherein said donor nucleic acid hybridizes to said extension sequence and / or processing sequence.

14. The nucleic acid of any one of claims 1-13, further comprising a target nucleic acid 3' of said Cpf1 crRNA.

15. The nucleic acid of any one of claims 1-14, wherein said extension sequence comprises less than about 60 nucleotides.

16. The nucleic acid of claim 15, wherein said extension sequence comprises less than about 20 nucleotides.

17. The nucleic acid of claim 15, wherein said extension sequence comprises about 2-20 nucleotides.

18. The nucleic acid of any one of claims 1-17, wherein said nucleic acid does not comprise a processing sequence.

19. The nucleic acid of claim 18, wherein said nucleic acid further comprises a donor nucleic acid covalently linked to the 5′ of said extension sequence.

20. The nucleic acid of claim 19, wherein said donor nucleic acid is linked via a linker group.

21. The nucleic acid of claim 18, wherein said nucleic acid further comprises a donor nucleic acid hybridized to said extension sequence.

22. The nucleic acid of any one of claims 18 - 21, further comprising a targeting nucleic acid of the 3' of said Cpf1 crRNA.

23. The nucleic acid of any one of claims 1 - 22, wherein said extension sequence comprises a self - hybridizing sequence.

24. The nucleic acid of any one of claims 1 - 23, wherein said extension sequence comprises a semi - stable hairpin structure, a stable hairpin structure, a pseudoknot structure, a G - quadruplex structure, a bulge loop structure, an internal loop structure, a branched loop structure, or a combination thereof.

25. The nucleic acid of claim 24, wherein said extension sequence comprises a repetitive trinucleotide motif.

26. The nucleic acid of any one of claim 25, wherein said repetitive trinucleotide motif is CAA, UUG, AAG, CUU, CCU, CCA, UAA, or a combination thereof.

27. The nucleic acid of any one of claim 25, wherein said repetitive trinucleotide motif is CAU, CUA, UUA, AUG, UAG, or a combination thereof.

28. The nucleic acid of any one of claim 25, wherein said repetitive trinucleotide motif is CGA, CGU, CGG, CAG, CUG, CCG, or a combination thereof.

29. The nucleic acid of any one of claim 25, wherein said repetitive trinucleotide motif is a CNG motif or a combination of CNG motifs, optionally with CGA or CGU.

30. The nucleic acid of any one of claim 25, wherein said repetitive trinucleotide motif is AGG, UGG, or a combination thereof.

31. The nucleic acid of any one of claims 1 - 25, wherein said extension sequence comprises a combination of repetitive trinucleotide motifs of any one of claims 26 - 30.

32. The nucleic acid of any one of claims 1 - 31, wherein said extension sequence or a portion thereof is resistant to nuclease degradation.

33. The nucleic acid of any one of claims 1 - 32, wherein said extension sequence comprises one or more modified internucleotide linkages.

34. The nucleic acid of claim 33, wherein a region of 4 or more adjacent nucleotides of said extension sequence, or the entire extension sequence, has modified internucleotide linkages.

35. The nucleic acid of claim 33 or 34, wherein said modified internucleotide linkages comprise phosphorothioate, dithiophosphonate, methylphosphonate, aminophosphonate, 2'-O - methyl, 2'-O - methoxyethyl, 2'-fluoro, bridged nucleic acid (BNA), or phosphotriester - modified linkages or a combination thereof.

36. The nucleic acid of any one of claims 1 - 35, wherein said extension sequence comprises one or more xeno nucleic acids (XNA).

37. The nucleic acid of any one of claims 1 - 36, wherein said nucleic acid further comprises a biotin and / or avidin or streptavidin molecule attached to the 5' end of said extension sequence.

38. The nucleic acid of claim 37, wherein said nucleic acid further comprises a biotin molecule attached to the 5' end of said extension sequence, an avidin or streptavidin molecule conjugated to said biotin molecule, and optionally, a targeting construct comprising a biotin molecule conjugated to avidin or streptavidin.

39. The nucleic acid of claim 38, wherein the targeting construct is a peptide comprising a biotin molecule attached thereto.

40. A composition comprising the nucleic acid of any one of claims 1-39 and a vector, and optionally further comprising a Cpf1 protein.

41. The composition of claim 40, wherein the composition is substantially free of divalent metal ions that promote Cpf1 cleavage.

42. The composition of claim 40 or 41, wherein the composition comprises a cationic lipid.

43. The composition of any one of claims 40-42, wherein the nucleic acid is in a liposome.

44. The composition of any one of claims 40-42, wherein the nucleic acid is partially or fully encapsulated by a polymeric nanoparticle or attached to a metal or polymeric nanoparticle.

45. A method of genetically modifying a eukaryotic target cell, comprising contacting the eukaryotic target cell with the nucleic acid of any one of claims 1-39 or the composition of any one of claims 40-44 to genetically modify a target nucleic acid.

46. The method of claim 45, wherein the Cpf1 crRNA comprises a targeting sequence that hybridizes to a target sequence in the target cell.

47. The method of claim 45 or 46, wherein the target cell is a mammalian cell, optionally a human cell.

48. The method of any one of claims 45-47, wherein the nucleic acid construct is cleaved after entering the cell, thereby releasing the Cpf1 crRNA from the other parts of the nucleic acid construct.

Citation Information

Patent Citations

  • Drug delivery device

    US4863457A

  • Sustained release drug delivery devices

    US5378475A

  • Biocompatible ocular implants

    US5443505A