CRISPR system having an engineered dual-guide nucleic acid
The dual-guide CRISPR-Cas system addresses the challenge of off-target editing by splitting the single-guide RNA into two nucleic acids, enhancing specificity and editing efficiency in genome editing applications.
Patent Information
- Application Number
- JP2022520664
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-03
- Filing Date
- 2020-10-02
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2040-10-02
AI Technical Summary
Current CRISPR-Cas systems face challenges in achieving high specificity and reducing off-target editing, particularly in applications involving genetically engineered cells for therapeutic use.
The development of a dual-guide CRISPR-Cas system that splits the single-guide RNA into two separate nucleic acids, a targeter nucleic acid and a modulator nucleic acid, allowing for greater flexibility and tunability in nucleic acid cleavage efficiency and specificity.
This dual-guide system enhances the specificity of genome editing by adjusting the hybridization length and affinity of the targeter and modulator nucleic acids, thereby reducing off-target editing and improving the effectiveness of nucleic acid editing.
Smart Images

Figure 0007689949000034 
Figure 0007689949000035 
Figure 0007689949000036
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 910,055, filed Oct. 3, 2019, the disclosure of which is incorporated herein by reference in its entirety for all purposes.
[0002] Field of the Invention The present invention relates to engineered dual-guide nucleic acids (e.g., RNAs) that can activate Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) nucleases that are clustered at regular intervals, methods of targeting, editing, and / or modifying nucleic acids using the engineered CRISPR system, and compositions and cells comprising the engineered CRISPR system.
Background Art
[0003] Background of the Invention Recently, sophisticated genome targeting technologies have advanced. For example, specific loci within genomic DNA can be targeted, edited, or modified in another manner by designer meganucleases, zinc finger nucleases, or transcription activator-like effector (TALE). Furthermore, the CRISPR-Cas systems of bacterial and archaeal adaptive immunity have been adapted for precise targeting of genomic DNA in eukaryotic cells. Compared to previous generations of genome editing tools, the CRISPR-Cas system is easy to set up, scalable, and suitable for targeting multiple positions within the genomes of eukaryotes, thereby providing a major means for new applications in genome engineering.
[0004] Two distinct classes of CRISPR-Cas systems have been identified. In class 1 CRISPR-Cas systems, effector complexes of multiple proteins are used, while in class 2 CRISPR-Cas systems, effector of a single protein is used (see Makarova et al. (2017) CELL, 168:328 (Non-Patent Document 1)). Among the three types of class 2 CRISPR-Cas systems, type II and type V systems typically target DNA, and type VI systems typically target RNA (ibid.). Naturally occurring type II effector complexes consist of Cas9, CRISPR RNA (crRNA), and trans-activating CRISPR RNA (tracrRNA), but in engineered systems, crRNA and tracrRNA can be fused as a single guide RNA for simplicity (see Wang et al. (2016) ANNU. REV. BIOCHEM., 85:227 (Non-Patent Document 2)). Some specific naturally occurring type V systems, such as type V-A, type V-C, and type V-D systems, do not require tracrRNA, and crRNA is used alone as a guide for cleavage of target DNA (see Zetsche et al. (2015) CELL, 163:759 (Non-Patent Document 3); Makarova et al. (2017) CELL, 168:328 (Non-Patent Document 1)).
[0005] CRISPR-Cas systems have been engineered for various purposes, such as for cleavage of genomic DNA, base editing, epigenome editing, and genome imaging (see, for example, Wang et al. (2016) ANNU. REV. BIOCHEM., 85:227 (Non-Patent Document 2) and Rees et al. (2018) NAT. REV. GENET., 19:770 (Non-Patent Document 4)). Considerable development has been carried out, but there is still a need for new and useful CRISPR-Cas systems as powerful and precise genome targeting tools.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Summary of the Invention
[0007] In part, the present invention is based on the design of a dual-guide CRISPR-Cas system that can activate a Cas nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system when a targeter nucleic acid and a modulator nucleic acid hybridize to form a complex. The engineered dual-guide CRISPR-Cas system described herein can be used to target, edit, or modify a target nucleic acid such as genomic DNA.
[0008] Type V-A, V-C, and V-D CRISPR-Cas systems naturally contain a Cas nuclease and a single-guide RNA (i.e., crRNA). By splitting this single-guide RNA into two different nucleic acids, the engineered systems described herein provide greater flexibility and tunability. For example, the efficiency of nucleic acid cleavage can be increased or decreased by adjusting the hybridization length and / or the affinity of the targeter nucleic acid and the modulator nucleic acid. Furthermore, considering the limitations on the length of nucleic acids that can be synthesized with high yield and accuracy, the use of dual-guide nucleic acids allows for the incorporation of more polynucleotide elements that can improve the effectiveness and / or specificity of editing.
[0009] In particular, this dual-guide system can be operated as an adjustable system for reducing off-target editing and can thus be used for editing nucleic acids with high specificity. The system can be used in a number of applications, such as the editing of cells for therapeutic use, such as mammalian cells. Reduction of off-target editing is particularly desirable when generating genetically engineered proliferative cells, such as stem cells, progenitor cells and immunological memory cells, that are administered to a subject in need of treatment. High specificity can be achieved using the dual-guide systems described herein, and the dual-guide systems optionally further comprise one or more chemical modifications to, for example, the targeter nucleic acid and / or the modulator nucleic acid, the editing enhancer sequence and / or the donor template recruitment sequence.
[0010] Thus, in one aspect, the invention provides (a) (i) a spacer sequence designed to hybridize to a target nucleotide sequence; and (ii) a targeter stem sequence comprising a targeter nucleic acid; and (b) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence comprising an engineered, non-naturally occurring system, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids and a complex comprising the targeter nucleic acid and the modulator nucleic acid can activate a CRISPR-associated (Cas) nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system. An engineered, non-naturally occurring system is provided.
[0011] In some particular embodiments, the Cas nuclease is a type V-A Cas nuclease.
[0012] In some specific embodiments, the targeter stem sequence and the modulator stem sequence are each 4 to 10 nucleotides in length. In some specific embodiments, the targeter stem sequence and the modulator stem sequence are each 5 nucleotides in length. In some specific embodiments, the targeter stem sequence and the modulator stem sequence hybridize by Watson-Crick base pairing.
[0013] In some specific embodiments, the spacer sequence is about 20 nucleotides in length. In some specific embodiments, the spacer sequence is 18 nucleotides in length or shorter. In some specific embodiments, the spacer sequence is 17 nucleotides in length or shorter.
[0014] In some specific embodiments, the targeter nucleic acid comprises, from 5' to 3', a targeter stem sequence, a spacer sequence, and any additional nucleotide sequence.
[0015] In some specific embodiments, the targeter nucleic acid comprises ribonucleic acid (RNA). In some specific embodiments, the targeter nucleic acid comprises modified RNA. In some specific embodiments, the targeter nucleic acid comprises a combination of RNA and DNA. In some specific embodiments, the targeter nucleic acid comprises chemical modifications. In some specific embodiments, the chemical modification is present in one or more nucleotides at the 3' end of the targeter nucleic acid. In some specific embodiments, the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof.
[0016] In some specific embodiments, the modulator nucleic acid further comprises an additional nucleotide sequence. In some specific embodiments, the additional nucleotide sequence is located on the 5' side of the modulator stem sequence. In some specific embodiments, the additional nucleotide sequence is 4 to 50 nucleotides in length. In some specific embodiments, the additional nucleotide sequence comprises a donor template recruit sequence that can hybridize with the donor template. In some specific embodiments, the engineered non-natural system further comprises a donor template. In some specific embodiments, the modulator nucleic acid comprises one or more nucleotides on the 3' side of the modulator stem sequence.
[0017] In some specific embodiments, the modulator nucleic acid comprises RNA. In some specific embodiments, the modulator nucleic acid comprises modified RNA. In some specific embodiments, the modulator nucleic acid comprises a combination of RNA and DNA. In some specific embodiments, the modulator nucleic acid comprises chemical modifications. In some specific embodiments, the chemical modification is present in one or more nucleotides at the 5' end of the modulator nucleic acid. In some specific embodiments, the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof.
[0018] In some specific embodiments, the targeter nucleic acid and the modulator nucleic acid are not covalently linked.
[0019] In some specific embodiments, the Cas nuclease comprises an amino acid sequence that is at least 80% identical to SEQ ID NO:1. In some specific embodiments, the Cas nuclease is Cpf1. In some specific embodiments, the engineered non-natural system further comprises a Cas nuclease. In some specific embodiments, the targeter nucleic acid, the modulator nucleic acid, and the Cas nuclease are present within a ribonucleoprotein (RNP) complex.
[0020] In another aspect, the present invention provides a eukaryotic cell comprising an engineered non-naturally occurring system disclosed herein.
[0021] In another aspect, the present invention provides a composition (e.g., a pharmaceutical composition) comprising an engineered non-naturally occurring system or eukaryotic cell disclosed herein.
[0022] In another aspect, the present invention provides a method of cleaving a target DNA having a target nucleotide sequence, comprising the step of contacting the target DNA with an engineered non-naturally occurring system disclosed herein, thereby resulting in cleavage of the target DNA.
[0023] In some specific embodiments, the contacting is performed in vitro.
[0024] In some specific embodiments, the contacting is performed ex vivo in a cell. In some specific embodiments, the target DNA is the genomic DNA of the cell. In some specific embodiments, the system is delivered into the cell as a pre-formed RNP complex. In some specific embodiments, the pre-formed RNP complex is delivered into the cell by electroporation.
[0025] In another aspect, the present invention provides a method of editing the genome of a eukaryotic cell, comprising the step of delivering into the eukaryotic cell an engineered non-naturally occurring system disclosed herein, thereby resulting in editing of the genome of the eukaryotic cell.
[0026] In some specific embodiments, the system is delivered into the cell as a pre-formed RNP complex. In some specific embodiments, the system is delivered into the cell by electroporation.
[0027] In some specific embodiments of the method involving a eukaryotic cell, the cell is an immune cell. In some specific embodiments, the immune cell is a T lymphocyte. [The present invention 1001] (a)(i) A spacer sequence designed to hybridize with a target nucleotide sequence; and (ii) A targeter stem sequence comprising a targeter nucleic acid; and (b) A modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence comprising an engineered, non-naturally occurring system, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and a complex comprising the targeter nucleic acid and the modulator nucleic acid can activate a CRISPR-associated (Cas) nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system, an engineered, non-naturally occurring system. [The present invention 1002] The engineered, non-naturally occurring system of the present invention 1001, wherein the Cas nuclease is a type V-A Cas nuclease. [The present invention 1003] The engineered, non-naturally occurring system of the present invention 1001 or 1002, wherein the targeter stem sequence and the modulator stem sequence are each 4 to 10 nucleotides in length. [The present invention 1004] The engineered, non-naturally occurring system of any one of the present inventions 1001 to 1003, wherein the targeter stem sequence and the modulator stem sequence are each 5 nucleotides in length. [The present invention 1005] The engineered, non-naturally occurring system of any one of the present inventions above, wherein the targeter stem sequence and the modulator stem sequence hybridize by Watson-Crick base pairing. [The present invention 1006] The engineered, non-naturally occurring system of any one of the present inventions 1001 to 1005, wherein the spacer sequence is about 20 nucleotides in length. [The present invention 1007] The engineered, non-naturally occurring system of any one of the present inventions 1001 to 1005, wherein the spacer sequence is 18 nucleotides in length or shorter. [The present invention 1008] The engineered, non-naturally occurring system of the present invention 1007, wherein the spacer sequence is 17 nucleotides in length or shorter. [The present invention 1009] The engineered, non-naturally occurring system of any one of the present inventions above, wherein the targeter nucleic acid comprises a targeter stem sequence, a spacer sequence, and any additional nucleotide sequence, in that order from 5' to 3'. [The present invention 1010] Any engineered, non-naturally occurring system of the present invention, wherein the targeter nucleic acid comprises ribonucleic acid (RNA). [Inventive concept 1011] An engineered, non-naturally occurring system of inventive concept 1010, wherein the targeter nucleic acid comprises modified RNA. [Inventive concept 1012] An engineered, non-naturally occurring system of inventive concept 1010 or 1011, wherein the targeter nucleic acid comprises a combination of RNA and DNA. [Inventive concept 1013] Any engineered, non-naturally occurring system of inventive concepts 1010 to 1012, wherein the targeter nucleic acid comprises chemical modifications. [Inventive concept 1014] An engineered, non-naturally occurring system of inventive concept 1013, wherein the chemical modification is present in one or more nucleotides at the 3'-end of the targeter nucleic acid. [Inventive concept 1015] An engineered, non-naturally occurring system of inventive concept 1013 or 1014, wherein the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof. [Inventive concept 1016] Any engineered, non-naturally occurring system of the present invention, wherein the modulator nucleic acid further comprises an additional nucleotide sequence. [Inventive concept 1017] An engineered, non-naturally occurring system of inventive concept 1016, wherein the additional nucleotide sequence is located 5' to the modulator stem sequence. [Inventive concept 1018] An engineered, non-naturally occurring system of inventive concepts 1016 or 1017, wherein the additional nucleotide sequence is 4 to 50 nucleotides in length. [Inventive concept 1019] An engineered, non-naturally occurring system of any of inventive concepts 1016 to 1018, wherein the additional nucleotide sequence comprises a donor template recruit sequence capable of hybridizing to a donor template. [Inventive concept 1020] An engineered, non-naturally occurring system of inventive concept 1019, further comprising a donor template. [Inventive concept 1021] Any engineered, non-naturally occurring system of the present invention, wherein the modulator nucleic acid comprises one or more nucleotides 3' to the modulator stem sequence. [Inventive concept 1022] Any engineered, non-naturally occurring system of the present invention, wherein the modulator nucleic acid comprises RNA. [Inventive concept 1023] An engineered, non-naturally occurring system of the invention 1022, wherein the modulator nucleic acid comprises a modified RNA. [The invention 1024] An engineered, non-naturally occurring system of the invention 1022 or 1023, wherein the modulator nucleic acid comprises a combination of RNA and DNA. [The invention 1025] An engineered, non-naturally occurring system of any one of the invention 1022 to 1024, wherein the modulator nucleic acid comprises a chemical modification. [The invention 1026] An engineered, non-naturally occurring system of the invention 1025, wherein the chemical modification is present in one or more nucleotides at the 5'-end of the modulator nucleic acid. [The invention 1027] An engineered, non-naturally occurring system of the invention 1025 or 1026, wherein the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof. [The invention 1028] An engineered, non-naturally occurring system of any of the above inventions, wherein the targeter nucleic acid and the modulator nucleic acid are not covalently linked. [The invention 1029] An engineered, non-naturally occurring system of any of the invention 1001 to 1028, wherein the Cas nuclease comprises an amino acid sequence that is at least 80% identical to SEQ ID NO:1. [The invention 1030] An engineered, non-naturally occurring system of any of the invention 1001 to 1028, wherein the Cas nuclease is Cpf1. [The invention 1031] An engineered, non-naturally occurring system of any of the above inventions, further comprising a Cas nuclease. [The invention 1032] An engineered, non-naturally occurring system of the invention 1031, wherein the targeter nucleic acid, the modulator nucleic acid, and the Cas nuclease are present within a ribonucleoprotein (RNP) complex. [The invention 1033] An engineered, non-naturally occurring system of any of the invention 1001 to 1032 comprising a eukaryotic cell. [The invention 1034] An engineered, non-naturally occurring system of any of the invention 1001 to 1032, or the eukaryotic cell of the invention 1033 comprising a composition. [The invention 1035] A method of cleaving a target DNA having a target nucleotide sequence, comprising the following steps: A step of contacting a target DNA with any of the engineered, non-naturally occurring systems of the present invention 1001 to 1032, thereby resulting in cleavage of the target DNA. [The present invention 1036] The method of the present invention 1035, wherein the contact is performed in vitro. [The present invention 1037] The method of the present invention 1035, wherein the contact is performed ex vivo in a cell. [The present invention 1038] The method of the present invention 1037, wherein the target DNA is the genomic DNA of the cell. [The present invention 1039] The method of the present invention 1037 or 1038, wherein the system is delivered into the cell as a pre-formed RNP complex. [The present invention 1040] The method of the present invention 1039, wherein the pre-formed RNP complex is delivered into the cell by electroporation. [The present invention 1041] A method for editing the genome of a eukaryotic cell, comprising the following steps: A step of delivering any of the engineered, non-naturally occurring systems of the present invention 1001 to 1032 into a eukaryotic cell, thereby resulting in editing of the genome of the eukaryotic cell. [The present invention 1042] The method of the present invention 1041, wherein the system is delivered into the cell as a pre-formed RNP complex. [The present invention 1043] The method of the present invention 1041 or 1042, wherein the system is delivered into the cell by electroporation. [The present invention 1044] The method of any of the present invention 1037 to 1043, wherein the cell is an immune cell. [The present invention 1045] The method of the present invention 1044, wherein the immune cell is a T lymphocyte.
Brief Description of the Drawings
[0028] [Figure 1A] It is a schematic diagram showing the structure of an exemplary dual-guide V-A type CRISPR-Cas system. [Figure 1B] Figures 1B to 1D are a series of schematic diagrams showing the incorporation of a protecting group (e.g., a protecting nucleotide sequence or a chemical modification moiety) (Figure 1B), a donor template recruitment sequence (Figure 1C), and an editing enhancer (Figure 1D) into the dual-guide V-A type CRISPR-Cas system. [Figure 1C] Refer to the description of Figure 1B. [Figure 1D] Refer to the description of Figure 1B. [Figure 2-1]Figure 2A is a series of schematic diagrams showing the predicted secondary structures of two crRNAs tested in in vitro cleavage experiments. Figure 2B is a photograph showing the results of gel electrophoresis of an in vitro cleavage experiment using MAD7 complexed with two different crRNAs, designated "crRNA1" and "crRNA2", chemically transcribed, and their corresponding sets of targeter RNA and modulator RNA. [Figure 2-2] Refer to the description of Figure 2-1. [Figure 3] A photograph showing the results of gel electrophoresis of an in vitro cleavage experiment using MAD7 complexed with three different crRNAs, designated "crRNA1", "crRNA3", and "crRNA4", prepared either by chemical synthesis or in vitro transcription, and their corresponding sets of targeter RNA and modulator RNA. [Figure 4-1] Figures 4A - 4H are a series of schematic diagrams showing the predicted secondary structures of hybridized targeter RNA and modulator RNA. The × mark (within the loop region) indicates the site where the RNA is split into targeter RNA and modulator RNA. In Figures 4A - 4F, RNA#1 is a single guide RNA. RNA#2, #4, #6, #8, and #10 are modulator RNAs, and RNA#3, #5, #7, #9, and #11 are targeter RNAs. In Figures 4G - 4H, RNA#12 and #14 are single guide RNAs containing hairpin sequences. RNA#13 is the modulator RNA corresponding to RNA#12, and RNA#15 is the targeter RNA corresponding to RNA#14. Figure 4I is a set of photographs showing the results of gel electrophoresis of an in vitro cleavage experiment using MAD7 complexed with a combination of targeter RNA and modulator RNA. [Figure 4-2] Refer to the description of Figure 4-1. [Figure 4-3] Refer to the description of Figure 4-1. [Figure 4-4] Refer to the description of Figure 4-1. [Figure 4-5] Refer to the description of Figure 4-1. [Figure 5-1]Figures 5A-5I are a series of schematic diagrams showing the putative secondary structures of crRNAs. When the crRNA is split into a combination of a modulator RNA and a targeter RNA, thick x marks (within loop regions, corresponding to combinations 3, 5, 7, 9, 11, 13, and 15) and thin x marks (within stem regions, corresponding to combinations 4, 6, 8, 10, 12, 14, and 16) indicate the sites where the crRNA is split. The Gibbs free energy change (ΔG) during the formation of the secondary structure of the corresponding crRNA, estimated by the RNAfold program, is noted for each construct or combination. Figures 5J-5K are photographs showing the results of gel electrophoresis of in vitro cleavage experiments using MAD7 complexed with a crRNA construct or a combination of a targeter RNA and a modulator RNA. The ratio of cleavage products in Figure 5J was determined by measuring the relative intensity of the bands. [Figure 5-2] Refer to the description of Figure 5-1. [Figure 5-3] Refer to the description of Figure 5-1. [Figure 5-4] Refer to the description of Figure 5-1. [Figure 5-5] Refer to the description of Figure 5-1. [Figure 5-6] Refer to the description of Figure 5-1. [Figure 6A] A bar graph showing the read number ratio of edited and unedited copies of target DNA by each of the tested crRNAs or corresponding combinations of a targeter RNA and a modulator RNA. "rep1" and "rep2" respectively mean the first and second replicates of the same experiment. [Figure 6B] A bar graph showing the number of sequencing reads obtained under each condition. The color indicates the quality of the reads. [Figure 7] A bar graph showing the percentage of edited copies of the target locus (shown on the x-axis) within the genome of Jurkat cells. [Figure 8]Bar graph showing the percentage of genomic copies edited in the CD52, PDCD1, or TIGIT gene of Jurkat cells after delivery of a dual-guide CRISPR system with crRNAs split at different positions (1st, 2nd, 3rd, 4th, or 5th nucleotide relative to the 5' end of the loop).
Mode for Carrying Out the Invention
[0029] Detailed Description of the Invention In part, the present invention is based on the design of a dual-guide CRISPR-Cas system that can activate a Cas nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system when a targeter nucleic acid and a modulator nucleic acid hybridize to form a complex. The engineered dual-guide CRISPR-Cas system described herein can be used to target, edit, or modify a target nucleic acid such as genomic DNA.
[0030] Type V-A, V-C, and V-D CRISPR-Cas systems naturally contain a Cas nuclease and a single-guide RNA (i.e., crRNA). Splitting this single-guide RNA into two different nucleic acids provides greater flexibility and tunability by the engineered systems described herein. For example, the efficiency of nucleic acid cleavage can be increased or decreased by adjusting the hybridization length and / or the affinity of the targeter nucleic acid and the modulator nucleic acid. Furthermore, considering the limitations on the length of nucleic acids that can be synthesized with high yield and accuracy, the use of dual-guide nucleic acids allows for the incorporation of more polynucleotide elements that can improve the efficacy and / or specificity of editing.
[0031] In particular, this dual-guide system can be operated as an adjustable system for reducing off-target editing and can thus be used for editing nucleic acids with high specificity. The system can be used in the editing of cells for several applications, such as for use in therapy, for example mammalian cells. Reducing off-target editing is particularly desirable when generating genetically engineered proliferating cells, such as stem cells, progenitor cells and immune memory cells, which are administered to a subject in need of treatment. High specificity can be achieved using the dual-guide systems described herein, and the dual-guide systems optionally further comprise one or more chemical modifications, for example, to the targeter nucleic acid and / or the modulator nucleic acid, the editing enhancer sequence and / or the donor template recruitment sequence.
[0032] The features and uses of the dual-guide CRISPR-Cas system are considered in the following sections.
[0033] I. Engineered, non-naturally occurring dual-guide CRISPR-Cas systems The engineered non-naturally occurring system of the invention comprises (a)(i) a spacer sequence designed to hybridize to a target nucleotide sequence; and (ii) a targeter stem sequence comprised in a targeter nucleic acid; and (b) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and the complex comprising the targeter nucleic acid and the modulator nucleic acid is capable of activating a Cas nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system.
[0034] The V-A, V-C, and V-D CRISPR-Cas systems are distinct subtypes of CRISPR-Cas systems based on the classification described in Makarova et al. (2017) CELL, 168:328. The naturally occurring CRISPR-Cas systems of these subtypes lack tracrRNA and rely on a single crRNA to guide the CRISPR-Cas complex to the target DNA. Naturally occurring V-A Cas proteins contain an RuvC-like nuclease domain but lack an HNH endonuclease domain and recognize a 5’T-rich protospacer adjacent motif (PAM) that is the 5’ orientation determined using the non-target strand (i.e., the strand that does not hybridize to the spacer sequence) as the coordinate strand. The naturally occurring V-A CRISPR-Cas system cleaves double-stranded DNA to produce double-stranded breaks that are staggered rather than blunt-ended. The cleavage site is distal to the PAM site (e.g., at least 10, 11, 12, 13, 14, or 15 nucleotides away from the PAM on the non-target strand and / or at least 15, 16, 17, 18, or 19 nucleotides away from the sequence complementary to the PAM on the target strand).
[0035] Thus, in another aspect, the present disclosure provides (a) (i) a spacer sequence designed to hybridize to a target nucleotide sequence; and (ii) a targeter stem sequence comprising a targeter nucleic acid; and (b) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence comprising an engineered, non-naturally occurring system, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and a complex comprising the targeter nucleic acid and the modulator nucleic acid is capable of activating a V-A, V-C, or V-D type Cas nuclease, providing an engineered, non-naturally occurring system. In some particular embodiments, the Cas nuclease is a V-A type Cas nuclease.
[0036] Cas protein As used interchangeably herein, the terms "CRISPR-associated protein", "Cas protein", and "Cas" refer to naturally occurring Cas proteins or engineered Cas proteins. Non-limiting examples of Cas protein engineering include, without limitation, Cas protein mutations and Cas protein modifications that modify Cas activity, modify PAM specificity, broaden the range of recognized PAMs, and / or reduce the ability to modify one or more off-target loci compared to the corresponding unmodified Cas. In some particular embodiments, modification of the activity of engineered Cas includes modification of the ability (e.g., specificity or kinetics) to bind to naturally occurring crRNAs or engineered dual-guide nucleic acids, modification of the ability (e.g., specificity or kinetics) to bind to target nucleotide sequences, modification of the processivity of nucleic acid scanning, and / or modification of effector (e.g., nuclease) activity. Cas proteins having nuclease activity are referred to as "CRISPR-associated nucleases" or "Cas nucleases", and these are used interchangeably herein.
[0037] In some particular embodiments, the Cas nuclease that can be activated by a complex comprising a targeter nucleic acid and a modulator nucleic acid is a type V-A, type V-C, or type V-D Cas nuclease. In some particular embodiments, the Cas nuclease is a type V-A nuclease.
[0038] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1. The Cpf1 protein is known in the art and is described in US Patent Nos. 9,790,490 and 10,113,179. Cpf1 orthologs can be found in the genomes of various bacteria and archaea.For example, in some specific embodiments, the Cpf1 protein is derived from Francisella novicida U112 (Fn), Acidaminococcus sp. BV3L6 (As), Ruminococcaceae bacterium ND2006 (Lb), Ruminococcaceae bacterium MA2020 (Lb2), Candidatus Methanoplasma termitum (CMt), Moraxella bovoculi 237 (Mb), Porphyromonas crevioricanis (Pc), Prevotella disiens (Pd), Francisella tularensis 1, the subspecies novicida of Francisella tularensis, Prevotella albensis, Ruminococcaceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Eubacterium eligens, Leptospira inadai, Porphyromonas macacae, Prevotella bryantii, Proteocatella sphenisci, Anaerovibrio sp. RM50, Moraxella caprae, Ruminococcaceae bacterium COE1 or Eubacterium coprostanoligenes.
[0039] In some specific embodiments, the V-A type Cas nuclease comprises AsCpf1 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:3. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:3. TIFF0007689949000001.tif174159
[0040] In some specific embodiments, the V-A type Cas nuclease comprises LbCpf1 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:4. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:4. TIFF0007689949000002.tif113160
[0041] In some specific embodiments, the V-A type Cas nuclease comprises FnCpf1 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:5. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:5. TIFF0007689949000003.tif118160
[0042] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of Prevotella bryantii or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:6. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:6. TIFF0007689949000004.tif114160
[0043] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of Proteocatella sphenisci or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:7. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:7. TIFF0007689949000005.tif104159
[0044] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of Anaerovibrio sp. RM50 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:8. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:8. TIFF0007689949000006.tif114160
[0045] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of Moraxella catarrhalis or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:9. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:9. TIFF0007689949000007.tif119159
[0046] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of the bacterium COE1 of the family Ruminococcaceae or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:10. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:10. TIFF0007689949000008.tif114159
[0047] In some specific embodiments, the V-A type Cas nuclease comprises Cpf1 of Butyribacterium coprostanoligenes or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:11. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:11. TIFF0007689949000009.tif118159
[0048] In some specific embodiments, the V-A type Cas nuclease is not Cpf1. In some specific embodiments, the V-A type Cas nuclease is not AsCpf1.
[0049] In some specific embodiments, the V-A type Cas nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19 or MAD20 or a variant thereof. MAD1-MAD20 are known in the art and are described in U.S. Patent No. 9,982,279.
[0050] In some specific embodiments, the V-A type Cas nuclease comprises MAD7 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:1. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:1. TIFF0007689949000010.tif114159
[0051] In some specific embodiments, the V-A type Cas nuclease comprises MAD2 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:2. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:2. TIFF0007689949000011.tif118160
[0052] In some particular embodiments, the type V-A Cas nuclease comprises Csm1. The Csm1 protein is known in the art and is described in U.S. Patent No. 9,896,696. Csm1 orthologs can be found in the genomes of various bacteria and archaea. For example, in some particular embodiments, the Csm1 protein is derived from the species SCADC(Sm) of the genus Sumicella, the species (Ss) of the genus Sulfuricurvum, or the Microgenomate (Roizmanbacteria) bacterium (Mb).
[0053] In some particular embodiments, the type V-A Cas nuclease comprises SmCsm1 or a variant thereof. In some particular embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:12. In some particular embodiments, the type V-A Cas protein comprises the amino acid sequence shown in SEQ ID NO:12. TIFF0007689949000012.tif99160
[0054] In some specific embodiments, the V-A type Cas nuclease comprises SsCsm1 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:13. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:13. TIFF0007689949000013.tif114160
[0055] In some specific embodiments, the V-A type Cas nuclease comprises MbCsm1 or a variant thereof. In some specific embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:14. In some specific embodiments, the V-A type Cas protein comprises the amino acid sequence shown in SEQ ID NO:14. TIFF0007689949000014.tif98159
[0056] Additional V-A type Cas nucleases and their corresponding naturally occurring CRISPR-Cas systems can be identified by methods using a computer and experimental methods known in the art, such as those described in U.S. Patent No. 9,790,490 and Shmakov et al. (2015) Mol. Cell, 60:385. Exemplary methods using a computer include analysis of putative Cas proteins by homology modeling, structural BLAST, PSI-BLAST, or HHPred, and analysis of putative CRISPR loci by identification of CRISPR arrays. Exemplary experimental methods include in vitro cleavage assays and intracellular nuclease assays (e.g., Surveyor assays) as described in Zetsche et al. (2015) Cell, 163:759.
[0057] In some particular embodiments, the Cas nuclease is directed to cleavage of one or both strands of the target locus, such as the target strand (i.e., the strand having the target nucleotide sequence that hybridizes with the single guide nucleic acid or dual guide nucleic acid) and / or the non-target strand. In some particular embodiments, the Cas nuclease is directed to cleavage of one or both strands within a range of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 50, about 100, about 200, about 500 or more nucleotides from the first or last nucleotide of the target nucleotide sequence or its complementary sequence. In some particular embodiments, the cleavage is staggered, i.e., generates sticky ends. In some particular embodiments, the cleavage generates a staggered cleavage site having a 5' overhang. In some particular embodiments, the cleavage generates a staggered cleavage site having a 5' overhang of 1 to 5 nucleotides, such as 4 or 5 nucleotides. In some particular embodiments, the cleavage site is distal from the PAM, e.g., the cleavage occurs behind the 18th nucleotide of the non-target strand and behind the 23rd nucleotide of the target strand.
[0058] In some particular embodiments, the engineered non-naturally occurring system of the invention further comprises a Cas nuclease that can be activated by a complex comprising a targeter nucleic acid and a modulator nucleic acid. In other embodiments, the engineered non-naturally occurring system of the invention further comprises a Cas protein associated with a Cas nuclease that can be activated by a complex comprising a targeter nucleic acid and a modulator nucleic acid. For example, in some particular embodiments, the Cas protein comprises an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) identical to the Cas nuclease. In some particular embodiments, the Cas protein comprises a nuclease-inactive variant of the Cas nuclease. In some particular embodiments, the Cas protein further comprises an effector domain.
[0059] In some particular embodiments, the Cas protein has substantially no DNA cleavage activity. Such Cas proteins can be produced by introducing one or more mutations into an active Cas nuclease (e.g., a naturally occurring Cas nuclease). A mutant Cas protein has a DNA cleavage activity of about 25% or less, about 10% or less, about 5% or less, about 1% or less, about 0.1% or less, about 0.01% or less, or lower than that of the corresponding non-mutant form of the protein, or is considered to have substantially no DNA cleavage activity if it is zero or negligible compared to the non-mutant form. Thus, the Cas protein can contain one or more mutations (e.g., mutations within the RuvC domain of a V-A type Cas protein) and can be used as a general DNA binding protein with or without fusion to an effector domain. Exemplary mutations include D908A, E993A, and D1263A with respect to the amino acid positions in AsCpf1; D832A, E925A, and D1180A with respect to the amino acid positions in LbCpf1; and D917A, E1006A, and D1255A with respect to the amino acid position numbering of FnCpf1. More mutations can be designed and produced according to the crystal structures described in Yamano et al. (2016) Cell, 165:949.
[0060] It will be understood that the Cas protein does not lose the nuclease activity to cleave any DNA, but only loses the ability to cleave the target strand of double-stranded DNA or only the ability to cleave the non-target strand, and thereby may function as a nickase (see Gao et al. (2016) CELL RES., 26:901). Thus, in some particular embodiments, the Cas nuclease is a Cas nickase. In some particular embodiments, the Cas nuclease has the activity to cleave the non-target strand, but has substantially no activity to cleave the target strand, for example, due to mutations within the Nuc domain. In some particular embodiments, the Cas nuclease has the cleavage activity to cleave the target strand but has substantially no activity to cleave the non-target strand.
[0061] In other embodiments, the Cas nuclease has the activity of cleaving double-stranded DNA, resulting in double-strand breaks.
[0062] Cas proteins that have substantially no DNA cleavage activity or only the ability to cleave one strand may sometimes be identified from naturally occurring systems. For example, some specific naturally occurring CRISPR-Cas systems may retain the ability to bind to target nucleotide sequences in eukaryotic (e.g., mammalian or human) cells but may have lost all or some DNA cleavage activity. Such type V-A proteins are disclosed, for example, in Kim et al. (2017) ACS SYNTH. BIOL. 6(7):1273-82 and Zhang et al. (2017) CELL DISCOV. 3:17018.
[0063] The activity of a Cas protein (e.g., a Cas nuclease) can be modified, thereby generating an engineered Cas protein. In some particular embodiments, the modification of the activity of the engineered Cas protein includes an increase in targeting efficiency and / or a reduction in off-target binding. Without wishing to be bound by theory, it is hypothesized that off-target binding can be recognized by a Cas protein due to the presence of one or more mismatches between a spacer sequence and a target nucleotide sequence that can affect, for example, the stability and / or conformation of the CRISPR-Cas complex. In some particular embodiments, the modification of the activity includes a modification of binding, e.g., an increase in binding to a target locus (e.g., a target strand or a non-target strand) and / or a reduction in binding to an off-target locus. In some particular embodiments, the modification of the activity includes a modification of the charge of a region of the protein that binds to a single-guide nucleic acid or a dual-guide nucleic acid. In some particular embodiments, the modification of the activity of the engineered Cas protein includes a modification of the charge of a region of the protein that binds to a target strand and / or a non-target strand. In some particular embodiments, the modification of the activity of the engineered Cas protein includes a modification of the charge of a region of the protein that binds to an off-target locus. Modifications of the charge can include a decrease in positive charge, a decrease in negative charge, an increase in positive charge, and an increase in negative charge. For example, a decrease in negative charge and an increase in positive charge can generally enhance binding to nucleic acid(s), while a decrease in positive charge and an increase in negative charge can weaken binding to nucleic acid(s). In some particular embodiments, the modification of the activity includes an increase or a reduction in steric hindrance between the protein and a single-guide nucleic acid or a dual-guide nucleic acid. In some particular embodiments, the modification of the activity includes an increase or a reduction in steric hindrance between the protein and a target strand and / or a non-target strand. In some particular embodiments, the modification of the activity includes an increase or a reduction in steric hindrance between the protein and an off-target locus. In some particular embodiments, the modification or mutation includes a substitution of Lys, His, Arg, Glu, Asp, Ser, Gly, or Thr. In some particular embodiments, the modification or mutation includes a substitution with Gly, Ala, Ile, Glu, or Asp.In some particular embodiments, the modification or variation includes amino acid substitutions in the groove between the WED and RuvC domains of a Cas protein (e.g., a type V-A Cas protein).
[0064] In some particular embodiments, the modification of the activity of the engineered Cas protein includes an increase in nuclease activity that cleaves the target locus. In some particular embodiments, the modification of the activity of the engineered Cas protein includes a decrease in nuclease activity that cleaves off-target loci. In some particular embodiments, the modification of the activity of the engineered Cas protein includes a modification of helicase dynamics. In some particular embodiments, the engineered Cas protein includes a modification that modifies the formation of the CRISPR complex.
[0065] In some particular embodiments, the binding of the Cas protein complex to the target locus is directed by a protospacer adjacent motif (PAM) or a PAM-like motif. Many Cas proteins have PAM specificity. The precise sequence and length requirements of the PAM vary depending on the Cas protein used. The PAM sequence is typically 2 to 5 base pairs in length and is adjacent to the target nucleotide sequence (but on a different target DNA strand from the target nucleotide sequence). The PAM sequence can be identified by testing the cleavage, targeting, or modification of oligonucleotides having various PAM sequences with the target nucleotide sequence using methods known in the art.
[0066] Exemplary PAM sequences are shown in Table 1. In one aspect, the Cas protein is MAD7 and the PAM is TTTN, where N is A, C, G, or T. In one aspect, the Cas protein is MAD7 and the PAM is CTTN, where N is A, C, G, or T. In another aspect, the Cas protein is AsCpf1 and the PAM is TTTN, where N is A, C, G, or T. In another aspect, the Cas protein is FnCpf1 and the PAM is 5' TTN, where N is A, C, G, or T. The PAM sequences of some other specific type V-A Cas proteins are disclosed in Zetsche et al. (2015) CELL, 163:759 and U.S. Patent No. 9,982,279. Furthermore, programming of PAM specificity can be enabled by engineering of the PAM-interacting (PI) domain of the Cas protein, the fidelity of target site recognition can be improved, and the versatility of engineered non-natural systems can be increased. An exemplary approach for modifying the PAM specificity of Cpf1 is described in Gao et al. (2017) NAT. BIOTECHNOL., 35:789.
[0067] In some specific aspects, the engineered Cas protein comprises a modification that modifies the specificity of the Cas protein in a state that cooperates with a modification to the targeting range. Cas variants can be designed, for example, by selecting a mutation that modifies PAM specificity (e.g., within the PI domain) and combining the mutation with an in-groove mutation that increases (or decreases if desired) the specificity of the on-target site as compared to the off-target site, to have an increase in target specificity and adaptation in the modification of PAM recognition. The Cas modifications described herein can be used to abrogate a loss of specificity resulting from a modification of PAM recognition, to enhance a gain of specificity resulting from a modification of PAM recognition, to abrogate a gain of specificity resulting from a modification of PAM recognition, or to enhance a loss of specificity resulting from a modification of PAM recognition.
[0068] In some specific embodiments, the engineered Cas protein comprises one or more nuclear localization signal (NLS) motifs. In some specific embodiments, the engineered Cas protein comprises at least two (e.g., at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs. Non-limiting examples of NLS motifs include the NLS of SV40 large T antigen having the amino acid sequence of PKKKRKV (SEQ ID NO:23); the NLS derived from nucleoplasmin, e.g., the bipartite NLS of nucleoplasmin having the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO:24); the c-myc NLS having the amino acid sequence of PAAKRVKLD (SEQ ID NO:25) or RQRRNELKRSP (SEQ ID NO:26); the hRNPA1 M9 NLS having the amino acid sequence of TIFF0007689949000015.tif4141; the importin-α IBB domain NLS having the amino acid sequence of TIFF0007689949000016.tif4149; the myogenic T protein NLS having the amino acid sequence of VSRKRPRP (SEQ ID NO:29) or PPKKARED (SEQ ID NO:30); the human p53 NLS having the amino acid sequence of PQPKKKPL (SEQ ID NO:31); the mouse c-abl IV NLS having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO:32); the influenza virus NS1 NLS having the amino acid sequence of DRLRR (SEQ ID NO:33) or PKQKKRK (SEQ ID NO:34); the hepatitis delta antigen NLS having the amino acid sequence of RKLKKKIKKL (SEQ ID NO:35); the mouse Mx1 protein NLS having the amino acid sequence of REKKKFLKRR (SEQ ID NO:36); The human poly(ADP-ribose) polymerase NLS having the amino acid sequence of TIFF0007689949000017;RKCLQAGMNLEARKTKK (SEQ ID NO:38), the human glucocorticoid receptor NLS having the amino acid sequence of, and synthetic NLS motifs such as PAAKKKKLD (SEQ ID NO:39).
[0069] Generally, one or more NLS motifs are of sufficient strength to drive the accumulation of Cas protein in detectable amounts within the nucleus of eukaryotic cells. The strength of the nuclear localization activity can be derived from the number of NLS motif(s) (one or more) within the Cas protein, the specific NLS motif(s) (one or more) used, the position(s) of the NLS motif(s) (one or more), or a combination of these elements. In some specific embodiments, the engineered Cas protein comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10) NLS motifs at the N-terminus or near the N-terminus (e.g., within about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, about 30, about 40, about 50 or more amino acids from the N-terminus along the polypeptide chain). In some specific embodiments, the engineered Cas protein comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10) NLS motifs at the C-terminus or near the C-terminus (e.g., within about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, about 30, about 40, about 50 or more amino acids from the C-terminus along the polypeptide chain). In some specific embodiments, the engineered Cas protein comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10) NLS motifs at or near the C-terminus and at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10) NLS motifs at or near the N-terminus. In some specific embodiments, the engineered Cas protein comprises 1, 2, or 3 NLS motifs at or near the C-terminus.In some specific embodiments, the engineered Cas protein contains one NLS motif at or near the N-terminus and one, two, or three NLS motifs at or near the C-terminus. In some specific embodiments, the engineered Cas protein contains the nucleoplasmin NLS at or near the C-terminus.
[0070] Detection of accumulation in the nucleus can be performed by any suitable method. For example, a detectable marker may be fused to the nucleic acid targeting protein so that the location within the cell can be visualized. Also, the cell nucleus may be isolated from the cell and then its contents may be analyzed by any suitable method for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. Also, detection of accumulation in the nucleus can be indirect, for example, by an assay (e.g., an assay for DNA cleavage or mutation at the target locus or an assay for modification of gene expression activity) that measures the effect of nuclear import of the Cas protein complex by comparing it to a control not exposed to the Cas protein or a control exposed to a Cas protein lacking one or more of the NLS motifs.
[0071] The Cas proteins in the present invention may include chimeric Cas proteins, such as Cas proteins with improved function due to being chimeric. A chimeric Cas protein can be a new Cas protein that contains fragments derived from more than one naturally occurring Cas protein or variants thereof. For example, fragments of multiple type V-A Cas homologs (e.g., orthologs) can be fused to form a chimeric Cas protein. In some specific embodiments, the chimeric Cas protein contains fragments of Cpf1 orthologs from multiple species and / or strains.
[0072] In some specific embodiments, the Cas protein comprises one or more effector domains. The one or more effector domains may be located at or near the N-terminus and / or at or near the C-terminus of the Cas protein. In some specific embodiments, the effector domains comprised in the Cas protein are a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., KRAB domain or SID domain), a heterologous nuclease domain (e.g., FokI), a deaminase domain (e.g., cytidine deaminase or adenine deaminase) or a reverse transcriptase domain (e.g., high-fidelity reverse transcriptase domain). Other activities of the effector domains include, without limitation, methylase activity, demethylase activity, transcriptional terminator activity, translation initiation activity, translational activation activity, translational repression activity, histone modification (e.g., acetylation or demethylation) activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity.
[0073] In some specific embodiments, the Cas protein comprises one or more protein domains that enhance homologous sequence-dependent repair (HDR) and / or inhibit non-homologous end joining (NHEJ). Exemplary protein domains having such functions are described in Jayavaradhan et al. (2019) Nat. Commun. 10(1):2866 and Janssen et al. (2019) Mol. Ther. Nucleic Acids 16:141-54. In some specific embodiments, the Cas protein comprises a dominant negative variant of p53 binding protein 1 (53BP1), e.g., a fragment of 53BP1 comprising the minimum focus forming region (e.g., amino acids 1231-1644 of human 53BP1). In some specific embodiments, the Cas protein comprises a motif targeted by APC-Cdh1, e.g., amino acids 1-110 of human geminin, thereby resulting in degradation of the fusion protein in the HDR-incompetent G1 phase of the cell cycle.
[0074] In some specific embodiments, the Cas protein comprises an inducible domain or a regulatory domain. Non-limiting examples of inducers or regulators include light, hormones, and small molecule drugs. In some specific embodiments, the Cas protein comprises a light-inducible domain or a light-regulatory domain. In some specific embodiments, the Cas protein comprises a chemical-inducible domain or a chemical-regulatory domain.
[0075] In some specific embodiments, the Cas protein comprises a tag protein or a tag peptide to facilitate tracking or purification. Non-limiting examples of tag proteins and tag peptides include fluorescent proteins (e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato), HIS tags (e.g., 6×His tag), hemagglutinin (HA) tags, FLAG tags, and Myc tags.
[0076] In some specific embodiments, the Cas protein is conjugated to a non-protein moiety, such as a fluorophore useful for imaging of the genome. In some specific embodiments, the Cas protein is covalently conjugated to a non-protein moiety. The use of the terms "CRISPR-associated protein", "Cas protein", "Cas", "CRISPR-associated nuclease", and "Cas nuclease" herein encompasses such conjugates even if one or more non-protein moieties are present.
[0077] Targeter nucleic acid and modulator nucleic acid The engineered non-naturally occurring system of the present invention comprises a targeter nucleic acid and a modulator nucleic acid that can activate the Cas nuclease disclosed herein when hybridized to form a complex. In some specific embodiments, the Cas nuclease is activated by a single crRNA without tracrRNA in a naturally occurring system. In some specific embodiments, the Cas nuclease is a type V-A, V-C, or V-D nuclease.
[0078] As used herein, the term "targeter nucleic acid" refers to a nucleic acid comprising (i) a spacer sequence designed to hybridize to a target nucleotide sequence; and (ii) a targeter stem sequence capable of hybridizing to a further nucleic acid to form a complex, the complex being capable of activating a Cas nuclease (e.g., a type V-A Cas nuclease) under appropriate conditions, and the targeter nucleic acid alone, without the further nucleic acid, being unable to activate the Cas nuclease under the same conditions.
[0079] As used herein in the context of a given targeter nucleic acid and its corresponding Cas nuclease, the term "modulator nucleic acid" refers to a nucleic acid capable of hybridizing to the targeter nucleic acid to form a complex, the complex being capable of activating a Cas nuclease type under appropriate conditions, but the modulator nucleic acid alone being unable to effect activation.
[0080] As used in the definitions of "targeter nucleic acid" and "modulator nucleic acid", the term "appropriate conditions" refers to conditions under which a naturally occurring CRISPR-Cas system is functional, e.g., within a prokaryotic cell, within a eukaryotic (e.g., mammalian or human) cell, or in an in vitro assay.
[0081] The target nucleic acid and / or the modulator nucleic acid can be chemically synthesized or can be produced by biological methods (e.g., those catalyzed by RNA polymerase in an in vitro reaction). In such reactions or methods, the lengths of the target nucleic acid and the modulator nucleic acid can be limited. In some specific embodiments, the target nucleic acid has a nucleotide length of 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, or 25 or less. In some specific embodiments, the target nucleic acid has a nucleotide length of at least 20, 25, 30, 40, 50, 60, 70, 80, or 90. In some specific embodiments, the target nucleic acid has a nucleotide length of 20 - 100, 20 - 90, 20 - 80, 20 - 70, 20 - 60, 20 - 50, 20 - 40, 20 - 30, 20 - 25, 25 - 100, 25 - 90, 25 - 80, 25 - 70, 25 - 60, 25 - 50, 25 - 40, 25 - 30, 30 - 100, 30 - 90, 30 - 80, 30 - 70, 30 - 60, 30 - 50, 30 - 40, 40 - 100, 40 - 90, 40 - 80, 40 - 70, 40 - 60, 40 - 50, 50 - 100, 50 - 90, 50 - 80, 50 - 70, 50 - 60, 60 - 100, 60 - 90, 60 - 80, 60 - 70, 70 - 100, 70 - 90, 70 - 80, 80 - 100, 80 - 90, or 90 - 100. In some specific embodiments, the modulator nucleic acid has a nucleotide length of 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, or 20 or less. In some specific embodiments, the modulator nucleic acid has a nucleotide length of at least 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90.In certain embodiments, the modulator nucleic acid is 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 15 to 20, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 100, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 100, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 100, 60 to 90, 60 to 80, 60 to 70, 70 to 100, 70 to 90, 70 to 80, 80 to 100, 80 to 90 or 90 to 100 nucleotides in length.
[0082] In naturally occurring type V-A CRISPR-Cas systems, the crRNA contains a backbone sequence (also referred to as a direct repeat sequence) and a spacer sequence that hybridizes to a target nucleotide sequence. In some specific naturally occurring type V-A CRISPR-Cas systems, the backbone sequence forms a stem-loop structure, and the stem within the structure consists of five consecutive base pairs. The dual-guide type V-A CRISPR-Cas system can be derived from a naturally occurring type V-A CRISPR-Cas system or a variant thereof in which the Cas protein is guided to the target nucleotide sequence by the crRNA alone, and such a system is referred to herein as a "single-guide type V-A CRISPR-Cas system". In the dual-guide type V-A CRISPR-Cas system disclosed herein, the targeter nucleic acid contains the strand of the stem sequence between the spacer and the loop ("targeter stem sequence") and the spacer sequence, and the modulator nucleic acid contains the stem sequence ("modulator stem sequence") and the other strand of the 5'-tail located on the 5'-side of the modulator stem sequence. The targeter stem sequence is 100% complementary to the modulator stem sequence. Therefore, the double-stranded complex of the targeter nucleic acid and the modulator nucleic acid retains the orientation of the 5'-tail, the modulator stem sequence, the targeter stem sequence, and the spacer sequence of the single-guide type V-A CRISPR-Cas system, but there is no loop structure between the modulator stem sequence and the targeter stem sequence. A schematic diagram of an exemplary double-stranded complex is shown in FIG. 1.
[0083] Despite the similarity in general structure, it has been found that the stem-loop structure of the crRNA of the naturally occurring type V-A CRISPR complex is not essential for the functionality of the CRISPR system. This finding is surprising because prior art has suggested that the stem-loop structure is extremely important (see Zetsche et al. (2015) CELL, 163:759) and that the activity of the AsCpf1 CRISPR system is lost when the loop structure is removed by "splitting" the crRNA (see Li et al. (2017) NAT. BIOMED. ENG., 1:0066).
[0084] The length of this double-strand is expected to be a factor in the provision of a functional dual-guide CRISPR system. In some specific embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 10 nucleotides that base pair with each other. In some specific embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 nucleotides that base pair with each other. In some specific embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4, 5, 6, 7, 8, 9, or 10 nucleotides. It is understood that the nucleotide composition within each sequence affects the stability of the double-strand, and C-G base pairs confer greater stability than A-U base pairs. In some specific embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the base pairs are C-G base pairs.
[0085] In some specific embodiments, the targeter stem sequence and the modulator stem sequence each consist of 5 nucleotides. Therefore, the targeter stem sequence and the modulator stem sequence form a double-stranded structure of 5 base pairs. In some specific embodiments, 0 to 4, 0 to 3, 0 to 2, 0 to 1, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 5, 2 to 4, 2 to 3, 3 to 5, 3 to 4, or 4 to 5 of the 5 base pairs are C-G base pairs. In some specific embodiments, 0, 1, 2, 3, 4, or 5 of the 5 base pairs are C-G base pairs. In some specific embodiments, the targeter stem sequence consists of 5'-GUAGA-3' (SEQ ID NO:21), and the modulator stem sequence consists of 5'-UCUAC-3'. In some specific embodiments, the targeter stem sequence consists of 5'-GUGGG-3' (SEQ ID NO:22), and the modulator stem sequence consists of 5'-CCCAC-3'.
[0086] Also, it is contemplated that the duplex compatibility with a given Cas nuclease can also be an element in the provision of a functional dual-guide CRISPR system. For example, the targeter stem sequence and the modulator stem sequence can be derived from a naturally occurring crRNA that can activate the Cas nuclease without tracrRNA. In some specific embodiments, the nucleotide sequences of the targeter stem sequence and the modulator stem sequence are identical to the corresponding stem sequences of the stem-loop structure of such a naturally occurring crRNA.
[0087] In some specific embodiments, the targeter nucleic acid comprises a targeter stem sequence and a spacer sequence from 5' to 3'. The spacer sequence is designed to hybridize with the target nucleotide sequence. To provide sufficient targeting to the target nucleotide sequence, the spacer sequence generally has a length of 16 or more nucleotides. In some specific embodiments, the spacer sequence has a length of at least 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or 75 nucleotides. In some specific embodiments, the spacer sequence is shorter than or equal to 75, 50, 45, 40, 35, 30, 25, or 20 nucleotide lengths. Shorter spacer sequences may be desirable for reducing off-target events. Thus, in some specific embodiments, the spacer sequence is shorter than or equal to 19, 18, or 17 nucleotides. In some specific embodiments, the spacer sequence has a length of 17 - 30 nucleotides, such as 20 - 30 nucleotides, 20 - 25 nucleotides, 20 - 24 nucleotides, 20 - 23 nucleotides, 23 - 25 nucleotides, 20 - 22 nucleotides or about 20 nucleotides. In some specific embodiments, the spacer sequence has a length of 20 nucleotides. In some specific embodiments, the spacer sequence is at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% complementary to the target nucleotide sequence. In some specific embodiments, the spacer sequence is 100% complementary to the target nucleotide sequence in the seed region (about 5 base pairs proximal to PAM). In some specific embodiments, the spacer sequence is 100% complementary to the target nucleotide sequence.DNA cleavage has been reported to be less tolerant of mismatches between the spacer sequence and the target nucleotide sequence compared to DNA binding (see Klein et al. (2018) CELL REPORTS, 22:1413). Thus, in certain embodiments, when the engineered non-naturally occurring system comprises a Cas nuclease, the spacer sequence is 100% complementary to the target nucleotide sequence.
[0088] Proper design of the spacer sequence depends on the selection of the target nucleotide sequence. For example, to select a target nucleotide sequence within a particular gene in a given genome, sequence analysis can be performed to minimize the likelihood of hybridization between the spacer sequence and any other locus in the genome. Also, binding of the target nucleotide sequence having a PAM recognized by the Cas protein is considered in many design methods. In type V-A CRISPR-Cas systems, when the non-target strand (i.e., the strand that does not hybridize to the spacer sequence) is used as the coordinating strand, the PAM is present immediately upstream of the target sequence. Computer-based models such as those disclosed in Doench et al. (2016) NAT. BIOTECHNOL., 34:184; Chuai et al. (2018) GENOME BIOLOGY, 19:80; and Klein et al. (2018) CELL REPORTS, 22:1413 have been developed to evaluate the targeting potential of the target nucleotide sequence and any potential off-target effects. Computer-based methods are useful for the selection of spacer sequences, but it is generally prudent to design multiple spacer sequences and select one or more having high efficiency and specificity based on the results of in vitro and / or in vivo experiments.
[0089] In some specific embodiments, the 3' end of the targeter stem sequence is linked to the 5' end of the spacer sequence by 1 or fewer, 2 or fewer, 3 or fewer, 4 or fewer, 5 or fewer, 6 or fewer, 7 or fewer, 8 or fewer, 9 or fewer, or 10 or fewer nucleotides. In some specific embodiments, the targeter stem sequence and the spacer sequence are adjacent to each other and are directly linked by a nucleotide bond. In some specific embodiments, the targeter stem sequence and the spacer sequence are linked by 1 nucleotide, such as uridine. In some specific embodiments, the targeter stem sequence and the spacer sequence are linked by 2 or more nucleotides. In some specific embodiments, the targeter stem sequence and the spacer sequence are linked by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.
[0090] In some specific embodiments, the targeter nucleic acid further comprises an additional nucleotide sequence on the 5' side of the targeter stem sequence. In some specific embodiments, the additional nucleotide sequence comprises at least 1 (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides. In some specific embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In some specific embodiments, the additional nucleotide sequence consists of 2 nucleotides. In some specific embodiments, the additional nucleotide sequence closely resembles the loop of the crRNA of the corresponding single-guide CRISPR-Cas system or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 3' end of the loop). It will be understood that the additional nucleotide sequence on the 5' side of the targeter stem sequence is not essential. Thus, in some specific embodiments, the targeter nucleic acid does not contain any additional nucleotides on the 5' side of the targeter stem sequence.
[0091] In some specific embodiments, the targeter nucleic acid further comprises an additional nucleotide sequence comprising one or more nucleotides at its 3'-end that does not hybridize with the target nucleotide sequence. The additional nucleotide sequence may protect the targeter nucleic acid from degradation by 3'-5' exonucleases. In some specific embodiments, the additional nucleotide sequence has a length of 100 nucleotides or less. In some specific embodiments, the additional nucleotide sequence has a length of 90 nucleotides or less, 80 nucleotides or less, 70 nucleotides or less, 60 nucleotides or less, 50 nucleotides or less, 40 nucleotides or less, 30 nucleotides or less, 20 nucleotides or less, or 10 nucleotides or less. In some specific embodiments, the additional nucleotide sequence has a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides. In some specific embodiments, the additional nucleotide sequence has a length of 5-100, 5-50, 5-40, 5-30, 5-25, 5-20, 5-15, 5-10, 10-100, 10-50, 10-40, 10-30, 10-25, 10-20, 10-15, 15-100, 15-50, 15-40, 15-30, 15-25, 15-20, 20-100, 20-50, 20-40, 20-30, 20-25, 25-100, 25-50, 25-40, 25-30, 30-100, 30-50, 30-40, 40-100, 40-50, or 50-100 nucleotides.
[0092] In some specific embodiments, additional nucleotide sequences form a spacer sequence and a hairpin portion. Such secondary structure can increase the specificity of engineered non-natural systems (see Kocak et al. (2019) NAT. BIOTECH. 37:657-66). In some specific embodiments, the free energy change during hairpin formation is greater than -20 kcal / mol, greater than -15 kcal / mol, greater than -14 kcal / mol, greater than -13 kcal / mol, greater than -12 kcal / mol, greater than -11 kcal / mol, or greater than -10 kcal / mol, or equal to -20 kcal / mol, -15 kcal / mol, -14 kcal / mol, -13 kcal / mol, -12 kcal / mol, -11 kcal / mol, or -10 kcal / mol. In some specific embodiments, the free energy change during hairpin formation is greater than -5 kcal / mol, greater than -6 kcal / mol, greater than -7 kcal / mol, greater than -8 kcal / mol, greater than -9 kcal / mol, greater than -10 kcal / mol, greater than -11 kcal / mol, greater than -12 kcal / mol, greater than -13 kcal / mol, greater than -14 kcal / mol, or greater than -15 kcal / mol, or equal to -5 kcal / mol, -6 kcal / mol, -7 kcal / mol, -8 kcal / mol, -9 kcal / mol, -10 kcal / mol, -11 kcal / mol, -12 kcal / mol, -13 kcal / mol, -14 kcal / mol, or -15 kcal / mol.In some specific embodiments, the free energy change during hairpin formation is in the range of -20 to -10 kcal / mol, -20 to -11 kcal / mol, -20 to -12 kcal / mol, -20 to -13 kcal / mol, -20 to -14 kcal / mol, -20 to -15 kcal / mol, -15 to -10 kcal / mol, -15 to -11 kcal / mol, -15 to -12 kcal / mol, -15 to -13 kcal / mol, -15 to -14 kcal / mol, -14 to -10 kcal / mol, -14 to -11 kcal / mol, -14 to -12 kcal / mol, -14 to -13 kcal / mol, -13 to -10 kcal / mol, -13 to -11 kcal / mol, -13 to -12 kcal / mol, -12 to -10 kcal / mol, -12 to -11 kcal / mol or -11 to -10 kcal / mol. In other embodiments, the targeter nucleic acid does not contain any nucleotides on the 3'-side of the spacer sequence.
[0093] In some specific embodiments, the modulator nucleic acid further comprises an additional nucleotide sequence on the 3'-side of the modulator stem sequence. In some specific embodiments, the additional nucleotide sequence comprises at least 1 (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45 or at least 50) nucleotides. In some specific embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some specific embodiments, the additional nucleotide sequence consists of 1 nucleotide (e.g., uridine). In some specific embodiments, the additional nucleotide sequence consists of 2 nucleotides. In some specific embodiments, the additional nucleotide sequence closely resembles the loop of the crRNA of the corresponding single-guide CRISPR-Cas system or a fragment thereof (e.g., 1, 2, 3 or 4 nucleotides at the 5'-end of the loop). It will be understood that the additional nucleotide sequence on the 3'-side of the modulator stem sequence is not essential. Thus, in some specific embodiments, the modulator nucleic acid does not contain any additional nucleotides on the 3'-side of the modulator stem sequence.
[0094] It will be understood that if present, the additional nucleotide sequence on the 5' side of the targeter stem sequence and the additional nucleotide sequence on the 3' side of the modulator stem sequence may interact with each other. For example, the nucleotide immediately 5' of the targeter stem sequence and the nucleotide immediately 3' of the modulator stem sequence do not form Watson-Crick base pairs (ordinarily, these could each constitute part of the targeter stem sequence and part of the modulator stem sequence, respectively), but other nucleotides within the additional nucleotide sequence on the 5' side of the targeter stem sequence and the additional nucleotide sequence on the 3' side of the modulator stem sequence may form one, two, three, or more base pairs (e.g., Watson-Crick base pairs). Such interactions can affect the stability of the complex comprising the targeter nucleic acid and the modulator nucleic acid.
[0095] The stability of a complex comprising a targeter nucleic acid and a modulator nucleic acid can be evaluated by the Gibbs free energy change (ΔG) upon complex formation, either computationally or experimentally. If all putative base pairings of the complex occur between bases in the targeter nucleic acid and bases in the modulator nucleic acid, i.e., when there is no intrastrand secondary structure, the ΔG upon complex formation generally correlates with the ΔG upon formation of the secondary structure in the corresponding single-guide nucleic acid. Methods for calculating or measuring ΔG are known in the art. An exemplary method is RNAfold (rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) as disclosed in Gruber et al. (2008) NUCLEIC ACIDS RES., 36(Web Server issue):W70-W74. Unless otherwise specified, the ΔG values in the present disclosure are calculated by RNAfold for the formation of the secondary structure in the corresponding single-guide nucleic acid. In some specific embodiments, ΔG is less than or equal to -1 kcal / mol, e.g., less than or equal to -2 kcal / mol, less than or equal to -3 kcal / mol, less than or equal to -4 kcal / mol, less than or equal to -5 kcal / mol, less than or equal to -6 kcal / mol, less than or equal to -7 kcal / mol, less than or equal to -7.5 kcal / mol, or less than or equal to -8 kcal / mol. In some specific embodiments, ΔG is greater than or equal to -10 kcal / mol, e.g., greater than or equal to -9 kcal / mol, greater than or equal to -8.5 kcal / mol, or greater than or equal to -8 kcal / mol. In some specific embodiments, ΔG ranges from -10 to -4 kcal / mol.In some specific embodiments, ΔG ranges from -8 to -4 kcal / mol, -7 to -4 kcal / mol, -6 to -4 kcal / mol, -5 to -4 kcal / mol, -8 to -4.5 kcal / mol, -7 to -4.5 kcal / mol, -6 to -4.5 kcal / mol, or -5 to -4.5 kcal / mol. In some specific embodiments, ΔG is about -8 kcal / mol, -7 kcal / mol, -6 kcal / mol, -5 kcal / mol, -4.9 kcal / mol, -4.8 kcal / mol, -4.7 kcal / mol, -4.6 kcal / mol, -4.5 kcal / mol, -4.4 kcal / mol, -4.3 kcal / mol, -4.2 kcal / mol, -4.1 kcal / mol, or -4 kcal / mol.
[0096] It will be understood that ΔG can be affected by sequences in the target nucleic acid that are not within the target stem sequence and / or sequences in the modulator nucleic acid that are not within the modulator stem sequence. For example, one or more base pairs (e.g., Watson-Crick base pairs) between additional sequences on the 5' side of the target stem sequence and additional sequences on the 3' side of the modulator stem sequence can reduce ΔG, i.e., can stabilize the nucleic acid complex. In some specific embodiments, the nucleotide immediately 5' of the target stem sequence contains uracil or is uridine, and the nucleotide immediately 3' of the modulator stem sequence contains uracil or is uridine, thereby forming a non-conventional U-U base pair.
[0097] In some specific embodiments, the modulator nucleic acid contains a nucleotide sequence herein referred to as the "5' tail" that is located 5' of the modulator stem sequence. When the CRISPR system is a type V-A CRISPR system, the 5' tail of the dual-guide system closely resembles the nucleotide sequence located 5' of the stem-loop structure of the backbone sequence in the crRNA (single guide). Thus, the 5' tail can contain the corresponding nucleotide sequence when operating the dual-guide system from a single-guide system.
[0098] Although not bound by theory, it is contemplated that the 5' tail may be involved in the formation of the CRISPR-Cas complex. For example, in some specific embodiments, the 5' tail forms a modulator stem sequence and a pseudoknot structure, which is recognized by the Cas protein (see Yamano et al. (2016) CELL, 165:949). In some specific embodiments, the 5' tail is at least 3 (e.g., at least 4 or at least 5) nucleotides in length. In some specific embodiments, the 5' tail is 3, 4, or 5 nucleotides in length. In some specific embodiments, the nucleotide at the 3' end of the 5' tail contains uracil or is uridine. In some specific embodiments, the nucleotide at the second position counted from the 3' end within the 5' tail contains uracil or is uridine. In some specific embodiments, the nucleotide at the third position counted from the 3' end within the 5' tail contains adenine or is adenosine. This third nucleotide may form a base pair (e.g., a Watson-Crick base pair) with the nucleotide on the 5' side of the modulator stem sequence. Thus, in some specific embodiments, the modulator nucleic acid contains a uridine or uracil-containing nucleotide on the 5' side of the modulator stem sequence. In some specific embodiments, the 5' tail contains the nucleotide sequence 5'-AUU-3'. In some specific embodiments, the 5' tail contains the nucleotide sequence 5'-AAUU-3'. In some specific embodiments, the 5' tail contains the nucleotide sequence 5'-UAAUU-3'. In some specific embodiments, the 5' tail is located immediately 5' to the modulator stem sequence.
[0099] In some specific embodiments, the targeter nucleic acid and / or the modulator nucleic acid are designed such that the degree of secondary structure other than hybridization between the targeter stem sequence and the modulator stem sequence is reduced. In some specific embodiments, when optimally folded, no more than about 75%, about 50%, about 40%, about 30%, about 25%, about 20%, about 15%, about 10%, about 5%, about 1% or less of the nucleotides of the targeter nucleic acid and / or the modulator nucleic acid are involved in self-complementary base pairing. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum value of Gibbs free energy. One example of such an algorithm is mFold as described in Zuker and Stiegler (Nucleic Acids Res. 9(1981), 133-148). An example of another folding algorithm is the online web server RNAfold developed by the Institute for Theoretical Chemistry at the University of Vienna and using the centroid structure prediction algorithm (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62).
[0100] The targeter nucleic acid is directed to a specific target nucleotide sequence, and the donor template is designed to modify the target nucleotide sequence or a sequence in its vicinity. Thus, it will be understood that binding of the targeter nucleic acid or modulator nucleic acid to the donor template can increase editing efficiency and reduce off-target effects. In multiplex methods (e.g., as disclosed in the subsection "Multiplex Methods" of Section II below), binding of the donor template to the modulator nucleic acid allows for combining a targeter nucleic acid library with a donor template library and making the design of screening or selection assays more efficient and flexible. Thus, in some specific embodiments, the modulator nucleic acid further comprises a donor template recruitment sequence that can hybridize to the donor template (see Figure 1C). The donor template is described in the subsection "Donor Template" of Section II below. The donor template and the donor template recruitment sequence can be designed to have sequence complementarity. In some specific embodiments, the donor template recruitment sequence is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) complementary to at least a portion of the donor template. In some specific embodiments, the donor template recruitment sequence is 100% complementary to at least a portion of the donor template. In some specific embodiments, when the donor template contains a engineered sequence that is not homologous to the sequence to be repaired, the donor template recruitment sequence can hybridize to the engineered sequence within the donor template. In some specific embodiments, the donor template recruitment sequence is at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides in length. In some specific embodiments, the donor template recruitment sequence is located at the 5' end of the modulator nucleic acid.In some specific embodiments, the donor template recruit array is linked to the 5' tail or the modulator stem sequence of the modulator nucleic acid, if present, by internucleotide linkages or via a nucleotide linker.
[0101] In some specific embodiments, the modulator nucleic acid further comprises an editing enhancer sequence that enhances the efficiency of gene editing and / or homology-directed repair (HDR) (see Figure 1D). Exemplary editing enhancer sequences are described in Park et al. (2018) NAT. COMMUN. 9:3313. In some specific embodiments, the editing enhancer sequence is located on the 5' side of the 5' tail or the 5' side of the modulator stem sequence, if present. In some specific embodiments, the editing enhancer sequence is 1 to 50, 4 to 50, 9 to 50, 15 to 50, 25 to 50, 1 to 25, 4 to 25, 9 to 25, 15 to 25, 1 to 15, 4 to 15, 9 to 15, 1 to 9, 4 to 9, or 1 to 4 nucleotides in length. In some specific embodiments, the editing enhancer sequence is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or about 55 nucleotides in length. The editing enhancer sequence is designed to minimize homology with the target nucleotide sequence or any other sequence that the engineered non-natural system may contact, e.g., the genomic sequence of the cell into which the engineered non-natural system is delivered. In some specific embodiments, the editing enhancer is designed to minimize the presence of hairpin structures. The editing enhancer may include one or more chemical modifications disclosed herein.
[0102] The modulator nucleic acid and / or the targeter nucleic acid may further comprise a protective nucleotide sequence that suppresses or reduces degradation of the nucleic acid. In some specific embodiments, the protective nucleotide sequence is at least 5 (e.g., at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45 or at least 50) nucleotides in length. Depending on the length of the protective nucleotide sequence, the time for the exonuclease to reach the 5'-tail, the modulator stem sequence, the targeter stem sequence and / or the spacer sequence is prolonged, thereby protecting these portions of the modulator nucleic acid and / or the targeter nucleic acid from degradation by the exonuclease. In some specific embodiments, the protective nucleotide sequence forms a secondary structure, such as a hairpin portion or a tRNA structure, and the rate of degradation by the exonuclease is reduced (see, e.g., Wu et al. (2018) CELL. MOL. LIFE SCI., 75(19):3593-3607). The secondary structure can be predicted by methods known in the art, such as the online web server RNAfold developed at the University of Vienna and using the centroid structure prediction algorithm (see Gruber et al. (2008) NUCLEIC ACIDS RES., 36:W70). Also, nucleic acid degradation can be suppressed or reduced by some specific chemical modifications that may be present in the protective nucleotide sequence, as disclosed in the following subsection "Modification of RNA".
[0103] The protective nucleotide sequence is typically at the 5'-end, 3'-end or both ends of the modulator nucleic acid or the targeter nucleic acid. In some specific embodiments, the modulator nucleic acid comprises the protective nucleotide sequence at the 5'-end, optionally via a nucleotide linker (see Figure 1B). In some specific embodiments, the modulator nucleic acid comprises the protective nucleotide sequence at the 3'-end. In some specific embodiments, the modulator nucleic acid comprises the protective nucleotide sequence at the 5'-end. In some specific embodiments, the modulator nucleic acid comprises the protective nucleotide sequence at the 3'-end.
[0104] As described above, various nucleotide sequences, such as, but not limited to, donor template recruit sequences, editing enhancer sequences, protective nucleotide sequences, and linkers that link such sequences, if present, to the 5' tail or to the modulator stem sequence, may be present in the 5' portion of the modulator nucleic acid. It will be understood that the functions of donor template recruitment, editing enhancement, protection against degradation, and binding are not mutually exclusive, and that one nucleotide sequence may have one or more of such functions. For example, in some particular embodiments, the modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruit sequence and an editing enhancer sequence. In some particular embodiments, the modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruit sequence and a protective sequence. In some particular embodiments, the modulator nucleic acid comprises a nucleotide sequence that is both an editing enhancer sequence and a protective sequence. In some particular embodiments, the modulator nucleic acid comprises a nucleotide sequence that is a donor template recruit sequence, an editing enhancer sequence, and a protective sequence. In some particular embodiments, the nucleotide sequence on the 5' side of the 5' tail or the 5' side of the modulator stem sequence, if present, is 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 90, 60 to 80, 60 to 70, 70 to 90, 70 to 80, or 80 to 90 nucleotides in length.
[0105] In some particular embodiments, the engineered non - naturally occurring system further comprises one or more compounds (e.g., small molecule compounds) that enhance HDR and / or inhibit NHEJ. Exemplary compounds having such functions are described in Maruyama et al. (2015) NAT BIOTECHNOL. 33(5):538 - 42; Chu et al. (2015) NAT BIOTECHNOL. 33(5):543 - 48; Yu et al. (2015) CELL STEM CELL 16(2):142 - 47; Pinder et al. (2015) NUCLEIC ACIDS RES. 43(19):9379 - 92; and Yagiz et al. (2019) COMMUN. BIOL. 2:198. In some particular embodiments, the engineered non - naturally occurring system further comprises one or more compounds selected from the group consisting of DNA ligase IV antagonists (e.g., SCR7 compound, Ad4 E1B55K protein and Ad4 E4orf6 protein), RAD51 agonists (e.g., RS - 1), DNA - dependent protein kinase (DNA - PK) antagonists (e.g., NU7441 and KU0060648), β3 - adrenergic receptor agonists (e.g., L755507), inhibitors of intracellular protein transport from the ER to the Golgi apparatus (e.g., brefeldin A), and any combination thereof.
[0106] The sequences of the modulator nucleic acid and the targeter nucleic acid may be compatible with the Cas protein. Exemplary sequences functional with some particular type V - A Cas proteins are shown in Table 1. These sequences are merely examples, and it will be understood that other guide nucleic acid sequences may also be used with these Cas proteins.
[0107] (Table 1) Type V - A Cas proteins and corresponding guide nucleic acid sequences TIFF0007689949000018.tif226160 1 The amino acid sequences of the Cas proteins are shown at the end of this specification. 2It will be understood that the "modulator array" listed herein can constitute the nucleotide sequence of the modulator nucleic acid. Alternatively, additional nucleotide sequences may be included on the 5' and / or 3' sides of the "modulator array" listed herein within the modulator nucleic acid. 3 In the consensus PAM sequence, N represents A, C, G, or T. When "5'" is indicated before the PAM sequence, this means that when the non-target strand (i.e., the strand that does not hybridize with the spacer sequence) is used as the coordinate strand, the PAM is present immediately upstream of the target sequence.
[0108] In some particular embodiments, the target nucleic acid of the engineered non-naturally occurring system comprises the target stem sequence listed in Table 1. In some particular embodiments, the target nucleic acid and the modulator nucleic acid of the engineered non-naturally occurring system each comprise the target stem sequence and the modulator sequence listed in the same order in Table 1. In some particular embodiments, the engineered non-naturally occurring system further comprises a Cas nuclease comprising the amino acid sequence shown in SEQ ID NO listed in the same order in Table 1. In some particular embodiments, the engineered non-naturally occurring system is useful for targeting, editing, or modifying a nucleic acid comprising a target nucleotide sequence proximal or adjacent to (e.g., immediately downstream of) the PAM listed in the same order in Table 1 when the non-target strand (i.e., the strand that does not hybridize with the spacer sequence) is used as the coordinate strand.
[0109] In some particular embodiments, an engineered non-naturally occurring system is adjustable or inducible. For example, in some particular embodiments, the targeter nucleic acid, the modulator nucleic acid, and / or the Cas protein can be introduced at different times to the target nucleotide sequence, and the system can be active only when all components are present. In some particular embodiments, the amounts of the targeter nucleic acid, the modulator nucleic acid, and / or the Cas protein can be adjusted to obtain desired efficiency and specificity. In some particular embodiments, a nucleic acid containing a targeter stem sequence or a modulator stem sequence can be added to the system in excess, thereby dissociating the complex of the targeter nucleic acid and the modulator nucleic acid and putting the system in an off state.
[0110] Modification of RNA The targeter nucleic acid can be composed of DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. The modulator nucleic acid can be composed of DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In some particular embodiments, the targeter nucleic acid is RNA and the modulator nucleic acid is RNA. The targeter nucleic acid in the form of RNA is also referred to as targeter RNA, and the modulator nucleic acid in the form of RNA is also referred to as modulator RNA. The nucleotide sequences disclosed herein are shown as DNA sequences by including thymidine (T), and / or as RNA sequences including uridine (U). It will be understood that corresponding DNA sequences, RNA sequences, and DNA / RNA chimeric sequences are also envisioned. For example, when a spacer sequence is shown as a DNA sequence, the nucleic acid as RNA containing this spacer sequence can be derived from the DNA sequences disclosed herein by replacing each T with U. As a result, T and U are used interchangeably herein in the interpretation of the description of nucleotide sequences.
[0111] In some specific embodiments, the targeter nucleic acid and / or the modulator nucleic acid is an RNA having one or more modifications within the ribose group, one or more modifications within the phosphate group, one or more modifications within the nucleobase, one or more terminal modifications, or a combination thereof. Exemplary modifications are disclosed in US Patent Application Publication Nos. 2016 / 0289675, 2017 / 0355985, 2018 / 0119140, Watts et al. (2008) Drug Discov. Today 13:842-55, and Hendel et al. (2015) NAT. BIOTECHNOL. 33:985.
[0112] Modifications in the ribose group include, without limitation, modifications at the 2'-position or 4'-position. For example, in some specific embodiments, the ribose contains 2'-O-C1-C4 alkyl, such as 2'-O-methyl (2'-OMe). In some specific embodiments, the ribose contains 2'-O-C1-C3 alkyl-O-C1-C3 alkyl, such as 2'-O-(2-methoxyethyl) or 2'-methoxyethoxy (2'-O-CH 2 CH 2 OCH 3 ) as also known as 2'-MOE. In some specific embodiments, the ribose contains 2'-O-allyl. In some specific embodiments, the ribose contains 2'-O-2,4-dinitrophenol (DNP). In some specific embodiments, the ribose contains 2'-halo, such as 2'-F, 2'-Br, 2'-Cl, or 2'-I. In some specific embodiments, the ribose contains 2'-NH 2 . In some specific embodiments, the ribose contains 2'-H (e.g., deoxynucleotide). In some specific embodiments, the ribose contains 2'-arabinose or 2'-F-arabinose. In some specific embodiments, the ribose contains 2'-LNA or 2'-ULNA. In some specific embodiments, the ribose contains 4'-thioglucose.
[0113] Modifications to the phosphate group include, but are not limited to, phosphorothioate nucleotide linkages, chiral phosphorothioate nucleotide linkages, phosphorodithioate nucleotide linkages, boranophosphonate nucleotide linkages, C 1~4 alkylphosphonate nucleotide linkages, such as methylphosphonate nucleotide linkages, boranophosphonate nucleotide linkages, phosphonocarboxylate nucleotide linkages, such as phosphonoacetate nucleotide linkages, phosphonocarboxylate ester nucleotide linkages, such as phosphonoacetate ester nucleotide linkages, amide linkages, thiophosphonocarboxylate nucleotide linkages, such as thiophosphonoacetate nucleotide linkages, thiophosphonocarboxylate ester nucleotide linkages, such as thiophosphonoacetate ester nucleotide linkages, and phosphodiester linkers or 2',5'-linkages having any of the above linkers. Also included are various salts, mixed salts, and free acid forms.
[0114] Modifications in nucleobases include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine, 5-methyluracil, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dihydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil, 5-allylcytosine, 5-aminoallyluracil, 5-aminoallyl-cytosine, 5-bromouracil, 5-iodouracil, diaminopurine, difluorotoluene, dihydrouracil, abasic nucleotide, Z base, P base, unstructured nucleic acid, isoguanine, isocytosine (see Piccirilli et al. (1990) NATURE, 343:33), 5-methyl-2-pyrimidine (see Rappaport (1993) BIOCHEMISTRY, 32:3047), x(A, G, C, T) and y(A, G, C, T).
[0115] Examples of terminal modifications include, but are not limited to, polyethylene glycol (PEG), hydrocarbon linkers (e.g., heteroatom (O, S, N)-substituted hydrocarbon spacers; halo-substituted hydrocarbon spacers; keto-, carboxyl-, amide-, thionyl-, carbamoyl-, thionocarbamaoyl-containing hydrocarbon spacers), spermine linkers, dyes such as fluorescent dyes (e.g., fluorescein, rhodamine, cyanine), quenchers (e.g., dabsyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides and / or proteins). In some specific embodiments, the terminal modification includes conjugation (or ligation) of RNA with another molecule including an oligonucleotide (e.g., deoxyribonucleotide and / or ribonucleotide), peptide, protein, sugar, oligosaccharide, steroid, lipid, folic acid, vitamin and / or other molecules. In some specific embodiments, the terminal modification incorporated within the RNA is located via a linker incorporated as a phosphodiester bond, such as 2-(4-butylamidofluorescein) propane-1,3-diol bis(phosphodiester) linker, anywhere between two nucleotides within the RNA.
[0116] The modifications disclosed above can be incorporated into the target nucleic acid and / or modulator nucleic acid in the form of RNA. In some specific embodiments, the modification in RNA is selected from the group consisting of incorporation of 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-halo-3'-phosphorothioate (e.g., 2'-fluoro-3'-phosphorothioate), 2'-halo-3'-phosphonoacetate (e.g., 2'-fluoro-3'-phosphonoacetate) and 2'-halo-3'-thiophosphonoacetate (e.g., 2'-fluoro-3'-thiophosphonoacetate).
[0117] In some specific embodiments, the stability of RNA is modified by the modification. In some specific embodiments, the modification improves the stability of RNA, for example, by increasing the nuclease resistance of RNA compared to the corresponding unmodified RNA. Non-limiting examples of stability-improving modifications include 2'-O-methyl, 2'-O-C 1~4 alkyl, 2'-halo (e.g., 2'-F, 2'-Br, 2'-Cl, or 2'-I), 2'-MOE, 2'-O-C 1~3 alkyl-O-C 1~3 alkyl, 2'-NH 2 , 2'-H (or 2'-deoxy), 2'-arabino, 2'-F-arabino, 4'-thiophosphoribosyl sugar moiety, 3'-phosphorothioate, 3'-phosphonoacetate, 3'-thiophosphonoacetate, 3'-methylphosphonate, 3'-boranophosphate, 3'-phosphorodithioate, incorporation of locked nucleic acid (「LNA」) nucleotides containing a methylene bridge group between the 2' and 4' carbons of the ribose ring, and unlocked nucleic acid (「ULNA」) nucleotides. Such modifications are suitable for use as protecting groups to suppress or reduce the degradation of the 5’ tail, modulator stem sequence, targeter stem sequence, and / or spacer sequence (see the above subsection 「Targeter Nucleic Acid and Modulator Nucleic Acid」).
[0118] In some specific embodiments, the modification modifies the specificity of an engineered non-natural system. In some specific embodiments, the modification improves the specifications of an engineered non-natural system, for example, by improving on-target binding and / or cleavage or reducing off-target binding and / or cleavage or a combination thereof. Non-limiting examples of specificity-improving modifications include 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, and pseudouracil.
[0119] In some specific embodiments, the immunostimulatory effect of the RNA is modified by the modification as compared to the corresponding unmodified RNA. For example, in some specific embodiments, the modification results in a decrease in the ability of the RNA to activate TLR7, TLR8, TLR9, TLR3, RIG-I, and / or MDA5.
[0120] In some particular embodiments, the targeter nucleic acid and / or the modulator nucleic acid comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39 or at least 40 modified nucleotides. The modification(s) may be made at one or more positions within these nucleic acids such that the targeter nucleic acid and / or the modulator nucleic acid retains functionality. For example, the modified nucleic acid can still direct the Cas protein to the target nucleotide sequence and enable the Cas protein to exert its effector function. It will be understood that the specific modification(s) at a given position can be selected based on the functionality of the nucleotide at that position. For example, specificity-improving modifications may be suitable for nucleotides within the spacer sequence, the targeter stem sequence or the modulator stem sequence. Stability-improving modifications may be suitable for one or more terminal nucleotides within the targeter nucleic acid and / or the modulator nucleic acid. In some particular embodiments, at least 1 (e.g., at least 2, at least 3, at least 4 or at least 5) terminal nucleotides at the 5' end and / or at least 1 (e.g., at least 2, at least 3, at least 4 or at least 5) terminal nucleotides at the 3' end of the targeter nucleic acid are modified nucleotides.In some particular embodiments, five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides at the 5'-end of the targeter nucleic acid and / or five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides at the 3'-end are modified nucleotides. In some particular embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides at the 5'-end of the modulator nucleic acid and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides at the 3'-end are modified nucleotides. In some particular embodiments, five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides at the 5'-end of the modulator nucleic acid and / or five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides at the 3'-end are modified nucleotides. The selection of the modification positions is described in U.S. Patent Application Publication Nos. 2016 / 0289675 and 2017 / 0355985. When used in this paragraph, if the targeter nucleic acid or the modulator nucleic acid is a combination of DNA and RNA, the nucleic acid is considered as RNA as a whole, and the DNA nucleotide(s) is / are considered as RNA modification(s), e.g., 2'-H modification of ribose and optionally modification of the nucleobase.
[0121] The target nucleic acid and the modulator nucleic acid are not present within the same nucleic acid, i.e., their ends are not linked by conventional inter-nucleotide bonds, but they may be conjugated to each other by covalent bonds through one or more chemical modifications introduced into these nucleic acids, whereby the stability of the double-stranded complex can be enhanced and / or other characteristics of the system can be improved.
[0122] II. Methods for targeting, editing and / or modifying genomic DNA The engineered non-naturally occurring systems disclosed herein are useful for targeting, editing and / or modifying target nucleic acids, such as DNA (e.g., genomic DNA), within cells or organisms. Thus, in one aspect, the present invention provides a method for modifying a target nucleic acid (e.g., DNA) having a target nucleotide sequence, the method comprising contacting the target nucleic acid with an engineered non-naturally occurring system disclosed herein, thereby effecting modification of the target nucleic acid.
[0123] An engineered non-naturally occurring system may be contacted with a target nucleic acid as a complex. Thus, in some particular embodiments, the method comprises contacting a target nucleic acid with (a) a targeter nucleic acid comprising (i) a spacer sequence designed to hybridize to a target nucleotide sequence and (ii) a targeter stem sequence; (b) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence; and (c) a dual-guide CRISPR-Cas complex comprising a Cas protein, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and the targeter nucleic acid and the modulator nucleic acid form a complex capable of activating a Cas nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system, thereby effecting modification of the target nucleic acid. In some particular embodiments, the Cas protein comprises an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) identical to a Cas nuclease.
[0124] The Cas protein and the Cas nuclease may be the same. Thus, in some particular embodiments, the present invention provides a method for cleaving a target nucleic acid (e.g., DNA) having a target nucleotide sequence, the method comprising contacting the target nucleic acid with an engineered non-naturally occurring system disclosed herein, thereby resulting in cleavage of the target DNA. In some particular embodiments, the method comprises contacting the target nucleic acid with: (a) a targeter nucleic acid comprising (i) a spacer sequence designed to hybridize to the target nucleotide sequence and (ii) a targeter stem sequence; (b) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence; and (c) a dual-guide CRISPR-Cas complex comprising a Cas nuclease, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and in a naturally occurring system, the Cas nuclease is activated by a single crRNA without tracrRNA, thereby resulting in cleavage of the target nucleic acid by the Cas nuclease.
[0125] In some particular embodiments, the Cas nuclease is a type V-A, V-C or V-D Cas nuclease. In some particular embodiments, the Cas nuclease is a type V-A Cas nuclease. In some particular embodiments, the target nucleic acid further comprises a cognate PAM positioned such that (a) the dual-guide CRISPR-Cas complex binds to the target nucleic acid; or (b) the Cas nuclease is activated when the dual-guide CRISPR-Cas complex binds to the target nucleic acid, with respect to the target nucleotide sequence.
[0126] The dual-guide CRISPR-Cas complex can be delivered to cells by introducing a pre-formed ribonucleoprotein (RNP) complex into the cells. Alternatively, one or more components of the dual-guide CRISPR-Cas complex may be expressed intracellularly. Exemplary delivery methods are known in the art and are described, for example, in U.S. Patent Nos. 10,113,167 and 8,697,359 and U.S. Patent Application Publications 2015 / 0344912, 2018 / 0044700, 2018 / 0003696, 2018 / 0119140, 2017 / 0107539, 2018 / 0282763, and 2018 / 0363009.
[0127] It will be understood that contacting intracellular DNA (e.g., genomic DNA) with the dual-guide CRISPR-Cas complex does not require delivery of all components of the complex into the cell. For example, one or more of the components may already be present intracellularly. In some particular embodiments, the cell (or its parental / ancestral cell) is engineered to express a Cas protein, and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) and a modulator nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the modulator nucleic acid) are delivered into the cell. In some particular embodiments, the cell (or its parental / ancestral cell) is engineered to express a modulator nucleic acid, and a Cas protein (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the Cas protein) and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) are delivered into the cell. In some particular embodiments, the cell (or its parental / ancestral cell) is engineered to express a Cas protein, and a modulator nucleic acid and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) are delivered into the cell.
[0128] In some specific embodiments, the target DNA is present within the genome of the target cell. Thus, in another aspect, the present invention provides cells comprising a non-naturally occurring system or the CRISPR expression system described herein. Further, the present invention provides cells whose genome has been modified by the dual-guide CRISPR-Cas system or complex disclosed herein.
[0129] The target cell can be a mitotic or post-mitotic cell from any organism, such as a bacterial cell, an archaeal cell, a cell of a unicellular eukaryote, a plant cell, an algal cell such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, etc., a fungal cell (e.g., a yeast cell), an animal cell, a cell derived from an invertebrate (e.g., Drosophila, an enidarian, an echinoderm, a nematode, etc.), a cell derived from a vertebrate (e.g., a fish, an amphibian, a reptile, a bird, a mammal), a cell derived from a mammal, a cell derived from a rodent, or a cell derived from a human. Non-limiting examples of the type of target cell include stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells), somatic cells (e.g., fibroblasts, hematopoietic cells, T lymphocytes (e.g., CD8 +T lymphocytes, NK cells, nerve cells, muscle cells, bone cells, hepatocytes, pancreatic cells), in vitro or in vivo embryonic cells at any stage of embryo (e.g., 1-cell stage, 2-cell stage, 4-cell stage, 8-cell stage; zebrafish embryos at the stage), etc. are included. The cells may be derived from an established cell line or may be primary cells (i.e., cells and cell cultures that are derived from a subject and have been proliferated by subculture a limited number of times in vitro). For example, the primary culture may be a culture that has been subcultured 0 times or less, 1 time or less, 2 times or less, 4 times or less, 5 times or less, 10 times or less, or 15 times or less but not enough to reach the crisis stage. Typically, the primary cell line of the present invention is maintained by subculture less than 10 times in vitro. When the cells are primary cells, they can be collected from an individual by any suitable method. For example, white blood cells can be collected by apheresis, leukapheresis, or density gradient separation, while cells derived from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, or stomach can be collected by biopsy. The collected cells may be used immediately, stored with a cryoprotectant under cryogenic conditions, and thawed in a manner generally known in the art at a later time.
[0130] Delivery of ribonucleoprotein (RNP) and delivery of "Cas RNA" The engineered non-naturally occurring systems disclosed herein can be delivered into cells by suitable methods known in the art, such as, but not limited to, the delivery of ribonucleoprotein (RNP) and the delivery of "Cas RNA" described below.
[0131] In some specific embodiments, a dual-guide CRISPR-Cas system comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein is integrated into an RNP complex, which can then be delivered into cells as a pre-formed complex. This method is suitable for the active modification of a cell's genetic or epigenetic information for a limited period. For example, if the Cas protein has nuclease activity for modifying the genomic DNA of a cell, this nuclease activity only needs to be retained for a period to enable DNA cleavage, and long-term nuclease activity can enhance off-target effects. Similarly, some specific epigenetic modifications can be maintained in cells once established and inherited by daughter cells.
[0132] "Ribonucleoprotein" or "RNP", as used herein, refers to a complex comprising a nucleoprotein and a ribonucleic acid. "Nucleoprotein", as indicated herein, refers to a protein that can bind to a nucleic acid (e.g., RNA, DNA). When a nucleoprotein binds to a ribonucleic acid, this is referred to as a "ribonucleoprotein". The interaction between a ribonucleoprotein and a ribonucleic acid may be direct, for example, by a covalent bond, or indirect, for example, by a non-covalent bond (e.g., electrostatic interactions (e.g., ionic bonds, hydrogen bonds, halogen bonds), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion forces), ring stacking (π effects), hydrophobic interactions, etc.). In some specific embodiments, the ribonucleoprotein comprises an RNA-binding motif that binds to the ribonucleic acid by a non-covalent bond. For example, an aromatic amino acid residue with a positive charge within the RNA-binding motif (e.g., a lysine residue) can form an electrostatic interaction with the negatively charged nucleic acid phosphate backbone of the RNA.
[0133] To ensure efficient loading of the Cas protein, the targeter nucleic acid and the modulator nucleic acid can be supplied in a molar excess (e.g., about 2-fold, about 3-fold, about 4-fold or about 5-fold) relative to the Cas protein. In some particular embodiments, the targeter nucleic acid and the modulator nucleic acid are annealed under appropriate conditions prior to complex formation with the Cas protein. In other embodiments, the targeter nucleic acid, the modulator nucleic acid and the Cas protein are directly mixed together to form the RNP.
[0134] The RNPs disclosed herein can be introduced into cells using a variety of delivery methods. Exemplary delivery methods or delivery vehicles include, but are not limited to, microinjection, liposomes (see, e.g., U.S. Patent Application Publication No. 2017 / 0107539), molecular trojan horse liposomes that deliver molecules across the blood-brain barrier (see Pardridge et al. (2010) COLD SPRING HARB. PROTOC., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, microvesicles (e.g., exosomes and ARMM), polycations, lipid:nucleic acid conjugates, electroporation, cell-penetrating peptides (see U.S. Patent Application Publication No. 2018 / 0363009), nanoparticles, nanowires (see Shalek et al. (2012) NANO LETTERS, 12:6498), exosomes, and perturbation of the cell membrane (e.g., by passing cells through constrictions in a microfluidic system, see U.S. Patent Application Publication No. 2018 / 0003696). When the target cells are proliferating cells, the efficiency of RNP delivery can be improved by cell cycle synchronization (see U.S. Patent Application Publication No. 2018 / 0044700).
[0135] In other aspects, the dual-guide CRISPR-Cas system is delivered into cells by the "Cas RNA" approach, i.e., by delivering RNA (e.g., messenger RNA (mRNA)) encoding a target nucleic acid, a modulator nucleic acid, and a Cas protein. The RNA encoding the Cas protein is translated intracellularly and can form a complex with the target nucleic acid and the modulator nucleic acid intracellularly. Similar to the RNP approach, the RNA has a limited half-life intracellularly even if one or more stability-enhancing modifications (s) are made to one or more of the RNAs. Thus, the "Cas RNA" approach is suitable for the active modification of cellular genetic or epigenetic information, such as DNA cleavage, for a limited period of time and has the advantage of reduced off-target effects.
[0136] mRNA can be produced by transcription of DNA containing regulatory elements operably linked to a Cas coding sequence. Considering that multiple copies of the Cas protein can result from one mRNA, the target nucleic acid and the modulator nucleic acid can generally be supplied in a molar excess (e.g., at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold or at least 100-fold) relative to the mRNA. In some particular aspects, the target nucleic acid and the modulator nucleic acid are annealed under appropriate conditions prior to delivery into cells. In other aspects, the target nucleic acid and the modulator nucleic acid are delivered into cells without annealing in vitro.
[0137] The "Cas RNA" system can be introduced into cells using various delivery systems. Non-limiting examples of delivery methods or delivery vehicles include microinjection, biolistic particles, liposomes (see, e.g., U.S. Patent Application Publication No. 2017 / 0107539), molecular Trojan horse liposomes that deliver molecules across the blood-brain barrier (see Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, polycations, lipid:nucleic acid conjugates, electroporation, nanoparticles, nanowires (see Shalek et al. (2012) NANO LETTERS, 12:6498), exosomes, and perturbation of the cell membrane (e.g., by passing cells through constrictions in a microfluidic system, see U.S. Patent Application Publication No. 2018 / 0003696). Specific examples of the "nucleic acid only" approach by electroporation are described in International (PCT) Publication No. 2016 / 164356.
[0138] In other embodiments, a dual-guide CRISPR-Cas system is delivered into cells in the form of DNA that includes a targeter nucleic acid, a modulator nucleic acid, and a regulatory element operably linked to a Cas coding sequence. The DNA can be provided in the form of a plasmid, a viral vector, or any other form described in the subsection "CRISPR expression systems". In such delivery methods, constitutive expression of the Cas protein may be achieved in the target cells (e.g., if the DNA is maintained as an episomal vector within the cell or integrated into the genome), and when the Cas protein has nuclease activity, the risk of unwanted off-target effects can be high. Nevertheless, this approach can be useful when the Cas protein includes a non-nuclease effector (e.g., a transcriptional activator or a transcriptional repressor). It is also useful for research purposes and for genome editing of plants.
[0139] CRISPR expression systems In another aspect, the present invention provides a CRISPR expression system comprising: a nucleic acid comprising a first regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid disclosed herein, the targeter nucleic acid comprising (a)(i) a spacer sequence designed to hybridize to a target nucleotide sequence and (ii) a targeter stem sequence; and a nucleic acid comprising a second regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid disclosed herein, the modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence, wherein the targeter nucleic acid and the modulator nucleic acid are expressed as separate nucleic acids, and the complex comprising the targeter nucleic acid and the modulator nucleic acid can activate a Cas nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system.
[0140] In some specific embodiments, the CRISPR expression system further comprises: a nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein disclosed herein. In some specific embodiments, the Cas protein comprises an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) identical to a Cas nuclease, thereby resulting in modification of a target nucleic acid (e.g., DNA). In some specific embodiments, the Cas protein is identical to the Cas nuclease, and cleavage of the target nucleic acid is thereby effected. In some specific embodiments, the Cas nuclease is a type V-A, type V-C or type V-D Cas nuclease. In some specific embodiments, the Cas nuclease is a type V-A Cas nuclease.
[0141] As used herein, the term "functionally linked" is intended to mean that the nucleotide sequence of interest is linked to a regulatory element in such a way as to permit expression of the nucleotide sequence (e.g., within an in vitro transcription / translation system or, if a vector is introduced into a host cell, within the host cell).
[0142] The forms of elements (a), (b), and (c) of the CRISPR expression system described above can be independently selected from a variety of nucleic acids, such as DNA (e.g., modified DNA) and RNA (e.g., modified RNA). In some particular embodiments, elements (a) and (b) are each in the form of DNA. In some particular embodiments, the CRISPR expression system further comprises an element (c) in the form of DNA. The third regulatory element can be a constitutive promoter or an inducible promoter that drives expression of the Cas protein. In other embodiments, the CRISPR expression system further comprises an element (c) in the form of RNA (e.g., mRNA).
[0143] Element (a), (b) and / or (c) may be supplied by one or more vectors. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Nucleic acids can be introduced into cells, such as prokaryotic cells, eukaryotic cells, mammalian cells or target tissues, using conventional virus-based and non-virus-based gene delivery methods. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and delivery vehicles, such as nucleic acids complexed with liposomes. Virus vector delivery systems include DNA viruses and RNA viruses that either become episomal or integrate into the genome after delivery to the cell. Gene therapy procedures are known in the art and are disclosed in Van Brunt (1988) BIOTECHNOLOGY, 6:1149; Anderson (1992) SCIENCE, 256:808; Nabel & Feigner (1993) TIBTECH, 11:211; Mitani & Caskey (1993) TIBTECH, 11:162; Dillon (1993) TIBTECH, 11:167; Miller (1992) NATURE, 357:455; Vigne, (1995) RESTORATIVE NEUROLOGY AND NEUROSCIENCE, 8:35; Kremer & Perricaudet (1995) BRITISH MEDICAL BULLETIN, 51:31; Haddada et al. (1995) CURRENT TOPICS IN MICROBIOLOGY AND IMMUNOLOGY, 199:297; Yu et al. (1994) GENE THERAPY, 1:13; and Doerfler and Bohm (Eds.) (2012) The Molecular Repertoire of Adenoviruses II: Molecular Biology of Virus-Cell Interactions. In some specific embodiments, at least one of the vectors is a DNA plasmid.In certain embodiments, at least one of the vectors is a viral vector (e.g., a retrovirus, an adenovirus, or an adeno-associated virus).
[0144] Certain vectors can autonomously replicate within the host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors and replication-defective viral vectors) do not autonomously replicate within the host cell. However, certain vectors may be integrated into the genome of the host cell and thereby replicated with the host genome. Those skilled in the art will recognize that a variety of vectors may be suitable for a variety of delivery methods, may have a variety of host tropisms, and that one or more vectors suitable for use can be selected.
[0145] As used herein, the term "regulatory element" refers to transcriptional control sequences and / or translational control sequences, such as promoters, enhancers, transcriptional termination signals (e.g., polyadenylation signals), internal ribosome entry sites (IRES), proteolytic signals, etc., that result in and / or regulate the transcription of non-coding sequences (e.g., targeter nucleic acids or modulator nucleic acids) or coding sequences (e.g., Cas proteins), and / or regulate the translation of the encoded polypeptides. Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY, 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of nucleotide sequences in many types of host cells and those that direct expression of nucleotide sequences only in some specific host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in the desired tissue of interest, such as muscle, nerve cells, bone, skin, blood, a particular organ (e.g., liver, pancreas) or a particular cell type (e.g., lymphocytes). Also, regulatory elements can direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, and the manner can also be tissue-type specific or cell-type specific or not. In some specific embodiments, the vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5 or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5 or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5 or more pol I promoters) or a combination thereof. Examples of pol III promoters include, but are not limited to, the U6 promoter and the H1 promoter.Examples of pol II promoters include, without limitation, the Rous sarcoma virus (RSV) LTR promoter of retrovirus (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. Also included in the term "regulatory element" are enhancer elements such as WPRE; CMV enhancer; R-U5' segment within the LTR of HTLV-I (see Takebe et al. (1988) MOL. CELL. BIOL., 8:466); SV40 enhancer; and the intron sequence between exon 2 and exon 3 of rabbit β-globin (see O'Hare et al. (1981) PROC. NATL. ACAD. SCI. USA., 78:1527). Those skilled in the art will recognize that the design of the expression vector can depend on factors such as, for example, the choice of host cell to be transformed, the desired expression level, etc. When the vector is introduced into the host cell, transcripts, proteins or peptides encoded by the nucleic acids as described herein, such as fusion proteins or fusion peptides (e.g., CRISPR transcripts, proteins, enzymes, their variant forms or their fusion proteins) can be produced.
[0146] In some specific embodiments, the nucleotide sequence encoding the Cas protein is codon-optimized for expression in a eukaryotic host cell, such as a yeast cell, a mammalian cell (e.g., a mouse cell, a rat cell or a human cell) or a plant cell. Different species exhibit a specific bias for certain codons of some specific amino acids. Codon bias (the difference in codon usage frequency among organisms) often correlates with the translation efficiency of messenger RNA (mRNA), and furthermore, this translation efficiency is considered to be dependent, inter alia, on the properties of the codons to be translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of tRNA selected within a cell is generally a reflection of the codons most frequently used in peptide synthesis. Thus, genes can be individually adjusted based on codon optimization for optimal gene expression in a given organism. A table of codon usage frequencies is readily available, for example, in the "Codon Usage Database" available at kazusa.or.jp / codon / , and this table can be adapted in several ways (see Nakamura et al. (2000) NUCL. ACIDS RES., 28:292). Also available are computer algorithms, such as Gene Forge (Aptagen; Jacobus, Pa.), for codon-optimizing a specific sequence for expression in a specific host cell. In some specific embodiments, codon optimization promotes or improves the expression of the Cas protein in the host cell.
[0147] Donor template The DNA damage pathway can be activated by cleavage of a target nucleotide sequence within the genome of a cell by the dual-guide CRISPR-Cas system or complex disclosed herein, whereby the cleaved DNA fragment can be religated by NHEJ or HDR. HDR requires either an endogenous or exogenous repair template for transferring sequence information from the repair template to the target.
[0148] In some specific embodiments, an engineered non-naturally occurring system or a CRISPR expression system further comprises a donor template. As used herein, the term "donor template" refers to a nucleic acid designed to function as a repair template in or near a target nucleotide sequence when introduced into a cell or an organism. In some specific embodiments, the donor template is complementary to a polynucleotide comprising the target nucleotide sequence or a portion thereof. When optimally aligned, the donor template may overlap with one or more nucleotides of the target nucleotide sequence (e.g., about 1, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40 or more, or more than about 1, more than about 5, more than about 10, more than about 15, more than about 20, more than about 25, more than about 30, more than about 35, more than about 40, or more than that number of nucleotides). The nucleotide sequence of the donor template is typically not identical to the genomic sequence to be replaced. Instead, the donor template may contain one or more substitutions, insertions, deletions, inversions or rearrangements relative to the genomic sequence, as long as there is sufficient homology to support homology-directed repair. In some specific embodiments, the donor template contains non-homologous sequences flanked by two regions of homology (i.e., homology arms), such that homology-directed repair between the target DNA region and the two flanking sequences results in the insertion of non-homologous sequences into the target region. In some specific embodiments, the donor template contains non-homologous sequences that are 10 to 100 nucleotides in length, 50 to 500 nucleotides in length, 100 to 1,000 nucleotides in length, 200 to 2,000 nucleotides in length or 500 to 5,000 nucleotides in length, located between the two homology arms.
[0149] Generally, the homologous region(s) of the donor template have at least 50% sequence identity with the genomic sequence where recombination is desired. The homology arms are designed or selected to be able to perform recombination with the nucleotide sequences flanking the target nucleotide sequence under intracellular conditions. In some specific embodiments, when HDR of the non-target strand is desired, the donor template includes a first homology arm homologous to the sequence on the 5' side of the target nucleotide sequence and a second homology arm homologous to the sequence on the 3' side of the target nucleotide sequence. In some specific embodiments, the first homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to the sequence on the 5' side of the target nucleotide sequence. In some specific embodiments, the second homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to the sequence on the 3' side of the target nucleotide sequence. In some specific embodiments, when the polynucleotide comprising the donor template sequence and the target nucleotide sequence is optimally aligned, the closest nucleotide of the donor template is within about 1 nucleotide, within about 5 nucleotides, within about 10 nucleotides, within about 15 nucleotides, within about 20 nucleotides, within about 25 nucleotides, within about 50 nucleotides, within about 75 nucleotides, within about 100 nucleotides, within about 200 nucleotides, within about 300 nucleotides, within about 400 nucleotides, within about 500 nucleotides, within about 1000 nucleotides, within about 2000 nucleotides, within about 3000 nucleotides, within about 4000 nucleotides or more nucleotides from the target nucleotide sequence.
[0150] In some particular embodiments, the donor template further includes an engineered array that is not homologous to the array to be repaired. Such an engineered array may have a barcode and / or an array that can hybridize with the donor template recruitment array disclosed herein.
[0151] In some particular embodiments, the donor template further includes one or more mutations to the genomic sequence, and the one or more mutations reduce or suppress the cleavage of the donor template by the same CRISPR-Cas system or the cleavage of the modified genomic sequence into which at least a portion of the donor template sequence is incorporated. In some particular embodiments, in the donor template, the PAM that is adjacent to the target nucleotide sequence and recognized by the Cas nuclease is mutated to a sequence not recognized by the same Cas nuclease. In some particular embodiments, in the donor template, the target nucleotide sequence (e.g., the seed region) is mutated. In some particular embodiments, the one or more mutations are silent with respect to the reading frame of the protein-coding sequence containing the mutation site.
[0152] The donor template can be supplied to the cell as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. It will be understood that the dual-guide CRISPR-Cas system disclosed herein can have nuclease activity that cleaves the target strand, the non-target strand, or both. When HDR of the target strand is desired, a donor template having a nucleic acid sequence complementary to the target strand is also envisioned.
[0153] Donor templates can be introduced into cells in linear or circular form. When introduced in linear form, the ends of the donor template can be protected by methods known to those of skill in the art (e.g., from degradation by exonucleases). For example, one or more dideoxynucleotide residues are added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides are ligated to one or both ends (see, e.g., Chang et al. (1987) PROC. NATL. ACAD SCI USA, 84:4959; Nehls et al. (1996) SCIENCE, 272:886; see also chemical modifications to enhance the stability and / or specificity of RNA disclosed above). Additional methods for protecting exogenous polynucleotides from degradation include, without limitation, the addition of terminal amino group(s) and the use of modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues. Instead of protecting the ends of the linear donor template, additional lengths of sequence that can be degraded without affecting recombination may be included outside the region of homology.
[0154] The donor template may be a component of the vectors described herein, may be included in a separate vector, or may be provided as a separate polynucleotide, such as an oligonucleotide, linear polynucleotide, or synthetic polynucleotide. In some particular embodiments, the donor template is DNA. In some particular embodiments, the donor template, where applicable, is present within the same nucleic acid as the sequence encoding the targeter nucleic acid, the sequence encoding the modulator nucleic acid, and / or the sequence encoding the Cas protein. In some particular embodiments, the donor template is provided as a separate nucleic acid. The donor template polynucleotide can be of any suitable length, such as about 50, about 75, about 100, about 150, about 200, about 500, about 1000, about 2000, about 3000, about 4000 or more, or at least about 50, at least about 75, at least about 100, at least about 150, at least about 200, at least about 500, at least about 1000, at least about 2000, at least about 3000, at least about 4000 or more nucleotide lengths.
[0155] The donor template can be introduced into cells as an isolated nucleic acid. Alternatively, the donor template can be introduced into cells as part of a vector (e.g., a plasmid) having additional sequences not intended for insertion into the DNA region of interest, such as an origin of replication, a promoter, and a gene encoding antibiotic resistance. Alternatively, the donor template can be delivered by a virus (e.g., an adenovirus, an adeno-associated virus (AAV)). In some particular embodiments, the donor template is introduced as AAV, such as pseudotyped AAV. The capsid protein of AAV can be selected by one skilled in the art based on the tropism of AAV and the target cell type. For example, in some particular embodiments, the donor template is introduced as AAV8 or AAV9 into hepatocytes. In some particular embodiments, the donor template is a hematopoietic stem cell, a hematopoietic progenitor cell, or a T lymphocyte (e.g., CD8 +It is introduced into T lymphocytes) as AAV6 or AAVHSC (see U.S. Patent No. 9,890,396). The sequence of the capsid protein (VP1, VP2 or VP3) may be modified from the wild-type AAV capsid protein, for example, having at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity with the wild-type AAV capsid sequence.
[0156] Donor templates can be delivered to cells (e.g., primary cells) by various delivery methods, such as the viral or non-viral methods disclosed herein. In some specific embodiments, non-viral donor templates are introduced into target cells as naked nucleic acids or as complexes with liposomes or poloxamers. In some specific embodiments, non-viral donor templates are introduced into target cells by electroporation. In other embodiments, viral donor templates are introduced into target cells by infection. Engineered non-natural systems can be delivered before, after, or simultaneously with the donor template (see International (PCT) Application Publication No. WO 2017 / 053729). One of ordinary skill in the art will be able to select an appropriate timing based on the form of delivery (e.g., the time required for transcription and translation of the RNA and protein components is considered) and the half-life of the molecule(s) within the cell. In a specific embodiment, when a dual-guide CRISPR-Cas system containing a Cas protein is delivered by electroporation (e.g., as an RNP), the donor template (e.g., as an AAV) is introduced into the cell within 4 hours (e.g., within 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, 6 minutes, 7 minutes, 8 minutes, 9 minutes, 10 minutes, 11 minutes, 12 minutes, 13 minutes, 14 minutes, 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, 60 minutes, 90 minutes, 120 minutes, 150 minutes, 180 minutes, 210 minutes, or 240 minutes) after the introduction of the engineered non-natural system.
[0157] In some specific embodiments, the donor template is conjugated to the modulator nucleic acid by a covalent bond. Covalent linkages suitable for this conjugation are known in the art and are described, for example, in U.S. Patent No. 9,982,278 and Savic et al. (2018) ELIFE 7:e33761. In some specific embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., the 5' end of the modulator nucleic acid) via an internucleotide bond. In some specific embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., the 5' end of the modulator nucleic acid) via a linker.
[0158] Efficiency and specificity The engineered non-natural system of the present invention has the advantage that the efficiency of nucleic acid targeting, cleavage or modification can be increased or decreased, for example, by adjusting the hybridization of dual guide nucleic acids and the length of the spacer sequence.
[0159] In some particular embodiments, engineered non - naturally occurring systems have high efficiency. For example, in some particular embodiments, upon contact with an engineered non - naturally occurring system, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% of a nucleic acid population having a target nucleotide sequence and cognate PAM is targeted, cleaved or modified. In some particular embodiments, upon contact with an engineered non - naturally occurring system, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% of the genome of a cell population is targeted, cleaved or modified.
[0160] It has been observed that the occurrence of on - target events and the occurrence of off - target events generally correlate. For some particular therapeutic purposes, low on - target efficiency may be tolerated and low off - target frequency is more desirable. For example, when editing or modifying proliferating cells that grow in vivo upon delivery to a subject, the tolerance for off - target events is low. However, prior to delivery, it is possible to evaluate on - target and off - target events and thereby select one or more colonies having the desired edit or modification and no undesired edits or modifications.
[0161] The methods disclosed herein are suitable for such use. In some particular embodiments, when a nucleic acid population having a target nucleotide sequence and a cognate PAM is contacted with the engineered non-naturally occurring systems disclosed herein, the frequency of off-target events (e.g., targeting, cleavage or modification depending on the function of the CRISPR-Cas system) is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% lower than the frequency of off-target events when using the corresponding CRISPR system comprising a single guide nucleic acid (e.g., a single crRNA consisting of the sequences of a targeter nucleic acid and a modulator nucleic acid) under the same conditions. In some particular embodiments, when genomic DNA having a target nucleotide sequence and a cognate PAM is contacted with the engineered non-naturally occurring systems disclosed herein in a cell population, the frequency of off-target events (e.g., targeting, cleavage or modification depending on the function of the CRISPR-Cas system) is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% lower than the frequency of off-target events when using the corresponding CRISPR system comprising a single guide nucleic acid (e.g., a single crRNA consisting of the sequences of a targeter nucleic acid and a modulator nucleic acid) under the same conditions.In some specific embodiments, when delivered to a population of cells containing genomic DNA having a target nucleotide sequence and cognate PAM, the frequency of off-target events (e.g., targeting, cleavage or modification depending on the function of the CRISPR-Cas system) in cells that have received the engineered non-naturally occurring system disclosed herein is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% lower than the frequency of off-target events in cells that have received a corresponding CRISPR system containing a single guide nucleic acid (e.g., a single crRNA consisting of the sequences of a targeter nucleic acid and a modulator nucleic acid) under the same conditions. Methods for assessing off-target events are summarized in Lazzarotto et al. (2018) Nat Protoc. 13(11):2615-42 and include in situ Cas off-target discovery and sequencing (DISCOVER-seq) validation as disclosed in Wienert et al. (2019) Science 364(6437):286-89; genome-wide unbiased identification of double-strand breaks (DSBs) enabled by sequencing (GUIDE-seq) as disclosed in Kleinstiver et al. (2016) Nat. Biotech. 34:869-74; and cyclization for in vitro reporting of cleavage effects by sequencing (CIRCLE-seq) as described in Kocak et al. (2019) Nat. Biotech. 37:657-66. In some specific embodiments, off-target events include targeting, cleavage or modification at a given off-target locus (e.g., the locus at which the most off-target events are detected). In some specific embodiments, off-target events include collective targeting, cleavage or modification at all loci having detectable off-target events.
[0162] Multiplex method The methods for targeting, editing, and / or modifying genomic DNA disclosed herein can be carried out in a multiplex manner. For example, a library of targeter nucleic acids can be used to target genomic loci; also, a library of donor templates can be used to generate multiple insertions, deletions, and / or substitutions. Multiplex assays can be carried out in screening methods, in which each individual cell culture (e.g., within the wells of a 96-well plate or a 384-well plate) is exposed to different targeter nucleic acids or different combinations of targeter nucleic acids and donor templates. Also, multiplex assays can be carried out in selection methods, in which cell cultures are exposed to a mixed population of different targeter nucleic acids and / or donor templates, and cells having a desired characteristic (e.g., functionality) are enriched or selected by, for example, favorable survival or proliferation, resistance to a particular drug, expression of a detectable protein (e.g., a fluorescent protein detectable by flow cytometry).
[0163] In some particular embodiments, in the multiplex method, a plurality of targeter nucleic acids capable of hybridizing to different target nucleotide sequences are used. In some particular embodiments, the plurality of targeter nucleic acids contain a common targeter stem sequence. In some particular embodiments, in the multiplex method, a single modulator nucleic acid capable of hybridizing to the plurality of targeter nucleic acids is used. In some particular embodiments, in the multiplex method, a single Cas protein (e.g., a Cas nuclease) disclosed herein is used.
[0164] In some specific embodiments, in the multiplexing method, a plurality of targeter nucleic acids that can hybridize to different target nucleotide sequences proximal or adjacent to different PAMs are used. In some specific embodiments, the plurality of targeter nucleic acids include different targeter stem sequences. In some specific embodiments, in the multiplexing method, a plurality of modulator nucleic acids, each of which can hybridize to a different targeter nucleic acid, are used. In some specific embodiments, in the multiplexing method, a plurality of Cas proteins (e.g., Cas nucleases) disclosed herein with different PAM specificities are used.
[0165] In some specific embodiments, the multiplexing method further includes introducing one or more donor templates into the cell population. In some specific embodiments, in the multiplexing method, a plurality of modulator nucleic acids, each of which includes a different donor template recruitment sequence, are used, where each donor template recruitment sequence can hybridize to a different donor template.
[0166] In some specific embodiments, the plurality of targeter nucleic acids and / or the plurality of donor templates are designed for saturation editing. For example, in some specific embodiments, each nucleotide position within the sequence of interest is systematically modified with each of the four conventional bases A, T, G, and C. In other embodiments, at least one sequence within each gene of a pool of genes of interest is modified, e.g., according to a CRISPR design algorithm. In some specific embodiments, each sequence of a pool of exogenous elements of interest (e.g., protein-coding sequences, non-protein-coding genes, regulatory elements) is inserted into one or more given loci of the genome.
[0167] It will be understood that multiplex methods suitable for performing screening or selection methods typically carried out for research purposes may differ from those suitable for therapeutic purposes. For example, constitutive expression of certain elements (e.g., Cas nucleases and / or modulator nucleic acids) may be undesirable for therapeutic purposes due to the potential for increased off-target effects. Conversely, for research purposes, constitutive expression of Cas nucleases and / or modulator nucleic acids may be desirable. For example, constitutive expression provides a broad time frame in which other elements can be introduced. Once a stable cell line is established for constitutive expression, the number of exogenous elements that need to be co-delivered into a single cell also decreases. Thus, constitutive expression of certain elements can increase the efficiency of the screening or selection process and reduce complexity. Also, inducible expression of certain elements of the systems disclosed herein can be used for research purposes where similar advantages are obtained. Expression can be induced by exogenous factors (e.g., small molecules) or by endogenous molecules or complexes present in a particular cell type (e.g., at a particular stage of differentiation). Methods known in the art, such as those described in the above subsection "CRISPR expression systems", can be used to constitutively or inducibly express one or more elements.
[0168] Although it is necessary to introduce at least three elements - a targeter nucleic acid, a modulator nucleic acid, and a Cas protein - it will be further understood that these three elements may be delivered into the cell as a single pre-formed complex of RNP. Thus, the efficiency of the screening or selection process can also be achieved in a multiplex manner by pre-organizing multiple RNP complexes.
[0169] In certain embodiments, the methods disclosed herein further comprise identifying a targeter nucleic acid, a modulator nucleic acid, a Cas protein, a donor template, or a combination of two or more of these elements by a screening process or a selection process. For example, identification may be facilitated by using a set of barcodes within the donor template between two homology arms. In certain embodiments, the methods comprise obtaining a cell population; selectively amplifying genomic DNA or RNA samples and / or barcodes comprising the target nucleotide sequence(s); and / or sequencing the selectively amplified genomic DNA or RNA samples and / or barcodes.
[0170] In another aspect, the invention provides a library comprising a plurality of targeter nucleic acids disclosed herein, optionally further comprising one or more modulator nucleic acids disclosed herein. In another aspect, the invention provides a library comprising a plurality of nucleic acids comprising regulatory elements each operably linked to a different targeter nucleic acid disclosed herein, optionally further comprising regulatory elements operably linked to a modulator nucleic acid disclosed herein. Such libraries can be used in screening or selection methods in combination with one or more Cas proteins or Cas-encoding nucleic acids disclosed herein and / or one or more donor templates as disclosed herein.
[0171] III. Pharmaceutical Compositions The invention provides a composition (e.g., a pharmaceutical composition) comprising an engineered non-naturally occurring system or eukaryotic cell disclosed herein. In certain embodiments, the composition comprises a complex of a targeter nucleic acid and a modulator nucleic acid. In certain embodiments, the composition comprises an RNP comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein (e.g., a Cas nuclease or related Cas protein that can be activated by the targeter nucleic acid and the modulator nucleic acid).
[0172] Also provided by the present invention is a method for preparing a composition, which includes incubating a target nucleic acid and a modulator nucleic acid of an engineered non-naturally occurring system disclosed herein under appropriate conditions, thereby producing a composition (e.g., a pharmaceutical composition) containing a complex of the target nucleic acid and the modulator nucleic acid. In some specific embodiments, the method further includes incubating the target nucleic acid and the modulator nucleic acid with a Cas protein (e.g., a Cas nuclease or related Cas protein that can be activated by the target nucleic acid and the modulator nucleic acid), thereby further producing a complex (e.g., an RNP) of the target nucleic acid, the modulator nucleic acid, and the Cas protein. In some specific embodiments, the method further includes purifying the complex (e.g., the RNP).
[0173] For therapeutic use, an engineered non-naturally occurring system, a CRISPR expression system, or a cell comprising such a system or modified by such a system as disclosed herein is combined with a pharmaceutically acceptable carrier. As used herein, the term “pharmaceutically acceptable” means that a compound, material, composition, and / or dosage form is suitable for use in contact with human and animal tissues within the scope of sound medical judgment, without undue toxicity, irritation, allergic response, or other problems or complications, and commensurate with a reasonable benefit / risk ratio.
[0174] As used herein, the term "pharmaceutically acceptable carrier" refers to buffers, carriers, and excipients that are suitable for use in contact with human and animal tissues without undue toxicity, hypersensitivity, allergic reactions, or other problems or complications, and that have a reasonable benefit / risk ratio. Pharmaceutically acceptable carriers include any standard pharmaceutical carrier, such as phosphate buffered saline solution, water, emulsions (such as oil / water or water / oil emulsions), and various types of wetting agents. Stabilizers and preservatives may also be included in the composition. For examples of carriers, stabilizers, and adjuvants, see, e.g., Martin, Remington’s Pharmaceutical Sciences, 15th Ed., Mack Publ. Co., Easton, PA (1975). Pharmaceutically acceptable carriers include buffers, solvents, dispersion media, coatings, isotonic agents, and absorption delaying agents that are compatible with the administration of pharmaceuticals. The use of such media and agents for pharmaceutically active substances is known in the art.
[0175] In some specific embodiments, salts such as NaCl, MgCl 2 , KCl, MgSO 4 etc.; buffers such as Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), MES sodium salt, 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.; solubilizing agents; detergents such as nonionic detergents such as Tween-20, etc.; nuclease inhibitors; etc. are included. For example, in some specific embodiments, the subject composition includes a buffer for stabilizing the subject DNA targeting RNA and nucleic acids.
[0176] In certain embodiments, the pharmaceutical composition may include formulation materials for modifying, maintaining, or preserving, for example, the pH, volume osmolarity, viscosity, clarity, color, isotonicity, sterility, stability, rate of dissolution or release, adsorption, or permeability of the composition.In such an embodiment, suitable formulation materials include, but are not limited to, amino acids (e.g., glycine, glutamine, asparagine, arginine or lysine); antibacterial agents; antioxidants (e.g., ascorbic acid, sodium sulfite or sodium bisulfite); buffers (e.g., borate, bicarbonate, Tris-HCl, citrate, phosphate or other organic acids); bulking agents (e.g., mannitol or glycine); chelating agents (e.g., ethylenediaminetetraacetic acid (EDTA)); complexing agents (e.g., caffeine, polyvinylpyrrolidone, beta-cyclodextrin or hydroxypropyl-beta-cyclodextrin); fillers; monosaccharides; disaccharides; and other carbohydrates (e.g., glucose, mannose or dextrin); proteins (e.g., serum albumin, gelatin or immunoglobulins); coloring agents, flavoring agents and diluents; emulsifying agents; hydrophilic polymers (e.g., polyvinylpyrrolidone); low molecular weight polypeptides; salt-forming counterions (e.g., sodium); preservatives (e.g., benzalkonium chloride, benzoic acid, salicylic acid, thimerosal, phenethyl alcohol, methylparaben, propylparaben, chlorhexidine, sorbic acid or hydrogen peroxide); solvents (e.g., glycerin, propylene glycol or polyethylene glycol); sugar alcohols (e.g., mannitol or sorbitol); suspending agents; surfactants or wetting agents (e.g., pluronics, PEG, sorbitan esters, polysorbates, e.g., polysorbate 20, polysorbate, triton, tromethamine, lecithin, cholesterol, tyloxapal); stability enhancers (e.g., sucrose or sorbitol); tonicity enhancers (e.g., alkali metal halides, preferably sodium chloride or potassium chloride, mannitol sorbitol); delivery vehicles; diluents; excipients and / or pharmaceutical adjuvants (see Remington’s Pharmaceutical Sciences, 18th ed. (Mack Publishing Company, 1990)).
[0177] In some specific embodiments, the pharmaceutical composition may contain nanoparticles, such as polymeric nanoparticles, liposomes or micelles (see Anselmo et al. (2016) BIOENG. TRANSL. MED. 1:10 - 29). In some specific embodiments, the pharmaceutical composition contains inorganic nanoparticles. Exemplary inorganic nanoparticles include, for example, magnetic nanoparticles (e.g., Fe 3 MnO 2 ) or silica. The outer surface of the nanoparticles may be conjugated with a polymer having a positive charge (e.g., polyethyleneimine, polylysine, polyserine) that enables the binding (e.g., conjugation or encapsulation) of the payload. In some specific embodiments, the pharmaceutical composition contains organic nanoparticles (e.g., the payload is encapsulated inside the nanoparticles). Exemplary organic nanoparticles include, for example, SNALP liposomes coated with polyethylene glycol (PEG) and containing a cationic lipid together with a neutral helper lipid, and complexes of protamine and nucleic acid coated with a lipid coating. In some specific embodiments, the pharmaceutical composition contains liposomes, such as the liposomes disclosed in International Application Publication No. WO 2015 / 148863.
[0178] In some specific embodiments, the pharmaceutical composition contains a targeting moiety that enhances the binding to target cells or the uptake of nanoparticles and liposomes. Exemplary targeting moieties include cell - specific antigens, monoclonal antibodies, single - chain antibodies, aptamers, polymers, sugar chains, and cell - penetrating peptides. In some specific embodiments, the pharmaceutical composition contains a peptide or polymer with membrane - fusogenic or endosome - destabilizing properties.
[0179] In certain embodiments, the pharmaceutical composition may comprise a sustained release formulation or a controlled release formulation. Methods for formulating sustained release or controlled release means, such as liposomal carriers, bioerodible microparticles or porous beads and depot injections, are also known to those skilled in the art. Sustained release preparations may include, for example, porous polymer microparticles or a semipermeable polymer matrix in the form of a shaped article, such as a film or a microcapsule. Examples of sustained release matrices include polyesters, hydrogels, polylactides, copolymers of L-glutamic acid and gamma ethyl-L-glutamate, poly(2-hydroxyethyl-inethacrylate), ethylene vinyl acetate or poly-D(-)-3-hydroxybutyric acid. Also, liposomes, which can be prepared by any of several methods known in the art, may be included as sustained release compositions.
[0180] The pharmaceutical compositions of the present invention can be administered by a variety of methods known in the art. The route and / or mode of administration will vary depending on the desired result. Administration can be intravenous, intramuscular, intraperitoneal or subcutaneous, or can be performed proximal to the target site. The pharmaceutically acceptable carrier is preferably suitable for intravenous, intramuscular, subcutaneous, parenteral, spinal or epidermal administration (e.g., by injection or infusion). Depending on the route of administration, the active compound, i.e., the multispecific antibody of the present invention, may be coated with a substance to protect the compound from the action of acids and other natural conditions that can inactivate the compound.
[0181] Formulation ingredients suitable for parenteral administration include sterile diluents, such as water for injection, saline solution, fixed oils, polyethylene glycols, glycerin, propylene glycol or other synthetic solvents; antibacterial agents, such as benzyl alcohol or methylparaben; antioxidants, such as ascorbic acid or sodium bisulfite; chelating agents, such as EDTA; buffers, such as acetate, citrate or phosphate; and agents for adjusting tonicity, such as sodium chloride or dextrose.
[0182] For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor ELTM (BASF, Parsippany, NJ), or phosphate-buffered saline (PBS). The carrier is preferably stable under the conditions of manufacture and storage and preserved against microorganisms. The carrier can be, for example, a solvent or dispersion medium containing water, ethanol, polyols (such as glycerol, propylene glycol, and liquid polyethylene glycol), and suitable mixtures thereof.
[0183] The pharmaceutical formulation is preferably sterile. Sterilization can be carried out by any suitable method, such as filtration through a sterile filtration membrane. If the composition is lyophilized, filtration sterilization can be carried out either before or after lyophilization and reconstitution. In some specific embodiments, the multispecific antibody is lyophilized and then reconstituted with buffered saline at the time of administration.
[0184] The pharmaceutical composition of the present invention can be prepared according to methods well-known and routinely practiced in the art. See, for example, Remington: The Science and Practice of Pharmacy, Mack Publishing Co., 20th ed., 2000; and Sustained and Controlled Release Drug Delivery Systems, J.R. Robinson, ed., Marcel Dekker, Inc., New York, 1978. The pharmaceutical composition is preferably manufactured under GMP conditions. Typically, a therapeutically effective dose or an effective dose of the multispecific antibody of the present invention is used in the pharmaceutical composition of the present invention. The multispecific antibody of the present invention is formulated into a pharmaceutically acceptable dosage form by conventional methods known to those skilled in the art. The dosage regimen is adjusted so as to provide the optimal desired response (e.g., a therapeutic response). For example, a single bolus may be administered, several divided doses may be administered over time, or the dose may be proportionally reduced or increased as dictated by the exigencies of the therapeutic situation. It is particularly advantageous to formulate parenteral compositions in unit dosage forms for ease of administration and uniformity of dosage. As used herein, a unit dosage form refers to a physically discrete unit suitable as a dosage unit for the subject to be treated; each unit contains a predetermined quantity of the active compound calculated to produce the desired therapeutic effect, together with the required pharmaceutical carrier.
[0185] The actual dosage level of the active ingredient in the pharmaceutical composition of the present invention may be varied so that an amount of the active ingredient that is effective to obtain the desired therapeutic response for a particular patient, composition, and mode of administration is obtained and is not toxic to the patient. The selected dosage level will depend upon a variety of pharmacokinetic factors, such as the activity of the specific composition of the present invention or its ester, salt or amide employed, the route of administration, the time of administration, the rate of excretion of the specific compound being used, the duration of the treatment, other drugs, compounds and / or substances used in combination with the specific composition employed, and the age, sex, weight, condition, general health and prior medical history of the patient being treated.
[0186] IV. Therapeutic Use The engineered non - naturally occurring systems and CRISPR expression systems disclosed herein are useful for targeting, editing, and / or modifying genomic DNA in cells or in the genome of an organism. Such systems, as well as cells containing one of the systems or cells in which the genome has been modified by an engineered non - naturally occurring system, can be used to treat diseases or disorders where modification of genetic or epigenetic information is desirable. Thus, in another aspect, the present invention provides a method of treating a disease or disorder, comprising administering to a subject in need thereof a non - naturally occurring system, a CRISPR expression system, or a cell disclosed herein.
[0187] The term "subject" includes humans and non - human animals. Non - human animals include any vertebrate, such as mammals and non - mammals, such as non - human primates, sheep, dogs, cows, chickens, amphibians, and reptiles. Unless otherwise noted, the terms "patient" or "subject" are used interchangeably herein.
[0188] The terms "treatment", "treating", "treat", "treated", etc., as used herein, refer to obtaining a desired pharmacological and / or physiological effect. The effect can be therapeutic in terms of partial or complete cure of a disease and / or adverse effects resulting from the disease, or in terms of delaying the progression of the disease. "Treatment" as used herein includes any treatment of a disease in a mammal, such as a human, and includes (a) suppressing the disease, i.e., arresting its development; and (b) alleviating the disease, i.e., causing regression of the disease. It will be understood that a disease or disorder can be identified by genetic methods and treated before any medical symptoms manifest.
[0189] For therapeutic purposes, the methods disclosed herein are particularly suitable for editing or modifying proliferating cells, such as stem cells (e.g., hematopoietic stem cells), progenitor cells (e.g., hematopoietic progenitor cells or lymphoid progenitor cells), or memory cells (e.g., memory T cells). Considering that such cells proliferate in vivo when delivered to a subject, the tolerance for off-target events is low. However, prior to delivery, it is possible to evaluate on-target and off-target events and thereby select one or more colonies that have the desired editing or modification but no unwanted editing or modification. Thus, a decreased editing or modification efficiency may be tolerated by such cells. The engineered non-naturally occurring systems of the invention have the advantage that the efficiency of nucleic acid cleavage can be increased or decreased, for example, by modulating the hybridization of dual guide nucleic acids. As a result, this can be used to minimize off-target events when generating genetically engineered proliferating cells.
[0190] To minimize toxicity and off-target effects, it is important to control the concentration of the dual guide CRISPR-Cas system delivered. The optimal concentration can be determined by testing various concentrations in cell models, tissue models, or non-human eukaryotic animal models and analyzing the extent of modification at potential off-target genomic loci using deep sequencing. The concentration that gives the highest on-target modification level with minimal off-target modification level should be selected for ex vivo or in vivo delivery.
[0191] Gene therapy It will be understood that the engineered non-naturally occurring systems and CRISPR expression systems disclosed herein can be used to treat genetic diseases or disorders, i.e., diseases or disorders associated with or mediated by an unwanted mutation in the genome of a subject.
[0192] Exemplary genetic diseases or disorders include age-related macular degeneration, adrenoleukodystrophy (ALD), Alagille syndrome, alpha-1-antitrypsin deficiency, argininemia, argininosuccinic aciduria, ataxia (e.g., Friedreich's ataxia, spinocerebellar ataxia, ataxia telangiectasia, essential tremor, spastic paraplegia), autism, biliary atresia, biotinidase deficiency, carbamoyl phosphate synthetase I deficiency, congenital disorder of glycosylation (CDGS), central nervous system (CNS)-related disorders (e.g., Alzheimer's disease, amyotrophic lateral sclerosis (ALS), Canavan disease (CD), ischemia, multiple sclerosis (MS), neuropathic pain, Parkinson's disease), Bloom syndrome, cancer, Charcot-Marie-Tooth disease (e.g., peroneal muscular atrophy, hereditary motor and sensory neuropathy), congenital hepatic porphyria, citrullinemia, Crigler-Najjar syndrome, cystic fibrosis (CF), dentatorubral-pallidoluysian atrophy (DRPLA), diabetes insipidus, Fabry disease, familial hypercholesterolemia (LDL receptor deficiency), Fanconi anemia, fragile X syndrome, fatty acid metabolism disorder, galactosemia, glucose-6-phosphate dehydrogenase (G6PD), glycogenosis (e.g., type I (glucose-6-phosphatase deficiency, von Gierke), type II (alpha-glucosidase deficiency, Pompe disease), type III (debranching enzyme deficiency, Cori disease), type IV (branching enzyme deficiency, Andersen disease), type V (muscle glycogen phosphorylase deficiency, McArdle disease), type VII (muscle phosphofructokinase deficiency, Tarui disease), type VI (liver phosphorylase deficiency,Ehlers-Danlos disease), type IX (hepatic glycogen phosphorylase kinase deficiency)), hemophilia A (associated with deficiency of factor VIII), hemophilia B (associated with deficiency of factor IX), Huntington's disease, glutaric aciduria, hypophosphatemia, Krabbe disease, lactic acidosis, Lafora disease, Leber congenital amaurosis, Lesch-Nyhan syndrome, lysosomal storage diseases, metachromatic leukodystrophy (MLD), mucopolysaccharidosis (MPS) (e.g., Hunter syndrome, Hurler syndrome, Maroteaux-Lamy syndrome, Sanfilippo syndrome, Scheie syndrome, Morquio syndrome, etc., MPSI, MPSII, MPSIII, MSIV, MPS VII), muscle / skeletal disorders (e.g., muscular dystrophy, Duchenne muscular dystrophy), myotonic dystrophy (DM), neovascularization, N-acetylglutamate synthase deficiency, ornithine transcarbamylase deficiency, phenylketonuria, primary open-angle glaucoma, retinitis pigmentosa, schizophrenia, severe combined immunodeficiency (SCID), spinal and bulbar muscular atrophy (SBMA), sickle cell anemia, Usher syndrome, Tay-Sachs disease, thalassemia (e.g., beta thalassemia), trinucleotide repeat diseases, tyrosinemia, Wilson's disease, Wiskott-Aldrich syndrome, X-linked chronic granulomatous disease (CGD), X-linked severe combined immunodeficiency, and xeroderma pigmentosum are included.,
[0193] Additional exemplary genetic diseases or disorders and related information are available on the World Wide Web at kumc.edu / gec / support, genome.gov / 10001200, and ncbi.nlm.nih.gov / books / NBK22183 / . Additional exemplary genetic diseases or disorders, related gene mutations, and gene therapy approaches for treating genetic diseases or disorders are described in International (PCT) Publication Nos. WO 2013 / 126794, WO 2013 / 163628, WO 2015 / 048577, WO 2015 / 070083, WO 2015 / 089354, WO 2015 / 134812, WO 2015 / 138510, WO 2015 / 148670, WO 2015 / 148860, WO 2015 / 148863, WO 2015 / 153780, WO 2015 / 153789, and WO 2015 / 153791, and U.S. Patent Application Publication Nos. 2009 / 0222937, 2009 / 0271881, 2009 / 0271881, 2010 / 0229252, 2010 / 0311124, 2011 / 0016540, 2011 / 0023139, 2011 / 0023144, 2011 / 0023145, 2011 / 0023145, 2011 / 0023146, 2011 / 0023153, 2011 / 0091441, 2011 / 0158957, 2011 / 0182867, 2011 / 0225664, 2012 / 0159653, 2012 / 0328580, 2013 / 0145487, and 2013 / 0202678.
[0194] Manipulation of immune cells It will be understood that the engineered non-naturally occurring systems and CRISPR expression systems disclosed herein can be used to engineer immune cells. Immune cells include, without limitation, lymphocytes (e.g., B lymphocytes or B cells, T lymphocytes or T cells, and natural killer cells), myeloid cells (e.g., monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes), and stem cells and progenitor cells that can differentiate into these cell types (e.g., hematopoietic stem cells, hematopoietic progenitor cells, and lymphoid progenitor cells). Cells can include autologous cells derived from the subject to be treated or also allogeneic cells derived from a donor.
[0195] In some particular embodiments, the immune cell is a T cell, which can be, for example, a cultured T cell, a primary T cell, a T cell derived from a cultured T cell line (e.g., Jurkat, SupTi), or a T cell isolated from a mammal, e.g., the subject to be treated. When isolated from a mammal, T cells can be isolated from a number of sources, e.g., without limitation, blood, bone marrow, lymph nodes, thymus, or other tissues or body fluids. Also, T cells can be enriched or purified. The T cell can be any type of T cell and at any stage of development, e.g., without limitation, CD4 + / CD8 + double positive T cells, CD4 + helper T cells (e.g., Th1 and Th2 cells), CD8 + T cells (e.g., cytotoxic T cells), tumor infiltrating lymphocytes (TILs), memory T cells (e.g., central memory T cells and effector memory T cells), regulatory T cells, naive T cells, etc.
[0196] In some particular embodiments, immune cells, such as T cells, are engineered to express foreign genes. For example, in some particular embodiments, the engineered CRISPR systems disclosed herein can be used to engineer immune cells to express foreign genes. For example, in some particular embodiments, the engineered CRISPR systems disclosed herein can catalyze DNA cleavage at a locus and enable site-specific integration of a foreign gene by HDR at the locus.
[0197] In certain embodiments, immune cells, such as T cells, are engineered to express a chimeric antigen receptor (CAR), i.e., the T cells are included with an exogenous nucleotide sequence encoding the CAR. As used herein, the term "chimeric antigen receptor" or "CAR" refers to any artificial receptor that includes an antigen-specific binding portion and one or more signaling chains derived from an immune receptor. The CAR may include, for example, a single-chain variable fragment (scFv) of an antibody specific for an antigen linked to a cytoplasmic domain of a T cell signaling molecule via a hinge and transmembrane region, such as a tandem T cell trigger domain (e.g., derived from CD3ζ) and a T cell co-stimulatory domain (e.g., derived from CD28, CD137, OX40, ICOS or CD27). T cells expressing a chimeric antigen receptor are referred to as CAR T cells. Exemplary CAR T cells include CD19-targeted CTL019 cells (see Grupp et al. (2015) BLOOD, 126:4983), 19-28z cells (see Park et al. (2015) J. CLIN. ONCOL., 33:7010) and KTE-C19 cells (see Locke et al. (2015) BLOOD, 126:3991). Further exemplary CAR T cells are described in U.S. Pat. Nos. 8,399,645, 8,906,682, 7,446,190, 9,181,527, 9,272,002 and 9,266,960, U.S. Patent Application Publication Nos. 2016 / 0362472, 2016 / 0200824 and 2016 / 0311917, and International (PCT) Publication Nos. 2013 / 142034, 2015 / 120180, 2015 / 188141, 2016 / 120220 and 2017 / 040945. Exemplary approaches for expressing a CAR using the CRISPR system are described in Hale et al. (2017) MOL THER METHODS CLIN DEV., 4:192, MacLeod et al. (2017) MOL THER, 25:949 and Eyquem et al. (2017) NATURE, 543:113.
[0198] In some particular embodiments, immune cells, such as T cells, bind via a T cell receptor (TCR) endogenous to an antigen, such as a cancer antigen. In some particular embodiments, immune cells, such as T cells, are engineered to express an exogenous TCR, such as an exogenous naturally occurring TCR or an exogenous engineered TCR. The T cell receptor comprises two chains referred to as the α-chain and the β-chain, which are integrated at the T cell surface and form a heterodimeric receptor capable of recognizing MHC-restricted antigens. Each of the α-chain and the β-chain contains a constant region and a variable region. Each variable region of the α-chain and the β-chain defines three loops known as complementarity-determining regions (CDRs) that confer antigen-binding activity and binding specificity to the T cell receptor 1 , CDR 2 and CDR 3 and define three loops referred to as complementarity-determining regions (CDRs).
[0199] In certain embodiments, the CAR or TCR binds to a cancer antigen selected from B cell maturation antigen (BCMA), mesothelin, prostate specific membrane antigen (PSMA), prostate stem cell antigen (PCSA), carbonic anhydrase IX (CAIX), carcinoembryonic antigen (CEA), CD5, CD7, CD10, CD19, CD20, CD22, CD30, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD70, CD74, CD123, CD133, CD138, epithelial glycoprotein 2 (EGP 2), epithelial glycoprotein-40 (EGP-40), epithelial cell adhesion molecule (EpCAM), receptor tyrosine protein kinase (FLT3), folate binding protein (FBP), fetal acetylcholine receptor (AChR), folate receptors -a and β (FRa and β), ganglioside G2 (GD2), ganglioside G3 (GD3), epidermal growth factor receptor 2 (HER-2 / ERB2), epidermal growth factor receptor vIII (EGFRvIII), ERB3, ERB4, human telomerase reverse transcriptase (hTERT), interleukin-13 receptor subunit alpha-2 (IL-13Ra2), K-light chain, kinase insert domain receptor (KDR), Lewis A (CA19.9), Lewis Y (LeY), LI cell adhesion molecule (LICAM), melanoma associated antigen 1 (Melanoma antigen family Al, MAGE-A1), mucin 16 (MUC-16), mucin 1 (MUC-1; e.g., cleaved MUC-1), KG2D ligand, cancer-testis antigen NY-ESO-1, tumor fetal antigen (h5T4), tumor associated glycoprotein 72 (TAG-72), vascular endothelial growth factor R2 (VEGF-R2), Wilms tumor protein (WT-1), receptor tyrosine kinase transmembrane receptor 1 (ROR1), B7-H3 (CD276), B7-H6 (Nkp30), chondroitin sulfate proteoglycan-4 (CSPG4), DNAX accessory molecule (DNAM-1), Ephrin type A receptor 2 (EpHA2), fibroblast activation protein (FAP), Gpl00 / HLA-A2, glypican 3 (GPC3), HA-IH, HERK-V, IL-1 IRa, latent membrane protein 1 (LMP1), neural cell adhesion molecule (N-CAM / CD56) and TRAIL receptor (TRAIL-R).
[0200] Loci suitable for insertion of a CAR coding sequence or a foreign TCR coding sequence include, without limitation, safe harbor loci (e.g., the AAVS1 locus), TCR subunit loci (e.g., the TCRα constant region (TRAC) locus), and other loci associated with certain advantages (e.g., the CCR5 locus, the inactivation of which can suppress or reduce HIV infection). It will be understood that insertion into the TRAC locus reduces tonic CAR signaling and improves the efficacy of T cells (see Eyquem et al. (2017) NATURE, 543:113). Furthermore, inactivation of the endogenous TRAC gene can reduce the graft-versus-host disease (GVHD) response, thereby potentially allowing allogeneic T cells to be used as a starting material for the preparation of CAR-T cells. Thus, in some particular embodiments, immune cells, such as T cells, are engineered to have reduced expression of an endogenous TCR or TCR subunit, such as the TCRα subunit constant region (TRAC). The cells may be engineered to have a partial reduction in the expression of the endogenous TCR or TCR subunit, or to have no expression of the endogenous TCR or TCR subunit. For example, in some particular embodiments, the expression of an endogenous TCR or TCR subunit in an immune cell, such as a T cell, is engineered to be less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10% or less than 5%) compared to the corresponding unmodified or parental cell. In some particular embodiments, immune cells, such as T cells, are engineered to have no detectable expression of an endogenous TCR or TCR subunit. Exemplary approaches for reducing TCR expression using the CRISPR system are described in U.S. Patent No. 9,181,527, Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, Cooper et al. (2018) LEUKEMIA, 32:1970 and Ren et al. (2017) ONCOTARGET, 8:17002.
[0201] In addition, some specific immune cells, such as T cells, also express major histocompatibility complex (MHC) genes or human leukocyte antigen (HLA) genes, and inactivation of these endogenous genes can reduce the GVHD response, thereby making it possible to use allogeneic T cells as starting materials for the preparation of CAR-T cells. It will be understood, therefore, that in some specific embodiments, immune cells, such as T cells, are engineered to have a reduced expression of one or more endogenous class I or class II MHC or HLA (e.g., beta2-microglobulin (B2M), class II major histocompatibility complex transactivator (CIITA), HLA-E and / or HLA-G). The cells may be engineered to have a partial reduction in the expression of endogenous MHC or HLA, or to have no expression of endogenous MHC or HLA. For example, in some specific embodiments, the expression of endogenous MHC (e.g., B2M, CIITA, HLA-E or HLA-G) in immune cells, such as T cells, is engineered to be less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10% or less than 5%) compared to the corresponding unmodified or parental cells. In some specific embodiments, immune cells, such as T cells, are engineered to have no detectable expression of endogenous MHC (e.g., B2M, CIITA, HLA-E or HLA-G). Exemplary approaches for reducing MHC expression using the CRISPR system are described in Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255 and Ren et al. (2017) ONCOTARGET, 8:17002.
[0202] Other genes that can be inactivated to reduce the GVHD response include, without limitation, CD3, CD52, and deoxythymidine kinase (DCK). For example, inactivation of CK can render immune cells (e.g., T cells) resistant to purine nucleotide analog (PNA) compounds that are often used to suppress the host immune system to reduce the GVHD response during immune cell therapy.
[0203] In some specific embodiments, immune cells, such as T cells, are engineered to have reduced expression of an endogenous gene. For example, in some specific embodiments, immune cells can be engineered to have reduced expression of an endogenous gene using the engineered CRISPR systems disclosed herein. For example, in some specific embodiments, DNA cleavage at a locus is effected by the engineered CRISPR systems disclosed herein, whereby the gene targeted can be inactivated. In other embodiments, the engineered CRISPR systems disclosed herein can be fused to an effector domain (e.g., a transcriptional repressor or a histone methyltransferase) to reduce the expression of a target gene.
[0204] It will be understood that the activity of immune cells (e.g., T cells) can be enhanced by inactivating or reducing the expression of immunosuppressive factors, such as immune checkpoint proteins. Thus, in some particular embodiments, immune cells, such as T cells, are engineered to have a reduced expression of immune checkpoint proteins. Exemplary immune checkpoint proteins expressed by wild-type T cells include, but are not limited to, PD-1, CTLA-4, A2AR, B7-H3, B7-H4, BTLA, KIR, LAG3, TIM-3, TIGIT, VISTA, PTPN6 (SHP-1), and FAS. The cells can be modified to have a partial reduction in the expression of immune checkpoint proteins or to have no expression of immune checkpoint proteins. For example, in some particular embodiments, the expression of immune checkpoint proteins in immune cells, such as T cells, is engineered to be less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) compared to the corresponding unmodified or parental cells. In some particular embodiments, immune cells, such as T cells, are engineered to have no detectable expression of immune checkpoint proteins. Exemplary approaches for reducing the expression of immune checkpoint proteins using the CRISPR system are described in International (PCT) Publication No. WO 2017 / 017184, Cooper et al. (2018) Leukemia, 32:1970, Su et al. (2016) Oncoimmunology, 6:e1249558, and Zhang et al. (2017) Front Med, 11:554.
[0205] In some particular embodiments, immune cells, such as T cells, are modified to express a dominant negative form of an immune checkpoint protein. In some particular embodiments, a dominant negative form of a checkpoint inhibitor binds to a natural ligand that would otherwise bind to and activate the wild-type immune checkpoint protein, or can function as a decoy receptor that sequesters the natural ligand. Examples of engineered immune cells, such as T cells, that contain a dominant negative form of an immunosuppressive factor are described, for example, in International (PCT) Publication No. WO 2017 / 040945.
[0206] In some particular embodiments, immune cells, such as T cells, are modified to express genes (such as transcription factors, cytokines or enzymes) that regulate the survival, proliferation, activity or differentiation (such as into memory cells) of the immune cells. In some particular embodiments, the immune cells are modified to express TET2, FOXO1, IL-12, IL-15, IL-18, IL-21, IL-7, GLUT1, GLUT3, HK1, HK2, GAPDH, LDHA, PDK1, PKM2, PFKFB3, PGK1, ENO1, GYS1 and / or ALDOA. In some particular embodiments, the modification is the insertion of a nucleotide sequence encoding a protein that is operably linked to a regulatory element. In some particular embodiments, the modification is the substitution of a single nucleotide polymorphism (SNP) site within an endogenous gene.
[0207] In some particular embodiments, immune cells, such as T cells, are modified to express proteins (such as cytokines or enzymes) that regulate a microenvironment (such as the tumor microenvironment) in which the immune cells are designed to migrate. In some particular embodiments, the immune cells are modified to express CA9, CA12, a V-ATPase subunit, NHE1 and / or MCT-1.
[0208] V. Kit It will be appreciated that the engineered non-naturally occurring systems, CRISPR expression systems and libraries disclosed herein may be packaged into kits suitable for use by a healthcare provider. Thus, in another aspect, the present invention provides a kit comprising any one or more of the elements disclosed in the above systems, libraries, methods and compositions. In some particular embodiments, the kit comprises an engineered non-naturally occurring system as disclosed herein and instructions for use for using the kit. The instructions for use may be specific to the uses and methods described herein. In some particular embodiments, one or more of the elements of the system are provided in solution. In some particular embodiments, one or more of the elements of the system are provided in lyophilized form and the kit further comprises a diluent. The elements may be provided individually or in combination and may be provided in any suitable container, such as a vial, bottle, tube, or may be immobilized on the surface of a solid phase substrate (e.g., a chip or microarray). In some particular embodiments, the kit comprises one or more of the nucleic acids and / or proteins described herein. In some particular embodiments, the kit comprises all of the elements of the system of the present invention.
[0209] In some particular embodiments of a kit comprising an engineered non-naturally occurring system, the targeter nucleic acid and the modulator nucleic acid are provided in separate containers. In other embodiments, the targeter nucleic acid and the modulator nucleic acid are pre-complexed and the complex is provided in a single container. In some particular embodiments, the kit comprises a nucleic acid encoding a Cas protein or a regulatory element operably linked to a nucleic acid encoding a Cas protein, provided in separate containers. In other embodiments, the kit comprises a Cas protein pre-complexed with the targeter nucleic acid and the modulator nucleic acid, and the complex is provided in a single container.
[0210] For targeting multiple target nucleotide sequences for use, for example, in a screening process or a selection process, a kit comprising a plurality of targeter nucleic acids can be provided. Thus, in some particular embodiments, the kit comprises a plurality of targeter nucleic acids as disclosed herein (e.g., immobilized in separate tubes or on the surface of a solid-phase substrate such as a chip or a microarray), optionally one or more modulator nucleic acids as disclosed herein, and optionally a Cas protein as disclosed herein or a regulatory element operably linked to a nucleic acid encoding a Cas protein. Such kits are useful for identifying the targeter nucleic acid having the highest efficiency and / or specificity for targeting a given gene, for identifying genes involved in a physiological or pathological pathway, or for manipulating cells such that a desired functionality is obtained in a multiplex assay. In some particular embodiments, the kit further comprises one or more donor templates provided in one or more separate containers. In some particular embodiments, the kit comprises a plurality of donor templates as disclosed herein (e.g., immobilized in separate tubes or on the surface of a solid-phase substrate such as a chip or a microarray), one or more targeter nucleic acids as disclosed herein, and one or more modulator nucleic acids as disclosed herein, and optionally a Cas protein as disclosed herein or a regulatory element operably linked to a nucleic acid encoding a Cas protein. Such kits are useful for identifying donor templates that introduce optimal gene modifications in a multiplex assay. Also, a CRISPR expression system as disclosed herein is suitable for use in a kit.
[0211] In some specific embodiments, the kit further comprises one or more reagents and / or buffers for use in a method of utilizing one or more of the elements described herein. The reagents can be provided in any suitable container and can be provided in a form that is usable in a particular assay or in a form that requires the addition of one or more other components prior to use (e.g., in concentrated form or lyophilized form). The buffers can be reaction buffers or storage buffers, such as, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, boric acid buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some specific embodiments, the buffer has a pH of from about 7 to about 10. In some specific embodiments, the kit further comprises a pharmaceutically acceptable carrier. In some specific embodiments, the kit further comprises one or more devices or other materials for administration to a subject.
[0212] Throughout this description, when a composition is described as having, including, or comprising specific components, or when a process and method are described as having, including, or comprising specific steps, it is also contemplated that there are compositions of the invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the invention that consist essentially of, or consist of, the recited process steps.
[0213] In this application, when an element or component is said to be included in, and / or selected from, a list of recited elements or components, it is to be understood that the element or component can be any one of the recited elements or components, or the element or component can be selected from a group consisting of two or more of the recited elements or components.
[0214] Furthermore, it should be understood that the elements and / or features of the compositions or methods described herein, whether explicit or implicit herein, can be combined in various ways without departing from the spirit and scope of the invention. For example, when referring to a particular compound, unless the context indicates otherwise, the compound can be used in various aspects of the compositions of the invention and / or in the methods of the invention. In other words, in this application, the aspects are described and illustrated in a manner that enables clear and concise application to be described and illustrated, but it is intended that the aspects can be variously combined and can be independent without departing from the present teachings and the invention(s), and this should be recognized. For example, it should be recognized that all of the features described and illustrated herein can be applicable to any aspect of the invention(s) described and illustrated herein.
[0215] In the context of describing the present invention (especially in the context of the following claims), the terms "a", "an", "the", and similar references should be construed to include both the singular and the plural unless specifically stated otherwise herein or clearly contradicted by the context. For example, the term "a cell" includes a plurality of cells including a mixture thereof. Also, when the plural form is used for compounds, salts, etc., this should be construed to also mean a single compound, salt, etc.
[0216] The expression "at least one" should be understood to include each individual of the objects described following this expression and various combinations of two or more of the objects described, unless the context and usage indicate otherwise. The expression "and / or" in the context of three or more objects described should be understood to have the same meaning unless the context indicates otherwise.
[0217] The use of the terms "comprising", "includes", "including", "having", "has", "having", "contain", "contains" or "containing", including their grammatically corresponding phrases, is generally open-ended and non-limiting, unless specifically stated otherwise or understood from the context not to be, i.e., it is understood not to exclude additional elements or steps not described.
[0218] When the term "about" is used before a numerical value, the present invention includes the specific numerical value itself, unless specifically stated otherwise. As used herein, the term "about" refers to a variation of ±10% from the nominal value, unless otherwise specified or inferred.
[0219] The order of steps or the order for performing a particular act is to be understood as not being important as long as the functionality of the present invention is maintained. Further, two or more steps or acts may be performed simultaneously.
[0220] Any example or exemplary language herein, such as the use of "such as" or "including", is intended only to better illustrate examples of the present invention and does not limit the scope of the present invention except as set forth in the claims. The language herein should not be construed as indicating that any element not recited in the claims is essential to the practice of the present invention.
Examples
[0221] The following examples are merely illustrative and are not intended to limit the scope or content of the present invention in any way.
[0222] Example 1. In Vitro Cleavage of Target DNA by the Dual-Guide MAD7 CRISPR-Cas System MAD7 is a V-A type Cas protein that has endonuclease activity when complexed with a single guide RNA, also known as crRNA, in a V-A type system (see U.S. Patent No. 9,982,279). In this example, cleavage of target DNA using MAD7 in the state of a complex with dual guide nucleic acids in an in vitro cleavage assay is described.
[0223] Briefly, two different crRNAs named crRNA1 and crRNA2 were designed to target the DNMT1 gene. In particular, crRNA2 has been reported to have a better ability to activate LbCas12a and FnoCas12a of zebrafish (see Liu et al. (2019) NUC. ACIDS RES. 47(8):4169 - 80). The predicted secondary structures of crRNA1 and crRNA2 are shown in FIG. 2A. Also, a set of targeter RNA and modulator RNA corresponding to crRNA1, named crRNA1_targeter1 and crRNA1_modulator1 respectively, and a set of targeter RNA and modulator RNA corresponding to crRNA2, named crRNA2_targeter1 and crRNA2_modulator1 respectively, were designed. Each set of dual guide RNAs indicates the split at the central position of the loop region of the corresponding single guide RNA. The nucleotide sequences of these guide RNAs are shown in Table 2.
[0224] (Table 2) Nucleotide sequences of single guide RNAs and dual guide RNAs tested TIFF0007689949000019.tif56170
[0225] These guide RNAs were chemically synthesized. Human DNMT1 target DNA was prepared by PCR, Included was the nucleotide sequence of TIFF0007689949000020.tif99159. The MAD7 protein contains a nucleoplasmin NLS at its C-terminus, was expressed in E. coli, and purified by fast protein liquid chromatography (FPLC) of the protein.
[0226] The single-guide CRISPR-Cas system and the dual-guide CRISPR-Cas system were tested in in vitro cleavage assays. Briefly, 1 μM of the MAD7 protein was incubated at room temperature for 10 minutes with 1 μM of crRNA1, 1 μM of crRNA1_modulator1, 1 μM of crRNA1_targeter1, the combination of 1 μM of crRNA1_modulator1 and 1 μM of crRNA1_targeter1, 1 μM of crRNA2, 1 μM of crRNA2_modulator1, 1 μM of crRNA2_targeter1, or the combination of 1 μM of crRNA2_modulator1 and 1 μM of crRNA2_targeter1 to form RNP complexes. Then, DNMT1 target DNA was added to the solution at a molar ratio of 10:1 or 1:1 of MAD7 to target DNA. After incubation at 37 °C for 10 minutes, the samples were analyzed by electrophoresis on an agarose gel.
[0227] As shown in Figure 2B, the nuclease activity of MAD7 was activated with crRNA1, crRNA2, and their corresponding dual-guide RNA pairs, and the DNMT1 target DNA was cleaved. In contrast, crRNA1_modulator1, crRNA1_targeter1, crRNA2_modulator1, or crRNA2_targeter1 alone did not show such activity. The ability of crRNA1 to activate the MAD7 nuclease under these conditions was greater than that of crRNA2. For each of crRNA1 and crRNA2, the ability of the single-guide RNA to activate the MAD7 nuclease was greater than that of the corresponding dual-guide system.
[0228] Extension of the modulator RNA at the 5'-end Next, it was evaluated whether the CRISPR-Cas system could allow the addition of nucleotide sequences to the 5' end of crRNA or modulator RNA. Two crRNA sequences named crRNA3 and crRNA4 were designed to contain additional nucleotide sequences at the 5' end of crRNA1. The corresponding dual-guide systems included modulator RNAs named crRNA3_modulator1 and crRNA4_modulator1 paired with crRNA1_targeter1 as the targeter RNA. The sequences of these newly designed guide RNAs are shown in Table 3. The additional nucleotide sequences at the 5' end of the RNA are underlined.
[0229] (Table 3) Nucleotide sequences of the tested crRNAs and modulator RNAs TIFF0007689949000021.tif53170
[0230] These guide RNAs were chemically synthesized. The in vitro cleavage assay was performed using the above method. Each guide RNA was used at a concentration of 1 μM when incubating with MAD7 to form RNP. The molar ratio of MAD7 to the target DNA was 10:1.
[0231] As shown in Figure 3, in all of crRNA1, crRNA3, and crRNA4, the nuclease activity of MAD7 was activated and the DNMT1 target DNA was cleaved. Furthermore, in each of crRNA1_modulator1, crRNA3_modulator1, and crRNA4_modulator1, in combination with crRNA1_targeter1, the MAD7 nuclease was activated. In contrast, neither the targeter RNA alone nor the modulator RNA alone showed such activity. Therefore, under this condition, the additional nucleotide sequences at the 5' end of crRNA or modulator RNA did not seem to have any negative effect on the ability of the guide RNA to activate the MAD7 nuclease.
[0232] In vitro transcribed modulator RNA Next, the activity of in vitro transcribed RNAs in the single-guide CRISPR-Cas system or dual-guide CRISPR-Cas system was evaluated. Briefly, crRNA1 and crRNA3 were transcribed in vitro from chemically synthesized double-stranded template DNA using the MegaScript kit (Ambion). The template DNA contained a T7 promoter immediately upstream of the sequence encoding the RNA of interest, which had the nucleotide sequence of TIFF0007689949000022.tif5128. As a result, the in vitro transcribed RNAs named crRNA1_T7 and crRNA3_T7 contained the nucleotide sequence GG at the 5' end of the transcribed RNA. This RNA was purified using the Oligo Clean and Concentration kit (Zymogen) and quantified with a Nanodrop. The quality of the in vitro transcribed RNA was evaluated on an agarose gel.
[0233] To generate the corresponding dual-guide system, template DNA containing a T7 promoter immediately upstream of the sequence encoding crRNA1_modulator1 or crRNA3_modulator1 was transcribed in vitro. The resulting RNAs were named crRNA1_modulator1_T7 and crRNA3_modulator1_T7, each of which contained the nucleotide sequence GG at the 5' end of the transcribed RNA. These RNA samples were purified and their quantity and quality were evaluated as described above. These in vitro transcribed modulator RNAs were used in combination with chemically synthesized crRNA1_targeter1.
[0234] The in vitro transcribed RNA was tested using the above method in an in vitro cleavage assay. Each guide RNA was used at a concentration of 1 μM when incubating with MAD7 to form an RNP. The molar ratio of MAD7 to target DNA was 10:1.
[0235] As shown in FIG. 3, crRNA1_T7 and crRNA3_T7 retained the ability to activate the MAD7 nuclease. Similarly, the combinations of (1) crRNA1_modulator1_T7 and crRNA1_targeter1 and (2) crRNA3_modulator1_T7 and crRNA1_targeter1 also retained the ability to activate the MAD7 nuclease. Therefore, under this condition, in vitro transcribed crRNAs and modulator RNAs, despite containing additional nucleotide sequences at the 5' end, were suitable for use in each of the single-guide CRISPR-Cas system and the dual-guide CRISPR-Cas system.
[0236] "Loop" ends of the modulator RNA and the targeter RNA The above dual-guide RNA was designed by splitting the single-guide RNA at the central position of the crRNA loop. Next, variants of the dual-guide RNA system with the single-guide RNA split at different positions within the loop were evaluated. As shown in FIGS. 4A-4F, crRNA1 (also referred to herein as RNA#1) was split at different positions within the loop to generate modulator RNAs named RNA#2, #4, #6, #8, and #10 and targeter RNAs named RNA#3, #5, #7, #9, and #11. The nucleotide sequences of these guide RNAs are shown in Table 4.
[0237] (Table 4) Nucleotide sequences of the single-guide RNAs and dual-guide RNAs tested TIFF0007689949000023.tif84170
[0238] These guide RNAs were chemically synthesized. The in vitro cleavage assay was performed using the method described above. Each guide RNA was used at a concentration of 1 μM when incubating with MAD7 to form RNP. The molar ratio of MAD7 to the target DNA was 10:1.
[0239] As shown in Fig. 4I, the nuclease activity of MAD7 was activated in the pairs of guide RNAs #2 and #3, #4 and #5, #6 and #7, #8 and #9, and #10 and #11, and the DNMT1 target DNA was cleaved. None of these targeter RNAs alone or modulator RNAs alone showed such activity. Therefore, under this condition, the position within the loop that split crRNA1 did not seem to affect the activity of the dual-guide RNA system.
[0240] Surprisingly, it was shown that the MAD7 nuclease was activated by a combination of any modulator RNA selected from RNA #2, #4, #6, #8, and #10 and any targeter RNA selected from RNA #3, #5, #7, #9, and #11 (Fig. 4I). In particular, the combination of RNA #4 and #11 did not contain the sequence derived from the loop of crRNA1, and the combination of RNA #10 and #5 contained the loop sequence of crRNA1 in both the modulator RNA and the targeter RNA. Therefore, under this condition, the loop or a fragment of the loop of the corresponding single-guide RNA was not essential for the dual-guide system. When the loop or loop fragment was present, its length did not seem to affect the activity of the dual-guide RNA system in either the targeter RNA or the modulator RNA.
[0241] Inclusion of additional hairpin sequences Next, a dual-guide RNA system containing a hairpin sequence at the 5’ end of the modulator RNA or the 3’ end of the targeter RNA was evaluated. As shown in FIGS. 4G - 4H, single-guide RNAs with a hairpin sequence added to the 5’ end or 3’ end of crRNA1, named RNA#12 and 14 respectively, were prepared. A modulator RNA corresponding to RNA#12 containing the hairpin sequence added to the 5’ end of crRNA1_modulator1 was designed and named RNA#13. A targeter RNA corresponding to RNA#14 containing the hairpin sequence added to the 3’ end of crRNA1_targeter1 was designed and named RNA#15. The nucleotide sequences of these guide RNAs are shown in Table 5. The hairpin sequences of the guide RNAs are underlined.
[0242] (Table 5) Nucleotide sequences of the single-guide RNAs and dual-guide RNAs tested TIFF0007689949000024.tif48170
[0243] These guide RNAs were chemically synthesized. The in vitro cleavage assay was performed using the method described above. Each guide RNA was used at a concentration of 1 μM when incubating with MAD7 to form an RNP. The molar ratio of MAD7 to the target DNA was 10:1.
[0244] As shown in Figure 4I, the nuclease activity of MAD7 was activated in single-guide RNAs #12 and #14 containing hairpins, and the DNMT1 target DNA was cleaved. In the corresponding modulator RNA #13 and targeter RNA #15 containing hairpin sequences at their 5'- and 3'-ends, respectively, such activity was not shown alone. However, when the modulator RNA #13 was combined with the targeter RNA #3 (described in the subsection "Loop" ends of modulator RNA and targeter RNA) to form a dual-guide system, the MAD7 nuclease was activated in this RNA pair. Similarly, when the targeter RNA #15 was combined with the modulator RNA #2 (described in the subsection "Loop" ends of modulator RNA and targeter RNA) to form a dual-guide system, the MAD7 nuclease was activated in this RNA pair. Notably, the combination of modulator RNA #13 and targeter RNA #15, each containing a pairpin sequence, also activated the MAD7 nuclease. Therefore, under this condition, the hairpin sequences added to the 5'-end of the modulator RNA or the 3'-end of the targeter RNA did not seem to negatively affect the activity of the dual-guide system.
[0245] Base pairing between modulator RNA and targeter RNA To evaluate the effect of base pairing between the modulator RNA and the targeter RNA on the activity of the dual-guide system, more single-guide systems and dual-guide systems were designed and tested. Specifically, the crRNA constructs were designed such that additional base pairing was introduced between the modulator RNA and the targeter RNA. The nucleotides in the modulator RNA that form such base pairs were placed on the 3' side of the modulator stem sequence, and the nucleotides in the targeter RNA that form such base pairs were placed on the 5' side of the targeter stem sequence. As shown in FIGS. 5A-5I, constructs 1 and 2 were identical to the above-described crRNA1 and crRNA2. The other constructs were made by splitting at any of the loop regions to create combinations 3, 5, 7, 9, 11, 13, and 15, or by splitting at any of the stem regions to create combinations 4, 6, 8, 10, 12, 14, and 16. The nucleotide sequences of these guide RNAs are shown in Table 6. The Gibbs free energy change (ΔG) of the corresponding crRNAs was calculated by the RNAfold program and is described in FIGS. 5A-5I.
[0246] (Table 6) Nucleotide sequences of the single-guide RNAs and dual-guide RNAs tested TIFF0007689949000025.tif202170
[0247] The guide RNAs were chemically synthesized. The in vitro cleavage assay was performed using the above method, except that the MAD7 protein was incubated with an equimolar amount of RNA(s) at 25°C for 20 minutes to form the RNP, and this RNP was incubated with the target DNA for 30 minutes. Each guide RNA was used at a concentration of 1 μM when incubating with MAD7 to form the RNP. The molar ratio of MAD7 to the target DNA was 10:1.
[0248] As shown in FIGS. 5J to 5K, when the crRNA was split within the stem region to form a dual guide, the activity of the CRISPR-Cas system was lost. However, when the crRNA was split within the loop region, the ability of the dual guide system to activate the MAD7 nuclease was reduced in a system that included additional base pairing between the modulator RNA and the targeter RNA.
[0249] Example 2. Cleavage of genomic DNA by the dual guide MAD7 CRISPR-Cas system In this example, the cleavage of genomic DNA of Jurkat cells using MAD7 in the form of a complex with single guide nucleic acid or dual guide nucleic acid is described.
[0250] Briefly, Jurkat cells were cultured in RPMI 1640 medium (Thermo Fisher Scientific, A1049101) supplemented with 10% fetal bovine serum at 37 °C in an environment of 5% CO 2 and split at a density of 100,000 cells / mL every 2 - 3 days. The MAD7 protein contained a nucleoplasmin NLS at the C-terminus, was expressed in E. coli, and purified by FPLC. The RNP complex was prepared by incubating 150 pmol of the MAD7 protein with 150 pmol of crRNA1 or a combination of 150 pmol of crRNA1_modulator1 and 150 pmol of crRNA1_targeter1 as described in Example 1 for 10 minutes at room temperature. This RNP was mixed with 200,000 Jurkat cells at a final volume of 25 μL. Electroporation was performed using program CA-137 in a 4D-Nucleofector (Lonza). After electroporation, the cells were cultured for 3 days.
[0251] The genomic DNA of the cells was extracted using Quick Extract DNA Extraction Solution 1.0 (Epicentre). The DNMT1 gene was amplified from the genomic DNA sample by PCR reaction. Amplified using a forward primer having the nucleotide sequence of TIFF0007689949000026 at position 12159 and a reverse primer having the nucleotide sequence of TIFF0007689949000027 at position 12158. The amplified DNA was purified and used as a template in a second PCR reaction using Nextera-indexed primers Index 1 and Index 2. The sequence of Index 1 is 4152, and the sequence of Index 2 is 5162, where i7 and i5 were used as barcodes for multiplexing. The PCR products were analyzed by next-generation sequencing, and the data were analyzed using the AmpliCan package (see Labun et al. (2019), Accurate analysis of genuine CRISPR editing events with ampliCan, GENOME RES., early online publication). The quality of the sequencing results was verified in Figure 6B. The editing efficiency was determined by the number of edited reads relative to the total number of reads obtained under each condition. The experiments were performed in duplicate.
[0252] As shown in Figure 6A, in the combination of crRNA1_modulator1 and crRNA1_targeter1 in the state of the complex with MAD7, 25-40% of the DNMT1 genomic locus was edited in the population of Jurkat cells. This observed efficiency was similar to the efficiency obtained by using crRNA1 and MAD7.
[0253] Example 3. Cleavage of other target sites by the dual-guide MAD7 CRISPR-Cas system Examples 1 and 2 describe the cleavage of target DNA having the sequence of the human DNMT1 gene. In this example, the cleavage of other target DNA using MAD7 in the state of the complex with dual-guide nucleic acids is described.
[0254] Briefly, crRNAs and corresponding targeter RNAs were designed to target other human genes. When such a targeter RNA is combined with crRNA1_modulator1, a dual-guide system can be created. The sequences of the guide RNAs used in this experiment are shown in Table 7. Also, guide RNAs targeting other human genes were designed.
[0255] (Table 7) Nucleotide sequences of exemplary single-guide RNAs and dual-guide RNAs TIFF0007689949000030.tif74170TIFF0007689949000031.tif207170
[0256] The guide RNAs were chemically synthesized. The intracellular cleavage assay was performed using the method described in Example 2.
[0257] As shown in Figure 7, at each target locus tested, the human genome was edited with the dual-guide RNA with the same efficiency as each single-guide RNA.
[0258] Example 4. Cleavage of other target sites by the dual-guide MAD7 CRISPR-Cas system using different cleavage sites within the crRNA loop This example describes the cleavage of DNA using MAD7 in the state of a complex with a dual-guide nucleic acid cleaved at different positions within the cRNA loop.
[0259] Briefly, crRNAs targeting CD52, PDCD1, and TIGIT, as well as the modulator RNA and targeter RNA of the dual-guide CRISPR system, were chemically synthesized. These RNA nucleotide sequences are shown in Table 8 below.
[0260] (Table 8) Nucleotide sequences of exemplary single-guide RNAs and dual-guide RNAs TIFF0007689949000032.tif171170TIFF0007689949000033.tif93170
[0261] In Table 8, crRNA_CD52, crRNA_PDCD1, and crRNA_TIGIT were used as single-guide RNAs targeting CD52, PDCD1, and TIGIT, respectively. crRNA_Modulator 1 was used as a dual-guide RNA corresponding to each single-guide RNA in combination with crRNA_CD52_Targeter 1, crRNA_PDCD1_Targeter 1, or crRNA_TIGIT_Targeter 1. In this case, the single-guide RNA was split by the first nucleotide bond from the 5'-end of the loop. crRNA_Modulator 2 was used as a dual-guide RNA corresponding to each single-guide RNA in combination with crRNA_CD52_Targeter 2, crRNA_PDCD1_Targeter 2, or crRNA_TIGIT_Targeter 2. In this case, the single-guide RNA was split by the second nucleotide bond from the 5'-end of the loop. crRNA_Modulator 3 was used as a dual-guide RNA corresponding to each single-guide RNA in combination with crRNA_CD52_Targeter 3, crRNA_PDCD1_Targeter 3, or crRNA_TIGIT_Targeter 3. In this case, the single-guide RNA was split by the third nucleotide bond from the 5'-end of the loop. crRNA_Modulator 4 was used as a dual-guide RNA corresponding to each single-guide RNA in combination with crRNA_CD52_Targeter 4, crRNA_PDCD1_Targeter 4, or crRNA_TIGIT_Targeter 4. In this case, the single-guide RNA was split by the fourth nucleotide bond from the 5'-end of the loop. crRNA_Modulator 5 was used as a dual-guide RNA corresponding to each single-guide RNA in combination with crRNA_CD52_Targeter 5, crRNA_PDCD1_Targeter 5, or crRNA_TIGIT_Targeter 5. In this case, the single-guide RNA was split by the fifth nucleotide bond from the 5'-end of the loop. The intracellular cleavage assay was performed using the method described in Example 1 above.
[0262] As shown in FIG. 8, for each of the target genes tested, in the dual-guide CRISPR system, the cell genome was edited with similar efficiency when the cleavage position was 2, 3, 4, or 5 in the intracellular cleavage assay, and with significantly lower efficiency when the cleavage position was 1 (i.e., cleavage at the first nucleotide bond from the 5' end of the loop). This result suggested that the modulator RNA may preferably contain at least one nucleotide (e.g., uridine) on the 3' side of the modulator stem sequence for optimal activity in cells.
[0263] Incorporation by reference The entire disclosure content of each of the patents and scientific documents mentioned herein is incorporated by reference for all purposes.
[0264] Equivalents The present invention can be embodied in other specific forms without departing from its spirit or essential characteristics. Therefore, the foregoing embodiments are to be considered illustrative rather than restrictive in all respects with respect to the invention described herein. Accordingly, the scope of the present invention is indicated by the appended claims rather than by the foregoing description, and it is intended that all modifications within the meaning and scope of equivalents of the claims be included in the claims.
Claims
**Claim 1** (a) (i) A spacer sequence designed to hybridize to a target nucleotide sequence; and (ii) A targeter stem sequence comprising a targeter nucleic acid; and (b) A modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence comprising an engineered, non-naturally occurring system, wherein the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids, and a complex comprising the targeter nucleic acid and the modulator nucleic acid can activate a type V-A CRISPR-associated (Cas) nuclease that is activated by a single crRNA without tracrRNA in a naturally occurring system, an engineered, non-naturally occurring system. **Claim 2** The engineered, non-naturally occurring system according to claim 1, wherein the targeter stem sequence and the modulator stem sequence are each 4 to 10 nucleotides in length. **Claim 3** The engineered, non-naturally occurring system according to claim 1 or 2, wherein the targeter stem sequence and the modulator stem sequence are each 5 nucleotides in length. **Claim 4** The engineered, non-naturally occurring system according to any one of claims 1 to 3, wherein the targeter stem sequence and the modulator stem sequence hybridize by Watson-Crick base pairing. **Claim 5** The engineered, non-naturally occurring system according to any one of claims 1 to 4, wherein the spacer sequence is about 20 nucleotides in length. **Claim 6** The engineered, non-naturally occurring system according to any one of claims 1 to 4, wherein the spacer sequence is 18 nucleotides in length or shorter. **Claim 7** The engineered, non-naturally occurring system according to claim 6, wherein the spacer sequence is 17 nucleotides in length or shorter. **Claim 8** The engineered, non-naturally occurring system according to any one of claims 1 to 7, wherein the targeter nucleic acid comprises a targeter stem sequence, a spacer sequence, and any additional nucleotide sequence, in that order from 5' to 3'. **Claim 9** The engineered, non-naturally occurring system according to any one of claims 1 to 8, wherein the targeter nucleic acid comprises ribonucleic acid (RNA). **Claim 10** The engineered, non-naturally occurring system according to claim 9, wherein the targeter nucleic acid comprises modified RNA. **Claim 11** The engineered, non-naturally occurring system of claim 9 or 10, wherein the targeter nucleic acid comprises a combination of RNA and DNA. **Claim 12** The engineered, non-naturally occurring system of any one of claims 9 to 11, wherein the targeter nucleic acid comprises a chemical modification. **Claim 13** The engineered, non-naturally occurring system of claim 12, wherein the chemical modification is present in one or more nucleotides at the 3' end of the targeter nucleic acid. **Claim 14** The engineered, non-naturally occurring system of claim 12 or 13, wherein the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof. **Claim 15** The engineered, non-naturally occurring system of any one of claims 1 to 14, wherein the modulator nucleic acid further comprises an additional nucleotide sequence. **Claim 16** The engineered, non-naturally occurring system of claim 15, wherein the additional nucleotide sequence is located 5' to the modulator stem sequence. **Claim 17** The engineered, non-naturally occurring system of claim 15 or 16, wherein the additional nucleotide sequence is 4 to 50 nucleotides in length. **Claim 18** The engineered, non-naturally occurring system of any one of claims 15 to 17, wherein the additional nucleotide sequence comprises a donor template recruit sequence that is capable of hybridizing to a donor template. **Claim 19** The engineered, non-naturally occurring system of claim 18, further comprising a donor template. **Claim 20** The engineered, non-naturally occurring system of any one of claims 1 to 19, wherein the modulator nucleic acid comprises one or more nucleotides 3' to the modulator stem sequence. **Claim 21** The engineered, non-naturally occurring system of any one of claims 1 to 20, wherein the modulator nucleic acid comprises RNA. **Claim 22** The engineered, non-naturally occurring system of claim 21, wherein the modulator nucleic acid comprises modified RNA. **Claim 23** The engineered, non-naturally occurring system of claim 21 or 22, wherein the modulator nucleic acid comprises a combination of RNA and DNA. **Claim 24** An engineered, non-naturally occurring system according to any one of claims 21 to 23, wherein the modulator nucleic acid comprises a chemical modification. **Claim 25** An engineered, non-naturally occurring system according to claim 24, wherein the chemical modification is present in one or more nucleotides at the 5' end of the modulator nucleic acid. **Claim 26** An engineered, non-naturally occurring system according to claim 24 or 25, wherein the chemical modification is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl, phosphorothioate, phosphorodithioate, pseudouridine, and any combination thereof. **Claim 27** An engineered, non-naturally occurring system according to any one of claims 1 to 26, wherein the targeter nucleic acid and the modulator nucleic acid are not covalently linked. **Claim 28** An engineered, non-naturally occurring system according to any one of claims 1 to 27, wherein the Cas nuclease comprises an amino acid sequence that is at least 80% identical to SEQ ID NO:
1. **Claim 29** An engineered, non-naturally occurring system according to any one of claims 1 to 28, wherein the Cas nuclease is Cpf1. **Claim 30** An engineered, non-naturally occurring system according to any one of claims 1 to 29, further comprising a Cas nuclease. **Claim 31** An engineered, non-naturally occurring system according to claim 30, wherein the targeter nucleic acid, the modulator nucleic acid, and the Cas nuclease are present within a ribonucleoprotein (RNP) complex. **Claim 32** An engineered, non-naturally occurring system according to any one of claims 1 to 31 comprising a eukaryotic cell that is not a human germ cell or a human embryonic cell. **Claim 33** An engineered, non-naturally occurring system according to any one of claims 1 to 31, or a eukaryotic cell according to claim 32 comprising a composition. **Claim 34** A method of cleaving a target DNA having a target nucleotide sequence, comprising: contacting the target DNA with an engineered, non-naturally occurring system according to any one of claims 1 to 31, thereby resulting in cleavage of the target DNA, wherein the contacting is performed in vitro. comprising a method. **Claim 35** A method of cleaving a target DNA having a target nucleotide sequence, comprising: A step of contacting a target DNA with an engineered, non-naturally occurring system according to any one of claims 1 to 31, thereby causing cleavage of the target DNA, wherein the contacting is performed ex vivo in a cell and the cell is not a human germ cell or a human embryonic cell. A method comprising this step. **Claim 36** The method according to claim 35, wherein the target DNA is genomic DNA of the cell. **Claim 37** The method according to claim 35 or 36, wherein the system is delivered into the cell as a pre-formed RNP complex. **Claim 38** The method according to claim 37, wherein the pre-formed RNP complex is delivered into the cell by electroporation. **Claim 39** An ex vivo or in vitro method for editing the genome of a eukaryotic cell, comprising: A step of delivering an engineered, non-naturally occurring system according to any one of claims 1 to 31 into a eukaryotic cell, thereby causing editing of the genome of the eukaryotic cell, wherein the cell is not a human germ cell or a human embryonic cell. A method comprising this step. **Claim 40** The method according to claim 39, wherein the system is delivered into the cell as a pre-formed RNP complex. **Claim 41** The method according to claim 39 or 40, wherein the system is delivered into the cell by electroporation. **Claim 42** The method according to any one of claims 35 to 41, wherein the cell is an immune cell. **Claim 43** The method according to claim 42, wherein the immune cell is a T lymphocyte.
Citation Information
Patent Citations
Modified site-directed modifying polypeptides and methods of use thereof
WO2017106569A1