Methods and compositions for increasing homologous directed repair
By combining the CRISPR/Cas system, CtBP interacting proteins, and 53BP1 protein inhibitors, CRISPR/Cas-mediated homology-directed repair was enhanced, solving the problem of difficulty in effectively editing genomes in existing technologies and achieving more efficient genome editing results.
Patent Information
- Application Number
- CN202480043139.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-23
AI Technical Summary
Existing CRISPR/Cas technology has difficulty effectively enhancing homology-directed repair (HDR) in mammalian cells, limiting the precision and efficiency of genome editing.
By combining the CRISPR/Cas system, the CtBP interacting protein (CtIP), and the inhibitor of 53BP1 protein (i53), and by using exogenous donor nucleic acid to perform homologous targeted repair with target DNA, targeted genetic modification is enhanced.
It significantly improves the efficiency of CRISPR/Cas-mediated homology-directed repair, thereby enhancing the accuracy and efficiency of genome editing.
Smart Images

Figure CN121399263A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims the benefit of U.S. Application No. 63 / 511,361, filed June 30, 2023, which is incorporated by reference herein in its entirety for all purposes.
[0002] Reference to a Sequence Listing Filed as XML document through EFS WEB The sequence listing written in file 614575SEQLIST.xml is 76,084 bytes, was created on June 27, 2024, and is hereby incorporated by reference. BACKGROUND
[0003] The development of CRISPR / Cas technology has provided an efficient method to introduce site-specific modifications in mammalian genomes, offering great potential for research and treatment of a wide range of genetic diseases. Currently, the most commonly used CRISPR system in genome engineering uses a single guide RNA (sgRNA) and a CRISPR-associated endonuclease (Cas9), which creates a double-strand break (DSB) at the targeted sequence. The two major DSB repair pathways in mammalian cells are (i) error-prone non-homologous end joining (NHEJ) and (ii) faithful homology-directed repair (HDR), which is limited to the S and G2 phases of the cell cycle and depends on the availability of a repair template carrying the modification to be introduced. Methods to enhance HDR would be useful to provide better and more efficient ways to perform precise genome editing. SUMMARY
[0004] Provided herein are combinations comprising a CRISPR / Cas system, an inhibitor of CtBP-interacting protein, and 53BP1, for use in enhancing homology-directed repair of a CRISPR / Cas-mediated cleavage of a target DNA by an exogenous donor nucleic acid. Also provided are methods of making targeted genetic modifications in a cell using such combinations by homology-directed repair of a CRISPR / Cas-mediated cleavage at a target genomic locus in the cell.
[0005] In one aspect, methods are provided for targeted genetic modification via homologous directed repair at target genomic loci in cells. Some such methods involve administering to cells: (a) a clustered regularly spaced short palindromic repeat (CRISPR)-associated (Cas) protein or nucleic acid encoding a Cas protein; (b) a guide RNA or one or more DNA sequences encoding a guide RNA, wherein the guide RNA contains one or more adaptor-binding elements to which an adaptor protein can specifically bind, and wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) a fusion protein or nucleic acid encoding a fusion protein, wherein the fusion protein contains a CtBP-interacting protein (CtIP) fused to the adaptor protein; (d) An inhibitor of the 53BP1 (i53) protein or a nucleic acid encoding the i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homologous arm hybridizing to a 5' target sequence at a target genomic locus and a 3' homologous arm hybridizing to a 3' target sequence at a target genomic locus, optionally wherein the 5' and 3' homologous arms are laterally inserted into the nucleic acid, wherein the Cas protein and the guide RNA form a complex, the Cas protein cleaves the guide RNA target sequence to produce a double-strand break, and the exogenous donor nucleic acid recombines with the target genomic locus via homologous directed repair to produce a targeted genetic modification.
[0006] In some such methods, the Cas protein is administered to cells in the form of a protein, optionally wherein the Cas protein is contained in lipid nanoparticles. In some such methods, a nucleic acid encoding the Cas protein is administered to cells, wherein the nucleic acid encoding the Cas protein comprises RNA encoding the Cas protein, optionally wherein the RNA encoding the Cas protein is contained in lipid nanoparticles. In some such methods, a nucleic acid encoding the Cas protein is administered to cells, wherein the nucleic acid encoding the Cas protein comprises DNA encoding the Cas protein, optionally wherein the DNA encoding the Cas protein is contained in a viral vector, optionally wherein the viral vector is a recombinant adeno-associated virus (AAV) vector. In some such methods, the Cas protein is the Cas9 protein. In some such methods, the Cas9 protein is *Streptococcus pyogenes* (…). Streptococcus pyogenes Cas9 protein, Campylobacter jejuni ( Campylobacter jejuni Cas9 protein or Staphylococcus aureus ( Staphylococcus aureus Cas9 protein, optionally wherein the Cas9 protein is the Streptococcus pyogenes Cas9 protein.
[0007] In some such methods, the guide RNA is administered in the form of RNA, optionally wherein the guide RNA is in a lipid nanoparticle. In some such methods, one or more DNAs encoding the guide RNA are administered to the cell, optionally wherein the one or more DNAs encoding the guide RNA are in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such methods, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind. In some such methods, the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA. In some such methods, the guide RNA is a single guide RNA comprising a CRISPR RNA (crRNA) portion partially fused to a trans-activating CRISPR RNA (tracrRNA) portion, and wherein the first loop is a four loop corresponding to residues 13 to 16 of SEQ ID NO: 11, 13, 15, or 16, and the second loop is a stem loop 2 corresponding to residues 53 to 56 of SEQ ID NO: 11, 13, 15, or 16. In some such methods, the adaptor binding elements comprise the sequences set forth in SEQ ID NO: 19 or 20. In some such methods, the guide RNA comprises the sequence set forth in SEQ ID NO: 21, 22, 23, 24, 25, or 26.
[0008] In some such methods, the fusion protein is administered to the cell in the form of a protein, optionally wherein the fusion protein is in a lipid nanoparticle. In some such methods, a nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, optionally wherein the RNA encoding the fusion protein is in a lipid nanoparticle. In some such methods, a nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises a DNA encoding the fusion protein, optionally wherein the DNA encoding the fusion protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such methods, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof. In some such methods, the adaptor protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 32. In some such methods, the adaptor protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33. In some such methods, the CtIP protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 34. In some such methods, the CtIP protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 36. In some such methods, the fusion protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 30. In some such methods, the fusion protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31.
[0009] In some such methods, the i53 protein is administered to the cell in the form of a protein, optionally wherein the i53 protein is in a lipid nanoparticle. In some such methods, a nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA messenger RNA encoding the i53 protein, optionally wherein the RNA encoding the i53 protein is in a lipid nanoparticle. In some such methods, a nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises a DNA encoding the i53 protein, optionally wherein the DNA encoding the i53 protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such methods, the i53 protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 40 or 42. In some such methods, the i53 protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 41 or 43.
[0010] In some such methods, the exogenous donor nucleic acid comprises an insert nucleic acid. In some such methods, the exogenous donor nucleic acid is in a viral vector. In some such methods, the viral vector is a recombinant AAV vector. In some such methods, the exogenous donor nucleic acid is a large targeting vector (LTVEC), wherein: (a) the LTVEC is at least 10 kb; (b) the sum of the 5' homology arm and the 3' homology arm of the LTVEC is at least 10 kb; (c) the LTVEC is about 50 kb to about 300 kb; or (d) the sum of the 5' homology arm and the 3' homology arm of the LTVEC is about 10 kb to about 200 kb.
[0011] In some such methods, a nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein a nucleic acid encoding the i53 protein is administered to the cell, the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle. In some such methods, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, a nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, a nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
[0012] In some such methods, the guide RNA comprises two adaptor binding elements with which the adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, and the RNA encoding the i53 protein are in a lipid nanoparticle, wherein the one or more DNAs encoding the guide RNA are administered to the cell, and wherein the one or more DNAs encoding the guide RNA and the exogenous donor nucleic acid are in a recombinant AAV vector.
[0013] In some such methods, the guide RNA comprises two adaptor binding elements with which the adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, wherein the guide RNA is administered to the cell in the form of RNA, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, the RNA encoding the i53 protein, and the guide RNA are in a lipid nanoparticle, wherein the exogenous donor nucleic acid is in a recombinant AAV vector.
[0014] In some such methods, the cell is a mammalian cell. In some such methods, the cell is a rodent cell. In some such methods, the cell is a mouse cell or a rat cell. In some such methods, the cell is a mouse cell. In some such methods, the cell is a human cell. In some such methods, the tissue is in vitro. In some such methods, the tissue is in vivo.
[0015] In another aspect, compositions or combinations are provided. Some such compositions or combinations comprise (a) a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) protein or a nucleic acid encoding the Cas protein; (b) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises one or more adaptor binding elements to which an adaptor protein is capable of specifically binding, and wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) a fusion protein or a nucleic acid encoding the fusion protein, wherein the fusion protein comprises a CtBP-interacting protein (CtIP) fused to an adaptor protein; (d) an inhibitor of a 53BP1 (i53) protein or a nucleic acid encoding the i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, optionally wherein the 5' homology arm and the 3' homology arm flank an insertion nucleic acid.
[0016] In some such compositions or combinations, the composition or combination comprises the Cas protein in the form of a protein, optionally wherein the Cas protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, optionally wherein the RNA encoding the Cas protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises a DNA encoding the Cas protein, optionally wherein the DNA encoding the Cas protein is in a viral vector, optionally wherein the viral vector is a recombinant adeno-associated virus (AAV) vector. In some such compositions or combinations, the Cas protein is a Cas9 protein. In some such compositions or combinations, the Cas9 protein is a S. pyogenes Cas9 protein, a C. jejuni Cas9 protein, or a S. aureus Cas9 protein, optionally wherein the Cas9 protein is a S. pyogenes Cas9 protein.
[0017] In some such compositions or combinations, the composition or combination comprises a guide RNA in the form of an RNA, optionally wherein the guide RNA is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises one or more DNAs encoding a guide RNA, optionally wherein the one or more DNAs encoding a guide RNA are in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such compositions or combinations, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind. In some such compositions or combinations, the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA. In some such compositions or combinations, the guide RNA is a single guide RNA comprising a CRISPR RNA (crRNA) portion partially fused to a trans-activating CRISPR RNA (tracrRNA) portion, and wherein the first loop is a four-loop corresponding to residues 13 to 16 of SEQ ID NO: 11, 13, 15, or 16, and the second loop is a stem loop 2 corresponding to residues 53 to 56 of SEQ ID NO: 11, 13, 15, or 16. In some such compositions or combinations, the adaptor binding elements comprise the sequences set forth in SEQ ID NO: 19 or 20. In some such compositions or combinations, the guide RNA comprises the sequence set forth in SEQ ID NO: 21, 22, 23, 24, 25, or 26.
[0018] In some such compositions or combinations, the composition or combination comprises a fusion protein in the form of a protein, optionally wherein the fusion protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, optionally wherein the RNA encoding the fusion protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises a DNA encoding the fusion protein, optionally wherein the DNA encoding the fusion protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such compositions or combinations, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof. In some such compositions or combinations, the adaptor protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 32. In some such compositions or combinations, the adaptor protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33. In some such compositions or combinations, the CtIP protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 34. In some such compositions or combinations, the CtIP protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 36. In some such compositions or combinations, the fusion protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 30. In some such compositions or combinations, the fusion protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31.
[0019] In some such compositions or combinations, the composition or combination comprises an i53 protein in the form of a protein, optionally wherein the i53 protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding an i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA messenger RNA encoding the i53 protein, optionally wherein the RNA encoding the i53 protein is in a lipid nanoparticle. In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding an i53 protein, wherein the nucleic acid encoding the i53 protein comprises a DNA encoding the i53 protein, optionally wherein the DNA encoding the i53 protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector. In some such compositions or combinations, the i53 protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 40 or 42. In some such compositions or combinations, the i53 protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 41 or 43.
[0020] In some such compositions or combinations, the exogenous donor nucleic acid comprises an insert nucleic acid. In some such compositions or combinations, the exogenous donor nucleic acid is in a viral vector. In some such compositions or combinations, the viral vector is a recombinant AAV vector. In some such compositions or combinations, the exogenous donor nucleic acid is a large targeting vector (LTVEC), wherein: (a) the LTVEC is at least 10 kb; (b) the sum of the 5' homology arm and the 3' homology arm of the LTVEC is at least 10 kb; (c) the LTVEC is about 50 kb to about 300 kb; or (d) the sum of the 5' homology arm and the 3' homology arm of the LTVEC is about 10 kb to about 200 kb.
[0021] In some such compositions or combinations, the composition or combination comprises a nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, the composition or combination comprises a nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle. In some such compositions or combinations, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises a nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, the composition or combination comprises a nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
[0022] In some such compositions or combinations, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the composition or combination comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, wherein the composition or combination comprises a nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the composition or combination comprises a nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, and the RNA encoding the i53 protein are in a lipid nanoparticle, wherein the composition or combination comprises one or more DNAs encoding the guide RNA, and wherein the one or more DNAs encoding the guide RNA and the exogenous donor nucleic acid are in a recombinant AAV vector.
[0023] In some such compositions or combinations, the guide RNA comprises two adaptor-binding elements to which an adaptor protein can specifically bind, wherein a first adaptor-binding element is within a first loop of the guide RNA and a second adaptor-binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 capsid protein or a functional fragment or variant thereof, wherein the composition or combination comprises a nucleic acid encoding a Cas protein, wherein the nucleic acid encoding the Cas protein comprises RNA encoding the Cas protein, wherein the composition or combination comprises a nucleic acid encoding a fusion protein, wherein the nucleic acid encoding the fusion protein comprises RNA encoding the fusion protein, wherein the composition or combination comprises a nucleic acid encoding an i53 protein, wherein the nucleic acid encoding the i53 protein comprises RNA encoding the i53 protein, wherein the composition or combination comprises a guide RNA in the form of RNA, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, the RNA encoding the i53 protein, and the guide RNA are in lipid nanoparticles, and wherein the exogenous donor nucleic acid is in a recombinant AAV vector. Attached Figure Description
[0024] Figure 1 Targeting was displayed LMNA A schematic diagram of the insertion of a CRISPR-mediated homology-directed repair (HDR) reporter into a gene. In this assay, the coding sequence for the mClover fluorescent protein was designed to integrate into lamin A via HDR. LMNA The mClover-LMNA fusion protein is located at the 5' end of the gene. The resulting expression and different localizations can then be observed and quantified microscopically. An AAV2 vector (AAVsgLMNA+mClover) containing an mClover sequence with an appropriate homologous arm (HA) side-joined to express sgRNA targeting LMNA (sgLMNA) was used as the HDR donor template. The donor template is side-joined with an sgLMNA recognition sequence, which allows for donor linearization during intracellular sgLMNA expression.
[0025] Figures 2A-2C Showing the use Figure 1 The assay shown is an assessment of baseline HDR in HEK293 cells expressing Cas9 as the MOI of the AAV2 HDR template increases. Figure 2A The mClover and Hoechst staining used to assess the percentage of mClover-LMNA positive cells is shown. Figure 2B The HDR efficiency is shown as measured by the percentage of mClover-LMNA positive cells. Figure 2C The data shown are from negative controls, including co-treatment with mirin or with donor templates without homologous arms.
[0026] Figures 3A-3BScreening potential HDR enhancers for selectively supporting HDR repair at CRISPR / Cas9-mediated double-strand breaks in Cas9-expressing HEK293 cells is shown. Figure 3A A schematic showing the use of the MS2 tagging method to recruit potential HDR enhancers to Cas9 cleavage sites is shown. In this method, the scaffold sequence of the sgRNA is modified to include MS2 phage aptamers (sgRNA2.0) and the HDR enhancement protein is fused to MS2 coat protein (MS2) to interact with sgRNA2.0. Figure 3B A schematic showing the use of the MS2 tagging method to recruit potential HDR enhancers to Cas9 cleavage sites is shown. In this method, the scaffold sequence of the sgRNA is modified to include MS2 phage aptamers (sgRNA2.0) and the HDR enhancement protein is fused to MS2 coat protein (MS2) to interact with sgRNA2.0. Figure 1 The effect on HDR efficiency as measured by the percentage of mClover-LMNA positive cells is shown using the assay shown in FIG. 1.
[0027] Figures 4A-4C The effect of MS2-CtIP mRNA on HDR efficiency in Cas9-expressing HEK293 cells as the amount of MS2-CtIP mRNA packaged into LNP is increased is shown, demonstrating that CtIP-mediated DNA end resection promotes cell patterning towards HDR. Figure 4A mClover and Hoechst staining to assess the percentage of mClover-LMNA positive cells is shown. Figure 4B HDR efficiency as measured by the percentage of mClover-LMNA positive cells and the percentage of cells with unwanted indels resulting from repair via non-homologous end joining is shown. All “HR efficiency” plots show absolute mCLOVER % from 3 experimental replicates. The box extends from the 25th to 75th percentile. The line in the middle of the box is drawn at the median. The “+” is drawn at the mean. Whiskers fall to the minimum and rise to the maximum. To ensure consistency, all microscopy images were taken from the same replicate (1 of 3 experimental replicates) and mCLOVER % from each condition is shown in the lower right of the corresponding image. Figure 4C The effect of MS2-CtIP on HDR efficiency using sgRNAs containing MS2 binding loops (2.0) or sgRNAs without MS2 binding loops (REG) is shown.
[0028] Figure 5 A comparison of HDR efficiency as measured by the percentage of mClover-LMNA positive cells when using plasmid delivery of MS2-CtIP or LNP delivery of MS2-CtIP mRNA in Cas9-expressing HEK293 cells is shown.
[0029] Figures 6A-6BThe study showed the effect of increased i53 mRNA levels packaged into LNPs on HDR efficiency in Cas9-expressing HEK293 cells, demonstrating that 53BP1 inhibition provides a pre-removal environment at double-strand breaks and further promotes HDR. Figure 6A The mClover and Hoechst staining used to assess the percentage of mClover-LMNA positive cells is shown. Figure 6B The HDR efficiency is shown as the percentage of mCLOVER-LMNA positive cells and the percentage of cells with unwanted insertion loss due to repair via non-homologous end joining. All “HR Efficiency” plots show the absolute mCLOVER cell percentage from 3 experimental replicates. The box extends from the 25th percentile to the 75th percentile. The line in the middle of the box is drawn at the median. “+” is drawn with the mean. Whiskers decrease to the minimum and increase to the maximum. To ensure consistency, all microscopic images were taken from the same replicate (1 of 3 experimental replicates), and the mCLOVER% from each condition is shown in the lower right of the corresponding image.
[0030] Figure 7 The combined effect of MS2-CtIP and i53 on enhanced HDR efficiency in Cas9-expressing HEK293 cells was shown, demonstrating that co-expression of i53 and MS2-CtIP significantly increased CRISPR-stimulated HDR. All “HR Efficiency” plots show the absolute mCLOVER% from three experimental replicates. The boxes extend from the 25th percentile to the 75th percentile. Lines in the middle of the boxes are drawn at the median. “+” is drawn with the mean. Whiskers decrease to the minimum and increase to the maximum. To ensure consistency, all microscopic images were taken from the same replicate (one of the three experimental replicates), and the mCLOVER% from each condition is shown in the lower right of the corresponding image.
[0031] Figure 8 The combined effects of MS2-CtIP and i53 on HDR efficiency in Cas9-expressing HEK293 cells were shown in relation to AZD7648 (a DNA-PKcs inhibitor).
[0032] Figures 9A-9B Displayed in AAV sgLMNA+mClover The combined effect of MS2-CtIP and i53 on HDR efficiency in HEK293 cells at increased doses. Two LNPs were used: LNPs encapsulating Cas9, MS2-CtIP, and i53 mRNA. booster And LNPs that encapsulate Cas9 and mCherry mRNA baseline .exist Figure 9AIn a first experiment, LNP transfection was performed into HEK293 cells transduced with AAV sgLMNA+mClover transduced HEK293 cells. In Figure 9B In a second experiment, guide RNAs targeting the C-terminus of two other genes HMGA1 and SEC61B were used in a repeat experiment.
[0033] Figure 10 The combined effect of MS2-CtIP and i53 on HDR efficiency in Cas9 expressing HEK293 cells in the context of non-viral donor delivery, in particular linear blunt ended dsDNA containing the mClover coding sequence, flanked by LMNA HA sequences as described above (dsDNA mclover-LMNA ), delivered via electroporation together with sgLMNA2.0 in the form of a plasmid, is shown.
[0034] Figure 11 The combined effect of MS2-CtIP and i53 on HDR efficiency in Cas9 expressing HEK293 cells in the context of non-viral donor delivery, in particular linear blunt ended dsDNA containing the mClover coding sequence, flanked by LMNA HA sequences as described above (dsDNA mclover-LMNA ), delivered via electroporation together with sgLMNA2.0 in the form of RNA, is shown.
[0035] Definitions The terms "protein," "polypeptide," and "peptide," used interchangeably herein, encompass polymeric forms of amino acids of any length, including coded and non-coded amino acids, and including those that have been modified, or those that have been derivatized, chemically or biologically. The terms also include polymers that have been modified, such as polypeptides with modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide that has a particular function or structure.
[0036] The terms "nucleic acid" and "polynucleotide," used interchangeably herein, encompass polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These terms include single-stranded, double-stranded, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers comprising purine, pyrimidine, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0037] The term "expression vector" or "expression construct" or "expression cassette" refers to a recombinant nucleic acid containing a desired coding sequence operably linked to the appropriate nucleic acid sequences necessary for expression of the coding sequence in a particular host cell or organism. Nucleic acid sequences necessary for expression in prokaryotes typically include a promoter, an operator (optional), and a ribosome binding site, among other sequences. It is well known that eukaryotic cells utilize promoters, enhancers, as well as termination signals and polyadenylation signals, although some elements can be deleted and others added without sacrificing essential expression.
[0038] The term "viral vector" refers to a recombinant nucleic acid comprising at least one element of viral origin and comprising elements sufficient or allowing for packaging into a viral vector particle. The vector and / or particle can be used for the purpose of transferring DNA, RNA, or other nucleic acids into cells, either ex vivo or in vivo. Various forms of viral vectors are known.
[0039] The term "isolated" with respect to proteins, nucleic acids, and cells includes proteins, nucleic acids, and cells that are relatively pure relative to other cellular or organismal components that would normally be present in situ, up to and including a substantially pure preparation of the protein, nucleic acid, or cell. The term "isolated" can include proteins and nucleic acids that do not have a naturally occurring counterpart, or proteins or nucleic acids that have been chemically synthesized and thus are substantially free of other proteins or nucleic acids. The term "isolated" can include proteins, nucleic acids, or cells that have been separated or purified from most other cellular components or organismal components that are naturally associated with the protein, nucleic acid, or cell, such as, but not limited to, other cellular proteins, nucleic acids, or cellular or extracellular components.
[0040] The term "wild type" includes an entity having a structure and / or activity as found in the normal (as compared to mutant, diseased, altered, etc.) state or condition. Wild type genes and polypeptides typically exist in a variety of different forms (e.g., alleles).
[0041] The term "endogenous sequence" refers to a nucleic acid sequence that naturally occurs in a cell or animal in vivo. For example, an endogenous sequence of an animal is a nucleic acid sequence that naturally occurs in a gene locus of the animal. Rosa26 The term "native sequence" refers to a native nucleic acid sequence that naturally occurs in a gene locus of an animal. Rosa26 The term "native sequence" refers to a native nucleic acid sequence that naturally occurs in a gene locus of an animal. Rosa26 The term "native sequence" refers to a native nucleic acid sequence that naturally occurs in a gene locus of an animal.
[0042] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in the cell in that form, or that is introduced into the cell from an external source. Normal presence includes presence with respect to a particular developmental stage of the cell and environmental conditions. For example, an exogenous molecule or sequence can include a mutated version of a corresponding endogenous sequence within the cell (such as a humanized version of an endogenous sequence), or can include a sequence that corresponds to an endogenous sequence within the cell but is not in the form (i.e., not within a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form in a particular cell at a particular developmental stage under particular environmental conditions.
[0043] The term "heterologous" when used in the context of a nucleic acid or protein indicates that the nucleic acid or protein comprises at least two segments that are not naturally found together in the same molecule. For example, the term "heterologous" when used with respect to a nucleic acid segment or a protein segment indicates that the nucleic acid or protein comprises two or more subsequences that are not found in the same relationship (e.g., joined together) in nature. As one example, a "heterologous" region of a nucleic acid vector is a nucleic acid segment that is within or attached to another nucleic acid molecule that is not associated with additional molecules in nature. For example, a heterologous region of a nucleic acid vector can comprise a coding sequence flanked by heterologous promoters that are not associated with the coding sequence in nature. Likewise, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule (e.g., a fusion protein or a protein with a tag) that is not associated with additional peptide molecules in nature. Similarly, a nucleic acid or protein can comprise a heterologous marker or a heterologous secretion or localization sequence.
[0044] " Codon optimization" takes advantage of the degeneracy of the code, as demonstrated by the variety of three-base pair codon combinations that specify an amino acid, and generally includes the process of modifying a nucleic acid sequence to enhance expression in a particular host cell by replacing at least one codon of the natural sequence with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the natural amino acid sequence. For example, a nucleic acid encoding a protein can be modified to replace codons that have a higher frequency of use in a given prokaryotic or eukaryotic cell (including a bacterial cell, a yeast cell, a human cell, a non-human cell, a mammalian cell, a rodent cell, a mouse cell, a rat cell, a hamster cell, or any other host cell) compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, at the "Codon Usage Database." These tables can be modified in a variety of ways. See Nakamura et al. (2000) Nucleic Acids Res. 28:292, which is incorporated by reference herein in its entirety for all purposes. Computer algorithms are also available for codon optimization of particular sequences for expression in a particular host (see, e.g., Gene Forge). Nucleic Acids Research 28:292, which is incorporated by reference herein in its entirety for all purposes. Computer algorithms are also available for codon optimization of particular sequences for expression in a particular host (see, e.g., Gene Forge).
[0045] The term "locus" refers to a specific location on a chromosome within an organism's genome, representing a gene (or significant sequence), DNA sequence, or polypeptide coding sequence. For example, " Rosa26 "Locus" can refer to Rosa26 Gene, Rosa26 Specific location of DNA sequence or Rosa26 Location on chromosomes within an organism's genome, a location that has been identified as the site where such sequences reside. Rosa26 "Locus" may include Rosa26 Regulatory elements of a gene include, for example, enhancers, promoters, 5' and / or 3' untranslated regions (UTRs), or combinations thereof.
[0046] The term "gene" refers to a DNA sequence in a chromosome that, if naturally present, may contain at least one coding region and at least one non-coding region. The DNA sequence of a chromosome encoding a product (e.g., but not limited to RNA products and / or polypeptide products) may contain a coding region interrupted by non-coding introns and a sequence located adjacent to the coding region at both the 5' and 3' ends such that the gene corresponds to a full-length mRNA sequence (containing the 5' and 3' untranslated sequences). Additionally, other non-coding sequences, including regulatory sequences (e.g., but not limited to promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions, may be present in a gene. These sequences may be located near the coding region of the gene (e.g., but not limited to within 10 kb) or at distant sites, and these sequences may influence the level or rate of transcription and translation of the gene.
[0047] A promoter is a regulatory region of DNA that typically contains a TATA box that guides RNA polymerase II to initiate RNA synthesis at the appropriate transcription start site of a specific polynucleotide sequence. Promoters may additionally contain other regions that influence the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of operatively linked polynucleotides. Promoters may be active in one or more cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, single-cell embryos, differentiated cells, or combinations thereof). Promoters may be, for example, constitutively active promoters, conditional promoters, inducible promoters, time-restricted promoters (e.g., developmentally regulated promoters), or spatially restricted promoters (e.g., cell-specific or tissue-specific promoters). Examples of promoters can be found, for example, in WO 2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.
[0048] A constitutive promoter is a promoter that is active in all tissues or in a particular tissue at all stages of development. Examples of constitutive promoters include the human cytomegalovirus immediate early (hCMV) promoter, the mouse cytomegalovirus immediate early (mCMV) promoter, the human elongation factor 1 alpha (hEF1a) promoter, the mouse elongation factor 1 alpha (mEF1a) promoter, the mouse phosphoglycerate kinase (PGK) promoter, the chicken beta actin hybrid (CAG or CBh) promoter, the SV40 early promoter, and the beta 2 tubulin promoter.
[0049] Examples of inducible promoters include, for example, chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include, for example, alcohol regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline regulated promoters (e.g., tetracycline-responsive promoters, tetracycline operator sequences (tetO), tet-On promoters, or tet-Off promoters), steroid regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoters, or ecdysone receptor promoters), or metal regulated promoters (e.g., metallothionein promoters). Physically regulated promoters include, for example, temperature regulated promoters (e.g., heat shock promoters) and light regulated promoters (e.g., light inducible promoters or light repressible promoters).
[0050] Tissue specific promoters can be, for example, neuron specific promoters or glia specific promoters or muscle specific promoters.
[0051] Developmentally regulated promoters include, for example, promoters that are active only during embryonic development or only in adult cells.
[0052] “Operably linked” or “operably connecting” includes positioning two or more components (e.g., a promoter and another sequence element) so that the two components function normally and allow at least one component to be able to mediate a function imposed on at least one other component. For example, a promoter can be operably linked to a coding sequence if the promoter controls the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include such sequences adjacent to each other or acting in trans (e.g., a regulatory sequence can act at a distance to control transcription of a coding sequence).
[0053] The methods and compositions provided herein employ a number of different components. Some components throughout the specification can have active variants and fragments. The term “functional” refers to the innate ability of a protein or nucleic acid (or fragment or variant thereof) to exhibit a biological activity or function. The biological function of a functional fragment or variant can be the same as or can actually be altered (e.g., with respect to its specificity or selectivity or efficacy) compared to the original molecule but retains the essential biological function of the molecule.
[0054] The term "variant" refers to a nucleotide sequence that differs from the most common sequence in the population (e.g., by one nucleotide) or a protein sequence that differs from the most common sequence in the population (e.g., by one amino acid).
[0055] When referring to proteins, the term "fragment" means a protein that is shorter than a full-length protein or has fewer amino acids. When referring to nucleic acids, the term "fragment" means a nucleic acid that is shorter than a full-length nucleic acid or has fewer nucleotides. When referring to protein fragments, a fragment can be, for example, an N-terminal fragment (i.e., a portion of the C-terminus of a protein has been removed), a C-terminal fragment (i.e., a portion of the N-terminus of a protein has been removed), or an internal fragment (i.e., a portion of both the N-terminus and C-terminus of a protein has been removed). When referring to nucleic acid fragments, a fragment can be, for example, a 5' fragment (i.e., a portion of the 3' end of a nucleic acid has been removed), a 3' fragment (i.e., a portion of the 5' end of a nucleic acid has been removed), or an internal fragment (i.e., a portion of both the 5' end and 3' end of a nucleic acid has been removed).
[0056] In the context of two polynucleotide or polypeptide sequences, "sequence identity" or "identity" refers to the same residues in two sequences when compared for maximum correspondence within a specified comparison window. When using a percentage of sequence identity relative to a protein, dissimilar residue positions are often distinguished by conserved amino acid substitutions, where an amino acid residue is replaced by another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the molecule's functional properties. When the conserved substitutions of sequences differ, the percentage of sequence identity can be adjusted upwards to correct for the conservatism of the substitution. Sequences that differ due to such conserved substitutions are considered to have "sequence similarity" or "similarity." The means of making this adjustment are well known. Typically, this involves counting conserved substitutions as partial mismatches rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, a score for a conserved substitution is between zero and 1 when the score for an identical amino acid is 1 and the score for a non-conservative substitution is zero. For example, the score for a conserved substitution is calculated, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).
[0057] "Percentage of sequence identity" includes the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window can comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percentage of sequence identity. The comparison window is the length of the shorter sequence over which the comparison is made. Unless otherwise specified, the sequence identity / similarity values include those generated using GAP Version 10 using a GAP weight of 50 and length weight of 3, and the nwsgapdna.cmp scoring matrix for nucleotide sequence % identity and similarity; a GAP weight of 8 and length weight of 2, and the BLOSUM62 scoring matrix for amino acid sequence % identity and similarity; or any equivalent program. "Equivalent program" includes any sequence comparison program that, when compared to the corresponding alignment generated by version 10 of GAP, generates an alignment having the same number of nucleotides or amino acid residues matched and the same percentage of sequence identity for any two sequences in question.
[0058] Unless otherwise specified, sequence identity / similarity values include those generated using GAP Version 10 using a GAP weight of 50 and length weight of 3, and the nwsgapdna.cmp scoring matrix for nucleotide sequence % identity and similarity; a GAP weight of 8 and length weight of 2, and the BLOSUM62 scoring matrix for amino acid sequence % identity and similarity; or any equivalent program. "Equivalent program" includes any sequence comparison program that, when compared to the corresponding alignment generated by version 10 of GAP, generates an alignment having the same number of nucleotides or amino acid residues matched and the same percentage of sequence identity for any two sequences in question.
[0059] The term "conservative amino acid substitution" refers to the replacement of an amino acid occurring in a sequence with a different amino acid having similar size, charge, or polarity. Examples of conservative substitutions include the replacement of one non-polar (hydrophobic) residue such as isoleucine, valine, or leucine with another non-polar residue. Likewise, examples of conservative substitutions include the replacement of one polar (hydrophilic) residue for another such as the replacement of a polar residue between arginine and lysine, between glutamine and asparagine, or between glycine and serine. Additionally, the replacement of one basic residue such as lysine, arginine, or histidine with another basic residue or the replacement of one acidic residue such as aspartic acid or glutamic acid with another acidic residue are further examples of conservative substitutions. Examples of non-conservative substitutions include the replacement of a non-polar (hydrophobic) amino acid residue such as isoleucine, valine, leucine, alanine, or methionine with a polar (hydrophilic) residue such as cysteine, glutamine, glutamic acid, or lysine and / or the replacement of a polar residue with a non-polar residue. A typical classification of amino acids is summarized as follows.
[0060] Table 1. Amino acid classification
[0061] “Homologous” sequences (e.g., nucleic acid sequences) include sequences that are identical or substantially similar to a known reference sequence, such that they are, for example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous sequences and paralogous sequences. For example, homologous genes are typically produced from a common ancestral DNA sequence by either a speciation event (orthologous genes) or a genetic duplication event (paralogous genes). “Orthologous” genes include genes in different species that have evolved from a common ancestor. Orthologs typically retain the same function over the course of evolution. “Paralogous” genes include genes associated with duplication within a genome. Paralogs can evolve new functions over the course of evolution.
[0062] The term “in vitro” includes artificial environments as well as processes or reactions that occur within artificial environments (e.g., test tubes or isolated cells or cell lines). The term “in vivo” includes natural environments (e.g., cells, organisms, or bodies) as well as processes or reactions that occur within natural environments. The term “ex vivo” includes cells that have been removed from an individual’s body as well as processes or reactions that occur within such cells.
[0063] Compositions or methods that “comprise” or “contain” one or more recited elements can include other non-recited elements. For example, a composition that “comprises” or “contains” a protein can contain the protein alone or the protein in combination with other ingredients. The transitional phrase “consisting essentially of” means that the scope of a claim should be interpreted as encompassing the specified elements recited in the claim and those that do not materially affect the basic and novel characteristic of the claimed application. Thus, the term “consisting essentially of’ should not be interpreted as equivalent to “comprising” when used in the claims of the present application.
[0064] “Optional” or “optionally” means that the subsequently described event or circumstance can or can not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0065] The specification of a numerical range includes all integers within or defining the range and all sub-ranges defined by the integers within the range. For example, 5-10 nucleotides is understood as 5, 6, 7, 8, 9, or 10 nucleotides, while 5%-10% is understood to include 5% and all possible values up to 10%.
[0066] At least 17 of the 20 nucleotides in a nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides in the provided sequence, providing an upper limit, even if an upper limit is not explicitly provided as would be clearly understood. Similarly, at most 3 nucleotides is understood to encompass 0, 1, 2, or 3 nucleotides, providing a lower limit, even if a lower limit is not explicitly provided. When "at least," "at most," or other like language modifies a number, that number is understood to be modified in this series.
[0067] As used herein, "no more than" or "less than" is understood to mean the value adjacent to the phrase, as well as logically lower values or integers down to zero, as would be reasonable from the context. For example, a duplex region of "no more than 2 nucleotide base pairs" has 2, 1, or 0 nucleotide base pairs. When "no more than" or "less than" precedes a series of numbers or a range, it should be understood that each number in that series or range is modified.
[0068] As used herein, it is understood that when a maximum amount of a value is expressed by 100% (e.g., 100% inhibition), the value is limited by the detection method. For example, 100% inhibition is understood to be inhibition to a level below the level of detection of the assay.
[0069] Unless otherwise clear from context, the term "about" encompasses values ±5% of the stated value. In certain embodiments, the term "about" is understood to encompass variations or errors tolerable within the art, for example, 2 standard deviations from the mean or sensitivity of the method used to make the measurement, or as a percentage of the value tolerable in the art, for example, with age. When "about" precedes the first value in a series, it is understood to modify each value in that series.
[0070] The term "and / or" means and encompasses any and all possible combinations of one or more of the associated listed items and the lack of combination in the alternative ("or").
[0071] The term "or" means either the member of the particular list or the absence of that list.
[0072] Unless otherwise clear from context, the singular forms "a," "an," and "the" include plural referents. For example, the term "protein" or "at least one protein" can include a plurality of proteins, including mixtures thereof.
[0073] Statistically significant means p < 0.05.
[0074] In the event of a conflict between a sequence in the present application and a specified accession number or position in an accession number, the sequence in the present application controls. DETAILED DESCRIPTION
[0075] I. Overview Provided herein are combinations comprising a CRISPR / Cas system, an inhibitor of CtBP-interacting protein (CtIP), and 53BP1 for use in enhancing homology directed repair of CRISPR / Cas-mediated cleavage of a target DNA by an exogenous donor nucleic acid. Also provided are methods of making targeted genetic modifications in a cell using such combinations by homology directed repair of CRISPR / Cas-mediated cleavage at a target genomic locus in the cell.
[0076] The compositions and methods disclosed herein improve the efficiency of precise CRISPR editing (precise gene knock-in) by stimulating DNA end resection in CRISPR-targeted mammalian cells, and thus improve the efficiency of homology directed repair (HDR). These compositions and methods are applicable to non-cycling cells (not amenable to HDR).
[0077] We observed that the combination of 53BP1 inhibition with CtIP localization resulted in additive or combinatorial enhancement of HDR efficiency. This result was unexpected, as their effects were expected to be redundant. Furthermore, in contrast to the expectation of hyper-resection leading to mutagenic single-strand annealing repair, an expected outcome of 53BP1 inhibition and CtIP-mediated resection, we observed an increase in error-free HDR. Our approach allows a stoichiometry of 1 :4 for sgRNA2.0:MS2-CtIP, while providing multiple copies of i53 as a global but transient block of 53BP1, ensuring their availability and activity are synchronized with CRISPR activity. In some embodiments, we employ LNP-mediated mRNA delivery, ensuring transient expression of our HDR enhancers, thus minimizing long-term effects of these factors on global genomic integrity.
[0078] II. Methods, compositions, and combinations for promoting homologous targeted repair Provided herein are compositions or combinations for facilitating homology directed repair. Such compositions or combinations can comprise: (a) a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) protein or a nucleic acid encoding a Cas protein; (b) a guide RNA or one or more DNAs encoding a guide RNA, wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) a CtBP-interacting protein (CtIP) or a nucleic acid encoding a CtIP protein; (d) an inhibitor of a 53BP1 (i53) protein or a nucleic acid encoding an i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus. Optionally, the 5' homology arm and the 3' homology arm flank an insert nucleic acid. For example, such compositions or combinations can comprise: (a) a Cas protein or a nucleic acid encoding a Cas protein; (b) a guide RNA or one or more DNAs encoding a guide RNA, wherein the guide RNA comprises one or more adaptor binding elements to which an adaptor protein is capable of specifically binding, and wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) a fusion protein or a nucleic acid encoding a fusion protein, wherein the fusion protein comprises a CtIP protein fused to an adaptor protein; (d) an i53 protein or a nucleic acid encoding an i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus. Optionally, the 5' homology arm and the 3' homology arm flank an insert nucleic acid. As used herein, the term "in combination with" means that some components can be administered prior to, concurrently with, or subsequent to the administration of other components. The different components of the combination can be formulated together, e.g., for simultaneous delivery, or separately formulated, e.g., as a kit comprising each component, e.g., with the additional agent in a separate formulation. Suitable CRISPR / Cas systems, including Cas proteins and guide RNAs, are described in more detail elsewhere herein. Likewise, suitable CtIP proteins, adaptor proteins, fusion proteins, i53 proteins, and exogenous donor nucleic acids are described in more detail elsewhere herein.
[0079] Also provided herein are methods of making a genetically modified cell. Such methods can comprise: (a) introducing into a cell a Cas protein or a nucleic acid encoding the Cas protein; (b) introducing into the cell a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) introducing into the cell a CtIP protein or a nucleic acid encoding the CtIP protein; (d) introducing into the cell an i53 protein or a nucleic acid encoding the i53 protein; and (e) introducing into the cell an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein the Cas protein and the guide RNA form a complex, the Cas protein cleaves the guide RNA target sequence to create a double-strand break, and the exogenous donor nucleic acid recombines with the target genomic locus via homology-directed repair to create the genetically modified cell. Optionally, the 5' homology arm and the 3' homology arm flank an insertion nucleic acid. For example, such methods can comprise: (a) introducing into a cell a Cas protein or a nucleic acid encoding the Cas protein; (b) introducing into the cell a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises one or more adaptor binding elements to which an adaptor protein is capable of specifically binding, and wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at a target genomic locus; (c) introducing into the cell a fusion protein or a nucleic acid encoding the fusion protein, wherein the fusion protein comprises a CtIP protein fused to an adaptor protein; (d) introducing into the cell an i53 protein or a nucleic acid encoding the i53 protein; and (e) introducing into the cell an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein the Cas protein and the guide RNA form a complex, the Cas protein cleaves the guide RNA target sequence to create a double-strand break, and the exogenous donor nucleic acid recombines with the target genomic locus via homology-directed repair to create the genetically modified cell. Optionally, the 5' homology arm and the 3' homology arm flank an insertion nucleic acid. Suitable CRISPR / Cas systems (including Cas proteins and guide RNAs) are described in more detail elsewhere herein. Also, suitable CtIP proteins, adaptor proteins, fusion proteins, i53 proteins, and exogenous donor nucleic acids are described in more detail elsewhere herein.
[0080] The Cas protein in the composition or combination or method can be any suitable Cas protein and can be in any form, such as in the form of a protein, in the form of an RNA encoding the Cas protein, or in the form of a DNA encoding the Cas protein (e.g., in a vector, such as a recombinant adeno-associated virus (AAV) vector described in more detail elsewhere herein). Likewise, the Cas protein or nucleic acid encoding the Cas protein can be in any form for delivery, such as in a lipid nanoparticle described in more detail elsewhere herein. The Cas protein can be any Cas protein described herein, such as a Cas9 protein. In one particular example, the Cas9 protein can be a S. pyogenes Cas9 protein, a C. jejuni Cas9 protein, or a S. aureus Cas9 protein (e.g., a S. pyogenes Cas9 protein). Cas proteins and CRISPR / Cas systems are described in more detail elsewhere herein.
[0081] The guide RNA in the composition or combination or method can be any suitable guide RNA and can be in any form, such as in the form of an RNA or in the form of one or more DNAs encoding the guide RNA (e.g., in a vector, such as a recombinant adeno-associated virus (AAV) vector described in more detail elsewhere herein). Likewise, the guide RNA or one or more DNAs encoding the guide RNA can be in any form for delivery, such as in a lipid nanoparticle described in more detail elsewhere herein. In one example, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind. For example, the guide RNA can comprise a first adaptor binding element within a first loop of the guide RNA, and a second adaptor binding element within a second loop of the guide RNA. For example, the guide RNA can be a single guide RNA comprising a CRISPR RNA (crRNA) portion partially fused to a trans-activating CRISPR RNA (tracrRNA) portion, wherein the first loop is a four-loop corresponding to residues 13 to 16 of SEQ ID NO: 11, 13, 15, or 16, and the second loop is a stem loop 2 corresponding to residues 53 to 56 of SEQ ID NO: 11, 13, 15, or 16. In one specific example, the adaptor binding elements comprise the sequences set forth in SEQ ID NO: 19 or 20. For example, the guide RNA can comprise the sequence set forth in any one of SEQ ID NOs: 21-26. Guide RNAs are described in more detail elsewhere herein.
[0082] The fusion protein or CtIP protein in the composition or combination or method can be in any form, such as in the form of a protein, in the form of an RNA encoding the protein, or in the form of a DNA encoding the protein (e.g., in a vector, such as a recombinant adeno-associated virus (AAV) vector described in more detail elsewhere herein). Likewise, the fusion protein or CtIP protein or nucleic acid encoding the protein can be in any form for delivery, such as in a lipid nanoparticle described in more detail elsewhere herein. In one specific example, the adaptor protein in the fusion protein comprises an MS2 coat protein or a functional fragment or variant thereof. For example, the adaptor protein can comprise a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 32, or can be encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33. In one example, the CtIP protein is a human CtIP protein. In one example, the CtIP protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 34, or is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 36. In one specific example, the fusion protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 30, or is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31. The CtIP protein, adaptor protein, and fusion protein are described in more detail elsewhere herein.
[0083] The i53 protein in the composition or combination or method can be in any form, such as in the form of a protein, in the form of an RNA encoding the protein, or in the form of a DNA encoding the protein (e.g., in a vector, such as a recombinant adeno-associated virus (AAV) vector described in more detail elsewhere herein). Likewise, the i53 protein or nucleic acid encoding the protein can be in any form for delivery, such as in a lipid nanoparticle described in more detail elsewhere herein. In one specific example, the i53 protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 40 or 42, or is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 41 or 43. The i53 protein is described in more detail elsewhere herein.
[0084] The exogenous donor nucleic acid in the composition or combination or method can be any suitable exogenous donor nucleic acid. In one example, the exogenous donor nucleic acid comprises an insert nucleic acid. The exogenous donor nucleic acid can be in a vector, such as a recombinant AAV vector. The exogenous donor nucleic acid can be any size nucleic acid. In some cases, it can be a large targeting vector (LTVEC). For example, it can be at least 10 kb in size or the sum of the 5' homology arm and the 3' homology arm can be at least 10 kb or can be about 10 kb to about 200 kb.
[0085] If a lipid nanoparticle is used, in one example, the Cas protein or nucleic acid encoding the Cas protein, the fusion protein or CtIP protein or nucleic acid encoding the fusion protein or CtIP protein, and the i53 protein or nucleic acid encoding the i53 protein can be in the same lipid nanoparticle. Alternatively, the Cas protein or nucleic acid encoding the Cas protein, the fusion protein or CtIP protein or nucleic acid encoding the fusion protein or CtIP protein, the i53 protein or nucleic acid encoding the i53 protein, and the guide RNA or DNA encoding the guide RNA can be in the same lipid nanoparticle. If a vector is used, in one example, the DNA encoding the guide RNA and the exogenous donor nucleic acid can be in the vector. Alternatively, only the exogenous donor nucleic acid can be in the vector.
[0086] In one specific example, the composition or combination comprises RNA encoding the fusion protein or CtIP protein and RNA encoding the i53 protein. In another specific example, the composition or combination comprises RNA encoding the fusion protein or CtIP protein and RNA encoding the i53 protein, wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle. In another specific example, the composition or combination comprises RNA encoding the fusion protein or CtIP protein and RNA encoding the i53 protein, wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle, and the exogenous donor nucleic acid is in a vector (e.g., a recombinant AAV vector).
[0087] In one specific example, the composition or combination comprises RNA encoding the Cas protein, RNA encoding the fusion protein or CtIP protein, RNA encoding the i53 protein, and one or more DNA encoding the guide RNA, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, and the RNA encoding the i53 protein are in a lipid nanoparticle, and the one or more DNA encoding the guide RNA and the exogenous donor nucleic acid are in a vector (e.g., a recombinant AAV vector).
[0088] In another specific example, the composition or combination comprises an RNA encoding a Cas protein, an RNA encoding a fusion protein or a CtIP protein, an RNA encoding an i53 protein, and a guide RNA in the form of an RNA, wherein the RNA encoding a Cas protein, the RNA encoding a fusion protein, the RNA encoding an i53 protein, and the guide RNA are in a lipid nanoparticle, and the exogenous donor nucleic acid is in a vector (e.g., a recombinant AAV vector).
[0089] In one specific example, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises an RNA encoding a fusion protein or a CtIP protein, and an RNA encoding an i53 protein. In another specific example, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises an RNA encoding a fusion protein or a CtIP protein, and an RNA encoding an i53 protein, and the RNA encoding a fusion protein and the RNA encoding an i53 protein are in a lipid nanoparticle. In another specific example, the guide RNA comprises two adaptor binding elements to which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises an RNA encoding a fusion protein or a CtIP protein, and an RNA encoding an i53 protein, and the RNA encoding a fusion protein and the RNA encoding an i53 protein are in a lipid nanoparticle, and the exogenous donor nucleic acid is in a vector (e.g., a recombinant AAV vector).
[0090] In one specific example, the guide RNA comprises two adaptor binding elements with which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises RNA encoding a Cas protein, the composition or combination comprises RNA encoding a fusion protein or a CtIP protein, the composition or combination comprises RNA encoding an i53 protein, the composition or combination comprises the guide RNA in the form of RNA, the RNA encoding a Cas protein, the RNA encoding a fusion protein, the RNA encoding an i53 protein, and the guide RNA is in a lipid nanoparticle, and the exogenous donor nucleic acid is in a vector (e.g., a recombinant AAV vector).
[0091] In another specific example, the guide RNA comprises two adaptor binding elements with which an adaptor protein can specifically bind, wherein the first adaptor binding element is within a first loop of the guide RNA and the second adaptor binding element is within a second loop of the guide RNA, the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, the composition or combination comprises RNA encoding a Cas protein, the composition or combination comprises RNA encoding a fusion protein or a CtIP protein, the composition or combination comprises RNA encoding an i53 protein, the composition or combination comprises the guide RNA in the form of RNA, the RNA encoding a Cas protein, the RNA encoding a fusion protein, the RNA encoding an i53 protein, and the guide RNA is in a lipid nanoparticle, and the exogenous donor nucleic acid is in a vector (e.g., a recombinant AAV vector).
[0092] The cells in these methods can be any suitable cell. Examples of different cells are disclosed in more detail elsewhere herein. For example, the cells can be mammalian cells, rodent cells, mouse cells, rat cells, or human cells. In some methods, the cells are non-cycling cells (i.e., non-dividing). In some methods, the cells are cycling (i.e., dividing) cells.
[0093] The nucleic acids or proteins used in the methods disclosed herein can be introduced into the cells by any suitable means. Various methods and compositions that allow for the introduction of molecules (e.g., nucleic acids or proteins) into cells or subjects are provided herein. Methods for introducing molecules into various cell types are known and include, for example, stable transfection methods, transient transfection methods, and viral-mediated methods.
[0094] The transfection protocol, as well as the protocol for introducing molecules into cells, can vary. Non-limiting transfection methods include chemical-based transfection methods using liposomes; nanoparticles; calcium phosphate (Graham et al., (1973)Virology 52 (2): 456-67, Bacchetti et al. (1977) Proc. Natl. Acad. Sci. U.S.A. 74(4): 1590-4, and Kriegler, M (1991). Transfer and Expression: A Laboratory Manual. New York: W. H. Freeman and Company. pp. 96-97); dendrimers; or cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation, sonoporation, and optical transfection. Particle-based transfection includes the use of a gene gun or magnet assisted transfection (Bertram (2006) Current Pharmaceutical Biotechnology 7, 277-28). Viral methods can also be used for transfection.
[0095] Introduction of nucleic acids or proteins into cells can also be mediated by electroporation, cytoplasmic injection, viral infection, adenovirus, adeno-associated virus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or nucleofection. Nucleofection is an improved electroporation technique that enables nucleic acid substrates to be delivered not only to the cytoplasm but also into the nucleus through the nuclear membrane. Additionally, the use of nucleofection in the methods disclosed herein typically requires much fewer cells than conventional electroporation (e.g., only about 2 million compared to 7 million for conventional electroporation). In one example, nucleofection is performed using the LONZA ® NUCLEOFECTOR ™ system.
[0096] Introduction of molecules (e.g., nucleic acids or proteins) into a cell (e.g., a zygote) can also be accomplished by microinjection. In a zygote (i.e., a single cell stage embryo), microinjection can be into the maternal and / or paternal pronucleus or cytoplasm. If microinjection is into only one pronucleus, the paternal pronucleus is preferred because it is larger in size. Microinjection of mRNA into the cytoplasm (e.g., delivering mRNA directly to the translation machinery) is preferred, while microinjection of Cas proteins or polynucleotides encoding Cas proteins or RNA into the nucleus / pronucleus is preferred. Alternatively, microinjection can be performed by injection into both the nucleus / pronucleus and cytoplasm: the needle can be first introduced into the nucleus / pronucleus and a first amount can be injected, and upon removal of the needle from the single cell stage embryo, a second amount can be injected into the cytoplasm. If Cas proteins are injected into the cytoplasm, the Cas proteins preferably include a nuclear localization signal to ensure delivery to the nucleus / pronucleus. Methods for performing microinjection are well known. See, e.g., Nagy et al. (Nagy A, Gertsenstein M, Vintersten K, Behringer R., 2003, Manipulating the Mouse Embryo. Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press); see also Meyer et al., (2010) Proc. Natl. Acad. Sci. U.S.A. 107: 15022-15026 and Meyer et al., (2012) Proc. Natl. Acad. Sci. U.S.A. 109: 9354-9359, each of which is incorporated by reference herein in its entirety for all purposes.
[0097] Other methods for introducing molecules (e.g., nucleic acids or proteins) into a cell or subject can include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. As specific examples, nucleic acids or proteins can be introduced into a cell or subject in a carrier such as a poly(lactic acid) (PLA) microsphere, a poly(D,L-lactic-co-glycolic acid) (PLGA) microsphere, a liposome, a micelle, an inverse micelle, a lipid helix, or a lipid microtube. Some specific examples of delivery to a subject include hydrodynamic delivery, virus-mediated delivery (e.g., adeno-associated virus (AAV)-mediated delivery), and lipid nanoparticle-mediated delivery.
[0098] Introduction of nucleic acids and proteins into cells or subjects can be accomplished by hydrodynamic delivery (HDD). For gene delivery to parenchymal cells, only the necessary DNA sequence needs to be injected via a selected blood vessel, eliminating safety concerns associated with current viral and synthetic vectors. When injected into the bloodstream, DNA is able to reach cells in different tissues accessible by blood. Hydrodynamic delivery addresses the problem of the physical barrier of endothelial and cell membranes that prevents large and impermeable membrane compounds from entering parenchymal cells using the force generated by the rapid injection of a large volume of solution into the non-compressible blood in circulation. In addition to delivering DNA, this method can also be used for efficient intracellular delivery of RNA, proteins, and other small compounds in vivo. See, e.g., Bonamassa et al. (2011) Pharm. Res. 28(4): 694-701, which is incorporated by reference herein in its entirety for all purposes.
[0099] The introduction of nucleic acids can also be accomplished via virus-mediated delivery, such as AAV-mediated delivery or lentivirus-mediated delivery. Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. Viruses can infect dividing cells, non-dividing cells, or both. Viruses can integrate into the host genome or, alternatively, not integrate into the host genome. Such viruses can also be engineered to have reduced immunity. Viruses may be capable of replication or may have replication defects (e.g., defects in one or more genes necessary for additional rounds of viral particle replication and / or packaging). Viruses can induce transient or more persistent expression. Viral vectors can be genetically modified from their wild-type counterparts. For example, viral vectors may contain insertions, deletions, or substitutions of one or more nucleotides to facilitate cloning or to alter one or more properties of the vector. Such properties may include packaging capacity, transduction efficiency, immunogenicity, genome integration, replication, transcription, and translation. In some examples, a portion of the viral genome may be deleted, enabling the virus to package exogenous sequences of a larger size. In some examples, the viral vector may have enhanced transduction efficiency. In some examples, it may reduce the virus-induced immune response in the host. In some examples, it may mutate viral genes (such as integrases) that promote the integration of viral sequences into the host genome, making the virus non-integrating. In some examples, the viral vector may be replication-defective. In some examples, the viral vector may contain exogenous transcriptional or translational control sequences to drive the expression of the coding sequence on the vector. In some examples, the virus may be helper-dependent. For example, the virus may require one or more helper components to provide the viral components (such as viral proteins) needed to amplify the vector and package the vector into viral particles. In this case, one or more helper components (including one or more vectors encoding viral components) may be introduced into a host cell or host cell population along with the vector system described herein. In other examples, the virus may be without helper components. For example, the virus may be able to amplify and package the vector without helper viruses. In some examples, the vector system described herein may also encode viral components needed for viral amplification and packaging.
[0100] Exemplary viral titers (e.g., AAV titers) include approximately 10 12 vg / mL to approximately 10 16 vg / mL. Other exemplary viral titers (e.g., AAV titers) include approximately 10 vg / mL. 12 vg / kg body weight approximately 10 16 vg / kg body weight.
[0101] Introduction of nucleic acids and proteins can also be accomplished by lipid nanoparticle (LNP)-mediated delivery. For example, LNP-mediated delivery can be used to deliver a combination of Cas mRNA and guide RNA or a combination of Cas protein and guide RNA. LNP-mediated delivery can be used to deliver guide RNA in the form of RNA. In a specific example, the guide RNA and Cas protein are each introduced in the form of RNA via LNP-mediated delivery into the same LNP. As discussed in greater detail elsewhere herein, one or more RNAs can be modified. Delivery by such methods results in transient Cas expression and / or transient presence of guide RNA, and biodegradable lipids improve clearance, improve tolerability, and reduce immunogenicity. Lipid formulations can protect biomolecules from degradation while improving their cellular uptake. Lipid nanoparticles are particles that include a plurality of lipid molecules that are physically associated with one another through intermolecular forces. These particles comprise microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), dispersed phases in emulsions, micelles, or internal phases in suspensions. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids can be used to deliver polyanions, such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time that a nanoparticle can exist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO 2016 / 010840 Al and WO 2017 / 173054 Al, which are incorporated by reference herein in their entireties for all purposes.
[0102] In certain LNPs, the cargo can comprise a guide RNA or a nucleic acid encoding a guide RNA. In certain LNPs, the cargo can comprise an mRNA encoding a Cas nuclease such as Cas9, and a guide RNA or a nucleic acid encoding a guide RNA. In certain LNPs, the cargo can comprise a nucleic acid construct. In certain LNPs, the cargo can comprise an mRNA encoding a Cas nuclease such as Cas9, a guide RNA or a nucleic acid encoding a guide RNA, and a nucleic acid construct. LNPs for use in these methods are described in greater detail elsewhere herein.
[0103] The mode of delivery can be selected to reduce immunogenicity. For example, Cas proteins and gRNAs can be delivered by different modes (e.g., dual mode delivery). These different modes can impart different pharmacodynamic or pharmacokinetic properties to the delivered molecules (e.g., Cas or nucleic acid encoding, gRNA or nucleic acid encoding, or nucleic acid construct encoding a polypeptide of interest) to the subject. For example, different modes result in different tissue distribution, different half-lives, or different temporal profiles. Some modes of delivery (e.g., delivery of nucleic acid vectors that persist in the cell by autonomous replication or genomic integration) result in more persistent expression and presence of the molecule, while other modes of delivery are transient and less persistent (e.g., delivery of RNA or protein). Delivery of Cas proteins in a more transient manner, e.g., as mRNA or protein, can ensure that the Cas / gRNA complex is present and active for only a short period of time, and can reduce immunogenicity caused by peptides from the bacterial source of the Cas enzyme displayed on the cell surface by MHC molecules. Such transient delivery can also reduce the likelihood of off-target modifications.
[0104] In vivo administration can be by any suitable route, including, for example, systemic routes of administration, such as parenteral administration, e.g., intravenous, subcutaneous, intra-arterial, or intramuscular. In one particular example, in vivo administration is intravenous.
[0105] A composition comprising a guide RNA and / or Cas protein (or nucleic acid encoding a guide RNA and / or Cas protein) can be formulated using one or more physiologically and pharmaceutically acceptable carriers, diluents, excipients or auxiliaries. The formulation can depend on the route of administration chosen. Pharmaceutically acceptable means that the carrier, diluent, excipient or auxiliary is compatible with the other ingredients of the formulation and not substantially deleterious to the recipient thereof. In one particular example, the route of administration and / or formulation is selected for delivery to the liver (e.g., hepatocytes).
[0106] The methods can further comprise identifying cells having a modified target genomic locus. Various methods can be used to identify cells and animals having targeted genetic modifications, such as PCR. The screening step can include, for example, a quantitative assay for assessing modification of an allele (MOA) of a parental chromosome. For example, the quantitative assay can be performed via quantitative PCR such as real-time PCR (qPCR). Real-time PCR can utilize a first primer set recognizing the target locus and a second primer set recognizing a non-targeted reference locus. The primer sets can include a fluorescent probe recognizing the amplified sequence. Other examples of suitable quantitative assays include fluorescence-mediated in situ hybridization (FISH), comparative genomic hybridization, isothermal DNA amplification, quantitative hybridization to immobilized probes, INVADER ® probe, TAQMAN ® molecular beacon probe, or ECLIPSE™ Probe technology (see, e.g., US 2005 / 0144655, which is incorporated by reference herein in its entirety for all purposes).
[0107] In some methods, the percentage of cells having a targeted genetic modification resulting from homology-directed repair is higher than in a control method in which a CtIP protein (or CtIP fusion protein) (in any form, such as a protein, RNA encoding, or DNA encoding) is administered but an i53 protein (in any form) is not administered. In some methods, the percentage of cells having a targeted genetic modification resulting from homology-directed repair is higher than in a control method in which an i53 (in any form) is administered but a CtIP protein (or CtIP fusion protein) (in any form) is not administered. In some methods, the percentage of cells having a targeted genetic modification resulting from homology-directed repair is higher than in a control method in which neither a CtIP protein (or CtIP fusion protein) nor an i53 protein (in any form) is administered. In some methods, the percentage of cells having a targeted genetic modification resulting from homology-directed repair is at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, or at least about 20-fold greater relative to a control method in which neither a CtIP protein (or CtIP fusion protein) nor an i53 protein (in any form) is administered. In some methods, the percentage of cells having a targeted genetic modification resulting from homology-directed repair is about 15-fold to about 25-fold, about 16-fold to about 24-fold, about 17-fold to about 23-fold, about 18-fold to about 22-fold, about 19-fold to about 21-fold, about 15-fold to about 20-fold, about 16-fold to about 20-fold, about 17-fold to about 20-fold, about 18-fold to about 20-fold, about 19-fold to about 20-fold, about 20-fold to about 25-fold, about 20-fold to about 24-fold, about 20-fold to about 23-fold, about 20-fold to about 22-fold, about 20-fold to about 21-fold, or about 20-fold greater relative to a control method in which neither a CtIP protein (or CtIP fusion protein) nor an i53 protein (in any form) is administered.
[0108] A. CRISPR / Cas system The methods and compositions disclosed herein can utilize a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) / CRISPR-associated (Cas) system or components of such a system to modify a genome within a cell. A CRISPR / Cas system comprises transcripts and other elements involved in expression of or directing activity of a Cas gene. A CRISPR / Cas system can be, for example, a Type I, Type II, Type III system, or a Type V system (e.g., Type V-A sub-type or Type V-B sub-type). The methods and compositions disclosed herein can employ a CRISPR / Cas system by utilizing a CRISPR complex (comprising a guide RNA (gRNA) complexed with a Cas protein) for site-directed binding or cleavage of a nucleic acid. A CRISPR / Cas system targeting a target locus comprises a Cas protein (or a nucleic acid encoding a Cas protein) and one or more guide RNAs (or DNA encoding one or more guide RNAs), wherein each of the one or more guide RNAs targets a different guide RNA target sequence in the target locus. Such a CRISPR / Cas system targeting a target locus can also comprise one or more exogenous donor sequences (e.g., targeting vectors) targeting the target locus.
[0109] A CRISPR / Cas system used in the compositions and methods disclosed herein can be non-naturally occurring. A "non-naturally occurring" system comprises anything that indicates involvement of human artifice, such as a component or components of the system being altered or mutated from its naturally occurring state, being at least substantially free of at least one other component with which it is naturally associated in nature, or being associated with at least one other component with which it is not naturally associated. For example, some CRISPR / Cas systems employ a non-naturally occurring CRISPR complex comprising a gRNA and a Cas protein that do not occur together in nature, employ a non-naturally occurring Cas protein, or employ a non-naturally occurring gRNA.
[0110] (1) Cas protein Cas proteins generally comprise at least one RNA recognition or binding domain that can interact with a guide RNA. Cas proteins can also include nuclease domains (e.g., DNase domains or RNase domains), DNA binding domains, helicase domains, protein-protein interaction domains, dimerization domains, and other domains. Some such domains (e.g., DNase domains) can be from a native Cas protein. Other such domains can be added to make a modified Cas protein. Nuclease domains have catalytic activity for nucleic acid cleavage, which comprises the breaking of a covalent bond of a nucleic acid molecule. Cleavage can produce a blunt end or a staggered end, and cleavage can be single-stranded or double-stranded. For example, wild-type Cas9 proteins generally make a blunt end cleavage product. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) can give a cleavage product with a 5-nucleotide 5' overhang, with cleavage occurring 18 base pairs after the PAM sequence on the non-targeted strand and after the 23rd base on the targeted strand. Cas proteins can have full cleavage activity to make a double-stranded break (e.g., a double-stranded break with a blunt end) at a target genomic locus, or they can be nickases that make a single-stranded break at a target genomic locus.
[0111] Examples of Cas proteins include Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8al, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), CaslO, CaslOd, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, as well as homologs or modified versions thereof.
[0112] An exemplary Cas protein is a Cas9 protein or a protein derived from a Cas9 protein. Cas9 proteins are from type II CRISPR / Cas systems and generally share four key motifs with conserved architecture. Motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins are from Streptococcus pyogenes (S. pyogenes) (SpCas9), Streptococcus thermophilus (StCas9), and Lachnospira pectinosolvens (LpCas9). Streptococcus pyogenes Streptococcus thermophilus Streptococcus ( Streptococcus sp. Staphylococcus aureus ( Staphylococcus aureus ), Nocardia dassonvillei ( Nocardiopsis dassonvillei ), Streptomyces coccidioides ( Streptomyces pristinaespiralis ), green-producing Streptomyces ( Streptomyces viridochromogenes ), green-producing Streptomyces, and Neurocystis ( Streptosporangium roseum ), Streptococcus, Cyclocycline ( Alicyclobacillus acidocaldarius ), Bacillus pseudomycosis ( Bacillus pseudomycoides ), selenized Bacillus ( Bacillus selenitireducens ), Siberian microbacteria ( Exiguobacterium sibiricum Lactobacillus delbrueckii (), Lactobacillus delbrueckii ), Lactobacillus salivarius ( Lactobacillus salivarius ), marine micro-vibrating cyanobacteria ( Microscilla marina Burkholderia ( ) Burkholderiales bacterium ), Naphthalene-eating polar monoclonal bacteria ( Polaromonas naphthalenivorans ), Polar Monoclonal bacteria ( Polaromonas sp. ), *Cyclocarya vulgaris* ( Crocosphaera watsonii ), Blue filamentous fungus ( Cyanothece sp. Microcystis aeruginosa ( Microcystis aeruginosa Synechococcus Synechococcus sp. ), Arabica acetate ( Acetohalobium arabaticum ), Degens ammonia-producing bacteria ( Ammonifex degensii ), pyrolytic cellulose bacteria ( Caldicelulosiruptor becscii ), candidate gold-mining bacteria ( Candidatus Desulforudis ), botulinum toxin ( Clostridium botulinum Clostridium difficile ( Clostridium difficile ), Griffon's bacterium ( Finegoldia magna ), thermophilic anaerobic bacilli ( Natranaerobius thermophilus ), propionic acid degrading bacteria ( Pelotomaculum thermopropionicum ), thermophilic acidophilic thiobacillus ( Acidithiobacillus caldus ), Acidophilic ferrous thiobacillus ( Acidithiobacillus ferrooxidans ), wine-colored heterochromatic bacteria ( Allochromatium vinosum ), seabacteria ( Marinobacter sp. ), halophilic nitrite cocci ( Nitrosococcus halophilus ), Nitrostrophus warwickii ( Nitrosococcus watsoni ), Pseudomonas alterniflora ( Pseudoalteromonas haloplanktis ), racemic fibrobacterium ( Ktedonobacter racemifer ), methanogenic bacteria ( Methanohalobium egestium ), Anabaena ( Anabaena variabilis ), Foamy Cladosporium ( Nodularia spumigena ), Nostoc ( Nostoc sp. ), Spirulina macrophylla ( Arthrospira maxima ), the largest arthrospira ( Arthrospira platensis ), Arthrospira (Arthrospira sp. ), Lin's algae ( Lyngbya sp. ), Prototype Microsheatha ( Microcoleus chthonoplastes Oscillatoria ( Oscillatoria sp. ), mobile Lithocarpus ( Petrotoga mobilis ), African thermocline bacteria ( Thermosipho africanus ), Deep-sea Akaro Tiger Tail Grass ( Acaryochloris marina ), Neisseria meningitidis ( Neisseria meningitidis ) or Campylobacter jejuni ( Campylobacter jejuni Further examples of Cas9 family members are described in WO 2014 / 131833, which is incorporated herein by reference in its entirety for all purposes. Cas9 (SpCas9) from *Streptococcus pyogenes* (e.g., assigned UniProt accession number Q99ZW2) is an exemplary Cas9 protein. Smaller Cas9 proteins (e.g., Cas9 proteins whose coding sequences are compatible with maximum AAV packaging capacity when combined with guide RNA coding sequences and regulatory elements of Cas9 and guide RNA, such as SaCas9, CjCas9, and Nme2Cas9) are other exemplary Cas9 proteins. For example, Cas9 (SaCas9) from *Staphylococcus aureus* (assigned UniProt accession number J7RUA5) is another exemplary Cas9 protein. Similarly, Cas9 (CjCas9) from *Campylobacter jejuni* (e.g., assigned UniProt accession number Q0P897) is another exemplary Cas9 protein. See, for example, Kim et al., (2017). Nat. Commun. 8:14500, this document is incorporated herein by reference in its entirety for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, for example, Edraki et al., (2019). Mol. Cell 73(4):714-726, which is incorporated herein by reference in its entirety for all purposes. Other exemplary Cas9 proteins are those from *Streptococcus thermophilus* (e.g., *Streptococcus thermophilus* LMD-9 Cas9 encoded by the CRISPR1 locus (St1Cas9) or *Streptococcus thermophilus* Cas9 encoded by the CRISPR3 locus (St3Cas9)). *Streptococcus thermophilus* from the novel killer *Francisella* (… Francisella novicida The Cas9 (FnCas9) of *F. francasus*, a novel RHA culprit that recognizes alternative PAMs (E1369R / E1449H / R1556A substitutions), is another exemplary Cas9 protein. These and other exemplary Cas9 proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017). Mamm. GenomeIn 28(7):247-261, the full text of which is incorporated herein by reference for all purposes. Examples of Cas9 coding sequences, Cas9 mRNA, and Cas9 protein sequences are provided in WO 2013 / 176772, WO 2014 / 065596, WO 2016 / 106121, and WO 2019 / 067910, each of which is incorporated herein by reference for all purposes. Specific examples of ORF and Cas9 amino acid sequences are provided in Table 30 of paragraph
[0449] of WO 2019 / 067910, and specific examples of Cas9 mRNA and ORF are provided in paragraphs
[0214] -
[0234] of WO 2019 / 067910. As an example, the Cas9 protein comprises, is substantially composed of, or is composed of the sequence shown in SEQ ID NO: 1. Such Cas9 protein sequences can be encoded by DNA including SEQ ID NO: 2, substantially composed of it, or composed of it.
[0113] Another example of a Cas protein is Cpf1 (from Prevotella). Prevotella ) and Francisella ( Francisella Cpf1 is a CRISPR protein. It is a large protein (approximately 1300 amino acids) containing a RuvC-like nuclease domain homologous to the corresponding domain of Cas9, as well as a counterpart to the characteristic arginine-rich Cas9 cluster. However, Cpf1 lacks the HNH nuclease domain present in Cas9 proteins, and the RuvC-like domain is continuous in the Cpf1 sequence, unlike Cas9, which contains a long insert containing the HNH domain. See, for example, Zetsche et al., (2015). Cell 163(3):759-771, this document is incorporated herein by reference in its entirety for all purposes. The exemplary Cpf1 protein is derived from *Tulafrancsis* (…). Francisella tularensis 1. A new culprit subspecies of *Tulafrancsis* ( Francisella tularensis subsp. novicida Prevostii Elbe ( Prevotella albensis ), bacteria of the family Trichophyceae ( Lachnospiraceae bacterium MC2017 1. Vibrio butyricum ( Butyrivibrio proteoclasticus ), Heterozoanthellae ( Peregrinibacteria bacterium GW2011_GWA2_33_10, Bacteria of the Superphylum of Thrift ( Parcubacteria bacterium GW2011_GWC2_44_17, genus *Smithia* ( Smithella sp. SCADC, amino acid cocci ( Acidaminococcus sp. BV3L6, bacteria of the Trichophyceae family ( Lachnospiraceae bacterium MA2020, candidate termite methane mycoplasma ( CandidatusMethanoplasma termitum ), picky eubacterium ( Eubacterium eligens ), Moraxella bubalana ( Moraxella bovoculi 237. Leptospira Inada ( Leptospira inadai ), bacteria of the family Trichophyceae ( Lachnospiraceae bacterium ND2006, Porphyromonas canis oralis ( Porphyromonas crevioricanis ) 3. Prevotella glycopeptone ( Prevotella disiens ) and Porphyromonas maculatus ( Porphyromonas macacae Cpf1 (FnCpf1; assigned UniProt accession number A0Q7Q2) from the novel culprit Francisella U112 is an exemplary Cpf1 protein.
[0114] Another example of a Cas protein is CasX (Cas12e). CasX is an RNA-guided DNA endonuclease that creates staggered double-strand breaks in DNA. CasX is less than 1000 amino acids in size. Exemplary CasX proteins are derived from *Deltaproteus* (*C. spp.*). Deltaproteobacteria (DpbCasX or DpbCas12e) and planktonic fungi ( Planctomycetes (PlmCasX or PlmCas12e). Similar to Cpf1, CasX uses a single RuvC active site for DNA cleavage. See, for example, Liu et al., (2019). Nature 566(7743):218-223, the full text of which is incorporated herein by reference for all purposes.
[0115] Another example of a Cas protein is CasΦ (CasPhi or Cas12j), which is uniquely found in bacteriophages. CasΦ is less than 1000 amino acids in size (e.g., 700–800 amino acids). CasΦ cleavage produces staggered 5' overhangs. A single RuvC active site in CasΦ enables crRNA processing and DNA cleavage. See, for example, Pausch et al., (2020). Science 369(6501):333-337, which is incorporated herein by reference in its entirety for all purposes.
[0116] A Cas protein can be a wild-type protein (i.e., those that occur in nature), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. A Cas protein can also be an active variant or fragment with respect to the catalytic activity of a wild-type or modified Cas protein. An active variant or fragment with respect to catalytic activity can include at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to a wild-type or modified Cas protein or portion thereof, where the active variant retains the ability to cleave at a desired cleavage site, and thus retains nick-inducing or double-strand break-inducing activity. Assays for nick-inducing or double-strand break-inducing activity are known, and generally measure the overall activity and specificity of a Cas protein on a DNA substrate containing a cleavage site.
[0117] One example of a modified Cas protein is a modified SpCas9-HFl protein, which is a high-fidelity variant of S. pyogenes Cas9 with alterations designed to reduce non-specific DNA contacts (N497A / R661A / Q695A / Q926A). See, e.g., Kleinstiver et al., (2016) Nature 529(7587): 490-495, which is incorporated by reference herein in its entirety for all purposes. Another example of a modified Cas protein is a modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, e.g., Slaymaker et al., (2016) Science 351(6268): 84-88, which is incorporated by reference herein in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A. These and other modified Cas proteins are reviewed in, e.g., Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7): 247-261, which is incorporated by reference herein in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, which is a SpCas9 variant that can recognize an expanded range of PAM sequences. See, e.g., Hu et al., (2018) Nature 556: 57-63, which is incorporated by reference herein in its entirety for all purposes.
[0118] A Cas protein can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. A Cas protein can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains unnecessary for protein function or to optimize (e.g., enhance or decrease) activity or properties of the Cas protein.
[0119] A Cas protein can include at least one nuclease domain, such as a DNase domain. For example, a wild-type Cpfl protein typically comprises a RuvC-like domain that cleaves both strands of a target DNA, which can be in a dimeric configuration. Likewise, CasX and CasΦ typically comprise a single RuvC-like domain that cleaves both strands of a target DNA. A Cas protein can also include at least two nuclease domains, such as DNase domains. For example, a wild-type Cas9 protein typically includes a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave a different strand of double-stranded DNA to form a double-stranded break in the DNA. See, e.g., Jinek et al., (2012) Science 337(6096): 816-821, which is incorporated by reference herein in its entirety for all purposes.
[0120] One or more or all of the nuclease domains in a nuclease domain can be deleted or mutated such that it is no longer active or has reduced nuclease activity. For example, if one of the nuclease domains in a Cas9 protein is deleted or mutated, the resulting Cas9 protein can be referred to as a nickase, and can generate a single-strand break but not a double-strand break within a double-stranded target DNA (i.e., it can cleave either the complementary strand or the non-complementary strand, but not both). If two of the nuclease domains are deleted or mutated, the resulting Cas protein (e.g., Cas9) will have reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein, or a catalytically dead Cas protein (dCas)). If none of the nuclease domains in a Cas9 protein are deleted or mutated, the Cas9 protein will retain double-strand break induction activity. An example of a mutation that converts Cas9 to a nickase is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Likewise, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert Cas9 to a nickase. Other examples of mutations that convert Cas9 to a nickase include corresponding mutations to Cas9 from S. thermophilus. See, e.g., Sapranauskas et al., (2011) Nucleic Acids Res. 39(21):9275-9282 and WO 2013 / 141680, each of which is incorporated by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Examples of other mutations that form a nickase can be found, e.g., in WO 2013 / 176772 and WO 2013 / 142578, each of which is incorporated by reference in its entirety for all purposes. If all of the nuclease domains in a Cas protein are deleted or mutated (e.g., both of the nuclease domains in a Cas9 protein are deleted or mutated), the resulting Cas protein (e.g., Cas9) will have reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein). One specific example is a D10A / H840A S. pyogenes Cas9 double mutant or corresponding double mutant in a Cas9 from another species when optimally aligned with S. pyogenes Cas9. Another specific example is a D10A / N863A S. pyogenes Cas9 double mutant or corresponding double mutant in a Cas9 from another species when optimally aligned with S. pyogenes Cas9.
[0121] Examples of inactivating mutations in the catalytic domain of xCas9 are the same as those described above for SpCas9. Examples of inactivating mutations in the catalytic domain of Staphylococcus aureus Cas9 protein are also known. For example, a Staphylococcus aureus Cas9 enzyme (SaCas9) can comprise a substitution at position N580 (e.g., an N580A substitution) and a substitution at position D10 (e.g., a D10A substitution) for generating a nuclease-inactivated Cas protein. See, e.g., WO 2016 / 106236, which is incorporated by reference herein in its entirety for all purposes. Examples of inactivating mutations in the catalytic domain of Nme2Cas9 are also known (e.g., a combination of D16A and H588A). Examples of inactivating mutations in the catalytic domain of St1Cas9 are also known (e.g., a combination of D9A, D598A, H599A, and N622A). Examples of inactivating mutations in the catalytic domain of St3Cas9 are also known (e.g., a combination of D10A and N870A). Examples of inactivating mutations in the catalytic domain of CjCas9 are also known (e.g., a combination of D8A and H559A). Examples of inactivating mutations in the catalytic domain of FnCas9 and RHA FnCas9 are also known (e.g., N995A).
[0122] Examples of inactivating mutations in the catalytic domain of Cpf1 proteins are also known. With reference to Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at position 908, 993, or 1263 of AsCpf1 or a corresponding position in a Cpf1 ortholog, or at position 832, 925, 947, or 1180 of LbCpf1 or a corresponding position in a Cpf1 ortholog. Such mutations can include, for example, one or more of mutations D908A, E993A, and D1263A of AsCpf1 or a corresponding mutation in a Cpf1 ortholog, or D832A, E925A, D947A, and D1180A of LbCpf1 or a corresponding mutation in a Cpf1 ortholog. See, e.g., US 2016 / 0208243, which is incorporated by reference herein in its entirety for all purposes.
[0123] Examples of inactivating mutations in the catalytic domain of CasX proteins are also known. With reference to CasX proteins from Deltaproteobacteria, D672A, E769A, and D935A (alone or in combination) or corresponding positions in other CasX orthologs are inactivating. See, e.g., Liu et al., (2019) Nature566 (7743):218-223, which is incorporated by reference herein in its entirety for all purposes.
[0124] Examples of inactivating mutations in the catalytic domain of CasΦ proteins are also known. For example, D371A and D394A, alone or in combination, are inactivating mutations. See, e.g., Pausch et al. (2020) Science 369 (6501):333-337, which is incorporated by reference herein in its entirety for all purposes.
[0125] Cas proteins can also be operably linked to a heterologous polypeptide as a fusion protein. For example, Cas proteins can be fused to cleavage domains. See WO 2014 / 089290, which is incorporated by reference herein in its entirety for all purposes. Cas proteins can also be fused to a heterologous polypeptide to provide increased or decreased stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or within the Cas protein.
[0126] As one example, Cas proteins can be fused to one or more heterologous polypeptides that provide subcellular localization. Such heterologous polypeptides can include, for example, one or more nuclear localization signals (NLS), such as a monopartite SV40 NLS and / or bipartite alpha-importin NLS for targeting to the nucleus, a mitochondrial localization signal for targeting to mitochondria, an ER retention signal, and the like. See, e.g., Lange et al. (2007) J. Biol. Chem. 282 (8):5101-5105, which is incorporated by reference herein in its entirety for all purposes. Such subcellular localization signals can be located at the N-terminus, C-terminus, or anywhere within the Cas protein. NLSs can include stretches of basic amino acids and can be monopartite or bipartite sequences. Optionally, Cas proteins can include two or more NLSs, including an NLS at the N-terminus (e.g., an alpha-importin NLS or monopartite NLS) and an NLS at the C-terminus (e.g., an SV40 NLS or bipartite NLS). Cas proteins can also include two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.
[0127] For example, a Cas protein can be fused to 1-10 NLSs (e.g., to 1-5 NLSs or to one NLS). Where one NLS is used, the NLS can be linked at the N- or C-terminus of the Cas protein sequence. It can also be inserted into the Cas protein sequence. Alternatively, a Cas protein can be fused to more than one NLS. For example, a Cas protein can be fused to 2, 3, 4, or 5 NLSs. In a specific example, a Cas protein can be fused to two NLSs. In some cases, the two NLSs can be the same (e.g., two SV40 NLSs) or different. For example, a Cas protein can be fused to two SV40 NLS sequences linked at the carboxy terminus. Alternatively, a Cas protein can be fused to two NLSs, one linked at the N-terminus and one linked at the C-terminus. In other examples, a Cas protein can be fused to 3 NLSs or not fused to an NLS. The NLS can be a monopartite sequence, such as an SV40 NLS, PKKKRKV (SEQ ID NO: 3), or PKKKRRV (SEQ ID NO: 4). The NLS can be a bipartite sequence, such as the NLS of nucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 5). In one specific example, a single PKKKRKV (SEQ ID NO: 3) NLS can be linked at the C-terminus of the Cas protein. One or more linkers are optionally included at the fusion site.
[0128] A Cas protein can also be operably linked to a cell-penetrating domain or protein transduction domain. For example, the cell-penetrating domain can be derived from the HIV-1 TAT protein, the TLM cell-penetrating motif from the human hepatitis B virus, MPG, Pep-1, VP22, a cell-penetrating peptide from the herpes simplex virus, or a polyarginine peptide sequence. See, e.g., WO 2014 / 089290 and WO 2013 / 176772 Each of these documents is incorporated herein by reference in its entirety for all purposes. The cell-penetrating domain can be located at the N-terminus, C-terminus, or anywhere within the Cas protein.
[0129] The Cas proteins can also be operably linked to a heterologous polypeptide to facilitate tracking or purification, such as a fluorescent protein, a purification tag, or an epitope tag. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, ZsGreenl, Azami Green, monomeric Azami Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent protein (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent protein (e.g., eBFP, eBFP2, Cypet, mKalama, GFPuv, CyPet, AmCyanl, Midoriishi-Cyan), cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent protein (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent protein (e.g., mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomato), and any other suitable fluorescent protein. Examples of tags include glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0130] The Cas proteins can also be tethered to a labeled nucleic acid. Such tethering (i.e., physical linkage) can be achieved through covalent or non-covalent interactions, and the tethering can be direct (e.g., through direct fusion or chemical conjugation, which can be achieved through modification of cysteine or lysine residues on the protein or intron modification) or can be through one or more intermediate linker or adaptor molecules such as streptavidin or aptamer. See, e.g., Pierce et al., (2005) Mini Rev. Med. Chem. 5(1): 41-55; Duckworth et al., (2007) Angew. Chem. Int.Ed. Engl. 46(46):8819-8822; Schaeffer and Dixon, (2009) Australian J. Chem. 62(10): 1328-1332, Goodman et al., (2009) Chembiochem .10(9): 1551-1557; and Khatwani et al., (2012) Bioorg. Med. Chem. 20(14):4532-4539, each of which is incorporated by reference herein in its entirety for all purposes. Non-covalent strategies for synthesizing protein-nucleic acid conjugates include biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by linking appropriately functionalized nucleic acids and proteins using a variety of chemical reactions. Some of these chemical reactions involve attaching oligonucleotides directly to amino acid residues on the surface of a protein (e.g., lysine amines or cysteine thiols), while other more complex schemes require post-translational modification of the protein or the involvement of catalytic or reactive protein domains. Methods for covalent linkage of proteins to nucleic acids can include, for example, chemical cross-linking of oligonucleotides to lysine or cysteine residues of a protein, expressed protein ligation, chemical enzymatic methods, and use of photoaptamers. The labeled nucleic acid can be tethered to a C-terminal, N-terminal, or internal region within the Cas protein. In one example, the labeled nucleic acid is tethered to a C-terminal or N-terminal of the Cas protein. Likewise, the Cas protein can be tethered to a 5’-end, 3’-end, or internal region within the labeled nucleic acid. That is, the labeled nucleic acid can be tethered in any orientation and polarity. For example, the Cas protein can be tethered to a 5’-end or 3’-end of the labeled nucleic acid.
[0131] The Cas protein can be provided in any form. For example, the Cas protein can be provided in the form of a protein, such as a Cas protein complexed with a gRNA. Alternatively, the Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as an RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon-optimized for efficient translation into a protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to substitute codons that have a higher frequency of usage in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, as compared to a naturally-occurring polynucleotide sequence. When the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein can be transiently, conditionally, or constitutively expressed in the cell.
[0132] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. An expression construct comprises any nucleic acid construct capable of directing the expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and can transfer such nucleic acid sequence of interest to a target cell. For example, the nucleic acid encoding the Cas protein can be in a vector that includes DNA encoding a gRNA. Alternatively, it can be in a vector or plasmid separate from a vector that includes DNA encoding a gRNA. Promoters that can be used in an expression construct include promoters active in one or more of, for example, a eukaryotic cell, a human cell, a non-human cell, a mammalian cell, a non-human mammalian cell, a rodent cell, a mouse cell, a rat cell, a pluripotent cell, an embryonic stem (ES) cell, an adult stem cell, a developmentally restricted progenitor cell, an induced pluripotent stem (iPS) cell, or a single cell stage embryo. Such a promoter can be, for example, a conditional promoter, an inducible promoter, a constitutive promoter, or a tissue-specific promoter. Optionally, the promoter can be a bi-directional promoter that drives expression of the Cas protein in one direction and the expression of the guide RNA in the other direction. Such a bi-directional promoter can consist of (1) a full, conventional, single-direction Pol III promoter, which contains 3 external control elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box; and (2) a second, essentially Pol III promoter, which includes a PSE and a TATA box fused in reverse orientation to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and TATA box, and the promoter can be bi-directional by creating a hybrid promoter, where reverse transcription is controlled by an additional PSE and TATA box derived from the U6 promoter. See, e.g., US 2016 / 0074535, which is incorporated by reference herein in its entirety for all purposes. Using a bi-directional promoter to express genes encoding the Cas protein and the guide RNA simultaneously allows for the generation of compact expression cassettes to facilitate delivery.
[0133] Different promoters can be used to drive Cas expression or Cas9 expression. In some methods, a small promoter is used so that the Cas or Cas9 coding sequence can fit into an AAV construct. For example, Cas or Cas9 and one or more gRNAs (e.g., 1 gRNA or 2 gRNAs or 3 gRNAs or 4 gRNAs) can be delivered via LNP-mediated delivery (e.g., in the form of RNA) or adeno-associated viral (AAV)-mediated delivery (e.g., AAV2-mediated delivery, AAV5-mediated delivery, AAV8-mediated delivery, or AAV7m8-mediated delivery). For example, Cas9 mRNA and gRNAs targeting a target genomic locus can be delivered via LNP-mediated delivery, or DNA encoding Cas9 and DNA encoding gRNAs targeting a target genomic locus can be delivered via AAV-mediated delivery. Cas or Cas9 and gRNAs can be delivered in a single AAV or via two separate AAVs. For example, a first AAV can carry a Cas or Cas9 expression cassette, and a second AAV can carry a gRNA expression cassette. Similarly, a first AAV can carry a Cas or Cas9 expression cassette, and a second AAV can carry two or more gRNA expression cassettes. Alternatively, a single AAV can carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and a gRNA expression cassette (e.g., a gRNA coding sequence operably linked to a promoter). Similarly, a single AAV can carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and two or more gRNA expression cassettes (e.g., gRNA coding sequences operably linked to a promoter). Different promoters can be used to drive expression of gRNAs, such as a U6 promoter or a small tRNA Gin. Likewise, different promoters can be used to drive Cas9 expression. For example, a small promoter is used so that the Cas9 coding sequence can fit into an AAV construct. Similarly, a small Cas9 protein (e.g., SaCas9 or CjCas9) is used to maximize AAV packaging capacity.
[0134] A Cas protein provided as mRNA can be modified to improve stability and / or immunogenicity properties. Modifications can be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1 -methyl-pseudouridine, and 5-methyl-cytidine. The mRNA encoding the Cas protein can also be capped. The cap can be, for example, a cap 1 structure in which the +1 ribonucleotide is methylated at the 2'0 position of the ribose. For example, capping can produce superior activity in vivo (e.g., by mimicking a natural cap), can produce a natural structure that reduces stimulation of the host's innate immune system (e.g., can reduce activation of pattern recognition receptors in the innate immune system). The mRNA encoding the Cas protein can also be polyadenylated (to include a poly(A) tail). The mRNA encoding the Cas protein can also be modified to include pseudouridines (e.g., can be completely replaced with pseudouridines). As another example, capped and polyadenylated Cas mRNA containing N1 -methyl pseudouridines can be used. As another example, Cas mRNA that is completely replaced with pseudouridines (i.e., all standard uracil residues are replaced with pseudouridines, which is an isomer of uridine in which the uracil is attached through a carbon-carbon bond rather than a nitrogen-carbon bond) can be used. Likewise, Cas mRNA can be modified by depleting uridine through the use of synonymous codons. For example, capped and polyadenylated Cas mRNA that is completely replaced with pseudouridines can be used.
[0135] A Cas mRNA can include modified uridines at at least one, multiple, or all uridine positions. A modified uridine can be a uridine that is modified at the 5 position (e.g., with a halogen, a methyl group, or an ethyl group). A modified uridine can be a pseudouridine that is modified at the 1 position (e.g., a pseudouridine modified with a halogen, a methyl group, or an ethyl group). A modified uridine can be, for example, a pseudouridine, a N1 -methyl pseudouridine, a 5-methoxyuridine, a 5-iodouridine, or a combination thereof. In some examples, the modified uridine is a 5-methoxyuridine. In some examples, the modified uridine is a 5-iodouridine. In some examples, the modified uridine is a pseudouridine. In some examples, the modified uridine is a N1 -methyl pseudouridine. In some examples, the modified uridine is a combination of a pseudouridine and a N1 -methyl pseudouridine. In some examples, the modified uridine is a combination of a pseudouridine and a 5-methoxyuridine. In some examples, the modified uridine is a combination of a N1 -methyl pseudouridine and a 5-methoxyuridine. In some examples, the modified uridine is a combination of a 5-iodouridine and a N1 -methyl pseudouridine. In some examples, the modified uridine is a combination of a pseudouridine and a 5-iodouridine. In some examples, the modified uridine is a combination of a 5-iodouridine and a 5-methoxyuridine.
[0136] The Cas mRNA disclosed herein can also include a 5' cap, such as Cap0, Cap1, or Cap2. A 5' cap is generally a 7-methylguanosine ribonucleotide (which can be further modified, e.g., with respect to ARCA) linked via a 5'-triphosphate to the 5' position of the first nucleotide of the 5' to 3' strand of the mRNA (i.e., the first cap proximal nucleotide). In Cap0, the riboses of the first and second cap proximal nucleotides of the mRNA both comprise a 2'-hydroxyl. In Cap1, the riboses of the first and second cap proximal nucleotides of the mRNA comprise a 2'-methoxy and a 2'-hydroxyl, respectively. In Cap2, the riboses of the first and second cap proximal nucleotides of the mRNA both comprise a 2'-methoxy. See, e.g., Katibah et al., (2014) Proc. Natl. Acad. Sci. U.S.A. 111(33): 12025-30 and Abbas et al., (2017) Proc. Natl. Acad. Sci. U.S.A. 114(11): E2106-E2115, each of which is incorporated by reference herein in its entirety for all purposes. Most endogenous higher eukaryotic mRNAs, including mammalian mRNAs such as human mRNAs, include Cap1 or Cap2. Cap0 and other cap structures that are different from Cap1 and Cap2 can be immunogenic in mammals such as humans because components of the innate immune system, such as IFIT-1 and IFIT-5, recognize them as non-self, which can result in elevated levels of cytokines including Type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, can also compete with eIF4E to bind mRNA with caps other than Cap1 or Cap2, potentially inhibiting translation of the mRNA.
[0137] A cap can be included co-transcriptionally. For example, ARCA (anti-reverse cap analog; Thermo Fisher Scientific catalog no. AM8045) is a cap analog that includes a 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of a guanosine ribonucleotide that can be incorporated into a transcript at the start in vitro. ARCA produces a Cap0 cap, in which the 2' position of the first cap proximal nucleotide is a hydroxyl. See, e.g., Stepinski et al., (2001) RNA 7: 1486-1495, which is incorporated by reference herein in its entirety for all purposes.
[0138] CleanCap ™ AG (m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnologies catalog no. N-7113) or CleanCap ™GG (m7G(5')ppp(5')(2'OMeG)pG; TriLink Biotechnologies catalog number N-7133) can be used to provide a Capl structure cotranscriptionally. A 3'-O-methylated version of CleanCap ™ AG and CleanCap ™ GG is also available from TriLink Biotechnologies, available under catalog numbers N-7413 and N-7433.
[0139] Alternatively, a cap can be added to an RNA post-transcriptionally. For example, vaccinia capping enzyme is commercially available (New England Biolabs catalog number M2080S) and has RNA triphosphatase and uridine transferase activity provided by its Dl subunit and guanine methyltransferase provided by its D12 subunit. As such, vaccinia capping enzyme can add a 7-methylguanine to an RNA in the presence of S-adenosylmethionine and GTP to provide Cap0. See, e.g., Guo and Moss (1990) Proc. Natl. Acad. Sci. U.S.A. 87:4023-4027 and Mao and Shuman, (1994) J. Biol. Chem. 269:24472-24479, each of which is incorporated by reference in its entirety for all purposes.
[0140] A Cas mRNA can further include a polyadenylated (poly-A or poly(A) or poly-adenine) tail. For example, a poly-A tail can include at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 adenines, and optionally up to 300 adenines. For example, a poly-A tail can include 95, 96, 97, 98, 99, or 100 adenine nucleotides.
[0141] Any of the above modifications to a Cas mRNA (e.g., to improve stability and / or immunogenicity properties) can also be used for an RNA encoding a CtIP fusion protein or i53 protein disclosed herein.
[0142] (2) wizard RNA A“guide RNA” or“gRNA” is an RNA molecule that binds to and targets a Cas protein (e.g., a Cas9 protein) to a particular location within a target DNA. A guide RNA can comprise two segments: a“DNA-targeting segment” (also referred to as a“guide sequence”) and a“protein-binding segment.” A“segment” includes a section or region of a molecule, such as a stretch of contiguous nucleotides in an RNA. Some gRNAs, such as those for Cas9, can comprise two separate RNA molecules: an“activator RNA” (e.g., a tracrRNA) and a“targeter RNA” (e.g., a CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which can also be referred to as“single-molecule gRNAs,”“single-guide RNAs,” or“sgRNAs.” See, e.g., WO 2013 / 176772, WO 2014 / 065596, WO 2014 / 089290, WO 2014 / 093622, WO 2014 / 099750, WO 2013 / 142578, and WO 2014 / 131833, each of which is incorporated by reference herein in its entirety for all purposes. A guide RNA can refer to a CRISPR RNA (crRNA) or a combination of a crRNA and a trans-activating CRISPR RNA (tracrRNA). The crRNA and tracrRNA can be associated as a single RNA molecule (a single-guide RNA or sgRNA) or in two separate RNA molecules (a dual-guide RNA or dgRNA). For example, for Cas9, a single-guide RNA can comprise a crRNA fused (e.g., via a linker) to a tracrRNA. For Cpf1 and CasF, only a crRNA is needed to achieve binding to a target sequence. The terms“guide RNA” and“gRNA” include both dual-molecule (i.e., modular) gRNAs and single-molecule gRNAs. In some methods and compositions disclosed herein, a gRNA is a S. pyogenes Cas9 gRNA or an equivalent thereof. In some methods and compositions disclosed herein, a guide RNA is a S. aureus Cas9 gRNA or an equivalent thereof.
[0143] An exemplary dual-molecule gRNA comprises a crRNA-like ("CRISPR RNA" or "targeting factor RNA" or "crRNA" or "crRNA repeat sequence") molecule and a corresponding tracrRNA-like ("trans-activating CRISPR RNA" or "activator RNA" or "tracrRNA") molecule. The crRNA comprises a DNA-targeting segment (single-stranded) of the gRNA and a nucleotide segment that forms half of a dsRNA duplex of the protein-binding segment of the gRNA. Examples of crRNA tails (e.g., for S. pyogenes Cas9) located downstream (3') of the DNA-targeting segment include, consist essentially of, or consist of GUUUUAGAGCUAUGCU (SEQ ID NO: 6) or GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 7). Any DNA-targeting segment can be joined to the 5' end of SEQ ID NO: 6 or SEQ ID NO: 7 to form a crRNA.
[0144] The corresponding tracrRNA (activator RNA) comprises a nucleotide segment that forms the other half of a dsRNA duplex of the protein-binding segment of the gRNA. The nucleotide segment of the crRNA is complementary to and hybridizes with the nucleotide segment of the tracrRNA to form a dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be considered to have a corresponding tracrRNA. Examples of tracrRNA sequences include any of, consist essentially of, or consist of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 8), AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 9), GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 10).
[0145] In systems requiring both a crRNA and a tracrRNA, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems requiring only a crRNA, the crRNA can be the gRNA. The crRNA additionally provides a single-stranded DNA targeting segment that hybridizes with the complementary strand of the target DNA. If used for modification within a cell, the exact sequence of a given crRNA or tracrRNA molecule can be designed to be specific to the species that will be using the RNA molecule. See, e.g., Mali et al., (2013) Science 339(6121):823-826; Jinek et al., (2012) Science 337(6096):816-821; Hwang et al., (2013) Nat. Biotechnol. 31(3):227-229; Jiang et al., (2013) Nat. Biotechnol. 31(3):233-239; and Cong et al., (2013) Science 339(6121):819-823, each of which is incorporated by reference in its entirety for all purposes.
[0146] The DNA targeting segment of a given gRNA (crRNA) comprises a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA targeting segment of a gRNA interacts with the target DNA in a sequence-specific manner via hybridization (i.e., base pairing). Thus, the nucleotide sequence of the DNA targeting segment can vary and determines the location within the target DNA with which the gRNA and the target DNA will interact. The DNA targeting segment of a subject gRNA can be modified to hybridize with any desired sequence within the target DNA. Naturally occurring crRNAs vary by CRISPR / Cas system and organism, but generally contain a targeting segment of between 21 and 72 nucleotides in length flanked by two direct repeat sequences (DRs) of between 21 and 46 nucleotides in length (see, e.g., WO 2014 / 131833, which is incorporated by reference in its entirety for all purposes). In the case of S. pyogenes, the DRs are 36 nucleotides in length and the targeting segment is 30 nucleotides in length. The DR located at the 3' is complementary to and hybridizes with the corresponding tracrRNA, which in turn binds to the Cas protein.
[0147] The length of the DNA-targeting segment can be, for example, at least about 12, at least about 15, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, or at least about 40 nucleotides. The length of such a DNA-targeting segment can be, for example, about 12 to about 100, about 12 to about 80, about 12 to about 50, about 12 to about 40, about 12 to about 30, about 12 to about 25, or about 12 to about 20 nucleotides. For example, the DNA-targeting segment can be about 15 to about 25 nucleotides (e.g., about 17 to about 20 nucleotides, or about 17, about 18, about 19, or about 20 nucleotides). See, e.g., US 2016 / 0024523, which is incorporated by reference herein in its entirety for all purposes. For Cas9 from S. pyogenes, a typical DNA-targeting segment is between 16 and 20 nucleotides in length or between 17 and 20 nucleotides in length. For Cas9 from S. aureus, a typical DNA-targeting segment is between 21 and 23 nucleotides in length. For Cpfl, a typical DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.
[0148] In one example, the length of the DNA-targeting segment can be about 20 nucleotides. However, shorter and longer sequences can also be used for the targeting segment (e.g., 15-25 nucleotides in length, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length). The degree of identity between the DNA-targeting segment and the corresponding guide RNA target sequence (or the degree of complementarity between the DNA-targeting segment and the other strand of the guide RNA target sequence) can be, for example, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%. The DNA-targeting segment and the corresponding guide RNA target sequence can include one or more mismatches. For example, the DNA-targeting segment of the guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches (e.g., where the total length of the guide RNA target sequence is at least 17, at least 18, at least 19, or at least 20 or more nucleotides). For example, the DNA-targeting segment of the guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches, where the total length of the guide RNA target sequence is 20 nucleotides.
[0149] A TracrRNA can be in any form (e.g., full-length tracrRNA or an activated partial tracrRNA) and have different lengths. It can comprise a primary transcript or a processed form. For example, a tracrRNA (as part of a single guide RNA or as a separate molecule as part of a dual-molecule gRNA) can comprise, consist essentially of, or consist of all or a portion of a wild-type tracrRNA sequence (e.g., about or more than about 20, about or more than about 26, about or more than about 32, about or more than about 45, about or more than about 48, about or more than about 54, about or more than about 63, about or more than about 67, about or more than about 85 or more nucleotides of a wild-type tracrRNA sequence). Examples of wild-type tracrRNA sequences from S. pyogenes include 171 -nucleotide, 89-nucleotide, 75-nucleotide, and 65-nucleotide versions. See, e.g., Deltcheva et al. (201 1 ) Nature 471 (7340):602-607; WO 2014 / 093661, each of which is incorporated by reference herein in its entirety for all purposes. Examples of tracrRNAs within a single guide RNA (sgRNA) include the tracrRNA segments present in +48, +54, +67, and +85 versions of sgRNAs, where "+n" indicates that up to +n nucleotides of a wild-type tracrRNA are included in the sgRNA. See US 8,697,359, incorporated by reference herein in its entirety for all purposes. Nature 471(7340):602-607; WO 2014 / 093661, each of which is incorporated by reference herein in its entirety for all purposes. Examples of tracrRNAs within a single guide RNA (sgRNA) include the tracrRNA segments present in +48, +54, +67, and +85 versions of sgRNAs, where "+n" indicates that up to +n nucleotides of a wild-type tracrRNA are included in the sgRNA. See US 8,697,359, incorporated by reference herein in its entirety for all purposes.
[0150] The percent complementarity between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). The percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be at least 60% over about 20 contiguous nucleotides. As one example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over 14 contiguous nucleotides at the 5' end of the complementary strand of the target DNA, and as low as 0% elsewhere. In this case, the DNA-targeting segment can be considered to be 14 nucleotides in length. As another example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over seven contiguous nucleotides at the 5' end of the complementary strand of the target DNA, and as low as 0% elsewhere. In this case, the DNA-targeting segment can be considered to be 7 nucleotides in length. In some guide RNAs, at least 17 nucleotides within the DNA-targeting segment are complementary to the complementary strand of the target DNA. For example, the DNA-targeting segment can be 20 nucleotides in length and can comprise 1, 2, or 3 mismatches to the complementary strand of the target DNA. In one example, the mismatches are not adjacent to the region of the complementary strand corresponding to the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (e.g., the mismatches are at the 5' end of the DNA-targeting segment of the guide RNA, or the mismatches and the region of the complementary strand corresponding to the PAM sequence are at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 base pairs apart).
[0151] The protein-binding segment of a gRNA can comprise two nucleotide segments that are complementary to each other. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of a subject gRNA interacts with a Cas protein, and the gRNA directs the bound Cas protein to a specific nucleotide sequence within a target DNA via the DNA-targeting segment.
[0152] A single guide RNA can include a DNA targeting segment and a scaffold sequence (i.e., a protein binding or Cas binding sequence of the guide RNA). For example, such a guide RNA can have a 5' DNA targeting segment joined to a 3' scaffold sequence. An exemplary scaffold sequence (e.g., for S. pyogenes Cas9) comprises, consists essentially of, or consists of: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU (Version 1; SEQ ID NO: 11); GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (Version 2; SEQ ID NO: 12); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (Version 3; SEQ ID NO: 13); GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (Version 4; SEQ ID NO: 14); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (Version 5; SEQ ID NO: 15); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (Version 6; SEQ ID NO: 16); GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (Version 7; SEQ ID NO: 17); or GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGGCACCGAGUCGGUGC (Version 8; SEQ ID NO: 18). In some guide sgRNAs, the four terminal U residues of Version 6 are absent. In some sgRNAs, only 1, 2, or 3 of the four terminal U residues of Version 6 are present.Guide RNA targeting can comprise any DNA-targeting segment on the 5' end of a guide RNA fused to any sequence in an exemplary guide RNA scaffold sequence on the 3' end of the guide RNA. That is, any DNA-targeting segment disclosed herein can be joined to the 5' end of any of the above-mentioned scaffold sequences to form a single guide RNA (chimeric guide RNA).
[0153] A guide RNA can comprise modifications or sequences that provide additional desirable features (e.g., modified or regulated stability; subcellular targeting; tracking with a fluorescent label; binding sites for proteins or protein complexes; etc.). A guide RNA can comprise one or more modified nucleosides or nucleotides, or one or more non-natural and / or natural components or configurations in addition to or instead of typical A, G, C, and U residues. Examples of such modifications include, for example, a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylated tail (i.e., a 3' poly(A) tail); a toehold switch sequence (e.g., to allow regulated stability and / or regulated protein and / or protein complex accessibility); a stability control sequence; a sequence that forms a dsRNA duplex (i.e., a hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows fluorescent detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, such as a protein that affects the process of homology-directed repair); and combinations thereof. Other examples of modifications include engineered stem-loop duplex structures, engineered bulge regions, engineered hairpin 3' of stem-loop duplex structures, or any combination thereof. See, e.g., US 2015 / 0376586, which is incorporated by reference herein in its entirety for all purposes. A bulge can be an unpaired region of nucleotides within a duplex comprised of a crRNA-like region and a minimal tracrRNA-like region. The bulge can comprise on one side of the duplex an unpaired 5'-XXXY-3', where X is any purine, and Y can be a nucleotide that can form a wobble pair with a nucleotide on the opposite strand; and on the other side of the duplex an unpaired region of nucleotides.
[0154] A guide RNA can comprise modified nucleosides and modified nucleotides, including, for example, one or more of: (1) alteration or replacement of one or both of the non- linking phospho oxygens and / or one or more of the linking phospho oxygens in a phosphodiester backbone linkage (exemplary backbone modifications); (2) alteration or replacement of a constituent of a ribose sugar, such as alteration or replacement of the 2’ hydroxyl on a ribose sugar (exemplary sugar modifications); (3) replacement (e.g., wholesale replacement) of a phosphoester moiety with a phosphoroamidate linker (exemplary backbone modifications); (4) modification or replacement of a naturally occurring nucleobase, including with a non-canonical nucleobase (exemplary base modifications); (5) replacement or modification of a ribose-phosphoester backbone (exemplary backbone modifications); (6) modification of the 3’ end or 5’ end of an oligonucleotide (e.g., removal, modification or replacement of a terminal phosphoester group or conjugation of a moiety, cap or linker (such 3’ or 5’ cap modifications can include sugar and / or backbone modifications)); and (7) modification or replacement of a sugar (exemplary sugar modifications). Other possible guide RNA modifications include modification or replacement of uracil or a polyuracil tract. See, e.g., WO 2015 / 048577 and US 2016 / 0237455, each of which is incorporated by reference herein in its entirety for all purposes. Similar modifications can be made to Cas-encoding nucleic acids, such as Cas mRNA. For example, Cas mRNA can be modified by depleting uridines using synonymous codons.
[0155] Chemical modifications such as those listed above can be combined to provide modified gRNAs and / or mRNAs comprising residues (nucleosides and nucleotides) that can have two, three, four or more modifications. For example, a modified residue can have a modified sugar and a modified nucleobase. In one example, every base of a gRNA is modified (e.g., all bases have a modified phosphate group, such as a phosphorothioate group). For example, all or substantially all of the phosphate groups of a gRNA can be replaced with phosphorothioate groups. Alternatively or additionally, a modified gRNA can comprise at least one modified residue at or near the 5’ end. Alternatively or additionally, a modified gRNA can comprise at least one modified residue at or near the 3’ end.
[0156] Some gRNAs comprise one, two, three or more modified residues. For example, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% of the positions in a modified gRNA can be modified nucleosides or nucleotides.
[0157] Unmodified nucleic acids can be susceptible to degradation. Foreign nucleic acids can also induce innate immune responses. Modifications can help to introduce stability and reduce immunogenicity. Some gRNAs described herein can comprise one or more modified nucleosides or nucleotides to introduce stability to intracellular or serum-based nucleases. Some modified gRNAs described herein can exhibit reduced innate immune responses when introduced into a population of cells.
[0158] The gRNAs disclosed herein can comprise backbone modifications, where the phosphate groups of the modified residues can be modified by replacing one or more of the oxygens with a different substituent. The modifications can include a wholesale replacement of unmodified phosphate moieties with modified phosphate groups as described herein. Backbone modifications of the phosphate backbone can also include alterations that result in uncharged linkers or charged linkers with asymmetric charge distribution.
[0159] Examples of modified phosphate groups include phosphorothioates, phosphoroselenoates, boranophosphates, borano phosphate esters, phosphorohydrothioates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters. The phosphorus atom in an unmodified phosphate group is achiral. However, replacing one of the non-bridging oxygens with one of the above atoms or groups of atoms can make the phosphorus atom chiral. The chiral phosphorus atom can have an "R" configuration (Rp) or an "S" configuration (Sp). The backbone can also be modified by replacing the bridging oxygens (that is, the oxygens that link the phosphate and the nucleoside) with nitrogen (bridging phosphoramidates), sulfur (bridging phosphorothioates), and carbon (bridging methylenephosphonates). The replacement can occur at either or both of the linking oxygens.
[0160] In certain backbone modifications, the phosphate groups can be replaced with non-phosphorous linkers. In some embodiments, the charged phosphate groups can be replaced with neutral moieties. Examples of moieties that can replace the phosphate groups can include, but are not limited to, for example, methylphosphonates, hydroxylamino, siloxanes, carbonates, carboxymethyl, carbamates, amides, sulfides, oxirane linkers, sulfonates, sulfonamides, thioformacetals, formacetals, oximes, methylenimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, and methylenoxymethylimino.
[0161] Nucleic acid-mimicking scaffolds can also be constructed, where phosphate linkers and ribose are replaced by nuclease-resistant nucleosides or nucleotide substitutes. Such modifications can include backbone modifications and sugar modifications. In some embodiments, nucleobases can be replaced by backbone tethers. Examples include, but are not limited to, morpholino, cyclobutyl, pyrrolidine, and peptide nucleic acid (PNA) nucleoside substitutes.
[0162] Modified nucleosides and modified nucleotides can contain one or more modifications to the glycosyl group (glycosyl modification). For example, the 2' hydroxyl group (OH) can be modified (e.g., replaced by multiple different oxygen or deoxy substituents). Modification of the 2' hydroxyl group can enhance the stability of nucleic acids because the hydroxyl group can no longer be deprotonated to form a 2'-alkoxide ion.
[0163] Examples of 2' hydroxyl modification may include alkoxy or aryloxy (OR, where "R" can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar); polyethylene glycol (PEG), O(CH2CH2O). n CH2CH2OR, where R can be, for example, H or an optionally substituted alkyl group, and n can be an integer from 0 to 20 (e.g., 0 to 4, 0 to 8, 0 to 10, 0 to 16, 1 to 4, 1 to 8, 1 to 10, 1 to 16, 1 to 20, 2 to 4, 2 to 8, 2 to 10, 2 to 16, 2 to 20, 4 to 8, 4 to 10, 4 to 16, and 4 to 20). The 2' hydroxyl modification can be 2'-O-Me. Similarly, the 2' hydroxyl modification can be a 2'-fluorine modification, where the 2' hydroxyl group is replaced with a fluoride. The 2' hydroxyl modification can include locked nucleic acids (LNAs), where the 2' hydroxyl group can be, for example, via C... 1-6 Alkylene or C 1-6 A heteroalkylene bridge is attached to the 4' carbon of the same ribose, wherein exemplary bridges may include methylene, propylene, ether, or amino bridges; O-amino (wherein the amino group may be, for example, NH2; alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino) and aminoalkoxy, O(CH2). n -Amino group (wherein the amino group can be, for example, NH2, alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino). The 2' hydroxyl modification can include an unlocked nucleic acid (UNA) where the ribose ring lacks the C2'-C3' bond. The 2' hydroxyl modification can include methoxyethyl (MOE), (OCH2CH2OCH3, for example, a PEG derivative).
[0164] Deoxy 2' modifications can include hydrogen (i.e., deoxyribose, e.g., at overhangs of partial dsRNAs); halogen (e.g., bromo, chloro, fluoro, or iodo); amino (wherein the amino group can be, e.g., NH2, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or an amino acid); NH(CH2CH2NH) n CH2CH2-amino (wherein the amino group can be, e.g., as described herein), -NHC(O)R (wherein R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), cyano; thiol; alkyl-thio-alkyl; thioalkoxy; and alkyl, cycloalkyl, aryl, alkenyl, and alkynyl, which can be optionally substituted with, e.g., amino as described herein.
[0165] Sugar modifications can include sugar groups that can also contain one or more carbons having a stereochemical configuration opposite that of the corresponding carbon in ribose. Thus, a modified nucleic acid can comprise a nucleotide containing, e.g., arabinose as a sugar. A modified nucleic acid can also comprise an abasic sugar. These abasic sugars can also be further modified at one or more of the constituent sugar atoms. A modified nucleic acid can also include one or more sugars in the L form (e.g., L-nucleosides).
[0166] The modified nucleosides and modified nucleotides described herein that can be incorporated into a modified nucleic acid can comprise a modified base, also referred to as a nucleobase. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleobases can be modified or completely replaced to provide a modified residue that can be incorporated into a modified nucleic acid. The nucleobase of a nucleotide can be independently selected from a purine, a pyrimidine, a purine analog, or a pyrimidine analog. In some embodiments, the nucleobases can include, e.g., naturally occurring and synthetic base derivatives.
[0167] In a bidirectional guide RNA, each of the crRNA and the tracrRNA can contain modifications. Such modifications can be at one or both ends of the crRNA and / or tracrRNA. In an sgRNA, one or more residues at one or both ends of the sgRNA can be chemically modified, and / or internal nucleosides can be modified, and / or the entire sgRNA can be chemically modified. Some gRNAs comprise a 5' end modification. Some gRNAs comprise a 3' end modification.
[0168] The guide RNAs disclosed herein can comprise one of the modification patterns disclosed in WO 2018 / 107028 Al, which is incorporated by reference herein in its entirety for all purposes. The guide RNAs disclosed herein can also comprise one of the structures / modification patterns disclosed in US 2017 / 0114334, which is incorporated by reference herein in its entirety for all purposes. The guide RNAs disclosed herein can also comprise one of the structures / modification patterns disclosed in WO 2017 / 136794, WO 2017 / 004279, US 2018 / 0187186, or US 2019 / 0048338, each of which is incorporated by reference herein in its entirety for all purposes.
[0169] As one example, the nucleotides at the 5' end or 3' end of the guide RNA can comprise phosphorothioate linkages (e.g., the bases can have modified phosphate groups that are phosphorothioate groups). For example, the guide RNA can comprise phosphorothioate linkages between 2, 3, or 4 terminal nucleotides at the 5' end or 3' end of the guide RNA. As another example, the nucleotides at the 5' and / or 3' end of the guide RNA can have 2'-O-methyl modifications. For example, the guide RNA can comprise 2'-O-methyl modifications at 2, 3, or 4 terminal nucleotides at the 5' end and / or 3' end (e.g., 5' end) of the guide RNA. See, e.g., WO 2017 / 173054 Al and Finn et al., (2018) Cell Rep. 22(9):2227-2235, each of which is incorporated by reference herein in its entirety for all purposes. Other possible modifications are described in greater detail elsewhere herein. In one specific example, the guide RNA comprises 2'-O-methyl analogs at the first three 5' terminal and 3' terminal RNA residues and 3' phosphorothioate internucleotide linkages. Such chemical modifications can provide, for example, greater stability and protection from exonucleases for the guide RNA, making it more long-lived in the cell than unmodified guide RNAs. Such chemical modifications can also, for example, prevent innate intracellular immune responses that can actively degrade RNA or trigger immune cascades that lead to cell death.
[0170] As an example, any of the guide RNAs described herein can comprise at least one modification. In one example, the at least one modification comprises a 2'-O-methyl (2'-O-Me) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, a 2'-fluoro (2'-F) modified nucleotide, or a combination thereof. For example, the at least one modification can comprise a 2'-O-methyl (2'-O-Me) modified nucleotide. Alternatively or additionally, the at least one modification can comprise a phosphorothioate (PS) linkage between nucleotides. Alternatively or additionally, the at least one modification can comprise a 2'-fluoro (2'-F) modified nucleotide. In one example, a guide RNA described herein comprises one or more 2'-O-methyl (2'-O-Me) modified nucleotides and one or more phosphorothioate (PS) linkages between nucleotides.
[0171] Modifications can occur anywhere in the guide RNA. As an example, a guide RNA comprises a modification at one or more of the first five nucleotides at the 5' end of the guide RNA, a modification at one or more of the last five nucleotides at the 3' end of the guide RNA, or a combination thereof. For example, a guide RNA can comprise a phosphorothioate linkage between the first four nucleotides of the guide RNA, a phosphorothioate linkage between the last four nucleotides of the guide RNA, or a combination thereof. Alternatively or additionally, a guide RNA can comprise a 2'-O-Me modified nucleotide at the first three nucleotides at the 5' end of the guide RNA, can comprise a 2'-O-Me modified nucleotide at the last three nucleotides at the 3' end of the guide RNA, or a combination thereof.
[0172] Another chemical modification that has been shown to affect nucleotide sugar rings is halogen substitution. For example, a 2'-fluoro (2'-F) substitution on a nucleotide sugar ring can increase oligonucleotide binding affinity and nuclease stability. Abasic nucleotides refer to those nucleotides that lack a nitrogenous base. Reverse bases refer to those reverse bases that have a linkage reversed from the normal 5' to 3' linkage (that is, a 5' to 5' linkage or a 3' to 3' linkage).
[0173] Abasic nucleotides can be attached with a reverse linkage. For example, an abasic nucleotide can be attached to a terminal 5' nucleotide via a 5' to 5' linkage, or an abasic nucleotide can be attached to a terminal 3' nucleotide via a 3' to 3' linkage. A reverse abasic nucleotide at a terminal 5' or 3' nucleotide can also be referred to as a reverse abasic end cap.
[0174] In one example, one or more of the first three, four, or five nucleotides at the 5' end, and one or more of the last three, four, or five nucleotides at the 3' end are modified. The modification can be, for example, 2'-0-Me, 2'-F, inverted abasic nucleotide, phosphorothioate linkage, or other well-known nucleotide modifications that increase stability and / or performance.
[0175] In another example, the first four nucleotides at the 5' end and the last four nucleotides at the 3' end can be linked with phosphorothioate linkages.
[0176] In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end can comprise 2'-0-methyl (2'-0-Me) modified nucleotides. In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end comprise 2'-fluoro (2'-F) modified nucleotides. In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end comprise inverted abasic nucleotides.
[0177] In some guide RNAs (e.g., single guide RNAs), at least one loop (e.g., both loops) of the guide RNA is modified by the insertion of a different RNA sequence that binds to one or more adaptors (i.e., adaptor proteins or domains). Such adaptor proteins can be used to further recruit one or more heterologous functional domains, such as proteins that influence the process of homology directed repair in a cell. Examples of fusion proteins that include such adaptor proteins (i.e., fusion proteins) are disclosed elsewhere herein. For example, the MS2 binding loop ggccAACAUGAGGAUCACCCAUGUCUGCAGggcc (SEQ ID NO: 19) can be substituted for the sgRNA scaffold (backbone) shown in SEQ ID NO: 11, 13, 15, or 16 or WO 2016 / 049258 and Konermann et al., (2015) Naturenucleotides + 13 to + 16 and nucleotides 53 to + 56 of the sgRNA backbone of the S. pyogenes CRISPR / Cas9 system described in Deltcheva et al., 2011, PNAS 108(3): 15179-15182; Hsu et al., 2014, Cell 157(7): 1262-1278; and Jinek et al., 2012, Science 337(6090): 816-821, each of which is incorporated by reference herein in its entirety for all purposes. The guide RNA numbering used herein refers to the nucleotide numbering in the guide RNA scaffold sequence (i.e., the sequence downstream of the DNA targeting segment of the guide RNA). For example, the first nucleotide of the guide RNA scaffold is +1, the second nucleotide of the scaffold is +2, and so on. The residues corresponding to nucleotides + 13 to + 16 in SEQ ID NO: 11, 13, 15, or 16 are the loop sequences in the region spanning nucleotides + 9 to + 21 in SEQ ID NO: 11, 13, 15, or 16 (referred to herein as the tetraloop region). The residues corresponding to nucleotides + 53 to + 56 in SEQ ID NO: 11, 13, 15, or 16 are the loop sequences in the region spanning nucleotides + 48 to + 61 in SEQ ID NO: 11, 13, 15, or 16 (referred to herein as the stem loop 2 region). The other stem loop sequences in SEQ ID NO: 11, 13, 15, or 16 include stem loop 1 (nucleotides + 33 to + 41) and stem loop 3 (nucleotides + 63 to + 75). The resulting structure is an sgRNA scaffold in which each of the tetraloop and stem loop 2 sequences are replaced with MS2 binding loops. The tetraloop and stem loop 2 are protruding from the Cas9 protein in such a way that the addition of the MS2 binding loops should not interfere with any Cas9 residues. Additionally, the proximity of the tetraloop and stem loop 2 sites to the DNA suggests that positioning to these positions can result in high interaction between the DNA and any recruited proteins, such as proteins that affect the process of homology directed repair. Thus, in some sgRNAs, the nucleotides corresponding to + 13 to + 16 and / or + 53 to + 56 of the guide RNA scaffold shown in SEQ ID NO: 11, 13, 15, or 16, or the corresponding residues when optimally aligned with any of these scaffolds / backbones, are replaced with a different RNA sequence that is capable of binding to one or more adaptor proteins or domains. Alternatively or additionally, an adaptor binding sequence can be added to the 5’ end or 3’ end of the guide RNA. An exemplary guide RNA scaffold comprising MS2 binding loops in the tetraloop and stem loop 2 regions can comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO: 21, 22, or 23 (e.g., SEQ ID NO: 23). An exemplary genetic single guide RNA comprising MS2 binding loops in the tetraloop and stem loop 2 regions can comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO: 24, 25, or 26 (e.g., SEQ ID NO: 26).
[0178] The guide RNA can be provided in any form. For example, the gRNA can be provided in the form of RNA, as two molecules (separate crRNA and tracrRNA) or as one molecule (sgRNA), and optionally in the form of a complex with a Cas protein. The gRNA can also be provided in the form of DNA encoding the gRNA. The DNA encoding the gRNA can encode a single RNA molecule (sgRNA) or separate RNA molecules (e.g., separate crRNA and tracrRNA). In the latter case, the DNA encoding the gRNA can be provided as one DNA molecule or as separate DNA molecules encoding the crRNA and tracrRNA, respectively.
[0179] When the gRNA is provided in the form of DNA, the gRNA can be transiently, conditionally, or constitutively expressed in the cell. The DNA encoding the gRNA can be stably integrated into the genome of the cell and operably linked to a promoter that is active in the cell. Alternatively, the DNA encoding the gRNA can be operably linked to a promoter in an expression construct. For example, the DNA encoding the gRNA can be in a vector comprising heterologous nucleic acid. Promoters that can be used in such expression constructs include promoters that are active in one or more cells in, for example, a eukaryotic cell, a human cell, a non-human cell, a mammalian cell, a non-human mammalian cell, a rodent cell, a mouse cell, a rat cell, a pluripotent cell, an embryonic stem (ES) cell, an adult stem cell, a developmentally restricted progenitor cell, an induced pluripotent stem (iPS) cell, or a single-cell stage embryo. Such promoters can be, for example, a conditional promoter, an inducible promoter, a constitutive promoter, or a tissue-specific promoter. Such promoters can also be, for example, a bi-directional promoter. Specific examples of suitable promoters include an RNA polymerase III promoter, such as a human U6 promoter, a rat U6 polymerase III promoter, or a mouse U6 polymerase III promoter.
[0180] Alternatively, the gRNA can be prepared by various other methods. For example, the gRNA can be prepared by in vitro transcription using, for example, a T7 RNA polymerase (see, e.g., WO 2014 / 089290 and WO 2014 / 065596, each of which is incorporated by reference herein in its entirety for all purposes). The guide RNA can also be a synthetically produced molecule prepared by chemical synthesis. For example, the guide RNA can be chemically synthesized to include 2'-0-methyl analogs and 3' thiophosphate internucleotide linkages at the first three 5' terminal and 3' terminal RNA residues.
[0181] A guide RNA (or nucleic acid encoding a guide RNA) can be in a composition comprising one or more guide RNAs (e.g., 1, 2, 3, 4, or more guide RNAs) and a carrier that increases stability of the guide RNA (e.g., extends the period of time that degradation products remain below a threshold value, such as less than 0.5% by weight of the starting nucleic acid or protein, under a given storage condition (e.g., -20°C, 4°C, or ambient temperature); or increases stability in vivo). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipidic spirals, and lipidic microtubes. Such compositions can also comprise a Cas protein, such as a Cas9 protein, or a nucleic acid encoding a Cas protein.
[0182] (3) Guide RNA Target sequence A target DNA for a guide RNA comprises a nucleic acid sequence present in DNA to which the DNA-targeting segment of the gRNA will bind, provided that sufficient binding conditions are present. Suitable DNA / RNA binding conditions include physiological conditions that are typically present in a cell. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 3rded. (Sambrook et al., Harbor Laboratory Press 2001), which is incorporated by reference herein in its entirety for all purposes). A target DNA strand that is complementary to and hybridizes with a gRNA can be referred to as the “complementary strand,” and a target DNA strand that is complementary to the “complementary strand” (and thus not complementary to the Cas protein or gRNA) can be referred to as the “non-complementary strand” or “template strand.”
[0183] A target DNA comprises a sequence on the complementary strand that hybridizes to a guide RNA and a corresponding sequence on the non-complementary strand (e.g., adjacent to a protospacer adjacent motif (PAM)). The term "guide RNA target sequence" as used herein refers specifically to the sequence on the non-complementary strand that corresponds to (i.e., is the reverse complement of) the sequence on the complementary strand to which the guide RNA hybridizes. That is, the guide RNA target sequence refers to the sequence on the non-complementary strand that is adjacent to (e.g., upstream or 5' of, in the case of Cas9) the PAM. The guide RNA target sequence is equivalent to the DNA targeting segment of the guide RNA, but with thymidines instead of uridines. As one example, the guide RNA target sequence for a SpCas9 enzyme can refer to the sequence on the non-complementary strand upstream of the 5'-NGG-3' PAM. The guide RNA is designed to be complementary to the complementary strand of the target DNA, with hybridization between the DNA targeting segment of the guide RNA and the complementary strand of the target DNA promoting formation of a CRISPR complex. Perfect complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex. If a guide RNA is referred to herein as targeting a guide RNA target sequence, it means that the guide RNA hybridizes to the complementary strand sequence of the target DNA that is the reverse complement of the guide RNA target sequence on the non-complementary strand.
[0184] A target DNA or guide RNA target sequence can comprise any polynucleotide and can be located, for example, in the nucleus or cytoplasm of a cell or within an organelle of a cell, such as a mitochondrion or chloroplast. A target DNA or guide RNA target sequence can be any nucleic acid sequence endogenous or exogenous to a cell. A guide RNA target sequence can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence) or can comprise both.
[0185] A target gene can include a gene expressed in a particular organ or tissue, such as the liver. A target gene can include a disease-associated gene. A disease-associated gene refers to any gene that produces a transcriptional or translational product at abnormal levels or in abnormal form in cells derived from a tissue affected by a disease, as compared to a tissue or cell of a control that is not diseased. It can be a gene that is expressed at an abnormally high level, where the altered expression is associated with the occurrence and / or progression of the disease. A disease-associated gene also refers to a gene with a mutation or genetic variation that contributes to the etiology of a disease. The transcriptional or translational product can be known or unknown and can be at normal or abnormal levels. A target gene can also be a gene involved in a pathway related to a disease or condition, or a gene that when overexpressed can mimic such a disease or condition. A target gene can also be a gene expressed or overexpressed in one or more types of cancer. See, e.g., Santarius et al., (2010) Nat. Rev. Cancer 10(1): 59-64, which is incorporated by reference herein in its entirety for all purposes.
[0186] Site-specific binding and cleavage of the Cas protein to the target DNA can occur at a location determined by both: (i) base pairing complementarity between the guide RNA and the complementary strand of the target DNA, and (ii) a short motif (referred to as a protospacer adjacent motif (PAM)) in the non-complementary strand of the target DNA. The PAM can be side-attached to the guide RNA target sequence. Optionally, the guide RNA target sequence can have a PAM side-attached at the 3' end (e.g., for Cas9). Alternatively, the guide RNA target sequence can have a PAM side-attached at the 5' end (e.g., for Cpf1). For example, the cleavage site of the Cas protein can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream of the PAM sequence (e.g., within the guide RNA target sequence). In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5'-N1GG-3', where N1 is any DNA nucleotide, and where the PAM is immediately adjacent to the 3' of the guide RNA target sequence on the non-complementary strand of the target DNA. Therefore, the sequence corresponding to PAM on the complementary strand (i.e., the reverse complementary sequence) would be 5'-CCN2-3', where N2 is any DNA nucleotide and is the 5' of the sequence immediately adjacent to the DNA targeting segment of the guide RNA that hybridizes with it on the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1=C and N2=G; N1=G and N2=C; N1=A and N2=T; N1=T and N2=A). For Cas9 from Staphylococcus aureus, PAM could be NNGRRT or NNGRR, where N can be A, G, C, or T, and R can be G or A. For Cas9 from Campylobacter jejuni, PAM could be, for example, NNNACAC or NNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (e.g., for FnCpf1), the PAM sequence may be located upstream of the 5' end and have the sequence 5'-TTN-3'. For DpbCasX, the PAM may have the sequence 5'-TTCN-3'. For CasΦ, the PAM may have the sequence 5'-TBN-3', where B is G, T, or C.
[0187] An example of a guide RNA target sequence is a DNA sequence consisting of 20 nucleotides immediately preceding the NGG motif recognized by the SpCas9 protein. For instance, two examples of guide RNA target sequences plus PAM are GN... 19 NGG or N 20NGG. See, e.g., WO 2014 / 165825, which is incorporated by reference herein in its entirety for all purposes. The guanine at the 5' end can facilitate transcription by RNA polymerase in cells. Other examples of guide RNA target sequences plus PAM can include two guanine nucleotides at the 5' end (e.g., GGN 20 NGG) to facilitate efficient transcription by T7 polymerase in vitro. See, e.g., WO 2014 / 065596, which is incorporated by reference herein in its entirety for all purposes. Other guide RNA target sequences plus PAM can be between 4 and 22 nucleotides in length, including a 5' G or GG and a 3' GG or NGG. Other guide RNA target sequences plus PAM can be between 14 and 20 nucleotides in length.
[0188] Formation of a CRISPR complex hybridized to a target DNA can result in cleavage of one or both strands of the target DNA within or near the region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and the reverse complement of the guide RNA hybridized thereto on the complementary strand). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a defined position relative to the PAM sequence). A "cleavage site" comprises the position of the target DNA at which a Cas protein produces a single-strand break or a double-strand break. The cleavage site can be on only one strand (e.g., when a nickase is used) or on both strands of double-stranded DNA. The cleavage site can be at the same position on both strands (producing a blunt end; e.g., Cas9) or can be at different sites on each strand (producing a staggered end (i.e., an overhang); e.g., Cpf1). A staggered end can be produced, for example, by using two Cas proteins, each producing a single-strand break at a different cleavage site on a different strand, thereby producing a double-strand break. For example, a first nickase can form a single-strand break on a first strand of double-stranded DNA (dsDNA) and a second nickase can form a single-strand break on a second strand of the dsDNA, such that an overhang sequence is formed. In some cases, the guide RNA target sequence or cleavage site for the nickase on the first strand is separated from the guide RNA target sequence or cleavage site for the nickase on the second strand by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 250, at least 500, or at least 1,000 base pairs.
[0189] B. CtBP-interacting protein (CtIP) fusion proteins The compositions or combinations and corresponding methods disclosed herein utilize a CtIP protein, which can be part of a fusion protein that can bind to a guide RNA disclosed elsewhere herein. The fusion proteins disclosed herein can be used in the methods described herein to bring CtIP in proximity to a cleaved target genomic locus to facilitate homology directed repair. The nucleic acids encoding the fusion proteins can be genomically integrated into a cell or animal (e.g., a cell or animal comprising a genomically integrated fusion protein expression cassette), or the fusion proteins or nucleic acids can be introduced into such cells and animals using the methods disclosed elsewhere herein (e.g., LNP-mediated delivery or AAV-mediated delivery).
[0190] Such fusions include: (a) an adaptor that specifically binds to an adaptor binding element within a guide RNA (i.e., an adaptor domain or adaptor protein); and (b) a CtIP protein. For example, a fusion protein can comprise: (a) an MS2 coat protein adaptor that specifically binds to one or more MS2 aptamers in a guide RNA (e.g., two MS2 aptamers in separate locations in a guide RNA); and (b) a CtIP protein.
[0191] CtIP can be fused directly to the adaptor. Alternatively, CtIP can be connected to the adaptor via a linker or combination of linkers or via one or more additional domains. Linkers that can be used in these fusion proteins can comprise any sequence that does not interfere with the function of the fusion protein. Exemplary linkers are short (e.g., 2-20 amino acids) and are typically flexible (e.g., comprising amino acids with high degrees of freedom, such as glycine, alanine, and serine). Some specific examples of linkers include one or more units consisting of GGGS (SEQ ID NO: 28) or GGGGS (SEQ ID NO: 29) (such as two, three, four, or more repeats of any combination of GGGS (SEQ ID NO: 28) or GGGGS (SEQ ID NO: 29)). Other linker sequences can also be used.
[0192] CtIP and adaptor can be in any order within the fusion protein. As an option, CtIP can be at the C-terminus of the adaptor and the adaptor can be at the N-terminus of the CtIP. For example, CtIP can be at the C-terminus of the fusion protein and the adaptor can be at the N-terminus of the fusion protein. However, CtIP can be at the C-terminus of the adaptor without being at the C-terminus of the fusion protein (e.g., if a nuclear localization signal is at the C-terminus of the fusion protein). Likewise, the adaptor can be at the N-terminus of the CtIP without being at the N-terminus of the fusion protein (e.g., if a nuclear localization signal is at the N-terminus of the fusion protein). As another option, CtIP can be at the N-terminus of the adaptor and the adaptor can be at the N-terminus of the CtIP. For example, CtIP can be at the N-terminus of the fusion protein and the adaptor can be at the C-terminus of the fusion protein.
[0193] The fusion proteins described herein can also be operably linked or fused to additional heterologous polypeptides. The fused or linked heterologous polypeptide can be located at the N-terminus, C-terminus, or anywhere within the fusion protein. For example, the CtIP protein can also include a nuclear localization signal. A specific example of such a protein includes an MS2 coat protein (adaptor) linked (directly or via a NLS) to the C-terminus of the CtIP protein. Such a protein can include, from N-terminus to C-terminus: the MCP; a nuclear localization signal; and the CtIP protein. In one example, the fusion protein can comprise an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the fusion protein sequence set forth in SEQ ID NO: 30. In another example, the fusion protein can consist essentially of an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the fusion protein sequence set forth in SEQ ID NO: 30. In another example, the fusion protein can consist of an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the fusion protein sequence set forth in SEQ ID NO: 30. In one example, the fusion protein can be encoded by a nucleic acid comprising a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31. In another example, the fusion protein can be encoded by a nucleic acid that consists essentially of a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31. In another example, the fusion protein can be encoded by a nucleic acid that consists of a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 31.
[0194] In one example, the fusion protein comprises the sequence set forth in SEQ ID NO: 30. In another example, the fusion protein consists essentially of the sequence set forth in SEQ ID NO: 30. In another example, the fusion protein consists of the sequence set forth in SEQ ID NO: 30. In one example, the nucleic acid encoding the fusion protein comprises the sequence set forth in SEQ ID NO: 31. In another example, the nucleic acid encoding the fusion protein consists essentially of the sequence set forth in SEQ ID NO: 31. In another example, the nucleic acid encoding the fusion protein consists of the sequence set forth in SEQ ID NO: 31.
[0195] The fusion protein can also be fused or linked to one or more heterologous polypeptides that provide for subcellular localization. Such heterologous polypeptides can include, for example, one or more nuclear localization signals (NLS), such as an SV40 NLS and / or an a-importin NLS for targeting to the nucleus, a mitochondrial localization signal for targeting to mitochondria, an ER retention signal, and the like. See, e.g., Lange et al. (2007) J. Biol. Chem. 282:5101-5105, which is incorporated by reference herein in its entirety for all purposes. The NLS can comprise a stretch of basic amino acids, and can be a one-part sequence or a two-part sequence. Optionally, the fusion protein comprises two or more NLS, including an NLS at the N-terminus (e.g., an a- importin NLS) and / or an NLS at the C-terminus (e.g., an SV40 NLS).
[0196] The fusion protein can also be operably linked to a cell-penetrating domain or protein transduction domain. For example, the cell-penetrating domain can be derived from the HIV-1 TAT protein, the TLM cell-penetrating motif from the human hepatitis B virus, MPG, Pep-1, VP22, a cell-penetrating peptide from the herpes simplex virus, or a polyarginine peptide sequence. See, e.g., WO 2014 / 089290 and WO 2013 / 176772, each of which is incorporated by reference herein in its entirety for all purposes. As another example, the fusion protein can be fused or linked to a heterologous polypeptide that provides for increased or decreased stability.
[0197] The fusion protein can also be operably linked to a heterologous polypeptide to facilitate tracking or purification, such as a fluorescent protein, a purification tag, or an epitope tag. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, ZsGreenl, Azami Green, monomeric Azami Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent protein (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent protein (e.g., eBFP, eBFP2, Cypet, mKalama, GFPuv, Cyblue, T-sapphire), cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent protein (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent protein (e.g., mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomato), and any other suitable fluorescent protein. Examples of tags include glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0198] The fusion protein can also be tethered to a labeled nucleic acid. Such tethering (i.e., physical linkage) can be achieved through covalent or non-covalent interactions, and the tethering can be direct (e.g., through direct fusion or chemical conjugation, which can be achieved through modification of cysteine or lysine residues on the proteins or intron modification) or can be through one or more intermediate linker or adaptor molecules such as streptavidin or aptamer. See, e.g., Pierce et al., (2005) Mini Rev. Med. Chem. 5(1): 41-55; Duckworth et al., (2007) Angew. Chem. Int.Ed. Engl. 46(46):8819-8822; Schaeffer and Dixon, (2009) Australian J. Chem. 62(10): 1328-1332, Goodman et al., (2009) Chembiochem .10(9): 1551-1557; and Khatwani et al., (2012) Bioorg. Med. Chem. 20(14):4532-4539, each of which is incorporated by reference in its entirety for all purposes. Non-covalent strategies for the synthesis of protein-nucleic acid conjugates include biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by using a variety of chemical reactions to link appropriately functionalized nucleic acids and proteins. Some of these chemical reactions involve the direct attachment of oligonucleotides to amino acid residues on the surface of the protein (e.g., lysine amines or cysteine thiols), while other more complex schemes require post-translational modification of the protein or the involvement of catalytic or reactive protein domains. Methods for covalent linkage of proteins to nucleic acids can include, for example, chemical cross-linking of oligonucleotides to lysine or cysteine residues of proteins, expressed protein ligation, chemical enzymatic methods, and the use of photoaptamers. The labeled nucleic acid can be tethered to a C-terminal, N-terminal, or internal region within a fusion protein. Likewise, the fusion protein can be tethered to a 5' end, 3' end, or internal region within the labeled nucleic acid. That is, the labeled nucleic acid can be tethered in any orientation and polarity.
[0199] (1) adaptor protein or adaptor domain An adaptor (i.e., an adaptor domain or an adaptor protein) is a nucleic acid binding domain (e.g., a DNA binding domain and / or an RNA binding domain) that specifically recognizes and binds to different sequences (e.g., binds to different DNA and / or RNA sequences, such as aptamers in a sequence-specific manner). An aptamer comprises a nucleic acid that, by virtue of its ability to adopt a specific three-dimensional conformation, can bind to a target molecule with high affinity and specificity. For example, such adaptors can bind to specific RNA sequences and secondary structures. These sequences (i.e., adaptor binding elements) can be engineered into a guide RNA. For example, an MS2 aptamer can be engineered into a guide RNA to specifically bind to an MS2 coat protein (MCP). In one example, an adaptor can comprise an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the MCP sequence set forth in SEQ ID NO: 32. In another example, an adaptor can essentially consist of an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the MCP sequence set forth in SEQ ID NO: 32. In another example, an adaptor can consist of an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the MCP sequence set forth in SEQ ID NO: 32. In one example, an adaptor can be encoded by a nucleic acid comprising a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33. In another example, an adaptor can be encoded by a nucleic acid that essentially consists of a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33. In another example, an adaptor can be encoded by a nucleic acid that consists of a sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 33.
[0200] In one example, the adaptor comprises the sequence set forth in SEQ ID NO: 32. In another example, the adaptor consists essentially of the sequence set forth in SEQ ID NO: 32. In another example, the adaptor consists of the sequence set forth in SEQ ID NO: 32. In one example, the nucleic acid encoding the adaptor comprises the sequence set forth in SEQ ID NO: 33. In another example, the nucleic acid encoding the adaptor consists essentially of the sequence set forth in SEQ ID NO: 33. In another example, the nucleic acid encoding the adaptor consists of the sequence set forth in SEQ ID NO: 33.
[0201] Some specific examples of adaptors and targets include RNA binding protein / aptamer combinations present within phage coat protein diversity. For example, the following adaptor proteins or functional fragments or variants thereof can be used: MS2 coat protein (MCP), PP7, Qp, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, Φ Cb8r, Φ Cb12r, ΦCb23r, 7s, and PRR1. See, e.g., WO 2016 / 049258, which is incorporated by reference herein in its entirety for all purposes. A functional fragment or functional variant of an adaptor protein is a functional fragment or functional variant that retains the ability to bind to a particular adaptor binding element (e.g., the ability to bind to a particular adaptor binding sequence in a sequence-specific manner). For example, a PP7 Pseudomonas phage coat protein variant in which amino acids 68-69 are mutated to SG and amino acids 70-75 are deleted from the wild-type protein can be used. See, e.g., Wu et al., (2012) Biophys. J. 102(12):2936-2944; and Chao et al., (2007) Nat. Struct. Mol. Biol. 15(1): 103-105, each of which is incorporated by reference herein in its entirety for all purposes. Likewise, MCP variants such as the N55K mutant can be used. See, e.g., Spingola and Peabody, (1994) J. Biol. Chem. 269(12):9006-9010, which is incorporated by reference herein in its entirety for all purposes.
[0202] Other examples of adaptor proteins that can be used include all or part of the endoribonuclease Csy4 or lambda N protein (e.g., its DNA binding). See, e.g., US 2016 / 0312198, which is incorporated by reference herein in its entirety for all purposes.
[0203] (2) CtIP The compositions or combinations and corresponding methods disclosed herein utilize a CtIP protein. In one example, the CtIP protein used in the methods, compositions, and combinations disclosed herein is a human CtIP protein. Human C-terminal binding protein (CtBP) interacting protein (CtIP) (also known as DNA endonuclease RBBP8, RBBP8, retinoblastoma binding protein 8, RBBP-8, retinoblastoma interacting protein, and myosin-like RIM, sporulation in the absence of SPO11 protein 2 homolog SAE2) is assigned UniProt reference number Q99807 and is encoded by the gene assigned NCBI GeneID 5932 CTIP (also known as RBBP8 ). The gene is located at position 18ql 1.2 on chromosome 18 (assembly: GRCh38.p14 (GCF_000001405.40); location: NC_000018.10 (22914139..23026486)). A canonical isoform of human CtIP (UniProt reference number Q99708-1, NCBI reference number NP_002885.1) is set forth in SEQ ID NO: 37. The mRNA encoding this canonical isoform is assigned NCBI reference number NM_002894.3 (SEQ ID NO: 38). The coding sequence of this canonical isoform is assigned CCDS number CCDS11875.1 (SEQ ID NO: 39). Another isoform of human CtIP (NCBI reference number AAC14371.1) is set forth in SEQ ID NO: 34. The mRNA encoding this canonical isoform is assigned NCBI reference number U72066.1 (SEQ ID NO: 35). The coding sequence of this canonical isoform is set forth in SEQ ID NO: 36. CtIP is an endonuclease that collaborates with the MRE11-RAD50-NBN (MRN) complex in DNA end resection, the first step in double-strand break repair by the homologous recombination pathway.
[0204] In one example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein comprises a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 34. In another example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein consists essentially of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 34. In another example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 34. In another example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein comprises the sequence set forth in SEQ ID NO: 34. In another example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein consists essentially of the sequence set forth in SEQ ID NO: 34. In another example, a CtIP protein for use in the methods, compositions, and combinations disclosed herein consists of the sequence set forth in SEQ ID NO: 34.
[0205] In one example, the CtIP protein is encoded by a nucleic acid comprising a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 36. In another example, the CtIP protein is encoded by a nucleic acid consisting essentially of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 36. In another example, the CtIP protein is encoded by a nucleic acid consisting of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 36. In another example, the CtIP protein is encoded by a nucleic acid comprising the sequence set forth in SEQ ID NO: 36. In another example, the CtIP protein is encoded by a nucleic acid consisting essentially of the sequence set forth in SEQ ID NO: 36. In another example, the CtIP protein is encoded by a nucleic acid consisting of the sequence set forth in SEQ ID NO: 36.
[0206] C. Inhibitors of 53BP1 The compositions or combinations and corresponding methods disclosed herein utilize an inhibitor of the 53BP1 (i53) protein. i53 is a variant of ubiquitin that blocks the accumulation of 53BP1 at DNA damage sites. i53 binds to and occludes the ligand binding site of the 53BP1 tudor domain, thereby blocking its ability to accumulate at DNA damage sites.
[0207] In one example, an i53 protein for use in the methods, compositions, and combinations disclosed herein comprises a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 40 or 42 (e.g., 40). In another example, an i53 protein for use in the methods, compositions, and combinations disclosed herein consists essentially of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 40 or 42 (e.g., 40). In another example, an i53 protein for use in the methods, compositions, and combinations disclosed herein consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 40 or 42 (e.g., 40). In another example, an i53 protein for use in the methods, compositions, and combinations disclosed herein comprises the sequence set forth in SEQ ID NO: 40 or 42 (e.g., 40). In another example, an i53 protein for use in the methods, compositions, and combinations disclosed herein consists essentially of the sequence set forth in SEQ ID NO: 40 or 42 (e.g., 40). In another example, an i53 protein for use in the methods, compositions, and combinations disclosed herein consists of the sequence set forth in SEQ ID NO: 40 or 42 (e.g., 40).
[0208] In one example, the i53 protein is encoded by a nucleic acid comprising a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 41 or 43 (e.g., 41). In another example, the i53 protein is encoded by a nucleic acid consisting essentially of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 41 or 43 (e.g., 41). In another example, the i53 protein is encoded by a nucleic acid consisting of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 41 or 43 (e.g., 41). In another example, the i53 protein is encoded by a nucleic acid comprising the sequence set forth in SEQ ID NO: 41 or 43 (e.g., 41). In another example, the i53 protein used in the methods, compositions, and combinations disclosed herein consists essentially of the sequence set forth in SEQ ID NO: 41 or 43 (e.g., 41). In another example, the i53 protein is encoded by a nucleic acid consisting of the sequence set forth in SEQ ID NO: 41 or 43 (e.g., 41).
[0209] D. Exogenous donor nucleic acids The methods and compositions disclosed herein utilize an exogenous donor nucleic acid to modify a target genomic locus after cleavage with a Cas protein. In such methods, the Cas protein cleaves the target genomic locus (e.g., to create a double-stranded break), and the exogenous donor nucleic acid recombines with the target nucleic acid through a homology directed repair event. Optionally, the exogenous donor nucleic acid repairs the guide RNA target sequence or Cas cleavage site that was removed or disrupted, such that the allele that has been targeted cannot be re-targeted by the Cas protein. In some cases, the exogenous repair template flanks the guide RNA target sequence that is cleaved by the Cas protein within the cell.
[0210] The exogenous donor nucleic acid can comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which can be single-stranded or double-stranded, and which can be in linear or circular form. For example, the exogenous donor nucleic acid can be a single-stranded oligodeoxynucleotide (ssODN). See, e.g., Yoshimi et al., (2016) Nat. Commun.7:10431, which is incorporated by reference herein in its entirety for all purposes. Exemplary exogenous donor nucleic acids are between about 50 nucleotides to about 5 kb, between about 50 nucleotides to about 3 kb, or between about 50 to about 1,000 nucleotides in length. Other exemplary exogenous donor nucleic acids are between about 40 to about 200 nucleotides in length. For example, an exogenous donor nucleic acid can be between about 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150, 150-160, 160-170, 170-180, 180-190, or 190-200 nucleotides in length. Alternatively, an exogenous donor nucleic acid can be between about 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 nucleotides in length. Alternatively, an exogenous donor nucleic acid can be between about 1 kb-1.5 kb, 1.5 kb-2 kb, 2 kb-2.5 kb, 2.5 kb-3 kb, 3 kb-3.5 kb, 3.5 kb-4 kb, 4 kb-4.5 kb, or 4.5 kb-5 kb in length. Alternatively, an exogenous donor nucleic acid can be, for example, no more than 5 kb, 4.5 kb, 4 kb, 3.5 kb, 3 kb, 2.5 kb, 2 kb, 1.5 kb, 1 kb, 900 nucleotides, 800 nucleotides, 700 nucleotides, 600 nucleotides, 500 nucleotides, 400 nucleotides, 300 nucleotides, 200 nucleotides, 100 nucleotides, or 50 nucleotides in length. Exogenous donor nucleic acids (e.g., targeting vectors) can also be longer.
[0211] In one example, the exogenous donor nucleic acid is an ssODN between about 80 nucleotides and about 200 nucleotides in length. In another example, the exogenous donor nucleic acid is an ssODN between about 80 nucleotides and about 3 kb in length. Such ssODNs can have, for example, homology arms that are each between about 40 nucleotides and about 60 nucleotides in length. Such ssODNs can also have, for example, homology arms that are each between about 30 nucleotides and 100 nucleotides in length. The homology arms can be symmetric (e.g., 40 nucleotides each or 60 nucleotides each in length), or they can be asymmetric (e.g., one homology arm is 36 nucleotides in length and one homology arm is 91 nucleotides in length).
[0212] The exogenous donor nucleic acid can comprise a modification or sequence that provides an additional desirable feature (e.g., improved or modulated stability; tracking or detection with a fluorescent label; a binding site for a protein or protein complex; and the like). The exogenous donor nucleic acid can include one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, the exogenous donor nucleic acid can comprise one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), such as at least 1, at least 2, at least 3, at least 4, or at least 5 fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and-6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A variety of fluorescent dyes are commercially available for labeling oligonucleotides (e.g., from Integrated DNA Technologies). Such fluorescent labels (e.g., internal fluorescent labels) can be used, for example, to detect an exogenous donor nucleic acid that has been directly integrated into a cleaved target nucleic acid having overhanging ends compatible with the ends of the exogenous donor nucleic acid. The tag or label can be at the 5' end, the 3' end, or internal to the exogenous donor nucleic acid. For example, the exogenous donor nucleic acid can be conjugated at the 5' end to an IR700 fluorophore from Integrated DNA Technologies (5' IRDYE 700). ® 700).
[0213] The exogenous donor nucleic acid disclosed herein comprises homology arms. If the exogenous donor nucleic acid also includes a nucleic acid insert, the homology arms can flank the nucleic acid insert. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This terminology relates to the relative position of the homology arms to the nucleic acid insert within the exogenous donor nucleic acid. The 5' and 3' homology arms correspond to regions within the target genomic locus that are referred to herein as the "5' target sequence" and the "3' target sequence," respectively.
[0214] When the homology arms and the target sequence share a sufficient level of sequence identity to one another, the two regions "correspond" to one another to serve as substrates for a homologous recombination reaction. The term "homology" includes DNA sequences that are identical or share sequence identity to the corresponding sequence. The sequence identity between a given target sequence and the corresponding homology arm present in the exogenous donor nucleic acid can be any degree of sequence identity that allows for homologous recombination to occur. Furthermore, the corresponding homologous regions between the homology arm and the corresponding target sequence can be of any length sufficient to promote homologous recombination. Exemplary homology arms are between about 25 nucleotides to about 2.5 kb, about 25 nucleotides to about 1.5 kb, or about 25 to about 500 nucleotides in length. For example, a given homology arm (or each of the homology arms) and / or the corresponding target sequence can include a corresponding homologous region that is between about 25-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 nucleotides in length, such that the homology arm has sufficient homology to undergo homologous recombination with the corresponding target sequence within the target nucleic acid. Alternatively, a given homology arm (or each homology arm) and / or the corresponding target sequence can include a corresponding homologous region that is between about 0.5 kb to about 1 kb, about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, or about 2 kb to about 2.5 kb in length. For example, the homology arms can each be about 750 nucleotides in length. The homology arms can be symmetrical (each about the same size in length), or can be asymmetrical (one longer than the other).
[0215] When the CRISPR / Cas system is used in conjunction with an exogenous donor nucleic acid, the 5' target sequence and the 3' target sequence are optionally positioned in a location that is in sufficient proximity to the Cas cleavage site (e.g., within a location that is in sufficient proximity to the guide RNA target sequence) to facilitate the occurrence of a homologous recombination event between the target sequence and the homology arm following a single-strand break (nick) or double-strand break at the Cas cleavage site. The term "Cas cleavage site" includes a DNA sequence in which a nick or double-strand break is formed by a Cas enzyme (e.g., a Cas9 protein in complex with a guide RNA). The target sequences corresponding to the 5' and 3' homology arms of the exogenous donor nucleic acid are "positioned in sufficient proximity" to the Cas cleavage site if such a distance is to facilitate the occurrence of a homologous recombination event between the 5' and 3' target sequences and the homology arm following a single-strand break or double-strand break at the Cas cleavage site. Thus, the target sequences corresponding to the 5' and / or 3' homology arms of the exogenous donor nucleic acid can be, for example, within at least 1 nucleotide of a given Cas cleavage site, or within at least 10 nucleotides to about 1,000 nucleotides of a given Cas cleavage site. As one example, the Cas cleavage site can be immediately adjacent to at least one or both of the target sequences.
[0216] The spatial relationship of the homology arm corresponding to the exogenous donor nucleic acid and the target sequence of the Cas cleavage site can vary. For example, the target sequence can be positioned 5' to the Cas cleavage site, the target sequence can be positioned 3' to the Cas cleavage site, or the target sequence can flank the Cas cleavage site.
[0217] The exogenous donor nucleic acid can also include a nucleic acid insert comprising a DNA segment to be integrated at the target genomic locus. Integration of the nucleic acid insert at the target genomic locus can result in the addition of a nucleic acid sequence of interest to the target genomic locus, the deletion of a nucleic acid sequence of interest at the target genomic locus, or the substitution of a nucleic acid sequence of interest at the target genomic locus (i.e., a deletion and an insertion). Some exogenous donor nucleic acids are designed to insert a nucleic acid insert at the target genomic locus without any corresponding deletion at the target genomic locus. Other exogenous donor nucleic acids are designed to delete a nucleic acid sequence of interest at the target genomic locus without any corresponding insertion of a nucleic acid insert. Yet other exogenous donor nucleic acids are designed to delete a nucleic acid sequence of interest at the target genomic locus and replace it with a nucleic acid insert.
[0218] The nucleic acid insert or corresponding nucleic acid at the deleted and / or replaced target genomic locus can be of various lengths. Exemplary nucleic acid inserts or corresponding nucleic acids at the deleted and / or replaced target genomic locus are between about 1 nucleotide to about 5 kb or between about 1 nucleotide to about 1,000 nucleotides in length. For example, the nucleic acid insert or corresponding nucleic acid deleted and / or replaced at the target genomic locus is between about 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150, 150-160, 160-170, 170-180, 180-190, or 190-120 nucleotides in length. Likewise, the nucleic acid insert or corresponding nucleic acid deleted and / or replaced at the target genomic locus is between 1-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 nucleotides in length. Likewise, the nucleic acid insert or corresponding nucleic acid deleted and / or replaced at the target genomic locus is between about 1 kb-1.5 kb, 1.5 kb-2 kb, 2 kb-2.5 kb, 2.5 kb-3 kb, 3 kb-3.5 kb, 3.5 kb-4 kb, 4 kb-4.5 kb, or 4.5 kb-5 kb or longer in length.
[0219] The nucleic acid insert can include sequences that are homologous or orthologous to all or a portion of the sequence intended for replacement. For example, the nucleic acid insert can include sequences that include one or more point mutations (e.g., 1, 2, 3, 4, 5, or more) compared to the sequence intended for replacement at the target genomic locus. Optionally, such point mutations can result in conservative amino acid substitutions in the encoded polypeptide (e.g., substitution of aspartic acid [Asp, D] with glutamic acid [Glu, E]).
[0220] In some cases, the exogenous donor nucleic acid can be a "large targeting vector" or "LTVEC" comprising targeting vectors comprising homology arms that correspond to and are derived from nucleic acid sequences that are larger than those typically used by other methods aimed at performing homologous recombination in a cell. LTVECs also include targeting vectors comprising nucleic acid inserts having nucleic acid sequences that are larger than those typically used by other methods aimed at performing homologous recombination in a cell. For example, LTVECs can make modifications to large loci that cannot be accommodated by traditional plasmid-based targeting vectors due to size limitations. For example, a targeted locus can be one that cannot be targeted or can only be mis-targeted or only targeted with significantly low efficiency using conventional methods in the absence of a nick or double-strand break induced by a nuclease agent (e.g., a Cas protein) (i.e., the 5' and 3' homology arms can correspond to). LTVECs can be of any length, and are typically at least 10 kb in length. The sum of the 5' homology arm and the 3' homology arm in a LTVEC is typically at least 10 kb. In one example, the length of a LTVEC can be between about 50 kb and about 300 kb, or the length of the sum of the 5' homology arm and the 3' homology arm can be about 10 kb to about 200 kb.
[0221] E. Vectors A nucleic acid disclosed herein (e.g., DNA encoding a CtIP fusion protein, DNA encoding an i53 protein, DNA encoding a Cas protein, DNA encoding a guide RNA, an exogenous donor nucleic acid, or a combination thereof, such as an exogenous donor nucleic acid and DNA encoding a guide RNA) can be provided in a vector. The vector can include additional sequences, such as, for example, an origin of replication, a promoter, and a gene encoding antibiotic resistance.
[0222] Some vectors can be circular. Alternatively, the vector can be linear. The vector can be packaged for delivery via a lipid nanoparticle, a liposome, a non-lipid nanoparticle, or a viral capsid. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0223] The vector can be, for example, a viral vector, such as an adeno-associated virus (AAV) vector. The AAV can be of any suitable serotype and can be single-stranded AAV (ssAAV) or self-complementary AAV (scAAV). Other exemplary viruses / viral vectors include retroviruses, lentiviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. The virus can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. The virus can or can not integrate into the host genome. Such viruses can also be engineered to have reduced immunogenicity. The virus can be replication-competent or replication-defective (e.g., have a defect in one or more genes necessary for additional rounds of virion replication and / or packaging). The virus can cause transient expression or more durable expression. The viral vector can be genetically modified from its wild-type counterpart. For example, the viral vector can comprise an insertion, deletion, or substitution of one or more nucleotides to facilitate cloning or such that one or more properties of the vector are altered. Such properties can include packaging capacity, transduction efficiency, immunogenicity, genome integration, replication, transcription, and translation. In some examples, a portion of the viral genome can be deleted such that the virus is able to package exogenous sequences of greater size. In some examples, the viral vector can have enhanced transduction efficiency. In some examples, the immune response induced by the virus in the host can be reduced. In some examples, a viral gene that promotes integration of viral sequences into the host genome, such as integrase, can be mutated such that the virus becomes non-integrating. In some examples, the viral vector can be replication-defective. In some examples, the viral vector can comprise exogenous transcriptional or translational control sequences to drive expression of the coding sequence on the vector. In some examples, the virus can be helper-dependent. For example, the virus can require one or more helper components to provide viral components, such as viral proteins, necessary for amplification of the vector and packaging of the vector into viral particles. In this case, one or more helper components, including one or more vectors encoding viral components, can be introduced into the host cell or population of host cells along with the vector system described herein. In other examples, the virus can be helper-component independent. For example, the virus is able to amplify and package the vector without a helper virus. In some examples, the vector system described herein can also encode viral components necessary for viral amplification and packaging.
[0224] Exemplary viral titers (e.g., AAV titers) include about 10 12 vg / mL to about 10 16 vg / mL. Other exemplary viral titers (e.g., AAV titers) include about 10 12 vg / kg body weight to about 10 16 vg / kg body weight.
[0225] Adeno-associated virus (AAV) is endemic in multiple species, including humans and non-human primates (NHPs). To date, at least 12 natural serotypes and hundreds of natural variants have been isolated and characterized. See, e.g., Li et al., (2020) Nat. Rev. Genet. 21 :255-272, which is incorporated by reference herein in its entirety for all purposes. AAV particles are naturally composed of an unenveloped icosahedral protein capsid containing a single-stranded DNA (ssDNA) genome. The DNA genome is flanked by two inverted terminal repeat sequences (ITRs), which serve as origins of viral replication and packaging signals. rep The gene encodes four proteins required for viral replication and packaging, while cap The gene encodes three structural capsid subunits that determine the AAV serotype and, in some serotypes, an assembly-activating protein (AAP) that facilitates virion assembly.
[0226] Recombinant AAV (rAAV) is one of the most commonly used viral vectors in gene therapy today, treating human diseases by delivering a therapeutic transgene to target cells in vivo. rAAV vectors are composed of an icosahedral capsid similar to native AAV, but rAAV virions do not encapsidate AAV protein-coding sequences or AAV replication sequences. These viral vectors are non-replicative. The only viral sequences required in rAAV vectors are the two ITRs, which are required to direct genome replication and packaging during rAAV vector manufacture. rAAV genomes lack AAV rep and cap genes, making them non-replicative in vivo. rAAV vectors are produced by expressing rep and cap genes in trans, in conjunction with a desired transgene cassette flanked by AAV ITRs.
[0227] In rAAV genomes, a gene expression cassette can be placed between the ITR sequences. Typically, rAAV genomic cassettes contain a promoter driving expression of a transgene, followed by a polyadenylation sequence. The ITRs flanking the rAAV expression cassette are typically derived from AAV2, the first serotype to be isolated and transformed into a recombinant viral vector. From that time, most rAAV manufacturing methods rely on AAV2 Rep based packaging systems. See, e.g., Colella et al., (2017) Mol. Ther. Methods Clin. Dev. 8:87-104, which is incorporated by reference herein in its entirety for all purposes.
[0228] The specific serotype of the recombinant AAV vector influences its in vivo tropism for specific tissues. AAV capsid proteins are responsible for mediating attachment and entry into target cells, followed by endosomal escape and transport to the nucleus. Thus, the choice of serotype will influence the cell types and tissues that the vector is most likely to bind and transduce when injected in vivo when developing rAAV vectors. Several rAAV serotypes, including rAAV8, are able to transduce the liver when delivered systemically in mice, NHPs, and humans. See, e.g., Li et al. (2020) Nat. Rev. Genet. 21 :255-272, which is incorporated by reference herein in its entirety for all purposes.
[0229] Once in the nucleus, the ssDNA genome is released from the virion and a complementary DNA strand is synthesized to generate a double-stranded DNA (dsDNA) molecule. The double-stranded AAV genome naturally circularizes via its ITRs and becomes an episome that will persist extrachromosomally in the nucleus. Thus, for episomal gene therapy procedures, rAAV-delivered rAAV episomes provide long-term, promoter-driven gene expression in non-dividing cells. However, this rAAV-delivered episomal DNA is diluted with cell division. In contrast, the gene therapy described herein is based on gene insertion to allow long-term gene expression.
[0230] The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow synthesis of a complementary DNA strand. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap can be provided in trans. In addition to Rep and Cap, AAV can require an adenovirus gene-containing helper plasmid. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing adenovirus genes El+ to produce infectious AAV particles. Alternatively, Rep, Cap, and adenovirus helper genes can be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.
[0231] Various serotypes of AAV have been identified. These serotypes differ in the cell types they infect (i.e., their tropism), allowing preferential transduction of specific cell types. The term AAV includes, for example, AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV. The genomic sequences of various serotypes of AAV, as well as the sequences of the natural terminal repeats (TRs), Rep proteins, and capsid subunits are known in the art. Such sequences can be found in the literature or in public databases such as GenBank. An “AAV vector” as used herein refers to an AAV vector comprising a heterologous sequence of non-AAV origin (i.e., a nucleic acid sequence heterologous to AAV), typically comprising a sequence encoding a foreign polypeptide of interest. The construct can comprise AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV capsid sequences. Generally, the heterologous nucleotide sequence (transgene) is flanked by at least one, typically two, AAV inverted terminal repeat sequences (ITRs). The AAV vector can be single-stranded (ssAAV) or self-complementary (scAAV). Examples of serotypes for liver tissue include AAV3B, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh.74, AAV-DJ, and AAVhu.37, and specifically AAV8. In one particular example, the AAV vector can be a recombinant AAV8 (rAAV8). A rAAV8 vector as described herein is a vector in which the capsid is from AAV8. For example, an AAV vector using ITRs from AAV2 and a capsid from AAV8 is considered a rAAV8 vector herein. In another particular example, the AAV vector can be a recombinant AAV2 (rAAV2).
[0232] Tropism can be further refined by pseudotypes, i.e., hybrid capsids and genomes from different viral serotypes. For example, AAV2 / 5 indicates a virus containing a serotype 2 genome packaged in a capsid from serotype 5. The use of pseudotyped viruses can increase transduction efficiency as well as alter tropism. Hybrid capsids derived from different serotypes can also be used to alter viral tropism. For example, AAV-DJ contains a hybrid capsid from eight serotypes and exhibits high infectivity in a broad range of cell types in vivo. AAV-DJ8 is another example that exhibits AAV-DJ properties but with enhanced brain uptake. AAV serotypes can also be modified by mutation. Examples of AAV2 mutation modifications include Y444F, Y500F, Y730F, and S662V. Examples of AAV3 mutation modifications include Y705F, Y731F, and T492V. Examples of AAV6 mutation modifications include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SAS TG.
[0233] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAVs rely on cellular DNA replication machinery to synthesize the complementary strand of the AAV single-stranded DNA genome, transgene expression can be delayed. To address this delay, scAAV containing complementary sequences that are able to anneal spontaneously after infection can be used, thereby eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0234] To increase packaging capacity, longer transgenes can be split between two AAV transfer plasmids, the first with a 3' splice donor and the second with a 5' splice acceptor. Upon cell co-infection, these viruses form concatemers, splice together, and the full-length transgene can be expressed. While this allows for longer transgene expression, the efficiency of expression is lower. A similar method for increasing capacity utilizes homologous recombination. For example, a transgene can be split between two transfer plasmids but with substantial sequence overlap, such that co-expression induces homologous recombination and expression of the full-length transgene.
[0235] F. Lipid nanoparticles The different components of the compositions or combinations disclosed herein (e.g., the CtIP fusion protein or DNA or RNA encoding, the i53 protein or DNA or RNA encoding, the Cas protein or DNA or RNA encoding, the guide RNA or DNA encoding, the exogenous donor nucleic acid, or combinations thereof, such as the RNA encoding the Cas protein, the RNA encoding the CtIP fusion protein, the RNA encoding the i53 protein, and optionally the guide RNA) can be provided in a lipid nanoparticle.
[0236] Lipid formulations can protect biomolecules from degradation while improving their cellular uptake. Lipid nanoparticles are particles that include a plurality of lipid molecules that are physically associated with one another through intermolecular forces. These particles comprise microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), dispersed phase in an emulsion, or internal phase in a suspension. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids can be used to deliver polyanions, such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time that a nanoparticle can exist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO 2016 / 010840 Al, which is incorporated by reference herein in its entirety for all purposes. Exemplary lipid nanoparticles can include a cationic lipid and one or more other components. In one example, the other components can include a helper lipid such as cholesterol. In another example, the other components can include a helper lipid (such as cholesterol) and a neutral lipid (such as DSPC). In another example, the other components can include a helper lipid (such as cholesterol), optionally a neutral lipid (such as DSPC), and a stealth lipid (such as S010, S024, S027, S031, or S033).
[0237] In one specific example, RNA encoding a CtIP fusion protein and RNA encoding an i53 protein are each introduced via LNP-mediated delivery into the same LNP. In another specific example, RNA encoding a Cas protein, RNA encoding a CtIP fusion protein, and RNA encoding an i53 protein are each introduced via LNP-mediated delivery into the same LNP. In another specific example, RNA encoding a Cas protein, RNA encoding a CtIP fusion protein, RNA encoding an i53 protein, and optionally a guide RNA are each introduced via LNP-mediated delivery into the same LNP. As discussed in greater detail elsewhere herein, one or more RNAs can be modified. Delivery by such methods results in transient Cas, CtIP fusion protein, or i53 protein expression and / or transient presence of a guide RNA, and biodegradable lipids improve clearance, improve tolerability, and reduce immunogenicity. Lipid formulations can protect biomolecules from degradation while improving their cellular uptake. Lipid nanoparticles are particles that include a plurality of lipid molecules that are physically associated with one another through intermolecular forces. These particles comprise microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), dispersed phase in emulsions, micelles, or internal phase in suspensions. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids can be used to deliver polyanions, such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time that a nanoparticle can exist in vivo. See, e.g., WO 2016 / 010840 Al and WO 2017 / 173054 Al, each of which is incorporated by reference herein in its entirety for all purposes. Exemplary lipid nanoparticles can include a cationic lipid and one or more other components.
[0238] In some LNP, the cargo can include a Cas mRNA (e.g., Cas9 mRNA) and a gRNA. The ratio of Cas mRNA and gRNA can vary. In some LNP, the cargo can include a nucleic acid construct encoding a product of interest (e.g., a polypeptide of interest) and a gRNA. The ratio of a nucleic acid construct encoding a product of interest (e.g., a polypeptide of interest) and a gRNA can vary.
[0239] Examples of suitable LNP can be found, e.g., in WO 2019 / 067992, WO 2020 / 082042, US 2020 / 0270617, WO 2020 / 082041, US 2020 / 0268906, WO 2020 / 082046 (see, e.g., pages 85-86), and US 2020 / 0289628, each of which is incorporated by reference herein in its entirety for all purposes.
[0240] The LNP can comprise one or more or all of: (i) a lipid for encapsulation and endosomal escape; (ii) a neutral lipid for stabilization; (iii) a helper lipid for stabilization; and (iv) a stealth lipid. See, e.g., Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO 2017 / 173054 Al, each of which is incorporated by reference herein in its entirety for all purposes. Specific examples of delivery to the brain using LNP are disclosed in Nabhan et al. (2016) Sci. Rep. 6:20019, which is incorporated by reference herein in its entirety for all purposes.
[0241] G. Cells or animals The cells targeted in the methods disclosed herein can be, for example, mammalian, non-human mammalian, or human. The mammal can be, for example, a non-human mammal, a human, a rodent, a rat, a mouse, or a hamster. Other non-human mammals include, for example, non-human primates, monkeys, apes, cats, dogs, rabbits, horses, oxen, deer, bison, livestock (e.g., bovine species such as cows, oxen, etc.; ovine species such as sheep, goats, etc.; and porcine species such as pigs and wild boars). The term “non-human” does not include humans. In one specific example, the cells are mammalian cells. In another example, the cells are rodent cells. In another example, the cells are mouse cells or rat cells. In another example, the cells are mouse cells. In another example, the cells are rat cells. In another example, the cells are human cells. In one example, the cells are non-cycling cells (i.e., non-dividing). In another example, the cells are cycling (i.e., dividing) cells.
[0242] The cells can be isolated cells (e.g., in vitro) or can be in vivo in a subject (e.g., an animal or a mammal). The cells can also be in any type of undifferentiated or differentiated state. In one example, the cells are liver cells. The cells provided herein can be normal, healthy cells, or can be diseased cells.
[0243] In some embodiments, the cells can be induced pluripotent stem cells (iPSCs), such as human iPSCs. In some embodiments, the cells can be hematopoietic stem cells (HSCs), such as human HSCs. In some embodiments, the cells can be embryonic stem cells (ES cells), such as mouse ES cells or rat ES cells. In some embodiments, the cells can be single cell stage embryos, such as mouse single cell stage embryos or rat single cell stage embryos.
[0244] In some cases, cells comprising a targeted genetic modification made by the methods disclosed herein can be used to make genetically modified organisms comprising the targeted genetic modification. Any convenient method or protocol for producing a genetically modified organism is suitable for producing such genetically modified non-human animals. See, e.g., Cho et al. (2009) Current Protocols in Cell Biology 42:19.11:19.11.1 - 19.11.22, and Gama Sosa et al. (2010) Brain Struct. Funct. 214(2-3): 91-109, each of which is incorporated by reference herein in its entirety for all purposes.
[0245] For example, a method of producing a non-human animal comprising a targeted genetic modification at a target genomic locus can comprise: (1) modifying the genome of a pluripotent cell to comprise a targeted genetic modification at a target genomic locus; (2) identifying or selecting a genetically modified pluripotent cell comprising a targeted genetic modification at a target genomic locus; (3) introducing the genetically modified pluripotent cell into a non-human animal host embryo; and (4) gestating the host embryo in a surrogate mother. Optionally, the host embryo comprising the modified pluripotent cell (e.g., a non-human ES cell) can be incubated until the blastocyst stage before implantation into and gestation in a surrogate mother to produce an F0 non-human animal. The surrogate mother can then produce an F0 generation non-human animal comprising a targeted genetic modification at a target genomic locus.
[0246] An example of a suitable pluripotent cell is an embryonic stem (ES) cell (e.g., a mouse ES cell or a rat ES cell). The modified pluripotent cell can be produced, for example, using the methods disclosed herein. The donor cell can be introduced into the host embryo at any stage, such as the blastocyst stage or the pre-tomoblast stage (i.e., the 4-cell stage or the 8-cell stage). Offspring capable of transmitting the genetic modification through the germline are produced. See, e.g., U.S. Patent No. 7,294,754, which is incorporated by reference herein in its entirety for all purposes.
[0247] Alternatively, a method of producing a non-human animal as described elsewhere herein can comprise: (1) modifying the genome of a single-cell stage embryo to comprise a targeted genetic modification at a target genomic locus; (2) selecting a genetically modified embryo; and (3) gestating the genetically modified embryo in a surrogate mother. Offspring capable of transmitting the genetic modification through the germline are produced.
[0248] Nuclear transfer techniques can also be used to produce non-human mammals. Briefly, methods for nuclear transfer can include the following steps: (1) enucleating or providing an enucleated oocyte; (2) isolating or providing a donor cell or nucleus for combination with the enucleated oocyte; (3) inserting the cell or nucleus into the enucleated oocyte to form a reconstructed cell; (4) implanting the reconstructed cell into the uterus of an animal to form an embryo; and (5) allowing the embryo to develop. In such methods, the oocyte is typically taken from a dead animal, but oocytes can also be isolated from the oviducts and / or ovaries of live animals. The oocyte can be matured in various well-known culture media prior to enucleation. Enucleation of the oocyte can be performed in a variety of well-known ways. The donor cell or nucleus can be inserted into the enucleated oocyte to form a reconstituted cell by microinjection of the donor cell under the zona pellucida prior to fusion. Fusion can be induced by applying a DC electrical pulse (electrofusion) across the plane of contact / fusion, by exposing the cells to a chemical that promotes fusion such as polyethylene glycol, or by means of an inactivated virus such as Sendai virus. The reconstituted cell can be activated by electrical and / or non-electrical means before, during, and / or after fusion of the nuclear donor and recipient oocyte. Activation methods include electrical pulses, chemically induced shock, sperm penetration, increasing the level of divalent cations in the oocyte, and decreasing the phosphorylation of cellular proteins in the oocyte such as by means of a kinase inhibitor. The activated reconstituted cell or embryo can be cultured in well-known culture media and then transferred to the uterus of an animal. See, e.g., US 2008 / 0092249, WO 1999 / 005266, US 2004 / 0177390, WO 2008 / 017234, and U.S. Patent No. 7,612,250, each of which is incorporated by reference herein in its entirety for all purposes.
[0249] The various methods provided herein allow for the production of genetically modified non-human F0 animals, wherein the cells of the genetically modified F0 animals comprise a targeted genetic modification at a target genomic locus. It will be appreciated that the number of cells in the F0 animal that have a targeted genetic modification at a target genomic locus will vary depending on the method used to generate the F0 animal. By way of example, VELOCIMOUSE® ®Methods introducing donor ES cells into pre-sphere stage embryos from a corresponding organism (e.g., an 8-cell stage mouse embryo) allows for a greater percentage of the F0 animal's cell population to include cells having a nucleotide sequence of interest including a targeted genetic modification. For example, at least 50%, 60%, 65%, 70%, 75%, 85%, 86%, 87%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cells of a non-human F0 animal's contribution can include a cell population having a targeted modification.
[0250] The cells of the genetically modified F0 animal can be heterozygous for the targeted genetic modification at the target genomic locus, or can be homozygous for the targeted genetic modification at the target genomic locus.
[0251] All patent applications, websites, other publications, accession numbers, and the like cited above or below are hereby incorporated by reference in their entireties for all purposes as if each individual item was specifically and individually indicated to be incorporated by reference. If different versions of a sequence are associated with different accession numbers at different times, the version associated with the accession number at the effective filing date of this application is intended. The effective filing date means the earlier of the actual filing date or the filing date of the priority application in which the accession number is mentioned, if applicable. Likewise, if different versions of a publication, website, and the like are published at different times, the version most recently published at the effective filing date of this application is intended, unless otherwise specified. Unless specifically stated otherwise, any feature, step, element, embodiment or aspect of the application can be used in combination with any other feature, step, element, embodiment or aspect. Although the application has been described in detail for the purpose of clarity and understanding, it will be apparent that certain modifications can be practised within the scope of the appended claims.
[0252] Sequence brief description Nucleotide and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotides bases and three letter code for amino acids. Nucleotide sequences follow the standard convention of starting at the 5' end of the sequence and continuing forward (i.e., from left to right in each line) to the 3' end. While only one strand of each nucleotide sequence is shown, it is understood that any reference to the shown strand includes the complementary strand. When a nucleotide sequence encoding an amino acid sequence is provided, it is understood that degenerate codon variants, which also encode the same amino acid sequence, are also provided. Amino acid sequences follow the standard convention of starting at the amino terminus of the sequence and continuing forward (i.e., from left to right in each line) to the carboxy terminus.
[0253] Table 2. Description of sequences
[0254] Examples Example 1. Combined expression of DNA repair modulators enhances CRISPR-mediated homologous recombination To stimulate homology directed repair (HDR) without suppressing other DSB repair mechanisms or endogenous functions of cell cycle regulation, we generated a HDR enhancer mix containing mRNAs encoding inhibitors of CTbP interacting protein (CtIP) and TP53 binding protein 1 (53BP1), i53. The former is involved in initiating DNA excision at DSB sites, a limiting step for HDR, while the latter suppresses 53BP1 recruitment to and action at damaged chromatin, in turn creating a favorable environment for HDR. We designed a dual delivery system that can efficiently deliver the combination of the two proteins together and transiently with Cas9 mRNA, sgRNA, and donor template. In one format, the mRNAs are all packaged in a lipid nanoparticle (LNP), while the donor template and sgRNA are delivered by an adeno-associated virus (AAV) vector. As shown below, our data indicate that the combination of the two enhancers (MS2-CtIP and i53) is a significant improvement of CRISPR-mediated HDR in human cells by about 20-fold, the rapid, transient, and integration-free expression of these proteins, allowing a better and more efficient way to perform precise genome editing.
[0255] via in LMNA precise knock-in at a locus mCLOVER , based on high-throughput microscopy assessment CRISPR mediated HDR efficiency To measure CRISPR-induced HDR efficiency, we used a microscope-based assay to quantify the percentage of cells that have undergone precise gene integration. The CRISPR reagents were designed to integrate the mCLOVER coding sequence into the 5' end of the gene encoding the lamin A / C protein. See FIG. 1. We targeted this gene because lamin A / C is prominently located in the nuclear envelope, allowing for the identification of cells that have undergone HDR (i.e., in-frame expression of the mCLOVER-lamin fusion). We designed a donor template (SEQ ID NO: 44) containing the mCLOVER coding sequence flanked by -600 bp homology arms corresponding to the region flanking the start codon. To prevent re-cutting of Cas9 after integration, the repair template contains a silent mutation at the gRNA PAM sequence. The repair template flanks the gRNA target site, allowing the template to be linearized when cut by Cas9 in the cell. LMNA Figure 1 To measure CRISPR-induced HDR efficiency, we used a microscope-based assay to quantify the percentage of cells that have undergone precise gene integration. The CRISPR reagents were designed to integrate the mCLOVER coding sequence into the 5' end of the gene encoding the lamin A / C protein. See FIG. 1. We targeted this gene because lamin A / C is prominently located in the nuclear envelope, allowing for the identification of cells that have undergone HDR (i.e., in-frame expression of the mCLOVER-lamin fusion). We designed a donor template (SEQ ID NO: 44) containing the mCLOVER coding sequence flanked by -600 bp homology arms corresponding to the region flanking the start codon. To prevent re-cutting of Cas9 after integration, the repair template contains a silent mutation at the gRNA PAM sequence. The repair template flanks the gRNA target site, allowing the template to be linearized when cut by Cas9 in the cell. LMNA
[0256] To implement LMNA-HDR, we transduced Cas9-expressing HEK293 cells with AAV2 packaged with a vector containing a donor template along with a U6-sgRNA2.0 expression cassette (AAV-sgRNA2.0-mClover). The Cas9 protein and DNA sequence are shown in SEQ ID NOs: 1 and 2, respectively, and LMNA The sgRNA sequence is shown in SEQ ID NO: 27. To optimize AAV delivery and establish a baseline HDR efficiency in these cells, we transduced at different MOIs and imaged the cells by confocal microscopy after 96 hours. We observed an increase in HDR events with increasing sgRNA and donor template availability, validating the sensitivity of our assay. See Figures 2A-2B AAV transduction at increasing MOIs into Cas9-expressing HEK293 cells (Cas9-HEK) resulted in a corresponding increase in the percentage of mClover-LMNA positive cells ( Figure 2B ). When cells were co-treated with mirin, a well-established MRE11, homologous recombination protein, inhibitor, or when a donor template without HA was used, expression of mClover-LMNA was reduced ( Figure 2C ). These data suggest that the presence of the mClover-LMNA fusion protein does represent successful HDR.
[0257] Screening for homologous recombination proteins to improve Cas9 efficiency of induction HDR of To selectively facilitate HDR repair at CRISPR-mediated DSBs, we sought to identify and recruit potential HDR enhancers to the Cas9 cleavage site using the MS2 tag method. In this method, the scaffold sequence of the sgRNA is modified to include MS2 bacteriophage aptamers (sgRNA2.0), and HDR increasing proteins are fused to MS2 coat proteins (MS2) to interact with sgRNA2.0 ( Figure 3A ). We hypothesized that the accumulation of these factors specifically at the CRISPR / Cas9 targeted locus would increase the frequency of precise edits without altering the global DSB repair landscape. In this study, all sgRNAs used contained the MS2 scaffold unless otherwise noted.
[0258] We selected nine homologous recombination proteins and assessed their effect on HDR frequency when expressed and recruited to the Cas9 targeted LMNA locus. To do this, we generated individual plasmids expressing these proteins fused to MS2 at the 5' end. We then co-delivered the plasmids and AAV sgLMNA+mClover containing sgRNAs to Cas9-HEK cells and assayed their corresponding HDR efficiency after 96 hours. While most candidates resulted in increased HDR, when co-delivered with AAVsgLMNA+mClover MS2-CtIP expression resulted in the most significant HDR improvement when compared to plasmid-induced basal HDR activity of MS2 alone Figure 3B ).
[0259] Single endonuclease enhancer LNP Delivery improved Cas9 Induced HDR Efficiency CtIP protein plays a key role in DNA repair, particularly in DNA end resection, which is a rate-limiting step in homologous recombination. While plasmid delivery is convenient and simple, it is limited to in vitro applications due to immune response and potential risk of non-specific DNA recombination with the genome. Moreover, while CtIP is generally considered a tumor suppressor, studies indicate its potential oncogenic role in promoting tumorigenesis. In particular, CtIP overexpression was found in gastric cancer; its amplification was also documented in several other cancers. See, e.g., Mozaffari et al. (2021) Semin. Cell. Dev. Biol. 113:47-56, which is incorporated by reference herein in its entirety for all purposes. Thus, we sought to explore a therapeutically relevant approach for delivering exogenous MS2-CtIP to enhance precise editing while minimizing its unwanted prolonged activity. Lipid nanoparticles (LNPs) have become an effective, clinically feasible delivery method for CRISPR editing systems (e.g., for delivering Cas9 mRNA and gRNA) due to their ability to efficiently encapsulate these components and deliver them into cells. This approach enables rapid and transient expression of Cas9 in cells, allowing for efficient editing while mitigating issues associated with long-term, off-target nuclease exposure. Based on this, we proposed that delivery of MS2-CtIP mRNA via LNPs could yield effective HDR boosting activity without any lasting negative effects.
[0260] To test whether CtIP localization can boost HDR, we tested whether LNP-mediated delivery of MS2-CtIP mRNA (LNP-MS2-CtIP) (amino acid and nucleotide sequences shown in SEQ ID NOs: 30 and 31, respectively) can affect HDR efficiency in Cas9-HEK293. We delivered LNP-MS2-CtIP and AAV-sgRNA2.0-mClover simultaneously to Cas9-HEK293 cells ( Figure 4A ). To maximize resolution of the effect of LNP-MS2-CtIP, we transduced cells with AAV at the lowest MOI we tested (1 x 105vg / cell) and treated cells in parallel with LNPs packaged with different amounts of mRNA. Our results indicated that expression of MS2-CtIP resulted in 4 , and cells were treated in parallel with LNPs packaged with different amounts of mRNA. Our results indicated that expression of MS2-CtIP resulted in LMNA up to about 13-fold increase in HDR at the locus (Figure 4B The co-delivery of MS2-CtIP with AAV containing a donor template and conventional gRNA but without aptamers (AAV-sgRNA-mClover) does not result in any change in HDR efficiency. Figure 4C This indicates the need to localize CtIP to the DSB to enable cells to repair breaks via HDR. Next, we analyzed the same samples using NGS amplicon sequencing. LMNA Insertion and deletion rates at loci were used to quantify the frequency of insertion or deletion mutations (insertions and deletions) around Cas9 cleavage sites, and MS2-CtIP treatment was found to reduce the efficiency of insertion and deletion formation, thus indicating that MS2-CtIP-induced end excision shifts the DSB repair balance from NHEJ to HDR. Figure 4B Our results indicate that enhanced availability of CtIP at CRISPR sites can significantly improve the efficiency of precise gene knock-in. See also Figures 4A-4C In summary, our data demonstrate that targeting the CtIP protein to the Cas9 cleavage site effectively shifts the repair pathway selection from NHEJ to HDR, and that LNP transfection is a suitable delivery method for transient MS2-CtIP expression.
[0261] We then compared LNP delivery of exogenous MS2-CtIP mRNA with plasmid delivery of MST-CtIP DNA and found that LNP delivery of mRNA worked better than plasmid expression. The transient nature of MS2-CtIP expression via LNP-mediated delivery raises concerns about its potential to attenuate its HDR-enhancing efficacy compared to stable expression of MS2-CtIP (such as from plasmids). Contrary to our expectations, when we compared the effect of MS2-CtIP on HDR efficiency via LNP delivery with plasmid expression, we observed a significantly stronger effect of LNP delivery. Specifically, the HDR efficiency using LNP increased by 12.5-fold compared to the baseline HDR efficiency when MS2-CtIP was not delivered, compared to a 2.5-fold increase using plasmid expression. Figure 5 ).
[0262] Given that LNP-mediated enhancement of end-resection can facilitate CRISPR-HDR, we next evaluated whether expression of an end-protection inhibitor (i.e., a 53BP1 inhibitory peptide i53) can also facilitate HDR upon CRISPR editing. 53BP1 plays a key role in DSB repair pathway choice by preventing early steps of DNA resection, in part by physically blocking resection nuclease activity and by mediating CtIP dephosphorylation. We packaged different amounts of i53 mRNA (amino acid and nucleotide sequences set forth in SEQ ID NOs: 42 and 43, respectively) in LNPs and delivered them to Cas9-HEK293 cells treated with AAV Figure 6A ). Increasing i53 expression resulted in a gradual and significant increase in HDR Figure 6B ). Accordingly, we observed LMNA indel rates at the locus, indicating that LNP-mediated i53 expression can facilitate HDR upon CRISPR editing Figure 6B ). Our data indicate that 53BP1 inhibition, which indirectly allows BRCA1 function and DNA end-resection, is a viable approach to enhance CRISPR-mediated precise gene knock-in. See Figures 6A-6B .
[0263] i53 and MS2-CtiP The combination expression further improves HDR rate We reasoned that while recombinant CtIP overexpression and localization to the CRISPR site can directly enhance DNA end-resection for HDR activation, the role of 53BP1 inhibition is to indirectly provide an HDR-permissive environment for CtIP and other HDR proteins to function without significantly down-regulating NHEJ. We therefore tested whether co-delivery of MS2-CtIP and i53 mRNA via LNPs can induce a combined HDR-enhancing effect in CRISPR-treated cells. We treated Cas9-expressing HEK293 cells with AAV-sgRNA2.0-mClover and LNPs encapsulating both MS2-CtIP and i53. Our data show that delivery of 7.5 pg MS2-CtIP and 15 pg i53 via LNPs in CRISPR-treated cells can enhance HDR efficiency in HEK cells by about 20-fold. See Figure 7 . A combined enhancement of HDR efficiency was also observed in Cas9-expressing Huh7 cells, where HDR efficiency was enhanced by about 5-fold relative to baseline (data not shown).
[0264] It is known that during the S / G2 phase of the cell cycle (when HDR is permissive), CtIP is phosphorylated by the cell cycle regulator CDKs, which in turn inhibits 53BP1 function and initiates HDR. See, e.g., Daley et al. (2014) Mol. Cell. Biol.34(8): 1380-1388 ("In S / G2, CtIP is phosphorylated by CDKs, inducing complex formation with BRCA1 and MRN. This complex displaces 53BP1 and initiates resection."), which is incorporated by reference herein in its entirety for all purposes. Given this established mechanism, it was surprising to find that, in our system, the addition of i53 concurrently with exogenous MS2-CtIP overexpression / localization could further increase HDR efficiency. It has been previously reported that silencing 53BP1 or depleting its ability to bind damaged chromatin changes limited DSB resection to hyper-resection and causes a switch from error-free gene conversion by RAD51 to mutagenic single-strand annealing by RAD52. See, e.g., Ochs et al. (2016) Nat. Struct. Mol. Biol. 23(8): 714-721, which is incorporated by reference herein in its entirety for all purposes. Based on this, it would be expected that our potentiator, which employs two easy-resection strategies, would result in the mutagenic editing outcome (hyper-resection) indicated by the study. However, contrary to this expectation, our potentiator successfully facilitated error-free HDR.
[0265] To our knowledge, our strategy produces the highest rate of HDR enhancement compared to other published strategies. Our strategy offers potential benefits in various research applications, such as disease model generation in cells and in vivo, as well as in therapeutic settings where the efficiency of precise gene knock-in can be significantly enhanced. Our strategy also offers benefits over methods where all components are fused to a single Cas9 protein. For example, our system facilitates the additional recruitment of CtIP molecules, which have been described to act as multimeric complexes, whereas a fusion protein would only allow one CtIP molecule as only one Cas9 molecule can bind to the target site. Similarly, i53 can have a more robust effect when delivered independently than as a fusion with limited stoichiometry.
[0266] NHEJ Chemical inhibition Chemical inhibition of NHEJ is often used as a strategy to enhance HDR in cells. In this context, we compared our HDR potentiator to AZD7648, one of the most potent NHEJ (DNA-PKcs) inhibitory small molecules. Our report indicates that the HDR potentiator significantly outperforms AZD7648 in enhancing HDR ( Figure 8 ). Notably, the combination of the HDR potentiator and AZD7648 did not result in further HDR enhancement ( Figure 8 ). We infer that our HDR potentiator maximized CRISPR-mediated HDR at the LMNA locus given the amount of sgRNA / donor template provided and the capacity of endogenous HDR machinery.
[0267] HDR The enhancer reduces HDR Precise gene integration is achieved with a template To simulate therapeutic gene editing, we next investigated the effects of HDR enhancers on precise gene editing when Cas9 was also transiently expressed by LNPs rather than being stably expressed in cell lines. Simultaneously, we tested whether HDR enhancers allowed AAV. sgLMNA+mClover The amount of [specific component] is reduced without affecting the target gene targeting efficiency. Therefore, we prepared two types of LNPs: LNPs encapsulating Cas9, MS2-CtIP, and i53 mRNA. booster And LNPs that encapsulate Cas9 and mCherry mRNA baseline Then, we transfect LNPs into AAVs with different MOIs. sgLMNA+mClover In transduced HEK293 cells, as expected, reducing AAV MOI led to decreased gene targeting efficiency. Significantly, regardless of AAV MOI, we observed a decrease compared to using LNP... baseline Compared to cells treated with LNP booster Consistent increase in HDR efficiency in treated cells ( Figure 9A This is accompanied by a corresponding reduction in the formation of insertional deletions (). Figure 9A ).
[0268] We repeated the targeting of two other genes. HMGA1 and SEC61B The experiment at the C-terminus yielded similar conclusions. Figure 9B In summary, our results demonstrate that HDR enhancers can be used in clinically relevant settings where Cas9 is transiently co-expressed with the enhancer. Importantly, this co-delivery allows for highly precise gene integration efficiency while reducing the amount of CRISPR reagent (AAV) required, thus minimizing off-target CRISPR activity and the chance of transgene integration.
[0269] HDR Strengtheners facilitate precise gene integration using different delivery methods The data above highlight the ability of HDR enhancers to improve precision editing using a co-delivery approach of LNP / AAV. While AAV remains the preferred donor template vector in both in vivo preclinical and clinical settings, non-viral templates such as linear dsDNA and ssDNA are emerging as promising alternatives. This is particularly evident in clinically relevant gene editing of primary human hematopoietic cells and in CRIPSR-mediated disease modeling using iPSCs, mESCs, and animal embryos.
[0270] Therefore, we sought to investigate whether HDR enhancers could enhance HDR in non-viral donor delivery scenarios. To achieve this, we obtained linearly closed-terminal dsDNA containing the mClover coding sequence, with side-joints as described above.LMNA homology arm sequences (dsDNA mclover-LMNA ), and delivered them via electroporation into Cas9-HEK cells along with a plasmid encoding sgLMNA2.0. Electroporated cells were then cultured in media containing or not containing LNP encapsulated HDR enhancer for 96 hours. Our data show that inclusion of HDR enhancer significantly increased the percentage of mClover-LMNA cells ( Figure 10 ).
[0271] We also tested whether HDR enhancer mRNA could be co-delivered via electroporation with the rest of the CRISPR reagents. To do so, we obtained synthetic sgLMNA2.0 and verified that they were as efficient as synthetic sgLMNA-REG. We then electroporated sgLMNA2.0, dsDNA mclover-LMNA and HDR enhancer mRNA into HEK-Cas9, and observed a significant increase in HDR efficiency ( Figure 11 ). In summary, our findings demonstrate that HDR enhancer can be effectively paired with different delivery methods, thus underscoring its broad versatility.
Claims
1. A method of targeted genetic modification by homology-directed repair at a target genomic locus in a cell, the method comprising administering to the cell: (a) a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) protein or a nucleic acid encoding the Cas protein; (b) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises one or more adaptor binding elements to which an adaptor protein is capable of specifically binding, and wherein the guide RNA is capable of forming a complex with and directing the Cas protein to a guide RNA target sequence at the target genomic locus; (c) a fusion protein or a nucleic acid encoding the fusion protein, wherein the fusion protein comprises a CtBP-interacting protein (CtIP) fused to the adaptor protein; (d) an inhibitor of a 53BP1 (i53) protein or a nucleic acid encoding the i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, optionally wherein the 5' homology arm and the 3' homology arm flank an insert nucleic acid, wherein the Cas protein and the guide RNA form a complex, the Cas protein cleaves the guide RNA target sequence to create a double-strand break, and the exogenous donor nucleic acid recombines with the target genomic locus via homology-directed repair to create the targeted genetic modification.
2. The method of claim 1, wherein the Cas protein is administered to the cell in the form of a protein, optionally wherein the Cas protein is in a lipid nanoparticle.
3. The method of claim 1, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, optionally wherein the RNA encoding the Cas protein is in a lipid nanoparticle.
4. The method of claim 1, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises a DNA encoding the Cas protein, optionally wherein the DNA encoding the Cas protein is in a viral vector, optionally wherein the viral vector is a recombinant adeno-associated virus (AAV) vector.
5. The method of any one of claims 1-4, wherein the Cas protein is a Cas9 protein.
6. The method of claim 5, wherein the Cas9 protein is a S. pyogenes Cas9 protein, a C. jejuni Cas9 protein, or a S. aureus Cas9 protein, optionally wherein the Cas9 protein is a S. pyogenes Cas9 protein.
7. The method of any one of claims 1-6, wherein the guide RNA is administered in the form of an RNA, optionally wherein the guide RNA is in a lipid nanoparticle.
8. The method of any one of claims 1 to 6, wherein the one or more DNAs encoding the guide RNA are administered to the cell, optionally wherein the one or more DNAs encoding the guide RNA are in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
9. The method of any one of claims 1 to 8, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding.
10. The method of claim 9, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA.
11. The method of claim 10, wherein the guide RNA is a single guide RNA comprising a CRISPR RNA (crRNA) portion fused to a trans-activating CRISPR RNA (tracrRNA) portion, and wherein the first loop is a four loop corresponding to residues 13 to 16 of SEQ ID NO: 11, 13, 15, or 16 and the second loop is a stem loop 2 corresponding to residues 53 to 56 of SEQ ID NO: 11, 13, 15, or 16.
12. The method of any one of claims 1 to 11, wherein the adaptor binding element comprises the sequence set forth in SEQ ID NO: 19 or 20.
13. The method of any one of claims 1 to 12, wherein the guide RNA comprises the sequence set forth in SEQ ID NO: 21, 22, 23, 24, 25, or 26.
14. The method of any one of claims 1 to 13, wherein the fusion protein is administered to the cell in the form of a protein, optionally wherein the fusion protein is in a lipid nanoparticle.
15. The method of any one of claims 1 to 13, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, optionally wherein the RNA encoding the fusion protein is in a lipid nanoparticle.
16. The method of any one of claims 1 to 13, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises a DNA encoding the fusion protein, optionally wherein the DNA encoding the fusion protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
17. The method of any one of claims 1 to 16, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof.
18. The method of any one of claims 1 to 17, wherein the adaptor protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
32. 19. The method of any one of claims 1-18, wherein the adaptor protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
33.
20. The method of any one of claims 1-19, wherein the CtIP protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
34.
21. The method of any one of claims 1-20, wherein the CtIP protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
36.
22. The method of any one of claims 1-21, wherein the fusion protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
30.
23. The method of any one of claims 1-22, wherein the fusion protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
31.
24. The method of any one of claims 1-23, wherein the i53 protein is administered to the cell in the form of a protein, optionally wherein the i53 protein is in a lipid nanoparticle.
25. The method of any one of claims 1-24, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA messenger RNA encoding the i53 protein, optionally wherein the RNA encoding the i53 protein is in a lipid nanoparticle.
26. The method of any one of claims 1-25, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises a DNA encoding the i53 protein, optionally wherein the DNA encoding the i53 protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
27. The method of any one of claims 1-26, wherein the i53 protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 40 or 42.
28. The method of any one of claims 1-27, wherein the i53 protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 41 or 43.
29. The method of any one of claims 1-28, wherein the exogenous donor nucleic acid comprises the insert nucleic acid.
30. The method of any one of claims 1-29, wherein the exogenous donor nucleic acid is in a viral vector.
31. The method of claim 30, wherein the viral vector is a recombinant AAV vector.
32. The method of any one of claims 1-31, wherein the exogenous donor nucleic acid is a large targeting vector (LTVEC), wherein: (a) the LTVEC is at least 10 kb; (b) the sum of the 5’ and 3’ homology arms of the LTVEC is at least 10 kb; (c) the LTVEC is about 50 kb to about 300 kb; or (d) the sum of the 5’ and 3’ homology arms of the LTVEC is about 10 kb to about 200 kb.
33. The method of any one of claims 1-32, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
34. The method of any one of claims 1-32, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
35. The method of any one of claims 1-32, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle. wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises RNA encoding the i53 protein, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, the RNA encoding the i53 protein, and the guide RNA are in a lipid nanoparticle, wherein the one or more DNA encoding the guide RNA is administered to the cell, and wherein the one or more DNA encoding the guide RNA and the exogenous donor nucleic acid are in a recombinant AAV vector.
36. The method of any one of claims 1-32, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the nucleic acid encoding the Cas protein is administered to the cell, wherein the nucleic acid encoding the Cas protein comprises RNA encoding the Cas protein, wherein the nucleic acid encoding the fusion protein is administered to the cell, wherein the nucleic acid encoding the fusion protein comprises RNA encoding the fusion protein, wherein the nucleic acid encoding the i53 protein is administered to the cell, wherein the nucleic acid encoding the i53 protein comprises RNA encoding the i53 protein, wherein the guide RNA is administered to the cell in the form of RNA, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, the RNA encoding the i53 protein, and the guide RNA are in a lipid nanoparticle, wherein the exogenous donor nucleic acid is in a recombinant AAV vector.
37. The method of any one of claims 1-36, wherein the cell is a mammalian cell.
38. The method of any one of claims 1-37, wherein the cell is a rodent cell.
39. The method of any one of claims 1-38, wherein the cell is a mouse cell or a rat cell.
40. The method of any one of claims 1-38, wherein the cell is a mouse cell.
41. The method of any one of claims 1-37, wherein the cell is a human cell.
42. The method of any one of claims 1-41, wherein the cell is in vitro.
43. The method of any one of claims 1-41, wherein the cell is in vivo.
44. A composition or combination comprising: (a) a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) protein or a nucleic acid encoding the Cas protein; (b) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises one or more adaptor binding elements to which the adaptor protein is capable of specifically binding, and wherein the guide RNA is capable of forming a complex with the Cas protein and directing it to a guide RNA target sequence at the target genomic locus; (c) a fusion protein or a nucleic acid encoding the fusion protein, wherein the fusion protein comprises a CtBP-interacting protein (CtIP) fused to the adaptor protein; (d) an inhibitor of a 53BP1 (i53) protein or a nucleic acid encoding the i53 protein; and (e) an exogenous donor nucleic acid comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, optionally wherein the 5' homology arm and the 3' homology arm flank an insert nucleic acid.
45. The composition or combination of claim 44, wherein the composition or combination comprises the Cas protein in the form of a protein, optionally wherein the Cas protein is in a lipid nanoparticle.
46. The composition or combination of claim 44, wherein the composition or combination comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, optionally wherein the RNA encoding the Cas protein is in a lipid nanoparticle.
47. The composition or combination of claim 44, wherein the composition or combination comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises a DNA encoding the Cas protein, optionally wherein the DNA encoding the Cas protein is in a viral vector, optionally wherein the viral vector is a recombinant adeno-associated viral (AAV) vector.
48. The composition or combination of any one of claims 44-47, wherein the Cas protein is a Cas9 protein.
49. The composition or combination of claim 48, wherein the Cas9 protein is a S. pyogenes Cas9 protein, a C. jejuni Cas9 protein, or a S. aureus Cas9 protein, optionally wherein the Cas9 protein is a S. pyogenes Cas9 protein.
50. The composition or combination of any one of claims 44-49, wherein the composition or combination comprises the guide RNA in the form of an RNA, optionally wherein the guide RNA is in a lipid nanoparticle.
51. The composition or combination of any one of claims 44-49, wherein the composition or combination comprises the one or more DNAs encoding the guide RNA, optionally wherein the one or more DNAs encoding the guide RNA is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
52. The composition or combination of any one of claims 44-51, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding.
53. The composition or combination of claim 52, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA.
54. The composition or combination of claim 53, wherein the guide RNA is a single guide RNA comprising a CRISPR RNA (crRNA) portion partially fused to a trans-activating CRISPR RNA (tracrRNA) portion, and wherein the first loop is a four-loop corresponding to residues 13 to 16 of SEQ ID NO: 11, 13, 15, or 16 and the second loop is a stem loop 2 corresponding to residues 53 to 56 of SEQ ID NO: 11, 13, 15, or 16.
55. The composition or combination of any one of claims 44-54, wherein the adaptor binding element comprises the sequence set forth in SEQ ID NO: 19 or 20.
56. The composition or combination of any one of claims 44-55, wherein the guide RNA comprises the sequence set forth in SEQ ID NO: 21, 22, 23, 24, 25, or 26.
57. The composition or combination of any one of claims 44-56, wherein the composition or combination comprises the fusion protein in the form of a protein, optionally wherein the fusion protein is in a lipid nanoparticle.
58. The composition or combination of any one of claims 44-56, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, optionally wherein the RNA encoding the fusion protein is in a lipid nanoparticle.
59. The composition or combination of any one of claims 44-56, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises a DNA encoding the fusion protein, optionally wherein the DNA encoding the fusion protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
60. The composition or combination of any one of claims 44-59, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof.
61. The composition or combination of any one of claims 44-60, wherein the adaptor protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
32.
62. The composition or combination of any one of claims 44-61, wherein the adaptor protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
33.
63. The composition or combination of any one of claims 44-62, wherein the CtIP protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
34.
64. The composition or combination of any one of claims 44-63, wherein the CtIP protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
36.
65. The composition or combination of any one of claims 44-64, wherein the fusion protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
30.
66. The composition or combination of any one of claims 44-65, wherein the fusion protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO:
31.
67. The composition or combination of any one of claims 44-66, wherein the composition or combination comprises the i53 protein in the form of a protein, optionally wherein the i53 protein is in a lipid nanoparticle.
68. The composition or combination of any one of claims 44-67, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises a RNA messenger RNA encoding the i53 protein, optionally wherein the RNA encoding the i53 protein is in a lipid nanoparticle.
69. The composition or combination of any one of claims 44-68, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises a DNA encoding the i53 protein, optionally wherein the DNA encoding the i53 protein is in a viral vector, optionally wherein the viral vector is a recombinant AAV vector.
70. The composition or combination of any one of claims 44-69, wherein the i53 protein comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 40 or 42.
71. The composition or combination of any one of claims 44-70, wherein the i53 protein is encoded by a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 41 or 43.
72. The composition or combination of any one of claims 44-71, wherein the exogenous donor nucleic acid comprises the intervening nucleic acid.
73. The composition or combination of any one of claims 44-72, wherein the exogenous donor nucleic acid is in a viral vector.
74. The composition or combination of claim 73, wherein the viral vector is a recombinant AAV vector.
75. The composition or combination of any one of claims 44-74, wherein the exogenous donor nucleic acid is a large targeting vector (LTVEC), wherein: (a) the LTVEC is at least 10 kb; (b) the sum of the 5’ and 3’ homology arms of the LTVEC is at least 10 kb; (c) the LTVEC is about 50 kb to about 300 kb; or (d) the sum of the 5’ and 3’ homology arms of the LTVEC is about 10 kb to about 200 kb.
76. The composition or combination of any one of claims 44-75, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
77. The composition or combination of any one of claims 44-75, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises an RNA encoding the fusion protein, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises an RNA encoding the i53 protein, and wherein the RNA encoding the fusion protein and the RNA encoding the i53 protein are in a lipid nanoparticle.
78. The composition or combination of any one of claims 44-75, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the composition or combination comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises an RNA encoding the Cas protein, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises RNA encoding the fusion protein, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises RNA encoding the i53 protein, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, and the RNA encoding the i53 protein are in a lipid nanoparticle, wherein the composition or combination comprises the one or more DNA encoding the guide RNA, and wherein the one or more DNA encoding the guide RNA and the exogenous donor nucleic acid are in a recombinant AAV vector.
79. The composition or combination of any one of claims 44-75, wherein the guide RNA comprises two adaptor binding elements to which the adaptor protein is capable of specifically binding, wherein a first adaptor binding element is within a first loop of the guide RNA and a second adaptor binding element is within a second loop of the guide RNA, wherein the adaptor protein comprises an MS2 coat protein or a functional fragment or variant thereof, wherein the composition or combination comprises the nucleic acid encoding the Cas protein, wherein the nucleic acid encoding the Cas protein comprises RNA encoding the Cas protein, wherein the composition or combination comprises the nucleic acid encoding the fusion protein, wherein the nucleic acid encoding the fusion protein comprises RNA encoding the fusion protein, wherein the composition or combination comprises the nucleic acid encoding the i53 protein, wherein the nucleic acid encoding the i53 protein comprises RNA encoding the i53 protein, wherein the composition or combination comprises the guide RNA in the form of RNA, wherein the RNA encoding the Cas protein, the RNA encoding the fusion protein, the RNA encoding the i53 protein, and the guide RNA are in a lipid nanoparticle, wherein the exogenous donor nucleic acid is in a recombinant AAV vector.
Citation Information
Patent Citations
Load-sustainer.
US1202530A
Method of nuclear transfer
US20040177390A1
Methods of modifying eukaryotic cells
US20050144655A1
Method of nuclear transfer
US20080092249A1
RNA Modification to Engineer Cas9 Activity
US20150376586A1