Compositions, systems and methods for eukaryotic gene editing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2026-03-16
AI Technical Summary
CRISPR-based genome editing often unpredictably produces large DNA deletions, posing safety risks and reducing the predictability of somatic cell genome editing.
A mammalian expression plasmid with a eukaryotic promoter linked to a non-viral nucleic acid sequence encoding a fusion protein comprising a DNA polymerase domain and a CRISPR-associated endonuclease, along with a guide RNA (gRNA) coding sequence, is used to reduce large deletions by enhancing non-homologous end joining (NHEJ) and templated insertions.
The solution effectively reduces large DNA deletions and increases the ratio of targeted 1 bp deletions and templated insertions, improving the safety and predictability of genome editing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Prior Related Applications This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 317,720, filed March 8, 2022, which is incorporated by reference herein in its entirety.
[0002] Sequence Listing This application contains a sequence listing in XML format. The sequence listing, named 095199-1375708.xml, was created on Mar. 7, 2023, is 175 kilobytes in size, and is incorporated by reference in its entirety. Field This disclosure describes compositions and methods of using same for eukaryotic gene editing. [Background technology]
[0003] background
[0001] CRISPR-based genome editing can unpredictably produce on-target deletions, such as long DNA deletions (e.g., >500bp) in the genome of cells. The possibility of generating these DNA deletions brings safety risks to somatic cell genome editing and makes the results of genome editing less predictable. Therefore, for safe and efficient genome editing, compositions and methods are needed to edit cell genomes while reducing on-target deletions. Summary of the Invention [Means for solving the problem]
[0004] overview
[0002] Provided herein is a mammalian expression plasmid comprising a eukaryotic promoter operably linked to a non-viral nucleic acid sequence, the non-viral nucleic acid sequence comprising (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) a guide RNA (gRNA) coding sequence comprising at least one aptamer coding sequence. In some embodiments, the polypeptide comprising a DNA polymerase domain comprises E. coli DNA polymerase I (DNA Pol I) or a fragment thereof. In some embodiments, the polypeptide comprising a DNA polymerase domain comprises a Klenow fragment of E. coli DNA polymerase I (DNA Pol I).
[0004] In some embodiments, the CRISPR-associated endonuclease coding sequence encodes a Cas9 protein or a Cpf1 protein.
[0005] In some embodiments, the at least one aptamer coding sequence encodes an aptamer sequence that is specifically bound by an ABP selected from the group consisting of MS2 coat protein, PP7 coat protein, lambda N RNA binding domain, or Com protein. In some embodiments, the aptamer is an MS2 aptamer sequence or a com aptamer sequence. In some embodiments, the sgRNA coding sequence comprises at least one aptamer coding sequence inserted into the tetraloop or ST2 loop of the sgRNA coding sequence. In some embodiments, the sgRNA code comprises at least one com aptamer inserted into the ST2 loop of the gRNA coding sequence.
[0006] Also provided is a lentiviral packaging system comprising: a) a packaging plasmid comprising a eukaryotic promoter operably linked to a Gag nucleotide sequence, the Gag nucleotide sequence comprising a nucleocapsid (NC) coding sequence and a matrix protein (MA) coding sequence, wherein one or both of the NC coding sequence or the MA coding sequence comprises at least one non-viral aptamer binding protein (ABP) nucleotide sequence, and wherein the packaging plasmid does not encode a functional integrase protein; b) at least one mammalian expression plasmid provided herein; and c) an envelope plasmid comprising an envelope glycoprotein coding sequence.
[0007] In some embodiments, the packaging plasmid further comprises a Rev nucleotide sequence and a Tat nucleotide sequence. In some embodiments, the system further comprises a second packaging plasmid comprising a Rev nucleotide sequence. In some embodiments, the at least one non-viral ABP nucleotide sequence encodes an MS2 coat protein, a PP7 coat protein, a lambda N peptide, or a Com protein.
[0008] Further provided is a lentiviral particle comprising: (a) a fusion protein comprising a nucleocapsid (NC) protein or a matrix (MA) protein, wherein the NC protein or the MA protein comprises at least one non-viral aptamer-binding protein (ABP); and (b) a ribonucleotide protein (RNP) complex, the ribonucleotide protein (RNP) complex comprising: (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) a guide RNA (gRNA) coding sequence, wherein the gRNA coding sequence comprises at least one aptamer coding sequence, and wherein the lentivirus-like particle does not comprise a functional integrase protein.
[0009] In some embodiments, the CRISPR-associated endonuclease coding sequence encodes a Cas9 protein or a Cpf1 protein.
[0010] In some embodiments, the fusion polypeptide comprises E. coli DNA Pol I. In some embodiments, the polypeptide comprises the Klenow fragment of DNA Pol I.
[0011] Also provided is a method of producing lentiviral particles, comprising: a) transfecting a plurality of eukaryotic cells with a packaging plasmid, at least one mammalian expression plasmid and an envelope plasmid of any of the systems provided herein; and b) culturing the transfected eukaryotic cells for a sufficient period of time such that lentiviral particles are produced.
[0012] In some embodiments, the lentiviral particle comprises: (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) a ribonucleotide-protein (RNP) complex comprising a guide RNA. In some embodiments, the plurality of eukaryotic cells are mammalian cells.
[0013] Also provided are cells that contain any of the plasmids, lentiviral packaging systems, or lentivirus-like particles described herein.Also provided are cells modified by any of the methods provided herein.
[0014] Further provided is a method of modifying a genomic target sequence in a cell, comprising transducing a plurality of eukaryotic cells with a plurality of viral particles, the plurality of viral particles comprising: (a) a fusion protein comprising an NC protein or an MA protein, wherein the NC protein or the MA protein comprises at least one non-viral ABP; and (b) an RNP complex comprising: (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain, and (b) a CRISPR-associated endonuclease coding sequence; and (ii) a gRNA coding sequence, wherein the gRNA coding sequence comprises at least one aptamer coding sequence, and the lentivirus-like particle does not comprise a functional integrase protein, wherein the RNP complex binds to the genomic target sequence in genomic DNA of the cell, and the CRISPR-associated endonuclease cleaves the genomic target sequence to generate a double-stranded break, thereby modifying the genomic target sequence.
[0015] In some embodiments, NJEH is increased compared to non-NHEJ end joining in cells compared to cells modified by CRISPR-associated endonuclease that is not fused to DNA polymerase domain.In some embodiments, non-NJEH is microhomology-mediated end joining (MMEJ) and / or single-strand annealing (SSA).
[0016] In some embodiments, the number of on-target deletions over 500 base pairs in size is reduced in the cell compared to the cell. In some embodiments, the ratio of on-target 1 base pair (1 bp) deletions to on-target deletions over 1 bp is increased in the cell. In some embodiments, the ratio of on-target 1 base pair (1 bp) deletions to deletions over 500 base pairs is increased in the cell. In some embodiments, the number of templated insertions (TIS) is increased in the cell. In some embodiments, the ratio of TIS to non-TIS is increased in the cell.
[0017] In some embodiments, the plurality of eukaryotic cells are cells present in a subject. In some embodiments, the subject is a human subject. In some embodiments, the subject is injected with a plurality of viral particles.
[0018] Also provided is a method for treating a disease in a subject, comprising: a) obtaining cells from a subject; b) modifying the cells of the subject using the method provided herein using lentivirus-like particles provided; and c) administering the modified cells to the subject. In some embodiments, the disease is cancer. In some embodiments, the disease is Duchenne muscular dystrophy. In some embodiments, the cells are T cells.
[0019] Description of the drawings This application includes the following drawings. The drawings are intended to illustrate certain embodiments and / or features of the compositions and methods, and to supplement any description(s) of the compositions and methods. The drawings do not limit the scope of the compositions and methods, unless the written description expressly indicates otherwise. [Brief description of the drawings]
[0020] [Figure 1A] FIG. 1 shows countering DNA resection by DNA polymerase I (DNA Pol I or pol I) fused to Cas9 (e.g., by MRE11). The expected outcome is suppression of microhomology-mediated end joining (MMEJ) and single-strand annealing (SSA) DNA repair pathways that require DNA resection.
[0021] [Figure 1B] Shows that DNA polymerase 1 generates a one base pair (1 bp) insertion (TIS) by filling in the 5' overhang. Red nucleotides are filled in by DNA polymerase. As an example, we used the target site of CLCN5 sgRNA ((GAGGACAAGTCGTACAATGGTGG) (SEQ ID NO: 111) and its complement (CTCCTGTTCAGCATGTTACCACC) (SEQ ID NO: 112)). The cleavage sites on both strands generate a one nucleotide (1 nt) 5' overhang, as indicated by the small arrows.
[0022] [Figure 1C] FIG. 13 is a graph showing that fusion of DNA pol I to Cas9 increased the percentage of 1 bp deletions and decreased >1 bp deletions targeting the CLCN5 gene in HEK293T cells. Two-way ANOVA followed by Bonferroni post-hoc test. Replicate numbers are listed in parentheses.
[0023] [Figure 1D]FIG. 13 is a graph showing that fusion of DNA pol I to Cas9 (Cas9-pol I) increased the ratio of 1 bp TIS to 1 bp non-TIS (two-tailed t-test).
[0024] [Figure 2A]
[0023] Figure 1 shows Cas9-pol I and various mutant fusion proteins tested in the studies described in the Examples. The dashed lines indicate the deleted regions.
[0025] [Figure 2B] Graph showing the effect of different DNA pol I mutants on the deletion profile targeted to the CLCN5 gene in HEK293T cells. Numbers in brackets indicate biological replicates. Groups above the red dashed line showed statistical significance compared to groups below the line.
[0026] [Figure 2C] Graph showing the effect of different DNA pol I mutants on targeted insertion of the CLCN5 gene in HEK293T cells. Numbers in brackets indicate biological replicates. Groups above the red dashed line showed statistical significance compared to groups below the line.
[0027] [Figure 2D] Graph showing the effect of different DNA pol I mutants on the deletion profile targeted to the CLCN5 gene in IMR90 cells. The number of replicates was 3 for all groups. * and ** indicate p<0.05 and p<0.01, respectively, between the indicated group and all other groups (2-way ANOVA followed by Bonferroni post-hoc test).
[0028] [Figure 2E]Graph showing the effect of different DNA pol I mutants on targeted insertion of the CLCN5 gene in IMR90 cells. The number of replicates was 3 for all groups. * and ** indicate p<0.05 and p<0.01, respectively, between the indicated group and all other groups (2-way ANOVA followed by Bonferroni post-hoc test).
[0029] [Figure 3A] Graph showing the effect of RBBP8 knockdown on Cas9-induced DNA mutation profile targeting CLCN5 in HEK293T cells. The number of replicates was 6 for Cas9 and 3 for the remaining groups. *, **, and *** indicate p<0.05, p<0.01, and p<0.001 between the indicated groups (2-way ANOVA followed by Bonferroni post-hoc test).
[0030] [Figure 3B] Graph showing the effect of RBBP8 knockdown on Cas9-induced insertion targeting CLCN5 in HEK293T cells. The number of replicates was 6 for Cas9 and 3 for the remaining groups. *, **, and *** indicate p<0.05, p<0.01, and p<0.001 between the indicated groups (2-way ANOVA followed by Bonferroni post-hoc test).
[0031] [Figure 3C] 1 is a graph showing the effect of RBBP8 knockdown on CLCN5-targeted deletions of various sizes in HEK293T cells. Single and double arrows indicate excision-dependent and -independent deletions, respectively.
[0032] [Figure 3D]1 shows the most frequently observed deletions generated by Cas9 targeting CLCN5 in HEK293T cells. A partial wild-type CLCN5 sequence (SEQ ID NO:113) is shown, as well as SEQ ID NO:113 with an 11 base pair deletion (SEQ ID NO:114), SEQ ID NO:113 with a different 11 base pair deletion (SEQ ID NO:115), SEQ ID NO:113 with an 8 base pair deletion (SEQ ID NO:116), SEQ ID NO:113 with a different 8 base pair deletion (SEQ ID NO:117), and SEQ ID NO:113 with a 16 base pair deletion (SEQ ID NO:118). Dashed lines indicate deletions. Microhomologies are underlined. Microhomologies away from the cleavage site are indicated with a caret symbol (^). Asterisks indicate microhomologies at the predicted cleavage size, which are indicated by vertical dashed lines.
[0033] [Figure 4A] Examination of suppressed deletions targeting the HBB gene in hematopoietic cells is shown. * and *** indicate p<0.05 and p<0.001 between Cas9 and Cas9-Klenow, respectively (Bonferroni post-hoc test after two-way ANOVA). Partial wild-type HBB sequence (SEQ ID NO:119) and SEQ ID NO:119 with a 3 base pair deletion (SEQ ID NO:120) and SEQ ID NO:119 with a 12 base pair deletion (SEQ ID NO:121) are shown. The top image shows the peak of deletions, and the bottom image shows the most frequently observed deletions. Microhomologies are underlined. Asterisks (*) indicate microhomologies at the predicted cleavage size, which are indicated by vertical dashed lines. Carets (^) indicate microhomologies away from the predicted cleavage site. Each group has three biological replicates.
[0034] [Figure 4B]Examination of suppressed deletions targeting DMD exon 53 in HEK293T cells. * and *** indicate p<0.05 and p<0.001 between Cas9 and Cas9-Klenow, respectively (Bonferroni post-hoc test after two-way ANOVA). Partial wild-type DMD exon 53 sequence (SEQ ID NO:122) and SEQ ID NO:122 with an 11 base pair deletion (SEQ ID NO:123), SEQ ID NO:122 with a 9 base pair deletion (SEQ ID NO:124) and SEQ ID NO:122 with a 6 base pair deletion (SEQ ID NO:125) are shown. The top image shows the peak of deletions and the bottom image shows the most frequently observed deletions. Microhomologies are underlined. Asterisks (*) indicate microhomologies at the predicted cleavage size, which are indicated by vertical dashed lines. Carets (^) indicate microhomologies away from the predicted cleavage site. Each group has three biological replicates.
[0035] [Figure 4C] Examination of suppressed deletions targeting the HBB gene in IMR90 cells. * and *** indicate p<0.05 and p<0.001 between Cas9 and Cas9-Klenow, respectively (Bonferroni post-hoc test after two-way ANOVA). A partial wild-type HBB sequence (SEQ ID NO:126) is shown, as well as SEQ ID NO:126 containing a 3 base pair deletion (SEQ ID NO:127) and SEQ ID NO:126 containing a 5 base pair deletion (SEQ ID NO:128). The top image shows the peak of deletions, and the bottom image shows the most frequently observed deletion. Microhomologies are underlined. Asterisks (*) indicate microhomologies at the predicted cleavage size, which are indicated by vertical dashed lines. Carets (^) indicate microhomologies away from the predicted cleavage site. Each group has three biological replicates.
[0036] [Figure 4D]Examination of suppressed deletions targeting DMD exon 44 in human myoblasts. * and *** indicate p<0.05 and p<0.001 between Cas9 and Cas9-Klenow, respectively (Bonferroni post-hoc test after two-way ANOVA). Shown is the partial wild-type DMD44 sequence (SEQ ID NO:129), as well as SEQ ID NO:129 with a 10 base pair deletion (SEQ ID NO:130), SEQ ID NO:129 with a 7 base pair deletion (SEQ ID NO:131), and SEQ ID NO:129 with a different 7 base pair deletion (SEQ ID NO:132). The top image shows the peak of deletions, and the bottom image shows the most frequently observed deletions. Microhomologies are underlined. Asterisks (*) indicate microhomologies at the predicted cleavage size, which are indicated by vertical dashed lines. Carets (^) indicate microhomologies away from the predicted cleavage site. Each group has three biological replicates.
[0037] [Figure 5A] Figure 1 shows a large deletion generated by Cas9 and Cas9-Klenow targeting the CLCN5 gene in HEK293T cells. Data was combined from three replicate experiments. Regions labeled 434 and 4534 indicate two 25 bp sequences used to calculate distance for deletion detection. Regions labeled 2471, 2481 and 2487 indicate sgRNA targets, and region labeled 2494 indicates PAM. Asterisks indicate identical deletions observed independently in two different experiments.
[0038] [Figure 5B] Figure 1 shows a large deletion generated by Cas9 and Cas9-Klenow targeting the CLCN5 gene in IMR90 cells. Data was combined from three replicate experiments. Regions labeled 434 and 4534 indicate two 25 bp sequences used to calculate distance for deletion detection. Regions labeled 2471, 2489 and 2488 indicate sgRNA targets, and region labeled 2494 indicates PAM. Asterisks indicate identical deletions observed independently in two different experiments.
[0039] [Figure 6] Next generation sequencing (NGS) analysis of integrated target sequences in GFP reporter cells treated with CLCN5 sgRNA and Cas9-pol. SEQ ID NOs: 133-141 are shown.
[0040] [Figure 7] 1 is a graph showing that Cas9 and various exonuclease fusions had similar mutational profiles as Cas9. Numbers in parentheses indicate repeat numbers.
[0041] [Figure 8A] Graph showing the effect of different pol I domains on 2bp TIS. Data for exo, 3'exo and 5'exo fusion proteins targeting CLCN5 in HEK293T cells were pooled into one group. Each dot represents one reference point. ** and *** indicate p<0.05 and p<0.001, respectively, compared to the Cas9 group. Tukey's multiple comparison test was performed following one-way ANOVA. Cas9-pol showed a trend toward an increase but did not reach statistical significance due to large within-group variation.
[0042] [Figure 8B] Graph showing the effect of different pol I domains on 3bp TIS. Data for exo, 3'exo and 5'exo fusion proteins targeting CLCN5 in HEK293T cells were pooled into one group. Each dot represents one reference point. ** and *** indicate p<0.05 and p<0.001, respectively, compared to the Cas9 group. Tukey's multiple comparison test was performed following one-way ANOVA. Cas9-pol showed a trend toward an increase but did not reach statistical significance due to large within-group variation.
[0043] [Figure 9]FIG. 13 is a graph showing that the overall INDEL rate did not affect the percentage of different variants.
[0044] [Figure 10A] FIG. 11 is a graph showing the effect of MRE11 or RBBP8 knockdown on CLCN5 mutation profile in IMR90 cells (% of all INDELs). Each group had three biological replicates. Two-way ANOVA followed by Bonferroni post-hoc test. ** and *** indicate p<0.01 and p<0.001, respectively, between the indicated groups.
[0045] [Figure 10B] FIG. 13 is a graph showing the effect of MRE11 or RBBP8 knockdown on CLCN5 mutation profile in IMR90 cells (% of total insertions). Each group had three biological replicates. Two-way ANOVA followed by Bonferroni post-hoc test. ** and *** indicate p<0.01 and p<0.001, respectively, between the indicated groups.
[0046] [Figure 11A] Figure 1 shows the deletions reduced by Cas9-Klenow when targeting HBB in HEK293T cells. The graph (top image) shows the percentage of deletion size. The bottom image shows the sequences of the most commonly observed deletions. Shown are the partial wild-type HBB sequence (SEQ ID NO:142), as well as SEQ ID NO:142 with a 3 base pair deletion (SEQ ID NO:143), SEQ ID NO:142 with a 5 base pair deletion (SEQ ID NO:144), and SEQ ID NO:142 with a 10 base pair deletion (SEQ ID NO:145). The underlined regions in green indicate microhomology at the predicted cleavage site, which is shown by the vertical dashed line. Two-way ANOVA was followed by Bonferroni post-hoc test. * and *** indicate p<0.05 and P<0.001 between the two groups, respectively.
[0047] [Figure 11B]Figure 1 shows reduced deletions by Cas9-Klenow when targeting intergenic 1 in IMR90 cells. The graph (top image) shows the percentage of deletion size. The bottom image shows the sequences of the most commonly observed deletions. A partial wild type intergenic 1 sequence (SEQ ID NO: 146) is shown, as well as SEQ ID NO: 146 containing a 4 base pair deletion (SEQ ID NO: 147). The sgRNA is shown shaded (grey). The PAM is overlined.
[0048] [Figure 11C] Figure 1 shows deletions reduced by Cas9-Klenow when targeting CLCN5 in IMR90 cells. The graph (top image) shows the percentage of deletion size. The bottom image shows the sequences of the most commonly observed deletions. It was not possible to enumerate sequences for CLCN5 / IMR90 as no major deletion peaks were observed. The sgRNA is shown shaded (grey). The PAM is shown overlined. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0049] definition As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0050] The use of any and all examples or exemplary language (e.g., "etc.") provided herein is intended merely to better describe the invention and does not limit the scope of the invention unless specifically claimed.
[0051] The terms "may," "may be," "can," and "can be," and related terms, unless the context clearly dictates otherwise, are intended to convey that the subject matter involved is optional (i.e., the subject matter may be present in some instances and absent in others), and are not a reference to the ability or probability of the subject matter.
[0052] The terms "optional" and "optionally" mean that the subsequently described event, circumstance, or material may or may not occur, or may or may not be present, and that the description includes instances where the event, circumstance, or material does occur or is present, as well as instances where it does not occur or is not present.
[0053] Use of the terms "including," "comprising," or "having" and variations thereof herein is meant to encompass the subsequently listed elements and equivalents thereof as well as additional elements. Embodiments recited as "including," "comprising," or "having" specific elements are also considered to "consist essentially of," and "consist of," those specific elements. As used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combination when interpreted in the alternative ("or").
[0054] As used herein, the transitional phrase "consisting essentially of" (and grammatical variations) is to be construed to include the recited materials or steps and "which do not materially affect the basic and novel characteristic(s)" of the claimed invention. See In re Herz, 537 F.2d 549,551-52,190 USPQ461,463 (CCPA 1976) (emphasis in original). See also MPEP §2111.03. Thus, the term "consisting essentially of" as used herein should not be construed as equivalent to "comprising."
[0055] The term "nucleic acid" or "nucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) (e.g., mRNA) and polymers thereof in single-stranded or double-stranded form. When RNA is described, its corresponding DNA is also described, and it is understood that uridine is represented as thymidine. Similarly, when DNA is described, its corresponding RNA is also described, and thymidine is represented by uridine. Unless specifically limited, the term encompasses nucleic acids that contain known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., Nucleic Acid Res. 19:5081 (1991); al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The polynucleotides of the present invention also encompass all forms of sequence, including, but not limited to, single-stranded forms, double-stranded forms, hairpins, stem and loop structures, and the like.
[0056] The term "gene" can refer to a segment of DNA involved in producing or encoding a polypeptide chain. It can include regions before and after the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). Alternatively, the term "gene" can refer to a segment of DNA involved in producing or encoding a non-translated RNA, such as rRNA, tRNA, guide RNA, or microRNA.
[0057] "Treat" refers to any indication of success in treating or improving or preventing a disease, condition or disorder, including any objective or subjective parameter, such as reduction; remission; reducing symptoms or making disease symptoms more tolerable to the patient; slowing down the rate of degeneration or decline; or making the end point of degeneration less debilitating. Treating or improving symptoms can be based on objective or subjective parameters, including the results of a physician's examination. Thus, the term "treat" includes the administration of a compound or agent of the present disclosure to prevent or delay, alleviate, or stop or inhibit the onset of symptoms or symptoms associated with a disease, condition or disorder described herein. The term "therapeutic effect" refers to the reduction, elimination, or prevention of a disease, a symptom of a disease, or a side effect of a disease in a subject. "Treating" or "treatment" using the methods of the present disclosure includes preventing the onset of symptoms in a subject who may be at high risk for a disease or disorder associated with the disease, condition, or disorder described herein, but has not yet experienced or exhibited symptoms, inhibiting (delaying or halting) the symptoms of the disease or disorder, relieving (including palliative) symptoms or side effects of the disease, and relieving (causing regression) symptoms of the disease. Treatment can be prophylactic (to prevent or delay the onset of the disease or to prevent the onset of its clinical or subclinical symptoms), or therapeutic suppression or alleviation of symptoms after the onset of the disease or condition. As used herein, the term "treatment" includes preventative (e.g., prophylactic), curative, or palliative treatment.
[0058] "Promoter" is defined as one or more nucleic acid control sequences that direct the transcription of a nucleic acid. As used herein, promoter includes the necessary nucleic acid sequence near the start site of transcription, for example, in the case of a polymerase II type promoter, a TATA element. Promoter also includes, if necessary, distal enhancer or repressor elements that may be located several thousand base pairs from the start site of transcription.
[0059] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. All three terms apply to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of the corresponding naturally occurring amino acids, as well as to natural and non-natural amino acid polymers. As used herein, the term encompasses full-length proteins, truncated proteins, and fragments thereof, as well as amino acid chains in which the amino acid residues are linked by covalent peptide bonds. The term "fusion polypeptide" or "fusion protein," as used throughout, is a polypeptide that includes two or more proteins or fragments thereof. In some embodiments, a linker containing about 3-10 amino acids can be placed between any two proteins or fragments thereof to help promote proper folding of the proteins upon expression.
[0060] The term "identity" or "substantial identity" as used in the context of polynucleotide or polypeptide sequences described herein refers to a sequence having at least 60% sequence identity with a reference sequence. Alternatively, the percent identity can be any integer between 60% and 100%. Exemplary embodiments include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% as compared to a reference sequence using the programs described herein; preferably, BLAST using standard parameters as described below. It is understood that any nucleotide or polypeptide sequence described herein, such as a sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to any one of SEQ ID NOs: 1-110, can be used in the compositions and methods provided herein. It is understood that a nucleic acid sequence can comprise, consist of, or consist essentially of any nucleic acid sequence described herein. Similarly, a polypeptide can comprise, consist of, or consist essentially of any polypeptide sequence described herein. For sequence comparison, typically one sequence acts as a reference sequence to which a test sequence is compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence relative to the reference sequence based on the program parameters.
[0061] As used herein, a "comparison window" includes reference to any one segment of a number of consecutive positions selected from the group consisting of 20-600, about 20-50, about 20-100, about 50-200, or about 100-150, and a sequence may be compared to a reference sequence of the same number of consecutive positions after the two sequences are optimally aligned. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be performed by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), by the homologous sequence alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988), by computerized implementations of these algorithms (e.g., BLAST), or by manual alignment and visual inspection.
[0062] Suitable algorithms for determining percent sequence identity and percent sequence similarity are described in Altschul et al. (1990) J. Mol. Biol. 215:403-410 and Altschul et al. (1977) Nucleic Acid The BLAST and BLAST 2.0 algorithms are described in Acids Res. 25:3389-3402. Software for performing BLAST analysis is publicly available through the website of the National Center for Biotechnology Information (NCBI). This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or meet some positive threshold score T when aligned with words of the same length in the database sequence. T is called the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds to initiate searches to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence as far as the cumulative alignment score can be increased. The cumulative score is calculated using the parameters M (reward score for a pair of matching residuals; always >0) and N (penalty score for mismatching residuals; always <0) for nucleotide sequences. Extension of the word hits in each direction is stopped when the cumulative alignment score falls by an amount X from its maximum achieved value. The cumulative score falls below 0 due to the accumulation of one or more negative scoring residue alignments, or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=1, N=-2, and a comparison of both strands.
[0063] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, for example, Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered to be similar to a reference sequence if the minimum sum probability in the comparison between the test nucleic acid and the reference nucleic acid is less than about 0.01, more preferably less than about 10-5, and most preferably less than about 10-20.
[0064] As used throughout, subject means an individual. For example, subject is a mammal, such as a primate, more specifically a human. Non-human primates are also subjects. The term subject includes domestic animals, such as cats, dogs, livestock animals (e.g., cows, horses, pigs, sheep, goats, etc.), and laboratory animals (e.g., ferrets, chinchillas, mice, rabbits, rats, gerbils, guinea pigs, etc.). Thus, veterinary uses as well as medical uses and formulations are contemplated herein. The term does not indicate a particular age or sex. Thus, living and newborn subjects are intended to be covered, whether male or female. As used herein, patient or subject can be used interchangeably and can refer to a subject suffering from a disease or disorder.
[0065] An "expression cassette" is a recombinantly or synthetically produced nucleic acid construct that has a set of specific nucleic acid elements that allow the transcription of a specific polynucleotide sequence in a host cell. An expression cassette can be part of a plasmid, a viral genome, or a nucleic acid fragment. Typically, an expression cassette contains a polynucleotide to be transcribed that is operably linked to a promoter and is followed by a transcription termination signal sequence. An expression cassette may or may not contain specific regulatory sequences, such as 5' or 3' untranslated regions from human globin genes.
[0066] A "reporter gene" encodes a protein that is easily detectable due to a biochemical characteristic, such as an enzymatic activity or a chemiluminescent property. These reporter proteins can be used as selectable markers. One specific example of such a reporter is green fluorescent protein. The fluorescence generated from this protein can be detected with a variety of commercially available fluorescence detection systems. Other reporters can be detected by staining. A reporter can also be an enzyme that generates a detectable signal when contacted with an appropriate substrate. A reporter can be an enzyme that catalyzes the formation of a detectable product. Suitable enzymes include, but are not limited to, proteases, nucleases, lipases, phosphatases and hydrolases. A reporter can encode an enzyme whose substrate is substantially impermeable to the eukaryotic plasma membrane, thus allowing for tight control of signal formation. Specific examples of suitable reporter genes encoding enzymes include, but are not limited to, CAT (chloramphenicol acetyltransferase; Alton and Vapnek (1979) Nature 282:864-869); luciferase (lux); β-galactosidase; LacZ; β-glucuronidase; and alkaline phosphatase (Toh, et al. (1980) Eur. J. Biochem. 182:231-238; and Hall et al. (1983) J. Mol. Appl. Gen. 2:101), each of which is incorporated herein by reference in its entirety. Other suitable reporters include those that encode specific epitopes that can be detected with a labeled antibody that specifically recognizes the epitope.
[0067] "CRISPR / Cas" systems refer to a broad class of bacterial systems for defense against foreign nucleic acids. CRISPR / Cas systems are found in a wide range of eubacterial and archaeal organisms. CRISPR / Cas systems include type I, type II and type III subtypes. The classification of CRISPR / Cas systems as described by Makarova, et al. (Nat Rev Microbiol. 2015 Nov;13(11):722-36) defines five types and 16 subtypes based on shared characteristics and evolutionary similarities. These are divided into two broad classes based on the structure of the effector complex that cleaves genomic DNA. Type II CRISPR / Cas systems were first used for genome engineering, followed by type V in 2015. Wild-type type II CRISPR / Cas systems utilize an RNA-mediated nuclease Cas protein or homolog (referred to herein as "CRISPR-associated endonucleases") that form a complex with a guide RNA to recognize and cleave foreign nucleic acids. Cas9 protein also uses activating RNA (also called transactivating RNA or tracr RNA). Guide RNA or guide RNA with both guide RNA and activating RNA activity is also known in the art depending on the type of CRISPR-associated endonuclease used with it. In some cases, such dual-activity guide RNA is called single guide RNA (sgRNA). Synthetic guide RNA that does not contain activating RNA sequence may also be called sgRNA. In this disclosure, the terms sgRNA and gRNA are used interchangeably to refer to RNA molecules that form a complex with CRISPR-associated endonuclease and localize CRISPR-associated endonuclease, for example, CRISPR-associated endonuclease in a ribonucleoprotein complex, to a target DNA sequence.
[0068] As used herein, "activity" in the context of sgRNA activity or RNP activity, i.e., the RNP activity of a complex comprising (1) a gRNA and (2) a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain, refers to the ability of the sgRNA to bind to a target genetic element. Typically, activity also refers to the ability of the RNP (i.e., the sgRNA complexed with a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain) to edit the genome of a cell. As used herein, the term "editing" in the context of editing the genome of a cell refers to inducing a structural change in the sequence of the genome at a target genomic region, for example, cleaving the genomic sequence and inserting a donor sequence at the cleavage site into the genome of the cell via homology-directed repair (HDR), or cleaving the sequence and allowing repair via non-homologous recombination end joining (NHEJ).
[0069] As used herein, terms such as "ribonucleoprotein complex," "RNP," and the like refer to a complex between: (1) a fusion protein comprising a CRISPR-associated endonuclease (e.g., Cas9) and a DNA polymerase domain, and a crRNA (e.g., a guide RNA or a single guide RNA), (2) a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain; and a trans-activating crRNA (tracrRNA), (3) a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain, and a guide RNA, or (4) a combination thereof (e.g., a complex containing a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain, a tracrRNA, and a crRNA guide).
[0070] As used herein, a "cell" can be any eukaryotic cell, such as a human T cell, or a cell that can differentiate into a T cell (e.g., a T cell that expresses a TCR receptor molecule). These include hematopoietic stem cells and cells derived from hematopoietic stem cells. Also provided is a population of cells, such as a population of cells that comprises viral particles or genetically modified cells produced by any of the genome editing methods provided herein.
[0071] As used herein, the phrase "hematopoietic stem cell" refers to a type of stem cell that can give rise to blood cells. Hematopoietic stem cells can give rise to cells of myeloid or lymphoid lineages, or a combination thereof. Hematopoietic stem cells are primarily found in bone marrow, but can be isolated from peripheral blood or a portion thereof. Various cell surface markers can be used to identify, select, or purify hematopoietic stem cells. In some cases, hematopoietic stem cells are identified as c-kit+ and lin-. In some cases, human hematopoietic stem cells are identified as CD34+, CD59+, Thy1 / CD90+, CD38lo / -, C-kit / CD117+, lin-. In some cases, human hematopoietic stem cells are identified as CD34-, CD59+, Thy1 / CD90+, CD38lo / -, C-kit / CD117+, lin-. In some cases, the human hematopoietic stem cells are identified as CD133+, CD59+, Thy1 / CD90+, CD38lo / -, C-kit / CD117+, lin-. In some cases, the mouse hematopoietic stem cells are identified as CD34lo / -, SCA-1+, Thy1+ / lo, CD38+, C-kit+, lin-. In some cases, the hematopoietic stem cells are CD150+CD48-CD244-.
[0072] As used herein, the phrase "hematopoietic cells" refers to cells derived from hematopoietic stem cells. Hematopoietic cells can be obtained or provided by isolation from an organism, system, organ or tissue (e.g., blood or a portion thereof). Alternatively, hematopoietic stem cells can be isolated and hematopoietic cells can be obtained or provided by differentiating the stem cells. Hematopoietic cells include cells that have limited potential to differentiate into additional cell types. Such hematopoietic cells include, but are not limited to, multipotent progenitor cells, lineage-restricted progenitor cells, common myeloid progenitor cells, granulocyte-macrophage progenitor cells, or megakaryocyte-erythroid progenitor cells. Hematopoietic cells include lymphoid and myeloid cells, such as lymphocytes, erythrocytes, granulocytes, monocytes, and platelets. In some embodiments, the hematopoietic cells are immune cells, such as T cells, B cells, macrophages, natural killer (NK) cells, or dendritic cells. In some embodiments, the cells are innate immune cells.
[0073] As used herein, the phrase "T cells" refers to lymphoid cells that express T cell receptor molecules. T cells include human alpha beta (αβ) T cells and human gamma delta (γδ) T cells. T cells include, but are not limited to, naive T cells, stimulated T cells, primary T cells (e.g., not cultured), cultured T cells, immortalized T cells, helper T cells, cytotoxic T cells, memory T cells, regulatory T cells, natural killer T cells, combinations thereof, or subpopulations thereof. T cells can be CD4+, CD8+, or CD4+ and CD8+. T cells can also be CD4-, CD8-, or CD4- and CD8-. T cells can be helper cells, such as TH1, TH2, TH3, TH9, TH17, or TFH types of helper cells. T cells can be cytotoxic T cells. Regulatory T cells can be FOXP3+ or FOXP3-. The T cells can be alpha / beta T cells or gamma / delta T cells. In some cases, the T cells are CD4+CD25hiCD127lo regulatory T cells. In some cases, the T cells are regulatory T cells selected from the group consisting of type 1 regulatory (Tr1), TH3, CD8+CD28-, Treg17, and Qa-1 restricted T cells, or a combination or subpopulation thereof. In some cases, the T cells are FOXP3+ T cells. In some cases, the T cells are CD4+CD25loCD127hi effector T cells. In some cases, the T cells are CD4+CD25loCD127hiCD45RAhiCD45RO- naive T cells. The T cells can be genetically engineered recombinant T cells.
[0074] As used herein, the term "primary" in the context of primary cells is a cell that has not been transformed or immortalized. Such primary cells can be cultured, subcultured, or passaged a limited number of times (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cultures). In some cases, primary cells are adapted to in vitro culture conditions. In some cases, primary cells are isolated from an organism, system, organ, or tissue, optionally sorted, and directly utilized without culture or subculture. In some cases, primary cells are stimulated, activated, or differentiated. For example, primary T cells can be activated by contacting with (e.g., culturing in the presence of) CD3, CD28 agonist, IL-2, IFN-γ, or a combination thereof. Detailed Description
[0075] The following description lists various aspects and embodiments of the present compositions and methods.Specific embodiments are not intended to define the scope of the compositions and methods.Rather, the embodiments merely provide non-limiting examples of various compositions and methods that are at least included within the scope of the disclosed compositions and methods.The description should be read from the perspective of those skilled in the art.Therefore, it does not necessarily include information that is well known to those skilled in the art.
[0076] The CRISPR / Cas9 (clustered regularly interspaced short palindromic repeats and CRISPR-associated protein 9) system has been used to create DNA double-strand breaks (DSBs) guided by a single guide RNA (sgRNA) using a single effector protein to effect specific genetic alterations in human cells. Streptococcus pyogenes Cas9 (SpCas9) primarily generates blunt ends by cleaving both DNA strands 3 nucleotides (nt) upstream of a "NGG" protospacer adjacent motif (PAM). It can also generate staggered ends with 1, 2 or 3 nt 5' overhangs by cleaving the target strand 3 nt and the non-target strand 4, 5 or 6 nt upstream of the PAM.
[0077] Human cells have multiple pathways to repair double-strand breaks (DSBs) created by CRISPR / Cas9 nucleases: homology-directed repair (HDR), canonical non-homologous end joining (cNHEJ or NHEJ), and alternative end-joining pathways including microhomology-mediated end joining (MMEJ) and single-strand annealing (SSA). In the cNHEJ pathway, the two ends of the DSB are processed for religation without resection or template involvement. Pol X family members (Pol λ, μ, and β and terminal transferase) can function to fill in the 5' overhang. This fill-in reaction can explain the widely observed predictable insertion or templated insertion (TIS) induced by SpCas9.
[0078] HDR, MMEJ and SSA repair pathways all depend on DNA resection to generate 3' overhangs. DNA resection is initiated by the MRE11-RAD50-NBS1 complex and stimulated by CtIP (encoded by RBBP8). The endonuclease activity of MRE11 generates a nick 3' to the DSB, and the 3'-5' exonuclease activity generates a 3' single-stranded DNA overhang from the nick. HDR faithfully repairs DSBs using long 3' overhangs and DNA templates. MMEJ and SSA use microhomology (2-5 nt) and short homology (10-15 nt), respectively, in the two 3' overhangs to drive DNA synthesis, generating DNA deletions with sizes depending on the distance between the homologous regions. Long 3' overhangs from excess DNA resection can also be filled in by DNA polα and associated complexes. This fill-in reaction explains the small tandem duplications observed at DSBs with 3' overhangs generated by Cas9 nickase.
[0079] Although CRISPR / Cas9-induced mutation profiles are somewhat predictable, with small insertions and deletions (INDELs) being the predominant mutation type, unpredictable on-target large deletions have been widely reported. Currently, methods to suppress the generation of large deletions are lacking. Although genome editing methods that do not generate DSBs, such as base editing and prime editing, have been developed, many applications, including eradication of integrated proviral DNA from human cells and removal of disease-causing DNA repeats, still involve DSB generation. Thus, methods that can generate improved mutation profiles would improve safety.
[0080] Provided herein are compositions, systems, methods of manufacture and methods for genome editing.Using the compositions and methods described herein, the genome of a cell can be efficiently edited using CRISPR-associated endonuclease, but the number of long deletions generated in the genome of a cell is reduced, for example, by increasing the non-homologous end joining (NHEJ) of the cell compared to the non-NHEJ of the cell (for example, microhomology-mediated end joining (MMEJ) or single-strand annealing (SSA)).Provided herein are exemplary components, systems, methods of manufacture and methods for editing the genome of a cell using (1) a fusion protein comprising a CRISPR-associated endonuclease and a polypeptide comprising a DNA polymerase domain; and (2) an RNP comprising sgRNA. Mammalian expression plasmids
[0081] Provided herein is a mammalian expression plasmid that is used to deliver the CRISPR component coding sequence, i.e., sgRNA and the fusion protein that comprises CRISPR-associated endonuclease and DNA polymerase domain, to the mammalian cell that is used to generate the lentivirus-like particle of the present disclosure.For example, provided herein is a mammalian expression plasmid that comprises a eukaryotic promoter operably linked to a non-viral nucleic acid sequence, the non-viral nucleic acid sequence comprising (i) a nucleic acid sequence that encodes a fusion protein, the fusion protein comprising (a) a polypeptide that comprises a DNA polymerase domain and (b) a sequence that encodes a CRISPR-associated endonuclease, and (ii) a guide RNA (gRNA) coding sequence that comprises at least one aptamer coding sequence.
[0082] In the mammalian expression plasmids provided herein, the polypeptide comprising the DNA polymerase domain can be fused or linked to the CRISPR-associated endonuclease. Optionally, the CRISPR-associated endonuclease is linked to the polypeptide comprising the DNA polymerase domain via a peptide linker. The linker can be between about 2 and about 25 amino acids in length.
[0083] In some embodiments, the CRISPR-associated endonuclease coding sequence encodes a Cas9 protein. Full-length Cas9 is an endonuclease that includes a recognition domain and two nuclease domains (HNH and RuvC, respectively) that generate double-stranded breaks in DNA sequences. Cas9 targets genomic sites in cells by interacting with guide RNAs that hybridize to a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by Cas9. This results in a double-stranded break in the genomic DNA of the cell. In some examples, Cas9 nucleases can be utilized that require a NGG protospacer adjacent motif (PAM) immediately 3' to the region targeted by the guide RNA. As another example, Cas9 proteins with orthogonal PAM motif requirements can be utilized to target sequences that do not have adjacent NGG PAM sequences. Exemplary Cas9 proteins with orthogonal PAM sequence specificity include, but are not limited to, those described in Esvelt et al., Nature Methods 10:1116-1121 (2013).
[0084] Various Cas9 nucleases can be utilized in the methods described herein. For example, Cas9 nucleases (e.g., SpCas9) that require a NGG protospacer adjacent motif (PAM) immediately 3' to the region targeted by the guide RNA can be utilized. Such Cas9 nucleases can target any region of the genome that contains a NGG sequence. In another example, Cas9 nucleases (e.g., SaCas9) that require a NNGRRT or NNGRR(N) PAM immediately 3' to the region targeted by the guide RNA can be utilized. As another example, Cas9 proteins with orthogonal PAM motif requirements can be utilized to target sequences that do not have adjacent NGG PAM sequences. Exemplary Cas9 proteins with orthogonal PAM sequence specificity include, but are not limited to, those described in Nature Methods 10, 1116-1121 (2013) and those described in Zetsche et al., Cell, Volume 163, Issue 3, p759-771, 22 October 2015. An exemplary amino acid sequence of a Cas9 protein is set forth herein as SEQ ID NO:29.
[0085] In some cases, the Cas9 protein is a nickase, so that when it binds to a target nucleic acid as part of a complex with a guide RNA, a single-strand break or nick is introduced into the target nucleic acid. A pair of Cas9 nickases, each bound to a structurally different guide RNA, can target two proximal sites in the target genomic region, thus introducing a pair of proximal single-strand breaks into the target genomic region. The nickase pair can provide enhanced specificity, since off-target effects may result in a single nick, which is generally repaired without damage by base excision repair mechanisms. Exemplary Cas9 nickases include Cas9 nucleases with D10A or H840A mutations.
[0086] In some embodiments, the CRISPR-associated endonuclease is a Cpf1 polypeptide. The Cpf1 protein is a class II, type V CRISPR / Cas system protein. Cpf1 is a smaller and simpler endonuclease than Cas9 (such as spCas9). The Cpf1 protein has a RuvC-like endonuclease domain that is similar to the RuvC domain of Cas9 but does not have an HNH endonuclease domain. The N-terminal domain of Cpf1 also does not have an alpha helix recognition lobe like the Cas9 protein. When cleaving DNA, Cpf1 introduces sticky end-like DNA double-strand breaks with 4 or 5 nucleotide overhangs. The Cpf1 protein does not require tracrRNA; rather, the Cpf1 protein functions only with crRNA. In the context of the present disclosure, when the CRISPR-associated endonuclease is a Cpf1 protein, the sgRNA does not contain a tracr sequence. The sgRNA used with the Cpf1 protein may only include the crRNA sequence (constant region). In some examples, a Cpf1 protein may be utilized that requires a TTTN or TTN PAM (where "N" is a nucleobase, depending on the species) immediately 5' to the region targeted by the guide RNA. Known Cpf1 proteins and derivatives thereof may be used in the context of the present disclosure. For example, in some cases, the CRISPR-associated endonuclease is FnCpf1p, the PAM is 5'TTN, and N is A / C / G or T. In some cases, the CRISPR-associated endonuclease is PaCpf1p, the PAM is 5'TTTV, and V is A / C or G. In some cases, the CRISPR-associated endonuclease is FnCpf1p, the PAM is 5'TTN, and N is A / C / G or T, and the PAM is located upstream of the 5' end of the protospacer. In some cases, CRISPR-associated endonuclease is FnCpf1p, PAM is 5'CTA, and is located upstream of the 5' end of protospacer or target locus.In one embodiment, CRISPR-associated endonuclease is AsCpf1p, and PAM is 5'TTTN.The exemplary amino acid sequence of Cpf1 is shown herein as SEQ ID NO:30.
[0087] As used throughout, a DNA polymerase domain is a fragment of a full-length DNA polymerase that catalyzes the 5'-3' polymerization of nucleotides into double-stranded DNA in the presence of a nucleic acid template. Any DNA polymerase domain or polypeptide that includes a DNA polymerase domain can be used to generate the fusion protein described herein. By fusing the DNA polymerase domain to a CRISPR-associated endonuclease, the DNA polymerase domain targets the site of the double-strand break created by the CRISPR-associated endonuclease in the genome of the cell. In some embodiments, the DNA polymerase domain is catalytically active. In some embodiments, the catalytic activity of the DNA polymerase domain is reduced (e.g., relative to the corresponding wild-type DNA polymerase domain), e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 80%, 90% or 100%. In some embodiments, the DNA polymerase domain includes one or more mutations that eliminate catalytic activity, e.g., the D705A mutation.
[0088] In some cases, a polypeptide comprising a DNA polymerase domain has polymerase activity and 3'-5' exonuclease activity. In some cases, a polypeptide comprising a DNA polymerase domain has 3'-5' exonuclease activity but reduced or absent polymerase activity (e.g., relative to the corresponding wild-type DNA polymerase domain).
[0089] An exemplary polypeptide comprising a DNA polymerase domain is DNA polymerase I (DNA Pol I). DNA Pol I is an enzyme having 5'-3' DNA-dependent DNA polymerase activity, 3'-5' exonuclease activity, 5'-3' exonuclease activity and 5'-3' RNA-dependent DNA polymerase activity. An exemplary sequence of DNA Pol I is provided herein as SEQ ID NO: 27. In some embodiments, a DNA pol I fragment comprising a DNA polymerase domain is fused to a CRISPR-associated endonuclease. In some cases, the fragment comprises an amino acid sequence having 5'-3' DNA polymerase activity and an amino acid sequence having 3'-5' exonuclease activity (e.g., a Klenow fragment of DNA Pol I). An exemplary sequence of a Klenow fragment of DNA Pol I is provided herein as SEQ ID NO: 28. Other exemplary DNA polymerases include, but are not limited to, Thermus aquaticus DNA Pol I and Bacillus stearothermophilus DNA Pol I.
[0090] The mammalian expression plasmids provided herein include coding sequences for CRISPR components, such as a fusion protein including a CRISPR-associated endonuclease and a DNA polymerase domain, and a gRNA. In some examples, the gRNA coding sequence includes at least one aptamer coding sequence. In some examples, the at least one aptamer coding sequence can be located at the 5' or 3' end of the gRNA. In some examples, the at least one aptamer coding sequence can be inserted at an internal location within the gRNA, such as one or more of the loops formed within the folded gRNA. For example, if the gRNA is for a Cas9 protein, the at least one aptamer coding sequence can be located in the tetraloop, stem loop 2 (ST2) or 3' end of the gRNA. In some examples, a spacer of 1-30 nucleotides can be located between the gRNA and the at least one aptamer coding sequence or adjacent to the at least one aptamer coding sequence.
[0091] In some examples, the mammalian expression vector includes at least one aptamer coding sequence that encodes an aptamer sequence that is specifically bound by an aptamer binding protein (ABP). In the context of the present disclosure, an aptamer sequence is an RNA sequence that forms a tertiary loop structure that is specifically bound by an ABP. An ABP is an RNA binding protein or an RNA binding protein domain. Suitable aptamer coding sequences include polynucleotide sequences that encode known bacteriophage aptamer sequences. Exemplary aptamer coding sequences include those that encode the aptamer sequences shown in Table 1 above. In some examples, the aptamer is bound by a dimer of an ABP. These aptamer sequences are RNA sequences to which a bacteriophage protein is known to specifically bind. In some situations, the at least one aptamer coding sequence encodes an aptamer sequence that is specifically bound by an ABP selected from the group consisting of MS2 coat protein, PP7 coat protein, lambda N RNA binding domain, or Com protein. [Table 1]
[0092] In some examples, the mammalian expression vector comprises an sgRNA that comprises one aptamer coding sequence downstream.In other examples, the gRNA can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 aptamer coding sequences.For example, in some examples, the gRNA can comprise two aptamer coding sequences in tandem.
[0093] As used throughout, a sgRNA is a single guide RNA sequence that specifically binds or hybridizes to a target nucleic acid (genomic target sequence) in the genome of a cell such that it interacts with a CRISPR-associated endonuclease (CRISPR site-specific nuclease) and the sgRNA and CRISPR-associated endonuclease co-localize to the target nucleic acid in the genome of the cell. Each sgRNA contains a DNA targeting sequence or protospacer sequence, approximately 10-50 nucleotides in length, that specifically binds or hybridizes to a target DNA sequence in the genome. For example, the DNA targeting sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. For example, the DNA targeting sequence can be about 15-30 nucleotides, about 15-25 nucleotides, about 10-25 nucleotides, or about 18-23 nucleotides. In one example, the DNA targeting sequence is about 20 nucleotides. In some embodiments, the sgRNA comprises a crRNA sequence and a transactivating crRNA (tracrRNA) sequence. In some embodiments, the sgRNA does not comprise a tracrRNA sequence.
[0094] Generally, the DNA targeting sequence is designed to be complementary (e.g., fully complementary) or substantially complementary (e.g., with 1-4 mismatches) to the target DNA sequence. In some cases, the DNA targeting sequence can incorporate wobble or degenerate bases to bind to multiple genetic elements. In some cases, the 19 nucleotides at the 3' or 5' end of the binding region are fully complementary to the target genetic element(s). In some cases, the binding region can be modified to increase stability. For example, non-natural nucleotides can be incorporated to increase RNA resistance to degradation. In some cases, the binding region can be modified or designed to avoid or reduce the formation of secondary structures in the binding region. In some cases, the binding region can be designed to optimize GC content. In some cases, the GC content is preferably about 40% to about 60% (e.g., 40%, 45%, 50%, 55%, 60%). In some cases, the binding region can be selected to begin with a sequence that promotes efficient transcription of the sgRNA. For example, the binding region may begin at the 5' end with a G nucleotide. In some cases, the binding region may include modified nucleotides, such as, but not limited to, methylated or phosphorylated nucleotides.
[0095] As used herein, the term "complementary" or "complementarity" refers to base pairing between nucleotides or nucleic acids, such as, but not limited to, base pairing between an sgRNA and a target sequence. Complementary nucleotides are generally A and T (or A and U), and G and C. Guide RNAs described herein can include a DNA targeting sequence that is fully complementary or substantially complementary (e.g., with 1-4 mismatches) to a sequence, e.g., a genomic sequence.
[0096] The sgRNA comprises a sgRNA constant region that interacts with or binds to a CRISPR-associated endonuclease. In the constructs provided herein, the constant region of the sgRNA can be about 75-250 nucleotides in length. In some examples, the constant region is a modified constant region that contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotide substitutions in the stem, stem loop, hairpin, region between the hairpin, and / or nexus of the constant region. In some examples, modified constant regions that have at least 80%, 85%, 90% or 95% activity compared to the activity of the native or wild-type sgRNA constant region from which the modified constant region is derived can be used in the constructs described herein. In particular, modifications should not be made at nucleotides that directly interact with the CRISPR-associated endonuclease or that are critical to the secondary structure of the constant region.
[0097] The mammalian expression plasmid comprises a eukaryotic promoter operably linked to a non-viral nucleic acid sequence.In some cases, an RNA polymerase II promoter is operably linked to the nucleic acid encoding the fusion protein comprising a CRISPR-associated endonuclease and a polypeptide comprising a DNA polymerase domain; an RNA polymerase III promoter is operably linked to a gRNA coding sequence.
[0098] The RNA polymerase II promoter sequence is selected from a mammalian species. The RNA polymerase III promoter sequence is selected from a mammalian species. For example, these promoter sequences can be selected from human, bovine, ovine, buffalo, porcine, or murine, to name a few. In some examples, the RNA polymerase II promoter sequence is a CMV, FE1α, or SV40 sequence. In some examples, the RNA polymerase III promoter sequence is a U6 or H1 sequence. In some examples, the RNA polymerase II sequence is a modified RNA polymerase II sequence. For example, an RNA polymerase II sequence having at least 80%, 85%, 90%, 95%, or 99% identity to a wild-type RNA polymerase II promoter sequence from any mammalian species can be used in the constructs provided herein. In some examples, the RNA polymerase III sequence is a modified RNA polymerase III sequence. For example, RNA polymerase III sequences with at least 80%, 85%, 90%, 95% or 99% identity to wild-type RNA polymerase III promoter sequences from any mammalian species can be used in the constructs provided herein. Those skilled in the art will easily understand how to determine the identity of two polypeptides or nucleic acids. For example, identity can be calculated after aligning the two sequences so that identity is at its highest level. Another method of calculating identity can be performed by published algorithms. For example, optimal alignment of sequences for comparison can be performed using the algorithm of Needleman and Wunsch, J. Mol. Biol. 48(3):443-453 (1970). In some examples, the eukaryotic promoter is an inducible or regulatable promoter.
[0099] Coding sequences transcribed from RNA pol II promoters contain a poly(A) signal and a transcription terminator sequence downstream of the coding sequence. Commonly used mammalian terminators (SV40, hGH, BGH, and rbGlob) contain the sequence motif AAUAAA (SEQ ID NO: 31), which promotes both polyadenylation and termination. Coding sequences transcribed from RNA pol III promoters contain a simple run of T residues downstream of the coding sequence as a terminator sequence. The role of a terminator, a sequence-based element, is to define the end of a transcription unit (such as a gene) and initiate the process of releasing newly synthesized RNA from the transcription machinery. Terminators are found downstream of the transcribed gene and typically occur immediately following any 3' regulatory elements, such as a polyadenylation signal or poly(A) signal.
[0100] In some examples, the mammalian expression plasmid may also include at least one polynucleotide sequence encoding an RNA stabilizing sequence located downstream of the sequence encoding the CRISPR components, or an aptamer coding sequence if located downstream of the CRISPR components coding sequence. The polynucleotide sequence encoding the RNA stabilizing sequence is transcribed downstream of the CRISPR / Cas system components coding sequence and stabilizes the life span of the transcribed RNA sequence. In one example, the polynucleotide sequence encoding the RNA stabilizing sequence is located downstream of the catalytically impaired CRISPR-associated endonuclease coding sequence. In another example, the polynucleotide sequence encoding the RNA stabilizing sequence is located downstream of the gRNA coding sequence. An exemplary RNA stabilizing sequence is the sequence of the 3'UTR of the human beta globin gene shown in SEQ ID NO:20 (DNA) and SEQ ID NO:21 (RNA). Another example of an RNA stabilizing sequence is SEQ ID NO:22, which contains two or more copies of SEQ ID NO:20. Other RNA stabilizing sequences are described in Hayashi, T. et al., Developmental Dynamics 239(7):2034-2040(2010) and Newbury, S. et al., Cell 48(2):297-310(1987). In some examples, a spacer of 1 to 30 nucleotides may be placed between the CRISPR component coding sequence and at least one polynucleotide sequence encoding an RNA stabilizing sequence.
[0101] In some examples, the mammalian expression plasmid may comprise one or more expression cassettes. In some examples, the mammalian expression plasmid comprises a first expression cassette that encodes any of the fusion proteins described herein and a second expression cassette that encodes a gRNA that comprises at least one aptamer. In some examples, the mammalian expression plasmid may also comprise a reporter gene. system
[0102] Another aspect of the present disclosure is a lentivirus packaging system. Such systems include the mammalian expression plasmids described in the present disclosure. These systems are useful for providing components for introduction into mammalian cells to generate the lentivirus-like particles described in the present disclosure.
[0103] In some cases, the system includes a lentiviral packaging plasmid comprising a eukaryotic promoter operably linked to a viral sequence, e.g., a Gag nucleotide sequence, wherein the Gag nucleotide sequence comprises a nucleocapsid (NC) coding sequence and a matrix protein (MA) coding sequence, wherein one or both of the NC coding sequence or the MA coding sequence comprises at least one non-viral aptamer binding protein (ABP) nucleotide sequence, and wherein the packaging plasmid does not encode a functional integrase protein.
[0104] For example, a lentiviral packaging system comprising: Provided herein is a lentiviral packaging system comprising: (a) a packaging plasmid comprising a eukaryotic promoter operably linked to a Gag nucleotide sequence, the Gag nucleotide sequence comprising a nucleocapsid (NC) coding sequence and a matrix protein (MA) coding sequence, wherein one or both of the NC coding sequence or the MA coding sequence comprises at least one non-viral aptamer binding protein (ABP) nucleotide sequence, and wherein the packaging plasmid does not encode a functional integrase protein; (b) at least one mammalian expression plasmid comprising (i) a nucleic acid sequence encoding a fusion protein comprising a CRISPR-associated endonuclease and a DNA polymerase domain; and (ii) a gRNA as described herein; and (c) an envelope plasmid comprising an envelope glycoprotein coding sequence.
[0105] The system may include a second generation or third generation packaging plasmid or modified versions thereof. In some cases, the packaging plasmid includes the Gag nucleotide sequence described above, and further includes a Rev nucleotide sequence and a Tat nucleotide sequence. In other cases, the system includes a first packaging plasmid that includes the Gag nucleotide sequence described above and a second packaging plasmid that includes a Rev nucleotide sequence. In each of the packaging plasmids, the viral protein coding sequences are operably linked to a eukaryotic promoter, for example, each individually or one promoter for multiple protein coding sequences. The system may include a second generation or third generation packaging plasmid or modified versions thereof.
[0106] In some cases, the ABP coding sequence is at the 5' or 3' end of the viral protein coding sequence, i.e., at the 5' or 3' end of the NC or MA coding sequence. In some cases, the ABP coding sequence can be inserted into the viral protein coding sequence so that the encoded ABP is fused to the viral protein. The ABP coding sequence can be inserted in frame at an internal position within the viral protein coding sequence. When located in frame at an internal position near the 5' or 3' end of the viral protein coding sequence, the ABP coding sequence is positioned so as not to disrupt the processing sequence as described in Tritch, RJ et al., J.Virol.65(2):922-30(1991) and Scarlata, S. and Carter, C., Biochimica et Biophysica Acta-Biomembranes 1614(1):62-72(2003), the entire contents of which are incorporated herein by reference. For example, the Gag nucleotide sequence encodes, among others, the NC and MA coding sequences, and the Gag precursor protein is processed into separate mature viral proteins by proteolytic cleavage. In-frame insertion of the ABP coding sequence does not destroy the nucleotides encoding the processing sequences for proteolytic cleavage. In some cases, nucleotides in the viral protein coding sequence may be replaced with the ABP protein coding sequence. In some cases, a linker sequence encoding 3-6 amino acids may be located between the viral protein coding sequence and the ABP coding sequence or adjacent to the ABP coding sequence to help promote proper folding of the protein domains upon expression.
[0107] In one example, the modified viral protein is NC, and the ABP coding sequence is inserted at the 5' or 3' end of the NC coding sequence. In another example, the modified viral protein is NC, and the ABP coding sequence is inserted before or after one of the zinc finger (ZF) domains. For example, the ABP coding sequence can be inserted after the last codon of the second ZF (ZF2) domain. In another example, the ABP coding sequence can be inserted before the first codon of the ZF2 domain. In another example, the ABP coding sequence can be inserted before the first codon of the first ZF (ZF1) domain. In another example, the ABP coding sequence can be inserted after the last codon of the first ZF (ZF1) domain. In some cases, the ABP coding sequence is inserted into the NC coding sequence in a manner that does not destroy the highly positive range of amino acids in the NC protein.
[0108] In another example, the modified viral protein is MA and the ABP coding sequence is inserted at the 5' or 3' end of the MA coding sequence. In another example, the ABP coding sequence is inserted in frame at an internal location within the MA coding sequence. In some cases, nucleotides in the MA coding sequence may be replaced with the ABP protein coding sequence. For example, nucleotides encoding amino acids 44-132 of the MA protein may be replaced with the ABP coding sequence. In another example, the ABP coding sequence is inserted before the codon encoding amino acid 44 of the MA protein. In another example, the ABP coding sequence is inserted after the codon encoding amino acid 132 of the MA protein.
[0109] In some cases, the system includes a packaging plasmid that includes a eukaryotic promoter operably linked to a NEF coding sequence or a VPR coding sequence, and the NEF coding sequence or the VPR coding sequence includes at least one non-viral ABP nucleotide sequence. The system may include a second generation packaging plasmid or a third generation packaging plasmid or a modified version thereof. In some cases, the packaging plasmid includes a Gag nucleotide sequence, a Rev nucleotide sequence and a Tat nucleotide sequence. In other cases, the system includes a first packaging plasmid that includes a Gag nucleotide sequence and a second packaging plasmid that includes a Rev nucleotide sequence.
[0110] In some cases, the modified viral protein is VPR and the ABP coding sequence is inserted at the 5' or 3' end of the VPR coding sequence. In one embodiment, the ABP coding sequence is inserted at the 5' end of the VPR coding sequence.
[0111] In other cases, the modified viral protein is NEF and the ABP coding sequence is inserted at the 5' or 3' end of the NEF coding sequence. In one example, the ABP coding sequence is inserted at the 3' end of the NEF coding sequence.
[0112] In some cases, the coding sequence of the viral protein can be any one of SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, or SEQ ID NO:25. In some cases, the amino acid sequence of the viral protein can be any one of SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:24, or SEQ ID NO:26. In some cases, the lentiviral packaging plasmid comprises a sequence encoding at least one of SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:24, or SEQ ID NO:26 operably linked to a eukaryotic promoter. In some cases, when the viral protein is NEF, the polypeptide can comprise three mutations that enhance packaging within the viral capsid, such as, for example, the following substitution mutations: G3C, V153L, and E177G.
[0113] In some cases, the plasmid may encode one or more viral proteins, including two or more aptamer-binding proteins fused thereto. In some cases, the Gag nucleotide sequence of the lentiviral packaging plasmid may include an NC coding sequence and an MA coding sequence, and one or both of the NC coding sequence or the MA coding sequence include a first non-viral ABP nucleotide sequence and a second non-viral ABP nucleotide sequence. The first non-viral ABP nucleotide sequence and the second non-viral ABP nucleotide sequence may both encode the same ABP. Alternatively, the first non-viral ABP nucleotide sequence and the second non-viral ABP nucleotide sequence encode different ABPs. In some cases, the Gag nucleotide sequence of the lentiviral packaging plasmid may include an NC coding sequence that includes at least one first non-viral ABP nucleotide sequence, and an MA coding sequence that includes at least one second non-viral ABP nucleotide sequence. The at least one first non-viral ABP nucleotide sequence and the at least one second non-viral ABP nucleotide sequence may both encode the same ABP. Alternatively, the at least one first non-viral ABP nucleotide sequence and the at least one second non-viral ABP nucleotide sequence encode different ABPs.
[0114] In some cases, the packaging plasmid may encode a VPR or NEF coding sequence, and the VPR or NEF coding sequence includes a first non-viral ABP nucleotide sequence and a second non-viral ABP nucleotide sequence. The first non-viral ABP nucleotide sequence and the second non-viral ABP nucleotide sequence may both encode the same ABP. Alternatively, the first non-viral ABP nucleotide sequence and the second non-viral ABP nucleotide sequence encode different ABPs.
[0115] Non-viral aptamer binding protein (ABP) nucleotide sequences encode polypeptide sequences that bind to RNA aptamer sequences. Several non-viral ABPs are suitable for use in the present disclosure. In particular, suitable ABPs include bacteriophage RNA binding proteins that specifically bind to RNA sequences that form a stem-loop structure called an RNA aptamer sequence. Exemplary non-viral aptamer binding proteins include MS2 coat protein, PP7 coat protein, lambda N peptide, and Com (Control of mom) protein. The lambda N peptide can be amino acids 1-22 of the lambda N protein, which is the RNA binding domain of the protein. In some cases, ABPs bind to their aptamers as dimers. Information regarding these ABPs and the aptamer sequences to which they bind is provided in Table 1. In some embodiments, at least one non-viral ABP nucleotide sequence encodes a polypeptide having a sequence set forth in any of SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, or SEQ ID NO:16. In some embodiments, at least one non-viral ABP nucleotide sequence comprises any of SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:13, or SEQ ID NO:15.
[0116] A characteristic of the lentiviral packaging plasmids provided herein is that they cannot code for a functional integrase protein. When packaging plasmids do not code for a functional integrase protein and they are used in the systems and methods described herein, there is a substantially reduced risk that the nucleic acid molecules carried by the lentivirus-like particles produced using these packaging plasmids will be integrated into the genome of the transduced eukaryotic cell. In some cases, the lentiviral packaging plasmids contain an integrase coding sequence with an integrase inactivating mutation therein. For example, the integrase inactivating mutation can be an aspartic acid to valine mutation (D64V) at amino acid position 64 of the integrase protein encoded by the integrase coding sequence. In some cases, the lentiviral packaging plasmids contain a deletion of all or part of the integrase coding sequence.
[0117] In some embodiments, the lentiviral packaging plasmid comprises a eukaryotic promoter operably linked to the Gag nucleotide sequence. In some embodiments, the mammalian expression plasmid comprises a eukaryotic promoter operably linked to the VPR coding sequence or the NEF coding sequence. In some cases, the eukaryotic promoter is an RNA polymerase II promoter. The RNA polymerase II promoter sequence is selected from a mammalian species. For example, the promoter sequence can be selected from human, bovine, ovine, buffalo, porcine, or murine, to name a few. In some examples, the RNA polymerase II promoter sequence is a CMV, FE1α, or SV40 sequence. In some examples, the RNA polymerase II sequence is a modified RNA polymerase II sequence. For example, an RNA polymerase II sequence having at least 80%, 85%, 90%, 95%, or 99% identity to a wild-type RNA polymerase II promoter sequence from any mammalian species can be used in the constructs provided herein. Those skilled in the art will readily understand how to determine the identity of the above two polypeptides or nucleic acids.
[0118] Coding sequences transcribed from an RNA pol II promoter contain a poly(A) signal and a transcription terminator sequence downstream of the coding sequence. Commonly used mammalian terminators (e.g., SV40, hGH, BGH, and rbGlob) contain the sequence motif AAUAAA, which promotes both polyadenylation and termination. The role of a terminator, a sequence-based element, is to define the end of a transcription unit (such as a gene) and initiate the process of releasing newly synthesized RNA from the transcription machinery. Terminators are found downstream of the gene being transcribed, typically occurring immediately following any 3' regulatory elements such as a polyadenylation signal or poly(A) signal.
[0119] In some cases, the lentiviral packaging plasmid may contain one or more expression cassettes.
[0120] The system may also include an envelope plasmid having an envelope coding sequence that encodes a viral envelope glycoprotein. For example, the Env nucleotide sequence may code for VSV-G. The envelope coding sequence is operably linked to a eukaryotic promoter. Suitable eukaryotic promoters are described above. In some cases, the eukaryotic promoter is an RNA pol II promoter.
[0121] The system may include any of a packaging plasmid, an envelope plasmid, and a mammalian expression plasmid, i.e., (i) a nucleic acid sequence encoding a fusion protein comprising a polypeptide comprising a CRISPR-associated endonuclease and a DNA polymerase domain; and (ii) a mammalian expression plasmid comprising a gRNA comprising at least one aptamer as described herein. When any of the packaging plasmid, mammalian expression plasmid, and envelope plasmid described herein is delivered as a system to a eukaryotic cell, the gRNA expressed by the mammalian expression plasmid forms a complex with the CRISPR-associated endonuclease expressed by the mammalian expression plasmid to form an RNP that is packaged by the viral particle produced by the eukaryotic cell through the interaction between the aptamer fused or linked to the gRNA and the ABP linked to the viral protein expressed by the packaging plasmid.
[0122] Also provided herein are kits that contain components of the systems described in this disclosure. In some embodiments, the kits contain one or more of the plasmids described herein. Lentivirus-like particles
[0123] In another aspect, a lentivirus-like particle is provided, for example, a lentivirus-like particle produced by any of the methods described herein. As used herein, a lentivirus-like particle is a multiprotein structure that mimics the organization and conformation of an authentic native virus but lacks a viral genome. A plurality of lentivirus-like particles are also provided. The lentivirus-like particle contains modified lentivirus proteins that are fusion proteins in which at least one aptamer-binding protein is fused to one or more viral proteins. In the context of the present disclosure, modified viral proteins can be structural or nonstructural. Exemplary structural proteins are lentivirus nucleocapsid (NC) protein and matrix (MA) protein. Exemplary nonstructural proteins are viral protein R (VPR) and negative regulatory factor (NEF). In some cases, the particle contains a fusion protein that includes NC protein and MA protein, one or both of which are fused to at least one nonviral aptamer-binding protein (ABP). The NC protein of the particle can have two functional zinc finger protein domains. In particular, retention of the second NC zinc finger domain may maintain the efficiency of viral assembly and budding. In some cases, the particle contains a fusion protein comprising a VPR protein or a NEF protein, and the VPR protein or the NEF protein is fused to at least one non-viral ABP. The particle also contains (i) a fusion protein comprising a CRISPR-associated endonuclease and a polypeptide comprising a DNA polymerase domain; and (ii) an RNP comprising a gRNA. Any of the mammalian expression plasmids described herein that contain a non-viral nucleic acid sequence in which at least one aptamer is bound or inserted into the gRNA sequence can be used to generate lentivirus-like particles containing RNP. In some cases, the lentivirus-like particle does not contain a functional integrase protein. These virus-like particles are useful for transducing eukaryotic cells of interest.
[0124] The particle may comprise a viral fusion protein comprising one or more ABPs. In some cases, the particle contains an NC protein, an MA protein, or both, and one or both of the NC protein or MA protein is fused to one or more non-viral ABPs. In some cases, the lentivirus-like particle comprises an NC protein fused to at least one non-viral ABP. In some cases, the lentivirus-like particle comprises an MA protein fused to at least one non-viral ABP. In some cases, the lentivirus-like particle may comprise an NC protein and an MA protein, and one or both of the NC protein or MA protein may be fused to two non-viral ABP proteins, a first non-viral ABP and a second non-viral ABP fused to the C'-terminus of the first non-viral ABP (i.e., in tandem). In some cases, the particle may comprise one or both of the NC protein or MA protein fused to a first non-viral ABP and a second non-viral ABP.
[0125] In some cases, the lentivirus-like particle contains a VPR protein or a NEF protein, and the VPR protein or the NEF protein is fused to one or more non-viral ABPs. In some cases, the lentivirus-like particle contains a VPR protein or a NEF protein fused to two non-viral ABPs, a first non-viral ABP and a second non-viral ABP fused to the C' terminus of the first non-viral ABP (i.e., in tandem). In some cases, the lentivirus-like particle contains a VPR protein or a NEF protein fused to a first non-viral ABP and a second non-viral ABP. The first non-viral ABP and the second non-viral ABP can both be the same ABP. Alternatively, the first non-viral ABP and the second non-viral ABP can be different ABPs. In some cases, the lentivirus-like particle can include an NC protein with at least one first non-viral ABP fused to an MA protein and at least one second non-viral ABP fused to its C' terminus. The at least one first non-viral ABP and the at least one second non-viral ABP are both the same ABP. Alternatively, the at least one first non-viral ABP protein and the at least one second non-viral ABP may be different ABPs. The first non-viral ABP and the second non-viral ABP may both be the same ABP. Alternatively, the first non-viral ABP and the second non-viral ABP may be different ABPs.
[0126] A non-viral ABP is a polypeptide sequence that binds to an RNA aptamer sequence. Several non-viral ABPs are suitable for use in the present disclosure. In particular, suitable ABPs include bacteriophage RNA binding proteins that specifically bind to known RNA aptamer sequences, which are RNA sequences that form a stem-loop structure. Exemplary non-viral aptamer binding proteins include MS2 coat protein, PP7 coat protein, lambda N peptide, and Com (Control of mom) protein. The lambda N peptide can be amino acids 1-22 of the lambda N protein, which is the RNA binding domain of the protein. Information regarding these ABPs and the aptamer sequences to which they bind is provided in Table 1 above.
[0127] A lentivirus-like particle may contain various lentivirus proteins. However, in some cases, a lentivirus-like particle does not contain all of the types of proteins or nucleic acids found in natural lentiviruses. In some cases, a particle may contain NC, MA, CA, SP1, SP2, P6, POL, ENV, TAT, REV, VIF, VPU, VPR and / or NEF proteins, or any derivative, combination or portion thereof. In some cases, a particle may contain NC, MA, CA, SP1, SP2, P6 and POL. In some cases, a lentivirus-like particle may only contain proteins that form the viral shell (capsid). In some cases, one or more lentivirus proteins may be completely or partially excluded from a lentivirus-like particle. For example, in some cases, a lentivirus-like particle may not contain a POL protein, or may contain a non-functional version of a POL protein, such as a POL protein with an inactivating point mutation or an inactivating truncation. In another example, the lentivirus-like particle may not contain an integrase protein or may include a non-functional version of the integrase protein, such as an integrase protein with an inactivating point mutation or inactivating truncation. For example, the lentivirus-like particle may contain a non-functional integrase protein that includes an aspartic acid to valine mutation at amino acid position 64 (D64V). In another example, the lentivirus-like particle may not contain a reverse transcriptase protein or may include a non-functional version of the reverse transcriptase protein, such as an integrase protein with an inactivating point mutation or inactivating truncation.
[0128] As mentioned above, gRNA generally comprises a DNA targeting sequence and a constant region that interacts with CRISPR-associated endonuclease. In some examples, gRNA may comprise a transactivating crRNA (tracrRNA) sequence. For example, gRNA may comprise a tracrRNA that is used with Cas9 protein or derivative. In other examples, gRNA does not comprise a tracrRNA sequence. For example, gRNA may not comprise a tracrRNA sequence that is used with Cpf1 protein or derivative.
[0129] In some examples, the gRNA includes at least one aptamer sequence. In some examples, the at least one aptamer sequence may be placed at the 5' or 3' end of the gRNA. In some examples, the at least one aptamer sequence may be inserted at an internal location within the gRNA, such as one or more of the loops formed within the folded gRNA. For example, if the gRNA is for a Cas9 protein, the at least one aptamer sequence may be placed in the tetraloop, stem loop 2 (ST2) or 3' end of the gRNA. In some examples, a spacer of 1-30 ribonucleotides may be placed between the gRNA and the at least one aptamer sequence or adjacent to the at least one aptamer sequence. In some cases, the at least one aptamer sequence does not interfere with lentivirus-like particle transduction of eukaryotic cells. For example, at least one non-viral ABP fused to one or more of the NC protein, the MA protein, the VPR protein, or the NEF protein may not interfere with lentivirus-like particle transduction of eukaryotic cells. Gene editing methods
[0130] The method of using the plasmid and system provided in the present disclosure in the CRISPR / Cas system for editing DNA target, for example gene, in the genome of eukaryotic cells is described herein.In the method provided herein, the eukaryotic cell that comprises the target genome sequence of interest to be modified is transduced with the lentivirus-like particle that contains the RNP that comprises a viral fusion protein that comprises a viral protein fused to at least one aptamer-binding protein (ABP), and (1) gRNA, and (2) a fusion protein that comprises a CRISPR-associated endonuclease and a polypeptide that comprises a DNA polymerase domain.
[0131] The method described herein can be used to edit the genome of a cell while reducing any size of on-target deletion. For example, it can reduce deletions of less than 100bp, 90bp, 80bp, 70bp, 60bp, 50bp, 40bp, 30bp, 20bp, 10bp, 9bp, 8bp, 7bp, 6bp, 5bp, 4bp, 3bp, 2bp or 1bp. In some embodiments, it can reduce large genomic deletions of more than 100bp, 200bp, 300bp, 400bp or 500bp in size. As used herein, the term "on-target deletion" refers to a deletion that occurs at or near the sgRNA target site in the genome of a cell, i.e., near the target genomic sequence of interest to be modified. In the methods provided herein, large genomic deletions can be reduced by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 90%, 95%, 99% or more compared to the number of large genomic deletions when cells are edited with a CRISPR-associated endonuclease that is not fused or linked to a polypeptide that includes a DNA polymerase domain.
[0132] In some cases, the provided methods increase non-homologous end joining (NHEJ) compared to non-NHEJ in cells edited with a CRISPR-associated endonuclease that is not fused or associated with a polypeptide comprising a DNA polymerase domain. In some embodiments, NJEH is increased compared to non-NHEJ end joining in cells. In some embodiments, non-NJEH is microhomology-mediated end joining (MMEJ) and / or single-strand annealing (SSA). In some embodiments, the increase is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, or 400% increase. In some embodiments, the increase is 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold or more increase.
[0133] In some embodiments, the ratio of NHEJ to non-NHEJ is increased in cells. In some methods, the ratio of on-target 1 base pair (1 bp) deletion to on-target deletion of more than 1 bp is increased in cells. For example, in some embodiments, the ratio of on-target 1 base pair (1 bp) deletion to on-target deletion of more than 500 base pairs is increased in cells.
[0134] In some embodiments, the number of non-templated insertions (non-TIS) is reduced in the cell. In some embodiments, the reduction is at least 10%, 20%, 30%, 40%, 50%, 60%, 80%, 90%, 95%, 99% or 100%;
[0135] In some embodiments, the number of 1 bp templated insertions (TIS) is increased in the cells. In some embodiments, the increase is at least a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, or 400% increase. In some embodiments, the increase is a 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold or more increase.
[0136] In some embodiments, the ratio of 1 bp TIS to 1 bp non-TIS is increased in cells. As used herein, "templated insertion" refers to the non-random insertion that occurs after Cas9-induced double-strand break leaves a 1 nt 5' overhang base, which is then filled in and ligated by DNA polymerase.
[0137] In some embodiments, the exonuclease activity of MRE11 double strand break repair nuclease (MRE11) is reduced (e.g., compared to wild-type MRE11). For example, in the methods provided herein, the exonuclease activity of MRE11 may be reduced by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 90%, 95%, 99% or more.
[0138] The use of lentivirus-like particles that lack integrase activity in this method reduces the risk that any of the nucleic acids carried by the particles will be integrated into the cell genome.In some cases, the lentivirus particles used lack the part of the lentivirus genome sequence that is essential for viral replication, thus reducing the risk of continued particle production.Another advantage of the provided components is that viral fusion proteins increase the packaging of RNP into lentivirus-like particles, which can increase genome editing efficiency.
[0139] In some examples, the transduced eukaryotic cells are mammalian cells. In some examples, the eukaryotic cells can be in vitro cultured cells. In some examples, the eukaryotic cells can be ex vivo cells obtained from a subject. In other examples, the eukaryotic cells are present within a subject. As used throughout, subject refers to an individual. For example, the subject is a mammal, such as a primate, more specifically a human. Non-human primates are also subjects. The term subject includes domestic animals, such as cats, dogs, livestock animals (e.g., cows, horses, pigs, sheep, goats, etc.), and laboratory animals (e.g., ferrets, chinchillas, mice, rabbits, rats, gerbils, guinea pigs, etc.). Thus, veterinary uses as well as medical uses and formulations are contemplated herein. The term does not indicate a particular age or sex. Thus, living and newborn subjects are intended to be covered, whether male or female. As used herein, patient or subject may be used interchangeably and may refer to a subject suffering from a disease or disorder. The lentivirus-like particles provided herein can be administered to a subject, e.g., injected into a subject, according to known routine methods. Exemplary modes of administration include oral, rectal, mucosal, topical, intranasal, inhalation (e.g., via aerosol), buccal (e.g., sublingual), vaginal, intrathecal, intraocular, transdermal, intradermal, intrapleural, intracerebral, and intraarticular), topical, etc., as well as direct tissue or organ injection. Administration can also be into a tumor. The most appropriate route in any given case depends on the nature and severity of the condition being treated and the nature of the particular lentivirus-like particle being used. In some cases, the lentivirus-like particles are injected intravenously (IV), intraperitoneally (IP), intramuscularly, or into a particular organ or tissue. In some embodiments, more than one administration (e.g., two, three, four or more administrations) can be used to achieve the desired level of gene editing over a variety of interval periods, e.g., daily, weekly, monthly, yearly, etc.
[0140] Effective amounts of any of the recombinant lentivirus-like particles described herein vary and can be determined by one of skill in the art through experimentation and / or clinical trials. For example, an effective dose may be about 10 6 ~about 10 15 lentivirus-like particles, e.g., about 10 6 ~about 10 14 pieces, about 10 6 ~about 10 13 pieces, about 10 6 ~about 10 12 lentivirus-like particles, approximately 10 6 ~about 10 12 pieces, about 10 6 ~about 10 11 Pieces or about 10 6 ~about 10 11 The effective dosage may be 1000 lentivirus-like particles. Other effective dosages may be easily established by those skilled in the art through routine testing to establish a dose-response curve. See, for example, Mangeot et al., "Genome editing in primary cells and in vivo using viral-derived Nanoblades loaded with Cas9-sgRNA ribonucleoproteins," Nat Commun 10, 45 (2019). doi.org / 10.1038 / s41467-018-07845-z.
[0141] In some cases, a method is provided for modifying a target locus of interest, the method comprising transducing a plurality of eukaryotic cells with a plurality of viral particles, the plurality of viral particles comprising: (i) a fusion protein comprising a viral protein, e.g., NC, MA, VRP or NEF, the viral protein comprising at least one non-viral aptamer-binding protein (ABP); and (ii) a ribonucleotide protein (RNP) complex comprising (1) a gRNA and (2) a fusion protein comprising a CRISPR-associated endonuclease and a polypeptide comprising a DNA polymerase domain, the RNP complex being capable of binding (e.g., preferentially binding) to a genomic target sequence in the genomic DNA of the cell via the gRNA, and the CRISPR-associated endonuclease altering the genomic DNA of the cell. As described above, the RNP complex is packaged into the viral particle via the interaction of an aptamer sequence bound to or inserted into the gRNA sequence that forms a complex with the CRISPR-associated endonuclease.
[0142] Generally, sgRNA targets a specific region of a gene or its vicinity.In some examples, sgRNA can target the region that requires a single base change, for example, to correct the single base mutation in the human β-globin gene that causes sickle cell anemia.SgRNA allows the RNP complex described herein to specific site in the genome sequence of a cell.
[0143] In some examples, the modifications to the system components described in this disclosure do not impair how the system components function after transduction into eukaryotic cells. Rather, the components may function similarly or better than the unmodified components when transducing into eukaryotic cells. For example, the viral fusion protein in the lentivirus-like particle may not interfere with the lentivirus-like particle transduction of eukaryotic cells. Similarly, if the RNP complex packaged in the lentivirus-like particle contains at least one aptamer sequence, the at least one aptamer sequence may not interfere with the lentivirus-like particle transduction of eukaryotic cells. In some examples, the lentivirus-like particle containing the viral fusion protein may result in greater gene editing when transducing into eukaryotic cells compared to the lentivirus-like particle that does not contain the viral fusion protein. In one illustrative example, the viral fusion protein may be a NC-ABP fusion protein, such as a NC-MS2 fusion protein or a NC-PP7 fusion protein. In one example, the NC fusion protein is fused to one or two ABPs, for example, one or two MS2 proteins, one or two PP7 proteins, or one MS2 protein and one PP7 protein.
[0144] The eukaryotic cells can be in vitro, ex vivo or in vivo. In some embodiments, the cells are primary cells (isolated from a subject). As used herein, primary cells are cells that are not transformed or immortalized. Such primary cells can be cultured, subcultured or passaged a limited number of times (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cultures). In some cases, the primary cells are adapted to in vitro culture conditions. In some cases, the primary cells are isolated from an organism, system, organ or tissue, optionally sorted, and utilized directly without culture or subculture. In some cases, the primary cells are stimulated, activated, or differentiated. In some embodiments, the cells are cultured under conditions effective to expand a population of modified cells. In some embodiments, the cells modified by any of the methods provided herein are purified. In some cases, cells are removed from the subject, modified using any of the methods described herein, and readministered to the patient.
[0145] In some examples, once cells are transduced with the viral particles, the cells are cultured for a sufficient time to allow gene editing to occur so that a pool of cells expressing a detectable phenotype can be selected from the plurality of transduced cells. The phenotype can be, for example, cell growth, survival or proliferation. In some examples, the phenotype is cell growth, survival or proliferation in the presence of an agent such as a cytotoxic agent, an oncogene, a tumor suppressor, a transcription factor, a kinase (e.g., a receptor tyrosine kinase), a gene (e.g., an exogenous gene) under the control of a promoter (e.g., a heterologous promoter), a checkpoint gene or a cell cycle regulator, a growth factor, a hormone, a DNA damaging agent, a drug, or a chemotherapeutic drug. The phenotype can also be protein expression, RNA expression, protein activity, or cell motility, migration or invasion. In some examples, selecting cells based on phenotype includes fluorescence-activated cell sorting, affinity purification of cells, or selection based on cell motility.
[0146] In some examples, selecting the cells includes analyzing the genomic DNA of the cells by amplification, sequencing, SNP analysis, etc. Sequencing methods include, but are not limited to, shotgun sequencing, bridge PCR, Sanger sequencing (including microfluidic Sanger sequencing), pyrosequencing, massively parallel signature sequencing, nanopore DNA sequencing, single molecule real-time sequencing (SMRT) (Pacific Biosciences, Menlo Park, CA), ion semiconductor sequencing, ligation sequencing, sequencing by synthesis (Illumina, San Diego, Ca), polony sequencing, 454 sequencing, solid-phase sequencing, DNA nanoball sequencing, heliscope single molecule sequencing, mass spectrometry sequencing, pyrosequencing, supported oligo-ligation detection (SOLiD) sequencing, DNA microarray sequencing, RNAP sequencing, tunneling current DNA sequencing, and any other DNA sequencing methods identified in the future. One or more of the sequencing methods described herein can be used in high-throughput sequencing methods. As used herein, the term "high-throughput sequencing" refers to any method related to sequencing of nucleic acids in which more than one nucleic acid sequence is sequenced at a given time. Treatment
[0147] Any of the methods and compositions described herein can be used to treat a disease in a subject, such as cancer, a blood disorder (e.g., sickle cell anemia or beta thalassemia), an infectious disease, an autoimmune disease, transplant rejection, graft-versus-host disease, or other inflammatory disorder. In some embodiments, the infectious disease is a viral infection, such as, but not limited to, human immunodeficiency virus (HIV), herpes simplex virus type 1 (herpetic stromal keratitis), and human papilloma virus (HPV). In some embodiments, the disease is a neurodegenerative or muscular degenerative disease, such as Duchenne muscular dystrophy.
[0148] In some methods, the cancer to be treated is selected from cancer of B-cell origin, breast cancer, cancer of the stomach, neuroblastoma, osteosarcoma, lung cancer, colon cancer, chronic myeloid cancer, leukemia (e.g., acute myeloid leukemia, chronic lymphocytic leukemia (CLL) or acute lymphocytic leukemia (ALL)), prostate cancer, colon cancer, renal cell carcinoma, liver cancer, kidney cancer, ovarian cancer, stomach cancer, testicular cancer, rhabdomyosarcoma, and Hodgkin's lymphoma. In some embodiments, the cancer of B-cell origin is selected from the group consisting of B-lineage acute lymphoblastic leukemia, B-cell chronic lymphocytic leukemia, and B-cell non-Hodgkin's lymphoma.
[0149] In some methods, the subject's cells are modified in vivo. In some methods, a method of treating a disease in a subject includes a) obtaining cells from the subject; b) modifying the cells using any of the methods provided herein; and c) administering the modified cells to the subject. See, for example, Milone and O'Doherty ''Clinical sue of lentiviral vectors,'' Leukemia 32, 1529-1541 (2018). Optionally, the disease is selected from the group consisting of cancer, blood disorders (e.g., sickle cell anemia or beta thalassemia), infectious diseases, autoimmune diseases, transplant rejection, graft-versus-host disease or other inflammatory disorders in the subject. In some methods for treating cancer, the cells obtained from the subject are modified to express a tumor-specific antigen. As used throughout, the phrase "tumor-specific antigen" refers to an antigen that is unique to cancer cells or that is more abundantly expressed in cancer cells than in non-cancerous cells. Optionally, the cells obtained from the subject are T cells. If necessary, the modified cells are expanded prior to administration to a subject.
[0150] The lentivirus-like particles or cells described herein can be formulated as a pharmaceutical composition. Thus, provided herein is a pharmaceutical composition comprising any of the lentivirus-like particles described herein. Also provided is a pharmaceutical composition comprising any of the modified cells described herein, and optionally the pharmaceutical composition can further comprise a carrier. The term "carrier" refers to a compound, composition, substance, or structure that, when combined with the lentivirus-like particle or cell, aids or facilitates the preparation, storage, administration, delivery, efficacy, selectivity, or any other characteristic of the lentivirus-like particle or cell for its intended use or purpose. For example, the carrier can be selected to minimize any degradation of the active ingredient and minimize any adverse side effects in the subject. Such pharmaceutically acceptable carriers include sterile biocompatible pharmaceutical carriers, including, but not limited to, saline, buffered saline, artificial cerebrospinal fluid, dextrose, and water. Pharmaceutically acceptable refers to a material that is not biologically or otherwise undesirable and can be administered to an individual together with a selected agent without causing unacceptable biological effects or interacting in an adverse manner with other components of the pharmaceutical composition in which it is contained. Embodiment
[0151] 1. A mammalian expression plasmid comprising a eukaryotic promoter operably linked to a non-viral nucleic acid sequence, said non-viral nucleic acid sequence comprising: (i) a nucleic acid sequence encoding a fusion protein comprising: (a) the polypeptide comprises a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) a guide RNA (gRNA) coding sequence that includes at least one aptamer coding sequence; 1. A mammalian expression plasmid comprising:
[0152] 2. The mammalian expression plasmid of embodiment 1, wherein the CRISPR-associated endonuclease coding sequence encodes a Cas9 protein.
[0153] 3. The mammalian expression plasmid of embodiment 1 or 2, wherein the polypeptide comprising the DNA polymerase domain is E. coli DNA polymerase I (DNA PolI).
[0154] 4. The mammalian expression plasmid of embodiment 1 or 2, wherein the polypeptide comprising the DNA polymerase domain is the Klenow fragment of DNA PolI.
[0155] 5. A mammalian expression plasmid according to any one of embodiments 1 to 4, wherein the polypeptide comprising the DNA polymerase domain has reduced polymerase activity.
[0156] 6. A mammalian expression plasmid according to any one of embodiments 1 to 5, wherein the at least one aptamer coding sequence encodes an aptamer sequence that is specifically bound by an ABP selected from the group consisting of MS2 coat protein, PP7 coat protein, lambda N RNA binding domain, or Com protein.
[0157] 7. A mammalian expression plasmid described in any one of embodiments 1 to 6, wherein the aptamer is an MS2 aptamer sequence or a com aptamer sequence.
[0158] 8. A mammalian expression plasmid according to any one of embodiments 1 to 7, wherein the sgRNA coding sequence comprises at least one aptamer inserted into the tetraloop or ST2 loop of the sgRNA coding sequence.
[0159] 9. The mammalian expression plasmid of embodiment 8, wherein the sgRNA code comprises at least one com aptamer inserted into the ST2 loop of the gRNA coding sequence.
[0160] 10. A lentiviral packaging system comprising: 10. A lentiviral packaging system comprising: a) a packaging plasmid comprising a eukaryotic promoter operably linked to a Gag nucleotide sequence, the Gag nucleotide sequence comprising a nucleocapsid (NC) coding sequence and a matrix protein (MA) coding sequence, wherein one or both of the NC coding sequence or the MA coding sequence comprises at least one non-viral aptamer binding protein (ABP) nucleotide sequence, and wherein the packaging plasmid does not encode a functional integrase protein; b) at least one mammalian expression plasmid according to any one of embodiments 1-9; and c) an envelope plasmid comprising an envelope glycoprotein coding sequence.
[0161] 11. The lentiviral packaging system of embodiment 10, wherein the packaging plasmid further comprises a Rev nucleotide sequence and a Tat nucleotide sequence.
[0162] 12. The lentiviral packaging system of embodiment 10 or 0, further comprising a second packaging plasmid comprising a Rev nucleotide sequence.
[0163] 13. The lentiviral packaging system of any one of embodiments 10 to 12, wherein the at least one non-viral ABP nucleotide sequence encodes an MS2 coat protein, a PP7 coat protein, a lambda N peptide, or a Com protein.
[0164] 14. A lentiviral particle comprising: A) a fusion protein comprising a nucleocapsid (NC) protein or a matrix (MA) protein, wherein the NC protein or the MA protein comprises at least one non-viral aptamer binding protein (ABP); and B) A ribonucleotide-protein (RNP) complex, (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) A lentiviral particle comprising a ribonucleotide-protein (RNP) complex comprising a guide RNA (gRNA) coding sequence, wherein the gRNA coding sequence comprises at least one aptamer coding sequence, and wherein the lentivirus-like particle does not comprise a functional integrase protein.
[0165] 15. The lentiviral particle of embodiment 14, wherein the CRISPR-associated endonuclease coding sequence encodes a Cas9 protein.
[0166] 16. A lentiviral particle described in embodiment 14 or 15, wherein the polypeptide comprising a DNA polymerase domain is Escherichia coli DNA polymerase I (DNA PolI).
[0167] 17. A lentiviral particle described in embodiment 14 or 15, wherein the polypeptide comprising a DNA polymerase domain is the Klenow fragment of DNA PolI.
[0168] 18. A lentiviral particle described in any one of embodiments 14 to 17, wherein the polypeptide comprising a DNA polymerase domain has reduced polymerase activity.
[0169] 19. A method for producing lentiviral particles, comprising: a) transfecting a plurality of eukaryotic cells with the packaging plasmid, the at least one mammalian expression plasmid and the envelope plasmid of the system according to any one of embodiments 10 to 13; and b) culturing the transfected eukaryotic cells for a period of time sufficient to allow lentiviral particles to be produced; The method includes:
[0170] 20. The method of embodiment 19, wherein the lentiviral particles comprise: (i) a nucleic acid sequence encoding a fusion protein comprising: (a) a polypeptide comprising a DNA polymerase domain; and (b) a CRISPR-associated endonuclease coding sequence; and (ii) Guide RNA The method comprises a ribonucleotide-protein (RNP) complex comprising:
[0171] 21. The method of embodiment 20, wherein the plurality of eukaryotic cells are mammalian cells.
[0172] 22. A lentiviral particle produced by a method according to any one of embodiments 19 to 21.
[0173] 23. A method for modifying a genomic target sequence in a cell, comprising transducing a plurality of eukaryotic cells with a plurality of viral particles, wherein the plurality of viral particles comprises a lentivirus-like particle described in any one of embodiments 14 to 18, wherein the RNP complex binds to the genomic target sequence in genomic DNA of the cells, and the CRISPR-associated endonuclease cleaves the genomic target sequence to generate a double-stranded break, thereby modifying the genomic target sequence.
[0174] 24. The method of embodiment 23, wherein non-homologous end joining (NHEJ) is increased in the cells compared to cells modified with a CRISPR-associated endonuclease that is not fused to a DNA polymerase domain.
[0175] 25. The method of embodiment 23, wherein the ratio of NHEJ to non-NHEJ is increased in the cell.
[0176] 26. The method of embodiment 25, wherein the non-NHEJ end joining is microhomology-mediated end joining (MMEJ) and / or single-stranded annealing (SSA).
[0177] 27. A method according to any one of embodiments 23 to 26, wherein the number of on-target deletions is reduced in the cells.
[0178] 28. The method of embodiment 27, wherein the number of on-target deletions greater than 500 base pairs in size is reduced in the cell.
[0179] 29. The method of any one of embodiments 23 to 27, wherein the ratio of on-target one base pair (1 bp) deletions to on-target deletions of greater than 1 bp is increased in the cell.
[0180] 30. The method of embodiment 29, wherein the ratio of on-target 1 base pair (1 bp) deletions to deletions of more than 500 base pairs is increased in the cell.
[0181] 31. A method according to any one of embodiments 23 to 30, wherein the number of templated insertions (TIS) is increased in the cell.
[0182] 32. The method of embodiment 31, wherein the ratio of TIS to non-TIS is increased in the cells.
[0183] 34. The method of any one of embodiments 23 to 32, wherein the plurality of eukaryotic cells are mammalian cells.
[0184] 35. The method of any one of embodiments 23 to 33, wherein the plurality of eukaryotic cells are cells present within a subject.
[0185] 36. The method of embodiment 34, wherein the subject is a human subject.
[0186] 37. The method of embodiment 35, wherein the subject is injected with the multiple viral particles.
[0187] 38. A cell containing a plasmid described in any one of embodiments 1 to 9.
[0188] 39. A cell containing the lentiviral packaging system described in any one of embodiments 10 to 13.
[0189] 40. A cell containing a lentiviral particle described in any one of embodiments 14 to 16.
[0190] 41. A cell modified using the method of any one of embodiments 23 to 36.
[0191] 42. A method for treating a disease in a subject, comprising: a) obtaining cells from said subject; b) modifying the cells of the subject using a method according to any one of embodiments 23 to 36; and c) administering the modified cells to the subject.
[0192] 43. The method of embodiment 41, wherein the disease is cancer.
[0193] 44. The method of embodiment 42, wherein the disease is Duchenne muscular dystrophy.
[0194] 45. The method of any one of embodiments 41 to 43, wherein the cell is a T cell.
[0195] All patents, patent publications, patent applications, journal articles, books, technical references, and the like, discussed in this disclosure are hereby incorporated by reference in their entirety for all purposes.
[0196] It is to be understood that the figures and descriptions of the present disclosure are simplified to show elements relevant to a clear understanding of the present disclosure. It is to be understood that the drawings are presented for illustrative purposes and are not presented as structural diagrams. The omitted details and modified or alternative embodiments are within the understanding of those skilled in the art.
[0197] It will be understood that in certain aspects of the disclosure, multiple components may be substituted for single components, and multiple components may be substituted for single components, to provide an element or structure or to perform a given function(s). Except where such substitution would not function to practice a particular embodiment of the disclosure, such substitutions are considered to be within the scope of the disclosure.
[0198] The examples presented herein are intended to illustrate potential and specific embodiments of the present disclosure. It will be understood that the examples are intended primarily to illustrate the present disclosure for those skilled in the art. There may be variations in these figures or the operations described herein without departing from the spirit of the present disclosure. For example, in certain cases, method steps or operations may be performed or executed in a different order, or operations may be added, deleted, or modified.
[0199] Where a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range is also specifically disclosed, to the smallest fraction of the unit of the lower limit, unless the context clearly dictates otherwise. Any smaller range between any stated or unstated intervening value in a stated range and any other stated or intervening value in that stated range is encompassed. The upper and lower limits of these smaller ranges may be independently included or excluded in the range, and each range in which either or both limits are included in the smaller range or neither limit is included, subject to any specifically excluded limit in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0200] Different arrangements of the components shown in the drawings or described above, as well as components and steps not shown or described, are possible. Similarly, some features and subcombinations are useful and can be employed without reference to other features and subcombinations. The embodiments of the present disclosure have been described for purposes of illustration and not limitation, and alternative embodiments will become apparent to the reader of this patent. Thus, the present disclosure is not limited to the embodiments described above or shown in the drawings, and various embodiments and modifications can be made without departing from the scope of the following claims. EXAMPLES
[0201] Working Example Escherichia coli DNA polymerase I (pol I) functions in DNA repair and replication of lagging strand chromosomal DNA. Its small N-terminal domain contains 5'-3' exonuclease activity and a large C-terminal Klenow fragment that can be generated by partial digestion of the full-length polymerase, with polymerase and 3'-5' exonuclease activity. Because MMEJ and SSA can generate large DNA deletions and both depend on DNA resection, it was hypothesized that countering DNA resection with DNA polymerase I could favor canonical NHEJ over MMEJ and SSA and reduce the generation of large deletions. Thus, E. coli pol I was targeted to DSM to counter DNA resection and reduce the likelihood of generating large deletions (Figure 1A). Studies were also performed to determine whether CRISPR / Cas9-generated 5' overhangs were filled in and the rate of TIS (Figure 1B) was increased to further refine INDELs. A. Materials and Methods
[0202] Constructs. The constructs used in this study are listed in Table 2. The envelope plasmid pMD2.G for lentiviral pseudotyping was purchased from Addgene (Addgene 12259, Watertown, MA). Some of the plasmids used in this study are available from Addgene (plasmid IDs: 176234, 76235, 176236, 176237, 176238). Others are available from the authors upon request. The sequences of the primers used in this study are listed in Table 3. The single guide RNA (sgRNA) sequences are listed in Table 4. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] [Table 2-13]
Table 2-14
Table 2-15
Table 2-16
Table 2-17
Table 2-18
Table 2-19
Table 2-20
Table 2-21
Table 3-1
Table 3-2
Table 3-3
Table 4
[0203] Cell culture. HEK293T (ATCC CRL-3216™) and recently reported HEK293T-derived CLCN5 GFP-reporter cells (Lu et al. Lentiviral Capsid-Mediated Streptococcus pyogenes Cas9 Ribonucleoprotein Delivery for Efficient and Safe Multiplex Genome Editing. CRISPR J 4(6):914-928(2021)) were cultured in DMEM containing 10% FBS, 2 mM L-glutamine, 100 U / ml penicillin and 100 μg / ml streptomycin (ThermoFisher Scientific, Waltham, MA) at 37°C in a 5% CO2 incubator. Human lung fibroblast IMR90 cells (ATCC CCL-186) were cultured in Eagle's Minimum Essential Medium supplemented with 10% FBS, 2 mM L-glutamine, 100 U / ml penicillin and 100 μg / ml streptomycin. Human CD34+ progenitor cells from mobilized peripheral blood (Lonza, Catalog No.: 4Y-101C, Basel, Switzerland) were cultured in serum-free medium (Stemcell Technology, Catalog No. 09605, Vancouver, CA) supplemented with 1× StemSpan™ CD34+ proliferation supplement (Stemcell Technology, Catalog No. 02691). Human skeletal muscle myoblasts (Lonza, Catalog No.: CC-2580) were cultured in SkGMTM-2 Skeletal Muscle Cell Growth Medium-2 BulletKit™ (Lonza, Catalog No.: CC-3245).
[0204] Transfection of HEK293T cells. HEK293T cells were transfected into 24-well plates using FuGENE HD (Promega, Cat#: E2312, Madison, WI). The day before transfection, 1.25 × 10 5Cells were seeded in 24-well plates. For DNA transfection, 0.5 μg of plasmid DNA was added to 50 μl of OPTI-MEM. In a separate tube, 1.5 μl of FuGENE HD was added to 50 μl of OPTI-MEM. The two mixtures were mixed and incubated at room temperature for 15 min before being added to the cells whose medium had been replaced with OPTI-MEM just before DNA transfection. 24 hours after transfection, the medium was replaced with normal growth medium and the cells were analyzed 72 hours after transfection.
[0205] Nucleofection of human primary cells. The Nucleofector™ 2b device (Lonza) was used for nucleofection of human primary cells. IMR90 cells, human myoblasts and human CD34+ hematopoietic cells were nucleofected with the Cell Line Nucleofector™ Kit R (Lonza, Catalog No.: VCA-1001, Program X-001), Human Dermal Fibroblast Nucleofector™ Kit (Lonza, Catalog No.: VPD-1001, Program P-022) and Human CD34+ Cell Nucleofector™ Kit (Lonza, Catalog No.: VPA-1003), respectively. The number of cells for each nucleofection was 2 × 10 5 4.5 μg of target plasmid DNA (expressing sgRNA / Cas9 or sgRNA / Cas9-Klenow) and 0.5 μg of GFP-expressing plasmid DNA (CmiR0001-MR03, GeneCopoeia, Inc., Rockville, MD) were used for each nucleofection, and the GFP-expressing plasmid DNA was used as an indicator of nucleofection efficiency. Before further experiments, the transfected cells were checked under a fluorescent microscope for a similar GFP-positive percentage.
[0206] Knockdown of MRE11 or RBBP8 in human HEK293T and IMR90 cells. MRE11 or RBBP8 was knocked down using CRISPR / Cas9 nuclease to observe its role in CLCN5 mutation profile induced by Cas9 or Cas9-Klenow. sgRNA sequences were confirmed as described in Shou et al.,''Precise and Predictable CRISPR Chromosomal Rearrangements Reveal Principles of Cas9-Mediated Nucleotide Insertion. Mol Cell, 71, 498-509 e494 (2018) and listed in Table 4. To increase the possibility of MRE11 or RBBP8 knockdown, the DNA ratio of CLCN5 sgRNA expression plasmid and MRE11 / RBBP8 sgRNA expression plasmid was 1:2. The DNA mixture was co-transfected into HEK293T cells by FuGENE HD and into IMR90 cells by nucleofection as described above.
[0207] Production of lentivirus-like particles. Lentivirus-like particles were produced as previously described (Lu et al. (2021)). The packaging plasmid pspAX2-D64V-NC-COM had the aptamer-binding protein Com inserted into the nucleocapsid protein, and the ST2 loop of the sgRNA was replaced by the com aptamer to enable packaging of Cas9 RNP into lentivirus capsids via the interaction between the aptamer com and the aptamer-binding protein Com. Briefly, 5 million HEK293T cells were seeded in 10 cm tissue culture dishes. 24 h after cell seeding, the following DNA mixture was added to 500 μl of OPTI-MEM: 7.5 μg of pspAX2-D64V-NC-COM, 7.5 μg of plasmid DNA expressing CLCN5 sgRNA and Cas9 (or Cas9-Klenow), and 3 μg of pMD2.G. Meanwhile, 500 μl of OPTI-MEM was mixed with 54 μl of 1 mg / ml polyethyleneimine (PEI, Polysciences Inc., Warrington, PA). The mixture was incubated at room temperature for 15 min before being added to the cells seeded the day before, and the medium was changed to OPTI-MEM immediately before DNA transfection. 24 hours after transfection, the medium was replaced with normal growth medium, and 48 hours after medium change, the medium containing virus-like particles was collected. The concentration of particles was quantified by p24-based ELISA (Cell Biolabs, QuickTiter™ Lentivirus Titer Kit Cat. No. VPK-107, San Diego, CA).
[0208] Transduction of lentivirus-like particles. To transduce HEK293T cells or CLCN5 GFP reporter cells with lentivirus-like particles, virus-like particles were added to the cells in the presence of 8 μg / ml polybrene. 2.5 × 10 4Up to 200 ng of p24 particles were added to the cells. 24 hours after transduction, the medium was replaced with normal growth medium. GFP expression in the CLCN5 GFP reporter cells could be detected 36 hours after transduction. Cells were harvested for DNA isolation 72 hours after transduction.
[0209] PCR amplification of target DNA for INDEL profile analysis. Genomic DNA was isolated using the DNeasy Blood & Tissue kit (Qiagen, Germantown, MD) according to the manufacturer's instructions. Primer sequences used for amplification of target DNA are shown in Table 3. The genomic DNA template input for PCR was a maximum of 0.5 μg. For samples with low DNA concentration, 0.2 μg of DNA was used. A minimum number of pre-determined cycles (25–30) was used to reduce amplification bias. Calibrated CloneAmp HiFi PCR Premix (Takara, Mountain View, USA; catalog #639298) was used for PCR.
[0210] Off-target analysis. Four potential off-targets of HBBsgRNA1 were analyzed to compare the off-target activity between Cas9 and Cas9-Klenow. These included G1-OT4, G1-OT5, HBD and Off-8. The regions of predicted off-targets were amplified with their respective specific primers (Table 3) and subjected to next-generation sequencing (NGS) analysis.
[0211] NGS and data analysis. NGS analysis was performed by Genewiz Inc. (Morrisville, NC) using their “Amplicon EZ” service. Approximately 50,000 reads were obtained per sample. After removing 3′ linker and 5′ barcode sequences, the obtained reads were subjected to online software Cas-Analyzer (Park et al., ''Cas-analyzer: an online tool for assessing genome editing results using NGS data. Bioinformatics, 33, 286-288 (2017)) and CRISPResso2 (Clement et al., CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol, 37, 224-226 (2019)) for mutation analysis. The two programs yielded similar results in most cases. However, CRISPResso2 did not perform well when the INDEL rate was smaller than 5%. Unless otherwise stated, data presented were analyzed with Cas-Analyzer.
[0212] Single Molecule Real-Time (SMRT) Sequencing (PacBio). Single Molecule Real-Time (SMRT) sequencing was performed to detect large deletions targeting the human CLCN5 gene. A region of 4862 bp was amplified by LongAmp® Hot Start Taq 2X Master Mix (New England Biolab Catalog Number: M0533, Ipswich, MA) with primers hCLCN5-F3 and hCLCN5-R3. DNA was submitted to GCB Sequencing and Genomic Technologies Shared Resource (Duke University, Durham, NC) for SMRT sequencing (sequence I). One SMRT cell was used for eight barcoded samples. Demultiplexed Circular Consensus (CCS) reads received from the sequencing center were used to search for large deletions by comparing with the reference sequence by R programming. Specifically, all sequences were searched for the presence of a 25 bp 5' sequence and a 25 bp 3' sequence (see Table 3 for sequences), allowing for 2 nucleotide differences (92% identity), respectively, to accommodate possible sequencing errors. These two regions are at the 5' and 3' ends of the sequenced DNA. Then, for each read, the distance between the two 25 bp regions was calculated. Reads with a distance different from that of the reference sequence and without an intact sgRNA target site were considered as reads with deletions. The counts and sequences of reads with deletions were enumerated.
[0213] Statistical analysis. GraphPad Prism software (V5) was used for t-tests, ANOVA and chi-square tests. P<0.05 was considered statistically significant. Means and sem (standard error of the mean) are reported. B. Results
[0214] Targeting E. coli DNA polymerase I to DNA DSBs increased 1-bp deletions and TIS and reduced large deletions induced by CRISPR / Cas9. E. coli DNA polymerase I (pol I) was targeted to DNA double-strand breaks (DSBs) by fusing it to the C-terminus of Streptococcus pyogenes Cas9 with the linker peptide used in EvolvR in E. coli and yeast (Figure 1B). To enhance fusion protein expression in human cells, the codons of the polA gene encoding pol I were optimized. The fusion protein was used to target the 5' untranslated region (5'UTR) of human chloride voltage-gated channel 5 (CLCN5), whose mutations cause a rare kidney disease called Dent's disease, in GFP reporter cells developed to sensitively detect genome editing activity (Lu et al. (2021)). These GFP reporter cells do not express GFP due to the disruption of the GFP reading frame by the insertion of the CLCN5 sgRNA target sequence between the start and second codons of the GFP coding sequence (from the CLCN5 5'UTR region). GFP is expressed only if genome editing restores the GFP reading frame by an in-frame INDEL (approximately 1 in 3 chance). Similar GFP-positive cells were observed in cells treated with CLCN5 gRNA / Cas9 and CLCN5 gRNA / Cas9-pol. Targeted deep sequencing confirmed genome editing at the target site by CLCN5 gRNA / Cas9-pol (Figure 6), demonstrating the functionality of the Cas9-pol fusion protein. Plasmid DNA expressing CLCN5 gRNA / Cas9 or CLCN5 gRNA / Cas9-pol (targeting human CLCN5 5' UTR) was transfected into HEK293T cells, and similar INDEL rates generated by Cas9 and Cas9-pol (11.6% ± 1.6%, N = 3 for Cas9; 14.5% ± 0.7%, N = 3 for Cas9-pol; p = 0.16) were observed, suggesting that the fusion does not impair Cas9 activity.
[0215] We analyzed the INDEL profiles induced by Cas9 and Cas9-pol (targeting CLCN5) and found that Cas9-pol caused a significant increase in 1 base pair (bp) deletions, a decrease in >1bp deletions (Figure 1C), and an increase in the 1bp TIS to 1bp non-TIS ratio (Figure 1D). This data supported the hypothesis that targeting polymerase I to DSBs favors the generation of short deletions over long deletions and favors TIS over non-TIS.
[0216] Targeting the polymerase to DSBs is necessary to modulate TIS but not deletion size. To test whether targeting the polymerase to DSBs is necessary for the observed effects, we generated the NLS-pol protein (Figure 2A), a deletion mutant of Cas9-pol that contains full-length pol I and a nuclear localization signal (NLS) but lacks most of the Cas9 functional domains (REC and part of RuvC, the entire HNH and part of the PAM interaction domain of Cas9). Since NLS-pol contained the NLS and linker sequences of Cas9-pol, it was expected to fold properly and be targeted to the nucleus. However, it should not have the DNA binding and nuclease activity of Cas9 due to the deletion of multiple Cas9 domains. Cas9, CLCN5 gRNA, and NLS-pol were co-expressed in HEK293T cells, and cells treated with Cas9 and NLS-pol were found to have a 1 bp deletion percentage between Cas9-treated cells and Cas9-pol-treated cells, and an increased DNA replacement rate in the region 20 bp 5' and 20 bp 3' of the predicted cleavage site for an unknown mechanism (Figure 2B). Importantly, co-expression of Cas9 with NLS-pol did not increase TIS to the level of Cas9-pol (Figure 2C).
[0217] To determine whether polymerase activity is required for the observed effects, Cas9 was transformed into pol 2 vectors in which the polymerase activity was inactivated. D705AThe polymerase activity of the Klenow fragment was inactivated. D705A , 5' exonuclease domain (5Exo), 3' exonuclease domain (3Exo), and both 5' exonuclease and 3' exonuclease domains (Exo) (Figure 2A). CLCN5 was targeted in HEK293T cells using these Cas9 fusion proteins, and it was found that all fusion proteins without DNA polymerase domain were unable to increase 1 bp deletion and decrease >1 bp deletion. All fusion proteins with DNA polymerase domain increased 1 bp deletion and decreased >1 bp deletion, regardless of whether the polymerase activity was inactivated (Figure 2B). Cas9-5exo, Cas9-3exo and Cas9-exo showed very similar mutation profiles to each other (Figure 7), so they were treated as one group (Cas9-all Exo) in Figure 2B. Thus, the polymerase domain, but not the polymerase activity, was sufficient to perturb the ratio of 1 bp and >1 bp deletions.
[0218] When TIS was analyzed, only DNA pol I and Klenow were able to increase 1 bp TIS and decrease 1 bp non-TIS (Figure 2C), consistent with the requirement of polymerase-mediated fill-in to generate TIS. The same conclusion was obtained when 2 bp TIS and 3 bp TIS were analyzed (Figure 8). Compared to pol I, Klenow showed a stronger effect on increasing 2 bp and 3 bp TIS. This may be because the Klenow fragment lacks the 5'>3' exonuclease domain present in pol I, which may remove 5' overhangs and promote the generation of deletions.
[0219] CLCN5 was similarly targeted in human primary IMR90 cells with various fusion proteins. In these cells, all fusion proteins with the Klenow domain (with or without polymerase activity) significantly increased the frequency of 1 bp deletions, but rather decreased the frequency of >1 bp deletions (Figure 2D). Again, targeting of polymerase activity to DSBs was required to increase TIS (Figure 2E).
[0220] We investigated the relationship between the mutation profile and the overall INDEL rate of targeting the CLCN5 locus by Cas9 and found that the overall INDEL rate did not affect the mutation profile (Figure 9). Thus, the observed effect of fusion of pol I or Klenow fragments to Cas9 on DNA mutation profile could not be explained by a possible effect on Cas9 cleavage activity. The most likely mechanism could be interfering with local cellular DNA repair machinery.
[0221] Overall, targeting the polymerase to DSBs was necessary to increase TIS, whereas neither fusion with Cas9 nor polymerase activity was required to increase 1-bp deletions. However, fusing the polymerase to Cas9 would still be beneficial, since it would increase the local polymerase concentration and reduce possible interference with other endogenous DSBs, especially since in many experimental settings, controlled amounts of genome editing effectors, rather than overexpressed ones, are used for genome editing.
[0222] DNA excision was involved in disruption of 1 bp versus >1 bp deletions. D705A and Klenow D705AThe observation that RBBP8 also favors the generation of 1-bp deletions over >1-bp deletions prompted us to examine whether DNA resection is involved in these effects. We therefore examined the mutational profile targeting the CLCN5 locus in cells with and without RBBP8 (expressing the CtIP protein) knockdown. We use the term knockdown rather than knockout to account for the possibility of disrupting expression from only one allele. A previously validated RBBP8 sgRNA (6) was used to mediate mutations in this gene. RBBP8 sgRNA-expressing DNA was co-transfected with CLCN5 sgRNA-expressing DNA into HEK293T cells. The sgRNA expression constructs also contained Cas9 or Cas9-Klenow expression cassettes (see Supplementary Table 1 for DNA constructs used). The INDEL rates of the co-transfected genes (RBBP8 and CLCN5) were highly similar (Supplementary Table S4), consistent with co-transfection. We observed quite different overall INDEL rates in Cas9- and Cas9-Klenow-treated cells for unknown reasons, however, this difference in INDEL rates is unlikely to affect our mutation profile analysis, as DNA mutation profiles were unrelated to the overall INDEL rates (Supplementary Figure S4).
[0223] Knockdown of RBBP8 did not affect the percentage of 1-bp deletions, but it significantly reduced the percentage of >1-bp deletions and increased the percentage of insertions when CLCN5 was targeted in HEK293T cells (Figure 3A). These results are consistent with the suppression of DNA resection, which may cause an increase in insertions. These observations are similar to those from Repair-Seq published while our paper was under revision (54). Importantly, in cells with RBBP8 knockdown, Klenow fragments reduced >1-bp deletions to a smaller extent than cells without RBBP8 knockdown, suggesting that RBBP8 is involved in the Klenow-mediated reduction of >1-bp deletions. The 1-bp TIS was not significantly changed in RBBP8 knockdown HEK293T cells, and Klenow fragments increased the 1-bp TIS (Figure 3B).
[0224] We also examined the CLCN5 mutation profile in IMR90 cells with and without RBBP8 knockdown. Knockdown of RBBP8 in IMR90 cells increased Cas9-induced 1-bp deletions, an effect similar to fusion of Klenow fragments to Cas9 in normal IMR90 cells (Figure 10). Doing so also increased 1-bp TISs and decreased 1-bp non-TISs. In MRE11 or RBBP8 knockdown cells, Klenow had no effect on 1-bp deletions, >1-bp deletions, or TISs. These observations suggest that knockdown of MRE11 or RBBP8 and fusion of Klenow to Cas9 may disrupt a similar pathway in IMR90 cells. Taken together, these data suggest that preventing DNA resection is at least one of the mechanisms of the effect of pol I or Klenow fragments on Cas9 DNA mutation profile.
[0225] For 1-30 bp deletions targeting CLCN5 in HEK293T cells, 11 bp and 8 bp deletions were most frequently observed, except for the 1 bp deletion (Figure 3C). The frequency of 11 bp deletions was not affected by RBBP8 knockdown (Figure 3C), suggesting that other mechanisms may also be involved in the generation of these deletions. These deletions were largely suppressed by fusing the Klenow fragment to Cas9. The frequency of the 8 bp deletion was reduced by RBBP8 knockdown and was not further reduced by the Klenow fragment. This observation further supports the notion that one way that the Klenow fragment affects the Cas9 DNA mutation profile is by interfering with DNA resection.
[0226] When examining the most frequently deleted sequences in Cas9-treated cells, microhomology was observed around the deletion (Figure 3D, underlined). The most frequently observed 11-bp deletion had microhomology (underlined green) at the extreme end of the predicted cleavage site (dashed line). In this case, DNA synthesis could be initiated without the need for a 3'-flap endonuclease such as XPF-ERCC1 (55,56) to remove the mismatched 3' flap during MMEJ. This likely explains why the 11-bp deletion was most frequently found in Cas9-treated cells. RBBP8 knockdown did not affect the frequency of this deletion, and there may be other unknown proteins that generate similar deletions under resection inhibition. In general, common deletions had microhomology around the deletion, one end at the predicted cleavage site, or both.
[0227] Cas9-Klenow increased 1-bp deletions on multiple loci in multiple human cell types. We investigated whether the effects of pol I or Klenow on DNA mutation profiles were target sequence-specific or cell type-specific. The Cas9-Klenow fusion protein was used in subsequent experiments given its smaller size and pronounced effect on increasing 1-bp deletions and TIS. Four more loci were tested, including Duchenne muscular dystrophy (DMD) exon 53 and DMD exon 44 (537 kb away from each other), the 5' coding region of HBB, and the intergenic locus intragenic 1 (GRCh38.p13, chromosome 20, 32752960-32752979). DMD exons 53 and 44 were selected because targeting these exons with single-cutting sgRNAs could restore dystrophin in DMD patients caused by exon deletions. The HBB5' coding region was selected for its potential application of genome editing in the treatment of sickle cell disease. Furthermore, this region was targeted with CRISPR / Cas9 to examine Cas9-induced gene conversion in human somatic cells (Parsijani et al., "CRISPR / Cas9 increases mitotic gene conversion in human cells," Gene Ther, 27, 281-296 (2020)). To exclude the possible contribution of the targeted gene product to the observed effects, the intragenic 1 was selected. In addition to HEK293T cells, various loci were targeted in human primary fibroblast IMR90 cells, human CD34+ hematopoietic stem cells or human primary myoblasts. A total of eight loci / cells were examined (Table 5). [Table 5] * , ** , ***indicates p<0.05, 0.01 and 0.0001, respectively, in a two-tailed t-test. Blue, red and black indicate increase, decrease and no change, respectively, in the Cas9-Klenow group. Cas9-Kle:Cas9-Klenow. Klenow fusion significantly altered the total insertion percentage at five loci / cell, so 1bp TIS was expressed as % of all insertions.
[0228] In all cases, targeting Klenow fragments to DSBs significantly increased 1-bp deletions (Table 5), increasing from an average of 18.80% ± 4.62% (N = 8) to 41.29% ± 7.39% (N = 8) across the eight loci / cells analyzed. In HBB / IMR90, 81.2% ± 5.47% of all INDELs generated by Cas9-Klenow were 1-bp deletions. The increase in 1bp deletions was accompanied by four possible phenomena: 1) a decrease in >1bp deletions (DMD44 / myoblasts, HBB / hematopoietic cells and HBB / IMR90); 2) a decrease in insertions (CLCN5 / IMR90, intragenic 1 / IMR90); 3) a decrease in both >1bp deletions and insertions (HBB / 293T); and 4) a decrease in >1bp deletions and an increase in total insertions (CLCN5 / 293T, DMD53 / 293T). The effect of Klenow on insertions was variable but caused a significant decrease in >1bp deletions in six of eight loci / cells. Targeting of the HBB locus in IMR90, HEK293T and hematopoietic cells showed different ways to explain the increase in 1bp deletions.
[0229] In all loci / cells except CLCN5 / IMR90, which did not have a clear deletion peak (Figure 11), Cas9-Klenow most significantly reduced the percentage of >1bp deletions with microhomology at the predicted cleavage site (underlined with green line), which was typically the highest >1bp deletion peak generated by Cas9 (Figure 4, Figure 11). Smaller >1bp deletion peaks did not have microhomology but were located at the predicted cleavage site (dashed vertical line) or had microhomology away from the cleavage site (underlined with red line). These observations suggest that these MMEJ events that do not require removal of the 3' flap were most sensitive to suppression by Klenow.
[0230] Cas9-Klenow increased TIS on multiple loci in multiple human cell types. We investigated the effect of Klenow fragments on TIS in these loci / cells. We focused on 1-bp TIS because 2- and 3-bp TIS were scarce in the ~50,000 reads analyzed in most loci / cells. In some cases, Klenow significantly altered the overall insertion percentage, so we compared the percentage of TIS in all insertions rather than all INDELs. With the exception of one locus / cell (HBB / IMR90) that had too few insertions to be analyzed, six of the seven loci / cells showed a significant increase in TIS occupancy in all insertions (Table 5). Only HBB / 293T showed a similar TIS percentage between Cas9 and Cas9-klenow, with ~80% of 1-bp insertions being TIS. This observation suggests that in the case of HBB / 293T, the endogenous polymerase was highly efficient at filling in the 5′ overhangs, explaining why Cas9-Klenow did not further increase TIS occupancy.
[0231] Fusing the Klenow fragment to Cas9 significantly reduced total insertions at three loci / cells (CLCN5 / IMR90, HBB / 293T, and intragenic 1 / IMR90). In two of the three loci / cells (CLCN5 / IMR90 and intragenic 1 / IMR90), non-TIS insertions contributed to 100% of the insertion reduction.
[0232] Fusing Klenow to Cas9 reduced CRISPR / Cas9-induced large DNA deletions. Targeting DNA polymerase to DSBs was observed to favor 1-bp deletions over >1-bp deletions, prompting us to examine whether doing so could reduce the generation of large (>500bp) on-target deletions. We analyzed large deletions targeting CLCN5 in HEK293T cells. We designed a pair of primers (Figure 4) that amplifies a 4862bp region with an sgRNA target sequence in the center of the amplicon. After generating lentivirus-like particles containing Cas9 RNP or Cas9-Klenow RNP and treating CLCN5 GFP reporter cells with the two types of RNPs, the percentage of GFP-positive cells was similar, suggesting that comparable genome editing activity was observed.
[0233] Then, HEK293T cells were treated with Cas9 RNP or Cas9-Klenow RNP. 72 hours after treatment, DNA was amplified and single molecule real-time (SMRT) sequencing (PacBio) was performed. More reads with deletions of >0.5kb, >1kb and >2kb were found in Cas9-treated cells than in Cas9-Klenow-treated cells (Table 6 and Figure 5). [Table 6] a Only CCSs containing both the 5' and 3' index sequences (boxes in Figure 5) were analyzed. CCS: Circular Consensus
[0234] CLCN5 was targeted in IMR90 cells by nucleofection of plasmid DNA expressing CLCN5 sgRNA / Cas9 or CLCN5 sgRNA / Cas9-Klenow. Although Cas9-treated cells had a lower short INDEL rate than Cas9-Klenow-treated cells (see Table 7, which used the same DNA samples for NGS analysis), they had more deletions of >0.5kb, >1kb, and >2kb compared to Cas9-Klenow-treated cells. [Table 7] * and *** indicates p<0.05 and p<0.0001, respectively, in the t-test (n=3).
[0235] When CCS reads with >0.5 kb deletions were examined, the deleted regions were found to span or involve the sgRNA target site (Figure 5), confirming that they were on-target deletions. Only limited types of deletions were observed. Cas9 generated more types of deletions and more CCS reads for each type of deletion than Cas9-Klenow. Both phenomena could be explained by Cas9 being more prone to generate larger deletions than Cas9-Klenow. Multiple CCS reads for the same type of deletion could be the result of a single deletion event in one cell or multiple deletion events in multiple cells. Our observations of the same deletion type in Cas9-treated and Cas9-Klenow-treated IMR90 cells (Figure 5B) * The results of the 14C-T cell lineages (indicated by ) supported the possibility of multiple deletion events for the same deletion type. Alternatively, Cas9 was able to induce larger deletions earlier after treatment than Cas9-Klenow, which increased the representation of deletions.
[0236] Cas9-Klenow increased the INDEL rate in human primary cells. Meanwhile, in HEK293T cells, Cas9 and Cas9-Klenow (or Cas9-pol) produced similar levels of INDEL rate, but in human primary cells, Cas9-Klenow produced significantly higher INDEL rate in 5 out of 6 loci / cells (Table 7). In general, the INDEL rate was relatively low in primary cells, and the reason was unclear. The observed INDEL was confirmed as a bona fide INDEL, since the background INDEL in cells treated with non-targeting sgRNA was very low and all INDEL was around the predicted cleavage site. 1 / 10 GFP-expressing plasmid DNA was included in the nucleofection experiment, and >50% GFP-positive cells were observed. Therefore, the low INDEL rate was not due to low nucleofection efficiency. Cell death was observed during the 72-h culture period after nucleofection. As observed in human ES cells, it is possible that many Cas9 or Cas9-Klenow positive cells may be lost due to the toxic effects of constantly generating DSBs (Happaniemie et al., ``CRISPR-Cas9 genome editing induces a p53-mediated DNA damage response,'' Nat Med, 24, 927-930 (2018)).
[0237] Cas9-Klenow did not increase DNA replacement rates or off-targets. We investigated whether targeting DNA polymerase to DSBs could increase DNA mutation rates near DSBs. We examined DNA replacement rates in a 40-bp region (20 bp on each side) around the predicted cleavage site for the following reasons: 1) MRE11 nicks 15–20 nt away from the DSB (18), 2) pol I has shown a processivity of 15–20 nucleotides (61), and 3) EvolR, a fusion protein between Cas9 nickase and error-prone DNA pol I, showed a mutation window of 15–20 bp (47). Cas9-Klenow did not cause an increase in DNA replacement rates (Table 4). Furthermore, a similar DNA replacement rate was also observed in the negative control with a non-targeting sgRNA, suggesting that the observed DNA replacements were mainly derived from cellular heterogeneity, PCR and sequencing errors. Thus, targeting DNA polymerase to DSBs did not increase DNA replacement rates.
[0238] We also tested the effect of fusion of the Klenow fragment to Cas9 on possible off-targets of Cas9. We targeted the HBB5' coding region in human IMR90 cells and detected INDEL rates at four potential off-targets predicted based on sequence similarity (Table 8; PAMs are underlined; italicized nucleotides indicate mismatches). No off-targets were observed in Cas9 and Cas9-Klenow treated cells using targeted NGS (<0.5%, detection limit of NGS), currently one of the most sensitive off-target detection methods. Thus, the data showed that fusion of Klenow to Cas9 does not increase off-targets to levels detectable by NGS. Because the Klenow fragment did not cause detectable off-targets on sequences similar to the bona fide target, it is unlikely to cause random off-targets on sequences not similar to the bona fide target. [Table 8]
[0239] The effect of Klenow on off-targets was examined in a third way. HEK293T-derived GFP reporter cells were used to detect CRISPR / Cas9-induced INDELs at the HBB sickle cell mutation sequence, which has a one nucleotide difference with the HBB sgRNA and is an "off-target" of the HBB sgRNA. These cells contain the HBB sgRNA bona fide target in the endogenous HBB gene and an off-target in an integrated GFP-reporter cassette. We targeted the endogenous HBB gene with a perfectly matched HBB sgRNA and examined the INDEL rate at the sickle cell mutation sequence as an off-target. Analysis showed that Cas9 and Cas9-Klenow had similar INDEL rates against the endogenous HBB target (21.47% ± 0.80% for Cas9, N = 3; 21.70% ± 1.10% for Cas9-Klenow, N = 3, p = 0.8723) and the integrated sickle cell mutation sequence (10.80% ± 0.80% for Cas9, N = 3; 12.73% ± 1.20% for Cas9-Klenow, N = 3, p = 0.2509). Collectively, these experiments demonstrate that targeting Klenow fragments to DSBs did not increase off-target or DNA mutations at DSBs.
[0240] As shown herein, targeting E. coli DNA pol I or Klenow fragments to DSBs with Cas9 fusion proteins increases the ratio of small to large deletions and the ratio of TIS to non-TIS. Importantly, doing so suppressed the generation of on-target deletions greater than 500 bp. These effects were observed in all loci analyzed (8 loci / cell), including four cell types (one cell line and three primary cell types) and five target sites. The effect of reducing deletion size and increasing TIS versus non-TIS is not cell type or target site specific. In primary cells, fusing Klenow to Cas9 significantly increased the overall INDEL rate in four out of five cases.
[0241] DNA resection is required for HDR, MMEJ and SSA. The latter two alternative NHEJ DNA repair pathways generate short and long deletions. The MRE11-RAD50-NBS1 complex is responsible for initiating DNA resection, while EXO1, BLM and DNA2 are responsible for extensive resection. Attempts have been made to suppress the generation of large deletions by counteracting DNA resection. The data provided herein suggest that preventing DNA resection is one of the mechanisms of Klenow's effect on deletion size. First, it was determined that knocking down MRE11 or CtIP, proteins involved in DNA resection, increased 1bp deletions and decreased >1bp deletions. Second, the effect of Klenow on deletion size was lost under MRE11 or CtIP knockdown. These observations are consistent with reports that inhibition of the MRE11 complex causes suppression of MMEJ (Hussmann et al., Mapping the genetic landscape of DNA double-strand break repair. Cell, 184, 5653-5669 e5625 (2021)). The frequently observed 11bp deletion in CLCN5 was not affected by MRE11 or CtIP knockdown but was suppressed by Cas9-Klenow fusion, suggesting that Klenow fusion also reduces MRE11-independent deletions. Fusing Klenow to Cas9 did not reduce the frequency of the 9bp deletion when targeting DMD exon 53 in HEK293T cells. This deletion may be generated by a process that Klenow could not interfere with.
[0242] pol with inactivated polymerase activity D705A and Klenow D705A It was determined that Cas9 has an effect on deletion size similar to pol I and Klenow fragment. These proteins may interfere with DNA resection through two non-exclusive mechanisms: 1) the addition of a bulk peptide (≥629 AA) to the C-terminus of Cas9 may prevent recruitment of the DNA resection complex or regulatory proteins, and 2) the addition of a bulk peptide (≥629 AA) to the C-terminus of Cas9 may prevent recruitment of the DNA resection complex or regulatory proteins. D705Aor Klenow D705A Residual DNA binding activity of pol can interfere with DNA excision. D705A and Klenow D705A does not affect the percentage of TIS and may be useful in cases where only small deletions and large deletions need to be increased. Although counteracting DNA resection is likely to be the mechanism underlying the observations, other cellular DNA damage repair mechanisms cannot be completely excluded.
[0243] The ability of Cas9-pol and Cas9-Klenow to increase the TIS / non-TIS ratio depended on the local availability of polymerase activity. This dependency is consistent with the observation that TIS is the result of polymerase filling in the 5' overhang ends generated by Cas9. In the studies provided herein, the Klenow fragment was more active than pol I in increasing 2-bp and 3-bp TIS (Figure 8). This suggests that the 5' exonuclease domain of pol I (absent in the Klenow fragment) may compete with the polymerase domain for the 5' overhang. The former removes the 5' overhang and promotes deletion, whereas the latter fills in the 5' overhang to produce a TIS. Removing the 1 nt 5' overhang also increased the activity of pol I to increase 1 bp deletion. D705A This may be one of the mechanisms.
[0244] Cas9-pol and Cas9-Klenow only increased TIS but not non-TIS. When targeting CLCN5 or intragenic site 1 in IMR90 cells, fusing Klenow to Cas9 significantly reduced overall insertions, but only non-TIS, not TIS. In human cells, DNA polymerase μ is required to generate both TIS and non-TIS, but it is unclear whether other proteins are required to generate non-TIS. Targeting Klenow fragments to Cas9 increased TIS but not non-TIS (or in some cases only reduced non-TIS but not TIS), suggesting that TIS and non-TIS are generated through different mechanisms. Fusing pol or Klenow fragments to Cas9 promotes filling of 5' overhangs but may have an inhibitory effect recruiting proteins involved in generating non-TIS.
[0245] Cas9-Klenow significantly increased the overall INDEL rate in four out of five cases in human primary cells but not in HEK293T cells. This can be explained by several non-exclusive mechanisms: 1) it counteracts DNA resection and inhibits homologous recombination to completely repair the DNA; 2) it fills in the 5' overhang before DNA ligase can ligate the complementary 5' overhang without generating INDEL; and 3) Cas9 induces P53-mediated DNA damage stress in primary cells but not in HEK293T cells, and fusing Klenow to Cas9 reduces the stress by preventing repeated futile editing. We did not examine the effect of fusing Klenow to Cas9 on HDR. In this way, it may have an inhibitory effect on HDR, given its effect on DNA resection required for HDR. Targeting the Klenow domain to DSBs may be useful in genome editing applications that are independent of HDR.
[0246] These studies indicate that fusing the Klenow domain to RNA-guided endonucleases, such as the Cas9 endonuclease, may improve the safety and efficiency of genome editing in in vitro and in vivo applications. Fusing the Klenow domain to Cas9 reduced the generation of unpredictable large on-target DNA deletions that have been observed by multiple groups (Owens et al., "Microhomologies are prevalent at Cas9-induced larger deletions. Nucleic Acids Res., 47, 7402-7417 (2019); Adikusuma et al., "Large deletions induced by Cas9 cleavage. Nature, 560, E8-E9 (2018); and Kosicki et al., "Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nat Biotechnol, 36, 765-771 (2018). It increased genome editing efficiency in primary cells and did not increase DNA replacement or off-target rates. Furthermore, the effect on 1bp deletion and TIS can be used to increase the percentage of the desired type of mutation to improve the efficiency of disrupting disease-causing genes or restoring genes disrupted by rearrangement. array [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5]
Claims
1. A mammalian expression plasmid comprising a eukaryotic promoter operably ligated to a nonviral nucleic acid sequence, wherein the nonviral nucleic acid sequence is (i) Nucleic acid sequences encoding a fusion protein including the following: (a) polypeptides containing a DNA polymerase domain; and (b) CRISPR-related endonuclease coding sequences; and (ii) A guide RNA (gRNA) coding sequence containing at least one aptamer coding sequence. A mammalian expression plasmid containing this plasmid.
2. The mammalian expression plasmid according to claim 1, wherein the CRISPR-related endonuclease coding sequence encodes a Cas9 protein.
3. The mammalian expression plasmid according to claim 1, wherein the polypeptide containing the DNA polymerase domain is Escherichia coli DNA polymerase I (DNA PolI).
4. The mammalian expression plasmid according to claim 1, wherein the polypeptide comprising the DNA polymerase domain is a Klenow fragment of DNA PolI.
5. The mammalian expression plasmid according to claim 1, wherein the polypeptide comprising the DNA polymerase domain has reduced polymerase activity.
6. The mammalian expression plasmid according to claim 1, wherein the at least one aptamer coding sequence codes for an aptamer sequence specifically bound by an ABP selected from the group consisting of an MS2 coat protein, a PP7 coat protein, a lambda N RNA binding domain, or a Com protein.
7. The mammalian expression plasmid according to claim 1, wherein the aptamer is an MS2 aptamer sequence or a com aptamer sequence.
8. The mammalian expression plasmid according to claim 1, wherein the sgRNA coding sequence comprises at least one aptamer inserted into a tetraloop or ST2 loop of the sgRNA coding sequence.
9. The mammalian expression plasmid according to claim 8, wherein the sgRNA code comprises at least one com aptamer inserted into the ST2 loop of the gRNA coding sequence.
10. A lentivirus packaging system, a) A packaging plasmid comprising a eukaryotic promoter operably linked to a Gag nucleotide sequence, wherein the Gag nucleotide sequence comprises a nucleocapsid (NC) coding sequence and a matrix protein (MA) coding sequence, and either or both of the NC coding sequence or the MA coding sequence comprises at least one nonviral aptamer-binding protein (ABP) nucleotide sequence, and the packaging plasmid does not encode a functional integrase protein; b) at least one mammalian expression plasmid according to any one of claims 1 to 9; and c) Envelope plasmid containing an envelope glycoprotein coding sequence A lentivirus packaging system, including...
11. The lentiviral packaging system according to claim 10, wherein the packaging plasmid further comprises a Rev nucleotide sequence and a Tat nucleotide sequence.
12. The lentiviral packaging system according to claim 10, further comprising a second packaging plasmid containing a Rev nucleotide sequence.
13. The lentiviral packaging system according to claim 10, wherein the at least one nonviral ABP nucleotide sequence encodes an MS2 coated protein, a PP7 coated protein, a lambda N peptide, or a Com protein.
14. It is a lentivirus particle, A) A fusion protein comprising a nucleocapsid (NC) protein or a matrix (MA) protein, wherein the NC protein or MA protein comprises at least one nonviral aptamer-binding protein (ABP); and B) Ribonucleotide protein (RNP) complex, (i) Nucleic acid sequences encoding a fusion protein including the following: (a) polypeptides containing a DNA polymerase domain; and (b) CRISPR-related endonuclease coding sequences; and (ii) A ribonucleotide protein (RNP) complex containing a guide RNA (gRNA) coding sequence. A lentiviral particle comprising, wherein the gRNA coding sequence comprises at least one aptamer coding sequence, and the lentiviral-like particle does not contain a functional integrase protein.
15. The lentiviral particle according to claim 14, wherein the CRISPR-related endonuclease coding sequence encodes the Cas9 protein.
16. The lentiviral particle according to claim 14, wherein the polypeptide containing the DNA polymerase domain is Escherichia coli DNA polymerase I (DNA PolI).
17. The lentiviral particle according to claim 14, wherein the polypeptide containing the DNA polymerase domain is a Klenow fragment of DNA PolI.
18. The lentiviral particle according to claim 14, wherein the polypeptide containing the DNA polymerase domain has reduced polymerase activity.
19. A method for producing lentiviral particles, a) transfecting a plurality of eukaryotic cells with the packaging plasmid, the at least one mammalian expression plasmid and the envelope plasmid of the system according to claim 10; and b) Culturing the transfected eukaryotic cells for a sufficient amount of time for lentiviral particles to be produced. A method that includes this.
20. The lentivirus particles are as follows: (i) Nucleic acid sequences encoding a fusion protein including the following: (a) polypeptides containing a DNA polymerase domain; and (b) CRISPR-related endonuclease coding sequences; and (ii) Guide RNA The method according to claim 19, comprising a ribonucleotide protein (RNP) complex containing the following.
21. The method according to claim 20, wherein the plurality of eukaryotic cells are mammalian cells.
22. Lentivirus particles prepared by the method described in claim 19.
23. A composition for use in a method for modifying an intracellular genomic target sequence, comprising lentiviral particles according to any one of claims 14 to 18, wherein the method comprises transduction of a plurality of eukaryotic cells with a plurality of viral particles, the plurality of viral particles comprising the lentiviral particles, the RNP complex binding to the genomic target sequence in the genomic DNA of the cells, and the CRISPR-associated endonuclease cleaving the genomic target sequence to produce a double-strand break, thereby modifying the genomic target sequence.
24. The composition according to claim 23, wherein non-homologous end joinings (NHEJs) are increased in cells compared to cells modified with a CRISPR-related endonuclease that is not fused to the DNA polymerase domain.
25. The composition according to claim 23, wherein the ratio of NHEJ to non-NHEJ increases in the cells.
26. The composition according to claim 25, wherein the non-NHEJ terminal bond is a microhomology-mediated terminal bond (MMEJ) and / or single-strand annealing (SSA).
27. The composition according to claim 23, wherein the number of on-target deletions is reduced in the cells.
28. The composition according to claim 27, wherein the number of on-target deletions larger than 500 base pairs is reduced in the cells.
29. The composition according to claim 23, wherein the ratio of on-target single base pair (1 bp) deletions to on-target deletions greater than 1 bp is increased in the cells.
30. The composition according to claim 29, wherein the ratio of on-target single base pair (1 bp) deletions to deletions exceeding 500 base pairs is increased in the cells.
31. The composition according to claim 23, wherein the number of template insertions (TIS) increases in the cells.
32. The composition according to claim 31, wherein the ratio of TIS to non-TIS is increased in the cells.
33. The composition according to claim 23, wherein the plurality of eukaryotic cells are mammalian cells.
34. The composition according to claim 23, wherein the plurality of eukaryotic cells are cells present within the target.
35. The composition according to claim 34, wherein the subject is a human subject.
36. The composition according to claim 35, characterized in that the plurality of virus particles are injected into the target.
37. A cell containing the plasmid described in claim 1.
38. Cells containing the lentiviral packaging system according to claim 10.
39. A cell containing lentiviral particles as described in claim 14.
40. Cells modified using the method of claim 23.
41. The method is a method for treating a target disease, a) Obtaining cells from the subject, b) Modifying the target cells using the composition, and c) Administering the modified cells to the subject, This method includes, The composition according to claim 23.
42. The composition according to claim 41, wherein the disease is cancer.
43. The composition according to claim 41, wherein the disease is Duchenne muscular dystrophy.
44. The composition according to claim 41, wherein the cells are T cells.