Target gene editing constructs and methods for using them
Nucleic acid constructs with engineered DNA-binding proteins and transposases or integrases facilitate site-specific integration of exogenous nucleic acids, addressing the inefficiencies and risks of current gene delivery methods, enhancing precision and safety in gene therapy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-06-11
- Publication Date
- 2026-03-16
AI Technical Summary
Current gene delivery strategies, such as those using lentiviruses, suffer from low efficiency and lack of specificity, leading to random integration and genotoxicity due to viral DNA insertion into host genomes, particularly at highly transcribed sites.
Development of nucleic acid constructs and fusion proteins, including engineered zinc finger proteins and Cas9 proteins, combined with highly active PiggyBac transposases or recombinant HIV integrases, for site-specific integration of exogenous nucleic acids into the genome.
Enables precise and controlled insertion of exogenous nucleic acids into specific genomic locations, reducing the risk of unintended mutations and genotoxicity, making it suitable for gene therapy applications.
Smart Images

Figure 0007830129000045 
Figure 0007830129000046 
Figure 0007830129000047
Abstract
Description
Detailed description of the invention
[0001] [Reference to electronically submitted sequence listings] The contents of the sequence listing submitted electronically in the ASCII text file (name: 4349.001PC01_Seqlisting_ST25; size: 389,120 bytes; creation date: June 11, 2020) filed with this application are incorporated herein by reference in their entirety.
[0002] 〔background〕 Many diseases, including cancer, developmental disorders, and some infectious diseases, share both congenital and acquired abnormalities. Gene therapy is designed to introduce genetic material into cells and directly target and edit the genome in order to correct genetically non-functional cells and thereby cure the associated disease. Zinc finger nucleases (ZFNs), Talen, and Crisp-Cas9 gene editing technologies represent some of the recently developed tools for DNA editing. Methods such as electroporation, cationic lipids, microinjection, or viruses have been used for the delivery of genetic material to the genome. Current strategies for gene delivery are generally based on adenoviruses, retroviruses, or naked DNA plasmids.
[0003] Lentiviruses, including HIV, are powerful tools when used as vectors for nucleic acid delivery. Lentiviruses have the ability to reliably infect both dividing and non-dividing cells. Lentiviral vectors are prone to random integration into the host genome, often at highly transcribed gene sites, increasing the risk of insertional mutations.
[0004] HIV-1 integrase catalyzes the insertion of viral DNA into the host genome. Generally, HIV-1 integrase consists of an N-terminal domain (NTD), a catalytic core domain (CCD), and a C-terminal domain (CTD). The NTD is used as a key cofactor to bind and coordinate to the Zn2+ cation, while the CTD is used for DNA binding. The CCD forms the catalytic core through which the insertion process is catalyzed. Challenges with insertion mechanisms used by viral vectors include low efficiency and lack of specificity, which can lead to unintended insertion mutations and genotoxicity.
[0005] 〔overview〕 Some aspects of the present invention provide constructs, plasmids, vectors, particles, fusion proteins, compositions, methods, and kits effective for targeted editing of nucleic acids, which include editing of a single site or region within a target genome, such as the human genome.
[0006] The examples herein provide detailed experimental data that credibly demonstrate the successful generation of programmable transposases and integrases and fusion protein constructs of Cas9 / zinc finger proteins. Furthermore, such constructs were able to induce site-specific integration of exogenous nucleic acid sequences into the genome of transfected cells. While not bound by theory, the inventors believe this is the first generation of such a fusion protein with the ability to site-specific integration of exogenous nucleic acids into the genome, making it particularly suitable for gene therapy involving large genes. The inventors also identified recombinant, highly active PiggyBac transposases that perform specific target transposition.
[0007] Consequently, one aspect of this disclosure relates to nucleic acid constructs. These nucleic acid constructs are a) A first polynucleotide sequence comprising a nucleic acid encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in the genome, wherein the first DNA-binding protein is a zinc finger protein or a Cas9 protein, b) A second polynucleotide sequence comprising a nucleic acid encoding a second DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein is i. Highly active PiggyBac transposase, or recombinant highly active PiggyBac with improved specificity for insertion of exogenous nucleic acids into the genome compared to highly active PiggyBac, or ii. Human immunodeficiency virus (HIV) integrase, or recombinant HIV integrase with improved specificity for insertion of exogenous nucleic acids into the genome compared to HIV integrase. The second polynucleotide sequence, c) any polynucleotide sequence including nucleic acids encoding a linker, A nucleic acid construct encodes a fusion protein comprising a first DNA-binding protein, a second DNA-binding protein, and an optional linker between the first and second DNA-binding proteins. Fusion proteins enable the insertion of exogenous nucleic acids into specific locations in the genome.
[0008] Also provided are compositions comprising a nucleic acid construct, vector, or fusion protein described herein and a polynucleotide sequence encoding an exogenous nucleic acid for insertion into a genome, wherein the composition is contained in or conjugated to a packaging vector.
[0009] This disclosure also provides a method for the controlled site-directed incorporation of one or more copies of an exogenous nucleic acid sequence into a cell. The method comprises (a) delivering a nucleic acid construct, vector, or fusion protein as described herein to a cell, and (b) delivering an exogenous nucleic acid to a cell, wherein the binding of the fusion protein to a specific genomic DNA sequence in the cell's genome causes genomic cleavage and incorporation of one or more copies of the exogenous nucleic acid into the cell's genome.
[0010] Another embodiment relates to providing a recombinant highly active PiggyBac transposase containing amino acid sequence number 9, where the amino acid at position 245 is A, the amino acid at position 275 is R or A, the amino acid at position 277 is R or A, the amino acid at position 325 is A or G, the amino acid at position 347 is N or A, the amino acid at position 351 is E, P or A, the amino acid at position 372 is R, the amino acid at position 375 is A, the amino acid at position 450 is D or N, the amino acid at position 465 is W or A, the amino acid at position 560 is T or A, the amino acid at position 564 is P or S, the amino acid at position 573 is S or A, the amino acid at position 592 is G or S, and the amino acid at position 594 is L or F.
[0011] In some embodiments, a fusion protein of (i) an integrase, recombinant integrase, transposase, or recombinant transposase, (ii) a Cas9 or zinc finger protein linked thereto, and a nucleic acid construct encoding the same are provided.
[0012] A particular aspect of this application relates to a nucleic acid construct, the nucleic acid construct comprising: (a) a first polynucleotide sequence encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in a genome; (b) a second polynucleotide sequence encoding a second DNA-binding protein enabling the insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein is (i) an integrase or an integrase recombinant over wild-type integrase, or (ii) a transposase or a transposase recombinant over wild-type transposase; and (c) a third polynucleotide sequence comprising a nucleic acid encoding a linker, wherein the nucleic acid construct encodes a fusion protein comprising the first DNA-binding protein, the second DNA-binding protein, and a linker between the first DNA-binding protein and the second DNA-binding protein.
[0013] In some embodiments, the nucleic acid construct comprises (a) a first polynucleotide sequence encoding the Cas9 protein, and (b) a second polynucleotide sequence encoding the transposase or recombinant highly active PiggyBac or a functional fragment thereof.
[0014] In some embodiments, the nucleic acid construct comprises (a) a first polynucleotide sequence encoding a zinc finger protein, and (b) a second polynucleotide sequence encoding an integrase or recombinant integrase or a functional fragment thereof of the Disclosure.
[0015] In some embodiments, this application relates to plasmids, vectors, or host cells comprising nucleic acid constructs of the present disclosure.
[0016] Some aspects of this application relate to a fusion protein, the fusion protein comprising a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in a genome, a second DNA-binding protein enabling the insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein is an integrase, a transposase, or a recombinant integrase or transposase, and a linker connecting the first protein and the second protein.
[0017] In some embodiments, the fusion protein comprises (a) a Cas9 protein and (b) the highly active PiggyBac of the Disclosure or recombinant highly active PiggyBac or a functional fragment thereof.
[0018] In some embodiments, the fusion protein comprises (a) a zinc finger protein and (b) an integrase or recombinant integrase or a functional fragment thereof of the present disclosure.
[0019] Some aspects of this application relate to lentiviral particles comprising the fusion protein of the present disclosure.
[0020] Some aspects of this application relate to a method for inserting an exogenous nucleic acid sequence into the genomic DNA of an organism. The method comprises administering a lentiviral particle containing a nucleic acid construct or fusion protein of the present disclosure to an organism to cause first and second DNA-binding proteins to bind to a specific genomic DNA sequence, thereby inserting the exogenous nucleic acid into the genomic DNA, where the exogenous nucleic acid becomes incorporated into the specific genomic DNA sequence.
[0021] Some aspects of the present disclosure relate to methods for the controlled site-specific integration of single or multiple copies of exogenous nucleic acid sequences into cells. The method includes (a) delivering a fusion protein of the present disclosure to a cell, and (b) delivering an exogenous nucleic acid to the cell, wherein the binding of the fusion protein to a specific genomic DNA sequence in the genome of the cell causes cleavage of the genome and integration of one or more copies of the exogenous nucleic acid into the genome of the cell, and the fusion protein is delivered to the cell by lentiviral particles.
[0022] Throughout the specification and claims, the term "comprising" and its conjugations are not intended to exclude other technical features, additives, components, or steps. Further objects, advantages, and features of the present invention will become apparent to those skilled in the art upon examination of this specification or may be learned by practice of the present invention. Furthermore, the present invention encompasses all possible combinations of the specific preferred embodiments described herein. The following examples and drawings are provided herein for illustrative purposes and are not intended to limit the present invention.
[0023] [Brief Description of the Drawings] Figures 1A and 1B show the percentage of cells having an exogenous nucleic acid sequence integrated into their genomes after transfection with (Figure 1A) a Cas9-PiggyBac fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and hyperactive PiggyBac (PB) transposase) and (Figure 1B) a Cas9-SB100 fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and hyperactive Sleeping Beauty (SB100) transposase). Vectors were constructed in which the 3'-end of Cas9 was linked to the 5'-end of each transposase by a GGS linker (SEQ ID NO: 48, 49) (hCas9PB, nCas9PB, dCas9PB, hCas9SB, nCas9SB, and dCas9SB). Other vectors were constructed in which the 3'-end of each transposase was linked to the 5'-end of Cas9 by a GGS linker (SEQ ID NO: 48, 49) (PBhCas9, PBnCas9, PBdCas9, SBhCas9, SbnCas9, and SBdCas9). "PiggyBac" (Figure 1A) and "SB100" (Figure 1B) were used as positive controls, and transposons encoding RFP (shown as "episomal RFP" in Figure 1A) and GFP (shown as "episomal GFP" in Figure 1B) alone were used as negative controls. Figure 1C is a different representation of Figure 1A and shows the transposition activity by PB and Cas9 in different configurations.
[0024] Figure 2A shows a plasmid construct encoding a Cas9 / PB fusion protein.
[0025] Figure 2B shows the percentage of cells with exogenous nucleic acid sequences integrated into the genome by fusion constructs formed by human Cas9-PiggyBac ("targeted HCas9") or nickase Cas9-PiggyBac ("targeted NCas9"). The 3' end of Cas9 was ligated to the 5' end of the transposase by a linker. "Non-targeted" is a control for total insertion (PiggyBac alone), and "episome" is a negative control for non-integration (transposon alone).
[0026] Figure 3 shows an exemplary ZFP-integrase fusion protein. The ZFP and integrase are linked by a GGS sequence. NLS refers to the nuclear localization sequence.
[0027] Figure 4 shows the lentiviral titers of wild-type integrase lentivirus (LV), empty virus particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZP-IN(AAVS1)), non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN), and wild-type integrase lentivirus with wild-type integrase (LV+IN). (') indicates a technical repeat.
[0028] Figure 5 shows the percentage of cells that incorporated the exogenous nucleic acid sequence into their genome (total integration) after infection with wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZP-IN(AAVS1)), non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN), and wild-type integrase lentivirus with wild-type integrase (LV+IN). For each condition, from left to right, the first column represents day 3, the second column represents day 5, the third column represents day 7, the fourth column represents day 10, and the fifth column represents day 12.
[0029] Figure 6 shows images of chromosomes containing representative AAVS1 integration and non-integration sites. The asterisks represent the AAVS1 sites on chromosome 19, the triangles represent non-target integration sites, and the diamonds represent targeted integration sites.
[0030] Figure 7A shows the viral titers generated by wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)).
[0031] Figure 7B shows the percentage of cells that have incorporated the exogenous nucleic acid sequence into their genome (total integration) after infection with wild-type integrase lentivirus (LV), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)).
[0032] Figure 7C shows the percentage of cells that incorporated exogenous nucleic acid sequences into their genome after infection with wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)).
[0033] Figure 7D shows the percentage of cells that incorporated exogenous nucleic acid sequences into their genome after infection with wild-type integrase lentivirus (LV), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)).
[0034] Figures 8A-8C show lentiviral titers (Figure 8A), the percentage of CAR-expressing cells on day 3 and day 14 (Figure 8B), and the percentage of CD3-expressing cells (Figure 8C). Jurkat cells were infected with lentiviruses under several conditions. Several conditions were identified: wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZFP-IN(TRCa-1)), and non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN). NILV showed a drastic decrease in titer, and the expression of IN WT in virus-producing cells or complementarity with fused ZNF-IN did not have a rescue effect on either titer or integration ability. Furthermore, when integration was directed toward the TCR gene locus (CD3 protein expression), the cells did not lose CD3 expression. This suggests the need for further factors for transcomplementation, such as VPR proteins (especially in the context of this cell line).
[0035] Figures 9A-9B show the titers for WT lentivirus and two different integrase-deficient viral systems (NILV and TAA, the latter indicating the introduction of a stop codon at the beginning of the IN-coding region in the lentiviral packaging plasmid) individually, or for those heterologously complemented by IN or VPR_IN fusion. Titers were detected by fluorescence cytometry analysis on post-infection day 3 (Figure 9A). Figure 9B shows the relative integration efficiency of heterologously complementary integration mechanisms (machineeries) demonstrating the advantages of VPR protein fusion to IN for heterologously complementary. WT: Lentivirus produced using WT IN; NILV: Lentivirus produced using non-integrated IN with two mutations in its catalytic center; TAA: Lentivirus produced using IN-deficient IN in which the protein is not expressed; +IN: Lentivirus heterologously complemented using IN; +VPR-IN: Lentivirus heterologously complemented using IN fused to VPR at the C-terminus.
[0036] Figure 10A shows the structure of a nucleic acid construct formed by an insertion domain having a DNA-binding domain and a programmable DNA-recognition domain fused by a linker. Figure 10B shows the fusion of Cas9 and transposase linked by a linker in a different configuration.
[0037] Figure 11 shows the results of Cas9 activity in Cas9 bound to hyPB using various linker sizes and compositions. Cas9 activity was measured by sequencing the gRNA target site and analyzing the indel frequency using CRISPR-GA. Two different gRNAs targeting the AAVS1 site were used. The linkers used were SEQ ID NOs. 50–63.
[0038] Figure 12 shows the results for programmable transposer zegene trap transfer efficiency. RFP fluorescence was measured by flow cytometry 10 days after transfection. The importance of linker length and composition in target insertion was determined using different linkers. Average of two independent experiments. Linkers used are SEQ ID NOs. 50–63.
[0039] Figure 13 shows the results for the hcas9_PB linker. Target translocation efficiency of different cas9-PB linker constructs was measured using fragmented GFP cell lines with two different gRNAs. GFP expression was measured by flow cytometry 72 hours after transfection.
[0040] Figure 14 outlines the split GFP reporter cell lines generated for high-throughput analysis of a library of various hyPB variants and for confirmation of individual variants. Using the Sleeping Beauty 100x system, a splicing acceptor (SA) downstream of the target region, followed by half of the GFP coding sequence (Ct-GFP), was introduced into the genome of Hek293T cells. For this screening, the PiggyBac transposon adjacent to the terminal repeat sequence (ITR) was either the complete RPF expression cassette, followed by the promoter and the remaining half of GFP (Nt-GFP) and the splicing donor (SD); or just half of the GFP fragment; as shown in the figure.
[0041] Figure 15 shows the results of targeted translocation of hcas9_PB selective mutants. Targeted translocation efficiency of hcas9_PB D450N and hcas9_PB R372A K375A D450. GFP expression was measured by flow cytometry 72 hours after transfection. Average of 4 independent experiments.
[0042] Figure 16 shows the results for random and targeted translocation of hcas9_PB selected mutants. Targeted translocation efficiency and random translocation efficiency for hcas9_PB D450N and hcas9_PB R372A K375A D450. GFP expression was measured by flow cytometry 72 hours after transfection, and RFP expression was measured by flow cytometry 15 days after transfection, normalized by RFP fluorescence 48 hours after transfection, which was assumed to be the transfection efficiency.
[0043] Figure 17 is a diagram illustrating the fusion of a ZFP with a transposase linked by linkers of various configurations.
[0044] Figure 18 shows the results of ZFP-PB fusion protein target transposition. Target transposition efficiency of ZFP_hyPB or ZFP_hyPBD450N in N and C-terminal conformations. GFP expression was measured by flow cytometry 5 days after transfection. More than 1 independent replicate. ZFP_PB: ZFP fused with hyPB in the C-terminal configuration using an XTEN linker; PB_ZFP: ZFP fused with hyPB in the N-terminal configuration using an XTEN linker; ZFP_450: ZFP fused with hyPB(D450N) in the C-terminal configuration using an XTEN linker; 450_ZFP: ZFP fused with hyPB(D450N) in the N-terminal configuration using an XTEN linker; hyPB: non-recombinant hyPB; 1 / 2GFP: control transposon alone.
[0045] Figure 19 shows a diagram of the analytical methods used in screening libraries of PiggyBac mutations.
[0046] In Figure 20, PiggyBac 1116 base pairs containing all library variants were sequenced using Illumina NGS technology. Except for variants 450 and 465, the I7 index primers were replaced with custom primers to enable complete sequencing of the different variants.
[0047] Figures 21A-21B show the results of hyPB library diversity generation. Figure 21A is an example of a sort plot. Positive target inclusion hits (GFP fluorescence) were selected at gate P4, and negative target inclusion hits (no GFP fluorescence) were selected at gate P5. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining. Figure 21B shows the results of dual plasmid transfection efficiency. Transfection efficiency was measured by transfecting GFP and RFP plasmids equimolarly with 1 / 2 GFP and gRNA transfections on the same day and under the same conditions. Gate P8 selects dual plasmid transfections. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining.
[0048] Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale.
[0049] Figure 23A shows the positive variants of Top1 and Top3 from three independent replicates. There is only a one-amino acid difference at position 254. Figure 23B shows the three top1 variants identified from three independent replicates. WT hyPB is also shown for reference.
[0050] Figure 24A shows the most overrepresented variants in GFP-positive and RFP-positive cells. It shows clustering of GPF, targeted insertions; RPF, random insertions, and negative populations. Figures 24B and 24C show variants observed among positive hits in more than one independent replicate. Rep: independent experimental replicate; Pos: positive cells with targeted insertion; Neg: negative cells where targeted insertion did not occur.
[0051] Figure 25 shows a histogram of variant covariates, representing the proportion of variants found together with others in positive samples divided by the proportion of negative samples. In addition to variants included in the library design, variants randomly introduced by lentiviral retrotranscriptase during viral library generation were analyzed. Some of these new variants are associated with positive hits and perform targeted incorporation for combinations. Examples of D450N and W465A are shown.
[0052] Figure 26 shows that recombinant hyPB showed greater enhancement in target integration compared to WT hyPB when fused with Cas9. Cas9 was fused to hyPB or different combinations of hyPB variants (Unilarge-A:D450N; Unilarge-B:R245A / D450N; Unilarge-C:R245A / G325A / D450N / S573P; Unilarge-D:R245A / G325A / S573P) using a 4GGS linker and reporter cell line system.
[0053] Figure 27 shows the results of heterogeneous complementation of integrase-deficient cells. Viral production efficiency measured on day 2 and integration capacity measured on day 7 were evaluated for different systems in Hek293T cells. Western blotting showed the presence of intergeneric IN in the viral particles. Viral production efficiency and its integration capacity were evaluated by infecting Hek293T cells with different conditions of integration-deficient virus and heterogeneous complementar virus. Cells were passaged for 7 days until episomal signaling was no longer detectable, and GFP signaling was analyzed by flow cytometry on days 2, 5, and 7. Since NILV is close to WT during production, different production efficiencies can be detected for different systems. In all cases, a clear rescue of integration activity was observed when heterogeneous complementation was performed using WT-HIV_IN. Proof of IN being loaded into the heterogeneous complementarized system was obtained by Western blotting. WT: Lentivirus produced with WT IN; NILV: Lentivirus produced with non-integrated IN and having two mutations in its catalytic center; TAA: IN-deficient lentivirus in which the protein is not expressed because a stop codon is present at the beginning of the IN coding sequence; TAAx3: Lentivirus produced with IN-deficient IN in which the protein is not expressed because three consecutive stop codons are present at the beginning of the IN coding sequence; Delta-IN: Lentivirus produced with IN-deficient IN in which the IN coding sequence has been removed; Delta-IN_cPPT: Lentivirus produced with IN-deficient IN in which the IN coding sequence has been replaced with a central polypyrimidine trac(cPPT) sequence; +VPR-IN: Lentivirus that is interspecies complementary to IN, fused with VPR at the C-terminus.
[0054] [Detailed description of the invention] [I. Definition] Where used herein, the singular forms "a," "an," and "the" include both singular and plural references unless the context explicitly indicates otherwise. For example, a reference to "an agent" includes both a singular agent and a plural such agent.
[0055] The terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” are used interchangeably and refer to deoxyribonucleotides or ribonucleotide polymers that are linear or cyclic in form and are either single-stranded or double-stranded. For the purposes of this disclosure, these terms should not be construed as limiting with respect to the length of the polymer. These terms may encompass known analogues of natural nucleotides, as well as nucleotides that have been recombinant in the base, sugar, and / or phosphate moieties (e.g., phosphorothioate skeletons). Generally, analogues of a particular nucleotide have the same base-pairing specificity; that is, an analogue of A forms a base pair with T.
[0056] The terms "polypeptide," "peptide," and "protein" are used interchangeably and refer to polymers of amino acid residues. This term also applies to amino acid polymers, where one or more amino acids are chemical analogs or recombinant derivatives of the corresponding naturally occurring amino acids.
[0057] As used herein, the term “binding protein” refers to a protein that can non-covalently bind to another molecule. Binding proteins can, for example, bind to DNA molecules (DNA-binding proteins), RNA molecules (RNA-binding proteins), and / or protein molecules (protein-binding proteins). In the case of protein-binding proteins, they can bind to themselves (forming homodimers, homotrimers, etc.), and / or they can bind to one or more molecules of one or more different types of proteins. Binding proteins can have two or more types of binding activity. For example, zinc finger proteins have DNA-binding activity, RNA-binding activity, and protein-binding activity.
[0058] As used herein, the term “zinc finger protein” refers to a protein or domain within a larger protein that binds to DNA in a sequence-specific manner via one or more zinc fingers. Here, a zinc finger is a region of amino acid sequence within the binding domain of a zinc finger protein whose structure is stabilized by the coordination of zinc ions. The term zinc finger protein is sometimes abbreviated as ZFP.
[0059] The term "zinc finger nuclease" refers to an artificial restriction enzyme produced by fusing a zinc finger DNA-binding domain to a DNA-cleaving domain. The zinc finger domain can be manipulated to target specific, desired DNA sequences. This allows zinc finger nucleases to target unique sequences within a complex genome. Zinc finger nucleases are sometimes abbreviated as ZFN or ZNP.
[0060] As used herein, the terms “nucleic acid sequence,” “polynucleotide sequence,” or “gene sequence” refer to a nucleotide sequence of any length. A nucleotide sequence may be DNA or RNA, may be linear, circular, or branched, and may be single-stranded or double-stranded.
[0061] As used herein, the terms “amino acid sequence,” “polypeptide,” or “protein” refer to a polymer of amino acid residues. Unless otherwise specified, a polymer of amino acid residues may be of any length.
[0062] As used herein, the term “exogenous” refers to a molecule that is not normally present in a cell but can be introduced into the cell by one or more genetic, biochemical, or other means. The normal presence in a cell is determined for the specific developmental stage and environmental conditions of the cell. For example, a molecule present only during muscle embryonic development is an exogenous molecule for adult muscle cells. Similarly, a molecule induced by heat shock is an exogenous molecule for non-heat-shock cells. Exogenous molecules may include, for example, a functional version of a dysfunctional endogenous molecule, or a dysfunctional version of a normally functioning endogenous molecule.
[0063] In contrast, "endogenous" molecules are molecules that are normally present in specific cells at specific developmental stages and under specific environmental conditions. For example, endogenous nucleic acids may include chromosomes, mitochondrial genomes, chloroplasts or other organelles, or naturally occurring episomal nucleic acids. Further endogenous molecules may include proteins, such as transcription factors and enzymes.
[0064] A "target site" or "target sequence" is a sequence that defines a portion of a nucleic acid or polypeptide to which a binding molecule can bind, provided that sufficient conditions for binding are present. For example, the sequence 5'-GAATTC-3' is the target site for EcoRI restriction endonuclease.
[0065] As used herein, the term “fusion” refers to a molecule in which two or more subunit molecules are linked, preferably covalently. The subunit molecules may be molecules of the same chemical type or molecules of different chemical types.
[0066] As used herein, the term “fusion protein” refers to a hybrid polypeptide comprising protein domains derived from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) portion, thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively.
[0067] As used herein, the terms "gene" or "genome" include the DNA region that codes for a gene product, as well as all DNA regions that regulate the production of the gene product. Such regulatory sequences are whether or not they are adjacent to the coding and / or transcription sequences. Thus, a gene includes, but is not limited to, promoter sequences, terminators, translation regulatory sequences (such as ribosome binding sites and internal ribosome entry sites), enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus regulatory regions.
[0068] The term "eukaryotic" cells include, but are not limited to, fungal cells (such as yeast), plant cells, animal cells, mammalian cells, and human cells (e.g., T cells).
[0069] As used herein, the term “linked” means that two or more components (such as array elements) are arranged in parallel, where both components function properly and enable at least one component to mediate a function exerted on at least one of the other components.
[0070] A “functional fragment” of a protein, polypeptide, or nucleic acid is a protein, polypeptide, or nucleic acid whose sequence is not identical to that of the full-length protein, polypeptide, or nucleic acid, but which retains the same function as the full-length protein, polypeptide, or nucleic acid. A functional fragment may have more, fewer, or the same number of residues as the corresponding native molecule, and / or may contain one or more amino acid or nucleotide substitutions.
[0071] As used herein, the term “transfect” refers to the introduction of nucleic acids (either DNA or RNA) into eukaryotic or prokaryotic cells or organisms.
[0072] As used herein, the term “breakage” refers to a rupture of the covalent backbone of a DNA molecule. Breakage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand breaks are possible, and a double-strand break may occur as a result of two different single-strand breaks. DNA breaks may result in the production of either blunt or adherent ends. In certain embodiments, a fusion polypeptide is used for targeted double-strand DNA breaks.
[0073] As used herein, the term “integrase” refers to an enzyme produced by a virus that enables genetic material to be incorporated into the DNA of an infected cell, such as genomic DNA.
[0074] As used herein, the term “specificity” refers to the ability to selectively bind to sequences that share a certain degree of sequence identity with a selected sequence.
[0075] As used herein, the terms “insertion” and “incorporation” refer to the addition of a nucleic acid sequence to a second nucleic acid sequence or genome.
[0076] The terms “specific,” “site-specific,” “targeted,” and “targeted” are used interchangeably herein to refer to insertions of nucleic acids, second nucleic acids, or genomes at specific sites. The terms “random,” “untargeted,” and “untargeted” refer to nonspecific and unintended gene insertions. The term “total” or “whole” refers to the total number of insertions.
[0077] As used herein, the term “mutation” refers to the substitution of one residue in a sequence, such as a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. In this specification, mutations are typically described by identifying the original residue, then the location of that residue in the sequence, and finally the newly substituted residue. Various methods for producing amino acid substitutions (mutations) provided herein are known in the art, for example, as provided in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0078] As used herein, the term “transposase” refers to an enzyme that binds to the end of a transposon and catalyzes the transposon’s movement to another part of the genome via a cut-and-paste or transmutation mechanism.
[0079] As used herein, the term “recombinant” refers to a protein or nucleic acid sequence that is different from the corresponding non-recombinant protein or nucleic acid sequence.
[0080] As used herein, the term “linker” refers to a chemical group or molecule that links two adjacent molecules or parts.
[0081] As used herein, the terms “vector” and “plasmid” refer to any polynucleotide capable of carrying, for example, a second polynucleotide of interest, and capable of transferring, for example, a gene sequence into a target cell. Thus, the terms include cloning and expression vehicles, as well as embedded vectors. In particular, as used herein, the term “expression vector” refers to any polynucleotide capable of directing the expression of a nucleic acid. In some embodiments, the terms “vector” and “plasmid” are used interchangeably with the term “nucleic acid construct.”
[0082] As used herein, the term “identity percentage” refers to the identity percentage of two sequences, regardless of whether they are nucleic acid sequences or amino acid sequences. The identity percentage is calculated by dividing the number of exact matches between two aligned sequences by the length of the shorter sequence and multiplying by 100.
[0083] As used herein, the terms “recombinant” or “manipulated” refer to artificially created proteins or nucleic acid sequences.
[0084] As used herein, the term “subject” refers to an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal.
[0085] Where used herein, the terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at reversing, reducing, delaying, progressing, or inhibiting one or more symptoms of a disease or disorder. Where used herein, the terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at reversing, reducing, delaying, progressing, or inhibiting one or more symptoms of a disease or disorder. In some embodiments, treatment may be administered after the onset of one or more symptoms and / or after the disease has been diagnosed. In other embodiments, treatment may be administered asymptomatically, for example, to prevent or reduce the likelihood of symptom onset, delay the onset of symptoms, or inhibit the onset or progression of the disease. For example, treatment may be administered to a susceptible individual before the onset of symptoms (for example, in light of symptom history and / or genetic or other susceptibility factors). Treatment may also be continued after symptoms have subsided, for example, to prevent or delay their recurrence.
[0086] [II. Nucleic acid construct] Targeted editing of nucleic acid sequences, such as the introduction of specific recombination into genomic DNA (e.g., insertion of exogenous nucleic acids), is a promising technique for treating human genetic diseases. To this end, the inventors aim to provide improved nucleic acid constructs for use in genome editing that are highly efficient in introducing desired recombination, minimal non-targeting activity, and the ability to be programmed to precisely edit sites within the human genome.
[0087] Certain aspects of this application relate to nucleic acid constructs for use in improving site-directed insertion of exogenous nucleic acids (e.g., genes of interest (GOIs)) into the genome. In some embodiments, the GOI is a therapeutic gene, such as a gene encoding a therapeutic protein. Examples of therapeutic genes of interest include the CFTR gene (Cystic Fibrosis Membrane Conductance Regulator) for treating cystic fibrosis; the SMN1 gene (Survival Motor Neuron 1) for treating spinal muscular atrophy (SMA); the LRP5 gene (LDL Receptor-Related Protein 5) variant G171V for preventing osteoporosis and fractures; and the APP gene (Amyloid-Beta Progenitor Protein) variant A673T for reducing predisposition to Alzheimer's disease.
[0088] In some embodiments, the exogenous nucleic acid for insertion (e.g., GOI) may have a length of up to about 10 kb, up to about 15 kb, up to about 20 kb, up to about 25 kb, up to about 30 kb, up to about 35 kb, or up to about 40 kb.
[0089] In some embodiments, the polynucleotide sequence encoding a DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome comprises an integrase, or an integrase recombinant to wild-type integrase. The exogenous nucleic acid for insertion may be up to 10 kb, up to 15 kb, or up to 20 kb in length, for example, about 1 kb to about 20 kb, about 1 kb to about 19 kb, about 1 to about 18 kb, about 1 kb to about 17 kb, about 1 kb to about 16 kb, or about 1 kb to about 15 kb.
[0090] In some embodiments, the polynucleotide sequence encoding a second DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome comprises a transposase, or a transposase recombinant to a wild-type transposase. The exogenous nucleic acid for insertion may be up to 10 kb, up to 15 kb, up to 20 kb, up to 25 kb, up to 30 kb, up to 35 kb, or up to 40 kb in length, for example, about 1 kb to about 40 kb, about 1 kb to about 39 kb, about 1 to about 38 kb, about 1 kb to about 37 kb, about 1 kb to about 36 kb, or about 1 kb to about 35 kb.
[0091] In some embodiments, the nucleic acid construct comprises a first DNA-binding protein, e.g., a polynucleotide sequence encoding a gene-editing polypeptide, and a second DNA-binding protein, e.g., an integrase or transposase, where the nucleic acid construct encodes the first and second binding proteins as a fusion protein. In some embodiments, the nucleic acid construct further comprises a nucleic acid sequence encoding a linker between the first and second binding proteins. In some embodiments, the nucleic acid construct encodes a fusion protein that enables and / or facilitates site-specific insertion of an exogenous nucleic acid into the genome. In some embodiments, the first or second binding protein is an integrase recombinant relative to the wild type. In some embodiments, the first or second binding protein is a transposase recombinant relative to the wild type. Some embodiments relate to vectors or plasmids comprising the nucleic acid construct of the present disclosure. In certain embodiments, the nucleic acid construct of the present disclosure encodes a fusion protein that enhances the specificity of insertion of a nucleic acid (e.g., GOI) into the genome. In some embodiments, the fusion protein and the exogenous nucleic acid are delivered to cells using lentiviral particles.
[0092] In some embodiments, the first and second binding proteins reside on separate nucleic acid constructs. For example, the transposase or integrase (e.g., a recombinant transposase and / or integrase relative to the wild type) resides on a separate nucleic acid construct from Cas9 or ZFP.
[0093] Certain embodiments relate to plasmids or vectors containing nucleic acid constructs disclosed herein. In some embodiments, the plasmid containing the nucleic acid construct is a packaging plasmid. In some embodiments, the plasmid containing the nucleic acid construct further comprises polynucleotides encoding capsid proteins, e.g., gag and pol. In some embodiments, (i) the plasmid containing the nucleic acid construct is combined with (ii) a plasmid containing polynucleotides encoding viral envelope proteins (envelope plasmid), and (iii) a plasmid containing an exogenous nucleic acid sequence (e.g., GOI). When this combination is introduced into a producing cell line (e.g., eukaryotic cells, prokaryotic cells and / or cell lines), a fusion protein is produced containing a viral particle with exogenous nucleic acid, e.g., GOI, as well as first and second binding proteins.
[0094] In some embodiments, (i) a plasmid containing a nucleic acid construct is combined with (ii) a plasmid containing a nucleic acid construct further comprising polynucleotides encoding capsid proteins, e.g., gag and pol (a packaging plasmid, the packaging plasmid lacking a functional integrase), (iii) a plasmid containing polynucleotides encoding proteins of the viral envelope (envelope plasmid), and (iv) a plasmid containing an exogenous nucleic acid sequence (e.g., GOI). When this combination is introduced into a producing cell line (e.g., eukaryotic cells, prokaryotic cells and / or cell lines), a fusion protein is produced comprising viral particles containing exogenous nucleic acids, e.g., GOI, as well as first and second binding proteins.
[0095] The nucleic acid construct comprises a first polynucleotide sequence encoding a first DNA-binding protein engineered to bind to a specific DNA sequence, a second polynucleotide sequence encoding a second DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome (wherein the second DNA-binding protein is an integrase or transposase (e.g., a transposase and / or integrase recombinant relative to the wild type)), and a third polynucleotide sequence comprising a nucleic acid sequence encoding a linker between the first and second polynucleotides. In some embodiments, the first DNA-binding protein is a zinc finger protein or a Cas9 protein.
[0096] In some embodiments, the nucleic acid construct includes a linker selected from the group consisting of (GGS)n, (GGGGS)n (SEQ ID NO: 133), (G)n, (EAAAK)n (SEQ ID NO: 134), XTEN-based linker, or (XP)n motif, or any combination thereof, where n is an independent integer between 1 and 50. In some embodiments, the nucleic acid encodes a linker containing an XTEN sequence or a GGS sequence. In some embodiments, the linker nucleic acid sequence is 3 to 150 nucleotides long. In some embodiments, the linker is 12 to 24 amino acids, or 36 to 72 nucleic acids long. In some embodiments, nucleic acid constructs have lengths of 6-120, 6-90, 6-78, 6-72, 9-120, 9-90, 9-78, 9-72, 12-120, 12-90, 12-78, 12-72, 15-120, 15-90, 15-78, 15-72, 18-120, 18-90, 18-78, 18-72, 21-120, 21-90, 21- The linker nucleic acid sequence comprises 78, 21-72, 24-120, 24-90, 24-78, 24-72, 27-120, 27-90, 27-78, 27-72, 30-120, 30-90, 30-78, 30-72, 33-120, 33-90, 33-78, 33-72, 36-120, 36-90, 36-78, or 36-72 nucleotides. In some embodiments, the nucleic acid encoding the linker has a length of 9-150 nucleotides. In some embodiments, the zinc finger protein is linked to the recombinant integrase of this disclosure using a linker containing a GGS sequence. In some embodiments, the linker has a length of 1-50 amino acids. In some embodiments, the linker has a length of 3-40, 3-30, 3-29, 3-24, 4-40, 4-30, 4-29, 4-24, 5-40, 5-30, 5-29, 5-24, 6-40, 6-30, 6-29, 6-24, 7-40, 7-30, 7-29, 7-24, 8-40, 8-30, 8-29, 8-24, 9-40, 9-30, 9-29, 9-24, 10-40, 10-30, 10-29, 10-24, 11-40, 11-30, 11-29, 11-24, 12-40, 12-30, 12-29, or 12-24 amino acids.
[0097] In some embodiments, the 3' end of the first polynucleotide sequence is linked to the 5' end of the second polynucleotide sequence by a linker-encoding nucleic acid. In some embodiments, the 5' end of the first polynucleotide sequence is linked to the 3' end of the second polynucleotide sequence by a linker-encoding nucleic acid. In some embodiments, the 3' end of the Cas9 protein is linked to the 5' end of the transposase by a linker. In some embodiments, the 5' end of the Cas9 protein is linked to the 3' end of the transposase by a linker. In some embodiments, the 3' zinc finger protein is linked to the 5' end of the integrase by a linker. In some embodiments, the 5' zinc finger protein is linked to the 3' end of the integrase by a linker.
[0098] In some embodiments, a linker is not required because the recombinant integrase or recombinant transposase is expressed from a plasmid separate from Cas9 or ZFP.
[0099] Certain aspects of the present disclosure relate to vectors or plasmids (e.g., expression vectors or packaging vectors) comprising nucleic acid constructs of the present disclosure that are suitable for expression in host cells, such as mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells.
[0100] In some embodiments, the nucleic acid construct comprises (a) a first polynucleotide sequence comprising a nucleic acid encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in the genome, wherein the first DNA-binding protein is a zinc finger protein or a Cas9 protein, and (b) a second polynucleotide sequence comprising a nucleic acid encoding a second DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein (i) enables the insertion of an exogenous nucleic acid into the genome compared to a highly active PiggyBac transposase or highly active PiggyBac. The nucleic acid construct comprises (ii) a recombinant high-activity PiggyBac with improved specificity for insertion of exogenous nucleic acids into the genome compared to human immunodeficiency virus (HIV) integrase, or recombinant HIV integrase with improved specificity for insertion of exogenous nucleic acids into the genome compared to HIV integrase, and (c) any polynucleotide sequence containing a nucleic acid encoding a linker, wherein the nucleic acid construct encodes a fusion protein containing a first DNA-binding protein, a second DNA-binding protein, and an arbitrary linker between the first DNA-binding protein and the second DNA-binding protein, the fusion protein enabling insertion of exogenous nucleic acids into a specific site in the genome.
[0101] In one embodiment, (a) the first DNA-binding protein is a Cas9 protein or a zinc finger protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase, or a recombinant highly active PiggyBac transposase with improved specificity for insertion of exogenous nucleic acids into the genome compared to a highly active PiggyBac transposase.
[0102] In another embodiment, (a) the first DNA-binding protein is the Cas9 protein, or a zinc finger protein, and (b) the second DNA-binding protein is HIV integrase, or recombinant HIV integrase with improved specificity for insertion of exogenous nucleic acids into the genome compared to HIV integrase.
[0103] In some embodiments, the Cas9 protein is one described herein, and is selected in particular from the group consisting of human Cas9, nickase Cas9, and inactive (dead) Cas9, and more particularly human Cas9 or nickase Cas9.
[0104] In one embodiment, when dCas9 is used, the second DNA-binding protein is not a Gin, Hin, or Tn3 recombinase catalytic domain, nor a FokI DNA cleavage domain. Such recombinases and FoKIs require a known site (an acceptor sequence in the genome) to which they can be incorporated. Therefore, the possibilities for target sites are far more limited. Furthermore, they need to form dimers, such as Gin, in order to be functional.
[0105] In another embodiment, the zinc finger protein is one described in this disclosure, and in particular, a C2H2 zinc finger protein comprising six binding domains.
[0106] In another embodiment, the linker is as described in the Disclosure, and in particular the linker includes an XTEN sequence (e.g., SEQ ID NO: 61, coded by SEQ ID NO: 60) or a GGS sequence, more particularly GGSx3 (SEQ ID NO: 49, coded by SEQ ID NO: 48), GGSx4 (SEQ ID NO: 51, coded by SEQ ID NO: 50), GGSx5 (SEQ ID NO: 53, coded by SEQ ID NO: 52), GGSx6 (SEQ ID NO: 55, coded by SEQ ID NO: 54), GGSx7 (SEQ ID NO: 57, coded by SEQ ID NO: 56), or GGSx8 (SEQ ID NO: 59, coded by SEQ ID NO: 58).
[0107] In another embodiment, the 3' end of the first polynucleotide sequence is ligated to the 5' end of the second polynucleotide.
[0108] In some embodiments, recombinant high-activity PiggyBac transposase is described in this disclosure. In other embodiments, recombinant HIV integrase is described in this disclosure.
[0109] In other embodiments, no linker is used. Instead, for example, the first and / or second polynucleotide sequences include nucleic acids encoding the first and second DNA-binding proteins, and further include an additional nucleic acid at at least one of their ends that functions as a linker.
[0110] In one embodiment, (a) the first DNA-binding protein is a Cas9 protein or a zinc finger protein, (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac with improved specificity for insertion of exogenous nucleic acids into the genome compared to highly active PiggyBac, and the nucleic acid construct comprises (c) a polynucleotide sequence containing a nucleic acid encoding a linker including an XTEN sequence or a GGS sequence, the 3' end of the first polynucleotide sequence being ligated to the 5' end of the second polynucleotide.
[0111] In one embodiment, (a) the first DNA-binding protein is a Cas9 protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac, provided that the linker is not KLAGGAPAVGGGPK (SEQ ID NO: 130) when Cas9 is inactive Cas9 (dcas9).
[0112] In one embodiment, (a) the first DNA-binding protein is a zinc finger protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac. Here, as described in this disclosure, the binding domain of the zinc finger protein can be manipulated to bind to a selected sequence, so that the zinc finger protein can recognize multiple recognition sites.
[0113] In one embodiment, (a) the first DNA-binding protein is a zinc finger protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac, and the linker is XTEN.
[0114] In one embodiment, (a) the first DNA-binding protein is a zinc finger protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac. Here, the zinc-binding protein does not have a Gal4 DNA-binding domain. Gal4 is CGG-N 11 -Binds to CCG, where N can be any base. This protein is a positive regulator of the gene expression of galactose-inducible genes such as GAL1, GAL2, GAL7, GAL10, and MEL1, which encode enzymes used to convert galactose to glucose. It recognizes a 17-base pair sequence (5'-CGGRNNRCYNYNCNCCG-3') (SEQ ID NO: 135) in the upstream activation sequence (UAS-G) of these genes. Therefore, Gal4 is not site-specific because it recognizes a short and very high-frequency sequence in the genome. In certain embodiments, the zinc-binding protein has a Gal4 DNA-binding domain that has been engineered to be site-specific.
[0115] In one embodiment, (a) the first DNA-binding protein is a zinc finger protein, and (b) the second DNA-binding protein is either a highly active PiggyBac transposase or a recombinant highly active PiggyBac transposase, provided that the linker is not EFGGGGSGGGGSGGGGSQF (SEQ ID NO: 131).
[0116] In another embodiment, (a) the first DNA-binding protein is a Cas9 protein or a zinc finger protein, (b) the second DNA-binding protein is HIV integrase or a recombinant HIV integrase with improved specificity for insertion of exogenous nucleic acids into the genome compared to HIV integrase, and the nucleic acid construct comprises (c) a polynucleotide sequence comprising a nucleic acid encoding a linker comprising an XTEN sequence or a GGS sequence, wherein the 3' end of the first polynucleotide sequence is ligated to the 5' end of the second polynucleotide.
[0117] In some embodiments, the nucleic acid construct is in the form of DNA or RNA.
[0118] This specification also provides vectors comprising any of the nucleic acid constructs provided herein. In particular, the vectors are suitable for expression in mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. This specification also provides host cells comprising any of the nucleic acid constructs or vectors provided herein.
[0119] [III. Integrases and Recombinant Integrases] Integrases are key enzymes for the stable integration of the viral genome into host cells. However, because the integration site by wild-type integrase is unpredictable, integrases are also associated with insertional mutagenesis. Integration has been shown to be favorable for highly transcribed genes, increasing the risk of mutations in important genes and regulators. Generally, HIV-1 integrase consists of an N-terminal domain (NTD), a catalytic core domain (CCD), and a C-terminal domain (CTD). The NTD contains Zn as a cofactor. 2+ While the CCD domain is used for binding and coordinating to cations, the CTD is used for DNA binding. The CCD domain forms a catalytic core that catalyzes the integration process. Upon entering the host cell and following reverse transcription of the viral RNA genome, four integrase molecules form a tetramer and attach to the ends of the viral DNA. This is called an intosome. The pre-integration complex (PIC) digests the 3'OH ends of the DNA to form 5'OH overhangs, which are later required for nucleophilic attack on host DNA. During the formation of this PIC, the complex is transported to the nucleus. After being transported to the nucleus, the PIC forms a complex with the host DNA called a chain transfer complex (STC). Here, both 3'OH overhangs of the viral DNA attack both sites of the host DNA backbone in a space of approximately 5 nucleotides. This leads to targeted replication of 5 nucleotides. After the nucleophilic attack, the viral DNA is integrated, and the single-stranded DNA portion is repaired by the host cell's DNA repair mechanisms.
[0120] This disclosure provides nucleic acid constructs comprising polynucleotides encoding integrase and recombinant integrase for insertion of exogenous nucleic acids into specific sites of the genome. In some embodiments, the exogenous nucleic acid for insertion may have a length of up to 10 kb, up to 15 kb, or up to 20 kb, for example, about 1 kb to about 20 kb, about 1 kb to about 19 kb, about 1 to about 18 kb, about 1 kb to about 17 kb, about 1 kb to about 16 kb, or about 1 kb to about 15 kb. In some embodiments, the polynucleotide sequence encoding a DNA-binding protein that enables insertion of exogenous nucleic acid into the genome includes integrase. The integrase may be recombinant from wild-type integrase. The exogenous nucleic acid for insertion may have a length of up to 10 kb or up to 15 kb.
[0121] Some aspects of this disclosure provide integrase fusion proteins designed using the methods and strategies described herein. Some embodiments of this disclosure provide nucleic acids encoding integrase or recombinant integrase and / or fusion proteins comprising the same. Some embodiments of this disclosure provide plasmids or expression vectors comprising such nucleic acid constructs encoding integrase or recombinant integrase and / or fusion proteins comprising the same.
[0122] The integrases or recombinant integrases of this disclosure may be any integrase capable of inserting exogenous nucleic acids into specific sites in the genome. Non-limiting examples of integrases include HIV integrase, lentiviral integrase, adenovirus integrase, retroviral integrase, and mammary mouse tumor virus integrase. In some embodiments, the integrase (e.g., recombinant integrase containing one or more recombinants relative to the wild type) is an HIV integrase sequence (sequence numbers 1 and 2, amino acid and nucleic acid sequences, respectively) corresponding to HIV integrase, particularly NC_001802.1. In some embodiments, the recombinant integrase contains one or more recombinants relative to wild-type HIV integrase (sequence numbers 1 and 2).
[0123] In some embodiments, the integrase is recombinant HIV integrase. Recombinant HIV integrase may contain one or more mutations in amino acids selected from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to the amino acid numbers of SEQ ID NO: 1. Recombinant HIV integrase mutations may contain one or more amino acid recombinations listed in Table 8. Recombinant HIV integrase mutations correspond to the amino acid numbers of SEQ ID NO: 1 or 3, specifically D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N120E, and N12 It may contain one or more amino acid recombinants selected from 0R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R.
[0124] In some embodiments, recombinant integrase may contain one or more mutations from the wild type that weaken DNA binding in amino acids 94, 117, 119, 120, 124, and / or 231, corresponding to the amino acid numbers of, for example, SEQ ID NO: 1 or SEQ ID NO: 4 (e.g., G94D, G94E, G94R, G94K, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N120E, N120R, N120K, A124D, A124E, A124R, A124K, R231G, R231K, R231D, R231E, and / or R231K).
[0125] In some embodiments, recombinant integrase may contain one or more mutations from the wild type that enhance DNA binding in amino acids 94, 117, 119, 120, 122, 124, and / or 231, corresponding to the amino acid numbers of, for example, SEQ ID NO: 1 or 5 (e.g., G94D, G94E, G94R, G94K, N117D, N117E, N117R (117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N120E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, R231G, R231K, R231D, R231E, and / or R231S).
[0126] In some embodiments, recombinant integrase may contain one or more mutations from the wild type in amino acids 264, 266, and / or 273, corresponding to the amino acid numbers of, for example, SEQ ID NO: 1 or SEQ ID NO: 6, which are involved in integrase acetylation by p300 (e.g., K264R, K266R, and / or K273R).
[0127] In some embodiments, recombinant integrase may contain one or more mutations in highly conserved amino acids important for retroviral integrase recombination, such as amino acid 10, 13, 64, 116, 128, 152, 168, and / or 170, corresponding to the amino acid numbers of, for example, SEQ ID NO: 1 or SEQ ID NO: 7 (e.g., D10K, E13K, D64A, D64E, D116A, D116E, D116E, A128T, E152A, E152D, Q168L, Q168A, and / or E170G).
[0128] In some embodiments, recombinant integrase may contain one or more mutations at amino acid 168, corresponding to the amino acid number of, for example, SEQ ID NO: 1 or SEQ ID NO: 8, that interfere with the interaction with LEDGF / p75 and weaken chromosomal tethering and HIV-1 replication (e.g., Q168L or Q168A).
[0129] In some embodiments, recombinant HIV integrase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence described in SEQ ID NO: 1. In some embodiments, recombinant HIV integrase comprises an amino acid sequence having one or more of the recombinants disclosed herein to SEQ ID NOs: 1, 3, 4, 5, 6, 7, or 8, each retaining at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the sequence described in SEQ ID NOs: 1, 3, 4, 5, 6, 7, or 8. In some embodiments, recombinant HIV integrase is selected compared to wild-type HIV integrase due to its high specificity for DNA integration into the genome.
[0130] Certain aspects of this disclosure relate to vectors or plasmids (e.g., expression vectors or packaging vectors) comprising nucleic acid constructs containing the integrase or recombinant integrase of this disclosure, suitable for expression in host cells, such as mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. In some embodiments, the integrase or recombinant integrase is expressed as a fusion protein with Cas9 or a zinc finger protein. In some embodiments, the integrase or recombinant integrase is co-expressed with Cas9 or a zinc finger protein from separate vectors but delivered to the same cells. In some embodiments, the integrase or recombinant integrase or a fusion protein containing the same is packaged in lentiviral particles for delivery to cells.
[0131] [IV. Transposases and Recombinant Transposases] Transposons are chromosomal fragments capable of translocation; for example, DNA that can translocate as a whole even without a complementary sequence in the host DNA. Transposons can be used to perform long-chain DNA manipulation in human cells. Common transposon systems used in mammalian cells include Sleeping Beauty (SB), which is reconstructed from an inactive transposon, and PiggyBac (PB), isolated from the moth Trichoplusia. PiggyBac has higher translocation activity than SB and can be excised without leaving a scar.
[0132] Natural DNA transposons typically contain a single gene encoding a transposase protein, flanked by terminal repeat sequences (ITRs) that carry transposase binding sites. During their transposition, the transposase protein recognizes these ITRs and randomly catalyzes the excision and subsequent reintegration of elements at other locations. Furthermore, some of these transposons can be adapted for use in gene therapy protocols by employing them as a binary system. Here, a plasmid contains an expression cassette that can introduce a DNA sequence located between transposon ITRs into the host genome. The introduction is directed by a co-transfected plasmid containing the sequence encoding the transposase enzyme, or its mRNA synthesized in vitro. In certain embodiments of this disclosure, transposon-based systems are used to efficiently mediate the stable integration and sustained expression of transgenes, such as therapeutic genes.
[0133] This disclosure provides nucleic acid constructs comprising polynucleotides encoding a transposase or recombinant transposase for insertion of exogenous nucleic acids into specific sites of the genome. In some embodiments, the exogenous nucleic acid for insertion may be up to 20 kb, up to 25 kb, up to 30 kb, or up to 40 kb in length, for example, about 1 kb to about 40 kb, about 1 kb to about 39 kb, about 1 kb to about 38 kb, about 1 kb to about 37 kb, about 1 kb to about 36 kb, about 1 kb to about 35 kb, about 1 kb to about 30 kb, about 1 kb to about 30 kb, or about 1 kb to about 25 kb. In some embodiments, the polynucleotide sequence encoding a DNA-binding protein that enables insertion of exogenous nucleic acids into the genome comprises a transposase or a transposase recombinant to a wild-type transposase. The exogenous nucleic acid for insertion may be up to 35kb or up to 40kb in length.
[0134] The transposases or recombinant transposases of this disclosure may be any transposases capable of inserting exogenous nucleic acids into specific sites in the genome. Some aspects of this disclosure provide transposase fusion proteins designed using the methods and strategies described herein. Some embodiments of this disclosure provide nucleic acids encoding such transposases or recombinant transposases and / or fusion proteins comprising them. Some embodiments of this disclosure provide plasmids or expression vectors comprising such nucleic acid constructs encoding transposases or recombinant transposases and / or fusion proteins comprising them.
[0135] Non-limiting examples of transposases include Frog Prince, Sleeping Beauty, High-Activity Sleeping Beauty, PiggyBac, and High-Activity PiggyBac. In some embodiments, the transposase is the High-Activity PiggyBac transposase corresponding to SEQ ID NOs. 9 and 67 (also referred to herein as hyPB or simply PB). In some embodiments, the recombinant transposase comprises one or more recombinants of the High-Activity PiggyBac transposase (SEQ ID NO: 9).
[0136] In some embodiments, the transposase is recombinant high-activity PiggyBac transposase. Recombinant high-activity PiggyBac transposase may contain one or more mutations in amino acids selected from amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, corresponding to the amino acid numbers of SEQ ID NO: 9. Recombinant high-activity PiggyBac mutations may contain one or more amino acid recombinations listed in Table 3. The recombinant high-activity PiggyBac transposase mutations correspond to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 10, specifically R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, and R37 It may contain one or more amino acid recombinations selected from 2A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0137] In some embodiments, the recombinant transposase may contain one or more mutations to hyPB at amino acids 268 and / or 346, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 11, which are involved in the conserved three catalytic residues (e.g., D268N and / or D346N).
[0138] In some embodiments, the recombinant transposase may contain one or more mutations to hyPB that are important for excision at amino acids 287, 287 / 290, and / or 460 / 461, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 12 (e.g., K287A, K287A / K290A, and / or R460A / K461A).
[0139] In some embodiments, the recombinant transposase may contain one or more mutations in hyPB that are involved in target binding at amino acids 351, 356, and / or 379, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 13 (e.g., S351E, S351P, S351A, and / or K356E).
[0140] In some embodiments, the recombinant transposase may contain one or more mutations to hyPB that are important for integration in amino acids 560, 564, 571, 573, 589, 592, and / or 594, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 14 (e.g., T560A, S564P, S571N, S573A, M589V, S592G, and / or F594L).
[0141] In some embodiments, the recombinant transposase may contain one or more mutations to hyPB involved in alignment at amino acids 325, 347, 350, 357 and / or 465, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 15 (e.g., G325A, N347A, N347S, T350A and / or W465A).
[0142] In some embodiments, the recombinant transposase may contain one or more well-conserved mutations for hyPB at amino acids 576 and / or 587, corresponding to the amino acid numbers of, for example, SEQ ID NO: 9 or SEQ ID NO: 16 (e.g., K576A and / or I587A).
[0143] In some embodiments, the recombinant transposase has Zn at amino acid number 586, which corresponds to amino acid number 9 or 17, for example. 2+ It may contain one or more mutations in hyPB that are involved in binding (e.g., H586A).
[0144] In some embodiments, the programmable transposase may contain one or more mutations to hyPB that are involved in incorporation, at amino acid numbers 315, 341, 372, and / or 375, for example, corresponding to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 18 (e.g., R315A, R341A, R372A, and / or K375A).
[0145] In some embodiments, recombinant high-activity PiggyBac comprises an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence described in SEQ ID NO: 9. In some embodiments, recombinant high-activity PiggyBac is selected compared to high-activity PiggyBac due to its high specificity for DNA integration into the genome. In some embodiments, recombinant high-activity PiggyBac comprises an amino acid sequence having one or more of the recombinants disclosed herein to SEQ ID NOs: 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18, each retaining at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the sequence described in SEQ ID NOs: 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18.
[0146] In some embodiments, the highly active PiggyBac transposase is encoded by a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 67. In some embodiments, the SB100 transposase is encoded by a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 68.
[0147] In some embodiments, the PB transposase contains an amino acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 72. In some embodiments, the SB100 transposase contains an amino acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 73.
[0148] In some embodiments, the recombinant transposase is a recombinant Sleeping Beauty transposase containing one or more mutations. In some embodiments, one or more mutations in the highly active Sleeping Beauty transposase or SB100 correspond to L25F, R36A, I42K, G59D, I212K, N245S, K252A, and Q271L of SEQ ID NO: 9 or SEQ ID NO: 73.
[0149] In certain embodiments, the recombinant transposase is not a Himar1C9 mutant.
[0150] Certain aspects of the present disclosure relate to vectors or plasmids (e.g., expression vectors or packaging vectors) comprising nucleic acid constructs containing the transposase or recombinant transposase of the present disclosure, suitable for expression in host cells, such as mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. In some embodiments, the transposase or recombinant transposase is expressed as a fusion protein with Cas9. In some embodiments, the transposase or recombinant transposase is co-expressed with Cas9 from separate vectors but delivered to the same cells. In some embodiments, the transposase or recombinant transposase or a fusion protein containing the same is packaged in lentiviral particles for delivery to cells.
[0151] As shown in Example 20, a newly developed high-activity PiggyBac transposase mutant library can be used to identify recombinant high-activity PiggyBac that perform specific target transposition. Using such a library, recombinant high-activity PiggyBac with aggressive target transposition was identified.
[0152] In some embodiments, recombinant highly active PiggyBac transposase may contain one or more mutations in amino acids selected from amino acids 245, 275, 277, 325, 347, 351, 372, 375, 388, 450, 465, 560, 564, 573, 589, 592, and 594, corresponding to the amino acid numbers of SEQ ID NO: 9.
[0153] In some embodiments, recombinant high-activity PiggyBac mutations may include one or more amino acid recombinations listed in Table 11.
[0154] In some embodiments, the recombinant high-activity PiggyBac transposase mutant may contain one or more amino acid recombinations selected from R245A, R275A, R277A, R275A / R277A, G325A, N347A, N347S, S351E, S351P, S351A, R372A, K375A, R388A, D450N, W465A, T560A, S564P, S573A, M589V, S592G, or F594L, corresponding to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 119.
[0155] In one embodiment, recombinant highly active PiggyBac transposase contains recombinant amino acid D450 corresponding to the amino acid number of SEQ ID NO: 9 or SEQ ID NO: 119.
[0156] In one embodiment, recombinant highly active PiggyBac transposase comprises recombinant amino acids R372A, K375A, and D450 corresponding to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 119.
[0157] In one embodiment, recombinant highly active PiggyBac transposase comprises recombinant amino acids R245A and D450 corresponding to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 119.
[0158] In one embodiment, the recombinant highly active PiggyBac transposase comprises amino acid recombinants R245A, G325A, and S573P corresponding to the amino acid numbers of SEQ ID NO: 9 or SEQ ID NO: 119.
[0159] In one embodiment, the recombinant highly active PiggyBac transposase comprises recombinant amino acids R245A, G325A, D450, and S573P corresponding to the amino acid numbering of SEQ ID NO: 9 or SEQ ID NO: 119.
[0160] As described above, recombinant high-activity PiggyBac transposases are provided herein. Recombinant high-activity PiggyBac transposases can be fused to the elements disclosed herein, but can also be used alone or in combination with different elements. The transposases were produced by the inventors. Thus, recombinant high-activity PiggyBac transposases comprising amino acid sequence SEQ ID NO: 9 are provided, i. The amino acid at position 245 is A. ii. The amino acid at position 275 is either R or A. iii. The amino acid at position 277 is either R or A. iv. The amino acid at position 325 is either A or G. The amino acid at position 347 is either N or A. The amino acid at position vi.351 is E, P, or A. The amino acid at position vii.372 is R, viii. The amino acid at position 375 is A. The amino acid at position ix.450 is either D or N. The amino acid at position x.465 is either W or A. The amino acid at position xi.560 is either T or A. The amino acid at position xii.564 is either P or S. The amino acid at position xiii.573 is either S or A. The amino acid at position xiv.592 is either G or S. The amino acid at position xv.594 is either L or F.
[0161] In some embodiments, recombinant highly active PiggyBac contains an amino acid sequence selected from the group consisting of SEQ ID NOs: 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129.
[0162] In some embodiments, recombinant high-activity PiggyBac comprises an amino acid sequence having one or more of the recombinations disclosed herein with respect to SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129, each retaining at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with respect to the sequences described in SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129. In some embodiments, recombinant high-activity PiggyBac is selected over high-activity PiggyBac due to its high specificity for DNA integration into the genome.
[0163] This disclosure also relates to recombinant, highly active PiggyBac transposases provided herein for use as pharmaceuticals, particularly in gene therapy, exovivo or in vivo.
[0164] [V.CAS9 and zinc finger gene editing] Current genome engineering tools, including genetically modified zinc finger proteins (ZFPs), transcriptional activator-like effector nucleases (TALENs), and more recently, the RNA-induced DNA endonuclease Cas9, result in sequence-specific DNA breaks in the genome. These programmable breaks can lead to DNA mutations at the break site via non-homologous end joining (NHEJ) or DNA replacement around the break site via homologous recombination repair (HDR).
[0165] Certain aspects of this disclosure relate to nucleic acid constructs comprising polynucleotides encoding DNA-binding proteins (e.g., Cas9 and ZFP) engineered to bind to specific genomic DNA sequences. In some embodiments, such DNA-binding proteins are fused to recombinant integrases or recombinant transposases disclosed herein for gene editing.
[0166] [i.Cas9] The CRISPR-Cas9 system is a highly effective tool for inactivating or recombining genes via sequence-specific double-strand breaks (DSBs). These DSBs are recognized by mechanisms that respond to cellular DNA damage and can be repaired by endogenous DSB repair pathways. The primary repair pathway is non-homologous end joining (NHEJ), which often results in small insertions and / or deletions that can lead to frameshift mutations and disrupt gene function. This pathway can be used to create gene knockout mutations. Alternatively, in the presence of a repair template, damage can be seamlessly repaired by homologous recombination repair (HDR). However, despite remarkable progress, HDR-mediated genome editing for introducing precise gene recombination is far less efficient than NHEJ-mediated gene disruption. Furthermore, large multi-kb substitutions via the HDR pathway present challenges and require selection and / or sorting of large populations of cells. Consequently, the primary application of the HDR pathway is localized substitution of critical regions within genes.
[0167] The terms "Cas9" and "Cas9 nuclease" refer to RNA-inducible nucleases containing the Cas9 protein or fragments thereof (e.g., proteins containing the active or inactive DNA cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). Cas9 nucleases are sometimes also called cason1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transmissible elements, and conjugative plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the precursor of the mobile element, and a target entry nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of the crRNA precursor requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. TracrRNA acts as a guide for the processing of crRNA precursors by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the spacer using an endonuclease. Target strands not complementary to the crRNA are first cleaved by an endonuclease and then by a 3'-5' exonuclease. Essentially, DNA binding and cleavage typically require both proteins and both RNAs. However, a single guide RNA ("sgRNA," or simply "gNRA") can be manipulated to incorporate both aspects of crRNA and tracrRNA into a single RNA species.
[0168] Cas9 recognizes short motifs (PAMs or protospacer-adjacent motifs) within CRISPR repeat sequences, helping to distinguish self from non-self. Cas9 nuclease sequences and structures are known to those skilled in the art. Cas9 orthologs have been described in various species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure. Such Cas9 nucleases and sequences include those from organisms and loci disclosed in Chylinski, et al., “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737 (the entire content of which is incorporated herein by reference).
[0169] In some embodiments, the Cas9 nuclease has an inactive (e.g., deactivated) DNA cleavage domain. Nuclease-inactivated Cas9 proteins are sometimes interchangeably referred to as “dCas9” proteins (in contrast to nuclease-“dead”Cas9). Methods for generating Cas9 proteins (or fragments thereof) with an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5):1173-83; the entire contents of each are incorporated herein by reference).
[0170] For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H841A completely inactivate the nuclease activity of S. pyogenes Cas9. Cas9 nickase is a variant of Cas9 nuclease with a point mutation (D10A) in the RuvC nuclease domain, which can introduce nicks into DNA but cannot cleave it.
[0171] The term “Cas9” also includes its variants and functional fragments. In some embodiments, proteins containing fragments of Cas9 are provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9, or (2) the DNA-cleaving domain of Cas9. In some embodiments, the protein containing Cas9 or a fragment of it is called a “Cas9 variant.” A Cas9 variant shares identity with Cas9 or a fragment of it. For example, a Cas9 variant may be at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, a Cas9 variant contains a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA-cleaving domain). The fragment is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or at least approximately 99.9% identical to the corresponding fragment of wild-type Cas9.In some embodiments, Cas9 refers to Cas9 derived from: Corynebacterium ulcerans (NCBI code: NC_015683.1, NC_017317.1) (SEQ ID NO: 19); Corynebacterium diphtheria (NCBI code: NC_016782.1, NC_016786.1) (SEQ ID NO: 20); Spiroplasma syrphidicola (NCBI code: NC_021284.1) (SEQ ID NO: 21); Spiroplasma intermedia (NCBI code: NC_017861.1) (SEQ ID NO: 22); Spiroplasma taiwanense (Spiroplasma taiwanense) (NCBI number: NC_021846.1) (Sequence ID 23); Streptococcus iniae (NCBI number: NC_021314.1) (Sequence ID 24); Belliella baltica (NCBI number: NC_018010.1) (Sequence ID 25); Psychroflexus torquisi (NCBI number: NC_018721.1) (Sequence ID 26); Streptococcus thermophilus (NCBI number: YP_820832.1) (Sequence ID 27); Listeria innocua (NCBI number: NP_472073.1) (Sequence ID 28); Campylobacter jejuni jejuni) (NCBI number: YP_002344900.1) (Sequence ID 29); or Neisseria meningitidis (NCBI number: YP_002342100.1) (Sequence ID 30). In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1) (Sequence ID 31).
[0172] Among known Cas9 proteins, S. pyogenes Cas9 is widely used as a tool in genome engineering. This Cas9 protein is a large multi-domain protein containing two different nuclease domains. Point mutations can be introduced into Cas9 to eliminate its nuclease activity, resulting in an inactive Cas9 (dCas9) that still retains the ability to bind to DNA in a manner programmed by sgRNA. Essentially, when fused to another protein or domain, dCas9 can target virtually any DNA sequence simply by co-expression with the appropriate sgRNA.
[0173] This disclosure provides nucleic acid constructs comprising a polynucleotide encoding a Cas9 protein for insertion of an exogenous nucleic acid into a specific site in the genome. Some aspects of this disclosure provide a fusion protein comprising the Cas9 protein and a recombinant integrase or recombinant transposase of this disclosure. Some embodiments of this disclosure provide nucleic acids encoding such a Cas9 protein or fusion protein. Some embodiments provide plasmids or expression vectors comprising such nucleic acids.
[0174] The Cas9 encoded by the nucleic acid constructs disclosed herein may be any Cas9 capable of binding to specific genomic DNA sequences in a genome. Non-limiting examples of Cas9 proteins include human Cas9 (hCas9), nickase Cas9 (nCas9), inactive Cas9 (dCas9), Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Cas12a, Cas12b, inactive Cas9 (dCas9), variants and their functional fragments. In some embodiments, Cas9 is human Cas9, or a variant or functional fragment thereof.
[0175] In some embodiments, hCas9 is encoded by a nucleic acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to SEQ ID NO: 64. In some embodiments, nCas9 is encoded by a nucleic acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to SEQ ID NO: 65. In some embodiments, dCas9 is encoded by a nucleic acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to sequence number 66.
[0176] In some embodiments, hCas9 contains an amino acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to SEQ ID NO: 69. In some embodiments, nCas9 contains an amino acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to SEQ ID NO: 70. In some embodiments, dCas9 contains an amino acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity with respect to SEQ ID NO: 71.
[0177] Certain aspects of this disclosure relate to vectors or plasmids (e.g., expression vectors or packaging vectors) comprising a nucleic acid construct containing Cas9, suitable for expression in host cells, such as mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. In some embodiments, the nucleic acid construct comprises a polynucleotide sequence encoding Cas9, which is expressed together with the recombinant transposase of this disclosure as a fusion protein.
[0178] [ii. Zinc finger proteins] The Disclosure also provides nucleic acid constructs comprising polynucleotides encoding zinc finger proteins (ZFPs) for insertion of exogenous nucleic acids into specific sites of the genome. Some embodiments of the Disclosure provide fusion proteins comprising ZFPs and recombinant integrases or recombinant transposases of the Disclosure. Some embodiments of the Disclosure provide nucleic acids encoding such ZFPs or fusion proteins. Some embodiments of the Disclosure provide plasmids or expression vectors comprising such encoding nucleic acids.
[0179] As used herein, zinc finger proteins are proteins that can bind to DNA in a sequence-specific manner. ZFPs are unevenly distributed in eukaryotes. ZFPs involved in DNA recognition, RNA binding, and protein binding have been identified. The specific classification of zinc finger proteins is based on the “fold group” in terms of the overall shape of the protein backbone within the folded domain. The most common “fold groups” of zinc fingers are C2H2, or Cys2His2-like (“classical zinc finger”), treble clefts, and zinc ribbons. A representative motif characterizing one class of these proteins (the C2H2 class) is -Cys-(X)2-4-Cys(X)12-His(X)3-5-His (where X is any amino acid).
[0180] The ZFPs of this disclosure may be any ZFP, its variants, or functional fragments that can bind to specific genomic DNA sequences in a genome. Non-limiting examples of ZFPs include ZFPs containing a group of folds or zinc finger motifs selected from C2H2, gag knuckle, treble cleft, zinc ribbon, Zn2 / Cys6-like, or TAZ2 domain-like, or any combination thereof. In some embodiments, the ZFP is a C2H2 zinc finger protein. In some embodiments, the ZFP is an engineered ZFP.
[0181] A manipulated zinc finger array can be fused to a DNA cleavage domain (typically the FokI cleavage domain) to generate a zinc finger nuclease. Such zinc finger-FokI fusions are useful reagents for genome manipulation.
[0182] The ZFPs of this disclosure may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more zinc finger domains. The ZFPs may comprise 2-12, 2-10, 2-8, 3-8, 4-8, or 5-8 zinc finger domains. In some embodiments, the ZFPs comprise six zinc finger domains.
[0183] A typical modular assembly process involves combining separate zinc fingers, each capable of recognizing a 3-base pair DNA sequence, to generate a 3-finger, 4-, 5-, or 6-finger array that recognizes target sites ranging in length from 9 to 18 base pairs. Alternatively, a 2-finger module is used to generate a zinc finger array with up to six individual zinc fingers.
[0184] In some embodiments, the binding domain of the ZFP can be manipulated to bind to a selected sequence. The manipulated zinc finger binding domain may have improved binding specificity compared to naturally occurring ZFPs. In some embodiments, the nucleic acid sequence encoding the ZFP corresponds to SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, or SEQ ID NO: 38. In some embodiments, the amino acid sequence of the ZFP corresponds to SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, or SEQ ID NO: 39. In some embodiments, the ZFP contains an amino acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any of SEQ ID NOs: 33, 35, 37, or 39.
[0185] Certain aspects of this disclosure relate to vectors or plasmids (e.g., expression vectors or packaging vectors) comprising a nucleic acid construct containing a ZFP suitable for expression in host cells, such as mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. In some embodiments, the nucleic acid construct comprises a polynucleotide sequence encoding a ZFP, which is expressed as a fusion protein in co-expression with a recombinant integrase or recombinant transposase of this disclosure.
[0186] [VII. Fusion Proteins] This disclosure provides fusion proteins for site-specific insertion of exogenous nucleic acids into a genome. In certain embodiments, the fusion protein comprises a first DNA-binding protein engineered to bind to a specific genomic DNA sequence, a second DNA-binding protein enabling insertion of the exogenous nucleic acid into the genome (wherein the second DNA-binding protein is an integrase or transposase as disclosed herein), and a linker connecting the first and second proteins. In some embodiments, the first DNA-binding protein is a Cas9 protein or a zinc finger protein. In some embodiments, the first DNA-binding protein is Cas9 and the second binding protein is a recombinant transposase as disclosed herein, and the first and second binding proteins can be oriented in either order within the construct. In some embodiments, the first DNA-binding protein is a zinc finger protein and the second binding protein is a recombinant integrase, and the first and second binding proteins can be oriented in either order within the construct.
[0187] In some embodiments, the fusion protein includes a linker between a first binding protein and a second binding protein. The linker includes (GGS)n, (GGGGS)n (SEQ ID NO: 133), (G)n, (EAAAK)n (SEQ ID NO: 134), XTEN-based, or (XP)n motifs, or any combination thereof, where n is an integer independently between 1 and 50. In some embodiments, the linker is encoded by a nucleic acid sequence that is 12 to 24 amino acids or has a length of 36 to 72 nucleic acids. In some embodiments, the linker includes an XTEN sequence or a GGS sequence. In some embodiments, the fusion protein includes a zinc finger protein linked to the recombinant integrase of the Disclosure, where the linker includes a GGS sequence or an XTEN sequence, and the recombinant integrase may be 5' or 3' relative to the linker. In some embodiments, the fusion protein includes a Cas9 protein linked to the recombinant transposase of the Disclosure. Here, the linker comprises a GGS sequence or an XTEN sequence, and the recombinant transposase may be 5' or 3' relative to the linker. In some embodiments, the linker is the linker shown in Table 1. In some embodiments, the linker comprises the amino acid sequence of SEQ ID NO: 49. In some embodiments, the linker comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 49, SEQ ID NO: 51, SEQ ID NO: 53, SEQ ID NO: 55, SEQ ID NO: 57, SEQ ID NO: 59, SEQ ID NO: 61, SEQ ID NO: 63, or any combination thereof. In some embodiments, the linker is encoded by a nucleic acid sequence comprising SEQ ID NO: 48. In some embodiments, the linker is encoded by a nucleic acid sequence comprising a sequence selected from the group consisting of SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 60, SEQ ID NO: 62, or any combination thereof. [Table 1]
[0188] In some embodiments, the 3' end of the first DNA-binding protein is linked to the 5' end of the second DNA-binding protein by a linker. In some embodiments, the 3' end of the second DNA-binding protein is linked to the 5' end of the first DNA-binding protein by a linker. In some embodiments, the 3' end of the Cas9 protein is linked to the 5' end of the transposase by a linker. In some embodiments, the 5' end of the Cas9 protein is linked to the 3' end of the transposase by a linker. In some embodiments, the 3' zinc finger protein is linked to the 5' end of the integrase by a linker. In some embodiments, the 5' zinc finger protein is linked to the 3' end of the integrase by a linker.
[0189] This specification also provides fusion proteins obtained from the expression of any of the nucleic acid constructs provided herein.
[0190] [VIII. Host cells / organisms] In some embodiments, the nucleic acid constructs of this disclosure are expressed in host cells. Suitable host cells include, but are not limited to, eukaryotic and prokaryotic cells, and / or cell lines. Non-limiting examples of such host cells, or cell lines generated from such cells, include COS, CHO (e.g., CHO-S, CHO-K1, CHO-DG44, CHO-DUXB11, CHO-DUKX, CHOK1SV), VERO, MDCK, WI38, V79, B14AF28-G3, BHK, HaK, NS0, SP2 / 0-Ag14, HeLa, HEK293 (e.g., HEK293-F, HEK293-H, HEK293-T), and perC6 cells, as well as insect cells such as Spodoptera fugiperda (Sf), or fungal cells such as Saccharomyces, Pichia, and Schizosaccharomyces.
[0191] In some embodiments, the host cell is of microbial origin. Microorganisms effective for certain methods disclosed herein include, for example, bacteria (e.g., Escherichia coli), yeast (e.g., Saccharomyces cerevisiae), and plants. The host cell may be prokaryotic or eukaryotic. In some embodiments, the host cell is a eukaryote. Suitable eukaryotic host cells include, but are not limited to, yeast cells, insect cells, plant cells, fungal cells, and algal cells.
[0192] In some embodiments, the host cells are competent host cells. In some embodiments, the host cells are naturally competent. In some embodiments, the host cells are made competent by methods such as using calcium chloride and heat shock. The cells used can be any competent cells, and in particular can be eukaryotic cells, especially mammalian cells, e.g., human or animal cells. The cells used can be somatic or embryonic stem cells or differentiated cells. In some embodiments, the cells include 293T cells, fibroblasts, hepatocytes, muscle cells (skeletal muscle, cardiac muscle, smooth muscle, vascular muscle, etc.), epithelial cells, nerve cells (neurons, glial cells, astrocytes) such as those in the kidney or eye, etc. They may also include insect cells, plant cells, yeast cells, or prokaryotic cells. Furthermore, primary cells may be isolated and used ex vivo for reintroduction into the subject being treated after treatment with a nuclease (e.g., ZFN or TALEN) or a nuclease system (e.g., CRISPR / Cas). Suitable primary cells include peripheral blood mononuclear cells (PBMCs) and other blood cell subsets, including but not limited to T lymphocytes (e.g., CD4+ T cells or CD8+ T cells). Suitable cells also include stem cells, such as embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells (CD34+), neuronal stem cells, and mesenchymal stem cells.
[0193] In some embodiments, host cells are transfected with a plasmid containing a nucleic acid construct disclosed herein. In some embodiments, the plasmid containing the nucleic acid construct is a packaging plasmid. In some embodiments, the plasmid containing the nucleic acid construct further comprises polynucleotides encoding capsid proteins, e.g., gag and pol. In some embodiments, host cells are transfected using: (i) the plasmid containing the nucleic acid construct in the host cell is combined with (ii) a plasmid containing polynucleotides encoding viral envelope proteins (envelope plasmid), and (iii) a plasmid containing an exogenous nucleic acid sequence (e.g., GOI). Here, a viral particle containing the exogenous nucleic acid, e.g., GOI, and a fusion protein containing first and second binding proteins are produced.
[0194] In some embodiments, host cells are transfected using: (i) a plasmid containing a nucleic acid construct; (ii) a plasmid containing a nucleic acid construct further comprising polynucleotides encoding capsid proteins, e.g., gag and pol (a packaging plasmid, the packaging plasmid lacking a functional integrase); (iii) a plasmid containing polynucleotides encoding proteins of the viral envelope (envelope plasmid); and (iv) a plasmid containing an exogenous nucleic acid sequence (e.g., GOI). Here, a fusion protein is produced, comprising a viral particle containing the exogenous nucleic acid, e.g., GOI, and first and second binding proteins.
[0195] In further embodiments, a vector, such as a lentiviral vector according to the Disclosure, can be used to deliver a fusion protein encoded by the nucleic acid construct and exogenous nucleic acid of the Disclosure to an organism, such as a mammal, and more particularly to a target mammalian cell of interest. A lentiviral vector containing the fusion protein of the Disclosure can transduce various cell types, such as liver cells (e.g., hepatocytes), muscle cells, brain cells, kidney cells, retinal cells, and hematopoietic cells. In some embodiments, the target cells of the Disclosure are “non-dividing” cells. These cells include cells such as neurons that do not normally divide. However, the Disclosure is not intended to be limited to non-dividing cells (including, but not limited to, muscle cells, leukocytes, spleen cells, hepatocytes, ophthalmic cells, epithelial cells, etc.).
[0196] In certain embodiments, the packaged fusion protein of this disclosure is administered to an organism for gene editing of the organism's DNA. In some embodiments, the organism is a human. In some embodiments, the organism is a non-human mammal. In some embodiments, the organism is a non-human primate. In some embodiments, the organism is a rodent. In some embodiments, the organism is a sheep, goat, cow, cat, or dog. In some embodiments, the organism is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the organism is a research animal. In some embodiments, the organism is a genetically modified, for example, a genetically modified non-human subject. The organism may be of any sex and at any stage of development.
[0197] [IX. Methods for insertion into the genome] Methods for inserting exogenous nucleic acids into the genome are described (e.g., Yusa et al. PNAS 4(108):1531-1536(2011); Feng et al. Nuc. Acid Res. 4(38):1204-1216(2009); Kettlun et al. Amer. Soc. Gene and Cell Ther. 9(19):1636-1644(2011); Skipper et al. 20(92):1-23(2013); Li et al. PNAS 25:E2279-E2287(2013); Mates et al. Nature Genetics 41(6):753-761(2009); Mali et al. Nat. Methods 10(10):957-963; Vargas et See al. J. Trans. Med. 14(288):1-15 (2016); Gersbach et al. Acc. Chem. Res. 47:2309-2318 (2014); Chandrasegaran et al. Cell Gene Ther. Ins. 3(1):33-41 (2017); Wilson et al. 649:353-363 (2010); Zhao Zhang, et al. Mol Ther Nucleic Acids. 9:230-241 (2017); Naldini L. EMBO Mol Med. 11(3) (2019); and Naldini L, et al. Hum Gene Ther. 27(10):727-728 (2016). Each is incorporated herein by reference.
[0198] This disclosure provides nucleic acid constructs encoding fusion proteins for the insertion of exogenous nucleic acids into specific sites in the genome. The present invention also provides fusion proteins for the insertion of exogenous nucleic acids into specific sites in the genome. In some embodiments, the exogenous nucleic acid for insertion may be up to 5kb, up to 10kb, up to 15kb, up to 20kb, up to 25kb, up to 30kb, up to 35kb, or up to 40kb in length.
[0199] In another embodiment, a method for site-specific nucleic acid insertion into a genome is provided. In some embodiments, the method involves contacting target DNA with one of the fusion proteins comprising Cas9 and a transposase as described herein. For example, in some embodiments, the method involves contacting DNA with a fusion protein comprising two linked polypeptides ((i) Cas9 and (ii) transposase), where active Cas9 binds to gRNA that hybridizes to a region of DNA, such as genomic DNA.
[0200] In some embodiments, the method involves contacting target DNA with one of the fusion proteins comprising Cas9 and integrase as described herein. For example, in some embodiments, the method involves contacting DNA with a fusion protein comprising two linked polypeptides ((i) Cas9 and (ii) integrase), where active Cas9 binds to gRNA that hybridizes to a region of DNA, such as genomic DNA.
[0201] In some embodiments, the method involves contacting target DNA with one of the fusion proteins comprising a ZFP and an integrase as described herein. For example, in some embodiments, the method involves contacting DNA with a fusion protein comprising two linked polypeptides ((i) ZFP and (ii) integrase), where the active ZFP binds to gRNA that hybridizes to a region of DNA, such as genomic DNA.
[0202] In some embodiments, the fusion protein is delivered to an organism and / or cell containing target DNA, such as genomic DNA, using a viral vector, such as a lentiviral particle.
[0203] [X. Lentivirus Packaging] Methods for lentiviral packaging are described (see Grandchamp et al. 9(6):1-13(2014); Voelkel et al. 107(17):7805-7810(2010); Tan et al. 80(4)1939-1948; Li et al. 9(8):1-9(2014); Mates et al. Nature Genetics 41(6):753-761(2009); and Robert H Kutner1, et al. NATURE PROTOCOLS 4(4):495(2009), each incorporated herein by reference).
[0204] Typically, lentiviral delivery systems use a split system that employs different lentiviral genes on separate plasmids used to produce a complete virus that does not contain the genetic elements necessary to cause a viral disease. For example, one plasmid (envelope plasmid) may encode proteins for the viral envelope (env). Another plasmid (packaging plasmid) may encode capsid proteins (e.g., gag and pol) as well as enzymes such as reverse transcriptase and / or integrase. A further plasmid (transfer plasmid) contains the gene of interest (GOI) flanked by long-terminal repeat sequences (for genomic integration) and psi sequences (indicating a signal for packaging the gene within the virus). When these plasmids are introduced into cells simultaneously, a virus is produced that contains a GOI without the viral genes necessary to cause the disease.
[0205] In certain embodiments of this disclosure, the lentiviral vector (or particle) of the present invention can be obtained by transfecting a permissible cell (such as a 293T cell) with a plasmid containing specific components of the lentiviral vector genome and at least one other plasmid, using a split system (e.g., an interspecies complementation system (vector / packaging system)). Herein, the at least one other plasmid provides the gag, pol, and env sequences encoding polypeptides GAG, POL, and envelope proteins in trans, or provides a portion of these polypeptides sufficient to enable the formation of a retroviral particle.
[0206] As an example, host cells are transfected with a) a packaging plasmid containing lentiviral gag and pol sequences, b) a second plasmid (envelope expression plasmid or pseudoenv plasmid) containing a gene encoding an envelope protein (such as VSV-G), c) a plasmid vector containing a 5'-3'LTR sequence, a psi capsidation sequence, and a transgene, and d) a plasmid vector containing a nucleic acid construct encoding an engineered fusion protein disclosed herein. In some embodiments, the nucleic acid construct encoding the engineered fusion protein disclosed herein is located on the packaging plasmid instead of a separate plasmid. Nucleic acids encoding gag, pol, and env cDNA can be advantageously prepared according to the prior art from viral gene sequences available in the prior art and databases.
[0207] In some embodiments, the lentiviral vector comprises a nucleic acid construct as described herein. In some embodiments, the lentiviral vector comprises a fusion protein as described herein.
[0208] The promoters used in the plasmid may be the same or different. In some embodiments, in a plasmid interspecies complementation system, the envelope plasmid and plasmid vector, respectively, may have the same or different promoters for the mRNA and transgene of the vector genome to promote the expression of the coat proteins gag and pol. Such promoters can be advantageously selected from ubiquitous promoters, or specifically, for example, viral promoters such as CMV, TK, RSV LTR promoters, and RNA polymerase III promoters (such as U6 or H1), or promoters of helper viruses encoding env, gag, and pol (i.e., adenovirus, baculovirus, herpesvirus).
[0209] For the production of the lentiviral vectors of this disclosure, the plasmids described herein may be introduced into host cells, where the virus is produced and collected. Suitable cells include, but are not limited to, eukaryotic and prokaryotic cells, and / or cell lines. Non-limiting examples of such cells, or cell lines generated from such cells, include, for example, COS, CHO (e.g., CHO-S, CHO-K1, CHO-DG44, CHO-DUXB11, CHO-DUKX, CHOK1SV), VERO, MDCK, WI38, V79, B14AF28-G3, BHK, HaK, NS0, SP2 / 0-Ag14, HeLa, HEK293 (e.g., HEK293-F, HEK293-H, HEK293-T), and perC6 cells, as well as insect cells such as Spodoptera fugiperda (Sf), or fungal cells such as Saccharomyces, Pichia, and Schizosaccharomyces.
[0210] Once host cells have been transfected with a plasmid and the lentiviral vector (or particles) of the Disclosure have been produced, the lentiviral vector (or particles) of the Disclosure can be purified from the cell supernatant. Purification of the lentiviral vector to increase its concentration can be achieved by any suitable method, such as density gradient purification (e.g., cesium chloride (CsCl)), chromatographic techniques (e.g., column or batch chromatography), or ultracentrifugation. For example, the vector of the Disclosure may be subjected to two or three CsCl concentration gradient purification steps. Preferably, the vector is purified from infected cells using a method comprising lysing the cells, pouring the lysate onto a chromatographic resin, eluting the virus from the chromatographic resin, and collecting a fraction containing the lentiviral vector of the Disclosure.
[0211] [XI. Delivery method] The delivery methods for lentiviral vectors are described (see, for example, Vargas et al. J.Trans.Med. 14(288):1-15 (2016); Mali et al. Nat.Methods 10(10):957-963; Mates et al. Nature Genetics 41(6):753-761 (2009); Skipper et al. 20(92):1-23 (2013)).
[0212] A lentiviral vector comprising a fusion protein encoded by the nucleic acid construct of this disclosure can be administered to a subject by any route. In some embodiments, the lentiviral vector of this disclosure can be delivered to the target cell either in vivo or ex vivo.
[0213] In some embodiments, the lentiviral vectors of the Disclosure may be delivered in vivo. In some embodiments, a lentiviral vector comprising a fusion protein encoded by a nucleic acid construct of the Disclosure may be used to deliver a GOI and / or to target a genetic defect in the DNA of a subject. In some embodiments, the lentiviral vector is administered to a subject parenterally, preferably intravascularly (including intravenously). When administered parenterally, the vector is preferably provided in an injection-suitable medicinal vehicle, such as a sterile aqueous solution or dispersion.
[0214] In some embodiments, the lentiviral vectors of this disclosure may be used exovivoically.
[0215] In some embodiments, lentiviral vectors containing fusion proteins encoded by the nucleic acid constructs of this disclosure may be used to deliver GOIs and / or to target genetic defects in the DNA of a subject. In some embodiments, cells are excised from a subject, and the lentiviral vector containing the fusion protein encoded by the nucleic acid constructs of this disclosure is administered to the cells ex vivo to recombinate the DNA of the cells. The cells, then, are grown to possess the recombinant DNA and reinjected into the subject. In certain embodiments, lentiviral vectors containing fusion proteins encoded by the nucleic acid constructs of this disclosure may be used for chimeric antigen receptor (CAR) T cell therapy to genetically recombinate a patient's autologous T cells to express CARs specific to tumor antigens. In further embodiments, the recombinant CAR-T cells are grown ex vivo and reinjected into the patient. In some embodiments, the modified T cells more specifically target cancer cells. Unlike antibody therapies, CAR-T cells can replicate in vivo and, as a result, are long-lasting.
[0216] After administration of the lentiviral vector of this disclosure or cells recombined ex vivo using the lentiviral vector of this disclosure, the subject can be monitored to detect the expression of the transgene. The therapeutic dose and duration are determined individually depending on the condition or disease being treated. Various conditions or diseases can be treated based on the gene expression produced by the administration of the target gene in the vector of this invention. The dose of the vector delivered using the method of this invention varies depending on the desired response by the host and the vector used.
[0217] In some gene therapy applications, it is desirable that gene therapy vectors be delivered with high specificity to specific tissue types. Therefore, viral vectors can be recombinant to exhibit specificity to a given cell type by expressing a ligand as a fusion protein with the viral coat protein on the outer surface of the virus. The ligand is selected to have affinity for a receptor known to be present on the target cell type.
[0218] Certain aspects of this disclosure relate to a method for inserting an exogenous nucleic acid sequence into the genomic DNA of an organism. The method includes: identifying a specific genomic DNA sequence in the genome of an organism; administering a lentiviral particle containing a nucleic acid construct of this disclosure to the organism to bind to the specific genomic DNA sequence, thereby inserting the exogenous nucleic acid into the genomic DNA; thereby the exogenous nucleic acid being incorporated into the specific genomic DNA sequence.
[0219] Certain aspects of the present disclosure relate to a method for controlled site-directed incorporation of one or more copies of an exogenous nucleic acid sequence into a cell, the method comprising: a) delivering a nucleic acid construct, vector, or fusion protein of the present disclosure into a cell; and b) delivering an exogenous nucleic acid into a cell, wherein the binding of the fusion protein to a specific genomic DNA sequence in the cell's genome causes genomic cleavage and incorporation of one or more copies of the exogenous nucleic acid into the cell's genome. In some embodiments, delivery to the cell is by lentiviral particles.
[0220] [XII. How to use / apply] Several strategies can be used to test the embedded parts and screen for the best mechanism for directional embedding.
[0221] For the analysis of recombinant integrases and transposons disclosed herein, reporter cell lines having a promoter, half of the coding sequence of GFP, and a splicing site donor downstream of the target insertion site in the genome may be used. For example, a lentiviral payload may have a fusion integrase variant, followed by a reverse splicing site acceptor, and the remaining half of GPF. GFP expression occurs when direct insertion takes place and splicing of GFP-containing mRNA generated from the insertion site and the integrated payload results in a complete GFP-CDS.
[0222] The VPR interspecies complementation system can also be used to screen and compare embedded variants. The interspecies complementation system can be used for targeted insertion of lentiviral payloads containing fusion integrase variants. These fusion integrase variants will be loaded into viral particles using VPR fusion, which promotes their own integration when expressed and loaded in particles. This complements the IN variants of the packaging vector used for particle production, which have poor integration encoding. Other methods that can be used for integration mapping include IC or FISH probes. Targeted insertions can also be screened by TCR or RFP target disruption, or by GFP activation through targeted splicing site integration.
[0223] For a FISH approach to simultaneous staining of chromatin insertions and target regions, fluorescence in situ hybridization can be performed to localize GOI transposons in the Hek293T genome. Hek293T can be transfected to PPP1R12 with 1) GOI transposons, 2) programmable transposases, and 3) gRNAs. Probes are designed to target the PPP1R12 gene, the CD46 gene (as a negative control), and GOIs, and can be synthesized using Nick Translation Mix (Sigma) derived from PCR-amplified DNA.
[0224] In some embodiments, the fusion proteins comprising recombinant transposase or recombinant integrase disclosed herein enhance the specificity of insertion of exogenous nucleic acids into the genome compared to fusion proteins comprising the corresponding wild-type protein, as determined, for example, by a gene trap assay. In some embodiments, HEK293T cells, or any other permissible cells, are transfected or transduced with lentiviral particles using the following plasmids or payloads: (i) a plasmid comprising a gRNA targeting a specific site of DNA; (ii) a plasmid comprising a nucleic acid construct of the Disclosure encoding a recombinant transposase fusion protein or recombinant integrase fusion protein; and (iii) a gene trap plasmid comprising a nucleic acid sequence encoding a promoter-less reporter protein (e.g., GFP). In some embodiments, the gene trap plasmid further comprises a transposon having a reverse repeat sequence.
[0225] In some embodiments, the percentage of cells containing GFP insertions can be determined by flow cytometry. In some embodiments, the programmable transposase fusion protein increases the percentage of cells containing GFP insertions by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, or at least 30% compared to the corresponding wild-type protein. In some embodiments, the programmable transposase fusion protein increases the percentage of cells containing GFP insertions by approximately 15–30%.
[0226] In some embodiments, the percentage of insertions at target sites and the percentage of coverage at target sites (number of reads per insertion site) can be determined by genomic DNA extraction and targeted sequencing using oligonucleotides specific to the viral LTR. In some embodiments, recombinant transposase fusion proteins increase the percentage of insertions at target sites by at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold compared to the corresponding wild-type protein. In some embodiments, the percentage of insertions at target sites increases by approximately 10 to 100-fold. In some embodiments, the recombinant transposase fusion protein increases the percentage of coverage at the target site (number of reads per insertion site) by at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least 110-fold, at least 120-fold, at least 130-fold, at least 140-fold, at least 150-fold, at least 160-fold, at least 170-fold, at least 180-fold, at least 190-fold, or at least 200-fold compared to the corresponding wild-type protein. In some embodiments, the percentage of coverage at the target site (number of reads per insertion site) is at least 100-fold.
[0227] In some embodiments, the recombinant integrase fusion protein, when quantified by GFP incorporation, enhances the specificity of the insertion of the exogenous nucleic acid into the genome compared to the corresponding wild-type protein. In some embodiments, lentiviruses containing recombinant integrase fusion proteins were generated by transfecting HEK293T cells or any other permissible cells with: (i) a plasmid containing a nucleic acid sequence encoding GFP, (ii) a plasmid containing a packaging protein, (iii) a plasmid containing an envelope protein, and (iv) a plasmid containing a nucleic acid construct encoding the recombinant integrase fusion protein. The supernatant containing the lentivirus was collected 48 hours after transfection.
[0228] For targeted insertion, HEK293T cells were infected with a lentivirus containing a recombinant integrase fusion protein. In some embodiments, the percentage of GFP-positive cells was quantified by flow cytometry on days 3, 5, 7, 10, and 12 post-infection. In some embodiments, the recombinant integrase fusion protein increased the percentage of cells containing GFP insertions by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, or at least 30% compared to the corresponding wild-type protein.
[0229] In some embodiments, the percentage of insertion at the target site and the percentage of coverage at the target site (number of reads per insertion site) can be determined by genomic DNA extraction and targeted sequencing using oligonucleotides specific to the viral insertion LTR. In some embodiments, the recombinant integrase fusion protein increases the percentage of insertion at the target site by at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold compared to the corresponding wild-type protein. In some embodiments, the recombinant integrase fusion protein increases the percentage of coverage at the target site (number of reads per insertion site) by at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least 110-fold, at least 120-fold, at least 130-fold, at least 140-fold, at least 150-fold, at least 160-fold, at least 170-fold, at least 180-fold, at least 190-fold, or at least 200-fold compared to the corresponding wild-type protein.
[0230] Possible applications of the lentiviral vectors containing the fusion proteins of this disclosure include gene therapy, i.e., gene transfer in any mammalian cells, particularly human cells. The cells may be dividing cells or quiescent cells, or cells belonging to central or peripheral organs such as the liver, pancreas, muscle, or heart. Gene therapy may enable the expression of proteins such as neurotrophic factors, enzymes, transcription factors, and receptors. The lentiviral vectors according to the present invention may also be particularly suitable for research purposes.
[0231] In some embodiments, the nucleic acid constructs, fusion proteins, and / or lentiviral vectors of this disclosure are administered to a subject to treat a disease. In some embodiments, the disease is a genetic disorder that may benefit from gene therapy.
[0232] In some embodiments, lentiviral vectors comprising the fusion protein according to the Disclosure may be used as pharmaceuticals. Lentiviral vectors according to the Disclosure may be particularly suitable for treating genetic disorders in subjects.
[0233] [XIII. Compositions and Kits] This disclosure also provides compositions for carrying out the methods of disclosure described herein. In some embodiments, the compositions include a nucleic acid construct or vector as defined herein, and a polynucleotide sequence encoding an exogenous nucleic acid for insertion into a genome, contained in or bound to a packaging vector.
[0234] In some embodiments, depending on the delivery method, the nucleic acid construct is in the form of RNA, DNA, or protein, and the polynucleotide sequence encoding the exogenous nucleic acid is in the form of RNA or DNA. In particular, the polynucleotide sequence encoding the exogenous nucleic acid is in the form of RNA.
[0235] In some embodiments, the composition does not contain viruses, and the packaging vector is nanoparticles, such as polymer nanoparticles or lipid nanoparticles. The packaging vector may also be a carrier bound to the elements of the composition. In some embodiments, the composition contains viral vectors, particularly lentiviral particles.
[0236] In some embodiments, the composition comprises (a) a nucleic acid construct in RNA form as described herein (e.g., including Cas9 and a transposase), (b) a guide RNA if necessary (e.g., as a separate linear single-stranded RNA molecule), and (c) a polynucleotide containing an exogenous gene for insertion into DNA form (e.g., a vector), which is contained in or bound to a packaging vector.
[0237] In some embodiments, the composition comprises (a) a fusion protein in protein form as described herein (e.g., including Cas9 and a transposase), (b) a guide RNA if necessary (e.g., as a separate linear single-stranded RNA molecule) (where the fusion protein and guide RNA form a ribonucleoprotein complex (RNP)), and (c) a polynucleotide containing an exogenous gene for insertion into DNA form (e.g., a vector), contained in or bound to a packaging vector.
[0238] In some embodiments, the composition comprises (a) a nucleic acid construct in DNA form as described herein (e.g., including Cas9 and a transposase), (b) guide RNA if necessary (e.g., as a separate linear RNA molecule or as DNA in the vector), and (c) a polynucleotide containing an exogenous gene for insertion of the DNA form (e.g., the vector), which is contained in or bound to a packaging vector.
[0239] In some embodiments, the composition comprises (a) a fusion protein described herein in protein form (e.g., comprising Cas9 and integrase), (b) guide RNA if necessary (e.g., as a separate RNA molecule that complexes with the fusion protein), and (c) a polynucleotide containing an exogenous gene for insertion, contained in or bound to a packaging vector. In certain embodiments, the packaging vector is a lentiviral particle. In some embodiments, (a) the fusion protein is bound to a lentiviral capsid by gag-pol or VPR (viral protein R). In some embodiments, (c) the polynucleotide is in RNA form as the payload for integrase.
[0240] In certain embodiments, when ZFPs are used, (b) guide RNA may not be required.
[0241] Kits for carrying out the methods disclosed herein are also provided herein. The kits may include nucleic acid constructs or fusion proteins described herein. In some embodiments, the kits may include lentiviral particles containing nucleic acid constructs or fusion proteins described herein.
[0242] The kit may further include instructions for using the components of the kit to carry out the method. Instructions for carrying out the method are generally recorded on a suitable recording medium. For example, instructions can be printed on a substrate such as paper or plastic. Thus, instructions may be present in the kit as an accompanying document, on a label of the kit or its components (i.e., related to packaging or subpackaging), etc. In other embodiments, instructions exist as an electronic storage data file located on a suitable computer-readable storage medium such as a CD-ROM or diskette. In yet another embodiment, the actual instructions are not present in the kit, but there is a means for obtaining the instructions from a remote source, for example, via the internet. An example of this embodiment is a kit that includes a web address from which instructions can be viewed and / or downloaded. Similar to the instructions, this means for obtaining the instructions is recorded on a suitable substrate.
[0243] [XIV. Embodiments] E1. A nucleic acid construct, a) A first polynucleotide sequence encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in the genome, b) A second polynucleotide sequence encoding a second DNA-binding protein that enables the insertion of exogenous nucleic acids into the genome, wherein the second DNA-binding protein is (i) an integrase recombinantly adapted for wild-type integrase, or (ii) a transposase recombinantly adapted for wild-type transposase, and c) A third polynucleotide sequence comprising a nucleic acid encoding a linker, Here, the nucleic acid construct encodes a fusion protein comprising a first DNA-binding protein, a second DNA-binding protein, and a linker between the first and second DNA-binding proteins. Nucleic acid constructs.
[0244] E2. The nucleic acid construct of Embodiment E1, wherein the second DNA-binding protein is recombinant to improve the specificity of insertion of exogenous nucleic acids into the genome compared to the corresponding wild-type protein.
[0245] E3. The exogenous nucleic acid for insertion is a nucleic acid construct of Embodiment E1 or E2, which may be up to approximately 20 kb in length.
[0246] E4. A nucleic acid construct of either Embodiment E1 or E3, wherein the first polynucleotide sequence encodes a protein selected from the group consisting of zinc finger proteins, Cas9 proteins, and any variant or functional fragment thereof.
[0247] E5. The Cas9 protein is a nucleic acid construct of Embodiment E4, selected from the group consisting of human Cas9, nickase Cas9, Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Cas12a, Cas12b, and inactive Cas9.
[0248] E6. The zinc finger protein is a C2H2 zinc finger protein, the nucleic acid construct of Embodiment E4.
[0249] E7. Recombinant integrase is a nucleic acid construct from any one of embodiments E1 to E6, which is recombinant human immunodeficiency virus (HIV) integrase or a functional fragment thereof.
[0250] E8. Recombinant HIV integrase is a nucleic acid construct of Embodiment E7, comprising one or more mutations from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1).
[0251] E9. Recombinant HIV integrase mutations correspond to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1), namely D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N A nucleic acid construct of Embodiment E8, comprising one or more of 120E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R.
[0252] E10. A recombinant HIV integrase comprising a nucleic acid construct of any one of Embodiments E7 to E9, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in Sequence ID No. 3.
[0253] E11. Recombinant transposase is a nucleic acid construct from any one of embodiments E1 to E6, selected from the group consisting of recombinant Frog Prince, recombinant Sleeping Beauty, recombinant high-activity Sleeping Beauty (SB100X), recombinant PiggyBac, recombinant high-activity PiggyBac, and functional fragments thereof.
[0254] E12. Recombinant transposase is recombinant high-activity PiggyBac or a functional fragment thereof, which is the nucleic acid construct of Embodiment E11.
[0255] E13. Recombinant high-activity PiggyBac is a nucleic acid construct of Embodiment E12, comprising one or more mutations from among amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, corresponding to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9).
[0256] E14. Recombinant high-activity PiggyBac mutations correspond to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9): R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, R A nucleic acid construct of Embodiment E13 comprising one or more of 372A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0257] E15. Recombinant highly active PiggyBac is a nucleic acid construct of any one of embodiments E12 to E14, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in SEQ ID NO: 10.
[0258] E16. The linker is a nucleic acid construct from any one of embodiments E1 to E15, comprising either an XTEN sequence or a GGS sequence.
[0259] E17. The linker encoding sequence is one nucleic acid construct from any one of embodiments E1 to E16, having a length of approximately 9 to approximately 150 nucleic acids.
[0260] E18. A nucleic acid construct of any one of embodiments E1 to E17, wherein the 3' end of a first polynucleotide sequence is connected to the 5' end of a second polynucleotide by a nucleic acid linker.
[0261] E19. A nucleic acid construct of any one of embodiments E1 to E17, wherein the 3' end of a second polynucleotide sequence is connected to the 5' end of a first polynucleotide sequence by a nucleic acid linker.
[0262] E20. A vector comprising one nucleic acid construct from any one of Embodiments E1 to E19, which is an expression vector suitable for expression in mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells.
[0263] E21. a) The first polynucleotide sequence encodes the Cas9 protein, and b) The second polynucleotide sequence encodes recombinant highly active PiggyBac recombinant transposase, or a functional fragment thereof. Nucleic acid construct of embodiment E1.
[0264] E22. The Cas9 protein is a nucleic acid construct of Embodiment E21, selected from the group consisting of human Cas9, nickase Cas9, Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Cas12a, Cas12b, and inactive Cas9.
[0265] E23. Recombinant high-activity PiggyBac is a nucleic acid construct of either Embodiment E21 or E22, comprising one or more mutations from among amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, corresponding to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9).
[0266] E24. Recombinant high-activity PiggyBac mutations correspond to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9): R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, R A nucleic acid construct of Embodiment E23 comprising one or more of 372A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0267] E25. Recombinant highly active PiggyBac is a nucleic acid construct of either Embodiment E21 or E22, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in Sequence ID No. 10.
[0268] E26. The nucleic acid encoding the linker is one of the nucleic acid constructs from Embodiments E21 to E25, comprising an XTEN sequence or a GGS sequence.
[0269] E27. The linker encoding sequence is one nucleic acid construct from any one of embodiments E21 to E26, having a length of 9 to 150 nucleic acids.
[0270] E28. A nucleic acid construct of any one of embodiments E22-E27, wherein the 3' end of a second polynucleotide sequence is connected to the 5' end of a first polynucleotide sequence by a linker.
[0271] E29. a) The first polynucleotide sequence encodes a zinc finger protein, and b) The second polynucleotide sequence encodes recombinant integrase or a functional fragment thereof. Nucleic acid construct of embodiment E1.
[0272] E30. The zinc finger protein is a C2H2 zinc finger protein, a nucleic acid construct of Embodiment E29.
[0273] E31. Recombinant integrase is recombinant human immunodeficiency virus (HIV) integrase, or a nucleic acid construct of either Embodiment E29 or E30, which is a functional fragment thereof.
[0274] E32. Recombinant HIV integrase is a nucleic acid construct of Embodiment E31, comprising one or more mutations from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1).
[0275] Embodiment E33. The nucleic acid construct of Embodiment E32, wherein the recombinant HIV integrase mutation comprises one or more of D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N120E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R, corresponding to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1).
[0276] Embodiment E34. The nucleic acid construct of any one of Embodiments E31 - E33, wherein the recombinant HIV integrase comprises an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO: 3.
[0277] Embodiment E35. The nucleic acid construct of any one of Embodiments E29 - E34, wherein the linker comprises an XTEN sequence or a GGS sequence.
[0278] Embodiment E36. The nucleic acid construct of any one of Embodiments E29 - E35, wherein the sequence encoding the linker is 9 - 150 nucleic acids in length.
[0279] Embodiment E37. The nucleic acid construct of any one of Embodiments E29 - E37, wherein the 3' end of the second polynucleotide sequence is connected to the 5' end of the first polynucleotide sequence by a linker.
[0280] Embodiment E38. A vector comprising the nucleic acid construct of any one of Embodiments E21 - E37, which is an expression vector suitable for expression in mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells.
[0281] E39. A host cell containing any one nucleic acid construct or vector from Embodiments E1 to E38.
[0282] E40. Fusion protein: A first DNA-binding protein engineered to bind to specific genomic DNA sequences within the genome; A second DNA-binding protein that enables the insertion of exogenous nucleic acids into the genome, wherein the second DNA-binding protein is an integrase or transposase that is recombinant relative to the wild type; and, A linker that connects the first protein and the second protein. A fusion protein containing [the specified ingredient].
[0283] E41. A fusion protein of Embodiment E40, wherein the second DNA-binding protein is recombinant to improve the specificity of insertion of exogenous nucleic acids into the genome compared to the corresponding wild-type protein.
[0284] E42. The exogenous nucleic acid is a fusion protein of either Embodiment E40 or E41, which may be up to approximately 20 kb in length.
[0285] E43. The first DNA-binding protein is a fusion protein selected from the group consisting of zinc finger protein, Cas9 protein, and any variant or functional fragment thereof, as in any one of embodiments E40 to E42.
[0286] E44.Cas9 protein is a fusion protein of Embodiment E43, selected from the group consisting of human Cas9, nickase Cas9, Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Cas12a, Cas12b, and inactive Cas9.
[0287] E45. The zinc finger protein is a C2H2 zinc finger protein, which is the fusion protein of Embodiment E43.
[0288] E46. Recombinant integrase is a fusion protein of any one of embodiments E40-E45, which is recombinant human immunodeficiency virus (HIV) integrase or a functional fragment thereof.
[0289] E47. Recombinant HIV integrase is a fusion protein of Embodiment E46, comprising one or more mutations from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1).
[0290] E48. Recombinant HIV integrase mutations correspond to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1), namely D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N1 A fusion protein of Embodiment E47 comprising one or more of 20E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R.
[0291] E49. Recombinant HIV integrase is a fusion protein of any one of embodiments E46 to E48, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in SEQ ID NO: 3.
[0292] E50. Recombinant transposase is a fusion protein of any one of embodiments E40 to E45, selected from the group consisting of recombinant Frog Prince, recombinant Sleeping Beauty, recombinant high-activity Sleeping Beauty (SB100X), recombinant PiggyBac, recombinant high-activity PiggyBac, and their functional fragments.
[0293] E51. Recombinant transposase is a fusion protein of embodiment E50, which is recombinant high-activity PiggyBac or a functional fragment thereof.
[0294] E52. Recombinant high-activity PiggyBac is a fusion protein of Embodiment E51, comprising one or more mutations from among amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, corresponding to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9).
[0295] E53. Recombinant high-activity PiggyBac mutations correspond to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9): R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, R3 A fusion protein of Embodiment E52 comprising one or more of the following: 72A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0296] The recombinant highly active PiggyBac is a fusion protein according to any one of Embodiments E50 to E53, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO: 10.
[0297] E55. The linker is a fusion protein according to any one of Embodiments E40 to E54, comprising an XTEN sequence or a GGS sequence.
[0298] E56. The linker is a fusion protein according to any one of Embodiments E40 to E55, having a length of 3 to 50 amino acids.
[0299] E57. a) The first DNA-binding protein is a Cas9 protein, b) The second DNA-binding protein is a recombinant highly active PiggyBac or a functional fragment thereof, The fusion protein of Embodiment E40.
[0300] E58. The Cas9 protein is selected from the group consisting of human Cas9, nickase Cas9, Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Cas12a, Cas12b, and inactive Cas9, and is the fusion protein of Embodiment E57.
[0301] E59. The recombinant highly active PiggyBac contains one or more mutations among amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, which correspond to the amino acid numbers of the highly active PiggyBac sequence (SEQ ID NO: 9), and is the fusion protein of any one of Embodiments E57 or E58.
[0302] E60. Recombinant high-activity PiggyBac mutations correspond to the amino acid numbers of the high-activity PiggyBac sequence (SEQ ID NO: 9): R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, R3 A fusion protein of Embodiment E59 comprising one or more of the following: 72A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0303] E61. Recombinant highly active PiggyBac is a fusion protein of any one of embodiments E57 to E60, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in SEQ ID NO: 10.
[0304] E62. a) The first DNA-binding protein is a zinc finger protein, b) The second DNA-binding protein is recombinant integrase or a functional fragment thereof. Fusion protein of embodiment E40.
[0305] E63. The zinc finger protein is a C2H2 zinc finger protein, which is the fusion protein of embodiment E62.
[0306] E64. Recombinant integrase is a fusion protein of either recombinant human immunodeficiency virus (HIV) integrase or a functional fragment thereof, as defined in Embodiments E62 or E63.
[0307] E65. Recombinant HIV integrase is a fusion protein of Embodiment E64, comprising one or more mutations from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1).
[0308] E66. Recombinant HIV integrase mutations correspond to the amino acid numbers of the wild-type HIV integrase sequence (SEQ ID NO: 1), namely D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N1 A fusion protein of Embodiment E65 comprising one or more of 20E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R.
[0309] E67. The recombinant HIV integrase is a fusion protein of Embodiment E62, comprising an amino acid sequence that is at least 85%, at least 90%, or at least 95% identical to the sequence described in SEQ ID NO: 3.
[0310] E68. The linker is a fusion protein of any one of embodiments E57 to E67, comprising either an XTEN sequence or a GGS sequence.
[0311] E69. The linker is a fusion protein of any one of embodiments E57-E68, having a length of 3-50 amino acids.
[0312] E70. A fusion protein of any one of embodiments E40-E69, wherein the 3' end of a second DNA-binding protein is linked to the 5' end of a first DNA-binding protein by a linker.
[0313] E71. A lentiviral particle containing one of the fusion proteins from Embodiments E40 to E69.
[0314] E72. a) A polynucleotide comprising any one nucleic acid construct of Embodiments E1 to E38, and b) Polynucleotides encoding proteins in the envelope of lentiviruses, A method for producing lentiviral particles for gene editing, comprising expressing a gene in a host cell.
[0315] E73.c) The method of Embodiment E72, further comprising expressing a polynucleotide sequence containing an exogenous nucleic acid.
[0316] E74. A polynucleotide comprising a nucleic acid construct further comprising a nucleic acid sequence encoding a lentiviral capsid protein, according to one embodiment of E72 or E73.
[0317] E75. Any one of embodiments E72-E74, further comprising recovering lentiviral particles from host cells.
[0318] E76. Any one of embodiments E72 to E75, further comprising purifying lentivirus particles.
[0319] E77. A method for inserting an exogenous nucleic acid sequence into the genomic DNA of an organism, the method comprising administering a lentiviral particle containing a nucleic acid construct of any of Embodiments E1 to E38 or a fusion protein of any of Embodiments E40 to E71 to an organism to cause first and second DNA-binding proteins to bind to a specific genomic DNA sequence, thereby inserting the exogenous nucleic acid into the genomic DNA, wherein the exogenous nucleic acid becomes incorporated into the specific genomic DNA sequence.
[0320] E78. A method for controlled site-specific integration of a single or multiple copies of an exogenous nucleic acid sequence into a cell, wherein the method is: a) Delivering one of the fusion proteins from Embodiments E40 to E71 to a cell, and b) Delivering exogenous nucleic acids to cells, Here, the binding of a fusion protein to a specific genomic DNA sequence in the cell's genome causes genomic cleavage and the incorporation of one or more copies of exogenous nucleic acids into the cell's genome, and the fusion protein is delivered to the cell by lentiviral particles.
[0321] E79. Nucleic acid constructs containing the following:
[0322] a) A first polynucleotide sequence comprising a nucleic acid encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in the genome; where the first DNA-binding protein is a zinc finger protein or a Cas9 protein.
[0323] b) A second polynucleotide sequence comprising a nucleic acid encoding a second DNA-binding protein that enables the insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein is (i) Highly active PiggyBac transposase, or recombinant highly active PiggyBac with improved specificity for insertion of exogenous nucleic acids into the genome compared to highly active PiggyBac, or (ii) Human immunodeficiency virus (HIV) integrase, or recombinant HIV integrase in which the specificity of insertion of exogenous nucleic acids into the genome is improved compared to HIV integrase.
[0324] c) Any polynucleotide sequence containing the nucleic acid that codes for the linker.
[0325] Here, the nucleic acid construct encodes a fusion protein comprising a first DNA-binding protein, a second DNA-binding protein, and an optional linker between the first and second DNA-binding proteins.
[0326] Here, the fusion protein enables the insertion of exogenous nucleic acids into specific locations in the genome.
[0327] E80.Cas9 protein is a nucleic acid construct of Embodiment E79, selected from the group consisting of human Cas9, nickase Cas9, and inactive Cas9.
[0328] E81. The zinc finger protein is a C2H2 zinc finger protein containing six domains, a nucleic acid construct of Embodiment E79.
[0329] E82. The linker is a nucleic acid construct from any one of embodiments E79 to E81, comprising either an XTEN sequence or a GGS sequence.
[0330] E83. A nucleic acid construct according to any one of embodiments E79-E82, wherein the 3' end of a first polynucleotide sequence is ligated to the 5' end of a second polynucleotide.
[0331] E84.A nucleic acid construct of any one of Embodiments E79-E83, where (a) the first DNA-binding protein is a Cas9 protein or a zinc finger protein, and (b) the second DNA-binding protein is a highly active PiggyBac transposase or recombinant highly active PiggyBac with improved specificity for insertion of exogenous nucleic acids into the genome compared to highly active PiggyBac, wherein the nucleic acid construct comprises (c) a polynucleotide sequence comprising a nucleic acid encoding a linker comprising an XTEN sequence or a GGS sequence, the 3' end of the first polynucleotide sequence being ligated to the 5' end of the second polynucleotide.
[0332] E85.A nucleic acid construct of any one of Embodiments E79-E83, wherein (a) the first DNA-binding protein is a Cas9 protein or a zinc finger protein, and (b) the second DNA-binding protein is HIV integrase or recombinant HIV integrase with improved specificity for insertion of exogenous nucleic acids into the genome compared to HIV integrase, wherein the nucleic acid construct comprises (c) a polynucleotide sequence comprising a nucleic acid encoding a linker comprising an XTEN sequence or a GGS sequence, the 3' end of the first polynucleotide sequence being ligated to the 5' end of the second polynucleotide.
[0333] E86. Recombinant highly active PiggyBac transposase is a nucleic acid construct of any one of Embodiments E79 to E84, comprising one or more mutations from amino acids 245, 268, 275, 277, 287, 290, 315, 325, 341, 346, 347, 350, 351, 356, 357, 372, 375, 388, 409, 412, 432, 447, 450, 460, 461, 465, 517, 560, 564, 571, 573, 576, 586, 587, 589, 592, and 594, corresponding to the amino acid sequence SEQ ID NO: 9 of highly active PiggyBac.
[0334] E87. Recombinant high-activity PiggyBac transposase mutations correspond to the amino acid numbers in SEQ ID NO: 9 of the amino acid sequence of high-activity PiggyBac: R245A, D268N, R275A / R277A, K287A, K290A, K287A / K290A, R315A, G325A, R341A, D346N, N347A, N347S, T350A, S351E, S351P, S351A, K356E, N357A, R372A, K375A, R A nucleic acid construct of Embodiment E86 comprising one of the amino acid recombinations selected from 372A / K375A, R388A, K409A, K412A, K409A / K412A, K432A, D447A, D447N, D450N, R460A, K461A, R460A / K461A, W465A, S517A, T560A, S564P, S571N, S573A, K576A, H586A, I587A, M589V, S592G, or F594L.
[0335] E88. Recombinant highly active PiggyBac transposase is a nucleic acid construct of any one of embodiments E79 to E84, comprising one or more mutations from among amino acids 245, 275, 277, 325, 347, 351, 372, 375, 388, 450, 465, 560, 564, 573, 589, 592, and 594, corresponding to the amino acid sequence number 9 of highly active PiggyBac.
[0336] E89. Recombinant high-activity PiggyBac transposase mutation is a nucleic acid construct of Embodiment E88, comprising one or more amino acid recombinations selected from R245A, R275A, R277A, R275A / R277A, G325A, N347A, N347S, S351E, S351P, S351A, R372A, K375A, R388A, D450N, W465A, T560A, S564P, S573A, M589V, S592G, or F594L, corresponding to amino acid sequence number 9 of high-activity PiggyBac.
[0337] E90. Recombinant highly active PiggyBac transposase comprising amino acid sequence number 9, where the amino acid at position 245 is A, the amino acid at position 275 is R or A, the amino acid at position 277 is R or A, the amino acid at position 325 is A or G, the amino acid at position 347 is N or A, the amino acid at position 351 is E, P or A, the amino acid at position 372 is R, the amino acid at position 375 is A, the amino acid at position 450 is D or N, the amino acid at position 465 is W or A, the amino acid at position 560 is T or A, the amino acid at position 564 is P or S, the amino acid at position 573 is S or A, the amino acid at position 592 is G or S, and the amino acid at position 594 is L or F, a nucleic acid construct of Embodiment E88.
[0338] E91. Recombinant highly active PiggyBac transposase is a nucleic acid construct of Embodiment E88 comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129.
[0339] E92. Recombinant high-activity PiggyBac transposase comprises an amino acid sequence that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129, wherein recombinant high-activity PiggyBac exhibits higher specificity for DNA integration into the genome compared to high-activity PiggyBac, as a nucleic acid construct of Embodiment E88.
[0340] E93. Recombinant HIV integrase is a nucleic acid construct of any one of embodiments E79-E83 or E85, comprising one or more mutations from amino acids 10, 13, 64, 94, 116, 117, 119, 120, 122, 124, 128, 152, 168, 170, 185, 231, 264, 266, or 273, corresponding to amino acid sequence number 1 of wild-type HIV integrase.
[0341] E94. Recombinant HIV integrase mutations correspond to the amino acid sequence SEQ ID NO: D10K, E13K, D64A, D64E, G94D, G94E, G94R, G94K, D116A, D116E, N117D, N117E, N117R, N117K, S119A, S119P, S119T, S119G, S119D, S119E, S119R, S119K, N120D, N A nucleic acid construct of E93 containing one or more of the following: 120E, N120R, N120K, T122K, T122I, T122V, T122A, T122R, A124D, A124E, A124R, A124K, A128T, E152A, E152D, Q168L, Q168A, E170G, F185K, R231G, R231K, R231D, R231E, R231S, K264R, K266R, or K273R.
[0342] E95. A vector comprising any one nucleic acid construct of Embodiments E79 to E95, wherein the vector is suitable for expression in mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells.
[0343] E96. A host cell comprising any one nucleic acid construct or vector from Embodiments E79 to E95.
[0344] E97. A fusion protein obtained from the expression of any one nucleic acid construct from Embodiments E79 to E94.
[0345] E98. A composition comprising one nucleic acid construct, vector or fusion protein according to Embodiments E79-E95 or E97, and a polynucleotide sequence encoding an exogenous nucleic acid for insertion into a genome, wherein the composition is contained in or conjugated to a packaging vector.
[0346] E99. The composition of Embodiment E98, wherein the nucleic acid construct is in the form of RNA, DNA, or protein, and the polynucleotide sequence encoding the exogenous nucleic acid is in the form of DNA or RNA.
[0347] E100. A composition of any one of embodiments E98 to E99, wherein the packaging vector is a nanoparticle or a lentiviral particle.
[0348] E101. A method for controlled site-directed incorporation of one or more copies of an exogenous nucleic acid sequence into a cell, the method comprising (a) delivering a nucleic acid construct, vector or fusion protein from any one of embodiments E79-E95 or E97 to a cell, and (b) delivering an exogenous nucleic acid to a cell, wherein the binding of the fusion protein to a specific genomic DNA sequence in the cell's genome causes genomic cleavage and incorporation of one or more copies of the exogenous nucleic acid into the cell's genome.
[0349] E102. Recombinant highly active PiggyBac transposase containing amino acid sequence number 9, where the amino acid at position 245 is A, the amino acid at position 275 is R or A, the amino acid at position 277 is R or A, the amino acid at position 325 is A or G, the amino acid at position 347 is N or A, the amino acid at position 351 is E, P or A, the amino acid at position 372 is R, the amino acid at position 375 is A, the amino acid at position 450 is D or N, the amino acid at position 465 is W or A, the amino acid at position 560 is T or A, the amino acid at position 564 is P or S, the amino acid at position 573 is S or A, the amino acid at position 592 is G or S, and the amino acid at position 594 is L or F.
[0350] E103. Recombinant highly active PiggyBac transposase of Embodiment E102, comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129.
[0351] E104. Recombinant highly active PiggyBac transposase of claim E012, comprising an amino acid sequence that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128 and 129, wherein recombinant highly active PiggyBac exhibits higher specificity for DNA integration into the genome compared with highly active PiggyBac.
[0352] The content of all cited references (including documents, patents, patent applications, and websites) that may be cited throughout this application is expressly incorporated by reference in their entirety for any purpose, as is the reference cited herein. The following examples are provided as illustrations and not as limitations.
[0353] [Examples] "PB" and "hyPB" are used interchangeably to indicate highly active PiggyBac transposases. Examples 1-3 below relate to the generation and performance of programmable transposase and Cas9 fusion protein constructs with respect to targeted integration. In Example 1, we successfully created various DNA constructs by fusing the highly active transposases PiggyBac and Sleeping Beauty with Cas9, and were able to integrate them into the genomes of transposon-transfected cells. Notably, the PiggyBac and Cas9 constructs were able to facilitate targeted integration into desired sites in the genome (Example 2). Example 3 provides a modified transposase generated to increase the specificity of exogenous nucleic acid sequences into the genome. Example 1: DNA vector for expression of programmable transposase fusion protein
[0354] This experiment aims to test different configurations of fusions of highly active PiggyBac transposase (referred to herein as hyPB or PB) and Sleeping Beauty (referred to herein as SB100x) to nuclease (h), nickase (n), and inactive (d) Cas9 for transposon integration performance. Programmable transposase fusion proteins were constructed by incorporating DNA sequences encoding wild-type human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) (sequence codes 64-66, respectively) and highly active PiggyBac (PB) or highly active Sleeping Beauty (SB100) transposase (sequence codes 67-68, respectively) into a pcDNA3.3-TOPO expression vector (Invitrogen plasmid backbone, Addgene plasmid #41815). Vectors were constructed by ligating the 3' end of Cas9 to the 5' end of each transposase using nucleic acid linker sequences (SEQ ID NO: 48) encoding GGS linkers (hCas9PB, nCas9PB, dCas9PB, hCas9SB, nCas9SB, and dCas9SB). Other vectors were constructed by ligating the 3' end of each transposase to the 5' end of Cas9 using nucleic acid linker sequences (SEQ ID NO: 48) encoding GGS linkers (PBhCas9, PBnCas9, PBdCas9, SBhCas9, SBnCas9, and SBdCas9). A summary of the fusion constructs is shown in Table 2. [Table 2]
[0355] Prior to transfection, frozen HEK293T cells were rapidly thawed at 37°C, then resuspended in 5 mL of pre-warmed medium, and pelletized by centrifugation at 1,000 rpm for 4 minutes. The pellet was resuspended in fresh medium and heated to approximately 1.6 × 10⁶. 6 The cells were seeded into a new T75 flask. When the cells reached 95% confluence, they were subcultured using trypsin and seeded again at 40% confluence. After two subcultures, the cells were used in the experiment.
[0356] For the trait transfer experiment, 5 × 10 per well 5 HEK293T cells were seeded on multiwell plates containing complete DMEM medium (Dulbecco's Modified Eagle Medium (DMEM)) supplemented with 10% fetal bovine serum, 2 mM glutamine, and 100 U penicillin / 0.1 mg / mL streptomycin. Prior to transfection, the medium was replaced with 2.7 mL of fresh complete DMEM medium. Opti-MEM I Reduced Serum Medium was mixed with each plasmid combination and 1 mg / mL of linear polyethyleneimine (PEI 25K) solution. A ratio of PEI 25K (μg):total DNA (μg) = 3:1 was used. The two solutions were mixed and incubated at room temperature for 15 minutes. After incubation, 300 μL of the mixture was added dropwise to the cells. 24 hours after transfection, the medium was replaced with fresh complete medium. Cells were harvested after transfection for flow cytometry or cell sorting and DNA extraction.
[0357] HEK293T cells were co-transfected with programmable transposable fusion proteins from Table 2, plasmids encoding the nucleic acids to be integrated, RFP (red fluorescent protein) or GFP (green fluorescent protein) transposons, and plasmids containing guide RNA targeting the AAVS1 site (adeno-associated virus integration site 1) in the human genome. Highly active PiggyBac and SB100 were used as positive controls, and transposons alone were used as negative controls for episomal expression detection (i.e., expression from non-inserted plasmids). Fluorescence was analyzed by flow cytometry up to day 14, after which episomal fluorescence could not be detected. Cells were then sorted by GFP expression, and two days after sorting, integration of the target DNA was quantified by counting the percentage of fluorescent cells.
[0358] Results and Conclusions: Results for the Cas9-PB fusion are shown in Figures 1A and 1C, and results for the Cas9-SB100 fusion are shown in Figure 1B. Human Cas9 fused to highly active PiggyBac (hCas9PB) and nicasse Cas9 fused to highly active PiggyBac (nCas9PB) increased the percentage of fluorescent cells by approximately 8% compared to the episomal RFP negative control after 14 days (Figures 1A, 1C). Therefore, the fusion proteins were able to successfully integrate exogenous DNA into the cellular genome. The tested Cas9-Sleeping Beauty fusion protein did not produce more fluorescent cells than the episomal GFP negative control after 14 days (Figure 1B). Example 2: Target transposition efficiency of programmable transposase fusion protein
[0359] Following the examples described above, we investigated whether there existed a targeted insertion (vs. non-target) with the best total insertion configuration in Example 1. For this purpose, HEK293T was co-transfected with lipofectamine 3000 along with a plasmid (pSico) encoding hCas9PB or nCas9PB, a gene trap plasmid encoding a transposon with inverted repeats and promoter-less GFP, and a guide RNA (gRNA) targeting an AAVS1 site or a site within the CD46 gene after the promoter on the human genome. The 3' end of Cas9 was ligated to the 5' end of the transposase by a linker (SEQ ID NO: 48). An example of a Cas9PB expression vector is shown in Figure 2A. The transposase contained a splicing acceptor and promoter-less GFP between the 3' and 5' ends. The gRNA and Cas9 instruct the transposase to incorporate the transposon into the promoter region. Using this approach, cells become fluorescent only when the transposon is inserted into the target site.
[0360] Results and Conclusions: Quantification of the percentage of GFP-expressing cells showed that the programmable transposase fusion protein Cas9-PiggyBac ("Target HCas9") and nickas Cas9-PiggyBac ("Target NCas9") had higher targeted DNA transduction ability compared to the control "untargeted" (control for overall insertion (PiggyBac alone)) and "episodes" (non-integrated negative control (transposon alone)) (Figure 2B). In this case, the 3-fold and 4-tine increase in signal above background was significant, in particular, considering that not all cells were efficiently transformed by all vectors required for transposon insertion; and the efficiency of random insertion for hyPB under unoptimized conditions as used herein is 10–15%. Example 3: Generation of recombinant highly active PiggyBac transposase
[0361] To enhance the specificity of exogenous nucleic acid sequence insertions into the genome, recombinant, highly active PiggyBac transposase was generated. Table 3 shows a list of amino acid mutations in the transposase. [Table 3] JPEG0007830129000004.jpg198169
[0362] In Example 4 below, several constructs were generated to create zinc finger proteins (ZFPs) capable of binding to chromosomal target sites for the introduction of the target gene. ZFPs constitute a substitute for Cas9 as a DNA-binding protein. Examples 5-13 generally relate to the generation and performance of HIV-1 integrase and Cas9 / ZFP fusion protein constructs with regard to targeted insertion. In particular, Example 5 generated a fusion protein of ZFP and integrase. In Example 11, Examples 6-10 provide various integrase-deficient packaging systems (i.e., non-integrated vectors) prepared to serve as a basis for in vitro studies to demonstrate the recovery of insertion function using the constructed integrase fusion proteins. In Example 12, it is observed that the targeted integrase fusion protein increases the rate of targeted insertion. Example 4: Generation of Target Zinc Finger Proteins (ZFPs)
[0363] The objective was to generate several ZFPs that bind to chromosomal target sites for the insertion of the target gene. A 6-domain zinc finger protein was generated to target the AAVS1 site (SEQ ID NO: 40) on the human genome. The target DNA sequence and the corresponding ZFP helix are shown in Table 4. Constructs encoding the target site and ZFP were prepared (AAVS1-6d-ZFP). The nucleic acid and amino acid sequences encoding the ZFP are SEQ ID NOs: 32 and 33, respectively. [Table 4] Example 5: Generation of ZFP-integrase fusion protein
[0364] An integrase fusion protein was generated with a ZFP having six domains (substantially sequence-specific). To generate site-specific integrase, the ZFP generated in Example 4 (AAVS1-6d-ZFP) was cloned into a pcDNA3.1 expression vector together with HIV-1 integrase (SEQ ID NO: 1) (pZFP-AAVS1-6d-IN). The sequence encoding the fusion protein encodes an N-terminal nuclear localization signal (SEQ ID NO: 47) and a GGS linker sequence (SEQ ID NO: 48) between the ZFP and the integrase (Figure 3).
[0365] Additional integrase fusion vectors, such as pZFP-TRCa-IN (containing SEQ ID NO: 38 and targeting the TRCa locus) and pZFP-AAV1-TEX-IN (containing the TEX linker (SEQ ID NO: 61)), were generated and prepared using a similar method. Example 6: Generation of an integrase-deficient DNA vector
[0366] The integrase-deficient packaging system was generated to serve as a basis for in vitro studies using manipulated integrases. A deficiency integrase construct was prepared from the non-integrated packing plasmid (NILV) psPAX2. The psPAX2 plasmid contains one N64D mutation and two N64D / N116D mutations. A deletion integrase (ΔIN) plasmid lacking the complete integrase coding region was prepared. A non-coding plasmid containing a stop codon before the integrase coding sequence (hereinafter, Example 8) was prepared. A plasmid containing the deletion integrase, including a construct with a C-terminal domain and DNA-binding domain that does not contain cPPT / CTS (hereinafter, Example 10), was prepared. A general cloning protocol is briefly described below. KAPA HiFi HotStart Protocol
[0367] For PCR experiments using KAPA HiFi HotStart, the PCR reaction mixture was prepared according to the protocol provided by the KAPA HiFi PCR Kit manufacturer. The KAPA HiFi PCR reaction was performed using a Mastercycler Pro. Plasmid DNA extraction
[0368] Plasmid DNA was extracted using the QIAprep Spin Miniprep Kit according to the manufacturer's protocol. Bacterial cultures were collected by centrifugation at 5,000 rpm for 3 minutes. The cell pellet was resuspended in 250 μL of buffer P1 and mixed by inverting the tube with 250 μL of buffer P2 4-6 times. 350 μL of buffer N3 was added and mixed by inverting the tube. The Eppendorf tube was centrifuged at 12,000 rpm for 10 minutes to remove cell debris and chromosomal DNA. The supernatant was transferred to a supplied QIAprep spin column and centrifuged for 1 minute (12,000 rpm). The sample was washed twice with 0.5 mL of buffer PB and 0.75 mL of buffer PE, centrifuged at 12,000 rpm for 1 minute each time. Residual wash buffer was removed by further centrifugation at 12,000 rpm for 1 minute. The QIAprep spin column was transferred to a new 1.5 ml microcentrifuge tube, 50 μL of water was added, the tube was allowed to stand for 1 minute, and then the plasmid was eluted by centrifugation at 12,000 rpm for 1 minute. The concentration was measured using NanoDrop One. Isolation and purification of plasmid DNA
[0369] Bacterial strains containing the desired plasmid (DH5α or DH10B) were grown overnight in LB medium containing 100 μg / mL carbenicillin. Plasmids were isolated using either a plasmid mini or maxi kit from NZYTech, according to the manufacturer's protocol. Plasmids were eluted with either 30 μL (miniprep) of 65°C warm water or 500 μL (maxiprep) of 65°C warm water. Plasmids were stored at -20°C. For PCR purification, the reaction mixture was treated with a PCR purification kit. DNA was eluted with 30 μL of 65°C warm water. DNA gel electrophoresis
[0370] Agarose was dissolved in 100 mL of TAE buffer by boiling. 4 μL of Greensafe was added to each 100 mL of agarose solution in the liquid gel, and the mixture was poured into a tray. To visualize the DNA preparations, the DNA was mixed with a 6× loading dye and loaded onto a 1% agarose gel. Additionally, 1 μL of gene ladder per 1 mm gel lane was loaded into one chamber. The gels were run at 100 V for 1.5 hours and visualized using a transilluminator. Transformation
[0371] For the DH5α transformation experiment, the plasmid was transformed into 50 μL of DH5α cells according to the manufacturer's protocol. After harvesting in soc medium, the bacteria were pelleted at 15,000 g for 30 seconds and resuspended in 50 μL of LB medium. The cells were spread on an LB agar plate containing 100 μg / mL carbenicillin and incubated overnight at 37°C. The culture was collected and incubated overnight in LB medium containing 100 μg / mL carbenicillin. The liquid culture was used again for plasmid isolation or as a glycerol stock. For the glycerol stock, 500 μL of liquid culture was mixed with 500 μL of 50% glycerol and stored at -80°C.
[0372] For the transformation experiment using XL-10 Gold super-competent cells, the cells were first thawed on ice, and 45 μL of cells were added to a pre-cooled 14 mL Falcon polypropylene round-bottom tube. 2 μL of the β-ME mixture supplied with the kit was added to the cells. The contents of the tube were gently swirled, and the cells were incubated on ice for 10 minutes (swirled every 2 minutes). 1.5 μL of DpnI-treated DNA was added to an aliquot of the cells, mixed, and incubated on ice for 30 minutes. The cell / DNA mixture was heat pulsed in the tube at 42°C for 30 seconds. The tubes were then incubated on ice for 2 minutes. Next, 0.5 mL of preheated (42°C) NZY+ broth was added to each tube, and then incubated at 37°C for 1 hour with shaking at 225-250 rpm. This mixture was then plated onto an agar plate containing appropriate antibiotics for the plasmid vector. Five colonies were selected for DNA extraction and their sequences were validated. Colony 1 was selected and maintained. Example 7: Generation of non-integrated vectors containing PPT or ZFP recombinant integration proteins
[0373] To construct a fully functional psPAX2 plasmid lacking integrase (IN), the polypyrimidine tract domain (PPT) (SEQ ID NO: 74, essential for the subsequent double-stranded cDNA formation of all retroviral RNA genomes, including lentiviruses) was cloned into an integrase-free psPAX2 vector (psPAX2-ΔIN). The AAVS1-targeting synthetic zinc finger construct (AAVS1-6d-ZFP-IN) generated in Example 4 was cloned into psPAX2-ΔIN. Two different forward primers and the same reverse primers (SEQ ID NOs: 75-77) were designed for PPTs with and without a stop codon (IN+PPT and IN+PPT(STOP)). Two different forward primers (SEQ ID NOs. 78-80) and the same reverse primer were designed for AVS1-6d-ZFP-IN with and without a nuclear localization signal (AAVS1-6d-ZFP-IN and AAVS1-6d-ZFP-IN(-NLS)). The inserts were amplified by PCR under Kappa standard conditions, with an annealing temperature of 62°C, a PPT extension time of 40 seconds, and an AAVS1-6d-ZFP-IN extension time of 90 seconds. The PCR products were separated by gel electrophoresis.
[0374] The amplification product was purified, and an assembly protocol was performed with a backbone:insert ratio of 1:2.5 and 5 cycles. 50 μL of competent cells were transformed with 4 μL of ligation product, and 60% of the competent cells were seeded on carbenicillin plates. Initial validation of the colonies was determined by restriction enzyme digestion and DNA gel electrophoresis. The following colonies were collected: colonies 1 and 2 (IN+PPT F1+R, AAVS1-6d-ZFP-IN F1+R, AAVS1-6d-ZFP-IN(-NLS)F2+R), and colonies 7 and 8 (IN+PPT(STOP)F2+R). Colony PCR was performed using 4 mM Mg, 62-STS, and NEB standard taq to further confirm that the colonies contained the correct inserts. Example 8: Generation of non-embedded vectors by stop codon insertion
[0375] A non-integrated vector was generated by inserting a stop code before the integrase open reading frame (psPAX2-TAA-IN). psPAX2-TAA-IN was synthesized by site-directed mutagenesis by adding two stop codons after the first protease cleavage site of integrase. psPAX2-TAA-IN was then synthesized using PCR conditions for site-directed mutagenesis.
[0376] After PCR, the reaction tubes were cooled on ice for 2 minutes. Next, 1 μL of DpnI was added directly to each amplification reaction, and the mixture was incubated at 37°C for 5 minutes to digest the parental (non-mutant) double-stranded DNA.
[0377] The plasmid DNA was digested to confirm that site-directed mutagenesis did not result in any undesirable recombination. Digestion of psPAX2 and psPAX2-TAA-IN with SacI and AgeI should yield three bands of 7,500, 1,900, and 1,300 bp. Digestion of psPax2-ΔIN with SacI and AgeI should yield three bands of 7,500, 1,300, and 800 bp. The digestion reactions were performed, and the digestion yielded the correct band patterns. Example 9: Reconstitution of wild-type integrase into an integrase-deficient vector
[0378] The objective was to develop a method to see if non-integrated vectors could recover insertion activity by expressing different forms of integrase fusion proteins. To confirm that psPAX2-ΔIN was fully functional, integrase was added to the vector using Gibson assembly. Furthermore, to test whether the assembly site was good for cloning the fusion "IN", wt-IN was cloned with an additional N-terminus of IN located in the scaffold preceding that site (containing a Leu that should not be present). This was also done using an extra protease target sequence to avoid this spurious N-terminal domain. PCR reactions were performed to amplify the IN-1, IN-2, and IN-3 fragments.
[0379] PCR amplification products were separated by DNA gel electrophoresis. The amplified bands were purified and assembled with a backbone:insert ratio of 1:2.5 for 5 cycles at 37°C. 50 μL of competent cells were transformed with 4 μL of ligation product and seeded on carbenicillin plates.
[0380] To generate constructs containing IN-3, Gibson assembly was performed according to the standard protocol for the Gibson Assembly HiFi 1-Step Kit (using CRG MM) (SGI-DNA, Inc., www.sgidna.com / products / gibson-assembly-reagents / ) to produce the reaction mixture. The reaction mixture was prepared and assembled at 50°C for 1 hour. Competent cells were transformed using 2 μL of the reaction mixture.
[0381] 50 μL of competent cells were transformed using 2 μL of ligation product and seeded on a carbenicillin plate. Example 10: Generation of a non-integrated vector containing a C-terminal domain deletion integrase
[0382] The C-terminal domain (CTD) (nucleic acid segments 83-118 of SEQ ID NO: 74) and the CppT+CTD (SEQ ID NO: 74) integrase fragment were cloned into the psPAX2 vector.
[0383] PCR amplification products were separated by DNA gel electrophoresis. CppT+CTD ligation was performed using the conditions used in Example 9.
[0384] Ligation was performed for 5 cycles at 65°C, and the ligation product was transformed. No colonies grew. Ligation and transformation were repeated, and three colonies were identified by sequencing using IN-fw primer (SEQ ID NO: 81). Example 11: Generation of integrase fusion protein
[0385] Targeted integrase fusion proteins were constructed by incorporating HIV-1 integrase and either targeted ZFP or human Cas9 into a pcDNA3.3 expression vector. A first vector was constructed in which the 3' end of ZFP or Cas9 was ligated to the 5' end of integrase by a nucleic acid linker. A second vector was constructed in which the 3' end of integrase was ligated to the 5' end of ZFP or Cas9 by a nucleic acid linker. The linkers used were XTEN or GGS with lengths ranging from 13, 16, 19, 22, 25, or 28 amino acids. ZFP-integrase fusion proteins were constructed to target the AAVS1 site or the T cell receptor α (TCRa) locus in the human genome. Cas9-integrase fusion proteins were used in combination with guide RNA targeting the AAVS1 site or the TCR locus in the human genome. A list of recombinant integrase fusion proteins is shown in Table 5. [Table 5] Example 12: CYS and interspecies complementation of integrase-deficient lentiviruses using targeted integrase fusion proteins
[0386] The target integrase fusion protein from Example 11 was used to compensate for the lack of integration ability of a non-integrated lentivirus expressing IN with two mutations in the catalytic domain (D64V / D116N). For this experiment, the target integrase fusion protein was cloned into the pcDNA3.1 vector. Lentiviruses were produced by co-transfecting with a pcDNA3.1 vector containing either pSICO (GFP expression payload), pmd2.g (VSVG for envelope expression), pax2 (containing packaging protein and integrase) or NILV-pax2 (containing packaging protein), and either wild-type integrase or target integrase (Table 6). [Table 6]
[0387] 6 x 10 per well 5 HEK293T cells (Passage 8) were seeded in 6-well plates and incubated overnight. Five hours before the start of virus production, the medium was replaced with 1.7 mL of medium containing 1:1000 chloroquine diphosphate (CD; stock = 25 mM). Plasmids were infected with a molar ratio of 1.6:1.32:0.72:3.32 (pSICO:pax2:VSVG:wtIN rescue). Polyethyleneimine (PEI; stock = 1 mg / mL) was used as the transfection reagent, while 3 μL of PEI was used for 1 μg of total DNA used for transfection. The DNA was diluted with 83 μL of Opti-MEM and 83 μL of PEI, mixed, and incubated at room temperature for 15–20 minutes. Each transfection mix was added dropwise to the cells along with CD medium. The cells were incubated overnight, and the medium was replaced with 2.5 mL of fresh medium the following day. The following day, the cell supernatant was centrifuged at 1,000 rpm for 5 minutes and passed through a 45 μM filter. The supernatant containing the virus was stored at -80°C.
[0388] The first step was to confirm that different lentiviral packages maintained the ability of infected cells independently of their contents. To determine viral titers, 75,000 HEK293T cells per well were seeded into a 6-well plate. The cells were infected with a mixture of 1 mL of medium containing 1:100 polyblen and 500 μL of pre-produced viral supernatant (1:3). The medium was changed the following day. The following day, the medium was aspirated and the cells were separated using 200 μL of trypsin. The reaction was stopped by adding 800 μL of normal medium and analyzed by flow cytometry. Viral titers were quantified for wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZP-IN(AAVS1)), non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN), and wild-type integrase lentivirus with wild-type integrase (LV+IN). LV and LVO were used as positive and negative controls, respectively. Viral titers were quantified by infecting HEK293T cells and counting the number of GFP-positive cells (Figure 4). Results: Viral titers were within the same order of magnitude for all conditions.
[0389] Next, the overall integration capacity of the target integrase fusion protein was determined by flow cytometry and next-generation sequencing of the target insert. HEK293T cells were infected with the same infection multiplicity under all conditions, and GFP fluorescence was measured at 3, 5, 7, 10, and 12 days post-infection. After 7 days post-infection, cells were sorted by GFP expression. Results: At day 12, cells infected with non-complementary NILV had a lower proportion of GFP-expressing cells (Figure 5), indicating a decrease in viral production capacity.
[0390] To evaluate the target integration ability of the tested integrase fusion proteins, genomic DNA was extracted on day 12 according to the DNeasy Blood and Tissue Kit Protocol (Qiagen). Cell cultures (maximum 5 × 10⁶) were centrifuged at 190 rpm for 5 minutes. 5 The ) was collected. The pellet was dissolved in 200 μL of PBS (phosphate-buffered saline). 20 μL of proteinase K was added along with 200 μL of buffer AL. After vortexing, the sample was incubated at 56°C for 10 minutes. 200 μL of ethanol (96-100%) was added, and after a short vortex, the mixture was transferred to a DNeasy Mini spin column, placed in a 3 mL collection tube, and centrifuged at 8,000 rpm for 1 minute. The spin column was transferred to a new 2 mL collection tube, and 500 μL of buffer AW1 was added. The tube was centrifuged at 8,000 rpm for 1 minute. This washing step was repeated for buffer AW2 (centrifugation for 3 minutes). Next, the spin was transferred to a new 1.5 mL microcentrifuge tube, 200 μL of buffer AE was added to the center of the spin column membrane, and the tube was left to stand for 1 minute to elute the DNA, after which it was centrifuged at 8,000 rpm for 1 minute. Genomic DNA concentration was quantified using NanoDrop One.
[0391] Inverse cloning was performed using oligonucleotides specific to the viral insertion LTR. Next-generation targeted sequencing was analyzed using the following parameters: Reads like both R1 and R2 were analyzed using the following parameters: Select R1 and R2-like reads containing the corresponding sequencing primers, restrict checks on the leftmost 5 bases of the read (as many base pairs as the primer has), tolerate two mismatches, cut out primer sequences (sequences 82-89), select R1 and R2-like reads containing the corresponding LTR bases, restrict checks on the leftmost 5 bases of the read, K=3 (k=3 for sequence ACTGA as follows) The five initial LTR bases (following the sequencing primer) were used to confirm the presence of the bases (ACT, CTG, TGA) in the read, allowing for two mismatches, the corresponding LTR base pairs were excised, the reads were mapped to the reference genome, coverage (number of reads per insertion site) was obtained, overlapping regions of R1 and R2 were split into two, if there was no overlap between R1 and R2, only one insertion site was added, a coverage threshold was applied, coverage per 10mb of the reference genome was calculated, a coverage plot was performed, and the percentage of coverage for each insertion site was calculated. Results: Targeted integrase fusion proteins increased the coverage and percentage of targeted insertions at the AAVS1 site (Table 7 and Figure 6). As seen in Table 7, when insertions are performed by integrase fusion proteins, the target site has a larger number of reads compared to IN WT showing targeted insertions. Figure 6 is representative of common target sites in the genome for IN and ZFP_IN(AAVS1), showing that the target insert is present only in the fusion state. [Table 7]
[0392] A second ZFP was also generated to target a nucleic acid segment within the CCR5 gene. This zinc finger protein was fused to HIV-1 integrase to create a CCR5-targeting integrase. A lentivirus containing this ZFP-IN was produced as described above and transduced into HEK293T cells (NILV+ZP-IN(CCR5)) (Table 6). Results: The viral titer of NILV+ZP-IN(CCR5) was similar to that of LV and NILV+IN (Figure 7A). This construct was able to produce viral particles with the same efficiency as the other ZFP_IN fusions tested (Figures 7B and C). Its ability to incorporate DNA in a site-specific manner was not tested for CCR5.
[0393] In another experiment, a newly cloned expression vector for a fusion ZFP-IN having the 6d-targeted TCR gene locus and a gRNA targeting the same site was used (see Example 11). This assay investigated whether the wild-type integrase and ZFP-integrase fusion could complement the NILV's ability to promote selective incorporation of CAR-T cassettes. Jurkat cells were infected with the same infection multiplicity for all TCR-targeted insertion particles. In this experiment, CD19 CAR-T cassettes were loaded onto viral particles to induce loss of CD3 (encoded by the TCR gene) protein expression after targeted insertion. The percentages of CD19-positive and CD3-negative cells were tracked over time. Lentiviral titers are shown in Figure 8A, and the percentages of CAR-expressing cells on days 3 and 14 are shown in Figure 8B. The percentages of CD3-expressing cells are shown in Figure 8C. This indicates that, in the absence of VPR, which is a key factor for efficient interspecies complementation of IN, interspecies complementation did not function in the context of this cell line. Example 13: Generation of recombinant integrase by site-directed mutagenesis and saturation mutagenesis
[0394] Recombinant HIV-1 integrase was generated by site-directed mutagenesis and saturated mutagenesis induction. For site-directed mutagenesis, recombinant HIV-1 integrase was produced by mutating amino acids via site-directed mutagenesis. Primers were designed using the QuikChange Lightning Multi Site-Directed Mutagenesis Kit according to the manufacturer's recommendations (SEQ ID NOs. 90-97). The plasmid to be mutated was approximately 7,000 bp in length. Approximately 5 colonies per approach were screened by sequencing. Glycerol stocks of colonies containing the desired plasmid were prepared.
[0395] Saturated mutations were introduced into HIV-1 integrase to generate a large combinatorial library of different HIV-1 integrase molecules. The protocol was adopted from Cornell et al. (Biochemistry, 57(5)604-613, 2018). Several forward primers containing degenerate NNS sequences at the mutation site were used, and one reverse primer was used per PCR reaction (SEQ ID NOs. 90-97). The entire plasmid was amplified to generate recombinant integrase molecules. The primers were optimized for a melting temperature of 68°C. During the cycle, the annealing temperature increased by 0.3°C per cycle. A list of amino acid mutations is shown in Table 8. [Table 8] JPEG0007830129000010.jpg227169 Example 14: Generation of pRRLVPR integrase constructs in HEK293T cells and testing of interspecies complementation efficiency.
[0396] pRRLIN, pRRLVPRIN, and pRRLINGFP vectors were generated for use in VPR heterologous complementation (Table 9). [Table 9]
[0397] The constructs were tested using a GFP expression assay. HEK293T cells were transfected with pSICO MAXI, pSICO MINI, and pRRL_INGFP, and pRRLINGFP episome expression was tested. Expression of the VPRINGFP construct in lentivirus-producing cells was positive. Next, the interspecies complementation efficiency in HEK293T cells was tested.
[0398] The LV medium was ultracentrifuged, allowed to stand for resuspension, and then seeded with cells. Infection was performed in a volume of 0.6 ml (1.5 × 0.4). Polybren was added. Titer was determined by cytometry. The titer (1:100) is shown in Figure 9.
[0399] We will compare recombinant integrase sequences for integration using the VPR heterologous complementation system.
[0400] In the following embodiments 15-19, various constructs of fusion proteins containing recombinant high-activity PiggyBac transposase were generated. The total transposability and targeted transposability activities of the constructs were determined, yielding results particularly relevant to the hcas9 mutant PB construct. Evidence for the generation of mutant PB and ZFP fusion protein constructs and the determination of their targeted transposability activities is also provided. Different linkers were tested, and XTEN demonstrated superior performance compared to the other linkers tested. 5GGS and 7GGS also performed well, demonstrating that linker length and its flexibility played a crucial role in their performance. Example 15: Method for generating a fusion protein containing recombinant high-activity PiggyBac transposase and determining its target transposition efficiency. Transfection:
[0401] Hek293T cells were seeded the day before transfection to achieve 70–80% confluence (typically 290,000 cells in a p12 well plate). Transfection was performed using lipofectamine 3000 reagent or PEI in OptiMem (1:3 DNA-PEI ratio) according to the manufacturer's instructions.
[0402] Programmable transposases (PTs), gRNAs, and transposon plasmids were transfected together in a ratio of 1 PT:2.5 gRNA:2.5 transposons.
[0403] The cells were passaged and maintained until the desired endpoint was reached, according to the experiment. Generation of PB mutants:
[0404] Site-directed mutagenesis was performed using the Quickchange Lightning Agilent Mutagenesis Kit to introduce different mutations into the hyPB sequence fused to Cas9 (hCas9_PB plasmid). Primers were designed using QuikChange Primer Design to achieve the following mutations: PB R245A, PB R275-277A, PB R388A, PB S351A, PB W465A, PB R372A-K375A, PB D450N (SEQ ID NOs: 100-106). Cas9 activity:
[0405] Programmable transposase plasmids containing the nuclease Cas9 and gRNa plasmid were transfected together in a 1:2.5 ratio. Cells were harvested after 48 hours and genomic DNA was extracted. PCR was performed using primers (NGS-aavs fw and NGS-aavs rv, SEQ ID NOs. 98-99) targeting 150-200 bp around the gRNA target site. Illumina adapters and barcodes were introduced into a second PCR, and miseq sequencing was performed using a standard 2×250 Nano flow cell. Results were analyzed using the CRISPR-GA web tool. Gene trap assay:
[0406] We first constructed promoter-less RFP transposons, designed splicing acceptors and gRNAs targeting PPR1α and CD46 intron 1, and cloned them under U6 promoter control. RFP fluorescence is detected only when the transposon is incidentally inserted into the target region or other promoter regions. For gene trap assays, Hek293T cells were transfected with gene trap transposons, programmable transposases, and gRNAs, and RFP signaling was analyzed by flow cytometry. Split GPF reporter cell line:
[0407] A 293T reporter cell line was generated for targeted transplantation evidence experiments. Briefly, this cell line possesses a target region (containing different gRNA and ZFP target sequences) and a splicing acceptor sequence, followed by half of the GFP coding sequence. This cell line was generated by random insertion of a reporter cassette using a highly active form of SB100X (Sleeping Beauty trasse). Targeted transposon introduction containing the first half of the GFP sequence with promoter and splicing donor results in a GFP signal detectable by flow cytometry.
[0408] To evaluate targeted insertion versus random insertion, a second transposon was generated containing half of the GFP sequence and the complete RFP sequence preceded by the EF1α constitutive promoter. Approximately 15 days after transfection, there was good decay of episomal signaling, enabling analysis of whole insertion (RFP signal) versus targeted insertion (GFP signal). Example 16: Generation of a fusion protein plasmid construct containing recombinant high-activity PiggyBac transposase
[0409] To achieve fusion of a programmable element targeting DNA (Cas9, ZNF) with a mammalian transposase (Piggybac, SB100), various plasmid constructs were cloned. The linker between the two modules was variable in different constructs and was selected from a linker library with sequence numbers 50–63. The constructs are shown in Table 10. [Table 10] Example 17: Transfer efficiency of various linkers
[0410] Hek 293T cells were transfected with hcas9_PB constructs containing linkers of different lengths and structures (linker library) and two different gRNAs (AAVS1 1 and AAVS1 2). Genomic DNA was extracted 48 days after transfection, the target region was amplified by PCR, and sequenced by Illumina miseq sequencing.
[0411] Results: Constructs with different linker lengths and structures did not interfere with Cas9 nuclease activity. The 4GGS linker gave higher Cas9 activity at both gRNA target sites compared to hCas9 activity (Figure 11). Example 18: Targeted transposition of fusion proteins using recombinant high-activity PiggyBac transposase 18.1. Gene Trap:
[0412] The target transposability activity of the hcas9_PB construct (hcas9 linked to hyPB using the different linkers described above) was evaluated using gene trap transposons. Gene trap transposons contain a promoter-less RFP sequence preceded by a splicing acceptor sequence that can only be expressed when inserted into the promoter region after a splicing donor.
[0413] Gene trap transposons were transfected with PPR1 intron 1 gRNA and programmable transposases containing different linker constructs. Ten days after transfection, the results were analyzed by RFP fluorescence using flow cytometry.
[0414] Results: Targeting activity was increased by programmable transposase compared to hyPB random insertions, which showed more fluorescence when transfected with programmable transposase than when transfected with wild-type hyPB. 8ggs, the XTEN linker increased gene trap targeting activity compared to other linkers (Figure 12). Split GFP reporter cell line:
[0415] 18.2 Target transfer hcas9_PB with different linkers
[0416] The target translocation activity of hcas9_PB constructs was evaluated using reporter cell lines. Cash9_PB constructs with different linkers were transfected with gRNA AAVS1 3 or TCR1α and half-GFP transposons. Results: No significant differences were observed in translocation between different linker constructs (Figure 13).
[0417] 18.3. Targeted translocation of selected mutants:
[0418] PB 450 and PB 372-375-450 were selected for further targeting transposition experiments due to their superior targeting transposition efficiency. Experiments were performed using gRNA aavs1 3 and tcr 1 as previously described. Results: Targeting transposition of hcas9_PB 450 and hcas9_PB 372-374-450 was 6–10 times higher compared to hcas9_PB with the hyPB WT sequence. Cash9+hyPB transfected into isolated plasmids showed some targeting activity, while hyPB without hCas9 showed 0 activity, indicating that split GFP reporter cell lines are a robust targeting insertion method for selecting functionally active variants over the noise of the Ther method, which is not sufficiently specific (Figure 15). 18.4. Selected PB variants with targeted and random translocation:
[0419] Targeted and random transposition were evaluated for the selected mutants of Example 19.4 using the RFP-GFP dual transposon described above. Red fluorescence indicates total insertion (RFP is constitutively expressed) approximately 15 days after transfection (to ensure non-episomal signaling), and GFP fluorescence indicates targeted transposition. Results: Figure 16 shows that both the hcas9_PB D450N and hcas9_PB R372A K375A D450 selected mutants showed higher targeted transposition than random transposition compared to hcas9:PB with wt hyPB sequences. Overall transposition efficiency was lower in both mutants, and the targeted results were consistent with Figure 15. 18.5. Targeted transfer ZFP-PB constructs:
[0420] Zinc finger-highly active PiggyBac fusion proteins were cloned using the ZFP target tcr4 sequence present in a split GFP reporter cell line, and hyPB with either hyPB or the D450N mutation. Cells were transfected with the ZFP-PB combination and 1 / 2 GFP transposon according to the protocol of Example 15. GFP signaling was analyzed 5 days after transfection. Results: Higher target transposition was observed in all constructs compared to the background (hyPB random insertion). Results: Target transposition was higher in ZFP at the N-terminal position for both hyPB and hyPB D450N (Figure 18). The ZFP sequences for these experiments correspond to 6-finger domain proteins with nucleic acid sequences and amino acid sequences SEQ ID NOs. 117 and 118, respectively.
[0421] In Example 20 below, a library of PB mutations was designed and subjected to a screening method to identify recombinant PBs with positive target translocations. Several hits of recombinant PBs with positive target translocations were identified and validated. Example 20: Preparation of a highly active PiggyBac mutant library and screening of target metastases. method:
[0422] We designed a library of hyPB mutations and purchased it from Twist Biosciences. [Table 11] Screening method:
[0423] A screening method was designed to identify Piggybac variants from a mutant library that is linked to a targetable DNA-binding protein such as cas9 and is designed to perform specific target transfer. The overview of the screening method is shown in Fig. 19. The PB library was cloned into a SIN transfer lentiviral plasmid containing hcas9 and an XTEN linker by Golden Gate assembly using the Esp3I enzyme to achieve the hcas9_XTEN_PB_NLS fusion protein under the regulation of the CMV promoter. Subsequently, an Esp3I cloning site was cloned before the NLS. After Invitrogen electroporation and ElectroMAX() TM Stbl4 TM Approximately 6,000,000 colonies were collected after competent cells, and the plasmid was extracted by maxiprep using the HiPure Maxiprep kit from LifeTechnologies. Lentiviruses were produced using the lentivirus production protocol from Addgene (using the pMD2.G and psPAX2 helper plasmids purchased from Addgene). The lentiviruses were ultracentrifuged and titrated by copy number analysis qPCR (using oligonucleotide sequence numbers 107 - 110). Briefly, 80,000 Hek293T cells were seeded in a p12 well plate the previous day. The cells were infected with the library lentivirus and a standard GFP lentivirus. For the library lentivirus, infections were carried out at dilutions of 1 / 2, 1 / 10, and for the GFP lentivirus at dilutions of 1 / 50, 1 / 100, 1 / 1000. The GFP signal was analyzed by flow cytometry 3 days after infection. The cells were harvested and gDNA was extracted. The qPCR method was designed to evaluate the WPRE gene copy number and was normalized by the RNAse gene copy number.
[0424] Hek293T reporter cells were plated in 500 cm using 1:1000 polybrene 210M cells were infected with a MOI of 0.8 in square dishes and cultured on plates the day before. Three to four days after infection, cells were transfected with 8.1 pmol of gRNA AAVS1 plasmid and 1 / 2 GFP transposon using PEI 1:3. 9M cells were plated in 15cm dishes the day before. Three to four days after transfection, cells were sorted using a FACSAria cytometer with a 0.70 μm nozzle. Transfection controls were performed in 10cm dishes using RFP and GFP plasmids at the same molar concentrations, and GFP-RFP positive cells were analyzed using a Fortessa cytometer. After sorting, gDNA was directly extracted.
[0425] We analyzed PB mutants with positive target translocation using different sequencing methods: PiggyBac library-based region-targeted sequencing:
[0426] The PiggyBac 1116bp region containing all library variants was PCR-amplified using KAPA HiFi Hotstart ReadyMix with primers (NGS cluster 1 fw and NGS cluster 2 rv). In the second PCR, an Illumina adapter and barcode were added, and NEBNext 9 primers and Illumina custom barcode were used (SEQ ID NOs. 111-114). Target sequencing was performed in v2 or v3 Illumina miseq flow cells. I7 index primers were replaced with custom primers to allow complete sequencing of different variants. Preparation and sequencing of Piggybac and Cas9 sequence shotgun libraries:
[0427] 6000 bp PCR was performed on genomic DNA from GFP-positive selected cells using primers CMV-F and SV40 pA rv (SEQ ID NOs. 115 and 132), and Cas9 and PB sequences were amplified using KAPA HiFi HotStart ReadyMix. The DNA was then purified using a Qiagen gel extraction kit and fragmented into 500 bp segments using Covaris S220 and microtube AFA fiber Crimp-Cap. Shotgun libraries were prepared using a KAPA Hyperprep kit according to the manufacturer's instructions. result: 20.1.hyPB Library Diversity Generation:
[0428] 1 / 2 GFP reporter cell lines were infected with a lentivirus containing hcas9_PB with a PB library mutation at an MOI of 0.8. Three days after infection, cells were transfected with gRNA AAVS13 and 1 / 2 GFP transposons with a transfection efficiency of 75-90%.
[0429] In the first experiment, a total of 254M cells were selected, yielding 185,757 positive cells, showing a 0.073% target metastasis positive variant. In the second experiment, 120M cells were selected, yielding 70,974 positive cells showing a 0.059% target metastasis positive variant (Figures 21A and 21B).
[0430] Genomic DNA was directly extracted from positive and negative sorted cells. Two-thirds of the obtained DNA was processed for targeted sequencing analysis, and one-third was processed as shotgun library sequencing as specified above in the Methods section of this example. 20.2. hyPB Library Screening Analysis by Target Sequencing of Variable Regions:
[0431] Positive and negative cell analysis of Cas9-PB variants was performed as follows: Reads from targeted sequencing were mapped to a reference sequence. All library variant locations were searched using two different approaches: one utilizing aligned reads by location, and the other utilizing pattern matching of surrounding sequences by sequence. The logarithmic change in all variant counts was calculated between positive (GFP-positive cells with targeted integration) and negative samples (non-targeted integration samples, regardless of whether integration occurred), and the top variant was retrieved. Furthermore, negative selection of these samples with random integration was performed using RFP-positive selection, where transposons were randomly inserted into the genome.
[0432] The results are shown in Figures 22A–22K. Thus, using an unsupervised high-throughput screening approach of the variant combinatorial library, a collection of Piggyback variants capable of site-directed insertion with high efficiency was identified, as demonstrated by the comparison of their presence in positive and negative cell populations.
[0433] Next, targeted and random transposition of the top positive hit in repeat 1 was evaluated using the RFP-GFP dual transposon described above. Red fluorescence indicates total insertion approximately 15 days after transfection (RFP is constitutively expressed), and GFP fluorescence indicates targeted transposition.
[0434] Results: Compared to hcas9_PB and wt hyPB, Top1 of the repeating 1 variant showed higher targeted transposition compared to random transposition (Figures 23A-23B). Independent validation of targeted insertion using our reporter cell line was performed, and significant targeting activity was observed compared to the WT version and the D450N mutation. 20.3. Identification of overexpression positive hits:
[0435] Screening identified several positive hits overexpressed in the GPF population against negative selection variants. Some of these were also not found in the RFP population, which represents the overall insertion, indicating increased insertion capacity. Furthermore, RFP includes both random and targeted insertions. Thus, a collection of combinatorial variants for PiggyBac capable of site-directed insertions with high efficiency was identified (Figures 24A-24C). 20.4. hyPB Library Screening Analysis using Shotgun Sequencing:
[0436] For shotgun sequencing, reads were mapped to a reference sequence, variant calling was performed to retrieve all variants from the reference, and Euclidean distances and correlation distances were calculated between positive and negative allele counts. The most distinct locations were identified as variants, and the relationships between these variants were calculated.
[0437] Results: In addition to the variants included in the library design, variants randomly introduced by lentiviral retrotranscriptase during viral library generation were analyzed. Some of these new variants were associated with positive hits and likely perform targeted incorporation in combination, and may need to be present as mutants in variant versions of hyPB to perform targeted incorporation. Examples of D450N and W465A are shown in Figure 25.
[0438] The mutant PB sequences identified in Example 20 are listed in Table 12 (Sequence IDs 120-129). 20.5.hyPB Library Screening Evaluation:
[0439] Targeted and random transpositions of several combinations of single mutations observed in Top1-1, identified by screening positive hits (Unilarge-A, -B, -C, and Unilarge-D), were evaluated using the aforementioned RFP-GFP dual transposon. Red fluorescence indicates total insertion approximately 15 days after transfection (RFP is constitutively expressed), while GFP fluorescence indicates targeted transposition.
[0440] Results: In all cases, compared to Cas9 fusion to the WT version of hyPB, an increase in targeted insertions was observed for Cas9 fused to different variant combinations of hyPB and 4GGS linkers (Unilarge-A:D450N; Unilarge-B:R245A / D450N; Unilarge-C:R245A / G325A / D450N / S573P; Unilarge-D:R245A / G325A / S573P) compared to Cas9 fusion to the WT version of hyPB, with an increase in targeted insertions compared to overall integration. Some of the variant combinations tested (R245A / G325A / D450N / S573P) showed a significant increase in targeted insertions to 30% of all integration events, unlike the 3% in hyPB fusion (Unilarge C) (Figure 26).
[0441] The following Example 21 provides an overview of the occurrence of various embedded deficiency viral vectors, the best interspecies complementation system, and data on interspecies complementation using IN fusion proteins. Example 21: Heterogeneous complementation of various integrase-deficient systems
[0442] To generate an efficient interspecies complementation system for testing IN fusion proteins, we evaluated viral production efficiency and integration capabilities by infecting Hek293T and Jurkats cells with different states of integrase-deficient viruses and interspecies complementation viruses. Cells were passaged for 7 days until episomal signaling was no longer detectable, and GFP signaling was analyzed by flow cytometry on days 2, 5, and 7.
[0443] Results: Since NILV is closed to WT during production, we were able to detect different production efficiencies for different systems. In all cases, when heterologous complementation was performed using WT-HIV_IN, a clear rescue of embedded activity was observed (Figure 27). Proof of IN being loaded into the heterologous complementary system was obtained by Western blotting. [Table 12] JPEG0007830129000015.jpg180169 JPEG0007830129000016.jpg215169 JPEG0007830129000017.jpg208169 JPEG0007830129000018.jpg221169 JPEG0007830129000019.jpg189169 JPEG0007830129000020.jpg144169 JPEG0007830129000021.jpg158169 JPEG0007830129000022.jpg149169 JPEG0007830129000023.jpg226169 JPEG0007830129000024.jpg233169 JPEG0007830129000025.jpg229169 JPEG0007830129000026.jpg227169 JPEG0007830129000027.jpg112169 JPEG0007830129000028.jpg235169 JPEG0007830129000029.jpg164169 JPEG0007830129000030.jpg231169 JPEG0007830129000031.jpg169169 JPEG0007830129000032.jpg236169 JPEG0007830129000033.jpg166169 JPEG0007830129000034.jpg183169 JPEG0007830129000035.jpg236169 JPEG0007830129000036.jpg148169 JPEG0007830129000037.jpg234169 JPEG0007830129000038.jpg234169 JPEG0007830129000039.jpg228169 JPEG0007830129000040.jpg182169 JPEG0007830129000041.jpg227169 JPEG0007830129000042.jpg233169 JPEG0007830129000043.jpg234169 JPEG0007830129000044.jpg58169 [Brief explanation of the drawing]
[0444] [Figure 1A]Figures 1A and 1B show (Figure 1A) Cas9-PiggyBac fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active PiggyBac (PB) transposase) and (Figure 1B) Cas9-SB100 fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active Sleeping This shows the percentage of cells with exogenous nucleic acid sequences integrated into their genomes after transfection using Beauty(SB100) transposases. Vectors were constructed in which the 3' end of Cas9 was ligated to the 5' end of each transposase by a GGS linker (SEQ ID NOs. 48, 49) (hCas9PB, nCas9PB, dCas9PB, hCas9SB, nCas9SB, and dCas9SB). Other vectors were constructed in which the 3' end of each transposase was ligated to the 5' end of Cas9 by a GGS linker (SEQ ID NOs. 48, 49). (PBhCas9, PBnCas9, PBdCas9, SBhCas9, SbnCas9, and SBdCas9). "PiggyBac" (Figure 1A) and "SB100" (Figure 1B) were used as positive controls, and transposons encoding RFP (shown as "Episome RFP" in Figure 1A) and GFP (shown as "Episome GFP" in Figure 1B) alone were used as negative controls. Figure 1C is a different representation of Figure 1A, showing the transposability activity by PB and Cas9 in different configurations. [Figure 1B]Figures 1A and 1B show (Figure 1A) Cas9-PiggyBac fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active PiggyBac (PB) transposase) and (Figure 1B) Cas9-SB100 fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active Sleeping This shows the percentage of cells with exogenous nucleic acid sequences integrated into their genomes after transfection using Beauty(SB100) transposases. Vectors were constructed in which the 3' end of Cas9 was ligated to the 5' end of each transposase by a GGS linker (SEQ ID NOs. 48, 49) (hCas9PB, nCas9PB, dCas9PB, hCas9SB, nCas9SB, and dCas9SB). Other vectors were constructed in which the 3' end of each transposase was ligated to the 5' end of Cas9 by a GGS linker (SEQ ID NOs. 48, 49). (PBhCas9, PBnCas9, PBdCas9, SBhCas9, SbnCas9, and SBdCas9). "PiggyBac" (Figure 1A) and "SB100" (Figure 1B) were used as positive controls, and transposons encoding RFP (shown as "Episome RFP" in Figure 1A) and GFP (shown as "Episome GFP" in Figure 1B) alone were used as negative controls. Figure 1C is a different representation of Figure 1A, showing the transposability activity by PB and Cas9 in different configurations. [Figure 1C]Figures 1A and 1B show (Figure 1A) Cas9-PiggyBac fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active PiggyBac (PB) transposase) and (Figure 1B) Cas9-SB100 fusion protein (human Cas9 (hCas9), nickase Cas9 (nCas9), or inactive Cas9 (dCas9) and highly active Sleeping This shows the percentage of cells with exogenous nucleic acid sequences integrated into their genomes after transfection using Beauty(SB100) transposases. Vectors were constructed in which the 3' end of Cas9 was ligated to the 5' end of each transposase by a GGS linker (SEQ ID NOs. 48, 49) (hCas9PB, nCas9PB, dCas9PB, hCas9SB, nCas9SB, and dCas9SB). Other vectors were constructed in which the 3' end of each transposase was ligated to the 5' end of Cas9 by a GGS linker (SEQ ID NOs. 48, 49). (PBhCas9, PBnCas9, PBdCas9, SBhCas9, SbnCas9, and SBdCas9). "PiggyBac" (Figure 1A) and "SB100" (Figure 1B) were used as positive controls, and transposons encoding RFP (shown as "Episome RFP" in Figure 1A) and GFP (shown as "Episome GFP" in Figure 1B) alone were used as negative controls. Figure 1C is a different representation of Figure 1A, showing the transposability activity by PB and Cas9 in different configurations. [Figure 2A] Figure 2A shows a plasmid construct encoding the Cas9 / PB fusion protein. [Figure 2B]Figure 2B shows the percentage of cells with exogenous nucleic acid sequences integrated into the genome by fusion constructs formed by human Cas9-PiggyBac ("targeted HCas9") or nickase Cas9-PiggyBac ("targeted NCas9"). The 3' end of Cas9 was ligated to the 5' end of the transposase by a linker. "Non-targeted" is a control for total insertion (PiggyBac alone), and "episome" is a negative control for non-integration (transposon alone). [Figure 3] Figure 3 shows an exemplary ZFP-integrase fusion protein. The ZFP and integrase are linked by a GGS sequence. NLS refers to the nuclear localization sequence. [Figure 4] Figure 4 shows the lentiviral titers of wild-type integrase lentivirus (LV), empty virus particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZP-IN(AAVS1)), non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN), and wild-type integrase lentivirus with wild-type integrase (LV+IN). (') indicates a technical repeat. [Figure 5] Figure 5 shows the percentage of cells that incorporated the exogenous nucleic acid sequence into their genome (total integration) after infection with wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZP-IN(AAVS1)), non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN), and wild-type integrase lentivirus with wild-type integrase (LV+IN). For each condition, from left to right, the first column represents day 3, the second column represents day 5, the third column represents day 7, the fourth column represents day 10, and the fifth column represents day 12. [Figure 6] Figure 6 shows images of chromosomes containing representative AAVS1 integration and non-integration sites. The asterisks represent the AAVS1 sites on chromosome 19, the triangles represent non-target integration sites, and the diamonds represent targeted integration sites. [Figure 7A] Figure 7A shows the viral titers generated by wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)). [Figure 7B] Figure 7B shows the percentage of cells that have incorporated the exogenous nucleic acid sequence into their genome (total integration) after infection with wild-type integrase lentivirus (LV), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)). [Figure 7C] Figure 7C shows the percentage of cells that incorporated exogenous nucleic acid sequences into their genome after infection with wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)). [Figure 7D]Figure 7D shows the percentage of cells that incorporated exogenous nucleic acid sequences into their genome after infection with wild-type integrase lentivirus (LV), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with a ZFP-IN fusion protein targeting the AAVS1 site (NILV+ZP-IN(AAVS1)), and non-integrated lentivirus with a ZFP-IN fusion protein targeting the CCR5 site (NILV+ZP-IN(CCR5)). [Figure 8A] Figures 8A-8C show lentiviral titers (Figure 8A), the percentage of CAR-expressing cells on day 3 and day 14 (Figure 8B), and the percentage of CD3-expressing cells (Figure 8C). Jurkat cells were infected with lentiviruses under several conditions. Several conditions were identified: wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZFP-IN(TRCa-1)), and non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN). NILV showed a drastic decrease in titer, and the expression of IN WT in virus-producing cells or complementarity with fused ZNF-IN did not have a rescue effect on either titer or integration ability. Furthermore, when integration was directed toward the TCR gene locus (CD3 protein expression), the cells did not lose CD3 expression. This suggests the need for further factors for transcomplementation, such as VPR proteins (especially in the context of this cell line). [Figure 8B]Figures 8A-8C show lentiviral titers (Figure 8A), the percentage of CAR-expressing cells on day 3 and day 14 (Figure 8B), and the percentage of CD3-expressing cells (Figure 8C). Jurkat cells were infected with lentiviruses under several conditions. Several conditions were identified: wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZFP-IN(TRCa-1)), and non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN). NILV showed a drastic decrease in titer, and the expression of IN WT in virus-producing cells or complementarity with fused ZNF-IN did not have a rescue effect on either titer or integration ability. Furthermore, when integration was directed toward the TCR gene locus (CD3 protein expression), the cells did not lose CD3 expression. This suggests the need for further factors for transcomplementation, such as VPR proteins (especially in the context of this cell line). [Figure 8C]Figures 8A-8C show lentiviral titers (Figure 8A), the percentage of CAR-expressing cells on day 3 and day 14 (Figure 8B), and the percentage of CD3-expressing cells (Figure 8C). Jurkat cells were infected with lentiviruses under several conditions. Several conditions were identified: wild-type integrase lentivirus (LV), empty viral particles (LVO), non-integrated lentivirus (NILV), non-integrated lentivirus with wild-type integrase (NILV+IN), non-integrated lentivirus with ZFP-integrase fusion protein (NILV+ZFP-IN(TRCa-1)), and non-integrated lentivirus with Cas9-integrase fusion protein (NILV+Cas-IN). NILV showed a drastic decrease in titer, and the expression of IN WT in virus-producing cells or complementarity with fused ZNF-IN did not have a rescue effect on either titer or integration ability. Furthermore, when integration was directed toward the TCR gene locus (CD3 protein expression), the cells did not lose CD3 expression. This suggests the need for further factors for transcomplementation, such as VPR proteins (especially in the context of this cell line). [Figure 9A]Figures 9A-9B show the titers for WT lentivirus and two different integrase-deficient viral systems (NILV and TAA, the latter indicating the introduction of a stop codon at the beginning of the IN-coding region in the lentiviral packaging plasmid) individually, or for those heterologously complemented by IN or VPR_IN fusion. Titers were detected by fluorescence cytometry analysis on post-infection day 3 (Figure 9A). Figure 9B shows the relative integration efficiency of heterologously complementary integration mechanisms (machineeries) demonstrating the advantages of VPR protein fusion to IN for heterologously complementary. WT: Lentivirus produced using WT IN; NILV: Lentivirus produced using non-integrated IN with two mutations in its catalytic center; TAA: Lentivirus produced using IN-deficient IN in which the protein is not expressed; +IN: Lentivirus heterologously complemented using IN; +VPR-IN: Lentivirus heterologously complemented using IN fused to VPR at the C-terminus. [Figure 9B] Figures 9A-9B show the titers for WT lentivirus and two different integrase-deficient viral systems (NILV and TAA, the latter indicating the introduction of a stop codon at the beginning of the IN-coding region in the lentiviral packaging plasmid) individually, or for those heterologously complemented by IN or VPR_IN fusion. Titers were detected by fluorescence cytometry analysis on post-infection day 3 (Figure 9A). Figure 9B shows the relative integration efficiency of heterologously complementary integration mechanisms (machineeries) demonstrating the advantages of VPR protein fusion to IN for heterologously complementary. WT: Lentivirus produced using WT IN; NILV: Lentivirus produced using non-integrated IN with two mutations in its catalytic center; TAA: Lentivirus produced using IN-deficient IN in which the protein is not expressed; +IN: Lentivirus heterologously complemented using IN; +VPR-IN: Lentivirus heterologously complemented using IN fused to VPR at the C-terminus. [Figure 10A]Figure 10A shows the structure of a nucleic acid construct formed by an insertion domain having a DNA-binding domain fused by a linker and a programmable DNA-recognition domain. [Figure 10B] Figure 10B shows the fusion of Cas9 and transposase linked by a linker in different configurations. [Figure 11] Figure 11 shows the results of Cas9 activity in Cas9 bound to hyPB using various linker sizes and compositions. Cas9 activity was measured by sequencing the gRNA target site and analyzing the indel frequency using CRISPR-GA. Two different gRNAs targeting the AAVS1 site were used. The linkers used were SEQ ID NOs. 50–63. [Figure 12] Figure 12 shows the results for programmable transposer zegene trap transfer efficiency. RFP fluorescence was measured by flow cytometry 10 days after transfection. The importance of linker length and composition in target insertion was determined using different linkers. Average of two independent experiments. Linkers used are SEQ ID NOs. 50–63. [Figure 13] Figure 13 shows the results for the hcas9_PB linker. Target translocation efficiency of different cas9-PB linker constructs was measured using fragmented GFP cell lines with two different gRNAs. GFP expression was measured by flow cytometry 72 hours after transfection. [Figure 14]Figure 14 outlines the split GFP reporter cell lines generated for high-throughput analysis of a library of various hyPB variants and for confirmation of individual variants. Using the Sleeping Beauty 100x system, a splicing acceptor (SA) downstream of the target region, followed by half of the GFP coding sequence (Ct-GFP), was introduced into the genome of Hek293T cells. For this screening, the PiggyBac transposon adjacent to the terminal repeat sequence (ITR) was either the complete RPF expression cassette, followed by the promoter and the remaining half of GFP (Nt-GFP) and the splicing donor (SD); or just half of the GFP fragment; as shown in the figure. [Figure 15] Figure 15 shows the results of targeted translocation of hcas9_PB selective mutants. Targeted translocation efficiency of hcas9_PB D450N and hcas9_PB R372A K375A D450. GFP expression was measured by flow cytometry 72 hours after transfection. Average of 4 independent experiments. [Figure 16] Figure 16 shows the results for random and targeted translocation of hcas9_PB selected mutants. Targeted translocation efficiency and random translocation efficiency for hcas9_PB D450N and hcas9_PB R372A K375A D450. GFP expression was measured by flow cytometry 72 hours after transfection, and RFP expression was measured by flow cytometry 15 days after transfection, normalized by RFP fluorescence 48 hours after transfection, which was assumed to be the transfection efficiency. [Figure 17] Figure 17 is a diagram illustrating the fusion of a ZFP with a transposase linked by linkers of various configurations. [Figure 18]Figure 18 shows the results of ZFP-PB fusion protein target transposition. Target transposition efficiency of ZFP_hyPB or ZFP_hyPBD450N in N and C-terminal conformations. GFP expression was measured by flow cytometry 5 days after transfection. More than 1 independent replicate. ZFP_PB: ZFP fused with hyPB in the C-terminal configuration using an XTEN linker; PB_ZFP: ZFP fused with hyPB in the N-terminal configuration using an XTEN linker; ZFP_450: ZFP fused with hyPB(D450N) in the C-terminal configuration using an XTEN linker; 450_ZFP: ZFP fused with hyPB(D450N) in the N-terminal configuration using an XTEN linker; hyPB: non-recombinant hyPB; 1 / 2GFP: control transposon alone. [Figure 19] Figure 19 shows a diagram of the analytical methods used in screening libraries of PiggyBac mutations. [Figure 20] In Figure 20, PiggyBac 1116 base pairs containing all library variants were sequenced using Illumina NGS technology. Except for variants 450 and 465, the I7 index primers were replaced with custom primers to enable complete sequencing of the different variants. [Figure 21A] Figures 21A-21B show the results of hyPB library diversity generation. Figure 21A is an example of a sort plot. Positive target inclusion hits (GFP fluorescence) were selected at gate P4, and negative target inclusion hits (no GFP fluorescence) were selected at gate P5. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining. Figure 21B shows the results of dual plasmid transfection efficiency. Transfection efficiency was measured by transfecting GFP and RFP plasmids equimolarly with 1 / 2 GFP and gRNA transfections on the same day and under the same conditions. Gate P8 selects dual plasmid transfections. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining. [Figure 21B]Figures 21A-21B show the results of hyPB library diversity generation. Figure 21A is an example of a sort plot. Positive target inclusion hits (GFP fluorescence) were selected at gate P4, and negative target inclusion hits (no GFP fluorescence) were selected at gate P5. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining. Figure 21B shows the results of dual plasmid transfection efficiency. Transfection efficiency was measured by transfecting GFP and RFP plasmids equimolarly with 1 / 2 GFP and gRNA transfections on the same day and under the same conditions. Gate P8 selects dual plasmid transfections. Non-viable cells and debris were negatively selected at the preliminary gate by DAPI staining. [Figure 22A] Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22B]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22C]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22D]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22E]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22F]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22G]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22H]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22I]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22J]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 22K]Figures 22A–22K show the results of library screening analysis comparing positive hits to negatives. Figures 22A–22B: Sequencing of the bulk library is shown as quality control; the majority of variants are shown only once. The logos of representative piggyback libraries from the bulk shown correspond to the following amino acid positions: 1-R245; 2-R275; 3-R277; 4-G325; 5-N347; 6-S351; 7-R372; 8-K375; 9-R388; 10-T560; 11-S564; 12-S573; 13-M589; 14-S592; 15-F594. Furthermore, the logos of selected negative cells are shown in a similar pattern to the bulk library. Figures 22C–22K correspond to three independent repeats of positive hits; the variant requiring a positive logo (bottom) and the Top1 variant after selection (top). The logos for the top 5 and top 10 variants are also shown. In the left panels of B and C, the relative abundance of the Piggyback variant in the positive and negative selected populations is shown on a log2 scale. [Figure 23A] Figure 23A shows the positive variants of Top1 and Top3 from three independent replicates. There is a difference of only one amino acid at position 254. [Figure 23B] Figure 23B shows the three top1 variants identified in three independent iterations. WT hyPB is also shown for reference. [Figure 24A] Figure 24A shows the most overrepresented variants in GFP-positive and RFP-positive cells. Clustering of GPF, targeted insertions; RPF, random insertions, and negative populations are shown. [Figure 24B] Figures 24B and 24C show variants observed among positive hits in more than one independent replicate. Rep: Independent experimental replicate; Pos: Positive cells with target integration; Neg: Negative cells where target integration did not occur. [Figure 24C]Figures 24B and 24C show variants observed among positive hits in more than one independent replicate. Rep: Independent experimental replicate; Pos: Positive cells with target integration; Neg: Negative cells where target integration did not occur. [Figure 25] Figure 25 shows a histogram of variant covariates, representing the proportion of variants found together with others in positive samples divided by the proportion of negative samples. In addition to variants included in the library design, variants randomly introduced by lentiviral retrotranscriptase during viral library generation were analyzed. Some of these new variants are associated with positive hits and perform targeted incorporation for combinations. Examples of D450N and W465A are shown. [Figure 26] Figure 26 shows that recombinant hyPB showed greater enhancement in target integration compared to WT hyPB when fused with Cas9. Cas9 was fused to hyPB or different combinations of hyPB variants (Unilarge-A:D450N; Unilarge-B:R245A / D450N; Unilarge-C:R245A / G325A / D450N / S573P; Unilarge-D:R245A / G325A / S573P) using a 4GGS linker and reporter cell line system. [Figure 27]Figure 27 shows the results of heterogeneous complementation of integrase-deficient cells. Viral production efficiency measured on day 2 and integration capacity measured on day 7 were evaluated for different systems in Hek293T cells. Western blotting showed the presence of intergeneric IN in the viral particles. Viral production efficiency and its integration capacity were evaluated by infecting Hek293T cells with different conditions of integration-deficient virus and heterogeneous complementar virus. Cells were passaged for 7 days until episomal signaling was no longer detectable, and GFP signaling was analyzed by flow cytometry on days 2, 5, and 7. Since NILV is close to WT during production, different production efficiencies can be detected for different systems. In all cases, a clear rescue of integration activity was observed when heterogeneous complementation was performed using WT-HIV_IN. Proof of IN being loaded into the heterogeneous complementarized system was obtained by Western blotting. WT: Lentivirus produced with WT IN; NILV: Lentivirus produced with non-integrated IN and having two mutations in its catalytic center; TAA: IN-deficient lentivirus in which the protein is not expressed because a stop codon is present at the beginning of the IN coding sequence; TAAx3: Lentivirus produced with IN-deficient IN in which the protein is not expressed because three consecutive stop codons are present at the beginning of the IN coding sequence; Delta-IN: Lentivirus produced with IN-deficient IN in which the IN coding sequence has been removed; Delta-IN_cPPT: Lentivirus produced with IN-deficient IN in which the IN coding sequence has been replaced with a central polypyrimidine trac(cPPT) sequence; +VPR-IN: Lentivirus that is interspecies complementary to IN, fused with VPR at the C-terminus.
Claims
1. nucleic acid constructs, a) A first polynucleotide sequence comprising a nucleic acid encoding a first DNA-binding protein engineered to bind to a specific genomic DNA sequence in the genome, wherein the first DNA-binding protein is a zinc finger protein or a Cas9 protein, b) A second polynucleotide sequence comprising a nucleic acid encoding a second DNA-binding protein that enables insertion of an exogenous nucleic acid into the genome, wherein the second DNA-binding protein is recombinant high-activity PiggyBac with improved specificity of insertion of the exogenous nucleic acid into the genome compared to high-activity PiggyBac, The nucleic acid construct encodes a fusion protein comprising the first DNA-binding protein and the second DNA-binding protein, The fusion protein enables the insertion of the exogenous nucleic acid into a specific site of the genome, The recombinant high-activity PiggyBac transposase contains one or more amino acid recombinations selected from R245A, R275A, R277A, R275A / R277A, G325A, N347A, N347S, S351E, S351P, S351A, R372A, K375A, R388A, D450N, W465A, T560A, S564P, S573A, M589V, S592G, or F594L, corresponding to the amino acid sequence SEQ ID NO: 9 of the high-activity PiggyBac. Nucleic acid constructs.
2. c) Containing a polynucleotide sequence that includes nucleic acids encoding a linker, The nucleic acid construct encodes a fusion protein comprising the first DNA-binding protein, the second DNA-binding protein, and the linker between the first DNA-binding protein and the second DNA-binding protein. The nucleic acid construct according to claim 1.
3. The Cas9 protein is selected from the group consisting of human Cas9, nickase Cas9, and inactive (dead) Cas9. The nucleic acid construct according to claim 1.
4. The Cas9 protein comprises the Cas9 active DNA cleavage domain and the Cas9 gRNA binding domain. The nucleic acid construct according to claim 1.
5. The aforementioned zinc finger protein contains six domains. 2 H 2 Zinc finger proteins, The nucleic acid construct according to claim 1.
6. The linker includes an XTEN sequence or a GGS sequence. The nucleic acid construct according to claim 2.
7. The 3' end of the first polynucleotide sequence is connected to the 5' end of the second polynucleotide. The nucleic acid construct according to claim 1.
8. a) The first DNA-binding protein is a Cas9 protein or a zinc finger protein, b) The second DNA-binding protein is a recombinant high-activity PiggyBac in which the specificity of insertion of the exogenous nucleic acid into the genome is improved compared to the high-activity PiggyBac. The nucleic acid construct includes a polynucleotide sequence comprising a nucleic acid encoding a linker containing an XTEN sequence or a GGS sequence, The 3' end of the first polynucleotide sequence is connected to the 5' end of the second polynucleotide. The recombinant high-activity PiggyBac transposase contains one or more amino acid recombinations selected from R245A, R275A, R277A, R275A / R277A, G325A, N347A, N347S, S351E, S351P, S351A, R372A, K375A, R388A, D450N, W465A, T560A, S564P, S573A, M589V, S592G, or F594L, corresponding to amino acid sequence number 9 of the high-activity PiggyBac. The nucleic acid construct according to claim 2.
9. The second DNA-binding protein is a recombinant, highly active PiggyBac transposase containing an amino acid sequence selected from the group consisting of SEQ ID NOs: 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129. The recombinant high-activity PiggyBac exhibits higher specificity for DNA integration into the genome compared to high-activity PiggyBac. The nucleic acid construct according to claim 1.
10. A vector comprising the nucleic acid construct described in claim 1, The vector is suitable for expression in mammalian cells, yeast cells, insect cells, plant cells, fungal cells, or algal cells. vector.
11. A host cell comprising the nucleic acid construct described in claim 1 or the vector described in claim 10.
12. A fusion protein obtained from the expression of the nucleic acid construct described in claim 1.
13. A composition comprising a nucleic acid construct according to claim 1, a vector according to claim 10, or a fusion protein according to claim 12, and a polynucleotide sequence encoding an exogenous nucleic acid for insertion into a genome, The composition is contained in or bound to the packaging vector. composition.
14. The nucleic acid construct is in the form of RNA or DNA, and the polynucleotide sequence encoding the exogenous nucleic acid is in the form of RNA or DNA. The composition according to claim 13.
15. The packaging vector is a nanoparticle or a lentiviral particle. The composition according to claim 13.
16. A composition comprising a nucleic acid construct according to claim 1, a vector according to claim 10, or a fusion protein according to claim 12, for preparing a pharmaceutical product for treating a genetic disorder in a subject requiring such treatment.
17. An ex vivo method for controlled site-specific incorporation of a single or multiple copies of an exogenous nucleic acid sequence into a cell, wherein the method is: a) A step of delivering the nucleic acid construct according to claim 1, the vector according to claim 10, or the fusion protein according to claim 12 to the cells, and b) The step of delivering the exogenous nucleic acid to the cell, The binding of the fusion protein to a specific genomic DNA sequence in the genome of the cell causes the cleavage of the genome and the incorporation of one or more copies of the exogenous nucleic acids into the genome of the cell. method.
18. A recombinant highly active PiggyBac transposase containing recombinant amino acid sequence SEQ ID NO: 9, wherein the amino acid number corresponding to SEQ ID NO: 9 i. The amino acid at position 245 is A, and ii. The amino acid at position 275 is R or A, and iii. The amino acid at position 277 is R or A, and iv. The amino acid at position 325 is A or G, and v. The amino acid at position 347 is N or A, and vi. The amino acid at position 351 is E, P or A, and vii. The amino acid at position 372 is R, and viiii. The amino acid at position 375 is A, and The amino acid at position 450 is either D or N, and x. The amino acid at position 465 is W or A, and xi. The amino acid at position 560 is T or A, and xi. The amino acid at position 564 is P or S, and xiiii. The amino acid at position 573 is S or A, and xiv. The amino acid at position 592 is G or S, and The amino acid at position xv. 594 is either L or F. Recombinant high-activity PiggyBac transposase.
19. It contains an amino acid sequence selected from the group consisting of SEQ ID NOs: 120, 121, 122, 123, 124, 125, 126, 127, 128, and 129. Recombinant high-activity PiggyBac exhibits higher specificity for DNA integration into the genome compared to high-activity PiggyBac. Recombinant highly active PiggyBac transposase according to claim 18.
Citation Information
Patent Citations
Cas9 Retro-integrase and Cas9 Recombinase Systems for Targeted Integration of DNA Sequences into the Genomes of Cells or Organisms
JP2018513681A
Improvements to eukaryotic transposase mutants and transposon end compositions for modifying nucleic acids and methods for production and use in the generation of sequencing libraries
WO2015157611A2
Methods of genome engineering by nuclease-transposase fusion proteins
WO2018175872A1