Synergistic NHEJ inhibition to obtain enrichment free sequential insertion of genes > 4 kb
By introducing double-strand breaks and using 53BP1 and DNA-PKcs inhibitors with dual AAV delivery, the method overcomes the 4 kb limit of AAV gene therapy, enhancing gene insertion efficiency and functional correction in airway stem cells.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- RES INST AT NATIONWIDE CHILDRENS HOSPITAL
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
Smart Images

Figure US2025055019_21052026_PF_FP_ABST
Abstract
Description
Atty. Dkt. No.: 106887-9660 SYNERGISTIC NHEJ INHIBITION TO OBTAIN ENRICHMENT FREE SEQUENTIAL INSERTION OF GENES > 4 KB CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority under 35 U. S. C. § 119(e) to U. S. Patent Provisional Application No. 63 / 720,051, filed November 13, 2024, the entirety of which is incorporated herein by reference in its entirety.STATEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with government support under Grant No. HL151900 awarded by the National Institutes of Health (NIH). The government has certain rights in the invention.BACKGROUND
[0003] Gene therapies have revolutionized the treatment of many genetic disorders. Many of these genetic therapies use the adeno-associated virus (AAV) to deliver the missing genes. AAV has also been used in combination with CRISPR-Cas9 to replace genes in primary stem cells ex-vivo. The creation of double-stranded breaks (DSB) in the genome can be used to boost targeted gene insertion mediated by homologous recombination (HR). DSBs can be induced in a site-specific manner using nucleases such as Cas9 ribonuclear protein (RNP) complexed to a single guide RNA (sgRNA). These DSBs are predominantly repaired using the non-homologous end joining pathway (NHEJ). However, the delivery of an HR template with left (LHA) and right (RHA) homology arms (HA) that resemble the DSB site enables the targeted insertion of exogenous sequences using HR. Compared to adenoviral and naked single stranded DNA templates, adeno-associated virus (AAV)-based HR templates result in highly efficient gene correction in primary stem cells. However, AAV capsids can only package up to 4.7 kb worth of genetic material. Thus, it is not possible to replace genes over 4kb using Cas9 and AAV. Several disorders such as cystic fibrosis (CF), epidermolysis bullosa (EB), Duchenne muscular dystrophy, DOCK8 immunodeficiency, Usher’s syndrome, Surfactant protein deficiency and others are caused by genes over 4 kb. Therefore, there is a need for methodologies to deliver such genes that are over 4kb in length.14921-0232-3317.3Atty. Dkt. No.: 106887-9660SUMMARY OF THE DISCLOSURE
[0004] In one aspect, provided herein is a method for inserting a polynucleotide between about 4.5Kb and about 8Kb and lacking a selection marker into a target genomic position in a cell, the method comprising, or alternatively consisting essentially of, or yet further consisting of (i) introducing a double strand break at the target genomic position, (ii) delivering to the cell a first fragment of the polynucleotide using a first adeno associated virus (AAV) and a second fragment of the polynucleotide using a second AAV, wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide comprise the polynucleotide, and (iii) culturing the cell in the presence of a p53-binding protein 1 (53BP1) inhibitor and a DNA-dependent protein kinase catalytic subunit (DNA-PKcs) inhibitor to complete sequentially inserting the first fragment of the polynucleotide and the second fragment of the polynucleotide into the target genomic position.
[0005] In some embodiments, the 53BP1 inhibitor is selected from any of SEQ ID NO: 1-6 or an equivalent of each thereof, or a dominant-negative mutant of 53BP1.
[0006] In some embodiments, the double strand break is introduced by a CRISPR-Cas system comprising, or alternatively consisting essentially of, or yet further consisting of a Cas protein and a guide RNA (gRNA) specific for the target genomic position.
[0007] In some embodiments, the Cas protein is fused to the 53BP1 inhibitor. In some embodiments, the 53BP1 inhibitor comprises, or alternatively consists essentially of, or yet further consists of DN1S as shown in SEQ ID NO: 13, or DN1 as shown in SEQ ID NO: 14, or an equivalent of each thereof.
[0008] In some embodiments, the DNA-PKcs inhibitor is selected from AZD7648, M3814, VX984, KU57788, BAY8400, or LTURM34.
[0009] In some embodiments, the 53BP1 inhibitor has a final concentration in culture between about 5μM and about 25μM. In some embodiments, the DNA-PKcs inhibitor has a final concentration in culture between about 0.25μM and about 0.5μM. In some embodiments, the culturing step of (iii) is between about 16 hours and about 48 hours.
[0010] In some embodiments, the first and the second AAVs are selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV-DJ, or a variant of each thereof, and the first and second AAVs are independently the same or different from each other.24921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0011] In some embodiments, the polynucleotide comprises, or alternatively consists essentially of, or yet further consists of a gene encoding Cystic fibrosis transmembrane conductance regulator (CFTR), Dedicator of cytokinesis 8 (DOCK8) or ATP binding cassette subfamily A member 3 (ABCA3).
[0012] In some embodiments, the double strand break is introduced by a CRISPR-Cas system. In some embodiments, the inserting is achieved by homology directed repair (HDR).
[0013] In some embodiments, the first fragment of the polynucleotide further comprises, or alternatively consists essentially of, or yet further consists of a 5’ homology arm at the 5’ end of the first fragment and a 3’ homology arm at the 3’ end of the first fragment, and wherein the second fragment of the polynucleotide further comprises, or alternatively consists essentially of, or yet further consists of a 5’ homology arm at the 5’ end of the second fragment and a 3’ homology arm at the 3’ end of the second fragment.
[0014] In some embodiments, the 5’ homology arm of the first fragment is homologous to a sequence 5’ of the double strand break, and wherein the 3’ homology arm of the first fragment is homologous to the 5’ homology arm of the second fragment, and wherein the 3’ homology arm of the second fragment is homologous to a sequence 3’ of the double strand break.BRIEF DESCRIPTION OF THE FIGURES
[0015] FIGS. 1A-1F. Combination of UbVa and AZD-7648 improves the efficiency of GFP insertion only modestly. UbVa is a peptide that blocks 53BP1 activity. AZD-7648 is a DNA-PKcs inhibitor. Applicant tested the insertion of GFP (a single polynucleotide) in the HBB locus in the absence or presence of 5μM UbVa and 0.5μM AZD-7648. (A) Mock insertion (no GFP, no inhibitor). (B) GFP insertion in the absence of any inhibitors. (C) GFP insertion in the presence of UbV-A alone. (D) GFP insertion in the presence of AZD-7648 alone. (E) GFP insertion in the presence of UbV-A and AZD-7648. Flow cytometry was used to measure the fraction of GFP+ cells. Compared to airway stem cells edited in the presence of neither compound (B), cells edited in the presence of UbVa (C) show more gene insertion (GFP+ cells). (D)The use of AZD-7648 is more effective in improving gene insertion than UbVa. (E) The combination of UbV-A and AZD-7648 only improves single gene insertion modestly. (F) Quantification of (A)-(E).
[0016] FIGS. 2A-2C. UbV-A treatment improved gene insertion of CFTR cDNA. (A) Applicant tested the use of UbVa to improve sequential gene insertion of the CFTR cDNA34921-0232-3317.3Atty. Dkt. No.: 106887-9660 and a truncated CD19 expression cassette packaged into two AAVs. (B) Normally <10% of edited airway stem cells are tCD19+ indicating the insertion of the sequences present in both AAVs. (C) The use of 5-25 pM UbVa improves gene insertion by -50-100%.
[0017] FIGS. 3A-3F. Combining DNA-PKcs and 53BP1 inhibition improves CFTR cDNA insertion by 10-fold. (A) Airway stem cells were edited using the two AAV system to insert the CFTR cDNA and tCD19 cassette. (B) Mock electroporated cells showed no tCD19+ cells. (C) Cells edited in the absence of any DNA-PKcs or 53BP1 inhibition showed -5% tCD19 cells. (D) The use of UbVa resulted in 10% tCD19+ cells while the use of (E) AZD-7648 resulted in 20% tCD19+ cells. (F) The use of both AZD-7648 and UbVa resulted in a surprising increase, where >50% of the cells were tCD19+.
[0018] FIGS. 4A-4D. Corrected CF airway stem cells produce epithelia with restored CFTR function. (A) CF airway stem cells edited as shown previously were differentiated in airliquid interface cultures. CFTR function was assessed by measuring their responses to Forskolin and CFTRinh-172. Cells edited in the presence of AZD-7648 and UbVa (which had the highest level of correction) showed restored CFTR function. (B) The drop in current after the addition of CFTRinh-172 was quantified and compared between different conditions. Cells edited in the presence of AZD-7648 and UbVa showed restored CFTR function comparable to WT controls. (C) Percent of off-target alleles with INDELs was not increased by NHEJ inhibition. (D) The drop in current after the addition of CFTRinh-172 was quantified and compared between different conditions. Cells edited in the presence of AZD-7648 and 53BP1 inhibiting peptide UbV-A (HDR Enhancer peptide or HEP) (5 pM) showed restored CFTR function comparable to WT controls.
[0019] FIG. 5. Blocking DNA-PKcs and 53BP1 is more effective than blocking DNA-PKcs and Pol-Theta. Gene insertion was compared in the presence of AZD-7648, PolQi2 and UbV-A alone and in combination. AZD-7648 combined with UbV-A improved gene insertion from -2% to 20% in these experiments. Experiment was repeated in cells from two different donors (2 replicates per donor)
[0020] FIGS. 6A-6B. UbV-A and AZD-7648 together do not improve insertion of ssDNA templates. WT seq is shown by SEQ ID NO: 7, Human AF508 sequence is shown by SEQ ID NO: 8 and the correction template is shown by SEQ ID NO: 9. (A) Single stranded DNA templates are also used as an alternative to AAV. Here, Applicant used 200 bp ssDNA templates to insert 22-25 bases targeting the F508del mutation in exon 11 of the CFTR locus.44921-0232-3317.3Atty. Dkt. No.: 106887-9660 Prior experiments targeted exon 1 of CFTR to replace the complete cDNA. Editing efficiency was measured 4 days after editing by PCR and Sanger sequencing. (B) Gene insertion was compared in the presence of 0.5 pM AZD-7648, and 0.5 pM AZD-7648 in combination with 5 pM UbV-A. AZD-7648 alone and AZD-7648 combined with UbV-A similarly improved gene insertion from ~2% to 20% in these experiments. Experiment was repeated in cells from two different donors (2 replicates per donor).
[0021] FIGS. 7A-7F. (A) Long genes can be split into two and packaged into two AAVs for gene insertion. The first template contains a copy of the original sgRNA site, which initiates a second DSB. (B) Simultaneous inhibition of DNA-PKcs and 53BP1 using AZD-7648 and HEP is more effective in improving sequential insertion compared to other combinations. (C) Sequential insertion of the CFTR cDNA and tCD19 (~6.5 kb) is possible in >40% of treated cells with simultaneous DNA-PKcs and 53BP1 inhibition. (D) Cells edited in the presence of AZD-7648 and HEP do not show a reduction in proliferation. (E) Expression of KRT5 and P63 (markers of sternness) in uABCs. (F) Similar improvement in sequential gene insertion by AZD-7648 and HEP is also observed in the CCR5 locus in uABCs. ****, ***;* indicate p < 0.0001, 0.005 and 0.05 after analysis using one way ANOVA followed by post-hoc test. Each dot represents a biological replicate (i.e. cells from a separate donor).
[0022] FIG. 8. Inhibition of both DNA-PKcs and 53BP1 reduces INDELS at an off-target locus compared to the inhibition of DNA-PKcs aloneDETAILED DESCRIPTIONDefinitions
[0023] As it would be understood, the section or subsection headings as used herein is for organizational purposes only and are not to be construed as limiting and / or separating the subject matter described.
[0024] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred methods, devices, and materials are now described. All technical and patent publications cited herein are incorporated herein by reference in their entirety. Nothing herein is to be construed as an54921-0232-3317.3Atty. Dkt. No.: 106887-9660 admission that the disclosure is not entitled to antedate such disclosure by virtue of prior disclosure.
[0025] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of tissue culture, immunology, molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd edition; the series Ausubel et al. eds. (2007) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N. Y.); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th edition; Gait ed. (1984) Oligonucleotide Synthesis; U. S. Patent No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London);Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology; Manipulating the Mouse Embryo: A Laboratory Manual, 3rd edition (Cold Spring Harbor Laboratory Press (2002)); Sohail (ed.) (2004) Gene Silencing by RNA Interference: Technology and Application (CRC Press).
[0026] As used in the specification and claims, the singular form “a,” “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes a plurality of cells, including mixtures thereof.
[0027] As used herein, the term “comprising” is intended to mean that the compounds, agents, compositions, and methods include the recited elements, but not exclude others.“Consisting essentially of’ when used to define compounds, agents, compositions, and methods, shall mean excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants, e.g., from the isolation and purification method and pharmaceutically acceptable carriers, preservatives, and the like. “Consisting of’ shall mean64921-0232-3317.3Atty. Dkt. No.: 106887-9660 excluding more than trace elements of other ingredients. Embodiments defined by each of these transition terms are within the scope of this technology.
[0028] All numerical designations, e.g., pH, temperature, time, concentration, and molecular weight, including ranges, are approximations which are varied (+) or (-) by increments of 1, 5, or 10%. It is to be understood, although not always explicitly stated that all numerical designations are preceded by the term “about.” It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art.
[0029] The term “about,” as used herein when referring to a measurable value such as an amount or concentration and the like, is meant to encompass variations of 20%, 10%, 5%, 1 %, 0.5%, or even 0.1 % of the specified amount.
[0030] As used herein, comparative terms as used herein, such as high, low, increase, decrease, reduce, or any grammatical variation thereof, can refer to certain variation from the reference. In some embodiments, such variation can refer to about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 80%, or about 90%, or about 1 fold, or about 2 folds, or about 3 folds, or about 4 folds, or about 5 folds, or about 6 folds, or about 7 folds, or about 8 folds, or about 9 folds, or about 10 folds, or about 20 folds, or about 30 folds, or about 40 folds, or about 50 folds, or about 60 folds, or about 70 folds, or about 80 folds, or about 90 folds, or about 100 folds or more higher than the reference. In some embodiments, such variation can refer to about 1%, or about 2%, or about 3%, or about 4%, or about 5%, or about 6%, or about 7%, or about 8%, or about 0%, or about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 75%, or about 80%, or about 85%, or about 90%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% of the reference.
[0031] As will be understood by one skilled in the art, for any and all purposes, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Furthermore, as will be understood by one skilled in the art, a range includes each individual member.
[0032] “Optional” or “optionally” means that the subsequently described circumstance may or may not occur, so that the description includes instances where the circumstance occurs and instances where it does not.74921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0033] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).
[0034] “Substantially” or “essentially” means nearly totally or completely, for instance, 95% or greater of some given quantity. In some embodiments, “substantially” or “essentially” means 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0035] The terms or “acceptable,” “effective,” or “sufficient” when used to describe the selection of any components, ranges, dose forms, etc. disclosed herein intend that said component, range, dose form, etc. is suitable for the disclosed purpose.
[0036] As used herein, the phrase “derived from” means isolated from, purified from, or engineered from, or any combination thereof.
[0037] As used herein, “treating” or “treatment” of a disease in a subject refers to (1) preventing the symptoms or disease from occurring in a subject that is predisposed or does not yet display symptoms of the disease; (2) inhibiting the disease or arresting its development; or (3) ameliorating or causing regression of the disease or the symptoms of the disease. As understood in the art, “treatment” is an approach for obtaining beneficial or desired results, including clinical results. For the purposes of the present technology, beneficial or desired results can include one or more, but are not limited to, alleviation or amelioration of one or more symptoms, diminishment of extent of a condition (including a disease), stabilized (i.e., not worsening) state of a condition (including disease), delay or slowing of condition (including disease), progression, amelioration or palliation of the condition (including disease), states and remission (whether partial or total), whether detectable or undetectable. When the disease is cancer, the following clinical end points are non-limiting examples of treatment: reduction in tumor burden, slowing of tumor growth, longer overall survival, longer time to tumor progression, inhibition of metastasis or a reduction in metastasis of the tumor. In one aspect, treatment excludes prophylaxis.
[0038] As used herein, the term “animal” refers to living multi-cellular vertebrate organisms, a category that includes, for example, mammals and birds. The term “mammal” includes both human and non-human mammals.
[0039] The term “subject,” “host,” “individual,” and “patient” are as used interchangeably herein to refer to animals, typically mammalian animals. Any suitable mammal can be treated by a method described herein. Non-limiting examples of mammals include humans, non-84921-0232-3317.3Atty. Dkt. No.: 106887-9660 human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, and the like), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs) and experimental animals (e.g., mouse, rat, rabbit, guinea pig). In some embodiments, a mammal is a human. A mammal can be any age or at any stage of development (e.g., an adult, teen, child, infant, or a mammal in utero). A mammal can be male or female. In some embodiments, a subject is a human. In some embodiments, a subject has or is diagnosed of having or is suspected of having a cancer.
[0040] The term “contacting” means direct or indirect binding or interaction between two or more. A particular example of direct interaction is binding. A particular example of an indirect interaction is where one entity acts upon an intermediary molecule, which in turn acts upon the second referenced entity. Contacting as used herein includes in solution, in solid phase, in vitro, ex vivo, in a cell and in vivo. Contacting in vivo can be referred to as administering, or administration.
[0041] “Administration” or “delivery” of an oncolytic virus or a composition containing same can be performed in one dose, continuously or intermittently throughout the course of treatment. Methods of determining the most effective means and dosage of administration are known to those of skill in the art and will vary with the composition used for therapy, the purpose of the therapy, the target cell being treated, and the subject being treated. Single or multiple administrations can be carried out with the dose level and pattern being selected by the treating physician or in the case of animals, by the treating veterinarian. Suitable dosage formulations and methods of administering the agents are known in the art. Route of administration can also be determined and method of determining the most effective route of administration are known to those of skill in the art and will vary with the composition used for treatment, the purpose of the treatment, the health condition or disease stage of the subject being treated, and target cell or tissue. Non-limiting examples of route of administration include oral administration, intraperitoneal, infusion, nasal administration, inhalation, injection, and topical application. In some embodiments, the administration is administration to a tissue microenvironment. In some embodiments, administering or a grammatical variation thereof also refers to more than one doses with certain interval. In some embodiments, the interval is 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 10 days, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year or longer. In some embodiments, one dose is repeated for once, twice, three times, four times, five times, six times, seven times, eight times, nine times, ten times or more.94921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0042] The term administration shall include without limitation, administration by oral, parenteral (e.g., intramuscular, intraperitoneal, intravenous, intravascular, intraperitoneal, intracerebroventricular (ICV), intrathecal, intracistemal injection or infusion, intracranial, ocular, intradermally, percutaneously, subcutaneous injection, or implant), intratum orally, by inhalation spray nasal, intratracheal, vaginal, rectal, sublingual, urethral (e.g., urethral suppository) or topical routes of administration (e.g., gel, ointment, cream, aerosol, etc.) and can be formulated, alone or together, in suitable dosage unit formulations containing conventional non-toxic pharmaceutically acceptable carriers, adjuvants, excipients, and vehicles appropriate for each route of administration. The disclosure is not limited by the route of administration, the formulation or dosing schedule.
[0043] An agent of the present disclosure can be administered for therapy by any suitable route of administration. It will also be appreciated that the optimal route will vary with the condition and age of the recipient, and the disease being treated.
[0044] Administration or treatment in “combination” refers to administering two agents such that their pharmacological effects are manifest at the same time. Combination does not require administration at the same time or substantially the same time, although combination can include such administrations.
[0045] The phrase “first line” or “second line” or “third line” refers to the order of treatment received by a patient. First line therapy regimens are treatments given first, whereas second or third line therapy are given after the first line therapy or after the second line therapy, respectively.
[0046] The terms "oligonucleotide" or "polynucleotide" or "portion," or "segment" thereof refer to a stretch of polynucleotide residues which is long enough to use in PCR or various hybridization procedures to identify or amplify identical or related parts of mRNA or DNA molecules. The polynucleotide compositions of this invention include RNA, cDNA, genomic DNA, synthetic forms, and mixed polymers, both sense and antisense strands, and may be chemically or biochemically modified or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those skilled in the art. Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analog, intemucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoamidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendent moieties (e.g., polypeptides),104921-0232-3317.3Atty. Dkt. No.: 106887-9660 intercalators (e.g., acridine, psoralen, etc.), chelators, alkylators, and modified linkages (e.g., alpha anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to a designated sequence via hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages substitute for phosphate linkages in the backbone of the molecule.
[0047] As used herein, the term “purified” does not require absolute purity; rather, it is intended as a relative term. Thus, for example, a purified nucleic acid, peptide, protein, biological complexes, cell, virus or other active compound is one that is isolated in whole or in part from proteins or other contaminants. Generally, substantially purified peptides, proteins, biological complexes, cell, virus or other active compounds for use within the disclosure comprise more than 80% of all macromolecular species present in a preparation prior to admixture or formulation of the peptide, protein, biological complex, cell, virus or other active compound with a pharmaceutical carrier, excipient, buffer, absorption enhancing agent, stabilizer, preservative, adjuvant or other co-ingredient in a complete pharmaceutical formulation for therapeutic administration. More typically, the peptide, protein, biological complex, cell, virus or other active compound is purified to represent greater than 90%, often greater than 95% of all macromolecular species present in a purified preparation prior to admixture with other formulation ingredients. In other cases, the purified preparation may be essentially homogeneous, wherein other macromolecular species are not detectable by conventional techniques.
[0048] In some embodiments, the term “engineered” or “recombinant” refers to having at least one modification not normally found in a naturally occurring protein, polypeptide, polynucleotide, strain, wild-type strain or the parental host strain of the referenced species. In some embodiments, the term “engineered” or “recombinant” refers to being synthetized by human intervention.
[0049] The term “a regulatory sequence” “a regulatory element” “an expression control element” or “promoter” as used herein, intends a polynucleotide that is operatively linked to a polynucleotide to be transcribed and / or replicated, and facilitates the expression and / or replication of the polynucleotide. Non-limiting examples of a regulatory sequence include a promoter, an enhancer, or a polyadenylation sequence.114921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0050] The term “promoter” as used herein refers to any sequence that regulates the expression of a coding sequence, such as a gene. Promoters may be constitutive, inducible, repressible, or tissue-specific, for example. A “promoter” is a control sequence that is a region of a polynucleotide sequence at which initiation and rate of transcription are controlled. It may contain genetic elements at which regulatory proteins and molecules may bind such as RNA polymerase and other transcription factors. Non-limiting examples of promoters include a cytomegalovirus CMV promoter or retroviral long terminal repeat (LTR) promoter. See, for example, Weber et al. Hum Gene Ther. 2007 Sep;18(9):849-60.
[0051] An enhancer is a regulatory element that increases the expression of a target sequence. A "promoter / enhancer" is a polynucleotide that contains sequences capable of providing both promoter and enhancer functions. For example, the long terminal repeats of retroviruses contain both promoter and enhancer functions. The enhancer / promoter may be "endogenous" or "exogenous" or "heterologous." An "endogenous" enhancer / promoter is one which is naturally linked with a given gene in the genome. An "exogenous" or "heterologous" enhancer / promoter is one which is placed in juxtaposition to a gene by means of genetic manipulation (i.e., molecular biological techniques) such that transcription of that gene is directed by the linked enhancer / promoter.
[0052] The term “express” refers to the production of a gene product, such as mRNA, peptides, polypeptides, or proteins. As used herein, “expression” refers to the process by which polynucleotides are transcribed into mRNA or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0053] The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom.
[0054] The term “vector” is used herein to refer to a nucleic acid molecule capable transferring or transporting another nucleic acid molecule. The transferred nucleic acid is generally linked to, e.g., inserted into, the vector nucleic acid molecule. A vector may include sequences that direct autonomous replication in a cell or may include sequences sufficient to124921-0232-3317.3Atty. Dkt. No.: 106887-9660 allow integration into host cell DNA. In some embodiments, the vector is a virus (i.e., a viral vector or oncolytic viral vector).
[0055] As used herein, the term “vector” refers to a nucleic acid construct deigned for transfer between different hosts, including but not limited to a plasmid, a virus, a cosmid, a phage, a BAC, a YAC, etc. In some embodiments, plasmid vectors may be prepared from commercially available vectors. In other embodiments, viral vectors may be produced from baculoviruses, retroviruses, adenoviruses, AAVs, etc. according to techniques known in the art. In one embodiment, the viral vector is a lentiviral vector.
[0056] The term “vector genome” refers to the nucleic acid component of a virus particle, which encodes the genome of the virus particle including any proteins required for replication and / or integration of the genome. In some embodiments, a viral genome acts as a viral vector and may comprise a heterologous gene operably linked to a regulatory sequence, such as a promoter. The promoter may be either native or heterologous to the gene and may be viral or non-viral in origin. The viral genomes described herein may be based on any virus, may be an RNA or DNA genome, and may be either single stranded or double stranded.
[0057] The term "adenovirus" as referred to herein indicates over 52 adenoviral subtypes isolated from humans, and as many from other mammals and birds. See, e.g., Strauss." Adenovirus Infections in Humans," in The Adenoviruses, Ginsberg, ed., Plenum Press, New York, N. Y., pp. 451-496 (1984). The term "adenovirus" can be referred to herein with the abbreviation " Ad" followed by a number indicating serotype, e.g., Ad5. As used herein, the term “oncolytic adenovirus” means an adenovirus that is an oncolytic virus, for example an adenovirus that can replicate or that it is replication-competent in a cancer cell. They are different from a non-replicating adenovirus because a non-replicating adenovirus is unable to replicate in the cancer cell. Non-replicating adenoviruses are used in gene therapy as carriers of genes to target cells, since the goal is to express the therapeutic gene within the intact cell and not the lysis of the cell. In contrast, the therapeutic action of oncolytic adenoviruses is based on the ability to replicate and to lyse the target cell, and thereby eliminate the cancer cell. As used herein, oncolytic adenoviruses include both replication-competent adenoviruses able to lyse cancer cells, even without selectivity, and oncolytic adenoviruses that replicate selectively in cancer cells. Non-limiting examples of an oncolytic adenoviruses include H101, a conditionally replicative adenovirus generated by both E1B and E3 gene deletion, which selectively infects and kills tumor cells through viral oncolysis (see, for example, US Patent Application Publication No. US20040202663A1) and <7 / 1520, a134921-0232-3317.3Atty. Dkt. No.: 106887-9660 variant adenovirus where a fragment of 827 bp in Elb region is deleted so that it does not express Elb-55kDa protein (see, for example, US Patent Application Publication No.US20030104625A1). The variant adenovirus dl1520 does not replicate in normal cells, but selectively replicate in cancer cells where the tumor-suppressor gene p53 is dysfunctional and eventually lyse cancer cells. See, more examples at US Patent No. US 11,000,560 B2.
[0058] Adeno-associated virus (AAV) is a non-pathogenic virus, so it is currently being investigated for many gene therapy applications including oncolytic cancer treatments due to its relatively safe nature. AAV can only package genomes between 2 - 5.2 kb in size when they are flanked with inverted terminal repeat sequences (ITRs), but optimally holds a genome of 4.1 to 4.9 kb in length. General information and reviews of AAV can be found in, for example, Carter, 1989, Handbook of Parvoviruses, Vol.l, pp.169-228, and Berns, 1990, Virology, pp.1743- 1764, Raven Press, (New York). However, it is fully expected that these same principles will be applicable to additional AAV serotypes since it is well known that the various serotypes are quite closely related, both structurally and functionally, even at the genetic level. (See, for example, Biacklowe, 1988, pp.165-174 of Parvoviruses and Human Disease, J. R. Pattison, ed.; and Rose, Comprehensive Virology 3:1-61 (1974)). For example, all AAV serotypes apparently exhibit very similar replication properties mediated by homologous rep genes; and all bear three related capsid proteins such as those expressed in AAV2. The degree of relatedness is further suggested by heteroduplex analysis which reveals extensive cross-hybridization between serotypes along the length of the genome; and the presence of analogous self-annealing segments at the termini that correspond to “inverted terminal repeat sequences” (ITRs). The similar infectivity patterns also suggest that the replication functions in each serotype are under similar regulatory control.
[0059] The term “?\AV vector” refers to a vector comprising one or more polynucleotides of interest (or transgenes) that are flanked by AAV terminal repeat sequences (ITRs). Such AAV vectors can be replicated and packaged into infectious viral particles when present in a host cell that has been transfected with a vector encoding and expressing rep and cap gene products. In one embodiment, the AAV vector is a vector derived from an adeno-associated virus serotype, including without limitation, AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, A AV-9, AAV- 10, AAV-11, AAV- 12, AAV- 13, AAV rhlO, and AAV rh74. AAV vectors, and variants thereof can have one or more of the AAV wild-type genes deleted in whole or part, preferably the rep and / or cap genes, but retain functional flanking ITR sequences. Functional ITR sequences are necessary for the rescue, replication144921-0232-3317.3Atty. Dkt. No.: 106887-9660 and packaging of the AAV virion. Thus, an AAV vector is defined herein to include at least those sequences required in cis for replication and packaging (e.g., functional ITRs) of the virus. The ITRs need not be the wild-type nucleotide sequences, and may be altered, e.g., by the insertion, deletion or substitution of nucleotides, so long as the sequences provide for functional rescue, replication and packaging.
[0060] By “recombinant virus” is meant a virus that has been genetically altered, e.g., by the addition or insertion of a heterologous nucleic acid sequence into the viral particle.
[0061] By “AAV virion” “AAV viral particle” or “AAV vector particle” refers to a viral particle composed of at least one AAV capsid protein and an encapsidated polynucleotide A AV vector. The AAV virion, in one embodiment, comprises a heterologous polynucleotide (i.e. a polynucleotide other than a wild-type AAV genome such as a transgene to be delivered to a mammalian cell). Production of AAV viral particles, in some embodiment, includes production of AAV vector, as such a vector is contained within an AAV vector particle. If the particle comprises a heterologous polynucleotide (i.e. a polynucleotide other than a wild-type AAV genome such as a transgene to be delivered to a mammalian cell), it is typically referred to as an “rAAV vector” or simply “rAAV particle ” Thus, production of AAV vector particle necessarily includes production of rAAV, as such a rAAV genome is contained within an rAAV vector particle.
[0062] For example, a wild-type (wt) AAV virus particle comprising a linear, single-stranded AAV nucleic acid genome associated with an AAV capsid protein coat. The AAV virion can be either a single-stranded (ss) AAV or self-complementary (SC) AAV. In one embodiment, a single-stranded AAV nucleic acid molecules of either complementary sense, e.g., “sense” or “antisense” strands, can be packaged into a AAV virion and both strands are equally infectious.
[0063] The term “heterologous” as it relates to nucleic acid sequences such as coding sequences and control sequences, denotes sequences that are not normally joined together, and / or are not normally associated with a particular cell. Thus, a “heterologous” region of a nucleic acid construct or a vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with the other molecule in nature. For example, a heterologous region of a nucleic acid construct could include a coding sequence flanked by sequences not found in association with the coding sequence in nature. Another example of a heterologous coding sequence is a construct where the coding sequence itself is154921-0232-3317.3Atty. Dkt. No.: 106887-9660 not found in nature (e.g., synthetic sequences having codons different from the native gene). Similarly, a cell transformed with a construct, which is not normally present in the cell would be considered heterologous for purposes of this invention. Allelic variation or naturally occurring mutational events do not give rise to heterologous DNA, as used herein.
[0064] A “coding sequence” or a sequence which “encodes” a particular protein, is a nucleic acid sequence which is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vitro or in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5’ (amino) terminus and a translation stop codon at the 3’ (carboxy) terminus. A coding sequence can include, but is not limited to, cDNA from prokaryotic or eukaryotic mRNA, genomic DNA sequences from prokaryotic or eukaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence will usually be located 3’ to the coding sequence.
[0065] A “nucleic acid” sequence refers to a DNA or RNA sequence. The nucleic acids include base analogues of DNA and RNA including, but not limited to 4-acetyl cytosine, 8-hydroxy-N6- methyladenosine, aziridinylcytosine, pseudoisocytosine, 5- (carboxyhydroxylmethyl)uracil, 5- fluorouracil, 5-bromouracil, 5-carboxymethylaminomethyl-2-thiouracil, 5- carboxymethylaminomethyluracil, dihydrouracil, inosine, N6-isopentenyl adenine, 1 -methyladenine, 1- methylpseudouracil, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2- methylguanine, 3-methylcytosine, 5-methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5- methoxy carbonylmethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5- oxyacetic acid methylester, uracil-5-oxyacetic acid, oxybutoxosine, pseudouracil, queosine, 2- thiocytosine, 5-methyl-2-thiouracil, 2 -thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5-oxyacetic acid, pseudouracil, queosine, 2-thiocytosine, and 2,6-diaminopurine.
[0084] The term DNA “control sequences” refers collectively to promoter sequences, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites (“IRES”), enhancers, and the like, which collectively provide for the replication, transcription and translation of a coding sequence in a recipient cell. Not all of these control sequences need always be present so long as the selected coding sequence is capable of being replicated, transcribed and translated in an appropriate host cell.164921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0066] The terms “equivalent” or “biological equivalent” are used interchangeably when referring to a particular molecule, biological, or cellular material and intend those having minimal homology while still maintaining desired structure or functionality. Non-limiting examples of equivalent polypeptides, include a polypeptide having at least 60%, or alternatively at least 65%, or alternatively at least 70%, or alternatively at least 75%, or alternatively 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95% identity thereto or for polypeptide sequences, or a polypeptide which is encoded by a polynucleotide or its complement that hybridizes under conditions of high stringency to a polynucleotide encoding such polypeptide sequences. Conditions of high stringency are described herein and incorporated herein by reference. Alternatively, an equivalent thereof is a polypeptide encoded by a polynucleotide or a complement thereto, having at least 70%, or alternatively at least 75%, or alternatively 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95% identity, or at least 97% sequence identity to the reference polynucleotide, e.g., the wild-type polynucleotide. In one aspect, the percent identity is determined using the BLAST program run under default parameters. In particular, exemplary programs include BLASTN and BLASTP, using the following default parameters: Genetic code=standard; filter=none; strand=both; cutoff=60; expect=10;Matrix=BLOSUM62; Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following Internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST. Sequence identity and percent identity can be determined by incorporating them into clustalO (available at the web address: ebi.ac.uk / jdispatcher / msa / clustalo, last accessed on November 8, 2024).
[0067] Non-limiting examples of equivalent polynucleotides having at least 60%, or alternatively at least 65%, or alternatively at least 70%, or alternatively at least 75%, or alternatively 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95%, or alternatively at least 97%, identity to a reference polynucleotide. An equivalent also intends a polynucleotide or its complement that hybridizes under conditions of high stringency to a reference polynucleotide.
[0068] A polynucleotide or polynucleotide region (or a polypeptide or polypeptide region) having a certain percentage (for example, 80%, 85%, 90%, or 95%) of “sequence identity” to another sequence means that, when aligned, that percentage of bases (or amino acids) are the same in comparing the two sequences. The alignment and the percent homology or sequence174921-0232-3317.3Atty. Dkt. No.: 106887-9660 identity can be determined using software programs known in the art, for example those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987) Supplement 30, section 7.7.18, Table 7.7.1. In certain embodiments, default parameters are used for alignment. A non-limiting exemplary alignment program is BLAST, using default parameters. In particular, exemplary programs include BLASTN and BLASTP, using the following default parameters: Genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sortby=HIGH SCORE;Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following Internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST. Sequence identity and percent identity can determined by incorporating them into clustalW (available at the web address:genome.jp / tools / clustalw / , last accessed on Jan. 13, 2017).
[0069] “Homology” or “identity” or “similarity” refers to sequence similarity between two peptides or between two nucleic acid molecules. Homology can be determined by comparing a position in each sequence that may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base or amino acid, then the molecules are homologous at that position. A degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. An “unrelated” or “non-homologous” sequence shares less than 40% identity, or alternatively less than 25% identity, with one of the sequences of the present disclosure. In some embodiments, BLAST (accessible at blast.ncbi.nlm.nih.gov / Blast.cgi) or Clustal Omega (accessible at ebi.ac.uk / Tools / msa / clustalo / ) are used in determining the identity. In further embodiments, default setting is applied.
[0070] “Homology” or “identity” or “similarity” can also refer to two nucleic acid molecules that hybridize under stringent conditions.
[0071] “Hybridization” refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding may occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi -stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction, or the enzymatic cleavage of a polynucleotide by a ribozyme.184921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0072] Examples of stringent hybridization conditions include: incubation temperatures of about 25° C to about 37° C; hybridization buffer concentrations of about 6× SSC to about 10× SSC; formamide concentrations of about 0% to about 25%; and wash solutions from about 4×SSC to about 8×SSC. Examples of moderate hybridization conditions include: incubation temperatures of about 40° C to about 50° C; buffer concentrations of about 9× SSC to about 2× SSC; formamide concentrations of about 30% to about 50%; and wash solutions of about 5×SSC to about 2×SSC. Examples of high stringency conditions include: incubation temperatures of about 55° C to about 68° C; buffer concentrations of about 1×SSC to about 0.1×SSC; formamide concentrations of about 55% to about 75%; and wash solutions of about 1×SSC, 0.1×SSC, or deionized water. In general, hybridization incubation times are from 5 minutes to 24 hours, with 1, 2, or more washing steps, and wash incubation times are about 1, 2, or 15 minutes. SSC is 0.15 M NaCl and 15 mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be employed.
[0073] As used herein, the term " HDR", or “homology directed repair,” refers to the process of repairing DNA damage using a homologous nucleic acid (e.g., a sister chromatid or an exogenous nucleic acid). In a normal cell, HDR typically involves a series of steps such as recognition of a double stranded break, stabilization of the break, resection, stabilization of single stranded DNA, formation of a DNA crossover intermediate, resolution of the crossover intermediate, and ligation.
[0074] As used herein, “AZD7648” refers to a DNA-PK inhibitor having a chemical formula:. (Selleckchem Catalog No. S8843)
[0075] As used herein, “M3814” refers to a DNA-PK inhibitor having a chemical formula:. (Selleckchem Catalog No. S8586)194921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0076] As used herein, “VX984” refers to a DNA-PK inhibitor having a chemical formula:OH Cl(Vertex Pharmaceuticals, further described in W02015058031A1).
[0077] As used herein, “KU57788” refers to a DNA-PK inhibitor having a chemical formula:
[0078] As used herein, “BAY8400” refers to a DNA-PK inhibitor having a chemical formula:(AbMole Catalog No. M21187).
[0079] As used herein, “LTURM34” refers to a DNA-PK inhibitor having a chemical formula:O. (Selleckchem Catalog No. S8427)
[0080] As used herein, “53BP1” refers to a p53-binding protein 1 (Genecards ID: TP53BP1, HGNC: 11999NCBI Gene: 7158 Ensembl: ENSG00000067369 OMIM®: 605230204921-0232-3317.3Atty. Dkt. No.: 106887-9660 UniProtKB / Swiss-Prot: Q12888). In some embodiments, 53BP1 has an amino acid sequence as shown in SEQ ID NO: 12, or an equivalent thereof.
[0081] As used herein, “DN1S” is a fragment of 53BP1 that inhibits 53BP1 by dominant negative activity against 53BP1. In some embodiments, DN1S comprises amino acids 1231-1644 of SEQ ID NO: 12, or an equivalent thereof. In some embodiments, DN1S has an amino acid sequence as shown in SEQ ID NO: 13, or an equivalent thereof.
[0082] As used herein, “DN1” is a fragment of 53BP1 that inhibits 53BP1 by dominant negative activity against 53BP1. In some embodiments, DN1 comprises amino acids 1231-1711 of SEQ ID NO: 12, or an equivalent thereof. In some embodiments, DN1S has an amino acid sequence as shown in SEQ ID NO: 14, or an equivalent thereof.CRISPR-Cas Systems
[0083] In general, a CRISPR-Cas or CRISPR system as used in herein and in documents, such as WO 2014 / 093622 (PCT / US2013 / 074667), refers collectively to transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding a Cas gene, a tracr (trans-activating CRISPR) sequence (e.g., tracrRNA or an active partial tracrRNA), a tracr-mate sequence (encompassing a “direct repeat” and a tracrRNA-processed partial direct repeat in the context of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer” in the context of an endogenous CRISPR system), or “RNA(s)” as that term is herein used (e.g., RNA(s) to guide Cas, such as Cas9, e.g. CRISPR RNA and transactivating (tracr) RNA or a single guide RNA (sgRNA) (chimeric RNA)) or other sequences and transcripts from a CRISPR locus. In general, a CRISPR system is characterized by elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of an endogenous CRISPR system). See, e.g., Shmakov et al. (2015) “Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems”, Molecular Cell, DOI: dx.doi.org / 10.1016 / j.molcel.2015.10.008, which is incorporated herein in its entirety.Class 1 Systems
[0084] The methods, systems, and tools provided herein may be designed for use with Class 1 CRISPR proteins. In certain example embodiments, the Class 1 system may be Type I, Type III or Type IV Cas proteins as described in Makarova et al. “Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants” Nature Reviews Microbiology, 18:67-81 (Feb 2020)., incorporated in its entirety herein by reference and214921-0232-3317.3Atty. Dkt. No.: 106887-9660 particularly as described in Figure 1, p. 326. The Class 1 systems typically use a multiprotein effector complex, which can, in some embodiments, include ancillary proteins, such as one or more proteins in a complex referred to as a CRISPR-associated complex for antiviral defense (Cascade), one or more adaptation proteins (e.g. Casl, Cas2, RNA nuclease), and / or one or more accessory proteins (e.g. Cas 4, DNA nuclease), CRISPR associated Rossman fold (CARF) domain containing proteins, and / or RNA transcriptase. Although Class 1 systems have limited sequence similarity, Class 1 system proteins can be identified by their similar architectures, including one or more Repeat Associated Mysterious Protein (RAMP) family subunits, e.g., Cas 5, Cas6, Cas7. RAMP proteins are characterized by having one or more RNA recognition motif domains. Large subunits (for example cas8 or caslO) and small subunits (for example, casl 1) are also typical of Class 1 systems. See, e.g., Figures 1 and 2. Koonin EV, Makarova KS. 2019 Origins and evolution of CRISPR-Cas systems. Phil. Trans. R. Soc. B 374: 20180087, DOI: 10.1098 / rstb.2018.0087. In one aspect, Class 1 systems are characterized by the signature protein Cas3. The Cascade in particular Classi proteins can comprise a dedicated complex of multiple Cas proteins that binds pre-crRNA and recruits an additional Cas protein, for example Cas6 or Cas5, which is the nuclease directly responsible for processing pre-crRNA. In one aspect, the Type I CRISPR protein comprises an effector complex comprises one or more Cas5 subunits and two or more Cas7 subunits. Class 1 subtypes include Type LA, LB, LC, LU, LD, LE, and LF, Type IV-A and IV-B, and Type IILA, IILD, III-C, and IILB. Class 1 systems also include CRISPR-Cas variants, including Type LA, LB, LE, LF and LU variants, which can include variants carried by transposons and plasmids, including versions of subtype LF encoded by a large family of Tn7-like transposon and smaller groups of Tn7-like transposons that encode similarly degraded subtype LB systems. Peters et al., PNAS 114 (35) (2017); DOI:10.1073 / pnas.1709035114; see also, Makarova et al, the CRISPR Journal, v. 1, n5, Figure 5.Class 2 Systems
[0085] The compositions, systems, and methods described in greater detail elsewhere herein can be designed and adapted for use with Class 2 CRISPR-Cas systems. Thus, in some embodiments, the CRISPR-Cas system is a Class 2 CRISPR-Cas system. Class 2 systems are distinguished from Class 1 systems in that they have a single, large, multi-domain effector protein. In certain example embodiments, the Class 2 system can be a Type II, Type V, or Type VI system, which are described in Makarova et al. “Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants” Nature Reviews Microbiology,224921-0232-3317.3Atty. Dkt. No.: 106887-9660 18:67-81 (Feb 2020), incorporated herein by reference. Each type of Class 2 system is further divided into subtypes. See Markova et al. 2020, particularly at Figure. 2. Class 2, Type II systems can be divided into 4 subtypes: II-A, II-B, II-C1, and II-C2. Class 2, Type V systems can be divided into 17 subtypes: V-A, V-Bl, V-B2, V-C, V-D, V-E, V-Fl, V-F1(V-U3), V-F2, V-F3, V-G, V-H, V-I, V-K (V-U5), V-Ul, V-U2, and V-U4. Class 2, Type IV systems can be divided into 5 subtypes: VI-A, VI-B1, VI-B2, VI-C, and VI-D.
[0086] The distinguishing feature of these types is that their effector complexes consist of a single, large, multi-domain protein. Type V systems differ from Type II effectors (e.g., Cas9), which contain two nuclear domains that are each responsible for the cleavage of one strand of the target DNA, with the HNH nuclease inserted inside the Ruv-C like nuclease domain sequence. The Type V systems (e.g., Casl2) only contain a RuvC-like nuclease domain that cleaves both strands. Type VI (Cast 3) are unrelated to the effectors of Type II and V systems and contain two HEPN domains and target RNA. Cast 3 proteins also display collateral activity that is triggered by target recognition. Some Type V systems have also been found to possess this collateral activity with two single-stranded DNA in in vitro contexts.
[0087] In some embodiments, the Class 2 system is a Type II system. In some embodiments, the Type II CRISPR-Cas system is a II-A CRISPR-Cas system. In some embodiments, the Type II CRISPR-Cas system is a II-B CRISPR-Cas system. In some embodiments, the Type II CRISPR-Cas system is a II-C1 CRISPR-Cas system. In some embodiments, the Type II CRISPR-Cas system is a II-C2 CRISPR-Cas system. In some embodiments, the Type II system is a Cas9 system. In some embodiments, the Type II system includes a Cas9.
[0088] In some embodiments, the Class 2 system is a Type V system. In some embodiments, the Type V CRISPR-Cas system is a V-A CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-Bl CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-B2 CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-C CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-D CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-E CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-Fl CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-Fl (V-U3) CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-F2 CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-F3 CRISPR-Cas system. In some embodiments, the234921-0232-3317.3Atty. Dkt. No.: 106887-9660 Type V CRISPR-Cas system is a V-G CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-H CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-I CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-K (V-U5) CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-Ul CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-U2 CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system is a V-U4 CRISPR-Cas system. In some embodiments, the Type V CRISPR-Cas system includes a Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl4, and / or Cas.
[0089] In some embodiments the Class 2 system is a Type VI system. In some embodiments, the Type VI CRISPR-Cas system is a VI-A CRISPR-Cas system. In some embodiments, the Type VI CRISPR-Cas system is a VI-B1 CRISPR-Cas system. In some embodiments, the Type VI CRISPR-Cas system is a VI-B2 CRISPR-Cas system. In some embodiments, the Type VI CRISPR-Cas system is a VI-C CRISPR-Cas system. In some embodiments, the Type VI CRISPR-Cas system is a VI-D CRISPR-Cas system. In some embodiments, the Type VI CRISPR-Cas system includes a Cast 3a (C2c2), Cast 3b (Group 29 / 30), Cast 3c, and / or Casl3d.High-fidelity Cas9As used herein, the phrase “High-fidelity Cas9” refers to engineered variants of the CRISPR-associated protein 9 (Cas9) that exhibit enhanced target specificity and reduced off-target cleavage compared to wild-type Cas9. These variants are designed to minimize unintended genomic modifications while maintaining efficient on-target activity. In some embodiments, a High-fidelity Cas9 comprises amino acid substitutions that alter the protein’s interaction with DNA, thereby improving discrimination between perfectly matched and mismatched target sequences. Such modifications typically reduce non-specific binding and cleavage events. Examples of High-fidelity Cas9 variants include, but are not limited to: SpCas9-HFl: a Streptococcus pyogenes Cas9 variant with multiple mutations (e.g., N497A, R661 A, Q695A, Q926A) that reduce off-target activity (Kleinstiver, Benjamin P., et al. " High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-targeteffects." Nature 529.7587 (2016): 490-495.); eSpCas9(1.1): an enhanced specificity Cas9 variant with mutations that weaken non-specific DNA interactions (Slaymaker, Ian M., et al. " Rationally engineered Cas9 nucleases with improved specificity." Science 351.6268 (2016): 84-88.); HypaCas9: a hyper-accurate Cas9 variant incorporating mutations in the REC3244921-0232-3317.3Atty. Dkt. No.: 106887-9660 domain to improve fidelity (Ikeda, Arisa, et al. " High-fidelity endonuclease variant HypaCas9 facilitates accurate allele-specific gene modification in mousezygotes." Communications biology 2 (2019): 371.); and Sniper-Cas9: a variant optimized for high specificity while retaining robust on-target cleavage efficiency (Lee, Jungjoon K., et al. " Directed evolution of CRISPR-Cas9 to increase its specificity." Nature communications 9.1 (2018): 3048 ). Guide Molecules
[0090] The CRISPR-Cas or Cas-Based system described herein can, in some embodiments, include one or more guide molecules. The terms guide molecule, guide sequence and guide polynucleotide refer to polynucleotides capable of guiding Cas to a target genomic locus and are used interchangeably as in foregoing cited documents such as International Patent Publication No. WO 2014 / 093622 (PCT / US2013 / 074667). In general, a guide sequence is any polynucleotide sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. The guide molecule can be a polynucleotide.
[0091] The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence may be assessed by any suitable assay. For example, the components of a nucleic acid-targeting CRISPR system sufficient to form a nucleic acid-targeting complex, including the guide sequence to be tested, may be provided to a host cell having the corresponding target nucleic acid sequence, such as by transfection with vectors encoding the components of the nucleic acid-targeting complex, followed by an assessment of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, such as by Surveyor assay (Qui et al. 2004. BioTechniques. 36(4)702-707). Similarly, cleavage of a target nucleic acid sequence may be evaluated in a test tube by providing the target nucleic acid sequence, components of a nucleic acid-targeting complex, including the guide sequence to be tested and a control guide sequence different from the test guide sequence, and comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art.
[0092] In some embodiments, the guide molecule is an RNA. The guide molecule(s) (also referred to interchangeably herein as guide polynucleotide and guide sequence) that are included in the CRISPR-Cas or Cas based system can be any polynucleotide sequence having sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid-targeting254921-0232-3317.3Atty. Dkt. No.: 106887-9660 complex to the target nucleic acid sequence. In some embodiments, the degree of complementarity, when optimally aligned using a suitable alignment algorithm, can be about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).
[0093] A guide sequence, and hence a nucleic acid-targeting guide, may be selected to target any target nucleic acid sequence. The target sequence may be DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), micro-RNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (IncRNA), and small cytoplasmatic RNA (scRNA). In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA, and IncRNA. In some embodiments, the target sequence may be a sequence within an mRNA molecule or a pre-mRNA molecule.
[0094] In some embodiments, a nucleic acid-targeting guide is selected to reduce the degree secondary structure within the nucleic acid-targeting guide. In some embodiments, about or less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or fewer of the nucleotides of the nucleic acid-targeting guide participate in self-complementary base pairing when optimally folded. Optimal folding may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimal Gibbs free energy. An example of one such algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is the online webserver RNAfold, developed at Institute for Theoretical Chemistry at the University of Vienna, using the centroid structure prediction algorithm (see e.g., A. R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151-62).264921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0095] In certain embodiments, a guide RNA or crRNA may comprise, consist essentially of, or consist of a direct repeat (DR) sequence and a guide sequence or spacer sequence. In certain embodiments, the guide RNA or crRNA may comprise, consist essentially of, or consist of a direct repeat sequence fused or linked to a guide sequence or spacer sequence. In certain embodiments, the direct repeat sequence may be located upstream (i.e., 5’) from the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3’) from the guide sequence or spacer sequence.
[0096] In certain embodiments, the crRNA comprises a stem loop, preferably a single stem loop. In certain embodiments, the direct repeat sequence forms a stem loop, preferably a single stem loop.
[0097] In certain embodiments, the spacer length of the guide RNA is from 15 to 35 nucleotides (nt). In certain embodiments, the spacer length of the guide RNA is at least 15 nt. In certain embodiments, the spacer length is from 15 to 17 nt, e.g., 15, 16, or 17 nt, from 17 to 20 nt, e.g., 17, 18, 19, or 20 nt, from 20 to 24 nt, e.g., 20, 21, 22, 23, or 24 nt, from 23 to 25 nt, e.g., 23, 24, or 25 nt, from 24 to 27 nt, e.g., 24, 25, 26, or 27 nt, from 27 to 30 nt, e.g., 27, 28, 29, or 30 nt, from 30 to 35 nt, e.g., 30, 31, 32, 33, 34, or 35 nt, or 35 nt or longer.
[0098] The “tracrRNA” sequence or analogous terms includes any polynucleotide sequence that has sufficient complementarity with a crRNA sequence to hybridize. In some embodiments, the degree of complementarity between the tracrRNA sequence and crRNA sequence along the length of the shorter of the two when optimally aligned is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher. In some embodiments, the tracr sequence is about or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and crRNA sequence are contained within a single transcript, such that hybridization between the two produces a transcript having a secondary structure, such as a hairpin.
[0099] In general, degree of complementarity is with reference to the optimal alignment of the sea sequence and tracr sequence, along the length of the shorter of the two sequences. Optimal alignment may be determined by any suitable alignment algorithm and may further account for secondary structures, such as self-complementarity within either the sea sequence or tracr sequence. In some embodiments, the degree of complementarity between the tracr sequence and sea sequence along the length of the shorter of the two when optimally aligned274921-0232-3317.3Atty. Dkt. No.: 106887-9660 is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher.
[0100] In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence can be about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100%; a guide or RNA or sgRNA can be about or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length; or guide or RNA or sgRNA can be less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length; and tracr RNA can be 30 or 50 nucleotides in length. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is greater than 94.5% or 95% or 95.5% or 96% or 96.5% or 97% or 97.5% or 98% or 98.5% or 99% or 99.5% or 99.9%, or 100%. Off target is less than 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% or 94% or 93% or 92% or 91% or 90% or 89% or 88% or 87% or 86% or 85% or 84% or 83% or 82% or 81% or 80% complementarity between the sequence and the guide, with it being advantageous that off target is 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% complementarity between the sequence and the guide.
[0101] In some embodiments according to the invention, the guide RNA (capable of guiding Cas to a target locus) may comprise (1) a guide sequence capable of hybridizing to a genomic target locus in the eukaryotic cell; (2) a tracr sequence; and (3) a tracr mate sequence. All (1) to (3) may reside in a single RNA, i.e., an sgRNA (arranged in a 5’ to 3’ orientation), or the tracr RNA may be a different RNA than the RNA containing the guide and tracr sequence. The tracr hybridizes to the tracr mate sequence and directs the CRISPR / Cas complex to the target sequence. Where the tracr RNA is on a different RNA than the RNA containing the guide and tracr sequence, the length of each RNA may be optimized to be shortened from their respective native lengths, and each may be independently chemically modified to protect from degradation by cellular RNase or otherwise increase stability.
[0102] Many modifications to guide sequences are known in the art and are further contemplated within the context of this invention. Various modifications may be used to increase the specificity of binding to the target sequence and / or increase the activity of the Cas protein and / or reduce off-target effects. Example guide sequence modifications are284921-0232-3317.3Atty. Dkt. No.: 106887-9660 described in International Patent Application No. PCT US2019 / 045582, specifically paragraphs
[0178] -
[0333] , which is incorporated herein by reference.Target Sequences, PAMs, and PFSs
[0103] In the context of formation of a CRISPR complex, “target sequence” refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between a target sequence and a guide sequence promotes the formation of a CRISPR complex. A target sequence may comprise RNA polynucleotides. The term “target RNA” refers to an RNA polynucleotide being or comprising the target sequence. In other words, the target polynucleotide can be a polynucleotide or a part of a polynucleotide to which a part of the guide sequence is designed to have complementarity with and to which the effector function mediated by the complex comprising the CRISPR effector protein and a guide molecule is to be directed. In some embodiments, a target sequence is located in the nucleus or cytoplasm of a cell.
[0104] The guide sequence can specifically bind a target sequence in a target polynucleotide. The target polynucleotide may be DNA. The target polynucleotide may be RNA. The target polynucleotide can have one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc. or more) target sequences. The target polynucleotide can be on a vector. The target polynucleotide can be genomic DNA. The target polynucleotide can be episomal. Other forms of the target polynucleotide are described elsewhere herein.
[0105] The target sequence may be DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), micro-RNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (IncRNA), and small cytoplasmatic RNA (scRNA). In some preferred embodiments, the target sequence (also referred to herein as a target polynucleotide) may be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA, and IncRNA. In some more preferred embodiments, the target sequence may be a sequence within an mRNA molecule or a pre-mRNA molecule.294921-0232-3317.3Atty. Dkt. No.: 106887-9660 PAM and PFS Elements
[0106] PAM elements are sequences that can be recognized and bound by Cas proteins. Cas proteins / effector complexes can then unwind the dsDNA at a position adjacent to the PAM element. It will be appreciated that Cas proteins and systems that include them that target RNA do not require PAM sequences (Marraffini et al. 2010. Nature. 463:568-571). Instead, many rely on PFSs, which are discussed elsewhere herein. In certain embodiments, the target sequence should be associated with a PAM (protospacer adjacent motif) or PFS (protospacer flanking sequence or site), that is, a short sequence recognized by the CRISPR complex. Depending on the nature of the CRISPR-Cas protein, the target sequence should be selected, such that its complementary sequence in the DNA duplex (also referred to herein as the nontarget sequence) is upstream or downstream of the PAM. In the embodiments, the complementary sequence of the target sequence is downstream or 3’ of the PAM or upstream or 5’ of the PAM. The precise sequence and length requirements for the PAM differ depending on the Cas protein used, but PAMs are typically 2-5 base pair sequences adjacent the protospacer (that is, the target sequence). Examples of the natural PAM sequences for different Cas proteins are provided herein below and the skilled person will be able to identify further PAM sequences for use with a given Cas protein.
[0107] The ability to recognize different PAM sequences depends on the Cas polypeptide(s) included in the system. See e.g., Gleditzsch et al. 2019. RNA Biology. 16(4): 504-517. Table 3 (from Gleditzsch et al. 2019) below shows several Cas polypeptides and the PAM sequence they recognize.Example PAM SequencesCas Protein PAM SequenceSpCas9 NGG / NRGSaCas9 NGRRT or NGRRN NmeCas9 NNNNGATTCjCas9 NNNNRYACStCas9 NNAGAAWCas 12a (Cpfl) (including TTTVLbCpfl and AsCpfl)304921-0232-3317.3Atty. Dkt. No.: 106887-9660 Cas 12b (C2cl) TTT, TTA, and TTC Casl2c (C2c3) TACas 12d (CasY) TACasl2e (CasX) 5'-TTCN-3'
[0108] In a specific embodiment, the CRISPR effector protein may recognize a 3’ PAM. In certain embodiments, the CRISPR effector protein may recognize a 3’ PAM which is 5’H, wherein H is A, C or U.
[0109] Further, engineering of the PAM Interacting (PI) domain on the Cas protein may allow programing of PAM specificity, improve target site recognition fidelity, and increase the versatility of the CRISPR-Cas protein, for example as described for Cas9 in Kleinstiver BP et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature. 2015 Jul 23;523(7561):481-5. doi: 10.1038 / naturel4592. As further detailed herein, the skilled person will understand that Casl3 proteins may be modified analogously. Gao et al, “Engineered Cpfl Enzymes with Altered PAM Specificities,” bioRxiv 091611; doi: dx.doi.org / 10.1101 / 091611 (Dec. 4, 2016). Doench et al. created apool of sgRNAs, tiling across all possible target sites of a panel of six endogenous mouse and three endogenous human genes and quantitatively assessed their ability to produce null alleles of their target gene by antibody staining and flow cytometry. The authors showed that optimization of the PAM improved activity and also provided an on-line tool for designing sgRNAs.
[0110] PAM sequences can be identified in a polynucleotide using an appropriate design tool, which are commercially available as well as online. Such freely available tools include, but are not limited to, CRISPRFinder and CRISPRTarget. Mojica et al. 2009. Microbiol. 155(Pt. 3):733-740; Atschul et al. 1990. J. Mol. Biol. 215:403-410; Biswass et al. 2013 RNA Biol. 10:817-827; and Grissa et al. 2007. Nucleic Acid Res. 35: W52-57. Experimental approaches to PAM identification can include, but are not limited to, plasmid depletion assays (Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Esvelt et al. 2013. Nat. Methods. 10:1116-1121; Kleinstiver et al. 2015. Nature. 523:481-485), screened by a high-throughput in vivo model called PAM-SCNAR (Pattanayak et al. 2013. Nat. Biotechnol. 31:839-843 and Leenay et al. 2016. Mol. Cell. 16:253), and negative screening (Zetsche et al. 2015. Cell. 163:759-771).314921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0111] As previously mentioned, CRISPR-Cas systems that target RNA do not typically rely on PAM sequences. Instead such systems typically recognize protospacer flanking sites (PFSs) instead of PAMs Thus, Type VI CRISPR-Cas systems typically recognize protospacer flanking sites (PFSs) instead of PAMs. PFSs represents an analogue to PAMs for RNA targets. Type VI CRISPR-Cas systems employ a Cast 3. Some Cast 3 proteins analyzed to date, such as Cast 3a (C2c2) identified from Leptotrichia shahii (LShCAsl3a) have a specific discrimination against G at the 3 ’end of the target RNA. The presence of a C at the corresponding crRNA repeat site can indicate that nucleotide pairing at this position is rejected. However, some Casl3 proteins (e.g., LwaCAsl3a and PspCasl3b) do not seem to have a PFS preference. See e.g., Gleditzsch et al. 2019. RNA Biology. 16(4): 504-517.
[0112] Some Type VI proteins, such as subtype B, have 5 '-recognition of D (G, T, A) and a 3 '-motif requirement of NAN or NNA. One example is the Cast 3b protein identified in Bergeyella zoohelcum (BzCasl3b). See e.g., Gleditzsch et al. 2019. RNA Biology. 16(4):504-517.
[0113] Overall Type VI CRISPR-Cas systems appear to have less restrictive rules for substrate (e.g., target sequence) recognition than those that target DNA (e.g., Type V and type II).Polynucleotide Fragments as Donor Templates
[0114] In some embodiments, the polynucleotide is delivered into cells as at least one fragment which is a template, e.g., a recombination / donor template. A template may be a component of a vector as described herein, contained in a separate vector, or provided as a separate polynucleotide. In some embodiments, a recombination template is designed to serve as a template in homologous recombination, such as within or near a target sequence nicked or cleaved by a nucleic acid-targeting effector protein as a part of a nucleic acidtargeting complex.
[0115] In an embodiment, the template nucleic acid alters the sequence of the target position. In an embodiment, the template nucleic acid results in the incorporation of a modified, or non-naturally occurring base into the target nucleic acid.
[0116] The template sequence may undergo a breakage mediated or catalyzed recombination with the target sequence. In an embodiment, the template nucleic acid may include sequence that corresponds to a site on the target sequence that is cleaved by a Cas protein mediated cleavage event. In an embodiment, the template nucleic acid may include a sequence that324921-0232-3317.3Atty. Dkt. No.: 106887-9660 corresponds to both, a first site on the target sequence that is cleaved in a first Cas protein mediated event, and a second site on the target sequence that is cleaved in a second Cas protein mediated event.
[0117] In certain embodiments, the template nucleic acid can include a sequence which results in an alteration in the coding sequence of a translated sequence, e.g., one which results in the substitution of one amino acid for another in a protein product, e.g., transforming a mutant allele into a wild type allele, transforming a wild type allele into a mutant allele, and / or introducing a stop codon, insertion of an amino acid residue, deletion of an amino acid residue, or a nonsense mutation. In certain embodiments, the template nucleic acid can include a sequence which results in an alteration in a non-coding sequence, e.g., an alteration in an exon or in a 5' or 3' non-translated or non-transcribed region. Such alterations include an alteration in a control element, e.g., a promoter, enhancer, and an alteration in a cis-acting or trans-acting control element.
[0118] A template nucleic acid having homology with a target position in a target gene may be used to alter the structure of a target sequence. The template sequence may be used to alter an unwanted structure, e.g., an unwanted or mutant nucleotide. The template nucleic acid may include a sequence which, when integrated, results in decreasing the activity of a positive control element; increasing the activity of a positive control element; decreasing the activity of a negative control element; increasing the activity of a negative control element; decreasing the expression of a gene; increasing the expression of a gene; increasing resistance to a disorder or disease; increasing resistance to viral entry; correcting a mutation or altering an unwanted amino acid residue conferring, increasing, abolishing or decreasing a biological property of a gene product, e.g., increasing the enzymatic activity of an enzyme, or increasing the ability of a gene product to interact with another molecule.
[0119] The template nucleic acid may include a sequence which results in a change in sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1, 12 or more nucleotides of the target sequence.
[0120] A template polynucleotide may be of any suitable length, such as about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000, 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides in length. In an embodiment, the template nucleic acid is about 10 to 4000, 1000 to 4000, 2000 to 4000, 2500-4000, or 3000-4000 nucleotides in length.
[0121] In some embodiments, the template polynucleotide is complementary to a portion of a polynucleotide comprising the target sequence. When optimally aligned, a template334921-0232-3317.3Atty. Dkt. No.: 106887-9660 polynucleotide might overlap with one or more nucleotides of a target sequences (e.g., about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more nucleotides). In some embodiments, when a template sequence and a polynucleotide comprising a target sequence are optimally aligned, the nearest nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides from the target sequence.
[0122] The exogenous polynucleotide template comprises a sequence to be integrated (e.g., a fragment of a gene / polynucleotide). The sequence for integration may be a sequence endogenous or exogenous to the cell. Examples of a sequence to be integrated include polynucleotides encoding a protein or a non-coding RNA (e.g., a microRNA). Thus, the sequence for integration may be operably linked to an appropriate control sequence or sequences. Alternatively, the sequence to be integrated may provide a regulatory function.
[0123] An upstream or downstream sequence may comprise from about 20 bp to about 2500 bp, for example, about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or 2500 bp. In some methods, the exemplary upstream or downstream sequence have about 200 bp to about 2000 bp, about 600 bp to about 1000 bp, or more particularly about 700 bp to about 1000.
[0124] An upstream or downstream sequence may comprise from about 20 bp to about 2500 bp, for example, about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or 2500 bp. In some methods, the exemplary upstream or downstream sequence have about 200 bp to about 2000 bp, about 600 bp to about 1000 bp, or more particularly about 700 bp to about 1000
[0125] In certain embodiments, one or both homology arms may be shortened to avoid including certain sequence repeat elements. For example, a 5' homology arm may be shortened to avoid a sequence repeat element. In other embodiments, a 3' homology arm may be shortened to avoid a sequence repeat element. In some embodiments, both the 5' and the 3' homology arms may be shortened to avoid including certain sequence repeat elements.
[0126] In some methods, the exogenous polynucleotide template may further comprise a marker. Such a marker may make it easy to screen for targeted integrations. Examples of suitable markers include restriction sites, fluorescent proteins, or selectable markers. The exogenous polynucleotide template of the disclosure can be constructed using recombinant techniques (see, for example, Sambrook et al., 2001 and Ausubel et al., 1996).344921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0127] In certain embodiments, a template nucleic acid for correcting a mutation may designed for use as a single-stranded oligonucleotide. When using a single-stranded oligonucleotide, 5' and 3' homology arms may range up to about 200 base pairs (bp) in length, e.g., at least 25, 50, 75, 100, 125, 150, 175, or 200 bp in length.
[0128] Suzuki et al. describe in vivo genome editing via CRISPR / Cas9 mediated homologyindependent targeted integration (2016, Nature 540:144-149), incorporated herein in its entirety.TALE Nucleases
[0129] In some embodiments, a TALE nuclease or TALE nuclease system is used to introduce the double strand break. In some embodiments, the methods provided herein use isolated, non-naturally occurring, recombinant or engineered DNA binding proteins that comprise TALE monomers or TALE monomers or half monomers as a part of their organizational structure that enable the targeting of nucleic acid sequences with improved efficiency and expanded specificity.
[0130] Naturally occurring TALEs or “wild type TALEs” are nucleic acid binding proteins secreted by numerous species of proteobacteria. TALE polypeptides contain a nucleic acid binding domain composed of tandem repeats of highly conserved monomer polypeptides that are predominantly 33, 34 or 35 amino acids in length and that differ from each other mainly in amino acid positions 12 and 13. In advantageous embodiments the nucleic acid is DNA. As used herein, the term “polypeptide monomers”, “TALE monomers” or “monomers” will be used to refer to the highly conserved repetitive polypeptide sequences within the TALE nucleic acid binding domain and the term “repeat variable di-residues” or “RVD” will be used to refer to the highly variable amino acids at positions 12 and 13 of the polypeptide monomers. As provided throughout the disclosure, the amino acid residues of the RVD are depicted using the IUPAC single letter code for amino acids. A general representation of a TALE monomer which is comprised within the DNA binding domain is Xl-11-(X12X13)-X14-33 or 34 or 35, where the subscript indicates the amino acid position and X represents any amino acid. X12X13 indicate the RVDs. In some polypeptide monomers, the variable amino acid at position 13 is missing or absent and in such monomers, the RVD consists of a single amino acid. In such cases the RVD may be alternatively represented as X*, where X represents X12 and (*) indicates that X13 is absent. The DNA binding domain comprises several repeats of TALE monomers and this may be represented as (Xl-11-(X12X13)-X14-354921-0232-3317.3Atty. Dkt. No.: 106887-9660 33 or 34 or 35)z, where in an advantageous embodiment, z is at least 5 to 40. In a further advantageous embodiment, z is at least 10 to 26.
[0131] The TALE monomers can have a nucleotide binding affinity that is determined by the identity of the amino acids in its RVD. For example, polypeptide monomers with an RVD of NI can preferentially bind to adenine (A), monomers with an RVD of NG can preferentially bind to thymine (T), monomers with an RVD of HD can preferentially bind to cytosine (C) and monomers with an RVD of NN can preferentially bind to both adenine (A) and guanine (G). In some embodiments, monomers with an RVD of IG can preferentially bind to T. Thus, the number and order of the polypeptide monomer repeats in the nucleic acid binding domain of a TALE determines its nucleic acid target specificity. In some embodiments, monomers with an RVD of NS can recognize all four base pairs and can bind to A, T, G or C. The structure and function of TALEs is further described in, for example, Moscou et al., Science 326:1501 (2009); Boch et al., Science 326:1509-1512 (2009); and Zhang et al., Nature Biotechnology 29:149-153 (2011).
[0132] The polypeptides used in methods of the invention can be isolated, non-naturally occurring, recombinant or engineered nucleic acid-binding proteins that have nucleic acid or DNA binding regions containing polypeptide monomer repeats that are designed to target specific nucleic acid sequences.
[0133] As described herein, polypeptide monomers having an RVD of HN or NH preferentially bind to guanine and thereby allow the generation of TALE polypeptides with high binding specificity for guanine containing target nucleic acid sequences. In some embodiments, polypeptide monomers having RVDs RN, NN, NK, SN, NH, KN, HN, NQ, HH, RG, KH, RH and SS can preferentially bind to guanine. In some embodiments, polypeptide monomers having RVDs RN, NK, NQ, HH, KH, RH, SS and SN can preferentially bind to guanine and can thus allow the generation of TALE polypeptides with high binding specificity for guanine containing target nucleic acid sequences. In some embodiments, polypeptide monomers having RVDs HH, KH, NH, NK, NQ, RH, RN and SS can preferentially bind to guanine and thereby allow the generation of TALE polypeptides with high binding specificity for guanine containing target nucleic acid sequences. In some embodiments, the RVDs that have high binding specificity for guanine are RN, NH RH and KH. Furthermore, polypeptide monomers having an RVD of NV can preferentially bind to adenine and guanine. In some embodiments, monomers having RVDs of H*, HA, KA, N*,364921-0232-3317.3Atty. Dkt. No.: 106887-9660 NA, NC, NS, RA, and S* bind to adenine, guanine, cytosine and thymine with comparable affinity.
[0134] The predetermined N-terminal to C-terminal order of the one or more polypeptide monomers of the nucleic acid or DNA binding domain determines the corresponding predetermined target nucleic acid sequence to which the polypeptides of the invention will bind. As used herein, the monomers and at least one or more half monomers are “specifically ordered to target” the genomic locus or gene of interest. In plant genomes, the natural TALE-binding sites always begin with a thymine (T), which may be specified by a cryptic signal within the non-repetitive N-terminus of the TALE polypeptide; in some cases, this region may be referred to as repeat 0. In animal genomes, TALE binding sites do not necessarily have to begin with a thymine (T) and polypeptides of the invention may target DNA sequences that begin with T, A, G or C. The tandem repeat of TALE monomers always ends with a half-length repeat or a stretch of sequence that may share identity with only the first 20 amino acids of a repetitive full-length TALE monomer and this half repeat may be referred to as a half-monomer. Therefore, it follows that the length of the nucleic acid or DNA being targeted is equal to the number of full monomers plus two.
[0135] As described in Zhang et al., Nature Biotechnology 29: 149-153 (2011), TALE polypeptide binding efficiency may be increased by including amino acid sequences from the “capping regions” that are directly N-terminal or C-terminal of the DNA binding region of naturally occurring TALEs into the engineered TALEs at positions N-terminal or C-terminal of the engineered TALE DNA binding region. Thus, in certain embodiments, the TALE polypeptides described herein further comprise an N-terminal capping region and / or a C-terminal capping region.
[0136] An exemplary amino acid sequence of a N-terminal capping region is:M D P I R S R T P S P A R E L L S G P Q P D G V Q P T A D R G V S P P A G G P L D G L P A R R T M S R T R L P S P P A P S P A F S A D S F S D L L R Q F D P S L F N T S L F D S L P P F G A H H T E A A T G E W D E V Q S G L R A A D A P P P T M R V A V T A A R P P R A K P A P R R R A A Q P S D A S P A A Q V D L R T L G Y S Q Q Q Q E K I K P K V R S T V A Q H H E A L V G H G F T H A H I V A L S Q H P A A L G T V A V K Y Q D M I A A L P E A T H E A I V G V G K Q W S G A R A L E A L L T V A G E L R G P P L Q L D T G Q L L K I A K R G G V T A V E A V H A W R N A L T G A P L N(SEQ ID NO: 10).4921 -0232-3317.3Atty. Dkt. No.: 106887-9660
[0137] An exemplary amino acid sequence of a C-terminal capping region is:
[0138] R P A L E S I V A Q L S R P D P A L A A L T N D H L V A L A C L G G R P A L D A V K K GL P H A P A L I K R T N RR I P E R T S H R V A D H A Q V V R V L GF F Q C H S H P A Q A F D D A M T Q F G M S R H G L L Q L F R R V G V T E L E A R S G T L P P A S Q R W D R I L Q A S G M K R A K P S P T S T Q T P D Q A S L H A F A D S L E R D L D A P S P M H E G D Q T R A S (SEQ ID NO: 11).
[0139] As used herein, the predetermined “N-terminus” to “C terminus” orientation of the N-terminal capping region, the DNA binding domain comprising the repeat TALE monomers and the C-terminal capping region provide structural basis for the organization of different domains in the d-TALEs or polypeptides of the invention.
[0140] The entire N-terminal and / or C-terminal capping regions are not necessary to enhance the binding activity of the DNA binding region. Therefore, in certain embodiments, fragments of the N-terminal and / or C-terminal capping regions are included in the TALE polypeptides described herein.
[0141] In certain embodiments, the TALE polypeptides described herein contain a N-terminal capping region fragment that included at least 10, 20, 30, 40, 50, 54, 60, 70, 80, 87, 90, 94, 100, 102, 110, 117, 120, 130, 140, 147, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260 or 270 amino acids of an N-terminal capping region. In certain embodiments, the N-terminal capping region fragment amino acids are of the C-terminus (the DNA-binding region proximal end) of an N-terminal capping region. As described in Zhang et al., Nature Biotechnology 29: 149-153 (2011), N-terminal capping region fragments that include the C-terminal 240 amino acids enhance binding activity equal to the full-length capping region, while fragments that include the C-terminal 147 amino acids retain greater than 80% of the efficacy of the full length capping region, and fragments that include the C-terminal 117 amino acids retain greater than 50% of the activity of the full-length capping region.
[0142] In some embodiments, the TALE polypeptides described herein contain a C-terminal capping region fragment that included at least 6, 10, 20, 30, 37, 40, 50, 60, 68, 70, 80, 90, 100, 110, 120, 127, 130, 140, 150, 155, 160, 170, 180 amino acids of a C-terminal capping region. In certain embodiments, the C-terminal capping region fragment amino acids are of the N-terminus (the DNA-binding region proximal end) of a C-terminal capping region. As described in Zhang et al., Nature Biotechnology 29: 149-153 (2011), C-terminal capping region fragments that include the C-terminal 68 amino acids enhance binding activity equal384921 -0232-3317.3Atty. Dkt. No.: 106887-9660 to the full-length capping region, while fragments that include the C-terminal 20 amino acids retain greater than 50% of the efficacy of the full-length capping region.
[0143] In certain embodiments, the capping regions of the TALE polypeptides described herein do not need to have identical sequences to the capping region sequences provided herein. Thus, in some embodiments, the capping region of the TALE polypeptides described herein have sequences that are at least 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical or share identity to the capping region amino acid sequences provided herein. Sequence identity is related to sequence homology.Homology comparisons may be conducted by eye, or more usually, with the aid of readily available sequence comparison programs. These commercially available computer programs may calculate percent (%) homology between two or more sequences and may also calculate the sequence identity shared by two or more amino acid or nucleic acid sequences. In some preferred embodiments, the capping region of the TALE polypeptides described herein have sequences that are at least 95% identical or share identity to the capping region amino acid sequences provided herein.
[0144] Sequence homologies can be generated by any of a number of computer programs known in the art, which include but are not limited to BLAST or FASTA. Suitable computer programs for carrying out alignments like the GCG Wisconsin Bestfit package may also be used. Once the software has produced an optimal alignment, it is possible to calculate % homology, preferably % sequence identity. The software typically does this as part of the sequence comparison and generates a numerical result.
[0145] In some embodiments described herein, the TALE polypeptides of the invention include a nucleic acid binding domain linked to the one or more effector domains. The terms “effector domain” or “regulatory and functional domain” refer to a polypeptide sequence that has an activity other than binding to the nucleic acid sequence recognized by the nucleic acid binding domain. By combining a nucleic acid binding domain with one or more effector domains, the polypeptides of the invention may be used to target the one or more functions or activities mediated by the effector domain to a particular target DNA sequence to which the nucleic acid binding domain specifically binds.
[0146] In some embodiments of the TALE polypeptides described herein, the activity mediated by the effector domain is a biological activity. For example, in some embodiments the effector domain is a transcriptional inhibitor (i.e., a repressor domain), such as an mSin394921-0232-3317.3Atty. Dkt. No.: 106887-9660 interaction domain (SID). SID4X domain or a Kriippel-associated box (KRAB) or fragments of the KRAB domain. In some embodiments, the effector domain is an enhancer of transcription (i.e., an activation domain), such as the VP16, VP64 or p65 activation domain. In some embodiments, the nucleic acid binding is linked, for example, with an effector domain that includes but is not limited to a transposase, integrase, recombinase, resol vase, invertase, protease, DNA methyltransferase, DNA demethylase, histone acetylase, histone deacetylase, nuclease, transcriptional repressor, transcriptional activator, transcription factor recruiting, protein nuclear-localization signal or cellular uptake signal.
[0147] In some embodiments, the effector domain is a protein domain which exhibits activities which include but are not limited to transposase activity, integrase activity, recombinase activity, resolvase activity, invertase activity, protease activity, DNA methyltransferase activity, DNA demethylase activity, histone acetylase activity, histone deacetylase activity, nuclease activity, nuclear-localization signaling activity, transcriptional repressor activity, transcriptional activator activity, transcription factor recruiting activity, or cellular uptake signaling activity. Other preferred embodiments of the invention may include any combination of the activities described herein.
[0148] Other preferred tools for genome editing for use in the context of this invention include zinc finger systems and TALE systems. One type of programmable DNA-binding domain is provided by artificial zinc-finger (ZF) technology, which involves arrays of ZF modules to target new DNA-binding sites in the genome. Each finger module in a ZF array targets three DNA bases. A customized array of individual zinc finger domains is assembled into a ZF protein (ZFP).Zinc Finger Nucleases
[0149] Zinc Finger proteins can comprise a functional domain. The first synthetic zinc finger nucleases (ZFNs) were developed by fusing a ZF protein to the catalytic domain of the Type IIS restriction enzyme Fokl. (Kim, Y. G. et al., 1994, Chimeric restriction endonuclease, Proc. Natl. Acad. Sci. U. S. A. 91, 883-887; Kim, Y. G. et al., 1996, Hybrid restriction enzymes: zinc finger fusions to Fok I cleavage domain. Proc. Natl. Acad. Sci. U. S. A. 93, 1156-1160). Increased cleavage specificity can be attained with decreased off target activity by use of paired ZFN heterodimers, each targeting different nucleotide sequences separated by a short spacer. (Doyon, Y. et al., 2011, Enhancing zinc-finger-nuclease activity with improved obligate heterodimeric architectures. Nat. Methods 8, 74-79). ZFPs can also be404921-0232-3317.3Atty. Dkt. No.: 106887-9660 designed as transcription activators and repressors and have been used to target many genes in a wide variety of organisms. Exemplary methods of genome editing using ZFNs can be found for example in U. S. Patent Nos. 6,534,261, 6,607,882, 6,746,838, 6,794,136, 6,824,978, 6,866,997, 6,933,113, 6,979,539, 7,013,219, 7,030,215, 7,220,719, 7,241,573, 7,241,574, 7,585,849, 7,595,376, 6,903,185, and 6,479,626, all of which are specifically incorporated by reference.Meganucleases
[0150] In some embodiments, a meganuclease or system thereof can be used to introduce the double strand break. Meganucleases, which are endodeoxyribonucleases characterized by a large recognition site (double-stranded DNA sequences of 12 to 40 base pairs). Exemplary methods for using meganucleases can be found in US Patent Nos. 8,163,514; 8,133,697; 8,021,867; 8,119,361; 8,119,381; 8,124,369; and 8,129,134; which are specifically incorporated in their entireties herein by reference.
[0151] As used herein, the term “HDR Enhancer peptide” or “HEP” refers to a peptide that promotes or increases the efficiency of homology-directed repair (HDR) in a cell. In some embodiments, a HEP is a 53BP1 inhibiting peptide.
[0152] As used herein, the term “53BP1 inhibiting peptide” refers to a peptide or polypeptide that reduces, blocks, or interferes with the activity or function of p53-binding protein 1 (53BP1) in a cell. In some embodiments, 53BP1 inhibiting peptide comprises peptides derived from the RIFl-binding region of 53BP1 that act as competitive inhibitors. In some embodiments, 53BP1 inhibiting peptide comprises peptides mimicking ubiquitin-dependent recruitment motifs to block 53BP1 localization. In some embodiments, a 53BP1 inhibiting peptide comprises UbV-A. In some embodiments, the 53BP1 inhibiting peptide is selected from any one of SEQ ID NOs: 1-6. In some embodiments, the 53BP1 inhibiting peptide is any one of the ubiquitin variants disclosed in US10808017B2.
[0153] As used herein, the term “enrichment / selection cassette” refers to a nucleic acid construct incorporated into a vector or genome editing system that enables the identification, isolation, or enrichment of cells that have been successfully modified or transduced. This cassette typically comprises one or more genetic elements that confer a selectable trait, such as expression of a marker protein or resistance to a specific agent, allowing for positive selection of desired cells.414921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0154] In some embodiments, an enrichment / selection cassette includes a gene encoding a cell-surface marker, such as truncated CD 19 (tCD19), which facilitates isolation of modified cells through antibody-based sorting or magnetic bead separation. In other embodiments, the cassette may contain a drug resistance gene, such as puromycin or neomycin resistance, enabling survival of transduced cells under selective conditions. The cassette may also include regulatory sequences, such as promoters and polyadenylation signals, to ensure proper expression of the selection marker. In certain embodiments, enrichment / selection cassettes are designed for transient or stable expression and may be flanked by recombination sites or cleavage sequences to allow subsequent removal after selection.
[0155] As described herein, “truncated CD 19” or “tCD19” refers to an enrichment / selection marker that comprises amino acids 1-54 the extracellular domain) of the CD 19 protein (the full length CD19 protein sequence is as shown in UniProt Accession No. P15391).
[0156] As used herein, the term “phosphoglycerate kinase promoter” or “PGK promoter” refers to a regulatory DNA sequence derived from the phosphoglycerate kinase (PGK) gene. In particular embodiments of any of the expression cassettes and gene delivery vectors described herein, the PGK promoter is a human phosphoglycerate kinase 1 (hPGK) promoter that comprises or consists of the following sequence, a functional fragment thereof, or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 15. As used herein, the phrase “Ussing chamber electrophysiological assay” refers to a method for evaluating ion transport and barrier properties across an epithelial tissue or cell monolayer mounted between two fluid-filled compartments. The assay employs electrodes to measure transepithelial potential difference, short-circuit current, and electrical resistance under controlled conditions, thereby enabling quantitative assessment of active and passive ion movement, epithelial integrity, and functional responses to pharmacological or genetic interventions.
[0157] As used herein, an “INDEL” (insertion-deletion) refers to the insertion or deletion of bases in a polynucleotide. Indels can represent a type of genetic variation where a specific nucleotide sequence is either present (insertion) or absent (deletion).
[0158] As used herein, the term “epidermolysis bullosa (EB) gene” refers to a gene whose mutations are associated with the inherited skin disorder epidermolysis bullosa, characterized by skin fragility and blister formation in response to minor mechanical trauma. These genes encode structural or functional proteins critical for dermal-epidermal adhesion and integrity.424921-0232-3317.3Atty. Dkt. No.: 106887-9660 In some embodiments, an EB gene comprises a gene encoding components of the basement membrane zone, keratin intermediate filaments, or associated anchoring structures. Mutations in these genes can lead to different EB subtypes, such as epidermolysis bullosa simplex (EBS), junctional EB (JEB), and dystrophic EB (DEB). Examples of EB genes include, but are not limited to: KRT5(keratin 5, Gene Cards: KRT5, HGNC: 6442, NCBI Gene: 3852, Ensembl: ENSG00000186081, OMIM®: 148040, UniProtKB / Swiss-Prot: P13647) and KRT14 (keratin 14) (Gene Cards: KRT14, HGNC: 6416 NCBI Gene: 3861 Ensembl:ENSG00000186847 OMIM®: 148066 UniProtKB / Swiss-Prot: P02533) - associated with EBS; COL7A1 (type VII collagen, Gene Cards: COL7A1 HGNC: 2214 NCBI Gene: 1294 Ensembl: ENSG00000114270 OMIM®: 120120 UniProtKB / Swiss-Prot: Q02388) -associated with DEB; LAMA3 (Laminin Subunit Alpha 3, Gene Cards: LAMA3, HGNC: 6483 NCBI Gene: 3909 Ensembl: ENSG00000053747 OMIM®: 600805 UniProtKB / Swiss-Prot: QI 6787), LAMB3 (Laminin Subunit Alpha 3, GeneCards: LAMB3, HGNC: 6490, NCBI Gene: 3914, Ensembl: ENSG00000196878, OMIM®: 150310, UniProtKB / Swiss-Prot: QI 3751), and LAMC2 (Laminin Subunit Gamma 2, Gene Cards: LAMC2, HGNC: 6493, NCBI Gene: 3918, Ensembl: ENSG00000058085, OMIM®: 150292, UniProtKB / Swiss-Prot: Q13753) - associated with JEB; ITGA6 (Integrin Subunit Alpha 6, Gene Cards:ITGA6, HGNC: 6142, NCBI Gene: 3655, Ensembl: ENSG00000091409, OMIM®: 147556, UniProtKB / Swiss-Prot: P23229) and ITGB4 (Integrin Subunit Beta 4, Gene Cards, ITGB4 HGNC: 6158, NCBI Gene: 3691, Ensembl: ENSG00000132470, OMIM®: 147557, UniProtKB / Swiss-Prot: P16144) - associated with JEB, and PLEC (plectin) - associated with EBS and other variants.
[0159] As used herein, the term “Duchenne muscular dystrophy gene” refers to a gene whose mutation or functional deficiency results in Duchenne muscular dystrophy (DMD), a severe X-linked neuromuscular disorder characterized by progressive muscle weakness and degeneration. The primary gene implicated in this condition is the dystrophin gene (DMD) (Gene Cards, DMD, HGNC: 2928 NCBI Gene: 1756 Ensembl: ENSG00000198947 OMIM®: 300377 UniProtKB / Swiss-Prot: Pl 1532), which encodes the dystrophin protein, a critical structural component that stabilizes the sarcolemma and maintains muscle fiber integrity during contraction.
[0160] As used herein, the term “CCR5” refers to the C-C chemokine receptor type 5 ( Gene Cards: CCR5, HGNC: 1606, NCBI Gene: 1234, Ensembl: ENSG00000160791, OMIM®:434921-0232-3317.3Atty. Dkt. No.: 106887-9660 601373, UniProtKB / Swiss-Prot: P51681), a G protein-coupled receptor expressed on the surface of various immune cells, including T cells, macrophages, and dendritic cells. CCR5 plays a critical role in mediating immune cell trafficking by binding to chemokines such as CCL3, CCL4, and CCL5, which regulate inflammatory responses and leukocyte migration. In some embodiments, CCR5 is recognized as a co-receptor for HIV-1 entry into host cells, where its interaction with the viral envelope glycoprotein facilitates viral fusion and infection. Genetic variations in CCR5, such as the A32 deletion, confer resistance to HIV infection by preventing functional receptor expression. In certain embodiments, CCR5 may be targeted for therapeutic purposes, including gene editing approaches to disrupt CCR5 expression in hematopoietic stem cells or T cells, thereby reducing susceptibility to HIV or modulating immune responses in autoimmune and inflammatory conditions.
[0161] As used herein, “CCR5 locus” refers to a locus is located on chromosome 3 at position 3p21.31. This locus encompasses the coding sequence for CCR5 as well as upstream and downstream regions that may include promoters, enhancers, introns, and untranslated regions.Modes for Carrying Out the Disclosure
[0162] In one aspect, provided herein is a method for inserting a polynucleotide between about 4.5Kb and about 8Kb and lacking a selection marker into a target genomic position in a cell, the method comprising, or alternatively consisting essentially of, or yet further consisting of (i) introducing a double strand break at the target genomic position, (ii) delivering to the cell a first fragment of the polynucleotide using a first adeno associated virus (AAV) and a second fragment of the polynucleotide using a second AAV, wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide comprise the polynucleotide, and (iii) culturing the cell in the presence of a p53-binding protein 1 (53BP1) inhibitor and a DNA-dependent protein kinase catalytic subunit (DNA-PKcs) inhibitor to complete sequentially inserting the first fragment of the polynucleotide and the second fragment of the polynucleotide into the target genomic position.
[0163] In some embodiments, the 53BP1 inhibitor is selected from any of SEQ ID NO: 1-6 or an equivalent thereof, or a dominant-negative mutant of 53BP1.
[0164] In some embodiments, the double strand break is introduced by a CRISPR-Cas system comprising, or alternatively consisting essentially of, or yet further consisting of a Cas protein and a guide RNA (gRNA) specific for the target genomic position.444921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0165] In some embodiments, the Cas protein is fused to the 53BP1 inhibitor. In some embodiments, the 53BP1 inhibitor comprises, or alternatively consists essentially of, or yet further consists of DN1S or DN1. In some embodiments, DN1S comprises, or alternatively consists essentially of, or yet further consists of an amino acid sequence as shown in SEQ ID NO: 13, or an equivalent thereof. In some embodiments, DN1 comprises, or alternatively consists essentially of, or yet further consists of an amino acid sequence as shown in SEQ ID NO: 14, or an equivalent thereof.
[0166] In some embodiments, the DNA-PKcs inhibitor is selected from AZD7648, M3814, VX984, KU57788, BAY8400, or LTURM34.
[0167] In some embodiments, the 53BP1 inhibitor has a final concentration in culture between about 5pM and about 25pM (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 pM, or any value therebetween).
[0168] In some embodiments, the DNA-PKcs inhibitor has a final concentration in culture between about 0.25μM and about 0.5μM (e.g., about 0.25, 0.3, 0.35, 0.4, 0.45, 0.5 pM, or any value therebetween).
[0169] In some embodiments, the culturing step of (iii) is between about 16 hours and about 48 hours (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38. 39, 40,41, 42, 43, 44, 45, 46, 47, 48 hours, or any value therebetween).
[0170] In some embodiments, the first and the second AAVs are selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV-DJ, or a variant of each thereof, and the first and second AAVs are independently the same or different from each other.
[0171] In some embodiments, the polynucleotide comprises, or alternatively consists essentially of, or yet further consists of a gene encoding a Cystic fibrosis transmembrane conductance regulator (Genecards: CFTR, HGNC: 1884 NCBI Gene: 1080 Ensembl:ENSG00000001626 OMIM®: 602421 UniProtKB / Swiss-Prot: P13569), Dedicator of cytokinesis 8 (Genecards: DOCK8, HGNC: 19191 NCBI Gene: 81704 Ensembl:ENSG00000107099 OMIM®: 611432 UniProtKB / Swiss-Prot: Q8NF50) or ATP binding cassette subfamily A member 3 (Genecards: ABCA3, HGNC: 33 NCBI Gene: 21 Ensembl: ENSG00000167972 OMIM®: 601615 UniProtKB / Swiss-Prot: Q99758).
[0172] In some embodiments, the double strand break is introduced by a CRISPR-Cas system. In some embodiments, the double strand break is introduced by a TALE nuclease.454921-0232-3317.3Atty. Dkt. No.: 106887-9660 In some embodiments, the double strand break is introduced by a Zinc Finger Nuclease (ZNF). In some embodiments, the double strand break is introduced by a meganuclease.
[0173] In some embodiments, the inserting is achieved by homology directed repair (HDR).
[0174] In some embodiments, the first fragment of the polynucleotide further comprises, or alternatively consists essentially of, or yet further consists of a 5’ homology arm at the 5’ end of the first fragment and a 3’ homology arm at the 3’ end of the first fragment, and wherein the second fragment of the polynucleotide further comprises, or alternatively consists essentially of, or yet further consists of a 5’ homology arm at the 5’ end of the second fragment and a 3’ homology arm at the 3’ end of the second fragment.
[0175] In some embodiments, the 5’ homology arm of the first fragment is homologous to a sequence 5’ of the double strand break, and wherein the 3’ homology arm of the first fragment is homologous to the 5’ homology arm of the second fragment, and wherein the 3’ homology arm of the second fragment is homologous to a sequence 3’ of the double strand break.
[0176] The following examples are included to demonstrate some embodiments of the disclosure. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the invention.EXAMPLESExperimental Methods
[0177] Splitting large genes into two halves and inserting sequentially using Cas9 and two AAVs can achieve successful gene insertion. However, it is very inefficient, resulting in less than 10% modified cells. Non-viral ssDNA templates also suffer from low efficiency. To improve yield, a separate cassette expressing a selection marker (e.g., surface proteins such as truncated CD 19 (tCD19)) are used to sort the cells using fluorescence or magnetic activated cell sorting (FIG. 2A). Since the tCD19 cassette is ~1.5 kb long, the strategy still only allows to replace genes that are around 5 kb long. This severely limits the ability to deliver longer genes to a target genome. Integrase-deficient lentiviruses and adenoviruses which can package up to 8 kb and 35 kb, respectively, enable less efficient gene insertion than AAV. Applicant has developed a method that increases the efficiency of long oligonucleotide insertion to a high degree where there is no need for a selection marker. The method464921-0232-3317.3Atty. Dkt. No.: 106887-9660 disclosed herein allows using the entire AAV cargo capacity to deliver larger genes in a sequential manner.
[0178] CRISPR-Cas9 or other gene editing / nuclease systems can be used to generate double stranded breaks (DSB) in the genome. Transgenes can be inserted into this DSB site if they are flanked by sequences homologous to this DSB. Insertion cassettes are delivered using AAV result in the efficient gene insertion. However, AAV vectors can only package genes up to 4000 base pairs (in addition to 800 bp homology arms). CFTR is a gene involved in cystic fibrosis and is about 4500 bp long. Applicant previously developed a sequential gene insertion strategy in which the CFTR gene was packaged into two AAVs and were sequentially insertion. However, the efficiency of insertion was low (less than 5%). Thus, Applicant had to include an enrichment / selection cassette that codes for anenrichment / selection marker. Although this strategy inserts up to 6.5 kb into the genome, about 2 kb is occupied by the enrichment cassette. The present disclosure provides data that the inhibition of DNA-PKcs and 53BP1, proteins involved in DNA repair, improve sequential gene insertion by about 10-fold. By inhibiting DNA-PKcs and 53BP1, Applicant observed >50% cells with CFTR cDNA insertion without the need for any enrichment. This allows to replace the enrichment cassette in the second AAV with genetic cargo. Hence, with the presently disclosed techniques genes up to 8 kb can be inserted into a genome.
[0179] Cystic fibrosis (CF) is a monogenic disease caused by impaired production and / or function of the CF transmembrane conductance regulator (CFTR) protein. There are many other pathogenic mutations throughout the CF gene. An autologous airway stem cell therapy in which the CFTR cDNA is precisely inserted into the CFTR locus may enable the development of a durable cure for almost all CF patients, irrespective of the causal mutation. Here, Applicant used CRISPR-Cas9 and two adeno-associated viruses (AAVs) carrying the two halves of the CFTR cDNA to sequentially insert the full CFTR cDNA in the presence of NHEJ inhibitors in upper airway basal stem cells (UABCs) and human bronchial epithelial cells (HBECs).Inhibition of proteins involved in DNA repair to improve gene insertion.
[0180] 53BP1 is a protein that biases DNA-repair towards NHEJ by limiting homologous recombination (HR). Initial studies showed that ubiquitin variants that prevent 53BP1 binding to DSB sites (i53) may improve single (i.e., not split, as a single oligonucleotide) gene insertion. Novel i53 variants that further improve single gene insertion have also been474921-0232-3317.3Atty. Dkt. No.: 106887-9660 reported. In addition, a dominant negative variant of 53BP1 fused to Cas9 has also been reported to improve single gene insertion. Apart from 53BP1, the inhibition of DNA-PKcs, a protein involved in NHEJ, (e.g., by using AZD-7648) can also improve single gene insertion. AZD-7648 combined with 53BP1 inhibition was shown to further enhance single gene insertion in hematopoietic cells (HSPCs). However, it is unclear if inhibiting 53BP1 and DNA-PKcs simultaneously may enhance CFTR cDNA insertion to levels that would help avoid the use of the tCD19 cassette. Applicant observed that DNA-PKcs inhibition combined with Microhomology-mediated end joining (MMEJ) inhibition improves gene insertion by 2-3 -fold but causes significant toxicity, showing that inhibiting different targets in gene repair pathways together do not always give positive results. See, Stack, Jacob T., et al., Molecular Therapy Nucleic Acids 35.4 (2024), which is incorporated herein in its entirety.Combination of UbVa and AZD-7648 Only Improves the Efficiency of GFP Insertion Modestly
[0181] UbVa (i53) is a peptide that blocks 53BP1 activity. AZD-7648 is a DNA-PKcs inhibitor. Applicant tested the insertion of GFP in the HBB locus in the presence of UbVa and AZD-7648. Flow cytometry was used to measure the fraction of GFP+ cells. Compared to airway stem cells edited in the presence of neither compound (FIG. IB), cells edited in the presence of UbVa (FIG. 1C) show more gene insertion (GFP+ cells). The use of AZD-7648 is more effective in improving gene insertion than UbVa (FIG. ID). The combination of UbV-A and AZD-7648 only improved single gene insertion modestly (FIG. IE).UbV-A improved gene insertion of CFTR cDNA
[0182] Applicant tested the use of UbVa to improve sequential gene insertion of the CFTR cDNA and a truncated CD19 expression cassette packaged into two AAVs (FIG.2A).Normally <10% of edited airway stem cells are tCD19+ indicating the insertion of the sequences present in both AAVs (FIG.2B). 5-25 pM UbVa treatment improved gene insertion by -50-100% (FIG.2C).Combined inhibition of DNA-PKcs and 53BP1 improves CFTR cDNA insertion by 10 fold in upper airway basal stem cells (UABCs) without increasing off-target activity.
[0183] Since the Cystic fibrosis transmembrane conductance regulator (CFTR) cDNA (4500 bp) with the left and right homology arms (LHA and RHA) necessary for HR is too large to be packaged in one AAV, Applicant used a sequential gene insertion approach to insert the CFTR cDNA with a truncated CD 19 (tCD19) enrichment tag under the control of a484921-0232-3317.3Atty. Dkt. No.: 106887-9660 phosphoglycerate kinase promoter (PGK) promoter into upper airway basal stem cells (uABCs) (FIG. 2A). Applicant evaluated gene insertion using high-fidelity Cas9 in the presence of AZD-7648 or i53 alone or in combination. When edited in the presence of both AZD-7648 (0.5 pM) and i53 (5 pM), > 50% of uABCs were positive for tCD19 and did not show reduced proliferation. By contrast, less than 5% of edited cells were tCD19+ in the absence of both NHEJ inhibitors (FIG. 3C). After differentiation on air-liquid interfaces, uABCs edited in the presence of both NHEJ inhibitors showed restored CFTR expression by immunoblotting. To assess restoration of CFTR function, the Ussing chamber electrophysiological assay was performed (see, ussingchamber.com / introduction-to-ussing-chamb er- systems / ). Consistent with immunoblotting results, only UABCs edited in the presence of both NHEJ inhibitors produced epithelial sheets with restored CFTR function comparable to non-CF controls (FIGS. 4A-4B). Applicant also measured off-target editing at a previously reported locus using next generation sequencing. Off-target editing was not increased by NHEJ inhibition (FIG. IF).
[0184] Since >50% edited cells are obtained using this strategy, the tCD19 cassette is no longer necessary. This will allow filling in the space occupied by the tCD19 cassette with therapeutically relevant genetic sequences. Thus, this technology will allow replacing insert genes up to ~6.5-7kb long without the need for selection. Example genes include D0CK8 (involved in D0CK8 deficiency) and ABCA3 (Surfactant protein deficiency).
[0185] Modulation of DNA repair to improve insertion of large genes: Gene insertion is improved by activating HDR or inhibiting non-homologous end joining (NHEJ) and microhomology -mediated end joining (MMEJ), which compete with HDR. Notably, a genome wide screen of DNA repair factors identified repair choices that improved HDR.32 Among NHEJ factors, the inhibition of DNA-dependent protein kinase catalytic subunit (DNA-PKcs) improves gene insertion most effectively (>200% increase). By contrast, inhibition of p53-binding protein 1 (53BP1), a negative regulator of HDR, using peptide ubiquitin variants, dominant negative variants or RAD 18 expression improves HDR by only <50%. The inhibition of other proteins (e.g. DNA ligase IV) has been limited by low specificity and potency of inhibitors. Delivery of HR proteins or their fragments alone or as a fusion also improves HR (e.g. RAD5225, RAD5128, CtIP29). Notably, some strategies are more effective when another repair protein is also modulated. For example, MMEJ inhibition is only effective when combined with DNA-PKcs inhibition. Applicant showed that DNA-PKcs inhibition only modestly improves sequential gene insertion (FIGS 1B-1E). However,494921-0232-3317.3Atty. Dkt. No.: 106887-9660 ongoing research shows that the simultaneous inhibition of DNA-PKcs and 53BP1 synergistically increases sequential gene insertion by >800% (FIGS. 7A-7G). Thus, the modulation of more than one DNA-repair protein to improve gene insertion needs further characterization.
[0186] Genomic integrity after genome editing: Aberrant genetic changes after gene editing pose serious safety concerns such as oncogenicity and this can be exacerbated by DNA-repair modulation. Editing in off-target sites that have sequence similarity to the target site can be reduced by >20-fold while preserving on-target activity using High-fidelity Cas proteins. By contrast, structural variants such as large on-target insertions and deletions (INDELs) and translocations caused by DNA-repair modulation needs to be reduced.However, assays to characterize genome wide structural variation with high resolution are only emerging. Next-generation sequencing of polymerase chain reaction (PCR) amplified target loci and chromosomal aberrations analysis by single targeted linker-mediated PCR sequencing (CAST-seq) do not detect genome wide changes. Targeted amplicon sequencing can also be inaccurate due to PCR bias towards amplifying short amplicons with large deletions. Alternative methods such as whole genome sequencing miss low frequency events. Single cell DNA sequencing of targeted loci and whole genomes are emerging but can sample only a limited number of cells. Applicant’s preliminary data shows that higher HDR mediated by AAV templates reduces structural variants. However, many prior studies characterized structural variants either without HDR templates or using sub-optimal ones. In addition, inhibiting MMEJ and DNA-PKcs simultaneously reduces large INDELs, indicating that modulating multiple DNA repair proteins may maximize gene insertion while minimizing aberrant editing. But a quantitative map of editing outcomes using combinatorial DNA repair modulation and different templates is not available.
[0187] Simultaneous inhibition of DNA-PKcs and 53BP1 synergistically improves sequential gene insertion. Applicant screened NHEJ and MMEJ inhibitors separately and in combination to improve sequential insertion of a ~7 kb sequence consisting of the CFTR cDNA and a tCD19 cassette in upper airway basal stem cells (uABCs) at the endogenous CFTR locus (FIG. 7A). When edited without any inhibitors, only 2 ± 0.5% uABCs were tCD19+. The DNA-PKcs inhibitor, AZD-7648 (0.5 pM) and a 53BP1 inhibiting peptide (HDR Enhancer peptide or HEP) (5 pM), individually improved the fraction of tCD19+by 20-100% (FIG. 7B). When used in combination the fraction of tCD19+cells increased 10-fold to 19 ± 3%. When uABCs from people with CF were edited in the presence of both504921-0232-3317.3Atty. Dkt. No.: 106887-9660 AZD-7648 and HEP, 30-50% of uABCs were tCD19+without any reduction in proliferation or loss of sternness markers (FIGs. 7C-7E). uABCs treated with AAV only (no Cas9) were used to control for expression from unintegrated AAV. After differentiation on air-liquid interfaces, cells edited in the presence of both NHEJ inhibitors showed restored CFTR function similar to non-CF controls, indicating seamless insertion (FIG. 7F). The approach was reproducible in the CCR5 locus using CCR5 specific constructs (FIG. 7G). The improved efficiency allows to replace the tCD19 tag with therapeutically relevant cargo.
[0188] Combinatorial DNA repair modulation may reduce aberrant editing. To assess if DNA repair inhibition increases off-target editing, we measured INDELs at an off-target locus associated with our using our CFTR cDNA insertion system using next-generation sequencing. Inhibiting both DNA-PKcs and 53BP1 using AZD-7648 and HEP reduced off-target editing relative to the use of AZD-7648 alone (FIG. 8). Further experiments are needed to verify that the low off-target activity is not because of the underestimation of small INDELs due to large INDELs.
[0189] Prior studies have investigated the combination of DNA-PKcs and 53BP1 inhibitors to improve insertion of single templates but not sequential insertion using two templates. The synergistic improvement in gene insertion by the simultaneous inhibition of DNA-PKcs and 53BP1 observed for sequential gene insertion is not observed when using a single AAV to insert a single gene cassette (see FIGS. 1A-1F).
[0190] The presently disclosed technology makes it easier to produce an airway stem cell therapy to treat cystic fibrosis but airway stem cell therapies have additional technical hurdles surrounding cell transplantation. The technology can be used for producing an autologous hematopoietic stem cell therapy for treating DOCK8 immunodeficiency syndrome which involves the gene DOCK8 (2099 aa, 6297 bp cDNA). Although such a therapy may be made using a lentivirus to over-express the gene, the approach will not maintain native gene regulation. In addition, lentiviral therapies pose risk of oncogenesis. This technology will allow to precisely replace DOCK8 in exon 1 of the native gene locus and thereby maintaining native gene regulation.Equivalents
[0191] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs.514921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0192] The present technology illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms “comprising,” “including,” “containing,” etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the present technology claimed.
[0193] Thus, it should be understood that the materials, methods, and examples provided here are representative of preferred aspects, are exemplary, and are not intended as limitations on the scope of the present technology.
[0194] It should be understood that although the present invention has been specifically disclosed by certain aspects, embodiments, and optional features, modification, improvement and variation of such aspects, embodiments, and optional features can be resorted to by those skilled in the art, and that such modifications, improvements and variations are considered to be within the scope of this disclosure.
[0195] The present technology has been described broadly and generically herein. Each of the narrower species and sub-generic groupings falling within the generic disclosure also form part of the present technology. This includes the generic description of the present technology with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.
[0196] In addition, where features or aspects of the present technology are described in terms of Markush groups, those skilled in the art will recognize that the present technology is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0197] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control.Sequences:
[0198] SEQ ID NO: 1, Protein, Artificial sequence524921-0232-3317.3Atty. Dkt. No.: 106887-9660 MLIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLAFAGKSLEDGRTLSDY NILKDSKLHPLLRLR
[0199] SEQ ID NO: 2, Protein, Artificial sequenceMQIYVKTFARKPITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAEMRLEDGRTLSDY NIKNDSTLFLVLKNSVT
[0200] SEQ ID NO: 3, Protein, Artificial sequenceMLIFVTTDMGMTISLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFGDKDLEDGRTLSD YNIQKES SLNLVLKLRGG
[0201] SEQ ID NO: 4, Protein, Artificial sequenceMQIFVTTDMWMRISLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFGDKDLEDGRTLSD YNIQKES SLNLVLNLRGG
[0202] SEQ ID NO: 5, Protein, Artificial sequenceMLIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKSLEDGRTLSDY NILKDSKLHPLLRLRGG
[0203] SEQ ID NO: 6, Protein, Artificial sequenceMRIIVKTFMRKPITLEVEPSDTIENVKAKIQDKEGIPPDQQRLYFAASQLEDGRTLSDY NIQKESTLLLVVRLLRV
[0204] SEQ ID NO: 7, DNA, Human WT CFTR sequenceATCATCTTTGGTGTTTCCTATGATGAATATAGATACAGA
[0205] SEQ ID NO: 8, DNA, Human AF508 sequenceATCATTGGTGTTTCCTATGATGAATATAGATACAGA
[0206] SEQ ID NO: 9, DNA, Artificial sequence (Correction template)ATCATCTTCGGCGTGTCTTACGACGAGTACAGATACAGA
[0207] SEQ ID NO: 10, Protein, Artificial Sequence (N-terminal capping region)534921-0232-3317.3Atty. Dkt. No.: 106887-9660 M D P I R S R T P S P A R E L L S GP Q P D G V Q P T A D R G V S P P A G GP L D G L P A R R T M S R T R L P S P P A P S P A F S A D S F S D L L R Q F D P S L F N T S L F D S L P P F G A H H T E A A T GE W D E V Q S G L R A A D A P P P T M R V A V T A A R P P R A K P A P R RR A A Q P S D A S P A A Q V D L R T L G Y S Q Q Q Q E K I K P K VR S T V A Q H HE A L V GH GF T H AH I V A L S Q H P A A L G T V A V K Y Q D M I A A L P E A T H E A I V G V GK Q W S G A R A L E A L L T V A G E L R G P P L Q L D T G Q L L K I A K R G G V T A V E A V H A W RN A L T G A P L N
[0208] SEQ ID NO: 11, Protein, Artificial Sequence (C-terminal capping region)R P A L E S I V A Q L S R P D P A L A A L T N D H L V A L A C L G GR P A L D A V K K G L P H A P A L I K R T N R R I P E R T S H R V A D H A Q V V R V L GF F Q C H S H P A Q A F D D A M T Q F GM S R H GL L Q L F RR V G V T E L E A R S G T L P P A S Q R W D R I L Q A S GM K R A K P S P T S T Q T P D Q A S L H A F A D S L E R D L D A P S P M H E G D Q T R A S
[0209] SEQ ID NO: 12, Protein, Homo sapiens 53BP1MDPTGSQLDSDFSQQDTPCLIIEDSQPESQVLEDDSGSHFSMLSRHLPNLQTHKENPV LDVVSNPEQTAGEERGDGNSGFNEHLKENKVADPVDSSNLDTCGSISQVIEQLPQPN RTSSVLGMSVESAPAVEEEKGEELEQKEKEKEEDTSGNTTHSLGAEDTASSQLGFGV LELSQSQDVEENTVPYEVDKEQLQSVTTNSGYTRLSDVDANTAIKHEEQSNEDIPIAE QSSKDIPVTAQPSKDVHVVKEQNPPPARSEDMPFSPKASVAAMEAKEQLSAQELMES GLQIQKSPEPEVLSTQEDLFDQSNKTVSSDGCSTPSREEGGCSLASTPATTLHLLQLSG QRSLVQDSLSTNSSDLVAPSPDAFRSTPFIVPSSPTEQEGRQDKPMDTSVLSEEGGEPF QKKLQSGEPVELENPPLLPESTVSPQASTPISQSTPVFPPGSLPIPSQPQFSHDIFIPSPSL EEQSNDGKKDGDMHSSSLTVECSKTSEIEPKNSPEDLGLSLTGDSCKLMLSTSEYSQS PKMESLSSHRIDEDGENTQIEDTEPMSPVLNSKFVPAENDSILMNPAQDGEVQLSQND DKTKGDDTDTRDDISILATGCKGREETVAEDVCIDLTCDSGSQAVPSPATRSEALSSV LDQEEAMEIKEHHPEEGSSGSEVEEIPETPCESQGEELKEENMESVPLHLSLTETQSQG LCLQKEMPKKECSEAMEVETSVISIDSPQKLAILDQELEHKEQEAWEEATSEDSSVVI VDVKEPSPRVDVSCEPLEGVEKCSDSQSWEDIAPEIEPCAENRLDTKEEKSVEYEGDL KSGTAETEPVEQDSSQPSLPLVRADDPLRLDQELQQPQTQEKTSNSLTEDSKMANAK QLSSDAEAQKLGKPSAHASQSFCESSSETPFHFTLPKEGDIIPPLTGATPPLIGHLKLEP KRHSTPIGISNYPESTIATSDVMSESMVETHDPILGSGKGDSGAAPDVDDKLCLRMKL544921 -0232-3317.3Atty. Dkt. No.: 106887-9660 VSPETEASEESLQFNLEKPATGERKNGSTAVAESVASPQKTMSVLSCICEARQENEAR SEDPPTTPIRGNLLHFPSSQGEEEKEKLEGDHTIRQSQQPMKPISPVKDPVSPASQKMV IQGPSSPQGEAMVTDVLEDQKEGRSTNKENPSKALIERPSQNNIGIQTMECSLRVPET VSAATQTIKNVCEQGTSTVDQNFGKQDATVQTERGSGEKPVSAPGDDTESLHSQGE EEFDMPQPPHGHVLHRHMRTIREVRTLVTRVITDVYYVDGTEVERKVTEETEEPIVE CQECETEVSPSQTGGSSGDLGDISSFSSKASSLHRTSSGTSLSAMHSSGSSGKGAGPLR GKTSGTEPADFALPSSRGGPGKLSPRKGVSQTGTPVCEEDGDAGLGIRQGGKAPVTP RGRGRRGRPPSRTTGTRETAVPGPLGIEDISPNLSPDDKSFSRVVPRVPDSTRRTDVG AGALRRSDSPEIPFQAAAGPSDGLDASSPGNSFVGLRVVAKWSSNGYFYSGKITRDV GAGKYKLLFDDGYECDVLGKDILLCDPIPLDTEVTALSEDEYFSAGVVKGHRKESGE LYYSIEKEGQRKWYKRMAVILSLEQGNRLREQYGLGPYEAVTPLTKAADISLDNLVE GKRKRRSNVS SPATPT AS S S S STTPTRKITESPRASMGVL SGKRKLIT SEEERSP AKRG RKSATVKPGAVGAGEFVSPCESGDNTGEPSALEEQRGPLPLNKTLFLGYAFLLTMAT TSDKLASRSKLPDGPTGSSEEEEEFLEIPPFNKQYTESQLRAGAGYILEDFNEAQCNTA YQCLLIADQHCRTRKYFLCLASGIPCVSHVWVHDSCHANQLQNYRNYLLPAGYSLE EQRILDWQPRENPFQNLKVLLVSDQQQNFLELWSEILMTGGAASVKQHHSSAHNKD IALGVFDVVVTDPSCPASVLKCAEALQLPVVSQEWVIQCLIVGERIGFKQHPKYKHD YVSH
[0210] SEQ ID NO: 13, Protein, Homo sapiens 53BP1 fragment DN1SPHGHVLHRHMRTIREVRTLVTRVITDVYYVDGTEVERKVTEETEEPIVECQECETEVS PSQTGGSSGDLGDISSFSSKASSLHRTSSGTSLSAMHSSGSSGKGAGPLRGKTSGTEPA DFALPSSRGGPGKLSPRKGVSQTGTPVCEEDGDAGLGIRQGGKAPVTPRGRGRRGRP PSRTTGTRETAVPGPLGIEDISPNLSPDDKSFSRVVPRVPDSTRRTDVGAGALRRSDSP EIPFQAAAGPSDGLDASSPGNSFVGLRVVAKWSSNGYFYSGKITRDVGAGKYKLLFD DGYECDVLGKDILLCDPIPLDTEVTALSEDEYFSAGVVKGHRKESGELYYSIEKEGQR KWYKRMAVILSLEQGNRLREQYGLGPYEAVTPLTKAADISLDNLVEGKRKRRSNVS SPATPTASSS
[0211] SEQ ID NO: 14, Protein, Homo sapiens 53BP1 fragment DN1PHGHVLHRHMRTIREVRTLVTRVITDVYYVDGTEVERKVTEETEEPIVECQECETEVS PSQTGGSSGDLGDISSFSSKASSLHRTSSGTSLSAMHSSGSSGKGAGPLRGKTSGTEPA DFALPSSRGGPGKLSPRKGVSQTGTPVCEEDGDAGLGIRQGGKAPVTPRGRGRRGRP554921-0232-3317.3Atty. Dkt. No.: 106887-9660 PSRTTGTRETAVPGPLGIEDISPNLSPDDKSFSRVVPRVPDSTRRTDVGAGALRRSDSP EIPFQAAAGPSDGLDASSPGNSFVGLRVVAKWSSNGYFYSGKITRDVGAGKYKLLFD DGYECDVLGKDILLCDPIPLDTEVTALSEDEYFSAGVVKGHRKESGELYYSIEKEGQR KWYKRMAVILSLEQGNRLREQYGLGPYEAVTPLTKAADISLDNLVEGKRKRRSNVS SPATPTASSSSSTTPTRKITESPRASMGVLSGKRKLITSEEERSPAKRGRKSATVKPGA VGAGEFVSPCESGDNTGE
[0212] SEQ ID NO: 15, DNA, Homo sapiens, hPGK promoter GGGGTTGGGGTTGCGCCTTTTCCAAGGCAGCCCTGGGTTTGCGCAGGGACGCGG CTGCTCTGGGCGTGGTTCCGGGAAACGCAGCGGCGCCGACCCTGGGTCTCGCAC ATTCTTCACGTCCGTTCGCAGCGTCACCCGGATCTTCGCCGCTACCCTTGTGGGCC CCCCGGCGACGCTTCCTGCTCCGCCCCTAAGTCGGGAAGGTTCCTTGCGGTTCGC GGCGTGCCGGACGTGACAAACGGAAGCCGCACGTCTCACTAGTACCCTCGCAGA CGGACAGCGCCAGGGAGCAATGGCAGCGCGCCGACCGCGATGGGCTGTGGCCA ATAGCGGCTGCTCAGCAGGGCGCGCCGAGAGCAGCGGCCGGGAAGGGGCGGTG CGGGAGGCGGGGTGTGGGGCGGTAGTGTGGGCCCTGTTCCTGCCCGCGCGGTGT TCCGCATTCTGCAAGCCTCCGGAGCGCACGTCGGCAGTCGGCTCCCTCGTTGACC GAATCACCGACCTCTCTCCCCAG.Additional Embodiments
[0213] Embodiment 1: A method for inserting a polynucleotide between about 4.5Kb and about 8Kb and lacking a selection marker into a target genomic position in a cell, the method comprising: (i) introducing a double strand break at the target genomic position, (ii) delivering to the cell a first fragment of the polynucleotide and a second fragment of the polynucleotide, wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide comprise the polynucleotide, and (iii) culturing the cell in the presence of a p53-binding protein 1 (53BP1) inhibitor and a DNA-dependent protein kinase catalytic subunit (DNA-PKcs) inhibitor to complete sequentially inserting the first fragment of the polynucleotide and the second fragment of the polynucleotide into the target genomic position.
[0214] Embodiment 2: The method of embodiment 1, wherein the 53BP1 inhibitor is selected from any of SEQ ID NO: 1-6 or an equivalent thereof, or a dominant-negative mutant of 53BP1.564921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0215] Embodiment 3: The method of embodiment 1 or embodiment 2, wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide do not comprise an enrichment cassette.
[0216] Embodiment 4: The method of any one of embodiments 1-3, wherein the double strand break is introduced by (a) a CRISPR-Cas system comprising a Cas protein and a guide RNA (gRNA) specific for the target genomic position, (b) a Transcription Activator-Like Effector (TALE) nuclease specific for the target genomic position, or (c) a Zinc Finger (ZNF) nuclease specific for the target genomic position.
[0217] Embodiment 5: The method of embodiment 4, wherein the Cas protein, the TALE nuclease or the ZNF nuclease is fused to the 53BP1 inhibitor.
[0218] Embodiment 6: The method of embodiment 5, wherein the 53BP1 inhibitor comprises DN1S as shown in SEQ ID NO: 13, or DN1 as shown in SEQ ID NO: 14, or an equivalent thereof.
[0219] Embodiment 7: The method of any one of embodiments 1-6, wherein the DNA-PKcs inhibitor is selected from AZD7648, M3814, VX984, KU57788, BAY8400, or LTURM34.
[0220] Embodiment 8: The method of any one of embodiments 1-4 and 7, wherein the 53BP1 inhibitor has a final concentration in culture between about 5pM and about 25pM.
[0221] Embodiment 9: The method of any one of embodiments 1-8, wherein the DNA-PKcs inhibitor has a final concentration in culture between about 0.25μM and about 0.5μM.
[0222] Embodiment 10: The method of any one of embodiments 1-9, wherein the culturing step of (iii) is between about 16 hours and about 48 hours.
[0223] Embodiment 11: The method of any one of embodiments 1-10, wherein the first fragment of the polynucleotide is delivered using a first adeno associated virus (AAV), and the second fragment of the polynucleotide is delivered using a second AAV.
[0224] Embodiment 12: The method of embodiment 11, wherein the first and the second AAVs are selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV-DJ, or a variant of each thereof, and the first and second AAVs are independently the same or different from each other.574921-0232-3317.3Atty. Dkt. No.: 106887-9660
[0225] Embodiment 13: The method of any one of embodiments 1-12, wherein the polynucleotide comprises a gene encoding Cystic fibrosis transmembrane conductance regulator (CFTR), Dedicator of cytokinesis 8 (DOCK8), ATP binding cassette subfamily A member 3 (ABCA3), a Duchenne muscular dystrophy gene, or an epidermolysis bullosa (EB) gene.
[0226] Embodiment 14: The method of any one of embodiments 1-13, wherein the double strand break is introduced by a CRISPR-Cas system.
[0227] Embodiment 15: The method of any one of embodiments 1-14, wherein the inserting is achieved by homology directed repair (HDR).
[0228] Embodiment 16: The method of any one of embodiments 1-15, wherein the first fragment of the polynucleotide further comprises a 5’ homology arm at the 5’ end of the first fragment and a 3’ homology arm at the 3’ end of the first fragment, and wherein the second fragment of the polynucleotide further comprises a 5’ homology arm at the 5’ end of the second fragment and a 3’ homology arm at the 3’ end of the second fragment.
[0229] Embodiment 17: The method of embodiment 16, wherein the 5’ homology arm of the first fragment is homologous to a sequence 5’ of the double strand break, and wherein the 3’ homology arm of the first fragment is homologous to the 5’ homology arm of the second fragment, and wherein the 3’ homology arm of the second fragment is homologous to a sequence 3’ of the double strand break.
[0230] Embodiment 18: The method of any one of embodiments 1-17, wherein off-target INDEL formation is reduced as compared to a treatment with the 53BP1 inhibitor or the DNA-PKcs inhibitor alone.
[0231] Embodiment 19: The method of any one of embodiments 1-18, wherein the target locus comprises a C-C chemokine receptor type 5 (CCR5) locus.
[0232] Other aspects are set forth within the following claims.584921-0232-3317.3
Claims
Atty. Dkt. No.: 106887-9660 WHAT IS CLAIMED IS:
1. A method for inserting a polynucleotide between about 4.5Kb and about 8Kb and lacking a selection marker into a target genomic position in a cell, the method comprising: (i) introducing a double strand break at the target genomic position,(ii) delivering to the cell a first fragment of the polynucleotide and a second fragment of the polynucleotide,wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide comprise the polynucleotide, and(iii) culturing the cell in the presence of a p53-binding protein 1 (53BP1) inhibitor and a DNA-dependent protein kinase catalytic subunit (DNA-PKcs) inhibitor to complete sequentially inserting the first fragment of the polynucleotide and the second fragment of the polynucleotide into the target genomic position.
2. The method of claim 1, wherein the 53BP1 inhibitor is selected from any of SEQ ID NO: 1-6 or an equivalent thereof, or a dominant-negative mutant of 53BP1.
3. The method of claim 1 or claim 2, wherein the first fragment of the polynucleotide and the second fragment of the polynucleotide do not comprise an enrichment cassette.
4. The method of any one of claims 1-3, wherein the double strand break is introduced by (a) a CRISPR-Cas system comprising a Cas protein and a guide RNA (gRNA) specific for the target genomic position, (b) a Transcription Activator-Like Effector (TALE) nuclease specific for the target genomic position, or (c) a Zinc Finger (ZNF) nuclease specific for the target genomic position.
5. The method of claim 4, wherein the Cas protein, the TALE nuclease or the ZNF nuclease is fused to the 53BP1 inhibitor.
6. The method of claim 5, wherein the 53BP1 inhibitor comprises DN1S as shown in SEQ ID NO: 13, or DN1 as shown in SEQ ID NO: 14, or an equivalent thereof.
7. The method of any one of claims 1-6, wherein the DNA-PKcs inhibitor is selected from AZD7648, M3814, VX984, KU57788, BAY8400, or LTURM34.
8. The method of any one of claims 1-4 and 7, wherein the 53BP1 inhibitor has a final concentration in culture between about 5μM and about 25μM.594921-0232-3317.3Atty. Dkt. No.: 106887-9660 9. The method of any one of claims 1-8, wherein the DNA-PKcs inhibitor has a final concentration in culture between about 0.25μM and about 0.5μM.
10. The method of any one of claims 1-9, wherein the culturing step of (iii) is between about 16 hours and about 48 hours.
11. The method of any one of claims 1-10, wherein the first fragment of the polynucleotide is delivered using a first adeno associated virus (AAV), and the second fragment of the polynucleotide is delivered using a second AAV.
12. The method of claim 11, wherein the first and the second AAVs are selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV-DJ, or a variant of each thereof, and the first and second AAVs are independently the same or different from each other.
13. The method of any one of claims 1-12, wherein the polynucleotide comprises a gene encoding Cystic fibrosis transmembrane conductance regulator (CFTR), Dedicator of cytokinesis 8 (DOCK8), ATP binding cassette subfamily A member 3 (ABCA3), a Duchenne muscular dystrophy gene, or an epidermolysis bullosa (EB) gene.
14. The method of any one of claims 1-13, wherein the double strand break is introduced by a CRISPR-Cas system.
15. The method of any one of claims 1-14, wherein the inserting is achieved by homology directed repair (HDR).
16. The method of any one of claims 1-15, wherein the first fragment of the polynucleotide further comprises a 5’ homology arm at the 5’ end of the first fragment and a 3’ homology arm at the 3’ end of the first fragment, andwherein the second fragment of the polynucleotide further comprises a 5’ homology arm at the 5’ end of the second fragment and a 3’ homology arm at the 3’ end of the second fragment.
17. The method of claim 16, wherein the 5’ homology arm of the first fragment is homologous to a sequence 5’ of the double strand break, and wherein the 3’ homology arm of the first fragment is homologous to the 5’ homology arm of the second fragment, and wherein the 3’ homology arm of the second fragment is homologous to a sequence 3’ of the double strand break.604921-0232-3317.3Atty. Dkt. No.: 106887-9660 18. The method of any one of claims 1-17, wherein off-target INDEL formation is reduced as compared to a treatment with the 53BP1 inhibitor or the DNA-PKcs inhibitor alone.
19. The method of any one of claims 1-18, wherein the target locus comprises a C-C chemokine receptor type 5 (CCR5) locus.614921-0232-3317.3