Compositions and methods for transgene expression from the albumin locus
Targeting the albumin gene intron 1 with guide RNAs and RNA-guided DNA-binding agents enhances gene insertion and expression, addressing inefficiencies in current gene therapy methods and achieving high expression levels in host cells.
Patent Information
- Application Number
- JP2024063210
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-29
- Filing Date
- 2024-04-10
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2039-10-18
AI Technical Summary
Current gene therapy approaches face challenges in efficiently inserting and expressing exogenous genes at safe harbor loci, such as the albumin locus, due to variations in insertion efficiency and expression levels, particularly in non-dividing cell types.
The use of guide RNAs targeting the albumin gene intron 1, combined with RNA-guided DNA-binding agents like Cas9, to facilitate precise insertion and expression of heterologous genes, utilizing bidirectional nucleic acid constructs and vectors like AAV to enhance gene delivery and expression.
This method achieves robust and efficient insertion and expression of heterologous polypeptides in host cells, including non-dividing cells, with increased expression levels up to 10% to 99% of the coding sequence, demonstrating improved therapeutic potential.
Smart Images

Figure 0007820434000053 
Figure 0007820434000054 
Figure 0007820434000055
Abstract
Description
[Technical Field]
[0001] This application claims the benefit of priority from U.S. Provisional Application No. 62 / 747,402, filed October 18, 2018, and U.S. Provisional Application No. 62 / 840,346, filed April 29, 2019. The specifications of each of the foregoing applications are incorporated herein by reference in their entirety. [Background technology]
[0002] Genome editing in gene therapy approaches arises from the idea that exogenous introduction of defective or otherwise impaired genetic material can correct genetic diseases. Gene therapy has long been recognized for its enormous potential in how practitioners approach and treat human diseases. Instead of relying on drugs or surgery, patients with underlying genetic causes can be treated by directly targeting the underlying cause. Furthermore, by targeting the underlying genetic cause, gene therapy may offer the possibility of effectively curing patients. However, the clinical application of gene therapy approaches still requires improvement in several aspects.
[0003] Provided herein are compositions and methods useful for inserting and expressing heterologous (exogenous) genes within genomic loci, e.g., safe harbor sites, of host cells. Several safe harbor loci have been described, including CCR5, HPRT, AAVS1, Rosa, and albumin. As described herein, targeting and inserting an exogenous gene at the albumin locus (e.g., intron 1) allows for the use of the endogenous promoter of albumin to drive robust expression of the exogenous gene. The present disclosure is based, in part, on the identification of guide RNAs that specifically target the albumin gene, e.g., a site within intron 1 of the albumin gene, and provide efficient insertion and / or expression of the exogenous gene. The following embodiments are provided:
[0004] In one aspect, the disclosure provides a method for inserting a nucleic acid encoding a heterologous polypeptide into the albumin locus of a host cell or population of cells, comprising: i) inserting a nucleic acid encoding a heterologous polypeptide into the albumin locus of a host cell or population of cells, the method comprising ... c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; e) a sequence that is at least one of the sequences selected from the group consisting of SEQ ID NOs: 2 to 33. f) a sequence selected from the group consisting of SEQ ID NOs: 34-97; g) a sequence that is complementary to 15 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2-33; h) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 98-119; i) a gRNA comprising a sequence selected from the group consisting of SEQ ID NOs: 98-119; and j) a sequence selected from the group consisting of SEQ ID NOs: 120-163; ii) an RNA-guided DNA-binding agent; and iii) a construct comprising a nucleic acid encoding a heterologous polypeptide, thereby inserting the nucleic acid encoding the heterologous polypeptide into the albumin locus of the host cell or cell population.
[0005] In another aspect, the disclosure provides a method of expressing a heterologous polypeptide from the albumin locus in a host cell or population of cells, comprising: i) expressing a heterologous polypeptide from the albumin locus in a host cell or population of cells that is at least 95%, 90%, 85%, 80%, or 75% identical to a) a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33; b) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33; c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence selected from the group consisting of SEQ ID NOs: 2 a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33; b) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33; c) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33; f) a sequence selected from the group consisting of SEQ ID NOs: 34-97; and g) a sequence that comprises 15 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2-33; ii) an RNA-guided DNA-binding agent; and iii) a construct that comprises a coding sequence for a heterologous polypeptide, thereby expressing the heterologous polypeptide in the host cell or cell population.
[0006] In another aspect, the disclosure provides a method of expressing a therapeutic agent in a non-dividing cell type or population of cells, comprising: i) expressing a therapeutic agent in a non-dividing cell type or population of cells comprising: a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 333; b) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33; c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence selected from the group consisting of SEQ ID NOs: 2-33; a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of: e) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33; f) a sequence selected from the group consisting of SEQ ID NOs: 34-97; and g) a sequence comprising 15 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2-33; ii) an RNA-guided DNA-binding agent; and iii) a construct comprising a coding sequence for a heterologous polypeptide, thereby expressing the therapeutic agent in the non-dividing cell type or cell population.
[0007] In some embodiments, the gRNA comprises a guide sequence selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and SEQ ID NO:33.
[0008] In some embodiments, the method is performed in vivo. In some embodiments, the method is performed in vitro.
[0009] In some embodiments, the gRNA binds to a region upstream of a propospacer adjacent motif (PAM). In some embodiments, the PAM is selected from NGG, NNGRRT, NNGRR(N), NNAGAAW, NNNNG(A / C)TT, and NNNNRYAC.
[0010] In some embodiments, the gRNA is a dual gRNA (dgRNA). In some embodiments, the gRNA is a single gRNA (sgRNA). In some embodiments, the sgRNA comprises one or more modified nucleosides. In some embodiments, the Cas nuclease is a Class 2 Cas nuclease. In some embodiments, the Cas nuclease is selected from the group consisting of S. pyogenes nuclease, S. aureus nuclease, C. jejuni nuclease, S. thermophilus nuclease, N. meningitidis nuclease, and variants thereof. In some embodiments, the Cas nuclease is Cas9. In some embodiments, the Cas nuclease is a nickase.
[0011] In some embodiments, the construct is a bidirectional nucleic acid construct. In some embodiments, the construct comprises: i. a first segment comprising a coding sequence for a heterologous polypeptide; and ii. a second segment comprising the reverse complement of the coding sequence for the heterologous polypeptide. In some embodiments, the construct comprises a polyadenylation signal sequence. In some embodiments, the construct comprises a splice acceptor site. In some embodiments, the construct does not comprise homology arms.
[0012] In some embodiments, the gRNA is administered in a vector and / or lipid nanoparticle. In some embodiments, the RNA-guided DNA-binding agent is administered in a vector and / or lipid nanoparticle. In some embodiments, the construct comprising a heterologous gene is administered in a vector and / or lipid nanoparticle. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is selected from the group consisting of an adeno-associated viral (AAV) vector, an adenoviral vector, a retroviral vector, and a lentiviral vector. In some embodiments, the AAV vector is selected from the group consisting of AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof.
[0013] In some embodiments, the construct comprising the gRNA, the RNA-guided DNA-binding agent, and the coding sequence of the heterologous polypeptide are administered simultaneously, individually or in any combination. In some embodiments, the construct comprising the gRNA, the RNA-guided DNA-binding agent, and the coding sequence of the heterologous polypeptide are administered sequentially in any order and / or in any combination. In some embodiments, the RNA-guided DNA-binding agent, or the combined RNA-guided DNA-binding agent and gRNA, are administered before providing the construct. In some embodiments, the construct comprising the coding sequence of the heterologous polypeptide is administered before the gRNA and / or the RNA-guided DNA-binding agent.
[0014] In some embodiments, the heterologous polypeptide is a secreted polypeptide. In some embodiments, the heterologous polypeptide is an intracellular polypeptide.
[0015] In some embodiments, the cells are liver cells. In some embodiments, the liver cells are hepatocytes.
[0016] In some embodiments, expression of a heterologous polypeptide in a host cell is increased by at least about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, or more relative to the level in the cell before administering the construct comprising the gRNA, the RNA-guided DNA-binding agent, and the coding sequence for the heterologous polypeptide.
[0017] In some embodiments, the gRNA comprises SEQ ID NO: 301.
[0018] In some embodiments, the gRNA mediates target-specific cleavage by the RNA-guided DNA binder, resulting in insertion of the coding sequence of the heterologous polypeptide within intron 1 of the albumin gene. In some embodiments, cleavage results in an insertion rate of at least about 10% of the heterologous nucleic acid in the cell population. In some embodiments, cleavage results in an insertion rate of about 30-35%, about 35-40%, about 40-45%, about 45-50%, about 50-55%, about 55-60%, about 60-65%, about 65-70%, about 70-75%, about 75-80%, about 80-85%, about 85-90%, about 90-95%, or about 95-99% of the coding sequence of the heterologous polypeptide.
[0019] In some embodiments, the RNA-guided DNA-binding protein is S. pyogenes Cas9 nuclease. In some embodiments, the nuclease is a cleavase or a nickase.
[0020] In some embodiments, the method further comprises administering an LNP comprising a gRNA. In some embodiments, the method further comprises administering an LNP comprising an mRNA encoding an RNA-guided DNA-binding agent. In some embodiments, the LNP comprises a gRNA and an mRNA encoding the RNA-guided DNA-binding agent. In some embodiments, the gRNA and the RNA-guided DNA-binding protein are administered as an RNP. In some embodiments, the construct is administered via a vector.
[0021] In one aspect, the disclosure provides a host cell produced by any one or more of the aforementioned methods.
[0022] In one aspect, the present disclosure provides a cell comprising a bidirectional nucleic acid construct encoding a heterologous polypeptide integrated within intron 1 of the albumin locus of a host cell. In some embodiments, the host cell is a liver cell. In some embodiments, the liver cell is a hepatocyte. [Brief explanation of the drawings]
[0023] [Figure 1] Construct formats represented in the AAV genome are shown: SA = splice acceptor, pA = polyA signal sequence, HA = homology arm, LHA = left homology arm, RHA = right homology arm. [Figure 2] This shows that vectors without homology arms are not effective in immortalized liver cell lines (Hepa1-6). The scAAV derived from plasmid P00204, which contains 200-bp homology arms, resulted in the expression of hFIX in dividing cells. The use of AAV vectors derived from P00123 (scAAV lacking homology arms) and P00147 (ssAAV bidirectional construct lacking homology arms) did not result in detectable expression of hFIX. [Figure 3] Figures 3A and 3B show the results of in vivo testing of insertion templates with and without homology arms using vectors derived from P00123, P00147, or P00204. Figure 3A shows that liver editing levels, as measured by indel formation, of approximately 60% were detected in each group of animals treated with LNP containing CRISPR / Cas9 components. Figure 3B shows that animals receiving a ssAAV vector without homology arms (derived from P00147) combined with LNP treatment resulted in the highest levels of hFIX expression in serum. [Figure 4A]Figure 4 shows the results of in vivo testing of ssAAV insertion templates with and without homology arms, comparing targeted insertion with vectors derived from plasmids P00350, P00356, P00362 (with asymmetric homology arms as shown), and P00147 (bidirectional construct as shown in Figure 4B). [Figure 4B] Figure 1 shows the results of in vivo testing of ssAAV insertion templates with and without homology arms, comparing insertion into the targeted second site with vectors derived from plasmids P00353, P00354 (with symmetric homology arms as shown), and P00147. [Figure 5A] Figure 1 shows the results of targeted insertion of bidirectional constructs across 20 target sites in primary mouse hepatocytes. Schematics of each of the vectors tested are shown. [Figure 5B] 1 shows the results of targeted insertion of bidirectional constructs across 20 target sites in primary mouse hepatocytes. Editing, as measured by indel formation, is shown for each treatment group across each combination tested. [Figure 5C] Figure 1 shows the results of targeted insertion of bidirectional constructs across 20 target sites in primary mouse hepatocytes. It shows that significant levels of editing (as indel formation at specific target sites) did not necessarily result in more efficient insertion or expression of the transgene. hSA = human F9 splice acceptor, mSA = mouse albumin splice acceptor, HiBit = tag for luciferase-based detection, pA = polyA signal sequence, Nluc = nanoluciferase reporter, GFP = green fluorescent reporter. [Figure 5D] Figure 1 shows the results of targeted insertion of bidirectional constructs across 20 target sites in primary mouse hepatocytes. It shows that significant levels of editing (as indel formation at specific target sites) did not necessarily result in more efficient insertion or expression of the transgene. hSA = human F9 splice acceptor, mSA = mouse albumin splice acceptor, HiBit = tag for luciferase-based detection, pA = polyA signal sequence, Nluc = nanoluciferase reporter, GFP = green fluorescent reporter. [Figure 6]
[0033] Figure 1 shows the results of an in vivo screen of targeted insertion with bidirectional constructs across 10 target sites using ssAAV derived from P00147. As shown, significant levels of indel formation do not necessarily result in high levels of transgene expression. [Figure 7A]
[0023] Figure 1 shows the results of an in vivo screen of bidirectional constructs across 20 target sites using ssAAV derived from P00147. Varying levels of editing, as measured by indel formation, were detected for each treatment group across each LNP / vector combination tested. [Figure 7B]
[0023] Figure 1 shows the results of an in vivo screening of bidirectional constructs across 20 target sites using ssAAV derived from P00147. Corresponding target insertion data is provided. The results show poor correlation between indel formation and insertion or expression of bidirectional constructs. [Figure 7C]
[0023] Figure 1 shows the results of an in vivo screen of bidirectional constructs across 20 target sites using ssAAV derived from P00147. The results demonstrate a positive correlation between in vitro and in vivo results. [Figure 7D]
[0023] Figure 1 shows the results of an in vivo screen of bidirectional constructs across 20 target sites using ssAAV derived from P00147. The results show poor correlation between indel formation and insertion or expression of bidirectional constructs. [Figure 8A] Figures 8A and 8B show the insertion of the bidirectional construct at the cellular level using in situ hybridization with a probe capable of detecting the junction between the hFIX transgene and mouse albumin exon 1 sequence (Figure 8A). Circulating hFIX levels correlated with the number of cells that were positive for the hybrid transcript (Figure 8B). [Figure 8B] This is a continuation of Figure 8A. [Figure 9] 1 shows the effect on target insertion of varying timing between delivery of ssAAV containing bidirectional hFIX constructs and delivery of LNP. [Figure 10]1 shows the effect of repeated administrations of LNP (eg, 1, 2, or 3 administrations) on target insertion after delivery of a bidirectional hFIX construct. [Figure 11A] 1 shows the durability of hFIX expression in vivo. [Figure 11B] This shows that expression from intron 1 of albumin was maintained. [Figure 12A] This shows that the amount of hFIX expression from intron 1 of the albumin gene in vivo can be modulated by varying the AAV or LNP dose. [Figure 12B] This shows that the amount of hFIX expression from intron 1 of the albumin gene in vivo can be modulated by varying the AAV or LNP dose. [Figure 13A]
[0023] Figure 1 shows the results of screening bidirectional constructs spanning target sites in primary cynomolgus monkey hepatocytes, showing the various levels of editing, as measured by indel formation, detected for each of the samples. [Figure 13B] 1 shows the results of screening bidirectional constructs spanning the target site in primary cynomolgus monkey hepatocytes, demonstrating that significant levels of indel formation were not predictive of insertion or expression of bidirectional constructs into intron 1 of the albumin gene. [Figure 13C] 1 shows the results of screening bidirectional constructs spanning the target site in primary cynomolgus monkey hepatocytes, demonstrating that significant levels of indel formation were not predictive of insertion or expression of bidirectional constructs into intron 1 of the albumin gene. [Figure 14A]
[0023] Figure 1 shows the results of screening bidirectional constructs spanning target sites in primary human hepatocytes, with editing measured by indel formation detected for each of the samples. [Figure 14B] 1 shows the results of screening bidirectional constructs spanning the target site in primary human hepatocytes, demonstrating that significant levels of indel formation were not predictive of insertion or expression of bidirectional constructs into intron 1 of the albumin gene. [Figure 14C] 1 shows the results of screening bidirectional constructs spanning the target site in primary human hepatocytes, demonstrating that significant levels of indel formation were not predictive of insertion or expression of bidirectional constructs into intron 1 of the albumin gene. [Figure 14D] 1 shows the results of screening bidirectional constructs spanning the target site in primary human hepatocytes, demonstrating that significant levels of indel formation were not predictive of insertion or expression of bidirectional constructs into intron 1 of the albumin gene. [Figure 15]
[0023] Figure 1 shows the results of an in vivo study in which non-human primates were administered LNPs with a bidirectional hFIX insertion template (derived from P00147). Systemic hFIX levels were achieved only in animals treated with both LNPs and AAV; there was no detectable hFIX using AAV or LNPs alone. [Figure 16A] Human Factor IX expression levels in plasma samples 6 weeks after injection are shown. [Figure 16B] Human Factor IX expression levels in plasma samples 6 weeks after injection are shown. [Figure 17] Serum levels and % positive cells at week 7 across multiple lobes for each animal are shown. DETAILED DESCRIPTION OF THE INVENTION
[0024] Reference will now be made in detail to certain embodiments of the invention, examples of which are illustrated in the accompanying drawings. While the present teachings will be described in conjunction with various embodiments, it is not intended to limit the present teachings to those embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those skilled in the art.
[0025] Before describing the present teachings in detail, it should be understood that the present disclosure is not limited to particular compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context dictates otherwise. Thus, for example, a reference to "a conjugate" includes a plurality of conjugates, a reference to "a cell" includes a plurality of cells, and so forth. As used herein, the term "include" and grammatical variations thereof are intended to be open-ended, such that the recitation of items in a list does not exclude other similar items that may be substituted for or added to the listed items.
[0026] Numerical ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with measurement. Additionally, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting. It should be understood that both the foregoing general description and the detailed description are exemplary and illustrative only and are not limiting of the teachings herein.
[0027] Unless otherwise stated herein, embodiments herein that list various components as "comprising" are also assumed to "consist of" or "consist essentially of" the listed components. Embodiments herein that list various components as "consisting of" are also assumed to "comprising" or "essentially consisting of" the listed components. Embodiments herein that list various components as "consisting essentially of" are also assumed to "consist of" or "comprising" the listed components (this interchangeability does not apply to the use of these terms in the claims). The term "or" is used in an inclusive sense, i.e., equivalent to "and / or," unless the context clearly dictates otherwise. The term "about," when used before a list, modifies each element of the list. The term "about" or "approximately" refers to an allowable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined.
[0028] The term "about," when used before a list, modifies each element of the list. The term "about" or "approximately" refers to an acceptable error for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined.
[0029] The section headings used herein are for organizational purposes only and should not be construed as limiting the desired subject matter in any way. In the event that any material incorporated by reference conflicts with any term defined herein or any other explicit content of this specification, the present specification shall control. While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those skilled in the art.
[0030] I. Definition Unless otherwise stated, the following terms and phrases used herein are intended to have the following meanings:
[0031] "Polynucleotide" and "nucleic acid" are used herein to refer to polymeric compounds comprising nucleosides or nucleoside analogs (including traditional RNA, DNA, mixed RNA-DNA, and polymers that are analogs thereof) with nitrogenous heterocyclic bases or base analogs linked together along the backbone. The nucleic acid "backbone" can be composed of various linkages, including one or more of sugar phosphodiester linkages, peptide-nucleic acid linkages ("peptide nucleic acid" or PNA, PCT Publication No. WO 95 / 32305), phosphorothioate linkages, methylphosphonate linkages, or combinations thereof. The sugar moiety of the nucleic acid can be ribose, deoxyribose, or similar compounds with any substitution, e.g., 2' methoxy or 2' halide substitutions. The nitrogenous bases can be the traditional bases (A, G, C, T, U), their analogs (e.g., modified uridines such as 5-methoxyuridine, pseudouridine, or N1-methylpseudouridine, or others), derivatives of inosine, purine, or pyrimidine (e.g., N 4 -methyldeoxyguanosine, deaza- or aza-purines, deaza- or aza-pyrimidines, pyrimidine bases with a substituent at the 5- or 6-position (e.g., 5-methylcytosine), purine bases with a substituent at the 2-, 6-, or 8-position, 2-amino-6-methylaminopurine, O 6 -methylguanine, 4-thio-pyrimidine, 4-amino-pyrimidine, 4-dimethylhydrazine-pyrimidine, and O 4 5,378,825 and PCT Publication WO 93 / 13121). For a general discussion, see The Biochemistry of the Nucleic Acids, vol. 5-36, Adams et al., ed., 11 thed., 1992). Nucleic acids may contain one or more "abasic" residues, where the backbone does not contain a nitrogenous base at one or more positions in the polymer (U.S. Pat. No. 5,585,481). Nucleic acids may contain only conventional RNA or DNA sugars, bases, and linkages, or may contain both conventional building blocks and substitutions (e.g., conventional nucleosides with 2' methoxy substituents, or polymers containing both conventional nucleotides and one or more nucleotide analogs). Nucleic acids include "locked nucleic acids" (LNAs), which are analogs containing one or more LNA nucleotide monomers that have a bicyclic furanose unit locked to an RNA-mimetic sugar structure and increase hybridization affinity for complementary RNA and DNA sequences (Vester and Wengel, 2004, Biochemistry 43(42):13233-41). RNA and DNA have different sugar moieties and may differ by the presence of uracil or its analogs in RNA and thymine or its analogs in DNA.
[0032] The terms "guide RNA," "gRNA," and simply "guide" are used interchangeably herein to refer to a guide comprising a guide sequence, e.g., either a crRNA (also known as CRISPR RNA) or a combination of crRNA and trRNA (also known as tracrRNA). The crRNA and trRNA can associate as a single-stranded RNA molecule (single guide RNA, sgRNA) or, for example, in two separate RNA molecules (dual guide RNA, dgRNA). "Guide RNA" or "gRNA" refers to each type. The trRNA can be a naturally occurring sequence or a trRNA sequence with a modification or mutation compared to the naturally occurring sequence. Guide RNAs, such as sgRNA or dgRNA, can include modified RNAs as described herein.
[0033] As used herein, a "guide sequence" refers to a sequence within a guide RNA that is complementary to a target sequence and functions to direct the guide RNA to the target sequence for binding or modification (e.g., cleavage) by an RNA-guided DNA-binding agent. A "guide sequence" may also be referred to as a "targeting sequence" or a "spacer sequence." A guide sequence may be 20 base pairs in length, for example, in the case of Streptococcus pyogenes (i.e., Spy Cas9) and related Cas9 homologs / orthologs. Shorter or longer sequences, e.g., 15, 16, 17, 18, 19, 21, 22, 23, 24, or 25 nucleotides in length, can also be used as a guide. In some embodiments, the target sequence is, for example, within a gene or on a chromosome, and is complementary to the guide sequence. In some embodiments, the degree of complementarity or identity between a guide sequence and its corresponding target sequence may be about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the guide sequence and target region may be 100% complementary or identical. In other embodiments, the guide sequence and target region may contain at least one mismatch. For example, the guide sequence and target sequence may contain 1, 2, 3, or 4 mismatches, and the total length of the target sequence is at least 17, 18, 19, 20, or more base pairs. In some embodiments, the guide sequence and target region may contain 1 to 4 mismatches, and the guide sequence comprises at least 17, 18, 19, 20, or more nucleotides. In some embodiments, the guide sequence and target sequence may contain 1, 2, 3, or 4 mismatches, and the guide sequence comprises 20 nucleotides.
[0034] Because the nucleic acid substrate of an RNA-guided DNA binding agent is a double-stranded nucleic acid, the target sequence of the RNA-guided DNA binding agent includes both the plus and minus strands of genomic DNA (i.e., the given sequence and the reverse complement of the sequence). Thus, when a guide sequence is said to be "complementary to a target sequence," it should be understood that the guide RNA can be directed to bind to the sense or antisense strand (e.g., the reverse complement) of the target sequence. Thus, in some embodiments, when the guide sequence binds to the reverse complement of the target sequence, the guide sequence is identical to certain nucleotides of the target sequence (e.g., the target sequence without the PAM) except for the substitution of U for T in the guide sequence.
[0035] As used herein, "RNA-guided DNA binding agent" refers to a polypeptide or polypeptide complex having RNA and DNA binding activity, or a DNA-binding subunit of such a complex, where the DNA-binding activity is sequence-specific and dependent on the sequence of the RNA. The term RNA-guided DNA binding agent also includes nucleic acids encoding such polypeptides. Exemplary RNA-guided DNA binding agents include Cas cleavase / nickases. Exemplary RNA-guided DNA binding agents may include inactivated forms thereof ("dCas DNA binding agents") (e.g., when these agents are modified to enable DNA cleavage, e.g., via fusion with a FokI cleavase domain). As used herein, "Cas nuclease" encompasses Cas cleavase and Cas nickases. Cas cleavase and Cas nickases include the Csm or Cmr complex of a type III CRISPR system, its Cas10, Csm1, or Cmr2 subunit, the Cascade complex of a type I CRISPR system, its Cas3 subunit, and class 2 Cas nucleases. As used herein, a "Class 2 Cas nuclease" is a single-chain polypeptide with RNA-guided DNA-binding activity. Class 2 Cas nucleases include Class 2 Cas cleavase / nickases that also have RNA-guided DNA cleavase or nickase activity (e.g., H840A, D10A, or N863A variants), and Class 2 dCas DNA binders with inactivated cleavase / nickase activity, where the agents have been modified to enable DNA cleavage. Class 2 Cas nucleases include, for example, Cas9, Cpf1, C2c1, C2c2, C2c3, HF Cas9 (e.g., N497A, R661A, Q695A, Q926A variants), HypaCas9 (e.g., N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g., K810A, K1003A, R1060A variants), and eSPCas9(1.1) (e.g., K848A, K1003A, R1060A variants) proteins, and modifications thereof.The Cpf1 protein (Zetsche et al., Cell, 163:1-13 (2015)) also contains a RuvC-like nuclease domain. Zetsche's Cpf1 sequences are incorporated by reference in their entirety. See, e.g., Tables S1 and S3 of Zetsche. See, e.g., Makarova et al., Nat Rev Microbiol, 13(11):722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015). As used herein, delivery of an RNA-guided DNA-binding agent (e.g., a Cas nuclease, a Cas9 nuclease, or an S. pyogenes Cas9 nuclease) includes delivery of a polypeptide or mRNA.
[0036] As used herein, "ribonucleoprotein" (RNP) or "RNP complex" refers to a guide RNA in combination with an RNA-guided DNA-binding agent, e.g., a Cas nuclease, e.g., a Cas cleavase, a Cas nickase, a Cas9 cleavase, or a Cas9 nickase. In some embodiments, the guide sequence directs an RNA-guided DNA-binding agent, such as Cas9, to a target sequence, where the guide RNA hybridizes to the target sequence and the agent binds to the target sequence, followed by cleavage or nicking.
[0037] As used herein, a first sequence is considered to "contain a sequence having at least X% identity" with a second sequence if alignment of the first sequence to the second sequence shows that X% or more of the positions of the second sequence are matched overall by the first sequence. For example, the sequence AAGA contains a sequence that is 100% identical to the sequence AAG, because the alignment gives 100% identity in that there is a match at all three positions in the second sequence. Differences between RNA and DNA (generally, the exchange of thymidine for uridine, or vice versa), and the presence of nucleoside analogs such as modified uridines, do not contribute to differences in identity or complementarity between polynucleotides, as long as the related nucleotides (such as thymidine, uridine, or modified uridine) have the same complement (e.g., adenosine for all thymidine, uridine, or modified uridine; another example is cytosine and 5-methylcytosine, both of which have guanosine or modified guanosine as their complement). Thus, for example, the sequence 5'-AXG, where X is any modified uridine, such as pseudouridine, N1-methylpseudouridine, or 5-methoxyuridine, is considered 100% identical to AUG, in that both are perfectly complementary to the same sequence (5'-CAU). Exemplary alignment algorithms are the Smith-Waterman and Needleman-Wunsch algorithms, which are well known in the art. Those skilled in the art will understand what choice of algorithm and parameter settings is appropriate for a given pair of sequences to be aligned. Generally, for sequences of similar length and predicted identity greater than 50% for amino acids or 75% for nucleotides, the Needleman-Wunsch algorithm, using the default settings of the Needleman-Wunsch algorithm interface provided by EBI at the www.ebi.ac.uk web server, is generally appropriate.
[0038] As used herein, a first sequence is considered to be "X% complementary" to a second sequence if X% of the bases in the first sequence base pair with the second sequence. For example, a first sequence 5'AAGA3' is 100% complementary to a second sequence 3'TTCT5', and the second sequence is 100% complementary to the first sequence. In some embodiments, a first sequence 5'AAGA3' is 100% complementary to a second sequence 3'TTCTGTGA5', while the second sequence is 50% complementary to the first sequence.
[0039] "mRNA" is used herein to refer to a polynucleotide comprising an open reading frame that can be translated into a polypeptide (i.e., can serve as a substrate for translation by ribosomes and aminoacylated tRNAs). mRNA comprises primarily RNA or modified RNA, which may include a phosphate sugar backbone containing ribose residues or analogs thereof, e.g., 2'-methoxyribose residues. In some embodiments, the sugars of the mRNA phosphate sugar backbone consist essentially of ribose residues, 2'-methoxyribose residues, or combinations thereof. The bases of the mRNA may be modified bases, such as pseudouridine, N-1-methyl-pseudouridine, or other naturally occurring or non-naturally occurring bases.
[0040] Exemplary guide sequences useful in the compositions and methods described herein are provided in Table 1 and throughout the application.
[0041] As used herein, "indel" refers to an insertion / deletion mutation consisting of several nucleotides that are either inserted or deleted at the site of a double-strand break (DSB) in a target nucleic acid.
[0042] As used herein, "polypeptide" refers to a wild-type protein or a variant protein (e.g., a mutant, fragment, fusion, or a combination thereof). A variant polypeptide may have at least or about 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the functional activity of the wild-type polypeptide. In some embodiments, a variant is at least 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of a wild-type polypeptide. In some embodiments, a variant polypeptide may be a hyperactive variant. In certain instances, the variant has from about 80% to about 120%, 140%, 160%, 180%, 200%, 300%, 400%, 500% or more of the functional activity of the wild-type polypeptide.
[0043] As used herein, "target sequence" refers to a nucleic acid sequence within a target gene that has complementarity to the guide sequence of a gRNA. The interaction between the target sequence and the guide sequence directs the RNA-guided DNA-binding agent to bind and potentially (depending on the activity of the agent) nick or cleave within the target sequence.
[0044] As used herein, a "heterologous gene" refers to a gene that has been introduced as an exogenous source into a site within a host cell genome (e.g., the albumin intron 1 site). That is, the introduced gene is heterologous with respect to its insertion site. A polypeptide expressed from such a heterologous gene is referred to as a "heterologous polypeptide." A heterologous gene can be naturally occurring or engineered and can be wild-type or variant. A heterologous gene can contain nucleotide sequences other than the sequence encoding the heterologous polypeptide (e.g., an internal ribosome entry site). A heterologous gene can be a gene that naturally occurs in the host genome as a wild-type or variant (e.g., mutant). For example, a host cell contains a gene of interest (either wild-type or variant), but the same gene or a variant thereof can be introduced as an exogenous source for expression, for example, at a highly expressed locus. A heterologous gene can also be a gene that does not naturally occur in the host genome or that expresses a heterologous polypeptide that does not naturally occur in the host genome. The terms "heterologous gene," "exogenous gene," and "transgene" are used interchangeably. In some embodiments, a heterologous gene or transgene comprises an exogenous nucleic acid sequence, e.g., the nucleic acid sequence is not endogenous to the recipient cell. In some embodiments, a heterologous gene or transgene comprises an exogenous nucleic acid sequence, e.g., a nucleic acid sequence that does not naturally occur in the recipient cell. For example, a heterologous gene can be heterologous with respect to its insertion site and with respect to the recipient cell.
[0045] A heterologous gene can be inserted into a safe harbor locus within a genome without significant adverse effects on host cells, e.g., hepatocytes, e.g., without causing apoptosis, necrosis, and / or senescence, or without causing more than 5%, 10%, 15%, 20%, 25%, 30%, or 40% apoptosis, necrosis, and / or senescence compared to control cells. See, e.g., Hsin et al., “Hepatocyte Death in Liver Inflammation, Fibrosis, and Tumorigenesis,” 2017. In some embodiments, a safe harbor locus allows overexpression of an exogenous gene without significant adverse effects on host cells, e.g., hepatocytes, e.g., without causing apoptosis, necrosis, and / or senescence, or without causing more than 5%, 10%, 15%, 20%, 25%, 30%, or 40% apoptosis, necrosis, and / or senescence compared to control cells. In some embodiments, a desirable safe harbor locus can be a locus in which expression of the inserted gene sequence is not perturbed by read-through expression from adjacent genes. In some embodiments, safe harbor loci allow expression of exogenous genes without significant adverse effects on host cells or cell populations, such as hepatocytes or liver cells, e.g., without causing apoptosis, necrosis, and / or senescence, or without causing more than 5%, 10%, 15%, 20%, 25%, 30%, or 40% apoptosis, necrosis, and / or senescence compared to control cells or cell populations.
[0046] In some embodiments, a heterologous gene may be inserted into a safe harbor locus and use the endogenous signal sequence of the safe harbor locus, e.g., the albumin signal sequence encoded by exon 1. For example, a coding sequence may be inserted into human albumin intron 1 such that it is downstream of and fused to the signal sequence of human albumin exon 1.
[0047] In some embodiments, a gene may include its own signal sequence, may be inserted into a safe harbor locus, or may even use the endogenous signal sequence of a safe harbor locus. For example, a coding sequence including its native signal sequence may be inserted into human albumin intron 1 such that it is downstream of and fused with the signal sequence of human albumin encoded by exon 1.
[0048] In some embodiments, a gene may contain its own signal sequence and internal ribosome entry site (IRES), may be inserted into a safe harbor locus, or may even use the endogenous signal sequence of the safe harbor locus. For example, a coding sequence containing its native signal sequence and IRES sequence may be inserted into human albumin intron 1 such that it is downstream of and fused with the signal sequence of human albumin encoded by exon 1.
[0049] In some embodiments, a gene may contain its own signal sequence and IRES, or may be inserted into a safe harbor locus and not use the endogenous signal sequence of the safe harbor locus. For example, a coding sequence containing its native signal sequence and IRES sequence may be inserted into human albumin intron 1 such that it is not fused to the signal sequence of human albumin encoded by exon 1. In these embodiments, the protein is translated from the IRES site and is not a chimera (e.g., an albumin signal peptide fused to a heterologous protein), which may advantageously be non-immunogenic or of low immunogenicity. In some embodiments, the protein is not secreted and / or transported extracellularly.
[0050] In some embodiments, a gene may be inserted into a safe harbor locus, may contain an IRES, and does not use any signal sequence. For example, a coding sequence containing an IRES sequence and not a native signal sequence may be inserted into human albumin intron 1 such that it is not fused to the signal sequence of human albumin encoded by exon 1. In some embodiments, the protein is translated from the IRES site without any signal sequence. In some embodiments, the protein is not secreted and / or transported extracellularly.
[0051] As used herein, a "bidirectional nucleic acid construct" (interchangeably referred to herein as a "bidirectional construct") comprises at least two nucleic acid segments, one segment (the first segment) comprising a coding sequence that encodes an agent of interest (the coding sequence may be referred to herein as a "transgene" or first transgene), and the other segment (the second segment) comprising a sequence whose complement encodes the agent of interest, or a second transgene.
[0052] In one embodiment, a bidirectional construct comprises at least two nucleic acid segments in cis, one segment (the first segment) comprising a coding sequence (sometimes referred to interchangeably herein as a "transgene") and the other segment (the second segment) comprising a sequence whose complement encodes the transgene. The first transgene and the second transgene may be identical or different. A bidirectional construct may comprise at least two nucleic acid segments in cis, one segment (the first segment) comprising a coding sequence encoding a heterologous gene in one orientation and the other segment (the second segment) comprising a sequence whose complement encodes a heterologous gene in the other orientation. That is, the first segment is the complement (not necessarily a perfect complement) of the second segment, and the complement of the second segment is the reverse complement of the first segment (both encoding the same heterologous protein, but not necessarily a perfect reverse complement). A bidirectional construct may comprise a first coding sequence encoding a heterologous gene linked to a splice acceptor, and a second coding sequence whose complement encodes a heterologous gene in the other orientation and is also linked to a splice acceptor.
[0053] The agent can be a therapeutic agent, such as a polypeptide, functional RNA, or mRNA. The transgene can encode an agent, such as a polypeptide, functional RNA, or mRNA. In some embodiments, the bidirectional nucleic acid construct comprises at least two nucleic acid segments, one segment (the first segment) comprising a coding sequence encoding a polypeptide of interest, and the other segment (the second segment) comprising a sequence whose complement encodes the polypeptide of interest, or a second transgene. That is, the at least two segments can encode the same or different polypeptides or the same or different agents. When two segments encode the same polypeptide, the coding sequence of the first segment need not be identical to the complement of the sequence of the second segment. In some embodiments, the sequence of the second segment is the reverse complement of the coding sequence of the first segment. The bidirectional construct can be single-stranded or double-stranded. The bidirectional constructs disclosed herein encompass constructs capable of expressing any polypeptide of interest. The bidirectional constructs are useful for genomic insertion of transgene sequences, particularly for targeted insertion of transgenes.
[0054] In some embodiments, a bidirectional nucleic acid construct comprises a first segment comprising a coding sequence encoding a first polypeptide (first transgene) and a second segment comprising a sequence whose complement encodes a second polypeptide (second transgene). In some embodiments, the first and second polypeptides are at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical. In some embodiments, the first and second polypeptides comprise amino acid sequences that are at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical, e.g., over 50, 100, 200, 500, 1000 or more amino acid residues.
[0055] As used herein, "reverse complement" refers to a sequence that is the complementary sequence of a reference sequence, where the complementary sequence is written in reverse. For example, for a hypothetical sequence 5'CTGGACCGA3' (SEQ ID NO: 500), a "perfect" complementary sequence would be 3'GACCTGGCT5' (SEQ ID NO: 501), and a "perfect" reverse complementary sequence would be 5'TCGGTCCAG3' (SEQ ID NO: 502). A reverse complementary sequence need not be "perfect" and may still encode the same or a similar polypeptide as the reference sequence. Due to redundancy in codon usage, a reverse complement may deviate from the reference sequence to encode the same polypeptide. As used herein, "reverse complement" also includes a sequence that is, for example, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the reverse complement sequence of a reference sequence.
[0056] II. Composition A. Compositions Comprising Guide RNA (gRNA) Provided herein are compositions and methods useful for inserting and expressing heterologous (exogenous) genes within genomic loci, e.g., safe harbor sites, of host cells. In particular, as exemplified herein, targeting and inserting an exogenous gene at the albumin locus (e.g., intron 1) allows for the use of the endogenous promoter of albumin to drive robust expression of the exogenous gene. The present disclosure is based, in part, on the identification of guide RNAs that specifically target a site within intron 1 of the albumin gene and provide efficient insertion and expression of the exogenous gene. As shown in the Examples and further described herein, the ability of an identified gRNA to mediate high levels of editing, as measured by indel formation activity, surprisingly does not necessarily correlate with the use of the same gRNA to mediate efficient insertion of a transgene, as measured, for example, by transgene expression. That is, a particular gRNA capable of achieving significant levels of indel formation is not necessarily capable of mediating efficient insertion; conversely, some gRNAs shown to achieve low levels of indel formation may mediate efficient insertion and expression of a transgene. Specifically, the data in the Examples show that gRNAs that effectively mediated indel formation (also referred to as % editing) did not have indel editing activity that correlated with insertion editing activity.
[0057] In some embodiments, provided herein are compositions and methods useful for inserting and expressing an exogenous gene within intron 1 of the albumin locus in a host cell. In some embodiments, disclosed herein are compositions and methods useful for introducing or inserting a heterologous nucleic acid into the albumin locus of a host cell, e.g., using a guide RNA disclosed herein in conjunction with an RNA-guided DNA binding agent and a construct comprising the heterologous nucleic acid (a "transgene"). In some embodiments, disclosed herein are compositions and methods useful for expressing a heterologous polypeptide at the albumin locus of a host cell, e.g., using a guide RNA disclosed herein in conjunction with an RNA-guided DNA binding agent and a construct comprising the heterologous nucleic acid (a "transgene"). In some embodiments, disclosed herein are compositions and methods useful for inducing a break (e.g., a double-strand break (DSB) or a single-strand break (nick)) in the albumin gene of a host cell, e.g., using a guide RNA disclosed herein in conjunction with an RNA-guided DNA binding agent (e.g., a CRISPR / Cas system). The compositions and methods can be used in vitro or in vivo, e.g., for therapeutic purposes.
[0058] In some embodiments, the guide RNAs disclosed herein comprise a guide sequence that binds or is capable of binding within an intron of the albumin locus. In some embodiments, the guide RNAs disclosed herein bind within a region of intron 1 of the human albumin gene (SEQ ID NO: 1). It will be understood that not all bases of the guide sequence must bind within the recited region. For example, in some embodiments, 15, 16, 17, 18, 19, 20, or more bases of the guide RNA sequence bind to the recited region. For example, in some embodiments, 15, 16, 17, 18, 19, 20, or more consecutive bases of the guide RNA sequence bind to the recited region.
[0059] In some embodiments, the guide RNAs disclosed herein mediate target-specific cleavage by an RNA-guided DNA-binding agent (e.g., a Cas nuclease) at a site within human albumin intron 1 (SEQ ID NO: 1). It will be understood that in some embodiments, the guide RNA comprises a guide sequence that binds to or is capable of binding to a region in SEQ ID NO: 1.
[0060] In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 164-196. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 98-119. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from Table 1. A gRNA may comprise one or more of the guide sequences shown in Table 1. A gRNA may comprise one or more of the sequences shown in Tables 1, 7, and 9. The gRNA may comprise one or more of the sequences shown in Tables 2, 8, and 10. The guide RNA may comprise one or more of SEQ ID NOs: 2-33. The gRNA may comprise one or more of SEQ ID NOs: 164-196. The gRNA may comprise one or more of SEQ ID NOs: 98-119.
[0061] In some embodiments, a guide RNA disclosed herein comprises a guide sequence having at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence having at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 164-196. In some embodiments, a guide RNA disclosed herein comprises a guide sequence having at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 98-119. In some embodiments, a guide RNA disclosed herein comprises a guide sequence having at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from Table 1.
[0062] In some embodiments, the guide gRNA comprises a sequence selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and SEQ ID NO:33.
[0063] In some embodiments, the albumin guide RNA (gRNA) is selected from the group consisting of: a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of: 2, 8, 13, 19, 28, 29, 31, 32, 33; b) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of: 2, 8, 13, 19, 28, 29, 31, 32, 33; c) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of: 2, 8, 13, 19, 28, 29, 31, 32, 33; and 97, d) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs:2-33, e) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs:2-33, f) a sequence selected from the group consisting of SEQ ID NOs:34-97, and g) a sequence that is complementary to 15 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs:2-33. In some embodiments, the albumin guide RNA comprises a sequence selected from the group consisting of SEQ ID NOs:2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, the guide RNA comprises a sequence selected from the group consisting of SEQ ID NOs:4, 13, 17, 19, 27, 28, 30, and 31.
[0064] In some embodiments, the guide RNA disclosed herein binds to a region upstream of a propospacer adjacent motif (PAM). As will be understood by those skilled in the art, the PAM sequence occurs on the strand opposite the strand containing the target sequence. That is, the PAM sequence is on the complementary strand of the target strand (the strand containing the target sequence to which the guide RNA binds). In some embodiments, the PAM is selected from the group consisting of NGG, NNGRRT, NNGRR(N), NNAGAAW, NNNNG(A / C)TT, and NNNNRYAC. In some embodiments, the PAM is NGG.
[0065] In some embodiments, the guide RNA sequences provided herein are complementary to sequences adjacent to the PAM sequence.
[0066] In some embodiments, the guide RNA sequence comprises a sequence that is complementary to a sequence within a genomic region selected from Table 1 according to coordinates within the human reference genome hg38. In some embodiments, the guide RNA sequence comprises a sequence that is complementary to a sequence comprising 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 contiguous nucleotides from within a genomic region selected from Table 1. In some embodiments, the guide RNA sequence comprises a sequence that is complementary to a sequence comprising 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 contiguous nucleotides spanning a genomic region selected from Table 1.
[0067] The guide RNAs disclosed herein mediate target-specific cleavage that results in a double-strand break (DSB). The guide RNAs disclosed herein mediate target-specific cleavage that results in a single-strand break (SSB or nick).
[0068] In some embodiments, a guide RNA disclosed herein mediates target-specific cleavage by an RNA-guided DNA binder (e.g., a Cas nuclease as disclosed herein), resulting in insertion of a heterologous nucleic acid within intron 1 of the albumin gene. In some embodiments, cleavage at the guide RNA and / or cleavage site results in an insertion rate of 30-35%, 35-40%, 40-45%, 45-50%, 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-99% of the heterologous gene. In some embodiments, the guide RNA and / or cleavage results in insertion of at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the heterologous nucleic acid. The insertion rate can be measured in vitro or in vivo. For example, in some embodiments, the insertion rate can be determined by detecting and measuring the inserted nucleic acid in a cell population and calculating the percentage of the population containing the inserted nucleic acid. Methods for measuring the insertion rate are known and available in the art. In some embodiments, the guide RNA allows for increased expression of the heterologous gene by 5-10%, 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, 45-50%, 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, 95-99%, or more. The increase in expression of the heterologous gene can be measured in vitro or in vivo. For example, in some embodiments, increased expression can be determined by, e.g., detecting and measuring polypeptide levels before treating the cell or before administration to a subject and comparing the levels to the polypeptide levels, hi some embodiments, increased expression can be determined by detecting and measuring heterologous polypeptide levels and comparing the levels to known polypeptide levels, e.g., normal levels of the polypeptide in a healthy subject.
[0069] Each of the guide sequences shown in Table 1 may further comprise additional nucleotides to form a crRNA, for example, having the following exemplary nucleotide sequence at its 3' end following the guide sequence: GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 300) in the 5' to 3' direction. Genomic coordinates are according to the human reference genome hg38. In the case of an sgRNA, the above-mentioned guide sequences may further comprise additional nucleotides to form an sgRNA, for example, having the following exemplary nucleotide sequence at the 3' end following the guide sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 301) or GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 302) in the 5' to 3' direction. Table 1: Human guide RNA sequences and chromosomal coordinates TIFF0007820434000001.tif246165TIFF0007820434000002.tif174165
[0070] The guide RNA may further comprise a trRNA. In each composition and method embodiment described herein, the crRNA and trRNA may be associated as a single RNA (sgRNA) or may be on separate RNAs (dgRNA). In the context of an sgRNA, the crRNA and trRNA components may be covalently linked, for example, via a phosphodiester bond or other covalent bond. In some embodiments, the sgRNA includes one or more linkages between nucleotides that are not phosphodiester bonds.
[0071] In each of the composition, use, and method embodiments described herein, the guide RNA may comprise two RNA molecules as a "dual guide RNA" or "dgRNA." The dgRNA comprises a first RNA molecule comprising a crRNA, e.g., comprising a guide sequence as shown in Table 1, and a second RNA molecule comprising a trRNA. The first and second RNA molecules may not be covalently linked, but may form an RNA duplex through base pairing between portions of the crRNA and trRNA.
[0072] In each of the embodiments of the compositions, uses, and methods described herein, the guide RNA may comprise a single RNA molecule, referred to as a "single guide RNA" or "sgRNA." The sgRNA may comprise a crRNA (or a portion thereof) comprising a guide sequence shown in Table 1 covalently linked to a trRNA. The sgRNA may comprise 15, 16, 17, 18, 19, or 20 consecutive nucleotides of a guide sequence shown in Table 1. In some embodiments, the crRNA and trRNA are covalently linked via a linker. In some embodiments, the sgRNA forms a stem-loop structure by base pairing between portions of the crRNA and the trRNA. In some embodiments, the crRNA and trRNA are covalently linked via one or more bonds that are not phosphodiester bonds.
[0073] In some embodiments, the trRNA may comprise all or a portion of a trRNA sequence from a naturally occurring CRISPR / Cas system. In some embodiments, the trRNA comprises a truncated or modified wild-type trRNA. The length of the trRNA depends on the CRISPR / Cas system used. In some embodiments, the trRNA comprises or consists of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 nucleotides. In some embodiments, the trRNA may comprise a particular secondary structure, such as, for example, one or more hairpin or stem-loop structures, or one or more bulges.
[0074] In some embodiments, a target sequence or a region within intron 1 of the human albumin locus (e.g., a nucleotide sequence corresponding to a region within SEQ ID NO: 1) may be complementary to the guide sequence of a guide RNA. In some embodiments, the degree of complementarity or identity between the guide sequence of a guide RNA and its corresponding target sequence may be at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the target sequence and the guide sequence of a gRNA may be 100% complementary or identical. In other embodiments, the target sequence and the guide sequence of a gRNA may contain at least one mismatch. For example, the target sequence and the guide sequence of a gRNA may contain 1, 2, 3, 4, or 5 mismatches, and the total length of the guide sequence is about 20 or 20 nucleotides. In some embodiments, the target sequence and the guide sequence of a gRNA may contain 1 to 4 mismatches, and the guide sequence is about 20 or 20 nucleotides.
[0075] As described and exemplified herein, albumin guide RNAs can be used to insert and express heterologous genes (e.g., transgenes) in intron 1 of the albumin gene. Accordingly, in some embodiments, the present disclosure includes compositions comprising one or more guide RNAs (gRNAs) that include a guide sequence that directs an RNA-guided DNA-binding agent (e.g., Cas9) to a target DNA sequence within the albumin gene.
[0076] In some embodiments, a composition or formulation disclosed herein comprises an mRNA comprising an open reading frame (ORF) encoding an RNA-guided DNA binder, e.g., a Cas nuclease described herein. As described below, the mRNA comprising the Cas nuclease may comprise a Cas9 nuclease, e.g., an S. pyogenes Cas9 nuclease with cleavase, nickase, and / or site-specific DNA-binding activity. In some embodiments, the ORF encoding the RNA-guided DNA nuclease is a "modified RNA-guided DNA binder ORF" or simply a "modified ORF," used as an abbreviation to indicate that the ORF is modified.
[0077] Cas9 ORFs, including modified Cas9 ORFs, are provided herein and are known in the art. As an example, a Cas9 ORF can be codon-optimized so that the coding sequence contains one or more alternative codons for one or more amino acids. As used herein, "alternative codons" refer to variations in codon usage for a given amino acid, which may or may not be preferred or optimized codons (codon optimization) for a given expression system. Preferred codon usage, or codons that are well tolerated in a given expression system, are known in the art. The Cas9 coding sequences, Cas9 mRNA, and Cas9 protein sequences of WO2013 / 176772, WO2014 / 065596, WO2016 / 106121, and WO2019 / 067910 are incorporated herein by reference. In particular, the ORF and Cas9 amino acid sequences in the table in paragraph
[0449] of WO2019 / 067910, and the Cas9 mRNA and ORF in paragraphs
[0214] to
[0234] of WO2019 / 067910 are incorporated herein by reference.
[0078] In some embodiments, an RNA-guided DNA-binding agent, e.g., an mRNA comprising an ORF encoding a Cas nuclease, is provided, used, or administered.
[0079] B. Modified gRNA and mRNA In some embodiments, the gRNA is chemically modified. A gRNA that includes one or more modified nucleosides or nucleotides is referred to as a "modified" gRNA or a "chemically modified" gRNA to describe the presence of one or more non-natural and / or naturally occurring components or arrangements used in place of, or in addition to, the standard A, G, C, and U residues. In some embodiments, the modified gRNA is synthesized using non-standard nucleosides or nucleotides and is referred to herein as "modified." Modified nucleosides and nucleotides can include one or more of the following: (i) an alteration in the phosphodiester backbone linkage, e.g., replacement of one or both of the non-bridging phosphate oxygens and / or one or more of the bridging phosphate oxygens (exemplary backbone modifications); (ii) an alteration, e.g., replacement, of a component of the ribose sugar, e.g., the 2' hydroxyl of the ribose sugar (exemplary sugar modifications); (iii) extensive replacement of a phosphate moiety with a "dephosphorylated" linker (exemplary backbone modifications); (iv) a modification or replacement of a naturally occurring nucleobase, including with a non-standard nucleobase (exemplary base modifications); (v) a replacement or modification of the ribose phosphate backbone (exemplary backbone modifications); (vi) a modification of the 3' or 5' end of the oligonucleotide, e.g., removal, modification, or replacement of a terminal phosphate group, or conjugation of a moiety, cap, or linker (such 3' or 5' cap modifications can include sugar and / or backbone modifications); and (vii) a modification or replacement of the sugar (exemplary sugar modifications).
[0080] Chemical modifications such as those listed above can be combined to provide modified gRNAs and / or mRNAs comprising nucleosides and nucleotides (collectively "residues") that may have two, three, four, or more modifications. For example, modified residues may have a modified sugar and a modified nucleobase. In some embodiments, each base of a gRNA is modified, e.g., all bases have a modified phosphate group, such as a phosphorothioate group. In certain embodiments, all, or substantially all, of the phosphate groups of a gRNA molecule are replaced with phosphorothioate groups. In some embodiments, a modified gRNA comprises at least one modified residue at or near the 5' end of the RNA. In some embodiments, a modified gRNA comprises at least one modified residue at or near the 3' end of the RNA. Certain gRNAs comprise at least one modified residue at or near the 5' and 3' ends of the RNA.
[0081] In some embodiments, the gRNA comprises one, two, three, or more modified residues. In some embodiments, at least 5% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%) of the positions in the modified gRNA are modified nucleosides or nucleotides.
[0082] Unmodified nucleic acids may be susceptible to degradation, for example, by intracellular nucleases or those found in serum. For example, nucleases can hydrolyze phosphodiester bonds of nucleic acids. Thus, in one aspect, the gRNAs described herein may contain one or more modified nucleosides or nucleotides to introduce stability against, for example, intracellular or serum-derived nucleases. In some embodiments, the modified gRNA molecules described herein may exhibit a reduced innate immune response when introduced into a cell population both in vivo and ex vivo. The term "innate immune response" includes cellular responses to exogenous nucleic acids, including single-stranded nucleic acids, including the induction of cytokine (particularly interferon) expression and release and cell death.
[0083] In some embodiments of backbone modification, the phosphate group of the modified residue can be modified by replacing one or more oxygen atoms with different substituents.In addition, modified residues, for example, modified residues present in modified nucleic acids, can include large-scale replacement of unmodified phosphate moieties with modified phosphate groups, as described herein.In some embodiments, backbone modification of phosphate backbone can include modifications that result in either uncharged linkers or charged linkers with asymmetric charge distribution.
[0084] Examples of modified phosphate groups include phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters. The phosphate atom in an unmodified phosphate group is achiral. However, replacing one of the non-bridging oxygens with one of the atoms or groups of atoms described above can make the phosphorus atom chiral. The stereogenic phosphorus atom can have either the "R" configuration (herein Rp) or the "S" configuration (herein Sp). The backbone can also be modified by replacing the bridging oxygen (i.e., the oxygen connecting the phosphate to the nucleoside) with nitrogen (bridging phosphoramidates), sulfur (bridging phosphorothioates), and carbon (bridging methylene phosphonates). Replacement can occur at either or both bridging oxygens.
[0085] The phosphate group can be replaced by a non-phosphorus-containing linking group in certain backbone modifications. In some embodiments, the charged phosphate group can be replaced by a neutral moiety. Examples of moieties that can replace the phosphate group can include, but are not limited to, methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, and methyleneoxymethylimino.
[0086] Nucleic acid-mimicking scaffolds can also be constructed, where the phosphate linker and ribose sugar are replaced with nuclease-resistant nucleoside or nucleotide surrogates. Such modifications can include backbone and sugar modifications. In some embodiments, the nucleobases can be tethered by surrogate backbones. Examples can include, but are not limited to, morpholino, cyclobutyl, pyrrolidine, and peptide nucleic acid (PNA) nucleoside surrogates.
[0087] Modified nucleosides and nucleotides can include one or more modifications to the sugar group, i.e., sugar modifications. For example, the 2' hydroxyl group (OH) can be modified, e.g., replaced with a number of different "oxy" or "deoxy" substituents. In some embodiments, modifications to the 2' hydroxyl group can enhance the stability of the nucleic acid because the hydroxyl can no longer be further deprotonated to form a 2'-alkoxide ion.
[0088] Examples of 2' hydroxyl group modifications include alkoxy or aryloxy (OR, where "R" can be alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), polyethylene glycol (PEG), O(CHCHO) n The 2' hydroxyl group modification may include CH2CH2OR, where R can be, for example, H or an optionally substituted alkyl, and n can be an integer from 0 to 20 (e.g., 0 to 4, 0 to 8, 0 to 10, 0 to 16, 1 to 4, 1 to 8, 1 to 10, 1 to 16, 1 to 20, 2 to 4, 2 to 8, 2 to 10, 2 to 16, 2 to 20, 4 to 8, 4 to 10, 4 to 16, and 4 to 20). In some embodiments, the 2' hydroxyl group modification can be 2'-O-Me. In some embodiments, the 2' hydroxyl group modification can be a 2'-fluoro modification, which replaces the 2'-hydroxyl group with fluorine. In some embodiments, the 2' hydroxyl group modification can be 2'-H, which replaces the 2'-hydroxyl group with hydrogen. In some embodiments, the 2' hydroxyl group modification can be a 2'-hydroxyl group modification, which replaces the 2'-hydroxyl group with a fluorine. 1-6 Alkylene or C 1-6They may also include "locked" nucleic acids (LNAs) that may be connected to the 4' carbon of the same ribose sugar via a heteroalkylene bridge, exemplary bridges being methylene, propylene, ether, or amino bridge, O-amino (amino may be, for example, NH, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino), and aminoalkoxy, O(CH) n -amino (amino can be, for example, NH, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino). In some embodiments, the 2' hydroxyl group modification can include an "unlocked" nucleic acid (UNA), in which the ribose ring lacks a C2'-C3' bond. In some embodiments, the 2' hydroxyl group modification can include a methoxyethyl group (MOE), (OCH2CHOCH3, e.g., a PEG derivative).
[0089] A "deoxy" 2' modification can be hydrogen (i.e., a deoxyribose sugar, e.g., in an overhang portion of a partial dsRNA), halo (e.g., bromo, chloro, fluoro, or iodo), amino (amino can be, e.g., NH, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid), NH(CHCHNH) n and alkyl, cycloalkyl, aryl, alkenyl, and alkynyl, optionally substituted with, for example, amino as described herein.
[0090] Sugar modifications can include sugar groups containing one or more carbons with the opposite stereochemical configuration to that of the corresponding carbon in ribose. Thus, modified nucleic acids can include nucleotides containing, for example, arabinose as the sugar. Modified nucleic acids can also include abasic sugars. These abasic sugars can also be further modified at one or more of the constituent sugar atoms. Modified nucleic acids can also include one or more sugars that are L-form, for example, L-nucleosides.
[0091] The modified nucleosides and modified nucleotides described herein that can be incorporated into modified nucleic acids can contain modified bases, also referred to as nucleobases. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleobases can be modified or completely replaced to result in modified residues that can be incorporated into modified nucleic acids. The nucleobases of a nucleotide can be independently selected from purines, pyrimidines, purine analogs, or pyrimidine analogs. In some embodiments, the nucleobases can include, for example, naturally occurring bases and synthetic derivatives of bases.
[0092] In embodiments using dual guide RNAs, each of the crRNA and tracrRNA can include modifications. Such modifications may be present at one or both ends of the crRNA and / or tracrRNAA. In embodiments including an sgRNA, one or more residues at one or both ends of the sgRNA may be chemically modified, and / or internal nucleosides may be modified, and / or the entire sgRNA may be chemically modified. Certain embodiments include 5'-end modifications. Certain embodiments include 3'-end modifications.
[0093] In some embodiments, the guide RNAs disclosed herein comprise one of the modification patterns disclosed in WO2018 / 107028A1, filed December 8, 2017, entitled "Chemically Modified Guide RNAs," the contents of which are incorporated herein by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structure / modification patterns disclosed in US20170114334, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structure / modification patterns disclosed in WO2017 / 136794, the contents of which are incorporated herein by reference in their entirety.
[0094] In some embodiments, sgRNAs of the present disclosure comprise the modification patterns shown in Table 2 below. "Full Sequence" in Table 2 refers to the sgRNA sequence for each of the guides listed in Table 1. "Full Sequence Modification" indicates the modification pattern for each sgRNA. Table 2: Modification patterns of sgRNA and human albumin guide sequence to sgRNA TIFF0007820434000003.tif244167TIFF0007820434000004.tif250162TIFF0007820434000005.tif250162 TIFF0007820434000006.tif249162TIFF0007820434000007.tif248162TIFF0007820434000008.tif173166
[0095] In some embodiments, the modified sgRNA comprises the following sequence: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 350), where "N" can be any natural or non-natural nucleotide, and the entirety of N comprises the albumin intron 1 guide sequence set forth in Table 1. For example, SEQ ID NO: 350 is encompassed herein, where N is omitted from SEQ ID NO: 350 but includes the modified conserved portion of the gRNA.
[0096] Any of the modifications described below can be present in the gRNAs and mRNAs described herein.
[0097] The terms "mA," "mC," "mU," or "mG" may be used to refer to nucleotides modified with 2'-O-Me.
[0098] The 2'-O-methyl modification can be depicted as follows: TIFF0007820434000009.tif55166
[0099] Another chemical modification that has been shown to affect the nucleotide sugar ring is halogen substitution. For example, 2'-fluoro (2'-F) substitution on the nucleotide sugar ring can increase oligonucleotide binding affinity and nuclease stability.
[0100] In this application, the terms "fA," "fC," "fU," or "fG" may be used to refer to a nucleotide substituted with a 2'-F.
[0101] The substitution of 2'-F can be depicted as follows: TIFF0007820434000010.tif56158
[0102] Phosphorothioate (PS) linkage or bond refers to a bond in which sulfur replaces one non-bridging phosphate oxygen in a phosphodiester linkage, such as a bond between nucleotide bases.When phosphorothioate is used to generate an oligonucleotide, the modified oligonucleotide can also be referred to as S-oligo.
[0103] "*" may be used to indicate a PS modification. In this application, the terms A*, C*, U*, or G* may be used to indicate a nucleotide that is linked to the next (e.g., 3') nucleotide with a PS bond.
[0104] In this application, the terms "mA*," "mC*," "mU*," or "mG*" may be used to refer to a nucleotide that is substituted with 2'-O-Me and linked to the next (e.g., 3') nucleotide with a PS bond.
[0105] The diagram below shows the substitution of S- for the non-bridging phosphate oxygen, resulting in a PS bond instead of a phosphodiester bond. TIFF0007820434000011.tif74152
[0106] An abasic nucleotide is one that lacks a nitrogenous base. The diagram below depicts an oligonucleotide with an abasic (also known as apurinic) site that lacks a base. TIFF0007820434000012.tif84151
[0107] Inverted bases refer to those having a linkage that is inverted from the normal 5' to 3' linkage (i.e., either a 5' to 5' linkage or a 3' to 3' linkage). For example, TIFF0007820434000013.tif67150
[0108] The abasic nucleotide can be connected with an inverted linkage. For example, the abasic nucleotide can be connected to the terminal 5' nucleotide via a 5' to 5' linkage, or the abasic nucleotide can be connected to the terminal 3' nucleotide via a 3' to 3' linkage. An inverted abasic nucleotide at either the terminal 5' or 3' nucleotide can also be referred to as an inverted abasic end cap.
[0109] In some embodiments, one or more of the first 3, 4, or 5 nucleotides at the 5' end and one or more of the last 3, 4, or 5 nucleotides at the 3' end are modified, hi some embodiments, the modifications are 2'-O-Me, 2'-F, inverted abasic nucleotides, PS linkages, or other nucleotide modifications known in the art to increase stability and / or performance.
[0110] In some embodiments, the first four nucleotides at the 5' end and the last four nucleotides at the 3' end are linked using phosphorothioate (PS) bonds.
[0111] In some embodiments, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end comprise 2'-O-methyl (2'-O-Me) modified nucleotides. In some embodiments, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end comprise 2'-fluoro (2'-F) modified nucleotides. In some embodiments, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end comprise inverted abasic nucleotides.
[0112] In some embodiments, the guide RNA comprises a modified sgRNA. In some embodiments, the sgRNA comprises the modification pattern set forth in SEQ ID NO: 350, where N is any natural or unnatural nucleotide, and the entirety of N comprises a guide sequence (e.g., as shown in Table 1) that directs a nuclease to a target sequence within human albumin intron 1.
[0113] In some embodiments, the guide RNA comprises an sgRNA set forth in any one of SEQ ID NOs: 34-97 or 120-163. In some embodiments, the guide RNA comprises an sgRNA set forth in any one of SEQ ID NOs: 197-229. In some embodiments, the guide RNA comprises a guide sequence set forth in SEQ ID NOs: 2-33 or 98-119 and an sgRNA comprising any one of the nucleotides of SEQ ID NO: 301, wherein the nucleotide of SEQ ID NO: 301 is at the 3' end of the guide sequence, and the sgRNA may be modified as set forth in Table 2 or SEQ ID NO: 350. In some embodiments, the guide RNA comprises a guide sequence set forth in SEQ ID NOs: 2-33 or 197-229 and an sgRNA comprising the nucleotide of SEQ ID NO: 301, wherein the nucleotide of SEQ ID NO: 301 is at the 3' end of the guide sequence, and the sgRNA may be modified as set forth in Table 2 or SEQ ID NO: 350, for example.
[0114] As mentioned above, in some embodiments, the compositions or formulations disclosed herein comprise an mRNA comprising an open reading frame (ORF) encoding an RNA-guided DNA binder, e.g., a Cas nuclease described herein. In some embodiments, an mRNA comprising an ORF encoding an RNA-guided DNA binder, e.g., a Cas nuclease, is provided, used, or administered. In some embodiments, the ORF encoding the RNA-guided DNA nuclease is a "modified RNA-guided DNA binder ORF" or simply a "modified ORF," used as an abbreviation to indicate that the ORF is modified.
[0115] In some embodiments, the modified ORF can include modified uridines at at least one, more than one, or all uridine positions. In some embodiments, the modified uridine is a uridine modified at the 5-position, e.g., with a halogen, methyl, or ethyl. In some embodiments, the modified uridine is a pseudouridine modified at the 1-position, e.g., with a halogen, methyl, or ethyl. The modified uridine can be, for example, pseudouridine, N1-methyl-pseudouridine, 5-methoxyuridine, 5-iodouridine, or a combination thereof. In some embodiments, the modified uridine is 5-methoxyuridine. In some embodiments, the modified uridine is 5-iodouridine. In some embodiments, the modified uridine is pseudouridine. In some embodiments, the modified uridine is N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and 5-methoxyuridine. In some embodiments, the modified uridine is a combination of N1-methylpseudouridine and 5-methoxyuridine. In some embodiments, the modified uridine is a combination of 5-iodouridine and N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and 5-iodouridine. In some embodiments, the modified uridine is a combination of 5-iodouridine and 5-methoxyuridine.
[0116] In some embodiments, the mRNA disclosed herein includes a 5' cap, e.g., Cap0, Cap1, or Cap2. The 5' cap is generally a 7-methylguanine ribonucleotide (e.g., with respect to ARCA, which may be further modified as discussed below) linked via a 5'-triphosphate to the 5' position of the first nucleotide (i.e., the first cap-proximal nucleotide) of the 5'-3' strand of the mRNA. In Cap0, the riboses of the first and second cap-proximal nucleotides of the mRNA both include a 2'-hydroxyl. In Cap1, the riboses of the first and second transcribed nucleotides of the mRNA include a 2'-methoxy and a 2'-hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the mRNA both include a 2'-methoxy. See, for example, Katibah et al. (2014) Proc Natl Acad Sci USA 111(33):12025-30 and Abbas et al. (2017) Proc Natl Acad Sci USA 114(11):E2106-E2115. Most endogenous higher eukaryotic mRNAs, including mammalian mRNAs such as human mRNAs, contain Cap1 or Cap2. Cap0 and other cap structures distinct from Cap1 and Cap2 are recognized as "non-self" by components of the innate immune system, such as IFIT-1 and IFIT-5, and can be immunogenic in mammals, including humans, and can cause elevated levels of cytokines, including type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, can also compete with eIF4E for binding to mRNAs with caps other than Cap1 or Cap2, potentially inhibiting mRNA translation.
[0117] A cap can be included co-transcriptionally. For example, ARCA (anti-reverse cap analog, Thermo Fisher Scientific catalog number AM8045) is a cap analog containing 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of guanine ribonucleotides and can be incorporated into transcripts in vitro at initiation. ARCA results in a Cap0 cap in which the 2' position of the first cap-proximal nucleotide is hydroxyl. See, for example, Stepinski et al. (2001) "Synthesis and properties of mRNAs containing the novel 'anti-reverse' cap analogs 7-methyl(3'-O-methyl)GpppG and 7-methyl(3'deoxy)GpppG," RNA 7:1486-1495. The structure of ARCA is shown below. TIFF0007820434000014.tif31150
[0118] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG, TriLink Biotechnologies catalog number N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG, TriLink Biotechnologies catalog number N-7133) can be used to co-transcriptionally provide the Cap1 structure. 3'-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also available from TriLink Biotechnologies as catalog numbers N-7413 and N-7433, respectively. The structure of CleanCap™ AG is shown below. TIFF0007820434000015.tif64150
[0119] Alternatively, a cap can be added to RNA after transcription. For example, Vaccinia capping enzyme is commercially available (New England Biolabs, catalog number M2080S), which possesses RNA triphosphatase and guanylyltransferase activities provided by its D1 subunit and guanine methyltransferase activity provided by its D12 subunit. Thus, in the presence of S-adenosylmethionine and GTP, 7-methylguanine can be added to RNA to give Cap0. See, e.g., Guo, P. and Moss, B. (1990) Proc. Natl. Acad. Sci. USA 87, 4023-4027; Mao, X. and Shuman, S. (1994) J. Biol. Chem. 269, 24472-24479.
[0120] In some embodiments, the mRNA further comprises a polyadenylation (poly-A) tail. In some embodiments, the poly-A tail comprises at least 20, 30, 40, 50, 60, 70, 80, 90, or 100 adenines, and optionally up to 300 adenines. In some embodiments, the poly-A tail comprises 95, 96, 97, 98, 99, or 100 adenine nucleotides.
[0121] C. RNA-guided DNA binders As described herein, guide RNAs of the present disclosure are used in combination with RNA-guided DNA binding agents to insert and express heterologous (exogenous) genes into genomic loci, such as safe harbor sites, of host cells. The RNA-guided DNA binding agent can be a protein or a nucleic acid encoding a protein, such as an mRNA. In some embodiments, methods of the present disclosure involve the use of a composition comprising a guide RNA comprising a guide sequence from Table 1 and an RNA-guided DNA binding agent, e.g., a nuclease, such as a Cas nuclease (e.g., Cas9), to form a ribonucleoprotein complex.
[0122] In some embodiments, the RNA-guided DNA binding agent, such as a Cas9 nuclease, has cleavase activity, which may also be referred to as double-stranded endonuclease activity. In some embodiments, the RNA-guided DNA binding agent, such as a Cas9 nuclease, has nickase activity, which may also be referred to as single-stranded endonuclease activity. In some embodiments, the RNA-guided DNA binding agent comprises a Cas nuclease. Examples of Cas nucleases include those from S. pyogenes, S. aureus, and other prokaryotic type II CRISPR systems (see, e.g., the list in the next paragraph), as well as variant or mutant (e.g., engineered, non-natural, natural, or other variant) versions thereof. See, e.g., US2016 / 0312198A1, US2016 / 0312199A1.
[0123] Non-limiting exemplary species from which Cas nuclease can be derived include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohlobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, Acidaminococcus sp.、Lachnospiraceae மாட்ட்டிந்துND2006、、、Acaryochloris marina is mentioned。.
[0124] In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas nuclease is a Cas9 nuclease from Neisseria meningitidis. In some embodiments, the Cas nuclease is a Cas9 nuclease from Staphylococcus aureus. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Francisella novicida. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Acidaminococcus sp. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Lachnospiraceae bacterium ND2006. In further embodiments, the Cas nuclease is a Cpf1 nuclease from Francisella tularensis, Lachnospiraceae bacteria, Butyrivibrio proteoclasticus, Peregrinibacteria bacteria, Parcubacteria bacteria, Smithella, Acidaminococcus, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevioricanis, Prevotella disiens, or Porphyromonas macacae. In certain embodiments, the Cas nuclease is a Cpf1 nuclease from Acidaminococcus or Lachnospiraceae.
[0125] In some embodiments, the gRNA combined with the RNA-guided DNA-binding agent is referred to as a ribonucleoprotein complex (RNP). In some embodiments, the RNA-guided DNA-binding agent is a Cas nuclease. In some embodiments, the gRNA combined with the Cas nuclease is referred to as a Cas RNP. In some embodiments, the RNP comprises type I, type II, or type III components. In some embodiments, the Cas nuclease is a Cas9 protein from a type II CRISPR / Cas system. In some embodiments, the gRNA combined with Cas9 is referred to as a Cas9 RNP.
[0126] Wild-type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 protein comprises more than one RuvC domain and / or more than one HNH domain. In some embodiments, the Cas9 protein is wild-type Cas9. In each of the composition, use, and method embodiments, Cas induces a double-strand break in the target DNA.
[0127] In some embodiments, chimeric Cas nucleases are used, in which one domain or region of a protein is replaced with a portion of a different protein. In some embodiments, the Cas nuclease domain may be replaced with a domain from a different nuclease, such as Fok1. In some embodiments, the Cas nuclease may be a modified nuclease.
[0128] In other embodiments, the Cas nuclease may be from a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a component of a cascade complex of a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a Cas3 protein. In some embodiments, the Cas nuclease may be from a type III CRISPR / Cas system. In some embodiments, the Cas nuclease may have RNA cleavage activity.
[0129] In some embodiments, the RNA-guided DNA binding agent has single-stranded nickase activity, i.e., it can cleave one DNA strand to generate a single-strand break (also known as a "nick"). In some embodiments, the RNA-guided DNA binding agent comprises a Cas nickase. A nickase is an enzyme that creates a nick in dsDNA, i.e., it cleaves one strand of the DNA double helix but not the other. In some embodiments, the Cas nickase is an aspect of a Cas nuclease (e.g., the Cas nucleases discussed above) in which the endonucleolytic active site has been inactivated, e.g., by one or more modifications (e.g., point mutations) in the catalytic domain. See, e.g., U.S. Patent No. 8,889,356 for a discussion of Cas nickases and exemplary catalytic domain modifications. In some embodiments, the Cas nickase, such as the Cas9 nickase, has an inactivated RuvC or HNH domain.
[0130] In some embodiments, the RNA-guided DNA binder is modified to contain only one functional nuclease domain. For example, the drug protein may be modified so that one of the nuclease domains is mutated or completely or partially deleted to reduce nucleic acid cleavage activity. In some embodiments, a nickase with a RuvC domain that has reduced activity is used. In some embodiments, a nickase with an inactive RuvC domain is used. In some embodiments, a nickase with an HNH domain that has reduced activity is used. In some embodiments, a nickase with an inactive HNH domain is used.
[0131] In some embodiments, conserved amino acids within the Cas protein nuclease domain are substituted to reduce or alter nuclease activity. In some embodiments, the Cas nuclease may comprise an amino acid substitution within the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include D10A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015) Cell Oct 22:163(3):759-771. In some embodiments, the Cas nuclease may comprise an amino acid substitution within the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include E762A, H840A, N863A, H983A, and D986A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015). Further exemplary amino acid substitutions include D917A, E1006A, and D1255A (based on the Francisella novicida U112 Cpf1 (FnCpH) sequence (UniProtKB-A0Q7Q2(CPF1_FRATN))).
[0132] In some embodiments, a nickase is provided in combination with a pair of guide RNAs that are complementary to the sense and antisense strands of the target sequence, respectively. In this embodiment, the guide RNAs direct the nickase to the target sequence by generating nicks on opposite strands of the target sequence (i.e., double nicking), thereby introducing a DSB. In some embodiments, a nickase is used with two separate guide RNAs that target opposite strands of DNA, generating a double nick in the target DNA. In some embodiments, a nickase is used with two separate guide RNAs selected to be in close proximity, generating a double nick in the target DNA. In some embodiments, the RNA-guided DNA binder comprises one or more heterologous functional domains (e.g., is or comprises a fusion polypeptide).
[0133] In some embodiments, the heterologous functional domain may facilitate the transport of the RNA-guided DNA binding agent to the cell nucleus. For example, the heterologous functional domain may be a nuclear localization signal (NLS). In some embodiments, the RNA-guided DNA binding agent may be fused with 1 to 10 NLS(s). In some embodiments, the RNA-guided DNA binding agent may be fused with 1 to 5 NLS(s). In some embodiments, the RNA-guided DNA binding agent may be fused with one NLS. When one NLS is used, the NLS may be linked at the N-terminal or C-terminal end of the sequence of the RNA-guided DNA binding agent. It may also be inserted within the sequence of the RNA-guided DNA binding agent. In other embodiments, the RNA-guided DNA binding agent may be fused with more than one NLS. In some embodiments, the RNA-guided DNA binding agent may be fused with 2, 3, 4, or 5 NLSs. In some embodiments, the RNA-guided DNA binding agent may be fused with two NLSs. In certain situations, the two NLSs may be the same (e.g., two SV40 NLSs) or different. In some embodiments, the RNA-guided DNA binding agent is fused to two SV40 NLS sequences linked at the carboxy-terminal end. In some embodiments, the RNA-guided DNA binding agent may be fused with two NLSs, one linked at the N-terminal end and one linked at the C-terminal end. In some embodiments, the RNA-guided DNA binding agent may be fused with three NLSs. In some embodiments, the RNA-guided DNA binding agent may be fused without an NLS. In some embodiments, the NLS may be a monokaryotic sequence, such as the SV40 NLS, PKKKRKV (SEQ ID NO: 600) or PKKKRRV (SEQ ID NO: 601). In some embodiments, the NLS may be a bikaryotic sequence, such as the nucleoplasmin NLS KRPAATKKAGQAKKKK (SEQ ID NO: 602). In certain embodiments, a single PKKKRKV (SEQ ID NO: 600) NLS may be linked at the C-terminal end of the RNA-guided DNA binding agent. One or more linkers are optionally included at the fusion site.
[0134] D. Donor construct / sequence The compositions and methods described herein involve the use of nucleic acid constructs containing sequences encoding heterologous genes that are inserted into the cleavage sites created by the guide RNAs and RNA-guided DNA-binding agents of the present disclosure. As used herein, such constructs are sometimes referred to as "donor constructs / templates." The constructs can encode any expressed nucleic acid (i.e., nucleic acid that can be expressed), such as DNA, messenger RNA (mRNA), functional RNA, small interfering RNA (siRNA), microRNA (miRNA), single-stranded RNA (ssRNA), long non-coding RNA, or antisense oligonucleotides.
[0135] The compositions and methods described herein include the use of non-bidirectional or unidirectional constructs, e.g., encoding a single transgene, encoding two transgenes in cis, etc. Unidirectional constructs may include a coding sequence linked to a splice acceptor.
[0136] The compositions and methods described herein include the use of bidirectional constructs described herein that contain at least two nucleic acid segments in cis, where one segment (the first segment) contains a coding sequence or transgene and the other segment (the second segment) contains a sequence whose complement encodes the transgene. A bidirectional construct can contain a first coding sequence that encodes a heterologous gene linked to a splice acceptor, and a second coding sequence whose complement encodes a heterologous gene in the other orientation and is also linked to a splice acceptor.
[0137] In some embodiments, the constructs disclosed herein include a splice acceptor site at either or both ends of the construct, e.g., 5' of the open reading frame in the first and / or second segment or 5' of one or both transgene sequences. In some embodiments, the splice acceptor site comprises NAG. In further embodiments, the splice acceptor site consists of NAG. In some embodiments, the splice acceptor is an albumin splice acceptor, e.g., the albumin splice acceptor used in splicing together exons 1 and 2 of albumin. In some embodiments, the splice acceptor is derived from the human albumin gene. In some embodiments, the splice acceptor is derived from the mouse albumin gene. In some embodiments, the splice acceptor is an F9 (or "FIX") splice acceptor, e.g., the F9 splice acceptor used in splicing together exons 1 and 2 of F9. In some embodiments, the splice acceptor is derived from the human F9 gene. In some embodiments, the splice acceptor is derived from the mouse F9 gene. Additional suitable splice acceptor sites useful in eukaryotes, including artificial splice acceptors, are known and can be found in the art. See, for example, Shapiro et al., 1987, Nucleic Acids Res., 15, 7155-7174; Burset et al., 2001, Nucleic Acids Res., 29, 255-259.
[0138] In some embodiments, a polyadenylation tail sequence is encoded at the 3' end of the first and / or second segment, e.g., as a "poly-A" stretch. In some embodiments, a polyadenylation tail sequence is provided co-transcriptionally as a result of a polyadenylation signal sequence encoded at or near the 3' end of the first and / or second segment. Methods for designing suitable polyadenylation tail sequences and / or polyadenylation signal sequences are well known in the art. Suitable splice acceptor sequences are disclosed and exemplified herein, including mouse albumin and human FIX splice acceptor sites. In some embodiments, the polyadenylation signal sequence AAUAAA (SEQ ID NO: 800) is commonly used in mammalian systems, although variants such as UAUAAA (SEQ ID NO: 801) or AU / GUAAA (SEQ ID NO: 802) have been identified. See, e.g., NJ Proudfoot, Genes & Dev. 25(17):1770-82, 2011. In some embodiments, a poly(A) tail sequence is included. The length of the construct can vary depending on the size of the gene to be inserted, for example, from 200 base pairs (bp) to about 5000 bp, such as from about 200 bp to about 2000 bp, or from about 500 bp to about 1500 bp. In some embodiments, the length of the DNA donor template is about 200 bp, or about 500 bp, or about 800 bp, or about 1000 bp, or about 1500 bp. In other embodiments, the length of the donor template is at least 200 bp, or at least 500 bp, or at least 800 bp, or at least 1000 bp, or at least 1500 bp.
[0139] The constructs can be DNA or RNA, single-stranded, double-stranded, or partially single-stranded and partially double-stranded, and can be introduced into host cells in linear or circular (e.g., minicircle) form. See, for example, U.S. Patent Publication Nos. 2010 / 0047805, 2011 / 0281361, and 2011 / 0207221. When introduced in linear form, the ends of the donor sequence can be protected (e.g., from exonuclease degradation) by methods known to those skilled in the art. For example, one or more dideoxynucleotide residues can be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides can be ligated to one or both ends. See, for example, Chang et al. (1987) Proc. Natl. Acad. Sci. USA 84:4959-4963; Nehls et al. (1996) Science 272:886-889. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. Constructs can be introduced into cells as part of vector molecules with additional sequences, such as origins of replication, promoters, and genes encoding antibiotic resistance. Constructs can be free of viral elements. Furthermore, donor constructs can be introduced as naked nucleic acid, as nucleic acid complexed with agents such as liposomes or poloxamers, or delivered by viruses (e.g., adenovirus, AAV, herpesvirus, retrovirus, lentivirus).
[0140] In some embodiments, although not required for expression, the constructs disclosed herein may also include transcriptional or translational control sequences, such as promoters, enhancers, insulators, internal ribosome entry sites, peptide coding sequences, and / or polyadenylation signals.
[0141] In some embodiments, a construct comprising a coding sequence for a polypeptide of interest may include one or more of the following modifications: codon optimization (e.g., to human codons) and / or the addition of one or more glycosylation sites. See, e.g., McIntosh et al. (2013) Blood(17):3335-44.
[0142] In some embodiments, the construct may be inserted such that its expression is driven by an endogenous promoter at the insertion site (e.g., the endogenous albumin promoter when the donor is integrated into the albumin locus of the host cell). In such cases, the transgene may lack the control elements (e.g., promoter and / or enhancer) that drive its expression (e.g., a promoterless construct). Nevertheless, it will be apparent that in other cases, the construct may include a promoter and / or enhancer, e.g., a constitutive promoter that drives expression of a functional protein upon integration, or an inducible or tissue-specific (e.g., liver- or platelet-specific) promoter. The construct may include a heterologous protein-encoding sequence downstream of a signal sequence encoding a signal peptide, e.g., the albumin signal peptide, a signal peptide from a hepatocyte-secreted protein, and may be operably linked to the signal sequence encoding the signal peptide, e.g., the albumin signal peptide, a signal peptide from a hepatocyte-secreted protein. The construct may include a heterologous protein-encoding sequence downstream of a signal sequence encoding a signal peptide from a heterologous protein, and may be operably linked to the signal sequence encoding the signal peptide from the heterologous protein. In some embodiments, the nucleic acid construct functions in homology-independent insertion of a nucleic acid encoding a transgene protein. In some embodiments, the nucleic acid construct functions in non-dividing cells, for example, cells in which NHEJ, rather than HR, is the primary mechanism by which double-stranded DNA breaks are repaired. The nucleic acid can be a homology-independent donor construct.
[0143] The construct may be a bidirectional nucleic acid construct comprising at least two nucleic acid segments, one segment (first segment) comprising a coding sequence encoding an agent of interest (the coding sequence may be referred to herein as a "transgene" or first transgene), and the other segment (second segment) comprising a sequence whose sequence complement encodes the agent of interest, or a second transgene. In some embodiments, the coding sequence encodes a therapeutic agent, such as a polypeptide, functional RNA, or enhancer. The at least two segments can encode the same or different polypeptides or the same or different agents. In some embodiments, the bidirectional constructs disclosed herein comprise at least two nucleic acid segments, one segment (first segment) comprising a coding sequence encoding a polypeptide of interest, and the other segment (second segment) comprising a sequence whose sequence complement encodes the polypeptide of interest. When used in combination with the gene editing systems described herein, the bidirectional nature of the nucleic acid construct allows for the construct to be inserted in either orientation within the target insertion site (rather than being limited to insertion in one direction), and as exemplified herein, allows for expression of a polypeptide of interest from either a) the coding sequence of one segment (e.g., the left-hand segment encoding "human F9" in the ssAAV construct at the top left of Figure 1 ) or 2) the complement of another segment (e.g., the complement of the right-hand segment encoding "human F9" shown upside down in the ssAAV construct at the top left of Figure 1 ), thereby improving insertion and expression efficiency. Targeted cleavage by the gene editing system can facilitate construct integration and / or transgene expression. Various known gene editing systems can be used in the practice of the present disclosure, including site-specific DNA cleavage systems, including, for example, CRISPR / Cas systems, zinc finger nuclease (ZFN) systems, or transcription activator-like effector nuclease (TALEN) systems.
[0144] In some embodiments, the bidirectional nucleic acid construct does not comprise a promoter driving expression of the agent or polypeptide. For example, expression of the polypeptide is driven by a host cell promoter (e.g., the endogenous albumin promoter when the transgene is integrated into the albumin locus of the host cell). In some embodiments, the bidirectional nucleic acid construct comprises a first segment and a second segment, each having a splice acceptor upstream of the transgene. In certain embodiments, the splice acceptor is compatible with the splice donor sequence of a safe harbor site in the host cell, e.g., the splice donor in intron 1 of the human albumin gene.
[0145] In some embodiments, a bidirectional nucleic acid construct comprises a first segment comprising a coding sequence for a polypeptide and a second segment comprising the reverse complement of the coding sequence for the polypeptide. The same is true for non-polypeptide drugs. Thus, the coding sequence in the first segment is capable of expressing a polypeptide, and the complement of the reverse complement in the second segment is also capable of expressing a polypeptide. As used herein, when referring to a second segment comprising a reverse complement sequence, "coding sequence" refers to the complementary (coding) strand of the second segment (i.e., the complementary coding sequence of the reverse complement sequence in the second segment).
[0146] In some embodiments, the coding sequence encoding polypeptide A in the first segment is less than 100% complementary to the reverse complement of the coding sequence also encoding polypeptide A. That is, in some embodiments, the first segment comprises coding sequence (1) for polypeptide A, and the second segment is the reverse complement of coding sequence (2) for polypeptide A, where coding sequence (1) is not identical to coding sequence (2). For example, coding sequence (1) and / or coding sequence (2) encoding polypeptide A may utilize different codons. In some embodiments, coding sequences (1) and (2) may be codon-optimized such that the reverse complements of coding sequences (1) and (2) have 100% or less than 100% complementarity. In some embodiments, the coding sequence of the second segment encodes a polypeptide using one or more alternative codons for one or more amino acids of the same polypeptide encoded by the coding sequence in the first segment. As used herein, "alternative codons" refers to variations in codon usage for a given amino acid, which may or may not be preferred or optimized codons for a given expression system (codon optimization). Preferred codon usage, or codons that are well tolerated in a given expression system, is known in the art.
[0147] In some embodiments, the second segment comprises a reverse complement sequence that employs different codon usage than the coding sequence of the first segment to reduce hairpin formation. Such a reverse complement forms fewer base pairs than all nucleotides of the coding sequence in the first segment, but still optionally encodes the same polypeptide. In such cases, the coding sequence of the first segment, e.g., of polypeptide A, can be homologous, but not identical, to the coding sequence of the second half of the bidirectional construct, e.g., of polypeptide A. In some embodiments, the second segment comprises a reverse complement sequence that is not substantially complementary (e.g., 70% or less complementary) to the coding sequence in the first segment. In some embodiments, the second segment comprises a reverse complement sequence that is highly complementary (e.g., at least 90% complementary) to the coding sequence in the first segment. In some embodiments, the second segment comprises a reverse complement sequence having at least about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, or about 99% complementarity to the coding sequence in the first segment.
[0148] In some embodiments, the second segment comprises a reverse complement sequence that is 100% complementary to the coding sequence in the first segment. That is, the sequence in the second segment is the exact reverse complement of the coding sequence in the first segment. For example, the first segment comprises the hypothetical sequence 5'CTGGACCGA3' (SEQ ID NO:500), and the second segment comprises the reverse complement of SEQ ID NO:1, i.e., 5'TCGGTCCAG3' (SEQ ID NO:502).
[0149] In some embodiments, the bidirectional nucleic acid construct comprises a first segment comprising a coding sequence for a polypeptide or agent (e.g., a first polypeptide) and a second segment comprising the reverse complement of the coding sequence for a polypeptide or agent (e.g., a second polypeptide). In some embodiments, the first polypeptide and the second polypeptide are the same, as described above. In some embodiments, the first therapeutic agent and the second therapeutic agent are the same, as described above. In some embodiments, the first polypeptide and the second polypeptide are different. In some embodiments, the first therapeutic agent and the second therapeutic agent are different. For example, the first polypeptide is polypeptide A and the second polypeptide is polypeptide B. As a further example, the first polypeptide is polypeptide A and the second polypeptide is a variant of polypeptide A (e.g., a fragment (such as a functional fragment), a mutant, a fusion (including the addition of as little as one amino acid at the end of the polypeptide), or a combination thereof). The coding sequence encoding the polypeptide may optionally include one or more additional sequences, such as a sequence encoding an amino- or carboxy-terminal amino acid sequence such as a signal sequence, a tag sequence (e.g., HiBit), or a heterologous functional sequence (e.g., a nuclear localization sequence (NLS) or a self-cleaving peptide) linked to the polypeptide. The coding sequence encoding the polypeptide may optionally include a sequence encoding one or more amino-terminal signal peptide sequences. Each of these additional sequences may be the same or different in the first and second segments of the construct.
[0150] The bidirectional constructs described herein can be used to express any polypeptide according to the methods disclosed herein. In some embodiments, the polypeptide is a secreted polypeptide. In some embodiments, the polypeptide is one whose function normally acts (e.g., is functionally active) as a secreted polypeptide. As used herein, a "secreted polypeptide" refers to a protein that is secreted by a cell and / or is functionally active as a soluble extracellular protein.
[0151] In some embodiments, the polypeptide is an intracellular polypeptide. In some embodiments, the polypeptide is one whose function normally operates within a cell (e.g., is functionally active). As used herein, "intracellular polypeptide" refers to a protein that is not secreted by a cell, including soluble cell solute polypeptides. In some embodiments, the polypeptide is a wild-type polypeptide.
[0152] In some embodiments, the polypeptide is a liver protein or a variant thereof. As used herein, a "liver protein" is, for example, a protein that is endogenously produced in the liver and / or functionally active in the liver. In some embodiments, the liver protein is a circulating protein produced by the liver or a variant thereof. In some embodiments, the liver protein is a protein or a variant thereof that is functionally active in the liver. In some embodiments, the liver protein exhibits elevated expression in the liver compared to one or more other tissue types. In some embodiments, the polypeptide is a non-liver protein.
[0153] In some embodiments, the bidirectional nucleic acid construct is linear. For example, the first and second segments are connected in a linear fashion through a linker sequence. In some embodiments, the 5' end of the second segment containing the reverse complement sequence is linked to the 3' end of the first segment. In some embodiments, the 5' end of the first segment is linked to the 3' end of the second segment containing the reverse complement sequence. In some embodiments, the linker sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 500, 1000, 1500, 2000, or more nucleotides in length. As will be appreciated by one of skill in the art, other structural elements in addition to or in place of a linker sequence may be inserted between the first and second segments.
[0154] The constructs disclosed herein can be modified to include any suitable structural features necessary for any particular use and / or that confer one or more desired functions. In some embodiments, the bidirectional nucleic acid constructs disclosed herein do not include homology arms. In some embodiments, the constructs, e.g., bidirectional nucleic acid constructs, can be inserted into a genomic locus by non-homologous end joining (NHEJ). In some embodiments, the constructs disclosed herein are homology-independent donor constructs. In some embodiments, due in part to the bidirectional function of the nucleic acid construct, the bidirectional construct can be inserted into a genomic locus in any of the orientations described herein to allow for efficient insertion and / or expression of a polypeptide of interest.
[0155] In some embodiments, the compositions described herein contain one or more internal ribosome entry sites (IRES). IRESs, first identified as a characteristic of picornavirus RNA, play an important role in initiating protein synthesis in the absence of a 5' cap structure. An IRES can act as the sole ribosome binding site or as one of multiple ribosome binding sites for a polynucleotide. Constructs containing more than one functional ribosome binding site can encode several peptides or polypeptides ("multicistronic nucleic acid molecules") that are independently translated by ribosomes. Alternatively, constructs can contain an IRES to express heterologous proteins that are not fused to an endogenous polypeptide (i.e., an albumin signal peptide). Examples of IRES sequences that can be utilized include, but are not limited to, those derived from picornaviruses (e.g., FMDV), vermin viruses (CFFV), polioviruses (PV), encephalomyocarditis viruses (ECMV), foot-and-mouth disease viruses (FMDV), hepatitis C viruses (HCV), classical swine fever viruses (CSFV), murine leukemia viruses (MLV), simian immunodeficiency viruses (SIV), or cricket paralysis viruses (CrPV).
[0156] In some embodiments, the nucleic acid construct includes a sequence encoding a self-cleaving peptide, such as a 2A sequence or a 2A-like sequence. The self-cleaving peptide can be a P2A peptide, a T2A peptide, or the like. In some embodiments, the self-cleaving peptide is located upstream of the polypeptide of interest. In one embodiment, the sequence encoding the 2A peptide can be used to separate the coding regions of two or more polypeptides of interest. In another embodiment, this sequence can be used to separate the coding sequences from the construct and separate the coding sequences from the endogenous locus (i.e., the endogenous albumin signal sequence). As a non-limiting example, the sequence encoding the 2A peptide can be between region A and region B (A-2A-B). The presence of the 2A peptide results in cleavage of a single long protein into protein A, protein B, and the 2A peptide. Protein A and protein B can be the same or different polypeptides of interest.
[0157] In some embodiments, one or both of the first and second segments comprises a polyadenylation tail sequence and / or a polyadenylation signal sequence downstream of the open reading frame, hi some embodiments, the polyadenylation tail sequence is encoded, e.g., as a "poly-A" stretch, at the 3' end of the first and / or second segment.
[0158] III. Delivery method The guide RNAs disclosed herein can be delivered to a host cell or subject in vivo or ex vivo using a variety of known and suitable methods available in the art. The guide RNAs can be delivered (individually or in combination) together with a construct comprising an RNA-guided DNA-binding agent, such as a nucleic acid encoding Cas or Cas9 (e.g., Cas9 or a nucleic acid encoding Cas9), as described herein, and a sequence encoding a heterologous gene to be inserted into the cleavage site created by the guide RNA of the present disclosure.
[0159] Conventional viral and non-viral gene delivery methods can be used to introduce the guide RNAs disclosed herein, as well as the RNA-guided DNA binders and donor constructs into cells (e.g., mammalian cells) and target tissues. As further provided herein, non-viral vector delivery systems, such as non-viral vectors and plasmid vectors, include naked nucleic acids and nucleic acids complexed with delivery vehicles such as liposomes, lipid nanoparticles (LNPs), or poloxamers. Viral vector delivery systems include DNA and RNA viruses.
[0160] Methods and compositions for non-viral delivery of nucleic acids include electroporation, lipofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, LNPs, polycation or lipid:nucleic acid complexes, naked nucleic acids (e.g., naked DNA / RNA), artificial virions, and drug-enhanced uptake of DNA. Sonoporation, for example, using the Sonitron 2000 system (Rich-Mar), can also be used to deliver nucleic acids.
[0161] Further exemplary nucleic acid delivery systems include those provided by AmaxaBiosystems (Cologne, Germany), Maxcyte, Inc. (Rockville, Md.), BTX Molecular Delivery Systems (Holliston, Ma.), and Copernicus Therapeutics Inc. (see, e.g., U.S. Patent No. 6,008,336). Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known in the art and is described herein.
[0162] Various delivery systems (e.g., vectors, liposomes, LNPs) containing guide RNAs, RNA-guided DNA binders, and donor constructs can also be administered alone or in combination to living organisms for delivery to cells in vivo, or to cells or cell cultures ex vivo. Administration can be by any of the routes commonly used to introduce molecules into blood, fluids, or ultimately into contact with cells, including but not limited to, injection, infusion, topical application, and electroporation. Suitable methods for administering such nucleic acids are available and well known to those skilled in the art.
[0163] In some embodiments, the guide RNA compositions described herein, alone or encoded in one or more vectors, are formulated in or administered via lipid nanoparticles; see, e.g., PCT / US2017 / 024973, the contents of which are incorporated herein by reference in their entirety. Any lipid nanoparticle (LNP) formulation known to those skilled in the art to be capable of delivering nucleotides to a subject can be utilized with the guide RNAs described herein, as well as mRNA encoding an RNA-guided DNA-binding agent such as Cas or Cas9, or the RNA-guided DNA-binding agent itself, such as Cas or Cas9 protein.
[0164] In some embodiments, the guide RNAs disclosed herein can be delivered to host cells (in vitro or in vivo) via LNPs. In some embodiments, the gRNA / LNPs are also associated with an mRNA encoding an RNA-guided RNA-binding agent, such as Cas9, or an RNA-guided DNA-binding agent, such as Cas9. In some embodiments, the gRNA / LNPs are also associated with a donor construct as described herein.
[0165] In some embodiments, the present disclosure includes methods of delivering a gRNA disclosed herein to a cell in vitro, wherein the gRNA is delivered via LNP. In some embodiments, the gRNA is delivered by a non-LNP means, such as via an AAV system, and the RNA-guided DNA-binding agent (e.g., Cas9) or mRNA encoding the RNA-guided DNA-binding agent (e.g., Cas9), and / or donor construct is delivered by LNP.
[0166] In some embodiments, the present disclosure provides compositions and LNPs comprising any one of the gRNAs disclosed herein. In some embodiments, the composition further comprises Cas9 or an mRNA encoding Cas9, or another RNA-guided DNA-binding agent described herein. In some embodiments, the composition further comprises a donor construct as described herein.
[0167] In some embodiments, the LNPs comprise a biodegradable cationic lipid, ie, (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate), also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate), or another ionic lipid. See, e.g., the lipids of PCT / US2018 / 053559 (filed September 28, 2018), WO / 2017 / 173054, WO2015 / 095340, and WO2014 / 136086, and the references provided therein. In some embodiments, in the context of LNP lipids, the terms cationic and ionizable are interchangeable, e.g., an ionizable lipid is cationic depending on the pH.
[0168] In some embodiments, any of the guide RNAs, RNA-guided DNA binders, and / or donor constructs (e.g., bidirectional constructs) described herein, alone or in combination, naked or as part of a vector, are formulated in or administered via lipid nanoparticles; see, e.g., WO / 2017 / 173054, the contents of which are incorporated herein by reference in their entirety.
[0169] Electroporation is also a well-known means for delivering cargo, and any electroporation methodology can be used to deliver any of the gRNAs disclosed herein. In some embodiments, electroporation can be used to deliver any of the gRNAs disclosed herein, optionally along with an RNA-guided DNA-binding agent such as Cas9 or an mRNA encoding an RNA-guided DNA-binding agent such as Cas9, by the same or different means. In some embodiments, electroporation can be used to deliver any one of the gRNAs disclosed herein and a donor construct as disclosed herein.
[0170] In certain embodiments, the present disclosure provides DNA or RNA vectors encoding any of the guide RNAs comprising any one or more of the guide sequences described herein. In certain embodiments, the present invention includes DNA or RNA vectors encoding any one or more of the guide sequences described herein. In some embodiments, in addition to the guide RNA sequence, the vector further comprises a nucleic acid that does not encode the guide RNA. The nucleic acid that does not encode the guide RNA includes, but is not limited to, a promoter, an enhancer, a regulatory sequence, a nucleic acid encoding an RNA-guided DNA-binding agent, which may be a nuclease such as Cas9, and a donor construct comprising a heterologous gene. In some embodiments, the vector comprises one or more nucleotide sequence(s) encoding a crRNA, a trRNA, or a crRNA and a trRNA, as disclosed herein.
[0171] In some embodiments, the vector comprises one or more nucleotide sequence(s) encoding an sgRNA and an mRNA encoding an RNA-guided DNA binder, which may be a Cas protein such as Cas9 or Cpf1. In some embodiments, the vector comprises one or more nucleotide sequence(s) encoding a crRNA, a trRNA, and an mRNA encoding an RNA-guided DNA binder, which may be a Cas protein such as Cas9 or Cpf1. In one embodiment, the Cas9 is derived from Streptococcus pyogenes (i.e., Spy Cas9). In some embodiments, the nucleotide sequence encoding the crRNA, trRNA, or crRNA and trRNA (which may be an sgRNA) comprises or consists of a guide sequence flanked in whole or in part by repeat sequences from a naturally occurring CRISPR / Cas system. A nucleic acid comprising or consisting of a crRNA, a trRNA, or a crRNA and a trRNA can comprise a vector sequence, which comprises or consists of a nucleic acid sequence that is not naturally found with the crRNA, the trRNA, or the crRNA and a trRNA.
[0172] In some embodiments, the crRNA and trRNA are encoded by non-contiguous nucleic acids within a single vector. In other embodiments, the crRNA and trRNA may be encoded by contiguous nucleic acids. In some embodiments, the crRNA and trRNA are encoded by opposite strands of a single nucleic acid. In other embodiments, the crRNA and trRNA are encoded by the same strand of a single nucleic acid.
[0173] In some embodiments, the vector may be circular. In other embodiments, the vector may be linear. In some embodiments, the vector may be delivered via a lipid nanoparticle, a liposome, a non-lipid nanoparticle, or a viral capsid. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0174] In some embodiments, the vector may be a viral vector. In some embodiments, the viral vector may be genetically modified from its wild-type counterpart. For example, a viral vector may contain one or more nucleotide insertions, deletions, or substitutions to facilitate cloning or to alter one or more characteristics of the vector. Such characteristics may include packaging ability, transduction efficiency, immunogenicity, genome integration, replication, transcription, and translation. In some embodiments, a portion of the viral genome may be deleted to allow the virus to package foreign sequences having a larger size. In some embodiments, the viral vector may have enhanced transduction efficiency. In some embodiments, the immune response elicited by the virus in the host may be reduced. In some embodiments, a viral gene that facilitates integration of viral sequences into the host genome (e.g., integrase, etc.) may be mutated to render the virus non-integrating. In some embodiments, the viral vector may be replication-deficient. In some embodiments, the viral vector may contain exogenous transcriptional or translational control sequences that drive expression of coding sequences on the vector. In some embodiments, the virus may be helper-dependent. For example, one or more helper viruses may be required to provide viral components (e.g., viral proteins, etc.) needed to amplify and package the vector into viral particles. In such cases, one or more helper components, including one or more vectors encoding viral components, may be introduced into a host cell along with the vector systems described herein. In some embodiments, the virus may be helper-free. For example, the virus may be capable of amplifying and packaging a vector without a helper virus. In some embodiments, the vector systems described herein may also encode viral components necessary for viral amplification and packaging.
[0175] Non-limiting exemplary viral vectors include adeno-associated viral (AAV) vectors, lentiviral vectors, adenoviral vectors, helper-dependent adenoviral vectors (HDAd), herpes simplex viral (HSV-1) vectors, bacteriophage T4, baculoviral vectors, and retroviral vectors. In some embodiments, the viral vector may be an AAV vector. In some embodiments, the viral vector may be a lentiviral vector.
[0176] In some embodiments, "AAV" refers to all serotypes, subtypes, and naturally occurring AAVs, as well as recombinant AAVs. "AAV" can be used to refer to the virus itself or its derivatives. The term "AAV" includes AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV. The genomic sequences of various serotypes of AAV, as well as the sequences of natural terminal repeats (TRs), Rep proteins, and capsid subunits, are known in the art. Such sequences can be found in the literature or public databases such as GenBank. As used herein, "AAV vector" refers to an AAV vector that contains heterologous sequences that are not of AAV origin (i.e., nucleic acid sequences that are heterologous to AAV), and typically contains a sequence encoding a heterologous polypeptide of interest. The construct may contain AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV capsid sequences. Generally, the heterologous nucleic acid sequence (transgene) is flanked by at least one, and generally two, AAV inverted terminal repeats (ITRs). AAV vectors may be either single-stranded (ssAAV) or self-complementary (scAAV).
[0177] In some embodiments, the lentivirus may be non-integrating. In some embodiments, the viral vector may be an adenovirus vector. In some embodiments, the adenovirus may be a high-cloning-capacity, or "gutless," adenovirus, in which all coding viral regions are removed from the 5' and 3' inverted terminal repeats (ITRs) and the packaging signal ("I") is deleted from the virus to increase its packaging capacity. In yet other embodiments, the viral vector may be an HSV-1 vector. In some embodiments, HSV-1-based vectors are helper-dependent, while in other embodiments, they are vector-independent. For example, amplicon vectors that retain only the packaging sequence require a helper virus with structural components for packaging, while a 30 kb-deleted HSV-1 vector that removes non-essential viral functions does not require a helper virus. In further embodiments, the viral vector may be bacteriophage T4. In some embodiments, bacteriophage T4 can package any linear or circular DNA or RNA molecule when the viral head is empty. In further embodiments, the viral vector may be a baculovirus vector. In yet further embodiments, the viral vector may be a retrovirus vector. In embodiments using AAV or lentivirus vectors with smaller cloning capacity, it may be necessary to use more than one vector to deliver all components of the vector system as disclosed herein. For example, one AAV vector may contain a sequence encoding an RNA-guided DNA binding agent such as Cas protein (e.g., Cas9), while a second AAV vector may contain one or more guide sequences.
[0178] In some embodiments, the vector may be capable of driving expression of one or more coding sequences in a cell. In some embodiments, the cell may be a eukaryotic cell, such as a yeast, plant, insect, or mammalian cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the eukaryotic cell may be a rodent cell. In some embodiments, the eukaryotic cell may be a human cell. Suitable promoters for driving expression in different types of cells are known in the art. In some embodiments, the promoter may be wild-type. In other embodiments, the promoter may be modified for more efficient or effective expression. In still other embodiments, the promoter may be shortened but still maintain its function. For example, the promoter may have a normal or reduced size suitable for proper packaging of the vector into a virus.
[0179] In some embodiments, the vector may include a nucleotide sequence encoding an RNA-guided DNA-binding agent, such as a Cas protein (e.g., Cas9) described herein. In some embodiments, the nuclease encoded by the vector may be a Cas protein. In some embodiments, the vector system may include one copy of the nucleotide sequence encoding the nuclease. In other embodiments, the vector system may include more than one copy of the nucleotide sequence encoding the nuclease. In some embodiments, the nucleotide sequence encoding the nuclease may be operably linked to at least one transcriptional or translational control sequence. In some embodiments, the nucleotide sequence encoding the nuclease may be operably linked to at least one promoter.
[0180] In some embodiments, the vector may comprise any one or more of the constructs comprising a heterologous gene described herein. In some embodiments, the heterologous gene may be operably linked to at least one transcriptional or translational regulatory sequence. In some embodiments, the heterologous gene may be operably linked to at least one promoter. In some embodiments, the heterologous gene is not linked to a promoter that drives expression of the heterologous gene.
[0181] In some embodiments, the promoter may be constitutive, inducible, or tissue-specific. In some embodiments, the promoter may be a constitutive promoter. Non-limiting exemplary constitutive promoters include the cytomegalovirus immediate early promoter (CMV), the simian virus (SV40) promoter, the adenovirus major late (MLP) promoter, the Rous sarcoma virus (RSV) promoter, the mouse mammary tumor virus (MMTV) promoter, the phosphoglycerin kinase (PGK) promoter, the elongation factor alpha (EF1a) promoter, the ubiquitin promoter, the actin promoter, the tubulin promoter, the immunoglobulin promoter, a functional fragment thereof, or a combination of any of the foregoing. In some embodiments, the promoter may be a CMV promoter. In some embodiments, the promoter may be a truncated CMV promoter. In some embodiments, the promoter may be the EF1a promoter. In some embodiments, the promoter may be an inducible promoter. Non-limiting exemplary inducible promoters include those inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter may be one that has a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0182] In some embodiments, the promoter may be a tissue-specific promoter, for example, a promoter specific for expression in the liver.
[0183] The vector may further comprise a nucleotide sequence encoding a guide RNA as described herein. In some embodiments, the vector comprises one copy of the guide RNA. In other embodiments, the vector comprises more than one copy of the guide RNA. In embodiments with more than one guide RNA, the guide RNAs may be non-identical, targeting different target sequences, or identical in that they target the same target sequence. In some embodiments, when the vector comprises more than one guide RNA, each guide RNA may have other distinct properties, such as activity or stability in a complex with an RNA-guided DNA nuclease, such as a Cas RNP complex. In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to at least one transcriptional or translational regulatory sequence, such as a promoter, a 3'UTR, or a 5'UTR. In one embodiment, the promoter is a tRNA promoter, e.g., a tRNA Lys3, or tRNA chimeras. See Mefferd et al., RNA. 2015 21:1683-9; Scherer et al., Nucleic Acids Res. 2007 35:2620-2628. In some embodiments, the promoter can be recognized by RNA polymerase III (Pol III). Non-limiting examples of Pol III promoters include the U6 and H1 promoters. In some embodiments, the nucleotide sequence encoding the guide RNA can be operably linked to a mouse or human U6 promoter. In other embodiments, the nucleotide sequence encoding the guide RNA can be operably linked to a mouse or human H1 promoter. In embodiments with more than one guide RNA, the promoters used to drive expression can be the same or different. In some embodiments, the nucleotides encoding the crRNA of the guide RNA and the nucleotides encoding the trRNA of the guide RNA can be provided on the same vector. In some embodiments, the nucleotides encoding the crRNA and the trRNA can be driven by the same promoter. In some embodiments, the crRNA and the trRNA can be transcribed into a single transcript. For example, crRNA and trRNA can be processed from a single transcript to form a dual-molecule guide RNA. Alternatively, crRNA and trRNA can be transcribed into a single-molecule guide RNA (sgRNA). In other embodiments, crRNA and trRNA can be driven by their corresponding promoters on the same vector. In yet other embodiments, crRNA and trRNA can be encoded by different vectors.
[0184] In some embodiments, the nucleotide sequence encoding the guide RNA may be located on the same vector as the nucleotide sequence encoding the RNA-guided DNA binder, such as a Cas protein. In some embodiments, the expression of the guide RNA and the RNA-guided DNA binder, such as a Cas protein, may be driven by their respective promoters. In some embodiments, the expression of the guide RNA may be driven by the same promoter that drives the expression of the RNA-guided DNA binder, such as a Cas protein. In some embodiments, the guide RNA and the RNA-guided DNA binder, such as a Cas protein transcript, may be included within a single transcript. For example, the guide RNA may be within the untranslated region (UTR) of the RNA-guided DNA binder, such as a Cas protein transcript. In some embodiments, the guide RNA may be within the 5' UTR of the transcript. In other embodiments, the guide RNA may be within the 3' UTR of the transcript. In some embodiments, the intracellular half-life of a transcript can be reduced by including the guide RNA within its 3' UTR, thereby shortening the length of the 3' UTR. In further embodiments, the guide RNA may be within an intron of the transcript. In some embodiments, suitable splice sites may be added in the intron within which the guide RNA is located so that the guide RNA is properly spliced out of the transcript.
[0185] In some embodiments, the composition comprises a vector system. In some embodiments, the vector system may comprise a single vector. In other embodiments, the vector system may comprise two vectors. In further embodiments, the vector system may comprise three vectors. When different guide RNAs are used for multiplexing or when multiple copies of a guide RNA are used, the vector system may comprise more than three vectors. In some embodiments, the vector system may further comprise a donor construct as described herein. In some embodiments, the vector system may further comprise a nucleic acid encoding a nuclease. In some embodiments, the vector system may further comprise a nucleic acid encoding a guide RNA and / or a nucleic acid encoding an RNA-guided DNA binder (which may be a Cas9 protein, such as Cas9). In some embodiments, the nucleic acid encoding the guide RNA and / or the nucleic acid encoding the RNA-guided DNA binder or nuclease are each or both on a vector separate from the vector comprising the donor construct disclosed herein. In any of the embodiments, the vector system may comprise other sequences, including but not limited to, promoters, enhancers, and regulatory sequences, as described herein. In some embodiments, the promoter in the vector system does not drive expression of a transgene in the donor construct (e.g., a bidirectional construct). In some embodiments, the vector system comprises one or more nucleotide sequence(s) encoding a crRNA, a trRNA, or a crRNA and a trRNA. In some embodiments, the vector system comprises one or more nucleotide sequence(s) encoding an sgRNA and an mRNA encoding an RNA-guided DNA binder, and the RNA-guided DNA nuclease may be a Cas nuclease (e.g., Cas9). In some embodiments, the vector system comprises one or more nucleotide sequence(s) encoding a crRNA, a trRNA, and an mRNA encoding an RNA-guided DNA binder, and the RNA-guided DNA nuclease may be a Cas nuclease, such as Cas9. In some embodiments, the Cas9 is derived from Streptococcus pyogenes (i.e., Spy Cas9).In some embodiments, the nucleotide sequence encoding the crRNA, trRNA, or crRNA and trRNA (which may be an sgRNA) comprises or consists of a guide sequence flanked by all or a portion of repeat sequences from a naturally occurring CRISPR / Cas system. The vector system may comprise a nucleic acid that comprises or consists of the crRNA, trRNA, or crRNA and trRNA, or the vector system comprises or consists of a nucleic acid that is not found in nature with the crRNA, trRNA, or crRNA and trRNA.
[0186] In some embodiments, the vector system may include an inducible promoter to initiate expression only after delivery to the target cell. Non-limiting exemplary inducible promoters include those that are inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter may have a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0187] In a further embodiment, the vector system may contain a tissue-specific promoter to initiate expression only after delivery to a specific tissue.
[0188] The vector or vector system can be delivered via liposomes, nanoparticles, exosomes, or microvesicles. The vector can also be delivered via lipid nanoparticles (LNPs). Donor constructs containing one or more guide RNAs, RNA-binding DNA binders (e.g., mRNA), or sequences encoding heterologous proteins can be delivered via liposomes, nanoparticles, exosomes, or microvesicles, individually or in any combination. Donor constructs containing one or more guide RNAs, RNA-binding DNA binders (e.g., mRNA), or sequences encoding heterologous proteins can be delivered via LNPs, individually or in any combination. Any of the LNPs and LNP formulations described herein are suitable for delivering constructs containing guides, Cas nucleases (or mRNAs encoding Cas nucleases), combinations thereof, and / or heterologous genes. Some embodiments include LNP compositions containing an RNA component and a lipid component, where the lipid component comprises an amine lipid, such as a biodegradable ionic lipid, and the RNA component comprises a guide RNA and / or mRNA encoding a Cas nuclease. In some examples, the lipid component includes a biodegradable, ionic lipid, cholesterol, DSPC, and PEG-DMG.
[0189] It will be apparent that the guide RNA, RNA-guided DNA binder (e.g., Cas nuclease or nucleic acid encoding a Cas nuclease), and donor construct disclosed herein can be delivered using the same or different systems. For example, the guide RNA, Cas nuclease, and construct can be carried by the same vector (e.g., AAV). Alternatively, the Cas nuclease (as a protein or mRNA) and / or gRNA can be carried by a plasmid or LNP, while the construct can be carried by a vector. Furthermore, the different delivery systems can be administered by the same or different routes.
[0190] In some embodiments, the method comprises administering a guide RNA and an RNA-guided DNA binder (such as an mRNA encoding a Cas9 nuclease) in an LNP. In further embodiments, the method comprises administering an AAV nucleic acid construct encoding a transgene protein, such as a bidirectional construct. CRISPR / Cas9 LNPs comprising a guide RNA and an mRNA encoding Cas9 can be administered intravenously. An AAV donor construct can be administered intravenously.
[0191] The different delivery systems can be delivered in vitro or in vivo simultaneously or in any sequential order. In some embodiments, the donor construct, guide RNA, and Cas nuclease can be delivered in vitro or in vivo simultaneously, e.g., in one vector, two vectors, individual vectors, one LNP, two LNPs, individual LNPs, or a combination thereof. In some embodiments, the donor construct can be delivered in vivo or in vitro as a vector and / or in association with LNPs (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 days or more) before the guide RNA and / or Cas nuclease are delivered as a vector and / or in association with LNPs alone or together as ribonucleoproteins (RNPs). In some embodiments, the donor construct can be delivered in multiple doses, e.g., daily, every 2 days, every 3 days, every 4 days, weekly, every 2 weeks, every 3 weeks, or every 4 weeks. In some embodiments, the donor construct may be delivered at weekly intervals, such as, for example, at week 1, week 2, and week 3. As a further example, the guide RNA and Cas nuclease may be delivered in vivo or in vitro as a vector and / or alone or together as a ribonucleoprotein (RNP) in association with an LNP (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 days or more) prior to delivery of the construct as a vector and / or in association with an LNP. In some embodiments, the albumin guide RNA may be delivered in multiple doses, for example, daily, every two days, every three days, every four days, weekly, every two weeks, every three weeks, or every four weeks. In some embodiments, the albumin guide RNA may be delivered at weekly intervals, such as, for example, at week 1, week 2, and week 3. In some embodiments, the Cas nuclease can be delivered in multiple doses, for example, every day, every two days, every three days, every four days, every week, every two weeks, every three weeks, or every four weeks. In some embodiments, the Cas nuclease can be delivered at weekly intervals, for example, at week 1, week 2, and week 3.
[0192] IV.How to use The gRNAs and related methods and compositions disclosed herein are useful for efficiently inserting heterologous (exogenous) genes into intron 1 of the human albumin locus of a host cell. In some embodiments, the present disclosure provides a method for inserting a heterologous gene into intron 1 of the human albumin locus of a host cell, comprising administering to the host cell (in vivo or in vitro) a donor construct comprising a guide RNA described herein (any one of SEQ ID NOS: 2-33), an RNA-guided DNA-binding agent (e.g., a Cas nuclease described herein), and a sequence encoding the heterologous polypeptide of interest.
[0193] The gRNAs and related methods and compositions disclosed herein are useful for expressing heterologous (exogenous) genes within intron 1 of the human albumin locus in a host cell. In some embodiments, the present disclosure provides a method for expressing a heterologous gene within intron 1 of the human albumin locus in a host cell, comprising administering to the host cell (in vivo or in vitro) a donor construct comprising a guide RNA described herein (any one of SEQ ID NOS: 2-33), an RNA-guided DNA-binding agent (e.g., a Cas nuclease described herein), and a sequence encoding the heterologous polypeptide of interest.
[0194] The gRNAs and related methods and compositions disclosed herein are useful for treating liver-related disorders in a subject, as described herein. In some embodiments, the present disclosure provides a method of treating a liver-related disorder, comprising administering to a host cell (in vivo or in vitro) a donor construct comprising a guide RNA (any one of SEQ ID NOS: 2-33) described herein, an RNA-guided DNA-binding agent (e.g., a Cas nuclease described herein), and a sequence encoding a polypeptide of interest.
[0195] The compositions and methods of the present disclosure are useful and applicable to a variety of host cells. In some embodiments, the host cell is a liver cell, a neuronal cell, or a muscle cell. In some embodiments, the host cell is any suitable non-dividing cell. As used herein, "non-dividing cell" refers to terminally differentiated, non-dividing cells, as well as quiescent cells that do not divide but retain the ability to re-enter cell division and proliferation. For example, liver cells retain the ability to divide (e.g., when injured or ablated), but typically do not divide. During mitotic cell division, homologous recombination is the mechanism by which the genome is protected and double-stranded breaks are repaired. In some embodiments, a "non-dividing" cell refers to a cell in which homologous recombination (HR) is not the primary mechanism by which double-stranded DNA breaks are repaired in the cell, e.g., compared to a control dividing cell. In some embodiments, a "non-dividing" cell refers to a cell in which non-homologous end joining (NHEJ) is the primary mechanism by which double-stranded DNA breaks are repaired in the cell, e.g., compared to a control dividing cell. Non-dividing cell types have been described in the literature, for example, by an active NHEJ double-strand DNA break repair mechanism. See, for example, Iyama, DNA Repair (Amst.) 2013, 12(8):620-636. In some embodiments, the host cell includes, but is not limited to, a liver cell, a muscle cell, or a neuronal cell. In some embodiments, the host cell is a liver cell, such as a mouse, cynomolgus monkey, or human hepatocyte. In some embodiments, the host cell is a muscle cell, such as a mouse, cynomolgus monkey, or human muscle cell. In some embodiments, provided herein is a host cell as described above, comprising a bidirectional construct disclosed herein. In some embodiments, the host cell expresses a transgene polypeptide encoded by a bidirectional construct disclosed herein. In some embodiments, provided herein is a host cell produced by the methods disclosed herein. In certain embodiments, the host cell is generated by administering or delivering to the host cell a bidirectional nucleic acid construct described herein and a gene editing system, such as a ZFN, TALEN, or CRISPR / Cas9 system.
[0196] In some embodiments, the method further comprises achieving a long-term effect, e.g., at least 1 month, 2 months, 6 months, 1 year, or 2 years. In some embodiments, the method further comprises achieving a therapeutic effect in a long-term and sustained manner, e.g., at least 1 month, 2 months, 6 months, 1 year, or 2 years. In some embodiments, the level of circulating Factor IX activity and / or levels stabilizes for at least 1 month, 2 months, 6 months, 1 year, or more. In some embodiments, steady-state activity and / or levels of FIX protein are achieved by at least 7 days, at least 14 days, or at least 28 days. In further embodiments, the method comprises maintaining Factor IX activity and / or levels for at least 1, 2, 4, or 6 months, or at least 1, 2, 3, 4, or 5 years after a single administration.
[0197] In further embodiments involving an insertion into the albumin locus, the individual's circulating albumin level is normal. The method may include maintaining the individual's circulating albumin level within ±5%, ±10%, ±15%, ±20%, or ±50% of the normal circulating albumin level. In certain embodiments, the individual's albumin level remains unchanged compared to the albumin level of an untreated individual by at least 4 weeks, 8 weeks, 12 weeks, or 20 weeks. In certain embodiments, the individual's albumin level temporarily decreases and then returns to normal levels. In particular, the method may include not detecting a significant change in plasma albumin levels.
[0198] In some embodiments, the present invention includes a method or use for modifying (e.g., creating a double-stranded break in) an albumin gene, such as the human albumin gene, comprising administering or delivering to a host cell or a population of host cells any one or more of the gRNA, donor construct (e.g., a bidirectional construct comprising a sequence encoding Factor IX), and RNA-guided DNA binding agent (e.g., a Cas nuclease) described herein. In some embodiments, the present invention includes a method or use for modifying (e.g., creating a double-stranded break in) an albumin intron 1 region, such as human albumin intron 1, comprising administering or delivering to a host cell or a population of host cells any one or more of the gRNA, donor construct (e.g., a bidirectional construct comprising a sequence encoding Factor IX), and RNA-guided DNA binding agent (e.g., a Cas nuclease) described herein. In some embodiments, the invention includes methods or uses for modifying (e.g., creating a double-strand break in) a human genomic locus, such as a safe harbor site, such as a liver cell or hepatocyte host cell, comprising administering or delivering to a host cell or host cell population any one or more of the gRNAs, donor constructs (e.g., bidirectional constructs comprising a sequence encoding Factor IX), and RNA-guided DNA binding agents (e.g., Cas nucleases) described herein. Insertion into a genomic locus, such as a safe harbor site, such as an albumin locus safe harbor site (e.g., intron 1), allows for overexpression of the Factor IX gene without significant deleterious effects on the host cell or cell population, such as hepatocytes or liver cells. In some embodiments, the invention includes methods or uses for modifying (e.g., creating a double-stranded break in) intron 1 of the human albumin locus, comprising administering or delivering to a host cell or population of host cells any one or more of the gRNAs, donor constructs (e.g., bidirectional constructs comprising sequences encoding Factor IX), and RNA-guided DNA-binding agents (e.g., Cas nucleases) described herein.In some embodiments, the guide RNA comprises a guide sequence comprising at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides that bind within intron 1 of the human albumin locus (SEQ ID NO: 1). In some embodiments, the guide RNA comprises at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, the guide RNA comprises a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, the guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33. In some embodiments, the guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34-97. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the donor construct is a bidirectional construct comprising a sequence encoding Factor IX. In some embodiments, the host cell is, for example, a liver cell. In a further embodiment, the liver cells are hepatocytes.
[0199] In some embodiments, the invention includes methods or uses for introducing Factor IX nucleic acids into a host cell or host cell population, comprising administering or delivering any one or more of the gRNA, donor construct (e.g., a bidirectional construct comprising a sequence encoding Factor IX), and RNA-guided DNA-binding agent (e.g., a Cas nuclease) described herein. In some embodiments, the guide RNA comprises a guide sequence comprising at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides capable of binding to a region within intron 1 of the human albumin locus (SEQ ID NO: 1). In some embodiments, the guide RNA comprises at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, the guide RNA comprises a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34-97.In some embodiments, the method is in vitro. In some embodiments, the method is in vivo. In some embodiments, the donor construct is a bidirectional construct comprising a sequence encoding Factor IX. In some embodiments, the host cell is a liver cell, or the host cell population is a liver cell, such as a hepatocyte.
[0200] In some embodiments, the invention includes a method or use of expressing Factor IX in a host cell or a population of host cells, comprising administering or delivering any one or more of the gRNA, donor construct (e.g., a bidirectional construct comprising a sequence encoding Factor IX), and RNA-guided DNA-binding agent (e.g., a Cas nuclease) described herein. In some embodiments, the guide RNA comprises a guide sequence comprising at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides capable of binding to a region within intron 1 of the human albumin locus (SEQ ID NO: 1). In some embodiments, the guide RNA comprises at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, the guide RNA comprises a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34-97.In some embodiments, the method is in vitro. In some embodiments, the method is in vivo. In some embodiments, the donor construct is a bidirectional construct comprising a sequence encoding Factor IX. In some embodiments, the host cell is a liver cell, or the host cell population is a liver cell, such as a hepatocyte.
[0201] In some embodiments, the present invention includes a method or use for treating hemophilia (e.g., hemophilia A or hemophilia B) comprising administering or delivering any one or more of the gRNA, donor construct (e.g., a bidirectional construct comprising a sequence encoding factor IX), and RNA-guided DNA binder (e.g., a Cas nuclease) described herein to a subject in need thereof. In some embodiments, the guide RNA comprises a guide sequence comprising at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides capable of binding to a region within intron 1 of the human albumin locus (SEQ ID NO: 1). In some embodiments, the guide RNA comprises at least 15, 16, 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, the guide RNA comprises a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, 33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence that is at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2-33. In some embodiments, a guide RNA disclosed herein comprises a guide sequence selected from the group consisting of SEQ ID NOs: 34-97.In some embodiments, the donor construct is a bidirectional construct comprising a sequence encoding a heterologous polypeptide, hi some embodiments, the host cell is a liver cell, or the host cell population is a liver cell, such as a hepatocyte.
[0202] As used herein, "hemophilia" refers to a disorder caused by a missing or defective factor IX gene or polypeptide. Hemophilia also refers to a disorder caused by a missing or defective factor VIII gene or polypeptide. Disorders include inherited and / or acquired (e.g., caused by spontaneous mutations in genes) conditions, including hemophilia A and hemophilia B. Hemophilia A is caused by factor VIII deficiency. Hemophilia B is caused by factor IX deficiency. In some embodiments, a defective factor IX gene or polypeptide results in reduced plasma factor IX levels and / or reduced factor IX clotting activity. As used herein, hemophilia includes mild, moderate, and severe hemophilia. For example, individuals with less than about 1% active factor are classified as having severe hemophilia, individuals with about 1-5% active factor are classified as having moderate hemophilia, and individuals with mild hemophilia are classified as having about 5-40% of normal levels of activated clotting factor.
[0203] In some embodiments, the donor construct comprises a sequence encoding factor IX, and the factor IX sequence is wild-type factor IX. In some embodiments, the sequence encodes a variant of factor IX. For example, the variant may have increased clotting activity relative to wild-type factor IX. For example, the variant factor IX may include one or more mutations, such as an amino acid substitution at position R338 (e.g., R338L), relative to wild-type factor IX. In some embodiments, the sequence is 80%, 85%, 90%, 93%, 95%, 97%, or 99% identical to wild-type factor IX and encodes a factor IX variant that has at least 80%, 85%, 90%, 92%, 94%, 96%, 98%, 99%, 100%, or more of the activity of wild-type factor IX. In some embodiments, the sequence encodes a fragment of Factor IX, wherein the fragment has at least 80%, 85%, 90%, 92%, 94%, 96%, 98%, 99%, 100% or more activity compared to wild-type Factor IX.
[0204] In some embodiments, the donor construct comprises a sequence encoding a factor IX variant that activates coagulation in the absence of its cofactor, factor VIII (expression results in therapeutically relevant FVIII-mimetic activity). Such a factor IX variant can further maintain the activity of wild-type factor IX. For example, such a factor IX variant can include an amino acid substitution at positions L6, V181, K265, I383, E185, or a combination thereof relative to wild-type factor IX (e.g., relative to SEQ ID NO: 701). For example, such a factor IX variant can include an L6F mutation, a V181I mutation, a K265A mutation, an I383V mutation, an E185D mutation, or a combination thereof relative to wild-type factor IX.
[0205] The compositions and methods of the present disclosure are useful for the efficient insertion of a heterologous gene of interest and the safe expression of a heterologous polypeptide (e.g., a therapeutic polypeptide). In some embodiments, the polypeptide is a secreted polypeptide. In some embodiments, the polypeptide is one whose function normally acts (e.g., is functionally active) as a secreted polypeptide. As used herein, a "secreted polypeptide" refers to a protein that is secreted by a cell and / or is functionally active as a soluble extracellular protein.
[0206] In some embodiments, the polypeptide is an intracellular polypeptide. In some embodiments, the polypeptide is one whose function normally operates within a cell (e.g., is functionally active). As used herein, "intracellular polypeptide" refers to a protein that is not secreted by a cell, including soluble cytosolic polypeptides. One or more IRES and / or self-cleaving peptide sequences may be adjacent to an intracellular polypeptide, e.g., at or near a terminus of the polypeptide, such as the amino terminus of the polypeptide.
[0207] In some embodiments, the polypeptide is a wild-type polypeptide. In some embodiments, the polypeptide is a variant (e.g., mutant) polypeptide (e.g., a hyperactive variant of a wild-type polypeptide). In some embodiments, the polypeptide is a liver protein. In some embodiments, the polypeptide is a non-liver protein. In some embodiments, the polypeptide is Factor IX or a variant thereof. In some embodiments, the liver polypeptide is a polypeptide for treating a liver disorder, such as, but not limited to, tyrosinemia, Wilson's disease, Tay-Sachs disease, hyperbilirubinemia (Crigler-Najjar), acute intermittent porphyria, type 1 citrullinemia, progressive familial intrahepatic cholestasis, or maple syrup urine disease.
[0208] In some embodiments, expression of the polypeptide by the host cell (whether in vitro or in vivo) is increased by at least 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, or more relative to the level expressed by the host cell prior to providing a composition disclosed herein. In further embodiments, expression of the heterologous polypeptide may be increased to at least a detectable or therapeutically effective level.
[0209] In some embodiments, expression of the polypeptide by the host cell (whether in vitro or in vivo) is increased by at least 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70% or more of a known normal level (e.g., the level of the polypeptide in a healthy subject).
[0210] In some embodiments, expression of the polypeptide by a host cell (whether in vitro or in vivo) is at least about 10 μg / ml, 15 μg / ml, 20 μg / ml, 25 μg / ml, 30 μg / ml, 35 μg / ml, 40 μg / ml, 45 μg / ml, 50 μg / ml, 55 μg / ml, 60 μg / ml, 65 μg / ml, 70 μg / ml, 75 μg / ml, 80 μg / ml, 85 μg / ml, 90 μg / ml, 95 μg / ml, 100 μg / ml, 120 μg / ml, 140 μg / ml, 160 μg / ml, 180 μg / ml, 200 μg / ml, 25 ...0 μg / ml, 350 μg / ml, 400 μg / ml, 450 μg / ml, 500 μg / ml, 550 μg / ml, 600 μg / ml, 650 μg / ml, 700 μg / ml, 750 / ml, 225μg / ml, 250μg / ml, 275μg / ml, 300μg / ml, 325μg / ml, 350μg / ml, 400μg / ml, 450 μg / ml, 500μg / ml, 550μg / ml, 600μg / ml, 650μg / ml, 700μg / ml, 750μg / ml, 800μg / ml, 85 0μg / ml, 900μg / ml, 1000μg / ml, 1100μg / ml, 1200μg / ml, 1300μg / ml, 1400μg / ml, 1500 μg / ml, 1600 μg / ml, 1700 μg / ml, 1800 μg / ml, 1900 μg / ml, 2000 μg / ml, or higher. Methods for detecting and measuring polypeptides in various samples are well known in the art.
[0211] In some embodiments, the compositions and methods of the present disclosure are useful for treating liver-related diseases. As used herein, "liver-related disorders" refers to diseases that directly cause damage to liver tissue, diseases resulting from damage to liver tissue, and / or disorders of non-liver organs or tissues resulting from liver defects. Examples of liver-related diseases include, but are not limited to, tyrosinemia, Wilson's disease, Tay-Sachs disease, Crigler-Najjar hyperbilirubinemia, acute intermittent porphyria, type 1 citrullinemia, progressive familial intrahepatic cholestasis, and maple syrup urine disease.
[0212] As described herein, any one or more of the guide RNA, RNA-guided DNA binder, and donor construct disclosed herein can be delivered using any suitable delivery system and method known in the art. The compositions can be delivered in vitro or in vivo simultaneously or in any sequential order. In some embodiments, the donor construct, guide RNA, and RNA-guided DNA binder can be delivered in vitro or in vivo simultaneously, for example, in one vector, two vectors, individual vectors, one LNP, two LNPs, individual LNPs, or a combination thereof. In some embodiments, the donor construct can be delivered in vivo or in vitro as a vector and / or in association with LNPs (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 days or more) before the guide RNA and / or RNA-guided DNA binder is delivered as a vector and / or in association with LNPs alone or together as ribonucleoproteins (RNPs). As a further example, the guide RNA and RNA-guided DNA binding agent can be delivered in vivo or in vitro as a vector and / or alone or together as a ribonucleoprotein (RNP) in association with an LNP (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 days or more) before delivery of the construct as a vector and / or in association with an LNP. In some embodiments, the guide RNA and RNA-guided DNA binding agent are associated with an LNP and delivered to the host cell prior to delivery of the donor construct.
[0213] In some embodiments, the donor construct comprises a sequence encoding factor IX, or a variant thereof. For example, the variant has increased activity relative to the wild-type polypeptide. In some embodiments, the sequence encodes a polypeptide variant that is 80%, 85%, 90%, 93%, 95%, 97%, or 99% identical to the wild-type polypeptide sequence and has at least 80%, 85%, 90%, 92%, 94%, 96%, 98%, 99%, 100%, or more activity relative to the wild-type polypeptide. In some embodiments, the sequence encodes a fragment of the wild-type polypeptide, and the fragment has at least 80%, 85%, 90%, 92%, 94%, 96%, 98%, 99%, 100%, or more activity relative to the wild-type factor IX.
[0214] In some embodiments, a single administration of a donor construct comprising a heterologous gene, a guide RNA, and an RNA-guided DNA binding agent is sufficient to increase expression of a polypeptide of interest to a desired level. In other embodiments, more than one administration of a composition comprising a donor construct comprising a heterologous gene, a guide RNA, and an RNA-guided DNA binding agent may be beneficial to maximize the therapeutic effect.
[0215] In some embodiments, the guide RNA, RNA-guided DNA binding agent, and donor construct, individually or in any combination, are administered intravenously. In some embodiments, the guide RNA, RNA-guided DNA binding agent, and donor construct, individually or in any combination, are administered into the hepatic circulation.
[0216] In some embodiments, the host or subject is a mammal. In some embodiments, the host or subject is a human. In some embodiments, the host or subject is a rodent (e.g., a mouse).
[0217] This description and exemplary embodiments should not be construed as limiting. For purposes of this specification and the appended claims, unless otherwise expressly stated, all numbers expressing quantities, percentages, or proportions, and other numerical values used in the specification and claims should be understood to be modified in each instance by the term "about," to the extent that they are not already so modified. Accordingly, unless otherwise indicated, the numerical parameters set forth in the following specification and appended claims are approximations that may vary depending upon the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. The present invention also relates to the following: [Item 1] 1. A method for inserting a nucleic acid encoding a heterologous polypeptide into the albumin locus of a host cell or cell population, comprising: i) a gRNA, a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; b) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; e) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; f) a sequence selected from the group consisting of SEQ ID NOs: 34 to 97; g) a sequence complementary to 15 consecutive nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2-33; h) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 98 to 119; i) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 98 to 119, and j) a sequence selected from the group consisting of SEQ ID NOs: 120 to 163 The gRNA comprises a sequence selected from ii) an RNA-guided DNA binder; and iii) a construct comprising a nucleic acid encoding said heterologous polypeptide; thereby inserting a nucleic acid encoding said heterologous polypeptide into the albumin locus of said host cell or cell population; The method. [Item 2] 1. A method for expressing a heterologous polypeptide from an albumin locus in a host cell or cell population, comprising: i) a gRNA, a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; b) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; e) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; f) a sequence selected from the group consisting of SEQ ID NOs: 34 to 97, and g) a sequence comprising 15 consecutive nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2 to 33 The gRNA comprises a sequence selected from ii) an RNA-guided DNA binder; and and iii) a construct comprising a coding sequence for said heterologous polypeptide, thereby expressing said heterologous polypeptide in said host cell or cell population. [Item 3] 1. A method for expressing a therapeutic agent in a non-dividing cell type or cell population, comprising: i) a gRNA, a) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; b) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2, 8, 13, 19, 28, 29, 31, 32, and 33; c) a sequence selected from the group consisting of SEQ ID NOs: 34, 40, 45, 51, 60, 61, 63, 64, 65, 66, 72, 77, 83, 92, 93, 95, 96, and 97; d) a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; e) at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2 to 33; f) a sequence selected from the group consisting of SEQ ID NOs: 34 to 97, and g) a sequence comprising 15 consecutive nucleotides ± 10 nucleotides of the genomic coordinates listed for SEQ ID NOs: 2 to 33 The gRNA comprises a sequence selected from ii) an RNA-guided DNA binder; and and iii) a construct comprising a coding sequence for said heterologous polypeptide, thereby expressing said therapeutic agent in said non-dividing cell type or cell population. The method. [Item 4] 4. The method of any one of Items 1 to 3, wherein the gRNA comprises a guide sequence selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and SEQ ID NO:33. [Item 5] 5. The method according to any one of items 1 to 4, wherein the method is carried out in vivo. [Item 6] 5. The method according to any one of items 1 to 4, wherein the method is carried out in vitro. [Item 7] 7. The method according to any one of items 1 to 6, wherein the gRNA binds to a region upstream of a protospacer adjacent motif (PAM). [Item 8] 8. The method of claim 7, wherein the PAM is selected from NGG, NNGRRT, NNGRR(N), NNAGAAW, NNNNG(A / C)TT, and NNNNRYAC. [Item 9] 9. The method according to any one of items 1 to 8, wherein the gRNA is a dual gRNA (dgRNA). [Item 10] 9. The method according to any one of items 1 to 8, wherein the gRNA is a single gRNA (sgRNA). [Item 11] 11. The method of claim 10, wherein the sgRNA comprises one or more modified nucleosides. [Item 12] 12. The method of any one of items 1 to 11, wherein the RNA-guided DNA binding agent is Cas9 or a nucleic acid encoding Cas9. [Item 13] 13. The method according to any one of items 1 to 12, wherein the RNA-guided DNA binding agent is a nucleic acid encoding the RNA-guided DNA binding agent. [Item 14] 14. The method of claim 13, wherein the nucleic acid encoding the RNA-guided DNA binding agent is mRNA. [Item 15] 15. The method of claim 14, wherein the mRNA is a modified mRNA. [Item 16] 16. The method according to any one of items 1 to 15, wherein the RNA-guided DNA binding agent is a Cas nuclease or a nucleic acid encoding the Cas nuclease. [Item 17] 17. The method of claim 16, wherein the Cas nuclease is a class 2 Cas nuclease. [Item 18] 18. The method of claim 16 or 17, wherein the Cas nuclease is selected from the group consisting of S. pyogenes nuclease, S. aureus nuclease, C. jejuni nuclease, S. thermophilus nuclease, N. meningitidis nuclease, and variants thereof. [Item 19] 19. The method of any one of items 16 to 18, wherein the Cas nuclease is Cas9. [Item 20] 20. The method of claim 19, wherein the Cas nuclease is S. pyogenes Cas9 nuclease. [Item 21] 21. The method according to any one of items 16 to 20, wherein the Cas nuclease has site-specific DNA binding activity. [Item 22] 22. The method of any one of items 16 to 21, wherein the Cas nuclease is a nickase. [Item 23] 22. The method of any one of items 16 to 21, wherein the Cas nuclease is a cleavase. [Item 24] 22. The method of any one of items 19 to 21, wherein the Cas nuclease does not have nickase or cleavase activity. [Item 25] 25. The method according to any one of items 19 to 24, wherein the nucleic acid construct is a homology-independent donor construct. [Item 26] 26. The method according to any one of items 1 to 25, wherein the construct is a bidirectional nucleic acid construct. [Item 27] The construct comprises: i. a first segment comprising a coding sequence for a heterologous polypeptide, and ii. a second segment comprising the reverse complement of the coding sequence of said heterologous polypeptide. 27. The method according to Item 26, comprising: [Item 28] 28. The method of any one of items 1 to 27, wherein the construct comprises a polyadenylation signal sequence. [Item 29] 29. The method of any one of items 1 to 28, wherein the construct comprises a splice acceptor site. [Item 30] 31. The method according to any one of items 1 to 29, wherein the construct does not contain homology arms. 31. The method of any one of items 1 to 30, wherein the gRNA is administered in a vector and / or lipid nanoparticle. [Item 32] 32. The method according to any one of items 1 to 31, wherein the RNA-guided DNA-binding agent is administered in a vector and / or lipid nanoparticle. [Item 33] 33. The method of any one of items 1 to 32, wherein the construct comprising the heterologous gene is administered in a vector and / or lipid nanoparticle. [Item 34] 34. The method according to any one of items 31 to 33, wherein the vector is a viral vector. [Item 35] 35. The method of claim 34, wherein the viral vector is selected from the group consisting of an adeno-associated viral (AAV) vector, an adenoviral vector, a retroviral vector, and a lentiviral vector. [Item 36] 37. The vector of Item 35, wherein the AAV vector is selected from the group consisting of AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof. 37. The method of any one of items 1 to 36, wherein the construct comprising the coding sequence of the gRNA, the RNA-guided DNA binding agent, and the heterologous polypeptide is administered simultaneously, individually or in any combination. [Item 38] 37. The method of any one of items 1 to 36, wherein the construct comprising the coding sequence of the gRNA, the RNA-guided DNA binding agent, and the heterologous polypeptide is administered sequentially in any order and / or in any combination. [Item 39] 39. The method of any one of items 1 to 36 and 38, wherein the RNA-guided DNA binding agent, or the combined RNA-guided DNA binding agent and gRNA, is administered before providing the construct. [Item 40] 39. The method of any one of items 1 to 36 and 38, wherein the construct comprising the coding sequence of the heterologous polypeptide is administered before the gRNA and / or RNA-guided DNA-binding agent. [Item 41] 41. The method of any one of items 1 to 40, wherein the bidirectional nucleic acid construct, the RNA-guided DNA binding agent, and the gRNA are administered in any combination within 1 hour of each other. [Item 42] 42. The method according to any one of items 1 to 41, wherein the heterologous polypeptide is a secreted polypeptide. [Item 43] 42. The method according to any one of items 1 to 41, wherein the heterologous polypeptide is an intracellular polypeptide. [Item 44] 44. The method according to any one of items 1 to 43, wherein the cells are liver cells. [Item 45] 45. The method of claim 44, wherein the liver cells are hepatocytes. [Item 46] 46. The method of any one of items 1 to 45, wherein expression of the heterologous polypeptide in the host cell is increased by at least about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, or more relative to the level in the cell before administering a construct comprising the gRNA, RNA-guided DNA binding agent, and coding sequence for the heterologous polypeptide. [Item 47] 47. The method of any of items 1 to 46, wherein the gRNA comprises SEQ ID NO: 301 or SEQ ID NO: 302. [Item 48] 48. The method of any one of items 1 to 47, wherein the gRNA mediates target-specific cleavage by an RNA-guided DNA-binding agent, resulting in insertion of the coding sequence of the heterologous polypeptide within intron 1 of the albumin gene. [Item 49] 49. The method of any one of items 1 to 48, wherein the cleavage results in an insertion rate of at least about 2%, about 5%, or about 10% of the heterologous nucleic acid in the cell population. [Item 50] 50. The method of claim 49, wherein the cleavage results in an insertion rate of about 30-35%, about 35-40%, about 40-45%, about 45-50%, about 50-55%, about 55-60%, about 60-65%, about 65-70%, about 70-75%, about 75-80%, about 80-85%, about 85-90%, about 90-95%, or about 95-99% of the coding sequence of the heterologous polypeptide. [Item 51] 51. The method of any one of items 1 to 50, wherein the RNA-guided DNA-binding protein is S. pyogenes Cas9 nuclease. [Item 52] 52. The method of claim 51, wherein the nuclease is a cleavase or a nickase. [Item 53] 53. The method of any one of items 1 to 52, further comprising administering an LNP comprising the gRNA. [Item 54] 54. The method of any one of items 1 to 53, further comprising administering an LNP comprising mRNA encoding the RNA-guided DNA binder. [Item 55] 55. The method of claim 54, wherein the LNP comprises the gRNA and the mRNA encoding the RNA-guided DNA-binding agent. [Item 56] 56. The method of any one of items 1 to 55, wherein the gRNA and the RNA-guided DNA-binding protein are administered as an RNP. [Item 57] 57. The method of any one of items 53 to 56, wherein the construct is administered via a vector. [Item 58] 58. A host cell produced by the method according to any one of items 1 to 57. [Item 59] A host cell comprising a bidirectional nucleic acid construct encoding a heterologous polypeptide integrated within intron 1 of the albumin locus of the host cell. [Item 60] 60. The host cell of item 58 or 59, wherein the host cell is a liver cell. [Item 61] 61. The host cell according to any one of items 58 to 60, wherein the liver cell is a hepatocyte. [Item 62] 62. The method or host cell according to any one of items 1 to 61, wherein the RNA-guided DNA binding agent is a nucleic acid encoding the RNA-guided DNA binding agent. [Item 63] 63. The method or host cell of any one of items 1 to 62, wherein the RNA-guided DNA binding agent is a Cas nuclease. [Item 64] 64. The method or host cell according to any one of items 1 to 63, wherein the RNA-guided DNA binding agent is a nucleic acid encoding a Cas nuclease. [Item 65] 65. The method or host cell of any one of items 1 to 64, wherein the RNA-guided DNA binding agent is an mRNA encoding a Cas nuclease. [Item 66] 66. The method or host cell of item 65, wherein the mRNA is a modified mRNA. [Item 67] 67. The method or host cell of any one of items 63 to 66, wherein the Cas nuclease is a class 2 Cas nuclease. [Item 68] 68. The method or composition of any one of items 63 to 67, wherein the Cas nuclease is Cas9. [Item 69] 69. The method or composition of any one of items 63 to 68, wherein the Cas nuclease is selected from the group consisting of S. pyogenes nuclease, S. aureus nuclease, C. jejuni nuclease, S. thermophilus nuclease, N. meningitidis nuclease, and variants thereof. [Item 70] 70. The method or composition of item 69, wherein the Cas nuclease is S. pyogenes Cas9 nuclease or a variant thereof. [Item 71] 71. The method or composition according to any one of items 63 to 70, wherein the Cas nuclease has site-specific DNA binding activity. [Item 72] 72. The method or composition of any one of items 63 to 71, wherein the Cas nuclease is a nickase. [Item 73] 73. The method or composition of any one of items 63 to 72, wherein the Cas nuclease is a cleavase. [Item 74] 74. The method of any one of items 63 to 73, wherein the Cas nuclease does not have nickase or cleavase activity. [Item 75] 75. The method of any one of items 1 to 74, further comprising achieving at least about 1% of normal heterologous polypeptide activity or heterologous polypeptide level, such as at least about 5% of normal. [Item 76] 76. The method of any one of items 1 to 75, wherein the heterologous polypeptide activity or heterologous polypeptide level is less than about 500% relative to normal. [Item 77] 77. The method of any one of items 1 to 76, further comprising achieving at least about 1% to 300% of normal heterologous polypeptide activity or heterologous polypeptide level. [Item 78] 78. The in vivo method of any one of paragraphs 5 to 77, further comprising achieving a long-term effect, for example, at least 1 month, 2 months, 6 months, 1 year, or 2 years of effect in the individual. [Item 79] 79. The in vivo method of any one of paragraphs 5 to 78, wherein the individual's circulating albumin levels are normal at least 1 month, 2 months, 6 months, or 1 year after administration of the nucleic acid construct. [Item 80] 80. The in vivo method of any one of items 5 to 79, wherein the individual's circulating albumin levels are maintained 4 weeks after administration of the bidirectional nucleic acid construct. [Item 81] 81. The in vivo method of any one of items 5 to 80, wherein the individual's circulating albumin levels are temporarily reduced and then return to normal. [Item 82] 82. The method of any one of items 1 to 81, wherein the guide RNA comprises at least 17, 18, 19, or 20 consecutive nucleotides of a sequence selected from the group consisting of SEQ ID NOs: 2 to 33. [Item 83] 83. The method of any one of items 1 to 82, wherein the guide RNA comprises a sequence that is at least 95%, 90%, 85%, 80%, or 75% identical to a sequence selected from the group consisting of SEQ ID NOs: 2 to 33. [Item 84] 84. The method according to any one of items 1 to 83, wherein the guide RNA comprises a sequence selected from the group consisting of SEQ ID NOs: 2 to 33. [Item 85] 85. The method or host cell of any one of items 1 to 84, wherein the guide RNA comprises SEQ ID NO: 301 or 302. [Item 86] 86. The method and composition according to any one of items 1 to 85, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 2. [Item 87] 87. The method and composition according to any one of items 1 to 86, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 3. [Item 88] 88. The method and composition according to any one of items 1 to 87, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 4. [Item 89] 89. The method and composition according to any one of items 1 to 88, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 5. [Item 90] 89. The method and composition according to any one of items 1 to 89, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 6. [Item 91] 91. The method and composition according to any one of items 1 to 90, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 7. [Item 92] 92. The method and composition according to any one of items 1 to 91, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 8. [Item 93] 93. The method and composition according to any one of items 1 to 92, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 9. [Item 94] 94. The method and composition according to any one of items 1 to 93, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 10. [Item 95] 95. The method and composition according to any one of items 1 to 94, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 11. [Item 96] 96. The method and composition according to any one of items 1 to 95, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 12. [Item 97] 97. The method and composition according to any one of items 1 to 96, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 13. [Item 98] 98. The method and composition according to any one of items 1 to 97, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 14. [Item 99] 99. The method and composition according to any one of items 1 to 98, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 15. [Item 100] 99. The method and composition according to any one of items 1 to 99, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 16. [Item 101] 101. The method and composition according to any one of items 1 to 100, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 17. [Item 102] 102. The method and composition according to any one of items 1 to 101, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 18. [Item 103] 103. The method and composition according to any one of items 1 to 102, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 19. [Item 104] 104. The method and composition according to any one of items 1 to 103, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 20. [Item 105] 105. The method and composition according to any one of items 1 to 104, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 21. [Item 106] 106. The method and composition according to any one of items 1 to 105, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 22. [Item 107] 107. The method and composition according to any one of items 1 to 106, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 23. [Item 108] 108. The method and composition according to any one of items 1 to 107, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 24. [Item 109] 109. The method and composition according to any one of items 1 to 108, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 25. [Item 110] 100. The method and composition according to any one of items 1 to 109, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 26. [Item 111] 111. The method and composition according to any one of items 1 to 110, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 27. [Item 112] 112. The method and composition according to any one of items 1 to 111, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 28. [Item 113] 13. The method and composition according to any one of items 1 to 112, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 29. [Item 114] 14. The method and composition according to any one of items 1 to 113, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 30. [Item 115] 115. The method and composition according to any one of items 1 to 114, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 31. [Item 116] 116. The method and composition according to any one of items 1 to 115, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 32. [Item 117] 117. The method and composition according to any one of items 1 to 116, wherein the guide RNA comprises the nucleic acid sequence of SEQ ID NO: 33. [Example]
[0218] The following examples are provided to illustrate certain disclosed embodiments and should not be construed as limiting the scope of the disclosure in any way.
[0219] Example 1 - Materials and Methods Cloning and plasmid preparation A bidirectional insert construct flanked by AAV2 ITRs was synthesized and cloned into pUC57-Kan by a commercial vendor. The resulting construct (P00147) was used as a parent cloning vector for other vectors. Other insert constructs (without ITRs) were also commercially synthesized and cloned into pUC57. Purified plasmids were digested with BglII restriction enzyme (New England BioLabs, catalog number R0144S), and the insert constructs were cloned into the parent vector. Plasmids were propagated in Stbl3™ Chemically Competent E. coli (Thermo Fisher, catalog number C737303).
[0220] AAV generation Triple transfection in HEK293 cells was used to package genomes with constructs of interest for AAV8 and AAV-DJ generation, and the resulting vectors were purified from both lysed cells and culture medium by iodixanol gradient ultracentrifugation (see, e.g., Lock et al., Hum Gene Ther. 2010 Oct;21(10):1259-71). Plasmids used in triple transfection, including genomes with constructs of interest, are referred to by "PXXXX" numbers in the Examples; see also, e.g., Table 9. Isolated AAV was dialyzed in storage buffer (PBS containing 0.001% Pluronic F68). AAV titers were determined by qPCR using primers / probes located within the ITR region.
[0221] In vitro transcription ("IVT") of nuclease mRNA Polyadenylated Streptococcus pyogenes ("Spy") Cas9 mRNA containing N1-methylpseudo-U capped mRNA was generated by in vitro transcription using a linearized plasmid DNA template and T7 RNA polymerase. Generally, plasmid DNA containing a T7 promoter and a 100-nt poly(A / T) tract was linearized by incubating with XbaI at 37°C to complete digestion, followed by complete heat inactivation of XbaI at 65°C. The linearized plasmid was purified from enzymes and buffer salts. The IVT reaction to generate Cas9-modified mRNA was incubated at 37°C for 4 hours with the following conditions: 50 ng / μL linearized plasmid, 2 mM each of GTP, ATP, CTP, and N1-methylpseudo-UTP (Trilink), 10 mM ARCA (Trilink), 5 U / μL T7 RNA polymerase (NEB), 1 U / μL mouse RNase inhibitor (NEB), 0.004 U / μL inorganic E. coli pyrophosphatase (NEB), and 1x reaction buffer. TURBO Dnase (ThermoFisher) was added to a final concentration of 0.01 U / μL, and the reaction was incubated for an additional 30 minutes to remove the DNA template. Cas9 mRNA was purified using the MegaClear Transcription Clean-up Kit according to the manufacturer's protocol (ThermoFisher). Alternatively, Cas9 mRNA was purified using LiCl, ammonium acetate, and sodium acetate precipitation, or using LiCl precipitation, followed by further purification by tangential flow filtration. Transcript concentrations were determined by measuring optical absorbance at 260 nm (Nanodrop), and transcripts were analyzed by capillary electrophoresis using a Bioanlayzer (Agilent).
[0222] The following Cas9 mRNAs comprise Cas9 ORF SEQ ID NO:703 or SEQ ID NO:704, or the sequences of Table 24 of PCT / US2019 / 053423 (incorporated herein by reference).
[0223] Lipid formulations for delivery of Cas9 mRNA and gRNA Cas9 mRNA and gRNA were delivered to cells and animals using lipid formulations containing the ionic lipid ((9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyloctadeca-9,12-dienoate), also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate), cholesterol, DSPC, and PEG2k-DMG.
[0224] For experiments utilizing premixed lipid formulations (referred to herein as "lipid packets"), the components were reconstituted in 100% ethanol at a molar ratio of 50:38:9:3 ionic lipid:cholesterol:DSPC:PEG2k-DMG before being mixed with RNA cargo (e.g., Cas9 mRNA and gRNA) at a lipid amine to RNA phosphophosphate (N:P) molar ratio of approximately 6.0, as further described herein.
[0225] For experiments utilizing components formulated as lipid nanoparticles (LNPs), the components were diluted in 100% ethanol at various molar ratios. RNA cargo (e.g., Cas9 mRNA and gRNA) was dissolved in 25 mM citrate, 100 mM NaCl (pH 5.0) to yield an RNA cargo concentration of approximately 0.45 mg / mL.
[0226] For the experiments described in Example 2, LNPs were formed by microfluidic mixing of lipid and RNA solutions using a Precision Nanosystems NanoAssemblr™ Benchtop Instrument according to the manufacturer's protocol. A 2:1 ratio of aqueous to organic solvent was maintained during mixing using differential flow rates. After mixing, the LNPs were collected, diluted in water (approximately 1:1 v / v), held at room temperature for 1 hour, and further diluted with water (approximately 1:1 v / v) before a final buffer exchange. A final buffer exchange into 50 mM Tris, 45 mM NaCl, 5% (w / v) sucrose, pH 7.5 (TSS) was completed using a PD-10 desalting column (GE). If necessary, the formulation was concentrated by centrifugation through an Amicon 100 kDa centrifugal filter (Millipore). The resulting mixture was then filtered using a 0.2 μm sterile filter. The final LNPs were stored at -80°C until further use. LNPs were formulated at a lipid amine to RNA phosphate (N:P) molar ratio of approximately 4.5 and a gRNA to mRNA weight ratio of 1:1, with an ionic lipid:cholesterol:DSPC:PEG2k-DMG molar ratio of 45:44:9:2.
[0227] For the experiments described in other examples, LNPs were prepared using cross-flow technology by impinging jet mixing of lipids in ethanol with two volumes of RNA solution and one volume of water. Lipids in ethanol were mixed with two volumes of RNA solution through a mixing cross. A fourth water stream was mixed with the outlet flow of the cross through an in-line tee (see Figure 2 in WO2016010840). The LNPs were held at room temperature for 1 hour and further diluted with water (approximately 1:1 v / v). The diluted LNPs were concentrated using tangential flow filtration on a flat-sheet cartridge (Sartorius, 100 kD MWCO) and then buffer-exchanged into 50 mM Tris, 45 mM NaCl, 5% (w / v) sucrose, pH 7.5 (TSS) by diafiltration. Alternatively, the final buffer exchange to TSS was completed using a PD-10 desalting column (GE). If necessary, the preparation was concentrated by centrifugation through an Amicon 100 kDa centrifugal filter (Millipore). The resulting mixture was then filtered using a 0.2 μm sterile filter. The final LNP was stored at 4°C or -80°C until further use. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.0 and a gRNA to mRNA weight ratio of 1:1, with an ionic lipid:cholesterol:DSPC:PEG2k-DMG molar ratio of 50:38:9:3.
[0228] Cell culture and in vitro delivery of Cas9 mRNA, gRNA, and insertion constructs Hepa1-6 cells Hepa1-6 cells were seeded in a 96-well plate at a density of 10,000 cells / well. 24 hours later, the cells were treated with LNP and AAV. Before treatment, the medium was aspirated from the wells. LNP was diluted to 4 ng / ul in DMEM + 10% FBS medium, and further diluted to 2 ng / ul in 10% FBS (in DMEM) and incubated at 37°C for 10 minutes (at a final concentration of 5% FBS). The target MOI of AAV was 1e6 and diluted in DMEM + 10% FBS medium. 50 μl of the diluted LNP at 2 ng / ul was added to the cells (delivering a total of 100 ng of RNA cargo), followed by the addition of 50 μl of AAV. The LNP and AAV treatments were spaced apart by several minutes. The total volume of medium in the cells was 100 μl. After 72 hours and 30 days of treatment, supernatants from these treated cells were collected for human FIX ELISA analysis as described below.
[0229] primary hepatocytes Primary mouse hepatocytes (PMH), primary cynomolgus monkey hepatocytes (PCH), and primary human hepatocytes (PHH) were thawed and resuspended in Hepatocyte Thawing Medium containing supplements (ThermoFisher), followed by centrifugation. The supernatant was discarded, and the pelleted cells were resuspended in Hepatocyte Seeding Medium and Supplement Pack (ThermoFisher). Cells were counted and seeded onto 96-well plates coated with Bio-coat collagen I at densities of 33,000 cells / well for PHH, 50,000 cells / well for PCH, and 15,000 cells / well for PMH. The seeded cells were allowed to adhere for 5 hours in a tissue culture incubator at 37°C and 5% CO2. After incubation, cells were examined for monolayer formation, washed three times with Hepatocyte Maintenance Medium, and incubated at 37°C.
[0230] In experiments using lipid packet delivery, Cas9 mRNA and gRNA were each diluted separately to 2 mg / ml in maintenance medium, and 2.9 μl of each was added to wells (in a 96-well Eppendorf plate) containing 12.5 μl of 50 mM sodium citrate, 200 mM sodium chloride (pH 5), and 6.9 μl of water. Then, 12.5 μl of lipid packet formulation was added, followed by 12.5 μl of water and 150 μl of TSS. Each well was diluted to 20 ng / μl (based on total RNA content) using hepatocyte maintenance medium, then diluted to 10 ng / μl (based on total RNA content) with 6% fresh mouse serum. The medium was aspirated from the cells prior to transfection, and 40 μl of the lipid packet / RNA mixture was added to the cells, followed by AAV (diluted in maintenance medium) at an MOI of 1e5. Media was collected 72 hours after treatment for analysis and cells were harvested for further analysis as described herein.
[0231] Luciferase assay For experiments involving NanoLuc detection in cell culture media, 1 volume of Nano-Glo® Luciferase Assay Substrate was combined with 50 volumes of Nano-Glo® Luciferase Assay Buffer. Assays were run on a Promega Glomax runner with a 0.5 second integration time using a 1:10 dilution of sample (50 μl reagent + 40 μl water + 10 μl cell culture media).
[0232] For experiments involving detection of HiBit tags in cell culture media, LgBiT protein and Nano-GloR HiBiT extracellular matrix were diluted 1:100 and 1:50, respectively, in Nano-GloR HiBiT extracellular buffer at room temperature. Assays were run on a Promega Glomax runner with a 1.0 second integration time using a 1:10 dilution of sample (50 μl reagent + 40 μl water + 10 μl cell culture media).
[0233] In vivo delivery of LNPs and / or AAVs Mice were administered AAV, LNP, both AAV and LNP, or vehicle (PBS + 0.001% Pluronic for AAV vehicle, TSS for LNP vehicle) via the lateral tail vein. AAV was administered in the amount described herein (vector genome / mouse, "vg / ms") in a volume of 0.1 mL per mouse. LNP was diluted in TSS and administered in the amount described herein, approximately 5 μl / gram body weight. Typically, mice were injected first with AAV, followed by LNP, if applicable. At various time points after treatment, serum and / or liver tissue were collected for certain analyses, as further described below.
[0234] Human Factor IX (hFIX) ELISA Assay For in vitro studies, total human factor IX levels secreted into cell culture media were determined using a human factor IX ELISA kit (Abcam, catalog no. ab188393) according to the manufacturer's protocol. Secreted hFIX levels were quantified from a standard curve using a four-parameter logistic fit and expressed as ng / ml of culture media.
[0235] For in vivo studies, blood was collected and serum or plasma was isolated as indicated. Total human factor IX levels were determined using a human factor IX ELISA kit (Abcam, catalog number ab188393) according to the manufacturer's protocol. Serum or plasma hFIX levels were quantified from a standard curve using a four-parameter logistic fit and expressed as μg / mL of serum.
[0236] Next-generation sequencing ("NGS") and analysis of on-target cleavage efficiency Deep sequencing was used to identify the presence of insertions and deletions introduced by gene editing (e.g., within intron 1 of albumin). PCR primers were designed around the target sites to amplify the genomic region of interest. Primer sequence design was performed as standard practice in the field.
[0237] An additional PCR was performed according to the manufacturer's protocol (Illumina) to add the necessary chemistry for sequencing. The amplicons were sequenced on an Illumina MiSeq instrument. After removing reads with low quality scores, the reads were aligned to the reference genome. The resulting file containing the reads was mapped to the reference genome (BAM file), reads that overlapped with the target region of interest were selected, and the number of wild-type reads versus the number of reads containing insertions or deletions ("indels") was calculated.
[0238] The editing percentage (e.g., "editing efficiency" or "editing rate") is defined as the total number of sequence reads containing insertions or deletions ("indels") relative to the total number of sequence reads containing the wild type.
[0239] In situ hybridization analysis BaseScope (ACDbio, Newark, CA) is a specialized RNA in situ hybridization technology that can provide specific detection of exon junctions in hybrid mRNA transcripts containing, for example, an inserted transgene (hFIX) and coding sequences from the insertion site (exon 1 of albumin). BaseScope was used to measure the percentage of liver cells expressing the hybrid mRNA.
[0240] To detect hybrid mRNAs, two probes for hybrid mRNAs that may result after insertion of the bidirectional construct were designed by ACDbio (Newark, CA). One probe was designed to detect hybrid mRNAs resulting from insertion of the construct in one orientation, and the other probe was designed to detect hybrid mRNAs resulting from insertion of the construct in the other orientation. Livers from different groups of mice were collected, fresh-frozen, and sectioned. BaseScope assays were performed using single or pooled probes according to the manufacturer's protocol. Slides were scanned and analyzed using HALO software. The background of this assay (saline-treated group) was 0.58%.
[0241] Example 2 - In vitro testing of albumin intron 1 insertion templates with and without homology arms In this example, Hepa1-6 cells were cultured and treated with insertion templates bearing various forms of AAV (e.g., bearing either a single-stranded genome ("ssAAV") or a self-complementary genome ("scAAV")) in the presence or absence of LNPs delivering Cas9 mRNA and G000551, e.g., as described in Example 1 (n=3). AAV and LNPs were prepared as described in Example 1. After treatment, media was collected for human factor IX levels, as described in Example 1.
[0242] Hepa1-6 cells are an immortalized mouse liver cell line that continues to divide in culture. As shown in Figure 2 (72 hours after treatment), a vector containing 200-bp homology arms (scAAV derived from plasmid P00204) resulted in detectable expression of hFIX after insertion into intron 1 of albumin in cycling cells. Use of AAV vectors derived from P00123 (scAAV lacking homology arms) and P00147 (ssAAV bidirectional construct lacking homology arms) did not result in detectable expression of hFIX in this experiment. Cells were maintained in culture, and these results were confirmed when re-assayed 30 days after treatment (data not shown).
[0243] Example 3 - In vivo testing of albumin intron 1 insertion templates with and without homology arms In this example, mice were treated with AAV derived from the same plasmids (P00123, P00204, and P00147) tested in vitro in Example 2. Dosage materials were prepared and administered as described in Example 1. C57B1 / 6 mice were each administered LNPs containing 3e11 vector genomes (vg / ms) followed by G000551 ("G551") at a dose of 4 mg / kg (based on total RNA cargo content) (n=5 for each group). Four weeks after administration, animals were euthanized, and liver tissue and serum were collected for editing and hFIX expression, respectively.
[0244] As shown in Figure 3A and Table 12, approximately 60% liver editing levels were detected in each group of animals treated with LNP containing a gRNA targeting intron 1 of mouse albumin. However, despite robust and consistent levels of editing in each treatment group, animals receiving a ssAAV vector without homology arms (vector derived from P00147) combined with LNP treatment resulted in the highest levels of hFIX expression in serum (Figure 3B and Table 13). Table 12: Indel % TIFF0007820434000016.tif43152Table 13: Factor IX levels (ug / mL) TIFF0007820434000017.tif48152
[0245] Example 4 - In vivo testing of ssAAV insertion templates of albumin intron 1 with and without homology arms The experiments described in this example investigated the effect of incorporating homology arms into ssAAV vectors in vivo.
[0246] The substances used in this experiment were prepared and administered as described in Example 1. C57Bl / 6 mice were administered LNPs containing 3e11vg / ms followed by G000666 ("G666") or G000551 ("G551") at a dose of 0.5mg / kg (based on total RNA cargo content) (n=5 for each group). Four weeks after administration, animal serum was collected for hFIX expression.
[0247] As shown in Figure 4A and Table 14, the use of ssAAV vectors with asymmetric homology arms (300 / 600 bp arms, 300 / 2000 bp arms, and 300 / 1500 bp arms for vectors derived from plasmids P00350, P00356, and P00362, respectively) for insertion into the albumin intron 1 site targeted by G551 resulted in levels of circulating hFIX below the lower limit of detection for the assay. However, the use of an ssAAV vector (derived from P00147) without homology arms and carrying two hFIX open reading frames (ORFs) in a bidirectional orientation resulted in detectable levels of circulating hFIX in each animal.
[0248] Similarly, use of ssAAV vectors with symmetric homology arms (derived from plasmids P00353 and P00354, respectively) for insertion into the albumin intron 1 site targeted by G666 resulted in lower, but detectable, levels compared to use of the bidirectional vector without homology arms (derived from P00147) (see Figure 4B and Table 15). Table 14:hFIX TIFF0007820434000018.tif43152Table 15: hFIX serum levels TIFF0007820434000019.tif37153
[0249] Example 5 - In vitro screening of bidirectional constructs spanning 20 target sites in intron 1 of albumin in primary mouse hepatocytes Having demonstrated that bidirectional constructs lacking homology arms outperformed vectors with other configurations for insertion into albumin intron 1, the experiments described in this example investigated the effect of altering the splice acceptor. These diverse bidirectional constructs were tested across a panel of target sites utilizing 20 different gRNAs targeting intron 1 of mouse albumin in primary mouse hepatocytes (PMH).
[0250] The ssAAV and lipid packet delivery agents tested in this example were prepared and delivered to PMH as described in Example 1, along with AAV at an MOI of 1e5. After treatment, isolated genomic DNA and cell culture medium were collected for editing and transgene expression analysis, respectively. Each of the vectors contained a reporter that could be measured by luciferase-based fluorescence detection, as described in Example 1, and the relative luciferase units ("RLU") are plotted in Figure 5C. For example, AAV vectors containing the hFIX ORF contained a HiBit peptide fused to their 3' ends, while AAV vectors containing only a reporter gene contained a NanoLuc ORF (in addition to GFP). A schematic diagram of each of the tested vectors is provided in Figure 5A. The gRNAs tested are shown in Figures 5B and 5C, using abbreviated numbers for those listed in Table 5 (e.g., leading zeros are omitted, e.g., "G551" corresponds to "G000551" in Table 5).
[0251] As shown in Figure 5B and Table 16, consistent but varying levels of editing were detected for each treatment group across each combination tested. Transgene expression using various combinations of template and guide RNA is shown in Figure 5C. As shown in Figure 5D, significant levels of indel formation did not necessarily result in more efficient expression of the transgene. Not all guides generated with significant indels resulted in high levels of protein with the same inserted template as measured by relative luciferase activity. Using the P00411- and P00418-derived templates, when no guides contained less than 10% editing, R 2 The values were 0.54 and 0.37 between indels and luciferase activity, respectively (Figure 5D). Interestingly, despite the different ORFs and splice acceptors, the relative expression levels, as measured by RLU, were consistent among the three vectors tested, demonstrating the robustness, reproducibility, and modularity of the bidirectional construct system for use, for example, in inserting a transgene of interest into intron 1 of albumin (see Figure 5C and Table 17). The mouse albumin splice acceptor and the human FIX splice acceptor each resulted in efficient transgene expression. Table 16: Indel % TIFF0007820434000020.tif155167Table 17: Luciferase expression TIFF0007820434000021.tif166166
[0252] Example 6 - In vivo screening of bidirectional constructs spanning the albumin intron 1 target site The ssAAV and LNPs tested in this example were prepared and delivered to C57B1 / 6 mice as described in Example 1 to evaluate the performance of the bidirectional constructs across target sites in vivo. Four weeks after administration, the animals were euthanized, and liver tissue and serum were collected for editing and hFIX expression, respectively.
[0253] In initial experiments, 10 different LNP formulations containing 10 different gRNAs targeting intron 1 of albumin were delivered to mice along with ssAAV derived from P00147. AAV and LNP were delivered at 3e11 vg / ms and 4 mg / kg (with respect to total RNA cargo content), respectively (n=5 for each group). The gRNAs tested in this experiment are shown in Figure 6. As shown in Figure 6 and Table 18 and observed in vitro, significant levels of indel formation were not predictive of transgene insertion or expression. Table 18: hFIX serum levels and % indels TIFF0007820434000022.tif77166
[0254] In a separate experiment, a panel of 20 gRNAs targeting 20 different target sites tested in vitro in Example 5 was tested in vivo. To this end, an LNP formulation containing 20 gRNAs targeting intron 1 of albumin was delivered to mice along with ssAAV derived from P00147. AAV and LNP were delivered at 3e11 vg / ms and 1 mg / kg (with respect to total RNA cargo content), respectively. The gRNAs tested in this experiment are shown in Figures 7A and 7B. Table 19. Liver TIFF0007820434000023.tif140162Table 20: Serum hFIX levels TIFF0007820434000024.tif238165TIFF0007820434000025.tif64165
[0255] As shown in Figure 7A and Table 19, varying levels of editing were detected for each treatment group across each LNP / vector combination tested. However, as shown in Figure 7B and consistent with the in vitro data described in Example 5, higher levels of editing did not necessarily result in higher levels of transgene expression in vivo, indicating a lack of correlation between editing and insertion / expression of the bidirectional construct. Indeed, there is little correlation between the amount of editing achieved and the amount of hFIX expression seen in the plots provided in Figure 7D and Table 20. Notably, when gRNAs achieving less than 10% editing are removed from the analysis, there is an R of only 0.34 between the editing and expression datasets for this experiment. 2 Interestingly, as shown in Figure 7C, a correlation plot is provided comparing the level of expression as measured in RLU from the in vitro experiment of Example 5 with the in vivo transgene expression levels detected in this experiment, showing that R 2 The value was 0.70, demonstrating a positive correlation between the primary cell screen and the in vivo treatment.
[0256] To assess the insertion of the bidirectional construct at the cellular level, liver tissue from treated animals was assayed using an in situ hybridization method (BaseScope) as described, for example, in Example 1. This assay utilized a probe capable of detecting the junction between the hFIX transgene and mouse albumin exon 1 sequence as a hybrid transcript. As shown in Figure 8A, cells positive for the hybrid transcript were detected in animals receiving both AAV and LNP. Specifically, when AAV alone was administered, less than 1.0% of cells were positive for the hybrid transcript. When LNP containing G011723, G000551, or G000666 was administered, 4.9%, 19.8%, or 52.3% of cells were positive for the hybrid transcript. Furthermore, as shown in Figure 8B, circulating hFIX levels correlated with the number of cells positive for the hybrid transcript. Finally, the assay utilized a pooled probe capable of detecting the insertion of the bidirectional hFIX construct in either orientation. However, when a single probe detecting only a single orientation was used, the amount of cells positive for the hybrid transcript was approximately half of that detected using pooled probes (4.46% vs. 9.68% in one example), suggesting that bidirectional constructs can indeed be inserted into albumin intron 1 in either orientation, resulting in expressed hybrid transcripts that correlate with the amount of transgene expression at the protein level. These data indicate that the circulating protein levels achieved are dependent on the guide used for insertion.
[0257] Example 7 - Timing of AAV and LNP delivery in vivo In this example, the timing between delivery of an ssAAV containing a bidirectional hFIX construct for targeted insertion into intron 1 of albumin and delivery of LNP was investigated.
[0258] The ssAAV and LNP tested in this example were prepared and delivered to mice as described in Example 1. The LNP formulation contained G000551, and the bidirectional template was delivered as ssAAV derived from P00147. AAV and LNP were delivered at 3e11 vg / ms and 4 mg / kg (based on total RNA cargo content), respectively (n=5 for each group). The "template only" cohort received AAV only, and the "PBS" cohort received no AAV or LNP. One cohort received AAV and LNP sequentially (minutes apart) on day 0 ("template + LNP day 0"), another cohort received AAV on day 0 and LNP on day 1 ("template + LNP day 1"), and the final cohort received AAV on day 0 and LNP on day 7 ("template + LNP day 7"). Plasma was collected at 1, 2, and 6 weeks for hFIX expression analysis.
[0259] As shown in Figure 9, hFIX was detected in each cohort at each time point assayed, except for the 1-week time point in the cohort that received LNP on the same day at week 1 post-AAV delivery.
[0260] Example 8 - Multiple administration of LNPs after delivery of AAV In this example, the effect of repeated administration of LNP after administration of ssAAV for targeted insertion into intron 1 of albumin was investigated.
[0261] The ssAAV and LNPs tested in this example were prepared and delivered to C57B1 / 6 mice as described in Example 1. The LNP formulation contained G000551, and the ssAAV was derived from P00147. AAV and LNPs were delivered at 3e11 vg / ms and 0.5 mg / kg (with respect to total RNA cargo content), respectively (n=5 for each group). The "template only" cohort received AAV only, and the "PBS" cohort received no AAV or LNPs. One cohort received AAV and LNP sequentially (minutes apart) on day 0 without further treatment ("Template + LNP (1)" in Figure 10), another cohort received AAV and LNP sequentially (minutes apart) on day 0 and a second dose on day 7 ("Template + LNP (2)" in Figure 10), and the final cohort received AAV and LNP sequentially (minutes apart) on day 0, a second dose of LNP on day 7 and a third dose of LNP on day 14 ("Template + LNP (3)" in Figure 10). Plasma was collected for hFIX expression analysis 1, 2, 4, and 6 weeks after AAV administration.
[0262] As shown in Figure 10, hFIX was detected in each cohort at each time point assayed, and multiple subsequent doses of LNP did not significantly increase the amount of hFIX expression.
[0263] Example 9 - Durability of hFIX expression in vivo In this example, we evaluated the durability of hFIX expression in treated animals over time after targeted insertion into albumin intron 1. To this end, hFIX was measured in the serum of treated animals as part of a 1-year durability study.
[0264] The ssAAV and LNPs tested in this example were prepared and delivered to C57B1 / 6 mice as described in Example 1. The LNP formulation contained G000551, and the ssAAV was derived from P00147. AAV was delivered at 3e11 vg / ms, and LNPs were delivered at either 0.25 or 1.0 mg / kg (with respect to total RNA cargo content) (n=5 for each group).
[0265] As shown in Figures 11A and 11B and Tables 21 and 22, hFIX expression from albumin intron 1 was maintained for both groups up to 12 weeks at each time point assessed. The decrease in levels observed at 8 weeks is likely due to variability in the ELISA assay. Serum albumin levels were measured by ELISA at 2 and 41 weeks, demonstrating that circulating albumin levels were maintained throughout the study. Table 21: FIX Levels TIFF0007820434000026.tif66168Table 22: hFIX Levels TIFF0007820434000027.tif67166
[0266] Example 10 - Effect of various doses of AAV and LNP on modulating hFIX expression in vivo In this example, the effect of varying doses of both AAV and LNP to modulate hFIX expression after targeted insertion into intron 1 of albumin was evaluated in C57B1 / 6 mice.
[0267] The ssAAV and LNPs tested in this example were prepared and delivered to mice as described in Example 1. The LNP formulation contained G000553, and the ssAAV was derived from P00147. AAV was delivered at 1e11, 3e11, 1e12, or 3e12 vg / ms, and LNPs were delivered at 0.1, 0.3, or 1.0 mg / kg (with respect to total RNA cargo content) (n=5 for each group). Two weeks after administration, animals were euthanized, and serum was collected for hFIX expression analysis.
[0268] As shown in Figure 12A (1 week) and Figure 12B (2 weeks) and Table 23, the amount of hFIX expression from intron 1 of albumin in vivo can be modulated by varying the dose of either AAV or LNP. Table 23: Serum hFIX TIFF0007820434000028.tif176148
[0269] Example 11 - In vitro screening of bidirectional constructs spanning target sites in primary cynomolgus monkey and primary human hepatocytes In this example, ssAAV vectors containing bidirectional constructs were tested across a panel of target sites using gRNAs targeting intron 1 of cynomolgus monkey ("cyno") and human albumin in primary cynomolgus monkey (PCH) and primary human hepatocytes (PHH), respectively.
[0270] The ssAAV and lipid packet delivery agents tested in this example were prepared and delivered to PCH and PHH as described in Example 1. After treatment, isolated genomic DNA and cell culture medium were collected for editing and transgene expression analysis, respectively. Each of the vectors contained a reporter (derived from plasmid P00415) that could be measured by luciferase-based fluorescence detection, as described in Example 1, and the RLU data were plotted as relative luciferase units ("RLU") in Figures 13B and 14B. The RLU data shown graphically in Figures 13B and 14B are numerically reproduced in Tables 3 and 4 below. For example, the AAV vector contained the NanoLuc ORF (in addition to GFP). Schematic diagrams of the tested vectors are provided in Figures 13B and 14B. The gRNAs tested are indicated in each figure using abbreviated numbers for those listed in Tables 1 and 3.
[0271] As shown in Figure 13A for PCH and Figure 14A for PHH, varying levels of editing were detected for each of the combinations tested (editing data for some combinations tested in the PCH experiment are not reported in Figure 13A and Table 3 due to the poor performance of certain primer pairs used for amplicon-based sequencing). The editing data shown graphically in Figures 13A and 14A are numerically reproduced in Tables 3 and 4 below. However, as shown in Figures 13B, 13C, and 14B and 14C, significant levels of indel formation were not predictive for transgene insertion or expression, indicating little correlation between editing and insertion / expression of bidirectional constructs in PCH and PHH, respectively. As one measure, the R calculated in Figure 13C 2 The value is 0.13, and the R 2 The value is 0.22. Table 3: Albumin intron 1 editing and transgene expression data for sgRNA delivered to primary cynomolgus monkey hepatocytes TIFF0007820434000029.tif176166 Table 4: Albumin intron 1 editing and transgene expression data for sgRNA delivered to primary human hepatocytes TIFF0007820434000030.tif207161
[0272] Additionally, ssAAV vectors containing the bidirectional constructs were tested across a panel of target sites using a single guide RNA targeting intron 1 of human albumin in primary human hepatocytes (PHH).
[0273] ssAAV and LNP materials were prepared and delivered to PHHs as described in Example 1. After processing, isolated genomic DNA and cell culture medium were collected for editing and transgene expression analysis, respectively. Each of the vectors contained a reporter (derived from plasmid P00415) that can be measured by luciferase-based fluorescence detection, as described in Example 1, and is plotted in Figure 14D and shown as relative luciferase units ("RLU") in Table 24. For example, the AAV vector contained the NanoLuc ORF (in addition to GFP). Schematics of the tested vectors are provided in Figures 13B and 14B. The tested gRNAs are shown in Figure 14D using abbreviated numbers for those listed in Tables 1 and 7. Table 24: Albumin intron 1 transgene expression data for sgRNA delivered to primary cynomolgus monkey hepatocytes TIFF0007820434000031.tif180152
[0274] Example 12 - In vivo testing of human Factor 9 gene insertion in non-human primates In this example, an 8-week study was conducted to evaluate human Factor 9 gene insertion and hFIX protein expression in cynomolgus monkeys following administration of adeno-associated viruses (AAV) and / or lipid nanoparticles (LNP) containing various guides. The study was performed with LNP and AAV formulations prepared as described above. Each LNP formulation contained Cas9 mRNA and guide RNA (gRNA) at a 2:1 mRNA:gRNA weight ratio. The ssAAV was derived from P00147.
[0275] Male cynomolgus monkeys were treated in cohorts of n=3. Animals were administered AAV by slow bolus injection or infusion at the doses listed in Table 5. After AAV treatment, animals received buffer or LNP as listed in Table 5 by slow bolus or infusion.
[0276] Two weeks after administration, liver specimens were collected by single ultrasound-guided percutaneous biopsy. Each biopsy specimen was flash-frozen in liquid nitrogen and stored at -86 to -60°C. Editing analysis of liver specimens was performed by next-generation sequencing (NGS) as previously described.
[0277] Blood samples were collected from animals on days 7, 14, 28, and 56 after administration for Factor IX ELISA analysis. After blood collection, blood samples were collected, processed to plasma, and stored at −86 to −60°C until analysis.
[0278] Total human factor IX levels were determined from plasma samples by ELISA. Briefly, Reacti-Bind 96-well microplates (VWR catalog no. PI15041) were coated with a capture antibody (mouse mAB against human factor IX antibody (HTI, catalog no. AHIX-5041)) at a concentration of 1 μg / ml and then blocked using 1x PBS containing 5% bovine serum albumin. Test samples or standards of purified human factor IX protein (ERL, catalog no. HFIX1009, lot no. HFIX4840) diluted in cynomolgus monkey plasma were then incubated in individual wells. The detection antibody (sheep anti-human factor IX polyclonal antibody, Abcam, catalog no. ab128048) was adsorbed at a concentration of 100 ng / ml. The secondary antibody (donkey anti-sheep IgG pAb with HRP, Abcam, catalog no. ab97125) was used at 100 ng / mL. Plates were developed using TMB substrate reagent set (BD OptEIA catalog number 555214). Optical density was assessed spectrophotometrically at 450 nm on a microplate reader (Molecular Devices i3 system) and analyzed using SoftMax pro 6.4.
[0279] Indel formation was detected, confirming that editing had occurred. NGS data demonstrated effective indel formation. Expression of hFIX from the albumin locus in NHPs was measured by ELISA and is shown in Table 6 and Figure 15. Plasma levels of hFIX reached levels previously described as therapeutically effective (George, et al., NEJM 377(23), 2215-27, 2017).
[0280] As measured, circulating hFIX protein levels were maintained over the 8-week study (see Figure 15, which shows mean levels of approximately 135, 140, 150, and 110 ng / mL on days 7, 14, 28, and 56, respectively), achieving protein levels ranging from approximately 75 ng / mL to approximately 250 ng / mL. Plasma hFIX levels were calculated using the approximately eight-fold higher specific activity for the R338L hyperactive hFIX variant (Simioni et al., NEJM 361(17), 1671-75, 2009), which reports a protein specific activity of 390 ± 28 U per milligram for hFIX-R338L and 45 ± 2.4 U per milligram relative to wild-type factor IX). Calculating the functionally normalized factor IX activity of the hyperactive factor IX variants tested in this example, the experiment achieved stable levels of human factor IX protein in NHPs over the 8-week study, corresponding to approximately 20-40% of wild-type factor IX activity (ranging from 12-67% of wild-type factor IX activity). Table 5: Editing in the liver TIFF0007820434000032.tif80153 Table 6: hFIX expression TIFF0007820434000033.tif105138
[0281] Example 13 In vivo testing of factor 9 insertions in non-human primates In this example, a study was conducted to evaluate factor 9 gene insertion and hFIX protein expression in cynomolgus monkeys following administration of CRISPR / Cas9 lipid nanoparticles (LNPs) containing various guides including ssAAV derived from P00147 and / or G009860 and various LNP components.
[0282] Indel formation was measured by NGS to confirm that editing had occurred. Total human factor IX levels were determined from plasma samples by ELISA using mouse mABs against human factor IX antibody (HTI, catalog number AHIX-5041), sheep anti-human factor IX polyclonal antibody (Abcam, catalog number ab128048), and donkey anti-sheep IgG pAb with HRP (Abcam, catalog number ab97125), as described in Example 12. More than three-fold higher human FIX protein levels than those achieved in the experiment in Example 12 were obtained from bidirectional templates using alternative CRISPR / Cas9 LNPs. In this study, ELISA assay results indicate that circulating hFIX protein levels above the normal range for human FIX levels (3-5 μg / mL, Amiral et al., Clin. Chem., 30(9), 1512-16, 1984) were achieved using G009860 in NHPs by at least the 14 and 28 day time points. Initial data show circulating human FIX protein levels of approximately 3-4 μg / mL at 14 days after a single dose, with levels (approximately 3-5 μg / mL) maintained over the first 28 days of the study. Human FIX levels were measured at the end of the study by the same method, and the data are presented in Table 25. Table 25: Serum human factor IX protein levels for Example 13 - ELISA method TIFF0007820434000034.tif59167
[0283] Circulating albumin levels were measured by ELISA and showed that baseline albumin levels were maintained at 28 days. In the study, the tested albumin levels in untreated animals varied by approximately ±15%. In treated animals, circulating albumin levels changed only slightly, did not fall out of the normal range, and levels returned to baseline within one month.
[0284] Circulating human FIX protein levels were also determined by a sandwich immunoassay with a larger dynamic range. Briefly, MSD GOLD 96-well Streptavidin SECTOR Plate (Meso Scale Diagnostics, catalog no. L15SA-1) was blocked with 1% ECL blocking agent (Sigma, GERPN2125). After tapping off the blocking solution, a biotinylated capture antibody (Sino Biological, 11503-R044) was immobilized on the plate. Recombinant human FIX protein (Enzyme Research Laboratories, HFIX1009) was used to prepare calibration standards in 0.5% ECL blocking agent. After washing, the calibration standards and plasma samples were added to the plate and incubated. After washing, a detection antibody conjugated with a sulfotag label (Haematologic Technologies, AHIX-5041) was added to the wells and incubated. After washing away any unbound detection antibody, read buffer T was applied to the wells. Plates were imaged without any additional incubation using an MSD Quick Plex SQ120 instrument, and data were analyzed using the Discovery Workbench 4.0 software package (Meso Scale Discovery). Concentrations are expressed as mean calculated concentrations in ug / m3. For samples, N=3 unless indicated with an asterisk, in which case N=2. Expression of hFIX from the albumin locus in the treated study groups, as measured by MSD ELISA, is shown in Table 26. Table 26: Serum human factor IX protein levels TIFF0007820434000035.tif67166
[0285] Example 14 Off-target analysis of albumin human guide Biochemical methods (see, e.g., Cameron et al., Nature Methods. 6, 600-606; 2017) were used to determine potential off-target genomic sites cleaved by albumin-targeting Cas9. In this experiment, 13 sgRNAs targeting human albumin and two control guides with known off-target profiles were screened using isolated HEK293 genomic DNA. The number of potential off-target sites detected using a guide concentration of 16 nM in the biochemical assay is shown in Table 27. This assay identified potential off-target sites for the sgRNAs tested. Table 27: Off-target analysis TIFF0007820434000036.tif198157
[0286] In known off-target detection assays, such as the biochemical method used above, many potential off-target sites are typically collected by design, so as to "cast a wide net" for potential sites that can be verified in other situations, such as primary cells of interest.For example, biochemical methods typically over-represent the number of potential off-target sites because the assay utilizes purified high-molecular-weight genomic DNA that does not contain cellular environment and depends on the amount of Cas9 RNP used.Therefore, the potential off-target sites identified by this method are verified by targeted sequencing of the identified potential off-target sites.
[0287] Example 15 Construction of constructs for expression of secreted or non-secreted proteins Constructs, such as bidirectional constructs, can be designed to express secreted or non-secreted proteins. For the production of secreted proteins, the constructs can include a signal sequence that assists in translocation of the polypeptide into the ER lumen. Alternatively, the constructs can utilize a signal sequence endogenous to the host cell (e.g., the endogenous albumin signal sequence when the transgene is integrated into the albumin locus of the host cell).
[0288] In contrast, constructs for the expression of non-secreted proteins can be designed so that they do not contain signal sequences and so that they do not utilize the host cell's endogenous signal sequence. Some ways in which this can be achieved include incorporating an internal ribosome entry site (IRES) sequence into the construct. IRES sequences, such as the EMCV IRES, allow translation initiation from any position within the mRNA immediately downstream from where the IRES is located. This allows the expression of proteins lacking the host cell's endogenous signal sequence from an insertion site that contains an upstream signal sequence (e.g., the signal sequence found in exon 1 of the albumin locus is not included in the expressed protein). In the absence of a signal sequence, the protein is not secreted. Examples of IRES sequences that can be used in the constructs include those derived from picornaviruses (e.g., FMDV), vermin viruses (CFFV), polioviruses (PV), encephalomyocarditis virus (ECMV), foot-and-mouth disease virus (FMDV), hepatitis C virus (HCV), classical swine fever virus (CSFV), murine leukemia virus (MLV), simian immunodeficiency virus (SIV), or cricket paralysis virus (CrPV).
[0289] An alternative approach for expressing non-secreted proteins is to include one or more self-cleaving peptides upstream of the polypeptide of interest in the construct. Self-cleaving peptides, such as 2A or 2A-like sequences, function as ribosomal skipping signals to generate multiple individual proteins from a single mRNA transcript. As shown in plasmid ID P00415 from Table 11, a self-cleaving peptide (e.g., P2A) can be used to generate a bicistronic vector expressing two transgenes (e.g., nanoluciferase and GFP). Alternatively, a self-cleaving peptide can be used to express a protein lacking the host cell's endogenous signal sequence (e.g., a 2A sequence located upstream of the protein of interest results in cleavage between the endogenous albumin signal sequence and the protein of interest). Representative 2A peptides that can be utilized are shown in Table 28. Additionally, a (GSG) residue may be added to the 5' end of the peptide to improve cleavage efficiency, as shown in Table 12. TIFF0007820434000037.tif67162
[0290] Example 16. Use of humanized albumin mice to screen guide RNAs for human F9 insertion in vivo We aimed to identify an effective guide RNA for hF9 insertion into the human albumin locus. To this end, we utilized mice in which the mouse albumin locus was replaced with the corresponding human albumin genomic sequence, including the first intron (ALB). hu / hu This allowed us to test the insertion efficiency of guide RNAs targeting the first intron of human albumin in the context of adult liver in vivo. hu / huTwo separate mouse experiments were conducted to screen a total of 11 guide RNAs, each targeting the first intron of the human albumin locus. All mice were weighed and injected via tail vein on day 0 of the experiment. Blood was collected via tail bleeding at weeks 1, 3, 4, and 6, and plasma was separated. Mice were sacrificed at week 7. Blood was collected via the vena cava, and plasma was separated. The liver and spleen were also dissected.
[0291] In the first experiment, six LNPs containing Cas9 mRNA and the following guides were prepared and tested as in Example 1: G009852, G009859, G009860, G009864, G009874, and G012764. LNPs were diluted to 0.3 mg / kg (using an average weight of 30 grams) and co-injected with AAV8 packaged with the bidirectional hF9 insertion template at a dose of 3E11 viral genomes per mouse. Five 12- to 14-week-old ALB mice were used per group. hu / hu Male mice were injected with LNP-G009874 alone. Five mice from the same cohort were injected with AAV8 packaged with the CAGG promoter operably linked to hF9, resulting in episomal expression of hF9 (3E11 viral genomes per mouse). There were three negative control groups, with three mice per group injected with buffer alone, AAV8 packaged with the bidirectional hF9 insert template alone, or LNP-G009874 alone.
[0292] In the experiment, the following LNPs containing Cas9 mRNA and the following guides were prepared and tested as in Example 1: G009860, G012764, G009844, G009857, G012752, G012753, and G012761. All were diluted to 0.3 mg / kg (using an average weight of 40 grams) and co-injected with AAV8 packaged with the bidirectional hF9 insertion template at a dose of 3E11 viral genomes per mouse. Five 30-week-old ALB mice per group were injected with LNPs containing Cas9 mRNA and the following guides: G009860, G012764, G009844, G009857, G012752, G012753, and G012761. All were diluted to 0.3 mg / kg (using an average weight of 40 grams) and co-injected with AAV8 packaged with the bidirectional hF9 insertion template at a dose of 3E11 viral genomes per mouse. hu / huMale mice were injected with LNP-G009874 alone. Five mice from the same cohort were injected with AAV8 packaged with the CAGG promoter operably linked to hF9, resulting in episomal expression of hF9 (3E11 viral genomes per mouse). There were three negative control groups, with three mice per group injected with buffer alone, AAV8 packaged with the bidirectional hF9 insert template alone, or LNP-G009874 alone.
[0293] For analysis, ELISA was performed to measure the levels of circulating hFIX in mice at each time point. For this purpose, a human Factor IX ELISA kit (ab188393) was used, and all plates were run with pooled human normal plasma from George King Bio-Medical as a positive assay control. Human Factor IX expression levels in plasma samples from each group at 6 weeks post-injection are shown in Figures 16A and 16B. Consistent with the in vitro insertion data, low to no Factor IX serum levels were detected when guide RNA G009852 was used. Consistent with the lack of adjacent PAM sequences in human albumin, Factor IX serum levels were undetectable when guide RNA G009864 was used. Liver editing was observed for groups using guide RNAs G009859, G009860, G009874, and G012764 (data not shown).
[0294] The spleen and a portion of the left lateral lobe of the entire liver were submitted for next-generation sequencing (NGS) analysis. NGS was used to assess the percentage of liver cells with insertions / deletions (indels) in the humanized albumin locus 7 weeks after injection of AAV-hF9 donor and LNP-CRISPR / Cas9. Consistent with the lack of adjacent PAM sequences in human albumin, no editing was detectable in the liver when guide RNA G009864 was used. Editing in the liver was observed for groups using guide RNAs G009859, G009860, G009874, and G012764 (data not shown).
[0295] The remaining liver was fixed in 10% neutral buffered formalin for 24 hours and then transferred to 70% ethanol. Four to five samples from separate lobes were cut and sent to HistoWisz for processing and embedding in paraffin blocks. Five-micron sections were then cut from each paraffin block and analyzed using the standard BASESCOPE™ procedure and reagents from Advanced Cell Diagnostics, and ALB when successful embedding and transfer were achieved. hu / hu BASESCOPE™ was performed on a Ventana Ultra Discovery (Roche) using a custom-designed probe targeting the unique mRNA junction formed between the human albumin signal sequence from the first intron of the albumin locus and the hF9 transgene. HALO imaging software (Indica Labs) was then used to quantify the percentage of positive cells in each sample. The average percentage of positive cells across multiple lobes for each animal was then correlated with hFIX levels in serum at week 7. The results are shown in Figure 17 and Table 29. Week 7 serum levels of hALB-hFIX mRNA and % positive cells were strongly correlated (r=0.89, R 2 =0.79). Table 29. Week 7 hFIX and BASESCOPE™ Data. TIFF0007820434000038.tif178167 Human albumin intron 1: (SEQ ID NO: 1) GTAAGAAATCCATTTTTCTATTGTTCAACTTTTATTCTATTTTCCCAGTAAAATAAAGTTTTAGTAAACTCTGCATCTTTAAAGAATTATTTTGGCATTTATTTCTAAAATGGCATAGTATTTTGTATTTGTGAAGTCTTACAAGGTTATCTTATTAATAAAATTCAAACATCCTAG GTAAAAAAAAAAAAAGGTCAGAATTGTTTAGTGACTGTAATTTTCTTTTGCGCACTAAGGAAAGTGCAAAGTAACTTAGAGTGACTGAAACTTCACAGAATAGGGTTGAAGATTGAATTCATAACTATCCCAAAGACCTATCCATTGCACTATGCTTTATTTAAAAACCACAAAACC TGTGCTGTTGATCTCATAAATAGAACTTGTATTTATATTTATTTTCATTTTAGTCTGTCTTCTTGGTTGCTGTTGATAGACACTAAAAGAGTATTAGATATTATCTAAGTTTGAATATAAGGCTATAAATATTTAATAATTTTTAAAATAGTATTCTTGGTAATTGAATTATTCTTC TGTTTAAAGGCAGAAGAAATAATTGAACATCATCCTGAGTTTTTCTGTAGGAATCAGAGCCCAATATTTTGAAACAAATGCATAATCTAAGTCAAATGGAAAGAAATATAAAAAGTAACATTATTACTTCTTGTTTTCTTCAGTATTTAACAATCCTTTTTTTTCTTCCCTTGCCCAG Table 7. Mouse albumin guide RNAs TIFF0007820434000039.tif152164Table 8. Mouse albumin sgRNA and modification patterns TIFF0007820434000040.tif233167TIFF0007820434000041.tif245167TIFF0007820434000042.tif234168Table 9. Cynomolgus monkey albumin guide RNA TIFF0007820434000043.tif219166 Table 10: Cynomolgus sgRNAs and modification patterns TIFF0007820434000044.tif244166TIFF0007820434000045.tif248161TIFF0007820434000046.tif248161TIFF0007820434000047.tif248161TIFF0007820434000048.tif248161TIFF0007820434000049.tif248160TIFF0007820434000050.tif78166Table 11: Vector components and sequences TIFF0007820434000051.tif202165TIFF0007820434000052.tif581655'ITR sequence (SEQ ID NO: 263): TTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT Mouse albumin splice acceptor (first orientation) (SEQ ID NO: 264): TAGGTCAGTGAAGAGAAGAACAAAAAGCAGCATATTACAGTTAGTTGTCTTCATCAATCTTTAAATATGTTGTGGTTTTTCTCTCCCTGTTTCCACAG Human Factor IX (R338L), first orientation (SEQ ID NO: 265): PolyA (first orientation) (SEQ ID NO: 266): CCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTC ATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCC PolyA (second orientation) (SEQ ID NO: 267): AAAAAACCTCCCACACCTCCCCCCTGAACCTGAAACATAAAATGAATGCAATTGTTGTTGTTAACTTGTTTATTGCAGCTTATAATGGTTACAAATAAAGCAATAGCATCACAAATTTCACAAATAAAGCATTTTTTTCACTGCATTCTAGTTGTGGTTTGTCCAAACTCATCAATGTATCTTATCATGTCTG Human Factor IX (R338L), second orientation (SEQ ID NO: 268): Mouse albumin splice acceptor (second orientation) (SEQ ID NO: 269): CTGTGGAAACAGGGAGAGAAAAACCACACAACATATTTAAAGATTGATGAAGACAACTAACTGTAATATGCTGCTTTTTGTTCTTCTCTTCACTGACCTA 3' ITR sequence (SEQ ID NO: 270): AGGAACCCCTAGTGATGGAGTTGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGGGAGTGGCCAA Human Factor IX splice acceptor (first orientation) (SEQ ID NO: 271): GATTATTTGGATTAAAAAACAAAGACTTTCTTAAGAGATGTAAAATTTTCATGATGTTTTCTTTTTTGCTAAAACTAAAGAATTATTCTTTTACATTTCAG Human Factor IX (R338L)-HiBit (first orientation) (SEQ ID NO: 272): Human Factor IX (R338L)-HiBit (second orientation) (SEQ ID NO: 273): Human Factor IX splice acceptor (second orientation) (SEQ ID NO: 274): CTGAAATGTAAAAGAATAATTCTTTAGTTTTAGCAAAAAAGAAAACATCATGAAAATTTTACATCTCTTAAGAAAGTCTTTGTTTTTAATCCAAATAATC Nluc-P2A-GFP (first orientation) (SEQ ID NO: 275): Nluc-P2A-GFP (second orientation) (SEQ ID NO: 276): P00147 complete sequence (ITR to ITR): (SEQ ID NO: 277) P00411 complete sequence (ITR to ITR): (SEQ ID NO: 278) P00415 complete sequence (ITR to ITR): (SEQ ID NO: 279) P00418 complete sequence (ITR to ITR): (SEQ ID NO: 280) P00123 complete sequence (ITR to ITR): (SEQ ID NO: 281) TAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCCACTAGTCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGA P00204 complete sequence (ITR to ITR): (SEQ ID NO: 282) CTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCCCTTAGGTGGTTATATTATTGATATATTTTTGGTATCTTTGATGACAATAATGGGGGATTTTGAAAGCTTAGCTTTAAATTTCTTTTAATTAAAAAAAAATGCTAGGCAGAATGACTCAAATTACGTTGGATACAGTTGAATTTATTACGGTCTCATAGGGCCTGCCTGCTCGACCATGCTATACTAAAAATTAAAAGTGTACTAGTCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGA Complete sequence of P00353 (from ITR to ITR): (SEQ ID NO:283) P00354 complete sequence (ITR to ITR): (SEQ ID NO: 284) P00350: 300 / 600bp HA F9 construct (for G551) (SEQ ID NO: 285) CTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGCAGAGGGGAGGATTGGGAAGACATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCCTTAGGTGGTTATATTATTGATATATTTTTGGTATCTTTGAATGACAAATGGGGATTTGAAAGCTTAGCTTTAAATTTCTTTAATTAAAAAAAAATGCTAGGCAGAATGACTCAAATTACGTTGG ATACAGTTGAATTATTACGGTCTCATAGGGCCTGCCTGCTCGACCATGCTATACTAAAAATTAAAAAGTGTGTGTTACTAATTTTATAAATGGAGTTTCCATTTATATTTACCTTATTTCTTATTTACCATTGTCTTAGATATTTACAAACATGACAGAAACACTAAAAGATCTAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCAGCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGATGGCCAA P00356:300 / 2000bpのHA F9 construct(G551) (sequence number 286) TCAAAGCCAATCATGACTCCATCACTTAAGGCCCCGGGAACACTGTGGCAGAGGGCAGCAGAGAGATTGATAAAGCCAGGGTGATGGGAATTTTCTGTGGGACTCCATTTCATAGTAATTGCAGAAGCTACAATACACTCAAAAAGTCTCACCACATGACTGCCCAAATGGGAGCTTGACAGTGACAGTGACAGTAGATATGCCAAAGTGGATGAGGGAAAGACCACAAGAGCTAAACCCTGTAAAAAGAACTGTAGGCAACTAAGGAATGCAGAGAGAAAGATCTAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAA P00362: HA F9 construct of 300 / 1500bp (for G551) (SEQ ID NO: 287) Cas9 ORF (SEQ ID NO: 703) CCGGAAGATAACGAACAGAAGCAGCTTTTCGTGGAGCAGCACAAGCATTATCTGGATGAAATCATCGAACAAATCTCCGAGTTTTCAAAGCGCGTGATCCTCGCCGACGCCAACCTCGACAAAGTCCTGTCGGCCTACAATAAGCATAGAGATAAGCCGATCAGAGAACAGGCCGAGAACATTATCCACTTGTTCACCCTGACTAACCTGGGAGCCCCAGCCGCCTTCAAGTACTTCGATACTACTATCGATCGCAAAAGATACACGTCCACCAAGGAAGTTCTGGACGCGACCCTGATCCACCAAAGCATCACTGGACTCTACGAAACTAGGATCGATCTGTCGCAGCTGGGTGGCGATU-dep Cas9 ORF (SEQ ID NO: 704) AGAAAGAGAATGCTGGCAAGCGCAGGAGAACTGCAGAAGGGAAACGAACTGGCACTGCCGAGCAAGTACGTCAACTTCCTGTACCTGGCAAGCCACTACGAAAAGCTGAAGGGAAGCCCGGAAGACAACGAACAGAAGCAGCTGTTCGTCGAACAGCACAAGCACTACCTGGACGAAATCATCGAACAGATCAGCGAATTCAGCAAGAGAGTCATCCTGGCAGACGCAAACCTGGACAAGGTCCTGAGCGCATACAACAAGCACAGAGACAAGCCGATCAGAGAACAGGCAGAAAACATCATCCACCTGTTCACACTGACAAACCTGGGAGCACCGGCAGCATTCAAGTACTTCGACACAACAATCGACAGAAAGAGATACACAAGCACAAAGGAAGTCCTGGACGCAACACTGATCCACCAGAGCATCACAGGACTGTACGAAACAAGAATCGACCTGAGCCAGCTGGGAGGAGACGGAGGAGGAAGCCCGAAGAAGAAGAGAAAGGTCTAG mRNA containing U dep Cas9 (SEQ ID NO: 705) GUACCUGGCAAGCCACUACGAAAAGCUGAAGGGAAGCCCGGAAGACAACGAACAGAAGCAGCUGUUCGUCGAACAGCACAAGCACUACCUGGACGAAAUCAUCGAACAGAUCAGCGAAUUCAGCAAGAGAGUCAUCCUGGCAGACGCAAACCUGGACAAGGUCCUGAGCGCAUACAACAAGCACAGAGACAAGCCGAUCAGAGAACAGGCAGAAAACAUCAUCCACCUGUUCACACUGACAAACCUGGGAGCACCGGCAGCAUUCAAGUACUUCGACACAACAAUCGACAGAAAGAGAUACACAAGCACAAAGGAAGUCCUGGACGCAACACUGAUCCACCAGAGCAUCACAGGACUGUACGAAACAAGAAUCGACCUGAGCCAGCUGGGAGGAGACGGAGGAGGAAGCCCGAAGAAGAAGAGAAAGGUCUAGCUAGCCAUCACAUUUAAAAGCAUCUCAGCCUACCAUGAGAAUAAGAGAAAGAAAAUGAAGAUCAAUAGCUUAUUCAUCUCUUUUUCUUUUUCGUUGGUGUAAAGCCAACACCCUGUCUAAAAAACAUAAAUUUCUUUAAUCAUUUUGCCUCUUUUCUCUGUGCUUCAAUUAAUAAAAAAUGGAAAGAACCUCGAGAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA [Sequence Listing] <110> INTELLIA THERAPEUTICS INC. REGENERON PHARMACEUTICALS, INC. <120> COMPOSITIONS AND METHODS FOR TRANSGENE EXPRESSION FROM AN ALBUMIN LOCUS <130> PA24-188 <141> 2019-10-18 <150> US 62 / 840,346 <151> 2019-04-29 <150> US 62 / 747,402 <151> 2018-10-18 <160> 1138 <170> PatentIn version 3.5 <210> 1 <211> 709 <212> DNA <213> Homo sapiens <400> 1 gtaagaaatc cattttcta ttgttcaact tttattctat tttcccagta aaataaagtt 60 ttagtaaact ctgcatcttt aaagaattat tttggcattt atttctaaaa tggcatagta 120 ttttgtattt gtgaagtctt acaaggttat cttattaata aaattcaaac atcctaggta 180 aaaaaaaaaa aaggtcagaa ttgtttagtg actgtaattt tcttttgcgc actaaggaaa 240 gtgcaaagta acttagagtg actgaaactt cacagaatag ggttgaagat tgaattcata 300 actatcccaa agacctatcc attgcactat gctttattta aaaaccacaa aacctgtgct 360 gttgatctca taaatagaac ttgtatttat atttattttc attttagtct gtcttcttgg 420 ttgctgttga tagacactaa aagagtatta gatattatct aagtttgaat ataaggctat 480 aaatatttaa taatttttaa aatagtattc ttggtaattg aattattctt ctgtttaaag 540 gcagaagaaa taattgaaca tcatcctgag tttttctgta ggaatcagag cccaatattt 600 tgaaacaaat gcataatcta agtcaaatgg aaagaaatat aaaaagtaac attattactt 660 cttgttttct tcagtattta acaatccttt tttttcttcc cttgcccag 709 <210> 2 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 2 gagcaaccuc acucuugucu 20 <210> 3 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 3 augcauuugu uucaaaauau 20 <210> 4 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 4 ugcauuuguu ucaaaauauu 20 <210> 5 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 5 auuuaugaga ucaacagcac 20 <210> 6 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 6 gaucaacagc acagguuuug 20 <210> 7 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 7 uuaaauaaag cauagugcaa 20 <210> 8 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 8 uaaagcauag ugcaauggau 20 <210> 9 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 9 uagugcaaug gauaggucuu 20 <210> 10 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 10 uacuaaaacu uuauuuuacu 20 <210> 11 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 11 aaaguugaac aauagaaaaa 20 <210> 12 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 12 aaugcauaau cuaagucaaa 20 <210> 13 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 13 uaauaaaauu caaacauccu 20 <210> 14 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 14 gcaucuuuaa agaauuauuu 20 <210> 15 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 15 uuuggcauuu auuucuaaaa 20 <210> 16 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 16 uguauuugug aagucuuaca 20 <210> 17 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 17 uccuagguaa aaaaaaaaaa 20 <210> 18 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 18 uaauuuucuu uugcgcacua 20 <210> 19 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 19 ugacugaaac uucacagaau 20 <210> 20 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 20 gacugaaacu ucacagaaua 20 <210> 21 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 21 uucauuuuag ucugucuucu 20 <210> 22 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 22 auuaucuaag uuugaauaua 20 <210> 23 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 23 aauuuuuaaa auaguauucu 20 <210> 24 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 24 ugaauuauuc uucuguuuaa 20 <210> 25 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 25 aucauccuga guuuuucugu 20 <210> 26 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 26 uuacuaaaac uuuauuuuac 20 <210> 27 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 27 accuuuuuuu uuuuuuaccu 20 <210> 28 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 28 agugcaaugg auaggucuuu 20 <210> 29 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 29 ugauuccuac agaaaaacuc 20 <210> 30 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 30 ugggcaaggg aagaaaaaaa 20 <210> 31 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 31 ccucacucuu gucugggcaa 20 <210> 32 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 32 accucacucu ugucugggca 20 <210> 33 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 33 ugagcaaccu cacucuuguc 20 <210> 34 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 34 gagcaaccuc acucuugucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 35 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 35 augcauuugu uucaaaauau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 36 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 36 ugcauuuguu ucaaaauauu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 37 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 37 auuuaugaga ucaacagcac guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 38 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 38 gaucaacagc acagguuuug guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 39 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 39 uuaaauaaag cauagugcaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 40 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 40 uaaagcauag ugcaauggau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 41 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 41 uagugcaaug gauaggucuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 42 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 42 uacuaaaacu uuauuuuacu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 43 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 43 aaaguugaac aauagaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 44 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 44 aaugcauaau cuaagucaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 45 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 45 uaauaaaauu caaacauccu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 46 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 46 gcaucuuuaa agaauuauuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 47 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 47 uuuggcauuu auuucuaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 48 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 48 uguauuugug aagucuuaca guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 49 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 49 uccuagguaa aaaaaaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 50 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 50 uaauuuucuu uugcgcacua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 51 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 51 ugacugaaac uucacagaau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 52 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 52 gacugaaacu ucacagaaua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 53 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 53 uucauuuuag ucugucuucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 54 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 54 auuaucuaag uuugaauaua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 55 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 55 aauuuuuaaa auaguauucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 56 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 56 ugaauuauuc uucuguuuaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 57 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 57 aucauccuga guuuuucugu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 58 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 58 uuacuaaaac uuuauuuuac guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 59 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 59 accuuuuuuu uuuuuuaccu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 60 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 60 agugcaaugg auaggucuuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 61 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 61 ugauuccuac agaaaaacuc guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 62 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 62 ugggcaaggg aagaaaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 63 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 63 ccucacucuu gucugggcaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 64 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 64 accucacucu ugucugggca guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 65 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 65 ugagcaaccu cacucuuguc guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 66 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 66 gagcaaccuc acucuugucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 67 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 67 augcauuugu uucaaaauau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 68 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 68 ugcauuuguu ucaaaauauu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 69 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 69 auuuaugaga ucaacagcac guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 70 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 70 gaucaacagc acagguuuug guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 71 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 71 uuaaauaaag cauagugcaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 72 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 72 uaaagcauag ugcaauggau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 73 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 73 uagugcaaug gauaggucuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 74 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 74 uacuaaaacu uuauuuuacu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 75 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 75 aaaguugaac aauagaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 76 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 76 aaugcauaau cuaagucaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 77 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 77 uaauaaaauu caaacauccu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 78 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 78 gcaucuuuaa agaauuauuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 79 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 79 uuuggcauuu auuucuaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 80 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 80 uguauuugug aagucuuaca guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 81 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 81 uccuagguaa aaaaaaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 82 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 82 uaauuuucuu uugcgcacua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 83 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 83 ugacugaaac uucacagaau guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 84 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 84 gacugaaacu ucacagaaua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 85 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 85 uucauuuuag ucugucuucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 86 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 86 auuaucuaag uuugaauaua guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 87 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 87 aauuuuuaaa auaguauucu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 88 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 88 ugaauuauuc uucuguuuaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 89 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 89 aucauccuga guuuuucugu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 90 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 90 uuacuaaaac uuuauuuuac guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 91 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 91 accuuuuuuu uuuuuuaccu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 92 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 92 agugcaaugg auaggucuuu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 93 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 93 ugauuccuac agaaaaacuc guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 94 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 94 ugggcaaggg aagaaaaaaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 95 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 95 ccucacucuu gucugggcaa guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 96 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 96 accucacucu ugucugggca guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 97 <211> 100 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 97 ugagcaaccu cacucuuguc guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 98 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 98 auuugcaucu gagaacccuu 20 <210> 99 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 99 aucgggaacu ggcaucuuca 20 <210> 100 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 100 guuacaggaa aaucugaagg 20 <210> 101 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 101 gaucgggaac uggcaucuuc 20 <210> 102 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 102 ugcaucugag aacccuuagg 20 <210> 103 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 103 cacucuuguc uguggaaaca 20 <210> 104 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 104 aucguuacag gaaaaucuga 20 <210> 105 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 105 gcaucuucag ggaguagcuu 20 <210> 106 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 106 caaucuuuaa auauguugug 20 <210> 107 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 107 ucacucuugu cuguggaaac 20 <210> 108 <211> 20 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Desc...
Claims
1. A single gRNA (sgRNA) comprising the sequence mG*mA*mG*CAACCUCACUCUUGUCUGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 66), "mA", "mC", "mU", and "mG" represent nucleotides substituted by 2'-O-Me, and * indicates a phosphorothioate (PS) linkage in the sgRNA.
2. i) the sgRNA of claim 1; ii) an RNA-guided DNA binder; and iii) Constructs containing nucleic acids encoding heterologous polypeptides 10. A composition comprising the sgRNA of claim 1 for use in a method for inserting a nucleic acid encoding a heterologous polypeptide into an albumin locus of a cell or cell population, comprising delivering an sgRNA to the cell or cell population, thereby inserting the nucleic acid encoding the heterologous polypeptide into the albumin locus of the cell or cell population.
3. The composition for use according to claim 2, wherein the RNA-guided DNA binding agent is a Cas nuclease or a nucleic acid encoding the Cas nuclease.
4. 4. The composition for use of claim 3, wherein the Cas nuclease is a class 2 Cas nuclease.
5. 5. The composition for use according to claim 3 or 4, wherein the Cas nuclease is Cas9.
6. The composition for use according to any one of claims 2 to 5, wherein the construct comprises a polyadenylation signal sequence.
7. The composition for use according to any one of claims 2 to 6, wherein the construct comprises a splice acceptor site.
8. The composition for use according to any one of claims 2 to 7, wherein said heterologous polypeptide is a secreted polypeptide or an intracellular polypeptide.
9. The composition for use according to any one of claims 2 to 8, wherein the RNA-guided DNA-binding agent is delivered as a nucleic acid encoding the RNA-guided DNA-binding agent.
10. 10. The composition for use according to any one of claims 2 to 9, wherein the construct comprising the nucleic acid encoding the sgRNA, the RNA-guided DNA-binding agent, and / or the heterologous polypeptide is delivered in a vector and / or lipid nanoparticle.
11. The composition for use according to claim 10, wherein the vector is a viral vector.
12. 12. The composition for use of claim 11, wherein the viral vector is selected from the group consisting of an adeno-associated viral (AAV) vector, an adenoviral vector, a retroviral vector, and a lentiviral vector.
13. 13. The composition for use of any one of claims 2 to 12, wherein the sgRNA mediates target-specific cleavage by the RNA-guided DNA-binding agent, resulting in insertion of the nucleic acid encoding the heterologous polypeptide within intron 1 of the albumin gene.
14. The composition for use according to claim 13, wherein the composition is for effecting an insertion rate of at least 2% of the nucleic acid encoding the heterologous polypeptide in the cell population by target-specific cleavage.
15. The composition for use according to any one of claims 2 to 14, wherein the cell or cell population is a liver cell or cell population.
16. A composition comprising a single gRNA (sgRNA) comprising the sequence mG*mA*mG*CAACCUCACUCUUGUCUGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 66), "mA", "mC", "mU", and "mG" represent nucleotides substituted by 2'-O-Me, * indicates a phosphorothioate (PS) bond; i) the single gRNA (sgRNA); ii) an RNA-guided DNA binder; and iii) Constructs containing nucleic acids encoding heterologous polypeptides 2. The composition of claim 1, wherein the nucleic acid encoding the heterologous polypeptide is inserted into the albumin locus of a cell or cell population in the subject by administering to the subject a nucleic acid encoding the heterologous polypeptide.
17. 17. The composition for use of claim 16, wherein the RNA-guided DNA binding agent is Cas9 or a nucleic acid encoding Cas9.
18. 18. The composition for use according to claim 16 or 17, wherein the RNA-guided DNA binding agent is administered as a nucleic acid encoding the RNA-guided DNA binding agent.
19. 19. The composition for use of claim 18, wherein the nucleic acid encoding the RNA-guided DNA binding agent is mRNA.
20. 20. The composition for use of claim 19, wherein the mRNA is a modified mRNA.
21. The composition for use according to any one of claims 16 to 20, wherein the RNA-guided DNA binding agent is a Cas nuclease or a nucleic acid encoding the Cas nuclease.
22. 22. The composition for use of claim 21, wherein the Cas nuclease is Cas9.
23. 23. The composition for use of claim 21 or 22, wherein the Cas nuclease is a S. pyogenes Cas9 nuclease.
24. The composition for use according to any one of claims 16 to 23, wherein the construct comprises a polyadenylation signal sequence.
25. The composition for use according to any one of claims 16 to 24, wherein the construct comprises a splice acceptor site.
26. The composition for use according to any one of claims 16 to 25, wherein the construct does not comprise homology arms.
27. The composition for use according to any one of claims 16 to 26, wherein the sgRNA is administered in a vector and / or lipid nanoparticle.
28. The composition for use according to any one of claims 16 to 27, wherein the RNA-guided DNA binding agent is administered in a vector and / or lipid nanoparticle.
29. The composition for use according to any one of claims 16 to 28, wherein the construct comprising the nucleic acid encoding the heterologous polypeptide is delivered in a vector and / or lipid nanoparticle.
30. The composition for use according to any one of claims 27 to 29, wherein the vector is a viral vector.
31. 31. The composition for use of claim 30, wherein the viral vector is selected from the group consisting of an adeno-associated viral (AAV) vector, an adenoviral vector, a retroviral vector, and a lentiviral vector.
32. 32. The composition for use of any one of claims 16 to 31, wherein the sgRNA mediates target-specific cleavage by the RNA-guided DNA-binding agent, resulting in insertion of the nucleic acid encoding the heterologous polypeptide within intron 1 of the albumin gene.
33. 33. The composition for use of claim 32, wherein the composition is for achieving an insertion rate of at least 2% of the nucleic acid encoding the heterologous polypeptide in the cell population by target-specific cleavage.
34. The composition for use according to any one of claims 16 to 33, wherein the sgRNA and / or the RNA-guided DNA binding agent is administered in a lipid nanoparticle.
35. 35. The composition for use of any one of claims 16 to 34, wherein the sgRNA and the RNA-guided DNA-binding agent are administered as a ribonucleoprotein (RNP).
36. A vector comprising the sgRNA of claim 1.
37. The vector of claim 36 , wherein the vector is a viral vector.
38. A lipid nanoparticle comprising the sgRNA of claim 1.
39. A cell comprising the sgRNA of claim 1.
40. 3. The composition of claim 2 for use in a method for inserting a nucleic acid encoding a heterologous polypeptide into the albumin locus of a cell or cell population.
41. 3. The composition of claim 2 for use in a method for inserting a nucleic acid encoding a heterologous polypeptide into the albumin locus of a cell or population of cells in a subject.
42. i) the sgRNA of claim 1; ii) an RNA-guided DNA binder; and iii) Constructs containing nucleic acids encoding heterologous polypeptides A kit or composition comprising:
Citation Information
Patent Citations
Method for obtaining sheep with different hair colors on basis of CRISPR / Cas9 and sgRNA of targeted ASIP gene
CN105950626A
Methods and compositions for controlling transgene expression
JP2014526279A
Methods and compositions for the treatment of lysosomal storage disorders
JP2015527881A
Optimized crispr-cas double nickase system, methods and compositions for sequence engineering
JP2016521994A
Crispr-CAS systems and methods for altering expression of gene products
US20150203872A1