Non-human animals comprising a humanized albumin locus
Non-human animals with a humanized albumin locus address the need for testing human albumin-targeting agents by accurately expressing human albumin, enabling effective evaluation and optimization of these agents in vivo.
Patent Information
- Application Number
- JP2021568885
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-17
- Filing Date
- 2020-06-05
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2040-06-05
AI Technical Summary
There is a need for suitable non-human animals that provide a true human genomic DNA target or a close approximation thereof for human albumin targeting reagents to enable testing of efficacy and mechanism of action in living animals, as well as pharmacokinetic and pharmacodynamic studies, particularly for human albumin targeting agents.
The development of non-human animals with a humanized albumin locus, where segments of the endogenous albumin locus are deleted and replaced with corresponding human albumin sequences, allowing for the expression of human albumin proteins, enabling the evaluation of human albumin-targeting reagents in vivo.
The humanized albumin locus in non-human animals accurately reflects human albumin expression and function, facilitating effective testing and optimization of human albumin-targeting reagents, such as CRISPR/Cas9 agents, by providing a more relevant model for pharmacokinetic and pharmacodynamic studies.
Smart Images

Figure 0007698587000010 
Figure 0007698587000011 
Figure 0007698587000012
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Application No. 62 / 858,589, filed Jun. 7, 2019, and U.S. Application No. 62 / 916,666, filed Oct. 17, 2019, each of which is hereby incorporated by reference in its entirety for all purposes.
[0002] Reference to a Sequence Listing Submitted as a Text File via EFS - WEB The sequence listing described in the file 548157SEQLIST.txt is 158 kilobytes, was created on May 27, 2020, and is hereby incorporated by reference in its entirety.
Background Art
[0003] Gene therapy is a promising therapeutic approach for several human diseases. One approach to gene therapy is to insert a transgene into a safe - harbor locus within the genome. Safe - harbor loci include chromosomal loci where a transgene or other exogenous nucleic acid insert can be stably and reliably expressed in all tissues of interest without overtly altering cellular behavior or phenotype. Often, a safe - harbor locus is a locus where the expression of the inserted gene sequence is not disrupted by read - through expression from adjacent genes. For example, a safe - harbor locus can include chromosomal loci where exogenous DNA can be integrated and function in a predictable manner without adversely affecting the structure or expression of endogenous genes. A safe - harbor locus can include extragenic or intragenic regions, for example, intragenic loci that can be disrupted without being essential, unnecessary, or resulting in overt phenotypic consequences.
[0004] An example of a safe harbor locus is albumin. However, there remains a need for suitable non-human animals that provide a true human genomic DNA target or a close approximation thereof for human albumin targeting reagents at the endogenous albumin locus in vivo, thereby enabling testing of the efficacy and mechanism of action of such agents in living animals as well as pharmacokinetic and pharmacodynamic studies in an environment that is the only version of albumin in which the humanized gene exists. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] Provided are non-human animals comprising a humanized albumin (ALB) locus, as well as methods of making and using such non-human animals. Also provided are non-human animal genomes or cells comprising a humanized albumin (ALB) locus. Also provided is a humanized albumin gene.
[0006] In one aspect, provided is a non-human animal genome, non-human animal cell, or non-human animal comprising a humanized albumin (ALB) locus. Such non-human animal genomes, non-human animal cells, or non-human animals may comprise a humanized endogenous albumin locus within their genomes, wherein a segment of the endogenous albumin locus has been deleted and replaced with a corresponding human albumin sequence.
[0007] In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus encodes a protein comprising a human serum albumin peptide. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus encodes a protein comprising a human albumin propeptide. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus encodes a protein comprising a human albumin signal peptide.
[0008] In some such non-human animal genomes, non-human animal cells, or non-human animals, regions of the endogenous albumin locus that include both coding and non-coding sequences are deleted and replaced with the corresponding human albumin sequence that includes both coding and non-coding sequences. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus includes the endogenous albumin promoter, and the human albumin sequence is operably linked to the endogenous albumin promoter. In some such non-human animal genomes, non-human animal cells, or non-human animals, at least one intron and at least one exon of the endogenous albumin locus are deleted and replaced with the corresponding human albumin sequence.
[0009] In some such non-human animal genomes, non-human animal cells, or non-human animals, the entire albumin coding sequence of the endogenous albumin locus is deleted and replaced with the corresponding human albumin sequence. Optionally, the region of the endogenous albumin locus from the start codon to the stop codon is deleted and replaced with the corresponding human albumin sequence.
[0010] In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus includes the human albumin 3' untranslated region. In some such non-human animal genomes, non-human animal cells, and non-human animals, the endogenous albumin 5' untranslated region is not deleted and not replaced with the corresponding human albumin sequence.
[0011] In some such non-human animal genomes, non-human animal cells, or non-human animals, the region of the endogenous albumin locus from the start codon to the stop codon is deleted and replaced with a human albumin sequence containing the corresponding human albumin sequence and the human albumin 3' untranslated region, and the endogenous albumin 5' untranslated region is not deleted, not replaced with the corresponding human albumin sequence, and the endogenous albumin promoter is not deleted and not replaced with the corresponding human albumin sequence.
[0012] In some such non-human animal genomes, non-human animal cells, or non-human animals, the human albumin sequence at the humanized endogenous albumin locus comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 35. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus encodes a protein that comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 5. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus comprises a coding sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 13. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus comprises a sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence set forth in SEQ ID NO: 17 or 18. In some such non-human animal genomes, non-human animal cells, or non-human animals, the human albumin sequence at the humanized endogenous albumin locus comprises a sequence that is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the sequence set forth in SEQ ID NO: 35. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus encodes a protein that comprises a sequence that is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the sequence set forth in SEQ ID NO: 5. In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus comprises a coding sequence that is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the sequence set forth in SEQ ID NO: 13.In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus comprises a sequence that is at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to the sequence set forth in SEQ ID NO: 17 or 18.
[0013] In some such non-human animal genomes, non-human animal cells, or non-human animals, the humanized endogenous albumin locus does not contain a selection cassette or a reporter gene.
[0014] In some such non-human animal genomes, non-human animal cells, or non-human animals, the non-human animal is homozygous for the humanized endogenous albumin locus. In some such non-human animal genomes, non-human animal cells, or non-human animals, the non-human animal contains the humanized endogenous albumin locus in its germline.
[0015] In some such non-human animal genomes, non-human animal cells, or non-human animals, the non-human animal is a mammal. Optionally, the non-human animal is a rat or a mouse. Optionally, the non-human animal is a mouse.
[0016] In some such non-human animal genomes, non-human animal cells, or non-human animals, the non-human animal comprises a serum albumin level of at least about 10 mg / mL. In some such non-human animal genomes, non-human animal cells, or non-human animals, the serum albumin level in the non-human animal is at least as high as the serum albumin level in a control non-human animal containing a wild-type albumin locus.
[0017] In some such non-human animal genomes, non-human animal cells, or non-human animals, the genome, cell, or animal is heterozygous for the humanized endogenous albumin locus. In some such non-human animal genomes, non-human animal cells, or non-human animals, the genome, cell, or animal is homozygous for the humanized endogenous albumin locus. In some such non-human animal genomes, non-human animal cells, or non-human animals, the genome, cell, or animal further comprises a coding sequence of an exogenous protein integrated into at least one allele of the humanized endogenous albumin locus in one or more cells of the non-human animal. Optionally, the coding sequence of the exogenous protein is integrated into intron 1 of at least one allele of the humanized endogenous albumin locus (e.g., in one or more cells of the non-human animal). In some such non-human animal genomes, non-human animal cells, or non-human animals, the genome, cell, or animal further comprises an inactivated endogenous locus that is not the endogenous albumin locus. Optionally, the non-human animal genome, non-human animal cell, or non-human animal further comprises a coding sequence of an exogenous protein integrated into at least one allele of the humanized endogenous albumin locus (e.g., in one or more cells of the non-human animal), and the exogenous protein replaces the function of the inactivated endogenous locus. Optionally, the inactivated endogenous locus is the inactivated F9 locus.
[0018] In another aspect, a targeting vector for generating the above non-human animal genome, non-human animal cell, or non-human animal is provided. Such a targeting vector is for generating a humanized endogenous albumin locus, in which a segment of the endogenous albumin locus is deleted and replaced with a corresponding human albumin sequence, and the targeting vector comprises an insertion nucleic acid comprising a corresponding human albumin sequence flanked by a 5' homology arm targeting the 5' target sequence of the endogenous albumin locus and a 3' homology arm targeting the 3' target sequence of the endogenous albumin locus.
[0019] In another aspect, a method for evaluating the activity of a human albumin-targeting reagent in vivo is provided. Some such methods include (a) administering a human albumin-targeting reagent to the non-human animal described above, and (b) evaluating the activity of the human albumin-targeting reagent in the non-human animal.
[0020] In some such methods, administration includes adeno-associated virus (AAV)-mediated delivery, lipid nanoparticle (LNP)-mediated delivery, or hydrodynamic delivery (HDD). Optionally, administration includes LNP-mediated delivery. Optionally, the LNP dose is from about 0.1 mg / kg to about 2 mg / kg. In some such methods, administration includes AAV8-mediated delivery.
[0021] In some such methods, step (b) includes isolating the liver from the non-human animal and evaluating the activity of the human albumin-targeting reagent in the liver.
[0022] In some such methods, the human albumin-targeting reagent is a genome editing agent, and the evaluation includes evaluating the modification of the humanized endogenous albumin locus. Optionally, the evaluation includes measuring the frequency of insertions or deletions within the humanized endogenous albumin locus.
[0023] In some such methods, the evaluation includes measuring the expression of albumin messenger RNA encoded by the humanized endogenous albumin locus. In some such methods, the evaluation includes measuring the expression of albumin protein encoded by the humanized endogenous albumin locus. Optionally, evaluating the expression of albumin protein includes measuring the serum level of albumin protein in the non-human animal. Optionally, evaluating the expression of albumin protein includes measuring the expression of albumin protein in the liver of the non-human animal.
[0024] In some such methods, the human albumin targeting reagent comprises a nuclease agent designed to target a region of the human albumin gene. In some such methods, the human albumin targeting reagent comprises a nuclease agent or a nucleic acid encoding a nuclease agent, and the nuclease agent is designed to target a region of the human albumin gene. Optionally, the nuclease agent comprises a Cas protein and a guide RNA designed to target a guide RNA target sequence of the human albumin gene. Optionally, the guide RNA target sequence is in intron 1 of the human albumin gene. Optionally, the Cas protein is a Cas9 protein.
[0025] In some such methods, the human albumin targeting reagent comprises an exogenous donor nucleic acid, the exogenous donor nucleic acid is designed to target the human albumin gene, and optionally, the exogenous donor nucleic acid is delivered via AAV. Optionally, the exogenous donor nucleic acid is a single-stranded oligodeoxynucleotide (ssODN). Optionally, the exogenous donor nucleic acid can be inserted into the humanized albumin locus by non-homologous end joining.
[0026] In some methods, the exogenous donor nucleic acid does not contain homology arms. In some methods, the exogenous donor nucleic acid comprises an insertion nucleic acid flanked by a 5' homology arm targeting a 5' target sequence of the humanized endogenous albumin locus and a 3' homology arm targeting a 3' target sequence of the humanized endogenous albumin locus. Optionally, each of the 5' target sequence and the 3' target sequence comprises a segment of intron 1 of the human albumin gene.
[0027] In some such methods, the exogenous donor nucleic acid encodes an exogenous protein. Optionally, the protein encoded by the humanized endogenous albumin locus targeted by the exogenous donor nucleic acid is a heterologous protein comprising a human albumin signal peptide fused to the exogenous protein. Optionally, the exogenous protein is a factor IX protein. Optionally, the evaluation includes measuring the serum level of the factor IX protein in a non-human animal and / or evaluating the activated partial thromboplastin time, or performing a thrombin generation assay. Optionally, the non-human animal further comprises an inactivated F9 locus, and the evaluation includes measuring the serum level of the factor IX protein in the non-human animal and / or evaluating the activated partial thromboplastin time, or performing a thrombin generation assay. Optionally, the human albumin targeting reagent comprises (1) a nuclease agent designed to target a region of the human albumin gene, and (2) an exogenous donor nucleic acid, the exogenous donor nucleic acid being designed to target the human albumin gene, the exogenous donor nucleic acid encoding an exogenous protein, and the protein encoded by the humanized endogenous albumin locus targeted by the exogenous donor nucleic acid being a heterologous protein comprising a human albumin signal peptide fused to the exogenous protein. Optionally, the evaluation includes measuring the expression of messenger RNA encoded by the exogenous donor nucleic acid. Optionally, the evaluation includes measuring the expression of the exogenous protein. Optionally, evaluating the expression of the heterologous protein includes measuring the serum level of the heterologous protein in a non-human animal. Optionally, evaluating the expression of the heterologous protein includes measuring the expression in the liver of a non-human animal.
[0028] In another aspect, methods are provided for optimizing the activity of a human albumin targeting reagent in vivo. Some such methods include: (I) first performing any of the above methods for evaluating the activity of a human albumin targeting reagent in vivo in a first non-human animal comprising a humanized endogenous albumin locus in its genome; (II) changing a variable element and second performing the method of step (I) in a second non-human animal comprising a humanized endogenous albumin locus in its genome using the changed variable element; and (III) comparing the activity of the human albumin targeting reagent of step (I) with the activity of the human albumin targeting reagent of step (II) and selecting the method that results in a higher activity.
[0029] In some such methods, the changed variable element in step (II) is the delivery method for introducing the human albumin targeting reagent into the non-human animal. Optionally, administration includes LNP-mediated delivery and the changed variable element in step (II) is the LNP formulation. In some such methods, the changed variable element in step (II) is the route of administration for introducing the human albumin targeting reagent into the non-human animal. In some such methods, the changed variable element in step (II) is the concentration or amount of the human albumin targeting reagent introduced into the non-human animal. In some such methods, the changed variable element in step (II) is the form of the human albumin targeting reagent introduced into the non-human animal. In some such methods, the changed variable element in step (II) is the human albumin targeting reagent introduced into the non-human animal.
[0030] In some such methods, the human albumin targeting reagent comprises a Cas protein and a guide RNA designed to target a guide RNA target sequence of the human albumin gene. In some such methods, the human albumin targeting reagent comprises a Cas protein or a nucleic acid encoding a Cas protein and a guide RNA or a DNA encoding a guide RNA, and the guide RNA is designed to target a guide RNA target sequence of the human albumin gene. Optionally, the variable element modified in step (II) is a guide RNA sequence or a guide RNA target sequence. Optionally, the Cas protein and the guide RNA are each administered in the form of RNA, and the variable element modified in step (II) is the ratio of Cas mRNA to guide RNA. Optionally, the variable element modified in step (II) is a guide RNA modification. Optionally, the human albumin targeting reagent comprises a messenger RNA (mRNA) encoding a Cas protein and a guide RNA, and the variable element modified in step (II) is the ratio of Cas mRNA to guide RNA.
[0031] In some such methods, the human albumin targeting reagent comprises an exogenous donor nucleic acid. Optionally, the variable element modified in step (II) is the form of the exogenous donor nucleic acid. Optionally, the exogenous donor nucleic acid comprises an insertion nucleic acid flanked by a 5' homology arm targeting a 5' target sequence of the humanized endogenous albumin locus and a 3' homology arm targeting a 3' target sequence of the humanized endogenous albumin locus, and the variable element modified in step (II) is the sequence or length of the 5' homology arm and / or the sequence or length of the 3' homology arm.
[0032] In another aspect, provided is a method of generating any of the above non-human animals. Some such methods include: (a) introducing into a non-human animal embryonic stem (ES) cell a targeting vector comprising: (i) a nuclease agent that targets a target sequence of an endogenous albumin locus, and (ii) a nucleic acid insert comprising a human albumin sequence flanked by a 5' homology arm corresponding to a 5' target sequence of the endogenous albumin locus and a 3' homology arm corresponding to a 3' target sequence of the endogenous albumin locus, wherein the targeting vector recombines with the endogenous albumin locus to produce a genetically modified non-human ES cell comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome; (b) introducing the genetically modified non-human ES cell into a non-human animal host embryo; and (c) implanting the non-human animal host embryo in a surrogate mother, wherein the surrogate mother gives birth to an F0 progeny genetically modified non-human animal comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome. In another aspect, provided is a method of generating any of the above non-human animals. Some such methods include: (a) introducing into a non-human animal embryonic stem (ES) cell a targeting vector comprising: (i) a nuclease agent or a nucleic acid encoding a nuclease agent (the nuclease agent targets a target sequence of an endogenous albumin locus), and (ii) a nucleic acid insert comprising a human albumin sequence flanked by a 5' homology arm corresponding to a 5' target sequence of the endogenous albumin locus and a 3' homology arm corresponding to a 3' target sequence of the endogenous albumin locus, wherein the targeting vector recombines with the endogenous albumin locus to produce a genetically modified non-human ES cell comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome; (b) introducing the genetically modified non-human ES cell into a non-human animal host embryo; and (c) implanting the non-human animal host embryo in a surrogate mother, wherein the surrogate mother gives birth to an F0 progeny genetically modified non-human animal comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome.Optionally, the targeting vector is a large targeting vector at least 10 kb in length or the total of the 5' and 3' homology arms is at least 10 kb in length.
[0033] Some such methods involve: (a) introducing into a one-cell stage embryo of a non-human animal a targeting vector comprising (i) a nuclease agent that targets a target sequence of the endogenous albumin locus and (ii) a nucleic acid insert comprising a human albumin sequence flanked by a 5' homology arm corresponding to the 5' target sequence of the endogenous albumin locus and a 3' homology arm corresponding to the 3' target sequence of the endogenous albumin locus, wherein the targeting vector recombines with the endogenous albumin locus to produce a genetically modified non-human one-cell stage embryo comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome; and (b) implanting the one-cell stage embryo of the genetically modified non-human animal in a surrogate mother to produce a genetically modified F0 generation non-human animal comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome. Some such methods involve: (a) introducing into a one-cell stage embryo of a non-human animal (i) a nuclease agent or a nucleic acid encoding a nuclease agent (the nuclease agent targeting a target sequence of the endogenous albumin locus) and (ii) a targeting vector comprising a nucleic acid insert comprising a human albumin sequence flanked by a 5' homology arm corresponding to the 5' target sequence of the endogenous albumin locus and a 3' homology arm corresponding to the 3' target sequence of the endogenous albumin locus, wherein the targeting vector recombines with the endogenous albumin locus to produce a genetically modified non-human one-cell stage embryo comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome; and (b) implanting the one-cell stage embryo of the genetically modified non-human animal in a surrogate mother to produce a genetically modified F0 generation non-human animal comprising a humanized endogenous albumin locus containing the human albumin sequence in its genome.
[0034] In some such methods, the nuclease agent comprises a Cas protein and a guide RNA. Optionally, the Cas protein is a Cas9 protein. Optionally, step (a) further comprises introducing a second guide RNA that targets a second target sequence within the endogenous albumin locus.
[0035] In some such methods, the non-human animal is a mouse or a rat. Optionally, the non-human animal is a mouse.
[0036] In another aspect, provided is a method of making any of the above non-human animals. Some such methods include: (a) modifying the genome of a pluripotent non-human animal cell to include a humanized endogenous albumin locus; (b) identifying or selecting a genetically modified pluripotent non-human animal cell that includes the humanized endogenous albumin locus; (c) introducing the genetically modified pluripotent non-human animal cell into a non-human animal host embryo; and (d) implanting the non-human animal host embryo into a surrogate mother. Some such methods include: (a) modifying the genome of a one-cell stage embryo of a non-human animal to include a humanized endogenous albumin locus; (b) selecting a genetically modified one-cell stage embryo of a non-human animal that includes the humanized endogenous albumin locus; and (c) implanting the genetically modified one-cell stage embryo of a non-human animal into a surrogate mother.
Brief Description of the Drawings
[0037]
FIG. 1A
FIG. 1B
FIG. 2
FIG. 3A
FIG. 3B
FIG. 4
FIG. 5
FIG. 6A
FIG. 6B
FIG. 7
FIG. 8
FIG. 9
FIG. 10A
FIG. 10B
FIG. 11
[0038] Definitions As used interchangeably herein, the terms "protein", "polypeptide", and "peptide" include polymeric forms of amino acids of any length, including coded and non-coded amino acids, as well as chemically or biochemically modified or derivatized amino acids. These terms also include modified polymers such as polypeptides having a modified peptide backbone. The term "domain" refers to any part of a protein or polypeptide having a specific function or structure.
[0039] As used interchangeably herein, the terms "nucleic acid" and "polynucleotide" include polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These include single-stranded, double-stranded, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers comprising purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, unnatural, or derivatized nucleotide bases.
[0040] The term "integrated into the genome" refers to a nucleic acid that has been introduced into a cell such that the nucleotide sequence is integrated into the genome of the cell. Any protocol may be used for stable integration of the nucleic acid into the genome of the cell.
[0041] The term "targeting vector" refers to a recombinant nucleic acid that can be introduced by homologous recombination, non-homologous end-joining-mediated ligation, or any other means of recombination to a target location in the genome of a cell.
[0042] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient or permissive for packaging into viral vector particles. The vector and / or particle can be utilized for the purpose of introducing DNA, RNA, or other nucleic acids into cells in vitro, ex vivo, or in vivo. Many forms of viral vectors are known.
[0043] The term "isolated" with respect to cells, tissues (e.g., liver samples), proteins, and nucleic acids includes cells, tissues (e.g., liver samples), proteins, and nucleic acids that are relatively purified with respect to other bacteria, viruses, cells, or other components that may be present in situ, up to substantially pure preparations of cells, tissues (e.g., liver samples), proteins, and nucleic acids, as well as substantially pure preparations thereof. The term "isolated" also includes proteins and nucleic acids that are chemically synthesized without having a naturally occurring counterpart and are thereby substantially free of contamination by other cells, tissues (e.g., liver samples), proteins, and nucleic acids, or are separated or purified from most other components (e.g., cellular components) that are naturally associated therewith (e.g., other cellular proteins, polynucleotides, or other components).
[0044] The term "wild-type" refers to an entity having a structure and / or activity as found in a normal state or context (as opposed to mutant, diseased, modified, etc.). Wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles).
[0045] The term "endogenous sequence" refers to a nucleic acid sequence that occurs naturally within a cell or non-human animal. For example, the endogenous albumin sequence of a non-human animal refers to the native albumin sequence that naturally exists at the albumin locus of the non-human animal.
[0046] "Exogenous" molecules or sequences include molecules or sequences that are not normally present in a cell in their form or location (e.g., genomic locus). Normal presence includes presence with respect to a particular developmental stage and environmental conditions of the cell. Exogenous molecules or sequences can include, for example, mutant versions of corresponding endogenous sequences within the cell, such as humanized versions of endogenous sequences, or sequences that are within the cell but correspond to endogenous sequences in a different form (i.e., not chromosomal). In contrast, endogenous molecules or sequences include molecules or sequences that are normally present in a particular cell, at a particular developmental stage, and under particular environmental conditions, in their form and location.
[0047] The term "heterologous" when used in the context of a nucleic acid or protein indicates that the nucleic acid or protein contains at least two segments that do not naturally occur together within the same molecule. For example, the term "heterologous" when used with respect to a segment of a nucleic acid or a segment of a protein indicates that the nucleic acid or protein contains two or more sub-sequences that are not found in nature in the same relationship (e.g., bound together) to one another. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid that is within or attached to a different nucleic acid molecule that is not found in nature in association with other molecules. For example, a heterologous region of a nucleic acid vector can include a coding sequence that is adjacent to a sequence that is not found in association with coding sequences in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids that is within or attached to a different peptide molecule that is not found in nature in association with other peptide molecules (e.g., a fusion protein, or a tagged protein). Similarly, a nucleic acid or protein can include a heterologous label or a heterologous secretion or localization sequence.
[0048] "Codon optimization" involves utilizing the degeneracy of codons, as indicated by the diversity of combinations of three-base pair codons that specify amino acids, and generally involves replacing at least one codon of a native sequence with a codon that is more frequently or most frequently used in the genes of a host cell, while maintaining the native amino acid sequence, to modify a nucleic acid sequence for enhanced expression in a particular host cell. For example, a nucleic acid encoding a Cas9 protein can be modified to alternative codons that are more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, in the "Codon Usage Database". These tables can be adapted in various ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, which is hereby incorporated by reference in its entirety for all purposes. Computer algorithms (e.g., see Gene Forge) for codon optimization of specific sequences for expression in a particular host are also available.
[0049] The term "locus" refers to a specific location of a gene (or significant sequence), DNA sequence, polypeptide coding sequence, or position on a chromosome of an organism's genome. For example, the "albumin locus" or "Alb locus" can refer to the albumin (Alb) gene, albumin DNA sequence, sequence encoding albumin, or the specific location of the position of albumin on the chromosome of the genome of an organism in which such a sequence has been identified as resident. The "albumin locus" can include regulatory elements of the albumin gene, including, for example, enhancers, promoters, 5' and / or 3' untranslated regions (UTRs), or combinations thereof.
[0050] The term "gene" refers to a DNA sequence in a chromosome that encodes a product (e.g., an RNA product and / or a polypeptide product), including a coding region interrupted by non-coding introns and sequences located adjacent to the coding region at both the 5' and 3' ends such that the gene corresponds to a full-length mRNA (including 5' and 3' untranslated sequences). The term "gene" also includes regulatory sequences (e.g., promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulator sequences, and other non-coding sequences including matrix attachment regions. These sequences may be close to (e.g., within 10 kb of) or distant from the coding region of the gene and they affect the level or rate of transcription and translation of the gene.
[0051] The term "allele" refers to variant forms of a gene. Some genes have various different forms located at the same position on a chromosome, or locus. A diploid organism has two alleles at each locus. Each pair of alleles represents the genotype of a particular locus. The genotype is described as homozygous when there are two identical alleles at a particular locus and heterozygous when the two alleles are different.
[0052] A "promoter" is a regulatory region of DNA that typically contains a TATA box and can direct RNA polymerase II to initiate RNA synthesis at an appropriate transcription start site for a particular polynucleotide sequence. The promoter may further contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of operably linked polynucleotides. The promoter can be active in one or more of the cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, one-cell stage embryos, differentiated cells, or combinations thereof). The promoter can be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO2013 / 176772, which is hereby incorporated by reference in its entirety for all purposes.
[0053] Examples of inducible promoters include, for example, chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include, for example, alcohol-regulated promoters (e.g., the alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., the tetracycline-responsive promoter, the tetracycline operator sequence (tetO), the tet-On promoter, or the tet-Off promoter), steroid-regulated promoters (e.g., the promoter of the rat glucocorticoid receptor, the estrogen receptor, or the ecdysone receptor), or metal-regulated promoters (e.g., the metallothionein promoter). Physically regulated promoters include, for example, temperature-regulated promoters (e.g., the heat shock promoter), and light-regulated promoters (e.g., the light-inducible promoter or the light-repressible promoter).
[0054] Tissue-specific promoters can be, for example, neuron-specific promoters, glia-specific promoters, muscle cell-specific promoters, heart cell-specific promoters, kidney cell-specific promoters, bone cell-specific promoters, endothelial cell-specific promoters, or immune cell-specific promoters (e.g., B cell promoters or T cell promoters).
[0055] Developmentally regulated promoters include, for example, promoters that are active only during embryonic stages of development or only in adult cells.
[0056] "Operably linked" or "operably connected" includes the juxtaposition of two or more components (e.g., a promoter and another sequence element) such that both components function properly and allow the possibility that at least one of the components can mediate a function that extends to at least one of the other components. For example, a promoter can be operably linked to a coding sequence if the promoter controls the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include the fact that such sequences are in proximity to each other or act in trans (e.g., a regulatory sequence can act remotely to control the transcription of a coding sequence).
[0057] The "complementarity" of nucleic acids means that the nucleotide sequence of one strand of a nucleic acid forms hydrogen bonds with another sequence on the opposite nucleic acid strand due to the orientation of its nucleobases. The complementary bases of DNA are usually A and T, and C and G. In RNA, they are usually C and G, and U and A. Complementarity can be complete or substantial / adequate. Complete complementarity between two nucleic acids means that the two nucleic acids can form a double strand, and all the bases of the double strand are bound to complementary bases by Watson-Crick base pairing. "Substantially" or "adequately" complementary means that the sequence of one strand is not completely and / or perfectly complementary to the sequence of the opposite strand, but sufficient binding occurs between the bases of the two strands to form a stable hybrid complex under a set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by using the sequence and standard mathematical calculations to predict the Tm (melting temperature) of the hybridized strands, or by empirically determining the Tm using routine methods. Tm includes the temperature at which 50% of the population of hybridization complexes formed between two nucleic acid strands denatures (i.e., the population of double-stranded nucleic acid molecules is half-dissociated into single strands). At temperatures lower than Tm, the formation of hybridization complexes is favored, while at temperatures higher than Tm, the melting or separation of the strands in the hybridization complex is favored. Other known Tm calculations take into account the structural features of nucleic acids, but Tm can be estimated for nucleic acids with a known G+C content in 1M aqueous NaCl solution, for example, by using Tm = 81.5 + 0.41(G+C%).
[0058] Hybridization requires that two nucleic acids contain complementary sequences, although mismatches between bases are possible. The conditions appropriate for hybridization between two nucleic acids depend on well-known variables, the length of the nucleic acid and the degree of complementarity. The greater the degree of complementarity between two nucleic acid sequences, the greater the value of the melting temperature (Tm) of the hybrid of the nucleic acids having those sequences. For hybridization between nucleic acids having short stretches of complementarity (e.g., complementarity over 35 or fewer, 30 or fewer, 25 or fewer, 22 or fewer, 20 or fewer, or 18 or fewer nucleotides), the position of the mismatch becomes important (see 11.7-11.8 of Sambrook et al. above). Typically, the length of the hybridizable nucleic acid is at least about 10 nucleotides. Exemplary minimum lengths of hybridizable nucleic acids include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Further, the temperature and the salt concentration of the washing solution can be adjusted as needed, depending on factors such as the length of the complementary region and the degree of complementarity.
[0059] The sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid in order to hybridize specifically. Further, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop structure, or a hairpin structure). A polynucleotide (e.g., gRNA) may contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to a target region within the target nucleic acid sequence to which it is targeted. For example, a gRNA that has 18 nucleotides out of 20 complementary to the target region and thus will hybridize specifically will exhibit 90% complementarity. In this example, the remaining non-complementary nucleotides may form clusters with the complementary nucleotides or may be scattered and need not be contiguous with each other or with the complementary nucleotides.
[0060] The percent complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be routinely determined using the default settings of the algorithms of Smith and Waterman (1981) Adv. Appl. Math. 2:482 - 489) by using the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al. (1990) J. Mol. Biol. 215:403 - 410, Zhang and Madden (1997) Genome Res. 7:649 - 656, which are hereby incorporated by reference in their entirety for all purposes) or by using the GAP program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.).
[0061] The methods and compositions provided herein use a variety of different components. Some of the components throughout the description can include active variants and fragments. Such components include, for example, Cas proteins, CRISPR RNAs, tracrRNAs, and guide RNAs. The biological activity of each of these components is described somewhere herein. The term "functional" refers to the inherent ability of a protein or nucleic acid (or a fragment, or variant thereof) that exhibits a biological activity or function. Such biological activities or functions can include, for example, the ability of a Cas protein to bind to a guide RNA and a target DNA sequence. The biological functions of functional fragments or variants can be the same as or actually altered (e.g., with respect to their specificity or selectivity or efficacy) compared to the original molecule, but retain the basic biological function of the molecule.
[0062] The term "variant" refers to a nucleotide sequence (e.g., 1 nucleotide) that is different from the most common sequence in a population, or a protein sequence (e.g., 1 amino acid) that is different from the most common sequence in a population.
[0063] The term "fragment", when referring to a protein, means a protein that has fewer amino acids than the full-length protein. The term "fragment", when referring to a nucleic acid, means a nucleic acid that has fewer nucleotides than the full-length nucleic acid. A fragment can be, for example, an N-terminal fragment (i.e., removal of a portion of the C-terminus of the protein), a C-terminal fragment (i.e., removal of a portion of the N-terminus of the protein), or an internal fragment (i.e., removal of portions of both the N-terminus and C-terminus of the protein). A fragment can be, for example, a 5' fragment (i.e., removal of a portion of the 3' end of the nucleic acid), a 3' fragment (i.e., removal of a portion of the 5' end of the nucleic acid), or an internal fragment (i.e., removal of portions of both the 5' end and 3' end of the nucleic acid) when referring to a nucleic acid fragment.
[0064] "Sequence identity" or "identity", in the context of two polynucleotide or polypeptide sequences, refers to residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. For proteins, when percentage of sequence identity is used, the positions of residues that are not identical often differ by conservative amino acid substitutions where the amino acid residue is substituted by another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), and thus do not change the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity". Means for making this adjustment are well known. Typically, this involves scoring conservative substitutions as partial matches rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, if identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, conservative substitutions are given a score between 0 and 1. Scoring of conservative substitutions is calculated, for example, as is done in the program PC / GENE (Intelligenetics, Mountain View, California).
[0065] "Percentage of sequence identity" includes a value determined by comparing two optimally aligned sequences (the maximum number of residues that exactly match) over a comparison window, where the portion of the polynucleotide sequence in the comparison window may include additions or deletions (i.e., gaps) when compared to a reference sequence (without additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleotide base or amino acid residue occurs in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence contains concatenated non-homologous sequences), the comparison window is the full length of the shorter of the two sequences being compared.
[0066] Unless otherwise stated, sequence identity / similarity values include the percentage identity and percentage similarity of nucleotide sequences using the following parameters: a GAP weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; the percentage identity and percentage similarity of amino acid sequences using a GAP weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or values obtained using any equivalent program and GAP version 10. An "equivalent program" includes any sequence comparison program that, when compared to the corresponding alignment generated by GAP version 10 for any two sequences in question, generates an alignment having the same nucleotide or amino acid residue matches and the same percent sequence identity.
[0067] The term "conservative amino acid substitution" refers to the substitution of an amino acid that is normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of one nonpolar (hydrophobic) residue, such as isoleucine, valine, or leucine, with another nonpolar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue with another, such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. In addition, substitution of one basic residue, such as lysine, arginine, or histidine, with another, or substitution of one acidic residue, such as aspartic acid or glutamic acid, with another acidic residue are additional examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a nonpolar (hydrophobic) amino acid residue, such as isoleucine, valine, leucine, alanine, or methionine, with a polar (hydrophilic) residue, such as cysteine, glutamine, glutamic acid, or lysine, and / or the substitution of a polar residue with a nonpolar residue. The classification of typical amino acids is summarized in Table 1 below. [Table 1]
[0068] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is identical or substantially similar to a known reference sequence, e.g., at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous sequences and paralogous sequences. For example, homologous genes typically derive from a common ancestral DNA sequence via either a speciation event (orthologous genes) or a gene duplication event (paralogous genes). "Ortholog" genes include genes of various species that have evolved from a common ancestral gene by speciation. Orthologs typically retain the same function during the process of evolution. "Paralogous" genes include genes that are related by duplication within the genome. Paralogs can evolve new functions during the process of evolution.
[0069] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube or isolated cells or cell lines). The term "in vivo" includes a natural environment (e.g., a cell or an organism or a body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.
[0070] The term "reporter gene" refers to a nucleic acid having a sequence that encodes a gene product (typically an enzyme) that can be readily and quantitatively assayed when a construct containing a reporter gene sequence operably linked to a heterologous promoter and / or enhancer element is introduced into a cell that contains (or can be made to contain) the factors necessary for the activation of the promoter and / or enhancer element. Examples of reporter genes include, but are not limited to, the gene encoding beta-galactosidase (lacZ), the bacterial chloramphenicol acetyltransferase (cat) gene, the firefly luciferase gene, the gene encoding beta-glucuronidase (GUS), and the gene encoding a fluorescent protein. A "reporter protein" refers to a protein encoded by a reporter gene.
[0071] As used herein, the term "fluorescent reporter protein" means a reporter protein that is detectable based on fluorescence, which fluorescence can be from either the reporter protein directly, the activity on the fluorescent generating substrate of the reporter protein, or a protein having an affinity for binding to a fluorescently tagged compound. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, and ZsGreen1), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, and ZsYellow1), blue fluorescent proteins (e.g., BFP, eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, and T-sapphire), cyan fluorescent proteins (e.g., CFP, eCFP, Cerulean, CyPet, AmCyan1, and Midoriishi-Cyan), red fluorescent proteins (e.g., RFP, mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, and Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, and tdTomato), and any other suitable fluorescent protein that can be detected for its presence in a cell by flow cytometry methods.
[0072] Repair in response to double-strand breaks (DSBs) occurs mainly via two conserved DNA repair pathways: homologous recombination (HR) and non-homologous end joining (NHEJ). See Kasparek & Humphrey (2011) Seminars in Cell & Dev. Biol. 22:886-897, which is hereby incorporated by reference in its entirety for all purposes. Similarly, repair of a target nucleic acid mediated by an exogenous donor nucleic acid can include any process of exchange of genetic information between two polynucleotides.
[0073] The term "recombination" includes any process of exchange of genetic information between two polynucleotides and can occur by any mechanism. Recombination can occur via homology-directed repair (HDR) or homologous recombination (HR). HDR or HR includes forms of nucleic acid repair that can require nucleotide sequence homology, use a "donor" molecule as a template for repair of a "target" molecule (i.e., a molecule that has experienced a double-strand break), and result in transfer of genetic information from the donor to the target. Without wishing to be bound by any particular theory, such transfer can include mismatch correction of heteroduplex DNA formed between the broken target and the donor, and / or synthesis-dependent strand annealing in which the donor is used to resynthesize genetic information that becomes part of the target, and / or related processes. In some cases, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide is incorporated into the target DNA. See Wang et al. (2013) Cell 153:910-918, Mandalos et al. (2012) PLOS ONE 7:e45768:1-9, and Wang et al. (2013) Nat Biotechnol. 31:530-532, each of which is hereby incorporated by reference in its entirety for all purposes.
[0074] Non-homologous end joining (NHEJ) involves the repair of double-strand breaks in nucleic acids by directly ligating the cut ends to each other or to an exogenous sequence without the need for a homologous template. Ligation of non-adjacent sequences by NHEJ can often result in deletions, insertions, or translocations near the site of the double-strand break. For example, NHEJ can also result in targeted integration of an exogenous donor nucleic acid by directly ligating the cut ends to the ends of the exogenous donor nucleic acid (i.e., NHEJ-based capture). Such NHEJ-mediated targeted integration may be preferred for insertion of an exogenous donor nucleic acid when the homologous-directed repair (HDR) pathway is not readily available (e.g., in non-dividing cells, primary cells, and cells that perform DNA repair based on homology poorly). Further, in contrast to homologous-directed repair, knowledge of large regions of sequence identity adjacent to the cut site is not required, which can be beneficial when attempting targeted insertion into genomes of organisms with limited knowledge of the genomic sequence. Integration can proceed via ligation of blunt ends between the exogenous donor nucleic acid and the cut genomic sequence, or via ligation of sticky ends (i.e., having 5' or 3' overhangs) of the exogenous donor nucleic acid adjacent to overhangs that are compatible with those generated by a nuclease agent in the cut genomic sequence. See, for example, US2011 / 020722, WO2014 / 033644, WO2014 / 089290, and Maresca et al. (2013) Genome Res. 23(3):539-546, each of which is hereby incorporated by reference in its entirety for all purposes. When blunt ends are ligated, resection of the target and / or donor may be required due to the region of microhomology generation required for fragment joining, which can result in unwanted modifications to the target sequence.
[0075] A composition or method that "comprises" or "includes" one or more of the recited elements may contain other elements not specifically recited. For example, a composition that "comprises" or "includes" a protein may contain the protein alone or in combination with other components. The transitional phrase "consisting essentially of" means that the scope of the claim is to be interpreted to encompass the specified elements in the claim, and those elements that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of" as used in the claims of the present invention is not intended to be interpreted as equivalent to "comprising".
[0076] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes both the case where the event or circumstance occurs and the case where it does not.
[0077] The specification of a range of values includes all integers within that range or defining that range, and all sub-ranges defined by the integers within that range.
[0078] Unless otherwise clear from the context, the term "about" includes values within the standard error of measurement of the specified value (e.g., SEM).
[0079] The term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the absence of combinations when interpreted in the alternative ("or").
[0080] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.
[0081] The singular forms of the articles "a," "an," and "the" include references to the plural unless the context clearly dictates otherwise. For example, the term "protein" or "at least one protein" can include multiple proteins including mixtures thereof.
[0082] Statistically significant means p ≤ 0.05. [Mode for Carrying Out the Invention]
[0083] I. Overview Provided herein are non-human animal genomes, non-human animal cells, and non-human animals comprising a humanized albumin (ALB) locus, and methods of using such non-human animal cells and non-human animals. A non-human animal cell or non-human animal comprising a humanized albumin locus expresses a chimeric albumin protein comprising a human albumin protein or one or more fragments of a human albumin protein. Such non-human animal cells and non-human animals can be used to evaluate the delivery or efficacy of a human albumin targeting agent (e.g., a CRISPR / Cas9 genome editing agent) in vitro, ex vivo, or in vivo, and can be used in methods for optimizing the delivery of the efficacy of such agents in vitro, ex vivo, or in vivo.
[0084] In some of the non-human animal cells and non-human animals disclosed herein, most or all of the human albumin genomic DNA is inserted into the corresponding orthologous non-human animal albumin locus. In some of the non-human animal cells and non-human animals disclosed herein, most or all of the non-human animal albumin genomic DNA is replaced on a one-to-one basis with the corresponding orthologous human albumin genomic DNA. Since the conserved regulatory elements are likely to remain intact and the spliced transcripts that undergo RNA processing are more stable than cDNA, the expression levels should be higher when the intron-exon structure and splicing mechanism are maintained compared to non-human animals into which cDNA has been inserted. In contrast, insertion of human albumin cDNA into the non-human animal albumin locus eliminates conserved regulatory elements such as those contained within the first exon and intron of non-human animal albumin. Replacing the non-human animal genomic sequence with the corresponding orthologous human genomic sequence or inserting the human albumin genomic sequence into the corresponding orthologous non-human albumin locus is likely to result in faithful expression of the transgene from the endogenous albumin locus. Similarly, transgenic non-human animals with transgenic insertion of the human albumin coding sequence into a random genomic locus rather than the endogenous non-human animal albumin locus do not accurately reflect the endogenous regulation of albumin expression. The humanized albumin alleles generated by replacing most or all of the non-human animal genomic DNA on a one-to-one basis with the corresponding orthologous human genomic DNA or inserting the human albumin genomic sequence into the corresponding orthologous non-human albumin locus provide an approximation of the true human target or the true human target of human albumin-targeting reagents (e.g., CRISPR / Cas9 reagents designed to target human albumin), thereby enabling testing of the efficacy and mechanism of action of such agents in living animals, as well as pharmacokinetic and pharmacodynamic studies in an environment where the humanized protein and the albumin in which the humanized gene is present are the only versions of albumin.
[0085] II. Non-human animals containing a humanized albumin (ALB) locus The non-human animal genomes, non-human animal cells, and non-human animals disclosed herein contain a humanized albumin (ALB) locus. A cell or non-human animal containing a humanized albumin locus expresses a human albumin protein or a partially humanized chimeric albumin protein, wherein one or more fragments of the native albumin protein are replaced with corresponding fragments from human albumin. Also disclosed herein are humanized non-human animal albumin genes in which segments of the non-human albumin gene have been deleted and replaced with the corresponding human albumin sequences.
[0086] The non-human animal genomes, non-human animal cells, and non-human animals disclosed herein may further comprise an inactivated (knocked out) endogenous gene that is not the albumin locus. Such non-human animal genomes, non-human animal cells, and non-human animals can be used, for example, to screen gene therapy reagents (e.g., transgenes) for insertion into a humanized albumin locus to replace the inactivated endogenous gene. Insertion into a humanized albumin locus to replace the inactivated endogenous gene can, for example, rescue the knockout. In one particular example, the non-human animal genomes, non-human animal cells, and non-human animals disclosed herein may further comprise an inactivated (knocked out) endogenous F9 gene (encoding coagulation factor IX). The inactivated (knocked out) endogenous F9 gene is a gene that does not express any coagulation factor IX (also known as Christmas factor, plasma thromboplastin component, or PTC). The wild-type human coagulation factor IX protein has been assigned the UniProt accession number P00740, and the human F9 gene has been assigned GeneID 2158. The wild-type mouse coagulation factor IX protein has been assigned the UniProt accession number P16294, and the mouse F9 gene has been assigned GeneID 14071. The wild-type rat coagulation factor IX protein has been assigned the UniProt accession number P16296, and the rat F9 gene has been assigned GeneID 24946.
[0087] The non-human animal genomes, non-human animal cells, and non-human animals disclosed herein may further comprise a coding sequence of an exogenous protein integrated into at least one allele of a humanized albumin locus (e.g., in one or more cells of a non-human animal such as one or more hepatocytes of a non-human animal). The coding sequence can be integrated, for example, into intron 1, intron 12, or intron 13 of the humanized albumin locus. In some cases, expression of human albumin from the humanized albumin locus is maintained at the same level after the coding sequence of the exogenous protein has been integrated into at least one allele of the humanized albumin locus (e.g., in one or more cells of a non-human animal such as one or more hepatocytes of a non-human animal). In one example, the non-human animal genome, cell, or animal further comprises an inactivated (knocked out) endogenous gene that is not the albumin locus, and the exogenous protein replaces the function of the inactivated endogenous gene (e.g., rescues the knockout). In one particular example, the exogenous protein is coagulation factor IX (e.g., human coagulation factor IX).
[0088] A. Albumin The cells and non-human animals described herein contain a humanized albumin (ALB) locus. Albumin is encoded by the ALB gene (also known as albumin, serum albumin, PRO0883, PRO0903, HSA, GIG20, GIG42, PRO1708, PRO2044, PRO2619, PRO2675, and UNQ696 / PRO1341). Albumin is synthesized in the liver as preproalbumin, which has an N-terminal peptide that is removed before the nascent protein is released from the rough endoplasmic reticulum. The resulting proalbumin is then cleaved in the Golgi vesicles to produce secreted albumin (serum albumin). Human serum albumin is the serum albumin found in human blood. This is the most abundant protein in human plasma and constitutes about half of the serum proteins. It is produced in the liver. It is water-soluble and monomeric. Albumin transports hormones, fatty acids, and other compounds, buffers pH, and maintains colloid osmotic pressure, among other functions. The concentration of human albumin in serum is typically about 35-50 g / L (3.5-5.0 g / dL). The serum half-life is about 20 days. The molecular weight is 66.5 kDa.
[0089] Albumin is considered a safe harbor locus of the genome due to its very high expression level and the ease of handling of the liver for gene delivery and in vivo editing compared to other tissues. A safe harbor locus includes chromosomal loci where a transgene or other exogenous nucleic acid insert can be stably and reliably expressed in all tissues of interest without overtly changing the behavior or phenotype of the cell. In many cases, a safe harbor locus is a locus where the expression of the inserted gene sequence is not disrupted by read-through expression from adjacent genes. For example, a safe harbor locus can include chromosomal loci where exogenous DNA can be integrated and function in a predictable manner without adversely affecting the structure or expression of endogenous genes. A safe harbor locus can include extragenic or intragenic regions, such as intragenic loci that can be disrupted without being essential, unnecessary, or resulting in an overt phenotypic consequence.
[0090] Since the albumin gene structure encodes a secretory peptide (signal peptide) whose first exon is cleaved from the final protein product, it is suitable for targeting transgenes into intron sequences. For example, the incorporation of a cassette without a promoter having a splice acceptor and a therapeutic transgene supports the expression and secretion of many different proteins.
[0091] Human ALB is located at human 4q13.3 on chromosome 4 (NCBI RefSeq Gene ID 213; Assembly GRCh38.p12 (GCF_000001405.38); location NC_000004.12 (73404239..73421484 (+))). This gene has been reported to have 15 exons. Of these, 14 exons are coding exons, and exon 15 is a non-coding exon that is part of the 3' untranslated region (UTR). The wild-type human albumin protein has been assigned the UniProt accession number P02768. At least three isoforms are known (P02768-1 to P02768-3). The sequence of one isoform, P02768-1 (identical to NCBI accession number NP_000468.1), is shown in SEQ ID NO: 5. The mRNA (cDNA) encoding the standard isoform has been assigned the NCBI accession number NM_000477.7 and is shown in SEQ ID NO: 37. The exemplary coding sequence (CDS) has been assigned the CCDS ID CCDS3555.1 and is shown in SEQ ID NO: 13. The full-length human albumin protein shown in SEQ ID NO: 5 has 609 amino acids, including a signal peptide (amino acids 1 to 18), a propeptide (amino acids 19 to 24), and serum albumin (amino acids 25 to 609). The delineation between these domains is as shown in UniProt. References to human albumin include the standard (wild-type) form as well as all allelic forms and isoforms. Any other form of human albumin has amino acids numbered relative to the maximum alignment with the wild-type form, and aligned amino acids are shown with the same number.
[0092] Mouse Alb is located at mouse 5 E1;5 44.7 cM on chromosome 5 (NCBI RefSeq Gene ID 11657; Assembly GRCm38.p4 (GCF_000001635.24); position NC_000071.6 (90,460,870..90,476,602 (+))). This gene has been reported to have 15 exons. Of these, 14 exons are coding exons and exon 15 is a non-coding exon that is part of the 3' untranslated region (UTR). The wild-type mouse albumin protein is assigned the UniProt accession number P07724. The sequence of mouse albumin (identical to NCBI accession number NP_033784.2) is set forth in SEQ ID NO:1. An exemplary mRNA (cDNA) isoform encoding the standard isoform is assigned the NCBI accession number NM_009654.4 and is shown in SEQ ID NO:36. An exemplary coding sequence (CDS) (CCDS ID CCDS19412.1) is shown in SEQ ID NO:9. The standard full-length mouse albumin protein shown in SEQ ID NO:1 has 608 amino acids and includes a signal peptide (amino acids 1-18), a propeptide (amino acids 19-24), and serum albumin (amino acids 25-608). The delineation between these domains is as shown in UniProt. References to mouse albumin include the standard (wild-type) form as well as all allelic forms and isoforms. Any other form of mouse albumin has amino acids numbered relative to the maximum alignment with the wild-type form, and aligned amino acids are indicated by the same number.
[0093] Albumin sequences of many other non-human animals are also known. These include, for example, bovine (UniProt accession number P02769; NCBI RefSeq Gene ID 280717), rat (UniProt accession number P02770; NCBI RefSeq Gene ID 24186), chicken (UniProt accession number P19121), Sumatran orangutan (UniProt accession number Q5NVH5; NCBI RefSeq Gene ID 100174145), horse (UniProt accession number P35747; NCBI RefSeq Gene ID 100034206), cat (UniProt accession number P49064; NCBI RefSeq Gene ID 448843), rabbit (UniProt accession number P49065; NCBI RefSeq Gene ID 100009195), dog (UniProt accession number P49822; NCBI RefSeq Gene ID 403550), pig (UniProt accession number P08835; NCBI RefSeq Gene ID 396960), mouse (UniProt accession number O35090), rhesus macaque (UniProt accession number Q28522); NCBI RefSeq Gene ID 704892), donkey (UniProt accession number Q5XLE4; NCBI RefSeq Gene ID 106835108), sheep (UniProt accession number P14639; NCBI RefSeq Gene ID 443393), Xenopus laevis (UniProt accession number P21847), golden hamster (UniProt accession number A6YF56); NCBI RefSeq Gene ID 101837229), and goat (UniProt accession number P85295).
[0094] B. Humanized albumin locus A humanized albumin locus is an albumin locus in which a segment of the endogenous albumin locus has been deleted and replaced with an orthologous human albumin sequence. The humanized albumin locus can be an albumin locus in which the entire albumin gene has been replaced with the corresponding orthologous human albumin sequence, or it can be an albumin locus in which only a portion of the albumin gene has been replaced with the corresponding orthologous human albumin sequence (i.e., humanized). For example, the entire albumin coding sequence of the endogenous albumin locus can be deleted and replaced with the corresponding human albumin sequence. The human albumin sequence corresponding to a particular segment of the endogenous albumin sequence refers to the region of human albumin that aligns with the particular segment of the endogenous albumin sequence when human albumin and the endogenous albumin are optimally aligned. Optimally aligned refers to the maximum number of residues that are exactly matched. The corresponding orthologous human sequence can include, for example, complementary DNA (cDNA) or genomic DNA. Optionally, the corresponding orthologous human albumin sequence is modified to be codon-optimized based on the codon usage frequency in the non-human animal. The replaced or inserted (i.e., humanized) region can include coding regions such as exons, non-coding regions such as introns, untranslated regions, or regulatory regions (e.g., promoters, enhancers, or transcription repressor binding elements), or any combination thereof. As an example, the exons corresponding to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all 15 exons of the human albumin gene can be humanized. For example, the exons corresponding to all exons (i.e., exons 1-15) of the human albumin gene can be humanized. As another example, the exons corresponding to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all 14 coding exons of the human albumin gene can be humanized. For example, the exons corresponding to all coding exons (i.e., exons 1-14) of the human albumin gene can be humanized.Similarly, can the introns corresponding to all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 introns of the human albumin gene be humanized, or can they be left endogenous? For example, the introns corresponding to all of the introns of the human albumin gene (i.e., introns 1-14) can be humanized. Similarly, can the introns corresponding to all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13 of the introns between the coding exons of the human albumin gene be humanized, or can they be left endogenous? For example, the introns corresponding to all of the introns between the coding exons of the human albumin gene (i.e., introns 1-13) can be humanized. The adjacent untranslated regions containing regulatory sequences can also be humanized or can remain endogenous. For example, the 5' untranslated region (UTR), 3' UTR, or both the 5' UTR and 3' UTR can be humanized, or the 5' UTR, 3' UTR, or both the 5' UTR and 3' UTR can be left endogenous. One or both of the human 5' and 3' UTRs can be inserted, and / or one or both of the endogenous 5' and 3' UTRs can be deleted. In certain examples, both the 5' UTR and 3' UTR remain endogenous. In another specific example, the 5' UTR remains endogenous and the 3' UTR is humanized. Depending on the degree of substitution by orthologous sequences, regulatory sequences such as promoters can be endogenous or can be provided by the human orthologous sequence being substituted. For example, the humanized albumin locus can include an endogenous non-human animal albumin promoter (i.e., the human albumin sequence can be operably linked to an endogenous non-human animal promoter).
[0095] One or more or all of the regions encoding the signal peptide, propeptide, or serum albumin can be humanized, or one or more of such regions can be left endogenous. Exemplary coding sequences for the mouse albumin signal peptide, propeptide, and serum albumin are set forth in SEQ ID NOs: 10-12, respectively. Exemplary coding sequences for the human albumin signal peptide, propeptide, and serum albumin are shown in SEQ ID NOs: 14-16, respectively.
[0096] For example, all or part of the albumin locus region encoding the signal peptide can be humanized, and / or all or part of the albumin locus region encoding the propeptide can be humanized, and / or all or part of the albumin locus region encoding serum albumin can be humanized. Alternatively or additionally, all or part of the albumin locus region encoding the signal peptide can be left endogenous, and / or all or part of the albumin locus region encoding the propeptide can be left endogenous, and / or all or part of the albumin locus region encoding serum albumin can be left endogenous. In one example, all or part of the albumin locus region encoding the signal peptide, propeptide, and serum albumin is humanized. Optionally, the CDS of the humanized region of the albumin locus comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 13 (or a degenerate thereof). Optionally, the CDS of the humanized region of the albumin locus comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 13 (or a degenerate thereof). Optionally, the humanized region of the albumin locus comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 35. Optionally, the humanized region of the albumin locus comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 35.Optionally, the humanized albumin locus encodes a protein that comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 5. Optionally, the humanized albumin locus encodes a protein that comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 5. Optionally, the humanized albumin locus comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 17 or 18. Optionally, the humanized albumin locus comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 17 or 18.
[0097] The albumin protein encoded by the humanized albumin locus can include one or more domains derived from the human albumin protein and / or one or more domains derived from the endogenous (i.e., native) albumin protein. Exemplary amino acid sequences of the mouse albumin signal peptide, propeptide, and serum albumin are shown in SEQ ID NOs: 2-4, respectively. Exemplary amino acid sequences of the human albumin signal peptide, propeptide, and serum albumin are shown in SEQ ID NOs: 6-8, respectively.
[0098] The albumin protein can comprise one or more or all of a human albumin signal peptide, a human albumin propeptide, and human serum albumin. Alternatively or additionally, the albumin protein can comprise one or more domains derived from an endogenous (i.e., native) non-human animal albumin protein. For example, the albumin protein can comprise a signal peptide derived from an endogenous (i.e., native) non-human animal albumin protein and / or a propeptide derived from an endogenous (i.e., native) non-human animal albumin protein and / or serum albumin derived from an endogenous (i.e., native) non-human animal albumin protein. As an example, the albumin protein can comprise a human signal peptide, propeptide, and serum albumin.
[0099] The domains of the chimeric albumin protein derived from a human albumin protein can be encoded by a fully humanized sequence (i.e., the entire sequence encoding the domain is replaced with an orthologous human albumin sequence) or can be encoded by a partially humanized sequence (i.e., a portion of the sequence encoding the domain is replaced with an orthologous human albumin sequence and the remaining endogenous (i.e., native) sequence encoding the domain encodes the same amino acids as the orthologous human albumin sequence such that the domain encoded is the same as that domain of the human albumin protein). Similarly, the domains of the chimeric protein derived from an endogenous albumin protein can be encoded by a fully endogenous sequence (i.e., the entire sequence encoding the domain is an endogenous albumin sequence) or can be encoded by a partially humanized sequence (i.e., a portion of the sequence encoding the domain is replaced with an orthologous human albumin sequence, but the orthologous human albumin sequence encodes the same amino acids as the replaced endogenous albumin sequence such that the domain encoded is the same as that domain of the endogenous albumin protein).
[0100] As an example, the albumin protein encoded by the humanized albumin locus may include the human albumin signal peptide. Optionally, the human albumin signal peptide comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 6. Optionally, the human albumin signal peptide comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 6. As another example, the albumin protein encoded by the humanized albumin locus may include the human albumin propeptide. Optionally, the human albumin propeptide comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 7. Optionally, the human albumin propeptide comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 7. As another example, the albumin protein encoded by the humanized albumin locus may include human serum albumin. Optionally, the human serum albumin comprises, consists essentially of, or consists of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 8. Optionally, the human serum albumin comprises, consists essentially of, or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 8. The albumin protein encoded by the humanized albumin locus may comprise, consist essentially of, or consist of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 5, or be essentially or consist of the same.For example, the albumin protein encoded by the humanized albumin locus may contain, consist essentially of, or consist of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 5. Optionally, the albumin CDS encoded by the humanized albumin locus may contain, consist essentially of, or consist of a sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 13 (or its degeneracy). Optionally, the albumin CDS encoded by the humanized albumin locus may contain, consist essentially of, or consist of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 13 (or its degeneracy).
[0101] The humanized albumin protein can retain the activity of the native albumin protein and / or the human albumin protein.
[0102] Optionally, the humanized albumin locus can include other elements. Examples of such elements can include a selection cassette, a reporter gene, a recombinase recognition site, or other elements. Alternatively, the humanized albumin locus may lack other elements (e.g., it may lack a selectable marker or a selection cassette). Examples of suitable reporter genes and reporter proteins are disclosed elsewhere in this specification. Examples of suitable selectable markers include neomycin phosphotransferase (neo r ), hygromycin B phosphotransferase (hyg r ), puromycin-N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r) contains xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k). Examples of recombinases include Cre, Flp, and Dre recombinases. An example of a Cre recombinase gene is Crei, where two exons encoding Cre recombinase are separated by an intron to prevent expression in prokaryotic cells. Such recombinases can further contain a nuclear localization signal to promote nuclear localization (e.g., NLS-Crei). Recombinase recognition sites contain nucleotide sequences that are recognized by site-specific recombinases and can function as substrates for recombination events. Examples of recombinase recognition sites include FRT, FRT11, FRT71, attp, att, rox, and lox sites such as loxP, lox511, lox2272, lox66, lox71, loxM2, and lox5171.
[0103] Other elements such as reporter genes or selection cassettes may be self-deleting cassettes adjacent to the recombinase recognition site. See, for example, US8,697,851 and US2013 / 0312129, which are hereby incorporated by reference in their entireties for all purposes. As an example, the self-deleting cassette can include a Crei gene (including two exons encoding Cre recombinase separated by an intron) operably linked to the mouse Prm1 promoter and a neomycin resistance gene operably linked to the human ubiquitin promoter. By using the Prm1 promoter, the self-deleting cassette can be specifically deleted in the male germ cells of F0 animals. The polynucleotide encoding the selectable marker can be operably linked to a promoter active in the target cell. Examples of promoters are described anywhere in this specification. As another specific example, the self-deleting selection cassette can include a hygromycin resistance gene coding sequence operably linked to one or more promoters (e.g., both the human ubiquitin and EM7 promoters), followed by a polyadenylation signal, followed by a Crei coding sequence operably linked to one or more promoters (e.g., the mPrm1 promoter), followed by another polyadenylation signal, with the entire cassette flanked by loxP sites.
[0104] The humanized albumin locus may also be a conditional allele. For example, a conditional allele may be a multifunctional allele as described in US2011 / 0104799, which is hereby incorporated by reference in its entirety for all purposes. For example, a conditional allele can include the following: (a) an operative sequence in the sense direction with respect to transcription of the target gene, (b) a drug selection cassette (DSC) in the sense or antisense direction, (c) a nucleotide sequence of interest (NSI) in the antisense orientation, (d) a conditional inversion module (COIN) in the reverse direction (utilizing modules such as exon-split introns and invertible gene traps). See, e.g., US2011 / 0104799. The conditional allele can further include a recombinable unit that, upon exposure to a first recombinase, recombines to form a conditional allele that (i) lacks the operative sequence and DSC, and (ii) includes the NSI in the sense direction and the COIN in the antisense direction. See, e.g., US2011 / 0104799.
[0105] One exemplary humanized albumin locus (e.g., a humanized mouse albumin locus) is a locus in which the region from the start codon to the stop codon is replaced with the corresponding human sequence. See FIGS. 1A and 1B and SEQ ID NOs: 17 and 18. In a particular example, the region from the ATG start codon to the stop codon (i.e., encoding exons 1-14) can be deleted from the non-human animal (e.g., mouse) albumin (Alb) locus, and in place of the deleted endogenous region, the corresponding region of human albumin (ALB) from the ATG start codon to about 100 bp downstream of the stop codon can be inserted.
[0106] C. Non-human animal genomes, non-human animal cells, and non-human animals comprising a humanized albumin (ALB) locus Provided are non-human animal genomes, non-human animal cells, and non-human animals that include a humanized albumin (ALB) locus as described elsewhere in this specification. The genome, cell, or non-human animal can be male or female. The genome, cell, or non-human animal can be heterozygous or homozygous with respect to the humanized albumin locus. Diploid organisms have two alleles at each locus. Each pair of alleles represents the genotype of a particular locus. Genotypes are described as homozygous when there are two identical alleles at a particular locus and as heterozygous when the two alleles are different. Non-human animals that include a humanized albumin locus can include the humanized endogenous albumin locus in their germline.
[0107] The non-human animal genomes or cells provided herein can be, for example, any non-human animal genome or cell that includes an albumin locus or genomic locus that is homologous or orthologous to the human albumin locus. The genome can be, for example, a eukaryotic cell including fungal cells (e.g., yeast), plant cells, animal cells, mammalian cells, non-human mammalian cells, and human cells. The term "animal" includes any member of the animal kingdom, including, for example, mammals, fish, reptiles, amphibians, birds, and insects. Mammalian cells can be, for example, non-human mammalian cells, rodent cells, rat cells, mouse cells, or hamster cells. Other non-human mammals include, for example, non-human primates, monkeys, apes, orangutans, cats, dogs, rabbits, horses, bulls, deer, bison, livestock (e.g., bovine species such as cows and steers, ovine species such as sheep and goats, and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, ducks, etc. Also included are domesticated animals and agricultural animals. The term "non-human" excludes humans.
[0108] The cells can also be in any type of undifferentiated or differentiated state. For example, the cells can be totipotent cells, pluripotent cells (e.g., human pluripotent cells, or non-human pluripotent cells such as mouse embryonic stem (ES) cells or rat ES cells), or non-pluripotent cells. Totipotent cells include undifferentiated cells that can give rise to any cell type, and pluripotent cells include undifferentiated cells that have the ability to develop into multiple differentiated cell types. Such pluripotent and / or totipotent cells can be, for example, ES cells or ES-like cells such as induced pluripotent stem (iPS) cells. ES cells include totipotent or pluripotent cells derived from the embryo that can contribute to any tissue of the developing embryo when introduced into the embryo. ES cells can be derived from the inner cell mass of the blastocyst and can differentiate into cells of any of the three vertebrate germ layers (endoderm, ectoderm, and mesoderm).
[0109] The cells provided herein can also be germ cells (e.g., sperm or oocytes). The cells can be mitotically competent cells or mitotically inactive cells, meiotically competent cells or meiotically inactive cells. Similarly, the cells can also be primary somatic cells or non-primary somatic cells. Somatic cells include any cells that are not gametes, germ cells, gametocytes, or undifferentiated stem cells. For example, the cells can be hepatocytes such as hepatoblasts or hepatocytes.
[0110] Suitable cells provided herein also include primary cells. Primary cells include cells or cell cultures directly isolated from an organism, organ, or tissue. Primary cells include cells that have not been transformed or immortalized. They include any cells obtained from an organism, organ, or tissue that have not been previously passaged in tissue culture or that have been previously passaged in tissue culture but cannot be passaged indefinitely in tissue culture. Such cells can be isolated by conventional techniques and include, for example, hepatocytes.
[0111] Other suitable cells provided herein include immortalized cells. Immortalized cells include cells from multicellular organisms that normally do not proliferate infinitely but, due to mutations or modifications, avoid normal cellular senescence and can instead continue to divide. Such mutations or modifications can occur naturally or be induced intentionally. A specific example of an immortalized cell line is the HepG2 human hepatocarcinoma cell line. Many types of immortalized cells are well known. Immortalized cells or primary cells typically include cells used for culturing or expressing recombinant genes or proteins.
[0112] The cells provided herein also include one-cell stage embryos (i.e., fertilized oocytes or zygotes). Such one-cell stage embryos can be derived from any genetic background (e.g., BALB / c, C57BL / 6, 129, or combinations thereof in the case of mice), can be fresh or frozen, and can be derived from natural reproduction or in vitro fertilization.
[0113] The cells provided herein can be normal and healthy cells or can be diseased cells or cells having mutations.
[0114] The non-human animals containing the humanized albumin described herein can be produced by the methods described elsewhere herein. The term "animal" includes any member of the animal kingdom, including, for example, mammals, fish, reptiles, amphibians, birds, and insects. In specific examples, the non-human animal is a non-human mammal. Non-human mammals include, for example, non-human primates, monkeys, apes, orangutans, cats, dogs, horses, bulls, deer, bison, sheep, rabbits, rodents (e.g., mice, rats, hamsters, and guinea pigs), and domestic animals (e.g., bovine species such as cows and steers, ovine species such as sheep and goats, and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, and ducks. Domestic and agricultural animals are also included. The term "non-human animal" excludes humans. Preferred non-human animals include rodents such as mice and rats.
[0115] The non-human animals can be from any genetic background. For example, suitable mice can be from the 129 strain, C57BL / 6 strain, a hybrid of 129 and C57BL / 6, BALB / c strain, or Swiss Webster strain. Examples of the 129 strain include 129P1, 129P2, 129P3, 129X1, 129S1 (e.g., 129S1 / SV, 129S1 / Svlm), 129S2, 129S4, 129S5, 129S9 / SvEvH, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1, and 129T2. See, for example, Festing et al. (1999) Mammalian Genome 10:836, which is hereby incorporated by reference in its entirety for all purposes. Examples of the C57BL strain include C57BL / A, C57BL / An, C57BL / GrFa, C57BL / Kal_wN, C57BL / 6, C57BL / 6J, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / 10ScSn, C57BL / 10Cr, and C57BL / Ola. Suitable mice can also be from a hybrid of the aforementioned 129 strain and the aforementioned C57BL / 6 strain (e.g., 50% 129 and 50% C57BL / 6). Similarly, suitable mice can be from a hybrid of the aforementioned 129 strain or a hybrid of the aforementioned BL / 6 strain (e.g., the 129S6 (129 / SvEvTac) strain).
[0116] Similarly, rats can be from, for example, the ACI rat strain, Dark Agouti (DA) rat strain, Wistar rat strain, LEA rat strain, Sprague Dawley (SD) rat strain, or a Fisher rat strain such as Fisher F344 or Fisher F6. Rats can also be obtained from a strain derived from a hybrid of two or more of the aforementioned strains. For example, suitable rats can be from the DA strain or the ACI strain. The ACI rat strain has a white belly and feet, and RT1 av1It is characterized by having a black agouti with a haplotype. Such strains are available from various sources including Harlan Laboratories. The Dark Agouti (DA) rat strain has an agouti coat and RT1 av1 is characterized by having a haplotype. Such rats are available from various sources including Charles River and Harlan Laboratories. Some suitable rats may be derived from inbred rat strains. See, for example, US2014 / 0235933, which is hereby incorporated by reference in its entirety for all purposes.
[0117] A non-human animal (e.g., a mouse or a rat) containing a humanized albumin locus (e.g., a homozygous humanized albumin locus) can express albumin from the humanized albumin locus such that the serum albumin level (e.g., serum human albumin level) is comparable to the serum albumin level in a control wild-type non-human animal. In one example, a non-human animal containing a humanized albumin locus (e.g., a homozygous humanized albumin locus) can have a serum albumin level (e.g., serum human albumin level) that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 100% of the serum albumin level in a control wild-type non-human animal. In another example, a non-human animal containing a humanized albumin locus (e.g., a homozygous humanized albumin locus) can have a serum albumin level (e.g., serum human albumin level) that is at least as high as the serum albumin level in a control wild-type non-human animal. In another example, a non-human animal containing a humanized albumin locus (e.g., a homozygous humanized albumin locus) can have a serum albumin level (e.g., serum human albumin level) that is higher than the serum albumin level in a control wild-type non-human animal. For example, a non-human animal containing a humanized albumin locus (e.g., a homozygous humanized albumin locus) can have a serum albumin level (e.g., serum human albumin level) of at least about 1 mg / mL, at least about 2 mg / mL, at least about 3 mg / mL, at least about 4 mg / mL, at least about 5 mg / mL, at least about 6 mg / mL, at least about 7 mg / mL, at least about 8 mg / mL, at least about 9 mg / mL, at least about 10 mg / mL, at least about 11 mg / mL, at least about 12 mg / mL, at least about 13 mg / mL, at least about 14 mg / mL, or at least about 15 mg / mL.In more specific examples, the non-human animal comprising a humanized albumin locus (e.g., a homozygous humanized albumin locus) may be a mouse and may have a serum albumin level (e.g., serum human albumin level) of at least about 1 mg / mL, at least about 2 mg / mL, at least about 3 mg / mL, at least about 4 mg / mL, at least about 5 mg / mL, at least about 6 mg / mL, at least about 7 mg / mL, at least about 8 mg / mL, at least about 9 mg / mL, at least about 10 mg / mL, at least about 11 mg / mL, at least about 12 mg / mL, at least about 13 mg / mL, at least about 14 mg / mL, or at least about 15 mg / mL. In certain examples, the non-human animal (e.g., a mouse) comprising a humanized albumin locus (e.g., a homozygous humanized albumin locus) may have a serum albumin level (e.g., serum human albumin level) of from about 10 mg / mL to about 15 mg / mL. In any of the above examples, the albumin encoded by the humanized albumin locus may comprise, for example, a human albumin signal peptide. For example, in one instance, the entire albumin coding sequence of the endogenous albumin locus has been deleted and replaced with the corresponding human albumin sequence, or the region of the endogenous albumin locus from the start codon to the stop codon has been deleted and replaced with the corresponding human albumin sequence.
[0118] III. Method of using a non-human animal comprising a humanized albumin locus to evaluate the efficacy of a human albumin targeting reagent in vivo or ex vivo To evaluate or optimize the delivery or efficacy of a human albumin-targeting reagent (e.g., a therapeutic molecule or complex) in vivo or ex vivo, various methods are provided for using a non-human animal comprising a humanized albumin locus described elsewhere herein. Because the non-human animal comprises a humanized albumin locus, the non-human animal more accurately reflects the efficacy of a human albumin-targeting reagent. The non-human animals disclosed herein comprise a humanized endogenous albumin locus rather than a transgenic insertion of a human albumin sequence at a random genomic locus, and the humanized endogenous albumin locus can comprise orthologous human genomic albumin sequences from both the coding and non-coding regions rather than an artificial cDNA sequence, and thus such non-human animals are particularly useful for testing genome editing reagents designed to target the human albumin gene.
[0119] A. Methods for testing the efficacy of a human albumin-targeting reagent in vivo or ex vivo As described elsewhere herein, various methods are provided for using a non-human animal comprising a humanized albumin locus to evaluate the delivery or efficacy of a human albumin-targeting reagent in vivo. Such methods can include (a) introducing a human albumin-targeting reagent into the non-human animal (i.e., administering the human albumin-targeting reagent to the non-human animal), and (b) evaluating the activity of the human albumin-targeting reagent.
[0120] A human albumin-targeting reagent can be any biological or chemical agent that targets the human albumin locus (human albumin gene), human albumin mRNA, or human albumin protein. Examples of human albumin-targeting reagents are disclosed elsewhere in this specification and include, for example, genome editing agents. For example, a human albumin-targeting reagent can be a nucleic acid encoding an albumin-targeting nucleic acid (e.g., a CRISPR / Cas guide RNA, a short hairpin RNA (shRNA), or a small interfering RNA (siRNA)) or an albumin-targeting protein (e.g., a Cas protein such as Cas9, ZFN, or TALEN). Alternatively, a human albumin-targeting reagent can be an albumin-targeting antibody or antigen-binding protein, or any other macromolecule or small molecule that targets human albumin. In one example, a human albumin-targeting reagent is a genome editing agent such as a nuclease agent and / or an exogenous donor nucleic acid (e.g., a targeting vector). In a specific example, the genome editing agent can target intron 1, intron 12, or intron 13 of the human albumin gene. For example, the genome editing agent can target intron 1 of the human albumin gene.
[0121] Such human albumin-targeting reagents can be administered by any delivery method (e.g., AAV, LNP, or HDD) and any route of administration as disclosed in more detail elsewhere in this specification. The means and routes of administration for delivering therapeutic complexes and molecules are disclosed in more detail elsewhere in this specification. In certain methods, the reagent is delivered via AAV-mediated delivery. For example, AAV8 can be used to target the liver. In other certain methods, the reagent is delivered by LNP-mediated delivery. In other certain methods, the reagent is delivered by hydrodynamic delivery (HDD). The dosage can be any suitable dosage. For example, in some methods where the reagent (e.g., Cas9 mRNA and gRNA) is delivered by LNP-mediated delivery, the dosage can be between about 0.01 to about 10 mg / kg, between about 0.01 to about 5 mg / kg, between about 0.01 to about 4 mg / kg, between about 0.01 to about 3 mg / kg, between about 0.01 to about 2 mg / kg, between about 0.01 to about 1 mg / kg, between about 0.1 to about 10 mg / kg, between about 0.1 to about 6 mg / kg; between about 0.1 to about 5 mg / kg, between about 0.1 to about 4 mg / kg, between about 0.1 to about 3 mg / kg, between about 0.1 to about 2 mg / kg, between about 0.1 to about 1 mg / kg, between about 0.3 to about 10 mg / kg, between about 0.3 to about 6 mg / kg; between about 0.3 to about 5 mg / kg, between about 0.3 to about 4 mg / kg, between about 0.3 to about 3 mg / kg, between about 0.3 to about 2 mg / kg, between about 0.3 to about 1 mg / kg, about 0.1 mg / kg, about 0.3 mg / kg, about 1 mg / kg, about 2 mg / kg, or about 3 mg / kg. In certain examples, the dosage is between about 0.1 to about 6 mg / kg; between about 0.1 to about 3 mg / kg, or between about 0.1 to about 2 mg / kg.
[0122] Methods for evaluating the activity of human albumin-targeting reagents are well known and provided elsewhere in this specification. The evaluation of activity can be in any cell type, any tissue type, or any organ type, as disclosed elsewhere in this specification. In some methods, the evaluation of activity is performed in hepatocytes. If the albumin-targeting reagent is a genome editing reagent (e.g., a nuclease agent), such methods can include evaluating the modification of the humanized albumin locus. As an example, the evaluation can include measuring non-homologous end joining (NHEJ) activity at the humanized albumin locus. This can include, for example, measuring the frequency of insertions or deletions within the humanized albumin locus. For example, the evaluation can include sequencing (e.g., next-generation sequencing) of the humanized albumin locus in one or more cells isolated from a non-human animal. The evaluation can include isolating a target organ or tissue (e.g., the liver) or tissue from a non-human animal and evaluating the modification of the humanized albumin locus in the target organ or tissue. The evaluation can also include evaluating the modification of the humanized albumin locus in two or more different cell types within the target organ or tissue. Similarly, the evaluation can include isolating a non-target organ or tissue (e.g., two or more non-target organs or tissues) from a non-human animal and evaluating the modification of the humanized albumin locus in the non-target organ or tissue.
[0123] Such methods can also include measuring the expression level of mRNA produced by the humanized albumin locus or measuring the expression level of the protein encoded by the humanized albumin locus. For example, the protein level can be measured in a particular cell, tissue, or organ type (e.g., the liver), or the secretion level can be measured in serum. Methods for evaluating the expression of albumin mRNA or protein expressed from the humanized albumin locus are provided elsewhere in this specification and are well known. As an example, the BASESCOPE™ RNA in situ hybridization (ISH) assay can be used to quantify, for example, cell-specifically edited transcripts.
[0124] In some methods, the human albumin targeting reagent comprises an exogenous donor nucleic acid (e.g., a targeting vector). Such an exogenous donor nucleic acid can encode an exogenous protein that is not encoded or expressed by the wild-type endogenous albumin locus (e.g., can contain an inserted nucleic acid encoding an exogenous protein). In one example, the exogenous protein can be a heterologous protein comprising a human albumin signal peptide fused to a protein that is not encoded or expressed by the wild-type endogenous albumin locus. In one example, the exogenous protein encoded by the exogenous donor nucleic acid can be a heterologous protein comprising a human albumin signal peptide fused to a protein that is not encoded or expressed by the wild-type endogenous albumin locus once integrated into the humanized albumin locus. In some methods, the assessment can include measuring the expression of messenger RNA encoded by the exogenous donor nucleic acid. The assessment can also include measuring the expression of the exogenous protein. For example, the expression of the exogenous protein can be measured in the liver of a non-human animal or the serum level of the exogenous protein can be measured.
[0125] In some methods, a non-human animal comprising a humanized albumin locus described elsewhere herein further comprises an inactivated (knocked out) endogenous gene that is not the albumin locus, and optionally, the human albumin targeting reagent comprises an exogenous donor nucleic acid (e.g., a targeting vector) encoding an exogenous protein to replace the function of the inactivated endogenous gene. In a specific example, the inactivated endogenous gene is F9 and the exogenous protein is coagulation factor IX (e.g., human coagulation factor IX).
[0126] In some methods, the human albumin targeting reagent comprises (1) a nuclease agent designed to target a region of the human albumin gene and (2) an exogenous donor nucleic acid, which is designed to target the human albumin gene. The exogenous donor nucleic acid can, for example, encode an exogenous protein, and optionally, the protein encoded by the humanized endogenous albumin locus targeted by the exogenous donor nucleic acid is a heterologous protein that includes a human albumin signal peptide fused to the exogenous protein.
[0127] As one specific example, when the human albumin targeting reagent is a genome editing reagent (e.g., a nuclease agent), the percentage of editing at the humanized albumin locus (e.g., the total number of insertions or deletions observed relative to the total number of sequences read by a PCR reaction from a pool of lysed cells) can be evaluated (e.g., in hepatocytes).
[0128] The various methods provided above for evaluating activity in vivo can also be used to evaluate the activity of a human albumin targeting reagent ex vivo, as described elsewhere herein.
[0129] In some methods, the human albumin targeting reagent is a nuclease agent such as a CRISPR / Cas nuclease agent that targets the human albumin gene. Such methods can include, for example, (a) introducing into a non-human animal a nuclease agent designed to cleave the human albumin gene (e.g., a Cas protein such as Cas9 and a guide RNA designed to target a guide RNA target sequence of the human albumin gene) and (b) evaluating the modification of the humanized albumin locus.
[0130] For example, in the case of CRISPR / Cas nuclease, when the guide RNA forms a complex with the Cas protein and directs the Cas protein to the humanized albumin locus, modification of the humanized albumin locus is induced, and the Cas / guide RNA complex cleaves the guide RNA target sequence, triggering repair by the cell (e.g., via non-homologous end joining (NHEJ) if no donor sequence is present).
[0131] Optionally, two or more guide RNAs can be introduced, each designed to target a different guide RNA target sequence within the human albumin gene. For example, two guide RNAs can be designed to excise the genomic sequence between the two guide RNA target sequences. Modification of the humanized albumin locus is induced when the first guide RNA forms a complex with the Cas protein and directs the Cas protein to the humanized albumin locus, the second guide RNA forms a complex with the Cas protein and directs the Cas protein to the humanized albumin locus, the first Cas / guide RNA complex cleaves the first guide RNA target sequence, and the second Cas / guide RNA complex cleaves the second guide RNA target sequence, resulting in excision of the intervening sequence.
[0132] Alternatively or additionally, an exogenous donor nucleic acid (e.g., a targeting vector) that can recombine with and modify the human albumin gene is also introduced into the non-human animal. Optionally, a nuclease agent or Cas protein can be tethered to the exogenous donor nucleic acid as described elsewhere herein. Modification of the humanized albumin locus occurs, for example, when a guide RNA forms a complex with a Cas protein and directs the Cas protein to the humanized albumin locus, the Cas / guide RNA complex cleaves the guide RNA target sequence, and the humanized albumin locus recombines with the exogenous donor nucleic acid to modify the humanized albumin locus. The exogenous donor nucleic acid can recombine with the humanized albumin locus, for example, via homology-directed repair (HDR) or NHEJ-mediated insertion. Any type of exogenous donor nucleic acid can be used, examples of which are provided elsewhere herein.
[0133] In some methods, the human albumin targeting reagent comprises an exogenous donor nucleic acid (e.g., a targeting vector). Such an exogenous donor nucleic acid can encode an exogenous protein not encoded or expressed by the wild-type endogenous albumin locus (e.g., can contain an inserted nucleic acid encoding an exogenous protein). In one example, the exogenous protein can be a heterologous protein comprising a human albumin signal peptide fused to a protein not encoded or expressed by the wild-type endogenous albumin locus. For example, the exogenous donor nucleic acid can be a promoterless cassette containing a splice acceptor, and the exogenous donor nucleic acid can target the first intron of human albumin.
[0134] B. Methods for Optimizing Delivery or Efficacy of Human Albumin Targeting Reagents In Vivo or Ex Vivo To optimize the delivery of a human albumin targeting reagent to cells or non-human animals, or to optimize the activity or efficacy of a human albumin targeting reagent in vivo, various methods are provided. Such methods can include, for example, (a) first performing a method of testing the efficacy of a human albumin targeting reagent as described above in a first non-human animal or first cell comprising a humanized albumin locus, (b) changing a variable element and second performing the method in a second non-human animal (i.e., of the same species) or second cell comprising a humanized albumin locus using the changed variable element, and (c) comparing the activity of the human albumin targeting reagent of step (a) with the activity of the human albumin targeting reagent of step (b) and selecting the method that results in a higher activity.
[0135] Methods for measuring the delivery, efficacy, or activity of a human albumin targeting reagent are disclosed elsewhere in this specification. For example, such methods may include measuring the modification of the humanized albumin locus. A more effective modification of the humanized albumin locus may mean different depending on the desired effect in a non-human animal or cell. For example, a more effective modification of the humanized albumin locus may mean one or more or all of a higher level of modification, higher accuracy, higher consistency, or higher specificity. A higher level of modification (i.e., higher potency) of the humanized albumin locus refers to a higher percentage of cells targeted within a particular target cell type, within a particular target tissue, or within a particular target organ (e.g., the liver). Higher accuracy refers to a more accurate modification of the humanized albumin locus (e.g., a higher percentage of target cells having the same modification or the desired modification without additional unintended insertions and deletions (e.g., NHEJ indels)). Higher consistency, when multiple types of cells, tissues, or organs are targeted, refers to a more consistent modification of the humanized albumin locus among different types of target cells, tissues, or organs (e.g., modification of more cell types within the liver). When a particular organ is targeted, higher consistency can also refer to a more consistent modification throughout all locations within the organ (e.g., the liver). Higher specificity can refer to higher specificity for the targeted genomic locus or loci, higher specificity for the targeted cell type, higher specificity for the targeted tissue type, or higher specificity for the targeted organ. For example, an increase in genomic locus specificity refers to fewer modifications of off-target genomic loci (e.g., a low percentage of target cells having modifications in unintended off-target genomic loci instead of or in addition to modifications of the target genomic locus). Similarly, an increase in cell type, tissue, or organ type specificity refers to fewer modifications of off-target cell types, tissue types, or organ types when a particular cell type, tissue type, or organ type is targeted. (For example, when a particular organ (e.g., the liver) is targeted, there are fewer modifications of cells in organs or tissues that are not the intended target).
[0136] The variable element to be changed can be any parameter. As an example, the changed variable element can be a human albumin targeting reagent or the packaging or delivery method by which a plurality of reagents are introduced into cells or non-human animals. Examples of delivery methods such as LNP, HDD, and AAV are disclosed elsewhere in this specification. For example, the changed variable element can be an AAV serotype. Similarly, administration can include LNP-mediated delivery and the changed variable element can be an LNP formulation. As another example, the changed variable element can be the route of administration for introducing a human albumin targeting reagent or a plurality of reagents into cells or non-human animals. Examples of routes of administration such as intravenous, intravitreal, intrasubstantial, and nasal instillation are disclosed elsewhere in this specification.
[0137] As another example, the changed variable element can be the concentration or amount of the introduced human albumin targeting reagent or a plurality of reagents. As another example, the changed variable element can be the concentration or amount of one introduced human albumin targeting reagent (e.g., guide RNA, Cas protein, or exogenous donor nucleic acid) compared to the concentration or amount of another introduced human albumin targeting reagent (e.g., guide RNA, Cas protein, or exogenous donor nucleic acid).
[0138] As another example, the changed variable element can be the timing of introducing a human albumin targeting reagent or a plurality of reagents compared to the timing of evaluating the activity or effectiveness of the reagent. As another example, the changed variable element can be the number or frequency of times a human albumin targeting reagent or a plurality of reagents are introduced. As another example, the changed variable element can be the timing of one introduced human albumin targeting reagent (e.g., guide RNA, Cas protein, or exogenous donor nucleic acid) compared to the timing of another introduced human albumin targeting reagent (e.g., guide RNA, Cas protein, or exogenous donor nucleic acid).
[0139] As another example, the modified variable element can be in the form in which a human albumin targeting reagent or a plurality of reagents are introduced. For example, the guide RNA can be introduced in the form of DNA or RNA. The Cas protein (e.g., Cas9) can be introduced in the form of DNA, RNA, or protein (e.g., forming a complex with the guide RNA). The exogenous donor nucleic acid can be DNA, RNA, single-stranded, double-stranded, linear, circular, etc. Similarly, each component can include various combinations of modifications for reasons such as stability, reducing off-target effects, facilitating delivery, etc.
[0140] As another example, the modified variable element can be the human albumin targeting reagent or a plurality of reagents to be introduced. For example, when the human albumin targeting reagent includes a guide RNA, the modified variable element can introduce different guide RNAs having different sequences (e.g., targeting different guide RNA target sequences). Similarly, when the human albumin targeting reagent includes a Cas protein, the modified variable element can introduce different Cas proteins (e.g., different Cas proteins having different sequences, or nucleic acids encoding the same Cas protein amino acid sequence with different sequences (e.g., codon optimization)). Similarly, when the human albumin targeting reagent includes an exogenous donor nucleic acid, the modified variable element can introduce different exogenous donor nucleic acids having different sequences (e.g., different inserted nucleic acids or different homology arms (e.g., longer or shorter homology arms or homology arms targeting different regions of the human albumin gene)).
[0141] In certain examples, a human albumin targeting reagent comprises a Cas protein and a guide RNA designed to target a guide RNA target sequence of the human albumin gene. In such methods, the variable element that is altered can be the guide RNA sequence and / or the guide RNA target sequence. In some such methods, the Cas protein and the guide RNA can each be administered in the form of RNA, and the variable element that is altered can be the ratio of Cas mRNA to the guide RNA (e.g., in an LNP formulation). In some such methods, the variable element that is altered can be a guide RNA modification (e.g., a guide RNA having a modification is compared to an unmodified guide RNA).
[0142] C. Human Albumin Targeting Reagents A human albumin targeting reagent can be any reagent that targets the human albumin gene, human albumin mRNA, or human albumin protein. For example, it can be a genome editing reagent such as a nuclease agent that cleaves a target sequence within the human albumin gene and / or an exogenous donor sequence that recombines with the human albumin gene, it can be an antisense oligonucleotide that targets human albumin mRNA, it can be an antigen-binding protein that targets an epitope of human albumin protein, or it can be a small molecule that targets human albumin. A human albumin targeting reagent in the methods disclosed herein can be a known human albumin targeting reagent, a putative albumin targeting reagent (e.g., a candidate reagent designed to target human albumin), or a reagent that is being screened for human albumin targeting activity.
[0143] (1) Nuclease Agents that Target the Human Albumin Gene The human albumin targeting reagent can be a genome editing reagent such as a nuclease agent that cleaves a target sequence within the human albumin gene. The nuclease target sequence includes a DNA sequence at which a nick or double-strand break is induced by the nuclease agent. The target sequence of the nuclease agent can be endogenous (or native) to the cell, or the target sequence can be exogenous to the cell. A target sequence that is exogenous to the cell does not naturally exist within the genome of the cell. The target sequence can also be exogenous to the polynucleotide of interest that is desired to be placed at the target locus. In some cases, the target sequence exists only once in the genome of the host cell. In certain examples, the nuclease target sequence can be intron 1, intron 12, or intron 13 of the human albumin gene. For example, the nuclease target sequence can be in intron 1 of the human albumin gene.
[0144] The length of the target sequence can vary. For example, in the case of a zinc finger nuclease (ZFN) pair, it can include a target sequence of about 30 - 36 bp (i.e., about 15 - 18 bp for each ZFN), in the case of a transcription activator-like effector nuclease (TALEN) it can be about 36 bp, or in the case of a CRISPR / Cas9 guide RNA it can be about 20 bp of target sequence.
[0145] Any nuclease agent that induces a nick or double-strand break in a desired target sequence can be used in the methods and compositions disclosed herein. As long as the nuclease agent induces a nick or double-strand break in the desired target sequence, naturally occurring or native nuclease agents can be used. Alternatively, modified or engineered nuclease agents can be used. An "engineered nuclease agent" includes a nuclease that has been engineered (modified or derived) from its native form and specifically recognizes and induces a nick or double-strand break in a desired target sequence. Thus, an engineered nuclease agent can be derived from a naturally occurring native nuclease agent or can be artificially created or synthesized. Engineered nucleases can, for example, induce a nick or double-strand break in a target sequence that is not a sequence that would be recognized by a native (non-engineered or unmodified) nuclease agent. Modification of a nuclease agent can be as little as 1 amino acid in a protein nuclease or 1 nucleotide in a nucleic acid nuclease. Causing a nick or double-strand break in a target sequence or other DNA can be referred to herein as "cutting" or "cleaving" the target sequence or other DNA.
[0146] Active variants and fragments of the exemplified target sequences are also provided. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a given target sequence, and the active variants retain biological activity and thus can be recognized and cleaved by a nuclease agent in a sequence-specific manner. Assays for measuring double-strand breaks of target sequences by nuclease agents are known. See, for example, Frendewey et al. (2010) Methods in Enzymology 476:295-307, which is incorporated herein by reference in its entirety for all purposes.
[0147] The target sequence of the nuclease agent may be anywhere within or near the albumin locus. The target sequence may be located within the coding region of the albumin gene or within a regulatory region that affects gene expression. The target sequence of the nuclease agent may be located in an intron, exon, promoter, enhancer, regulatory region, or any non-protein-coding region.
[0148] One type of nuclease agent is the transcription activator-like effector nuclease (TALEN). TAL effector nucleases are a class of sequence-specific nucleases that can be used to introduce double-strand breaks at specific target sequences within the genomes of prokaryotes or eukaryotes. TAL effector nucleases are created by fusing a native or engineered transcription activator-like (TAL) effector, or a functional portion thereof, to the catalytic domain of an endonuclease such as FokI. The unique modular TAL effector DNA-binding domain enables the design of proteins with specific DNA recognition specificities. Thus, the DNA-binding domain of the TAL effector nuclease can be engineered to recognize a specific DNA target site and thus can be used to introduce double-strand breaks at a desired target sequence. See WO2010 / 079430, Morbitzer et al. (2010) PNAS 10.1073 / pnas.1013133107, Scholze & Boch (2010) Virulence 1:428-432, Christian et al. Genetics (2010) 186:757-761, Li et al. (2010) Nuc. Acids Res. (2010) doi:10.1093 / nar / gkq704, and Miller et al. (2011) Nature Biotechnology 29:143-148, each of which is hereby incorporated by reference in its entirety.
[0149] Examples of suitable TAL nucleases, and methods for preparing suitable TAL nucleases, are disclosed, for example, in US2011 / 0239315A1, US2011 / 0269234A1, US2011 / 0145940A1, US2003 / 0232410A1, US2005 / 0208489A1, US2005 / 0026157A1, US2005 / 0064474A1, US2006 / 0188987A1, and US2006 / 0063231A1, each of which is incorporated herein by reference in its entirety. In various embodiments, for example, a TAL effector nuclease that cleaves within or near a target nucleic acid sequence at a locus of interest or a genomic locus of interest is engineered, where the target nucleic acid sequence is the sequence that is modified by a targeting vector or near it. TAL nucleases suitable for use in the various methods and compositions provided herein include those that are specifically designed to bind to a target nucleic acid sequence modified by the targeting vectors described herein or near it.
[0150] In some TALENs, each monomer of the TALEN contains 33 - 35 TAL repeats that recognize a single base pair via two hypervariable residues. In some TALENs, the nuclease agent is a chimeric protein that includes a TAL repeat-based DNA binding domain operably linked to an independent nuclease such as the FokI endonuclease. For example, the nuclease agent can include a first TAL repeat-based DNA binding domain and a second TAL repeat-based DNA binding domain, each of the first and second TAL repeat-based DNA binding domains is operably linked to a FokI nuclease, the first and second TAL repeat-based DNA binding domains recognize two adjacent target DNA sequences in each strand of a target DNA sequence separated by a spacer sequence of varying lengths (12 - 20 bp), and the FokI nuclease subunits dimerize to create an active nuclease that makes a double-strand break at the target sequence.
[0151] The nuclease agents used in the various methods and compositions disclosed herein can further include zinc finger nucleases (ZFNs). In some ZFNs, each monomer of the ZFN contains three or more zinc finger-based DNA binding domains, and each zinc finger-based DNA binding domain binds to a 3bp subsite. In other ZFNs, the ZFN is a chimeric protein that includes a zinc finger-based DNA binding domain operably linked to an independent nuclease such as FokI endonuclease. For example, the nuclease agent can include a first ZFN and a second ZFN, each of the first and second ZFNs being operably linked to a FokI nuclease subunit, and the first and second ZFNs recognizing two adjacent target DNA sequences in each strand of a target DNA sequence separated by a spacer of about 5-7bp, and the FokI nuclease subunits dimerizing to produce an active nuclease that makes a double-strand break. For example, each of these can be referred to US2006 / 0246567, US2008 / 0182332, US2002 / 0081614, US2003 / 0021776, WO / 2002 / 057308A2, US2013 / 0123484, US2010 / 0291048, WO / 2011 / 017293A2, and Gaj et al. (2013) Trends in Biotechnology, 31(7):397-405, which are incorporated herein by reference.
[0152] Another type of nuclease agent is an engineered meganuclease. Meganucleases are classified into four families based on conserved sequence motifs, the families being the LAGLIDADG, GIY-YIG, H-N-H, and His-Cys box families. These motifs are involved in metal ion coordination and hydrolysis of the phosphodiester bond. Meganucleases are notable for their long target sequences and their tolerance of some sequence polymorphisms in their DNA substrates. Meganuclease domains are known in structure and function; see, for example, Guhan and Muniyappa (2003) Crit Rev Biochem Mol Biol 38:199-248, Lucas et al., (2001) Nucleic Acids Res 29:960-9, Jurica and Stoddard, (1999) Cell Mol Life Sci 55:1304-26, Stoddard, (2006) Q Rev Biophys 38:49-95, and Moure et al., (2002) Nat Struct Biol 9:764. In some instances, naturally occurring variants and / or engineered derivatives of meganucleases are used. Methods for modifying kinetics, cofactor interactions, expression, optimal conditions, and / or target sequence specificity, and screening for activity are known.For example, each of these is incorporated herein by reference in its entirety: Epinat et al., (2003) Nucleic Acids Res 31:2952-62, Chevalier et al., (2002) Mol Cell 10:895-905, Gimble et al., (2003) Mol Biol 334:993-1008, Seligman et al., (2002) Nucleic Acids Res 30:3870-9, Sussman et al., (2004) J Mol Biol 342:31-41, Rosen et al., (2006) Nucleic Acids Res 34:4791-800, Chames et al., (2005) Nucleic Acids Res 33:e178, Smith et al., (2006) Nucleic Acids Res 34:e149, Gruen et al., (2002) Nucleic Acids Res 30:e29, Chen and Zhao, (2005) Nucleic Acids Res 33:e154, WO2005 / 105989, WO2003 / 078619, WO2006 / 097854, WO2006 / 097853, WO2006 / 097784, and WO2004 / 031346.
[0153] For example, any meganuclease can be used that includes I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-CeuI, I-CeuAIIP, I-CreI, I-CrepsbIP, I-CrepsbIIP, I-CrepsbIIIP, I-CrepsbIVP, I-TliI, I-PpoI, PI-PspI, F-SceI, F-SceII, F-SuvI, F-TevI, F-TevII, I-AmaI, I-AniI, I-ChuI, I-CmoeI, I-CpaI, I-CpaII, I-CsmI, I-CvuI, I-CvuAIP, I-DdiI, I-DdiII, I-DirI, I-DmoI, I-HmuI, I-HmuII, I-HsNIP, I-LlaI, I-MsoI, I-NaaI, I-NanI, I-NcIIP, I-NgrIP, I-NitI, I-NjaI, I-Nsp236IP, I-PakI, I-PboIP, I-PcuIP, I-PcuAI, I-PcuVI, I-PgrIP, I-PobIP, I-PorI, I-PorIIP, I-PbpIP, I-SpBetaIP, I-ScaI, I-SexIP, I-SneIP, I-SpomI, I-SpomCP, I-SpomIP, I-SpomIIP, I-SquIP, I-Ssp6803I, I-SthPhiJP, I-SthPhiST3P, I-SthPhiSTe3bP, I-TdeIP, I-TevI, I-TevII, I-TevIII, I-UarAP, I-UarHGPAIP, I-UarHGPA13P, I-VinIP, I-ZbiIP, PI-MtuI, PI-MtuHIP, PI-MtuHIIP, PI-PfuI, PI-PfuII, PI-PkoI, PI-PkoII, PI-Rma43812IP, PI-SpBetaIP, PI-SceI, PI-TfuI, PI-TfuII, PI-ThyI, PI-TliI, PI-TliII, or an active variant or fragment thereof.
[0154] Meganucleases can recognize, for example, double-stranded DNA sequences of 12 to 40 base pairs. In some cases, the meganuclease recognizes one target sequence that exactly matches within the genome.
[0155] Some meganucleases are homing nucleases. One type of homing nuclease is the LAGLIDADG family of homing nucleases, including, for example, I-SceI, I-CreI, and I-Dmol.
[0156] The nuclease agent can further comprise a CRISPR / Cas system, as described in more detail below.
[0157] Also provided are active variants and fragments of the nuclease agent (i.e., engineered nuclease agent). Such active variants may have at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to the native nuclease agent, and the active variant retains the ability to cleave at a desired target sequence and thus retains nick or double-strand break-inducing activity. For example, any of the nuclease agents described herein may be modified from a native endonuclease sequence and designed to recognize and induce nicks or double-strand breaks at target sequences not recognized by the native nuclease agent. Thus, some engineered nucleases have specificity for inducing nicks or double-strand breaks at target sequences different from the corresponding native nuclease agent target sequences. Assays for nick or double-strand break-inducing activity are known and generally measure the overall activity and specificity of the endonuclease on a DNA substrate containing the target sequence.
[0158] The nuclease agent can be introduced into cells or non-human animals by any known means. The polypeptide encoding the nuclease agent may be introduced directly into cells or non-human animals. Alternatively, a polynucleotide encoding the nuclease agent may be introduced into cells or non-human animals. When the polynucleotide encoding the nuclease agent is introduced, the nuclease agent can be expressed intracellularly transiently, conditionally, or constitutively. The polynucleotide encoding the nuclease agent may be included in an expression cassette and may be operably linked to a conditional promoter, an inducible promoter, a constitutive promoter, or a tissue-specific promoter. Examples of promoters are discussed in more detail elsewhere in this specification. Alternatively, the nuclease agent can be introduced into cells as mRNA encoding the nuclease agent.
[0159] The polynucleotide encoding the nuclease agent can be stably integrated into the genome of the cell and operably linked to an active promoter intracellularly. Alternatively, the polynucleotide encoding the nuclease agent can be within a targeting vector.
[0160] When the nuclease agent is provided to the cell by introduction of a polynucleotide encoding the nuclease agent, such a polynucleotide encoding the nuclease agent can be modified to substitute codons that have a higher usage frequency in the target cell as compared to the polynucleotide sequence encoding the naturally occurring nuclease agent. For example, the polynucleotide encoding the nuclease agent can be modified to alternative codons that are used more frequently in a given eukaryotic cell of interest, including human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, as compared to the naturally occurring polynucleotide sequence.
[0161] (2) CRISPR / Cas system targeting the human albumin gene Certain types of human albumin targeting reagents can be a clustered regularly interspaced short palindromic repeat (CRISPR) / CRISPR-associated (Cas) system that targets the human albumin gene. The CRISPR / Cas system includes transcripts and other elements that are involved in the expression of the Cas gene or direct its activity. The CRISPR / Cas system can be, for example, a type I, type II, type III system, or a type V system (e.g., subtype V-A or subtype V-B). The CRISPR / Cas systems used in the compositions and methods disclosed herein can be non-naturally occurring. A "non-naturally occurring" system includes those that are modified or mutated from their naturally occurring state, or are at least substantially free from at least one other component with which they are naturally associated, or are associated with at least one other component with which they are not naturally associated, such as one or more components of the system that indicate the involvement of human assistance. For example, some CRISPR / Cas systems use a non-naturally occurring CRISPR complex that includes a non-naturally occurring gRNA and a Cas protein together, use a non-naturally occurring Cas protein, or use a non-naturally occurring gRNA.
[0162] Polynucleotides encoding Cas proteins and chimeric Cas proteins. Cas proteins generally include at least one RNA recognition or binding domain capable of interacting with guide RNA (gRNA). Cas proteins can also include nuclease domains (e.g., DNase domain or RNase domain), DNA binding domains, helicase domains, protein-protein interaction domains, dimerization domains, and other domains. Some such domains (e.g., the DNase domain) can be derived from native Cas proteins. Other such domains can be added to create modified Cas proteins. The nuclease domain has catalytic activity against nucleic acid cleavage, including cleavage of covalent bonds in nucleic acid molecules. The cleavage can generate blunt ends or staggered ends, which can be single-stranded or double-stranded. For example, wild-type Cas9 protein usually generates blunt-end cleavage products. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) can result in cleavage products containing a 5-nucleotide 5’ overhang, and the cleavage occurs 18 base pairs after the PAM sequence of the non-targeted strand and 23 bases after the targeted strand. The Cas protein can have full cleavage activity and be able to create double-strand breaks at target genomic loci (e.g., double-strand breaks including blunt ends), or it can be a nickase that creates single-strand breaks at target genomic loci.
[0163] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, as well as homologs or modified versions thereof.
[0164] Exemplary Cas proteins are Cas9 proteins or proteins derived from Cas9 proteins. Cas9 proteins are derived from the type II CRISPR / Cas system and typically share four important motifs with a conserved architecture. Motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins include those from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., derived from Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of Cas9 family members are described in WO2014 / 131833, which is hereby incorporated by reference in its entirety for all purposes. Cas9 derived from S. pyogenes (SpCas9, assigned SwissProt accession number Q99ZW2) is an exemplary Cas9 protein. S.Cas9 derived from Staphylococcus aureus (SaCas9, UniProt accession number J7RUA5 assigned) is another exemplary Cas9 protein. Cas9 derived from Campylobacter jejuni (CjCas9, UniProt accession number Q0P897 assigned) is another exemplary Cas9 protein. See, for example, Kim et al. (2017) Nat. Comm. 8:14500, which is hereby incorporated by reference in its entirety for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Exemplary Cas9 proteins include, consist essentially of, or consist of SEQ ID NO: 38. Exemplary DNA encoding a Cas9 protein includes, consist essentially of, or consist of SEQ ID NO: 39.
[0165] Another example of a Cas protein is the Cpf1 (CRISPR from Prevotella and Francisella 1) protein. Cpf1 is a large protein (about 1300 amino acids) that contains a RuvC-like nuclease domain homologous to the corresponding domain of Cas9, along with a counterpart of the characteristic arginine-rich cluster of Cas9. However, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein, and in contrast to Cas9, which contains a long insertion including the HNH domain, the RuvC-like domains are adjacent in the Cpf1 sequence. See, for example, Zetsche et al. (2015) Cell 163(3):759-771, which is hereby incorporated by reference in its entirety for all purposes. Exemplary Cpf1 proteins are derived from Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae. Cpf1 from Francisella novicida U112 (FnCpf1; assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.
[0166] A Cas protein can be a wild-type protein (i.e., that which occurs naturally), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. A Cas protein can also be an active variant or fragment with respect to the catalytic activity of a wild-type or modified Cas protein. An active variant or fragment with respect to catalytic activity can include at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a wild-type or modified Cas protein or a portion thereof, and the active variant retains the ability to cleave at a desired cleavage site and thus retains nick-inducing or double-strand break-inducing activity. Assays for nick-inducing or double-strand break-inducing activity are known and generally measure the overall activity and specificity of a Cas protein on a DNA substrate containing a cleavage site.
[0167] A Cas protein can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. A Cas protein can also be modified to alter other activities or properties of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be cleaved to remove domains that are not essential for the function of the protein, or the activity or properties of the Cas protein can be optimized (e.g., enhanced or reduced).
[0168] An example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity variant (N497A / R661A / Q695A / Q926A) of Streptococcus pyogenes Cas9 that incorporates modifications designed to reduce non-specific DNA contacts. See, for example, Kleinstiver et al. (2016) Nature 529(7587):490-495, which is hereby incorporated by reference in its entirety for all purposes. Another example of a modified Cas protein is the modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, for example, Slaymaker et al. (2016) Science 351(6268):84-88, which is hereby incorporated by reference in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A.
[0169] A Cas protein can include at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 protein generally contains a RuvC-like domain that cleaves both strands of the target DNA, perhaps in a dimeric conformation. A Cas protein can also include at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 protein generally contains a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave different strands of double-stranded DNA to create a double-strand break in the DNA. See, for example, Jinek et al. (2012) Science 337:816-821, which is hereby incorporated by reference in its entirety for all purposes.
[0170] By deleting or mutating one or more or all of the nuclease domains, they can be made non-functional or the nuclease activity can be reduced. For example, if one of the nuclease domains is deleted or mutated in the Cas9 protein, the resulting Cas9 protein is called a nickase and can generate single-strand breaks within double-stranded target DNA but cannot generate double-strand breaks (i.e., it can cleave the complementary or non-complementary strand but not both). When both nuclease domains are deleted or mutated, the resulting Cas protein (e.g., Cas9) has a reduced ability to cleave both strands of double-stranded DNA (e.g., nuclease-null or nuclease-inactive Cas protein, or Cas protein without catalytic activity (dCas)). An example of a mutation that converts Cas9 to a nickase is the D10A (from aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Similarly, H939A (from histidine to alanine at amino acid position 839), H840A (from histidine to alanine at amino acid position 840), or N863A (from asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert Cas9 to a nickase. Other examples of mutations that convert Cas9 to a nickase include the corresponding mutations in Cas9 from S. thermophilus. See, for example, Sapranauskas et al. (2011) Nucleic Acids 39):9275-9282 and WO2013 / 141680, each of which is incorporated herein by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, mutagenesis via PCR, or total gene synthesis. Other examples of mutations that generate nickases can be found, for example, in WO2013 / 176772 and WO2013 / 142578, each of which is incorporated herein by reference in its entirety for all purposes.When all of the nuclease domains are deleted or mutated in the Cas protein (e.g., when both nuclease domains are deleted or mutated in the Cas9 protein), the resulting Cas protein (e.g., Cas9) has a reduced ability to cleave both strands of double-stranded DNA (e.g., nuclease-null or nuclease-inactive Cas protein). One specific example is the D10A / H840A S. pyogenes Cas9 double mutant, or the corresponding double mutant of Cas9 from another species when optimally aligned with S. pyogenes Cas9. Another specific example is the D10A / N863A S. pyogenes Cas9 double mutant, or the corresponding double mutant of Cas9 from another species when optimally aligned with S. pyogenes Cas9.
[0171] Examples of inactivating mutations in the catalytic domain of the Cas9 protein of Staphylococcus aureus are also known. For example, the Cas9 enzyme of Staphylococcus aureus (SaCas9) can include a substitution at position N580 (e.g., N580A substitution) and a substitution at position D10 (e.g., D10A substitution), resulting in a nuclease-inactive Cas protein. See, for example, WO2016 / 106236, which is hereby incorporated by reference in its entirety for all purposes.
[0172] Examples of inactivating mutations in the catalytic domain of Cpf1 proteins are also known. For Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at position 908, 993, or 1263 of AsCpf1 or the corresponding positions of Cpf1 orthologs, or at position 832, 925, 947, or 1180 of LbCpf1 or the corresponding positions of Cpf1 orthologs. Such mutations can include, for example, one or more of the mutations D908A, E993A, and D1263A of AsCpf1 or the corresponding mutations of Cpf1 orthologs, or D832A, E925A, D947A, and D1180A of LbCpf1 or the corresponding mutations of Cpf1 orthologs. See, for example, US2016 / 0208243, which is hereby incorporated by reference in its entirety for all purposes.
[0173] A Cas protein (e.g., a nuclease-active Cas protein or a nuclease-inactive Cas protein) can also be operably linked to a heterologous polypeptide as a fusion protein. For example, a Cas protein can be fused to a cleavage domain or an epigenetic modification domain. See WO2014 / 089290, which is hereby incorporated by reference in its entirety for all purposes. A Cas protein can also be fused to a heterologous polypeptide that provides an increase or decrease in stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally within the Cas protein.
[0174] As an example, the Cas protein may be fused to one or more heterologous polypeptides that provide intracellular localization. Such heterologous polypeptides may include, for example, one or more nuclear localization signals (NLSs) such as a monopartite SV40 NLS and / or a bipartite alpha-importin NLS for targeting the nucleus, a mitochondrial localization signal for targeting mitochondria, an ER retention signal, and the like. See, for example, Lange et al. (2007) J. Biol. Chem. 282:5101-5105, which is incorporated herein by reference in its entirety for all purposes. Such intracellular localization signals may be located at the N-terminus, C-terminus, or anywhere within the Cas protein. The NLS can contain a stretch of basic amino acids and can be a monopartite or bipartite sequence. Optionally, the Cas protein can contain two or more NLSs, with an NLS (e.g., an alpha-importin NLS or a monopartite NLS) at the N-terminus and an NLS (e.g., an SV40 NLS or a bipartite NLS) at the C-terminus. The Cas protein can also contain two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.
[0175] The Cas protein can also be operably linked to a cell-permeable domain or protein transduction domain. For example, the cell-permeable domain can be derived from an HIV-1 TAT protein, a TLM cell-permeable motif from hepatitis B virus, MPG, Pep-1, VP22, a cell-permeable peptide from herpes simplex virus, or a polyarginine peptide sequence. See, for example, WO2014 / 089290 and WO2013 / 176772, each of which is incorporated herein by reference in its entirety for all purposes. The cell-permeable domain can be located at the N-terminus, C-terminus, or anywhere within the Cas protein.
[0176] The Cas protein can also be operably linked to a heterologous polypeptide such as a fluorescent protein, a purification tag, or an epitope tag to facilitate tracking or purification. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent protein (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent protein (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent protein (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato), and any other suitable fluorescent protein. Examples of tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0177] The Cas protein may also be tethered to an exogenous donor nucleic acid or a labeled nucleic acid. Such tethering (i.e., physical binding) can be achieved via covalent or non-covalent interactions, and the tethering can be direct (e.g., via direct fusion or chemical bonding that can be achieved by modification of cysteine or lysine residues on the protein or intein modification), or via one or more intervening linker or adapter molecules such as streptavidin or an aptamer. See, for example, Pierce et al. (2005) Mini Rev. Med. Chem. 5(1):41-55, Duckworth et al. (2007) Angew. Chem. Int. Ed. Engl. 46(46):8819-8822, Schaeffer and Dixon (2009) Australian J. Chem. 62(10):1328-1332, Goodman et al. (2009) Chembiochem. 10(9):1551-1557, and Khatwani et al. (2012) Bioorg. Med. Chem. 20(14):4532-4539, each of which is hereby incorporated by reference in its entirety for all purposes. Non-covalent strategies for synthesizing protein-nucleic acid conjugates include the biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by connecting appropriately functionalized nucleic acids and proteins using various chemicals. Some of these chemicals involve direct binding of oligonucleotides to amino acid residues on the protein surface (e.g., lysine amine or cysteine thiol), while other more complex schemes require post-translational modification of the protein or the involvement of catalytic or reactive protein domains. Methods for covalent attachment of proteins to nucleic acids can include, for example, chemical cross-linking of oligonucleotides to protein lysine or cysteine residues, expressed protein ligation, chemoenzymatic methods, and the use of photoaptamers. The exogenous donor nucleic acid or labeled nucleic acid may be tethered to the C-terminus, N-terminus, or internal region of the Cas protein.In one example, the exogenous donor nucleic acid or labeled nucleic acid is tethered to the C-terminus or N-terminus of the Cas protein. Similarly, the Cas protein may be tethered to the 5'-end, 3'-end, or internal region of the exogenous donor nucleic acid or labeled nucleic acid. That is, the exogenous donor nucleic acid or labeled nucleic acid may be tethered in any orientation and polarity. For example, the Cas protein may be tethered to the 5'-end or 3'-end of the exogenous donor nucleic acid or labeled nucleic acid.
[0178] The Cas protein may be provided in any form. For example, the Cas protein may be provided in the form of a protein such as a Cas protein complexed with a gRNA. Alternatively, the Cas protein may be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon-optimized for efficient translation into protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to alternative codons that are used more frequently in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the naturally occurring polynucleotide sequence. When the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein can be expressed intracellularly transiently, conditionally, or constitutively.
[0179] The Cas protein provided as mRNA can be modified to improve stability and / or immunogenicity properties. The modification can be performed on one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. For example, a capped and polyadenylated Cas mRNA containing N1-methylpseudouridine can be used. Similarly, the Cas mRNA can be modified by depleting uridine using synonymous codons.
[0180] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably linked to an active promoter within the cell. Alternatively, the nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. The expression construct includes any nucleic acid construct that can direct the expression of a gene or other nucleic acid sequence of interest (e.g., the Cas gene) and transfer such a nucleic acid sequence of interest to a target cell. For example, the nucleic acid encoding the Cas protein may be present within a targeting vector containing a nucleic acid insert and / or a vector containing DNA encoding a gRNA. Alternatively, it may be a vector or plasmid separate from the targeting vector containing the nucleic acid insert and / or separate from the vector containing DNA encoding a gRNA. Promoters that can be used in the expression construct include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, embryonic stem (ES) cells, or zygotes. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Optionally, the promoter may be a bidirectional promoter that drives the expression of both the Cas protein in one direction and the guide RNA in the other direction. Such a bidirectional promoter may be composed of (1) a complete conventional unidirectional Pol III promoter containing three external control elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box, and (2) a second basic Pol III promoter containing a PSE and a TATA box fused in the reverse direction to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and the TATA box, and the promoter can be made bidirectional by creating a hybrid promoter in which the reverse transcription is controlled by adding a PSE and a TATA box derived from the U6 promoter. See, for example, US2016 / 0074535, which is hereby incorporated by reference in its entirety for all purposes.By using a bidirectional promoter to simultaneously express genes encoding Cas protein and guide RNA, a compact expression cassette can be generated to facilitate delivery.
[0181] The guide RNA, also known as "guide RNA" or "gRNA", is an RNA molecule that binds to a Cas protein (e.g., Cas9 protein) and targets the Cas protein to a specific position within the target DNA. The guide RNA can include two segments: a "DNA targeting segment" and a "protein binding segment". A "segment" includes a section or region of a molecule, such as a continuous stretch of nucleotides in an RNA. Some gRNAs, such as those for Cas9, can include two separate RNA molecules: an "activator RNA" (e.g., tracrRNA) and a "targeter RNA" (e.g., CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which can also be referred to as "single molecule gRNAs", "single guide RNAs", or "sgRNAs". See, for example, WO2013 / 176772, WO2014 / 065596, WO2014 / 089290, WO2014 / 093622, WO2014 / 099750, WO2013 / 142578, and WO2014 / 131833, each of which is hereby incorporated by reference in its entirety for all purposes. For example, in the case of Cas9, a single guide RNA can include a crRNA fused to a tracrRNA (e.g., via a linker). For example, in the case of Cpf1, only the crRNA is required to effect binding to and / or cleavage of the target sequence. The terms "guide RNA" and "gRNA" include both bimolecular (i.e., modular) gRNAs and single molecule gRNAs.
[0182] Exemplary two-molecule gRNAs include a crRNA-like (“CRISPR RNA” or “targeter RNA” or “crRNA” or “crRNA repeat”) molecule and a corresponding tracrRNA-like (“trans-acting CRISPR RNA” or “activator RNA” or “tracrRNA”) molecule. The crRNA includes both a DNA targeting segment (single-stranded) of the gRNA and a stretch of nucleotides (i.e., the crRNA tail) that forms half of the dsRNA duplex of the protein-binding segment of the gRNA. Examples of crRNA tails located downstream (3’) of the DNA targeting segment include, consist essentially of, or consist of GUUUUAGAGCUAUGCU (SEQ ID NO: 40). Any of the DNA targeting segments disclosed herein can bind to the 5’ end of SEQ ID NO: 40 to form a crRNA.
[0183] The corresponding tracrRNA (activator RNA) includes a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. The stretch of nucleotides of the crRNA is complementary to the stretch of nucleotides of the tracrRNA and hybridizes therewith to form the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be said to have a corresponding tracrRNA. Examples of tracrRNA sequences include, consist essentially of, or consist of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 41).
[0184] In systems where both crRNA and tracrRNA are required, the crRNA and the corresponding tracrRNA hybridize to form the gRNA. In systems where only crRNA is required, the crRNA can be the gRNA. The crRNA further provides a single-stranded DNA targeting segment that hybridizes to the complementary strand of the target DNA. When used for intracellular modification, the exact sequence of a given crRNA or tracrRNA molecule can be designed to be specific to the species in which the RNA molecule is used. See, for example, Mali et al. (2013) Science 339:823-826, Jinek et al. (2012) Science 337:816-821, Hwang et al. (2013) Nat. Biotechnol. 31:227-229, Jiang et al. (2013) Nat. Biotechnol. 31:233-239, and Cong et al. (2013) Science 339:819-823, each of which is hereby incorporated by reference in its entirety for all purposes.
[0185] The DNA targeting segment (crRNA) of a given gRNA contains a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA targeting segment of the gRNA interacts with the target DNA in a sequence-specific manner via hybridization (i.e., base pairing). Thus, the nucleotide sequence of the DNA targeting segment may vary and determines the position within the target DNA where the gRNA and the target DNA interact. The DNA targeting segment of a desired gRNA can be modified to hybridize to any desired sequence within the target DNA. Naturally occurring crRNAs often contain a targeting segment 21-72 nucleotides in length flanked by two direct repeats (DRs) 21-46 nucleotides in length, although this varies depending on the CRISPR / Cas system and organism (see, e.g., WO2014 / 131833, which is hereby incorporated by reference in its entirety for all purposes). In the case of S. pyogenes, the DR is 36 nucleotides long and the targeting segment is 30 nucleotides long. The DR located 3’ is complementary to the corresponding tracrRNA, hybridizes to it, which in turn binds to the Cas protein.
[0186] DNA targeting segments can have a length of, for example, at least about 12, 15, 17, 18, 19, 20, 25, 30, 35, or 40 nucleotides. Such DNA targeting segments can have a length of, for example, from about 12 to about 100, from about 12 to about 80, from about 12 to about 50, from about 12 to about 40, from about 12 to about 30, from about 12 to about 25, or from about 12 to about 20 nucleotides. For example, a DNA targeting segment can be from about 15 to about 25 nucleotides (e.g., from about 17 to about 20 nucleotides, or about 17, 18, 19, or 20 nucleotides). For example, reference is made to US2016 / 0024523, which is hereby incorporated by reference in its entirety for all purposes. In the case of Cas9 from S. pyogenes, typical DNA targeting segments are 16 - 20 nucleotides in length, or 17 - 20 nucleotides in length. In the case of Cas9 from S. aureus, typical DNA targeting segments are 21 - 23 nucleotides in length. In the case of Cpf1, typical DNA targeting segments are at least 16 nucleotides in length or at least 18 nucleotides in length.
[0187] The tracrRNA can be in any form (e.g., full-length tracrRNA or active partial tracrRNA) and can be of various lengths. They can include primary transcripts or processed forms. For example, the tracrRNA (as part of a single guide RNA or as a separate molecule as part of a two-molecule gRNA) can contain, consist essentially of, or consist of all or part of a wild-type tracrRNA sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracrRNA sequence). Examples of wild-type tracrRNA sequences from S. pyogenes include 171-nucleotide, 89-nucleotide, 75-nucleotide, and 65-nucleotide versions. See, for example, Deltcheva et al. (2011) Nature 471:602-607, WO2014 / 093661, each of which is hereby incorporated by reference in its entirety for all purposes. Examples of tracrRNA within a single guide RNA (sgRNA) include tracrRNA segments found within the +48, +54, +67, and +85 versions of the sgRNA, where “+n” indicates that the maximum +n nucleotides of the wild-type tracrRNA are included in the sgRNA. See US8,697,359, which is hereby incorporated by reference in its entirety for all purposes.
[0188] The percentage of complementarity between the DNA targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). The percentage of complementarity between the DNA targeting segment and the complementary strand of the target DNA can be at least 60% over about 20 consecutive nucleotides. As an example, the percentage of complementarity between the DNA targeting segment and the complementary strand of the target DNA can be 100% over 14 consecutive nucleotides at the 5' end of the complementary strand of the target DNA, and can be as low as 0% over the remaining portion. In such a case, the DNA targeting segment can be considered to be 14 nucleotides in length. As another example, the percentage of complementarity between the DNA targeting segment and the complementary strand of the target DNA can be 100% over 7 consecutive nucleotides at the 5' end of the complementary strand of the target DNA, and can be as low as 0% over the remaining portion. In such a case, the DNA targeting segment can be considered to be 7 nucleotides in length. In some guide RNAs, at least 17 nucleotides within the DNA targeting segment are complementary to the complementary strand of the target DNA. For example, the DNA targeting segment can be 20 nucleotides in length and can contain 1, 2, or 3 mismatches with the complementary strand of the target DNA. In one example, the mismatch is not adjacent to the region of the complementary strand corresponding to the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (e.g., the mismatch is at the 5' end of the DNA targeting segment of the guide RNA, or the mismatch is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 base pairs away from the region of the complementary strand corresponding to the PAM sequence).
[0189] The protein-binding segment of the gRNA can include two stretches of nucleotides that are complementary to each other. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of the gRNA of interest interacts with the Cas protein, and the gRNA orients the bound Cas protein to a specific nucleotide sequence within the target DNA via the DNA-targeting segment.
[0190] A single guide RNA can include a DNA target segment bound to a scaffold sequence (i.e., the protein-binding or Cas-binding sequence of the guide RNA). For example, such a guide RNA can have a 5’ DNA targeting segment and a 3’ scaffold sequence. Exemplary scaffold sequences include, consist essentially of, or consist of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU (version 1; SEQ ID NO: 42), GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 2; SEQ ID NO: 43), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 3; SEQ ID NO: 44), and GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 4; SEQ ID NO: 45). A guide RNA that targets any guide RNA target sequence can include, for example, a 5’ DNA targeting segment of the guide RNA fused to any of the exemplary guide RNA scaffold sequences at the 3’ end of the guide RNA. That is, any of the DNA targeting segments disclosed herein can bind to the 5’ end of any one of SEQ ID NOS: 42-45 to form a single guide RNA (chimeric guide RNA). Guide RNA versions 1, 2, 3, and 4 disclosed elsewhere herein refer to DNA target segments (i.e., guide sequences or guides) bound to scaffold versions 1, 2, 3, and 4, respectively.
[0191] The guide RNA may contain modifications or sequences that provide additional desirable features (e.g., modified or regulated stability; intracellular targeting; tracking by fluorescent labeling; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5’ cap (e.g., 7-methylguanylic acid cap (m7G)), a 3’ polyadenylation tail (i.e., 3’ poly(A) tail), riboswitch sequences (e.g., enabling regulation of stability and / or accessibility by proteins and / or protein complexes), stability control sequences, sequences that form dsRNA duplexes (i.e., hairpins), modifications or sequences that target the RNA to an intracellular location (e.g., nucleus, mitochondria, chloroplast, etc.), modifications or sequences that provide tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, sequences that enable fluorescent detection, etc.), modifications or sequences that provide a binding site for a protein (e.g., a protein that acts on DNA, including DNA methyltransferase, DNA demethylase, histone acetyltransferase, histone deacetylase, etc.), and combinations thereof. Other examples of modifications include engineered stem-loop duplex structures, engineered bulge regions, engineered hairpins at the 3’ of stem-loop duplexes, or any combination thereof. See, for example, US2015 / 0376586, which is hereby incorporated by reference in its entirety for all purposes. A bulge can be a region of non-paired nucleotides within a duplex composed of a crRNA-like region and a minimal tracrRNA-like region. A bulge can include, on one side of the duplex, 5’-XXXY-3’ that is not paired, where X is any purine and Y is a nucleotide that can form a wobble pair with a nucleotide in the opposite strand and a region of unpaired nucleotides on the opposite side of the duplex.
[0192] Unmodified nucleic acids can be prone to degradation. Exogenous nucleic acids can also induce innate immune responses. Modifications can help introduce stability and reduce immunogenicity. Guide RNAs can include, for example: (1) modifications or substitutions of one or both of the unbridged phosphate oxygens and / or one or more of the bridging phosphate oxygens in the phosphodiester backbone linkage, (2) modifications or substitutions of components of the ribose sugar such as modification or substitution of the 2'-hydroxyl of the ribose sugar, (3) substitution by a dephospho linker of the phosphate moiety, (4) modification or substitution of naturally occurring nucleobases, (5) substitution or modification of the ribose-phosphate backbone, (6) modification of the 3' or 5' end of the oligonucleotide (e.g., removal, modification, or substitution of the terminal phosphate group, or complexation of the moiety), and (7) one or more of the sugar modifications, and can include modified nucleosides and modified nucleotides. Other possible guide RNA modifications include modification or substitution of uracil or poly-uracil tracts. See, for example, WO2015 / 048577 and US2016 / 0237455, each of which is hereby incorporated by reference in its entirety for all purposes. Similar modifications can be made to Cas-encoding nucleic acids such as Cas mRNA.
[0193] As an example, the nucleotides at the 5' or 3' end of the guide RNA may contain phosphorothioate linkages (e.g., the bases may have modified phosphate groups that are phosphorothioate groups). For example, the guide RNA may contain phosphorothioate linkages between 2, 3, or 4 terminal nucleotides at the 5' or 3' end of the guide RNA. As another example, the nucleotides at the 5' and / or 3' end of the guide RNA may have 2'-O-methyl modifications. For example, the guide RNA may contain 2'-O-methyl modifications on 2, 3, or 4 terminal nucleotides at the 5' and / or 3' end (e.g., the 5' end) of the guide RNA. For example, reference is made to WO2017 / 173054A1 and Finn et al. (2018) Cell Reports 22:1-9, each of which is incorporated herein by reference in its entirety for all purposes. In one specific example, the guide RNA contains 2'-O-methyl analogs and 3' phosphorothioate nucleotide internucleotide linkages in the first 3 5' and 3' terminal RNA residues. In another specific example, the guide RNA has all 2'OH groups that do not interact with the Cas9 protein replaced with 2'-O-methyl analogs, and the tail region of the guide RNA that minimally interacts with Cas9 is modified with 5' and 3' phosphorothioate nucleotide internucleotide linkages. For example, reference is made to Yin et al. (2017) Nat. Biotech. 35(12):1179-1187, which is incorporated herein by reference in its entirety for all purposes. Other examples of modified guide RNAs are provided, for example, in WO2018 / 107028A1, which is incorporated herein by reference in its entirety for all purposes.
[0194] The guide RNA can be provided in any form. For example, the gRNA can be provided in the form of RNA as either two molecules (separate crRNA and tracrRNA) or one molecule (sgRNA), and optionally in the form of a complex with a Cas protein. The gRNA can also be provided in the form of DNA encoding the gRNA. The DNA encoding the gRNA can encode a single RNA molecule (sgRNA) or separate RNA molecules (e.g., separate crRNA and tracrRNA). In the latter case, the DNA encoding the gRNA can be provided as one DNA molecule or as separate DNA molecules encoding crRNA and tracrRNA, respectively.
[0195] When the gRNA is provided in the form of DNA, the gRNA can be expressed intracellularly transiently, conditionally, or constitutively. The DNA encoding the gRNA can be stably integrated in the genome of the cell and operably linked to an active promoter in the cell. Alternatively, the DNA encoding the gRNA can be operably linked to a promoter in an expression construct. For example, the DNA encoding the gRNA can be within a vector containing a heterologous nucleic acid such as a nucleic acid encoding a Cas protein. Alternatively, it can be a separate vector or plasmid from the vector containing the nucleic acid encoding the Cas protein. Promoters that can be used in such expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, progenitor cells with restricted development, induced pluripotent stem (iPS) cells, or one-cell stage embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Such promoters can also be, for example, bidirectional promoters. Specific examples of suitable promoters include RNA polymerase III promoters such as the human U6 promoter, the rat U6 polymerase III promoter, or the mouse U6 polymerase III promoter.
[0196] Alternatively, the gRNA can be prepared in a variety of other ways. For example, the gRNA can be prepared by in vitro transcription using, for example, T7 RNA polymerase (see, for example, WO2014 / 089290 and WO2014 / 065596, each of which is incorporated herein by reference in its entirety for all purposes). The guide RNA can also be a synthetically produced molecule prepared by chemical synthesis.
[0197] The guide RNA (or nucleic acid encoding the guide RNA) can be a composition comprising one or more guide RNAs (e.g., 1, 2, 3, 4, or more guide RNAs) and a carrier that increases the stability of the guide RNA (e.g., extends the period during which degradation products remain below a threshold, e.g., less than 0.5% by weight of the starting nucleic acid or protein, under certain storage conditions (e.g., -20 °C, 4 °C, or ambient temperature), or increases stability in vivo). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, and lipid nanotubes. Such compositions can further comprise a Cas protein such as Cas9 protein, or a nucleic acid encoding the Cas protein.
[0198] Guide RNA target sequence. The target DNA of the guide RNA includes the nucleic acid sequence present in the DNA to which the DNA target segment of the gRNA binds when sufficient conditions for binding exist. Suitable DNA / RNA binding conditions include the physiological conditions normally present in cells. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), which is hereby incorporated by reference in its entirety for all purposes). The strand of the target DNA that is complementary to and hybridizes with the gRNA can be referred to as the "complementary strand," and the strand of the target DNA that is complementary to the "complementary strand" (and thus not complementary to the Cas protein or gRNA) can be referred to as the "non-complementary strand" or "template strand."
[0199] The target DNA includes both the sequence on the complementary strand to which the guide RNA hybridizes and the corresponding sequence on the non-complementary strand (e.g., adjacent to the protospacer adjacent motif (PAM)). As used herein, the term "guide RNA target sequence" specifically refers to the sequence on the non-complementary strand that corresponds to (i.e., is the reverse complement of) the sequence to which the guide RNA hybridizes on the complementary strand. That is, the guide RNA target sequence refers to the sequence on the non-complementary strand adjacent to the PAM (e.g., upstream or 5' of the PAM in the case of Cas9). The guide RNA target sequence is equivalent to the DNA targeting segment of the guide RNA but contains thymine instead of uracil. As an example, the guide RNA target sequence of the SpCas9 enzyme may refer to the sequence upstream of the 5'-NGG-3' PAM on the non-complementary strand. The guide RNA is designed to have complementarity to the complementary strand of the target DNA, and hybridization between the DNA targeting segment of the guide RNA and the complementary strand of the target DNA promotes the formation of the CRISPR complex. Complete complementarity is not necessarily required as long as there is sufficient complementarity to cause hybridization and promote the formation of the CRISPR complex. When a guide RNA is referred to herein as targeting a guide RNA target sequence, it means that the guide RNA hybridizes to the complementary strand sequence of the target DNA that is the reverse complement of the guide RNA target sequence on the non-complementary strand.
[0200] The target DNA or guide RNA target sequence can include any polynucleotide and can be located, for example, within the nucleus or cytoplasm of a cell or within an organelle of a cell such as a mitochondrion or chloroplast. The target DNA or guide RNA target sequence can be any nucleic acid sequence that is endogenous or exogenous to the cell. The guide RNA target sequence can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence), or can include both. In certain examples, the guide RNA target sequence can be intron 1, intron 12, or intron 13 of the human albumin gene. For example, the guide RNA target sequence can be in intron 1 of the human albumin gene.
[0201] Site-specific binding and cleavage of target DNA by Cas proteins can occur at positions determined by both (i) the base pairing complementarity between the guide RNA and the complementary strand of the target DNA, and (ii) a short motif called the protospacer adjacent motif (PAM) in the non-complementary strand of the target DNA. A PAM can be adjacent to the guide RNA target sequence. Optionally, the guide RNA target sequence can have a PAM adjacent at the 3’ end (e.g., in the case of Cas9). Alternatively, the guide RNA target sequence can have a PAM adjacent at the 5’ end (e.g., in the case of Cpf1). For example, the cleavage site of a Cas protein can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream (e.g., within the guide RNA target sequence) of the PAM sequence. In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5’-N1GG-3’, where N1 is any DNA nucleotide, and the PAM is immediately adjacent to the 3’ of the guide RNA target sequence on the non-complementary strand of the target DNA. Thus, the sequence corresponding to the PAM on the complementary strand (i.e., the reverse complement) is 5’-CCN2-3’, where N2 is any DNA nucleotide, and is immediately adjacent to the 5’ of the sequence where the DNA targeting segment of the guide RNA hybridizes to the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1 = C and N2 = G, N1 = G and N2 = C, N1 = A and N2 = T, or N1 = T and N2 = A). In the case of Cas9 from S. aureus, the PAM can be NNGRRT or NNGRR, where N can be A, G, C, or T, and R can be G or A. In the case of Cas9 from C. jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (e.g., in the case of FnCpf1), the PAM sequence is upstream of the 5’ and can have the sequence 5’-TTN-3’.
[0202] An example of a guide RNA target sequence is a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by the SpCas9 protein. For example, two examples of guide RNA target sequence + PAM are GN 19 NGG (SEQ ID NO: 46) or N 20 NGG (SEQ ID NO: 47). For example, see WO2014 / 165825, which is hereby incorporated by reference in its entirety for all purposes. The guanine at the 5' end can facilitate transcription by RNA polymerase in the cell. Other examples of guide RNA target sequence + PAM may include two guanine nucleotides at the 5' end to facilitate efficient transcription by T7 polymerase in vitro (e.g., GGN 20 NGG; SEQ ID NO: 48). For example, see WO2014 / 065596, which is hereby incorporated by reference in its entirety for all purposes. Other guide RNA target sequence + PAM may have a length of 4 to 22 nucleotides of SEQ ID NOs: 46-48, including 5' G or GG and 3' GG or NGG. Still other guide RNA target sequence PAM may have a length of 14 to 20 nucleotides of SEQ ID NOs: 46-48.
[0203] The formation of a CRISPR complex hybridized to a target DNA can result in cleavage of one or both strands of the target DNA within or near a region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and on the reverse complement of the complementary strand to which the guide RNA hybridizes). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a position defined relative to the PAM sequence). A "cleavage site" includes the position on the target DNA where the Cas protein generates a single-strand break or a double-strand break. The cleavage site can be on only one strand (e.g., when a nickase is used) or on both strands of double-stranded DNA. The cleavage site can be at the same position on both strands (generating blunt ends; e.g., Cas9), or at different sites on each strand (generating staggered ends (i.e., overhangs); e.g., Cpf1). Staggered ends can be generated, for example, by using two Cas proteins that each generate a single-strand break at a different cleavage site on a different strand, thereby generating a double-strand break. For example, a first nickase can create a single-strand break on the first strand of double-stranded DNA (dsDNA), and a second nickase can create a single-strand break on the second strand of dsDNA such that an overhang sequence is created. In some cases, the guide RNA target sequence or cleavage site of the first-strand nickase is separated from the guide RNA target sequence or cleavage site of the second-strand nickase by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.
[0204] (3) Exogenous donor nucleic acid targeting the human albumin gene The methods and compositions disclosed herein can utilize exogenous donor nucleic acids to modify the humanized albumin locus after cleavage of the humanized albumin locus by a nuclease agent or independently of cleavage of the humanized albumin locus by a nuclease agent. In such methods using a nuclease agent, the nuclease agent protein cleaves the humanized albumin locus to create a single-strand break (nick) or a double-strand break, and the exogenous donor nucleic acid rejoins the humanized albumin locus via non-homologous end joining (NHEJ)-mediated ligation or a homology-directed repair event. Optionally, repair by the exogenous donor nucleic acid removes or disrupts the nuclease target sequence, such that the targeted allele cannot be re-targeted by the nuclease agent.
[0205] The exogenous donor nucleic acid can target any sequence of the human albumin gene. Some exogenous donor nucleic acids contain homology arms. Other exogenous donor nucleic acids do not contain homology arms. The exogenous donor nucleic acids can be inserted into the humanized albumin locus by homologous recombination repair and / or they can be inserted into the humanized albumin locus by non-homologous end joining. In one example, the exogenous donor nucleic acid (e.g., a targeting vector) can target intron 1, intron 12, or intron 13 of the human albumin gene. For example, the exogenous donor nucleic acid can target intron 1 of the human albumin gene.
[0206] Exogenous donor nucleic acids can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which can be single-stranded or double-stranded, and which can be in linear or circular form. For example, exogenous donor nucleic acids can be single-stranded oligodeoxynucleotides (ssODNs). See, for example, Yoshimi et al. (2016) Nat. Commun. 7:10431, which is incorporated herein by reference in its entirety for all purposes. Exogenous donor nucleic acids may be naked nucleic acids or may be delivered by viruses such as AAV. In a specific example, exogenous donor nucleic acids can be delivered via AAV and inserted into the humanized albumin locus by non-homologous end joining (e.g., exogenous donor nucleic acids may not contain homology arms).
[0207] Exemplary exogenous donor nucleic acids are about 50 nucleotides to about 5 kb in length, or about 50 nucleotides to about 3 kb in length, or about 50 to about 1,000 nucleotides in length. Other exemplary exogenous donor nucleic acids are about 40 to about 200 nucleotides in length. For example, the exogenous donor nucleic acid can be about 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 110, 110 - 120, 120 - 130, 130 - 140, 140 - 150, 150 - 160, 160 - 170, 170 - 180, 180 - 190, or 190 - 200 nucleotides in length. Alternatively, the exogenous donor nucleic acid can be about 50 - 100, 100 - 200, 200 - 300, 300 - 400, 400 - 500, 500 - 600, 600 - 700, 700 - 800, 800 - 900, or 900 - 1000 nucleotides in length. Alternatively, the exogenous donor nucleic acid can be about 1 - 1.5, 1.5 - 2, 2 - 2.5, 2.5 - 3, 3 - 3.5, 3.5 - 4, 4 - 4.5, or 4.5 - 5 kb in length. Alternatively, the exogenous donor nucleic acid can be, for example, 5 kb, 4.5 kb, 4 kb, 3.5 kb, 3 kb, 2.5 kb, 2 kb, 1.5 kb, 1 kb, 900 nucleotides, 800 nucleotides, 700 nucleotides, 600 nucleotides, 500 nucleotides, 400 nucleotides, 300 nucleotides, 200 nucleotides, 100 nucleotides, or 50 nucleotides or less in length. The exogenous donor nucleic acid (e.g., the targeting vector) can also be longer.
[0208] In one example, the exogenous donor nucleic acid is an ssODN having a length of from about 80 nucleotides to about 200 nucleotides. In another example, the exogenous donor nucleic acid is an ssODN having a length of from about 80 nucleotides to about 3 kb. Such ssODNs may also have homology arms that are, for example, each from about 40 nucleotides to about 60 nucleotides in length. Such ssODNs may also have homology arms that are, for example, each from about 30 nucleotides to 100 nucleotides in length. The homology arms may be symmetric (e.g., each 40 nucleotides or each 60 nucleotides in length), or they may be asymmetric (e.g., one homology arm that is 36 nucleotides in length and one homology arm that is 91 nucleotides in length).
[0209] The exogenous donor nucleic acid can include modifications or sequences that provide additional desirable features (e.g., modified or regulated stability; tracking or detection by fluorescent labeling; binding sites for proteins or protein complexes, etc.). The exogenous donor nucleic acid may include one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, the exogenous donor nucleic acid may include one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes) such as at least 1, at least 2, at least 3, at least 4, or at least 5 fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and -6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide variety of fluorescent dyes for labeling oligonucleotides are commercially available (e.g., from Integrated DNA Technologies). Such fluorescent labels (e.g., internal fluorescent labels) can be used, for example, to detect exogenous donor nucleic acids directly incorporated into cleaved target nucleic acids having overhangs compatible with the ends of the exogenous donor nucleic acid. The label or tag may be at the 5' end, 3' end, or internal of the exogenous donor nucleic acid. For example, the exogenous donor nucleic acid can be conjugated to the 5' end having the IR700 fluorophore (5’IRDYE(登録商標) 700 ) from Integrated DNA Technologies.
[0210] The exogenous donor nucleic acid can also include a nucleic acid insert that contains a segment of DNA to be integrated into the humanized albumin locus. Integration of the nucleic acid insert at the humanized albumin locus can result in the addition of the nucleic acid sequence of interest to the humanized albumin locus, deletion of the nucleic acid sequence of interest at the humanized albumin locus, or replacement (i.e., deletion and insertion) of the nucleic acid sequence of interest at the humanized albumin locus. Some exogenous donor nucleic acids are designed for insertion of the nucleic acid insert at the humanized albumin locus without a corresponding deletion at the humanized albumin locus. Other exogenous donor nucleic acids are designed to delete the nucleic acid sequence of interest at the humanized albumin locus without insertion of any corresponding nucleic acid insert. Still other exogenous donor nucleic acids are designed to delete the nucleic acid sequence of interest at the humanized albumin locus and replace it with a nucleic acid insert.
[0211] The nucleic acid insert or corresponding nucleic acid at the humanized albumin locus to be deleted and / or replaced can be of various lengths. Exemplary nucleic acid inserts or corresponding nucleic acids at the humanized albumin locus to be deleted and / or replaced are of a length of about 1 nucleotide to about 5 kb, or of a length of about 1 nucleotide to about 1,000 nucleotides. For example, the nucleic acid insert or corresponding nucleic acid at the humanized albumin locus to be deleted and / or replaced can be of a length of about 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150, 150-160, 160-170, 170-180, 180-190, or 190-120 nucleotides. Similarly, the nucleic acid insert or corresponding nucleic acid at the humanized albumin locus to be deleted and / or replaced can be of a length of 1-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 nucleotides. Similarly, the nucleic acid insert or corresponding nucleic acid at the humanized albumin locus to be deleted and / or replaced can be of a length of about 1-1.5, 1.5-2, 2-2.5, 2.5-3, 3-3.5, 3.5-4, 4-4.5, or 4.5-5 kb, or longer.
[0212] The nucleic acid insert can include a sequence that is homologous or orthologous to all or part of the sequence to be replaced. For example, the nucleic acid insert can include a sequence that contains one or more point mutations (e.g., 1, 2, 3, 4, 5, or more) compared to the sequence targeted for replacement at the humanized albumin locus. Optionally, such point mutations can result in conservative amino acid substitutions (e.g., substitution of aspartic acid [Asp, D] with glutamic acid [Glu, E]) in the encoded polypeptide.
[0213] Some exogenous donor nucleic acids can encode exogenous proteins that are not encoded or expressed by the wild-type endogenous albumin locus (e.g., may include an inserted nucleic acid encoding an exogenous protein). In one example, a humanized albumin locus targeted by an exogenous donor nucleic acid can encode a heterologous protein that includes a human albumin signal peptide fused to a protein that is not encoded or expressed by the wild-type endogenous albumin locus. For example, the exogenous donor nucleic acid can be a promoterless cassette that includes a splice acceptor, and the exogenous donor nucleic acid can target the first intron of human albumin.
[0214] Donor nucleic acids for non-homologous end joining-mediated insertion. Some exogenous donor nucleic acids can be inserted into the humanized albumin locus by non-homologous end joining. In some cases, such exogenous donor nucleic acids do not contain homology arms. For example, such exogenous donor nucleic acids can be inserted into blunt-ended double-strand breaks after cleavage by a nuclease agent. In a specific example, the exogenous donor nucleic acid can be delivered via AAV and inserted into the humanized albumin locus by non-homologous end joining (e.g., the exogenous donor nucleic acid may not contain homology arms). In a specific example, the exogenous donor nucleic acid can be inserted via homology-independent targeted integration. For example, the insertion sequence of the exogenous donor nucleic acid inserted into the humanized albumin locus can be adjacent to the target sites of the nuclease agent on each side (e.g., the same target site as the humanized albumin locus and the same nuclease agent used to cleave the target site of the humanized albumin locus). The nuclease agent can then cleave the target sites adjacent to the insertion sequence. In a specific example, the exogenous donor nucleic acid is delivered by AAV-mediated delivery, and cleavage of the target sites adjacent to the insertion sequence can remove the inverted terminal repeats (ITRs) of AAV. In some methods, when the insertion sequence is inserted into the humanized albumin locus in the correct orientation, the target site of the humanized albumin locus (e.g., the gRNA target sequence containing the adjacent protospacer adjacent motif) is no longer present, but is re-formed when the insertion sequence is inserted into the humanized albumin locus in the opposite direction. This can help ensure that the insertion sequence is inserted in the correct orientation for expression.
[0215] Other exogenous donor nucleic acids have short single-stranded regions at the 5' end and / or 3' end, which are complementary to one or more overhangs created by nuclease-mediated cleavage at the humanized albumin locus. These overhangs are sometimes referred to as 5' and 3' homology arms. For example, some exogenous donor nucleic acids have short single-stranded regions at the 5' end and / or 3' end, which are complementary to one or more overhangs created by nuclease-mediated cleavage at the 5' and / or 3' target sequences of the humanized albumin locus. Some such exogenous donor nucleic acids have complementary regions only at the 5' end or only at the 3' end. For example, some such exogenous donor nucleic acids have complementary regions only at the 5' end complementary to the overhang created at the 5' target sequence of the humanized albumin locus, or only at the 3' end complementary to the overhang created at the 3' target sequence of the humanized albumin locus. Other such exogenous donor nucleic acids have complementary regions at both the 5' and 3' ends. For example, other such exogenous donor nucleic acids have complementary regions (e.g., complementary to the first and second overhangs, respectively) at both the 5' and 3' ends generated by nuclease-mediated cleavage at the humanized albumin locus. For example, when the exogenous donor nucleic acid is double-stranded, the single-stranded complementary regions can extend from the 5' end of the upper strand of the donor nucleic acid and the 5' end of the lower strand of the donor nucleic acid to create 5' overhangs at both ends. Alternatively, the single-stranded complementary regions can extend from the 3' end of the upper strand of the donor nucleic acid and from the 3' end of the lower strand of the template to create 3' overhangs.
[0216] The complementary region can be of any length sufficient to promote ligation between the exogenous donor nucleic acid and the target nucleic acid. Exemplary complementary regions are from about 1 to about 5 nucleotides in length, from about 1 to about 25 nucleotides in length, or from about 5 to about 150 nucleotides in length. For example, the complementary region can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. Alternatively, the complementary region can be from about 5 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, 90 to 100, 100 to 110, 110 to 120, 120 to 130, 130 to 140, or 140 to 150 nucleotides, or longer.
[0217] Such complementary regions can complement overhangs created by pairs of nickases. By using a first and second nickase that cleave opposite strands of DNA to create a first double-strand break, and a third and fourth nickase that cleave opposite strands of DNA to create a second double-strand break, two double-strand breaks with staggered ends can be created. For example, Cas proteins can be used to nick at first, second, third, and fourth guide RNA target sequences corresponding to the first, second, third, and fourth guide RNAs. The first and second guide RNA target sequences can be arranged to create a first cleavage site such that nicks created by the first and second nickases on the first and second strands of DNA create a double-strand break (i.e., the first cleavage site includes the nicks within the first and second guide RNA target sequences). Similarly, the third and fourth guide RNA target sequences can be arranged to create a second cleavage site such that nicks created by the third and fourth nickases on the first and second strands of DNA create a double-strand break (i.e., the second cleavage site includes the nicks within the third and fourth guide RNA target sequences). Preferably, the nicks within the first and second guide RNA target sequences and / or the third and fourth guide RNA target sequences can be offset nicks that create overhangs. The offset window can be, for example, at least about 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, or more. See Ran et al. (2013) Cell 154:1380-1389, Mali et al. (2013) Nat. Biotech. 31:833-838, and Shen et al. (2014) Nat. Methods 11:399-404, each of which is hereby incorporated by reference in its entirety for all purposes. In such cases, a double-stranded exogenous donor nucleic acid can be designed with single-stranded complementary regions that are complementary to the overhangs created by the nicks within the first and second guide RNA target sequences and the nicks within the third and fourth guide RNA target sequences.Subsequently, such exogenous donor nucleic acids can be inserted by non-homologous end joining-mediated ligation.
[0218] Donor nucleic acids for insertion by homology-directed repair. Some exogenous donor nucleic acids contain homology arms. When the exogenous donor nucleic acid also contains a nucleic acid insert, the homology arms can flank the nucleic acid insert. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This terminology relates to the relative position of the homology arms with respect to the nucleic acid insert within the exogenous donor nucleic acid. The 5' and 3' homology arms correspond to regions within the humanized albumin locus, which are referred to herein as the "5' target sequence" and "3' target sequence", respectively.
[0219] A homology arm and a target sequence are "corresponding" or "matched" to each other if the two regions share a level of sequence identity sufficient for them to act as substrates for homologous recombination reactions. The term "homology" includes DNA sequences that are identical to or share sequence identity with the corresponding sequences. The sequence identity between a given target sequence and the corresponding homology arm found in an exogenous donor nucleic acid can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of sequence identity shared by the homology arm of an exogenous donor nucleic acid (or a fragment thereof) and the target sequence (or a fragment thereof) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, whereby the sequences undergo homologous recombination. Further, the corresponding homologous region between the homology arm and the corresponding target sequence can be of any length sufficient to promote homologous recombination. Exemplary homology arms are about 25 nucleotides to about 2.5 kb in length, about 25 nucleotides to about 1.5 kb in length, or about 25 to about 500 nucleotides in length. For example, a given homology arm (or each homology arm) and / or the corresponding target sequence can include a corresponding homologous region that is about 25 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 nucleotides in length, such that the homology arm has sufficient homology to undergo homologous recombination with the corresponding target sequence within the target nucleic acid. Alternatively, a given homology arm (or each homology arm) and / or the corresponding target sequence can include a corresponding homologous region that is about 0.5 kb to about 1 kb, about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, or about 2 kb to about 2.5 kb in length. For example, the homology arms can each be about 750 nucleotides in length.The homology arms may be symmetric (each of approximately the same size) or asymmetric (one longer than the other).
[0220] When using a nuclease agent in combination with an exogenous donor nucleic acid, the 5' and 3' target sequences are preferably placed in close proximity to the nuclease cleavage site (e.g., in close proximity to the nuclease target sequence) so as to promote the occurrence of a homologous recombination event between the target sequence and the homology arms upon single-strand cleavage (nick) or double-strand cleavage at the nuclease cleavage site. The term "nuclease cleavage site" includes the DNA sequence at which a nick or double-strand cleavage is created by a nuclease agent (e.g., a Cas9 protein complexed with a guide RNA). Target sequences within the targeted locus corresponding to the 5' and 3' homology arms of the exogenous donor nucleic acid are "positioned in close proximity" to the nuclease cleavage site if the distance is such that it promotes the occurrence of a homologous recombination event between the 5' and 3' target sequences and the homology arms upon single-strand or double-strand cleavage at the nuclease cleavage site. Thus, the target sequences corresponding to the 5' and / or 3' homology arms of the exogenous donor nucleic acid can be, for example, within at least 1 nucleotide of a given nuclease cleavage site, or within at least 10 nucleotides to about 1,000 nucleotides of a given nuclease cleavage site. As an example, the nuclease cleavage site can be directly adjacent to at least one or both of the target sequences.
[0221] The spatial relationship between the homology arms of the exogenous donor nucleic acid and the target sequences corresponding to the nuclease cleavage site can vary. For example, the target sequence may be located 5' to the nuclease cleavage site, the target sequence may be located 3' to the nuclease cleavage site, or the target sequence may be adjacent to the nuclease cleavage site.
[0222] (4) Other human albumin targeting reagents The activity of any other known or putative human albumin targeting reagent can also be evaluated using the non-human animals disclosed herein. Similarly, any other molecule can be screened for human albumin targeting activity using the non-human animals disclosed herein.
[0223] Examples of other human albumin targeting reagents include antisense oligonucleotides (e.g., siRNA or shRNA) that act via RNA interference (RNAi). Antisense oligonucleotides (ASO) or antisense RNA are short synthetic strings of nucleotides designed to prevent the expression of a target protein by selectively binding to the RNA encoding the target protein and thereby preventing translation. These compounds bind to RNA with high affinity and selectivity via well-characterized Watson-Crick base pairing (hybridization). RNA interference (RNAi) is an endogenous cellular mechanism for controlling gene expression, in which small interfering RNAs (siRNAs) that bind to the RNA-induced silencing complex (RISC) mediate the cleavage of target messenger RNA (mRNA).
[0224] Other human albumin targeting reagents include antibodies or antigen-binding proteins designed to specifically bind to human albumin epitopes. Other human albumin targeting reagents include small molecule reagents.
[0225] D. Administration of Human Albumin Targeting Reagents to Non-Human Animals or Cells The methods disclosed herein can include introducing various molecules (e.g., human albumin targeting reagents such as therapeutic molecules or complexes), including for example nucleic acids, proteins, nucleic acid-protein complexes, or protein complexes, into non-human animals or cells. "Introducing" includes presenting a molecule (e.g., a nucleic acid or protein) to a cell or non-human animal in such a way as to gain access to the interior of the cell or the interior of the non-human animal cell. Introducing can be accomplished by any means, and two or more components (e.g., two components, or all components) can be introduced into the cell or non-human animal simultaneously or sequentially in any combination. For example, a Cas protein can be introduced into a cell or non-human animal before or after introduction of the guide RNA, or vice versa. As another example, an exogenous donor nucleic acid can be introduced before or after introduction of the Cas protein and guide RNA (e.g., the exogenous donor nucleic acid can be administered about 1, 2, 3, 4, 8, 12, 24, 36, 48, or 72 hours before or after introduction of the Cas protein and guide RNA). See, e.g., US 2015 / 0240263 and US 2015 / 0110762, each of which is incorporated herein by reference in its entirety for all purposes. Further, two or more components can be introduced into the cell or non-human animal by the same delivery method or different delivery methods. Similarly, two or more components can be introduced into the non-human animal by the same administration route or different administration routes.
[0226] In some methods, the components of the CRISPR / Cas system are introduced into non-human animals or cells. The guide RNA can be introduced into non-human animals or cells in the form of RNA (e.g., in vitro transcribed RNA) or in the form of DNA encoding the guide RNA. When introduced in the form of DNA, the DNA encoding the guide RNA can be operably linked to an active promoter within the cells of the non-human animal. For example, the guide RNA can be delivered via AAV and expressed in vivo under the U6 promoter. Such DNA can be present in one or more expression constructs. For example, such an expression construct can be a component of a single nucleic acid molecule. Alternatively, they can be separated in any combination among two or more nucleic acid molecules (i.e., the DNA encoding one or more CRISPR RNAs and the DNA encoding one or more tracrRNAs can be components of separate nucleic acid molecules).
[0227] Similarly, the Cas protein can be provided in any form. For example, the Cas protein can be provided in the form of a protein such as a Cas protein complexed with the gRNA. Alternatively, the Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon optimized for efficient translation to the protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to alternative codons that are used more frequently in mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest compared to the naturally occurring polynucleotide sequence. When the nucleic acid encoding the Cas protein is introduced into a non-human animal, the Cas protein can be expressed intracellularly transiently, conditionally, or constitutively.
[0228] The nucleic acid encoding the Cas protein or the guide RNA can be operably linked to a promoter in an expression construct. The expression construct includes any nucleic acid construct that can direct the expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and transfer such a nucleic acid sequence of interest into a target cell. For example, the nucleic acid encoding the Cas protein may be present within a vector containing DNA encoding one or more gRNAs. Alternatively, it may be a separate vector or plasmid from the vector containing DNA encoding one or more gRNAs. Suitable promoters that can be used in the expression construct include, for example, promoters that are active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, progenitor cells with restricted development, induced pluripotent stem (iPS) cells, or one-cell stage embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Optionally, the promoter may be a bidirectional promoter that drives the expression of both the Cas protein in one direction and the guide RNA in the other direction. Such a bidirectional promoter may be composed of (1) a complete conventional unidirectional Pol III promoter containing three external control elements, a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box, and (2) a second basic Pol III promoter containing the PSE and TATA box fused in the reverse direction to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and the TATA box, and the promoter can be made bidirectional by creating a hybrid promoter in which the transcription in the reverse direction is controlled by adding the PSE and TATA box derived from the U6 promoter. See, for example, US2016 / 0074535, which is hereby incorporated by reference in its entirety for all purposes.By using a bidirectional promoter to simultaneously express the genes encoding the Cas protein and the guide RNA, a compact expression cassette can be generated to facilitate delivery.
[0229] Molecules introduced into non-human animals or cells (e.g., Cas protein or guide RNA) can be provided in a composition comprising a carrier that increases the stability of the introduced molecule (e.g., extending the period during which degradation products remain below a threshold, e.g., less than 0.5 wt% of the starting nucleic acid or protein, under given storage conditions (e.g., -20 °C, 4 °C, or ambient temperature), or increasing stability in vivo). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, and lipid nanotubes.
[0230] To enable the introduction of molecules (e.g., nucleic acids or proteins) into cells or non-human animals, various methods and compositions are provided herein. Methods for introducing molecules into various cell types are known and include, for example, stable transfection methods, transient transfection methods, and virus-mediated methods.
[0231] The transfection protocol, as well as the protocol for introducing molecules into cells, can be different. Non-limiting transfection methods include chemical-based transfection methods using liposomes, nanoparticles, calcium phosphate (Graham et al. (1973) Virology 52(2):456-67, Bacchetti et al. (1977) Proc. Natl. Acad. Sci. USA 74(4):1590-4, and Kriegler, M (1991). Transfer and Expression: A Laboratory Manual. New York: W.H. Freeman and Company. pp.96-97), dendrimers, or cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation, sonoporation, and optical transfection. Particle-based transfection includes the use of a gene gun or magnet-assisted transfection (Bertram (2006) Current Pharmaceutical Biotechnology 7, 277-28). Viral methods can also be used for transfection.
[0232] The introduction of molecules (e.g., nucleic acids or proteins) into cells can also be mediated by electroporation, cytoplasmic injection, viral infection, adenovirus, adeno-associated virus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or nucleofection. Nucleofection is an improved electroporation technology that enables the delivery of nucleic acid substrates to the nucleus through not only the cytoplasm but also the nuclear membrane. Furthermore, the use of nucleofection in the methods disclosed herein typically requires far fewer cells than conventional electroporation (e.g., only about 2 million compared to 7 million by conventional electroporation). In one example, nucleofection is performed using the LONZA® NUCLEOFECTOR™ system.
[0233] The introduction of molecules (e.g., nucleic acids or proteins) into cells (e.g., zygotes) can also be achieved by microinjection. In zygotes (i.e., one-cell stage embryos), microinjection can be performed into the maternal and / or paternal pronuclei or cytoplasm. When microinjection is performed into only one pronucleus, the paternal pronucleus is more suitable because of its larger size. Microinjection of mRNA is preferably into the cytoplasm (e.g., to directly deliver the mRNA to the translation machinery), while microinjection of Cas protein or a polynucleotide encoding Cas protein or encoding RNA is preferably into the nucleus / pronucleus. Alternatively, microinjection can be performed by injection into both the nucleus / pronucleus and the cytoplasm. The needle can be first introduced into the nucleus / pronucleus, the first volume can be injected, and while removing the needle from the one-cell stage embryo, the second volume can be injected into the cytoplasm. When the Cas protein is injected into the cytoplasm, the Cas protein preferably contains a nuclear localization signal to ensure delivery to the nucleus / pronucleus. Methods for performing microinjection are well known. See, for example, Nagy et al. (Nagy A, Gertsenstein M, Vintersten K, Behringer R., 2003, Manipulating the Mouse Embryo. Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press), and also see Meyer et al. (2010) Proc. Natl. Acad. Sci. USA 107:15022-15026, and Meyer et al. (2012) Proc. Natl. Acad. Sci. USA 109:9354-9359.
[0234] Other methods for introducing a molecule (e.g., a nucleic acid or a protein) into a cell or a non-human animal can include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. As specific examples, a nucleic acid or a protein can be introduced into a cell or a non-human animal in a carrier such as poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid vortices, or lipid nanotubes. Some specific examples of delivery to non-human animals include hydrodynamic delivery, virus-mediated delivery (e.g., adeno-associated virus (AAV)-mediated delivery), and lipid nanoparticle-mediated delivery.
[0235] The introduction of nucleic acids and proteins into cells or non-human animals can be achieved by hydrodynamic delivery (HDD). In gene delivery to parenchymal cells, it is necessary to inject from a blood vessel that has selected only the essential DNA sequences, eliminating the safety concerns associated with current viral and synthetic vectors. When DNA is injected into the bloodstream, it can reach the cells of various tissues that have access to the blood. Hydrodynamic delivery utilizes the force generated by the rapid injection of a large volume of solution into the non-compressible blood in circulation to overcome the physical barriers of the endothelium and cell membrane that prevent large and membrane-impermeable compounds from entering parenchymal cells. In addition to the delivery of DNA, this method helps to efficiently deliver RNA, proteins, and other small compounds intracellularly in vivo. See, for example, Bonamassa et al. (2011) Pharm. Res. 28(4):694-701, which is hereby incorporated by reference in its entirety for all purposes.
[0236] Introduction of nucleic acids can also be achieved by viral-mediated delivery, such as AAV-mediated delivery or lentiviral-mediated delivery. Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia virus, poxvirus, and herpes simplex virus. Viruses can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. The virus may or may not integrate into the host genome. Such viruses can also be engineered to have a reduced immune response. The virus may be replication-competent or replication-deficient (e.g., defective in one or more genes required for virion replication and / or additional rounds of packaging). The virus can cause transient expression, long-term expression (e.g., for at least 1 week, 2 weeks, 1 month, 2 months, or 3 months), or persistent expression (e.g., of Cas9 and / or gRNA). Exemplary viral titers (e.g., AAV titers) are 10 12 、10 13 、10 14 、10 15 、および10 16 vector genomes / mL.
[0237] The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two terminal inverted repeats that allow synthesis of complementary DNA strands. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap are supplied in trans. In addition to Rep and Cap, AAV may require a helper plasmid containing genes derived from adenovirus. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenovirus gene E1+ to generate infectious AAV particles. Alternatively, Rep, Cap, and adenovirus helper genes can be combined into a single plasmid. Similar packaging cells and methods can also be used for other viruses such as retroviruses.
[0238] Multiple serotypes of AAV have been identified. These serotypes differ in the types of cells they infect (i.e., their tropism) and enable preferential transduction of specific cell types. Serotypes for CNS tissue include AAV1, AAV2, AAV4, AAV5, AAV8, and AAV9. Serotypes for heart tissue include AAV1, AAV8, and AAV9. The serotype for kidney tissue includes AAV2. Serotypes for lung tissue include AAV4, AAV5, AAV6, and AAV9. The serotype for pancreatic tissue includes AAV8. Serotypes for photoreceptor cells include AAV2, AAV5, and AAV8. Serotypes for retinal pigment epithelial tissue include AAV1, AAV2, AAV4, AAV5, and AAV8. Serotypes for skeletal muscle tissue include AAV1, AAV6, AAV7, AAV8, and AAV9. Serotypes for liver tissue include AAV7, AAV8, and AAV9, particularly AAV8.
[0239] Directivity can be further refined by pseudotyping, which is the mixing of genomes from different viral serotypes than the capsid. For example, AAV2 / 5 refers to a virus containing the genome of serotype 2 packaged in the capsid of serotype 5. The use of pseudotyped viruses can not only improve transduction efficiency but also change directivity. Hybrid capsids derived from different serotypes can also be used to modify the directivity of the virus. For example, AAV-DJ contains hybrid capsids derived from eight serotypes and exhibits high infectivity across a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the characteristics of AAV-DJ but with enhanced uptake into the brain. AAV serotypes can also be modified by mutation. Examples of mutated modifications of AAV2 include Y444F, Y500F, Y730F, and S662V. Examples of mutated modifications of AAV3 include Y705F, Y731F, and T492V. Examples of mutated modifications of AAV6 include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SASTG.
[0240] To accelerate the expression of the transgene, self-complementary AAV (scAAV) variants can be used. AAV depends on the cell's DNA replication machinery to synthesize the complementary strand of the single-stranded DNA genome of AAV, so the expression of the transgene may be delayed. To address this delay, scAAV containing complementary sequences that can spontaneously anneal upon infection can be used, eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0241] To increase packaging capacity, longer transgenes can be split between two AAV transfer plasmids, the first being the 3' splice donor and the second being the 5' splice acceptor. Upon co-infection of cells, these viruses can form concatemers and be spliced together to express the full-length transgene. This allows for the expression of longer transgenes, but the efficiency of expression is reduced. A similar method for increasing capacity utilizes homologous recombination. For example, a transgene can be split between two transfer plasmids, but there is substantial sequence overlap such that co-expression induces homologous recombination and expression of the full-length transgene.
[0242] The introduction of nucleic acids and proteins can also be achieved by lipid nanoparticle (LNP)-mediated delivery. For example, LNP-mediated delivery can be used to deliver a combination of Cas mRNA and guide RNA, or a combination of Cas protein and guide RNA. Delivery by such methods results in transient expression of Cas, and the biodegradable lipids improve clearance, improve tolerability, and reduce immunogenicity. Lipid formulations can protect biomolecules from degradation while improving cellular uptake. Lipid nanoparticles are particles that contain multiple lipid molecules physically bound to each other by intermolecular forces. These include microspheres (including monolayer and multilayer vesicles such as liposomes), the dispersed phase in emulsions, micelles, or the internal phase in suspensions. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids are useful for delivering polyanions such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the time the nanoparticles can exist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO2016 / 010840A1, which is hereby incorporated by reference in its entirety for all purposes. Exemplary lipid nanoparticles can contain a cationic lipid and one or more other components. In one example, the other components can include a helper lipid such as cholesterol. In another example, the other components can include a helper lipid such as cholesterol and a neutral lipid such as DSPC. In another example, the other components can include a helper lipid such as cholesterol, any neutral lipid such as DSPC, and a stealth lipid such as S010, S024, S027, S031, or S033.
[0243] The LNP may include one or more or all of the following: (i) lipids for encapsulation and endosomal escape, (ii) neutral lipids for stabilization, (iii) helper lipids for stabilization, and (iv) stealth lipids. See, for example, Finn et al. (2018) Cell Reports 22:1-9 and WO2017 / 173054A1, each of which is hereby incorporated by reference in its entirety for all purposes. In certain LNPs, the cargo can include a guide RNA or a nucleic acid encoding a guide RNA. In certain LNPs, the cargo can include an mRNA encoding a Cas nuclease such as Cas9, and a guide RNA or a nucleic acid encoding a guide RNA.
[0244] The lipids for encapsulation and endosomal escape can be cationic lipids. The lipids can also be biodegradable lipids such as biodegradable ionizable lipids. An example of a suitable lipid is lipid A or LP01, which is also known as (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-(((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, or 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-(((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate. See, for example, Finn et al. (2018) Cell Reports 22:1-9 and WO2017 / 173054A1, each of which is hereby incorporated by reference in its entirety for all purposes. Another example of a suitable lipid is lipid B, which is also known as ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate), or ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate). Another example of a suitable lipid is lipid C, which is 2-((4-(((3-(dimethylamino)propoxy)carbonyl)oxy)hexadecanoyl)oxy)propane-1,3-diyl(9Z,9’Z,12Z,12’Z)-bis(octadeca-9,12-dienoate). Another example of a suitable lipid is lipid D, which is 3-(((3-(dimethylamino)propoxy)carbonyl)oxy)-13-(octanoyloxy)tridecyl 3-octylundecanoate. Other suitable lipids include heptatriaconta-6,9,28,31-tetraene-19-yl 4-(dimethylamino)butanoate (also known as Dlin-MC3-DMA (MC3))).
[0245] Some such lipids suitable for use in the LNPs described herein are biodegradable in vivo. For example, LNPs containing such lipids include those in which at least 75% of the lipid is removed from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. As another example, at least 50% of the LNP is removed from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days.
[0246] Such lipids may be ionizable depending on the pH of the medium in which they are contained. For example, in a slightly acidic medium, the lipid can be protonated and thus carry a positive charge. Conversely, in a slightly basic medium such as blood, where the pH is about 7.35, the lipid may not be protonated and thus may not carry a charge. In some embodiments, the lipid can be protonated at a pH of at least about 9, 9.5, or 10. The ability of such lipids to carry a charge is related to their intrinsic pKa. For example, the lipids can independently have a pKa in the range of about 5.8 to about 6.2.
[0247] Neutral lipids function to stabilize the LNPs and improve their processing. Examples of suitable neutral lipids include various neutral, uncharged, or zwitterionic lipids. Examples of neutral phospholipids suitable for use in the present disclosure are 5-heptadecylbenzene-1,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine (DSPC), phosphatidylcholine (DOPC), dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauroylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DMPC), 1-myristoyl-2-palmitoylphosphatidylcholine (MPPC), 1-palmitoyl-2-myristoylphosphatidylcholine (PMPC), 1-palmitoyl-2-stearoylphosphatidylcholine (PSPC), 1,2-diarachidonoyl-sn-glycero-3-phosphocholine (DBPC), 1-stearoyl-2-palmitoylphosphatidylcholine (SPPC), 1,2-dieicosenoyl-sn-glycero-3-phosphocholine (DEPC), palmitoyloleoylphosphatidylcholine (POPC), lysophosphatidylcholine, dioleoylphosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine distearoylphosphatidylethanolamine (DSPE), dimyristoylphosphatidylethanolamine (DMPE), dipalmitoylphosphatidylethanolamine (DPPE), palmitoylrolenoylphosphatidylethanolamine (POPE), lysophosphatidylethanolamine and combinations thereof, including but not limited to these. For example, the neutral phospholipid can be selected from the group consisting of distearoylphosphatidylcholine (DSPC), and dimyristoylphosphatidylethanolamine (DMPE).
[0248] Helper lipids include lipids that promote transfection. The mechanism by which helper lipids enhance transfection may include enhancing particle stability. In some cases, helper lipids can increase membrane fusibility. Helper lipids include steroids, sterols, and alkylresorcinols. Examples of suitable helper lipids include cholesterol, 5-heptadecylresorcinol, and cholesteryl hemisuccinate. In one example, the helper lipid can be cholesterol or cholesteryl hemisuccinate.
[0249] Stealth lipids include lipids that alter the length of time nanoparticles can exist in vivo. Stealth lipids are useful in the formulation process, for example, by reducing particle aggregation and controlling particle size. Stealth lipids can modulate the pharmacokinetic properties of LNPs. Suitable stealth lipids include lipids having a hydrophilic head group linked to a lipid moiety.
[0250] The hydrophilic head group of the stealth lipid may include a polymer moiety selected from polymers based on, for example, PEG (often called poly(ethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N-vinylpyrrolidone), polyamino acids, and poly N-(2-hydroxypropyl)methacrylamide. The term PEG means polyethylene glycol or other polyalkylene ether polymers. In certain LNP formulations, PEG is PEG-2K, also called PEG2000, with an average molecular weight of about 2,000 daltons. See, for example, WO2017 / 173054A1, which is incorporated herein by reference in its entirety for all purposes.
[0251] The lipid portion of the stealth lipid can be derived from, for example, a dialkylglycerol or a dialkylglycamide group having an alkyl chain length of about C4 to about C40 saturated or unsaturated carbon atoms independently, where the chain may contain one or more functional groups such as amide or ester. The dialkylglycerol or dialkylglycamide group may further contain one or more substituted alkyl groups.
[0252] As an example, the stealth lipid may be selected from PEG-dilauroylglycerol, PEG-dimyristoyl glycerol (PEG-DMG), PEG-dipalmitoyl glycerol, PEG-distearoyl glycerol (PEG-DSPE), PEG-dilauryl glycamide, PEG-dimyristyl glycamide, PEG-dipalmitoyl glycamide, and PEG-distearoyl glycamide, PEG-cholesterol (I-[8’-(cholest-5-en-3[beta]-oxy) carboxamide-3’,6’-dioxaoctanyl] carbamoyl-[omega]-methyl-poly(ethylene glycol)), PEG-DMB (3,4-ditetradecyloxybenzyl-[omega]-methyl-poly(ethylene glycol) ether), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DMG), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSPE), 1,2-distearoyl-sn-glycerol, methoxypolyethylene glycol (PEG2k-DSG), poly(ethylene glycol)-2000-dimethacrylate (PEG2k-DMA)), and 1,2-distearyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSA). In one specific example, the stealth lipid may be PEG2k-DMG.
[0253] The LNP can contain different respective molar ratios of the component lipids in the formulation. The molar % of the CCD lipid can be, for example, about 30 mol% to about 60 mol%, about 35 mol% to about 55 mol%, about 40 mol% to about 50 mol%, about 42 mol% to about 47 mol%, or about 45%. The molar % of the helper lipid can be, for example, about 30 mol% to about 60 mol%, about 35 mol% to about 55 mol%, about 40 mol% to about 50 mol%, about 41 mol% to about 46 mol%, or about 44%. The molar % of the neutral lipid can be, for example, about 1 mol% to about 20 mol%, about 5 mol% to about 15 mol%, about 7 mol% to about 12 mol%, or about 9 mol%. The molar % of the stealth lipid can be, for example, about 1 mol% to about 10 mol%, about 1 mol% to about 5 mol%, about 1 mol% to about 3 mol%, about 2 mol%, or about 1 mol%.
[0254] The LNP can have different ratios between the positively charged amine group of the biodegradable lipid (N) and the negatively charged phosphate group (P) of the encapsulated nucleic acid. This can be mathematically represented by the equation N / P. For example, the N / P ratio can be about 0.5 to about 100, about 1 to about 50, about 1 to about 25, about 1 to about 10, about 1 to about 7, about 3 to about 5, about 4 to about 5, about 4, about 4.5, or about 5. The N / P ratio can also be about 4 to about 7 or about 4.5 to about 6. In certain examples, the N / P ratio can be 4.5 or 6.
[0255] In some LNPs, the cargo can include Cas mRNA and gRNA. The Cas mRNA and gRNA can be in different ratios. For example, the LNP formulation can include a ratio of Cas mRNA to gRNA in the range of about 25:1 to about 1:25, in the range of about 10:1 to about 1:10, in the range of about 5:1 to about 1:5, or a ratio of about 1:1. Alternatively, the LNP formulation can include a ratio of Cas mRNA to gRNA nucleic acids in the range of about 1:1 to about 1:5, or about 10:1. Alternatively, the LNP formulation can include a ratio of Cas mRNA to gRNA nucleic acids of about 1:10, 25:1, 10:1, 5:1, 3:1, 1:1, 1:3, 1:5, 1:10, or 1:25. Alternatively, the LNP formulation can include a ratio of Cas mRNA to gRNA nucleic acids in the range of about 1:1 to about 1:2. In a specific example, the ratio of Cas mRNA to gRNA can be about 1:1 or about 1:2.
[0256] In some LNPs, the cargo can include exogenous donor nucleic acid and gRNA. The exogenous donor nucleic acid and gRNA can be in different ratios. For example, the LNP formulation can include a ratio of exogenous donor nucleic acid to gRNA in the range of about 25:1 to about 1:25, in the range of about 10:1 to about 1:10, in the range of about 5:1 to about 1:5, or a ratio of about 1:1. Alternatively, the LNP formulation can include a ratio of exogenous donor nucleic acid to gRNA in the range of about 1:1 to about 1:5, in the range of about 5:1 to about 1:1, about 10:1, or about 1:10. Alternatively, the LNP formulation can include a ratio of exogenous donor nucleic acid to gRNA nucleic acids of about 1:10, 25:1, 10:1, 5:1, 3:1, 1:1, 1:1:3, 1:5, 1:10, or 1:25.
[0257] Specific examples of suitable LNPs include those having a nitrogen to phosphate (N / P) ratio of 4.5 and containing a biodegradable cationic lipid, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 45:44:9:2. The biodegradable cationic lipid can be (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl (9Z,12Z)-octadeca-9,12-dienoate. See, for example, Finn et al. (2018) Cell Reports 22:1-9, which is incorporated herein by reference in its entirety for all purposes. The Cas9 mRNA can be in a weight ratio of 1:1 with respect to the guide RNA. Another specific example of a suitable LNP contains Dlin-MC3-DMA (MC3), cholesterol, DSPC, and PEG-DMG in a molar ratio of 50:38.5:10:1.5.
[0258] Another specific example of a suitable LNP has a nitrogen to phosphate (N / P) ratio of 6 and contains a biodegradable cationic lipid, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50:38:9:3. The biodegradable cationic lipid can be (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl (9Z,12Z)-octadeca-9,12-dienoate. The Cas9 mRNA can be in a weight ratio of 1:2 with respect to the guide RNA.
[0259] To reduce immunogenicity, the delivery mode can be selected. For example, the Cas protein and gRNA can be delivered by different modes (e.g., bi-modal delivery). These different modes may confer different pharmacokinetic or pharmacodynamic properties on the delivered molecules in the subject (e.g., Cas or nucleic acid encoding gRNA or nucleic acid encoding nucleic acid, or exogenous donor nucleic acid / repair template). For example, different modes may result in different tissue distributions, different half-lives, or different temporal distributions. Some delivery modes (e.g., delivery of nucleic acid vectors that persist intracellularly by self-replication or genomic integration) result in more persistent expression and presence of the molecule, while other delivery modes are transient and less persistent (e.g., delivery of RNA or protein). For example, delivery of the Cas protein in a more transient manner as mRNA or protein ensures that the Cas / gRNA complex is present and active only for a short period, reducing the immunogenicity caused by peptides from bacterial-derived Cas enzymes displayed on the cell surface by MHC molecules. Such transient delivery can also reduce the likelihood of off-target modifications.
[0260] In vivo administration can be by any suitable route, including, for example, parenteral, intravenous, oral, subcutaneous, intra-arterial, intracranial, intrathecal, intraperitoneal, topical, intranasal, or intramuscular. Systemic administration modes include, for example, oral and parenteral routes. Examples of parenteral routes include intravenous, intra-arterial, intraosseous, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. A specific example is intravenous injection. Intranasal injection and intravitreal injection are other specific examples. Topical administration modes include, for example, intrathecal, intraventricular, intracerebral (e.g., local intracerebral delivery to the striatum (e.g., to the caudate nucleus or putamen), cerebral cortex, precentral gyrus, hippocampus (e.g., dentate gyrus or CA3 region), temporal cortex, tonsil, prefrontal cortex, thalamus, cerebellum, medulla, thalamus, tectum, putamen, or substantia nigra), intraocular, intraorbital, subconjunctival, intravitreal, subretinal, and transscleral routes. A significantly smaller amount of the component can exert an effect when administered locally (e.g., intracerebral or intravitreal) compared to when administered systemically (e.g., intravenously). The topical administration mode can also reduce or eliminate the incidence of potentially toxic side effects that can occur when a therapeutically effective amount of the component is administered systemically.
[0261] In vivo administration can be by any suitable route, including, for example, parenteral, intravenous, oral, subcutaneous, intra-arterial, intracranial, intrathecal, intraperitoneal, topical, intranasal, or intramuscular. A specific example is intravenous injection. Compositions containing guide RNA and / or Cas protein (or nucleic acids encoding guide RNA and / or Cas protein) can be formulated using one or more physiologically and pharmaceutically acceptable carriers, diluents, excipients, or adjuvants. The formulation may depend on the selected route of administration. The term "pharmaceutically acceptable" means that the carrier, diluent, excipient, or adjuvant is compatible with the other components of the formulation and is not substantially harmful to its recipient.
[0262] The frequency and number of administrations can depend, among other factors, on the half-life and route of administration of the exogenous donor nucleic acid, guide RNA, or Cas protein (or nucleic acid encoding the guide RNA or Cas protein). The introduction of nucleic acid or protein into cells or non-human animals can be carried out once or multiple times over a period of time. For example, the introduction can be carried out at least 2 times over a period of time, at least 3 times over a period of time, at least 4 times over a period of time, at least 5 times over a period of time, at least 6 times over a period of time, at least 7 times over a period of time, at least 8 times over a period of time, at least 9 times over a period of time, at least 10 times over a period of time, at least 11 times over a period of time, at least 12 times over a period of time, at least 13 times over a period of time, at least 14 times over a period of time, at least 15 times over a period of time, at least 16 times over a period of time, at least 17 times over a period of time, at least 18 times over a period of time, at least 19 times over a period of time, or at least 20 times over a period of time.
[0263] E. Measurement of delivery, activity, or efficacy of a human albumin-targeting reagent in vivo or ex vivo The methods disclosed herein may further include detecting or measuring the activity of a human albumin-targeting reagent. For example, when the human albumin-targeting reagent is a genome editing reagent (e.g., CRISPR / Cas designed to target the human albumin locus), the measurement may include evaluating the humanized albumin locus for modification.
[0264] A variety of methods can be used to identify cells having a targeted genetic modification. Screening can include quantitative assays to evaluate the modification of alleles (MOA) of parental chromosomes. For example, see US2004 / 0018626, US2014 / 0178879, US2016 / 0145646, WO2016 / 081923, and Frendewey et al. (2010) Methods Enzymol. 476:295-307, each of which is hereby incorporated by reference in its entirety for all purposes. For example, the quantitative assay can be performed via quantitative PCR such as real-time PCR (qPCR). Real-time PCR can utilize a first primer set that recognizes the target locus and a second primer set that recognizes a non-target reference locus. The primer sets can include a fluorescent probe that recognizes the amplified sequence. Other examples of suitable quantitative assays include fluorescence in situ hybridization (FISH), comparative genomic hybridization, isothermal DNA amplification, quantitative hybridization to all immobilized probes, INVADER® probes, TAQMAN® molecular beacon probes, or Eclipse™ probe technology (see, e.g., US2005 / 0144655, which is hereby incorporated by reference in its entirety for all purposes).
[0265] Next-generation sequencing (NGS) can also be used for screening. Next-generation sequencing is sometimes referred to as "NGS" or "ultra-parallel sequencing" or "high-throughput sequencing". NGS can be used as a screening tool in addition to the MOA assay to define the exact nature of the targeted genetic modification and whether it is consistent between cell types, tissue types, or organ types.
[0266] Evaluating the modification of the humanized albumin locus in non-human animals can be in any cell type from any tissue or organ. For example, the evaluation can be performed on multiple cell types from the same tissue or organ, or on cells from multiple locations within a tissue or organ. This can provide information about which cell types within the target tissue or organ are targeted, or which sections of the tissue or organ are reached by the human albumin targeting reagent. As another example, the evaluation can be performed on multiple types of tissues or multiple organs. In the way a particular tissue, organ, or cell type is targeted, this can provide information about how effectively that tissue or organ is targeted and whether there are off-target effects on other tissues or organs.
[0267] If the reagent is designed to inactivate the humanized albumin locus, affect the expression of the humanized albumin locus, prevent the translation of humanized albumin mRNA, or remove the humanized albumin protein, the measurement can include an assessment of humanized albumin mRNA or protein expression. This measurement can be performed within the liver or within a particular cell type or region within the liver, or can include measuring the serum levels of the secreted humanized albumin protein.
[0268] When the reagent is an exogenous donor nucleic acid encoding an exogenous protein that is not encoded or expressed by the wild-type endogenous albumin locus, the measurement may include assessing the expression of mRNA encoded by the exogenous donor nucleic acid or assessing the expression of the exogenous protein. This measurement can be performed in the liver or within a specific cell type or region within the liver, or may include measuring the serum level of the secreted exogenous protein. In a specific example, the exogenous protein is factor IX protein. Optionally, the assessment includes measuring the serum level of factor IX protein in a non-human animal and / or assessing the activated partial thromboplastin time or performing a thrombin generation assay. Optionally, the non-human animal further comprises an inactivated F9 locus, and the assessment includes measuring the serum level of factor IX protein in the non-human animal and / or assessing the activated partial thromboplastin time (aPTT) or performing a thrombin generation assay (TGA). These assays are described in more detail in the examples.
[0269] One example of an assay that can be used is the BASESCOPE™ RNA in situ hybridization (ISH) assay, which, in the context of intact fixed tissue, can quantify cell-specifically edited transcripts that contain single nucleotide changes. The BASESCOPE™ RNA ISH assay can complement NGS and qPCR in the characterization of gene editing. While NGS / qPCR can provide quantitative averages of wild-type and edited sequences, they do not provide information regarding the heterogeneity or percentage of edited cells within a tissue. The BASESCOPE™ ISH assay can provide a landscape view of the entire tissue and quantification of wild-type versus edited transcripts at single-cell resolution that can quantify the actual number of cells within the target tissue that contain the edited mRNA transcript. The BASESCOPE™ assay uses paired oligo (“ZZ”) probes to amplify signals without non-specific background, enabling single molecule RNA detection. However, the design of the BASESCOPE™ probe and signal amplification system allows for the detection of single molecule RNA using the ZZ probe, enabling differential detection of single nucleotide editing and mutations in intact fixed tissue.
[0270] The production and secretion of humanized albumin protein or exogenous protein can be evaluated by any known means. For example, expression can be evaluated by measuring the level of encoded mRNA in the liver of a non-human animal or the level of encoded protein in the liver of a non-human animal using known assays. The secretion of humanized albumin protein or exogenous protein can be evaluated by measuring the level of encoded humanized albumin protein or exogenous protein or plasma level or serum level in a non-human animal using known assays.
[0271] IV. Method for Producing a Non-Human Animal Containing a Humanized Albumin Locus As disclosed elsewhere in this specification, various methods are provided for generating non-human animal genomes, non-human animal cells, or non-human animals that include the humanized albumin (ALB) locus. Any convenient method or protocol for generating a genetically modified organism is suitable for generating such genetically modified non-human animals. See, for example, Cho et al. (2009) Current Protocols in Cell Biology 42:19.11:19.11.1-19.11.22 and Gama Sosa et al. (2010) Brain Struct. Funct. 214(2-3):91-109, each of which is incorporated herein by reference in its entirety for all purposes. Such genetically modified non-human animals can be generated, for example, via gene knock-in at the targeted albumin locus.
[0272] For example, a method for producing a non-human animal that includes the humanized albumin locus can include: (1) modifying the genome of a pluripotent cell to include the humanized albumin locus; (2) identifying or selecting the genetically modified pluripotent cell that includes the humanized albumin locus; (3) introducing the genetically modified pluripotent cell into a non-human animal host embryo; and (4) implanting and gestating the host embryo in a surrogate mother. For example, a method for producing a non-human animal that includes the humanized albumin locus can include: (1) modifying the genome of a pluripotent cell to include the humanized albumin locus; (2) identifying or selecting the genetically modified pluripotent cell that includes the humanized albumin locus; (3) introducing the genetically modified pluripotent cell into a non-human animal host embryo; and (4) gestating the host embryo in a surrogate mother. Optionally, the host embryo that includes the modified pluripotent cell (e.g., non-human ES cells) can be incubated to the blastocyst stage prior to implanting and gestating in a surrogate mother to produce an F0 non-human animal. The surrogate mother can then produce a non-human animal of the F0 generation that includes the humanized albumin locus.
[0273] This method may further include identifying a cell or animal having a modified target genomic locus. Various methods can be used to identify cells having a targeted genetic modification.
[0274] The step of modifying the genome can, for example, utilize an exogenous donor nucleic acid (e.g., a targeting vector) to modify the albumin locus to include the humanized albumin locus disclosed herein. As an example, the targeting vector can be for generating a humanized albumin gene at an endogenous albumin locus (e.g., an endogenous non-human animal albumin locus), and the targeting vector includes a 5' homology arm that targets a 5' target sequence of the endogenous albumin locus and a 3' homology arm that targets a 3' target sequence of the endogenous albumin locus. The exogenous donor nucleic acid can also include a nucleic acid insert that includes a segment of DNA to be integrated into the albumin locus. Integration of the nucleic acid insert at the albumin locus can result in the addition of a nucleic acid sequence of interest at the albumin locus, the deletion of a nucleic acid sequence of interest at the albumin locus, or the substitution (i.e., deletion and insertion) of a nucleic acid sequence of interest at the albumin locus. The homology arms can flank the inserted nucleic acid containing the human albumin sequence to generate a humanized albumin locus (e.g., to delete a segment of the endogenous albumin locus and replace it with an orthologous human albumin sequence).
[0275] The exogenous donor nucleic acid can be for insertion via non-homologous end joining or for homologous recombination. The exogenous donor nucleic acid can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which can be single-stranded or double-stranded, and which can be in linear or circular form. For example, the repair template can be a single-stranded oligodeoxynucleotide (ssODN).
[0276] Heterologous donor nucleic acids can also contain heterologous sequences that are not present at the non-targeted endogenous albumin locus. For example, an exogenous donor nucleic acid can contain a selection cassette such as a selection cassette adjacent to a recombinase recognition site.
[0277] Some exogenous donor nucleic acids contain homology arms. When the exogenous donor nucleic acid also contains a nucleic acid insert, the homology arms can be adjacent to the nucleic acid insert. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This term is related to the relative position of the homology arms with respect to the nucleic acid insert within the exogenous donor nucleic acid. The 5' and 3' homology arms correspond to regions within the albumin locus, which are referred to herein as the "5' target sequence" and the "3' target sequence", respectively.
[0278] The homology arms and the target sequences "correspond" or "are corresponding" to each other when the two regions share a level of sequence identity sufficient for the two regions to act as substrates for a homologous recombination reaction. The term "homology" includes DNA sequences that are identical to or share sequence identity with the corresponding sequences. The sequence identity between a given target sequence and the corresponding homology arm found in the exogenous donor nucleic acid can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of sequence identity shared by the homology arm of the exogenous donor nucleic acid (or a fragment thereof) and the target sequence (or a fragment thereof) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, whereby the sequences undergo homologous recombination. Further, the corresponding homology regions between the homology arms and the corresponding target sequences can be of any length sufficient to promote homologous recombination. In some expression vectors, the intended mutation of the endogenous albumin locus is contained in the inserted nucleic acid adjacent to the homology arms.
[0279] In cells other than one-cell stage embryos, the exogenous donor nucleic acid may be a "targeting vector" or "LTVEC", which includes a targeting vector corresponding to a nucleic acid sequence larger than those typically used in other approaches aimed at performing homologous recombination within the cell and containing homology arms derived therefrom. LTVEC also includes a targeting vector containing a nucleic acid insert having a nucleic acid sequence larger than those normally used in other approaches aimed at performing homologous recombination within the cell. For example, LTVEC enables alteration of large loci that cannot be accommodated by conventional plasmid-based targeting vectors due to size limitations. For example, the targeted locus may be a locus of a cell that cannot be targeted using conventional methods or can only be inaccurately or with significantly low efficiency targeted in the absence of a nick or double-strand break induced by a nuclease agent (e.g., a Cas protein) (i.e., the 5' and 3' homology arms may correspond). LTVEC can be of any length, usually at least 10 kb in length. The total of the 5' and 3' homology arms of LTVEC is typically at least 10 kb.
[0280] The screening step can include, for example, a quantitative assay for evaluating modification of alleles (MOA) of the parental chromosome. For example, the quantitative assay can be performed via quantitative PCR such as real-time PCR (qPCR). Real-time PCR can utilize a first primer set that recognizes the target locus and a second primer set that recognizes a non-target reference locus. The primer set can include a fluorescent probe that recognizes the amplified sequence.
[0281] Other examples of suitable quantitative assays include fluorescence-mediated in situ hybridization (FISH), comparative genomic hybridization, isothermal DNA amplification, quantitative hybridization to immobilized probes, INVADER® probes, TAQMAN® molecular beacon probes, or ECLIPSE™ probe technology (see, e.g., US2005 / 0144655, which is incorporated herein by reference in its entirety for all purposes).
[0282] Examples of suitable pluripotent cells are embryonic stem (ES) cells (e.g., mouse ES cells or rat ES cells). Modified pluripotent cells can be generated by recombination, for example, by (a) introducing into the cell one or more exogenous donor nucleic acids (e.g., a targeting vector) comprising an inserted nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites, respectively, wherein the inserted nucleic acid comprises a human albumin sequence for generating a humanized albumin locus, and (b) identifying at least one cell comprising the inserted nucleic acid integrated into the endogenous albumin locus in its genome (i.e., identifying at least one cell comprising a humanized albumin locus). Modified pluripotent cells can be generated by recombination, for example, by (a) introducing into the cell one or more targeting vectors comprising an inserted nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites, respectively, wherein the inserted nucleic acid comprises a humanized albumin locus, and (b) identifying at least one cell comprising the inserted nucleic acid integrated into the target genomic locus in its genome.
[0283] Alternatively, the modified pluripotent cells can be generated by: (a) introducing into the cells (i) a nuclease agent that induces a nick or double-strand break at a target site within the endogenous albumin locus, and (ii) optionally, one or more exogenous donor nucleic acids (e.g., a targeting vector) comprising an insertion nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites located in sufficient proximity to the nuclease target site, wherein the insertion nucleic acid comprises a human albumin sequence for generating a humanized albumin locus; and (c) identifying at least one cell in its genome that contains the insertion nucleic acid integrated into the endogenous albumin locus (i.e., identifying at least one cell containing a humanized albumin locus). Alternatively, the modified pluripotent cells can be generated by: (a) introducing into the cells (i) a nuclease agent or a nucleic acid encoding a nuclease agent that induces a nick or double-strand break at a target site within the endogenous albumin locus, and (ii) optionally, one or more exogenous donor nucleic acids (e.g., a targeting vector) comprising an insertion nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites located in sufficient proximity to the nuclease target site, wherein the insertion nucleic acid comprises a human albumin sequence for generating a humanized albumin locus; and (c) identifying at least one cell in its genome that contains the insertion nucleic acid integrated into the endogenous albumin locus (i.e., identifying at least one cell containing a humanized albumin locus).Alternatively, the modified pluripotent cells can be generated by: (a) introducing into the cells (i) a nuclease agent that induces a nick or double-strand break at a recognition site within the target genomic locus, and (ii) one or more targeting vectors comprising an inserted nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites located in sufficient proximity to the recognition site, wherein the inserted nucleic acid comprises a humanized albumin locus; and (c) identifying at least one cell comprising a modification (e.g., integration of the inserted nucleic acid) at the target genomic locus. Any nuclease agent that induces a nick or double-strand break at a desired recognition site can be used. Alternatively, the modified pluripotent cells can be generated by: (a) introducing into the cells (i) a nuclease agent or a nucleic acid encoding a nuclease agent that induces a nick or double-strand break at a recognition site within the target genomic locus, and (ii) one or more targeting vectors comprising an inserted nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites located in sufficient proximity to the recognition site, wherein the inserted nucleic acid comprises a humanized albumin locus; and (c) identifying at least one cell comprising a modification (e.g., integration of the inserted nucleic acid) at the target genomic locus. Any nuclease agent that induces a nick or double-strand break at a desired recognition site can be used. Examples of suitable nucleases include transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), meganucleases, and clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems (e.g., the CRISPR / Cas9 system) or components of such systems (e.g., CRISPR / Cas9). See, for example, US 2013 / 0309670 and US 2015 / 0159175, each of which is incorporated herein by reference in its entirety for all purposes.
[0284] Donor cells can be introduced into the host embryo at any stage, such as the blastocyst stage or the pre-morula stage (i.e., the 4-cell stage or the 8-cell stage). Progeny capable of transmitting the genetic modification through the germ cell line are produced. See, for example, U.S. Patent No. 7,294,754, which is hereby incorporated by reference in its entirety for all purposes.
[0285] Alternatively, a method of making a non-human animal described elsewhere herein can include (1) modifying the genome of a one-cell stage embryo to include a humanized albumin locus using the methods described above for modifying pluripotent cells, (2) selecting the genetically modified embryo, and (3) implanting and gestating the genetically modified embryo in a surrogate mother. Alternatively, a method of making a non-human animal described elsewhere herein can include (1) modifying the genome of a one-cell stage embryo to include a humanized albumin locus using the methods described above for modifying pluripotent cells, (2) selecting the genetically modified embryo, and (3) gestating the genetically modified embryo in a surrogate mother. Progeny capable of transmitting the genetic modification through the germ cell line are produced.
[0286] Nuclear transfer technology can also be used to generate non-human mammals. Briefly, the method of nuclear transfer can include the steps of: (1) enucleating an oocyte or providing an enucleated oocyte; (2) isolating or providing a donor cell or nucleus to combine with the enucleated oocyte; (3) inserting the cell or nucleus into the enucleated oocyte to form a reconstructed cell; (4) implanting the reconstructed cell into the uterus of an animal to form an embryo; and (5) developing the embryo. In such a method, oocytes are generally recovered from dead animals, but can also be isolated from the oviducts and / or ovaries of living animals. Oocytes can be matured in various well-known media prior to enucleation. Enucleation of oocytes can be performed by many well-known methods. Insertion of a donor cell or nucleus into an enucleated oocyte to form a reconstructed cell can be performed by microinjecting the donor cell under the zona pellucida prior to fusion. Fusion can be induced by application of a DC electrical pulse across the contact / fusion plane (electrical fusion), exposure of the cells to a fusion-promoting chemical such as polyethylene glycol, or an inactivated virus such as Sendai virus. The reconstructed cell can be activated by electrical and / or non-electrical means before, during, and / or after fusion of the nuclear donor and recipient oocyte. Activation methods include electrical pulses, chemically induced shocks, penetration by sperm, an increase in the level of divalent cations in the oocyte, and a reduction in phosphorylation of cellular proteins in the oocyte (by a kinase inhibitor). The activated reconstructed cell, or embryo, can be cultured in a well-known medium and then transferred to the uterus of an animal. See, for example, US2008 / 0092249, WO1999 / 005266, US2004 / 0177390, WO2008 / 017234, and U.S. Patent No. 7,612,250, each of which is hereby incorporated by reference in its entirety for all purposes.
[0287] The various methods provided herein enable the generation of genetically modified non-human F0 animals, and the cells of the genetically modified F0 animals contain a humanized albumin locus. It is recognized that the number of cells in the F0 animals having a humanized albumin locus varies depending on the method used to generate the F0 animals. For example, via the VELOCIMOUSE® method, by introducing donor ES cells into pre-morula stage embryos (e.g., 8-cell stage mouse embryos) from the corresponding organism, a greater proportion of the cell population of the F0 animal can be made to contain cells having the nucleotide sequence of interest that includes a targeted genetic modification. For example, at least 50%, 60%, 65%, 70%, 75%, 85%, 86%, 87%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cell contribution of the non-human F0 animal can include a cell population having a targeted modification.
[0288] The cells of the genetically modified F0 animals can be heterozygous with respect to the humanized albumin locus or can be homozygous with respect to the humanized albumin locus.
[0289] All patent applications, websites, other publications, accession numbers, etc., cited above or below are hereby incorporated by reference in their entirety for all purposes to the same extent as if each item were specifically and individually indicated to be incorporated by reference. If different versions of a sequence are associated with an accession number at different times, it means the version associated with the accession number on the effective filing date of the present application. The effective filing date, where applicable, means the actual filing date for the accession number or the filing date of the priority application prior thereto. Similarly, if different versions of a publication, website, etc. are published at different times, unless otherwise specified, it means the version published most recently on the effective filing date of the present application. Unless otherwise indicated, any feature, step, element, embodiment, or aspect of the present invention can be used in combination with any other. Although the present invention has been described in some detail through diagrams and examples for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be made within the scope of the appended claims.
[0290] Brief Description of the Sequences The nucleotide and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases and the three-letter notations for amino acids. The nucleotide sequences follow the standard convention starting at the 5' end of the sequence and proceeding forward (i.e., left to right in each row) to the 3' end. Only one strand of each nucleotide sequence is shown, but the complementary strand is understood to be included by reference to the strand shown. When a nucleotide sequence encoding an amino acid sequence is provided, it is understood that its codon degeneracy variants encoding the same amino acid sequence are also provided. The amino acid sequences follow the standard convention starting at the amino terminus of the sequence and proceeding forward (i.e., left to right in each row) to the carboxy terminus. [Table 2-1] [Table 2-2]
Example
[0291] Example 1. Generation of a mouse containing a humanized albumin (ALB) locus A large targeting vector (LTVEC) containing a 5' homology arm containing a 20 kb mouse albumin (Alb) locus (derived from bMQ-127G8) and a 3' homology arm containing a 127 kb mouse albumin (Alb) locus (derived from bMQ-127G8) was generated, and a 14.4 kb (14,376 bp) region of the mouse albumin (Alb) gene was replaced with a 17.3 kb (17,335 bp) (derived from RP11-31P12) of the corresponding human albumin sequence (ALB). Information regarding mouse and human albumin is shown in Table 3. VELOCIGENE (登録商標) The generation and use of large targeting vectors (LTVECs) derived from bacterial artificial chromosome (BAC) DNA by bacterial homologous recombination (BHR) using genetic engineering techniques are described, for example, in US 6,586,251 and Valenzuela et al. (2003) Nat. Biotechnol. 21(6):652-659, each of which is hereby incorporated by reference in its entirety for all purposes. The generation of LTVEC by in vitro assembly methods is referred to, for example, in US 2015 / 0376628 and WO2015 / 200334, each of which is hereby incorporated by reference in its entirety for all purposes.
Table 3
[0292] Specifically, the region from the ATG start codon to the stop codon (i.e., coding exons 1-14) was deleted from the mouse albumin (Alb) locus. Instead of the deleted mouse region, the corresponding region of human albumin (ALB) from the ATG start codon to 100 bp downstream of the stop codon was inserted. The loxP-mPrm1-Crei-pA-hUb1-em7-Neo-pA-loxP cassette (4,766 bp) was inserted downstream of the human 3’UTR along with a buffer of approximately 100 bp of 3’ human sequence after the 3’UTR immediately preceding the cassette. This is the MAID7626 allele. See Figure 1A. After deletion of the cassette, the loxP and cloning sites (38 bp) remained downstream of the human 3’UTR along with a buffer of approximately 100 bp of 3’ human sequence after the 3’UTR immediately preceding the remaining loxP site. This is the MAID7627 allele. See Figure 1B.
[0293] The sequences of the mouse albumin signal peptide, propeptide, and serum albumin are shown in SEQ ID NOs: 2-4, respectively, and the corresponding coding sequences are shown in SEQ ID NOs: 10-12, respectively. The sequences of the human albumin signal peptide, propeptide, and serum albumin are shown in SEQ ID NOs: 6-8, respectively, and the corresponding coding sequences are shown in SEQ ID NOs: 14-16, respectively. The predicted encoded humanized albumin protein is identical to the human albumin protein. See Figures 1A and 1B. An alignment of the mouse and human albumin proteins along with the humanized albumin protein is shown in Figures 3A-3B. The coding sequences of the mouse and human Alb / ALB are shown in SEQ ID NOs: 9 and 13, respectively. The albumin protein sequences of the mouse and human are shown in SEQ ID NOs: 1 and 5, respectively. The predicted humanized ALB coding sequence and the sequence of the predicted humanized albumin protein are shown in SEQ ID NOs: 13 and 5, respectively.
[0294] To generate mutant alleles, the above large targeting vector was introduced into F1H4 mouse embryonic stem cells. The F1H4 mouse ES cells were derived from hybrid embryos produced by mating female C57BL / 6NTac mice with male 12956 / SvEvTac mice. See, for example, US2015 / -0376651 and WO2015 / 200805, each of which is incorporated herein by reference in its entirety for all purposes. After antibiotic selection, colonies were picked, expanded, and screened by TAQMAN (登録商標) See Figure 2. Allele loss assays were performed using the primers and probes listed in Table 4 to detect loss of the endogenous mouse allele, and allele gain assays were performed to detect acquisition of the humanized allele.
Table 4
[0295] Allele modification (MOA) assays, including allele loss (LOA) and allele gain (GOA) assays, are described, for example, in US2014 / 0178879, US2016 / 0145646, WO2016 / 081923, and Frendewey et al. (2010) Methods Enzymol. 476:295-307, each of which is incorporated herein by reference in its entirety for all purposes. The allele loss (LOA) assay reverses the logic of conventional screening and quantifies the copy number of genomic DNA samples at the native locus targeted by the mutation. In clones of correctly targeted heterozygous cells, the LOA assay detects one of the two native alleles (for genes not on the X or Y chromosome), and the other allele is disrupted by the targeted modification. The same principle as the allele gain (GOA) assay can be applied in reverse to quantify the copy number of the inserted targeting vector in genomic DNA samples.
[0296] F0 mice were generated using VELOCIMOUSE (登録商標)Generated from ES cells modified using the method. Specifically, a clone of mouse ES cells containing the above-described humanized albumin locus selected by the above-described MOA assay was injected into an 8-cell stage embryo using the VELOCIMOUSE (登録商標) method. For example, see US7,576,259, US7,659,442, US7,294,754; US2008 / 0078000, and Poueymirou et al. (2007) Nat. Biotechnol. 25(1):91-99, each of which is hereby incorporated by reference in its entirety for all purposes. VELOCIMOUSE (登録商標) In the method, targeted mouse embryonic stem (ES) cells are injected, for example, via laser-assisted injection into a pre-morula stage embryo, such as an 8-cell stage embryo, that efficiently gives rise to F0 generation mice that are entirely ES cell-derived. VELOCIMOUSE (登録商標) In the method, the injected pre-morula stage embryo is cultured until the blastocyst stage, the blastocyst stage embryo is introduced into a surrogate mother and allowed to gestate to produce F0 generation mice. Starting with a clone of mouse ES cells that is homozygous for the target modification, F0 mice that are homozygous for the target modification are produced. Starting with a clone of mouse ES cells that is heterozygous for the target modification, subsequent breeding can be performed to produce mice that are homozygous for the target modification.
[0297] Example 2. Verification of mice containing the humanized albumin (ALB) locus To verify the humanized albumin mice, mouse and human serum albumin ELISA kits (Abcam ab179887 and ab207620, respectively) were used to measure the albumin levels of mouse and human in plasma samples. The humanized mice used for verification were F1 mice in which the self-deleting selection cassette had self-deleted. Human albumin protein was detected in normal human plasma and humanized albumin mouse plasma samples, but not in wild-type (WT) mouse or VelocImmune (VI) mouse plasma samples. See Figure 4. Mouse albumin protein was detected in wild-type mouse plasma samples and VI mouse plasma samples, but not in humanized albumin mouse plasma samples. See Figure 5. In particular, pooled normal human plasma (purchased from George King-Biomedical Inc.) contained approximately 30 - 40 mg / mL of human albumin. Humanized albumin mouse plasma contained approximately 10 - 15 mg / mL of human albumin, but mouse albumin was not detected. Normal VI and WT mouse plasma contained approximately 7 - 13 mg / mL of mouse albumin.
[0298] Example 3. Verification of Mice Containing Guide RNAs Targeting Human Albumin for Humanized Albumin (ALB) Locus - F9 Insertion To further verify the humanized albumin mice, the humanized albumin mice were used to evaluate the use of CRISPR / Cas9 technology to integrate the F9 transgene into the albumin locus. Specifically, the integration and expression of the integrated human F9 Padua variant (hF9-R338L) in homozygous humanized albumin mice were tested. Various guide RNAs were designed against intron 1 of the human albumin locus. ALB hu / huTwo separate mouse experiments were set up using mice, and a total of 11 guide RNAs targeting the first intron of the human albumin locus were screened. On day 0 of the experiment, the body weights of all mice were measured and they were injected via the tail vein. Blood was collected by tail bleeding at weeks 1, 3, 4, and 6, and plasma was separated. The mice were sacrificed at week 7. Blood was collected from the vena cava and plasma was separated. The liver and spleen were also dissected. The guide sequences (DNA targeting segments) of these guide RNAs are shown in Table 5. [Table 5]
[0299] In the first experiment, LNPs separately containing Cas9 mRNA and each of the following six guide RNAs: G009852, G009859, G009860, G009864, G009874, and G012764 were tested. The LNPs were diluted to 0.3 mg / kg (using an average weight of 30 grams) and co-injected with AAV8 packaged with a bidirectional hF9 insertion template (SEQ ID NO: 63; ITR - splice acceptor - hF9 (exon 2 - 8) - bGH - SV40 polyA - codon - optimized hF9 - pLac - pMB - splice acceptor - Kan resistance) at a dose of 3E11 viral genomes per mouse. Five ALB hu / hu male mice aged 12 - 14 weeks per group were injected. Five mice from the same cohort were injected with AAV8 packaged with a CAGG promoter (SEQ ID NO: 64; CAGG - ITR - hF9 - WPRE - bGH - ITR - pLac - pMB - Amp resistance) operably linked to hF9 that results in episomal expression of hF9 (at 3E11 viral genomes per mouse). There were three negative control groups, each containing three mice injected with buffer only, AAV8 packaged with the bidirectional hF9 insertion template only, or LNP - G009874 only.
[0300] In the second experiment, LNPs containing Cas9 mRNA and each of the following six guide RNAs: G009860, G012764, G009844, G009857, G012752, G012753, and G012761 were tested. The LNPs were diluted to 0.3 mg / kg (using an average weight of 40 grams) and co-injected with AAV8 packaging the bidirectional hF9 insertion template (SEQ ID NO: 63) at a dose of 3E11 viral genomes per mouse. Five ALB hu / hu male mice at 30 weeks of age per group were injected. Five mice from the same cohort were injected with AAV8 packaging the CAGG promoter (SEQ ID NO: 64) operably linked to hF9, which results in episomal expression of hF9 (at 3E11 viral genomes per mouse). There were three negative control groups, each containing three mice injected with buffer only, AAV8 packaging the bidirectional hF9 insertion template only, or LNP-G009874 only.
[0301] For analysis, ELISA was performed to measure the hFIX circulating levels in the mice at each time point. For this purpose, a human factor IX ELISA kit (ab188393) was used, and all plates were run with human pooled normal plasma from George King Bio-Medical as a positive assay control. The expression levels of human factor IX in the plasma samples of each group at 6 weeks post-injection are shown in Figures 6A and 6B. Consistent with the in vitro insertion data, when guide RNA G009852 was used, factor IX serum levels were not detected above background. Consistent with the lack of an adjacent PAM sequence in human albumin, when guide RNA G009864 was used, factor IX serum levels were not detected. The guide sequence (DNA targeting segment) of G009864 is UACUUUGCACUUUCCUUAGU (SEQ ID NO: 61), targeting cyno genomic coordinates (mf5) chr5: 61199187-61199207. Expression of factor IX in serum was observed with some of the other guide RNAs, including G009857, G009859, G009860, G009874, and G0012764.
[0302] The spleen and a part of all the left lobes of the liver were submitted for next-generation sequencing (NGS) analysis. Using NGS, the percentage of hepatocytes with insertions / deletions (indels) at the human albumin locus was evaluated 7 weeks after injection of AAV-hF9 donor and LNP-CRISPR / Cas9. Consistent with the lack of adjacent PAM sequences in human albumin, no editing was detected in the liver when using guide RNA G009864. Editing in the liver was observed in the groups using guide RNAs G009859, G009860, G009874, and G012764 (data not shown).
[0303] The remaining liver was fixed in 10% neutral buffered formalin for 24 hours and transferred to 70% ethanol. Four to five samples from separate lobes were excised, sent to HistoWisz, processed, and embedded in paraffin blocks. Next, 5-micron sections were cut from each paraffin block, and Universal BASESCOPE™ procedures and reagents by Advanced Cell Diagnostics, and ALB hu / hu BASESCOPE™ was performed on a Ventana Ultra Discovery (Roche) using a custom-designed probe targeting the unique mRNA junction formed between the human albumin signal sequence and the hF9 transgene from the first intron of the albumin locus. Next, the percentage of positive cells in each sample was quantified using HALO imaging software (Indica Labs). Next, the average percentage of positive cells across multiple lobes of each animal was correlated with the hFIX levels in the serum at 7 weeks. The results are shown in Figure 7 and Table 6. The serum levels at 7 weeks and the % positive cells of hALB-hFIX mRNA were strongly correlated (r = 0.89; R 2 = 0.79).
Table 6
[0304] Example 4. Verification of Mice Containing a Humanized Albumin (ALB) Locus-F9 Insertion in F9 KO Mice To further verify the humanized albumin mice, the humanized albumin mice were mated with F9 knockout mice to obtain ALB m / hu xF9 - / - mice (heterozygous and homozygous F9 knockout for humanization of the albumin locus), and the use of CRISPR / Cas9 technology to integrate the F9 transgene into the albumin locus was evaluated.
[0305] Next, humanized albumin F9 KO mice were used to test the insertion of the human F9 Padua variant (hF9-R338L) transgene into intron 1 of the humanized albumin locus. On day 0 of the experiment, the body weights of all mice were measured and they were injected via the tail vein. Blood was collected by tail bleeding at weeks 1 and 3, and plasma was separated. The mice were sacrificed at week 4. Blood was collected from the vena cava and plasma was separated. The liver and spleen were also dissected.
[0306] LNPs separately containing Cas9 mRNA and the following two guide RNAs: G009860 (targeting the first intron of the human albumin locus) and G000666 (targeting the first intron of the mouse albumin locus) were tested. The guide sequence (DNA targeting segment) of G009860 is shown in Table 5. The guide sequence of G000666 is CACUCUUGUCUGUGGAAACA (SEQ ID NO: 62), targeting the mouse genomic coordinates (mm10) chr5:90461709-90461729. G009860 was diluted to 0.3 mg / kg and G000666 was diluted to 1.0 mg / kg (using an average weight of 31.2 grams), and both were co-injected with AAV8 packaging the bidirectional hF9 insertion template (SEQ ID NO: 63) at a dose of 3E11 viral genomes per mouse. Five ALB per group ms / hu xF9 - / -Injected into male mice (16 weeks old). Five mice from the same cohort were injected with AAV8 packaged with the CAGG promoter (SEQ ID NO: 64) operably linked to hF9, which results in episomal expression of hF9 (at 3E11 viral genomes per mouse). One mouse per group was injected with AAV8 packaged with buffer only or with the bidirectional hF9 insertion template only, and there were six negative control animals, two mice per group injected with 0.3 mg / kg and 1.0 mg / kg of LNP-G009860 or LNP-G000666 only.
[0307] For analysis, ELISA was performed to measure the hFIX circulating levels in the mice at each time point. For this purpose, a human factor IX ELISA kit (ab188393) was used, and all plates were run with human pooled normal plasma from George King Bio-Medical as a positive assay control. A portion of the spleen and a part of the left lobe of all livers were submitted for NGS analysis.
[0308] The expression levels of human factor IX in plasma samples of each group at 1, 2, and 4 weeks post-injection are shown in Figure 8 and Table 7. Additionally, the NGS results showing the insertion and deletion (indel) levels at the albumin locus in the liver and spleen are shown in Table 7. As shown in Figure 8 and Table 7, hFIX was detected in the plasma of Alb + / hu / F9 - / - mice, and ELISA showed expression values of 0.5 - 10 μg / mL at 1, 3, and 4 weeks.
Table 7
[0309] The remaining liver was fixed in 10% neutral buffered formalin for 24 hours and transferred to 70% ethanol. Four to five samples from separate leaves were excised, sent to HistoWiz, processed, and embedded in paraffin blocks. Next, 5-micron sections were cut from each paraffin block for analysis via BASESCOPE™ using custom-designed probes targeting unique mRNA junctions formed between either the human or mouse albumin signal sequence from the first intron of each respective albumin locus in mice and the hF9 transgene, with the Universal BASESCOPE™ procedure and reagents by Advanced Cell Diagnostics, and Ventana Ultra Discovery (Roche). The percentage of positive cells in each sample was quantified using HALO imaging software (Indica Labs). ms / hu Using custom-designed probes targeting unique mRNA junctions formed between either the human or mouse albumin signal sequence from the first intron of each respective albumin locus in mice and the hF9 transgene, 5-micron sections were cut from each paraffin block for analysis via BASESCOPE™ with Ventana Ultra Discovery (Roche). The percentage of positive cells in each sample was quantified using HALO imaging software (Indica Labs).
[0310] Next, terminal blood was used to evaluate functional coagulation activity by activated partial thromboplastin time (aPTT) and thrombin generation assay (TGA). Activated partial thromboplastin time (aPTT) is a clinical measurement of the intrinsic pathway coagulation activity in plasma. The addition of ellagic acid or kaolin induces plasma to coagulate, both of which activate factor XII in the intrinsic pathway of coagulation (known as the contact pathway), and then thrombin is activated and fibrin is generated from fibrinogen. The aPTT assay provides an estimate of an individual's ability to form a blood clot, and this information can be used to determine the risk of bleeding or thrombosis. To test aPTT, a semi-automatic benchtop system (Diagnostica Stago STart 4) equipped with an electromechanical blood clot detection method (viscosity-based detection system) was used to evaluate blood clots in plasma. 50 μL of citrated plasma was added to each cuvette containing steel balls, incubated at 37 °C for 5 minutes, and then 50 μL of ellagic acid (final concentration 30 μM) was added at 37 °C for 300 seconds to initiate coagulation. 50 μL of 0.025 M calcium chloride (final concentration 8 mM) was added to each cuvette to finally activate coagulation, and then the steel balls began to vibrate back and forth between two drive coils. The movement of the balls was detected by the receiving coil. The generation of fibrin increased the plasma viscosity until the balls stopped moving, and this was recorded as the clotting time. The only parameter measured was the clotting time. The runs were performed in duplicate.
[0311] The Thrombin Generation Assay (TGA) is a non-clinical evaluation of the dynamics of thrombin generation in activated plasma. Thrombin is involved in the activation of other coagulation factors and the propagation of further thrombin (via FXI activation) for the conversion of fibrinogen to fibrin, so thrombin generation is essential for the coagulation process. The thrombin generation assay provides an estimate of an individual's ability to generate thrombin, and this information can be used to determine the risk of bleeding or thrombosis. To perform the TGA, a calibrated automated thrombinogram was used to evaluate thrombin generation levels with a spectrophotometer (Thrombinograph™, Thermo Scientific). For high-throughput experiments, 96-well plates (Immulon II HB) were used. 55 μL of citrated plasma (4-fold diluted with saline for mouse plasma) was added to each well and incubated at 37 °C for 30 minutes. Thrombin generation was initiated by adding 15 μL of 2 μM ellagic acid (final concentration 0.33 μM) at 37 °C for 45 minutes. Thrombin generation was determined after automatically injecting 15 μL of a fluorogenic substrate containing 16 mM CaCl2 (FluCa; Thrombinoscope BV) into each well. The fluorogenic substrate reacts with the generated thrombin, which was continuously measured at 460 nm every 33 seconds for 90 minutes in plasma. Fluorescence intensity was proportional to the proteolytic activity of thrombin. The main parameters measured in the tracing were lag time, peak thrombin generation, time to peak thrombin generation, and endogenous thrombin potential (ETP). Lag time provides an estimate of the time required for the initial detection of thrombin in plasma. The peak is the maximum amount of thrombin generated at a given time after activation. The time to peak thrombin generation is the time from the start of the coagulation cascade to the peak of thrombin generation. ETP is the total amount of thrombin generated during the 60 minutes measured. Runs were performed in duplicate.
[0312] As shown in Figure 9 and Table 8, the insertion of the hF9 transgene using either mouse albumin gRNA or human albumin gRNA showed restoration of coagulation function in the aPTT assay. Negative control samples of saline, AAV only, and LNP only showed an extended aPTT time of 45 - 60 seconds. The positive control CAGG and the test sample (AAV8 + LNP) were close to normal human aPTT of 28 - 34 seconds.
Table 8
[0313] As shown in Figures 10A, 10B, and 11 and Table 8, the insertion of the hF9 transgene using either mouse albumin gRNA or human albumin gRNA showed an increase in thrombin generation in the TGA - EA analysis. The thrombin concentration was higher in the positive control CAGG and AAV8 + LNP compared to the negative control samples.
[0314] In conclusion, hFIX was detected in the plasma of Alb + / hu / F9 - / - mice at 1, 3, and 4 weeks, thrombin was generated in the TGA assay, and the aPTT clotting time was improved, indicating that the expressed hFIX - R338L is functional. The present invention provides, for example, the following items. (Item 1) A non-human animal comprising a humanized endogenous albumin locus in its genome, wherein a segment of the endogenous albumin locus is deleted and replaced with a corresponding human albumin sequence. (Item 2) The non-human animal according to Item 1, wherein the humanized endogenous albumin locus encodes a protein comprising a human serum albumin peptide. (Item 3) The non-human animal according to Item 1 or 2, wherein the humanized endogenous albumin locus encodes a protein comprising a human albumin propeptide. (Item 4) The non-human animal according to any one of the preceding items, wherein the humanized endogenous albumin locus encodes a protein comprising a human albumin signal peptide. (Item 5) The non-human animal according to any one of the preceding items, wherein the region of the endogenous albumin locus comprising both coding and non-coding sequences is deleted and replaced with a corresponding human albumin sequence comprising both coding and non-coding sequences. (Item 6) The non-human animal according to any one of the preceding items, wherein the humanized endogenous albumin locus comprises the endogenous albumin promoter, and the human albumin sequence is operably linked to the endogenous albumin promoter. (Item 7) The non-human animal according to any one of the preceding items, wherein at least one intron and at least one exon of the endogenous albumin locus are deleted and replaced with the corresponding human albumin sequence. (Item 8) The non-human animal according to any one of the preceding items, wherein the entire albumin coding sequence of the endogenous albumin locus is deleted and replaced with the corresponding human albumin sequence. (Item 9) The non-human animal according to Item 8, wherein the region of the endogenous albumin locus from the start codon to the stop codon is deleted and replaced with the corresponding human albumin sequence. (Item 10) The non-human animal according to any one of the preceding items, wherein the humanized endogenous albumin locus comprises a human albumin 3' untranslated region. (Item 11) A non-human animal according to any one of the preceding items, wherein the endogenous albumin 5'-untranslated region is not deleted and is not replaced with the corresponding human albumin sequence. (Item 12) The region of the endogenous albumin locus from the start codon to the stop codon is deleted and replaced with a human albumin sequence including the corresponding human albumin sequence and the human albumin 3'-untranslated region, and the endogenous albumin 5'-untranslated region is not deleted and is not replaced with the corresponding human albumin sequence, and the endogenous albumin promoter is not deleted and is not replaced with the corresponding human albumin sequence, a non-human animal according to any one of the preceding items. (Item 13) (i) The human albumin sequence at the humanized endogenous albumin locus comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 35, or (ii) The humanized endogenous albumin locus encodes a protein comprising a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the sequence set forth in SEQ ID NO: 5, (iii) The humanized endogenous albumin locus comprises a coding sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 13, or (iv) The humanized endogenous albumin locus comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 17 or 18, a non-human animal according to any one of the preceding items. (Item 14) A non-human animal according to any one of the preceding items, wherein the humanized endogenous albumin locus does not contain a selection cassette or a reporter gene. (Item 15) A non-human animal according to any one of the preceding items, wherein the non-human a...
Claims
**Claim 1** A non-human animal comprising a humanized endogenous albumin locus in its genome, wherein a segment from the start codon to the stop codon of the endogenous albumin locus is deleted and replaced with a corresponding human albumin sequence comprising a region from the start codon to the stop codon of the human albumin sequence and the human albumin 3' untranslated region, wherein the humanized endogenous albumin locus comprises an endogenous albumin promoter, and the human albumin sequence is operably linked to the endogenous albumin promoter, the endogenous albumin 5' untranslated region is not deleted and not replaced with a corresponding human albumin sequence, the non-human animal is homozygous for the humanized endogenous albumin locus, and the non-human animal is a mouse. **Claim 2** (i) the human albumin sequence in the humanized endogenous albumin locus comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 35, or (ii) the humanized endogenous albumin locus encodes a protein comprising a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the sequence set forth in SEQ ID NO: 5, or (iii) the humanized endogenous albumin locus comprises a coding sequence comprising a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 13, or (iv) the humanized endogenous albumin locus comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 17 or 18, the non-human animal according to claim 1. **Claim 3** The non-human animal according to claim 1 or 2, wherein the humanized endogenous albumin locus does not contain a selection cassette or a reporter gene. **Claim 4** The non-human animal according to any one of claims 1 to 3, wherein the non-human animal comprises the humanized endogenous albumin locus in its germline. **Claim 5** (i) the non-human animal comprises a serum albumin level of at least about 10 mg / mL; or (ii) the serum albumin level in the non-human animal is at least as high as the serum albumin level in a control non-human animal comprising a wild-type albumin locus. The non-human animal according to any one of claims 1 to 4.
6. The non-human animal according to any one of claims 1 to 5, further comprising a coding sequence of an exogenous protein integrated into at least one allele of the humanized endogenous albumin locus in one or more cells of the non-human animal.
7. The non-human animal according to claim 6, wherein the coding sequence of the exogenous protein is integrated into intron 1 of the at least one allele of the humanized endogenous albumin locus in the one or more cells of the non-human animal.
8. The non-human animal according to any one of claims 1 to 7, further comprising an inactivated endogenous locus that is not the endogenous albumin locus.
9. The non-human animal according to claim 8, further comprising a coding sequence of an exogenous protein integrated into at least one allele of the humanized endogenous albumin locus in one or more cells of the non-human animal, wherein the exogenous protein replaces the function of the inactivated endogenous locus.
10. The non-human animal according to claim 8 or 9, wherein the inactivated endogenous locus is an inactivated F9 locus.
11. A non-human animal cell, comprising a humanized endogenous albumin locus in its genome, wherein a segment from the start codon to the stop codon of the endogenous albumin locus is deleted and replaced with a corresponding human albumin sequence comprising a region from the start codon to the stop codon of the human albumin sequence and the human albumin 3' untranslated region, the humanized endogenous albumin locus comprises an endogenous albumin promoter, and the human albumin sequence is operably linked to the endogenous albumin promoter, the endogenous albumin 5' untranslated region is not deleted and not replaced with a corresponding human albumin sequence, the non-human animal cell is homozygous for the humanized endogenous albumin locus, and the non-human animal is a mouse.
12. A method for evaluating the activity of a human albumin-targeting reagent in vivo, comprising: (a) administering the human albumin-targeting reagent to a non-human animal according to any one of claims 1 to 10; and (b) evaluating the activity of the human albumin-targeting reagent in the non-human animal.
13. The method according to claim 12, wherein the administering comprises adeno-associated virus (AAV)-mediated delivery, lipid nanoparticle (LNP)-mediated delivery, or hydrodynamic delivery (HDD).
14. The method according to claim 13, wherein the administering comprises LNP-mediated delivery.
15. The method according to claim 14, wherein the dose of the LNP is about 0.1 mg / kg to about 2 mg / kg.
16. The method according to claim 13, wherein the administering comprises AAV8-mediated delivery.
17. The method according to any one of claims 12 to 16, wherein step (b) comprises isolating the liver from the non-human animal and evaluating the activity of the human albumin-targeting reagent in the liver.
18. The method according to any one of claims 12 to 17, wherein the human albumin-targeting reagent is a genome editing agent, and the evaluating comprises evaluating the modification of the humanized endogenous albumin locus.
19. The method according to claim 18, wherein the evaluating comprises measuring the frequency of insertions or deletions within the humanized endogenous albumin locus.
20. (i) the evaluating comprises measuring the expression of albumin messenger RNA encoded by the humanized endogenous albumin locus; or (ii) the evaluating comprises measuring the expression of albumin protein encoded by the humanized endogenous albumin locus.
21. The method according to claim 20, wherein evaluating the expression of the albumin protein comprises measuring the serum level of the albumin protein in the non-human animal or measuring the expression of the albumin protein in the liver of the non-human animal.
22. The method according to any one of claims 12 to 21, wherein the human albumin targeting reagent comprises a nuclease agent or a nucleic acid encoding the nuclease agent, and the nuclease agent is designed to target a region of the human albumin gene.
23. The method according to claim 22, wherein the nuclease agent comprises a Cas protein and a guide RNA designed to target a guide RNA target sequence of the human albumin gene.
24. The method according to claim 23, wherein the Cas protein is a Cas9 protein.
25. The method according to claim 23 or 24, wherein the guide RNA target sequence is in intron 1 of the human albumin gene.
26. The method according to any one of claims 12 to 25, wherein the human albumin targeting reagent comprises an exogenous donor nucleic acid, and the exogenous donor nucleic acid is designed to target the human albumin gene.
27. The method according to claim 26, wherein the exogenous donor nucleic acid is delivered via AAV.
28. The method according to claim 26 or 27, wherein the exogenous donor nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
29. The exogenous donor nucleic acid does not contain a homology arm, or The method according to any one of claims 26 to 28, wherein the exogenous donor nucleic acid comprises an insertion nucleic acid adjacent to a 5' homology arm targeting a 5' target sequence of the humanized endogenous albumin locus and a 3' homology arm targeting a 3' target sequence of the humanized endogenous albumin locus.
30. The method according to claim 29, wherein each of the 5' target sequence and the 3' target sequence comprises a segment of intron 1 of the human albumin gene.
31. The method according to any one of claims 26 to 30, wherein the exogenous donor nucleic acid encodes an exogenous protein.
32. The method according to claim 31, wherein the protein encoded by the humanized endogenous albumin locus targeted by the exogenous donor nucleic acid is a heterologous protein comprising a human albumin signal peptide fused to the exogenous protein.
33. The method according to claim 31 or 32, wherein the exogenous protein is a factor IX protein.
34. The method according to claim 33, wherein said evaluating comprises measuring the serum level of said Factor IX protein in said non-human animal and / or evaluating the activated partial thromboplastin time or performing a thrombin generation assay.
35. The human albumin targeting reagent comprises (1) a nuclease agent designed to target a region of the human albumin gene, and (2) an exogenous donor nucleic acid, wherein the exogenous donor nucleic acid is designed to target the human albumin gene, wherein the exogenous donor nucleic acid encodes an exogenous protein, The method according to any one of claims 12 to 34, wherein the protein encoded by the humanized endogenous albumin locus targeted by the exogenous donor nucleic acid is a heterologous protein comprising a human albumin signal peptide fused to the exogenous protein.
36. The method according to any one of claims 31 to 35, wherein said evaluating comprises measuring the expression of messenger RNA encoded by said exogenous donor nucleic acid.
37. The method according to claim 36, wherein said evaluating comprises an in situ hybridization assay for quantifying the expression of messenger RNA encoded by said exogenous donor nucleic acid at single cell resolution.
38. The method according to claim 36 or 37, wherein said evaluating comprises measuring the expression of messenger RNA encoded by said exogenous donor nucleic acid in a plurality of lobes derived from the liver of said non-human animal.
39. The method according to any one of claims 31 to 38, wherein said evaluating comprises measuring the expression of said exogenous protein.
40. The method according to claim 39, wherein evaluating the expression of said heterologous protein comprises measuring the serum level of said heterologous protein in said non-human animal or measuring the expression of said heterologous protein in the liver of said non-human animal.
41. A method for optimizing the activity of a human albumin targeting reagent in vivo, comprising: (I) first performing the method according to any one of claims 12 to 40 in a first non-human animal comprising a humanized endogenous albumin locus in its genome; (II) modifying the variable element and re-performing the method of step (I) a second time in a second non-human animal comprising a humanized endogenous albumin locus in its genome using the modified variable element; (III) comparing the activity of the human albumin targeting reagent of step (I) with the activity of the human albumin targeting reagent of step (II) and selecting a method that results in a higher activity.
42. (i) the modified variable element in step (II) is a delivery method for introducing the human albumin targeting reagent into the non-human animal; (ii) the modified variable element in step (II) is an administration route for introducing the human albumin targeting reagent into the non-human animal; (iii) the modified variable element in step (II) is the concentration or amount of the human albumin targeting reagent introduced into the non-human animal; (iv) the modified variable element in step (II) is the form of the human albumin targeting reagent introduced into the non-human animal; (v) the modified variable element in step (II) is the human albumin targeting reagent introduced into the non-human animal; (vi) the human albumin targeting reagent comprises a Cas protein or a nucleic acid encoding the Cas protein and a guide RNA or a DNA encoding the guide RNA, and the guide RNA is designed to target a guide RNA target sequence of the human albumin gene, and: (A) the modified variable element in step (II) is the guide RNA sequence or the guide RNA target sequence; (B) the human albumin targeting reagent comprises a messenger RNA (mRNA) encoding the Cas protein and the guide RNA, and the modified variable element in step (II) is the ratio of Cas mRNA to guide RNA; (C) the modified variable element in step (II) is a guide RNA modification; or (vii) the human albumin targeting reagent comprises an exogenous donor nucleic acid, and: (A) the modified variable element in step (II) is the form of the exogenous donor nucleic acid; (B) The exogenous donor nucleic acid includes a 5' homology arm targeting the 5' target sequence of the humanized endogenous albumin locus and an insertion nucleic acid adjacent to a 3' homology arm targeting the 3' target sequence of the humanized endogenous albumin locus, and the modified variable element in step (II) is the sequence and / or length of the 5' homology arm and / or the sequence and / or length of the 3' homology arm. The method according to claim 41.
43. The method according to claim 42, wherein the administering includes LNP-mediated delivery, and the modified variable element in step (II) is an LNP formulation.
44. A method for producing a non-human animal according to any one of claims 1 to 10, comprising: (I) (a) modifying the genome of a pluripotent non-human animal cell to include the humanized endogenous albumin locus; (b) identifying or selecting the genetically modified pluripotent non-human animal cell comprising the humanized endogenous albumin locus; (c) introducing the genetically modified pluripotent non-human animal cell into a non-human animal host embryo; (d) implanting the non-human animal host embryo into a non-human animal surrogate mother, or (II) (a) modifying the genome of a one-cell stage embryo of a non-human animal to include the humanized endogenous albumin locus; (b) selecting the genetically modified one-cell stage embryo of a non-human animal comprising the humanized endogenous albumin locus; (c) implanting the genetically modified one-cell stage embryo of a non-human animal into a non-human animal surrogate mother.
Citation Information
Patent Citations
Non-human transgenic animals for pharmacological and toxicological studies
JP2004533826A