High-efficiency site-specific knock-in method for long-chain nucleic acid sequence

The use of circular DNA vectors and CRISPR-Cas9-mediated double-strand cleavage in the described method enables efficient and specific knock-in of long-chain nucleic acid sequences into non-human mammals, overcoming the limitations of traditional techniques.

WO2025126797A1PCT designated stage expired Publication Date: 2025-06-19THE UNIV OF TOKYO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041314
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-11-21
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for knocking in long-chain nucleic acid sequences into the genome of non-human mammals are inefficient and often result in non-specific gene insertion, making it difficult to achieve high-throughput genotyping and maintain pluripotency in embryonic stem cells.

Method used

A method using circular DNA vectors and CRISPR-Cas9-mediated double-strand cleavage to facilitate high-efficiency, locus-specific knock-in of human orthologous full-length genes into non-human mammals, utilizing homologous recombination and drug resistance gene selection to enhance specificity and efficiency.

Benefits of technology

This method achieves nearly 100% hit clone establishment efficiency for inserting human orthologous full-length genes of up to 10 kbp into specific loci, significantly improving the efficiency and specificity of gene knock-in compared to traditional linear DNA vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JPOXMLDOC01-APPB-T000001
    Figure JPOXMLDOC01-APPB-T000001
  • Figure JPOXMLDOC01-APPB-T000002
    Figure JPOXMLDOC01-APPB-T000002
  • Figure JPOXMLDOC01-APPB-T000003
    Figure JPOXMLDOC01-APPB-T000003
Patent Text Reader

Abstract

The purpose of the present invention is to provide a method for substituting a full-length gene of a non-human mammal with a human ortholog full-length gene with high reproducibility and high efficiency by using a homologous recombination technique. According to the present invention, there is provided a method for the knock-in of a human ortholog full-length gene into a corresponding locus in a human mammal, the method being characterized by using: a first cyclic vector that targets a non-human mammalian ortholog full-length gene in a non-human mammalian ortholog full-length gene locus; and a second cyclic vector that includes a human ortholog full-length gene and a human homologous arm sequence located at each of the 5' end and the 3' end of the human ortholog full-length gene.
Need to check novelty before this filing date? Find Prior Art

Description

Highly efficient site-specific knock-in of long nucleic acid sequences

[0001] The present invention relates to a method for highly efficiently knocking in a long nucleic acid sequence into the genome of a non-human mammal.

[0002] The development of genetically engineered mouse models and their phenotypic characterization are robust techniques used to precisely analyze gene function or physiological behavior of specific cells within an animal. Several important biological phenomena have been elucidated using various genetically engineered animal models. A locus-specific exogenous gene integration technique called gene knock-in (KI) has several advantages, such as precise integration site, controlled copy number, and controlled KI gene expression, which more closely resembles the locus-specific gene expression pattern of a specific cell type (Non-Patent Documents 1 and 2). The conventional method for developing genetic KI mouse models is to use embryonic stem cell (ES cell) targeting technology, followed by ES cell injection into preimplantation embryos to generate chimeras (Non-Patent Documents 3-5).

[0003] Induction of double-strand breaks in the genome increases the efficiency of transgene integration via homologous recombination in mammalian cells (Non-Patent Document 6). More recently, site-specific endonucleases such as ZFNs (Non-Patent Documents 7 and 8), TALENs (Non-Patent Documents 9 and 10), and CRISPR-Cas9 (Non-Patent Document 11), which induce double-strand breaks at specific sequences, have been introduced, making it possible to directly modify the genome of fertilized eggs. Among these genome editing techniques, CRISPR-Cas9 is currently used in many laboratories as a powerful tool for fertilized egg genome editing due to its simplicity and accuracy. CRISPR-Cas9-mediated locus-specific gene manipulation in mouse embryos was first reported by introducing single guide RNA and Cas9 mRNA into the pronuclei of fertilized eggs using micromanipulation, and several papers have reported that analysis of F0 mice generated by CRISPR-Cas9-mediated fertilized egg genome editing revealed several important gene functions in several fields, such as spermatogenesis and spermatogenesis. In addition to indel mutations, site-specific short-nucleotide KI, such as single-base substitution or peptide tag insertion into mouse zygotes, can also be achieved by using single-stranded oligonucleotides (ssODNs) as KI donors via homology-directed repair. In addition to short fragment KI, DNA sequences longer than several kilobase pairs (kbp) can be integrated into the zygote genome by CRISPR-Cas9-mediated genome editing using double-stranded plasmids, long single-stranded DNA (lssDNA), or adeno-associated viruses (AAV) as KI donors.

[0004] Although zygote genome editing has been reported using several mouse strains, such as C3H / HeJ or B6D2F1, the C57BL / 6 strain is commonly used due to its flexibility in in vitro manipulation, including superovulation, in vitro fertilization, or in vitro embryo culture, as well as pure genetic backgrounds. On the other hand, zygote genome editing remains challenging in certain strains, such as BALB / c, which is an important strain for certain research areas, such as immunology, due to its poor response to superovulation or high susceptibility to in vitro embryo manipulation (Non-Patent Document 12).

[0005] Compared with direct genome editing of zygotes, genetic manipulation using ES cells is believed to have several advantages for generating genetically engineered mouse models. For example, high-throughput genotyping screening can be performed in vitro, and ES cell lines established from various inbred strains, including BALB / c, DBA2, or C57BL / 6, can be used for genetic manipulation. Another major advantage is that sequential KI can be performed to establish KI ES cell clones carrying multiple exogenous genes, such as multiple fluorescent reporters or CreERT2 and loxP-STOP-loxP, at different genomic loci. However, modifying multiple genes is time-consuming, and the pluripotency of ES cells is known to decrease with longer culture periods. Therefore, it is necessary to develop a highly efficient technique that can KI long DNA fragments at multiple loci in a single gene transfer.

[0006] Furthermore, more than 15,000 orthologous genes are known between humans and mice, and humanized mammal models in which specific human genes are introduced into small, easily kept mammals such as mice and the human genes are expressed are extremely important tools in medical research, drug discovery, and other fields. Conventionally, when introducing a foreign gene (e.g., a human gene) into a genetic locus in a non-human mammal, it has been common to knock-in the human gene (cDNA, excluding untranslated regions) into the existing genetic locus of the non-human mammal by homologous recombination (substitution). The nucleic acid sequence of this insertable foreign gene is at most about 10 kbp long, and contains only the minimum genetic information required for protein expression.

[0007] Another known transgenic mouse model uses a bacterial artificial chromosome (BAC) to generate a full-length human gene. Unlike the homologous recombination knock-in technique described above, this method has the advantage that the inserted foreign gene can contain untranslated regions. Furthermore, although the length of the inserted nucleic acid sequence is 100 kbp or more, the recombination efficiency is low and, because homologous recombination is not used, the desired human gene is inserted randomly into the mouse genome, and its copy number and chromosome location cannot be controlled. Furthermore, because homologous recombination is not used, not only the desired human gene but also the corresponding mouse gene remains in the mouse genome.

[0008] Wang, F. and Qi, LS, Trends Cell Bil., 26, 875-888 (2016)Webster, JD, et al., Cell Tissue Res., 380, 325-340 (2020)Mansour, SL, et al. Nature, 336, 348-352 (1988)Johnson, RS, et al., Science, 245, 1234-1236 (1989)Hasty, P., et al., Nature, 350, 243-246 (1991)Rouet, P., et al., Mol. Cell Biol., 14, 8096-8106 (1994)Geurts, AM, et al., Science, 325, 433 (2009)Carbery, ID, et al., Genetics, 186, 451-459 (2010)Tesson, L., et al., Nat. Biotechnol., 29, 695-696 (2011)Sung, YH, et al., Nat. Biotechnol., 31, 23-24 (2013)Wang, H., et al., Cell, 153, 910-918 (2013) Golkar-Narenji, A., et al., Reprod. Med. Biol., 11, 185-192 (2012)

[0009] In view of the above-mentioned drawbacks of the conventional techniques, the present invention aims to provide a method for replacing a full-length gene of a non-human mammal with a full-length human orthologous gene with good reproducibility and high efficiency using homologous recombination techniques.

[0010] Previously, gene knock-in techniques for host cells typically involved the use of linearized DNA as a targeting vector and the selection of knock-in strains using drugs. However, the use of such linearized DNA as a targeting vector has been problematic due to low integration efficiency and high rates of nonspecific gene insertion. Furthermore, such knock-in techniques have resulted in low clone yields of approximately 1-10%. Therefore, the present inventors have avoided the use of linearized DNA (also known as "linear DNA") and instead used circular DNA (plasmid) as a targeting vector. Furthermore, by combining this with double-strand breaks via genome editing, they have successfully suppressed nonspecific gene insertion and significantly improved integration specificity. Furthermore, they have successfully obtained cells in which a full-length human orthologous gene of approximately 10 kbp has been inserted into the corresponding mouse locus with an extremely high clone establishment efficiency of nearly 100%, thereby completing the present invention.

[0011] That is, the present invention is as follows: [1] A method for knocking in a full-length human orthologous gene into a corresponding gene locus in a non-human mammal, the method comprising: (a) preparing a first circular vector targeting the full-length human orthologous gene at the full-length human orthologous gene locus, the first circular vector comprising a gene encoding a drug resistance gene, and at its 5'-end and 3'-end, respectively, non-human mammal homologous arm sequences flanking the full-length human orthologous gene, and further comprising human homologous arm sequences flanking each of the non-human mammal homologous arm sequences, (b) knocking in a gene encoding a drug resistance gene sandwiched between human homologous sequences into the full-length orthologous gene locus in a non-human mammal cell by homologous recombination using the first circular vector prepared in (a), (c) preparing a second circular vector comprising the full-length human orthologous gene and human homologous arm sequences at its 5'-end and 3'-end, respectively, (d) knocking in a gene encoding a full-length human orthologous gene and a drug resistance gene into the knock-in cell obtained in step (b) by homologous recombination using the second circular vector prepared in step (c). [2] The method described in [1], wherein the knock-in is performed by a genome editing method using homologous recombination. [3] The method described in [2], wherein the genome editing method uses CRISPR-Cas9, Cas12a, ZFN, or TALEN. [4] The method described in any of [1] to [3], wherein the length of the full-length human orthologous gene is 10 kbp or more. [5] The method described in any of [1] to [4], wherein the step of selecting cells with an antibiotic is performed after steps (b) and (d). [6] The method described in any of [1] to [5], wherein the first circular vector is derived from a plasmid vector, a viral vector, or a cosmid vector. [7] The method described in any of [1] to [6], wherein the second circular vector is derived from a bacterial artificial chromosome (BAC).[8] The method according to any one of [1] to [7], wherein the drug resistance gene is one or more selected from the group consisting of a neomycin resistance gene, a puromycin resistance gene, a blasticidin resistance gene, a hygromycin resistance gene, a chloramphenicol resistance gene, a tetracycline resistance gene, an erythromycin resistance gene, a spectinomycin resistance gene, a kanamycin resistance gene, a zeocin resistance gene, and a phleomycin resistance gene. [9] A kit for knocking in a full-length human orthologous gene into a corresponding gene locus in a non-human mammal, comprising: (i) a first circular vector targeting the full-length non-human mammalian orthologous gene at the full-length non-human mammalian orthologous gene locus, the first circular vector comprising a gene encoding a drug resistance gene, and at its 5'-end and 3'-end, respectively, non-human mammalian homologous arm sequences flanking the full-length non-human mammalian orthologous gene, and further comprising human homologous arm sequences flanking each of the non-human mammalian homologous arm sequences; and (ii) a second circular vector comprising the full-length human orthologous gene and human homologous sequences at its 5'-end and 3'-end, respectively.

[10] The kit according to [9], wherein the length of the full-length human orthologous gene is 10 kbp or more.

[0012] By using the method of the present invention for highly efficient knock-in of long DNA sequences into the genome of non-human mammals, it is possible to introduce each of more than 15,000 human orthologous genes into the genome of small mammals with a life cycle of approximately two years, thereby easily providing so-called "humanized mammals."

[0013] CRISPR-Cas9 ribonucleoprotein (RNP)-mediated integration of circular plasmids into the ES cell genome is highly locus-specific. (A, B) Schematic diagrams of the circular plasmid introduction strategy are shown. A linear vector containing a 1-kbp CAG-EGFP cassette flanked by gRNA targets was transduced into ES cells via electroporation (A). The circular plasmid as a targeting vector was introduced into ES cells by single electroporation (B) using a CRISPR-Cas9 expression vector (left) or CRISPR-Cas9 ribonucleoprotein (RNP, right). (C) Representative ES cell colonies are shown under a fluorescent microscope. The upper panel shows an EGFP image, and the lower panel shows a merged image of EGFP and bright field. Left: No electroporation (No-EP); Center-Left: Transduction of targeting vector only; Center-Right: Transduction of targeting vector with CRISPR-Cas9 all-in-one plasmid; Right: Transduction of targeting vector with CRISPR-Cas9 RNP. Scale bar, 50 μm. (D) Flow cytometry of ES cells. Gates represent EGFP-positive fractions. (E) EGFP-positive cell ratios are shown as mean ± SEM. Asterisks indicate significant differences (n = 3, P < 0.01). (F, G) Genomic PCR analysis of EGFP-positive ES cell clones collected by flow cytometry. Gel images or KI ratios are shown in F or G, respectively. Red arrows in the upper and lower panels indicate 5' or 3' Rosa26 KI, respectively, and black arrows indicate genomic DNA (gDNA) PCR controls. Both 5' and 3' PCR-positive clones are indicated with red numbers. NC negative control used genomic DNA from wild-type B6 mouse tail. F, G) Genomic PCR analysis of EGFP-positive ES cell clones collected by flow cytometry. Gel images or KI ratios are shown in F or G, respectively. Red arrows in the upper or lower panels indicate 5' or 3' Rosa26 KI, respectively, and black arrows indicate genomic DNA (gDNA) PCR controls. Both 5' and 3' PCR-positive clones are indicated with red numbers. NC negative control used genomic DNA from wild-type B6 mouse tail. Cas9 RNP-mediated plasmid DNA KI is due to homologous recombination.(A) Schematic diagram of site-specific Cas9-RNP-mediated genome editing of the Rosa26 or tyrosinase locus. A circular vector containing a CAG-EGFP cassette flanked by the 5' and 3' homologous arms of Rosa26 was transduced into ES cells via electroporation with either gRosa26-RNP or gTyr-RNP. (B) T7 endonuclease I mismatch cleavage (T7EI) analysis was performed to verify the efficiency of double-strand break induction by the tyrosinase gene-directed CRISPR-Cas9 RNP. The top, middle, and bottom arrows indicate the uncut wild-type genome, the long half of the mutant genome, or the short half of the mutant genome, respectively. The numbers at the bottom of the image indicate the induction ratio of indel mutations at the tyrosinase locus in targeted ES cells. The indel mutation ratio was calculated using the formula mentioned elsewhere (http: / / crispr.technology / resources / quantification.html). (C) Representative ES cell colonies under a fluorescent microscope are shown. The upper panel shows an EGFP image, and the lower panel shows a merged image of EGFP and bright field. Left: No electroporation (No-EP); center-left: Transduction of pR26-CE plasmid alone; center-right: Transduction of pR26-CE targeting vector with tyrosinase locus-specific CRISPR-Cas9 RNP; right: Transduction of pR26-CE targeting vector with Rosa26 locus-specific CRISPR-Cas9 RNP; right: Transduction of pR26-CE targeting vector with Rosa26 locus-specific CRISPR-Cas9 RNP. Scale bar, 50 μm. (D) Flow cytometry of ES cells. The gate represents the EGFP-positive fraction. (E) The GFP-positive cell ratio is shown as mean ± SEM. Asterisks indicate significant differences (n = 3, P < 0.01). (F) Schematic diagram of genome editing at the Rosa26 locus using site-specific Cas9-RNP. Targeting vectors with or without the 5' and 3' homologous arms of Rosa26 were transduced into ES cells with gRosa26-RNP via electroporation. (G) Flow cytometry of ES cells. The gate represents the EGFP-positive fraction.(H) GFP-positive cell ratios are shown as mean ± SEM. Asterisks indicate significant differences (n = 3, P < 0.001). KI efficiency was dramatically increased by the targeting vector carrying a drug resistance gene cassette. (A, B) Schematic diagrams of the Rosa26 locus KI and drug selection strategy. (C) Representative ES cell colonies under a fluorescent microscope. The upper panel shows an EGFP image, and the lower panel shows a merged image of EGFP and bright field. Left: no electroporation (No-EP); middle: transduction of the targeting vector with gRosa26-RNP without drug selection; right: transduction of the targeting vector with gRosa26-RNP followed by G418 selection. Scale bar, 50 μm. (D, E) Genomic PCR analysis of ES cell clones. Gel images and KI ratios are shown in D and E, respectively. Red arrows in the upper and lower panels indicate 5' or 3' Rosa26 KI, respectively, and black arrows indicate genomic DNA (gDNA) PCR controls. Both 5' and 3' PCR-positive clones are indicated with red numbers. NC: negative control using genomic DNA from wild-type B6 mouse tails. (F) Schematic diagram of the Stra8 locus KI strategy. (G) Representative ES cell colonies under a fluorescent microscope. The upper panel shows a GFP image, and the lower panel shows a bright-field image. Left: transduction of the targeting vector without gStra8-RNP, followed by puromycin selection; right: transduction of the targeting vector with gStra8-RNP, followed by puromycin selection. Scale bar, 50 μm. (H) Genomic PCR analysis of ES cell clones. Gel images or KI ratios are indicated by H or I, respectively. Red arrows in the upper and lower panels indicate 5' or 3'Stra8 KI, respectively. (I) Genomic PCR analysis of ES cell clones. Gel images or KI ratios are indicated by H or I, respectively. The upper or lower red arrows indicate 5' or 3' Stra8 KI, respectively. The black arrows indicate genomic DNA (gDNA) PCR controls. Both 5' and 3' PCR-positive clones are indicated by red numbers. A negative control was performed using genomic DNA from wild-type B6 mouse tails. Biallelic KI in ES cells could be achieved by a single electroporation. (A) Schematic diagram of the Rosa26 locus KI strategy using two independent targeting vectors.(B) Genomic PCR analysis of ES cell clones. The upper and lower red arrows indicate Bsd-GOI-A KI or NeoR-GOI-B, respectively. The black arrow indicates a genomic DNA (gDNA) PCR control. Both 5' and 3' PCR-positive clones are indicated with red numbers. NC1: Negative control using genomic DNA from wild-type B6 mouse tail. NC2: Negative control using a targeting vector. (C) Schematic diagram of the combined Rosa26 and Cd6 locus KI strategy using two independent targeting vectors and gRNA-RNP. (D) Genomic PCR analysis of ES cell clones. Gel images or KI ratios are shown in D and E, respectively. (E) Genomic PCR analysis of ES cell clones. Gel images or KI ratios are shown in D and E, respectively. The red arrows indicate the 5' or 3' KI-specific bands of the Rosa26 or Cd6 locus, respectively. Both Rosa26 KI and Cd6 KI clones are indicated by red numbers. NC: Negative control using genomic DNA from wild-type B6 mouse tail. (F) Schematic diagram of the gene construct for tamoxifen-inducible Cre / loxP recombination at the Rosa26 and Cd6 loci. (G) Representative double KI ES cell colonies are shown under a fluorescent microscope with (upper panel) or without (lower panel) 4-hydroxytamoxifen (4OHT). The left panel shows GFP, the middle panel shows RFP, and the left panel shows bright field. Scale bar, 50 μm. (H) Schematic diagram of tissue sampling after tamoxifen administration in chimeric mice carrying both Rosa26 KI and Cd6 KI alleles. (I) RFP expression in tissues of chimeric mice is shown. Mice were intraperitoneally administered tamoxifen five times and used for fluorescent microscopy.

[0014] Other terms and concepts in the present invention are based on the meanings of terms commonly used in the relevant technical field, and various techniques used to carry out the present invention can be easily and reliably carried out by those skilled in the art based on known literature, etc., except for techniques whose sources are particularly specified. For example, genetic engineering techniques can be carried out by the methods described in J. Sambrook, E. F. Fritsch & T. Maniatis, Molecular Cloning: A Laboratory Manual (2nd edition), Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1989); D. M. Glover et al. ed., DNA Cloning, 2nd ed., Vol. 1 to 4, (The Practical Approach Series), IRL Press, Oxford University Press (1995), etc., or by substantially similar methods or modified methods thereto. Information regarding the various proteins and peptides used in the present invention, or the DNAs encoding them, can be obtained from the existing TAIR or NCBI GenBank database (URL: http: / / www.arabidopsis.org / or http: / / www.ncbi.nlm.nih.gov / sites / entrez?db=pubmed, etc.). The contents of the technical literature, patent publications, and patent application specifications cited in this specification are to be referenced as the contents of the present invention.

[0015] DEFINITIONS The terms "nucleic acid" and "polynucleotide" are used interchangeably herein and include polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These terms include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.

[0016] Nucleic acids are said to have a "5' end" and a "3' end" because mononucleotides react to form oligonucleotides such that the 5' phosphate of one mononucleotide pentose ring is attached unidirectionally to the nearby 3' oxygen by a phosphodiester linkage. An end of an oligonucleotide is referred to as the "5' end" if its 5' phosphate is not linked to the 3' oxygen of a mononucleotide pentose ring. An end of an oligonucleotide is referred to as the "3' end" if its 3' oxygen is not linked to the 5' phosphate of another mononucleotide pentose ring. Nucleic acid sequences, even if internal to a larger oligonucleotide, can also be said to have 5' and 3' ends. In either a linear or circular DNA molecule, distinct elements are referred to as being "upstream" or 5' of the "downstream" or 3' element.

[0017] A "genome" is the entirety of the genetic material contained in each cell of an organism. Unless otherwise indicated, a particular nucleic acid sequence of the present invention also encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences as well as the explicitly indicated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. The term nucleic acid molecule is used interchangeably with gene, cDNA, and mRNA encoded by a gene.

[0018] The term "genomically integrated" refers to a nucleic acid introduced into a cell such that the nucleic acid (nucleotide) sequence is integrated into the genome of the cell. Any protocol can be used for stable integration of a nucleic acid into the genome of a cell. According to the present invention, a long nucleic acid (genomic) sequence can be knocked into the genome of a non-human mammal with high efficiency.

[0019] As noted above, a "transgene" refers to a nucleic acid molecule that is introduced into a genome by transformation and stably maintained. A transgene may contain at least one expression cassette, typically contains at least two expression cassettes, and may contain ten or more expression cassettes. A transgene may include, for example, a gene that is orthologous to the particular gene being transformed. In general, the term "endogenous gene" refers to a native gene in its natural location in the genome of an organism. In contrast, a "foreign" gene refers to a gene not normally found in the host organism but that is introduced into an organism by gene transfer.

[0020] "Intron" refers to an intervening segment of DNA that is found almost exclusively in eukaryotic genes but is not translated into an amino acid sequence in the gene product. Introns are removed from the premature mRNA through a process called splicing, which leaves the exons intact to form the mRNA. "Exon" also refers to a segment of DNA that contains the coding sequence for a protein or part thereof. Exons are separated by intervening non-coding sequences (introns).

[0021] The terms "protein," "polypeptide," and "peptide" are used interchangeably herein and include polymeric forms of amino acids of any length, including coded and non-coded amino acids, and amino acids that are chemically or biochemically modified or derivatized.

[0022] Proteins have an "N-terminus" and a "C-terminus." The term "N-terminus" refers to the free amine group (-NH 2 The term "C-terminus" refers to the end of a chain of amino acids (protein or polypeptide) that ends with a free carboxyl group (-COOH).

[0023] The terms "transduction" and "transfect" refer to the introduction of a molecule, such as a nucleic acid (viral vector, plasmid), into a cell. A cell is "transduced" or "transfected" when an exogenous nucleic acid (according to the present invention, a "human orthologous gene") is introduced inside the cell membrane. Thus, a "transduced cell" is a cell into which a "nucleic acid" or "polynucleotide" has been introduced, or its progeny into which an exogenous nucleic acid has been introduced. In certain embodiments, a "transduced" cell (e.g., in a mammal, e.g., a cell or tissue or organ cell) may have a genetic change following incorporation of an exogenous molecule, e.g., a nucleic acid (e.g., a transgene). The "transduced" cell may be propagated and may transcribe the introduced nucleic acid and / or express a protein.

[0024] Nucleic acids (sometimes referred to as "plasmids") can be single-stranded, double-stranded, or triple-stranded, linear or circular, and can be of any length. When discussing nucleic acids (plasmids), the sequence or structure of a particular polynucleotide may be described herein according to the convention of providing the sequence in the 5' to 3' direction.

[0025] "Plasmid" refers to a form of nucleic acid or polynucleotide that typically has additional elements for expression (e.g., transcription, replication, etc.) or propagation (replication) of the plasmid. As used herein, plasmid can also be used to refer to nucleic acid and polynucleotide sequences.

[0026] The term "vector (or expression vector)" or "expression construct" refers to a recombinant nucleic acid containing a desired coding sequence operably linked to appropriate nucleic acid sequences necessary for the expression of the operably linked coding sequence in a particular host cell or organism. In eukaryotic cells, it is generally known to use promoters, enhancers, and termination and polyadenylation signals, although some elements can be deleted and others added without sacrificing the necessary expression.

[0027] As used herein, "vector" refers to a nucleic acid construct designed for transfer between different hosts, such as, but not limited to, a plasmid, virus, cosmid, phage, BAC, or YAC. A "viral vector" is defined as a recombinantly produced virus or viral particle containing a polynucleotide to be delivered into a host cell, either in vivo, ex vivo, or in vitro. In some embodiments, a plasmid vector can be prepared from a commercially available vector. In other embodiments, a viral vector can be produced from a baculovirus, retrovirus, adenovirus, AAV, or the like, according to techniques known in the art. In one embodiment, the viral vector is a lentiviral vector. Examples of viral vectors include retroviral vectors, adenoviral vectors, adeno-associated viral vectors, alphaviral vectors, and the like. Vectors containing both a promoter and a cloning site into which a polynucleotide can be operably linked are also well known in the art. Such vectors are capable of transcribing RNA in vitro or in vivo, and are commercially available from sources such as Agilent Technologies (Santa Clara, Calif.) and Promega Biotech (Madison, Wis.).

[0028] "Targeting vector" or "target vector" are used interchangeably and refer to a recombinant nucleic acid that can be introduced by homologous recombination, non-homologous end joining-mediated ligation, or any other means of recombination into a target location in the genome of a cell.

[0029] Transformation of host cells (nucleic acid introduction) can be performed using commonly used known methods, such as, but not limited to, electroporation (Mackenxie, DA et al., Appl. Environ. Microbiol., vol. 66, pp. 4655-4661, 2000), heat shock (U.S. Pat. No. 2,765,299), lipofection (PNAS, 1989, 86: 6077; PNAS, 1987, 84: 7413), particle delivery (JP 2005-287403 A), spheroplast method (Proc. Natl. Acad. Sci. USA, vol. 75, pp. 1929-1978), and lithium acetate method (J. Bacteriology, vol. 153, pp. 163, 1983).

[0030] The term "isolated" with respect to proteins, nucleic acids, and cells includes proteins, nucleic acids, and cells that are relatively purified with respect to other cellular or biological components that may normally be present in situ, up to and including substantially pure preparations of the protein, nucleic acid, or cell. The term "isolated" also includes proteins and nucleic acids that have no naturally occurring counterpart or proteins or nucleic acids that are chemically synthesized and thus are substantially free from contaminating other proteins or nucleic acids. The term "isolated" also includes proteins, nucleic acids, or cells that have been separated or purified from many other cellular or biological components with which they are naturally associated (e.g., other cellular proteins, nucleic acids, or cellular or extracellular components).

[0031] The term "wild-type" includes entities having a structure and / or activity found in a normal state or context (as opposed to a mutant, diseased, altered state or context, etc.). Wild-type genes and polypeptides often exist in many different forms (e.g., alleles).

[0032] The term "endogenous sequence" refers to a nucleic acid sequence that is naturally present in a cell or non-human animal. For example, an endogenous Rosa26 sequence of a non-human mammal refers to the native Rosa26 sequence that is naturally present at the Rosa26 locus in the non-human mammal.

[0033] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normally present includes present with respect to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence may include, for example, a mutated version of a corresponding endogenous sequence in a cell, such as a humanized version of an endogenous sequence, or may include a sequence that corresponds to an endogenous sequence in a cell but in a different form (i.e., not in a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form in a particular cell at a particular developmental stage under particular environmental conditions.

[0034] The term "heterologous," when used with respect to a nucleic acid or a protein, indicates that the nucleic acid or protein comprises at least two segments that are not naturally found together in the same molecule. For example, the term "heterologous," when used with respect to a segment of a nucleic acid or a segment of a protein, indicates that the nucleic acid or protein comprises two or more subsequences that are not found in the same relationship (e.g., joined) to each other in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that other molecule in nature. For example, a heterologous region of a nucleic acid vector can include a coding sequence flanked by sequences not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with that other peptide molecule in nature (e.g., a fusion protein, or a protein with a tag). Similarly, a nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.

[0035] "Codon optimization" takes advantage of codon degeneracy, as indicated by the multiplicity of three-base pair codon combinations that specify amino acids, and generally involves modifying a nucleic acid sequence by replacing at least one codon of the native sequence with a codon more frequently or most frequently used in the host cell's genes while maintaining the native amino acid sequence, for enhanced expression in a particular host cell. For example, a nucleic acid encoding a Cas9 protein can be modified to replace a codon with a codon more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, in the "Codon Usage Database." These tables can be adapted in several ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a particular sequence for expression in a particular host (see, eg, Gene Forge).

[0036] A "promoter" is a regulatory region of DNA that typically contains a TATA box that can direct RNA polymerase II to begin RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. Promoters may further contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein modulate the transcription of an operably linked polynucleotide. The promoter may be active in one or more of the cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, one-cell stage embryos, differentiated cells, or combinations thereof). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO 2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.

[0037] A constitutive promoter is one that is active in all tissues or in specific tissues at all developmental stages. Examples of constitutive promoters include the human cytomegalovirus immediate early (hCMV) promoter, the mouse cytomegalovirus immediate early (mCMV) promoter, the human elongation factor 1 alpha (hEF1α) promoter, the mouse elongation factor 1 alpha (mEF1α) promoter, the mouse phosphoglycerate kinase (PGK) promoter, the chicken beta actin hybrid (CAG or CBh) promoter, the SV40 early promoter, and the beta 2 tubulin promoter.

[0038] Examples of inducible promoters include chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include, for example, alcohol-regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., tetracycline-responsive promoter, tetracycline operator sequence (tetO), tet-On promoter, or tet-Off promoter), steroid-regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-regulated promoters (e.g., metalloprotein promoters). Physically regulated promoters include, for example, temperature-regulated promoters (e.g., heat shock promoters) and light-regulated promoters (e.g., light-inducible promoters or light-repressible promoters).

[0039] The tissue-specific promoter may be, for example, a neuronal-specific promoter, a glial-specific promoter, a muscle cell-specific promoter, a cardiac cell-specific promoter, a kidney cell-specific promoter, a bone cell-specific promoter, an endothelial cell-specific promoter, or an immune cell-specific promoter (e.g., a B cell promoter or a T cell promoter).

[0040] Developmentally-regulated promoters include, for example, promoters that are active only during the embryonic stage of development or only in adult cells.

[0041] "Operable linkage" or "operably linked" includes proximity of two or more components (e.g., a promoter and another sequence element) that permits the possibility that both components can function normally and that at least one of the components can mediate a function exerted by at least one of the other components. For example, a promoter may be operably linked to a coding sequence if it regulates the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include sequences that are contiguous or trans-acting with each other (e.g., regulatory sequences can act at a distance to regulate transcription of a coding sequence).

[0042] "Complementarity" of a nucleic acid means that a nucleotide sequence in one strand of a nucleic acid forms hydrogen bonds with another sequence on an opposing nucleic acid strand due to the orientation of its nucleobase groups. Complementary bases in DNA are generally A-T and C-G. In RNA, complementary bases are generally C-G and U-A. Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which every base in the duplex binds with a complementary base through Watson-Crick pairing. "Substantial" or "sufficient" complementarity means that the sequence in one strand is not completely and / or completely complementary to the sequence in the opposing strand, but sufficient binding occurs between the bases on the two strands under a set of hybridization conditions (e.g., salt concentration and temperature) to form a stable hybrid complex. Such conditions can be predicted by using sequences and standard mathematical calculations to predict the Tm (melting temperature) of hybridized strands, or by empirically determining the Tm using routine methods. The Tm comprises the temperature at which a population of hybridization complexes formed between two nucleic acid strands becomes 50% denatured (i.e., the population of double-stranded nucleic acid molecules becomes half-dissociated into single strands). Temperatures below the Tm favor the formation of hybridization complexes, while temperatures above the Tm favor melting or separation of the strands within the hybridization complexes. For a nucleic acid of known G+C content in aqueous 1 M NaCl solution, the Tm can be estimated, for example, by using Tm = 81.5 + 0.41 (% G+C), although other known Tm computations take into account structural properties of nucleic acids.

[0043] "Hybridization conditions" include the cumulative environment in which one nucleic acid strand binds with a second nucleic acid strand through complementary interactions and hydrogen bonding to form a hybridization complex. Such conditions include the chemical components of the aqueous or organic solution containing the nucleic acid and their concentrations (e.g., salts, chelating agents, formamide), as well as the temperature of the mixture. Other factors, such as the length of incubation time or the dimensions of the reaction chamber, may contribute to the environment. See, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed., pp. 1.90-1.91, 9.47-9.51, 11.47-11.57 (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is incorporated herein by reference in its entirety for all purposes.

[0044] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. Appropriate conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, well-known variables. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) of hybrids of nucleic acids having those sequences. For hybridization between nucleic acids having short stretches of complementarity (e.g., complementarity over 35 nucleotides or less, 30 nucleotides or less, 25 nucleotides or less, 22 nucleotides or less, 20 nucleotides or less, or 18 nucleotides or less), the position of mismatches becomes important (see Sambrook et al., supra, pp. 11.7-11.8). Generally, the length for a hybridizable nucleic acid is at least about 10 nucleotides. Exemplary minimum lengths for hybridizable nucleic acids include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Additionally, the temperature and salt concentration of the wash solutions can be adjusted as needed, depending on factors such as the length of the complementary region and the degree of complementarity.

[0045] The sequence of a polynucleotide does not need to be 100% complementary to the polynucleotide sequence of its target nucleic acid to be specifically hybridizable. Furthermore, a polynucleotide can hybridize across one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop or hairpin structure). A polynucleotide (e.g., a gRNA) can contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to a target region within a targeted nucleic acid sequence. For example, a gRNA in which 18 of 20 nucleotides are complementary to a target region and therefore specifically hybridizes thereto would be 90% complementary. In this example, the remaining non-complementary nucleotides can be clustered or interspersed with complementary nucleotides and do not need to be contiguous with each other or with complementary nucleotides.

[0046] The percent complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using the BLAST program (basic local alignment search tool) and the PowerBLAST program (Altschul et al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res. 7:649-656), or using the algorithm of Smith and Waterman (Adv. Appl. Math. 1981, 2:482-489), using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University of Wisconsin Research) using default settings. The concentration of hydroxybenzoates can be routinely determined by using a hydroxybenzoate analyzer (H.P. Park, Madison Wis.).

[0047] "Sequence identity" or "identity," in the context of two polynucleotide or polypeptide sequences, refers to the residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When percentages of sequence identity are used with respect to proteins, non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), thereby not altering the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Methods for making this adjustment are well known. Generally, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, where identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of zero, conservative substitutions are assigned a score between zero and 1. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0048] "Percentage of sequence identity" includes a value determined by comparing two optimally aligned sequences (maximum number of perfectly matched residues) over a comparison window, where a portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue occurs to obtain the number of matching positions, dividing the number of matching positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes a linked heterologous sequence), the comparison window is the entire length of the shorter of the two sequences being compared.

[0049] Unless otherwise specified, sequence identity / similarity values ​​include values ​​obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a gap weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a gap weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program thereof. "Equivalent program" includes any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and the same percent sequence identity for any two sequences at issue compared to corresponding alignments produced by GAP version 10.

[0050] "Conservative amino acid substitution" refers to the substitution of an amino acid normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another polar (hydrophilic) residue, such as the substitution between arginine and lysine, the substitution between glutamine and asparagine, or the substitution between glycine and serine. Furthermore, the substitution of a basic residue such as lysine, arginine, or histidine for another basic residue, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another acidic residue, are further examples of conservative substitutions. Examples of non-conservative substitutions include substitution of a polar (hydrophilic) residue such as cysteine, glutamine, glutamic acid or lysine with a nonpolar (hydrophobic) amino acid residue such as isoleucine, valine, leucine, alanine, or methionine and / or substitution of a nonpolar residue with a polar residue.

[0051] "Site-specific recombination" refers to a recombination event achieved by cognate site-specific recombinases between two compatible sequence-specific recombination sites on a single nucleic acid molecule. "Site-specific recombinase" refers to an enzyme that mediates site-specific recombination. Examples of site-specific recombinases include, but are not limited to, Cre recombinase and FLP recombinase. Bacteriophage P1 Cre recombinase recognizes lox recombination sites, while yeast FLP recombinase recognizes FRT recombination sites.

[0052] "Cre recombinase" refers to a protein capable of mediating the site-specific recombinase activity of the Cre protein of bacteriophage P1 (see Hamilton, DL, et al., J. Mol. Biol., 178:481-486 (1984)). The Cre protein of bacteriophage P1 mediates site-specific recombination between specific recombination sequences known as "loxP" sequences (see, e.g., Hoess et al., Proc. Natl. Acad. Sci. USA, 79:3398-3402 (1982)). The term encompasses any derivative of the naturally occurring Cre recombinase that retains the ability to effect recombination between two compatible loxP sites.

[0053] The recombination products depend on the position and relative orientation of the lox sites. When two lox sequences with the same orientation are present in the same DNA molecule, the DNA sequence flanked by the two lox sequences is excised by Cre recombinase to form a circular molecule (excision reaction). On the other hand, when two lox sequences are present in different DNA molecules, one of them becomes a circular DNA, which is inserted into the other DNA molecule via the lox sequence (insertion reaction).

[0054] "FLP recombinase" refers to a protein with yeast FLP (Saccharomyces cerevisiae) site-specific recombinase activity. The FLP protein, encoded by yeast 2μ DNA, mediates site-specific recombination between a pair of compatible FRT recombination sites (see Babineau et al., J. Biol. Chem., 260, 12313-12319 (1985)). The term encompasses any derivative of the naturally occurring FLP recombinase that retains the ability to achieve recombination between two compatible FRT sites. The recombination product depends on the position and relative orientation of the FRT sites. When two FRT sequences with identical orientations are present in the same DNA molecule, the DNA sequences flanking the two FRT sequences are excised by the FLP recombinase to form a circular molecule (excision reaction). On the other hand, when two FRT sequences are present in different DNA molecules, one of which is a circular DNA, the circular DNA is inserted into the other DNA molecule via the FRT sequence (intercalation reaction).

[0055] An "adeno-associated virus" is a replication-deficient parvovirus that can replicate only in cells in which certain viral functions are provided by a coinfecting helper virus, such as adenovirus, herpesvirus, and, in some cases, a poxvirus such as vaccinia. An "adeno-associated virus vector" or "AAV" is a vector that can replicate in any cell line, such as of human, monkey, or rodent origin, equipped with the appropriate helper virus functions (see Berns and Bohensky, Advances in Virus Research, Academic Press., 32:243-306 (1987)).

[0056]

[0013] The present invention provides a method for knocking in a human long nucleic acid sequence with high efficiency, more specifically, a method for knocking in a full-length human orthologous gene into a corresponding locus in a human mammal, the method comprising the steps of: (a) preparing a first circular vector targeting the full-length human orthologous gene in a non-human mammal orthologous locus, the first circular vector comprising a gene encoding a drug resistance gene, and at its 5'-end and 3'-end, respectively, non-human mammal homologous arm sequences flanking the full-length non-human mammal orthologous gene, and further comprising human homologous arm sequences flanking each of the non-human mammal homologous arm sequences; (b) knocking in a gene encoding a drug resistance gene flanked by human homologous sequences into the full-length orthologous locus in a non-human mammal cell by homologous recombination using the first circular vector prepared in step (a); (c) preparing a second circular vector comprising the full-length human orthologous gene and human homologous arm sequences at its 5'-end and 3'-end, respectively; (d) knocking in genes encoding the full-length human orthologous gene and the drug resistance gene by homologous recombination into the knock-in cells obtained in step (b) using the second circular vector prepared in step (c).

[0057] (1) Human Orthologous Genes to be Knocked Into Non-Human Mammals Generally, "ortholog" is intended to mean a gene that is derived from a common ancestral gene through speciation and is found in different species as a result of speciation. Genes found in different species are considered orthologs if their nucleotide sequences and / or the protein sequences they encode share at least about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more sequence identity. It is well known that the functions of orthologs are well conserved between species.

[0058] According to the present invention, the human orthologous gene to be knocked into the corresponding locus in a non-human mammal is not particularly limited. For example, information on "human orthologous genes" is available, for example, from Evola (http: / / hinv.jp / evola). Evola is a database that provides orthologs and gene family information between humans and other vertebrates (primates, fish, etc.). Evola identifies orthologs by (1) creating genome alignments between humans and other organisms and comparing the genes on the alignments at the exon level (splice variant level), (2) comparing the amino acid sequences, and (3) comparing the gene and species phylogenetic trees (Manual Curation). Furthermore, gene families (overlapping genes) are identified based on homology (identity or similarity) between human genes. Evola's gene family browser allows for interspecies comparison of orthologs and paralogs (genes resulting from duplication) for each gene family between humans and four other organisms. It is possible to view the chromosomal distribution of these genes and examine the genomic region of each gene. Evola 7.5 (released December 2010) provides ortholog information for 22,768 human genes in 14 organisms.

[0059] In Example 6 described below, the KIT gene located on the long arm of human chromosome 4 (4q12) was used as a typical example of the human orthologous locus to be introduced into the mouse orthologous locus. The term "KIT" refers to a human tyrosine kinase sometimes referred to as the obesity / stem cell growth factor receptor (SCFR), the proto-oncogene c-KIT, the tyrosine protein kinase Kit, or CD117. The term "KIT gene" is used to refer to a gene encoding a polypeptide having KIT kinase activity, for example, the sequence of which is located at nucleotides 54,657,957 and 54,740,715 on chromosome 4 of the human genome reference hg19. The term "KIT transcript" refers to a transcription product of the KIT gene, an example of which has the sequence of NCBI Reference Sequence NC 000004.12.

[0060] Examples of "non-human mammals" into which a human orthologous gene is knocked in include, but are not limited to, mice, rats, rabbits, monkeys, dogs, cats, pigs, and cows, with mice and rats being preferred.

[0061] According to the present invention, when introducing a human orthologous gene into a non-human mammalian cell (hereinafter sometimes simply referred to as a "host cell"), it is preferable to use a technique based on homologous recombination using "homologous arms." When using homologous recombination for such targeted integration (see Doetschman, T., et al., Nature 330 (1987) 576-578; Thomas, KR and Capecchi, MR, Cell 51 (1987) 503-512; Thompson, S., et al., Cell 56 (1989) 313-321; Zijlstra, M., et al., Nature 342 (1989) 435-438; Bouabe, H. and Okkenhaug, K., Meth. Mol. Biol. 1064 (2013) 315-336), the nucleic acid sequence for recombination is a sequence homologous to the exogenous nucleic acid sequence and is referred to herein as a "homologous arm." In this case, the deoxyribonucleic acid introduced into the host cell contains, as a first recombination sequence, a sequence 5' (upstream) homologous to the exogenous nucleic acid sequence (i.e., landing site), and, as a second recombination sequence, a sequence 3' (downstream) homologous to the exogenous nucleic acid sequence. Generally, the frequency of targeted integration increases with the length and isogenicity of the homologous arms. Ideally, the homologous arms are derived from genomic DNA prepared from the host cell. The length of the homologous arms to be used or targeted is not limited, but is preferably 1 kbp or longer.

[0062] According to the present invention, the plasmid used to knock-in a human orthologous locus into a non-human mammalian orthologous locus is preferably a circular plasmid. As described below, the present invention is characterized by knocking-in a human orthologous locus in two steps. In both steps, a circular plasmid is preferably used for introducing a foreign gene into the locus.

[0063] (2) Knock-in method of human orthologous genes into non-human mammals The knock-in method of the present invention is mainly characterized by integrating a human orthologous gene locus into a predetermined locus of a host cell in two stages. The circular vector used in each stage can be prepared based on a known knock-in vector. Such vectors include, but are not limited to, plasmid vectors, cosmid vectors, bacterial artificial chromosome (BAC) vectors, yeast artificial chromosome (YAC) vectors, viral vectors (e.g., retroviral vectors, lentiviral vectors, adeno-associated viral vectors, adenoviral vectors), etc. In this specification, the vectors used in each stage are distinguished and referred to as a "first circular vector" and a "second circular vector," respectively.

[0064] The "first circular vector" is a vector that targets a full-length non-human mammal orthologous gene at a full-length non-human mammal orthologous gene locus, and preferably comprises a gene encoding a drug resistance gene, non-human mammal homologous arm sequences adjacent to the full-length non-human mammal orthologous gene at its 5'-end and 3'-end, respectively, and further comprises human homologous arm sequences adjacent to each of the non-human mammal homologous arm sequences. The arrangement of each gene is not limited, but a typical example of a transgene introduced into a circular vector is 5'-(non-human mammal homologous arm sequence)-(human homologous arm sequence)-(drug resistance gene)-(human homologous arm sequence)-(non-human mammal homologous arm sequence)-3'.

[0065] The "second circular vector" is a vector containing a full-length human orthologous gene and human homologous arm sequences at its 5' and 3' ends, respectively. The arrangement of each gene is not limited, but a typical example of a transgene introduced into a circular vector is 5'-(human homologous arm sequence)-(human gene sequence)-(human homologous arm sequence)-3'. However, a drug resistance gene may be included before or after the human gene sequence. Since the second circular vector is intended to incorporate a long gene, a BAC vector is preferably used. As known to those skilled in the art, a "BAC vector" is a plasmid constructed using the F plasmid of E. coli and is capable of stably maintaining and propagating large DNA fragments of approximately 300 kb or more in bacteria such as E. coli. According to the present invention, a full-length human orthologous gene of 10 kbp or more can be introduced into non-human mammalian cells using the knock-in method of the present invention. The length of the human orthologous full-length gene is at least 10 kbp, but is preferably 20 kbp, 30 kbp, 40 kbp, 50 kbp, 60 kbp, 70 kbp, 80 kbp, 90 kbp, 100 kbp or more, more preferably 100 kbp, and even more preferably 200 kbp or more.

[0066] As used herein, the term "drug resistance gene" refers to a gene that can confer resistance to drugs such as antibiotics to cells that express it, and any drug resistance gene can be used in the knock-in method of the present invention. Examples of drug resistance genes include, but are not limited to, neomycin resistance genes, puromycin resistance genes, blasticidin resistance genes, hygromycin resistance genes, chloramphenicol resistance genes, tetracycline resistance genes, erythromycin resistance genes, spectinomycin resistance genes, kanamycin resistance genes, zeocin resistance genes, and phleomycin resistance genes, and these drug resistance genes are well known to those skilled in the art.

[0067] The knock-in method of the present invention is characterized by using the above-mentioned two circular vectors in two steps to introduce (replace) a human orthologous gene locus into a corresponding gene locus in a host cell. In the first knock-in step, the corresponding gene locus can be replaced with "5'-(human homologous arm sequence)-(drug resistance gene)-(human homologous arm sequence)-3'" by homologous recombination using homologous arms ("non-human mammal homologous arm sequences") adjacent to the gene in the host cell. Subsequently, in the second knock-in step, the human gene locus can be introduced by homologous recombination using the two human homologous arm sequences already incorporated into the host cell.

[0068] The method used to introduce a human orthologous gene locus into a corresponding orthologous gene agent in a host cell using the knock-in method of the present invention is preferably, but not limited to, genome editing, which is a method for modifying a target gene using a site-specific nuclease (e.g., a DNA double-strand cleavage enzyme such as zinc finger nuclease (ZFN), transcription activation-like effector nuclease (TALEN), or CRISPR-Cas9 (Csn1)). For example, fusion proteins such as ZFN (U.S. Patent Nos. 6,265,196, 8,524,500, 7,888,121, European Patent No. 1,720,995), TALEN (U.S. Patent Nos. 8,470,973, 8,586,363), and PPR (pentatricopeptide repeat) fused with a nuclease domain (Nakamura et al., Plant Cell Physiol 53:1171-1179(2012)), CRISPR-Cas9 (U.S. Patent No. 8,697,359, International Publication No. 2013 / 176772), CRISPR-Cas12a (Cpf1) (Zetsche B. et al., Cell, 163(3):759-71, (2015)), and Target-AID (K. Nishida et al., Targeted nucleotide editing using hybrid Examples of such methods include a method using a complex of guide RNA and a protein, such as the Cas9 and Cas12a described above, as well as Cas12b (C2c1), Cas13a (C2c2), Cas13b (C2c6), and Cas13c (C2c7).

[0069] The gene encoding the nuclease can be delivered to cells by plasmid DNA, viral vector, or in vitro transcribed mRNA. Transfection of the plasmid DNA or mRNA can be carried out by electroporation or cationic lipid-based reagents. For example, AAV vectors can be used to deliver the nuclease.

[0070] "CRISPR" refers to a sequence-specific genetic engineering technique that relies on the clustered regularly interspaced short palindromic repeats pathway. CRISPR can be used to perform gene editing and gene regulation, as well as simply targeting proteins to specific genomic locations. "Gene editing" refers to a type of genetic engineering in which the nucleotide sequence of a target polynucleotide is modified through the introduction of deletions, insertions, single- or double-strand breaks, or base substitutions into the polynucleotide sequence. In some embodiments, CRISPR-mediated gene editing uses non-homologous end joining (NHEJ) or homologous recombination pathways to perform editing. Gene regulation refers to increasing or decreasing the production of specific gene products, such as proteins or RNA.

[0071] As used herein, "gRNA" or "guide RNA" refers to a guide RNA sequence used to target a specific polynucleotide sequence for gene editing using CRISPR technology. The gRNA comprises a fusion polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA); or a polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA); or consists essentially of a fusion polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA); or a polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA); or consists of a fusion polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA); or a polynucleotide comprising CRISPR RNA (crRNA) and a trans-activated CRISPR RNA (tracrRNA).

[0072] "Cas9" refers to the CRISPR-associated endonuclease represented by this name. Non-limiting examples of Cas9 include Staphylococcus aureus Cas9, nuclease-dead Cas9, and their respective orthologs and biological equivalents. Generally, wild-type endonuclease CRISPR-Cas9 proteins induce site-specific double-strand breaks, thereby inactivating genes by non-homologous end joining, which is useful for genome editing, or introducing heterologous genes by homologous recombination.

[0073] The present invention will be described in more detail below with reference to examples. However, the present invention is not limited to the following examples, and can be implemented in any form within the scope of the present invention.

[0074] (1) Animals and Ethics. C57BL / 6J or BALB / c mice were purchased from Clea-Japan (Tokyo, Japan), and ICR mice were purchased from Japan SLC (Shizuoka, Japan) or Clea-Japan. Mice were housed in a disease-free environment in the Laboratory Animal Facility of the Institute of Medical Science, The University of Tokyo, with free access to food and water and under a 12-hour light / 12-hour dark photoperiod. All mouse experiments were approved by the University of Tokyo Animal Care and Use Committee (approval number PA17-63) and conducted in accordance with the ARRIVE guidelines (https: / / arriveguidelines.org).

[0075] (2) Mouse Embryonic Fibroblast Growth and ES Cell Culture To obtain mouse embryonic fibroblasts (MEFs), BALB / c and C57BL / 6J mice were mated. Pregnant females at E14.5 were euthanized by cervical dislocation. The collected fetuses were homogenized, and these cells were then cultured in Dulbecco's modified Eagle's medium (Nacalai, Kyoto, Japan) containing 10% (v / v) fetal bovine serum (FBS) (Sigma-Aldrick, St. Louis, MO, USA), 100 U / mL penicillin, and 100 μg / mL streptomycin (Fujifilm Wako Pure Chemical Industries, Osaka, Japan) under humidified 5% CO. 2MEFs were isolated by culturing in a medium at 37°C. After 12-14 days of culture, the proliferated MEFs were irradiated with X-rays (50 Gy) to arrest the cell cycle, and the mitotically inactivated MEFs were used for ES cell culture under feeder culture conditions. ES cells with a C57BL / 6J or BALB / c background were developed by the inventors with some modifications (Yagi, M., et al., Nature, 548, 224-227 (2017)). Briefly, mated females were euthanized by cervical dislocation, and embryos were collected by flushing the oviduct the day after observation of nulliparous oocyte plugs. The collected embryos were cultured in simple potassium-optimized medium (KSOM) (Merck-Millipore, Darmstadt, Germany) for 2 days to produce blastocysts. Each blastocyst was cultured with irradiated MEFs (1.2 × 10 5 / cm 2 The cells were transferred to a 24-well culture plate (BD Biosciences, Bedford, MA, USA) containing ES culture medium (ESCM; Thermo Fisher Scientific, Waltham, MA, USA), 15% (v / v) FBS (Thermo Fisher Scientific), 100 U / mL penicillin (Fujifilm Wako Pure Chemical Industries, Ltd.), 0.1 mM 2-mercaptoethanol (Thermo Fisher Scientific), leukemia inhibitory factor and t2i (0.2 μM PD0325901 (Sigma-Aldrich), and 3 μM CHIR99021 (Axon The blastocysts were cultured in ESCM (Medchem, Groningen, Netherlands). Five to seven days after blastocyst seeding, the inner cell mass outgrowth was digested into single cells by treatment with 0.25% (w / v) trypsin-ethylenediaminetetraacetic acid (EDTA) (Nacalai), then passaged onto fresh MEFs and cultured in ESCM to generate self-renewing ES cells. In addition to these in-house developed ES cell lines, JM8.A3 (C57BL / 6N) or V6.5 (B6-129 F1) ES cell lines were also used in the experiments. Each ES cell line was passaged at a 1:10 dilution every 2 to 3 days. For feeder-free ES cell culture, ES cells were cultured in ESCM using dishes precoated with 0.1% (w / v) gelatin solution (Sigma-Aldrich).

[0076] (3) Plasmids. The targeting vector pAAV-mRosa26-CAG-EFGP (pR26-CE), which has a 5' homologous arm (962 bp) and a 3' homologous arm (1,006 bp) to Rosa26 or pAAV-CAG-EGFP (pCE), was kindly provided by Dr. Mizuno. To generate the pRosa26-CAG-EGFP-PGK-NeoR plasmid (pR26-CE-PN), a DNA fragment encoding PGK-NeoR-bGHpA was PCR-amplified using pL452 as a template and cloned between the EGFP and 3' homologous arm sequences of pR26-CE using an In-Fusion HD Cloning Kit (TaKaRa, Shiga, Japan). pNanog-T2A-mCherry (pNmC) or pStra8-CreERT2-EF1-copGFP2aPuro (pStra8-CE-E1GP) was constructed using pUC19 as a backbone. To generate pNmC, the 5' and 3' homologous arms of the Nanog locus were PCR-amplified using the C57BL / 6J genome as a template. Each homologous arm and T2A-mCherry-bGHpA were cloned into pUC19 using the In-Fusion HD Cloning Kit (TaKaRa). To prepare pStra8-CE-E1GP, DNA encoding CreERT2-rabbit globin poly-A was purchased from Genewiz (Genewiz Japan, Saitama, Japan). The 5' and 3' homology arms of the Stra8 locus and the EF1-copGFP2aPuro DNA fragment were amplified by PCR using the C57BL / 6 J genome or PB513 (System Biosciences, Palo Alto, CA, USA) as templates, respectively, and each fragment was cloned into pUC19 using the In-Fusion HD Cloning Kit (TaKaRa).To generate pAAV-mRosa26-CAG-loxP-Neo-loxP-RFP (pR26-lsl-RFP), the loxP-Neo-loxP fragment or the RFP fragment was PCR-amplified using pROSA26-DEST (#21189, Addgene) or PB514 (System Biosciences), respectively, as a template and cloned between the CAG promoter and 3′ homologous arm of pR26-CE using the In-Fusion HD Cloning Kit (Takara). To generate pCd6-CAG-CreERT2-EF1-copGFP2aPuro (pCd6-CE-E1GP), the homologous arms of the Cd6 locus or the EF1-copGFP2aPuro fragment were PCR-amplified using genomic DNA or PB514 (System Biosciences), respectively, as templates. The CreERT2 fragment was purchased from Genewiz (Genewiz Japan). All DNA fragments were cloned into pAAV-MCS2 (#46954, Addgene) using the In-Fusion HD Cloning Kit (Takara) to generate pCd6-CE-E1GP. All-in-one plasmids expressing Cas9 mRNA and guide RNA (gRNA) targeting specific genes of interest were prepared by cloning double-stranded oligos into the BbsI site of pX459 (Addgene, #48139). The gRNA targets (each 20 nucleotides long) were as follows: Rosa26, 5'-AAGGGATTCTCCCAGGCCCA-3' (SEQ ID NO: 1); tyrosinase, 5'-GGTCATCCACCCCTTTGAAG-3' (SEQ ID NO: 2); Nanog, 5'-TATGAGACTTACGCAACATC-3' (SEQ ID NO: 3); Stra8, 5'-TAGATTATAATGGCCACCCC-3' (SEQ ID NO: 4); and Cd6, 5'-ACAAGTTGGGAAAGGTTTAT-3' (SEQ ID NO: 5). A summary of the targeting vectors used in this study is shown in Table 1.

[0077]

[0078] Preparation of CRISPR-Cas9 Ribonucleoprotein Complexes (RNPs) and Electroporation of ES Cells. TracrRNA, crRNA, and Cas9 protein were purchased from IDT (Coralville, IA, USA). TracrRNA and crRNA were dissolved in Duplex Buffer (IDT, 200 μM each) and annealed in a thermal cycler at 95°C for 10 minutes, followed by a step-down cycle of -1°C / min to 25°C. The annealed RNA (100 μM) was then incubated with 3 μg / μL Cas9 protein at 37°C for 20 minutes to form Cas9-RNPs. Plasmids and Cas9-RNPs were electroporated into ES cells using the Neon transfection system (MPK5000, ThermoFisher). Briefly, 1 x 10 cells with 1 μg of targeting vector 5 ES cells were electroporated with or without the all-in-one CRISPR-Cas9 vector (0.5 μg) or Cas9-RNP (at a final concentration of 10 μM annealed RNA and 0.3 μg / μL Cas9 protein) in a 10 μL tip (MPK1096, Thermo). The Neon system used two pulses of 1200 V and 20 ms, or a single pulse of 1400 V and 30 ms, to perform all-in-one vector transfection or Cas9-RNP transfection, respectively. Electroplated ES cells were cultured in ESCM with or without MEFs.

[0079] Drug selection and in vitro induction of tamoxifen-induced Cre-loxP recombination. For drug selection using puromycin or blasticidin resistance genes, electroporated ES cells were treated with G418 (400 ng / mL), puromycin (0.5 μg / mL), and / or blasticidin (20 μg / mL) to develop stable transfectants. For tamoxifen-mediated in vitro Cre / loxP recombination, ES cell clones were treated with 4-hydroxytamoxifen (4OHT, 1 μM) for 2 days, and fluorescent reporter expression was observed using a fluorescence microscope (BZ-X710, Keyence).

[0080] For flow cytometry analysis to detect fluorescent reporter expression, ES cell colonies were digested into single cells with 0.25% (w / v) trypsin-EDTA, and the single cells were resuspended in PBS, 2 mM EDTA, and 1% (w / v) BSA. Fluorescence expression was monitored using a FACSCalibur system (BD Biosciences, San Jose, CA, USA), and collected data were analyzed using FlowJo software (BD Biosciences).

[0081] Genotyping and indel mutation rate determination. The KI genotype of each ES cell clone was determined by PCR using genomic DNA as a template. Single ES cell colonies were manually selected as clones and transferred to new culture dishes. The expanded ES cell clones were then harvested and lysed at 65°C for at least 2 hours using Tail Lysis Buffer (Nacalai) containing 7 U / mL proteinase K (#9034, Takara). Genomic DNA was extracted from the crude lysate and purified by conventional phenol / chloroform / isoamyl alcohol treatment (all Nakarai) followed by ethanol precipitation for PCR analysis. KODONE polymerase (Toyobo, Osaka, Japan) was used for DNA fragment amplification. The indel mutation rate was measured by T7 endonuclease I mismatch cleavage assay using the Alt-R Genome Editing Detection Kit (IDT). Indel mutation ratios were calculated as described (http: / / crispr.technology / resources / quantification.html). Table 2 shows the primer sequences used for KI genotyping.

[0082]

[0083] Chimeric mice were developed by blastocyst injection of gene-targeted ES cells. Briefly, 8-week-old female mice were superovulated by intraperitoneal administration of 7.5 U equine chorionic gonadotropin (Serotropin, ASKA Animal Health, Tokyo, Japan) and 7.5 U human chorionic gonadotropin (Gonatropin, ASKA Animal Health) at 48-hour intervals. The mice were then mated with adult males of the same strain to obtain embryos. The mated females were euthanized the day after vaginal plug observation, and two-cell embryos were collected by oviductal irrigation with modified Whitten medium (mWM). The collected embryos were cultured in KSOM (Merck-Millipore) for 2 days to develop into blastocysts. Six to ten ES cells were injected into a single blastocyst, and the injected blastocysts were transferred to the uterus of pseudopregnant ICR female surrogates (20 to 25 blastocysts per female) 2.5 days after mating to obtain chimeric offspring. Blastocysts with an ICR background (albino) were used to generate chimeras using B6J (black), B6N (JM8.A3, Agouti), or B6-129 F1 (V6.5, Agouti). For BALB / c (albino) ES cell injection, blastocysts with a B6J background (black) were used. For in vivo tamoxifen-induced Cre / loxP recombination, 5-week-old chimeric mice were intraperitoneally injected with tamoxifen (Sigma-Aldrich, 2 mg / body diluted in corn oil) or corn oil alone (control) once daily for 5 times, and were harvested for tissue collection 3 days after the final tamoxifen administration. The harvested tissues were observed for fluorescence expression analysis using a fluorescent stereomicroscope (Leica Microsystems, Tokyo, Japan).

[0084] Statistical analysis: All numerical data are presented as the mean ± SEM of three independent replicates. Differences between treatments or genotypes were tested using a two-tailed Student's t-test. A P value of <0.05 was considered significant.

[0085] The experimental methods used in the following examples are described in a paper previously submitted by the present inventors (Ozawa, M., wt al., Scientific Reports, (2022) 12:21558, https: / / doi.org / 10.1038 / s41598-022-26107-z), which can be referenced as needed.

[0086] Example 1: Integration of Linear DNA Fragments Occurs Primarily Randomly in the Embryonic Stem Cell Genome The Rosa26 locus is known as a safe harbor locus for stable gene expression and is frequently used to generate knock-in (KI) mice carrying a fluorescein reporter or Cre / loxP conditional recombination. Thus, we used this locus to determine KI based on homologous recombination using large donor DNA. Because linear plasmids are routinely used as DNA donors for conventional gene targeting, we first compared the frequency of DNA integration into the genome between linear and circular vectors. ES cells were transduced via electroporation with linear or circular plasmids (5' and 3' ends, respectively) carrying an 1,850-bp CAG-EGFP cassette flanked on either side of the homologous arms (962 bp and 1,006 bp 5' and 3' ends, respectively) of the Rosa26 locus (pR26-CE). After 7-10 days, EGFP expression was measured via microscopy or flow cytometry. The number of EGFP-positive cells was very low with both transductions (0.7±0.1% and 0.4±0.1% for linear and circular vectors, respectively), but the percentage of EGFP-positive cells significantly increased when the linear vector was used (data not shown).

[0087] Next, we selected ES cells that had integrated the linear vector into their genome via drug resistance to verify whether the vector integration was a homology arm-dependent site-specific KI. ES cells were transduced with the linearized vector pR26-CE-PN, which contained a PGK-NeoR cassette subcloned at the 3' end of CAG-EGFP in pR26-CE. 24 hours after electroporation, G418 was added to the medium, and the cells were cultured for an additional 7 to 10 days to select stable drug-resistant clones. After G418 selection, almost all ES cell colonies were EGFP-positive (data not shown). Analysis of individual colonies for Rosa26 locus-specific KI via genomic PCR showed that none of the 47 independent clones contained either the 5' or 3' KI-specific bands. These results indicated that integration of linear targeting vectors into the genome was mainly random and not site-specific, whereas circular plasmids, which carry homologous arms and are thought to have a lower frequency of random genome integration than linear plasmids, were used as targeting vectors below.

[0088] Example 2: CRISPR-Cas9 ribonucleoprotein (RNP)-mediated integration of circular plasmids into ES cell genomes was highly locus-specific. Recent studies have shown that Cas9-RNP-mediated genome editing exhibits lower cytotoxicity and higher genome rearrangement induction efficiency than that mediated by CRISPR-Cas9 expression plasmids (Seidler, B., et al., Proc. Natl. Acad. Sci., 105, 10137-10142 (2008)). Therefore, we next compared the efficiency of gene KI in ES cells with that of an all-in-one plasmid expressing both gRNA and Cas9 mRNA, or Cas9-RNP consisting of crRNA / tracrRNA / Cas9 protein. We transduced ES cells harboring the circular pR26-CE by electroporation with the CRISPR-Cas9 all-in-one plasmid or Cas9-RNP (Fig. 1A, B) and measured EGFP expression 7–10 days later. In this experiment, no drug resistance cassette was used for positive or negative selection. When ES cells were transduced with pR26-CE alone, almost no EGFP-positive ES cells emerged (Fig. 1C, D). Introduction of pR26-CE and the all-in-one vector significantly increased the frequency of EGFP-positive ES cells, but the percentage was low (2.7 ± 0.1%, Fig. 1D, E). In contrast, when ES cells were transduced with pR26-CE and Cas9-RNP, the percentage of EGFP-positive cells increased approximately 10-fold compared to the all-in-one plasmid (27.7 ± 0.7%, Fig. 1D, E). Next, ES cells that became EGFP-positive were sorted by FACS, transferred to fresh medium, and cloned colonies were selected and used for KI genotyping by PCR. Of the selected clones, 41 out of 47 were positive for both the 5' and 3' KI bands (Figure 1F), with a KI ratio of 87.2% (Figure 1G). These results indicated that Cas9-RNP-mediated double-stranded DNA KI at the Rosa26 locus of ES cells was more pronounced than that found with the all-in-one plasmid. Next, the Rosa26 locus was amplified by PCR to determine whether this plasmid KI represented a homozygous or heterozygous integration event, and 11 clones were analyzed.One of these clones lacked a PCR amplicon, while the remaining 10 clones contained indel mutations (data not shown), suggesting that most cases of Rosa26 KI are heterologous integration KI events, but that genetic rearrangements frequently occur on both chromosomes after Cas9-RNP induction.

[0089] Example 3: Cas9 RNP-mediated plasmid DNA KI is applicable to various loci in ES cells. Next, we determined whether Cas9-RNP-mediated circular plasmid KI results from homologous arm-mediated homologous recombination or Cas9-RNP-induced double-strand break site-specific recombination. First, Cas9-RNPs recognizing either the Rosa26 locus (gRosa26-RNP) or the tyrosinase locus (gTyr-RNP) were transduced into ES cells using pR26-CE as the KI donor by electroporation (Figure 2A). The indel mutation efficiency of gTyr-RNP via the T7 endonuclease 1 (T7E1) assay was an average of 90.5% (n = 3 for each replicate, 88.6%, 90.4%, or 92.5%) (Figure 2B). Seven to 10 days after electroporation, EGFP-positive ES cells were identified using microscopy (Figure 2C) or flow cytometry (Figure 2D). The percentage of EGFP-positive ES cells in the pR26-CE and gRosa26-RNP transfection group was 23.8 ± 3.2% (Figure 2E). In contrast, when pR26-CE was transduced with gTyr-RNP, EGFP-positive cells rarely appeared (Figure 2E). Next, ES cells were transduced with gRosa26-RNP and CAG-EGFP plasmids flanked by Rosa26 homology arms, with or without pR26-CE, and GFP expression was measured by flow cytometry 7 to 10 days after electroporation (Figure 2F). The percentage of EGFP-positive ES cells in the pR26-CE and gRosa26-RNP transduced groups was 19.2 ± 1.2%, whereas EGFP-positive cells rarely appeared when gRosa26-RNP was transduced with homology armless pCE (0.68 ± 0.1%) (Figure 2G, H). These results suggested that circular plasmid DNA was integrated into the genome via homology arm-mediated homologous recombination. Rosa26 is known to form stable open chromatin in virtually all tissues, including ES cells; therefore, the integration efficiency is likely to be relatively higher than that of other genomic loci in mice. Next, we investigated whether this CRISPR-Cas9-mediated plasmid KI without drug selection could function efficiently at various loci in mouse ES cells.First, we targeted a promoterless T2A-mCherry cassette to the C-terminus of the Nanog gene, which is highly expressed in ES cells. Because this targeting vector lacks an exogenous promoter, mCherry fluorescent signals were only observed when the T2A-mCherry cassette was properly integrated at the Nanog locus due to endogenous promoter activity. Transduction of pNanog-T2A-mCherry (pNmC) and gNanog-RNP into ES cells resulted in 37.5% (18 / 48) of ES cell clones showing mCherry signals. Eight mCherry-positive clones were randomly selected and analyzed for integration by PCR. The results confirmed that all analyzed clones were integrated with mCherry at the Nanog locus (data not shown). Next, we targeted nine independent loci on six different chromosomes using various ES cell lines from C57BL / 6J (B6J), C57BL / 6N (B6N), BALB / c, or B6-129 F1 backgrounds, successfully developing precise KIs at all nine attempted loci. KI ratios ranged from 6.8 to 59.1% (Table 3). The expression levels of each endogenous gene at the KI loci in ES cells were analyzed using a public database (https: / / www.ebi.ac.uk / gxa / experiments / E-GEOD-27843 / Results) and are shown in Table 3. All ES cell lines used in this study were able to contribute to chimeric offspring after blastocyst injection and embryo transfer into surrogate mothers. These results suggest that the drug-selection-free plasmid KI method is applicable to various loci in ES cells.

[0090]

[0091] Example 4: KI efficiency was significantly increased with a targeting vector carrying a drug resistance gene cassette. For Rosa26 targeting (Figures 1E and 2E), the overall rate of accurate KI clones using the drug selection-free plasmid KI method was 20-30%, but the KI ratio among EGFP-positive ES cell clones was 87.2% (41 / 47) (Figures 1F and 1G), suggesting that the frequency of random integration of the targeting vector into the genome was relatively low. Therefore, we hypothesized that KI efficiency could be dramatically improved using a targeting vector incorporating a drug resistance cassette (Figure 3A). pR26-CE-PN, in which a PGK promoter-driven neomycin resistance cassette was subcloned downstream of the EGFP sequence of pR26-CE, was targeted to the Rosa26 locus with gRosa26-RNP (Figure 3B). Although some ES cell clones became EGFP-positive without G418 selection, as we hypothesized, almost all clones became EGFP-positive when ES cells were treated with G418 (Fig. 3C). PCR genotyping of all selected clones (47 clones) revealed positive results for both the 5' and 3' KI PCR bands (Fig. 3D, E).

[0092] Next, we investigated whether drug selection mediates enhanced KI efficiency at other genomic loci. We created another targeting vector carrying a promoterless CreERT2 vector, followed by the EF1 promoter-driven copGFP2aPuro vector, which is predicted to be KI at the Stra8 locus (pStra8-CE-E1GP) (Figure 3F). Stra8 is a germ cell-specific gene, and its expression in ES cells is quite low (TPM = 6, compared with 389 for well-known pluripotency genes such as Nanog and 1154 for Pou5f1; Table 1) (https: / / www.ebi.ac.uk / gxa / experiments / E-GEOD-27843 / Results). Introducing pStra8-CE-E1GP into ES cells that do not cleave the Stra8 locus (gStra8-RNP) did not produce stable ES cell clones after puromycin selection (Figure 3G). In contrast, 17.5 ± 3.0% of ES cells became GFP-positive after introducing pStra8-RNP, which combines pStra8-CE-E1EP and gStra8-RNP, and almost all colonies showed GFP signals after puromycin selection (Figure 3G). Site-specific KI was confirmed by PCR, and all selected clones showed KI PCR bands on both the 5' and 3' ends (46 independent clones) (Figures 3H, I). Eleven clones were randomly selected and further sequenced for their KI site. These transgenes contained a precisely in-frame KI at the Stra8 locus. These results demonstrated that the efficiency of CRISPR-Cas9-mediated plasmid KI can achieve a high KI ratio of up to 100% in ES cells by drug selection.

[0093] Example 5: Single electroporation could achieve dual-allelic KI in ES cells. Because KI efficiency can be increased using a plasmid KI with drug selection, we hypothesized that it might be possible to simultaneously obtain multiple alleles of KI by using targeting vectors with different drug selection cassettes. To test this hypothesis, we first targeted two different gene cassettes to the same locus, Rosa26, by single electroporation (Figure 4A). PCR genotyping after electroporation and subsequent G418 and blasticidin selection revealed that all selected clones (22 clones) contained KI in both cassettes (Figure 4B). We next determined whether a single electroporation could achieve dual KI at different genomic loci. To this end, a CAG-CreERT2-EF1-copGFP2aPuro cassette was targeted into ES cells by single electroporation into the Cd6 locus (pCd6-CE-E1GP) and a CAG-loxP-Neo-loxP-Red fluorescent protein (RFP) cassette into the Rosa26 locus (pR26-lsl-RFP) using gCd6-RNP and gRosa26-RNP (Figure 4C). The Cd6 locus is another safe harbor locus for broad and stable gene expression in various mouse tissues. After selection with G418 and puromycin for 7 days, 17 ES cell clones were selected for PCR analysis to determine the KIs of the Cd6 and Rosa26 loci. Surprisingly, 11 of the 17 ES cell clones showed both Cd6 and Rosa26 KI bands (64.7%, Figure 4D, E). Using the double KI ES clones, we demonstrated that the two KI alleles functioned in vitro, and RFP signals were only observed after 4OHT was added to the culture medium (Figure 4F, G). We further investigated whether these KI alleles functioned correctly in vivo. Chimeric mice were generated by injecting the ES cell clones into preimplantation blastocysts and then transferring the chimeric embryos into pseudo-regenerated surrogate mothers.Tamoxifen was administered intraperitoneally to chimeric mice 28 days after birth once daily for 5 days, and tissues were collected and subjected to microscopic analysis to evaluate the expression of the red fluorescent protein (Fig. 4H). RFP signals were observed in various tissues of the tamoxifen-treated chimeric mice (Fig. 4I), demonstrating that the integrated transgene functions as expected in vivo.

[0094] Discussion: We aimed to develop a simple and highly efficient KI method in mouse ES cells to efficiently develop genetically engineered mouse models. We found that the introduction of circular plasmid DNA as a targeting vector, in addition to the KI site-specific Cas9-RNP, enabled efficient DNA KI without drug selection. Furthermore, by incorporating a drug resistance gene cassette into the circular targeting vector, we were able to generate multiple KIs with extremely high efficiency by a single electroporation. The maximum size of the gene cassette used for KI in this example was 11.7 kbp without drug selection and 6.2 kbp with drug selection. The lengths of DNA fragments frequently used to generate genetically engineered mice, such as CreERT2, EGFP, or Cas9, are approximately 2.5, 1, or 4 kbp, respectively. Therefore, CRISPR-Cas9-mediated plasmid KI is universally applicable to ES cells and is expected to be highly beneficial for the efficient generation of genetically engineered mice.

[0095] Previous studies using mouse ES cells and fibroblasts have shown that introduction of plasmid DNA into cells results in very limited genomic integration. These studies are consistent with our data showing that circular plasmids, even those with homologous arms, rarely integrate into the genome of ES cells (Figure 1C, D). Therefore, linear DNA fragments with homologous arms are integrated into genomic loci by homology-dependent repair mechanisms and have traditionally been used for site-specific KI. However, while linearization of plasmid DNA significantly increases the frequency of genomic integration, the majority of genomic integration is nonspecific, even if the DNA contains homologous arms. Similar findings were observed in this study, where linear DNA fragments carrying drug resistance genes were introduced into ES cells in the absence of CRISPR-Cas9, and all ES cell clones that became drug-resistant exhibited inaccurate KI. Meanwhile, it has been reported that CRISPR-Cas9 genome editing can increase the efficiency of long DNA KI in mouse ES cells. In this method, a circular plasmid was introduced into ES cells and subsequently linearized with CRISPR-Cas9. Nevertheless, the exact KI ratio was approximately half of the integrated ES cells, with the remainder exhibiting incomplete partial KI or random integration. Furthermore, the efficiency of CRISPR-Cas9-mediated long-chain DNA KI in fertilized eggs was improved compared to circular plasmids when the double-stranded target plasmid was linearized intracellularly by CRISPR-Cas9. However, no significant difference in KI efficiency was observed between circular and linearized plasmids in mouse ES cells. Our results demonstrate that when using CRISPR-Cas9-RNP, circular plasmids KI in various chromosomes even without drug selection, with a ratio of up to approximately 60%. Therefore, at least in the case of KI in ES cells, a sufficiently high KI efficiency can be achieved by using Cas9-RNP and a circular plasmid as the target vector.

[0096] In this study, the KI efficiency in ES cells was significantly higher with Cas9-RNP than with Cas9 plasmid. Unlike the Cas9 plasmid, which requires intracellular transcription and translation, Cas9-RNP rapidly enters the nucleus and cleaves the genome after transfection. Previous reports have shown that Cas9-RNP significantly increases genome editing efficiency compared to Cas9-plasmid. Therefore, it is suggested that the increased rate of genome double-strand breaks by Cas9-RNP led to the increased KI efficiency in ES cells observed in this study.

[0097] It is also worth noting that Cas9-RNPs are degraded more rapidly in cells than Cas9 plasmids, resulting in less frequent off-target cleavage. Furthermore, in this example, high-fidelity (HiFi) Cas9 protein was used as the RNP component. Several studies have reported that off-target cleavage by HiFi Cas9 significantly reduces to undetectable levels through next-generation sequence analysis. Therefore, this KI method using HiFi Cas9 protein requires further analysis to clarify the presence or absence of off-target cleavage, but site specificity will also be considered. On the other hand, PCR KI bands were sometimes observed only on either the 5' or 3' side in this experiment, suggesting that this method resulted in some degree of incomplete KI. Therefore, performing both 5' and 3' PCR is essential for KI screening. It should also be noted that the 5' or 3' side may be KI for different chromosomes. Therefore, it may be useful to confirm KI by PCR using primers designed outside each HA to amplify the entire target allele.

[0098] Furthermore, by using a plasmid carrying different drug resistance gene cassettes as a targeting vector, we achieved simultaneous gene KI at multiple loci with high efficiency (Figure 4). This theoretically halves the development time for double-gene-modified ES cell clones compared to sequential manipulation of two genes. Long-term culture of ES cells is known to result in more abnormal DNA methylation and increased chromosomal instability, resulting in reduced pluripotency. Therefore, modifying multiple genes with a single electroporation in a shorter culture period would be advantageous for developing chimeric mice while maintaining the pluripotent quality of ES cells.

[0099] Generally, mice carrying double mutations are generated by crossing with each mutation, which requires several months to obtain individuals of an appropriate age for phenotypic analysis. The existing plasmid KI method with drug selection allows for the short-term manipulation of multiple genes in chimeric mice and is therefore very useful, especially in areas where phenotypic analysis can be performed on chimeric animals.

[0100] The contribution of blastocyst-injected ES cells to germ cells, i.e., sperm or oocytes, in chimeras (germline transmission, GLT) is a key issue in generating stable genetically engineered mice through chimeric mice. In particular, chimeric mice generated using ES cells from inbred strains such as C57BL / 6 and BALB / c often have poor GLT. GLT can be variable and inefficient, partly due to competition between host and donor cells in chimeras, which is affected by the quality of the ES cells. Some recent studies have suggested that the blastocyst complementation method, in which ES cells are injected into genetically germline-deprived blastocysts, can significantly improve the efficiency of GLT. Although six out of nine mouse strains achieved GLT in the current experiment (Table 3, supra), future studies will be worthwhile combining our gene targeting technology with the blastocyst complementation method to establish a more efficient technique for generating genetically modified animals that can consistently achieve GLT.

[0101] Another concern is backcrossing. Certain research areas require analysis using specific pure-bred mice. Therefore, if ES cell-based chimeric mice or zygote genome-edited mice are generated using an undesirable strain, backcrossing must be performed 6 to 10 times to obtain congenic strains before phenotypic analysis can be performed. This process requires a long time, lasting more than a year. In this example, we demonstrated that precise and highly efficient gene targeting is possible using ES cells from various genetic backgrounds, such as a hybrid of the B6-129 F1 strain or a pure strain of B6N and BALB / c. Injecting these ES cells into 4n tetraploid blastocysts enables in vivo gene function analysis using pure inbred mouse strains in the F0 generation. Future studies may require the establishment of high-quality ES cells from additional pure mouse strains to avoid time-consuming backcrossing.

[0102] In conclusion, our study demonstrated that multiple KI editing can be performed in a single step with high efficiency using circular plasmids with Cas9 / RNP genome editing in mouse ES cells. We believe that this technology will be useful in substantially reducing both the time required for traditional generation of genetically engineered chimeric mouse models and the number of animals currently used in this process.

[0103] Example 6: Knock-in of a long nucleic acid sequence. Using a method similar to that used in Example 4, a sequence homologous to human KIT and a puromycin resistance gene were introduced into mouse ES cells, followed by selection with puromycin to obtain ES cell clones in which the KIT sequence had been integrated into the Rosa26 locus. The nucleic acid sequence of human KIT is disclosed as Gene ID: 3815, and sequence information can be obtained by referring to locus NC 000004.12 (https: / / www.ncbi.nlm.nih.gov / nuccore / NC_000004.12?report=fasta&from=54657957&to=54740715). Subsequently, a human KIT BAC incorporating a blasticidin S resistance gene was introduced into these ES cells using a method similar to that used in Example 4 to obtain ES cells in which the full-length human KIT had been integrated into the Rosa26 locus. The success or failure of KI was confirmed by PCR using the following primers, which specifically amplify only when human KIT is KI'd at the 5' and 3' ends: This example demonstrates that the knock-in method of the present invention is an effective means for replacing a long (approximately 100 kbp) full-length human orthologous gene with the corresponding full-length mouse orthologous gene. 5' side primer: 5'-TCTCTGCTGCCTCCTGGCTTCTGAGGACCGCCCTGGGCCTGGGAGAATCCCTTCCCCCTCTTCCCTCGTGATCTGCAACTCCAGTCTTATTTAGAACTTTGGTAGACGTTGTCAGTTCTTATTATGATTACAAACTCCCGCTGATGCTTTTCTGCCCTTATCTTATTTGACCTCCCTATG-3' (SEQ ID NO: 20) 3' side primer: 5'-ATTTCCACCCTCTTCGAGTTTCCATTTTTTGCTATTGAGAAGTCTGTTCAGGTTGTCATTCTCCTGAAGATATTATTTTTTTGTTTTCATCTAGAAGATGGGCGGGAGTCTTCTGGGCAGGCTTAAAGGCTAACCTGGTGTGGGCGTTGTCCTGCAGGGGAATTGAACAGGTGTAAA-3' (SEQ ID NO: 21)

[0104] Example 7: Examination of knock-in efficiency depending on the length of the homologous sequence of the vector used in the first stage to induce homologous recombination. Aiming to achieve further efficiency in knock-in, we newly investigated knock-in efficiency, focusing on the length of the homologous sequence. Specifically, targeting vectors with no homologous sequence, 100 bp, 1 kbp, and 3 kbp homologous sequences were used in the first stage knock-in, followed by a similar 90 kbp human gene knock-in in the second stage. As a result, human gene knock-in was unsuccessful when first-stage vectors with no homologous sequence or a 100 bp homologous sequence were used. On the other hand, human gene knock-in was confirmed with first-stage vectors with 1 kbp and 3 kbp homologous sequences, and the efficiency was higher when a first-stage vector with a 3 kbp homologous sequence was used. This indicates that the use of a first-stage vector with a homologous sequence of approximately 3 kbp improves the work efficiency in generating humanized mice (data not shown).

[0105] Example 8: Extension of DNA length that can be knocked into the mouse genome Using the first-stage vector containing the 3 kbp homologous sequence identified in Example 7, we attempted to knock in a 200 kbp human gene region (full-length human APOBEC3 cluster + Bsd cassette), more than twice the length of the previously successful 90 kbp knock-in. The knock-in was performed using the two-stage knock-in method described above. As a result, we succeeded in precisely knocking in a 200 kbp gene region containing seven genes into a specific locus in the mouse genome (data not shown).

[0106] The knock-in method of the present invention can be used as an effective tool for the development of new antibody drugs (for cancer, immune diseases, etc.) and for the search for new compounds that target specific molecules (for cancer, immune diseases, etc.). Furthermore, the present invention can be an excellent means for producing highly accurate animal models. It has been pointed out that the same compound can react differently depending on the animal species, and dosing based on the results of inappropriate animal models can sometimes cause serious drug-related harm, such as the thalidomide-related harm in the 1970s. This method is expected to prevent such problems from occurring in the future.

Claims

1. A method for knocking-in a full-length human orthologous gene into a corresponding gene locus in a non-human mammal, comprising the steps of: (a) preparing a first circular vector targeting the full-length human orthologous gene at the full-length human orthologous gene locus in the non-human mammal, the first circular vector comprising a gene encoding a drug resistance gene, and at its 5'-end and 3'-end, a non-human mammal homologous arm sequence adjacent to the full-length human orthologous gene, and a human homologous arm sequence adjacent to each of the non-human mammal homologous arm sequences; (b) knocking-in a gene encoding a drug resistance gene flanked by human homologous sequences into the full-length orthologous gene locus in a non-human mammal cell by homologous recombination using the first circular vector prepared in (a); (c) preparing a second circular vector comprising the full-length human orthologous gene and a human homologous arm sequence at its 5'-end and 3'-end, respectively; (d) knocking in a gene encoding a full-length human orthologous gene and a drug resistance gene by homologous recombination into the knock-in cell obtained in step (b) using the second circular vector prepared in step (c).

2. The method of claim 1, wherein knock-in is performed by genome editing via homologous recombination.

3. The method according to claim 2, wherein the genome editing method uses CRISPR-Cas9, Cas12a, ZFN, or TALEN.

4. The method according to claim 1 or 2, wherein the length of the full-length human orthologous gene is 10 kbp or more.

5. The method according to claim 1 or 2, further comprising a step of selecting the cells with an antibiotic after steps (b) and (d).

6. The method of claim 1 or 2, wherein the first circular vector is derived from a plasmid vector, a viral vector, or a cosmid vector.

7. The method of claim 1 or 2, wherein the second circular vector is derived from a bacterial artificial chromosome (BAC) vector.

8. The method according to claim 1 or 2, wherein the drug resistance gene is one or more selected from the group consisting of a neomycin resistance gene, a puromycin resistance gene, a blasticidin resistance gene, a hygromycin resistance gene, a chloramphenicol resistance gene, a tetracycline resistance gene, an erythromycin resistance gene, a spectinomycin resistance gene, a kanamycin resistance gene, a zeocin resistance gene, and a phleomycin resistance gene.

9. A kit for knocking in a full-length human orthologous gene into a corresponding locus of a non-human mammal, comprising: (i) a first circular vector targeting a full-length non-human mammalian orthologous gene at a full-length non-human mammalian orthologous gene locus, the first circular vector comprising a gene encoding a drug resistance gene and at its 5' end and 3' end, respectively, a non-human mammalian homologous arm sequence adjacent to the full-length non-human mammalian orthologous gene, and further a human homologous arm sequence adjacent to each of the non-human mammalian homologous arm sequences; and (ii) a second circular vector comprising a full-length human orthologous gene and a human homologous sequence at its 5' end and 3' end, respectively.

10. The kit according to claim 9, wherein the length of the full-length human orthologous gene is 10 kbp or more.

Citation Information

Patent Citations

  • Methods, cells & organisms

    US20150079680A1

  • Gene knockin method and kit for gene knockin

    US20190225989A1

  • Scarless genome editing through two-step homology directed repair

    WO2019018534A1