Genome modification methods and genome modification kits

The genome modification method and kit address the limitations of conventional technologies by using sequence-specific nucleic acid cleavage and selection marker donor DNAs to efficiently modify multiple alleles and large genomic regions, achieving seamless modifications up to 8 kbp.

JP2026086768APending Publication Date: 2026-05-26LOGOMIX INC(JP) +1

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LOGOMIX INC(JP)
Filing Date
2026-02-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Conventional genome modification technologies, such as HDR, struggle to efficiently modify two or more alleles simultaneously and are limited in the size of genomic regions that can be modified, making large-scale modifications difficult.

Method used

A genome modification method and kit that utilizes a sequence-specific nucleic acid cleavage molecule and selection marker donor DNAs with homology arms to introduce distinct selection marker genes into multiple alleles, followed by selective cell sorting and recombinant donor DNA introduction to achieve large-scale modifications.

Benefits of technology

Efficiently modifies two or more alleles and allows for the modification of relatively large genomic regions up to 8 kbp, enabling seamless linking of upstream and downstream sequences without insertions or deletions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086768000001
    Figure 2026086768000001
  • Figure 2026086768000002
    Figure 2026086768000002
  • Figure 2026086768000003
    Figure 2026086768000003
Patent Text Reader

Abstract

The present invention provides a genome modification method and a genome modification kit that can efficiently modify two or more alleles and modify relatively large regions. [Solution] A genome modification method for modifying two or more alleles of a chromosome genome, comprising: (a) introducing the following (i) and (ii) into a cell containing the chromosome; (i) a genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosome genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule; (ii) two or more selection marker donor DNAs having mutually different selection marker genes (the number of types of selection marker donor DNAs is equal to or greater than the number of alleles to be modified); and (b) selecting the cell based on all the selection marker genes possessed by the two or more selection marker donor DNAs.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a genome modification method and a genome modification kit. [Background technology]

[0002] Since the CRISPR / Cas system was reported as a novel genome editing tool, various studies using the CRISPR / Cas system have been conducted (for example, Patent Document 1). In genome editing using the CRISPR / Cas system, the target region targeted by guide RNA is double-stranded by the Cas9 nuclease. It is known that double-stranded DNA is repaired by homologous recombination repair (HDR) or non-homologous end-joining repair (NHEJ). In HDR, any sequence can be incorporated into the target region by introducing donor DNA having a sequence homologous to the surrounding region of the target region into cells together with the CRISPR / Cas system. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] International Publication No. 2014 / 093661 [Overview of the project] [Problems that the invention aims to solve]

[0004] Conventional genome modification technologies, such as HDR, could not efficiently modify two or more alleles simultaneously. Furthermore, HDR had limitations in the size of the genomic region that could be modified, making it difficult to efficiently perform large-scale genome modifications (e.g., 10 kbp).

[0005] Therefore, the object of the present invention is to provide a genome modification method and a genome modification kit that can efficiently modify two or more alleles and modify relatively large regions. [Means for solving the problem]

[0006] The present invention includes, as an example, the following embodiments. [1] A genome modification method for modifying two or more alleles of a chromosomal genome, (a) A step of introducing (i) and (ii) below into cells containing the chromosome, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more selection marker donor DNAs containing the base sequence of a selection marker gene between an upstream homology arm having a base sequence homologous to the base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to the base sequence adjacent to the downstream side of the target region, wherein the two or more selection marker donor DNAs each contain a different selection marker gene, and the number of types of selection marker donor DNAs is equal to or greater than the number of alleles targeted for genome modification, (b) A genome modification method comprising the step of selecting the cells after step (a) based on all of the selection marker genes contained in the two or more selection marker donor DNAs. [2] The genome modification method according to [1], wherein the selection marker gene is a positive selection marker gene, and step (b) is a step of selecting cells that express the same number of positive selection marker genes as the number of alleles. [3] The genome modification method according to [2], wherein the donor DNA for selection markers further has a negative selection marker gene between the upstream homology arm and the downstream homology arm. [4] (c) After step (b), a step of introducing recombinant donor DNA containing a desired base sequence into the cell between an upstream homology arm having a base sequence homologous to a base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to a base sequence adjacent to the downstream side of the target region; and (d) After step (c), a step of selecting cells that do not express the negative selection marker gene, according to [3]. [5] The genome modification method according to [3] or [4], wherein the positive selection marker gene is a drug resistance gene and the negative selection marker gene is a fluorescent protein gene. [6] The genome modification method according to any one of [1] to [5], wherein the sequence-specific nucleic acid cleavage molecule is a sequence-specific endonuclease. [7] The genome modification method according to [6], wherein the genome modification system comprises a Cas protein and a guide RNA having a base sequence homologous to the base sequence in the target region. [8] A genome modification kit for modifying two or more alleles of a chromosomal genome, comprising (i) and (ii) below. (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more selection marker donor DNAs, each containing the base sequence of a selection marker gene between an upstream homology arm having a base sequence homologous to a base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to a base sequence adjacent to the downstream side of the target region, wherein the two or more selection marker donor DNAs each contain a different selection marker gene, and the number of types of selection marker donor DNAs is equal to or greater than the number of alleles targeted for genome modification. [9] The genome modification kit according to [8], wherein the selection marker gene is a positive selection marker gene.

[10] The genome modification kit according to [9], wherein the selection marker donor DNA further has a negative selection marker gene between the upstream homology arm and the downstream homology arm.

[11] The genome modification kit described in any one of [8] to

[10] , wherein the sequence-specific nucleic acid cleavage molecule is a sequence-specific endonuclease.

[12] A genome modification kit according to any one of [8] to

[11] , comprising a Cas protein and a guide RNA having a base sequence homologous to the base sequence in the target region.

[0007] The present invention includes, as an example, the following embodiments. [1] A method for producing cells in which two or more alleles of the chromosomal genome have been modified, (a) The steps of introducing (i) and (ii) below into cells containing two or more alleles to introduce a selection marker gene into each of the two or more alleles, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving target regions in two or more alleles of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, and containing the base sequence of the select marker gene between the upstream homology arm and the downstream homology arm, wherein each of the two or more types of select marker donor DNA has a select marker gene that is distinguishable from each other, the select marker gene is unique to each type of select marker donor DNA, and the number of types of select marker donor DNA is equal to or greater than the number of alleles targeted for genome modification, (b) After the step (a), different types of donor DNAs for selection markers respectively undergo homologous recombination with the two or more alleles, whereby distinct and unique selection marker genes are introduced into the two or more alleles, and a step of selecting cells that express all of the introduced and distinct selection marker genes (a step for positive selection); A method comprising the above. [2] A method for modifying two or more alleles of a chromosomal genome, comprising: (a) A step of introducing the following (i) and (ii) into a cell containing two or more alleles to introduce a selection marker gene into each of the two or more alleles, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting a target region in two or more alleles of the chromosomal genome and cleaving the target region, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule, (ii) Two or more types of donor DNAs for selection markers, each having an upstream homology arm having a base sequence capable of homologous recombination with the base sequence upstream of the target region and a downstream homology arm having a base sequence capable of homologous recombination with the base sequence downstream of the target region, and having a base sequence of a selection marker gene between the upstream homology arm and the downstream homology arm. The two or more types of donor DNAs for selection markers each have a distinct and unique selection marker gene, the selection marker gene is unique for each type of donor DNA for selection marker, and the number of types of donor DNAs for selection marker is equal to or greater than the number of alleles to be genome-modified; (b) After the step (a), different types of donor DNAs for selection markers respectively undergo homologous recombination with the two or more alleles, whereby distinct and unique selection marker genes are introduced into the two or more alleles, and a step of selecting cells that express all of the introduced and distinct selection marker genes (a step for positive selection); A method comprising the above. 〔3〕The method according to 〔1〕 or 〔2〕 above, wherein the target region has a length of 5 kbp or more. 〔4〕The method according to 〔3〕 above, wherein the target region has a length of 8 kbp or more. 〔5〕Each of the donor DNAs for two or more selection markers has, between the upstream homology arm and the downstream homology arm, a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence. Here, when the selection marker gene is used for both positive selection and negative selection, it may not have another selection marker gene for negative selection. (c) After the step (b), introducing into the selected cells the following (iii) and (iv) to introduce the donor DNA for recombination into the two or more alleles: (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) A donor DNA for recombination containing a desired base sequence, having an upstream homology arm having a base sequence capable of homologous recombination with the base sequence on the upstream side of the target region, and a downstream homology arm having a base sequence capable of homologous recombination with the base sequence on the downstream side of the target region. Donor DNA for recombination (d) After the step (c), a step of selecting cells that do not express the marker gene for negative selection (step for negative selection). The method according to any one of 〔1〕 to 〔4〕 above, further comprising the above steps. 〔6〕Each of the donor DNAs for two or more selection markers has, between the upstream homology arm and the downstream homology arm, a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence. Here, when the selection marker gene is used for both positive selection and negative selection, it may not have another selection marker gene for negative selection. (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired base sequence, comprising an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, Recombinant donor DNA, (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection), The method described in [3] above, further comprising: [7] Each of two or more selection marker donor DNAs has a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence between the upstream homology arm and the downstream homology arm, wherein if the selection marker gene is used for both positive and negative selection, it does not need to have another selection marker gene for negative selection. (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired base sequence, comprising an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection), The method described in [4] above, further comprising: [8] The method according to any one of [5] to [7] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more. [9] The method according to [6] or [7] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

[10] The method according to [7] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

[11] The method according to [8] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

[12] A genome modification kit for modifying two or more alleles of a chromosomal genome, comprising (i) and (ii) below. (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule, (ii) Two or more donor DNAs for selection markers, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, wherein the base sequence of the selection marker gene is included between the upstream homology arm and the downstream homology arm. Two or more types of selection marker donor DNA, each having distinct selection marker genes, each selection marker gene being unique to each type of selection marker donor DNA, and the number of types of selection marker donor DNA being equal to or greater than the number of alleles targeted for genome modification.

[13] The kit described in

[12] above, wherein the target region has a length of 5 kbp or more.

[14] The kit described in

[13] above, wherein the target region has a length of 8 kbp or more.

[15] A kit according to any of

[12] to

[14] above, further comprising recombinant donor DNA.

[16] A kit according to any one of

[12] to

[15] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

[17] A kit according to any one of

[12] to

[16] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

[18] The kit according to

[12] above, wherein the target region has a length of 5 kbp or more, and the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

[19] The kit according to

[18] above, wherein the target region has a length of 8 kbp or more, and the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

[20] The method according to [5] above, wherein the recombinant donor DNA does not have a base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA, and in the two or more alleles of the chromosomal genome obtained after modification, the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

[21] The method according to [6] or [7] above, wherein the recombinant donor DNA has no base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA, and in the two or more alleles of the chromosomal genome obtained after modification, the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

[22] The method according to [3] above, wherein the two or more alleles of the chromosomal genome obtained after modification do not contain target sequences for site-directed recombinant enzymes.

[23] The method according to [4] above, wherein the two or more alleles of the chromosomal genome obtained after modification do not contain target sequences for site-directed recombinant enzymes.

[24] The method according to [5] above, wherein the two or more alleles of the chromosomal genome obtained after modification do not contain target sequences for site-directed recombinant enzymes.

[25] The method according to [6] or [7] above, wherein the two or more alleles of the chromosomal genome obtained after modification do not contain target sequences for site-directed recombinant enzymes.

[26] The method according to any of [1] to

[11] and

[20] to

[25] above, wherein single-cell cloning is not performed in the process up to selecting cells in which two or more alleles have been modified in step (b).

[27] A cell having two or more alleles on the chromosomal genome with respect to a target region, wherein the target region of each of the two or more alleles is deleted, and the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

[28] The cell described in

[27] above, wherein the target region has a length of 5 kbp or more.

[29] The cells described in

[27] or

[28] above, which do not have a target sequence for site-directed recombinant enzymes in their genome.

[30] The method according to [6] or [7] above, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

[31] The kit according to

[15] above, wherein the recombinant donor DNA does not have a base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA. [Effects of the Invention]

[0008] The present invention provides a genome modification method and a genome modification kit that can efficiently modify two or more alleles and modify relatively large regions. [Brief explanation of the drawing]

[0009] [Figure 1]The structure of the donor DNA prepared in Experimental Example 1 is shown. [Figure 2] This shows the process of introducing the donor DNA prepared in Experimental Example 1 into cells. It also shows the design position of the junction primer used to confirm the knock-in of the donor DNA. [Figure 3] The results of junction PCR performed after knock-in of donor DNA are shown. [Figure 4] This section shows the process of introducing donor DNA, prepared in the reference experiment, into cells. It also shows the design position of the junction primer used to confirm donor DNA knock-in, and the results of junction PCR performed after donor DNA knock-in. [Figure 5] This section describes the process of introducing the wild-type TP53 plasmid as donor DNA into cells prepared using the donor DNA prepared in Experimental Example 1 (cells prepared in Experimental Example 2). It also shows the design positions of the primers used to confirm the knock-in of the wild-type TP53 plasmid. [Figure 6] The results of PCR performed using the primers shown in Figure 5 after knock-in of the wild-type TP53 plasmid are shown. [Figure 7] This section shows the process of introducing the four types of donor DNA prepared in Experimental Example 4 into the cells prepared in Experimental Example 2. It also shows the design positions of the primers used to confirm the knock-in of the four types of donor DNA. [Figure 8] The results of PCR performed using the primers shown in Figure 7 after knock-in of the four types of donor DNA prepared in Experimental Example 4 are shown. [Figure 9] The results of the experiment in which deletions were induced in both alleles of the TP53 gene in human iPS cells, as performed in Example 5, are shown. [Figure 10] The results of the experiment in which deletions were induced in both alleles of the MLH1 gene, as performed in Example 6, are shown. [Figure 11] The results of the experiment in which deletions were induced in both alleles of the CD44 gene, as performed in Example 6, are shown. [Figure 12]The results of the experiment in which deletions were induced in both alleles of the MET gene, as performed in Example 6, are shown. [Figure 13] The results of the experiment in which deletions were induced in both alleles of the APP gene, as performed in Example 6, are shown. [Figure 14A] This is a schematic diagram illustrating that in the step of introducing donor DNA for selection markers in the method of the present invention, two cleavage sites are induced so as to sandwich the target region of the genome. [Figure 14B] This is a schematic diagram illustrating the induction of a single cleavage near the target region of the genome during the step of introducing donor DNA for selection markers in the method of the present invention. [Modes for carrying out the invention]

[0010] [Definition] The terms "polynucleotide" and "nucleic acid" are used interchangeably and refer to nucleotide polymers in which nucleotides are linked by phosphodiester bonds. "Polynucleotides" and "nucleic acids" may be DNA, RNA, or a combination of DNA and RNA. Furthermore, "polynucleotides" and "nucleic acids" may be polymers of natural nucleotides, polymers of natural nucleotides and non-natural nucleotides (analogs of natural nucleotides, nucleotides in which at least one of the base, sugar, and phosphate parts is modified (e.g., a phosphorothioate skeleton)), or polymers of non-natural nucleotides.

[0011] The base sequences of "polynucleotides" or "nucleic acids" are written using commonly accepted single-letter codes unless otherwise specified. Unless otherwise specified, base sequences are written from the 5' end to the 3' end. The nucleotide residues that make up "polynucleotides" or "nucleic acids" may be written simply as adenine, thymine, cytosine, guanine, or uracil, or by their single-letter codes.

[0012] The term "gene" refers to a polynucleotide containing at least one open reading frame that codes for a particular protein. Genes can contain both exons and introns.

[0013] The terms "polypeptide," "peptide," and "protein" are interchangeable and refer to polymers of amino acids linked by amide bonds. A "polypeptide," "peptide," or "protein" may be a polymer of natural amino acids, a polymer of natural amino acids and non-natural amino acids (chemical analogs, modified derivatives, etc. of natural amino acids), or a polymer of non-natural amino acids. Unless otherwise specified, amino acid sequences are written from the N-terminus to the C-terminus.

[0014] The term "allele" refers to a set of base sequences located at the same locus on a chromosomal genome. In some embodiments, diploid cells have two alleles at the same locus, and triploid cells have three alleles at the same locus. In other embodiments, additional alleles may be formed by abnormal copies of the chromosome or abnormal additional copies of the locus in question.

[0015] The terms "genome modification" and "genome editing" are interchangeable and refer to inducing mutations at a desired location (target region) on the genome. Genome modification may include the use of sequence-specific nucleic acid cleavage molecules designed to cleave target region DNA. In a preferred embodiment, genome modification may include the use of nucleases engineered to cleave target region DNA. In a preferred embodiment, genome modification may include the use of nucleases engineered to cleave target sequences having a specific base sequence within the target region (e.g., TALENs or ZFNs). In a preferred embodiment, genome modification may include the use of restriction enzymes that have only one cleavage site in the genome, such as meganucleases (e.g., 16-base sequence-specific restriction enzymes (theoretically 4)) to cleave target sequences having a specific base sequence within the target region. 16 Restriction enzymes with 17-base sequence specificity (theoretically 4) (present at a ratio of one per base)17 Restriction enzymes that have 18-base sequence specificity (theoretically 4 18Sequence-specific endonucleases, such as those present at a rate of one per base, can also be used. Typically, the use of site-specific nucleases induces double-strand breaks (DSBs) in the DNA of the target region, and the genome is subsequently repaired by endogenous cellular processes such as homologous recombination repair (HDR) and non-homologous end-joining repair (NHEJ). NHEJ is a repair method that joins the ends of double-strand breaks without using donor DNA, and insertions and / or deletions (indels) are frequently induced during the repair. HDR is a repair mechanism that uses donor DNA, and it is also possible to introduce desired mutations into the target region. As for genome modification technologies, the CRISPR / Cas system is a preferred example.Examples of meganucleases include I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-CeuI, I-CeuAIIP, I-CreI, I-CrepsbIP, I-CrepsbIIP, I-CrepsbIIIP, I-CrepsbIVP, I-TliI, I-PpoI, PI-PspI, F-SceI, F-SceII, F-SuvI, F-TevI, F-TevII, I-AmaI, I-AniI, I-ChuI, I-CmoeI, I-CpaI, I-CpaII, I-CsmI, I-CvuI , I-CvuAIP, I-DdiI, I-DdiII, I-DirI, I-DmoI, I-HmuI, I-HmuII, I-HsNIP, I-LlaI, I-MsoI, I-NaaI, I-Na nI, I-NclIP, I-NgrIP, I-NitI, I-NjaI, I-Nsp236IP, I-PakI, I-PboIP, I-PcuIP, I-PcuAI, I-PcuVI, I-Pg rIP, I-PobIP, I-PorI, I-PorIIP, I-PbpIP, I-SpBetaIP, I-ScaI, I-SexIP, I-SneIP, I-SpomI, I-SpomCP, I-SpomIP, I-SpomIIP, I-SquIP, I-Ssp68031, I-SthPhiJP, I-SthPhiST3P, I-SthPhiSTe3bP, I-TdeIP, I- TevI, I-TevII, I-TevIII, I-UarAP, I-UarHGPAIP, I-UarHGPA13P, I-VinIP, I-ZbiIP, PI-Mtul, PI-MtuHIP A meganuclease and its cleavage site (or recognition site) selected from the group consisting of PI-MtuHIIP, PI-PfuI, PI-PfuII, PI-PkoI, PI-PkoII, PI-Rma43812IP, PI-SpBetaIP, PI-SceI, PI-TfuI, PI-TfuII, PI-ThyI, PI-TliI, and PI-TliII, and functional derivative restriction enzymes thereof, preferably a meganuclease and its cleavage site (or recognition site) that is a restriction enzyme having sequence specificity of 18 bases or more, in particular a meganuclease and its cleavage site that does not cleave the cell genome at one or more locations, can be used.

[0016] The term "target region" refers to the genomic region that is the target of genome modification.

[0017] The term "donor DNA" refers to DNA used to repair double-strand breaks in DNA, which is homologous to the DNA surrounding the target region. Donor DNA includes homology arms consisting of a base sequence upstream of the target region and a base sequence downstream of the target region (e.g., a base sequence adjacent to the target region). In this specification, a homology arm consisting of a base sequence upstream of the target region (e.g., a base sequence adjacent to the upstream side) may be referred to as an "upstream homology arm," and a homology arm consisting of a base sequence downstream of the target region (e.g., a base sequence adjacent to the downstream side) may be referred to as a "downstream homology arm." Donor DNA may include a desired base sequence between the upstream homology arm and the downstream homology arm. The length of each homology arm is preferably 300 bp or longer, and is usually around 500 to 3000 bp. The lengths of the upstream and downstream homology arms may be the same or different. If homologous recombination is successfully induced between the target region and donor DNA after sequence-dependent cleavage, the sequences between the upstream and downstream base sequences of the target region will be replaced with sequences from the donor DNA.

[0018] The "upstream" region of a target region refers to the DNA region located at the 5' end of the reference nucleotide strand in the double-stranded DNA of the target region. The "downstream" region refers to the DNA located at the 3' end of the reference nucleotide strand. If the target region contains a protein-coding sequence, the reference nucleotide strand is usually the sense strand. Generally, the promoter is located upstream of the protein-coding sequence. The terminator is located downstream of the protein-coding sequence.

[0019] The term "sequence-specific nucleic acid cleavage molecule" refers to a molecule that recognizes a specific nucleic acid sequence and can cleave the nucleic acid at that specific sequence. A sequence-specific nucleic acid cleavage molecule is a molecule that has the activity to cleave nucleic acids in a sequence-specific manner (sequence-specific nucleic acid cleavage activity).

[0020] The term "target sequence" refers to the DNA sequence in the genome that is targeted for cleavage by a sequence-specific nucleic acid cleavage molecule. When the sequence-specific nucleic acid cleavage molecule is a Cas protein, the target sequence refers to the DNA sequence in the genome that is targeted for cleavage by the Cas protein. When using the Cas9 protein as the Cas protein, the target sequence must be a sequence adjacent to the 5' end of a protospacer adjacent motif (PAM). Typically, the target sequence is selected from a sequence of 17 to 30 bases (preferably 18 to 25 bases, more preferably 19 to 22 bases, and even more preferably 20 bases) adjacent to the 5' end of a PAM. Known design tools such as CRISPR DESIGN (crispr.mit.edu / ) can be used to design the target sequence.

[0021] The term "Cas protein" refers to a CRISPR-associated protein. In a preferred embodiment, the Cas protein forms a complex with guide RNA and exhibits endonuclease activity or nickase activity. Examples of Cas proteins, though not particularly limited, include Cas9 protein, Cpf1 protein, C2c1 protein, C2c2 protein, and C2c3 protein. The Cas protein includes wild-type Cas proteins and their homologs (paralogs and orthologs), as well as their variants, insofar as they cooperate with guide RNA to exhibit endonuclease activity or nickase activity. In a preferred embodiment, the Cas protein is involved in a class 2 CRISPR / Cas system, and more preferably in a type II CRISPR / Cas system. A preferred example of a Cas protein is the Cas9 protein.

[0022] The term "Cas9 protein" refers to the Cas protein involved in the type II CRISPR / Cas system. The Cas9 protein forms a complex with guide RNA and exhibits activity in cooperation with the guide RNA to cleave DNA in a target region. The Cas9 protein includes the wild-type Cas9 protein and its homologs (paralogs and orthologs), as well as their variants, as long as they possess the aforementioned activity. The wild-type Cas9 protein has a RuvC domain and an HNH domain as nuclease domains, but the Cas9 protein used herein may have either the RuvC domain or the HNH domain inactivated. Cas9 with either the RuvC domain or the HNH domain inactivated introduces single-strand breaks (nicks) into double-stranded DNA. Therefore, when using Cas9 with either the RuvC domain or the HNH domain inactivated to cleave double-stranded DNA, a modified system can be constructed in which target sequences for Cas9 are set for both the sense strand and the antisense strand, so that nicks in the sense strand and antisense difference occur at sufficiently close positions, thereby inducing double-strand breaks. The species from which the Cas9 protein originates is not particularly limited, but bacteria belonging to the genera Streptococcus, Staphylococcus, Neisseria, or Treponema are preferred examples. More specifically, Cas9 proteins derived from S. pyogenes, S. thermophilus, S. aureus, N. meningitidis, or T. denticola are preferred examples. In a preferred embodiment, the Cas9 protein is a Cas9 protein derived from S. pyogenes.

[0023] Information on the amino acid sequences and coding sequences of various Cas proteins can be obtained from various databases such as GenBank, UniProt, and Addgene. For example, the amino acid sequence of the Cas9 protein of S. pyogenes can be used from the one registered in Addgene as plasmid number 42230. An example of the amino acid sequence of the Cas9 protein of S. pyogenes is shown in Sequence ID No. 1.

[0024] The terms “guide RNA” and “gRNA” are used interchangeably and refer to RNA that can form a complex with the Cas protein and guide the Cas protein to a target region. In a preferred embodiment, the guide RNA includes CRISPR RNA (crRNA) and trans-activated CRISPR RNA (tracrRNA). crRNA is involved in binding to the target region on the genome, and tracrRNA is involved in binding to the Cas protein. In a preferred embodiment, the crRNA includes a spacer sequence and a repeat sequence, the spacer sequence binding to the complementary strand of the target sequence in the target region. In a preferred embodiment, the tracrRNA includes an anti-repeat sequence and a 3' tail sequence. The anti-repeat sequence has a sequence complementary to the repeat sequence of the crRNA and forms base pairs with the repeat sequence, and the 3' tail sequence typically forms three stem-loops. The guide RNA may be a single guide RNA (sgRNA) formed by ligating the 5' end of tracrRNA to the 3' end of crRNA, or it may be a separate RNA molecule with base pairings formed by repeat and anti-repeat sequences. In a preferred embodiment, the guide RNA is sgRNA.

[0025] The repeat sequences of crRNA and tracrRNA can be appropriately selected depending on the type of Cas protein, and those derived from the same bacterial species as the Cas protein can be used. For example, when using Cas9 protein derived from S. pyogenes, the length of the sgRNA can be approximately 50 to 220 nucleotides (nt), preferably 60 to 180 nt, and more preferably 80 to 120 nt. The length of the crRNA, including the spacer sequence, can be approximately 25 to 70 base pairs, preferably 25 to 50 nt. The length of the tracrRNA can be approximately 10 to 130 nt, preferably 30 to 80 nt. The repeat sequence of crRNA may be the same as that in the bacterial species from which the Cas protein originates, or it may have a portion of its 3' end removed. TracrRNA may have the same sequence as mature tracrRNA in the bacterial species from which the Cas protein originates, or it may be a truncated form obtained by cutting the 5' and / or 3' ends of the mature tracrRNA. For example, tracrRNA may be a truncated form obtained by removing approximately 1 to 40 nucleotide residues from the 3' end of mature tracrRNA. Alternatively, tracrRNA may be a truncated form obtained by removing approximately 1 to 80 nucleotide residues from the 5' end of mature tracrRNA. Furthermore, tracrRNA may be a truncated form obtained by removing approximately 1 to 20 nucleotide residues from the 5' end and approximately 1 to 40 nucleotide residues from the 3' end. Various crRNA repeat sequences and tracrRNA sequences have been proposed for sgRNA design, and those skilled in the art can design sgRNAs based on known techniques (e.g., Jinek et al. (2012) Science, 337, 816-21; Mali et al. (2013) Science, 339: 6121, 823-6; Cong et al. (2013) Science, 339: 6121, 819-23; Hwang et al. (2013) Nat. Biotechnol. 31: 3, 227-9; Jinek et al. (2013) eLife, 2, e00471).

[0026] The terms "protospacer adjacency motif" and "PAM" are used interchangeably and refer to the sequence recognized by the Cas protein during DNA cleavage by the Cas protein. The sequence and position of the PAM vary depending on the type of Cas protein. For example, in the case of the Cas9 protein, the PAM must be adjacent to the target sequence immediately after the 3' end. The PAM sequence corresponding to the Cas9 protein varies depending on the bacterial species from which the Cas9 protein originates. For example, the PAM corresponding to the Cas9 protein of S. pyogenes is "NGG", the PAM corresponding to the Cas9 protein of S. thermophilus is "NNAGAA", the PAM corresponding to the Cas9 protein of S. aureus is "NNGRRT" or "NNGRR(N)", the PAM corresponding to the Cas9 protein of N. meningitidis is "NNNNGATT", and the PAM corresponding to the Cas9 protein of T. denticola is "NAAAAC" ("R" is A or G; "N" is A, T, G or C).

[0027] The terms "spacer sequence" and "guide sequence" are used interchangeably and refer to sequences included in the guide RNA that can bind to the complementary strand of the target sequence. Typically, the spacer sequence is identical to the target sequence (except where T in the target sequence becomes U in the spacer sequence). In embodiments of the present invention, the spacer sequence may contain one or more nucleotide mismatches with respect to the target sequence. If it contains multiple nucleotide mismatches, the mismatches may be adjacent or distant. In a preferred embodiment, the spacer sequence may contain 1 to 5 nucleotide mismatches with respect to the target sequence. In a particularly preferred embodiment, the spacer sequence may contain one nucleotide mismatch with respect to the target sequence. In guide RNA, the spacer sequence is located at the 5' end of the crRNA.

[0028] The term "functionally linked" as used in relation to polynucleotides means that the first base sequence is positioned close enough to the second base sequence that the first base sequence can influence the second base sequence or a region under the control of the second base sequence. For example, functionally linked polynucleotides to a promoter mean that the polynucleotide is linked in such a way that it is expressed under the control of the promoter.

[0029] The term "expression-ready state" refers to a state in which a polynucleotide can be transcribed within a cell into which the polynucleotide has been introduced. The term "expression vector" refers to a vector containing a target polynucleotide and equipped with a system that enables the expression of the target polynucleotide within the cell into which the vector is introduced. For example, a "Cas protein expression vector" means a vector that enables the expression of the Cas protein within the cell into which the vector is introduced. Similarly, a "guide RNA expression vector" means a vector that enables the expression of guide RNA within the cell into which the vector is introduced.

[0030] In this specification, sequence identity (or homology) between nucleotide sequences or amino acid sequences is determined by juxtaposing the two nucleotide sequences or amino acid sequences, inserting gaps in the areas corresponding to insertions and deletions so that the corresponding nucleotides or amino acids are most likely to match, and then determining the proportion of matching nucleotides or amino acids relative to the entire nucleotide sequence or amino acid sequence, excluding the gaps in the resulting alignment. Sequence identity between nucleotide sequences or amino acid sequences can be determined using various homology search software known in the art. For example, the sequence identity value of a nucleotide sequence can be obtained by calculation based on the alignment obtained by the known homology search software BLASTN, and the sequence identity value of an amino acid sequence can be obtained by calculation based on the alignment obtained by the known homology search software BLASTP.

[0031] [Genome modification methods] In one embodiment, the present invention provides a genome modification method for modifying two or more alleles of a chromosomal genome. The genome modification method includes the following steps (a) and (b): (a) the step of introducing (i) and (ii) below into cells containing the chromosome; and (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more selection marker donor DNAs containing the base sequence of a selection marker gene between an upstream homology arm having a base sequence homologous to the base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to the base sequence adjacent to the downstream side of the target region, wherein the two or more selection marker donor DNAs each contain a different selection marker gene, and the number of types of selection marker donor DNAs is equal to or greater than the number of alleles targeted for genome modification, (b) A step of selecting the cells after step (a) based on all the selection marker genes present in the two or more selection marker donor DNAs. In this embodiment, the selection marker genes may be unique to each type of selection marker donor DNA. In this embodiment, step (b) may also be a step of selecting cells that express all of the introduced uniquely distinct selection marker genes, after step (a), by homologous recombination of different types of selection marker donor DNAs with respect to the two or more alleles (a step for positive selection). The above method may also be a method for producing cells in which two or more alleles of the chromosomal genome have been modified.

[0032] In one embodiment, the present invention is a method for producing cells in which two or more alleles of the chromosomal genome have been modified, (a) The steps of introducing (i) and (ii) below into cells containing two or more alleles to introduce a selection marker gene into each of the two or more alleles, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving target regions in two or more alleles of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, and containing the base sequence of the select marker gene between the upstream homology arm and the downstream homology arm, wherein each of the two or more types of select marker donor DNA has a select marker gene that is distinguishable from each other, the select marker gene is unique to each type of select marker donor DNA, and the number of types of select marker donor DNA is equal to or greater than the number of alleles targeted for genome modification, (b) After step (a), a step of selecting cells that express all of the introduced, distinctly different, unique selection marker genes, by homologous recombination of different types of selection marker donor DNA with respect to the two or more alleles (a step for positive selection), This could include methods such as:

[0033] (Step (a)) In step (a), (i) and (ii) are introduced into cells containing chromosomes.

[0034] The cells used in the genome modification method of this embodiment are not particularly limited and may be any cells having a chromosome genome of diploid or greater. The cells may be diploid, triploid, or tetraploid or greater. The cells are not particularly limited, but eukaryotic cells are an example. The cells may be plant cells, animal cells, or fungal cells. The animal cells are not particularly limited, but may be human, non-human mammals, birds, reptiles, amphibians, fish, insects, or other invertebrates. The cells used in the genome modification method of this embodiment are not cells that do not have alleles (for example, cells having a haploid chromosome genome, such as prokaryotic cells).

[0035] The target region for genome modification can be any region on the genome that has one or more alleles. The size of the target region is not particularly limited. The genome modification method of this embodiment can modify regions of a larger size than conventional methods. The target region may be, for example, 10 kbp or larger. The target region may be, for example, 100 bp or larger, 200 bp or larger, 400 bp or larger, 800 bp or larger, 1 kbp or larger, 2 kbp or larger, 3 kbp or larger, 4 kbp or larger, 5 kbp or larger, 8 kbp or larger, 10 kbp or larger, 20 kbp or larger, 40 kbp or larger, 80 kbp or larger, 100 kbp or larger, 200 kbp or larger, 300 kbp or larger, or 400 kbp or larger. The target region may be, for example, a region containing one to several genes, a region containing one gene, or a region that is a part of one gene. In one embodiment, the target region is deleted in the modified cell.

[0036] <(i) Genome modification system> A "genome modification system" refers to a molecular mechanism capable of modifying a desired target region. The genome modification system includes a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome, or a polynucleotide that encodes the said sequence-specific nucleic acid cleavage molecule.

[0037] The sequence-specific nucleic acid cleavage molecule is not particularly limited as long as it is a molecule that has sequence-specific nucleic acid cleavage activity, and may be a synthetic organic compound or a biomolecular compound such as a protein. Examples of synthetic organic compounds with sequence-specific nucleic acid cleavage activity include pyrrole-imidazole polyamides. Examples of proteins with sequence-specific site cleavage activity include sequence-specific endonucleases.

[0038] Sequence-specific endonucleases are enzymes that can cleave nucleic acids at a predetermined sequence. Sequence-specific endonucleases can cleave double-stranded DNA at a predetermined sequence. Examples of sequence-specific endonucleases are not limited to zinc finger nucleases (ZFNs), TALENs (Transcription activator-like effector nucleases), and Cas proteins, but are not limited to these.

[0039] ZFNs are artificial nucleases containing a nucleic acid cleavage domain conjugated to a binding domain containing a zinc finger array. Examples of cleavage domains include the cleavage domain of the type II restriction enzyme FokI. The design of zinc finger nucleases capable of cleaving target sequences can be carried out using known methods.

[0040] TALENs are artificial nucleases that contain a DNA-binding domain of a transcription activator-like (TAL) effector in addition to a DNA-cleaving domain (e.g., a FokI domain). Designing TALE constructs capable of cleaving target sequences can be done using known methods (e.g., Zhang, Feng et. al. (2011) Nature Biotechnology 29 (2)).

[0041] When a Cas protein is used as the sequence-specific nucleic acid cleavage molecule, the genome modification system includes a CRISPR / Cas system. That is, the genome modification system preferably includes a Cas protein and a guide RNA having a base sequence homologous to the base sequence in the target region. The guide RNA only needs to include a spacer sequence homologous to the sequence in the target region (target sequence). The guide RNA only needs to be able to bind to the DNA in the target region and does not need to have a sequence that is completely identical to the target sequence. This binding should be formed under physiological conditions in the cell nucleus. The guide RNA can, for example, include a mismatch of 0 to 3 bases with respect to the target sequence. The number of mismatches is preferably 0 to 2 bases, more preferably 0 to 1, and even more preferably no mismatches. The guide RNA can be designed based on known methods. The genome modification system is preferably a CRISPR / Cas system and preferably includes a Cas protein and a guide RNA. The Cas protein is preferably a Cas9 protein.

[0042] Sequence-specific endonucleases may be introduced into cells as proteins or as polynucleotides encoding them. For example, mRNA of the sequence-specific endonuclease may be introduced, or an expression vector for the sequence-specific endonuclease may be introduced. In the expression vector, the coding sequence of the sequence-specific endonuclease (sequence-specific endonuclease gene) is functionally linked to a promoter. The promoter is not particularly limited, and various pol II-type promoters can be used. Examples of pol II-type promoters are not particularly limited, but include the CMV promoter, EF1 promoter (EF1α promoter), SV40 promoter, MSCV promoter, hTERT promoter, β-actin promoter, CAG promoter, and CBh promoter.

[0043] The promoter may be an inductive promoter. An inductive promoter is a promoter that can induce the expression of a polynucleotide functionally linked to it only in the presence of an inductive factor that drives the promoter. Examples of inductive promoters include promoters that induce gene expression by heating, such as heat shock promoters. Inductive promoters also include promoters in which the inductive factor that drives the promoter is a drug. Examples of such drug-inductive promoters include Cumate operator sequences, λ operator sequences (e.g., 12×λOp), and tetracycline-based inductive promoters. Examples of tetracycline-based inductive promoters include promoters that drive gene expression in the presence of tetracycline or its derivatives (e.g., doxycycline), or reverse tetracycline-regulating transactivator (rtTA). An example of a tetracycline-based inductive promoter is the TRE3G promoter.

[0044] Any known expression vector can be used without particular restriction. Examples of expression vectors include plasmid vectors and viral vectors. When the sequence-specific endonuclease is a Cas protein, the expression vector may include a guide RNA coding sequence (guide RNA gene) in addition to the Cas protein coding sequence (Cas protein gene). In this case, it is preferable that the guide RNA coding sequence (guide RNA gene) is functionally configured as a pol III promoter. Examples of pol III promoters include mouse and human U6-snRNA promoters, human H1-RNase P RNA promoters, and human valine-tRNA promoters.

[0045] (ii) Donor DNA for selection markers Selective marker donor DNA is donor DNA used to knock in a selective marker into a target region. Selective marker donor DNA contains the base sequence of one or more selective marker genes between an upstream homology arm having a base sequence homologous to the adjacent base sequence upstream of the target region and a downstream homology arm having a base sequence homologous to the adjacent base sequence downstream of the target region.

[0046] The donor DNA for selection markers is not particularly limited, but may have lengths of, for example, 1kb or more, 2kb or more, 3kb or more, 4kb or more, 5kb or more, 6kb or more, 7kb or more, 8kb or more, 9kb or more, 9.5kb or more, or 10kb or more. The donor DNA for selection markers is not particularly limited, but may have lengths of, for example, 50kb or less, 45kb or less, 40kb or less, 35kb or less, 30kb or less, 25kb or less, 20kb or less, 15kb or less, 14kb or less, 13kb or less, 12kb or less, 11kb or less, 10kb or less, 9kb or less, 8kb or less, 7kb or less, 6kb or less, 5kb or less, or 4kb or less.

[0047] A "selection marker" refers to a protein that can be used to select cells based on whether or not it is expressed. A selection marker gene is a gene that codes for a selection marker. In a cell population containing both cells expressing and not expressing a selection marker, when selecting cells that express the selection marker, the selection marker is called a "positive selection marker" or "selection marker for positive selection." In a cell population containing both cells expressing and not expressing a selection marker, when selecting cells that do not express the selection marker, the selection marker is called a "negative selection marker" or "selection marker for negative selection." Selection markers being different from each other means that they can be distinguished from each other (for example, they are distinguishably different), meaning that they can be distinguished from each other at least in physiological properties such as the drug resistance properties or other physicochemical properties that the selection marker confers to cells into which it has been introduced. In other words, when selection markers are different from each other, it means that multiple different selection markers can be detected distinguishably from other selection markers, or that drugs can be selected distinguishably from other selection markers. Furthermore, the statement that the selection marker gene is unique to each type of selection marker donor DNA means that a selection marker gene present in one type of selection marker donor DNA is not present in any other type of selection marker donor DNA, or, if present in multiple types of donor DNA, it is configured so that it is not expressed simultaneously from two or more types of donor DNA. In this case, the two or more types of donor DNA may be identical except for the selection marker, or they may differ in the sequence and / or composition other than the selection marker.

[0048] Positive selection markers are not particularly limited as long as they allow for the selection of cells that express them. Examples of positive selection marker genes include drug resistance genes, fluorescent protein genes, luminescent enzyme genes, and chromogenic enzyme genes.

[0049] Negative selection markers are not particularly limited as long as they can select cells that do not express them. Examples of negative selection marker genes include suicide genes (such as thymidine kinase), fluorescent protein genes, luminescent enzyme genes, and chromogenic enzyme genes. If a negative selection marker gene is a gene that negatively affects cell survival (e.g., a suicide gene), it can be functionally linked to an inductive promoter. By functionally linking to an inductive promoter, the negative selection marker gene can be expressed only when it is desired to remove cells that possess the negative selection marker gene. If the negative selection marker gene is an optically detectable marker gene (visible marker gene) such as a fluorescence, luminescence, or chromogenic gene, and has little negative impact on cell survival, it may be constitutively expressed.

[0050] Examples of drug resistance genes include, but are not limited to, puromycin resistance genes, blastisidin resistance genes, geneticin resistance genes, neomycin resistance genes, tetracycline resistance genes, kanamycin resistance genes, zeosin resistance genes, hygromycin resistance genes, and chloramphenicol resistance genes. Examples of fluorescent protein genes include, but are not limited to, the green fluorescent protein (GFP) gene, the yellow fluorescent protein (YFP) gene, and the red fluorescent protein (RFP) gene. Examples of luminescent enzyme genes include, but are not limited to, the luciferase gene. Examples of chromogenic enzyme genes include, but are not limited to, the β-galactosidase gene, the β-glucuronidase gene, and the alkaline phosphatase gene. Examples of suicide genes include, but are not limited to, the herpes simplex virus thymidine kinase (HSV-TK) and inducible caspase 9.

[0051] The selection marker gene contained in the donor DNA for selection markers is preferably a positive selection marker gene. That is, cells expressing the selection marker can be selected as cells in which the selection marker gene has been knocked in.

[0052] The upstream homology arm has a sequence that can homologously recombine with the sequence upstream of the target region in the genome to be modified, for example, a sequence homologous to the sequence adjacent to the upstream side of the target sequence. The downstream homology arm has a sequence that can homologously recombine with the sequence upstream of the target region in the genome to be modified, for example, a sequence homologous to the sequence adjacent to the downstream side of the target sequence. The length and sequence of the upstream and downstream homology arms are not particularly limited, as long as they can homologously recombine with the surrounding region of the target region. The upstream and downstream homology arms do not necessarily have to be perfectly identical to the upstream or downstream sequence of the target region, as long as homologous recombination is possible. For example, the upstream homology arm can be a sequence that has 90% or more sequence identity (homology) with the sequence adjacent to the upstream side of the target region, and it is preferable that it has 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more sequence identity. For example, the downstream homology arm can be a sequence having 90% or more sequence identity (homology) with a nucleotide sequence adjacent to the downstream side of the target region, and preferably has 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more sequence identity. Furthermore, the efficiency of allele modification can be further increased if at least one of the upstream homology arm and the downstream homology arm is closer to the cleavage site in or near the target region. Here, "close" may mean that the distance between the two sequences is 100 bp or less, 50 bp or less, 40 bp or less, 30 bp or less, 20 bp or less, or 10 bp or less. In some embodiments, there is one cleavage site. In other embodiments, there can be two or more cleavage sites. When introducing multiple cleavage sites into the genome, one cleavage site can be set upstream of the target region and the other can be set downstream of the target region. When the entire or nearly entire target region is excised from the genome by two cuts in the presence of a select marker donor DNA, the select marker donor DNA undergoes homologous recombination repair upstream and downstream of the target region, and the deleted region is replaced by the sequence of the select marker donor DNA.This approach increases the efficiency of recombination compared to a single cut. Furthermore, the efficiency of allele modification can be further enhanced if at least one, preferably both, of the upstream and downstream homology arms are closer to the cut site in or near the target region.

[0053] In the donor DNA for selection markers, the selection marker gene is located between the upstream and downstream homology arms. This means that when the donor DNA for selection markers is introduced into cells along with the genome modification system described in (i) above, the selection marker gene is introduced into the target region via HDR (if this results in gene disruption, it is called gene knockout; if this results in the introduction of a desired gene, it is called gene knock-in, which allows for the knockout of one gene while simultaneously knocking in another).

[0054] The selection marker gene is preferably functionally linked to a promoter so that it is expressed under the control of an appropriate promoter. The promoter can be appropriately selected depending on the type of cell into which the donor DNA is introduced. Examples of promoters include the SRα promoter, SV40 initial promoter, retroviral LTR, CMV (cytomegalovirus) promoter, RSV (Rous sarcoma virus) promoter, HSV-TK (herpes simplex virus thymidine kinase) promoter, EF1α promoter, metallothionein promoter, and heat shock promoter. The donor DNA for the selection marker may have any regulatory sequences such as enhancers, poly(A) addition signals, or terminators.

[0055] The donor DNA for selection markers may have an insulator sequence. An "insulator" is a sequence that blocks or mitigates the influence of the adjacent chromosomal environment, ensuring or enhancing the independence of transcriptional regulation of the DNA sandwiched within that region. Insulators are defined by their enhancer blocking effect (the effect of blocking the influence of an enhancer on promoter activity by inserting it between an enhancer and a promoter) and their positional effect suppression effect (the effect of preventing the expression of a transgene from being affected by its position on the genome where it is inserted by sandwiching both sides of the transgene with insulators). The donor DNA for selection markers may have an insulator sequence between the upstream arm and the selection marker gene (or between the upstream arm and the promoter that controls the selection marker gene). The donor DNA for selection markers may have an insulator sequence between the downstream arm and the selection marker gene.

[0056] The donor DNA for the selection marker may be linear or circular, but is preferably circular. Preferably, the donor DNA for the selection marker is a plasmid. In addition to the above sequence, the donor DNA for the selection marker may contain any other sequence. For example, spacer sequences may be included between all or part of the sequences of the upstream homology arm, insulator, selection marker gene, and downstream homology arm.

[0057] In step (a), the cells are introduced with a number of selection marker donor DNAs equal to or greater than the number of alleles targeted for genome modification. Different types of selection marker donor DNAs have different (distinguishable) types of selection marker genes. In some embodiments, different types of selection marker donor DNAs do not have completely identical selection marker genes or sets. That is, the first type of selection marker donor DNA has the first type of selection marker gene, the second type of selection marker donor DNA has the second type of selection marker gene, the third type of selection marker donor DNA has the third type of selection marker gene, and so on for subsequent types of selection marker donor DNA. If there are two alleles targeted for genome modification, there are two or more types of selection marker donor DNAs. If there are three alleles targeted for genome modification, there are three or more types of selection marker donor DNAs. In some embodiments, a single selection marker donor DNA may have two or more mutually distinct (distinguishable) selection markers (even in this case, different types of selection marker donor DNA must have mutually distinct (distinguishable) types (e.g., unique) selection marker genes). In some embodiments, the selection marker donor DNA does not have site-directed recombinase recombinant sequences (e.g., loxP sequences and their variants that are recombined by Cre recombinase). Also, in some embodiments, the method of the present invention does not use site-directed recombinase and its recombinant sequences (e.g., loxP sequences and their variants that are recombined by Cre recombinase). When site-directed recombinase is used, one site-directed recombinase recombinant sequence usually remains in the edited genome. In contrast, in some embodiments, the modified genome of cells obtained by the method of the present invention does not have site-directed recombinase recombinant sequences (which are foreign).

[0058] The number of types of donor DNA for selection markers should be equal to or greater than the number of alleles targeted for genome modification, and there is no particular upper limit. By using as many or more types of donor DNA for selection markers as the number of alleles targeted for genome modification, two or more alleles can be stably modified. From the viewpoint of the selection operation in step (b) described below, the number of types of donor DNA for selection markers is preferably equal to or 1 to 2 more than the number of alleles targeted for genome modification, and more preferably equal to the number of alleles targeted for genome modification.

[0059] The method for introducing (i) and (ii) into cells is not particularly limited, and known methods can be used without particular restriction. Examples of methods for introducing (i) and (ii) into cells include, but are not limited to, viral infection, lipofection, microinjection, calcium phosphate, DEAE-dextran, electroporation, and particle gun. By introducing (i) and (ii) into cells, the DNA in the target region is cleaved by the sequence-specific nucleic acid cleavage molecule of (i), and then the selection marker in the selection marker donor DNA of (ii) is knocked into the target region by HDR. In this case, if two or more selection marker donor DNAs have the same upstream homology arm and downstream homology arm, they can be randomly knocked into two or more alleles of the target region. However, two or more donor DNAs for selection markers can modify each of the two or more alleles as long as they each have homology arm sequences that are homologously recombinable with the upstream and downstream sequences of the target regions of each of the two or more alleles; therefore, they do not need to have completely identical homology arm sequences. In some embodiments, the upstream and downstream homology arm sequences of two or more donor DNAs for selection markers may have sequences that are more identical to the upstream and downstream sequences of the target regions of each allele (for example, they may be optimized in this way).

[0060] In one embodiment, the donor DNA for selection markers has an upstream homology arm and a downstream homology arm, and between the upstream and downstream homology arms, it has a selection marker gene, which may further preferably have a target sequence of an endonuclease (a sequence-specific nucleic acid cleavage molecule), such as a meganuclease cleavage site. In this embodiment, in one preferred embodiment, the selection marker includes a selection marker gene for positive selection and a marker gene for negative selection. In another preferred embodiment, the selection marker includes a selection marker for positive selection, but does not necessarily include a separate negative selection marker gene. In one preferred embodiment, the selection marker gene for positive selection may also be used for negative selection, and such a marker gene is a visualization marker gene. A set of two or more selection marker donor DNAs is a combination of the above selection marker donor DNAs, and each has a selection marker gene for positive selection that is distinguishable from each other. The above set may further have target sequences of endonucleases (nucleotide sequence-specific nucleic acid cleavage molecules), such as meganuclease cleavage sites, and these target sequences may be different from each other, but it is preferable that they are the same (or can be cleaved by the same nucleotide sequence-specific nucleic acid cleavage molecule). The length of the selection marker donor DNA is as described above, but for example, it may be 5kbp or longer, 8kbp or longer, or 10kbp or longer.

[0061] (Step (b)) After step (a) above, step (b) is performed. In step (b), cells in which two or more alleles each have distinctly different selection marker genes or combinations thereof are introduced are selected based on the expression of said distinctly different selection marker genes. More specifically, in step (b), cells are selected in which two or more alleles each have distinctly different unique selection marker genes introduced by homologous recombination of different types of selection marker donor DNA, and which express all of the introduced distinctly different selection marker genes. In one embodiment, in step (b), cells are selected in which each allele has been modified by the introduction of different selection marker donor DNA, based on the expression of all selection marker genes present in the two or more selection marker donor DNAs that have been integrated into the chromosomal genome. In another embodiment, in step (b), cells are selected based on all of the selection marker genes present in the two or more selection marker donor DNAs. In one embodiment, step (b) selects cells in which each allele has been modified by the introduction of distinguishable selection marker donor DNA, based on the expression of all selection marker genes (positive selection marker genes) that are incorporated into the chromosomal genome and are present in the two or more selection marker donor DNAs. In one embodiment, the cells obtained in step (b) have different positive selection marker genes for each allele. In one embodiment, the cells obtained in step (b) have a common positive selection marker gene for each allele. Here, in one embodiment, single-cell cloning is not performed in step (b) {however, this may or may not include single-cell cloning after selecting cells in which two or more alleles have been modified in step (b)}. In one embodiment, cell selection in step (b) is performed based on the expression of multiple distinguishable positive selection marker genes introduced into each allele. In one embodiment, step (b) is not performed in a manner that estimates the number of modified alleles based on the expression intensity of a single selection marker gene (e.g., the expression intensity or fluorescence intensity of a fluorescent protein).When selecting cells using a method that estimates the number of modified alleles based on the strength of expression of a single selection marker gene, variations in gene expression levels occur in each cell, making it difficult to completely isolate cells with two or more modified alleles from cells with only one modified allele. Therefore, single-cell cloning becomes necessary in step (b).

[0062] Step (b) should involve selecting cells as appropriate, depending on the type of selection marker gene used in step (a). In this case, cells should be selected based on the expression of all the selection marker genes used in step (a).

[0063] For example, if the selection marker gene is a positive selection marker gene, cells expressing all selection marker genes incorporated into (or already incorporated into) the chromosome genome to be modified can be selected. For example, cells expressing the same number of positive selection markers as the number of alleles to be modified can be selected. If the positive selection marker gene is a drug resistance gene, cells expressing the positive selection marker can be selected by culturing the cells in a medium containing the drug. If the positive selection marker gene is a fluorescent protein gene, a luminescent enzyme gene, or a chromogenic enzyme gene, cells expressing the positive selection marker can be selected by selecting cells that exhibit fluorescence, luminescence, or color development due to the fluorescent protein, luminescent enzyme, or chromogenic enzyme. In this process, if the same number of selection marker donor DNAs as the number of alleles to be modified are incorporated into the genome, that number of alleles are modified. In n-ploid cells, the number of alleles to be modified is n or less, and if the number of selection marker donor DNAs between n and n is incorporated into the genome, at least the alleles to be modified (two or more alleles) are modified. In one embodiment, the number of alleles to be modified is n, and this number of types of selection marker donor DNA are incorporated into the chromosomal genome, thereby modifying all alleles. In another embodiment, since this step uses the same number or more types of selection marker donor DNA as the number of alleles to be modified, the number of positive selection markers expressed by the cells means that this number of alleles have been reliably modified. From the viewpoint of improving the cell selection efficiency in step (b), it is preferable that the number of alleles to be modified is the same as the number of types of selection marker donor DNA.

[0064] As described above, in the genome modification method of this embodiment, by inducing HDR using n types of selection marker donor DNA to modify n alleles in n-ploid cells, cells in which all alleles have been modified can be efficiently obtained. Furthermore, since cells in which all alleles have been modified can be reliably obtained, even if the target region is large (e.g., 10 kbp or larger), cells in which the target region has been modified can be efficiently obtained. Therefore, large-scale genome modification becomes possible.

[0065] In one embodiment, in step (b), modified cells can be selected from a pool containing cells obtained in step (a) without cloning the cells. By omitting the cloning step, the time required for the step can be reduced. In one embodiment, the pool is 10 5 The above 10 6 The above 10 7 or more, or 10 8 The above cells may be included.

[0066] (Optional process) The genome modification method of this embodiment may include any additional steps in addition to steps (a) and (b) described above. Examples of optional steps include steps (c) and (d) below: (c) After step (b), a step of introducing recombinant donor DNA containing a desired base sequence into the cell between an upstream homology arm having a base sequence homologous to the base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to the base sequence adjacent to the downstream side of the target region; and (d) A step of selecting cells that do not express the negative selection marker after step (c).

[0067] In some embodiments, the genome modification method of this embodiment may include any additional steps in addition to steps (a) and (b) described above. In some embodiments, the genome modification method of this embodiment or the method for obtaining cells with a modified genome may include two or more selection marker donor DNAs, each having a selection marker gene for positive selection, a separate marker gene for negative selection, and a target sequence between the upstream homology arm and the downstream homology arm, wherein if the selection marker gene is used for both positive and negative selection, it does not need to have the separate selection marker gene for negative selection, and for example, the method may include steps (c) and (d) below: (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets and can cleave the further target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired nucleotide sequence, comprising an upstream homology arm having a nucleotide sequence homologously recombinable with the nucleotide sequence upstream of the target region, and a downstream homology arm having a nucleotide sequence homologously recombinable with the nucleotide sequence downstream of the target region {the recombinant donor DNA may or may not contain the desired nucleotide sequence between the upstream homology arm and the downstream homology arm}. (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection).

[0068] <Process (c)> Step (c) may be performed after step (b). In one embodiment, step (c) involves introducing recombinant donor DNA containing or not containing the desired nucleotide sequence between an upstream homology arm and a downstream homology arm into the cells selected in step (b). In one embodiment, step (c) involves introducing recombinant donor DNA containing the desired nucleotide sequence between an upstream homology arm having a nucleotide sequence homologous to the nucleotide sequence adjacent to the upstream side of the target region and a downstream homology arm having a nucleotide sequence homologous to the nucleotide sequence adjacent to the downstream side of the target region into the cells selected in step (b).

[0069] Recombinant donor DNA Recombinant donor DNA may contain the desired nucleotide sequence to be knocked in. The desired nucleotide sequence is not particularly limited. For example, if the purpose of genome modification is to knock out the function of a gene contained in a target region, a nucleotide sequence in which part or all of the nucleotide sequence of the target region is deleted can be used as the desired nucleotide sequence. Also, if an exogenous gene is to be incorporated into the target region, a nucleotide sequence containing that gene can be used as the desired nucleotide sequence. The size of the desired nucleotide sequence is not particularly limited and can be any size. The desired base sequence can be, for example, 10 bp or larger, 20 bp or larger, 40 bp or larger, 80 bp or larger, 200 bp or larger, 400 bp or larger, 800 bp or larger, 1 kbp or larger, 2 kbp or larger, 3 kbp or larger, 4 kbp or larger, 5 kbp or larger, 6 kbp or larger, 7 kbp or larger, 8 kbp or larger, 9 kbp or larger, 10 kbp or larger, 15 kbp or larger, 20 kbp or larger, 40 kbp or larger, 80 kbp or larger, 100 kbp or larger, or 200 kbp or larger. The method of this embodiment allows for efficient selection of cells in which the desired base sequence has been knocked into two or more alleles. Therefore, even large DNAs of, for example, 5 kbp or larger, 8 kbp or larger, or 10 kbp or larger can be knocked in. The recombinant donor DNA may be, for example, shorter in length than the selection marker donor DNA.

[0070] The upstream and downstream homology arms of the recombinant donor DNA may be the same as or different from those of the selection marker donor DNA. For convenience, the upstream and downstream homology arms included in the selection marker donor DNA may be referred to as the "first upstream homology arm" and the "first downstream homology arm," and the upstream and downstream homology arms included in the recombinant donor DNA may be referred to as the "second upstream homology arm" and the "second downstream homology arm." The length and sequence of the second upstream and second downstream homology arms are not particularly limited, for example, as long as they are homologously recombinable with the first upstream homology arm or a region upstream thereof, and homologously recombinable with the first downstream homology arm or a region downstream thereof (in some embodiments, their length and sequence are not particularly limited as long as they are homologously recombinable with the region surrounding the target region). After recombination with recombinant donor DNA, it is permissible for some of the base sequence of the selective marker donor DNA to remain on the genome, but preferably, the base sequence of the selective marker donor DNA is completely removed from the genome by recombination with the recombinant donor DNA. Various genes carried on the selective marker donor DNA are removed by recombination with the recombinant donor DNA. In some embodiments, this allows two or more alleles of a cell to be replaced by the recombinant donor DNA. In some embodiments, the recombinant donor DNA may have a desired base sequence, so that a cell in which two or more alleles have been modified will have the desired base sequence in the modified alleles.

[0071] In recombinant donor DNA, the desired base sequence is located between the second upstream homology arm and the second downstream homology arm. If the recombinant donor DNA contains a foreign gene, it is preferable that the foreign gene is functionally ligated to the promoter. The recombinant donor DNA may have any regulatory sequences such as enhancers, poly(A) addition signals, or terminators. Furthermore, if the recombinant donor DNA contains a foreign gene, it may have insulator sequences upstream and downstream of the foreign gene. In one embodiment, the recombinant donor DNA includes a spacer sequence between the second upstream homology arm and the second downstream homology arm. In one embodiment, when removing cells that have a negative selection marker gene contained in the selection marker donor DNA, if a gene identical (or indistinguishable) to the said gene is expressed under conditions in which its toxicity is exerted, it is not possible to select cells in which homologous recombination has occurred with the recombinant donor DNA. Therefore, the recombinant donor DNA is configured such that, when cells possessing the negative selection marker gene of the selection marker donor DNA are removed, the same (or indistinguishable) gene as that gene is not expressed under conditions in which its toxicity is exerted. For example, in one embodiment, the recombinant donor DNA does not have a negative selection marker gene and a second target sequence between the second upstream homology arm and the second downstream homology arm.

[0072] It is preferable that the recombinant donor DNA is introduced into the cells together with (i) above. By introducing the recombinant donor DNA into the cells together with (i) above, the DNA in the target region is cleaved by the sequence-specific nucleic acid cleavage molecule in (i) above, and then the desired base sequence in the recombinant donor DNA is knocked into the target region by HDR. Since the cells into which the recombinant donor DNA is introduced in this step are the cells selected in step (b), the base sequence of the selection marker donor DNA has been knocked into the target region. Therefore, the target sequence of the genome modification system in (i) is the base sequence contained in the target region after the selection marker donor DNA knock-in. For convenience, the target sequence of the genome modification system in step (a) may be referred to as the "first target sequence," and the target sequence of the genome modification system in step (c) may be referred to as the "second target sequence." The second target sequence can be any sequence contained in the target region of the cell after step (b). In some embodiments, the second target sequence in the selection marker donor DNA may be a sequence that does not exist on the genome of the cell in question. In one embodiment, the second target sequence in the selection marker donor DNA is a sequence not present in the cell's genome and is distinct from other sequences to the extent that it does not cleave other genomic sequences due to off-target reactions. In another embodiment, the second target sequence in the selection marker donor DNA may be a meganuclease cleavage site not present in the genome. In yet another embodiment, the second target sequence is a region other than the negative selection marker gene in step (d). Naturally, the recombinant donor DNA is configured such that homologous recombination by the recombinant donor DNA is not significantly inhibited. If the first target sequence remains in the target region of the cell after step (b) or if the first target sequence is reintroduced by the selection marker donor DNA, the second target sequence may be the same as or different from the first target sequence.

[0073] Recombinant donor DNA does not need to contain a base sequence between the upstream homology arm and the downstream homology arm, but it may contain a base sequence of 10 bp or less, 20 bp or less, 30 bp or less, 40 bp or less, 50 bp or less, 60 bp or less, 70 bp or less, 80 bp or less, 90 bp or less, 100 bp or less, 200 bp or less, 300 bp or less, 400 bp or less, 500 bp or less, 600 bp or less, 700 bp or less, 800 bp or less, 900 bp or less, or 1 kbp or less between the upstream and downstream homology arms. Recombinant donor DNA may contain a base sequence of 1 kbp or more, 2 kbp or more, 3 kbp or more, 4 kbp or more, 5 kbp or more, 6 kbp or more, 7 kbp or more, 8 kbp or more, 9 kbp or more, or 10 kbp or more between the upstream and downstream homology arms.

[0074] Recombination donor DNA contains, or does not contain, one or more or all of, selected from the group consisting of a selection marker gene, a target sequence for site-directed recombinant enzyme, a gene encoding a physiologically active factor, a gene encoding a cytotoxic factor, and a promoter sequence between the upstream homology arm and the downstream homology arm.

[0075] In step (c), recombinant donor DNA is introduced into the cells selected in step (b). The cells selected in step (b) have a selection marker gene knocked into the target region. Step (c) can also be described as a step of removing the selection marker gene that has been knocked into the target region or replacing it with a desired base sequence.

[0076] (Step (d)) After step (c), step (d) may be performed. In step (d), cells that do not express the negative selection marker are selected.

[0077] When performing step (d), the donor DNA for selection markers used in step (a) may each contain a positive selection marker gene and a negative selection marker gene. That is, the donor DNA for selection markers used in step (a) may contain a positive selection marker gene and a negative selection marker gene between the upstream homology arm and the downstream homology arm. The positional relationship between the positive selection marker gene and the negative selection marker gene is not particularly limited; the positive selection marker gene may be upstream of the negative selection marker gene, or vice versa. When the donor DNA for selection markers contains a positive selection marker gene and a negative selection marker gene, a nucleotide sequence encoding a self-cleaving peptide or an IRES (internal ribozyme entry site) sequence may be interposed between the positive selection marker gene and the negative selection marker gene. By interposing these sequences, the positive selection marker gene and the negative selection marker gene can be expressed independently from a single promoter. Examples of 2A peptides include 2A peptide derived from foot-and-mouth disease virus (FMDV) (F2A), 2A peptide derived from equine rhinitis A virus (ERAV) (E2A), 2A peptide derived from Porcine teschovirus (PTV-1) (P2A), and 2A peptide derived from Thosea asigna virus (TaV) (T2A).

[0078] Alternatively, the same selection marker gene may be used as a positive selection marker in step (a) and as a negative selection marker in step (d). For example, if the selection marker gene is a marker gene involved in fluorescence or color development such as a fluorescent protein gene, a luminescent enzyme gene, or a chromogenic enzyme gene (visualization marker gene), then in step (a), cells exhibiting fluorescence, emission, or color development due to the expression of the fluorescent protein, luminescent enzyme, or chromogenic enzyme may be selected, and in step (c), cells in which these fluorescence, emission, or color development have disappeared may be selected. The case where the same selection marker gene serves as both a positive and negative selection marker is also included when the selection marker donor DNA further contains a negative selection marker in addition to the positive selection marker.

[0079] The negative selection marker gene may be different or the same for each type of donor DNA used as a selection marker. Using a common negative selection marker gene simplifies the cell selection process in step (d).

[0080] Step (d) should involve selecting cells appropriately according to the type of negative selection marker gene used in step (a). In this case, cells that do not express any of the negative selection marker genes used in step (a) should be selected.

[0081] For example, when the negative selection marker gene is a visualization marker gene such as a fluorescent protein gene, a luminescent enzyme gene, or a chromogenic enzyme gene, cells in which the visualization marker such as fluorescence, luminescence, or chromogenesis has disappeared may be selected. Further, when the negative selection marker gene is a suicide gene, cells that do not express the negative selection marker can be selected by culturing the cells in a medium containing a drug that exhibits toxicity due to the expression of the suicide gene. For example, when the thymidine kinase gene is used as the suicide gene, the cells may be cultured in a medium containing ganciclovir. The disappearance of the expression of the negative selection marker gene means that the negative selection marker gene integrated into the target region in step (a) has been replaced with a polynucleotide containing the desired nucleotide sequence of the donor DNA for recombination. At this time, it is considered that the replacement of the polynucleotide occurs in the entire nucleotide sequence knocked in in step (a). Therefore, by selecting cells in which the expression of the negative selection marker gene has disappeared, cells in which the nucleotide sequence knocked in in step (a) has been replaced with the desired nucleotide sequence of the donor DNA for recombination can be efficiently selected. The negative selection marker gene such as a suicide gene may be functionally linked to an inducible promoter, and the cells are cultured in the presence of a drug that drives the inducible promoter so that the negative selection marker gene is expressed under conditions where its toxicity is exerted, whereby cells that do not express the negative selection marker can be selected. In this case, the negative selection marker gene may be a gene encoding a cytotoxin (e.g., ricin and diphtheria toxin) that causes toxicity to cells only when expressed.

[0082] In one aspect, in step (d), from the pool containing the cells obtained in step (c), without cloning the cells, cells in which two or more alleles have been modified (cells in which the negative selection marker gene is absent) can be selected. In one aspect, the above pool contains 10 5 or more, 10 6 or more, 10 7 or more, or 10 8 or more cells may be included.

[0083] As described above, by performing steps (c) and (d), cells in which all alleles have been modified to the desired sequence can be efficiently obtained. Furthermore, since cells in which all alleles have been modified can be reliably obtained, even if the desired base sequence is large in size (e.g., 10 kbp or more), cells in which the base sequence has been knocked into the target region can be efficiently obtained. In one embodiment, by performing steps (c) and (d), the target region is deleted in all alleles of the cell, and the sequences before and after the deletion (i.e., the sequences in which the upstream homology arm and downstream homology arm undergo homologous recombination, respectively) are seamlessly linked without one or more selected from the group consisting of base insertions, substitutions, and deletions (e.g., without base insertions, substitutions, and deletions). In another embodiment, in the obtained cells, the base sequences on the upstream and downstream sides of the deleted region are seamlessly linked.

[0084] Furthermore, if the number of viable cells is small or no viable cells are obtained in step (b), it is shown that the upstream homology arm and the target region removed from the genome by homologous recombination contain genes that affect cell proliferation or survival. Therefore, it is possible to investigate whether the target region contains genes that affect cell proliferation or survival. In this case, by changing the design positions of the upstream and downstream homology arms, the genes that are removed from the genome by homologous recombination can be changed to identify genes that affect cell proliferation or survival. Therefore, the present invention allows step (e) to be performed after step (b). That is, step (e) includes identifying genes that affect cell proliferation or survival by narrowing the target region and reducing the number of genes removed from the genome in step (b) if the number of viable cells is small or no viable cells are obtained. If the target region contains only one gene, it is shown that this gene is a gene that affects cell proliferation or survival. Then, if a gene that affects cell proliferation or survival has been identified, step (f) can be performed. Step (f) includes knocking in a gene that affects the proliferation or survival of the identified cells to another region of the genome (e.g., a safe harbor region) {recombinant donor DNA may be used for the knock-in}. This expands the region removed by the method of the present invention (extending the target region upstream and / or downstream). The low number of viable cells can be confirmed by comparing it to the number of cells obtained when steps (a) and (b) are performed on a region that does not affect cell survival or proliferation. In some embodiments, step (a) may not target a region that would cause loss of cell proliferation or survival.

[0085] [Genome modification kit] In one embodiment, the present invention provides a genome modification kit for modifying two or more alleles of a chromosomal genome. The genome modification kit comprises (i) and (ii) below: (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule; and (ii) Two or more selection marker donor DNAs, each containing the base sequence of a selection marker gene between a downstream homology arm having a base sequence homologous to a base sequence adjacent to the upstream side of the target region and a downstream homology arm having a base sequence homologous to a base sequence adjacent to the downstream side of the target region, wherein the two or more selection marker donor DNAs each contain a different selection marker gene, and the number of types of selection marker donor DNAs is equal to or greater than the number of alleles targeted for genome modification.

[0086] In one embodiment, the present invention relates to a genome modification kit for modifying two or more alleles of a chromosomal genome, comprising (i) and (ii) below. (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving a target region of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule, (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, with the base sequence of the select marker gene included between the upstream and downstream homology arms, the two or more types of select marker donor DNA each having a select marker gene that is distinguishable from each other, and the number of types of select marker donor DNA being equal to or greater than the number of alleles targeted for genome modification. In this embodiment, the select marker gene may be unique to each type of select marker donor DNA. The above kit may be used in the method of the present invention. The above kit may be used in a method for producing cells in which two or more alleles of the chromosomal genome have been modified.

[0087] The kit of this embodiment includes (i) and (ii), which are the same as those described in the section on [Genome Modification Method] above. The genome modification method described above can be easily performed by using the kit of this embodiment.

[0088] In one embodiment, the present invention provides cells in which two or more alleles of a chromosomal genome have been modified, and each of the two or more alleles has a mutually distinct (distinguishable) selection marker gene. In one embodiment, the cells may be cells of a single-celled organism. In one embodiment, the cells may be isolated cells. In one embodiment, the cells may be cells selected from the group consisting of pluripotent cells and pluripotent stem cells (such as embryonic stem cells and induced pluripotent stem cells). In one embodiment, the cells may be tissue stem cells. In one embodiment, the cells may be somatic cells. In one embodiment, the cells may be germline cells (e.g., germ cells). In one embodiment, the cells may be cell lines. In one embodiment, the cells may be immortalized cells. In one embodiment, the cells may be cancer cells. In one embodiment, the cells may be non-cancer cells. In one embodiment, the cells may be cells of a diseased patient. In one embodiment, the cells may be cells of a healthy person. In one embodiment, the cells may be selected from a group consisting of animal cells (e.g., human cells), insect cells (e.g., silkworm cells), HEK293 cells, HEK293T cells, Expi293F® cells, FreeStyle® 293F cells, Chinese hamster ovary cells (CHO cells), CHO-S cells, CHO-K1 cells, and ExpiCHO cells, as well as cells derived from these cells. In one preferred embodiment, in the above cells, all alleles of the target region of the chromosomal genome are modified, and each modified region has a distinct (distinguishable) selection marker gene.

[0089] In one embodiment, a method for culturing cells is provided in which two or more alleles of the chromosomal genome have been modified, and each of the two or more alleles has mutually different (distinguishable) selection marker genes. If the selection marker genes are drug resistance marker genes, the culture may be carried out in the presence of drugs for each drug resistance marker gene. The culture may be carried out under conditions suitable for cell maintenance or proliferation.

[0090] In one embodiment, the present invention provides a non-human organism having a chromosomal genome in which two or more alleles are modified, wherein each of the two or more alleles has a mutually distinct (distinguishable) selection marker gene. In one embodiment, the cell may be a cell of a single-celled organism. In one embodiment, the non-human organism may be a yeast (e.g., fission yeast or budding yeast, e.g., Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces fragilis, Saccharomyces rouxii, etc., Candida utilis, Candida tropicalis) The organism may be selected from the group consisting of yeasts such as Candida (e.g., tropicalis), Pichia, Kluyveromyces, Yarrowia, Hansenula, and Endomyces. In one embodiment, the non-human organism may be a filamentous fungus (e.g., Aspergillus, Trichoderma, Humicola, Acremonium, Fusarium, and Penicillium species). In one embodiment, the non-human organism may be a multicellular organism. In one embodiment, the non-human organism may be a non-human animal. In one embodiment, the non-human organism may be a plant. In one preferred embodiment, in the above non-human organism, all alleles of a target region of the chromosomal genome are modified, and the modified regions each have a different (distinguishable) selection marker gene.

[0091] In the cells described above, one or more genes necessary for cell survival and proliferation may be located in or clustered in other regions of the chromosomal genome. These other regions may be, for example, safe harbor regions (e.g., the AAVS1 region). [Examples]

[0092] The present invention will be explained below with reference to experimental examples, but the present invention is not limited to the following examples.

[0093] [Experimental Example 1] Donor DNA production To select cells edited with two alleles, two types of donor DNA with different positive selection markers were created (Figure 1; Puromycin). R plasmid, Blasticidin R Plasmid). The GFP gene was used as a negative selection marker. Puromycin R In the plasmid, the GFP gene is ligated downstream of the EF1 promoter, and the puromycin resistance gene is ligated across the T2A sequence. R In the plasmid, the GFP gene is ligated downstream of the EF1 promoter, and the blastisidin resistance gene is ligated across the T2A sequence.

[0094] The HR110PA-1 plasmid (Funakoshi) was used as the backbone sequence for the donor plasmid (donor DNA). This HR110PA-1 plasmid was cleaved with the restriction enzymes EcoRI and BamHI to extract the sequence containing the origin of replication and the ampicillin resistance gene, and then purified. This purified sequence was used as the backbone sequence for all donor plasmids used in the future.

[0095] The plasmid selection marker sequences and homology arm sequences were amplified by the following PCR. The EF1 promoter sequence was constructed by amplifying from the EcoRI recognition sequence on HR110PA-1 to the EF1 promoter sequence. The GFP sequence was constructed using the CS-CDF-CG-PRE plasmid (dnaconda.riken.jp / search / RDB_clone / RDB04 / RDB04379.html) as a template. The puromycin expression sequence from the T2A sequence was constructed by amplifying from the T2A sequence on HR110PA-1 to the BamHI recognition sequence. The upstream homology arm (810 bp: SEQ ID NO: 2) and downstream homology arm (792 bp: SEQ ID NO: 3) were constructed using the HCT116 genome as a template.

[0096] During PCR, 20 bp homologous sequences corresponding to adjacent sequences were added to the selection marker sequence and homology arm sequence. These homologous sequences were used to ligate the backbone sequence, selection marker sequence, and homology arm sequence using NEBuilder HiFi DNA Assembly (NEW ENGLAND BioLabs). The resulting DNA was then introduced into E. coli and cultured in ampicillin-supplemented medium to clone the puromycin plasmid.

[0097] The blastisidin plasmid was prepared by modifying the puromycin plasmid. The blastisidin ORF sequence, amplified by PCR, and the sequence obtained by amplified everything except the puromycin ORF sequence using the puromycin plasmid as a template were joined using NEBuilder HiFi DNA Assembly, and cloning was performed in the same manner as when the puromycin plasmid was prepared.

[0098] [Experimental Example 2] Introduction of a selection marker to the first intron of the TP53 gene (Cells and target regions) The cells used were HCT116 cells, a cell line derived from colorectal cancer. Most of these cells are diploid. The first intron of the tumor suppressor gene TP53 was selected as the target region for genome editing.

[0099] (Step (a)) 10 5 To each HCT116 cell, Puromycin was used as the donor plasmid (donor DNA). R plasmid and blastidin R Plasmids (12.5 ng each) were introduced.

[0100] The day before donor plasmid introduction, 10 per well in a 24-well plate. 5 HCT116 cells were seeded. For culture, 500 μL of McCoy's 5A medium (with FBS added to make a 10% concentration) was used per well.

[0101] The day after seeding the cells, a plasmid introduction solution was prepared with the following composition. Opti-MEM 100μL Cas9 & gRNA co-expression plasmid (pX330-U6-Chimeric_BB-CBh-hSpCas9;Addgene, plasmid number 42230) 475ng Puromycin R plasmid 12.5 ng Blasticidin R plasmid 12.5 ng FuGENE HD 1.5μL

[0102] The target sequences of the gRNAs used in the Cas9 & gRNA co-expression plasmid are shown below. CTCAGAGGGGGCTCGACGCTAGG (Sequence No. 4)

[0103] After mixing the plasmid introduction solution described above by pipetting, it was incubated at room temperature for 10 minutes and then added to a 24-well plate.

[0104] (Step (b)) After introducing both plasmids, the cells were cultured for 8 days in McCoy's 5A medium supplemented with 1 μg / mL puromycin and 10 μg / mL blatisidin (with FBS added to make a 10% concentration). After culturing, cells were collected from independent colonies by pipetting.

[0105] (Confirmation of knock-in of selected marker gene) DNA was extracted from the 27 recovered cell clones using phenol / chloroform extraction, and junction PCR was performed to confirm the knock-in of the selected marker gene.

[0106] HCT116 cells were collected in a 1.5 mL tube and centrifuged at 150 g for 3 minutes at room temperature. The supernatant was removed and the cells were resuspended in 173 μL of TE buffer (100 mM NaCl). After incubation at 95°C for 5 minutes, the cells were cooled to room temperature. 20 μL of TE buffer (100 mM NaCl, 0.5% SDS), 5 μL of Proteinase K (5 mg / ml), and 2 μL of RNase A (100 mg / ml) were added and mixed, then inverted and stirred at 37°C for 2 hours. After incubation at 85°C for 25 minutes, 200 μL of phenol / chloroform / isoamyl alcohol (25:24:1) was added and inverted and stirred for 10 minutes. After centrifugation at 15,000 g for 10 minutes at room temperature, the upper layer was transferred to a new 1.5 mL tube. The same amount of chloroform as the transferred upper layer was added to the 1.5 mL tube and inverted and stirred for 10 minutes. After centrifugation at 15,000 g for 10 minutes at room temperature, the upper layer was transferred to a new 1.5 mL tube. An equal volume of isopropanol was added to the 1.5 mL tube and centrifuged at 15,000 g for 30 minutes at 4°C. The supernatant was removed, 150 μL of 70% ethanol was added, and the mixture was centrifuged at 15,000 g for 5 minutes at 4°C. The supernatant was completely removed, and the precipitate was dissolved in 30 μL of TE buffer.

[0107] Using the DNA recovered as described above as a template, junction PCR was performed using the junction primers shown in Figure 2. The junction primers are primers that amplify across the junctions on both the 5' and 3' sides of the puromycin resistance gene or the blasticidin resistance gene. Agarose gel electrophoresis of the PCR products was performed to confirm the amplified DNA fragments. The PCR conditions and primer sequences are shown below.

[0108] <PCR Conditions> Composition of PCR Solution (Reaction System 25 μL) 1x KOD one (TOYOBO) 0.2 μM Forward Primer 0.2 μM Reverse Primer 10 ng Genomic DNA

[0109] Thermal Cycler Conditions 98°C, 30 seconds 98°C, 10 seconds; 68°C, 20 seconds for 50 cycles 4°C

[0110] <Primer Sequences> PCR Primer for Puromycin 5’ junction Forward Primer: TGCCCCGTTGTTATCCTTAC (SEQ ID NO: 5) Reverse Primer: GCTCGTAGAAGGGGAGGTTG (SEQ ID NO: 6) PCR Primer for Puromycin 3’ junction Forward Primer: GTCACCGAGCTGCAAGAAC (SEQ ID NO: 7) Reverse Primer: GAAGACGGCAGCAAAGAAAC (SEQ ID NO: 8) PCR Primer for Blasticidin 5’ junction Forward Primer: TGCCCCGTTGTTATCCTTAC (SEQ ID NO: 9) Reverse Primer: GCTTCAATATGTACTGCCGAAA (SEQ ID NO: 10) PCR primers for Blasticidin 3' junction Forward primer: GAAGCCATTGCGATTGGTAG (SEQ ID NO: 11) Reverse primer: GAAGACGGCAGCAAAGAAAC (SEQ ID NO: 12)

[0111] The results are shown in Figure 3. In DNA extracted from the wild-type HCT116, no bands of the expected size were detected for either the puromycin resistance gene or the blastisidin resistance gene (far right lane: WT). On the other hand, in DNA extracted from cell clones obtained by steps (a) and (b) above, bands of the expected size were detected for both the puromycin resistance gene and the blastisidin resistance gene in 12 out of 27 clones (clones shown underlined). This result indicates that in these 12 clones, one allele of the first intron of TP53 was replaced with the puromycin resistance gene, and the other allele was replaced with the blastisidin resistance gene. For clone 13, sequence analysis of the PCR product was performed. As a result, both sequences in which the first intron of TP53 was replaced with the puromycin resistance gene and sequences in which the first intron of TP53 was replaced with the blastisidin resistance gene were confirmed.

[0112] [Example experiment] Blasticidin R Except for not using plasmid, the HCT116 cells were treated with puromycin using the same method as above. R A plasmid was introduced.

[0113] Puromycin R After plasmid introduction, cells were cultured for 8 days in McCoy's 5A medium supplemented with 1 μg / mL puromycin (with FBS added to make a 10% concentration). After culturing, cells were collected from independent colonies by pipetting.

[0114] DNA was extracted from the 10 recovered cell clones using the same method as described above. Junction PCR was performed on the puromycin resistance gene using a junction primer (see Figure 4, left) that amplified the gene across both the 5' and 3' junctions, as described above. The amplified DNA fragments were identified by agarose gel electrophoresis of the PCR products.

[0115] As a result, no clones were obtained in which both alleles in the first intron of TP53 were replaced with puromycin resistance genes. Furthermore, the efficiency of obtaining clones in which only one allele was replaced with a puromycin resistance gene was also low (Figure 4, right).

[0116] [Experimental Example 3] Remove selection markers (Process (c)) A plasmid containing the wild-type TP53 sequence (wild-type TP53 plasmid) was constructed as donor DNA. The wild-type TP53 plasmid was introduced into the cell clone obtained in Experimental Example 2 (clone #13 in Figure 3) using the same method as described above, except that a plasmid containing the sequence of the first intron and its surrounding regions (5' arm and 3' arm) of wild-type TP53 was used as the donor DNA. The target sequence of the gRNA used in the Cas9&gRNA co-expression plasmid is shown below. GGCGCAACGCGATCGCGTAAGGG (Sequence No. 13)

[0117] (Step (d)) After introducing the wild-type TP53 plasmid, cell clones with lost GFP expression were selected using a cell sorter (SH800, Sony), and each cell was collected in a 96-well plate.

[0118] (Confirmation of marker gene removal) DNA was extracted from the recovered cells using the same method as described above. PCR was performed using the primer shown in Figure 5. Agarose gel electrophoresis of the PCR product was performed to confirm the amplified DNA fragment. The composition of the PCR solution was the same as described above. The thermal cycler conditions and primer sequences are shown below.

[0119] Thermal cycler conditions 98℃, 30 seconds 98°C, 10 seconds; 68°C, 2 minutes 30 seconds, 50 cycles 4℃

[0120] <Primer Sequence> TP53 PCR primer for amplification of the first intron Forward primer: TGCCCCGTTGTTATCCTTAC (Sequence No. 14) Reverse primer: GAAGACGGCAGCAAAGAAAC (SEQ ID NO: 15)

[0121] The results are shown in Figure 6. Of the 8 clones, bands of the expected size were detected in 7 clones (clones underlined). This result indicates that in these 7 clones, the selection markers (puromycin resistance gene or blastisidin resistance gene, and GFP gene) were replaced with the first intron sequence of the wild-type TP53 gene. Furthermore, in these 7 clones, the 5.6kb band that would be detected if the selection marker remained in one allele was not detected. This result suggests that the selection markers were replaced with the wild-type first intron in both alleles.

[0122] [Experimental Example 4] Modification of the first intron of the TP53 gene Steps (c) and (d) were performed in the same manner as in Experimental Example 3, except that the plasmids shown in Figures 7A to D were used as donor DNA.

[0123] DNA was extracted from the recovered cells using the same method as described above. PCR was performed using the primers shown in Figures 7A-D. Agarose gel electrophoresis of the PCR products was performed to confirm the amplified DNA fragments. The composition of the PCR solution was the same as described above. The thermal cycler conditions and primer sequences are shown below.

[0124] Thermal cycler conditions 98℃, 30 seconds 98°C, 10 seconds; 68°C, 2 minutes 30 seconds, 50 cycles 4℃

[0125] <Primer Sequence> TP53 PCR primer for amplification of the first intron Forward primer: TGCCCCGTTGTTATCCTTAC (Sequence No. 14) Reverse primer: GAAGACGGCAGCAAAGAAAC (SEQ ID NO: 15)

[0126] The results are shown in Figure 8. In all cases using donor DNA, cell clones were obtained with a high probability in which both alleles were replaced with sequences from the donor DNA (clones shown underlined).

[0127] The first intron of the human TP53 gene is a relatively large genomic region of 10,762 bp. The method described above not only allows for efficient genome editing of both alleles, but also demonstrates the ability to efficiently edit even such a relatively large genomic region.

[0128] [Example 5] In Example 4, deletion of the first intron of the human TP53 gene was attempted in HCT116 cells. In this example, however, instead of HCT116 cells, deletion of the first intron of the human TP53 gene was attempted in human pluripotent stem cells. Specifically, human induced pluripotent stem cells (iPS cells) were used.

[0129] In this example, as shown in Figure 14A, cleavage was induced using a CRISPR / Cas9 system with gRNAs designed to specifically induce cleavage upstream and downstream of the target gene in the presence of a selection marker donor DNA. The cleavage sites were designed between the region where the upstream homology arm hybridizes and the target gene, and between the region where the downstream homology arm hybridizes and the target gene. This induced the first step of recombination. The sequence of the gRNA used was as follows.

[0130] <Guide RNA> Human TP53 upstream gRNA: CUCAGAGGGGGCUCGACGCU (Sequence ID 16) Human TP53 downstream gRNA: GGUGCUUUAAGAAUUACCGC (Sequence ID 17)

[0131] In the presence of the two types of selection marker donor DNA used in Example 4, cuts were made upstream and downstream of the first intron of the human TP53 gene in the genome of human iPS cells. Subsequently, cells with genomes in which the puromycin resistance gene and the blastisidin resistance gene were introduced into each allele in the presence of puromycin and blastisidin were obtained and cloned. The TP53 gene, puromycin resistance gene, and blastisidin resistance gene were amplified using the primers described above. The results are shown in Figure 9. As shown in Figure 9, cells were obtained in which the first intron of the human TP53 gene was replaced with the donor DNA sequence in both alleles (i.e., cells in which the first intron of the human TP53 gene was deleted in both alleles).

[0132] [Example 6] Similar to TP53, the above method was applied to human genes encoding MLH1, CD44, MET, and APP (hereinafter referred to as target genes) in HCT116 cells to induce gene deletion. As shown in Figures 10-13, the sizes of the target genes were 58kb for MLH1, 94kb for CD44, 126kb for MET, and 290kb for APP. Upstream and downstream homology arms were designed for each of these target genes, and select marker donor DNAs were prepared, each carrying different select marker genes and visualization marker genes that are distinguishable from each other between the arms. Specifically, one select marker donor DNA contained GFP and a puromycin resistance gene, and the other select marker donor DNA contained GFP and a blastisidin resistance gene.

[0133] As shown in Figure 14A, cleavage was induced using a CRISPR / Cas9 system with gRNAs designed to specifically induce cleavage upstream and downstream of the target gene in the presence of a selection marker donor DNA. The cleavage sites were designed between the region where the upstream homology arm hybridizes and the target gene, and between the region where the downstream homology arm hybridizes and the target gene. This induced the first step of recombination.

[0134] <Guide RNA sequence> The guide RNA sequences for the upstream and downstream of each target gene were as follows: MLH1 upstream gRNA: GGCCUGACGUCGCGUUCGC (Sequence ID 18) MLH1 downstream gRNA: GGAGGCCUUGGCACGGGUUC (Sequence ID 19) CD44 upstream gRNA: CGAGGAUGGCGGACCGAACC (Sequence ID 20) CD44 downstream gRNA: GCCAAGUGGACUCAACGGAG (Sequence ID 21) MET upstream gRNA:GGGCCGCGCGCGCCGAUGCC (Sequence ID 22) MET1 downstream gRNA:GUUCCCACCUCGCAAGCAAU (SEQ ID NO: 23) APP upstream gRNA:CUCCCGGGGGUGUCGUAUAA (Sequence ID 24) APP downstream gRNA: UUCUAUAAAUGGACACCGAU (Sequence ID 25)

[0135] <Primer Sequence> Deletions of each target gene and introduction of drug resistance genes were detected using the following primers. PCR primers for MLH1 amplification PCR primers for MLH1 5' junction Forward primer: CCAAGAACGCTTCCATTTCT (SEQ ID NO: 26) Reverse primer: CCCTGTGCCTGGTCTGTC (SEQ ID NO: 27) PCR primers for MLH1 3' junction Forward primer: TTCTGAGGTCTCCAGCAAGT (SEQ ID NO: 28) Reverse primer: AAGTTGAAGATGAATTGAAAGCAG (SEQ ID NO: 29) PCR primers for Puromycin 5' junction Forward primer: CCAAGAACGCTTCCATTTCT (SEQ ID NO: 30) Reverse primer: GCTCGTAGAAGGGGAGGTTG (SEQ ID NO: 31) PCR primers for Puromycin 3' junction Forward primer: GTCACCGAGCTGCAAGAAC (SEQ ID NO: 32) Reverse primer: AAGTTGAAGATGAATTGAAAGCAG (SEQ ID NO: 33) PCR primers for Blasticidin 5' junction Forward primer: CCAAGAACGCTTCCATTTCT (SEQ ID NO: 34) Reverse primer: GCTTCAATATGTACTGCCGAAA (SEQ ID NO: 35) PCR primers for Blasticidin 3' junction Forward primer: GAAGCCATTGCGATTGGTAG (SEQ ID NO: 36) Reverse primer: AAGTTGAAGATGAATTGAAAGCAG (SEQ ID NO: 38) PCR NMR for CD44 amplification PCR primers for CD44 5' junction Forward primer: AGTGGATGGACAGGAGGATG (Sequence No. 39) Reverse primer: GCGAAAGGAGCTGGAGGA (SEQ ID NO: 40) PCR primers for CD44 3' junction Forward primer: ATGGAGCTGTGGAGGACAGA (SEQ ID NO: 41) Reverse primer: GAGTGGGTCTGAGTGGGAAC (SEQ ID NO: 42) PCR primers for Puromycin 5' junction Forward primer: AGTGGATGGACAGGAGGATG (Sequence No. 43) Reverse primer: GCTCGTAGAAGGGGAGGTTG (SEQ ID NO: 44) PCR primers for Puromycin 3' junction Forward primer: GTCACCGAGCTGCAAGAAC (SEQ ID NO: 45) Reverse primer: GAGTGGGTCTGAGTGGGAAC (SEQ ID NO: 46) PCR primers for Blasticidin 5' junction Forward primer: AGTGGATGGACAGGAGGATG (SEQ ID NO: 47) Reverse primer: GCTTCAATATGTACTGCCGAAA (SEQ ID NO: 48) PCR primers for Blasticidin 3' junction Forward primer: GAAGCCATTGCGATTGGTAG (SEQ ID NO: 49) Reverse primer: GAGTGGGTCTGAGTGGGAAC (SEQ ID NO: 50) PCR primers for MET amplification PCR primers for MET 5' junction Forward primer: TGAAATCACTCTTATGTAACCTCTGG (SEQ ID NO: 51) Reverse primer: AAGGGGCTGCAATTTTACCT (SEQ ID NO: 52) PCR primers for MET 3' junction Forward primer: GGGTGGATGGATTGAAAAGA (SEQ ID NO: 53) Reverse primer: TGCAGGTATAGGCAGTGACAAG (SEQ ID NO: 54) PCR primers for Puromycin 5' junction Forward primer: TGAAATCACTCTTATGTAACCTCTGG (SEQ ID NO: 55) Reverse primer: GCTCGTAGAAGGGGAGGTTG (SEQ ID NO: 56) PCR primers for Puromycin 3' junction Forward primer: GTCACCGAGCTGCAAGAAC (SEQ ID NO: 57) Reverse primer: TGCAGGTATAGGCAGTGACAAG (SEQ ID NO: 58) PCR primers for Blasticidin 5' junction Forward primer: TGAAATCACTCTTATGTAACCTCTGG (SEQ ID NO: 59) Reverse primer: GCTTCAATATGTACTGCCGAAA (SEQ ID NO: 60) PCR primers for Blasticidin 3' junction Forward primer: GAAGCCATTGCGATTGGTAG (SEQ ID NO: 61) Reverse primer: TGCAGGTATAGGCAGTGACAAG (SEQ ID NO: 62) PCR primers for APP amplification PCR primers for APP 5' junction Forward primer: GGGGAGCTGGTACAGAAATG (SEQ ID NO: 63) Reverse primer: CAGGATCAGGGAAAGGTGAG (SEQ ID NO: 64) PCR primers for APP 3' junction Forward primer: GAACGGCTACGAAAATCCAA (SEQ ID NO: 65) Reverse primer: CTCTTCTCCCCACCCAAAA (SEQ ID NO: 66) PCR primers for Puromycin 5' junction Forward primer: GGGGAGCTGGTACAGAAATG (SEQ ID NO: 67) Reverse primer: GCTCGTAGAAGGGGAGGTTG (SEQ ID NO: 68) PCR primers for Puromycin 3' junction Forward primer: GTCACCGAGCTGCAAGAAC (SEQ ID NO: 69) Reverse primer: CTCTTCTCCCCACCCAAAA (SEQ ID NO: 70) PCR primers for Blasticidin 5' junction Forward primer: GGGGAGCTGGTACAGAAATG (SEQ ID NO: 71) Reverse primer: GCTTCAATATGTACTGCCGAAA (SEQ ID NO: 72) PCR primers for Blasticidin 3' junction Forward primer: GAAGCCATTGCGATTGGTAG (SEQ ID NO: 73) Reverse primer: CTCTTCTCCCCACCCAAAA (SEQ ID NO: 74)

[0136] As a result, as shown in "1st step" in Figures 10-13, multiple clones lacking the target genes of both alleles were obtained. Furthermore, as shown in Figure 14A, introducing cuts both upstream and downstream of the target gene in the presence of the selection marker donor DNA resulted in a higher efficiency of obtaining clones lacking the target genes of both alleles compared to inducing cuts only downstream or only upstream of the target gene in the presence of the selection marker donor DNA, as shown in Figure 14B. In this way, the target region of the various genes mentioned above could be replaced with the sequence of the selection marker donor DNA.

[0137] Next, the selection marker, which was introduced instead of deleting the target gene, was removed by a second-step recombination. Here, GFP, a visualization marker, was used as an indicator, and cells lacking the region containing the selection marker were selected based on their lack of GFP expression. As a result, as shown in "2nd step" in Figures 10-13, the selection marker containing GFP was successfully removed from the genome in both alleles. [Industrial applicability]

[0138] The present invention provides a genome modification method and a genome modification kit that can efficiently modify two or more alleles and modify relatively large regions.

Claims

1. A method for producing cells in which two or more alleles of the chromosomal genome have been modified, (a) The steps of introducing (i) and (ii) below into cells containing two or more alleles to introduce a selection marker gene into each of the two or more alleles, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving target regions in two or more alleles of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, and containing the base sequence of the select marker gene between the upstream homology arm and the downstream homology arm, wherein the two or more types of select marker donor DNA each have a select marker gene that is distinguishable from each other, the select marker gene is unique to each type of select marker donor DNA, and the number of types of select marker donor DNA is equal to or greater than the number of alleles targeted for genome modification, (b) After step (a) above, a step of selecting cells that express all of the introduced, distinctly different, unique selection marker genes, by homologous recombination of different types of selection marker donor DNA with respect to the two or more alleles (a step for positive selection), Methods that include...

2. A method for modifying two or more alleles of a chromosomal genome, (a) The steps of introducing (i) and (ii) below into cells containing two or more alleles to introduce a selection marker gene into each of the two or more alleles, (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule capable of targeting and cleaving target regions in two or more alleles of the chromosomal genome, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, and containing the base sequence of the select marker gene between the upstream homology arm and the downstream homology arm, wherein the two or more types of select marker donor DNA each have a select marker gene that is distinguishable from each other, the select marker gene is unique to each type of select marker donor DNA, and the number of types of select marker donor DNA is equal to or greater than the number of alleles targeted for genome modification, (b) After step (a) above, a step of selecting cells that express all of the introduced, distinctly different, unique selection marker genes, by homologous recombination of different types of selection marker donor DNA with respect to the two or more alleles (a step for positive selection), Methods that include...

3. The method according to claim 1 or 2, wherein the target region has a length of 5 kbp or more.

4. The method according to claim 3, wherein the target region has a length of 8 kbp or more.

5. Each of two or more selection marker donor DNAs has a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence between the upstream homology arm and the downstream homology arm, wherein if the selection marker gene is used for both positive and negative selection, it does not need to have another selection marker gene for negative selection. (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired base sequence, comprising an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection), The method according to any one of claims 1 to 4, further comprising:

6. Each of two or more selection marker donor DNAs has a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence between the upstream homology arm and the downstream homology arm, wherein if the selection marker gene is used for both positive and negative selection, it does not need to have another selection marker gene for negative selection. (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired base sequence, comprising an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection), The method according to claim 3, further comprising:

7. Each of two or more selection marker donor DNAs has a selection marker gene for positive selection, a marker gene for negative selection, and a target sequence between the upstream homology arm and the downstream homology arm, wherein if the selection marker gene is used for both positive and negative selection, it does not need to have another selection marker gene for negative selection. (c) After step (b) above, a step of introducing (iii) and (iv) below into selected cells to introduce recombinant donor DNA into the two or more alleles, (iii) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets the target sequence and can cleave the target sequence, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (iv) Recombinant donor DNA containing a desired base sequence, comprising an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, (d) After step (c), a step of selecting cells that do not express the negative selection marker gene (a step for negative selection), The method according to claim 4, further comprising:

8. The method according to any one of claims 5 to 7, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

9. The method according to claim 6 or 7, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

10. The method according to claim 7, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

11. The method according to claim 8, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

12. A genome modification kit for modifying two or more alleles of a chromosomal genome, including (i) and (ii) below. (i) A genome modification system comprising a sequence-specific nucleic acid cleavage molecule that targets a target region of the chromosomal genome and can cleave the target region, or a polynucleotide encoding the sequence-specific nucleic acid cleavage molecule. (ii) Two or more types of select marker donor DNA, each having an upstream homology arm having a base sequence homologously recombinable with the base sequence upstream of the target region, and a downstream homology arm having a base sequence homologously recombinable with the base sequence downstream of the target region, with the base sequence of the select marker gene included between the upstream homology arm and the downstream homology arm, the two or more types of select marker donor DNA each having a select marker gene that is distinguishable from each other, the select marker gene is unique to each type of select marker donor DNA, and the number of types of select marker donor DNA is equal to or greater than the number of alleles targeted for genome modification.

13. The kit according to claim 12, wherein the target region has a length of 5 kbp or more.

14. The kit according to claim 13, wherein the target region has a length of 8 kbp or more.

15. The kit according to any one of claims 12 to 14, further comprising recombinant donor DNA.

16. The kit according to any one of claims 12 to 15, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

17. The kit according to any one of claims 12 to 16, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

18. The kit according to claim 12, wherein the target region has a length of 5 kbp or more, and the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 5 kbp or more.

19. The kit according to claim 18, wherein the target region has a length of 8 kbp or more, and the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

20. The method according to claim 5, wherein the recombinant donor DNA does not have a base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA, and in the two or more alleles of the chromosomal genome obtained after modification, the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

21. The method according to claim 6 or 7, wherein the recombinant donor DNA does not have a base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA, and in the two or more alleles of the modified chromosomal genome, the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

22. The method according to claim 3, wherein the target sequences of the site-directed recombinant enzyme are absent in two or more alleles of the chromosomal genome obtained after modification.

23. The method according to claim 4, wherein the target sequences of the site-directed recombinant enzyme are absent in two or more alleles of the chromosomal genome obtained after modification.

24. The method according to claim 5, wherein the target sequences of the site-directed recombinant enzyme are absent in two or more alleles of the chromosomal genome obtained after modification.

25. The method according to claim 6 or 7, wherein, in the two or more alleles of the chromosomal genome obtained after modification, there are no target sequences for site-directed recombinant enzymes.

26. The method according to any one of claims 1 to 11 and 20 to 25, wherein single-cell cloning is not performed in the process up to selecting cells in which two or more alleles have been modified in step (b).

27. A cell having two or more alleles on the chromosomal genome with respect to a target region, wherein the target region of each of the two or more alleles is deleted, and the upstream and downstream sequences of the target region are seamlessly linked without base insertions, substitutions, or deletions.

28. The cell according to claim 27, wherein the target region has a length of 5 kbp or more.

29. The cell according to claim 27 or 28, wherein its genome does not contain a target sequence for site-directed recombinant enzymes.

30. The method according to claim 6 or 7, wherein the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA has a length of 8 kbp or more.

31. The kit according to claim 15, wherein the recombinant donor DNA does not have a base sequence in the region between the upstream homology arm and the downstream homology arm of the recombinant donor DNA.