Method for manufacturing cell including modified genome DNA, and method for producing gene product
By specifically cleaving and joining genomic DNA sequences to create modified genes without foreign DNA, the method addresses regulatory challenges of genetically modified organisms, enabling efficient production of cells with desired gene expression capabilities.
Patent Information
- Application Number
- JP2024026212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-02-26
AI Technical Summary
Genetically modified organisms are subject to stringent regulations under the Cartagena Protocol, making their distribution, management, and commercial use disadvantageous, while existing genome editing methods struggle to create cells with modified genes that do not contain foreign DNA.
A method involving specific cleavage of genomic DNA at multiple target sequences using sequence-specific nucleases and DNA double-strand break repair to delete regions and join sequence fragments, creating a modified gene that differs from the existing gene without introducing foreign DNA.
Produces cells with modified genomic DNA capable of expressing polypeptides or functional RNA, subject to less stringent legal restrictions, facilitating their distribution and use.
Smart Images

Figure 2025129527000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for producing cells containing modified genomic DNA that is free of foreign DNA, cells produced by the method, and methods for producing gene products using the cells or organisms containing the cells. [Background technology]
[0002] Genome editing techniques that directly manipulate genomic DNA are known to modify the genome of cells or organisms. For example, the CRISPR / Cas system, zinc finger nucleases (ZFNs), transcription activation-like effector nucleases (TALENs), and meganucleases can recognize and selectively cleave specific sequences in the genome. The ends of the cleaved genomic DNA are known to be rejoined by the cell's inherent DNA double-strand break (DSB) repair mechanism. By cleaving genomic DNA at two sites, the region between the two sites is excised and rejoined, allowing for the artificial creation of genomic DNA lacking the excised region. It is known that the DSB repair process can result in deletions, insertions, or substitutions of several bases at the sites of the cuts and ligations.
[0003] Mutations introduced by genome editing through deletion manipulation of genomic DNA and DSB repair are essentially indistinguishable from mutations caused by DNA double-strand breaks and repair that occur naturally within cells. Therefore, organisms (including cells) obtained through genome editing that do not introduce exogenous nucleotides are sometimes treated as organisms that do not fall under the category of "genetically modified organisms" (referred to as "non-genetically modified organisms") under the Cartagena Protocol on Biosafety to the Convention on Biological Diversity (Cartagena Protocol). Non-genetically modified organisms are subject to less stringent regulations than genetically modified organisms, which may offer advantages in terms of distribution, management, and commercial use.
[0004] Patent Document 1 discloses a method for creating a novel gene in an organism, comprising the steps of simultaneously creating DNA breaks at two or more different specific sites in the genome of the organism, the specific sites being genomic sites that can separate different genetic elements or different protein domains, and the DNA breaks creating a new combination of various genetic elements or various protein domains that differ from the original genomic sequence and are linked to each other by non-homologous end joining (NHEJ) or homology repair, thereby creating a new gene. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Special Publication No. 2022-553598 [Non-patent literature]
[0006] [Non-Patent Document 1] Liu G. et al. Molecular Cell, 2022, Vol.82, pp.333-347 [Non-patent document 2] Gao C. Cell, 2021, Vol.184, pp.1621-1635 Summary of the Invention [Problem to be solved by the invention]
[0007] To constitutively produce proteins that are not naturally present in cells or that are expressed in low amounts in cells or organisms, foreign DNA is often introduced into cells to create genetically modified organisms. However, genetically modified organisms are subject to legal restrictions under the Cartagena Protocol, and strict regulations are imposed on their distribution, management, and use, making them commercially disadvantageous.
[0008] Patent Document 1 describes a technical concept of linking promoters or different protein domains of different genes present on existing genomic DNA solely through deletion manipulation of genomic DNA to create a new functional gene. While Patent Document 1 describes cutting genomic DNA at two or more locations (i.e., excising fragments), it does not specify repeating deletion manipulations more than twice, and does not specifically describe creating a sequence dissimilar to any gene on genomic DNA through repeated deletion manipulations, which differs from the present invention. Rather, Patent Document 1 argues that gene editing tools such as CRISPR / Cas9 can be used for knockout, but because they mutate existing genes without generating new genes, they are considered difficult to meet certain production needs (see paragraph
[0004] of Patent Document 1).
[0009] On the other hand, the objective of the present invention is to provide a means for constructing modified DNA having a nucleotide sequence different from an existing genomic DNA sequence, which contains a modified gene capable of expressing a polypeptide, functional RNA, or both, and which does not contain foreign DNA. [Means for solving the problem]
[0010] As a result of extensive investigation, the present inventors have found that (a) specifically cleaving the genomic DNA of a cell at at least two target sequences with sequence-specific nucleases and joining the sequences by DNA double-strand break repair by the cell; at least twice to delete at least two regions from the existing gene and the non-coding regions adjacent to the existing gene to generate a modified gene. The present inventors have found that it is possible to produce cells containing a modified gene capable of expressing a polypeptide, a functional RNA, or both, and containing genomic DNA that does not contain any foreign DNA, thereby completing the present invention.
[0011] The present invention includes, but is not limited to, the following aspects. [1] 1. A method for generating a cell containing modified genomic DNA that is free of exogenous DNA, comprising: (a) specifically cleaving the genomic DNA of a cell at at least two target sequences with sequence-specific nucleases and joining the sequences by DNA double-strand break repair by the cell; at least twice to delete at least two regions from the existing gene and the non-coding regions adjacent to the existing gene, thereby generating a modified gene that differs from the nucleotide sequence of the existing gene; The genomic DNA of the cell in step (a) At least two deleted regions that are not included in the modified genomic DNA due to the deletion; comprising at least one linked fragment of the region between the two deleted regions, which is the portion remaining in the modified genomic DNA; at least one of the linked fragments is 1 to 150 base pairs in length; the promoter region, the transcribed region, or both of the modified gene comprises the full length or a portion of at least one of the linked fragments; The method is configured to express a polypeptide, a functional RNA, or both from the modified gene. [2] The method according to [1], wherein at least two of the steps (a) are carried out simultaneously or consecutively using different combinations of the target sequences. [3] (b) The method according to [1] or [2], further comprising the step of selecting the cells that do not contain off-target mutations. [4] The method according to any one of [1] to [3], wherein at least one of the target sequences includes the sequence of at least one linked fragment (first linked fragment), and at least one of the ends of the first linked fragment is connected to the whole or part of the sequence of the linked fragment or genomic DNA that is linked to the first linked fragment in the step (a). [5] The method according to any one of [1] to [4], which does not include a step of inserting foreign DNA into the genomic DNA of the cell. [6] The method further comprises a step of inserting, by cellular recombination repair, a foreign DNA having a first homology arm and a second homology arm attached to both ends thereof, capable of homologous recombination with the genomic DNA, into at least one of the deleted regions, inserting the foreign DNA so as to be adjacent to at least one of the deleted regions, or inserting the foreign DNA so as to replace a part or all of at least one of the deleted regions; In the step (a), the foreign DNA is removed from the genomic DNA by the deletion; and The method according to any one of [1] to [4], further comprising the step of selecting the cells that are free of exogenous DNA. [7] The method according to [6], wherein the foreign DNA comprises a cell selection marker gene. [8] the first or second homology arm comprises the sequence of at least one linked fragment (first linked fragment); The method according to [6] or [7], wherein all or part of the sequence of the linked fragment or genomic DNA to be linked to the first linked fragment in step (a) is connected to at least one of the ends of the first linked fragment. [9] The method according to any one of [1] to [8], wherein the total length of the linked fragments is 9 to 180 base pairs.
[10] The method according to any one of [1] to [8], wherein the total length of the linked fragments is 30% or less of the total length of the deleted regions.
[11] The method according to any one of [1] to
[10] , wherein the existing gene and the modified gene encode polypeptides, and in a wild-type cell of the cell, the polypeptide encoded by the existing gene accounts for 10% or more of the mass of all proteins.
[12] the existing gene and the modified gene encode a polypeptide, and the coding region of the modified gene comprises the full length or a portion of at least one of the linked fragments; The method according to any one of [1] to
[11] , wherein the entire length or a portion of at least one of the linked fragments encodes a modified polypeptide region of 3 to 60 amino acids that is not encoded by the existing gene.
[13] The method according to
[12] , wherein the modified polypeptide region comprises a functional peptide sequence, an enzyme cleavage sequence, or both.
[14] The method described in any one of [1] to
[13] , wherein the existing gene and the modified gene encode a polypeptide, and the full-length polypeptide encoded by the modified gene is 20% or less of the full-length polypeptide encoded by the existing gene.
[15] The method described in any one of [1] to
[13] , wherein the existing gene and the modified gene encode a polypeptide, and the full-length polypeptide encoded by the modified gene is 50% or more in length of the full-length polypeptide encoded by the existing gene.
[16] The method according to any one of [1] to
[15] , wherein the cells are cells used in food or for producing food.
[17] The method according to any one of [1] to
[16] , wherein the sequence-specific nuclease is carried out by a CRISPR / Cas system.
[18] The method according to
[17] , wherein the CRISPR / Cas system comprises SpCas9-NG or a functional analog thereof.
[19] A cell having modified genomic DNA comprising: The genomic DNA is At least two deleted regions (deleted regions) in an existing gene and its adjacent non-coding region of the wild-type genomic DNA of the cell, as compared with the wild-type genomic DNA; At least one linked fragment corresponding to a region sandwiched between deleted regions (deleted regions) in the existing gene; a modified gene whose nucleotide sequence is altered relative to the existing gene by the deletion; and at least one of the linked fragments is 1 to 150 base pairs in length; the promoter region, the transcribed region, or both of the modified gene comprises the full length or a portion of at least one of the linked fragments; A cell configured to express a polypeptide, a functional RNA, or both from the modified gene.
[20] The cell according to
[19] , wherein the total length of the linked fragments is 9 to 180 base pairs. [twenty one] The cell according to
[19] or
[20] , wherein the total length of the linked fragments is 30% or less of the total length of the deleted regions. [twenty two] The cell according to any one of
[19] to
[21] , wherein the existing gene and the modified gene encode polypeptides, and in a wild-type cell of the cell, the amount of protein encoded by the existing gene accounts for 10% or more of the mass of total protein. [twenty three] the existing gene and the modified gene encode a polypeptide, and the coding region of the modified gene comprises the full length or a portion of at least one of the linked fragments; The cell according to any one of
[19] to
[22] , wherein the entire length or a portion of at least one of the linked fragments encodes a modified polypeptide region of 3 to 60 amino acids that is not encoded by the existing gene. [twenty four] The cell according to
[23] , wherein the modified polypeptide region comprises a functional peptide, an enzyme cleavage sequence, or both. [twenty five] A cell described in any one of
[19] to
[24] , wherein the existing gene and the modified gene encode a polypeptide, and the full length of the polypeptide encoded by the modified gene is 20% or less of the full length of the polypeptide encoded by the existing gene.
[26] The cell according to any one of
[19] to
[24] , wherein the existing gene and the modified gene encode a polypeptide, and the full length of the polypeptide encoded by the modified gene is 50% or more of the full length of the polypeptide encoded by the existing gene.
[27] The cell according to any one of
[19] to
[26] , wherein the cell is used in food or in the production of food.
[28] An organism comprising a cell produced by the method according to any one of [1] to
[18] , or a cell according to any one of
[19] to
[27] .
[29] A method for producing a gene product, comprising a step of expressing a gene product encoded by the modified gene in a cell produced by the method described in any one of [1] to
[18] , a cell described in any one of
[19] to
[27] , or an organism containing such a cell. [Effects of the Invention]
[0012] The cell production method of the present invention can produce cells that contain modified genomic DNA containing a modified gene capable of expressing a polypeptide, functional RNA, or both, without containing any foreign DNA. Such cells and organisms containing the cells may be subject to less stringent legal restrictions in terms of distribution, storage, use, etc. than genetically modified organisms. Furthermore, the cells or organisms can be used to produce gene products encoded by the modified genes. [Brief explanation of the drawings]
[0013] [Figure 1A] (a) Schematic diagram showing an example of a method for producing modified genomic DNA when each step is performed once. [Figure 1B] FIG. 1 is a schematic diagram showing an example of a method for producing modified genomic DNA in which step (a) is carried out multiple times at once. [Figure 2] FIG. 1 shows an example of a method for producing modified genomic DNA, in which a modified gene is produced by deleting part of the coding region of an existing gene. [Figure 3A]FIG. 1 is a schematic diagram showing an example of a method for producing modified genomic DNA, including an embodiment in which foreign DNA containing a marker gene (GFP, ampicillin resistance gene (Amp)) is inserted when the deleted region is excised, and then the foreign DNA is removed. [Figure 3B] FIG. 1 is a schematic diagram showing an example of a method for producing modified genomic DNA, including an embodiment in which foreign DNA containing a marker gene (RFP, kanamycin resistance gene (Kan)) is inserted and then the foreign DNA and the deleted region are removed together. [Figure 4A] FIG. 1 shows the sequences of the portions of the sense strand corresponding to the sense strand PAM sequence and antisense strand PAM sequence of SpCas9-NG (PAM sequence: NG) on the sense strand. [Figure 4B] This figure shows specific examples of the nucleotide restriction sites (italicized G or C) for the PAM sequence at each cleavage site (↓ or ↑) when cleaving both ends of linked fragments of various lengths using SpCas-NG. The linked fragment is indicated by the underlined area surrounded by the two cleavage sites. [Figure 4C] This figure shows specific examples of the nucleotide restriction sites (italicized G or C) for the PAM sequence at each cleavage site (↓ or ↑) when cleaving both ends of linked fragments of various lengths using SpCas-NG. The linked fragment is indicated by the underlined area surrounded by the two cleavage sites. [Figure 4D] The following shows classification (Types 1 to 4) based on the position of the PAM sequence relative to the ligated fragment when both ends of the ligated fragment are cleaved. The ligated fragment is indicated by an underlined area surrounded by two cleavage sites. [Figure 4E] The classification (types 5 and 6) is based on the position of the PAM sequence relative to the cleavage site when the cleavage site closest to the 5' or 3' end is cleaved. [Figure 5A] FIG. 1 is a schematic diagram showing an example of an embodiment in which a first ligated fragment is included in a target sequence of CRISPR / Cas9. [Figure 5B] FIG. 1 is a schematic diagram showing an example of an embodiment in which step (a) is carried out multiple times simultaneously using a guide RNA in which the first linked fragment is included in the target sequence. [Figure 6A] FIG. 1 shows a schematic diagram illustrating the positional relationship of linked fragments in existing genes and modified genes in Example 1, and a diagram illustrating the mapped search sequence (a fragment consisting of linked fragments and a PAM sequence). [Figure 6B] FIG. 10 is a schematic diagram showing the positional relationship of the linked fragments in the existing gene and the modified gene, and a diagram showing the mapped search sequence (a fragment consisting of the linked fragments and the PAM sequence) in Example 2. [Figure 6C] FIG. 10 is a schematic diagram showing the positional relationship of the linked fragments in the existing gene and the modified gene, and a diagram showing the mapped search sequence (a fragment consisting of the linked fragments and the PAM sequence) in Example 3. [Figure 6D] FIG. 1 is a schematic diagram showing the positional relationship of ligated fragments in an existing gene and a modified gene in Example 4. [Figure 6E] FIG. 10 shows the mapped search sequences (fragments consisting of linked fragments and PAM sequences) in Example 4. [Figure 6F] FIG. 10 is a schematic diagram showing the positional relationship of the linked fragments in the existing gene and the modified gene, and a diagram showing the mapped search sequence (a fragment consisting of the linked fragments and the PAM sequence) in Example 5. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, the drawings are merely examples, and the present invention is not limited to the embodiments shown in the drawings.
[0015] Unless otherwise specified, nucleotide sequences are described herein from the 5' to the 3' end, and amino acid sequences are described herein from the N-terminus to the C-terminus.
[0016] As used herein, "N" in a nucleotide sequence represents any one of the bases adenine, guanine, cytosine, and uracil in the case of RNA, and any one of the bases adenine, guanine, cytosine, and thymine in the case of DNA.
[0017] As used herein, the term "gene" refers to a polynucleotide that is operably linked to an appropriate control sequence (e.g., a promoter, an enhancer, etc.) and contains a transcribed region that is transcribed from DNA to RNA. A gene in which a polypeptide is expressed by translating the transcribed mRNA may be referred to as a "polypeptide gene" herein. A polypeptide gene contains a ribosome binding sequence and at least one open reading frame (ORF) that is translated. A region that is not transcribed or translated, such as a pseudogene, does not qualify as a gene. As used herein, the sequence of a gene is represented by the sequence of the sense strand, unless otherwise specified.
[0018] That is, gene products resulting from gene expression can include polypeptides as well as functional RNA. Examples of functional RNA include, but are not limited to, small RNAs (such as siRNA, miRNA, and piRNA), tRNA, rRNA, functional tRNA-derived RNA fragments (tRFs; see Alves CS et al., Front. Mol. Biosci., 2021, 8:638911), antisense RNA, ribozymes, and RNA aptamers. A gene may express two or more polypeptides, functional RNAs, or both.
[0019] As used herein, the term "existing gene" refers to a specific gene that is present in the genomic DNA of a cell before modification and whose sequence is to be modified in the method of producing a cell containing modified genomic DNA of the present invention.
[0020] As used herein, the term "modified gene" refers to a gene that results from the modification of the nucleotide sequence of an existing gene in the method of producing a cell containing modified genomic DNA of the present invention.
[0021] As used herein, the "5'-end" or "5'-end" representing a position on a gene or a position relative to a gene refers to the 5'-end or 5'-end of the sense strand of the gene, unless otherwise specified. As used herein, the "3'-end" or "3'-end" representing a position on a gene or a position relative to a gene refers to the 3'-end or 3'-end of the sense strand of the gene, unless otherwise specified.
[0022] As used herein, the term "non-coding region" refers to a region of genomic DNA that does not correspond to a gene. For example, a region encoding a pseudogene is included in the non-coding region.
[0023] As used herein, the term "control sequence" refers to a nucleotide sequence that is necessary for or regulates the expression (transcription or translation) of a gene. Examples of control sequences in prokaryotes include promoters, operator sequences, ribosomal binding sequences, and transcription termination sequences (terminators). Examples of control sequences in eukaryotic cells include promoters, polyadenylation signals, enhancers, transcription termination sequences, internal ribosome entry sites (IRES), and the like.
[0024] As used herein, the term "5'-untranslated region" or "5'-UTR" refers to the region of nucleotides that encodes the untranslated region at the 5' portion of an mRNA molecule. As used herein, the term "3'-untranslated region" or "3'-UTR" refers to the region of nucleotides that encodes the untranslated region at the 3' portion of an mRNA molecule.
[0025] As used herein, a "coding region" refers to a region that encodes the amino acid sequence of a polypeptide and is translated into a polypeptide when placed under the control of appropriate control sequences, including a promoter. Unless otherwise specified, the "coding region" of a particular polypeptide refers to a region that encodes the entire amino acid sequence of the polypeptide or a region that encodes a part of the amino acid sequence of the polypeptide.
[0026] As used herein, an "exogenous" or "foreign" gene or nucleotide refers to a gene or nucleotide that is not found in the cell prior to genetic manipulation and that has been or will be introduced into the cell by genetic manipulation.
[0027] As used herein, "genetically modified organisms" refers to organisms (including organisms and cells) that contain nucleic acids or their replicates obtained by techniques for processing nucleic acids outside cells (so-called recombinant DNA techniques) or techniques for fusing living organism cells, as defined in the Cartagena Protocol. Therefore, even if foreign DNA is inserted into genomic DNA or foreign plasmid DNA is introduced into a cell, it does not fall under the category of genetically modified organisms as long as the foreign DNA is completely removed from the final cell and the organism comprising the cell.
[0028] As used herein, the term "non-genetically modified organism" refers to an organism that does not fall under the category of genetically modified organisms.
[0029] As used herein, the term "functional analog" refers to a polypeptide that has an amino acid sequence similar to that of a given polypeptide (e.g., amino acid identity of 95% or more, 98% or more, 99% or more, or 99.9% or more) and exhibits the same qualitative function. Specific examples include polypeptides into which amino acid mutations that do not affect activity have been introduced into a given polypeptide.
[0030] As used herein, the identity (%) of an amino acid sequence or a nucleotide sequence is the "identity" value in an alignment performed in Protein BLAST or Nucleotide BLAST of NCBI BLAST (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi) with Align two or more sequences selected and default parameters.
[0031] [Method for producing cells containing modified genomic DNA] One embodiment of the present invention is a method for producing a cell containing modified genomic DNA that does not contain foreign DNA (sometimes referred to herein as the "production method of the present invention"), which comprises: (a) specifically cleaving the genomic DNA of a cell at at least two target sequences with sequence-specific nucleases and joining the sequences by DNA double-strand break repair by the cell; at least twice to delete at least two regions from the existing gene and its adjacent non-coding regions to generate a modified gene.
[0032] According to the production method of the present invention, by linking sequence fragments scattered throughout an existing gene or its adjacent non-coding region, it is possible to construct a modified gene that contains a nucleotide sequence significantly different from the sequence contained in the existing gene and is capable of expressing a polypeptide, functional RNA, or both.
[0033] ((a) process) In one (a) step, genomic DNA is cut at two sites and the region between them is excised. The excised region is then deleted from the genomic DNA that is joined by DNA double-strand break (DSB) repair in the cell. In some cases, a substitution, insertion, or deletion of one or more base pairs in length is introduced at the site of the cut and ligation during DSB repair. Herein, the excised region is referred to as the "deleted region." Furthermore, the sites at both ends of the deleted region that are cleaved by the sequence-specific endonuclease are referred to as the "cleavage site."
[0034] By performing step (a) two or more times, two or more deletion regions are excised from the genomic DNA. Herein, the regions flanked by the deletion regions and remaining in the modified genomic DNA are referred to as "ligated fragments." Finally, by performing step (a) k times, k deletion regions are lost from the genomic DNA, and k-1 ligated fragments are ligated to construct a modified genomic DNA containing a modified promoter region and / or transcribed region. Figure 1A shows an example of a deletion region and a cleavage site.
[0035] The deletion site on the genomic DNA in the operation differs between one step (a) and another step (a). Thus, in one embodiment, the two target sequences in each step (a) are different from each other.
[0036] Two or more steps (a) may be performed simultaneously or consecutively by a single genetic engineering operation. For example, in Figure 1B, of the three deletion regions, two regions are excised by two simultaneous steps (a), and the other region is excised by a separate step (a). In this way, when two or more steps (a) are performed simultaneously, it is not necessary for all of the steps (a) to be performed simultaneously or consecutively.
[0037] In this specification, one or more steps (a) that are carried out simultaneously or successively by a single genetic engineering operation may be referred to as a "set of steps (a)."
[0038] In the production method of the present invention, the cleavage site does not need to be located at a site on the genomic DNA that encodes a site capable of separating protein domains of an existing gene (ie, a domain boundary). In one embodiment, the existing gene encodes a protein, and in step (a), at least one cleavage and ligation occurs at a position other than a site encoding a domain boundary of the protein. In a particular embodiment, in step (a), all cleavages occur at positions other than a site encoding a domain boundary of the protein.
[0039] In one embodiment, the deletion region is present in one existing gene and its adjacent non-coding region in the genomic DNA of the cell. This configuration is preferable because it allows the use of a portion of the regulatory sequence and transcribed region (e.g., coding region) of the existing gene in the construction of the modified gene. In one embodiment, the deletion region includes the non-coding region located 5' to the existing gene encoding the polypeptide, and at least one region selected from the group consisting of the regulatory region, coding region, 5' untranslated region, intron, and 3' untranslated region of the existing gene.
[0040] In one embodiment, the deleted region consists of the non-coding region 5' of the existing gene, the promoter region, or a combination thereof. This embodiment is advantageous in that it allows the promoter sequence to be modified and allows the expression of the existing gene to be regulated without affecting other regulatory sequences, splicing factors, and coding regions (see Example 5).
[0041] In one embodiment, the deleted region consists of a transcribed region of an existing gene. When the existing gene encodes a polypeptide, the deleted region consists of, for example, at least one region selected from the group consisting of a coding region, an intron, and a 3'-UTR. This embodiment is preferable because it facilitates the construction of a modified gene while minimizing the introduction of mutations into the regulatory sequence of the existing gene.
[0042] The possible locations of the linked fragments can be understood based on the possible locations of the deleted regions described above. In one embodiment, the linked fragments are present in one existing gene and its adjacent non-coding regions in the genomic DNA of the cell. In a more specific embodiment, the linked fragments include at least one selected from the group consisting of the 5' non-coding region and 3' non-coding region adjacent to the existing gene, and the transcribed region. In one embodiment, the existing gene encodes a polypeptide, and the linked fragments include at least one selected from the group consisting of the 5' non-coding region and 3' non-coding region adjacent to the existing gene, the regulatory region of the existing gene, the coding region, the 5' untranslated region, introns, and 3' untranslated region.
[0043] In the production method of the present invention, the length of each linked fragment is not particularly limited, as long as the desired modified gene sequence can be produced by deletion alone from the genomic DNA sequence. Longer linked fragments reduce the number of steps (a), thereby saving labor and time. On the other hand, if the linked fragments are too long, it may be difficult to produce a modified gene having a sequence unrelated to existing genes.
[0044] In one embodiment, the length of at least one linked fragment is 1 to 150 base pairs (bp), for example, 1 bp or more, 2 bp or more, 3 bp or more, 4 bp or more, 5 bp or more, 6 bp or more, 8 bp or more, 10 bp or more, 20 bp or more, 30 bp or more, 40 bp or more, 50 bp or more, 60 bp or more, 70 bp or more, 80 bp or more, 90 bp or more, 100 bp or more, 110 bp or more, 120 bp or more, 130 bp or more, or 140 bp or more, and may be 140 bp or less, 130 bp or less, 120 bp or less, 110 bp or less, 100 bp or less, 90 bp or less, 80 bp or less, 70 bp or less, 60 bp or less, 50 bp or less, 40 bp or less, 30 bp or less, 20 bp or less, 10 bp or less, 8 bp or less, 6 bp or less, or 3 bp or less. In one embodiment, the length of at least one linked fragment is 1 to 150 bp, and may be, for example, 1 to 120 bp, 1 to 100 bp, 1 to 80 bp, 1 to 50 bp, 1 to 30 bp, 1 to 20 bp, 1 to 10 bp, 1 to 8 bp, or 1 to 5 bp.
[0045] In one embodiment, the length of at least two, three, four, or five linked fragments is between 1 and 150 base pairs (bp), for example, 1 bp or more, 2 bp or more, 3 bp or more, 4 bp or more, 5 bp or more, 6 bp or more, 8 bp or more, 10 bp or more, 20 bp or more, 30 bp or more, 40 bp or more, 50 bp or more, 60 bp or more, 70 bp or more, 80 bp or more, 90 bp or more, 100 bp or more, 110 bp or more, 120 bp or more, 130 bp or more, or 140 bp or more, and may be 140 bp or less, 130 bp or less, 120 bp or less, 110 bp or less, 100 bp or less, 90 bp or less, 80 bp or less, 70 bp or less, 60 bp or less, 50 bp or less, 40 bp or less, 30 bp or less, 20 bp or less, 10 bp or less, 8 bp or less, 6 bp or less, or 3 bp or less. In one embodiment, the length of at least two, three, four, or five linked fragments is 1 to 150 bp, and may be, for example, 1 to 120 bp, 1 to 100 bp, 1 to 80 bp, 1 to 50 bp, 1 to 30 bp, 1 to 20 bp, 1 to 10 bp, 1 to 8 bp, or 1 to 5 bp.
[0046] In one embodiment, the length of all linked fragments is 1 to 150 base pairs (bp), for example, 1 bp or more, 2 bp or more, 3 bp or more, 4 bp or more, 5 bp or more, 6 bp or more, 8 bp or more, 10 bp or more, 20 bp or more, 30 bp or more, 40 bp or more, 50 bp or more, 60 bp or more, 70 bp or more, 80 bp or more, 90 bp or more, 100 bp or more, 110 bp or more, 120 bp or more, 130 bp or more, or 140 bp or more, and may be 140 bp or less, 130 bp or less, 120 bp or less, 110 bp or less, 100 bp or less, 90 bp or less, 80 bp or less, 70 bp or less, 60 bp or less, 50 bp or less, 40 bp or less, 30 bp or less, 20 bp or less, 10 bp or less, 8 bp or less, 6 bp or less, or 3 bp or less. In one embodiment, the length of all linked fragments is 1 to 150 bp, and may be, for example, 1 to 120 bp, 1 to 100 bp, 1 to 80 bp, 1 to 50 bp, 1 to 30 bp, 1 to 20 bp, 1 to 10 bp, 1 to 8 bp, or 1 to 5 bp.
[0047] In one embodiment, the total length of the linked fragments of the modified genomic DNA is, for example, 1 to 500 base pairs, preferably 9 to 180 base pairs (bp), and may be, for example, 9 bp or more, 10 bp or more, 20 bp or more, 30 bp or more, 40 bp or more, 50 bp or more, 60 bp or more, 70 bp or more, 80 bp or more, 90 bp or more, 100 bp or more, 110 bp or more, 120 bp or more, 130 bp or more, 140 bp or more, 150 bp or more, 160 bp or more, or 170 bp or more, and 170 bp or less, 160 bp or less, 150 bp or less, 140 bp or less, 130 bp or less, 120 bp or less, 110 bp or less, 100 bp or less, 90 bp or less, 80 bp or less, 70 bp or less, 60 bp or less, 50 bp or less, 40 bp or less, 30 bp or less, 20 bp or less, 10 bp or less, 8 bp or less, 6 bp or less, or 3 bp or less. In one embodiment, the total length of all linked fragments is 9 to 180 bp, and may be, for example, 9 to 150 bp, 9 to 120 bp, 9 to 100 bp, 9 to 80 bp, 9 to 50 bp, 9 to 30 bp, or 9 to 20 bp. Although it depends on the target sequence of the sequence-specific endonuclease used, when the CRISPR / Cas9 system is used, even if the total length of the linked fragments of the modified genomic DNA is 150 bp or more, construction by only deleting existing genes is possible (see Example 4).
[0048] In one embodiment, 50% or more, 60% or more, 70% or more, or 80% or more of the linked fragments have a length of 30 bp or less (preferably, 10 to 30 bp or 10 to 20 bp). As shown in Examples 1 to 5, when constructing a sequence different from an existing gene using linked fragments, linked fragments of these lengths tend to be used frequently.
[0049] In one embodiment, the total length of the linked fragments of the modified genomic DNA is 30% or less of the total length of the deleted region, and may be, for example, 20% or less, 10% or less, 5% or less, or 2% or less.
[0050] DSB repair includes non-homologous end joining (NHEJ), in which ends without homologous regions are joined; microhomology-mediated end joining (MMEJ), in which ends with a complementary sequence of about 5 to 20 bases are joined; and homology-directed repair (HDR). When no foreign DNA is inserted, repair in step (a) of the present invention is preferably carried out by either NHEJ or MMEJ. Among these, ligation by NHEJ repair is more preferred due to its high degree of flexibility in target sequences. As described below, when foreign DNA is inserted along with deletion of genomic DNA, MMEJ or HDR can be used for DSB repair.
[0051] During DSB repair, several additional nucleotides may be deleted or inserted at the ends of the resulting break. Even if such DSB repair-related mutations exist in the modified genomic DNA, they do not involve the insertion of foreign DNA, and therefore cells or organisms containing the modified genomic DNA are not considered genetically modified organisms.
[0052] In one embodiment, the modified genomic DNA contains mutations associated with DSB repair. The number of mutated bases is defined as the number of substitutions or gaps in sequence alignment with the original genomic DNA. In a more specific embodiment, the modified genomic DNA contains, for example, 1 to 10, 1 to 5, 1 to 3, 1, 2, or 3 mutations within 20 base pairs on both sides of at least one cleavage site. However, these mutations may inhibit the intended cleavage or impair the function of the intended gene. Therefore, in the production method of the present invention, it is preferable that the modified genomic DNA does not contain mutations associated with DSB repair.
[0053] Therefore, it is more preferable that the production method of the present invention includes a step of selecting cells that do not have mutations associated with DSB repair.
[0054] (Existing gene) The existing gene is not particularly limited as long as it is possible to create a modified gene sequence by deletion manipulation, and is preferably a gene that is not essential for cell survival, or a gene that is essential for cell survival but has redundancy due to the existence of multiple similar genes.
[0055] In one embodiment, the existing gene is a gene for a highly expressed protein. This embodiment is preferable because the promoter region and ribosome binding sequence can be used as is, and the desired modified gene can be produced with a reduced number of operations in step (a). Examples of such highly expressed protein genes include genes that express proteins in wild-type cells at an amount that accounts for 10% or more, 15% or more, 20% or more, or 30% or more of the total protein amount. Examples of such genes include structural proteins (e.g., collagen) in animal cells and storage proteins (e.g., seed storage proteins) in plant cells.
[0056] (modified gene) The modified genomic DNA includes a modified gene configured to express a polypeptide, a functional RNA, or both. The region of the genomic DNA before modification that is included in the modified gene is not particularly limited. In one embodiment, the modified gene comprises or consists of at least one selected from the group consisting of the 5'- and 3'-terminal non-coding regions of an existing gene, a regulatory region and a transcribed region of an existing gene. In one embodiment, the modified gene comprises or consists of at least one selected from the group consisting of the 5'- and 3'-terminal non-coding regions of an existing gene, a regulatory region, a coding region, an intron, a 5'-UTR and a 3'-UTR of an existing gene.
[0057] The region of the genomic DNA before modification contained in the promoter region of the modified gene is not particularly limited. In one embodiment, it consists of at least one region selected from the group consisting of the promoter region of an existing gene and the non-coding region adjacent to the 5' side thereof. In such a case, the risk of deleting sequences downstream of the promoter region of the existing gene (e.g., ribosome binding sequence, transcription regulatory sequence) can be reduced, and modification can be achieved by performing step (a) a fewer number of times.
[0058] The region of the genomic DNA before modification contained in the transcribed region of the modified gene is not particularly limited. In one embodiment, the transcribed region of the modified gene consists of the transcribed region of an existing gene. In such a case, it is preferable because it may be possible to use the promoter region of the existing gene as is.
[0059] In one embodiment, the existing gene and the modified gene encode a polypeptide, and the coding region of the modified gene comprises at least one selected from the group consisting of the coding region, intron, 5'-UTR, and 3'-UTR of the existing gene, more preferably at least one selected from the group consisting of the coding region, intron, and 3'-UTR. This is preferable because it may be possible to utilize the promoter region and ribosome binding sequence of the existing gene as is. In one embodiment, the promoter region, coding region, or both of the modified gene comprise the full length or a portion of at least one linked fragment. In one embodiment, the coding region of the modified gene comprises the full length or a portion of at least one linked fragment (see, for specific examples, FIG. 2 and Examples 1 to 4).
[0060] When the modified gene encodes a polypeptide, a start codon or stop codon of the modified gene may be generated by the ligation in step (a) (see, for example, Example 2-2 in Figure 2 and Examples 1 to 4). The ligation may result in a frameshift, and the frameshift may generate a start codon or stop codon in a region different from that of the existing gene. When a new start codon or stop codon of the modified gene is generated, the start codon or stop codon of the existing gene may be removed by deletion, or the start codon or stop codon of the existing gene may be replaced with another codon during ligation (see, for example, Example 2-2 in Figure 2).
[0061] It is also possible to form a coding region longer than the coding region of an existing gene by step (a) (e.g., Example 2-2 in Figure 2, Example 2). In such an embodiment, if the codon reading frame of the existing gene is maintained, the modified gene encodes a fusion protein of the existing gene.
[0062] It is also possible to generate a coding region that is shortened from an existing gene, as shown in Examples 2-3 of Figure 2 and Examples 1, 3, and 4.
[0063] In one embodiment, the promoter region of the modified gene comprises the entire length or a portion of at least one of the linked fragments (see Example 5 for an example). In this embodiment, a part or all of the promoter region of an existing gene is deleted to obtain a promoter consisting of a different nucleotide sequence. To construct a modified promoter region, for example, an existing promoter region and its adjacent 5'-terminal non-coding region can be used.
[0064] In one embodiment, the sequence of the promoter region of the modified gene can be a promoter with higher transcription activity than that of the existing gene. This embodiment is preferable because the modified gene obtained by the production method of the present invention has higher transcription activity and is capable of producing a modified gene product. The transcription level of the modified gene can be, for example, 1.1 times or more, 1.2 times or more, 1.5 times or more, 1.8 times or more, 2 times or more, or 3 times or more than the transcription level of the existing gene.
[0065] In one embodiment, the promoter of the modified gene has a sequence different from that of the promoter of the existing gene, in order to increase the expression level of the modified gene product. In the case of human cells, CMV promoter, EF1α promoter, SV40 promoter, RSV promoter, etc. In the case of plant cells, the RuBisco promoter, the ADH promoter, the cauliflower mosaic virus (CaMV) 35S promoter, the CaMV 19S promoter, the El2 omega promoter, the NOS promoter, etc. In the case of yeast cells, Gal1 / 10 promoter, PGK promoter, ADH promoter, PHO5 promoter, etc. In the case of E. coli, examples include the T7 promoter, trp promoter, lpp promoter, and recA promoter.
[0066] As shown in the following embodiments, if the length of the region formed by the linked fragments in the modified DNA (i.e., the total length of the linked fragments) is a nucleotide sequence encoding approximately 60 amino acid residues or less, or a nucleotide sequence of 180 base pairs or less, the number of cleavage steps and the number of (a) steps required to obtain modified genomic DNA can be reduced. For example, the number of cleavage sites can be 100 or less, and the number of (a) steps can be 50 or less; the number of cleavage sites can be 50 or less, and the number of (a) steps can be 25 or less; the number of cleavage sites can be 50 or less, and the number of (a) steps can be 20 or less; the number of cleavage sites can be 20 or less, and the number of (a) steps can be 10 or less; or the number of cleavage sites can be 10 or less, and the number of (a) steps can be 5 or less.
[0067] When an existing gene and an altered gene encode a polypeptide, and the coding region of the altered gene contains the entire length or a portion of at least one of the linked fragments, it is preferable that the entire length or a portion of the linked fragment encodes an altered polypeptide region that is not present in the polypeptide region of the existing gene. For example, when a linked fragment of m base pairs is present in the coding region of the altered gene, the length of the altered polypeptide region is expressed as an integer length of 1 / 3m, rounded down to the nearest integer. In certain embodiments, the length of the altered polypeptide region is 3 to 60 or 3 to 50 amino acids, preferably 3 to 40, 3 to 30, 3 to 20, 3 to 15, or 3 to 10 amino acids.
[0068] In one embodiment, the modified polypeptide region comprises, for example, a functional peptide sequence, an enzyme cleavage sequence, or both. Here, the functional peptide sequence is not particularly limited, but includes, for example, growth factors (epidermal growth factor (EGF), platelet-derived growth factor (PDGF) peptide, etc.), peptide hormones (insulin, glucagon, vasopressin, oxytocin, growth hormone, gastrin, cholecystokinin, secretin, atrial natriuretic peptide (ANP), melanocyte-stimulating hormone (MSH), adrenocorticotropic hormone (ACTH), angiotensin, orexin, endothelin, ghrelin, etc.), peptide pheromones; plant stem cells; Examples of such peptides include signal peptides (CLE41 / 44 (TDIF), EPFL4 / 6, CLE9 / 10, CLE25, CLE46, etc.); peptide sequences with functions such as blood pressure lowering (angiotensin-converting enzyme (ACE) inhibitor peptides, etc.), mineral absorption promotion, cholesterol regulation, immunomodulation, antioxidant, anti-fatigue, anti-stress, beta-amyloid toxicity alleviation, bifidobacterial growth promotion, flavor, vaccine (e.g., peptides containing epitope sequences of tumor-associated antigens or viral antigens); and affinity (6xHis tag, etc.).
[0069] The modified polypeptide region may further comprise a linker sequence.
[0070] In one embodiment, the full-length polypeptide encoded by the modified gene is 50% or more, for example, 60% or more, 70% or more, 80% or more, or 90% or more, of the full-length polypeptide encoded by the existing gene. A typical example of such an embodiment is one in which the modified gene encodes a fusion protein in which a short modified polypeptide region is added to the N-terminus, C-terminus, or internal region of a protein or domain thereof encoded by the existing gene.
[0071] In another embodiment, the full length of the polypeptide encoded by the modified gene is 50% or less, e.g., 40% or less, 30% or less, 20% or less, or 10% or less, of the full length of the polypeptide encoded by the existing gene. A typical example of such an embodiment is one in which the full length of the coding region of the modified gene consists of a nucleotide sequence that encodes a short modified polypeptide that differs from the sequence of the existing gene.
[0072] (cell) The cells are eukaryotic or prokaryotic cells, preferably eukaryotic cells. As used herein, the cells produced by the production method of the present invention include proliferated cells (replicates) obtained by culturing or the like after production.
[0073] Eukaryotes include, for example, animals, plants, fungi, protists, etc., and are preferably animals or plants.
[0074] Examples of animals include mammals such as humans, mice, rats, rabbits, monkeys (chimpanzees, gorillas, orangutans, rhesus monkeys, green monkeys, etc.), sheep, goats, cows, horses, pigs, guinea pigs, dogs, cats, and hamsters; birds such as parakeets and parrots; reptiles such as lizards and snakes; amphibians such as frogs and salamanders; fish such as salmon, tuna, bonito, sea bream, yellowtail, eels, and killifish; and animals of the phylum Arthropoda, such as insects and crustaceans.
[0075] Examples of plants include grasses such as wheat, rice, barley, oats, rye, corn, sugarcane, foxtail millet, and barnyard millet; legumes such as soybeans, adzuki beans, peas, and kidney beans; solanaceae plants such as tobacco, tomato, eggplant, chili pepper, and potato; cucurbits such as pumpkin, watermelon, cucumber, melon, and Japanese cantaloupe; seed plants such as Arabidopsis, buckwheat, cassava, sweet potato, taro, mulberry, pine, cedar, cypress, ginkgo, and eucalyptus; ferns; and mosses.
[0076] Examples of fungi include yeasts such as those of the genera Saccharomyces, Pichia, Schizosaccharomyces, and Candida; filamentous fungi such as those of the genera Rhizopus and Aspergillus; dimorphic fungi such as those of the genus Penicillium; and mushrooms.
[0077] Examples of protists include algae (green algae, red algae, brown algae, cyanobacteria, Euglenophyta, Haptophyta, Cryptophyta, etc.), ciliates, and amoeba.
[0078] Examples of prokaryotes include bacteria (Escherichia coli, Bacillus subtilis, thermophilic bacteria (such as the genus Thermus), rhizobia, cyanobacteria, etc.) and archaea.
[0079] The cell is not particularly limited as long as it is genome-editable. In one embodiment, the cell is a cell in tissue isolated from the living body of a human or non-human organism (ex vivo cell; for example, a cell derived from an excised organ, plant leaf, stem, etc.), or an in vitro cell (primary culture cell, passaged cell, cultured cell differentiated from stem cell such as iPS cell, etc.). In one embodiment, the cell is an in vitro cell. In another embodiment, the cell is a cell in the living body of a non-human organism (in vivo). In another embodiment, the cell is a cell in the living body of a human.
[0080] The type of cell is also not particularly limited as long as it is genome-editable. When creating a genome-edited organism, germ cells, pluripotent cells (iPS cells, ES cells, etc.), or cells that can be dedifferentiated (e.g., plant cells) are preferred.
[0081] In one embodiment, the cells are cells used in food or in the production of food. Such cells are preferably (i) those contained in food, such as those generally consumed as food under Article 7, Paragraph 2 of the Food Sanitation Act of Japan, or (ii) those generally used in the production of such food. Here, "food" includes not only so-called general foods but also food additives. Specific examples of (ii) include cells used in the production of fermented foods such as soy sauce and sake, as well as cells used for the fermentation production of amino acids, enzymes for food production, other food additives, and the like.
[0082] (sequence-specific endonuclease, target sequence) To delete a region in genomic DNA in step (a), the genomic DNA is cleaved with a sequence-specific endonuclease specific to a target sequence. As used herein, the term "target sequence" refers to a sequence that the sequence-specific endonuclease needs to identify and cleave.
[0083] The sequence-specific endonuclease is not particularly limited as long as it can cleave a target sequence in a target cell, but it is preferable that it recognizes and cleaves a target sequence that is unique to the genomic DNA of the cell (e.g., a specific target sequence of 16 or more bases or 20 or more bases). Examples of such endonucleases include the CRISPR / Cas system, zinc finger nucleases (ZFNs), TALENs (transcription activator-like effector nucleases), meganucleases, etc. The sequence-specific endonuclease may be a wild-type enzyme or a modified mutant.
[0084] The sequence-specific endonuclease used may be of a different type or the same type for each target sequence and each step (a). From the viewpoint of simplifying the procedure, it is preferable that all the sequence-specific endonucleases are of the same type.
[0085] If the cell is a eukaryotic cell, the sequence-specific endonuclease preferably contains at least one nuclear localization signal (NLS).
[0086] Among these, the CRISPR / Cas system is preferred as a sequence-specific endonuclease because it can easily impart specificity to a specific target sequence. The CRISPR / Cas system includes a Cas protein with endonuclease activity and a guide RNA that specifies the target sequence. The Cas pairs with the guide RNA and cleaves nucleotides containing a protospacer adjacent motif (PAM) sequence at a specific position. The Cas protein and guide RNA may be naturally occurring or may be a combination that does not occur in nature.
[0087] Cas9 or Cas12a (Cpf1) are preferred Cas proteins because they have DNA cleavage activity and pinpoint cleavage at target sequences. CRISPR / Cas9 forms blunt ends regardless of whether the PAM sequence is present on the sense or antisense strand. Therefore, Cas9 has the advantage of being less restricted in target sequences than other Cas proteins.
[0088] In nature, CRISPR / Cas9 contains crRNA and tracrRNA as guide RNA components, but the production method of the present invention more preferably uses a system using single-stranded guide RNA (sgRNA), in which Cas, tracrRNA, and the target sequence are combined into a single RNA.
[0089] When using the CRISPR / Cas9 system, Cas9 proteins, guide RNAs, and other components can be selected based on published literature to be compatible with the target cells. Cas9 proteins are preferably Cas9 proteins from Staphylococcus, more preferably Cas9 proteins from Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus. These Cas9 proteins may be wild-type or mutant, as long as they have target sequence specificity and DNA double-strand cleavage activity.
[0090] The restriction of nucleotides by PAM sequences complicates the design of guide RNA. Therefore, Cas9 with a small number of positions restricted to specific nucleotides in the PAM sequence is particularly suitable for use. Examples of such Cas9 include Cas9 (SpCas9) derived from Streptococcus pyogenes that recognizes NGG as a PAM sequence, a variant of SpCas9 (SpCas-NG) that recognizes NG as a PAM sequence (Nishimasu, H. et al., 2018, Science, Vol. 361, pp. 1259-1262), xCas9-3.7 (Hu, JH et al., 2018, Nature, Vol. 556, pp. 57-63), ScCas9 that recognizes NNG (Chatterjee, P. et al., 2018, Sci. Adv. Vol. 4, eaau0766), and ScCas9. ++ (Chatterjee, P. et al. Nat. Biotechnol., 2020, Vol. 38, pp. 1154-1158), SpG, which recognizes NGN (Walton RT et al., 2020, Science, Vol. 368, pp. 290-296), or functional analogs thereof.
[0091] In SpCas9 and SpCas-NG, the target sequence consists of 5'-N(17)-(Cas cleavage site)-NNN-PAM sequence-3', where N(17) represents any 17-nucleotide sequence. The three nucleotides immediately preceding the PAM sequence are referred to herein as the spacer sequence.
[0092] When the sequence-specific nuclease is a zinc finger nuclease (ZFN), TALEN (transcription activation-like effector nuclease), or meganuclease, the cleavage in step (a) can be carried out by allowing two nucleases with different target sequences to coexist in a cell.
[0093] When target sequence-specific cleavage is performed using the CRISPR / Cas system, for example, one type of Cas and two types of guide RNAs that pair with the two target sequences to be cleaved can be used in a single step (a) in a cell. However, the Cas used in each step (a) does not need to be the same. When using multiple Cas, target sequences containing appropriate PAM sequences can be selected depending on the Cas.
[0094] (Step of selecting cells that do not contain off-target mutations, step of removing off-target mutations) In addition to the above step (a), the production method of the present invention preferably further comprises step (b) of selecting cells that do not contain off-target mutations. As used herein, the term "off-target mutation" refers to a mutation that occurs as a result of cleavage and DSB repair at a site that was not originally intended to be cleaved.
[0095] In one embodiment, step (b) is performed after step (a) has been performed one or more times, for example, in a single operation. In one embodiment, step (b) is performed after all deleted regions have been excised from the genomic DNA.
[0096] The presence or absence of off-target mutations is determined, for example, by genomic sequencing of the modified genomic DNA.
[0097] The production method of the present invention may further include a step of removing off-target mutations in addition to the above steps (a) and (b). Specific examples of methods for removing off-target mutations include the step of producing an organism containing a modified genome, which will be described later, and backcrossing (backcrossing) an organism containing a modified genome with a non-recombinant organism of the same species that does not contain the off-target mutations (e.g., a wild-type organism having cells before modification).
[0098] (Insertion and removal of foreign DNA into genomic DNA) The production method of the present invention may further comprise the steps of inserting and removing foreign DNA into genomic DNA. As used herein, removal of foreign DNA refers to eliminating foreign DNA and its replicas from the cell. In one embodiment, removal of foreign DNA is carried out by step (a).
[0099] To insert foreign DNA into genomic DNA, for example, a technique can be used that combines genomic DNA cleavage with sequence-specific endonuclease and DNA repair by recombinational repair. Genomic DNA cleavage can be performed at, for example, one or two locations depending on the purpose, as shown in Figures 3A and 3B described below. The two cleavages can be performed, for example, by the cleavage in step (a). Homologous recombination repair (HRD) is preferred for recombinational repair. Regions homologous to sequences on genomic DNA (referred to as the first and second homology arms) are added to both ends of the foreign DNA so that recombinational repair can occur at appropriate locations.
[0100] The length of the first homology arm and the second homology arm is not particularly limited as long as it allows foreign DNA to be inserted, and can be, for example, 5 base pairs (bp) or more, 10 bp or more, 20 bp or more, 50 bp or more, 100 bp or more, 200 bp or more, or 500 bp or more, and can be, for example, 10,000 bp or less, 5,000 bp or less, 2,000 bp or less, or 1,000 bp or less.
[0101] In one embodiment, the foreign DNA is inserted into the deleted region or adjacent to the deleted region, so that the entire deleted region can be removed in step (a).
[0102] In one embodiment, the foreign DNA is inserted to replace part or all of the deleted region, i.e., when the deleted region is excised by cleavage with a sequence-specific endonuclease in step (a), the foreign DNA is inserted by recombinational repair (see, for example, Figure 3A).
[0103] To enable cleavage of the foreign DNA in step (a), the foreign DNA may be provided with a sequence essential for cleavage of the target sequence, such as a PAM sequence.
[0104] In one embodiment, the first or second homology arm encompasses the sequence of at least one linked fragment (referred to as a "first linked fragment"), and all or part of the sequence of the linked fragment or genomic DNA (referred to as a "ligated portion") to be ligated to the first linked fragment in step (a) is connected to at least one of the ends of the first linked fragment. According to this embodiment, insertion of foreign DNA can be regulated depending on the ligation between the first linked fragment and the ligated portion (for a similar example regarding guide RNA, see the section <Design of guide RNA encompassing linked fragments>).
[0105] The length of the first linked fragments is not particularly limited as long as it is within the range of the length of the homology arms. However, from the viewpoint of reducing the risk of recombination, not limited to the ligation of the first linked fragments, the length of the first linked fragments is preferably, for example, 5% or more of the length of the homology arms, more preferably 10% or more, even more preferably 15% or more, and even more preferably 20% or more. From the same viewpoint, it is also preferable that the ends of the first linked fragments are relatively close to the ends of the foreign DNA. On the other hand, from the viewpoint of reducing the risk of losing the sequence of the first linked fragments due to cutting away by nuclease, it is preferable that the ends of the first linked fragments are separated from the ends of the foreign DNA. Therefore, it is preferable that the ends of the first linked fragments are located, for example, 1 to 30 bp, more preferably 3 to 20 bp, and even more preferably 5 to 15 bp from the ends of the foreign DNA.
[0106] In a more specific embodiment, the foreign DNA includes a marker gene for cell selection. In such an embodiment, the deleted region containing or adjacent to the foreign DNA can be labeled with the marker gene.
[0107] Cells into which the foreign DNA has been inserted can be selected based on the expression of a marker gene.
[0108] Then, after excising the foreign DNA containing the marker gene, or after excising the region containing the marker gene and the deleted region together, cells that do not express the marker gene can be easily selected to select cells in which the deleted region has been excised. Among the embodiments of the present invention, in embodiments containing a short deleted region (e.g., 50 bp or less, 30 bp or less, 20 bp or less, 10 bp or less, 5 bp or less, 3 bp or less, or 1 bp), it may be difficult to distinguish cells in which the deleted region has been excised using PCR-based methods. However, by labeling the deleted region with a marker gene, cells in which the deleted region has been excised can be selected without undergoing cumbersome processes such as sequencing.
[0109] Both positive and negative selection marker genes can be used as cell selection marker genes. Positive selection marker genes are genes that allow cells to be selected based on their presence, such as fluorescent proteins (GFP, YFP, CFP, etc.), drug resistance genes (neomycin resistance gene, tetracycline resistance gene, chloramphenicol resistance gene, ampicillin resistance gene, kanamycin resistance gene, sulfonylurea resistance gene (ALS), glyphosate resistance gene (EPSPS), etc.), and reporter enzyme genes (luciferase, β-galactosidase, β-glucuronidase (GUS), dihydrofolate reductase (DHFR), etc.). Negative selection marker genes are genes that allow cells to be selected based on their absence, such as genes encoding toxic proteins and suicide genes (HSV-TK, iCasp9, etc.).
[0110] When a fluorescent protein is used as a marker gene, cells can be easily selected on a large scale using a cell sorter or the like. Furthermore, when a drug resistance gene is used as a marker gene, cells having the resistance gene can be selected by culturing the cells in a medium containing the drug. Therefore, a fluorescent protein or a drug resistance gene is more preferable as the marker gene for cell selection used in the production method of the present invention. In particular, when the marker gene is a fluorescent protein, its absence can be confirmed by the fluorescence intensity of the cells, and both the presence and absence of foreign DNA can be confirmed relatively easily, making it particularly preferable.
[0111] Figure 3A shows an example of an embodiment in which the deletion region (the second region in the figure) is deleted together with the insertion of foreign DNA containing a marker gene. The insertion of foreign DNA can be performed during step (a). If the foreign DNA used is one that is not maintained intracellularly, cells containing the marker gene can be selected after cell culture to obtain cells with genomic DNA in which the second deletion region has been replaced with the marker gene. Next, the foreign DNA is removed by cleaving the genomic DNA at two sites with a sequence-specific nuclease, and then cells lacking the marker are selected to select cells from which the foreign DNA has been removed.
[0112] Figure 3B shows an example of an embodiment in which foreign DNA containing a marker gene is inserted adjacent to the deleted region (the second region in the figure), and then the foreign DNA and the deleted region are removed together. The insertion is achieved by single-site cleavage with a sequence-specific nuclease and HDR. The foreign DNA and the deleted region are then excised and removed together in step (a). Cell selection can be performed as in Figure 3A.
[0113] In one embodiment, two or more marker genes are inserted within or adjacent to the deleted region. In such an embodiment, the marker genes are preferably different. The marker genes are preferably inserted by inserting a single foreign DNA containing multiple marker genes into the genomic DNA (see Figures 3A and 3B). However, the marker genes can also be inserted into the genomic DNA by dividing the foreign DNA into multiple marker genes.
[0114] (Step of introducing sequence-specific endonuclease and / or its functionally associated factor) In one embodiment, the production method of the present invention can further include a step of introducing foreign DNA containing the target sequence-specific endonuclease or its gene into a cell in order to allow the above-mentioned target sequence-specific endonuclease to function in the cell, and a step of removing the foreign DNA.
[0115] When the target sequence-specific endonuclease is a CRISPR / Cas system, the method may further include the step of introducing an appropriate guide RNA or a foreign DNA capable of expressing the guide RNA in cells. These steps may be performed simultaneously or separately.
[0116] When the sequence-specific endonuclease and guide RNA are directly introduced into cells, it is preferable to carry out this before each set of steps (a).
[0117] In one embodiment, the foreign DNA is present in the cell in a form separated from the genomic DNA, making it easier to completely remove it from the cell. Such foreign DNA can be introduced into the cell as a vector, such as a plasmid, cosmid, or artificial chromosome. When the foreign DNA is separated from the genomic DNA, the foreign DNA is naturally lost, and the foreign DNA can sometimes be removed by selecting cells that do not contain the foreign DNA.
[0118] In another embodiment, foreign DNA is inserted into genomic DNA. Such foreign DNA can be introduced into cells as, for example, linear DNA, a viral vector, etc. In this embodiment, the production method of the present invention further comprises a step of removing the foreign DNA. Optionally, it may also comprise a step of selecting cells from which the foreign DNA has been removed.
[0119] To confirm that a sample does not contain foreign DNA, for example, primers capable of specifically amplifying foreign DNA can be designed and used for PCR amplification, and the foreign DNA can be detected by various PCR methods. The absence of foreign DNA in genomic DNA can also be confirmed by genome sequencing, for example.
[0120] The gene encoding the target sequence-specific endonuclease of the above-mentioned foreign DNA may be adjusted to have a codon usage frequency similar to that of the cell (so-called codon optimization) in order to improve expression in the cell.
[0121] In addition to the gene of interest, for example, a promoter, an enhancer, an insulator, an intron, a terminator, a poly(A) addition signal, a selection marker gene, etc. can be ligated to the vector.
[0122] The target gene to be inserted into a vector may be one or more types per vector.
[0123] As used herein, the introduction of substances such as nucleotides and proteins into cells is not particularly limited as long as it is a means capable of delivering RNA and proteins to living cells, and can be carried out by, for example, the liposome method (lipofection, etc.), particle gun (gene gun) method, electroporation method, polyethylene glycol (PEG) method, plasma method (see, for example, WO2018016217), whisker method, laser injection method, etc. When the cells are plant cells, the particle gun method is preferred.
[0124] (Design process of ligated fragments and target sequences) By determining the location of cleavage by a sequence-specific endonuclease in a genomic DNA region containing an existing gene, the target sequence recognized by each endonuclease can be determined. The cleavage site can be determined, for example, by the following steps (i) to (v), without any particular limitation. The design of the target sequence, including the following steps, can be performed manually or based on a computer program or a trained machine learning or artificial intelligence (AI) algorithm.
[0125] (i) First, the DNA sequence to be included in the modified gene is prepared. When creating a coding region, multiple sequences with different codon combinations can be used as candidates. Furthermore, because frameshifts can occur due to deletions in step (a), sequences can be generated with a partially frameshifted reading frame. Therefore, if the linked fragments cannot be mapped in order in step (iv) below for the DNA sequence of a coding region, steps (ii) to (iv) can be repeated for substitution with synonymous codons or frameshifted sequences.
[0126] (ii) The sequence of (i) is fragmented into two or more linked fragments. The boundary between the linked fragments is the cutting point. The smaller the linked fragments, the easier it is to find the target existing gene, but this has the disadvantage of increasing the number of steps (a). On the other hand, the larger the linked fragments, the more limited the number of target existing genes, but the greater the possibility of suppressing the increase in step (a).
[0127] (iii) If there is a sequence essential for cleaving the target sequence of the sequence-specific nuclease to be used (e.g., a PAM sequence in the CRISPR / Cas system), add the sequence to the resulting ligated fragment to design a search sequence.
[0128] (iv) The search sequence designed in (iii) above is mapped onto genomic DNA. For a given existing gene, it is examined whether the search sequence can be mapped onto the existing gene and its adjacent non-coding regions so that the linked fragments are arranged in the same order as the linked fragments of the modified gene. For example, multiple linked fragments may be mapped at once, and one in which the linked fragments are arranged in the same order as the modified gene may be selected. However, sequential mapping of the search sequences one by one is preferable because it reduces the amount of calculation required for the search. That is, if a second linked fragment exists adjacent to a first linked fragment, the first search sequence corresponding to the first linked fragment is first mapped, and then it is examined whether the second search sequence corresponding to the second linked fragment can be mapped near the mapped first search sequence. The vicinity may be set, for example, within 1000 base pairs (bp), 500 bp, 300 bp, 200 bp, 100 bp, or 50 bp of the search sequence.
[0129] (v) If all the linked fragments can be mapped to the existing gene in the same order as the modified gene, the cleavage sites on the genomic DNA will be identified. The target sequence can then be determined so that cleavage occurs at the identified cleavage sites. Among the mappable fragments, the one with the fewest number of linked fragments is preferred in terms of minimizing the number of steps (a). In one embodiment, the target sequence is determined based on the genomic DNA sequence before modification. In another embodiment, the target sequence is determined based on the genomic DNA sequence during modification, i.e., the genomic DNA from which one or more deletion regions have been excised. Note that one of these embodiments can be selected independently for each target sequence corresponding to each cleavage site.
[0130] A more specific embodiment of (iii) using CRISPR / Cas9 can be carried out, for example, by the following procedure.
[0131] (iii)-1: A sequence (search sequence) is generated by adding a PAM sequence to each ligated fragment. Since Cas9 forms blunt ends, the PAM sequence can be present on either the sense or antisense strand. Therefore, the positional relationship with the cleavage site on the sense strand can be either (1) 5'-(cleavage site)-NNN-(PAM sequence)-3' or (2) 5'-(complementary sequence of the PAM sequence)-NNN-(cleavage site)-3' (Figure 4A).
[0132] When SpCas-NG was used, search sequences generated by placing PAM sequences so as to cleave both ends of ligated fragments of various lengths are shown in Figures 4B and 4C.
[0133] (iii)-2: When one of the two cleavage sites is cleaved and ligated, the PAM sequence for the other cleavage site may be lost. For example, if the PAM sequence is located on the right side of the ligated fragments, as in Type 1 in Figure 4D, cleavage and ligation on the right side of the ligated fragments may prevent cleavage on the left side. Conversely, if the PAM sequence is located on the left side of the ligated fragments, as in Type 2, cleavage and ligation on the left side may prevent cleavage on the right side. Therefore, when using these types of PAM sequences, it is preferable to limit the cleavage order, or if the cleavage order is not limited, to apply Type 1 to the 5'-end and Type 2 to the 3'-end of the ligated fragments. On the other hand, when two PAM sequences are not contained in the ligated fragments (Type 3 in Figure 4D) or when two PAM sequences are present in the ligated fragments (Type 4 in Figure 4D), cleavage on the other side is possible even after the first cleavage, and this is more preferable because there is no restriction on the cleavage order. Among the possible cleavage sites on the genomic DNA, the cleavage site closest to the 5' or 3' end can be designed to contain either the sequence (1) or (2) in Figure 4A (see types 5 and 6 in Figure 4E).
[0134] Based on the cleavage site and PAM sequence identified in steps (i) to (iv) above, guide RNAs for the CRISPR / Cas system can be designed. When using Cas9, guide RNAs can be designed that contain the PAM sequence and a 20-nucleotide sequence 5' upstream of it (17 nucleotides + 3 nucleotide spacer sequence) in the target sequence.
[0135] Optionally, the method may further include a step of predicting off-target sequences based on the target sequences obtained, and a step of selecting target sequences that are less likely to be off-targets. Off-target prediction can be performed using, for example, a known prediction service or software.
[0136] <Design of guide RNA containing linked fragments> When a certain linked fragment (referred to as the first linked fragment) is 16 base pairs or less, it is preferable to design and use a guide RNA such that the sequence of at least one linked fragment is included in the target sequence, and the linked fragment, genomic DNA, or foreign DNA (referred to as the "ligated portion of the first linked fragment") to be ligated to the first linked fragment in step (a) is located at either end of the first linked fragment. The first linked fragment may be included within the 17 nucleotides on the 5' side of the cleavage site, or within the 6 nucleotides consisting of the spacer sequence and PAM sequence on the 3' side. The ligated portion may be located on the 5' side, the 3' side, or both the 5' and 3' sides of the first linked fragment. Multiple linked fragments (e.g., two, three, four, or five) may be included.
[0137] A guide RNA containing the above-mentioned linked fragments can be used not only in step (a) of the present invention, but also for inserting foreign DNA into genomic DNA using a sequence-specific endonuclease.
[0138] Figure 5A (I) and (II) show examples of regions where the ligated portion is present on the 5' and 3' sides of the first ligated fragment, respectively, and consists of the target sequence of the guide RNA and the PAM sequence. The portion on the 5' side that is not the ligated region (e.g., the left end portion of Figure 5A (II)) may be derived from the (adjacent) ligated fragment, a genomic DNA region, a deleted region, or foreign DNA inserted into the genomic DNA. In (II), the 3' side of the first ligated fragment further includes the sequence of another ligated fragment.
[0139] The guide RNA containing the above-mentioned linked fragment can cleave the target cleavage site only after the first linked fragment and its ligated portion are ligated. This reduces the possibility of unintended ligation occurring when multiple steps (a) are performed simultaneously or consecutively using three or more types of guide RNA. Furthermore, by using this guide RNA, step (a) proceeds only if the included linked fragment does not have an unintended ligation (e.g., ligation with another linked fragment or inverted ligation) or a mutation associated with DSB repair at the ligated portion. Therefore, the use of such a guide RNA reduces the risk of unintended mutations remaining.
[0140] FIG. 5B shows a case where cleavage and ligation are carried out in the coexistence of three types of guide RNAs (guide RNAs 1 to 3), including a guide RNA (guide RNA 3) that encompasses the ligated fragment. In the first cleavage, guide RNAs 1 and 2 are cleaved and joined, but cleavage by guide RNA 3 does not occur. This ensures that fragments 1 and 3 are joined in the correct combination after the first cleavage. Then, in the second cleavage, guide RNA 3 cleaves the genomic DNA, with fragments 1 and 3 joined in the correct orientation, between fragments 3 and 4. On the other hand, when three guide RNAs capable of cleaving the sequence before ligation are used without including the ligated fragment, cleavage occurs with guide RNAs 1 to 3, generating fragments 1 to 4. Since ligation between fragments 1 to 4 can occur in various combinations, the efficiency of producing genomic DNA in which the ligated fragments are correctly ligated decreases.
[0141] (Other processes) The production method of the present invention preferably includes a step of determining the sequence of the obtained modified genomic DNA. The production method of the present invention preferably further includes a step of selecting cells having the modified genomic DNA of interest based on the determined sequence. By including these steps, cells having the modified genomic DNA of interest can be isolated and concentrated.
[0142] The production method of the present invention may further include a step of growing cells containing the modified genomic DNA. The cell growth can be performed using known methods used in growing the original cells (e.g., in vitro cell culture using a medium that can be used to culture the cells). The step of growing cells containing the modified genomic DNA can be performed before, during, or after the production of the modified genomic DNA, but is preferably performed after the production of the cells.
[0143] In the production method of the present invention, the cells may be subjected to an appropriate dedifferentiation step, differentiation step, etc. depending on the intended use. These steps can be carried out before, during, or after the production of modified genomic DNA.
[0144] [Cells containing modified genomic DNA] In one embodiment of the present invention, the cell whose genomic DNA has been modified comprises: The genomic DNA of the cell At least two deleted regions (deleted regions) in an existing gene and its adjacent non-coding region of the wild-type genomic DNA of the cell, as compared with the wild-type genomic DNA; At least one linked fragment corresponding to a region sandwiched between deleted regions (deleted regions) in the existing gene; a modified gene whose nucleotide sequence is altered relative to said existing gene by said deletion; The modified gene is characterized in that it is configured to express a polypeptide, a functional RNA, or both.
[0145] In a more specific aspect, a cell having modified genomic DNA, which is one embodiment of the present invention, comprises: A cell having modified genomic DNA comprising: The genomic DNA is At least two deleted regions (deleted regions) in an existing gene and its adjacent non-coding region of the wild-type genomic DNA of the cell, as compared with the wild-type genomic DNA; At least one linked fragment corresponding to a region sandwiched between deleted regions (deleted regions) in the existing gene; a modified gene whose nucleotide sequence is altered relative to the existing gene by the deletion; and at least one of the linked fragments is 1 to 180 base pairs in length; the promoter region, the transcribed region, or both of the modified gene comprises the full length or a portion of at least one of the linked fragments; The modified gene is characterized in that it is configured to express a polypeptide, a functional RNA, or both.
[0146] Specific embodiments of the above-mentioned cells, existing genes, modified genes, deleted regions, and linked fragments conform to the respective embodiments described in the above-mentioned production methods of the present invention. Specific embodiments of the positional relationship between the existing gene or modified gene and the linked fragments or deleted regions also conform to the respective embodiments described in the above-mentioned production methods of the present invention.
[0147] The method for producing the cell of the present invention in which genomic DNA has been modified is not particularly limited, and the cell can be produced, for example, by the production method of the present invention described above.
[0148] [Biological body] An organism according to one embodiment of the present invention comprises cells produced by the method of the present invention or cells containing modified DNA of the present invention. These cells may be present throughout the organism or in parts thereof. If present in parts, they may be present locally or scattered. In one embodiment, these cells are present throughout the organism.
[0149] In one embodiment, the organism of the present invention is a non-genetically modified organism.
[0150] The method for producing an organism according to the present invention includes, for example, the following aspects. (I) Cells containing modified DNA are produced from cells in a living organism using the production method of the present invention. (II) Cells produced by the production method of the present invention (including cells replicated after production) or cells containing the modified DNA of the present invention are subjected to known methods for producing transgenic organisms or cloned organisms to produce organisms.
[0151] [Method for producing gene products] The method for producing a gene product of the present invention includes a step of expressing a gene product encoded by a modified gene contained in the modified genomic DNA using cells obtained by the manufacturing method of the present invention, cells containing the modified genomic DNA of the present invention, or an organism of the present invention.
[0152] In one embodiment, the gene product is a polypeptide.
[0153] Expression of the gene product may be appropriately induced using a regulatory sequence such as a promoter of the modified gene.
[0154] The method for producing a gene product of the present invention may further include a step of growing the cells or growing or propagating the organisms used. Growing the cells or growing or propagating the organisms can be carried out, for example, using a medium, feed, fertilizer, or the like that is commonly used for cells or organisms, based on a commonly used method.
[0155] The method for producing a gene product of the present invention may further include a step of isolating the gene product. Examples of the step of isolating the gene product include disruption of cells or organisms, centrifugation, filtration, solubilization, concentration, separation, and purification. The method for producing a gene product of the present invention may further include steps of sterilization, drying, packaging, and the like. Known techniques can be used as appropriate for these steps. [Example]
[0156] The present invention will be specifically explained below with reference to examples, but the present invention is not limited to these examples.
[0157] The following examples show specific examples of combinations of cleavage sites and PAM sequences that can be used to construct specific modified genes from specific genomic sequences using the production methods of the present invention. These were all manually specified based on the genomic DNA sequence. By designing a target sequence based on the specified cleavage site and PAM sequence information and using a CRISPR / Cas9 system appropriate for each cell, those skilled in the art can prepare cells containing modified DNA.
[0158] In each figure in the examples, the wedge above the sequence represents the cleavage site caused by the sense strand PAM sequence, and the wedge below the sequence represents the cleavage site caused by the antisense strand PAM sequence. The sequences enclosed in boxes and written in italics represent the PAM sequence and spacer sequence. The underlined portion represents the ligated fragment or the genomic DNA sequence maintained in the modified genomic DNA. The deleted region is indicated as the region between the two cleavage sites and includes the dashed omission. The numbers in the table indicate the corresponding positions in the genome sequence listed in NCBI (https: / / www.ncbi.nlm.nih.gov / ) using the reference numbers listed in each example.
[0159] Example 1: Modification of the coding region of an existing gene The coding region of the β-conglycinin α subunit 2 (CG-2) gene on chromosome 20 of the soybean (Glycine max) genome (NCBI reference number: NC_038256.2) was modified solely by deletion using SpCas9 (PAM sequence: NGG) to create a gene encoding a polypeptide with four new amino acid residues added to the N-terminal 57 residues of CG-2. To identify the cleavage site, a search sequence mapped based on the design procedure described above is shown in Figure 6A. The modified gene can be constructed through five (a) steps. The stop codon of the existing gene is eliminated by deletion, and a new stop codon is created by ligation. When the polypeptide encoded by the modified gene is produced in soybean cells, it is cleaved at the C-terminus of the asparagine residue by vacuolar protease (VPE) in ripening seeds, producing an umami-enhancing peptide consisting of Ala-Glu-Asp (AED) (Maehashi K. et al., Biosci. Biotechnol. Biochem., 1999, Vol. 63, No. 3, pp. 555-559). The sequence surrounding the modified region shown in Figure 6A, the modified polypeptide, and the identified cleavage site and PAM sequence are shown in Table 1.
[0160] [Table 1]
[0161] [Example 2: Modification of the coding region of an existing gene 2] The 3'-UTR region of lipoxygenase-3 (LOX-3) on chromosome 15 of the soybean (Glycine max) genome (NCBI Reference Number: NC_038251.2) was modified solely by deletion using SpCas9-NG (PAM sequence: NG) to create a gene encoding a polypeptide with a linker and His tag attached to the C-terminus of LOX3. To identify the cleavage site, a search sequence mapped based on the design procedure described above is shown in Figure 6B. The modified gene can be constructed through five steps (a). A new stop codon is generated by ligation. The sequence surrounding the modified site shown in Figure 6B, the modified polypeptide, and the identified cleavage site and PAM sequence are listed in Table 2.
[0162] [Table 2]
[0163] [Example 3: Modification of the coding region of an existing gene 3] A gene encoding a new polypeptide was created by deleting the β-conglycinin α-subunit (CG-2) from the coding region to the 3′-terminal non-coding region of soybean (Glycine max) chromosome 20 (NCBI reference number: NC_038256.2) using only SpCas9-NG (PAM sequence: NG) deletions. The new polypeptide is an N-terminal fragment of CG-2 with four tandem tripeptide sequences possessing angiotensin-converting enzyme (ACE) inhibitory activity. Figure 6C shows the search sequence mapped based on the design procedure described above to identify the cleavage site. The modified gene can be constructed through seven steps (a). The stop codon and terminator sequence of the existing gene are deleted, and a new stop codon is generated by ligation. From the resulting polypeptide, two molecules of Ile-Pro-Pro (IPP) and Val-Pro-Pro (VPP) peptides are obtained using thermolysin. The sequences around the modified portion shown in Figure 6C, the modified polypeptide, and the identified cleavage site and PAM sequence are shown in Table 3.
[0164] [Table 3]
[0165] [Example 4: Modification of the coding region of an existing gene 4] The soybean (Glycine max) genome chromosome 20 (NCBI Reference Number: NC_038256.2) was modified solely by deletion using SpCas9-NG (PAM sequence: NG) from the coding region of the β-conglycinin α-subunit (CG-2) to the 3′-terminal non-coding region to create a gene encoding a protein in which human epidermal growth factor (EGF) is fused to the N-terminal fragment of CG-2. To identify the cleavage site, the location of the ligated fragments mapped based on the above design procedure is shown in Figure 6D, and the search sequence is shown in Figure 6E. The modified gene can be constructed through 33 iterations of step (a). The stop codon and terminator sequence of the existing gene are deleted, and a new stop codon is generated by ligation. The sequence surrounding the modified site shown in Figure 6D, the modified polypeptide, and the identified cleavage site and PAM sequence are listed in Tables 4-1 and 4-2.
[0166] [Table 4-1]
[0167] [Table 4-2]
[0168] [Example 5: Modification of the promoter region of an existing gene] The lipoxygenase-3 (LOX-3) gene on chromosome 15 of the soybean (Glycine max) genome (NCBI Reference Number: NC_038251.2) was modified using only SpCas9-NG (PAM sequence: NG) deletions from the 5' non-coding region to the promoter region to create a gene in which the promoter was changed to a minimal CaMV 35S promoter. To identify the cleavage site, a search sequence mapped based on the design procedure described above is shown in Figure 6F. The modified gene can be constructed through nine steps (a). The sequence surrounding the modified site shown in Figure 6F, the modified polypeptide, and the identified cleavage site and PAM sequence are listed in Table 5.
[0169] [Table 5]
[0170] As described above, even when the existing genes are limited to the CG-2 gene and the LOX-3 gene, modified genes with various sequences can be produced. Therefore, it is easy to understand that various modified genes can be constructed from the genomic DNA sequence by deletion manipulation alone.
[0171] The CG-2 and LOX-3 genes are highly expressed in soybean seeds, and therefore it is expected that the modified polypeptides can be mass-produced from the modified genes of Examples 1 to 4, which have the same promoters as these genes, and from the modified genes of Example 5, which have an even stronger promoter.
Claims
1. 1. A method for generating a cell containing modified genomic DNA that is free of exogenous DNA, comprising: (a) specifically cleaving the genomic DNA of a cell at at least two target sequences with sequence-specific nucleases and joining the sequences by DNA double-strand break repair by the cell; at least twice to delete at least two regions from the existing gene and the non-coding regions adjacent to the existing gene, thereby generating a modified gene that differs from the nucleotide sequence of the existing gene; The genomic DNA of the cell in step (a) At least two deleted regions that are not included in the modified genomic DNA due to the deletion; at least one linked fragment, which is a region between the two deleted regions and which is a portion remaining in the modified genomic DNA; at least one of the linked fragments is 1 to 150 base pairs in length; the promoter region, the transcribed region, or both of the modified gene comprises the full length or a portion of at least one of the linked fragments; The method is configured to express a polypeptide, a functional RNA, or both from the modified gene.
2. The method according to claim 1, wherein at least two of the steps (a) are carried out simultaneously or consecutively using different combinations of the target sequences.
3. 2. The method of claim 1, further comprising the step of (b) selecting the cells that do not contain off-target mutations.
4. The method according to claim 1, wherein at least one of the target sequences comprises the sequence of at least one linked fragment (first linked fragment), and at least one of the ends of the first linked fragment is connected to the sequence of all or part of the linked fragment or genomic DNA that is linked to the first linked fragment in step (a).
5. The method of claim 1, which does not include a step of inserting foreign DNA into the genomic DNA of the cell.
6. The method further comprises a step of inserting, by cellular recombination repair, a foreign DNA having a first homology arm and a second homology arm attached to both ends thereof, capable of homologous recombination with the genomic DNA, into at least one of the deleted regions, inserting the foreign DNA so as to be adjacent to at least one of the deleted regions, or inserting the foreign DNA so as to replace a part or all of at least one of the deleted regions; In the step (a), the foreign DNA is removed from the genomic DNA by the deletion; and The method of claim 1, further comprising the step of selecting said cells for the absence of exogenous DNA.
7. The method of claim 6 , wherein the foreign DNA comprises a cell selection marker gene.
8. the first or second homology arm comprises the sequence of at least one linked fragment (first linked fragment); The method according to claim 6, wherein the first linked fragment has connected to it at least one of its ends a sequence of the linked fragment or genomic DNA that is linked to the first linked fragment in step (a).
9. 2. The method of claim 1, wherein the total length of the linked fragments is 9 to 180 base pairs.
10. The method of claim 1, wherein the total length of the linked fragments is 30% or less of the total length of the deleted region.
11. The method of claim 1, wherein the existing gene and the modified gene encode polypeptides, and in a wild-type cell of the cell, the polypeptides encoded by the existing genes account for 10% or more of the mass of total proteins.
12. the existing gene and the modified gene encode a polypeptide, and the coding region of the modified gene comprises the full length or a portion of at least one of the linked fragments; 2. The method of claim 1, wherein at least one full length or portion of the linked fragment encodes a modified polypeptide region of 3 to 60 amino acids that is not encoded by the existing gene.
13. 13. The method of claim 12, wherein the modified polypeptide region comprises a functional peptide sequence, an enzyme cleavage sequence, or both.
14. The method of claim 1, wherein the existing gene and the modified gene encode polypeptides, and the total length of the polypeptide encoded by the modified gene is 20% or less of the total length of the polypeptide encoded by the existing gene.
15. The method of claim 1, wherein the existing gene and the modified gene encode polypeptides, and the total length of the polypeptide encoded by the modified gene is 50% or more of the total length of the polypeptide encoded by the existing gene.
16. 10. The method of claim 1, wherein the cells are cells used in or for the manufacture of food products.
17. The method of claim 1, wherein the sequence-specific nuclease is performed by a CRISPR / Cas system.
18. 18. The method of claim 17, wherein the CRISPR / Cas system comprises SpCas9-NG or a functional analog thereof.
19. 1. A cell having modified genomic DNA comprising: The genomic DNA is at least two deleted regions (deleted regions) in an existing gene and its adjacent non-coding region of the wild-type genomic DNA of the cell, as compared with the wild-type genomic DNA; At least one linked fragment corresponding to a region sandwiched between deleted regions (deleted regions) in the existing gene; a modified gene whose nucleotide sequence is altered relative to the existing gene by the deletion; and at least one of the linked fragments is 1 to 150 base pairs in length; the promoter region, the transcribed region, or both of the modified gene comprises the full length or a portion of at least one of the linked fragments; A cell configured to express a polypeptide, a functional RNA, or both from said modified gene.
20. The cell of claim 19, wherein the total length of the linked fragments is 9 to 180 base pairs.
21. The cell of claim 19, wherein the total length of the linked fragments is 30% or less of the total length of the deleted region.
22. The cell of claim 19, wherein the existing gene and the modified gene encode polypeptides, and in a wild-type cell of the cell, the amount of protein encoded by the existing gene accounts for 10% or more of the mass of total protein.
23. the existing gene and the modified gene encode a polypeptide, and the coding region of the modified gene comprises the full length or a portion of at least one of the linked fragments; 20. The cell of claim 19, wherein the full length or a portion of at least one linked fragment encodes a modified polypeptide region of 3 to 60 amino acids that is not encoded by the existing gene.
24. 24. The cell of claim 23, wherein the modified polypeptide region comprises a functional peptide, an enzyme cleavage sequence, or both.
25. The cell of claim 19, wherein the existing gene and the modified gene encode polypeptides, and the total length of the polypeptide encoded by the modified gene is 20% or less of the total length of the polypeptide encoded by the existing gene.
26. The cell described in claim 19, wherein the existing gene and the modified gene encode polypeptides, and the total length of the polypeptide encoded by the modified gene is 50% or more of the total length of the polypeptide encoded by the existing gene.
27. 20. The cell of claim 19, wherein the cell is a cell used in a food product or used in the production of a food product.
28. An organism comprising a cell produced by the method of any one of claims 1 to 18, or a cell according to any one of claims 19 to 27.
29. A method for producing a gene product, comprising a step of expressing a gene product encoded by the modified gene in a cell produced by the method of any one of claims 1 to 18, a cell of any one of claims 19 to 27, or an organism containing such a cell.
Citation Information
Patent Citations
Methods for creating novel genes in organisms and uses thereof
JP2022553598A