Gene recombinant arguloida insect

JP2025012136A5Active Publication Date: 2026-04-14NAT AGRI & FOOD RES ORG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NAT AGRI & FOOD RES ORG
Filing Date
2023-07-12
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing methods for creating a stable and large-scale protein expression system in silkworms using the GAL4/UAS system are inefficient due to random genome integration, leading to fluctuating protein expression levels and difficulty in establishing reliable production strains.

Method used

A new expression system is developed by fusing the target protein sequence with the signal peptide of endogenous genes within the exon sequence, utilizing the promoter and enhancer activities directly, and employing genome editing to introduce the target gene sequence into the intron sequence, specifically using TAL-PITCh and homologous recombination methods.

Benefits of technology

This approach allows for stable and high-level production of target proteins in silkworms, overcoming the limitations of the GAL4/UAS system by ensuring consistent and predictable protein expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000030_0001
    Figure 00000030_0001
  • Figure 00000030_0002
    Figure 00000030_0002
Patent Text Reader

Abstract

To provide a new expression system for generating target protein stably in large volume in Arguloida insects such as silkworms.SOLUTION: A gene recombinant Arguloida insect includes a target genetic sequence coding target protein or its fragment in an exon alignment that codes signal peptide of endogenous gene or its functional fragment. The target protein or its fragment is fused with a C terminal side of the signal peptide or its functional fragment.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a genetically modified lepidopteran insect and a method for producing the same. [Background technology]

[0002] The silk gland of the silkworm (Bombyx mori) has the ability to synthesize large amounts of protein in a short period of time. In addition, the silk gland of the silkworm is a large organ, making it easy to extract, and the synthesized protein is stored in the lumen of the silk gland, making it easy to recover. Therefore, transgenic silkworms that express a target protein in the silk gland are considered promising as a mass production system for proteins.

[0003] The silk gland of the silkworm is a pair of organs, one on the left and one on the right, each of which is composed of three regions: the anterior silk gland, the middle silk gland, and the posterior silk gland. In the posterior silk gland cells, three major proteins that constitute fibroin, the fiber component of silk thread, are expressed: fibroin H chain (hereinafter often abbreviated as "Fib H"), fibroin L chain (hereinafter often abbreviated as "Fib L"), and fibrohexamarin (also called p25 / FHX). In addition, in the middle silk gland cells, sericin, a gelatin-like protein that is a coating component of silk thread, is expressed. The three proteins expressed in the posterior silk gland cells form a complex (silk fibroin elementary unit; SFEU complex) in a ratio of Fib H:Fib L:p25=6:6:1, and are secreted into the lumen of the posterior silk gland. In contrast, sericin is secreted into the lumen of the middle silk gland after expression. The fibroin secreted into the lumen of the posterior silk gland then migrates to the lumen of the middle silk gland, where it is coated with sericin and spun into silk threads (Non-Patent Document 1). Therefore, when using the silkworm silk gland as a protein expression system, a gene expression system that is specifically expressed in the middle or posterior silk gland may be used.

[0004] When using silkworm silk glands as a protein expression system, the GAL4 / UAS system (Non-Patent Document 2) and a mass expression method using a system combining the sericin 1 promoter and Hr3 enhancer (Non-Patent Document 3) have been reported as recombinant protein expression systems, but currently the GAL4 / UAS system is widely used due to its superiority in terms of protein expression levels.

[0005] The GAL4 / UAS system is a gene control system that uses a combination of the yeast-derived transcription factor GAL4 and the regulatory sequence UAS. In the GAL4 / UAS system used as a protein production system in the silk gland of silkworms, a GAL4 line expressing the GAL4 gene under the promoter control of a gene specifically expressed in the middle or posterior silk gland, and a UAS line expressing a target protein gene under the control of the regulatory sequence UAS are independently established by genetic recombination using piggyBac, and then the two lines are crossed to construct an expression system that expresses a target protein in the silk gland.

[0006] In the GAL4 / UAS system, it is necessary to establish the GAL4 line and the UAS line separately and then cross them, so it takes time to construct the expression system, and the GAL4 gene and the UAS regulatory sequence are introduced at random positions on the genome, so the expression level of the target protein may vary, which are obstacles to establishing new GAL4 and UAS lines. In addition, the expression level in the GAL4 / UAS system is thought to have reached its technical limit. Therefore, there is a need for new methods for stable and large-scale production of target proteins. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Inoue S. et al., 2000, The Journal of Biological Chemistry, 275 (51): 40517-40528. [Non-Patent Document 2] Tatematsu K. et al., 2010, Transgenic Research, 19(3):473-87. [Non-Patent Document 3] Tomita M. et al., 2007, Transgenic Research, 16 (4):449-465. Summary of the Invention [Problem to be solved by the invention]

[0008] An object of the present invention is to provide a new expression system for stable and large-scale production of a target protein in a lepidopteran insect such as a silkworm. [Means for solving the problem]

[0009] In order to solve the above problems, the present inventors came up with the idea of ​​constructing a new expression system in which a target gene encoding a target protein is fused to the C-terminus of an endogenous signal peptide in an exon sequence encoding a signal peptide in a sericin gene, a fibroin gene, or the like, thereby utilizing the promoter activity and enhancer activity of the endogenous gene to express the target gene.

[0010] In general, to efficiently knock-in an exogenous gene into the silkworm genome, it is necessary to cut the target gene locus using a genome editing enzyme or the like. Therefore, the present inventors set a genome cutting position within the exon sequence into which the target gene sequence is to be introduced, and attempted to knock-in the gene into the exon sequence. However, as a result of carrying out this method, they were faced with the result that 95% or more of the individuals in the generation injected could not produce normal cocoons, and 98% or more of the individuals could not develop into mating-capable adults. Therefore, it was found that it was extremely difficult to establish a lineage using this method.

[0011] Therefore, the present inventors attempted to knock-in a target gene sequence into an exon sequence by cutting the genome not in the exon sequence into which the target gene sequence is introduced, but in an intron sequence adjacent to the exon sequence. As a result, they found that almost all individuals of the injected generation developed into fertile adults and were able to produce a target protein in amounts far exceeding those of the GAL4 / UAS system of the prior art, and thus completed the present invention. The present invention is based on the above research results and provides the following.

[0012] (1) A genetically modified lepidopteran insect, In an exon sequence encoding a signal peptide or a functional fragment thereof of an endogenous gene, the exon sequence includes a gene sequence encoding a protein of interest or a fragment thereof, the genetically modified lepidopteran insect, wherein the target protein or a fragment thereof is fused to the C-terminus of the signal peptide or a functional fragment thereof; (2) The genetically modified lepidopteran insect according to (1), wherein the endogenous gene encodes fibroin, sericin, and / or fibrohexamarin. (3) The genetically modified lepidopteran insect according to (2), wherein the fibroin is a fibroin H chain and / or a fibroin L chain. (4) The endogenous gene is fibroin H chain and fibroin L chain, fibroin heavy chain and sericin 1, or Fibroin H chain, fibroin L chain, and sericin 1 A genetically modified lepidopteran insect according to (1), which encodes (5) The genetically modified lepidopteran insect according to any one of (1) to (4), wherein the exon sequence comprises a transcription termination sequence on the 3'-terminal side of the target gene sequence. (6) A double-stranded circular DNA for introducing a target gene sequence into a genome cleavage site in an intron sequence in an endogenous gene of a genetically modified lepidopteran insect, comprising: The endogenous gene is (a) a first spacer sequence adjacent to the 5'-terminus of the genome cleavage site; (b) a second spacer sequence adjacent to the 3' end of the genome cleavage site; (c) a first recognition sequence recognized by a first genome editing enzyme at the 5' end side of the first spacer sequence; and (d) a second recognition sequence recognized by a second genome editing enzyme on the 3'-end side of the second spacer sequence; Including, the double-stranded circular DNA comprises the first recognition sequence, the second spacer sequence, the first spacer sequence, a genome homologous sequence, and a target gene sequence in this order; the genome-homologous sequence comprises a base sequence homologous to a genome sequence from the second recognition sequence to an exon sequence or a partial sequence thereof located on the 3'-terminal side of the intron sequence, The double-stranded circular DNA, wherein the target gene sequence encodes a target protein or a fragment thereof that is fused to the C-terminus of a signal peptide or a functional fragment thereof of the endogenous gene. (7) A donor nucleic acid for producing a genetically modified lepidopteran insect by homologous recombination, comprising: The homologous recombination method includes cleaving a genomic cleavage site within an intron sequence in an endogenous gene with a genome editing enzyme, The donor nucleic acid is (a) a first genomic homologous sequence and a second genomic homologous sequence derived from the endogenous gene; and (b) a gene sequence of interest placed therebetween Including, The first genome homologous sequence is a base sequence homologous to a genome sequence from a base located on the 5'-end side of the genome cleavage position on the genome to an exon sequence or a partial sequence thereof located on the 3'-end side of the intron sequence, and has a mutation in a recognition sequence for the genome editing enzyme; the second genomic homologous sequence is a base sequence homologous to a genomic sequence located on the 3'-terminal side of the exon sequence or a partial sequence thereof on the genome, The donor nucleic acid, wherein the gene sequence of interest encodes a protein of interest or a fragment thereof that is fused to the C-terminus of a signal peptide or a functional fragment thereof of the endogenous gene. (8) A donor nucleic acid for producing a genetically modified lepidopteran insect by homologous recombination, comprising: The homologous recombination method includes cleaving a genomic cleavage site within an intron sequence in an endogenous gene with a genome editing enzyme, The donor nucleic acid is (a) a first genomic homologous sequence and a second genomic homologous sequence derived from the endogenous gene; and (b) a gene sequence of interest placed therebetween Including, the first genomic homologous sequence is a nucleotide sequence homologous to a genomic sequence from a nucleotide located on the 5'-terminal side of the intron sequence to an exon sequence or a partial sequence thereof located on the 5'-terminal side of the intron sequence, The second genome homologous sequence is a base sequence homologous to a genome sequence from a base located on the 3'-terminal side of the exon sequence or a partial sequence thereof and on the 5'-terminal side of the genome cleavage position to a base located on the 3'-terminal side of the genome cleavage position, and has a mutation in a recognition sequence for the genome editing enzyme; The donor nucleic acid, wherein the gene sequence of interest encodes a protein of interest or a fragment thereof that is fused to the C-terminus of a signal peptide or a functional fragment thereof of the endogenous gene. (9) The donor nucleic acid according to (7) or (8), comprising a nuclease recognition sequence at the end opposite the target gene sequence of the first genome homologous sequence and / or the second genome homologous sequence. (10) The donor nucleic acid described in (9), wherein the nuclease recognition sequence is a recognition sequence for the genome editing enzyme or a restriction enzyme recognition sequence. (11) A method for producing a genetically modified lepidopteran insect, comprising the steps of: (6) The double-stranded circular DNA according to The first genome editing enzyme or a nucleic acid encoding the first genome editing enzyme in an expressible state; and The second genome editing enzyme or a nucleic acid encoding the second genome editing enzyme in an expressible state. A process of introducing the vector into lepidopteran eggs by microinjection. The method comprising: (12) A method for producing a genetically modified lepidopteran insect, comprising the steps of: A donor nucleic acid according to any one of (7) to (10), and The genome editing enzyme, or a nucleic acid encoding the genome editing enzyme in an expressible state. A process of introducing the vector into lepidopteran eggs by microinjection. The method comprising: (13) A method for producing a target protein or a fragment thereof using a genetically modified lepidopteran insect according to any one of (1) to (5) or a genetically modified lepidopteran insect produced by the method according to (11) or (12). (14) The method according to (13), wherein the lepidopteran insect is a silkworm, and the target protein or a fragment thereof is produced in the silk gland of the silkworm. Effect of the Invention

[0013] According to the present invention, a target protein can be stably and mass-produced in a lepidopteran insect such as a silkworm. [Brief description of the drawings]

[0014] [Figure 1] Introduction of a target gene sequence having a stop codon into an endogenous gene. The target gene sequence is introduced into an exon sequence that encodes a signal peptide in the endogenous gene. The target protein encoded by the target gene sequence is fused to the C-terminal side of the signal peptide encoded by the endogenous gene. [Diagram 2] This is a diagram showing the genomic cleavage site in an endogenous gene in the production of a knock-in line. The genomic cleavage site is designed within an intron sequence located at the 5' end of an exon sequence into which a target gene sequence is introduced. As examples of endogenous genes into which a target gene sequence is introduced, FIG. 2A shows the sericin 1 gene, FIG. 2B shows the fibroin H gene, and FIG. 2C shows the fibroin L gene. [Diagram 3]This is a diagram showing an outline of gene knock-in using the TAL-PITCh method. Figure 3A shows the position in an endogenous gene from which each sequence used in constructing a donor nucleic acid used in the TAL-PITCh method originates. Figure 3B shows the structure of a double-stranded circular DNA used in the TAL-PITCh method. Figure 3C shows the structure of a knock-in gene obtained by knocking in a target gene sequence using the TAL-PITCh method. [Figure 4] FIG. 1 shows a method for gene knock-in using homologous recombination. [Diagram 5] 1 shows the results of observation of silk glands and cocoons in 5th instar larvae of the SP(FibH)-EGFP knock-in line. [Figure 6] The graph shows the average of the measured values ​​(n=1 to 4), and the error bars show the standard error. [Figure 7] 7 shows GM-CSF production in a silkworm line in which the GM-CSF gene sequence was knocked into the second exon of the endogenous fibroin H gene. Figure 7A shows the knock-in of the GM-CSF gene sequence into the fibroin H gene. Figure 7B shows the results of detecting GM-CSF by Western blotting. [Figure 8] 8 shows IgG production in silkworm strains in which gene sequences encoding IgG H chain and IgG L chain were knocked in to the second exon of the endogenous fibroin H gene and the third exon of the endogenous fibroin L gene, respectively. Figure 8A shows the knock-in of the IgG H chain gene sequence into the fibroin H gene. Figure 8B shows the knock-in of the IgG L chain gene sequence into the fibroin L gene. Figure 8C shows the amount of IgG produced. [Figure 9] The results of knock-in were obtained by designing the genome cleavage site within an intron sequence or an exon sequence. Figure 9A shows the results of homologous recombination after genome cleavage within the intron sequence of the Fib H gene. Figure 9B shows the results of homologous recombination after genome cleavage within the exon sequence of the Fib H gene. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] 1. Transgenic Lepidoptera 1-1. Overview A first aspect of the present invention is a genetically modified lepidopteran insect. The genetically modified lepidopteran insect of the present invention contains a target gene sequence in an exon sequence encoding a signal peptide or a functional fragment thereof of an endogenous gene, and expresses a target protein or a fragment thereof fused to the C-terminus of the signal peptide or the functional fragment thereof. The genetically modified lepidopteran insect of this aspect can stably and mass-produce a target protein.

[0016] 1-2.Definition The following terms frequently used in this specification are defined below. As used herein, the term "lepidoptera insects" refers to insects that belong to the taxonomic order Lepidoptera, and includes butterflies and moths. Butterflies include insects that belong to the families Nymphalidae, Papilionidae, Pieridae, Lycaenidae, and Hesperiidae. Moths include insects belonging to the families Saturniidae, Bombycidae, Brahmaeidae, Eupterotidae, Lasiocampidae, Psychidae, Geometridae, Archtiidae, Noctuidae, Pyralidae, and Sphingidae. For example, moths include species belonging to the genera Bombyx, Samia, Antheraea, Saturnia, Attacus, and Rhodinia, specifically, silkworms, Bombyx mandarina, Samia cynthia (including Samia cynthia ricini and hybrids of Samia cynthia and Samia ricini), Antheraea yamamai, Antheraea pernyi, Saturnia japonica, Actias gnoma, etc. Lepidoptera insects as hosts for the transformant of the present invention are not limited to these, but silkworms, which have high industrial applicability, are particularly preferred as hosts.

[0017] "Genetically modified lepidopteran insect" refers to a genetically modified lepidopteran insect carrying foreign DNA produced using recombinant gene technology, or its progeny. In this specification, the genetically modified lepidopteran insect particularly refers to a genetically modified insect obtained by introducing foreign DNA into lepidopteran insect eggs by microinjection.

[0018] As used herein, the term "silk gland" refers to a modified tubular organ of a salivary gland that has the function of producing, accumulating, and secreting liquid silk. Silk glands are usually present in pairs on the left and right sides of insects capable of spinning silk threads, mainly along the digestive tract of the larvae, and each silk gland is composed of three regions: the anterior, middle, and posterior silk gland. The posterior silk gland produces and secretes fibroin, a fiber component of silk thread. The middle silk gland also produces and secretes sericin, a coating component, and accumulates in its lumen together with fibroin transferred from the posterior silk gland.

[0019] As used herein, the term "endogenous gene" refers to a gene derived from a lepidopteran insect that is present a priori on the genome of that lepidopteran insect. In the present invention, an endogenous gene is, in principle, a gene that encodes a protein having a signal peptide. Therefore, in the present specification, an endogenous gene is, in principle, a gene that encodes a secretory protein or a membrane protein. The secretory protein may be, for example, any protein that constitutes silk thread. In the present specification, any protein that constitutes silk thread is often referred to as a "silk protein," and a gene that encodes a silk protein is referred to as a "silk gene." Specific examples of endogenous genes in lepidopteran insects include genes that encode fibroin, sericin, and fibrohexamarin.

[0020] In addition, in this specification, "exogenous gene" or "foreign gene" refers to a foreign gene that is acquired later through artificial manipulation or the like and is not present in the genome of a wild-type lepidopteran insect.

[0021] "Fibroin" is a protein that constitutes the fiber component of silk thread. Silkworm fibroin is mainly composed of three proteins, namely, fibroin H chain (Fib H), fibroin L chain (Fib L), and fibrohexamerin. As mentioned above, fibrohexamerin is also called p25 / FHX.

[0022] "Sericin" is a protein that covers the outer layer of the fibers formed by fibroin in silk threads. In silkworms, sericin is synthesized in the cells of the middle silk gland and secreted into the lumen of the middle silk gland after synthesis. The functions of sericin are known to be adhesion between fibroin fibers and protection of fibroin fibers from external stimuli. Silkworms can spin silk immediately after hatching, but the protein components of silk threads spun at each stage and silk threads from cocoons are different, and the sericin variant composition contained therein is also different. In general, about six types of sericin protein variants (sericin 1A', sericin 1C, sericin 1D, sericin 2, sericin 3, and sericin 4) are known to be biosynthesized from four types of sericin genes (Ser1, Ser2, Ser3, and Ser4) in silkworms. Of these, there are four main sericin variants contained in cocoons: sericin 1A', sericin 1C, sericin 1D, and sericin 3. In this specification, when the term "sericin" is used, it means a general term for sericin unless otherwise specified.

[0023] As used herein, a "signal peptide" or a "secretion signal" refers to an extracellular transport signal required for secreting a protein biosynthesized by gene expression outside the cell. After translation, the signal peptide is cleaved and removed by a signal peptidase before being secreted outside the cell. In this specification, the signal peptide is often written as "SP" and the endogenous gene name from which the signal peptide is derived is written in parentheses. A signal peptide is usually a relatively short peptide sequence of several tens of amino acids or less, and is characterized by a highly hydrophobic sequence. The sequence of a signal peptide can be predicted based on the amino acid sequence of a protein using a prediction tool such as signalP, but a structure prediction provided on a database can also be used. For example, in the case of silkworms, the sequence region of the signal peptide can be determined based on the sequence annotation provided on databases such as KAIKObase and KAIKOcDNA available in the Agrigenomics Information Database.

[0024] As used herein, a "functional fragment" of a signal peptide refers to a fragment consisting of a partial sequence of a signal peptide and retaining extracellular localization signal activity. A functional fragment of a signal peptide may retain, for example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 98% or more, or the same or more, of the extracellular localization signal activity of the full-length signal peptide. The amino acid length of a functional fragment is not particularly limited as long as it retains the activity of the full-length signal peptide, and may be, for example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the full length.

[0025] In the present specification, the term "full length" refers to the entire amino acid sequence corresponding to a protein that is synthesized and functions in a living body, or the entire base sequence in a gene that codes for it. In principle, in the case of a gene, the full length gene corresponds to the sequence from the initiation codon to the termination codon, and in the case of a protein, the full length protein corresponds to a polypeptide or peptide consisting of the amino acid sequence encoded by the full length gene. However, in the case of a secretory protein, the endogenous signal peptide contained on the N-terminus side is cleaved and removed during the secretion process and is not ultimately included. Therefore, in the case of a secretory protein, the "full length" does not need to include the signal peptide. In the present specification, the full length protein before the signal peptide is cleaved and removed during the secretion process is called a "precursor protein", and the full length protein after the signal peptide is cleaved and removed is called a "mature protein".

[0026] In this specification, "exon" means a region of the base sequence of a gene that remains in a mature transcript. In general, in eukaryotes, after a gene is transcribed as a primary transcript, an intervening region called an "intron" is removed by splicing, and exons are linked to each other to form a mature transcript. In this specification, "exon sequence" means a base sequence corresponding to an exon, and "intron sequence" means a base sequence corresponding to an intron. The exon sequence and intron sequence of any gene can be determined by comparing the genome sequence and cDNA sequence of the gene, but the exon / intron structure can also be predicted by obtaining sequence information published in databases such as the National Center for Biotechnology Information (NCBI), or by using genome analysis tools available in the technical field. For example, in the case of silkworm gene information, exon sequences and intron sequences can be searched for using databases such as KAIKObase and KAIKOcDNA available in the Agrigenomics Information Database.

[0027] As used herein, the term "target gene sequence" refers to a gene sequence that codes for a target protein or a fragment thereof. The target gene sequence may be a gene sequence derived from a genome or a gene sequence consisting of cDNA, and may or may not contain an intron within the target gene sequence. Furthermore, the target gene sequence may or may not contain a stop codon in addition to the gene sequence that codes for the target protein or a fragment thereof, and may or may not contain a transcription termination sequence downstream of the stop codon.

[0028] As used herein, the term "target protein" refers to a desired protein encoded by a target gene. The type of target protein is not important. It may be either a structural protein or a functional protein. Examples of structural proteins include fibrous proteins such as collagen, actin, myosin, and fibroin, keratin, and histone. Examples of functional proteins include peptide hormones (insulin, calcitonin, parathormone, growth hormone, etc.), cytokines (granulocyte-macrophage colony-stimulating factor (GM-CSF), epidermal growth factor (EGF), fibroblast growth factor (FGF), interleukin (IL), interferon (IFN), tumor necrosis factor α (TNF-α), transforming growth factor β (TGF-β), etc.), transcription factors (including GAL4), antibodies (immunoglobulins, etc.), serum albumin, hemoglobin, enzymes, fluorescent proteins, pigment synthesis proteins, and luminescent proteins. The immunoglobulin may be of any class (e.g., IgG, IgE, IgM, IgA, IgD, and IgY) or any subclass (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, IgA2). The fluorescent protein is not limited and may be, for example, CFP, AmCyan, RFP, DsRed, YFP, GFP (including derivatives such as EGFP and EYFP). The pigment synthesis protein may be, for example, a protein involved in the biosynthesis of melanin pigments (including dopamine melanin), ommochrome pigments, or pteridine pigments. The luminescent protein may be, for example, aequorin or luciferase. The target protein may be either a wild-type protein or a mutant protein.

[0029] As used herein, a "fragment" of a protein refers to a polypeptide or peptide that includes a partial region of a full-length protein. The fragment preferably retains activity. For example, the fragment may retain 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 98% or more, or the same or more, of the activity of the full-length protein. The amino acid length of the fragment is not particularly limited as long as it retains the activity of the full-length protein, and may be, for example, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the full length.

[0030] As used herein, the term "transcription termination sequence" refers to a sequence capable of terminating gene transcription, and is also called a terminator. The type of transcription termination sequence is not particularly limited. Preferably, the terminator is derived from the same species as the genetically modified Lepidoptera insect. For example, in the case of an insect such as a silkworm, an hsp70 terminator, an SV40 terminator, etc. can be used. In this specification, the term "plurality" refers to an integer of 2 or more, for example, an integer of 2 to 10, 2 to 7, 2 to 5, 2 to 4, or 2 to 3.

[0031] As used herein, "identity" of a base sequence refers to the percentage (%) of matching bases in the entire length of two base sequences when the two base sequences are aligned by inserting appropriate gaps into one or both of the sequences as necessary to maximize the number of matching bases.

[0032] As used herein, "homologous sequence" refers to a base sequence having about 60% or more identity to a reference sequence. The identity of a homologous sequence to a reference sequence may be, for example, 70% or more, 80% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 99.9% or more. Note that "genome homologous sequence" refers to a base sequence having any of the above-mentioned identities with the genome sequence of a lepidopteran insect as a reference sequence, and "genomic sequence" refers to a sequence having 100% identity to the corresponding base sequence in the genome.

[0033] As used herein, "amino acid identity" refers to the percentage (%) of matching amino acid residues out of the total number of amino acid residues in the amino acid sequences of two polypeptides being compared, when the sequences are aligned by inserting appropriate gaps into one or both sequences as necessary to maximize the number of matching amino acid residues.

[0034] As used herein, the term "amino acid substitution" refers to substitution within a conservative amino acid group that has similar properties such as charge, side chain, polarity, aromaticity, etc., among the 20 types of amino acids that constitute natural proteins. Examples include substitution within the uncharged polar amino acid group with a low polarity side chain (Gly, Asn, Gln, Ser, Thr, Cys, Tyr), the branched chain amino acid group (Leu, Val, Ile), the neutral amino acid group (Gly, Ile, Val, Leu, Ala, Met, Pro), the neutral amino acid group with a hydrophilic side chain (Asn, Gln, Thr, Ser, Tyr, Cys), the acidic amino acid group (Asp, Glu), the basic amino acid group (Arg, Lys, His), and the aromatic amino acid group (Phe, Tyr, Trp).

[0035] In this specification, the terms "5'-end" and "3'-end" refer to the 5'-end and 3'-end, respectively, of a transcription product transcribed from an endogenous gene, unless otherwise specified. In addition, in this specification, the terms "upstream" and "downstream" refer to the upstream and downstream directions of a gene, respectively, based on the transcription direction of an endogenous gene, unless otherwise specified.

[0036] 1-3.Configuration The genetically modified lepidopteran insect of the present invention comprises a gene sequence of interest encoding a protein of interest or a fragment thereof in an exon sequence encoding a signal peptide or a functional fragment thereof of an endogenous gene, the gene sequence of interest being contained in the exon sequence such that the protein of interest or a fragment thereof is fused to the C-terminus of the signal peptide or the functional fragment thereof.

[0037] As used herein, the term "exon sequence encoding a signal peptide or a functional fragment thereof of an endogenous gene" (hereinafter often referred to as "target exon sequence") is not limited to an exon sequence encoding a signal peptide in an endogenous gene. In an endogenous gene, a signal peptide is usually encoded by the first exon located most upstream in an mRNA transcribed from the endogenous gene, or by a plurality of exon sequences including the first exon, but the target exon sequence may be any exon sequence. For example, the target exon sequence may be the first exon, the second exon, the third exon, or the fourth exon. The target exon sequence may be, for example, an exon encoding the C-terminal amino acid residue of a signal peptide, or an exon adjacent to the 5'-terminal side thereof.

[0038] In the genetically modified lepidopteran insect of the present invention, the gene sequence of interest is inserted into the target exon sequence so that the protein of interest or a fragment thereof encoded by the gene sequence of interest is fused to the C-terminus of the signal peptide or a functional fragment thereof of the endogenous gene. More specifically, the gene sequence of interest is linked in frame to the 3'-terminus of the base sequence encoding the signal peptide or a functional fragment thereof in the target exon sequence of the endogenous gene. Thus, the N-terminus of the protein of interest or a fragment thereof is fused to the C-terminus of the signal peptide or a functional fragment thereof of the endogenous gene, and a fusion gene encoding a fusion polypeptide containing the signal peptide or a functional fragment thereof of the endogenous gene and the protein of interest or a fragment thereof is constructed in the locus of the endogenous gene. In this fusion polypeptide, the signal peptide or a functional fragment thereof and the protein of interest or a fragment thereof may be directly linked, or an amino acid sequence other than the signal peptide encoded by the target exon sequence (for example, an amino acid sequence located at the N-terminus of the mature protein described below) may be inserted between them.

[0039] In one embodiment, the endogenous gene encodes a protein that constitutes silk thread. The protein that constitutes silk thread is not particularly limited, and may be, for example, fibroin, sericin, and / or fibrohexamarin. Fibroin may be fibroin H chain and / or fibroin L chain. Sericin is not particularly limited, and may be, for example, sericin 1.

[0040] In the silkworm fibroin H chain, the precursor protein including the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 1, and the mature protein excluding the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 2. The signal peptide of the fibroin H chain consists of the amino acid sequence of positions 1 to 21 in SEQ ID NO: 1.

[0041] In the silkworm fibroin H-chain gene, the signal peptide is encoded by the first and second exons, and the C-terminal amino acid residue of the signal peptide is encoded by the second exon (FIG. 2B). In the genomic sequence of the fibroin H-chain gene shown in SEQ ID NO: 3, the first exon is at positions 1001 to 1042, the first intron is at positions 1043 to 2013, the second exon includes positions 2014 to at least 17763, and the region encoding the signal peptide in the second exon is at positions 2014 to 2034.

[0042] In the silkworm fibroin L chain, the precursor protein including the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 4, and the mature protein excluding the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 5. The signal peptide of the fibroin L chain consists of the amino acid sequence of positions 1 to 16 in SEQ ID NO: 4.

[0043] In the silkworm fibroin L chain gene, the signal peptide is encoded by the first, second, and third exons, and the C-terminal amino acid residue of the signal peptide is encoded by the third exon (FIG. 2C). In the genome sequence shown in SEQ ID NO: 6, the first exon of the fibroin L chain gene is at positions 574 to 889, the first intron is at positions 890 to 966, the second exon is at positions 967 to 1036, the second intron is at positions 1037 to 8976, the third exon is at positions 8977 to 9059, and the region encoding the signal peptide in the third exon is at positions 8977 to 8988.

[0044] In silkworm sericin 1, multiple isoforms are generated by selective splicing. In one example of an isoform, a precursor protein including a signal peptide has the amino acid sequence shown in SEQ ID NO: 7, and a mature protein excluding the signal peptide has the amino acid sequence shown in SEQ ID NO: 8. In the above isoforms of sericin 1, the signal peptide has the amino acid sequence of positions 1 to 19 in SEQ ID NO: 7.

[0045] In the silkworm sericin 1 gene, the signal peptide of the above isoform is encoded by the first and second exons, and the C-terminal amino acid residue of the signal peptide is encoded by the second exon (FIG. 2A). In the genome sequence shown in SEQ ID NO: 9, in the sericin 1 gene, the first exon of the above isoform is at positions 947 to 1039, the first intron is at positions 1040 to 3051, the second exon is at positions 3052 to 3082, and the region encoding the signal peptide in the second exon is at positions 3052 to 3069.

[0046] In silkworm fibrohexamarin, the precursor protein including the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 10, and the mature protein excluding the signal peptide consists of the amino acid sequence shown in SEQ ID NO: 11. The signal peptide of fibrohexamarin consists of the amino acid sequence of positions 1 to 17 in SEQ ID NO: 10.

[0047] In the silkworm fibrohexamarin gene, the signal peptide is encoded by exon 1, and the C-terminal amino acid residue of the signal peptide is encoded by exon 1. In the genomic sequence shown in SEQ ID NO: 12, the first exon of the fibrohexamarin gene is at positions 918 to 1052, the first intron is at positions 1053 to 1536, the second exon is at positions 1537 to 1756, and the region encoding the signal peptide in the first exon is at positions 1001 to 1051.

[0048] In the genetically modified lepidopteran insect of the present invention, the target gene sequence may be introduced into a single endogenous gene or into multiple endogenous genes. In addition, the genetically modified lepidopteran insect of the present invention may have an exon sequence containing the target gene sequence in a heterozygous or homozygous form. When the target gene sequence is introduced into multiple endogenous genes, the types of target gene sequences introduced into the multiple endogenous genes may be the same or different.

[0049] In one embodiment, the multiple endogenous genes into which the gene sequence of interest has been introduced may be genes encoding a fibroin H chain and a fibroin L chain, or genes encoding a fibroin H chain and sericin 1, or genes encoding a fibroin H chain, a fibroin L chain, and sericin 1.

[0050] In one embodiment, in the transgenic lepidopteran insect of the present invention, the gene of interest sequence has a stop codon. In a further embodiment, in the transgenic lepidopteran insect of the present invention, the targeted exon sequence comprises a transcription termination sequence 3' to the stop codon of the gene of interest sequence.

[0051] In one embodiment, the protein of interest is a fluorescent protein, an antibody, an antigen polypeptide, an enzyme, a cytokine, or an antimicrobial polypeptide. For example, when the protein of interest is an antibody, the heavy chain gene and the light chain gene constituting the antibody may be introduced into different endogenous genes.

[0052] 1-4.Effects The genetically modified lepidopteran insect of the present invention is capable of stably and abundantly producing a target protein encoded by a target gene introduced into a target exon sequence.

[0053] In the conventional GAL4 / UAS system, the GAL4 gene and UAS regulatory sequence are introduced into random locations on the genome, which can result in large variations in the expression level of the target protein. However, in the genetically modified lepidopteran insect of the present invention, the promoter activity and enhancer activity of the endogenous gene can be directly utilized to express the target gene, thereby enabling more reliable control of the expression level.

[0054] 2. Double-stranded circular DNA 2-1. Overview The second aspect of the present invention is a double-stranded circular DNA. The double-stranded circular DNA of this aspect can introduce a target gene sequence into a target exon sequence located on the 3'-end side of a genomic cleavage site in an intron sequence in an endogenous gene of a lepidopteran insect. The double-stranded circular DNA of this aspect can be used for knocking in a target gene based on, for example, the TAL-PITCh (precise integration into target chromosome) method.

[0055] 2-2.Definition In this embodiment, the term "double-stranded circular DNA" refers to a circular double-stranded DNA molecule that contains at least a gene sequence of interest for introducing the gene sequence into an endogenous gene of a lepidopteran insect. The double-stranded circular DNA is preferably a vector that can be maintained and / or replicated in bacterial cells such as E. coli, and may contain, for example, a sequence necessary for maintenance or replication in the cell (such as a replication origin and / or a gene encoding an antibiotic resistance protein). The double-stranded circular DNA may be, for example, a plasmid vector.

[0056] In this specification, "genome editing" refers to a gene targeting technique that uses a DNA repair mechanism associated with double strand break (DSB) caused by a DNA cleavage enzyme to insert a foreign gene (knock-in) or destroy a target gene (knock-out) at any position on the genome. Known genome editing techniques include the zinc finger nuclease (ZFN) method, the TALEN method, and the CRISPR / Cas method, and any of these methods may be used in this specification.

[0057] The "TALEN (Transcription Activator-Like Effector Nuclease) method" is a genome editing technology using an artificial DNA cleavage enzyme that combines a TAL effector (TALE) protein derived from the plant pathogenic bacterium Xanthomonas with a non-specific endonuclease domain. TALEN is a protein consisting of a TALE domain that contains repeated DNA binding units as a DNA binding domain and a non-specific endonuclease domain such as the nuclease domain of FokI. Since the nuclease domain, which has the enzymatic activity of cleaving DNA, functions as a dimer, TALEN functions as a dimer consisting of a polypeptide that recognizes the DNA sequence near the upstream (5') side of the double-strand break (DSB) site in the target base sequence (often referred to as "Left-TALEN" in this specification) and a polypeptide that recognizes the DNA sequence near the downstream (3') side of the DSB site (often referred to as "Right-TALEN" in this specification). The DNA-binding unit constituting the TALE domain has a mutation in the amino acid residues at positions 12 and 13 from the N-terminus, and each pair of two amino acids can specifically recognize each of the four bases constituting DNA (A: adenine, G: guanine, C: cytosine, T: thymine). For example, the amino acid residues at positions 12-13 recognize adenine when NI or NN, guanine when NN, cytosine when HD, and thymine when NG. The number of repeats of the DNA-binding unit can be varied depending on the base length of the target base sequence. By manipulating the TALE domain, gene targeting targeting any DNA sequence on the genome becomes possible. A gene knockout method using the TALEN method in lepidopteran insects such as silkworms is a known technique. For example, the method described in Takasu Y., et al., 2013, PLoS One 8, e73458 may be referred to.

[0058] The "Zinc Finger Nuclease (ZFN) method" is a genome editing technology that uses an artificial DNA cleaving enzyme consisting of a zinc finger domain as a DNA binding domain and a non-specific endonuclease domain such as the nuclease domain of FokI. Since one zinc finger motif can recognize three bases and bind to a target nucleic acid, by linking multiple zinc finger motifs, it specifically recognizes and binds to three times the number of bases linked. It functions as a dimer, and after binding to the target site, it uses its endonuclease activity to create a double-strand break (DSB) at a specific site in the target nucleic acid.

[0059] The "CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats / CRISPR associated proteins) method" is a genome editing technology that utilizes the adaptive immune system that evolved in bacteria and archaea to eliminate foreign DNA or RNA such as viruses and plasmids. In addition to the CRISPR / Cas9 method that uses the Cas9 protein, variations using other Cas proteins such as Cpf1 and Cas13a have been reported. Bacteria and archaea fragment the invaded foreign DNA or RNA, insert it into the CRISPR region in the genome, and use it as a template to synthesize CRISPR RNA (crRNA) of about 40 bp. The crRNA binds to a Cas protein with nuclease activity directly or via a trans-activating RNA (tracrRNA) to form a CRISPR / Cas complex. The CRISPR / Cas complex binds to and cleaves a target DNA or RNA sequence that has a complementary base sequence via the crRNA. When a double-stranded nuclease such as Cas9 or Cpf1 is used as the Cas protein, a DSB is induced at the target site.

[0060] As used herein, the term "genome editing enzyme" refers to a protein having the activity of specifically cleaving and editing a target site on a genome. Examples of genome editing proteins include TALEN (Transcription activator-like effector nuclease), Cas9 (CRISPR associated protein 9), and ZFN (zinc finger nuclease), which can be used for the above-mentioned genome editing. When the genome editing protein is TALEN, Left TALEN and Right TALEN can be used as TALENs that can function as dimers. When the genome editing protein is Cas9, a guide RNA such as the above-mentioned crRNA is required to perform genome editing.

[0061] 2-3.Configuration The double-stranded circular DNA of this embodiment includes a first recognition sequence, a second spacer sequence, a first spacer sequence, a genome homologous sequence, and a target gene sequence in this order. Here, the first recognition sequence, the second spacer sequence, the first spacer sequence, and the genome homologous sequence are derived from the base sequence of a genome region including an endogenous gene that is a target for introducing the target gene sequence. The double-stranded circular DNA of this embodiment includes a second recognition sequence on the 5'-end side of the genome homologous sequence.

[0062] In the introduction of a gene of interest using the double-stranded circular DNA of this embodiment, a double-stranded cleavage site is generated in an intron sequence on the genome using two genome editing enzymes (hereinafter referred to as the "first genome editing enzyme" and the "second genome editing enzyme") in a target endogenous gene. In this specification, this double-stranded cleavage site is referred to as the "genome cleavage position". The genome cleavage position can be set at any position within the intron sequence of the endogenous gene, but it is preferable to place it at a position other than the functional sequences such as the splice donor sequence and splice acceptor sequence required for splicing, and the branch site. Note that, depending on the type of genome editing enzyme, the genome cleavage position may not be accurately specified. Even in such a case, when designing the double-stranded circular DNA of the present invention, each element sequence constituting the double-stranded circular DNA is specified based on the position assumed as the genome cleavage position, and the position where the above-mentioned two genome editing enzymes actually cleave the genome and the double-stranded circular DNA is not limited to that position, but may be a position nearby (for example, any position in the first spacer sequence and / or the second spacer sequence).

[0063] The first recognition sequence, the second spacer sequence, and the first spacer sequence contained in the double-stranded circular DNA of this embodiment, as well as the second recognition sequence located at the 5'-end in the genome homologous sequence, are derived from a base sequence located near the genome cleavage site in the genome sequence of the endogenous gene. The "first recognition sequence" and the "second recognition sequence" are located at the 5'-end and 3'-end of the genome cleavage site in the genome sequence of the endogenous gene, respectively, and are identical to the base sequences recognized and bound by the first genome editing enzyme and the second genome editing enzyme, respectively. The base length of the first recognition sequence and the second recognition sequence varies depending on the type of genome editing enzyme, but is usually 8 to 30 bases long, for example, 10 to 25 bases long, 12 to 20 bases long, or 14 to 18 bases long. The "first spacer sequence" is derived from a sequence adjacent to the 5'-end of the genome cleavage site in the genome sequence of the endogenous gene, and is derived from a base sequence located between the first recognition sequence and the genome cleavage site. The "second spacer sequence" is derived from a sequence adjacent to the 3'-end of the genome cleavage site in the genome sequence of the endogenous gene, and is derived from a base sequence located between the genome cleavage site and the above-mentioned second recognition sequence. The base length of the first spacer sequence and the second spacer sequence varies depending on the type of genome editing enzyme, but is usually 6 to 30 bases long, for example, 8 to 25 bases long, 10 to 20 bases long, or 12 to 15 bases long. Here, the endogenous gene includes the first recognition sequence, the first spacer sequence, the second spacer sequence, and the second recognition sequence in order from the upstream side of the gene, whereas the double-stranded circular DNA of this embodiment is characterized in that the arrangement of the first spacer sequence and the second spacer sequence is reversed.

[0064] The genome-homologous sequence contained in the double-stranded circular DNA of this embodiment includes the above-mentioned second recognition sequence on its 5'-end side, and is composed of a base sequence homologous to the genome sequence from the second recognition sequence to the exon sequence or a partial sequence thereof located on the 3'-end side of the intron sequence including the genome cleavage position. Here, the exon sequence or a partial sequence thereof located on the 3'-end side of the intron sequence including the genome cleavage position may be an exon sequence adjacent to the 3'-end side of the intron sequence including the genome cleavage position. In the double-stranded circular DNA of this embodiment, the exon sequence or a partial sequence thereof included on the 3'-end side of the genome-homologous sequence encodes the C-terminal side of a signal peptide or a functional fragment thereof, and is linked in frame to the target gene sequence located on the 3'-end side thereof. In the double-stranded circular DNA of this embodiment, the genome-homologous sequence is not particularly limited as long as it includes the second recognition sequence of the genome editing enzyme. The base length of the genome homologous sequence may be, for example, 15 to 20,000 bases, 20 to 10,000 bases, 50 to 5,000 bases, 100 to 2,000 bases, or 500 to 1,000 bases.

[0065] In one embodiment, in the double-stranded circular DNA of this aspect, the region from the first recognition sequence to the second recognition sequence in the genome homologous sequence consists of the first recognition sequence, the second spacer sequence, the first spacer sequence, and the second recognition sequence in the genome homologous sequence.

[0066] In one embodiment, the gene of interest sequence in the double-stranded circular DNA of this embodiment contains a stop codon. In a further embodiment, the double-stranded circular DNA of this embodiment contains a transcription termination sequence on the 3'-terminal side of the gene of interest sequence.

[0067] In one embodiment, the double-stranded circular DNA of this aspect contains a marker gene for identifying an individual into which a gene sequence of interest has been introduced into an endogenous gene. For example, the marker gene can be located on the 3'-end side of the transcription termination sequence located on the 3'-end side of the gene sequence of interest.

[0068] The type of genome editing enzyme that recognizes the first recognition sequence and the second recognition sequence contained in the double-stranded circular DNA of this embodiment is not limited, and may be TALEN, ZFN, and / or Cas9. For example, the genome editing enzyme that recognizes the first recognition sequence and the second recognition sequence may be TALEN. In this case, the two genome editing enzymes that recognize the first recognition sequence and the second recognition sequence may be Left-TALEN and Right-TALEN that function as a dimer. In one embodiment, in the endogenous gene according to this aspect, multiple exons, including the first exon, encode a signal peptide.

[0069] 2-4.Effects By introducing the double-stranded circular DNA of this embodiment into the eggs of a lepidopteran insect together with the first and second genome editing enzymes, microhomology-mediated end-joining is possible between the first spacer sequence adjacent to the genome cleavage site on the genome and the first spacer sequence in the double-stranded circular DNA, and the target gene sequence can be inserted into the target exon sequence located on the 3' end side of the genome cleavage site within the intron sequence in the endogenous gene of the lepidopteran insect.

[0070] In one embodiment, the double-stranded circular DNA of this aspect can be used in the TAL-PITCh method. For details of the TAL-PITCh method, see the known technical literature (Nature communications, 2014, 5:5560).

[0071] 3. Donor Nucleic Acid 3-1. Overview The third aspect of the present invention is a donor nucleic acid. The donor nucleic acid of this aspect can introduce a target gene sequence into a target exon sequence located on the 3'-end or 5'-end side of a genome cleavage site in an intron sequence in an endogenous gene of a lepidopteran insect. The donor nucleic acid of this aspect can be used, for example, for knocking in a target gene based on homologous recombination.

[0072] 3-2.Configuration As used herein, the term "donor nucleic acid" refers to a nucleic acid for introducing a target gene sequence into an endogenous gene of a lepidopteran insect. The form of the donor nucleic acid is not limited, and it may be, for example, a double-stranded circular DNA such as a plasmid vector or a linear DNA.

[0073] The donor nucleic acid of this embodiment includes a first genome homologous sequence, a second genome homologous sequence, and a target gene sequence disposed therebetween. The first genome homologous sequence and the second genome homologous sequence are derived from the base sequence of a genomic region including an endogenous gene to be a target for introducing the target gene sequence.

[0074] In the introduction of a gene of interest using the donor nucleic acid of this embodiment, a genome editing enzyme that recognizes a sequence in the vicinity of an intron sequence on the genome is used to generate a double-stranded cleavage site in the intron sequence on the genome in the target endogenous gene. As in the second embodiment, in this embodiment, the double-stranded cleavage site is referred to as a "genomic cleavage site". As described above, depending on the type of genome editing enzyme, the genome cleavage site may not be accurately identified. Even in such a case, when designing the donor nucleic acid of the present invention, each element sequence constituting the donor nucleic acid is specified based on the position assumed as the genome cleavage site, and the position where the genome editing enzyme actually cleaves the genome is not limited to that position, but may be a position nearby. The genome cleavage site can be set at any position within the intron sequence of the endogenous gene, but it is preferable to place it at a position other than the splice donor sequence, splice acceptor sequence, branch site, and other sequences necessary for splicing. In this embodiment, the sequence recognized by this genome editing enzyme is referred to as a "genome editing enzyme recognition sequence" or simply a "recognition sequence".

[0075] The specific configurations of the first genomic homologous sequence and the second genomic homologous sequence contained in the donor nucleic acid of this embodiment differ between when the genomic cleavage site is located on the 5' end side of the target exon sequence into which the target gene sequence is introduced and when they are located on the 3' end side of the genomic cleavage site, and therefore will be described separately below.

[0076] (1) An embodiment in which the genomic cleavage site is located on the 5' end side of the target exon sequence In an embodiment in which the genome cleavage position is located on the 5'-end side of the target exon sequence, the first genome homologous sequence is composed of a base sequence homologous to the genome sequence from a base located on the 5'-end side of the genome cleavage position on the genome (e.g., a base located 10 bases, 20 bases, 50 bases, 100 bases, 500 bases, 1,000 bases or more upstream from the genome cleavage position) to a target exon sequence or a partial sequence thereof located on the 3'-end side of (or adjacent to) an intron sequence including the genome cleavage position. In this embodiment, the genome sequence corresponding to the first genome homologous sequence has a mutation in its recognition sequence so as not to be cleaved by a genome editing enzyme that cleaves the above-mentioned genome cleavage position. The type of the mutation is not particularly limited, and may be, for example, a substitution, deletion, and / or insertion of a base in the recognition sequence. The number of mutated bases in the recognition sequence is not particularly limited. For example, one or more bases may be substituted, deleted, and / or inserted, more specifically, one or more, two or more, three or more, four or more, five or more, or six or more bases may be substituted, deleted, and / or inserted. In this embodiment, the second genome-homologous sequence consists of a base sequence that is homologous to a genome sequence located on the 3'-terminal side of the target sequence or a partial sequence thereof.

[0077] In this embodiment, the base length of the first and second genome homologous sequences is not particularly limited and may be, for example, 100 to 20,000 bases, 200 to 10,000 bases, 500 to 5,000 bases, or 1,000 to 2,000 bases.

[0078] (2) An embodiment in which the genomic cleavage site is located on the 3'-end side of the target exon sequence In an embodiment in which the genomic cleavage site is located on the 3'-terminal side of the target exon sequence, the first genomic homologous sequence consists of a base sequence homologous to the genomic sequence from a base located on the 5'-terminal side of the intron sequence containing the genomic cleavage site (specifically, a base located on the 5'-terminal side of the position in the target exon sequence where the target gene sequence is inserted, for example, a base located 100 bases, 200 bases, 500 bases, 1,000 bases, 5,000 bases, or 10,000 bases or more upstream from the position where the target gene sequence is inserted) to a target exon sequence or a partial sequence thereof located on the 5'-terminal side of (or adjacent to) the intron sequence (for example, to a base encoding the C-terminal amino acid residue of a signal peptide or a functional fragment thereof in the target exon sequence, or to a base encoding an amino acid residue further C-terminal than the C-terminal amino acid residue of a signal peptide or a functional fragment thereof in the target exon sequence).

[0079] In this embodiment, the second genome homologous sequence is a nucleotide sequence homologous to the genome sequence from a nucleotide located on the 3'-terminal side of the target exon sequence or a partial sequence thereof and on the 5'-terminal side of the genome cleavage position (e.g., a nucleotide encoding an amino acid residue further C-terminal than the C-terminal amino acid residue of a signal peptide or a functional fragment thereof in the target exon sequence) to a nucleotide located on the 3'-terminal side of the genome cleavage position (e.g., a nucleotide located 10 nucleotides, 20 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, or 10,000 nucleotides or more downstream from the genome cleavage position). In this embodiment, the genome sequence corresponding to the second genome homologous sequence has a mutation in its recognition sequence so as not to be cleaved by the genome editing enzyme that cleaves the above-mentioned genome cleavage position. The type of the mutation and the number of mutated bases are not particularly limited, and may be, for example, a substitution, deletion, and / or insertion of one or more bases in the recognition sequence, as in (1) above.

[0080] In this embodiment, the base length of the first and second genome homologous sequences is not particularly limited and may be, for example, 100 to 20,000 bases, 200 to 10,000 bases, 500 to 5,000 bases, or 1,000 to 2,000 bases.

[0081] In the donor nucleic acid of this embodiment, the target exon sequence or a partial sequence thereof contained in either the first or second genomic homologous sequence encodes a signal peptide or a functional fragment thereof or a sequence on the C-terminal side thereof, and is linked in frame to a target gene sequence located on the 3'-terminal side thereof via a base sequence encoding, in some cases, multiple amino acid residues.

[0082] In one embodiment, the gene of interest sequence in the donor nucleic acid of this aspect comprises a stop codon. In a further embodiment, the donor nucleic acid of this aspect comprises a transcription termination sequence on the 3' end of the gene of interest sequence. In one embodiment, the donor nucleic acid of this aspect contains a marker gene for identifying a transgenic Lepidoptera insect in which a gene sequence of interest has been introduced into an endogenous gene. The type of genome editing enzyme used in the homologous recombination method using the donor nucleic acid of this embodiment is not limited, and may be TALEN, ZFN, and / or Cas9.

[0083] In one embodiment, the donor nucleic acid of this aspect comprises a nuclease recognition sequence at the end opposite to the target gene sequence of the first genome homologous sequence and / or the second genome homologous sequence. The nuclease recognition sequence is not particularly limited, and may be a recognition sequence of a genome editing enzyme that cleaves the above-mentioned genome cleavage position, or may be a restriction enzyme recognition sequence that can be cleaved by any restriction enzyme different from the genome editing enzyme. When the nuclease recognition sequence is a TALEN recognition sequence, the nuclease recognition sequence may be a combination of two recognition sequences recognized by Left TALEN and Right TALEN.

[0084] Effects By introducing the donor nucleic acid of this embodiment together with a genome editing enzyme into an egg of a lepidopteran insect, homologous recombination can be induced in an endogenous gene of the lepidopteran insect, and a target gene sequence can be inserted into the target exon sequence.

[0085] 4. Methods for producing transgenic lepidopteran insects 4-1. Overview The fourth aspect of the present invention is a method for producing a genetically modified lepidopteran insect, which comprises introducing the double-stranded circular DNA according to the second aspect or the donor nucleic acid according to the third aspect into an egg of a lepidopteran insect by microinjection to introduce a target gene sequence into an exon sequence of an endogenous gene, thereby producing a genetically modified lepidopteran insect.

[0086] 4-2. Method The production method of this embodiment includes an essential step of introducing a double-stranded circular DNA or a donor nucleic acid into an egg of a lepidopteran insect by microinjection, and includes an egg obtaining step and a genetically modified lepidopteran insect selecting step as selection steps. The configuration of each step will be described below.

[0087] (1) Egg acquisition process The "egg obtaining step" is a step of obtaining eggs by allowing adult female parent silkworms to lay eggs. The method of obtaining eggs may be a method commonly used in the art. Egg laying is initiated by providing an egg-laying mat to the mated female parent silkworm. The temperature during egg laying is 23 to 28°C, preferably around 25°C. Usually, female silkworms start laying eggs several hours after mating. In order to incorporate the DNA introduced into the eggs into the nucleus, microinjection must be performed within 2 to 8 hours, preferably 3 to 6 hours, after laying.

[0088] (2) Introduction process The "introduction step" refers to a step of introducing a double-stranded circular DNA or a donor nucleic acid into an egg of a lepidopteran insect by microinjection.

[0089] In an embodiment in which the introduction step of the production method of this aspect involves introducing a double-stranded circular DNA into an egg, the configuration of the double-stranded circular DNA is as described in aspect 2. In the introduction step of this embodiment, the double-stranded circular DNA described in aspect 2, the first genome editing enzyme described in aspect 2 or a nucleic acid encoding the first genome editing enzyme in an expressible state, and the second genome editing enzyme described in aspect 2 or a nucleic acid encoding the second genome editing enzyme in an expressible state are introduced into an egg of a lepidopteran insect by microinjection.

[0090] In an embodiment in which the introduction step of the production method of this aspect involves introducing a donor nucleic acid into an egg, the configuration of the donor nucleic acid is as described in aspect 3. In the introduction step of this embodiment, the donor nucleic acid described in aspect 3, the genome editing enzyme described in aspect 3, or a nucleic acid encoding the genome editing enzyme in an expressible state is introduced into an egg of a lepidopteran insect by microinjection.

[0091] As used herein, the term "expressible state" refers to a state in which a gene to be expressed is located downstream of a promoter under the control of the promoter. The nucleic acid encoding the genome editing enzyme in an expressible state may be RNA such as mRNA, or DNA such as plasmid DNA or linear DNA. The DNA encoding the genome editing enzyme in an expressible state contains a promoter that can be expressed in the eggs of a lepidopteran insect in addition to a base sequence encoding the first or second genome editing enzyme, and may contain components such as a marker gene (selection marker), an enhancer, a terminator, a replication origin, and a polyA signal, as necessary.

[0092] The microinjection method may be performed by a method known in the art. For example, the double-stranded circular DNA or donor nucleic acid, and the genome editing enzyme or the nucleic acid encoding the genome editing enzyme in an expressible state are dissolved or diluted with a solvent such as water or a buffer to an appropriate concentration to prepare an injection solution. Then, the fertilized egg is microinjected 3 to 6 hours after laying. The amount of nucleic acid to be introduced is not particularly limited. It may be appropriately determined depending on the type, properties, and purpose of the nucleic acid. Usually, 50 nL to 30 nL is sufficient. After the introduction, the lepidopteran insect egg may be incubated under appropriate conditions, for example, at 25°C, until hatching.

[0093] (3) Genetically modified lepidopteran insect selection process The "selection step of genetically modified lepidopteran insects" refers to a step of selecting genetically modified lepidopteran insects from hatched lepidopteran insects. This step may also be carried out by a method known in the art. For example, when the double-stranded circular DNA or donor nucleic acid used in the introduction step contains a marker gene, the desired genetically modified lepidopteran insect can be easily selected based on the expression of the marker gene. As used herein, the term "marker gene" refers to a polynucleotide consisting of a base sequence that codes for a marker protein, also called a selection marker.

[0094] As used herein, the term "marker protein" refers to a protein that can confer a new trait not present in a host lepidopteran insect upon expression of a marker gene, and includes enzymes, fluorescent proteins, pigment synthesis proteins, luminescent proteins, etc. Based on the activity of the marker protein, it becomes possible to easily distinguish a transformant that has an introduced nucleic acid.

[0095] 5. Methods for Producing a Protein or Fusion Protein of Interest 5-1. Overview The fifth aspect of the present invention is a method for producing a target protein, a fragment thereof, or a fusion protein containing the same. According to the production method of this aspect, the target protein, a fragment thereof, or a fusion protein containing the same can be mass-produced using the genetically modified lepidopteran insect of the first aspect or the genetically modified lepidopteran insect produced by the production method according to the fourth aspect.

[0096] 5-2.Production method The production method of the present invention includes a rearing step and a recovery step. Each step will be described below.

[0097] (1) Rearing process The "rearing step" refers to a step of rearing the genetically modified lepidopteran insect of the first embodiment or the genetically modified lepidopteran insect produced by the production method described in the fourth embodiment. The genetically modified lepidopteran insect may be reared by a technique known in the art for each lepidopteran insect. For example, if the lepidopteran insect is a silkworm, "General Theory of Silkworms; Takami Tsuyoshi, published by the National Silkworm Association" may be referred to. The feed may be, for example, natural leaves of food tree species such as leaves of the genus Morus for silkworms and mulberry silkworms, leaves of Ricinus communis or Ailanthus altissima for Eri silkworms, and leaves of Fagaceae for Scutellaria japonica, or artificial feed such as Silkmate L4M or for 1-3 instars of original silkworm species (Nihon Nosan Kogyo). Considering that it is possible to suppress the occurrence of diseases, to provide stable quality and amount of food, and to rear the silkworms in a sterile environment as necessary, artificial feed is preferable. Below, a simple rearing method will be explained using silkworms as an example.

[0098] Sweeping is performed using eggs laid by an appropriate number of females of the same genetically modified lepidopteran insect (e.g., 4 to 10). The hatched larvae are transferred from the egg carrier to a container lined with dry-proof paper (paraffin-treated paper) that serves as a silkworm bed, and are fed with artificial feed such as Silkmate arranged on the dry-proof paper. As a rule, the feed is replaced once for the 1st and 2nd instars, and once to three times for the 3rd instar. If there is a lot of leftover old feed, it is removed to prevent decay. For rearing 4th to 5th instar adult silkworm larvae, they are transferred to a large container and the number of larvae per container is adjusted appropriately. Depending on the humidity and conditions inside the container, the container may be covered with dry-proof paper, acrylic, or mesh lids. The rearing temperature is 25 to 28°C throughout all instars.

[0099] (2) Recovery process The "recovery process" refers to a process of recovering the target protein or a fragment thereof, or a fusion protein containing the same, which has been expressed and secreted in the silk gland cells of the larvae of a genetically modified lepidopteran insect and then accumulated in the silk gland lumen.

[0100] The genetically modified lepidopteran insect used in this embodiment expresses in silk gland cells a precursor protein in which a signal peptide or a functional fragment thereof is fused to the N-terminus of a fusion protein (hereinafter referred to as "target protein, etc.") containing a mature protein or a C-terminal fragment thereof encoded by an endogenous gene at the C-terminus of the target protein or fragment thereof. The precursor protein expressed in silk gland cells is transported to the endoplasmic reticulum by the action of the signal peptide or a functional fragment thereof, and the signal peptide or a functional fragment thereof is cleaved by the action of an enzyme such as a peptidase in the endoplasmic reticulum, and then secreted into the lumen of the silk gland. The target protein, etc. after the signal peptide or the functional fragment thereof is cleaved is secreted from the anterior silk gland to the outside of the individual during the pupation stage and spun. Therefore, methods for recovering the target protein, etc. include a method of recovering it from a cocoon, or a method of directly recovering the target protein, etc. by extracting the silk gland from the insect body at the late final instar to the pre-pupal stage. In particular, the method of recovering it from a cocoon is advantageous in that the target protein, etc. can be recovered easily.

[0101] The method for recovering the target protein from the cocoon is as follows: first, the final stage larvae are transferred to the cocoon and allowed to spin a cocoon. Next, the target protein is extracted from the cocoon. There is no particular limitation on the extraction method. For example, the target protein can be recovered by simply immersing the cocoon in water or an appropriate neutral extraction buffer that does not contain a protein denaturant (e.g., phosphate-buffered saline, pH 7.2, containing or not containing 1% Tween-20 and 0.05% sodium azide). To enhance the extraction effect, the cocoon may be cut or crushed before immersion. The extraction temperature is low, 0 to 10°C, preferably 0 to 5°C, to prevent thermal denaturation of the target protein. However, if the target protein is not a heat-sensitive peptide, the extraction may be performed at 10 to 40°C. The extraction liquid may be stirred as necessary. The extraction time may be appropriately determined depending on the extraction conditions, such as the state of the cocoons (e.g., uncut or powdered), the amount of the extract, the extraction temperature, the presence or absence of stirring, etc. Insoluble components such as fibroin can be removed from the extract by centrifugation or filtration as necessary.

[0102] The method of extracting the silk gland from the insect body in the late final instar to prognathic stage and recovering the target protein, etc., can be achieved by a method known in the art. For example, silkworms immediately before spinning on the 6th day of the final instar (5th instar) may be anesthetized on ice, the dorsal side may be incised, and the silk gland may be extracted with tweezers without damaging it (see Mori Yasushi, ed., New Biological Experiments with Silkworms, Sanseido, 1970, pp. 249-255). The extracted silk gland may be gently shaken in the extraction buffer at a temperature of 0 to 10°C, preferably 0 to 5°C, to elute the target protein, etc., in the buffer. If the target protein, etc. is not a heat-sensitive peptide, the extraction may be performed at a temperature of 10 to 40°C. Then, impurities such as tissue fragments may be removed by centrifugation or filtration, and the supernatant containing the target protein, etc. may be recovered.

[0103] Effects According to the production method of the present invention, by using the genetically modified lepidopteran insect of the first aspect or the larvae of a genetically modified lepidopteran insect produced by the production method described in the fourth aspect as a protein production system, it is possible to produce a large amount of a target protein, etc., and to easily recover it, compared to the case where the GAL4 / UAS system is used. EXAMPLES

[0104] <Example 1: Creation of knock-in lines for useful protein production> (the purpose) The GAL4 / UAS system of the prior art is a gene control system that utilizes a combination of the yeast-derived transcription factor GAL4 and the regulatory sequence UAS. In the GAL4 / UAS system used as a protein production system in silkworms, an expression system that expresses a target protein in the silk gland or the like is constructed by crossing a GAL4 line that expresses the GAL4 gene under the control of a promoter such as a silk gene with a UAS line that expresses a target gene under the control of the regulatory sequence UAS.

[0105] In the GAL4 / UAS system, it is necessary to establish the GAL4 line and the UAS line separately and then cross them, so it takes time to construct the expression system. In addition, the GAL4 gene and the UAS regulatory sequence are introduced at random positions on the genome, so the expression level of the target protein varies and is difficult to predict. Therefore, after creating a large number of lines, it is necessary to examine the expression level in the individuals obtained by crossing the GAL4 line and the UAS line, and select good parent lines based on the results.

[0106] Therefore, the present inventors came up with the idea of ​​constructing a new expression system for stable mass production of a target protein by knocking in a target gene sequence into an endogenous gene. More specifically, if a target gene encoding a target protein is knocked in so as to be fused with a signal peptide into an exon sequence encoding a signal peptide in an endogenous sericin gene or fibroin gene, it may be possible to highly express the target gene by utilizing the promoter activity and enhancer activity of the endogenous gene as is.

[0107] However, efficient knock-in to endogenous genes requires cutting the target gene using genome editing enzymes.

[0108] In the comparative example 1 described below, the inventors attempted knock-in by designing a genome cleavage site in an exon sequence, but found that more than 95% of the individuals of the injected generation were unable to produce normal cocoons, and more than 98% were unable to develop into mating-capable adults. Therefore, it is extremely difficult to establish a lineage using this method.

[0109] In this Example, therefore, it will be examined whether the above problems can be overcome by designing a genomic cleavage site within an intron sequence and performing knock-in to an exon sequence.

[0110] (Methods and Results) In this example, a gene of interest encoding an EGFP protein is knocked in as a target protein fused to the C-terminus of a signal peptide in an exon sequence encoding a signal peptide in an endogenous silk gene. The TAL-PITCh (precise integration into target chromosome) method and homologous recombination method are used to knock in the gene of interest, and the intron sequence adjacent to the 5' end of the target exon sequence is cleaved with the genome editing enzyme TALEN (transcription activator-like effector nuclease) (Figure 2). The EGFP gene sequence, which is the gene of interest, has a stop codon and a transcription termination sequence at the 3' end (Figure 1).

[0111] Knock-in to the fibroin H (FibH) gene, fibroin L (FibL) gene, and sericin 1 (Ser1) gene was performed by the following method. In the following examples, the vector expressing TALEN was constructed with reference to Y. Takasu, S. Sajwan, T. Daimon, M. Osanai-Futahashi, K. Uchino, H. Sezutsu, T. Tamura, M. Zurovec (2013): Efficient TALEN construction for Bombyx mori gene targeting, PLoS One, 8, e73458. In addition, TALEN mRNA was synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen) using each TALEN expression vector as a template.

[0112] (1) Knock-in of fibroin H gene using TAL-PITCh method In the fibroin H gene (hereinafter referred to as "FibH gene"), the first and second exons encode a signal peptide. The EGFP gene sequence was introduced into the second exon using the TAL-PITCh method so that the EGFP protein was fused to the C-terminus of the FibH signal peptide.

[0113] Specifically, in the genomic sequence of the FibH gene (SEQ ID NO: 3), the position in the first intron between the first and second exons (the position between positions 1945 and 1964 in SEQ ID NO: 3, FIG. 2A) was set as the genomic cleavage position, and a double-stranded circular DNA shown in FIG. 3B was constructed as a donor nucleic acid for use in the TAL-PITCh method (hereinafter, referred to as "SP(FibH)-EGFP donor nucleic acid for TAL-PITCh"). The SP(FibH)-EGFP donor nucleic acid for TAL-PITCh contains the first recognition sequence, the second spacer sequence, the first spacer sequence, the genomic homologous sequence, the target gene sequence, the transcription termination sequence, and the marker gene in this order (FIG. 3B).

[0114] Here, the first spacer sequence is adjacent to the 5'-end side of the genome cleavage position, and is a 10-base-long sequence consisting of positions 1945 to 1954 in SEQ ID NO: 3. The second spacer sequence is adjacent to the 3'-end side of the genome cleavage position, and is a 10-base-long sequence consisting of positions 1955 to 1964 in SEQ ID NO: 3. The first recognition sequence is a 20-base-long sequence consisting of positions 1925 to 1944 in SEQ ID NO: 3 that is recognized by Left TALEN at the 5'-end side of the first spacer sequence. The genome homologous sequence is a base sequence homologous to the genome sequence from the second recognition sequence to the codon encoding the C-terminal residue of the signal peptide in the second exon sequence, and is a 70-base-long sequence consisting of positions 1965 to 2034 in SEQ ID NO: 3. Here, the second recognition sequence is a 20-base-long sequence consisting of positions 1965 to 1984 in SEQ ID NO: 3 that is recognized by Right TALEN at the 3'-end side of the second spacer sequence. The target gene sequence was the EGFP gene sequence, and the transcription termination sequence used was the transcription termination sequence of the silkworm-derived sericin 1 gene. The base sequence from the first recognition sequence to the transcription termination sequence of the EGFP gene in the SP(FibH)-EGFP donor nucleic acid for TAL-PITCh is shown in SEQ ID NO: 13. In addition, the amino acid sequence of a protein in which a FibH-derived signal peptide is fused to the N-terminus of the C-terminal fragment excluding the initiation methionine in EGFP, which is encoded by the gene after knock-in (FIG. 3C), is shown in SEQ ID NO: 14 (hereinafter referred to as "SP(FibH)-EGFP fusion protein").

[0115] SP(FibH)-EGFP donor nucleic acid for TAL-PITCh was injected into silkworm eggs of the white-eyed, white-egg, non-diapause strain w1-pnd maintained at the National Agriculture and Food Research Organization 2 to 8 hours after oviposition together with mRNAs encoding Left and Right TALENs synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen). After injection, the eggs were incubated at 25°C in a humidified environment until they hatched. The silkworms of the current generation after injection were crossed with the parent strain, and the resulting next-generation larvae were selected by the fluorescence of DsRed2 expressed in the whole body or EGFP expressed in the silk gland to obtain knock-in silkworm strains. Hereinafter, the knock-in strain obtained is referred to as the "SP(FibH)-EGFP knock-in strain." In addition, the Fib H gene after this knock-in is referred to as the "SP(FibH)-EGFP knock-in gene."

[0116] Of the larvae that developed from the 105 microinjected embryos, 103 spun cocoons normally, meaning that no cocoon-spinning defects were observed in the injected larvae.

[0117] In addition, among the adults of the injected generation, which had successfully spun cocoons, 34 out of 36 female silkworms mated normally with wild-type male silkworms and laid eggs, and 50 out of 53 male silkworms mated normally with wild-type female silkworms and laid eggs. Therefore, no mating problems were observed in the adults of the injected generation. These results demonstrate that knock-in lines capable of normal spinning and mating can be produced extremely efficiently by disrupting the genome within an intron sequence.

[0118] In the SP(FibH)-EGFP knock-in line, strong EGFP fluorescence was observed in the middle and posterior silk glands from the first instar larvae. Figure 5 shows the results of observation of the silk glands and cocoons of fifth instar larvae. Specifically, the larvae were anesthetized on ice just before spinning on the sixth day of the fifth instar, and the dorsal side was incised and the middle and posterior silk glands were removed with tweezers without damaging them, and observed under a fluorescence microscope without fixing. As a result, extremely strong EGFP fluorescence was observed (left side of Figure 5). Furthermore, the cocoons of this line were clearly yellow-green under normal white light (right side of Figure 5).

[0119] (2) Knock-in of the fibroin H gene by homologous recombination The EGFP gene sequence was introduced into the second exon of the FibH gene by homologous recombination so that the EGFP protein was fused to the C-terminus of the signal peptide of FibH.

[0120] Specifically, as in (1) above, a position in the first intron (position between positions 1945 and 1964 in SEQ ID NO: 3, FIG. 2A) was used as the genome cleavage position, and a double-stranded circular DNA shown in FIG. 4 was constructed as a donor nucleic acid to be used in homologous recombination (hereinafter referred to as "SP(FibH)-EGFP donor nucleic acid for homologous recombination"). The SP(FibH)-EGFP donor nucleic acid for homologous recombination contains a first genome homologous sequence and a second genome homologous sequence, as well as an EGFP gene sequence, a transcription termination sequence, and a marker gene, which are the target gene sequence arranged therebetween, and contains a TALEN recognition sequence at the end opposite to the EGFP gene sequence of the first genome homologous sequence and the second genome homologous sequence.

[0121] Here, the first genome homologous sequence is a 1034-base-long sequence homologous to the genome sequence from the base located on the 5'-end side of the genome cleavage position on the genome (position 1001 in SEQ ID NO: 3) to the codon encoding the C-terminal residue of the signal peptide in the second exon sequence. The first genome homologous sequence has a mutation in the vicinity of the above-mentioned genome cleavage position so as not to be recognized by the Left TALEN and Right TALEN described in (1) above (specifically, the base sequence AACTTCGATTGAATGTGCGAAATTTATAGCTCAATATTTTAGCACTTATCGTATTGATTT (SEQ ID NO: 33) located at positions 1925 to 1984 in SEQ ID NO: 3 is replaced with the base sequence AtCTaCGATTGAAaGaGCGtAATTTATAGCTCAATATTTTAtGCtCaTAaCGTATTGATTT (SEQ ID NO: 34)). The second genome homologous sequence is a 377-base sequence homologous to the genome sequence located on the 3'-terminal side of the codon encoding the C-terminal residue of the signal peptide in the second exon sequence on the genome. In the SP(FibH)-EGFP donor nucleic acid for homologous recombination, the TALEN recognition sequence and the base sequence from the first genome homologous sequence to the second genome homologous sequence and the TALEN recognition sequence are shown in SEQ ID NO: 15. In addition, the protein in which the FibH-derived signal peptide is fused to the N-terminal side of EGFP, which is encoded by the gene after knock-in (FIG. 4), consists of the amino acid sequence shown in SEQ ID NO: 14 as in (1) above (hereinafter, referred to as "SP(FibH)-EGFP fusion protein" as in (1) above).

[0122] The SP(FibH)-EGFP donor nucleic acid for homologous recombination was injected into silkworm eggs of the w1-pnd strain 2 to 8 hours after spawning together with mRNAs encoding the Left TALEN and Right TALEN synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen). After injection, the eggs were incubated in a humidified environment at 25°C until hatching. The silkworms of the current generation after injection were crossed with the parent strain, and the resulting next-generation larvae were selected based on the fluorescence of DsRed2 expressed throughout the body or EGFP expressed in the silk gland to obtain knock-in silkworm strains. Hereinafter, the knock-in strain obtained is referred to as the "SP(FibH)-EGFP knock-in strain" as in (1) above.

[0123] As in (1) above, no defects in spinning or mating were observed in the individuals developed from the microinjected embryos. This indicates that knock-in strains capable of normal spinning and mating can also be efficiently produced by cutting the genome within the intron sequence using homologous recombination.

[0124] (3) Knock-in of the fibroin L (FibL) gene by homologous recombination The EGFP gene sequence was introduced into the third exon of the fibroin L (hereinafter referred to as "FibL") gene by homologous recombination so that the EGFP protein was fused to the C-terminus of the signal peptide of FibL.

[0125] Specifically, a position in the second intron of the FibL gene (position between positions 8937 and 8954 in SEQ ID NO: 6) was used as the genome cleavage position, and a double-stranded circular DNA was constructed as a donor nucleic acid to be used in homologous recombination (hereinafter referred to as "SP(FibL)-EGFP donor nucleic acid for homologous recombination"). The SP(FibL)-EGFP donor nucleic acid for homologous recombination includes a first genome homologous sequence and a second genome homologous sequence, as well as an EGFP gene sequence, a transcription termination sequence, and a marker gene, which are the target gene sequence arranged therebetween, and includes a TALEN recognition sequence at the end opposite to the EGFP gene sequence of the first genome homologous sequence and the second genome homologous sequence.

[0126] Here, the above-mentioned genomic cleavage site is cleaved by a Left TALEN that recognizes a 20-base-long sequence consisting of positions 8917 to 8936 in SEQ ID NO: 6, and a Right TALEN that recognizes a 20-base-long sequence consisting of positions 8955 to 8974 in SEQ ID NO: 6.

[0127] The first genome homologous sequence is a 1128-base-long sequence homologous to the genome sequence from the base located at the 5'-end side of the genome cleavage position on the genome (position 7864 in SEQ ID NO: 6) to the codon encoding the C-terminal residue of the signal peptide in the third exon sequence. The first genome homologous sequence has a mutation in the vicinity of the above-mentioned genome cleavage position so that it is not recognized by the Left TALEN and Right TALEN (specifically, the base sequence CCCGAGAAAACAATTTGTTGTGTATAATTTAAACCAAAACCCGAATTTAATTTTTCGC (SEQ ID NO: 35) located at positions 8917 to 8974 in SEQ ID NO: 6 is replaced with the base sequence CCCGAGAAAAgAATTcGTTcTGTATAATTTAAACCAAAAttCGAATTTAATTTTTCGC (SEQ ID NO: 36)). The second genome homologous sequence is a 1362-base sequence homologous to the genome sequence located on the 3'-terminal side of the codon encoding the C-terminal residue of the signal peptide in the third exon sequence on the genome. In the SP(FibL)-EGFP donor nucleic acid for homologous recombination, the TALEN recognition sequence and the base sequence from the first genome homologous sequence to the second genome homologous sequence and the TALEN recognition sequence are shown in SEQ ID NO: 16. In addition, the protein in which the FibL signal peptide is fused to the N-terminal side of the C-terminal fragment of EGFP excluding the initiating methionine, which is encoded by the gene after knock-in, consists of the amino acid sequence shown in SEQ ID NO: 17 (hereinafter referred to as "SP(FibL)-EGFP fusion protein").

[0128] The SP(FibL)-EGFP donor nucleic acid for homologous recombination was injected into silkworm eggs of the w1-pnd strain 2 to 8 hours after spawning together with mRNAs encoding Left TALEN and Right TALEN synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen). The eggs after injection were incubated in a humidified state at 25°C until hatching. A knock-in silkworm line was obtained in the same manner as in (2) above. Hereinafter, the knock-in line obtained is referred to as the "SP(FibL)-EGFP knock-in line" as in (1) above. The FibL gene after this knock-in is referred to as the "SP(FibL)-EGFP knock-in gene". As in (1) above, no cocoon spinning or mating problems were observed in the individuals developed from the microinjected embryos.

[0129] (4) Knock-in of sericin 1 gene by homologous recombination The EGFP gene sequence was introduced into the second exon of the sericin 1 (hereinafter referred to as "Ser1") gene by homologous recombination so that the EGFP protein was fused to the C-terminus of the signal peptide of Ser1.

[0130] Specifically, a position in the first intron of the Ser1 gene (position between positions 3020 and 3033 in SEQ ID NO: 9) was used as the genome cleavage position, and a double-stranded circular DNA was constructed as a donor nucleic acid to be used in homologous recombination (hereinafter referred to as "SP(Ser1)-EGFP donor nucleic acid for homologous recombination"). The SP(Ser1)-EGFP donor nucleic acid for homologous recombination contains a first genome homologous sequence and a second genome homologous sequence, as well as an EGFP gene sequence, a transcription termination sequence, and a marker gene, which are the target gene sequence arranged therebetween, and contains a TALEN recognition sequence at the end opposite to the EGFP gene sequence of the first genome homologous sequence and the second genome homologous sequence.

[0131] Here, the above-mentioned genomic cleavage site is cleaved by a Left TALEN that recognizes a 19-base-long sequence consisting of positions 3001 to 3019 in SEQ ID NO:9, and a Right TALEN that recognizes a 16-base-long sequence consisting of positions 3034 to 3049 in SEQ ID NO:9.

[0132] The first genome homologous sequence is a 2069-base sequence homologous to the genome sequence from the base located on the 5'-end side of the genome cleavage position on the genome (position 947 in SEQ ID NO: 9) to the codon encoding the C-terminal residue of the signal peptide in the second exon sequence. The first genome homologous sequence has a mutation in the vicinity of the above-mentioned genome cleavage position so as not to be recognized by Left TALEN and Right TALEN (specifically, the base sequence TATATTGTAAAGCACAACATATATATTAATGAATTTTTTATTTATTTTTC (SEQ ID NO: 37) located at positions 3000 to 3049 in SEQ ID NO: 9 is replaced with the base sequence agTATTGagAAGCACAAgtaATATATTAATGAATTTTTTcTTTcTTTTTC (SEQ ID NO: 38)). The second genome homologous sequence is a 2000-base sequence homologous to the genome sequence located on the 3'-end side of the codon encoding the C-terminal residue of the signal peptide in the second exon sequence on the genome. In the SP(Ser1)-EGFP donor nucleic acid for homologous recombination, the base sequence from the restriction enzyme recognition sequence and the first genome homologous sequence to the second genome homologous sequence and the restriction enzyme recognition sequence is shown in SEQ ID NO: 18. Furthermore, the protein in which the Ser1 signal peptide is fused to the N-terminus of the C-terminal fragment of EGFP excluding the initiator methionine, which is encoded by the gene after knock-in, consists of the amino acid sequence shown in SEQ ID NO: 19 (hereinafter referred to as "SP(Ser1)-EGFP fusion protein").

[0133] SP(Ser1)-EGFP donor nucleic acid for homologous recombination was injected into silkworm eggs of the w1-pnd strain 2 to 8 hours after spawning together with mRNAs encoding Left TALEN and Right TALEN synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen). After injection, the eggs were incubated in a humidified state at 25°C until hatching. A knock-in silkworm line was obtained in the same manner as in (2) above. Hereinafter, the knock-in line obtained is referred to as the "SP(Ser1)-EGFP knock-in line" as in (1) above. The Ser1 gene after this knock-in is referred to as the "SP(Ser1)-EGFP knock-in gene".

[0134] As in (1) above, no cocoon spinning or mating problems were observed in the individuals developed from the microinjected embryos.

[0135] Example 2: Production of EGFP protein (the purpose) The expression level of EGFP protein in the silk gland of each knock-in line prepared in Example 1 is measured.

[0136] (Methods and Results) (1) Silkworm strains A line having a combination of multiple knock-in genes was produced by crossbreeding each knock-in line produced in (2) to (4) of Example 1. In this Example, for example, a line having two knock-in genes, an SP(FibH)-EGFP knock-in gene and an SP(FibL)-EGFP knock-in gene, is referred to as an SP(FibH)-EGFP / SP(FibL)-EGFP knock-in line, etc. Note that in the silkworms used in this Example, all of the knock-in genes are heterozygous.

[0137] In this example, a line obtained by crossing a FibH+Ser1-GAL4 line expressing the GAL4 gene under the control of the FibH gene promoter and the Ser1 gene promoter with a UAS-EGFP line expressing the EGFP gene under the control of the UAS regulatory sequence (hereinafter referred to as the "EGFP-producing GAL4 / UAS line") was used as a control group.

[0138] (2) Rearing conditions Silkworms were reared as follows. All stages of larvae were reared on artificial diet (Silkmate Original Species 1-3 Stages S, Nosan Corporation) in a rearing room at 25-27°C. The artificial diet was changed every 2-3 days (Uchino K. et al., 2006, J Insect Biotechnol Sericol, 75:89-97).

[0139] (3) Measurement of EGFP protein expression level Just before spinning on the 6th day of the 5th instar, the 5th instar was anesthetized on ice, and the dorsal side was incised and the silk gland was removed with tweezers without damaging it. The middle and posterior silk glands were removed. Each of these was placed in 10 mL of PBS (pH 7.2) / 1% Tween20 / 0.05% sodium azide and shaken at room temperature for 24 hours to extract water-soluble proteins. The water-soluble protein extract obtained was centrifuged at 2,000 × g for 10 minutes, and the supernatant was collected. The EGFP protein concentration in the water-soluble protein contained in the supernatant was measured by ELISA. Specifically, 100 μL of the supernatant was added to a 96-well plate coated with an anti-GFP antibody (Aves GFP-1010, Cosmo Bio) and left to stand at room temperature for 1 hour. After washing three times with PBS / 0.05% Tween 20, horseradish peroxidase-conjugated anti-GFP antibody (Rockland Immunochemicals) was added and incubated at room temperature for 1 hour. After washing three times with PBS / 0.05% Tween 20, a color reaction was performed using a TMB Peroxidase EIA Substrate Kit (Bio-Rad) and the reaction was stopped by adding 1N sulfuric acid. The color development was quantified using a plate reader (SpectraMax iD3; Molecular Devices). A standard curve was prepared using serial dilutions (1-400 pg / μL) of recombinant GFP protein (Takara Bio; Z2373N). FIG. 6 shows the results of measuring the amount of EGFP expression per silkworm.

[0140] The amount of EGFP expression in the SP(Ser1)-EGFP knock-in line (Fig. 6, SP(Ser1)-EGFP) was slightly higher than that of the EGFP-producing GAL4 / UAS line (Fig. 6, Control(GAL4 / UAS)). In the EGFP-producing GAL4 / UAS line, GAL4 protein is expressed by two promoters, the Ser1 gene promoter and the FibH gene promoter, and the amount of expression derived from the Ser1 gene promoter out of the 3.3 mg of EGFP expression is estimated to be less than 1 mg. Therefore, the amount of EGFP expression in the SP(Ser1)-EGFP knock-in line is considered to be overwhelmingly higher.

[0141] The amount of EGFP expression in the SP(FibH)-EGFP knock-in line (Fig. 6, SP(FibH)-EGFP) was more than twice as high as that in the SP(FibL)-EGFP knock-in line (Fig. 6, SP(FibL)-EGFP), and more than four times as high as that in the SP(Ser1)-EGFP knock-in line (Fig. 6, SP(Ser1)-EGFP). This result was unexpected, because the molar amounts of FibH and FibL proteins produced in the silk gland of silkworms are equal in terms of the composition of fibroin, and the amount of fibroin that constitutes silk thread is three times that of sericin by weight.

[0142] The SP(FibH)-EGFP / SP(FibL)-EGFP / SP(Ser1)-EGFP knock-in line (Figure 6, far right) expressed 33.7 mg of EGFP, which was significantly higher than the expected amount as the sum of the EGFP expression levels in the SP(FibH)-EGFP knock-in line, the SP(FibL)-EGFP knock-in line, and the SP(Ser1)-EGFP knock-in line.

[0143] Example 3: Production of GM-CSF protein (the purpose) We performed knock-in by homologous recombination so that granulocyte-macrophage colony-stimulating factor (GM-CSF) was fused to the C-terminus of the signal peptide of FibH, and evaluated the amount of GM-CSF produced in the silk gland.

[0144] (Methods and Results) (1) Silkworm strains The homologous recombination method described in Example 1(2) was carried out by replacing the EGFP gene sequence with the GM-CSF gene sequence. The GM-CSF gene sequence consists of the base sequence shown in SEQ ID NO: 20, and encodes a GM-CSF protein consisting of the amino acid sequence shown in SEQ ID NO: 21. The resulting knock-in line is called the "SP(FibH)-GM-CSF knock-in line." In addition, the protein encoded by the knock-in gene (FIG. 7A), in which the FibH-derived signal peptide is fused to the N-terminus of the mature amino acid sequence of GM-CSF excluding the GM-CSF-derived signal peptide, consists of the amino acid sequence shown in SEQ ID NO: 22 (hereinafter referred to as the "SP(FibH)-GM-CSF fusion protein").

[0145] In addition, a line obtained by crossing the Ser1-GAL4 line expressing the GAL4 gene under the control of the Ser1 gene promoter with the UAS-GM-CSF line expressing the GM-CSF gene under the control of the UAS regulatory sequence (hereinafter referred to as the "middle silk gland GM-CSF-producing GAL4 / UAS line"), and a line obtained by crossing the FibH-GAL4 line expressing the GAL4 gene under the control of the FibH gene promoter with the UAS-GM-CSF line (hereinafter referred to as the "posterior silk gland GM-CSF-producing GAL4 / UAS line") were used as control groups.

[0146] (2) Western blotting According to the method described in Example 2, the middle and posterior silk glands were removed from each line, and proteins in each silk gland were extracted. Next, the undiluted or 2- to 128-fold diluted silk gland extract was mixed with a specified amount of NuPAGE LDS Sample Buffer (Thermo Fisher) and NuPAGE Sample Reducing Agent (Thermo Fisher), and SDS-converted by heating at 70°C for 10 minutes. The SDS-converted sample was electrophoresed for about 90 minutes at a constant current of 20 mA using an 8 cm x 13 cm 4-12% SDS-PAGE gel. The gel was transferred to a PVDF membrane using a semi-dry transfer device (iBlot2, Thermo Fisher). After the transfer, the membrane was gently shaken for 5 minutes with EZ wash (AE-1480, ATTO), and then reacted with the primary antibody (anti-GM-CSF antibody, 3000-fold dilution, Immundiagnostik, product number AS1021.2) at 4°C overnight. The membrane was washed three times for 10 minutes with EZ wash, and then reacted with the secondary antibody (Anti-Rabbit IgG, HRP-Linked Whole Ab Donkey, Cytiva, product number NA934-100UL, 50,000-fold dilution) at room temperature for 1 hour. The membrane was washed three times for 10 minutes with EZ wash, and then reacted with ECL prime (RPN2232, GE Healthcare) for 5 minutes, and the signal was detected with Fusion FX (Vilber Bio Imaging).

[0147] The results are shown in Figure 7B. In the combined extract of the middle and posterior silk glands of the SP(FibH)-GM-CSF knock-in line, approximately 13-fold and 3-fold more GM-CSF was detected compared to the middle silk gland GM-CSF-producing GAL4 / UAS line and the posterior silk gland GM-CSF-producing GAL4 / UAS line, respectively. This demonstrates that GM-CSF is expressed and secreted extremely efficiently in the silk glands of the SP(FibH)-GM-CSF knock-in line.

[0148] Example 4: Antibody production (the purpose) Knock-in is performed by homologous recombination so that the IgG H chain is fused to the C-terminus of the signal peptide of FibH. Furthermore, knock-in is performed by homologous recombination so that the IgG L chain is fused to the C-terminus of the signal peptide of FibL. The two knock-in strains obtained are crossed to produce antibody molecules containing the IgG H chain and the IgG L chain.

[0149] (Methods and Results) (1) Silkworm strains The homologous recombination method described in Example 1(2) was carried out by replacing the EGFP gene sequence with an IgG H chain gene sequence. The IgG H chain gene sequence consists of the nucleotide sequence shown in SEQ ID NO: 23 and encodes an IgG H chain consisting of the amino acid sequence shown in SEQ ID NO: 24. The resulting knock-in line is called the "SP(FibH)-IgG H chain knock-in line." In addition, the protein encoded by the knock-in gene (FIG. 8A) in which a FibH-derived signal peptide is fused to the N-terminus of the IgG H chain consists of the amino acid sequence shown in SEQ ID NO: 25 (hereinafter referred to as the "SP(FibH)-IgG H chain fusion protein").

[0150] Furthermore, the homologous recombination method described in Example 1(3) was performed by replacing the EGFP gene sequence with an IgG L chain gene sequence. The IgG L chain gene sequence consists of the nucleotide sequence shown in SEQ ID NO: 26, and encodes an IgG L chain consisting of the amino acid sequence shown in SEQ ID NO: 27. The resulting knock-in line is referred to as an "SP(FibL)-IgG L chain knock-in line." Furthermore, the protein in which a FibL-derived signal peptide is fused to the N-terminus of the IgG L chain, encoded by the knock-in gene (FIG. 8B), consists of the amino acid sequence shown in SEQ ID NO: 28 (hereinafter referred to as an "SP(FibL)-IgG L chain fusion protein").

[0151] By crossing the above two knock-in lines, we generated a line that has two knock-in genes and can produce IgG molecules containing an IgG H chain and an IgG L chain (hereinafter referred to as the "SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in line").

[0152] In this example, a line obtained by crossing a FibH+Ser1-GAL4 line expressing the GAL4 gene under the control of the FibH gene promoter and the Ser1 gene promoter, a UAS-IgG H chain line expressing the IgG H chain gene under the control of the UAS regulatory sequence, and a UAS-IgG L chain line expressing the IgG L chain gene under the control of the UAS regulatory sequence (hereinafter referred to as an "antibody-producing GAL4 / UAS line") was used as a control group.

[0153] (2) Quantification of IgG expression levels The middle and posterior silk glands were excised from each line, and proteins in each silk gland were extracted according to the method described in Example 2. Subsequently, IgG was purified using Ab SpinTrap (Cytiva), and the amount of IgG in the extract was quantified using a Protein Assay BCA Kit (Nacalai).

[0154] The results are shown in Figure 8C. It was shown that the SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in line was able to produce approximately 4.6 times more IgG than the middle and posterior silk glands of the antibody-producing GAL4 / UAS line. It was revealed that the silk glands of the SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in line expressed and secreted antibody molecules containing IgG H chain and IgG L chain with extremely high efficiency.

[0155] <Comparative Example 1: Creation of knock-in lines when genome cleavage site is designed within exon sequence> (the purpose) By creating a genomic cleavage site within an exon sequence, not an intron sequence, the EGFP gene is knocked into the fibroin H gene by homologous recombination. The efficiency of generating knock-in lines is compared with that of cleaving the genome within an intron sequence.

[0156] (Methods and Results) In the homologous recombination method described in (2) of Example 1, the method was modified so that the genomic cleavage position was set within the second exon of the FibH gene, and the EGFP gene sequence was knocked into the fibroin H gene. Specifically, a mutation was introduced into the donor nucleic acid for homologous recombination near the genomic cleavage position in the second exon so that it would not be recognized by TALEN. The donor nucleic acid for homologous recombination was injected into 384 eggs together with mRNA encoding TALEN, which was synthesized using mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen) into silkworm eggs of the w1-pnd strain 2 to 8 hours after spawning. The eggs after injection were incubated in a humidified state at 25°C until hatching. The silkworms of the current generation after injection were raised to adulthood and crossed with the parent strain to determine whether they had mating ability.

[0157] The results are shown in Figure 9B. In the case of homologous recombination where the genome is cut within the exon sequence, of the 118 larvae that grew to the 5th instar, 113 were incapable of pupation or had naked pupae or thin cocoons, and even when they emerged, they showed mating failure, and only 5 larvae produced normal cocoons. However, of the 5 normal cocoons, only 2 were capable of mating, and only 1 of the 2 females and 1 of the 3 males were capable of mating. The reason why abnormalities occur when the exon sequence is cut may be that in the injected embryos, mutations such as frameshifts occur at the genome cut site, resulting in the failure to express normally functioning proteins.

[0158] This result is in contrast to the result in Example 1(2) where almost no cocoon spinning or mating defects were observed when the genome was cut within the intron sequence (FIG. 9A), and indicates that the knock-in silkworm line can be produced with overwhelmingly high efficiency by the method described in Example 1. This is thought to be because even if some mutations are introduced into the intron sequence, the effect on normal protein expression is negligible.

Claims

1. Genetically modified lepidopteran insects, In the endogenous gene of the genetically modified lepidopteran insect, the exon sequence encoding the signal peptide or a functional fragment thereof of the endogenous gene includes a target gene sequence encoding the target protein or a fragment thereof. The genetically modified lepidopteran insect wherein the target protein or a fragment thereof is fused to the C-terminal side of the signal peptide or a functional fragment thereof.

2. The genetically modified lepidopteran insect according to claim 1, wherein the endogenous gene encodes fibroin, sericin, and / or fibrohexamarin.

3. The genetically modified lepidopteran insect according to claim 2, wherein the fibroin is a fibroin H chain and / or a fibroin L chain.

4. The aforementioned endogenous gene, Fibroin H chain and fibroin L chain, Fibroin H chain and sericin 1, or Fibroin H chain, fibroin L chain, and sericin 1 A genetically modified lepidopteran insect according to claim 1, which codes for the insect.

5. The genetically modified lepidopteran insect according to claim 1, wherein the exon sequence includes a transcription termination sequence at the 3' end of the target gene sequence.

6. A double-stranded circular DNA for introducing a target gene sequence at a genome break site within an intron sequence in an endogenous gene of a genetically modified lepidopteran insect, The aforementioned endogenous gene is, (a) A first spacer sequence adjacent to the 5' end of the genome cleavage site, (b) A second spacer sequence adjacent to the 3' end of the genome cleavage site, (c) A first recognition sequence recognized by a first genome editing enzyme at the 5' end of the first spacer sequence, and (d) A second recognition sequence recognized by a second genome editing enzyme at the 3' end of the second spacer sequence. Includes, The double-stranded circular DNA comprises the first recognition sequence, the second spacer sequence, the first spacer sequence, a genome homologous sequence, and a target gene sequence in this order. The genome homologous sequence consists of a nucleotide sequence homologous to the genome sequence from the second recognition sequence to the exon sequence or a subsequence thereof located at the 3' end of the intron sequence, and the target gene sequence is the double-stranded circular DNA encoding the target protein or a fragment thereof to be fused to the C-terminal side of the signal peptide or functional fragment thereof of the endogenous gene.

7. A donor nucleic acid for producing genetically modified lepidopteran insects using homologous recombination, The homologous recombination method described above includes cutting the genome break site within the intron sequence of an endogenous gene with a genome editing enzyme, The donor nucleic acid is, (a) A first genome homologous sequence and a second genome homologous sequence derived from the endogenous gene, and (b) Target gene sequence located between them Includes, The first genome homologous sequence consists of a base sequence homologous to the genome sequence from a base located 5' end-side to the genome break site on the genome to an exon sequence or a subsequence thereof located 3' end-side to the intron sequence, and has a mutation in the recognition sequence of the genome editing enzyme. The second genome homologous sequence consists of a base sequence homologous to a genome sequence located 3' end to the exon sequence or a subsequence thereof on the genome. The target gene sequence is the donor nucleic acid that encodes the target protein or fragment thereof to be fused to the C-terminal side of the signal peptide or functional fragment thereof of the endogenous gene.

8. A donor nucleic acid for producing genetically modified lepidopteran insects using homologous recombination, The homologous recombination method described above includes cutting the genome break site within the intron sequence of an endogenous gene with a genome editing enzyme, The donor nucleic acid is, (a) A first genome homologous sequence and a second genome homologous sequence derived from the endogenous gene, and (b) Target gene sequence located between them Includes, The first genome homologous sequence consists of a base sequence homologous to the genome sequence from a base located 5' end-side to the intron sequence to an exon sequence or a subsequence thereof located 5' end-side to the intron sequence. The second genome homologous sequence consists of a base sequence homologous to the genome sequence from a base located 3' end-to-5 The target gene sequence is the donor nucleic acid that encodes the target protein or fragment thereof to be fused to the C-terminal side of the signal peptide or functional fragment thereof of the endogenous gene.

9. The donor nucleic acid according to claim 7 or 8, wherein the end of the first genome homologous sequence and / or the second genome homologous sequence opposite to the target gene sequence includes a nuclease recognition sequence.

10. The donor nucleic acid according to claim 9, wherein the nuclease recognition sequence is the recognition sequence of the genome editing enzyme or the restriction enzyme recognition sequence.

11. A method for creating genetically modified lepidopteran insects, Double-stranded circular DNA according to claim 6, The first genome editing enzyme or a nucleic acid encoding the first genome editing enzyme in a state capable of expression, and The second genome editing enzyme or a nucleic acid encoding the second genome editing enzyme in a state capable of expression The introduction process involves introducing the substance into the eggs of lepidopteran insects using the microinjection method. The method, including the method described above.

12. A method for creating genetically modified lepidopteran insects, The donor nucleic acid according to claim 7 or 8, and The genome editing enzyme, or a nucleic acid encoding the genome editing enzyme in a state capable of expression. The introduction process involves introducing the substance into the eggs of lepidopteran insects using the microinjection method. The method, including the method described above.

13. A method for producing the target protein or a fragment thereof using a genetically modified lepidopteran insect according to any one of claims 1 to 5, or a genetically modified lepidopteran insect produced by the method described in claim 11.

14. The method according to claim 13, wherein the lepidopteran insect is a silkworm, and the target protein or a fragment thereof is produced in the silk gland of the silkworm.