Genetically recombinant lepidoptera insects
By knocking the target gene into the exon sequence of the silkworm's endogenous gene and utilizing the promoter and enhancer activities of the endogenous gene, the problems of unstable and inefficient expression in the silk gland were solved, and the stable and large-scale production of the target protein, especially the high expression of fluorescent protein, was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies for expression systems in silkworm glands suffer from unstable and inefficient expression levels. In particular, the GAL4/UAS system requires mating for construction and exhibits significant variations in expression levels. Furthermore, the fusion protein loses its activity, making it difficult to achieve stable and large-scale production of the target protein.
A new expression system was constructed by knocking in the exon sequence of the signal peptide or functional fragment of the silkworm's endogenous gene, utilizing the promoter and enhancer activities of the endogenous gene, and combining the genome editing enzyme to insert the target gene at the cleavage site within the intron sequence.
Stable and high-volume production of the target protein in silkworms was achieved, significantly increasing the expression level. Fluorescent cocoons and fluorescent filaments with strong fluorescence were successfully produced, solving the problems of unstable and inefficient expression.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to recombinant lepidopteran insects and methods for their preparation. Background Technology
[0002] The silk glands of the silkworm (Bombyx mori) possess the ability to synthesize large quantities of protein in a short period. Furthermore, the silk glands are large organs, making them easy to remove, and the synthesized proteins are stored within the gland lumen, facilitating their recovery. Therefore, recombinant silkworms expressing target proteins through their silk glands are considered promising systems for mass protein production.
[0003] The silk glands of silkworms are a pair of organs, consisting of three regions: the anterior, middle, and posterior silk glands. In the posterior silk gland cells, three main proteins constituting the fibrous components of silk—fib H (often simply referred to as "Fib H"), fib L (often simply referred to as "Fib L"), and fibrohexamerin (also known as p25 / FHX)—are expressed. In the middle silk gland cells, sericin, a gelatinous protein that forms the silk coating, is expressed. These three proteins expressed in the posterior silk gland cells form a complex (SFEU complex) in a ratio of Fib H:Fib L:p25 = 6:6:1, which is then secreted into the lumen of the posterior silk gland. Conversely, sericin, after expression, is secreted into the lumen of the middle silk gland. The sericin secreted into the lumen of the posterior silk gland is then transferred to the lumen of the middle silk gland, where it is coated with sericin and spun into silk (Non-Patent Literature 1). Therefore, when using the silk gland as a protein expression system, a gene expression system specifically expressed in the middle or posterior silk gland can be utilized.
[0004] In the case of using silk glands as a protein expression system, large-scale expression methods using the GAL4 / UAS system as a recombinant protein expression system (Non-Patent Document 2) and the system combining the sericin 1 promoter and the Hr3 enhancer have been reported to date (Non-Patent Document 3). However, the current situation is that the GAL4 / UAS system, which is superior in terms of protein expression levels, is widely used.
[0005] The GAL4 / UAS system is a gene control system that combines the yeast transcription factor GAL4 and the control sequence UAS. In the GAL4 / UAS system used as a protein production system in silkworm silk glands, a GAL4 system expressing the GAL4 gene under the promoter control of a gene specifically expressed in the middle or posterior silk glands, and a UAS system expressing the target protein gene under the control of the control sequence UAS, were independently established through gene recombination with piggyBac. The two systems were then interbred to construct an expression system for the target protein in the silk glands.
[0006] In the GAL4 / UAS system, separate GAL4 and UAS systems need to be established before mating, thus the construction of the expression system takes time. Furthermore, the expression level of the target protein varies because the GAL4 gene and UAS control sequence are introduced into random locations on the genome, posing obstacles to the establishment of new GAL4 and UAS systems. Additionally, it is believed that the expression level in the GAL4 / UAS system has reached a technical limit.
[0007] Non-Patent Literature 4 discloses various methods previously implemented for expressing recombinant proteins in the silk glands of silkworms. However, the methods reported in the past have problems such as low expression levels of fusion proteins and loss of activity of proteins fused with silk proteins.
[0008] Therefore, there is a need for new methods for the stable and large-scale production of target proteins, and new methods for the efficient production of silk proteins fused with target proteins.
[0009] Existing technical documents
[0010] Non-patent literature
[0011] Non-patent literature 1 Inoue S. et al., 2000, The Journal of Biological Chemistry, 275 (51): 40517-40528.
[0012] Non-patent literature 2 Tatematsu K. et al., 2010, Transgenic Research, 19(3):473-87.
[0013] Non-patent literature 3 Tomita M. et al., 2007, Transgenic Research, 16 (4):449-465.
[0014] Non-patent literature 4 Tomita M. et al., 2011, Biotechnol Lett, 33:645-654. Summary of the Invention
[0015] The object of this invention is to provide a novel expression system for the stable and large-scale production of a target protein in lepidopteran insects such as silkworms. Additionally, a novel method for producing functional filaments containing filament proteins fused with the target protein is also a subject of this invention.
[0016] In order to solve the above-mentioned problems, the inventors conceived of constructing a new expression system that directly utilizes the promoter and enhancer activities of the endogenous gene to express the target gene by knocking into the exon sequence encoding the signal peptide in the sericin gene, filamentin gene, etc.
[0017] Generally, to effectively knock foreign genes into the silkworm genome, it is necessary to cut the target gene locus using genome editing enzymes or similar methods. Therefore, the inventors of this invention set a genome cutting position within the exon sequence of the target gene sequence and attempted to knock it into that exon sequence. However, the result of implementing this method was that over 95% of the injected individuals could not spin cocoons normally, and over 98% of the individuals could not develop into mating adults. Therefore, it was determined that establishing a system using this method is extremely difficult.
[0018] Therefore, the inventors attempted to knock-in the target gene sequence into the exon sequence by cutting the genome not within the exon sequence to which the target gene sequence is to be introduced, but within the intron sequence near the exon sequence. The results showed that almost all contemporary individuals injected with the gene developed into mating-capable adults, producing the target protein at a much higher rate than the existing GAL4 / UAS system.
[0019] Furthermore, the inventors used the aforementioned knock-in technique to introduce a fluorescent protein gene into the exon sequence of an endogenous filamentin gene, creating a cocoon composed of a fusion protein containing both fluorescent protein and full-length filamentin. The result was the successful creation of novel fluorescent cocoons and filaments that, under white light, exhibit an overwhelmingly stronger color compared to the strongest fluorescent cocoons and filaments in existing technologies, thus completing this invention. Based on the above research findings, this invention provides the following solution.
[0020] (1) Genetic recombination in Lepidoptera insects, The recombinant lepidopteran insects contain a target gene sequence encoding a target protein or a fragment thereof in the exon sequence encoding the signal peptide or functional fragment of the endogenous gene. The target protein or a fragment thereof is fused to the C-terminus of the signal peptide or its functional fragment.
[0021] (2) The recombinant lepidopteran insect according to (1), wherein the endogenous gene encodes filamentin, sericin and / or fibrous hexamer protein.
[0022] (3) According to the gene recombinant lepidopteran insect described in (2), the filoprotein is filoprotein H chain and / or filoprotein L chain.
[0023] (4) According to the gene recombination lepidopteran insect described in (1), the endogenous gene encodes
[0024] Silk core protein H chain and silk core protein L chain, Silken H chain and sericin 1, or Silken H chain, silken L chain and sericin 1.
[0025] (5) The recombinant lepidopteran insect according to any one of (1) to (4), wherein the exon sequence contains a transcription termination sequence at the 3' end of the target gene sequence.
[0026] (6) Double-stranded circular DNA is a type of double-stranded circular DNA used to introduce the target gene sequence into the genome cleavage site within the intron sequence of endogenous genes in lepidopteran insects for gene recombination. The endogenous genes include (a) The first spacer sequence adjacent to the 5' end of the genome cleavage site, (b) The second spacer sequence adjacent to the 3' end of the genome cleavage site, (c) The first recognition sequence recognized by the first genome editing enzyme at the 5' end of the first spacer sequence, and (d) The second recognition sequence, which is recognized by the second genome editing enzyme at the 3' end of the second spacer sequence. The double-stranded circular DNA sequentially comprises the first recognition sequence, the second spacer sequence, the first spacer sequence, a genomic homologous sequence, and the target gene sequence. The genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the second recognition sequence to the exon sequence or a portion thereof located at the 3' end of the intron sequence. The target gene sequence encodes a target protein or fragment thereof that is fused to the C-terminus of the signal peptide or functional fragment thereof of the endogenous gene.
[0027] (7) Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the 5' end of the genome cut site to the 3' end of the exon sequence or a portion thereof, and the recognition sequence of the genome editing enzyme has a mutation. The second genomic homologous sequence consists of a base sequence homologous to a genomic sequence located on the 3' end of the exon sequence or a portion thereof. The target gene sequence encodes a target protein or fragment thereof that is fused to the C-terminus of the signal peptide or functional fragment thereof of the endogenous gene.
[0028] (8) Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the bases located at the 5' end of the intron sequence to the exon sequence or a portion thereof located at the 5' end of the intron sequence. The second genomic homologous sequence consists of a base sequence homologous to the genomic sequence from the bases located at the 3' end of the exon sequence or a portion thereof and at the 5' end of the genomic cut site, to the bases located at the 3' end of the genomic cut site, and the recognition sequence of the genome editing enzyme has a mutation. The target gene sequence encodes a target protein or fragment thereof that is fused to the C-terminus of the signal peptide or functional fragment thereof of the endogenous gene.
[0029] (9) The donor nucleic acid according to (7) or (8) contains a nuclease recognition sequence at the end of the first genomic homologous sequence and / or the second genomic homologous sequence opposite to the target gene sequence.
[0030] (10) According to the donor nucleic acid described in (9), the nuclease recognition sequence is the recognition sequence of the genome editing enzyme or the restriction endonuclease recognition sequence.
[0031] (11) A method for producing recombinant lepidopteran insects, the method comprising: […].
[0032] (6) The double-stranded circular DNA, The first genome editing enzyme, or the nucleic acid encoding the first genome editing enzyme in an expressible state, and The second genome editing enzyme, or the nucleic acid encoding the second genome editing enzyme in an expressible state. The process of introducing microinjection into the eggs of lepidopteran insects.
[0033] (12) A method for producing recombinant lepidopteran insects, the method comprising: […].
[0034] The donor nucleic acid described in any one of (7) to (10), and
[0035] The genome editing enzyme, or the nucleic acid encoding the genome editing enzyme in an expressible state.
[0036] The process of introducing microinjection into the eggs of lepidopteran insects.
[0037] (13) A method for producing the target protein or a fragment thereof using the recombinant lepidopteran insect described in any one of (1) to (5), or the recombinant lepidopteran insect prepared by the method described in (11) or (12).
[0038] (14) According to the method of (13), the lepidopteran insect is a silk-producing insect, and the target protein or a fragment thereof is produced in the silk gland of the silk-producing insect.
[0039] The present invention further provides the following solutions.
[0040] (1) Genetic recombination in Lepidoptera insects, The recombinant lepidopteran insects contain a target gene sequence encoding a target protein or a fragment thereof in the exon sequence encoding the signal peptide or functional fragment of the endogenous gene. The target protein or a fragment thereof is fused between the signal peptide or a functional fragment thereof and a mature protein or its C-terminal fragment obtained by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
[0041] (2) According to the gene recombinant lepidopteran insect described in (1), the mature protein is filamentin, sericin, and / or fibrous hexamer protein.
[0042] (3) According to the gene recombinant lepidopteran insect described in (2), the filoprotein is filoprotein H chain and / or filoprotein L chain.
[0043] (4) The recombinant lepidopteran insect according to any one of (1) to (3), wherein the target protein is selected from fluorescent proteins, antibodies, antigenic peptides, enzymes, cytokines and antimicrobial peptides.
[0044] (5) The recombinant lepidopteran insect according to any one of (1) to (4) homozygously contains the exon sequence containing the target gene sequence.
[0045] (6) Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the 5' end of the genome cut site to the 3' end of the exon sequence or a portion thereof, and the recognition sequence of the genome editing enzyme has a mutation. The second genomic homologous sequence consists of a base sequence homologous to a genomic sequence located on the 3' end of the exon sequence or a portion thereof. The target gene sequence encodes a target protein or a fragment thereof that is fused between a signal peptide or a functional fragment thereof of the endogenous gene and a mature protein or a C-terminal fragment thereof formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
[0046] (7) Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the bases located at the 5' end of the intron sequence to the exon sequence or a portion thereof located at the 5' end of the intron sequence. The second genomic homologous sequence consists of a base sequence homologous to the genomic sequence from the bases located at the 3' end of the exon sequence or a portion thereof and at the 5' end of the genomic cut site, to the bases located at the 3' end of the genomic cut site, and the recognition sequence of the genome editing enzyme has a mutation. The target gene sequence encodes a target protein or a fragment thereof that is fused between a signal peptide or a functional fragment thereof of the endogenous gene and a mature protein or a C-terminal fragment thereof formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
[0047] (8) The donor nucleic acid according to (6) or (7) contains a nuclease recognition sequence at the end of the first genomic homologous sequence and / or the second genomic homologous sequence opposite to the target gene sequence.
[0048] (9) According to the donor nucleic acid described in (8), the nuclease recognition sequence is the recognition sequence of the genome editing enzyme or the recognition sequence of the restriction endonuclease.
[0049] (10) A method for producing a genetically recombinant lepidopteran insect, the method comprising: […].
[0050] The donor nucleic acid described in (6) or (7), and
[0051] The genome editing enzyme, or the nucleic acid encoding the genome editing enzyme in an expressible state.
[0052] The process of introducing microinjection into the eggs of lepidopteran insects.
[0053] (11) A method for producing a fusion protein is a method of producing a fusion protein comprising the target protein or a fragment thereof, and the mature protein or a C-terminal fragment thereof, using any one of the recombinant lepidopteran insects described in (1) to (5) or a recombinant lepidopteran insect prepared by the method described in (10).
[0054] (12) According to the method of (11), the lepidopteran insect is a silk-producing insect, and the fusion protein is produced in the silk gland of the silk-producing insect.
[0055] (13) Fusion protein, starting from the N-terminus, sequentially contains
[0056] Target protein or fragment thereof, and
[0057] Silken protein, sericin, or fibrohexamer protein.
[0058] (14) According to the fusion protein of (13), the filoprotein is filoprotein H chain and / or filoprotein L chain.
[0059] (15) A cocoon or silk containing the fusion protein described in (13) or (14).
[0060] (16) Cocoon or silk, The silk protein contained in the cocoon or silk is selected from any one or more of the following: silk core protein H chain, silk core protein L chain, sericin 1, sericin 2, sericin 3 and fibrous hexamer protein, and is composed of the fusion protein described in (13) or (14).
[0061] (17) The cocoon or silk according to (16), the recombinant lepidopteran insect from any one of (1) to (5), or the recombinant lepidopteran insect produced by the method described in (10).
[0062] (18) The target protein is selected from fluorescent proteins, antibodies, antigenic peptides, enzymes, cytokines and antimicrobial peptides according to the cocoon or silk described in (16) or (17).
[0063] This specification contains the disclosure of Japanese Patent Application No. 2023-114712, which forms the basis of the priority claim of this application. Attached Figure Description
[0064] Figure 1 This illustrates the introduction of a target gene sequence with a stop codon into an endogenous gene. The target gene sequence is then introduced into the exon sequence encoding a signal peptide within the endogenous gene. The target protein encoded by the target gene sequence is fused to the C-terminus of the signal peptide encoded by the endogenous gene.
[0065] Figure 2 This diagram illustrates the genomic cut location in an endogenous gene during the creation of the knock-in system. The genomic cut location is designed within an intron sequence located at the 5' end of the exon sequence of the introduced target gene sequence. As an example of an endogenous gene for introducing the target gene sequence, Figure 2 A shows the sericin 1 gene. Figure 2 B shows the filamentin H gene. Figure 2 C shows the L gene of filamentin.
[0066] Figure 3 This is a diagram showing an overview of gene knock-in using the TAL-PITCh method. Figure 3 A shows the location of the sequences used in the construction of the donor nucleic acids used in the TAL-PITCh method within the endogenous gene. Figure 3 B shows the structure of the double-stranded circular DNA used in the TAL-PITCh method. Figure 3 C shows the structure of the knock-in gene obtained by knocking in the target gene sequence using the TAL-PITCh method.
[0067] Figure 4 This is a diagram showing the gene knock-in method using homologous recombination.
[0068] Figure 5 The results of observations on silk glands and cocoons in 5th instar larvae of the SP(FibH)-EGFP knock-in system are shown.
[0069] Figure 6 This chart displays the EGFP expression levels of each silkworm in each knock-in system. The bar chart shows the average values for n=1 to 4, and the error bars represent the standard error.
[0070] Figure 7 This demonstrates GM-CSF production in the silkworm system where the GM-CSF gene sequence was knocked into the second exon of the endogenous silkheart protein H gene. Figure 7 A shows the knock-in of the GM-CSF gene sequence into the filoprotein H gene. Figure 7 B shows the results of GM-CSF detection using Western blotting.
[0071] Figure 8 This demonstrates IgG production in the silkworm system by knocking in the gene sequences encoding the IgG H chain and IgG L chain into the second exon of the endogenous filamentin H gene and the third exon of the endogenous filamentin L gene, respectively. Figure 8 A shows the knock-in of the IgG H chain gene sequence into the filoprotein H gene. Figure 8 B shows the knock-in of the IgG L chain gene sequence into the filoprotein L gene. Figure 8 C indicates IgG production levels.
[0072] Figure 9This demonstrates the introduction of a target gene sequence without a stop codon into an endogenous gene. The target gene sequence is then introduced into the exon sequence encoding a signal peptide within the endogenous gene. The target protein encoded by the target gene sequence is fused between the secretory peptide and the mature protein formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
[0073] Figure 10 This is a diagram showing the gene knock-in method using homologous recombination.
[0074] Figure 11 The results show the detection of EGFP protein using Western blotting. Figure 11 A shows the results of the wild-type system (WT) and the EGFP-FibL knock-in system. Figure 11 B shows the results of the wild-type system (WT), SP(Ser1)-EGFP knock-in system, and EGFP-Ser1 knock-in system.
[0075] Figure 12 The image shows a photograph taken under normal white light of a cocoon obtained from an EGFP-FibL knock-in system containing the EGFP-FibL knock-in gene, either heterozygous or homozygous.
[0076] Figure 13 The images show cocoons produced by heterozygotes and homozygotes of the EGFP-FibL knock-in system prepared by the method of the present invention, as well as cocoons prepared using the piggyBac system of the prior art, observed under white light or by fluorescence.
[0077] Figure 14 The results show the knock-in effect achieved by designing the genome cut-off site within an intron or exon sequence. Figure 14 A shows the result of homologous recombination through genomic cleavage within the intron sequence of the Fib H gene. Over 99% of the individuals developed into mating-capable adults, and no poor cocooning or mating problems were observed. Figure 14 B shows the result of homologous recombination through genomic cleavage within the exon sequence of the Fib H gene. Less than 2% of individuals develop into mating-capable adults, and over 95% of contemporary individuals injected with the gene show incomplete pupation, naked pupae, or thin cocoons. Detailed Implementation
[0078] 1. Genetic recombination in Lepidoptera insects
[0079] 1-1. Overview
[0080] The first embodiment of this invention is a recombinant lepidopteran insect. This recombinant lepidopteran insect contains a target gene sequence in the exon sequence of a signal peptide or a functional fragment thereof encoding an endogenous gene, and expresses a target protein or fragment thereof fused to the C-terminus of the signal peptide or its functional fragment. This recombinant lepidopteran insect is able to stably and abundantly produce the target protein.
[0081] 1-2. Definition
[0082] The following terms, which are frequently used in this specification, are defined.
[0083] In this instruction manual, "lepidoptera" refers to insects belonging to the taxonomic order Lepidoptera, commonly known as butterflies or moths. Butterflies include insects belonging to the families Nymphalidae, Papilionidae, Pieridae, Lycaenidae, and Hesperiidae. Moths include insects belonging to the families Saturniidae, Bombycidae, Brahmaeidae, Euptotidae, Lasiocampidae, Psychidae, Geometridae, Archtiidae, Noctuidae, Pyralidae, and Sphingidae. For example, moths can be listed as species belonging to the genera *Bombyx*, *Samia*, *Antheraea*, *Saturnia*, *Attacus*, and *Rhodinia*, specifically silkworms, *Bombyx mandarina*, *Samia cynthia* (including *Samia cynthia ricini* and hybrids of *Samia cynthia* and *Samia cynthia ricini*), *Antheraea yamamai*, *Antheraea pernyi*, *Saturnia japonica*, and *Actias gnoma*. The lepidopteran insects that serve as hosts for the transformants of this invention are not limited to these, but silkworms, which have high industrial applicability, are particularly preferred as hosts.
[0084] The term "recombinant lepidopteran insect" refers to a recombinant lepidopteran insect or its offspring containing foreign DNA, created using gene recombination technology. In this specification, recombinant lepidopteran insects specifically refer to recombinant lepidopteran insects obtained by introducing foreign DNA into the eggs of lepidopteran insects via microinjection.
[0085] The term "silk gland" in this instruction manual refers to a tubular organ derived from a salivary gland, which has the function of producing, accumulating, and secreting liquid silk. Silk glands typically exist in pairs along the digestive tract of silk-producing insects, primarily in the larval stage. Each silk gland consists of three regions: anterior, middle, and posterior. The posterior silk gland produces and secretes filamentin, a fibrous component of silk. Additionally, the middle silk gland produces and secretes sericin, a covering component, which, along with the filamentin transferred from the posterior silk gland, accumulates in its lumen.
[0086] In this specification, "endogenous gene" refers to a gene of lepidopteran origin that is innately present in the genome of the lepidopteran insect. In principle, endogenous genes in this invention are genes encoding proteins with signal peptides. Therefore, in principle, endogenous genes in this specification are genes encoding secretory proteins or membrane proteins. Secretory proteins can be, for example, any protein that constitutes silk. Furthermore, in this specification, any protein that constitutes silk is often simply referred to as "silk protein," and the gene encoding silk protein is called a "silk gene." Specific examples of endogenous genes in lepidopteran insects include genes encoding filoin, sericin, and fibrous hexamer.
[0087] In addition, the term "exogenous gene" or "foreign gene" used in this specification refers to foreign genes acquired through artificial manipulation or other means, which are genes that do not exist in the genome of wild-type Lepidoptera insects.
[0088] "Fibrin" is a protein that makes up the fibrous components of silk. Silkworm fibrin is mainly composed of three proteins: fibrin H chain (Fib H), fibrin L chain (Fib L), and fibrohexamerin. Fibrohexamerin, as mentioned above, is also known as p25 / FHX.
[0089] "Serin" is a protein that forms a layered outer layer covering the fibers formed by sericin in silk. In silkworms, sericin is synthesized in the central silk gland cells and then secreted into the lumen of the central silk gland. Besides the adhesion of sericin fibers to each other, sericin is known to protect these fibers from external stimuli. Silkworms begin spinning silk immediately after hatching, but the protein composition of the silk and cocoon differs at different instars, and the composition of sericin variants also varies. Generally, about six sericin protein variants (seriin 1A', sericin 1C, sericin 1D, sericin 2, sericin 3, and sericin 4) are known to be biosynthesized by four sericin genes (Ser1, Ser2, Ser3, and Ser4). The cocoon contains four main sericin variants (seriin 1A', sericin 1C, sericin 1D, and sericin 3). In this manual, when the term is simply used to refer to "sericin," it refers to all types of sericin unless otherwise specified.
[0090] The term "signal peptide" or "secretion signal" used in this specification refers to the extracellular transfer signal required for the secretion of proteins biosynthesized through gene expression into the extracellular space. The signal peptide is cleaved by a signal peptidase after translation and before being secreted into the extracellular space. Furthermore, the signal peptide is frequently referred to as "SP" in this specification, with the name of the endogenous gene from which the signal peptide originates listed in parentheses. Signal peptides are typically short peptide sequences of a few tens of amino acids or less, characterized by their hydrophobicity. The sequence of the signal peptide can be predicted using tools such as signalP based on the amino acid sequence of the protein, or using structural predictions provided in databases. For example, in the case of silkworms, the sequence region of the signal peptide can be determined based on sequence annotations available in databases such as KAIKObase and KAIKOcDNA, which are accessible through the Agricultural Genomics Information Database.
[0091] In this specification, the term "functional fragment" of a signal peptide refers to a fragment consisting of a portion of the signal peptide's sequence that maintains its extracellular signal transduction activity. The functional fragment of a signal peptide can be a fragment that maintains, for example, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or equivalent or greater of the extracellular signal transduction activity of the full-length signal peptide. The amino acid length of the functional fragment is not particularly limited as long as it maintains the activity of the full-length signal peptide, but for example, it can be 50%, 60%, 70%, 80%, 90%, 95%, or 98% or greater of the full-length signal peptide.
[0092] In this specification, "full-length" refers to the entire amino acid sequence of a protein that is synthesized and functions in an organism, or the entire base sequence of the gene encoding that amino acid sequence. In principle, in the case of genes, the sequence from the start codon to the stop codon corresponds to a full-length gene; in the case of proteins, a polypeptide or peptide composed of the amino acid sequence encoded by the full-length gene corresponds to a full-length protein. However, in the case of secretory proteins, the endogenous signal peptide contained at the N-terminus is cleaved and removed during secretion and is ultimately not included in the protein. Therefore, in the case of secretory proteins, "full-length" may not include the signal peptide. Furthermore, this specification distinguishes between full-length proteins before the signal peptide is cleaved and removed during secretion as "precursor proteins" and those after the signal peptide is cleaved and removed as "mature proteins."
[0093] In this specification, "exon" refers to the region in a gene's base sequence that remains in the mature transcript. Generally, in eukaryotes, after a gene is transcribed as a primary transcript, the intermediate region called an "intron" is removed through splicing, and the exons connect to form the mature transcript. In this specification, "exon sequence" refers to the base sequence equivalent to an exon, and "intron sequence" refers to the base sequence equivalent to an intron. In any gene, exon and intron sequences can be determined by comparing the gene's genomic sequence with its cDNA sequence. Alternatively, sequence information publicly available in databases such as the National Center for Biotechnology Information (NCBI) can be obtained, or exon / intron structure can be predicted using genomic analysis tools available in this technical field. For example, for silkworm gene information, exon and intron sequences can be retrieved using databases such as KAIKObase and KAIKOcDNA, which are accessible through the Agricultural and Livestock Genomics Information Database.
[0094] The term "target gene sequence" in this specification refers to the gene sequence encoding the target protein or a fragment thereof. The target gene sequence can be a genome-derived gene sequence or a gene sequence composed of cDNA. The target gene sequence may or may not contain introns. In addition to the gene sequence encoding the target protein or a fragment thereof, the target gene sequence may or may not contain a stop codon. Downstream of the stop codon may or may not contain a transcription termination sequence.
[0095] The term "target protein" in this specification refers to the desired protein encoded by the target gene. The type of target protein is not limited. It can be either a structural protein or a functional protein. Examples of structural proteins include fibrous proteins such as collagen, actin, myosin, and filamentin, as well as keratin and histones. Examples of functional proteins include peptide hormones (insulin, calcitonin, parathyroid hormone, growth hormone, etc.), cytokines (granulocyte-macrophage colony-stimulating factor (GM-CSF), epidermal growth factor (EGF), fibroblast growth factor (FGF), interleukin (IL), interferon (IFN), tumor necrosis factor-α (TNF-α), transforming growth factor-β (TGF-β), etc.), transcription factors (including GAL4), antibodies (immunoglobulins, etc.), serum albumin, hemoglobin, enzymes, fluorescent proteins, pigment-synthesizing proteins, and luminescent proteins. Immunoglobulins can be any class (e.g., IgG, IgE, IgM, IgA, IgD, and IgY) or any subclass (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2). Fluorescent proteins are not limited and can be, for example, CFP, AmCyan, RFP, DsRed, YFP, and GFP (including derivatives such as EGFP and EYFP). Pigment synthesis proteins can be, for example, proteins involved in the biosynthesis of melanin-based pigments (including dopamine-melanin), ocular pigments, or pteridine-based pigments. Luminescent proteins can be, for example, jellyfish luminescent proteins or luciferases. Furthermore, the target protein can be either a wild-type protein or a mutant protein.
[0096] In this specification, the term "fragment" of a protein refers to a polypeptide or peptide that contains a portion of the full-length protein. The fragment preferably retains activity. For example, it may be a fragment that retains 50%, 60%, 70%, 80%, 90%, 95%, 98%, or equivalent or more of the activity of the full-length protein. The amino acid length of the fragment is not particularly limited as long as it retains the activity of the full-length protein; for example, it may be 50%, 60%, 70%, 80%, 90%, 95%, or 98% or more of the full-length protein.
[0097] The term "transcription termination sequence" in this specification refers to a sequence capable of terminating gene transcription, also known as a terminator. The type of transcription termination sequence is not particularly limited. Preferably, it is a terminator derived from the same biological species as the gene-recombining lepidopteran insect. For example, for insects such as silkworms, the hsp70 terminator, SV40 terminator, etc., can be used.
[0098] In this specification, "multiple" refers to integers of 2 or more, such as 2 to 10, 2 to 7, 2 to 5, 2 to 4, or 2 to 3 integers.
[0099] The term "identity" of a base sequence in this specification refers to the percentage (%) of identical bases in the entire length of the base sequence when comparing (aligning) two base sequences with the maximum number of identical bases, by inserting appropriate gaps between one or both as needed.
[0100] In this specification, "homologous sequence" refers to a base sequence that shares approximately 60% or more identity with a reference sequence. The identity of a homologous sequence with respect to a reference sequence can be, for example, 70% or more, 80% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 99.9% or more. Furthermore, "genomic homologous sequence" refers to a base sequence that shares any of the above-mentioned identity with a lepidopteran insect genome sequence as a reference sequence, and "genomic sequence" refers to a sequence that shares 100% identity with its corresponding base sequence on the genome.
[0101] The term "amino acid identity" in this specification refers to the percentage (%) of the total number of amino acid residues when comparing (aligning) two polypeptides by maximizing the number of identical amino acid residues in their amino acid sequences, with appropriate gaps inserted in one or both as needed.
[0102] The term "amino acid substitution" in this specification refers to substitutions among the 20 amino acids that constitute natural proteins within conserved groups of amino acids with similar properties such as charge, side chain, polarity, and aromaticity. Examples include substitutions within groups of charge-free polar amino acids with low-polarity side chains (Gly, Asn, Gln, Ser, Thr, Cys, Tyr), branched-chain amino acids (Leu, Val, Ile), neutral amino acids (Gly, Ile, Val, Leu, Ala, Met, Pro), neutral amino acids with hydrophilic side chains (Asn, Gln, Thr, Ser, Tyr, Cys), acidic amino acids (Asp, Glu), basic amino acids (Arg, Lys, His), and aromatic amino acids (Phe, Tyr, Trp).
[0103] Unless otherwise specified, the terms "5' end" and "3' end" in this specification refer to the 5' and 3' ends of the transcripts transcribed from the endogenous gene, respectively, defining directionality. Furthermore, unless otherwise specified, "upstream" and "downstream" in this specification refer to the direction of transcription of the endogenous gene, respectively, representing the upstream and downstream directions of the gene.
[0104] 1-3. Composition
[0105] The recombinant lepidopteran insect of the present invention contains a target gene sequence encoding a target protein or a fragment thereof in the exon sequence of a signal peptide or functional fragment thereof encoding an endogenous gene. The target gene sequence is included in the above-mentioned exon sequence in such a way that the target protein or a fragment thereof is fused to the C-terminus of the signal peptide or functional fragment thereof.
[0106] The term "exon sequence encoding a signal peptide or a functional fragment thereof of an endogenous gene" (hereinafter often referred to as "target exon sequence") used in this specification is not limited to any exon sequence encoding a signal peptide in an endogenous gene. Typically, in an endogenous gene, the signal peptide is encoded by the first exon, or a sequence of multiple exons containing the first exon, located upstream in the mRNA transcribed from the endogenous gene. However, the target exon sequence can be any exon sequence. For example, the target exon sequence can be exon 1, exon 2, exon 3, or exon 4. The target exon sequence can, for example, be an exon encoding a C-terminal amino acid residue of the signal peptide, or an exon adjacent to its 5' end.
[0107] In the recombinant lepidopteran insects of the present invention, the target gene sequence is inserted into the target exon sequence in such a manner that the target protein or a fragment thereof encoded by the target gene sequence is fused to the C-terminus of the signal peptide or a functional fragment thereof of the endogenous gene. More specifically, the target gene sequence is inframed within the target exon sequence of the endogenous gene and attached to the 3' end of the base sequence encoding the signal peptide or a functional fragment thereof. Thus, the N-terminus of the target protein or a fragment thereof is fused to the C-terminus of the signal peptide or a functional fragment thereof of the endogenous gene, and a fusion gene encoding a fusion polypeptide comprising the signal peptide or a functional fragment thereof of the endogenous gene and the target protein or a fragment thereof is constructed in the locus of the endogenous gene. Furthermore, in the fusion polypeptide, the signal peptide or a functional fragment thereof may be directly linked to the target protein or a fragment thereof, or an amino acid sequence other than the signal peptide encoded by the target exon sequence (e.g., the N-terminal amino acid sequence in the mature protein described later) may be inserted between them.
[0108] In one embodiment, an endogenous gene encodes proteins that constitute the filament. The proteins constituting the filament are not particularly limited and may be, for example, sericin, sericin, and / or fibrohexamer proteins. Sericin may be a sericin H chain and / or a sericin L chain. Sericin is not particularly limited and may be, for example, sericin 1.
[0109] In the silken H chain of silkworm fibroin, the precursor protein containing the signal peptide consists of the amino acid sequence shown in Serial No. 1, while the mature protein, excluding the signal peptide, consists of the amino acid sequence shown in Serial No. 2. The signal peptide of the silken H chain in Serial No. 1 consists of amino acid sequences from position 1 to 21.
[0110] Regarding the silk core protein H chain gene in silkworms, the signal peptide is encoded by exons 1 and 2, and the C-terminal amino acid residue of the signal peptide is encoded by exon 2. Figure 2 B). In the genomic sequence shown in sequence number 3 of the filoprotein H chain gene, exon 1 is from position 1001 to 1042, intron 1 is from position 1043 to 2013, exon 2 contains from position 2014 to at least position 17763, and the region encoding the signal peptide in exon 2 is from position 2014 to 2034.
[0111] Regarding the silkworm filament protein L-chain, the precursor protein containing the signal peptide consists of the amino acid sequence shown in SEQ ID NO. 4, while the mature protein, excluding the signal peptide, consists of the amino acid sequence shown in SEQ ID NO. 5. The signal peptide of the filament protein L-chain consists of amino acids from position 1 to position 16 in SEQ ID NO. 4.
[0112] Regarding the silk core protein L-chain gene in silkworms, the signal peptide is encoded by exons 1, 2, and 3, and the C-terminal amino acid residue of the signal peptide is encoded by exon 3. Figure 2 C). In the genomic sequence shown in sequence number 6, the first exon of the filoprotein L chain gene is located at positions 574–889, the first intron at positions 890–966, the second exon at positions 967–1036, the second intron at positions 1037–8976, and the third exon at positions 8977–9059. The region encoding the signal peptide in the third exon is located at positions 8977–8988.
[0113] Silkworm sericin 1 generates multiple isoforms through alternative splicing. In one example of an isoform, the precursor protein containing the signal peptide consists of the amino acid sequence shown in Serial No. 7, while the mature protein containing the signal peptide consists of the amino acid sequence shown in Serial No. 8. In the aforementioned isoforms of sericin 1, the signal peptide in Serial No. 7 consists of amino acid sequences from position 1 to position 19.
[0114] Regarding the silkworm sericin 1 gene, the signal peptide of the aforementioned isotype is encoded by exon 1 and exon 2, and the C-terminal amino acid residue of the signal peptide is encoded by exon 2. Figure 2A). In the genomic sequence shown in sequence number 9, the first exon of the above isotype in the sericin 1 gene is located at positions 947–1039, the first intron is located at positions 1040–3051, and the second exon is located at positions 3052–3082. The region encoding the signal peptide in the second exon is located at positions 3052–3069.
[0115] Regarding the silkworm fibrohexamethylenetetramer protein, the precursor protein containing the signal peptide consists of the amino acid sequence shown in SEQ ID NO. 10, while the mature protein, excluding the signal peptide, consists of the amino acid sequence shown in SEQ ID NO. 11. The signal peptide of the fibrohexamethylenetetramer protein consists of amino acids from position 1 to position 17 of SEQ ID NO. 10.
[0116] Regarding the silkworm fibrohexamer protein gene, the signal peptide is encoded by exon 1, and the C-terminal amino acid residues of the signal peptide are also encoded by exon 1. In the genome sequence shown in sequence number 12, exon 1 of the fibrohexamer protein gene is located at positions 918–1052, intron 1 at positions 1053–1536, and exon 2 at positions 1537–1756. The region encoding the signal peptide in exon 1 is located at positions 1001–1051.
[0117] In the recombinant lepidopteran insects of the present invention, the target gene sequence can be introduced into a single endogenous gene or into multiple endogenous genes. Furthermore, the recombinant lepidopteran insects of the present invention can heterozygously possess exon sequences containing the target gene sequence or can homozygously possess exon sequences containing the target gene sequence. When the target gene sequence is introduced into multiple endogenous genes, the types of target gene sequences introduced into the multiple endogenous genes can be the same or different.
[0118] In one embodiment, the multiple endogenous genes into which the target gene sequence is introduced may be genes encoding the filamentin H chain and the filamentin L chain, genes encoding the filamentin H chain and sericin 1, or genes encoding the filamentin H chain, the filamentin L chain and the sericin 1.
[0119] In one embodiment, in the recombinant lepidopteran insect of the present invention, the target gene sequence has a stop codon. In a further embodiment, in the recombinant lepidopteran insect of the present invention, the target exon sequence contains a transcription termination sequence at the 3' end of the stop codon of the target gene sequence.
[0120] In other embodiments, the target gene sequence does not have a stop codon. In this embodiment, the target protein or a fragment thereof is fused between a signal peptide or a functional fragment thereof, and a mature protein encoded by an endogenous gene (a protein from which the signal peptide has been cleaved from a precursor protein) or its C-terminal fragment (e.g., a fragment in the mature protein where a portion of the N-terminal sequence encoded by the target exon sequence is deleted). As a result, the N-terminus of the target protein or a fragment thereof is fused to the C-terminal side of the signal peptide or functional fragment thereof of the endogenous gene, and the C-terminus of the target protein or a fragment thereof is fused to the N-terminus of the mature protein or its C-terminal fragment thereof. Therefore, a fusion gene encoding a fusion protein comprising the signal peptide or functional fragment thereof of the endogenous gene, the target protein or a fragment thereof, and the mature protein or its C-terminal fragment thereof is constructed at the endogenous gene locus.
[0121] In one embodiment, the target protein is a fluorescent protein, antibody, antigenic peptide, enzyme, cytokine, or antimicrobial peptide. For example, when the target protein is an antibody, the heavy chain gene and light chain gene constituting the antibody can be introduced into different endogenous genes.
[0122] 1-4. Effects
[0123] The recombinant lepidopteran insects of this invention can stably and massively produce the target protein encoded by the target gene introduced into the target exon sequence.
[0124] In existing GAL4 / UAS systems, the expression levels of the target protein can vary significantly due to the random placement of the GAL4 gene and UAS control sequence on the genome. However, in the recombinant lepidopteran insects of this invention, the promoter and enhancer activities of the endogenous gene can be directly utilized to express the target gene, thus enabling more precise control over the expression level.
[0125] 2. Double-stranded circular DNA
[0126] 2-1. Overview
[0127] The second variant of this invention is a double-stranded circular DNA. Based on this variant of the double-stranded circular DNA, a target gene sequence can be introduced into an endogenous gene in a lepidopteran insect within a target exon located at the 3' end of a genome cleavage site within an intron sequence. This variant of the double-stranded circular DNA can, for example, be used for knock-in of target genes based on the TAL-PITCh (precise integration into target chromosome) method.
[0128] 2-2. Definition
[0129] In this context, "double-stranded circular DNA" refers to a circular double-stranded DNA molecule that contains at least the target gene sequence for the purpose of introducing the target gene sequence into the endogenous genes of lepidopteran insects. The double-stranded circular DNA is preferably a vector capable of being maintained and / or replicated within bacterial cells such as *E. coli*, and may contain sequences necessary for maintenance and replication within the cell (such as origin of replication and / or genes encoding antibiotic resistance proteins). The double-stranded circular DNA can be, for example, a plasmid vector.
[0130] The term "genome editing" as used in this manual refers to gene-targeting technology that utilizes DNA repair mechanisms, including double-strand break (DSB) cleavage by DNA cutting enzymes, to insert (knock in) foreign genes or destroy (knock out) target genes at any location on the genome. Zinc finger nuclease (ZFN) methods, TALEN methods, and CRISPR / Cas methods are known in genome editing technology, but any method may be used in this manual.
[0131] The TALEN (Transcription Activator-Like Effector Nuclease) method is a genome editing technology that utilizes an artificial DNA-cutting enzyme fused with a TAL effector (TALE) protein derived from the plant pathogenic bacterium Xanthomonas, forming a non-specific endonuclease domain. TALEN is a protein composed of a repeating TALE domain containing a DNA-binding unit and a non-specific endonuclease domain such as the FokI nuclease domain. The nuclease domain, which has enzymatic activity for cutting DNA, functions as a dimer. Therefore, TALEN functions as a dimer consisting of a polypeptide recognizing the DNA sequence near the upstream (5' side) of the double-strand cleavage (DSB) site in the target sequence (often referred to as "Left-TALEN" in this specification) and a polypeptide recognizing the DNA sequence near the downstream (3' side) of the DSB site (often referred to as "Right-TALEN" in this specification). The DNA-binding unit constituting the TALE domain has mutations at amino acid residues 12 and 13, starting from the N-terminus, enabling it to specifically recognize the four DNA bases (A: adenine, G: guanine, C: cytosine, T: thymine) in groups of two amino acids. For example, adenine is recognized when the amino acid residues at positions 12-13 are NI or NN, guanine when NN, cytosine when HD, and thymine when NG. The number of repeats of the DNA-binding unit can vary depending on the base length of the target base sequence. By manipulating the TALE domain, gene targeting can be achieved using any DNA sequence in the genome. Gene knockout using the TALEN method in lepidopteran insects such as silkworms is a well-known technique. For example, the method described in Takasu Y., et al., 2013, PLoS One 8, e73458 can be referenced.
[0132] Zinc-finger nucleases (ZFNs) are a genome editing technique that uses artificial DNA-cutting enzymes composed of a zinc finger domain (which acts as a DNA-binding domain) and a non-specific endonuclease domain such as FokI. Each zinc finger motif recognizes three bases and can bind to the target nucleic acid. By linking multiple zinc finger motifs, binding can be achieved by specifically recognizing and binding to a multiple of three bases. Functioning as a dimer, after binding to the target site, the enzyme performs double-strand cleavage (DSB) at a specific location within the target nucleic acid through endonuclease activity.
[0133] The "CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats / CRISPR-associated proteins) method" utilizes a genome editing technology evolved in bacteria and archaea to acquire an immune system in order to exclude foreign DNA or RNA such as viruses and plasmids. This report describes the CRISPR / Cas9 method using the Cas9 protein, as well as methods utilizing other Cas proteins such as Cpf1 and Cas13a. Bacteria and archaea fragment invading foreign DNA or RNA and insert it into a CRISPR region in their genome, using this as a template to synthesize approximately 40 bp of CRISPR RNA (crRNA). The crRNA binds directly or via trans-activating RNA (tracrRNA) to a Cas protein with nuclease activity, forming a CRISPR / Cas complex. The CRISPR / Cas complex then binds to and cleaves a target DNA or RNA sequence with a complementary base sequence via the crRNA. When using double-stranded nucleases such as Cas9 and Cpf1 as Cas proteins, DSB is induced at the target site.
[0134] The term "genome editing enzyme" in this specification refers to proteins that have the activity of specifically cutting and editing target sites on the genome. Examples of genome editing proteins include TALEN (transcription activator-like effector nuclease), Cas9 (CRISPR-associated protein 9), and ZFN (zinc finger nuclease), which can be used in the aforementioned genome editing processes. When the genome editing protein is a TALEN, Left TALEN and Right TALEN can be used as TALENs that function as dimers. When the genome editing protein is Cas9, guide RNAs such as crRNA are required for genome editing.
[0135] 2-3. Composition
[0136] This type of double-stranded circular DNA sequentially comprises a first recognition sequence, a second spacer sequence, a first spacer sequence, a genomic homologous sequence, and a target gene sequence. The first recognition sequence, the second spacer sequence, the first spacer sequence, and the genomic homologous sequence are derived from the base sequences of a genomic region containing an endogenous gene that serves as the target for the introduced gene sequence. Furthermore, this type of double-stranded circular DNA includes a second recognition sequence at the 5' end of the genomic homologous sequence.
[0137] Regarding the introduction of the target gene using this type of double-stranded circular DNA, two genome editing enzymes (hereinafter referred to as "first genome editing enzyme" and "second genome editing enzyme") are used to generate a double-stranded cleavage site within an intron sequence of the genome of the target endogenous gene. In this description, this double-stranded cleavage site is referred to as the "genome cleavage site". The genome cleavage site can be set at any position as long as it is within an intron sequence of the endogenous gene, but it is preferred to be located outside of functional sequences such as splice donor sequences, splice acceptor sequences, and branch points required for splicing. Furthermore, depending on the type of genome editing enzyme, the genome cleavage site may not be accurately specified. In such cases, when designing the double-stranded circular DNA of the present invention, the sequences of each element constituting the double-stranded circular DNA are specified based on the assumed position as the genome cleavage site. The actual cleavage positions of the two genome editing enzymes in cutting the genome and the double-stranded circular DNA are not limited to this position, but can also be at nearby positions (e.g., any position in the first spacer sequence and / or the second spacer sequence).
[0138] The first recognition sequence, the second spacer sequence, and the second recognition sequence located at the 5' end in the genomic homologous sequence of this double-stranded circular DNA are derived from the base sequence near the genomic cleavage site in the endogenous gene's genomic sequence. The "first recognition sequence" and the "second recognition sequence" are located at the 5' and 3' ends of the genomic cleavage site in the endogenous gene's genomic sequence, respectively, and are identical to the base sequences recognized and bound by the first and second genome editing enzymes, respectively. The base lengths of the first and second recognition sequences vary depending on the type of genome editing enzyme, typically ranging from 8 to 30 bases, for example, 10 to 25 bases, 12 to 20 bases, or 14 to 18 bases. The "first spacer sequence" originates from the sequence adjacent to the 5' end of the genomic cleavage site in the endogenous gene's genomic sequence, derived from the base sequence located between the aforementioned first recognition sequence and the genomic cleavage site. The "second spacer sequence" originates from the sequence adjacent to the 3' end of the genome cleavage site in the endogenous gene's genomic sequence, specifically from the base sequence located between the genomic cleavage site and the aforementioned second recognition sequence. The base lengths of the first and second spacer sequences vary depending on the type of genome editing enzyme, typically ranging from 6 to 30 bases, for example, 8 to 25 bases, 10 to 20 bases, or 12 to 15 bases. A key feature here is that the endogenous gene, starting from the upstream side, sequentially contains the first recognition sequence, the first spacer sequence, the second spacer sequence, and the second recognition sequence; conversely, in this type of double-stranded circular DNA, the arrangement of the first and second spacer sequences is reversed.
[0139] The double-stranded circular DNA of this type contains a genomic homologous sequence with the aforementioned second recognition sequence at its 5' end, consisting of a base sequence homologous to the genomic sequence from the second recognition sequence to the exon sequence or a portion thereof located at the 3' end of the intron sequence containing the genomic cleavage site. Here, the exon sequence or a portion thereof located at the 3' end of the intron sequence containing the genomic cleavage site can be an exon sequence adjacent to the 3' end of the intron sequence containing the genomic cleavage site. In this type of double-stranded circular DNA, the exon sequence or a portion thereof located at the 3' end of the genomic homologous sequence, encoding a signal peptide or its functional fragment, is inframed with the target gene sequence located further 3' to its end. In this type of double-stranded circular DNA, the genomic homologous sequence is not particularly limited as long as it contains the second recognition sequence of the genome editing enzyme. The base length of genomic homologous sequences can be, for example, 15 bases to 20,000 bases, 20 bases to 10,000 bases, 50 bases to 5,000 bases, 100 bases to 2,000 bases, or 500 bases to 1,000 bases.
[0140] In one embodiment, the region from the first recognition sequence to the second recognition sequence in the genomic homologous sequence in this double-stranded circular DNA is composed of the first recognition sequence, the second spacer sequence, the first spacer sequence, and the second recognition sequence in the genomic homologous sequence.
[0141] In one embodiment, the target gene sequence in the double-stranded circular DNA of this type contains a stop codon. In a further embodiment, the double-stranded circular DNA of this type contains a transcription termination sequence at the 3' end of the target gene sequence.
[0142] In one embodiment, the double-stranded circular DNA of this type contains a marker gene for identifying an individual in which the target gene sequence has been introduced into an endogenous gene. For example, the marker gene may be positioned 3' further to the transcription termination sequence located at the 3' end of the target gene sequence.
[0143] The type of genome editing enzyme that recognizes the first and second recognition sequences contained in the double-stranded circular DNA of this type is not limited and can be TALEN, ZFN, and / or Cas9. For example, the genome editing enzyme that recognizes the first and second recognition sequences can be TALEN. In this case, the two genome editing enzymes that recognize the first and second recognition sequences can be Left-TALEN and Right-TALEN, which function as dimers.
[0144] In one embodiment, the endogenous gene in this pattern contains multiple exons of exon 1 encoding signal peptides.
[0145] 2-4. Effects
[0146] By introducing this type of double-stranded circular DNA along with the first and second genome editing enzymes into the eggs of lepidopteran insects, end rejoining (microhomology-mediated end-joining) can be achieved based on the small homology between the first spacer sequence adjacent to the genome cut site on the genome and the first spacer sequence in the double-stranded circular DNA. This allows the insertion of the target gene sequence into the target exon sequence located at the 3' end of the genome cut site within the intron sequence of the lepidopteran insect's endogenous gene.
[0147] In one embodiment, the double-stranded circular DNA of this sample can be used in the TAL-PITCh method. Furthermore, for information on the TAL-PITCh method, please refer to the well-known technical literature (Nature communications, 2014, 5:5560).
[0148] 3. Donor nucleic acid
[0149] 3-1. Overview
[0150] The third aspect of this invention is a donor nucleic acid. Based on this donor nucleic acid, a target gene sequence can be introduced into a target exon sequence located at the 3' or 5' end of a genome cleavage site within an intron sequence of an endogenous gene in lepidopteran insects. This donor nucleic acid can be used, for example, in target gene knock-in based on homologous recombination.
[0151] 3-2. Composition
[0152] The term "donor nucleic acid" in this specification refers to nucleic acid used to introduce the target gene sequence into the endogenous genes of lepidopteran insects. The method of using donor nucleic acid is not limited; it can be, for example, double-stranded circular DNA or straight-stranded DNA such as plasmid vectors.
[0153] The donor nucleic acid in this sample contains a first-genomic homologous sequence and a second-genomic homologous sequence, as well as a target gene sequence positioned between them. The first-genomic homologous sequence and the second-genomic homologous sequence are derived from the base sequences of genomic regions containing endogenous genes that become the targets for the introduced target gene sequence.
[0154] In the introduction of the target gene using the donor nucleic acid of this type, a genome editing enzyme that recognizes sequences in the vicinity of the target endogenous gene is used to generate a double-stranded cleavage site within an intron sequence of the genome. Similar to the second type, this double-stranded cleavage site is also referred to as the "genome cleavage site" in this type. Furthermore, as mentioned above, the genome cleavage site cannot always be accurately specified depending on the type of genome editing enzyme. However, in such cases, when designing the donor nucleic acid of this invention, the sequences constituting the donor nucleic acid are specified based on the assumed location as the genome cleavage site. The actual location where the genome editing enzyme cuts the genome is not limited to this location; it can also be a nearby location. The genome cleavage site can be set at any location as long as it is within an intron sequence of the endogenous gene, but it is preferably located outside of the splice donor sequence, splice acceptor sequence, branch point, and other sequences required for splicing. Additionally, the sequence recognized by the genome editing enzyme in this type is referred to as the "recognition sequence of the genome editing enzyme" or simply as the "recognition sequence."
[0155] The specific composition of the first and second genomic homologous sequences contained in the donor nucleic acid of this type differs depending on whether the genome cut position is located at the 5' end of the target exon sequence into which the target gene sequence is introduced or at the 3' end of the genome cut position. These differences will be explained separately below.
[0156] (1) Implementation method where the genome cleavage site is located at the 5' end of the target exon sequence
[0157] In an embodiment where the genome cleavage site is located at the 5' end of the target exon sequence, the first genomic homologous sequence consists of a base sequence homologous to the genome sequence of the target exon sequence or a portion thereof, extending from the bases located at the 5' end of the genome cleavage site (e.g., upstream of the cleavage site at bases of 10, 20, 50, 100, 500, or 1,000 bases or more) to the 3' end (or adjacent to) of the intron sequence containing the genome cleavage site. In this embodiment, the genome sequence corresponding to the first genomic homologous sequence has a mutation in its recognition sequence in a manner that prevents it from being cleaved by a genome editing enzyme that cuts the aforementioned genome cleavage site. The type of mutation is not particularly limited and can be, for example, a substitution, deletion, and / or insertion of bases within the recognition sequence. The number of mutated bases within the recognition sequence is not particularly limited. For example, one or more bases may be substituted, deleted, and / or inserted; more specifically, one or more bases may be substituted, deleted, and / or inserted, or more than two, three, four, five, or six bases may be substituted, deleted, and / or inserted.
[0158] In this embodiment, the second genomic homologous sequence consists of a base sequence that is homologous to the genomic sequence located at the 3' end of the target sequence or a portion thereof.
[0159] In this embodiment, the base length of the homologous sequences of the first and second genomes is not particularly limited, and can be, for example, 100 bases to 20,000 bases, 200 bases to 10,000 bases, 500 bases to 5,000 bases, or 1,000 bases to 2,000 bases.
[0160] (2) Implementation method where the genome cleavage site is located at the 3' end of the target exon sequence
[0161] In an embodiment where the genome cut site is located at the 3' end of the target exon sequence, the first genomic homologous sequence consists of a base sequence homologous to the genome sequence from the bases located at the 5' end of the intron sequence containing the genome cut site (specifically, the bases at the 5' end of the target exon sequence into which the target gene sequence is inserted, for example, 100, 200, 500, 1,000, 5,000, or more than 10,000 bases upstream of the insertion site), to the target exon sequence or a portion thereof located at (or adjacent to) the 5' end of the intron sequence (e.g., to the bases to the C-terminal amino acid residues encoding the signal peptide or a functional fragment thereof in the target exon sequence, or to the bases further C-terminus amino acid residues encoding the C-terminal amino acid residues of the signal peptide or a functional fragment thereof in the target exon sequence).
[0162] In this embodiment, the second genomic homologous sequence consists of a base sequence homologous to the genomic sequence located at the 3' end of the target exon sequence or a portion thereof and at the 5' end of the genomic cleavage site (e.g., bases further C-terminus of the C-terminal amino acid residues encoding the signal peptide or its functional fragment in the target exon sequence), up to the 3' end of the genomic cleavage site (e.g., bases downstream of the genomic cleavage site of 10, 20, 50, 100, 500, 1,000, 5,000, or more than 10,000 bases). In this embodiment, the genomic sequence corresponding to the second genomic homologous sequence has a mutation in its recognition sequence in a manner that is not subject to cleavage by the genome editing enzyme that cuts the aforementioned genomic cleavage site. The type and number of mutated bases are not particularly limited, and similarly to (1) above, can be, for example, substitution, deletion, and / or insertion of one or more bases within the recognition sequence.
[0163] In this embodiment, the base length of the homologous sequences of the first and second genomes is not particularly limited, and can be, for example, 100 bases to 20,000 bases, 200 bases to 10,000 bases, 500 bases to 5,000 bases, or 1,000 bases to 2,000 bases.
[0164] In the donor nucleic acid of this type, the target exon sequence or a portion thereof contained in either the first or second genomic homologous sequence, encoding a signal peptide or a functional fragment thereof or a sequence on its C-terminus, is inframed with the target gene sequence located on its 3' end via a base sequence encoding multiple amino acid residues, as appropriate.
[0165] In one embodiment, the target gene sequence in the donor nucleic acid of this type contains a stop codon. In a further embodiment, the donor nucleic acid of this type contains a transcription termination sequence at the 3' end of the target gene sequence.
[0166] In one embodiment, the donor nucleic acid of this type contains a marker gene for identifying recombinant lepidopteran insects in which the target gene sequence has been introduced into the endogenous gene.
[0167] The type of genome editing enzyme used in homologous recombination of donor nucleic acids of this type is not limited and can be TALEN, ZFN, and / or Cas9.
[0168] In one embodiment, the donor nucleic acid of this type contains a nuclease recognition sequence at its end opposite to the target gene sequence on the side of the homologous sequence to the first genome and / or the homologous sequence to the second genome. The nuclease recognition sequence is not particularly limited and can be a recognition sequence of a genome editing enzyme that cuts the aforementioned genome cutting site, or a restriction endonuclease recognition sequence that can be cleaved by any restriction endonuclease different from that genome editing enzyme. If the nuclease recognition sequence is a TALEN recognition sequence, it can be a combination of two recognition sequences recognized by Left TALEN and Right TALEN.
[0169] 3-3. Effects
[0170] By introducing donor nucleic acid of this type along with genome editing enzyme into the eggs of lepidopteran insects, homologous recombination can be induced in the endogenous genes of lepidopteran insects, inserting the target gene sequence into the target exon sequence.
[0171] 4. Methods for producing recombinant lepidopteran insects
[0172] 4-1. Overview
[0173] The fourth aspect of this invention is a method for producing a recombinant lepidopteran insect. This method involves introducing the double-stranded circular DNA described in the second aspect or the donor nucleic acid described in the third aspect into the eggs of a lepidopteran insect via microinjection, thereby enabling the introduction of the target gene sequence into the exon sequence of the endogenous gene, thus producing a recombinant lepidopteran insect.
[0174] 4-2. Methods
[0175] The method for preparing this specimen includes, as a necessary step, the introduction of double-stranded circular DNA or donor nucleic acid into the eggs of lepidopteran insects using microinjection, and as optional steps, the egg acquisition step and the selection step for gene recombination lepidopteran insects. The composition of each step is described below.
[0176] (1) Egg retrieval process
[0177] The "egg acquisition process" refers to the process of obtaining eggs by inducing the female parent silkworm to lay eggs. The egg acquisition method can be carried out using conventional methods in this field. An oviposition paper is given to the mated female parent silkworm to induce oviposition. The oviposition temperature is 23–28°C, preferably around 25°C. Typically, the female silkworm begins oviposition several hours after mating. To ensure that the DNA introduced into the egg enters the nucleus, microinjection is performed 2–8 hours after egg laying, preferably 3–6 hours.
[0178] (2) Introduction process
[0179] The so-called "introduction process" is the process of introducing double-stranded circular DNA or donor nucleic acid into the eggs of lepidopteran insects using microinjection.
[0180] In the embodiment of the preparation method of this sample, where the double-stranded circular DNA is introduced into the egg in the introduction step, the composition of the double-stranded circular DNA is as described in the second sample. In the introduction step of this embodiment, the double-stranded circular DNA described in the second sample, the first genome editing enzyme described in the second sample or the nucleic acid encoding the first genome editing enzyme in an expressible state, and the second genome editing enzyme described in the second sample or the nucleic acid encoding the second genome editing enzyme in an expressible state are introduced into the egg of a lepidopteran insect using a microinjection method.
[0181] In the method for preparing this sample, the donor nucleic acid is introduced into the egg via the introduction step, and the composition of the donor nucleic acid is as described in the third sample. In the introduction step of this embodiment, the donor nucleic acid described in the third sample, the genome editing enzyme described in the third sample, or a nucleic acid encoding a genome editing enzyme in an expressible state is introduced into the egg of a lepidopteran insect using microinjection.
[0182] In this specification, "expressible state" refers to the configuration of the gene to be expressed in the downstream region of the promoter, under the control of the promoter. Furthermore, the nucleic acid encoding the genome editing enzyme in the expressible state can be RNA such as mRNA, or DNA such as plasmid DNA or linear DNA. In addition to the base sequence encoding the first or second genome editing enzyme, the DNA encoding the genome editing enzyme in the expressible state also contains a promoter that can be expressed in the eggs of lepidopteran insects, and may include, as needed, marker genes (selection markers), enhancers, terminators, origins of replication, and poly-A signals.
[0183] Microinjection can be performed using methods known in the field. For example, an injection solution is prepared by dissolving or diluting double-stranded circular DNA or donor nucleic acid, and genome editing enzymes or nucleic acids encoding genome editing enzymes in an expressible state, at appropriate concentrations, using solvents such as water or buffer solutions. Then, fertilized eggs laid 3–6 hours prior are microinjected. The amount of nucleic acid introduced is not particularly limited and can be appropriately determined based on the type, properties, and purpose of the nucleic acid. Typically, 50 nL–30 nL is sufficient. The introduced lepidopteran eggs can be incubated under appropriate conditions, such as at 25°C, until hatching.
[0184] (3) Selection process for Lepidoptera insects undergoing gene recombination
[0185] The so-called "selection process for recombinant lepidopteran insects" refers to the process of selecting recombinant lepidopteran insects from hatched lepidopteran insects. This process can also be carried out using methods known in the field. For example, if the double-stranded circular DNA or donor nucleic acid used in the introduction process contains a marker gene, the target recombinant lepidopteran insects can be easily screened based on the expression of the marker gene.
[0186] The term "marker gene" in this specification refers to a polynucleotide consisting of a base sequence that encodes a marker protein known as a screening marker.
[0187] The term "marked protein" in this specification refers to a protein that can confer new traits not found in the host lepidopteran insect through the expression of a marker gene. These proteins include enzymes, fluorescent proteins, pigment-synthesizing proteins, or luminescent proteins. Based on the activity of the marked protein, transformants retaining the introduced nucleic acid can be easily identified.
[0188] 5. Methods for producing target proteins or fusion proteins
[0189] 5-1. Summary
[0190] The fifth aspect of the invention is a method for producing a target protein or a fragment thereof, or a fusion protein comprising them. According to the production method of this aspect, the target protein or a fragment thereof, or a fusion protein comprising them, can be mass-produced using recombinant lepidopteran insects of the first aspect or recombinant lepidopteran insects produced by the method described in the fourth aspect.
[0191] 5-2. Production Method
[0192] The production method of the present invention includes a feeding process and a recycling process. The following describes each process.
[0193] (1) Feeding process
[0194] The term "breeding process" refers to the process of raising genetically recombinant lepidopteran insects of type 1 or those produced using the methods described in type 4. Regarding the breeding methods for genetically recombinant lepidopteran insects, various lepidopteran insects can be raised using techniques known in the field. For example, if the lepidopteran insect is a silkworm, one can refer to "General Introduction to Silkworm Breeds; by Takeo Takami, published by the National Silkworm Breeding Association." Regarding feed, for example, for silkworms or wild mulberry silkworms, it can be leaves of the genus *Morus*; for castor silkworms, it can be leaves of castor beans (*Ricinus communis*) or ailanthus saltissima; for tussah silkworms, it can be natural leaves from trees belonging to the family Fagaceae, which are used by herbivores; it can also be... Artificial feeds such as L4M or those for 1st-3rd instar silkworm eggs (Nippon Nippon Kogyo) are preferred if considering factors such as disease suppression, stable quality and quantity of feeding, and the ability to maintain sterile rearing as needed. The following section uses silkworms as an example to illustrate a simple rearing method.
[0195] The process involves cleaning eggs obtained from the oviposition of a suitable number (e.g., 4–10) of genetically recombinant lepidopteran insects from the same lineage of females. After hatching, the larvae are transferred from the egg tray to a container lined with parchment paper (paraffin-processed paper) used as a drying tray. Artificial feed should be placed on absorbent paper before feeding. Feed should be changed once each for the first and second instars, and 1-3 times for the third instar. Uneaten feed should be removed to prevent spoilage. Strong silkworm larvae of the fourth and fifth instars should be moved to larger containers, with the number of larvae per container adjusted as needed. Depending on humidity and the condition of the containers, absorbent paper, acrylic materials, or mesh lids can be added. The rearing temperature should be maintained between 25-28℃ throughout the entire instar range.
[0196] (2) Recycling process
[0197] The so-called "recovery process" refers to the process of recovering the target protein or its fragments, or the fusion protein containing them, from the silk glands of recombinant lepidopteran insect larvae after they have been expressed, secreted, and accumulated in the lumen of the silk glands.
[0198] In this model, the recombinant lepidopteran insects express a target protein or a fragment thereof within their silk gland cells, or a precursor protein (hereinafter referred to as "target protein, etc.") containing a signal peptide or a functional fragment thereof fused to the N-terminus of a fusion protein (hereinafter referred to as "target protein, etc.") encoded by an endogenous gene at its C-terminus. The precursor protein expressed within the silk gland cells is transported to the endoplasmic reticulum by the action of the signal peptide or its functional fragment. After the signal peptide or its functional fragment is cleaved by enzymes such as peptidases within the endoplasmic reticulum, it is secreted into the lumen of the silk gland. The target protein, etc., after the signal peptide or its functional fragment has been cleaved, is secreted from the anterior silk gland and spun out of the individual during pupation. Therefore, methods for recovering the target protein, etc., include methods of recovery from the cocoon, or methods of direct recovery by removing the silk gland from the insect during the late instar to prepupal stage. In particular, the method of recovery from the cocoon is superior in terms of its ease of recovery of the target protein, etc.
[0199] The method for recovering target proteins from cocoons involves first transferring late-stage larvae to a cocoon (the larvae are moved to a cocoon) and causing them to spin cocoons. Then, the target proteins are extracted from the cocoons. The extraction method is not particularly limited. For example, the target proteins can be recovered simply by soaking the cocoons in water or a suitable neutral extraction buffer (e.g., phosphate-buffered saline with or without 1% Tween-20 and 0.05% sodium azide, pH 7.2) that does not contain protein denaturants. To improve extraction efficiency, the cocoons can be cut or pulverized before soaking. As for the extraction temperature, to prevent heat denaturation of the target proteins, it is performed at a low temperature of 0–10°C, preferably 0–5°C. However, if the target proteins are not heat-sensitive peptides, it can be performed at 10–40°C. The extraction solution can be stirred as needed. The extraction time varies depending on the state of the cocoons (e.g., uncuttered or powdered), the volume of the extraction solution, the extraction temperature, and whether stirring is present or absent; therefore, it should be set appropriately according to the conditions. Insoluble components such as fibroin can be removed from the extract by centrifugation or filtration as needed.
[0200] Methods for extracting silk glands and recovering target proteins from silkworms in the late final instar to prepupal stage can be achieved using methods known in the field. For example, silkworms that are about to spin silk on the 6th day of their final instar (5th instar) can be anesthetized on ice, and their silk glands can be extracted by making a lateral incision and using forceps without damaging them (see Mori Yasushi, ed.). (New Biological Experiments Using Silkworms), Sanseido, 1970, pp. 249-255). The extracted silk glands can be slowly agitated in the extraction buffer at a temperature of 0–10°C, preferably 0–5°C, to dissolve the target protein, etc., in the buffer. If the target protein, etc., is not a heat-sensitive peptide, it can also be performed at a temperature of 10–40°C. Then, impurities such as tissue fragments can be removed by centrifugation or filtration, and the supernatant containing the target protein, etc., can be recovered.
[0201] 5-3. Effects
[0202] According to the production method of the present invention, by using the larvae of recombinant lepidopteran insects prepared using the method described in the first type or the fourth type as a protein production system, it is possible to produce a large quantity of the target protein, etc., compared with the case of using the GAL4 / UAS system, and it can also be easily recycled.
[0203] 6. Fusion proteins
[0204] The sixth embodiment of the present invention is a fusion protein. This fusion protein, starting from the N-terminus, sequentially comprises the target protein or a fragment thereof, and any protein constituting a filament. The filament-constituting proteins in this fusion protein can be sericin, sericin, or fibrohexamer protein. Furthermore, the sericin can be a sericin H chain and / or a sericin L chain. Additionally, the filament-constituting proteins in this fusion protein can be full-length mature proteins.
[0205] The target protein or its fragments contained in the fusion protein of this type are not particularly limited. For example, it can be a fluorescent protein, antibody, antigenic peptide, enzyme, cytokine, or antimicrobial peptide.
[0206] 7. Cocoon or silk
[0207] The seventh aspect of the present invention is a cocoon or silk. The cocoon or silk of this aspect contains the fusion protein of the sixth aspect. In a further embodiment, in the cocoon or silk of this aspect, any protein constituting the silk is composed of the fusion protein of the sixth aspect. For example, any one or more proteins selected from sericin H chain, sericin L chain, sericin 1, sericin 2, sericin 3, and fibrohexamer protein contained in the cocoon or silk of this aspect are composed of the fusion protein described in the sixth aspect. For example, when the sericin H chain is composed of the fusion protein described in the sixth aspect, in the cocoon or silk of this aspect, all the sericin H chains become the fusion protein described in the sixth aspect, and substantially do not contain any protein that has not been fused with the target protein.
[0208] In one embodiment, the cocoon or silk of this type originates from a recombinant lepidopteran insect of type 1 or a recombinant lepidopteran insect produced by the method described in type 4. This recombinant lepidopteran insect may be a homozygous recombinant lepidopteran insect containing an exon sequence with the target gene sequence.
[0209] In cases where the fusion protein contained in the cocoon or silk of this type is a fluorescent protein fused with fibroin, fibrous hexamer, or similar proteins, the cocoon or silk exhibits extremely strong fluorescence, for example, a very strong color that is visible to the naked eye under normal white light. This color is overwhelmingly stronger than the fluorescent cocoons or filaments that can be produced using existing technologies, thus providing a material with high utilization value.
[0210] Example
[0211] <Example 1: Fabrication of a knock-in system for useful protein production>
[0212] (Purpose)
[0213] The existing GAL4 / UAS system is a gene control system that combines the yeast transcription factor GAL4 and the control sequence UAS. For the GAL4 / UAS system used as a protein production system in silkworms, an expression system for expressing the target protein in silk glands is constructed by crossbreeding a GAL4 system that expresses the GAL4 gene under the control of a promoter of a silk gene, etc., with a UAS system that expresses the target gene under the control of the control sequence UAS.
[0214] In the GAL4 / UAS system, the GAL4 and UAS systems need to be constructed separately before mating, thus the construction of the expression system takes time. Furthermore, the GAL4 gene and UAS control sequence are introduced into random locations on the genome, making the expression level of the target protein unpredictable. Therefore, after constructing multiple systems, it is necessary to study the expression levels in individuals obtained by mating the GAL4 and UAS systems, and select a suitable parental system based on the results.
[0215] Therefore, the inventors conceived of constructing a new expression system for the stable and large-scale production of target proteins by knocking the target gene sequence into endogenous genes. More specifically, if the target gene encoding the target protein is knocked into the exon sequence encoding the signal peptide in the endogenous sericin gene or fibroin gene in a manner fused with the signal peptide, it is possible to directly utilize the promoter and enhancer activities of the endogenous gene to achieve high expression of the target gene.
[0216] However, effective knock-in of endogenous genes requires the use of genome editing enzymes to cut target genes.
[0217] In Comparative Example 1, described later, the inventors attempted to knock in the genome by designing the cleavage site within the exon sequence. The result was that over 95% of the injected individuals in the current generation failed to form cocoons normally, and over 98% failed to develop into mating adults. Therefore, establishing a system using this method is extremely difficult.
[0218] Therefore, this embodiment investigates whether the above-mentioned problems can be overcome by designing the genome cut-off site within the intron sequence to implement knock-in into the exon sequence.
[0219] (Methods and Results)
[0220] In this embodiment, the target gene encoding the EGFP protein is knocked into the exon sequence encoding the signal peptide in the endogenous silk gene, serving as the target protein fused to the C-terminus of the signal peptide. The knock-in of the target gene utilizes the TAL-PITCh (precise integration into target chromosome) method and homologous recombination. The intron sequence adjacent to the 5' end of the target exon sequence is cleaved using the genome editing enzyme TALEN (transcription activator-like effector nuclease). Figure 2 Furthermore, in this embodiment, the EGFP gene sequence, which is the target gene, has a stop codon and a transcription termination sequence at the 3' end, which differs from embodiments 5-6 described later. Figure 1 ).
[0221] Knock-in into the fibH (FibH), fibL (FibL), and sericin 1 (Ser1) genes was performed using the following method. Furthermore, in subsequent embodiments, the vectors expressing TALENs were constructed according to Y. Takasu, S. Sajwan, T. Daimon, M. Osanai-Futahashi, K. Uchino, H. Sezutsu, T. Tamura, M. Zurovec (2013): Efficient TALEN construction for Bombyx mori gene targeting, PLoS One, 8, e73458. Additionally, TALEN mRNAs were synthesized using the respective TALEN expression vectors as templates via the mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen).
[0222] (1) Knock-in of the filoprotein H gene using the TAL-PITCh method
[0223] In the fibH gene (hereinafter referred to as the "FibH gene"), exons 1 and 2 encode a signal peptide. The EGFP gene sequence was introduced into exon 2 using the TAL-PITCh method by fusing EGFP protein to the C-terminus of the FibH signal peptide.
[0224] Specifically, in the genomic sequence of the FibH gene (Sequence No. 3), the position within the first intron between exon 1 and exon 2 (positions 1945 and 1964 in Sequence No. 3) is... Figure 2 A) As a genome cleavage site, it was used as a donor nucleic acid for construction in the TAL-PITCh method. Figure 3 The double-stranded circular DNA shown in B (hereinafter referred to as "TAL-PITCh SP(FibH)-EGFP donor nucleic acid") contains, in sequence, a first recognition sequence, a second spacer sequence, a first spacer sequence, a genomic homologous sequence, a target gene sequence, a transcription termination sequence, and a marker gene. Figure 3 B).
[0225] The first spacer sequence, adjacent to the 5' end of the genome cleavage site, is a 10-base sequence consisting of positions 1945-1954 in sequence number 3. The second spacer sequence, adjacent to the 3' end of the genome cleavage site, is also a 10-base sequence consisting of positions 1955-1964 in sequence number 3. The first recognition sequence, identified by Left TALEN at the 5' end of the first spacer sequence, is a 20-base sequence consisting of positions 1925-1944 in sequence number 3. The genomic homologous sequence is a 70-base sequence homologous to the genomic sequence of the codon encoding the C-terminal residue of the signal peptide from the second recognition sequence to the second exon sequence, consisting of positions 1965-2034 in sequence number 3. The second recognition sequence, identified by Right TALEN at the 3' end of the second spacer sequence, is a 20-base sequence consisting of positions 1965-1984 in sequence number 3. The target gene sequence is the EGFP gene sequence, and the transcription termination sequence used is the transcription termination sequence of the silkworm-derived sericin 1 gene. The base sequence from the first recognition sequence to the transcription termination sequence of the EGFP gene in the SP(FibH)-EGFP donor nucleic acid of TAL-PITCh is designated as sequence number 13. Additionally, the knock-in gene ( Figure 3C) The amino acid sequence of the protein encoding a signal peptide derived from FibH fused to the N-terminal side of the C-terminal fragment of EGFP except for the initiation methionine is designated by sequence number 14 (hereinafter referred to as "SP(FibH)-EGFP fusion protein").
[0226] TAL-PITCh, using SP(FibH)-EGFP donor nucleic acid, along with mRNA encoding Left TALEN and Right TALEN synthesized using the mMESSAGE mMACHINE T7ULTRA Transcription Kit (Invitrogen), was injected into silkworm eggs 2–8 hours after oviposition in the w1-pnd system, a white-eyed, white-egg, non-dormant system maintained at the National Research Institute of Agriculture and Food Science. The injected eggs were incubated at 25°C under humidification until hatching. The injected silkworms were mated with the parent system, and the next generation of larvae were selected using fluorescence screening with DsRed2 expressed systemically or EGFP expressed in the silk glands to obtain the knock-in silkworm system. Hereinafter, the obtained knock-in system will be referred to as the "SP(FibH)-EGFP knock-in system." Furthermore, the knock-in Fib H gene will be referred to as the "SP(FibH)-EGFP knock-in gene."
[0227] Of the 105 larvae that developed from the microinjected embryos, 103 formed cocoons normally. Therefore, no poor cocooning was observed in the injected larvae.
[0228] Furthermore, among the injected adult generation that developed from normally spun cocoons, 34 out of 36 female silkworms mated normally with wild-type male silkworms and laid eggs, and 50 out of 53 male silkworms mated normally with wild-type female silkworms and laid eggs. Therefore, no mating defects were found in the injected adult generation.
[0229] The results above show that by cutting the genome within intron sequences, it is possible to create knock-in systems that can cocoon and mate normally.
[0230] In the SP(FibH)-EGFP knock-in system, strong EGFP fluorescence was observed in the middle and posterior silk glands of first-instar larvae. Figure 5 The observation results of the silk glands and cocoons of 5th instar larvae are shown. Specifically, the larvae were anesthetized on ice on the 6th day of the 5th instar, just before they were about to spin silk. The dorsal side was cut open, and the silk glands were removed with forceps without damaging the middle and posterior parts. They were not fixed and observed under a fluorescence microscope. The results showed extremely strong EGFP fluorescence. Figure 5 (Left side). Additionally, the cocoons of this system exhibit a distinct yellowish-green color under normal white light (…). Figure 5 (Right side)
[0231] (2) Knocking into the filamentin H gene using homologous recombination.
[0232] The EGFP gene sequence was introduced into the second exon of the FibH gene by fusing the EGFP protein to the C-terminus of the FibH signal peptide using homologous recombination.
[0233] Specifically, similar to (1) above, the position within the first intron (the position between positions 1945 and 1964 in sequence number 3) is used. Figure 2 A) As a genome cleavage site, it served as the donor nucleic acid used in homologous recombination. Figure 4 The double-stranded circular DNA shown is referred to as "SP(FibH)-EGFP donor nucleic acid for homologous recombination". The SP(FibH)-EGFP donor nucleic acid for homologous recombination contains a first-genomic homologous sequence and a second-genomic homologous sequence, an EGFP gene sequence as the target gene sequence arranged between them, a transcription termination sequence, and a marker gene. A TALEN recognition sequence is contained at the ends of the first-genomic homologous sequence and the second-genomic homologous sequence on the opposite side from the EGFP gene sequence.
[0234] The first genomic homologous sequence is a 1034-base-long sequence homologous to the genomic sequence from the bases located at the 5' end of the genomic cut site (position 1001 in sequence number 3) to the codon encoding the C-terminal residue of the signal peptide in the second exon sequence. Furthermore, the first genomic homologous sequence has a mutation near the aforementioned genomic cut site in a manner not recognized by the Left TALEN and Right TALEN described in (1) above (specifically, the base sequence AACTTCGATTGAATGTGCGAAATTTATAGCTCAATATTTTAGCACTTATCGTATTGATTT (sequence number 33) located at positions 1925–1984 in sequence number 3 is replaced by the base sequence AtCTaCGATTGAAaGaGCGtAATTTATAGCTCAATATTTTAtGCtCaTAaCGTATTGATTT (sequence number 34)). The second genomic homologous sequence is a 377-base sequence homologous to the genomic sequence located on the 3' end of the codon encoding the signal peptide in exon 2 of the genome. In the SP(FibH)-EGFP donor nucleic acid used for homologous recombination, the sequence from the TALEN recognition sequence and the first genomic homologous sequence to the second genomic homologous sequence and the TALEN recognition sequence is indicated by sequence number 15. Additionally, the sequence derived from the knock-in gene ( Figure 4The protein encoded by SP(FibH)-EGFP, which has a signal peptide derived from FibH fused to the N-terminus of EGFP, is composed of the amino acid sequence shown in sequence number 14, just as described in (1) above (hereinafter, it is referred to as “SP(FibH)-EGFP fusion protein”, just as described in (1) above).
[0235] Homologous recombination using SP(FibH)-EGFP donor nucleic acid, along with mRNA encoding Left TALEN and Right TALEN synthesized using the mMESSAGE mMACHINE T7 ULTRATranscription Kit (Invitrogen), was injected into silkworm eggs of the w1-pnd system 2–8 hours after oviposition. The injected eggs were incubated in a humidified state at 25°C until hatching. The injected contemporary silkworm parent system was mated, and the next generation of larvae were selected by fluorescence screening using DsRed2 expressed systemically or EGFP expressed in the silk glands, thus obtaining the knock-in silkworm system. Hereinafter, the obtained knock-in system will be referred to as the "SP(FibH)-EGFP knock-in system," as described in (1) above.
[0236] In individuals developed from embryos that underwent the aforementioned microinjection, no cocooning or mating defects were observed, similar to those observed in (1) above. This demonstrates that, in homologous recombination, knock-in systems capable of normal cocooning and mating can also be efficiently created by cutting the genome within intron sequences.
[0237] (3) Knock-in of the fibL gene using homologous recombination
[0238] The EGFP gene sequence was introduced into the third exon of the FibL gene by means of fusing the EGFP protein to the C-terminus of the signal peptide of fibL (hereinafter referred to as "FibL").
[0239] Specifically, a double-stranded circular DNA (hereinafter referred to as "SP(FibL)-EGFP donor nucleic acid for homologous recombination") was constructed using the location within the second intron of the FibL gene (between positions 8937 and 8954 in sequence number 6) as the genome cleavage site, as the donor nucleic acid used in homologous recombination. The SP(FibL)-EGFP donor nucleic acid for homologous recombination contains a first-genomic homologous sequence and a second-genomic homologous sequence, as well as an EGFP gene sequence, a transcription termination sequence, and a marker gene positioned between them. A TALEN recognition sequence is contained at the ends of the first-genomic homologous sequence and the second-genomic homologous sequence opposite to the EGFP gene sequence.
[0240] Specifically, the aforementioned genome cutting positions are identified by Left TALEN, which identifies a 20-base-length sequence consisting of positions 8917 to 8936 in sequence number 6, and Right TALEN, which identifies a 20-base-length sequence consisting of positions 8955 to 8974 in sequence number 6.
[0241] Additionally, the first genomic homologous sequence is a 1128-base sequence homologous to the genomic sequence from the bases located at the 5' end of the genomic cut (position 7864 in sequence number 6) to the codon encoding the C-terminal residue of the signal peptide in the third exon sequence. Furthermore, the first genomic homologous sequence has a mutation near the aforementioned genomic cut position in a manner not recognized by Left TALEN and Right TALEN (specifically, the base sequence CCCGAGAAAACAATTTGTTGTGTATAATTTAAACCAAAACCCGAATTTAATTTTTCGC (sequence number 35) located at positions 8917–8974 in sequence number 6 is replaced by the base sequence CCCGAGAAAAgAATTcGTTcTGTATAATTTAAACCAAAAttCGAATTTAATTTTTCGC (sequence number 36)). The second genomic homologous sequence is a 1362-base-long sequence homologous to the genomic sequence located on the 3' end of the codon encoding the signal peptide in the third exon of the genome. In the SP(FibL)-EGFP donor nucleic acid used for homologous recombination, the base sequence from the TALEN recognition sequence and the first genomic homologous sequence to the second genomic homologous sequence and the TALEN recognition sequence is indicated by sequence number 16. In addition, the protein encoded by the knock-in gene, which has the FibL signal peptide fused to the N-terminus of the C-terminal fragment of EGFP except for the initiation methionine, consists of the amino acid sequence shown in sequence number 17 (hereinafter referred to as the "SP(FibL)-EGFP fusion protein").
[0242] Homologous recombination SP(FibL)-EGFP donor nucleic acid, along with mRNA encoding Left TALEN and Right TALEN synthesized using the mMESSAGE mMACHINE T7 ULTRATranscription Kit (Invitrogen), was injected into silkworm eggs of the w1-pnd system 2–8 hours after spawning. The injected eggs were incubated at 25°C under humidification until hatching. The knock-in silkworm system was obtained using the same method as described in (2) above. Hereinafter, the obtained knock-in system will be referred to as the "SP(FibL)-EGFP knock-in system," as in (1) above. Furthermore, the knock-in FibL gene will be referred to as the "SP(FibL)-EGFP knock-in gene."
[0243] In individuals developed from embryos that underwent the aforementioned microinjection, no poor cocooning or poor mating was found, similar to (1) above.
[0244] (4) Knock-in of the sericin 1 gene using homologous recombination.
[0245] The EGFP gene sequence was introduced into the second exon of the Ser1 gene by homologous recombination by fusing the EGFP protein to the C-terminus of the signal peptide of sericin 1 (hereinafter referred to as "Ser1").
[0246] Specifically, a double-stranded circular DNA (hereinafter referred to as "SP(Ser1)-EGFP donor nucleic acid for homologous recombination") was constructed using the location within the first intron of the Ser1 gene (between positions 3020 and 3033 in sequence number 9) as the genome cleavage site, as the donor nucleic acid used in homologous recombination. The SP(Ser1)-EGFP donor nucleic acid for homologous recombination contains a first-genomic homologous sequence and a second-genomic homologous sequence, as well as an EGFP gene sequence, a transcription termination sequence, and a marker gene positioned between them. A TALEN recognition sequence is contained at the ends of the first-genomic homologous sequence and the second-genomic homologous sequence opposite to the EGFP gene sequence.
[0247] The aforementioned genome cutting positions are identified by Left TALEN, which identifies a 19-base sequence consisting of positions 93001 to 3019, and Right TALEN, which identifies a 16-base sequence consisting of positions 3034 to 3049 in sequence 9.
[0248] Additionally, the first genomic homologous sequence is a 2069-base-long sequence homologous to the genomic sequence from the bases located at the 5' end of the genomic cut (position 947 in SEQ ID NO: 9) to the codon encoding the C-terminal residue of the signal peptide in the second exon sequence. Furthermore, the first genomic homologous sequence contains a mutation near the aforementioned genomic cut in a manner unrecognized by Left TALEN and Right TALEN (specifically, the base sequence TATATTTGTAAAGCACAACATATATATTAATGAATTTTTTATTTTTTTC (SEQ ID NO: 37) located at positions 3000–3049 in SEQ ID NO: 9 is replaced by the base sequence agTATTGagAAGCACAAgtaATATATTAATGAATTTTTTcTTTcTTTTTC (SEQ ID NO: 38)). The second genomic homologous sequence is a 2000-base-long sequence homologous to the genomic sequence located at the 3' end of the codon encoding the C-terminal residue of the signal peptide in the second exon sequence. In the SP(Ser1)-EGFP donor nucleic acid used for homologous recombination, the base sequence from the restriction endonuclease recognition sequence and the first genomic homologous sequence to the second genomic homologous sequence and the restriction endonuclease recognition sequence is indicated by sequence number 18. Additionally, the protein encoded by the knock-in gene, in which the signal peptide of Ser1 is fused to the N-terminus of the C-terminal fragment of EGFP except for the initiating methionine, consists of the amino acid sequence shown in sequence number 19 (hereinafter referred to as the "SP(Ser1)-EGFP fusion protein").
[0249] Homologous recombination using SP(Ser1)-EGFP donor nucleic acid, along with mRNA encoding Left TALEN and Right TALEN synthesized using the mMESSAGE mMACHINE T7 ULTRATranscription Kit (Invitrogen), was injected into silkworm eggs of the w1-pnd system 2–8 hours after spawning. The injected eggs were incubated at 25°C under humidification until hatching. The knock-in silkworm system was obtained using the same method as described in (2) above. Hereinafter, the obtained knock-in system will be referred to as the "SP(Ser1)-EGFP knock-in system," as in (1) above. Furthermore, the knock-in Ser1 gene will be referred to as the "SP(Ser1)-EGFP knock-in gene."
[0250] In individuals developed from embryos that underwent the aforementioned microinjection, no poor cocooning or poor mating was found, similar to (1) above.
[0251] <Example 2: Production of EGFP Protein>
[0252] (Purpose)
[0253] The expression level of EGFP protein in the silk glands of each knock-in system prepared in Example 1 was determined.
[0254] (Methods and Results)
[0255] (1) Silkworm System
[0256] By crossbreeding the various knock-in systems prepared in (2) to (4) of Example 1, a system combining multiple knock-in genes was created. In this example, for instance, a system having two knock-in genes, SP(FibH)-EGFP and SP(FibL)-EGFP, is referred to as the SP(FibH)-EGFP / SP(FibL)-EGFP knock-in system. Furthermore, in the silkworms used in this example, all knock-in genes are heterozygous.
[0257] In this embodiment, a system obtained by mating the FibH+Ser1-GAL4 system, which expresses the GAL4 gene under the control of the FibH gene promoter and the Ser1 gene promoter, and the UAS-EGFP system, which expresses the EGFP gene under the control of the UAS control sequence (hereinafter referred to as the "EGFP-generated GAL4 / UAS system") was used as a control group.
[0258] (2) Feeding conditions
[0259] Silkworm rearing is carried out using the following method. In a rearing room at 25–27°C, larvae of all ages are fed artificial feed (…). The original 1-3 year old S-type animals were raised by Japanese agricultural industry. The artificial feed was changed every 2-3 days (Uchino K. et al., 2006, J Insect Biotechnol Sericol, 75:89-97).
[0260] (3) Determination of EGFP protein expression level
[0261] On the 6th day of the 5th year, before the silk-producing stage, the silk glands were anesthetized on ice. The dorsal side was cut open, and the silk glands were removed with forceps without damaging them. The middle and posterior silk glands were removed. Each gland was placed in 10 mL of PBS (pH 7.2) / 1% Tween 20 / 0.05% sodium azide and shaken at room temperature for 24 hours to extract water-soluble proteins. The resulting water-soluble protein extract was centrifuged at 2000×g for 10 minutes, and the supernatant was collected. The concentration of EGFP protein in the water-soluble proteins contained in the supernatant was determined by ELISA. Specifically, the supernatant was coated with anti-GFP antibody (Aves GFP-1010, ...). Add 100 μL of supernatant to a 96-well plate and incubate at room temperature for 1 hour. Wash three times with PBS / 0.05% Tween 20, then add horseradish peroxidase-conjugated anti-GFP antibody (Rockland Immunochemicals) and incubate at room temperature for 1 hour. After washing three times with PBS / 0.05% Tween 20, perform a colorimetric reaction using a TMB Peroxidase EIA Substrate Kit (Bio-Rad), and stop the reaction by adding 1N sulfuric acid. Quantify the colorimetric results using a microplate reader (SpectraMax iD3; Molecular Devices). Use recombinant GFP protein (… A standard curve was prepared using serially diluted solutions (1–400 pg / μL) of Z2373N.
[0262] The results of EGFP expression level determination for each silkworm are shown below. Figure 6 .
[0263] SP(Ser1)-EGFP knock-in system Figure 6 The EGFP expression level in SP(Ser1)-EGFP was slightly higher than that in the EGFP-producing GAL4 / UAS system. Figure 6 (Control (GAL4 / UAS)). In the EGFP-generated GAL4 / UAS system, GAL4 protein is expressed through two promoters: the Ser1 gene promoter and the FibH gene promoter. Since the amount of EGFP expressed in 3.3 mg is estimated to be less than 1 mg from the Ser1 gene promoter, the EGFP expression level in the SP(Ser1)-EGFP knock-in system is considered to be overwhelmingly high.
[0264] SP(FibH)-EGFP knock-in system Figure 6 EGFP expression levels of SP(FibH)-EGFP and SP(FibL)-EGFP knock-in system ( Figure 6 Compared to SP(FibL)-EGFP, it is more than twice as effective, and compared to the SP(Ser1)-EGFP knock-in system ( Figure 6 Compared to SP(Ser1)-EGFP, it is more than 4 times higher. This result is unexpected, given that the molar number of FibH protein and FibL protein produced in the silk gland of silkworms is equal in terms of the composition of filamentin, and that filamentin, which makes up silk, is 3 times more abundant than sericin by weight.
[0265] SP(FibH)-EGFP / SP(FibL)-EGFP / SP(Ser1)-EGFP knock-in system Figure 6The right end shows an EGFP expression level of 33.7 mg, which is significantly greater than the expected amount as a sum of the EGFP expression levels in the SP(FibH)-EGFP knock-in system, the SP(FibL)-EGFP knock-in system, and the SP(Ser1)-EGFP knock-in system.
[0266] <Example 3: Production of GM-CSF Protein>
[0267] (Purpose)
[0268] Homologous recombination knock-in was performed by fusing granulocyte-macrophage colony-stimulating factor (GM-CSF) to the C-terminus of the FibH signal peptide. GM-CSF production in the serine glands was evaluated.
[0269] (Methods and Results)
[0270] (1) Silkworm System
[0271] The homologous recombination method described in Example 1 (2) was performed by replacing the target gene from the EGFP gene sequence with the GM-CSF gene sequence. The GM-CSF gene sequence consists of the base sequence shown in sequence number 20 and encodes the GM-CSF protein consisting of the amino acid sequence shown in sequence number 21. The resulting knock-in system is called the "SP(FibH)-GM-CSF knock-in system". Furthermore, the knock-in gene ( Figure 7 A) A protein encoding a signal peptide derived from FibH that is fused to the N-terminus of the mature amino acid sequence of the signal peptide derived from GM-CSF in GM-CSF, consisting of the amino acid sequence shown in sequence number 22 (hereinafter referred to as "SP(FibH)-GM-CSF fusion protein").
[0272] In addition, a system obtained by mating the Ser1-GAL4 system, which expresses the GAL4 gene under the control of the Ser1 gene promoter, with the UAS-GM-CSF system, which expresses the GM-CSF gene under the control of the UAS control sequence (hereinafter referred to as the "Middle serous gland GM-CSF producing GAL4 / UAS system") and a system obtained by mating the FibH-GAL4 system, which expresses the GAL4 gene under the control of the FibH gene promoter, with the UAS-GM-CSF system (hereinafter referred to as the "Posterior serous gland GM-CSF producing GAL4 / UAS system") were used as control groups.
[0273] (2) Protein blotting
[0274] Following the method described in Example 2, the middle and posterior silk glands were extracted from each system, and proteins were extracted from each gland. Next, the undiluted or diluted (2-128 times) silk gland extract was mixed with a specified amount of NuPAGE LDSSample Buffer (Thermo Fisher) and NuPAGE Sample Reducing Agent (Thermo Fisher), and heated at 70°C for 10 minutes to achieve SDS-PAGE. The SDS-PAGE sample was then used in an 8cm × 13cm 4-12% SDS-PAGE gel and sputtered at a constant current of 20mA for approximately 90 minutes. The gel was then transferred to a PVDF membrane using a semi-dry transfer apparatus (iBlot2, Thermo Fisher). After the transferred membrane was gently shaken in EZ wash (AE-1480, ATTO) for 5 minutes, the primary antibody (anti-GM-CSF antibody, 3000-fold dilution, manufactured by Immundiagnostik, product number AS1021.2) was incubated overnight at 4°C. The membrane was washed three times with EZwash for 10 minutes each time, and then the secondary antibody (Anti-Rabbit IgG, HRP-Linked Whole AbDonkey, manufactured by Cytiva, product number NA934-100UL, 50,000-fold dilution) was allowed to react at room temperature for 1 hour. The membrane was then washed three times with EZwash for 10 minutes each time, and then reacted with ECL prime (RPN2232, GE HealthCare) for 5 minutes. The signal was detected using Fusion FX (Vilber Bio Imaging).
[0275] The results are shown in Figure 7 B. In the extracts from the combined middle and posterior serine glands of the SP(FibH)-GM-CSF knock-in system, approximately 13-fold and 3-fold GM-CSF were detected compared to the GAL4 / UAS production system of the middle serine gland and the GAL4 / UAS production system of the posterior serine gland, respectively. This indicates highly efficient expression and secretion of GM-CSF in the serine glands of the SP(FibH)-GM-CSF knock-in system.
[0276] <Example 4: Antibody Production>
[0277] (Purpose)
[0278] Knock-in was performed using homologous recombination by fusing an IgG H chain to the C-terminus of the FibH signal peptide. Similarly, knock-in was performed using homologous recombination by fusing an IgG L chain to the C-terminus of the FibL signal peptide. The resulting two knock-in systems were then cross-linked to produce an antibody molecule containing both an IgG H chain and an IgG L chain.
[0279] (Methods and Results)
[0280] (1) Silkworm System
[0281] The homologous recombination method described in Example 1(2) was performed by replacing the target gene sequence from the EGFP gene sequence with the IgG H chain gene sequence. The IgG H chain gene sequence consists of the base sequence shown in sequence number 23, encoding an IgG H chain composed of the amino acid sequence shown in sequence number 24. The resulting knock-in system is called the "SP(FibH)-IgG H chain knock-in system". Furthermore, the knock-in gene ( Figure 8 A) A protein encoding a signal peptide derived from FibH fused to the N-terminus of the IgG H chain, consisting of the amino acid sequence shown in Serial No. 25 (hereinafter referred to as "SP(FibH)-IgG H chain fusion protein").
[0282] In addition, the homologous recombination method described in Example 1 (3) was performed by replacing the target gene from the EGFP gene sequence with the IgG L chain gene sequence. The IgG L chain gene sequence consists of the base sequence shown in sequence number 26 and encodes an IgG L chain consisting of the amino acid sequence shown in sequence number 27. The resulting knock-in system is called the "SP(FibL)-IgG L chain knock-in system". Furthermore, the knock-in gene ( Figure 8 B) A protein encoding a signal peptide derived from FibL fused to the N-terminus of an IgG L chain, consisting of the amino acid sequence shown in Serial No. 28 (hereinafter referred to as "SP(FibL)-IgG L chain fusion protein").
[0283] By mating the two knock-in systems mentioned above, a system with two knock-in genes is created that can produce IgG molecules containing IgG H and IgG L chains (hereinafter referred to as the "SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in system").
[0284] In this embodiment, a system obtained by mating the FibH+Ser1-GAL4 system, which expresses the GAL4 gene under the control of the FibH gene promoter and the Ser1 gene promoter, the UAS-IgG H chain system, which expresses the IgG H chain gene under the control of the UAS control sequence, and the UAS-IgG L chain system, which expresses the IgG L chain gene under the control of the UAS control sequence (hereinafter referred to as the "antibody-generating GAL4 / UAS system") was used as a control group.
[0285] (2) Quantitative analysis of IgG expression
[0286] Following the method described in Example 2, the middle and posterior silk glands of each system were removed, and proteins were extracted from each silk gland. Next, IgG was purified using Ab SpinTrap (Cytiva), and proteins were analyzed using a protein assay kit (BCA). The amount of IgG in the extract was quantified.
[0287] The results are shown in Figure 8 C. The SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in system demonstrated the ability to produce approximately 4.6 times more IgG than the middle and posterior serine glands in the GAL4 / UAS antibody-producing system. This indicates highly efficient expression and secretion of antibody molecules containing both IgG H and IgG L chains within the serine glands of the SP(FibH)-IgG H chain / SP(FibL)-IgG L chain knock-in system.
[0288] <Example 5: Fabrication of a knock-in system for fusion protein production>
[0289] (Purpose)
[0290] In existing technologies for producing functional filaments, such as fluorescent filaments containing filament proteins fused with fluorescent proteins and functional peptides, methods are known to replace the repetitive sequence portion in the central part of FibH with the target protein, and to fuse the target protein to the C-terminus of FibL. However, existing technologies suffer from problems such as low expression levels and loss of activity of the target protein.
[0291] The inventors conceived that by knocking in the target gene into the exon sequence encoding the signal peptide of silk genes such as sericin gene and filamentin gene, and fusing the target protein between the peptide and the mature protein obtained by cleaving the signal peptide from the precursor protein encoded by the endogenous silk gene, it is possible to directly utilize the promoter activity and enhancer activity of the endogenous silk gene to highly express the silk protein fused with the target protein.
[0292] Therefore, in this embodiment, similar to the embodiments described above, a genome cleavage site is designed within the intron sequence located at the 5' end of the exon sequence encoding the signal peptide, and homologous recombination is applied to create a knock-in system expressing the fusion silk protein.
[0293] (Methods and Results)
[0294] In this embodiment, a target gene encoding EGFP protein is knocked into the exon sequence encoding a signal peptide in an endogenous silk gene, serving as the target protein fused to the C-terminus of the signal peptide. The knock-in of the target gene utilizes homologous recombination, where the intron sequence adjacent to the 5' end of the target exon sequence is cleaved using the genome editing enzyme TALEN. Figure 10 Furthermore, in this embodiment, the EGFP gene sequence, which is the target gene, does not contain a stop codon at the 3' end, and the EGFP protein is inframed with the filament protein (a mature protein obtained by cleaving the signal peptide from a precursor protein encoded by an endogenous filament gene), which differs from the embodiments 1 to 4 described above.
[0295] The donor nucleic acids used in the knock-in of the fibL and Ser1 genes were prepared by the following method.
[0296] (1) EGFP-FibL donor nucleic acid
[0297] In the homologous recombination method described in Example 1 (3), the transcription termination sequence and marker gene were removed from the SP(FibL)-EGFP donor nucleic acid described in Example 1. The EGFP gene sequence and the second genome homologous sequence were modified in such a way that the C-terminal amino acid residues of the EGFP protein could fuse with the N-terminal amino acid residues of FibL after the signal peptide was removed, and a donor nucleic acid (double-stranded circular DNA) for homologous recombination (hereinafter referred to as "EGFP-FibL donor nucleic acid") was constructed. The genome cleavage position, Left TALEN and Right TALEN, first genome homologous sequence, TALEN recognition sequence, etc. in the endogenous FibL gene were constructed in accordance with Example 1. In the constructed donor nucleic acid, the base sequence from the TALEN recognition sequence and the first genome homologous sequence to the second genome homologous sequence and the restriction endonuclease recognition sequence is the sequence in which the stop codon TAA at positions 1821 to 1823 of the SP(FibH)-EGFP donor nucleic acid shown in Serial No. 15 was deleted. Additionally, the EGFP-FibL precursor protein, encoded by the knock-in gene, has a FibL signal peptide fused to the N-terminus of EGFP and a mature FibL protein cleaved from the signal peptide fused to the C-terminus of EGFP, and consists of the amino acid sequence shown in sequence number 29. The mature EGFP-FibL protein, obtained by cleaving the signal peptide from the EGFP-FibL precursor protein, consists of the amino acid sequence shown in sequence number 30.
[0298] (2) EGFP-Ser1 donor nucleic acid
[0299] In the homologous recombination method described in Example 1 (4), the transcription termination sequence and marker gene were removed from the SP(Ser1)-EGFP donor nucleic acid described in Example 1. The EGFP gene sequence and the second genome homologous sequence were modified in such a way that the C-terminal amino acid residues of the EGFP protein could fuse with the N-terminal amino acid residues of the mature Ser1 after the signal peptide was cleaved, and a donor nucleic acid (double-stranded circular DNA) for homologous recombination (hereinafter referred to as "EGFP-Ser1 donor nucleic acid") was constructed. The genome cleavage position, Left TALEN and Right TALEN, first genome homologous sequence, TALEN recognition sequence, etc. in the endogenous Ser1 gene were constructed in accordance with Example 1. In the constructed donor nucleic acid, the base sequence from the restriction endonuclease recognition sequence and the first genome homologous sequence to the second genome homologous sequence and the restriction endonuclease recognition sequence is the sequence in which the stop codon TAA at positions 2841 to 2843 of the SP(Ser1)-EGFP donor nucleic acid shown in Serial No. 18 was deleted. Additionally, the EGFP-Ser1 precursor protein, encoded by the knock-in gene, has the Ser1 signal peptide fused to the N-terminus of EGFP and the mature Ser1 protein cleaved from the signal peptide fused to the C-terminus of EGFP, and consists of the amino acid sequence shown in SEQ ID NO. 31. The mature EGFP-Ser1 protein, obtained by cleaving the signal peptide from the EGFP-Ser1 precursor protein, consists of the amino acid sequence shown in SEQ ID NO. 32.
[0300] The donor nucleic acids prepared in (1) and (2) above were injected into silkworm eggs together with the mRNA encoding TALEN according to the methods described in (3) and (4) of Example 1 to create a knock-in system. Hereinafter, they are referred to as the "EGFP-FibL knock-in system" and the "EGFP-Ser1 knock-in system", respectively. In addition, the knock-in FibL gene and Ser1 gene are referred to as the "EGFP-FibL knock-in gene" and the "EGFP-Ser1 knock-in gene", respectively.
[0301] No poor cocooning or poor mating was found in individuals developed from embryos that underwent the aforementioned microinjection.
[0302] <Example 6: Production of a fusion protein containing EGFP protein and silk protein>
[0303] (Purpose)
[0304] The knock-in system developed in Example 5 was used to produce a fusion protein containing EGFP protein and silk protein.
[0305] (Methods and Results)
[0306] Following the method described in Example 2, the middle and posterior silk glands were extracted from each knock-in system prepared in Example 5, and proteins were extracted from each silk gland. Next, Western blotting was performed according to the method described in (2) of Example 3. Furthermore, in this example, peroxide-conjugated anti-GFP antibody (4000-fold dilution, manufactured by ROCKLAND, product number 600-103-215) was used as the detection antibody. Additionally, in this example, the SP(Ser1)-EGFP knock-in system and the w1-pnd system prepared in Example 1 (4) were used as control groups.
[0307] The results of the protein blotting are shown in Figure 11 Mature EGFP-FibL protein was detected in the EGFP-FibL knock-in system. Figure 11 A). Mature EGFP-Ser1 protein was detected in the EGFP-Ser1 knock-in system. Figure 11 B).
[0308] <Example 7: Homozygote with Knock-in Gene>
[0309] (Purpose)
[0310] Individuals homozygous for the EGFP-FibL knock-in gene were created and cocooned.
[0311] (Methods and Results)
[0312] EGFP-FibL knock-in systems heterozygous for the EGFP-FibL knock-in gene were crossbred to create homozygous individuals containing the EGFP-FibL knock-in gene. Photographs of the cocoons obtained from the homozygous and heterozygous individuals, taken under normal white light, are shown below. Figure 12 The homozygote displays a very bright yellow-green color.
[0313] Fluorescent cocoons produced using the piggyBac system according to the method described in the literature (Insect Biochemistry and Molecular Biology, 2005, 35: 51-59) exhibit the strongest fluorescence among existing techniques. The results of observing wild-type silkworm cocoons, FibL-fused EGFP cocoons produced using the piggyBac system, and cocoons produced by the aforementioned heterozygotes and homozygotes side-by-side under white light or fluorescence are shown below. Figure 13 Furthermore, in fluorescence observation, fluorescence was detected using blue light excitation and an EGFP filter. The heterozygotes of the EGFP-FibL knock-in system prepared by the method of this invention showed strong fluorescence compared to the FibL-fused EGFP cocoons prepared using the piggyBac system, indicating that the homozygotes exhibited overwhelmingly strong fluorescence.
[0314] <Comparative Example 1: Creation of a knock-in system with genome cut-off sites designed within exon sequences>
[0315] (Purpose)
[0316] Instead of designing genomic cleavage sites within intron sequences, we knocked the EGFP gene into the filocene H gene using homologous recombination. The efficiency of the knock-in system was compared with that of cleaving the genome within intron sequences.
[0317] (Methods and Results)
[0318] In the homologous recombination method described in Example 1 (2), the EGFP gene sequence was knocked into the fibrin H gene by altering the method in a manner that sets the genome cut-off site within the second exon of the FibH gene. Specifically, a mutation was introduced into the donor nucleic acid for homologous recombination near the genome cut-off site within the second exon in a manner that is not recognized by TALEN. In 384 eggs, the aforementioned donor nucleic acid for homologous recombination, together with mRNA encoding TALEN synthesized using the mMESSAGE mMACHINE T7 ULTRATranscription Kit (Invitrogen), was injected into silkworm eggs of the w1-pnd system 2–8 hours after oviposition. The injected eggs were incubated in a humidified state at 25°C until hatching. The presence or absence of mating ability was determined by raising the injected contemporary silkworms to adulthood and mating them with the parent system.
[0319] The results are shown in Figure 14 B. In homologous recombination with exon sequence-based genome cutting, of the 118 larvae that developed into 5th instar larvae, 113 formed incomplete pupation or naked pupae and thin cocoons; even those that emerged showed incomplete mating, with only 5 producing normal cocoons. However, of the 5 normal cocoons, only 2 individuals were capable of mating: only 1 out of 2 females and only 1 out of 3 males. As a reason for the abnormalities occurring with exon sequence cutting, it is considered that the injected embryos may have undergone mutations such as frameshifts at the genome cutting sites, resulting in the inability to express normally functioning proteins.
[0320] This result is similar to the result in Example 1 (2) where no poor cocooning or mating was found when the genome was cut within the intron sequence. Figure 14 A) is a control, showing that the method for preparing the knock-in silkworm system described in Example 1 can produce the knock-in system with overwhelmingly high efficiency. This is likely because even if some mutations are added to the intron sequence, the effect on the expression of normal proteins is slight.
[0321] Industry availability
[0322] According to the present invention, the target protein can be stably and massively produced in lepidopteran insects such as silkworms.
[0323] All publications, patents and patent applications cited in this specification are incorporated herein by reference.
Claims
1. Genetically recombinant Lepidoptera insects, The recombinant lepidopteran insects contain a target gene sequence encoding a target protein or a fragment thereof in the exon sequence encoding the signal peptide or functional fragment of the endogenous gene. The target protein or a fragment thereof is fused between the signal peptide or a functional fragment thereof and a mature protein or a C-terminal fragment thereof, the mature protein being formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
2. The recombinant lepidopteran insect according to claim 1, wherein the mature protein is filamentin, sericin, and / or fibrous hexamer protein.
3. The recombinant lepidopteran insect according to claim 2, wherein the filoprotein is a filoprotein H chain and / or a filoprotein L chain.
4. The recombinant lepidopteran insect according to claim 1, wherein the target protein is selected from fluorescent proteins, antibodies, antigenic peptides, enzymes, cytokines, and antimicrobial peptides.
5. The recombinant lepidopteran insect according to claim 1, homozygously comprising the exon sequence containing the target gene sequence.
6. Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the 5' end of the genome cut site to the 3' end of the exon sequence or a portion thereof, and the recognition sequence of the genome editing enzyme has a mutation. The second genomic homologous sequence consists of a base sequence homologous to a genomic sequence located on the 3' end of the exon sequence or a portion thereof. The target gene sequence encodes a target protein or a fragment thereof fused between a signal peptide or a functional fragment thereof of the endogenous gene and a mature protein or a C-terminal fragment thereof, wherein the mature protein is formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
7. Donor nucleic acid is the donor nucleic acid used to produce recombinant lepidopteran insect genes using homologous recombination. The homologous recombination method involves cutting the genome at a cleavage site within an intron sequence in an endogenous gene using a genome editing enzyme. The donor nucleic acid contains (a) Homologous sequences of the first and second genomes derived from the endogenous gene, and (b) The target gene sequence positioned between the homologous sequence of the first genome and the homologous sequence of the second genome. The first genomic homologous sequence consists of a sequence of bases homologous to the genomic sequence from the bases located at the 5' end of the intron sequence to the exon sequence or a portion thereof located at the 5' end of the intron sequence. The second genomic homologous sequence consists of a base sequence homologous to the genomic sequence from the bases located at the 3' end of the exon sequence or a portion thereof and at the 5' end of the genomic cut site, to the bases located at the 3' end of the genomic cut site, and the recognition sequence of the genome editing enzyme has a mutation. The target gene sequence encodes a target protein or a fragment thereof fused between a signal peptide or a functional fragment thereof of the endogenous gene and a mature protein or a C-terminal fragment thereof, wherein the mature protein is formed by cleaving the signal peptide from a precursor protein encoded by the endogenous gene.
8. The donor nucleic acid according to claim 6 or 7, wherein the end of the first genomic homologous sequence and / or the second genomic homologous sequence opposite to the target gene sequence contains a nuclease recognition sequence.
9. The donor nucleic acid according to claim 8, wherein the nuclease recognition sequence is the recognition sequence of the genome editing enzyme or the recognition sequence of the restriction endonuclease.
10. A method for producing recombinant lepidopteran insects, the method comprising: Will The donor nucleic acid as described in claim 6 or 7, and The genome editing enzyme, or the nucleic acid encoding the genome editing enzyme in an expressible state. The process of introducing microinjection into the eggs of lepidopteran insects.
11. A method for producing a fusion protein, comprising a genetically recombinant lepidopteran insect according to any one of claims 1 to 5, and a fusion protein comprising the target protein or a fragment thereof, and the mature protein or its C-terminal fragment thereof.
12. The method according to claim 11, wherein the lepidopteran insect is a silk-producing insect, and the fusion protein is produced in the silk gland of the silk-producing insect.
13. Fusion proteins, starting from the N-terminus, sequentially contain... Target protein or fragment thereof, and Silken protein, sericin, or fibrohexamer protein.
14. The fusion protein according to claim 13, wherein the filoprotein is a filoprotein H chain and / or a filoprotein L chain.
15. A cocoon or silk comprising the fusion protein of claim 13.
16. Cocoon or silk, The silk protein contained in the cocoon or silk, selected from any one or more of the following: filamentin H chain, filamentin L chain, sericin 1, sericin 2, sericin 3, and fibrous hexamer protein, is composed of the fusion protein described in claim 13.
17. The cocoon or silk according to claim 16, derived from a recombinant lepidopteran insect according to any one of claims 1 to 5.
18. The cocoon or silk according to claim 16, wherein the target protein is selected from fluorescent proteins, antibodies, antigenic peptides, enzymes, cytokines, and antimicrobial peptides.
Citation Information
Patent Citations
Ultrapure water production device, and operation management method of ultrapure water production device
JP2023114712A