Peptide synthetase library construction method

By introducing non-natural restriction enzyme recognition sequences into NRPS DNA, the method addresses limitations in NRP production, enabling efficient synthesis of diverse NRPs with tailored NRPS systems.

CN120322560APending Publication Date: 2025-07-15KOBE UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380070358.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-07-29
Filing Date
2023-07-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently synthesize non-ribosome peptides (NRPs) with specific structures, especially in the synthesis of peptides containing non-protein amino acids, and there is a lack of efficient NRP production methods.

Method used

By introducing non-natural restriction enzyme recognition sequences into nucleic acids encoding NRPS, using restriction enzymes to treat and ligate broken nucleic acid fragments, we can construct modified NRPS to achieve efficient combination of multiple NRPS modules and the production of new NRPs.

Benefits of technology

The efficient production of new NRPs has been achieved, the types and quantities of NRPs have been expanded, and the efficiency and diversity of NRP production have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BPA0000372218130000411
    Figure BPA0000372218130000411
  • Figure BPA0000372218130000412
    Figure BPA0000372218130000412
  • Figure BPA0000372218130000413
    Figure BPA0000372218130000413
Patent Text Reader

Abstract

The invention relates to a novel NRPS. In the present invention, a restriction enzyme recognition sequence is successfully introduced into a nucleic acid encoding an NRPS while maintaining the function of producing an NRP of interest. Various NRPS can be produced using the introduced restriction enzyme recognition sequence. In one aspect, the present invention provides NRPS-encoding nucleic acids comprising a non-natural restriction enzyme recognition sequence. In other aspects, the present invention provides a method for efficiently producing a nucleic acid encoding a novel NRPS using a non-natural restriction enzyme recognition sequence. In addition, the present invention also provides a novel NRPS and a novel NRP.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a modified non-ribosomal peptide synthetase (NRPS) and its use. The present invention also provides a plasmid or plasmid library encoding the modified NRPS, or a method for using or producing these. In addition, the present invention also provides a non-ribosomal peptide (NRP) produced by NRPS. Background Art

[0002] NRPS is known as an enzyme that synthesizes a specific structure (Non-patent Document 1). Since peptide synthesis using NRPS does not rely on ribosomes, peptides containing non-proteinogenic amino acids (ornithine, D-configured amino acids, etc.) can be synthesized, and it is known that there are more than 800 types of amino acids that can be contained in NRP. NRPs synthesized by NRPS are used in various fields such as antibiotics and pesticides. Development of new NRPS is expected to provide efficient production of NRP and new NRPs. Prior Art Documents Non-Patent Documents

[0003] Non-Patent Document 1: Angew Chem Int Ed Engl. 2017 Mar 27; 56(14): 3770 - 3821. Summary of the Invention Means for Solving the Problem

[0004] The present inventors have successfully introduced a restriction enzyme recognition sequence into a nucleic acid encoding NRPS while maintaining the function of producing the target NRP. Various NRPS can be produced using the introduced restriction enzyme recognition sequence. In one aspect, the present invention provides a nucleic acid encoding NRPS containing an unnatural restriction enzyme recognition sequence. In other aspects, the present invention provides a method for efficiently producing a nucleic acid encoding a new NRPS using an unnatural restriction enzyme recognition sequence. In addition, the present invention also provides a new NRPS and a new NRP.

[0005] Therefore, the present invention provides the following. (Item 1) A method for producing a production plasmid containing a nucleic acid sequence encoding a non-ribosomal peptide synthetase (NRPS), comprising: a step of preparing a starting plasmid group, each of which contains a base sequence encoding at least 1 NRPS module and at least 1 restriction enzyme recognition sequence; a step of treating the starting plasmid group with a restriction enzyme to prepare a mixture containing a plurality of resulting fragmented nucleic acid fragments; a step of ligating the plurality of fragmented nucleic acid fragments to form a ligated nucleic acid; and a step of bringing the ligated nucleic acid into contact with a transformation organism to form a production plasmid. (Item 2) According to the method of the above item, the multiple fragmented nucleic acid fragments include a group of fragmented nucleic acid fragments, the fragmented nucleic acid fragments have the same overhanging sequence, and contain a base sequence encoding a variety of NRPS modules. (Item 3) According to any one of the methods of the above items, the multiple fragmented nucleic acid fragments include a variety of groups of fragmented nucleic acid fragments, the fragmented nucleic acid fragments have the same overhanging sequence, and contain a base sequence encoding a variety of the NRPS modules that capture amino acids. (Item 4) According to any one of the methods of the above items, each of the fragmented nucleic acid fragments has an overhanging sequence with a length of 3 to 5 bases. (Item 5) According to any one of the methods of the above items, in the starting plasmid, the restriction enzyme recognition sequence is located between adjacent base sequences encoding NRPS modules. (Item 6) According to any one of the methods of the above items, in the starting plasmid, the restriction enzyme recognition sequence is located between the C domain coding region and the A domain coding region of the NRPS module. (Item 7) According to any one of the methods of the above items, in the starting plasmid, the restriction enzyme recognition sequence is located in the region between the C domain coding region and the A domain coding region of the NRPS module, excluding the C domain coding region and the A domain coding region. (Item 8) According to any one of the methods of the above items, at least one of the restriction enzyme recognition sequences is present within the base sequence encoding the NRPS module. (Item 9) According to any one of the methods of the above items, at least one of the restriction enzyme recognition sequences is present outside the base sequence encoding the NRPS module. (Item 10) According to any one of the methods of the above items, the restriction enzyme recognition sequence is a non-natural sequence in the NRPS module. (Item 11) According to any one of the methods of the above items, each of the fragmented nucleic acid fragments contains at most one NRPS module. (Item 12) According to any one of the methods of the above items, the restriction enzyme recognition sequence is cleaved at a site different from the specific recognition site of the restriction enzyme. (Item 13) According to the method of any one of the above items, the restriction enzyme recognition sequence includes the recognition sequences of SfiI, BglI, AlwNI, DraIII, PflMI, BstAPI, AarI, BbsI, BsaI, BsmBI or BspQI. (Item 14) According to the method of any one of the above items, each of the starting plasmids contains a base sequence encoding 4 to 12 NRPS modules. (Item 15) According to the method of any one of the above items, in the step of preparing the mixture containing the fragmented nucleic acid fragments, each of the starting plasmids generates fragmented nucleic acid fragments having mutually different protruding sequences. (Item 16) According to the method of any one of the above items, the mixture containing the fragmented nucleic acid fragments is a solution. (Item 17) According to the method of any one of the above items, each of the starting plasmids independently contains the same or different selection marker sequences. (Item 18) According to the method of any one of the above items, the selection marker sequence includes a drug resistance marker sequence. (Item 19) According to the method of any one of the above items, each of the starting plasmids independently contains the same or different replication origins that can function in Bacillus subtilis. (Item 20) According to the method of any one of the above items, each of the starting plasmids independently contains a promoter region upstream of the same or different base sequences encoding the NRPS. (Item 21) A production plasmid or plasmid library is produced by the method of any one of the above items. (Item 22) A non-ribosomal peptide synthetase (NRPS) is produced from any one of the above production plasmids or plasmid libraries. (Item 23) A non-ribosomal peptide (NRP) is produced from any one of the above NRPSs. (Item 24) A nucleic acid, which is a nucleic acid encoding a non-ribosomal peptide synthetase (NRPS), contains a base sequence encoding at least 1 NRPS module and a restriction enzyme recognition sequence that is not natural for the NRPS module. (Item 25) According to the nucleic acid of any one of the above items, the restriction enzyme recognition sequence is located between adjacent base sequences encoding the NRPS module. (Item 26) For any nucleic acid of the above items, the restriction enzyme recognition sequence is located between the coding region of the C domain and the coding region of the A domain of the NRPS module. (Item 27) For any nucleic acid of the above items, the restriction enzyme recognition sequence is located in the region between the coding region of the C domain and the coding region of the A domain of the NRPS module, excluding the coding region of the C domain and the coding region of the A domain. (Item 28) For any nucleic acid of the above items, the restriction enzyme recognition sequence is the recognition sequence of a restriction enzyme that recognizes a sequence containing a region where any type of base can be present. (Item 29) For any nucleic acid of the above items, the restriction enzyme recognition sequence is the recognition sequence of SfiI. (Item 30) For any nucleic acid of the above items, the number of the NRPS modules encoded by the nucleic acid is 4 to 12. (Item 31) A method for producing a product nucleic acid containing a nucleic acid sequence encoding a non-ribosomal peptide synthetase (NRPS), comprising the steps of preparing any nucleic acid of the above items as a starting nucleic acid, treating the nucleic acid with a restriction enzyme that recognizes the restriction enzyme recognition sequence to prepare fragmented nucleic acid fragments, and ligating the fragmented nucleic acid fragments to form a product nucleic acid different from the starting nucleic acid. (Item 32) A product nucleic acid produced by the method according to any one of the above items.

[0006] In the present invention, except for the explicit combinations, one or more of the above features can be further combined and provided. Those skilled in the art can recognize more embodiments and advantages of the present invention by understanding the following detailed description as needed. Effects of the Invention

[0007] According to the present invention, a novel NRPS can be efficiently produced. In this way, an NRPS capable of efficiently producing NRP and an NRPS providing a novel NRP can be provided. Brief Description of the Drawings

[0008] Figure 1 shows an overview of the structure of plipastatin (A) and the NRPS (B) that produces it. A: Adenylation domain, T: Thiolation domain, C: Condensation domain, TE: Thioesterase domain, E: Epimerization domain.​ Figure 2 represents the alignment of the C-A linker region between the NRPS modules of pulcherrimin and the introduction point of the SfiI recognition sequence (boxed part). Figure 3 represents the outline of the introduction of the SfiI recognition sequence. Figure 4 represents the outline of the accumulation unit DNA by the OGAB method. Figure 5 represents the comparison of the gel electrophoresis bands when the plasmids ppsW03 and ppsW10 are treated with various restriction enzymes. Figure 6 represents the HPLC analysis results of the extracts of the culture broths of each strain. The results of (a) BEST8628, (b) ppsW03 / BUSY9261, (c) ppsW10 / BUSY9261, and (d) BUSY9261 strains are shown respectively. The vertical axis represents the absorbance at 205 nm, and the horizontal axis represents the retention time (minutes). Figure 7 represents the results of the determination of the substance produced by the ppsW03 / BUSY9621 strain using a mass spectrometer. Peaks a to d represent pulcherrimin. Figure 8 represents the pulcherrimin production of each strain. The results of the BEST8628, ppsW10 / BUSY9261, and ppsW03 / BUSY9261 strains are shown starting from the left. The vertical axis represents the pulcherrimin production (μg / mL), and the horizontal axis represents the culture time (hours). Figure 9 represents the outline of the chimeric NRPS gene. Figure 10A represents the MS / MS analysis results of the novel chimeric peptide found in BUSY9261 / chmera1. It is the result of chimeric peptide 1. Figure 10B represents the MS / MS analysis results of the novel chimeric peptide found in BUSY9261 / chmera2. It is the result of chimeric peptide 2. Figure 10C represents the MS / MS analysis results of the novel chimeric peptide found in BUSY9261 / chmera3. It is the result of chimeric peptide 3. Figure 11 represents the outline of the introduction of the SfiI recognition sequence into the surfactin synthetase. Figure 12 represents the outline of the construction of multiple NRPSs by the Combi-OGAB method. Detailed implementation mode ​​​​​​​​​​​​​

[0009] Hereinafter, the present invention will be described while showing the best mode. Throughout this specification, unless otherwise specifically mentioned, expressions in the singular form should be understood to also include the concept of their plural forms. Therefore, unless otherwise specifically mentioned, articles in the singular form (for example, "a", "an", "the", etc. in the case of English) should be understood to also include the concept of their plural forms. In addition, unless otherwise specifically mentioned, terms used in this specification should be understood to be applicable in accordance with the meanings commonly used in the art. Therefore, unless otherwise defined, all technical terms and scientific and technical terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present invention pertains. In case of contradiction, this specification (including definitions) shall prevail.

[0010] Hereinafter, the definitions and / or basic technical contents of terms specifically used in this specification will be appropriately described.

[0011] In the present specification, "non-ribosomal peptide synthetase" (also referred to as NRPS in the present specification) refers to a protein that synthesizes peptides (which may be referred to as non-ribosomal peptides (NRP)) using amino acids as substrates without mRNA. Generally, multiple molecules of NRPS protein form a complex to generate peptides, but in the present specification, NRPS may not form a protein complex. The NRPS molecule can bind to other NRPS via the COM domain located at the end to form an NRPS protein complex. In addition to the 20 essential amino acids, non-proteinogenic amino acids such as ornithine can be used as substrates for NRPS. Peptides with modifications other than amino acids such as fatty acids or sugar chains, and cyclic or branched peptides can also be synthesized by NRPS. It is known that NRPS is involved in the synthesis of peptides such as cyclosporin and vancomycin in microorganisms. Representatively, NRPS is known to be present in organisms such as the genus Bacillus (e.g., Bacillus subtilis, Bacillus amyloliquefaciens, Bacillus licheniformis), Brevibacillus brevis, Paenibacillus polymyxa, Photorhabdus asymbiotica, Photorhabdus luminescens, the genus Xenorhabdus (e.g., Xenorhabdus bovienii, Xenorhabdus budapestensis, Xenorhabdus doucetiae, Xenorhabdus indica, Xenorhabdus miraniensis, Xenorhabdus nematophila, Xenorhabdus stockiae, Xenorhabdus szentirmaii), the genus Streptomyces (e.g., Streptomyces verticillus, Streptomyces roseosporus), the genus Penicillium, the genus Aspergillus, the genus Tolypocladium (e.g., Tolypocladium inflatum, Tolypocladium niveum), the genus Fusarium, the genus Microcystis, etc.

[0012] In one embodiment, the fact that a certain protein is an NRPS can be determined by the protein containing one or more motifs selected from the group consisting of an A1 core motif sequence (L(TS)YxEL: SEQ ID NO: 44), an A2 core motif sequence (LKAGxAYL(VL)P(LI)D: SEQ ID NO: 45), an A3 core motif sequence (LAYxxYTSG(ST)TGxPKG: SEQ ID NO: 46), an A4 core motif sequence (FDxS), an A5 core motif sequence (NxYGPTE: SEQ ID NO: 47), an A6 core motif sequence (GELxIxGxG(VL)ARGYL: SEQ ID NO: 48), an A7 core motif sequence (Y(RK)TGDL: SEQ ID NO: 49), an A8 core motif sequence (GRxDxQVKIRGxRIELGEIE: SEQ ID NO: 50), an A9 core motif sequence (LPxYM(IV)P: SEQ ID NO: 51), an A10 core motif sequence (NGK(VL)DR: SEQ ID NO: 52), a T core motif sequence (DxFFxxLGG(HD)S(LI): SEQ ID NO: 53), a C1 core motif sequence (SxAQxR(LM)(WY)xL: SEQ ID NO: 54), a C2 core motif sequence (RHExLRTxF: SEQ ID NO: 55), a C3 core motif sequence (MHHxISDG(WV)S: SEQ ID NO: 56), a C4 core motif sequence (YxD(FY)AVW: SEQ ID NO: 57), a C5 core motif sequence ((IV)GxFVNT(QL)(CA)xR: SEQ ID NO: 58), a C6 core motif sequence ((HN)QD(YV)PFE: SEQ ID NO: 59), and a C7 core motif sequence (RDxSRNPL: SEQ ID NO: 60). Preferably, the NRPS contains consecutive (e.g., A8 - A10 core motif sequences) motif sequences among the A1 - A10 core motif sequences, the T core motif sequence, and the C1 - C7 core motif sequences in the same order. When containing multiple motif sequences, 1 - 3 mutations (e.g., conservative amino acid substitutions) may also be contained in the entire motif sequence. In the present specification, "non - ribosomal peptide (NRP)" refers to a peptide that can be generated or has been generated by NRPS, and also includes a peptide that can be synthesized by ribosome (composed of the natural 20 amino acids). Typically, an NRP is a peptide that cannot be synthesized by ribosome.

[0014] In this specification, the "NRPS module" is a partial structure of NRPS. As a principle, one NRPS module in NRPS is involved in the binding of one amino acid, but there are also cases where it is involved in the binding of substrates other than amino acids (such as β-amino fatty acids, β-hydroxy fatty acids, polyketides, etc.). Typically, one NRPS module contains one A domain, one T domain, and one C domain, and may also contain other domains such as TE. When the middle part of the region encoding any of the A, T, and C domains is located at the end of the nucleic acid molecule fragment, if at least a part of the regions encoding the A, T, and C domains respectively are present on one nucleic acid molecule fragment, then this nucleic acid molecule fragment contains one NRPS module. For example, one NRPS module contains domains of (A, T, and C), (T, C, and A), or (C, A, and T) from the N-terminus to the C-terminus. Typically, when encoding multiple NRPS modules on one molecule of nucleic acid, the number of regions containing the A, T, and C domains in the same order from the N-terminus to the C-terminus is regarded as one NRPS module to consider the maximum number (for example, for a nucleic acid encoding domains of A-T-C-A-T-C-A-T-C, when counting T-C-A or C-A-T as one NRPS module, the number of NRPS modules is 2, and when counting A-T-C as one NRPS module, the number of NRPS modules is 3, so the number of NRPS modules at this time is 3 。 )). In this specification, typically, one NRPS module is configured between two restriction enzyme recognition regions. A region having only one or two of the A, T, and C domains may be referred to as an incomplete NRPS module.

[0015] In this specification, the "A domain" refers to the adenylation domain of NRPS, which is considered to capture a specific amino acid and activate the amino acid through ATP, etc. The A domain is considered to have a size of about 500 - 550 amino acids and has substrate specificity for a specific amino acid. The A domain contains an A 核心 subdomain of about 400 amino acids having a binding pocket for the amino acid and an A 核心 subdomain having a lysine involved in catalytic activity and the substrate specificity for a specific amino acid. In one embodiment, the A domain may refer to the region from the 70th amino acid upstream of the A1 core motif sequence (L(TS)YxEL) to the 40th amino acid downstream of the A10 core motif sequence (NGK(VL)DR).

[0016] In this specification, the "T domain" (also referred to as the PCP domain) refers to the thiolation domain of NRPS, and the thiol moiety of the cysteine residue in this domain is considered to capture the amino acid received from the A domain. The T domain is considered to have a size of approximately 70 to 90 amino acids. In one embodiment, the T domain may refer to the region from the 20th amino acid upstream to the 80th amino acid downstream of the T core motif sequence (DxFFxxLGG(HD)S(LI)).

[0017] In this specification, the "C domain" is the condensation domain of NRPS and is considered to have the function of binding the peptide or amino acid bound to the above NRPS module to the amino acid captured by the T domain. The C domain is considered to have a size of approximately 450 amino acids. In one embodiment, the C domain may refer to the region from the 20th amino acid upstream of the C1 core motif sequence (SxAQxR(LM)(WY)xL) to the 70th amino acid downstream of the C7 core motif sequence (RDxSRNPL).

[0018] In this specification, the "TE domain" refers to the thioesterase domain of NRPS and is considered to have the function of releasing NPR from NRPS. There are also cases where the TE domain has the function of cyclizing NPR.

[0019] In this specification, the "linker region" refers to a relatively short region (for example, less than 20 amino acids in length) existing between the modules of NRPS and between the domains of NRPS. Since the linker region exists between the extensions of the modules and domains, the linker region can be determined by determining the extensions of these modules and domains, etc.

[0020] In this specification, the description of "between" two regions (for example, nucleic acid regions encoding domains), in English called "between", also includes the two regions themselves. Typically, between two regions refers to the region from the bond that bisects the number of amino acid or nucleotide residues (N) of the region constituting one (when N is even) or the amino acid or nucleotide residue (when N is odd) to the bond that bisects the number of amino acid or nucleotide residues (M) of the region constituting the other (when M is even) or the amino acid or nucleotide residue (when M is odd).

[0021] In this specification, the description of the "boundary" of two regions (for example, nucleic acid regions encoding domains), in English called "boundary", refers to the bond between two adjacent amino acid residues in a peptide or the bond between two adjacent nucleotides in a polynucleotide, and exists between the two regions. In one embodiment, the "boundary" of two regions exists between the two regions and does not exist within the two regions.

[0022] "Unnatural" base sequence refers to a region of a base sequence that contains an amino acid region different from that of the natural corresponding peptide (such as an NRPS module) that encodes the peptide to be encoded by the base sequence when compared with the natural corresponding peptide (such as an NRPS module) showing the highest identity thereto.

[0023] "Restriction enzyme recognition sequence" refers to a base sequence that is specific depending on the type of restriction enzyme and is required for cutting a nucleic acid using the restriction enzyme. The restriction enzyme recognition sequence may also contain a portion specifying any base (usually represented as "N" in a motif). In particular, there is a case where a portion that needs to be any one of G, C, T, or A among the restriction enzyme recognition sequences is called a restriction enzyme recognition specific site. For example, SfiI (motif 5'-GGCCNNNN / NGGCC-3') cuts the site sandwiched between two restriction enzyme specific recognition sites, and AarI (motif 5'-CACCTGCNNNN / NNNN-3') cuts the site separated from the restriction enzyme specific recognition site.

[0024] When used in this specification, "plasmid" refers to a circular DNA that exists separately from the chromosome in a cell or exists separately from the chromosome when introduced into a cell.

[0025] As used herein, the term "nucleic acid sequence that promotes plasmid replication" refers to any nucleic acid sequence that, when introduced into a host cell (e.g., Bacillus subtilis), promotes the replication (which is the same as amplification in this context) of a plasmid present in the host cell. The nucleic acid sequence that promotes the replication of the plasmid is preferably operably linked to the nucleic acid sequence encoding the subject plasmid, but is not limited thereto. Details such as the nucleic acid sequence that promotes plasmid replication, the subject plasmid, and their relationship are described elsewhere in this specification. Examples of the nucleic acid sequence that promotes plasmid replication include a nucleic acid sequence containing an origin of replication that functions in the host cell of interest (e.g., Bacillus subtilis). For example, a nucleic acid sequence that promotes plasmid replication in Bacillus subtilis has the ability to promote plasmid replication in Bacillus subtilis, but may further have the ability to promote plasmid replication in microorganisms other than Bacillus subtilis (e.g., Escherichia coli). The nucleic acid sequence that promotes plasmid replication in Bacillus subtilis and other organisms (e.g., Escherichia coli) may be a nucleic acid sequence in which the same region is used to promote plasmid replication in both Bacillus subtilis and other organisms (e.g., Escherichia coli), or may be sequences located at separate positions. For example, plasmids pSTK1 (Issay Narumt et al., BIOTECHNOLOGY LETTERS, Volume 17 No.5 (May 1995) pp.475-480) and pBS195 (Molekuliarnaia Genetika et al., 01 May 1991, (5):26-29, PMID:1896058) have been shown to replicate in the same replication origin region in both Bacillus subtilis and Escherichia coli, and such nucleic acid sequences can be used as nucleic acid sequences that promote plasmid replication in both Bacillus subtilis and Escherichia coli. In addition, a shuttle plasmid prepared by ligating BR322 (an Escherichia coli plasmid) and pC194 (a Bacillus subtilis plasmid) (P Trieu-Cuot et al., EMBO J.1985 Dec 16;4(13A):3583-3587.) and the plasmid of the B. subtilis Secretory Protein Expression System containing pUBori and ColEl ori provided by Takara Bio (Shiga) can replicate in both Bacillus subtilis and Escherichia coli, and such nucleic acid sequences located separately can also be used as nucleic acid sequences that promote plasmid replication in both Bacillus subtilis and Escherichia coli.

[0026] As the replication mechanisms of nucleic acids (plasmids) in Bacillus subtilis, rolling circle type, theta type, oriC, replication using phages, etc. can be cited. As the sequences replicated in Bacillus subtilis, sequences using any mechanism can also be used. The rolling circle type replication mechanism is a mechanism in which single-stranded replication of one side of double-stranded DNA is carried out first and then replication of the other strand is carried out. Since the nucleic acid exists as single-stranded DNA for a long time, plasmids tend to be unstable. The theta type replication mechanism is a mechanism in which double-stranded DNA (plasmid) replication starts simultaneously in two directions from the replication origin in the same way as when replicating the bacterial chromosome. It is preferably used when replicating long-chain DNA (for example, 10 kb or more) as in the present invention (Janniere, L, A. Gruss, and S.D. Ehrlich. 1993. Plasmids, p. 625-644. In A.L. Sonenshein, J.A.Hoch, and R.Losick (ed.), Bacillus subtilis and other gram-positive bacteria: biochemistry, physiology and molecular genetics. American Society for Microbiology, Washington, D.C.). The oriC replication mechanism operates in the same way as when replicating the chromosome of the host bacterium (for example, Bacillus subtilis). As the nucleic acid sequences that promote plasmid replication in Bacillus subtilis, plasmids or their parts or replication origins or modified forms thereof that are known to carry out replication mechanisms such as rolling circle type, theta type, oriC, and using phages can be cited. As rolling circle type plasmids, pUB110, pC194, pE194, pT181, etc. are known; as theta type plasmids, pAMβ1, pTB19, pLS32, pLS20, etc. are known.

[0027] The nucleic acid sequence that promotes plasmid replication in Bacillus subtilis may have a sequence identical or similar to the known replication origin (e.g., sequence identity of 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more). For example, after introducing a DNA fragment ligating a candidate DNA fragment and a selection marker gene effective in Bacillus subtilis such as a drug resistance gene into Bacillus subtilis, and after culturing (by adding reagents, etc.), when it is replicated extrachromosomally, it can be determined that the candidate DNA fragment has a nucleic acid sequence that promotes plasmid replication in Bacillus subtilis. The nucleic acid sequence that promotes plasmid amplification is a DNA fragment having a replication origin, referring to a replicable DNA fragment independent of chromosomal DNA. It can be confirmed whether these DNA fragments are replicable by ligating a selection marker gene such as a drug resistance gene to these DNA fragments, culturing under selection conditions using reagents, etc., purifying the plasmid DNA, and observing the DNA bands by electrophoresis. In addition to the replication origin, the plasmid may also contain a partitioning mechanism gene for effectively distributing the plasmid in the mother cell, a selection marker gene, etc.

[0028] As used in this specification, the "replication origin" refers to a fragment of a nucleic acid sequence where, by binding of a protein that recognizes the nucleic acid sequence (such as the initiator DnaA protein, etc.) or synthesizing RNA, a part of the DNA double helix is unwound and replication begins.

[0029] In this specification, a "transformed organism" refers to an organism that can be used in the OGAB method, can take up non-circular long-chain nucleic acids, and can generate plasmids. As a transformed organism, a microorganism that spontaneously takes up nucleic acids can be used. A representative transformed organism is Bacillus subtilis. In this specification, "Bacillus subtilis" is an aerobic, Gram-positive, catalase-positive bacterium that is commonly present in soil, plants, the gastrointestinal tracts of ruminants and humans, with a size of 0.7 - 0.8 × 2 - 3 μm, mesophilic, and an optimal cultivation temperature of 25 - 40°C. It forms spores and includes any non-pathogenic bacterium with natural transformation ability, such as the scientific name Bacillus subtilis or Bacillus amyloliquefaciens, Bacillus licheniformis, Bacillus pumilus, etc., which are Bacillus species closely related to it. In a preferred embodiment, the Bacillus subtilis is strain Marburg 168 or its derivative strain RM125, which have high natural transformation ability among Bacillus subtilis strains. Strain 168 is the Gram-positive bacterium most studied genetically and can be obtained from ATCC (American Type Culture Collection) etc. together with strain RM125. Since it is judged to be harmless to humans and livestock and can adapt to methods such as the OGAB method, it can also be preferably used in the present invention.

[0031] In this specification, "repeated sequence" or "tandem repeat" (of nucleic acid sequences) is a general term for sequences in which the same sequence is repeatedly (especially several times or more) observed in the nucleic acid sequences of a biological genome. Any repeated sequence used in this field can be used in the present invention. Typically, promoter sequences occur repeatedly.

[0032] When used in this specification, "host cell" refers to a cell (including the progeny of such a cell) into which foreign nucleic acid, protein, or virus has been introduced. In this specification, "protein", "polypeptide", and "peptide" are used with the same meaning and refer to polymers of amino acids of any length. The polymer can be linear, branched, or cyclic. The amino acids can be natural, non-natural, or modified amino acids. This term also includes naturally or artificially modified polymers. Such modifications include, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other operation or modification (such as conjugation with a labeling component).

[0033] In this specification, "polynucleotide" and "nucleic acid" are used interchangeably and refer to polymers of nucleotides of any length. Examples of nucleic acids include: DNA, RNA, cDNA, mRNA, rRNA, tRNA, microRNA (miRNA), lncRNA. This term also encompasses "polynucleotide derivatives". A "polynucleotide derivative" refers to a polynucleotide containing nucleotide derivatives or internucleotide linkages that are different from the normal ones. A "nucleotide derivative" refers to a nucleotide having a structure different from the normal nucleotides used in natural DNA or RNA. Examples include: locked nucleic acid (LNA), 2′-O,4'-C-ethylene bridged nucleic acid (ENA) and other ethyl-bridged nucleic acids, other bridged nucleic acids (BNA), hexitol nucleic acid (HNA), amido-bridged nucleic acid (AmNA), morpholino nucleic acid, tricyclic-DNA (tcDNA), polyether nucleic acid (see, for example, U.S. Patent No. 5,908,845), cyclohexene nucleic acid (CeNA), etc. Examples of internucleotide linkages different from the normal ones include, for example: an internucleotide linkage in which a phosphodiester bond is converted to a phosphorothioate bond, an internucleotide linkage in which a phosphodiester bond is converted to an N3′-P5′ aminophosphonate bond, an internucleotide linkage in which a ribose and a phosphodiester bond are converted to a peptide nucleic acid bond, etc.

[0034] Unless otherwise indicated to the contrary, a particular nucleic acid sequence is conceived to include its conservatively modified variants (e.g., degenerate codon substitutions) and complementary sequences as if explicitly shown. Specifically, degenerate codon substitutions can be achieved by generating a sequence in which the third position of one or more selected (or all) codons is substituted with a mixture of bases and / or deoxyinosine residues. For example, variants of a specific wild-type sequence such as the A domain of a particular NRPS also include nucleic acids containing sequences having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% sequence identity to the original sequence.

[0035] In this specification, "gene" refers to a nucleic acid portion that can achieve a certain biological function. As this biological function, the following can be cited: encoding a polypeptide or protein, encoding a protein non-coding functional RNA (rRNA, tRNA, microRNA (miRNA), lncRNA, etc.), specifically binding to a specific protein, controlling the cleavage and replication of nucleic acids. Therefore, in this specification, a gene includes, in addition to the nucleic acid portion encoding a protein or a protein non-coding functional RNA, a promoter, a terminator, an enhancer, an insulator, a silencer, and an origin of replication. In this specification, "gene product" can refer to a polypeptide, a protein, or a protein non-coding functional RNA encoded by a gene.

[0036] When used in this specification, a "deleted" gene means that the nucleic acid does not contain the gene or contains a gene modified so as not to exhibit the normal function of the gene (for example, the function of producing a functional protein).

[0037] When used in this specification, "operably linked" means being disposed under the control of a transcriptional and translational regulatory sequence (such as a promoter, an enhancer, etc.) or a translational regulatory sequence for the expression (operation) of a desired sequence. For a promoter to be operably linked to a gene, the promoter is usually disposed immediately upstream of the gene, but it is not necessarily required to be adjacently disposed.

[0038] In this specification, the "homology" of nucleic acids refers to the degree of relative identity of two or more nucleic acid sequences. Generally, having "homology" means a high degree of identity or similarity. Therefore, the higher the homology of two nucleic acids, the higher the identity or similarity of these sequences. "Similarity" is a value calculated by also including similar bases in addition to identity. Similar bases refer to cases where a part is identical in degenerate bases (for example, R = A + G, M = A + C, W = A + T, S = C + G, Y = C + T, K = G + T, H = A + T + C, B = G + T + C, D = G + A + T, V = A + C + G, N = A + C + G + T). Whether two nucleic acids have homology can be investigated by direct comparison of the sequences or by a hybridization method under stringent conditions. When directly comparing two nucleic acid sequences, when they are typically at least 50% identical, preferably at least 70% identical, more preferably at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical between the nucleic acid sequences, these genes have homology.

[0039] Amino acids can be referred to in this specification by either their commonly known three-letter symbols or the single-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleic acids can similarly be referred to by the commonly known single-letter codes. In this specification, comparisons of the similarity, identity, and homology of amino acid sequences and base sequences can be calculated using BLAST as a sequence analysis tool with default parameters. Retrieval of identity can be performed, for example, using BLAST 2.10.1+ (published on June 18, 2020) of NCBI. The value of identity in this specification refers to the value obtained when performing sequence alignment under default conditions using the above-mentioned BLAST. However, when a higher value appears by changing the parameters, the highest value is taken as the value of identity. When evaluating identity for a duplicated region, the highest value among them is taken as the value of identity. Similarity is a value that takes into account similar amino acids in addition to identity.

[0040] In this specification, "conservative amino acid substitution" means substituting an amino acid in a peptide with another amino acid with similar properties, and it can be predicted that the properties of the peptide are similar before and after the conservative substitution. In one embodiment, conservative amino acid substitutions can use other amino acids within the same group to substitute the amino acids in the following amino acid groups. · Group 1: Glycine, Alanine, Valine, Leucine, and Isoleucine · Group 2: Serine and Threonine · Group 3: Phenylalanine, Tyrosine, and Tryptophan · Group 4: Lysine, Hydroxylysine, Arginine, and Histidine · Group 5: Cysteine, Methionine, Methionine Sulfoxide, and Homocysteine · Group 6: Proline and Hydroxyproline · Group 7: Glutamic Acid and Aspartic Acid · Group 8: Glutamine and Asparagine

[0041] In this specification, unless otherwise specified, references to a biological substance (such as a protein, nucleic acid, gene) are to be understood as also referring to variants of that biological substance (such as variants with modifications in the amino acid sequence) that perform the same function (which may not be to the same extent) as the biological function of that biological substance. Among such variants, when compared to the original molecular sequence being aligned, over a fragment of the original molecule, the amino acid sequence or nucleic acid sequence of the original biological substance of the same size, or through computer homology programs well-known in the art, it may include molecules that are at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% identical. Among the variants, it may include molecules with modified amino acids (such as modifications caused by disulfide bond formation, glycosylation, lipidation, acetylation or phosphorylation) or modified nucleotides (such as modifications caused by methylation).

[0042] In this specification, "stringent conditions" refer to conditions well-known and commonly used in the art. Stringent conditions can be, for example, the following conditions. (1) For washing, use low ionic strength and high temperature (for example, at 50 °C, using 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate), (2) use modifiers such as formamide in hybridization (for example, at 42 °C, using 50% (v / v) formamide and 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer at pH 6.5 and 750 mM sodium chloride, 75 mM sodium citrate) or (3) incubate overnight at 37 °C in a solution containing 20% formamide, 5×SSC, 50 mM sodium phosphate (pH 7.6), 5×Denhardt's Solution, 10% dextran sulfate, and 20 mg / ml of sheared and modified salmon sperm DNA, and then wash the filter membrane at about 37 - 50 °C with 1×SSC. Additionally, the formamide concentration can also be 50% or higher. The washing time can be 5, 15, 30, 60 or 120 minutes or can also be longer than these times. As factors affecting the stringency of the hybridization reaction, various factors such as temperature and salt concentration can be considered, and for detailed information, reference can be made to Ausubel et al., Current Protocols in Molecular Biology, Wiley Interscience Publishers, (1995). Examples of "highly stringent conditions" are 0.0015 M sodium chloride, 0.0015 M sodium citrate, at 65 - 68 °C; or 0.015 M sodium chloride, 0.0015 M sodium citrate, and 50% formamide, at 42 °C. Hybridization can be carried out according to the methods described in experimental books such as Molecular Cloning 2nd ed., Current Protocols in Molecular Biology, Supplement 1 - 38, DNA Cloning 1: Core Techniques, A Practical Approach, Second Edition, Oxford University Press (1995), etc. Among them, sequences hybridized under stringent conditions preferably exclude sequences containing only A sequences or only T sequences. Moderately stringent conditions, for example, can be easily determined by those skilled in the art based on the length of DNA, as shown in Sambrook et al., Molecular Cloning: A Laboratory Manual, No. 3, Vol. 1, 7.42 - 7.45 Cold Spring Harbor Laboratory Press, 2001, and regarding nitrocellulose filters, including a pre-washing solution using 5×SSC, 0.5% SDS, 1.0 mM EDTA (pH 8.0), at about 40 - 50 °C, hybridization conditions of about 50% formamide, 2×SSC - 6×SSC (or other equivalent hybridization solutions such as Stark's solution in about 50% formamide at about 42 °C), and washing conditions of 0.5×SSC, 0.1% SDS at about 60 °C. Therefore, the polypeptides used in the present invention also include polypeptides encoded by nucleic acid molecules that hybridize to the nucleic acid molecules encoding the polypeptides specifically described in the present invention under highly or moderately stringent conditions.

[0043] In this specification, a "corresponding" amino acid or nucleic acid refers to an amino acid or nucleic acid that has or is predicted to have the same function as a specified amino acid or nucleotide in a polypeptide or polynucleotide molecule used as a comparison reference in a certain polypeptide molecule or polynucleotide molecule. Particularly in the case of an enzyme molecule, it refers to an amino acid that is present at the same position in the active site and equally contributes to the catalytic activity. For example, in the case of an antisense strand molecule, it can be the same part in an ortholog corresponding to a specific part of the antisense strand molecule. A corresponding amino acid can be, for example, a specific amino acid that has been cysteinylated, glutathionylated, formed an S-S bond, oxidized (such as oxidation of the methionine side chain), formylated, acetylated, phosphorylated, glycosylated, tetradecylated, etc. Or a corresponding amino acid can be an amino acid that undertakes dimerization. Such a "corresponding" amino acid or nucleic acid can be over a certain range of regions or domains. Therefore, in such a case, the region or domain referred to as "corresponding" in this specification. As a result of alignment between two sequences, a nucleic acid or amino acid located at the same position can be a corresponding nucleic acid or amino acid.

[0044] In this specification, a "corresponding" gene (such as a polynucleotide sequence or molecule) refers to a gene (such as a polynucleotide sequence or molecule) that has or is predicted to have the same function as a specified gene in a species used as a comparison reference in a certain species. When there are multiple genes having such a function, it refers to genes having the same evolutionary origin. Therefore, a gene corresponding to a certain gene can be an ortholog of that gene. For example, a gene corresponding to an NRPS module related to the binding of a specific amino acid in a certain bacterium can be found by using the gene sequence of the NRPS module that serves as the reference for the corresponding gene as a query sequence to search the sequence database of the target bacterium.

[0045] According to the present invention, in this specification, the term "activity" refers to the function of a molecule in the broadest sense. The activity is not intended to be limiting, but generally includes the biological function, biochemical function, physical function, or chemical function of the molecule. The activity includes, for example, enzyme activity, the ability to interact with other molecules, or the ability to activate, promote, stabilize, hinder, inhibit, or destabilize the functions of other molecules, stability, and the ability to be confined to a specific intracellular location. When applicable, the term also relates to the function of a protein complex in the broadest sense.

[0046] In this specification, "biological function", when referring to a certain gene or a nucleic acid molecule or polypeptide related thereto, means a specific function that the gene, nucleic acid molecule or polypeptide can have in a living organism. For this, for example, specific cell surface structure recognition ability, enzyme activity, the ability to bind to a specific protein, etc. can be cited, but it is not limited to these. In the present invention, for example, the function that a certain promoter is recognized in a specific host cell can be cited, but it is not limited thereto. In this specification, the biological function can be exerted through "biological activity". In this specification, "biological activity" means the activity that a certain factor (such as a polynucleotide, protein, etc.) can have in a living organism, including the activity of exerting various functions (such as transcription promoting activity), and for example, also includes the activity of activating or inactivating other molecules due to the interaction with a certain molecule. For example, when a certain factor is an enzyme, its biological activity includes its enzyme activity. In other examples, when a certain factor is a ligand, it includes the binding to the receptor corresponding to the ligand. Such biological activity can be measured by techniques well-known in the art. Therefore, "activity" means various measurable indicators that show binding or definite binding (either direct or indirect); or affect the response (that is, have a measurable effect on a slight exposure or stimulus response), for example, the amount of upstream or downstream proteins in a host cell or other similar functional scales can be cited.

[0047] When used in this specification, unless otherwise specifically mentioned, the terms "transformation", "transduction" and "transfection" can be used interchangeably, and refer to the introduction of nucleic acid into a host cell (optionally via a virus or a virus-derived construct). As a transformation method, any method can be used as long as it is a method for introducing nucleic acid into a host cell. For example, various well-known techniques such as the use of competent cells, electroporation, the method using a particle gun, and the calcium phosphate method can be cited.

[0048] In this specification, a "purified" substance or biological factor (such as a nucleic acid or a protein, etc.) means that at least a part of the factors that are naturally associated with the biological factor is removed. Therefore, generally, the purity of the biological factor in the purified biological factor is higher than the state in which the biological factor usually exists (that is, it has been concentrated). The term "purified" used in this specification means that the homologous biological factor preferably exists at least 75% by weight, more preferably at least 85% by weight, still more preferably at least 95% by weight, and then most preferably at least 98% by weight. The substances used in the present invention are preferably "purified" substances.

[0049] In this specification, the terms "reagent", "agent" or "factor" (which are all equivalent to "agent" in English) can be used interchangeably in a broad sense, and can be any substance or other element (such as energy sources like light, radiation, heat, electricity, etc.) as long as it can achieve the intended purpose. As such substances, for example, proteins, polypeptides, oligopeptides, peptides, polynucleotides, oligonucleotides, nucleotides, nucleic acids (such as DNA including cDNA and genomic DNA, RNA such as mRNA), polysaccharides, oligosaccharides, lipids, organic low-molecules (such as hormones, ligands, signaling substances, organic low-molecules, molecules synthesized by combinatorial chemistry, low-molecules that can be used as drugs (such as low-molecular ligands, etc.)), composite molecules of these, and mixtures of these can be mentioned, but are not limited to these.

[0050] When used in this specification, the term "complex" or "composite molecule" refers to any construct containing two or more parts. For example, when one part is a polypeptide, the other part can be a polypeptide or a substance other than that (such as a substrate, sugar, lipid, nucleic acid, other hydrocarbons, etc.). In this specification, two or more parts constituting the complex can be bonded by covalent bonds or by other bonds (such as hydrogen bonds, ionic bonds, hydrophobic interactions, van der Waals forces, etc.).

[0051] In this specification, a "kit" refers to a unit that is usually divided into two or more compartments and provides the parts to be provided (such as constructs derived from NRPS, instructions, etc.). For the sake of stability, etc., when it is intended to provide a composition that is preferably mixed just before use rather than being provided as a mixture, the form of such a kit is preferred. Such a kit preferably facilitates the provision of instructions or a manual describing how to use the provided parts or how to handle the reagents. When used as a kit in this specification, the kit usually includes instructions, etc., describing the usage method of plasmids, etc.

[0052] The term "about" means plus or minus 10% of the indicated value. When "about" is used with respect to "temperature", it means plus or minus 5 °C of the indicated temperature; when "about" is used with respect to "pH", it means plus or minus 0.5 of the indicated pH.

[0053] (Preferred embodiments) The preferred embodiments of the present invention are described below. The embodiments provided below can be understood to be provided for better understanding of the present invention, and the scope of the present invention should not be limited to the following description. Therefore, it is obvious that those skilled in the art can make appropriate changes within the scope of the present invention with reference to the description in this specification. In addition, it can be understood that the following embodiments can be used alone or in combination.

[0054] (Modifying NRPS) In one aspect, the present invention provides a nucleic acid encoding a modified NRPS comprising one or more non-natural restriction enzyme recognition sequences. In one embodiment, the present invention provides a modified NRPS encoded by such a nucleic acid. The inventors introduced a restriction enzyme recognition sequence into the nucleic acid encoding NRPS to produce a modified NRPS, and found that the modified NRPS could produce the target NRP in the same manner as the original NRPS. By introducing the restriction enzyme recognition sequence, the modification of NRPS becomes easier, and various NRPSs can be produced.

[0055] The modified NRPS may contain one or more NRPS modules. Each NRPS module can independently be an NRPS from organisms such as Bacillus (e.g., Bacillus subtilis, Bacillus amyloliquefaciens, Bacillus licheniformis), Brevibacillus brevis, Paenibacillus polymyxa, Photorhabdus asymbiotiea, Photorhabdus luminescens, Xenorhabdus (e.g., Xenorhabdus bovienii, Xenorhabdus budapestensis, X. doucetiae, Xenorhabdus indica, X. miraniensis, Xenorhabdus nematophila, X. stockiae, X. szentirmaii), Streptomyces (e.g., Streptomyces verticillus, Streptomyces roseosporus), Penicillium, Aspergillus, Tolypocladium (e.g., Tolypocladium inflatum, Tolypocladium niveum), Fusarium, Microcystis, etc. For example, one or more NRPS modules in the modified NRPS may have at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% amino acid sequence identity with a natural (from a specific species or strain) NRPS module.

[0056] Since one species or strain may have multiple NRPS modules, the modified NRPS of the present invention may also contain multiple NRPS modules from one species or strain in a specific order.

[0057] In one embodiment, each domain of the NRPS module (A domain, T domain, C domain, etc.) can also be replaced with the corresponding domain of an NRPS module from another organism (species or strain) related to the linkage of the same type of amino acid.

[0058] In one embodiment, in addition to the A domain, T domain, and C domain of the NRPS module, the modified NRPS of the present invention may also contain other domains such as the COM domain, TE domain, E (epimerization) domain, M (methylation) domain, R (reduction) domain, F (formylation) domain, Cy (cyclization) domain, Ox (oxidation) domain, KR (ketoacyl reductase) domain, MOx (monooxygenase) domain (for example, refer to Angew Chem Int Ed Engl. 2017 Mar 27; 56(14): 3770 - 3821.), etc.

[0059] In one embodiment, the modified NRPS of the present invention may also contain a desired number of NRPS modules, for example, it may contain 3 to 100 NRPS modules, such as 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 NRPS modules.

[0060] In one embodiment, the nucleic acid encoding the modified NRPS of the present invention may contain non - natural restriction enzyme recognition sequences between nucleic acid sequences encoding NRPS modules from different organisms (species or strains). In one embodiment, the nucleic acid encoding the modified NRPS of the present invention may contain non - natural restriction enzyme recognition sequences between nucleic acid sequences encoding NRPS modules that are not naturally adjacent.

[0061] (Restriction enzyme recognition sequence) In one embodiment, the nucleic acid encoding the NRPS of the present invention may contain one or more non - natural restriction enzyme recognition sequences. In the introduction of non - natural restriction enzyme recognition sequences, insertions, deletions, and substitutions of bases can be used in any combination. The non - natural restriction enzyme recognition sequences are introduced without forming a stop codon and without deviation of the codon reading frame. In one embodiment, the non - natural restriction enzyme recognition sequences are designed to encode a small number of amino acids (such as alanine, glycine, etc.).

[0062] In one embodiment, in the alignment of the NRPS module, a non-natural restriction enzyme recognition sequence is designed outside the consensus sequence portion imported to high homology. In one embodiment, when obtaining and aligning the entire amino acid sequence of the region from the A domain to the C domain of the NRPS module possessed by a certain biological strain, a non-natural restriction enzyme recognition sequence is imported into a region with an amino acid length of 3, 4, 5, 6, 7, 8, 9, or 10 amino acids at amino acid positions where the proportion of the same amino acids present in the corresponding amino acid positions does not exceed 30%, 40%, 50%, 60%, 70%, 80%, or 90%.

[0063] In one embodiment, the non-natural restriction enzyme recognition sequence can be imported into the portion where the amino acid specified by the capital letter amino acid can be imported without changing the restriction enzyme recognition sequence in each amino acid sequence of the A1 core motif sequence (L(TS)YxEL), A2 core motif sequence (LKAGxAYL(VL)P(LI)D), A3 core motif sequence (LAYxxYTSG(ST)TGxPKG), A4 core motif sequence (FDxS), A5 core motif sequence (NxYGPTE), A6 core motif sequence (GELxIxGxG(VL)ARGYL), A7 core motif sequence (Y(RK)TGDL), A8 core motif sequence (GRxDxQVKIRGxRIELGEIE), A9 core motif sequence (LPxYM(IV)P), A10 core motif sequence (NGK(VL)DR), T core motif sequence (DxFFxxLGG(HD)S(LI)), C1 core motif sequence (SxAQxR(LM)(WY)xL), C2 core motif sequence (RHExLRTxF), C3 core motif sequence (MHHxISDG(WV)S), C4 core motif sequence (YxD(FY)AVW), C5 core motif sequence ((IV)GxFVNT(QL)(CA)xR), C6 core motif sequence ((HN)QD(YV)PFE), C7 core motif sequence (RDxSRNPL), or into a portion of any sequence that does not belong to these core motif sequences.

[0064] In one embodiment, the non-natural restriction enzyme recognition sequence is located between the base sequences encoding adjacent NRPS modules. In one embodiment, the non-natural restriction enzyme recognition sequence is located between the C domain coding region and the A domain coding region of the NRPS module. In one embodiment, the non-natural restriction enzyme recognition sequence is located in a region that does not contain the two domain coding regions (such as the C domain coding region and the A domain coding region) between the coding regions of two of the A, T, and C domains of the NRPS module.

[0065] In one embodiment, the non-natural restriction enzyme recognition sequence is a sequence that is cleaved at the site flanked by two restriction enzyme specific recognition sites. For example, it is the recognition sequence of Type II restriction enzymes such as SfiI (5’-GGCCNNNN / NGGCC-3’), BglI (5’-GCCNNNN / NGGC-3’), AlwNI (5’-CAGNN / CTG-3’), DraIII (5’-CACNNN / GTG-3’), PflMI (5’-CCANNNN / NTGG-3’), BstAPI (5’-GCANNNN / NTGC-3’), etc. In one embodiment, the non-natural restriction enzyme recognition sequence is a sequence that is cleaved at the site separated from the restriction enzyme specific recognition site. For example, it is the recognition sequence of Type IIS restriction enzymes such as AarI (5’-CACCTGCNNNN / NNNN-3’), BbsI (5’-GAAGACNN / NNNN-3’), BsaI (5’-GGTCTCN / NNNN-3’), BsmBI (5’-CGTCTCN / NNNN-3’), BspQI (5’-GCTCTTCN / NNN-3’), etc. Since the restriction enzyme specific recognition site is different from the cleavage site, the overhang sequences of the nucleic acid fragments generated by restriction enzyme treatment become variable, and multiple overhang sequences can be formed by the treatment with one restriction enzyme. Therefore, Type II or Type IIS restriction enzymes can be preferably used in methods that utilize multiple overhang sequences (such as the OGAB method).

[0066] In one embodiment, for the non-natural restriction enzyme recognition sequence, when a nucleic acid encoding an NRPS containing it is treated with a restriction enzyme that recognizes the restriction enzyme recognition sequence, it can be introduced such that the multiple fragmented nucleic acid fragments do not have the same overhang sequence.

[0067] Any type of restriction enzyme recognition sequence can be introduced into the introduction site of any of the above non-natural restriction enzyme recognition sequences.

[0068] (Production of a novel NRPS using a restriction enzyme recognition sequence) In one aspect, the present invention provides a method for producing a product nucleic acid containing a nucleic acid sequence encoding a non-ribosomal peptide synthetase (NRPS), which includes the steps of preparing a nucleic acid containing a non-natural restriction enzyme recognition sequence as a starting nucleic acid, treating the nucleic acid with a restriction enzyme that recognizes the restriction enzyme recognition sequence to prepare fragmented nucleic acid fragments, and ligating the fragmented nucleic acid fragments to form a product nucleic acid different from the starting nucleic acid.

[0069] In one embodiment, in the step of ligating fragmented nucleic acid segments, the reaction mixture (solution, etc.) for ligation contains a plurality of fragmented nucleic acid segments. The plurality of fragmented nucleic acid segments have a plurality of overhang sequences, and at least a part of the overhang sequences are generated by cleavage of non-natural restriction enzyme recognition sequences. Among the plurality of overhang sequences, if there are complementary overhang sequence pairs, the fragmented nucleic acid segments having complementary overhang sequences are preferentially ligated. Therefore, when there are a plurality of fragmented nucleic acid segments having the same type of overhang sequence, it is random which fragmented nucleic acid segment is introduced at a specified position according to the type of the overhang sequence, and it is possible to easily obtain a modified NRPS that replaces a specific region (for example, one NRPS module) of NRPS.

[0070] In one embodiment, the plurality of fragmented nucleic acid segments include a group of fragmented nucleic acid segments having the same overhang sequence and containing base sequences encoding a plurality of NRPS modules. In one embodiment, the plurality of fragmented nucleic acid segments include a group of fragmented nucleic acid segments having the same overhang sequence and containing base sequences encoding a plurality of NRPS modules that capture amino acids. In one embodiment, each of the fragmented nucleic acid segments has an overhang sequence with a length of 3 to 5 bases.

[0071] In NRPS, one NRPS module is involved in the ligation of one amino acid. Therefore, by having an integer number of NRPS modules (the same number of A domains, T domains, and C domains) in the fragmented nucleic acid segments, the design of the NRP produced by the modified NRPS thus produced becomes easy. In one embodiment, each of the fragmented nucleic acid segments independently contains 0 to 5 (0, 1, 2, 3, 4, or 5) NRPS modules. In one embodiment, each of the fragmented nucleic acid segments contains at most one NRPS module.

[0072] In one embodiment, in the reaction mixture for ligation, a group of fragmented nucleic acid segments having the same overhang sequence and encoding at least a part of the NRPS module is preferably cleaved at the corresponding position of the NRPS module. The so-called corresponding position can be, for example, a position that deviates by no more than 20 amino acids (for example, no more than 15, 10, 5, 4, 3, 2, or 1 amino acid) when aligning the amino acid sequences of NRPS modules such as the A-T linker region or a specific region of the A domain of each NRPS module. For example, when there is a fragment cleaved in the A domain coding region and a fragment cleaved in the T domain coding region among the fragmented nucleic acid segments having the same overhang sequence, continuous occurrence or deletion of the A domain coding region may occur in the NRPS-encoding nucleic acid obtained by ligation, and thus the generation of NRP may fail.

[0073] In one embodiment, the fragmented nucleic acid may also contain domains such as the COM domain, TE domain, E (epimerization) domain, M (methylation) domain, R (reduction) domain, F (formylation) domain, Cy (cyclization) domain, Ox (oxidation) domain, KR (ketoacyl reductase) domain, MOx (monooxygenase) domain of the NRPS module (for example, refer to Angew Chem Int Ed Engl. 2017 Mar 27; 56(14): 3770 - 3821, etc.). For these domains, those skilled in the art can appropriately select according to the types of amino acids to be introduced, for example.

[0074] In one embodiment, when a nucleic acid encoding an NRPS module is treated with a restriction enzyme, the nucleic acid may also contain a natural restriction enzyme recognition sequence based on the restriction enzyme. In this case, the protruding sequence of the nucleic acid fragment generated by treating the nucleic acid with the restriction enzyme is preferably different from those of other nucleic acid fragments to be mixed with the nucleic acid fragment. If there are only nucleic acid fragments with the same protruding sequence, the nucleic acid fragments generated by cleaving the natural restriction enzyme recognition sequence can be restored to the state before cleavage by ligation.

[0075] In one embodiment, when connecting NRPS modules that are not naturally adjacent, it is necessary to connect them in a form that does not damage the function of the peptide chain possessed by each stretched NRPS module. For example, since the C domain specifies the connected amino acids, it is necessary not to disrupt this function.

[0076] In one embodiment, the present invention provides a method for producing a generation plasmid containing a nucleic acid sequence encoding a non - ribosomal peptide synthetase (NRPS), which includes the steps of preparing a starting plasmid set, treating the starting plasmid set with a restriction enzyme to prepare a mixture containing a plurality of resulting fragmented nucleic acids, ligating the plurality of fragmented nucleic acids to form a ligated nucleic acid, and contacting the ligated nucleic acid with a transformation organism to form a generation plasmid. Each of the starting plasmid sets contains a base sequence encoding at least 1 NRPS module and at least 1 restriction enzyme recognition sequence. In one embodiment, the mixture containing the fragmented nucleic acids is a solution.

[0077] In one embodiment, each starting plasmid independently contains the same or different selection marker sequences. In one embodiment, the selection marker sequence includes a drug - resistance marker sequence. In one embodiment, each starting plasmid independently contains the same or different replication origins that can function in Bacillus subtilis. In one embodiment, each starting plasmid contains a promoter region upstream of the base sequence encoding NRPS that is independently the same or different.

[0078] In one embodiment, the starting plasmid contains a nucleic acid sequence that promotes plasmid replication in a transformed organism (such as Bacillus subtilis). In one embodiment, the nucleic acid sequence that promotes plasmid replication in Bacillus subtilis may contain an origin of replication that functions in Bacillus subtilis, such as the origins of replication contained in plasmids such as oriC and pTB19 (Imanaka, T., et al. J. Gen. Microbiol. 130, 1399-1408. (1984)), pLS32 (Tanaka, T and Ogra, M. FEBS Lett. 422, 243-246. (1998)), pAMβ1 (Swinfield, T. J., et al. Gene 87, 79-90. (1990)), etc. As the nucleic acid sequence that promotes plasmid replication in Bacillus subtilis, there can be mentioned: plasmids or parts thereof or origins of replication or modified forms of these that are known to operate replication mechanisms such as rolling circle type, theta type, oriC, phage, etc. In one embodiment, the nucleic acid sequence that promotes plasmid replication in Bacillus subtilis may have a sequence that is the same as or similar to a known origin of replication (for example, a sequence identity of 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more).

[0079] In one embodiment, the starting plasmid may have a promoter and / or enhancer that functions in Bacillus subtilis. For example, as a promoter of Bacillus subtilis, there can be mentioned Pspac (Yansura, D. and Henner, D. J. Pro. Natl. Acad. Sci, USA 81, 439-443. (1984.)) or Pr promoter (Itaya, M. Biosci. Biotechnol. Biochem. 63, 602-604. (1999)) that can control expression with IPTG (isopropyl-β-D-thiogalactopyranoside), etc. The nucleic acid elements that function in Bacillus subtilis do not need to be derived from Bacillus subtilis, and highly functional nucleic acid elements, etc., can be selected.

[0080] In one embodiment, the present invention provides a method for producing a NRPS-derived construction plasmid of the present invention, including a step of operably linking a plurality of nucleic acid sequences. In one embodiment, since Bacillus subtilis may have the ability to form a plasmid from nucleic acids taken up from the outside, in this method, the introduced nucleic acid may not be a plasmid either. For example, when Bacillus subtilis comes into contact with a nucleic acid (such as a non-circular nucleic acid having a tandemly repeated nucleic acid sequence), it can take up the nucleic acid and form a plasmid within Bacillus subtilis (for example, refer to the OGAB method described in this specification).

[0081] In one embodiment, the method for introducing nucleic acid into Bacillus subtilis can be any method, for example, the following can be cited: using competent Bacillus subtilis, electroporation method, method using a gene gun, calcium phosphate method, etc. "Competent" refers to a state in which cells become more permeable to foreign substances (such as nucleic acids). In order to make Bacillus subtilis competent, any well-known method can be used. For example, the method described in Anagnostopoulou, C. and Spizizen, J. J. Bacteriol., 81, 741-746 (1961) can be used. In one embodiment, the NRPS-derived construction plasmid of the present invention can also be prepared in Bacillus subtilis and directly amplified in the same Bacillus subtilis.

[0082] (OGAB method) In one embodiment, the NRPS-encoding plasmid or plasmid library of the present invention can be prepared by the OGAB method. The OGAB (Ordered Gene Assembly in Bacillus subtilis) method is a method of generating a circular plasmid in a transformed organism by introducing an assembled nucleic acid into the transformed organism. By the OGAB method, a plasmid containing a large-sized nucleic acid can be easily prepared. In this specification, the nucleic acid to be introduced into the transformed organism is referred to as "assembled nucleic acid". Typically, the assembled nucleic acid has a tandem repeat structure in which a unit plasmid and a set of unit nucleic acids repeat in the same direction. By using the assembled nucleic acid having such a structure, the generation of plasmids in the transformed organism can be promoted. In this specification, "unit nucleic acid" refers to a nucleic acid molecule or a part of a nucleic acid molecule having a part of the sequence constituting the assembled nucleic acid. As will be described in detail below, a variety of unit nucleic acids are prepared as unit vectors, and then an assembled nucleic acid is constructed.

[0083] The following describes the steps for preparing the assembled nucleic acid for introduction into the transformed organism.

[0084] · Preparation of unit nucleic acid Prepare the unit nucleic acid to be introduced into the assembled nucleic acid. The unit nucleic acid can be prepared by any known method. For example, it can be prepared by polymerase chain reaction (PCR) or chemical synthesis. The unit nucleic acid can have any desired sequence, such as a sequence encoding a desired protein (such as NRPS) or a part thereof, a sequence of a control gene (promoter, enhancer, etc.), a sequence for operating the nucleic acid (restriction enzyme recognition sequence, etc.). When introduced into the assembled nucleic acid, the ends of each unit nucleic acid can be constructed to produce a specific overhanging sequence so that a variety of unit nucleic acids are arranged in a specific order and / or characteristic direction.

[0085] Since many unit nucleic acids can ultimately aggregate on a plasmid, one or more unit nucleic acids can be designed to encode one or more genes with a longer coding base length. As genes with a longer coding base length, for example, a gene cluster constituting an NRPS complex can be cited.

[0086] · Preparation of unit vector A unit vector can be prepared by linking a unit nucleic acid and an additional nucleic acid different therefrom. By using a unit vector, it becomes possible to more easily handle the unit nucleic acid.

[0087] The additional nucleic acid can be a linear nucleic acid or a circular plasmid. When a circular plasmid is used as the additional nucleic acid, the unit vector will also have a circular structure, and thus can be used for transformation of, for example, Escherichia coli. In one embodiment, the additional nucleic acid can contain an origin of replication such that the unit vector is replicated in the introduced host. In one embodiment, all the unit nucleic acids for constructing a certain aggregated nucleic acid can be linked to the same type of additional nucleic acid. In this way, the size difference between unit vectors can be reduced, and the operation of multiple unit vectors becomes easy. In one embodiment, the unit nucleic acids for constructing a certain aggregated nucleic acid can be linked to different types of additional nucleic acids. In one embodiment, for one or more unit nucleic acids for constructing a certain aggregated nucleic acid, the ratio of (the base length of the unit nucleic acid) / (the base length of the unit vector) or its average can be 50% or less, 40% or less, 30% or less, 20% or less, 15% or less, 10% or less, 7% or less, 5% or less, 2% or less, 1.5% or less, 1% or less, or 0.5% or less. The larger the size of the additional nucleic acid compared to the unit nucleic acid, the more uniformly different types of unit vectors can be operated. The linking of the unit nucleic acid and the additional nucleic acid can be carried out by any method such as ligation using DNA ligase, TA cloning, etc. In one embodiment, for one or more unit nucleic acids for constructing a certain aggregated nucleic acid, the ratio of (the base length of the unit nucleic acid) / (the base length of the unit vector) or its average can be 1% or more, 0.3% or more, 0.1% or more, 0.03% or more, 0.01% or more, 0.003% or more, or 0.001% or more, and the operation of the unit vector will become easy.

[0088] In one embodiment, the length of the unit nucleic acid can be 10 bp or more, 20 bp or more, 50 bp or more, 70 bp or more, 100 bp or more, 200 bp or more, 500 bp or more, 700 bp or more, 1000 bp or more, or 1500 bp or more, and 5000 bp or less, 5000 bp or less, 2000 bp or less, 1500 bp or less, 1200 bp or less, 1000 bp or less, 700 bp or less, or 500 bp or less.

[0089] In one embodiment, the aggregated nucleic acid can be constructed from more than 2, more than 4, more than 6, more than 8, more than 10, more than 15, more than 20, more than 30, more than 40, more than 50, more than 60, more than 70, more than 80, more than 90, or more than 100 unit nucleic acids, and is 1000 or less, 700 or less, 500 or less, 200 or less, 120 or less, 100 or less, 80 or less, 70 or less, 60 or less, or 50 or less unit nucleic acids. By adjusting the molar amounts of the respective unit nucleic acids (or unit vectors) to be substantially the same, the desired aggregated nucleic acid having a tandem repeat-like structure can be efficiently produced.

[0090] In one embodiment, the unit nucleic acid can have a base length that roughly divides one set of repeat sequences in the aggregated nucleic acid by the number of unit nucleic acids. In this way, the operation of making the molar amounts of the respective unit nucleic acids (or unit vectors) consistent becomes easy. In one embodiment, the unit nucleic acid can have a base length that is increased or decreased by 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, 7% or less, or 5% or less from the base length that roughly divides one set of repeat sequences in the aggregated nucleic acid by the number of unit nucleic acids.

[0091] In one embodiment, the unit nucleic acid can be designed to have a non-palindromic sequence (a sequence that is not a palindrome) at the end of the unit nucleic acid. For a unit nucleic acid designed in this way, when the non-palindromic sequence is a protruding sequence, a structure in which the unit nucleic acids in the aggregated nucleic acid are linked in a state of maintaining their order with each other can be easily provided.

[0092] · Production of Aggregated Nucleic Acid An aggregated nucleic acid can be constructed by linking unit nucleic acids to each other. In one embodiment, the unit nucleic acids can be prepared by cutting them out from unit vectors using restriction enzymes or the like. The aggregated nucleic acid can contain 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more repeat sequence groups. The repeat sequences in the aggregated nucleic acid can contain the sequences of the unit nucleic acids and, if necessary, the sequences of the aggregation vector nucleic acids. The aggregated nucleic acid can have a sequence that enables the replication of the nucleic acid in the transformed organism. In one embodiment, the sequence that enables the replication of the nucleic acid in the transformed organism can contain an origin of replication effective in the transformed organism (e.g., a bacterium belonging to the genus Bacillus (Bacillus subtilis)). The sequence of the origin of replication effective in Bacillus subtilis is not particularly limited, and as a sequence having a theta-type replication mechanism, for example, the origin of replication contained in plasmids such as pTB19 (Imanaka, T. et al. J. Gen. Microbiol. 130, 1399-1408. (1984)), pLS32 (Tanaka, T and Ogra, M. FEBS Lett. 422, 243-246. (1998)), pAMβ1 (Swinfield, T. J. et al. Gene 87, 79-90. (1990)) can be mentioned.

[0093] The aggregated nucleic acid can contain, if necessary, an additional base sequence in addition to the unit nucleic acids. In one embodiment, the aggregated nucleic acid can contain base sequences that control transcription and translation, such as promoters, operators, activators, and terminators. As a promoter in the case where Bacillus subtilis is used as a host, specifically, Pspac (Yansura, D. and Henner, D. J. Proc. Natl. Acad. Sci, USA 81, 439-443. (1984)), which can control expression using IPTG (isopropyl s-D-thiogalactopyranoside), or the Pr promoter (Itaya, M. Biosci. Biotechnol. Biochem. 63, 602-604. (1999)) can be mentioned.

[0094] The unit nucleic acids can form a repeating structure in the aggregated nucleic acid that maintains a certain order and orientation. In one embodiment, by constructing the unit nucleic acids such that the base sequences of the overhanging ends of the unit nucleic acids cut out from the unit vector are complementary to each other, a repeating structure that maintains a certain order and orientation in the aggregated nucleic acid can be formed. In one embodiment, by making the structures of the overhanging ends of each different unit nucleic acid unique, a repeating structure that maintains a certain order and orientation can be efficiently formed. In one embodiment, the overhanging ends can have a non-palindromic sequence and can be either 5'-end overhanging or 3'-end overhanging.

[0095] In one embodiment, a unit nucleic acid with overhanging ends can be excised from a unit vector using a restriction enzyme. In this embodiment, the unit vector may have one or more restriction enzyme recognition sequences. When the unit vector has multiple restriction enzyme recognition sequences, each restriction enzyme recognition sequence can be recognized by the same restriction enzyme or by different restriction enzymes. In one embodiment, the unit vector may contain a pair of regions recognized by the same restriction enzyme, such that the complete unit nucleic acid region is contained between these regions. The restriction enzyme used is not particularly limited, and type II restriction enzymes can be used. For example, AarI, BbsI, BbvI, BcoDI, BfuAI, BsaI, BsaXI, BsmAI, BsmBI, BsmFI, BspMI, BspQI, BtgZI, FokI, SfaNI, etc. can be cited. These restriction enzymes can create overhanging ends at positions separated by a certain distance from the recognition sequence. If an IIS type restriction enzyme (e.g., alone) is used, the sequences of the overhanging ends of the excised unit nucleic acids will be different for each unit nucleic acid. Therefore, it will be beneficial to assemble multiple unit nucleic acids in a certain order and direction. In one embodiment of using an IIS type restriction enzyme, the unit vector does not contain the region recognized by the IIS type restriction enzyme within the unit nucleic acid region. In the embodiment of using a restriction enzyme that cuts the recognition region, the unit vector may contain the region recognized by the restriction enzyme at the end of the unit nucleic acid region.

[0096] In one embodiment, when the same type of restriction enzyme is used to excise unit nucleic acids from multiple unit vectors, the restriction enzyme treatment can be carried out in a solution containing the multiple unit vectors, which can improve the efficiency of the operation. The number of types of restriction endonucleases used to produce a certain assembled nucleic acid can be, for example, 5 or less, 4 or less, 3 or less, 2 or less, or 1. By using fewer types of restriction enzymes, the deviation in the molar amounts between unit nucleic acids can be reduced. In one embodiment, the unit nucleic acids excised from the unit vector can be easily purified using any known fractionation method such as agarose gel electrophoresis.

[0097] Unit nucleic acids and, if necessary, the aggregation carrier nucleic acids can be linked (connected) to each other using DNA ligase or the like. Thus, aggregated nucleic acids can be produced. For example, the linking of the unit nucleic acids and, if necessary, the aggregation carrier nucleic acids can be carried out in the presence of components such as polyethylene glycol (e.g., PEG2000, PEG4000, PEG6000, PEG8000, etc.) and salts (e.g., monovalent alkali metals, sodium chloride, etc.). The concentration of each unit nucleic acid in the ligation reaction solution is not particularly limited and can be 1 fmol / μl or more, etc. The reaction temperature and time for ligation are not particularly limited and can be, for example, at 37°C for 30 minutes or more. In one embodiment, before linking the unit nucleic acids and, if necessary, the aggregation carrier nucleic acids, a composition containing the unit nucleic acids and, if necessary, the aggregation carrier nucleic acids can be subjected to any conditions for inactivating restriction enzymes (e.g., phenol-chloroform treatment).

[0098] The unit nucleic acids can be adjusted to approximately the same molar amount using the methods described in WO2015 / 111248 or the like. By adjusting the unit nucleic acids to approximately the same molar amount, the desired aggregated nucleic acids having a tandem repeat-like structure can be efficiently produced. By measuring the concentration of the unit vector or unit nucleic acids, the molar amount of the unit nucleic acids can be adjusted.

[0099] · Production of plasmids from aggregated nucleic acids By contacting the aggregated nucleic acids with a transformed organism, plasmids can be formed in the transformed organism. In one embodiment, examples of the transformed organism include bacteria of the genus Bacillus, bacteria of the genus Streptococcus, bacteria of the genus Haemophilus, Neisseria, etc. Examples of bacteria of the genus Bacillus include Bacillus subtilis, Bacillus megaterium, Bacillus stearothermophilus, etc. In a preferred embodiment, the transformed organism is Bacillus subtilis. In one embodiment, the transformed organism that has taken up the aggregated nucleic acids is in a competent state and can actively take up nucleic acids. For example, competent Bacillus subtilis cleaves the double-stranded nucleic acid that is the substrate on the cell surface, degrades any one of the two strands from the cleavage point, and takes up the other strand into the bacterial cell. The single strand taken up will be repaired into a circular double-stranded nucleic acid in the bacterial cell. To make the transformed organism in a competent state, any known method can be used. For example, Bacillus subtilis can be made competent using the method described in Anagnostopoulou, C. and Spizizen, J. J. Bacteriol., 81, 741-746 (1961). The transformation method can use known methods suitable for various transformed organisms.

[0100] Plasmids produced from the transformed cells can be purified using any well-known method, and the present invention also provides plasmids purified in this way. In one embodiment, the case where the purified plasmid has the desired nucleic acid sequence can be confirmed by investigating the size pattern of the fragments generated by restriction enzyme digestion, the PCR method, the base sequence determination method, etc. In one embodiment, a Bacillus subtilis containing the plasmid of the present invention is provided.

[0101] In one embodiment, the present invention provides a plasmid library encoding various NRPSs produced by the method of the present invention. In one embodiment, the present invention provides a library of Bacillus subtilis containing plasmids encoding various NRPSs produced by the method of the present invention.

[0102] In one embodiment, the present invention provides an NRPS produced from a plasmid encoding an NRPS produced by the method of the present invention. In one embodiment, the present invention provides a library of NRPSs produced from plasmids encoding various NRPSs produced by the method of the present invention.

[0103] (Composition) The peptide (NRP) produced by the NRPS described in this specification or a composition containing the same can be used in various applications such as treatment, public health, and bioengineering depending on the peptide. In one embodiment, the present invention provides a composition containing the peptide (NRP) produced by the NRPS described in this specification. The composition can be provided in various forms. As the form of the composition, for example, it can be an injection, a capsule, a tablet, a granule, an inhalant, etc. The aqueous solution for injection can be stored in a vial or a stainless steel container. In addition, the aqueous solution for injection can also be mixed with, for example, physiological saline, sugar (such as trehalose), NaCl, or NaOH.

[0104] In one embodiment, the composition of the present invention comprises a pharmaceutically acceptable carrier or excipient. Such a carrier may also be a sterile liquid, such as water and oils, including substances of petroleum, animal, vegetable or synthetic origin, although not limited thereto, including peanut oil, soybean oil, mineral oil, sesame oil and the like. For excipients, it includes light anhydrous silicic acid, crystalline cellulose, mannitol, starch, glucose, lactose, sucrose, gelatin, malt, rice, wheat flour, chalk, silica gel, sodium stearate, glycerol monostearate, talc, sodium chloride, skim milk powder, glycerol, propylene, ethylene glycol, water, ethanol, calcium carboxymethyl cellulose, sodium carboxymethyl cellulose, hydroxypropyl cellulose, hydroxypropyl methyl cellulose, polyvinyl acetal diethylaminoacetate, polyvinylpyrrolidone, gelatin, medium-chain triglycerides, polyoxyethylene hydrogenated castor oil 60, poloxamer, granulated sugar, carboxymethyl cellulose, corn starch, inorganic salts and the like. Ideally, the composition may also contain a small amount of wetting agent or emulsifier, or also contain a pH buffer. These compositions may also take the form of solutions, suspensions, emulsions, tablets, pills, capsules, powders, sustained-release complexes and the like. Conventional binders and carriers, such as triglycerides, may also be used to formulate the composition into suppositories. Oral formulations may also contain standard carriers such as pharmaceutical-grade mannitol, lactose, starch, magnesium stearate, sodium saccharin, cellulose, magnesium carbonate and the like. Examples of suitable carriers are described in E.W. Martin, Remington’s Pharmaceutical Sciences (Mark Publishing Company, Easton, U.S.A.). In addition, it may also contain, for example, surfactants, excipients, colorants, flavors, preservatives, stabilizers, buffers, suspending agents, isotonic agents, binders, disintegrants, lubricants, flow promoters, flavoring agents and the like. In one embodiment, any component of the composition of the present invention may be provided as a pharmaceutically acceptable salt.

[0105] In this specification, "or" is used when "one or more" of the items listed in the text are adopted. The same applies to "or". When the "range" of "two values" is clearly stated in this specification, the two values themselves are also included within the range.

[0106] References such as scientific literatures, patents, patent applications and the like cited in this specification are hereby incorporated by reference in this specification to the same extent as each specific description.

[0107] The preferred embodiments are shown above for ease of understanding to illustrate the present invention. The present invention will be described below based on examples, but the above description and the following examples are provided only for illustrative purposes and not for the purpose of limiting the present invention. Therefore, the scope of the present invention is not limited to the embodiments and examples specifically described in this specification, but is only limited by the claims. Example

[0108] The products described in the specific use examples of reagents were used, but equivalent products from other manufacturers (such as Sigma-Aldrich, Wako Pure Chemical Industries, Nacalai, R&D Systems, USCN Life Science INC, etc.) can also be used as substitutes.

[0109] Strain For Escherichia coli used for gene cloning, DH5α, TOP10, and JM109 were used. For the host of Bacillus subtilis gene recombination, a derivative strain of strain 168 was used. Specifically, when aggregating by the OGAB method of pristinamycin NRPS, strain BEST9731 lacking ppsABCDE and srfAA, AB, and AC was used. In the recombinant NRPS expression host, strain BUSY9621 lacking ppsABCDE and complementing the sfp gene deletion of strain 168 was used. In the cloning of the wild-type ppsABCDE gene by the BReT method, strain BEST8628 producing pristinamycin was used respectively. Culture medium For LB medium, it is prepared by dissolving 10 g of bacteria tryptone (Bactotryptone), 5 g of yeast extract, and 5 g of sodium chloride in 1 L of water. When making an agar plate, 15 g of Bacto Agar is further added and prepared by autoclaving (121 °C, 20 minutes). As needed, carbenicillin (final concentration 100 μg / mL) or tetracycline (final concentration 10 μg / mL) is added for use. The TF-I medium and TF-II medium for Bacillus subtilis transformation are prepared as follows. First, 10×Spizizen (containing 140 g of K2HPO4 (anhydrous), 60 g of KH2PO4 (anhydrous), 20 g of (NH4)2SO4, 10 g of Na3-citrate·2H2O per 1 L), 50% glucose, 2% MgSO4·7H2O, 2% casein amino acids, and water are individually prepared by autoclaving. In addition, aqueous solutions (5 mg / mL) of each amino acid of tryptophan, arginine, leucine, and threonine are prepared by membrane filtration sterilization. 50 mL of 10×Spizizen, 50% glucose, 2% MgSO4·7H2O, 2% casein amino acids, 5 mL of each of the amino acids of 5 mg / mL tryptophan, arginine, leucine, and threonine are mixed, and finally 415 mL of sterilized water is mixed and TF-I medium (500 mL) is prepared by membrane filtration sterilization and stored at 4 °C until use. Except that 2% casein amino acids are set to 2.5 mL, each amino acid solution (5 mg / mL) is set to 0.5 mL, and sterilized water is set to 435.5 mL, the same formula as TF-I medium is used to prepare TF-II medium (500 mL) by membrane filtration sterilization and stored at 4 °C until use. The ACS medium for Bacillus subtilis to produce peptides is prepared as follows. 100 g of sucrose, 11.7 g of citric acid, 4 g of sodium sulfate, 5 g of yeast extract, 4.2 g of diammonium hydrogen phosphate, 0.76 g of potassium chloride, 0.42 g of magnesium chloride hexahydrate, 0.0104 g of zinc chloride, 0.0245 g of ferric chloride hexahydrate, and 0.0181 g of manganese chloride tetrahydrate are dissolved in each 1 L of medium, adjusted to pH 6.9 with ammonia water, and then autoclaved.

[0111] In-tube gene manipulation For other general DNA manipulations, they are carried out according to the standard protocol (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1989)).

[0112] Bacillus subtilis transformation method Add 2 mL of LB medium to a 14 mL test tube (Falcon 2051). Inoculate it with a Bacillus subtilis strain from glycerol stocks stored at -70 °C, and culture it at 37 °C for 17 hours while rotating using a rotary culture device. Inject 900 μL of TF-I medium into separate new 14 mL test tubes (Falcon 2051), add 25 μL of 2% casein amino acids, add 50 μL of the pre-culture medium, and culture it at 37 °C for 4 hours while rotating using a rotary culture device. Then, inject 900 μL of TF-II medium into separate new 14 mL test tubes, add 100 μL of the TF-I culture medium, and culture it at 37 °C for 1.5 hours while rotating using a rotary culture device. Add 100 μL of the TF-II culture medium to a 1.5 mL centrifuge tube, add 8 μL of DNA. Culture it at 37 °C for 30 minutes while rotating using a rotary culture device, then add 300 μL of LB medium, and further culture it at 37 °C for 1 hour while rotating using a rotary culture device. Then, spread it on an LB medium agar plate containing 10 μg / mL tetracycline and incubate it at 37 °C overnight to obtain transformants.

[0113] Amplification of a single gene by PCR method and cloning into a plasmid Add 2.5 μL of TAKARA Ex-Taq with 10× enzyme, 2 μL of 2.5 mM dNTP solution, 0.25 μL each of 10 pmol / μL primer sets, 1 μL of genomic DNA of Bacillus subtilis as a template, 16 μL of sterilized water, and 0.5 μL of Ex-Taq HS. After incubating at 94 °C for 5 min, perform 30 cycles of 20 s at 98 °C, 30 s at 58 °C, and 1 min / kb at 72 °C to amplify the DNA. For the obtained DNA fragment, use the pCR-XL-TOPO Kit (Invitrogen) according to the manual. Or a TA cloning vector prepared by cutting the plasmid pBR-delTypeIIS-XcmI with XcmI. When cloning according to the TA cloning method, after purifying the PCR product using the MiniElute PCR purification kit, mix 1 μL of it with 1 μL of the 10 ng / μL above-mentioned vector fragment, add 2 μL of TAKATA Ligation (Mighty) Mix, and incubate overnight at 16 °C. Add 4 μL of this ligation solution to 40 μL of chemically competent cells of Escherichia coli DH5α or JM109. After incubating at a constant temperature on ice for 15 min, give a heat shock at 42 °C for 30 s. After placing on ice for 2 min, add 200 μL of LB medium. After incubating at a constant temperature of 37 °C for 1 h, spread it on an LB plate containing 1.5% agar with a concentration of 100 μg / mL carbenicillin and culture overnight at 37 °C to obtain a plasmid transformant.

[0114] Purification of Escherichia coli plasmid and restriction enzyme digestion Respectively culture the Escherichia coli transformants with plasmids cloning DNA fragments containing the desired sequences overnight at 37 °C and 120 spm in 2 mL of LB medium containing 100 μg / mL carbenicillin. For the obtained bacterial cells, use the QIA spin miniprep kit (QLAGEN) to purify the plasmid according to the manual. Add 50 μL of 10× Cut smart buffer and 5 μL of SfiI restriction enzyme (NEB) to 40 μL of DNA solution containing approximately 20 μg of plasmid, and cut out the unit DNA fragment from the plasmid vector by reacting at 50 °C for 2 h. Size fractionation of DNA using low melting point agarose gel Using a 1% low melting point agarose gel (2-Hydroxyethyl Agarose Type VII, Sigma), in the presence of 1×TAE buffer (Tris-Acetate-EDTA Buffer), a voltage of 50 V (about 4 V / cm) was applied in a general agarose gel electrophoresis apparatus (i-MyRun.N nucleic acid electrophoresis system, COSMO BIO) for 1 h of electrophoresis to separate the plasmid vector and unit DNA from the restriction enzyme-treated DNA solution. The electrophoresis gel was stained for 30 min with 100 mL of 1×TAE buffer containing 1 μg / mL of ethidium bromide (Sigma), visualized by irradiation with long-wavelength ultraviolet light (366 mn), and a DNA fragment of a specified size was cut out using a spatula and recovered into a 1.5 mL tube. To the recovered low melting point agarose gel (about 300 mg) was added 1×TAE buffer to make the total volume about 700 μL, and the gel was melted by incubating it at 65 °C for 10 min. Thereafter, an equal volume of TE-saturated phenol (NACALAI TESQUE) was added, and the restriction enzyme was inactivated by good mixing. It was separated into a phenol phase and an aqueous phase by centrifugation (20,000×g, 10 min), and the aqueous phase (about 900 μL) was recovered into a new 1.5 mL tube. 500 μL of 1-butanol (Wako Pure Chemical Industries, Ltd.) was added thereto, and after thorough mixing, it was separated by centrifugation (20,000×g, 1 min), and the operation of removing water-saturated 1-butanol was repeated until the volume of the aqueous phase became 450 μL or less to reduce the volume of the aqueous phase. 50 μL of 3M potassium acetate-acetate buffer (pH 5.2) and 900 μL of ethanol were added thereto, and DNA was precipitated by centrifugation (20,000×g, 10 min), rinsed with 70% ethanol, and dissolved in 20 μL of TE.

[0116] DNA Concentration Adjustment To 25 μL of the DNA to be measured adjusted to approximately several ng / μL concentration, add 125 μL of the working solution of SYBR green I (Molecular Probe) diluted 1 / 12,500 with TE, and mix well to obtain the measurement sample. By diluting DNA with known concentration (TOYOBO λ / HindIII DNA (250 μg / mL)) with TE, a 12-step two-fold dilution series starting from 31.25 ng / μL was prepared. Similarly, add 125 μL of the working solution to 25 μL of these, and mix well to obtain the standard dilution series. Transfer the measurement sample and the dilution series standards to a black 96-well plate, use a fluorescence plate reader (Thermofisher Fluoroskan FL), measure the fluorescence intensity at an excitation wavelength of 485 nm and a fluorescence wavelength of 535 nm, and determine the concentration of the DNA to be measured from the calibration curve made from the standard dilution series.

[0117] Gene clustering using the OGAB method Mix the gene clustering vector and each DNA fragment to be clustered in an equimolar amount containing 10 fmol, and adjust the total volume to 10 μL with TE. Add 11 μL of 2× ligation buffer (20% (w / v) PEG6000, 132 mM Tris-HCl (pH 7.6), 13.2 mM MgCl2, 20 mM DTT, 0.2 mM ATP, 300 mM NaCl), incubate the whole at 37 °C for 5 min, then add 1 μL of T4 DNA ligase (Takara), and incubate at 37 °C for 4 h. Take out 10 μL for electrophoresis, confirm after ligation, extract 10 μL of it into a new test tube, add 100 μL of Bacillus subtilis competent cells, and rotate and culture at 37 °C using a rotary mixer (duckrotor) for 30 min. Then, add 300 μL of LB medium, rotate and culture at 37 °C using a rotary mixer for 1 h, and then spread the culture solution on an LB plate containing 10 μg / mL of tetracycline and culture overnight at 37 °C.

[0118] Large-scale preparation of plasmid by ultracentrifugation method After the cultivation, 50 mL of each was separately injected into 4 50-mL test tubes (Falcon 2070) and centrifuged at 5,000×g for 10 minutes. The supernatant was discarded, and the bacterial mass was completely dispersed by vortexing. In the case of Escherichia coli, P1 Buffer (QIAGEN) was used alone; in the case of Bacillus subtilis, a P1 Buffer (QIAGEN) solution supplemented with 10 mg / mL of lysozyme and 10 mg / mL of ribonuclease A was used, and 5 mL of it was added to each of the 4 test tubes containing bacteria and mixed well. It was incubated at room temperature for 5 minutes. 5 mL of P2 Buffer (QIAGEN) was added to each of the 4 test tubes, mixed gently, and incubated at room temperature for 5 minutes. Further, P3 Buffer (QIAGEN) was added at 5-mL increments and mixed vigorously to a certain extent so that the turbidity could be evenly dispersed. It was centrifuged at 5,000×g for 10 minutes, and the supernatant was aspirated using a pipette and transferred to 4 new 50-mL screw-cap test tubes (Falcon 2070). 5 mL of phenol-saturated TE was added to each test tube and mixed vigorously. It was centrifuged at 5,000×g for 10 minutes, and the supernatant was aspirated using a pipette and transferred to 4 new 50-mL screw-cap test tubes (Falcon 2070). 20 mL of 100% ethanol was added to each for mixing, and it was centrifuged at 5,000×g for 10 minutes to remove the supernatant. 5.4 mL of TE was added to the precipitate and dissolved completely. Then, 6.40 g of cesium chloride was added and dissolved completely. Further, 2.6 mL of a 1.1 g / mL cesium chloride solution (a solution prepared by mixing 1.1 g of cesium chloride with 1 mL of water without adjusting the volume) was added. Finally, 600 μL of a 10 mg / mL ethidium bromide solution was added and mixed well. The contents were transferred to 1 ultracentrifuge tube (Beckman 362181). Water or a 1.1 g / mL cesium chloride solution (specific gravity about 1.5 g / mL) was added to fine-tune the weight so that the weight difference from the balance was within 20 mg. Centrifugation was performed using an ultracentrifuge (Beckman Coulter) under the following conditions: Temperature: 18°C; Speed: 50,000 rpm; Acceleration: Max; Deceleration: Max. Centrifugation was carried out for more than 15 hours.

[0119] After centrifugation, under ultraviolet (365 nm) observation, insert a 1 mL syringe equipped with a needle (21G×5 / 8”) into the ccc-type plasmid band, recover the plasmid solution, and transfer it to a 15 mL test tube. Add 500 μL of P3 to it, and then add water to make the total volume reach 3 mL. Further, add 9 mL of 100% ethanol. Centrifuge at 5,000×g for 10 minutes to remove the supernatant. Add 700 μL of TE to the resulting precipitate to dissolve the DNA. Transfer it to a 1.5 mL test tube, add 600 μL of 1-butanol and mix. Centrifuge at 20,000×g for about 10 seconds to separate into two layers, and discard the upper butanol layer. Add 600 μL of fresh 1-butanol and mix, centrifuge at 20,000×g for about 10 seconds to separate into two layers, and discard the upper butanol layer. Continue this operation until the aqueous layer becomes less than 450 μL. If the aqueous layer becomes less than 450 μL, add 50 μL of P3 Buffer (QIAGEN), and further add 900 μL of ethanol and mix. Then centrifuge at 20,000×g for 10 minutes. Remove the supernatant, wash the resulting clumped DNA with 900 μL of 70% ethanol, and dissolve it in 22 μL of TE.

[0120] Analysis of Lipopeptides by Analytical HPLC Pour the culture solution into a 50 mL centrifuge tube, add 12N hydrochloric acid to pH 2, and centrifuge the resulting precipitate at 9,160×g for 10 min to recover it as a pellet. Add 4 mL of 95% ethanol to it, mix with a magnetic stirrer for 1 hour, and then centrifuge at 9,160×g for 10 min to recover the supernatant. Pass it through a 0.45 μm PTFE filter membrane to obtain an HPLC sample. The HPLC column used is Inertsil ODS-2 [diameter 4.6 mm × length 250 mm; GLSciences]. The eluent uses two organic solvents: (A) 0.05% aqueous trifluoroacetic acid solution and (B) acetonitrile:isopropanol = 7:3 containing 0.02% trifluoroacetic acid. Usually, the total of the two eluents flows at 1 mL / min. The time variation of each eluent is as follows. At the beginning (0 minutes), A = 60% and B = 40%. Then, it linearly changes in a way that A decreases and B increases. After 30 minutes, it becomes A = 0% and B = 100%, and then maintains this state for 5 minutes. After that, it maintains the state of A = 60% and B = 40% for 10 minutes. Detection is carried out by absorption at 205 nm.

[0121] Purification of Peptides by Preparative HPLC Add 200 mL of ACS medium supplemented with tetracycline (10 μg / mL) to a 500 mL Erlenmeyer flask, inoculate the strain into it, and culture at 30 °C and 120 spm for 4 days. Adjust the culture broth to pH 2 using 12N hydrochloric acid, and recover the resulting precipitate by centrifugation at 9,160×g for 10 min. Add 500 mL of 95% ethanol thereto, dissolve by stirring, and filter through a 0.45 μm filter membrane (Corning (registered trademark) 150 mL Filter System, 0.45 μm 13.6 cm 2 Cellulose Acetate). Among the obtained filtrates, concentrate 300 mL to 100 mL using a rotary evaporator. Again, filter through a 0.45 μm filter membrane and concentrate to 10 - 15 mL using a rotary evaporator. Further, concentrate using a vacuum centrifugal concentrator (Labconco Centrivap concentrator), and filter through a 0.45 μm PTFE filter membrane (htslabs.com, Standard / Filtervial PTFE 0.45 μM W / Pre-slit septum Blue snap cap PK100) as a sample for preparative HPLC. Mount a Phenomenex, Luna (registered trademark) 5 μm C18(2) LC Column 250x10mm, Ea chromatographic column in an HPLC apparatus manufactured by Shimadzu Corporation equipped with a fraction collector (FRC-10A), and use conditions of a column temperature of 40 °C and a flow rate of 3 mL / min. The eluent used was the same A and B as in analytical HPLC, and the time change in the ratio of A to B also used the same conditions as in analytical HPLC. Detection was performed at 205 nm.

[0122] Example 1: Introduction of SfiI site into Prilipastatin NRPS Prilipastatin is a cyclic lipopeptide composed of 10 amino acids produced by Bacillus subtilis and biosynthesized by NRPS. Prilipastatin NRPS is as Figure 1It is composed of five enzymes, PpsA, PpsB, PpsC, PpsD, and PpsE, each related to the extension of two, two, two, three, and one specified amino acids, respectively. There are a total of ten modules between PpsA and PpsE, and the NRPS genes are also arranged in the same order on the Bacillus subtilis genome. In this example, a restriction enzyme SfiI site was introduced into the DNA sequence of the region between the C domain at the module boundary and the A domain of the adjacent module (hereinafter referred to as the C-A linker). First, among the ten C-A linker regions between PpsA and PpsE, the amino acid sequences between the C5 consensus sequence ((IV)GxFVNT(QL)(CA)xR) in the C domain and the A1 consensus sequence (L(TS)YxEL) in the A domain were collected and subjected to multiple alignment analysis by ClustalW( Figure 2 ). As a result, 25 amino acids away from the A1 consensus sequence in the direction of the C5 consensus sequence, there is a part where diversity can be observed in the amino acid sequence, so this was noted. The SfiI site, 5'-GGCCNNNN / NGGCC-3' ( / is the cleavage site, and N can be any one of A, C, T, G), affects 13 bases and the maximum continuous five amino acids in the NRPS coding region. In order to meet the requirement that the overhang sequences generated by each SfiI cleavage are unique after considering the complementary strand in the OGAB method, and further meet the requirement that at least one base of G or C is contained in the overhang sequence of three bases, while minimizing the change from the wild-type sequence of the coding region amino acid sequence caused by the introduction of the SfiI site, the modified DNA sequence was determined according to the following rules( Figure 3 ). First, among the 13 bases of the SfiI site, by placing the G base closest to the C domain at the third base of the codon encoding the most upstream amino acid among the five amino acids that can be affected, the possibility of conversion to a synonymous codon was increased. Through GCC at the second to fourth bases, alanine can be encoded regardless of the wild-type amino acid sequence. The fifth to seventh bases ensure the specificity of the overhang sequence to be formed and make it as similar as possible to the wild-type sequence. The eighth to tenth bases ensure the specificity of the overhang sequence to be formed and attempt to encode the same amino acid as the wild-type. Even if it is difficult to do so, the second base of the codon is made the same as the wild-type, so that the hydrophobicity / hydrophilicity of the encoded amino acid is not changed significantly. The eleventh to thirteenth bases can encode alanine regardless of the wild-type amino acid sequence. Based on such a design principle, as Figure 3As shown in SEQ ID NOS: 1 to 22, primers designed such that the SfiI site designed at the 5'-end of the primer was positioned were used, and wild-type Bacillus subtilis genomic DNA was used as a template, amplified by PCR, and cloned into Escherichia coli DH5α using the pCR-XL-TOPO kit (Invitrogen) according to the manual to obtain DNA fragments of pps01 to pps09. For pps10, which is the last module, primers were designed that could amplify in such a way as to include the consecutive part (upstream side) of pps09 in the C-A linker to the putative ρ-factor-independent terminator sequence (downstream side) present downstream of the ppsE gene. In addition, for the pps00 region upstream of pps01, primers were designed that could amplify starting from the consecutive part of pps00 and pps01 to include the promoter region of the pps operon. After the pps00 fragment and the pps10 fragment were also cloned in pCR-XL-TOPO, and after confirming the cloning direction of the vector, at the SpeI site of the plasmid cloning pps00, a fragment of pps10 obtained by cutting the plasmid cloning pps10 with SpeI and XbaI was ligated, and a clone ligated with pps00 and pps10 maintaining the positional relationship in the genome was selected to construct pCR-pps00-pps10. This plasmid was partially digested with the restriction enzyme SilI, and the 4.3-kb DNA fragment ligated with pps00 and pps10 was size-fractionated by low-melting-point agarose gel electrophoresis. After purification, it was ligated to the DraIII cloning site of pGETS109DraIII, a shuttle plasmid vector for Escherichia coli - Bacillus subtilis. This plasmid was cut with SfiI, and the long DNA fragment was excised and purified by using low-melting-point agarose gel electrophoresis to obtain a vector DNA fragment that could be used in the OGAB method. The vector DNA fragment and the fragments of pps01 to pps09 obtained by cutting the plasmid DNA of the DNA fragments cloned from pps01 to pps09 with SfiI were mixed at equimolar concentrations and assembled by the OGAB method ( Figure 4 ). After transformation, Bacillus subtilis was spread on an LB plate containing tetracycline, and 57 transformants were obtained by overnight culture at 37°C. After culturing randomly selected 12 colonies in an LB medium containing tetracycline, the plasmids were purified in small amounts and cut with SfiI and other restriction enzymes, and the cleavage patterns were analyzed by agarose gel electrophoresis ( Figure 5 ). As a result, 5 out of 12 contained the ligated DNA of pps00-pps01-pps02-pps03-pps04-pps05-pps06-pps07-pps08-pps09-pps10. One of them was named ppsW03 (SEQ ID NO: 23).

[0123] Confirmation of the production of plipastatin using an NRPS into which an SfiI site has been introduced BUSY9261, the NRPS production host, was transformed by the Bacillus subtilis competent cell method using ppsW03, and the resulting strain was named ppsW03 / BUSY9261. After pre-culturing this strain in LB medium containing tetracycline, it was cultured in ACS medium at 30 °C for 5 days. The acidic precipitate of the culture broth of this strain was extracted with ethanol and analyzed by analytical HPLC. As a result, a peak was observed at the same retention time as that of plipastatin of the wild-type BEST8628 strain ( Figure 6 ). When this peak substance was recovered by preparative HPLC and subjected to mass spectrometry analysis, the production of plipastatin was confirmed ( Figure 7 ). By transforming the BEST8628 strain with pGETS109DraIII-pps00-10 digested with SfiI, a plasmid ppsW10 having the same sequence as ppsW03 but having a wild-type NRPS without an SfiI site was constructed by the BReT method (see Tomita, S., Tsuge, K., Kikuchi, Y., Itaya, M., 2004. Targeted isolation of a designated region of the Bacillus subtilis genome by recombinational transfer. Appl. Environ. Microbiol. 70, 2508 - 2513.) ( Figure 5 ). The ppsW10 / BUSY9261 strain into which this was introduced into the BUSY9261 strain was cultured as a comparison object. As a result, the ppsW03 / BUSY9261 strain showed the same amount of plipastatin production as the ppsW10 / BUSY9261 strain ( Figure 8 ). It was confirmed that the introduction of the SfiI site did not affect the amount of plipastatin production.

[0124] Example 2: Construction of a chimeric NRPS A novel recombinant NRPS gene was constructed by replacing the pps07 fragment encoding the 7th module and the pps08 fragment encoding the 8th module of the plipastatin NRPS prepared in Example 1 with DNA fragments of modules derived from other NRPSs ( Figure 9) Replace the modules of surfactin NRPS of Bacillus subtilis. To encode the srf07 fragment of the fourth module of surfactin NRPS, an SfiI site was designed in the same manner as in Example 1 at a position approximately 25 amino acids upstream of the A1 consensus sequence to form an overhang sequence identical to pps07. In addition, for the srf08 fragment encoding the fifth module of surfactin NRPS, an SfiI site was also designed in the same manner at a position approximately 25 amino acids upstream of the A1 consensus sequence to form an overhang sequence identical to pps08. Using the primers of SEQ ID NOs: 24 to 27 in the same manner as in Example 1, the DNA fragments of srf07 and srf08 were amplified with the genomic DNA of Bacillus subtilis 168 strain as a template and cloned into pCR-XL-TOPO. The srf07 fragment and srf08 fragment were obtained by cleaving the resulting plasmid with SfiI.

[0125] For the chimeric NRPS gene, the following three patterns of chimeras 1 to 3 were prepared ( Figure 9 ). Chimera 1: pps00 - pps01 - pps02 - pps03 - pps04 - pps05 - pps06 - srf07 - pps08 - pps09 - pps10 (SEQ ID NO: 28) Chimera 2: pps00 - pps01 - pps02 - pps03 - pps04 - pps05 - pps06 - pps07 - srf08 - pps09 - pps10 (SEQ ID NO: 29) Chimera 3: pps00 - pps01 - pps02 - pps03 - pps04 - pps05 - pps06 - srf07 - srf08 - pps09 - pps10 (SEQ ID NO: 30)

[0126] These plasmids were assembled using the OGAB method with the Bacillus subtilis host strain BEST8666. Each accurately assembled plasmid was introduced into the Bacillus subtilis host strain BUSY9261 lacking ppsABCDE using the competent cell method, and the resulting strains were named BUSY9261 / chimeral, BUSY9261 / chimera2, and BUSY9261 / chimera3, respectively. After pre-culturing each strain using LB medium, they were cultured using ACS medium for 5 days and analyzed by analytical HPLC to confirm that these strains produced compounds with different mobilities from those of known plipastatin and surfactin. For more detailed structural analysis, large-scale culturing was performed, and then, the new substances were purified by preparative HPLC for large-scale preparation. As a result of determining the structure by LC-MS / MS, it was suggested that · Chimeric peptide 1: β-hydroxy fatty acid - Glu - Leu - Leu - Val -Gln-Tyr-Ile-OH (SEQ ID NO: 61), produced by BUSY9261 / chmera2 and BUSY9261 / chmera3 · Chimeric peptides 2 and 3: β-hydroxy fatty acid - Glu - Leu - Leu the peptide -Val-Asp-Tyr-Ile-OH (SEQ ID NO: 62) (Figure 10).

[0127] The peptide products predicted from the NRPS genes encoded in chimeras 1 to 3 were: Derived from chimera 1: β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val - Val -Gln-Tyr-Ile-OH (SEQ ID NO: 63) (or a lactone ring formation between the hydroxyl group on the fatty acid chain at the N-terminus of this peptide or the hydroxyl group of Tyr near the N-terminus and the carboxyl group at the C-terminus); Derived from chimera 2: β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val -Pro-Asp-Tyr-Ile-OH (SEQ ID NO: 64) (or a lactone ring formation between the hydroxyl group on the fatty acid chain at the N-terminus of this peptide or the hydroxyl group of Tyr near the N-terminus and the carboxyl group at the C-terminus); Derived from chimera 3: β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val-Val-Asp-Tyr-Ile-OH (SEQ ID NO: 65) (or a lactone ring formed between the hydroxyl group on the fatty acid chain at the N-terminus of the peptide or the hydroxyl group of Tyr near the N-terminal side and the carboxyl group at the C-terminus). Contrary to expectations, with the introduced srf07 or srf08 as the boundary, a new peptide is formed where the first half is derived from surfactin NRPS and the second half is derived from plipastatin NRPS. Chimeric peptides 2 and 3 are the same peptide. The reason is speculated to be that the plasmids of chimeras 1-3 introduced into BUSY9261 were introduced as a result of the wild-type surfactin NRPS gene existing in the host genomic DNA being introduced using its homology with srf07 or srf08 for some reason. In short, new peptides were obtained.

[0128] Example 3 (Hypothetical Example): Introduction of an SfiI Recognition Sequence into Surfactin Synthase Similar to the plipastatin NRPS of Example 1, as Figure 11 shown, a recognition site for the restriction enzyme SfiI was introduced into the portion 25 amino acids upstream of the C1 sequence of the consensus sequence of the C domain. Using the primer DNAs shown in SEQ ID NOS: 31-42, each brick (unit DNA) for the OGAB method was amplified by PCR. The gene fragments were cloned into pBR-delTypeIIS-3-AarI to prepare clones with accurate base sequences. srf00, srf04-srf10, which are the total unit DNAs of srf07 and srf08 containing the unit DNAs prepared in Example 2, were prepared. Similar to Example 1, an assembly plasmid (SEQ ID NO: 43) linked with srf00-srf04-srf05-srf06-srf07-srf08-srf09-srf10 was constructed using the BEST9731 strain by the OGAB method. This plasmid produces surfactin.

[0129] Example 4: Construction of a Chimeric NRPS Library Using the plasmid ppsW03 with an SfiI site introduced into the plipastatin NRPS gene prepared in Example 1 and the plasmid srf-SfiI with an SfiI site introduced into the surfactin NRPS gene prepared in Example 3, a chimeric plasmid library was constructed by the Combi-OGAB method shown in WO2020 / 203496. The ppsW03 and srf-SfiI plasmids were purified to high purity by ultracentrifugation. 1 fmol of each was digested with the restriction enzyme SfiI, and the DNA fragments were purified by phenol treatment, butanol treatment, and ethanol precipitation. The resulting DNAs were mixed, ligated into a pseudo-tandem repeat form using 2× ligation buffer and T4 DNA ligase, and introduced into Bacillus subtilis carrying the sfp gene in which the NRPS for the contents of plipastatin and surfactin was deleted, BUSY9731, by the competent cell method, thereby constructing a chimeric NRPS library( Figure 12 ).

[0130] Example 5: Introduction of a restriction enzyme recognition site at different positions In Example 1, an SfiI site was introduced into the upstream part 25 amino acids from the C1 consensus sequence of plipastatin NRPS. However, in Example 5, differently, a restriction enzyme PflMI site was introduced at the position 25 amino acids downstream from the C3 consensus sequence. 3’-CCANNNN / NTGG-5’ recognized by PflMI ( / indicates the cleavage site, and N indicates any one of A, C, T, G) is present only once in plipastatin NRPS. Each unit DNA with an attached PflMI site was amplified by PCR in the same manner as in Example 1. After cloning into an Escherichia coli plasmid vector, it was digested with PflMI, and each unit DNA was purified by low melting point agarose gel electrophoresis. These unit DNAs and the OGAB assembly vector DNA were mixed to an equimolar concentration, and gene assembly was performed by the OGAB method using Bacillus subtilis strain BEST9731. By introducing the resulting plasmid DNA into strain BUSY9621, plipastatin was produced using the plipastatin NRPS generated from the NRPS gene into which the PflMI site was introduced. Example 6: Introduction of a restriction enzyme site into iturin NRPS Iturin A produced by Bacillus subtilis strain RB14 is an NRP in which a β-amino fatty acid and seven amino acids are linked in a ring and is synthesized by NRPS encoded by three ORFs, ituA, ituB, and ituC. These three ORFs are consecutive in this order, and there is no restriction enzyme SflI site within the gene. An SfiI site should be introduced into the upstream part 25 amino acids from the C1 consensus sequence of the NRPS of iturin A. Each unit DNA with the SfiI site added is amplified by PCR in the same manner as in Example 1. After cloning into an Escherichia coli plasmid vector, it is cut with SfiI, and each unit DNA is purified by low melting point agarose gel electrophoresis. These unit DNAs and the OGAB assembly vector DNA are mixed to an equimolar concentration, and the Bacillus subtilis strain BEST9731 is used for gene assembly by the OGAB method. By introducing the resulting plasmid DNA into the BUSY9621 strain, iturin A is produced by the iturin A NRPS generated from the NRPS gene into which the SfiI site has been introduced.

[0132] (Annotation) As described above, although the present invention has been exemplified by preferred embodiments of the present invention, the present invention should be understood to be interpreted only by the claims. It is understood that the patents, patent applications, and literature cited in this specification should be incorporated by reference in their entirety as if their contents were the same as the specific descriptions in this specification. This application claims priority to Japanese Patent Application No. 2022-121822 filed with the Japan Patent Office on July 29, 2022, and the entire contents thereof are incorporated by reference as if they were the same as the contents constituting the present application. Industrial Applicability

[0133] According to the present invention, a novel NRPS can be efficiently produced. In this way, an NRPS capable of efficiently producing NRP and an NRPS for providing a novel NRP can be provided. Sequence Listing Free Text

[0134] Sequence Number 1: Primer 00-F Sequence Number 2: Primer 00-R Sequence Number 3: Primer 01-F Sequence Number 4: Primer 01-R Sequence Number 5: Primer 02-F Sequence Number 6: Primer 02-R SEQ ID No. 7: Primer 03-F SEQ ID No. 8: Primer 03-R SEQ ID No. 9: Primer 04-F SEQ ID No. 10: Primer 04-R SEQ ID No. 11: Primer 05-F SEQ ID No. 12: Primer 05-R SEQ ID No. 13: Primer 06-F SEQ ID No. 14: Primer 06-R SEQ ID No. 15: Primer 07-F SEQ ID No. 16: Primer 07-R SEQ ID No. 17: Primer 08-F SEQ ID No. 18: Primer 08-R SEQ ID No. 19: Primer 09-F SEQ ID No. 20: Primer 09-R SEQ ID No. 21: Primer 10-F SEQ ID No. 22: Primer 10-R SEQ ID No. 23: ppsW03 As described in the Sequence Listing. SEQ ID No. 24: Primer srf07-F SEQ ID No. 25: Primer srf07-R Sequence number 26: Primer srf08-F Sequence number 27: Primer srf08-R Sequence number 28: Chimera 1 As described in the sequence listing. Sequence number 29: Chimera 2 As described in the sequence listing. Sequence number 30: Chimera 3 As described in the sequence listing. Sequence number 31: Primer srf00-F Sequence number 32: Primer srf00-R Sequence number 33: Primer srf04-F Sequence number 34: Primer srf04-R Sequence number 35: Primer srf05-F Sequence number 36: Primer srf05-R Sequence number 37: Primer srf06-F Sequence number 38: Primer srf06-R Sequence number 39: Primer srf09-F Sequence number 40: Primer srf09-R Sequence number 41: Primer srf10-F Sequence number 42: Primer srf10-R Sequence number 43: srfAA-AC As described in the sequence listing. Sequence number 44: A1 core motif sequence L(TS)YxEL Serial number 45: A2 core motif sequence LKAGxAYL(VL)P(LI)D Serial number 46: A3 core motif sequence LAYxxYTSG(ST)TGxPKG Serial number 47: A5 core motif sequence NxYGPTE Serial number 48: A6 core motif sequence GELxIxGxG(VL)ARGYL Serial number 49: A7 core motif sequence Y(RK)TGDL Serial number 50: A8 core motif sequence GRxDxQVKIRGxRIELGEIE Serial number 51: A9 core motif sequence LPxYM(IV)P Serial number 52: A10 core motif sequence NGK(VL)DR Serial number 53: T core motif sequence DxFFxxLGG(HD)S(LI) Serial number 54: C1 core motif sequence SxAQxR(LM)(WY)xL Serial number 55: C2 core motif sequence RHExLRTxF Serial number 56: C3 core motif sequence MHHxISDG(WV)S Serial number 57: C4 core motif sequence YxD(FY)AVW Serial number 58: C5 core motif sequence (IV)GxFVNT(QL)(CA)xR Serial number 59: C6 core motif sequence (HN)QD(YV)PFE Serial number 60: C7 core motif sequence RDxSRNPL Serial number 61: Chimeric peptide 1 β-hydroxy fatty acid - Glu - Leu - Leu - Val - Gln - Tyr - Ile Serial number 62: Chimeric peptides 2 and 3 β-hydroxy fatty acid - Glu - Leu - Leu - Val - Asp - Tyr - Ile Sequence number 63: Chimeric 1 putative peptide β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val - Val - Gln - Tyr - Ile Sequence number 64: Chimeric 2 putative peptide β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val - Pro - Asp - Tyr - Ile Sequence number 65: Chimeric 3 putative peptide β-hydroxy fatty acid - Glu - Orn - Tyr - Thr - Glu - Val - Val - Asp - Tyr - Ile.

Claims

1. A method, which is a method for preparing a production plasmid containing a nucleic acid sequence encoding a non-ribosomal peptide synthetase NRPS, comprising: a step of preparing a starting plasmid group, wherein each of the starting plasmid groups contains a base sequence encoding at least 1 NRPS module and at least 1 restriction enzyme recognition sequence; a step of treating the starting plasmid group with a restriction enzyme to prepare a mixture containing a plurality of obtained fragmented nucleic acid fragments; a step of ligating the plurality of fragmented nucleic acid fragments to form a ligated nucleic acid and a step of bringing the ligated nucleic acid into contact with a transformed organism to form a production plasmid.

2. The method according to claim 1, wherein the plurality of fragmented nucleic acid fragments include a group of fragmented nucleic acid fragments, the fragmented nucleic acid fragments have the same protruding sequence, and contain a base sequence encoding a plurality of NRPS modules.

3. The method according to claim 1 or 2, wherein the plurality of fragmented nucleic acid fragments include a group of fragmented nucleic acid fragments, the fragmented nucleic acid fragments have the same protruding sequence, and contain a base sequence encoding a plurality of the NRPS modules that capture amino acids.

4. The method according to any one of claims 1 to 3, wherein each of the fragmented nucleic acid fragments has a protruding sequence with a length of 3 to 5 bases.

5. The method according to any one of claims 1 to 4, wherein in the starting plasmid, the restriction enzyme recognition sequence is located between adjacent base sequences encoding NRPS modules.

6. The method according to any one of claims 1 to 5, wherein in the starting plasmid, the restriction enzyme recognition sequence is located between the C-domain coding region and the A-domain coding region of the NRPS module.

7. The method according to any one of claims 1 to 6, wherein in the starting plasmid, the restriction enzyme recognition sequence is located in a region between the C-domain coding region and the A-domain coding region of the NRPS module, excluding the C-domain coding region and the A-domain coding region.

8. The method according to any one of claims 1 to 7, wherein at least 1 of the restriction enzyme recognition sequences is present within the base sequence encoding the NRPS module.

9. The method according to any one of claims 1 to 8, wherein at least 1 of the restriction enzyme recognition sequences is present outside the base sequence encoding the NRPS module.

10. The method according to any one of claims 1 to 9, wherein the restriction enzyme recognition sequence is a non-natural sequence in the NRPS module.

11. The method according to any one of claims 1 to 10, wherein each of the fragmented nucleic acid fragments contains at most 1 NRPS module.

12. The method according to any one of claims 1 to 11, wherein the restriction enzyme recognition sequence is cleaved at a site different from the specific recognition site of the restriction enzyme.

13. The method according to any one of claims 1 to 12, wherein the restriction enzyme recognition sequence includes the recognition sequence of SfiI, BglI, AlwNI, DraIII, PflMI, BstAPI, AarI, BbsI, BsaI, BsmBI or BspQI.

14. The method according to any one of claims 1 to 13, wherein each of the starting plasmids contains a base sequence encoding 4 to 12 NRPS modules.

15. The method according to any one of claims 1 to 14, wherein in the step of preparing the mixture containing the fragmented nucleic acid fragments, each of the starting plasmids generates fragmented nucleic acid fragments having mutually different overhanging sequences.

16. The method according to any one of claims 1 to 15, wherein the mixture containing the fragmented nucleic acid fragments is a solution.

17. The method according to any one of claims 1 to 16, wherein each of the starting plasmids independently contains the same or different selection marker sequences.

18. The method according to claim 17, wherein the selection marker sequence includes a drug resistance marker sequence.

19. The method according to any one of claims 1 to 18, wherein each of the starting plasmids independently contains the same or different replication origins that can function in Bacillus subtilis.

20. The method according to any one of claims 1 to 19, wherein each of the starting plasmids independently contains a promoter region upstream of the same or different base sequences encoding the NRPS.

21. A method for generating a production plasmid or a plasmid library, which is produced by the method according to any one of claims 1 to 20.

22. A non-ribosomal peptide synthetase (NRPS) generated from the production plasmid or plasmid library according to claim 21.

23. A non-ribosomal peptide (NRP) generated from the non-ribosomal peptide synthetase according to claim 22.

24. A nucleic acid, which is a nucleic acid encoding a non-ribosomal peptide synthetase (NRPS), and includes a base sequence encoding at least 1 NRPS module and a restriction enzyme recognition sequence that is non-natural for the NRPS module.

25. The nucleic acid according to claim 24, wherein the restriction enzyme recognition sequence is located between adjacent base sequences encoding the NRPS module.

26. The nucleic acid according to claim 24 or 25, wherein the restriction enzyme recognition sequence is located between the C domain coding region and the A domain coding region of the NRPS module.

27. The nucleic acid according to any one of claims 24 to 26, wherein the restriction enzyme recognition sequence is located in a region between the C domain coding region and the A domain coding region of the NRPS module, excluding the C domain coding region and the A domain coding region.

28. The nucleic acid according to any one of claims 24 to 27, wherein the restriction enzyme recognition sequence is a recognition sequence of a restriction enzyme that recognizes a sequence containing a region where any kind of base can exist.

29. The nucleic acid according to any one of claims 24 to 28, wherein the restriction enzyme recognition sequence is a recognition sequence of SfiI.

30. The nucleic acid according to any one of claims 24 to 29, wherein the number of NRPS modules encoded by the nucleic acid is 4 to 12.

31. A method, which is a method for producing a product nucleic acid containing a nucleic acid sequence encoding a non-ribosomal peptide synthetase (NRPS), and includes: A step of preparing the nucleic acid according to any one of claims 24 to 30 as a starting nucleic acid, A step of treating the nucleic acid with a restriction enzyme that recognizes the restriction enzyme recognition sequence to prepare a fragmented nucleic acid fragment, and a step of ligating the fragmented nucleic acid fragments to form a product nucleic acid different from the starting nucleic acid.

32. A product nucleic acid obtained by the method according to claim 31.

Citation Information

Patent Citations

  • Foreign object protection device and foreign object protection method

    JP2022121822A

  • Polyether nucleic acids

    US5908845A

  • Method for preparing DNA unit composition, and method for creating concatenated DNA

    WO2015111248A1

  • Method for constructing chimeric plasmid library

    WO2020203496A1