Random mutation library

The method introduces random mutations into plasmid DNA populations for direct combinatorial library construction, addressing limitations of existing methods by enabling rapid screening and optimizing gene clusters for enhanced microbial strain production.

WO2026048921A1PCT designated stage Publication Date: 2026-03-05SYNPLOGEN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/030254
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-29
Filing Date
2025-08-28
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing methods for constructing combinatorial DNA libraries, such as the Combi-OGAB® method, struggle with handling random mutations and require time-consuming secondary library preparation via PCR, limiting their applicability to gene clusters with long chains and many genes, especially for natural product biosynthetic enzyme genes.

Method used

A method involving the introduction of random mutations into plasmid DNA populations, followed by direct combinatorial library construction using the Combi-OGAB® method, allowing for rapid screening and reuse of plasmids without additional preparation steps, and optimizing promoters and genes for improved production strains.

Benefits of technology

Enables the efficient creation of microbial strains with desired traits by rapidly constructing combinatorial libraries and selecting mutations that enhance target substance production, particularly in hosts like Bacillus subtilis, facilitating improved yield and novel substance generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JPOXMLDOC01-APPB-T000001
    Figure JPOXMLDOC01-APPB-T000001
  • Figure JPOXMLDOC01-APPB-T000002
    Figure JPOXMLDOC01-APPB-T000002
  • Figure JPOXMLDOC01-APPB-T000003
    Figure JPOXMLDOC01-APPB-T000003
Patent Text Reader

Abstract

The present disclosure provides a method for creating, by mutating a gene group, a plasmid that exhibits desired properties such as high production and by-product suppression. The present disclosure provides: a method for producing a combinatorial library; or a method for producing a combinatorial library (including Golden Gate), the method comprising 1) a step for providing source plasmid DNA (single type or multiple types), 2) a step for introducing mutations into the source plasmid DNA and generating a plasmid DNA population (seed plasmid population) (seed plasmids) including mutation-introduced plasmid DNA (seed plasmids) into which the mutations have been introduced, and 3) a step for subjecting the plasmid DNA population to a combinatorial library formation procedure.
Need to check novelty before this filing date? Find Prior Art

Description

Random mutation library

[0001] The present disclosure relates to methods for selecting random mutations to exhibit desired properties.

[0002] The Combi-OGAB® method makes it possible to construct large-scale combinatorial DNA libraries and handle a large number of library fragments without being affected by the length or complexity of the DNA to be library-generated. Furthermore, the plasmid DNA harbored by strains that exhibit a favorable phenotype can be directly used in the next library preparation, allowing for screening cycles without the need for a process of re-preparing fragments using methods such as PCR (Patent Document 1). As an example, a combinatorial library was created in which promoters with different transcription initiation times and transcription strengths were arranged monocistronically for a group of natural product biosynthetic genes. A screening cycle of promoter combinations was then performed to identify promoter combinations that would produce strains capable of producing target substances. As a result, a transcriptionally regulated strain was successfully established, resulting in a heterologous production strain (Miyamoto, N., et al., ACS Omega, 2024, 9, 6, 6873-6879, Non-Patent Document 1).

[0003] However, while the conventional Combi-OGAB® method allows for the construction of combinatorial libraries from prepared plasmids by selecting and designing target DNA sequences for library construction in advance, it is difficult to construct combinatorial libraries containing random mutations and screen for favorable combinations of mutations. Functional improvement through unpredictable mutations is an important tool for constructing strains with desired traits.

[0004] There have been reports of constructing random mutation libraries using widely used DNA fragment assembly techniques such as the Golden Gate method and the Gibson Assembly method (Non-Patent Documents 2 and 3). However, these methods require screening from a library once constructed, and then re-preparing fragments by PCR to create a secondary library, which takes more time than the Combi-OGAB (registered trademark) method. Furthermore, there are technical limitations in that the range of DNA lengths and fragment numbers that can be handled is narrow, making these methods unsuitable for gene clusters with long chains and many constituent genes, such as natural product biosynthetic enzyme genes.

[0005] Patent No. 7101431

[0006] Miyamoto, N. , et al. , ACS Omega, 2024, 9, 6, 6873-6879. Pullmann, P. , et al. , Sci. Rep. , 2019, 9, 10932. Olszakier, S. , et al. , BMC Biotechnol. , 2022, 13, 22 (1), 10.

[0007] From the above perspectives, the present disclosure has been completed, in which a variety of plasmids containing random mutations are prepared, and a library is constructed and screened using the Combi-OGAB (registered trademark) method, thereby selecting mutations that lead to improved production of a target substance in host cells such as Bacillus subtilis.

[0008] That is, the gist of the present disclosure is as follows: (Item 1) A method for producing a combinatorial library, comprising: 1) providing starting plasmid DNA(s); 2) introducing mutations into the starting plasmid DNA to generate a plasmid DNA population (also referred to as a "seed plasmid population") containing mutated plasmid DNAs (also referred to as "seed plasmids") into which the mutations have been introduced; and 3) subjecting the plasmid DNA population to a combinatorial library construction procedure. (Item 2) The method according to any one of the above items, wherein the seed plasmid population contains or does not contain the starting plasmid. (Item 3) The method according to any one of the above items, wherein the combinatorial library construction procedure includes the Combi-OGAB (registered trademark) method, the Golden Gate method, the Gibson Assembly method, or overlap extension PCR. (Item 4) The method according to any one of the above items, wherein the library can be recombinatorialized (i.e., the resulting plasmids can be directly used next time). (Item 5) The method of any one of the above items, wherein the combinatorial library construction procedure includes Combi-OGAB (registered trademark) or the like. (Item 6) The method of any one of the above items, comprising a step of confirming, prior to step 3), that the plasmid DNA population generated in step 2) has a structure appropriate for the combinatorial library construction procedure. (Item 7) The method of any one of the above items, comprising a step of repeating step 3), or steps 2) and 3), at least once. (Item 8) The method of any one of the above items, comprising a step of selecting plasmid DNA having desired properties. (Item 9) The method of any one of the above items, wherein the selection of the plasmid DNA is performed after step 2), after step 3), or both steps 2) and 3). (Item 10) The method of any one of the above items, wherein the method is a method for producing plasmid DNA that confers desired properties. (Item 11) The method of any one of the above items, further comprising a step of optimizing a promoter. (Item 12) The method according to any one of the above items, further comprising optimizing the homologous gene.(Item 13) The method according to any one of the preceding items, further comprising a step of optimizing the base sequence in the structural gene. (Item 14) The method according to any one of the preceding items, wherein the optimization is performed to improve the function or stability of the gene or gene product. (Item 15) The method according to any one of the preceding items, wherein the optimization step is performed before step 1) or after step 2). (Item 16) The method according to any one of the preceding items, further comprising a step of preparing the starting plasmid DNA in a structure suitable for the combinatorial library construction procedure. (Item 17) The method according to any one of the preceding items, further comprising a step of preparing the starting plasmid DNA in a structure suitable for the combinatorial library construction procedure. (Item 18) The method according to any one of the preceding items, wherein the step of generating a seed plasmid population further comprises a step of preparing the starting plasmid DNA in a structure suitable for the combinatorial library construction procedure. (Item 19) A method for producing a desired host organism strain, comprising a step of transforming a host organism with the combinatorial library produced by the method according to any one of the preceding items. (Item 20) A method for producing a desired host organism strain, comprising transforming a host organism with a plasmid DNA that confers a desired trait to produce the desired strain. (Item 21) A method for producing a desired host organism strain described in any one of the above items, wherein the transformation comprises integrating a gene that confers the desired trait into the genome of the host cell. (Item 22) A method for improving a desired trait by random mutation, comprising: 1) inducing mutations in plasmid DNA carrying a gene cluster containing at least one gene, 2) constructing a combinatorial library based on diverse plasmid DNAs containing random mutations, and 3) selecting plasmid DNA that expresses the desired trait. (Item 22A) A method for improving a desired trait by mutation, comprising: 1) inducing mutations in plasmid DNA carrying a gene cluster containing at least one gene, 2) constructing a combinatorial library based on the plasmid DNAs containing the mutations, and 3) selecting plasmid DNA that expresses the desired trait.(Item 23) The method of any one of the above items, wherein the desired property includes the ability to be recombinatorial. (Item 24) The method of any one of the above items, wherein the plasmid after recombinatorialization is used for recombinatorialization without further manipulation. (Item 25) The method of any one of the above items, wherein the mutation is introduced into the entire plasmid DNA (raw plasmid DNA and / or mutated plasmid DNA). (Item 26) The method of any one of the above items, wherein the mutation is induced by PCR or a mutagenic chemical. (Item 27) The method of any one of the above items, wherein the mutation is introduced into a specific region within the plasmid DNA. (Item 28) The method of any one of the above items, wherein the mutation is induced by genome editing, base editing, or prime editing. (Item 29) The method of any one of the above items, wherein the plasmid DNA containing the mutation is prepared by the OGAB® method. (Item 30) The method according to any one of the preceding items, wherein the step of constructing a combinatorial library is the Combi-OGAB (registered trademark) method. (Item 31) A method for producing a substance in a host organism, the method comprising: (A) inducing mutations in plasmid DNA, (B) creating a combinatorial library of the mutations by the Combi-OGAB (registered trademark) method, (C) arbitrarily selecting and culturing host organisms harboring the plasmid DNA created in the combinatorial library, and selecting a plurality of strains based on arbitrary criteria, and (D) repeating any of (A), (A) to (C), or (B) to (C) in any order. (Item 32) The method according to any one of the preceding items, wherein the host organism is Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacteria, Corynebacterium, animal cells, or plants. (Item 33) The method according to any one of the preceding items, wherein the gene cluster carried by the plasmid DNA is related to substance production. (Item 34) The method according to any one of the preceding items, wherein the substance production includes DNA production.(Item 35) The method according to any one of the above items, wherein the gene cluster carried by the plasmid DNA is a nonribosomal peptide synthetase (NRPS) or a polyketide synthase (PKS). (Item 36) The method according to any one of the above items, wherein the plasmid DNA is an adeno-associated virus (AAV) production plasmid.

[0009] According to one embodiment of the present disclosure, a desired gene group, such as a biosynthetic enzyme gene group, can be successfully constructed, and a combinatorial library of promoters for each gene can be created, with appropriate mutations included.

[0010] It is contemplated that the present disclosure may provide one or more of the above-described features in combinations other than those explicitly stated. Still further embodiments and advantages of the present disclosure will be recognized by those skilled in the art upon reading and understanding the following detailed description, if necessary.

[0011] According to the present disclosure, it is possible to rapidly and efficiently obtain strains that exhibit desired properties. The present disclosure also provides a method for creating microbial strains that exhibit desired properties, such as improved production yield, by utilizing a desired gene cluster, such as genes constituting a gene cluster for natural product biosynthetic enzymes, and by incorporating appropriate mutations, and thereby providing potentially novel substances, etc.

[0012] FIG. 1 is a diagram showing the process of preparing a mutated plasmid from a starting plasmid through PCR and enrichment using Bacillus subtilis. FIG. 2 is a diagram showing the convergence of Gramicidin S production after introducing mutations by PCR and undergoing a screening cycle with Combi-OGAB (registered trademark). FIG. 3-1 is a diagram showing the effectiveness of crossing mutations with Combi-OGAB (registered trademark). FIG. 3-2 is a diagram showing the effectiveness of crossing mutations with Combi-OGAB (registered trademark). FIG. 4 is a diagram showing the convergence of Gramicidin S production after reintroducing mutations by PCR and undergoing a screening cycle with Combi-OGAB (registered trademark). FIG. 5 is a diagram comparing the production amounts of each strain after reverting the mutations in each SfiI-treated fragment of the P2nd-B2 strain. Figure 6 shows the increase in Pyrone-OMe production after introducing mutations by PCR and creating a combinatorial library using Combi-OGAB (registered trademark). Figure 7 shows the effects on the accumulation efficiency by Bacillus subtilis and the mutation rate in the plasmid when Mn in the PCR reaction solution was 0 μM, 250 μM, or 500 μM. Figure 8 shows the process of newly introducing two SfiI recognition sequences into grsB of pGETS151-GS 3rd-C2.

[0013] The present disclosure will now be described with reference to the best mode. Throughout this specification, singular expressions should be understood to include the plural concept unless otherwise specified. Therefore, singular articles (e.g., "a," "an," "the," etc. in English) should be understood to include the plural concept unless otherwise specified. Furthermore, it should be understood that terms used in this specification are used in the sense commonly used in the art unless otherwise specified. Therefore, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. In the event of conflict, the present specification (including definitions) will prevail.

[0014] The following provides definitions of terms particularly used in this specification and / or explains basic technical content as appropriate.

[0015] As used herein, "combinatorial library construction" refers to the creation of a library by creating a comprehensive combination of multiple substances (e.g., multiple promoters). Combinatorial library construction is a technique for consolidating diverse gene sequences into a single library. When using plasmid DNA, the procedure is as follows: First, a target gene is selected based on the research objective. Next, each mutant of this target gene is synthesized or amplified by PCR. In this step, various mutagenesis techniques are used to ensure genetic diversity. For example, combinatorial library construction involves constructing a combinatorial plasmid library in which the promoter sequence portion of the plasmid is shuffled. In an exemplary technique, vector plasmid DNA is digested with a restriction enzyme to prepare it for ligation. Each mutant gene is ligated into a plasmid, inserting a different gene sequence into each plasmid. This ligation reaction creates a library of plasmids carrying diverse genes. Next, the ligated plasmids are transformed into Escherichia coli and can be selected using a selection technique such as antibiotic resistance. The selected clones can be individually cultured, and plasmid DNA can be extracted from each. The extracted plasmid DNA is pooled to construct a combinatorial library, which contains many clones with genetic diversity.

[0016] After mutagenesis, sequencing can be performed. The gene sequence of the seed plasmid is sequenced to confirm the genetic mutations in each clone. Finally, screening is performed to select clones with the desired properties from the library. This allows for efficient identification of genetic mutants with specific functions or properties.

[0017] As used herein, the term "combinatorial library generation procedure" refers to any procedure for generating a combinatorial library. Examples include, but are not limited to, the OGAB (registered trademark) method, Combi-OGAB (registered trademark) method, Golden Gate method, overlap extension PCR, DNA shuffling, iPac method, and synthesis techniques using a DNA synthesizer.

[0018] As used herein, the term "desired properties" refers to properties that are beneficial in substance production. Desired properties include the ability to produce a target substance, the ability to produce a high level of a target substance, the ability to retain the target substance within the bacterial cell and not secrete it, the ability to secrete the target substance outside the bacterial cell, the ability to be easily cultured with the production of a target substance (for example, growing under general culture conditions (medium, temperature, culture time, etc.)), the ability to produce multiple target substances in a desired ratio, the ability to produce a target substance without producing harmful substances (no toxicity, etc.), the ability to produce a robust amount of target substance, the ability to prevent defects in genes that produce the target substance during the culture period, the ability to control the production of a target substance, the ability to suppress the production of substances other than the target substance, affinity, Examples of such a property include, but are not limited to, the ability to suppress the production of by-products and / or impurities that are difficult to separate by ion exchange, gel filtration, reverse phase, normal phase, chiral chromatography, distillation, crystallization, etc.; the ability to produce a low amount of empty capsids in viral vector production (for example, an empty capsid rate of 80% or less, preferably 70% or less, more preferably 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, and most preferably 0%); stability in which no deletions occur in plasmids or genes; and stability in which gene products such as mRNA and proteins are produced and function correctly.

[0019] As used herein, the term "the property of overproducing a target substance" refers to a property of producing a higher amount of target substance than a naturally occurring strain. A naturally occurring strain refers to a strain in which a substance-producing gene cluster without sequence engineering (without promoter replacement, etc.) has been introduced into a production host. The property of overproducing a target substance can be, but is not limited to, a property of producing a higher amount of target substance than the naturally occurring strain, usually by at least 10% or more, preferably by 20% or more, 30% or more, 40% or more, or 50% or more.

[0020] As used herein, the term "production host" refers to an organism used to produce a target substance using genetic engineering techniques or a heterologous expression system. Production hosts range from a variety of sources, including microorganisms (e.g., bacteria, yeast, and filamentous fungi), plant cells, and animal cells, and are selected depending on the type of target substance and the production process. After a target gene is introduced into these hosts, they are capable of expressing the gene and efficiently producing proteins, metabolites, or other products. Production hosts are widely used for industrial-scale substance production, including the manufacture of pharmaceuticals, biofuels, and enzymes.

[0021] As used herein, the term "ease of culturing resulting in the production of a target substance" refers to the ability of the strain to grow and produce a target substance under culture conditions commonly used for the strain being cultured. For example, the term indicates that the strain can be cultured under common culture conditions, such as 24 hours of shaking culture at 37°C using a medium commonly used for microbial culture, such as LB medium, a non-selective medium, and is not "difficult to cultivate." Furthermore, the ease of culturing also includes conditions under which the strain can be cultured in a general laboratory (BSL1), rather than in a Biosafety Level (BSL) 2 or BSL 3 laboratory, which requires special equipment for handling the strain in terms of pathogenicity, etc.

[0022] As used herein, the term "robust production of a target substance" refers to a highly reproducible property that always shows the same production amount when production is examined. As one method for verifying robustness, for example, the robustness of experimental results can be evaluated by calculation using the Fano factor described in ACS Synth. Biol., 11, 1686-1691. (2022).

[0023] As used herein, the term "the property of not causing gene deficiency" refers to the property of not causing a gene or a part thereof to be defective due to the deletion of all or part of the nucleic acid encoding the gene, nor causing a loss of gene function due to a mutation in part of the nucleic acid encoding the gene.

[0024] As used herein, the term "starting plasmid" refers to a plasmid that serves as a starting material for preparing a seed plasmid.

[0025] As used herein, the term "seed plasmid (DNA)" is used interchangeably with "mutation-introduced plasmid (DNA)" and refers to a plasmid obtained by subjecting a starting plasmid to some kind of mutation-introducing treatment. Although a plasmid containing no mutation may be produced by the mutation-introducing treatment, such a plasmid is included in the seed plasmid because it is a plasmid produced by the treatment. The seed plasmid refers to a plasmid DNA that is subsequently subjected to a combinatorial library creation procedure.

[0026] As used herein, the term "seed plasmid (DNA) population" is used interchangeably with the term "mutagenesis plasmid (DNA) population" and refers to a mixture of seed plasmids (DNA).

[0027] As used herein, the term "a step of preparing a structure suitable for a combinatorial library construction procedure" refers to an optional step of preparing an appropriate structure that is considered necessary for plasmid DNA when performing a specific combinatorial library construction procedure. Specifically, this optional step refers to a step of preparing a structure that allows plasmid DNA to be divided into multiple DNA fragments, more specifically, to introducing restriction enzyme recognition sites into the plasmid DNA, and even more specifically, includes a step of preparing the restriction enzyme recognition sites so that the cohesive end sequences obtained as a result of cleaving the restriction enzyme recognition sites with a restriction enzyme are unique sequences at each restriction enzyme recognition site.

[0028] As used herein, "generating a seed plasmid population" refers to generating a seed plasmid population by some method. In a specific example, this refers to mixing multiple types of mutated plasmids (seed plasmids) obtained by introducing an SfiI recognition sequence and / or random mutations into a starting plasmid. Examples of mixing methods include, but are not limited to, mixing the plasmid DNAs in equimolar amounts, mixing at any molar ratio, mixing without considering the molar ratio, and collecting and mixing transformants. The method may further include a step of preparing a structure suitable for the combinatorial library construction procedure, or the step of generating the seed plasmid population may further include a step of preparing a structure suitable for the combinatorial library construction procedure.

[0029] As used herein, the term "gene cluster" refers to a region of the genome in which multiple genes related to the same biological function are contiguously arranged. These genes are often expressed in a coordinated manner under a common regulatory mechanism, and are responsible for a specific biological process or metabolic pathway. For example, an operon in bacteria is an example of a gene cluster in which multiple genes controlled by a promoter are transcribed as a single mRNA. Gene clusters involved in the biosynthesis of antibiotics are also well known, and an example of this is the streptomycin biosynthetic gene cluster in bacteria of the genus Streptomyces. This enables the efficient production of required metabolic products under specific conditions.

[0030] As used herein, "recombinatorially viable" refers to the ability to rearrange existing genes or gene sequences into new combinations. This allows for the creation of greater diversity and the generation of gene variants with new functions or properties. This technology is expected to have a variety of applications, such as improving drug resistance or enzyme activity, or developing new metabolic pathways, by rearranging or recombining genes for specific purposes. Recombinatorially viable is particularly important in the fields of evolutionary molecular engineering and biotechnology. In the present disclosure, non-limiting examples include the ability to directly use the plasmids of strains obtained after library construction and strain selection using the Combi-OGAB (registered trademark) method for the construction of a next library. This reusability of the plasmids is an example of the property of being "recombinatorially viable."

[0031] As used herein, the term "mutation" refers to a different nucleotide sequence in a seed plasmid compared to the nucleotide sequence of a starting plasmid. When a seed plasmid is prepared by incorporating a structure suitable for a combinatorial library construction procedure into a starting plasmid, the structure is also included in the mutation.

[0032] As used herein, the term "random mutation" refers to a change in a gene or genetic material that occurs randomly due to external factors or natural processes. Examples include randomly mutating one or more target bases in a plasmid, randomly mutating a target sequence region, and randomly mutating an entire plasmid. Random mutation typically encompasses changes in the nucleotide sequence of a gene (including insertions, deletions, and substitutions), but also encompasses other changes, such as changes in methylation status. Random mutations are typically unpredictable, and mutations can result in alterations to the DNA base sequence. For example, in genetic engineering, methods exist for introducing random mutations into genes using chemical mutagens or radiation. This method enables functional analysis of specific genes and the creation of novel proteins. PCR can also be used to amplify specific gene regions and introduce mutations resulting from DNA polymerase amplification errors. Screening of these randomly introduced mutations can then be used to create a wide variety of mutants. Random mutations are widely used in genetic engineering and biotechnology as an important tool for discovering the functions and properties of novel organisms.

[0033] As used herein, "entire plasmid DNA" refers to the base sequence of the entire plasmid DNA molecule. Plasmids are circular, double-stranded DNA molecules that exist within the cells of bacteria and other organisms and replicate independently of the genetic information of the host cell. In other words, "introducing a mutation into the entire plasmid" refers to introducing not only the target gene / gene cluster and its corresponding transcription factors (promoter, enhancer, terminator, etc.), but also all base sequences on the plasmid, including the vector, as the target for mutation introduction.

[0034] As used herein, the term "specific region within plasmid DNA" refers to a specific base sequence or gene on a plasmid DNA molecule. This specific region within plasmid DNA includes various functional sequences, such as an origin of replication (ori), a selectable marker gene (e.g., an antibiotic resistance gene), a promoter, a reporter gene, a cloning site, a multiple cloning site (MCS), and a gene or gene cluster of interest. For example, the origin of replication (ori) is the replication origin of the plasmid, and a selectable marker gene confers antibiotic resistance, allowing for the selection of transformants. Furthermore, a multiple cloning site (MCS) is a region containing multiple consecutive restriction enzyme sites and is used to facilitate the insertion and deletion of genes. Specific regions within plasmid DNA play an important role in genetic engineering and molecular biology experiments. Targeting these regions for gene insertion, deletion, substitution, mutagenesis, and other purposes enables gene function analysis, gene expression control, and the introduction of novel genes.

[0035] As used herein, "PCR" stands for polymerase chain reaction and refers to a technology for exponentially amplifying a specific DNA sequence. PCR is a fundamental technology for mass-replicating minute amounts of DNA in a sample, facilitating analysis and manipulation. The basic procedure proceeds as follows: First, the double-stranded target DNA is dissociated into single strands by thermal denaturation (e.g., about 95°C). Next, if necessary, the temperature is lowered to allow annealing (e.g., about 50-65°C) in which primers bind to specific sequences. Finally, elongation (e.g., about 72°C) is performed, in which new DNA strands are extended using a thermostable DNA polymerase. By repeating this cycle, it is possible to amplify a specific region of the target DNA millions of times, but this is not limiting.

[0036] One application of PCR is DNA amplification. This technology allows minute amounts of DNA to be amplified to sufficient amounts, making it easy to obtain materials needed for treatment, diagnosis, analysis, and research. For example, in the medical field, minute amounts of DNA can be amplified to detect pathogens, produce nucleic acid molecules for genetic diagnosis, and produce therapeutics. Furthermore, in molecular biology research, specific gene sequences can be amplified and then used for cloning and sequencing, contributing to the elucidation of gene structure and function. Furthermore, random replication errors during PCR can be exploited to introduce random mutations into the target gene sequence (error-prone PCR). This also contributes to the functional analysis and improvement of the gene. PCR has become an essential technology for these applications.

[0037] As used herein, the term "mutagenic chemical" refers to a chemical substance capable of inducing mutations in nucleic acids such as DNA. These substances induce mutations by inducing mutations in gene sequences and altering the genetic information of cells or organisms. Mutagenic chemicals can induce various forms of genetic changes, such as specific base pair substitutions, deletions, or insertions, or chromosomal rearrangements. Representative mutagenic chemicals include ethyl methanesulfonate (EMS), nitrosamines, benzopyrene, and acridine orange. These chemicals induce mutations by interfering with DNA replication and repair processes. For example, ethyl methanesulfonate (EMS) binds to guanine bases, causing alkylation and introducing mutations through mismatch repair. Mutagenic chemicals are used as important tools in genetic research, carcinogenicity research, drug development, and other fields. These substances can be used to study the function of specific genes and evaluate the effectiveness of new drugs.

[0038] As used herein, "genome editing" refers to a technique for precisely and efficiently modifying specific DNA sequences. This technique allows for the insertion, deletion, modification, or replacement of genes. Genome editing uses nucleases that target the sequence of a desired gene and cleave the DNA at specific sites. The most common genome editing tools include CRISPR-Cas9, TALEN, and ZFN (zinc finger nuclease).

[0039] As used herein, "base editing" refers to a technique for changing specific bases in nucleic acids (e.g., DNA). There are four types of bases: adenine (A), cytosine (C), guanine (G), and thymine (T) (or uracil (U) in the case of RNA). Base editing technology, based on gene editing technologies such as CRISPR-Cas9, makes it possible to change specific base pairs to other base pairs. For example, C can be changed to T, and A to G. A major feature of this technology is its ability to directly change specific bases without cleaving the DNA double strand. This allows for targeted changes to be made with greater precision than conventional gene editing technologies, reducing the risk of off-target effects (unintentional gene modification). Furthermore, base editing is expected to be a powerful tool for correcting specific genetic mutations that cause disease.

[0040] As used herein, "prime editing" refers to an innovative gene editing technology for precisely and efficiently modifying specific sequences in nucleic acids (e.g., DNA). This technology is based on the conventional CRISPR-Cas9 system and enables highly accurate DNA modification by using reverse transcriptase. The prime editing process begins by designing a prime editing guide RNA (pegRNA), which binds to the target site in the DNA. Next, Cas9 nickase cuts a portion of the DNA, and reverse transcriptase inserts a new DNA sequence into that site. This method allows for precise replacement of the target gene sequence.

[0041] As used herein, the term "OGAB® method" stands for "Ordered Gene Assembly in Bacillus subtilis" and refers to a technique for efficiently assembling multiple DNA fragments in an orderly manner using Bacillus subtilis. This technique allows multiple gene fragments to be combined in a precise order, facilitating the construction of gene clusters and large-scale genetic engineering. The OGAB® method process begins by ligating the target DNA fragments to each fragment in a tandem repeat configuration at overlapping protruding ends. Next, utilizing the natural transformation ability of Bacillus subtilis, the tandem repeat ligation product is taken up into the cell, where plasmid DNA is assembled. This process allows for the creation of plasmid DNA in which multiple gene fragments are simultaneously linked in the correct order (Tsuge, K., et al., Sci. Rep., 2015, 5, 10655). The advantages of the OGAB® method include high assembly efficiency and precision, short setup times, and rapid construction of large gene clusters. This technology is expected to be applied in a wide range of fields, including synthetic biology, genetic engineering, pharmaceutical development, and the improvement of microbial metabolic pathways. Actual applications include the reconstruction of complex metabolic pathways, the development of novel biologics, and detailed analysis of gene function.

[0042] As used herein, the "Combi-OGAB® method" refers to a technique that further improves on the OGAB® method, a technology for efficiently assembling multiple genes, and enables the construction of a chimeric plasmid library (Japanese Patent No. 7101431). The Combi-OGAB® method allows for simultaneous testing of different gene combinations, making it extremely useful for analyzing gene function and constructing new metabolic pathways. This process involves first mixing equimolar amounts of multiple seed plasmids and cleaving them with restriction enzymes. Subsequently, ligation is performed in the same manner as described above to obtain tandem repeat ligation products, which are shuffled DNA fragments with identical overlapping protruding ends. Next, utilizing the transformation ability of Bacillus subtilis, these tandem repeat ligation products are incorporated into cells, and a chimeric plasmid library is assembled through natural recombination. Advantages of the Combi-OGAB® method include the ability to rapidly test diverse gene combinations, high assembly efficiency, and the ability to utilize the excellent transformation ability of Bacillus subtilis. This technology is a particularly powerful tool for analyzing complex gene networks and designing new metabolic pathways, and its practical applications include the optimization of industrial enzymes, the development of new biopharmaceuticals, and even the genetic improvement of agricultural crops.

[0043] As used herein, "substance production" refers to the general process of producing specific organic compounds (e.g., nucleic acids (e.g., DNA, RNA, etc.), proteins, etc.), chemicals, or biomass. This includes bioprocesses that use microorganisms or cells to produce a target substance, chemical processes that manufacture compounds through chemical synthesis, and even methods for extracting specific substances from plants or animals. When a plasmid is used, it refers to a process of producing a specific organic compound, chemical, or biomass using plasmid DNA. Plasmid DNA is a small, circular DNA molecule found in microorganisms such as bacteria and is widely used as a vector for introducing a target gene through genetic engineering. Plasmid design: Plasmid DNA containing a gene for producing a target substance is designed. This includes constructions that include a promoter (a sequence that controls gene expression), a target gene (a gene encoding the substance to be produced), and a selection marker (a gene used to select transformed cells). Transformation: The designed plasmid DNA is introduced into host cells such as bacteria (e.g., E. coli). The host cells then incorporate the plasmid DNA and begin to express the target gene. Fermentation and cultivation: The transformed bacteria are cultured in an appropriate medium and fermentation is carried out. During the fermentation process, bacteria express genes encoded by the plasmid DNA to produce the desired substance. Harvesting and Purification: After fermentation, the produced substance is harvested from the host cells and purified. This process involves cell disruption, filtration, chromatography, and other methods to obtain a highly purified product. The advantages of using plasmid DNA to produce substances include the ability to produce large quantities of the desired substance with high efficiency, the relative ease of genetic manipulation, and the ease of scaling up the production process. Practical applications include the production of pharmaceuticals such as insulin, biofuels, and industrial enzymes. Advances in this technology enable more efficient and cost-effective substance production, playing an important role in fields such as pharmaceuticals, energy, and chemical industries.Gene clusters related to substance production can be genes for biosynthetic enzymes of polyketide synthases (PKS), non-ribosomal peptide synthases (NRPS), ribosomal translation system ribosomally synthesized and post-translationally modified peptides (RiPPs), or terpene biosynthetic enzymes. Therefore, virus production is also included in "substance production," as well as "desired DNA production" such as plasmid DNA for mRNA or adeno-associated virus production, DNA vaccines, DNA storage, and DNA sensors.

[0044] The present disclosure provides a method for producing constructs related to adeno-associated viruses. Adeno-associated viruses (AAVs) are linear, single-stranded DNA viruses of the Parvoviridae family, Dependovirus genus, with viral particles measuring 20-26 nm in diameter. Adenovirus elements are required for AAV propagation. T-shaped hairpin structures called inverted terminal repeats (ITRs) are present at both ends of the AAV genome. These ITRs serve as replication initiation sites and act as primers. These ITRs are also required for packaging into viral particles and integration into the chromosomal DNA of host cells. The left half of the genome contains rep genes encoding nonstructural proteins, i.e., regulatory proteins (Rep78, Rep76, Rep52, and Rep40) that control replication and transcription, while the right half of the genome contains cap genes encoding three structural capsid proteins (VP1, VP2, and VP3). The life cycle of AAV can be divided into latent infection and lytic infection. The former occurs when AAV is infected alone and is characterized by integration into the AAVS1 region (19q13.3-qter) of the long arm of chromosome 19 of the host cell. This integration occurs through non-homologous recombination and involves Rep. It has been reported that Rep78 / Rep76 bind to a base sequence (GAGC repeat sequence) that is commonly present in the AAVS1 region and the Rep binding region of the ITR. Therefore, it is thought that when wild-type AAV infects a target cell, Rep binds to the ITR and AAVS1 of AAV, and site-specific integration of the AAV genome into chromosome 19 occurs via Rep. When a helper virus such as adenovirus is simultaneously infected, and a cell latently infected with AAV is further superinfected with the helper virus, AAV replication occurs, and large amounts of virus are released due to cell destruction (lytic infection).

[0045] In one embodiment, the AAV-derived construct plasmid contains at least one of the nucleic acid sequences necessary for constructing an AAV-derived construct and a desired gene. In one embodiment, the AAV-derived construct plasmid contains the desired gene between the 5' ITR and the 3' ITR. In one embodiment, the AAV-derived construct plasmid contains the desired gene, promoter, and terminator between the 5' ITR and the 3' ITR. In one embodiment, the AAV-derived construct plasmid does not contain one or more (e.g., all) of the nucleic acid sequences necessary for constructing an AAV-derived construct between the 5' ITR and the 3' ITR. In one embodiment, the AAV-derived construct plasmid is configured to function in a producer cell such that the virus-derived construct plasmid and the producer cell contain the nucleic acid sequences necessary for constructing an AAV-derived construct. In one embodiment, one or more (e.g., all) of the genes of the nucleic acid sequences necessary for constructing an AAV-derived construct that are not contained in the virus-derived construct plasmid may be encoded in the chromosome of the producer cell.

[0046] In one embodiment, the nucleic acid sequences necessary to construct an AAV-derived construct may include the 5' ITR, rep, cap, AAP (assembly activating protein), MAAP (membrane-associated accessory protein), and 3' ITR of AAV, and E1A, E1B, E2A, VA, and E4 of adenovirus. In one embodiment, the nucleic acid sequences necessary to construct an AAV-derived construct may include a nucleic acid sequence encoding an AAV structural protein (cap), a nucleic acid sequence encoding an AAV packaging factor (rep), two AAV long terminal repeat sequences (5' ITR, 3' ITR), and nucleic acid sequences encoding functional accessory factors (E1A, E1B, E2A, VA, and E4 of adenovirus). In one embodiment, the AAV-derived construct plasmid may not include one or more of rep, cap, VA, E1A, E1B, E2A, E4, and in one embodiment, the nucleic acid sequences necessary to construct the adenovirus-derived construct include these.

[0047] AAV serotypes 1 to 12 have been reported based on the capsid. The following serotypes have also been reported: rh10, DJ, DJ / 8, PHP.eB, PHP.S, AAV2-retro, AAV2-QuadYF, AAV2.7m8, AAV6.2, rh.74, AAV2.5, AAV-TT, and Anc80. The AAV-derived constructs of the present disclosure may be prepared based on an AAV of a suitable serotype depending on the target tissue. For example, the (serotype):(target tissue) relationship may be selected as follows: (AAV1):(muscle, liver, airway, nerve cells), (AAV2):(muscle, liver, nerve cells), (AAV3):(muscle, liver, nerve cells), (AAV4):(muscle, ependymal cells), (AAV5):(muscle, liver, nerve cells, glial cells, airway), (AAV6):(muscle, liver, airway, nerve cells), (AAV7):(muscle, liver), (AAV8):(muscle, liver), (AAV9):(muscle, liver, airway). Producer cells for producing AAV-derived constructs include cells such as HEK293, HEK293T, HEK293F, Hela, and Sf9. Examples of AAV capsids include wild-type capsids, as well as capsids with targeted mutations (AAV2i8, AAV2.5, AAV-TT, AAV9.HR, etc.), capsids with random mutations (AAV-PHP.B, etc.), and capsids designed in silico (Anc80, etc.). In one embodiment, the AAV-derived constructs of the present disclosure may contain these modified capsids, and the AAV-derived construct plasmids of the present disclosure may be constructed so as to encode these modified capsids. When a nucleic acid sequence is described herein as being derived from a particular serotype, it is intended that the nucleic acid sequence may encode a wild-type capsid or a capsid that has been modified as described above based on the wild-type capsid.

[0048] As used herein, the term "host cell" refers to a cell used to produce a desired substance by introducing a specific gene or plasmid. Host cells have the ability to incorporate foreign DNA through genetic manipulation and produce proteins or other compounds based on that genetic information. It refers to cells (including the progeny of such cells) into which foreign nucleic acids or proteins, or viruses or viral vectors, have been introduced. Host cells are selected based on the type of desired substance and the requirements of the production process. Commonly used host cells include Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacterium, Corynebacterium sp., animal cells, and plants, including Escherichia coli, yeast (Saccharomyces cerevisiae), mammalian cells (e.g., CHO cells, HEK293 cells), plant cells (e.g., tobacco BY-2 cells), and insect cells (e.g., Sf9 cells). Host cells play a central role in each step of gene introduction, expression, and product collection and purification. Selection of appropriate host cells and optimized culture conditions are important for achieving efficient and high-quality production of substances.

[0049] As used herein, the term "arbitrary criteria" refers to criteria or standards established according to specific purposes or requirements for plasmids. These criteria are not necessarily standardized but are used for evaluation and judgment. These criteria are applied in the context of plasmid design, construction, introduction, and functional evaluation. Criteria are flexibly set depending on the use and purpose of the plasmid. For example, in plasmid design, criteria include the selection of specific promoters and selectable markers. In plasmid construction, criteria may be established for the use of specific enzyme sites or reporter genes. In addition, in plasmid introduction, criteria are used to evaluate transformation efficiency and intracellular stability. Criteria are not necessarily established by standardization organizations or regulatory agencies; they can be independently set by researchers and engineers and adjusted according to their purposes. Setting arbitrary criteria can ensure consistency in the performance and function of plasmids under specific conditions and aid in the interpretation and application of results. For example, in a certain experiment, a high expression level in a specific cell line can be used as a criterion. Criteria may also be established to evaluate plasmid stability and sustained gene expression. Using arbitrary criteria can result in more reliable results in research and development using plasmids.

[0050] As used herein, the term "virus-derived construct" refers to a nucleic acid construct that has at least a partial structure derived from a virus or a construct that contains a protein produced therefrom. Examples of virus-derived constructs include viral vectors, virus-like particles (VLPs), oncolytic viruses (including those modified to grow only in cancer cells), and viral replicons (self-replicating viral genomes that have been deficient in the ability to produce virus particles).

[0051] As used herein, the term "viral vector" refers to a construct that has at least a partial structure derived from a virus and is capable of introducing nucleic acid into a target cell. Typically, a viral vector is in the form of a viral particle that contains viral structural proteins and a nucleic acid containing a heterologous gene.

[0052] As used herein, the term "nucleic acid sequence necessary for constructing a virus-derived construct" refers to a nucleic acid sequence (e.g., a combination of nucleic acid sequences) that enables a producer cell containing the nucleic acid sequence to produce a virus-derived construct. In one embodiment, the nucleic acid sequence necessary for constructing a virus-derived construct includes a nucleic acid sequence encoding a viral structural protein, a nucleic acid sequence encoding a packaging factor, two viral terminal repeat sequences, and a nucleic acid sequence encoding a functional cofactor. A "nucleic acid sequence encoding a functional cofactor" may, for example, include or essentially consist of a nucleic acid sequence necessary for constructing any virus-derived construct other than a nucleic acid sequence encoding a viral structural protein, a nucleic acid sequence encoding a packaging factor, and two viral terminal repeat sequences. Each of these nucleic acid sequences may differ for each virus-derived construct. Details of these sequences are described elsewhere in this specification. In a preferred embodiment, the nucleic acid sequence necessary for constructing a virus-derived construct utilizes advantageous sequences common to AAV and adenovirus, and more preferably, advantageous sequences for AAV. Such advantageous sequences common to AAV and adenovirus are characterized by including a desired gene between inverted terminal repeats (ITRs) and optionally including functional cofactors rep and cap or L1, L2, L3, L4, and L5, while advantageous sequences for AAV include a desired gene between ITRs and optionally including functional cofactors rep and cap. The functional cofactors rep and cap may be located outside the two ITRs. Examples of such sequences include, but are not limited to, the sequences of specific functional cofactors E1A, E1B, E2A, E2B, E3, and E4.

[0053] As used herein, the term "packaging factor" includes proteins that encapsulate the viral genome in structural proteins, and examples thereof include rep and ψ. Packaging factors that can be used include, but are not limited to, those derived from adenovirus serotypes 1 to 52 or adeno-associated virus serotypes 1 to 12, or variants thereof (rh10, DJ, DJ / 8, PHP.eB, PHP.S, AAV2-retro, AAV2-QuadYF, AAV2.7m8, AAV6.2, rh.74, AAV2.5, AAV-TT, Anc80, etc.). These proteins may be naturally occurring or may have artificially mutated proteins. Those that have had "artificial mutations" introduced can be prepared, for example, by introducing appropriate mutations into the sequence of a naturally occurring protein described in a known literature.

[0054] As used herein, "structural protein" refers to a protein produced from a viral gene, at least a portion of which is present on the surface of the virus or a virus-derived construct (e.g., across the envelope). Structural proteins may be responsible for cell infectivity.

[0055] As used herein, "capsid" refers to a protein produced from a viral gene that is present on the surface of a virus or virus-derived construct (optionally enclosed in an envelope), and is also referred to as a "capsid protein."

[0056] As used herein, the term "repeated sequence" or "tandem repeat" (such as a nucleic acid sequence) refers to a nucleic acid sequence in an organism genome in which the same sequence is found repeatedly (especially several times or more). Any repetitive sequence used in the art can be used in the present disclosure. Typically, a promoter sequence appears repeatedly.

[0057] As used herein, the term "terminal repeat sequence" refers to a nucleic acid sequence in an organism's genome in which the same sequence is repeated (especially several times or more), and is present at the end of the nucleic acid sequence. Examples of terminal repeat sequences include, but are not limited to, inverted terminal repeats (ITRs) or long terminal repeats (LTRs). Terminal repeat sequences may be derived from adenovirus serotypes 1 to 52 or adeno-associated virus serotypes 1 to 12, or variants thereof (rh10, DJ, DJ / 8, PHP.eB, PHP.S, AAV2-retro, AAV2-QuadYF, AAV2.7m8, AAV6.2, rh.74, AAV2.5, AAV-TT, Anc80, etc.). Terminal repeat sequences may be naturally occurring or may be those into which mutations have been artificially introduced. Those that have been "artificially mutated" can be prepared, for example, by introducing appropriate mutations into the sequence of a natural product described in a known literature.

[0058] As used herein, the term "functional cofactor" refers to a gene that supports the amplification of a virus that cannot grow alone. In the present disclosure, functional cofactors may include, but are not limited to, nucleic acid sequences encoding the packaging factors of the virus, as well as nucleic acid sequences necessary for the construction, growth, activity enhancement, or toxicity reduction of any virus-derived construct other than the two terminal repeat sequences of the virus. Examples of functional cofactors that can be used include, but are not limited to, E1A, E1B, E2A, E2B, E4, RPE, WRPE, PPT, oPRE, enhancer, insulator, silencer sequences, and the like.

[0059] As used herein, "loaded nucleic acid" refers to nucleic acid carried by a viral vector. Whether or not a nucleic acid is loaded can be confirmed by treating a viral vector preparation with DNase to degrade nucleic acids outside the capsid, inactivating the DNase, and then examining the nucleic acid extracted from the capsid. Whether or not a nucleic acid is loaded can also be confirmed by examining the sequence of the resulting nucleic acid product (preferably full-length or nearly full-length).

[0060] As used herein, "plasmid" refers to a circular DNA that exists separately from chromosomes in a cell or that exists separately from chromosomes when introduced into a cell.

[0061] As used herein, the term "nucleic acid sequence that promotes plasmid replication" is used interchangeably with the term "nucleic acid sequence that promotes plasmid amplification" and refers to any nucleic acid sequence that, when introduced into a host cell (e.g., Bacillus subtilis), promotes the replication (in this context, the same as amplification) of a plasmid present in the host cell. Preferably, but not limited to, this nucleic acid sequence that promotes plasmid replication is operably linked to a nucleic acid sequence encoding a target plasmid. Details of the nucleic acid sequence that promotes plasmid replication, the target plasmid, and their relationship are described elsewhere in this specification. Examples of nucleic acid sequences that promote plasmid replication include nucleic acid sequences containing an origin of replication that operates in the target host cell (e.g., Bacillus subtilis). A nucleic acid sequence that promotes plasmid amplification refers to a DNA fragment that has an origin of replication and can replicate independently of chromosomal DNA. Whether these DNA fragments are replicable can be confirmed by, for example, linking a selectable marker gene such as a drug resistance gene to these DNA fragments, culturing them under selective conditions such as a drug, purifying the plasmid DNA, and observing the DNA band by electrophoresis. In addition to the replication origin, the plasmid may contain a rep protein gene that induces a host DNA replication enzyme at the replication origin, a partitioning mechanism gene for ensuring partitioning of the plasmid into daughter cells, a selection marker gene, etc. The replication origin may be fused within the rep protein gene.

[0062] As used herein, "packaging cells" refers to cells for producing vector plasmids.

[0063] As used herein, the terms "vector plasmid" or "viral vector plasmid" refer to a plasmid containing a gene to be carried in the viral vector and used to produce the viral vector in a host cell. In one embodiment, the vector plasmid contains, between two terminal repeat sequences, a sequence of a gene (desired gene) to be expressed in the target cell of the viral vector, and may also contain a promoter and other elements (e.g., an enhancer, a terminator, etc.).

[0064] In this disclosure, the term "all-in-one plasmid" refers to an integrated plasmid that consolidates multiple genes into a single plasmid. This plasmid contains all the necessary genes, such as a replication origin, expression control elements, selection markers, and all of the genes of interest, making it possible to easily and efficiently conduct gene transfer and expression experiments.

[0065] The method for generating the mutagenesis plasmid (= Combi-OGAB® seed plasmid) of the present disclosure can be used to produce an "all-in-one plasmid" for use in producing AAV vectors by single transfection (see Examples).

[0066] As used herein, the terms "protein," "polypeptide," and "peptide" are used interchangeably to refer to a polymer of amino acids of any length. The polymer may be linear, branched, or cyclic. The amino acids may be natural, non-natural, or modified. The term also includes naturally or artificially modified polymers. Such modifications include, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification (e.g., conjugation with a labeling component).

[0067] Unless otherwise indicated, a particular nucleic acid sequence is also intended to include the explicitly indicated sequence, as well as conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences. Specifically, degenerate codon substitutions can be achieved by creating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. For example, variants based on specific wild-type sequences, such as adenovirus serotypes 1 to 52 and adeno-associated virus serotypes 1 to 12, include not only known variants (e.g., rhlO, DJ, DJ / 8, PHP.eB, PHP.S, AAV2-retro, AAV2-QuadYF, AAV2.7m8, AAV6.2, rh.74, AAV2.5, AAV-TT, Anc80, etc.), but also nucleic acids comprising a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the original sequence.

[0068] As used herein, the term "gene" refers to a nucleic acid moiety that performs a specific biological function. These biological functions include encoding a polypeptide or protein, encoding non-protein-coding functional RNA (e.g., rRNA, tRNA, microRNA (miRNA), lncRNA), controlling the production of polypeptides, proteins, or non-protein-coding functional RNA, specifically binding to a specific protein, and controlling nucleic acid cleavage or replication. Therefore, as used herein, genes include not only nucleic acid moieties that encode proteins or non-protein-coding functional RNA, but also transcriptional / translational regulatory sequences such as promoters, terminators, enhancers, insulators, silencers, replication origins, and internal ribosome entry sites, as well as nucleic acid moieties required for packaging into viral particles. As used herein, "gene product" may refer to a polypeptide, protein, or non-protein-coding functional RNA encoded by a gene.

[0069] When referring to the number of genes herein, one gene refers to a gene having a contiguous sequence in the form that is normally present (highest frequency or probability of 50% or more) in the genome of a certain organism. For example, two exons encoding a certain protein can be two genes. For example, if a promoter sequence and a protein-encoding sequence form a contiguous sequence, the nucleic acid portion including the promoter sequence and the protein-encoding sequence can be one gene. For example, if a protein that becomes functional when cleaved is encoded by a contiguous sequence on the genome, this protein can be encoded by one gene. When referring to a gene from the perspective of function, the nucleic acid sequence does not need to be a contiguous sequence; for example, multiple exons encoding a certain protein can be collectively referred to as the gene for that protein.

[0070] As used herein, "operably linked" means that the expression (operation) of a desired sequence is placed under the control of a transcriptional / translational regulatory sequence (e.g., a promoter, an enhancer, etc.) or a translational regulatory sequence. In order for a promoter to be operably linked to a gene, the promoter is usually placed immediately upstream of the gene, but is not necessarily placed adjacent to the gene.

[0071] As used herein, the term "transcriptional / translational regulatory sequence" collectively refers to promoter sequences, polyadenylation signals, transcription termination sequences, upstream regulatory domains, replication origins, enhancers, IRES, and the like, which cooperate to enable the replication, transcription, and translation of a coding sequence in a recipient cell. Not all of these transcriptional / translational regulatory sequences necessarily need to be present, as long as the selected coding sequence is capable of being replicated, transcribed, and translated in an appropriate host cell. Those skilled in the art can easily identify regulatory nucleic acid sequences from publicly available information. Furthermore, those skilled in the art can identify transcriptional / translational regulatory sequences that are applicable to the intended use, for example, in vivo, ex vivo, or in vitro.

[0072] As used herein, "promoter" refers to a segment of a nucleic acid sequence that controls the transcription of an operably linked nucleic acid sequence. A promoter contains a specific sequence that is sufficient for recognition, binding, and transcription initiation by RNA polymerase. A promoter may also contain a sequence that regulates RNA polymerase recognition, binding, or transcription initiation.

[0073] As used herein, "enhancer" refers to a segment of a nucleic acid sequence that functions to increase the efficiency of expression of a gene of interest.

[0074] As used herein, the term "silencer" refers to a segment of a nucleic acid sequence that has the function of reducing the expression efficiency of a gene of interest, in contrast to an enhancer.

[0075] As used herein, "insulator" refers to a segment of nucleic acid sequence that has a cis-regulatory function of regulating the expression of genes located at distant positions in the DNA sequence.

[0076] As used herein, "terminator" refers to a segment of nucleic acid sequence located downstream of a protein-coding region and involved in terminating transcription when the nucleic acid is transcribed into mRNA.

[0077] As used herein, "origin of replication" refers to a segment of a nucleic acid sequence at which replication is initiated by the binding of a protein that recognizes that nucleic acid sequence (e.g., the initiator DnaA protein) or by the synthesis of RNA, which partially unwinds the DNA double helix.

[0078] As used herein, an internal ribosome entry site ("IRES") refers to a nucleic acid segment that facilitates the entry or retention of a ribosome during translation of a downstream nucleic acid sequence.

[0079] As used herein, the term "homology" of nucleic acids refers to the degree of identity between two or more nucleic acid sequences, and generally, "homology" refers to a high degree of identity or similarity. Therefore, the higher the homology between two nucleic acids, the higher the identity or similarity between their sequences. "Similarity" is a numerical value that takes into account not only identity but also similar bases, where similar bases refer to partial matches in mixed bases (e.g., R = A + G, M = A + C, W = A + T, S = C + G, Y = C + T, K = G + T, H = A + T + C, B = G + T + C, D = G + A + T, V = A + C + G, N = A + C + G + T). Whether two types of nucleic acids are homologous can be determined by direct sequence comparison or by hybridization under stringent conditions. When two nucleic acid sequences are compared directly, the genes are homologous if there is typically at least 50% identity between the nucleic acid sequences, preferably at least 70% identity, and more preferably at least 80%, 90%, 95%, 96%, 97%, 98% or 99% identity between the nucleic acid sequences.

[0080] Amino acids may be referred to herein by either their commonly known three-letter symbols or the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides may also be referred to by their commonly accepted one-letter codes. Herein, comparisons of amino acid sequence and base sequence similarity, identity, and homology are calculated using the sequence analysis tool BLAST with default parameters. Identity searches can be performed, for example, using NCBI's BLAST 2.10.1+ (published June 18, 2020). Identity values ​​herein generally refer to values ​​obtained when aligned using the above-mentioned BLAST under default conditions. However, if a higher value is obtained by changing parameters, the highest value is used as the identity value. When identity is evaluated in multiple regions, the highest value among them is used as the identity value. Similarity is a numerical value that takes into account not only identity but also similar amino acids.

[0081] Unless otherwise specified, reference herein to a biological substance (e.g., a protein, nucleic acid, or gene) is understood to also refer to variants of that biological substance (e.g., variants with modifications in the amino acid sequence) that perform biological functions similar to, but not necessarily to, the biological function of the biological substance. Such variants may include fragments of the original molecule, molecules that are at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identical to the amino acid or nucleic acid sequence of the original biological substance over the same size, or to the sequence of the original molecule as aligned by computer homology programs known in the art. Variants may include molecules with modified amino acids (e.g., modifications due to disulfide bond formation, glycosylation, lipidation, acetylation, or phosphorylation) or modified nucleotides (e.g., modifications due to methylation).

[0082] As used herein, a "corresponding" amino acid or nucleic acid refers to an amino acid or nucleotide in a polypeptide or polynucleotide molecule that has or is predicted to have the same function as a given amino acid or nucleotide in a reference polypeptide or polynucleotide. In particular, in the case of an enzyme molecule, this refers to an amino acid that is located at a similar position in the active site and contributes similarly to catalytic activity. For example, in the case of an antisense molecule, this may be a similar portion in an orthologue corresponding to a specific portion of the antisense molecule. A corresponding amino acid may be, for example, a specific amino acid that is cysteinylated, glutathionylated, forms an S-S bond, oxidized (e.g., methionine side chain oxidation), formylated, acetylated, phosphorylated, glycosylated, myristylated, or the like. Alternatively, the corresponding amino acid may be an amino acid responsible for dimerization. Such a "corresponding" amino acid or nucleic acid may be a region or domain spanning a certain range. Therefore, in such cases, it is referred to herein as a "corresponding" region or domain.

[0083] As used herein, a "corresponding" gene (e.g., a polynucleotide sequence or molecule) refers to a gene (e.g., a polynucleotide sequence or molecule) that has or is predicted to have the same function in a certain species as a given gene in a species used as a reference for comparison. When multiple genes with such function exist, the term refers to genes that have the same evolutionary origin. Thus, a gene corresponding to a certain gene may be an ortholog of that gene. For example, the cap of serotype 1 AAV may correspond to the cap of serotype 2 AAV. For example, a corresponding gene in a certain virus can be found by searching a sequence database for that virus using the gene sequence of the reference virus for the corresponding gene as a query sequence.

[0084] In accordance with the present disclosure, the term "activity" as used herein refers to the function of a molecule in the broadest sense. Activity generally includes, but is not limited to, the biological, biochemical, physical, or chemical function of a molecule. Activity includes, for example, enzymatic activity, the ability to interact with other molecules, and the ability to activate, promote, stabilize, inhibit, suppress, or destabilize the function of other molecules, stability, and the ability to localize to a specific subcellular location. Where applicable, the term also relates to the function of a protein complex in the broadest sense.

[0085] As used herein, the term "biological function," when referring to a gene or a nucleic acid molecule or polypeptide related thereto, refers to a specific function that the gene, nucleic acid molecule, or polypeptide may have in a living organism. Examples of such functions include, but are not limited to, the ability to recognize a specific cell surface structure, enzymatic activity, and the ability to bind to a specific protein. In the present disclosure, examples of such functions include, but are not limited to, the ability of a promoter to be recognized in a specific host cell. As used herein, a biological function can be exerted through "biological activity." As used herein, "biological activity" refers to the activity that a certain factor (e.g., a polynucleotide, a protein, etc.) may have in a living organism, including the ability to exert various functions (e.g., transcription-promoting activity), including, for example, the activity of activating or inactivating another molecule through interaction with another molecule. For example, if a certain factor is an enzyme, its biological activity includes its enzymatic activity. In another example, if a certain factor is a ligand, its biological activity includes the binding of the ligand to its corresponding receptor. Such biological activity can be measured by techniques well known in the art. Thus, "activity" refers to various measurable indicators that demonstrate or reveal binding (either directly or indirectly); affect a response (i.e., have a measurable effect in response to some exposure or stimulus), including, for example, a measure of the amount of an upstream or downstream protein or other similar function in a host cell.

[0086] As used herein, the terms "transformation," "transduction," and "transfection" are used interchangeably unless otherwise specified, and refer to the introduction of nucleic acid into a host cell (optionally via a virus or viral vector). Any method for introducing nucleic acid into a host cell can be used as the transformation method, and various well-known techniques, such as the use of competent cells, electroporation, particle gun (gene gun) methods, and calcium phosphate methods, can be used.

[0087] As used herein, a "purified" substance or biological factor (e.g., a nucleic acid or protein) refers to a biological factor from which at least a portion of the factors naturally associated with the biological factor have been removed. Thus, the purity of the biological factor in a purified biological factor is typically higher (i.e., more concentrated) than in the state in which the biological factor normally exists. As used herein, the term "purified" preferably means that at least 75% by weight, more preferably at least 85% by weight, even more preferably at least 95% by weight, and most preferably at least 98% by weight of the same type of biological factor is present. The substance used in the present disclosure is preferably a "purified" substance.

[0088] As used herein, the terms "drug," "agent," or "factor" (all of which correspond to the English term "agent") are used interchangeably in a broad sense and may refer to any substance or other element (e.g., energy such as light, radioactivity, heat, or electricity) that can achieve the intended purpose. Examples of such substances include, but are not limited to, proteins, polypeptides, oligopeptides, peptides, polynucleotides, oligonucleotides, nucleotides, nucleic acids (e.g., DNA such as cDNA and genomic DNA, and RNA such as mRNA), polysaccharides, oligosaccharides, lipids, small organic molecules (e.g., hormones, ligands, signaling substances, small organic molecules, molecules synthesized by combinatorial chemistry, small molecules that can be used as pharmaceuticals (e.g., small molecule ligands), etc.), composite molecules thereof, and mixtures thereof.

[0089] As used herein, the terms "complex" or "complex molecule" refer to any construct comprising two or more moieties. For example, if one moiety is a polypeptide, the other moiety may be a polypeptide or another substance (e.g., a substrate, a sugar, a lipid, a nucleic acid, another carbohydrate, etc.). As used herein, the two or more moieties constituting a complex may be bonded by a covalent bond or by other bonds (e.g., hydrogen bonds, ionic bonds, hydrophobic interactions, van der Waals forces, etc.).

[0090] The term "about" refers to the indicated value plus or minus 10% or significant figures if indicated. When "about" is used in reference to temperature, it refers to the indicated temperature plus or minus 5°C, and when "about" is used in reference to pH, it refers to the indicated pH plus or minus 0.5.

[0091] (Preferred Embodiments) Preferred embodiments of the present disclosure will be described below. The embodiments provided below are provided for a better understanding of the present disclosure, and it is understood that the scope of the present disclosure should not be limited to the following description. Therefore, it is clear that those skilled in the art can make appropriate modifications within the scope of the present disclosure in light of the description in this specification. It is also understood that the following embodiments can be used alone or in combination.

[0092] In one aspect, the present disclosure provides a method for producing a combinatorial library, comprising the steps of: 1) providing starting plasmid DNA(s); 2) introducing mutations into the starting plasmid DNA(s) to generate a plasmid DNA population (seed plasmid population) containing mutated plasmid DNAs (seed plasmids) into which mutations have been introduced; and 3) subjecting the plasmid DNA population to a combinatorial library generation procedure.

[0093] In one embodiment, the step 1) of providing starting plasmid DNA(s) can be carried out as follows. First, a method can be used in which a plasmid is mass-amplified using bacteria such as E. coli, followed by extraction and purification. This method is easy to scale up and cost-effective. Second, a method can be used in which the desired plasmid is produced using eukaryotic organisms such as yeast and then purified. Eukaryotic organisms such as yeast have the advantage of high stability of specific plasmid sequences. Furthermore, it is also possible to synthesize a plasmid in vitro using enzymes and provide it. This method is effective when creating a plasmid containing a special sequence that is toxic to a host such as E. coli during cloning, but does not pose a problem in the final host organism. These methods are appropriately selected depending on the desired properties and application of the plasmid. Alternatively, a plasmid can be synthesized chemically.

[0094] In another embodiment, the step of introducing mutations into starting plasmid DNA to generate a plasmid DNA population (seed plasmid population) containing mutated plasmid DNA (seed plasmid) into which mutations have been introduced can be carried out as follows. First, in the step of introducing mutations, methods such as saturation mutagenesis (a method of introducing all possible mutations into a specific gene site), random mutagenesis (a method of randomly introducing mutations into the entire DNA sequence (chemical, physical, or enzymatic methods are used)), error-prone PCR (a method of introducing random mutations by performing PCR under conditions that intentionally increase errors), LFEAP (Ligation of Fragments After PCR) (a method of performing PCR using a primer that creates an overhang after inverse PCR and then ligating the overhangs to generate a mutant plasmid), or gene cloning with degenerated fragments can be used. Oligonucleotides (a technique that uses partially randomized oligonucleotides to introduce mutations into specific regions of genes), recombinase-mediated gene recombination (a technique that uses recombinase to recombine specific sites in DNA to introduce mutations), phosphoramidite synthesis (a technique that uses chemically synthesized oligonucleotides to create libraries containing diverse sequences), diversification technologies (a technique that introduces diversity by applying various chemical modifications to specific DNA sequences), transposon mutagenesis (a technique that uses transposons to perform random insertions into genes), and coding loop cloning. (Cloning; a technique for introducing various amino acid mutations by randomizing specific codons) can be used. Plasmid DNA prepared using these methods is called a seed plasmid. Next, the step of generating a population of seed plasmids can be carried out as follows.Multiple seed plasmids selected according to any criteria can be mixed in equimolar amounts. Liquid solutions of plasmid-containing strains can also be mixed in equal volumes. Colonies of transformants can also be scraped from a transformation plate using a cell scraper and mixed. In this way, a population of seed plasmids can be generated.

[0095] The step of subjecting the seed plasmid DNA population to the combinatorial library construction procedure can be carried out as follows. In this step, various methods are used, including the Combi-OGAB (registered trademark) method, the Golden Gate method (a method in which fragments are assembled using the protruding ends that appear at both ends of DNA fragments using Type IIS restriction enzymes), the Gibson Assembly method (a method in which 3' overhangs are generated at both ends of DNA fragments using exonuclease, and fragments are assembled according to homologous sequences), the iPac method (an in vitro packaging-assisted DNA assembly method in which the ligation product of DNA fragments containing cos sequences is packaged into phage particles and a plasmid is constructed in E. coli using the cos sequences), the overlap-extension PCR method (a method in which homologous sequences in PCR primers and fragments are used to construct a hybrid gene), and DNA shuffling (a method in which DNA fragments are assembled using the cos sequences). Shuffling (a technique in which fragments of different mutants of the same gene are recombined to create new combinations) can be used. The resulting combinatorial library then forms a population of plasmid DNA with a large amount of variation, with each plasmid carrying a different genetic mutation. This method generates a diverse gene pool designed to efficiently search for plasmids with desired properties and functions. Furthermore, the resulting combinatorial library can be used in a process to select and amplify plasmids with desired properties by applying specific selection pressures.

[0096] In one embodiment, the seed plasmid collection used in the present disclosure may or may not include a starting plasmid. If the starting plasmid is included, it is advantageous that the starting plasmid is used directly to generate the combinatorial library. The starting plasmid may not be included.

[0097] In one embodiment, the combinatorial library construction procedure used in the present disclosure includes the Combi-OGAB method, the Golden Gate method, the Gibson Assembly method, or the overlap extension PCR and iPac method. These methods are described in Miyamoto, N., et al., ACS Omega, 2024, 9, 6, 6873-6879; Pullmann, P., et al., Sci. Rep., 2019, 9, 10932; Olszakier, S., et al., BMC Biotechnol., 2022, 13, 22(1), 10; Williams, E. M., et al. , Methods Mol. Biol. , 2014, 1179, 83-101. , Nozaki, S. , ACS Synth. Biol. , 2022, 11 (12), 4113-4122. It can be carried out based on the description of the present specification with reference to the above.

[0098] In one embodiment, the library of the present disclosure is capable of recombinatorialization. An advantage of such a recombinatorial library is that the resulting plasmids can be directly reused. Multiple common restriction enzyme SfiI recognition sequences are contained between the plasmid molecules. The SfiI recognition sequence is GGCCNNNNNGGCC, with the central three bases, NNN, forming a protruding end upon SfiI digestion. Therefore, by varying these three bases for each SfiI recognition sequence within a plasmid molecule, fragments with different protruding ends can be generated upon digestion with a single restriction enzyme, allowing the ligation order of each fragment to be specified. Because the SfiI recognition sequence, including this protruding end, is preserved throughout all steps from library construction to screening, recombinatorialization can be performed continuously during the process.

[0099] In one embodiment, the combinatorial library construction procedure used in the present disclosure includes Combi-OGAB (registered trademark) or the like. Combi-OGAB (registered trademark) is described in Miyamoto, N., et al., ACS Omega, 2024, 9, 6, 6873-6879, and is advantageously used in terms of highly efficient construction of libraries composed of long, multi-fragment sequences and rapid screening cycles due to the reusability of plasmids (the property of enabling recombinatorial construction).

[0100] In another embodiment, the method of the present disclosure further comprises a step of confirming that the plasmid DNA population generated in step 2) has a structure appropriate for the combinatorial library generation procedure prior to step 3). Procedures for confirming that the plasmid DNA population has a structure appropriate for the combinatorial library generation procedure are known in the art and may include checking the fragment pattern generated by restriction enzymes or DNA sequencing.

[0101] In one embodiment, the method of the present disclosure may include repeating step 3) or steps 2) and 3) at least once. This is because the step of repeating the library synthesis cycle can include a procedure of performing PCR again to create a library. Another advantage of repeating the cycle is that it allows multiple combinations of mutations to be tested and the mutations to be optimized to improve desired properties.

[0102] In one embodiment, the method of the present disclosure may include a step of selecting a plasmid DNA having desired properties. The method for selecting a plasmid DNA having desired properties can be any method known in the art, for example, selecting a DNA molecule that exhibits strong binding activity to a protein or nucleic acid to be detected in a DNA sensor.

[0103] In one embodiment, in the method of the present disclosure, selection of plasmid DNA is performed after step 2), after step 3), or in both steps 2) and 3). Selecting plasmid DNA in step 2) is significant in that it allows for the selection of plasmid DNA that does not lose desired properties after mutagenesis.

[0104] The significance of selecting plasmid DNA in step 3) is that it is possible to select plasmid DNA in which desired properties have been enhanced by a combination of introduced mutations. The significance of selecting plasmid DNA in steps 2) and 3) is that it is possible to select plasmid DNA in which desired properties have been enhanced by a combination of introduced mutations.

[0105] In one embodiment, the method of the present disclosure can be said to be a method for producing a plasmid DNA that confers a desired property. Therefore, the method of producing a plasmid that confers a desired property of the present disclosure includes the steps of: 1) providing a starting plasmid DNA(s); 2) introducing mutations into the starting plasmid DNA to generate a plasmid DNA population (seed plasmid population) containing mutated plasmid DNA (seed plasmid) into which mutations have been introduced; and 3) subjecting the plasmid DNA population to a combinatorial library construction procedure. It can be said that strain construction has already been performed to confirm the desired property. The method of the present disclosure can also include a method for producing a strain carrying a plasmid (i.e., a desired strain). Such a method can include transforming a production host organism with the plasmid DNA (for example, but not limited to, integrating it into the genome).

[0106] In one embodiment, the method further includes a step of optimizing the promoter. Plasmid optimization can be performed as follows, but is not limited to: For example, the promoter optimization step is performed to efficiently control the expression of a target gene and achieve maximum expression under specific conditions. This step requires the selection or design of an optimal promoter, taking into account various techniques and conditions. First, promoters of different strengths are evaluated to regulate the gene expression level, and a promoter with the optimal strength is selected depending on the experimental objectives. Furthermore, by using a temperature-sensitive promoter, gene expression can be designed to change in a specific temperature environment. This method makes it possible to control expression depending on temperature. In addition, the timing of gene expression can be adjusted by introducing an inducible promoter whose expression is induced in response to chemicals or specific environmental conditions. Inducible promoters initiate expression in response to external stimuli, making them effective when expression is desired only under specific conditions.

[0107] In one embodiment, the method of the present disclosure further includes a step of optimizing a homologous gene. Such optimization can be performed using any process known in the art. Examples include, but are not limited to, the following. The step of optimizing a homologous gene involves comparing and adjusting genes encoding proteins with the same function that are conserved across different biological species to achieve the most efficient expression and function for a specific purpose. First, the amino acid sequences of homologous genes obtained from different biological species are compared to identify highly homologous regions. This comparison establishes a foundation for optimization while preserving functionally important regions. Next, codon usage appropriate for the target biological species and experimental conditions is considered. Each biological species has a codon usage frequency that improves gene translation efficiency. Therefore, optimizing the codons of a homologous gene can improve protein translation efficiency. Furthermore, by adjusting the exon-intron arrangement and splicing pattern, a homologous gene can be optimized for accurate and efficient expression. Furthermore, factors that control gene expression, such as promoters and enhancers, can also be taken into consideration, and the homologous gene can be designed to express at a timing and intensity appropriate for the purpose. In this process, it is important to properly position the binding site for the transcription factor, which improves expression efficiency. Furthermore, in optimizing a homologous gene, adjustments can be made without introducing mutations into the amino acid sequence so that the protein structure and function are maintained. Finally, after homologous gene optimization is complete, test expression and functional analysis are performed to confirm that the gene functions correctly in the target environment. Thus, homologous gene optimization is a multi-step process, with each step requiring ingenuity to improve gene function and expression.

[0108] In one embodiment, the method of the present disclosure may further include a step of optimizing the base sequence in the structural gene. Various methods for optimizing such base sequences are contemplated, including any method known in the art. For example, but not limited to, optimizing the base sequence in the structural gene is a method for maximizing gene expression efficiency and ensuring accurate production of a target protein. In the first stage of optimization, the base sequence of the target gene is modified taking into account codon usage in the target organism. Because certain codons are translated more efficiently in certain biological species, adjusting the sequence based on codon usage improves protein production efficiency. Next, the GC content is adjusted. Because GC content affects gene stability and translation efficiency, optimizing the sequence to a GC content appropriate for the target biological species contributes to improved gene expression. Furthermore, ribosome binding site stability and avoidance of secondary structures are also important factors, and sequence modifications are performed to suppress secondary structure formation so that the base sequence is properly recognized by the ribosome. Furthermore, to improve translation speed and accuracy, the sequence is adjusted to avoid rare codons and enable smooth translation. This process requires designing the gene so that protein folding is accurate and the translation rate is appropriately controlled at each stage of protein synthesis. Next, signal peptides and targeting sequences may also be optimized, and the base sequence is adjusted to ensure appropriate intracellular localization and secretion of the protein. This optimization enables the protein encoded by the structural gene to be delivered into the cell in a functional form. Finally, experimental expression tests are conducted to verify whether the optimized base sequence is effective in an actual expression system. Through this series of steps, the base sequence of the structural gene is optimized according to the purpose, enabling efficient and accurate protein expression.

[0109] In one embodiment, optimization in the present disclosure is performed to improve the function or stability of a gene or gene product. Optimization for the purpose of improving the function or stability of a gene or gene product is performed so that the gene product's function (transcription efficiency into mRNA, translation efficiency from mRNA, and function of the translated protein) is not impaired and the gene sequence on the plasmid DNA is not deleted, enabling the gene to exist in a stable state. It is also performed to enable the target protein to function more efficiently and exist in a stable state. For example, codon usage can be adjusted to improve translation efficiency and increase protein production. Amino acid substitution can also be performed to stabilize the protein structure and increase its resistance to heat and pH fluctuations. Furthermore, various techniques that lead to improved stability can be employed, such as modifying the protein's degradation signal to suppress degradation.

[0110] In one embodiment, the optimization step of the present disclosure may be performed before step 1) or after step 2).

[0111] In one embodiment, the method of the present disclosure further includes a step of preparing the starting plasmid DNA into a structure suitable for the combinatorial library construction procedure. The step of preparing the starting plasmid DNA into a structure suitable for the combinatorial library construction procedure involves optimizing the plasmid DNA structure to achieve library diversity and efficient gene expression. In this step, first, the necessary gene regions and functional elements (e.g., promoter, origin, selection marker) are identified and positioned. Next, the DNA is cleaved and recombined using enzymes such as restriction enzymes and recombinases to ensure that these elements are combined in a manner suitable for the combinatorial library construction. Furthermore, the insertion of cloning sites and reporter genes is designed to enable effective screening of mutant genes in the library. Furthermore, the vector size is adjusted to ensure that it is not too large and maintains high replication efficiency. Finally, the prepared plasmid DNA is evaluated through trial cloning and expression tests to determine whether it is suitable for the desired gene expression and functional analysis in the combinatorial library construction. In this way, a plasmid structure optimal for combinatorial library construction is prepared.

[0112] In one embodiment, the method of the present disclosure further includes a step of preparing a structure suitable for combinatorial library construction. For example, but not limited to, the step of preparing a structure suitable for combinatorial library construction is an important step for optimizing the starting plasmid DNA to maximize the diversity and gene expression efficiency in the library. In this step, necessary genetic elements (e.g., promoters, replication origins, selection markers, etc.) in the plasmid are first identified, and the structure of these elements is adjusted so that they are located in a position suitable for combinatorial library construction. Furthermore, the structure of the plasmid DNA is optimized by designing restriction enzyme recognition sequences and adding cloning sites to enable efficient insertion and recombination of diverse gene sequences. Next, the plasmid size is optimized, and the overall size of the plasmid is adjusted to avoid a decrease in replication efficiency. This process includes deleting unnecessary sequences and introducing additional functional sequences. Furthermore, after a structure suitable for library construction is prepared, the plasmid is cut and ligated into the desired structure using recombinases or restriction enzymes, enabling diverse combinations in the library. Finally, experimental gene expression and mutation screening are performed to confirm that the prepared plasmid DNA functions properly in the library construction procedure. In this way, the optimal plasmid structure for combinatorial library construction is established, enabling efficient library construction.

[0113] In one embodiment, the step of generating a seed plasmid population further includes a step of preparing a structure suitable for the combinatorial library construction procedure. This preparation step can be performed by any method known in the art. For example, but not limited to, the step of preparing a structure suitable for the combinatorial library construction procedure is a step of modifying the structure of the starting plasmid DNA to optimize the diversity and gene expression efficiency within the library. First, basic genetic elements contained in the plasmid DNA, such as promoters, replication origins, and selection markers, are designed to be arranged in a manner suitable for a combinatorial library. During this process, restriction enzyme recognition sequences and cloning sites are added or rearranged as necessary to ensure efficient function of each element. Next, the structure of the plasmid DNA is adjusted to ensure diversity of gene sequences in the library construction. Specifically, unnecessary sequences are deleted, and the overall size of the plasmid is adjusted to increase replication efficiency. Furthermore, the plasmid structure is modified using recombinases or restriction enzymes to enable random recombination and insertion of gene sequences. Finally, test gene expression and screening are performed to confirm whether the prepared plasmid DNA functions appropriately in the library construction procedure. In this way, a plasmid structure suitable for the combinatorial library creation procedure is established, enabling efficient library generation. Specifically, for example, if the starting plasmid does not have an SfiI recognition sequence, an SfiI recognition sequence can be inserted by PCR to create a plasmid suitable for the Combi-OGAB (registered trademark) method. In this case, two approaches are possible: first introducing an SfiI site and then introducing a mutation; or introducing a mutation while inserting the SfiI site.

[0114] In another aspect, the present disclosure provides a method for producing a desired host organism strain, comprising transforming a host organism with a combinatorial library produced by the method of the present disclosure or a plasmid contained therein. Any method and embodiment known in the art can be used for such production. Specifically, but not limited to, a method for producing a desired host organism strain using a plasmid involves introducing a plasmid containing a gene of interest into a host organism, designed to efficiently express the gene. First, the gene of interest is cloned into a plasmid along with an appropriate promoter and selection marker. This plasmid is then introduced into a host such as Escherichia coli, yeast, insect cells using baculovirus as a host, HEK293 cells (human embryonic kidney 293 cells), CHO cells (Chinese hamster ovary cells), or plant cells. Next, the plasmid is inserted into the host organism using electroporation or chemical transformation. The plasmid is replicated within the host organism, and the gene of interest is expressed, resulting in the desired protein or function. Examples include the production of insulin using Escherichia coli and the production of bioethanol-producing enzymes using yeast. Other examples include the production of adeno-associated virus (AAV) using HEK293 cells, the production of antibody drugs using CHO cells, the production of vaccines using plant cells, and the generation of virus-like particles (VLPs) using baculovirus-based insect cells.

[0115] In one embodiment, transformation includes integration into the genome. For example, but not limited to, transformation is the process of introducing foreign genetic material into a host cell, which includes integrating the introduced genetic material into the genome. The integration of the introduced plasmid or DNA fragment into the genome of the host cell enables stable gene expression. In transformation, foreign genetic material is inserted into the cell through the cell membrane using electroporation, chemical methods, or the like, and is expected to be integrated into the genome by the cell's homologous recombination mechanism, or the like. As a result, the transformed cell acquires the ability to permanently express the function of the foreign gene.

[0116] In one aspect, the present disclosure provides a method for improving a desired property by random mutation, comprising the steps of: 1) inducing mutations in a plasmid DNA carrying a gene cluster containing at least one gene, 2) constructing a combinatorial library based on various plasmid DNAs containing random mutations, and 3) selecting plasmid DNAs that express the desired property.

[0117] In one embodiment, the step of inducing mutations in plasmid DNA carrying a gene cluster containing at least one or more genes can be performed in any manner based on the description herein, with reference to techniques known in the art. For example, but not limited to, the step of inducing mutations in plasmid DNA carrying a gene cluster containing at least one or more genes is a procedure performed to modify the function or expression of a target gene cluster. In this step, plasmid DNA containing a gene cluster is first prepared, and the mutagenesis method disclosed herein, as well as chemical mutagens, physical methods (e.g., ultraviolet irradiation, radiation), or site-specific mutagenesis techniques using enzymes (e.g., the CRISPR / Cas9 system) can be used. Random or targeted mutations are then introduced into the gene cluster within the plasmid. The induced mutations affect the intensity of gene expression and the function and structure of proteins, potentially resulting in the production of genes or gene products with novel properties. The mutated plasmid DNA is then introduced into host cells by transformation, and the effects of the mutations are evaluated. This process enables the discovery and improvement of genes and proteins with novel functions.

[0118] In one embodiment, the process of constructing a combinatorial library based on diverse plasmid DNAs containing random mutations can be carried out in any manner based on the description herein, with reference to techniques known in the art. For example, but not limited to, the process of constructing a combinatorial library based on diverse plasmid DNAs containing random mutations is a process for maximizing genetic diversity and discovering gene sequences with novel functions or properties. Using plasmids into which mutations have been introduced using the above-described method, combinatorial libraries can be created using the Combi-OGAB (registered trademark) method, Golden Gate method, Gibson Assembly method, overlap extension PCR method, iPac method, or the like. These methods generate diverse plasmids containing combinations of mutations for each library unit fragment. Next, these mutant plasmids are transformed into host cells to produce cell populations each containing different mutations. The cell populations obtained in this process contain a variety of genetic mutations and therefore function as a combinatorial library. The cell library is then used to screen for plasmids that express desired properties, such as cells that survive under specific environmental conditions or cells that produce specific proteins. This process creates a combinatorial library with extensive genetic diversity, which can be used to discover genes and proteins with desired properties.

[0119] In one embodiment, the step of selecting plasmid DNA expressing a desired property can be performed as desired based on the description herein, with reference to techniques known in the art. For example, but not limited to, the step of selecting plasmid DNA expressing a desired property is a step for efficiently identifying a plasmid expressing a gene or characteristic of interest. In this step, plasmid DNA containing a gene of interest or a marker gene is first introduced into host cells by transformation. Next, a plasmid containing a selectable marker, such as an antibiotic resistance gene or a fluorescent protein gene, is used for selection, and the cells are cultured under conditions in which only cells that have correctly incorporated the plasmid survive or express the marker. For example, cells introduced with a plasmid containing an antibiotic resistance gene are cultured in a medium containing an antibiotic, and only cells resistant to the antibiotic survive. Furthermore, when a plasmid expressing a fluorescent protein is used, fluorescent cells can be selected using a fluorescence microscope or the like. In this way, cells harboring a plasmid expressing the desired property are selected and used for subsequent research or production processes.

[0120] In one embodiment, in the method of the present disclosure, the desired property includes the ability to be recombinatorialized. Specifically, the plasmid after recombinatorialization includes one that can be used again for recombinatorialization without further manipulation, such as re-preparation of fragments by PCR. This means that the recombinatorial process using genetic material (e.g., plasmid DNA, gene clusters, etc.) can be performed multiple times. Specifically, this refers to a process of increasing genetic diversity by introducing new combinations or mutations based on existing combinatorial libraries or gene sequences. This "recombinatorialization" allows further mutations and recombinations to be performed within the same library, resulting in the creation of more diverse gene sequences and combinations with novel gene functions. For example, by recombinatorializing the candidate sequences obtained in the initial library, it is possible to verify new mutation combinations and select genes with more desirable properties or functions.

[0121] In one embodiment of the present disclosure, the mutations employed are introduced into the entire plasmid DNA, which refers to all base sequences on the plasmid DNA, including not only the gene / gene cluster of interest and its corresponding transcription factors (promoter, enhancer, terminator, etc.), but also the vector.

[0122] In one embodiment, in the present disclosure, mutations are induced by PCR or chemical mutagens. Such techniques are known in the art and can be performed by those skilled in the art.

[0123] In one embodiment, the present disclosure provides a method for introducing mutations into specific regions of plasmid DNA. In this process, mutations are concentrated in specific regions of the plasmid DNA, aiming to modify a gene sequence with a desired function or property. For example, by introducing mutations into regions important for gene expression or function, such as the active site or regulatory sequence of a specific protein, it is possible to improve the function or confer novel properties. The introduction of mutations is accomplished using techniques such as error-prone PCR or site-directed mutagenesis, which alters specific base sequences within the target region.

[0124] In one embodiment, in the present disclosure, mutation induction is achieved by genome editing, base editing, or prime editing. Genome editing, without limitation, uses technologies such as CRISPR / Cas9 to cut DNA at specific locations, and mutations are introduced during the repair process. Base editing is a technique that directly converts specific bases without cutting DNA, thereby enabling highly accurate mutation introduction. Prime editing is a technique that uses guide RNA and modified reverse transcriptase to change a specific DNA sequence to an arbitrary base sequence, allowing for complex mutation introduction. These techniques allow precise mutations to be introduced into specific regions within plasmid DNA, generating genes or sequences with desired functions.

[0125] In one embodiment, in the present disclosure, the diverse plasmid DNA containing mutations is prepared by the OGAB® method. Preferably, the method for constructing a combinatorial library is the Combi-OGAB® method.

[0126] In another aspect, the present disclosure provides a method for producing a substance in a host organism. This method comprises the steps of: (A) inducing mutations in plasmid DNA; (B) creating a combinatorial library of the mutations using the Combi-OGAB® method; (C) arbitrarily selecting and culturing host organisms harboring the combinatorial library of plasmid DNAs and selecting multiple strains based on arbitrary criteria; and (D) repeating any of steps (A), (A) to (C), or (B) to (C) in any order. Here, it is understood that A can be arbitrarily incorporated as well as A(BC)(BC)(BC) ..., A(BC)(ABC)(ABC) ..., and the like.

[0127] In one embodiment, the step of inducing mutations in plasmid DNA can be performed in any manner based on the description herein, with reference to techniques known in the art. For example, but not limited to, the step of inducing mutations in plasmid DNA is a procedure performed to modify the function or characteristics of a target gene. In this step, mutations are introduced into specific regions of the plasmid DNA using techniques such as genome editing, base editing, and prime editing. In genome editing, plasmid DNA is cut at specific sites using techniques such as CRISPR / Cas9, and random or targeted mutations are introduced during the repair process. In base editing, precise mutations can be introduced by directly converting specific bases without cutting the plasmid DNA. Prime editing is used to achieve more complex base sequence changes and introduce specific mutations into the target sequence. These methods effectively induce mutations that confer new properties and functions to plasmid DNA.

[0128] In one embodiment, the step of creating a combinatorial library of the mutations using the Combi-OGAB® method can be performed in any manner based on the description herein, with reference to techniques known in the art. For example, but not limited to, the step of creating a combinatorial library of the mutations using the Combi-OGAB® method is a method for efficiently combining DNA fragments with multiple mutations to construct a combinatorial library for exploring novel gene functions and characteristics. First, DNA fragments containing mutations introduced into a gene of interest or plasmid DNA are prepared. Next, these DNA fragments are ligated with high efficiency using the Combi-OGAB® method to construct a large-scale gene library. In this method, DNA fragments with multiple mutations are combined using specific restriction enzyme sites or recombinases to generate plasmid DNA with random combinations. The generated combinatorial library contains diverse gene sequences and is introduced into a host organism for gene expression and functional analysis. This library is useful for comprehensively evaluating the various functions of genes into which mutations have been introduced and for searching for genes and proteins with novel functions and properties.

[0129] In one embodiment, the process of arbitrarily selecting and culturing host organisms carrying plasmid DNA in a combinatorial library and selecting multiple strains based on arbitrary criteria can be performed in any manner based on the description herein, with reference to techniques known in the art. For example, but not limited to, the process of arbitrarily selecting and culturing host organisms carrying plasmid DNA in a combinatorial library and selecting multiple strains based on arbitrary criteria is performed to efficiently obtain strains with desired properties and functions. In this process, first, host organisms (e.g., bacteria or yeast) into which the plasmid DNA in a combinatorial library has been introduced are cultured, and conditions are established to select strains expressing the desired gene or function. Selection is performed based on arbitrary criteria, such as specific antibiotic resistance, fluorescent protein expression, or production of a specific metabolite. For example, strains carrying plasmids containing antibiotic resistance genes are selected by culturing them in a medium containing an antibiotic. Furthermore, strains expressing fluorescent proteins can be visually confirmed and selected using a fluorescence microscope. Other criteria include the ability to survive in a specific environment, or the analysis of the metabolites produced by the strains. In this way, multiple strains with specific properties based on arbitrary criteria can be selected and further cultivated for the desired application.

[0130] In one embodiment, the production host organism of the present disclosure is Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacteria, corynebacterium, animal cells, or plants. In this process, the production host organism to which the constructed combinatorial library is applied may be, but is not limited to, Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacteria, corynebacterium, animal cells, or plants. First, DNA fragments containing target mutations are compiled into a library using the Combi-OGAB® method, and various combinations of these fragments are then introduced into a host organism. Each host organism is suitable for promoting the expression of a target gene or protein under specific conditions and improving production efficiency and characteristics. For example, Bacillus subtilis and E. coli are used for industrial protein production, actinomycetes and filamentous fungi for the production of antibiotics and secondary metabolites, and yeast for fermentation processes. Furthermore, hydrogen bacteria and corynebacterium are used for the production of specific compounds, and animal cells and plants for the production of complex proteins and pharmaceuticals. In this way, the combinatorial library generated by the Combi-OGAB® method is evaluated in various host organisms, and the optimal strains and cell lines are selected for each production application.

[0131] In one embodiment, the gene cluster carried by the plasmid DNA is related to substance production. In the step of creating a combinatorial library of mutations using the Combi-OGAB® method, the gene cluster carried by the plasmid DNA may be related to substance production. This gene cluster functions to produce a specific compound, protein, or enzyme, and mutations can be introduced to improve their expression or function. Such gene clusters often contain enzyme genes responsible for the production of secondary metabolites or genes controlling specific metabolic pathways. In this step, random or specific mutations are introduced into the plasmid DNA carrying the gene cluster involved in substance production, and a combinatorial library is constructed based on the mutations. The library plasmids are then introduced into an appropriate host organism (e.g., Bacillus subtilis, actinomycetes, yeast, Escherichia coli, etc.) to optimize the production efficiency and characteristics of the target substance. The selected bacterial strains or cell lines may be used for industrial-scale substance production.

[0132] In one embodiment, the gene cluster carried by the plasmid DNA in the present disclosure may be specifically a nonribosomal peptide synthetase (NRPS) or a polyketide synthase (PKS). It is advantageous for the gene cluster carried by the plasmid DNA to be related to a nonribosomal peptide synthetase (NRPS) or a polyketide synthase (PKS). NRPSs and PKSs are involved in the synthesis of complex compounds produced as secondary metabolites by microorganisms, and these enzymes play an essential role in the production of substances such as antibiotics, immunosuppressants, and antitumor agents. In this process, random or specific mutations are introduced into the gene cluster encoding the NRPS or PKS, and a combinatorial library is constructed based on the mutations. This library aims to improve the function and properties of the NRPS or PKS and to produce novel compounds or increase the production efficiency of existing compounds. The plasmid DNA library is then introduced into an appropriate host organism, such as an actinomycete, a filamentous fungus, Bacillus subtilis, or Escherichia coli, and the production of the desired secondary metabolite is evaluated. This will enable optimization of substance production by NRPS or PKS and the search for new substances.

[0133] The plasmid DNA of the present disclosure may be an adeno-associated virus (AAV) production plasmid. In this case, in this process, plasmid DNA carrying AAV components such as capsid proteins, replicons, and gene clusters for constructing the viral genome is used for the purpose of efficient AAV production. This enables mass production of AAV vectors, which are applicable to medical applications such as gene therapy and vaccine development. Using the Combi-OGAB® method, mutations are introduced into AAV production plasmids to create combinatorial libraries, enabling more efficient AAV vector production and the design of novel AAV capsids with specific functions and properties. The library-formed plasmid DNA is introduced into appropriate host cells (e.g., HEK293 cells), and the production efficiency and function of AAV vectors are evaluated. This process optimizes the AAV production process and promotes applications such as gene therapy.

[0134] (Construct Creation) A nucleic acid having a sequence that replicates within a bacterium can be introduced into a host cell to form a starting plasmid in the host cell. The OGAB® method can be suitably used to form the starting plasmid. In one embodiment, a strain such as Bacillus subtilis can have the ability to form a starting plasmid from an exogenously introduced nucleic acid, so in this method, the introduced nucleic acid does not have to be a plasmid. For example, when Bacillus subtilis comes into contact with a nucleic acid (e.g., a non-circular nucleic acid having a tandem repeat nucleic acid sequence), it can take up the nucleic acid and form a plasmid within Bacillus subtilis (see, for example, the OGAB® method described herein). For construct creation, see, for example, International Publication No. WO 2022 / 097646.

[0135] (Preparation of Mutation-Introduced Plasmid) A method for preparing a mutation-introduced plasmid is described below. The following is an example, and other methods can also be used. Using the starting plasmid as a template, primers are designed to divide the plasmid into lengths amplifiable by the DNA polymerase used, and DNA fragments are prepared by PCR. Type IIS restriction enzyme recognition sequences, such as AarI, BsaI, BbsI, and BsmBI, are added to both ends of the fragments. After fragment preparation, the fragments are treated with a restriction enzyme appropriate to the added recognition sequences, thereby creating protruding ends that seamlessly determine the joining order. The PCR reaction can be performed according to the procedure provided by the supplier, or an optimal concentration of manganese ions can be added to increase the mutation rate. The resulting PCR fragments are mixed in equimolar amounts, treated with the restriction enzyme, and then subjected to tandem repeat ligation. The ligation product can then be used to construct a mutation-introduced plasmid (seed plasmid) using the OGAB® method. A plurality of such seed plasmids are selected based on an arbitrary criterion, and an equimolar mixture is used as a seed plasmid population for creating a combinatorial library.

[0136] (Combinatorial library construction) Combinatorial library construction will now be described. The following is an example, and other methods may also be used. A construct is constructed to which a promoter insertion site and a sequence for reconstruction (a structure suitable for combinatorial construction) are added. As described above, the OGAB (registered trademark) method can be suitably used to construct a seed plasmid. The sequences for reconstruction may be SfiI, AarI, AlwNI, BbsI, BbvI, BcoDI, BfuAI, BglI, BsaI, BsaXI, BslI, BsmAI, BsmBI, BsmFI, BspMI, BspQI, BstAPI, BstXI, BtgZI, DraIII, EarI, Esp3I, FokI, HgaI, MwoI, PflMI, SfaNI, or SapI recognition sequences, and similar sequences may also be used, with N residues generated upon cleavage. A seed plasmid for combinatorial library creation is constructed by introducing various promoters upstream of each gene in the gene cluster. Multiple promoters with different transcription initiation times and / or transcription strengths are selected. After creating a combinatorial library, each construct in the library may be transduced into a host cell, one or more strains having desired properties may be selected, and then another combinatorial library may be created based on the selected strains. By repeating the cycle of creating a combinatorial library, transduction, and selecting strains multiple times, strains with better properties and / or multiple properties may be obtained (fine tuning). For example, by fine tuning the enzymes described in D. B. Basnet, et al., J. Biotechnol., 2008, 135, 92-96 (EryCIII enzyme reaction, EryK / EryG enzyme reaction); Appl. Microbiol. Biotechnol., 2017, 101, 5951, the production yield (content in a fraction) of a desired substance may be improved. The cycle may be performed more than once, for example, two, three, four or five times.

[0137] The selection of strains to be used in the second cycle of combinatorial library creation may be performed based on any criteria. For example, the top five strains for desired properties (e.g., high production of a target substance) may be selected, or strains that satisfy a predetermined property (e.g., production of a target substance exceeding that of a natural strain) may be selected. The strains selected for the second cycle may be 2 or more, 50 or more, 100 or more, 1,000 or more, or 10,000 or more. Preferably, 2 or more strains are selected for the second cycle. The number of strains selected for the second cycle is the maximum number of transformants that can be obtained, and may be, for example, 10, 15, 20, 30, 40, 50, or 100. The minimum and maximum number of strains selected may be any number between the above values.

[0138] (Preparation of integrating nucleic acids) Below, the procedure for preparing integrating nucleic acids to be incorporated into transformed organisms used in the present disclosure will be described.

[0139] Preparation of Nucleic Acid Units: Nucleic acid unit molecules to be incorporated into the nucleic acid assembly are prepared. The nucleic acid unit molecules can be produced by any known method, such as polymerase chain reaction (PCR) or chemical synthesis. The nucleic acid unit molecules can have any desired sequence, such as a sequence or part thereof encoding a desired protein (e.g., metabolic enzyme, natural product biosynthetic enzyme group, artificial enzyme, therapeutic protein, protein constituting a viral vector), a gene-regulating sequence (e.g., promoter, enhancer), or a sequence for manipulating nucleic acids (e.g., restriction enzyme recognition sequence). The ends of each nucleic acid unit molecule can be configured to produce specific overhanging sequences, so that multiple types of nucleic acid unit molecules are arranged in a specific order and / or specific orientation when incorporated into the nucleic acid assembly.

[0140] Since a large number of nucleic acid unit molecules can ultimately be assembled on a plasmid, one or more nucleic acid unit molecules may be designed to encode one or more long genes, such as a group of genes constituting a series of metabolic pathways or a protein complex composed of multiple subunits.

[0141] - Preparation of unit vectors Unit vectors can be prepared by linking unit nucleic acids with additional nucleic acids that are different from the unit nucleic acids. Use of unit vectors may make it easier to handle unit nucleic acids.

[0142] The additional nucleic acid may be a linear nucleic acid or a circular plasmid. When a circular plasmid is used as the additional nucleic acid, the unit vector may also have a circular structure, and therefore may be used, for example, for transformation of E. coli or the like. In one embodiment, the additional nucleic acid may contain an origin of replication so that the unit vector is replicated in the host into which it is introduced. In one embodiment, all of the unit nucleic acids for constructing a certain integrated nucleic acid may be linked to the same type of additional nucleic acid, thereby reducing the size difference between the unit vectors and making it easier to handle multiple types of unit vectors. In one embodiment, the unit nucleic acids for constructing a certain integrated nucleic acid may be linked to different types of additional nucleic acids. In one embodiment, for one or more unit nucleic acids for constructing a certain integrated nucleic acid, the ratio or average of (base length of unit nucleic acid) / (base length of unit vector) may be 50% or less, 40% or less, 30% or less, 20% or less, 15% or less, 10% or less, 7% or less, 5% or less, 2% or less, 1.5% or less, 1% or less, or 0.5% or less. The larger the size of the additional nucleic acid is than the unit nucleic acid, the more uniform the handling of different types of unit vectors may be. The unit nucleic acid and the additional nucleic acid can be linked by any method, for example, ligation using DNA ligase, TA cloning, etc. In one embodiment, for one or more unit nucleic acids used to construct a certain assembled nucleic acid, the ratio or average of (base length of unit nucleic acid) / (base length of unit vector) may be 1% or more, 0.3% or more, 0.1% or more, 0.03% or more, 0.01% or more, 0.003% or more, or 0.001% or more, which may facilitate the manipulation of the unit vector.

[0143] In one embodiment, the length of the unit nucleic acid may be 10 bp or more, 20 bp or more, 50 bp or more, 70 bp or more, 100 bp or more, 200 bp or more, 500 bp or more, 700 bp or more, 1000 bp or more, or 1500 bp or more, and 5000 bp or less, 5000 bp or less, 2000 bp or less, 1500 bp or less, 1200 bp or less, 1000 bp or less, 700 bp or less, or 500 bp or less.

[0144] In one embodiment, the nucleic acid assembly can be constructed from 2 or more, 4 or more, 6 or more, 8 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more types of nucleic acid units, and 1,000 or less, 700 or less, 500 or less, 200 or less, 120 or less, 100 or less, 80 or less, 70 or less, 60 or less, or 50 or less types of nucleic acid units. By adjusting the molar numbers of each nucleic acid unit (or vector unit) to be approximately the same, a desired nucleic acid assembly having a tandem repeat structure can be efficiently produced.

[0145] In one embodiment, the nucleic acid unit may have a base length that divides one set of repeat sequences in the nucleic acid assembly approximately equally by the number of nucleic acid unit sequences. This can facilitate the operation of adjusting the molar number of each nucleic acid unit (or unit vector). In one embodiment, the nucleic acid unit may have a base length that is increased or decreased by 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, 7% or less, or 5% or less from the base length that divides one set of repeat sequences in the nucleic acid assembly approximately equally by the number of nucleic acid unit sequences.

[0146] In one embodiment, the nucleic acid unit molecules can be designed to have non-palindromic sequences (sequences that are not palindromic sequences) at the ends of the nucleic acid unit molecules. When the non-palindromic sequences of the nucleic acid unit molecules designed in this way are made into protruding sequences, they can easily give rise to a structure in which the nucleic acid unit molecules are linked to each other while maintaining their order in the assembled nucleic acid.

[0147] Preparation of Assembly Nucleic Acids Assembly nucleic acids can be constructed by linking unit nucleic acids to each other. In one embodiment, unit nucleic acids can be prepared by excising them from unit vectors using a restriction enzyme or the like. Assembly nucleic acids can contain one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more sets of repeat sequences. The repeat sequences in the assembly nucleic acid can include the sequences of the unit nucleic acids and, if necessary, the sequences of the assembly vector nucleic acids. Assembly nucleic acids can have a sequence that enables replication of the nucleic acid in a transformed organism. In one embodiment, the sequence that enables replication of the nucleic acid in a transformed organism can include an origin of replication that is effective in the transformed organism (e.g., Bacillus bacteria (Bacillus subtilis)). The sequence of a replication origin effective in Bacillus subtilis is not particularly limited, and examples of those having a θ-type replication mechanism include sequences of replication origins contained in plasmids such as pTB19 (Imanaka, T., et al. J. Gen. Microbioi. 130, 1399-1408. (1984)), pLS32 (Tanaka, T and Ogra, M. FEBS Lett. 422, 243-246. (1998)), and pAMβ1 (Swinfield, T.J., et al. Gene 87, 79-90. (1990)).

[0148] The nucleic acid assembly may contain additional base sequences as necessary in addition to the nucleic acid units. In one embodiment, the nucleic acid assembly may contain base sequences that control transcription / translation, such as a promoter, operator, activator, or terminator. Specific examples of promoters when Bacillus subtilis is used as a host include Pspac (Yansura, D. and Henner, D.J. Pro. Natl. Acad. Sci. USA 81, 439-443 (1984)), whose expression can be controlled by IPTG (isopropyl s-D-thiogalactopyranoside), and Pr promoter (Itaya, M. Biosci. Biotechnol. Biochem. 63, 602-604 (1999)).

[0149] The nucleic acid unit molecules can form a repeat structure in the assembled nucleic acid, maintaining a specific order and orientation. In one embodiment, by constructing the nucleic acid unit molecules such that the base sequences of the protruding ends of the nucleic acid unit molecules excised from the unit vectors are complementary to each other, it may be possible to form a repeat structure in the assembled nucleic acid, maintaining a specific order and orientation. In one embodiment, by making the structure of the protruding end unique for each different nucleic acid unit molecule, it is possible to efficiently form a repeat structure in a specific order and orientation. In one embodiment, the protruding end may have a non-batch sequence and may be either a 5'-end protruding or a 3'-end protruding.

[0150] In one embodiment, a nucleic acid unit having a cohesive end can be excised from a unit vector using a restriction enzyme. In this embodiment, the unit vector may have one or more restriction enzyme recognition sequences. When a unit vector has multiple restriction enzyme recognition sequences, the respective restriction enzyme recognition sequences may be recognized by the same restriction enzyme or by different restriction enzymes. In one embodiment, the unit vector may contain a pair of regions recognized by the same restriction enzyme, such that the complete nucleic acid unit region is contained between these regions. In an embodiment in which a restriction enzyme is used to cleave a recognition region, the unit vector may contain a region recognized by that restriction enzyme at the end of the nucleic acid unit region.

[0151] In one embodiment, when the same type of restriction enzyme is used to excise the nucleic acid units from multiple unit vectors, restriction enzyme treatment can be performed in a solution containing these multiple unit vectors, thereby improving work efficiency. The types of restriction enzymes used to prepare a certain nucleic acid assembly can be, for example, five or fewer, four or fewer, three or fewer, two or fewer, or one type, and using a small number of restriction enzymes can reduce variation in the number of moles between the nucleic acid units. In one embodiment, the nucleic acid units excised from the unit vectors can be easily purified by any known fractionation method, such as agarose gel electrophoresis.

[0152] The nucleic acid units and, if necessary, the integrated vector nucleic acid can be ligated to each other using DNA ligase or the like. This allows the production of an integrated nucleic acid. For example, the ligation of the nucleic acid units and, if necessary, the integrated vector nucleic acid can be carried out in the presence of components such as polyethylene glycol (e.g., PEG2000, PEG4000, PEG6000, PEG8000, etc.) and salt (e.g., monovalent alkali metal, sodium chloride, etc.). The concentration of each nucleic acid unit in the ligation reaction solution is not particularly limited and may be 1 fmol / μl or more. The reaction temperature and time for the ligation are not particularly limited and may be 37°C or more for 30 minutes or more. In one embodiment, before ligating the nucleic acid units and, if necessary, the integrated vector nucleic acid, a composition containing the nucleic acid units and, if necessary, the integrated vector nucleic acid may be subjected to any conditions that inactivate the restriction enzyme (e.g., phenol-chloroform treatment).

[0153] The nucleic acid unit molecules can be adjusted to approximately the same number of moles using, for example, the method described in WO 2015 / 111248. By adjusting the nucleic acid unit molecules to approximately the same number of moles, a desired assembled nucleic acid having a tandem repeat structure can be efficiently produced. The number of moles of the nucleic acid unit molecules can be adjusted by measuring the concentration of the unit vector or nucleic acid unit.

[0154] (Preparation of Plasmid from Accumulated Nucleic Acid) Plasmids can be formed in transformed organisms by contacting the accumulated nucleic acid with the transformed organism. In one embodiment, transformed organisms include bacteria of the genus Bacillus, Streptococcus, Haemophilus, and Neisseria. Examples of Bacillus bacteria include B. subtilis, B. megaterium, and B. stearothermophilus, but any transformed organism capable of forming a plasmid from the accumulated nucleic acid can be used. In one embodiment, the transformed organism into which the accumulated nucleic acid is introduced is competent and can actively take up nucleic acid. For example, Bacillus subtilis in a competent state cleaves the double-stranded nucleic acid substrate on the cell, degrades one of the two strands from the cleavage point, and incorporates the other single strand into the cell. The incorporated single strand can be repaired to a circular double-stranded nucleic acid within the cell. Any known method can be used to make a transformed organism competent, and for example, Bacillus subtilis can be made competent using the method described in Anagnostopoulou, C. and Spizizen, J. J. Bacteriol., 81, 741-46 (1961). As the transformation method, a known method suitable for each transformed organism can be used.

[0155] In one embodiment, the present disclosure provides a vector plasmid produced from packaging cells. In one embodiment, the vector plasmid produced from packaging cells can be purified using any known method, and the present disclosure also provides a vector plasmid purified in this manner. In one embodiment, the presence of the desired nucleic acid sequence in the purified vector plasmid can be confirmed by examining the size pattern of fragments generated by restriction enzyme cleavage, PCR, nucleotide sequencing, or the like. In one embodiment, a vector plasmid-containing composition prepared by the vector plasmid production method of the present disclosure may contain a small amount of endotoxin. In one embodiment, a Bacillus subtilis containing the vector plasmid of the present disclosure is provided.

[0156] In the present disclosure, the transformed organism (host cell) used may be selected from Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacterium, coryne bacteria, animal cells, or plant cells. A preferred transformed organism is Bacillus subtilis. These transformed organisms are preferably capable of producing plasmids from the accumulated nucleic acids. The transformed organisms used to generate plasmids from the accumulated nucleic acids and create a combinatorial library may be the same or different from the transformed organisms used to produce the substance. For example, Bacillus subtilis may be used as the transformed organism for creating a combinatorial library, and the same Bacillus subtilis may be used as the transformed organism for producing the substance, or another transformed organism, such as actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacterium, coryne bacteria, animal cells, or plants, may be used.

[0157] In one embodiment, the gene cluster associated with substance production may be a gene cluster for biosynthetic enzymes of polyketide synthase (PKS), non-ribosomal peptide synthase (NRPS), ribosomal translation system post-translationally modified peptides (RiPPs), or terpene biosynthetic enzymes.

[0158] Promoter combinations in constructing strains having desired properties in the present disclosure include PabrB, PvalS, Phag, PspoVG, PyvyD, Phema, Pffh, PsdhB, PiolS, PspoOA, PglyA, PyceC, PsrfAA, PyrxA, PytxG, PytvI, PspoVB, PphrK, PphrF, PftsH, Pasd, Pspo0F, PmenB, PgltX, Pmi nC, PphrC, PspoVS, PlicH, Pcdd, PflgM, PcheV, PcitZ, PclpQ, PphrG, PlytR, Phbs, PodhA, PsecA, PphrE, PftsZ, PxpaC, PcsbA, PywjC, PyknW, PysdB, PrapI, PyteJ, PureA, PtrxA, Pdps, PxynA, PrsbW, PspoIVCA, PyqfY, PyqeZ, PyabR, PkinA, PnadE, PsigH, PresE, PspsA, PgspA, PyflG, Pveg, Pgap, PahpC, PyvrE, PytkL, PyfkJ, PdegS, PpyrAB, PyacL, PywkA, Psp oVT, PywtG, PpdhC, PyuzA, PyhdF, PalsD, PpurE, PythP, PcomK, PptsI, PgtaB, PkatX, PydbP, PyqgZ, PydjO, PlytE, Pgsi Examples of promoters that can be used include any combination selected from the group consisting of Bacillus subtilis, PywrE, PopuE, PyoxA, PthrS, PbltD, PyfhL, PyraD, PydaD, PywnJ, PrelA, PydfK, PgerBC, PyitG, PybfO, PcotH, PylnF, PribT, PrapH, PmmgA, PyqfD, and PsigW, but are not limited to these, and other promoters that are growth phase-dependent of Bacillus subtilis can also be used as appropriate.

[0159] In this specification, "or" is used when "at least one or more" of the items listed in the sentence can be employed. The same applies to "alternative." In this specification, when it is specified that "within a range" of "two values," the range includes the two values ​​themselves.

[0160] All references cited herein, including scientific literature, patents, patent applications, and the like, are incorporated by reference in their entirety to the same extent as if each were specifically set forth.

[0161] The present disclosure has been described above by showing preferred embodiments for ease of understanding. The present disclosure will be described below based on examples. However, the above description and the following examples are provided for illustrative purposes only and are not intended to limit the present disclosure. Therefore, the scope of the present disclosure is not limited to the embodiments or examples specifically described herein, but is limited only by the scope of the claims.

[0162] Mutation introduction methods that can be used in the present disclosure include, but are not limited to, PCR, which is a method that can efficiently introduce mutations into a plasmid. Other methods (such as mutagenic chemicals and genome editing) can also be used.

[0163] The reasons why PCR is preferable include, but are not limited to, the following: Long-chain plasmids such as those exemplified in Example 1 exceed the limit of PCR amplification (generally several kb) at one time. Therefore, by using the technology disclosed herein to separate the plasmid into fragments, perform PCR, and then assemble the fragments using the OGAB (registered trademark) method, a full-length plasmid can be constructed. In particular, fragmentation is essential for the "introduction of a new SfiI recognition sequence" described below, which is difficult to achieve using methods other than PCR. Furthermore, the mutation rate can be varied by adding an appropriate concentration of Mn ions to the PCR reaction solution. For example, the mutation rate can be increased by increasing the Mn ion concentration (Mol. Biotechnol., 2016, 58(8-9), 551-557). Furthermore, unmutated fragments can be obtained by treating the plasmid with a restriction enzyme and purifying it through gel excision. The OGAB (registered trademark) method and Combi-OGAB (registered trademark) method can assemble a larger number of fragments and can assemble longer DNA chain lengths than the existing Golden Gate method and Gibson Assembly method, and therefore can construct plasmids even from long DNA chains such as natural product biosynthetic enzyme gene clusters and viral vector constructs.

[0164] If an SfiI recognition sequence is not present in the starting plasmid or if you want to add one, you can design primers that include the SfiI recognition sequence and then introduce / add the SfiI recognition sequence by PCR. In this case, you can first introduce the SfiI recognition sequence and then introduce the mutation, or you can introduce the SfiI recognition sequence and introduce the mutation simultaneously.

[0165] In this disclosure, a seed plasmid for Combi-OGAB® can be prepared, which can be said to be a plasmid into which random (unplanned) mutations have been introduced. PCR is used to create this seed plasmid, but this differs from the two purposes of PCR: (1) obtaining a DNA fragment with 100% correct sequence, and (2) adding a DNA barcode.

[0166] A DNA fragment containing no mutation can be excised from a plasmid, and such a fragment can also be prepared by PCR.

[0167] Examples of library construction include the Golden Gate method and the Gibson Assembly method, which are other DNA fragment accumulation methods that can be used in the present disclosure.

[0168] The Golden Gate method (Methods Mol. Biol., 2013, 1073, 141-156.; Golden Mutagenesis: An efficient multi-site-saturation mutagenesis approach by Golden Gate cloning with automated primer design. (Sci. Rep., 2019, 9, 10932.)) will be described.

[0169] In this method, by making the protruding ends (f1 to fn+1) that appear after restriction enzyme treatment identical, sequences at the same position (Cn-x) can be created into a combinatorial library. However, since the restriction enzyme recognition sequences disappear upon accumulation, it is necessary to add the restriction enzyme recognition sequences again by PCR to construct a secondary library (plasmids cannot be reused).

[0170] This paper describes a method for constructing a combinatorial library by using the Golden Gate method to enrich for mutation-containing fragments prepared by PCR. Conventional methods that directly enrich PCR-generated fragments (e.g., "Direct in Expression Vector" and "Direct") are considered difficult to achieve with PCR products containing multiple mutations. Using the method disclosed herein, transformation can be achieved in one day regardless of the number, length, or sequence complexity of fragments (plasmid extraction → SfiI treatment → tandem repeat ligation → transformation of Bacillus subtilis).

[0171] Gibson Assembly is a genetic engineering technique that joins multiple DNA fragments together. Compared to classical techniques using restriction enzymes and DNA ligase, this method requires fewer manipulations, can be completed in a shorter time, and has the advantage of not leaving any extraneous sequences from the restriction sites (Gibson et al. (2009). "Enzymatic assembly of DNA molecules up to several hundred kilobases" (PDF). Nature Methods 6(5):343-345. doi:10.1038 / nmeth.1318).

[0172] To determine the order of ligation of adjacent fragments, linear double-stranded DNA with homologous sequences at both ends is prepared. The DNA is treated with T5-exonuclease to convert both ends into single strands, which are then annealed according to the homologous sequences. DNA polymerase then synthesizes complementary strands to fill the gap, and ligase fills the nick, completing the ligation. As with the Golden Gate method, the plasmid cannot be reused for secondary library construction, and a step for preparing fragments for Gibson Assembly is essential (see BMC Biotechnol., 2022, 13, 22(1), 10.). A method has been shown in which mutation-containing fragments prepared by PCR are pooled using the Gibson Assembly method to construct a library. This method also directly pools fragments prepared by PCR, without subcloning. The entire process is specified as taking 4 days, which is thought to be longer than using Combi-OGAB®.

[0173] The Golden Gate method and the Gibson Assembly method are both capable of accumulating PCR fragments, as demonstrated in the two examples above, and are applicable to the present disclosure. However, the accumulated plasmids cannot be directly reused for the next library construction, and constructing a secondary library for proceeding with the screening cycle requires re-preparation of the accumulation fragments by PCR. Therefore, the Combi-OGAB® method, which allows this, is preferred. For example, the plasmid extracted from the strain obtained in the primary screening can be reused directly for the next library construction, allowing for a rapid screening cycle (e.g., one day; plasmid extraction → SfiI treatment → tandem repeat ligation → transformation of Bacillus subtilis). Due to this property, Combi-OGAB® can shorten each cycle by one to three days when screening PCR fragments containing mutations.

[0174] (General Method of Use) The mRNA provided by the plasmid and production method of the present disclosure is used as an active ingredient in mRNA pharmaceuticals, particularly pharmaceuticals such as infectious disease vaccines, cancer vaccines, and replacement protein therapy.

[0175] The present disclosure can be used, for example, in the production of mRNA drugs. Using the method of the present disclosure, a designed and constructed mRNA drug production plasmid is used for transformation. The resulting transformed strain is used as an industrial production strain, and seed culture and production culture are performed in a culture size appropriate for the purpose. The cells are then collected and lysed, and the plasmid is recovered and purified. After linearization by enzymatic treatment, RNA is produced by a method such as in vitro transcription using T7 RNA polymerase. This is purified, subjected to the necessary specification analysis, and can be used as a drug substance for mRNA drugs.

[0176] The present disclosure has been described above by showing preferred embodiments for ease of understanding. The present disclosure will be described below based on examples. However, the above description and the following examples are provided for illustrative purposes only and are not intended to limit the present disclosure. Therefore, the scope of the present disclosure is not limited to the embodiments or examples specifically described herein, but is limited only by the scope of the claims.

[0177] The present disclosure will be specifically described in the following examples, but the present disclosure is not limited to these examples. The reagents used were specifically products described in the examples, but equivalent products from other manufacturers (Sigma-Aldrich, Wako Pure Chemical Industries, Nakarai, R&D Systems, USCN Life Science INC, etc.) can also be used.

[0178] Reagents and Test Methods: The reagents and test methods used in the examples are as follows. The Bacillus subtilis host used was the RM125 strain (Uozumi, T., et al., Molec. Gen. Genet., 152, 65-69 (1977)), and the BUSY9797 and Marburg168 strains, which are derivatives of the RM125 strain. The plasmid vector pGETS151ΔBsmBI (Non-Patent Document 1) was used as a replicable vector in Bacillus subtilis. The antibiotics carbenicillin, tetracycline, and kanamycin were manufactured by Nacalai Tesque. Lysozyme was purchased from Fujifilm Wako Pure Chemical Industries, Ltd. Restriction enzymes were manufactured by New England Biolabs. T4 DNA ligase was purchased from Takara Bio Inc. For PCR to introduce mutations into the starting plasmid, Ex-Taq HS manufactured by Takara Bio Inc. was used. When error-prone PCR was performed to increase the mutation rate, MnSO was added to the PCR reaction solution. 4 An aqueous solution of 2-hydroxyethyl agarose (manufactured by Nacalai Tesque) was mixed to the desired concentration. For gel excision and purification, 2-hydroxyethyl agarose (manufactured by Sigma), a low-melting-point agarose gel for DNA electrophoresis, was used. Another common agarose used for electrophoresis was UltraPure Agarose (manufactured by Invitrogen). TE saturated phenol (containing 8-xylinol) was manufactured by Nacalai Tesque. LB medium components and agar were manufactured by Becton Dickinson. Reagents used for extraction from the bacterial cells or culture medium after cultivation and for HPLC analysis were manufactured by Nacalai Tesque.

[0179] DNA was purified from the PCR reaction solution using Qiagen's MinElute Reaction Cleanup Kit. Thermo's Nano Drop One ultraspectrophotometer was used. The base sequence was determined using Illumina's Miseq and Applied Biosystems' 3500xL Genetic Analyzer, a fluorescent automated sequencer.

[0180] Other general DNA manipulations were performed according to standard protocols (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1989)). Transformation of Bacillus subtilis using the OGAB® method and plasmid extraction, as well as the creation and screening of combinatorial libraries using the Combi-OGAB® method, were performed as previously reported (Non-Patent Document 1, Tsuge, K., et al., Sci. Rep., 2015, 5, 10655, Patent Document 1, Japanese Patent No. 6440636).

[0181] Example 1: Investigation into Improvement of Nonribosomal Peptide Production Ability (1) Preparation of Mutation-Introduced Plasmid for Gramicidin S Production In Non-Patent Document 1, monocistronic promoter screening was performed to construct a heterologous Gramicidin S (GS)-producing strain in Bacillus subtilis. In the present invention, the plasmids harbored by the two strains obtained during this screening, 3rd-C2, which is the highest GS-producing strain, and 2nd-1E9, which has high culture reproducibility, were used as starting plasmids (SEQ ID NOs: 1 and 2). In this example, we aimed to construct a strain exhibiting GS production ability exceeding that of 3rd-C2.

[0182] Primers were designed to separate the starting plasmid (22,138 bp) into four fragments for PCR (SEQ ID NOS: 3-10). The primer sets for each fragment were designed to contain the restriction enzyme AarI recognition sequence at the end of the fragment. The target fragments were ligated together via the protruding ends generated by digesting the PCR product with AarI (Figure 1). PCR reactions for each fragment were performed using Ex Taq HS (Takara Bio) according to the standard protocol provided with the kit. The resulting fragments were digested with AarI and ligated. The ligated products were then used for enrichment in Bacillus subtilis BUSY9797, which had been transformed with the plasmid pUB8 (Tsuge, K., et al., Arch. Microbiol., 1996, 165, 243-251), encoding the phosphopantethenylate transferase lpa-8. The above steps were performed separately for 3rd-C2 and 2nd-1E9. Twelve colonies from each strain were selected, for a total of 24 strains, and the plasmid sequences were confirmed to have introduced mutations at an average rate of approximately 0.77 substitutions / kb. The GS productivity of the 24 strains was then measured, and strains were confirmed to exhibit productivity comparable to or even superior to that of the 3rd-C2 strain (Figure 2). The four plasmids from these strains were mixed to create a DNA population of mutated plasmids. This DNA population was then used to create a combinatorial library.

[0183] (2) Screening of high-GS-producing strains using Combi-OGAB®: The DNA population was treated with SfiI, the enzyme was inactivated with TE-saturated phenol, the aqueous phase was concentrated with 1-butanol, and then ethanol precipitated. The resulting precipitate was dissolved in TE. This SfiI-treated fragment solution was religated with T4 DNA ligase and enriched in the BUSY9797 strain into which pUB8 had been introduced. The fragments were then plated on LB-10 μg / mL tetracycline-kanamycin-containing plates and incubated overnight at 30°C. The resulting colonies were inoculated into 300 μL of LB-10 μg / mL tetracycline-kanamycin-containing liquid medium, cultured overnight at 30°C, and then frozen and stored at -70°C. For culturing to evaluate GS productivity, a frozen stock was inoculated into 2 mL of YTG (Bacto Yeast Extract 50 g / L, Bacto Tryptone 50 g / L, glucose 5 g / L)-10 μg / mL tetracycline-kanamycin-containing liquid medium and cultured at 30°C for 72 hours with shaking. 2 mL of ethyl acetate was added to the culture, and the mixture was stirred for 30 seconds using a vortex mixer. The mixture was centrifuged at 10,000 × g for 10 minutes, and the organic layer was collected. The organic layer was evaporated to dryness and dissolved in 200 μL of 70% methanol-0.05% formic acid. This solution was centrifuged at 15,000 × g for 5 minutes, and 20 μL of the supernatant was injected into an HPLC to analyze GS productivity. The HPLC measurement conditions were as follows: Column: COSMOSIL® 5C18-AR-II Packed Column 4.6 mm ID x 150 mm Mobile phase A: 0.1% formic acid - purified water, Mobile phase B: methanol Gradient: (60% B, 5 minutes) - (60% → 78% B, 9 minutes) - (78% → 100% B, 0.1 minutes) - (100% B, 5 minutes) Flow rate: 1 mL / min Detection: 210 nm After analysis, several high-producing strains were selected from the obtained data, and after re-culture, equal amounts of the cultures were mixed and plasmids were extracted. Thereafter, the library was reconstructed by SfiI treatment, and the productivity of each strain was similarly measured using the Combi-OGAB® cycle until the production amount converged. During the culture for evaluating GS productivity, 3rd-C2 was also cultured in each cycle, and GS was extracted and analyzed by HPLC.The GS productivity of this 3rd-C2 was used as an internal standard for relative comparison.

[0184] (3) Changes in GS production per cycle In the first cycle (PCR-1st_1st cycle), the GS production ability of 30 randomly selected strains was analyzed, and equal amounts of plasmids extracted from the top 10 strains were mixed together before proceeding to the second cycle. In the second cycle (PCR-1st_2nd cycle), the GS production ability of 20 strains was analyzed, and equal amounts of plasmids extracted from the top 5 strains were mixed together before proceeding to the third cycle. In the third cycle (PCR-1st_3rd cycle), the GS production ability of 19 strains was analyzed. Since the average GS production ability in the second and third cycles was almost flat, it was considered that the production amount had converged in the second cycle, and screening was terminated (Table 1, Figure 2).

[0185]

[0186] Among the strains analyzed, the second cycle B2 strain (P2nd-B2) showed the highest productivity, with GS productivity 1.50 times higher than that of 3rd-C2.

[0187] (4) Verification of the effectiveness of combining mutagenesis and Combi-OGAB®. The mixture of plasmids extracted from the top 10 strains in the first cycle, used in the second cycle of library preparation, was used as the starting plasmid and mutagenesis was performed by PCR. The GS productivity of 30 randomly selected strains was then analyzed, and no strains were found to have a GS productivity exceeding that of the P2nd-B2 strain (Figure 3). These results suggest that screening the combination of introduced mutations with Combi-OGAB® may be more efficient than attempting to improve productivity by mutagenesis using PCR alone.

[0188] (5) Further Mutation Introduction and Investigation of Productivity Improvement Using Combi-OGAB (Registered Trademark) The aim was to create a strain with a GS production level exceeding that of the P2nd-B2 strain by introducing further mutations. A mixture of the extracted plasmids from the top 10 strains in the second cycle was used as the starting plasmid, and a mutated plasmid was similarly prepared by PCR and enriched in Bacillus subtilis (PCR-2nd). Twelve strains were selected, and their GS productivity was analyzed. Then, equal amounts of the plasmids extracted from the top four strains were mixed to form a DNA population, which was then transferred to the first cycle (PCR-2nd_1st cycle). The GS productivity of 24 strains was analyzed, and the top 10 strains were transferred to the second cycle (PCR-2nd_2nd cycle). As a result, GS production ability converged already in the first cycle, and the PP1st-A8 strain, which showed the highest GS production ability, showed GS production ability 1.52 times higher than the production amount of 3rd-C2, and showed higher GS production ability than the P2nd-B2 strain (Table 2, Figure 4).

[0189]

[0190] (6) Confirmation of converged mutations in strains P2nd-B2 and PP1st-A8 The sequences of the plasmids of the two strains obtained this time (P2nd-B2, PP1st-A8) were analyzed using Miseq (Illumina) to confirm the mutations. The number of mutations introduced for each promoter, gene, and vector is shown in Table 3.

[0191]

[0192] It is impossible to predict whether these mutations will contribute to improved productivity. The process of the present invention does not require the design of the position or sequence of the mutation. As in this example, it was shown that it is possible to randomly introduce mutations into the entire 22 kb long plasmid and verify the combination of mutations in each gene region and vector, thereby achieving an improvement in productivity for the nonribosomal peptide Gramicidin S.

[0193] (7) Mutation Analysis of Plasmids Carried by the P2nd-B2 Strain Finally, the mutations introduced into the P2nd-B2 strain plasmid obtained this time were repaired for each SfiI fragment (grsA: 3 mutations, grsB: 16 mutations, vector: 4 mutations) to create constructs, and strains carrying each plasmid were prepared. Next, GS production ability was analyzed for each strain (N=3) (Figure 5). Strains carrying plasmids from which all mutations had been repaired (mutation-free plasmid carriers) exhibited lower GS production ability compared to the P2nd-B2 strain. Strains carrying a single mutation fragment (grsA mutation fragment carrier, grsB mutation fragment carrier, vector mutation fragment carrier) also exhibited lower GS production ability compared to the P2nd-B2 strain, with no significant difference observed between them and the mutation-free strains. Furthermore, the two-fragment mutation repair mutants (grsA + grsB mutation fragment carriers, grsA + vector mutation fragment carriers, and grsB + vector mutation fragment carriers) exhibited higher GS productivity than strains with a single mutation fragment, but their average GS productivity was lower than that of the P2nd-B2 strain. This suggests that the improvement in GS productivity in the P2nd-B2 strain was not achieved by mutations accumulated in a single fragment, but rather that the mutations contained in each fragment contributed cooperatively to the improvement in GS productivity. Therefore, by screening for combinations of mutations using the Combi-OGAB (registered trademark) method, improved GS productivity was achieved, and the P2nd-B2 strain was successfully obtained.

[0194] Example 2: Investigation into Improvement of Polyketide Productivity (1) Preparation of Mutation-Introducing Plasmid for Triketide Pyrone Production The Bacillus subtilis genome contains a dormant Triketide Pyrone biosynthetic gene cluster (composed of two genes, bpsA and bpsB), a polyketide biosynthetic enzyme gene cluster. To construct a Bacillus subtilis strain capable of activating and producing this cluster, a combinatorial library containing 11 Bacillus subtilis growth-phase-dependent promoters (PabrB, PvalS, Phag, PspoVG, PsrfAA, Pcdd, PlytR, Pveg, PmmgA, PyqfD, and PsigW) arranged monocistronically was constructed using the Combi-OGAB® method and screened. As a result, the 3rd-A1 strain exhibited high productivity. This strain was screened for the PvalS promoter for bpsA and the PsrfAA promoter for bpsB. In this example, the plasmid contained in the 3rd-A1 strain was used as the starting plasmid (SEQ ID NO: 11).

[0195] The entire starting plasmid (5,915 bp) was subjected to PCR using Ex Taq HS with primers of SEQ ID NOs: 3 and 10 from Example 1. Both ends of the PCR product contain the restriction enzyme AarI recognition sequence, and tandem repeat ligation products can be generated using the protruding ends that appear when the PCR product is treated with AarI. PCR reactions were performed according to the standard protocol for Ex Taq HS (Takara Bio). The resulting fragment was treated with AarI, ligated, and enriched in Bacillus subtilis Marburg168. Twelve strains were selected, and the plasmid sequences were confirmed, revealing an average of approximately 0.72 substitutions / kb of mutations. The productivity of methylated triketide pyrone (Pyrone-OMe) with a molecular weight of 350 for each strain was measured, and equal amounts of the top four strains and the template 3rd-A1 strain, totaling five plasmids, were mixed to create a DNA population of mutated plasmids. This DNA population was then used to create a combinatorial library.

[0196] (2) Screening of Pyrone-OMe high-producing strains using Combi-OGAB (registered trademark) The DNA population was treated with SfiI, the enzyme was inactivated with TE-saturated phenol, the aqueous phase was concentrated with 1-butanol, and then precipitated with ethanol. The precipitate was dissolved in TE. This SfiI-treated fragment solution was religated with T4 DNA ligase and enriched in Marburg168 strain. The resulting fragments were plated on LB-10 μg / mL tetracycline-containing plates and incubated overnight at 30°C. The resulting colonies were inoculated into 300 μL of LB-10 μg / mL tetracycline-containing liquid medium, cultured overnight at 30°C, and then frozen and stored at -70°C. To evaluate the Pyrone-OMe productivity of each strain, a frozen stock was inoculated into 10 mL of LB-10 μg / mL tetracycline-containing liquid medium and cultured at 37°C for 24 hours with shaking. The culture was centrifuged at 8,000 × g for 5 minutes, and the precipitate was collected. The precipitate was then suspended in 15 mL of purified water and centrifuged again. The precipitate was suspended in 5 mL of 150 mM NaCl and 20 mM HCl and incubated at 80°C for 30 minutes. After that, 5 mL of ethyl acetate was added and the mixture was stirred for 30 seconds using a vortex mixer. The mixture was then centrifuged at 10,000 × g for 10 minutes, and the organic layer was collected. The collected organic layer was evaporated to dryness and dissolved in 100 μL of methanol. This solution was centrifuged at 15,000 × g for 5 minutes, and 20 μL of the supernatant was injected into an HPLC to analyze the Pyrone-OMe productivity. The HPLC measurement conditions were as follows. Column: COSMOSIL (registered trademark) 5C18-MS-II Packed Column 4.6 mm ID x 150 mm. Mobile phase A: 0.1% formic acid - purified water. Mobile phase B: methanol. Gradient: (80% B, 3 minutes) - (80% → 98% B, 4 minutes) - (98% B, 18 minutes). Flow rate: 1 mL / min. Detection: 280 nm. During the cultivation to evaluate Pyrone-OMe productivity, 3rd-A1 was also cultivated, and Pyrone-OMe extraction and HPLC analysis were performed. The productivity of this 3rd-A1 was used as an internal standard for relative comparison.

[0197] (3) Validation of Combi-OGAB® for Improving Pyrone-OMe Productivity The Pyrone-OMe productivity of 12 randomly selected strains was analyzed. As a result, strain E4 (P1st-E4) exhibited 1.36-fold higher Pyrone-OMe productivity than strain 3rd-A1, the highest-producing strain obtained by promoter optimization (Figure 6).

[0198] From the above, it was suggested that the process of the present invention can be applied to nonribosomal peptide synthetases (NRPSs) and polyketide synthases (PKSs), and is useful for improving the production yield of target substances.

[0199] Example 3 Analysis of Accumulation Efficiency and Mutation Rate of PCR Products by Bacillus subtilis in the Presence of Mn Ions (1) Error-Prone PCR of Gene Cluster Region The biosynthetic gene cluster region containing the promoter in the plasmid of 3rd-A1, the template strain of Example 2, was amplified by PCR using MnSO in the absence of Mn ions. 4 In the presence of 250 μM, MnSO 4 PCR was performed under three conditions in the presence of 500 μM of ribosomal DNA. The vector fragment was prepared by treating pGETS151ΔBsmBI with SfiI, followed by gel excision and purification.

[0200] (2) Measurement of transformants and calculation of mutation rate The PCR products obtained under the three conditions above were treated with SfiI, ligated with the pGETS151 vector, and enriched in Bacillus subtilis RM125. As a result, it was confirmed that the number of transformants tended to decrease with increasing Mn ion concentration. Five strains were then selected from each line, and the plasmid sequences were analyzed. It was confirmed that increasing the Mn ion concentration increased the mutation rate ( FIG. 7 ).

[0201] Example 4 Additional Introduction of SfiI Recognition Sequences into Starting Plasmids (1) Introduction of New SfiI Recognition Sequences and Mutations into pGETS151-GS 3rd-C2 To run the screening cycle using the Combi-OGAB (registered trademark) method, the restriction enzyme SfiI recognition sequences are effectively used as in Examples 1 and 2. These SfiI sequences were already present in the starting plasmids. In this example, it was verified that a library could be created using the SfiI sequences newly introduced into the starting plasmids.

[0202] The pGETS151-GS 3rd-C2 plasmid, which was the starting plasmid of Example 1, was used. Primers (SEQ ID NOs: 12 to 17) capable of newly introducing two SfiI recognition sequences into the grsB region were designed, and PCR was performed using these primers to simultaneously add a new SfiI sequence and introduce a mutation. PCR was performed under conditions without Mn addition or with MnSO 4 Three fragments (grsB-1, grsB-2, grsB-3) were prepared under conditions of 150 μM addition. Next, the SfiI fragment of the grsT region, the SfiI fragment of the grsA region, and the SfiI fragment of the vector region were gel excised and purified from the plasmid pGETS151-GS 3rd-C2. The excised fragments were mixed with the three PCR fragments, ligated, and then enriched in Bacillus subtilis strain BUSY9797 (Figure 8). Two strains were obtained from the PCR product enrichment system without Mn addition, and one strain was enriched in MnSO. 4 Four strains were selected from the PCR product accumulation system under the 150 μM addition condition, and the extracted plasmids from each strain were mixed in equal amounts to prepare a DNA population of mutated plasmids.

[0203] (2) Library Preparation with New SfiI Recognition Sequences Using Combi-OGAB (Registered Trademark) The DNA population was digested with SfiI, ligated, and then enriched in Bacillus subtilis strain BUSY9797. Ten strains were selected, and the plasmid sequences of the six strains comprising the DNA population, for a total of 16 strains, were confirmed. Mutations were confirmed by analyzing a plasmid (SEQ ID NO: 18) containing two new SfiI recognition sequences added to pGETS151-GS 3rd-C2 as a reference sequence. As a result, it was confirmed that the grsB-1, grsB-2, and grsB-3 fragments of the selected 10 strains were organized into a combinatorial library of grsB-1, grsB-2, and grsB-3 fragments of the six plasmids comprising the DNA population (Table 4).

[0204]

[0205] From the above, it was demonstrated that a combinatorial library can be prepared using a new SfiI recognition sequence introduced into the starting plasmid by PCR.

[0206] Example 5: Improvement of adeno-associated virus production All-in-one plasmid for producing adeno-associated virus (AAV) as disclosed in WO 2023 / 214578 A1 TM Using the starting plasmid, the Cap region is subjected to PCR in the presence of Mn, and the Helper region, Rep region, and Bacillus subtilis vector pGETS118 region are subjected to PCR in the absence of Mn. The rAAV region is prepared by treating the starting plasmid with SfiI, followed by gel excision and purification. These fragments are mixed, ligated, and enriched using Bacillus subtilis to prepare a mutated plasmid. Equal amounts of these mutated plasmids are selected and mixed to prepare a DNA population, and a combinatorial library is prepared using Combi-OGAB (registered trademark) as in Example 1. Plasmids are extracted from the resulting strains and introduced into HEK293 cells. Virus particles produced by each cell clone are recovered, and the AAV titer and empty capsid rate are measured. Multiple plasmids introduced into clones with high AAV titers are selected and subjected to a screening cycle as in Example 1, allowing screening for mutations in the plasmids involved in improving AAV production capacity.

[0207] (Note) While the present disclosure has been illustrated using preferred embodiments thereof, it is understood that the scope of the present disclosure should be construed solely by the claims. It is understood that the patents, patent applications (including PCT / JP2024 / 020950), and literature cited in this specification are incorporated by reference into this specification in their entirety as if the contents themselves were specifically set forth herein. This application claims priority to Japanese Patent Application No. 2024-148029, filed on August 29, 2024, with the Japan Patent Office, the entire contents of which are incorporated by reference herein.

[0208] The present invention provides a method for producing a microbial strain exhibiting desired properties by mutating a desired gene cluster, such as a gene cluster constituting a natural product biosynthetic enzyme gene cluster, and potentially novel substances, which can be used in bioindustries such as pharmaceuticals, agriculture, fisheries, fragrances, food, textiles, chemicals, and the environment.

[0209]

Claims

1. A method for producing a combinatorial library, comprising: 1) providing one or more starting plasmid DNAs; 2) introducing mutations into the starting plasmid DNAs to generate a plasmid DNA population (also referred to as a "seed plasmid population") containing mutated plasmid DNAs (also referred to as "seed plasmids") into which the mutations have been introduced; and 3) subjecting the plasmid DNA population to a combinatorial library creation procedure.

2. The method of claim 1, further comprising, prior to step 3), a step of confirming that the plasmid DNA population produced in step 2) has a structure appropriate for the combinatorial library construction procedure.

3. The method according to claim 1 or 2, comprising repeating step 3), or step 2) and step 3), at least once.

4. The method according to any one of claims 1 to 3, which comprises a step of selecting a plasmid DNA having desired properties.

5. The method according to any one of claims 1 to 4, wherein the selection of the plasmid DNA is carried out after step 2), after step 3), or after both steps 2) and 3).

6. The method according to any one of claims 1 to 5, wherein the method is a method for producing a plasmid DNA that confers a desired property.

7. The method according to any one of claims 1 to 6, further comprising the step of optimizing the promoter.

8. The method according to any one of claims 1 to 7, further comprising the step of optimizing the homologous gene.

9. The method according to any one of claims 1 to 8, further comprising the step of optimizing the base sequence in the structural gene.

10. The method according to any one of claims 1 to 9, further comprising the step of preparing the starting plasmid DNA in a structure suitable for the combinatorial library construction procedure.

11. The method of any one of claims 1 to 10, further comprising the step of preparing a structure suitable for the combinatorial library generation procedure.

12. The method of any one of claims 1 to 11, wherein the step of generating a seed plasmid population further comprises the step of preparing constructs suitable for the combinatorial library generation procedure.

13. A method for producing a desired host organism strain, comprising the step of transforming a host organism with the combinatorial library produced by the method of any one of claims 1 to 12.

14. A method for improving a desired property by mutation, comprising: 1) inducing mutations in a plasmid DNA carrying a gene cluster containing at least one gene; 2) constructing a combinatorial library based on the plasmid DNA containing the mutations; and 3) selecting a plasmid DNA that expresses the desired property.

15. The method of any one of claims 4 to 14, wherein the desired property comprises the ability to recombinatorially combine.

16. The method of any one of claims 1 to 15, wherein the plasmid DNA containing the mutation is prepared by the OGAB (registered trademark) method.

17. The method according to any one of claims 1 to 16, wherein the step of constructing the combinatorial library is the Combi-OGAB (registered trademark) method.

18. A method for producing a substance in a host organism, the method comprising: (A) a step of inducing mutations in plasmid DNA; (B) a step of creating a combinatorial library of the mutations by the Combi-OGAB (registered trademark) method; (C) a step of arbitrarily selecting and culturing host organisms carrying the plasmid DNA created in the combinatorial library, and selecting multiple strains based on arbitrary criteria; and (D) a step of repeating any of (A), (A) to (C), or (B) to (C) in any order.

19. The method according to claims 13 to 18, wherein the host organism is Bacillus subtilis, actinomycetes, filamentous fungi, yeast, Escherichia coli, hydrogen bacteria, Corynebacterium, animal cells or plants.

20. The method according to claim 18 or 19, wherein the gene cluster carried by the plasmid DNA is related to substance production.

21. The method according to any one of claims 18 to 20, wherein the gene cluster carried by the plasmid DNA is a nonribosomal peptide synthetase (NRPS) or a polyketide synthase (PKS).

22. The method of any one of claims 18 to 21, wherein the plasmid DNA is an adeno-associated virus (AAV) production plasmid.

Citation Information

Patent Citations

  • Methods for Producing Polynucleotides with Desired Characteristics by Iterative Selection and Recombination

    JP2000500981A

  • Methods and compositions for cell and metabolic engineering

    JP2000507444A