Preparation method of unit DNA composition, and manufacturing method of DNA connected body

By linking additional sequences and using precise enzyme treatments, the method achieves uniform molar distribution and accurate ligation of DNA units, addressing the OGAB method's precision issues and improving concatemer production efficiency.

JP2025169974APending Publication Date: 2025-11-14SYNPLOGEN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025140092
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-01-21
Filing Date
2025-08-26
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for assembling multiple DNA unit fragments, such as the OGAB method, face challenges in precisely controlling the molar ratio due to inaccuracies in DNA measurement and length distribution, leading to errors in DNA concatemer production.

Method used

A method involving linking an additional sequence to each DNA unit molecule, measuring and aliquoting to uniformize the molar numbers, using restriction enzymes to remove the additional sequence, and ligating vector DNA with DNA units to form DNA concatemers, with specific enzyme choices and conditions to maintain order and uniformity.

Benefits of technology

This approach results in a more uniform molar distribution of DNA units, reducing errors and enabling efficient ligation into concatemers, enhancing transformation efficiency and accuracy in microbial cells like Bacillus subtilis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169974000003
    Figure 2025169974000003
  • Figure 2025169974000004
    Figure 2025169974000004
  • Figure 2025169974000005
    Figure 2025169974000005
Patent Text Reader

Abstract

To provide a preparation method of a unit DNA composition in which the mole number of a plurality of unit DNA are arrayed better, and a manufacturing method of a DNA connected body.SOLUTION: A preparation method of a unit DNA composition has: a process of preparing solution including a plurality of unit DNA, to which an additional sequence is connected, for each kind of the unit DNA; a process of, after preparation of each solution, measuring the concentration of the unit DNA in each solution in the state where an additional sequence is connected to the unit DNA, dispensing each solution on the basis of its result, and making the mole numbers of the unit DNA in each solution become closer to the same with each other. A manufacturing method of a DNA connected body has: a process of preparing a unit DNA composition; a process of preparing a vector DNA; a process of removing each additional sequence from the unit DNA in which an additional sequence in solution after preparation is connected by using restriction enzyme; and a process of connecting the vector DNA and each unit DNA with each other after the removal process.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for preparing a DNA unit composition and a method for producing a DNA concatemer. [Background technology]

[0002] In recent years, there has been active development of DNA synthesis technologies aimed at constructing long-chain DNA at the genome level. Known DNA synthesis techniques include chemically synthesized DNA and the assembly of DNA units amplified by PCR. However, it is known that the synthesis process of chemically synthesized DNA and PCR introduce random mutations into the synthesized DNA. Therefore, in gene assembly, it is necessary to confirm the DNA sequence at some stage up to the final stage and select the one with the desired sequence.

[0003] To confirm the base sequence, an automated fluorescent sequencer using the Sanger method is usually used. This method makes it possible to confirm a base sequence of approximately 800 consecutive bases in length in a single run. When confirming the base sequence of DNA unit fragments synthesized by chemical synthesis or PCR before gene assembly, reducing the number of base sequence determinations reduces time and financial costs. Therefore, the shorter the DNA unit fragments synthesized by chemical synthesis or PCR for gene assembly, the better.

[0004] However, if the DNA unit fragments used for gene assembly are short, it is necessary to assemble many DNA unit fragments.

[0005] Currently, one of the known methods for assembling multiple DNA unit fragments is a gene assembling method (OGAB method) that utilizes the plasmid transformation system of Bacillus subtilis. For example, Patent Document 1 discloses a method for preparing plasmid DNA for transforming Bacillus subtilis cells using the OGAB method.

[0006] The OGAB method uses so-called multimeric plasmids, in which multiple plasmid units exist in a single DNA molecule through homologous recombination between plasmid molecules. According to the OGAB method, the DNA molecule used for transformation does not need to be circular; as long as it is in a tandem repeat configuration, in which one plasmid unit and each DNA unit used for integration appear repeatedly in the same direction, plasmid transformation is possible.

[0007] As described above, the OGAB method requires the preparation of DNA molecules ligated in tandem repeats. Therefore, when multiple DNA unit fragments are used, they must be ligated to a plasmid. However, the more DNA unit fragments there are, the more difficult it becomes to ligate them to form tandem repeats. Therefore, in order to ligate a large number of DNA unit fragments, it is desirable to keep the molar ratio of each DNA unit fragment close to 1 during ligation. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Patent No. 4479199 Summary of the Invention [Problem to be solved by the invention]

[0009] However, in reality, it is difficult to precisely control the number of moles of each DNA unit. One reason for this is that DNA measurement methods using fluorescent double-stranded DNA intercalators such as SYBR Geen I can only measure the number of molecules with a reproducibility of about ±20% due to factors such as fading of the fluorescent substance during measurement. In addition, these measurement methods are difficult to precisely control the number of moles of each DNA unit. The OGAB method measures the weight per unit volume, but since the amount of DNA unit is calculated based on the number of moles per unit volume, it is necessary to convert the DNA weight measured by the above measurement method into a molar concentration. Therefore, when the distribution of lengths of DNA unit molecules is wide, even for DNA molecules with the same number of moles, the weight is proportional to the length of the DNA unit, and particularly when the measured weights differ by several times or more, the value calculated based on the measured values ​​often contains a large error. The method of Patent Document 1 also attempts to adjust the molar ratio of each DNA unit to 1, but because the length distribution of each DNA unit is wide, it is not possible to strictly control the molar ratio.

[0010] The present invention has been made in view of the above circumstances, and aims to provide a method for preparing a DNA unit composition in which the molar numbers of multiple DNA unit molecules are more uniform, and a method for producing a DNA concatemer. [Means for solving the problem]

[0011] The present inventors have found that measuring the number of moles of DNA unit molecules while each DNA unit molecule is linked to an additional sequence reduces errors in the measurement results, and have thus completed the present invention. More specifically, the present invention provides the following.

[0012] (1) preparing a solution containing a plurality of DNA unit fragments each having an additional sequence linked thereto, for each type of DNA unit fragment; after preparing each of the solutions, measuring the concentration of the DNA unit molecules in each of the solutions while the additional sequence is linked to the DNA unit molecules, and then aliquoting each of the solutions based on the results to make the number of moles of the DNA unit molecules in each solution similar to each other.

[0013] (2) The method for preparing a DNA unit composition according to (1), wherein the DNA unit to which the additional sequence is linked has a circular structure, and the additional sequence is a plasmid DNA sequence having an origin of replication.

[0014] (3) A method for preparing a DNA unit composition according to (1) or (2), wherein the standard deviation of the distribution of the total length of each DNA unit and the total length of the additional sequence linked to each DNA unit is within ±20% of the average total length.

[0015] (4) A method for preparing a DNA unit composition according to any one of (1) to (3), wherein the average base length of the additional sequences linked to each DNA unit is at least twice the average base length of the DNA unit.

[0016] (5) A method for preparing a DNA unit composition according to any one of (1) to (4), wherein the length of each of the DNA unit fragments is 1600 bp or less.

[0017] (6) The DNA unit fragments are used to prepare DNA concatemers containing DNA assemblies composed of the DNA unit fragments, The method for preparing a DNA unit composition according to any one of (1) to (5), wherein the step of preparing a solution containing the DNA unit comprises a step of designing the DNA unit to have non-palindromic sequences at its ends, with non-palindromic sequences near sequences at positions obtained by equally dividing the DNA assembly as boundaries, so that when the base length of the sequence of the DNA assembly is divided by the number of types of DNA unit, the base lengths of the respective DNA unit segments are equal.

[0018] (7) A method for preparing a DNA concatenate for transforming a microbial cell, which contains more than one DNA assembly unit consisting of a vector DNA containing an origin of replication effective in a host microorganism and an assembly DNA, comprising: Preparing a DNA unit composition by the method according to any one of (1) to (6); providing the vector DNA; removing each additional sequence from the DNA unit fragments to which the additional sequences have been linked in the prepared solution using a restriction enzyme; a step of ligating the vector DNA and each of the DNA unit fragments to each other after the removal step, the vector DNA and each of the DNA unit fragments have a structure that allows them to be repeatedly linked to each other while maintaining their order, The method for producing a DNA concatemer, wherein the DNA assembly is composed of DNA in which each of the DNA unit molecules is linked to another DNA unit molecule.

[0019] (8) A method for producing a DNA concatemer according to (7), comprising a step of adjusting the coefficient of variation of the concentrations of the vector DNA and each of the DNA unit molecules in the ligation step based on a relational expression between the yield of DNA fragments of the target ligation number, which is represented by the product of the number of DNA unit molecules constituting an assembly unit and the number of assembly units, and the coefficient of variation of the concentrations of the DNA fragments.

[0020] (9) The method for preparing a DNA concatemer according to (7) or (8), wherein the restriction enzyme is a type II restriction enzyme.

[0021] (10) The method for preparing a DNA concatemer according to any one of (7) to (9), further comprising the step of mixing solutions containing two or more types of DNA unit fragments among the solutions containing the prepared DNA unit fragments before the removal step.

[0022] (11) The method for producing a DNA concatemer according to any one of (7) to (10), further comprising the step of inactivating the restriction enzyme after the removal step and before the ligation step.

[0023] (12) The method for producing a DNA concatemer according to any one of (7) to (11), wherein the microorganism is Bacillus subtilis. [Effects of the Invention]

[0024] According to the present invention, there are provided a method for preparing a DNA unit composition in which the molar numbers of multiple DNA unit molecules are more uniform, and a method for producing a DNA concatemer. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 shows a vector DNA according to one embodiment of the present invention. [Figure 2] FIG. 1 shows the structures of synonymous codon mutants of the 10th fragment of the DNA unit fragment in Example 1 of the present invention. [Figure 3]FIG. 1 shows a photograph of electrophoresis of a crude plasmid containing fragment 01 or fragment 21 of the DNA unit fragments in Example 1 of the present invention, and a highly purified plasmid obtained by purifying the crude plasmid after restriction enzyme treatment. [Figure 4] FIG. 1 shows electrophoresis photographs of the purified plasmids containing the DNA unit fragments in Example 1 of the present invention, after simultaneous treatment with each restriction enzyme for a group treated with BbsI, a group treated with AarI, and a group treated with BsmBI. [Figure 5] FIG. 1 shows a photograph of electrophoresis after merging groups of purified plasmids containing the DNA unit fragments in Example 1 of the present invention, including a group treated with BbsI, a group treated with AarI, and a group treated with BsmBI, which were digested together with the respective restriction enzymes. [Figure 6] FIG. 1 shows the distribution of the number of DNA unit molecules before and after size fractionation after treating the plasmids containing the DNA unit fragments in Example 1 of the present invention collectively with the respective restriction enzymes (groups treated with BbsI, group treated with AarI, and group treated with BsmBI) and merging the groups. [Figure 7] FIG. 1 shows the rate of change in the number of molecules of each DNA unit fragment before and after size fractionation after treating the plasmids containing the DNA unit fragments in Example 1 of the present invention collectively with the respective restriction enzymes (groups treated with BbsI, group treated with AarI, and group treated with BsmBI) and merging the groups. [Figure 8] FIG. 1 is a photograph showing electrophoresis after ligation of DNA unit fragments and vector DNA in Example 1 of the present invention. [Figure 9] FIG. 1 shows a photograph of electrophoresis of a DNA concatenate obtained by ligating a DNA unit cell and a vector DNA into Bacillus subtilis, followed by transformation of the DNA concatenate into Bacillus subtilis, and then treatment of the plasmid with a restriction enzyme from a plurality of transformed Bacillus subtilis cells in Example 1 of the present invention. [Figure 10]FIG. 1 shows a photograph of electrophoresis after restriction enzyme treatment of a Bacillus subtilis clone containing the desired DNA assembly, selected from the results of electrophoresis photographs taken after restriction enzyme treatment of the plasmids extracted from a plurality of transformed Bacillus subtilis in Example 1 of the present invention. [Figure 11] FIG. 1 shows that the selected DNA aggregates formed lambda phage DNA plaques in Example 1 of the present invention. [Figure 12] FIG. 1 shows photographs of the genome of the selected DNA assembly and wild-type lambda phage after restriction enzyme treatment with AvaI in Example 1 of the present invention. [Figure 13] FIG. 10 is a photograph showing electrophoresis of purified plasmids containing DNA fragments after en bloc restriction enzyme digestion with AarI in Example 2 of the present invention. [Figure 14] FIG. 10 is a photograph showing electrophoresis after ligation of DNA unit fragments and vector DNA in Example 2 of the present invention. [Figure 15] FIG. 10 is a photograph showing electrophoresis of a DNA concatenate obtained by ligating a DNA unit cell and a vector DNA into Bacillus subtilis, followed by transformation of the DNA concatenate into Bacillus subtilis, extraction of plasmids from a plurality of transformed Bacillus subtilis cells, and treatment of the plasmids with restriction enzymes in Example 2 of the present invention. [Figure 16] FIG. 10 is a photograph showing the electrophoresis results of Example 2 of the present invention, in which plasmids were extracted from multiple transformed Bacillus subtilis strains and treated with restriction enzymes. From the results of the electrophoresis photographs taken after the plasmids were treated with restriction enzymes, a Bacillus subtilis clone containing the target accumulated DNA was selected and then treated with restriction enzymes. [Figure 17] FIG. 1 shows photographs of electrophoresis of DNA (A) to (H) used in Test Example 1. [Figure 18] 1 is a graph showing the number of transformants appearing in Bacillus subtilis competent cells transformed with the DNAs (A) to (H) used in Test Example 1. [Figure 19]1 shows graphs showing the relationship between the CV (%) of the variation in the concentration of DNA unit fragments and the relative amount of each DNA unit fragment at each gene accumulation scale in Simulation 1. (a) shows a graph for 6-fragment accumulation, (b) shows a graph for 13-fragment accumulation, (c) shows a graph for 26 fragments, and (d) shows a graph for 51-fragment accumulation. [Figure 20] 10 is a graph showing the relationship between N (the number of DNA units contained in one ligation product) and the number of ligation product molecules when CV=20% in the accumulation of 6 fragments in Simulation 1. [Figure 21] (a) is a graph of the λ function for the CV (%) of the concentration variation of the DNA unit fragments, derived from fitting to an exponential distribution curve in Simulation 1, and (b) is a graph of the λ function for the CV (%) of the concentration variation of the DNA unit fragments, calculated from the average N value of the hypothetical ligation product. [Figure 22] FIG. 1 shows misligation sites for #1, #2, #5, #7, #8, #9, #10, and #11 of the aggregates obtained in the λ phage genome reconstruction experiment in Simulation 1. [Figure 23] FIG. 10 is a photograph of pulsed-field gel electrophoresis of ligation products of unit DNA51 fragments with varying numbers of fragments with CV=6.6% in an experiment of reconstructing the λ phage genome in Simulation 1. [Figure 24]Graphs comparing actual ligation efficiency with ligation efficiency from ligation simulation in Simulation 1. (a) shows a graph compared with a simulation with a potential ligation number of 95%, (b) shows a graph compared with a simulation with a potential ligation number of 96%, (c) shows a graph compared with a simulation with a potential ligation number of 97%, (d) shows a graph compared with a simulation with a potential ligation number of 98%, (e) shows a graph compared with a simulation with a potential ligation number of 99%, and (f) shows a graph compared with a simulation with a potential ligation number of 100%. [Figure 25] This graph shows the relationship between the variation CV (%) of the unit DNA concentration fragments and the relative amount of each unit DNA fragment at each gene accumulation scale in Simulation 1, obtained using the general formula f(N) = 0.0058 * CV (%) * exp(-0.058 * CV (%) * N). [Figure 26] 1 is a graph showing the relationship between the variation in concentration of DNA unit fragments and the average number of DNA unit fragments per ligation product, obtained using the general formula f(N)=0.0058*CV(%)*exp(-0.0058*CV(%)*N). DETAILED DESCRIPTION OF THE INVENTION

[0026] Hereinafter, an embodiment of the present invention will be described, but the present invention is not limited to this.

[0027] <Method for preparing DNA unit compositions> The method for preparing a DNA unit composition of the present invention comprises the steps of: preparing a solution containing a plurality of DNA unit molecules, each of which has an additional sequence linked thereto; and measuring the concentration of the DNA unit molecules in each solution after preparing the solutions while the additional sequence is still linked to the DNA unit molecules; and, based on the results, aliquoting each solution so that the number of moles of DNA unit molecules in each solution is similar. In this specification, the types of "DNA unit molecules" are distinguished by their respective base sequences. Furthermore, "DNA unit molecules" include those with and without restriction enzyme recognition sites added.

[0028] In the present invention, when measuring the concentration of each DNA unit in a solution containing DNA units, an additional sequence is linked to each DNA unit. This reduces the distribution of base sequence lengths when measuring the solution concentration due to the linked additional sequence. This reduces the error in the molar number of each DNA unit calculated based on the measurement results. Therefore, by aliquoting each solution based on the measurement results and adjusting the molar number of DNA units in each solution to be identical, it is easy to bring the molar ratio in each solution closer to 1. The "concentration of DNA units in solution" measured above refers to the molar concentration of the DNA units. The method for measuring the molar concentration of DNA units in a solution is not particularly limited. For example, the molar concentration of DNA units in a solution may be calculated from the measured value of the mass % of DNA units in the solution. The molar concentration of DNA units in a solution is preferably measured using a means capable of measuring DNA weight concentration with an accuracy of within ±20%, and more specifically, ultraviolet absorption spectroscopy using a microspectrophotometer is preferred.

[0029] The step of preparing a solution containing a DNA unit fragment linked to an additional sequence is not particularly limited, and may be carried out, for example, by preparing a DNA unit fragment and then linking the additional sequence to the DNA unit fragment. good.

[0030] The DNA unit fragments may be prepared by using pre-synthesized DNA fragments or by preparing DNA unit fragments. DNA unit fragments can be prepared by conventional methods, such as polymerase chain reaction (PCR) or chemical synthesis. When restriction enzyme recognition sequences are added to the DNA unit fragments, they can be prepared by PCR using primers containing restriction enzyme recognition sequences that generate protruding ends in the base sequence of a template DNA, or by chemical synthesis after incorporating restriction enzyme recognition sequences into the ends to generate desired protruding sequences. The base sequence of the prepared DNA unit fragments can be confirmed by conventional methods, such as by incorporating the DNA unit fragments into a plasmid and determining the base sequence using an automated fluorescent sequencer using the Sanger method.

[0031] The additional sequence is not particularly limited and may be a linear DNA or a circular plasmid. When a circular plasmid DNA sequence is used, the DNA unit fragment to which the additional sequence is linked has a circular structure, making it possible to transform a host such as Escherichia coli.

[0032] The type of plasmid DNA is not particularly limited, but in order to replicate the plasmid DNA in a transformed host, it is preferable that the plasmid DNA sequence has an origin of replication. Specifically, pUC19, a high-copy plasmid vector for E. coli, or a plasmid derived therefrom is preferred. Furthermore, it is preferable that all DNA unit fragments be cloned into the same type of plasmid vector, as this reduces the length distribution between DNA fragments to which additional sequences are linked and allows the molar numbers of DNA unit fragments to be closer to the same.

[0033] The additional sequence may be linked to the DNA unit fragment by ligation using DNA ligase, for example, or by TA cloning when linking to a plasmid DNA.

[0034] The standard deviation of the distribution of the total length of each DNA unit base and the base length of the additional sequence linked to each DNA unit base is not particularly limited, but the smaller the standard deviation, the smaller the error in the number of moles of each DNA unit base calculated based on the measurement results of the DNA concentration in solution, and therefore the number of moles of DNA unit bases in each solution can be made closer to each other. Specifically, the standard deviation of the distribution of the total length of each DNA unit base and the base length of the additional sequence linked to each DNA unit base is preferably within ±20% of the average total length, more preferably within ±15%, even more preferably within ±10%, even more preferably within ±5%, even more preferably within ±1%, and most preferably within ±0.5%.

[0035] The average base length of the additional sequences linked to each DNA unit cell is not particularly limited. However, the longer the average base length of the additional sequences linked to each DNA unit cell, the smaller the error in the number of moles of each DNA unit cell calculated based on the measurement results of the DNA concentration in the solution, and the closer the number of moles of DNA unit cells in each solution to each other can be. Specifically, the average base length of the additional sequences linked to each DNA unit cell is preferably at least twice the average base length of the DNA unit cells, more preferably at least five times, even more preferably at least ten times, and most preferably at least 20 times. Furthermore, if the average base length of the additional sequences linked to the DNA unit cells is too long, it becomes difficult to manipulate the DNA unit cells to which the additional sequences are linked. Therefore, the average base length of the additional sequences linked to each DNA unit cell is preferably 10,000 times or less (specifically, 5,000 times or less, 3,000 times or less, 1,000 times or less, 500 times or less, 250 times or less, 100 times or less, etc.) the average base length of the DNA unit cells.

[0036] The length of each DNA unit is not particularly limited, but when confirming the base sequence of the DNA unit, Furthermore, fewer sequencing runs are preferable, as this reduces time and financial costs. Therefore, the length of each DNA unit is preferably shorter; specifically, 1600 bp or less is preferred, and 1200 bp or less is even more preferred. In particular, when sequencing is performed using an automated fluorescent sequencer using the Sanger method, a sequence of approximately 800 consecutive bases can be confirmed in a single sequencing run. Therefore, the length of each DNA unit is most preferably 800 bp or less (specifically, 600 bp or less, 500 bp or less, 400 bp or less, 200 bp or less, 100 bp or less, etc.). Thus, if each DNA unit is short, a large number of DNA unit fragments will be required when used to prepare a DNA concatemer, as described below. However, when the DNA unit fragments prepared by the method of the present invention are used, a large number of DNA unit fragments can be concatenated, as described below. Furthermore, if each DNA unit fragment is too short, the number of DNA unit fragments will increase, resulting in reduced operational efficiency. Therefore, the length of each DNA unit fragment is preferably 20 bp or more, more preferably 30 bp or more, and even more preferably 50 bp or more.

[0037] The use of the DNA unit composition prepared by the preparation method of the present invention is not particularly limited, but the DNA unit composition prepared by the preparation method of the present invention can be used to prepare a DNA concatemer containing a DNA assembly composed of the DNA unit. When a DNA unit composition prepared by the preparation method of the present invention is used to prepare a DNA concatemer by the method described below, many DNA unit molecules (e.g., 50 or more types) can be ligated. This is thought to be because the molar numbers of each DNA unit molecule in the DNA unit composition prepared by the preparation method of the present invention are more accurately close to the same.

[0038] In the present invention, the step of preparing a solution containing unit DNA may include the step of designing unit DNA. The design of unit DNA is not particularly limited. For example, when a unit DNA composition is used to produce a DNA conjugate containing integrated DNA, when the base length of the sequence of the integrated DNA is divided by the number of types of unit DNA, the respective base lengths may be made equal by using non-palindromic sequences near the sequences at the positions where the integrated DNA is equally divided as boundaries. When the unit DNA is designed in this way, the length of each unit DNA becomes substantially the same length. Therefore, when used in the production of the DNA conjugate described later, in the size fractionation after electrophoresis after removing the additional sequence with a restriction enzyme, it appears as a band at substantially the same position, so the unit DNA can be recovered by one size fractionation, which is preferable in terms of improving work efficiency. The "near the sequence at the position where the integrated DNA is equally divided" is not particularly limited, but may be appropriately set according to the length of the base sequence. For example, when the base length of each unit DNA is 1000 bp, it may be within 100 bp (specifically, within 90 bp, within 80 bp, within 70 bp, within 60 bp, within 50 bp, within 30 bp, within 20 bp, within 10 bp, within 5 bp, etc.) from the "position where the integrated DNA is equally divided".

[0039] Also, when designing the unit DNA as described above, when producing a DNA conjugate containing the target integrated DNA, it is preferable to design the unit DNA to have a non-palindromic sequence (a sequence that is not a palindrome sequence) at the end of the unit DNA. When the unit DNA designed in this way has its non-palindromic sequence as a protruding sequence, it has a structure that can be repeatedly ligated while maintaining the order with each other as described later.

[0040] <Method for Producing DNA Conjugate> The present invention also includes a method for producing a DNA conjugate. The method for producing a DNA conjugate of the present invention includes the step of preparing a unit DNA composition by the above-described method, the step of preparing vector DNA, the step of removing each additional sequence from the unit DNA to which the additional sequence in the prepared solution is ligated using a restriction enzyme, and the step of ligating the vector DNA and each unit DNA to each other after the removal step.

[0041] The DNA concatenator contains more than one DNA assembly unit and is used for transforming microbial cells. The DNA assembly unit consists of a vector DNA and an assembly DNA. The number of DNA assembly units in the DNA concatenator is not particularly limited as long as it is more than one, but to increase transformation efficiency, it is preferably 1.5 or more, more preferably 2 or more, even more preferably 3 or more, and most preferably 4 or more.

[0042] The vector DNA contains an origin of replication that is effective in the host microorganism to be transformed. The vector DNA is not particularly limited as long as it has a sequence that enables DNA replication in the microorganism that can be transformed with the DNA concatenate, and examples thereof include the sequence of an origin of replication that is effective in the Bacillus bacteria (Bacillus subtilis) described below. The sequence of an origin of replication that is effective in Bacillus subtilis is not particularly limited, and examples of those having a θ-type replication mechanism include sequences such as the origin of replication contained in plasmids such as pTB19 (Imanaka, T., et al. J. Gen. Microbioi. 130, 1399-1408 (1984)), pLS32 (Tanaka, T. and Ogra, M. FEBS Lett. 422, 243-246 (1998)), and pAMβ1 (Swinfield, TJ, et al. Gene 87, 79-90 (1990)).

[0043] The DNA assembly is composed of DNA in which the above-mentioned DNA unit fragments are linked together. The DNA in the present invention refers to DNA to be cloned, and its type and size are not particularly limited. Specifically, it may be a naturally occurring sequence from a prokaryote, eukaryote, virus, or the like, or an artificially designed sequence. As described above, in the method of the present invention, it is preferable to use DNA with a long base length, since a large number of DNA unit fragments can be linked onto a plasmid. Examples of DNA with a long base length include a gene cluster constituting a metabolic pathway, and the entire or partial genomic DNA of a phage or the like.

[0044] The DNA assembly unit may or may not contain an appropriate base sequence other than the vector DNA or the assembled DNA, as necessary. For example, when preparing a plasmid for expressing a gene contained in the assembled DNA, the DNA assembly unit may contain a base sequence that controls transcription / translation, such as a promoter, operator, activator, or terminator. Specific examples of promoters when Bacillus subtilis is used as a host include Pspac (Yansura, D. and Henner, D.J. Pro. Natl. Acad. Sci. 2012; 10:111-113, 2012), whose expression can be controlled with IPTG (isopropyl sD-thiogalactopyranoside). Sci. USA 81, 439-443 (1984)), or the Pr promoter (Itaya, M. Biosci. Biotechnol. Biochem. 63, 602-604 (1999)).

[0045] The vector DNA and each DNA unit have a structure that allows them to be repeatedly ligated to each other while maintaining their order. As used herein, "ligating to each other while maintaining their order" refers to DNA unit or vector DNA having adjacent sequences in the DNA assembly unit being ligated while maintaining their order and orientation. Furthermore, "repeatedly ligated" refers to the ligation of the 5'-end of a DNA unit or vector DNA having a 5'-end nucleotide sequence with the 3'-end of a DNA unit or vector DNA having a 3'-end nucleotide sequence. Specific examples of such DNA unit include those having ends that allow them to be repeatedly ligated to each other while maintaining their order, utilizing the complementarity of the nucleotide sequences at the overhanging ends of the fragments. The structure of this overhang is not particularly limited, including the shape of the 5'-end overhang and the 3'-end overhang, as long as it is a non-batch sequence.

[0046] The cohesive ends are preferably prepared by removing each additional sequence from the DNA unit fragments using a restriction enzyme. In this case, the DNA unit fragments preferably have a restriction enzyme recognition sequence so that the additional sequence can be removed by the restriction enzyme. Vector DNA can be prepared, for example, by treating the vector DNA with a restriction enzyme so as to provide cohesive ends that allow repeated ligation of the vector DNA with the DNA unit fragments while maintaining the order.

[0047] The restriction enzyme used to remove the additional sequence is not particularly limited, but is preferably a type II restriction enzyme, more preferably a type IIS restriction enzyme such as AarI, BbsI, BbvI, BcoDI, BfuAI, BsaI, BsaXI, BsmAI, BsmBI, BsmFI, BspMI, BspQI, BtgZI, FokI, or SfaNI, which can generate cohesive ends of any sequence at a fixed distance outside the recognition sequence. Use of a type IIS restriction enzyme allows the cohesive ends of the DNA unit fragments to be different at each ligation site, thereby maintaining the order of ligation. Similarly to the preparation of the DNA unit fragments, the vector DNA is preferably prepared using a type IIS restriction enzyme to generate cohesive ends that can be repeatedly ligated to the DNA unit fragments while maintaining the order.

[0048] When the DNA unit fragments are classified into groups based on the type of restriction enzyme used to remove the additional sequence, solutions containing two or more types of DNA unit fragments can be mixed for each group before the removal step. This eliminates the need for restriction enzyme treatment for each DNA unit fragment, allowing restriction enzyme treatment for each group to be performed in a single step. Furthermore, when fractionating DNA unit fragments by electrophoresis, for example, DNA unit fragments can be recovered in a single fractionation, improving work efficiency. Furthermore, when there are multiple DNA unit fragment groups, fractionation into each group and recovery of DNA unit fragments can result in differences in the amount recovered between groups, which can lead to errors in the number of moles matched between DNA unit fragments. Therefore, when classifying DNA unit fragments into groups based on the type of restriction enzyme used to remove the additional sequence, it is preferable to have as few groups as possible; that is, it is preferable to use as few types of restriction enzymes as possible to remove the additional sequence. Therefore, the number of types of restriction enzymes used is preferably five or fewer, more preferably three or fewer, and most preferably one. In other words, when only one type of restriction enzyme is used, solutions containing all the DNA unit fragments can be mixed, dramatically improving work efficiency and reducing the likelihood of errors in the number of moles matched between DNA unit fragments. Since the restriction enzymes are mixed in approximately equimolar amounts, a large number of DNA unit fragments can be ligated even when such a mixture is used.

[0049] When a Type IIS restriction enzyme recognition sequence is added to a DNA unit, the restriction enzyme recognition sequence of the DNA unit is designed so that it does not recognize the sequence of each DNA unit. That is, when a certain Type IIS restriction enzyme is to be used, if that Type IIS restriction enzyme does not recognize the sequence of one DNA unit but recognizes the sequence of the other DNA unit, the restriction enzyme recognition sequence of each DNA unit is designed so that a different Type IIS restriction enzyme is used for the other DNA unit. When designed in this way, a different Type IIS restriction enzyme is used for each DNA unit, and the DNA unit can be classified into the above groups depending on the type of Type IIS restriction enzyme. If there is a Type IIS restriction enzyme that does not recognize the sequence of any of the DNA unit molecules, DNA unit molecules can be designed with the recognition sequence of that restriction enzyme added, and the added sequences can be removed from the DNA unit molecules all at once using a single Type IIS restriction enzyme.

[0050] The step of ligating the vector DNA and each DNA unit fragment to each other is not particularly limited, but can be carried out by following restriction enzyme treatment, fractionating the additional sequence after restriction enzyme treatment and the DNA unit fragments, and then ligating the fractionated DNA unit fragments to the vector DNA using DNA ligase or the like. This produces a DNA concatenate for microbial transformation. Note that the DNA unit fragments used in the ligation step do not have restriction enzyme recognition sequences added to them.

[0051] The method for fractionating the additional sequence and DNA units is not particularly limited, but is preferably a method that does not disrupt the molar ratio relationship between the DNA units after restriction enzyme treatment, and specifically, agarose gel electrophoresis is preferred.

[0052] The method for ligating the DNA unit fragments and the vector DNA is not particularly limited, but is preferably carried out in the presence of polyethylene glycol and a salt. The salt is preferably a monovalent alkali metal salt. More specifically, the ligation is preferably carried out in a ligation reaction solution containing 10% polyethylene glycol 6000 and 250 mM sodium chloride. The concentration of each DNA unit fragment in the ligation reaction solution is not particularly limited, but is preferably 1 fmol / μl or more. The reaction temperature and time for the ligation are not particularly limited, but are preferably 37°C for 30 minutes or more. The concentration of the vector DNA in the ligation reaction solution is preferably measured before the reaction, and adjusted to be equimolar to the DNA unit fragments.

[0053] The method for preparing a DNA concatemer according to the present invention may or may not include a step of adjusting the coefficient of variation (hereinafter referred to herein as "coefficient of variation 2") of the concentrations of the vector DNA and each DNA unit fragment in the ligation step based on a relational expression (hereinafter referred to herein as "relational expression") between the yield of DNA fragments with a target number of ligations, which is represented by the product of the number of DNA unit fragments constituting an assembly unit and the number of assembly units, and the coefficient of variation (hereinafter referred to herein as "coefficient of variation 1") of the concentration of the DNA fragments. Note that coefficient of variation 1 is a coefficient of variation used for convenience in the relational expression, and coefficient of variation 2 is the coefficient of variation of the concentrations of the DNA unit fragments and vector DNA in the actual ligation step. By including this adjustment step, when attempting to ligate a desired number of DNA unit fragments (e.g., 50 DNA unit fragments) in the ligation step, the desired ligation can be achieved by adjusting coefficient of variation 2 to fall within the range indicated by the above relational expression.

[0054] The target ligation number refers to the number of desired DNA fragments to be ligated in the ligation step, and is more specifically expressed as the product of the number of DNA unit fragments constituting the ligated assembly unit and the number of assembly units. The "yield of DNA fragments corresponding to the target ligation number" refers to the ratio of the number of DNA fragments constituting the assembly DNA unit after ligation to the total number of DNA fragments used for ligation.

[0055] The relational equation according to the present invention is an equation showing the relationship between the yield of DNA fragments with a target ligation number and the coefficient of variation 1, and can be, for example, an equation obtained from a computer-based ligation simulation. More specifically, the relational equation can be set by, for example, performing a ligation simulation on a population of DNA unit fragments (e.g., 10-30 populations) in which the coefficient of variation 1 is changed in 1% increments from 0 to 20%, examining the distribution (e.g., exponential distribution) of the number of DNA unit fragments in the DNA concatenation complex, creating a fitting curve for this distribution, and using this fitting curve. The specific tool used for the ligation simulation is not particularly limited, and conventional, well-known means can be used. For example, VBA (Visual Basic for Applications) in the spreadsheet software Excel® 2007 can be used. Simulations can be performed by programming using this and setting the required algorithm. Furthermore, the fitting curve can be created, for example, using the exponential approximation curve function of the spreadsheet software Excel® 2007 software. The coefficient of variation of 2 based on the relational equation can be adjusted, for example, by designing the relational equation, substituting the yield of the desired DNA fragment into the relational equation, and adjusting the operations in each step before ligation so that the DNA fragments after ligation have the calculated coefficient of variation of 1. The adjustment method is not particularly limited, but for example, when measuring the concentration of vector DNA or each DNA unit in each step such as the step of preparing vector DNA, the step of preparing DNA unit, and the step of ligating vector DNA and each DNA unit to each other, the measurement equipment (spectrophotometer, spectrofluorophotometer, When selecting a measurement instrument (e.g., real-time PCR instrument), it is possible to select a measurement instrument whose measurement error is known in advance so that the coefficient of variation 2 becomes the desired coefficient of variation.

[0056] Although there are no particular limitations on the coefficient of variation 2, the smaller the error in the concentration of DNA unit molecules during ligation, the more DNA unit molecules can be ligated. For this reason, the coefficient of variation 2 is preferably 20% or less, more preferably 15% or less, even more preferably 10% or less, even more preferably 8% or less, and most preferably 5% or less.

[0057] The method for preparing a DNA concatemer of the present invention may further include a step of inactivating a restriction enzyme after the removal step and before the ligation step. If a DNA unit contains a restriction enzyme cleavage site used to cleave the additional sequence of another DNA unit, it is difficult to mix DNA unit populations containing the additional sequence before inactivation of the restriction enzyme. Therefore, it is not possible to fractionate DNA units en masse by merging them. However, by inactivating the restriction enzyme, DNA unit populations containing the additional sequence can be merged after inactivation, enabling the DNA unit to be fractionated en masse. This facilitates the preparation of DNA concatemers containing a larger number of assembled DNA units during the ligation step, thereby facilitating the transformation of Bacillus subtilis. Restriction enzyme inactivation can be performed by conventional methods known in the art, such as phenol-chloroform treatment.

[0058] The host microorganism to be transformed is not particularly limited as long as it has natural transformation ability. Examples of natural transformation ability include those that process DNA into single-stranded DNA before incorporating it. Specific examples include bacteria of the genus Bacillus, Streptococcus, Haemophilus, and Neisseria. Examples of Bacillus bacteria include B. subtilis, B. megaterium, and B. stearothermophilus. Of these, the most preferred microorganism is Bacillus subtilis, which has excellent natural transformation ability and recombination ability.

[0059] The DNA concatenate prepared by the method of the present invention can be used to transform microbial cells. A known method suitable for each microorganism can be selected as the method for rendering the microorganism to be transformed competent. Specifically, for example, in the case of Bacillus subtilis, the method described in Anagnostopoulou, C. and Spizizen, JJ Bacteriol., 81, 741-746 (1961) is preferably used. Furthermore, a known method suitable for each microorganism can be used as the transformation method. There are no particular limitations on the volume of the ligation product to be added to the competent cells. Preferably, the volume is 1 / 20 to the same volume as the competent cell culture medium, and more preferably, half the volume. Known methods can also be used to purify plasmids from transformants.

[0060] The presence of the integrated DNA in the plasmid purified from the transformant can be confirmed by the size pattern of fragments generated by restriction enzyme digestion, PCR, or base sequencing. If the inserted DNA has a function such as substance production, it can be confirmed by detecting that function. [Example]

[0061] The present invention will be described in more detail with reference to the following examples. However, the following examples are merely illustrative of the present invention, and the scope of the present invention is not limited to the following examples.

[0062] (material) Bacillus subtilis was used as the microbial cell to be transformed. The RM125 strain (Uozumi, T., et al. Moi. Gen. Genet., 152, 65-69 (1977)) and its derivative, BUSY9797, were used. pGET118 (Kaneko, S., et al. Nucleic Acids Res. 31, e112 (2003)) was used as the vector DNA replicable in Bacillus subtilis, and pGETS118-AarI-pBR (see SEQ ID NO: 1) and pGETS151-pBR (see SEQ ID NO: 2), constructed as described below, were used. Lambda phage DNA (Toyobo Co., Ltd.) (see SEQ ID NO: 3) and the mevalonate pathway artificial operon (see SEQ ID NO: 4), described below, were used as the integrating DNA. The antibiotic carbenicillin (Wako Pure Chemical Industries, Ltd.) was used to select E. coli carrying the plasmid DNA incorporating the unit DNA. The antibiotic tetracycline (Sigma) was used to select Bacillus subtilis. The type IIS restriction enzymes used were AarI (Thermo), BbsI (NEB), BsmBI (NEB), and SfiI (NEB). The restriction enzymes HindIII, PvuII, and T4 DNA were used. The ligase used was manufactured by Takara Bio. For general ligation to construct E. coli plasmids, the Takara Ligation Kit (Mighty) (Takara Bio) was used. For PCR reactions to prepare DNA fragments, Toyobo's KOD plus polymerase was used. For colony PCR to determine the base sequence of DNA cloned into the plasmid, Takara Bio's Ex-Taq HS was used. The plasmid DNA used as the additional sequence to incorporate the DNA fragments was pMD-19 (simple) (Takara Bio). Plasmid Safe, an enzyme used to purify circular plasmids, was manufactured by EPICENTER. For agarose gel electrophoresis, 2-Hydroxyethyl agarose (Sigma), a low-melting-point agarose gel for DNA electrophoresis, or UltraPure Agarose (Invitrogen) was used. Phenol:chloroform:isoamyl alcohol (25:24:1) and TE-saturated phenol (containing 8-quinolinol) were used to inactivate restriction enzymes and were manufactured by Nacalai Tesque. Lambda terminase was manufactured by EPICENTER. Lambda phage was packaged using Agilent Technologies' Gigapack III Plus Packaging Extract. Lysozyme was manufactured by Wako Pure Chemical Industries, Ltd. LB medium components and agar were manufactured by Becton Dickinson. IPTG (isopropyl sD-thiogalactopyranoside) was manufactured by Wako Pure Chemical Industries, Ltd. All other medium components and biochemical reagents were manufactured by Wako Pure Chemical Industries, Ltd. Unless otherwise specified, E. coli DH5α, JM109, or TOP10 strains were used for constructing plasmids. The constructed plasmids were purified from E. coli using Qiagen's QIAprep Spin Miniprep Kit for small-scale purification, and Qiagen's QIAfilter Midi Kit for large-scale purification. DNA was cleaned up from the enzyme reaction solution using Qiagen's MinElute Reaction Cleanup Kit or Qiagen's QIAquick PCR purification Kit.The Qiagen MinElute Gel Extraction Kit was used to purify gel blocks obtained by standard agarose gel electrophoresis. A Thermo Nano-Drop 2000 ultra-microspectrophotometer was used. For sequencing, a fluorescent automated sequencer, the 3130xl Genetic Analyzer, manufactured by Applied Biosystems, was used. Other general DNA manipulations were performed according to standard protocols (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1989)). Transformation of Bacillus subtilis and plasmid extraction were performed as previously described (Tsuge, K., et al., Nucleic Acids Res. 31, e133.(2003)).

[0063] (Construction of vector DNA used for integration) The vector DNA pGETS118-AarI-pBR (SEQ ID NO: 1) used for accumulating lambda phage DNA was constructed through a multi-step process based on the E. coli-B. subtilis shuttle plasmid vector pGETS118 (Kaneko, et al., Nucleic Acids Res., 31, e112. (2003)), which contains the replication origin oriS of the E. coli F factor and the replication origin repA that functions in B. subtilis. The structure is shown in Figure 1. The cloning site for the integrated gene is located between two AarI cleavage sites, which are removed during integration. Between these two AarI cleavage sites, the replication origin of the multicopy E. coli plasmid pBR322 and the ampicillin resistance gene have been inserted to facilitate vector isolation in E. coli. Furthermore, the natural AarI cleavage site present in the tetracycline resistance gene within pGETS118 was eliminated by introducing a single base mutation that did not affect the amino acid sequence of the tetracycline resistance gene (tetL). pGETS151-pBR (SEQ ID NO: 2), the vector DNA used to assemble the mevalonate pathway artificial operon, was synthesized using the above-mentioned pGETS118-AarI-pBR DNA as a template, with three pairs of primers: Part A (5'-TAGGGTCTCAaagcggccgcaagctt-3' (see SEQ ID NO: 5) and 5'-TAGGGTCTCAGCggccaagaaggcc-3' (see SEQ ID NO: 6)), Part B (5'-TAGGGTCTCAccGCCCTTCCCGGTCGATAT-3' (see SEQ ID NO: 7) and 5'-TAGGGTCTCAtaTTAGC This vector DNA was constructed by ligating fragments amplified with pGETS118-AarI-pBR (5'-TAGGGTCTCAAAtaactggaaaaaattagtgtctcatggttcg-3' (see SEQ ID NO: 9) and 5'-TAGGGTCTCAgcttaagtggtgggtagttgacc-3' (see SEQ ID NO: 10)), and compared to the original plasmid, gene regions that function only in E. coli (between cat and oriS, and between parA and parC) have been deleted (Figure 1). This vector DNA can only replicate in B. subtilis when genes are integrated, but it exhibits the same properties as pGETS118-AarI-pBR during gene integration. Approximately 10 μl (equivalent to 5 μg) of these plasmid solutions was added with 29 μl of sterile water, 5 μl of 10× Buffer_for_AarI (included with the restriction enzyme), 1 μl of 50× Oligocleotide (for cleavage activation) (also included with the restriction enzyme), and 5 μl of the restriction enzyme AarI (Thermo), and the reaction was carried out at 37°C for 2 hours.The resulting liquid was separated by low-melting-point agarose gel electrophoresis, and the approximately 15 kb fragment (in the case of pGETS118-AarI-pBR) or 4.3 kb fragment (in the case of pGETS151-pBR) of the vector main body was excised from the gel. The target vector DNA was purified and dissolved in 20 μl of TE. The concentration of the vector DNA was measured by taking 1 μl of this TE solution and measuring it using an ultramicrospectrophotometer.

[0064] (Setting the unit DNA division region) The diversity of four-base overhangs is 4 to the fourth power, or 256 possible combinations. Of these, overhang sequences for use in the present invention were selected based on the following criteria. First, all 16 palindromic sequences (Group 0) (AATT, ATAT, TATA, TTAA, CCGG, CGCG, GCGC, GGCC, ACGT, AGCT, TCGA, TGCA, CATG, CTAG, GATC, GTAC) were excluded because their complementary sequences are identical, allowing for ligation of identical fragments. The remaining 240 sequences encompass one sequence (e.g., CCTA) and its complementary sequence (TAGG). Therefore, theoretically, there are 120 combinations of overhang sequences that can be used for DNA ligation: 240 / 2 = 120. Based on differences in GC content and the order in which the GC bases appear, the overhang combinations were grouped based on the following criteria:

[0065] (Group I) Six combinations of protrusions consisting of only A and T (AAAA / TTTT, TAA A / TTTA,ATAA / TTAT,AATA / TATT,AAAT / ATTT,ATTA / TAAT). (Group II) All 32 combinations where the total of A and T is 3 and the total of C and G is 1 (CAAA / TTTG, ACAA / TTGT, AACA / TGTT, AAAC / GTTT, GAAA / TTTC, AGAA / TTCT, AAGA / TCTT, AAAG / CTTT, CAAT / ATTG, ACAT / ATGT, AACT / AGTT, AATC / GATT, GAAT / ATTC, AGAT / ATCT, AAGT / ACTT, AATG / CATT, CATA / TATG, ACTA / TAGT, ATCA / TGAT, ATAC / GTAT, GATA / TATC, AGTA / TACT, ATGA / TCAT, ATAG / CTAT, CTTA / TAAG, TCTA / TAGA, TTCA / TGAA, TTAC / GTAA, GTTA / TAAC, TGTA / TACA, TTGA / TCAA, TTAG / CTAA). (Group III) 44 combinations excluding 8 palindrome combinations out of all 52 combinations where the total of A and T is 2 and the total of C and G is 2 (AACC / GGTT, AACG / CGTT, AAGC / GCTT, AAGG / CCTT, ACAC / GTGT, ACAG / CTGT, ACCA / TGGT, ACCT / AGGT, ACGA / TCGT, ACTC / GAGT, ACTG / CAGT, AGAC / GTCT, AGAG / CTCT, AGCA / TGCT, AGGA / TCCT, AGTC / GACT, AGTG / CACT, ATCC / GGAT, ATCG / CGAT, ATGC / GCAT, ATGG / CCAT, CAAC / GTTG, CAAG / CTTG, CACA / TGTG, CAGA / TCTG, CATC / GATG, CCAA / TTGG, CCTA / TAGG, CGAA / TTCG, CGTA / TACG, CTAC / GTAG, CTCA / TGAG, CTGA / TCAG, CTTC / GAAG, GAAC / GTTC, GACA / TGTC, GAGA / TCTC, GCAA / TTGC, GCTA / TAGC, GGAA / TTCC, GGTA / TACC, GTCA / TGAC, GTGA / TCAC, TCCA / TGGA). (Group IV) Of the 32 combinations in which A and T total 1 and C and G total 3, there are 16 combinations in which C and G do not occur three times in a row (CACC / GGTG, CCAC / GTGG, CTCC / GGAG, CCTC / GAGG, CACG / CGTG, CCAG / CTGG, CTCG / CGAG, CCTG / CAGG, CAGC / GCTG, CGAC / GTCG, CTGC / GCAG, CGTC / GACG, GAGC / GCTC, GGAC / GTCC, GTGC / GCAC, GGTC / GACC). (Group V) Of the 32 combinations with a total of 1 A and 1 T and a total of 3 C and G, there were 16 combinations with 3 consecutive Cs and Gs (ACCC / GGGT, CCCA / TGGG, TCCC / GGGA, CCCT / AGGG, ACCG / CGGT, CCGA / TCGG, TCCG / CGGA, CCGT / ACGG, ACGC / GCGT, CGCA / TGCG, TCGC / GCGA, CGCT / AGCG, AGGC / GCCT, GGCA / TGCC, TGGC / GCCA, GGCT / AGCC). (Group VI) Six combinations in which the overhanging C and G are the only C and G (CCCC / GGGG, GCCC / GGGC, CGCC / GGCG, CCGC / GCGG, CCCG / CGGG, CGGC / GCCG).

[0066] In Examples 1 and 2, the boundary between the vector DNA and the DNA unit fragments was selected from Group 1 out of the above divided groups. Furthermore, a total of 60 overhanging combinations were selected as candidates for the boundaries between the DNA unit fragments, from Group III (44 combinations) and Group IV (16 combinations). Each overhanging combination was specified by first determining the complete base sequence of the target sequence to be accumulated, and then setting ideal dividing boundaries that equally divide this entire length. The base sequence used in Example 1 is explained below as a specific example.

[0067] In Example 1, the target of reconstruction was a 48,522 bp fragment of the lambda phage genome, which was the full length of 48,502 bp, to which a 16 bp cos site and a 4 bp protruding sequence required for assembly were added. Table 1 below shows the ideal division unit, actual division unit, and protruding base sequence of the assembled DNA in Example 1. This table shows the results. The DNA unit fragments, excluding the plasmid vector for assembly, were divided into 50 fragments of approximately equal size and then ligated together to attempt reconstitution. Ideally, all 50 fragments would be divided to equal lengths. However, to avoid introducing any base changes into the sequence to be assembled, it is necessary to create a 4-base 5'-end overhang for assembly depending on the original sequence. However, the probability that the above-described overhang sequences are present in the exact same amount on all ideal division boundaries is virtually zero, making it impossible to equally divide the DNA unit fragments at ideal division boundaries. In this example, a simulation was performed to assign overhang combinations to achieve lengths as close as possible to the ideal division unit. First, 970 bp, obtained by dividing the total length (48,522 bp) by 50, was used as the ideal division unit. The DNA unit fragments with the smallest absolute base numbers were named fragment 01, fragment 02, fragment 03, ..., fragment 50. Starting from the absolute position of this ideal division boundary (i.e., between bases 970 and 971, between bases 1940 and 1941, between bases 2910 and 2911, ..., between bases 47530 and 47531), the ideal division boundary was expanded by one base each to the left and right, centered on this ideal division boundary: 4 bases, 6 bases, 8 bases, 10 bases, 12 bases, 14 bases, 16 bases, 18 bases, 20 bases, 22 bases, and 24 bases, and the presence of the four-base sequence that could be the protruding candidate was examined. A specific example is explained below (Table 1). The ideal division boundary between fragments 01 and 02 is between bases 970 and 971. For the 16 bases centered on this ideal division boundary (the base sequence 5'-ATGCTGCTGGGTGTTT-3' from base 963 to base 988), there are seven candidate protruding combinations (ACAC / GTGT, AGCA / TGCT, ATGC / GCAT, CACC / GGTG, CAGC / GCTG, CCAG / CTGG, CTGC / GCAG). This operation was performed on all 49 base sequences around the ideal division unit extracted at a specified width, and it was examined whether there was at least one protruding candidate sequence within. As a result, when the extraction width was expanded to 24 bp, it was confirmed that there was at least one protruding candidate sequence of four bases in all extracted base sequences.To select specific protrusions from the protrusion candidates present in each sequence, we first prioritized the protrusion combination sequences that appeared least frequently among all the extracted sequences, as described above, from the extracted sequences with the smallest total number of protrusion combination candidates, thereby allocating unique protrusion combinations to all division units.

[0068] [Table 1]

[0069] Example 1: Preparation of lambda phage point mutants by integrating 50 DNA fragments and vector DNA <Lambda phage> Lambda phage is a bacteriophage that infects Escherichia coli and is the most extensively studied phage in molecular biology. Its genome consists of a double-stranded DNA with a total length of 48,502 bp, and the entire base sequence has been elucidated. The existence of various mutants has also been identified. In this example, we attempted to create lambda phage point mutants from short DNA units of approximately 1 kb.

[0070] <Design of fragmented lambda phage genome> The lambda phage DNA used was manufactured by Toyobo Co., Ltd. This product is linearized at the cos site. When the full-length base sequence of the phage genome was examined, six base sequences (g.138delG, g.14266_14267insG, g.37589C>) were found to be different from the base sequence registered in the database (accession number J02459.1). The differences were (T, g.37743C>T, g.43082G>A, g.45352G>A) (SEQ ID NO: 3) (the total length of SEQ ID NO: 3 is 48,526 bp, including the above 48,522 bp plus the other 4 bases at the overhanging end). Using the obtained base sequence, ideal division boundaries were set at 970 bp intervals to divide the total length of 48,522 bp (including the overlapping cos site) into approximately equal lengths. As a result of the above-mentioned (method for setting DNA unit division regions), the 5'-end overhangs composed of the 4 bases to the right of the cleavage site shown in Table 1 could be allocated so that they were unique within each DNA unit population.

[0071] <Selection of the type of restriction enzyme that generates the overhang> Type IIS restriction enzymes that generate any four-base overhang sequence include AarI (5'-CACCTGC(N)4 / -3',5'- / (N)8GCAGGTG-3'), BbsI (5'-GAAGAC(N)2 / -3',5'- / (N)6GTCTTC-3'), BbvI (5'-GCAGC(N)8 / -3',5'- / (N)12GCTGC-3'), BcoDI (5'-GTCTCN / -3',5'- / (N)5GAGAC-3'), BfuAI (5'-ACCTGC(N)4 / -3',5'- / (N)8GCAGGT-3'), and BsaI (5'-GGTCTCN / -3',5'- / (N)5 Examples include BsmAI (5'-CGTCTCN / -3',5'- / (N)5GAGACG-3'), BsmFI (5'-GGGAC(N)10 / -3',5'- / (N)14GTCCC-3'), BspMI (5'-GCGATG(N)10 / -3',5'- / (N)14CATCGC-3'), FokI (5'-GGATG(N)9 / -3',5'- / (N)13CATCC-5'), and SfaNI (5'-GCATC(N)9 / -3',5'- / (N)13GATGC-5'). Among these restriction enzymes, we investigated whether they were present in the E. coli plasmid vectors (pMD19, Simple, TAKARA) used to subclone the gene fragments, or whether the restriction enzymes that were present but only generated fragments sufficiently larger and sufficiently smaller than the ideal cleavage unit.We found five types (AarI, BbsI, BfuAI, BsmFI, BtgZI) that had no cleavage site at all, and one restriction enzyme (BsmBI) that had a recognition sequence within the vector but generated fragments sufficiently larger and sufficiently smaller than the ideal cleavage unit, for a total of six candidates. When the distribution of these candidate restriction enzyme sites was examined throughout the entire lambda phage from fragments 01 to 50, 12 sites were found for AarI, 24 sites for BbsI, 41 sites for BfuAI, 38 sites for BsmFI, 45 sites for BtgZI, and 14 sites for BsmBI. For each restriction enzyme, there were no restriction enzyme recognition sites that did not exist in the lambda phage genome.Therefore, we decided to select and use restriction enzymes that do not cut internally for each DNA unit. In order to minimize the number of restriction enzymes used, we investigated combinations of restriction enzymes and found that using only three types, BbsI, AarI, and BsmBI, was sufficient. The allocation of type IIS restriction enzymes used to excise each DNA unit is as follows: The group cleaved with BbsI contained a total of 33 fragments: fragments 01-08, 12, 16-22, 24, 27, 28, 33-39, 43, and 45-50. The group cleaved with AarI contained a total of 9 fragments: fragments 09-11, 13, 23, 25, 30, 32, and 44. The group cleaved with BsmBI contained a total of 8 fragments: fragments 14, 15, 26, 29, 31, and 40-42.

[0072] <Cloning of gene fragments> All 50 fragments, from fragment 01 to fragment 50, were amplified from the entire lambda phage genome using PCR. First, the restriction enzyme recognition site determined above was added to the 5' end of the primer for amplifying the DNA sequence between the overhang combinations determined above at the position where the desired overhang would be excised, and a primer with a TAG sequence added to the 5' end was used. These primer sets were used to amplify DNA fragments of the specified region from the entire lambda phage genome. PCR reaction conditions were as follows: per reaction (50 μl), 25 μl of KOD Plus 10x buffer Ver., 3 μl of 25 mM MgSO, 5 μl of dNTPs (2 mM each), 1 μl of KOD Plus (1 unit / μl), 48 pg of lambda phage DNA (TOYOBO), 15 pmol of primers (F primer and R primer each), and sterile water. PCR was performed using a GeneAmp PCR System 9700 (Applied Biosystems) according to the following program:

[0073] After incubation at 94°C for 2 minutes, 30 cycles of 98°C for 10 seconds, 55°C for 30 seconds, and 68°C for 1 minute were repeated, followed by incubation at 68°C for 7 minutes. The amplified DNA fragments were separated on a 1% agarose gel (UltraPure Agarose, Invitrogen) prepared in 1x TAE buffer (prepared by diluting 50x concentrated Tris-acetate-EDTA buffer (pH 8.3 at 25°C) manufactured by Nacalai Tesque) containing 2 mg / ml Crystal Violet (Wako Pure Chemical Industries, Ltd.) at 100 V for 10 minutes using an electrophoresis apparatus (i-MyRun.NC, Cosmo Bio). The target DNA bands were then separated using a razor blade as gel fragments of approximately 200 mg. DNA fragments were purified from the gel fragments using a Concert Rapid Gel Extraction System (Life Technologies). Specifically, three volumes of L1 Buffer were added to the gel fragments and dissolved in a block incubator at 45°C for approximately 10 minutes. The solution was then added to the provided spin column cartridge (a 2 ml centrifuge tube with a spin column attached) and centrifuged at 20,000 × g for 1 minute. The flow-through was discarded. After that, 750 μl of L2 Buffer was added to the spin column and centrifuged at 20,000 × g for 1 minute. The flow-through was discarded. To more thoroughly remove any residual L2 Buffer remaining in the spin column, the column was centrifuged again at 20,000 × g for 1 minute. The 2 ml centrifuge tube was then discarded and the spin column was transferred to a new 1.5 ml centrifuge tube. The spin column was added with 30 μl of TE buffer (10 mM Tris-HCl, 1 mM EDTA, pH 8.0) and left for 2 minutes. The DNA solution was then centrifuged at 20,000 × g for 1 minute to recover the DNA. The recovered DNA was stored at -20°C until use. The resulting DNA fragments were cloned into E. coli plasmid vectors using the TA cloning method described below.

[0074] To 8 μl of the DNA unit solution, 1 μl of 10×Ex-Taq Buffer, 0.5 μl of 100 mM dATP, and 0.5 μl of Ex-Taq, which are included with TAKARA's PCR enzyme Ex-Taq, were added, and the mixture was incubated at 65°C for 10 minutes to add an A overhang to the 3' end of the DNA unit. 1 μl of this DNA unit solution was mixed with 1 μl of TAKARA's pMD19-Simple and 3 μl of sterile water, and 5 μl of TAKATA Ligation (Mighty) Mix was added. The mixture was then incubated at 16°C for 30 minutes. Five microliters of this ligation solution was added to 50 μl of chemically competent E. coli DH5α cells, incubated on ice for 15 minutes, heat-shocked at 42°C for 30 seconds, and then left on ice for 2 minutes. After adding 200 μl of LB medium, the cells were incubated at 37°C for 1 hour, and then plated onto an LB plate containing carbenicillin (100 μg / ml) and 1.5% agar. The cells were then cultured overnight at 37°C to obtain plasmid transformants. The resulting colonies were prepared using PCR template DNA preparation reagents (Cica Geneus DNA Preparation Reagent, Kanto Chemical Co., Ltd.). Specifically, a small amount of colony material from the plate, picked with a toothpick, was suspended in 2.5 μl of a solution prepared by mixing reagents A and B in a 1:10 ratio. The suspension was then incubated at 72°C for 6 minutes, followed by incubation at 94°C for 3 minutes. To the resulting liquid, 2.5 μl of TAKARA Ex-Taq 10X enzyme, 2 μl of 2.5 mM dNTP solution, 0.25 μl of 10 pmol / μl M13F primer, 0.25 μl of 10 pmol / μl M13R primer, 17 μl of sterile water, and 0.5 μl of Ex-TaqHS were added, and the mixture was incubated at 94°C for 5 minutes, followed by one cycle of 98°C for 20 seconds, 55°C for 30 seconds, and 72°C for 1 minute. DNA was amplified by 30 cycles of PCR, and the nucleotide sequence of the PCR product was examined to confirm whether it matched the desired sequence. Finally, the correct sequence was obtained from all clones. During this process, one of the mutants obtained for fragment 10 was a synonymous substitution mutant (g.9515G>C) in the coding region of gene V. This mutation results in the appearance of a new restriction enzyme AvaI recognition site in the phage genome (Figure 2). In this example, we decided to use this synonymous substitution mutant (g.9515G>C) for fragment 10 instead of the wild type to clearly demonstrate that the constructed phage was artificially produced.

[0075] <High-purity purification of plasmids containing DNA units> A total of 50 E. coli transformants carrying plasmids cloning fragments 01-50 containing the desired sequences were each cultured overnight in 50 ml of LB medium containing 100 μg / ml carbenicillin at 37°C and 120 μm spm. The resulting cells were purified using a QIAfilter Plasmid Midi Kit (Qiagen). 50 μl of the resulting crude plasmid solution was added with 5 μl of 3 M potassium acetate-acetic acid buffer (pH 5.2) and 125 μl of ethanol, and centrifuged at 20,000 × g for 10 min to precipitate the DNA. The resulting precipitate was rinsed with 70% ethanol, the residue removed, and redissolved in 50 μl of TE (pH 8.0). A 1 μl aliquot of this crude plasmid solution was used to measure the DNA concentration using an ultramicrospectrophotometer (ND-2000, Thermo). At this point, the DNA content of the crude plasmid solution was approximately 0.5-4 μg / μl. Based on the measured values, 5 μg of DNA was collected from each crude plasmid solution in a 1.5 ml tube, and sterile water was added to bring the total volume to 50 μl. This was mixed with 6 μl of Plasmid Safe (Epicenta) 10x reaction buffer, 2.4 μl of 25 mM ATP solution, and 2 μl of Plasmid Safe enzyme solution. The mixture was incubated at 37°C for 1 h in a programmable block incubator BI-526T (ASTEC), followed by 75°C for 30 min to inactivate the enzyme. The resulting solution was purified using a PCR purification kit (Qiagen). In the final purification step, the DNA adsorbed to the column was eluted with 25 μl of TE buffer (pH 8.0) instead of the elution buffer provided with the kit, yielding a highly purified plasmid solution. DNA electrophoresis (UltraPure Agarose, Invitrogen) was performed on the plasmid containing fragment 01 before and after purification, and the plasmid containing fragment 21, and it was confirmed that the target fragment (unit DNA) had been incorporated (Figure 3).

[0076] <Precise concentration adjustment and equimolar synthesis of plasmids containing unit DNA> The resulting DNA solutions were again measured using an ultra-microspectrophotometer to determine the concentrations of the high-purity plasmid solutions. The concentrations of each sample ranged from approximately 100 ng / μl to 200 ng / μl, reflecting the degree of purification of the crude plasmid solution, relative to the theoretical maximum of 200 ng / μl. Based on the measurement results, 15 μl of each plasmid solution was placed in a 1.5 ml tube, and TE was added to each solution to achieve a concentration of 100 ng / μl. The concentrations of the resulting high-purity plasmid solutions were again measured using an ultra-microspectrophotometer. Deviations from the target value of 100 ng / μl were found to be within a few percent. Therefore, for each high-purity plasmid, the volume of 500 ng of DNA was calculated to two decimal places (μl accuracy). This volume (approximately 5 μl) of each DNA solution was then aliquoted and combined to approximately equimolar amounts for the type of restriction enzyme (BbsI group, AarI group, BsmBI group) used for subsequent excision.

[0077] <Bulk digestion of equimolar integrated plasmids with restriction enzymes> The total volume of the combined equimolar plasmid solution was approximately 165 μl for the BbsI group, 45 μl for the AarI group, and 40 μl for the BsmBI group. Two volumes of sterile water were added to each group to obtain 495, 135, and 120 μl of the highly purified plasmid solution, respectively. The resulting fragments were digested with the following restriction enzymes:

[0078] The BbsI group added 55 μl of 10× NEB buffer #2 and 27.5 μl of the restriction enzyme BbsI (NEB), and reacted a total of approximately 577 μl at 37°C for 2 hours. The AarI group added 15 μl of the 10× Buffer_for_AarI provided with the restriction enzyme, 3 μl of 50× Oligocleotide (for cleavage activation) provided with the restriction enzyme, and 7.5 μl of the restriction enzyme AarI (Thermo), and reacted a total of approximately 160 μl at 37°C for 2 hours. The BsmBI group added 13.3 μl of 10× NEB buffer #3 and 6.3 μl of the restriction enzyme BsmBI (NEB), and reacted a total of approximately 140 μl at 55°C for 2 hours. After 2 hours, 33 μl of plasmid solution was taken from each sample, 9 μl from the AarI group, and 8 μl from the BsmBI group, without compromising the equimolar relationship. Five μl of each was subjected to DNA electrophoresis to confirm that the plasmids had been cleaved by each restriction enzyme (Figure 4).

[0079] <Simultaneous fractionation and purification of 50 DNA units by agarose gel electrophoresis> After confirmation, an equal volume of phenol-chloroform-isoamyl alcohol (25:24:1) (Nacalai Tesque) was added to each group and mixed thoroughly to inactivate the restriction enzymes. The phenol-chloroform-isoamyl alcohol (25:24:1) mixtures from each group were then combined into a single tube and centrifuged (20,000 × g, 10 min) to separate the phenol and aqueous phases. The aqueous phase (approximately 900 μl) was collected in a separate 1.5 ml tube. 500 μl of 1-butanol (Wako Pure Chemical Industries) was added, mixed thoroughly, and centrifuged (20,000 × g, 1 min) to separate the phenol and aqueous phases. The saturated 1-butanol was removed. This procedure was repeated until the volume of the aqueous phase was reduced to less than 450 μl. To this mixture, 50 μl of 3M potassium acetate-acetic acid buffer (pH 5.2) and 900 μl of ethanol were added, and the mixture was centrifuged (20,000 × g, 10 min) to precipitate the DNA. The DNA was then rinsed with 70% ethanol and dissolved in 20 μl of TE. Two μl of 10× dye for electrophoresis was added, and the entire volume was passed through a 0.7% low-melting-point agarose gel (2-hydroxyethyl agarose Type VII, Sigma) in the presence of 1× TAE (Tris-Acetate-EDTA Buffer) using a general-purpose agarose gel electrophoresis system (i-MyRun.N Nucleic Acid Electrophoresis System, Cosmo Bio) at a voltage of 35 V (approximately 2 V / cm) for 4 hours to separate fragments 01-50 and the plasmid vector (Figure 5). The electrophoresis gel was stained with 100 ml of 1x TAE buffer containing 1 μg / ml ethidium bromide (Sigma) for 30 minutes and visualized by illuminating with long-wavelength ultraviolet light (366 nm). The bands (approximately 1 kb) representing fragments 01-50 were excised with a razor and collected in a 1.5 ml tube. 1x TAE buffer was added to the collected low-melting-point agarose gel (approximately 300 mg) to bring the total volume to approximately 700 μl, and the gel was dissolved by incubating at 65°C for 10 minutes.500 μl of 1-butanol was added to the resulting gel solution, and the aqueous and butanol phases were separated by centrifugation (20,000 × g, 1 min). The water-saturated butanol was discarded, and this process was repeated until the volume of the aqueous phase was reduced to 450 μl or less. 50 μl of 3 M potassium acetate-acetic acid buffer (pH 5.2) and 900 μl of ethanol were added to the resulting liquid, and the mixture was centrifuged (20,000 × g, 1 min) to obtain a DNA precipitate. This was rinsed with 70% ethanol and then dissolved in 20 μl of TE. 1 μl of the aliquot was used to measure the concentration using an ultramicrospectrophotometer.

[0080] Quantitative PCR was performed to confirm the number of moles of each group before and after size fractionation. Figure 6 shows the distribution of the number of molecules of each DNA unit before and after size fractionation, and Figure 7 shows the rate of change in the number of molecules of each DNA unit. This confirmed that the 50 fragments were collected in approximately equimolar ratios and without losing their molar ratio.

[0081] <Gene accumulation> The DNA weight concentration of the equimolar mixture of fragments 01-50 was 98 ng / μL, with a total base sequence of 48,522 bp, while the vector DNA (pGETS118-AarI / AarI) was 190 ng / μL and had a total length of 15,139 bp. To obtain an equimolar mixture of the two DNAs, 6.21 μl of the equimolar mixture of fragments 01-50 was mixed with 1.00 μl of vector DNA. To 7.2 μl of the resulting equimolar mixture, 8.2 μl of 2x ligation buffer was added, and the whole was incubated at 37°C for 5 min. Then, 1 μl of T4 DNA ligase (Takara) was added and the mixture was incubated at 37°C for 4 h. Ligation was confirmed by electrophoresis of a portion of the mixture (Figure 8). 8 μl of this was transferred to a new tube, and 100 μl of Bacillus subtilis competent cells was added. The mixture was then cultured in a duck rotor at 37°C for 30 minutes. 300 μl of LB medium was then added, and the mixture was cultured in a duck rotor at 37°C for 1 hour. The culture was then spread onto an LB plate containing 10 μg / ml tetracycline and cultured overnight at 37°C. 250 colonies were obtained.

[0082] <Confirmation of the plasmid structure of the transformant> Twelve colonies were randomly selected and cultured overnight in 2 ml of LB medium containing 10 μg / ml tetracycline. To amplify the internal plasmid copy number, IPTG was added to a final concentration of 1 mM, and the culture was further incubated at 37°C for 3 hours. Plasmids were extracted from the resulting cells and double-digested with the restriction enzymes HindIII and SfiI. Confirmation by electrophoresis revealed that four of the 12 strains showed the desired cleavage pattern (Figure 9). Plasmids from these four strains were mass-prepared by cesium chloride-ethidium bromide density gradient ultracentrifugation. The structures of the plasmids were analyzed using 13 restriction enzymes and confirmed by electrophoresis. All sequences matched the expected fragments (Figure 10). Furthermore, the entire plasmid, excluding the vector portion, was sequenced, and all four strains' plasmids were found to be identical to the expected sequences.

[0083] <Confirmation of the function of the accumulated genes> To confirm the lambda phage function of the four plasmids, we examined their plaque-forming ability as follows. First, we digested each of the integrated plasmids (#3, #4, #6, and #12) with lambda terminase (Epicentre) to separate the vector and integrated gene portion. This was then added to a lambda packaging extract (Gigapack III Plus Packaging Extract, Agilent Technologies). This was then used to infect E. coli (VCS257 strain), spread onto an LB plate, and incubated overnight at 37°C. Plaques were observed. The morphology of the resulting plaques was confirmed to be similar to that obtained with TOYOBO lambda phage DNA (Figure 11). Phage DNA was purified from plaques obtained from each plasmid and digested with the restriction enzyme AvaI to check for the presence of the introduced mutations. As shown in Figure 12, the digestion pattern was different from that of TOYOBO lambda phage DNA, confirming the presence of the AvaI site as expected in all phages. This confirmed that the lambda phage genome prepared by assembling all 50 fragments, fragments 01 to 50, was complete in both base sequence and plaque-forming ability.

[0084] These results indicated that a total of 51 DNA fragments, consisting of 50 DNA units constituting lambda phage DNA and the vector DNA (pGETS118-AarI / AarI), were successfully ligated.

[0085] Example 2: Construction of an artificial mevalonate pathway operon by assembling 55 DNA fragments and vector DNA Many isoprenoids, which have an isoprene unit as a backbone, are known, but they are all synthesized from a common starting material, isopentenyl diphosphate (IPP). Two pathways leading to IPP from glycolysis are known to exist: the mevalonate pathway and the non-mevalonate pathway. Some organisms possess both pathways, but Escherichia coli only possesses the non-mevalonate pathway. To enhance the IPP production ability of E. coli, we attempted to construct artificial genes for some of the genes in the mevalonate pathway of eukaryotic yeast, adapted to the codon usage of E. coli, by assembling synthetic DNA fragments.

[0086] <Sequence design of artificial mevalonate operon> We attempted to create an artificial operon (5,951 bp) (SEQ ID NO: 4) consisting of three artificial genes (ERG10 (1.2 kb), ERG13 (1.5 kb), and HMG1 (3.2 kb)) required for the metabolic pathway from acetyl-CoA to mevalonate, which constitutes the first half of the yeast mevalonate pathway. The codons were converted based on the codon usage frequency in E. coli. (Note: The total length of SEQ ID NO: 4 is 5,955 bp, including the four-base sequence used for the overhang.) The yeast genes were converted to E. coli codons by ranking the frequency of occurrence among all yeast synonymous codons, and then by ranking the frequency of occurrence among all E. coli synonymous codons. The conversion of yeast genes to E. coli codons was performed by converting codons with the same ranking.

[0087] <Design of DNA units> The 5,951-bp DNA sequence that had undergone synonym codon conversion was searched for restriction enzyme sites that would not cleave it, as in Example 1. It was found that there was no recognition sequence for the restriction enzyme AarI, and therefore AarI could not cleave it. Therefore, it was decided to prepare all clones using AarI. Dividing the full length 5,951 bp into 55 fragments yielded fragments with an average size of 108 bp. This size was used as the ideal division unit. We examined whether any one of the specific sequences appeared in all of the ideal division units among the 60 specific sequences (44 combinations (Group III) mentioned above, excluding the 8 palindromic combinations out of the 52 combinations in which A and T totaled two and C and G totaled two, and 16 combinations (Group IV) mentioned above, in which there were no three consecutive C and G sequences out of the 32 combinations in which A and T totaled one and C and G totaled three) and found that any one specific sequence appeared within a ±7 bp range from the ideal division unit. Based on this, the full length was divided into 55 fragments of 98 to 115 bp. Table 2 shows the division units and overhanging base sequences of the assembled DNA in Example 2. Note that overhangs consisting only of A and T (ATTA and AAAA) were used for the boundary between the mevalonate gene cluster and the gene assembly vector.

[0088] [Table 2]

[0089] <Preparation of DNA units using synthetic DNA> Each fragment obtained by division was analyzed by the method of Rossi et al. It was prepared using two 80-base chemically synthesized DNA fragments according to the method described in Itakura, K. 1982. J. Biol. Chem. 257, 9226-9229 (1982)). Specifically, the two chemically synthesized DNA fragments were hybridized at their 3' ends by several tens of base pairs, and a recognition site was added to the 5' end of the DNA fragment, closer to the AarI cleavage site than the AarI cleavage site, so that the overhang designed above would appear upon AarI digestion. To amplify the double-stranded DNA fragments obtained by hybridization of these two synthetic DNA fragments and subsequent template-dependent extension reaction using PCR, three types of DNA fragments, including one PCR primer designed to hybridize to the AarI recognition sites at both ends, were added. A DNA fragment surrounded by AarI cleavage sites was obtained by PCR, and this fragment was ligated to the E. coli plasmid vector pMD19 using the TA cloning method and transformed into E. coli for cloning. By sequencing these, clones with desirable base sequences were selected for each fragment.

[0090] <Equimolar mixture of plasmids containing unit DNA> Fifty-five E. coli strains containing the desired clones were cultured, and 50 μl of crude plasmid solution was obtained from each strain using Plasmid Mini-Prep (QIAGEN). DNA concentrations of 1 μl of each sample were measured using a microvolume spectrophotometer, yielding values ​​ranging from 82 to 180 ng / μl. Approximately 5 μg of each plasmid was treated with Plasmid Safe, heat-inactivated, and purified using a Mini-Elute PCR Purification Kit (QIAGEN) to yield 25 μl of highly purified plasmid solution. Concentrations of 1 μl of each sample were measured using a microvolume spectrophotometer, yielding values ​​ranging from 108 to 213 ng / μl. From this, 20 μl of the highly purified plasmid solution was transferred to separate tubes, and the tubes were diluted with TE to a calculated concentration of 100 ng / μl. The concentration of this purified plasmid solution was again calculated using a microspectrophotometer. Based on this concentration, the volume required to obtain 500 ng of each highly purified plasmid was calculated to two decimal places (to the nearest μl). This volume (approximately 5 μl) was then aliquoted from each plasmid solution and pooled into a single tube. To a total of approximately 275 μl of equimolar plasmid mixture, 2x sterile water, 137.5 μl of 10x Buffer for AarI, and 67.5 μl of the restriction enzyme AarI were added and incubated overnight at 37°C.

[0091] <Single size fractionation of 55 DNA units> The reaction mixture was inactivated by adding an equal volume of phenol, chloroform, and isoamyl alcohol (25:24:1) and then centrifuged. The supernatant was purified by ethanol precipitation, and the precipitate was dissolved in 20 μl of TE. Xylene cyanol was added as a dye for electrophoresis, and the vector DNA pMD19 and the insert DNA fragments were separated by electrophoresis on a 2.5% agarose gel in TAE buffer at 100 V for 30 minutes (Figure 13). The gel was then cut with a razor blade, and a portion was stained with ethidium bromide. The bands of the desired 55 equimolar fragments were excised from the unstained gel, while the positions of the bands were confirmed.

[0092] <Purification of equimolar DNA population> DNA was purified from the resulting gel fragment using a MiniElute Gel Extraction Kit (QIAGEN) as follows.

[0093] After measuring the volume of the gel by weight, 15 volumes of CG Buffer were added and the gel was dissolved by incubating at 50°C for 10 minutes. A volume of isopropyl alcohol 5 times the volume of the gel was added, and the liquid was poured into the attached column and centrifuged to adsorb the DNA to the column. The column was washed by adding 500 μl of CG Buffer and centrifuging, and then further washed by adding 750 μl of PE Buffer and centrifuging. After centrifuging once to completely remove residue, 10 μl of TE buffer was added to the column and centrifuged to obtain a mixed solution of approximately equimolar 55 fragment DNA units.

[0094] <Addition of DNA containing a replication origin to an equimolar DNA mixture> The DNA concentration was measured using an ultramicrospectrophotometer and found to be 20 ng / μL. The concentration of pGET151 / AarI, which was prepared in parallel, was 67 ng / μL. Taking into account the length ratio of these fragments (5955 bp:4306 bp), an equimolar mixture of the 55 fragments and pGETS151 / AarI was mixed at a ratio of 4.63:1.

[0095] <Gene accumulation> To 5.63 μl of the resulting equimolar mixture, 6.63 μl of 2x ligation buffer was added. The mixture was incubated at 37°C for 5 min, followed by the addition of 1 μl of T4 DNA ligase (Takara) and incubation at 37°C for 4 h. An aliquot was electrophoresed to confirm whether the fragment DNA and vector DNA had been ligated in a tandem repeat configuration (Figure 14). After the ligation, 8 μl of the solution was transferred to a separate tube, to which 100 μl of B. subtilis competent cells were added. The mixture was then cultured in a Duck rotor at 37°C for 30 min. 300 μl of LB medium was then added, and the mixture was cultured in a Duck rotor at 37°C for 1 h. The culture was then spread onto an LB plate containing 10 μg / ml tetracycline.

[0096] <Transformation and structural confirmation of the aggregate> From the 154 colonies obtained, 24 clones were randomly selected and inoculated into LB containing 10 μg / ml tetracycline. IPTG was added to a final concentration of 1 mM during logarithmic growth, and the cells were cultured until stationary phase. Plasmid DNA was extracted, digested with the restriction enzyme PvuII, and the digestion pattern was examined by electrophoresis (Figure 15). Two clones (#10 and #20) were confirmed to have the expected nucleotide sequence. Further digestion with other restriction enzymes and detailed structural confirmation by electrophoresis confirmed the desired structure (Figure 16). Sequencing of these plasmids confirmed that clones 10 and 20 had the intended nucleotide sequence.

[0097] These results indicated that a total of 56 DNA fragments, including 55 DNA units constituting the mevalonate pathway artificial operon and the vector DNA (pGETS151-pBR), were successfully ligated.

[0098] From the above, it was confirmed that the method for preparing a DNA concatemer of the present invention can ligate more than 50 DNA fragments. The reason why such a large number of DNA fragments can be ligated is thought to be that the molar numbers of each DNA fragment are more accurately closer to the same in the DNA fragment composition prepared by the method of the present invention.

[0099] The reason why the molar numbers of each DNA unit molecule are closer to being exactly the same in the DNA unit molecule composition prepared by the method of the present invention is presumed to be as follows.

[0100] In Examples 1 and 2 above, when measuring the concentration of each DNA unit in a solution containing a DNA unit, an additional sequence (specifically, a circular plasmid DNA) is linked to each DNA unit. Thus, even if the distribution of base sequence lengths among the different types of DNA unit is large, the distribution of base sequence lengths is narrowed when measuring the concentration of the solution due to the addition of an additional sequence. This reduces the error in the number of moles of each DNA unit calculated based on the measurement results. Therefore, by aliquoting each solution based on the measurement results and adjusting the number of moles of DNA unit in each solution to be the same, it is possible to bring the molar ratio in each solution closer to 1, and it is presumed that the DNA unit is more accurately made approximately equimolar.

[0101] In Example 1, the standard deviation of the distribution of the combined length of each DNA unit base and the combined length of the additional sequence linked to each DNA unit base was 3691.4±6.6 bp, or ±0.18% of the average combined length. In Example 2, the standard deviation of the distribution of the combined length of each DNA unit base and the combined length of the additional sequence linked to each DNA unit base was 2828.2±4.5 bp, or ±0.16% of the average combined length. In Examples 1 and 2, the standard deviation was so small relative to the average combined length that it is believed that the error in the number of moles of each DNA unit base calculated based on the measurement results of DNA concentration in solution was reduced.

[0102] The ratio of the average base length of the additional sequences linked to each DNA unit to the average base length of the DNA unit is about 2.7 in Example 1 and about 27 in Example 2. As described above, because the average base length of the additional sequences linked to each DNA unit is longer than the average base length of the DNA unit, it is thought that the error in the number of moles of each DNA unit calculated based on the measurement results of the DNA concentration in solution is further reduced, further reducing the error in the number of moles of each DNA unit calculated based on the measurement results of the DNA concentration in solution.

[0103] Furthermore, in Examples 1 and 2, the DNA unit fragments were designed using non-palindromic sequences near the DNA assembly at equally divided positions as boundaries so that when the base length of the DNA assembly sequence was divided by the number of DNA unit fragments, the base lengths would be equal. When DNA unit fragments were designed in this way, the lengths of the individual DNA unit fragments were approximately the same. As a result, when additional sequences were removed with restriction enzymes and then electrophoresed and size fractionated, the DNA unit fragments appeared as bands at approximately the same positions, enabling DNA unit fragments to be recovered in a single size fractionation, improving work efficiency.

[0104] In Example 1 above, the restriction enzymes used to remove the additional sequences were classified into groups according to their type (three types in Example 1, one type in Example 2). Therefore, before the removal step, solutions containing two or more types of DNA unit fragments could be mixed for each group, eliminating the need to treat each DNA unit fragment with a restriction enzyme and allowing each group of restriction enzymes to be treated with a restriction enzyme in a single step. This confirmed that the efficiency of the DNA concatemer production process was improved. Furthermore, even when such a mixture was used, it was confirmed that a large number of DNA unit fragments could be ligated, as the mixture was mixed in a substantially equimolar state, as described above.

[0105] (Test Example 1: Confirmation of the redundancy (r) of the repeating unit of the number of assembled DNA units required for transformation of Bacillus subtilis plasmid DNA) To confirm the repeat number (redundancy) r of the number of assembled DNA units required for Bacillus subtilis plasmid transformation, the following test was carried out.

[0106] The following DNAs (A) to (H) were prepared using the plasmid pGETS118-t0-Pr-SfiI-pBR (SEQ ID NO: 1) which has an effective replication origin in Bacillus subtilis.

[0107] <Preparation of DNA (A)> DNA (A) is a circular monomer plasmid DNA with redundancy r = 1. First, pGETS118-t0-Pr-SfiI-pBR was transformed into E. coli. The plasmid obtained from this transformant mainly contained DNA (A), but also contained a small amount of multimers. To remove these, the plasmid was subjected to DNA size fractionation by low-melting-point agarose gel electrophoresis, and only the monomer plasmid DNA region was excised from the gel and purified to prepare DNA (A).

[0108] <Preparation of DNA (B)> The DNA in (B) is a linear monomeric plasmid DNA with redundancy r = 1. This DNA in (B) was prepared by treating the DNA in (A) with the restriction enzyme BlpI (recognition site: 5'-GC / TNAGC-3').

[0109] <Preparation of DNA (C)> The DNA in (C) is a linear multimeric plasmid DNA with tandem repeats with redundancy r>1. The BlpI used to prepare the DNA in (B) forms a non-palindromic 3-base overhang at the 5' end. Therefore, by ligating the DNA in (B) with DNA ligase, a linear multimeric plasmid DNA with consecutive plasmid units in the same direction is obtained. The DNA of (C), which is a mer plasmid DNA, was prepared.

[0110] <Preparation of DNA (D)> The DNA (D) is a linear monomeric plasmid DNA with redundancy r = 1. This DNA (D) was prepared by treating the DNA (A) with the restriction enzyme EcoRI (the recognition site is 5'-G / AATTC-3').

[0111] <Preparation of DNA (E)> DNA (E) is a linear multimeric plasmid DNA in which the DNA (D) above is ligated in random orientation, partially containing a redundancy r>1 region. The EcoRI used to prepare DNA (D) above forms a palindrome with a three-base overhang at the 5' end. Ligation of EcoRI-cleaved plasmid DNA makes it possible to create a multimeric plasmid DNA in which the plasmid units are ligated in random orientation. DNA (E) was prepared by ligating DNA (D) above using DNA ligase.

[0112] <Preparation of DNA (F)> The DNA in (F) is a linear quasi-monomer mixture with r ≒ 1, which is obtained by cleaving the DNA in (A) with the restriction enzyme KasI, which cleaves only one site, dephosphorylating it, and then cleaving it with the nearby BlpI. This is an equal mixture of DNA fragments cleaved with the restriction enzyme AfeI, which cleaves only one site, dephosphorylating it, and then cleaving it with the nearby BlpI. The redundancy r of each DNA fragment in this mixture (F) is slightly less than 1.

[0113] <Preparation of DNA (G)> The DNA in (G) is a linear quasi-dimeric plasmid DNA with redundancy r = 1.98, which was created by joining the two DNA fragments in (F) above using DNA ligase, with the directional linkage specified only by the BlpI site.

[0114] <Preparation of DNA (H)> The cleavage sites of (B) and (D) are far apart. The DNA of (H) is an equimolar mixture of the DNAs of (B) and (D) without ligation.

[0115] <Transformation of DNA (A) to (H) into Bacillus subtilis competent cells> The DNAs (A) to (H) above were transformed into Bacillus subtilis competent cells, and the number of transformants per μg was calculated based on the number of tetracycline-resistant strains obtained. The DNAs (A) to (H) were dissolved in ligation buffer and used for transformation, regardless of whether or not a ligation reaction was performed. Figure 17 shows an electrophoretic image of the DNAs (A) to (H), and Figure 18 shows the number of transformants in Bacillus subtilis competent cells for the DNAs (A) to (H). In the electrophoretic image of Figure 17, the DNAs (C) and (E) are difficult to distinguish because various sizes are distributed over a wide range on the lane. In Figure 17, the upper band in "G" represents the DNA (G) with redundancy r = 1.97, and the lower band represents the DNA (G) with redundancy r = 0.95 that was mixed in with the DNA (G).

[0116] These results confirmed that, excluding the circular DNA A, the only transformants obtained were the DNAs (C), (E), and (G) that were formed by ligating DNA. This indicates that when the redundancy is r=1 or r<1, no transformants are obtained even if two types of linear plasmid molecules with different cleavage sites that complement each other's cleavage site sequences are mixed, and that at least for linear DNA, the minimum redundancy must be r>1.

[0117] (Simulation 1 Ligation Simulation) <Ligation simulation algorithm settings> The simulation was programmed using VBA in the spreadsheet software Excel (registered trademark) 2007. The DNA fragment F in the virtual ligation is determined by three parameters F i (N i ,L i ,R i) where "i" refers to the fragment identification number, more specifically, the i-th row cell in Excel. "N" indicates the number of unit DNA fragments contained in one molecule of ligated DNA fragment during virtual ligation, "L" indicates the sequence of the left-hand overhanging end of the ligation product quantified as an arbitrary natural number, and "R" indicates the sequence of the right-hand overhanging end of the ligation product quantified as an arbitrary natural number, just like "L". Here, when L = R, the two overhanging sequences are complementary, and L = R defines that ligation is possible. The ligation simulation was performed as follows.

[0118] F i (N i ,L i ,R i ) fragment, a random number j satisfying i≠j is generated by multiplying a uniform random number between 0 and 1 generated by the RAND() method by m (described later) and rounding it to an integer, thereby generating F j (N j ,L j ,R j The following discriminant was used to determine whether these two fragments could be linked, and if so, the parameters of the fragments were changed as shown below.

[0119] L i =R j If the left end of Fi and the right end of Fj can be connected, then F i(new) (N i(old) +N j(old) ,L j(old) ,R i(old) ) fragments and F j(new) (0,0,0) fragment. i =L j If F i The right end of the fragment and F j If the left ends of the fragments can be ligated, i(new) (N i(old) +N j(old) ,L i(old) ,R j(old) ) fragments and F j(new) The (0,0,0) fragment was transformed.i ≠R j And R i ≠L j In the case of , no conversion (F i(new )(N i(old) ,L i(old) ,R i(old) ) Fragment and F j(new) (N j(old ),L j(old) ,R j(old) ) fragments), and virtual ligation was assumed not to occur. One cycle of virtual ligation was defined as performing these calculations from i = 1 to m. Here, m is a variable that indicates the total number of DNA fragments in the virtual ligation cycle, and in the case of the first cycle of the simulation, it means the total number of initial unit DNA fragments. After calculating one cycle of virtual ligation, the L i F so that the values ​​are in descending order i By rearranging the fragments, F other than the F(0,0,0) fragment can be i The total number of fragments was counted, and this total number was inserted as a new m, and virtual ligation was performed in the next cycle. The virtual ligation cycle was the minimum number of fragments m where there were no more complementary overhanging fragments and virtual ligation could not be performed any more. min I went to m min The value was calculated using the information of the initial DNA unit fragment before the start of ligation. m min = (total number of DNA unit fragments) - (total number of the smaller DNA unit fragments in the entire system among the two types of DNA unit fragments that satisfy the relationship L = R)

[0120] <Ligation simulation> For each accumulation scale of 6, 13, 26, and 51 fragments, a virtual unit DNA fragment population with an average number of identical unit DNA fragments of 640 and a coefficient of variation (CV) set in 1% increments in the range of 0 to 20% was created as follows.

[0121] Generate a random number group of 0 to 1 corresponding to each cluster size generated by Excel's uniform random number command RAND(), standardize the random number group to a mean value of 0 and a variance of 1, and then add the fragment mean value to each standardized random number. * Each provisional unit DNA fragment population was created by multiplying the CV (%) / 100 and adding the average number of fragments to the resulting value. This random number population was used to create each CV (%) of each accumulation scale. For the value, 20 independent groups were created, and the above simulation was performed for each random number group to obtain m min Virtual ligation was performed up to 1000 nucleotides. The resulting 20 virtual ligation fragments were combined, and the number of ligated fragments was tallied for each N value. The percentage of this (N value × number of ligated fragments) relative to the total number of DNA unit fragments used in ligation was calculated. A 100% stacked graph was created, with larger N values ​​at the bottom. The graph is shown in Figure 19. Figure 19 shows the distribution of the size of the initial DNA unit fragments ultimately incorporated into the ligation products. In Figure 19, (a) shows a graph for the accumulation of 6 fragments, (b) shows a graph for the accumulation of 13 fragments, (c) shows a graph for the accumulation of 26 fragments, and (d) shows a graph for the accumulation of 51 fragments. In the case of accumulation of 6 fragments, n = 6, and redundancy r = 1. The region with redundancy r < 1 is shown in the upper right corner. These results show that in the case of 6-fragment accumulation, most of the unit DNA fragments are incorporated into DNA fragments in the redundancy r>1 region, even at that CV value. Conversely, in the case of 51-fragment accumulation, most regions are in the redundancy region of less than 1, except for the region where the CV is close to 0%.

[0122] <Derivation of theoretical ligation formula from ligation simulation> From the numerical analysis of the above ligation simulation, in obtaining a generalized equation for the ligation mechanism, it was verified whether it was possible to calculate the fitting curve of the distribution of ligation products at each CV value for each integration scale. The distribution of the number of unit DNA fragments contained in the ligation products at CV = 20% for the case of 6 - fragment integration with an average of 640 fragments is shown in FIG. 20. In FIG. 20, the pattern of each bar graph in each redundancy (0 < r < 1, 1 < r < 2, 2 < r < 5, 5 < r < 10) indicates the type of component (in the case of 6 - fragment integration, out of the 6 components with remainders of 0, 1, 2, 3, 4, 5, excluding the component with a remainder of 0 for which logarithmic conversion cannot be performed due to an N value of 0, 5 components) that is distinguished when N of each redundancy is divided by r. Among different redundancies, bar graphs with the same pattern indicate that the components distinguished when N is divided by r are of the same type. Also, the linear approximation curves in the graph of FIG. 20 are obtained for each type of component pattern.

[0123] Overall, the histogram of the number of ligation product molecules for each N value showed an exponentially decreasing trend as the N value increased, and at first glance, it showed a distribution similar to a geometric distribution, which is a type of discrete probability distribution. However, microscopically, this histogram showed a periodic structure with the number of genetic integration modules or a period of 1 redundancy (6 in the case of 6 - fragment integration), and particularly at parts where the N value corresponded to an integer multiple of the genetic integration scale, it showed a characteristic structure where no fragments appeared. This characteristic does not completely match the geometric distribution or the exponential distribution when regarded as a continuous probability distribution. However, when the ligation product molecule number axis is converted to logarithmic display, and each component of its microscopic periodic structure, that is, the components distinguished by the remainder when N is divided by the integration scale 6, are taken out from each period and a linear approximation curve is obtained, it shows a very high squared correlation coefficient value (0.94 or more). Therefore, it is considered that there is no problem even if each component approximates an exponential distribution. This distribution was also observed in examples of other integration scales and other CV values other than the example with a CV of 20% for the concentration variation of 6 - fragment integration shown in FIG. 20. Hereafter, it is assumed heuristically that this mechanism can be approximated by an exponential distribution, and the exponential distribution function (f(n)=λ* exp(-λ * n)) was fitted.

[0124] Specifically, (1) first, in all the simulations for each of the integration scales described above, for each component distinguished by the remainder when N is divided by the integration scale, the slope of the line obtained by linear approximation from the logarithmically transformed molecular values ​​of about three periods was determined as -λ, and λ was calculated from this value. (2) Next, since the reciprocal of the parameter λ in an exponential distribution function, 1 / λ, is the average value of f(N), the average N value of the ligation product in the 20 random number populations was determined, and λ (CV(%)) for each CV(%) was calculated from the reciprocal of this average value. The results of (1) and (2) above were plotted on a graph with the horizontal axis representing the CV(%) value of the concentration variation of the unit DNA fragment and the vertical axis representing the λ value. When the plots were plotted, it was confirmed that all plots lay on a straight line of direct proportion passing through a certain origin, regardless of the scale of gene accumulation. Linear approximation curves were calculated for each of these plots, and the graphs are shown in Figure 21. In Figure 1, (a) is a graph showing the relationship between the slope λ calculated from the three cycles in (1) above and the variation in the concentration of the DNA unit fragment, and (b) is a graph showing the relationship between λ calculated from the reciprocal of the average N value in (2) above and the variation in the concentration of the DNA unit fragment. The general formula for this linear approximation curve, obtained from (2) with higher accuracy, is f(N) = 0.0058 * CV(%) * exp(-0.0058 * CV(%) * N). Also, from Figure 21, λ = 0.0058 * It was found that the squared value of the correlation coefficient for CV (%) can be expressed with a high correlation of 0.99. Therefore, it was confirmed that there are no problems with this general formula.

[0125] <Qualitative analysis of ligation reaction rate> The above ligation simulation assumed that all ligations between regular overhanging pairs were completed. To examine how close the ligation reaction conditions in actual gene assembly are to the reaction conditions in the simulation, we investigated the kinetics of the ligation reaction.

[0126] First, in a ligation reaction performed under actual gene assembly conditions, ligation products were sampled at various reaction times. The actual ligation rate was then qualitatively analyzed for all 51 ligation sites using the 51-fragment ligation of λ phage reconstitution. The average concentration of the DNA unit fragments used in the ligation was approximately 0.2 fmol / μl. After adding T4 DNA ligase to the DNA unit fragment solution, aliquots of each reaction solution were sampled at 0, 1.25, 2.5, 5, 10, 20, 40, 80, 160, and 320 minutes at 37°C. Commercially available λ phage genomic DNA (Toyobo) or a pre-assembled plasmid was digested with restriction enzymes using a primer set designed to amplify DNA spanning the junction between two properly ligated DNA unit fragments, and a primer set designed to amplify only the interior of each DNA unit fragment. The progress of ligation of each fragment was assessed using a dilution series of linearized DNA as an index. As a result, under these reaction conditions, ligation was completed in approximately 10 minutes at all ligation sites, and it was confirmed that ligation was nearly complete over the actual reaction time of 4 hours (240 minutes). Furthermore, after sufficient time had passed (40 minutes or later), the reaction rate at most ligation sites was approximately 1 relative to the value based on the smaller number of each DNA unit fragment, confirming that most of the ligated DNA fragments were ligated to their correct ligation partners.

[0127] <Estimation of misligation rate> To confirm the ligation status in more detail, we sequenced all clones (#1, #2, #5, #7, #8, #9, #10, #11) obtained from the λ phage genome reconstruction experiment, except for #3, #4, #6, and #12, for which the entire base sequence was determined, in order to identify misligation sites. The misligation sites for each clone are shown in Figure 22. From these results, it was confirmed that all seven clones, excluding #11, contained one or two misligations, and all misligations were identified. On the other hand, for #11, the same DNA unit fragment was repeatedly present within the DNA assembly, and although complete structural confirmation was not possible, the presence of all six misligation sites was confirmed. Since the exact number of DNA fragments in #11 is unknown, #11 was excluded and the frequency of misligation in all clones was calculated. As a result, it was confirmed that misligation occurred at a rate of about 1 in 46 ligation points, and the misligation rate was relatively low at about 2.2%. This result was consistent with the quantitative PCR results mentioned above.

[0128] <Verification of the consistency between the size distribution of actual ligation products and the simulation> Based on the above two verifications, the estimation of the misligation rate and the qualitative analysis of the ligation reaction rate, it was estimated that in actual ligation reactions, after a sufficient time of 4 hours, almost all ligation is completed and the probability of misligation occurring is low. Therefore, the following method was used to verify whether it is possible to predict the size distribution of actual ligation products by simulation for an actual population of DNA unit fragments in which the initial DNA unit fragments have varying concentrations.

[0129] First, we used the population used in the λ phage genome reconstruction experiment, in which a 7.5% CV variation in DNA unit fragment concentration was observed by quantitative PCR. Since the observed value of this quantitative PCR included a measurement error of CV = 3.6%, it was estimated that the CV of the true DNA unit fragment variation might be lower than CV = 7.5%. Therefore, we calculated the possible true CV by simulation, and estimated that if the true CV was 6.6%, the observed value might be CV = 7.5% due to the measurement error of CV = 3.6%. Therefore, we set the true DNA unit fragment variation of this population to CV = 6.6%, and simulated the ligation of 51 types of initial DNA unit fragments that satisfies the conditions of an average of 640 fragments and CV = 6.6% prepared for the RAND() method. In this simulation, m, which indicates a 100% reaction rate, was used. minIn addition, 100 independent random number populations were prepared for m values ​​obtained when ligation was performed at 95%, 96%, 97%, 98%, and 99% of the ligation potential. The reaction was continued until the specified m value was reached. The resulting 100 ligation product distributions were combined, and the DNA length of each hypothetical ligation product was calculated in bp units using the F(N,L,R) parameter. Next, the actual DNA unit fragment populations from the λ phage genome reassembly experiment, which exhibited a CV of 6.6% in fragment number, were reacted at 37°C for 4 hours, as described in the "Qualitative Analysis of Ligation Reaction Rate" section above. After the reaction, the actual molecular weight distribution of the DNA was determined for 16 hours using a CHEF pulsed-field gel electrophoresis system (Biocraft) under conditions of 0.5x TBE, 5 V / cm, and a 30-second cycle. A photograph of the electrophoresis is shown in Figure 23. The DNA density distribution of the electrophoretic photographs was obtained using NIH Image software, and the predicted DNA distribution maps obtained by simulation for each ligation efficiency were overlaid and compared. The results are shown in Figure 24. Figure 24 confirms that the DNA molecular weight distribution obtained by electrophoresis generally matches the predicted DNA distribution map for 98%-100% ligation efficiency, and in particular, the molecular weight of the maximum value indicating the highest concentration matches well with the 98% ligation efficiency. This indicates that the simulation results show that almost all ligation is completed within a 4-hour ligation reaction, and that although approximately 2% misligation may occur, the overall ligation process can be roughly reproduced.

[0130] <Generalization of ligation simulation> The DNA unit fragments actually used for assembly inevitably have variations in concentration. However, the extent to which the variation in concentration of the DNA unit fragments must be suppressed can be calculated by using the general formula f(N)=0.0058 obtained above. * CV(%) * exp(-0.0058 * CV(%) *Figure 25 shows a summary of the results using the CV(%). In the current gene accumulation experiments, the variation in DNA concentration is generally about CV(%) = 6.6, but Figure 25 shows that when CV(%) = 6.6, in the accumulation of 51 fragments, about 40% of the unit DNA fragments used are incorporated into ligation products with an r value of greater than 1. Furthermore, Figure 25 shows that if a new gene accumulation is planned using a 102 fragment population, which is twice the accumulation scale, and an accumulation efficiency similar to that of 51 fragments is expected, a CV(%) = 3.3 must be achieved. It was also shown that the general formula f(N) = 0.0058 * CV(%) * exp(-0.0058 * CV(%) * The relationship between the variation in the concentration of DNA unit fragments and the average number of DNA unit fragments per ligation product was calculated using the general formula f(N)=0.0058. The results are shown in Figure 26. * CV(%) * exp(-0.0058 * CV(%) * It was shown that by using the CV (%) value, it is possible to easily estimate the average number of unit DNA fragments contained in one ligation product.

Claims

[Claim 1] The invention described in the examples.

Citation Information

Patent Citations

  • Method for manufacturing plasmids containing inserted DNA units

    JP4479199B2